Vocal Separation

Split speech or singing away from the background and download isolated stems for further editing.

Usage

Use Vocal Separation to isolate the vocal track from background music, ambience, or mixed source audio.

This helps when you need a cleaner voice track before editing, dubbing, or downstream delivery.

Before you start

Vocal Separation supports three input sources: a direct upload, a Net Video source, or a completed session from your workspace.

For session-based input, the source session must already be completed and in the postprocess stage.

Uploads must be audio or video files. Text-only or unsupported files are rejected by the API.

If the input comes from Net Video, extraction must succeed first before the separation task can proceed.

Step-by-step instructions

  1. Open Vocal Separation

    Go to Sound → Vocal Separation.

  2. Choose the source type

    Pick Upload, Net Video, or Session depending on where the input already lives.

  3. Submit the task

    For uploads, VoxStudio gives you a signed upload target first. For sessions and Net Video, the task starts from the existing source metadata.

  4. Wait for asynchronous processing

    The task runs in the background and updates live status through SSE events while it is queued or processing.

  5. Download the results

    When finished, use the available download actions to save the separated vocals and background outputs.

Tips and examples

For the cleanest result, start from a higher-quality source file instead of a heavily compressed social export.

If the mix is extremely dense, expect some bleed between the vocal and background stems.

If you already have a completed session, use Session input so you do not have to re-upload the same media.

Common problems and fixes

Why can’t I choose my session?

Only completed sessions in the postprocess stage are valid audio-tool sources. Wait until the original session finishes first.

Why does the upload option fail immediately?

The API only accepts audio or video media uploads for audio tools. Re-export the file to a supported media type and retry.

Why is the result still noisy?

Separation isolates layers but does not guarantee perfect cleanup. Use Vocal Repair afterward if the vocal stem still needs improvement.

Why did the task disappear?

Expired tasks are cleaned up after a retention period. Download the outputs once they are ready instead of leaving them unclaimed.

FAQ

Can I run Vocal Separation on a supported online video link?

Yes, if the source platform is supported and the media can be extracted. Unsupported links still fail before processing starts.

Do I get both vocals and background?

Yes. The page exposes separate downloads for the isolated vocal and background outputs when the task completes.

Can I use Vocal Separation for speech-only recordings?

Yes. It can still help remove background layers from a spoken recording.

Should I use this before dubbing?

Use it when the original source audio needs cleanup or stem isolation before you create a new voice layer.

Related guides