Separating Vocals for Better Syncing

Separating a song creates a version where the voice is easier to hear without the instruments covering it.

This can help you hear:

  • When each word starts and ends

  • Quiet or slurred words

  • Breaths and pauses

  • Ad-libs

  • Background vocals

  • Overlapping singers

You can sync the song using the separated vocal stem. It must come from the exact same source as the original, and you must perform the final timing review using the original song.

What you need

  • A Windows or Mac computer

  • An NVIDIA or AMD GPU, if available

  • Pymss Studio

  • The song you want to sync

  • The becruily_deux.ckpt model

  • Headphones

FLAC or WAV is preferred, but a good-quality MP3 or M4A can also work.

1. Install Pymss Studio

Open the Pymss Studio releases and download the version matching your computer.

Computer

Version

Windows with an NVIDIA GPU

Windows CUDA

Windows with an AMD GPU

Windows ROCm

Windows without an NVIDIA GPU

Windows CPU

Windows with a stable connection and limited storage

Windows Online

Apple Silicon Mac

macOS MLX

Use the CUDA or ROCm version if you have a supported GPU. It will normally process songs much faster.

Install or extract Pymss Studio, then open it.

2. Download the starter model

Open the model browser in Pymss Studio and search for:

becruily_deux.ckpt

Download or import that exact model.

This model is a good starting point because it creates both:

song
├── vocals
└── instrumental

You do not need to download every model in the browser.

3. Check your song

Make sure you have the exact version you want to sync.

Watch out for:

  • Clean and explicit versions

  • Remasters

  • Deluxe or album versions

  • Radio edits

  • Music-video audio

  • Live recordings

  • Extra silence at the beginning

Two versions of the same song can have different timing.

4. Separate the song

In Pymss Studio:

  1. Add the original song.

  2. Select becruily_deux.ckpt.

  3. Choose an output folder and format.

  4. Start the separation.

  5. Wait for processing to finish.

The output folder should contain a vocal file and an instrumental file.

Give them clear names if necessary:

song-original.flac
song-vocals-deux.flac
song-instrumental-deux.flac

5. Listen to the vocal stem

Play the separated vocal file.

You should hear the singer much more clearly, although some instruments or audio artifacts may remain.

Listen for:

  • The first sound of each word

  • The moment one word changes into the next

  • The final consonant of a word

  • Pauses between words

  • Repeated or stretched syllables

  • Quiet background vocals

  • Ad-libs hidden behind the lead singer

The result does not need to sound perfect. It only needs to make the performance easier to understand.

6. Use it while syncing

Import the vocal stem into your timing tool and use it as your main syncing audio. Because the instruments are quieter or removed, word boundaries are usually much easier to hear.

Keep the original song available in case the model removes a vocal, creates an artifact, or makes a transition sound misleading.

While syncing:

  1. Play the vocal stem.

  2. Mark the beginning of the first word.

  3. Mark each transition into the next word.

  4. End the final word when the singer stops or pauses.

  5. Replay the line using the vocal stem and adjust it as needed.

  6. Check the original only when something sounds missing, unclear, or unnatural.

You do not need to switch back to the original after every line. Syncing the whole song against the vocal stem is fine as long as the stem stays aligned and the final review uses the original.

Example

A line may sound like this in the original:

I don't know what you're saying

The instruments might make don't know sound like one continuous word.

The vocal stem may reveal:

I | don't | know | what | you're | saying

You can use those clearer transitions to time the line directly against the vocal stem.

7. Check background vocals and ad-libs

Separation is not always perfect. A model may accidentally remove a background vocal or place it in the instrumental file.

If something appears to be missing:

  1. Listen to the original song.

  2. Check the vocal stem.

  3. Check the instrumental stem.

  4. Follow the original song when the stems disagree.

Never remove a lyric only because it is missing from the separated vocal.

8. Perform the final review

After syncing the song:

  1. Stop using the separated vocal.

  2. Play the TTML against the original song.

  3. Review every line.

  4. Check the beginning and end of the song.

  5. Make sure the timing still feels natural with the instruments present.

The listener will hear the original song, not your isolated vocal stem.

This final review is required even when the entire song was easy to sync using the vocal stem. The instruments can change how a transition feels, and the separation model may slightly distort or remove sounds.

If the first result is bad

Start with Deux. Only try another model if it fails in a noticeable way.

The vocal contains too many instruments

Try:

mel_band_roformer_vocals_fv7b_gabox.ckpt

This model focuses specifically on creating a cleaner vocal stem.

FV7b removes words or background vocals

Try:

BS-Roformer-Resurrection.ckpt

Compare it with Deux and FV7b. Use whichever result makes that part of the song easiest to hear.

You need a cleaner instrumental

Try:

Inst_GaboxFv9.ckpt

An instrumental is mainly useful for checking whether a sound belongs to the singer or the music. You normally do not need it for basic lyric syncing.

You want individual instruments

Use:

BS-Roformer-SW.ckpt

This can separate stems such as vocals, drums, bass, guitar, piano, and other sounds.

Most TTML makers do not need this. Use it only when a normal vocal separation cannot reveal a difficult section.

Which model should I use?

Situation

Model

First attempt

becruily_deux.ckpt

Cleaner vocals

mel_band_roformer_vocals_fv7b_gabox.ckpt

Alternative vocals

BS-Roformer-Resurrection.ckpt

Cleaner instrumental

Inst_GaboxFv9.ckpt

Individual instruments

BS-Roformer-SW.ckpt

The simple rule is:

Start with Deux.
If the vocals are unclear, try FV7b.
If FV7b removes something, try Resurrection.

Common problems

Processing is very slow

Make sure you installed the CUDA or ROCm version if your computer has a GPU. Close games and other applications using the GPU.

Pymss Studio runs out of GPU memory

Close other GPU-heavy applications and try again. Larger models require more video memory.

The vocal sounds metallic or watery

This is an AI separation artifact. Try another model, but ignore minor artifacts if the words remain easy to hear.

Part of the vocal is missing

Check the original and instrumental files. The model may have mistaken the vocal for an instrument.

The separated audio is early or late

Confirm that it came from the same original file. Do not use a stem created from a different release of the song.

The timing works locally but not in Spotify

Your source may not match Spotify’s release. Check for a different master, edit, clean version, or silence at the beginning.

Recommended workflow

Get the correct song
        ↓
Run becruily_deux
        ↓
Listen to the vocal stem
        ↓
Sync the song using the vocal stem
        ↓
Check suspicious or missing parts against the original
        ↓
Perform the complete final review with the original

Final checklist

Before finishing:

  • The source is the correct version of the song.

  • The vocal stem was created from that exact source.

  • Missing words were checked against the original.

  • Background vocals and ad-libs were reviewed.

  • The complete song received one final uninterrupted review.

You do not need the cleanest vocal extraction ever created. You only need a stem that helps you hear the performance more clearly.