Separating Vocals for Better Syncing
Separating a song creates a version where the voice is easier to hear without the instruments covering it.
This can help you hear:
When each word starts and ends
Quiet or slurred words
Breaths and pauses
Ad-libs
Background vocals
Overlapping singers
You can sync the song using the separated vocal stem. It must come from the exact same source as the original, and you must perform the final timing review using the original song.
What you need
A Windows or Mac computer
An NVIDIA or AMD GPU, if available
The song you want to sync
The
becruily_deux.ckptmodelHeadphones
FLAC or WAV is preferred, but a good-quality MP3 or M4A can also work.
1. Install Pymss Studio
Open the Pymss Studio releases and download the version matching your computer.
Computer | Version |
|---|---|
Windows with an NVIDIA GPU | Windows CUDA |
Windows with an AMD GPU | Windows ROCm |
Windows without an NVIDIA GPU | Windows CPU |
Windows with a stable connection and limited storage | Windows Online |
Apple Silicon Mac | macOS MLX |
Use the CUDA or ROCm version if you have a supported GPU. It will normally process songs much faster.
Install or extract Pymss Studio, then open it.
2. Download the starter model
Open the model browser in Pymss Studio and search for:
becruily_deux.ckptDownload or import that exact model.
This model is a good starting point because it creates both:
song
├── vocals
└── instrumentalYou do not need to download every model in the browser.
3. Check your song
Make sure you have the exact version you want to sync.
Watch out for:
Clean and explicit versions
Remasters
Deluxe or album versions
Radio edits
Music-video audio
Live recordings
Extra silence at the beginning
Two versions of the same song can have different timing.
4. Separate the song
In Pymss Studio:
Add the original song.
Select
becruily_deux.ckpt.Choose an output folder and format.
Start the separation.
Wait for processing to finish.
The output folder should contain a vocal file and an instrumental file.
Give them clear names if necessary:
song-original.flac
song-vocals-deux.flac
song-instrumental-deux.flac5. Listen to the vocal stem
Play the separated vocal file.
You should hear the singer much more clearly, although some instruments or audio artifacts may remain.
Listen for:
The first sound of each word
The moment one word changes into the next
The final consonant of a word
Pauses between words
Repeated or stretched syllables
Quiet background vocals
Ad-libs hidden behind the lead singer
The result does not need to sound perfect. It only needs to make the performance easier to understand.
6. Use it while syncing
Import the vocal stem into your timing tool and use it as your main syncing audio. Because the instruments are quieter or removed, word boundaries are usually much easier to hear.
Keep the original song available in case the model removes a vocal, creates an artifact, or makes a transition sound misleading.
While syncing:
Play the vocal stem.
Mark the beginning of the first word.
Mark each transition into the next word.
End the final word when the singer stops or pauses.
Replay the line using the vocal stem and adjust it as needed.
Check the original only when something sounds missing, unclear, or unnatural.
You do not need to switch back to the original after every line. Syncing the whole song against the vocal stem is fine as long as the stem stays aligned and the final review uses the original.
Example
A line may sound like this in the original:
I don't know what you're sayingThe instruments might make don't know sound like one continuous word.
The vocal stem may reveal:
I | don't | know | what | you're | sayingYou can use those clearer transitions to time the line directly against the vocal stem.
7. Check background vocals and ad-libs
Separation is not always perfect. A model may accidentally remove a background vocal or place it in the instrumental file.
If something appears to be missing:
Listen to the original song.
Check the vocal stem.
Check the instrumental stem.
Follow the original song when the stems disagree.
Never remove a lyric only because it is missing from the separated vocal.
8. Perform the final review
After syncing the song:
Stop using the separated vocal.
Play the TTML against the original song.
Review every line.
Check the beginning and end of the song.
Make sure the timing still feels natural with the instruments present.
The listener will hear the original song, not your isolated vocal stem.
This final review is required even when the entire song was easy to sync using the vocal stem. The instruments can change how a transition feels, and the separation model may slightly distort or remove sounds.
If the first result is bad
Start with Deux. Only try another model if it fails in a noticeable way.
The vocal contains too many instruments
Try:
mel_band_roformer_vocals_fv7b_gabox.ckptThis model focuses specifically on creating a cleaner vocal stem.
FV7b removes words or background vocals
Try:
BS-Roformer-Resurrection.ckptCompare it with Deux and FV7b. Use whichever result makes that part of the song easiest to hear.
You need a cleaner instrumental
Try:
Inst_GaboxFv9.ckptAn instrumental is mainly useful for checking whether a sound belongs to the singer or the music. You normally do not need it for basic lyric syncing.
You want individual instruments
Use:
BS-Roformer-SW.ckptThis can separate stems such as vocals, drums, bass, guitar, piano, and other sounds.
Most TTML makers do not need this. Use it only when a normal vocal separation cannot reveal a difficult section.
Which model should I use?
Situation | Model |
|---|---|
First attempt |
|
Cleaner vocals |
|
Alternative vocals |
|
Cleaner instrumental |
|
Individual instruments |
|
The simple rule is:
Start with Deux.
If the vocals are unclear, try FV7b.
If FV7b removes something, try Resurrection.Common problems
Processing is very slow
Make sure you installed the CUDA or ROCm version if your computer has a GPU. Close games and other applications using the GPU.
Pymss Studio runs out of GPU memory
Close other GPU-heavy applications and try again. Larger models require more video memory.
The vocal sounds metallic or watery
This is an AI separation artifact. Try another model, but ignore minor artifacts if the words remain easy to hear.
Part of the vocal is missing
Check the original and instrumental files. The model may have mistaken the vocal for an instrument.
The separated audio is early or late
Confirm that it came from the same original file. Do not use a stem created from a different release of the song.
The timing works locally but not in Spotify
Your source may not match Spotify’s release. Check for a different master, edit, clean version, or silence at the beginning.
Recommended workflow
Get the correct song
↓
Run becruily_deux
↓
Listen to the vocal stem
↓
Sync the song using the vocal stem
↓
Check suspicious or missing parts against the original
↓
Perform the complete final review with the originalFinal checklist
Before finishing:
The source is the correct version of the song.
The vocal stem was created from that exact source.
Missing words were checked against the original.
Background vocals and ad-libs were reviewed.
The complete song received one final uninterrupted review.
You do not need the cleanest vocal extraction ever created. You only need a stem that helps you hear the performance more clearly.