# Separating Vocals for Better Syncing

Separating a song creates a version where the voice is easier to hear without the instruments covering it.

This can help you hear:

* When each word starts and ends
* Quiet or slurred words
* Breaths and pauses
* Ad-libs
* Background vocals
* Overlapping singers

> You can sync the song using the separated vocal stem. It must come from the exact same source as the original, and you must perform the final timing review using the original song.

## What you need

* A Windows or Mac computer
* An NVIDIA or AMD GPU, if available
* [Pymss Studio](https://github.com/pymss-project/pymss-studio)
* The song you want to sync
* The `becruily_deux.ckpt` model
* Headphones

FLAC or WAV is preferred, but a good-quality MP3 or M4A can also work.

## 1. Install Pymss Studio

Open the [Pymss Studio releases](https://github.com/pymss-project/pymss-studio/releases) and download the version matching your computer.

| Computer | Version |
|----------|---------|
| Windows with an NVIDIA GPU | Windows CUDA |
| Windows with an AMD GPU | Windows ROCm |
| Windows without an NVIDIA GPU | Windows CPU |
| Windows with a stable connection and limited storage | Windows Online |
| Apple Silicon Mac | macOS MLX |

Use the CUDA or ROCm version if you have a supported GPU. It will normally process songs much faster.

Install or extract Pymss Studio, then open it.

## 2. Download the starter model

Open the model browser in Pymss Studio and search for:

```text
becruily_deux.ckpt
```

Download or import that exact model.

This model is a good starting point because it creates both:

```text
song
├── vocals
└── instrumental
```

You do not need to download every model in the browser.

## 3. Check your song

Make sure you have the exact version you want to sync.

Watch out for:

* Clean and explicit versions
* Remasters
* Deluxe or album versions
* Radio edits
* Music-video audio
* Live recordings
* Extra silence at the beginning

Two versions of the same song can have different timing.

## 4. Separate the song

In Pymss Studio:


1. Add the original song.
2. Select `becruily_deux.ckpt`.
3. Choose an output folder and format.
4. Start the separation.
5. Wait for processing to finish.

The output folder should contain a vocal file and an instrumental file.

Give them clear names if necessary:

```text
song-original.flac
song-vocals-deux.flac
song-instrumental-deux.flac
```

## 5. Listen to the vocal stem

Play the separated vocal file.

You should hear the singer much more clearly, although some instruments or audio artifacts may remain.

Listen for:

* The first sound of each word
* The moment one word changes into the next
* The final consonant of a word
* Pauses between words
* Repeated or stretched syllables
* Quiet background vocals
* Ad-libs hidden behind the lead singer

The result does not need to sound perfect. It only needs to make the performance easier to understand.

## 6. Use it while syncing

Import the vocal stem into your timing tool and use it as your main syncing audio. Because the instruments are quieter or removed, word boundaries are usually much easier to hear.

Keep the original song available in case the model removes a vocal, creates an artifact, or makes a transition sound misleading.

While syncing:


1. Play the vocal stem.
2. Mark the beginning of the first word.
3. Mark each transition into the next word.
4. End the final word when the singer stops or pauses.
5. Replay the line using the vocal stem and adjust it as needed.
6. Check the original only when something sounds missing, unclear, or unnatural.

You do not need to switch back to the original after every line. Syncing the whole song against the vocal stem is fine as long as the stem stays aligned and the final review uses the original.

### Example

A line may sound like this in the original:

```text
I don't know what you're saying
```

The instruments might make `don't know` sound like one continuous word.

The vocal stem may reveal:

```text
I | don't | know | what | you're | saying
```

You can use those clearer transitions to time the line directly against the vocal stem.

## 7. Check background vocals and ad-libs

Separation is not always perfect. A model may accidentally remove a background vocal or place it in the instrumental file.

If something appears to be missing:


1. Listen to the original song.
2. Check the vocal stem.
3. Check the instrumental stem.
4. Follow the original song when the stems disagree.

Never remove a lyric only because it is missing from the separated vocal.

## 8. Perform the final review

After syncing the song:


1. Stop using the separated vocal.
2. Play the TTML against the original song.
3. Review every line.
4. Check the beginning and end of the song.
5. Make sure the timing still feels natural with the instruments present.

The listener will hear the original song, not your isolated vocal stem.

This final review is required even when the entire song was easy to sync using the vocal stem. The instruments can change how a transition feels, and the separation model may slightly distort or remove sounds.

## If the first result is bad

Start with Deux. Only try another model if it fails in a noticeable way.

### The vocal contains too many instruments

Try:

```text
mel_band_roformer_vocals_fv7b_gabox.ckpt
```

This model focuses specifically on creating a cleaner vocal stem.

### FV7b removes words or background vocals

Try:

```text
BS-Roformer-Resurrection.ckpt
```

Compare it with Deux and FV7b. Use whichever result makes that part of the song easiest to hear.

### You need a cleaner instrumental

Try:

```text
Inst_GaboxFv9.ckpt
```

An instrumental is mainly useful for checking whether a sound belongs to the singer or the music. You normally do not need it for basic lyric syncing.

### You want individual instruments

Use:

```text
BS-Roformer-SW.ckpt
```

This can separate stems such as vocals, drums, bass, guitar, piano, and other sounds.

Most TTML makers do not need this. Use it only when a normal vocal separation cannot reveal a difficult section.

## Which model should I use?

| Situation | Model |
|-----------|-------|
| First attempt | `becruily_deux.ckpt` |
| Cleaner vocals | `mel_band_roformer_vocals_fv7b_gabox.ckpt` |
| Alternative vocals | `BS-Roformer-Resurrection.ckpt` |
| Cleaner instrumental | `Inst_GaboxFv9.ckpt` |
| Individual instruments | `BS-Roformer-SW.ckpt` |

The simple rule is:

```text
Start with Deux.
If the vocals are unclear, try FV7b.
If FV7b removes something, try Resurrection.
```

## Common problems

### Processing is very slow

Make sure you installed the CUDA or ROCm version if your computer has a GPU. Close games and other applications using the GPU.

### Pymss Studio runs out of GPU memory

Close other GPU-heavy applications and try again. Larger models require more video memory.

### The vocal sounds metallic or watery

This is an AI separation artifact. Try another model, but ignore minor artifacts if the words remain easy to hear.

### Part of the vocal is missing

Check the original and instrumental files. The model may have mistaken the vocal for an instrument.

### The separated audio is early or late

Confirm that it came from the same original file. Do not use a stem created from a different release of the song.

### The timing works locally but not in Spotify

Your source may not match Spotify’s release. Check for a different master, edit, clean version, or silence at the beginning.

## Recommended workflow

```text
Get the correct song
        ↓
Run becruily_deux
        ↓
Listen to the vocal stem
        ↓
Sync the song using the vocal stem
        ↓
Check suspicious or missing parts against the original
        ↓
Perform the complete final review with the original
```

## Final checklist

Before finishing:

* The source is the correct version of the song.
* The vocal stem was created from that exact source.
* Missing words were checked against the original.
* Background vocals and ad-libs were reviewed.
* The complete song received one final uninterrupted review.

You do not need the cleanest vocal extraction ever created. You only need a stem that helps you hear the performance more clearly.