4. Sync the lyrics
Key | Use it when… |
|---|---|
| the selected word begins |
| the next word begins; it ends the current word and starts the next one |
| the selected word ends and no next word should start yet |
| you need to select the previous / next word |
| you need to rewind / fast-forward audio by 5 seconds |
Syncing Process
This is a fairly lengthy process and requires precision. Here is a detailed breakdown:
Go to the
Timetab.Select the first word of a line.
Press
Spaceto start the audio.Press
Fat the exact moment the singer starts singing the first word. This sets its start time.Press
Gas the singer transitions to the next word. PressingGsets the end time of the current word and the start time of the next word simultaneously.Keep pressing
Gat the beginning of each subsequent word in the line.Press
Honly when the singer reaches a pause, stops to breathe, or reaches the end of a line. PressingHsets the end time of the current word without starting a new one.When the singer starts singing again after a pause, press
Fon the next word to set its start time, then resume pressingGfor the following words.Use
A(previous word) orD(next word) to adjust selection, and the arrow keys←/→to jump backward or forward in the audio by 5 seconds if you make a mistake.Double-check the synchronization for the line before moving on to the next one.
When you have finished synchronizing all lines, perform a final check and proceed to the next step.
Avoid using
FandHfor every word unless there is an actual pause after that word. Otherwise, you should useGto ensure seamless transitions between words.If a vocal passage is too fast, hover over the music icon and reduce playback speed. Precision matters more than speed.
Ensure you have close to no audio latency (use wired headphones if possible) to maintain high accuracy.
Use the spectrogram to refine timing
A spectrogram shows sound over time:
Left to right is time.
Bottom to top is frequency, from lower to higher sounds.
Brighter or stronger colors show more energy at that time and frequency.
It does not identify vocals automatically. Drums, cymbals, guitars, and other instruments can create shapes that resemble consonants or word boundaries. Use the spectrogram to support careful listening—not to replace it.
Open and navigate it
Load the exact audio used for the TTML and go to the
Timeview.Select the line you want to inspect.
Use the Expand / collapse spectrogram chevron at the far right of the audio bar.
Scroll normally to move along the timeline. Hold
Ctrlwhile scrolling to zoom around the pointer.Drag the top edge of the panel if you need more vertical space.
Drag a visible word or line boundary to adjust its timestamp. Right-click a timed word to audition only that word.
While a spectrogram selection is active, the default audition shortcuts are Q for the 500 ms before it, S for the selection, and W for the 500 ms after it. Keybindings can be changed, so the tool’s displayed bindings remain authoritative.
The settings button inside the panel controls gain, frequency resolution, height, and appearance. Start with the defaults. Increase gain when quiet details are invisible; change the FFT size only when you need a different balance between time detail and frequency detail.
Find the beginning of a word
Place the start at the first audible part of the performed word, which may be a quiet consonant before the vowel.
Common visual clues include:
Vowels and other voiced sounds: sustained horizontal bands or a harmonic stack.
Sibilants such as s, sh, and f: diffuse energy higher in the display, sometimes before the vowel becomes obvious.
Plosives such as p, t, k, b, d, and g: a brief burst or sharp transition.
Breathy entrances: faint, spread-out energy that may be hard to distinguish from the mix.
Do not snap every start to the brightest vowel. That often makes words beginning with quiet consonants appear late.
Find transitions between words
For continuously sung words, the boundary is often a change in articulation rather than silence. Begin with G during normal syncing so the previous word’s end and the next word’s start share a timestamp, then refine the shared divider in the spectrogram.
Listen for the point where the next word becomes perceptible. A consonant may belong to the beginning of the next word even when the vowel energy from the previous word continues underneath it.
Preserve a real pause when one exists. Do not force adjacent timestamps merely because two shapes touch visually.
Find the end of a word or line
Follow the vocal until the performed sound actually ends. A held vowel may remain visible as fading harmonic bands after its loudest point.
Do not automatically extend the lyric through:
instrumental sustain that continues after the voice;
cymbals or percussion at the same moment;
an echo or reverb tail that no longer sounds like the active performed word;
background vocals belonging to another lyric line.
Move the end boundary, audition the word and its surrounding audio, and check whether the transition feels early, late, or natural.
A practical refinement loop
Sync the passage by ear first.
Zoom into one questionable boundary.
Compare the visible pattern with the audio before, inside, and after the current timing.
Move the boundary by a small amount.
Audition it again.
Check the result in Preview at normal playback speed.
For example, if stay visibly begins with faint high-frequency energy before its strong vowel, listen for the initial s. Move the start earlier only if that faint region is actually the consonant. If the timing continues across a bright snare after the vocal ends, pull the end back to the audible end of the voice.
Difficult mixes and vocal stems
Reverb, doubled vocals, harmonies, and overlapping instruments can make one word appear to have several beginnings or endings. Compare neighboring repetitions and use context, but do not copy a boundary blindly—the performances may differ.
A separated vocal stem can make quiet consonants and background vocals easier to hear. Separation may also add artifacts or soften transients, so sync and perform the final check against the original recording.
Common mistakes
Timing from the image without listening.
Starting at the brightest vowel and cutting off the opening consonant.
Ending at the end of an instrument instead of the voice.
Treating every visible gap as a pause between words.
Using extreme zoom and losing the phrase’s musical context.
Making millisecond-level changes that do not improve the audible result.
Assuming copied repetitions have identical delivery.
The spectrogram is strongly recommended for difficult sections but is not required for approval. Accuracy in the final playback matters; using a particular visual tool does not.