Splitting Guidelines

This page defines what is accepted, recommended, and incorrect when splitting words into timed parts. To learn the tool controls first, read Basics Of Splitting.

Main standard

Splits should be based on how the singer pronounces the word in the recording, not only on its spelling, grammar, or usual dictionary pronunciation.

A word may have different correct splits in different recordings. Listen to the version being timed and judge the syllables the singer actually performs.

The written lyric must stay the same. Splitting divides the existing text; it must not respell the word to imitate the pronunciation.

Accepted splitting choices

The following are valid:

  • Leaving a word unsplit.

  • Splitting every syllable the singer clearly performs.

  • Using only some correctly placed syllable splits. This is allowed, but consistent splitting across the song is recommended.

  • Using fewer splits for words that are sung quickly.

  • Choosing whether emphasis changes the placement of a split. This is left to the maker.

  • Keeping a real audible gap between split parts.

  • Using a shared boundary when split parts flow continuously with no audible gap.

Accurately splitting every performed syllable is valid. A TTML must not be rejected merely because it contains many accurate splits or produces detailed animation.

Incorrect splitting

A split is incorrect when it separates one sung syllable into smaller sounds.

For example, thi / nk is incorrect when think is sung as one syllable. Holding the th, changing pitch, or stretching the word across several notes does not create another syllable.

This is what oversplitting means in this guide: creating more timed parts than the singer performs as syllables. It does not mean accurately splitting every clearly performed syllable.

A split is also incorrect when it follows a pronunciation that is not used in the recording.

Examples

These examples are valid only when they match the pronunciation in the recording.

Pronunciation in the recording

Accepted split

two-syllable “fire”

fi / re

two-syllable “every”

eve / ry

two-syllable “different”

diffe / rent

two-syllable “chocolate”

choco / late

ordinary “making”

ma / king

ordinary “running”

run / ning

ordinary “hundred”

hun / dred

two-syllable “react”

re / act

syllabic-consonant “little”

lit / tle

three-syllable “family”

fa / mi / ly

compressed two-syllable “family”

fami / ly

The two family examples are intentionally different. Neither spelling should be applied automatically to every performance.

Automatic splitting

Automatic splitting is a starting point, not final proof that the splits are correct.

Every automatic result must be checked against the recording. Correct any split that uses the wrong pronunciation or places a boundary somewhere the singer does not perform one.

Timing standards

Timing must follow the recording:

  • When two split parts flow continuously, they should touch at a shared boundary.

  • When the singer leaves a real audible gap, preserve that gap.

  • Do not stretch either part across silence just to make the timings touch.

A spectrogram or isolated vocals may be used as supporting evidence, but the recording remains the source of truth. If the pronunciation or boundary is unclear, ask the maker about the intended pronunciation or ask an experienced reviewer for another opinion.