Speaker diarisation: identifying who spoke when
Speaker diarisation estimates when the voice changes; it does not inherently know that a voice belongs to Dmitry, the moderator or an audience member.
Separate words from voices
Speech recognition creates words, while diarisation creates voice intervals. Combining them does not repair recognition errors.
Time-aligned voice clusters and recognised words are separate outputs; combining them can introduce boundary errors without changing either model.
Explain unknown labels
SPEAKER_00 is a normal output without enrolled voices or participant metadata. Assign real names from introductions, channels, video or manual review—not from topic guesses.
Names require evidence such as introductions, isolated channels or manual confirmation, and unresolved identities should remain explicitly unresolved.
Test hard transitions
Overlap, echo, short interjections, similar voices and edited sources can split one person or merge several people. Tune on representative samples.
Measure both false speaker splits and accidental merges on overlap, short responses and transitions rather than reporting only a speaker count.
Prefer source separation
Independent microphone channels are often more reliable than recovering speakers from a mixed track. Verify all changes of speaker and unresolved identities.
Where capture equipment provides one channel per microphone, retain that structure instead of reconstructing it later from a mixed recording.
Continue with the underlying material
Need an estimate for your project?
Tell us about the project. We will break it into stages and explain the budget drivers.