The job starts before voice generation
Translating a video means rebuilding its audio layer. The service must first extract speech, mark the beginning and end of each line, distinguish the speakers and prepare a contextual translation. A good generated voice cannot rescue errors made at those earlier stages.
We built GistiQ Dub as a standalone project workspace. A user uploads a file or supplies a link, then works with one connected set of source text, translation, speakers, generated clips and final outputs rather than moving unrelated files between tools.
The service first identifies who speaks and when
After upload, the service separates the soundtrack and transcribes speech with timing information. It then identifies which lines belong to different participants. This matters for interviews, podcasts and courses, where a host and guest should not suddenly speak with the same voice.
Line boundaries remain editable. If a short reaction has been attached to the next speaker, or two people talk over one another, the problem can be corrected before it spreads into translation and voice generation.
A translation must preserve the thought and fit the scene
The text is translated with neighbouring lines as context. That helps preserve references, pronouns and the meaning of short replies. Timing matters too: an accurate translation may be much longer than the original and fail to fit the scene.
Names, product terms and recurring expressions are collected for review. When a line overruns its available time, the editor can shorten the wording, adjust that line’s pace or deliberately retain part of the original rather than hiding the conflict.
Each speaker keeps a consistent voice
Different licensed voices can be assigned to the detected participants. The choice is stored for the speaker across the whole video, and changing it only requires the associated clips to be generated again.
A supplied voice sample can also be used when the owner has consented and the intended use is clear. The capability is not permission to impersonate anyone. For ordinary localisation, voices with established usage rights remain the safer default.
Editing happens at line level
Names, product terms, humour and tone routinely need human correction. The workspace therefore keeps the source line, translation, speaker and time range together. An editor can see exactly what will be heard and where.
The editor can change the wording, select another voice, replace a generated clip with a recording, or regenerate a single line. One mistake in an hour-long video should not restart transcription and translation for the entire project.
Assembly includes the rest of the soundtrack
Prepared lines are placed back into their source time ranges and combined into a new track. Music, room sound and other ambience still matter: replacing all original audio would make even a good translation feel detached from the scene.
If a translated line is too long, the editor sees the conflict before final assembly and can shorten the wording or adjust that line. The finished file is produced only after the text and sound have been reviewed.
Long processing remains understandable
A long recording with several speakers cannot be processed instantly. The interface shows completed stages and specific failures, so users do not have to keep a page open and guess whether a task has disappeared.
Intermediate work is preserved. If one line fails or final assembly stops, the service resumes from the necessary stage instead of repeating work that has already completed successfully.
The deliverable
The user receives a finished video with its translated soundtrack and an editable project that can be revisited. Courses, interviews, presentations and recurring programmes can therefore use a repeatable editorial process rather than a one-off manual assembly.
Automation removes mechanical work but not language and ethics review. A fluent speaker should listen before publication, names and claims must be checked, and any use of a recognisable person’s voice requires explicit permission.