Guide · The tools
Subtitles and dubbing
Subtitles from your own video, translated and kept in time, and a dubbed audio track in another language — all on your device.
Updated · 2 min read
Video has words in it too. Subtitles turns the speech in your video into a subtitle file, translated, with the timing preserved; Dub goes one step further and speaks the translation over the video. Both run in your browser, on your device.
Subtitles
Drop a video or an existing subtitle file (.srt or .vtt). Speech is transcribed on your device by the speech model, which is downloaded once the first time you use it from a public model host and kept locally; the page names the host and the size before it starts, and asks. Each cue is translated with its neighbours in view, so a sentence split across two cues is translated as one sentence and put back in two, and the timestamps are never touched. Export as .srt or .vtt, in the target language or as a bilingual file.
A long video takes a while — transcription is the slow part — and the status line counts cues and shows a measured estimate. Stop keeps what is finished.
Dub
Dub takes the translated cues and speaks them with an on-device neural voice for the target language, timed to the original cues, and mixes the result over the video's own sound with the original speech lowered. You get a preview to play, and exports: the dubbed audio track (.wav), the translated subtitles (.srt) and a bilingual .srt. Combining them into a new video file arrives with the desktop engine; until then a video editor does that step. Voices are downloaded once per language and kept locally; where no on-device voice ships for a language yet, the page says so and asks you to pick another.
Limits, honestly
- Transcription quality follows audio quality. Music under speech, crosstalk and heavy accents produce more errors; review the cues before dubbing.
- A translation that is much longer than its source has to be spoken faster to fit its cue. The page shows which cues are tight so you can shorten them.
- Speaker identification is not attempted: one voice per language.
- These tools are in the browser app and the desktop app. The phone apps carry neither.
- Is my video uploaded?
- No. Transcription, translation, speech and mixing all run in your browser. The only downloads are the speech model and the voice, once each, from a public model host the page names.
- Which formats?
- Any video your browser can play. Subtitles in and out are .srt and .vtt.