Speech-to-text built into the workflow
There is no transcript to copy between tools before you can see captions on the clip. Transcription, editing, styling, and export are one continuous path.
Upload a talking video and get timed captions ready to review, restyle, reposition, and export on the footage itself. Correction is part of the path, not an afterthought.
Drop a video here
MP4, MOV, WebM or MKV · up to 500 MB · 10 minutes
No signup. No watermark. Guest projects stay editable for 24 hours, or 7 days once you sign in.
Upload the clip from your browser and let the app prepare its audio for transcription. The file goes straight to storage rather than through a server that re-encodes it.
AI converts the speech into caption segments with timing for each spoken word, not merely for each line. That detail is what later lets a style follow the voice.
Review the transcript, fix the words the model misheard, adjust line breaks, and apply a caption style against the real footage.
Download a captioned MP4, or a subtitle file for another editor or platform.
There is no transcript to copy between tools before you can see captions on the clip. Transcription, editing, styling, and export are one continuous path.
Caption styles can emphasise the current word because the transcript carries per-word timing. That is what makes karaoke and active-word treatments possible at all.
Correction is a normal step rather than a hidden one. The product does not pretend transcription is infallible, because it is not.
Fixing a word rewrites that word alone. The timings on either side are preserved, so a correction does not push the rest of the caption out of sync.
A file is recognised by its checksum, so re-uploading the same video does not transcribe it again, and does not cost you a second time.
Upload, transcription, styling, and preview all run without signing in. The wall is at download, once you have already seen the result.
Plenty of tools will hand you a transcript. The gap between a transcript and a publishable video is where most of the work actually is.
Raw output arrives as a wall of text with line breaks in arbitrary places. Captions need line lengths that fit the frame and breaks that fall where someone pauses.
Speech models handle ordinary sentences well and stumble on names, products, and jargon: precisely the words your audience recognises. Correction is one click, not a re-run.
Contrast and placement can only be judged against the actual shot. A caption that reads perfectly on white is invisible over a bright sky.
If the transcript is in a script none of the bundled fonts can draw, the export is blocked rather than returned as rows of empty boxes after the wait.
Automatic transcription is not universally appropriate. These are the cases where it will disappoint you.
The speech transcribes, but no bundled font can draw those characters, so a burned-in export would come back as empty boxes. The export is refused instead. SRT and VTT still work and keep the text intact.
Court, medical, and compliance work needs a human transcriptionist and an accuracy warranty. This produces captions for video, not certified records.
Heavy overlap between speakers, loud music over dialogue, or very distant microphones give any model little to work with. Fix the audio first and the captions follow.
It transcribes speech from a video, creates timed caption segments with per-word timing, and opens them in an editor for review and styling.
Yes. Open the script panel or double-click a word on the video. The surrounding timing is preserved, so a fix does not desynchronise the rest.
Yes. SRT, VTT, and TXT downloads are available alongside burned-in MP4 export, and they are free on every plan.
No. Neither the free nor the Starter plan adds a watermark. The free limit is how many videos you can download, not how good they look.
No. Upload, transcription, styling, and preview work without signing in. An account is required only at download.
It is recognised by its checksum and the cached transcript is reused, so the second upload is effectively instant and does not transcribe again.
Speech in many languages transcribes, but burning captions in requires a bundled font that can draw the script. Latin, Cyrillic, Greek, Devanagari, Arabic and Hebrew are covered; CJK and Thai are not yet.
7 days on the free plan and 90 on Starter, after which the source file and its project are deleted. Guest uploads are removed after one day.