Fix crosstalk by hand
When two people speak at once, words from both can land in one caption or be misheard. Correct the words, then drag a caption's start or end on the timeline so it lines up with who is audible.
Start from the clip you already cut from a recorded episode. Generate the captions, correct the names and the moments where both people talk at once, keep the text clear of faces, then export the MP4.
Drop a video here
MP4, MOV, WebM or MKV · up to 500 MB · 10 minutes
No signup. No watermark. Guest projects stay editable for 24 hours, or 7 days once you sign in.
Export the moment you want to share from your editing or recording software as MP4, MOV, WebM or MKV. This page does not find highlights or trim an episode.
Correct names, show titles and jargon in the script panel. Listen again wherever both speakers overlap.
Drag and resize the caption block on the frame so it sits below or between the speakers rather than across a mouth.
Download the burned-in MP4 for social feeds, or SRT or VTT when a player draws its own captions.
When two people speak at once, words from both can land in one caption or be misheard. Correct the words, then drag a caption's start or end on the timeline so it lines up with who is audible.
The script panel lists the whole transcript in order. Read it through, fix each spelling of a name or brand, and the timing of the surrounding words is kept.
Two-shot and split-screen podcast layouts leave different space free. Move the caption block on your actual footage rather than trusting a fixed bottom position.
Word-level timing lets styles follow the words being spoken. Every style can be previewed; Clean Minimal and Documentary export on Free.
Burn captions into one MP4 and post it where you publish clips. The platform guides cover each app's interface and canvas.
Some podcast tools start from a full episode, find clips and label speakers. This workflow starts later, from a clip you already chose.
No automatic highlight detection or trimming. That keeps the edit in your hands and means the clip has to exist before you upload.
Captions are not split or labelled by voice. If you need names on screen for each speaker, that is not available in this editor.
An audio-only MP3 or WAV cannot be uploaded, and there is no RSS or episode-link import. Export a video clip from your recording first.
The time goes into reading the transcript on the real footage, which is where overlapping speech and names are caught.
Worth knowing before you upload an episode.
Free accepts videos up to 10 minutes. A full episode needs Starter, which accepts up to 120 minutes and 5 GB; anything longer has to be split first.
There is no highlight detection, automatic trimming, social account connection or scheduled posting. The output is a file for you to publish.
Captions are in the spoken language and are not labelled by speaker. Use a transcription service built for interviews when a named, speaker-separated transcript is the deliverable.
No. The uploader accepts MP4, MOV, WebM and MKV video files. Export a video clip from your recording or editor first.
No. Captions are generated from the speech without identifying speakers. Correct the words and timing yourself where two voices overlap.
No. It captions a clip you have already cut. Choose and trim the moment in your editing software, then upload that clip.
Up to 10 minutes and 500 MB on Free, and up to 120 minutes and 5 GB on Starter.
No. Free counts 3 distinct video projects a month. Downloading a corrected version of the same clip that month uses the slot it already has.
Drag the caption block on the video and resize it until it sits in free space, usually below both speakers or between them in a split-screen layout. Check the result in the preview before exporting.