Skip to main content
Back to homeYour media stays on your computer

Free AI video caption generator

AI Transcription and Multi-Speaker Identification

No signup · 99 languages

See captions in action

Multiple speakers · 16:9

Multiple speakers · 9:16

Word highlighting

Just need a transcript? Open the audio transcriber.

FAQ

How do I add captions?

Choose audio or video, then generate captions. Review the text, timing, and speakers. For video, choose a style and create a captioned MP4 or WebM. You can also download a transcript or subtitles.

Do my edits apply to every download?

Yes. MP4, WebM, TXT, SRT, VTT, CSV, and JSON use your corrected caption text. Downloads that contain timestamps or speaker labels also use your timing and speaker corrections. Deleted captions are left out of every download. RTTM contains only speaker labels and times, so there is no caption text to update.

Can I correct words without losing word highlighting?

Yes. A word correction keeps its original timing. Adding or removing words estimates timings around that edit; changing a caption’s start or end adjusts its word timings to fit. Highlighting stays available in the preview and captioned video. JSON includes the corrected words and marks estimated timings. Review the result after larger edits.

Can I caption just part of a video?

Yes. After choosing a file, use Start and End under Section to transcribe to choose the part you want. Enter times in seconds. By default, the tool selects the whole file or its first 10 minutes, whichever is shorter. Your captioned video contains only that section.

Where do subtitle timestamps start?

Zero means the start of the section you captioned. If you set Start to 120 seconds (2:00), speech at 2:05 in the original video appears at 0:05 in your subtitles and captioned video. JSON also records where that section starts in the original video.

Is it free and private?

Yes. No signup, payment, or added watermark. Your media and transcript stay on your device. Processing runs in your browser.

Can I keep captions clear of TikTok, Reels, and Shorts controls?

For vertical video, choose a platform under Keep captions clear of. This adjusts caption placement to leave room for that platform’s controls. Vertical videos start with TikTok margins. Horizontal and square videos use standard margins and do not show the platform selector. You can show a safe-area outline and platform overlay in the preview; neither guide appears in your exported video. App controls vary, so check the result before posting.

Can it tell different speakers apart?

Turn on Identify multiple speakers before generating. Review the labels, rename speakers, or add a missed speaker and reassign captions. It supports up to eight speakers. It does not identify who a person is. The transcript follows one speech stream, so overlapping voices may need correction.

Can I change caption and speaker colors?

Choose White or Black under Caption color to use one color for everyone. Choose By speaker, or enable Use a different caption color for each speaker, to use speaker colors. Change each speaker’s color beside their name. Colors stay attached to that speaker and appear in preview, video downloads, and compatible VTT players. JSON saves your color choices; SRT and TXT use plain text. Dark text gets a light outline or background so it stays readable.

How does word highlighting work with speaker colors?

Speaker colors include speaker names and replace word highlighting. To highlight each spoken word, choose the Word highlight style and set Caption color to White or Black, with Show speaker names turned off.

Which languages are supported?

The Whisper Tiny and Base transcription models support 99 spoken languages, with automatic detection available. Captions stay in the spoken language; they are not translated. Accuracy varies by language, accent, and recording quality. This language support applies to transcription; speaker identification uses a separate model.

Can I identify speakers in every caption language?

Speaker identification uses Nemotron 3 Diarization to tell voices apart and label who spoke when. You can enable it with any transcription language. The model was trained and evaluated on multilingual audio, but accuracy is not guaranteed across all 99 languages. Review the speaker labels, especially when voices overlap.

What files and clip lengths can I use?

Choose audio or video up to 600 MB and a section up to 10 minutes. Common formats include MP4, MOV, WebM, MP3, M4A, and WAV. If your browser cannot read a file, use Prepare compatible copy. Video downloads preserve the aspect ratio at up to 1080p.

Which download format should I choose?

MP4 or WebM adds captions to your video. SRT and VTT are subtitle files for other editors and players. TXT is a plain transcript. CSV includes caption text, speakers, and times. JSON includes caption settings and timing data. RTTM is for tools that use speaker timelines.

What does Save video as it’s created do?

Choose a filename and save location before creating the video. The tool writes the video to that file as it processes, using less browser memory for larger exports. When processing finishes, the file is already saved. Leave this off to create the video first, then use Download video. This option is available in browsers that support it.

Which browser and transcription model should I use?

Use desktop Chrome or Edge. Tiny is faster; Base can improve recognition. Mobile processing may be slow or fail. Keep this tab open while processing.

Why does the first run take longer?

The first run downloads the speech model. Tiny uses about 44–154 MB; Base uses about 80–294 MB, depending on your device. Speaker identification adds about 106 MB. Your browser caches downloads when storage allows. Processing can take several minutes.