Free Caption Generator for Video
Drop in a video and get automatic captions, three words at a time with the spoken word highlighted, the style short-form video uses. It is transcribed on your device, you watch it play back with the captions on, and then you download it or open it in the free Video Editor to style the captions and export without a watermark. No sign-up, and your video never leaves your device.
Drop a video file
or click to browse — MP4, MOV, WebM. Audio files work too (you get a caption video).
100% in your browser. Your video stays on your device. The Whisper AI model downloads once (~75 MB) from our servers, then runs locally for every transcription. We can't access your audio because it never leaves your computer. Privacy policy.
Transcript
How long will your file take to transcribe?
Free runs 100% in your browser — private, $0, great for most files. For a long file or top accuracy, Pro runs on cloud GPUs in seconds.
What's the catch with free?
There isn't a hidden one — it's genuinely free, unlimited, and private, because it runs on your own device instead of our servers, so there's nothing for us to charge for. The only trade-off is time: your computer isn't as fast as a cloud GPU, so transcription takes a few minutes. For short clips that's barely noticeable. For long files the wait adds up — and that's the one case where Pro earns its keep: the same job on cloud GPUs in seconds, at higher accuracy, on any device.
Free vs Pro — typical transcription times
| File length | Free — in your browser* | Pro — cloud |
|---|---|---|
| 10 min | ~2–3 min | ~22 sec |
| 30 min | ~7–8 min | ~45 sec |
| 1 hour | ~15 min | ~1.5 min |
| 2 hours | split into ≤60-min parts first | ~2.5 min |
*Free times on a typical laptop with a WebGPU browser (Chrome/Edge); older devices without a GPU can be 2–4× slower. Free is capped at 60 minutes per file. Pro runs on cloud GPUs at the same speed on any device, using the larger large-v3 model for higher accuracy.
What makes free transcription faster?
- Use a WebGPU browser — Chrome or Edge (and the latest Firefox) run it on your GPU, roughly 2× faster than a CPU-only browser. Brave disables WebGPU by default.
- Keep your laptop plugged in — battery-saver mode throttles your CPU/GPU and can nearly double the time.
- Close other heavy tabs and apps — transcription uses your CPU/GPU and memory; free them up.
- A device with a GPU is much faster — a recent phone or laptop with WebGPU beats an old CPU-only machine by a wide margin (a modern phone can even out-run a no-GPU laptop).
- The first run downloads the model once — a one-time wait; every transcription after that skips it.
Speed by device & browser — a 16-minute file
| Your setup | Time to transcribe (16 min) |
|---|---|
| Laptop or phone with WebGPU (Chrome, Edge, latest Firefox) | ~4 min |
| Phone, CPU only (no WebGPU) | ~5–6 min |
| Laptop without WebGPU (Brave default, older browser) | ~8 min |
| Old laptop, single-core fallback | up to ~18 min |
| Any device with Pro (cloud GPUs) | ~29 sec |
Measured on our own benchmarks with the Whisper base model in-browser (large-v3-turbo for Pro). Your exact time varies with your device's chip and the audio itself.
How do I auto caption a video for free?
Drop the video on this page and click Transcribe. The speech is transcribed in your browser with word-level timing, and the video plays back with captions on: three words at a time, with the spoken word highlighted. Download that video, or open it in SnipSound’s free Video Editor to style the captions and export without a watermark. There is no sign-up, and your video never leaves your device.
Style your captions in the Video Editor
- No watermark: the quick download here carries a small snipsound.com stamp; the editor’s export does not.
- Font, size, colour and position: every caption is a clip you can restyle and retime.
- Trim, cut and add music: free background music from the built-in library, lowered automatically under the speech.
- Export for TikTok, Reels, Shorts or YouTube, rendered on your device.
Auto captions made for short-form video
Most short-form video is watched with the sound off, so captions decide whether people keep watching. This AI caption generator runs a speech model in your browser that times every word, so the captions follow the speaker word by word instead of appearing as long blocks of text. Most online caption generators process your video on their servers; here the transcription, the preview and the final export all happen on your own device.
Made for
- TikTok, Reels and Shorts creators: burned-in captions for viewers scrolling on mute.
- Podcast and interview clips: turn a talking-head clip into something people read along with.
- Course makers and trainers: captions make lessons easier to follow.
- Accessibility: captions make your video usable for deaf and hard-of-hearing viewers, and for anyone somewhere loud.
Captions or subtitles?
They come from the same transcript. Captions here means text burned onto the video for social clips. If you want subtitles for a longer video, or a file viewers can switch on and off, use Add Subtitles to Video or the SRT Generator. Only need the words? The Audio Transcription tool gives you an editable transcript. For platform-by-platform steps, see how to add captions to a video.