Free AI Audio & Video Transcription
SnipSound's free transcription tool turns audio or video into text in 99 languages, right in your browser — turn a voice recording, podcast, interview or voice memo into text (transcribe audio to text, voice to text, convert audio to text, sound to text). Get word-level timestamps with karaoke highlighting, make a captioned video, bleep any word, and export .txt, .srt, or .vtt. Files never leave your device. No sign-up.
Drop an audio or video file
or click to browse.
Don't have a file? Record one with our voice recorder to test how transcription works.
100% in your browser. Audio stays on your device. The Whisper AI model downloads once (~75 MB) from our servers, then runs locally for every transcription. We can't access your audio because it never leaves your computer. Privacy policy.
Transcript
How long will your file take to transcribe?
Free runs 100% in your browser — private, $0, great for most files. For a long file or top accuracy, Pro runs on cloud GPUs in seconds.
What's the catch with free?
There isn't a hidden one — it's genuinely free, unlimited, and private, because it runs on your own device instead of our servers, so there's nothing for us to charge for. The only trade-off is time: your computer isn't as fast as a cloud GPU, so transcription takes a few minutes. For short clips that's barely noticeable. For long files the wait adds up — and that's the one case where Pro earns its keep: the same job on cloud GPUs in seconds, at higher accuracy, on any device.
Free vs Pro — typical transcription times
| File length | Free — in your browser* | Pro — cloud |
|---|---|---|
| 10 min | ~2–3 min | ~22 sec |
| 30 min | ~7–8 min | ~45 sec |
| 1 hour | ~15 min | ~1.5 min |
| 2 hours | split into ≤60-min parts first | ~2.5 min |
*Free times on a typical laptop with a WebGPU browser (Chrome/Edge); older devices without a GPU can be 2–4× slower. Free is capped at 60 minutes per file. Pro runs on cloud GPUs at the same speed on any device, using the larger large-v3 model for higher accuracy.
What makes free transcription faster?
- Use a WebGPU browser — Chrome or Edge (and the latest Firefox) run it on your GPU, roughly 2× faster than a CPU-only browser. Brave disables WebGPU by default.
- Keep your laptop plugged in — battery-saver mode throttles your CPU/GPU and can nearly double the time.
- Close other heavy tabs and apps — transcription uses your CPU/GPU and memory; free them up.
- A device with a GPU is much faster — a recent phone or laptop with WebGPU beats an old CPU-only machine by a wide margin (a modern phone can even out-run a no-GPU laptop).
- The first run downloads the model once — a one-time wait; every transcription after that skips it.
Speed by device & browser — a 16-minute file
| Your setup | Time to transcribe (16 min) |
|---|---|
| Laptop or phone with WebGPU (Chrome, Edge, latest Firefox) | ~4 min |
| Phone, CPU only (no WebGPU) | ~5–6 min |
| Laptop without WebGPU (Brave default, older browser) | ~8 min |
| Old laptop, single-core fallback | up to ~18 min |
| Any device with Pro (cloud GPUs) | ~29 sec |
Measured on our own benchmarks with the Whisper base model in-browser (large-v3-turbo for Pro). Your exact time varies with your device's chip and the audio itself.
Free, private AI audio & video transcription — how it works
SnipSound's transcription tool uses OpenAI's open-source Whisper speech-recognition model running entirely in your browser via WebAssembly. Drop in audio or video — for a video we extract the audio track automatically. The first time you click Transcribe, your browser downloads a ~75 MB model file from our servers; after that, every transcription is fully local. Your file never gets uploaded to any server — not ours, not OpenAI's, not anyone's.
Working specifically with video? Try Video to Text. Just need a subtitle file? The SRT Generator exports ready-to-use .srt/.vtt.
What it's good for
- Transcribing podcast interviews, meeting recordings, voice memos, lectures, or any clear-speech audio.
- Adding subtitles to a video — generate automatic subtitles as .srt and .vtt downloads with accurate timestamps that drop into YouTube, Vimeo, or any editor.
- A free, private alternative to Otter, Rev, and Descript for journalists, researchers, students, and creators — no subscription, no upload, no per-file time cap.
- Privacy-sensitive audio you don't want on a third-party server — therapy notes, confidential interviews, internal meetings.
- Word-level timestamps — karaoke word highlighting and click-to-jump, so you can pull an exact quote or build precise captions.
- Turn audio into a captioned video — one-click word-by-word caption video for podcasts, voice notes and social clips, downloadable and shareable.
- Bleep or censor words — search any word and mute every occurrence, then download clean audio (pairs with our Profanity Remover).
Edit the transcript — fix mistakes, find & replace
Auto-transcription is never perfect, so the transcript is fully editable. Hit ✎ Edit and click any line to retype a misheard word — your changes save automatically and flow into every export (.txt, .srt, .vtt, and the rest). Use Find to jump to every place a word appears, and Replace all to fix it everywhere at once — perfect for correcting a name, brand, or term Whisper spelled wrong throughout (say, every "Mark" → "Marc"). The matching is whole-word, so replacing "a" won't touch the letter a inside other words. Prefer the raw output? Flip between the Original and Edited tabs anytime — the untouched original stays saved with the file.
What it's not so good for
- Heavy background noise, music behind voice, or multiple overlapping speakers — tiny Whisper struggles with these.
- Heavy accents or non-mainstream dialects — bigger Whisper models handle these better but are too heavy for a browser.
- Speaker diarization ("who said what") — not supported by Whisper-tiny.
- Files longer than 60 minutes — we cap input length to keep browser RAM under control.
Translate audio to English
Tick "Translate to English" and Whisper renders any non-English audio as English text. Spanish podcast → English transcript. Mandarin interview → English notes. Dedicated Audio Translator tool here if translation is your primary need.
Transcribe audio in 99 languages
SnipSound transcribes speech in 99 languages, with a searchable language picker and real automatic language detection — it runs a genuine language-ID pass, which most in-browser tools can't. On clear speech, accuracy is strong across major languages including English, Spanish, French, Portuguese, Italian, German, Russian, Chinese and Japanese — so it handles "transcribir audio a texto", "transcrire audio en texte", "trascrizione audio", "音声文字起こし" and more, right in your browser.
A couple of honest notes so you know what to expect: accuracy figures are for clear speech — heavy accents or noisy audio reduce accuracy on any on-device model — and a few low-resource languages (Hindi/Urdu among them) are weaker on the free model, which is why the language picker marks a limited-accuracy tier for them; for those, the Pro model does better. For Chinese and Japanese, transcription and captions are excellent; word-by-word click precision is strongest on Latin-script languages. Localized pages for major languages are rolling out — or just open the tool and pick your language (or let auto-detect do it).
More than a transcript
Beyond speech-to-text, the same free tool reads long files aloud as they transcribe (streaming results), lets you edit a word while keeping its exact caption timing, shows a most-said-words analysis, and finds and plays every instance of a word. When you're done, one click sends your video plus the word-timed transcript straight to the free video editor with captions already on the timeline — and re-opening a file later restores its transcript instantly.
Review mode is the standout: the AI underlines any word it wasn't fully confident about, and a Review button walks you through each one, replaying just that word's audio so you can keep or fix it in a tap — and your verdicts are saved. There's more: search filters the transcript to matching sentences and shades where they fall on the waveform, find & replace-all fixes a misheard name everywhere at once (timings kept), edits auto-save and restore on re-upload, the caption video matches your source's aspect ratio with its own playback controls, and there's a fullscreen transcript view with sentence-level timestamps.
SnipSound vs Otter.ai, Rev & Descript
People usually weigh SnipSound against paid cloud transcribers like Otter.ai, Rev and Descript. The core difference: those upload your audio to their servers and charge a monthly subscription or per-minute fee, while SnipSound runs free in your browser with nothing uploaded — only the optional Pro tier costs anything, and it is pay-per-use with no subscription.
| Feature | SnipSound | Otter.ai | Rev | Descript |
|---|---|---|---|---|
| Cost | Free — Pro is pay-per-use, no subscription | Freemium, paid monthly | Paid per-minute / subscription | Paid monthly |
| Your audio | Stays in your browser — never uploaded | Uploaded to their servers | Uploaded | Uploaded |
| Account required | None | Required | Required | Required |
| Languages (model support) | 99 | Fewer | Fewer | Fewer |
| Word-level timestamps | Yes | Yes | Yes | Yes |
| Make a captioned video | Yes, built-in | No | No | Yes (paid) |
| Free per-file cap | 60 min, unlimited files | Limited free minutes | Trial only | Limited free |
A free, private Otter.ai alternative for anyone who does not want their audio on a third-party server. For the hardest audio, SnipSound Pro runs the largest model on cloud GPUs, pay per use.
How accurate is it — and when is Pro worth it?
On clean, single-speaker audio the free in-browser model (OpenAI’s Whisper base) is very accurate. In our own tests, English and Spanish came back around 100%, Chinese and Japanese ~100% content-accurate, and French around 88%. On a standard clean-speech benchmark that is roughly 95%; the Pro model (Whisper large-v3) is around 98%. Real-world audio with background noise, music, strong accents or overlapping speakers is harder for any model, so accuracy there is lower — we never claim 99%.
What the free model struggles with (and where Pro helps)
- Strong accents, proper nouns, names, jargon and acronyms — the smaller free model mis-hears rare words more often; Pro’s larger model gets them right more of the time.
- Background music or noise under the voice, and overlapping speakers — both are hard for a browser-sized model, and neither model labels who said what.
- Some languages. The free model is excellent for major European languages plus Chinese and Japanese, but weaker on others. For example, Hindi on the free model can come back in the related Urdu script with phonetic drift — Hindi and Urdu are the same spoken language, and this is a known limitation of the smaller model. The word timing stays accurate; it is the text that suffers. Pro’s larger model handles Hindi and other harder languages much better.
| Language (clean speech) | Free in-browser model |
|---|---|
| English, Spanish | ~100% |
| Chinese, Japanese | ~100% content-accurate |
| French | ~88% |
| Hindi | Word timing accurate; text weaker on free — use Pro |
Word-synced captions were verified across English, Spanish, French, Hindi, Chinese and Japanese (including non-Latin scripts) — every word timed, 98–100% audio coverage. "99 languages" refers to the model’s language support; accuracy varies by language and audio quality, and free is ~95% / Pro ~98% on clean speech.