Voice Dictation
Speak — or drop an audio file — and get accurate text with on-device Whisper. 90+ languages, edit, fix recurring words, export TXT/Markdown/SRT. Nothing uploaded.
Transcribe an audio file or your voice to text in 90+ languages, then export TXT, SRT or Markdown. Runs on your device — audio never leaves your browser.
About Voice Dictation
Voice Dictation is a free, browser-based speech-to-text tool that turns your microphone or an existing audio file into editable text using on-device Whisper (the multilingual whisper-base model running via Transformers.js / ONNX Runtime Web). Use it to dictate notes and drafts, or to transcribe an mp3/m4a/wav/etc. recording, then edit, fix recurring words, and export to TXT, Markdown, or timed SRT subtitles. The whole pipeline runs locally in your browser — your audio is never uploaded and there is no signup. It is desktop-focused, since the Whisper model is too heavy to run reliably on phones.
How to use Voice Dictation
- Pick your spoken language from the Language dropdown (Auto-detect can mis-default to English, so choosing accurately improves results).
- Click 'Start recording' and speak, then 'Stop & transcribe' — or use 'Hold to talk' to record only while you hold the button. On first use the ~145 MB Whisper model downloads once from a public CDN and is then cached.
- Alternatively, drop an audio file (mp3, m4a, wav, webm, ogg, aac, flac) onto the 'Or transcribe a recording' zone to transcribe it on your device.
- Edit the resulting text directly in the Transcript box; record again or drop another file to append more text.
- Open 'Corrections & replace' to add whole-word auto-corrections for recurring names/jargon (saved and applied to future transcriptions) or to do a one-off find & replace.
- Export with the TXT or Markdown button — and SRT (timed) when the transcript came from a recording or file, since subtitle timing comes from those cues.
Frequently asked questions
- Is my audio uploaded to a server?
- No. Speech-to-text runs entirely on your device with on-device Whisper, so your microphone audio and any audio file you drop are processed locally and never uploaded. The Whisper model weights are downloaded once from a public CDN (Hugging Face/jsdelivr) and cached in your browser — those CDNs see your IP from that download, but your audio stays on your device.
- Is it free and do I need an account?
- Yes, it's completely free with no signup or account. The only one-time cost is downloading the ~145 MB Whisper model the first time you use it, after which it's cached locally.
- Which audio file formats can I transcribe?
- You can drop common formats including mp3, m4a, wav, webm, ogg, aac, and flac. The file is decoded and transcribed entirely on your device.
- How do the auto-corrections work?
- In 'Corrections & replace' you add 'heard as → should be' pairs that are applied as whole-word fixes to new transcriptions — ideal for recurring names or jargon Whisper mishears. These are saved in your browser, and you can also apply them to the current text. The separate Find & Replace is a one-off literal substring replacement.
- Can I export subtitles, and why is SRT sometimes missing?
- You can export TXT and Markdown anytime there's text, and SRT (timed) when the transcript came from a recording or an audio file. SRT needs per-segment timing cues, which only exist for recorded/uploaded audio — text you type by hand has no timestamps, so the SRT button only appears when cues are present.
- Does it work on my phone, and how is it different from Live Transcription?
- This tool is built for desktop browsers — the Whisper model is too heavy for phones and can crash the tab. It favors accuracy and editing over speed; if you want captions to appear as you speak (and on mobile), use the separate Live Transcription tool, which uses a lighter model.
People also search for
Voice Dictation is also known as voice to text dictation software, convert speech to text, voice typing online, transcribe audio file to text, voice dictation app, audio to text converter, dictation to word.