Comparisons
Browser dictation vs. Whisper upload
Free browser dictation compared with a Whisper upload transcript, and when to use each.
By GistCite team, Product and engineering · For podcasters, journalists, developers
Published · Reviewed
Key takeaways
- Browser dictation uses the device's own speech engine, is free, and nothing leaves the browser.
- Whisper upload sends a recording to a server model and returns word-level timestamps and translation.
- Dictation works only for live speech in real time; upload works on an existing recording up to 500 MB.
- Upload is priced per started minute; dictation costs nothing and needs no account.
Live dictation
Browser dictation uses the speech recognition already built into your browser. Speak, and the text appears as you go. It costs nothing, needs no account, and nothing you say is sent to a server.
It works well for drafting quickly, capturing a thought before it is lost, or filling in a form field by voice instead of typing.
Whisper upload
Uploading a recording, up to 500 MB, sends it to Whisper for a transcript with word-level timestamps. It can translate the result to English and take a vocabulary hint for names and jargon a general model would otherwise misspell. Results are cached by content hash, so re-running the same file costs one unit the second time.
The vocabulary hint is especially useful for an interview or lecture with proper nouns, technical terms or brand names that a general-purpose model has never encountered.
Real time versus after the fact
Dictation only works on speech happening now, into a microphone that is listening. Upload works on anything already recorded: a meeting, an interview, a lecture or a voice memo, regardless of when it happened.
A recording made on a phone, pulled from a video call, or captured months ago all work the same way through upload, since none of them depend on a live microphone connection.
Cost and accuracy
Dictation is free; upload is priced per started minute, with a small minimum, and a higher-quality mode at a higher rate per minute. Whisper handles clear speech well and degrades with background noise and crosstalk, which is where the vocabulary hint helps most.
Because pricing is per minute rather than per file, a short voice memo and a long lecture are charged proportionally to their length, not a flat rate regardless of size.
