Skip to main content

Comparisons

Browser dictation vs. Whisper upload

Free browser dictation compared with a Whisper upload transcript, and when to use each.

By GistCite team, Product and engineering · For podcasters, journalists, developers

Published · Reviewed

Key takeaways

  • Browser dictation uses the device's own speech engine, is free, and nothing leaves the browser.
  • Whisper upload sends a recording to a server model and returns word-level timestamps and translation.
  • Dictation works only for live speech in real time; upload works on an existing recording up to 500 MB.
  • Upload is priced per started minute; dictation costs nothing and needs no account.

Live dictation

Browser dictation uses the speech recognition already built into your browser. Speak, and the text appears as you go. It costs nothing, needs no account, and nothing you say is sent to a server.

It works well for drafting quickly, capturing a thought before it is lost, or filling in a form field by voice instead of typing.

Whisper upload

Uploading a recording, up to 500 MB, sends it to Whisper for a transcript with word-level timestamps. It can translate the result to English and take a vocabulary hint for names and jargon a general model would otherwise misspell. Results are cached by content hash, so re-running the same file costs one unit the second time.

The vocabulary hint is especially useful for an interview or lecture with proper nouns, technical terms or brand names that a general-purpose model has never encountered.

Real time versus after the fact

Dictation only works on speech happening now, into a microphone that is listening. Upload works on anything already recorded: a meeting, an interview, a lecture or a voice memo, regardless of when it happened.

A recording made on a phone, pulled from a video call, or captured months ago all work the same way through upload, since none of them depend on a live microphone connection.

Cost and accuracy

Dictation is free; upload is priced per started minute, with a small minimum, and a higher-quality mode at a higher rate per minute. Whisper handles clear speech well and degrades with background noise and crosstalk, which is where the vocabulary hint helps most.

Because pricing is per minute rather than per file, a short voice memo and a long lecture are charged proportionally to their length, not a flat rate regardless of size.

References

The tools this article uses

Speech to Text

Dictate live in the browser for free, or upload a recording up to 500 MB for a Whisper transcript with word timestamps and SRT export.

About Speech to Text

Put it into practice

The step this article describes is one action in the toolkit.