Skip to main content

Free to start

Speech to Text

Turn a meeting, interview, lecture or voice memo into an editable transcript with timestamps and subtitles.

Published · Reviewed

Who it is for

  • Journalists and researchers
  • Podcasters
  • Students
  • Anyone who records meetings

The problem

Recordings pile up unlistened. Typing them out takes four times their length, and consumer dictation stops at the first pause.

How GistCite solves it

Free in-browser dictation uses your browser's speech engine with nothing sent to a server. Upload mode sends the recording to Whisper for a transcript with word timestamps, optional translation to English and a vocabulary hint for names and jargon; results are cached by content hash and downloadable as TXT or SRT.

Capabilities

Live dictation

Speak and watch the text appear; runs in the browser, free.

Benefit: Draft by voice with nothing uploaded.

Dictate in the browser

Whisper upload

Recordings up to 500 MB, word timestamps, translate-to-English and a names-and-jargon hint.

Benefit: A transcript you can search and quote.

Transcribe a recording

Benefits

  • Priced per started minute with a small minimum, and shown before the upload runs.
  • Cached by content hash: the same recording twice costs one unit the second time.
  • The transcript joins your history, so it can be summarized, questioned and turned into drafts.

How it works

  1. Step 1

    Record or upload

    Dictate live, record a browser tab, or upload audio or video.

  2. Step 2

    Transcribe

    Whisper returns text with timestamps; high-quality mode costs more per minute.

  3. Step 3

    Edit and export

    Fix names, then copy or download TXT or SRT.

Integrations and formats

  • OpenAI Whisper
  • SRT / TXT export
  • AI Summary, Ask and Create
  • REST API

Security and data handling

Read the full picture on the security and compliance page.

  • Live dictation never leaves the browser.
  • Uploads are processed in memory and never written to disk or object storage; only the extracted text is kept.
  • Every extraction endpoint is rate-limited per IP; signed-in use is metered against the plan's monthly quota.
  • TLS in transit; passwords bcrypt-hashed; API keys stored only as SHA-256 digests.

Evidence

  • Upload transcription runs on Whisper; the engine is named on the pricing and guide pages. Check this

Worked examples

Independent podcaster

From a one-hour podcast episode to a newsletter, captions and three clips

One public episode on YouTube becomes a transcript, a chaptered summary, a newsletter draft, SRT captions and three short clips, all from one upload.

Read the independent podcaster walkthrough

Pricing

Plans
Free (dictation), Free and Pro (uploads)
Cost per action
1 unit per started minute, minimum 3; 2 units per minute at high quality; cached repeats 1 unit.
How it is priced
Live dictation costs nothing and needs no account.

Compare the Free, Pro and API plans

Speech to Text questions

Is dictation really free?
Yes. It uses the browser's own speech recognition and sends nothing to GistCite.
What formats can I upload?
Common audio and video files up to 500 MB.
How accurate is it?
Whisper handles clear speech well and degrades with noise and crosstalk; the vocabulary hint improves names and jargon.

Try Speech to Text now

No account is needed to try it; an account adds history, API keys and a monthly quota.

Prefer to talk first? Contact the team · Read the API documentation