Free to start
Speech to Text
Turn a meeting, interview, lecture or voice memo into an editable transcript with timestamps and subtitles.
Published · Reviewed
Who it is for
- Journalists and researchers
- Podcasters
- Students
- Anyone who records meetings
The problem
Recordings pile up unlistened. Typing them out takes four times their length, and consumer dictation stops at the first pause.
How GistCite solves it
Free in-browser dictation uses your browser's speech engine with nothing sent to a server. Upload mode sends the recording to Whisper for a transcript with word timestamps, optional translation to English and a vocabulary hint for names and jargon; results are cached by content hash and downloadable as TXT or SRT.
Capabilities
Live dictation
Speak and watch the text appear; runs in the browser, free.
Benefit: Draft by voice with nothing uploaded.
Dictate in the browserWhisper upload
Recordings up to 500 MB, word timestamps, translate-to-English and a names-and-jargon hint.
Benefit: A transcript you can search and quote.
Transcribe a recordingSRT and TXT export
Subtitles for video or plain text for notes.
Benefit: Caption an interview without re-timing it.
Download subtitles from a recordingBenefits
- Priced per started minute with a small minimum, and shown before the upload runs.
- Cached by content hash: the same recording twice costs one unit the second time.
- The transcript joins your history, so it can be summarized, questioned and turned into drafts.
How it works
Step 1
Record or upload
Dictate live, record a browser tab, or upload audio or video.
Step 2
Transcribe
Whisper returns text with timestamps; high-quality mode costs more per minute.
Step 3
Edit and export
Fix names, then copy or download TXT or SRT.
Integrations and formats
- OpenAI Whisper
- SRT / TXT export
- AI Summary, Ask and Create
- REST API
Security and data handling
Read the full picture on the security and compliance page.
- Live dictation never leaves the browser.
- Uploads are processed in memory and never written to disk or object storage; only the extracted text is kept.
- Every extraction endpoint is rate-limited per IP; signed-in use is metered against the plan's monthly quota.
- TLS in transit; passwords bcrypt-hashed; API keys stored only as SHA-256 digests.
Evidence
- Upload transcription runs on Whisper; the engine is named on the pricing and guide pages. Check this
Worked examples
Independent podcaster
From a one-hour podcast episode to a newsletter, captions and three clips
One public episode on YouTube becomes a transcript, a chaptered summary, a newsletter draft, SRT captions and three short clips, all from one upload.
Read the independent podcaster walkthroughPricing
- Plans
- Free (dictation), Free and Pro (uploads)
- Cost per action
- 1 unit per started minute, minimum 3; 2 units per minute at high quality; cached repeats 1 unit.
- How it is priced
- Live dictation costs nothing and needs no account.
Speech to Text questions
- Is dictation really free?
- Yes. It uses the browser's own speech recognition and sends nothing to GistCite.
- What formats can I upload?
- Common audio and video files up to 500 MB.
- How accurate is it?
- Whisper handles clear speech well and degrades with noise and crosstalk; the vocabulary hint improves names and jargon.
Try Speech to Text now
No account is needed to try it; an account adds history, API keys and a monthly quota.
Prefer to talk first? Contact the team · Read the API documentation
