Journalists, researchers, and podcasters: turn interview recordings into accurate, quotable text — free, unlimited, and completely private.
Files are processed locally and never uploaded.
Transfer your interview recording (MP3, WAV, or M4A) to your computer and drop it into the uploader above.
Set the language explicitly and pick the small tier for maximum accuracy on important interviews — base is fine for clear audio.
Copy the transcript into your editor, verify every quote against the audio, and download the TXT for your records.
An interview transcript isn't notes — it's evidence. Journalists quote from it, researchers code it, podcasters cut from it. That means accuracy requirements are higher than for a casual voice memo: a misheard name becomes a printed error, a dropped "not" reverses a meaning. Free, unlimited, private transcription removes the cost barrier that used to force people to transcribe only the "important" interviews.
Because there's no per-minute fee and no quota, the economics flip: transcribe everything, decide what's quotable later. The interview you almost didn't transcribe is invariably the one with the perfect quote buried at minute forty.
Transcription quality is decided at recording time. Put the recorder close to the subject — a phone on the table between you beats a phone in your pocket by a mile — and prefer quiet rooms over cafés. If you're doing this regularly, a cheap lapel mic is the best money you'll ever spend on accuracy; even wired earbuds as a mic outperform a distant phone.
Do a ten-second test at the start of every important interview: record, play back, listen for clarity. And state names and spellings on the recording itself ("Could you spell your surname for the transcript?") — future-you, hunting for the correct spelling of a source's name at midnight, will be grateful.
For interviews that will be quoted, use the small tier if your machine handles it — it's meaningfully better on the hard cases interviews produce: accented English, overlapping speech, domain jargon, emotional speech that speeds up and trails off. For clear, close-mic'd conversations, base is honestly fine and much faster.
Set the language explicitly rather than trusting auto-detect; interviews in a known language should be tagged as such. And transcribe the full recording rather than excerpts — context helps Whisper resolve ambiguous words, and you'll want the surrounding material when you verify quotes anyway.
Treat the transcript as a 95%-accurate draft, because that's what it is. The editing pass has three jobs: verify every quote you'll publish against the original audio (proper nouns especially — Whisper guesses phonetically and guesses wrong), add speaker labels yourself since the output has none, and clean up the verbal tics that read terribly in print but sound natural in speech.
Build the habit of keeping audio and text side by side during the edit. The transcript tells you where to look; the audio tells you what's true. Never publish a quote you haven't re-heard — this is journalism's oldest rule, and AI transcription doesn't retire it.
Recording someone creates obligations. Get consent before you record — in many jurisdictions it's legally required, and everywhere it's ethically required. Tell sources how the recording will be used, and honor off-the-record requests completely, not performatively.
Local transcription helps on the privacy side: your source's voice and words are processed on your device, never uploaded to a transcription company's servers. For sensitive interviews — whistleblowers, vulnerable populations, embargoed material — that architectural privacy isn't a nice-to-have; it's part of protecting your sources. Store recordings and transcripts securely, and know your outlet's or institution's data-handling rules.
This is Whisper's known weak spot — it renders unfamiliar names phonetically and sometimes invents plausible-sounding ones. Always verify names, places, and technical terms against the audio before quoting. Asking sources to spell their names on the recording is the best prevention.
No — there's no speaker diarization, so the transcript is one continuous text. For two-person interviews, adding speaker labels during your editing pass is quick; for multi-person panels, verify against the audio as you label.
Yes — Whisper supports 90+ languages. Set the language explicitly instead of auto-detect for the most reliable results, especially for shorter clips where auto-detection is less certain.
Really. Everything runs on your device, so there are no per-minute fees, no monthly quotas, and no accounts. Transcribing a three-hour interview costs exactly as much as a three-minute one: nothing but your computer's time.