Transcribe Interviews to Text — Free

Journalists, researchers, and podcasters: turn interview recordings into accurate, quotable text — free, unlimited, and completely private.

Drop Audio File Here

Supports MP3, WAV, M4A

OR
Model

"Balanced by default. Fast saves data, Accurate is best on desktop."

Language

Transcription will appear here...

Files are processed locally and never uploaded.

How to use

  1. 1

    Transfer your interview recording (MP3, WAV, or M4A) to your computer and drop it into the uploader above.

  2. 2

    Set the language explicitly and pick the small tier for maximum accuracy on important interviews — base is fine for clear audio.

  3. 3

    Copy the transcript into your editor, verify every quote against the audio, and download the TXT for your records.

Why interviews demand better transcription

An interview transcript isn't notes — it's evidence. Journalists quote from it, researchers code it, podcasters cut from it. That means accuracy requirements are higher than for a casual voice memo: a misheard name becomes a printed error, a dropped "not" reverses a meaning. Free, unlimited, private transcription removes the cost barrier that used to force people to transcribe only the "important" interviews.

Because there's no per-minute fee and no quota, the economics flip: transcribe everything, decide what's quotable later. The interview you almost didn't transcribe is invariably the one with the perfect quote buried at minute forty.

Recording for transcription, not just for listening

Transcription quality is decided at recording time. Put the recorder close to the subject — a phone on the table between you beats a phone in your pocket by a mile — and prefer quiet rooms over cafés. If you're doing this regularly, a cheap lapel mic is the best money you'll ever spend on accuracy; even wired earbuds as a mic outperform a distant phone.

Do a ten-second test at the start of every important interview: record, play back, listen for clarity. And state names and spellings on the recording itself ("Could you spell your surname for the transcript?") — future-you, hunting for the correct spelling of a source's name at midnight, will be grateful.

Choosing accuracy settings that matter

For interviews that will be quoted, use the small tier if your machine handles it — it's meaningfully better on the hard cases interviews produce: accented English, overlapping speech, domain jargon, emotional speech that speeds up and trails off. For clear, close-mic'd conversations, base is honestly fine and much faster.

Set the language explicitly rather than trusting auto-detect; interviews in a known language should be tagged as such. And transcribe the full recording rather than excerpts — context helps Whisper resolve ambiguous words, and you'll want the surrounding material when you verify quotes anyway.

The editing pass: from transcript to quotable text

Treat the transcript as a 95%-accurate draft, because that's what it is. The editing pass has three jobs: verify every quote you'll publish against the original audio (proper nouns especially — Whisper guesses phonetically and guesses wrong), add speaker labels yourself since the output has none, and clean up the verbal tics that read terribly in print but sound natural in speech.

Build the habit of keeping audio and text side by side during the edit. The transcript tells you where to look; the audio tells you what's true. Never publish a quote you haven't re-heard — this is journalism's oldest rule, and AI transcription doesn't retire it.

Frequently asked questions

How accurate is it on names and proper nouns?▼

This is Whisper's known weak spot — it renders unfamiliar names phonetically and sometimes invents plausible-sounding ones. Always verify names, places, and technical terms against the audio before quoting. Asking sources to spell their names on the recording is the best prevention.

Can it identify different speakers in the interview?▼

No — there's no speaker diarization, so the transcript is one continuous text. For two-person interviews, adding speaker labels during your editing pass is quick; for multi-person panels, verify against the audio as you label.

Does it handle non-English interviews?▼

Yes — Whisper supports 90+ languages. Set the language explicitly instead of auto-detect for the most reliable results, especially for shorter clips where auto-detection is less certain.

Is there really no cost or limit?▼

Really. Everything runs on your device, so there are no per-minute fees, no monthly quotas, and no accounts. Transcribing a three-hour interview costs exactly as much as a three-minute one: nothing but your computer's time.

Related tools