Drop in any MP3 — podcasts, lectures, voice recordings — and get a text transcript in your browser. Nothing is uploaded, ever.
Files are processed locally and never uploaded.
Drop your MP3 into the uploader above — the Whisper model loads once, then stays cached in your browser.
Pick a tier: tiny for quick drafts, base for the best balance, small for maximum accuracy on difficult audio.
Copy the transcript or download it as a TXT file. Your audio never left your device.
Nearly every recording in the world ends up as an MP3: podcast episodes, lecture captures, phone call recordings, voice memos exported from apps, downloaded audio from the web. It's the format everything agrees on — which is why a transcription tool that handles MP3 natively covers the vast majority of real-world jobs with zero conversion friction.
This tool accepts MP3 directly — no converting to WAV first, no renaming tricks. The browser decodes the file locally and feeds it straight to Whisper. If your recording lives in another supported format (WAV, M4A), that works too, but MP3 is the path of least resistance and the format this page is optimized around.
Three Whisper sizes are available, and the choice depends on your audio, not your ambition. Tiny is the sprinter: fast to load, fast to transcribe, and perfectly good for clear single-speaker recordings — quick voice notes, clean dictation, well-recorded narration. Base is the balanced default most people should start with: noticeably better on accents, background noise, and imperfect recordings.
Small is the precision instrument for difficult audio — overlapping voices, heavy accents, technical jargon, distant microphones. It downloads a larger model and wants a capable device; older phones will chug. A practical rule: start with base, and only re-run the tricky files on small. Your choice is remembered, so the tool adapts to your hardware once.
On clear, single-speaker MP3s — a podcast with a decent mic, a lecture recorded near the speaker — Whisper's accuracy is excellent, comparable to commercial services, because it is the same model family they use. Proper nouns, brand names, and unusual terms are the weak spot: Whisper guesses phonetically, so verify names against the audio before quoting them anywhere public.
Difficult audio degrades gracefully but visibly: background music beds confuse it, crosstalk merges speakers into one voice, and heavy compression artifacts (low-bitrate MP3s, speakerphone recordings) cost accuracy. The single biggest lever is the recording itself — a clean source on the tiny tier beats a noisy source on small every time. Set the language explicitly instead of auto-detect when you know it; auto-detection is convenient but less reliable.
The output is plain running text — a faithful rendering of what was said, which you can copy to the clipboard or download as a TXT file. There are no timestamps and no speaker labels; this is a deliberate simplicity, and it shapes the workflow: the transcript is a draft for notes, quotes, show notes, and searchable archives, not a finished subtitle file.
For most MP3 jobs that's exactly right. Podcasters paste it into their notes app to write show notes; students turn lecture audio into study documents; anyone with a folder of old voice recordings finally gets a searchable archive. Edit lightly, verify quotes against the audio, and you're done — no account, no per-minute fee, no waiting room.
People transcribe MP3s that are often sensitive: unreleased podcast episodes, recorded calls, personal voice journals, client interviews. The standard workflow — uploading that audio to a stranger's server — is a privacy compromise most people accept only because they don't know there's an alternative. Here the alternative is the whole product: the file is decoded and transcribed by your own device's CPU.
You can verify the claim yourself: open your browser's network inspector, drop in an MP3 after the model has loaded once, and watch — no audio data leaves. For pre-release content, confidential recordings, and anything you'd rather not have sitting in a third party's logs, that architectural difference is the entire point.
No hard limit, but be realistic: the whole audio is decoded into memory and processed on your device, so a multi-hour MP3 will take a while and needs a decent machine. For very long files, split them into 30–60 minute parts for smoother processing.
Somewhat. Standard 128kbps+ MP3s transcribe fine; heavily compressed low-bitrate files or speakerphone recordings lose accuracy. Don't re-compress an already-compressed MP3 before transcribing — you can't add quality back, and each re-encode throws away a little more detail.
Not currently — the tool outputs plain text (copy or TXT download), with no timestamps. It works well for notes, quotes, and searchable archives; timestamped subtitles are on the roadmap.
Yes, after the first use. The Whisper model files download once and are cached in your browser; after that, you can disconnect from the internet and transcription keeps working entirely on-device.