Transcribe Audio: 4 Real Paths (2026)

Upload to a cloud service, use a free tier, run Whisper locally, or use your OS built-in — the honest decision framework with real accuracy figures.

Four paths exist:upload to a cloud service (easiest, paid above a free tier), self-host Whisper (free, technical), use a free tier (has limits), or use your operating system’s built-in transcription (macOS Voice Memos, Microsoft Word, Pixel Recorder). Pick based on the file you have, the accuracy you need, and whether you can upload the audio. DeluxeScribeis one of the cloud options — 99 languages, files up to 5 GB, speaker labels, six export formats. Free tier is 60 minutes; paid plans start at $10/month. Below: the decision framework, honest accuracy by audio condition (not the usual “99%” blanket claim), pricing across the field, and format-specific routing for your exact case.
  • 60 minutes free
  • No credit card
  • 99 languages
  • Speaker labels

Last verified July 8, 2026

TL;DR — pick your path

Match your situation to a path. Each is explained in detail below.

Your situationBest path
Any audio file, want it done fastCloud upload — DeluxeScribe (free 60 min)
Sensitive content — can’t upload to cloudSelf-hosted Whisper (local, free, same accuracy)
Long recording (over 3 hours) or ongoing volumePaid cloud tier or self-hosted Whisper on GPU
iPhone (iOS 18+) with a short English clipVoice Memos built-in transcription (on-device)
You have a Microsoft 365 subscriptionWord Transcribe (free with subscription)
Format-specific need (MP3, WAV, M4A, MP4)Jump to the format page (linked below)
Building transcription into your productAssemblyAI, Deepgram, or self-hosted Whisper API

How to transcribe audio (3 steps)

  1. Upload — drop your file into DeluxeScribe or any cloud transcription service. MP3, M4A, WAV, MP4, and 20+ other formats work directly.
  2. Select language — pick the spoken language or leave on auto-detect (runs on the first ~30 seconds).
  3. Export — download as TXT, DOCX, PDF, SRT, VTT, or JSON with speaker labels and timestamps.

That’s the common case. If your situation is different (sensitive content, ongoing volume, iOS built-in), the sections below cover it.

Transcribe any audio file in 99 languages

60 minutes free, no credit card. Speaker labels, six export formats, files up to 5 GB.

The 4 paths, compared

PathCostSpeedPrivacyMax lengthLanguagesSpeaker labels
Cloud service (DeluxeScribe et al.)Free tier → $10+/mo5–10 min per hourCloud-encrypted5 GB / ~90 hrs99Automatic
Self-hosted WhisperFreeCPU: 10–30× realtime, GPU: near realtimeLocal (never uploads)Your disk99Add-on (WhisperX)
Free web toolsFree (with caps)5–15 min per hourCloud (varies)Usually 30 min10–40Sometimes
Native OS (iOS/macOS/Word/Pixel)Free / includedRealtime (background)On-device (mostly)Varies (iOS: ~30 min)1–12No

Path 1 — Cloud transcription service

Upload audio in a browser, get a transcript back in minutes. The easiest path for most use cases. Free tiers exist; paid plans start around $10/month.

When cloud wins: multiple files, non-English audio, need speaker labels without setup, need in-browser editing to fix errors, want SRT/VTT/DOCX/JSON exports, need speed.

When cloud loses: sensitive content (medical, legal, NDA), offline requirement, cost sensitivity at high volume.

Path 2 — Self-hosted Whisper

OpenAI’s Whisper model runs locally with one command if you have Python. Audio never leaves your machine.

pip install openai-whisper
whisper recording.mp3 --model large-v3 --output_format srt

Model size tradeoff: tiny and base are fast but rough; medium is a good default; large-v3 is state-of-the-art but ~10× slower on CPU.

Speed reality:on a typical CPU, expect 10–30× real-time (a 1-hour file takes 10–30 hours). On a recent GPU it’s real-time or faster. For speaker labels, pair with WhisperX (adds word-level timestamps + diarization) or faster-whisper for a 4× speedup.

When Whisper wins: medical, legal, or confidential audio; high volume where cloud cost adds up; you already have a GPU.

Path 3 — Free web tools (with caveats)

Sites like audiototext.com, freepodcasttranscription.com, audioconvert.ai offer browser-based transcription with no signup. Convenient for a one-off short file.

The catch:“free” almost always hides a paywall. Common patterns: file-length cap (10–30 minutes), transcript truncation after the free portion, watermarked output, paywall appearing after processing.

Honest DeluxeScribe note:our free tier gives you 60 minutes upfront with no card required — no after-the-fact paywall, no watermark, all export formats included. It’s the fairest free tier we’ve found; we’re not saying that’s neutral, we’re saying that’s why the “free web tools” category is often worse than a legitimate free tier.

Path 4 — Native OS options

Your operating system probably has built-in transcription. All free, all language-limited relative to cloud services.

  • iPhone Voice Memos (iOS 18+) — on-device transcription, English + growing language list, iPhone 12+. See our iPhone Voice Memo guide.
  • macOS Voice Memos — same engine as iOS, syncs via iCloud, Apple Silicon Macs.
  • Microsoft Word Transcribe — included with Microsoft 365, web-only, supports MP3/WAV/M4A/MP4, ~10 languages. Records live or accepts file upload.
  • Google Docs voice typing — live only (no file upload), 100+ languages, free with Google account.
  • Pixel Recorder (Android) — on-device on Pixel 3+, English + a handful of other languages.

Accuracy reality — WER by audio condition

Every vendor claims “99% accuracy.” That number is real — on the easiest possible condition. On everything else, accuracy drops in predictable ways.

Word Error Rate (WER) is the standard metric: the percentage of words wrong (substitutions + insertions + deletions). 5% WER means 5 out of 100 words are wrong. Under 15% is workable with light editing.

Audio conditionTypical WERReads like
Studio mic, one speaker, clear English2–5%Almost perfect; light proofreading
Conference room, 2–3 speakers, good mics5–10%Useful as-is; spot-check names
Zoom call, mixed mics, 4+ speakers10–20%Readable, needs editing for publication
Phone call (8 kHz narrow-band)15–30%Gist clear, details unreliable
Heavy accent + background noise20–40%Re-listen to verify key points
Music + speech overlapOften failsModel may hallucinate lyrics or skip sections

For the full breakdown by language and Whisper model size, see How Accurate Is Whisper.

Where all providers fail the same way

  • Proper nouns — names, businesses, product names
  • Phone numbers — fast or non-standard groupings
  • Technical jargon — drug names, legal terms, science
  • Homophones — their/there/they’re, principal/principle
  • Numbers with units — “15 mg” vs “50 mg”

Spot-check these regardless of vendor.

Pricing reality across the field

What you actually pay per hour of transcribed audio, across the main cloud services. Verified July 8, 2026.

ServiceFree tierPaid fromApprox. $/hr audioNotes
DeluxeScribe60 min one-time$10/mo · 1,200 min~$0.50/hr99 languages, speaker labels, 6 export formats
TurboScribe Unlimited3 files/day$10/moEffectively $0 at volumeUnlimited hours; single-tier pricing
Otter300 min/mo$17/mo · 1,200 min~$0.85/hrMeeting bot, calendar integration
Descript1 hour/mo$24/mo · 30 hrs~$0.80/hrEdit audio by editing transcript text
Rev AI (self-serve)None$0.25/min$15/hrPay-per-minute; optional human tier at $1.50/min
TrintNone$48/mo · 7 hrs~$6.85/hrNewsroom editor for journalists
Self-hosted WhisperFree$0 (electricity)Free forever; needs hardware + Python

For a full comparison of TurboScribe alternatives, see our TurboScribe alternatives rankings. For Whisper-specific alternatives (developer/API focus), see Whisper alternatives.

File formats — what works and what needs converting

Most modern services accept every common audio and video format. The exceptions are rare and easy to convert.

Upload directly (works everywhere)

FormatCommon sourceFormat-specific guide
MP3Podcasts, music apps, universalMP3 to Text
M4AiPhone Voice Memos, WhatsApp voice notesM4A to Text
WAVUncompressed studio, pro audio appsWAV to Text
MP4Zoom exports, screen recordings, phone videoMP4 to Text
MOV, WebM, FLAC, AAC, OGG, OpusQuickTime, browser recordings, audiophileUpload directly to DeluxeScribe

Convert first

SourceConvert toFFmpeg command
MKV videoMP4 (re-mux, no re-encode)ffmpeg -i input.mkv -c copy output.mp4
AMR (old Android recorder)M4Affmpeg -i input.amr -c:a aac output.m4a
WMA (old Windows)MP3ffmpeg -i input.wma output.mp3
ALAC (Apple Lossless in .m4a)AAC in M4Affmpeg -i input.m4a -c:a aac output.m4a

By source — jump to your exact use case

This page is the overview. For specific formats, sources, or workflows, we have dedicated guides with exact steps and gotchas.

By format

By recording source

Podcasts + long-form

Video, subtitles & captions

Alternatives + comparisons

Privacy — one honest paragraph

Cloud services:your audio is uploaded and processed on the vendor’s servers. DeluxeScribe encrypts in transit (TLS) and at rest, doesn’t use your audio to train models, and lets you delete recordings anytime. Self-hosted Whisper: nothing leaves your machine. Native OS (iPhone Voice Memos, Pixel Recorder): processes on-device.

We are not HIPAA-compliant — do not upload Protected Health Information. For PHI, use a vendor with a signed BAA, or self-host Whisper. For attorney-client privileged content, check with your firm before uploading to any cloud service.

How this page was verified

Word Error Rate (WER) ranges come from published benchmarks on the LibriSpeech corpus (clean studio audio) and the AMI Meeting Corpus (multi-speaker meeting audio), generalized across modern transformer-based ASR models. Whisper accuracy ranges align with Radford et al. (2022). Whisper command-line behavior from the OpenAI Whisper repository. Pricing figures verified against vendor pricing pages on July 8, 2026. We don’t cite “99% accurate” marketing claims that appear across competitor copy — they’re not sourced to a published study.

Frequently Asked Questions

How accurate is AI audio transcription?

On clean studio audio with one English speaker, modern services hit 95–98% word accuracy (2–5% WER). On a Zoom call with 4 speakers and mixed mics, expect 80–90%. On phone call audio (narrow-band codec), 70–85%. On audio with music or heavy background noise, often much worse. The '99% accurate' marketing claim is real only on the easiest condition — see the WER-by-condition table below.

Is there a truly free option?

Yes. iPhone Voice Memos on iOS 18+ transcribes on-device for free (English, plus a growing list of languages). Pixel Recorder is free on Pixel phones. Microsoft 365 subscribers get Word's Transcribe feature included. Self-hosted Whisper is free forever (requires Python and decent hardware). Cloud services offer free tiers — DeluxeScribe gives 60 minutes free with no credit card. "Free web tools" typically cap at 5–10 minutes before paywalling.

Can I transcribe an audio file over 3 hours long?

Yes. DeluxeScribe accepts files up to 5 GB — roughly 90 hours of audio at typical bitrates. Most cloud services handle multi-hour files without splitting. Self-hosted Whisper has no length limit (limited only by your disk and time). The bottleneck is usually upload speed on cloud services, not processing.

What file formats work?

Direct upload: MP3, M4A, WAV, MP4, MOV, FLAC, AAC, OGG, Opus, WebM, and 15+ others. May need conversion first: MKV (re-mux to MP4), AMR (convert to M4A), some proprietary formats. The original file format barely affects accuracy — audio quality at recording time matters more than the container.

Can I get speaker labels?

Yes, on cloud services and with Whisper + a diarization add-on (WhisperX or whisper-diarization). DeluxeScribe includes automatic speaker labels. Expect 85–95% attribution accuracy on clean recordings, dropping to 70–85% on Zoom-style mixed-mic setups. iOS 18's Voice Memos transcription does NOT include speaker labels.

How long does transcription take?

Cloud services typically finish a 1-hour file in 3–10 minutes. Self-hosted Whisper on a CPU takes 10–30× real-time (a 1-hour file takes 10–30 hours). On a recent GPU, near real-time. Native OS transcription (iPhone, Pixel) processes in the background at roughly real-time speed.

Does audio quality matter more than bitrate?

Yes, dramatically. A 64 kbps recording from a close microphone in a quiet room outperforms a 320 kbps recording from a phone across a noisy room. Mic distance, room acoustics, and background noise dominate accuracy. Bitrate matters only below ~32 kbps where compression artifacts start hurting speech recognition.

Is my audio kept private?

Depends on the path. Cloud services: audio is uploaded, processed, and stored per the vendor's retention policy. DeluxeScribe encrypts in transit and at rest, doesn't train on customer audio, and lets you delete anytime. Self-hosted Whisper: nothing leaves your machine. Native OS on iPhone/Pixel: processes on-device. For medical, legal, or NDA-covered content, self-host — no cloud service (including ours) is HIPAA-compliant.

Can I transcribe non-English audio?

Yes. DeluxeScribe supports 99 languages. Whisper (self-hosted) supports 99+. Most cloud services support 30–60 languages. Native OS options are more limited: iOS 18 launched with English and added more in point releases; Word Transcribe supports ~10 languages; Pixel Recorder supports a handful.