Audio to Text: Convert Audio to TXT Online — Free
Upload any audio file → get text, transcript, subtitles, or a summary in minutes. Honest accuracy by recording type, all export formats free from the first minute.
- 60 minutes free
- No credit card
- 99 languages
- Speaker labels
Last verified August 4, 2026
TL;DR — pick your path
| Your situation | Best path |
|---|---|
| Have an audio file, want text back | Upload workflow (DeluxeScribe, 60 min free) |
| Multi-speaker audio (interview, meeting, podcast) | Upload + speaker labels enabled |
Need subtitles too (.srt, .vtt) | Upload once, export both |
| File is MP3 | Format page: MP3 to Text |
| File is WAV | Format page: WAV to Text |
| File is M4A | Format page: M4A to Text |
| File is MP4 (video) | Format page: MP4 to Text |
| Sensitive content (medical, legal, unreleased) | Self-hosted Whisper — no upload |
| 40+ hours per month, cost-first | Heavy-volume alternative — see below |
| Broadcast-standard captions | Human captioner (Rev, Verbit) |
Convert audio to text in 3 steps
The section the SERP skips: you don’t re-upload for subtitles or a summary. One transcription produces text, timestamped subtitles, and an AI summary from the same file in the same editor.
- Sign upfor DeluxeScribe (60 minutes free, no credit card). No trial timer — the 60 minutes are yours to use across as many files as you want.
- Upload the audio file.Drag any MP3, WAV, M4A, OGG, OPUS, FLAC, AAC, or WebM. Files up to 5 GB accepted — enough for ~40 hours of typical voice audio.
- Wait 1–3 minutes for a typical file. Language auto-detects; speaker labels default on for multi-speaker audio. Then export as TXT, DOCX, PDF, SRT, VTT, or JSON.
“Convert audio to text,” “convert audio file to text,” “convert sound to text,” and “turn audio into text” are all the same operation. Google groups them into one intent; the same workflow serves all of them.
Transcribe audio — verb form
“Transcribe audio” and “convert audio to text” describe the same operation. Some searchers use the verb (“transcribe”), some use the phrase (“convert to text”); both land on the same SERP with the same tools. This section uses the verb form for readers who search that way.
How to transcribe audio for free
Three real options. DeluxeScribe’s 60-minute one-time free tier covers a typical interview or short podcast without a credit card. Self-hosted Whisper is free forever if you can run pip install openai-whisperand wait for processing. Apple Voice Memos on iPhone 12+ (iOS 18+) transcribes on-device for English. Most other “free audio transcription” web tools cap at 5–15 minutes per file or hide a paywall behind download.
Transcribe an audio file
Same workflow, any format. Upload the file, wait for transcription, export the text. DeluxeScribe accepts MP3, WAV, M4A, MP4, OGG, OPUS, FLAC, AAC, and WebM directly — no format conversion needed before upload. Files up to 5 GB per upload; longer files, split with FFmpeg first.
Transcribe an audio recording
Any recording that plays back as audio can be transcribed — voice memos, phone calls, interviews, lectures, podcasts, meeting recordings, dictation. Recording quality matters more than format: a clean recording at 128 kbps transcribes better than a noisy recording at 320 kbps. See the accuracy section below for realistic expectations by recording type.
Audio transcription — free & paid options
“Audio transcription” is the noun form of “transcribe audio.” Same operation, same tools. Below: honest comparison of the real free and paid options.
| Option | Cost | Volume ceiling | Best for |
|---|---|---|---|
| DeluxeScribe (us) | 60 min free → $10/mo | 1,200 min/mo Pro | Speaker labels + all export formats free |
| Self-hosted Whisper | Free forever (MIT) | Unlimited | Privacy + volume, if you can run CLI |
| TurboScribe | 90 min/day free → $10/mo unlimited | Unlimited on paid | 40+ hrs/mo raw transcription |
| Otter | 300 min/mo free | Paid tiers vary | Live meeting capture |
| HappyScribe | 10 min free/mo | Paid per-minute | European workflows |
| Rev (human tier) | $1.50/audio min | Unlimited | Legal / evidentiary accuracy |
AI audio transcriptionrefers specifically to the cloud + self-hosted options above (everything except Rev human tier). All AI options run Whisper large-v3 or a similar model — accuracy on the same source is roughly identical across services. What differs is monthly limits, export formats included on free, speaker diarization quality, and editor features.
Free audio transcriptionhonestly means one of: DeluxeScribe’s 60-minute credit, self-hosted Whisper, or Apple Voice Memos on-device. Beyond that, “free” usually means “free trial” or “free with watermark.” See Best Free Transcription Software for the machine-readable caps table.
What to look for in an audio-to-text converter
Quick disambiguation: an audio-to-text converter is not an audio format converter. The former uses speech recognition to produce text; the latter (CloudConvert, Zamzar, FreeConvert) changes MP3 to WAV or WAV to FLAC without producing any text. If you want to change your file’s format, you want a format converter. If you want the words as text, you want a transcription tool — this page.
Four criteria that actually matter when comparing audio-to-text converters:
- Accuracy on your audio.All top cloud options run Whisper large-v3 or a comparable model — accuracy is roughly identical on the same source. Vendor marketing (“99% accurate”) is typically unsourced. Test with a real file from your own workflow.
- Supported formats. Baseline: MP3, WAV, M4A, MP4. Advanced: OPUS (WhatsApp), FLAC (audiophile), WebM (browser recordings). Anything missing means an extra FFmpeg step before upload.
- Export formats on the free tier. Many services gate SRT / VTT / DOCX behind paid plans and give free users only TXT. DeluxeScribe unlocks all six (TXT, DOCX, PDF, SRT, VTT, JSON) from the first minute.
- Honest free-tier limits.“Free audio-to-text converter” often hides caps: watermark on video export, 5-minute per-file limit, or silent rate-limiting after 5 uploads. Read the free-tier terms before committing.
Realistic accuracy expectations
Every transcription product claims “99% accurate.” None of them are, in the general case. Accuracy depends on the recording conditions, not the vendor — the same Whisper large-v3 model hits 97% on a clean podcast and 75% on a phone call regardless of which wrapper (DeluxeScribe, TurboScribe, HappyScribe) serves it.
| Recording type | Realistic accuracy | Notes |
|---|---|---|
| Studio podcast (single mic, post-production) | 95–98% | Best case for ASR. |
| Zoom meeting (good headset) | 90–95% | Compression costs a few points. |
| Phone interview / VoIP | 80–90% | Bandlimited audio drops hard consonants. |
| Lecture hall from back row | 75–85% | Reverb and distance hurt more than noise. |
| Outdoor / on-the-street | 60–80% | Wind and traffic dominate. |
The honest takeaway: AI transcription is great as a first draftfor most content. Expect to spend 10–20% of the audio runtime reviewing and correcting — much less than typing from scratch, but not zero. Full WER-by-condition breakdown at How Accurate Is Whisper.
Which audio formats work?
DeluxeScribe accepts every common audio and video format directly — no pre-conversion needed:
| Format | Typical source | Format-specific page |
|---|---|---|
MP3 | Podcasts, downloads, generic recordings | MP3 to Text |
WAV | Uncompressed studio recordings | WAV to Text |
M4A | iPhone Voice Memos, AAC recordings | M4A to Text |
MP4 | Video files with audio | MP4 to Text |
OGG, OPUS | WhatsApp voice notes, some browsers | — upload directly |
FLAC | Lossless music, archival audio | — upload directly |
AAC, WebM | Browser recordings, some phone apps | — upload directly |
For unusual formats not listed above (AIFF, WMA, AU), convert to MP3 or WAV first with FFmpeg: ffmpeg -i input.xxx -acodec libmp3lame output.mp3.
Wrong format? Sibling pages
If your file is a specific format, the format-specific page has deeper coverage (bitrate reality for MP3, iOS Voice Memos workflow for M4A, uncompressed audio physics for WAV):
.mp3(compressed audio) → MP3 to Text.wav(uncompressed audio) → WAV to Text.m4a(iPhone Voice Memos, AAC container) → M4A to Text.mp4(video with audio) → MP4 to Text- Any video source(YouTube, Zoom recording, screen capture) → Video to Text
- iPhone voice memos specifically → iPhone Voice Memo Transcription
AI transcription — what it actually means for your audio
“AI transcription” means Whisper large-v3 (or a similar architecture) doing the speech recognition. Most modern cloud services — DeluxeScribe, TurboScribe, HappyScribe, ElevenLabs, and even most Otter transcripts — run OpenAI’s Whisper large-v3 or a derivative under the hood. On the same source audio, accuracy is nearly identical across these services. What differs is:
- Speaker diarization pipeline— the add-on that identifies who spoke each segment. Varies meaningfully between vendors.
- Editor and export— whether you get a browser editor for corrections, and which formats are unlocked on the free tier.
- Limits and price— monthly minutes, per-file size caps, and per-minute cost after free.
For a genuinely different ASR engine (not a Whisper wrapper), the real alternatives are Deepgram Nova-3, AssemblyAI Universal-2, Speechmatics, and Google Cloud Speech-to-Text. See our Whisper Alternatives page for the honest ranked comparison.
When DeluxeScribe is the right tool
The right answer is us when you have multiple audio files, need speaker labels included by default, want every export format (TXT, DOCX, PDF, SRT, VTT, JSON) unlocked from the first minute, work in a language other than English, or want a browser transcript editor plus AI summary in one flow. Full capability list on our features page, and background on the team at about. Our 60-minute free tier (no card) covers a typical interview series or podcast episode without paying anything.
When another tool fits better
Being honest about the sub-cases where DeluxeScribe isn’t the best fit:
- 40+ hours of audio per month, cost-first. TurboScribe’s Unlimited plan at $120/year is hard to beat on raw volume. Our Pro plan caps at 1,200 min/month — genuine tradeoff. See TurboScribe Alternatives.
- Sensitive content that can’t be uploaded. Self-hosted Whisper is the correct answer — same engine, on your own machine, no cloud policy required. Court audio, PHI without a BAA, unreleased music masters.
- Legally-critical accuracy. Rev human tier ($1.50/audio minute) or Verbit for education / legal / government. AI is good enough for most content; when a judge might read the transcript, humans still win.
- Live meeting workflows.Otter, Notta, or Fireflies. Different product category — a bot joins your Zoom / Meet / Teams call and transcribes in real time. If you’re just exporting the recording afterward, that’s a hint the meeting-bot flow would save you a step.
- Broadcast / OTT delivery with QC. BBC iPlayer, Netflix, HBO Max spec compliance requires human captioning. AI needs manual segmentation cleanup at minimum.
How this page was verified
Related guides
- MP3 to TextFormat-specific sibling — MP3 workflow with bitrate reality and file-check step.
- WAV to TextUncompressed audio — highest accuracy ceiling and the FFmpeg compression command.
- M4A to TextiPhone Voice Memos and other M4A files. Covers iOS 18's built-in transcription.
- MP4 to TextVideo files with audio — MP4, MOV, AVI. Audio track extracted automatically.
- How to Transcribe Audio (pillar)The pillar — every path (SaaS, free tools, self-hosted Whisper, native OS) and how to pick.
- Best Free Transcription SoftwareMachine-readable free-tier caps table — Fathom, Otter, TurboScribe, Whisper, and more.