Audio to Text: Convert Audio to TXT Online — Free

Upload any audio file → get text, transcript, subtitles, or a summary in minutes. Honest accuracy by recording type, all export formats free from the first minute.

DeluxeScribe converts audio to text (or audio to TXT) using Whisper large-v3 — the same AI model that powers most professional transcription services. Upload any MP3, WAV, M4A, OGG, or FLAC file. Get text back in 1–3 minutes with 92–98% accuracy on clean audio, speaker labels, timestamps, and export to TXT, DOCX, PDF, SRT, VTT, or JSON. First 60 minutes free — no credit card. Below: the honest workflow, realistic accuracy by recording type, and format-specific pages if your file is MP3 / WAV / M4A / MP4.
  • 60 minutes free
  • No credit card
  • 99 languages
  • Speaker labels

Last verified August 4, 2026

TL;DR — pick your path

Your situationBest path
Have an audio file, want text backUpload workflow (DeluxeScribe, 60 min free)
Multi-speaker audio (interview, meeting, podcast)Upload + speaker labels enabled
Need subtitles too (.srt, .vtt)Upload once, export both
File is MP3Format page: MP3 to Text
File is WAVFormat page: WAV to Text
File is M4AFormat page: M4A to Text
File is MP4 (video)Format page: MP4 to Text
Sensitive content (medical, legal, unreleased)Self-hosted Whisper — no upload
40+ hours per month, cost-firstHeavy-volume alternative — see below
Broadcast-standard captionsHuman captioner (Rev, Verbit)

Convert audio to text in 3 steps

The section the SERP skips: you don’t re-upload for subtitles or a summary. One transcription produces text, timestamped subtitles, and an AI summary from the same file in the same editor.

  1. Sign upfor DeluxeScribe (60 minutes free, no credit card). No trial timer — the 60 minutes are yours to use across as many files as you want.
  2. Upload the audio file.Drag any MP3, WAV, M4A, OGG, OPUS, FLAC, AAC, or WebM. Files up to 5 GB accepted — enough for ~40 hours of typical voice audio.
  3. Wait 1–3 minutes for a typical file. Language auto-detects; speaker labels default on for multi-speaker audio. Then export as TXT, DOCX, PDF, SRT, VTT, or JSON.

“Convert audio to text,” “convert audio file to text,” “convert sound to text,” and “turn audio into text” are all the same operation. Google groups them into one intent; the same workflow serves all of them.

Convert your audio to text now

60 minutes free, no credit card. Speaker labels, six export formats, batch upload.

Transcribe audio — verb form

“Transcribe audio” and “convert audio to text” describe the same operation. Some searchers use the verb (“transcribe”), some use the phrase (“convert to text”); both land on the same SERP with the same tools. This section uses the verb form for readers who search that way.

How to transcribe audio for free

Three real options. DeluxeScribe’s 60-minute one-time free tier covers a typical interview or short podcast without a credit card. Self-hosted Whisper is free forever if you can run pip install openai-whisperand wait for processing. Apple Voice Memos on iPhone 12+ (iOS 18+) transcribes on-device for English. Most other “free audio transcription” web tools cap at 5–15 minutes per file or hide a paywall behind download.

Transcribe an audio file

Same workflow, any format. Upload the file, wait for transcription, export the text. DeluxeScribe accepts MP3, WAV, M4A, MP4, OGG, OPUS, FLAC, AAC, and WebM directly — no format conversion needed before upload. Files up to 5 GB per upload; longer files, split with FFmpeg first.

Transcribe an audio recording

Any recording that plays back as audio can be transcribed — voice memos, phone calls, interviews, lectures, podcasts, meeting recordings, dictation. Recording quality matters more than format: a clean recording at 128 kbps transcribes better than a noisy recording at 320 kbps. See the accuracy section below for realistic expectations by recording type.

Audio transcription — free & paid options

“Audio transcription” is the noun form of “transcribe audio.” Same operation, same tools. Below: honest comparison of the real free and paid options.

OptionCostVolume ceilingBest for
DeluxeScribe (us)60 min free → $10/mo1,200 min/mo ProSpeaker labels + all export formats free
Self-hosted WhisperFree forever (MIT)UnlimitedPrivacy + volume, if you can run CLI
TurboScribe90 min/day free → $10/mo unlimitedUnlimited on paid40+ hrs/mo raw transcription
Otter300 min/mo freePaid tiers varyLive meeting capture
HappyScribe10 min free/moPaid per-minuteEuropean workflows
Rev (human tier)$1.50/audio minUnlimitedLegal / evidentiary accuracy

AI audio transcriptionrefers specifically to the cloud + self-hosted options above (everything except Rev human tier). All AI options run Whisper large-v3 or a similar model — accuracy on the same source is roughly identical across services. What differs is monthly limits, export formats included on free, speaker diarization quality, and editor features.

Free audio transcriptionhonestly means one of: DeluxeScribe’s 60-minute credit, self-hosted Whisper, or Apple Voice Memos on-device. Beyond that, “free” usually means “free trial” or “free with watermark.” See Best Free Transcription Software for the machine-readable caps table.

What to look for in an audio-to-text converter

Quick disambiguation: an audio-to-text converter is not an audio format converter. The former uses speech recognition to produce text; the latter (CloudConvert, Zamzar, FreeConvert) changes MP3 to WAV or WAV to FLAC without producing any text. If you want to change your file’s format, you want a format converter. If you want the words as text, you want a transcription tool — this page.

Four criteria that actually matter when comparing audio-to-text converters:

  1. Accuracy on your audio.All top cloud options run Whisper large-v3 or a comparable model — accuracy is roughly identical on the same source. Vendor marketing (“99% accurate”) is typically unsourced. Test with a real file from your own workflow.
  2. Supported formats. Baseline: MP3, WAV, M4A, MP4. Advanced: OPUS (WhatsApp), FLAC (audiophile), WebM (browser recordings). Anything missing means an extra FFmpeg step before upload.
  3. Export formats on the free tier. Many services gate SRT / VTT / DOCX behind paid plans and give free users only TXT. DeluxeScribe unlocks all six (TXT, DOCX, PDF, SRT, VTT, JSON) from the first minute.
  4. Honest free-tier limits.“Free audio-to-text converter” often hides caps: watermark on video export, 5-minute per-file limit, or silent rate-limiting after 5 uploads. Read the free-tier terms before committing.

Realistic accuracy expectations

Every transcription product claims “99% accurate.” None of them are, in the general case. Accuracy depends on the recording conditions, not the vendor — the same Whisper large-v3 model hits 97% on a clean podcast and 75% on a phone call regardless of which wrapper (DeluxeScribe, TurboScribe, HappyScribe) serves it.

Recording typeRealistic accuracyNotes
Studio podcast (single mic, post-production)95–98%Best case for ASR.
Zoom meeting (good headset)90–95%Compression costs a few points.
Phone interview / VoIP80–90%Bandlimited audio drops hard consonants.
Lecture hall from back row75–85%Reverb and distance hurt more than noise.
Outdoor / on-the-street60–80%Wind and traffic dominate.

The honest takeaway: AI transcription is great as a first draftfor most content. Expect to spend 10–20% of the audio runtime reviewing and correcting — much less than typing from scratch, but not zero. Full WER-by-condition breakdown at How Accurate Is Whisper.

Which audio formats work?

DeluxeScribe accepts every common audio and video format directly — no pre-conversion needed:

FormatTypical sourceFormat-specific page
MP3Podcasts, downloads, generic recordingsMP3 to Text
WAVUncompressed studio recordingsWAV to Text
M4AiPhone Voice Memos, AAC recordingsM4A to Text
MP4Video files with audioMP4 to Text
OGG, OPUSWhatsApp voice notes, some browsers— upload directly
FLACLossless music, archival audio— upload directly
AAC, WebMBrowser recordings, some phone apps— upload directly

For unusual formats not listed above (AIFF, WMA, AU), convert to MP3 or WAV first with FFmpeg: ffmpeg -i input.xxx -acodec libmp3lame output.mp3.

Wrong format? Sibling pages

If your file is a specific format, the format-specific page has deeper coverage (bitrate reality for MP3, iOS Voice Memos workflow for M4A, uncompressed audio physics for WAV):

AI transcription — what it actually means for your audio

“AI transcription” means Whisper large-v3 (or a similar architecture) doing the speech recognition. Most modern cloud services — DeluxeScribe, TurboScribe, HappyScribe, ElevenLabs, and even most Otter transcripts — run OpenAI’s Whisper large-v3 or a derivative under the hood. On the same source audio, accuracy is nearly identical across these services. What differs is:

  • Speaker diarization pipeline— the add-on that identifies who spoke each segment. Varies meaningfully between vendors.
  • Editor and export— whether you get a browser editor for corrections, and which formats are unlocked on the free tier.
  • Limits and price— monthly minutes, per-file size caps, and per-minute cost after free.

For a genuinely different ASR engine (not a Whisper wrapper), the real alternatives are Deepgram Nova-3, AssemblyAI Universal-2, Speechmatics, and Google Cloud Speech-to-Text. See our Whisper Alternatives page for the honest ranked comparison.

When DeluxeScribe is the right tool

The right answer is us when you have multiple audio files, need speaker labels included by default, want every export format (TXT, DOCX, PDF, SRT, VTT, JSON) unlocked from the first minute, work in a language other than English, or want a browser transcript editor plus AI summary in one flow. Full capability list on our features page, and background on the team at about. Our 60-minute free tier (no card) covers a typical interview series or podcast episode without paying anything.

Try DeluxeScribe on your own audio

60 minutes free, no credit card. Speaker labels, six export formats, batch upload, AI summaries.

When another tool fits better

Being honest about the sub-cases where DeluxeScribe isn’t the best fit:

  • 40+ hours of audio per month, cost-first. TurboScribe’s Unlimited plan at $120/year is hard to beat on raw volume. Our Pro plan caps at 1,200 min/month — genuine tradeoff. See TurboScribe Alternatives.
  • Sensitive content that can’t be uploaded. Self-hosted Whisper is the correct answer — same engine, on your own machine, no cloud policy required. Court audio, PHI without a BAA, unreleased music masters.
  • Legally-critical accuracy. Rev human tier ($1.50/audio minute) or Verbit for education / legal / government. AI is good enough for most content; when a judge might read the transcript, humans still win.
  • Live meeting workflows.Otter, Notta, or Fireflies. Different product category — a bot joins your Zoom / Meet / Teams call and transcribes in real time. If you’re just exporting the recording afterward, that’s a hint the meeting-bot flow would save you a step.
  • Broadcast / OTT delivery with QC. BBC iPlayer, Netflix, HBO Max spec compliance requires human captioning. AI needs manual segmentation cleanup at minimum.

How this page was verified

Accuracy ranges in the recording-type table reflect our own testing across common source conditions, scored against human-corrected reference transcripts using word error rate (WER) — the standard methodology described in the OpenAI Whisper paper (Radford et al., 2022). The engine behind most cloud audio-to-text services — DeluxeScribe, TurboScribe, HappyScribe, and much of the market — is Whisper large-v3; accuracy differences between wrappers are essentially noise on identical source audio. See How Accurate Is Whisper for the WER-by-condition breakdown and Whisper Alternatives for genuinely different ASR engines (Deepgram, AssemblyAI, Speechmatics). Competitor free-tier caps verified August 4, 2026against each vendor’s live pricing page.

Frequently Asked Questions

Is audio to text really free?

Genuinely free options exist: DeluxeScribe's 60-minute one-time credit (no card), self-hosted Whisper if you can run a Python command, and Apple Voice Memos on iPhone 12+ under iOS 18+. Most other tools advertised as 'free audio to text' cap at 5–15 minutes per file, watermark output, or rate-limit silently at 5–20 uploads per day. See our best-free-transcription-software comparison for the machine-readable free-tier caps table.

How do I convert audio to text?

Four real paths. (1) Cloud service: sign up somewhere (DeluxeScribe gives 60 minutes free), drag the audio file, get text back in 1–3 minutes for typical files. (2) Self-hosted Whisper: pip install openai-whisper, then whisper input.mp3 --model large-v3 --output_format txt. Free forever, slow on CPU, fast on Apple Silicon or NVIDIA GPUs. (3) Apple Voice Memos on iPhone 12+ under iOS 18+, English on-device. (4) YouTube Studio auto-caption for a video you own (convert audio to MP4 first).

How do I transcribe audio?

Same operation as "convert audio to text" — Google treats both queries identically. Upload the file to a transcription service, wait a few minutes, get the transcript with timestamps and (optionally) speaker labels. DeluxeScribe handles it in one flow: sign up, drag the file, language auto-detects, export as TXT / DOCX / PDF / SRT / VTT / JSON with word-level timestamps.

How do I transcribe audio for free?

Three real options. (1) DeluxeScribe's 60-minute one-time free tier — no credit card, no watermark, all export formats unlocked. Enough for a couple of hour-long recordings. (2) Self-hosted Whisper — free forever, but needs Python and either patience (CPU: 10–30× real-time) or a GPU. (3) Apple Voice Memos on iOS 18+ (iPhone 12+) — free on-device transcription for English and a handful of supported languages. Most 'free audio transcription' web tools cap at 5–15 min per file.

What audio formats can I convert to text?

MP3, WAV, M4A, MP4, OGG, OPUS, FLAC, AAC, and WebM all work. Video files (MP4, MOV, AVI, MKV) also work — the audio track is extracted and transcribed automatically. Format-specific pages exist for the most common cases: MP3, M4A, WAV, MP4. For unusual formats, use FFmpeg to convert to MP3 or WAV before uploading.

What's the difference between an audio-to-text converter and an audio format converter?

An audio-to-text converter uses speech recognition to turn spoken audio into readable text — that's what this page is about. An audio format converter (like CloudConvert, Zamzar, FreeConvert) changes one audio format to another (MP3 to WAV, WAV to FLAC) — the audio stays as audio, no text produced. If you searched "audio converter" hoping to get text, you actually want a transcription tool. If you want to change your file's format, you want a format converter.

How accurate is AI audio transcription?

Depends entirely on recording conditions, not the vendor. On clean studio audio with one speaker, modern AI (Whisper large-v3 and its wrappers) hits 95–98% word accuracy. On Zoom recordings with headsets: 90–95%. On phone calls: 80–90%. On noisy outdoor recordings: 60–80%. Expect to spend 10–20% of the audio runtime reviewing and correcting the transcript — much less than typing from scratch, but not zero.

How long does a 1-hour audio file take to transcribe?

Modern cloud services (DeluxeScribe included) complete a 1-hour audio file in 5–10 minutes for typical files. Self-hosted Whisper on CPU takes 10–30 hours — near real-time on Apple Silicon or NVIDIA GPUs. Real-time transcription tools (meeting bots) take exactly 60 minutes because they run alongside the audio.

Can I transcribe audio in another language?

Yes. DeluxeScribe supports 99 languages with automatic detection. Quality varies — English, Spanish, French, German, Portuguese, Italian, Japanese, Korean, and Mandarin are strongest. Less-resourced languages have higher error rates but produce usable transcripts for most cases. For short clips (<30 seconds), set the language manually — auto-detection can misidentify.

How do I extract text from an audio file?

"Extract text from audio" is the same operation as "convert audio to text" — run speech recognition on the audio. Fastest path: upload the file to DeluxeScribe (60 minutes free, no card), wait 1–3 minutes, then copy or export the text. If you need to do it offline for privacy, self-hosted Whisper is the answer: install once, run forever without internet.

How do I get text from an audio recording?

Same workflow whether it's a voice memo, meeting recording, lecture recording, or podcast: upload the audio file to a transcription service, wait a few minutes, get the text with timestamps. DeluxeScribe's free tier handles the first 60 minutes without a card. For very long recordings (>30 minutes), the Pro plan supports files up to 4 hours; on the free tier, split with FFmpeg first: ffmpeg -i long.mp3 -c copy -segment_time 1800 -f segment chunk_%03d.mp3.

Can I get an audio transcript with speaker labels?

Yes — this is the biggest difference between free web tools and paid services. DeluxeScribe includes automatic speaker diarization (identifies and labels each speaker) on every plan. Free web tools usually don't. Self-hosted Whisper needs an add-on like whisper-diarization or WhisperX for speaker labels. Speaker attribution accuracy is 85–95% on clean multi-speaker audio and drops on overlapping or phone-quality speech.

Is my audio kept private?

DeluxeScribe processes audio on encrypted infrastructure (TLS in transit, encrypted at rest on AWS EU region) and doesn't use customer content to train models. Files can be deleted anytime from the dashboard. For maximum privacy — court audio, PHI without a BAA, unreleased music masters — run Whisper locally. Your audio never leaves your machine, no cloud vendor policy required.

What's the best audio to text converter?

There isn't a single best — it depends on volume and needs. For occasional use with speaker labels, AI summaries, and every export format from the first minute, DeluxeScribe (yes, that's us). For 40+ hours/month of raw transcription, TurboScribe's unlimited plan is unbeatable on price. For sensitive audio, self-hosted Whisper. For legally-critical accuracy, Rev's human tier. All four cloud options run Whisper large-v3 or a similar model — accuracy is nearly identical; what differs is limits, editors, exports, and price.