How to Measure Transcription Accuracy (Without Fooling Yourself)

The WER formula, the normalization rules that make the number mean something, and a calculator that shows both.

Count three things: words swapped, words missing, words invented. Divide by the words in the reference transcript. That's WER — word error rate, the standard accuracy metric for transcription. The formula takes thirty seconds; what separates a real measurement from a misleading one is everything around it: the reference transcript, the normalization rules, the sample size, and knowing what a “good” number looks like for your audio. This page covers all of it, with a calculator that runs in your browser.
  • Formula + in-browser calculator
  • Normalization rules spelled out
  • Thresholds by audio condition
  • No vendor leaderboard — a method you can run

Last verified: October 4, 2026

The 60-second method

In short:transcribe a representative sample, correct it by hand into a reference, fix your text-normalization rules, align the two transcripts, and divide the error count by the reference word count. Interpret the result against your audio condition — not against a vendor's marketing number.

  1. Pick representative audio — the audio you actually process, not the cleanest clip you have. If your real traffic is Zoom calls, test Zoom calls.
  2. Build the reference transcript — a human-verified, word-exact transcript of that sample. This is the expensive step, and the one most evaluations skip.
  3. Decide normalization before scoring — case, punctuation, numbers, contractions, fillers. Same rules, both sides, decided in advance.
  4. Align and count — substitutions, deletions, insertions over a minimum-edit-distance alignment. The calculator below does this for you.
  5. Interpret by condition — 8% on phone audio is a good result; 8% on studio audio is a poor one. The thresholds table has honest ranges.

What WER actually counts

Word error rate compares the machine transcript (the hypothesis) against a verified human transcript (the reference), over the alignment that minimizes total edits:

WER = (S + D + I) / N

  • S — substitutions:a reference word replaced by a different word (“Alvarez” → “Suarez”)
  • D — deletions: a reference word the machine dropped entirely
  • I — insertions: a word the machine invented that is not in the reference
  • N — the number of words in the reference, not in the hypothesis

Because N comes from the reference and insertions add errors without adding reference words, WER can exceed 100%. Reference: “stop the recording”. Hypothesis: “please stop um the long recording now”. All three reference words are present, but three inserted words make it 3/3 — 100% WER on a transcript containing every correct word. Audio that triggers hallucination (long silences, music) does this in practice.

One more consequence worth internalizing: WER is computed on an alignment, so it does not care whichwords were wrong. Getting “um” wrong costs exactly as much as getting a drug name, a price, or the word “not” wrong. More on that in when WER is the wrong metric.

Prepare your reference transcript — the expensive step

Every error you count is counted against the reference, so the reference has to be right. This is the part of the method that costs real time: a careful word-exact reference for one hour of multi-speaker audio is several hours of human work. Three rules keep it honest:

  • Don't build the reference by lightly editing the ASR output you are about to score.You will unconsciously accept the machine's phrasing wherever it is plausible, and your measured WER will come out flattering. Either transcribe from the audio directly, or have the correction done by someone who hasn't seen the hypothesis — and budget a second listening pass purely for small function words (“a” vs “the”, dropped “not”), which are the words correctors skim past.
  • Write down the transcription conventions. Does the reference include filler words? How are numbers written? Are contractions kept as spoken? Whatever you decide becomes part of your normalization policy in the next step — undocumented reference conventions are where irreproducible WER numbers come from.
  • Human disagreement is the floor. Linguistics Data Consortium studies estimate careful human transcribers disagree with each other on roughly 4.1–4.5% of words. A measured WER near that range on hard audio may be measuring your reference as much as the machine.

Normalization: decide the rules first

Two teams can score the same transcript pair and report 9% and 3% — both honestly — because they made different text-normalization choices. This is the single biggest source of incomparable WER numbers, and most published figures don't state their policy. Here is ours, in full. Use it, or write your own — but write it down before you look at any scores, and apply every rule to both transcripts.

The DeluxeScribe normalization policy (v1)

  1. Unicode and whitespace — normalize Unicode (NFKC), trim, collapse repeated whitespace. Always applied; never a judgment call.
  2. Case— lowercase both sides. “Friday” vs “friday” is not a transcription error.
  3. Punctuation — strip it, unless punctuation quality is itself what you are evaluating (captions, readability). Keep intra-word apostrophes consistent with rule 4.
  4. Contractions— expand to one canonical form on both sides: don't → do not, we'll → we will, can't → cannot. Expansion beats contraction because ASR engines emit either form unpredictably.
  5. Numbers— one policy, both sides: either all digits to words or all words to canonical digits. Your policy must cover ordinals, decimals, currencies, dates, times, and phone numbers — decide whether $5 is “five dollars”, “5 dollars”, or “$5” before scoring, not after.
  6. Filler words— remove um, uh, er from both sides for content transcription. Keep them and score them explicitly if hesitation detection matters to your product. Do not remove “like”, “well”, or “you know” without a documented rule separating filler use from meaningful use.
  7. Abbreviations and spelling — optionally normalize Dr. → doctor, colour → color, and hyphenation. Less important than rules 1–6, but whatever you choose, keep it fixed across every experiment you intend to compare.

A WER number published without its normalization policy is a number you cannot compare to anything. That cuts both ways: it is why this page teaches a method instead of publishing a vendor leaderboard.

WER calculator

Paste a reference and a hypothesis, toggle the normalization rules, and watch both numbers: the raw WER and the WER under your policy. The gap between them is rule-dependence you'd otherwise never see.

Runs entirely in your browser — nothing you paste here is uploaded.

Normalization policy (applied to both sides)

Paste both transcripts (or load the sample pair) to see the raw and normalized scores side by side.

Worked example: one pair, four scores

The calculator's sample pair is one spoken sentence — a reference with normal written formatting, and a plausible all-lowercase ASR output that says “we will” for “we'll”, “doctor” for “Dr.”, adds one “um”, and renders $1,450 as “fourteen fifty”. Here is that identical pair scored under four normalization regimes:

RegimeSDINWER
Raw — exact tokens, case and punctuation count10032065.0%
+ lowercase7032050.0%
+ strip punctuation6022138.1%
Full policy (contractions, numbers, fillers)3502729.6%

Computed with the calculator above — load the sample pair and reproduce every row by toggling the rules.

Same audio, same transcripts: 65% or 30% depending on normalization alone.And the residual errors under the full policy are instructive too. “dr” vs “doctor” survives because the v1 policy leaves abbreviations optional (rule 7) — add that rule and the score drops further. The number mismatch survives because the speaker said “fourteen fifty” colloquially while the reference wrote $1,450, which canonicalizes to “one thousand four hundred fifty” — a genuine ambiguity your number policy has to take a position on. This is why the policy comes before the scoring.

What's a good WER?

Only answerable per audio condition. These ranges match our measured Whisper accuracy report and hold approximately for current large ASR models in English:

Audio conditionRealistic WERReading
Clean studio speech (audiobook, podcast with good mic)2–5%At or near the human-disagreement floor
Clean meeting audio, one speaker at a time5–10%Publication-grade after light cleanup
Zoom/Teams calls, mixed mics10–15%Usable transcript; review quotes before citing
Phone audio (narrow-band 8 kHz)15–25%Gist-level; expect real correction work
Speech over music, heavy crosstalk30%+Spot-check everything; hallucinations likely
Human transcriber vs human transcriber (LDC estimate)4.1–4.5%The floor “100% accurate” pretends not to have

Two consequences. First, a vendor's “99% accurate” claim is either a human-reviewed service (a person fixes the machine output) or a best-case benchmark on clean audio under favorable normalization — it is not a prediction about your Zoom recordings. Second, a 3% WER claim on arbitrary audio would beat the measured human floor, which should make you ask for the dataset and the normalization policy, not applaud.

Want the measurement done for you first?

Upload 60 minutes free, export the transcript, and score it against your own reference with the calculator above. No card required.

When WER is the wrong metric

WER measures text similarity, not meaning, and that produces known blind spots:

  • Every word costs the same.Dropping “um” and dropping “not” are each one deletion. A transcript can have excellent WER and still invert the one sentence that mattered — or terrible WER from fillers while every fact survived. If specific terms are critical (names, drugs, amounts, commands), measure them separately: build a term list and report recall on those terms alongside WER.
  • Valid paraphrases are penalized.“it's $5” vs “it is five dollars” can be three errors or zero depending on normalization — and genuinely different-but-equivalent phrasing always scores as error. WER rewards verbatim fidelity, which is usually what you want from transcription, but it makes WER unsuitable for judging summaries or translations.
  • CER for boundary-free languages and near-misses. For Japanese, Chinese, or Thai, word segmentation is itself ambiguous — score characters (CER), not words. CER is also the fairer lens on near-miss spellings of names: “Alvarez” → “Alvares” is one full word error but a single character error.
  • Formatting quality is a separate axis. Punctuation, capitalization, paragraphing, and speaker labels determine whether a transcript is pleasant to read — and standard WER deliberately normalizes all of that away. If readability matters, evaluate it explicitly; don't expect the WER number to carry it.

Measuring speaker accuracy (DER) — separately

“Who said it” fails independently of “what was said”, so it gets its own metrics. Benchmark diarization separately from recognition — blending them produces numbers nobody can act on.

DER = (missed speech + false alarm + speaker confusion) / total reference speech time

  • Report DER by component.An identical 15% DER can be mostly missed speech (the system didn't hear anyone) or mostly speaker confusion (it heard everything and attributed it wrong) — completely different failures demanding different fixes.
  • cpWER— concatenate each speaker's words, then compute WER under the optimal permutation of speaker labels. Forgives arbitrary speaker numbering (machine's “Speaker 2” = reference's “Alice”).
  • SA-WER — every word must be attributed to the correct reference speaker, no permutation rescue. The strictest combined metric.
  • ΔcpWER = cpWER − WER — the single most useful derived number: it isolates how much error comes from speaker attribution rather than recognition.

Why Δ matters in practice: pipeline wiring can cost more accuracy than model choice. In our own end-to-end test of a popular open-source Whisper + pyannote recipe on the AMI meeting corpus, the diarization model alone scored 8.7% DER — consistent with its published benchmarks — while the combined pipeline's speaker-attributed error landed near 39%. The segment-merging and word-to-speaker assignment glue between the two models cost roughly thirty points that neither model's standalone benchmark would ever show. If you evaluate a speaker-attributed transcription product, measure the pipeline you will actually run, never the component benchmarks.

The mistakes that invalidate your test

Everything above is the method. These are the failure modes we keep seeing — several of them lessons from our own benchmark work, learned the expensive way:

  • Testing on audio that doesn't match production. Clean read-aloud test clips predict almost nothing about call audio. Synthetic or re-recorded audio behaves measurably differently from organically captured speech — if your traffic is phone calls, your eval set is phone calls.
  • Deciding normalization after seeing the scores. Once you know which rules help which system, every toggle is a thumb on the scale. Policy first, scores second — that is the entire reason the policy section precedes the calculator on this page.
  • Building the reference from the output under test — the anchoring problem from the reference section; it silently flatters whichever system generated the draft.
  • Sample sizes that can't support the verdict. Published vendor guidance converges on ~30 minutes for a quick screen, ~1 hour baseline for uniform speech, 2–3 hours when accents, crosstalk, jargon, or noise are in play — and roughly an hour per audio condition, not one mixed pool. On five minutes of audio, a single mangled sentence moves WER by whole points.
  • Calling a tie a win. Two systems at 8.1% and 8.4% on one hour of audio are statistically indistinguishable. If the gap is smaller than what re-running on a different same-condition sample would produce, report a tie.
  • Comparing numbers across different datasets or dates. A WER from vendor A's blog on their dataset versus vendor B's on theirs is not a comparison. Models also update silently — date-stamp every result and note the engine version where you can.

Running a fair vendor comparison

  1. Pick 1–3 audio conditions that match your production traffic.
  2. Build one reference set per condition (~1 hour each).
  3. Write the normalization policy down (start with v1 above).
  4. Run every candidate on the identical files in the same week; record the date.
  5. Score with the same tool and policy; report S/D/I by condition, plus your critical-term recall if you defined one.
  6. Treat overlapping results as ties; decide those on price, speed, export formats, and data handling instead.

This is also why cross-vendor WER comparisons you find online are rarely meaningful: without identical audio, identical references, and an identical stated normalization policy, the numbers are not comparable — which is the standard position in the field, not a convenient one. It is why this page hands you a method instead of a leaderboard.

And it applies to us. DeluxeScribe runs on Whisper large-v3 — our measured condition ranges are published in the Whisper accuracy report. Don't take either page's word for it: run this method on our output too. The free tier (60 minutes, no card) is enough for a one-hour eval against your own reference.

Tools for the counting step

  • This page's calculator — in-browser, nothing uploaded, normalization toggles with raw-vs-policy side-by-side. Fine for pasted pairs up to ~2,000 words.
  • jiwer — pip install jiwer; the standard Python library for scripted evaluation, with its own composable normalization transforms.
  • NIST SCTK / sclite — the reference implementation used in formal ASR benchmarking; heavier setup, canonical output reports.
  • asr-evaluation — pip install asr-evaluation, then wer reference.txt hypothesis.txt for quick command-line runs with per-sentence breakdowns.
  • For published model-level numbers, the Open ASR Leaderboard scores open models on shared datasets with a shared normalizer — the comparability this page keeps insisting on, applied at dataset scale. Several commercial vendors also publish their own benchmarking tooling (AssemblyAI maintains an SDK for it); remember any vendor's harness embeds that vendor's normalization choices.

How this page was verified

The WER definition follows the NIST scoring convention implemented in SCTK/sclite. Human inter-transcriber disagreement figures are Linguistics Data Consortium estimates. Accuracy ranges by audio condition match our Whisper accuracy report, which cites its benchmark sources. Every number in the worked example is reproducible in the calculator on this page.

Frequently asked questions

What is an acceptable word error rate?

It depends entirely on the audio condition — there is no single acceptable number. On clean studio English, modern models score 2-5% WER; on clean meeting audio 5-10% is normal; Zoom calls land around 10-15%; narrow-band phone audio 15-25%; speech over music 30% or worse. For context, careful human transcribers disagree with each other by roughly 4-4.5% (Linguistics Data Consortium estimates), so a 3% WER on clean audio is effectively at the human floor. Judge any number against the matching condition, never against a vendor's benchmark figure.

Can WER be over 100%?

Yes. WER divides errors by the number of words in the reference, and insertions add errors without adding reference words. If the reference is 'stop the recording' and the ASR outputs 'please stop um the long recording now', you get 3 insertions on a 3-word reference even with every reference word correct — 100% WER. Noisy audio that triggers hallucinated text routinely pushes WER past 100%.

Does punctuation count as an error in WER?

Only if your normalization policy says so. Standard practice is to strip punctuation and lowercase both transcripts before scoring, so 'friday' vs 'Friday.' is not an error. But this is a choice, not a law — if you are evaluating readability or caption quality, you may deliberately score punctuation. The non-negotiable part: apply the same rules to both the reference and the hypothesis, and decide the rules before you look at any scores.

What is the difference between WER and CER?

WER counts errors at the word level; CER (character error rate) counts them at the character level. CER is the better metric for languages without clear word boundaries (Japanese, Chinese, Thai) and is more forgiving of near-miss spellings — 'Alvarez' vs 'Alvares' is one full word error in WER but a single character error in CER. For English content transcription, WER is the standard; use CER when word segmentation itself is ambiguous.

How many samples do I need for a fair accuracy test?

Published evaluation guidance from ASR vendors converges on: about 30 minutes for quick screening of uniform audio, about 1 hour as a baseline for ordinary single-language speech, and 2-3 hours when accents, multiple speakers, jargon, or noise are involved. If your production audio spans several conditions (clean, accented, phone, noisy), test roughly an hour per condition rather than one big mixed pool — a single blended number hides exactly the failures you need to see.

What does '99% accurate' actually mean in transcription marketing?

Usually one of two things: a human-reviewed service guaranteeing 99% after a person corrects the machine output, or a best-case benchmark number measured on clean audio with favorable normalization. It is not what you should expect on your own recordings — real-world WER varies from 2-5% on studio audio to 30%+ on speech over music. When you see a flat accuracy percentage with no stated audio condition, dataset, or normalization policy, treat it as marketing, not measurement.