The Direct Answer: “85% Accuracy” Usually Is Not One Accuracy Number
A quoted 85% real-world ASR result usually combines several conditions that are absent from a clean lab test: accents, background noise, overlapping speakers, technical vocabulary, clipped audio, multiple dialects, and long-form recording defects. It may also represent an average across recordings rather than the score of a particular audio file, and “accuracy” itself may mean word error rate, word accuracy, exact-match accuracy, or a business-specific quality score. Those measures are not interchangeable, so comparing 85% with a lab claim above 95% can be misleading unless both numbers use the same corpus, language, scoring method, and failure policy. The honest conclusion is not that laboratory systems are fraudulent or that 85% is a universal ceiling; it is that production audio is substantially more difficult than curated evaluation audio.
Also worth reading: How Do You Evaluate German Speech-to-Text Systems for Accuracy in 2026? · How Do You Benchmark AI Transcription Systems for Accuracy, Speed, and Cost? · What Are the True Real-Time ASR Accuracy Tradeoffs in 2026?
As of September 27, 2026, a strong production target is often expressed as an overall word error rate below 10%, which corresponds to roughly 90% word accuracy under a simple 1-WER conversion. For clean, single-speaker English with a common vocabulary, 2–5% WER can be realistic. For crowded, multilingual, telephone, or domain-heavy material, even 10–20% WER may be normal before correction. Organizations should therefore set thresholds by use case rather than adopt a universal percentage: legal verbatim transcripts, search indexing, call analytics, and rough content drafts have different costs for the same recognition error.
Why Laboratory Scores Exceed Production Results
High-scoring benchmarks commonly use segmented recordings with controlled microphones, limited background noise, known speakers, and transcripts prepared in advance. They may also exclude proper names, hesitations, false starts, music, and code-switching, or evaluate short utterances that never accumulate drift across an hour of audio. A model can perform exceptionally well on such data while still losing accuracy when the channel changes between a studio microphone and a laptop, room microphone, Bluetooth headset, or telephone codec. Compression and packet loss introduce errors that are not primarily semantic model failures, yet they affect the transcript users ultimately receive.
Production systems also face a broader input distribution than benchmark designers anticipated. Users speak different English accents, switch between languages, place words in noisy environments, and use specialized terms that appear rarely in general training data. In customer support, for example, product names, account identifiers, serial numbers, and regulatory language can dominate useful information. Word error rate treats every substitution as equal, so missing “Acme Industrial Systems” may matter more than correctly reproducing several ordinary function words. Research on Indic ASR likewise argues that WER alone is insufficient because grammaticality, meaning, and task success can diverge, particularly in multilingual or code-switched speech.
| Condition | Curated lab audio | Real-world audio | Expected effect |
|---|---|---|---|
| Word error rate | 2–5% on clean English | 8–15% on ordinary business speech | More substitutions, deletions, and insertions |
| Audio bandwidth | Often wideband, 16 kHz | Often 8 kHz or variable | Lower recognition of consonants and names |
| Background noise | Usually low or controlled | Vehicles, offices, music, crowds | More insertions and substitutions |
| Speaker overlap | Often absent | Common in meetings and calls | Attribution and recognition errors rise sharply |
| Vocabulary | General and known | Industry-specific and unpredictable | Proper names are often misheard |
| Reporting unit | Short, clean utterances | Hours of continuous audio | Small errors accumulate |
| Scoring | Often plain WER | WER plus semantic or task metrics | A single percentage hides useful differences |
Audio Quality Often Matters More Than Model Choice
Channel quality is one of the simplest controls. Speech below 16 kHz can still be transcribed correctly, especially with modern models trained on varied audio, but an 8 kHz telephone channel removes acoustic information and makes similarly sounding consonants harder to distinguish. Sampling rate, bitrate, codec, clipping, gain, and noise are therefore worth recording in every test. A 48 kHz studio file sent through a poorly configured conferencing chain may be less usable than a 16 kHz file captured cleanly, because hardware quality and preprocessing can be more consequential than nominal resolution.
Microphone placement also changes the difficulty of a recording. A headset positioned two to four centimeters from the mouth generally provides a better signal-to-noise ratio than a laptop microphone across a desk, although exact distances depend on the device. In meetings, one microphone per participant is preferable to placing a single device in the center of a table. Teams should use wired headsets where practical, disable aggressive noise suppression when it creates musical or watery artifacts, and avoid automatic gain settings that amplify room noise after quiet passages. For high-stakes interviews, redundancy matters: maintain a local high-quality recording even if the conferencing platform also creates a compressed cloud recording.
Silence, breath sounds, and long pauses should be handled explicitly. Aggressive voice activity detection can remove useful words at the beginning or end of a turn, while automatic segmentation can cut a sentence in half. Modern systems usually improve when they receive enough surrounding context, but some pipelines make the opposite mistake by sending isolated five-second chunks to an API. Longer context can resolve a technical term, yet it can also increase latency and cost. A practical balance is short-turn streaming for interactive applications and wider context for post-call transcription, while measuring both systems on complete files.
How Evaluation Changes the Result
The phrase “real-world ASR” has no standardized test protocol. Some evaluations use random samples from support calls, others use complaint cases, internal meetings, or a fixed duration per speaker. Sampling can bias the result substantially: a production system may look excellent if it handles short, well-recorded calls and poor if every evaluated item is a difficult voicemail. Report sample size, duration, language, source, recording channel, overlap rate, and confidence intervals. A score based on 20 clips is not a dependable basis for a contract, while 100 hours sampled across channels and accents can support a much stronger purchasing decision.
For standard comparative testing, WER remains useful because it is reproducible. Its basic formula divides substitutions, deletions, and insertions by the number of words in the reference transcript. Deletions are omitted words, substitutions replace reference words with incorrect ones, and insertions add words that were not spoken. Exact-match accuracy is stricter and useful for short commands, while semantic similarity can detect whether a transcript preserves meaning despite different wording. Diarization error rate should be measured separately when the task needs to know who said what, and named-entity accuracy is more relevant when customer, medical, or product names carry the information.
Human and language-model judgments can help assess grammar, adequacy, or task success, but they should not replace basic transcript metrics. A grader can sometimes recognize that “the invoice was sent Tuesday” and “we sent the invoice Tuesday” communicate a similar fact, yet a less capable judge may be inconsistent across a long corpus. Human and LLM-based evaluations are strongest when calibrated against trained reviewers, given explicit rubrics, and sampled for disagreement. Research published by Hasegawa Johnson in August 2025 on intelligibility and pronunciation assessment illustrates why phonetic, semantic, and natural-language-inference approaches can complement one another, but combining scores must remain methodologically transparent.
A Practical Workflow for Measuring Production Accuracy
Begin with a representative gold set assembled from actual use cases. Include easy and difficult audio rather than filtering exclusively for dramatic failures. Aim initially for at least 10–30 hours if the budget permits, and stratify results by channel, language, accent where appropriate, speech rate, noise level, overlap, and business domain. Every item should have a corrected reference transcript, with transcription conventions fixed for punctuation, numbers, spellings, and partial words. The reference should describe what was reasonably intelligible, not contain facts the speaker never said, and disagreements should be adjudicated rather than settled by whichever evaluator ran last.
Then test the complete pipeline. Request a transcript through the same API, upload behavior, normalization settings, language detection, diarization, and post-processing used in production. Do not benchmark a clean model output and ignore the product’s proprietary vocabulary layer, or measure only the engine while excluding human review. Record latency, API failure rate, missing-audio rate, speaker-attribution error, processing time, and cost per audio minute. These operational values can matter more than a small WER gain: a 4% versus 5% WER difference is rarely decisive if one option loses calls, requires manual reprocessing, or returns unstable results for a rare dialect.
| Decision goal | Preferred threshold | Why it matters |
|---|---|---|
| General business transcription | WER at or below 10% | Usually a practical target for readable drafts and search |
| Clean single-speaker English | 2–5% WER | Strong quality is often attainable on controlled audio |
| Telephone or noisy calls | 8–15% WER for baseline | May require vocabulary, diarization, or human review |
| Exact commands and short answers | At least 90–95% exact-match accuracy | Word-level averaging can hide critical short-utterance failures |
| Speaker attribution | Diarization error near or below 10% for many workflows | Required when speaker identity changes the meaning |
| Critical entity accuracy | At least 95% for agreed high-value fields | Measures product names, amounts, and identifiers directly |
| Live voice interaction | Response latency below roughly 800 ms for natural turn-taking | Transcript accuracy alone does not make a usable agent |
Common Measurement Mistakes and How to Avoid Them
One common error is treating accuracy as character-level agreement. Character accuracy can look favorable because spelling mistakes add or remove only one character, but it overweights long words and does not measure word substitutions cleanly. Another is using WER without a scoring standard: case, punctuation, numbers, contractions, and spelling normalization can move a score by several points. If a system omits punctuation while the reference includes it, the evaluation must define whether punctuation counts as words and whether normalization occurs before or after semantic comparison.
Another mistake is averaging everything into one number. A 95% score on ten hours of mostly clean audio should not hide 35% WER on the two hours containing heavy overlap and low bandwidth. Report a total score and category scores, with minimum thresholds for critical groups. Organizations also make the mistake of using synthetic or read speech as a substitute for field recordings. Read sentences are valuable for regression tests, but they lack natural timing, interruptions, and spontaneous phrasing, so they cannot establish real-world performance by themselves.
Finally, avoid test-set contamination. If a vendor repeatedly tunes on the same public audio samples and reports the best result, the number may reflect familiarity rather than generalization. Use a private holdout set that is not used during prompt, model, or vendor configuration. Freeze the holdout until a major revision, and inspect failures manually. If post-processing improves the score, count that cost and latency too; “ASR accuracy” delivered after a large language model rewrite may be semantically better but less faithful, which is unacceptable for strict verbatim use.
Model, API, and Human Review Options
There is no universally best provider because performance depends on language, audio, deployment constraints, and scoring rules. Open-source systems such as Whisper can be self-hosted and offer strong general transcription, but the operator still pays for computing, monitoring, updates, and engineering. Cloud APIs simplify operations and commonly provide higher throughput, but pricing, retention rules, regional availability, and model changes create tradeoffs. Enterprise platforms may add diarization, vocabulary controls, review tools, and compliance features that a raw transcription endpoint does not include.
| Feature | General cloud or open model | Specialized domain workflow | Human-assisted service |
|---|---|---|---|
| Typical cost structure | Per audio minute or infrastructure cost | Per minute plus configuration or subscription | Per minute or per completed task |
| Setup | Low to moderate | Moderate | Low for the customer, higher for provider operations |
| Custom vocabulary | Model-dependent | Often configurable | Expert can correct names directly |
| Privacy control | Varies by hosting choice | Often includes contractual options | Provider handles files under agreed terms |
| Best fit | Drafting, search, broad coverage | Calls, meetings, regulated terminology | Legal, medical, media, and high-risk material |
| Main limitation | Hidden settings or operational burden | Configuration can be complex | More expensive and slower |
Human review remains appropriate when verbatim fidelity is legally or financially material. A sensible escalation policy sends low-confidence passages, overlapping speech, very low audio quality, and segments containing key entities to a trained reviewer. It also keeps random samples of high-confidence output in human review, because confidence scores are not guarantees. This arrangement lets people focus on difficult passages instead of retyping every word while still measuring residual error.
When to Act and What Good Performance Looks Like
Act immediately when transcription errors affect money, compliance, safety, accessibility, or customer trust. In those cases, collect a failure sample before changing providers and name the harm explicitly. If amounts, medication names, legal admissions, or speaker identities are wrong, overall WER is too broad a target; measure those entities and review workflows directly. A 97% whole-corpus score can still be unacceptable if a small number of monetary substitutions occur systematically. Conversely, occasional punctuation errors may be tolerable in an internal search index and do not justify a costly project.
A defensible purchasing decision uses at least two candidate systems, the same holdout audio, the same reference transcripts, and a scoring script reviewed by the buyer. It examines clean and difficult strata rather than only the aggregate, and it prices the final workflow including diarization, storage, post-processing, and review. Organizations should also verify what happens when the API is unavailable, how long audio is retained, whether data is used for training under the selected contract, and where processing occurs. Those facts can outweigh a one-point benchmark difference.
There is no evidence that all real-world ASR is “stuck” at 85%. Performance can improve through better microphones, lower noise, appropriate chunks, stronger language identification, domain adaptation, phonetic or LLM-based evaluation, and workflow-specific correction. The apparent gap mainly exposes differences in difficulty and measurement. Treat 95% as a conditional claim about a defined test, and treat 85% as another conditional claim until both are placed under the same test design. For an audio-to-text buyer, that discipline is more useful than searching for the largest percentage on a model card.