| Takeaway | Detail |
|---|---|
| For accurate transcripts, select 1—ASR by the lowest WER on representative overlapped audio. | The WER evaluation must include speech during overlap, not just single-speaker segments. |
| Judge diarization separately with overlap-aware DER. | DER combines three errors over total speech time: missed speech, false alarm, and speaker confusion. |
| Do not treat published diarization error rates as overlap performance. | Those rates usually exclude overlapping speech and speech near turn boundaries—the moments that matter most for voice agents. |
| Add diarization only when speaker attribution is required. | If the deliverable is an accurate transcript rather than an authoritative speaker ledger, diarization remains a separate attribution layer. |
This guide explains how to evaluate transcription systems when multiple people speak at once. It establishes inclusive WER as the standard for 1—ASR and overlap-aware DER as the separate test for diarization.

Measure Words, Not Speaker Turns
For a transcript intended to recover what was said, use a single end-to-end 1—ASR system and score its output with word error rate (WER) over the complete recording, including simultaneous speech. Define 1—ASR once for the evaluation: one end-to-end system whose output is compared with the reference transcript across all scored audio, without first removing overlap-heavy regions. The primary deliverable is the words and their sequence, not a complete ledger of who spoke when.
Compute WER as (substitutions + deletions + insertions) ÷ reference words. Apply one documented normalization policy before counting errors, covering punctuation, casing, and filler conventions. For example, decide whether contractions remain separate tokens, whether interjections and filled pauses count as words, and whether punctuation affects tokenization. Apply the same policy to every system and region; otherwise, a formatting difference can look like a recognition error.
Report WER separately for isolated turns, adjacent turns, and simultaneous speech, while retaining an inclusive WER for the entire scored recording. Each region must contain enough reference words for the result to be interpretable, so report the reference-word count alongside the error total. A system that performs well on isolated speech but substitutes or omits words during overlap has not demonstrated transcript accuracy for the actual audio. Do not evaluate only a diarization-clean subset: that would conceal the cases the transcript must handle.
Speaker names should be treated as optional labels attached to an otherwise accurate transcript. A correct word assigned to the wrong speaker may still be correctly transcribed, while an omitted or substituted word cannot be repaired by assigning it to another speaker. If attribution matters, evaluate that layer independently with an overlap-aware measure such as diarization error rate (DER), rather than allowing speaker labels to substitute for lexical evaluation.
A practical acceptance check is to inspect the three region-level WER reports and the inclusive WER, then review every high-error overlap segment against the reference. Confirm that normalization rules were applied consistently, that simultaneous speech was not silently discarded, and that any speaker labels are reported as attribution metadata. On this basis, select the 1—ASR system with the lowest inclusive WER on representative overlapped audio, and add diarization only when speaker attribution is itself required.

Why Overlap Breaks Clean DER
Published diarization scores often look strongest because they exclude overlapping speech and speech near turn boundaries—the exact intervals a voice agent must process in real conversations. Cekura makes this limitation explicit: overlap and boundary speech are typically removed before diarization performance is reported. The practical check is to inspect the score’s evaluation protocol and ask what share of simultaneous speech was excluded. A result based only on clean, single-speaker regions is not evidence of performance on the difficult part of the recording.
Standard diarization error rate (DER) aggregates missed speech, false alarms, and speaker confusion over the reference speech included in the evaluation. If a benchmark removes overlap-heavy regions, the denominator no longer represents those regions, and substantial failures there can remain invisible. In that setting, DER can approach zero even while the system assigns no reliable structure to the portions where conversational meaning may depend on two voices. Treat a clean-region DER as the misleading state that excludes the overlap-heavy audio where conversational meaning often depends on both voices.
When source-separated audio and a matching reference are available, a useful diagnostic is to segment the recording, transcribe each separated stream, and compare its words with the corresponding reference words. Keep this as a source-separation check, not as a general transcript-quality score: ask whether words assigned to each stream are recovered, whether the streams together cover the complete audio, and whether overlap causes one voice to suppress or absorb the other. A low WER on separated streams does not excuse a complete transcript that drops simultaneous speech.
When source separation is unavailable, do not manufacture a word-level comparison against a reference that cannot identify which separated stream should have produced each word. Instead, document the limitation and evaluate the transcript against the full reference, including overlapping speech. Record whether the test set contains true simultaneous speech, speech at turn boundaries, and multiple voices with similar pitch or accent. Also report any passages that were segmented by the system before transcription, because hidden preprocessing can make the evaluation less comparable than the label suggests.
The rule is simple: before accepting a published DER, demand the inclusion rules, not just the headline number. Confirm that overlap and turn-boundary speech are scored, identify the reference speech time used as the denominator, and inspect at least one overlap-heavy example manually. If the evaluation removes those intervals, use the number only for the narrower clean-region task it actually measures—not as a proxy for whether a voice agent can handle the conversation as heard.

Compare Three Production Choices
Evaluate all three pipelines against the same corpus of recordings, the same complete-audio reference transcripts, and the same scoring script. Include natural two-person overlap, interruptions, interruptions with correction, and background noise. The decisive comparison is word error rate on the full reference, not performance on selected clips where speakers happen to alternate cleanly. Freeze normalization rules, punctuation handling, number formatting, and the treatment of fillers before running any system.
| Pipeline | What it optimizes | Required check | Verdict |
|---|---|---|---|
| Single end-to-end 1—ASR | Recovery of all words in the complete recording | Inclusive WER on the same overlapped references | Winner for transcript accuracy |
| ASR followed by diarization | Adding speaker labels to an existing transcript | Confirm that attribution does not drop, merge, or replace words | Secondary option when attribution is required |
| Overlap detection or source separation before transcription | Resolving simultaneous speech into streams for attribution or downstream processing | Re-score the combined transcript after concatenation, and inspect overlap regions separately | Secondary option for speaker attribution or specialized processing |
Set the production threshold before reviewing vendor output. For the highest-ranked option, require the lowest inclusive WER on the representative test set, with a predefined margin for operational concerns such as latency, reliability, and data handling. If a pipeline removes overlapping audio to make diarization easier, restore the complete recording and rescore it: a cleaner speaker ledger is not an acceptable substitute for missing or duplicated words. A useful audit is to compare insertions, deletions, and substitutions within overlap intervals as well as across the full program.
Add a diarization layer only when the deliverable requires answers to questions such as who said a particular line or who interrupted whom. Evaluate that layer independently with overlap-aware DER, while retaining inclusive WER as the transcription acceptance test. For a separation-first design, concatenate all recognized streams back into one transcript before scoring; otherwise, improvements in stream-level clarity can conceal losses at speaker handoffs. The combination is production-ready only when both checks pass.
Repeat the comparison under the exact conditions expected in production, including representative accents, channel quality, microphone placement, and maximum simultaneous speakers. Preserve failure logs for every overlap-heavy failure, and require manual review of the worst examples. This creates a repeatable decision rule: select the lowest-WER transcript engine, then add attribution only for a defined speaker-facing requirement rather than paying the extra complexity by default.

Budget for Accuracy, Not Headlines
Budget against the words that reach the transcript, not against a headline rate. The cited batch diarization prices—$0.288 for 60 minutes, $4.80 for 1,000 minutes, and $480 for 100,000 minutes—illustrate linear volume scaling: $0.288 × 1,000 ÷ 60 = $4.80, and $4.80 × 100 = $480. They do not establish transcription accuracy, and this section treats those operating points and prices as non-comparable until overlap coverage, audio hours, and output scope are disclosed.
For each pipeline, calculate cost per correctly transcribed reference word:
(total pipeline cost) ÷ (reference words − inclusive-WER errors)
A cheaper service can still be more expensive per accepted word if it produces substantially more errors. Compare the ASR-only option and any ASR-plus-diarization option using the same representative recordings, the same inclusive WER evaluation, and the same billing assumptions. Include model fees, per-minute or per-hour charges, minimum durations, and any charges for retries or extra output fields. Then ask a simple control question: did the price buy better words, or merely speaker labels?
Before accepting a quote, require it to state the number of audio minutes, e

Know What These Numbers Cannot Prove
ElevenLabs Scribe-v1’s reported DIHARD III performance is useful evidence about that system under that benchmark’s conditions, not a guarantee that it will lead on your calls, meetings, or multilingual conversations. The decisive limitation is comparability: vendor accuracy tables cannot establish cross-domain superiority when overlap prevalence, language, and scoring exclusions differ. Check whether the published result includes overlapping speech, which languages and accents are represented, how many speakers occur, and whether any intervals or reference words were excluded. A benchmark win transfers cleanly only when the test population and scoring policy resemble the audio on which you will rely.
Set a minimum acceptance test of 30 minutes of your own audio. Do not choose a merely “representative” clip: include every interval in that sample where two or more people speak simultaneously, even if those intervals make the vendor’s result look worse. Preserve the full reference transcript, document the language mix and recording conditions, and compare outputs under the same rules. This test exposes a vendor weakness that a clean-speech sample can conceal and gives you a threshold for deciding whether the reported benchmark advantage is practically relevant to your workload.
Keep transcription quality and speaker attribution separate in the test report. For the transcript, compare the words recovered across the complete sample, including every simultaneous utterance. For attribution, report overlap-aware diarization performance and note where speakers were confused or omitted. If a system wins the benchmark but fails this test because it drops a word during overlap, its DIHARD III standing does not make it the better choice for recovering what your participants actually said.
Do not infer that Italian—or any language—is intrinsically harder from diarization errors alone. The Italian Transcription Guide’s observation that meeting recordings can produce more diarization errors is a diagnostic prompt: inspect overlap, turn-taking behavior, recording quality, participant count, and scoring exclusions before offering an explanation. Then run the same 30-minute overlap-inclusive comparison for that language. A result can reflect your meetings rather than a language-wide causal difference, and adding an attribution layer cannot repair words the transcription system failed to recover.

Audit One-Minute Overlapped Call
Start the audit with a 60-second, manually verified reference transcript. Include isolated speech, at least one turn boundary, and a 20-second interval in which two people speak simultaneously. Count all 150 spoken reference words, including both voices during overlap. Do not collapse the overlap into one dominant-speaker track, omit false starts, or count speaker turns in place of words. The acceptance rule is simple: every spoken word must be represented once in the reference and every extra, missing, or substituted word in the hypothesis must be visible in the WER calculation.
Suppose the single end-to-end ASR candidate produces 9 substitutions, 3 deletions, and 6 insertions. The error count is 9 + 3 + 6 = 18, so inclusive WER is 18 ÷ 150 = 0.12, or 12.0%. Verify the same result with the standard expansion: (9 + 3 + 6) ÷ 150 × 100 = 12.0%. Accept this candidate for the transcript only if its 12.0% WER is lower than the incumbent’s WER on the same 150-word reference and the complete 60-second audio. Segment-level averages, isolated-speech scores, or a diarization result cannot replace that comparison.
Now review the speaker-attribution result. Suppose diarization reports 6.0% DER because its clean-region evaluation excludes the overlapping interval, even though nine reference words occur there and are missing from the attributed output. This is the audit’s crucial distinction: a low DER can coexist with lost conversational content. DER measures speaker-assignment mistakes over speech time; it does not certify that every word was recovered. Before accepting the diarization result, inspect the overlap segment itself and identify every omitted word.
Record those nine words as DER-only omissions attributable to the attribution layer, not as additional ASR errors already counted above. They are not ignored, but they are not double-counted either: the ASR WER assesses transcription accuracy, while the overlap inspection exposes a separate failure to deliver an attributed transcript. The check is to reconcile the final attributed output against the original 150-word reference and investigate every missing word, especially throughout the 20 seconds of simultaneous speech.
The decision follows the deliverable. If the goal is the most accurate transcript of what was said, the 12.0% inclusive-WER candidate advances over any incumbent scoring worse than 12.0%, even when another system advertises 6.0% DER. Select the diarization layer only if its output passes the separate overlap-aware attribution check. Reject any apparent DER winner that leaves nine words missing from the overlapping region.
Apply Five Operational Decisions
If the transcript must show what every participant said, require inclusive WER below 10% before accepting the 1—ASR output for editorial-quality transcription. Calculate the score against the complete reference transcript, retaining words spoken during overlap and checking substitutions, deletions, and insertions over the full recording. If 10% is too strict—or too lenient—for the material, set a task-specific threshold before testing vendors, and document the sample composition, language mix, and overlap level. This makes the acceptance test depend on recovered words rather than on the number of speaker turns.
If the output also needs reliable speaker names, run a separate diarization comparison and require both inclusive WER and overlap-aware DER. Report the two results side by side: WER establishes whether the words were recovered, while DER evaluates attribution on the same audio, including overlapping speech. Do not treat a low DER as evidence that the transcript is accurate, or a low WER as evidence that speaker labels are reliable. Names should be added only after the chosen diarization layer meets the project’s attribution requirements.
If overlap occupies less than 5% of reference words, use inclusive WER to choose the transcriber and add only lightweight attribution. Count the reference words inside intervals with simultaneous speech and divide by the total reference-word count. Below that cutoff, prioritize the system with the best inclusive WER, then apply a modest labeling pass for names or roles. Still inspect a sample for obvious speaker mix-ups, but do not make a full diarization deployment the price of a mostly single-speaker transcript.
If overlap reaches 5% or more of reference words, treat attribution as a demonstrated evaluation requirement. Compare diarization systems separately, preserve their overlap scores, and test them on the recordings where simultaneous speech actually occurs. Because published DER figures often exclude overlapping speech and turn-boundary regions, the comparison should use a reference that includes those intervals rather than relying on a headline score. Keep diarization in an attribution layer and keep transcript accuracy in the WER decision.
If the deliverable is only a faithful record of what was said, do not add diarization merely because it is available. Gate the extra layer on a specific need for speaker names, accountability, or individualized retrieval. Run the final acceptance check twice: once on overall inclusive WER and, when attribution is required, once on overlap-aware DER. That separation lets the transcript remain the primary standard while ensuring that speaker labels are judged only when they matter.
What to do next
| Step | Action | Why it matters |
|---|---|---|
| 1 | Use representative audio with speech occurring during overlap to compare 1—ASR systems by word error rate (WER). | WER must measure what the transcript says, including overlapped speech, rather than only single-speaker segments. |
| 2 | Select 1—ASR with the lowest WER on that overlapped-audio set. | For an accurate transcript, transcription performance is the decision criterion—not the diarization score. |
| 3 | Evaluate diarization separately using overlap-aware DER. | DER combines missed speech, false alarms, and speaker confusion over total speech time, so it measures attribution independently. |
| 4 | Do not treat published diarization error rates as proof of overlap performance. | Those rates commonly exclude overlapping speech and speech near turn boundaries, exactly where voice-agent attribution is most difficult. |
| 5 | Add diarization only if the deliverable requires authoritative speaker attribution. | Diarization is a separate attribution layer, not a substitute for choosing 1—ASR by lowest WER. |
| 6 | Document the WER, overlap-aware DER, and speaker-attribution requirement before finalizing the transcription setup. | Separate evidence keeps transcript accuracy distinct from the need for a speaker ledger. |
Frequently Asked Questions
How should a 1—ASR system be selected for accurate transcription of overlapped speech?
Select the system with the lowest WER on representative overlapped audio, with the evaluation including speech during overlap.
Should overlap-heavy regions be removed before computing WER?
No, 1—ASR should be evaluated across all scored audio without first removing overlap-heavy regions.
Which three errors are combined in overlap-aware DER?
DER combines missed speech, false alarm, and speaker confusion over total speech time.
Why should published diarization error rates not be treated as overlap-performance results?
Those rates usually exclude overlapping speech and speech near turn boundaries.
When should diarization be added to a transcription workflow?
Add diarization only when speaker attribution is required.
How is 1—ASR defined for this evaluation?
It is one end-to-end system whose output is compared with the reference transcript across all scored audio.
Quick answers
| How should an 1—ASR system be selected for accurate transcripts? | Select 1—ASR by the lowest WER on representative overlapped audio. |
| What audio must be included in the WER evaluation? | The WER evaluation must include speech during overlap, not just single-speaker segments. |
| How should diarization be evaluated? | Judge diarization separately with overlap-aware DER. |
| Which three errors does DER combine? | DER combines missed speech, false alarm, and speaker confusion over total speech time. |
| When should diarization be added? | Add diarization only when speaker attribution is required. |
Also worth reading: Convert Video to Audio for Transcription: Mono WAV Cuts Word Error Rate 25% vs MP4: Convert Video to Audio for · How classic algorithms power the next generation of speech recognition: How classic algorithms power the · Overlapping speech in meetings: Two-pass stack cuts 29-31% vs single pass: Overlapping speech in meetings: Two-pass