# Cisco Support Services 2026: 10% WER Threshold vs Speaker Diarization

Piper Bowen · October 6, 2026

> Takeaway Detail Treat 10% WER as the transcription-accuracy threshold for the 2025 Cisco support-transcription decision. The guide's thesis ties the Cisco Suppo

| Takeaway | Detail |
| --- | --- |
| Treat 10% WER as the transcription-accuracy threshold for the 2025 Cisco support-transcription decision. | The guide's thesis ties the Cisco Support Services GenAI-era decision to a 10% WER threshold and speaker diarization. |
| Require speaker diarization that returns participant speaker labels out of the box for real-time and async transcription. | Recall.ai docs state perfect diarization returns participant speaker labels by default and is supported for both real-time and async transcription. |
| Do not accept a single DER benchmark without specified audio conditions. | Gladia says acceptable diarization error rate thresholds depend on audio conditions, and vendor benchmarks quoting one DER figure without them are not actionable. |
| Check all transcription parameters, including speaker diarization, in index.js via TRANSCRIPTION_CONFIG. | kossakovsky/transcriber configures transcription parameters in index.js by modifying TRANSCRIPTION_CONFIG at lines 28-95, where speaker diarization is listed. |

This guide is a practical verify-before-you-commit reference for the 2025 support-transcription decision, centered on a 10% WER threshold and speaker diarization. It shows how to verify the live, complete option and compare like-for-like totals and terms before committing.

![Cisco Support Services 2026](https://static.mm-ais.com/article-images-ai/cisco-support-services-2026-10-wer-thres-ai-d7dd522e.jpg)

## How It Works

First, confirm that the transcription pipeline is set up to extract audio from the source video and then run speech‑to‑text with speaker diarization enabled. In the kossakovsky/transcriber repository this is done by editing the TRANSCRIPTION_CONFIG object in index.js (lines 28‑95) and setting speakerDiarization: true 【kossakovsky/transcriber】. When this flag is present the system calls the ElevenLabs Scribe API (or another backend) to produce a transcript that includes speaker turn information.

Next, verify that the output actually contains speaker labels. According to docs.recall.ai, when “perfect diarization” is enabled the transcript is returned with participant speaker labels out of the box, and this works for both real‑time and asynchronous processing 【docs.recall.ai】. A practical check is to inspect the first few lines of the transcript: each sentence should be prefixed with a speaker identifier (e.g., [Speaker 1] or [Host]). If labels are missing, the diarization step did not run or was disabled.

Then, check the Word Error Rate (WER) of the transcript against your internal quality target. WER measures the proportion of words that are inserted, deleted, or substituted relative to a reference transcript and is a well‑established metric for speech‑to‑text accuracy. You should run a WER evaluation (using a tool such as jiwer or the vendor’s built‑in scorer) and confirm that the result is below the threshold you have defined for your support workflow. No specific number is prescribed here; the rule is simply to verify that your chosen WER limit is satisfied.

After that, compute the Diarization Error Rate (DER) to assess speaker attribution quality. DER is defined as the sum of missed speech (E_miss), false‑alarm speech (E_fa), and speaker‑confusion duration (E_conf) divided by the total ground‑truth speech time (T_total) 【Gladia - Diarization error rate (DER) explained】. In practice you obtain E_miss, E_fa, and E_conf by aligning the system’s speaker‑segment output with a manually annotated reference and summing the mismatched time intervals.

Finally, apply a rule: accept the transcription only if the calculated DER is below the maximum error rate you have established for your use case. Because DER captures errors that WER ignores—such as missing speaker turns or assigning words to the wrong speaker—it is essential for multi‑speaker support transcripts where downstream tasks (e.g., LLM summarization or ticket routing) depend on correct speaker attribution. If DER exceeds your limit, return to the configuration step, adjust diarization sensitivity or model choice, and re‑run the checks.

![How It Works — Cisco Support Services 2026](https://static.mm-ais.com/article-images-ai/cisco-support-services-2026-10-wer-thres-ai-ad3a871f.jpg)

## Key Factors to Consider

This section lists the top decision criteria and the numbers that matter when evaluating a live transcription option before you commit. Verify that the service meets each criterion on a like‑for‑like basis; only then should you proceed with a purchase or integration.

**Word Error Rate (WER)** – WER expresses the proportion of words that are incorrectly transcribed, presented as a percentage. Before committing, check that the provider’s reported WER aligns with your internal quality target; a lower WER indicates higher lexical accuracy. The Gladia DER explainer notes that WER alone captures only half of the picture for multi‑speaker workflows, so it must be evaluated alongside speaker attribution metrics.

**Speaker Diarization Accuracy** – Diarization performance is quantified by the Diarization Error Rate (DER), which measures the share of audio time where speech is missed, falsely detected, or assigned to the wrong speaker. DER is calculated as (E_miss + E_fa + E_conf) / T_total, yielding a percentage of total speech duration. Verify that the service returns speaker‑labeled segments (as confirmed by the Recall.ai diarization documentation) and that its DER falls within the range you deem acceptable for your use case.

**Configuration and Integration Effort** – Assess whether the transcription pipeline can be readily adapted to your environment. The kossakovsky/transcriber repository shows that all transcription parameters, including speaker diarization toggles, are exposed via a TRANSCRIPTION_CONFIG object, allowing you to enable or disable features without code changes. Additionally, consider user‑friendly adoption: screenapp.io highlights that podcasters and content creators routinely pull diarized transcripts directly into show notes, indicating a low‑friction workflow.

| Criterion | What to Verify | Source |
| --- | --- | --- |
| Word Error Rate (WER) | Measured WER meets your internal quality target (percentage) | Gladia DER explainer |
| Speaker Diarization Accuracy | DER within acceptable range; output includes speaker labels | Gladia DER explainer; Recall.ai diarization docs |
| Configuration & Integration Effort | Adjustable via TRANSCRIPTION_CONFIG; minimal workflow changes | kossakovsky/transcriber; screenapp.io |

![Key Factors to Consider — Cisco Support Services 2026](https://static.mm-ais.com/article-images-pixabay/cisco-support-services-2026-10-wer-thres-9fb779a9.jpg)

## Common Mistakes

The first mistake is reading a diarization score as a property of the service rather than of the audio it was measured on. Gladia's guide to diarization error rate puts it bluntly: vendor benchmarks quoting a single DER figure without specifying audio conditions are not actionable. A concrete case: a team validates on a clean, two-person podcast-style recording — the exact use case ScreenApp describes, where the raw episode comes back split by host and guest — then commits to a support line dominated by hold music, crosstalk, and speakerphone audio. DER is a time-based metric, so non-speech labeled as speech (a false alarm) and speech that goes undetected (a miss) both inflate it, on top of outright speaker confusion. Test against a clip that sounds like your worst real recording, not your best.

The second mistake is using word error rate as a stand-in for attribution quality. Gladia warns that measuring WER is only half the picture and that a transcript can post excellent WER while remaining operationally useless because speaker turns are misattributed. Picture a refund call: the agent says "I'll credit that back," the customer says "I'll send the receipt." Swap the labels and your summarizer assigns the customer's action to your team. Verify attribution directly — mark speaker turns on a short sample by hand and count how many land on the wrong person.

The third mistake is validating only one delivery mode. Recall's diarization documentation states that perfect diarization returns participant speaker labels out of the box and is supported for both real-time and async transcription. Teams routinely prove the async batch path on a recorded call, then commit to a live-agent deployment without repeating the test in the streaming path — or the reverse. Whichever mode your production traffic will use, run the diarization check in that mode before you commit.

A fourth mistake is committing against a demo or spec sheet instead of the live option. Re-run your own audio on the exact account and configuration you intend to buy, and confirm the output actually carries speaker labels where you expect them. If the live result comes back unlabeled, treat that as a configuration or tier gap rather than a flaw in your test clip — and resolve it before the commitment, not after.

The working rule: score your audio, in your mode, on the live account, and read attribution separately from word accuracy. A clean benchmark sample and an unspecific demo are both evidence of very little about your workload.

![Common Mistakes — Cisco Support Services 2026](https://static.mm-ais.com/article-images-pixabay/cisco-support-services-2026-10-wer-thres-551de491.jpg)

## Insider Tactics

The non-obvious move is to stop evaluating demos and start running a vendor-style pipeline on your own audio. In the kossakovsky/transcriber repository, every transcription parameter lives in one TRANSCRIPTION_CONFIG object in index.js (lines 28–95), including the speaker diarization switch. That single control point is the tactic: point that pipeline at your recordings, toggle diarization yourself, and produce transcripts under your conditions instead of judging a curated sample. When you request a pilot, ask for the config file or the equivalent settings panel — not a link to a sample transcript.

Build the pilot set from your worst calls, not your cleanest ones. Pull recordings with crosstalk, speakerphone audio, and background noise, hand-label who spoke when for a representative portion, then score the vendor's speaker-labeled output against your labels — recomputing from the segment timestamps rather than accepting a summary figure. Then test the actual downstream use: paste the labeled transcript into the summarizer, ticketing system, or QA process you really run. A transcript that reads as an unattributed wall of text is the failure you want to catch before signing, not after.

Control the mode when you compare. Diarization output is produced on both real-time and async paths — Recall.ai notes perfect diarization is supported in both — so a real-time transcript measured against an async one is not a like-for-like comparison. Run both passes over the same audio with settings frozen between runs, then decide which mode you are buying. If the vendor only quotes one mode, get the other mode's output in writing before you treat the pilot as complete.

On timing: schedule the pilot to close before you sign anything with a term, a price lock, or an auto-renewal clause. Verification only carries leverage while the complete option is still live on the table; after commitment, the pilot becomes documentation. Ask in writing for the terms attached to the tier you are piloting, and confirm that the tier you tested is the tier you will be billed for.

Then re-time after any change. Model versions, diarization settings, and processing mode all alter output, so a pilot that passes under one configuration is not a verified result under another. Treat the configuration itself as a term of the deal: pin it, request notice before it changes, and retain the right to re-run the pilot on the same labeled audio. Never commit on a configuration you cannot rerun.

![Insider Tactics — Cisco Support Services 2026](https://static.mm-ais.com/article-images-pixabay/cisco-support-services-2026-10-wer-thres-121133aa.jpg)

## Comparison

The three live options I compare most often are an async speech-to-text API with diarization switched on (Gladia publishes the DER methodology), a real-time service that returns participant labels out of the box (Recall.ai documents perfect diarization for both real-time and async transcription), and the self-hosted kossakovsky/transcriber tool, a Node.js application that runs on the ElevenLabs Scribe API and keeps every transcription parameter in the TRANSCRIPTION_CONFIG object at lines 28–95 of index.js. The table normalizes what each one hands you so the totals line up before you commit.

| Option | What lands in your inbox | Like-for-like check |
| --- | --- | --- |
| Async diarization API | Speaker-labeled chunks plus a DER you can decompose | Request missed speech, false alarm, and speaker confusion separately; Gladia states that a single DER figure quoted without audio conditions is not actionable |
| Real-time diarization service | Participant speaker labels streamed during the session | Confirm both real-time and async are covered, since Recall.ai supports perfect diarization in both modes |
| Self-hosted transcriber repo | Your own audio pipeline and your own configuration file | Verify you can edit TRANSCRIPTION_CONFIG yourself and that the ElevenLabs Scribe dependency is acceptable under your data terms |

Put the numbers on one denominator. Gladia defines DER as (E_miss + E_fa + E_conf) divided by T_total, the total ground-truth speech duration. If your scored call contains 30 seconds of missed speech, 10 seconds of false alarm, and 60 seconds of speaker confusion, the numerator is 100 seconds; against 1,000 seconds of ground-truth speech that is a 10% DER. Keep the same 100 seconds of error and drop the speech total to 500 seconds, and the DER doubles to 20%. Nothing about the system changed — only the audio you scored it on.

That denominator is why a Word Error Rate number and a diarization number belong in different columns. WER is a share of words; DER is a share of speaking time. Gladia's own explainer notes a transcript can post excellent WER and still be operationally useless when speaker turns are misattributed. Before you sign, ask each vendor for its component breakdown on your audio, then hold every option to the WER ceiling you already fixed and to the same recorded sample set.

Each option wins in a different situation. Real-time diarization wins when labels are needed while the conversation is still happening and an async pass would arrive too late. An async API wins when you can re-score against ground truth, tune the pipeline, and compare error components across several runs before buying seats. The self-hosted repository wins when the configuration itself has to stay under your control and you would rather maintain the extraction and transcription steps than delegate them.

The rule that survives all three: compare like-for-like totals and terms. Same audio, same speech-duration denominator, component-level error figures, and the full subscription term in the quote. If a vendor will not break the number down on your own recordings, the comparison is incomplete — treat that as a reason to keep evaluating, not as a reason to commit.

## What to do next

| Step | Action | Why it matters |
| --- | --- | --- |
| 1 | Before committing, pull the live, complete Cisco support-transcription option set for the GenAI-era decision rather than a cached datasheet or an older proposal version — confirm the transcription capability you are scoring is the one currently shipping for both real-time and async use. | This is the first half of the guide's canonical rule: verify the live, complete option before committing. A retired or partial SKU can look identical on a feature grid while failing the actual decision criteria. |
| 2 | Force the vendor to state, on the record, whether speaker diarization returns participant speaker labels out of the box for both real-time and async transcription — the exact bar set in the Recall.ai docs referenced above. | Diarization that only emits anonymized speaker turns does not satisfy the guide's requirement, and a claim that holds in async mode may not hold in real time. The two modes must be confirmed separately. |
| 3 | Reject any proposal that quotes a single diarization error rate figure without the audio conditions behind it; ask for the DER measured on the audio profile closest to your Cisco support calls (channel count, noise, overlap, accent mix). | Gladia states that acceptable diarization error rate thresholds depend on audio conditions. A lone DER benchmark is not actionable and cannot be compared like-for-like against another vendor's number. |
| 4 | Re-run the comparison table above so the WER result and the DER result are read from the same rows: same audio, same speaker count, same mode (real-time versus async), same evaluation window. | Mixed-basis numbers are the most common way a vendor clears the 10% WER threshold on paper while failing it in production. Like-for-like totals are the second half of the canonical rule. |
| 5 | Check all remaining transcription parameters alongside WER and speaker labels — punctuation, timestamp granularity, custom vocabulary, redaction, and retention — and confirm each one is present in the live option you are about to sign. | A transcription service that clears the accuracy threshold but drops a required parameter still fails the Cisco support-transcription decision. Verify completeness before commitment, not after. |
| 6 | Put the final like-for-like totals and terms side by side — identical scope, identical modes, identical measurement basis — and only then commit to the option that survives every check above. | Terms that shift between the evaluation and the contract (mode coverage, diarization labels, measurement conditions) invalidate the comparison the guide's decision rests on. |

## Quick answers

| What is the transcription-accuracy threshold for the Cisco support-transcription decision? | The transcription-accuracy threshold for the Cisco support-transcription decision is the 10% WER threshold. |
| --- | --- |
| What does the guide's thesis tie the Cisco Support Services GenAI-era decision to? | The guide's thesis ties the Cisco Support Services GenAI-era decision to the 10% WER threshold and speaker diarization. |
| What does Recall.ai docs state about perfect diarization? | Recall.ai docs state that perfect diarization returns participant speaker labels by default and is supported for both real-time and async transcription. |
| Why should you not accept a single DER benchmark without specified audio conditions? | You should not accept a single DER benchmark without specified audio conditions because acceptable diarization error rate thresholds depend on audio conditions, and vendor benchmarks quoting one DER figure without them are not actionable. |
| Where can you check all transcription parameters, including speaker diarization? | You can check all transcription parameters, including speaker diarization, in index.js via TRANSCRIPTION_CONFIG. |

Also worth reading: **Diarization Cuts Podcast WER by 18.4%: Pre vs Post ASR**: [Diarization Cuts Podcast WER by](https://transcribeall.io/blog/diarization-cuts-podcast-wer-by-184-pre-vs-post-asr.php) · **Whisper's 2026 WER: Evidence, Decision Matrix, and Variance**: [Whisper's 2026 WER: Evidence, Decision](https://transcribeall.io/blog/whispers-2026-wer-evidence-decision-matrix-and-variance.php) · **Whisper large-v3 Fine-Tuning: 18% WER Cut on Indian English Calls**: [Whisper large-v3 Fine-Tuning: 18% WER](https://transcribeall.io/blog/whisper-large-v3-fine-tuning-18-wer-cut-on-indian-english-calls.php)

### Related reading

- [7 Real-World Applications of Speaker Diarization in Audio Transcription Technology](https://transcribeall.io/blog/7_real_world_applications_of_speaker_diarization_in_audio_tr.php)
- [Overlapping Speech Error Rates: Use 1—ASR Word Error Rate, Not Diarization](https://transcribeall.io/blog/overlapping-speech-error-rates-use-1asr-word-error-rate-not-diarization.php)
- [MIT Audit: Diarization Costs Valid; Three-Zone ROI Exposes Risks](https://transcribeall.io/blog/mit-audit-diarization-costs-valid-three-zone-roi-exposes-risks.php)
- [MIT SLP 2026 Diarization: Embedding Drift & Pipeline Architecture](https://transcribeall.io/blog/mit-slp-2026-diarization-embedding-drift-pipeline-architecture.php)
- [Diarization Cuts Podcast WER by 18.4%: Pre vs Post ASR](https://transcribeall.io/blog/diarization-cuts-podcast-wer-by-184-pre-vs-post-asr.php)
- [Whisper VAD: 32% Diarization Error Reduction Is Conditional](https://transcribeall.io/blog/whisper-vad-32-diarization-error-reduction-is-conditional.php)

### Latest

- [Overlapping Speech Error Rates: Use 1—ASR Word Error Rate, Not Diarization](https://transcribeall.io/blog/overlapping-speech-error-rates-use-1asr-word-error-rate-not-diarization.php)
- [Podcast transcription errors 2026: $0.12 auto vs $1.79 human review](https://transcribeall.io/blog/podcast-transcription-errors-2026-012-auto-vs-179-human-review.php)
- [Speech Recognition Error Rates: Whisper Wins Off-Domain at Zero Labels](https://transcribeall.io/blog/speech-recognition-error-rates-whisper-wins-off-domain-at-zero-labels.php)

Canonical: https://transcribeall.io/blog/cisco-support-services-2026-10-wer-threshold-vs-speaker-diarization.php
Markdown: https://transcribeall.io/blog/cisco-support-services-2026-10-wer-threshold-vs-speaker-diarization.php/index.md
