Back to News & insightsResearch

Speech recognition quality: what word error rate leaves out

Use word error rate alongside speaker attribution, critical facts, and correction effort to evaluate transcripts for real work.

Editorial guide · Updated September 27, 2026 · 4 min read
A silver acoustic wave passes through glass and becomes separated luminous beads.

Two transcripts can have similar error counts and very different consequences. One misses a filler word; the other changes a deadline. If a team evaluates speech recognition only by how natural the transcript looks, those differences can disappear beneath readable prose.

Word error rate is a useful starting point because it compares a transcript with a reference. A product decision also needs to ask which errors matter, who said each statement, and how much work is required to correct the result.

What word error rate measures

WER counts substitutions, deletions, and insertions relative to a reference transcript, divided by the number of reference words. Hugging Face's evaluation metric documentation describes this conventional calculation. The metric depends on alignment and text-normalization choices, so those choices belong with the reported score.

In an illustrative hundred-word reference, three substitutions, two deletions, and one insertion produce a WER of six percent. This arithmetic describes that invented example only. It says nothing about whether the six errors were harmless or whether one reversed the meaning of an instruction.

Establish the transcript's purpose

Imagine a fictional project team using transcripts to draft meeting actions. It needs the responsible person, the agreed task, and the deadline. It may not need every hesitation, but it must preserve corrections and negations.

Decide whether the reference should be verbatim or lightly normalized. Specify how to handle numbers, acronyms, unfinished sentences, and overlapping speech. Reviewers cannot produce a consistent ground truth if each applies a different transcription style.

Keep the audio available to authorized evaluators. A polished reference can contain human mistakes too. When a model and reference disagree about an important detail, return to the recording rather than automatically treating the reference as infallible.

Build a critical-fact layer

Alongside WER, label the facts the downstream workflow relies on. For the meeting assistant, these could include names, dates, task ownership, explicit cancellations, and changes of plan. Score preservation of those items separately.

Include a correction such as a speaker changing a deadline later in the conversation. The transcript should preserve enough evidence for the summarizer to identify the final agreement. A downstream summary that chooses the earlier date may be a reasoning error even when the transcript is correct.

This separation is valuable during debugging. Test summarization with a human-corrected transcript as well as the automatic one. If the action item remains wrong, improving the speech recognizer alone will not solve the product failure.

Speaker attribution is its own task

A transcript can contain the right words under the wrong speaker. In a meeting workflow, that can assign responsibility incorrectly. Evaluate speaker diarization and any mapping from anonymous speaker labels to actual people separately from word accuracy.

Use recordings with interruptions, similar voices, remote microphones, and a person joining late. Do not infer identity from a voice unless the application has an appropriate, validated identity process. An anonymous label can be more honest than a confident but unsupported name.

When attribution is uncertain, carry that uncertainty into the action draft. Asking a reviewer to confirm ownership is better than quietly treating an uncertain speaker assignment as a verified commitment.

Evaluate representative recording conditions

Create groups for the microphones, accents, languages, room acoustics, and background noise the product will encounter. Obtain recordings through an appropriate consent and data-handling process. Avoid judging an international service only on clean studio speech.

Measure per-group behavior and inspect poor cases. An overall average can look acceptable while one recurring setting remains unusable. Also test silence and non-speech audio, where an invented transcript can create facts that no one said.

Keep timestamps and segmentation quality in the review if users need to jump back to the audio. A transcript that is accurate but difficult to verify can still impose substantial correction effort.

Choose a workflow people can trust

Track how long reviewers spend producing an accepted meeting record. Include the time needed to correct names, resolve speakers, and verify deadlines. This connects model quality to the actual work the tool is intended to save.

Publish a transcript as a draft until its critical details have passed the required review. A useful speech system preserves a path back to evidence; readability alone should not turn uncertain recognition into an authoritative record.

Technical background

Hugging Face: Technical documentation

An original editorial guide. Provider capabilities and documentation can change. Follow the linked sources and test the exact model or service before relying on it.