Technical explainer

Word error rate is not medical accuracy.

WER tells you how many words changed. It does not tell you whether the changed word was “Tuesday,” “metoprolol,” or “50 milligrams.” Medical speech needs separate clinical error views.

Omi Health ResearchPublished 1 September 2026Benchmark v7
5.99%Omi overall WER on the current sealed board.
0.94%Medical-term WER on 2,872 medical terms.
0 / 442Omi drug-name errors on 442 mentions.
97.7Dosage F1 across 89 aligned dosage events.

Omi medical API, benchmark v7. The board contains 1,513 clinical clips, 7.18 hours and 63.8k reference words. The leading three systems statistically tie on medical-term error; these figures do not establish universal accuracy.

The short answer

Word error rate is the number of substitutions, deletions and insertions divided by the number of words in the reference transcript. It is useful for overall transcription quality. For healthcare, add medical-term error, drug-name error and dosage precision/recall so clinically important failures cannot disappear inside the average.

What is word error rate?

Word error rate, usually shortened to WER, compares an automatic transcript with a reference transcript. An alignment identifies three kinds of mistakes:

  • Substitution: the system produced a different word.
  • Deletion: a reference word is missing from the transcript.
  • Insertion: the transcript contains a word that is not in the reference.
WER = (substitutions + deletions + insertions) ÷ reference wordsMultiply by 100 to express the result as a percentage. Lower is better.

If a 100-word reference contains two substitutions, one deletion and one insertion, the WER is 4%. The metric is easy to reproduce and useful for broad regressions, but every changed word has the same weight.

Why ordinary WER can hide medical risk

Consider the reference “start amoxicillin 500 milligrams.” A transcript that changes “start” to “begin” and a transcript that changes “500” to “50” may each receive one substitution. They do not have the same consequence for a downstream clinical note or medication workflow.

WER also mixes a large number of ordinary conversational words with a much smaller number of diagnoses, anatomy, medicines and measurements. A system can improve the common words enough to post a strong overall score while still making more of the mistakes a healthcare product cares about.

Use WER. Do not use it alone.

Overall WER remains a valuable quality signal. The mistake is treating it as a complete medical evaluation.

Four views for medical speech recognition

1. Overall WER

Use overall WER to understand the transcript as a whole and to catch broad changes in acoustic or language performance. Keep the normalization and scorer identical across systems.

2. Medical-term WER

Medical-term WER, or M-WER, focuses the aligned comparison on the clinical terms in the reference. Omi's board uses a fixed lexicon covering diagnoses, medications, anatomy, procedures and other medical vocabulary. The denominator matters: the current board contains 2,872 reference medical terms.

M-WER does not declare every medical word equally important, and it depends on the term set. Its value is that a result cannot look medically strong solely by getting “the,” “and” and “patient” right.

3. Drug-name error

Drug names deserve a separate count because they are easy to confuse and unusually important downstream. Report both the percentage and the raw numerator/denominator—for example, zero errors across 442 mentions—so the evidence is visible.

4. Dosage precision, recall and F1

Dosage evaluation should align canonical number-and-unit events. Recall measures how many reference doses were recovered. Precision penalizes dose events introduced by the system. F1 combines both, which prevents an invented dose from disappearing behind otherwise strong recall.

A real example: the lowest WER did not win every clinical view

On Omi's current sealed board, Azure MAI-Transcribe-1.5 recorded the lowest overall WER at 5.97%, fractionally ahead of Omi at 5.99%. Omi recorded lower medical-term error and a higher dosage F1 on the same audio. That does not make one metric “correct” and another “wrong.” It shows that they answer different questions.

The top three medical-term results—Omi, ElevenLabs and Google—statistically tie under the current paired analysis. Omi therefore does not claim an outright accuracy win over systems that the benchmark cannot reliably separate. Use the complete board and methodology before drawing a product conclusion.

How to evaluate medical speech-to-text on your own audio

  1. Define the workflow. Dictation, consultations, patient calls and far-field rooms produce different audio and different error costs.
  2. Build a representative acceptance set. Include real specialties, accents, microphones, interruptions, abbreviations and medication language. Obtain the appropriate rights and handle PHI safely.
  3. Create a reviewed reference. A scorer cannot be more reliable than the transcript it treats as truth.
  4. Blind the system names. Review output before the evaluator knows which provider produced it.
  5. Freeze normalization. Decide how punctuation, casing, numbers and abbreviations are handled before comparing systems.
  6. Score overall and clinical errors separately. Add drug names and dosages rather than only publishing WER.
  7. Measure product behavior separately. Speaker attribution, latency, throughput, formatting and failures are not WER.
  8. Inspect examples. A single average can still hide a recurring failure mode in one specialty or environment.

How to read close leaderboard results

A difference in displayed percentages is not automatically a meaningful difference. The same clips should be scored for every system, and uncertainty should be estimated with a paired method. When intervals overlap or the paired comparison cannot separate systems, describe the result as a tie.

Then let product requirements decide: your audio, supported language, speaker behavior, privacy, retention, capacity and total price may matter more than a few hundredths of a benchmark point.

Where Omi fits

Omi builds medical speech-to-text as a hosted batch and live API and as an open self-deployable model. The hosted API includes medical vocabulary, speaker labels and timestamps on every plan. The Builder plan includes 25 pooled audio-hours each month without a card.

Start with the medical ASR buyer guide, inspect speaker diarization as a separate evaluation problem, or compare the current result with a specific provider on the comparison hub.

Measure the words that matter.

Inspect all 30 systems, then test the shortlist on your own clinical audio.