Open benchmark

Medical speech-to-text, tested where it matters.

The same 1,513 clinical clips, scorer and rules across 29 systems. Medical terms and dosage are scored separately.

Updated 1 August 2026 · v6 · 29 systems
5.99%Word errorFull transcript
0.94%Medical-term errorTied #1 of 29
0.00%Drug-name error0 of 442
97.7%Dosage F1#1 of 30
Clinical accuracy

Three views of the medical record.

Overall wording, medical terminology and dosage events answer different clinical questions. We score all three.

See all 29 systems ↓

Overall word error

Closest major APIs · lower is better
Azure5.97
Omi5.99
AWS6.12
ElevenLabs6.17
Google6.30
WER measures every word, medical or not.

Medical-term error

Top full-coverage systems · lower is better
Omi0.94
ElevenLabs0.97
Google1.11
Speechmatics1.36
Deepgram2.19
Omi and ElevenLabs are statistically tied here.

Dosage F1

Major hosted APIs · higher is better
Omi97.7
Deepgram86.8
ElevenLabs85.4
Azure83.3
Google80.7
Omi: 96.6% recall · 98.9% precision · 0 extra doses.
Full leaderboard

29 systems, one board.

Select any accuracy column to reorder the board. Medical-term error is the default because the benchmark ranks medical accuracy first.

#SystemType
1omi-medical-1 (live API)5.99%0.94%0.00%ours · API / self-host
2ElevenLabs Scribe v26.17%0.97%0.00%API
3Google Chirp 36.30%1.11%1.13%API
4VibeVoice-ASR 9B7.47%1.29%1.36%open
5Gemini 3.1 Pro Preview6.91%1.34%0.23%API
6Speechmatics Enhanced Medical6.42%1.36%1.81%API · medical
7Azure MAI-Transcribe-1.55.97%1.43%0.90%API
8AssemblyAI Universal-3.5 Pro (Medical)6.74%1.43%1.13%API · medical
9OpenAI gpt-transcribe7.26%1.43%1.36%API
10Soniox STT Async v46.63%1.60%3.39%API
11Gemini 3.5 Flash7.70%1.73%0.45%API
12Gladia Solaria-37.94%1.98%4.98%API
13OpenAI Whisper-17.89%2.05%3.17%API
14omi-medical-edge-16.64%2.16%5.43%ours · open · GPU Parakeet + curated dictionary; on-device unvalidated
15Deepgram Nova-3 Medical6.98%2.19%2.26%API · medical
16Reson8 Prerecorded6.33%2.26%6.56%API
17Voxtral Mini Transcribe v27.82%2.47%5.66%open
18Whisper Large v3 Turbo12.17%2.54%6.11%open
19Qwen3-ASR 1.7B7.01%2.79%6.11%open
20Qwen3-ASR 0.6B7.20%3.13%7.92%open
21OpenAI GPT-4o Mini Transcribe10.02%3.24%3.39%API
22Parakeet TDT 0.6B v311.93%3.66%9.50%open
23Parakeet TDT 0.6B v212.48%3.80%8.60%open
24Amazon Transcribe Medical6.12%4.28%1.81%API · medical
25NVIDIA Canary 1B Flash12.86%4.32%13.12%open
26Corti Transcripts9.60%5.05%11.31%API · medical
27Azure gpt-4o-transcribe12.52%5.57%6.33%API
28GCP medical_conversation17.70%5.57%12.44%API · medical
29Google MedASR32.21%10.83%14.48%open · medical
29 full-coverage systems · dosage “—” means that system is not in the published dosage summary · highlighted rows are Omi.
Method

Simple enough to audit.

Open any step for the high-level run convention.

1Same clinical audioEvery full-coverage system receives the same 1,513 clips from 57 consultations.

Every system receives the same files, not a provider-specific subset. Full coverage means a transcript for all 1,513 clips. Together they contain 7.2 hours of clinical English and 63.8k reference words.

2Same scorerOne frozen normalization and alignment pipeline scores all transcripts.

Raw transcripts pass through one versioned normalizer and alignment pipeline. Casing, punctuation, medical terms and drug names are handled consistently, and the public scorer records the exact run convention.

3Medical-first rankingMedical terms, drug names and dosage events are measured independently from overall WER.

The board ranks M-WER first, then drug-name error and overall WER. Dosage is a separate event-level board with both precision and recall, so invented doses are penalised. Lower is better for WER, M-WER and drug error; higher is better for dosage F1.

Dataset and scoring definitions

The full-coverage board contains 1,513 clips, 7.2 hours and 63.8k reference words in clinical English. It includes roughly 2,872 medical reference tokens, including 442 drug-name occurrences.

WER measures errors across the whole transcript. M-WER measures errors on the medical lexicon. Drug M-WER isolates drug names. The separate dosage board aligns canonicalised number-and-unit events and reports recall, precision and F1.

The dosage set contains 89 reference events across 63 clips. Omi recovered 86, added no extra doses and missed two after alignment: 96.6% recall, 98.9% precision and 97.7% F1.

Coverage and run convention

“Full coverage” means a system returned a scoreable transcript for every clip in the main board. API systems were called through their available product interfaces; open systems were run under the published harness. The saved model or endpoint name, configuration and scorer version form part of the result.

The primary ranking is M-WER → Drug M-WER → WER. The interactive table can also be ordered by dosage F1, but missing dosage results remain at the bottom rather than being treated as zero.

Interpretation and limitations
  • The 0.94% versus 0.97% M-WER difference is about one medical token and is not statistically significant. The two leading systems are a statistical tie on that metric.
  • Omi's dosage F1 is significant against every tested competitor under paired cluster bootstrap. Dosage claims always include the underlying 86/89 count and 89-event sample size.
  • A result describes the tested model, endpoint and configuration—not every product or deployment a provider offers.
  • This board measures prerecorded clinical English. It is not evidence for realtime latency, speaker diarization or multilingual quality.
Reproduce and cite

The public scorer and run instructions are in medical-STT-eval on GitHub. The frozen normalizer, alignment rules and public board generator keep result updates reviewable. Full reruns use the same manifest and row identifiers.

Omi Health. (2026). Medical Speech-to-Text Benchmark (v6): 29 Systems Ranked. https://omi.health/benchmark
Version history

Changelog

  • v6 — August 2026: 29 full-coverage systems, medical-first ranking and the 30-system dosage-event board with precision and F1.
  • v5 — July 2026: the 1,513-clip board and frozen v3 scorer became the publication baseline.
  • Earlier revisions: retained as historical experiment records; current public claims resolve to v6.

Test Omi on your audio.

25 audio-hours every month. No card required.