Benchmark update · 29 August 2026

Three new transcription models, tested on medical audio.

Azure MAI-Transcribe-1.5, OpenAI gpt-transcribe and Google Gemini 3.5 Transcribe joined the board. We ran every system through the same 1,513 clinical clips and the same scorer.

Metric
Omi
Azure
OpenAI
Google
Medical-first board rank
1st
7th
9th
17th
Word error ↓
5.99%
5.97%
7.26%
7.89%
Medical-term error ↓
0.94%
1.43%
1.43%
2.44%
Drug-name errors / 442
0
4
6
6
Dosage F1 ↑
97.7
83.3
80.7
75.8

“Google” here means the tested Gemini 3.5 Transcribe row. “Azure” means the tested MAI-Transcribe-1.5 preview row. Board rank is medical-first and includes all 30 systems.

The overall-WER winner did not win the medical task

Azure recorded the best plain word accuracy on the board at 5.97% WER, fractionally ahead of Omi's 5.99%. But Azure ranked 14th of 30 on dosage, while Omi recovered 86 of 89 aligned dose events with no invented doses.

That is the core lesson: getting the words right and getting the medicine right are not the same optimisation target. A model can improve common conversational words without improving the tokens that drive a clinical record.

Why we publish four views

Overall WER remains useful. It describes the transcript as a whole and catches broad regressions. Medical-term WER narrows the score to clinical vocabulary. Drug-name error isolates medication names. Dosage F1 aligns canonical number-and-unit events so both omissions and inventions matter.

No single metric decides production suitability. Together, the four views make the trade-off visible enough for a builder to decide what to test next.

What this result does not prove

The board is clinical English, not every language, specialty, room or microphone. It does not measure latency, throughput, formatting preference or your security architecture. It narrows a shortlist. Your own representative acceptance set still makes the deployment decision.

Reproduce the reasoning

Open the full benchmark and methodology, compare Omi with Azure or Omi with OpenAI, then run the same product audio through the Omi playground.