What “best medical ASR” should mean
Speech recognition leaderboards usually report word error rate (WER): the percentage of inserted, deleted or substituted words. That is useful, but a clinical transcript can have a respectable WER while still changing a drug name or dose.
We therefore rank systems with three separate views: overall WER, medical-term WER (M-WER), and dosage-event F1. Drug-name error is reported independently as well. This makes the trade-off visible instead of hiding it inside one average.
1. Medical terminology
Use encounters that contain diagnoses, anatomy, abbreviations and medications. A dedicated M-WER score tells you how often the system changes the vocabulary clinicians care about. Do not infer medical accuracy from a general podcast or meeting benchmark.
2. Drug names and dosage
Drug-name error should be visible as its own metric. Dosage testing should align canonical number-and-unit events and report precision as well as recall, so invented doses are penalised rather than disappearing from the score.
3. Your actual workflow
Prerecorded consultation transcription, live dictation and far-field meetings are different products. Ask whether the published result covers your audio shape, language and speaker count. Test your own audio before signing a larger contract.
4. Privacy and control
For healthcare, check where audio is processed, how long it is retained, whether customer data is used for training, and how a BAA or DPA is signed. If audio must remain inside your environment, an open model or supported private deployment may matter more than a hosted benchmark lead.
5. Total product cost
Compare the actual configuration you need. Speaker diarization, vocabulary, language support, long-audio jobs and compliance may be bundled, metered separately or available only on enterprise plans.
Shortlist two or three systems using the benchmark, then run a blinded test on your own representative audio. Score medical terms and dosage before reviewing style or punctuation.
Where Omi fits
Omi is designed for medical transcription and offers both a hosted API and an open edge model. The hosted API starts with 25 free audio-hours each month, includes a self-serve BAA, and is priced at $0.20 per audio-hour beyond the free allowance.
Compare Omi with a specific provider, inspect the complete medical speech-to-text benchmark, or choose the hosted API or open model.