One medical product, from API to edge.
Use the flagship model through the hosted API or deploy the open edge model on your own hardware.
Production accuracy through one API.
Built for consultations, dictation and long clinical recordings, with vocabulary, eight-language transcription and speaker labels on the same endpoint.
Tied for first of 29 systems by M-WER (the top two are statistically inseparable); the dosage result is significant against every competitor tested.
Everything around the transcript.
One model for clinical workflows across eight supported languages.
Medical vocabulary
Send encounter names, products and terminology with each request.
Eight languages
English, Spanish, Portuguese, French, German, Dutch, Arabic and Hindi.
Speakers and timing
Speaker-labelled segments and word timing where supported.
Long audio
Async jobs, polling, webhooks and model-build provenance.
Call the hosted API.
Short clips return inline. Long audio returns an async job.
Medical transcription inside your product.
Run omi-medical-edge-1 locally on Mac, CUDA or CPU. Audio can stay entirely inside your environment.
Device numbers are kept separate from the faster GPU-served leaderboard row. See the model card for runtime-specific accuracy and memory.
Own the complete runtime.
Open weights and a local runtime for products that cannot send audio away.
Open weights
Inspect, benchmark and adapt under CC-BY-4.0.
Fully local
No audio upload or hosted API dependency.
Three runtimes
Run on Mac, NVIDIA CUDA or CPU.
CLI and SDK
Start from the CLI or embed the runtime in your product.
Run it offline.
Install the runtime and transcribe without sending audio anywhere.
Choose your model.
Start with the API or bring the open model into your stack.