Research & evaluation
Medical accuracy.
Measured in the open.
We evaluate speech models on clinical audio, with attention to medical terms, medications, dosages and response speed.
Batch transcription
Selected models from the September 2026 edition, ordered by medical word error.
| Model | Medical WER ↓ | WER ↓ | Dose / unit matches ↑ | Batch responsemedian / p95 |
|---|---|---|---|---|
| omi-medical-1 | 0.87% | 5.84% | 93.62% | 0.177 / 0.292 s |
| ElevenLabs Scribe v2 Medical | 1.08% | 5.87% | 86.17% | 0.400 / 0.675 s |
| Google Chirp 3 | 1.11% | 6.31% | 81.91% | 0.751 / 1.014 s |
| ElevenLabs Scribe v2 | 1.15% | 6.49% | 81.91% | 0.699 / 1.308 s |
| Microsoft MAI-Transcribe-2 | 1.18% | 5.64% | 85.11% | 0.434 / 0.873 s |
Response time runs from request start until the completed transcript is available. Timing uses short clips, one request at a time, from Frankfurt.
Realtime transcription
Accuracy uses confirmed segments. Speed measures confirmation after speech ends.
| Model | Medical WER ↓ | WER ↓ | Dose / unit matches ↑ | Confirmed after speech endsmedian / p95 |
|---|---|---|---|---|
| Speechmatics Linden-1 | 1.29% | 6.67% | 72.34% | 0.237 / 0.381 s |
| omi-medical-1 | 1.32% | 6.55% | 87.23% | 0.489 / 0.649 s |
| OpenAI GPT Live Transcribe | 1.32% | 6.83% | 82.98% | — Manual commit¹ |
| ElevenLabs Scribe v2 Realtime | 1.39% | 6.62% | 80.85% | 1.644 / 1.863 s |
| AssemblyAI Universal 3.6 Pro | 1.39% | 6.83% | 82.98% | 0.380 / 1.530 s |
Confirmation is timed after a shared speech-end boundary on short clips; only observed confirmations are included. ¹ OpenAI used a client commit, so no comparable automatic-confirmation time is shown. See full results and coverage.
Our open medical model
omi-medical-edge-1 is evaluated on the same English recordings.
Self-hosted inference. Hosted API response times do not apply to this model; speed depends on the deployment.
Read more about the dataset, scoring and timing methodology.
Read the methodology