Deepgram Extends Medical Speech Recognition to 10 Languages
In short: Deepgram has expanded its medical-specialized speech model, Nova-3 Medical, from English-only to 10 languages including Spanish, German, Hindi, Japanese and Russian, with automatic code switching in both batch and streaming modes. On a 115-hour curated medical audio test set, the multilingual medical model cut word error rate by 25-30% and entity error rate (drugs, doses, conditions) by 37-40% compared to Deepgram's general-purpose multilingual model. It also beat competing models from ElevenLabs and Microsoft on average word error rate across all 10 languages.
This summary was generated automatically by AI from Deepgram's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.
What changed?
- 1Nova-3 Medical now supports 10 languages: Dutch, English, French, German, Hindi, Italian, Japanese, Portuguese, Russian, Spanish
- 2Automatic code switching between supported languages in both batch and streaming
- 3Batch WER dropped from 6.62% to 4.96% (25% reduction) vs general multilingual model
- 4Streaming WER dropped from 8.32% to 5.86% (30% reduction)
- 5Batch entity error rate (drugs, doses, conditions) cut from 14.46% to 9.11% (37% reduction)
- 6Dose recognition errors reduced by 64% in batch and 81% in streaming
- 7Lowest average batch WER (4.96%) among models tested, ahead of ElevenLabs Scribe v2 Medical (5.49%) and Microsoft MAI 1.5 (7.80%)
- 8Available via single API parameter: model=nova-3-medical&language=multi, plus self-hosted deployment option
| Parameter | Before | Now |
|---|---|---|
| Languages supported | English only (Nova-3 Medical) | 10 languages: Dutch, English, French, German, Hindi, Italian, Japanese, Portuguese, Russian, Spanish |
| Batch Word Error Rate (vs general multilingual model) | 6.62% | 4.96% |
| Streaming Word Error Rate (vs general multilingual model) | 8.32% | 5.86% |
| Batch Entity Error Rate | 14.46% | 9.11% |
| Streaming Entity Error Rate | 16.31% | 9.78% |
| Code switching | Not specified | Automatic code switching between supported languages in both batch and streaming |
Why it matters
Accurate transcription of drug names, doses and conditions is critical for any healthcare voice agent or ambient documentation tool; this release shows that domain-specialized models still substantially outperform general multilingual STT even when expanded across many languages, which matters for intake bots, dictation tools and clinical note generation across language barriers.
What it means for AI agents and contact centers
While this is a healthcare-specific model, the benchmark methodology (comparing specialized vs general multilingual STT on word and entity error rate, in both batch and streaming) is directly useful for evaluating speech-to-text vendors for any non-English voice agent deployment, including assessing how well a provider handles numbers, names and domain vocabulary in smaller languages.
🧪 Worth Testing
Testing Nova-3 Medical Multilingual's streaming accuracy and code-switching behavior could inform STT vendor selection for multilingual voice agents, particularly where domain vocabulary (names, numbers, technical terms) accuracy matters more than generic transcription quality.
Sources
- DeepgramOfficialPrimary sourceOriginal article →„Article · · Announcements Introducing Nova-3 Medical Multilingual: Medical Speech Recognition in 10 Languages“7 Oct 2026, 03:00
- Published by source
- 7 Oct 2026, 03:00
- Found by our system
- 7 Oct 2026, 08:36
- Summary generated
- 7 Oct 2026, 08:37
This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy