AssemblyAI launches Universal-3.5 Pro with native code-switching and improved speaker diarization
In short: AssemblyAI released Universal-3.5 Pro, a new flagship async speech-to-text model priced at $0.21/hr. It natively transcribes code-switched conversations across 18 languages without separate configuration, introduces jointly modeled speaker diarization built directly into the transcript (rather than stitched from a separate system), and supports contextual prompting to prime the model with domain knowledge. Benchmarks show lower word error rate on code-switched audio and higher cpWER accuracy on diarization compared to several competitor models and AssemblyAI's own prior Universal-3 Pro.
This summary was generated automatically by AI from AssemblyAI's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.
What changed?
- 1Native code-switching transcription across 18 languages, no separate configuration needed
- 2New joint diarization approach that ties words to speakers directly rather than via timestamp alignment
- 3Measured using cpWER instead of DER, which the company says better reflects real transcript accuracy
- 4Contextual prompting lets users supply domain context (clinical notes, meeting agendas, product names) to improve accuracy
- 5Lower code-switching WER (7.69% avg vs 9.07% for Universal-3 Pro) and best cpWER (30.17 avg) among compared models
- 6Priced at $0.21/hr; available via API parameter universal-3-5-pro and as the new realtime default
| Parameter | Before | Now |
|---|---|---|
| Code-switching WER (avg, 5 language pairs) | 9.07% (Universal-3 Pro) | 7.69% |
| Languages supported at full accuracy | Not specified | 18 languages with native mid-sentence code-switching |
| Diarization accuracy (cpWER avg) | Not specified | 30.17% (lowest/best among compared models) |
| Price | Not specified | $0.21/hr |
Why it matters
For teams building contact-center analytics or ambient voice products, more accurate speaker attribution and native handling of mixed-language calls reduces downstream errors in QA, compliance and analytics pipelines, while contextual prompting can cut domain-specific transcription errors without retraining.
What it means for AI agents and contact centers
If your company runs multilingual or code-switched calls, or needs reliable speaker-level transcripts for call QA and analytics, this model is worth benchmarking against your current STT provider on real call recordings, checking both diarization accuracy (cpWER) and cost per hour.
🧪 Worth Testing
Improved diarization and code-switching accuracy at a stated price point could meaningfully improve call-center transcript quality and downstream analytics if validated on your own audio and languages.
Sources
- AssemblyAIOfficialPrimary sourceOriginal article →„Universal-3.5 Pro: native code switching, our most accurate speaker diarization yet, and expanded language support“
- Published by source
- —
- Found by our system
- 6 Oct 2026, 20:35
- Summary generated
- 6 Oct 2026, 20:36
This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy