NVIDIA Publishes Fine-Tuning Recipe for Dialect-Specific Speech Recognition with Nemotron
In short: NVIDIA released a detailed workflow for fine-tuning its Nemotron 3.5 ASR multilingual streaming model on Saudi Arabic dialects (Najdi and Hijazi). Using minimal data curation, replay mixing with FLEURS data, duration-based bucketing, and partial encoder unfreezing, they reduced word error rate on the target dialects from 55.05% to 29.96% while also slightly improving English and other Arabic performance. The post also covers decoding tweaks and speaker diarization extensions for multi-speaker transcription.
This summary was generated automatically by AI from NVIDIA Developer's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.
What changed?
- 1Fine-tuned on 133.7 hours of Najdi/Hijazi speech (82.5% of curated corpus) cut WER from 55.05% to 29.96% on target dialects
- 2English WER improved from 11.04% to 10.42% and Arabic WER from 12.67% to 11.41%, showing no catastrophic forgetting
- 3Replay mixing used 90% target dialect data, 7% English FLEURS, 3% Arabic FLEURS to retain prior language skills
- 4Unfreezing top 8 of 24 encoder layers (230.4M trainable params) balanced accuracy and training cost
- 5Larger attention context (13 lookahead frames) plus beam-8 MALSD decoding cut WER by 2.71 points, adding ~800ms latency suitable for batch (not realtime) use
- 6Nemotron 3 Diarization extends the pipeline to speaker-attributed transcription for up to 8 speakers
| Parameter | Before | Now |
|---|---|---|
| WER (Najdi+Hijazi test) | 55.05% | 29.96% |
| CER (Najdi+Hijazi test) | 31.63% | 12.18% |
| Full SADA WER | 58.84% | 35.61% |
| FLEURS English WER | 11.04% | 10.42% |
| FLEURS Arabic WER | 12.67% | 11.41% |
| Trainable parameters (top-8 layers unfrozen) | Not specified | 230.4M of 638M total |
| Language coverage | 40 language-locales (base model) | Same, plus specialized Najdi/Hijazi dialects |
Why it matters
This gives developers building multilingual voice systems a validated, reproducible method to specialize a streaming ASR model for an underrepresented dialect or accent without sacrificing performance on previously supported languages — a common challenge for contact-center and voice-agent deployments serving mixed-language populations.
What it means for AI agents and contact centers
Your company could apply the same recipe (minimal curation, replay mixing, partial encoder unfreezing, bucketing) to fine-tune Nemotron or similar NeMo-based ASR models for Lithuanian or regional Lithuanian accents, improving transcription accuracy for call analytics and voice agents without degrading support for other languages used in multilingual deployments.
🧪 Worth Testing
The recipe is concrete, reproducible, and open (NeMo framework, public notebook referenced), making it practical to test for improving STT accuracy in Lithuanian or other underrepresented languages/dialects relevant to your company's voice agent pipelines.
NVIDIA NeMo ASR fine-tuning recipe (replay mixing + partial encoder unfreezing + bucketing)· New
Sources
- NVIDIA DeveloperOfficialPrimary sourceOriginal article →„Fine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects, with a Path to Other Languages“1 Oct 2026, 08:00
- Published by source
- 1 Oct 2026, 08:00
- Found by our system
- 2 Oct 2026, 19:51
- Summary generated
- 2 Oct 2026, 20:07
This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy