Skip to content
Sunday, 4 October 2026
Tenesys AI News
Subscribe
Voice AI· Important· 🧪 Worth Testing

NVIDIA Publishes Fine-Tuning Recipe for Dialect-Specific Speech Recognition with Nemotron

In short: NVIDIA released a detailed workflow for fine-tuning its Nemotron 3.5 ASR multilingual streaming model on Saudi Arabic dialects (Najdi and Hijazi). Using minimal data curation, replay mixing with FLEURS data, duration-based bucketing, and partial encoder unfreezing, they reduced word error rate on the target dialects from 55.05% to 29.96% while also slightly improving English and other Arabic performance. The post also covers decoding tweaks and speaker diarization extensions for multi-speaker transcription.

Source: NVIDIA DeveloperNVIDIANemotron 3.5 ASROriginal article ↗

This summary was generated automatically by AI from NVIDIA Developer's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.

What changed?

  • 1Fine-tuned on 133.7 hours of Najdi/Hijazi speech (82.5% of curated corpus) cut WER from 55.05% to 29.96% on target dialects
  • 2English WER improved from 11.04% to 10.42% and Arabic WER from 12.67% to 11.41%, showing no catastrophic forgetting
  • 3Replay mixing used 90% target dialect data, 7% English FLEURS, 3% Arabic FLEURS to retain prior language skills
  • 4Unfreezing top 8 of 24 encoder layers (230.4M trainable params) balanced accuracy and training cost
  • 5Larger attention context (13 lookahead frames) plus beam-8 MALSD decoding cut WER by 2.71 points, adding ~800ms latency suitable for batch (not realtime) use
  • 6Nemotron 3 Diarization extends the pipeline to speaker-attributed transcription for up to 8 speakers
Nemotron 3.5 ASR
ParameterBeforeNow
WER (Najdi+Hijazi test)55.05%29.96%
CER (Najdi+Hijazi test)31.63%12.18%
Full SADA WER58.84%35.61%
FLEURS English WER11.04%10.42%
FLEURS Arabic WER12.67%11.41%
Trainable parameters (top-8 layers unfrozen)Not specified230.4M of 638M total
Language coverage40 language-locales (base model)Same, plus specialized Najdi/Hijazi dialects

Why it matters

This gives developers building multilingual voice systems a validated, reproducible method to specialize a streaming ASR model for an underrepresented dialect or accent without sacrificing performance on previously supported languages — a common challenge for contact-center and voice-agent deployments serving mixed-language populations.

What it means for AI agents and contact centers

Your company could apply the same recipe (minimal curation, replay mixing, partial encoder unfreezing, bucketing) to fine-tune Nemotron or similar NeMo-based ASR models for Lithuanian or regional Lithuanian accents, improving transcription accuracy for call analytics and voice agents without degrading support for other languages used in multilingual deployments.

🧪 Worth Testing

The recipe is concrete, reproducible, and open (NeMo framework, public notebook referenced), making it practical to test for improving STT accuracy in Lithuanian or other underrepresented languages/dialects relevant to your company's voice agent pipelines.

NVIDIA NeMo ASR fine-tuning recipe (replay mixing + partial encoder unfreezing + bucketing)· New

Sources

  • NVIDIA DeveloperOfficialPrimary source
    „Fine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects, with a Path to Other Languages“
    1 Oct 2026, 08:00
    Original article →
Published by source
1 Oct 2026, 08:00
Found by our system
2 Oct 2026, 19:51
Summary generated
2 Oct 2026, 20:07

This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy