Skip to content
Sunday, 4 October 2026
Tenesys AI News
Subscribe

Voice AI

Speech-to-Text, Text-to-Speech, realtime models and voice AI agents.

Voice AI· Important· 🧪 Worth Testing

NVIDIA Publishes Fine-Tuning Recipe for Dialect-Specific Speech Recognition with Nemotron

NVIDIA released a detailed workflow for fine-tuning its Nemotron 3.5 ASR multilingual streaming model on Saudi Arabic dialects (Najdi and Hijazi). Using minimal data curation, replay mixing with FLEURS data, duration-based bucketing, and partial encoder unfreezing, they reduced word error rate on the target dialects from 55.05% to 29.96% while also slightly improving English and other Arabic performance. The post also covers decoding tweaks and speaker diarization extensions for multi-speaker transcription.

NVIDIA Developer