Skip to content
Sunday, 4 October 2026
Tenesys AI News
Subscribe

Voice AI

Speech-to-Text, Text-to-Speech, realtime models and voice AI agents.

Voice AI· Important· 🧪 Worth Testing

NVIDIA Publishes Fine-Tuning Recipe for Dialect-Specific Speech Recognition with Nemotron

NVIDIA released a detailed workflow for fine-tuning its Nemotron 3.5 ASR multilingual streaming model on Saudi Arabic dialects (Najdi and Hijazi). Using minimal data curation, replay mixing with FLEURS data, duration-based bucketing, and partial encoder unfreezing, they reduced word error rate on the target dialects from 55.05% to 29.96% while also slightly improving English and other Arabic performance. The post also covers decoding tweaks and speaker diarization extensions for multi-speaker transcription.

NVIDIA Developer
Voice AI· 🧪 Worth Testing

Suno launches 'Speech' feature combining AI voiceovers with background music

AI music generator Suno has launched a public beta feature called Speech, which generates synthetic voiceovers alongside AI-composed background music as a single track. Users can choose Simple mode (prompt-based) or Advanced mode (custom script with voice gender, style, and variation controls). The feature supports a maximum duration of about eight minutes and the background music can be toggled off for clean speech output.

The Verge