NVIDIA Publishes Fine-Tuning Recipe for Dialect-Specific Speech Recognition with Nemotron
| Parameter | Before | Now |
|---|---|---|
| WER (Najdi+Hijazi test) | 55.05% | 29.96% |
| CER (Najdi+Hijazi test) | 31.63% | 12.18% |
| Full SADA WER | 58.84% | 35.61% |
| FLEURS English WER | 11.04% | 10.42% |
| FLEURS Arabic WER | 12.67% | 11.41% |
| Trainable parameters (top-8 layers unfrozen) | Not specified | 230.4M of 638M total |
| Language coverage | 40 language-locales (base model) | Same, plus specialized Najdi/Hijazi dialects |