Fine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects, with a Path to Other Languages
Sep 30, 2026, 10:00 PM · NVIDIA Developer

NVIDIA engineers cut Najdi/Hijazi ASR word error from 55% to ~30% by fine-tuning Nemotron 3.5 — without tanking other Arabic dialects or English.
Why it matters
A NeMo-framework cookbook shows how to adapt multilingual Nemotron 3.5 ASR (40 language-locales) to Saudi Najdi and Hijazi with 133.7 hours of target speech, weighted replay mixing with FLEURS, duration bucketing, and partial encoder unfreezing. Target-split WER fell from 55.05% to 29.96%, while English improved slightly to 10.42% WER and other Arabic dialects held.
Unfreezing all 24 encoder layers hit the best accuracy at 230.4M trainable parameters; freezing trades accuracy for cheaper fine-tunes. Decoding tweaks (13-frame lookahead, beam-8 MALSD) cut another 2.71 absolute WER at about 800 ms extra latency for batch jobs. Nemotron 3 Diarization extends the path to speaker-attributed transcripts.
From the desk
We’re cheering dialect work that doesn’t erase the multilingual base.
Generic “Arabic” models fail real users in Najdi and Hijazi speech; a documented recipe with replay mixing is how useful speech AI localizes. Holding other dialects steady is the hard part — catastrophic forgetting is the usual tax.
I’m watching whether ministries, telcos, and contact centers pick up the recipe, and how diarization behaves in noisy multi-speaker Saudi settings. Open playbooks beat closed regional models nobody can adapt.
We’ll treat this as a template for other under-served locales. Fine-tuning with care is still the fastest path to speech equity.
Context
NVIDIA Technical Blog, September 30, 2026, Nemotron Saudi dialect fine-tuning.
Who feels it
- Speech engineers
- Reuse the NeMo weighted-replay and encoder-unfreeze pattern for other dialects.
- Enterprises in the Gulf
- 30% WER isn’t human parity — budget human review for critical transcripts.
- Language communities
- Documented dialect adaptation is a lever for inclusion beyond MSA.
What to watch
- Public release of the 133.7-hour training recipe artifacts
- Production latency tradeoffs for the lookahead/beam decoding gains
- Ports of the workflow to other Arabic and non-Arabic dialects
Companies: NVIDIA