AI · Oct 5, 2026
Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the NuanceOne Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO
Oct 7, 2026, 5:45 AM · Hugging Face

NVIDIA's olympiad medals matter less as trophies than as a recipe it handed out, showing open models can be tuned into elite specialists by people outside a frontier lab.
Why it matters
NVIDIA says it fine-tuned its Nemotron 3 models into systems that reached gold-medal level at both the 2026 International Olympiad in Informatics and the 2026 International Mathematical Olympiad. At IOI, a specialist called Nemotron-3-Ultra-CC scored 535.4 out of 600, above the 361.12 gold threshold and the top human score of 498.27. At IMO, a system combining several checkpoints scored 30 of 42, one point above the gold line, with full credit on four of six problems.
The bigger news for builders: NVIDIA published the checkpoints, training datasets, a new 200-problem benchmark called Nemotron-IMO-Bench, papers and inference code on Hugging Face and in its NeMo-Skills repository.
From the desk
We have spent weeks covering math results from closed labs that arrive as announcements about models nobody outside can use. This is the other model of progress, and we like it better. The claim here is not that NVIDIA has a secret system. It's that a fairly standard approach, supervised fine-tuning, reinforcement learning where useful, and a loop that generates, checks and revises answers, can turn an open base model into a specialist. Then it gave away the pieces so others can test that claim.
The details support a sober read. For the IOI work, NVIDIA curated 22,000 programming problems. Its smaller Nano model went from 130 points on the 2025 test before post-training to 468 with its GenCorrect feedback loop, crossing that year's gold line. For the IMO, the training set held 414,890 filtered examples across 15,818 proof problems, and the whole system worked in natural language with no formal prover, external tools or internet. NVIDIA's own conclusion is the one we'd repeat: the medals came from designing the model, data and search loop together, not from fine-tuning or brute-force sampling alone.
Now the caveats. This is a vendor writing about its own models. The IOI run was live and under contestant conditions, which helps, but it was unofficial and unsupervised and does not appear in the official rankings. The IMO proofs were graded by official graders, which is stronger evidence. NVIDIA also calls the compute "substantial" without a public figure in this post, and that number matters for whether a university lab could actually reproduce it.
The longer trajectory is worth naming. Olympiads were built to find and encourage talented students. As they become standard AI benchmarks, the risk is that the contest becomes a marketing stage and the kids become the baseline. We'd rather the field use these results to build better tutoring and verification tools than to declare human competition solved.
Context
The IOI rewards code that passes hidden tests under strict limits; the IMO rewards rigorous written proofs. Doing well at both with one model family is NVIDIA's argument that Nemotron is a flexible foundation, not a one-off.
Who feels it
- Researchers and open-source developers
- Released datasets, checkpoints and inference code make it possible to reproduce, audit or extend the work rather than take it on faith.
- Enterprises
- The recipe suggests specialized open models can be built for hard domains without training a new foundation model.
- Educators and contest organizers
- AI systems now clearing gold thresholds raise fresh questions about how olympiads are used and protected.
What to watch
- Independent reproductions using the released Nemotron data and code
- Disclosure of the compute behind the competition runs
- Results from other teams on the new Nemotron-IMO-Bench
- Whether olympiad organizers set formal rules for AI participation