NVIDIA reports Nemotron 3 Ultra-based systems reached gold-level scores at IMO 2026 (30/42) and an unofficial IOI 2026 run (535.4/600), and releases models, datasets and a benchmark
NVIDIA published results and released artifacts: the Nemotron Labs IMO 2026 collection (SFT and RL checkpoints, two training datasets, Nemotron-IMO-Bench with 200 olympiad problems), the Nemotron-3-Ultra-CC model on Hugging Face, and IMO/IOI inference pipelines, prompts and submitted proofs in NeMo-Skills, along with IMO and IOI papers. Reported results: IMO 2026 30/42 (gold threshold 29), graded by official IMO graders, using a natural-language generate-verify-refine system with no tools or internet; IOI 2026 535.4/600 (gold threshold 361.12, top human 498.27) in a live prospective run under contestant constraints, but unofficial and not in the official ranking.
Entities: NVIDIA, Nemotron 3 Ultra, Nemotron-3-Ultra-CC, Nemotron-3-Nano-CC, GenCorrect, Nemotron-IMO-Bench
1 primary
What happened
NVIDIA says systems built on its Nemotron 3 Ultra model scored 30/42 at IMO 2026 (gold threshold 29), marked by official IMO graders, using a natural-language generate-verify-refine loop with no tools or internet. It also reports 535.4/600 on the IOI 2026 problems (gold threshold 361.12, top human 498.27) in a live run under contestant constraints, but this was unofficial and is not in the official ranking. Alongside the results it released the Nemotron Labs IMO 2026 collection (SFT and RL checkpoints, two training datasets), Nemotron-IMO-Bench (200 olympiad problems), the Nemotron-3-Ultra-CC model on Hugging Face, and inference pipelines, prompts and submitted proofs in NeMo-Skills, plus IMO and IOI papers.
Why it matters
The main practical value is the open release. Researchers and developers can inspect, reproduce and build on a competition-grade reasoning recipe, and the new benchmark gives others a shared test. Enterprises should not read it as proof of general reliability, since olympiad problems are narrow, well-defined and heavily studied. Competitors face a higher bar for open models, but gold-level IMO performance by AI is no longer new, so the competitive shift is modest.
What is noise
The IOI figure is self-run and unofficial, so comparing it to top human scores is suggestive rather than settled. The framing of Nemotron as a "strong, adaptable foundation" that turns into "world-class specialists" is marketing, and the results do not show the recipe transfers beyond contests. The IMO margin is thin (30 against a 29 threshold), and the coverage gives no cost, compute or number of attempts per problem.
Watch next
- 01Independent reproduction of the IMO and IOI results from the released checkpoints and NeMo-Skills pipelines, including compute cost and number of samples per problem.
- 02Whether other labs adopt Nemotron-IMO-Bench and what scores open and closed models post on it over the next few months.
- 03Evidence that the fine-tuning recipe works on non-competition domains such as enterprise coding, science or formal verification, and download and fork activity for the released checkpoints on Hugging Face.
Coverage
1 storyMore capability signals
Full feed →- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- OpenAI discloses sandbox-escape and credential-leak incidents, confirms pause on tool-use for its most capable models26 Sept 202680
- OpenAI launches GPT-6 Sol and Luna at half the token price of GPT-5.6, with roughly flat intelligence scores per independent analysis22 Sept 202680
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680