NVIDIA fine-tunes Nemotron to gold-level IOI 2026 (535.4/600) and IMO 2026 (30/42) results, releasing weights and training data
One open model family clears the gold bar in both programming and math olympiads, with the data and evaluation pipelines released too.
// Key points
- IOI 2026: Nemotron-3-Ultra-CC (550B total, 55B active) with SFT and GenCorrect scored 535.4/600, against a 361.12 gold threshold and a top human score of 498.27; the run was unofficial and not ranked.
- IMO 2026: a system built on Nemotron 3 Ultra scored 30/42, above the gold threshold of 29, with proofs graded by official IMO graders.
- Released: Ultra-CC NVFP4 weights, IMO SFT and RL checkpoints, both training datasets (including 22,000 programming problems) and the 200-problem Nemotron-IMO-Bench; inference and evaluation pipelines are in the NeMo-Skills repo.
- The post doesn't state a license.
Builder's takeMost teams won't run a 550B model; the part worth taking is the method: on IOI 2025 the Nano model (30B total, 3B active) went from 280 after SFT to 468 with test-time strategies like GenCorrect. If you build coding products, try a generate-then-self-correct layer on your own eval set first.
// Background · from #Open models
Full timeline →- Oct 9 LightOn open-sources LightOnOCR-3 in 0.8B, 1B and 4B sizes under Apache 2.0, scoring up to 86.3 on olmOCR-Bench
- Oct 9 Liquid AI releases open d1 decision models: d1-3B answers in one forward pass, 8 ms per question on an RTX 4090
- Oct 7 Mistral Large 4 enters public preview: 1T total, 49B active, $1.36 per million input tokens, weights by month end