Research #Open models Oct 9, 2026 · Originally published Oct 7

NVIDIA fine-tunes Nemotron to gold-level IOI 2026 (535.4/600) and IMO 2026 (30/42) results, releasing weights and training data

One open model family clears the gold bar in both programming and math olympiads, with the data and evaluation pipelines released too.

Primary source NVIDIA on the Hugging Face blog · huggingface.co Read the original ↗

// Key points

  • IOI 2026: Nemotron-3-Ultra-CC (550B total, 55B active) with SFT and GenCorrect scored 535.4/600, against a 361.12 gold threshold and a top human score of 498.27; the run was unofficial and not ranked.
  • IMO 2026: a system built on Nemotron 3 Ultra scored 30/42, above the gold threshold of 29, with proofs graded by official IMO graders.
  • Released: Ultra-CC NVFP4 weights, IMO SFT and RL checkpoints, both training datasets (including 22,000 programming problems) and the 200-problem Nemotron-IMO-Bench; inference and evaluation pipelines are in the NeMo-Skills repo.
  • The post doesn't state a license.

Builder's takeMost teams won't run a 550B model; the part worth taking is the method: on IOI 2025 the Nano model (30B total, 3B active) went from 280 after SFT to 468 with test-time strategies like GenCorrect. If you build coding products, try a generate-then-self-correct layer on your own eval set first.

// Background · from #Open models

Full timeline →
  1. Oct 9 LightOn open-sources LightOnOCR-3 in 0.8B, 1B and 4B sizes under Apache 2.0, scoring up to 86.3 on olmOCR-Bench
  2. Oct 9 Liquid AI releases open d1 decision models: d1-3B answers in one forward pass, 8 ms per question on an RTX 4090
  3. Oct 7 Mistral Large 4 enters public preview: 1T total, 49B active, $1.36 per million input tokens, weights by month end