explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • The one caveat to read first
  • The four-part specialization recipe
  • IOI: from general coding ability to 535.4
  • IMO: teaching Nemotron to prove, check, and revise
  • Fine-tuning and test-time compute work together
  • What is open
  • What this means if you build or fine-tune models
  • What to do this week
  • Related reading
← Back to blog

explainx / blog

NVIDIA Nemotron Hits Gold-Level at IOI and IMO 2026: The Fine-Tuning Recipe and What Is Open

NVIDIA, Nemotron, Fine-Tuning, Competitive Programming, Mathematics

NVIDIA fine-tuned Nemotron 3 to gold-level results at IOI 2026 (535.4/600) and IMO 2026 (30/42). The SFT, RL and GenCorrect recipe, plus what is open.

Oct 7, 2026·11 min read·Yash Thakker
add explainx.ai
go deep
NVIDIA Nemotron Hits Gold-Level at IOI and IMO 2026: The Fine-Tuning Recipe and What Is Open

NVIDIA says one model family reached gold-medal level at two very different contests in the same year. At IOI 2026, a fine-tuned Nemotron-3-Ultra-CC system scored 535.4 out of 600. At IMO 2026, a system built on Nemotron 3 Ultra scored 30 out of 42. NVIDIA published the details on October 7, 2026 in a Hugging Face post and two arXiv papers.

The headline is not only the scores. NVIDIA presents a reusable four-step recipe and releases a large set of files. This post separates what NVIDIA measured, what it published, and what is still unconfirmed.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR

table · 3 cols
ItemIOI 2026 (coding)IMO 2026 (math proofs)
Score535.4 / 60030 / 42
Gold threshold361.1229
Top human score498.27Not stated in the source
ModelNemotron-3-Ultra-CC (550B total, 55B active)Nemotron 3 Ultra: general, SFT, and RL checkpoints
MethodSFT plus GenCorrect test-time loopGenerate-verify-refine search plus a high-compute selection stage
GradingUnofficial; not in the official IOI rankingGraded by official IMO graders
PaperarXiv:2609.02849arXiv:2609.10712

The one caveat to read first

The IOI number is not an official IOI result. NVIDIA writes that the run came "under the same time, internet-access, and submission constraints as human contestants," and then adds that it "was an unofficial, unsupervised benchmark and was not included in the official IOI ranking."

So "gold-level" means the score is above the gold threshold for the problem set. It does not mean the International Olympiad in Informatics awarded a medal. The IMO case is different in one respect. NVIDIA says the submitted proofs were graded by official IMO graders. We have not seen a separate statement from the IMO about this entry, so treat the grading claim as NVIDIA's report.

The four-part specialization recipe

NVIDIA frames the work as a test of whether Nemotron is "easy to fine-tune." Its answer is a recipe with four parts:

  1. Start with a strong Nemotron base model.
  2. Curate domain-specific problems and high-quality reasoning traces.
  3. Apply standard post-training methods such as supervised fine-tuning (SFT) and, where useful, reinforcement learning (RL).
  4. Pair the specialist model with an inference loop that generates, evaluates, and improves candidate answers.

The authors write that "the underlying approach is familiar and reproducible" and that they "did not need to build a new foundation model for every challenge." The base model matters. For background on the model, see our Nemotron 3 Ultra 550B guide.

The fourth step is the one many teams skip. A fine-tuned model alone did not produce the medals. NVIDIA says they "were not produced by fine-tuning alone, and they were not produced by brute-force sampling alone." They came from co-designing the model, the data, and the inference loop.

IOI: from general coding ability to 535.4

The data and the two models

For competitive programming, NVIDIA curated 22,000 problems and generated synthetic reasoning traces. It trained two specialists:

  • Nemotron-3-Nano-CC: 30 billion total parameters, 3 billion active. It received SFT and RL.
  • Nemotron-3-Ultra-CC: 550 billion total parameters, 55 billion active. It received SFT only.

The "CC" suffix refers to competitive coding. The model page for the Ultra-CC checkpoint lists a context length of up to 262,144 tokens and an NVFP4 format.

The IOI 2025 ablation

NVIDIA used the 2025 contest to measure what each step adds. The numbers below come from the blog post and the arXiv abstract.

table · 2 cols
Configuration (IOI 2025)Points
Nano before post-training130
Nano after SFT280
Nano after RL291
Nano with GenCorrect468 (gold threshold: 438.3)
Ultra-CC with GenCorrect502

Two patterns stand out. First, SFT delivered most of the Nano gain: 130 to 280 points. RL added a smaller but consistent step to 291. Second, test-time compute added the largest single jump: 291 to 468.

What GenCorrect does

GenCorrect is described as an "iterative generate-evaluate-refine strategy." The arXiv abstract calls it "a feedback-driven test-time compute strategy that iteratively generates, evaluates, and refines diverse solutions." NVIDIA says it "turned the gains from fine-tuning into larger improvements over multiple feedback rounds."

In plain terms, the system writes many candidate programs, tests them, reads the feedback, and revises. A better fine-tuned model writes better first drafts and better revisions, so the loop pays off more.

Why Ultra needed only SFT

The post makes a scale-dependent point. For Nano, RL helped a little. For the stronger Ultra model, "one SFT epoch was enough to outperform the fully post-trained Nano model across IOI, ICPC, and LiveCodeBench Pro." That finding guided the competition-specific Ultra-CC system used at IOI 2026.

For builders, this is a cost signal. If your base model is strong, a single SFT pass on clean reasoning traces may beat a heavier pipeline on a smaller model.

The IOI 2026 result

The competition-specific Ultra-CC system scored 535.4 out of 600. The thresholds were 361.12 for gold and 498.27 for the top human score. The arXiv abstract adds a bold claim: "To our knowledge, this is the first AI system to outscore the highest-scoring human contestant on an IOI problem set."

That sentence is the authors' own claim. We did not check it against other labs' results. Remember the unofficial status above.

IMO: teaching Nemotron to prove, check, and revise

The training data

The math project started from Nemotron 3 Ultra and trained two specialists, one with SFT and one with RL.

table · 2 cols
StageSize reported
SFT corpus414,890 quality-filtered examples
Unique proof problems in SFT15,818
RL problems9,597 proof problems chosen near the model's capability frontier

The SFT data did more than teach final answers. According to NVIDIA, it covered "proof generation, refinement, verification, and meta-verification," so the model learned to build arguments, find gaps, respond to critiques, and judge whether a proof was complete.

The system at test time

NVIDIA describes the pipeline this way. For each problem, the models generated candidate proofs, scored them, produced critiques, and refined the most promising attempts. A separate high-compute stage then selected the final submission.

Three checkpoints powered the search: the general-availability Nemotron 3 Ultra and the two post-trained specialists. In development experiments, "both post-trained checkpoints outperformed the general-availability model." The SFT checkpoint was strongest in the first search round. The RL checkpoint gave the best overall single-checkpoint result. Because the strengths were complementary, the final system used both plus the general model.

The whole system worked in natural language. NVIDIA says it used "no formal prover, external tools, or internet access."

The IMO 2026 result

The system scored 30 out of 42, with full credit on four of the six problems. The official gold threshold was 29. That is a one-point margin, so a single grading difference could change the label. NVIDIA reports the margin as above the threshold, and we do the same.

A related note on proof quality: natural-language proofs are not machine-checked. If you want the contrast with formal verification, read our piece on the Lean 4 formalization cost collapse.

Fine-tuning and test-time compute work together

The most useful section of the NVIDIA post is the shortest. It says better specialization "gives the inference system better candidates, better critics, and better refinements." Two findings support this:

  • At IOI, GenCorrect converted fine-tuning gains into larger gains across several feedback rounds.
  • At IMO, mixing complementary SFT and RL checkpoints was "more valuable than simply drawing more samples from one checkpoint."

If you build agent systems, the lesson is practical. Do not pick between a better model and a better loop. Measure both. A weak critic limits a strong generator, and a weak generator wastes a strong critic.

For a wider view of how these tests fit into model evaluation, see our AI benchmarks guide.

What is open

NVIDIA lists these resources. We checked the pages that were reachable on October 7, 2026.

table · 3 cols
ResourceWhat NVIDIA says it containsLink
Nemotron Labs IMO 2026 collectionSFT and RL checkpoints, both training datasets, Nemotron-IMO-Bench (200 olympiad-level problems)Hugging Face collection
IMO paperTraining approach and the generate-verify-refine systemarXiv:2609.10712
NeMo-Skills (IMO recipe)IMO inference pipeline, prompts, submitted proofs, reproducible quickstartNVIDIA-NeMo/Skills
Nemotron-3-Ultra-CCThe IOI model checkpointHugging Face model
IOI paperTraining recipe and GenCorrectarXiv:2609.02849
NeMo-Skills (IOI)IOI evaluation and inference pipelineNVIDIA-NeMo/Skills

The IMO arXiv abstract adds that NVIDIA releases "the training and inference code, the submitted solutions, and Nemotron-IMO-Bench, a new benchmark of 200 novel olympiad-level problems."

What is not open or not confirmed

  • IMO checkpoint and dataset licenses: not confirmed. The collection page we read lists six items (three models and three datasets) but did not show license terms. Check each repository page before you use the files.
  • IOI model license: the Ultra-CC model page states the OpenMDW License Agreement, version 1.1. Its Hugging Face tag reads "license:other." Read the full terms before commercial use. We did not review the license text.
  • IOI training data and SFT traces: NVIDIA's blog does not say that the 22,000-problem dataset or its synthetic traces are released. We found no such statement. Treat the data as not confirmed open.
  • The IOI 2026 submissions: we did not see a statement that the IOI 2026 run logs are published.
  • Independent replication: no outside team has reproduced these scores in anything we read.

What this means if you build or fine-tune models

The recipe is ordinary on purpose. SFT on curated reasoning traces, optional RL, and a verify-and-refine loop. NVIDIA's pitch is that you can repeat this for your own domain. The hard parts are the data curation and the compute, not a secret method.

Data size is a real benchmark for your own plans. The IMO SFT set had 414,890 examples over 15,818 problems. That is about 26 examples per problem on average (our arithmetic, not NVIDIA's). The IOI set had 22,000 problems. These are not toy datasets.

Plan for inference cost. Both systems are test-time-compute heavy. The IMO system had a separate high-compute selection stage. NVIDIA's post does not give wall-clock time, GPU counts, or dollar cost for the final runs in the text we read. Do not assume you can match the scores at small scale.

Hardware is a gate. The Ultra-CC checkpoint is 550 billion parameters. The model page lists a file size of about 352 GB. Most readers will use hosted inference or a multi-GPU server.

Use the smaller path first. The Nano-CC result is the practical one. A 30B model with 3B active parameters reached 468 points on IOI 2025 with GenCorrect, above that year's 438.3 gold threshold. That is the version a small team can study.

For earlier explainx.ai coverage of NVIDIA's smaller research releases, see the Nemotron Labs TwoTower diffusion LLM post. For how other labs handle AI-made math claims, see Gemini Cogentic's five math problems and Meta Muse Spark's six math problems.

What to do this week

  1. Read the two abstracts first. arXiv:2609.02849 for IOI and arXiv:2609.10712 for IMO. They are short and state the claims precisely.
  2. Open the IMO collection and check each license before you plan any reuse.
  3. Try Nemotron-IMO-Bench as a held-out set of 200 problems for your own proof-writing models, if the license allows it.
  4. Copy the loop, not only the model. Build a generate, critique, refine cycle around whatever model you use, and measure the gain per round.
  5. Keep the IOI caveat in every summary you write. The result is unofficial.

Related reading

  • Nemotron 3 Ultra 550B: the open-weight MoE model
  • Gemini Cogentic: five open math results
  • Meta Muse Spark: six open math problems
  • AGMAI: responsible release of AI-generated mathematics
  • Lean 4 formal proofs and the cost collapse
  • AI benchmarks: the complete guide

Primary sources: NVIDIA on Hugging Face · IOI paper · IMO paper · IMO 2026 collection · Nemotron-3-Ultra-CC model page · NeMo-Skills


Details reflect NVIDIA's October 7, 2026 Hugging Face post, the two arXiv abstracts, and the model and collection pages as read on October 7, 2026. The IOI 2026 result is unofficial and is not in the official IOI ranking. explainx.ai did not re-run any model or re-grade any proof. License terms for the IMO checkpoints and datasets were not confirmed.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 24, 2026

NVIDIA Nemotron 3 Diarization: An Open 100M-Parameter Model That Labels Up to Eight Overlapping Speakers

[Speaker diarization](/dictionary/speaker-diarization), working out who spoke when, is the quiet piece behind every good meeting transcript and voice agent. NVIDIA's new Nemotron 3 Diarization is a 100-million-parameter open-weight model that handles up to eight speakers, overlapping speech and four streaming latency settings, and tops VoiceArena's diarization leaderboard at 14.72% DER.

Aug 11, 2026

NVIDIA Nemotron 3.5 Lightning: A 30B Open MoE Built for Always-On Agents

On August 11, 2026, NVIDIA shipped Nemotron 3.5 Lightning — 30B total parameters, 3B active, interleaved Mamba-2 and MoE layers, up to 1M tokens of context, and a permissive OpenMDW-1.1 license. Here's what the benchmark table actually says, why the released checkpoint is already quantized, and where this model is the wrong choice.

Jul 2, 2026

NVIDIA Nemotron-Labs-TwoTower: Split a 30B Model in Two for 2.42× Faster Diffusion Generation

@NVIDIAAI took a 30B Nemotron Nano and split it in two — one tower holds context, the other denoises blocks of tokens in parallel. No training from scratch: 98.7% of AR quality at 2.42× generation speed. Model on Hugging Face July 2, 2026.