Open recipe targets IMO gold with post-trained Nemotron math models
A new arXiv paper examines how post-training choices and test-time inference setups influence a model's ability to write natural-language proofs for difficult olympiad problems. Using Nemotron 3 Ultra as a base, the authors produce two specialist checkpoints via supervised fine-tuning and reinforcement learning, and release the training approach publicly.