DeepSeekMath-V2 debuts with IMO gold-level performance
On November 27, 2025, DeepSeek quietly launched its DeepSeekMath-V2 model, a self-verifiable mathematical reasoning framework based on DeepSeek V3.2 Exp Base. The model uses an LLM validator to automatically review generated proofs and optimize performance with high-difficulty samples.
The code and weights are open source and available on Hugging Face and GitHub. In benchmark tests, DeepSeekMath-V2 reached gold-level scores in the International Mathematical Olympiad (IMO 2025) and Chinese Mathematical Olympiad (CMO 2024), and scored 118/120 on the Putnam 2024 exam.
In basic reasoning tests, it scored 99 points - ahead of Claude Sonnet4, GPT-5, and Gemini 2.5 Pro. In advanced tests, it scored 65.7, second only to Google's Gemini DeepThink.
DeepSeek emphasized that while the model is still evolving, these results validate the feasibility of self-verifiable mathematical reasoning and lay the foundation for more powerful math-focused AI systems.
DeepSeek is a Chinese AI research company specializing in large language models and reasoning frameworks. The DeepSeekMath-V2 model is trained on DeepSeek V3.2 Exp Base and uses a multi-stage validation pipeline. Benchmark scores are reported in competition-standard formats and converted to U.S. equivalents where applicable. The model is open source and free to use. Deployment costs are estimated at ¥10 million (about $1.4 million USD), with inference latency under 1 second for most tasks.