AI reaches gold at the International Mathematical Olympiad (2025)

In July 2025, language models from Google DeepMind and OpenAI solve five of six IMO problems for 35 of 42 points, writing proofs in natural language within the exam time limit.

Event date: Published:
HistoryResearch

In July 2025, the International Mathematical Olympiad met on Australia’s Sunshine Coast. Within three days, two labs announced that their language models had reached gold-medal level on the same six problems: 35 points out of 42, five tasks solved, proofs written in natural language under the same time rules as the human contestants.

What makes the IMO so hard

The IMO is not a calculator contest. Each year, roughly the top 8 percent of participants earn gold on six problems drawn from algebra, combinatorics, geometry, and number theory. Contestants sit two exams of 4.5 hours each and must produce complete, creative proofs, not just numeric answers.

That puts the IMO closer to what people mean by “thinking” than bounded games like chess or Go. Deep Blue (1997) and AlphaGo (2016) won in closed worlds with fixed rules. The IMO demands open-ended reasoning in words.

From silver in Lean to gold in language

In 2024, Google DeepMind’s AlphaProof and AlphaGeometry 2 reached silver with 28 points, four problems solved. Experts first translated the tasks into the formal language Lean; the run took two to three days.

In 2025, both leading systems worked end to end in natural language within the 4.5-hour limit, with no tools and no internet access. DeepMind described Gemini Deep Think as using “parallel thinking” (pursuing several solution paths at once rather than a single linear chain of thought), new reinforcement learning techniques, and curated training on high-quality solutions. That is the same broad recipe behind o1 and DeepSeek-R1: more test-time compute at the moment you ask the question.

Two paths to gold

On July 19, OpenAI researcher Alexander Wei reported on X that an experimental reasoning model (not GPT-5) had scored 35/42 under the official conditions. Each proof was graded by three former IMO medalists who had to agree; the results were published on GitHub. OpenAI said it did not plan to release a public model at that maths level for several months.

On July 21, Google DeepMind announced that Gemini Deep Think had achieved the same score. IMO president Gregor Dolinar confirmed that DeepMind had earned 35 of 42 possible points, a gold-medal score, under grading by IMO coordinators using the same criteria as for students. DeepMind also noted a limit: the IMO confirmed the correctness of the solutions but did not validate the system, process, or model behind them.

Ars Technica stressed the difference: OpenAI did not take part in the official grading process, while DeepMind waited for coordinator confirmation.

Why it matters

The result extends the line from AlphaFold in science to competition mathematics, and it lands in the middle of the reasoning boom that o1 opened and DeepSeek-R1 widened. It also revives the evaluation question the o1 article raised: what counts as “solving” when one score is officially certified and another is peer-reviewed outside the contest?

Contest maths is not research maths. Gold at the IMO is a milestone, not a proof that machines have mastered open problems. Still, crossing the same threshold as the world’s best young mathematicians, in language rather than in a bespoke formal pipeline, showed how far general-purpose reasoning had moved in a single year.

#reasoning#llm#reinforcement-learning