Reactions from ranked influencers
7 postsThe International Math Olympiad (IMO) 2026, the hardest math contest for high schoolers, just ended. I ran Fable (high), Sol (xhigh), K3 (max) and Axiom against it and all got a perfect score of 42/42 (repo below if you want to check their solutions): — Claude Fable 5 was the solved it in 1 attempt, and was the fastest. — GPT 5.6 Sol took 1 more attempts, and was cheapest. — Kimi K3 did it but took 4 more attempts, and took a LOT of tokens. — Axiom Math actually proved everything in Lean. P3 and P6 were the hardest followed by P2, judging by attempts + num tokens. Students had 9hrs to solve these 6 problems, and Fable and Sol were under 4hrs. The frontier of AI has officially moved well past IMO math.
Results like this suggest that human-designed tests may soon be insufficient to measure the capabilities of frontier models, if they are not already. In a few years, we may need an “IMO for agents”: problems generated and filtered by multiple frontier models using very large compute budgets, then given to individual agents to solve under strict, limited compute. Humans would define the rules and verify correctness, but the challenge generation itself may need to be AI-assisted to stay ahead of the systems being evaluated.
The International Math Olympiad (IMO) 2026, the hardest math contest for high schoolers, just ended. I ran Fable (high), Sol (xhigh), K3 (max) and Axiom against it and all got a perfect score of 42/42 (repo below if you want to check their solutions): — Claude Fable 5 was the solved it in 1 attempt, and was the fastest. — GPT 5.6 Sol took 1 more attempts, and was cheapest. — Kimi K3 did it but took 4 more attempts, and took a LOT of tokens. — Axiom Math actually proved everything in Lean. P3 and P6 were the hardest followed by P2, judging by attempts + num tokens. Students had 9hrs to solve these 6 problems, and Fable and Sol were under 4hrs. The frontier of AI has officially moved well past IMO math.
even K3 can ace IMO It's using a lot of tokens but not THAT many tokens. without trivial network/billing failures and at 80tps, it'd be within a factor of 2 of 5.6 Sol. More compute and RL needed. Almost there.
The International Math Olympiad (IMO) 2026, the hardest math contest for high schoolers, just ended. I ran Fable (high), Sol (xhigh), K3 (max) and Axiom against it and all got a perfect score of 42/42 (repo below if you want to check their solutions): — Claude Fable 5 was the solved it in 1 attempt, and was the fastest. — GPT 5.6 Sol took 1 more attempts, and was cheapest. — Kimi K3 did it but took 4 more attempts, and took a LOT of tokens. — Axiom Math actually proved everything in Lean. P3 and P6 were the hardest followed by P2, judging by attempts + num tokens. Students had 9hrs to solve these 6 problems, and Fable and Sol were under 4hrs. The frontier of AI has officially moved well past IMO math.
of note: 3 people on Chinese team could do 42/42 likewise for Russia 2 for the US, UK, Canada, Bulgaria, Taiwan, Ukraine, Thailand… (I have to congratulate India though, they came far) Anyway, this is mostly an indicator of how many grindcels a nation can field
Source: Lean (from Axiom): https://github.com/AxiomMath/IMO2026 Fable: https://github.com/deedy/imo-2026/blob/main/solutions/claude-fable-5.pdf Sol: https://github.com/deedy/imo-2026/blob/main/solutions/gpt-5.6-sol-xhigh.pdf Kimi K3: https://github.com/deedy/imo-2026/blob/main/solutions/kimi-k3.pdf Repo: https://github.com/deedy/imo-2026/blob/main/REPORT.md
This is the first time IMO has been conclusively solved. Last year, the best models was an unreleased Gemini Deep Think and a experimental OpenAI model which both did 35/42. This year, 3 models that are open for public use, including a (soon) open weight one, did it for $10-50!
Combined views
306K
7 posts, first seen 10h ago