Aristotle: IMO-level Automated Theorem Proving
DOI:
https://doi.org/10.70777/si.v3i3.18745Keywords:
Aristotle, automated theorem proving, formal verification, Lean 4, mathematical reasoning, Monte Carlo Graph Search, International Mathematical OlympiadAbstract
We introduce Aristotle, an AI system that combines formal verification with informal reasoning, achieving gold-medal-equivalent performance on the 2025 International Mathematical Olympiad problems. Aristotle integrates three main components: a Lean proof search system, an informal reasoning system that generates and formalizes lemmas, and a dedicated geometry solver. Our system demonstrates state-of-the-art performance with favorable scaling properties for automated theorem proving.
References
[1] Ekin Akyürek, Mehul Damani, Adam Zweiger, Linlu Qiu, Han Guo, Jyothish Pari, Yoon Kim, and Jacob Andreas. The Surprising Effectiveness of Test‑Time Training for Few‑Shot Learning, 2024. URL https://arxiv.org/abs/2411.07279.
[2] Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba. Hindsight Experience Replay. In Advances in Neural Information Processing Systems (NeurIPS) 30, pages 5055–5065, 2017. URL https://papers.nips.cc/paper/7090-hindsight-experience-replay.pdf.
[3] Thomas Anthony, Zheng Tian, and David Barber. Thinking Fast and Slow with Deep Learning and Tree Search. In Advances in Neural Information Processing Systems 30 (NeurIPS 2017), pages 5360–5370, 2017.
[4] Zhangir Azerbayev, Bartosz Piotrowski, Hailey Schoelkopf, Edward W. Ayers, Dragomir Radev, and Jeremy Avigad. ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics, 2023. URL https://arxiv.org/abs/2302.12433.
[5] Luoxin Chen, Jinming Gu, Liankai Huang, Wenhao Huang, Zhicheng Jiang, Allan Jie, Xiaoran Jin, Xing Jin, Chenggang Li, Kaijing Ma, Cheng Ren, Jiawei Shen, Wenlei Shi, Tong Sun, He Sun, Jiahui Wang, Siran Wang, Zhihong Wang, Chenrui Wei, Shufa Wei, Yonghui Wu, Yuchen Wu, Yihang Xia, Huajian Xin, Fan Yang, Huaiyuan Ying, Hongyi Yuan, Zheng Yuan, Tianyang Zhan, Chi Zhang, Yue Zhang, Ge Zhang, Tianyun Zhao, Jianqiu Zhao, Yichi Zhou, and Thomas Hanwen Zhu. Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving, 2025. URL https://arxiv.org/abs/2507.23726.
[6] Yuri Chervonyi, Trieu H. Trinh, Miroslav Olšák, Xiaomeng Yang, Hoang Nguyen, Marcelo Menegali, Junehyuk Jung, Vikas Verma, Quoc V. Le, and Thang Luong. Gold‑medalist Performance in Solving Olympiad Geometry with AlphaGeometry2, 2025. URL https://arxiv.org/abs/2502.03544.
[7] Adrien Couëtoux, Jean-Baptiste Hoock, Nataliya Sokolovska, Olivier Teytaud, and Nicolas Bonnard. Contin- uous upper confidence trees. In Proceedings of the 5th International Conference on Learning and Intelligent Optimization (LION’05), pages 433–445, Berlin, Heidelberg, 2011. Springer-Verlag.
[8] Johannes Czech, Patrick Korus, and Kristian Kersting. Improving alphazero using monte-carlo graph search. In Proceedings of the 31st International Conference on Automated Planning and Scheduling (ICAPS 2021), pages 103–111, 2021.
[9] Kefan Dong and Tengyu Ma. STP: Self-Play LLM Theorem Provers with Iterative Conjecturing and Proving, 2025. URL https://arxiv.org/abs/2502.00212.
[10] Emily First, Markus N. Rabe, Talia Ringer, and Yuriy Brun. Baldur: Whole-Proof Generation and Repair with Large Language Models. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE 2023), pages 1229–1241. Association for Computing Machinery, 2023.
[11] Fabian Gloeckle, Jannis Limperg, Gabriel Synnaeve, and Amaury Hayat. ABEL: Sample Efficient Online Reinforcement Learning for Neural Theorem Proving. In NeurIPS 2024 Workshop MATH-AI, 2024. URL https://openreview.net/forum?id=kk3mSjVCUO.
[12] Google DeepMind. AlphaProof: when reinforcement learning meets formal mathematics. YouTube video, March 2025. URL https://www.youtube.com/watch?v=TFBzP78Jp6A. Presenter: Thomas Hubert.
[13] Jesse Michael Han, Jason Rute, Yuhuai Wu, Edward W. Ayers, and Stanislas Polu. Proof Artifact Co-Training for Theorem Proving with Language Models, 2021. URL https://arxiv.org/abs/2102.06203.
[14] Harmonic. IMO2025: Harmonic’s IMO 2025 Results (Problems & Proofs in Lean). GitHub repository, 2025. URL https://github.com/harmonic-ai/IMO2025. Contains Lean statements and proofs for Problems 1–5 on IMO 2025.
[15] Harmonic. Aristotle Achieves Gold Medal Performance at the IMO. Blog post, July 28, 2025. URL https://harmonic.fun/news#blog-post-imo.
[16] Harmonic. Running Lean at Scale. Blog post, September 11, 2025. URL https://harmonic.fun/news# blog-post-lean.
[17] Albert Qiaochu Jiang, Wenda Li, Jesse Michael Han, and Yuhuai Wu. LISA: Language Models of Isabelle Proofs. In Proceedings of the 6th Conference on Artificial Intelligence and Theorem Proving (AITP 2021), volume 83 of EPiC Series in Computing, pages 378–392. EasyChair, 2021.
[18] Albert Qiaochu Jiang, Wenda Li, Szymon Tworkowski, Konrad Czechowski, Tomasz Odrzygóźdź, Piotr Miłoś, Yuhuai Wu, and Mateja Jamnik. THOR: Wielding Hammers to Integrate Language Models and Automated Theorem Provers. In Advances in Neural Information Processing Systems 35 (NeurIPS 2022), pages 8360–8373, 2022.
[19] Albert Qiaochu Jiang, Sean Welleck, Jin Peng Zhou, Wenda Li, Jiacheng Liu, Mateja Jamnik, Timothée Lacroix, Yuhuai Wu, and Guillaume Lample. Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs. In International Conference on Learning Representations (ICLR), 2023. URL https://arxiv.org/abs/2210.12283.
[20] Guillaume Lample, Marie-Anne Lachaux, Thibaut Lavril, Xavier Martinet, Amaury Hayat, Gabriel Ebner, Au- rélien Rodriguez, and Timothée Lacroix. HyperTree Proof Search for Neural Theorem Proving. In Advances in Neural Information Processing Systems 35 (NeurIPS 2022), pages 34651–34664, 2022.
[21] Yang Li, Dong Du, Linfeng Song, Chen Li, Weikang Wang, Tao Yang, and Haitao Mi. HunyuanProver: A Scalable Data Synthesis Framework and Guided Tree Search for Automated Theorem Proving, 2024. URL https://arxiv.org/abs/2412.20735.
[22] Haohan Lin, Zhiqing Sun, Yiming Yang, and Sean Welleck. Lean-STaR: Learning to Interleave Thinking and Proving, 2024. URL https://arxiv.org/abs/2407.10040.
[23] Yong Lin, Shange Tang, Bohan Lyu, Jiayun Wu, Hongzhou Lin, Kaiyu Yang, Jia Li, Mengzhou Xia, Danqi Chen, Sanjeev Arora, and Chi Jin. Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving, 2025. URL https://arxiv.org/abs/2502.07640.
[24] Xiaoyang Liu, Kangjie Bao, Jiashuo Zhang, Yunqi Liu, Yu Chen, Yuntian Liu, Yang Jiao, and Tao Luo. ATLAS: Autoformalizing Theorems through Lifting, Augmentation, and Synthesis of Data. 2025. URL https://arxiv. org/abs/2502.05567.
[25] Alexander Meiburg. Two lemmas for approxhom. GitHub pull request #251 in teorth/pfr, 2025. URL https://github.com/teorth/pfr/pull/251.
[26] Alexander Meiburg. Roots of Matrix.charpoly are the eigenvalues. GitHub pull request #27118 in leanprover-community/mathlib4, 2025. URL https://github.com/leanprover-community/mathlib4/pull/27118.
[27] Alexander Meiburg. limsup/liminf of f + g when either f or g tends to zero. GitHub pull request #27115 in leanprover-community/mathlib4, 2025. URL https://github.com/leanprover-community/mathlib4/pull/27115.
[28] Alexander Meiburg. Niven’s theorem. GitHub pull request #26371 in leanprover-community/mathlib4, 2025. URL https://github.com/leanprover-community/mathlib4/pull/26371.
[29] Alexander Meiburg, Leonardo A Lessa, and Rodolfo R. Soldati. Quantum information in Lean, 2025. URL https://github.com/Timeroot/Lean-QuantumInfo.
[30] Maciej Mikuła, Szymon Tworkowski, Szymon Antoniak, Bartosz Piotrowski, Albert Qiaochu Jiang, Jin Peng Zhou, Christian Szegedy, Łukasz Kuciński, Piotr Miłoś, and Yuhuai Wu. MagnusHammer: A Transformer-Based Approach to Premise Selection, 2023. URL https://arxiv.org/abs/2303.04488.
[31] Azim Ospanov, Farzan Farnia, and Roozbeh Yousefzadeh. APOLLO: Automated LLM and Lean Collaboration for Advanced Formal Reasoning. 2025. URL https://arxiv.org/abs/2505.05758.
[32] Stanislas Polu and Ilya Sutskever. Generative Language Modeling for Automated Theorem Proving, 2020. URL https://arxiv.org/abs/2009.03393.
[33] Stanislas Polu, Jesse Michael Han, Kunhao Zheng, Mantas Baksys, Igor Babuschkin, and Ilya Sutskever. For- mal Mathematics Statement Curriculum Learning. In The Eleventh International Conference on Learning Representations (ICLR 2023), 2023. URL https://openreview.net/forum?id=6CE2L64c0de.
[34] Z. Z. Ren, Zhihong Shao, Junxiao Song, Huajian Xin, Haocheng Wang, Wanjia Zhao, Liyue Zhang, Zhe Fu, Qihao Zhu, Dejian Yang, Z. F. Wu, Zhibin Gou, Shirong Ma, Hongxuan Tang, Yuxuan Liu, Wenjun Gao, Daya Guo, and Chong Ruan. DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition. 2025. URL https://arxiv.org/abs/2504.21801.
[35] Christopher D. Rosin. Multi-armed bandits with episode context. Annals of Mathematics and Artificial Intelligence, 61:203–230, 2011. doi: 10.1007/s10472-011-9258-6. URL https://link.springer.com/article/10.1007/s10472-011-9258-6.
[36] Ziju Shen, Naohao Huang, Fanyi Yang, Yutong Wang, Guoxiong Gao, Tianyi Xu, Jiedong Jiang, Wanyi He, Pu Yang, Mengzhou Sun, Haocheng Ju, Peihao Wu, Bryan Dai, and Bin Dong. REAL‑Prover: Retrieval Aug- mented Lean Prover for Mathematical Reasoning, 2025. URL https://arxiv.org/abs/2505.20613.
[37] David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanc- tot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis. A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play. Science, 362 (6419):1140–1144, 2018. doi: 10.1126/science.aar6404. URL https://www.science.org/doi/10.1126/science.aar6404.
[38] Terence Tao. Analysis I, volume 37 of Texts and Readings in Mathematics. Hindustan Book Agency, New Delhi, India, 3 edition, 2014.
[39] Terence Tao, 2025. URL https://github.com/teorth/analysis.
[40] Terence Tao. Some corrections. GitHub pull request #113 in teorth/analysis, 2025. URL https://github. com/teorth/analysis/pull/113.
[41] Terence Tao. Machine‐assisted proof. Notices of the American Mathematical Society, 72(1):6–13, 2025.
[42] Terence Tao, Yaël Dillies, et al. Repository for formalization of the Polynomial Freiman-Ruzsa conjecture, 2024. URL https://teorth.github.io/pfr/.
[43] Amitayush Thakur, George Tsoukalas, Yeming Wen, Jimmy Xin, and Swarat Chaudhuri. An In-Context Learning Agent for Formal Theorem-Proving. In Proceedings of the 2024 Conference on Language Modeling (COLM 2024), 2024. URL https://openreview.net/forum?id=V7HRrxXUhN.
[44] The Lean Prover Community. A read-eval-print-loop for Lean 4. GitHub repository, 2023. URL https://github.com/leanprover-community/repl.
[45] Trieu H. Trinh, Yuhuai Wu, Quoc V. Le, He He, and Thang Luong. Solving olympiad geometry without human demonstrations. Nature, 625(7961):476–482, 2024. doi: 10.1038/s41586-023-06747-5. URL https://www.nature.com/articles/s41586-023-06747-5.
[46] George Tsoukalas, Jasper Lee, John Jennings, Jimmy Xin, Michelle Ding, Michael Jennings, Amitayush Thakur, and Swarat Chaudhuri. PutnamBench: Evaluating Neural Theorem‑Provers on the Putnam Mathematical Com- petition, 2024. URL https://arxiv.org/abs/2407.11214.
[47] Haiming Wang, Ye Yuan, Zhengying Liu, Jianhao Shen, Yichun Yin, Jing Xiong, Enze Xie, Han Shi, Yujun Li, Lin Li, et al. DT-Solver: Automated Theorem Proving with Dynamic-Tree Sampling Guided by Proof-Level Value Function. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12632–12646, Toronto, Canada, 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.703. URL https://aclanthology.org/2023.acl-long.703/.
[48] Haiming Wang, Huajian Xin, Chuanyang Zheng, Lin Li, Zhengying Liu, Qingxing Cao, Yinya Huang, Jing Xiong, Han Shi, Enze Xie, Jian Yin, Zhenguo Li, Heng Liao, and Xiaodan Liang. LEGO-Prover: Neural Theorem Proving with Growing Libraries. In Proceedings of the Twelfth International Conference on Learning Representations (ICLR 2024), 2024. URL https://proceedings.iclr.cc/paper_files/paper/2024/file/85dca46374dc0f27b4bb5f265b3d17f0-Paper-Conference.pdf.
[49] Haiming Wang, Mert Unsal, Xiaohan Lin, Mantas Baksys, Junqi Liu, Marco Dos Santos, Flood Sung, Marina Vinyes, Zhenzhe Ying, Zekai Zhu, Jianqiao Lu, Hugues de Saxcé, Bolton Bailey, Chendong Song, Chenjun Xiao, Dehao Zhang, Ebony Zhang, Frederick Pu, Han Zhu, Jiawei Liu, Jonas Bayer, Julien Michel, Longhui Yu, Léo Dreyfus-Schmidt, Lewis Tunstall, Luigi Pagani, Moreira Machado, Pauline Bourigault, Ran Wang, Stanislas Polu, Thibaut Barroyer, Wen-Ding Li, Yazhe Niu, Yann Fleureau, Yangyang Hu, Zhouliang Yu, Zihan Wang, Zhilin Yang, Zhengying Liu, and Jia Li. Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning. 2025. URL http://arxiv.org/abs/2504.11354.
[50] David J. Wu. Monte-Carlo Graph Search from First Principles. KataGo GitHub Documentation. URL https://github.com/lightvector/KataGo/blob/master/docs/GraphSearch.md.
[51] Yuhuai Wu, Albert Qiaochu Jiang, Wenda Li, Markus Rabe, Charles Staats, Mateja Jamnik, and Christian Szegedy. Autoformalization with Large Language Models. In Advances in Neural Information Processing Systems 35 (NeurIPS 2022), pages 32353–32368, 2022.
[52] Zijian Wu, Suozhi Huang, Zhejian Zhou, Huaiyuan Ying, Jiayu Wang, Dahua Lin, and Kai Chen. InternLM2.5- StepProver: Advancing Automated Theorem Proving via Expert Iteration on Large-Scale LEAN Problems, 2024. URL https://arXiv.org/abs/2410.15700.
[53] Huajian Xin, Daya Guo, Zhihong Shao, Zhizhou Ren, Qihao Zhu, Bo Liu, Chong Ruan, Wenda Li, and Xiaodan Liang. DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data, 2024. URL https://arxiv.org/abs/2405.14333.
[54] Huajian Xin, Z. Ren, Z. Junxiao Song, Zhihong Shao, Wanjia Zhao, Haocheng Wang, Bo Liu, Liyue Zhang, Xuan Lu, Qiushi Du, Wenjun Gao, Qihao Zhu, Dejian Yang, Zhibin Gou, F. Wu, Z. Fuli Luo, and Chong Ruan. DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte- Carlo Tree Search. 2025. URL https://openreview.net/forum?id=I4YAIwrsXa. Poster presentation; see OpenReview.
[55] Ran Xin, Chenguang Xi, Jie Yang, Feng Chen, Hang Wu, Xia Xiao, Yifan Sun, Shen Zheng, and Ming Ding. BFS‑Prover: Scalable Best‑First Tree Search for LLM‑based Automatic Theorem Proving, July 2025.
[56] Kaiyu Yang, Aidan M. Swope, Alex Gu, Rahul Chalamala, Peiyang Song, Shixing Yu, Saad Godil, Ryan Prenger, and Anima Anandkumar. LeanDojo: Theorem Proving with Retrieval-Augmented Language Mod- els. In Proceedings of the 2023 Conference on Neural Information Processing Systems Datasets & Benchmarks Track, pages 4668–4685, 2023.
[57] Zhouliang Yu, Ruotian Peng, Keyi Ding, Yizhe Li, Zhongyuan Peng, Minghao Liu, Yifan Zhang, Zheng Yuan, Huajian Xin, Wenhao Huang, Yandong Wen, Ge Zhang, and Weiyang Liu. FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models, 2025. URL https://arxiv.org/abs/2505.02735.
[58] Jingyuan Zhang, Qi Wang, Xingguang Ji, Yahui Liu, Yang Yue, Fuzheng Zhang, Di Zhang, Guorui Zhou, and Kun Gai. LeanAbell-Prover: Posttraining Scaling in Formal Reasoning, 2025. URL https://arxiv.org/abs/2504.06122.
[59] Kunhao Zheng, Jesse Michael Han, and Stanislas Polu. miniF2F: a cross-system benchmark for formal Olympiad-level mathematics. International Conference on Learning Representations (ICLR), 2022. URL https://openreview.net/forum?id=9ZPegFuFTFv.
[60] Yichi Zhou, Jianqiu Zhao, Yongxin Zhang, Bohan Wang, Siran Wang, Luoxin Chen, Jiahui Wang, Haowei Chen, Allan Jie, Xinbo Zhang, Haocheng Wang, Luong Trung, Rong Ye, Phan Nhat Hoang, Huishuai Zhang, Peng Sun, and Hang Li. Solving Formal Math Problems by Decomposition and Iterative Reflection, 2025. URL https://arxiv.org/abs/2507.15225.
[61] Thomas Zhu, Joshua Clune, Jeremy Avigad, Albert Qiaochu Jiang, and Sean Welleck. Premise Selection for a Lean Hammer. 2025. URL https://arxiv.org/abs/2506.07477.
Published
How to Cite
Issue
Section
Categories
License
Copyright (c) 2026 Tudor Achim, Alex Best, Alberto Bietti, Kevin Der, Mathïs Fédérico, Sergei Gukov, Daniel Halpern-Leistner, Kirsten Henningsgard, Yury Kudryashov, Alexander Meiburg, Martin Michelsen, Riley Patterson, Eric Rodriguez, Laura Scharff, Vikram Shanker, Vladmir Sicca, Hari Sowrirajan, Aidan Swope, Matyas Tamas, Vlad Tenev, Jonathan Thomm, Harold Williams, Lawrence Wu

This work is licensed under a Creative Commons Attribution 4.0 International License.