AI for Mathematics: Progress, Challenges, and Prospects
DOI:
https://doi.org/10.70777/si.v3i3.18747Keywords:
AI for mathematics, mathematical discovery, automated theorem proving, formal reasoning, autoformalization, mathematical information retrieval, foundation models, machine learningAbstract
AI for Mathematics (AI4Math) has emerged as a distinct field that leverages machine learning to navigate mathematical landscapes historically intractable for early symbolic systems. While mid-20th-century symbolic approaches successfully automated formal logic, they faced severe scalability limitations due to the combinatorial explosion of the search space. The recent integration of data-driven approaches has revitalized this pursuit. In this review, we provide a systematic overview of AI4Math, highlighting its primary focus on developing AI models to support mathematical research. Crucially, we emphasize that this is not merely the application of AI to mathematical activities; it also encompasses the development of stronger AI systems where the rigorous nature of mathematics serves as a premier testbed for advancing general reasoning capabilities. We categorize existing research into two complementary directions: problem-specific modeling, involving the design of specialized architectures for distinct mathematical tasks, and general-purpose modeling, focusing on foundation models capable of broader reasoning, retrieval, and exploratory workflows. We conclude by discussing key challenges and prospects, advocating for AI systems that go beyond facilitating formal correctness to enabling the discovery of meaningful results and unified theories, recognizing that the true value of a proof lies in the insights and tools it offers to the broader mathematical landscape.
References
[1] Mohammed Abouzaid, Andrew J Blumberg, Martin Hairer, Joe Kileel, Tamara G Kolda, Paul D Nelson, Daniel Spielman, Nikhil Srivastava, Rachel Ward, Shmuel Weinberger, et al. First Proof. arXiv preprint arXiv:2602.05192, 2026.
[2] Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774, 2023.
[3] Tudor Achim, Alex Best, Alberto Bietti, Kevin Der, Mathı̈s Fédérico, Sergei Gukov, Daniel Halpern-Leistner, Kirsten Henningsgard, Yury Kudryashov, Alexander Meiburg, et al. Aristotle: IMO-Level Automated Theorem Proving. arXiv preprint arXiv:2510.01346, 2025.
[4] Ayush Agrawal, Siddhartha Gadgil, Navin Goyal, Ashvni Narayanan, and Anand Tadi- patri. Towards a Mathematics Formalisation Assistant Using Large Language Models. arXiv preprint arXiv:2211.07524, 2022.
[5] Luke Alexander, Eric Leonen, Sophie Szeto, Artemii Remizov, Ignacio Tejeda, Jarod Alper, Giovanni Inchiostro, and Vasily Ilin. Semantic search over 9 million mathematical theorems. arXiv preprint arXiv:2602.05216, 2026.
[6] Alberto Alfarano, François Charton, and Amaury Hayat. Global Lyapunov Functions: A Long-Standing Open Problem in Mathematics, with Symbolic Transformers. Advances in Neural Information Processing Systems, 37:93643–93670, 2024.
[7] Malik Amir, Yang-Hui He, Kyu-Hwan Lee, Thomas Oliver, and Eldar Sultanow. Machine Learning Class Numbers of Real Quadratic Fields. International Journal of Data Science in the Mathematical Sciences, 01(02):107–134, 2023.
[8] Simon Arridge, Peter Maass, Ozan Öktem, and Carola-Bibiane Schönlieb. Solving Inverse Problems Using Data-Driven Models. Acta Numerica, 28:1–174, 2019.
[9] Justin Asher. LeanExplore: A Search Engine for Lean 4 Declarations. arXiv preprint arXiv:2506.11085, 2025.
[10] Axiom Math. From Seeing Why to Checking Everything. https://axiommath.ai/ territory/from-seeing-why-to-checking-everything, 2025.
[11] Zhangir Azerbayev, Bartosz Piotrowski, Hailey Schoelkopf, Edward W Ayers, Dragomir Radev, and Jeremy Avigad. Proofnet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics. arXiv preprint arXiv:2302.12433, 2023.
[12] Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen Marcus McAleer, Albert Q. Jiang, Jia Deng, Stella Biderman, and Sean Welleck. Llemma: An Open Language Model for Mathematics. In The Twelfth International Conference on Learning Representations, 2024.
[13] Kshitij Bansal, Sarah Loos, Markus Rabe, Christian Szegedy, and Stewart Wilcox. Holist: An Environment for Machine Learning of Higher Order Logic Theorem Proving. In In- ternational Conference on Machine Learning, pages 454–463. PMLR, 2019.
[14] Timothée Bénard and Weikun He. Effective Brascamp–Lieb Inequalities. arXiv preprint arXiv:2511.11091, 2025. 21 SuperIntelligence - Safety & Alignment 2026 V3N3 Toward AGI Science - Math - ATP
[15] Gergely Bérczi, Honglu Fan, and Mingcong Zeng. An ML Approach to Resolution of Singularities. In Topological, Algebraic and Geometric Learning Workshops 2023, pages 469–487. PMLR, 2023.
[16] Per Berglund, Ben Campbell, and Vishnu Jejjala. Machine Learning Kreuzer–Skarke Calabi–Yau Threefolds. International Journal of Modern Physics A, 40(28):2550097, 2025.
[17] Per Berglund, Yang-Hui He, Elli Heyes, Edward Hirst, Vishnu Jejjala, and Andre Lukas. New Calabi–Yau Manifolds from Genetic Algorithms. Physics Letters B, 850:138504, 2024.
[18] David S. Berman, Yang-Hui He, and Edward Hirst. Machine Learning Calabi-Yau Hy- persurfaces. Phys. Rev. D, 105:066002, Mar 2022.
[19] Charles Blundell, Lars Buesing, Alex Davies, Petar Veličković, and Geordie Williamson. Towards Combinatorial Invariance for Kazhdan-Lusztig Polynomials. Representation The- ory of the American Mathematical Society, 26(37):1145–1191, 2022.
[20] Jonathan W Bober, Andrew R Booker, M Lee, and David Lowry-Duda. Murmurations of Modular Forms in the Weight Aspect. Algebra and Number Theory, 2025.
[21] Steven L Brunton and J Nathan Kutz. Promising Directions of Machine Learning for Partial Differential Equations. Nature Computational Science, 4(7):483–494, 2024.
[22] Jim Bryan, Balázs Elek, Freddie Manners, George Salafatinos, and Ravi Vakil. The Motivic Class of the Space of Genus 0 Maps to the Flag Variety. arXiv preprint arXiv:2601.07222, 2026.
[23] Gergely Bérczi, Baran Hashemi, and Jonas Klüver. Flow-Based Extremal Mathematical Structure Discovery, 2026.
[24] Paul-Jean Cahen, Marco Fontana, Sophie Frisch, and Sarah Glaz. Open problems in commutative ring theory. In Commutative Algebra: Recent Advances in Commutative Rings, Integer-Valued Polynomials, and Polynomial Functions, pages 353–375. Springer, 2014.
[25] Jonathan Carifio, James Halverson, Dmitri Krioukov, and Brent D Nelson. Machine Learning in the String Landscape. Journal of High Energy Physics, 2017(9):1–36, 2017.
[26] François Charton, Jordan S Ellenberg, Adam Zsolt Wagner, and Geordie Williamson. Patternboost: Constructions in Mathematics with a Little Help from AI. arXiv preprint arXiv:2411.00566, 2024.
[27] Jiangjie Chen, Wenxiang Chen, Jiacheng Du, Jinyi Hu, Zhicheng Jiang, Allan Jie, Xiao- ran Jin, Xing Jin, Chenggang Li, Wenlei Shi, Zhihong Wang, Mingxuan Wang, Chenrui Wei, Shufa Wei, Huajian Xin, Fan Yang, Weihao Gao, Zheng Yuan, Tianyang Zhan, Zeyu Zheng, Tianxi Zhou, and Thomas Hanwen Zhu. Seed-Prover 1.5: Mastering Undergraduate-Level Theorem Proving via Learning from Experience. arXiv preprint arXiv:2512.17260, 2025.
[28] Tianlong Chen, Xiaohan Chen, Wuyang Chen, Howard Heaton, Jialin Liu, Zhangyang Wang, and Wotao Yin. Learning to Optimize: A Primer and A Benchmark. Journal of Machine Learning Research, 23(189):1–59, 2022.
[29] Yuri Chervonyi, Trieu H Trinh, Miroslav Olšák, Xiaomeng Yang, Hoang Nguyen, Marcelo Menegali, Junehyuk Jung, Vikas Verma, Quoc V Le, and Thang Luong. Gold-Medalist Performance in Solving Olympiad Geometry with AlphaGeometry2. arXiv preprint arXiv:2502.03544, 2025. 22 SuperIntelligence - Safety & Alignment 2026 V3N3 Toward AGI Science - Math - ATP
[30] A Chervov, A Soibelman, S Lytkin, I Kiselev, S Fironov, A Lukyanenko, A Dolgorukova, A Ogurtsov, F Petrov, S Krymskii, et al. CayleyPy RL: Pathfinding and Reinforcement Learning on Cayley Graphs. arXiv preprint arXiv:2502.18663, 2025.
[31] Alexander Chervov, Kirill Khoruzhii, Nikita Bukhal, Jalal Naghiyev, Vladislav Zamkovoy, Ivan Koltsov, Lyudmila Cheldieva, Arsenii Sychev, Arsenii Lenin, Mark Obozov, et al. A Machine Learning Approach That Beats Large Rubik’s Cubes. arXiv preprint arXiv:2502.13266, 2025.
[32] Shang-Ching Chou and Xiao-Shan Gao. Automated Reasoning in Geometry. Handbook of automated reasoning, 1:707–749, 2001.
[33] Shang-Ching Chou, Xiao-Shan Gao, and Jing-Zhong Zhang. A Deductive Database Ap- proach to Automated Geometry Theorem Proving and Discovering. Journal of Automated Reasoning, 25(3):219–246, 2000.
[34] D. A. Clarke. Martin Davis. A program for Presburger’s algorithm. Summaries of talks presented at the Summer Institute for Symbolic Logic, Cornell University, 1957, 2nd edn., Communications Research Division, Institute for Defense Analyses, Princeton, N.J., 1960, pp. 215–223. Journal of Symbolic Logic, 31(1):138–138, 1966.
[35] Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training Verifiers to Solve Math Word Problems. arXiv preprint arXiv:2110.14168, 2021.
[36] Katherine M Collins, Albert Q Jiang, Simon Frieder, Lionel Wong, Miri Zilka, Umang Bhatt, Thomas Lukasiewicz, Yuhuai Wu, Joshua B Tenenbaum, William Hart, et al. Evaluating Language Models for Mathematics Through Interactions. Proceedings of the National Academy of Sciences, 121(24):e2318124121, 2024.
[37] Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities. arXiv preprint arXiv:2507.06261, 2025.
[38] MARIANA COOK. Mathematicians: An Outer View of the Inner World. Princeton University Press, 2009.
[39] Lukasz Czajka and Cezary Kaliszyk. Hammer for Coq: Automation for Dependent Type Theory. Journal of Automated Reasoning, 61(1):423–453, 2018.
[40] Alex Davies, András Juhász, Marc Lackenby, and Nenad Tomašev. The Signature and Cusp Geometry of Hyperbolic Knots. Geometry & Topology, 28(5):2313–2343, 2024.
[41] Alex Davies, Petar Veličković, Lars Buesing, Sam Blackwell, Daniel Zheng, Nenad Tomašev, Richard Tanburn, Peter Battaglia, Charles Blundell, András Juhász, et al. Advancing Mathematics by Guiding Human Intuition with AI. Nature, 600(7887):70–74, 2021.
[42] Martin Davis. The Prehistory and Early History of Automated Deduction. Automation of Reasoning 1: Classical Papers on Computational Logic 1957-1966, pages 1–28, 1983.
[43] Leonardo de Moura and Sebastian Ullrich. The Lean 4 Theorem Prover and Programming Language. In André Platzer and Geoff Sutcliffe, editors, Automated Deduction - CADE 28 - 28th International Conference on Automated Deduction, Virtual Event, July 12-15, 2021, Proceedings, volume 12699 of Lecture Notes in Computer Science, pages 625–635. Springer, 2021. 23 SuperIntelligence - Safety & Alignment 2026 V3N3 Toward AGI Science - Math - ATP
[44] Leonardo Mendonça de Moura, Soonho Kong, Jeremy Avigad, Floris van Doorn, and Jakob von Raumer. The Lean Theorem Prover (System Description). In Amy P. Felty and Aart Middeldorp, editors, Automated Deduction - CADE-25 - 25th International Conference on Automated Deduction, Berlin, Germany, August 1-7, 2015, Proceedings, volume 9195 of Lecture Notes in Computer Science, pages 378–388. Springer, 2015.
[45] Bin Dong. On Mathematical Modeling in Image Reconstruction and Beyond. In Proc. Int. Cong. Math, volume 7, pages 5420–5449, 2022.
[46] Bin Dong, Xuhua He, Pengfei Jin, Felix Schremmer, and Qingchao Yu. Machine Learning Assisted Exploration for Affine Deligne–Lusztig Varieties. Peking Mathematical Journal, pages 1–50, 2024.
[47] Kefan Dong and Tengyu Ma. STP: Self-Play LLM Theorem Provers with Iterative Con- jecturing and Proving. In Forty-second International Conference on Machine Learning, 2025.
[48] Michael R Douglas and Kit Fraser-Taliente. Diffusion Models for Cayley Graphs. arXiv preprint arXiv:2503.05558, 2025.
[49] Oliver Dressler. Lean LSP MCP: Tools for Agentic Interaction with the Lean Theorem Prover. https://github.com/oOo0oOo/lean-lsp-mcp, 2025.
[50] Boyan Duan, Xiao Liang, Shuai Lu, Yaoxiang Wang, Yelong Shen, Kai-Wei Chang, Ying Nian Wu, Mao Yang, Weizhu Chen, and Yeyun Gong. Gold-Medal-Level Olympiad Geometry Solving with Efficient Heuristic Auxiliary Constructions. arXiv preprint arXiv:2512.00097, 2025.
[51] Harold Erbin and Riccardo Finotello. Machine Learning for Complete Intersection Calabi- Yau Manifolds: A Methodological Study. Phys. Rev. D, 103:126014, Jun 2021.
[52] Tony Feng. Eigenweights for arithmetic hirzebruch proportionality. arXiv preprint arXiv:2601.23245, 2026.
[53] Tony Feng, Junehyuk Jung, Sang-hyun Kim, Carlo Pagano, Sergei Gukov, Chiang-Chiang Tsai, David Woodruff, Adel Javanmard, Aryan Mokhtari, Dawsen Hwang, et al. Aletheia tackles FirstProof autonomously. arXiv preprint arXiv:2602.21201, 2026.
[54] Tony Feng, Trieu H Trinh, Garrett Bingham, Dawsen Hwang, Yuri Chervonyi, June- hyuk Jung, Joonkyung Lee, Carlo Pagano, Sang-hyun Kim, Federico Pasqualotto, et al. Towards Autonomous Mathematics Research. arXiv preprint arXiv:2602.10177, 2026.
[55] Deborah Ferreira and André Freitas. Premise Selection in Natural Language Mathematical Texts. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 7365–7374. Association for Computational Linguistics, 2020.
[56] Emily First, Markus N Rabe, Talia Ringer, and Yuriy Brun. Baldur: Whole-Proof Gen- eration and Repair with Large Language Models. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Soft- ware Engineering, pages 1229–1241, 2023.
[57] Simon Frieder, Luca Pinchetti, Chevalier Chevalier, Ryan-Rhys Griffiths, Tommaso Sal- vatori, Thomas Lukasiewicz, Philipp Petersen, and Julius Berner. Mathematical Capa- bilities of ChatGPT. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36, pages 27699–27744. Curran Associates, Inc., 2023. 24 SuperIntelligence - Safety & Alignment 2026 V3N3 Toward AGI Science - Math - ATP
[58] Siddhartha Gadgil, Anand Rao Tadipatri, Ayush Agrawal, Ashvni Narayanan, and Navin Goyal. Towards Automating Formalisation of Theorem Statements Using Large Language Models. In 36th Conference on Neural Information Processing Systems (NeurIPS 2022) Workshop on MATH-AI, 2022.
[59] Guoxiong Gao, Haocheng Ju, Jiedong Jiang, Zihan Qin, and Bin Dong. A Semantic Search Engine for Mathlib4. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024, pages 8001– 8013, Miami, Florida, USA, November 2024. Association for Computational Linguistics.
[60] Guoxiong Gao, Yutong Wang, Jiedong Jiang, Qi Gao, Zihan Qin, Tianyi Xu, and Bin Dong. Herald: A Natural Language Annotated Lean 4 Dataset. In The Thirteenth International Conference on Learning Representations, 2025.
[61] Liangcai Gao, Ke Yuan, Yuehan Wang, Zhuoren Jiang, and Zhi Tang. The Math Retrieval System of ICST for NTCIR-12 MathIR Task. In Noriko Kando, Tetsuya Sakai, and Mark Sanderson, editors, Proceedings of the 12th NTCIR Conference on Evaluation of Information Access Technologies, National Center of Sciences, Tokyo, Japan, June 7-10, 2016. National Institute of Informatics (NII), 2016.
[62] Bogdan Georgiev, Javier Gómez-Serrano, Terence Tao, and Adam Zsolt Wagner. Mathe- matical Exploration and Discovery at Scale. arXiv preprint arXiv:2511.02864, 2025.
[63] Elliot Glazer, Ege Erdil, Tamay Besiroglu, Diego Chicharro, Evan Chen, Alex Gunning, Caroline Falkman Olsson, Jean-Stanislas Denain, Anson Ho, Emily de Oliveira Santos, et al. Frontiermath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI. arXiv preprint arXiv:2411.04872, 2024.
[64] Fabian Gloeckle, Alex Gu, Gabriel Synnaeve, and Amaury Hayat. Reinforcement Learning for Hierarchical Proof Generation in Lean 4. In The 5th Workshop on Mathematical Reasoning and AI at NeurIPS 2025, 2025.
[65] Fabian Gloeckle, Jannis Limperg, Gabriel Synnaeve, and Amaury Hayat. ABEL: Sam- ple Efficient Online Reinforcement Learning for Neural Theorem Proving. In The 4th Workshop on Mathematical Reasoning and AI at NeurIPS’24, 2024.
[66] Kurt Gödel. On Formally Undecidable Propositions of Principia Mathematica and Related Systems. 1931.
[67] Zhibin Gou, Zhihong Shao, Yeyun Gong, Yujiu Yang, Minlie Huang, Nan Duan, Weizhu Chen, et al. ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solv- ing. In The Twelfth International Conference on Learning Representations.
[68] Sergei Gukov, James Halverson, Ciprian Manolescu, and Fabian Ruehle. Searching for Ribbons with Machine Learning. Machine Learning: Science and Technology, 6(2):025065, jun 2025.
[69] Sergei Gukov and Rak-Kyeong Seong. Machine Learning BPS Spectra and the Gap Conjecture. Phys. Rev. D, 110:046016, Aug 2024.
[70] Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, et al. DeepSeek-R1 Incentivizes Reasoning in LLMs Through Reinforcement Learning. Nature, 645(8081):633–638, 2025.
[71] Baran Hashemi, Roderic Guigo Corominas, and Alessandro Giacchetto. Can Transformers Do Enumerative Geometry? In The Thirteenth International Conference on Learning Representations, 2025. 25 SuperIntelligence - Safety & Alignment 2026 V3N3 Toward AGI Science - Math - ATP
[72] Yang-Hui He, Kyu-Hwan Lee, Thomas Oliver, and Alexey Pozdnyakov. Murmurations of Elliptic Curves. Experimental Mathematics, 34(3):528–540, 2025.
[73] Yiming He, Jia Zou, Xiaokai Zhang, Na Zhu, and Tuo Leng. FGeo-TP: A Language Model-Enhanced Solver for Euclidean Geometry Problems. Symmetry, 16(4):421, 2024.
[74] Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring Mathematical Problem Solving With the MATH Dataset. NeurIPS, 2021.
[75] Kryštof Hoder and Andrei Voronkov. Sine Qua Non for Large Theory Reasoning. In International Conference on Automated Deduction, pages 299–314. Springer, 2011.
[76] Daniel Huang, Prafulla Dhariwal, Dawn Song, and Ilya Sutskever. Gamepad: A Learning Environment for Theorem Proving. arXiv preprint arXiv:1806.00608, 2018.
[77] Yanxing Huang, Xinling Jin, Sijie Liang, Fuwen Luo, Peng Li, and Yang Liu. FormaRL: Enhancing Autoformalization with No Labeled Data. In Second Conference on Language Modeling, 2025.
[78] Yichen Huang and Lin F Yang. Winning Gold at IMO 2025 with a Model-Agnostic Verification-and-Refinement Pipeline. arXiv preprint arXiv:2507.15855, 2025.
[79] Thomas Hubert, Rishi Mehta, Laurent Sartran, Miklós Z Horváth, Goran Žužić, Eric Wieser, Aja Huang, Julian Schrittwieser, Yannick Schroecker, Hussain Masoom, et al. Olympiad-Level Formal Mathematical Reasoning with Reinforcement Learning. Nature, pages 1–3, 2025.
[80] Geoffrey Irving, Christian Szegedy, Alexander A. Alemi, Niklas Eén, François Chollet, and Josef Urban. DeepMath: Deep Sequence Models for Premise Selection. Advances in Neural Information Processing Systems (NeurIPS), 29, 2016.
[81] Uijeong Jang and Ernest K Ryu. Point Convergence of Nesterov’s Accelerated Gradient Method: An AI-Assisted Proof. arXiv preprint arXiv:2510.23513, 2025.
[82] Albert Q Jiang, Wenda Li, and Mateja Jamnik. Multilingual Mathematical Autoformal- ization. arXiv preprint arXiv:2311.03755, 2023.
[83] Albert Qiaochu Jiang, Wenda Li, Szymon Tworkowski, Konrad Czechowski, Tomasz Odrzygóźdź, Piotr Miloś, Yuhuai Wu, and Mateja Jamnik. Thor: Wielding Hammers to Integrate Language Models and Automated Theorem Provers. Advances in Neural Information Processing Systems, 35:8360–8373, 2022.
[84] Albert Qiaochu Jiang, Sean Welleck, Jin Peng Zhou, Timothee Lacroix, Jiacheng Liu, Wenda Li, Mateja Jamnik, Guillaume Lample, and Yuhuai Wu. Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs. In The Eleventh International Conference on Learning Representations, 2023.
[85] Jiedong Jiang, Wanyi He, Yuefeng Wang, Guoxiong Gao, Yongle Hu, Jingting Wang, Nailing Guan, Peihao Wu, Chunbo Dai, Liang Xiao, et al. FATE: A Formal Benchmark Series for Frontier Algebra of Multiple Difficulty Levels. arXiv preprint arXiv:2511.02872, 2025.
[86] Haocheng Ju, Leheng Chen, Peihao Wu, Bryan Dai, and Bin Dong. Matlas: A Semantic Search Engine for Mathematics. arXiv preprint arXiv:2604.17484, 2026. 26 SuperIntelligence - Safety & Alignment 2026 V3N3 Toward AGI Science - Math - ATP
[87] Haocheng Ju, Guoxiong Gao, Jiedong Jiang, Bin Wu, Zeming Sun, Leheng Chen, Yutong Wang, Yuefeng Wang, Zichen Wang, Wanyi He, et al. Automated Conjecture Resolution with Formal Verification. arXiv preprint arXiv:2604.03789, 2026.
[88] George Em Karniadakis, Ioannis G Kevrekidis, Lu Lu, Paris Perdikaris, Sifan Wang, and Liu Yang. Physics-Informed Machine Learning. Nature Reviews Physics, 3(6):422–440, 2021.
[89] Omar Khattab and Matei Zaharia. ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR), pages 39–48, 2020.
[90] Daniel Klaewer and Lorenz Schlechter. Machine Learning Line Bundle Cohomologies of Hypersurfaces in Toric Varieties. Physics Letters B, 789:438–443, 2019.
[91] Alexander Koldobsky. Intersection Bodies in R4. Advances in Mathematics, 136(1):1–14, 1998.
[92] Alexander Koldobsky. Intersection Bodies, Positive Definite Distributions, and the Busemann-Petty Problem. American journal of mathematics, 120(4):827–840, 1998.
[93] János Kollár. Non-Quasi-Projective Moduli Spaces. Annals of mathematics, pages 1077– 1096, 2006.
[94] Kriste Krstovski and David M Blei. Equation Embeddings. arXiv preprint arXiv:1803.09123, 2018.
[95] Guillaume Lample, Timothee Lacroix, Marie-Anne Lachaux, Aurelien Rodriguez, Amaury Hayat, Thibaut Lavril, Gabriel Ebner, and Xavier Martinet. Hypertree Proof Search for Neural Theorem Proving. Advances in neural information processing systems, 35:26337– 26349, 2022.
[96] Thomas Lanard and Alberto Mı́nguez. An Algorithm for Aubert–Zelevinsky Dualitya la M {oe} glin–Waldspurger. arXiv preprint arXiv:2509.13231, 2025.
[97] Robert Tjarko Lange, Yuki Imajuku, and Edoardo Cetin. Shinkaevolve: Towards Open- Ended and Sample-Efficient Program Evolution. arXiv preprint arXiv:2509.19349, 2025.
[98] Joonkyung Lee and Jaehyeon Seo. Lower bounds for multivariate independence polyno- mials and their generalisations. arXiv preprint arXiv:2602.02450, 2026.
[99] Kyu-Hwan Lee. Data-Scientific Study of Kronecker Coefficients. Experimental Mathemat- ics, 0(0):1–14, 2025.
[100] Kyu-Hwan Lee and Seewoo Lee. Machines Learn Number Fields, But How? The Case of Galois Groups. arXiv preprint arXiv:2508.06670, 2025.
[101] Kyu-Hwan Lee, Thomas Oliver, and Alexey Pozdnyakov. Murmurations of Dirichlet Characters. International Mathematics Research Notices, 2025(1):rnae277, 2025.
[102] Chenyi Li, Ziyu Wang, Wanyi He, Yuxuan Wu, Shengyang Xu, and Zaiwen Wen. Formal- ization of Convergence Rates of Four First-Order Algorithms for Convex Optimization. Journal of Automated Reasoning, 69(4):28, 2025. 27 SuperIntelligence - Safety & Alignment 2026 V3N3 Toward AGI Science - Math - ATP
[103] Jia Li, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Huang, Kashif Rasul, Longhui Yu, Albert Q Jiang, Ziju Shen, et al. Numinamath: The Largest Public Dataset in Ai4maths With 860k Pairs of Competition Math Problems and Solu- tions. Hugging Face repository, 13(9):9, 2024.
[104] Yang Li, Dong Du, Linfeng Song, Chen Li, Weikang Wang, Tao Yang, and Haitao Mi. HunyuanProver: A Scalable Data Synthesis Framework and Guided Tree Search for Au- tomated Theorem Proving. arXiv preprint arXiv:2412.20735, 2024.
[105] Zhaoyu Li, Jialiang Sun, Logan Murphy, Qidong Su, Zenan Li, Xian Zhang, Kaiyu Yang, and Xujie Si. A Survey on Deep Learning for Theorem Proving. In First Conference on Language Modeling, 2024.
[106] Yong Lin, Shange Tang, Bohan Lyu, Jiayun Wu, Hongzhou Lin, Kaiyu Yang, Jia Li, Mengzhou Xia, Danqi Chen, Sanjeev Arora, et al. Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving. arXiv preprint arXiv:2502.07640, 2025.
[107] Gang Liu, Yihan Zhu, Jie Chen, and Meng Jiang. Scientific Algorithm Discovery by Augmenting Alphaevolve with Deep Research. arXiv preprint arXiv:2510.06056, 2025.
[108] Junqi Liu, Zihao Zhou, Zekai Zhu, Marco Dos Santos, Weikun He, Jiawei Liu, Ran Wang, Yunzhou Xie, Junqiao Zhao, Qiufeng Wang, et al. Numina-Lean-Agent: An Open and General Agentic Reasoning System for Formal Mathematics. arXiv preprint arXiv:2601.14027, 2026.
[109] Qi Liu, Xinhao Zheng, Xudong Lu, Qinxiang Cao, and Junchi Yan. Rethinking and Im- proving Autoformalization: Towards a Faithful Metric and a Dependency Retrieval-Based Approach. In The Thirteenth International Conference on Learning Representations, 2025.
[110] Xiaoyang Liu, Tao Zhu, Zineng Dong, Yuntian Liu, Qingfeng Guo, Zhaoxuan Liu, Yu Chen, and Tao Luo. ASSESS: A Semantic and Structural Evaluation Framework for Statement Similarity. arXiv preprint arXiv:2509.22246, 2025.
[111] Jialin Lu, Kye Emond, Weiran Sun, and Wuyang Chen. Lean Finder: Semantic Search for Mathlib That Understands User Intents. In 2nd AI for Math Workshop @ ICML 2025, 2025.
[112] Jianqiao Lu, Yingjia Wan, Zhengying Liu, Yinya Huang, Jing Xiong, Chengwu Liu, Jian- hao Shen, Hui Jin, Jipeng Zhang, Haiming Wang, et al. Process-Driven Autoformalization in Lean 4. arXiv preprint arXiv:2406.01940, 2024.
[113] Behrooz Mansouri, Vı́t Novotný, Anurag Agarwal, Douglas W. Oard, and Richard Zanibbi. Overview of ARQMath-3 (2022): Third CLEF Lab on Answer Retrieval for Questions on Math (Working Notes). In Proceedings of CLEF 2022 Working Notes (Con- ference and Labs of the Evaluation Forum), 2022.
[114] Behrooz Mansouri, Douglas W. Oard, and Richard Zanibbi. Contextualized Formula Search Using Math Abstract Meaning Representation. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, CIKM ’22, page 4329–4333, New York, NY, USA, 2022. Association for Computing Machinery.
[115] Behrooz Mansouri, Shaurya Rohatgi, Douglas W. Oard, Jian Wu, C. Lee Giles, and Richard Zanibbi. Tangent-CFT: An Embedding Model for Mathematical Formulas. In Proceedings of the 2019 ACM SIGIR International Conference on Theory of Information Retrieval (ICTIR), pages 11–18. ACM, 2019. 28 SuperIntelligence - Safety & Alignment 2026 V3N3 Toward AGI Science - Math - ATP
[116] Behrooz Mansouri, Richard Zanibbi, Douglas W. Oard, and Anurag Agarwal. Overview of ARQMath-2 (2021): Second CLEF Lab on Answer Retrieval for Questions on Math. In K. Selçuk Candan, Bogdan Ionescu, Lorraine Goeuriot, Birger Larsen, Henning Müller, Alexis Joly, Maria Maistro, Florina Piroi, Guglielmo Faggioli, and Nicola Ferro, editors, Experimental IR Meets Multilinguality, Multimodality, and Interaction, pages 215–238, Cham, 2021. Springer International Publishing.
[117] Math, Inc. Introducing Gauss, an Agent for Autoformalization. https://www.math.inc/ gauss, 2025.
[118] Jia Meng and Lawrence C Paulson. Lightweight Relevance Filtering for Machine- Generated Resolution Problems. Journal of Applied Logic, 7(1):41–57, 2009.
[119] A. Newell and H. Simon. The logic theory machine–A complex information processing system. IRE Transactions on Information Theory, 2(3):61–79, 1956.
[120] Alexander Novikov, Ngân Vũ, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco JR Ruiz, Abbas Mehrabian, et al. AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery. arXiv preprint arXiv:2506.13131, 2025.
[121] Aditya Paliwal, Sarah Loos, Markus Rabe, Kshitij Bansal, and Christian Szegedy. Graph Representations for Higher-Order Logic and Theorem Proving. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 2967–2974, 2020.
[122] Anand Patel. The simplicity of the hodge bundle. arXiv preprint arXiv:2603.19052, 2026.
[123] Shuai Peng, Ke Yuan, Liangcai Gao, and Zhi Tang. MathBERT: A Pre-trained Model for Mathematical Formula Understanding. arXiv preprint arXiv:2105.00377, 2021.
[124] Auguste Poiroux, Gail Weiss, Viktor Kunčak, and Antoine Bosselut. Reliable Evaluation and Benchmarks for Statement Autoformalization. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 17958–17980, 2025.
[125] Stanislas Polu, Jesse Michael Han, Kunhao Zheng, Mantas Baksys, Igor Babuschkin, and Ilya Sutskever. Formal Mathematics Statement Curriculum Learning. In The Eleventh International Conference on Learning Representations, 2023.
[126] Stanislas Polu and Ilya Sutskever. Generative Language Modeling for Automated Theorem Proving. arXiv preprint arXiv:2009.03393, 2020.
[127] Anja Reusch, Maik Thiele, and Wolfgang Lehner. TU DBS in the ARQMath Lab 2021, CLEF. In Working Notes of CLEF 2021, pages 107–124, 2021.
[128] Stephen Robertson, S. Walker, S. Jones, M. M. Hancock-Beaulieu, and M. Gatford. Okapi at TREC-3. In Overview of the Third Text REtrieval Conference (TREC-3), pages 109– 126. Gaithersburg, MD: NIST, January 1995.
[129] Stephen Robertson and Hugo Zaragoza. The Probabilistic Relevance Framework: BM25 and Beyond. Found. Trends Inf. Retr., 3(4):333–389, apr 2009.
[130] Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Ba- log, M Pawan Kumar, Emilien Dupont, Francisco JR Ruiz, Jordan S Ellenberg, Pengming Wang, Omar Fawzi, et al. Mathematical Discoveries from Program Search with Large Language Models. Nature, 625(7995):468–475, 2024. 29 SuperIntelligence - Safety & Alignment 2026 V3N3 Toward AGI Science - Math - ATP
[131] Gerard Salton and Christopher Buckley. Term-Weighting Approaches in Automatic Text Retrieval. Information Processing & Management, 24(5):513–523, 1988.
[132] Georg Schumacher and Hajime Tsuji. Quasi-Projectivity of Moduli Spaces of Polarized Varieties. Annals of mathematics, pages 597–639, 2004.
[133] Georg Schumacher and Hajime Tsuji. Retraction: “Quasi-Projectivity of Moduli Spaces of Polarized Varieties”. Annals of Mathematics, 198(3):1305–1305, 2023.
[134] Michael Shalyt, Uri Seligmann, Itay Beit Halachmi, Ofir David, Rotem Elimelech, and Ido Kaminer. Unsupervised Discovery of Formulas for Mathematical Constants. Advances in Neural Information Processing Systems, 37:113156–113190, 2024.
[135] Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the Limits of Math- ematical Reasoning in Open Language Models. arXiv preprint arXiv:2402.03300, 2024.
[136] Asankhaya Sharma. OpenEvolve: An Open-Source Evolutionary Coding Agent, 2025.
[137] Ali Shehper, Anibal M. Medina-Mardones, Lucas Fagan, Bartlomiej Lewandowski, Angus Gruen, Yang Qiu, Piotr Kucharski, Zhenghan Wang, and Sergei Gukov. What Makes Math Problems Hard for Reinforcement Learning: A Case Study. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025.
[138] Ziju Shen, Naohao Huang, Fanyi Yang, Yutong Wang, Guoxiong Gao, Tianyi Xu, Jiedong Jiang, Wanyi He, Pu Yang, Mengzhou Sun, et al. REAL-Prover: Retrieval Augmented Lean Prover for Mathematical Reasoning. arXiv preprint arXiv:2505.20613, 2025.
[139] Stephen G Simpson. Partial Realizations of Hilbert’s Program. The Journal of Symbolic Logic, 53(2):349–363, 1988.
[140] Grzegorz Swirszcz, Adam Zsolt Wagner, Geordie Williamson, Sam Blackwell, Bogdan Georgiev, Alex Davies, Ali Eslami, Sebastien Racaniere, Theophane Weber, and Pushmeet Kohli. Advancing Geometry with AI: Multi-Agent Generation of Polytopes. arXiv preprint arXiv:2502.05199, 2025.
[141] Yicheng Tao, Haotian Liu, Shanwen Wang, and Hongteng Xu. Learning an Effective Premise Retrieval Model for Efficient Mathematical Formalization. In Proceedings of the 2nd AI for Math Workshop at ICML 2025, 2025.
[142] Qwen Team. QwQ-32B: Embracing the Power of Reinforcement Learning, March 2025.
[143] Trieu H Trinh, Yuhuai Wu, Quoc V Le, He He, and Thang Luong. Solving Olympiad Geometry without Human Demonstrations. Nature, 625(7995):476–482, 2024.
[144] S. Ulam. John von Neumann 1903–1957. Bulletin of the American Mathematical Society, 64(3.P2):1 – 49, 1958.
[145] Stanislaw M. Ulam. Adventures of a Mathematician. Charles Scribner’s Sons, New York, 1976.
[146] Sumanth Varambally, Thomas Voice, Yanchao Sun, Zhifeng Chen, Rose Yu, and Ke Ye. Hilbert: Recursively Building Formal Proofs with Informal Reasoning. arXiv preprint arXiv:2509.22819, 2025.
[147] Adam Zsolt Wagner. Constructions in Combinatorics via Neural Networks. arXiv preprint arXiv:2104.14516, 2021. 30 SuperIntelligence - Safety & Alignment 2026 V3N3 Toward AGI Science - Math - ATP
[148] Haiming Wang, Mert Unsal, Xiaohan Lin, Mantas Baksys, Junqi Liu, Marco Dos Santos, Flood Sung, Marina Vinyes, Zhenzhe Ying, Zekai Zhu, et al. Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning. arXiv preprint arXiv:2504.11354, 2025.
[149] Haiming Wang, Huajian Xin, Chuanyang Zheng, Zhengying Liu, Qingxing Cao, Yinya Huang, Jing Xiong, Han Shi, Enze Xie, Jian Yin, Zhenguo Li, and Xiaodan Liang. LEGO- Prover: Neural Theorem Proving with Growing Libraries. In The Twelfth International Conference on Learning Representations, 2024.
[150] Haiming Wang, Ye Yuan, Zhengying Liu, Jianhao Shen, Yichun Yin, Jing Xiong, Enze Xie, Han Shi, Yujun Li, Lin Li, Jian Yin, Zhenguo Li, and Xiaodan Liang. DT-solver: Auto- mated theorem proving with dynamic-tree sampling guided by proof-level value function. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12632–12646, Toronto, Canada, July 2023. Association for Computational Linguistics.
[151] Hanyu Wang, Ruohan Xie, Yutong Wang, Guoxiong Gao, Xintao Yu, and Bin Dong. Aria: An Agent For Retrieval and Iterative Auto-Formalization via Dependency Graph. arXiv preprint arXiv:2510.04520, 2025.
[152] Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. A Survey on Large Language Model Based Autonomous Agents. Frontiers of Computer Science, 18(6):186345, 2024.
[153] Mingzhe Wang, Yihe Tang, Jian Wang, and Jia Deng. Premise Selection for Theorem Proving by Deep Graph Embedding. Advances in Neural Information Processing Systems (NeurIPS), 30, 2017.
[154] Qingxiang Wang, Chad Brown, Cezary Kaliszyk, and Josef Urban. Exploration of Neural Machine Translation in Autoformalization of Mathematics in Mizar. In Proceedings of the 9th ACM SIGPLAN International Conference on Certified Programs and Proofs, pages 85–98, 2020.
[155] Qingxiang Wang, Cezary Kaliszyk, and Josef Urban. First Experiments with Neural Translation of Informal to Formal Mathematics. In International Conference on Intelligent Computer Mathematics, pages 255–270. Springer, 2018.
[156] Zichen Wang, Anjie Dong, and Zaiwen Wen. Tree-Based Premise Selection for Lean4. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025.
[157] Ziyu Wang, Bowen Yang, Shihao Zhou, Chenyi Li, Yuan Zhang, Bin Dong, and Zaiwen Wen. Translating Informal Proofs into Formal Proofs Using a Chain of States. arXiv preprint arXiv:2512.10317, 2025.
[158] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. Advances in neural information processing systems, 35:24824–24837, 2022.
[159] E Weinan et al. The Dawning of a New Era in Applied Mathematics. Notices of the American Mathematical Society, 68(4):565–571, 2021.
[160] Wu Wen-Tsun. Basic Principles of Mechanical Theorem Proving in Elementary Geome- tries. Journal of automated Reasoning, 2(3):221–252, 1986. 31 SuperIntelligence - Safety & Alignment 2026 V3N3 Toward AGI Science - Math - ATP
[161] Daniel Whalen. Holophrasm: A Neural Automated Theorem Prover for Higher-Order Logic. arXiv preprint arXiv:1608.02644, 2016.
[162] Wen-tsün Wu. Mechanical Theorem Proving in Geometries: Basic Principles. Springer Science & Business Media, 2012.
[163] Wenjun Wu and Xiaoshan Gao. Mathematics Mechanization and Applications After Thirty Years. Frontiers of Computer Science in China, 1(1):1–8, 2007.
[164] Yuhuai Wu, Albert Qiaochu Jiang, Wenda Li, Markus Rabe, Charles Staats, Mateja Jam- nik, and Christian Szegedy. Autoformalization with Large Language Models. Advances in neural information processing systems, 35:32353–32368, 2022.
[165] Huajian Xin, Daya Guo, Zhihong Shao, Zhizhou Ren, Qihao Zhu, Bo Liu, Chong Ruan, Wenda Li, and Xiaodan Liang. Deepseek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data. arXiv preprint arXiv:2405.14333, 2024.
[166] Huajian Xin, Z.Z. Ren, Junxiao Song, Zhihong Shao, Wanjia Zhao, Haocheng Wang, Bo Liu, Liyue Zhang, Xuan Lu, Qiushi Du, Wenjun Gao, Haowei Zhang, Qihao Zhu, Dejian Yang, Zhibin Gou, Z.F. Wu, Fuli Luo, and Chong Ruan. DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search. In The Thirteenth International Conference on Learning Representations, 2025.
[167] Ran Xin, Chenguang Xi, Jie Yang, Feng Chen, Hang Wu, Xia Xiao, Yifan Sun, Shen Zheng, and Ming Ding. Bfs-Prover: Scalable Best-First Tree Search for LLM-Based Au- tomatic Theorem Proving. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 32588–32599, 2025.
[168] Ran Xin, Zeyu Zheng, Yanchen Nie, Kun Yuan, and Xia Xiao. Scaling Up Multi- Turn Off-Policy RL and Multi-Agent Tree Search for LLM Step-Provers. arXiv preprint arXiv:2509.06493, 2025.
[169] Yu Xuejun, Jianyuan Zhong, Zijin Feng, Pengyi Zhai, Roozbeh Yousefzadeh, Wei Chong Ng, Haoxiong Liu, Ziyi Shou, Jing Xiong, Yudong Zhou, et al. Mathesis: Towards Formal Theorem Proving from Natural Languages. arXiv preprint arXiv:2506.07047, 2025.
[170] An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 Technical Report. arXiv preprint arXiv:2505.09388, 2025.
[171] An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li, Dayi- heng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, et al. Qwen2. 5-Math Techni- cal Report: Toward Mathematical Expert Model via Self-Improvement. arXiv preprint arXiv:2409.12122, 2024.
[172] Kaiyu Yang, Aidan M Swope, Alex Gu, Rahul Chalamala, Peiyang Song, Shixing Yu, Saad Godil, Ryan Prenger, and Anima Anandkumar. LeanDojo: Theorem Proving with Retrieval-Augmented Language Models. In Thirty-seventh Conference on Neural Infor- mation Processing Systems Datasets and Benchmarks Track, 2023.
[173] Huaiyuan Ying, Zijian Wu, Yihan Geng, Jiayu Wang, Dahua Lin, and Kai Chen. Lean Workbook: A Large-Scale Lean Problem Set Formalized from Natural Language Math Problems. Advances in Neural Information Processing Systems, 37:105848–105863, 2024.
[174] Huaiyuan Ying, Shuo Zhang, Linyang Li, Zhejian Zhou, Yunfan Shao, Zhaoye Fei, Yichuan Ma, Jiawei Hong, Kuikun Liu, Ziyi Wang, et al. Internlm-Math: Open Math Large Language Models Toward Verifiable Reasoning. arXiv preprint arXiv:2402.06332, 2024. 32 SuperIntelligence - Safety & Alignment 2026 V3N3 Toward AGI Science - Math - ATP
[175] Richard Zanibbi and Dorothea Blostein. Recognition and Retrieval of Mathematical Ex- pressions. Int. J. Document Anal. Recognit., 15(4):331–357, 2012.
[176] Richard Zanibbi, Douglas W. Oard, Anurag Agarwal, and Behrooz Mansouri. Overview of ARQMath 2020: CLEF Lab on Answer Retrieval for Questions on Math. In Avi Arampatzis, Evangelos Kanoulas, Theodora Tsikrika, Stefanos Vrochidis, Hideo Joho, Christina Lioma, Carsten Eickhoff, Aurélie Névéol, Linda Cappellato, and Nicola Ferro, editors, Experimental IR Meets Multilinguality, Multimodality, and Interaction, pages 169–193, Cham, 2020. Springer International Publishing.
[177] Chi Zhang, Jiajun Song, Siyu Li, Yitao Liang, Yuxi Ma, Wei Wang, Yixin Zhu, and Song- Chun Zhu. Proposing and Solving Olympiad Geometry with Guided Tree Search. arXiv preprint arXiv:2412.10673, 2024.
[178] Gaoyong Zhang. Intersection Bodies and the Busemann-Petty Inequalities in R4. Annals of Mathematics, 140(2):331–346, 1994.
[179] Gaoyong Zhang. A Positive Solution to the Busemann-Petty Problem in R4. Annals of Mathematics, pages 535–543, 1999.
[180] Jingyuan Zhang, Qi Wang, Xingguang Ji, Yahui Liu, Yang Yue, Fuzheng Zhang, Di Zhang, Guorui Zhou, and Kun Gai. Leanabell-Prover: Posttraining Scaling in Formal Reasoning. arXiv preprint arXiv:2504.06122, 2025.
[181] Xiaokai Zhang, Na Zhu, Yiming He, Jia Zou, Qike Huang, Xiaoxiao Jin, Yanjun Guo, Chenyang Mao, Zhe Zhu, Dengfeng Yue, et al. FormalGeo: The First Step toward Human- Like IMO-Level Geometric Automated Reasoning. arXiv preprint arXiv:2310.18021, 2023.
[182] Xiaokai Zhang, Na Zhu, Cheng Qin, Yang Li, Zhenbing Zeng, and Tuo Leng. FGeo- HyperGNet: Geometric Problem Solving Integrating Formal Symbolic System and Hy- pergraph Neural Network. arXiv preprint arXiv:2402.11461, 2024.
[183] Kunhao Zheng, Jesse Michael Han, and Stanislas Polu. MiniF2F: A Cross-System Bench- mark for Formal Olympiad-Level Mathematics. arXiv preprint arXiv:2109.00110, 2021.
[184] Wei Zhong, Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. One Blade for One Purpose: Advancing Math Information Retrieval using Hybrid Search. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Infor- mation Retrieval, SIGIR ’23, page 141–151, New York, NY, USA, 2023. Association for Computing Machinery.
[185] Wei Zhong, Shaurya Rohatgi, Jian Wu, C. Lee Giles, and Richard Zanibbi. Accelerating Substructure Similarity Search for Formula Retrieval. In Advances in Information Re- trieval: 42nd European Conference on IR Research, ECIR 2020, Lisbon, Portugal, April 14–17, 2020, Proceedings, Part I, page 714–727, Berlin, Heidelberg, 2020. Springer-Verlag.
[186] Wei Zhong and Richard Zanibbi. Structural Similarity Search for Formulas Using Leaf- Root Paths in Operator Subtrees. In Leif Azzopardi, Benno Stein, Norbert Fuhr, Philipp Mayr, Claudia Hauff, and Djoerd Hiemstra, editors, Advances in Information Retrieval - 41st European Conference on IR Research, ECIR 2019, Cologne, Germany, April 14-18, 2019, Proceedings, Part I, volume 11437 of Lecture Notes in Computer Science, pages 116–129. Springer, 2019.
[187] Yichi Zhou, Jianqiu Zhao, Yongxin Zhang, Bohan Wang, Siran Wang, Luoxin Chen, Jiahui Wang, Haowei Chen, Allan Jie, Xinbo Zhang, et al. Solving Formal Math Problems by Decomposition and Iterative Reflection. arXiv preprint arXiv:2507.15225, 2025. 33 SuperIntelligence - Safety & Alignment 2026 V3N3 Toward AGI Science - Math - ATP
[188] Jia Zou, Xiaokai Zhang, Yiming He, Na Zhu, and Tuo Leng. FGeo-DRL: Deductive Reasoning for Geometric Problems through Deep Reinforcement Learning. Symmetry, 16(4):437, 2024.
[189] Nina Zubrilina. Murmurations. Inventiones mathematicae, 241(3):627–680, 2025. 34 SuperIntelligence - Safety & Alignment 2026 V3N3 Toward AGI Science - Math - ATP
Downloads
Published
How to Cite
Issue
Section
Categories
License
Copyright (c) 2026 Haocheng Ju, Bin Dong

This work is licensed under a Creative Commons Attribution 4.0 International License.