Mathematics in the Age of AI: From Theorem Proving to Theorem Mining
Some thoughts, intended to generate discussion!
Mathematics in the Age of AI: From Theorem Proving to Theorem Mining
There is a profound shift underway in mathematics—one that is not about faster proofs or larger collaborations, but about who (or what) is doing the mathematics.
Until now, most mathematical research—pure or applied—has followed a recognizable pattern. A researcher studies a body of literature, identifies a tractable gap, proves a result extending prior work, and publishes it. In engineering and applied mathematics, this often means sharpening bounds, extending stability results, or adapting known techniques to new models. This system has worked remarkably well. It trains researchers, produces cumulative knowledge, and sustains a professional ecosystem.
But much of this activity is, in a precise sense, incremental. And incremental reasoning is exactly the kind of thing that modern AI systems are beginning to do well.
The Automation of Incremental Mathematics
We are no longer speculating about a distant future. We are already close to a world in which a significant fraction of routine mathematical research can be automated.
Consider typical tasks:
Extending a known theorem under slightly weaker assumptions
Verifying technical lemmas in a long proof
Exploring parameter regimes in a model
Generating counterexamples or testing conjectures numerically
These are structured, local, and often guided by well-established heuristics. They resemble what machine learning systems already excel at: pattern recognition, search, and guided optimization.
What has changed in recent years is not just capability but infrastructure. Formal proof systems such as Lean, together with large formalized libraries, are turning mathematical reasoning into something that can be both checked and increasingly generated by machines. Once mathematics is formalized, proofs can be verified automatically, results can be composed reliably, and AI systems can search over structured arguments rather than informal text.
Toward “Theorem Mining”
A useful analogy may be drawn from cryptocurrency: instead of proof mining, we may see the rise of theorem mining.
In this model, large-scale AI systems continuously generate mathematical results: lemmas, propositions, conjectures, and connections across fields. These would live in evolving, queryable repositories rather than static papers.
As others have remarked, the role of the mathematician will begin to resemble that of an orchestra conductor.
A conductor does not produce sound directly. Instead, they shape the performance: choosing tempo, emphasizing certain voices, coordinating sections, and interpreting the piece. The musicians produce the notes. The conductor contributes structure, timing, and meaning.
Similarly, in a theorem-mining ecosystem, AI systems generate the notes: candidate theorems, proofs, counterexamples, and conjectures. The human mathematician acts as conductor—deciding direction, emphasizing themes, coordinating interactions between areas, and judging which results matter.
The output may be vast, redundant, and uneven in quality. But it could be comprehensive in a way that current human-driven mathematics is not.
Why Not Just Do Mathematics On Demand?
If AI becomes powerful enough, why not generate results on demand? Why build databases of theorems at all?
Mathematics is not just a collection of answers; it is a layered structure of conceptual frameworks. Major advances depend on prior frameworks:
General relativity relies on differential geometry and manifolds
Quantum mechanics relies on Hilbert spaces and spectral theory
Signal processing relies on harmonic analysis
Coding theory and cryptography rely on algebraic geometry and group theory
One way to understand the importance of such frameworks is to look at how they arise. Ideas that are almost trivial in simple settings often lead, through abstraction, to entirely new fields. The intermediate value theorem in one dimension—almost obvious when viewed geometrically—leads, when abstracted, to the foundations of topology. Similarly, what we now call Hilbert spaces were not initially conceived in full generality; the abstraction to infinite-dimensional inner product spaces was clarified and elevated in the work of John von Neumann, extending earlier, more concrete work of Hilbert on function spaces. In algebra, Cayley’s abstraction of permutations into the concept of a group transformed Galois’ insights into a general structural theory.
These examples illustrate a recurring pattern: mathematics advances not only by solving problems, but by creating the languages in which problems can be formulated.
These frameworks were not created on demand. They emerged from long lines of inquiry. Even with formal systems, correctness is not the same as conceptual insight. Abstractions often precede applications by decades.
This suggests that precomputation—in the form of theorem mining—may be essential. AI systems may need to build and maintain conceptual frameworks before answering specific questions. Even machines may require a kind of mathematical culture.
A Possible Division of Labor
AI systems generate, verify, and organize large bodies of results
Human researchers guide direction, shape abstractions, and evaluate meaning
This role includes something that is difficult to formalize, and perhaps even harder to automate: the judgment of depth and aesthetic value. Not all theorems are equal. Some results, even if technically correct, feel routine or inevitable; others reveal unexpected structure, unify disparate ideas, or introduce concepts that reshape an area. Mathematicians have long valued elegance—proofs that are not merely correct, but illuminating, economical, or surprising. It is not clear that such judgments can be reduced to formal criteria, or whether they will remain, at least in part, a distinctly human contribution. A system may generate thousands of correct theorems; it takes judgment to recognize the few that matter.
This emphasis on aesthetics is not new. Early in my career, Samuel Eilenberg—one of the founders of homological algebra—remarked to me that elegance should be the only criterion by which mathematics is judged. This is, of course, an extreme view. But it is one that has stayed with me, and it captures something essential about how mathematicians evaluate their work.
This raises a fundamental question: what happens to the current training model of mathematicians? If incremental work is automated, the traditional pathway may no longer define expertise.
A natural follow-up question is what this implies for training. If the mathematician’s role becomes closer to that of a conductor, how much must a conductor know about playing each instrument? In other words, which skills should we be teaching our students? Technical competence will remain essential, but perhaps the emphasis will shift toward recognizing structure, formulating the right abstractions, and developing the taste needed to judge depth, significance, and elegance.
The End of Journals?
If theorem mining and formal verification become central, it is difficult to see how traditional journals survive. One can make a stronger claim: mathematical journals, as we know them, are likely to disappear.
Journals serve several roles today:
Certification (peer review)
Dissemination
Curation
Career signaling
Each of these roles is being eroded. Formal verification replaces certification. Online repositories replace dissemination. AI and community filtering replace curation. And new forms of influence may replace traditional career signals.
Instead of static papers, we may see living mathematical structures: definitions as nodes, theorems as relations, proofs as composable objects, and explanations as overlays.
The unit of contribution shifts from the paper to the mathematical object itself.
Closing Thoughts
We may be moving toward a world in which mathematics is no longer primarily a human-driven enterprise of theorem proving, but a hybrid system of theorem mining, formal verification, and conceptual guidance.
The question is not whether AI will do mathematics—it already is.
The real questions are: What kind of mathematics will still require humans? What kind of mathematics will humans want to do? And perhaps: who, or what, will be writing the score?
Postscript
I should add a personal note. I have been interested in, and to some extent working on, artificial intelligence and neural networks since 1969, including writing a book on the subject in 1972. For many years, like others, I saw progress as uneven and often overhyped. Only recently have I become convinced that something fundamentally different is happening—that the current trajectory is real in a way that earlier waves were not.
I am also reminded of a conversation from much earlier. Around 1999, I suggested to my Rutgers colleague Felix Browder—then president of the American Mathematical Society—that mathematics might benefit from an effort analogous to the genome project, or to large-scale initiatives in physics such as plans for a supercollider. My thought was that we should aim to build a comprehensive, hyperlinked database of mathematical results. The idea, at the time, did not gain traction.
In retrospect, it may simply have been too early. What now seems technically and conceptually within reach required not just an idea, but the surrounding ecosystem—computational, cultural, and institutional—that did not yet exist. The current moment feels different.
Acknowledgment
This essay has itself benefited from the tools it discusses. In particular, I used ChatGPT to help polish the exposition and refine the presentation of ideas. The views expressed here are, of course, my own, but the process of writing was already a small example of the kind of human–AI interaction described above. I would be very interested in hearing reactions, criticisms, or alternative perspectives from readers.


You raise very good questions. I will add another one:
-- What is the impact of this transition on the younger generations?
I can see many senior mathematicians becoming excellent conductors (in the spirit of your analogy), but it is less clear that this new system will be conducive to the education of a new generation of senior conductors.
A source of friction - grains of sand in the use of AI for mathematics - is that chatbots generate errors. Humans do to, and we have a mental and institutional framework for dealing with possible errors.
But theorem-mining may overwhelm the guardrails. Chatbots can generate so many theorems, that finding the hallucinated ones takes up more time that looking for correct theorems the old-fashioned way does. I have read studies saying that AI increases personal productivity (in science, number of publications per researcher), while not necessarily increasing collective productivity - that is, scientific progress, which is harder to measure.
A version of this already occurs with the increase in the number of published papers, including those on questionable journals. We may be missing some jewels but who has time to read them? AI may exponentiate this problem.