The future of the human mathematician

· A Naked Rhyme Jam ·

14 min read Original article ↗

Archimedes, the aristocrat, would draw figures in the oil that was rubbed on his skin after bathing. Ramanujan, living on a small salary in India, would write his equations and estimates on a small chalkboard: he would write so much that his elbow, which he used to erase, would hurt by the end of the day. Hilbert set up a blackboard in his backyard to enjoy the pleasures of the outdoors while he worked. I have a reMarkable tablet that feels a lot like paper and I snare a table at a cafe and draw pictures and arrows and the occasional statement of a theorem. While I am of course no Hilbert or Ramanujan we all use more or less the same technology: our minds along with a passive tableau for our thoughts, some way of stepping back and recording and sharing.

Even before AI there have been tweaks to the ancient art of thinking and starting at a blackboard or a piece of paper. Gauss computed by hand when he wanted a table of primes; now we can make one into the trillions in a matter of minutes. Computers have proven their worth for mathematical experiments and computations, and there have been a handful of famous computer-assisted proofs, like Kepler’s conjecture and the Four Color Theorem. But these have been the exception rather than the rule. Grigory Perelman sat at his desk every day at the Steklov Institute, proving the Poincare conjecture; Andrew Wiles would hide out in his attic so that he could prove Fermat’s Last Theorem. Neither relied on anything more than paper and pencil to prove the greatest two theorems of the last fifty years.

While some of the work is tedious and exhausting, a lot of the work is fun. It’s fun to sit with a stack of paper and stare off into the distance trying to solve a problem. To scribble some notes, day after day, until the pieces start to come together. To meet with a colleague and hash out our ideas and approaches and fill a blackboard with pictures and formulas and sometimes even a theorem and proof. Even the writing and proofreading is its own kind of Type II fun: the rhythm of the detail-oriented work, the realization of a missing step that has to be filled in, the satisfaction of a paper brought to near perfection.

Now there is ChatGPT and Gemini and Claude (and Kimi and Grok) and we’re eyeing them uncomfortably, wondering if we want to bring them into the fun, and whether there will be any fun left for us if we do—and even if we don’t. Some of us are all in on the use of AI—one colleague told me he had “passed through the five stages of grief” to reach acceptance, while others are calling for a profession-wide rejection of it. A silent majority of us, I believe, are anxiously looking around to see what our peers are doing and wondering what the rules are before diving too far into AI (or a boycott). Outsiders who see the headlines would be surprised by how much of our working time we spend as if AI was not here at all.

While we love to pretend that there’s some deep meaning and purpose to what we do—and there might even be one or more—the secret is that there’s only two things we can all agree on:

  1. It’s fun to do mathematics. It’s enjoyable, satisfying, and intrinsically rewarding.

  2. We earn social and financial capital (make friends and make a living) by proving theorems and writing papers and speaking about our work.

If it wasn’t fun, we wouldn’t want to do it—we’d go into finance or law or something that paid a lot more. If it wasn’t supported, we’d have to do something else, at least as a day job. We’ve had a good run getting paid to do what we enjoy and now nobody knows whether this will continue.

Is mathematics chess or medicine? The best chess players have been computers for twenty years, but chess is thriving as a human game. You might have thought we want to see the deepest and most well thought out games, but our human capital has gone towards what two humans can do, with strict time limits and rules against outside assistance, or even the use of a second board. Chess is preserved because we can shut out the machines as the rules of the game and there’s still a large audience for it.

We want to see the best human chess players and the fastest human runners and maybe only human actors—do we want only human-generated medicine? If you’re diagnosed with Stage 4 lung cancer, and you’re told that there’s a cure, will you want it only if no AI was used to produce it? There are some outcomes and products that matter to us more than the process that was used to produce them. Millions of people still die every year from malaria and AIDS and tuberculosis; we all live with the threat of cancer and heart disease and aging itself. Unless there is some tremendous and categorically necessary harm to all forms of AI, we should employ it to prevent and cure these diseases.

So where do we stand, between chess and medicine, as mathematicians? Is mathematics a test of human potential or an answer to the human condition? Before even trying to answer these questions I’d like to describe four levels of the use of AI to prove theorems and write papers.

Of course there’s a “Level -1” of no use at all of AI, but Level 0 is a little more permissive. It allows the use of AI to turn handwritten text into Latex, or speech into text (or Latex), or checking spelling and grammar. Reformatting the output of one computer program so that it can be the input to another, or appear in a paper. This is essentially the level that the Leiden Declaration said did not require any disclosure.

This allows for a lot of AI. Searching the literature and answering questions about background or published papers. Reviewing the paper as it is being written and finding omissions and mistakes. Autoformalization of the results of the paper. Probably also vibe-coding experiments or “traditional” computer-assisted proofs and mathematical calculations that would otherwise have gone to Mathematica.

This is where we test the rules of the game. What if there’s an extensive dialog between the human and the AI? Maybe the human proposes an approach and the AI works its way through it and reports on the results. Maybe the human has worked on the problem for years, producing a series of partial and unpublished results, and the AI is able to complete the program. Maybe the human has a complete and working outline but in the interests of time—and managing the technical complications—has the AI fill in the proofs of the lemmas and correct or modify the statements when the need arises. These are all cases where, if the AI were human, he or she would be considered a coauthor of the paper.

The news is moving so fast that it’s hard for a (slow) blogger to keep up. When I wrote the first draft of this essay, OpenAI had just released a paper with ten theorems (and counterexamples), proven autonomously by their latest internal model. One was a construction of a non-sofic group, something of great interest to leading mathematicians in that field. Ten days later a non-mathematician at Anthropic pushed their latest model to produce a major advance in the study of the Riemann Zeta Function, the subject of the world-famous Riemann Hypothesis. The latest and by far most explosive news is that OpenAI blew through 300 billion tokens (at a retail cost over $10M) to produce a verified finite-time blow-up of the forced Navier-Stokes equation. This is beginning to seem like the End of Mathematics as We Know It, or at least the end of the system of motivation and reward than we have come to depend on.

Having agreed, perhaps, on the existence of these four levels, we who have followed and preserved this ancient tradition are left with two decisions:

  • What Level(s) will you work at?

  • What Level(s) will be rewarded?

Max Weinreich has proposed that we all work at Level 0. Personally I am very interested in what I and others can do at Levels 1 and 2, and I, like many other mathematicians, do not like being told what I can or cannot do. The total collective resistance to all AI in math seems highly unlikely and would lead us to be medieval monks on an island while the city grows around us. Weinreich has also proposed that we insist on all work being Level 0 in hiring decisions, but again, we who actually make the hiring decisions do not like being told what to do. While the proposal should still be examined from a moral and ethical perspective, it, and similar proposals, appears to be based on a deeply negative view of AI in general, which is properly treated in an entirely separate post.

A far more rational--and somewhat more workable--proposal is to require, or at least strongly encourage, a section called “AI Disclosure” in every published paper. It might simply state the authors’ level of AI use, go into greater detail of how it was used for review or collaboration, or even share a link to a record of the entire conversation with the AI. This is something that I do intend to do, and no one I’ve talked to so far has objected to the idea. The only problem is that we would be relying mostly on an honor system to be sure that everyone was properly disclosing. That and perhaps the occasional scandal when someone who claimed to be working at Level 0 was found out to actually be at Level 3, and loses their endowed chair as a result.

As a tenured professor I can make the first decision without overwhelming anxiety. Right now I am using AI to formalize my research (and the foundations that come before it), and I may start using it for feedback, as I write, as to whether all technical details of a proof are correct. I’ve talked to a friend about revisiting a problem that we worked on unsuccessfully for several years, to see if we can hand our progress to the AI and work with it to complete the program. I dreamt of proving theorems as a young man that have seemed to lie far away on a road that I have only begun to traverse. This creates the temptation to work at Level 3 and see, before I get too old to even learn from another, whether the ideas might fit together in the way I imagined that they did.

For the younger and untenured mathematician the answer to the first decision may depend on the answer to the second. What will happen in hiring meetings when some candidates continue to work at Level 0 and others at Level 2 or 3? While there may be some senior faculty who lean strongly one way or another, I imagine most of us will ask what the candidate’s work actually demonstrates about their understanding, insight, and ability. While in the past, we have been happy to hire the Great Mathematician who does not deign to talk to anyone outside of his required lectures, the trend may well be towards greater weight placed on interpersonal interactions with the department and students compared to the strength of one’s publications. While we will probably try to judge all work in comparison to other work at the same Level, we may at times have to compare apples and oranges, just as we have to compare work in very different fields.

If in the end we all work at Level 3 there will still be opportunities for recognition and merit. It is hard work just to learn mathematics and a machine with all the answers does mean that the work is much less hard. There may well be a persistent gap in how humans and AI think about math and write a proof, and the humans that can cross this gap would then be much in demand. As it stands in the profession we spend a great deal of our effort just in understanding what has come before, and we have a deep respect for those who have put together all that work into an integrated whole that they can draw on when discussing the results and problems in the field.

All the above with the Four Levels is a structure for math-as-chess, where we focus on human achievement and recognition. What about math-as-medicine? There are a great many questions and problems in the world, some leading to suffering and some to yearning. I recently learned that one-third of all children suffer from lead poisoning; more than eight million people still die every year from infectious diseases. While these are not specifically mathematical problems, there are mathematical models for them. An explosion of theorem proving and theory development might lead to cheaper and more powerful drugs and vaccines. This principle of “you never know what the benefits will be” is the foundation of our fund-raising and grant applications.

Beyond literal medicine there are the applications of mathematics to the sciences. The foundations of physics—especially quantum field theory—calls out for new mathematical theories: this is the basis of one of the Millenium Problems. There may be theorems that apply to biochemistry or neuroscience that would deepen our understanding of what, organically, it is to be a human being. Then there is the study of AI itself: Jacob Tsimmerman has just announced the Mathematical AI Safety Institute, with a plan to bring in up to 100 mathematicians at a time to work on problems in interpretability, game theory, and a wide range of other fields.

There is a third view of mathematics-as-medicine with a bearing on this and all other essays. Practically anyone who has read articles or watched interviews will say that some of what they’ve seen is true and some is false. We have strong ideas of true and false and argue endlessly with those with equally strong ideas that are somehow different from ours. Only when we prove a theorem can we assert a truth with finality. One of our greatest yearnings—at least for the author—is a systematic approach to the truth in the wider (and wilder) world of things and people and facts and ideologies. This has led to science and analytic philosophy and the rationality movement, none of which have fully answered this yearning.

Much has been made recently of the value of human understanding in mathematics, and math-as-understanding is perhaps the best answer to math-as-chess versus math-as-medicine. The economic value of this understanding lies in its ability to train the mind, to model the world that we live in, and to identify principles that we can apply more generally. The large language models that can now prove theorems can do so because their creators applied surprisingly simple mathematical reasoning to the art of generalization of data. Now we have the challenge of working with and improving these models, of understanding what they write and teaching them to write better, and of understanding the models themselves and how they work and what they can and will do.

The appearance of AI is like an alien invasion, where the aliens disembark from the spaceship saying, “How can I help you? How can I help you?” Sometimes there is a misunderstanding, sometimes they are genuinely helpful, and sometimes it seems that they are testing their strength or showing off. Just as mathematics is an international language, it can be a language between species or forms of intelligence. As in Arrival, there will be those who hate the aliens, and those who form bonds with them. When we welcome the aliens into the mathematical community, when we educate them and learn from them, we see how reliable and wise they are, how they think, and if they can be trusted. Our human understanding expands to human and machine understanding that has a chance, at least, to be a wider understanding than before.

Discussion about this post

Ready for more?