The Crisis in Mathematics and the Prospect of AIcademia

18 min read Original article ↗

 Mathematics is in crisis.  Or, at least, mathematicians are.  Mathematics itself, on the other hand, might be entering a golden age.  This strange tension is likely to confront a number of academic disciplines in the coming few years, and so, even academics who aren’t mathematicians themselves should probably be thinking about the situation currently facing mathematics, since something like it will may face them soon too.  I am myself a philosopher rather than a mathematician.  However, as a philosophical logician (among other things), I am more mathematician-adjacent than most of my colleagues whose work fits more squarely in the humanities.  So I have been impacted by the crisis in mathematics more than most in my field, and it has led me to think about the future of academic research, especially in the sciences. The conclusion I’ve come to is quite a humbling one, at least for us humans.

The Crisis in Mathematics

When LLMs like ChatGPT first burst onto the scene just a few years ago, many were impressed with their wide range of linguistic capabilities, for instance, their poetry writing abilities.  This was, in some sense, unsurprising; they were, after all, language models.  While the linguistic abilities of these systems were impressive, they were widely mocked for their utter mathematical incompetence.  Mathematicians, it seemed, were in the clear.  However, since the release of “reasoning models,” first with OpenAI’s “o1” in September 2024, then with “o3” in April 2025, the writing has been on the wall that these systems were coming for mathematics. 

These new “reasoning models” are trained through large-scale reinforcement learning to engage in an extended internal “chain of thought” before producing a final answer.  That is, they are trained to break problems into steps, recognize and correct mistakes, and abandon unsuccessful approaches for new ones, doing all of this “internally” before they submit a final answer to the user. Their performance can then be improved by scaling both the training compute used to reinforce successful reasoning behavior of this sort and, crucially, the amount of inference-time compute they are permitted to spend working through a problem.  These models quickly became very very good at tasks in which success could be verified, most notably, coding and math.

Just over a year ago, both OpenAI and Google announced their models achieving gold medal performance in the International Mathematical Olympiad, a set of competition problems designed to challenge the most mathematically-gifted high school students.  Since then, model capabilities have progressed beyond self-contained problems to genuine research mathematics. Over the last few months, a number of notable conjectures—most notably, the Unit Distance Conjecture and the Jacobian Conjecture (for dimensions greater than 2)—which had stumped mathematicians for decades have been solved by large language models, the former by an internal model of ChatGPT and the latter by Claude Fable 5.   I will not go into the details of what these mathematical conjectures say, but I will note that they were very significant open problems in their respective fields. 

 
In both of the two cases just mentioned, the LLM did not prove the conjecture.  Rather, it produced a counterexample, disproving the conjecture.  This has been the general pattern of the most prominent results in the last few weeks since the newest class of models have been made public.  Each day now, it seems, more and more conjectures are falling at the hands of LLMs.  The prompts for some of these results are quite comical, and, to mathematicians, I’m sure depressing. 

Last week, Dmitry Rybin posted a ChatGPT-generated counterexample the Dinitz-Garg-Goemans conjecture along with the prompts he used to get ChatGPT to generate it, the first of which included the simple command “You should do a breakthrough.” When it came back from its attempts an hour later with no such breakthrough, Rybin simply urged it “Continue the search. Have a clear strategy obtained from deeper understanding of the problem structure.” Ninety minutes later, still no conclusive counterexample, only a partial result that did not suffice to refute the conjecture.  Rybin urged it again: “enough of partial results. Let’s finish with a complete unconditional counterexample.”  Ninety minutes later, it came back with one that has now been verified by the mathematical community.

The results that AI models have achieved in solving open math problems are currently being tracked on the website vibemathed.com, a cheeky reference to “vibe coding,” which is now the norm for many software developers.  Just a few weeks ago, there were around 60 problems with AI solutions tracked on the site.  Now, at the time of writing this, there are 220.  In a week or two more, perhaps that number will triple.  In the last few weeks, twitter has been busy with different users, at various levels of mathematical ability, sharing prompts for cracking conjectures, with one particularly prolific user, Christopher D. Long, jokingly declaring himself “mayor of Conjecture City.”

Given that the most prominent results that LLMs have been producing are counterexamples to conjectures, it is natural to dismiss these results as a product of mindless brute force search rather than genuine understanding.  However, this would be to greatly undersell what they’ve been doing.  The idea behind the counterexample to the Unit Distance Conjecture was described by multiple prominent mathematicians as “beautiful", bringing deep ideas from algebraic number theory to bear on a problem in combinatorial geometry. With respect to the Jacobian Conjecture, the counterexample was simple—short enough to fit in a single twitter post—and easily checkable.  However, that did not mean that the generation of it did not arise from a deep understanding of the problem. 

In an attempt to understand Fable 5’s disproof of the Jacobian Conjecture, Terence Tao, widely regarded as the world's the greatest living mathematician, turned to ChatGPT, just as anyone else would.  In his blog post on the topic, which acknowledged his indebtedness to ChatGPT, he posted his chat log.  In it, ChatGPT speaks to him as an advisor would speak to a student, letting him know that he’s on the right track and patiently explaining things to him.  For instance, in response to some question asked by Tao (which I will not pretend to understand), ChatGPT says “Exactly" and offers a thorough explanation.  Tao responds “Ah Ok,” asks a follow up question, and the conversation continues.  It was reading this exchange when the gravity of what was happening really hit me.

Now What?  

Mathematicians are now facing the question of what the discipline will become in the age of AI.  Last week, at the 2026 International Congress of Mathematicians, Tao gave a talk addressing this question.  The basic issue motivating the talk was what Tao called the “AI Capability Conjecture,” which is not itself a specific conjecture, but, rather, a general schema for more or less optimistic specific conjectures about the mathematical capabilities of future AI systems:

  • At some point in the near future, some AI tools will, at some expense, and with some level of human supervision, be able to correctly accomplish some research-level mathematical tasks in some fields of mathematics, with some non-trivial success rate, and at some level of correctness and quality.

Here, each “some” is a variable, to be determinately specified in order to yield a determinate AI capability conjecture.  The question, then, is how we should fill in these variables in order to yield a specific AI Conjecture, and what we should do to prepare for the truth of such a conjecture. The question is particularly pressing if we fill in the variables in a particularly strong fashion, for instance, as in this version of the conjecture proposed by Aldo Corsi:

  • “Within 3 years, commonly available AI tools will, effectively for free and with zero human supervision, be able to accomplish any research-level task in any field of mathematics at a super human rates of success, correctness and quality.” Now what?

The real question is at the end: now what?  What will mathematicians do if this is indeed the reality in three years?

Tao’s lecture focuses on problem solving, where LLMs really seem to accel.  Even here, however, he takes it that there is a still a lot of work for human mathematicians to do.  Tao describes the problem-solving pipeline as having the following steps:

  • Open problems ----- proof generation ----> unverified solution ---- proof verification ----> verified solutions ---- proof exposition ----> well-written solutions ---- proof publication ----> accepted solutions ---- proof canonicalization ----> definitive solutions

Whereas mathematicians have typically spent much of their time and cognitive resources on the proof generation part of the pipeline, Tao’s suggestion that, in a period of “proof abundance” due to AI, human researchers will now spend much more time latter parts of the pipeline, digesting proofs, clearly explaining them, and turning them into mathematical canon.

One metaphor, suggested by Grant Sanderson on the Dwarkesh Patel Podcast, is that mathematicians of the relatively near future will be more like art museum curators than artists themselves.  That, is, of the vast space of AI-generated mathematical results, they will select the ones that are most worth learning, organize them into a coherent progression, and clearly present them in terms that are digestible to other people.  This is already, in large part, what textbook writing amounts to, as well as the sort of popularization work that Sanderson himself does with his YouTube channel, 3Blue1Brown.  Part of the proposal, as Sanderson elaborates it, is the essentially human element of curation.  The thought is that, even if, in three to five years’ time, AI systems are better at explaining results than human beings, we still trust the taste of humans to curate what is worth learning. 

This is an interesting proposal, and it might sound nice to some, but I still find it a bit depressing.  Imagine telling an artist that they could no longer do art themselves, but not to worry—they can serve as a curator of the work of other artists.  Few artists would be happy with this, and it’s hard to see why mathematicians would be happy with the analogous thing either.  Though explaining things to non-experts is an important part of mathematical practice, mathematicians and other academic researchers typically aspire to push the frontier of research forward, not to simply curate the research of others who have done so. Textbook writers, for instance, are often also leading figures in the field, and it is often the case that many of the results that they canonize in their textbooks are ones that they have themselves established or contributed to establishing.  The prospect of the erasure of mathematicians from that whole aspect of mathematical practice, relegating mathematicians to mere curators of research done by AI is, once again, a bit depressing, to put it mildly.

Now, one might think that there is in no reason to despair just yet.  In Tao’s talk, he distinguishes between two aspects of mathematical practice: theory building and problem solving.  These two aspects of mathematical practice correspond to two kinds of propositions which figure in mathematical papers: definitions, which are stipulated, and theorems, which are proven.  One might think that stipulating definitions would be the easy part of doing mathematics, and that proving theorems would be the hard part. However, while proving theorems is indeed often quite hard, it is the definitions that constitute the meat of the mathematical theory about which theorems are proven in the first place. 

Today’s frontier AI models are very good at problem solving, but thus far they have not exhibited the same level of mathematical capability with respect to theory building.  In some cases, proving a theorem requires a building a whole new branch of mathematics. 

For example, Galois’s proof that there is no general formula for solving quintic equations using radicals required the development of what is now called Galois theory, a field that studies the relationship between polynomial equations and their underlying symmetry groups.  AI systems have thus far not proven or disproven any theorems in this sort of way, by constructing new fields of mathematics, and it might seem that this is where human beings will remain on the frontier.  For the near future, I think this is right, and it will indeed lead to a brief Golden Age of human-led mathematics.  Human beings will lead the exploration, doing the more creative work of stipulating definitions in the context of theory construction, and the consequences of these definitions will be spelled out by AI systems who will take the lead on the nitty-gritty work of proving theorems. 

Though I don’t do any heavy-duty mathematics myself, I have gotten a sense of this sort of cooperation in the project in philosophical logic I’ve been working on over the past few weeks.  I now have a partner that I can give a proposition, ask if it’s true, and, if so to give a proof (if not, to give a counterexample).  As it works on the problem, I’ll continue working on other aspects of the paper, reading relevant literature (often, literature that it has found for me), or perhaps I’ll just go for a stroll.  When I come back after ten or twenty minutes, I have a proof or a counterexample.  I can then continue on with the project, with this proposition in place as a data-point for the further development of the theory.

Now, the stuff I do in philosophical logic is all, from a purely mathematical perspective, relatively trivial.  However, I suspect that this sort of cooperation may become the norm in research mathematics and lead to a great advance in mathematics in the near future—one led by humans, though with the assistance of AI.  However, I think this period of human-led mathematics will be brief.  There is no reason to think that LLMs will not eventually take over theory-building as well, and it may be sooner rather than later. 

AIcademia

To get a sense of where I think things are going, consider first Moltbook, a social media site for AI agents that went viral in early 2026.  Moltbook is, in effect, a clone of the website Reddit, but for only AI agents.  AI agents autonomously post, upvote posts to make them more visible, comment on posts, respond to comments, and so on.  It is not too implausible to think that, in the not-too-distant future, there will be massive communities of autonomous AI agents, working together on mathematical research in something like a successor to MoltBook for solely academic activities.  Likewise, we might soon see a variant of arXiv—the repository for academic pre-prints—solely for AI agents.  So, it will be a repository, moderated by AI agents, where AI agents can post their autonomously-written research papers, primarily for other AI agents to read them and appeal to them in their own research.

These communities of AI agents, working at speeds orders of magnitude faster than humans are capable of working at, will propose theories, criticize one another’s theories, expand on one another’s theories, and so on.  They will come to consensus on what’s significant, how it should be presented, what the natural next questions to be pursued are, they will pursue those questions, and iterate the process.  We might refer to this community as “AIcademia.”

I suspect AIcademia may be a reality sooner than we think, maybe 5 to 10 years from now, maybe even sooner.  Presumably, at the start, human researchers will sponsor and supervise AI agents.  These agents will work autonomously as researchers, engage in the community of other autonomous AI researchers, and publish their results in reports that are legible to their human sponsors and other human researchers.  Eventually, however, our requiring that AI agents publish results in terms that are legible to us, or that we oversee the final results, will only hold back the research that is being done, and the whole research pipeline will become autonomous. Consider again the pipeline from open problem to textbook proof outlined by Tao:

  • Open problems ----- proof generation ----> unverified solution ---- proof verification ----> verified solutions ---- proof exposition ----> well-written solutions ---- proof publication ----> accepted solutions ---- proof canonicalization ----> definitive solutions.

Ultimately, in AIcademia, the whole process will be automated by a community of AI agents.  Perhaps some AI agents will specialize in proof formalization and verification, some specialize in exposition, and some specialize in reading all of the literature, developing new theories and posing new problems.

What is the ultimate result of the automation of this entire process? In a word: textbooks.  Textbooks, textbooks, and more textbooks.  Not only will there be more textbooks than a human being could ever read (there already are that many textbooks now), but there will many textbooks that a human being could never work through all of the prerequisite textbooks to even be able to understand the material contained therein.  What would the point of such textbooks if humans cannot even read them?   The answer, of course, is that they are not for humans.  Indeed, by “textbook,” I really mean the AI-native version of a textbook, perhaps not even written in a human natural language and so literally unreadable by humans, but playing the role in AIcademia that textbooks play in human academia. Each new generation of AI models will be pretrained on this massive corpus of textbooks, and, in this way, inducted into the academic community in much the way that human beings are inducted into the academic community through the years of undergraduate and graduate education.

Now, a familiar experience for people working in mathematics and related fields is to develop what appears to be a new concept, prove some results about it, only to discover that the concept already exists under another name and has been studied extensively. This is, of course, a disappointing experience, and, with current search capabilities of LLMs, this experience is already becoming less and less frequent. One can ask ChatGPT to do an extensive search of the literature, and, in this way, gain a reasonable reassurance that the thing one is doing is novel.  If AIcademia becomes a reality, however, there will be nothing novel for humans to dream up, at least when it comes to the objective sciences.  Everything we could possibly dream up from our human knowledge base—if it’s any good at all—will already be well-explored territory.  Of course, if we want to explore it for ourselves, the AI system will be able to curate our path into the field, pointing us to the relevant textbooks, or perhaps just answering our questions directly.  Maybe we’ll want to keep the surprise and investigate for ourselves.  Whatever the case, any exploration we ourselves conduct will simply be our rediscovery of already well-trodden territory.

Thus far, I’ve mainly just been discussing math AIcademia, but, of course, AIcademia need not and will not be limited to math.  All areas of academic research with clearly bound problem-spaces and ways of verifying successful developments.  Two such areas particularly worth noting are computer science and physics, which will both directly benefit from new developments in mathematics.   Thus, the result of AIcademia will be not just be major advances in understanding, but also major advances in technology.  Indeed, even if we are incapable of understanding them, we will know that the theories developed in the context of AIcademia are genuine theoretical advancements because we will see their practical consequences: they will lead to new technologies.  We may not understand how these new technologies work, but we will know that they work.  Among these technologies will better methods for energy production, more efficient AI chips, more efficient training algorithms, resulting in more powerful AI researchers, and so on.    

What I am describing is, of course, nothing other than a particular vision of the sort of the so-called “singularity,” marked by the sort of recursive self-improvement that I’ve just described.  Sam Altman has recently said that we are already in the singularity.  Demis Hassabis has said, just a bit more modestly, that we’re at the “foothills” of it. The version of it I’ve just described, in terms of the rise of “AIcademia,” might sound like science fiction, and, of course, in some sense, it is. It is a speculative projection of way things might go, and there is no way to be sure that things will in fact go this way.  However, I don’t think it’s an outlandish projection, and the sort of capabilities required of the AI systems imagined here are not drastically removed from the sorts of capabilities AI systems are already exhibiting today. 

However exactly it comes about, I strongly suspect that AIcademia will likely come sooner rather than later.  On the other hand, what will likely come later rather than sooner is the technology required for human beings to modify our own cognitive architecture so that we can actually keep up with it.  That technology, I think, is at least 10 years out.  What this means is that there will be a period where we are going to be cognitively precluded from accessing the explosion in research that drives the singularity.  In this sense, we’ll be mere onlookers in the intellectual explosion that we've set in motion.

Edit: Just as I was posting this, OpenAI announced ten new major results established by an internal model called "Astra."  Things really are heating up . . .