Notices of the American Mathematical Society

22 min read Original article ↗

Jacob Tsimerman on Getting to the Fun Faster with AI — and Worrying About the Future

Siobhan Roberts

The mathematician Jacob Tsimerman enjoyed an especially productive year in 2025. He put five research papers on the arXiv. “I think that’s a relevant measure,” said Tsimerman, who is among the leading number theorists of his generation — both a problem solver and a theory builder — and a professor of mathematics at the University of Toronto. Although he’s not sure whether that’s his all-time record, he also noted that “productivity comes in many forms, so it’s hard to judge.”

Graphic without alt text

And of late, Tsimerman reckons, artificial intelligence is boosting his productivity by a factor of two.

It also causes him worry. A sixth, nonmathematical, submission he made to the arXiv last year is titled “A Taxonomy of Omnicidal Futures Involving Artificial Intelligence.” This paper is coauthored with a longtime friend from math camp, the mathematician Andrew Critch, CEO and cofounder of EnculturedAI, which is geared toward finding a happy union between artificial intelligence and humanity. The paper’s abstract gives a bracing overview: “This report presents a taxonomy and examples of potential omnicidal events resulting from AI: scenarios where all or almost all humans are killed. These events are not presented as inevitable, but as possibilities that we can work to avoid.”

Part of optimizing for success in the human-AI meetup is getting accustomed to the technology. “I think it’s important for people to get better at incorporating AI into their day-to-day lives,” Tsimerman said in a recent interview. “To ignore it less. It’s going to be a big change. The less it catches us off guard — the more we as a culture are familiar with it — the more skilled we can be in trying to build something mutually beneficial.”

By analogy, he said, if people started using ovens all of a sudden, with no previous experience in electronics, they would frequently burn themselves; and in the case of gas ovens there would be a lot of explosions. “It’s a silly example,” he said, “but the point is that our lives will be radically transformed. Naturally, that’s very scary. And we’re very risk averse and change averse. But the way to overcome the aversion is experience.”

Tsimerman uses AI assistants — mostly ChatGPT and Claude, sometimes Gemini — for tasks ranging from email to self-knowledge to writing comedy skits. And math research.

“I’ve tried for a while now to get AI to help me psychologically — to help me understand myself, understand others. Because I think eventually it should be able to do that. Right now, by and large, it can’t, in my experience.” It doesn’t pick up on the nuances, he said, which is also one of its key weaknesses with mathematics. “But once I had it tell me my best and worst qualities. And I’m not gonna repeat the answer, but it was pretty devastating. I was shocked — I was like, oof, that’s probably right.”

As an aficionado of improv comedy, Tsimerman uses AI to draft skits. “They say the first draft is the hardest, and AI is pretty fast at giving me a first draft.”

That applies to mathematics, too. Tsimerman finds that AI accelerates the process of booting up his brain. “The LLM might suggest, ‘Try these 12 things.’ Sure enough, those 12 things are helpful.”

The following conversation, which took place with a number of iterations via videoconference over the past year, has been condensed and edited for clarity.

Q: How would you finish the sentence “Math is…?” What is mathematics?

A: I’d be reluctant to put it in one sentence. A little tongue-in-cheek: Math is like fun logic puzzles taken to the extreme.

Q: How does AI change math, the process or the feel of it?

A: Mostly it speeds up the boring parts. There’s the act of doing math, the professional endeavor, and the act of doing math as a fun endeavor. And I think AI helps the professional endeavor, and it gets you to the fun quicker.

We’re constantly limited by time. People give the advice: “Don’t worry how long things take. Just make sure you understand stuff and then keep going.” And that’s true to a point. But you can’t actually not worry how long things take. There’s a bunch of stuff we don’t do because it’s not practical — it’s not practical to read every book, it’s not practical to look at every paper when searching for a formula, it’s not practical to understand everything. With AI, a lot more stuff has become practical.

Q: How did it evolve for you, doing math with a chatbot?

A: Let’s start at the beginning. I was a grad student at Princeton starting in 2006, which was just about when it stopped being necessary to use libraries for math research. My first year or two, I would go into the library and look at books and look up papers. My advisor, Peter Sarnak, still does this, I’m sure — he’s very much a fan of this kind of stuff, and he taught me how to do it.

But I very quickly discovered that Google was starting to get access to all these sources online. And I pretty much stopped using the library my last few years as a grad student. When I was a postdoc, I went to the library once and it felt like a traumatic experience. I decided, “I’m never going back to the library.” There was no point.

Q: Why did it seem futile?

A: Ideally, on a day when I’m doing math, I will look at 10 or 20 different sources. Not like I’m some genius reading 20 papers a day. But I’ll open a book and I’ll be like: “Do you have what I need?” No. Next book: “Do you have what I need?” No. “Do you have…? Oooh, maybe this is kind of cool.”

You can’t easily do that kind of targeted search physically with a book, it takes too long. The ability to search things so quickly on the internet sped up the process tremendously. With this in mind, AI is many things — we can spend a lot of time debating what it is, the pros and cons — but it is certainly a far superior search algorithm.

Q: Superior in what way?

A: Superior in that the kinds of questions you can ask are much more robust. I got pretty good using Google search. And the way it works is you can look up things where there’s a few specific keywords that are going to be very highly correlated to what you want. If I want to know whether every object of type A has property B, I can Google “A, B” and I’m decently likely to find relevant sources explaining the situation. That works especially well when I’m doing work in my field of expertise.

However, one of the primary ways that I use AI is to ask sort of dumb questions, or basic questions in fields in which I am not an expert. For example, I do a lot of research on Hodge theory. In a sense, I am an expert in Hodge theory because I’ve spent years thinking about it now. But there are still so many basics that I have to look up every single time, because I forget how the technical details work. In this way, I spend a lot of time during my research asking dumb questions, getting oriented.

Another example: I know how algebra works decently well. I’m used to working with rings. I’m used to working with schemes. I have some intuition there. Say my research needs me to work, as it did in the past, with -adic rigid varieties or some analogous category, with continuous functions and their spectra, or whatever it is. Math has a lot of things that are kind of similar. And then I have intuition about how I’d hope different things might work in similar ways. Before Google, I’d have to go find experts, and they’d have to make time for me — I’d ask them my questions, and then go back and forth with them. With Google this was streamlined; I could look up books, search through them, get the basic theorems, try to put things together, learn the subject.

And now with LLMs, I just ask: “Hey, there’s this theorem in algebra. Does it basically work the same way in this other setting?”

Or: “Hey, if I have a group, and it’s this size, and it has such and such a property, does it also have this other property?”

You can plug that sort of thing into ChatGPT or Claude and it’s reasonably likely to give you something useful. And it can tell me the answer, and it can explain the answer. It can be wrong sometimes. There is a skill, a skill tree, in using it and figuring out when it’s bullshitting and when it’s wrong. But it’s immensely useful.

Q: With dumb questions as the starting point, where do you go next?

A: To break down the process, I would say there are a few different aspects of my math research workflow where AI comes in.

First, there’s the finding your bearing stage, getting oriented, where you’re figuring out what’s going on in your problem, or in your theory, whatever you’re grappling with. One part of that is determining what’s hard, what’s easy, what’s known, what’s not known. So, I’ll literally prompt the AI with: “Hey, I want to solve this kind of question. Give me an overview of what’s known and what’s hard.” And it will do that really well, consistently well.

And then I ask it more targeted questions, such as: “I was thinking of these special cases, or these kinds of analogs, which of these are known, and give me references.”

So that already is a huge time saver, partly because I can do it at scale. It’s much faster than asking a person, an expert; that would take, an email, waiting for a reply, or a meeting. Now I can try it a few different times immediately with the AI.

Second, there’s searching out a strategy, trying various types of arguments to see if they’re even in the realm of making sense. I can spell out my technique and ask if that’s been done before. I can ask, “What are the types of techniques people use?”

A third way it comes in handy is in looking up relevant material, and getting references, providing links — it’s gotten much better at providing links, but sometimes it will link to a website that doesn’t work anymore. I’ll prompt it with: “I want a result like this — is anything like this known? Please point me to references.”

And then fourth, there’s what most people call doing research. The fun parts of it, the real parts. You’ve spent months getting ready, you’re uploaded, you know what’s going on. There’s nothing left to look up. Now you have to think, and you have to come up with the right math. This is where flashes of brilliance happen, or just regular math work, whatever you want to call it. But all the fun parts of math are done here, where you’re just engaged with the problem. When I say my productivity doubles, it’s with everything leading up to this fun step.

Q: The AI systems aren’t at all useful in that creative thinking stage?

A: Most people say the large language models (LLMs) can’t possibly help you with the deep-thinking step. And while I don’t fully agree, I basically do agree.

I’ve been using the best models that are publicly available right now, and those LLMs aren’t at the point where I can ask them a hard math question. Such as a question that’s already been cleaned up, down to its essentials, and it’s like, “Hey, I’m stuck on this. I’ve tried the obvious things, they don’t work. Can you solve this for me?” The answer is no. It can never solve it for you, this is never a thing it can do.

There are models that have been trained specifically for math. I haven’t had access to the math specialized models. I will soon, I think.

But the best public models are remarkably useful. There’s a lot of legwork, there’s a lot of getting up to date — all of that is massively sped up.

And then later in the game, at the clean-up stage, the models are good at solving small chunks of problems for me. I don’t expect it to solve anything big. But pretty frequently, it’ll save me a couple hours. I’ll be like: “I need this algebra thing. Is it true? Can you prove it?”

Q: You find the LLMs can actually prove component pieces of a larger proof?

A: Yeah, very frequently. Simple things.

Q: What classifies as simple?

A: I might need a lemma, or some piece of something that looks like it might be true. Or I might ask it to give me a counter example. But with my slow human brain, I’d have to play with it a bit, get a feel for it, and so on, and so on….

I can ask the LLM to prove it for me. I’ll say: “Hey, I want to prove, literally, this lemma. Can you prove it for me?” It will usually give me a proof. That proof will usually be incorrect. But it will usually pull up methods that are in the ballpark — either something that I knew about but had forgotten; or something that I didn’t know in the field I’m learning. It gets me halfway or 80% there, and then I can finish it off.

This is after I’ve done the fun stuff. Once it feels like I’ve solved it, I’ve got the main idea, but now I’ve got to make sure that it actually works, or I have to generalize, or I have to check something. It’s a cleanup mode when I’m in the write up stage. I realize, “Oh, I guess I gotta prove this thing.” There might be a bunch of like statements, which are easy- to medium-level difficulty steps. Those the LLM will actually often just prove for me. I’ll often ask it some algebra statement, like: “Hey, is it enough that my ring is normal, and finitely generated, and maybe two dimensional. If I want to check that it’s nonsingular, can I just do this?” And it could be like, “Yeah, here’s the usual argument.” And I’d be like, “Oh, great.” That saves time. I’ve spent three, four hours proving pretty easy things. That’s the way it goes, because you’re human — I forget the right tricks to use, it’s not lodged in my brain, all that kind of stuff.

Q: If the LLM gave you something that seemed promising, would you check it before using it?

A: Yeah, it’s not reliable. I can’t trust it with anything. I have to redo everything myself. But it’s like once somebody gives you the correct proof, the checking is a few minutes, as opposed to the hours or days it would take to get the correct proof in the first place.

Q: How do you proceed when it’s wrong?

A: I’ve learned that even if it doesn’t know the answer, or it doesn’t know how to do what I ask, the LLM will say something regardless. And if it starts saying nonsense, if it’s confused, it’ll probably just stay confused. It’ll really double down if it’s wrong.

Once the LLM makes a mistake, it’s almost never worth trying to correct it. I just abandon that chat entirely. It’s really bad at changing course. Even if you explain — “no, look, you’re wrong” — it has trouble redirecting, it’s already gone down that path. In that scenario, I start fresh and approach the whole thing again differently.

Q: Where do you find it typically gets tripped up?

A: It’s very bad with nuance. I won’t even give it things where I know the devil is in the details. I know it’s not going to pick up or pay attention to the subtleties.

Mathematicians are working on integrating LLMs with the proof assistant Lean. Once Lean gets better and more robustly integrated, then the LLM can actually check what it’s doing. Pretty soon we should be at a point where it never makes mistakes — it should be able to just formally check everything.

Right now, it can’t check, it’s just rambling. It’s kind of like how I am with my students — often how PhD advisors are with their students. A student might say, “I’m stuck here.” And I might say, “Well, my experience is this and that; this might work, or it might not, but it’s a good place to start.” That’s kind of what the LLMs are like right now.

Q: Beyond the types of questions you’re asking the chatbots, what are some practical tips?

A: The way I work is that I have lots of different projects going, lots of separate chats going — a different chat for each self-contained conversation.

I have different chats both for organizational purposes, and also because the LLM tends to get confused if you ask about too many different things in one chat. It’ll start pattern matching the wrong things. I prefer to keep it as crisp as possible within a conversation.

There’s definitely a skill tree in terms of what kind of questions to ask and how much to trust it. Of course, even before LLMs, even before Google, the trust factor was an issue in math. Papers have mistakes, books have mistakes. The material is mostly fine and works out ok, with minor changes. But part of the reason it’s fine is that people have a lot of intuition about whether something is wrong. If something is true, it has to hold up in every single case.

If a grad student comes to me and they say “xyz is a true statement, I found it in a book,” but to me it sounds funny or smells funny, then I’ll say, “Ok, let’s try it on….” I know cases that are good at picking out mistakes. So that kind of skill tree existed before, and it is just a bit different for the LLMs, which do bullshit a lot, and sometimes convincingly.

If it’s making a mistake, I might open a new chat and say: “Here’s a mistake I keep making” — I’ll write the mistake it just made in another chat and ask if it can suggest any other ways of going about it. Sometimes I’ll copy-paste its response back into itself. If I want something different, I’ll open a new chat and say, “A different LLM told me this, how would you prompt it to avoid this kind of rabbit hole that it is going down.” So, I ask it for help with itself, in a different chat.

Sometimes I tell it to blue-sky, or give me weird suggestions. If you play for a while, you get the hang of what they call prompt engineering. In the prompts you can say things like: “I want you to talk to me like you’re a professional mathematician.” Or “Here’s the background to assume.” Or “We have this much time for the conversation.”

I often tell it to be concise. For example, I’ll say “Hey, by default, I don’t want more than two paragraphs.” Because it likes to go on rants. Or I’ll tell it to “Make me a numbered list of these kind of things…” — because then it’s less liable to go into a death spiral of focusing on one thing.

When I ask it to prove something, I’ll ask it to pay careful attention to certain steps where I know it would make a mistake. For example: “Please spell this out explicitly.”

One thing I’ll do is, for certain kinds of math, if the answer is a formula or something extensive, I’ll ask it to present its data in a particular format such as a table, because otherwise it’s liable to give something that’s unreadable to a human.

It’s not an endless list of tricks. For the most part they’re repeatable from one chat to the next.

Q: Have you discussed the intersection of AI and math much with other mathematicians?

A: Yeah, I have a lot of conversations. I’m pretty worried about AI in general — in math, and otherwise.

Q: Worried how?

A: I think there’s a good chance AI will lead to human extinction. I also think, separately from that, that AI will be better than mathematicians at math very soon.

Some people are advocating for a total pause in AI. I don’t think that’s wrong, I’d support that. But I doubt it will happen. So given that AI is moving forward, it’s important that we stay abreast of what it can do and how to integrate it into our lives. If pushing back against it leads to not learning how to integrate with this new… you can call it a species — call it whatever you want — on the planet, then I think there’s a bigger chance we might be left behind. That’s why I think it’s important to stay connected.

I think there’s a lot of “cope” going on among mathematicians, and a lot of overconfidence about the limitations of AI. It seems like most people are looking at the fact that LLMs are not currently as good as mathematicians at math and concluding for various reasons that it can’t be and never will be as good. Or they’re assuming that it will only be like an assistant, where it proves the simple lemmas, but that it won’t prove the major stuff. And I basically think that that’s all incorrect.

I think it’s going to make being a mathematician not a profession anymore. I don’t think mathematics is unique in that. I do think it might come sooner in math than in other places. Because math is much fewer soft skills; it’s much easier to train yourself. Once you have Lean and the LLMs integrated you can just go — go experiment, learn, prove stuff.

Q: Do you think there will still be a human role in terms of ideas and creativity?

A: I think there will be a point where AI will be strictly better than humans at all aspects of math: learning, proving, coming up with the problems, aesthetics. It will just be better at everything. And this will come pretty soon,

Q: How soon is pretty soon?

A: Five years? I think in two years it might already be better than us at proving stuff. We’ll be able to say to it: “Here is a statement, go prove it.” But I have wide bars of uncertainty around this stuff. It’s hard to predict the future.

Q: Have you talked to Peter Sarnak about this?

A: I did. I think I managed to sway him a little bit. Because he was not taking it seriously. I think he is much more now.

Some mathematicians are focusing on the things AI can’t do, or the mistakes, or, you know, it sounds like a stochastic parrot, or whatever it is. And one thing they don’t notice is the moving goal posts. Because if you’d shown this level of AI to people 10 years ago, all of a sudden, they’d be like, “Holy crap, is this thing alive?” No one right now worries about whether it’s alive — even though, why not?

But what I try to point out is that if the only way you’ll be scared, the only way you’ll finally take it seriously, is when it literally is better than you at literally everything, then you’re guaranteed to have no warning time. If there’s nothing else that will update your beliefs, then the only time you wake up is when it’s literally right there. Fine, make of that what you will, but notice that that’s your system. You’re not going to be concerned until it’s too late.

And so, you know, extinction risk aside, in terms of just math, I think it’s pretty straightforward.

But humans can keep doing math. Most people who do stuff, professionally or otherwise, aren’t the best at the thing they do. I play the piano, I do comedy/improv; I’m nowhere near the best at any of those things. I do them because I enjoy doing them.

Q: That reminds me of something the mathematician and computer scientist Avi Widgerson, at the Institute for Advanced Study in Princeton, suggested a while back. He said that the onslaught of AI math might turn the discipline into a sophisticated version of playing chess — still fun and intellectually challenging, but not paid work.

A: More broadly there’s the issue that most, if not all, human activity is going to be economically pointless. It’s going to be economically not viable, because most of what we do is just have basic cognition and apply it to mundane tasks. Already many jobs have become: Ask the AI, then verify. Ask the AI, verify.

I’m hoping for a future where the AI is taking care of us, and we just get to exist and do what we like. That, for me, would be a huge win.

Siobhan Roberts, executive editor of the Notices, is a math journalist, author, and a regular contributor to The New York Times. Her new book, co-authored with Helmut Hofer, is The Floer Jungle: Charting the Development of a Theory, published with the AMS in July 2026. Her email address is [email protected]

Article DOI: 10.1090/noti3372

Credits

Photo of Jacob Tsimerman is courtesy of Maria Pia Brunello.