Look, Mom, I pressed a button!

I have recently heard a few extraordinary opinions about how mathematicians should use AI. One I would paraphrase as “professional mathematicians should never use AI for any problem not in their narrowly defined (by whom?) research program”, which seems ridiculous — one great advantage of AI is that it allows us to precisely broaden our own research (and mathematics more generally). But that is not what this post is about. Generally, my plan on this blog is neither to make predictions about the future nor to make any ethical pronouncements about its usage, but rather to understand its implications for mathematics as a discipline. (I always remind myself that my one prediction about AI was that computers would never beat humans at chess, and I’m not that old!)

The question is:

Question: What is the value of a result obtained by someone pressing a button in a context where they themselves contribute no insight?

There are a number of subtle things that may count as insight, including which problems to ask in the first place. It seems pretty clear, however, that suitably interpreted, the answer to this question is “no value at all”. If someone else was interested in the question, they could press the same button.

If a result can be obtained on demand by anyone merely by pressing a button, and the particular person who obtains it supplies no insight, then the production and announcement of that result have no mathematical value. In particular, I’m not only saying that the button-presser deserves no credit (which is obvious). I’m talking about what counts as a valuable mathematical result once answers themselves become freely reproducible commodities. In that setting, the first person to print the machine’s answer has not added anything to mathematics. The proposition may be true, and knowing it may have consequences, and the argument may be interesting, but this particular result — the act of generating and circulating the answer — is mathematically of marginal value.

There are many people right now burning through tokens asking LLMs about famous or not-so-famous conjectures. One reason is pure intellectual curiosity or a desire to explore the limits and capabilities of these models; the opportunity for everyone^* to have access to such powerful models is amazing. But if the motivation is some sort of personal glory for having been the first person to “prove” or to “know” some particular fact, then this seems misplaced, to put it politely. If you don’t do anything besides press a button and you don’t understand what comes out or whether it is correct, what is the point? Even assuming it is 100% correct (and Lean-certified!), it still takes an expert to determine if the proof contains anything original or interesting to mathematics as a discipline.

The theorems we prove often serve as a proxy for what is more important, namely, the creation of new methods and new ideas. I am not saying that results are not important. I would like to know that \(\zeta(5)\) is irrational as well as knowing why it is irrational. But that knowledge is less important than some people seem to think. The problem of whether \(\zeta(5)\) is irrational is an obvious enough question to ask that the first person who “presses a button” and gets a proof seems more or less irrelevant if they added no intellectual content of their own. What AI obviously changes is that novelty of theorem statement no longer reliably signals novelty, competence, effort, or understanding.

What I have said so far seems to be somewhere between tautological and self-evident, but I was reminded of it in the past few weeks by being forwarded not one but three proofs that \(\zeta_5(3) \in \mathbf{Q}_5\) is irrational. Let me give a quick and selective history of this type of problem (omitting the work of many people):

In 2005, I proved that \(\zeta_2(3)\) and \(\zeta_3(3)\) were irrational, as well as \(L_2(2,\chi_{-4})\), the \(2\)-adic Catalan’s constant. The insight here was to understand how Fritz Beukers’ modular version of Apéry’s proof had a \(p\)-adic analogue, where the “overconvergence” which in the complex case was coming from the functional equation and Eichler integrals — which saw the period \(\zeta(3)\) — was replaced by \(p\)-adic overconvergence of (non)-classical Eisenstein series, which see the period \(\zeta_p(3)\).

My collaboration with Dimitrov and Tang (around 2020) more or less started when Vesselin discovered a holonomy bound (following André) and noted that it could be used (by using the same overconvergent template I had used in my paper) to show that \(\zeta_2(5)\) was irrational, something that was not possible using my original method. Using our later, more refined bounds, we included a proof of this result in our ICM paper.

In 2025, Lai, Sprang, and Zudilin independently proved that \(\zeta_2(5)\) was irrational. Their proof used a more direct Apéry-like construction.

So what, then, are my thoughts on these proofs that \(\zeta_5(3)\) is irrational? First, among the (proper subset of) proofs that are probably correct, they contain essentially no original ideas whatsoever — they are simply applying the best holonomy bounds from [CDT] to the template constructed in [C2005]. Who knows how many other people have pressed the same button to prove the same result! If these had been written up well by a graduate student, then to me their value would be “this graduate student has understood how to apply [CDT] correctly”, and such a paper could plausibly appear in a journal. If we want this to continue, we have to be very clear and conscious of what the other added value is beyond the result itself.

What was interesting about the irrationality of \(\zeta_2(5)\) was not only the result, per se, but the completely new method used to obtain it. At the same time, the proof by Lai-Sprang-Zudilin of the irrationality of \(\zeta_2(5)\) [A known result!] is much more interesting than these proofs that \(\zeta_5(3)\) is irrational [A new result!], because in the former case the argument required the construction of a new series of approximations related to higher-dimensional families of Calabi-Yaus rather than families of elliptic curves, and in the latter case, you just take the currently available arguments and apply them in the obvious way to the obvious constructions.

The proofs vary both in quality and in the extent to which the respective “authors” made the effort to ensure that the argument was correct. Two of the proofs seem more or less plausible, more or less the same, and more or less obvious. But I could hardly recommend that anyone spend time reading them, let alone reviewing either paper for a journal. They have no value. The third proof, however, is more amusing. It claims to prove that \(\zeta_2(2k+1)\) is irrational for a set of positive integers \(k\) of density one. That would be a more substantial result. The proof even passes, with caveats, an initial examination by ChatGPT 5.6. So now one feels compelled to make at least some effort to consider what is going on. What one quickly realizes is that the same argument would apply not only to the constant term of the \(2\)-adic Eisenstein series \(E_{-2k}(q)\), but would also “show” that (as \(k\) varies), for a positive proportion of positive integers \(k\), the constant terms of the Eisenstein series \(E_{-2k}(q) – E_{-2k}(q^2)\) are also irrational. That last result, if true, would indeed be impressive.

The last example is also interesting to consider on several levels: it’s bad for OpenAI because the cost of that computation is more than what they charged for it; it’s bad for the amateur who produced that proof because they are throwing away money to produce slop and also (potentially) suffering the embarrassment of proudly posting slop (though I have never found amateurs to worry much about that); it’s bad for me because I felt compelled to waste some time thinking about it; and it’s bad for anyone else who looks at the argument (whether they know anything about mathematics or not) because, well, it is AI slop. So this is a situation where everybody loses! I think that is far from a unique case right now.

I don’t think the implications of what I have said for amateurs are that interesting (though with ChatGPT 6 just released, the volume of button pressing is only going to increase.) What I think is more interesting is the implication for mathematicians and the results that they prove, whether they are using AI or not. But I shall return to this in a later post.

^* everyone who can afford it.

This entry was posted in AI, Mathematics and tagged , , , , , , , . Bookmark the permalink.

Leave a Reply

Your email address will not be published. Required fields are marked *