Ways to improve Mathematics (an introduction)

The goal of this continuing sequence of posts is to recognize the reality in how our subject is changing and how we should adapt.

Even in a hypothetical world where AI was infallible, essentially omniscient, benevolent, a wonderful expositor, and freely available, I believe that many people would still want humans to maintain and develop a deep understanding of mathematics. They would not want to leave mathematics entirely to the machines, any more than they would want to abandon other large swathes of human thought. For this post, I will take that desire as a starting assumption.

My core belief is that understanding mathematics is hard and AI is not going to fundamentally change that, even if it makes proving theorems in mathematics much easier. I remember driving to Wisconsin and listening to Jordan Ellenberg on Lex Fridman’s podcast. What struck me most from the podcast was the discussion of Fermat’s Last Theorem. Fridman was essentially arguing that since the statement of Fermat’s Last Theorem was so simple, there must inherently be a simple explanation of why it was true. This reflects a philosophical idea about science and mathematics that I think is fundamentally untrue: that any truth that is simple to state will ultimately be true for a simple reason. The easiest proof of Fermat may well not be the one found by Wiles (or maybe it will be). But consider instead something much older and established in mathematics, namely class field theory. If I hold any position in this post with conviction, it would be that, even with superhuman exposition, a human could not acquire a good understanding of the statements and proofs of class field theory without years of dedicated study^*. This is not an isolated example. Maintaining and developing human understanding across mathematics requires people who can devote substantial parts of their lives to it. If we value that understanding, there is a case for supporting those people and the communities in which they work, rather than leaving the whole enterprise to whoever wants to work on it in their spare time. That is a central part of the case for mathematics as a profession.

What is important to recognize, however, is that mathematics as a profession will surely be changing rapidly. Daniel Litt has told us that we need to be honest not only about the aspects of our field that will break with AI, but also about the aspects that are already broken. This is an amazing opportunity to fix some of these things! If we are going to defend mathematics, at least in part, as a means of increasing human understanding, then we ought to be rather more demanding about whether our own practices actually contribute to this understanding.

To me, there are (at least) four natural systemic issues in mathematics that we have to address. Addressing any one of them will require consensus building to achieve the massive shift necessary in our community norms. I plan to devote a blog post to each of the issues, so for now I will just introduce them. I am not attempting to be prescriptive, but rather simply to think aloud and hope for suggestions.

One thing we need to do in particular is ask whether our institutions and practices reward the production of mathematical understanding, or merely activities that we have come to treat as evidence of it. With AI, activities such as producing proofs will no longer be quite so closely coupled with understanding.

  • Journal Articles: The journal system was, to put it politely, already struggling before AI. There are too many papers, and the refereeing system is bursting at the seams. At least in mathematics we have a large number of high-quality journals, which helps diffuse the power of editorial boards, given the career importance of publications. [Imagine a field where success could come only from publishing in a single journal; such fields exist.] As we go forward, we very much have an opportunity to reconsider how the publication system works in a radical way. If we don’t abandon it altogether, then we need to redefine what we hope to get out of it. I already have anecdotal evidence that submission rates to top journals are rising sharply, and I doubt that the current refereeing system can sustain such an increase. I don’t think it is as simple as demanding better exposition — it’s not unreasonable to expect AI to vastly improve in this dimension as well. If mathematics is a conversation, then what we would like to achieve are strands of interesting conversations that people are both invested in and listening to. This touches, in part, on the question of insularity raised below.
  • Seminar Talks: I would say that the median mathematics seminar could (perhaps harshly) be described as a waste of time for both the participants and the speaker. If we are to claim that fostering mathematical understanding is one of our main goals, we certainly haven’t made much of an effort to reward good talks. One institutional obstruction has always been the expectation that people talk about their own work. What often gets lost when one does this is an explanation of why the broader question being addressed is interesting in the first place, the methods that have been used most successfully in the field in the past, and the most promising questions to consider in the future. (One can do this in a talk about one’s own work, but people frequently do not.)
  • Insularity: The past few decades have seen an explosion in mathematics. But I feel this has come at the cost of mathematicians being less able to communicate; not only with people in other areas of mathematics, but sometimes even within their own field. Some have argued that this is an inevitable consequence of the growing difficulty of mathematics, but I suspect that once a mathematical subcommunity reaches a certain size, the impetus to reach out diminishes, to all our detriment.
  • Ego: Perhaps the thorniest question of all: how do we shape the incentives in our field to produce better outcomes? For all that we emphasize understanding, it would be insane not to acknowledge the importance of ego, and the way that the desire to be the first person to prove something has motivated many of us. This moment is going to require a great deal of humility. If human understanding of deep mathematics is what we want to defend, it ought also to be what we reward.

^*: now I have the following in my head:

Posted in Mathematics | Tagged , , , , , , , | 3 Comments

Look, Mom, I pressed a button! (Go edition)

Apropos of nothing, I was reminded recently of a fascinating story I heard from Geordie Williamson about AI and go — perhaps this story was the original source. It is a story about how a technology which has the power to greatly increase knowledge can rather create barriers to actually acquiring that knowledge. To give just some excerpts:

I started my career as a Go teacher in 2020 … I now estimate that about half our students had used AI in at least one game and one in ten were chronic users. We were originally baffled … It didn’t make sense that players would just throw away their practice games to have AI win on their behalf.

None of these reasons [for cheating] were surprising to us … What personally shocked me, however, was the way our students conceptualised their AI use. In this, Carlo Metta was also a surprisingly predictive case. The original reddit thread discussing his ban featured a comment from a user called “carlo_metta”, which read:

I never let Leela choose move. I just decide myself which one is better, for this reason i think i can find my own style with Leela. Go is an art and Leela help me tyo [sic] express my skill

That account was a burner, quite possibly a troll. However, I couldn’t help but recall the comment when I heard identical arguments coming from our cheating students’ accounts.

I think this story has a number of themes relevant to math, including how AI use by experts can make us worse mathematicians. I think when mathematicians use AI they need to be extremely careful not to fall into the trap of imagining that they (rather than some autonomous agent) are doing mathematics. Amateurs at least don’t have this problem!

I know good mathematicians who have used AI to prove interesting theorems, and they swear that the result started with their own original ideas which were then combined with the insights of AI. And I believe them! But it is very easy to start taking ownership of ideas that are not your own. I think mathematicians need to be very conscious that button pressing has the possibility of warping one’s perception of how much you actually contributed yourself.

For me, I think it helps to have at least one project where you simply do not use AI at all. And for projects that do use AI, make every effort to be honest to yourself about what your contribution is.

In my professional work that has appeared online so far, I have only used AI to “review” one paper after it was written. I have found this very helpful, and I certainly continue to do this. In my experience, AI reviews papers incredibly thoroughly. Ironically, what it seems most likely to miss are arguments that are so poorly written that it’s not even clear exactly what the argument is, but when you actually include some details it has an opportunity to find the hole in the argument.

That said, I definitely do plan to use AI for future research. I have used AI to try to better optimize the choice of the functions \(\psi\) we used in [CDT] (for example, look at Figure A.4.5). I was utterly confused during that paper how to optimize the (very non-linear) holonomy bounds as \(\psi\) ranged over all holomorphic maps \(\varphi: D(0,1) \rightarrow D(0,1-\varepsilon)\). Not only is this a complicated non-linear optimization problem, but there is also the issue of being able to actually compute the answer quickly and rigorously. After using ChatGPT 5.6, I remain utterly confused as to what the optimal choice looks like, but at least it *could* do better, and found a certain Ansatz of functions to try which improved the numerology. The improvements were not quite good enough (yet) to simplify our proof of the irrationality of \(L(2,\chi_{-3})\) in any meaningful way, but they were good enough to prove that the \(\mathbf{Q}(x)\)-vector space generated by functions \(f(x) \in \mathbf{Q}[[x]]\) on \(\mathbf{P}^1 \setminus \{0,1,\infty\}\) which have denominator type \([1,2,\ldots,n]^2\) has dimension at most \(8\) rather than dimension at most \(9\). (We still suspect the actual answer is \(5\).)

My “biggest” (in terms of tokens) use of AI so far is a somewhat quixotic attempt to construct a new finite simple group. Again, more on this later, but as the computation continues, it is my obligation to remain crystal clear about what exactly my contribution is. Note that although the headline goal of this project will certainly end in failure, there is the hope that interesting mathematics will nonetheless come out of this.

Posted in Mathematics | Tagged , , | 4 Comments

Look, Mom, I pressed a button!

I have recently heard a few extraordinary opinions about how mathematicians should use AI. One I would paraphrase as “professional mathematicians should never use AI for any problem not in their narrowly defined (by whom?) research program”, which seems ridiculous — one great advantage of AI is that it allows us to precisely broaden our own research (and mathematics more generally). But that is not what this post is about. Generally, my plan on this blog is neither to make predictions about the future nor to make any ethical pronouncements about its usage, but rather to understand its implications for mathematics as a discipline. (I always remind myself that my one prediction about AI was that computers would never beat humans at chess, and I’m not that old!)

The question is:

Question: What is the value of a result obtained by someone pressing a button in a context where they themselves contribute no insight?

There are a number of subtle things that may count as insight, including which problems to ask in the first place. It seems pretty clear, however, that suitably interpreted, the answer to this question is “no value at all”. If someone else was interested in the question, they could press the same button.

If a result can be obtained on demand by anyone merely by pressing a button, and the particular person who obtains it supplies no insight, then the production and announcement of that result have no mathematical value. In particular, I’m not only saying that the button-presser deserves no credit (which is obvious). I’m talking about what counts as a valuable mathematical result once answers themselves become freely reproducible commodities. In that setting, the first person to print the machine’s answer has not added anything to mathematics. The proposition may be true, and knowing it may have consequences, and the argument may be interesting, but this particular result — the act of generating and circulating the answer — is mathematically of marginal value.

There are many people right now burning through tokens asking LLMs about famous or not-so-famous conjectures. One reason is pure intellectual curiosity or a desire to explore the limits and capabilities of these models; the opportunity for everyone^* to have access to such powerful models is amazing. But if the motivation is some sort of personal glory for having been the first person to “prove” or to “know” some particular fact, then this seems misplaced, to put it politely. If you don’t do anything besides press a button and you don’t understand what comes out or whether it is correct, what is the point? Even assuming it is 100% correct (and Lean-certified!), it still takes an expert to determine if the proof contains anything original or interesting to mathematics as a discipline.

The theorems we prove often serve as a proxy for what is more important, namely, the creation of new methods and new ideas. I am not saying that results are not important. I would like to know that \(\zeta(5)\) is irrational as well as knowing why it is irrational. But that knowledge is less important than some people seem to think. The problem of whether \(\zeta(5)\) is irrational is an obvious enough question to ask that the first person who “presses a button” and gets a proof seems more or less irrelevant if they added no intellectual content of their own. What AI obviously changes is that novelty of theorem statement no longer reliably signals novelty, competence, effort, or understanding.

What I have said so far seems to be somewhere between tautological and self-evident, but I was reminded of it in the past few weeks by being forwarded not one but three proofs that \(\zeta_5(3) \in \mathbf{Q}_5\) is irrational. Let me give a quick and selective history of this type of problem (omitting the work of many people):

In 2005, I proved that \(\zeta_2(3)\) and \(\zeta_3(3)\) were irrational, as well as \(L_2(2,\chi_{-4})\), the \(2\)-adic Catalan’s constant. The insight here was to understand how Fritz Beukers’ modular version of Apéry’s proof had a \(p\)-adic analogue, where the “overconvergence” which in the complex case was coming from the functional equation and Eichler integrals — which saw the period \(\zeta(3)\) — was replaced by \(p\)-adic overconvergence of (non)-classical Eisenstein series, which see the period \(\zeta_p(3)\).

My collaboration with Dimitrov and Tang (around 2020) more or less started when Vesselin discovered a holonomy bound (following André) and noted that it could be used (by using the same overconvergent template I had used in my paper) to show that \(\zeta_2(5)\) was irrational, something that was not possible using my original method. Using our later, more refined bounds, we included a proof of this result in our ICM paper.

In 2025, Lai, Sprang, and Zudilin independently proved that \(\zeta_2(5)\) was irrational. Their proof used a more direct Apéry-like construction.

So what, then, are my thoughts on these proofs that \(\zeta_5(3)\) is irrational? First, among the (proper subset of) proofs that are probably correct, they contain essentially no original ideas whatsoever — they are simply applying the best holonomy bounds from [CDT] to the template constructed in [C2005]. Who knows how many other people have pressed the same button to prove the same result! If these had been written up well by a graduate student, then to me their value would be “this graduate student has understood how to apply [CDT] correctly”, and such a paper could plausibly appear in a journal. If we want this to continue, we have to be very clear and conscious of what the other added value is beyond the result itself.

What was interesting about the irrationality of \(\zeta_2(5)\) was not only the result, per se, but the completely new method used to obtain it. At the same time, the proof by Lai-Sprang-Zudilin of the irrationality of \(\zeta_2(5)\) [A known result!] is much more interesting than these proofs that \(\zeta_5(3)\) is irrational [A new result!], because in the former case the argument required the construction of a new series of approximations related to higher-dimensional families of Calabi-Yaus rather than families of elliptic curves, and in the latter case, you just take the currently available arguments and apply them in the obvious way to the obvious constructions.

The proofs vary both in quality and in the extent to which the respective “authors” made the effort to ensure that the argument was correct. Two of the proofs seem more or less plausible, more or less the same, and more or less obvious. But I could hardly recommend that anyone spend time reading them, let alone reviewing either paper for a journal. They have no value. The third proof, however, is more amusing. It claims to prove that \(\zeta_2(2k+1)\) is irrational for a set of positive integers \(k\) of density one. That would be a more substantial result. The proof even passes, with caveats, an initial examination by ChatGPT 5.6. So now one feels compelled to make at least some effort to consider what is going on. What one quickly realizes is that the same argument would apply not only to the constant term of the \(2\)-adic Eisenstein series \(E_{-2k}(q)\), but would also “show” that (as \(k\) varies), for a positive proportion of positive integers \(k\), the constant terms of the Eisenstein series \(E_{-2k}(q) – E_{-2k}(q^2)\) are also irrational. That last result, if true, would indeed be impressive.

The last example is also interesting to consider on several levels: it’s bad for OpenAI because the cost of that computation is more than what they charged for it (edit according to the comments this is probably wrong!); it’s bad for the amateur who produced that proof because they are throwing away money to produce slop and also (potentially) suffering the embarrassment of proudly posting slop (though I have never found amateurs to worry much about that); it’s bad for me because I felt compelled to waste some time thinking about it; and it’s bad for anyone else who looks at the argument (whether they know anything about mathematics or not) because, well, it is AI slop. So this is a situation where everybody loses! I think that is far from a unique case right now.

I don’t think the implications of what I have said for amateurs are that interesting (though with ChatGPT 6 just released, the volume of button pressing is only going to increase.) What I think is more interesting is the implication for mathematicians and the results that they prove, whether they are using AI or not. But I shall return to this in a later post.

^* everyone who can afford it.

Posted in AI, Mathematics | Tagged , , , , , , , | 10 Comments

OpenAI, updated

An update on this post, from my inside sources:

word on the math streets of SF is that the OpenAI team tried something like 500 problems to get their 10 solutions.

I don’t know how much of an insider this source is (or this sources sources, etc), but (allowing for the possibility of confirmation bias) this is within the expected range.

Posted in Mathematics | Tagged , , | 6 Comments

Google Alert!

I have my google alert set for the phrase “Galois Representations”. Every six months or so it pops up with a suggestion, and I can’t quite work out what algorithm is using. Here was today’s breaking news: On the conductors of mod \(\ell\) Galois representations coming from modular forms.

Posted in Uncategorized | Tagged , , | Leave a comment

The inverse Galois challenge, part II

This is a sequence to this post. The SAIR competition (Round I) has been completed! 98.4% of the possible signatures were obtained, with only 39 non-solvable cases missing.

Some thoughts.

First, my timing in the last post of dissing the problem of realizing \(M_{23}\) as a Galois group was not so great. It seems to me that the delightful paper does an excellent job of combining human and AI thoughts but also clearly and concisely explaining the ideas, especially distinguishing between what is known, what is clever, and what is lucky. Nicely done!

Moving on to the competition. I thought that it would be better to get a precise sense of the difficulty by trying it myself. The approach I used was purely to tell CODEX to do 6 obvious things, but not to either look at any literature myself, not to write any code, and just to come back and complain when it failed. This quickly produced around 40,000 pairs, but then stalled. One approach that wasn’t successful at all was as follows. There were around 80,000 pairs or so could be realized as coming from the Galois closure of degree 12 extensions of quadratic fields. But alas, my suggestions for how to construct these were not taken up sensibly, and I didn’t pursue it.

Certainly my personal explorations produced no meaning mathematical content at all. The only mathematical idea I had that was not completely obvious was one I learnt entirely from David Roberts. In situations where one has a Galois extension \(L/K/\mathbf{Q}\) where \(K\) has Galois group \(G\) and \(L\) has Galois group a central extension of \(G\) of degree \(2\), then one can write \(L\) as the splitting field of a polynomial of the form \(f(x^2)\) where \(f(x)\) has splitting field \(K\) and one root of \(f(x)\) generates a field \(E\). But now, given \(g(x)\) with \(E \simeq \mathbf{Q}[x]/g(x)\), how does one find \(f(x)\)? The observation is that one can often find \(f(x)\) by applying \(\texttt{polred}\) to \(g(x)\).

That said, having done some of these experiments, it did help me appreciate what type of problem this was. It certainly seemed to be the case that real skill and knowledge working with explicit polynomials and explicit Galois theory would be genuinely useful, and simply a purely theoretical knowledge of (say) the general solvable case is not sufficient. It is no surprise then that Klüners and Malle (the leading team) were so successful.

But where does it lead us? I don’t think the conclusion is so far from my original prediction. I think there might be a new second round coming, and after that is done, it really could be the case that the only pairs remaining are \((G,r)\) where \(G = \mathrm{PSL}_2(\mathbf{F}_{23})\) and also \(G = \mathrm{PGL}_2(\mathbf{F}_{23})\) with \(r=0\) (which are hard for the same reasons), and then possibly some cases of \(M_{24}\) (say with \(r=0\)) as well. We shall see!

Posted in Mathematics | Tagged , , , , , , , , , , , , | 3 Comments

Putting ChatGPT through its paces

We are all aware of what ChatGPT can do. I think we would be better informed if we also learn what it cannot do (at least right now!). In order to better understand the current capability of the latest publicly available AI models, I decided to curate a set of problems and test how well ChatGPT (running 5.6 sol ultra) could do. While this certainly is not a scientific experiment, I had some rules I set for myself in advance:

  1. The problems should be, as far as I know, generally be open problems.
  2. I should have at least some original thought or idea on how to approach the problem which I can suggest to the model, however stupid it might be.
  3. I will limit myself to one ChatGPT pro subscription and the time between the ICM and the Emerton-Kisin conference (a bit under two weeks) to address all of these questions.

The strategy I employed was as follows. I gave ChatGPT one to two hours on each problem (ultra think in the chat window). If it made no progress at all, then I didn’t pursue the problem any further. If it did make progress, then I used codex on goal mode to push towards the problem, or at least towards some interesting intermediate goal, with a maximum run of two days.

The goal of this project is not to test the limits of what can be done by these models, but a much more practical test of how it might be to use these models as a working mathematician. If you take this experiment, scale up the amount of compute, and the amount of mathematicians giving (limited) direction to the machines, the result is (to my mind) broadly consistent with what OpenAI achieved, of course assuming that there were a significant number of problems on which they made no progress and then abandoned.

This quarter at Chicago I will be running a “ChatGPT seminar” precisely to explore these questions. The scope of that seminar will be somewhat broader than pure problem solving, and also include typesetting, coming up with interesting conjectures, and many other things. That said, it will certainly involve tests such as these.

So how well did it do? Let us see.

  1. Compute the slopes of all finite slope overconvergent modular forms with \(p=2\), level \(N=1\), and integral weight \(k\).
  2. This is a special case of the Ghost Conjecture of Bergdall and Pollack, which has now been solved by Liu-Truong-Xiao-Zhao (see also this post. Note, however, that that proof excludes this particular case when \(p=2\). I specifically pointed the model towards Conjecture 2 of this paper. In this very special case, the problem reduces to computing the Newton Polygon of a very explicit matrix where one can take \(k\) to be a non-negative integer. Our paper answer the case when \(k=0\). We also worked out the case \(k=-12\) and \(k=-72\) (the latter in part for proof of concept of the approach we were using).

    Level of interest: Kevin and I certainly spent some time trying to prove it! This special case probably now mostly of historical interest in light of more recent approaches.

    Level of difficulty: I would not be surprised if it could be solved by some elementary arguments.

    Result: No progress in the initial time period; not pursued. I was a little surprised, but with the time constraints this did not seem worth devoting extra time to this question. Time spent: about 90 minutes. I’m going to get on a plane in a few hours, I’m going to give it another go for 6 hours or so this evening, then update tomorrow if anything changes. (Update: my flight is delayed and I’m waiting at the airport, but I’m stopping it after looking what it has done so far.)

  3. Prove that there exists a constant \(N\) such that, if \(G\) is a finite group with \(H^i(G,\mathbf{Z})=0\) for \(i=1,2,\ldots,N\), then \(G\) is trivial.
  4. This perhaps the one problem I felt I had the least insight. I think I learnt it from a mathoverflow question in the long past (yes, I looked it up and found it here).

    Level of interest: Hard for me to say. One imagines this problem should be more or less a computation plus a literature search for the case of finite simple groups (assuming CFSG), and then it becomes some inductive problem which may or may not be about facts concerning the cohomology of almost simple groups. But this is not my area.

    Level of difficulty: I have no idea.

    Result: Partial progress. It knows enough to answer the case of finite simple groups, which is the first step in the obvious induction argument. It does cover quite a few non-trivial cases, but then gets bogged down, and comes up against what it calls difficult problems. Time spent: about 48 hours.

  5. Determine the slopes of all periodic billiard paths in the regular heptagon.

    An equivalent formulation is to take the \((2,7,\infty)\) triangle group and ask for a classification of its cusps in \(\mathbf{P}^1(K)\) where \(K = \mathbf{Q}(\zeta_7)^{+}\). This is a thin group inside \(\mathrm{SL}_2(\mathcal{O}_K)\). I heard about this problem from Curt McMullen, who also gave a possible answer (who he attributed to someone else, but since this was just a conversation I apologize that I did not remember at the time).

    Level of interest: I think if you answered this question then Curt would be impressed. What more could you ask for?

    Level of difficulty: One reason I considered this problem is that I had a sense that the answer should involve some mix of algebraic number theory and or Arakelov theory, and at the same time some complex analysis in the form of Hodge Theory. This could exactly be the type of situation where there might be a simple answer just by combining ideas from different fields.

    Result: ChatGPT had sloppy thoughts on this one! My first reading is that it did not have any crucial insights beyond fleshing out a little what I had suggested. It certainly diligently tried to push things as far as it could, but I think it is still missing a (or the) key idea. Time spent: around 48 hours.

  6. Let \(M\) and \(N\) be two finite volume hyperbolic \(3\)-manifolds with isomorphic pro-finite completions. Prove that \(M \simeq N\).
  7. I first learnt about this problem from Martin Bridson and Alan Reid in Ventotene in 2015. At the time, I had some idea about approaching this via the representation variety, but it was sufficiently far from things I knew that I didn’t pursue it.

    Level of interest: Definitely there are people interested in this problem.

    Level of difficulty: Too difficult for me to say, but it’s a well–known problem,
    and (as I learnt during this process) significant progress has been made over the past few years.

    Result: Claimed Solution. This perhaps might be the most interesting positive case. ChatGPT informed me of a recent paper of Liu in which he proved that, for closed hyperbolic \(3\)-manifolds, the profinite completion determined the volume modulo a conjecture about the injectivity of a certain regulator map. I suggested that one could bypass this using the mod-\(p\) Chern class maps discussed in Calegari-Garoufalidis-Zagier. With that, ChatGPT was very quickly able to write a \(5\)-page paper using Liu’s result giving an unconditional proof (in this class of manifolds) that volume was determined by the pro-finite completion. I think that this could have lead to a nice short note that I could reasonably post under my name with suitable AI assistance disclaimers. But then I learnt from Alan Reid and Martin Bridson (who I sent a draft to) that the full result had recently been proved by Xu! At this point, I “pressed another button” and asked ChatGPT to prove the full result, which it did. In particular, the notes below were produced completely independently from the work of Xu, now available here, but they were produced with knowledge that such a paper existed. This surely (?) would have influenced the strategy that ChatGPT decided to pursue. In fact, while there are similarities in the argument they are certainly not the same; After Xu’s paper was posted on the arXiV, I asked ChatGPT to compare the proofs, and it came up with the following: When it was first done, I asked ChatGPT to referee and revise its own work back and forth in until it claimed it was ready to be submitted. It modestly suggested that it should be submitted to the Annals of Mathematics. Now it is not the main point of this post, but obviously the question of how we evaluate work going forward is going to be an extremely important one. While the paper posted above does contain at least one idea of mine, I certainly do not intend to publish it, nor am I willing to take responsibility for its contents. Time spent: about 12 hours. Added: I was asked for my estimate of the chances that this proof is correct, and my response was “over 75% … Perhaps higher”.

  8. Let \(\Delta = \langle x,y | x^p, y^q, (xy)^r \rangle \) be a hyperbolic triangle group, and let \(B/K\) be the associated quaternion algebra over the invariant trace field. Let \(g(p,q,r)\) be the density of real places such that \(B\) is non-split. Prove that either \(g(p,q,r)=0\) or \(g(p,q,r) \ge 1/12\).
  9. Level of interest: This is a question of Curt McMullen raised in this paper.

    Level of difficulty: Note that in this paper here we prove that \(g(p,q,r)=0\) for precisely \(14\) explicit hyperbolic triangle groups, also answering a conjecture of Curt from that same paper. It was definitely clear to me during the writing of this paper that it could certainly be possible to prove this result. I actually started a project with University of Chicago undergrads towards it, but none of the people who signed up seemed actually willing to do any work so it petered out. The one difficulty that was certainly possible was that some eventual argument might be effective, but not effectively effective. Two improvements were needed from the previous paper; the first was to optimize the Fourier analysis aspect. The second one was to replace the Jacobsthal function argument which produced a single interesting conjugate to something more flexible that could produce a positive density of interesting conjugates.

    Result: Solved. Here ChatGPT did a number of things I expected, which was to choose a much more elaborate test function in the Fourier argument than we used, since it would obviously be much better handling much more complicated expressions. This was the first problem I asked, and for some time I actually was going to get ChatGPT to formalize the proof in Lean, which it felt completely capable of doing. But the time frame was going to be several weeks, and I didn’t want to waste the tokens. But this might possibly be worth doing. Time spent: about 6 hours.

  10. All the problems listed in my current NSF proposal draft.
  11. They are all ChatGPT hard, right now!

  12. Construct a new finite sporadic simple group not in the current classification.
  13. Level of interest: A lot. This sounds like a trolling question, but I do actually have one not entirely stupid idea, which should hopefully at least produce some interesting mathematics.

    Level of difficulty: Probably quite hard.

    Result: We are now 10 days into various computations that are making progress on something. But it didn’t come under budget, so I will talk about it later instead.

Posted in Mathematics | Tagged , , , , , , , , , , , , , | 5 Comments

Observations from the ICM, Part 1

My recent post generated a surprising amount of personal emails defending of Philadelphia. I can happily report that, although I didn’t really get a chance to explore the city in any depth, it has at least one excellent cafe. “Thank You, Thank You” is one of the best cafes I have been to in the US; I wish there were something half as decent near Hyde Park. (Hat tip to Toby Gee for finding this in his research.)

Returning to the conference, Terry Tao made the point in his public talk — as others have made elsewhere — that, more than ever, we should insist on rewarding aspects of mathematics beyond simply proof, in particular good exposition, particularly of the deepest and most difficult ideas.

It was interesting, in this light, to see Dennis Gaitsgory’s plenary talk^*. The fundamental problem with the ICM is that there is a contradiction behind the entire concept of an invited talk, particularly a plenary talk. It is simultaneously supposed to be an honor for great work and an opportunity to communicate those ideas. On the one hand, there is no question that Dennis deserves the first honor. On the other hand, the talk was, shall we say, somewhat challenging for a mainstream mathematical audience. I don’t blame Dennis; he has his style of giving talks^**, and this one went more or less exactly as anticipated. But it seems to me that a very simple solution would have been to have asked David Ben-Zvi to talk about the work of Dennis (and his collaborators). I hope for the next ICM the struture committee will consider more radical changes than what they have done so far. Would inviting someone to talk on the work of X be any less of an honour for X than asking X to talk? If they did that with *every* plenary talk, and at the same time made an effort to choose the right speakers (who could even collaborate with X), I think that would be a great improvement.

As the community moves ever so slightly towards demanding better exposition, it is interesting to look back on the Mochizuki circus. If we are going to be honest now, we should also be honest about the past. “Inter-universal Teichmüller theory” is nonsense, and that was more or less obvious to everyone at the time I wrote that post (which was five years after the announcement). The saving grace of those papers is that they were written in “Mochislop,” an almost comically embarrassing STYLE that screams “I am a crackpot.” Were it not for Mochizuki’s reputation, they would have been immediately dismissed. I am glad that, at the time, there were people (Scholze and Stix) willing to donate their time in a heroic effort to actually engage with the mathematics. In today’s world of AI slop, I think people would feel more comfortable not even bothering, and this is a good thing.

^*: The reason I single Dennis out here is not because his talk was in any way particularly exceptional compared to some other plenary talks, but simply because I know him well enough to be able to criticize his talk and feel confident that he will be OK.

^**: A student onced asked Dennis whether, given two correspondences \(X \leftarrow Z \rightarrow Y\) and \(Y \leftarrow W \rightarrow X\), if the fixed point groupoids of \(W \circ Z\) and \(Z \circ W\) are canonically equivalent. Dennis immediately gave the correct answer and said the proof was simple. The student replies “amazing! Most people end up struggling through the technical aspects of \((2,\infty)\)-categories”. Dennis replied “is there any another way?”. OK, well this might not have happened.

Posted in Mathematics | Tagged , , , , , , , , , , , | 1 Comment

Look at everything that OpenAI failed to prove!

I was heartened to see a recent list of results proved by OpenAI here. Heartened because we are now clearly in an age in which AI can truly contribute to serious research mathematics. But also heartened because one can only imagine how many thousand of open problems OpenAI tried and failed to solve, assuming these are the ones they consider the most impressive. It would be vastly more informative to see the list of problems they tried to solve and didn’t, as well as possible partial progress that was made on other problems.

Posted in Mathematics | Tagged , | 8 Comments

Xena on Counterexamples

There is an interesting post here on Xena’s blog.

I don’t centre “proof” in mathematics to anywhere near the extent that Kevin does, but I think emphasizing that distinction obscures the fact that our opinions about mathematics, mathematicians, and, I suspect, AI are quite similar.

It is fascinating to observe the progress of AI in mathematics (other people might use different adjectives). The result that a finite group scheme of order \(n\) need not be annihilated by \(n\) is the first AI-assisted result that I actually “care about.” I most closely associate this problem with a remark Hendrik Lenstra (the human equivalent of an LLM in the 90s) once made to me in the elevators of Evans Hall: namely, that there is a weaker result asserting that every finite flat group scheme \(G\) of order \(n\) is annihilated by \(n^{c(n)}\), for some integer \(c(n)\) independent of \(G\). The point is that there is, in an appropriate sense, a universal family of group schemes of order \(n\) over a finite-type base \(S\), where \(S\) is more or less the parameter space required to write down all the necessary Hopf-algebra data.

I absolutely agree with Kevin’s reaction to reading LLM-generated mathematics. Reading mathematics is already extremely hard; we are used to the fact that we might have to spend weeks understanding a few lines in a difficult paper. But what gives us the spirit to persevere is that we have been trained to give the author the benefit of the doubt: it is we who are missing the idea, and if only we think a little more about it, we will work it out. For better or worse, this is a cultural truth of mathematics. Once that is taken away (and it absolutely should be, for now, when reading LLM-generated mathematics) it becomes almost impossible to read mathematics. If there is one thing that LLMs have mastered, it is writing with the confidence and airs of someone possessing great authority.

Low-Hanging Fruit:

Some people have dismissed the recent AI-assisted breakthroughs because they were “obvious in retrospect” or because “not enough people tried to do them.” The first criticism seems clearly ridiculous, since “obvious in retrospect” is very far from “obvious.”

The second criticism reflects, in part, the way humans solve conjectures. In my experience, the process goes as follows. First, you learn about the conjecture and why it might be interesting. Second, you learn a little about why it is hard. You might think about it for a while and fail, and then file it away in the back of your mind. Later, when you encounter new ideas, you perform some pattern recognition and ask whether those ideas have any relation to previous problems you care about. If you are lucky, you “see a connection,” and then you can begin. Sometimes the first inspirational connection is most of the work; sometimes it is only the beginning, and there is much more work to do. In either case, that initial step is crucial.

An LLM, on the other hand (to anthropomorphize), can conduct an extremely thorough literature search and then throw a thousand different ideas at the problem to see whether anything sticks. If one of those ideas does seem relevant, and if the journey from that realization to the end of the proof is not a long one, then the LLM can solve the entire problem. As it becomes possible to chain together longer and longer stretches of reasoning, the only problems that can resist attack are those requiring a genuinely new idea (whatever that means).

Finally, there is the observation that these results are all counterexamples to conjectures that, as far as I know, were generally regarded as true. This phenomenon has a clear diagnosis, given by Deligne: “All problems in mathematics are psychological.” Well, to be precise, that quote was ascribed to Deligne* by Kisin in a lecture at Luminy, but it captures the essence of something clearly true. And the good news is that AI is an insane psychofreak with no hangups.

Addressing the more interesting—and controversial—question of what we, as a profession, are to do about all of this will have to wait until later. This week, I will be attending my first ICM in person. It is a much-maligned conference whose most exciting moment has already been blown by incompetent website design and which is being held in a city with slightly less appeal to me personally than St. Petersburg; but we shall see!

* Since I always check the original source, here is a slightly more nuanced version of that quote:


Dear Calegari,

I don’t remember the exact words, but I expect it is roughly accurate. Of course, it is not always true. The meaning was that we often have blocks which prevent us from seeing things which later will seem obvious to us.

Best,
Pierre Deligne

Posted in Mathematics | Tagged , , , , , , , , , , | 3 Comments