A Brief Digest of Some New OpenAI Results

I thought it might be worth discussing some of the recent OpenAI results in my field, or at least in algebraic number theory. For each of them, I’d like to offer a few thoughts, if I have any, on what seems to be new and what constitutes the breakthrough. My hands are currently shot with RSI, so I am narrating this to ChatGPT and asking to convert it to html, so let’s see how that goes

Added: I heard that OpenAI has already withdrawn some of their papers on (partial results towards) the Hodge Conjecture due to mistakes being found. Apparently they found a sign error! That may be the best evidence of sentience yet!

  1. Modularity of elliptic curves over imaginary quadratic fields.

    This problem was almost already solved by Caraiani and Newton in their paper. They prove the modularity of elliptic curves over any imaginary quadratic field \(F\) such that the elliptic curve \(X_0(15)\) has rank zero over \(F\). In particular, this applies to a positive proportion of imaginary quadratic fields.

    The problem is what to do with elliptic curves which, for example, have reducible mod \(3\) and mod \(5\) representations. This is the problem that OpenAI overcame, by what I think is a very nice trick.

    As already indicated in the comments on my last post, Bao Le Hung has understood this trick and has posted a corresponding paper on Hexagon Math. This argument, as Bao points out, seems quite robust, and is already enough to prove modularity of elliptic curves over all totally real fields and all CM fields not containing \(\zeta_5\) (and no doubt more tricks could be applied to that case as well).

    I view this result as definitely something a human could come up with. It’s a really nice trick rather than a fundamental breakthrough. I’m sure variations of it will be useful in other ways, and this is definitely worth exploring.

  2. Fontaine–Mazur modularity at the prime \(2\).

    It’s a little bit hard for me to see exactly what the new idea is in this 70-plus-page paper. Superficially, I’m not sure I find this particularly surprising in light of recent results of Lue Pan and others over the past few years. My first thought is that the main difficulties at \(p=2\) are technical. At the same time, I wouldn’t exactly say there’s a lot of money in proving this particular result, and so the few people most capable of possibly proving it have probably been thinking about other things. It is very clean, which is useful for applications.

    That said, upon closer inspection this paper may turn out to have some really clever new trick, but I can’t see what that is on a superficial reading. I would actually say I’m surprised that this is a problem OpenAI would have tried to solve, except for the fact that it gets used as an input for Hilbert’s tenth problem.

  3. Hilbert’s tenth problem over the rational numbers.

    This one threw me for a loop. I remember that a number of people, perhaps starting with Mazur and others, thought about this problem in the context of Hilbert’s tenth problem over \(\mathbf{Z}\). In particular, one of the issues became whether \(\mathbf{Z}\) was existentially definable over \(\mathbf{Q}\). All the talks I saw on this problem seemed to suggest that the answer to this was negative.

    As for Hilbert’s tenth problem itself, it seemed plausible to me that it was actually solvable over the rational numbers. I don’t think I would have put money on that, however. I still certainly hope that Hilbert’s tenth problem has a positive answer over the rational numbers for algebraic curves, and I think that this should be a consequence of conjectures that we believe.

    I hope to be able to encourage someone I know to write a guest blog post on this paper at some point. In fact, that might be true for a number of the other papers as well.

  4. Artin’s conjecture on primitive roots.

    When I first saw this, my impression was: well, of course Hooley’s argument should work if you knew the quasi-Riemann hypothesis for Dedekind \(L\)-functions. Except that OpenAI doesn’t prove the quasi-Riemann hypothesis for all Dedekind \(L\)-functions, just for Dirichlet \(L\)-functions. So instead, the argument is a little bit different, but also not at all surprising.

    The cheapest way to try to prove Artin’s conjecture is as follows. For example, let’s concentrate on \(2\) as the element one wants to be a primitive root. Take primes \(p\) of the form \(8k+3\). For each such prime, \(2\) will be a quadratic nonresidue. But now let’s imagine that \(p-1=2q\), where \(q\) is prime. If \(2\) is not a quadratic residue, the only way it could fail to be a primitive root is if it were a \(q\)-th power, in which case \(2^2\equiv 1 \pmod p\), which is ridiculous unless \(p=3\).

    So the “only thing” you need to prove is the existence of infinitely many pairs of primes \(p\) and \(q\) with \(p-1=2q\) and \(p\equiv 3\pmod 8\). Of course, this immediately returns you to another quite difficult problem. On the other hand, this is not so far in spirit from the work on bounded gaps between primes by Zhang, Maynard, and others.

    Instead, we could try to write \(p-1=2aq\), where now, instead of insisting that \(a=1\), we require that all the prime factors of \(a\) are reasonably large, for example of size at least \(\exp((\log p)^\delta)\) for some positive \(\delta\). The exact details of the size restrictions will be a combination of what one needs in Hooley’s argument and what one can actually prove. Here \(q\) is some large prime, itself of size at least some fixed power of \(p\). I think it might literally be larger than \(p^{9/10}\) in the actual paper.

    The point is that because all the prime factors of \(a\) are large, one doesn’t have to use GRH for Dirichlet \(L\)-functions to eliminate the possibility that \(2\) is a cube, a fifth power, a seventh power, and so on. One has reduced the problem to one of the other regimes in Hooley’s argument, where one doesn’t need the Riemann hypothesis, but instead can use elementary sieving arguments.

    Of course, one is now left with the difficult problem of finding primes \(p=2aq+1\), where \(a\) is not necessarily \(1\), but has prime factors of some restrictive shape. How does one do that? Understanding the prime divisors of \(p-1\) and finding primes of this form is certainly not easy. But this is where one is really choosing primes satisfying various congruences, and for this one can use the quasi-Riemann hypothesis for Dirichlet \(L\)-functions rather than for Dedekind \(L\)-functions. That, at least, is my first guess as to what is going on.

    So again, this is an interesting type of argument which I don’t find particularly surprising. But no one was going to write this paper, which depends on the quasi-Riemann hypothesis for Dirichlet \(L\)-functions, without knowing anything about general Dedekind \(L\)-functions. And this brings us to the next problem.

  5. The quasi-Riemann hypothesis.

    WTF.

    I think if you had asked me what was more likely, that AI would prove the Riemann hypothesis or that it would prove the quasi-Riemann hypothesis, I might honestly have thought that the first was more likely.

    My belief, and I suspect this is shared by a certain number of other people, is that the actual proof of the Riemann hypothesis is not going to be an analytic proof at all, but rather something much more arithmetic, involving ideas that we don’t necessarily understand yet. I think that in many situations in number theory, analytic techniques are not the right way of thinking about the problem. They may give the first nontrivial results, but are ultimately insufficient for the final proof.

    One example here is the Ramanujan conjecture concerning the size of the coefficients of \(\Delta\). There is a bound of order \(O(p^{6})\) first proved by Hardy. A significant improvement came from Rankin \(O(p^{6-1/5})\), using what we would now call Rankin’s method, involving the symmetric square in automorphic terms, although Rankin certainly didn’t think about it this way. I also just learnt right now that Kloosterman had also improved Hardy’s result to \(O(p^{6-1/8})\). These were analytic approaches. Of course, Langlands understood that one could also prove the Ramanujan bound from functoriality, but I don’t think of functoriality as a problem in analytic number theory either.

    When it comes to the Riemann zeta function, there are many results going back a long time concerning zero-free regions. Perhaps my point of view was that these represented something close to the limits of analytic arguments, and that one would need a more algebraic perspective to make a truly significant improvement. Obviously, I was totally wrong, and something was clearly missed. But I have no idea of the structure of what is going on. I still believe, although perhaps with slightly less conviction, that improvements to these analytic arguments are not going to be sufficient to prove the full Riemann hypothesis.

    Another thing that would be surprising is the idea that one might be able to prove the Riemann hypothesis for \(\zeta(s)\) without proving the corresponding hypothesis for Dirichlet \(L\)-functions.

    This surprising result puts us in an unusual situation. We now have a very strong intermediate result that no one ever imagined would be proved without also getting GRH. People have proved results assuming the generalized Riemann hypothesis, but no one has thought to prove results of the form: suppose you knew the quasi-Riemann hypothesis for Dirichlet \(L\)-functions, but nothing about Dedekind \(L\)-functions.

    What this means is that there may be quite a few results which have previously been proved under GRH that might now be provable using only the quasi-Riemann hypothesis for Dirichlet \(L\)-functions. Artin’s primitive root conjecture is exactly one of these things. I would imagine that, once an expert knew what had been proved, they would see that this application was not surprising at all.

    One of my favorite applications of the generalized Riemann hypothesis concerns the discriminant bounds of Stark and Odlyzko. But for all the applications I have in mind, which concern general number fields, one really needs the corresponding result for all Dedekind \(L\)-functions rather than merely Dirichlet \(L\)-functions. So we have no improvement on those bounds, at least as yet.

  6. Results on BSD.

    There are some results proving BSD under certain rank conditions. My first impression is that these are technical improvements on results already in the literature, and that they represent no genuine progress towards the higher-rank version of BSD. At this point I don’t have so much interest in looking at them myself, but I would be happy to have a guest blog post explaining anything that is new, or what the new trick or idea might be.

  7. Catalan’s constant is irrational.

    As we saw in the AI proof that \(\zeta(5)\) is irrational, there are clearly a bunch of new ideas, or at least new optimizations, that we were not previously aware of. I feel that some of these papers are among the least well written, in the sense that certain optimizations are taking place without being motivated at all.

    I am helping to run a conference at SLMath on irrationality in January, and one of our tasks will clearly be to try to understand what these new methods are, and whether they are simply new tricks or evidence of something more substantial. Of course, Apéry’s result from almost 50 years ago is still something we don’t really understand.

    I don’t know to what extent this is true of the other papers, but it’s amusing to see this paper cite two crank papers, both of which give nonsense proofs that Catalan’s constant is rational, and which AI casts some shade on.

    One thing I’m confused, or perhaps bemused, by is that I haven’t yet seen OpenAI give an alternative proof of the irrationality of \(L(2,\chi_{-3})\), although it wouldn’t be too surprising if such a proof existed.

  8. The irrationality measure of \(\pi\) is \(2\).

    This is an amazing result. I’ve yet to have time to look at it at all, but Wadim Zudilin has already noted to me (perhaps unsurprisingly) the close similarity between the arguments here and Roth’s argument for algebraic irrationals. Again, this is something that I hope we shall address at our upcoming conference.

Broadly speaking, for the problems in the Langlands program that I am interested in, while OpenAI did an amazing job, nothing it did at this point seems at all superhuman in the way that the proof of the quasi-Riemann hypothesis seems, and perhaps other results in other fields seem, although I can’t comment on those.

This entry was posted in Mathematics, Uncategorized and tagged , , , , , , , , , , , , , , , , , . Bookmark the permalink.

Leave a Reply

Your email address will not be published. Required fields are marked *