Are business schools intellectually bankrupt? (Part Two)

[I just found this draft of a blog post from 2017, I thought it might be a light diversion.]

There are essentially three types of people who claim proofs of the Riemann Hypothesis.

First, there are the cranks. The crank often writes something which is utterly incoherent — possibly invoking something from physics, possibly also simultaneously proving RH and Fermat’s Last Theorem at the same time. Alternatively, there are cranks who have some basic knowledge of mathematical formalism, and manage to scribble down something which at least shares some of the grammatical structure of a mathematical argument. My mental image of a crank used to be a retired 60 year old engineer, but, like so many other things, the internet has expanded my horizons, and nowadays cranks come in many different flavours.

Second, there are the amateurs with a sufficient amount of hubris that they somehow believe they can make a contribution a problem which — after a century of study — is clearly very deep and almost everyone believes will require fundamentally new ideas to crack. These people are often successful in their own careers — including Nobel prize winners in Chemistry (or was that a proof of Fermat?). This appears to be related to a (no doubt well-known) variation of the Dunning–Kruger effect — people with a high level of competence in one domain mistakenly overrate their competence in another (in this case number theory). This is usually a very bad mistake when it comes to higher mathematics. (Although don’t imagine that the shoe is never on the other foot — try to imagine for a moment otherwise smart mathematicians pontificating about biology.)

Third, there is the mentally ill. (Obviously there is often a non-trivial intersection between some or all of these classes.)

It is really not worth my while bothering to look at any purported proof of RH — it’s fairly clear that attempting to interact with any character of type one, two, or three above is not really worth one’s time. However, recent circumstances have unfortunately brought me to do so. It began when a probabilist was asked to review a short paper for MathSciNet in a Brazilian probability journal. The reviewer then noticed that the main claim of the paper was a proof of the Riemann hypothesis. Naturally he was confused! The author of the paper was a professor at the University of Chicago. Not of the mathematics department, of course, nor of the statistics department, but of the business school! (For more on business schools, see this post.) The reviewer decided to consult Reddit, who suggested forwarding it to me:

Seeing as the author is at U Chicago (but not in the math department), just email Calegari and inform him that a “colleague” of his has somehow managed to get a proof of RH published in a low-tier probability journal. He’ll deal with it (or get someone to), that’s a professional embarrassment for U Chicago.

Well, of course, I do whatever Reddit tells me to do (/s)

I place the argument — to the extent that I can tell — in the second class. Half an hour of study was sufficient to determine a hole big enough to drive a lorry through. It was the type of mistake that I might have made (but didn’t) when I was 15. There’s a certain amount of probabilistic window dressing in the paper which is syntactically related to real mathematics but tangential to RH. Once this window dressing is removed, the argument is essentially as follows:

  1. Take the function \( 1/\zeta(s)\), then take its inverse Mellin transform.
  2. Assume without comment properties of this inverse Mellin transform which require RH.
  3. Deduce RH by considering the Mellin transform again.

In fact, the good news is that the argument can be upgraded in a smaller number of pages so that it proves that all the zeros of the function \((1/4-s)(3/4-s)\) lie on the critical line. Now that would be a spectacular result!

Let’s prove it! We can start at around equation (2.35) in the paper, where the function

\[ \displaystyle{G(x) = \sum_{k=1}^{\infty} \frac{ (-1)^{k+1}}{\Gamma(k)} \frac{\xi(1/2) \xi(k + 3/2)}{\xi(k + 1/2)} x^k,}\]

is defined, with corresponding Laplace transform

\[\displaystyle{m_G(s) = \int_{0}^{\infty} x^{s-1} G(x) dx.}\]

Here \(\xi(s)\) is the part of the zeta function coming from the Hadamard factorization (i.e. without the Gamma factors). Now let’s imagine instead that \(\xi(s)\) is just the function \(3/4 – s\).
A standard calcuation (for either the \(\xi\) coming from Riemann or from \(3/4 – s)\) gives

\[ \displaystyle{m_{G}(s) = \int_{0}^{\infty} x^{s-1} G(x) dx = \frac{\xi(3/2 – s) \Gamma(1+s) \xi(1/2)}{\xi(1/2 – s)}}.\]

In particular, this is a purely “formal” calculation (Ramanujan Master Theorem Style). The crux of the argument, the part where something “gets done” and one gets access to \(\xi\) in the critical strip is where he “eliminates” \(m_{G}(s)\) by using the identity

\[\displaystyle{
\frac{1}{\xi(1/2 – s)} = \frac{m_G(s)}{\xi(1/2)\Gamma(1+s) \xi(3/2 – s)}.}\]

He really wants to use this identity (see the line before (2.45), where variables have been changed slightly) in the range

\[ \displaystyle{s \in (-1/2,0).}\]

After all, he wants to understand \(\xi\) — even if just on the real line — in the critical strip. In order to get this, you need to know something about the growth of \(G(x)\) as \(x\) goes to infinity, because you want the Mellin transform to be well defined. In order to get convergence for real negative s close to zero (which he uses), you certainly want to assume that

\[\displaystyle{G(x) = O(x^{\epsilon})}.\]

So let’s assume exactly this. And now let me prove that \(3/4 – s\) has no zeroes for \(1/2 \le \mathrm{Re}(s)\le 1\) which is nonsense. From our equation above, \(m_G(s)\) is now well defined for \(\mathrm{Re}(s)\) in \((-1/2,0)\). But then the RHS is well defined in this range, so the LHS has no poles, so \(\xi(1/2 – s)\) has no zeroes for \(-1/2 \le \mathrm{Re}(s) \le 0\), or \(\xi(s)\) has no zeroes for \(1/2 \le \mathrm{Re}(s) \le 1\).. Done!

So what is wrong with this argument? Obviously one actually has to understand the growth of \(G(x)\). In the case of \(\xi(s) = r – s\), one can compute \(G(x)\) explicitly, and at least away from \(1/2\) where symmetry forces \(G(x)\) to vanish one can compute it explicitly in terms of incomplete Gamma functions and get

\[\displaystyle{G(x) \sim x^{\mathrm{Re}(s) – 1/2}}.\]

Of course, this is no surprise, since this exactly eliminates the contradiction. So now let us return to the paper. We have the function \(G(x)\), and we need to say something
about its growth at infinity. In order to prove RH we need to show that it grows slower than any power of \(x\). So what is going to happen? Well, we are going to have that

\[\displaystyle{G(x) = O(x^{\rho – 1/2 + \epsilon})},\]

where \(\rho\) is the supremum over the real parts of all the non-trivial zeros of \(\zeta(s)\). So, in order to prove RH, one only needs to prove … the Riemann Hypothesis!

Note that understanding the growth of functions like \(G(x)\) and their link to RH is not new. In fact, already over 100 years ago, Riesz proves the following. Let

\[ \displaystyle{ F(x) = \sum_{n=1}^{\infty} (-1)^n \frac{x^n}{\zeta(2n) \Gamma(n)}.}\]

Then

\[ \displaystyle{F(x) = O(x^{1/4 + \epsilon})}\]

is equivalent to RH. What I have sketched above is basically already a moral explanation of this argument — poles of a function imply growth of the inverse Mellin transform and vice versa. Indeed, Grosswald (in the paper cited by Polson!) proves that the rate of growth (up to \(\epsilon\)) of \(F(x)\) is exactly \(x^{\theta/2}\) where \(\theta\) is the supremum of the real part of zeros of \(\zeta(s)\).

Well, now at least we know the answer to “what does it take to get you to look at my proof of RH?” The answer: you have to have tenure at Chicago, and you have to have a published proof of RH.

One month later: This, at least, was the original story as of a month ago. But there was a twist. Greg Lawler and I actually contacted the author of this paper. Communications via email were not particularly successful. He actually produced a second purported proof (!?) which was worst than the first — basically writing down integrals for a complex parameter s related to \(1/\zeta(s)\) paying no attention to the domains of applicability, and then using (in effect) precisely facts about convergence of these integrals which require being careful about the domain of applicability. But then we met in person, and he was very polite, and seemed to realize that both approaches were flawed. The original published paper was retracted by the author, and balance was restored. Success! Or at least I thought so, until I just found out that he recently updated his second paper with more of the same claims! (A little effort — more than it is worth — shows that simply by changing some 1/2s to 1/3rds or 2/3rds one can prove that \(\zeta(s)\) doesn’t have any zeroes at all.)

So where does this lead us with respect to the question in the title? I guess in the context where producing banal observations about human behavior that have been repeatedly observed by others gets you a (not really a) Nobel Prize, paying someone $400000 a year (Note: guesstimate) to produce quisquilian proofs of the Riemann Hypothesis sounds perfectly sane. C’est la vie.

Nine years later: Apparently Polson is back in the news for authoring 258 papers in 2026. Well, I guess when you can prove the RH, nothing is beyond you!

Posted in Uncategorized | Tagged , , , , , , | 3 Comments

zeta(5) is irrational

So one day after an amazing result by humans, the computers have struck back! This time, a proof that \(\zeta(5)\) is irrational. The paper I saw gives me the impression of being almost entirely AI generated (if so, it then comes with a very dishonest disclosure), but never mind, thanks for the compute!

I spent some effort trying to see what was going on. Since the argument has now been verified in Lean, it seemed more useful to try to identify the key ideas and where they could be traced, and where they were new.

One way I tried to understand this was to insist that ChatGPT try to find in the literature where similar ideas had been used before. I am sure that I did not do a great job with the literature, but I did my best in the several hours I had available. I sure hope Jacob Tsimerman’s suggestion that AI might reach superhuman exposition will come soon. The discussion below connects the arguments to the work of Prévost (not cited in the paper), but there may well be other sources that ChatGPT learnt these ideas from. Certainly Hankel determinants are also considered by Zudilin, and I suspect that he would do a better job unearthing the ideas behind this paper (some of which are no doubt in his own papers) than I have done. Later on I quote from ChatGPT directly on what it thinks.

Put \(D_0(t)=1\) and \(D_m(t)=\prod_{j=1}^m(t+j^2)\) for \(m\geq1\). Fix an integer \(k\geq2\), and let \(H_n^{(k)} =\sum_{v=1}^n v^{-k}\) with \(H_0^{(k)}=0\), and define the Bernoulli numbers as usual by \(u/(e^u-1)=\sum_{m\geq0}B_mu^m/m!\), so \(B_1=-1/2\).

Definition. Let \(t\) and \(X\) be independent indeterminates. Define the \(\mathbf{Q}\)-vector space

\[\begin{gathered} \mathcal R=\bigcup_{m\geq0}\left\{\frac{P(t)}{D_m(t)}:P(t)\in\mathbf{Q}[t]\right\}\subset\mathbf{Q}(t),\\ \mathbf{Q}[X]_{\leq1}=\{aX+b:a,b\in\mathbf{Q}\}. \end{gathered}\]

Thus \(\mathcal R\) consists of the rational functions of \(t\) whose finite poles, if any, are simple and belong to \(\{-1^2,-2^2,\ldots\}\). There is no restriction on the polynomial part. For the fixed integer \(k\geq2\), define the \(\mathbf{Q}\)-linear map \(\mu_{k,X}:\mathcal R\longrightarrow\mathbf{Q}[X]_{\leq1}\) by

\[\begin{aligned}\mu_{k,X}(t^e)&=(-1)^eB_{2e+2}\frac{(2e+k)!}{(k-1)!(2e+2)!}&& (e\geq0),\\ \mu_{k,X}\left(\frac1{t+j^2}\right)&=j^{k-1}(X-H_j^{(k)})-\frac1{k-1}+\frac1{2j}&& (j\geq1).\end{aligned}\]

Write \(\mu_k:\mathbf{Q}[t]\to\mathbf{Q}\) for its restriction to polynomials, and abbreviate \(\mu_{k,X}((t+j^2)^{-1})\) to \(\phi_{k,j}(X)\). The fact that this is well-defined follows from the partial fraction expansion. If one writes \(R\in\mathcal R\) as \(R(t)=P(t)+\sum_{j=1}^m c_j/(t+j^2)\), with \(P(t)=\sum_e p_et^e\in\mathbf{Q}[t]\) and \(c_j\in\mathbf{Q}\), then

\[\begin{aligned} \mu_{k,X}(R)={}&\left(\sum_{j=1}^m c_jj^{k-1}\right)X+\sum_e p_e\mu_k(t^e)\\ &+\sum_{j=1}^m c_j\left(-j^{k-1}H_j^{(k)}-\frac1{k-1}+\frac1{2j}\right). \end{aligned}\]

For a real number \(\xi\), let \(\operatorname{ev}_\xi:\mathbf{Q}[X]\to\mathbf{R}\) be evaluation at \(X=\xi\), and define \(\mu_{k,\xi}=\operatorname{ev}_\xi\circ\mu_{k,X}\). Thus \(\mu_{k,\xi}(R)\) is a real number. For any field \(F\) containing \(\mathbf{Q}\), the same basis and formulas define an \(F\)-linear map on \(\mathcal R_F=\bigcup_m D_m(t)^{-1}F[t]\), with values in \(F[X]_{\leq1}\). We keep the notation \(\mu_{k,X}\) for this extension and \(\mu_k\) for its restriction to \(F[t]\). When \(F=\mathbf{R}\), we can specialize \(X=\xi\in\mathbf{R}\) as above. In particular, this specifies the map on rational functions with real coefficients.

The formulas of Euler and Hermite

Let \(f(y)=(e^{2\pi y}-1)^{-1}\) for \(y>0\), and put

\[\begin{aligned} w_k(y)&=\frac{2(-1)^{k-1}y^k}{(k-1)!}f^{(k-1)}(y)\\ &=\frac{2(2\pi)^{k-1}y^k}{(k-1)!}\sum_{l\geq1}l^{k-1}e^{-2\pi l y}. \end{aligned}\]

The density \(w_k(y)\) is positive for \(y>0\), tends to \(1/\pi\) as \(y\to0\), and satisfies \(w_k(y)=O_k(y^ke^{-2\pi y})\) at infinity. Euler’s and Hermite’s formulas give the following integral representation.

Proposition. For \(R\in\mathcal R_{\mathbf{R}}\),

\[\mu_{k,\zeta(k)}(R)=\int_0^\infty R(y^2)w_k(y)\mathrm{d} y.\]

For \(a>0\), write \(\zeta(k,a)=\sum_{v\geq0}(v+a)^{-k}\). The corresponding transform identity is

\[\int_0^\infty\frac{w_k(y)}{y^2+a^2}\mathrm{d} y=a^{k-1}\zeta(k,a)-\frac1{2a}-\frac1{k-1}.\]

The transform identity is the integer-weight specialization of Prévost’s Stieltjes representation [3] (Theorem 1).

Hankel Determinants

Choose integers \(0\leq N<K\), \(r\geq1\), and \(h\geq1\). Let \(\mathcal P_h=\mathbf{Q}[t]_{<h}\) be the \(h\)-dimensional vector space of polynomials of degree less than \(h\). Multiplication by \(W(t)=D_N(t)^r/D_K(t)\) sends every polynomial to \(\mathcal R\). Therefore

\[\beta_{k,X}:\mathcal P_h\times\mathcal P_h\longrightarrow\mathbf{Q}[X]_{\leq1},\qquad \beta_{k,X}(p,q)=\mu_{k,X}(Wpq)\]

is a well-defined symmetric bilinear map. Its matrix in the basis \(1,t,\ldots,t^{h-1}\) and its determinant are

\[\begin{gathered} G_k(X)=\left[\mu_{k,X}\left(\frac{D_N(t)^rt^{i+j}}{D_K(t)}\right)\right]_{0\leq i,j<h},\\ \Delta_k(X)=\det G_k(X)\in\mathbf{Q}[X]. \end{gathered}\]

Such a matrix is called Hankel because its entries depend on \(i+j\). (The notation suppresses \(K,N,r,h\).) Since the entries have degree at most one in \(X\), the determinant has degree at most \(h\).

Lemma. After scalar extension to \(\mathbf{R}\) and specialization \(X=\zeta(k)\), the bilinear map \(\beta_{k,\zeta(k)}\) is an inner product. For every nonzero real polynomial \(q\) of degree less than \(h\),

\[\beta_{k,\zeta(k)}(q,q)=\int_0^\infty W(y^2)q(y^2)^2w_k(y)\mathrm{d} y>0.\]

Thus \(G_k(\zeta(k))\) is a positive-definite Gram matrix, meaning the matrix of pairwise inner products of a linearly independent list. In particular \(\Delta_k(\zeta(k))>0\).

Allowing for a general rational \(W\), Prévost’s arguments for \(\zeta(2)\) and \(\zeta(3)\) fit this same Hankel construction, using the one-pole choices \(W_{2,n}(t)=1/(t+(2n+1)^2)\) and \(W_{3,n}(t)=t/(t+(n+1)^2)\), respectively, with matrices of size \(n+1\) [3]. These determinants are affine polynomials in \(X\), and their roots are the corresponding rational approximations to \(\zeta(k)\): a change of polynomial basis expresses each determinant as a rational factor times a Padé approximation error. Prévost’s coefficient and error estimates allow these determinants to be rescaled into integer linear forms whose nonzero values at \(\zeta(2)\) or \(\zeta(3)\) tend to zero, proving irrationality. On the other hand, in this new argument, the single pole of \(W\) is replaced by the \(K-N\) poles of \(D_N(t)^r/D_K(t)\), producing a determinant of degree \(K-N\) in \(X\).

Proposition (Degree of the determinant). Take \(h=K-N\). Then \(\Delta_k\) has degree \(h\), with

\[[X^h]\Delta_k(X)=(-1)^{h(h-1)/2}\prod_{j=N+1}^Kj^{k-1}D_N(-j^2)^{r-1}\ne0.\]

For a nonzero polynomial \(\Delta\in\mathbf{Q}[X]\), a positive rational number \(c\) with \(c\Delta\in\mathbf{Z}[X]\) will be called a normalizing factor. For example, one can multiply by the least common multiple of the coefficient denominators. One can also divide out a common factor of the resulting integer coefficients; such division is important in this construction. The required estimate concerns the value of the normalized polynomial, including both effects.

Lemma. Let \(\xi\in\mathbf{R}\). Suppose \(P_n\in\mathbf{Z}[X]\), \(\deg P_n\leq Cn\), and \(0<P_n(\xi)\leq\exp(-cn^2+o(n^2))\), where \(C,c>0\) are fixed. Then \(\xi\) is irrational.

The strategy

The goal is to apply the integer-polynomial lemma with \(\xi=\zeta(k)\) and with \(P_n\) a positive rational multiple of the determinant \(\Delta_k(X)\) in the Hankel matrix. “All” that has been done is to adjust Prévost’s positive measure by \(D_N^r/D_K\) for suitable parameters \(N\) and \(K\). Increasing the numerator power actually makes the determinant larger, but it turns out that one obtains extra cancellations in the coefficient denominators that compensate. Establishing those cancellations uniformly is required for the proof, of course, but finding the right construction is the key step. Now it might seem as though this new factor has come out of nowhere, but even this is not actually new, and the rows of this matrix can be identified with well-poised hypergeometric integrals considered by Zudilin. Once the required arithmetic estimates are proven, one deduces the irrationality of \(\zeta(k)\) with \(k=2,3,4,5\), but not (at least directly) for larger \(k\).

Prévost’s moments and the introduction of the zeta value

We now attempt some mathematical archaeology. Prévost’s Theorem 1 in [3] gives the Hurwitz-zeta representation above. Write \(\mathrm{d}\nu_k(t)=w_k(\sqrt t)\,\mathrm{d}t/(2\sqrt t)\) for \(t>0\). Its polynomial moments are

\[m_j:=\int_0^\infty t^j\mathrm{d}\nu_k(t)=\mu_k(t^j)=(-1)^jB_{2j+2}\frac{(2j+k)!}{(k-1)!(2j+2)!}\in\mathbf{Q}.\]

These are rational for every integer \(k\geq2\). Their Hankel determinants compute rational orthogonal polynomials. Put \(H_s=\det[m_{i+j}]_{0\leq i,j<s}\), with \(H_0=1\). The determinant formula in Prévost’s equation (2.11), written for this moment sequence, is

\[p_s(z)=\frac{1}{H_s} \det\begin{pmatrix} m_0&m_1&\cdots&m_s\\ m_1&m_2&\cdots&m_{s+1}\\ \vdots&\vdots&&\vdots\\ m_{s-1}&m_s&\cdots&m_{2s-1}\\ 1&z&\cdots&z^s \end{pmatrix}.\]

The polynomial is monic of degree \(s\) and satisfies \(\int t^j p_s(t)\mathrm{d}\nu_k(t)=0\) for \(j<s\); these are its orthogonality conditions. Both \(H_s\) and the coefficients of \(p_s\) are rational. The zeta value enters through a subsequent evaluation of a transform, as follows.

Define \(\mathcal F_k(z)=\int_0^\infty(z+t)^{-1}\mathrm{d}\nu_k(t)\) for \(z>0\). At a positive integer \(a\), we get

\[\mathcal F_k(a^2)=a^{k-1}\bigl(\zeta(k)-H_a^{(k)}\bigr)-\frac1{k-1}+\frac1{2a}=\phi_{k,a}(\zeta(k)).\]

Consequently a rational approximation to \(\mathcal F_k(a^2)\) gives a rational approximation to \(\zeta(k)\). A Padé approximation matches finitely many coefficients of the asymptotic expansion \(\sum_{j\geq0}(-1)^jm_jz^{-j-1}\), as \(z\to+\infty\), by a rational function. The coefficients of that rational function are determined by the moment equations above. Prévost–Rivoal give effective approximation bounds when the Padé degree and \(a\) increase together [4].

Replacing \(\mathrm{d}\nu_k(t)\) by \(\mathrm{d}\nu_k(t)/(t+a^2)\), polynomial division gives the affine moment

\[s_{a,j}(X):=\mu_{k,X}\left(\frac{t^j}{t+a^2}\right) =\sum_{l=0}^{j-1}(-a^2)^l m_{j-1-l}+(-a^2)^j\phi_{k,a}(X),\]

where the sum is empty at \(j=0\). At \(X=\zeta(k)\) this is exactly the \(j\)th moment of the corresponding measure.

The measure in the new argument is the following explicit modification of the same \(\nu_k\):

\[\mathrm{d}\sigma_{k,K,N,r}(t)=W(t)\mathrm{d}\nu_k(t),\qquad W(t)=\frac{D_N(t)^r}{D_K(t)} =\frac{\displaystyle\prod_{a=1}^N(t+a^2)^{r-1}} {\displaystyle\prod_{j=N+1}^K(t+j^2)}.\]

Here \(0\leq N<K\) and \(r\geq1\) are integers. All factors are positive on the integration interval \(t\geq0\). For fixed \(K,N,r\) the moments are \(\mu_{k,\zeta(k)}(Wt^j)\), and their Hankel determinant of size \(h=K-N\) is \(\Delta_k(\zeta(k))\). These are the moments used throughout the argument.

ChatGPT tells me: Uvarov studies the effect of precisely this type of rational multiplication on orthogonal polynomials [7]. Krattenthaler’s Theorem 1 and Proposition 13 express the modified Hankel determinant through the original \(p_s\) and their Cauchy transforms \(C_s(y)=\int p_s(t)/(y-t)\mathrm{d}\nu_k(t)\), for \(y<0\) [8]. For the displayed modifier, each numerator node \(-a^2\) gives \(r-1\) rows \(p_s^{(d)}(-a^2)/d!\) for \(0\leq d<r-1\), and each denominator node gives a Cauchy-transform row. Thus the measure considered here belongs to an established family of rational modifications. The starting measure and remainder approximation come from Prévost and Prévost–Rivoal [2], [3], [4]. The exact algebra for multiplying that measure by \(D_N^r/D_K\) is supplied by Uvarov’s rational modifications and Krattenthaler’s determinant identity [7], [8]. Zudilin then supplies the irrationality criterion for a positive moment determinant with sufficiently small coefficient denominators [5]. These results identify the starting approximation problem, the modified determinant, and the final integer-polynomial argument, respectively. The closest precedent for improving the coefficient normalization of the matrix is Brown’s Sections 7.4–7.5 [6]. There one first clears row denominators, then removes column common factors and uses congruences between rows or columns to reduce the total multiplier. Here the concrete mechanism uses congruences between square nodes, while also accounting for the polynomial parts in the partial-fraction expansions. Applying \(\mu_k\) to those polynomial parts can introduce Bernoulli denominators. The required quantitative input is that the assigned divisibility holds simultaneously for every mixed pairing, with these contributions included.

To come back to the argument, one can (uniformly, for \(k=2,3,4,5\)) take \(K=12n\), \(N=n\), \(r=6\), and \(h=11n\), and write \(\Delta_{k,n}\) for the corresponding determinant. (This is different from what is done in the Lean verification, but whatever.) Set

\[S_n=\frac{(K!)^{2h}4^{h-1}}{(N!)^{12h}\prod_{i=1}^{h-1}((2i)!)^2},\qquad F_{k,n}(X)=S_n\Delta_{k,n}(X).\]

One then has to estimate this at the real place and then control the denominators. In particular:

Theorem: \(\zeta(k)\) is irrational for \(k \in \{2,3,4,5\}\). For all sufficiently large \(n\),

\[\begin{aligned}P_{k,n} & \in\mathbf{Z}[X], \\
\deg P_{k,n} & =11n \\ P_{k,n}(\zeta(k))& \le e^{-\gamma_k n^2} \\ (\gamma_2,\gamma_3,\gamma_4,\gamma_5)& =(30,96,6,3/2).\end{aligned}\]

References

[1] F. Beukers, A note on the irrationality of \(\zeta(2)\) and \(\zeta(3)\), Bull. London Math. Soc. 11 (1979), 268–272. doi:10.1112/blms/11.3.268.

[2] M. Prévost, A new proof of the irrationality of \(\zeta(2)\) and \(\zeta(3)\) using Padé approximants, J. Comput. Appl. Math. 67 (1996), 219–235. doi:10.1016/0377-0427(95)00019-4.

[3] M. Prévost, Remainder Padé approximants for the Hurwitz zeta function, Results Math. 74 (2019), article 51; arXiv:1709.05389v1.

[4] M. Prévost and T. Rivoal, Diagonal convergence of the remainder Padé approximants for the Hurwitz zeta function, J. Number Theory 222 (2021), 346–361. doi:10.1016/j.jnt.2020.10.019.

[5] W. Zudilin, A determinantal approach to irrationality, Constr. Approx. 45 (2017), 301–310. arXiv:1507.05697v2; doi:10.1007/s00365-016-9333-7.

[6] F. Brown, Mellin transforms, transfinite diameter and rational approximations of integrals, 2026, arXiv:2604.20741v1.

[7] V. B. Uvarov, The connection between systems of polynomials that are orthogonal with respect to different distribution functions, U.S.S.R. Comput. Math. Math. Phys. 9, no. 6 (1969), 25–36. doi:10.1016/0041-5553(69)90124-4.

[8] C. Krattenthaler, A determinant identity for moments of orthogonal polynomials that implies Uvarov’s formula for the orthogonal polynomials of rationally related densities, 2021. arXiv:2103.03969v1.

Posted in Mathematics | Tagged , , , , , | 2 Comments

Real quadratic fields and finite quantum dilogarithms I

Danylo Radchenko and Campbell Wheeler have posted an extraordinary new paper in which they prove that Stark units for real quadratic fields are algebraic numbers. (Not yet a Shimura reciprocity law that shows these generate abelian extensions but hey, this is only part I.) This, to me, is arguably the most significant mathematical result of the past month. There is a lot of rich mathematics going on here, and I apologize for not having the time to even summarize it. But I didn’t want it to go unmentioned!

Posted in Mathematics | Tagged , , , , , , , | 3 Comments

An even PSL_2(F_23) Galois extension

I was wrong!

I have said on a number of occasions that my favourite family of groups for the inverse Galois problem is \(G=\mathrm{SL}_2(\mathbf{F}_p)\). This is still true, and the reason is still the same; the fact that \(G\) has no non-trivial non-central involutions means that any such extension must be totally real up to twist, and thus it cannot come (for example) as the mod-\(p\) reduction of some strongly compatible system (assuming \(p > 5\)). For a similar reason, the group \(\mathrm{PSL}_2(\mathbf{F}_p)\) should also present difficulties under the additional assumption that the field is totally real, and I used this to argue that a totally real extension with Galois group \(\mathrm{PSL}_2(\mathbf{F}_{23})\) would be more interesting to me than any other Galois closure of a degree \(24\) field. It was then a delight to receive an email from Eray Karabiyik (a recent student of Zywina from Cornell, soon to be a postdoc at HIMIS, Shenzhen) giving me an explicit degree \(24\) polynomial with all roots real whose splitting field had Galois group \(\mathrm{PSL}_2(\mathbf{F}_{23})\)!

Here is an introduction to his construction. The goal is to construct (more or less) a representation

\[\rho: G_{\mathbf{Q}} \rightarrow \mathrm{GL}_2(\mathbf{F}_{23})\]

which is even, has image containing \(\mathrm{SL}_2(\mathbf{F}_{23})\), and whose determinant lands in the squares \((\mathbf{F}_{23}^{\times})^2\). Its projectivization then has image exactly \(\mathrm{PSL}_2(\mathbf{F}_{23})\). As mentioned above, this should not come from a regular rank \(2\) (with coefficients) motive. But there is no such parity obstruction to its appearing as a constituent of the mod-\(23\) reduction of a rank \(4\) motive! So, for example, one could hope to find an abelian surface \(A\) with

\[A[23]^{\mathrm{ss}} \simeq \rho \oplus \rho^{\vee}(1)\]

where \(\rho\) is even and \((1)\) denotes the mod-\(23\) cyclotomic twist. Suppose for now that \(A\) and a principal polarization are defined over \(\mathbf{Q}\).

One way to find \(A\) so that \(A[p]\) breaks up is for its geometric endomorphism ring to be the ring of integers \(\mathcal{O}_K\) of a real quadratic field \(K\). If this action is defined over \(\mathbf{Q}\), however, then one just gets (once again) rank \(2\) compatible systems which are odd. Suppose instead that the action of \(K\) is only defined over a quadratic field \(E\). Now, for a prime \(p\) that splits in \(K\), \(A[p]\) decomposes over \(E\) into two factors which are interchanged by the action of \(\mathrm{Gal}(E/\mathbf{Q})\). On the other hand, at an inert prime, one gets a two-dimensional representation of \(G_E\) over \(\mathbf{F}_{p^2}\), with the non-trivial coset of \(G_E\) acting semilinearly through the field automorphism, and generically with large image.

But if \(p\) is ramified in \(K\), with \((p)=\pi^2\), then the subspace \(A[\pi]\) does give a representation \(\rho:G_{\mathbf{Q}}\rightarrow\mathrm{GL}_2(\mathbf{F}_p)\). Writing \(\omega\) for the mod-\(p\) cyclotomic character and \(\chi\) for the quadratic character of \(E\), one obtains

\[0\longrightarrow\rho\longrightarrow A[p]\longrightarrow\rho\otimes\chi\longrightarrow0,\qquad \det\rho=\omega\chi.\]

Thus \(\rho\) is even iff \(E\) is imaginary. The condition that the determinant be a square is precisely \(\chi=\omega^{(p-1)/2}\) with \(p=23\). So one is indeed in with a chance assuming:

  1. \(K\) is a real quadratic field with discriminant divisible by \(23\),
  2. \(E=\mathbf{Q}(\sqrt{-23})\).

In this case \(\det\rho=\omega^{12}\) and \(\rho\otimes\chi\simeq\rho^{\vee}(1)\), as desired. Of course, large image still has to be checked.

To find surfaces \(A\) with endomorphisms by \(K\), it helps to have access to the corresponding Hilbert modular surface. One can also insert \(E\) into the story by suitable twisting. This is analogous to how one can look for \(\mathbf{Q}\)-curves defined over a quadratic field \(E\), admitting an isogeny of degree \(q\) to their conjugate, by twisting \(X_0(q)\) by the Atkin–Lehner involution and the character of \(E\). The quotient \(X_0^+(q)\) forgets which conjugate one started with.

Here Eray takes \(K=\mathbf{Q}(\sqrt{23})\), for which Elkies and Kumar give an explicit model of \(Y_-(92)\). There are two relevant involutions: one exchanges the two RM embeddings, while the other switches the two principal polarization classes on the generic unpolarized RM surface. The latter correspond to totally positive units modulo squares (the fundamental unit \(\epsilon\) is totally positive): starting with a principal polarization \(\lambda\) and RM embedding \(\iota\), the other class is represented by \(\lambda\circ\iota(\epsilon)\). Both leave the unpolarized surface unchanged, corresponding to “\(q=1\)”. Eray searches on the appropriate non-K3 involution quotient, twisted by \(E\) using the remaining involution. If I understood him, this turns out to be (birational to) a genus two fibration over \(\mathbf{P}^1\). He then finds a point which gives rise to such an \(A\).

Now I have simplified Eray’s example, because his \(A\) is defined over an auxiliary quadratic field \(H=\mathbf{Q}(\sqrt{18193})\), with the RM defined over \(EH\). This is presumably either related to field of moduli versus field of definition issues at low level, or (somewhat relatedly) lifting rational points from the quotient by the involution back to the cover. It affects some of the discussion of Galois representations above, but (in favorable situations) is no longer visible when one considers projective representations.

As \(p\) gets larger, one will either have to make \(K\) larger (so that \(p\) is ramified) or take a non-maximal order like \(\mathbf{Z} + p \mathcal{O}_K\). In either case, the corresponding Hilbert modular surfaces (and all their quotients by involutions) will have general type for large enough \(p\), and my guess is that rational points away from the special loci on the relevant Hilbert modular surfaces and their twisted quotients will become increasingly scarce and most likely eventually empty. So I suspect this is unlikely to produce totally real \(\mathrm{PGL}_2(\mathbf{F}_p)\) or \(\mathrm{SL}_2(\mathbf{F}_p)\) extensions for all (or even infinitely many) \(p\), but it’s a very nice example nonetheless.

Posted in Mathematics, Uncategorized | Tagged , , , , , , , | Leave a comment

Ways to improve Mathematics (an introduction)

The goal of this continuing sequence of posts is to recognize the reality in how our subject is changing and how we should adapt.

Even in a hypothetical world where AI was infallible, essentially omniscient, benevolent, a wonderful expositor, and freely available, I believe that many people would still want humans to maintain and develop a deep understanding of mathematics. They would not want to leave mathematics entirely to the machines, any more than they would want to abandon other large swathes of human thought. For this post, I will take that desire as a starting assumption.

My core belief is that understanding mathematics is hard and AI is not going to fundamentally change that, even if it makes proving theorems in mathematics much easier. I remember driving to Wisconsin and listening to Jordan Ellenberg on Lex Fridman’s podcast. What struck me most from the podcast was the discussion of Fermat’s Last Theorem. Fridman was essentially arguing that since the statement of Fermat’s Last Theorem was so simple, there must inherently be a simple explanation of why it was true. This reflects a philosophical idea about science and mathematics that I think is fundamentally untrue: that any truth that is simple to state will ultimately be true for a simple reason. The easiest proof of Fermat may well not be the one found by Wiles (or maybe it will be). But consider instead something much older and established in mathematics, namely class field theory. If I hold any position in this post with conviction, it would be that, even with superhuman exposition, a human could not acquire a good understanding of the statements and proofs of class field theory without years of dedicated study^*. This is not an isolated example. Maintaining and developing human understanding across mathematics requires people who can devote substantial parts of their lives to it. If we value that understanding, there is a case for supporting those people and the communities in which they work, rather than leaving the whole enterprise to whoever wants to work on it in their spare time. That is a central part of the case for mathematics as a profession.

What is important to recognize, however, is that mathematics as a profession will surely be changing rapidly. Daniel Litt has told us that we need to be honest not only about the aspects of our field that will break with AI, but also about the aspects that are already broken. This is an amazing opportunity to fix some of these things! If we are going to defend mathematics, at least in part, as a means of increasing human understanding, then we ought to be rather more demanding about whether our own practices actually contribute to this understanding.

To me, there are (at least) four natural systemic issues in mathematics that we have to address. Addressing any one of them will require consensus building to achieve the massive shift necessary in our community norms. I plan to devote a blog post to each of the issues, so for now I will just introduce them. I am not attempting to be prescriptive, but rather simply to think aloud and hope for suggestions.

One thing we need to do in particular is ask whether our institutions and practices reward the production of mathematical understanding, or merely activities that we have come to treat as evidence of it. With AI, activities such as producing proofs will no longer be quite so closely coupled with understanding.

  • Journal Articles: The journal system was, to put it politely, already struggling before AI. There are too many papers, and the refereeing system is bursting at the seams. At least in mathematics we have a large number of high-quality journals, which helps diffuse the power of editorial boards, given the career importance of publications. [Imagine a field where success could come only from publishing in a single journal; such fields exist.] As we go forward, we very much have an opportunity to reconsider how the publication system works in a radical way. If we don’t abandon it altogether, then we need to redefine what we hope to get out of it. I already have anecdotal evidence that submission rates to top journals are rising sharply, and I doubt that the current refereeing system can sustain such an increase. I don’t think it is as simple as demanding better exposition — it’s not unreasonable to expect AI to vastly improve in this dimension as well. If mathematics is a conversation, then what we would like to achieve are strands of interesting conversations that people are both invested in and listening to. This touches, in part, on the question of insularity raised below.
  • Seminar Talks: I would say that the median mathematics seminar could (perhaps harshly) be described as a waste of time for both the participants and the speaker. If we are to claim that fostering mathematical understanding is one of our main goals, we certainly haven’t made much of an effort to reward good talks. One institutional obstruction has always been the expectation that people talk about their own work. What often gets lost when one does this is an explanation of why the broader question being addressed is interesting in the first place, the methods that have been used most successfully in the field in the past, and the most promising questions to consider in the future. (One can do this in a talk about one’s own work, but people frequently do not.)
  • Insularity: The past few decades have seen an explosion in mathematics. But I feel this has come at the cost of mathematicians being less able to communicate; not only with people in other areas of mathematics, but sometimes even within their own field. Some have argued that this is an inevitable consequence of the growing difficulty of mathematics, but I suspect that once a mathematical subcommunity reaches a certain size, the impetus to reach out diminishes, to all our detriment.
  • Ego: Perhaps the thorniest question of all: how do we shape the incentives in our field to produce better outcomes? For all that we emphasize understanding, it would be insane not to acknowledge the importance of ego, and the way that the desire to be the first person to prove something has motivated many of us. This moment is going to require a great deal of humility. If human understanding of deep mathematics is what we want to defend, it ought also to be what we reward.

^*: now I have the following in my head:

Posted in Mathematics | Tagged , , , , , , , | 6 Comments

Look, Mom, I pressed a button! (Go edition)

Apropos of nothing, I was reminded recently of a fascinating story I heard from Geordie Williamson about AI and go — perhaps this story was the original source. It is a story about how a technology which has the power to greatly increase knowledge can rather create barriers to actually acquiring that knowledge. To give just some excerpts:

I started my career as a Go teacher in 2020 … I now estimate that about half our students had used AI in at least one game and one in ten were chronic users. We were originally baffled … It didn’t make sense that players would just throw away their practice games to have AI win on their behalf.

…

None of these reasons [for cheating] were surprising to us … What personally shocked me, however, was the way our students conceptualised their AI use. In this, Carlo Metta was also a surprisingly predictive case. The original reddit thread discussing his ban featured a comment from a user called “carlo_metta”, which read:

I never let Leela choose move. I just decide myself which one is better, for this reason i think i can find my own style with Leela. Go is an art and Leela help me tyo [sic] express my skill

That account was a burner, quite possibly a troll. However, I couldn’t help but recall the comment when I heard identical arguments coming from our cheating students’ accounts.

I think this story has a number of themes relevant to math, including how AI use by experts can make us worse mathematicians. I think when mathematicians use AI they need to be extremely careful not to fall into the trap of imagining that they (rather than some autonomous agent) are doing mathematics. Amateurs at least don’t have this problem!

I know good mathematicians who have used AI to prove interesting theorems, and they swear that the result started with their own original ideas which were then combined with the insights of AI. And I believe them! But it is very easy to start taking ownership of ideas that are not your own. I think mathematicians need to be very conscious that button pressing has the possibility of warping one’s perception of how much you actually contributed yourself.

For me, I think it helps to have at least one project where you simply do not use AI at all. And for projects that do use AI, make every effort to be honest to yourself about what your contribution is.

In my professional work that has appeared online so far, I have only used AI to “review” one paper after it was written. I have found this very helpful, and I certainly continue to do this. In my experience, AI reviews papers incredibly thoroughly. Ironically, what it seems most likely to miss are arguments that are so poorly written that it’s not even clear exactly what the argument is, but when you actually include some details it has an opportunity to find the hole in the argument.

That said, I definitely do plan to use AI for future research. I have used AI to try to better optimize the choice of the functions \(\psi\) we used in [CDT] (for example, look at Figure A.4.5). I was utterly confused during that paper how to optimize the (very non-linear) holonomy bounds as \(\psi\) ranged over all holomorphic maps \(\varphi: D(0,1) \rightarrow D(0,1-\varepsilon)\). Not only is this a complicated non-linear optimization problem, but there is also the issue of being able to actually compute the answer quickly and rigorously. After using ChatGPT 5.6, I remain utterly confused as to what the optimal choice looks like, but at least it *could* do better, and found a certain Ansatz of functions to try which improved the numerology. The improvements were not quite good enough (yet) to simplify our proof of the irrationality of \(L(2,\chi_{-3})\) in any meaningful way, but they were good enough to prove that the \(\mathbf{Q}(x)\)-vector space generated by functions \(f(x) \in \mathbf{Q}[[x]]\) on \(\mathbf{P}^1 \setminus \{0,1,\infty\}\) which have denominator type \([1,2,\ldots,n]^2\) has dimension at most \(8\) rather than dimension at most \(9\). (We still suspect the actual answer is \(5\).)

My “biggest” (in terms of tokens) use of AI so far is a somewhat quixotic attempt to construct a new finite simple group. Again, more on this later, but as the computation continues, it is my obligation to remain crystal clear about what exactly my contribution is. Note that although the headline goal of this project will certainly end in failure, there is the hope that interesting mathematics will nonetheless come out of this.

Posted in Mathematics | Tagged , , | 4 Comments

Look, Mom, I pressed a button!

I have recently heard a few extraordinary opinions about how mathematicians should use AI. One I would paraphrase as “professional mathematicians should never use AI for any problem not in their narrowly defined (by whom?) research program”, which seems ridiculous — one great advantage of AI is that it allows us to precisely broaden our own research (and mathematics more generally). But that is not what this post is about. Generally, my plan on this blog is neither to make predictions about the future nor to make any ethical pronouncements about its usage, but rather to understand its implications for mathematics as a discipline. (I always remind myself that my one prediction about AI was that computers would never beat humans at chess, and I’m not that old!)

The question is:

Question: What is the value of a result obtained by someone pressing a button in a context where they themselves contribute no insight?

There are a number of subtle things that may count as insight, including which problems to ask in the first place. It seems pretty clear, however, that suitably interpreted, the answer to this question is “no value at all”. If someone else was interested in the question, they could press the same button.

If a result can be obtained on demand by anyone merely by pressing a button, and the particular person who obtains it supplies no insight, then the production and announcement of that result have no mathematical value. In particular, I’m not only saying that the button-presser deserves no credit (which is obvious). I’m talking about what counts as a valuable mathematical result once answers themselves become freely reproducible commodities. In that setting, the first person to print the machine’s answer has not added anything to mathematics. The proposition may be true, and knowing it may have consequences, and the argument may be interesting, but this particular result — the act of generating and circulating the answer — is mathematically of marginal value.

There are many people right now burning through tokens asking LLMs about famous or not-so-famous conjectures. One reason is pure intellectual curiosity or a desire to explore the limits and capabilities of these models; the opportunity for everyone^* to have access to such powerful models is amazing. But if the motivation is some sort of personal glory for having been the first person to “prove” or to “know” some particular fact, then this seems misplaced, to put it politely. If you don’t do anything besides press a button and you don’t understand what comes out or whether it is correct, what is the point? Even assuming it is 100% correct (and Lean-certified!), it still takes an expert to determine if the proof contains anything original or interesting to mathematics as a discipline.

The theorems we prove often serve as a proxy for what is more important, namely, the creation of new methods and new ideas. I am not saying that results are not important. I would like to know that \(\zeta(5)\) is irrational as well as knowing why it is irrational. But that knowledge is less important than some people seem to think. The problem of whether \(\zeta(5)\) is irrational is an obvious enough question to ask that the first person who “presses a button” and gets a proof seems more or less irrelevant if they added no intellectual content of their own. What AI obviously changes is that novelty of theorem statement no longer reliably signals novelty, competence, effort, or understanding.

What I have said so far seems to be somewhere between tautological and self-evident, but I was reminded of it in the past few weeks by being forwarded not one but three proofs that \(\zeta_5(3) \in \mathbf{Q}_5\) is irrational. Let me give a quick and selective history of this type of problem (omitting the work of many people):

In 2005, I proved that \(\zeta_2(3)\) and \(\zeta_3(3)\) were irrational, as well as \(L_2(2,\chi_{-4})\), the \(2\)-adic Catalan’s constant. The insight here was to understand how Fritz Beukers’ modular version of Apéry’s proof had a \(p\)-adic analogue, where the “overconvergence” which in the complex case was coming from the functional equation and Eichler integrals — which saw the period \(\zeta(3)\) — was replaced by \(p\)-adic overconvergence of (non)-classical Eisenstein series, which see the period \(\zeta_p(3)\).

My collaboration with Dimitrov and Tang (around 2020) more or less started when Vesselin discovered a holonomy bound (following André) and noted that it could be used (by using the same overconvergent template I had used in my paper) to show that \(\zeta_2(5)\) was irrational, something that was not possible using my original method. Using our later, more refined bounds, we included a proof of this result in our ICM paper.

In 2025, Lai, Sprang, and Zudilin independently proved that \(\zeta_2(5)\) was irrational. Their proof used a more direct Apéry-like construction.

So what, then, are my thoughts on these proofs that \(\zeta_5(3)\) is irrational? First, among the (proper subset of) proofs that are probably correct, they contain essentially no original ideas whatsoever — they are simply applying the best holonomy bounds from [CDT] to the template constructed in [C2005]. Who knows how many other people have pressed the same button to prove the same result! If these had been written up well by a graduate student, then to me their value would be “this graduate student has understood how to apply [CDT] correctly”, and such a paper could plausibly appear in a journal. If we want this to continue, we have to be very clear and conscious of what the other added value is beyond the result itself.

What was interesting about the irrationality of \(\zeta_2(5)\) was not only the result, per se, but the completely new method used to obtain it. At the same time, the proof by Lai-Sprang-Zudilin of the irrationality of \(\zeta_2(5)\) [A known result!] is much more interesting than these proofs that \(\zeta_5(3)\) is irrational [A new result!], because in the former case the argument required the construction of a new series of approximations related to higher-dimensional families of Calabi-Yaus rather than families of elliptic curves, and in the latter case, you just take the currently available arguments and apply them in the obvious way to the obvious constructions.

The proofs vary both in quality and in the extent to which the respective “authors” made the effort to ensure that the argument was correct. Two of the proofs seem more or less plausible, more or less the same, and more or less obvious. But I could hardly recommend that anyone spend time reading them, let alone reviewing either paper for a journal. They have no value. The third proof, however, is more amusing. It claims to prove that \(\zeta_2(2k+1)\) is irrational for a set of positive integers \(k\) of density one. That would be a more substantial result. The proof even passes, with caveats, an initial examination by ChatGPT 5.6. So now one feels compelled to make at least some effort to consider what is going on. What one quickly realizes is that the same argument would apply not only to the constant term of the \(2\)-adic Eisenstein series \(E_{-2k}(q)\), but would also “show” that (as \(k\) varies), for a positive proportion of positive integers \(k\), the constant terms of the Eisenstein series \(E_{-2k}(q) – E_{-2k}(q^2)\) are also irrational. That last result, if true, would indeed be impressive.

The last example is also interesting to consider on several levels: it’s bad for OpenAI because the cost of that computation is more than what they charged for it (edit according to the comments this is probably wrong!); it’s bad for the amateur who produced that proof because they are throwing away money to produce slop and also (potentially) suffering the embarrassment of proudly posting slop (though I have never found amateurs to worry much about that); it’s bad for me because I felt compelled to waste some time thinking about it; and it’s bad for anyone else who looks at the argument (whether they know anything about mathematics or not) because, well, it is AI slop. So this is a situation where everybody loses! I think that is far from a unique case right now.

I don’t think the implications of what I have said for amateurs are that interesting (though with ChatGPT 6 just released, the volume of button pressing is only going to increase.) What I think is more interesting is the implication for mathematicians and the results that they prove, whether they are using AI or not. But I shall return to this in a later post.

^* everyone who can afford it.

Posted in AI, Mathematics | Tagged , , , , , , , | 11 Comments

OpenAI, updated

An update on this post, from my inside sources:

word on the math streets of SF is that the OpenAI team tried something like 500 problems to get their 10 solutions.

I don’t know how much of an insider this source is (or this sources sources, etc), but (allowing for the possibility of confirmation bias) this is within the expected range.

Posted in Mathematics | Tagged , , | 6 Comments

Google Alert!

I have my google alert set for the phrase “Galois Representations”. Every six months or so it pops up with a suggestion, and I can’t quite work out what algorithm is using. Here was today’s breaking news: On the conductors of mod \(\ell\) Galois representations coming from modular forms.

Posted in Uncategorized | Tagged , , | Leave a comment

The inverse Galois challenge, part II

This is a sequence to this post. The SAIR competition (Round I) has been completed! 98.4% of the possible signatures were obtained, with only 39 non-solvable cases missing.

Some thoughts.

First, my timing in the last post of dissing the problem of realizing \(M_{23}\) as a Galois group was not so great. It seems to me that the delightful paper does an excellent job of combining human and AI thoughts but also clearly and concisely explaining the ideas, especially distinguishing between what is known, what is clever, and what is lucky. Nicely done!

Moving on to the competition. I thought that it would be better to get a precise sense of the difficulty by trying it myself. The approach I used was purely to tell CODEX to do 6 obvious things, but not to either look at any literature myself, not to write any code, and just to come back and complain when it failed. This quickly produced around 40,000 pairs, but then stalled. One approach that wasn’t successful at all was as follows. There were around 80,000 pairs or so could be realized as coming from the Galois closure of degree 12 extensions of quadratic fields. But alas, my suggestions for how to construct these were not taken up sensibly, and I didn’t pursue it.

Certainly my personal explorations produced no meaning mathematical content at all. The only mathematical idea I had that was not completely obvious was one I learnt entirely from David Roberts. In situations where one has a Galois extension \(L/K/\mathbf{Q}\) where \(K\) has Galois group \(G\) and \(L\) has Galois group a central extension of \(G\) of degree \(2\), then one can write \(L\) as the splitting field of a polynomial of the form \(f(x^2)\) where \(f(x)\) has splitting field \(K\) and one root of \(f(x)\) generates a field \(E\). But now, given \(g(x)\) with \(E \simeq \mathbf{Q}[x]/g(x)\), how does one find \(f(x)\)? The observation is that one can often find \(f(x)\) by applying \(\texttt{polred}\) to \(g(x)\).

That said, having done some of these experiments, it did help me appreciate what type of problem this was. It certainly seemed to be the case that real skill and knowledge working with explicit polynomials and explicit Galois theory would be genuinely useful, and simply a purely theoretical knowledge of (say) the general solvable case is not sufficient. It is no surprise then that Klüners and Malle (the leading team) were so successful.

But where does it lead us? I don’t think the conclusion is so far from my original prediction. I think there might be a new second round coming, and after that is done, it really could be the case that the only pairs remaining are \((G,r)\) where \(G = \mathrm{PSL}_2(\mathbf{F}_{23})\) and also \(G = \mathrm{PGL}_2(\mathbf{F}_{23})\) with \(r=0\) (which are hard for the same reasons), and then possibly some cases of \(M_{24}\) (say with \(r=0\)) as well. We shall see!

Posted in Mathematics | Tagged , , , , , , , , , , , , | 3 Comments