zeta(5) is irrational

So one day after an amazing result by humans, the computers have struck back! This time, a proof that \(\zeta(5)\) is irrational. The paper I saw gives me the impression of being almost entirely AI generated (if so, it then comes with a very dishonest disclosure), but never mind, thanks for the compute!

I spent some effort trying to see what was going on. Since the argument has now been verified in Lean, it seemed more useful to try to identify the key ideas and where they could be traced, and where they were new.

One way I tried to understand this was to insist that ChatGPT try to find in the literature where similar ideas had been used before. I am sure that I did not do a great job with the literature, but I did my best in the several hours I had available. I sure hope Jacob Tsimerman’s suggestion that AI might reach superhuman exposition will come soon. The discussion below connects the arguments to the work of Prévost (not cited in the paper), but there may well be other sources that ChatGPT learnt these ideas from. Certainly Hankel determinants are also considered by Zudilin, and I suspect that he would do a better job unearthing the ideas behind this paper (some of which are no doubt in his own papers) than I have done. Later on I quote from ChatGPT directly on what it thinks.

Put \(D_0(t)=1\) and \(D_m(t)=\prod_{j=1}^m(t+j^2)\) for \(m\geq1\). Fix an integer \(k\geq2\), and let \(H_n^{(k)} =\sum_{v=1}^n v^{-k}\) with \(H_0^{(k)}=0\), and define the Bernoulli numbers as usual by \(u/(e^u-1)=\sum_{m\geq0}B_mu^m/m!\), so \(B_1=-1/2\).

Definition. Let \(t\) and \(X\) be independent indeterminates. Define the \(\mathbf{Q}\)-vector space

\[\begin{gathered} \mathcal R=\bigcup_{m\geq0}\left\{\frac{P(t)}{D_m(t)}:P(t)\in\mathbf{Q}[t]\right\}\subset\mathbf{Q}(t),\\ \mathbf{Q}[X]_{\leq1}=\{aX+b:a,b\in\mathbf{Q}\}. \end{gathered}\]

Thus \(\mathcal R\) consists of the rational functions of \(t\) whose finite poles, if any, are simple and belong to \(\{-1^2,-2^2,\ldots\}\). There is no restriction on the polynomial part. For the fixed integer \(k\geq2\), define the \(\mathbf{Q}\)-linear map \(\mu_{k,X}:\mathcal R\longrightarrow\mathbf{Q}[X]_{\leq1}\) by

\[\begin{aligned}\mu_{k,X}(t^e)&=(-1)^eB_{2e+2}\frac{(2e+k)!}{(k-1)!(2e+2)!}&& (e\geq0),\\ \mu_{k,X}\left(\frac1{t+j^2}\right)&=j^{k-1}(X-H_j^{(k)})-\frac1{k-1}+\frac1{2j}&& (j\geq1).\end{aligned}\]

Write \(\mu_k:\mathbf{Q}[t]\to\mathbf{Q}\) for its restriction to polynomials, and abbreviate \(\mu_{k,X}((t+j^2)^{-1})\) to \(\phi_{k,j}(X)\). The fact that this is well-defined follows from the partial fraction expansion. If one writes \(R\in\mathcal R\) as \(R(t)=P(t)+\sum_{j=1}^m c_j/(t+j^2)\), with \(P(t)=\sum_e p_et^e\in\mathbf{Q}[t]\) and \(c_j\in\mathbf{Q}\), then

\[\begin{aligned} \mu_{k,X}(R)={}&\left(\sum_{j=1}^m c_jj^{k-1}\right)X+\sum_e p_e\mu_k(t^e)\\ &+\sum_{j=1}^m c_j\left(-j^{k-1}H_j^{(k)}-\frac1{k-1}+\frac1{2j}\right). \end{aligned}\]

For a real number \(\xi\), let \(\operatorname{ev}_\xi:\mathbf{Q}[X]\to\mathbf{R}\) be evaluation at \(X=\xi\), and define \(\mu_{k,\xi}=\operatorname{ev}_\xi\circ\mu_{k,X}\). Thus \(\mu_{k,\xi}(R)\) is a real number. For any field \(F\) containing \(\mathbf{Q}\), the same basis and formulas define an \(F\)-linear map on \(\mathcal R_F=\bigcup_m D_m(t)^{-1}F[t]\), with values in \(F[X]_{\leq1}\). We keep the notation \(\mu_{k,X}\) for this extension and \(\mu_k\) for its restriction to \(F[t]\). When \(F=\mathbf{R}\), we can specialize \(X=\xi\in\mathbf{R}\) as above. In particular, this specifies the map on rational functions with real coefficients.

The formulas of Euler and Hermite

Let \(f(y)=(e^{2\pi y}-1)^{-1}\) for \(y>0\), and put

\[\begin{aligned} w_k(y)&=\frac{2(-1)^{k-1}y^k}{(k-1)!}f^{(k-1)}(y)\\ &=\frac{2(2\pi)^{k-1}y^k}{(k-1)!}\sum_{l\geq1}l^{k-1}e^{-2\pi l y}. \end{aligned}\]

The density \(w_k(y)\) is positive for \(y>0\), tends to \(1/\pi\) as \(y\to0\), and satisfies \(w_k(y)=O_k(y^ke^{-2\pi y})\) at infinity. Euler’s and Hermite’s formulas give the following integral representation.

Proposition. For \(R\in\mathcal R_{\mathbf{R}}\),

\[\mu_{k,\zeta(k)}(R)=\int_0^\infty R(y^2)w_k(y)\mathrm{d} y.\]

For \(a>0\), write \(\zeta(k,a)=\sum_{v\geq0}(v+a)^{-k}\). The corresponding transform identity is

\[\int_0^\infty\frac{w_k(y)}{y^2+a^2}\mathrm{d} y=a^{k-1}\zeta(k,a)-\frac1{2a}-\frac1{k-1}.\]

The transform identity is the integer-weight specialization of Prévost’s Stieltjes representation [3] (Theorem 1).

Hankel Determinants

Choose integers \(0\leq N<K\), \(r\geq1\), and \(h\geq1\). Let \(\mathcal P_h=\mathbf{Q}[t]_{<h}\) be the \(h\)-dimensional vector space of polynomials of degree less than \(h\). Multiplication by \(W(t)=D_N(t)^r/D_K(t)\) sends every polynomial to \(\mathcal R\). Therefore

\[\beta_{k,X}:\mathcal P_h\times\mathcal P_h\longrightarrow\mathbf{Q}[X]_{\leq1},\qquad \beta_{k,X}(p,q)=\mu_{k,X}(Wpq)\]

is a well-defined symmetric bilinear map. Its matrix in the basis \(1,t,\ldots,t^{h-1}\) and its determinant are

\[\begin{gathered} G_k(X)=\left[\mu_{k,X}\left(\frac{D_N(t)^rt^{i+j}}{D_K(t)}\right)\right]_{0\leq i,j<h},\\ \Delta_k(X)=\det G_k(X)\in\mathbf{Q}[X]. \end{gathered}\]

Such a matrix is called Hankel because its entries depend on \(i+j\). (The notation suppresses \(K,N,r,h\).) Since the entries have degree at most one in \(X\), the determinant has degree at most \(h\).

Lemma. After scalar extension to \(\mathbf{R}\) and specialization \(X=\zeta(k)\), the bilinear map \(\beta_{k,\zeta(k)}\) is an inner product. For every nonzero real polynomial \(q\) of degree less than \(h\),

\[\beta_{k,\zeta(k)}(q,q)=\int_0^\infty W(y^2)q(y^2)^2w_k(y)\mathrm{d} y>0.\]

Thus \(G_k(\zeta(k))\) is a positive-definite Gram matrix, meaning the matrix of pairwise inner products of a linearly independent list. In particular \(\Delta_k(\zeta(k))>0\).

Allowing for a general rational \(W\), Prévost’s arguments for \(\zeta(2)\) and \(\zeta(3)\) fit this same Hankel construction, using the one-pole choices \(W_{2,n}(t)=1/(t+(2n+1)^2)\) and \(W_{3,n}(t)=t/(t+(n+1)^2)\), respectively, with matrices of size \(n+1\) [3]. These determinants are affine polynomials in \(X\), and their roots are the corresponding rational approximations to \(\zeta(k)\): a change of polynomial basis expresses each determinant as a rational factor times a Padé approximation error. Prévost’s coefficient and error estimates allow these determinants to be rescaled into integer linear forms whose nonzero values at \(\zeta(2)\) or \(\zeta(3)\) tend to zero, proving irrationality. On the other hand, in this new argument, the single pole of \(W\) is replaced by the \(K-N\) poles of \(D_N(t)^r/D_K(t)\), producing a determinant of degree \(K-N\) in \(X\).

Proposition (Degree of the determinant). Take \(h=K-N\). Then \(\Delta_k\) has degree \(h\), with

\[[X^h]\Delta_k(X)=(-1)^{h(h-1)/2}\prod_{j=N+1}^Kj^{k-1}D_N(-j^2)^{r-1}\ne0.\]

For a nonzero polynomial \(\Delta\in\mathbf{Q}[X]\), a positive rational number \(c\) with \(c\Delta\in\mathbf{Z}[X]\) will be called a normalizing factor. For example, one can multiply by the least common multiple of the coefficient denominators. One can also divide out a common factor of the resulting integer coefficients; such division is important in this construction. The required estimate concerns the value of the normalized polynomial, including both effects.

Lemma. Let \(\xi\in\mathbf{R}\). Suppose \(P_n\in\mathbf{Z}[X]\), \(\deg P_n\leq Cn\), and \(0<P_n(\xi)\leq\exp(-cn^2+o(n^2))\), where \(C,c>0\) are fixed. Then \(\xi\) is irrational.

The strategy

The goal is to apply the integer-polynomial lemma with \(\xi=\zeta(k)\) and with \(P_n\) a positive rational multiple of the determinant \(\Delta_k(X)\) in the Hankel matrix. “All” that has been done is to adjust Prévost’s positive measure by \(D_N^r/D_K\) for suitable parameters \(N\) and \(K\). Increasing the numerator power actually makes the determinant larger, but it turns out that one obtains extra cancellations in the coefficient denominators that compensate. Establishing those cancellations uniformly is required for the proof, of course, but finding the right construction is the key step. Once the required arithmetic estimates are proven, one deduces the irrationality of \(\zeta(k)\) with \(k=2,3,4,5\), but not (at least directly) for larger \(k\).

Prévost’s moments and the introduction of the zeta value

We now attempt some mathematical archaeology. Prévost’s Theorem 1 in [3] gives the Hurwitz-zeta representation above. Write \(\mathrm{d}\nu_k(t)=w_k(\sqrt t)\,\mathrm{d}t/(2\sqrt t)\) for \(t>0\). Its polynomial moments are

\[m_j:=\int_0^\infty t^j\mathrm{d}\nu_k(t)=\mu_k(t^j)=(-1)^jB_{2j+2}\frac{(2j+k)!}{(k-1)!(2j+2)!}\in\mathbf{Q}.\]

These are rational for every integer \(k\geq2\). Their Hankel determinants compute rational orthogonal polynomials. Put \(H_s=\det[m_{i+j}]_{0\leq i,j<s}\), with \(H_0=1\). The determinant formula in Prévost’s equation (2.11), written for this moment sequence, is

\[p_s(z)=\frac{1}{H_s} \det\begin{pmatrix} m_0&m_1&\cdots&m_s\\ m_1&m_2&\cdots&m_{s+1}\\ \vdots&\vdots&&\vdots\\ m_{s-1}&m_s&\cdots&m_{2s-1}\\ 1&z&\cdots&z^s \end{pmatrix}.\]

The polynomial is monic of degree \(s\) and satisfies \(\int t^j p_s(t)\mathrm{d}\nu_k(t)=0\) for \(j<s\); these are its orthogonality conditions. Both \(H_s\) and the coefficients of \(p_s\) are rational. The zeta value enters through a subsequent evaluation of a transform, as follows.

Define \(\mathcal F_k(z)=\int_0^\infty(z+t)^{-1}\mathrm{d}\nu_k(t)\) for \(z>0\). At a positive integer \(a\), we get

\[\mathcal F_k(a^2)=a^{k-1}\bigl(\zeta(k)-H_a^{(k)}\bigr)-\frac1{k-1}+\frac1{2a}=\phi_{k,a}(\zeta(k)).\]

Consequently a rational approximation to \(\mathcal F_k(a^2)\) gives a rational approximation to \(\zeta(k)\). A Padé approximation matches finitely many coefficients of the asymptotic expansion \(\sum_{j\geq0}(-1)^jm_jz^{-j-1}\), as \(z\to+\infty\), by a rational function. The coefficients of that rational function are determined by the moment equations above. Prévost–Rivoal give effective approximation bounds when the Padé degree and \(a\) increase together [4].

Replacing \(\mathrm{d}\nu_k(t)\) by \(\mathrm{d}\nu_k(t)/(t+a^2)\), polynomial division gives the affine moment

\[s_{a,j}(X):=\mu_{k,X}\left(\frac{t^j}{t+a^2}\right) =\sum_{l=0}^{j-1}(-a^2)^l m_{j-1-l}+(-a^2)^j\phi_{k,a}(X),\]

where the sum is empty at \(j=0\). At \(X=\zeta(k)\) this is exactly the \(j\)th moment of the corresponding measure.

The measure in the new argument is the following explicit modification of the same \(\nu_k\):

\[\mathrm{d}\sigma_{k,K,N,r}(t)=W(t)\mathrm{d}\nu_k(t),\qquad W(t)=\frac{D_N(t)^r}{D_K(t)} =\frac{\displaystyle\prod_{a=1}^N(t+a^2)^{r-1}} {\displaystyle\prod_{j=N+1}^K(t+j^2)}.\]

Here \(0\leq N<K\) and \(r\geq1\) are integers. All factors are positive on the integration interval \(t\geq0\). For fixed \(K,N,r\) the moments are \(\mu_{k,\zeta(k)}(Wt^j)\), and their Hankel determinant of size \(h=K-N\) is \(\Delta_k(\zeta(k))\). These are the moments used throughout the argument.

ChatGPT tells me: Uvarov studies the effect of precisely this type of rational multiplication on orthogonal polynomials [7]. Krattenthaler’s Theorem 1 and Proposition 13 express the modified Hankel determinant through the original \(p_s\) and their Cauchy transforms \(C_s(y)=\int p_s(t)/(y-t)\mathrm{d}\nu_k(t)\), for \(y<0\) [8]. For the displayed modifier, each numerator node \(-a^2\) gives \(r-1\) rows \(p_s^{(d)}(-a^2)/d!\) for \(0\leq d<r-1\), and each denominator node gives a Cauchy-transform row. Thus the measure considered here belongs to an established family of rational modifications. The starting measure and remainder approximation come from Prévost and Prévost–Rivoal [2], [3], [4]. The exact algebra for multiplying that measure by \(D_N^r/D_K\) is supplied by Uvarov’s rational modifications and Krattenthaler’s determinant identity [7], [8]. Zudilin then supplies the irrationality criterion for a positive moment determinant with sufficiently small coefficient denominators [5]. These results identify the starting approximation problem, the modified determinant, and the final integer-polynomial argument, respectively. The closest precedent for improving the coefficient normalization of the matrix is Brown’s Sections 7.4–7.5 [6]. There one first clears row denominators, then removes column common factors and uses congruences between rows or columns to reduce the total multiplier. Here the concrete mechanism uses congruences between square nodes, while also accounting for the polynomial parts in the partial-fraction expansions. Applying \(\mu_k\) to those polynomial parts can introduce Bernoulli denominators. The required quantitative input is that the assigned divisibility holds simultaneously for every mixed pairing, with these contributions included.

To come back to the argument, one can (uniformly, for \(k=2,3,4,5\)) take \(K=12n\), \(N=n\), \(r=6\), and \(h=11n\), and write \(\Delta_{k,n}\) for the corresponding determinant. (This is different from what is done in the Lean verification, but whatever.) Set

\[S_n=\frac{(K!)^{2h}4^{h-1}}{(N!)^{12h}\prod_{i=1}^{h-1}((2i)!)^2},\qquad F_{k,n}(X)=S_n\Delta_{k,n}(X).\]

One then has to estimate this at the real place and then control the denominators. In particular:

Theorem: \(\zeta(k)\) is irrational for \(k \in \{2,3,4,5\}\). For all sufficiently large \(n\),

\[\begin{aligned}P_{k,n} & \in\mathbf{Z}[X], \\
\deg P_{k,n} & =11n \\ P_{k,n}(\zeta(k))& \le e^{-\gamma_k n^2} \\ (\gamma_2,\gamma_3,\gamma_4,\gamma_5)& =(30,96,6,3/2).\end{aligned}\]

References

[1] F. Beukers, A note on the irrationality of \(\zeta(2)\) and \(\zeta(3)\), Bull. London Math. Soc. 11 (1979), 268–272. doi:10.1112/blms/11.3.268.

[2] M. Prévost, A new proof of the irrationality of \(\zeta(2)\) and \(\zeta(3)\) using Padé approximants, J. Comput. Appl. Math. 67 (1996), 219–235. doi:10.1016/0377-0427(95)00019-4.

[3] M. Prévost, Remainder Padé approximants for the Hurwitz zeta function, Results Math. 74 (2019), article 51; arXiv:1709.05389v1.

[4] M. Prévost and T. Rivoal, Diagonal convergence of the remainder Padé approximants for the Hurwitz zeta function, J. Number Theory 222 (2021), 346–361. doi:10.1016/j.jnt.2020.10.019.

[5] W. Zudilin, A determinantal approach to irrationality, Constr. Approx. 45 (2017), 301–310. arXiv:1507.05697v2; doi:10.1007/s00365-016-9333-7.

[6] F. Brown, Mellin transforms, transfinite diameter and rational approximations of integrals, 2026, arXiv:2604.20741v1.

[7] V. B. Uvarov, The connection between systems of polynomials that are orthogonal with respect to different distribution functions, U.S.S.R. Comput. Math. Math. Phys. 9, no. 6 (1969), 25–36. doi:10.1016/0041-5553(69)90124-4.

[8] C. Krattenthaler, A determinant identity for moments of orthogonal polynomials that implies Uvarov’s formula for the orthogonal polynomials of rationally related densities, 2021. arXiv:2103.03969v1.

This entry was posted in Mathematics and tagged , , , , , . Bookmark the permalink.

Leave a Reply

Your email address will not be published. Required fields are marked *