We are all aware of what ChatGPT can do. I think we would be better informed if we also learn what it cannot do (at least right now!). In order to better understand the current capability of the latest publicly available AI models, I decided to curate a set of problems and test how well ChatGPT (running 5.6 sol ultra) could do. While this certainly is not a scientific experiment, I had some rules I set for myself in advance:
- The problems should be, as far as I know, generally be open problems.
- I should have at least some original thought or idea on how to approach the problem which I can suggest to the model, however stupid it might be.
- I will limit myself to one ChatGPT pro subscription and the time between the ICM and the Emerton-Kisin conference (a bit under two weeks) to address all of these questions.
The strategy I employed was as follows. I gave ChatGPT one to two hours on each problem (ultra think in the chat window). If it made no progress at all, then I didn’t pursue the problem any further. If it did make progress, then I used codex on goal mode to push towards the problem, or at least towards some interesting intermediate goal, with a maximum run of two days.
The goal of this project is not to test the limits of what can be done by these models, but a much more practical test of how it might be to use these models as a working mathematician. If you take this experiment, scale up the amount of compute, and the amount of mathematicians giving (limited) direction to the machines, the result is (to my mind) broadly consistent with what OpenAI achieved, of course assuming that there were a significant number of problems on which they made no progress and then abandoned.
This quarter at Chicago I will be running a “ChatGPT seminar” precisely to explore these questions. The scope of that seminar will be somewhat broader than pure problem solving, and also include typesetting, coming up with interesting conjectures, and many other things. That said, it will certainly involve tests such as these.
So how well did it do? Let us see.
- Compute the slopes of all finite slope overconvergent modular forms with \(p=2\), level \(N=1\), and integral weight \(k\).
- Prove that there exists a constant \(N\) such that, if \(G\) is a finite group with \(H^i(G,\mathbf{Z})=0\) for \(i=1,2,\ldots,N\), then \(G\) is trivial.
- Determine the slopes of all periodic billiard paths in the regular heptagon.
An equivalent formulation is to take the \((2,7,\infty)\) triangle group and ask for a classification of its cusps in \(\mathbf{P}^1(K)\) where \(K = \mathbf{Q}(\zeta_7)^{+}\). This is a thin group inside \(\mathrm{SL}_2(\mathcal{O}_K)\). I heard about this problem from Curt McMullen, who also gave a possible answer (who he attributed to someone else, but since this was just a conversation I apologize that I did not remember at the time).
Level of interest: I think if you answered this question then Curt would be impressed. What more could you ask for?
Level of difficulty: One reason I considered this problem is that I had a sense that the answer should involve some mix of algebraic number theory and or Arakelov theory, and at the same time some complex analysis in the form of Hodge Theory. This could exactly be the type of situation where there might be a simple answer just by combining ideas from different fields.
Result: ChatGPT had sloppy thoughts on this one! My first reading is that it did not have any crucial insights beyond fleshing out a little what I had suggested. It certainly diligently tried to push things as far as it could, but I think it is still missing a (or the) key idea. Time spent: around 48 hours.
- Let \(M\) and \(N\) be two finite volume hyperbolic \(3\)-manifolds with isomorphic pro-finite completions. Prove that \(M \simeq N\).
- Let \(\Delta = \langle x,y | x^p, y^q, (xy)^r \rangle \) be a hyperbolic triangle group, and let \(B/K\) be the associated quaternion algebra over the invariant trace field. Let \(g(p,q,r)\) be the density of real places such that \(B\) is non-split. Prove that either \(g(p,q,r)=0\) or \(g(p,q,r) \ge 1/12\).
- All the problems listed in my current NSF proposal draft.
- Construct a new finite sporadic simple group not in the current classification.
This is a special case of the Ghost Conjecture of Bergdall and Pollack, which has now been solved by Liu-Truong-Xiao-Zhao (see also this post. Note, however, that that proof excludes this particular case when \(p=2\). I specifically pointed the model towards Conjecture 2 of this paper. In this very special case, the problem reduces to computing the Newton Polygon of a very explicit matrix where one can take \(k\) to be a non-negative integer. Our paper answer the case when \(k=0\). We also worked out the case \(k=-12\) and \(k=-72\) (the latter in part for proof of concept of the approach we were using).
Level of interest: Kevin and I certainly spent some time trying to prove it! This special case probably now mostly of historical interest in light of more recent approaches.
Level of difficulty: I would not be surprised if it could be solved by some elementary arguments.
Result: No progress in the initial time period; not pursued. I was a little surprised, but with the time constraints this did not seem worth devoting extra time to this question. Time spent: about 90 minutes. I’m going to get on a plane in a few hours, I’m going to give it another go for 6 hours or so this evening, then update tomorrow if anything changes.
This perhaps the one problem I felt I had the least insight. I think I learnt it from a mathoverflow question in the long past (yes, I looked it up and found it here).
Level of interest: Hard for me to say. One imagines this problem should be more or less a computation plus a literature search for the case of finite simple groups (assuming CFSG), and then it becomes some inductive problem which may or may not be about facts concerning the cohomology of almost simple groups. But this is not my area.
Level of difficulty: I have no idea.
Result: Partial progress. It knows enough to answer the case of finite simple groups, which is the first step in the obvious induction argument. It does cover quite a few non-trivial cases, but then gets bogged down, and comes up against what it calls difficult problems. Time spent: about 48 hours.
I first learnt about this problem from Martin Bridson and Alan Reid in Ventotene in 2015. At the time, I had some idea about approaching this via the representation variety, but it was sufficiently far from things I knew that I didn’t pursue it.
Level of interest: Definitely there are people interested in this problem.
Level of difficulty: Too difficult for me to say, but it’s a well–known problem,
and (as I learnt during this process) significant progress has been made over the past few years.
Result: Claimed Solution. This perhaps might be the most interesting positive case. ChatGPT informed me of a recent paper of Liu in which he proved that, for closed hyperbolic \(3\)-manifolds, the profinite completion determined the volume modulo a conjecture about the injectivity of a certain regulator map. I suggested that one could bypass this using the mod-\(p\) Chern class maps discussed in Calegari-Garoufalidis-Zagier. With that, ChatGPT was very quickly able to write a \(5\)-page paper using Liu’s result giving an unconditional proof (in this class of manifolds) that volume was determined by the pro-finite completion. I think that this could have lead to a nice short note that I could reasonably post under my name with suitable AI assistance disclaimers. But then I learnt from Alan Reid and Martin Bridson (who I sent a draft to) that the full result had recently been proved by Xu! At this point, I “pressed another button” and asked ChatGPT to prove the full result, which it did. In particular, the notes below were produced completely independently from the work of Xu, now available here, but they were produced with knowledge that such a paper existed. This surely (?) would have influenced the strategy that ChatGPT decided to pursue. In fact, while there are similarities in the argument they are certainly not the same; After Xu’s paper was posted on the arXiV, I asked ChatGPT to compare the proofs, and it came up with the following: When it was first done, I asked ChatGPT to referee and revise its own work back and forth in until it claimed it was ready to be submitted. It modestly suggested that it should be submitted to the Annals of Mathematics. Now it is not the main point of this post, but obviously the question of how we evaluate work going forward is going to be an extremely important one. While the paper posted above does contain at least one idea of mine, I certainly do not intend to publish it, nor am I willing to take responsibility for its contents. Time spent: about 12 hours.
Level of interest: This is a question of Curt McMullen raised in this paper.
Level of difficulty: Note that in this paper here we prove that \(g(p,q,r)=0\) for precisely \(14\) explicit hyperbolic triangle groups, also answering a conjecture of Curt from that same paper. It was definitely clear to me during the writing of this paper that it could certainly be possible to prove this result. I actually started a project with University of Chicago undergrads towards it, but none of the people who signed up seemed actually willing to do any work so it petered out. The one difficulty that was certainly possible was that some eventual argument might be effective, but not effectively effective. Two improvements were needed from the previous paper; the first was to optimize the Fourier analysis aspect. The second one was to replace the Jacobsthal function argument which produced a single interesting conjugate to something more flexible that could produce a positive density of interesting conjugates.
Result: Solved. Here ChatGPT did a number of things I expected, which was to choose a much more elaborate test function in the Fourier argument than we used, since it would obviously be much better handling much more complicated expressions. This was the first problem I asked, and for some time I actually was going to get ChatGPT to formalize the proof in Lean, which it felt completely capable of doing. But the time frame was going to be several weeks, and I didn’t want to waste the tokens. But this might possibly be worth doing. Time spent: about 6 hours.
They are all ChatGPT hard, right now!
Level of interest: A lot. This sounds like a trolling question, but I do actually have one not entirely stupid idea, which should hopefully at least produce some interesting mathematics.
Level of difficulty: Probably quite hard.
Result: We are now 10 days into various computations that are making progress on something. But it didn’t come under budget, so I will talk about it later instead.