how to read an ai mathematics announcement
In October 2025 an OpenAI vice president posted that GPT-5 had "found solutions to 10 previously unsolved Erdős problems". What had happened was literature search: the model had located existing papers, under different vocabulary, for problems the database maintainer was simply not aware had been solved. Real capability, wrong label, deleted post, and a year of hardened priors. In August 2026 Anthropic announced that an unreleased Claude had raised the proportion of zeta zeros provably on the critical line from 41.67% to 67.25%. Two mathematicians as authors of record, the holders of the previous records as reviewers, a Lean formalization, session transcripts, and three weeks later an independent human reproof. Maynard called the company "remarkably restrained in avoiding overhyping".
Those two are the calibration points. Between them, a checklist.
What did the machine output? An object, a pattern, or an argument. Objects are checked in seconds and the interesting question is the classical baseline. Patterns become theorems only when a human proves one. Arguments need a referee, and the announcement should say who.
Was the problem open, or just unattended? "Open" on erdosproblems.com means Bloom is not aware of a solution. Many problems there were never seriously attacked. Tao's line on problem #728, "the win says more about speed than difficulty", applies widely. Erdős problem #4 on prime gaps, where the 2018 argument is not a soft target, is a different kind of claim from a 1979 "investigate" question nobody had looked at.
Is it Lean-verified, and of what? A zero-sorry build proves the formal statement. It says nothing about whether the formal statement is the theorem as the field understands it, or whether the argument was in the literature already. Astra's ten certificates were trustlessly checkable and two of them still drew named prior-art complaints. Palomar's second check, an informal description compared against the formal one, exists precisely for this.
Is the model released? Most 2026 headline results came from unreleased internal models. Nobody outside the lab can reproduce the discovery, only the verification. That is fine for correctness and bad for science.
Who is speaking? The unit distance disproof was announced by nine external mathematicians writing a digestion paper. The IMO 2026 perfect scores split into two tiers, two systems graded by the IMO and four graded by Claude-based agents, and most coverage merged them. The ICPC and IMO 2025 results were announced by the labs, one of them before official grading concluded.
How much digestion did it need? The Sendov proof arrived as 90,000 lines of Lean and left Tao's hands as 15,000, after several days of his time. The zeta argument was "technically intricate" with a mechanism "not immediately transparent" until Lamzouri found the Hilbert-space version. Digestion debt is real, it is paid by the scarcest people in the field, and it does not scale.
What does the tier breakdown say? Epoch's open-problems set is the most rigorous attribution standard in use: the core ideas must be unambiguously the machine's, and a result with an unclear human share is marked human. By its count, six of fifty problems fell to AI by August 2026, none of them in the top two significance tiers.
Is there a claim at all? Four days after the checklist above was written, a rumour that Claude had solved Navier–Stokes went round the world on the strength of one user's prediction and a hypothetical Tao post. No paper, no repo, no lab statement. The real work it attached itself to, Alpöge and Buckmaster's smooth-forcing blow-up for Euler, Boussinesq and porous media, was a genuine and Lean-checked step that its own assessor described as a direction with enormous remaining difficulties. Two verified results in a month buy a third claim credibility it has not earned.
None of this is a reason for scepticism about the aggregate. Nobody at the ICM panel disputed that AI now produces new mathematics, and the unit distance technique seeded human breakthroughs on the sum-product problem within weeks. It is a reason to read each claim as the specific thing it is.