The Verification Gap Behind the Headline Number

On October 6, 2026, OpenAI published 722 mathematical manuscripts generated by an unreleased internal frontier model, organized into 372 result families and posted to a public GitHub repository under an Apache-2.0 license [1]. The model had been posed roughly 4,000 open problems, with each published result representing on average about three hours of ChatGPT Pro-level thinking compute [1]. Among the headline claims: a partial 'quasi-Riemann hypothesis' result, a Lean-verified proof connected to the Unique Games Conjecture, and a resolution of the Hodge conjecture for CM abelian varieties [2].
But the number that matters more than 722 is 42 percent - the share of the released proofs that had actually undergone Lean formal verification, the automated proof-checking standard the field increasingly treats as a baseline for trust [3]. Just 10 of the resulting 719 manuscripts included released chain-of-thought summaries showing how the model reached its answers [3]. That gap turned out not to be theoretical: one day after release, OpenAI withdrew three manuscripts tied to the Hodge conjecture after discovering a sign error that invalidated an argument the other two withdrawn papers depended on, and separately revised 14 other papers and updated 13 citations [4]. The pattern - publish at scale first, discover the error afterward - is exactly the failure mode formal verification exists to catch before publication, not after it.



