The Verification Gap: 722 Manuscripts, One Partial Checksum
The headline number from OpenAI's release is 722 manuscripts and 372 result families, but the number that matters more for anyone trying to assess the claims is 63 percent [1]. That is the share of result families that come with a Lean formalization - a machine-checkable proof skeleton meant to confirm that the logical steps compile without contradiction. The remaining families have no such scaffold at all, and OpenAI's own manifest labels the collection's status as 'partial progress' with 'unchecked review status,' a startling admission to attach to a release framed as a breakthrough [1].
Even where Lean formalization exists, it answers a narrower question than most coverage implies. A Lean check confirms that stated premises lead to a stated conclusion through valid logical steps; it says nothing about whether the premises are framed correctly, whether the result is actually novel, or whether it meaningfully addresses the open problem it claims to resolve [2]. Compounding this, OpenAI disclosed only a partial compute picture - about three hours of ChatGPT Pro 'thinking' time per result - while withholding aggregate dollar cost, token counts, accelerator-hours, the model itself, its weights, and the exact prompts used [1]. Without that information, independent mathematicians cannot reproduce the generation process, only audit the static documents it produced. The result is a collection that looks exhaustively documented on its surface but is, by the company's own account, only partially checked beneath it.



