OpenAI's 722 AI-generated math proofs spark a mathematician backlash
TECH

OpenAI's 722 AI-generated math proofs spark a mathematician backlash

31+
Signals

Strategic Overview

  • 01.
    On October 6, 2026, OpenAI published 722 mathematical manuscripts generated by an unreleased internal frontier model, organized into 372 result families and posted to a public GitHub repository under an Apache-2.0 license; the model had been posed roughly 4,000 open problems, with each result representing on average about three hours of ChatGPT Pro-level thinking compute.
  • 02.
    One day after release, OpenAI withdrew three manuscripts connected to the Hodge conjecture after discovering a sign error that invalidated an argument the other two withdrawn papers depended on, reducing the total from 722 to 719, and separately revised 14 other papers and updated 13 citations.
  • 03.
    The release included claimed progress on the Riemann hypothesis (a partial 'quasi-Riemann hypothesis' result), a Lean-verified proof connected to the Unique Games Conjecture, and a resolution of the Hodge conjecture for CM abelian varieties.
  • 04.
    Only about 42% of the released proofs had undergone Lean formal verification and just 10 of the 719 manuscripts included released chain-of-thought summaries, and OpenAI kept the identity of the model that produced them undisclosed, both falling short of recommendations from its own mathematics advisory group.

Deep Analysis

The Verification Gap Behind the Headline Number

The Verification Gap Behind the Headline Number
OpenAI released 722 math manuscripts; only 42% had Lean formal verification, and three were withdrawn within a day over a sign error.

On October 6, 2026, OpenAI published 722 mathematical manuscripts generated by an unreleased internal frontier model, organized into 372 result families and posted to a public GitHub repository under an Apache-2.0 license [1]. The model had been posed roughly 4,000 open problems, with each published result representing on average about three hours of ChatGPT Pro-level thinking compute [1]. Among the headline claims: a partial 'quasi-Riemann hypothesis' result, a Lean-verified proof connected to the Unique Games Conjecture, and a resolution of the Hodge conjecture for CM abelian varieties [2].

But the number that matters more than 722 is 42 percent - the share of the released proofs that had actually undergone Lean formal verification, the automated proof-checking standard the field increasingly treats as a baseline for trust [3]. Just 10 of the resulting 719 manuscripts included released chain-of-thought summaries showing how the model reached its answers [3]. That gap turned out not to be theoretical: one day after release, OpenAI withdrew three manuscripts tied to the Hodge conjecture after discovering a sign error that invalidated an argument the other two withdrawn papers depended on, and separately revised 14 other papers and updated 13 citations [4]. The pattern - publish at scale first, discover the error afterward - is exactly the failure mode formal verification exists to catch before publication, not after it.

A Transparency Standard OpenAI Helped Write, Then Didn't Meet

The verification gap compounds a second, arguably more consequential one: disclosure. In September, following an earlier controversy over claims that its models had solved more than 100 longstanding math and theoretical computer science problems, OpenAI announced the formation of the nine-member Advisory Group on Mathematics and Artificial Intelligence (AGMAI), hosted at the Institute for Advanced Study in Princeton, to help rebuild trust before any large-scale release [5]. A week later, AGMAI issued release guidelines recommending that frontier labs disclose the specific model used, the prompts given, a summarized chain of thought, time taken, and compute cost for every claimed result [6].

OpenAI's October 6 release did not meet that bar. The company kept the identity of the model that produced all 722 (later 719) manuscripts undisclosed, publishing them through a company-controlled repository with only selective reasoning summaries [7]. University of Toronto mathematician Daniel Litt argued there is no justification for the secrecy: 'If we want to know the answers to these math questions, I see no reason why we should ask the company to keep them secret from us' [7]. MIT's Andrew Sutherland went further, arguing that without model release and independent replication, 'you should treat any claims about one-shotting problems with a single agent as unverified' [6]. OpenAI effectively wrote the rulebook mathematicians are now holding it to - and missed its own bar on the two dimensions, model identity and reasoning transparency, the advisory group called out specifically.

Math by Press Release: The Cost of Racing a Machine

The release's impact wasn't confined to retrospective grading of OpenAI's own rigor - it reshaped how human mathematicians were working in real time. As OpenAI's model closed in on a result related to the Unique Games Conjecture, MIT's Dor Minzer and his students Yumou Fei and Shuo Wang rushed to finish and publish their own related result rather than risk being overtaken by the company's release schedule [8]. Minzer described the asymmetry bluntly: 'You are human, right? You need to sleep, you need to eat, you have moods. You don't know if you're going to get scooped by the trillion-dollar company' [8].

Princeton's Mark Braverman framed the broader shift this competitive dynamic represents: 'Math by press release is not that healthy for math' [8]. The concern isn't only about credit - it's that a field's incentive structure, built around peer review, slow verification, and shared discovery, is being bent by a company that can post hundreds of claimed results at once and let the community sort out later which ones hold up. For early-career researchers especially, a problem they'd been working toward for months can be rendered moot overnight, regardless of whether the AI-generated version eventually survives scrutiny.

Tao's 'Strip-Mining' Critique and the Push for a Boycott

The sharpest institutional response came from the Association for Human Mathematics (AHM), whose statement - shared via Fields Medalist Terence Tao's blog - called the release 'not a demonstration of scholarship, but a demonstration of power' and urged mathematicians to discontinue collaboration with OpenAI [9]. Tao's own critique goes further than questions of verification or credit: he argues industrial-scale AI solving of open problems permanently changes the character of a field, because 'once a problem is considered solved, it can't be made unsolved again' [10]. In this framing, mass-producing proofs for hundreds of open problems in one release doesn't just risk errors - it strip-mines the research value out of those problems for good, foreclosing the years of community exploration, false starts, and cross-pollination that typically accompany a hard result finally falling.

That's the real stakes of the boycott call: not a one-time dispute over a sign error, but a disagreement over whether mathematics should optimize for solved problems at all, or for the slower process - what Tao calls 'Math 2.0' - of exposition, teaching, and communal understanding that a single company's release schedule cannot replicate no matter how fast its model runs.

The Credit Dispute Simmering Online

Beneath the formal advisory critique, a more personal dispute has been driving reaction across social platforms. NYU mathematician Tristan Buckmaster has said OpenAI's release has 'destroyed the careers of early career mathematicians' - a framing that resonated widely with online commentators treating the drop less as a research contribution than as a devaluation of the work young mathematicians depend on to build a career.

That sentiment is tangled up with an earlier, unresolved dispute in which Buckmaster and Anthropic mathematician Levent Alpoge alleged that OpenAI's model reproduced a novel, unpublished approach to the Navier-Stokes equations the pair had spent roughly a year feeding into the company's tools, and that OpenAI pushed back on their authorship claims rather than crediting them; OpenAI has denied using their session data. Broader online discussion of the 722-manuscript release has skewed in a similar skeptical register - less focused on any single proof's correctness than on the optics of posting hundreds of unvetted results at once without the kind of contextual commentary a journal submission would require.

Historical Context

2026-09-21
OpenAI announced formation of the nine-member Advisory Group on Mathematics and Artificial Intelligence, hosted at the Institute for Advanced Study in Princeton, to help rebuild trust after an earlier claim that its models had solved 100+ longstanding problems.
2026-09-29
The advisory group issued release guidelines recommending disclosure of model names, prompts, summarized chain-of-thought, time taken, and compute cost for every claimed result.
2026-10-06
OpenAI publicly released the 722 math manuscripts on GitHub, prompting MIT's Dor Minzer and students to rush publication of their own related result to avoid being scooped.
2026-10-07
OpenAI withdrew three manuscripts over a sign error, and the Association for Human Mathematics published its statement, shared via Terence Tao's blog, calling for a boycott.
2026-10-08
Reporting detailed how the release fell short of AGMAI's standards - only 42% Lean-verified and just 10 chain-of-thought summaries released - widening the controversy over AI math verification norms.

Power Map

Key Players
Subject

OpenAI's 722 AI-generated math proofs spark a mathematician backlash

OP

OpenAI

Released the 722 (later 719) AI-generated manuscripts from an unnamed internal frontier model, then withdrew three papers over a sign error one day later - its credibility and disclosure practices are now the center of the controversy.

AD

Advisory Group on Mathematics and Artificial Intelligence (AGMAI)

Nine-member independent panel hosted at Princeton's Institute for Advanced Study that set release guidelines - model identity, compute cost, reasoning summaries - which OpenAI's October 6 release did not fully meet.

AS

Association for Human Mathematics (AHM)

Mathematician advocacy group, whose statement was shared via Terence Tao's blog, publicly called the release 'a demonstration of power' rather than scholarship and urged mathematicians to discontinue collaboration with OpenAI.

TE

Terence Tao

Fields Medalist whose 'Math 2.0' critique - that industrial-scale AI problem-solving permanently reduces a field's research value - became the backlash's most-quoted argument.

DO

Dor Minzer (MIT)

Rushed to finish and publish his own Unique Games Conjecture-related result with students Yumou Fei and Shuo Wang to avoid being overtaken by OpenAI's release, illustrating the direct competitive pressure the drop put on working mathematicians.

AN

Andrew Sutherland (MIT) and Daniel Litt (University of Toronto)

Mathematicians who publicly challenged OpenAI's unverified claims and the withheld model identity, arguing results should be treated as unverified without replication and disclosure.

Fact Check

10 cited
  1. [1] OpenAI Releases 722 Math Manuscripts From an Unreleased AI Model
  2. [2] OpenAI Releases AI-Generated Proofs Of Quasi Riemann Hypothesis, Unique Games Conjecture & Hodge Conjecture For CM Abelian Varieties
  3. [3] OpenAI's Math Solutions Aren't Meeting the Field's Standards, Yet
  4. [4] OpenAI Withdraws Preprints Among 722 Manuscripts on Unsolved Math Problems
  5. [5] OpenAI Launches Independent Mathematics Advisory Group Following Backlash Over AI Math Claims
  6. [6] OpenAI Says Secret AI Model Solved Advanced Math Problems
  7. [7] OpenAI's 722 Maths Papers Are Open. The Model Behind Them Is Not.
  8. [8] As AI Closed In on Unique Games Proof, Researchers Raced to Beat the Machines
  9. [9] AHM Statement on OpenAI's October 6 Release of Mathematical Documents
  10. [10] Some Mathematicians Call for OpenAI Boycott After AI-Generated Proofs Flood Their Field

Source Articles

Top 5

THE SIGNAL.

Analysts

“Argues industrial-scale AI solving of open problems 'strip-mines' a field because a solved problem can't be unsolved again, and calls for a 'Math 2.0' valuing exposition and community understanding over raw problem-solving volume: 'Once a problem is considered solved, it can't be made unsolved again.'”

Terence Tao
Fields Medalist, UCLA

“Says unverified, single-prompt claims should not be trusted absent model release and independent replication: 'Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified.'”

Andrew Sutherland
Mathematician, MIT

“Argues there is no justification for OpenAI keeping the model's identity secret from the mathematicians who need to evaluate its output: 'If we want to know the answers to these math questions, I see no reason why we should ask the company to keep them secret from us.'”

Daniel Litt
Mathematician, University of Toronto

“Describes the psychological and competitive pressure of racing an AI lab's release schedule: 'You are human, right? You need to sleep, you need to eat, you have moods. You don't know if you're going to get scooped by the trillion-dollar company.'”

Dor Minzer
Mathematician, MIT

“Warns that mathematics is becoming driven by corporate press releases rather than peer-reviewed process: 'Math by press release is not that healthy for math.'”

Mark Braverman
Mathematician, Princeton
The Crowd

“NYU math professor Tristan Buckmaster says OpenAI's Tuesday drop of 722 AI-generated math papers has "destroyed the careers of early career mathematicians."”

@@unusual_whales3260

“Mathematicians are still working through OpenAI's 160+ page Navier–Stokes proof a month later. I asked Claude Fable 5.5 to explain it. It made this 5-min 3D video: the proof in 8 steps. Caveats: Lean-checked, not yet fully human-verified. Needs an outside force.”

@@imjustnewatai517

“OPENAI DROPPED 722 MATH PAPERS WRITTEN BY ITS SECRET MODEL. Every manuscript came from an internal system that hasn't been released yet. They're grouped into 372 families across different fields. Among the boldest claims: - a zeta zero-free region, billed as the quasi-Riemann hypothesis...”

@@Mikadzyki_NFT94

“OpenAI unleashes hundreds more math results upon a field already in shock”

@u/ResultBackground24506000
Broadcast
OpenAI's biggest math breakthrough is getting ugly...

OpenAI's biggest math breakthrough is getting ugly...

BIGGEST Proof Dump in History! | OpenAI Release 722 Manuscripts

BIGGEST Proof Dump in History! | OpenAI Release 722 Manuscripts

OpenAI Just Broke Math With Its Most Powerful AI Yet

OpenAI Just Broke Math With Its Most Powerful AI Yet