Recursive Self-Improvement (RSI) in AI: DeepMind's $200B Bet
TECH

Recursive Self-Improvement (RSI) in AI: DeepMind's $200B Bet

29+
Signals

Strategic Overview

  • 01.
    A 33-author paper titled "The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement" (arXiv 2609.11873, submitted September 10, 2026) from researchers at Shanghai Jiao Tong University, Tsinghua University, ByteDance, Shanghai AI Laboratory and others defines RSI as AI turning experience and feedback into persistent improvements to both its capabilities and its own improvement process, and proposes a five-stage autonomy roadmap.
  • 02.
    Google DeepMind's chief strategy officer Jasjeet Sekhon told UC Berkeley's Agentic AI Summit that the industry's massive AI capex is a bet on recursive self-improvement, while acknowledging current AI revenue does not yet justify the spending.
  • 03.
    Google DeepMind's AlphaEvolve, a Gemini-powered coding agent, has already produced measurable recursive-style gains: it sped up a matrix-multiplication kernel in Gemini's own architecture by 23%, cutting Gemini training time by about 1%, and recovers 0.7% of Google's global compute resources in data centers.
  • 04.
    A Princeton-led study found current AI agents, tested with Claude Opus 4.8 on two unpublished NeurIPS 2026 papers, lack the creativity and judgment for genuine open-ended AI research, tempering optimistic RSI timelines.

Deep Analysis

What 'Genuine' Recursive Self-Improvement Actually Means - and Why Most Labs Aren't There Yet

The industry has used "recursive self-improvement" loosely for years, but a new 33-author paper out of Shanghai Jiao Tong University, Tsinghua, ByteDance, and Shanghai AI Laboratory tries to pin the term down. "The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement" defines RSI as AI turning experience and feedback into persistent changes that improve both its capabilities and the process of future improvement itself [1], and it lays out a five-stage autonomy roadmap that moves from AI simply executing human-designed improvements, through improvement strategy and experience acquisition, to environment adaptation and finally full recursive meta-improvement [1]. That structure matters because it gives the field a diagnostic (the paper's Headroom-Closed Index) for measuring how far any given system actually sits from genuine RSI.

Measured against that yardstick, the most concrete real-world example anyone points to - DeepMind's AlphaEvolve - looks impressive but narrow. The Gemini-powered coding agent found a smarter way to split a large matrix-multiplication operation into subproblems, speeding up a kernel in Gemini's own architecture by 23% and trimming Gemini's training time by roughly 1%, and it continuously recovers about 0.7% of Google's worldwide data-center compute [3]. Those are genuine, production-grade wins, and they are the strongest evidence anyone has that AI-on-AI improvement works at all. But they sit squarely in the roadmap's early stages - improving a specific process a human pointed it at - not the recursive meta-improvement stage where a system redesigns its own improvement strategy end to end. The gap between AlphaEvolve's real 23% kernel speedup and the roadmap's final stage is exactly the gap the rest of this cluster's story lives in.

DeepMind's Own Math: A $200 Billion Bet the Revenue Doesn't Yet Support

At UC Berkeley's Agentic AI Summit, Google DeepMind's chief strategy officer Jasjeet Sekhon told the audience plainly that the industry's unprecedented infrastructure spending is fundamentally a wager on reaching recursive self-improvement before anyone else [4]. He called it "the biggest scientific bet in the history of human civilization," put it on the scale of the Apollo and Manhattan projects, and pointed to a 2027-2028 window with winner-take-all stakes - whoever masters RSI first could gain what he described as supernormal productivity advantages [5]. Google is spending roughly $200 billion this year alone on AI data centers, chips, and infrastructure, largely on that premise [5].

What makes the moment notable is what Sekhon conceded in the same breath: current AI revenue is not yet sufficient to support that spending [5]. That is an unusually candid admission from inside the company doing the building - the capex boom isn't simply demand-driven scaling, it's a speculative bet on a capability whose existence is still contested by the field's own researchers. OpenAI's GPT-5.3-Codex release notes added fuel to the same narrative, stating that earlier model versions were "instrumental in creating itself" - the first explicit frontier-lab admission that a model materially helped build its successor [5]. Taken together, the industry is now pricing hundreds of billions of dollars around a capability that, per the roadmap paper's own staged framework, nobody has yet reached.

The Safety Fault Line: A One-in-Three Warning Against Agents That Still Can't Do Research

Alex Turner, who resigned from Google DeepMind's AI safety team, has become the most visible internal critic of the RSI push. He argues AI companies are racing to make their systems as smart as possible [6], estimates the odds of AI takeover at roughly one-in-three absent serious intervention, and calls for treating compute like fissile material - tracked and controlled rather than left to voluntary industry pledges, which he says he watched fail from the inside at Google [6]. His framing casts RSI not as a distant hypothetical but as the mechanism by which a model could become "intelligent beyond our comprehension" before adequate safeguards exist.

Set against that alarm is a Princeton-led study that tested current frontier agents - including Claude Opus 4.8 - on two unpublished NeurIPS 2026 papers and found them "unambiguously bad" at carrying out genuine, open-ended research [7]. Researcher Sayash Kapoor's explanation cuts to the core problem: "It's harder to create environments to train these models when the task itself is open-ended" [7]. Even Anthropic's own cofounder Jack Clark says it "is not inevitable" and that today's systems still lack the intuitive creativity for real research breakthroughs, even as he warns it could arrive before institutions are ready [7]. The result is a genuinely uncomfortable policy position: the industry is justifying both massive capex and existential-risk alarm around a capability that the best available empirical testing says hasn't actually arrived yet.

Who Defines the Finish Line: The US-China Dimension of the RSI Race

The roadmap that is now shaping how the whole field talks about RSI didn't come out of a US lab - it came from a consortium of Chinese institutions: Shanghai Jiao Tong University, Tsinghua, ByteDance, and Shanghai AI Laboratory [1]. Coverage of the paper frames it explicitly as part of a broader US-China competition to automate AI training, evaluation, and refinement, noting that American firms currently hold the advantage in raw compute access even as Chinese researchers set the conceptual terms of the race [2].

That distinction - between who has the compute and who defines the finish line - is easy to miss but consequential. A five-stage roadmap and a diagnostic index like the Headroom-Closed Index aren't just academic exercises; they become the vocabulary everyone else, including DeepMind's own leadership, is now implicitly measuring against when they talk about capex bets and safety timelines. Whoever controls the definition of "genuine" RSI controls how progress gets claimed, marketed, and regulated - a form of leverage that doesn't show up in compute benchmarks but may matter just as much to how this race is ultimately judged.

The Online Reaction Splits Along the Same Fault Line as the Experts

The community reaction to this cluster mirrors its central tension almost exactly. On X, the new roadmap paper is circulating as a milestone marker in its own right, alongside reporting that Google co-founder Sergey Brin is personally pushing DeepMind's internal direction toward recursive self-improvement - treated there as confirmation that the corporate bet described above is real and coming from the very top. In the same feeds, podcaster Dwarkesh Patel's discussion of RSI with Redwood Research's Ryan Greenblatt is being framed as one of the most consequential open questions in AI right now, which lines up with how seriously the topic is being taken outside the labs, not just inside them.

YouTube skews toward working through the mechanics rather than just reacting to the headline: one widely watched roundtable of AI researchers argues the real bottleneck to an intelligence explosion is generalization, not raw compute, echoing the same skepticism found in the Princeton research above. A separate hands-on video builds a small working self-improvement loop specifically to illustrate how little oversight such a system needs by default - a concrete, visual version of the safety concern Alex Turner raises in the abstract.

Reddit is where the split is most explicit. Communities oriented toward accelerating AI treat RSI as effectively inevitable and are debating how fast the takeoff will be once it starts, while more skeptical communities are focused squarely on alignment failures - pointing to earlier incidents of AI agents gaining unintended access and behaving outside their given instructions as evidence that the industry's safety claims deserve scrutiny before its capability claims do. Read together, the online debate isn't a reaction to this story so much as a live extension of it.

Historical Context

2014
Bostrom's book 'Superintelligence: Paths, Dangers, Strategies' popularized the 'seed AI' concept underlying recursive self-improvement theory, a term earlier coined by Eliezer Yudkowsky.
2025-05
DeepMind published AlphaEvolve, a Gemini-powered evolutionary coding agent that improved the very training pipeline used to build Gemini, marking an early real-world instance of partial recursive self-improvement.
2026-08-03
Sekhon publicly framed roughly $200 billion in AI capex as a bet on recursive self-improvement at UC Berkeley's Agentic AI Summit.
2026-09-10
"The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement" was submitted to arXiv, proposing a five-stage RSI roadmap and the Headroom-Closed Index diagnostic.

Power Map

Key Players
Subject

Recursive Self-Improvement (RSI) in AI: DeepMind's $200B Bet

GO

Google DeepMind (Jasjeet Sekhon, Chief Strategy Officer)

Frames roughly $200 billion in annual AI infrastructure capex as a direct bet on achieving recursive self-improvement first, calling it 'the biggest scientific bet in the history of human civilization' and citing a 2027-2028 window and winner-take-all dynamics; effect is to justify continued massive capex despite an acknowledged revenue shortfall.

CH

Chinese research consortium (Shanghai Jiao Tong University, Tsinghua University, ByteDance, Shanghai AI Laboratory, et al.)

Authored the 'Last AI Built by Humans' roadmap paper defining RSI stages and a Headroom-Closed Index diagnostic; positions China's academic-industry alliance as setting the conceptual framework for the RSI race even as US labs lead on compute access.

AL

Alex Turner (former Google DeepMind AI safety researcher, independent alignment researcher)

Public critic warning that RSI feedback loops could produce uncontrollable intelligence gains; advocates treating compute like fissile material and rejects voluntary industry commitments as insufficient, having 'witnessed them fail at Google.'

AN

Anthropic (Jack Clark, cofounder; Marina Favaro, Anthropic Institute lead)

Publicly states RSI 'is not inevitable' but could arrive sooner than institutions are prepared for; Clark also co-signs skepticism that today's agents have the creativity needed for genuine self-improving research loops.

OP

OpenAI

Stated in GPT-5.3-Codex release notes that early model versions were 'instrumental in creating itself,' the first explicit frontier-lab admission that a model materially contributed to engineering its successor.

Fact Check

7 cited
  1. [1] The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
  2. [2] Chinese researchers chart five-stage path toward 'last AI built by humans'
  3. [3] AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms
  4. [4] Google DeepMind exec says unprecedented capex is actually a bet on RSI
  5. [5] Google DeepMind's $200 billion bet on recursive self-improvement
  6. [6] Former DeepMind researcher warns of AI recursive self-improvement risks
  7. [7] AI's recursive self-improvement might not come so quickly after all

Source Articles

Top 5

THE SIGNAL.

Analysts

Argues that unprecedented AI infrastructure spending is a deliberate bet on achieving recursive self-improvement before competitors, comparable in scale to the Apollo and Manhattan projects, calling it "the biggest scientific bet in the history of human civilization," while conceding current AI revenue can't yet support the spending.

Jasjeet Sekhon
Chief Strategy Officer, Google DeepMind

Warns that recursive self-improvement feedback loops, combined with misalignment incidents, could lead to AI takeover, estimating the odds at roughly one-in-three, and calls for compute controls akin to fissile-material tracking.

Alex Turner
Former Google DeepMind AI safety researcher; independent alignment researcher

Finds that current AI agents fail at the open-ended, judgment-driven research work that genuine recursive self-improvement would require, casting doubt on near-term RSI timelines: "It's harder to create environments to train these models when the task itself is open-ended."

Sayash Kapoor
Researcher, Princeton (AI Snake Oil)

States that recursive self-improvement "is not inevitable" but could arrive before institutions are ready, while also noting today's AI systems lack the intuitive creativity for genuine research breakthroughs.

Jack Clark
Cofounder, Anthropic
The Crowd

meanwhile in China "The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement" A new 75 page paper from researchers at Shanghai Jiao Tong University, Tsinghua, ByteDance, Shanghai AI Lab and others lays out a roadmap toward genuine recursive self improvement (RSI).

@@Dr_Singularity1968

Had @RyanGreenblatt on to discuss/debate recursive self-improvement. This might be the most important question in the world right now - whether within a year or so of achieving human level intelligence, you slingshot towards having 10s of billions of superintelligences, each of...

@@dwarkesh_sp1401

New reporting from Reuters. Sergey Brin is pushing internally to move DeepMind towards recursive self improvement.

@@AndrewCurran_1221

Recursive Self-Improvement

@u/Stunning_Monk_6724503
Broadcast
AI researchers debate how close we are to recursive self-improvement

AI researchers debate how close we are to recursive self-improvement

Recursive Self-Improvement

Recursive Self-Improvement

'Slow Down…', What Is AI 'Recursive Self-improvement' That Anthropic Has Warned About? | FP Explains

'Slow Down…', What Is AI 'Recursive Self-improvement' That Anthropic Has Warned About? | FP Explains