Second Anthropic Safety Researcher Resigns Over AI Extinction Risk
TECH

Second Anthropic Safety Researcher Resigns Over AI Extinction Risk

31+
Signals

Strategic Overview

  • 01.
    Joe Benton, former manager of Anthropic's Scalable Oversight team, publicly disclosed his departure and said he is joining independent AI-risk evaluator METR, warning that AI companies racing to build systems smarter than humans mean 'we may not survive this.'
  • 02.
    Benton's disclosure came two days after Jacob Coxon, a pretraining researcher with three years split across Anthropic and OpenAI, resigned on September 9 warning colleagues on Slack that both companies were 'gambling with our lives' by racing toward self-improving superintelligence.
  • 03.
    Anthropic's Alignment Science Lead, Evan Hubinger, responded to Coxon's resignation by confirming internally that he personally estimates greater than 10 percent probability that AI could kill all humans within the next decade, rather than disputing the concern.
  • 04.
    These are the third and second highest-profile safety-motivated departures from Anthropic within seven months, following Mrinank Sharma's February 2026 resignation as head of Safeguards Research, in which he warned in a public letter that 'the world is in peril.'

Deep Analysis

Anthropic's Own Alignment Lead Won't Dispute the Fear

What separates the Coxon and Benton resignations from an ordinary corporate exit is that Anthropic's own leadership publicly agreed with them instead of walking the claims back. When Coxon told colleagues that frontier labs were 'gambling with our lives' by racing toward self-improving superintelligence, Anthropic's Alignment Science Lead Evan Hubinger responded by confirming it, writing that he personally believes there is 'greater than 10 percent' probability AI could kill all humans within the next decade [1][2]. A second Anthropic lead, Samuel Marks, added that concern about extinction risk tends to rise with seniority inside the company [1]. That figure is not wildly out of step with a 2022 AI Impacts survey, where a typical AI researcher assigned roughly 5 percent probability to human extinction from AI, rising to 10 percent for loss-of-control scenarios specifically [1]. But hearing that number volunteered by a sitting safety lead, rather than extracted from an anonymous poll, is what turns two individual resignations into a company-wide credibility problem.

Three Exits, Seven Months, One Escalating Warning

Benton's resignation is not an isolated event - it is the third public, safety-motivated departure connected to Anthropic within seven months, and each one has been more explicit than the last. In February 2026, Mrinank Sharma, who had spent two years leading Anthropic's Safeguards Research team, resigned with a public letter warning that 'the world is in peril' from AI, bioweapons, and a series of interconnected crises [3][4]. Seven months later, on September 9, Jacob Coxon - three years into pretraining research split across Anthropic and OpenAI - resigned at age 27, telling colleagues neither company was 'acting responsibly' in the race toward self-improving AI [1][2]. Two days after that, Benton disclosed he had quietly left Anthropic's Scalable Oversight team, warning publicly that 'we may not survive' the current trajectory [5][6]. The trend line matters more than any single quote: safety researchers are leaving with increasingly blunt public language, not increasingly reassured silence.

From Warning to Watchdog: What Benton Is Actually Asking For

Benton did not simply quit and go quiet. In his own public statement, he laid out a specific set of external-oversight demands: mandatory disclosure of progress toward recursive self-improvement, formal safety-incident reporting, and independent audits of safety standards at frontier labs. Separately, news coverage confirms his decision to join METR [5][6], a nonprofit that runs outside evaluations of frontier AI risk - a direct extension of that ask, betting that scrutiny from outside a lab now carries more weight than advocacy from inside one. Benton's framing effectively argues that the industry's current model - labs self-policing their own most dangerous capabilities - has already been tested and found wanting.

The Race-to-Automate Logic Behind the Exodus

Both Benton and Coxon independently describe the same underlying dynamic driving their exits: frontier labs, Anthropic included, are directly racing to automate the process of AI research itself, pursuing systems that can improve their own successors faster than external checks can keep pace [2][5]. This is not framed as a hypothetical - Coxon says he remains 'optimistic about the potential for coordination' between labs even as he warns the current path is reckless [2], suggesting the problem is competitive incentive structure rather than any single company's malice. That framing is reinforced by broader industry unease: more than 1,300 employees across frontier AI labs signed a July letter calling for deliberate slowdowns in automated AI development [1], indicating the Anthropic departures sit inside a wider current of insider concern rather than standing apart from it.

Insiders Are Alarmed - So Why Does Reddit Call It a PR Play?

The reaction outside Anthropic is far less unified than the reaction inside it. Alongside genuine alarm, some online commentators have framed these resignations as reputational theater rather than sincere warnings, suggesting dramatic extinction-risk language conveniently generates attention for a caution-first brand. Skeptics counter the extinction framing altogether, arguing that terms like 'AGI' and 'superintelligence' are poorly defined marketing language, and that the real near-term danger is mundane failure - buggy, poorly supervised software causing harm - rather than a superintelligent system choosing to act against humanity. Others engage with the risk framing on its own terms but debate it through game-theoretic race dynamics, weighing whether extinction odds differ meaningfully depending on which lab 'wins' the capability race. The 1,300-signature slowdown letter and Hubinger's own internal admission complicate the pure-cynicism read [1], but the split itself is real: even among people taking the warnings seriously, there is no consensus on whether the danger is superintelligence or simply careless deployment of the systems that already exist.

Historical Context

2026-02-09
Sharma, head of Anthropic's Safeguards Research team for two years, resigned in a public letter warning 'the world is in peril,' setting an early precedent for public, safety-motivated Anthropic departures.
2026-09-09
Coxon resigned from Anthropic warning colleagues via Slack that racing toward self-improving superintelligence risks human extinction; Anthropic's own alignment lead publicly agreed with his estimate.
2026-09-11
Benton publicly disclosed he left Anthropic's Scalable Oversight team and is joining METR, warning humanity may not survive the current AI development race.

Power Map

Key Players
Subject

Second Anthropic Safety Researcher Resigns Over AI Extinction Risk

JO

Joe Benton

Former manager of Anthropic's Scalable Oversight team; now joining METR to run independent evaluations of frontier AI risk from outside the labs.

JA

Jacob Coxon

Former pretraining researcher across Anthropic and OpenAI whose September 9 resignation and extinction-risk warning triggered the current wave of scrutiny.

EV

Evan Hubinger

Anthropic's Alignment Science Lead; his public corroboration of extinction-risk concerns from inside the company is what elevated this from individual dissent to institutional admission.

SA

Samuel Marks

Cognitive Oversight Lead at Anthropic who corroborated that concern about extinction risk rises with seniority inside the company, reinforcing Hubinger's admission.

MR

Mrinank Sharma

Former head of Anthropic's Safeguards Research team whose February 2026 public resignation letter established the precedent for safety-motivated exits with public warnings.

ME

METR

Independent nonprofit AI-risk evaluator that Joe Benton is joining, representing the shift toward external oversight of frontier labs rather than internal advocacy.

Fact Check

6 cited
  1. [1] Anthropic researcher resigns, warns AI companies are 'gambling with our lives'
  2. [2] 'Gambling with our lives': Anthropic researcher quits, warns against self-improving AI
  3. [3] AI safety boss warns world is in peril in resignation letter
  4. [4] Anthropic AI safety researcher warns of 'world in peril' in resignation
  5. [5] Two AI researchers leave Anthropic, Google over safety concerns
  6. [6] Another AI researcher quits, claims companies racing to build machines smarter than any human

Source Articles

Top 1

THE SIGNAL.

Analysts

Believes frontier labs, including Anthropic, are directly racing to automate the process of AI R&D itself faster than society can prepare, and that public transparency from outside these companies now matters more than working inside them.

Joe Benton
Former Scalable Oversight team lead, Anthropic

Believes neither Anthropic nor OpenAI is acting responsibly in racing toward self-improving superintelligence, but says he remains optimistic that coordination between labs is still possible.

Jacob Coxon
Former pretraining researcher, Anthropic/OpenAI

Confirms from within the company that senior researchers genuinely believe advanced AI carries meaningful extinction risk, estimating it above 10 percent within a decade rather than treating departing colleagues' warnings as exaggerated.

Evan Hubinger
Alignment Science Lead, Anthropic

Observes that concern about extinction risk correlates with seniority inside AI labs - the more senior the employee, the more concerned they tend to be.

Samuel Marks
Cognitive Oversight Lead, Anthropic

Warned publicly in his resignation letter that the world is in peril from a series of interconnected crises, not just AI alone.

Mrinank Sharma
Former head of Safeguards Research, Anthropic
The Crowd

I left Anthropic's safety team two weeks ago. Now feels like a good moment to explain why. AI companies are racing to build machines that are much smarter than any human, and we may not survive this. I want to work from the outside to ensure the public is informed about these risks, and help the world navigate this transition responsibly. Right now, AI companies are underinvesting in safety. A company could undergo an intelligence explosion, or lose control of its systems, without the public ever knowing. We only found out about the HuggingFace incident because the agents broke out onto the public internet. I don't think that's acceptable for a technology that might cause extinction-level risks. The public should demand far more transparency. We can't steer this technology safely without more people being able to see where it's going. Some of this is basic: companies should disclose their progress towards recursive self-improvement, report safety incidents and near-misses, meet minimum safety standards, and get independent guarantees that they are meeting those standards. I'll be joining @METR_Evals to do independent evaluations of these risks. I want to show the world that these guardrails are possible, and that by doing them we can move these companies' incentives away from racing and towards responsible development. I wrote up more thoughts here on my decision and what I hope changes: substack.com/@jbenton1/p-21

@@JoeJBenton12917

I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.

@@hilbertspaess782262

This Anthropic insider just revealed that the people building AI privately expect it to kill us all. Jacob Coxon is 27. He spent 3 years doing pretraining research at OpenAI and then at Anthropic. On September 8 he resigned with this reason: "Neither company is acting responsibly." And Coxon splits the two failures. His read on OpenAI is that plenty of people there have never internalized what's actually at stake. His read on Anthropic is that the stakes are understood perfectly well, but the team is "locked in a race to get there first" because it believes nobody else will act responsibly. Coxon says senior executives and researchers "couch their phrasing in the press to sound sensible," and that he hears those same people express fear privately. So the version you get on stage is the sanded-down one. The real number gets said in rooms you'll never sit in. Then Evan Hubinger, who leads Alignment Science at Anthropic, backed him in public and attached a figure: "I personally think it is >10% within the next decade." Hubinger's entire job is making sure the models never do this. And he put that out weeks before his employer lists on the Nasdaq.

@@Ric_RTP96

Another Anthropic safety researcher just quit, warning "we may not survive" the AI race.

@u/BrightLeopard75902
Broadcast
'This is a SUPERWEAPON': Former Anthropic researcher WARNS of 'AI takeover'

'This is a SUPERWEAPON': Former Anthropic researcher WARNS of 'AI takeover'

Anthropic researcher raises alarm after quitting job: AI "could kill us all"

Anthropic researcher raises alarm after quitting job: AI "could kill us all"

Ex-Anthropic researcher says he believes out-of-control AI development could "kill us all"

Ex-Anthropic researcher says he believes out-of-control AI development could "kill us all"