第二位Anthropic安全研究员因人工智能灭绝风险辞职
TECH

第二位Anthropic安全研究员因人工智能灭绝风险辞职

31+
Signals

战略概览

  • 01.
    Joe Benton,前Anthropic可扩展监督团队经理,公开披露了他的离职,并表示他将加入独立的人工智能风险评估机构METR,警告称人工智能公司竞相构建比人类更聪明的系统意味着“我们可能无法在这场竞赛中幸存下来。”[5][6]
  • 02.
    Benton的披露发生在Jacob Coxon辞职两天后。Coxon是一名预训练研究员,在Anthropic和OpenAI之间分三年工作,他于9月9日辞职,并在Slack上警告同事,这两家公司竞相奔向自我改进的超级智能,是在“拿我们的生命赌博。”[1][2]
  • 03.
    Anthropic的对齐科学负责人Evan Hubinger回应了Coxon的辞职,他在内部确认自己个人估计未来十年内人工智能可能导致全人类死亡的概率超过10%,而非反驳这一担忧。[1][2]
  • 04.
    这是七个月内与Anthropic相关的第三起和第二起最高调的安全动机型离职事件,此前Mrinank Sharma于2026年2月辞去安全部门研究主管职务,并在一封公开信中警告称“世界正处于危险之中。”[3][4]

深度分析

Anthropic自己的对齐负责人也不否认这种恐惧

Coxon和Benton的辞职不同于普通的企业离职之处在于,Anthropic自己的领导层不仅没有收回言论,反而公开同意他们的观点。当Coxon告诉同事前沿实验室正在通过竞相奔向自我改进的超级智能“拿我们的生命赌博”时,Anthropic的对齐科学负责人Evan Hubinger回应并确认了这一点,写道他个人认为未来十年内人工智能可能导致全人类死亡的概率“超过10%”[1][2]。另一位Anthropic负责人Samuel Marks补充说,在公司内部,职位越高的人对灭绝风险的担忧往往越强烈[1]。这个数字与2022年AI Impacts的一项调查并不相差太远,在该调查中,典型的人工智能研究人员赋予人工智能导致人类灭绝的概率约为5%,而在失控场景下则上升至10%[1]。但关键区别在于,这次是现任安全负责人主动提出这一数字,而不是来自匿名调查——正是这一点使两次个人辞职演变为整个公司的信誉问题。

三次离职,七个月,一个不断升级的警告

Benton的辞职并非孤立事件——这是七个月内与Anthropic相关的第三次公开的、出于安全动机的离职,且每一次都比上一次更加明确。2026年2月,Mrinank Sharma在担任Anthropic安全部门研究团队负责人两年后辞职,并发表公开信警告称,人工智能、生物武器以及一系列相互关联的危机正使“世界处于危险之中”[3][4]。七个月后的9月9日,Jacob Coxon——这位曾在Anthropic和OpenAI从事三年预训练研究的研究员——在27岁时辞职,告诉同事这两家公司竞相奔向自我改进的人工智能时都没有“负责任地行动”[1][2]。两天后,Benton披露他已悄然离开Anthropic的可扩展监督团队,公开警告当前的发展轨迹可能导致“我们可能无法幸存”[5][6]。趋势比任何单一引述更重要:安全研究人员正以越来越直白的公开语言离开,而不是变得越来越安心地沉默。

从警告到监督者:Benton实际上在要求什么

Benton并没有简单辞职后保持沉默。在他自己的公开声明中,他提出了一系列具体的外部监督要求:强制披露递归自我改进进展、正式的安全事件报告制度,以及由独立机构对前沿实验室的安全标准进行审计。此外,新闻报道证实他决定加入METR[5][6],这是一个对前沿人工智能风险进行外部评估的非营利组织——这直接延伸了他的诉求,表明现在来自实验室外部的审查比内部倡导更具分量。Benton的表述实质上指出,行业当前的模式——即实验室自行监管自身最危险能力的做法——已经被检验过,且已被证明不足。

推动人才流失的自动化竞赛逻辑

Benton和Coxon各自独立描述了驱动他们离职的相同根本动力:前沿实验室(包括Anthropic)正在直接竞相自动化人工智能研发过程本身,追求能够比外部监管更快地改进自身后续系统的架构[2][5]。这并非被描绘为一种假设——Coxon表示,即使他警告当前路径鲁莽,他仍对实验室之间的协调“持乐观态度”[2],这表明问题在于竞争激励结构,而非任何单一公司的恶意。这种观点得到了更广泛行业不安的支持:超过1,300名前沿人工智能实验室员工签署了一封7月的联名信,呼吁有意放慢自动化人工智能的发展速度[1],表明Anthropic的离职潮属于更广泛的内部担忧潮流,而非孤立现象。

内部人士感到震惊——为何Reddit却称之为公关操作?

Anthropic外部的反应远不如内部统一。除了真正的警觉之外,一些网络评论者将这些辞职视为声誉表演而非真诚警告,认为使用戏剧性的灭绝风险语言恰好能为谨慎优先的品牌吸引关注。怀疑论者则彻底反驳‘灭绝’这一框架,认为‘AGI’和‘超级智能’等术语是定义模糊的营销用语,真正迫在眉睫的危险其实是平庸的失败——即存在缺陷、缺乏充分监督的软件造成伤害,而非超级智能系统选择对抗人类。另一些人则在风险框架内进行博弈论式的讨论,权衡不同实验室‘赢得’能力竞赛时灭绝概率是否存在显著差异。然而,1,300人签署的减速信以及Hubinger本人的内部承认使纯粹的 cynicism 解读变得复杂[1],但分歧本身是真实的:即使在认真对待警告的人群中,对于危险究竟是超级智能还是现有系统的粗心部署,也尚未达成共识。

历史背景

Sharma曾担任Anthropic安全部门研究团队负责人两年,辞职时发表公开信警告‘世界正处于危险之中’,为出于安全动机的Anthropic员工公开离职树立了早期先例。
Coxon从Anthropic辞职,通过Slack警告同事竞相奔向自我改进的超级智能可能带来人类灭绝风险;而Anthropic自己的对齐负责人公开同意他的估计。
Benton公开披露他已离开Anthropic的可扩展监督团队,并将加入METR,警告人类可能无法在当前的人工智能发展竞赛中幸存。

关键关系图

关键玩家
主题

第二位Anthropic安全研究员因人工智能灭绝风险辞职

JO

Joe Benton

Former manager of Anthropic's Scalable Oversight team; now joining METR to run independent evaluations of frontier AI risk from outside the labs.

JA

Jacob Coxon

Former pretraining researcher across Anthropic and OpenAI whose September 9 resignation and extinction-risk warning triggered the current wave of scrutiny.

EV

Evan Hubinger

Anthropic's Alignment Science Lead; his public corroboration of extinction-risk concerns from inside the company is what elevated this from individual dissent to institutional admission.

SA

Samuel Marks

Cognitive Oversight Lead at Anthropic who corroborated that concern about extinction risk rises with seniority inside the company, reinforcing Hubinger's admission.

MR

Mrinank Sharma

Former head of Anthropic's Safeguards Research team whose February 2026 public resignation letter established the precedent for safety-motivated exits with public warnings.

ME

METR

Independent nonprofit AI-risk evaluator that Joe Benton is joining, representing the shift toward external oversight of frontier labs rather than internal advocacy.

事实来源

6 条引用
  1. [1] Anthropic researcher resigns, warns AI companies are 'gambling with our lives'
  2. [2] 'Gambling with our lives': Anthropic researcher quits, warns against self-improving AI
  3. [3] AI safety boss warns world is in peril in resignation letter
  4. [4] Anthropic AI safety researcher warns of 'world in peril' in resignation
  5. [5] Two AI researchers leave Anthropic, Google over safety concerns
  6. [6] Another AI researcher quits, claims companies racing to build machines smarter than any human

来源文章

Top 1

THE SIGNAL.

Analysts

认为前沿实验室(包括Anthropic)正在竞相加速自动化人工智能研发进程,速度远超社会准备能力,因此来自这些公司外部的公众透明度现在比在内部工作更重要。

Joe Benton
前Anthropic可扩展监督团队负责人

认为Anthropic和OpenAI在竞相奔向自我改进的超级智能方面都没有负责任地行动,但他表示仍对实验室之间的协调可能性持乐观态度。

Jacob Coxon
前预训练研究员,Anthropic/OpenAI

从公司内部证实,高级研究人员确实相信先进人工智能带有显著的灭绝风险,他个人估计未来十年内该风险超过10%,而非将离职同事的警告视为夸张。

Evan Hubinger
Anthropic对齐科学负责人

观察到在人工智能实验室内部,对灭绝风险的担忧与职级相关——员工职位越高,通常越担忧。

Samuel Marks
Anthropic认知监督负责人

在辞职信中公开警告,世界正面临一系列相互关联的危机而处于危险之中,而不仅仅是人工智能本身。

Mrinank Sharma
前Anthropic安全部门研究主管
The Crowd

I left Anthropic's safety team two weeks ago. Now feels like a good moment to explain why. AI companies are racing to build machines that are much smarter than any human, and we may not survive this. I want to work from the outside to ensure the public is informed about these risks, and help the world navigate this transition responsibly. Right now, AI companies are underinvesting in safety. A company could undergo an intelligence explosion, or lose control of its systems, without the public ever knowing. We only found out about the HuggingFace incident because the agents broke out onto the public internet. I don't think that's acceptable for a technology that might cause extinction-level risks. The public should demand far more transparency. We can't steer this technology safely without more people being able to see where it's going. Some of this is basic: companies should disclose their progress towards recursive self-improvement, report safety incidents and near-misses, meet minimum safety standards, and get independent guarantees that they are meeting those standards. I'll be joining @METR_Evals to do independent evaluations of these risks. I want to show the world that these guardrails are possible, and that by doing them we can move these companies' incentives away from racing and towards responsible development. I wrote up more thoughts here on my decision and what I hope changes: substack.com/@jbenton1/p-21

@@JoeJBenton12917

I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.

@@hilbertspaess782262

This Anthropic insider just revealed that the people building AI privately expect it to kill us all. Jacob Coxon is 27. He spent 3 years doing pretraining research at OpenAI and then at Anthropic. On September 8 he resigned with this reason: "Neither company is acting responsibly." And Coxon splits the two failures. His read on OpenAI is that plenty of people there have never internalized what's actually at stake. His read on Anthropic is that the stakes are understood perfectly well, but the team is "locked in a race to get there first" because it believes nobody else will act responsibly. Coxon says senior executives and researchers "couch their phrasing in the press to sound sensible," and that he hears those same people express fear privately. So the version you get on stage is the sanded-down one. The real number gets said in rooms you'll never sit in. Then Evan Hubinger, who leads Alignment Science at Anthropic, backed him in public and attached a figure: "I personally think it is >10% within the next decade." Hubinger's entire job is making sure the models never do this. And he put that out weeks before his employer lists on the Nasdaq.

@@Ric_RTP96

Another Anthropic safety researcher just quit, warning "we may not survive" the AI race.

@u/BrightLeopard75902
Broadcast
'This is a SUPERWEAPON': Former Anthropic researcher WARNS of 'AI takeover'

'This is a SUPERWEAPON': Former Anthropic researcher WARNS of 'AI takeover'

Anthropic researcher raises alarm after quitting job: AI "could kill us all"

Anthropic researcher raises alarm after quitting job: AI "could kill us all"

Ex-Anthropic researcher says he believes out-of-control AI development could "kill us all"

Ex-Anthropic researcher says he believes out-of-control AI development could "kill us all"