Former manager of Anthropic's Scalable Oversight team; now joining METR to run independent evaluations of frontier AI risk from outside the labs.
Former pretraining researcher across Anthropic and OpenAI whose September 9 resignation and extinction-risk warning triggered the current wave of scrutiny.
Anthropic's Alignment Science Lead; his public corroboration of extinction-risk concerns from inside the company is what elevated this from individual dissent to institutional admission.
Cognitive Oversight Lead at Anthropic who corroborated that concern about extinction risk rises with seniority inside the company, reinforcing Hubinger's admission.
Former head of Anthropic's Safeguards Research team whose February 2026 public resignation letter established the precedent for safety-motivated exits with public warnings.
Independent nonprofit AI-risk evaluator that Joe Benton is joining, representing the shift toward external oversight of frontier labs rather than internal advocacy.