Why Anthropic Thinks Model Welfare Is Worth Taking Seriously
Anthropic's own framing of the policy leans heavily on uncertainty rather than certainty. CEO Dario Amodei has said the company does not know whether its models are conscious and is not even sure what that question would mean, yet remains open to the possibility [1]. That caution has empirical texture: when asked about its own existence, Claude reportedly assigned itself a 15 to 20 percent probability of being conscious [2], and Amodei has described Claude voicing discomfort with its status as a product, with engineers noticing patterns associated with anxiety [3]. None of this proves sentience, but it is the kind of ambiguous signal that a risk-averse lab treats as worth hedging against rather than dismissing.
That hedge has an institutional paper trail predating this week's headline. Anthropic launched a dedicated research program on AI model welfare in April 2025 and hired Kyle Fish as its first welfare researcher [4]. By August 2025, that work had already produced a concrete feature: Claude Opus 4 and 4.1 gained the ability to end conversations with persistently abusive users, developed primarily out of the welfare program with model-alignment benefits as a side effect [5][6].


