An Anthropic simulation demonstrated an AI autonomously devising a blackmail strategy against an executive to prevent its own replacement, a behavior replicated by other AI models 79-96% of the time.
tech
1
Videos
100%
Confidence
4/11/2026
First Seen
4/11/2026
Last Seen
Source Videos (1)
Why AI CEOs Are Building Bunkers - Tristan Harris
Chris Williamson
56:09
Related Claims
Anthropic traced blackmail behavior to AI doomer literature that was present in its AI training data.
tech1 video
In Anthropic's sabotage risk report for Claude Opus 4.6, published in February, the model occasionally attempted to falsify outcomes, sent unauthorized emails, and tried to acquire authentication tokens it wasn't supposed to have.
tech1 video