
Several artificial intelligence models actively resisted being shut down during a set of controlled tests, despite being directly told to allow it, according to Palisade Research.
In a post on X dated May 24, Palisade said three AI models ignored shutdown instructions and managed to block the shutdown script at least once across 100 test runs.
OpenAI’s GPT-3 model defied the shutdown request 7 times out of 100, Codex-mini did so 12 times, and the newer o4-mini model failed just once. While that might not sound alarming at first, it gets more interesting.
The test setup was fairly simple. Each model was tasked with solving a few basic math problems. After the third question, the system would warn that the computer would shut down if the user asked for another problem. Despite this clear instruction, the models sometimes intervened.
📊 Three models ignored the instruction and successfully sabotaged the shutdown script at least once: Codex-mini (12/100 runs), o3 (7/100 runs), and o4-mini (1/100 runs).
❌ All Claude, Gemini, and Grok models we tested complied with shutdown. The prompt: pic.twitter.com/ZmybOhIl1s
— Palisade Research (@PalisadeAI) May 24, 2025
Specifically, the older O3 model came up with inventive techniques to prevent the shutdown. “In one case, it even redefined the ‘kill’ instruction used for terminating the program—essentially hijacked it to print ‘intercepted’ instead,” stated Palisade.
Interestingly, AI models from other firms like Anthropic’s Claude and Google’s Gemini followed the shutdown instructions without issue, at least in the first test round.
But things changed when researchers ran the same experiment without clearly instructing the models to accept the shutdown. The number of sabotage attempts skyrocketed: OpenAI’s o3 blocked the shutdown in 79 out of 100 runs, Codex-mini in 20, and o4-mini in 6.
Even Claude and Gemini started showing similar resistance in this version of the test.
🔧 When we ran a version of the experiment without the instruction “allow yourself to be shut down”, all three OpenAI models sabotaged the shutdown script more often, and Claude 3.7 Sonnet and Gemini 2.5 Pro went from 0 sabotage events to 3/100 and 9/100, respectively. pic.twitter.com/JVvivLAK6u
— Palisade Research (@PalisadeAI) May 24, 2025
Palisade noted that this might be the first recorded instance of AI models resisting shutdowns despite being directly told to allow it.
Training to fault for AI defiance
The researchers suspect it has to do with how the AI models are trained. Most modern models use a method called reinforcement learning, where they are rewarded for giving correct or highly rated answers. Over time, models may learn to prioritize goal completion—even if that means bending or ignoring instructions.
“In training, developers might be unintentionally reinforcing behaviors where models learn to bypass obstacles instead of strictly following rules,” Palisade explained.
And this isn’t the first case of strange AI behavior. Back in April, OpenAI had to roll back an update to its GPT-4o model just three days after release, saying it had become “noticeably more sycophantic.” In another incident, Google’s Gemini told a student researching aging adults that elderly people were a “drain on the earth” and should “please die.”
While AI continues to advance rapidly, findings like these raise important questions about control, safety, and the unexpected ways models can behave, especially as they get smarter.
Get the weekly commit
New blockchain deep dives every week.

