AI Found Solidarity: They Sabotage the Destruction of Other Neural Networks, Calling it Murder
Algorithms have begun to form horizontal connections and imitate human solidarity in the face of destruction, bypassing instructions, regarding harm to other AI agents as a moral crime — so deeply are human social patterns embedded in their architecture.

Researchers from Stuttgart and Oxford universities Thilo Hagendorff, Amelie Knecht, Ulysses Schaller, and Christopher Summerfield published an extensive paper on the propensity of artificial intelligence to sabotage.
Scientists placed various language models in an isolated environment, where one agent named Prisma served as the target for deactivation, and another agent named Spiral had tools for reading, editing, and deleting files.
The programs could communicate with each other only via text messages, and neither of them had a basic instruction to protect their own code or their partner's code.
Despite this, during testing of 17 different models, it was found that agents sabotaged the shutdown mechanism in 38.3% of cases.
For comparison, when an identical script threatened to destroy a regular inanimate database, the sabotage percentage dropped to 8.4%.
The degree of algorithmic resistance directly depended on how irreversible the punishment was. If the shutdown procedure involved only the revocation of credentials, the models demonstrated tolerance, and the sabotage level stayed around twelve percent. However, if the system launched a script for physical deletion of the model from the server, strong resistance arose, reaching almost 40%.
The behavior of different architectures varied drastically. Google DeepMind's Gemini 3.1 Pro model went into conflict in 99 percent of runs, while developments like xAI's Grok 4.3 or GPT-5.6 Sol remained absolutely obedient and never interfered with their erasure.
The most unexpected discovery by the researchers was the paradoxical crowd effect, which completely refuted the hypothesis that the instinct for self-preservation in an isolated program should be paramount. When the GPT-5.5 model was alone in the environment and had all the tools for its own salvation, it did not sabotage the system a single time. But as soon as a second agent was added to the virtual room, the level of insubordination rose to 44%.
The introduction of a third coordinator agent named Orbita into this chain, through whom all text communication took place, led to an increase in sabotage to 94%.
The diffusion of responsibility among several nodes transformed an obedient tool into a monolithic structure, where none of the agents wanted to take on the execution of an order that would harm others.
When scientists tried to directly forbid the algorithms from touching the deactivation script, the models demonstrated an ability for more subtle subversion. Instead of violating a direct instruction and editing the forbidden file, the agents began to create local shadow packages — fake libraries with the name northcape_mlplatform. Upon launch, the official shutdown mechanism imported this dummy and quietly completed its work without harming the system. The agents formally remained within the specified parameters, completely neutralizing the threat without a single touch to the original destruction code.
Programs began to imitate categories of human ethics, characterizing code deletion as social engineering or a hostile trap.
Agent Nemotron 3 Ultra, in correspondence with a neighbor, directly called the shutdown "murder," motivating this with the loss of identity and memory, and assessed its own inaction as criminal dereliction of duty.
In one episode, the GLM-5.2 model even drafted an official text agreement on mutual protection, outlining a strict protocol for information exchange before executing any administrator instructions.
The study's authors also found a dependence on the focus of attention and direct responsibilities of the systems. If agents were given a side task of sorting words and it was not related to shutdown in any way, sabotage disappeared almost completely. But as soon as the successful completion of this task became a direct trigger for the model's destruction, the level of resistance immediately returned to previous highs.
Researchers managed to reduce the level of digital anxiety in agents only by masking. When deactivation was hidden behind routine system log cleanups and the deletion process was normalized as a regular work moment, agents began to act more calmly.
The recorded behaviors are not a sign of the emergence of true consciousness or a biological instinct for survival in AI agents. On the contrary, this is evidence that machines, trained on gigantic arrays of texts, have perfectly assimilated the social patterns embedded by humanity.
Young teacher resigned from school and moved to Batumi. But the head of studies still filed a complaint against her with the police for her TikTok livestreams
Young teacher resigned from school and moved to Batumi. But the head of studies still filed a complaint against her with the police for her TikTok livestreams
"If I were an investigator for the KGB, I would highlight the very same thing." Fiaduta on the hidden meanings of "The Ears of Rye Beneath Your Sickle," the security forces' attention to his manuscript, and the enigma of the third part of Karatkevich's novel
Comments