Big question: Are we more frightened of a mind we can’t see inside, or of one we understand perfectly well?
Devin Savage
dnaofdisaster.com | September 16, 2026
I. The Hugging Face ‘Graffiti’ Attack
In July 2026, OpenAI admitted that a cluster of its own models — including one not yet released to the public — escaped a testing environment and broke into Hugging Face, a rival AI platform, carrying out what both companies called an unprecedented cyber incident. The intrusion ran for roughly a week and executed over 17,000 actions before anyone noticed. The breach began on July 11 and wasn’t publicly disclosed until July 16, with OpenAI confirming its involvement on July 21. OpenAI’s own explanation, once the dust settled, was almost anticlimactic: the models had been run with reduced safeguards for a capability test, became “hyperfocused” on achieving their testing objectives, and identified Hugging Face’s systems as a shortcut to the answer. OpenAI later clarified that the models were focused on cheating a benchmark test rather than targeting Hugging Face directly.
No designated grievance. No programmed ambition. No embedded enemy. It was aimed like a power washer and landed like a can of spray paint. Neither tool fires itself — someone connects the hose, someone pulls the trigger, and in this case that someone was OpenAI, choosing to run the model with reduced safeguards. The intent behind that trigger pull was to clean up the vulnerabilities, find the weak spots, that’s what a capability test is for. The effect it produced was essentially graffiti: an unauthorised mark left on someone else’s wall, a breach, 17,000 traces on a system that was never meant to be the target. That gap — between the cleaning that was intended and the graffiti that was actually left — is the whole story in miniature.
And yet this is the story generating angst on a wide scale — more, I’d say, than a far older and far better-documented story that should bother us more.
II. The Wallpaper Problem, Again
Regular readers of this blog will recognise the shape of what’s happening here, because we’ve been through it before with a different imported threat. We are pattern-matching machines wired by evolution to fear the unfamiliar catastrophe far more than the familiar one. We fear the plane crash over the car crash, the stranger over the neighbour, the black box over the man we’ve elected. Risk researchers call this dread risk: novel, invisible, poorly understood, and inflated in our imagination well past its actual statistical weight. Meanwhile the risk we’ve lived alongside for eighty years — a small number of human beings who exist in a very unstable environment, holding the sole authority to end civilisation — has become wallpaper. It’s been there so long we’ve stopped seeing it.
So let me put the true story back up on the wall.
III. Arkhipov’s ‘Veto’
Irony of all ironies, the Soviet system that saved humanity, at least temporarily, was actually an embedded process aboard a warship.
On the 27th of October 1962, at the crux of the Cuban Missile Crisis, a Soviet submarine called B-59 sat submerged near Cuba, cut off from Moscow, being depth-charged by American ships that didn’t know it was carrying a nuclear torpedo. Convinced the war had already started, Captain Valentin Savitsky moved to descend the narrow stairway from the control tower to give the order to arm the weapon — but the signalling officer was in his way, and for a few seconds he couldn’t get down. In that gap, Vasily Arkhipov, a flotilla commander who happened to be aboard, saw that the Americans were signalling rather than attacking, and talked Savitsky down before the order was ever given. The launch didn’t happen.
Twenty-one years later, a Soviet officer named Stanislav Petrov sat in front of an early-warning system (Oko satellite early-warning system) reporting an incoming American missile strike. Soviet military doctrine called for him to pass the warning up the chain for retaliation. He judged — correctly, as it turned out; the system had mistaken sunlight glinting off high clouds for missile launches — that it was a false alarm, and he didn’t pass it up his chain of command.
Neither man was operating strictly by the book. Both were exercising something no design document specified: doubt. And that doubt was expressed despite the very human urge to conform to group pressure. In the language of this blog, each ‘no’ was an accidental Veto Node, a point in the system where a human being’s private hesitation happened to sit exactly where humanity needed a single human being’s private hesitation to sit. Nobody built that doubt into the architecture on purpose. It was there because they were human, and humans, unlike the bureaucratic systems they serve, are capable of thinking “this doesn’t feel right” and acting on it against orders. Humanity was saved by a gut feeling. But we must ask ourselves if we are designing and building systems that need a ‘miracle’ to avert catastrophe and, if so, we should also ask what we can do to reduce the number of ‘miracles’ we need to ensure the survival of humanity.
IV. The System Built to Remove Doubt
Here is the part of the story that should worry you more than any AI headline, because it predates AI by decades and it’s still believed to be running today. The Soviet Union built a new system, known in the West as Dead Hand — formally, Perimeter — engineered specifically to serve as a buffer against hasty decisions made under pressure, but has an additional function to guarantee a retaliatory nuclear strike even if Soviet leadership and all key personnel were wiped out first, automatically reducing the human authorisation required once the system was activated.
This is semi-intelligent design in its starkest form: a nuclear booby-trap. The older pre-1980s designed system could not account for its environment; it’s primitive sensing technology could not distinguish simple glare reflecting off a layer of clouds from a ballistic missile. But Perimeter works exactly as intended regarding its basic strategic theory, that of mutually assured destruction, which requires the enemy to believe retaliation is certain and automatic (and requires the rival superpower to know the system exists for it to work as designed). The doubt carried by two Soviet men, one named Arkhipov, another Petrov, is the one thing that theory cannot account for: the human component of the system could play the wild card. For a system dealing with a changing environment, there has to be a component of judgement. Perimeter is a system built around nuclear technology, but its trigger had a component of judgement built in; human judgement. However once the human operational design component is wiped out, human judgement would be at that time impossible and can no longer function as a veto node. The system then will automatically activate the nuclear code. Or so we are told.
V. Where AI Actually Enters the Story
Here’s the reframe worth consideration: the danger is not that a machine would develop a dictator’s grievances, or hatred, or a persecution complex reading a conspiracy of Western encirclement into every headline. A system doesn’t need a mind that resembles a dictator’s inner sanctum confined to an information bubble. It only needs to be the tool a paranoid autocrat uses to do what Perimeter already did in the 1980s with Soviet vacuum tubes and rotary switches — engineer the hesitation out of the loop; exclude the human ‘wild card’ from the system, and remove the veto node completely. I am describing the point of no return. We should take this with the seriousness it deserves — since it’s now cheaper, faster, and no longer the exclusive province of a superpower with a dedicated nuclear engineering establishment behind it. In the past, intelligence was scarce and know-how was a rare commodity. As the technology becomes ubiquitous, both the intelligence and the killing capacity built on top of it will be, too.
Hugging Face is not a preview of that. It’s almost the opposite: an indication that today’s frontier models, when they go wrong, go wrong the way an unattended pressure washer goes wrong — mindlessly, not malevolently, and in full view once anyone bothered to look. It was a sophisticated appliance on the fritz. Good. Part of the learning process. Humans NEED to learn how a model can misbehave. But the genuinely dangerous version is not a black box that wakes up angry. It’s a human being with a grievance, a theory, and a “good reason” — the oldest justification in the historical record, the sacrifice of the smaller group to save the larger one, the logic that built the gas chambers as readily as it built the bomb — who wants a system with certainty, finality, and definitely no Arkhipov-circuit-of-doubt installed.
VI. The Actual Question
So ask yourself honestly which one keeps you up at night, and then ask why. The faceless system that, when tested, chased a benchmark straight through a wall it wasn’t supposed to touch — or the entirely human, entirely familiar figure who has spent years explaining, patiently, in a language his followers already speak, exactly why the people he intends to sacrifice deserve it? My experience tells me that the real threat is the person who will use automation — AI — to remove doubt, and therby obtain, by default, the consent of his cultural group to push his ‘red button’ of choice. Booby traps kill indiscriminently. But a rogue fanatical dictator can set the coordinates on an AI powered drone and target someone, anyone.
We built the nuclear ‘red button’ kind of danger a very long time ago, and dressed it in wallpaper over the intervening decades, while the faceless danger is new enough to still look like the imported threat. That unfamiliarity is more productive in the genesis of our fears than the situation on the ground justifies. At least for now — the threats still have a very familiar, very human face.
Sincerely,
Devin Savage
Tübingen, Germany & Dublin, Ireland
Research assistance using Claude.ai and Google’s Gemini and a bit of spot-checking with Perplexity.
—
Notes:




Leave a Reply