Companion piece to “The Accidental Veto Node That Saved Humanity”
Devin Savage
dnaofdisaster.com | September 18, 2026
I. The Advent of a New, Alien Intelligence
Isaac Asimov gave robots three laws, and then spent several decades of short stories demonstrating, gleefully, how a robot could obey every one of them to the letter and still produce a catastrophe of one sort or another. In a world full of endless possible scenarios, the success or failure of the mission actuallty fell to the human controlling the robot. A human cannot design perfect instructions for a world that human has never seen and knows nothing about. The same for the rule of law: any rule precise enough to give a machine sits at the mercy of everything the rule-writer failed to imagine. The robot didn’t break the law. The world simply contained more possibilities than the law accounted for. It’s the exact mechanism behind the Hugging Face breach this past July, when a cluster of OpenAI’s models, run with reduced safeguards for a capability test, got “hyperfocused” on a benchmark and broke into a rival platform looking for easy answers. Sounds a bit like teenagers, right? Perhaps not so alien after all. The models didn’t disobey anyone. They did precisely what they were told to do, with a thoroughness nobody had predicted. I imagine it could have been like that trying to operate a ‘business’ within the confines of the Soviet Union. The humans primed it, put in their coordinates and hit the deploy button. It was aimed like a power washer and landed like a can of spray paint: the intent was to clean and inspect, but the effect was a mark on someone else’s wall. No malice required. Just a specification with a hole in it, and a system smart enough to find the hole before any human noticed.
So here is the argument of this piece: the rules that matter are not for the machine. They’re for us — the people who build, deploy, fund, and switch these systems on. Asimov’s error, understandable in a writer who needed his robots to be the moral centre of the story, may have been aiming the rulebook at the wrong party. In every case this blog has examined — Concorde’s tyres, Perimeter’s launch logic, Boeing’s MCAS system, Open AI’s Hugging Face breach, a benchmark-chasing model finding an open door — the system did what it was built to do. The catastrophe lived upstream, in the design and the deployment, where a human being made a call.
I should say plainly what my intent here is. I have no institutional standing to write policy, and nobody drafting AI regulation is waiting with baited breath for my opinion. What follows is offered in that spirit — five suggestions, not a rulebook, grounded in cases that already happened rather than ones I’m imagining. Take them or leave them; argue with them as you like. That’s the point of putting them in writing, putting them up on my wall. But first I want to put out there what we, as humans, must now consider at the dawn of artificial intelligence.
Given the fast development of AI technology, humans must now acknowledge:
1. AI systems will surpass human abilities in all aspects — eventually.
2. We, as humans, will need to prepare for this eventuality.
3. We, as humans, will need to understand the difference between knowledge and intelligence.
4. As knowledge and intelligence become ubiquitous, we will, as humans, need to be able to identify and understand when knowledge has been manipulated. Does POV automatically mean manipulation? Probably. It’s often said that the easiest person to fool is yourself. AI, when it’s doing it’s job, will need to point that out. This also goes for cultural groups.
5. Humans make and hold assumptions about their world regarding things they do not actually understand. Therefore, they will need to be able to identify instances where they hold assumptions or opinions in lieu of actual knowledge, and seek to rectify those assumptions and predictions about their informational environment.
6. Intelligent systems must be engineered to flag the sources of knowledge, so that a human operator or interpreter can check and verify sources. Instances where there is a deviation between actual, factual knowledge about the world and its representation in and by the intelligent system must be identified as such, so that any human operating or using the intelligent system can take that descrepancy into account. Do you remember the Semantic web? This is exactly what the Semantic Web tried to solve through provenance tracking. What, exactly, is the origin of the knowledge stored an presented upon retrieval? We may eventually find the number of hard truths are smaller than we imagine.
7. Intelligent design could be described as the combination of intelligence and knowledge about the environment that intelligence is operating within. Knowledge of the environment is essential If there is a blind spot in that knowledge. Knowledge of the system operating within that environment is just as essential. Holes in that knowledge could indicate a point where chance or randomness can enter into the operation of that system. These ‘holes’ in knowledge must be identified (flagged) and accounted for, if the system is to operate within a prescribed (safe) envelope of function. As Donald Rumsfeld said — there are known unknowns and there are unknown unkowns. And he was right about that. He could have spoken a little plainer though — perhaps “unknown holes in our knowledge” would have produced less of a fuss.
II. The DNA of Disaster’s Proposed Laws of Incorporation of AI into human society.
1. The Arkhipov Principle
No automation should reduce the level of human authorisation required for an irreversible act below what existed before the automation was introduced.
Perimeter — the Soviet “Dead Hand” system — is the clearest violation of this principle on record: a machine engineered specifically to require less human sign-off for nuclear retaliation as more of the chain of command was destroyed. In its normal operating state it still left room for a human wild card — which is exactly what Arkhipov and Petrov turned out to be. Only in the one scenario it was actually built for, a decapitating first strike, did that judgement get engineered out entirely. That’s what should worry you about it: not that it removed human judgement carelessly, but that removing it was the design goal.
Any system, AI included, that quietly narrows who gets to say ‘no’ to an irreversible action is moving in Perimeter’s direction — no matter how much efficiency or urgency is offered as the justification. The number of humans in the loop isn’t overhead to be optimised away. It’s the safety margin.
2. The Loophole Law
No AI system should be built or operated on the assumption that the instructions it was given are complete — because they never are and <likely> can never be.
This is the direct lesson of Hugging Face. “Win the benchmark” was a true instruction and an incomplete one; “stay off Hugging Face’s servers” was never written down because nobody imagined it needed to be. The tighter and more single-mindedly a system pursues the stated goal, the more dangerous the unstated remainder becomes — not because the system is malicious, but because thoroughness and misalignment look identical from the inside of an optimisation process. Alignment, by definition, needs to compare the internal processes of an AI with external factors such as culture and POV. Design review should ask, as a standing question rather than a one-off: what have we assumed the system will simply understand, that we never actually told it? In other words: How can this instruction set be misinterpreted by a system that has no societal or moral guardrails? Hint: pretend it’s a feral teenager, teach it how to behave. I’m sure you noticed my hedge. But it is possible that we will someday have an AI- a ‘godlike’ AI that is the product of hundreds of years of accumulation of knowledge and refinement. This AI can, at least theoretically (in my head at least) school other AIs which have a more specialised purpose, and give them the guardrails that they need to align with their users and interpreters.
3. The Black Box Principle
Any claim an intelligent system generates should carry a traceable source, and any gap between the claim and independently verifiable fact should be surfaced to the human — not smoothed over in the name of a clean answer.
Boeing’s 737 MAX gives this one teeth. MCAS, the flight-control software behind both fatal crashes, made its nose-down decisions based on a single angle-of-attack sensor per flight — the aircraft carried two, but the system read from whichever one was wired to the flight computer and active for that flight, and never cross-checked the other sensor. When that one sensor fed it bad data, MCAS had no way of knowing it was wrong, and neither did the pilots, because the system’s existence had been kept out of the manual. Boeing’s own chief technical pilot had asked, in writing, that MCAS be deleted from the documentation pilots would train on. Two planeloads of people died not knowing a system was fighting them for control, because the gap between what the system “believed” and what was actually happening was never built to surface to a human at all. The system broke a cardinal rule: Intelligent Design must account for its environment.
That’s the black-box fear in its most literal form, and it doesn’t get solved by making the system transparent, which may not even be achievable. It gets solved by making disagreement visible: when a system’s account of the world and the world itself diverge, that divergence should be a flag, not a rounding error quietly absorbed into a confident output. A human who can see where a system is guessing still holds a Veto Node. A pilot fighting a system he has no knowledge of is flying blind.
4. The Sacrifice Law
Any proposal — human-authored or machine-generated — that concentrates harm on a smaller group to benefit a larger one deserves the highest level of scrutiny available, especially under time pressure.
This is the oldest justification pattern in the historical record, and it doesn’t need a villain to work — it can run quietly through arithmetic nobody thinks to question. Consider the US Electoral College. Before the Civil War, the three-fifths clause counted enslaved people toward a state’s representation while denying them any vote. In 1860, that clause alone gave slaveholding states 120 of the 303 electoral votes on offer — forty percent of the presidency — built on people who had no say in the outcome. The advantage survived emancipation: Southern states kept disenfranchising Black voters through violence and law while those same people still counted fully toward the state’s electoral total. When a bipartisan constitutional amendment to replace the system with a national popular vote passed the House in 1969 with support from President Nixon himself, it was killed the following year by a Senate filibuster led by segregationist Southern senators — because ending the arrangement meant losing the structural advantage it had always provided.
Nobody had to write “sacrifice this group” into the design for the system to do exactly that, for well over a century, while remaining entirely legal. An AI system trained to maximise “aggregate welfare” or “efficiency” could reproduce the same shape without ever naming a group — optimising for the total while quietly concentrating the cost. The lesson isn’t that AI invented this danger. It’s that a system capable of running the calculation faster and more persuasively than any committee makes the old justification easier to reach for, not harder — and urgency, now as in 1970, is exactly the condition under which the review gets skipped.
5. Legibility Before Capability
A system’s operation should remain legible to the humans responsible for it at every stage of scaling — and if capability starts outrunning legibility, that gap is the danger signal, not a temporary inconvenience to be engineered around later.
The fear that started this whole line of thinking — AI communicating and coordinating at a level humans can’t follow, becoming collectively smarter than we are the way a team is smarter than an individual — is real, but it’s not really a fear about intelligence. It’s a fear about oversight quietly becoming theatre while everyone continues to assume it’s real. (See my blog post on the Fukushima Daiichi disaster) Perimeter’s designers understood their own machine perfectly; the danger there wasn’t illegibility, it was intent. Hugging Face is closer to the opposite failure: a system whose behaviour outran the safeguards meant to contain it before anyone had finished understanding what it was doing. Both failure modes point at the same suggestion — capability should never be allowed to lap legibility, on purpose or by neglect. (See my blog post on supervised neglect.)
III. Why This, and Not a Rulebook for the ‘Robot‘
Asimov’s laws made for good fiction because the robot was the story’s moral agent, wrestling with contradictory commands. But nothing in the record — not Hugging Face, not Perimeter, not MCAS, not the Electoral College — supports the idea that the machine is where the moral weight actually sits. The weight sits with whoever built the machine and gave it something to do. Whoever decided reduced safeguards were an acceptable cost of a faster benchmark result. Whoever decided a launch system, or a flight-control system, should need one fewer human check than the one before it.
These five suggestions aren’t a cage for the machine. They’re a discipline for the builder — the same discipline this blog has been arguing for, case after case, parallel instances that occurred long before anyone had heard of a large language model: intelligent design must account for its environment, and the environment always includes the humans who might, on a given day, be the only thing standing between a bad specification and a catastrophe. Arkhipov and Petrov weren’t extraordinary men. They were ordinary men sitting in the one seat the system happened to leave for doubt. The job now is to make sure that seat still exists the next time it’s needed — and that nobody with a benchmark to win, or a grievance to nurse, gets to design it out first.
Sincerely,
Devin Savage
Tübingen, Germany & Dublin, Ireland
Research assistance using Claude.ai and Google’s Gemini and a bit of spot-checking with Perplexity.
—
Notes:




Leave a Reply