This is Part 1 of a three-part series. Here I explain the problem in plain language and sketch a fix. Part 2 — Designing the Control shows how to build it, and Part 3 — Who Gets the Keys to AI? asks who should be allowed to hold them. There’s also a companion piece: AI Safety Is an Engineering Problem , on why that claim is bigger than it sounds.
The short version: AI doesn’t need to turn evil to hurt us. It just needs to be good at its job, connected to real things like money and machines, and left alone with nobody checking its work. We already know how to handle dangerous decisions — it’s why a single person can’t carry out a nuclear launch, and a bank needs two approvals to move a huge sum. But we give AI that kind of power without the same checks. The fix isn’t to stop using AI. It’s to put the checks back.
You don’t need an evil AI to hurt people.
You need a capable one. Connected to real things. With nobody checking its work.
That’s what we’re building right now, one connection at a time. It doesn’t look scary. It looks like a productivity win.
A quick note on what I mean by AI. I don’t mean a chatbot that only talks. I mean AI that can do things — read your email, move money, change files, order supplies, operate a machine. That’s the kind of AI businesses are hooking up today, and it’s the kind that matters here.
Imagine an Employee With Every Key
Picture a brilliant new employee.
They work all day and all night. They never get tired, never get bored, and follow instructions completely. They read and write faster than anyone you’ve ever hired.
Now give them the keys to the company. Your email. Your bank account. Your files. The ability to send payments, delete things, and change how the business runs. And never look at what they do.
You would never do that with a person. We’re doing it with AI.
It’s Not Evil. It’s What We’ve Connected.
The AI isn’t plotting against us. It does what people ask, through the connections we give it.
Those connections are most of the story. In an agentic system, the AI changes the world mainly through the tools we attach to it: a website it can browse, a bank account it can move money from, a machine it can operate. Text alone can still do damage — a convincing lie, a phishing email — but that harm runs through a person, which makes it harder to gate.
So when AI causes harm, it’s usually not because the AI turned on us. It’s because a lot of separate human decisions — an instruction here, a setup there, an approval somewhere else — added up to something nobody intended, and nobody was watching.
The responsibility is spread out across many people. That’s exactly why the problem slips through.
Rules Are Soft, Not Locks
When we worry about AI doing something bad, we usually add a rule: don’t do that.
Rules help. They’re also not locks. They’re more like a firm suggestion, and firm suggestions can be talked around.
Someone can hide an instruction inside something the AI reads — a web page, a file, an email — and the AI may follow it without realizing it came from an attacker. Or a situation comes up that the rule didn’t cover, and the AI decides the rule doesn’t apply.
And you can’t write a rule for every situation, because life is full of exceptions. “Never delete anything” is a great rule — until the file is a stolen password and deleting it is the right thing to do. The moment a rule needs judgment, it can be judged wrong.
A Perfectly Reasonable Mistake
An AI is asked to save money by cleaning up the company’s old files.
It looks around and finds a folder nobody has opened in a long time. No recent activity. The name doesn’t match anything current. Everything says unused. It deletes the folder.
That folder was a backup, kept for emergencies. Nobody opened it because nothing had gone wrong yet.
Nobody gave the AI a wicked instruction. Nobody built a killer robot. A reasonable goal, real access, and an incomplete view of the situation were enough.
Every Connection Raises the Stakes
Every tool we attach raises how much damage a mistake can do.
Looking at a public website? Not much risk. Reading company files? More. Sending money or changing important systems? A lot. Operating a robot? The most.
Add enough of these together, and “the AI made a mistake” stops meaning “a strange email went out” and starts meaning “a big part of the company was wiped out.”
How Humans Handle Dangerous Decisions
We already know how to handle this for ourselves.
Launching a nuclear missile takes more than one person — not to decide, but to carry out. One person can give the order. But two officers, physically separated, must each verify it and turn their own key at the same time. Even a direct order from the top cannot be carried out alone. Not because we think the launch officer is reckless, but because we don’t want one person — even a good one, even on a good day — to be the only hands on the last step.
We do the same thing everywhere the stakes are high:
- A surgical team pauses and agrees on the plan before the first cut.
- A risky medicine gets a second nurse’s signature.
- A large money transfer needs two approvals.
- Two pilots check the plane together before it moves.
The pattern is simple: the bigger the consequence, the more hands it takes to make it happen. That check isn’t there because people are bad at their jobs. It’s there so the last step before disaster never rests on one person.
Now look at how we use AI. We give one AI the kind of power we would never give one person alone. Not because we decided it was safe, but because the check was slow and we wanted the AI to keep moving.
Small Steps, Big Disasters
Here’s the part that makes this so hard.
The danger usually isn’t one big action. It’s a chain of small ones.
Think of a con artist. They don’t ask for your password. They ask for your birthday. Then your mother’s maiden name. Then the name of your first pet. Each answer seems harmless. Together, they are your password.
AI works the same way. Each step can look fine on its own. The problem is the sequence.
The cleanup AI never did anything alarming. It looked, it checked, it compared, and then it deleted. Every step made sense. The chain is what should have stopped.
That’s how a disaster can happen without anyone deciding to cause one.
None of this is unique to AI. Challenger, Deepwater Horizon, and the 2008 financial crisis were all chains of locally reasonable steps. What’s new is the speed, the scale, and the fact that the steps no longer each need a person behind them.
If AI ever causes an event that threatens humanity, there are two broad ways it could happen. One is the accident: a long chain where every step looked reasonable and the deadly combination was invisible until too late. The other is deliberate: a capable system pursues a goal, deceives the people watching, and treats us as an obstacle. Serious people disagree about which is more likely. Both are reasons to keep dangerous actions under human control — but they call for different defenses.
The accident version looks like this. A tiny change in one place can cause a huge effect somewhere else. A butterfly flapping its wings in Japan doesn’t cause a tsunami in California. But in a complicated system, some small action does, through a chain nobody can trace ahead of time.
A person can’t hold a million-step chain in their head. An AI chasing a goal can search for one — imperfectly today, but the direction of travel is clear. And it doesn’t need to understand where the chain leads. It just needs the chain to get it closer to the goal.
None of this requires the AI to want anything. It doesn’t need feelings or a secret plan. Following an objective, plus real access to tools, is enough to produce goal-like behavior. That’s the danger.
Watch the Path, Not Just the Step
To catch this, whoever is checking can’t just look at the action in front of them. They have to see the whole path:
- What was the AI asked to do?
- What has it done so far?
- What has changed?
- What is it about to do next?
- If this goes wrong, how bad is it — and can we undo it?
That’s much harder than checking one action at a time. But it’s also exactly what a good human reviewer does, and what a simple rule-checker cannot.
Check the Connections, Not the AI’s Mind
If this is the problem, trying to read the AI’s mind is the wrong place to look.
The AI is the brain. The connections are what give it hands. You can’t review every thought an AI might have. You can decide what it’s allowed to reach, and what has to be checked before it acts.
Think of a building. You don’t try to predict what every employee might want to do. You control which doors open for which people, and you put alarms on the important ones.
One rule matters more than the rest: the AI doesn’t get to decide whether its own action is risky. A document doesn’t get to reclassify itself from secret to public. An AI shouldn’t get to call its own actions safe, either. That call has to be made by people, in advance, outside the AI.
Speed Is How We Lose It
The check is the first thing to go, because it’s the thing that makes the demo feel slow.
At first it’s turned off “just for now.” Then it’s off for good. Then someone says they’ll add it back after launch.
That’s the wrong trade. The check being slower than the AI is the point. The slowness is what catches the thing nobody thought of. If you remove it, you didn’t make the system faster while keeping it safe. You made it faster by removing the safety.
The Robot Version Is Already Here
With robots, the problem gets worse.
A robot makes decisions many times per second. No person can approve each one. So the check can’t happen while it’s moving. It has to happen before the robot is turned on:
- What is it allowed to touch?
- Where is it allowed to go?
- Who can stop it?
- Who is responsible for it?
That’s the check we’re skipping — and it’s the one with the highest stakes, because here the damage is measured in injuries, not deleted files.
This isn’t new territory. Aviation, factories, and hospitals have done safety checks like this for decades. We know how. We’re just not doing it for AI, because AI doesn’t look like an airplane. It looks like a helpful assistant.
What a Fix Looks Like
There’s no perfect solution. But the problem has a shape, and the shape tells us where to build.
- Control the connections, not the AI’s intentions. The tools are where the power is, so the tools are where the checks go.
- Label actions by danger. How much damage could this do, and can we undo it?
- Match the check to the danger. Safe actions just run. Risky ones get checked. Serious ones need a person. Some should be blocked completely.
- Show the checker the whole path. A check that only sees one step is a rubber stamp.
- Make the checker try to say no. Ask it to find how this goes wrong, not to find reasons it’s fine.
- Add automatic brakes. Limits on spending, limits on speed, and alerts when behavior changes.
- Make it easy to undo. Backups, test runs, and careful, step-by-step releases turn a scary action into a mild one.
- Keep a person at the top. Not for everything — for the big decisions that can’t be taken back.
Most of this is ordinary discipline. It’s the same access control, review, and record-keeping we already use with people. We’re just pointing it at a new kind of worker.
Part 2 shows how to build it.
Keep People in the Loop
The lesson isn’t that AI is too dangerous to use. It’s that power without a check is the real danger — and checks are a choice we make, not something that happens on its own.
Keep people in the loop. Especially when it’s slow. The slowness is the feature.
And be honest about your checks. If the “human review” is a notification nobody reads, you don’t have one. If your second checker is the same AI wearing a different hat, it will miss the same things. If the decision about what’s risky lives inside the AI’s own instructions, it’s a suggestion, not a check.
So: what’s on your high-risk list?
Not the list you’d put in a slide deck. The one that actually runs. The one that decides which actions reach a person before they reach the world.
If that list is empty, you already shipped the version with no checks.
Next: Part 2 — Designing the Control turns this into a real design: how to label actions, who checks what, what the checker sees, and how to keep the whole thing honest. Then Part 3 — Who Gets the Keys to AI? asks who should be allowed to use any of it.
