5 min readAI Agents
How Do You Stop an AI Agent Before It Acts?
By Michelle Savage, Experience Design Director, PayPal

TL;DR: You stop an AI agent before it acts with a gate outside the model, at the tool boundary. Let the agent reason and draft freely. Block the high-impact, hard-to-undo actions until policy and a person say yes. A prompt that says "never do X" isn't a control, because the rule never sits on the path the action takes.
Last updated October 5, 2026. Rebuilt for verifiedagents.ai: why prompts can't stop agents, where the gate goes, what it checks, and the ladder that decides what waits for a human.
Why can't a prompt stop an agent?
Because prompt rules live inside the model, and actions leave through tools. I've watched teams paste long "never do X" lists into a system prompt and call it governance. The agent could still send email. It could still buy things. The model can agree with your rule and fire the tool anyway, because the rule never sat on the path the action takes.
I write rules for a living, so hear me on this: words only govern when something enforces them. Security already knows the pattern. You don't stop a bad API call by hoping the app remembers policy. You put a control on the request. Agents need the same move. The control lives at the tool boundary, not in the chat history.
Where does the gate go?
In front of the write, the send, the purchase, and the permission change. An action gateway sits at the boundary and does one of two things with each call: deny it, or queue it for a human. The agent keeps thinking and drafting. The gate decides what leaves the sandbox.
And the gate fails closed. If the approval service is down, the agent doesn't get a free pass. It waits, or the call fails. Silent success while the gate is broken is how "we thought we had approvals" ends up in an incident review.
Should a human approve every action?
No, and this is the part teams get backwards. If a person must click yes on every tool call, one of two things happens: the process dies, or people approve without reading. Neither is control.
Sort actions by two questions instead: how big is the impact, and how hard is it to undo?
Impact and reversibility | What the agent does | Example |
Low impact, reversible | Act and log | Drafting, sorting, internal lookups |
Low impact, hard to undo | Act with a delay window | Archiving records |
High impact, reversible | Act and alert | Updating a staging config |
High impact, hard to undo | Propose and wait | Money moving, customer messages leaving, permission changes, production writes |
People stay on the decisions that count. They come off the rubber-stamp clicks.
What belongs in front of an external action?
A boring checklist, and boring is what leadership can fund.
A verified agent identity with a named human owner.
An allowlist of tools and parameters, scoped to least privilege.
A risk class tagged on the action, with human approval required when the class is high.
Fail closed when the gate is down.
In ATF terms: Identity Management says who the agent is and Segmentation says how far it reaches, while Incident Response says how fast you revoke it. The permission split behind all of this is in access control versus action control.
What does a gate look like when it works?
There's a scene in Josh's book about Kevin's procurement agent. A bulk discount looked, to the agent, like permission to buy. It was ready to commit $1.4 million of floor cleaner. A confirmation gate held the spend until a person looked.
That save wasn't a smarter prompt. It was a gate that still asked before money moved. Without it, "we could have stopped it" becomes a line in the postmortem instead.
What about the day the gate fails?
That's what the kill switch is for, and it has a design test: pulling the switch stops one agent, not the business. You revoke that agent's identity and cut its tool rights in seconds. Peer services keep running. If your only "stop" is emailing a vendor or turning off a shared cloud key, you have a fire drill, not a switch. The full build spec is in how to build an AI agent kill switch.
Frequently asked questions
Is turning off the LLM vendor account a kill switch?
No. That stops many systems at once and may not revoke one agent's local credentials or tool tokens. A real switch targets one agent and leaves the rest of the environment running.
Who should be able to pull the switch?
Someone other than the agent's builder, on call, with a named backup. If only the builder can stop it, you don't have Incident Response. You have hope.
Do chat-only copilots need a gate?
Lower priority, as long as they truly can't write, send, pay, or change systems. The moment one can act outside the chat, treat it as an agent with tools.
What's the fastest test of whether our approvals are real?
Pick one live agent with a high-impact tool. Ask who can deny that tool call today, and time how long revoking the agent's identity takes. If the answer involves a ticket queue, the approval story isn't real yet.
Key takeaways
Prompts are wishes. Gates are controls. Enforcement lives where the action happens.
Sort by impact and reversibility: act-and-log for the small stuff, propose-and-wait for the irreversible.
Fail closed. A broken gate never means a free pass.
The kill switch test: stop one agent without stopping the business.
The $1.4 million floor cleaner never shipped because a gate asked first.
Score your stop story
The free ATF assessment covers Incident Response alongside the other four elements, in about ten minutes.
Your agent will draft something expensive eventually. The gate decides whether that's a story you tell or a loss you report.