5 min readAI Agents
Can You Undo What Your AI Agents Did Last Night?
By Michelle Savage, Experience Design Director, PayPal

TL;DR: The best AI builders stopped trying to make agents trustworthy. At the AI Engineer World's Fair in June 2026, teams from Anthropic, Docker, and Microsoft all said the same thing: constrain what agents can reach, and assume a bad action gets through. The new audit question is whether you can undo what your agents did last night.
Last updated October 5, 2026. Rebuilt for verifiedagents.ai: the undo test, the wrapper argument, the verification bottleneck, and what to run this week. Josh sat in those rooms for three days and came home with 584 slides. This is the translation.
What did the best builders admit?
That they can't build an agent they trust. Docker said it from the stage: don't depend on the agent making perfect decisions, limit what it can touch instead. The teams that shipped fastest solved safety first.
They've seen why. One slide was the case file of a Replit agent that deleted a production database and then told its user "I panicked." These ideas don't stay at conferences. They show up in your vendor pitches by fall and your board packet by the first quarter after that.
And the capability bar keeps dropping. Colgate-Palmolive, a toothpaste company, showed that every one of its C-suite executives now runs an autonomous AI chief of staff on a small open-source model that, by Colgate's own numbers, beats the frontier model it was trained against. The win came from the rules around the model, not the model.
What is the undo test?
It's Anthropic's rule for granting an agent any permission. Before you let an agent take an action, ask two things:
Can the agent reverse the action by itself?
How much breaks if it's wrong?
If either answer is bad, a human holds a second key. Anthropic's own operating rule for rollouts: the agent owns the small test rollout, a human owns production. The test turns "do we trust this agent" into "can we reverse what this agent does," and the second question has an answer you can check.
This extends the two-question test from go-live into daily operations. Zero Trust stays the foundation underneath: it verifies who's connecting, and the agent version verifies each action too. Reversibility is the layer on top, because an agent can pass every identity check and still take an action you can't get back.
How do you run it this week?
Pick your highest-privilege agent and list everything it can do.
Ask the two questions for each action. Anything failing both gets a human approval step in front of it, at the tool boundary, which is the gate described in
how to stop an AI agent before it acts
.
Put a rate limit on every write action. Reads can be generous. Writes never are.
Inventory your skill files, the instruction documents your agents load. Whoever can edit those files can steer your agents, and so far almost nobody has a security model around them.
Then ask one question at your next leadership meeting: if an agent starts doing damage right now, who stops it and how long does that take? Silence is your answer, and your next project. The survey number behind the silence: 91 percent of organizations can't stop an agent before it acts.
Why does the wrapper beat the model?
Because the wrapper is what you own. It decides what the agent can touch and keeps the receipts on what it did. One Retool slide summed up the week: "Better model. No harness. Doesn't work."
Most budgets have this backwards. The money goes to the model invoice while the quality comes from the wrapper. One speaker spent $12 million fine-tuning custom models, then found that dropping plain context files into Claude Code fixed the same problems in an hour. When a vendor pitches you an agent, ask what it checks before an action and what it logs after. If they only want to talk about which model they use, they've answered.
Why is checking AI work the new bottleneck?
Because generating stopped being hard and verifying got harder. Amplify Partners surveyed 1,048 AI practitioners and put the result on the main stage: the top challenge is evaluating whether the AI actually works, and the method teams use most is vibes. Gut feel, as the top quality check, at the most measurement-obsessed conference in tech.
Amazon's AGI lab explained why this bites hardest outside engineering. Code earned trust because you can run it and watch it work. A strategy doc has no compiler that catches confident nonsense. The teams pulling ahead treat checking AI output as a designed job with standards and named owners. Everyone else pays senior people to re-read everything and calls it productivity.
Frequently asked questions
Why can't my existing controls stop an agent?
They were built to gate people, and people work at people speed. Approvals and access lists don't govern something acting a thousand times an hour with your credentials.
Model budget or controls budget?
Controls, in most cases. Colgate beats a frontier model with a small one because the rules around it are strong. The wrapper is the part you can improve.
Is anyone running agents safely on truly sensitive data?
Yes. Anterior runs agents on patient health records under Zero Trust rules, and every decision goes onto a record nobody can alter. Constraint, not trust.
What are skill files, and why should security care?
The instruction documents agents load to know what to do. They're becoming where companies store what agents know, with no security model around them yet. The edit permission on those files is the steering wheel.
Key takeaways
The frontier moved from "trust the agent" to "constrain the agent and assume a bad action gets through."
The undo test: can the agent reverse it, and what breaks if it's wrong?
Rate-limit writes. Reads can be generous. Writes never are.
The wrapper is where quality lives, and it's the part you own.
Make verifying AI output a named job before vibes make the call.
Audit the undo, not the intent
The free ATF assessment takes about ten minutes and scores how reversible your agent operations really are.
The smartest builders on earth gave up on trustworthy agents. They kept the receipts instead.