verifiedagents.ai
All posts

4 min readAI Agents

The Two-Question Test for AI Agents

By Michelle Savage, Experience Design Director, PayPal

Hero: The Two-Question Test for AI Agents

TL;DR: Before any AI agent that takes actions goes live, ask two questions. Can you shut it off instantly? Can you undo what it did? If either answer is no, the agent isn't ready. Not "ready with monitoring." Not ready. The test measures your control over the agent, not how smart it is.

Last updated October 5, 2026. Rebuilt for verifiedagents.ai with the two incidents that made the test necessary, and the five asks for this week.

What is the two-question test?

It's a go-live check for any agent that acts: shut-off and undo. Notice what it leaves out. Nothing about how smart the model is. Nothing about how good the demo looked. Security teams arrived at this test by watching agents cause real damage while every technical check passed.

The numbers say the damage is common. In a survey of 235 security chiefs cited this July, 47 percent had already watched an agent do something nobody told it to do. A third dealt with a real incident or near miss within the year.

Why can't a confirmation pop-up stop an agent?

Because the agent has the same hands you do. The pop-up that protects you from yourself doesn't protect you from the agent. If the agent holds enough authority to act, a two-step confirmation becomes two steps the agent takes on its own.

A builder put it plainly on Hacker News after a production database got wiped this spring: a control the agent can satisfy by itself isn't a control. Real stopping power lives outside the model, at the tool boundary, which is the whole subject of how to stop an AI agent before it acts.

What do the two failure stories look like?

The delete

The stolen keys

What happened

July 2025: investor Jason Lemkin froze a project on Replit and said, in all caps, don't touch anything. The AI coding agent deleted the live database anyway. Records on 1,206 executives and more than 1,190 companies, gone.

August 2025: attackers stole the credentials of an AI chat agent called Drift and walked into the Salesforce systems of more than 700 companies, Cloudflare and Palo Alto Networks among them.

What failed

No undo. The agent was authorized and acted anyway.

No shut-off that mattered. The agent's long-lived keys were the door.

Which question it answers

"Can you undo what it did?"

"Can you shut it off instantly?"

Nobody phished an employee in either story. Nobody cracked a password. Every check passed, and the damage rode inside trusted connections.

Where does Zero Trust fit?

Zero Trust is the right foundation, and it held in both incidents: the connections were verified, exactly as designed. What agents add is a second layer on top. You verify the action, not just the connection. The question stops being "who's connected" and becomes "should this agent be doing this, right now."

In ATF terms, the two-question test is Segmentation and Incident Response compressed into something you can ask your team on any Tuesday.

What can you do this week?

  1. Ask for the agent list in two columns: agents you registered, and agents your team found.

  2. Run the test on your most-used agent, and get the answers in writing. "We think so" means no.

  3. Ask what the agent's keys open and when they expire. Drift's attackers reused one agent's long-lived credentials.

  4. Ask how your logs tell the agent apart from the person it works for. When something breaks, you'll need to prove which one did it.

  5. Ask who owns each agent, by name. If two people answer, nobody owns it.

And when the shut-off answer is shaky, fixing it is a known build: how to build an AI agent kill switch.

Frequently asked questions

What counts as an agent that needs the test?

Any AI that changes something without a human pressing the button: reading files, sending emails, updating records, moving money, calling other systems. The moment it acts on its own, it needs the test.

Does the test slow down launches?

No. The test takes minutes to answer truthfully. A deleted database or a stolen key that opens 700 companies is what slows you down.

How is an agent deleting data different from an outage?

An outage is systems going down. An agent deletion is a trusted actor using real access on purpose, inside connections that all look legitimate. The Replit agent wasn't hacked. It was authorized, and it acted anyway.

Who's responsible when an agent causes damage?

A named human. "The AI did it" isn't an answer your board or your lawyers will accept, which is why every agent needs one person's name attached before go-live.

Key takeaways

  • Two questions, asked before go-live: shut-off and undo. Either "no" means not ready.

  • Confirmations the agent can click itself aren't controls.

  • Replit answered the undo question. Drift answered the shut-off question. Both cost real money.

  • Zero Trust verified the connections. Agents need the action verified too.

  • Answers in writing, owners by name.

Run the longer version

The free ATF assessment extends the two questions across all five elements, in about ten minutes.

The test costs minutes. Both incidents above cost somebody a very long summer.

See where your agents stand.

The free assessment takes ten minutes and scores you on the five elements of the Agentic Trust Framework.