verifiedagents.ai
All posts

7 min readAI Agents · Cybersecurity

How Security Says Yes to AI Agents Without Losing Control

By Michelle Savage, Experience Design Director, PayPal

Hero: How Security Says Yes to AI Agents Safely

TL;DR: Security leaders don't have to pick between blocking AI agents and absorbing the risk. Approve every agent with four written questions: owner, scope, failure definition, kill switch. Then let it earn autonomy through four levels, Intern to Principal, with evidence at each step. That's how you say yes safely.

Last updated October 5, 2026. This piece was rebuilt from the ground up around the operating model, so you can run the four questions and the Autonomy Ladder this week.

Here's the squeeze every security leader knows right now. The board wants AI shipped yesterday. Product teams are building agents without looping security in. Say yes without controls and you own whatever goes wrong. Say no and you killed the AI initiative, and the business moves without you anyway. One security leader put it plainly: "Security says no, the business moves anyway. The choice is to engage and influence, or watch from the sidelines."

There's a third option, and the leaders using it are the ones getting budget instead of blame. It fits in four written questions and one ladder.

How bad is the visibility problem really?

Worse than most boards know. At the RSAC Conference in 2026, a discovery scan at a single Fortune 500 found 600 AI agents nobody had approved, running with access to AWS, Snowflake, GitHub, and production code. Nobody was monitoring them. Nobody could shut one down on command. Nobody owned the consequences. And that was 24 hours of looking at one company.

The broader numbers match. Only 14.4 percent of organizations report that all their AI agents went live with full security and IT approval. The rest are already running, under-secured or not secured at all. If you haven't counted yours yet, start with AI agent sprawl, because you can't govern what you can't see.

And the damage doesn't need an attacker. A customer-service agent at a mid-size company started issuing refunds outside policy after a customer learned how to prompt it, and it ran for 11 days before anyone noticed. No breach. Just an agent doing what it was built to do, with nobody checking whether that matched what the business wanted.

What four questions should approve every AI agent?

Four questions, answered in writing, before any agent touches a production system. The whole exercise takes under thirty minutes, and it happens before anyone writes a line of code.

  1. Who owns this agent?

    One named human accountable for its behavior. Not a team. No named owner, no go-live. We cover why in

    who owns your AI agent

    .

  2. What's its scope?

    A written list of actions, not systems. "This agent has access to the CRM" isn't a scope. "This agent can read contact records and log call notes, and can't modify deal values or touch billing" is. Regulators want approved actions, so write it that way from the start.

  3. What does it do wrong that triggers human review?

    A named behavior, not a generic error threshold. "If the agent sends an external message nobody asked for." "If it modifies a record outside its scope." Almost nobody writes this one down, and it's the question that separates governing an agent from describing it after something breaks.

  4. How do you kill it?

    A named person who can shut it down in under five minutes, with a log entry proving they tested it in the last 30 days. Not "we have a process." A log entry. The build is in

    how to build an AI agent kill switch

    .

Most reviews answer the first two. Almost none have a written failure definition, and the kill switch test almost never happens until someone asks. That's where the liability lives.

What is the Autonomy Ladder?

A path between yes and no. Most companies treat agent approval as binary: it runs or it doesn't. The ladder replaces that with four levels, and each promotion requires written evidence. Think about hiring. You don't give an intern root access to production. Trust gets earned through demonstrated behavior, and agents work the same way.

Level

What the agent may do

What promotion requires

Intern

Observes and recommends. Executes nothing. Every output reviewed by a human.

Every agent starts here. No exceptions for a good demo.

Junior

Executes low-risk, reversible actions within scope.

Reversible means undone in under five minutes without data loss.

Senior

Executes moderate-risk actions with escalation triggers and automated logging.

30 days of action logs and a documented error rate, plus two examples of correct escalation at the edge of scope.

Principal

Operates with an established track record.

Clear ownership and auditable logs, plus a formal review every 90 days, because the business context keeps changing.

The ladder is also a negotiation tool. When the business pushes to ship faster, you're not saying no. You're saying: Intern this week, Junior in 30 days if the logs look right, Senior in 90 if we can show the board a clean record. Most leaders take that deal. Blanket refusal is the one they won't, and blanket refusal is what ends your seat at the table.

Where do AI agents need more than standard Zero Trust?

Zero Trust is the right foundation, and agents stress it in four specific places it was never asked to cover. The Agentic Trust Framework extends it there rather than replacing it.

Memory drift. An agent at 80 percent memory capacity reasons differently than it did at 20. Identity gets verified at session start, but nothing watches behavior shift inside a session. Practical fix: checkpoints. At 60 percent capacity, save state. At 70, require human review before continuing.

Least privilege at task level. An agent approved to read customer records for contract renewals uses that same access for everything else it runs. Real least privilege scopes access to the task running right now, with expiry. Most identity platforms don't support that for non-human identities yet. Write the policy anyway and make the technology catch up.

Agent-to-agent calls. When agents call other agents, most architectures inherit the highest privilege in the chain by default. Nobody decided that. Treat every agent-to-agent call like an external API call: no contract, no action.

Behavioral baselines. Normal for a human is stable. Normal for an agent shifts with every prompt edit and every model or data change. Log every action for 30 days before granting autonomy above Intern, re-baseline after every change, and send anything outside two standard deviations to a human.

Frequently asked questions

What's a fast way to find shadow AI agents in my company?

Ask your IT and engineering leads, and each business unit head, the same question separately, before they compare notes: what AI agents are running, and what can they reach? The differences between their answers are your finding. Then check network logs against known AI provider endpoints and audit finance for unexplained AI platform charges.

What does a real failure definition look like?

A specific named behavior a monitoring system can trigger on. "If the agent sends an external message that wasn't requested." "If it issues a refund above a set amount without approval." Generic error thresholds don't count, because the 11-day refund agent never threw an error.

How do I decide what level an agent sits at?

Evidence, not vibes. Intern is the default for every new agent. Junior requires actions reversible in under five minutes. Senior requires 30 days of logs and a documented error rate, plus two correct escalations. Principal requires a track record and a 90-day review cadence.

What is the Agentic Trust Framework?

The open governance standard that extends Zero Trust to AI agents, published by the Cloud Security Alliance in February 2026. Its five elements are Identity Management, Behavioral Monitoring, Data Governance, Segmentation, and Incident Response.

Key takeaways

  • A discovery scan at the RSAC Conference in 2026 found 600 unapproved agents at one Fortune 500 in 24 hours, with access to AWS, Snowflake, GitHub, and production code.

  • Only 14.4 percent of organizations say all their agents went live with full security and IT approval.

  • Approve every agent with four written answers: owner, scope, failure definition, kill switch. Thirty minutes, before any code.

  • The Autonomy Ladder, Intern through Principal, turns "no" into "yes, with evidence," and keeps security in the conversation.

  • Agents need four additions on top of Zero Trust: memory checkpoints, task-level least privilege, agent-to-agent policy, and per-agent behavioral baselines.

To see where your own program stands against the five elements, the free self assessment takes about ten minutes.

AI agents are already inside your company. Governance decides whether they become the reason AI shipped or the reason it got pulled, and the teams that wrote four answers down before the first line of code are the ones still moving.

See where your agents stand.

The free assessment takes ten minutes and scores you on the five elements of the Agentic Trust Framework.