verifiedagents.ai
All posts

7 min readAI Agents · Cybersecurity

Four AI Agent Failures and the Governance That Fixes Them

By Michelle Savage, Experience Design Director, PayPal

Hero: Four AI Agent Failures

TL;DR: AI agent failures almost never come from the model. They come from the missing human layer around it. Four stories from the lab of the man who wrote the governance framework show how agents fail while doing exactly what they were told, and the written rules that would have caught each one.

Last updated October 5, 2026. This piece was rebuilt from the ground up around the failure stories, because the lessons stick better than the checklist.

Here's the uncomfortable setup. Josh Woodruff wrote the Agentic Trust Framework for the Cloud Security Alliance and co-authored Agentic AI + Zero Trust with John Kindervag. He advises CISOs on agent security for a living. He still got burned by his own agents, more than once. If the person who built the governance model can make these mistakes, think about what's happening inside companies that don't have a governance model at all. These four failures, and their fixes, are the fastest education in agent governance you can get.

Why do AI agent rollouts fail when the demo worked?

Because the human layer around the agent was never built. A consultant we worked with had a client who did everything right on paper: capable model, strong use case, impressive demos. Then they went straight to full autonomy. The agent had a 2 percent error rate on customer-facing emails. Sounds small, until real customers called and executives asked who approved this. Six months rebuilding internal trust.

The model worked exactly as designed. Nobody had defined what the agent was allowed to do wrong before someone had to answer for it. That's the pattern in all four stories below.

What do four real agent failures look like?

Failure 1: $300 gone overnight from uncontrolled spawning. Josh's manager agent, Atti, checks in every 30 minutes. A vague instruction made it hand a research task to Scout, the research agent. Scout returned a result. Atti decided it needed more and created another copy of Scout, then another. By 6 AM there were 47 copies of Scout running. No alarms, no spending limits. The fix: agents can't create copies of themselves, every check-in has a spending cap, and daily costs hitting a threshold pause everything and ring his phone. None of it required buying anything new.

Failure 2: security so tight the agent couldn't work. Forge, the coding agent, ran in a locked-down environment. No internet, no files outside its workspace. Perfect security, until a real task needed a standard software package. Blocked, workaround, blocked again. A governance system that stops your AI from doing its job doesn't improve security. It creates pressure to bypass the controls, and "just give it full access this one time" becomes permanent. That's how shadow IT was born in the 1990s, and the same pattern is repeating with agents. The fix was an approved allowlist: Forge requests a new source and gets a one-time approval, and the source stays on the list for good.

Failure 3: six hours of work, gone by morning. Josh and Atti worked six hours on a strategic initiative. Next morning, Atti had no idea any of it happened. Agents hold limited working memory, and long conversations get compressed into vague summaries. Now picture that in a vendor negotiation: a conditional clause the agent agreed to gets lost, and your company is operating on a commitment it can't verify. The fix: every decision, completed action, or status change gets saved to a permanent file immediately. Working memory is a scratchpad. The file is the record of truth.

Failure 4: the model that failed silently. A local model handled Forge's simple tasks fine, then produced professional-looking code that didn't work on a complex job, and sometimes stopped mid-task with no error. It looked finished. It wasn't. The fix: simple tasks go to the cheap model, complex work goes to the capable one, and any output notably shorter or simpler than the task required gets marked for human review instead of treated as done. The free model that fails silently costs more than the paid model that works.

All four failures share one trait: the agents were doing exactly what they were told. The technology worked. The governance didn't.

What four questions would have caught these?

The same four that should precede any agent touching production. Written down, thirty minutes, before anyone writes code.

  1. Who owns this agent?

    One named human who gets the call at 2 AM. No name, no go-live.

  2. What can it do?

    A written list of actions, not systems. "Reads contact records and logs call notes, can't modify deal values or touch billing." That precision counts enormously when a regulator asks what the agent was authorized to do.

  3. What does failure look like?

    A named behavior that triggers human review, like an external message nobody requested. The 47 Scout copies would have tripped this in the first hour.

  4. Who can shut it down, and how fast?

    If the answer involves physically running to a computer, you're hoping, and hope isn't governance. The full build is in

    how to build an AI agent kill switch

    .

When a request can answer all four, approval becomes a form instead of a fight. We cover that intake process in AI agent approval criteria.

How is agent governance different from traditional IT security?

Traditional IT security

AI agent governance

Focus

Controls access to systems

Controls which actions an agent takes inside systems it already reaches

Decision-maker

Humans

Agents, autonomously

How failures appear

Usually visible

Silent and confident-looking

Scope

Defined by role

Defined by specific permitted actions

Records

Built into systems

Agents lose context through memory compression

Shutdown

Revoke credentials

Instant, remote, tested before go-live

The numbers say most companies haven't made this shift. Gravitee found 86 percent of AI agents shipped without security approval in a February 2026 survey of 919 people. A CSA survey presented at the RSAC Conference in 2026 found only 26 percent of organizations have AI governance policies. Token Security found 600 ungoverned agents in 24 hours at a single Fortune 500.

How do you start on Monday morning?

One conversation, ten minutes, more revealing than any audit. Ask your IT and engineering leads, and each business unit head, the same question separately, before they compare notes: what AI agents or automations are running in our environment right now, and what can they reach? You'll get different answers, and the distance between those answers is your finding. Then apply the four questions to every agent on the combined list, thirty minutes each.

Frequently asked questions

What is AI agent governance?

The set of written rules that define who owns an agent, what actions it can take, what counts as failure, and how to shut it down. It's the human layer around autonomous systems that keeps a working demo from becoming a production incident.

What is the golden path principle?

The secure way to work must also be the easy way to work. If following the rules is harder than breaking them, people break them, the same dynamic that created shadow IT in the 1990s. Forge's allowlist fix is the principle in practice: requesting access takes less effort than working around the block.

What caused the $300 overnight failure?

An orchestrator agent interpreted a vague instruction as requiring repeated research and created 47 copies of a research agent overnight, with no spending cap and no circuit breaker. The agents did exactly what they were told, which is the worst kind of failure, because nothing looked broken.

Why can't traditional security tools catch these failures?

Because nothing was breached. Every agent used valid access. The failures were behavioral: too many copies, lost context, confident wrong output. Catching those takes action-level scopes and saved records, plus named failure definitions, not more access control.

Key takeaways

  • Agent failures come from the missing human layer, not the model. In all four lab stories, the agents did exactly what they were told.

  • Four written questions before go-live: owner, scope, failure definition, kill switch. Thirty minutes each.

  • Over-tight controls backfire. The golden path principle says the secure way has to be the easy way, or people route around it.

  • Working memory isn't a record. Decisions get saved to a permanent file the moment they happen.

  • 86 percent of agents ship without security approval, and only 26 percent of organizations have AI governance policies at all.

Want to know which of these failures your own setup would catch? The free self assessment takes about ten minutes and scores you across all five elements of the framework.

The technology almost never causes the failure. The missing governance does, and the companies that write four answers down this week are the ones that won't be rebuilding trust for six months next year.

See where your agents stand.

The free assessment takes ten minutes and scores you on the five elements of the Agentic Trust Framework.