verifiedagents.ai
All posts

7 min readAI Agents · Cybersecurity

What Does Your AI Agent Believe to Be True?

By Michelle Savage, Experience Design Director, PayPal

Hero: What Does Your AI Agent Believe to Be True?

TL;DR: Identity proves who your AI agent is. It can't prove that what the agent believes is true. A belief audit closes that second half: list every source your agents treat as fact and put a name and a date on each one. Then stop any handoff that can't show a working source.

Last updated October 5, 2026. This piece was rebuilt from the ground up around the audit itself, so you can run it this week.

Your identity controls answer one question: is this agent allowed to act? They say nothing about whether the facts behind that action are real. Those are two different questions, and almost all the money spent so far has gone to the first one. The second one barely has a name yet. I call it the Context Layer, and my team has to govern it on every product space we run. This piece shows you how to audit yours, starting this afternoon.

Why isn't AI agent identity enough on its own?

Because an agent can pass every identity check while acting on a fact that's wrong. Josh Woodruff, my co-author, learned this the uncomfortable way. One of his research agents made up a source overnight, and by morning his other agents had built a full day of work on top of it. Every one of them passed the identity check. Every permission was correct. The information was still fake, and it moved through his lab while he slept.

When he told me about that morning, I wasn't surprised. What an AI agent is allowed to believe is the least governed document in most companies. Nobody signs it. Nobody dates it. Somebody pasted it in because they were closest to the problem, and then everyone moved on.

John Kindervag, who created Zero Trust, says it plainly: identity is crackable. Identity work is real work and it counts. It's also only half the problem, and it's the half you've already funded. If you're still working through that half, start with the AI agent identity problem.

What is the belief set your AI agents run on?

Everything your agents treat as true. The knowledge base someone built a year ago. The folder of shared docs. The block of instructions pasted into a settings panel. Any live pull from the open web. That pile decides what every agent downstream accepts as fact, and in most companies no human has ever looked at it as one thing.

My team calls the governed version the Context Layer. Same pile, but with an owner's name on it and a place it lives. That's the whole difference, and it's a bigger difference than it sounds.

Here's the uncomfortable math. Gravitee surveyed 919 people in February 2026 and found 86 percent of AI agents shipped with no security sign-off at all. The belief set behind those agents got even less review than the agents did.

How do you run a belief audit in one week?

Five steps. None of it is a purchase. It's a list, a name, a date, and one good afternoon.

  1. Write down what one chain treats as true.

    Pick the agent chain closest to a customer. List every place its agents pull facts from. That list is the belief set, and most leaders have never seen it on one page. Seeing it is half the work.

  2. Put a name and a date on each source.

    For every item, ask who approved it and when. If nobody signed it and nobody dated it, it isn't governed. You'll find agents running with full confidence on a doc no one has opened since last year.

  3. Find the one belief that hurts most if it's wrong.

    A price, a policy, a legal fact, a discount limit. Pick the single fact that would cost you most if an agent had it backward, then make one check confirm it against your system of record before the next agent builds on it.

  4. Make every agent show its work.

    Each fact gets a live link or a record number attached. "According to industry research" isn't a source. A link that opens to a real page is. No source, no handoff.

  5. Run the ten-minute test.

    Take one thing your agents produced last week and trace every claim back to a real, working source. Time yourself. If you can't do it in ten minutes on a quiet afternoon, you can't do it with an auditor or your own CFO standing over your shoulder.

Do it once and you won't look at your AI agents the same way again.

Where does Zero Trust stop and belief governance start?

Zero Trust is the right strategy, and it does its job well. It checks the caller on every request and continuously verifies the connection. It was never built to check whether the thing being passed is true. In an agent chain, the request can be clean while the content is a lie, and identity waves it through correctly, because identity was doing its job. We walk through the strategy itself in Zero Trust for AI agents in plain English.

Kindervag's first move is always the protect surface: name the small, knowable thing you're defending before you buy anything. What one agent accepts from another as true belongs inside that protect surface. Almost nobody draws the line there, which is exactly why the damage lives in that seam.

Map Josh's bad morning to the Agentic Trust Framework, which the Cloud Security Alliance published in February 2026, and the element that failed was Data Governance. The framework asks every agent what it's eating and serving. His belief set was ungoverned, so the lie traveled at machine speed and nothing was watching the one thing that was actually wrong.

What does a working belief rule look like?

Write rules a machine can check. That's the test that separates a real control from a wish.

Rule as written

Can a second agent enforce it?

Why

"Never make up a source"

No

An agent that invents a citation doesn't know it did

"Only use trustworthy information"

No

Trustworthy is an opinion, and the agent grades its own

"Every fact carries a link that opens to a real page on the approved list"

Yes

A second agent can click the link and check the list

"No working source, no handoff"

Yes

The check happens at the handoff, before the damage spreads

That's the rule Josh added to his own lab after the fake source: a cited page has to actually open, or the work waits for a human. It's slower, and it holds up some research that was fine. He takes that trade every time, because a held handoff costs a few minutes and the fake source cost most of a day plus two documents rebuilt from scratch.

Frequently asked questions

Isn't strong identity management enough to govern AI agents?

No. Identity tells you an agent is who it claims to be and that its permissions check out. It can't tell you the facts behind the agent's action are real. Agents can pass every identity check while building a full day of work on an invented source. Identity is necessary. It isn't sufficient.

What's the difference between a belief set and a knowledge base?

A knowledge base is one source. A belief set is all of them at once: the knowledge base, the pasted instructions, the shared docs, and any live web access an agent has. The Context Layer is the governed version of that pile, treated as one thing with an owner and a date.

How long does a belief audit actually take?

The first pass takes an afternoon for one agent chain. Listing the sources is the slow part. Steps two and three go quickly once the list exists, and the ten-minute traceability test is the one worth repeating monthly.

Who should own the belief audit?

Whoever owns the agent chain can start it. Most of the work is listing sources and asking who approved them, which doesn't need a security team. Bring security in at step three, when you decide which fact gets checked against your system of record before an agent acts on it.

Key takeaways

  • Identity proves who your AI agent is, not whether what it believes is true.

  • The belief set is every source your agents treat as fact, and most companies have never written it down as one thing.

  • Zero Trust continuously verifies the connection. Checking the truth of the content is a separate job that belongs inside your protect surface.

  • A rule an agent checks against its own opinion isn't a rule. Write rules a second agent can enforce.

  • The first belief audit takes an afternoon and costs nothing to start.

If you want to see where your own agents stand first, the free self assessment walks you through the framework in about ten minutes.

Your agents will keep passing the identity test you built. They'll pass it the day before the loss and the morning of. What catches up with everyone is what the agent believed, and how far that belief traveled before anyone checked. An afternoon with a list closes most of that distance.

See where your agents stand.

The free assessment takes ten minutes and scores you on the five elements of the Agentic Trust Framework.