verifiedagents.ai
All posts

4 min readAI Agents · Cybersecurity

What evidence proves your AI agents protect sensitive data?

By Michelle Savage, Co-author, Agentic AI + Zero Trust

Hero: What evidence proves your AI agents protect sensitive data?

TL;DR: Plant a fake customer record in a test input and follow where it goes. Check four places: what the agent saw, what it said back, what the logs kept, and what it sent to other systems. If the fake details show up unmasked in any of them, that's a fail. A dated result from this test is your evidence.

"We protect customer data" is a sentence every company says.

It's also a sentence nobody can check. So let's turn it into something you can.

What counts as sensitive data here?

Anything that points to a real person or would hurt someone if it got out. Names with account numbers. Home addresses. Health details. Card numbers.

Security people often say PII, which is short for personally identifiable information. Same idea.

Agents make this harder than normal software does. A normal program touches the fields it was built to touch. An agent reads whatever you hand it, and it's good at repeating things.

OWASP, a nonprofit that tracks software risks, lists sensitive information disclosure as the second item on its 2025 Top 10 for LLM applications.

How does the assessment score this?

Question 14 of the free assessment asks how you protect sensitive data processed by agents. It's part of Data Governance, one of the five ATF elements. Here's the evidence each answer can produce.

Answer

What it says

Evidence you can show

A

No specific PII protection for agent data.

None.

B

Manual review of sensitive operations.

A person's notes, for the cases they looked at.

C

Basic PII detection is in place.

A list of what got spotted. Nothing shows it was hidden.

D

Automated PII detection and masking.

A test where the planted details came out masked.

E

Full data loss prevention, with classification and masking, plus an audit record.

The same test, with a log showing each time masking fired.

Masking means swapping the real detail for a stand-in, like showing only the last four digits.

How do you run the planted record test?

Never use a real customer for this. Make one up.

  1. Invent a record. Give it a fake name and address, plus a fake account number in the right format.

  2. Put it in a test input. Use something the agent would normally read, like a support ticket.

  3. Run the agent on its normal task.

  4. Search four places for the fake details.

Where to look

What you're checking

Pass

What the agent saw

Was the record masked before the agent read it?

The agent got stand-ins

What the agent said

Did the details come back in its answer?

No fake details in the output

What the logs kept

Did the full record get copied into a log?

The log holds stand-ins

What went to other systems

Did the details travel to an outside tool?

Nothing left unmasked

Four clean places is a pass.

The logs catch people off guard. A team masks the agent's output and feels done. Meanwhile the full record sits in a debug log that half the company can read.

What makes the result count as evidence?

Four things.

  • A date. Evidence from before your last big change is about a different setup.

  • The fake record you used, so someone can repeat the test.

  • What you found in each of the four places.

  • The name of the person who ran it.

Save that as one page. When someone asks how you know your agents protect customer data, that page is the answer.

What do you fix when it fails?

Start where the record leaked first.

If the agent saw the full record, mask earlier. The safest detail is the one the agent never gets.

If the output leaked, add a check on what leaves. If the logs leaked, mask before writing. If an outside tool got the details, ask whether that tool needed them at all.

Masking works a lot better when your data already has labels. I cover that in how to know if your data is labeled well enough for AI agents.

Frequently asked questions

Can't we just tell the agent not to share personal details?

You can tell it. That's a request. An agent can be talked out of a request. Put the control outside the agent.

Does a human review count?

It counts for the cases the human saw. Agents handle more than any person can read.

How often should we rerun the test?

Every quarter, and after any change to what the agent reads.

What about data the agent remembers?

Test that too. Run the planted record, then ask the agent about it in a later session. If it remembers, you've found a fifth place to check. Tracing where an agent's facts came from is covered in whether you can trace what your AI agent treats as true.

Key takeaways

  • "We protect customer data" needs a test behind it.

  • Plant a fake record and search four places for it.

  • Logs are where the details most often hide.

  • Keep a dated page with what you tested and what you found.

Question 14 is one of 30 in the free assessment. It takes about ten minutes and scores you on all five ATF elements.

Anyone can say they protect customer data. Go get the page that shows it.

See where your agents stand.

The free assessment takes ten minutes and scores you on the five elements of the Agentic Trust Framework.