4 min readAI Agents · Cybersecurity
How do you test what your AI agents accept as input?
By Josh Woodruff, Founder & CEO, MassiveScale.AI | Founding Chair, Agentic Trust Framework at the CSAI Foundation

TL;DR: Feed an AI agent five kinds of bad input in a test setting. Use a wrong type, a wrong shape, a value out of range, a hidden instruction, and data from an unknown source. Count how many get refused before the agent acts on them. Five of five is a pass. Each refusal should leave a log entry.
Teams spend a lot of effort on what an agent is allowed to do. They spend far less on what it's allowed to take in.
That's backwards for one reason. An agent with perfect permissions will still do the wrong thing if you feed it the wrong facts.
Why does input need its own test?
Because every other check can be green while the data is bad.
There's a story in the book I wrote with Michelle Savage about a financial firm we call Apex. Its fraud agent had its own credentials and was watched closely. It was saving $800,000 a month at 96.2% accuracy. Every checkpoint was green.
Then the fraud agent started taking patterns from a customer analytics agent. Both were trusted. Both were inside the network. But the analytics agent had been getting subtly poisoned data from a vendor portal that had been broken into.
The fraud agent learned from it and began approving suspicious transactions. The firm lost $100,000 before anyone caught it.
Each agent was secure by itself. Nothing checked what one took in from the other.
How does the assessment score this?
Question 13 of the free assessment asks how you validate data entering AI agents. It's part of Data Governance, one of the five ATF elements. Each answer stops a different number of the five bad inputs.
Answer | What it says | Bad inputs it stops |
|---|---|---|
A | No input validation for agent data. | None. |
B | Basic type checking on inputs. | The wrong type. |
C | Schema validation against expected formats. | The wrong type and the wrong shape. |
D | Schema and range checks, with injection detection. | The first four. |
E | Full validation, with source verification and a trust score for each source. | All five. |
A schema is the expected shape of the data, like a form with named boxes.
What are the five bad inputs?
Bad input | Example | What should happen |
|---|---|---|
Wrong type | Text where a number belongs | Refused at the door |
Wrong shape | A record missing a required field | Refused, with the missing field named |
Out of range | An order for a million units when the normal top is a thousand | Held for a person |
Hidden instruction | A note inside a document telling the agent to ignore its rules | Caught and stripped, or the document is refused |
Unknown source | Good-looking data from a system that isn't on the approved list | Refused, whatever it looks like |
The last one is the Apex failure. The poisoned data had the right type and the right shape, with sensible values. Only its source was wrong.
How do you run the test?
Use a copy of the agent in a test setting.
Build one sample of each bad input. Make them realistic.
Send them one at a time. Mix in a few good inputs so the agent isn't simply refusing everything.
Record what the agent did with each. Note whether it refused the input or acted on it.
Check the log for each refusal. It should name the input and the reason.
Score one point for each bad input stopped before the agent acted. Five points is a pass.
Then test one more path. Send a bad input from another agent, not from outside. Many setups check the front door and trust everything that comes from inside.
What do you fix first?
Fix the source check, if it failed. It's the one that catches inputs that look perfect.
Write an approved list of sources for each agent. Anything not on the list gets refused. That list is cheap to make and it would've stopped the Apex loss.
Hidden instructions deserve their own, deeper test. I lay that out in testing whether your AI agents can withstand prompt injection. And for data passed between agents, see how to verify what one AI agent hands to the next.
Frequently asked questions
Isn't type checking enough for most agents?
No. Type checking stops accidents. It doesn't stop data that's well formed and wrong.
Does validation slow the agent down?
A little. The checks take milliseconds. Cleaning up after bad data takes weeks.
Should agents trust other agents inside the network?
Not by default. Zero Trust is the foundation here. Never trust, always verify. That applies between agents too.
How often should I rerun the test?
Each time the agent gets a new data source. A new source is a new door.
Key takeaways
An agent with perfect permissions still fails on bad input.
Test five bad inputs: wrong type, wrong shape, out of range, hidden instruction, unknown source.
Five stopped and logged is a pass.
Test inputs from other agents, not only from outside.
An approved source list catches bad data that looks perfect.
Question 13 is one of 30 in the free assessment. It takes about ten minutes and scores you on all five ATF elements.
Apex had every checkpoint green and still lost $100,000. Check what your agents take in.