4 min readAI Agents · Cybersecurity
What evidence proves your AI agent logs would hold up in an audit?
By Josh Woodruff, Founder & CEO, MassiveScale.AI | Founding Chair, Agentic Trust Framework at the CSAI Foundation

TL;DR: Pick one action an AI agent took last week and rebuild it from the logs alone. You need four things for that action: what came in, what the agent reasoned, what it did, and what came back. If any one is missing, the log won't hold up. Then check that the agent can't edit its own log.
In the first chapter of the book I wrote with Michelle Savage, I describe a demo I watched at the RSAC Conference in 2025. A dashboard showed the logins at an ordinary company of 400 people. About 60% of them weren't human.
One agent on that screen had handled 900 customer service tickets in 24 hours. It had made 56 pricing decisions along the way.
Now picture an auditor pointing at one of those 56 and asking why. That's the test.
What does hold up mean?
It means someone who wasn't there can rebuild what happened from the log alone.
An ordinary app log says the agent called a tool at 2:14. That proves something happened. It doesn't say what the agent was looking at or why it chose that tool.
For a person, you can ask. For an agent, the log is the only witness. If the log is thin, there's no one else to question.
How does the assessment score this?
Question 7 of the free assessment asks how agent actions are logged. It's part of Behavioral Monitoring, one of the five ATF elements. Each answer gives you a different amount to work with on a replay.
Answer | What it says | What you can rebuild |
|---|---|---|
A | Minimal or no logging of agent actions. | Nothing. |
B | Basic application logs that capture some agent activity. | That something happened, and when. |
C | Structured logging that records each action with its inputs and outputs. | What went in and what came out. The why is missing. |
D | Comprehensive logging of inputs and outputs, with the agent's reasoning, kept as a full audit trail. | The whole action, start to finish. |
E | A full audit trail with lineage tracking, fit for forensics. | The whole action, plus where its data came from. |
How do you run the replay test?
Pick one real action from last week. Choose one that changed something, like a refund or a record update.
Give the log to someone who wasn't involved. Don't explain anything.
Ask them four questions. What did the agent receive? What did it reason? What did it do? What came back?
Time them. Stop at 30 minutes.
Score one point for each question they can answer from the log. Four points is a pass.
Many first runs score two. The log shows what the agent did and what came back. The input and the reasoning are gone.
What else does the evidence need?
Two more checks. Both are quick.
Where the log lives. Try to change a log entry using the agent's own access, in a test setting. It should fail. A log the agent can edit is a diary. An agent that's been turned will lie in it.
How long the log lasts. Ask how far back you can go. Then ask your auditor how far back they'll look. If the first number is smaller than the second, you have a hole in your evidence.
How common is thin logging?
Very. Gravitee's State of AI Agent Security 2026 report found that more than half of all agents run with no security oversight or logging.
So if your replay scores two, you're ahead of many. You're still two short.
Frequently asked questions
Do I need to log the agent's reasoning?
For any action that changes something, yes. The reasoning is how you tell a mistake from an attack. Without it, both look the same.
Won't full logging hold sensitive data?
It can. Treat the log like the data it records. Limit who can read it, and mask what you'd mask anywhere else.
What do I do with good logs once I have them?
Use them to build a baseline of normal. That's the next test, in whether your monitoring would catch an AI agent drifting slowly.
How does this help in an attack?
It's how you find one. The same four fields are what you check in whether you would catch a compromised AI agent.
Key takeaways
A log holds up when a stranger can rebuild one action from it.
Four fields per action: what came in, the reasoning, the action, the result.
Four out of four on the replay test is a pass.
The agent must not be able to edit its own log.
Question 7 is one of 30 in the free assessment. It takes about ten minutes and scores you on all five ATF elements.
One agent made 56 pricing decisions in a day. Pick one of yours and see if you can explain it.