4 min readAI Agents · Cybersecurity
How do you score your AI agent observability tooling?
By Josh Woodruff, Founder & CEO, MassiveScale.AI | Founding Chair, Agentic Trust Framework at the CSAI Foundation

TL;DR: Score your tooling on four jobs, one point each. Can it follow one request from start to finish? Can it search every agent's records in one place? Can it see prompts and tool calls, the things only agents have? Can it alert a person? Then prove the first one by tracing a single request in ten minutes.
Observability is a long word for a simple idea. It means you can see what's happening inside a running system from the outside.
For normal software, the tools are mature. For agents, most teams are still using the normal-software tools and hoping. So the first thing to score is whether your tools can see an agent at all.
What makes agent tooling different?
Ordinary monitoring answers one question. Is it up?
An agent can be up and wrong. It can be fast and healthy while it does the wrong job. Ordinary tools show green the whole time.
Agent-aware tooling sees things ordinary tools don't. It sees the prompt that went in and the tools the agent called. It sees the order of the steps and what each one cost. Those are the facts you need when someone asks what happened.
How does the assessment score this?
Question 12 of the free assessment asks what observability tooling you have for AI agents. It belongs to Behavioral Monitoring, one of the five ATF elements. Each answer covers a different number of the four jobs.
Answer | What it says | Jobs it usually covers |
|---|---|---|
A | Standard application monitoring only, with nothing agent-aware. | None of the four. It tells you the agent is up. |
B | Custom logging we built ourselves. | Whatever you built it for, and only for the agents you wired in. |
C | Log aggregation with search. | Search in one place. |
D | A dedicated LLM or agent observability platform. | Search and agent-level detail, and usually end-to-end traces. |
E | A full stack with tracing, metrics, alerts, and dashboards. | All four. |
I don't sell products, and I'm not going to name a winner. Score the jobs. The brand doesn't change the score.
What are the four jobs?
Job | The question it answers | How to check it |
|---|---|---|
Trace | What path did this one request take? | Follow one request across every agent and tool it touched |
Search | What did any agent do, and when? | Find one agent's action from last Tuesday without logging in to a second system |
Agent detail | What was the agent told, and what did it call? | Open one step and read the prompt and the tool call |
Alert | Who finds out when something's off? | Check that an alert goes to a named person |
One point each. A score of four means your tools can do the work the rest of your monitoring depends on.
How do you run the one-request test?
This proves the first job, which is the one that fails most.
Pick a real request from yesterday. Choose one that went through more than one agent or tool.
Start a ten-minute clock.
Follow it. Find each step, in order, with what went in and what came out.
Mark where you lost the thread.
If you reach the end inside ten minutes, the trace job passes. If you had to open three systems and line up timestamps by hand, it doesn't.
The place you lost the thread is usually a handoff between two agents. That's worth knowing. It's also where bad data slips through, which is why I test it separately in how to verify what one AI agent hands to the next.
What do you fix first?
Fix the lowest job that blocks the others.
Search comes first. If records sit in five places, nothing else works well. Agent detail comes next, because a trace with no prompts is a list of timestamps. Trace follows. Alerts come last, since an alert is only as good as what it can see.
Good tooling is the base for the tests that come after it. Once you can see, check that you'd be told, using the test for whether you'd see an AI agent misbehave in real time.
Frequently asked questions
Do we need a dedicated agent platform?
Not always. What you need is the four jobs done. Some teams get there by extending tools they own.
Is custom logging a bad answer?
No. It's a fragile one. It covers the agents someone remembered to wire in, and it depends on the person who built it.
How often should we rescore?
Each time you add an agent platform. New platforms often keep their records somewhere new.
Does tooling replace a baseline?
No. Tooling lets you see. A baseline tells you what normal looks like.
Key takeaways
Ordinary monitoring says the agent is up. It can't say the agent is right.
Score four jobs: trace, search, agent detail, and alert.
Prove the trace job by following one request in ten minutes.
Fix search first, since the other jobs lean on it.
Question 12 is one of 30 in the free assessment. It takes about ten minutes and scores you on all five ATF elements.
You can't check what you can't see. Score the tools before you trust what they tell you.