verifiedagents.ai
All posts

4 min readAI Agents · Cybersecurity

How do you test whether an AI agent can explain its decisions?

By Michelle Savage, Co-author, Agentic AI + Zero Trust

Hero: How do you test whether an AI agent can explain its decisions?

TL;DR: Pick five decisions an AI agent made last week. For each one, ask why. You want one plain sentence a customer would understand. Then check that sentence against what was logged when the decision happened. Five reasons that are plain and match the record is a pass. A nice story written afterward doesn't count.

Ask a person why they did something and you'll get an answer. It might be a bad answer. But you'll get one.

Ask an AI agent and you might get nothing. Or you might get a smooth paragraph that sounds great and was made up on the spot.

Both are a problem. The second one is sneakier.

Why does an explanation count for so much?

Because someone will ask. A customer or a manager will. So will an auditor.

Here's a small example. A customer says, "Cancel my subscription." The agent says, "Done. You won't be charged again." But there's a 14-day processing window, and one more charge is coming.

The customer doesn't think "billing mix-up." They think "they lied to me." I call that a polite breach of trust. The agent did its job, and the customer still got hurt.

When that customer calls, somebody has to explain what the agent did and why. If nobody can, you've got a second problem on top of the first.

How does the assessment score this?

Question 11 of the free assessment asks whether your agents can explain their decisions. It's part of Behavioral Monitoring, one of the five ATF elements. Here's how each answer sounds when someone asks why.

Answer

What it says

What you can tell the person asking

A

No explainability capability.

"We don't know."

B

Basic logging of decision inputs.

"Here's what it was looking at. We can't say why it chose that."

C

Decision rationale is logged for review.

"Here's the reason it recorded at the time."

D

On-demand explanation queries are supported.

"Let me pull that up for you right now."

E

Real-time explanations, with what would have changed the outcome.

"Here's why, and here's what would've made it go the other way."

How do you run the test?

You need five decisions and someone who doesn't work on the agent.

  1. Pick five decisions from last week. Choose ones that affected a customer or cost money.

  2. Get the reason for each. Use whatever your setup offers, whether that's a log entry or a question to the system.

  3. Hand each reason to your outsider. Ask them to say it back in one sentence.

  4. Check each reason against the record. Look at what was logged at the moment of the decision.

Then grade each decision on two things.

Check

Pass

Fail

Is it plain?

Your outsider can say it back in one sentence

They need a glossary

Is it true?

It matches what was logged at the time

It was written afterward and nothing backs it up

Five decisions that pass both checks is a pass.

What's wrong with an explanation written afterward?

It might be fiction.

If you ask an agent today why it did something last Tuesday, it'll give you an answer. It's good at answers. That doesn't mean the answer is what really drove the choice.

So the reason has to be saved when the decision is made. Then a later explanation has something to be checked against. Josh covers what a solid record needs in what evidence proves your AI agent logs would hold up in an audit.

What does a plain explanation look like?

Short, and in the customer's words.

Hard to use

Easy to use

"Anomalous pattern recognition in the classification layer."

"It treated the new scanner's images as a warning sign."

"The request exceeded policy parameters."

"The refund was over the $200 limit, so it asked a person."

If you wouldn't say it out loud to a customer, rewrite it.

Frequently asked questions

Does every decision need an explanation?

No. Start with decisions that touch a customer or move money. Then add the ones that can't be undone and the ones people keep asking about.

Who should write the plain version?

Someone who talks to customers. Engineers know what happened. The person on the phone knows how to say it.

Can we just ask the agent to explain itself?

You can ask. Don't trust the answer by itself. Check it against what was recorded.

How fast do we need the explanation?

Before the person asking loses patience. If something's going wrong right now, minutes count, and Josh has a test for that in whether you'd see an AI agent misbehave in real time.

Key takeaways

  • An explanation has to be plain and true. One without the other fails.

  • Test five real decisions with someone who doesn't work on the agent.

  • Save the reason when the decision is made. Later is too late.

  • If you wouldn't say it to a customer, rewrite it.

Question 11 is one of 30 in the free assessment. It takes about ten minutes and scores you on all five ATF elements.

Your agent made a few thousand decisions this week. Somebody's going to ask about one of them. Have the sentence ready.

See where your agents stand.

The free assessment takes ten minutes and scores you on the five elements of the Agentic Trust Framework.