verifiedagents.ai
All posts

4 min readAI Agents · Cybersecurity

What evidence proves you learn from AI agent incidents?

By Michelle Savage, Co-author, Agentic AI + Zero Trust

Hero: What evidence proves you learn from AI agent incidents?

TL;DR: Pick your last AI agent incident and look for four things. A dated report. A cause written in one plain sentence. A fix with an owner's name that's marked done. And a rerun, where you staged the same problem again and it got caught. If the fourth one is missing, you wrote about the incident. You haven't shown you learned from it.

"We'll make sure this never happens again."

Everybody says it after something goes wrong. It's a nice sentence. It's also a promise with no proof attached.

So here's how to check whether the promise was kept.

What does learning from an incident look like?

It looks like the same problem failing to happen twice.

That sounds obvious. But most of what teams do after an incident is talk about it. They hold a meeting and someone writes notes. People feel better.

Feeling better isn't the same as being safer. Learning leaves something behind that you can point to. A rule that changed, or a test that didn't exist before.

How does the assessment score this?

Question 30 of the free assessment asks how you learn from agent incidents. It's part of Incident Response, one of the five ATF elements. Each answer leaves a different amount of proof.

Answer

What it says

What you can show afterward

A

No formal post-incident process.

Nothing.

B

Informal discussion when something happens.

People's memories of a meeting.

C

We document incident reports.

A report with a date on it.

D

We do root cause analysis with remediation tracking.

A report with a cause, and a list of fixes with owners.

E

A full post-incident review, with policy updates and prevention measures.

All of that, plus a changed rule and proof the change works.

Root cause is the real reason something happened, underneath the first thing that looked wrong.

What are the four pieces of evidence?

Evidence

What good looks like

What doesn't count

A dated report

Written within a week, saying what happened and what it cost

A chat thread

A plain cause

One sentence a new hire would understand

"Human error" or "the model misbehaved"

A finished fix

An owner's name and a due date, marked done

A list of ideas nobody was given

A rerun

The same problem staged again, and caught this time

"We're confident it's fixed"

The second row is where I'd push hardest. "The agent made a mistake" isn't a cause. Why was it able to? What was missing that would've stopped it?

Here's the difference in practice.

Weak cause

Useful cause

"The agent sent the wrong email."

"Nothing checked outgoing emails against the approved template."

"The agent used old prices."

"The price file had no owner, so nobody updated it."

A useful cause points at something you can fix.

How do you run the check?

Give yourself an hour.

  1. Pick the most recent incident. A near miss works too.

  2. Look for each of the four pieces. Don't accept "I think we did that."

  3. Score one point for each one you can open and read.

  4. If the rerun is missing, do it now. Stage the same problem in a test setting and see what happens.

Four points is a pass.

The rerun is the honest part. If you stage the same problem and it slips through again, you've saved yourself from finding that out the hard way. Josh covers how to stage a problem safely in testing your AI agent incident response before the first incident.

What if we haven't had an incident?

Lucky you. Maybe.

It could mean your agents are well run. It could also mean nobody's noticed anything yet. Those look the same from the inside.

Either way, you can still build the habit. Run a drill, then treat the drill like a real incident. Write the report. Name the cause of whatever went badly. Assign the fix. Run it again.

Frequently asked questions

Who should write the report?

Someone who was there, with a second person reading it for plain language. If the second reader can't follow it, rewrite it.

Should the review name who made the mistake?

Name the missing control. If a person could make the mistake that easily, the next person will too.

How long should a review take?

Days. If it takes a month, the details are gone and the fixes have gone stale.

What should change in our go-live checks?

Add the lesson. If an incident showed you couldn't undo something, that goes into the check every new agent has to pass. I describe that check in what evidence proves an AI agent passes the shut-off and undo test.

Key takeaways

  • A promise that it won't happen again needs proof.

  • Look for four things: a dated report, a plain cause, a finished fix, and a rerun.

  • "The agent made a mistake" isn't a cause. Ask what was missing.

  • The rerun is what shows the fix works.

Question 30 is one of 30 in the free assessment. It takes about ten minutes and scores you on all five ATF elements.

Anyone can write up what went wrong. Stage it again and see if it still does.

See where your agents stand.

The free assessment takes ten minutes and scores you on the five elements of the Agentic Trust Framework.