verifiedagents.ai
All posts

4 min readAI Agents · Cybersecurity

How do you measure how fast you detect an AI agent incident?

By Josh Woodruff, Founder & CEO, MassiveScale.AI | Founding Chair, Agentic Trust Framework at the CSAI Foundation

Hero: How do you measure how fast you detect an AI agent incident?

TL;DR: Take your last few AI agent incidents and near misses. For each one, write down the time of the first bad action and the time a person first knew. Then write down who told you. The time between is your detection time. If a customer told you, it's a fail, however short the time was.

Teams love to measure how fast they fix things. Fewer measure how long the problem ran before anyone knew.

With agents, that second number is the one that hurts. An agent doesn't get tired or stop to think. Every minute it runs wrong is more wrong work to clean up.

What exactly are you measuring?

Two times and one name.

The first time is when the agent took its first bad action. You get that from the logs, after the fact. It's almost always earlier than people guess.

The second time is when a person who could act first knew something was wrong.

The name is who told you. A customer, a coworker who happened to look, a scheduled review, or your own alert. Those four sources are very different, and the source says more about your setup than the clock does.

How does the assessment score this?

Question 27 of the free assessment asks how quickly you detect agent incidents. It's part of Incident Response, one of the five ATF elements. Each answer lines up with a source.

Answer

What it says

Who usually tells you

A

We only know about problems when users report them.

A customer.

B

We find incidents during regular log reviews.

A review, days later.

C

We have threshold-based alerts.

An alert, if the problem was loud.

D

We have real-time anomaly alerts.

An alert, even when the problem was subtle.

E

We have continuous monitoring with proactive detection.

An alert, before the damage starts.

How do you measure from your own history?

Go back through the last six months.

  1. List every incident and near miss that involved an agent. Include the small ones.

  2. Find the first bad action in the logs. Don't use the time on the ticket.

  3. Find when a person first knew. A chat message or a ticket usually shows it.

  4. Write down who told them.

  5. Subtract. That's the detection time for that incident.

Put it in a table like this one.

What to record

Where to find it

Why it counts

First bad action

The agent's action log

It's the true start, not the ticket time

First human awareness

The first message or ticket

It's when a response could begin

Source

Ask the person who raised it

It shows whether your alerts did the work

Detection time

The second time minus the first

It's the number you report

If you can't find the first bad action, write that down. It means your logs can't tell you when an incident began, and that's a finding too.

What's a good result?

Judge the source first and the clock second.

Source

Verdict

Your own alert, in minutes

Good

Your own alert, in hours

The alert works. It's too slow or too quiet.

A person who happened to look

Luck. Don't count on it twice.

A customer

Fail. Someone outside saw it before you did.

My bar is simple. Your own alert tells you, and it tells you in minutes.

One number will hide the truth, so don't average. A two-minute detection and a nine-day detection don't make a four-and-a-half-day problem. They make one good day and one very bad one. Report the worst.

What if you have no history?

Then you measure with drills. Trigger a fake problem in a test setting and time how long until someone knows. I lay out that drill in whether you'd see an AI agent misbehave in real time.

Be careful what you conclude, though. A drill tells you the alert path works. It doesn't tell you whether you'd notice a kind of problem nobody thought to test.

Having no incidents on record can also mean nobody's looking. Check that before you celebrate.

Frequently asked questions

Isn't response time the number leadership wants?

Often. Give them both. Detection time is usually the bigger share of the damage.

Should near misses count?

Yes. A near miss is a free lesson. It has a first bad action and a moment someone noticed, same as a real incident.

How do we get the number down?

Move up one source at a time. If customers tell you, get to reviews. If reviews tell you, get to alerts. Then make the alerts faster.

What happens after detection?

The response clock starts. I cover that hour in testing your AI agent incident response before the first incident.

Key takeaways

  • Detection time runs from the first bad action to the first person who knew.

  • Record who told you. The source says more than the clock.

  • If a customer told you, it's a fail.

  • Don't average. Report your worst case.

Question 27 is one of 30 in the free assessment. It takes about ten minutes and scores you on all five ATF elements.

You can't respond to something you haven't noticed. Find out how long that usually takes you.

See where your agents stand.

The free assessment takes ten minutes and scores you on the five elements of the Agentic Trust Framework.