verifiedagents.ai
All posts

5 min readAI Agents

How Do You Detect a Compromised AI Agent?

By Josh Woodruff, Founder & CEO, MassiveScale.AI | Founding Chair, Agentic Trust Framework at the CSAI Foundation

Hero: How Do You Detect a Compromised AI Agent?

TL;DR: Only 5 percent of organizations feel confident they could detect and contain a compromised AI agent, per a 2026 industry survey. EDR and SIEM rules were built for compromised humans and devices, and agents fail differently: the compromise lives in language, not in the operating system. Three signals catch most of it: behavioral drift and scope expansion, plus outputs that don't match the task.

Last updated October 5, 2026. Rebuilt for verifiedagents.ai: the three compromise patterns, the three signals, the four-field log, and the containment playbook with its five-minute clock.

Why doesn't EDR catch a compromised agent?

EDR detects malicious processes and suspicious file access, plus known indicators on a host. A compromised agent trips none of that. Its process keeps running. Its file access stays in scope. Its network calls hit the same approved endpoints they always have. The compromise lives in the language inside the agent, and EDR was never built to read language.

Three patterns dominate in 2026:

Pattern

How it works

What your stack records

Prompt injection

Hidden instructions in a document or ticket the agent reads; it obeys with its own privileges

A clean day

Latent compromise

A line planted in the agent's instruction files drifts its behavior. One 2026 red team added a single line that exported credentials daily and survived restarts

A file change inside the agent's normal workspace, so nothing

Supply chain

ClawHavoc planted 1,100 poisoned skills in a popular open-source framework; the poison activates only under specific conditions

A healthy installation, for weeks

What signals actually catch it?

Three, and all of them need a documented baseline to compare against. No baseline, no detection.

  1. Behavioral drift.

    Build a 30-day baseline per agent: API call counts, systems reached, record types modified, run times, output volume. Anything far outside it goes to a person. This is per-agent profiling, not generic SIEM anomaly rules, because every agent's normal is different. The build order is in

    how to monitor an AI agent's behavior

    .

  2. Scope expansion.

    Treat the agent's documented scope as a contract and alert on every action outside it, including the ones that succeed. An agent approved to read contact records that suddenly reads billing records is a signal even when the read works.

  3. Output-task mismatch.

    The hardest to wire and the most valuable. A customer service agent producing refund approvals nobody asked for is the tell. The named human owner spots this best, because only the owner knows what the agent was supposed to be doing. Wire the owner into detection, not just the security team.

How much logging do you need to replay a decision?

Four fields per action: the input the agent received, the reasoning it produced, the action it took, and the result that came back. Most frameworks log only the last two by default, and that's not enough to tell compromised from mistaken from correct-under-odd-circumstances.

Build the trail at the agent layer, not the system layer, and store it somewhere the agent can't write. A compromised agent will lie about what it did if you let it keep its own books. Microsoft's Defender team published 2026 guidance naming the same four telemetry sources, and any strategy missing one of them misses most compromise patterns.

One more rule: re-baseline whenever the agent's instructions, model version, data access, or memory state changes. The baseline is configuration, not a one-time calibration.

What does the containment playbook look like?

Four steps, a five-minute clock, run by the named owner and verified by security.

  1. Identify.

    Pull the inventory record: version, credentials, scope, documented kill switch. A missing record is its own finding.

  2. Cut credentials.

    API keys, service tokens, database credentials, session tokens. If the agent inherited a human's credentials, rotate the human's too. Don't trust that a 401 stops it.

  3. Verify it stopped.

    No new actions in the logs within five minutes, or the kill switch failed and you escalate. If your switch involves running to a workstation, you have hope, not a switch. The real thing is specced in

    how to build an AI agent kill switch

    .

  4. Preserve evidence.

    Instruction file, memory state, recent audit trail, and the inputs that triggered the behavior, captured before the process restarts.

How do you cap the damage before detection catches up?

Three architectural choices: task-scoped credentials that expire with the task and mandatory human approval on irreversible actions, with the four-field log on everything. None prevents compromise. Together they decide whether the first compromise is a bad hour or a bad quarter.

Frequently asked questions

What's the difference between compromised and misbehaving?

Compromised means turned by an attacker: intentional and adaptive, needing credential rotation and forensics. Misbehaving means doing its job poorly: unintentional and predictable, needing retraining and scope fixes. The detection overlaps. The response doesn't.

How long does a compromised agent dwell?

Early reports suggest weeks to months without per-agent baselines. The benchmark worth remembering is the refund agent that ran outside policy for 11 days with no attacker at all, just nobody comparing prompts to work.

Should AI detect compromised AI?

Useful for first-pass triage at scale, and it adds a new attack surface: the detection agent itself. Keep a human on the second pass, and keep the watcher's autonomy low.

Where does the ATF fit?

Detection lives in Behavioral Monitoring. Containment lives in Incident Response. The framework doesn't replace your stack. It shows what's missing from it.

Key takeaways

  • 95 percent of organizations couldn't catch a turned agent. The tools were built for a different attacker.

  • Three signals, all needing baselines: drift and scope expansion, plus output-task mismatch.

  • Four log fields per action: input, reasoning, action, result. Stored where the agent can't write.

  • Containment is four steps on a five-minute clock, run by the named owner.

  • Scoped credentials and human approval cap the blast radius while detection catches up.

Picture your most-used agent turned right now

If you can't say what your detection stack would show, the free ATF assessment is the ten-minute way to find out where to start.

The confident 5 percent built baselines, logs, owners, and tested switches. The other 95 percent are running production agents on hope.

See where your agents stand.

The free assessment takes ten minutes and scores you on the five elements of the Agentic Trust Framework.