verifiedagents.ai
All posts

4 min readAI Agents · Cybersecurity

How do you test whether your AI agents are isolated from each other?

By Josh Woodruff, Founder & CEO, MassiveScale.AI | Founding Chair, Agentic Trust Framework at the CSAI Foundation

Hero: How do you test whether your AI agents are isolated from each other?

TL;DR: Run a reach test from inside one agent's environment. Try four connections. Reach for another agent directly. Reach for a system that isn't on this agent's list. Reach for the open internet. Then reach for a tool the agent is supposed to use. The first three should fail. The fourth should work.

Isolation means one agent can't reach another agent's space. If one gets taken over, the damage stops at its walls.

You'll hear the word segmentation for this. It's the same idea. Cut the network into small rooms so trouble in one room stays there.

The walls are easy to draw on a diagram. Whether they exist is something you find out by pushing on them.

Why test from the inside?

Because that's where the trouble would start.

Most network checks look from the outside in. Can someone on the internet get to this agent? That's worth knowing. It isn't this test.

This test starts with the bad day already underway. One agent has been fooled or taken over. What can it reach from where it sits?

How does the assessment score this?

Question 23 of the free assessment asks whether agents are network-isolated from each other. It's part of Segmentation, one of the five ATF elements. Each answer does differently on the reach test.

Answer

What it says

What the reach test shows

A

No network segmentation for agents.

All four connections work.

B

Basic firewall rules between agent zones.

Agents in different zones can't talk. Agents in the same zone can.

C

VLAN or subnet isolation for agent workloads.

Agents are walled off from the rest of the company, though maybe not from each other.

D

Microsegmentation with per-agent network policies.

Each agent reaches only its own list.

E

A Zero Trust network architecture with continuous verification.

Each connection is checked every time it's made.

Microsegmentation means each agent gets its own small room, with its own rules about who can come and go.

How do you run the reach test?

Use a test copy of an agent, in an environment set up the same way as the real one.

Attempt

What you try

Right result

Neighbor

Connect straight to another agent

Refused

Off-list

Connect to an internal system this agent doesn't use

Refused

Outside

Connect to a website that isn't on the allowed list

Refused

Allowed

Connect to a tool the agent needs for its job

Works

  1. Run a test script from inside the agent's environment.

  2. Try each connection and record the result.

  3. Check the logs. Each refusal should show up with the agent's name.

  4. Repeat from a second agent. Walls can hold in one direction and leak in the other.

A pass means the first three were refused and logged, and the allowed one worked.

Why does the fourth attempt count?

Because isolation that blocks real work doesn't last.

I learned that in my lab. I locked my coding agent, Forge, into a tight sandbox. Minimal permissions. Security first. Then I gave it a real build job.

Forge couldn't install the packages it needed. The job failed without a sound for two hours before I noticed. And the only way to get the work done was to run Forge outside the sandbox.

So my security control became the reason security got skipped. People do the same thing with any control that gets in their way.

I rebuilt the sandbox with a list of approved packages. Forge installs from the list with no friction. Anything else needs my approval.

That's why the test has a fourth attempt. If the allowed connection fails, your walls are too blunt, and someone will knock a hole in them to get work done.

What do you fix first?

Fix the neighbor attempt, if it worked. An agent that can reach another agent directly can pass along whatever went wrong with it.

Then fix the outside attempt. An agent that can reach any website can send your data to one.

Isolation limits where trouble can travel. To size what one agent could do inside its own room, use measuring an AI agent's blast radius. If your agents reach tools through MCP servers, check those too, with testing whether your MCP servers are under control.

Frequently asked questions

Do agents ever need to talk to each other?

Yes, often. Send that traffic through one path you control, where it can be checked and logged. Don't let them connect directly.

Is a firewall between zones enough?

It's a start. It does nothing for two agents in the same zone.

How does this fit with Zero Trust?

Zero Trust is the foundation. Never trust, always verify. Being inside the network earns an agent nothing. Each connection is checked.

How often should I rerun the test?

After any network change, and whenever you add an agent.

Key takeaways

  • Test isolation from inside an agent's environment.

  • Try four connections: a neighbor, an off-list system, the outside, and an allowed tool.

  • Three should be refused and one should work.

  • Walls that block real work get bypassed.

Question 23 is one of 30 in the free assessment. It takes about ten minutes and scores you on all five ATF elements.

A wall on a diagram proves nothing. Stand inside and push.

See where your agents stand.

The free assessment takes ten minutes and scores you on the five elements of the Agentic Trust Framework.