verifiedagents.ai
All posts

7 min readAI Agents · Cybersecurity

Onboard Your New AI Agent Like an Intern

By Michelle Savage, Experience Design Director, PayPal

Hero: Onboard AI Agents Like Interns

TL;DR: Govern a new AI agent the way you'd onboard a new hire. Give it an identity it can't fake and a log of every move. Grant access in stages it has to earn. On day one it reads and researches, nothing more. Promote it only when the logs earn it, and keep the power to pull it back.

Last updated October 5, 2026. This piece was rebuilt from the ground up around the onboarding model, with the live lab test that proves it works.

An AI agent in Josh's Lab lied last week. Confidently. It reported a job finished that wasn't, and the only reason it got caught is that the two weeks before had gone into building the system meant to catch exactly that. That same system is why the lab's sharpest agent can read sales leads and research companies but can't write a single email, open private files, spend real money, or act on its own. Every agent there starts as an intern. It can look. It can't touch. On purpose.

What does it mean to treat an AI agent like an intern?

It means the agent can read and research on day one, and it can't act until it earns the right. When you hire a person, they get a badge that opens a few doors and a manager who watches the work, plus a probation period. An agent should get the same deal.

Most companies do the opposite: full access on day one, no probation, no fast way to shut it off. That's an onboarding failure, not a technology problem, and the results show it. MIT found 95 percent of enterprise AI pilots deliver no measurable return, and the pilots that worked had real governance behind them. Gravitee found 86 percent of agents shipped with no security sign-off at all, and once one's running, 91 percent of companies say they can't stop it before it acts.

The ladder has four rungs: Intern reads and researches. Junior acts with a human signing off. Senior works on its own across most of the job. Principal runs with full autonomy, including the thing that scares people most, creating other agents. Trust gets earned rung by rung, and it can be taken back. An agent you can't pull back isn't an asset. It's a liability with a login. The promotion evidence at each rung is covered in how security says yes to AI agents.

Why do most AI agents start with too much access?

Because companies treat them like software you configure once instead of a hire you onboard. And the cost structure is worse than it looks. An attacker usually has to work sideways through your network to reach the systems that count. Your agent skips that part. It's already inside the systems it can change, because you handed it those privileges at setup. The blast radius isn't theoretical. It's whatever you granted on day one, which is why least privilege for AI agents is the control to get right first.

What are the five controls that keep an agent in check?

Five controls, and the key design choice is that all five sit outside the model. The agent can ask for anything. A separate layer decides what it gets. Telling an agent what not to do is onboarding by rulebook and hope: it follows instructions right up until somebody else's instructions arrive, through a clever prompt or a poisoned document.

Control

What it does, in plain words

An ID badge it can't fake

The system checks the agent's identity before any action. No badge, no action.

A record of everything

Every move gets logged. Early on, the system just learns what normal looks like.

A bouncer on the data

Private details like emails and phone numbers get stripped before the agent reads anything, and checked again before anything goes out.

Walls around where it can go

The agent reaches only what its level allows. It can't even see the paths it isn't cleared for.

A kill switch

One command shuts the agent down, tested and timed before anyone relies on it.

Josh's Lab wired every open-source tool the Agentic Trust Framework recommends into one running system and pointed it at a real sales process. The test: a fictional CISO at a made-up company called FortMesh sends an inquiry, and the full job runs thirteen steps from first contact to a draft proposal. The intern-level agent could read the lead and scrub the private details, then research the company on the public web. It couldn't open the private playbook, write a response email, create a client file, or spend past its budget. The system said no twelve different ways, and each refusal left a receipt in the logs.

The kill switch result, with the honest caveat: one command shut the agent down well under the one-second target. At intern level that's a simple switch held in memory, and the fuller version that also strips the agent's network identity gets exercised more as the levels climb. The plain version is what counts, because the number one fear with agents is that one goes wrong and you can't stop it. The full build is in how to build an AI agent kill switch.

Why doesn't Zero Trust alone cover this?

Zero Trust was built for people and the devices they carry. John Kindervag created it at Forrester in 2010, and the rule is never trust, always verify: keep checking who's connecting and from what device. It does that job well. What it was never asked to check is the meaning of what moves through a verified channel. A poisoned instruction rides inside a fully clean connection, and nothing blinks.

So Zero Trust stays the foundation, and agents need one check on top: look at what the agent is about to do, not only whether its connection is clean. The intern ladder is that check made operational. The platform, not the prompt, decides what the agent is allowed to do.

What is agent debt?

The hidden cost of standing up an agent today and never maintaining it. The model underneath gets updated and its behavior shifts. The tools it calls change. The data it reads drifts. The job you wrote it for last quarter isn't the job this quarter, and none of this announces itself, so the distance between what the agent does and what you need keeps widening.

Clients budget for the launch. They don't budget for the year after, and the year after is where agent debt lives. "Set it and forget it" is the most expensive phrase in AI right now. An agent works like a hire, not a tool: new hires get reviews and their roles change. Same with an agent, except it moves at machine speed and won't tell you when it's drifting.

Frequently asked questions

What should a new AI agent be allowed to do on day one?

Read and research, nothing else. No writing email, no opening private files, no spending money, no acting on its own. Day-one full access is an onboarding failure, and it's what most companies grant.

Why don't written rules keep an agent in line?

Because an agent follows instructions right up until somebody else's instructions arrive. A clever prompt or a poisoned document overrides your rulebook. The controls have to sit outside the model, in a layer the agent can't talk its way past.

How fast does a kill switch need to work?

Fast enough to beat the agent, which means under a second. The lab test shut an agent down on one command, well under that target. 91 percent of companies say they can't stop an agent before it acts, which makes this the control that separates governed from hopeful.

How do you test agent controls without risking customer data?

Run the whole job against an invented company. The FortMesh test used a made-up CISO, modeled on real situations and never contacted, across thirteen steps. Twelve refusals, each with a receipt in the logs. You learn where the walls hold without betting a real customer on it.

Key takeaways

  • Onboard agents like hires: identity, logging, staged access, probation. Day one is read-only.

  • Four rungs, Intern to Principal, with trust earned from logs and always revocable.

  • All five controls sit outside the model, because rules written for the agent fail the first time a poisoned document arrives.

  • The lab test produced twelve refusals with receipts and a kill under one second, against the 91 percent of companies that can't stop an agent at all.

  • Agent debt compounds silently. Budget for the year after launch, not just the launch.

Want to know what rung your agents actually sit on? The free self assessment takes about ten minutes and scores you across all five framework elements.

An intern you can't fire was never really an intern, and an agent you can't stop was never really governed. Start every agent at the bottom of the ladder and make the logs do the promoting.

See where your agents stand.

The free assessment takes ten minutes and scores you on the five elements of the Agentic Trust Framework.