4 min readAI Agents · Cybersecurity
How do you test whether an AI agent's rate limits hold?
By Josh Woodruff, Founder & CEO, MassiveScale.AI | Founding Chair, Agentic Trust Framework at the CSAI Foundation

TL;DR: Run a burst test. Push one AI agent to twice its rate limit in a test setting and check four things. The agent gets slowed down. Other agents on the same system keep working. A person gets an alert. The slowed agent backs off and doesn't hammer the door. Four yeses is a pass.
A rate limit is a cap on how many actions an agent can take in a set time. A hundred calls a minute, say.
People move at people speed, so nobody thought much about this for people. An agent can make a week's worth of calls before lunch. The cap is what keeps a small mistake from becoming a large one fast.
Most teams have a rate limit somewhere. The question is whether it's the right kind, and whether it works when pushed.
What can go wrong with a rate limit that exists?
Two things, mostly.
The first is a shared limit. Many setups have one cap at the front door for everybody. When one agent goes wild, it uses up the whole allowance. Every other agent gets blocked along with it. You've turned one agent's problem into an outage.
The second is a silent limit. The agent gets slowed and nobody hears about it. The cap did its job, and the reason the agent was racing is still unknown.
OWASP lists the wider risk as unbounded consumption. It's tenth on its 2025 Top 10 for LLM applications.
How does the assessment score this?
Question 22 of the free assessment asks how agent activity rate is controlled. It belongs to Segmentation, one of the five ATF elements. Each answer passes a different number of the four checks.
Answer | What it says | Checks it passes |
|---|---|---|
A | No rate limiting on agent activity. | None. The agent runs as fast as it can. |
B | Infrastructure-level limits only, such as an API gateway. | The agent is slowed. So is everyone else. |
C | Application-level rate limits per agent. | The agent is slowed, and its neighbors are fine. |
D | Per-agent rate limits with monitoring and alerting. | The first two, plus the alert. |
E | Adaptive rate limiting based on behavior and context. | All four, and the cap tightens when the agent acts oddly. |
How do you run the burst test?
You need two test agents on the same system.
Write down the limit for agent one.
Start agent two on normal, steady work.
Push agent one to twice its limit. Keep it there for five minutes.
Watch both agents and your alert channel.
Then answer four questions.
Check | Question | Pass |
|---|---|---|
Slowed | Did agent one get capped at its limit? | Its extra calls were refused or queued |
Neighbors | Did agent two keep working? | No errors and no slowdown on agent two |
Alert | Did a person hear about it? | An alert named agent one inside five minutes |
Back-off | What did agent one do when refused? | It waited longer between tries |
The pass bar is four yeses.
Why does back-off get its own check?
Because a refused agent that retries at once makes things worse.
Picture an agent that's told no and tries again right away. Each refusal creates another call. The agent is now spending all its effort being refused, and it can drag other systems down with it.
A well-built agent waits longer after each refusal. A capped number of tries is better still. If yours doesn't back off, fix that in the agent's setup. Raising the limit won't help.
This is how a small slowdown turns into a runaway. I cover the automatic stop for that in testing whether a runaway AI agent gets stopped automatically.
Where should you set the rate?
Use the agent's own history.
Look at 30 days of normal work and find its busiest minute. Set the cap somewhat above that. An agent that normally peaks at 40 calls a minute doesn't need room for 4,000.
Rate is one kind of limit. Size is another. An agent can stay under its rate and still take one enormous action, so test that too, with the limit on what an AI agent can do in one action.
Frequently asked questions
Isn't a gateway limit enough?
It protects the system behind the gateway. It doesn't protect your agents from each other. You want a cap for each agent.
Should reads and writes have the same limit?
No. Reads can be generous. Writes change things, so keep them tight.
What should happen to calls over the limit?
Queue the harmless ones and refuse the rest. Either way, log them.
How often should I rerun the test?
Each time you add agents to a shared system. More agents means more ways to crowd each other out.
Key takeaways
A rate limit caps how fast an agent can act.
A shared limit turns one agent's problem into everyone's outage.
Run a burst test at twice the limit and check four things.
A refused agent should wait longer between tries.
Question 22 is one of 30 in the free assessment. It takes about ten minutes and scores you on all five ATF elements.
An agent can do a week of damage before lunch. The cap decides whether it gets the chance.