How many AI agents can you run at once before you hit a rate limit?
Unity Horizon8 min read
On the Claude API only uncached input counts toward the limit, so caching lifts one agent's ceiling about 22 times. Checked 6 October 2026.
About three at a time on the Start tier, or about fifty-seven. Same model, same tier, same limits, same agent, and the only variable is whether its prefix is cached. On the Claude API, only input_tokens and cache_creation_input_tokens draw on your input limit. Cache reads do not count at all.
The mechanism is that an agent re-sends its whole history on every turn. Uncached, your input limit is charged the sum of every turn. Cached, it is charged roughly the final context, once.
What counts against an AI agent's rate limit?
Three limits, counted separately for each model class: requests per minute, plus input tokens per minute (ITPM) and output tokens per minute (OTPM) in buckets of their own. The documentation is blunt about what fills the input bucket. "For most Claude models, only uncached input tokens count toward your ITPM rate limits." Tokens written to the cache count at full value. Tokens read from it do not count at all, on every model in the current tables except the retired Claude Haiku 3.5, which carries a footnote saying it does count them.
Two consequences that never make it into a capacity plan. The input_tokens field in the response reports only what sits after your last cache breakpoint, so a 200,000-token cached document with a 50-token question reports 50. And the limiter is a token bucket replenishing continuously rather than resetting on the minute, so a generous per-minute allowance still rejects a burst: the docs warn that 60 requests per minute "might be enforced as 1 request per second."
Checked by Unity Horizon on 6 October 2026.
Why does caching raise an agent's ceiling when trimming lowers it?
Take an ordinary tool-using agent. Estimated, not measured: 12,000 tokens of system prompt and tool definitions, 40 turns, each turn adding about 2,500 tokens of tool result and model reply, each reply about 700 output tokens. Your numbers will differ. The shape will not.
Uncached, turn 1 sends 12,000 input tokens and turn 40 sends 109,500. Summed across the run, one agent session draws 2,430,000 tokens from your input limit.
Cached, with the breakpoint moved to the end of the history each turn, each turn draws only the material added since the last one: 12,000 at the start, then 2,500 a turn. Total 109,500, which is the length of the final context. The whole run costs your input limit what its last request would have cost on its own.
| One 40-turn session | Uncached | Cached |
|---|---|---|
| Input tokens drawn from the per-minute limit | 2,430,000 | 109,500 |
| Average per turn | 60,750 | 2,738 |
The cache saves money on the bill. It saves minutes on the limit.
How many agent sessions does one usage tier support?
The Start tier on Claude Opus 5 is 2,000,000 input tokens per minute, 400,000 output tokens per minute and 1,000 requests per minute. Run the session above through each one.
| Limit | Uncached | Cached |
|---|---|---|
| Turns a minute at 2,000,000 input tokens | 32 | 730 |
| Turns a minute at 400,000 output tokens, 700 a turn | 571 | 571 |
| Turns a minute at 1,000 requests | 1,000 | 1,000 |
| What binds | input | output |
Uncached, input binds at 32 turns a minute while 539 turns of output capacity go unused. Cached, input stops binding and output takes over at 571. Put a turn at roughly 6 seconds of wall clock and a single session makes 10 turns a minute, which is about 3 concurrent sessions in the first column and about 57 in the second.
The ceiling moved eighteen times and nobody raised a limit.
Does the same agent get the same ceiling on Amazon Bedrock?
No, and the advice inverts. Bedrock counts input and output into one TPM bucket, deducts Total input tokens + max_tokens at the start of a request, then settles to InputTokenCount + CacheWriteInputTokenCount + (OutputTokenCount x burndown rate) at the end. Cache reads are excluded there too. The burndown rate is 10x for output on Claude Opus 5.5, Opus 5, Sonnet 5 and Fable 5.1, 15x on Claude 4.8, and 5x on 4.7 and earlier.
One cached turn of our agent is 2,738 input and 700 output. On the Claude API that draws 2,738 from a 2,000,000 bucket and 700 from a 400,000 bucket. The same turn on Bedrock draws 9,738 from a single bucket, and 7,000 of it is the output.
| Claude API | Bedrock bedrock-runtime |
|
|---|---|---|
| Buckets | Input and output per minute, separate, per model | One combined tokens per minute |
| Cache reads | Excluded | Excluded |
| Cache writes | Counted | Counted |
| Output weighting | 1x | 10x on Opus 5.5, Opus 5, Sonnet 5 and Fable 5.1 |
max_tokens |
"no rate limit downside to setting a higher max_tokens value" |
"deducted from your quota at the beginning of each request" |
Read the last row twice. Both quotes are vendor documentation and both describe Claude Opus 5. They instruct you to do opposite things. Anthropic's page says output per minute counts "only the actual tokens generated", so declaring a high ceiling costs nothing. AWS reserves the ceiling you declared and refunds the unused part at the end, which is harmless for the bill and expensive for concurrency: their own worked example reserves 36,000 tokens for a request that settles at 9,000.
Two rows we could not fill in. Google's Gemini rate-limit page defines tokens per minute as input only and says nothing about cached tokens. OpenAI's rate-limit page names TPM as a metric without saying what fills it, and does not mention caching anywhere on the page.
What does a screenshot do to an agent's input limit?
It makes pruning expensive in a second currency. A 1920x1080 screenshot is 2,691 visual tokens on the high-resolution tier used by Claude 4.7 and later models, against 1,560 on the standard tier, because an image costs one visual token per 28 by 28 pixel patch. Declaring computer_toolset_20260801 adds about 4,500 input tokens to every request before any screenshot arrives.
Cached, a computer-use turn draws roughly 3,000 uncached input tokens, which puts it alongside the text agent above. Then somebody prunes the old screenshots to hold the context down. Adding or removing an image anywhere in the prompt changes the message blocks, the message cache goes with them, and a 30-turn history of about 102,000 tokens returns as uncached input: one request eating 5% of a minute's entire input budget on the Start tier. We priced that same invalidation in money in Does trimming an AI agent's context save money?. Against a rate limit it reads worse, because the bill is monthly and the 429 is now.
We built OFFSET, which files declarations directly to the ΑΑΔΕ portal, and a portal agent has exactly this shape: a browser driven one screenshot at a time.
Who does this arrangement suit, and who pays for it?
Then there is the tokenizer. Claude 4.7 and later models use a newer one that "produces approximately 30% more tokens for the same text", so identical traffic draws about 30% more input per minute after that upgrade, with no price change to show for it.
Two kinds of team sit on either side of this line. A team with one fixed system prompt, one tool catalogue, one model and a conversation that only grows at the end keeps nearly all of its input off the limit, and collects the capacity without asking anyone for it. A team that templates a system prompt per customer pays twice: the prefix differs in every session, so it gets written instead of read, and writes count at full value.
A demo runs one session. A rate limit needs a second one.
What should you check before turning on more agents?
Read the response headers rather than the tier table. anthropic-ratelimit-input-tokens-remaining reports what is left, rounded to the nearest thousand, and the console's input chart plots hourly maximum uncached input tokens per minute beside your cache rate. Those two numbers together are the capacity plan.
Do not ramp in one step. Anthropic warns of acceleration limits that return a 429 "if your organization has a sharp increase in usage". OpenAI documents a slow_down error that "can occur even when your traffic is within its requests-per-minute and tokens-per-minute limits", and a rule of thumb with a number in it: past 1 million input tokens per minute, "increase it by no more than 50% every 15 minutes."
Spread the work across models. Limits apply separately for each model, so a verification read sent to Haiku 4.5 draws on a different bucket from the Opus 5 turn it is checking. That is the cheapest capacity you can add without filing a request for more.
Then check what kind of 429 you are reading. A response carrying enforced_spend_limit_reached is the monthly spend cap rather than a rate limit, it sends no retry-after header, and retrying fails until the first of the month. We pulled those two apart in Can you put a hard spending cap on an AI agent?.
Every figure above is arithmetic over published limits and documented counting rules rather than a load test. We have not published a measured 429 rate per thousand agent turns at any tier, and with a token bucket the arrival shape matters as much as the average, so a session sitting well under its input limit across a minute can still be rejected in the second it bursts. If you have that 429 rate from your own fleet with its cache hit rate beside it, we would rather quote your number than our arithmetic.
What share of your agent's input tokens came out of the cache last week, and if nobody on the team can read that figure off the usage page, that is the kind of work we scope.
Common questions
What counts toward an AI agent's rate limit?
On the Claude API, input tokens per minute and output tokens per minute are counted separately for each model class, and only uncached input counts. Tokens written to the cache draw on the input limit at full value. Tokens read from the cache do not count at all, on every model in the current tables except the retired Claude Haiku 3.5.
Does prompt caching increase your rate limit?
It does not raise the limit, it lowers what you draw against it. Because cache reads are excluded from input tokens per minute, a cached agent session draws roughly its final context length once instead of the sum of every turn. On a 40-turn session that is about 110,000 input tokens instead of about 2,430,000.
Should you lower max_tokens to avoid rate limits?
It depends on the platform, and the two rules are opposite. Anthropic's documentation states there is no rate limit downside to a higher max_tokens because output tokens per minute counts only the tokens actually generated. Amazon Bedrock deducts max_tokens from your quota at the start of every request and advises reducing it if you hit quotas earlier than expected.
Why does an AI agent get a 429 when it is under its token limit?
Three causes. The limiter is a token bucket, so a burst inside one second can be rejected while the per-minute average is fine. Acceleration limits reject a sharp increase in traffic regardless of your ceiling. And a 429 carrying enforced_spend_limit_reached is the monthly spend cap, which sends no retry-after header.