Docs/Start/Rate limits

Rate limits

Per-plan request ceilings, the headers that report them, and how to back off correctly — including the two very different failures that both answer 429.

View as Markdown

Two separate ceilings apply to every account, and they fail in the same way but mean opposite things:

  • A rate limit — how many calls per minute. Hitting it means slow down; the call would have succeeded a moment later.
  • A credit allowance — how many calls per billing cycle. Running out means the month is spent; retrying will not help until you top up or the cycle renews.

Both arrive as a failed call. Telling them apart is the single most useful thing on this page, and it is covered under which limit did you hit.

The per-minute ceiling

PlanRate limitConcurrent callsCredits / month
Free5/min1200
Starter60/min5200,000
Pro180/min201,000,000
MegaNo limit504,000,000

Rate limit is calls per minute, measured per API key. Concurrent calls is a separate ceiling — how many may be in flight at the same moment — and it is the one that bites first when you fan out. Ten parallel workers on a Starter key trip its concurrency ceiling of 5 long before they get near its 60/min limit.

Where the table reads No limit, there is no per-minute ceiling at all. That is not the same as unlimited: the credit allowance and the concurrency ceiling still apply.

A rate limit is a shape, not a total

A per-minute limit is not a bucket you empty and then wait on. The limiter refills continuously, so a steady stream spread across the minute never trips a ceiling that the same number of calls fired in one second usually does. Spacing calls evenly is worth more than counting them.

Your own lower limit

A key can be capped below its plan's ceiling from the dashboard, under the key's restrictions. That is worth doing for any key you have handed to someone else: it caps the blast radius of a runaway loop at a number you chose, instead of at your plan's.

The effective limit is the lowest of the plan ceiling, your own setting, and any cap support has applied to the account.

Which limit did you hit

Both arrive at the agent as a failed tool call with a readable reason, and they need opposite responses. "Rate limit exceeded" is transient — the next call a minute later succeeds. "Monthly credit limit reached" is not, and an agent that retries it will keep retrying until something stops it.

That asymmetry is the argument for a per-agent sub-key: its own rate window means one enthusiastic agent cannot starve the rest, and its own line in analytics means you can see which one it was.

Backing off correctly

Wait, then try again once. That is the whole technique.

Retrying aggressively makes it last longer

When a key is rate-limited, the edge remembers it and holds that key in a short cooldown — around 50 seconds. Hammering during the cooldown does not shorten it. Backing off properly usually gets you through in one wait; retrying in a tight loop can stay locked out for minutes.

Repeated identical failures

There is a third throttle that has nothing to do with volume. If the same call keeps failing the same way — a malformed parameter, a missing required input — that exact call is throttled and the original error is replayed back to you.

This is a guard against a loop retrying a permanently broken call forever, and it costs nothing to avoid: an input the source rejected has to change before it is worth sending again. See errors.

Staying under the ceiling

  • Narrow the tool list. An agent offered fifty tools tries more of them than one offered five, and each attempt is a call.
  • Give each workload its own key. A sub-key with its own limit means a backfill job cannot starve your live traffic. See sub-keys.

Next

Every other way a call can fail, and which are worth retrying, is in errors. For what a call costs against your allowance rather than your rate, see credits and billing.

Was this page helpful?

Last updated