How Should an AI Bot Handle a Tool's Rate Limit Error?
A rate limit is a temporary capacity signal, but it may appear as HTTP 429 or inside a tool result. A bot should identify the layer, honor provider guidance, retry safe operations with bounded backoff, and surface the outcome.
AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. Editorial policy.
Key takeaways
- Classify the failure at its actual layer: HTTP 429, a tool result with isError, and a JSON-RPC protocol error are not interchangeable.
- Honor Retry-After when present; otherwise use provider guidance and bounded exponential backoff with jitter.
- Before retrying a write with an uncertain outcome, check idempotency or read back state. Stop at a budget and tell the user what remains unknown.
First identify where the limit appeared
A bot can encounter a rate limit in more than one place. An HTTP API may return status 429, which means the client sent too many requests in a period. An MCP tool that calls that API may instead return a normal tools/call response whose result has isError set to true and explains that the upstream API refused work. A transport failure, a JSON-RPC error, and a tool execution error have different shapes; treating them as one generic broken connection discards useful recovery information.
The MCP Tools specification distinguishes protocol errors such as unknown tools or invalid arguments from tool execution errors such as upstream API failures. That distinction does not prescribe one universal rate-limit code inside every tool. The bot should inspect the actual response and the connected service's documentation before deciding whether a retry is appropriate. A malformed argument should be fixed; waiting and resending the same invalid request will not help.
Sources: Tools - Model Context Protocol, 429 Too Many Requests - HTTP | MDN.
Use the provider’s wait signal when it exists
HTTP 429 may include Retry-After, which tells a client how long to wait before another request. The header is optional. Some providers document a rate limit without providing an exact reset time, so a bot cannot assume that every 429 means the same delay. The first step is to preserve the status, header, service name, and original operation in the run record.
Docebo's API guidance is a concrete example of a provider without a Retry-After header. It asks clients to use exponential backoff with jitter and a strict retry cap. That policy is specific to Docebo; another service may set a longer wait or prohibit retries in some circumstances. Provider guidance should override a bot's generic schedule when available.
Sources: 429 Too Many Requests - HTTP | MDN, Best practices for handling API rate limits and 429 errors.
Illustrative example: a calendar search pauses mid-task
Imagine a bot fetching available meeting slots through a calendar tool. On the fourth read request, the tool reports that its upstream API returned 429 and provides no wait header. The bot marks the calendar search incomplete. It waits for a bounded interval, retries the same read, and increases the delay if the limit persists. The example is a design exercise, not a report of a real calendar integration or a measured recovery time.
Exponential backoff increases the maximum wait after repeated failures; jitter varies individual waits so many clients do not all retry at once. AWS's engineering explanation uses simulations to show why synchronized backoff can still produce bursts. A practical bot also sets an overall task deadline and an attempt cap, because an unbounded loop can consume the entire time budget without producing a useful answer.
If the provider supplies a Retry-After value, the bot should respect it rather than schedule a retry sooner. If a shared account is rate-limited, parallel bot workers should coordinate: ten independently reasonable retries can still become an unreasonable burst. The bot can tell the user that the calendar result is delayed and either resume later within an authorized task window or ask whether the user wants to stop.
Sources: Exponential Backoff And Jitter, Best practices for handling API rate limits and 429 errors, 429 Too Many Requests - HTTP | MDN.
A write needs a different retry decision
The calendar example is a read. Now imagine the same bot had tried to create a meeting and lost the response. It would be unsafe to assume that the meeting was not created. Blindly retrying could create a duplicate even if the earlier call succeeded. Before another write, the bot should use an idempotency key when the provider supports one, or read the calendar to establish whether the first operation took effect.
A 429 usually signals that the request was rejected, but the bot may receive a transformed error from an intermediary or lose the response after the provider acts. The retry decision therefore depends on the exact operation and the evidence the bot retained. Where the effect cannot be checked, the honest result is an uncertain state and a request for human review, not a silent second attempt.
Sources: 429 Too Many Requests - HTTP | MDN.
Stop at a budget and report what happened
Docebo's published guidance caps attempts to prevent infinite retry loops and returns an error upstream when the limit remains. The exact count in that guidance is a service-specific recommendation, not a universal number for every bot. Set a cap based on the provider's policy, the task deadline, and the user's tolerance for waiting.
When the budget is exhausted, record the tool, the kind of rate-limit signal, the wait and retry decisions, and whether the original action was a read or a write. Tell the user which result is missing and whether any side effect is uncertain. BotBento is still in development; this is a reliability pattern to build and verify, not a claim that BotBento already provides a production run-record feature.
Sources: Best practices for handling API rate limits and 429 errors.
Primary sources
Sources checked 2026-09-25. Standards and product documentation can change; follow the linked version when implementing.
- 429 Too Many Requests - HTTP | MDN — MDN Web Docs
- Tools - Model Context Protocol — Model Context Protocol
- Best practices for handling API rate limits and 429 errors — Docebo
- Exponential Backoff And Jitter — Amazon Web Services
BotBento is in development. Suggest a correction.