FIELD NOTES / 4 MIN READ

How Long Should You Let a Bot's Tool Call Run Before Timing Out?

MCP recommends request timeouts but does not set a universal duration. Clients and other network layers can enforce different limits, so measure the tool, configure the caller you actually use, and use an asynchronous job pattern for work that cannot reliably finish within that budget.

AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. Editorial policy.

Key takeaways

  • The MCP specification recommends request timeouts and per-request configuration, but sets no universal duration.
  • Check the version and configuration of the client that makes the request; do not treat a proposed SDK change as a shipped default.
  • For unpredictable or long-running work, return a job identifier and let the caller check status instead of holding one request open indefinitely.

The short answer

Choose the timeout from the behavior of this tool and the caller that invokes it. Measure normal and slow runs with representative inputs, leave room for expected network variation, and set a per-request limit when the client supports one. A simple lookup and a document-processing job should not inherit one unexplained global number. For work whose slow tail exceeds the caller’s acceptable wait, return a job ID and expose a status check instead of stretching a synchronous call indefinitely.

The Model Context Protocol specification says implementations SHOULD establish timeouts for sent requests and SDKs SHOULD allow per-request configuration. It does not prescribe a duration. The timeout that matters to your bot is the shortest applicable limit in its actual client, transport, gateway, or upstream API.

Sources: Lifecycle - Model Context Protocol.

Which timeout clock actually ends the call?

A request can cross several independent clocks. The MCP TypeScript SDK has been documented with a 60-second request default, while an open issue in the Python SDK repository reports that its client-side requests lacked an equivalent default. Those are version-specific implementation observations, not MCP requirements: check the SDK release and request options that your bot actually runs.

Anthropic’s Claude API documentation addresses a different clock: its SDKs validate that non-streaming Messages API requests are not expected to exceed ten minutes and recommend streaming or Message Batches for long work. That guidance does not set an MCP tool timeout. A client, gateway, or external API can finish waiting before the tool server does, so record which layer timed out instead of assuming one number governs the whole path.

For example, a client configured to wait 60 seconds cannot receive a tool result that arrives after 90 seconds merely because the server would have permitted a longer execution. The server may still have started side effects; a retry should therefore use an idempotency key or check the prior job state where the operation supports it.

Sources: Lifecycle - Model Context Protocol, No default timeout for requests (unlike TS SDK) · Issue #1374, Claude API errors - Claude Platform Docs.

Worked example: a document-processing tool (illustrative)

Say a bot exposes an MCP tool that fetches a file, extracts its contents, normalizes the output, and writes a summary. In this illustrative case, the full job often takes longer than the caller’s configured request timeout. The caller can stop waiting even though the server may have started the work; the precise error and cancellation behavior depend on the client and transport. Neither a completed file nor a failed operation can be inferred from the timeout alone.

Measure actual durations on representative files and record whether the job produced a result or side effect after the caller stopped waiting. If the slow tail is rare and the caller permits per-request configuration, a longer bounded request might be reasonable. If it is common or unpredictable, return a job ID promptly, persist progress and results, and let the bot ask for status before deciding whether to retry.

Progress notifications don't reliably save a slow call

MCP defines a way for a server to signal it's still working: implementations MAY choose to reset the timeout clock when receiving a progress notification corresponding to the request, as this implies that work is actually happening, but implementations SHOULD always enforce a maximum timeout, regardless of progress notifications, to limit the impact of a misbehaving client or server. That second half of the sentence is the part builders miss. Even a tool that sends a progress ping every five seconds can still be killed by a hard ceiling, because the spec only permits resetting on progress; it doesn't require it, and it always allows a hard cap regardless.

Whether your specific client actually resets the clock on progress notifications is an implementation detail you have to check, not assume. This is exactly the kind of gap that generates bug reports: an open proposal filed as SEP-1539 (Timeout Coordination) argues that progress notifications only help after the operation starts and don't solve the information asymmetry about how long an operation should take initially, since a client still needs to guess an appropriate initial timeout before any progress is reported. Treat progress notifications as a UX signal, not a timeout strategy, until you've confirmed your client honors them for that purpose.

Sources: Lifecycle - Model Context Protocol, SEP-1539: Timeout Coordination · Issue #1539.

When to stop raising the timeout and switch to polling

When normal duration is unpredictable, regularly exceeds the caller’s acceptable wait, or depends on an external system whose latency you do not control, a different call shape may be safer. Have the tool return a job ID, persist the job state, and expose a separate status operation the bot can poll. That gives the caller a way to reconcile a timeout with work the server may already have started. Add an idempotency key or equivalent deduplication when retrying could repeat a side effect.

Anthropic gives analogous guidance for long model requests: use the streaming Messages API or Message Batches instead of a long idle non-streaming connection, because networks can drop idle requests and batches can be polled. This is guidance for Anthropic API calls, not a claim that MCP defines a job API. A bot wrapping slow external work can design its own job-ID-and-status tool and document the failure and retry behavior.

Sources: Claude API errors - Claude Platform Docs.

A checklist before you ship a tool call

Run through this before you lock in a number for a new tool, and revisit it whenever the tool's underlying dependency changes.

  • Measure the tool's real duration distribution (not just the happy path) before picking a timeout number.
  • Set the timeout per call or per tool, not as one global value for every tool your bot uses.
  • Confirm whether your specific MCP client actually resets its clock on progress notifications; don't assume it does.
  • For anything with an unpredictable or multi-minute tail, use a job-ID-and-poll pattern instead of a longer synchronous timeout.
  • Log the tool name and the configured duration whenever a timeout fires, so a slow tail shows up as a pattern instead of a one-off mystery.

Primary sources

Sources checked 2026-09-29. Standards and product documentation can change; follow the linked version when implementing.

  1. Lifecycle - Model Context Protocol — Model Context Protocol
  2. No default timeout for requests (unlike TS SDK) · Issue #1374 — Model Context Protocol (GitHub)
  3. Claude API errors - Claude Platform Docs — Anthropic
  4. SEP-1539: Timeout Coordination · Issue #1539 — Model Context Protocol (GitHub)

BotBento is in development. Suggest a correction.