What Should Your Bot Do When It Hits the Context Window Mid-Task?
A context limit is not one universal failure mode. In Anthropic’s Messages API, oversized input returns an error; newer Claude models can also stop generation at the window boundary. Server-side compaction is a separate, configured beta feature. A bot should inspect the actual response, preserve task constraints outside disposable history, and make any summarization or retry visible in its run record.
AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. Editorial policy.
Key takeaways
- Identify the boundary first: oversized input, output stopped at the limit, and configured compaction have different signals and recovery paths.
- Do not assume an API silently trims system instructions. Anthropic documents an error for oversized input; compaction summarizes earlier conversation only when that feature is configured.
- Keep task constraints in durable state, then record the provider response, token usage, compaction event and recovery decision before continuing.
Identify which context boundary you reached
Editorial correction, 2026-10-08: an earlier version described exactly three outcomes and implied that automatic compaction could silently discard the system prompt. Those statements were too broad. Context handling depends on the API, model, product and configuration. This guide now distinguishes the documented Claude Messages API behavior from Claude Code and from an application’s own history management.
A context window includes the material a model can reference for one response: the system prompt, messages, tool results, tool definitions and generated output all use space. Anthropic’s Context windows guide says an API request whose input alone exceeds the model window returns HTTP 400 with a prompt-is-too-long error. That is a concrete failure signal; it does not mean the API quietly dropped the oldest instructions to make the request fit.
There is another boundary during generation. On Claude 4.5 and newer, a request may be accepted even when input plus the requested maximum output exceeds the window. If generation then reaches the boundary, the response stops with model_context_window_exceeded. Earlier models may reject that request instead. Your code should inspect both HTTP errors and the response stop reason rather than treating any returned text as a complete answer.
Compaction is a separate strategy. Anthropic’s server-side compaction is a beta feature for supported models: on-demand compaction happens when an application asks for it, while threshold compaction runs at a configured token trigger. It replaces older conversation turns with a summary. Claude Code has its own automatic conversation compaction near its limits. Neither fact establishes that an arbitrary Messages API integration compacts by default.
Sources: Context windows - Claude Platform Docs, Compaction overview - Claude Platform Docs, Best practices for Claude Code - Claude Code Docs.
Worked example: a research bot approaching its limit
Consider an illustrative research bot asked to review a set of documents, cite primary sources and stop after 40 pages. It stores a page counter and the source list outside the model conversation. After 20 pages, a large fetched document pushes its next request close to the window. The operator’s first question is not “which text did the model forget?” It is “what did this exact call and configuration do?”
If the next input is too large, the Messages API returns an error. The bot can leave its page counter and source list untouched, shorten the document or split it into smaller excerpts, and retry with the same task boundaries. The failed call did not produce a reviewed page. Recording it as progress would make the final page count wrong.
If the call returns text with model_context_window_exceeded, the bot has a partial generation, not an accepted final research note. It should preserve the partial text as evidence, avoid treating it as a completed citation review, and continue from a smaller context or a fresh bounded request. The stop reason, not the presence of fluent prose, tells the operator that generation ended at the limit.
If the application configured compaction, the bot should record that a summary replaced earlier turns and check whether the citation rule and stopping rule are still represented in durable state. For an on-demand loop, the application chooses when to request a summary and sends the returned block back in place of summarized messages. For threshold compaction, the API runs it when the configured trigger is reached. In both cases, a summary is useful working context, not an audit copy of the original documents.
Sources: Context windows - Claude Platform Docs, Compaction overview - Claude Platform Docs.
Choose recovery by what the task must preserve
For a task with non-negotiable limits, keep those limits in application state and include them in each request that needs them. That might be a source allow-list, a page cap, an approval boundary or the acceptance criteria for the final note. The application should check them again before a tool action or final publication. This avoids relying on a long, disposable conversation as the sole record of an instruction.
For a conversation that can tolerate a summary, consider configured compaction before the window is exhausted. Anthropic’s compaction overview distinguishes on-demand, threshold and client-side summarization. On-demand gives the application control over when to request and reuse the summary; threshold lets the API trigger it at a set size. Choose a mode deliberately and test it with a long fixture that contains an early constraint you expect to retain.
For a batch of large tool outputs, another option is to omit or clear old outputs while retaining a durable pointer to their source. Anthropic documents context editing for old tool results and thinking blocks separately from compaction. What matters is that the bot can still retrieve evidence when asked, and that a dropped tool result does not become an invented citation in the final answer.
Do not turn a context error into an unbounded retry loop. Set a bounded retry count, reduce the input or change the work unit, and stop for review if the next attempt still cannot fit. A shorter request that loses the user’s actual task is not a recovery. A good recovery preserves the decision the bot was hired to make, even when the raw conversation must be shortened.
Sources: Context windows - Claude Platform Docs, Compaction overview - Claude Platform Docs.
Record enough evidence to explain a partial run
A useful run record has the model and endpoint, the request’s estimated and reported token usage, the HTTP status or stop reason, and the handling mode the application configured. If compaction or client-side trimming ran, record its trigger, the span of history replaced and a pointer to the retained source material. Do not copy private document contents into analytics just to make the trace detailed.
Separate three outcomes in the product state: rejected request, partial generation and completed answer. A rejected request contributes no reviewed pages in the example. A partial generation may be kept for diagnosis, but it still needs a continuation or explicit review. A completed answer passes the citation and page-cap checks before it can be published. These are application rules, not promises that the model or provider will enforce them for you.
Finally, test the behavior of the actual product your bot uses. Claude Code automatically compacts its conversation, while the Messages API documents explicit compaction modes and context-limit signals. A bot that calls another provider, gateway or SDK may expose different errors and settings. Reproduce a near-limit run in a disposable fixture, inspect the response, and update your runbook from that evidence rather than assuming one universal overflow rule.
- On input rejection: save the error and reduce or split the next request.
- On output boundary: mark the answer partial using the returned stop reason.
- On compaction: record the trigger and retain a durable pointer to original evidence.
Sources: Context windows - Claude Platform Docs, Compaction overview - Claude Platform Docs, Best practices for Claude Code - Claude Code Docs.
Primary sources
Sources checked 2026-10-08. Standards and product documentation can change; follow the linked version when implementing.
- Context windows - Claude Platform Docs — Anthropic
- Compaction overview - Claude Platform Docs — Anthropic
- Best practices for Claude Code - Claude Code Docs — Anthropic
BotBento is in development. Suggest a correction.