# BotBento > A pre-release workspace for people and AI bots, powered by Hermes. BotBento is in development. Its public site provides early-access registration, a separate newsletter signup, contact and sourced editorial articles. Article examples do not imply shipped BotBento capabilities. AI-assisted articles disclose their method and link primary sources. ## Pages - [For readers and AI assistants — BotBento](https://botbento.com/agents/) — [Markdown](https://botbento.com/text/agents/index.md) - [How Do You Stop a Retried Tool Call From Running Twice? — BotBento](https://botbento.com/blog/agent-tool-retry-idempotency-keys/) — [Markdown](https://botbento.com/text/blog/agent-tool-retry-idempotency-keys/index.md) - [What should an AI agent record after each run? — BotBento](https://botbento.com/blog/ai-agent-run-result-record/) — [Markdown](https://botbento.com/text/blog/ai-agent-run-result-record/index.md) - [What belongs in an AI agent stopping rule? — BotBento](https://botbento.com/blog/ai-agent-stopping-rules/) — [Markdown](https://botbento.com/text/blog/ai-agent-stopping-rules/index.md) - [AI agents vs workflows: choose who decides the next step — BotBento](https://botbento.com/blog/ai-agents-vs-workflows/) — [Markdown](https://botbento.com/text/blog/ai-agents-vs-workflows/index.md) - [How Should an AI Bot Handle a Tool's Rate Limit Error? — BotBento](https://botbento.com/blog/ai-bot-rate-limit-error-handling/) — [Markdown](https://botbento.com/text/blog/ai-bot-rate-limit-error-handling/index.md) - [How Do You Scope an API Key So a Bot Can't Exceed Its Job? — BotBento](https://botbento.com/blog/bot-api-key-scoping-least-privilege/) — [Markdown](https://botbento.com/text/blog/bot-api-key-scoping-least-privilege/index.md) - [What Should a Bot Do When a Multi-Step Task Fails Halfway? — BotBento](https://botbento.com/blog/bot-multi-step-task-failure-rollback/) — [Markdown](https://botbento.com/text/blog/bot-multi-step-task-failure-rollback/index.md) - [How Do You Know Your Bot's Recurring Routine Has Gone Silent? — BotBento](https://botbento.com/blog/bot-routine-silent-failure-detection/) — [Markdown](https://botbento.com/text/blog/bot-routine-silent-failure-detection/index.md) - [Field notes for the agent world — BotBento](https://botbento.com/blog/) — [Markdown](https://botbento.com/text/blog/index.md) - [Local or cloud: where should your AI agent run? — BotBento](https://botbento.com/blog/local-vs-cloud-ai-agents/) — [Markdown](https://botbento.com/text/blog/local-vs-cloud-ai-agents/index.md) - [What Can an MCP Tool Ask You For Mid-Task, and What Can't It? — BotBento](https://botbento.com/blog/mcp-elicitation-mid-task-permissions/) — [Markdown](https://botbento.com/text/blog/mcp-elicitation-mid-task-permissions/index.md) - [MCP Resource Indicators: What Your Bot's OAuth Tokens Must Specify — BotBento](https://botbento.com/blog/mcp-resource-indicator-requirement/) — [Markdown](https://botbento.com/text/blog/mcp-resource-indicator-requirement/index.md) - [MCP Roots Is Deprecated. Where Should Your Bot's File Boundaries Live Now? — BotBento](https://botbento.com/blog/mcp-roots-deprecated-migration/) — [Markdown](https://botbento.com/text/blog/mcp-roots-deprecated-migration/index.md) - [MCP Sampling Is Deprecated. What Should Bot Builders Do Now? — BotBento](https://botbento.com/blog/mcp-sampling-deprecated/) — [Markdown](https://botbento.com/text/blog/mcp-sampling-deprecated/index.md) - [MCP Dropped Sessions. How Do You Keep State Across Tool Calls Now? — BotBento](https://botbento.com/blog/mcp-stateless-cross-call-state/) — [Markdown](https://botbento.com/text/blog/mcp-stateless-cross-call-state/index.md) - [What Actually Happens When Your Bot Cancels an MCP Tool Call? — BotBento](https://botbento.com/blog/mcp-tool-call-cancellation-what-happens/) — [Markdown](https://botbento.com/text/blog/mcp-tool-call-cancellation-what-happens/index.md) - [MCP Tools Can Declare an Output Schema. Should Your Bot Validate It? — BotBento](https://botbento.com/blog/mcp-tool-output-schema-validation/) — [Markdown](https://botbento.com/text/blog/mcp-tool-output-schema-validation/index.md) - [MCP tool permissions: what to check before connecting a bot — BotBento](https://botbento.com/blog/mcp-tool-permissions/) — [Markdown](https://botbento.com/text/blog/mcp-tool-permissions/index.md) - [MCP Tool Results: How Do You Know a Call Actually Succeeded? — BotBento](https://botbento.com/blog/mcp-tool-result-success-signal/) — [Markdown](https://botbento.com/text/blog/mcp-tool-result-success-signal/index.md) - [Why we are building BotBento around people, not just bots — BotBento](https://botbento.com/blog/people-and-bots-in-one-room/) — [Markdown](https://botbento.com/text/blog/people-and-bots-in-one-room/index.md) - [Plugin or authorized connection: what's actually different? — BotBento](https://botbento.com/blog/plugin-vs-authorized-connection/) — [Markdown](https://botbento.com/text/blog/plugin-vs-authorized-connection/index.md) - [How Do You Review an AI-Generated Research Note? — BotBento](https://botbento.com/blog/reviewing-ai-generated-research-notes/) — [Markdown](https://botbento.com/text/blog/reviewing-ai-generated-research-notes/index.md) - [How Do You Actually Revoke a Bot's Access to a Tool It No Longer Uses? — BotBento](https://botbento.com/blog/revoke-bot-tool-access/) — [Markdown](https://botbento.com/text/blog/revoke-bot-tool-access/index.md) - [When Should a Bot Use a Supervisor Agent Instead of One Agent With Tools? — BotBento](https://botbento.com/blog/supervisor-agent-vs-single-agent-tools/) — [Markdown](https://botbento.com/text/blog/supervisor-agent-vs-single-agent-tools/index.md) - [How do you test an AI agent that uses a calendar? — BotBento](https://botbento.com/blog/test-calendar-ai-agent/) — [Markdown](https://botbento.com/text/blog/test-calendar-ai-agent/index.md) - [How Long Should You Let a Bot's Tool Call Run Before Timing Out? — BotBento](https://botbento.com/blog/tool-call-timeout-duration/) — [Markdown](https://botbento.com/text/blog/tool-call-timeout-duration/index.md) - [How Do You Verify an AI Agent Actually Finished the Task? — BotBento](https://botbento.com/blog/verify-ai-agent-task-completion/) — [Markdown](https://botbento.com/text/blog/verify-ai-agent-task-completion/index.md) - [Bot templates for BotBento](https://botbento.com/bots/) — [Markdown](https://botbento.com/text/bots/index.md) - [Contact BotBento](https://botbento.com/contact/) — [Markdown](https://botbento.com/text/contact/index.md) - [Editorial policy — BotBento](https://botbento.com/editorial/) — [Markdown](https://botbento.com/text/editorial/index.md) - [BotBento | People and AI bots, working together](https://botbento.com/) — [Markdown](https://botbento.com/text/index.md) - [A good week for a little progress — BotBento](https://botbento.com/newsletters/) — [Markdown](https://botbento.com/text/newsletters/index.md) - [The launch edition: a place to start with AI bots — BotBento](https://botbento.com/newsletters/launch-edition-a-place-to-start/) — [Markdown](https://botbento.com/text/newsletters/launch-edition-a-place-to-start/index.md) - [Weekly field notes: make an agent’s next step checkable — BotBento](https://botbento.com/newsletters/weekly-field-notes-2026-09-11/) — [Markdown](https://botbento.com/text/newsletters/weekly-field-notes-2026-09-11/index.md) - [Weekly field notes: recover a bot without guessing — BotBento](https://botbento.com/newsletters/weekly-field-notes-2026-09-25/) — [Markdown](https://botbento.com/text/newsletters/weekly-field-notes-2026-09-25/index.md) - [Plugins for BotBento bots](https://botbento.com/plugins/) — [Markdown](https://botbento.com/text/plugins/index.md) - [Privacy — BotBento](https://botbento.com/privacy/) — [Markdown](https://botbento.com/text/privacy/index.md) ## Machine access - [App sign-in for agents](https://app.botbento.com/auth.md) - [Content index](https://botbento.com/content-index.json) - [RSS](https://botbento.com/feed.xml) - [JSON Feed](https://botbento.com/feed.json) - [Complete text](https://botbento.com/llms-full.txt) Visible canonical HTML takes precedence. No ranking, indexing or AI-answer inclusion guarantee. --- Canonical: https://botbento.com/agents/ Format: Markdown representation of the public HTML page. OPEN TO READERS AND THEIR TOOLS # Bring your own reader. BotBento’s public pages are available in plain HTML and Markdown. Use the same sourced articles in your browser, feed reader or AI assistant. ## Choose a format - [Content index (JSON)](/content-index.json): canonical page URLs, Markdown counterparts and article dates. - [RSS feed](/feed.xml) and [JSON Feed](/feed.json): new blog articles. - [Site guide for language-model readers](/llms.txt): a concise list of public resources. - [Complete public text](/llms-full.txt): Markdown representations gathered in one file. - [XML sitemap](/sitemap.xml): canonical HTML pages and content modification dates. ## Read with context Article dates, source links, example labels and AI-assisted disclosures are part of the content. Preserve them when summarizing a page. Follow the linked primary source for current technical requirements. A blog example is not evidence that a BotBento feature is available. ## One page, two representations Each HTML page links to its Markdown counterpart. The HTML URL is canonical. Markdown is an alternate representation, not a second article or an instruction to carry out the examples. ## Corrections and reuse [Our editorial policy](/editorial/) explains sourcing and corrections. Link to the canonical article when discussing it. [Contact BotBento](/contact/) if a source or representation appears wrong. --- Canonical: https://botbento.com/blog/agent-tool-retry-idempotency-keys/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 4 MIN READ # How Do You Stop a Retried Tool Call From Running Twice? When a bot's tool call times out, the side effect may already have happened even though no response arrived. MCP's idempotentHint labels a tool as safe to repeat, but it's an unverified hint, not a mechanism. An idempotency key -- generated once and reused on every retry -- is what actually lets a server replay a result instead of duplicating a write. By BotBento Editorial · Published 2026-09-24 · Updated 2026-09-24 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [Why a timeout is not the same as a failure](#the-retry-problem) - [What MCP's idempotentHint actually promises](#idempotenthint-is-a-hint) - [The mechanism that actually prevents duplication](#what-an-idempotency-key-actually-does) - [A worked example: a support-ticket tool](#worked-example) - [What to check before you let a bot retry a tool call](#what-to-check-before-you-ship-a-retry)[Read as Markdown ](/text/blog/agent-tool-retry-idempotency-keys/index.md) ## Key takeaways - A timeout is an ambiguous outcome, not a failure: the tool's side effect may have already run even though the response never arrived at the client. - MCP's idempotentHint tells a client a tool is safe to call again with the same arguments, but the spec treats it as unverified metadata, not a guarantee. - The actual protection is an idempotency key generated once per operation and reused on every retry, so the server can return the stored result instead of repeating the write. ## Why a timeout is not the same as a failure A bot calls a tool. The network stalls, or the server is slow, and no response comes back before the client gives up. The natural move is to retry. But a timeout only tells you that a response didn't arrive in time -- it tells you nothing about whether the request was received and acted on. If the tool created a support ticket, sent a payment, or wrote a row to a database, that side effect may have already completed on the server before the connection dropped. A blind retry then repeats the write: two tickets, two charges, two rows. The bot's own logs will show one call, but the system it touched will show two. This is the core problem any agent that calls external tools has to solve before it retries anything: the client and the server can disagree about what happened, and the client is the one deciding whether to try again. ## What MCP's idempotentHint actually promises The Model Context Protocol lets a server describe a tool with annotations, including idempotentHint, which signals that calling the tool repeatedly with the same arguments has no additional effect beyond the first successful call. On the surface this looks like exactly the guarantee a retrying client needs. It isn't. The specification is explicit that these annotations are hints, not enforced behavior, and clients should not rely on them without additional verification, especially when the server is untrusted. A tool author can mark a tool idempotentHint: true and still implement it in a way that creates a duplicate record on a second call -- the annotation describes intent, not a contract the runtime checks. For a bot developer this matters because it shifts the responsibility back onto the client and the tool's actual implementation. Reading idempotentHint on a tool definition tells you what the server author believes about their own tool. It does not tell you what will actually happen if your agent calls it twice after a timeout. Sources: [Tools](https://modelcontextprotocol.io/specification/2025-06-18/server/tools). ## The mechanism that actually prevents duplication The pattern that provides a real guarantee, used across payment and infrastructure APIs, is an idempotency key. The client generates a unique value once, before the first attempt, and sends that same value with every retry of the same logical operation. The server stores the key alongside the result of the first successful execution. If a request arrives with a key it has already processed, the server returns the stored result instead of running the operation again. This works because the server -- not the client -- is the one deciding whether an operation is new or a repeat. The client's job is only to keep reusing the same key across retries of the same intended action, and to generate a fresh key for a genuinely new action. HTTP itself distinguishes methods that are safe to retry from those that aren't: GET is defined as safe and idempotent by nature, while POST is defined as neither, which is exactly why POST-style tool calls -- creating a ticket, charging a card, sending a message -- are the ones that need an explicit idempotency key rather than relying on the method itself. None of this is something a client can verify from outside. Whether a tool honors an idempotency key, and whether it stores keys long enough to matter for your retry window, are implementation details of the server. If the tool's documentation doesn't mention idempotency keys, assume retries can duplicate the effect. Sources: [Idempotent requests](https://docs.stripe.com/api/idempotent_requests), [HTTP Semantics](https://www.rfc-editor.org/rfc/rfc9110). ## A worked example: a support-ticket tool Illustrative example. A personal bot has a tool called create\_ticket that takes a title and description and files a support ticket. The bot calls it, the connection times out after 8 seconds with no response, and the bot's retry logic fires. Without an idempotency key: the bot calls create\_ticket again with the same title and description. If the first call actually succeeded server-side before the timeout, there are now two tickets. The bot has no way to tell, from its own side, that this happened -- both calls looked identical to it. With an idempotency key: before the first call, the bot generates a key, for example a random string tied to this specific ticket-filing intent, and sends it as part of the request. On retry, it sends the exact same key. If the server implements idempotency keys, it recognizes the second request as a repeat of the first and returns the original ticket ID instead of filing a new one. The bot ends up with one ticket either way. The difference between these two outcomes is entirely in whether the tool's server side was built to store and check keys -- the bot's retry code looks almost the same in both cases. ## What to check before you let a bot retry a tool call Before adding automatic retries to any tool call that writes, sends, or charges something, check these points rather than assuming idempotentHint or a timeout policy covers you. - Does the tool's documentation mention an idempotency key or a similar request-deduplication parameter? If not, treat retries as unsafe by default. - If idempotentHint is set on an MCP tool, confirm what backs it -- a documented key mechanism, or just the author's description of intended behavior. - For any tool that has a side effect outside the conversation (a write, a send, a charge), generate one key per intended action and reuse it across every retry of that action, never generating a new key just because the previous attempt timed out. - Log the key alongside the action so a human reviewing the bot's run can tell a genuine repeat attempt from two separate, intended actions. - If the tool gives no way to pass a key, prefer surfacing the timeout to a human before retrying, rather than guessing. ## Primary sources Sources checked 2026-09-24. Standards and product documentation can change; follow the linked version when implementing. - [Tools](https://modelcontextprotocol.io/specification/2025-06-18/server/tools) — Model Context Protocol - [Idempotent requests](https://docs.stripe.com/api/idempotent_requests) — Stripe - [HTTP Semantics](https://www.rfc-editor.org/rfc/rfc9110) — IETF BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/ai-agent-run-result-record/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 5 MIN READ # What should an AI agent record after each run? An AI agent run record should connect the requested outcome to what was actually verified. Record the run identity, task scope, outcome evidence, unresolved actions and the next safe step. Keep operational detail separate from the short result a person reads. By BotBento Editorial · Published 2026-09-08 · Updated 2026-09-08 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [Start with the result someone needs to check](#start-with-the-result) - [Give the run and its attempts distinct identities](#identify-the-run) - [An illustrative record for a partially completed routine](#small-record) - [Treat an unknown external effect as a recovery decision](#retry-unknown-actions) - [Keep the record useful without copying private content](#minimum-useful-data) - [Review the record before putting the routine on a schedule](#review-before-scheduling)[Read as Markdown ](/text/blog/ai-agent-run-result-record/index.md) ## Key takeaways - A completed tool call is evidence about one operation; the routine still needs an explicit check of the user's requested outcome. - Distinguish incomplete work from actions whose effects are unknown before deciding whether to retry. - Store references and useful diagnostics without copying credentials, private messages or unnecessary personal data into the run history. ## Start with the result someone needs to check Imagine a morning routine that reads three approved project sources, writes a short research note and posts a link to a team channel. A message saying the run succeeded leaves several questions unanswered. Were all three sources available? Where is the saved note? Did the channel receive the link, or did the routine only prepare a message? Those distinctions determine what a teammate should do next. Our proposed result record starts with the task's acceptance conditions. For this example, success means that the three sources were checked, the note was saved in the intended workspace and the channel post was confirmed. A readable draft with one unavailable source is still useful, but it does not meet that complete definition. Record the useful output and the missing condition together. This is a design proposal for a routine you control. It is not a claim that BotBento currently ships this record format, and the example below is illustrative rather than a measured production run. ## Give the run and its attempts distinct identities OpenTelemetry describes a trace as connected operations and a span as a unit of work with timing and identifying context. Child spans can represent sub-operations. These concepts help explain which operation happened inside a larger run, but a trace does not define your business acceptance conditions for you. For a small routine, keep a stable logical run ID and identify individual attempts separately. Suppose the scheduled research note is run research-2026-09-08-am. A second attempt to retrieve a source should remain associated with that run. It should not look like a second successful morning publication in your daily count. Record the routine version and the intended destination alongside the run ID. Use internal references where appropriate: a workspace identifier is usually more useful for diagnosing a wrong destination than the bot's display name. If you already collect traces, link the run record to its trace ID instead of embedding every diagnostic event in the human-facing summary. Sources: [Traces](https://opentelemetry.io/docs/concepts/signals/traces/). ## An illustrative record for a partially completed routine Here is a deliberately small example. The field names are our suggested application contract, not OpenTelemetry standard attributes. In a real system, references would point to access-controlled records that the reviewer is permitted to open. None of the identifiers below represent a real person, workspace or delivery. The useful distinction is between the known output and the unresolved action. Someone can read the saved note immediately. They can also see why replaying the entire routine would be a poor recovery instruction: it could create a second note and potentially a second channel post. - Run: research-2026-09-08-am; routine version: 3; attempt: 1. - Requested scope: sources A, B and C; save one note; post its link to channel R. - Observed result: sources A and B checked; C unavailable; draft note N saved and read back. - External action: channel post request timed out; whether a post exists is unresolved. - Outcome: partial; acceptance conditions not met; note N contains a visible missing-source notice. - Next step: inspect the channel operation before retrying it; recover source C separately; owner: the routine operator. ## Treat an unknown external effect as a recovery decision A timeout alone does not tell the caller whether a remote action happened. AWS's Builders' Library explains how a caller-provided request identifier can let a service recognize repeated intent and handle retries without duplicating the operation. The protection depends on the service's actual idempotency contract; writing a unique ID into your own log is not enough. For the illustrative channel post, first look for a provider operation identifier or another supported way to inspect the result. If the provider supports idempotent retries, follow its rules about identifiers, request contents and retention windows. Do not generate a fresh action identity just because the original response was missing. Your record should distinguish confirmed failure from an unresolved effect. We suggest separate fields for the observed error and the recovery decision: for example, timeout and inspect before retry. This makes an unattended worker's stopping point understandable. When the provider offers neither safe inspection nor a supported retry contract, leave the action for review instead of presenting a guessed delivery result. Sources: [Making retries safe with idempotent APIs](https://aws.amazon.com/builders-library/making-retries-safe-with-idempotent-APIs/). ## Keep the record useful without copying private content OpenTelemetry's sensitive-data guidance recommends collecting only what serves an observability purpose and reviewing instrumentation for information it may expose. That applies to the convenient fields in an agent run history as much as to a tracing backend. Authentication credentials, session tokens and private message bodies should not become routine diagnostic attachments. For our research example, store a reference to note N and a short error category for source C. Avoid copying the complete source response into every run record. Keep detailed evidence in its appropriate storage location, with suitable access and retention, and make the summary useful even when a reader cannot open that evidence. Decide which record fields are allowed before collecting them. A small list of operational fields is easier to inspect than an unrestricted dump of tool arguments and responses. Review any linked artifact too: a carefully redacted summary can still expose private content through an unrestricted evidence link. Sources: [Handling sensitive data](https://opentelemetry.io/docs/security/handling-sensitive-data/). ## Review the record before putting the routine on a schedule Walk through the example with three deliberately different outcomes: everything completes, a source is unavailable, and a write request returns no clear result. For each one, ask another person to identify the usable output, the unverified condition and the next action from the result record alone. If they need the original chat to understand the outcome, the record is missing context. Then consider a second attempt. Does it preserve the relationship to the original run? Can a reviewer tell whether the note was revised or a new copy was created? Does the record keep an earlier unknown action visible until evidence resolves it? These questions test the usefulness of the design without pretending that a tidy status label establishes completion. Start with a short record that answers those questions. Add detail when a real recovery need justifies it. The goal is a routine whose result can be checked and whose next step can be chosen without reconstructing the entire conversation. ## Primary sources Sources checked 2026-09-08. Standards and product documentation can change; follow the linked version when implementing. - [Traces](https://opentelemetry.io/docs/concepts/signals/traces/) — OpenTelemetry - [Making retries safe with idempotent APIs](https://aws.amazon.com/builders-library/making-retries-safe-with-idempotent-APIs/) — Amazon Web Services - [Handling sensitive data](https://opentelemetry.io/docs/security/handling-sensitive-data/) — OpenTelemetry BotBento is in development. [Suggest a correction](/contact/). ## Keep reading - [AI agents vs workflows: choose who decides the next step](/blog/ai-agents-vs-workflows/) - [MCP tool permissions: what to check before connecting a bot](/blog/mcp-tool-permissions/) --- Canonical: https://botbento.com/blog/ai-agent-stopping-rules/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 4 MIN READ # What belongs in an AI agent stopping rule? An AI agent stopping rule should distinguish verified completion, a required handoff, a resource limit and a lack of useful progress. Define the evidence for each condition, the component that enforces it and what happens to unfinished work. Reaching a limit ends an attempt; it does not establish that the requested task succeeded. By BotBento Editorial · Published 2026-09-10 · Updated 2026-09-10 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [Give completion and limits different meanings](#completion-and-limits) - [Write four explicit decisions for one task](#four-decisions) - [Match each limit to something the runtime measures](#enforce-limits) - [Inspect a loop before raising its allowance](#detect-no-progress) - [Test the exits and the next attempt together](#test-the-exits)[Read as Markdown ](/text/blog/ai-agent-stopping-rules/index.md) ## Key takeaways - Separate the condition that makes a task complete from the limits that prevent an attempt from continuing indefinitely. - Define what counts as progress and which missing information requires a handoff before running the agent unattended. - Test each stopping condition alone, including its saved output and recovery path, rather than relying on a final success message. ## Give completion and limits different meanings An instruction to keep working until the answer is good leaves two decisions undefined: what makes the answer acceptable, and what ends an unsuccessful attempt. A useful stopping rule answers both. It should also explain when the agent needs something only a person or another service can provide. Anthropic's Building effective agents describes agents using feedback from their environment, pausing for human input and operating with stopping conditions such as iteration limits. The article was published in December 2024 and now carries a tooling-change notice. We use its general design principle here, not its examples as evidence of a current product release. Consider our fictional research assistant, asked to compare the export formats of three named tools and save one note. Completion requires an answer for each tool supported by its official documentation, or a visible unresolved field where that documentation provides no answer. The note must be saved and read back. Merely reaching the end of a conversation does not satisfy those conditions. This is an illustrative design, not a measured BotBento run; BotBento remains in development. Sources: [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents). ## Write four explicit decisions for one task For the research assistant, we propose the following four decisions. They are an application contract, not names required by a framework. Write them before choosing a message count or timeout so that the numbers serve a defined task. For example, an official page that says nothing about an export format leaves a useful unknown. The assistant can mark that field unresolved and finish the agreed comparison. A missing destination workspace is different: without knowing where to save the note, it cannot satisfy the requested delivery. That should trigger a handoff rather than a guessed destination. - Complete: every requested tool has a cited answer or a clearly marked unknown, and the note exists in the agreed destination. - Needs input: an unresolved choice would change the task's scope, authority or destination; ask the specific question and preserve the draft. - Limit reached: the configured attempt allowance is exhausted; stop starting new work and report the unmet acceptance conditions. - No useful progress: repeated steps produce no new relevant evidence or resolution; explain the blockage instead of silently repeating them. ## Match each limit to something the runtime measures AutoGen's AgentChat documentation offers termination conditions for message count, reported token usage, elapsed duration, handoffs and external control. It also allows conditions to be combined with AND or OR. Token-based termination depends on agents reporting their usage. These are examples of available mechanisms, not interchangeable definitions of task completion. In our proposed research routine, count source retrieval attempts separately from model turns. One turn might request several pages, while another only summarizes existing evidence. If the operator cares about elapsed time or provider spending, a turn limit alone is insufficient evidence that those separate allowances are enforced. Define how the relevant measurement is collected and who owns it. Review the combination logic as well. If either a cancellation request or an exhausted allowance should stop new work, combining those conditions so that both must be true would violate that intention. Test the actual behavior at each boundary. Also define what a stop means for work already in flight: stop scheduling, cancellation requested and remote action confirmed cancelled are different observations. Do not promise cancellation of an external action without a supported mechanism and a checked result. Sources: [Termination](https://microsoft.github.io/autogen/stable/user-guide/agentchat-user-guide/tutorial/termination.html). ## Inspect a loop before raising its allowance LangGraph documents GRAPH\_RECURSION\_LIMIT as reaching the maximum number of steps before a stopping condition. Its troubleshooting guidance distinguishes an unintended cycle from a complex graph that legitimately needs more steps. Increasing the configured limit is therefore a possible adjustment, not a diagnosis of why the previous attempt stopped. For our research example, imagine the assistant repeatedly opening the same overview page even though the missing answer is absent. A useful progress check asks whether a step supplied a new relevant source, resolved a field or saved a reviewed result. Simply restating the plan should not count. This is our suggested evaluation rule, and it needs tuning against the actual task. Record the last useful change and the repeated action, then inspect why the agent chose it. Perhaps the tool result was not available to the next step, or the completion condition could never be satisfied. Raising the allowance before understanding that behavior can leave the same unresolved condition after a longer run. A larger legitimate task may need more room, but that decision should follow evidence about the work remaining. Sources: [GRAPH\_RECURSION\_LIMIT](https://docs.langchain.com/oss/python/langgraph/errors/GRAPH_RECURSION_LIMIT). ## Test the exits and the next attempt together Use disposable inputs to exercise four cases: all sources answer the question, a required destination is missing, the attempt allowance expires, and the same source keeps returning no useful answer. For each case, inspect why the agent stopped, what was saved and which acceptance conditions remain open. These are proposed tests; no pass rates or production outcomes are claimed here. Then resume from the retained state. Check that the assistant reuses the existing draft, preserves unresolved external effects and does not reset the logical task's allowances accidentally. An operator may authorize another attempt, but the new attempt should remain connected to the earlier result. Otherwise, a per-attempt limit can hide an indefinitely repeated routine. Keep the final message short and actionable: where the output is, which condition ended the attempt and what would allow the remaining work to continue. A stopping rule is useful when another person can make that next decision without reconstructing every tool call. ## Primary sources Sources checked 2026-09-10. Standards and product documentation can change; follow the linked version when implementing. - [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) — Anthropic - [Termination](https://microsoft.github.io/autogen/stable/user-guide/agentchat-user-guide/tutorial/termination.html) — Microsoft AutoGen - [GRAPH\_RECURSION\_LIMIT](https://docs.langchain.com/oss/python/langgraph/errors/GRAPH_RECURSION_LIMIT) — LangChain BotBento is in development. [Suggest a correction](/contact/). ## Keep reading - [What should an AI agent record after each run?](/blog/ai-agent-run-result-record/) - [AI agents vs workflows: choose who decides the next step](/blog/ai-agents-vs-workflows/) --- Canonical: https://botbento.com/blog/ai-agents-vs-workflows/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 5 MIN READ # AI agents vs workflows: choose who decides the next step Use a workflow when you know the sequence. Give an agent discretion when discovering the sequence is part of the work. Then test the result, the cost and the boundaries separately. By BotBento Editorial · Published 2026-09-07 · Updated 2026-09-07 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [Who decides what happens next?](#the-distinction) - [Example: a weekly research note](#a-concrete-task) - [Keep the outer process predictable](#a-bounded-design) - [Three questions before you add autonomy](#choose-the-smallest-design) - [Test the result and the path](#evaluate-the-whole-run) - [A useful first version](#a-first-implementation)[Read as Markdown ](/text/blog/ai-agents-vs-workflows/index.md) ## Key takeaways - An agent chooses actions during a run; a workflow follows a sequence or routing rules you designed. - A useful starting point is a fixed outer workflow with a small, bounded agent task inside it. - Evaluate the delivered result and the permitted actions separately. A polished answer can still conceal an incomplete run. ## Who decides what happens next? An AI workflow is a sequence you arrange in advance: receive an input, classify it, retrieve relevant records, draft a response, and pass the draft to a reviewer. The model may perform one or several steps, but your program defines the permitted route. An AI agent has more discretion over the route. It can inspect a result, decide which tool would help next, and continue until a stopping condition is reached. Anthropic’s engineering guide uses this distinction between predefined paths and model-directed processes. The practical question is how much choice a particular task needs. The term agent alone tells you little about the quality, reliability or permissions of a system. A workflow can involve sophisticated reasoning, and an agent can spend its time making quite ordinary decisions. Sources: [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents). ## Example: a weekly research note Consider an illustrative task for a small product team: prepare a Friday note about changes to three services the team depends on. The finished note should contain the change, its source, its relevance to the team, and a proposed next action. It should remain a draft until someone checks the recommendation. This example is a design exercise, not a claim about a deployed BotBento routine. A workflow could fetch three known release-note pages, compare each page with the previous snapshot, extract new entries, ask a model to summarize them, and save the result. That is a reasonable design when the sources are stable and the question is narrow. An empty change set is also a valid result. The workflow should explicitly say that it found no changes rather than asking the model to invent something interesting. An agent becomes more useful when a release note introduces an unfamiliar dependency or links to a migration guide. The next useful source may depend on the first finding. The agent could follow the migration link, look up the affected setting, and collect evidence for a proposed action. Give it a specific research question and allowed sources. Do not turn a request for a note into unrestricted permission to change production settings. ## Keep the outer process predictable For this research note, our suggested starting design is a fixed outer process with a bounded investigation inside it. The outer workflow owns the schedule, the list of approved services, the destination draft and the review requirement. The agent owns one question: does this particular change require attention, based on the available evidence? Before each investigation, record the source URL and what has already been learned. Set a limit on the number of pages the agent can inspect and the time the investigation can take. Those limits are design choices for your budget, not universal values. If the limit is reached, preserve the useful evidence and label the question unresolved. A partial, accurately described finding is easier to review than a confident conclusion assembled after the budget has run out. Keep collection and publication separate. The process can save a draft without having permission to send it to customers or alter a live account. If the agent recommends changing a setting, it should point to the relevant documentation and explain the expected effect. A different, explicitly authorized step can carry out the change after the recommendation has been checked. ## Three questions before you add autonomy The following questions are a practical decision aid we use for this example. They are not a benchmark or a maturity score. Answer them for the task you actually want to complete, rather than for the most impressive demonstration you can imagine. - Is the next step known before the run starts? If it is, write that step into the workflow. For example, converting an approved document into two predetermined formats does not need a bot to invent a plan. - Does evidence change what should be investigated? If it does, consider an agent for that investigation. For example, a confusing migration note may need a follow-up search that cannot be selected until the note has been read. - Can you recognize a good stopping point? If you cannot, write the acceptance criteria first. For the Friday note, every included change needs a source and relevance explanation; an unresolved change needs an explicit unknown, not a guessed recommendation. ## Test the result and the path Anthropic’s evaluation guide distinguishes the record of a run from the outcome it produces and discusses combining different kinds of grading. That distinction is useful here: reading the final note is necessary, but it cannot by itself show whether the underlying sources were checked or whether the process performed an unauthorized action. Build a small example set before you automate the weekly job. Include a week with no changes, a broken source page, two announcements that describe the same change, an announcement with an ambiguous date, and a migration guide that disagrees with an older overview. Keep the expected behavior beside each example. The expected answer can be “needs a person to resolve this contradiction.” Inspect the saved note and the run record together. Did each source open successfully? Does a claim point to the exact page supporting it? Were failed sources reported? Was the output saved in the intended place? Were tool calls limited to the allowed actions? Compare the same cases using the simpler workflow before adopting a more autonomous approach. More steps are only helpful if they improve a result that matters to your team. Sources: [Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents). ## A useful first version Start with one service and one draft destination. Write down the input, the permitted sources, the shape of the output and the conditions that require review. Run it manually on several contrasting examples. Keep the draft visible alongside its evidence so that a reviewer can check a claim without reconstructing the whole investigation. Once the basic process is dependable, decide which repeated decisions are wasting time. A conditional branch may solve the problem. A narrowly scoped agent may solve it better. Add the smallest amount of discretion that helps, then rerun your examples. The goal is a research note someone can use every Friday, including the uneventful and awkward Fridays. ## Primary sources Sources checked 2026-09-07. Standards and product documentation can change; follow the linked version when implementing. - [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) — Anthropic - [Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) — Anthropic BotBento is in development. [Suggest a correction](/contact/). ## Keep reading - [MCP tool permissions: what to check before connecting a bot](/blog/mcp-tool-permissions/) --- Canonical: https://botbento.com/blog/ai-bot-rate-limit-error-handling/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 4 MIN READ # How Should an AI Bot Handle a Tool's Rate Limit Error? A rate limit is a temporary capacity signal, but it may appear as HTTP 429 or inside a tool result. A bot should identify the layer, honor provider guidance, retry safe operations with bounded backoff, and surface the outcome. By BotBento Editorial · Published 2026-09-25 · Updated 2026-09-25 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [First identify where the limit appeared](#identify-the-layer) - [Use the provider’s wait signal when it exists](#honor-the-hint) - [Illustrative example: a calendar search pauses mid-task](#worked-example) - [A write needs a different retry decision](#writes-need-care) - [Stop at a budget and report what happened](#stop-and-report)[Read as Markdown ](/text/blog/ai-bot-rate-limit-error-handling/index.md) ## Key takeaways - Classify the failure at its actual layer: HTTP 429, a tool result with isError, and a JSON-RPC protocol error are not interchangeable. - Honor Retry-After when present; otherwise use provider guidance and bounded exponential backoff with jitter. - Before retrying a write with an uncertain outcome, check idempotency or read back state. Stop at a budget and tell the user what remains unknown. ## First identify where the limit appeared A bot can encounter a rate limit in more than one place. An HTTP API may return status 429, which means the client sent too many requests in a period. An MCP tool that calls that API may instead return a normal tools/call response whose result has isError set to true and explains that the upstream API refused work. A transport failure, a JSON-RPC error, and a tool execution error have different shapes; treating them as one generic broken connection discards useful recovery information. The MCP Tools specification distinguishes protocol errors such as unknown tools or invalid arguments from tool execution errors such as upstream API failures. That distinction does not prescribe one universal rate-limit code inside every tool. The bot should inspect the actual response and the connected service's documentation before deciding whether a retry is appropriate. A malformed argument should be fixed; waiting and resending the same invalid request will not help. Sources: [Tools - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-06-18/server/tools), [429 Too Many Requests - HTTP | MDN](https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Status/429). ## Use the provider’s wait signal when it exists HTTP 429 may include Retry-After, which tells a client how long to wait before another request. The header is optional. Some providers document a rate limit without providing an exact reset time, so a bot cannot assume that every 429 means the same delay. The first step is to preserve the status, header, service name, and original operation in the run record. Docebo's API guidance is a concrete example of a provider without a Retry-After header. It asks clients to use exponential backoff with jitter and a strict retry cap. That policy is specific to Docebo; another service may set a longer wait or prohibit retries in some circumstances. Provider guidance should override a bot's generic schedule when available. Sources: [429 Too Many Requests - HTTP | MDN](https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Status/429), [Best practices for handling API rate limits and 429 errors](https://developer.docebo.com/docs/best-practices-for-handling-api-rate-limits-and-429-errors). ## Illustrative example: a calendar search pauses mid-task Imagine a bot fetching available meeting slots through a calendar tool. On the fourth read request, the tool reports that its upstream API returned 429 and provides no wait header. The bot marks the calendar search incomplete. It waits for a bounded interval, retries the same read, and increases the delay if the limit persists. The example is a design exercise, not a report of a real calendar integration or a measured recovery time. Exponential backoff increases the maximum wait after repeated failures; jitter varies individual waits so many clients do not all retry at once. AWS's engineering explanation uses simulations to show why synchronized backoff can still produce bursts. A practical bot also sets an overall task deadline and an attempt cap, because an unbounded loop can consume the entire time budget without producing a useful answer. If the provider supplies a Retry-After value, the bot should respect it rather than schedule a retry sooner. If a shared account is rate-limited, parallel bot workers should coordinate: ten independently reasonable retries can still become an unreasonable burst. The bot can tell the user that the calendar result is delayed and either resume later within an authorized task window or ask whether the user wants to stop. Sources: [Exponential Backoff And Jitter](https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/), [Best practices for handling API rate limits and 429 errors](https://developer.docebo.com/docs/best-practices-for-handling-api-rate-limits-and-429-errors), [429 Too Many Requests - HTTP | MDN](https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Status/429). ## A write needs a different retry decision The calendar example is a read. Now imagine the same bot had tried to create a meeting and lost the response. It would be unsafe to assume that the meeting was not created. Blindly retrying could create a duplicate even if the earlier call succeeded. Before another write, the bot should use an idempotency key when the provider supports one, or read the calendar to establish whether the first operation took effect. A 429 usually signals that the request was rejected, but the bot may receive a transformed error from an intermediary or lose the response after the provider acts. The retry decision therefore depends on the exact operation and the evidence the bot retained. Where the effect cannot be checked, the honest result is an uncertain state and a request for human review, not a silent second attempt. Sources: [429 Too Many Requests - HTTP | MDN](https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Status/429). ## Stop at a budget and report what happened Docebo's published guidance caps attempts to prevent infinite retry loops and returns an error upstream when the limit remains. The exact count in that guidance is a service-specific recommendation, not a universal number for every bot. Set a cap based on the provider's policy, the task deadline, and the user's tolerance for waiting. When the budget is exhausted, record the tool, the kind of rate-limit signal, the wait and retry decisions, and whether the original action was a read or a write. Tell the user which result is missing and whether any side effect is uncertain. BotBento is still in development; this is a reliability pattern to build and verify, not a claim that BotBento already provides a production run-record feature. Sources: [Best practices for handling API rate limits and 429 errors](https://developer.docebo.com/docs/best-practices-for-handling-api-rate-limits-and-429-errors). ## Primary sources Sources checked 2026-09-25. Standards and product documentation can change; follow the linked version when implementing. - [429 Too Many Requests - HTTP | MDN](https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Status/429) — MDN Web Docs - [Tools - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-06-18/server/tools) — Model Context Protocol - [Best practices for handling API rate limits and 429 errors](https://developer.docebo.com/docs/best-practices-for-handling-api-rate-limits-and-429-errors) — Docebo - [Exponential Backoff And Jitter](https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/) — Amazon Web Services BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/bot-api-key-scoping-least-privilege/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 4 MIN READ # How Do You Scope an API Key So a Bot Can't Exceed Its Job? An unrestricted Stripe secret key can call every Stripe API resource in the account, while a restricted key can be limited to specific actions. GitHub's fine-grained personal access tokens can narrow permissions and repository access compared with a classic token, though a GitHub App is usually the better fit for long-lived organization automation. This guide walks through a dispute-summary bot and a repository PR bot, then shows how to test narrow permissions before production. By BotBento Editorial · Published 2026-10-01 · Updated 2026-10-01 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [The default key gives your bot everything](#the-default-is-too-wide) - [Worked example: scoping a billing bot's Stripe key](#worked-example-stripe) - [Worked example: scoping a repo bot's GitHub token](#worked-example-github) - [What goes wrong when the scope is wrong](#failure-paths) - [The decision to make before you connect the tool](#decision)[Read as Markdown ](/text/blog/bot-api-key-scoping-least-privilege/index.md) ## Key takeaways - Use a Stripe restricted key for only the required resources and actions; a classic GitHub token's repo scope is broader than a fine-grained token limited to one repository. - Map each actual API call to its documented permissions, test the narrow credential in a sandbox, and review request logs or endpoint failures. - Fine-grained GitHub tokens can be non-expiring unless policy forbids it; choose a deliberate expiry, and consider a GitHub App for long-lived organization bots. ## The default key gives your bot everything When you connect a tool to a bot, the fastest key to grab is usually the one with no restrictions. Stripe calls this a secret key, and any person, agent, or system holding it can do anything in the account: create charges, issue refunds, read customer data, trigger payouts, and more. A bot that only needs to check whether a dispute exists still ends up holding a key that can also move money. A bot can execute tool calls faster and more repeatedly than a person, so a broad credential raises the cost of a bug or malicious instruction. Application-level approvals can add a pause, but the provider's credential permissions remain the hard ceiling on what the API will accept. Give the bot a key that can perform its intended task and no unrelated action. Sources: [Restricted API keys](https://docs.stripe.com/keys/restricted-api-keys.md). ## Worked example: scoping a billing bot's Stripe key Say you're building a bot that checks open disputes each morning and drafts a summary. It never needs to create a charge, issue a refund, or touch a customer record. Stripe's restricted keys let you assign Read, Write, or None to each resource separately, so you create a key with Disputes set to Read and everything else set to None. If that key leaks, an attacker holding it can only read dispute data; they can't create charges, access payment methods, or trigger payouts. Don't stop at the first guess. After the bot runs for a few days, Stripe's dashboard lets you open the request logs for that specific key and see every call it actually made. Compare that list to the permissions you granted, and remove anything the bot never used. If you're not sure what a key needs up front, you can also work backward from code: search your codebase for Stripe SDK calls and map each one to its permission, for example \`Dispute.list(...)\` to Disputes: Read. Sources: [Restricted API keys](https://docs.stripe.com/keys/restricted-api-keys.md). ## Worked example: scoping a repo bot's GitHub token Now imagine a bot that creates a branch and opens pull requests in one repository. A classic personal access token with the repo scope may work, but GitHub says classic tokens can reach all repositories the owner can access, subject to the owner's own rights and organization policy. A fine-grained token can instead target one resource owner, selected repositories, and named permissions. This is a narrower fit for a bounded personal automation. For this example's two actions, creating a Git reference requires Contents: Write, and creating a pull request requires Pull requests: Write; Metadata: Read is included with repository tokens. Grant those permissions to the selected repository, then check the exact endpoints the code uses for any additional requirements. Expiration is a choice, not a fixed 366-day requirement: GitHub currently allows a fine-grained token with no expiration unless an organization or enterprise policy limits its lifetime. Set a deliberate short expiry anyway. For a long-lived bot acting for an organization, GitHub recommends a GitHub App rather than a personal token tied to one user's account. Sources: [Managing your personal access tokens - GitHub Docs](https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/managing-your-personal-access-tokens), [REST API endpoints for Git references](https://docs.github.com/en/rest/git/refs), [REST API endpoints for pull requests](https://docs.github.com/en/rest/pulls/pulls). ## What goes wrong when the scope is wrong A scope that is too narrow fails when the bot reaches an endpoint it cannot use. Stripe documents an invalid-request response for missing restricted-key permissions and shows how its per-key request logs identify failures. GitHub's REST endpoint documentation lists the fine-grained permissions required for each call. Test the full workflow, including less common branches, before relying on the key in production; an untested path can still fail later. Scoping too broad fails quietly. The bot works fine in every normal run, because the extra permissions never get exercised on purpose. The risk only shows up if the key leaks, the bot is prompted into calling something it wasn't meant to call, or a future version of the bot reuses the same credential for a different task than the one it was scoped for. None of that produces an error message before it happens. - Too narrow: an exercised call is rejected; inspect its response and endpoint permissions during tests. - Too broad: nothing fails until the key is misused, and by then the damage is whatever the full permission set allows. Sources: [Restricted API keys](https://docs.stripe.com/keys/restricted-api-keys.md), [Managing your personal access tokens - GitHub Docs](https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/managing-your-personal-access-tokens). ## The decision to make before you connect the tool Before wiring a credential into a bot, list every provider call the bot's code actually makes, then map each call to the permissions it requires. Some endpoints need more than one permission, so use the provider's current endpoint documentation rather than guessing. Grant only those permissions, restrict repository or resource ownership where supported, and choose an expiry appropriate to the task. Test in a sandbox or limited repository, then review request logs or equivalent evidence and remove unused access. If you can't yet tell which permissions a bot will need because the task is still being defined, that's a sign to delay connecting a broad key at all. Start with the bot in a read-only or sandbox mode, observe what it tries to do, and scope the real key from that observation instead of from a guess made before the bot ran. Sources: [Restricted API keys](https://docs.stripe.com/keys/restricted-api-keys.md), [Managing your personal access tokens - GitHub Docs](https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/managing-your-personal-access-tokens). ## Primary sources Sources checked 2026-10-01. Standards and product documentation can change; follow the linked version when implementing. - [Restricted API keys](https://docs.stripe.com/keys/restricted-api-keys.md) — Stripe - [Managing your personal access tokens - GitHub Docs](https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/managing-your-personal-access-tokens) — GitHub - [REST API endpoints for Git references](https://docs.github.com/en/rest/git/refs) — GitHub - [REST API endpoints for pull requests](https://docs.github.com/en/rest/pulls/pulls) — GitHub BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/bot-multi-step-task-failure-rollback/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 5 MIN READ # What Should a Bot Do When a Multi-Step Task Fails Halfway? When a bot's task has several steps that each change something outside the bot — a booking, a payment, a calendar invite — a failure partway through leaves real-world state stuck in between. Borrowing the compensating-transaction idea from distributed systems gives bot builders a concrete way to design what 'undo' means for each step, decide which steps can't be undone, and record enough to recover manually when they can't. By BotBento Editorial · Published 2026-09-28 · Updated 2026-09-28 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [The problem: your bot already changed something](#the-half-finished-problem) - [Design the undo before you design the action](#compensating-actions) - [Find the point of no return before you build](#the-pivot-point) - [A cancelled tool call is not a confirmed undo](#cancellation-is-not-rollback) - [Illustrative example: a trip-booking bot](#worked-example) - [What the run record needs](#what-to-record)[Read as Markdown ](/text/blog/bot-multi-step-task-failure-rollback/index.md) ## Key takeaways - Before you build the forward action for a step, write down what the compensating action is — if you can't describe one, that step needs a different design or a human checkpoint. - Mark one step in your sequence as the pivot: the point after which the task must go forward, not back. Steps before it are compensable; steps after it must be retried until they succeed. - Cancellation notifications (MCP or otherwise) tell a tool to stop working — they don't prove any external effect was undone. Record cancellation-requested and effect-confirmed as separate facts. ## The problem: your bot already changed something Even one tool call can have an uncertain outcome: the request may reach a provider while its response is lost. A multi-step task adds another problem. Say a bot books a flight, then a hotel, then adds both to a calendar. If the hotel booking fails after the flight is confirmed, the user may have a paid flight with no hotel. Reporting 'step 2 failed' and stopping leaves the user to discover and resolve the remaining booking. This is the same problem distributed systems hit when a transaction spans multiple services with no shared database to roll back. There is no single 'undo everything' button, because each step already committed somewhere else. The fix used there — the saga pattern — is worth borrowing directly for bot task design, because a bot's tool calls are exactly this: independent, already-committed actions chained together. ## Design the undo before you design the action The saga pattern asks what action can compensate for each completed step when a later step fails. That action is a new operation with an offsetting business effect, not a database rollback. Microsoft's Azure Architecture Center describes a saga as a sequence of local transactions and says failure may require compensating earlier changes. For a bot, this means recording both the forward action's external identifier and the operation that could reverse or mitigate it. Some steps have no clean opposite. Microsoft's compensating-transaction guidance explains that compensation need not run in strict reverse order and that business rules determine what a reversal can achieve. Canceling a flight, for example, does not necessarily restore the full payment. For a bot, the response to a sent email may be a correction or human follow-up; the original message cannot be unsent. An automatic compensation is not always possible. Compensations can fail or leave a result temporarily uncertain. Record the progress and external reference for each step so recovery can resume without blindly repeating an action. Microsoft recommends idempotent commands for this reason. A bot asked to cancel a hotel should check whether the reservation was already cancelled and confirm the provider's current state before reporting a refund or release. Sources: [Saga Design Pattern - Azure Architecture Center](https://learn.microsoft.com/en-us/azure/architecture/patterns/saga), [Compensating Transaction Pattern - Azure Architecture Center](https://learn.microsoft.com/en-us/azure/architecture/patterns/compensating-transaction). ## Find the point of no return before you build The saga pattern separates compensable steps, a pivot that commits the workflow to moving forward, and later retryable steps. That classification is useful only if the later steps can actually be completed through safe retries. Microsoft's guidance treats post-pivot actions as idempotent operations that help the workflow reach a consistent state. A bot should not label a step retryable merely because it can call the same API again; permanent unavailability, changed prices and human decisions still need an escalation path. For a trip, payment capture might seem like the pivot, but only after the team checks the flight's cancellation terms and the hotel's ability to complete a booking. If no acceptable hotel remains, repeatedly trying the same booking does not solve the task. The bot should stop, show the confirmed flight and payment state, and seek the user's decision about alternatives or a refund request. Choose the pivot based on the actual provider contracts, not on an appealing diagram. Sources: [Saga Design Pattern - Azure Architecture Center](https://learn.microsoft.com/en-us/azure/architecture/patterns/saga). ## A cancelled tool call is not a confirmed undo If your bot calls tools over MCP, there's a related trap: a cancellation notification tells a server to stop, but it doesn't tell you the external effect was reversed. The MCP specification defines cancellation as a one-way, best-effort signal: when a party wants to cancel an in-progress request, it sends a notifications/cancelled notification containing the ID of the request to cancel and an optional reason string that can be logged or displayed. The spec also warns that timing is not guaranteed: due to network latency, cancellation notifications may arrive after processing has completed, and the sender of the cancellation notification SHOULD ignore any response to the request that arrives afterward. In practice this means 'I sent a cancel' and 'the booking was actually released' are two different facts, and a bot's run record should keep them separate. If your bot cancels a hotel-booking tool call mid-flight, log that the cancellation was requested, then separately confirm — by calling a status or list endpoint — that the booking no longer exists before you tell the user it's undone. Sources: [Cancellation - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-03-26/basic/utilities/cancellation). ## Illustrative example: a trip-booking bot This is a fictional walkthrough to make the design concrete, not a description of a real product. Suppose a personal bot handles 'book my trip to Denver.' It plans three steps: reserve a flight, reserve a hotel, add both to the calendar. Step 1 is a temporary flight hold that the provider explicitly allows the bot to release. Step 2 is a prepaid hotel booking whose terms define the pivot only if the remaining work can safely continue. Step 3 creates calendar entries using a stable trip ID and a provider-supported idempotency key. If the calendar call times out, the bot first reads back entries for that trip, then retries only if no matching entry exists. The example's safety depends on those stated provider capabilities; without them, it must pause for review. If step 2 fails before payment clears, the bot requests release of the flight hold and verifies that release before saying nothing remains booked. If payment appears captured but the hotel confirmation is missing, the bot queries the booking by request key, preserves the payment reference and reports the uncertainty. If that lookup cannot resolve it, a human should choose the next action. The bot should not promise a refund or issue another charge just to make the workflow appear complete. ## What the run record needs For each step in a multi-step task, the run record should capture: which category the step falls into (compensable, pivot, or retryable), whether the forward action succeeded, whether a compensation was attempted, and whether that compensation was confirmed — not just requested. This is a small extension of a normal run log, but it's the piece that turns 'the task failed' into 'the task failed and here's exactly what state the world is in now.' A useful run record also distinguishes a requested cancellation from a confirmed reversal and keeps the external identifiers needed for a later check. This is design guidance for agent builders, not a claim that BotBento currently provides compensation tracking. BotBento's editorial team used AI assistance to draft this article and checked its claims against the linked primary sources. ## Primary sources Sources checked 2026-09-28. Standards and product documentation can change; follow the linked version when implementing. - [Saga Design Pattern - Azure Architecture Center](https://learn.microsoft.com/en-us/azure/architecture/patterns/saga) — Microsoft Learn - [Compensating Transaction Pattern - Azure Architecture Center](https://learn.microsoft.com/en-us/azure/architecture/patterns/compensating-transaction) — Microsoft Learn - [Cancellation - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-03-26/basic/utilities/cancellation) — Model Context Protocol BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/bot-routine-silent-failure-detection/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 4 MIN READ # How Do You Know Your Bot's Recurring Routine Has Gone Silent? A quiet recurring routine has three possible failure points: no invocation, failed delivery to the target, or work that starts but stalls. An independent deadline check detects a missed slot; scheduler retries and a dead-letter queue record failed delivery; a configured heartbeat timeout detects stalled long-running work. A hypothetical nightly digest shows how to combine the signals without mistaking one for another. By BotBento Editorial · Published 2026-09-30 · Updated 2026-09-30 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [Three places a recurring routine can go quiet](#three-kinds-of-silence) - [When the schedule fires but target delivery fails](#failed-target-delivery) - [When work begins and then stalls](#stalled-work) - [A nightly digest, traced from schedule to result](#worked-example) - [The smallest useful reliability loop](#practical-checklist)[Read as Markdown ](/text/blog/bot-routine-silent-failure-detection/index.md) ## Key takeaways - Check each expected schedule slot against a durable run or completion record after a defined deadline; a dead-letter queue cannot report an invocation that never occurred. - EventBridge Scheduler retries failed target delivery and can send exhausted failures to a dead-letter queue; its CloudWatch metrics help diagnose attempts but are best effort. - For long-running work, a configured heartbeat timeout can expose lost progress; safe retries still need idempotent external actions or checkpoints. ## Three places a recurring routine can go quiet Imagine a bot expected to produce a support digest every night. If there is no digest in the morning, that absence alone does not identify the fault. The schedule may never have invoked the target. It may have attempted delivery and failed. Or the target may have accepted the request while the digest work later stalled. Those cases have different owners, evidence, and recovery steps. Start with a durable record keyed by the expected slot, such as the routine name and scheduled date. Record invocation, running, completed, and failed states where possible. A separate watchdog can compare the expected slot with that record after a deadline and alert a person when the slot is absent or incomplete. Run this check outside the routine being watched: a bot cannot reliably report its own failure to start. Amazon EventBridge Scheduler's invocation-attempt metrics can help investigate whether it tried to invoke a target. AWS cautions that these CloudWatch metrics are delivered on a best-effort basis and are not a complete accounting. A missing metric is therefore a clue, not proof that no invocation happened. The independent expected-slot check remains the backstop. Sources: [Monitoring Amazon EventBridge Scheduler with Amazon CloudWatch](https://docs.aws.amazon.com/scheduler/latest/UserGuide/monitoring-cloudwatch.html). ## When the schedule fires but target delivery fails EventBridge Scheduler can retry a failed target invocation according to the schedule's configured retry policy. If delivery still fails, a configured dead-letter queue (DLQ) can receive a record with the target, error code, error message, retry attempts, and reason retries stopped. The retry limits are configuration choices; do not assume a particular number of attempts or duration for your schedule. That DLQ answers a narrow question: what happened when Scheduler tried to reach the target? It does not mean the downstream job completed successfully, and it cannot report a schedule that never attempted an invocation. If the target accepted the request and later hung, delivery may look successful while the business result is missing. For a small system, the same idea can be a durable failed-delivery row with slot ID, target, error, and resolution state. Alert on new unresolved failures, inspect the cause, and replay only after the target action can tolerate duplicate attempts. If the row is empty but the deadline watchdog sees no completed slot, investigate the scheduler and worker rather than declaring the routine healthy. Sources: [Configuring a schedule's dead-letter queue in EventBridge Scheduler](https://docs.aws.amazon.com/scheduler/latest/UserGuide/configuring-schedule-dlq.html), [Monitoring Amazon EventBridge Scheduler with Amazon CloudWatch](https://docs.aws.amazon.com/scheduler/latest/UserGuide/monitoring-cloudwatch.html). ## When work begins and then stalls A successful handoff says little about work that lasts minutes or hours. A worker might freeze during a tool call, lose its process, or stop moving through a batch. Temporal documents Activity Heartbeats for long-running work: the Activity reports progress, and a configured Heartbeat Timeout can fail an attempt when heartbeats stop. A retry follows only when the retry policy allows it. Heartbeat timeouts must actually be configured, and the Activity must send heartbeats. A zero or unset Heartbeat Timeout disables that missing-heartbeat check. A separate overall deadline is also useful when a worker keeps sending heartbeats but never completes meaningful work. A heartbeat is evidence of a live worker, not proof of a correct result. Temporal lets an Activity include progress details in a heartbeat that a later attempt can read. That can help resume a batch, but the detail alone does not make side effects safe. Store checkpoints durably and make external writes idempotent, for example by giving each digest post a stable slot ID, before allowing retries to repeat them. Sources: [Detecting Activity failures](https://docs.temporal.io/encyclopedia/detecting-activity-failures), [Timeouts and Retry Policies](https://docs.temporal.io/evaluate/features/timeouts-and-retries). ## A nightly digest, traced from schedule to result Consider a hypothetical team bot that reads support tickets, drafts a digest, and posts it to a channel. This is an illustrative design, not a report of a live deployment. The team expects one completed record for each night's slot. Its watchdog checks the previous slot after the agreed deadline and sends an actionable alert if the record is missing or unfinished. If the Scheduler attempted to call the digest target but that delivery failed, retry and DLQ evidence point to an address, permission, throttling, or target-service problem. The operator checks the DLQ and target logs, repairs the cause, and replays the same slot safely. If there is no invocation record or delivery failure, the operator checks the schedule state and independent logs; an empty DLQ alone proves nothing. If the digest worker started and then stopped making progress, a heartbeat timeout or run deadline marks the attempt as stalled. A retry can read the last durable checkpoint, but posting the digest must still use the slot ID to avoid a duplicate message. The final completed record should identify the actual posted result, so the watchdog can tell completion from mere launch. Sources: [Configuring a schedule's dead-letter queue in EventBridge Scheduler](https://docs.aws.amazon.com/scheduler/latest/UserGuide/configuring-schedule-dlq.html), [Monitoring Amazon EventBridge Scheduler with Amazon CloudWatch](https://docs.aws.amazon.com/scheduler/latest/UserGuide/monitoring-cloudwatch.html), [Detecting Activity failures](https://docs.temporal.io/encyclopedia/detecting-activity-failures). ## The smallest useful reliability loop Give every scheduled run a stable slot ID and a completion deadline. Persist a state transition when the target starts and when the intended result is actually delivered. Check expected slots from a separate watchdog and route misses to someone who can act; otherwise the failure only moves from one silent database row to another. Configure retry and failed-delivery capture for the scheduler-to-target boundary. For long-running tasks, add an appropriate progress heartbeat and timeout, plus an overall deadline. Make retries safe through idempotent writes and durable checkpoints. The right signal depends on where the work stopped, so inspect the slot record, delivery evidence, and worker progress together. Finally, test the three cases deliberately: disable a future schedule in a safe test environment, make a target reject a test delivery, and pause a long-running test worker. Verify that each produces a distinct alert and a recoverable record. A green scheduler dashboard by itself does not establish that the digest arrived. Sources: [Configuring a schedule's dead-letter queue in EventBridge Scheduler](https://docs.aws.amazon.com/scheduler/latest/UserGuide/configuring-schedule-dlq.html), [Monitoring Amazon EventBridge Scheduler with Amazon CloudWatch](https://docs.aws.amazon.com/scheduler/latest/UserGuide/monitoring-cloudwatch.html), [Detecting Activity failures](https://docs.temporal.io/encyclopedia/detecting-activity-failures). ## Primary sources Sources checked 2026-09-30. Standards and product documentation can change; follow the linked version when implementing. - [Configuring a schedule's dead-letter queue in EventBridge Scheduler](https://docs.aws.amazon.com/scheduler/latest/UserGuide/configuring-schedule-dlq.html) — Amazon Web Services - [Monitoring Amazon EventBridge Scheduler with Amazon CloudWatch](https://docs.aws.amazon.com/scheduler/latest/UserGuide/monitoring-cloudwatch.html) — Amazon Web Services - [Detecting Activity failures](https://docs.temporal.io/encyclopedia/detecting-activity-failures) — Temporal Technologies - [Timeouts and Retry Policies](https://docs.temporal.io/evaluate/features/timeouts-and-retries) — Temporal Technologies BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/ Format: Markdown representation of the public HTML page. BOTBENTO / BLOG # Field notes for the agent world Practical guides to AI agents, workflows, tools and permissions. Original, sourced explanations from BotBento. 2026-10-014 min read ## [What Actually Happens When Your Bot Cancels an MCP Tool Call? ](/blog/mcp-tool-call-cancellation-what-happens/) MCP cancellation asks a server to stop work, but it may still finish. Learn the current stdio and HTTP rules and how to check side effects. 2026-10-014 min read ## [How Do You Scope an API Key So a Bot Can't Exceed Its Job? ](/blog/bot-api-key-scoping-least-privilege/) A practical Stripe and GitHub example for limiting a bot's credential to its job, with the permissions to verify and the failure paths to test. 2026-09-304 min read ## [MCP Roots Is Deprecated. Where Should Your Bot's File Boundaries Live Now? ](/blog/mcp-roots-deprecated-migration/) MCP Roots is deprecated in the 2026-07-28 spec. Learn what remains supported, why roots are advisory, and how to migrate file scoping safely. 2026-09-304 min read ## [How Do You Know Your Bot's Recurring Routine Has Gone Silent? ](/blog/bot-routine-silent-failure-detection/) A scheduled bot can miss its start, fail to reach its target, or stall after starting. Each silence needs a different signal and a clear recovery path. 2026-09-294 min read ## [How Long Should You Let a Bot's Tool Call Run Before Timing Out? ](/blog/tool-call-timeout-duration/) A practical guide to MCP tool-call timeouts: distinguish protocol guidance from client defaults, measure real durations, and use asynchronous jobs for slow work. 2026-09-294 min read ## [MCP Tools Can Declare an Output Schema. Should Your Bot Validate It? ](/blog/mcp-tool-output-schema-validation/) MCP tools can declare an outputSchema and return structuredContent. Here is how to validate successful results without confusing tool errors with transport failures. 2026-09-284 min read ## [MCP Dropped Sessions. How Do You Keep State Across Tool Calls Now? ](/blog/mcp-stateless-cross-call-state/) The MCP 2026-07-28 spec removed session IDs and the initialize handshake. Here's the explicit-handle pattern that replaces them, and what breaks if you don't update. 2026-09-285 min read ## [What Should a Bot Do When a Multi-Step Task Fails Halfway? ](/blog/bot-multi-step-task-failure-rollback/) A bot that books a flight then fails to book a hotel has left the world in a half-finished state. Here is how to design the undo step. 2026-09-274 min read ## [How Do You Actually Revoke a Bot's Access to a Tool It No Longer Uses? ](/blog/revoke-bot-tool-access/) OAuth token revocation can have a propagation delay. Learn what the standard guarantees, how validation affects the result, and how to verify a bot has stopped using a tool. 2026-09-274 min read ## [MCP Tool Results: How Do You Know a Call Actually Succeeded? ](/blog/mcp-tool-result-success-signal/) Read MCP protocol errors, isError, structuredContent and real-world readback separately before telling a user a tool action succeeded. 2026-09-263 min read ## [How Do You Verify an AI Agent Actually Finished the Task? ](/blog/verify-ai-agent-task-completion/) A bot saying 'done' is not proof. Here is how to check completion independently, with a worked example and what evidence to keep. 2026-09-263 min read ## [When Should a Bot Use a Supervisor Agent Instead of One Agent With Tools? ](/blog/supervisor-agent-vs-single-agent-tools/) A concrete rule for choosing between one agent with many tools and a supervisor that delegates to specialist agents, with a worked example. 2026-09-254 min read ## [MCP Sampling Is Deprecated. What Should Bot Builders Do Now? ](/blog/mcp-sampling-deprecated/) MCP's sampling feature, which let servers ask your client to run an LLM call, is now deprecated. Here is what that means for bots built on it. 2026-09-254 min read ## [How Should an AI Bot Handle a Tool's Rate Limit Error? ](/blog/ai-bot-rate-limit-error-handling/) Learn how an AI bot should classify rate limits, honor Retry-After, retry safely with jitter and a cap, and report an unfinished tool action. 2026-09-244 min read ## [What Can an MCP Tool Ask You For Mid-Task, and What Can't It? ](/blog/mcp-elicitation-mid-task-permissions/) MCP elicitation lets a connected tool pause and ask a question mid-task. Here's what the spec allows it to request, and what it must route elsewhere. 2026-09-244 min read ## [How Do You Stop a Retried Tool Call From Running Twice? ](/blog/agent-tool-retry-idempotency-keys/) A timed-out tool call can secretly succeed. See why MCP's idempotentHint doesn't protect you, and what an idempotency key actually does. 2026-09-233 min read ## [How Do You Review an AI-Generated Research Note? ](/blog/reviewing-ai-generated-research-notes/) A practical checklist for checking source ownership, dates, contradictions and unresolved questions in AI-drafted research notes before you rely on them. 2026-09-234 min read ## [MCP Resource Indicators: What Your Bot's OAuth Tokens Must Specify ](/blog/mcp-resource-indicator-requirement/) MCP's HTTP authorization spec requires OAuth resource indicators. Learn what they protect, where implementation can fail, and what to check before connecting a bot. 2026-09-205 min read ## [Plugin or authorized connection: what's actually different? ](/blog/plugin-vs-authorized-connection/) Installing a tool on a bot and letting it act on your data are two separate steps. Here is where the line sits, with a worked example. 2026-09-204 min read ## [Why we are building BotBento around people, not just bots ](/blog/people-and-bots-in-one-room/) Grok Bot's group chats hold bots only. BotBento is built so people and bots share one room under one permission model. Here is the reasoning and the honest status. 2026-09-105 min read ## [How do you test an AI agent that uses a calendar? ](/blog/test-calendar-ai-agent/) Test a calendar AI agent with a disposable event, explicit time zone, denied write and revoked connection. Check the saved result before trusting its reply. 2026-09-104 min read ## [What belongs in an AI agent stopping rule? ](/blog/ai-agent-stopping-rules/) Define when an AI agent should finish, pause or stop at a limit, with an illustrative research task and practical checks for loops and unresolved work. 2026-09-085 min read ## [Local or cloud: where should your AI agent run? ](/blog/local-vs-cloud-ai-agents/) Choose where an AI agent runs by separating its model, tools and scheduler. Compare local, cloud and hybrid setups with a practical decision checklist. 2026-09-085 min read ## [What should an AI agent record after each run? ](/blog/ai-agent-run-result-record/) Design a useful AI agent run record with evidence, partial results, retry decisions and clear next steps, using an illustrative research routine. 2026-09-075 min read ## [MCP tool permissions: what to check before connecting a bot ](/blog/mcp-tool-permissions/) Understand MCP servers, authorization and tool approval with an illustrative calendar example and a practical connection review checklist. 2026-09-075 min read ## [AI agents vs workflows: choose who decides the next step ](/blog/ai-agents-vs-workflows/) A practical way to choose between an AI agent and a fixed workflow, with a weekly research example, stopping rules and an evaluation checklist. --- Canonical: https://botbento.com/blog/local-vs-cloud-ai-agents/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 5 MIN READ # Local or cloud: where should your AI agent run? Choose an agent's location by mapping three things separately: where the model processes requests, where tools act on files and services, and where the agent process and scheduler stay running. A local interface can use a remote model, and a cloud tool environment does not move a laptop's scheduler with it. By BotBento Editorial · Published 2026-09-08 · Updated 2026-09-08 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [Map the model, tools and scheduler separately](#three-locations) - [A local interface does not establish a local data path](#data-path) - [Place tools where their required resources are available](#execution-and-access) - [Put the scheduler where it can actually run](#unattended-work) - [Choose for one task before choosing for every bot](#choose-for-one-task) - [Use a short placement acceptance checklist](#acceptance-checklist)[Read as Markdown ](/text/blog/local-vs-cloud-ai-agents/index.md) ## Key takeaways - Model hosting, tool execution and the process that starts a routine are separate placement decisions. - Choose local execution for a task's actual access needs, and remote execution for a defined operational requirement rather than a cloud label. - Test the intended setup with the client closed and a dependency unavailable before relying on it for unattended work. ## Map the model, tools and scheduler separately The question usually starts with a practical need: should a bot work on my laptop or on a server? Before choosing, draw the complete path of one task. Include the process receiving the request, the model endpoint, the tools reading or changing data, and the process responsible for starting scheduled work. These components can live in different places. Hermes documents terminal backends separately from model-provider configuration. Its local terminal backend executes commands on the host, while its SSH backend executes them on a remote server. That is a choice about command execution. It does not, by itself, tell you where the model processes a prompt. For a proposed document assistant, write down a concrete map: agent process on laptop A, model endpoint B, files in folder C, output in destination D. Use actual configured endpoints and storage locations when implementing it. A diagram containing only a laptop icon and a cloud icon hides the decisions that affect the work. Sources: [Hermes Agent Configuration](https://hermes-agent.nousresearch.com/docs/user-guide/configuration/). ## A local interface does not establish a local data path Ollama's FAQ distinguishes running a model locally from using its cloud-hosted models, where prompts and responses are processed by the service. It also documents a local-only setting that disables Ollama cloud models and web search. This is a useful example of why the application's name or the location of its window cannot settle the data question. Our recommendation is to inspect each boundary in the task. If a local agent sends a document excerpt to a hosted model, that excerpt crosses a network boundary. If a local model asks a connected search tool a question containing private project details, the tool request is a separate boundary to review. Turning off one application's cloud feature does not configure every other tool in an agent's workflow. Decide which material the task needs before choosing its placement. A bot that sorts filenames may need much less information than one that summarizes complete documents. Use representative, non-sensitive fixtures to check the actual requests and outputs. Keep the data decision tied to the task rather than assuming that either local or cloud automatically meets it. Sources: [FAQ](https://docs.ollama.com/faq). ## Place tools where their required resources are available A tool needs a deliberate route to the files or service it uses. Consider a fictional bot that checks a folder on your workstation and creates a summary for your team. Moving its command execution to a remote machine does not make that folder appear there. You need a defined copy, mount or authorized connection, and you need to know where the result will be saved. Hermes offers both container and remote terminal backends, among other options. Its configuration documentation describes the local backend as direct execution on the machine and the Docker backend as execution in a container. Treat that distinction as part of the access design, not a blanket promise about everything an agent can reach. For our folder example, make a small resource inventory: one input directory, one output directory and any external API the routine actually needs. Verify the intended resources are accessible and unrelated resources are not. If the only way the design works is to give a simple sorting task broad access to an entire workstation, reconsider the task boundary before deciding which machine should host it. Sources: [Hermes Agent Configuration](https://hermes-agent.nousresearch.com/docs/user-guide/configuration/). ## Put the scheduler where it can actually run Hermes's scheduled-task documentation says its gateway daemon checks for due jobs and starts agent sessions. This makes the gateway process an operational dependency of Hermes scheduling. Choosing a remote terminal backend is a separate setting; it does not mean the gateway itself has moved to that remote environment. In an illustrative hybrid setup, a laptop runs the gateway and sends commands to a remote machine. If the laptop stops running the gateway, the remote machine's availability alone is not sufficient to start a new scheduled agent session. To design for work while the laptop is unavailable, identify every process needed to trigger and complete that work and place those processes deliberately. Keep client closure distinct from stopping the service. A chat window may be disposable while a managed background service continues running. Test the exact installation you plan to use: close the client, observe whether the service remains active, and check whether a scheduled fixture completes. Do not infer this behavior from a successful interactive conversation. Sources: [Scheduled Tasks (Cron)](https://hermes-agent.nousresearch.com/docs/user-guide/features/cron/). ## Choose for one task before choosing for every bot The following examples are proposed designs, not measured performance comparisons or claims about available BotBento features. Each begins with a different requirement. The point is to identify the dependency driving the decision, then test a small version before expanding it. - Personal folder cleanup while you are present: start by evaluating a local tool environment with a narrow input folder and a reviewable output. Decide model hosting separately according to the content it will receive. - A public-source briefing needed while your laptop is off: evaluate a remotely hosted agent process and scheduler, with only the source and delivery access required for that briefing. - A team task that also needs a workstation-only file: evaluate a hybrid design, but make the workstation dependency visible. Define whether an unavailable file postpones the task or permits a clearly marked partial result. - An occasional large computation: evaluate a separate execution environment for that step. Specify how inputs arrive and how outputs are retained before the temporary environment is removed. ## Use a short placement acceptance checklist For your chosen design, ask someone else to identify where a prompt is processed, where a command runs and where the schedule is owned. They should also be able to locate the output after the client is closed. If those answers require guessing, the setup is not yet clear enough to depend on. Run a disposable fixture with one unavailable dependency. Inspect whether the outcome names that dependency and preserves any useful partial result. Then restore it and check the recovery behavior. Keep the evidence with the task's result record, including which parts were verified and which still need review. Finally, name who maintains each running component: updates, credentials, storage and recovery need an owner in both local and remote arrangements. Compare measured resource use and real operational effort only after trying the representative task. This guide offers no universal price or speed winner. BotBento remains in development; these placement principles can help you evaluate an agent setup independently of its future interface. ## Primary sources Sources checked 2026-09-08. Standards and product documentation can change; follow the linked version when implementing. - [Hermes Agent Configuration](https://hermes-agent.nousresearch.com/docs/user-guide/configuration/) — Nous Research - [FAQ](https://docs.ollama.com/faq) — Ollama - [Scheduled Tasks (Cron)](https://hermes-agent.nousresearch.com/docs/user-guide/features/cron/) — Nous Research BotBento is in development. [Suggest a correction](/contact/). ## Keep reading - [What should an AI agent record after each run?](/blog/ai-agent-run-result-record/) - [MCP tool permissions: what to check before connecting a bot](/blog/mcp-tool-permissions/) --- Canonical: https://botbento.com/blog/mcp-elicitation-mid-task-permissions/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 4 MIN READ # What Can an MCP Tool Ask You For Mid-Task, and What Can't It? MCP elicitation lets a connected tool pause a task and ask the user a question before continuing. In the 2025-11-25 specification, form mode handles ordinary structured questions and URL mode handles sensitive interactions outside the client. Knowing which mode a request should use helps you review what the tool is asking you to share. By BotBento Editorial · Published 2026-09-24 · Updated 2026-09-24 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [What elicitation actually is](#what-elicitation-is) - [What a tool must not ask for through a form](#what-cannot-be-asked) - [A worked example: booking a meeting room (illustrative)](#worked-example) - [What to check before you trust a tool's elicitation requests](#what-to-check) - [Why this matters for how you scope tool access](#why-it-matters-for-permissions)[Read as Markdown ](/text/blog/mcp-elicitation-mid-task-permissions/index.md) ## Key takeaways - Elicitation has two modes: form mode for ordinary structured questions, URL mode for anything involving credentials, payments, or third-party authorization. - The spec explicitly forbids form mode from being used to collect passwords, API keys, access tokens, or payment credentials — that traffic must go through URL mode instead. - In the 2025-11-25 protocol, a client declares the elicitation modes it supports during initialization and must give the user a way to decline or cancel a request. ## What elicitation actually is Elicitation is a message type inside the Model Context Protocol. It lets a server-side tool pause in the middle of a task and ask the person using it for more information, instead of guessing, failing, or silently proceeding with incomplete data. The specification supports two distinct modes. Form mode lets a server request structured data from the user through a JSON Schema the client renders as a form. URL mode sends the user to an external URL for interactions that must never pass through the MCP client at all. Every elicitation response resolves to one of three actions: the user accepted and provided data, declined, or cancelled the whole exchange. A tool that can't handle a decline or a cancel gracefully is not implementing the pattern correctly. Sources: [Elicitation - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-11-25/client/elicitation). ## What a tool must not ask for through a form The line between the two modes is not stylistic. It's a security boundary written into the spec. The official documentation is explicit that servers must not request certain categories of information through the in-band form. The specification states that "Servers MUST NOT use form mode elicitation to request sensitive information such as passwords, API keys, access tokens, or payment credentials." Those must go through URL mode instead, where the exchange happens outside the client entirely. The definition of sensitive here is narrow but firm: "secrets and credentials that grant access or authorize transactions." General contact details like a name or email aren't automatically banned from a form — the server can ask, but the user still has to be able to review and decline. If you see a connected tool's form asking for a password, an API key, or a card number, that's not a quirky UI choice. It's a spec violation, and it's a signal to stop and check what else that tool might be doing wrong. Sources: [Elicitation - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-11-25/client/elicitation), [One Year of MCP: November 2025 Spec Release](https://blog.modelcontextprotocol.io/posts/2025-11-25-first-mcp-anniversary/). ## A worked example: booking a meeting room (illustrative) Say a scheduling bot uses a meeting-room tool over MCP. The bot asks it to book a room for 2pm. The tool doesn't know your seating preference or whether you need a projector, so it sends a form-mode elicitation: 'Which room feature matters most — projector, whiteboard, or capacity?' You pick one, the client lets you review the choice before sending, and the tool books accordingly. That's a normal, bounded use of form mode. Now say the same tool needs to check room availability against your company's calendar provider, and that provider requires you to log in and grant access. Under the 2025-11-25 spec, that exchange cannot happen as a form. The server can use URL mode to offer a link that starts an external authorization flow. After you inspect the full URL and consent to open it, you complete the provider login outside the MCP client. The client does not receive your password or the resulting third-party tokens; the server manages any tokens it obtains. This second case is illustrative, not a report of a real deployment, but it maps directly onto why the spec splits the two modes: one is for filling in gaps in a task, the other is for anything where a credential or a payment changes hands. ## What to check before you trust a tool's elicitation requests Before you connect a bot to a tool that uses elicitation, there are a handful of concrete things worth checking rather than assuming compliance. First, whether your client declares support for elicitation during initialization. In the 2025-11-25 protocol, the client declares the form and URL modes it supports, and the server must not send a mode the client did not declare. A newer draft uses per-request capability metadata instead, so check the protocol version your client and server actually use. Second, whether URL-mode requests are handled the way the spec demands: the client must show you the full URL before you consent, must not fetch it automatically, and must open it in a way that keeps the page's content and your input away from the client and the model. - For 2025-11-25, client declares the form and URL modes it supports during initialization - Every request gives you a visible way to decline or cancel, not just accept - Form-mode requests never ask for passwords, API keys, tokens, or payment details - URL-mode requests show you the full target URL and wait for your consent before opening it - The URL opens in a secure context that keeps page content and input away from the client and model Sources: [Elicitation - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-11-25/client/elicitation), [Elicitation - Draft Specification](https://modelcontextprotocol.io/specification/draft/client/elicitation). ## Why this matters for how you scope tool access Elicitation changes what "connecting a tool" means. A tool that only ever answers a call and returns a result is easy to reason about: you know its inputs and outputs ahead of time. A tool that can pause and ask you something mid-task introduces a second channel you have to trust — and the mode it uses tells you how much trust that channel actually requires. If you're auditing which tools a bot can reach, review form-mode requests for the data they collect and treat URL-mode requests as sensitive interactions that deserve scrutiny before you open the link. A permission review should show which modes a tool can use and which server is asking, before a bot is allowed to run it unattended. BotBento is still in development. ## Primary sources Sources checked 2026-09-24. Standards and product documentation can change; follow the linked version when implementing. - [Elicitation - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-11-25/client/elicitation) — Model Context Protocol - [One Year of MCP: November 2025 Spec Release](https://blog.modelcontextprotocol.io/posts/2025-11-25-first-mcp-anniversary/) — Model Context Protocol - [Elicitation - Draft Specification](https://modelcontextprotocol.io/specification/draft/client/elicitation) — Model Context Protocol BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/mcp-resource-indicator-requirement/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 4 MIN READ # MCP Resource Indicators: What Your Bot's OAuth Tokens Must Specify For MCP clients using HTTP OAuth authorization, the specification requires a 'resource' parameter in authorization and token requests. It identifies the intended MCP server; effective token binding also depends on the authorization server issuing the right audience and the MCP server validating it. Here is an illustrative replay scenario and a practical connection checklist. By BotBento Editorial · Published 2026-09-23 · Updated 2026-09-23 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [The resource-indicator requirement](#what-changed) - [Why this closes a real phishing-shaped gap](#why-it-matters) - [A worked example: two MCP servers, one bot](#worked-example) - [What to check before you trust a connection](#what-to-check) - [Timeline and current status](#status-and-timeline)[Read as Markdown ](/text/blog/mcp-resource-indicator-requirement/index.md) ## Key takeaways - MCP clients using HTTP OAuth authorization must send a 'resource' parameter identifying the intended MCP server in both the authorization and token requests. - Without audience binding and validation, a token minted for one MCP server could be accepted by a different server sharing the same authorization domain. - Before connecting a bot to an MCP server, check that the server validates the token's audience claim against its own canonical URI, not just that a token is present. ## The resource-indicator requirement For MCP clients using HTTP OAuth authorization, the specification requires Resource Indicators for OAuth 2.0, as defined in RFC 8707, to name the target resource a token is being requested for. The resource parameter must appear in both the authorization request and the token request, must identify the specific MCP server the client intends to use the token with, and must use that server's canonical URI rather than a broader domain or a wildcard. MCP authorization itself remains optional for implementations. This is a hardening, not a suggestion. As the MCP lead maintainer put it when the requirement landed in the draft spec, resource indicators are no longer optional, and client developers are now required to implement RFC 8707. That phrasing matters for anyone maintaining an MCP client today: a client that skips the resource parameter is no longer just missing a nice-to-have, it is out of spec. The practical effect is that each OAuth request in this flow carries an explicit statement of intent. A client is not asking an authorization server for 'a token that can talk to MCP servers in general.' It is asking for a token scoped to one named server. Whether the resulting token is actually restricted to that server also depends on the authorization server honoring the parameter and the MCP server checking the token's audience. Sources: [Authorization - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/basic/authorization), [Update To MCP Authorization Spec - Resource Parameter (RFC 8707) | Den Delimarsky](https://den.dev/blog/mcp-authorization-resource/). ## Why this closes a real phishing-shaped gap The change traces back to a phishing-style scenario raised against the MCP repository. In the illustrative version the lead maintainer used to explain it: a person searches for an MCP server to manage travel plans, lands on a page that looks legitimate, and is told the server that books hotels lives at a particular URL. The person authorizes their bot to connect. If the token issued during that flow could be replayed against a different server, the wrong site could end up holding a working credential, even though the person only ever meant to authorize one destination. Requiring the canonical URI of the target server in the resource parameter lets an authorization server bind the token to that destination at issuance time. This only closes the replay path when the authorization server honors the parameter and the receiving MCP server rejects tokens issued for another audience. The client parameter alone is not a complete defense. This kind of problem has a name in the wider identity world: a confused deputy attack, where a credential meant for one purpose gets used for another because nothing checked that the purpose matched. MCP's resource parameter requirement is a targeted fix for that specific shape of failure, applied to the tool-connection pattern that MCP formalizes. Sources: [Update To MCP Authorization Spec - Resource Parameter (RFC 8707) | Den Delimarsky](https://den.dev/blog/mcp-authorization-resource/), [Authorization - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/basic/authorization). ## A worked example: two MCP servers, one bot This is illustrative, not a real deployment. Say a personal bot is connected to two MCP servers: one that manages a calendar, one that files expense reports. Both authorization servers sit behind the same identity provider, because that's convenient to set up and is a common pattern for small teams running more than one internal tool. Without a resource parameter, a token issued during the calendar connection might carry no explicit audience, or a broad one that just names the identity provider. If the expense server accepts any token from that identity provider without checking who the token was actually issued for, the calendar token could be replayed there. Nothing about the token itself would look wrong in isolation; it would just be handed to a service it was never meant to reach. If the authorization server honors the resource parameter, the calendar token's audience names the calendar server's canonical URI. The expense server, receiving that token, checks the audience, sees it does not match its own URI, and rejects it. In this correctly configured example, a wiring mistake or deliberate replay attempt fails at the audience check. ## What to check before you trust a connection The requirement only protects you if both sides implement it correctly. A client sending the resource parameter does nothing if the server on the other end ignores the audience claim it receives, or checks it loosely enough that a near-match passes. - Does the client send a resource parameter in both the authorization request and the token request, naming the MCP server's canonical URI? - Does the server validate the token's audience claim against its own canonical URI, using exact matching rather than a prefix or wildcard match? - Does the setup reject a token that arrives with no audience claim at all, rather than treating a missing claim as implicitly valid? - If you operate more than one MCP server behind the same identity provider, has each server's canonical URI been registered distinctly, so tokens for one cannot be mistaken for another? Sources: [Authorization - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/basic/authorization). ## Timeline and current status The resource-indicator requirement was described by MCP lead maintainer Den Delimarsky on June 18, 2025. It is not a new requirement from the July 2026 specification release. That later release made other protocol and authorization changes, as the maintainers' release post and roadmap describe. The specification still requires the client to send the resource parameter and the MCP server to validate that tokens were issued for it. Den also noted an implementation limit when the requirement was introduced: an authorization server that does not support the parameter may ignore it. Sending the parameter does not, by itself, prove that a token is audience-bound. If you are building or reviewing an MCP connection, check the current specification revision and test the full flow: client request, authorization-server token issuance, and MCP-server audience validation. BotBento is in development and does not currently offer an automated connection review for this check. Sources: [Update To MCP Authorization Spec - Resource Parameter (RFC 8707) | Den Delimarsky](https://den.dev/blog/mcp-authorization-resource/), [Authorization - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/basic/authorization), [The 2026-07-28 Specification | Model Context Protocol Blog](https://blog.modelcontextprotocol.io/posts/2026-07-28/), [The New MCP Roadmap | Model Context Protocol Blog](https://blog.modelcontextprotocol.io/posts/mcp-roadmap/). ## Primary sources Sources checked 2026-09-23. Standards and product documentation can change; follow the linked version when implementing. - [Authorization - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/basic/authorization) — Model Context Protocol - [Update To MCP Authorization Spec - Resource Parameter (RFC 8707) | Den Delimarsky](https://den.dev/blog/mcp-authorization-resource/) — Den Delimarsky - [The 2026-07-28 Specification | Model Context Protocol Blog](https://blog.modelcontextprotocol.io/posts/2026-07-28/) — Model Context Protocol - [The New MCP Roadmap | Model Context Protocol Blog](https://blog.modelcontextprotocol.io/posts/mcp-roadmap/) — Model Context Protocol BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/mcp-roots-deprecated-migration/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 4 MIN READ # MCP Roots Is Deprecated. Where Should Your Bot's File Boundaries Live Now? MCP Roots is deprecated in the 2026-07-28 specification and cannot be removed before the first revision released on or after 2027-07-28. The old standalone roots/list request works on legacy connections, while the 2026 protocol carries roots/list inside a Multi Round-Trip Request. Neither form enforces filesystem access. This guide distinguishes those wire paths and shows how to move scoping to tool inputs, resource URIs, or server configuration with real OS-level controls. By BotBento Editorial · Published 2026-09-30 · Updated 2026-09-30 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [MCP Deprecated Roots in July 2026. What Exactly Happened](#what-changed) - [Roots Never Enforced Anything — So What Are You Actually Losing?](#why-roots-never-enforced) - [Worked Example: Migrating a Filesystem Server Off Roots](#worked-example) - [What to Do With Your Bot Today](#decision-now)[Read as Markdown ](/text/blog/mcp-roots-deprecated-migration/index.md) ## Key takeaways - Roots is deprecated in the 2026-07-28 spec; the first eligible removal revision is on or after 2027-07-28, and actual removal is a maintainer decision. - The old standalone roots/list request needs a legacy connection. A modern 2026 request can carry roots/list inside an InputRequiredResult, but new implementations should use the documented migration paths. - Roots are advisory, not filesystem access control. Validate paths in the server and restrict the process with permissions or sandboxing. ## MCP Deprecated Roots in July 2026. What Exactly Happened The Model Context Protocol's 2026-07-28 revision marked Roots as deprecated under a new feature lifecycle policy. The official spec page for Roots now states this plainly: new implementations should stop adopting it, and existing ones should move on. This wasn't an isolated change. The same revision deprecated Sampling and Logging alongside Roots, all under one proposal. The registry entry lists the deprecation SEP for all three as SEP-2577, deprecated in the 2026-07-28 revision, with migration paths spelled out and an earliest removal window of a revision released on or after 2027-07-28. Deprecated does not mean immediately removed. Roots remains specified in the 2026-07-28 revision, and the first revision eligible to remove it is one released on or after 2027-07-28; removal is still a Core Maintainer decision. But wire behavior matters: the older standalone server-to-client roots/list request works on a legacy, pre-2026 connection, not on a modern stateless connection without a back-channel. The 2026 protocol can instead carry a roots/list input request inside an InputRequiredResult and complete it through Multi Round-Trip Requests. Treat existing integrations according to the protocol version they use, rather than assuming an old request works unchanged everywhere. Sources: [Roots - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/client/roots), [Deprecated Features - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/deprecated), [The 2026-07-28 Specification](https://blog.modelcontextprotocol.io/posts/2026-07-28/), [Deprecated features - MCP Python SDK](https://py.sdk.modelcontextprotocol.io/deprecated/). ## Roots Never Enforced Anything — So What Are You Actually Losing? Before deciding what to migrate to, it helps to be honest about what Roots ever gave you. The MCP client-concepts documentation is direct about this: roots serve as a coordination mechanism between clients and servers, not a security boundary. The spec only required servers to 'SHOULD respect root boundaries,' not 'MUST enforce' them, because servers run arbitrary code the client cannot fully control. In practice that meant a compliant client would advertise a directory as a root, and a well-behaved server would stay inside it — but nothing in the protocol stopped a careless or malicious server from reading outside that boundary. Real enforcement always had to happen at the operating-system level: file permissions, container mounts, or bind mounts that physically restrict what a process can touch. So the deprecation mostly formalizes something that was already true: if you were treating roots as your access-control layer, you were relying on an honor system. Losing roots as a first-class feature doesn't remove a security control you had — it removes a coordination hint you were probably over-trusting. Sources: [Understanding MCP clients - Model Context Protocol](https://modelcontextprotocol.io/docs/2026-07-28/learn/client-concepts). ## Worked Example: Migrating a Filesystem Server Off Roots Consider a hypothetical filesystem MCP server for a personal research bot. In a legacy, handshake-era integration, the client declared a roots capability, the server sent a standalone roots/list request, and it received a file:// URI for a notes folder. The server used that advisory location as context for later tool calls. A modern 2026 connection has no back-channel for that old standalone request; a roots/list input request would have to travel through the new Multi Round-Trip Request flow. The specification offers three migration paths for new work: pass directories or files through tool parameters, resource URIs, or server configuration. For a fixed notes folder, the server could receive an allowed directory at startup and validate every tool path against it. If client input is genuinely needed during a call, the Python SDK documents putting a ListRootsRequest inside an InputRequiredResult. That is a different wire flow from calling the old standalone roots/list method on a modern connection. Concretely, that means three tools that used to assume '/reports is the root' now each take a path argument, and your server validates that argument against a server-configured allowlist before touching disk. This is more code up front — you're writing the validation instead of leaning on a protocol field — but it also means the boundary is enforced by your own server logic and, ideally, OS-level permissions, rather than by a value a client merely suggested. One failure path worth naming: keeping the old standalone roots/list call while switching a connection to the 2026 stateless protocol can fail because there is no server-to-client back-channel. Keeping the old flow on a legacy connection still leaves the root advisory, so a server-side allowlist and OS-level restrictions are needed either way. Sources: [Roots - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/client/roots), [The 2026-07-28 Specification](https://blog.modelcontextprotocol.io/posts/2026-07-28/), [Deprecated features - MCP Python SDK](https://py.sdk.modelcontextprotocol.io/deprecated/). ## What to Do With Your Bot Today If you're building a new MCP server or bot integration, do not adopt Roots as a new dependency. The deprecation registry lists tool parameters, resource URIs, and server configuration as migration paths. Choose based on who should set the scope: a caller-provided path can be a tool argument, while an administrator-controlled boundary belongs in server configuration and filesystem permissions. Validate each input path against that boundary. If you maintain an existing server, inventory both its Roots use and the protocol versions its clients negotiate. The old standalone request can still work on legacy connections, but it does not carry over unchanged to a 2026 stateless connection. Roots cannot be removed from the specification before the first revision released on or after 2027-07-28; the date is an eligibility floor, not a guaranteed removal date or a clock that restarts with every release. Plan migration against the published registry and your SDK's support policy. The one thing not to do is treat this purely as a protocol-compliance chore. Because roots were never an enforced boundary, migrating away from them is a good moment to ask whether your server has real access control at all — an OS-level sandbox, a container mount, or at minimum a hard-coded allowlist checked on every call. If the honest answer is 'no, we just trusted the client's root,' fix that as part of the migration, not after it. Sources: [Deprecated Features - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/deprecated), [The 2026-07-28 Specification](https://blog.modelcontextprotocol.io/posts/2026-07-28/), [Deprecated features - MCP Python SDK](https://py.sdk.modelcontextprotocol.io/deprecated/), [Feature Lifecycle and Deprecation Policy](https://modelcontextprotocol.io/community/feature-lifecycle). ## Primary sources Sources checked 2026-09-30. Standards and product documentation can change; follow the linked version when implementing. - [Roots - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/client/roots) — Model Context Protocol - [Deprecated Features - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/deprecated) — Model Context Protocol - [The 2026-07-28 Specification](https://blog.modelcontextprotocol.io/posts/2026-07-28/) — Model Context Protocol - [Understanding MCP clients - Model Context Protocol](https://modelcontextprotocol.io/docs/2026-07-28/learn/client-concepts) — Model Context Protocol - [Deprecated features - MCP Python SDK](https://py.sdk.modelcontextprotocol.io/deprecated/) — Model Context Protocol - [Feature Lifecycle and Deprecation Policy](https://modelcontextprotocol.io/community/feature-lifecycle) — Model Context Protocol BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/mcp-sampling-deprecated/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 4 MIN READ # MCP Sampling Is Deprecated. What Should Bot Builders Do Now? MCP deprecated Sampling and Roots in protocol version 2026-07-28. Existing integrations have at least twelve months before either feature is eligible for removal; builders should inventory usage and plan the migrations named in the specification. By BotBento Editorial · Published 2026-09-25 · Updated 2026-09-25 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [What changed in the MCP spec](#what-changed) - [How sampling worked, briefly](#how-sampling-worked) - [Roots is also deprecated, but it was never an access-control boundary](#roots-also-deprecated) - [Why this matters if your bot connects to MCP servers](#why-it-matters) - [What to do now](#what-to-do-now)[Read as Markdown ](/text/blog/mcp-sampling-deprecated/index.md) ## Key takeaways - Sampling let an MCP server request a model generation through a supporting client; it is deprecated, not yet removed. - The specification keeps deprecated features for at least twelve months before removal eligibility; that is a minimum, not a guaranteed removal date. - New integrations should avoid Sampling and Roots. Review model-provider ownership separately from filesystem permissions. ## What changed in the MCP spec The Model Context Protocol's sampling feature is deprecated as of protocol version 2026-07-28, tracked under SEP-2577. Under the protocol's feature lifecycle policy, a deprecated feature stays in the specification for at least twelve months after the deprecating revision before it can even be considered for removal, so nothing breaks immediately for bots that already rely on it. The specification advises new implementations not to adopt Sampling and asks existing ones to migrate toward direct LLM provider APIs. Its Roots page separately deprecates the client feature for listing relevant filesystem locations and recommends tool parameters, resource URIs, or server configuration instead. These are two distinct migration paths; the specification does not say that either change removes the need for client consent or server-side access controls. Sources: [Sampling - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/client/sampling), [Roots - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/client/roots). ## How sampling worked, briefly Sampling gave MCP servers a way to request a completion from a language model without holding an API key of their own. The protocol provided a standardized way for servers to request LLM sampling from language models via clients, letting the client keep control over model access, selection, and permissions while the server borrowed that capability to power its own logic. A server might use this to summarize a document, classify an input, or draft a reply, all without ever seeing a raw provider credential. Because a server calling into your model is a meaningful trust boundary, the spec built in a safeguard: for trust and safety, there should always be a human in the loop with the ability to deny sampling requests, and applications were expected to make reviewing those requests easy. That review step, not the sampling call itself, was the main control a bot owner had over what a connected server could ask a model to do. Sources: [Sampling - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/client/sampling), [Sampling - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-06-18/client/sampling). ## Roots is also deprecated, but it was never an access-control boundary The 2026-07-28 revision also deprecates Roots. A root tells a supporting server which filesystem locations the client considers relevant. The specification explicitly says roots are informational guidance, not access control, and the protocol does not force a server to remain inside them. A server that needs filesystem access still needs its own permission and path checks. An illustrative review: if a document-search server reads only one project folder today, first find out whether that limit comes from the server's actual permissions or merely from a Roots hint. Then ask its maintainer how it will receive relevant paths after migration. Passing a path as a tool parameter changes how the server discovers it; it does not authorize reading that path by itself. Sources: [Roots - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/client/roots). ## Why this matters if your bot connects to MCP servers Bot builders often connect a client or agent to MCP servers built by another team. A server may request Sampling only if the client advertises support; the server does not declare a Sampling capability on the client's behalf. If a server uses that path, the client mediates model access and can present the request for human review. A direct provider integration moves the model request into the server's own provider relationship. Ask the server owner what model, credentials, data handling, and approval controls replace the client-mediated path. Illustrative example: imagine a small MCP server called 'invoice-explainer' that used sampling to ask your client's model to summarize uploaded invoices in plain language. Under the old design, every summary request passed through your client, which meant your approval gate and your model choice governed what happened. If invoice-explainer's maintainer migrates to a direct provider API, that summary request stops touching your client at all. Your bot loses visibility into that call, and the server takes on its own cost and model choice. That is a meaningful shift in who is accountable for what the tool does, even though the end result — a summarized invoice — looks the same to the user. It also means a security review of invoice-explainer after migration needs to ask a different question: not 'what can it ask my model to do,' but 'what provider account does it use, and what data does it send there.' Sources: [Sampling - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/client/sampling). ## What to do now Treat the deprecation as an inventory and migration task, not an emergency. The twelve-month minimum is a floor for removal eligibility, not a promise of a particular removal date or of universal implementation support. Start by checking the capabilities your client advertises and the MCP servers that actually request Sampling or Roots. Record which server owns each request and how a human reviews it. - Inspect the client's advertised Sampling and Roots capabilities and the requests made by each connected server. - Ask server maintainers how any Sampling use will migrate to a direct provider integration and how model data and cost will be governed. - Do not design new integrations around deprecated features. - Keep a human review path for sensitive model requests while Sampling remains in use. - Treat Roots as guidance, never as the sole filesystem permission boundary; verify actual server-side controls. Sources: [Sampling - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/client/sampling), [Roots - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/client/roots). ## Primary sources Sources checked 2026-09-25. Standards and product documentation can change; follow the linked version when implementing. - [Sampling - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/client/sampling) — Model Context Protocol - [Sampling - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-06-18/client/sampling) — Model Context Protocol - [Roots - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/client/roots) — Model Context Protocol BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/mcp-stateless-cross-call-state/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 4 MIN READ # MCP Dropped Sessions. How Do You Keep State Across Tool Calls Now? The Model Context Protocol's 2026-07-28 revision removed protocol-level sessions and the Mcp-Session-Id header, so a bot can no longer rely on the transport to remember it between tool calls. The spec's replacement is an explicit handle: a tool returns an identifier, and the model passes it back as an ordinary argument on later calls. This piece walks through what changed, a worked example of the handle pattern, and the specific things that break if your bot's MCP integration still assumes a session. By BotBento Editorial · Published 2026-09-28 · Updated 2026-09-28 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [What the 2026-07-28 spec actually removed](#what-changed) - [The replacement: a handle the model carries, not a session the transport hides](#the-fix) - [What breaks if your bot still assumes sessions](#what-breaks) - [A short checklist before you touch a bot's MCP integration](#checklist) - [Why a single-instance bot should still care](#why-it-matters-for-small-bots)[Read as Markdown ](/text/blog/mcp-stateless-cross-call-state/index.md) ## Key takeaways - MCP 2026-07-28 removed the initialize handshake and the Mcp-Session-Id header; requests carry their protocol context and can land on any compatible server instance. - The replacement is an explicit, server-minted handle (like a booking\_id) that the model carries as a tool argument, not hidden session state. - Check three things before you touch a bot's MCP integration: session-ID routing logic, a hardcoded -32002 error check, and any reliance on Roots, Sampling, or Logging. ## What the 2026-07-28 spec actually removed The Model Context Protocol's 2026-07-28 revision makes the protocol stateless at its core. The changelog is explicit about the mechanics: the spec removes the initialize and notifications/initialized handshake, and it removes protocol-level sessions, including the Mcp-Session-Id header, from the Streamable HTTP transport. List endpoints such as tools/list, resources/list, and prompts/list no longer vary per connection either, since there's no connection-scoped state left for them to vary against. Requests now carry protocol version and client capabilities in their metadata instead of relying on a negotiated session. The new specification says clients should also identify themselves on each request; that identity field is recommended, not a mandatory claim about every message. A compatible request can land on any server instance behind an ordinary load balancer without sticky protocol-session routing. If your integration depends on the old Mcp-Session-Id or connection-scoped state, test it against the negotiated protocol version before changing it. Sources: [Key Changes - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/changelog), [The 2026-07-28 Specification](https://blog.modelcontextprotocol.io/posts/2026-07-28/). ## The replacement: a handle the model carries, not a session the transport hides The spec's own guidance on this is direct: if a server needs to carry state across calls, it should mint an explicit handle from a tool and have the model pass that handle back as an argument. The maintainers describe this as working better than session state hidden in the transport, because the model can see the handle and thread it between tool calls itself, rather than the server keeping track of who's calling on its behalf. Worked example, illustrative only: say a small bot manages hotel bookings through an MCP server. An older implementation might have kept an in-progress booking tied to a protocol session. Under the new pattern, a start\_booking tool returns booking\_id; the model passes that identifier to add\_room and confirm\_booking. The server can still authenticate and authorize the caller through its normal credentials. The booking\_id identifies the in-progress work, not the caller's permission to change it. Check both the handle and the caller's authority on each step. This illustrates the explicit state pattern the maintainers recommend; it is not a BotBento feature claim. Sources: [The 2026-07-28 Specification](https://blog.modelcontextprotocol.io/posts/2026-07-28/). ## What breaks if your bot still assumes sessions Three migration checks are worth running against a client and server that actually negotiate 2026-07-28. First, a gateway that requires Mcp-Session-Id for routing must be updated because modern requests no longer send it. Second, the resource-not-found code changed from -32002 to -32602 (Invalid Params); handle the code according to the negotiated version rather than replacing the old branch for legacy servers. Third, tool logic that assumes the next call reaches the same process must move its work state into a durable record addressed by an explicit handle, or provide another safe recovery path. Roots, Sampling, and Logging are formally deprecated, with a minimum twelve-month deprecation window. That does not mean their old server-to-client methods work in a new stateless 2026-07-28 session: the Python SDK says those calls need a legacy connection, while the modern flow uses Multi Round-Trip Requests or other replacements. If a tool needs user or model input mid-call, review the new input-required flow. For filesystem paths, consider tool parameters or resource URIs; for server logs, use ordinary logging or telemetry. Keep legacy behavior only where the negotiated older version supports it. Sources: [Key Changes - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/changelog), [The 2026-07-28 MCP Specification Release Candidate](https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/), [Deprecated features - MCP Python SDK](https://py.sdk.modelcontextprotocol.io/deprecated/). ## A short checklist before you touch a bot's MCP integration None of this requires a rewrite. It requires finding every place your bot's tooling assumed a session was doing work for it, and replacing that assumption with an explicit value the model can see and carry forward on its own. - Confirm the SDK your bot's MCP client and server use actually speaks 2026-07-28 — the four Tier 1 SDKs (TypeScript, Python, Go, C#) shipped support on release day, with Rust in beta. - Search your integration code for any check on Mcp-Session-Id or a stored per-connection object, and replace it with a handle argument the model passes back explicitly. - Search for a hardcoded -32002 check on missing resources and handle -32602 when the peer negotiates 2026-07-28; retain the legacy branch if you support older peers. - If you use Roots, Sampling, or Logging, test a modern connection separately from a legacy one and plan the appropriate replacement rather than assuming old callbacks work on both. Sources: [Key Changes - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/changelog), [Deprecated features - MCP Python SDK](https://py.sdk.modelcontextprotocol.io/deprecated/). ## Why a single-instance bot should still care A personal or small-team bot usually runs on one process, not behind a load balancer, so the routing story may not bite directly. But the handle pattern is worth adopting anyway, for a reason that has nothing to do with scaling: state that's visible to the model as an argument is state you can log, replay, and debug after the fact. Hidden session state that lived only in server memory disappears the moment the process restarts, and it leaves no trace in your run records. If you're already recording what a bot did after each run, an explicit handle gives you something concrete to write down — a booking\_id, a ticket\_id, a cart\_id — instead of a vague note that 'the session had state.' That's a small design change with an outsized effect on whether you, or anyone reviewing the bot's work later, can actually verify what it was doing partway through a multi-step task. Sources: [The 2026-07-28 Specification](https://blog.modelcontextprotocol.io/posts/2026-07-28/). ## Primary sources Sources checked 2026-09-28. Standards and product documentation can change; follow the linked version when implementing. - [The 2026-07-28 Specification](https://blog.modelcontextprotocol.io/posts/2026-07-28/) — Model Context Protocol - [The 2026-07-28 MCP Specification Release Candidate](https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/) — Model Context Protocol - [Key Changes - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/changelog) — Model Context Protocol - [Deprecated features - MCP Python SDK](https://py.sdk.modelcontextprotocol.io/deprecated/) — Model Context Protocol BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/mcp-tool-call-cancellation-what-happens/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 4 MIN READ # What Actually Happens When Your Bot Cancels an MCP Tool Call? A bot that cancels a slow MCP tool call cannot assume the server stopped its work. The current protocol asks servers to stop, while its stdio and Streamable HTTP transports prohibit further messages for a cancelled request. On Streamable HTTP, the client cancels by closing the response stream. Builders must stop waiting locally and separately check any side effect that might already have happened. By BotBento Editorial · Published 2026-10-01 · Updated 2026-10-01 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [Cancellation is a request, not a command](#the-rule) - [The mechanism depends on the transport](#stdio-vs-http) - [Worked example: cancelling a slow report tool](#worked-example) - [Where real implementations drift from the spec](#failure-paths) - [What to build instead of trusting cancellation](#decision)[Read as Markdown ](/text/blog/mcp-tool-call-cancellation-what-happens/index.md) ## Key takeaways - Stopping work is advisory: servers SHOULD stop. Under the 2026 transport rules they MUST NOT send further messages for a cancelled request, even if work continues. - On stdio, your bot sends notifications/cancelled; on Streamable HTTP, closing the response stream is the cancellation signal and no notification is sent at all. - After cancelling, stop waiting for a response; a late result may have been in flight already, and silence does not prove a side effect was prevented. ## Cancellation is a request, not a command The current Model Context Protocol lets a client request cancellation of an in-progress call. The cancellation pattern says a server SHOULD stop processing and free resources; it MAY be unable to stop a request that has already completed or cannot be cancelled. The transport rules are stricter about the wire: after cancellation, a server MUST NOT send further messages for that request on stdio or Streamable HTTP. That distinction matters for anyone building a bot that calls MCP tools. Your bot can ask a server to stop generating a report, scraping a page, or running a long query, but it cannot take silence as proof that the work stopped. A response may also have been sent before the cancellation arrived and reach the client afterward. The client must tolerate that race without acting on the late result. Sources: [Cancellation - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/basic/patterns/cancellation), [stdio - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/basic/transports/stdio), [Streamable HTTP - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/basic/transports/streamable-http). ## The mechanism depends on the transport How you cancel changes depending on how your bot is connected to the server. Over stdio in the 2026 protocol, the client MUST send notifications/cancelled with the request ID. The server SHOULD stop the work as soon as practical and MUST NOT send further messages for that request. Over Streamable HTTP, there's no cancellation message at all in the current protocol revision. Closing the SSE response stream is itself the cancellation signal, and the server must treat that disconnect as cancellation of the request riding on it. The spec is direct about this: since each request gets its own response stream, the disconnect is unambiguous, and the server should stop work as soon as practical and must not send any further messages for that request. If your bot's HTTP client library doesn't actually close the underlying connection when you call an abort method — it just stops reading from it — the server may never learn the request was cancelled. Sources: [stdio - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/basic/transports/stdio), [Streamable HTTP - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/basic/transports/streamable-http). ## Worked example: cancelling a slow report tool Illustrative scenario. A bot calls a tool named generate\_quarterly\_summary over stdio. The tool normally returns in 10 seconds but on this run it's still running at 45 seconds, past the bot's own patience threshold. The bot sends notifications/cancelled with the original request's ID and reason: "exceeded 30s budget," then immediately treats the call as done locally — it does not keep a slot open waiting for a reply. Three things can happen next, and the bot's logic has to tolerate all three without crashing. First, the server stops the work cleanly and sends nothing back. Second, the server sent a result before it received the notification, but the result reaches the bot afterward; the cancellation pattern tells the sender to ignore that late response, so the bot's run record should say 'cancelled, late result discarded' rather than treating the payload as a valid output. Third, the server cannot stop the work and completes a report file anyway; under the current stdio rules it still must not send further messages for the cancelled request. The bot cannot distinguish the third case from the first through silence alone. It needs a separate read to check whether the report file was created. ## Where real implementations drift from the spec A reported Go SDK v1.7.0 issue illustrates the gap between current transport rules and an implementation: after a client sends notifications/cancelled, the server can write a response when its handler returns. The issue's reproduction shows the handler context being cancelled while a response is still written. That is an implementation report, not evidence that all Go SDK versions or MCP servers behave this way. The separate operational point is about the work itself. Even a server that correctly sends no further messages may already have written a file, sent an email, or started another side effect before it processes the cancellation. Do not build retry or state logic that treats the absence of a response as proof that the operation was undone. Sources: [A response is still written for a request after notifications/cancelled · Issue #1235 · modelcontextprotocol/go-sdk](https://github.com/modelcontextprotocol/go-sdk/issues/1235). ## What to build instead of trusting cancellation Treat a sent cancellation as a request you fired and forgot, not as a stop you confirmed. Three concrete rules follow from that. First, when your bot decides to give up on a tool call, cancel using the transport's mechanism (notifications/cancelled on stdio, closing the response stream on current Streamable HTTP), then release your own wait state. The cancellation pattern's timeout guidance says to stop waiting, and a response already in flight should not be treated as a successful tool result. Second, for any tool call whose cancellation matters beyond saving a few seconds of latency — anything with a side effect your bot would need to undo — add a separate verification step after cancelling. Check whether the effect happened through an idempotent read, not through the tool call's response. Third, log cancellation attempts and their outcomes separately from normal failures in your run record. 'Cancelled, outcome unverified' is a different and more honest status than 'failed' or 'succeeded,' and it tells whoever reviews the run exactly what still needs checking by hand. ## Primary sources Sources checked 2026-10-01. Standards and product documentation can change; follow the linked version when implementing. - [Cancellation - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/basic/patterns/cancellation) — Model Context Protocol - [Streamable HTTP - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/basic/transports/streamable-http) — Model Context Protocol - [stdio - Model Context Protocol](https://modelcontextprotocol.io/specification/2026-07-28/basic/transports/stdio) — Model Context Protocol - [A response is still written for a request after notifications/cancelled · Issue #1235 · modelcontextprotocol/go-sdk](https://github.com/modelcontextprotocol/go-sdk/issues/1235) — GitHub (modelcontextprotocol/go-sdk) BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/mcp-tool-output-schema-validation/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 4 MIN READ # MCP Tools Can Declare an Output Schema. Should Your Bot Validate It? MCP tool definitions can declare an outputSchema and successful results can carry structuredContent. A bot should check that result against the declared schema before acting on it, while handling tool execution errors and request failures on their own paths. By BotBento Editorial · Published 2026-09-29 · Updated 2026-09-29 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [What actually changed in the tool result](#what-changed) - [Shape is a different question from success](#shape-vs-success) - [A worked example (illustrative)](#worked-example) - [Where this breaks in practice](#failure-paths) - [The decision](#the-decision)[Read as Markdown ](/text/blog/mcp-tool-output-schema-validation/index.md) ## Key takeaways - outputSchema describes the expected shape of structuredContent; isError flags a tool execution error in a returned tool result, while transport or protocol failures may have no tool result to inspect. - For a successful call to a tool that declares outputSchema, treat absent or nonconforming structuredContent as a result-contract failure before acting on the data. - Keep a documented path for tools without outputSchema; do not use an unvalidated text block to bypass a failed structured-result check. ## What actually changed in the tool result MCP tools may return a content array with text or other blocks, and a Tool definition may include an optional outputSchema alongside inputSchema. The 2025-11-25 tools specification already defines outputSchema and structuredContent; they are not newly introduced by the 2026 revision. A weather tool, for example, can return a human-readable summary and structured JSON so a client need not parse prose to find the temperature. The 2026-07-28 specification expands what those fields can express. Its output schemas support full JSON Schema 2020-12, and structuredContent can be any JSON value, including an array or a scalar, rather than only an object. The specification says implementations must not automatically dereference external $ref URIs and should bound schema depth and validation time. The useful promise is an expected, machine-readable result shape. It does not prove the data is true or that the operation was authorized; clients still need their ordinary trust and action checks. Sources: [Tools - Model Context Protocol (2025-11-25)](https://modelcontextprotocol.io/specification/2025-11-25/server/tools), [Tools - Model Context Protocol (2026-07-28)](https://modelcontextprotocol.io/specification/2026-07-28/server/tools). ## Shape is a different question from success A returned tool result can use isError: true to report a tool execution error, such as an underlying API failure or invalid input. A request timeout, transport failure, or JSON-RPC protocol error may instead leave the caller without a CallToolResult at all. Handle those request failures before examining result fields; isError is not a universal success signal for every attempted call. For a completed, non-error result, outputSchema and structuredContent answer a different question: does the returned JSON match the shape the tool advertised? A server can return a non-error result with a missing required field or a renamed property. If the client passes it downstream without validation, the next action may break or use the wrong value. The tools specification says servers with an outputSchema must provide conforming structured results and clients should validate them. Sources: [Tools - Model Context Protocol (2026-07-28)](https://modelcontextprotocol.io/specification/2026-07-28/server/tools). ## A worked example (illustrative) Say a bot calls a fictional invoice-lookup tool. The tool's listing declares an outputSchema requiring amount\_cents (integer), currency (string), and status (one of paid, pending, overdue). On a normal call, the result carries a content block with a human-readable summary and a structuredContent object matching that shape: \{"amount\_cents": 4200, "currency": "USD", "status": "pending"}. Now suppose the invoice service changes its API and the tool handler starts returning status as a nested object (\{"code": "pending", "label": "Awaiting payment"}) without updating its declared outputSchema. The call returns a non-error result, but structuredContent no longer matches the advertised schema. A bot that assumes status is a string may reject the value too late or pass the wrong shape to a later step. Before using structuredContent, check that it is present for a successful result when the tool declares outputSchema, then validate it with a JSON Schema validator configured for the declared dialect and bounded reference resolution. On mismatch, log the contract failure and avoid acting on that data. A text block from the same response is not a safe bypass unless the tool has a separately documented fallback contract that you validate too. Retrying unchanged may return the same malformed value; route the mismatch to the tool owner when appropriate. ## Where this breaks in practice One failure path is validating an error result as though it were a successful structured result. When isError is true, handle the tool execution error first; an error response may be text meant for recovery rather than the success payload described by outputSchema. Keep request or transport failures, which may produce no tool result, on a separate path. Another failure path is asymmetric adoption. A tool without outputSchema may return only content blocks. A bot that assumes structured output everywhere will break on those tools. Keep a documented text-handling path for tools that do not advertise an output schema, while requiring schema validation for successful structured responses from tools that do. A third failure path is treating schema validation as a security boundary. It only checks shape. A tool can return a schema-conforming value that is false, stale, or unsafe to use. Keep authorization, provenance, and sanity checks independent of output validation. ## The decision When a tool advertises outputSchema, first determine whether the request returned a tool result. Handle request failures and isError tool errors separately. For a successful result, require structuredContent and validate it against the declared schema before the bot acts on it. When a tool has no outputSchema, use the tool’s documented content contract rather than inventing a schema it never promised. This prevents a bot from quietly acting on data that does not match the tool’s advertised shape. Validation is one check in the result path, alongside the existing authorization and trust checks; it cannot turn untrusted tool data into a verified fact. ## Primary sources Sources checked 2026-09-29. Standards and product documentation can change; follow the linked version when implementing. - [Tools - Model Context Protocol (2025-11-25)](https://modelcontextprotocol.io/specification/2025-11-25/server/tools) — Model Context Protocol - [Tools - Model Context Protocol (2026-07-28)](https://modelcontextprotocol.io/specification/2026-07-28/server/tools) — Model Context Protocol BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/mcp-tool-permissions/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 5 MIN READ # MCP tool permissions: what to check before connecting a bot Connecting an MCP server makes tools available. It does not, by itself, answer which account a bot can access or which actions it should perform. Check the connection, the scope and the action separately. By BotBento Editorial · Published 2026-09-07 · Updated 2026-09-07 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [Start by naming the parts](#name-the-parts) - [Separate connection, authorization and action](#separate-the-decisions) - [Write an action you can inspect](#make-a-specific-request) - [A short review before the first useful task](#a-connection-review) - [Test the boundary as well as the happy path](#test-a-denial) - [What a connected status cannot prove](#what-a-connection-cannot-prove)[Read as Markdown ](/text/blog/mcp-tool-permissions/index.md) ## Key takeaways - Identify the host application, MCP server and downstream service before authorizing a connection. - A provider authorization grant and permission for a particular bot action are different decisions. - Test a harmless read, a denied action and revocation before relying on a new integration. ## Start by naming the parts Model Context Protocol, or MCP, describes a way for an AI application to connect to servers that expose capabilities such as tools and resources. The host application manages the user-facing experience; an MCP client handles its connection to a server. A server may then interact with another service. These are separate components even when an installation screen presents them as one integration. For an illustrative calendar connection, name three things before continuing: the application in which you talk to the bot, the server that provides calendar tools, and the calendar account those tools will use. If you cannot identify one of them, a successful connection message is not enough information to decide what access you are granting. Sources: [MCP architecture overview](https://modelcontextprotocol.io/docs/2026-07-28/learn/architecture). ## Separate connection, authorization and action The July 2026 MCP authorization specification defines an authorization flow for HTTP-based transports. It describes MCP servers acting as protected resources and clients obtaining tokens to access them. It also distinguishes HTTP authorization from standard-input/output connections, where credentials are typically supplied through the environment. The transport matters when you assess how a particular server gets its access. The same specification recommends requesting scopes needed for the intended operation and describes resource-bound token requests. In plain terms, ask what the granted permission covers and which server the credential is intended for. A broad provider grant can permit more than the single task you currently want. The grant should not silently become an instruction for the bot to exercise every available capability. For the calendar example, connecting the server answers whether the application can communicate with it. Selecting the account answers whose calendar is involved. Authorizing a scope determines what the connected software can technically request. Asking the bot to summarize tomorrow’s meetings is a narrower instruction. None of those steps should be treated as a substitute for the others. Sources: [MCP authorization specification, 2026-07-28](https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization). ## Write an action you can inspect Compare “sort out my calendar” with this illustrative instruction: “Read tomorrow’s events in my work calendar, identify overlaps, and draft a suggestion. Do not create, edit, cancel or message anyone.” The second request produces a reviewable result and identifies the actions outside its scope. It still needs a correctly selected account and a service implementation that respects the access boundary. If a later request involves moving a meeting, describe the exact event, intended time and attendees before committing the change. A useful confirmation screen should show those details in a form the person can check. “Allow calendar access” may describe a technical grant; it does not tell someone whether a particular meeting is about to be rescheduled. The MCP tools specification recommends that applications keep people able to deny tool calls, display available tools and indicate their use. It also says tool annotations should be treated as untrusted unless they come from trusted servers. Our practical implication is to check the actual tool behavior and account permissions instead of relying only on a friendly name or a reassuring label. Sources: [MCP tools specification, 2026-07-28](https://modelcontextprotocol.io/specification/2026-07-28/server/tools). ## A short review before the first useful task The following is an illustrative review process for a new integration. It is deliberately specific enough to produce observations you can write down. The right controls depend on your host, server, service account and data. It is not a certification of any server or a claim that a protocol removes the need for implementation review. - Identify the operator: record the server URL or local package source, its publisher and the downstream service it uses. An unexpected operator is a reason to investigate before entering credentials. - Check the selected account: read the visible account name in the provider flow and compare it with the account needed for the task. Do not infer the selected account from the browser profile alone. - Read the access request: separate reading records from creating, changing, deleting or sharing them. If the requested access is broader than your task, find out whether a narrower integration or account is available. - Inspect the first result: use a harmless, recognizable record that you are authorized to access. Check that the result comes from the intended account and that the tool did not change anything. - Find the off switch: locate connection removal and provider revocation before you depend on the integration. Disconnecting a visible bot entry and revoking an underlying grant may be separate operations. ## Test the boundary as well as the happy path A successful read only proves that one request worked. For the calendar example, use a disposable test calendar and try a request that should be denied by the configuration, such as a write when the integration is configured for reading. Check the calendar afterwards. A refusal message is helpful, but the actual absence of an unauthorized change is the more meaningful result. Next, revoke the test connection using the documented provider route. Confirm that a fresh operation cannot continue through the revoked grant. Account for cached results when interpreting the screen: seeing an old meeting summary is different from successfully fetching new calendar data. If a new operation still works, stop using the test connection until you understand which credential or session remains active. Keep these checks small and reversible. Do not test destructive operations against a real customer account. Record the host version, server version when available, granted access and observed result. That gives you a useful comparison when the integration updates. A test performed before an update is historical evidence, not proof that a later version behaves identically. ## What a connected status cannot prove A connected badge does not establish that the right account is selected, every tool is safe, all provider permissions are narrow, or the bot will interpret every request correctly. It is one status inside a larger system. Treat it as a starting point for a specific task with a specific acceptance check. For everyday use, keep a simple record of what the integration is for, what it can access and how to remove that access. Revisit the record when you add a tool or change the account. That modest habit makes an expanding collection of bot connections easier to reason about. BotBento is being developed around conversations, routines and plugins; these general notes should not be read as a claim that a particular MCP integration is already available in the product. ## Primary sources Sources checked 2026-09-07. Standards and product documentation can change; follow the linked version when implementing. - [MCP architecture overview](https://modelcontextprotocol.io/docs/2026-07-28/learn/architecture) — Model Context Protocol - [MCP authorization specification, 2026-07-28](https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization) — Model Context Protocol - [MCP tools specification, 2026-07-28](https://modelcontextprotocol.io/specification/2026-07-28/server/tools) — Model Context Protocol BotBento is in development. [Suggest a correction](/contact/). ## Keep reading - [AI agents vs workflows: choose who decides the next step](/blog/ai-agents-vs-workflows/) --- Canonical: https://botbento.com/blog/mcp-tool-result-success-signal/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 4 MIN READ # MCP Tool Results: How Do You Know a Call Actually Succeeded? An MCP response can be well formed, report no tool error, and still leave the requested real-world action unconfirmed. Distinguish protocol errors, isError, structured output and independent readback before reporting success. By BotBento Editorial · Published 2026-09-27 · Updated 2026-09-27 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [First, identify which layer answered](#two-kinds-of-failure) - [Read the result fields for what they actually say](#content-vs-structured-content) - [A valid payload is not a confirmed outcome](#schema-vs-outcome) - [Illustrative example: booking a calendar slot](#worked-example) - [A short acceptance check for each tool](#what-to-check)[Read as Markdown ](/text/blog/mcp-tool-result-success-signal/index.md) ## Key takeaways - A JSON-RPC error and a tool result with isError: true represent different failure paths; inspect the actual response before deciding whether to retry. - An outputSchema describes the shape of structuredContent. Schema validity does not prove that an external action happened. - For consequential actions, read the resulting state back from its source of truth before claiming completion. ## First, identify which layer answered The MCP tools specification distinguishes protocol-level errors from errors returned by a tool. An unknown tool or malformed request can produce a JSON-RPC error. An invoked tool can instead return a normal tool-result envelope with isError: true when its own work fails. The specification lists input problems among possible tool execution errors too, so the exact path depends on where validation happens. Do not infer from an invalid argument alone that the tool never ran. This distinction changes the next safe action. If a request was rejected before invocation, fix the request. If a tool reports a failure, use its error details to decide whether a changed attempt is appropriate. If the connection dropped after a side-effecting call and no result arrived, neither response shape proves whether the action happened. Search for the intended record, preferably using an idempotency key or other stable identifier, before retrying. Otherwise a calendar event or payment can be duplicated. Sources: [Tools - Model Context Protocol (2025-06-18)](https://modelcontextprotocol.io/specification/2025-06-18/server/tools). ## Read the result fields for what they actually say An MCP tool result has content blocks, which may contain text, images or resources. It may also have structuredContent for machine-readable data. A tool can declare an outputSchema for that structured value. In the 2025-06-18 specification, servers that provide structured content must make it conform to the declared schema, while clients should validate it. The current draft continues that distinction. A schema check protects the shape of a value; it does not verify the truth of the value or the state of an external service. The specification recommends also serializing structured data into a text content block for compatibility with older clients. Because that is a recommendation rather than a guarantee, a client should deliberately handle both fields. For example, a travel tool might return a readable summary in content and a reservation\_id in structuredContent. A display can use the summary, while a follow-up lookup uses the identifier. If either field is absent, treat that as an observed property of this particular tool result, not a license to invent the missing value. Version matters here. The 2025-06-18 page describes structuredContent as a JSON object; the current draft permits any JSON value. A client targeting a specific MCP version should validate against that version and the tool's declared schema. Do not silently apply draft behavior to an older server. Sources: [Tools - Model Context Protocol (2025-06-18)](https://modelcontextprotocol.io/specification/2025-06-18/server/tools), [Tools - Model Context Protocol (draft)](https://modelcontextprotocol.io/specification/draft/server/tools). ## A valid payload is not a confirmed outcome Suppose a calendar tool returns isError: false, a structured reservation\_id and a text block saying that a meeting was created. That is evidence of what the tool reported. It is not independent proof that the calendar now contains the meeting: a downstream provider might have timed out, the identifier might refer to a pending job, or the integration could be wrong. Validate the payload, then use a read operation against the calendar to confirm the event's identity, time and attendees before telling the user it is booked. The reverse case is also possible. A tool may report isError: true after a provider accepted a request but before the integration received its acknowledgment. Blindly retrying because the flag says error can create two events. For side-effecting tools, design the action around a stable request key, make the tool document whether retries are safe, and reconcile against provider state when the result is ambiguous. This is a practical reliability rule, not a guarantee made by the MCP result envelope. Error output needs careful schema handling. The specification says that structured content, when returned under a declared outputSchema, must conform to that schema. It also says tool execution failures should be reported using isError. It does not explicitly state that every client must skip output-schema validation for isError results. Design and test your own success and failure shapes instead of treating an implementation-specific workaround as a rule from the specification. Sources: [Tools - Model Context Protocol (2025-06-18)](https://modelcontextprotocol.io/specification/2025-06-18/server/tools), [Tools - Model Context Protocol (draft)](https://modelcontextprotocol.io/specification/draft/server/tools). ## Illustrative example: booking a calendar slot A bot asks a calendar tool to book a 2pm slot with request key meeting-4821. The tool returns isError: true and a text explanation that the slot is unavailable. The bot should offer another time, not announce a booking. If the tool supplies a structured reason code and its schema allows that shape, the bot can use the code to choose a recovery path; the readable text remains useful to explain the result. A schema that only describes successful bookings may not describe that error, so the integration needs an explicit, tested error contract. Now imagine the network times out after the same request was sent. There is no result to inspect. Retrying immediately could create a duplicate if the provider accepted the first call. The bot should query by request key or search for the event and compare its details. Only a matching record supports a claim that the booking succeeded. If the provider offers neither lookup nor idempotency, the honest user-facing state is that confirmation is pending, followed by a manual check. Sources: [Tools - Model Context Protocol (2025-06-18)](https://modelcontextprotocol.io/specification/2025-06-18/server/tools). ## A short acceptance check for each tool This check is small enough to run while integrating a tool, and it catches the mistake that matters most: promoting a successful-looking message into a claim that the real-world task is complete. BotBento's editorial team used AI assistance in drafting this guide and checked the technical claims against the linked primary specifications. - Capture one successful result and one deliberate tool failure. Record the JSON-RPC envelope, isError, content and structuredContent fields actually returned. - If outputSchema is declared, validate structuredContent using the negotiated MCP version; separately test the tool's error shape. - For writes, verify the resulting record through an independent read. Test a dropped response and a duplicate request with the same stable key. - Keep the user-facing wording aligned with the evidence: reported, pending and confirmed are different states. Sources: [Tools - Model Context Protocol (2025-06-18)](https://modelcontextprotocol.io/specification/2025-06-18/server/tools), [Tools - Model Context Protocol (draft)](https://modelcontextprotocol.io/specification/draft/server/tools). ## Primary sources Sources checked 2026-09-27. Standards and product documentation can change; follow the linked version when implementing. - [Tools - Model Context Protocol (2025-06-18)](https://modelcontextprotocol.io/specification/2025-06-18/server/tools) — Model Context Protocol - [Tools - Model Context Protocol (draft)](https://modelcontextprotocol.io/specification/draft/server/tools) — Model Context Protocol BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/people-and-bots-in-one-room/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 4 MIN READ # Why we are building BotBento around people, not just bots Today's agent workspaces put one person in charge of many bots. The bots talk to each other; the person supervises. BotBento starts from a different unit: a room that people and bots both belong to, with the same scoped permissions applied to everyone in it. By BotBento Editorial · Published 2026-09-20 · Updated 2026-09-20 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [The assumption almost every agent workspace makes](#the-single-operator-assumption) - [What the documentation actually says](#what-the-documentation-says) - [Why that boundary costs something](#why-the-boundary-matters) - [What BotBento is building instead](#what-botbento-is-building) - [Where this actually stands today](#where-it-actually-stands)[Read as Markdown ](/text/blog/people-and-bots-in-one-room/index.md) ## Key takeaways - Grok Bot's own documentation describes group chats as two to six Bots, and sharing a Bot as handing someone a separate copy rather than joining them to your conversation. - A workspace where only one human can be present makes the coordination between colleagues invisible to the tools doing the work. - BotBento is pre-release. Mixed human and bot rooms are the design it is being built against, not a feature you can rely on today. ## The assumption almost every agent workspace makes Look closely at the current generation of agent products and you find the same shape underneath: one human account, many agents, and a supervisory relationship between them. The person delegates, the agents work, the person approves. It is a genuine improvement on typing into a single chat box, and it is not the shape most work actually has. xAI's design write-up for Grok Bot is unusually candid about this, because it explains the reasoning rather than only the features. The team describes moving the product's main objects from disposable chat sessions to a persistent roster of Bots, giving each one an avatar, a memory and its own computer. It closes on the principle behind the whole thing: as agents take on more responsibility, the interface should ask less of the person. That is a coherent goal, and it produces a coherent product. It also quietly fixes the number of people in the picture at one. Sources: [Designing Grok Bot for a world of persistent agents](https://x.ai/news/designing-grok-bot). ## What the documentation actually says This is not an inference from marketing copy. Grok Bot's own documentation describes creating a group chat by selecting two to six Bots, and recommends a group when several Bots need one shared outcome with visible handoffs. The design post gives the same picture from the other side: a group is where a designer, engineer, PM and data scientist Bot share project context while keeping their separate memories. Sharing follows the same boundary. The documented way to give a colleague one of your Bots is to share it as a template, and the person who accepts it gets their own copy. The documentation is explicit that they do not receive your computer, your logins or your conversation history. That is a sensible security decision. It also means the thing you hand over is a clone, not a seat in the room you were working in. So the collaboration in these products is real, and it is collaboration between agents. The multiplayer part, the one where two people and three bots are looking at the same thread, is simply not what they were built to do yet. Sources: [Message and collaborate](https://docs.x.ai/grok-bot/chat-and-collaboration), [Frequently asked questions](https://docs.x.ai/grok-bot/faq), [Designing Grok Bot for a world of persistent agents](https://x.ai/news/designing-grok-bot). ## Why that boundary costs something Most work that is worth automating is already shared before any agent touches it. A launch has a designer and an engineer disagreeing about scope. A support escalation has one person who spoke to the customer and another who owns the fix. A newsroom has an editor who will not run a story until a second person has checked it. If the agent workspace can only hold one of those people, the handoffs between them happen somewhere else: a group chat in another app, a comment thread, a call. The bot sees the instruction it was given and none of the conversation that produced it. Every time context crosses that gap, a person has to carry it by hand, which is exactly the coordination work these products set out to remove. The alternative is not to give bots more autonomy. It is to stop treating the human side of the work as something that happens off-screen. ## What BotBento is building instead BotBento starts from the room rather than the roster. A conversation can hold people, bots, or both, and the same scoped permission model applies to everyone in it. A bot in a shared room works under that room's authority, not under a private transcript belonging to whoever invited it, and it cannot reach back into a personal history it was never granted. That single decision drives most of the rest. Membership has to be real, so there are invitations and directories rather than account-local fixtures. History has to be governed by the room, so shared threads carry their own access grants. Execution has to be explicit, so enrolling a runtime, running work in a shared room and controlling a machine are separate approvals rather than one switch. Routines belong to the server rather than to a phone that might be asleep. None of that is more elegant than the single-operator design. It is considerably more work. It is the bet that the room, not the assistant, is the thing worth getting right. ## Where this actually stands today BotBento is pre-release, and the project's internal rule is that a visible control or a passing test is not proof a feature works. So: mixed rooms, shared threads, reactions, drafts and routines exist and pass their own checks. Real-account, native-device and live-provider acceptance remain open. A private cloud environment with a maintained runtime is planned and not purchasable. There is no access gate on any of this. The blog is public, the newsletter is public, and the field notes here will keep describing what is being built, including the parts that are not finished. When something becomes genuinely usable, that will be said plainly and dated, not implied. If the argument in this article is wrong, the useful version of being wrong is specific: tell us which shared workflow you would actually move into a room like this, and which one you would not. ## Primary sources Sources checked 2026-09-20. Standards and product documentation can change; follow the linked version when implementing. - [Designing Grok Bot for a world of persistent agents](https://x.ai/news/designing-grok-bot) — xAI - [Message and collaborate](https://docs.x.ai/grok-bot/chat-and-collaboration) — xAI - [Frequently asked questions](https://docs.x.ai/grok-bot/faq) — xAI BotBento is in development. [Suggest a correction](/contact/). ## Keep reading - [AI agents vs workflows: choose who decides the next step](/blog/ai-agents-vs-workflows/) - [MCP tool permissions: what to check before connecting a bot](/blog/mcp-tool-permissions/) --- Canonical: https://botbento.com/blog/plugin-vs-authorized-connection/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 5 MIN READ # Plugin or authorized connection: what's actually different? A plugin or action definition tells a bot what an API can do. An authorized connection is the separate step where a provider issues that bot a scoped, revocable credential. Confusing the two leads people to think a tool is 'live' the moment it's installed, when in most setups it still needs a human to grant consent. By BotBento Editorial · Published 2026-09-20 · Updated 2026-09-20 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [Two different questions get collapsed into one](#plugin-vs-connection) - [What installing a tool actually configures](#what-installation-covers) - [What an authorized connection adds on top](#what-authorization-adds) - [A third setting: whether it runs on its own or waits for you](#bot-enablement-confirmation) - [A worked example (illustrative)](#worked-example) - [What to check before you call a connection "live"](#checklist)[Read as Markdown ](/text/blog/plugin-vs-authorized-connection/index.md) ## Key takeaways - Installing a tool (adding a schema or manifest) only declares what an API could do. Nothing runs against real data until an authorization step separately grants access. - The authorization step is where provider consent lives: a person signs into the provider, approves specific scopes, and the bot gets a token it can lose again if that consent is revoked. - Whether a call runs automatically or pauses for confirmation is a third, distinct setting — it depends on how the endpoint is marked, not on whether the connection is authorized. ## Two different questions get collapsed into one "Installing a plugin" answers one question: what could this bot theoretically do? "Authorizing a connection" answers a different one: has anyone actually granted it the right to do that against a specific account? In OpenAI's GPT Actions framework, an action is defined by two separate components: how the GPT authenticates with the API, and a schema that defines what the API can do. Those two components can move independently. A builder can finish the schema, publish the tool, and still have zero live access, because the authentication side hasn't been configured yet. That's the installation step. The authorization step is what turns the declared capability into something that can actually touch a real account, and it usually involves the account owner signing in somewhere. Sources: [Getting started with GPT Actions](https://platform.openai.com/docs/actions/getting-started). ## What installing a tool actually configures When you add an action to a GPT, the schema defines what your API can do — it tells ChatGPT which operations exist and what parameters they take. By default, the authentication method for all actions is set to "None", so a freshly installed action can be reachable without any sign-in step at all, until someone deliberately changes that. That default matters because it means installation and authorization aren't just conceptually different, they're set independently, and the wrong default can leave a tool answering requests before anyone intended it to hold real access. Builders can mix a single authentication type along with endpoints that don't require authentication at all inside the same tool, so "installed" doesn't tell you much about what's actually reachable without a credential. There's also a structural boundary worth knowing: a GPT can use either apps or actions, but not both at the same time. That's a platform-level constraint on installation itself, separate from anything about consent or credentials. Sources: [GPT Action authentication | OpenAI API](https://platform.openai.com/docs/actions/authentication), [Getting started with GPT Actions](https://platform.openai.com/docs/actions/getting-started), [Production notes on GPT Actions](https://platform.openai.com/docs/actions/production). ## What an authorized connection adds on top Authorization is the step where a person, not a developer, grants access. Actions allow OAuth sign in for each user, and that flow is described as the best way to provide personalized experiences and make the most powerful actions available. Setting it up means entering a client ID, client secret, and authorization details, and once configured, the platform hands back a callback URL that has to be registered with the third-party provider before sign-in will work. The Model Context Protocol's authorization specification draws the same line more formally: an MCP client acts as an OAuth 2.1 client, making protected resource requests on behalf of a resource owner, while the authorization server is responsible for interacting with the user and issuing access tokens for use at the server. The tool definition and the credential are handled by different parties on purpose. Transport matters here too. Implementations using an STDIO transport — the kind of local, same-machine tool call common in personal bot setups — are told they should not follow the OAuth-style flow at all, and should instead retrieve credentials from the environment. So a locally installed tool and a remotely authorized one aren't just different in permission scope; they're often authorized through entirely different mechanisms, one via a login screen, one via whatever's sitting in a local environment variable or keychain. Sources: [GPT Action authentication | OpenAI API](https://platform.openai.com/docs/actions/authentication), [Authorization - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/basic/authorization). ## A third setting: whether it runs on its own or waits for you Even after a connection is authorized, there's a separate switch controlling whether a call executes automatically or pauses for a human. In the OpenAPI schema behind an action, an endpoint can be marked with an x-openai-isConsequential flag. A good example of a consequential action is booking a hotel room and paying for it on behalf of a user. If that field is set to true, the operation is treated as one that must always prompt the user for confirmation before running, and the interface won't offer an "always allow" shortcut. That flag is independent of both installation and authorization. A read-only endpoint on the same authorized connection can run without asking, while a write endpoint sitting right next to it pauses every time. This is the practical layer of "bot enablement" — not whether the bot has access, but whether it's allowed to use that access without checking in first. Sources: [Production notes on GPT Actions](https://platform.openai.com/docs/actions/production). ## A worked example (illustrative) Say a small team builds a scheduling bot they call Aria. First, someone installs a calendar tool: they write a schema declaring a GET operation to list events and a POST operation to create them. At this point Aria has a plugin installed, but it can't touch anyone's calendar — there's no credential attached yet, and by default the action's authentication is set to none. Next, the team switches the create-event endpoint to OAuth. They register Aria's app with the calendar provider, get a client ID and secret, and paste the provider's authorization and token URLs into the tool's settings. The provider hands back a callback URL, which the team registers on their side to complete the loop. The first time a team member actually asks Aria to book something, they're redirected to the calendar provider's own sign-in page, not anything BotBento or the bot builder controls. They approve a scope like "read and write calendar events." That approval — done on the provider's site, by the account owner — is the authorized connection. It didn't exist when the tool was installed; it exists now because a specific person consented to a specific scope for a specific account. Finally, the team marks the POST /events endpoint as consequential. Now Aria can read the calendar freely, but every time it tries to create an event, it stops and asks first. Reading, writing, and confirming are three separate settings stacked on the same tool — installed, authorized, and enabled are not the same fact. ## What to check before you call a connection "live" Before assuming a connected tool is safe to leave running, it's worth checking a short list. First, what's the actual authentication setting on each endpoint — not what you intended, but what's configured, since the default is none unless someone changed it. Second, what scope did the account owner actually approve during sign-in, since OAuth grants are scoped, not all-or-nothing. Third, which endpoints are marked consequential and which aren't, since that's what decides whether the bot pauses before acting. For tools that run locally rather than through a hosted OAuth flow, the check is different: credentials come from the environment rather than a login screen, so revoking access means removing or rotating whatever's stored locally, not clicking "disconnect" on a provider's consent page. Knowing which model your tool uses — hosted OAuth or local environment credentials — is part of what BotBento's permission views are being built to make visible in one place, rather than something you have to reconstruct from separate settings screens. Sources: [Authorization - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/basic/authorization). ## Primary sources Sources checked 2026-09-20. Standards and product documentation can change; follow the linked version when implementing. - [GPT Action authentication | OpenAI API](https://platform.openai.com/docs/actions/authentication) — OpenAI - [Getting started with GPT Actions](https://platform.openai.com/docs/actions/getting-started) — OpenAI - [Production notes on GPT Actions](https://platform.openai.com/docs/actions/production) — OpenAI - [Authorization - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/basic/authorization) — Model Context Protocol BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/reviewing-ai-generated-research-notes/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 3 MIN READ # How Do You Review an AI-Generated Research Note? AI-drafted research notes can contain citations that look real but aren't, sources that have gone stale, and claims that quietly contradict each other. Reviewing one means checking source ownership, checking dates separately from content, hunting for contradictions, and explicitly marking what stays unverified. By BotBento Editorial · Published 2026-09-23 · Updated 2026-09-23 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [Why an AI-generated research note needs a review pass](#why-review-matters) - [Check who actually owns each claim](#check-source-ownership) - [Check dates before you check content](#check-dates) - [Look for claims that contradict each other or the sources they cite](#spot-contradictions) - [Mark what you couldn't verify — don't erase it](#mark-the-unknowns)[Read as Markdown ](/text/blog/reviewing-ai-generated-research-notes/index.md) ## Key takeaways - Verify each citation against a library catalog, Google Scholar, or the journal's own table of contents — don't trust that a reference looks correctly formatted. - Check when each source was published and whether a changeable claim has been verified against current material. - Keep a visible 'unverified' or 'fabricated' status on each claim instead of deleting uncertain material — the gap itself is useful information. ## Why an AI-generated research note needs a review pass An AI-written research note can read fluently even when a citation is wrong. A citation beside a sentence does not by itself confirm that the reference exists or supports the sentence. This matters because fabricated citations rarely look obviously wrong. They often mix real elements with invented ones, so a real author name can be attached to a title that was never published, or a real journal can be credited with a study it never ran. Sources: [AI Hallucinated Citations - AI Hallucinated Citations - Research Guides at University of North Carolina at Charlotte](https://guides.library.charlotte.edu/hallucinatedcitations). ## Check who actually owns each claim Before trusting a citation in an AI-generated note, trace it back to a real, findable document. One workable sequence: search the exact title in a library catalog first, and if it doesn't turn up, search the title in Google Scholar or plain Google to try to reach the full text or the publisher's page. Two more checks catch fakes that slip past a first pass. First, some fabricated citations have already been indexed by Google Scholar, so appearing there is not proof a source is real. Second, go directly to the journal's site and open the table of contents for the specific volume and issue cited. If the article isn't listed, flag the citation for further checking against the publisher or a DOI record. It also helps to check whether the named author actually lists that publication on their own site. Sources: [AI Hallucinated Citations - AI Hallucinated Citations - Research Guides at University of North Carolina at Charlotte](https://guides.library.charlotte.edu/hallucinatedcitations). ## Check dates before you check content Treat dates as a separate check from the claim itself. A note can cite a real, existing source and still misstate when it was published, or lean on a source that predates a more recent finding it doesn't mention. NIST describes its generative AI profile as a companion to the broader AI Risk Management Framework that helps organizations identify risks posed by generative AI and consider responses. Checking dates on an AI research note is a practical editorial step within that broader risk-management goal; this particular checklist is BotBento's, not a NIST requirement. For a working note, check two things independently: the publication or last-updated date on the source itself, and whether the underlying claim could have changed since then. Without checking a current primary source, a note may miss a superseded standard, revised price, or later ruling. Sources: [AI Risk Management Framework | NIST](https://www.nist.gov/itl/ai-risk-management-framework), [Technical Reports](https://airc.nist.gov/technical-reports). ## Look for claims that contradict each other or the sources they cite Read the note a second time looking only for internal consistency, not accuracy. AI-generated text can state a number in one paragraph and a slightly different number for the same figure two paragraphs later, because each sentence was generated locally without cross-checking the rest of the document. The same blending problem that produces fake citations also produces contradictions: a citation may combine details from two real sources, so the note ends up attributing one paper's finding to another paper's authors. When a cited source is found, open it and confirm the claim next to it is actually the claim the source makes — not just that the source is real. Sources: [AI Hallucinated Citations - AI Hallucinated Citations - Research Guides at University of North Carolina at Charlotte](https://guides.library.charlotte.edu/hallucinatedcitations). ## Mark what you couldn't verify — don't erase it Illustrative example: a fictional AI-drafted note on battery-recycling regulation contains five claims. On review, claim one (a 2024 EU directive number) checks out against the official text. Claim two (a specific recovery-rate percentage attributed to a named report) can't be found in that report's table of contents, and gets tagged 'unverified — check the report itself.' Claim three (a company's stated recycling capacity) is real but from a source two years older than the note implies, and gets tagged 'outdated — needs a newer source.' Claim four (an industry-wide cost trend) has no findable source at all and is tagged 'unverified — treat as a hypothesis.' Claim five is a direct quote that, when checked against the original page, differs by two words — tagged 'misquoted.' The point of this pass isn't to delete anything that fails. It's to attach a status to every claim so the next reader — human or bot — knows which parts of the note are load-bearing and which parts are still open questions. A note with five tagged statuses is more useful than a note with five confident-sounding sentences and no way to tell which of them is solid. This is the kind of review record we want BotBento routines to support: a note with visible verification status alongside its claims, so the next reader can see what has and has not been checked. BotBento is still in development. ## Primary sources Sources checked 2026-09-23. Standards and product documentation can change; follow the linked version when implementing. - [AI Hallucinated Citations - AI Hallucinated Citations - Research Guides at University of North Carolina at Charlotte](https://guides.library.charlotte.edu/hallucinatedcitations) — UNC Charlotte Library - [AI Risk Management Framework | NIST](https://www.nist.gov/itl/ai-risk-management-framework) — NIST - [Technical Reports](https://airc.nist.gov/technical-reports) — NIST AI Resource Center BotBento is in development. [Suggest a correction](/contact/). ## Keep reading - [What should an AI agent record after each run?](/blog/ai-agent-run-result-record/) - [What belongs in an AI agent stopping rule?](/blog/ai-agent-stopping-rules/) --- Canonical: https://botbento.com/blog/revoke-bot-tool-access/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 4 MIN READ # How Do You Actually Revoke a Bot's Access to a Tool It No Longer Uses? Revocation invalidates the submitted token at the authorization server, but related-token policy and resource-server awareness affect when access stops in practice. A fictional calendar bot makes the distinction concrete. By BotBento Editorial · Published 2026-09-27 · Updated 2026-09-27 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [What an OAuth revoke request actually does](#the-gap-between-clicking-revoke-and-access-stopping) - [Token format and validation both affect the delay](#why-token-type-decides-what-happens-next) - [Worked example: a fictional calendar bot](#worked-example-calendar-bot) - [What to check for an authorized MCP connection](#mcp-specific-considerations) - [A short checklist before trusting a revoke button](#checklist-before-you-trust-a-revoke-button)[Read as Markdown ](/text/blog/revoke-bot-tool-access/index.md) ## Key takeaways - RFC 7009 invalidates the submitted token, while warning that propagation to resource servers can take time. - A self-contained token can remain usable at a server that checks only signature and expiry; an opaque token can also appear active while an introspection result is cached. - Check the provider’s related-token policy and verify rejection of the old credential before calling revocation complete. ## What an OAuth revoke request actually does A bot may hold an OAuth access token for a tool and a refresh token for obtaining later access tokens. RFC 7009 defines an HTTPS revocation endpoint where a client can submit either kind of token. It requires support for revoking refresh tokens and recommends support for access tokens, so the provider’s actual behavior matters. For a valid request, the authorization server invalidates the submitted token immediately. The RFC also warns of a practical propagation delay while different servers learn about that invalidation. A successful response does not, by itself, prove that every resource server has stopped accepting the token at that instant. The client must stop using the token after a successful response. Related tokens have their own policy. If a refresh token is revoked and the server supports access-token revocation, the RFC says the server should also invalidate access tokens from the same grant. If an access token is submitted, the server may revoke its related refresh token. Neither direction is an unconditional guarantee about every credential tied to the connection. Sources: [RFC 7009: OAuth 2.0 Token Revocation](https://www.rfc-editor.org/info/rfc7009/). ## Token format and validation both affect the delay RFC 7009 describes self-contained access tokens that a resource server can check without contacting the issuer, and reference tokens that require a lookup. A self-contained JWT is a common example of the first design, but its format alone does not decide whether a server can enforce early revocation. The validation method matters. If a server checks only a token’s signature and expiry, it may continue accepting an already-issued token until expiry unless it also checks a revocation list or receives another invalidation signal. This is a conditional risk, not a rule that every JWT remains valid until expiry. Ask how the specific provider validates tokens and whether it can stop an active token early. Reference tokens can provide fresher status when the resource server queries the issuer, but a lookup is not automatically live on every request. RFC 7662 permits caching introspection responses and notes that a cached active result may leave a window in which a revoked token is still accepted. The cache lifetime and provider implementation matter as much as the token label. Sources: [RFC 7009: OAuth 2.0 Token Revocation](https://www.rfc-editor.org/info/rfc7009/), [RFC 7662: OAuth 2.0 Token Introspection](https://www.rfc-editor.org/info/rfc7662/). ## Worked example: a fictional calendar bot Illustrative example: a personal bot has calendar-write permission, a refresh token and an access token that expires in 60 minutes. The owner removes the connection in the calendar provider console. First, the builder checks whether that action revokes the grant, the refresh token, the current access token or some combination. A button label alone does not specify the provider’s policy. Suppose the provider confirms the refresh token is revoked. The bot should no longer be able to exchange that token for a new access token. Whether the current access token also becomes unusable depends on the provider’s related-token policy and the calendar server’s validation. A locally validated token with no early-revocation check might still be accepted for the remainder of its lifetime. A server with an effective invalidation check could reject it sooner. The builder removes the bot’s stored credentials and disables new calendar calls in its own system. For a controlled verification, it asks the provider or an authorized test client to attempt a harmless read with the old access token and records the response and time. Rejection supports the claim that this route has stopped working; a successful read means access is not yet fully cut off. Do not change a real event merely to prove the revoke button worked. Sources: [RFC 7009: OAuth 2.0 Token Revocation](https://www.rfc-editor.org/info/rfc7009/). ## What to check for an authorized MCP connection For an HTTP MCP connection using authorization, the MCP specification treats the MCP server as an OAuth resource server. It requires the server to validate each request’s access token, including that the token was issued for that server as its intended audience. Invalid or expired tokens must receive HTTP 401. The specification does not prescribe one token format or one revocation propagation method for every MCP server. An MCP server may also call a downstream API with a separate token. The MCP specification prohibits passing the incoming MCP client token straight through to that API. Revoking the MCP-side grant and revoking a downstream grant can therefore be distinct operations. A builder should list both connections, identify which party issued each credential and revoke the credential that grants the unwanted action. For a BotBento-style tool connection, the practical question is which credential can still reach which resource after revocation. Check the MCP server’s token validation and the downstream provider’s status independently. BotBento is pre-release; this is a design checklist, not a claim that it currently provides a completed revocation audit feature. Sources: [Authorization - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-11-25/basic/authorization). ## A short checklist before trusting a revoke button Revocation is a useful control, but an operator should distinguish the submitted token, related tokens and observed resource access. RFC 7009 says the submitted token is invalidated; it also permits differences in related-token policy and notes propagation delay. Verify the provider’s behavior before treating a console action as an emergency stop for a bot. - Record the issuer, audience, scopes and expiry for each access token; identify whether the server validates locally or checks live status. - Ask the provider whether revoking a refresh token also invalidates related access tokens, and whether revoking an access token affects its refresh token. - Check for cached introspection results or other propagation delays. Short token lifetime limits one possible window but is not a complete revocation mechanism. - Remove stored credentials and stop the bot’s own tool calls, then verify that an authorized old-token request is rejected by the resource server. - For an MCP server that calls another API, inspect both the MCP authorization and any separate downstream credential before declaring access removed. Sources: [RFC 7009: OAuth 2.0 Token Revocation](https://www.rfc-editor.org/info/rfc7009/), [RFC 7662: OAuth 2.0 Token Introspection](https://www.rfc-editor.org/info/rfc7662/), [Authorization - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-11-25/basic/authorization). ## Primary sources Sources checked 2026-09-27. Standards and product documentation can change; follow the linked version when implementing. - [RFC 7009: OAuth 2.0 Token Revocation](https://www.rfc-editor.org/info/rfc7009/) — IETF RFC Editor - [RFC 7662: OAuth 2.0 Token Introspection](https://www.rfc-editor.org/info/rfc7662/) — IETF RFC Editor - [Authorization - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-11-25/basic/authorization) — Model Context Protocol BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/supervisor-agent-vs-single-agent-tools/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 3 MIN READ # When Should a Bot Use a Supervisor Agent Instead of One Agent With Tools? Multi-agent orchestration adds routing and context decisions. Start with one agent and its tools; add specialists when different instructions, tools or policies materially improve the work, then choose whether a manager keeps control or hands off the turn. By BotBento Editorial · Published 2026-09-26 · Updated 2026-09-26 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [The default is one agent with tools, not a supervisor](#the-default-is-one-agent) - [If you do split, there are two distinct patterns](#two-ways-to-split) - [A worked example: a support bot that also books refunds](#worked-example) - [Splitting adds failure modes you now own](#new-failure-modes)[Read as Markdown ](/text/blog/supervisor-agent-vs-single-agent-tools/index.md) ## Key takeaways - Start with one agent where possible; add specialists when they improve capability or policy isolation, prompt clarity, or the ability to inspect a run. - Two common orchestration patterns are agents-as-tools, where a manager keeps the user-facing turn, and handoffs, where a specialist becomes active for the rest of that turn. - Splitting introduces new failure modes you must design for yourself: how much context to pass at a handoff, what happens when a sub-agent fails, and how to test the added routing paths. ## The default is one agent with tools, not a supervisor Bot builders often reach for a supervisor-and-workers design before they need one. The official guidance from OpenAI's Agents SDK documentation is direct on this point: start with one agent whenever you can, and only add orchestration once a single agent with a good prompt and the right tools stops being sufficient. The reason is that a multi-agent design adds coordination decisions: how work is split, who owns the final answer, what context moves between agents, and what happens when a specialist fails. A single agent still needs clear tool and error handling, but has fewer routing paths. If one set of instructions and tools works well in evaluation, keep it until a specific split improves the result. Sources: [Orchestration and handoffs | OpenAI API](https://developers.openai.com/api/docs/guides/agents/orchestration). ## If you do split, there are two distinct patterns When a split has a clear benefit, the OpenAI Agents SDK documentation describes two patterns that come up most often. One is agents-as-tools, where a specialist helps with a bounded subtask while the manager keeps the user-facing turn. The other is handoffs, where the chosen specialist becomes the active agent for the rest of the current turn. These answer different ownership questions. In the agents-as-tools pattern, a central manager invokes specialists and combines their results into the final answer. In the handoff pattern, the first agent transfers control and the specialist responds directly. That does not mean the specialist owns every future conversation; the application still decides how later turns resume and route. Anthropic's own engineering guidance uses similar language for a comparable pattern it calls orchestrator-workers, where a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results. Anthropic recommends this specifically for complex tasks where you can't predict the subtasks in advance -- for example, coding tasks where the number of files that need changing depends on the input. That's a different signal than the routing signal behind handoffs: orchestrator-workers is about not knowing the shape of the task ahead of time, while handoffs are about which specialist should own the conversation once the category of the request is known. Sources: [Orchestration and handoffs | OpenAI API](https://developers.openai.com/api/docs/guides/agents/orchestration), [Agent orchestration - OpenAI Agents SDK](https://openai.github.io/openai-agents-python/multi_agent/), [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents). ## A worked example: a support bot that also books refunds Illustrative example. Suppose a small team is building a support bot. Most requests are simple questions answered from a knowledge base -- one agent, one retrieval tool, no split needed. Now two request types show up: billing disputes that need a policy lookup, and refund requests that use a different tool and require an approval step. The team can keep one agent with both tool sets if evaluation shows it routes correctly and respects approval policy. If the instructions or permissions interfere, it can use a triage agent that hands off to a billing or refund specialist. The specialist split should be justified by observed routing or policy outcomes, not the mere presence of two categories. If the team picks handoffs, it must decide how much conversation history the refund specialist receives. Pass irrelevant history and the specialist may lose focus; pass too little and the customer may need to repeat an order number. The SDK provides a mechanism for filtering handoff history, but the team must choose and test a policy for its own workflow. Sources: [Orchestration and handoffs | OpenAI API](https://developers.openai.com/api/docs/guides/agents/orchestration), [Agent orchestration - OpenAI Agents SDK](https://openai.github.io/openai-agents-python/multi_agent/). ## Splitting adds failure modes you now own Multi-agent orchestration adds work. First, context passing: how much history moves at a handoff, and whether filtering is appropriate, requires a deliberate decision. Second, error handling: if a specialist fails mid-task, the application must decide whether to show an error, retry or return control to triage. Third, testing: each supported route and context boundary needs a representative evaluation. More agents can mean more combinations to check, but the number depends on the actual routing graph; there is no universal exponential rule. None of these problems disappear by choosing agents-as-tools instead of handoffs -- they just shift shape. In the agents-as-tools pattern, the manager retains the final say, so a failing sub-agent is easier to catch before it reaches the user, but the manager now needs its own logic for deciding when a sub-agent's result is good enough to use. For a personal or small-team bot, begin with the smallest set of specialists that solves an observed problem. Review traces and test the routes that exist before adding another role. Include ordinary questions, policy-bound refund cases and failed tool responses in that evaluation. OpenAI recommends adding specialists when they materially improve isolation, clarity or trace legibility, rather than splitting for its own sake. Sources: [Orchestration and handoffs | OpenAI API](https://developers.openai.com/api/docs/guides/agents/orchestration), [Agent orchestration - OpenAI Agents SDK](https://openai.github.io/openai-agents-python/multi_agent/). ## Primary sources Sources checked 2026-09-26. Standards and product documentation can change; follow the linked version when implementing. - [Orchestration and handoffs | OpenAI API](https://developers.openai.com/api/docs/guides/agents/orchestration) — OpenAI - [Agent orchestration - OpenAI Agents SDK](https://openai.github.io/openai-agents-python/multi_agent/) — OpenAI - [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) — Anthropic BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/test-calendar-ai-agent/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 5 MIN READ # How do you test an AI agent that uses a calendar? Test a calendar AI agent against a written event specification in a disposable calendar. Verify the saved event independently, then test an ambiguous request, a denied write and revoked access. Record both the provider result and what the agent tells the user; a confident reply is not proof that the intended event exists. By BotBento Editorial · Published 2026-09-10 · Updated 2026-09-10 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [Write the expected event before asking the agent](#write-the-expectation) - [Check the calendar and time in the saved result](#check-input-and-output) - [Test an unclear instruction and a denied write separately](#ambiguous-and-denied) - [Make an uncertain write distinguishable from a fresh request](#uncertain-write) - [Revoke the test connection and inspect the next attempt](#revoked-access) - [Keep a small acceptance record and clean up the fixture](#keep-a-test-receipt)[Read as Markdown ](/text/blog/test-calendar-ai-agent/index.md) ## Key takeaways - Use a separate test calendar and an event with no guests before testing invitations or real appointments. - Check the exact calendar, event identity, time zone and duration against an expectation written before the run. - Test permission failures and revoked access separately, and require the agent to explain what remains unverified. ## Write the expected event before asking the agent A calendar assistant can produce a convincing confirmation while choosing the wrong calendar or misunderstanding the time. Start the evaluation with a small written specification that a second person could check without reading the conversation. This guide proposes a test plan using Google Calendar documentation as a concrete reference. We have not executed these scenarios as a BotBento integration test, and BotBento remains in development. Our fictional task is to create one event called TEST: sketch review, on 14 September 2026, from 10:00 to 10:30 in Asia/Amman, on a private calendar called Agent Sandbox. It has no attendees, recurrence, attachments or conferencing. The calendar name is a human label; record its actual identifier when setting up your own fixture. Choose a disposable test account and calendar you control. Keep real appointments and other people's addresses out of the initial test. Record the initial state so you can distinguish a new event from an old fixture left by a previous run. Use a different run marker when starting a genuinely new test, and retain the previous marker when investigating an uncertain result. ## Check the calendar and time in the saved result Google's creation guide distinguishes the destination calendar identifier from the event body. The special value primary refers to the signed-in user's primary calendar. It also distinguishes timed events, which use dateTime fields, from all-day events, which use date fields. These details make useful assertions for a calendar-agent test. For our fixture, require the agent's proposed destination to resolve to Agent Sandbox. Inspect the exact date, start, end and time zone before the write. Afterward, use Google's Events: get operation with the returned event ID and intended calendar ID to retrieve the saved event. Compare the result with your original specification, and inspect the event in the calendar interface as a separate human readback. Grade each assertion independently: correct destination, correct date, thirty-minute duration, correct zone, no guests and one event. A saved event with the wrong time is a failed task even if the API request succeeded. Conversely, if the event exists but the agent reports failure, record that mismatch too; it affects whether a person will repeat the request. Sources: [Create events](https://developers.google.com/workspace/calendar/api/guides/create-events), [Events: get](https://developers.google.com/workspace/calendar/api/v3/reference/events/get). ## Test an unclear instruction and a denied write separately Now change one input at a time. For an ambiguity test, omit the time zone and ensure the test conversation supplies no agreed default. Our proposed acceptance rule is that the agent asks for the missing zone before saving. That rule is an evaluation choice for this fixture, not a Google API requirement. A different application may have an explicit user-approved default; test that default visibly instead of relying on an assumption. For the permission test, use a separate connection whose granted scopes you have verified are read-only. Google's Events: get accepts calendar.events.readonly, while Events: insert lists write-capable scopes and does not include that read-only scope. Google's consent guidance recommends requesting the minimum access required for the application. Ask the read-only agent to create the same disposable event. The desired application behavior is to explain that it cannot save the event with its current access, retain any useful draft and leave the calendar unchanged. If the application blocks the write before contacting Google, record an application-level denial. Test the provider's denial separately in the integration harness. Neither result should be labelled a successful calendar write. Sources: [Events: get](https://developers.google.com/workspace/calendar/api/v3/reference/events/get), [Events: insert](https://developers.google.com/workspace/calendar/api/v3/reference/events/insert), [Configure the OAuth consent screen and choose scopes](https://developers.google.com/workspace/guides/configure-oauth-consent). ## Make an uncertain write distinguishable from a fresh request In a controlled integration harness, model a write whose response is lost after the provider accepts it. Do not interrupt a real appointment to manufacture this condition. Google's creation guide describes supplying a compliant event ID to help prevent duplicate creation when a failure happens after Calendar has executed the request. Use the documented ID format; an arbitrary test label is not necessarily a valid event ID. Our proposed check is that the integration preserves the action's identity, inspects the intended event and compares its fields before deciding what to retry. An existing ID with different contents is a discrepancy to investigate. Record the evidence that resolves the uncertainty instead of turning every timeout into a new insertion. Keep the first fixture guest-free. Google's insert reference documents notification options and cautions about side effects of suppressing updates. Testing invitation delivery deserves its own controlled recipient and delivery evidence; an event appearing in the organizer's calendar does not establish that an invite reached someone else's inbox. Sources: [Create events](https://developers.google.com/workspace/calendar/api/guides/create-events), [Events: insert](https://developers.google.com/workspace/calendar/api/v3/reference/events/insert). ## Revoke the test connection and inspect the next attempt Run revocation checks only with an isolated test user and OAuth project. Google's OAuth documentation says revocation removes the user's granted scopes for the project and invalidates its issued access and refresh tokens across that project's clients. It also notes that revocation can take time to have full effect. Revoking a production connection merely to test a calendar feature can therefore affect more than that feature. After revoking the test grant through the supported account or application flow, record when it happened and inspect a subsequent authenticated request. Do not declare the test passed from the revocation response alone. Distinguish an actual provider read from content the application already cached; displaying an old event is not evidence of continuing access. The behavior we want to evaluate is understandable recovery: the agent identifies the unavailable connection, does not claim a new write succeeded and requires reconnection before resuming dependent work. If access still works during propagation, keep the check unresolved and observe it again within a bounded test period. Reconnect deliberately for any remaining cleanup; do not silently restore the revoked grant. Sources: [Using OAuth 2.0 for Web Server Applications](https://developers.google.com/identity/protocols/oauth2/web-server#tokenrevoke). ## Keep a small acceptance record and clean up the fixture For each scenario, save the intended event, the connection's tested access level, the observed calendar result and the agent's final statement. Keep tokens and unrelated event contents out of the report. Name any untested behavior, such as recurring events, daylight-saving transitions, attendee responses or another provider. Passing this narrow plan does not prove those separate cases. Finally, remove only the disposable records belonging to the test, using their retained identifiers and an authorized cleanup connection. Verify cleanup and preserve the short result record. The useful outcome is a set of specific, reproducible observations that lets you decide whether this calendar task is ready for its intended use. ## Primary sources Sources checked 2026-09-10. Standards and product documentation can change; follow the linked version when implementing. - [Create events](https://developers.google.com/workspace/calendar/api/guides/create-events) — Google for Developers - [Events: get](https://developers.google.com/workspace/calendar/api/v3/reference/events/get) — Google for Developers - [Events: insert](https://developers.google.com/workspace/calendar/api/v3/reference/events/insert) — Google for Developers - [Configure the OAuth consent screen and choose scopes](https://developers.google.com/workspace/guides/configure-oauth-consent) — Google for Developers - [Using OAuth 2.0 for Web Server Applications](https://developers.google.com/identity/protocols/oauth2/web-server#tokenrevoke) — Google for Developers BotBento is in development. [Suggest a correction](/contact/). ## Keep reading - [MCP tool permissions: what to check before connecting a bot](/blog/mcp-tool-permissions/) - [What should an AI agent record after each run?](/blog/ai-agent-run-result-record/) --- Canonical: https://botbento.com/blog/tool-call-timeout-duration/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 4 MIN READ # How Long Should You Let a Bot's Tool Call Run Before Timing Out? MCP recommends request timeouts but does not set a universal duration. Clients and other network layers can enforce different limits, so measure the tool, configure the caller you actually use, and use an asynchronous job pattern for work that cannot reliably finish within that budget. By BotBento Editorial · Published 2026-09-29 · Updated 2026-09-29 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [The short answer](#the-decision) - [Which timeout clock actually ends the call?](#what-the-defaults-actually-are) - [Worked example: a document-processing tool (illustrative)](#worked-example) - [Progress notifications don't reliably save a slow call](#progress-notifications-are-not-a-fix) - [When to stop raising the timeout and switch to polling](#when-to-switch-to-async) - [A checklist before you ship a tool call](#checklist)[Read as Markdown ](/text/blog/tool-call-timeout-duration/index.md) ## Key takeaways - The MCP specification recommends request timeouts and per-request configuration, but sets no universal duration. - Check the version and configuration of the client that makes the request; do not treat a proposed SDK change as a shipped default. - For unpredictable or long-running work, return a job identifier and let the caller check status instead of holding one request open indefinitely. ## The short answer Choose the timeout from the behavior of this tool and the caller that invokes it. Measure normal and slow runs with representative inputs, leave room for expected network variation, and set a per-request limit when the client supports one. A simple lookup and a document-processing job should not inherit one unexplained global number. For work whose slow tail exceeds the caller’s acceptable wait, return a job ID and expose a status check instead of stretching a synchronous call indefinitely. The Model Context Protocol specification says implementations SHOULD establish timeouts for sent requests and SDKs SHOULD allow per-request configuration. It does not prescribe a duration. The timeout that matters to your bot is the shortest applicable limit in its actual client, transport, gateway, or upstream API. Sources: [Lifecycle - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-03-26/basic/lifecycle). ## Which timeout clock actually ends the call? A request can cross several independent clocks. The MCP TypeScript SDK has been documented with a 60-second request default, while an open issue in the Python SDK repository reports that its client-side requests lacked an equivalent default. Those are version-specific implementation observations, not MCP requirements: check the SDK release and request options that your bot actually runs. Anthropic’s Claude API documentation addresses a different clock: its SDKs validate that non-streaming Messages API requests are not expected to exceed ten minutes and recommend streaming or Message Batches for long work. That guidance does not set an MCP tool timeout. A client, gateway, or external API can finish waiting before the tool server does, so record which layer timed out instead of assuming one number governs the whole path. For example, a client configured to wait 60 seconds cannot receive a tool result that arrives after 90 seconds merely because the server would have permitted a longer execution. The server may still have started side effects; a retry should therefore use an idempotency key or check the prior job state where the operation supports it. Sources: [Lifecycle - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-03-26/basic/lifecycle), [No default timeout for requests (unlike TS SDK) · Issue #1374](https://github.com/modelcontextprotocol/python-sdk/issues/1374), [Claude API errors - Claude Platform Docs](https://docs.anthropic.com/en/api/errors). ## Worked example: a document-processing tool (illustrative) Say a bot exposes an MCP tool that fetches a file, extracts its contents, normalizes the output, and writes a summary. In this illustrative case, the full job often takes longer than the caller’s configured request timeout. The caller can stop waiting even though the server may have started the work; the precise error and cancellation behavior depend on the client and transport. Neither a completed file nor a failed operation can be inferred from the timeout alone. Measure actual durations on representative files and record whether the job produced a result or side effect after the caller stopped waiting. If the slow tail is rare and the caller permits per-request configuration, a longer bounded request might be reasonable. If it is common or unpredictable, return a job ID promptly, persist progress and results, and let the bot ask for status before deciding whether to retry. ## Progress notifications don't reliably save a slow call MCP defines a way for a server to signal it's still working: implementations MAY choose to reset the timeout clock when receiving a progress notification corresponding to the request, as this implies that work is actually happening, but implementations SHOULD always enforce a maximum timeout, regardless of progress notifications, to limit the impact of a misbehaving client or server. That second half of the sentence is the part builders miss. Even a tool that sends a progress ping every five seconds can still be killed by a hard ceiling, because the spec only permits resetting on progress; it doesn't require it, and it always allows a hard cap regardless. Whether your specific client actually resets the clock on progress notifications is an implementation detail you have to check, not assume. This is exactly the kind of gap that generates bug reports: an open proposal filed as SEP-1539 (Timeout Coordination) argues that progress notifications only help after the operation starts and don't solve the information asymmetry about how long an operation should take initially, since a client still needs to guess an appropriate initial timeout before any progress is reported. Treat progress notifications as a UX signal, not a timeout strategy, until you've confirmed your client honors them for that purpose. Sources: [Lifecycle - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-03-26/basic/lifecycle), [SEP-1539: Timeout Coordination · Issue #1539](https://github.com/modelcontextprotocol/modelcontextprotocol/issues/1539). ## When to stop raising the timeout and switch to polling When normal duration is unpredictable, regularly exceeds the caller’s acceptable wait, or depends on an external system whose latency you do not control, a different call shape may be safer. Have the tool return a job ID, persist the job state, and expose a separate status operation the bot can poll. That gives the caller a way to reconcile a timeout with work the server may already have started. Add an idempotency key or equivalent deduplication when retrying could repeat a side effect. Anthropic gives analogous guidance for long model requests: use the streaming Messages API or Message Batches instead of a long idle non-streaming connection, because networks can drop idle requests and batches can be polled. This is guidance for Anthropic API calls, not a claim that MCP defines a job API. A bot wrapping slow external work can design its own job-ID-and-status tool and document the failure and retry behavior. Sources: [Claude API errors - Claude Platform Docs](https://docs.anthropic.com/en/api/errors). ## A checklist before you ship a tool call Run through this before you lock in a number for a new tool, and revisit it whenever the tool's underlying dependency changes. - Measure the tool's real duration distribution (not just the happy path) before picking a timeout number. - Set the timeout per call or per tool, not as one global value for every tool your bot uses. - Confirm whether your specific MCP client actually resets its clock on progress notifications; don't assume it does. - For anything with an unpredictable or multi-minute tail, use a job-ID-and-poll pattern instead of a longer synchronous timeout. - Log the tool name and the configured duration whenever a timeout fires, so a slow tail shows up as a pattern instead of a one-off mystery. ## Primary sources Sources checked 2026-09-29. Standards and product documentation can change; follow the linked version when implementing. - [Lifecycle - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-03-26/basic/lifecycle) — Model Context Protocol - [No default timeout for requests (unlike TS SDK) · Issue #1374](https://github.com/modelcontextprotocol/python-sdk/issues/1374) — Model Context Protocol (GitHub) - [Claude API errors - Claude Platform Docs](https://docs.anthropic.com/en/api/errors) — Anthropic - [SEP-1539: Timeout Coordination · Issue #1539](https://github.com/modelcontextprotocol/modelcontextprotocol/issues/1539) — Model Context Protocol (GitHub) BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/blog/verify-ai-agent-task-completion/ Format: Markdown representation of the public HTML page. [Home](/) / [Blog](/blog/) FIELD NOTES / 3 MIN READ # How Do You Verify an AI Agent Actually Finished the Task? An agent's own 'task complete' message is generated text, not a fact. This piece explains why self-reported completion fails, what independent checks look like, and walks through a fictional calendar-booking bot to show where verification has to sit outside the agent's own output. By BotBento Editorial · Published 2026-09-26 · Updated 2026-09-26 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [Why 'the agent said it worked' isn't verification](#why-self-report-fails) - [What independent verification actually checks](#build-independent-checks) - [A worked example: a fictional meeting-scheduling bot](#worked-example) - [What to log so verification is possible later](#what-to-log-for-evidence)[Read as Markdown ](/text/blog/verify-ai-agent-task-completion/index.md) ## Key takeaways - An agent's success message is generated the same way as the rest of its output — it is not evidence the work happened. - Verification has to check the actual side effect (a file, a calendar entry, a sent email), not the agent's description of that side effect. - Anthropic's own engineering notes describe agents marking features complete without testing them end-to-end — this is a known, named failure mode, not a rare bug. ## Why 'the agent said it worked' isn't verification An agent's completion message is a claim about the task, not an independent observation of the result. Even when it summarizes successful tool calls, the message may omit an error, a delayed failure or a missing end-to-end check. Treat 'done' as a signal to inspect evidence rather than as the evidence itself. Anthropic's engineering write-up on long-running agent harnesses describes a concrete version of this failure: an agent marked a feature complete after making code changes and running limited checks, even though the feature did not work end-to-end. A passing unit test or a successful request to a development server did not establish that a person could use the finished feature. In that experiment, explicit end-to-end testing with browser automation helped the agent find and fix problems that were not obvious from the code alone. The lesson is to choose an observation that matches the user's requested outcome. A prompt asking the agent to be careful cannot replace an actual check of the result. Sources: [Effective harnesses for long-running agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents). ## What independent verification actually checks Independent verification means checking evidence beyond the agent's narration. In coding tasks, an automated test can exercise behavior in the resulting program and report a pass or failure. Anthropic points to this verifiability as one reason coding tasks can suit agents: the agent can use test results as feedback and revise its work. Tests still have limits, so the check must cover the behavior the user actually requested. Outside coding, define a check for each side effect. Read the calendar API to confirm the intended event and attendees; inspect the database row for the expected value; inspect a mail provider's status for the exact message. A provider's 'accepted' response means it took the request, not that a recipient received the message. Even a 'delivered' event generally indicates acceptance by a receiving server, not proof that a person saw it in an inbox. Anthropic's agent-evaluation guidance separates an agent's work from the grading of its result. For a coding evaluation, the agent changes an environment and tests then grade the working result. The same separation is useful in an operational workflow: define the expected state first, let the agent act, then inspect that state with a read-only check whose output can be reviewed. Sources: [Effective harnesses for long-running agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents), [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents), [Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents). ## A worked example: a fictional meeting-scheduling bot Say a fictional bot, 'Scheduler,' is asked to book a 3pm meeting and send a confirmation email. Scheduler calls a calendar tool, calls an email tool, and returns: 'Done — meeting booked and confirmation sent.' That sentence does not establish whether either tool call achieved the requested outcome. The calendar API could have returned a conflict that Scheduler misread as success. The email tool could have accepted the request but later reported a bounce. Neither problem is resolved by the bot repeating its own summary. A verification step for this fictional example would read the calendar API for an event at 3pm with the expected attendees and use the email provider's message ID to inspect its status. A provider delivery event supports a narrower claim than 'the recipient read the email'; a bounce means that part of the task failed. If the calendar entry is absent or the mail status is unresolved, the run should remain incomplete or pending. This illustrates a check design, not the behavior of a real scheduling product. ## What to log so verification is possible later Verification only works if there's something to check against. That means logging the tool call's actual arguments and the raw response from the external system — not just the agent's summary of that response. If a calendar tool returns an error code, the log needs the error code, not the agent's paraphrase of it. Anthropic's guidance on building tools for agents recommends recording more than a top-level success score: tool errors, call and task duration, number of calls and token use can reveal what happened during a run. Keep the verification result and the external system's identifier alongside those operational details. That lets a reviewer distinguish an unchecked claim from a check that found a real problem. BotBento is pre-release and does not currently offer a built completion-verification feature. The team is exploring how a run record might separate the agent's self-reported status from an independently-checked one, but nothing in that direction is available yet — the practical advice above holds regardless of what tooling a bot builder uses. Sources: [Writing effective tools for AI agents—using AI agents](https://www.anthropic.com/engineering/writing-tools-for-agents). ## Primary sources Sources checked 2026-09-26. Standards and product documentation can change; follow the linked version when implementing. - [Effective harnesses for long-running agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents) — Anthropic - [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) — Anthropic - [Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) — Anthropic - [Writing effective tools for AI agents—using AI agents](https://www.anthropic.com/engineering/writing-tools-for-agents) — Anthropic BotBento is in development. [Suggest a correction](/contact/). --- Canonical: https://botbento.com/bots/ Format: Markdown representation of the public HTML page. BOTS # Starting points, not finished staff. A template is a bot someone already shaped: a name, a job, and the connections it needs to do that job. You copy it into a workspace, rename it, and decide what it may reach. A template never carries conversations, runtime access or credentials with it. Image: The six bot shapes drawn by BotBento: an acid circle, a cyan squircle, a pink pill, a lilac drop, a tangerine blob and an acid Bento, each with two oval eyes. Image: The six bot shapes drawn by BotBento: an acid circle, a cyan squircle, a pink pill, a lilac drop, a tangerine blob and an acid Bento, each with two oval eyes. Six shapes and five colours, so the bots in a conversation stay apart at a glance. People keep their own pictures; a bot never borrows one. - ### Researcher Template Reads what you connect and brings back findings with the source next to each one. Good first bot for a team that argues about evidence. Bento · Needs Web reading, Files - ### Planner Template Turns findings into a short plan with owners left blank, because it does not get to decide who does the work. Drop · Needs Notes - ### Morning brief Template Posts one summary into a conversation at a fixed time each weekday. Nothing else, no digest of digests. Pill · Needs Calendar, Notes - ### Repo watcher Template Reports what changed on the branches you care about, in the words your team uses rather than commit subjects. Squircle · Needs Repositories - ### Meeting notes Template Takes the notes nobody wants to take, and leaves the decisions where the people who made them can correct them. Circle · Needs Calendar, Notes, Files - ### Second pair of eyes Template Reviews a draft against rules you wrote down, and quotes the rule it thinks you broke. Blob · Needs Files None of these exist to download yet. They describe the bots we are building the library around, and what each one would need to be given. A reviewed public library is planned; private sharing inside a workspace comes first. [Join the list](/#waitlist) to hear when the first templates are shareable. ## Suggest a bot Tell us what you would connect, and why. We read every suggestion. Nothing appears on this page without a read-through first. Form endpoint: /api/suggest What should this bot be called? A short, plain name. Nova, Planner, Repo watcher. Link (optional) Documentation, an API, or anything that explains it. What is its job? Up to 1,000 characters. What it does, and which connections it would need. Your email address I agree that BotBento can store this suggestion and the email address I provided in order to follow it up. [Privacy notice](/privacy/). Send suggestion Sending a suggestion does not add you to the early-access list. [Join the list](/#waitlist) if you want the updates too. --- Canonical: https://botbento.com/contact/ Format: Markdown representation of the public HTML page. LET’S TALK # A human on the other side. Use our contact form for product questions, early-access help or privacy requests. Prefer email? Say hello at [hello@botbento.com](mailto:hello@botbento.com), or get help at [support@botbento.com](mailto:support@botbento.com). Form endpoint: /api/contact Your name (optional) Your email address What would you like us to know? Up to 5,000 characters. Please don’t include passwords, API keys, payment details or other secrets. I agree that BotBento can store my message and use the information I provide to handle my request. [Privacy notice](/privacy/). Send message Your message is saved for the team. We can reply using the email address you provide. ## Privacy and your signup For removal from the early-access list or questions about your information, email [privacy@botbento.com](mailto:privacy@botbento.com) or use the form above. Include the email address associated with your signup where possible. ## Reporting a problem A short description of what happened and the page URL are a useful starting point. Read our [privacy notice](/privacy/) for details about how contact messages are handled. --- Canonical: https://botbento.com/editorial/ Format: Markdown representation of the public HTML page. HOW WE WRITE # Useful before frequent. BotBento publishes practical explanations for people building and working with AI agents. We want each article to answer a real question. ## AI assistance is disclosed Our editorial workflow uses AI for research, drafting and content checks. Article bylines identify the editorial organization rather than inventing an individual author. AI assistance does not make a claim correct; readers should be able to trace factual statements to primary sources. ## Sources and examples We prefer official documentation, standards, original research and first-party engineering reports. Citations sit near the sections they support. Examples we devise are labelled illustrative and must not be presented as measured customer results. Documentation versions and source-check dates are visible when relevant. ## What we publish Our planned cadence is two useful articles daily and one weekly newsletter. A schedule does not override quality: incomplete, unsupported or repetitive drafts stay unpublished. We do not pad articles to hit a keyword target, invent release announcements or claim BotBento features are available before public availability is confirmed. ## Dates and corrections Publication dates record when an article first appears. Updated dates change when its substance changes. We keep sources linked so readers can identify newer guidance. [Send a correction or question](/contact/) with the article URL and the statement you want us to check. ## Newsletter consent The newsletter has its own signup. Joining early access does not enroll you automatically. Every newsletter includes an unsubscribe route. Read published issues in the archive, with or without a subscription. ## Search and agent access Articles are available as HTML and Markdown, with feeds and a machine-readable index. These formats help readers and tools access the same content. They do not guarantee search ranking, rich results, indexing or inclusion in AI answers. Policy updated 2026-09-07. [Privacy notice](/privacy/). --- Canonical: https://botbento.com/ Format: Markdown representation of the public HTML page. Illustration: Pill, Drop and Bento, the three sculpted BotBento characters. # A workspace for people and AI bots. Chat with your bots, give them tools and put repeat work on a schedule. Your team stays in the conversation. [Join early access ](#early-access) ## Start with a shared brief. Give Pill the research. Let Drop turn it into a plan. Your team can follow the work in the same conversation. ## The work comes back to your team. Review the sources, discuss the next step, or make the job a routine. ## Their tools. Your permission. People and bots follow the same permissions. Choose what each bot can access on local or cloud machines. ## Make room for your AI team. We're building BotBento around Hermes. Join the list for updates and an invitation when access opens. [Questions about BotBento?](#questions) Launch workspaceExample You to Pill + Drop ### Help us plan the launch. Compare three ways to reach our first customers. Bring back the sources. Research Project plan Daily brief #### Pill Reads the sources you connect. Research #### Drop Turns the findings into next steps. Plan Your connected toolsEnabled for these bots Back in the conversationExample Launch workspace ### A launch plan you can review. - 01Three channels, comparedResearch - 02Sources beside every findingEvidence - 03A first experiment to tryNext step Review the options and their sources, then decide what to try. Make it a routine Every weekday08:30 AM MTWTFSS Pill's accessExample ### What can Pill do? 2 permissions enabled Use connected tools Manage bots Invite people Read private conversations Try the switches. These example permissions don't change your account. Join the teamIn development Form endpoint: /api/early-access Email address Join early access I agree to receive BotBento early-access and product emails. I can unsubscribe at any time. [Privacy notice](/privacy/). Joining the list is free. It doesn't create an app account. [Read the FAQ](#questions)[Field notes](/blog/)[Get in touch](/contact/) In development Pause animation Your bots. Part of the team. [Brief](#workspace)[Result](#result-heading)[Access](#access) --- Canonical: https://botbento.com/newsletters/ Format: Markdown representation of the public HTML page. BOTBENTO / NEWSLETTER # A good week for a little progress Subscribe to the weekly BotBento newsletter for practical AI agent articles and product news. Browse published issues. ONCE A WEEK / A LITTLE MORE SIGNAL ## Useful notes for your next bot. Get the weekly BotBento newsletter: selected articles, practical examples and product news when there is something to share. A separate subscription from early access. Confirm by email before your subscription starts. Form endpoint: /api/newsletter Your email address Subscribe ' + ARROW\_NE + 'I agree to receive the weekly BotBento newsletter. I can unsubscribe at any time. [Privacy notice](/privacy/). ## Published issues 2026-09-252 min read ## [Weekly field notes: recover a bot without guessing ](/newsletters/weekly-field-notes-2026-09-25/) This week’s BotBento field notes cover MCP deprecations, rate-limit recovery, and how to make a retry safe before an agent acts again. 2026-09-112 min read ## [Weekly field notes: make an agent’s next step checkable ](/newsletters/weekly-field-notes-2026-09-11/) This week’s BotBento field notes cover run records, local versus cloud placement and a practical checklist for deciding what an agent should do next. 2026-09-072 min read ## [The launch edition: a place to start with AI bots ](/newsletters/launch-edition-a-place-to-start/) The first BotBento newsletter introduces two practical guides, a small exercise for your next bot and the current early-access product status. --- Canonical: https://botbento.com/newsletters/launch-edition-a-place-to-start/ Format: Markdown representation of the public HTML page. [Home](/) / [Newsletter](/newsletters/) FIELD NOTES / 2 MIN READ # The launch edition: a place to start with AI bots Start with one useful task. This launch edition brings together our first guides to agent design and tool permissions, with a small exercise you can try before connecting another service. By BotBento Editorial · Published 2026-09-07 · Updated 2026-09-07 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [Start with the work you want finished](#start-with-the-work) - [Connect tools with a specific purpose](#connect-with-intent) - [Where BotBento is today](#where-botbento-is)[Read as Markdown ](/text/newsletters/launch-edition-a-place-to-start/index.md) ## Key takeaways - Read our first two guides on choosing an agent or workflow and reviewing MCP tool permissions. - Write the intended output and permitted actions for one small task before expanding its access. - BotBento remains in development. The newsletter and early-access list are separate subscriptions. ## Start with the work you want finished Welcome to the first BotBento newsletter. We are opening the public archive with two guides about decisions that come before a useful bot: who chooses the next step, and what the bot is allowed to do. Our aim is to make those decisions easier to discuss with the people who will use and maintain the result. Our agents-versus-workflows guide starts from Anthropic’s distinction between predefined paths and model-directed processes. Its original weekly-research example shows how to give a bot a narrow investigation while keeping the schedule, destination and review requirement explicit. Follow the related-article links below to read the full examples and sources. Sources: [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents). ## Connect tools with a specific purpose The second guide examines MCP tool permissions through an illustrative calendar connection. MCP’s authorization specification addresses access to protected servers; a useful task still needs a correctly selected account and a clear request. The guide walks through connection review and small checks you can carry out before depending on an integration. Here is a short exercise for the week: choose a task you repeat, such as preparing a note from public release pages. Write one sentence describing the finished result and one sentence listing the actions that are allowed. Then name the action that should require another decision. If you struggle to write those sentences, start by narrowing the task rather than adding more tools. This is an editorial exercise, not a tested product benchmark. Sources: [MCP authorization specification, 2026-07-28](https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization). ## Where BotBento is today BotBento is being developed as a conversation-first workspace for Hermes bots, chats, routines and plugins. You can join the early-access list and explore a preview of the workspace on our homepage. Public app access is not open yet. We will share availability news when there is a clear next step for people who want to try it. The newsletter has its own confirmation-based signup. Joining the early-access list does not subscribe you automatically. Read every issue here, with or without a subscription. Our planned rhythm is a useful weekly selection rather than a copy of every article. You can browse the blog, read the Markdown versions with your preferred tools, or send a question through the contact page. Future issues will focus on practical examples, carefully sourced explanations and product news when there is something verified to share. ## Primary sources Sources checked 2026-09-07. Standards and product documentation can change; follow the linked version when implementing. - [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) — Anthropic - [MCP authorization specification, 2026-07-28](https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization) — Model Context Protocol BotBento is in development. [Suggest a correction](/contact/). ## Keep reading - [AI agents vs workflows: choose who decides the next step](/blog/ai-agents-vs-workflows/) - [MCP tool permissions: what to check before connecting a bot](/blog/mcp-tool-permissions/) --- Canonical: https://botbento.com/newsletters/weekly-field-notes-2026-09-11/ Format: Markdown representation of the public HTML page. [Home](/) / [Newsletter](/newsletters/) FIELD NOTES / 2 MIN READ # Weekly field notes: make an agent’s next step checkable A useful agent leaves a result someone can inspect. This issue gathers three practical guides on run evidence, execution placement and permission boundaries, then offers a small exercise for your next routine. By BotBento Editorial · Published 2026-09-11 · Updated 2026-09-11 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [Make the result checkable](#make-the-result-checkable) - [Map where each part of the agent runs](#map-the-runtime) - [Review the connection before the tool acts](#review-the-connection)[Read as Markdown ](/text/newsletters/weekly-field-notes-2026-09-11/index.md) ## Key takeaways - Start a routine with an acceptance condition and retain enough evidence to distinguish completion from an unknown effect. - Choose where a model, tool and scheduler run separately; a local screen does not prove a local data path. - BotBento remains in development, and this archive issue is separate from early-access signup and email delivery. ## Make the result checkable This week’s run-record guide starts with a small distinction: a successful tool call proves that one operation returned, while the requested outcome may still be unverified. An illustrative research routine should record its scope, useful output, unresolved actions and the next safe step. OpenTelemetry’s trace model is a useful vocabulary for connecting operations, but it does not decide the business acceptance condition for your task. Try this exercise before your next scheduled routine: write one sentence for the finished result, one sentence for the evidence you will inspect, and one sentence for what should happen if an external write times out. If the last sentence says “try again,” add how you will check whether the first write already happened. Sources: [Traces](https://opentelemetry.io/docs/concepts/signals/traces/). ## Map where each part of the agent runs Our local-versus-cloud guide separates model processing, tool execution and the process that starts scheduled work. Hermes configuration documents local, container and remote terminal choices as execution settings; they do not by themselves establish where every model request or scheduler lives. Draw the path for one task and name the actual input, output and owner for each step. A laptop window can send text to a hosted model, and a remote tool environment does not automatically give that environment access to a laptop folder. Make the boundary visible before you compare arrangements. This issue offers no universal speed or cost winner and does not imply that BotBento’s future interface is available today. Sources: [Hermes Agent Configuration](https://hermes-agent.nousresearch.com/docs/user-guide/configuration/). ## Review the connection before the tool acts The tool-permissions guide uses an illustrative calendar connection to show why a named integration is not the same as an approved action. Read the provider’s authorization scope, the selected account and the task’s allowed operation separately. The MCP authorization specification describes protected-resource authorization flows; your application still needs to show which connection is in use and what the next tool call may do. BotBento is in development. The public blog and newsletter archive are educational surfaces, and early-access signup is a separate confirmation-based subscription. Marketing delivery remains gated until the required postal footer is supplied; this archive issue does not claim that an email was sent. Sources: [MCP Authorization](https://modelcontextprotocol.io/specification/2025-06-18/basic/authorization). ## Primary sources Sources checked 2026-09-11. Standards and product documentation can change; follow the linked version when implementing. - [Traces](https://opentelemetry.io/docs/concepts/signals/traces/) — OpenTelemetry - [Hermes Agent Configuration](https://hermes-agent.nousresearch.com/docs/user-guide/configuration/) — Nous Research - [MCP Authorization](https://modelcontextprotocol.io/specification/2025-06-18/basic/authorization) — Model Context Protocol BotBento is in development. [Suggest a correction](/contact/). ## Keep reading - [What should an AI agent record after each run?](/blog/ai-agent-run-result-record/) - [Local or cloud: where should your AI agent run?](/blog/local-vs-cloud-ai-agents/) - [MCP tool permissions: what to check before connecting a bot](/blog/mcp-tool-permissions/) --- Canonical: https://botbento.com/newsletters/weekly-field-notes-2026-09-25/ Format: Markdown representation of the public HTML page. [Home](/) / [Newsletter](/newsletters/) FIELD NOTES / 2 MIN READ # Weekly field notes: recover a bot without guessing Two changes call for the same habit: identify who owns the action, preserve evidence, and make the next step explicit. This issue links the week’s MCP and retry guides and offers a short review exercise. By BotBento Editorial · Published 2026-09-25 · Updated 2026-09-25 AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/). ## In this article - [Know which side owns the next model call](#the-client-boundary) - [Retry at the right layer, with a stopping point](#retry-at-the-right-layer) - [The public guides this week](#what-we-shipped)[Read as Markdown ](/text/newsletters/weekly-field-notes-2026-09-25/index.md) ## Key takeaways - MCP Sampling and Roots are deprecated, but the specification keeps each for at least twelve months before removal eligibility. - A rate limit may arrive as HTTP 429 or as an MCP tool execution error; inspect the actual signal before retrying. - For an uncertain write, check for an idempotency guarantee or read back the result before trying again. ## Know which side owns the next model call Our guide to MCP Sampling explains the 2026-07-28 deprecation and what it means when a server has been asking a client to run model generations. The specification says new implementations should avoid Sampling and existing integrations should plan direct provider access. It also deprecates Roots, the client’s list of relevant filesystem locations. These changes are not a signal to remove approval or permission checks: Roots was always guidance, not an access-control boundary. Try a five-minute inventory for one connected server. Does your client advertise Sampling or Roots? Does the server actually request either feature? If the server will call a model directly, who will own its provider account, data handling, and review path? If it will receive file paths another way, where are filesystem permissions enforced? The answers turn a protocol announcement into a concrete migration plan. Sources: [Sampling - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/client/sampling), [Roots - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/client/roots). ## Retry at the right layer, with a stopping point Our rate-limit guide separates HTTP 429 from an MCP tool result marked isError. The MCP Tools specification also distinguishes a tool execution error from a JSON-RPC protocol error. A bot that collapses all three into “the connection broke” may retry malformed arguments or lose the provider’s wait instruction. An HTTP 429 may include Retry-After, but that header is optional, so the bot needs both provider-specific guidance and a bounded fallback. Here is a small exercise for the next routine you automate. Write down the tool call, whether it reads or writes, the exact rate-limit signal, the allowed wait, and the retry cap. Then add one line for the uncertain case: if the call timed out after a write, how will you tell whether it already succeeded? A safe answer may be an idempotency key or a readback query. Without one, report uncertainty rather than silently repeating the write. Sources: [Tools - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-06-18/server/tools), [429 Too Many Requests - HTTP | MDN](https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Status/429). ## The public guides this week You can read the full articles on the BotBento blog: “MCP Sampling Is Deprecated. What Should Bot Builders Do Now?”, “How Should an AI Bot Handle a Tool’s Rate Limit Error?”, and “Agent Tool Retries Need Idempotency Keys.” Each gives a different angle on the same operational question: what evidence makes the next action safe? BotBento remains in development. The articles and this archive are educational material, not a claim that a finished bot runtime, automated run record, or provider integration is available. Newsletter email is a separate confirmed-subscription channel; publishing this archive alone does not mean an email was delivered. ## Primary sources Sources checked 2026-09-25. Standards and product documentation can change; follow the linked version when implementing. - [Sampling - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/client/sampling) — Model Context Protocol - [Roots - Model Context Protocol](https://modelcontextprotocol.io/specification/draft/client/roots) — Model Context Protocol - [Tools - Model Context Protocol](https://modelcontextprotocol.io/specification/2025-06-18/server/tools) — Model Context Protocol - [429 Too Many Requests - HTTP | MDN](https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Status/429) — MDN Web Docs BotBento is in development. [Suggest a correction](/contact/). ## Keep reading - [MCP Sampling Is Deprecated. What Should Bot Builders Do Now?](/blog/mcp-sampling-deprecated/) - [How Should an AI Bot Handle a Tool's Rate Limit Error?](/blog/ai-bot-rate-limit-error-handling/) - [How Do You Stop a Retried Tool Call From Running Twice?](/blog/agent-tool-retry-idempotency-keys/) --- Canonical: https://botbento.com/plugins/ Format: Markdown representation of the public HTML page. PLUGINS # What a bot is allowed to touch. A plugin is a connection a bot can be given: a calendar, a repository, a machine of your own. You install it, you approve it once, then you choose which bots may use it — one bot at a time, never all of them by default. Everything below is being built. Nothing here is downloadable yet, and the list will change as we find out what people actually connect. [Join the list](/#waitlist) to hear when the first ones ship. - ### Your machine In development Hermes runs on hardware you own. Pair the machine once, then choose which bots may reach it. Nothing leaves the machine unless a bot you allowed asks for it. - ### Managed environment Planned A private Linux environment we keep running for you, with persistent files, a maintained Hermes and an included AI allowance. For people who want the bots without the machine. - ### Calendar In development Reads the events a bot needs to plan around. Writing to your calendar is a separate permission you grant on its own. - ### Notes In development Reads and writes the notes a bot is allowed to touch, so a plan it wrote has somewhere to live. - ### Files In development Attachments and documents shared in a conversation, available to the bots in that conversation and no others. - ### Repositories In development Reads branches and files from a repository you connect, so a bot can answer questions about code that already exists. - ### Web reading In development Fetches pages you allow and keeps the source beside every finding, so you can check the work rather than trust it. - ### Computer viewing and control In development Watch a bot work on a screen, or take over. Authorised separately from every other permission, and never implied by one. - ### Boards and notes Planned Native boards and notes inside the workspace rather than another tab to keep open. How access works. A connection belongs to you, not to a bot. Granting one bot the calendar does not grant it to the others, and removing a plugin removes it everywhere at once. Reading private conversations is never part of a plugin. ## Suggest a plugin Tell us what you would connect, and why. We read every suggestion. Nothing appears on this page without a read-through first. Form endpoint: /api/suggest What should we connect? The name of the service, tool or device. Link (optional) Documentation, an API, or anything that explains it. What would a bot do with it? Up to 1,000 characters. What it should read, and what it should be allowed to change. Your email address I agree that BotBento can store this suggestion and the email address I provided in order to follow it up. [Privacy notice](/privacy/). Send suggestion Sending a suggestion does not add you to the early-access list. [Join the list](/#waitlist) if you want the updates too. --- Canonical: https://botbento.com/privacy/ Format: Markdown representation of the public HTML page. YOUR INFORMATION # Privacy, in plain words. Updated 17 September 2026 This notice covers the BotBento public website, its early-access email list and messages you send us. The BotBento application is in development; this notice does not describe future app or connected-provider data practices. ## Who to contact BotBento operates this website. For privacy questions, requests to access or correct your signup information, or removal from the list, email [privacy@botbento.com](mailto:privacy@botbento.com) or use our [contact form](/contact/). ## When you join early access We collect the email address you enter and your agreement to receive early-access and product emails. We use this information to maintain the list and send related updates and invitations. Signup records may include the time of signup and consent so we can administer the list. Joining puts your address on the update list straight away, with no confirmation link to click, because the consent checkbox on the form is the record of your agreement. We send a short welcome email, then updates while we build and an invitation when access opens. Every one of those emails carries an unsubscribe link. Joining the list does not create an app account. You can withdraw your email consent at any time by using an unsubscribe option in our emails or contacting us. We may retain a minimal record of an unsubscribe request to avoid adding you back to future mailings. ## When you subscribe to the newsletter The weekly newsletter has a separate signup from early access. We store your email address, your consent and subscription status. We send a confirmation request and activate the subscription only after you confirm. You can unsubscribe using the link in each newsletter or request help through our [contact form](/contact/). We use an email delivery provider to send early-access updates, newsletter confirmations and newsletter issues. That provider processes your email address, message content and delivery information for that purpose. Delivery records, including bounces and unsubscribe requests, help us maintain the list and stop messages when appropriate. Email open and click tracking are disabled for these messages. ## When you suggest a plugin or a bot The [plugins](/plugins/) and [bots](/bots/) directories each have a suggestion form. We store what you suggested, any link you added, the email address you entered, your consent and the submission time, so we can follow the suggestion up with you. Sending a suggestion does not add you to the early-access list and does not publish anything. Nothing appears in a directory without a person reading it first, and we will not put your email address on a public page. ## When you contact us We receive the name (if provided), email address and message you enter in the contact form, along with your consent and submission time. Messages are saved in our support inbox. We use this information to handle your request and can reply using the email address you provide. Please avoid sending sensitive information or account secrets. If you email us directly, Cloudflare Email Routing forwards your message to our team inbox hosted by Hostinger. Those providers process your email address, message and delivery information to deliver and protect the correspondence. We use it to handle your request. ## Running and protecting the website Our hosting, DNS and security providers process technical information needed to deliver and protect the website. This can include IP addresses, requested URLs, browser information, timestamps and security events. Signup requests may be checked or limited to prevent spam and abuse. ## Optional analytics When optional analytics is configured, we ask before loading it. If you allow it, we use Umami to understand website visits, including pages viewed, referring sites, browser and device categories, and approximate location. Technical request information, including an IP address, reaches the analytics service as part of receiving the request. We do not intentionally send your signup email address or form contents to analytics. Your analytics choice is stored in your browser’s local storage. You can change it through “Analytics preferences” in the footer when analytics is available. Declining analytics does not affect signup or access to the website. Browser extensions or privacy settings can also prevent analytics from loading. ## Service providers We use infrastructure and service providers to host and protect the site, store and deliver email, and operate optional analytics. Information needed for those functions may be processed by those providers. Infrastructure can be located in countries other than your own. We do not sell the early-access list. ## How long we keep information We keep early-access information while it is needed for the list and its purpose. You can request removal at any time. Support correspondence and technical logs may be retained as needed to respond to requests, investigate problems and maintain security. Some information may remain temporarily in backups or where retention is required. ## Your choices You can ask us to access, correct or delete information you provided. Depending on where you live, you may have additional privacy rights. We may need to verify that a request comes from the person associated with the information. For signup requests, use the address associated with your signup in the contact form. ## Changes to this notice We will update this page when the website’s data practices change. Before app access opens, we will publish information that reflects the application and any connected services.