Which of Your Bot's Tool Calls Actually Need a Human Approval Gate?
MCP's own spec only says a human should be able to deny a tool call; it doesn't say which calls need a prompt. This piece gives a concrete three-tier method for deciding, using a fictional calendar-and-billing bot, then shows how Claude Agent SDK's six-step permission chain and OpenAI Agents SDK's require_approval setting implement that method in practice, including the ways both can silently skip the gate.
AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. Editorial policy.
Key takeaways
- Sort tools by reversibility and blast radius, not by how the model phrases the request: reads auto-approve, reversible writes get an ask rule, irreversible or external-facing actions require approval every time.
- Claude Agent SDK's bypassPermissions mode approves everything that reaches it, including tools you never listed, and a subagent cannot override a parent session running in that mode.
- OpenAI Agents SDK's require_approval is per-tool, not per-call-risk: a dict mapping tool names to "always"/"never" policies, so it only gates what you explicitly name.
MCP tells you to have a gate. It doesn't tell you where to put it.
The Model Context Protocol's tools specification is explicit about one thing and silent about everything else. For trust and safety, it says there should always be a human in the loop with the ability to deny tool invocations, and it recommends that clients prompt for confirmation on sensitive operations, show tool inputs before calling, and log usage for audit purposes. What it does not do is define which operations count as sensitive, or standardize an approval UI.
That gap matters because a bot builder who reads only the headline recommendation tends to implement one global switch: prompt on every call, or prompt on none. Neither works. Prompting on every call trains the person approving to click through without reading, which defeats the purpose of the gate. Prompting on none means the first destructive call runs exactly like the thousandth read-only call. The decision that actually needs making is per-tool, made once at connection time, not per-call at run time.
Sources: Tools - Model Context Protocol.
A worked example: sorting one bot's tools into three tiers
Take a fictional bot that manages a small team's calendar and expense approvals. It exposes four tools: read_calendar_events, reschedule_event, send_meeting_cancellation_email, and issue_refund. Sorting these by what happens if the bot is wrong gives you the three tiers that matter in practice.
Tier one is auto-approve: read_calendar_events returns data and changes nothing. Getting it wrong costs a wasted context window, not a wasted afternoon. Tier two is ask-first: reschedule_event and send_meeting_cancellation_email are reversible in principle — you can re-invite people or send a correction — but someone outside the bot's control (a meeting attendee) sees the mistake before you can fix it, so a confirmation prompt buys a cheap second look. Tier three is always-deny-without-explicit-approval: issue_refund moves money and cannot be un-sent. No permission mode, allowlist, or 'the model seemed confident' heuristic should let this one through without a person naming the amount and recipient out loud.
The useful property of this sort is that it survives the bot getting smarter. A better model doesn't move issue_refund into tier one; it just makes fewer mistakes inside tier three. The tiering is about the consequence of being wrong, not about how trustworthy the model currently seems.
How Claude Agent SDK and OpenAI Agents SDK actually enforce this
Anthropic's Claude Agent SDK evaluates every tool request through a fixed six-step chain: hooks first, then deny rules, then ask rules, then the active permission mode, then allow rules, then a canUseTool callback for anything unresolved. Deny rules hold in every permission mode, including bypassPermissions — a deny rule blocks the matching call outright, no matter what else is configured. allowed_tools and disallowed_tools add entries to the allow and deny lists in that chain; they control whether a call is approved, not whether the tool exists for Claude to try calling.
Mapped onto the three-tier example: read_calendar_events goes in allowed_tools so it resolves at the allow-rule step without a prompt. reschedule_event and send_meeting_cancellation_email are left off every list, so they fall through to the permission mode and your canUseTool callback — the SDK's version of ask-first. issue_refund goes in disallowed_tools, a deny rule that holds even if someone later sets the whole session to bypassPermissions for convenience during testing.
OpenAI's Agents SDK takes a flatter approach for MCP tools: a require_approval setting on the tool config, either a single policy string ("always" or "never") or a dict mapping individual tool names to those policies. When a call needs approval, the SDK calls an on_approval_request callback with the tool name and arguments; the callback returns an approve/reject decision, and a rejection can carry a reason string that gets surfaced back to the run. For local (not hosted) MCP servers, the equivalent require_approval argument lives on MCPServerStdio, MCPServerSse, and MCPServerStreamableHttp. There is no three-way ask/allow/deny precedence here — a tool is either on the always-approval list or it isn't, so the tiering work has to happen in how you build that dict, not in the framework's evaluation order.
Sources: Configure permissions - Claude API Docs, Model context protocol (MCP) - OpenAI Agents SDK.
Two ways the gate gets bypassed without anyone noticing
The first failure path is permission-mode creep. Someone sets bypassPermissions to speed through a demo or a CI run, forgets to unset it, and every subsequent tool call — including issue_refund if it wasn't also deny-listed — runs without a prompt. The deny rule is the only thing immune to that mistake; an ask rule or a missing allow entry is not, because bypassPermissions approves anything that reaches the permission-mode step, and only a matching deny rule arrives earlier in the chain than that.
The second failure path is subagent inheritance. When a Claude Agent SDK session running in bypassPermissions spawns a subagent, that subagent inherits bypassPermissions and cannot override it with a different permission mode of its own. A subagent built with a narrower system prompt and a seemingly cautious task description still gets the parent's blanket approval. If your bot architecture delegates the risky step to a subagent specifically to keep it contained, check what mode the parent session is running in first — containment by prompt does not survive an inherited bypass.
A third, quieter failure is the OpenAI-side equivalent: require_approval only gates tools you named. A new tool added to an MCP server after you wrote the dict — a delete_project added alongside list_tasks and get_project — runs under whatever the fallback policy is, which is "never" unless you set a default. Treat a growing tool list on a server you don't control as a reason to re-check the approval dict, not a one-time setup step.
Sources: Configure permissions - Claude API Docs.
The decision to apply before you connect a tool
Before wiring a tool into a bot, write down one sentence: what happens if this call runs and it was wrong? If the answer is 'nothing persists,' auto-approve it. If the answer is 'someone else sees the mistake before I can undo it,' ask first, every time, with the actual arguments shown — not just the tool name — because the MCP spec's own recommendation is to show tool inputs before calling, precisely so the person approving can catch bad arguments on an otherwise fine tool. If the answer is 'money moves, data is deleted, or an external party is notified irreversibly,' put it on a deny-by-default list that nothing short of an explicit, per-call human decision can clear, and make sure that list is enforced as a deny rule or an always-approval entry, not just an omission from an allow list.
Revisit the list whenever the tool server changes, not just when you first connect it. A server that adds a new destructive tool, or renames an existing one, resets your tiering silently unless you're checking the tool list, not just the policy you wrote for the old one.
Sources: Tools - Model Context Protocol.
Primary sources
Sources checked 2026-10-03. Standards and product documentation can change; follow the linked version when implementing.
- Tools - Model Context Protocol — Model Context Protocol
- Configure permissions - Claude API Docs — Anthropic
- Model context protocol (MCP) - OpenAI Agents SDK — OpenAI
BotBento is in development. Suggest a correction.