Canonical: https://botbento.com/blog/bot-tool-call-tracing-otel/
Format: Markdown representation of the public HTML page.

[Home](/) / [Blog](/blog/)

FIELD NOTES / 5 MIN READ

# How Do You Trace a Bot's Tool Calls to Debug a Failed Run?

When a bot run fails, related tool, model and agent spans can make the execution path easier to inspect. OpenTelemetry's GenAI conventions describe their names and attributes, but the conventions are still at Development stability. A trace can show reported errors and timing; it does not establish that the bot achieved the requested result.

By BotBento Editorial · Published 2026-10-06 · Updated 2026-10-06

AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/).

## In this article

- [A Failed Bot Run Looks Like Scattered Log Lines](#the-problem)
- [What OpenTelemetry's GenAI Conventions Actually Standardize](#what-the-spec-defines)
- [A Worked Example: Tracing a Failed Run (Illustrative)](#worked-example)
- [What a Trace Span Can't Tell You](#what-tracing-does-not-tell-you)
- [What to Set on Every Tool-Call Span](#decision)[Read as Markdown ](/text/blog/bot-tool-call-tracing-otel/index.md)

## Key takeaways

- OpenTelemetry describes execute\_tool, chat and invoke\_agent operations that can appear in a related trace when the application propagates context and instruments those steps.
- For a GenAI execute-tool span, gen\_ai.operation.name and gen\_ai.tool.name are required attributes and the recommended span kind is INTERNAL. An MCP client/server exchange has its own CLIENT and SERVER spans.
- A trace can show a reported tool error and timing, but a clean span is not proof that the bot's task succeeded. Verify the tool result and intended outcome separately.

## A Failed Bot Run Looks Like Scattered Log Lines

A bot that plans, calls three tools, and hands part of the job to a sub-step produces one log line per step if you're lucky, and no connection between those lines. When the run fails, you're left guessing which call was the problem: did the model ask for the wrong tool, did the tool return an error, or did a downstream API just time out? Without a single structured record of the whole run, you reconstruct the failure by timestamp-matching separate log statements, which gets slower every time the bot adds a tool.

Tracing offers one practical way to connect those steps: create a span for each instrumented operation and propagate trace context so the spans can be viewed together. The trace is only as complete as the instrumentation. If a tool or downstream call does not propagate context, its work can still be missing from the tree; a dashboard cannot fill that gap by matching timestamps alone.

## What OpenTelemetry's GenAI Conventions Actually Standardize

OpenTelemetry's developing GenAI semantic conventions give instrumented model, agent and tool operations a shared vocabulary under the gen\_ai namespace. For an execute-tool span, gen\_ai.operation.name and gen\_ai.tool.name are required attributes; the operation name should be execute\_tool. The guidance encourages developers to instrument tool calls made by their own code when automatic instrumentation does not cover them. These attributes make different steps easier to search and compare, but they do not create a trace or propagate context on their own.

Span kind depends on which operation is being described. The GenAI execute-tool span guidance says its kind should be INTERNAL, even when the tool implementation later calls a remote API. A separate HTTP client span can describe that outbound request. The GenAI model-inference guidance discusses CLIENT and, for an in-process model, INTERNAL; applying that inference rule to tool-execution spans would be a category error. MCP instrumentation has a different boundary again: its tools/call client and server spans use CLIENT and SERVER respectively.

The execute-tool guidance says span status should follow OpenTelemetry's error-recording rules, and error.type is conditionally required when the operation ends in an error. gen\_ai.agent.name is conditionally required when applicable. Those fields help identify a reported failure and the associated agent when the instrumentation records them. They do not guarantee that the tool's returned data was correct, or that every downstream failure was classified and attached to the span.

The MCP conventions recommend gen\_ai.operation.name=execute\_tool for a tool call and mcp.method.name for the protocol method, such as tools/call. This lets consumers relate an MCP call to other tool activity while retaining the protocol boundary. If an MCP request is part of a session, mcp.session.id is recommended; a stateless call has no session ID to add. Correlating the MCP span with a downstream API span still requires propagated trace context and appropriate instrumentation.

Sources: [semantic-conventions-genai/docs/gen-ai/gen-ai-spans.md at main · open-telemetry/semantic-conventions-genai](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-spans.md), [semantic-conventions-genai/docs/gen-ai/mcp.md at main · open-telemetry/semantic-conventions-genai](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/mcp.md).

## A Worked Example: Tracing a Failed Run (Illustrative)

Imagine a fictional InvoiceBot reconciling a vendor invoice with a purchase order. It makes a model call, invokes lookup\_po, then invokes flag\_mismatch. If those operations are instrumented and trace context is carried between them, a trace could show an invoke\_agent parent, a model chat span, and two execute\_tool spans. The flag\_mismatch tool span might have error.type=rate\_limited if that operation reports a rate-limit failure. The span names and attributes help you find where the reported error occurred; the example does not claim a measured BotBento run.

That tree narrows the investigation, but it does not prove the model chose the right tools or that lookup\_po returned the right purchase order. A lookup\_po span without a reported error says only that its instrumentation did not record one. Inspect the tool response and the bot's validation result before treating the lookup as successful. If flag\_mismatch reports a rate limit, inspect the error detail and any child span for its downstream API call before deciding where the limit originated.

If flag\_mismatch runs over MCP, an MCP tools/call CLIENT span and its SERVER counterpart can carry mcp.method.name and, only when the request belongs to a session, mcp.session.id. A downstream HTTP span or server log may reveal whether a rate limit came from the MCP server's own policy or an API it called. The method and session fields alone cannot identify that origin, and missing trace propagation can leave the downstream span outside the tree.

Sources: [semantic-conventions-genai/docs/gen-ai/gen-ai-spans.md at main · open-telemetry/semantic-conventions-genai](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-spans.md), [semantic-conventions-genai/docs/gen-ai/mcp.md at main · open-telemetry/semantic-conventions-genai](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/mcp.md).

## What a Trace Span Can't Tell You

A recorded span shows an instrumented operation, its observed duration and any error the instrumentation attached. It cannot by itself establish that the operation produced the right business result. flag\_mismatch might return a syntactically valid response while flagging the wrong invoice, or a tool might quietly perform no update. A trace without a recorded error can therefore coexist with a failed task. Keep an application-level result check alongside tracing, such as comparing the intended invoice ID with the record actually changed.

The GenAI span document labels these semantic conventions Development rather than Stable. Treat exact attribute-based dashboards and alerts as maintained integrations: pin the convention version your instrumentation emits, test your queries after SDK upgrades, and keep the original tool result available for debugging without placing secrets or sensitive arguments in spans. The published convention is guidance for instrumentation, not evidence that any particular bot framework emits all of its fields automatically.

Sources: [semantic-conventions-genai/docs/gen-ai/gen-ai-spans.md at main · open-telemetry/semantic-conventions-genai](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-spans.md).

## What to Set on Every Tool-Call Span

For a bot you operate, instrument the tool execution with gen\_ai.operation.name=execute\_tool and gen\_ai.tool.name, and record the operation's error status and error.type when it fails. The GenAI execute-tool span should be INTERNAL. If the tool makes an outbound HTTP call, let HTTP instrumentation describe that call with its own client span, then propagate trace context so it remains connected. For an MCP tool call, record mcp.method.name on the MCP client/server spans and mcp.session.id only if the call is part of a session.

Use the trace to ask where an observed error or delay appeared; use the bot's result record and a separate verification step to ask whether the task was completed correctly. In the invoice example, verify the invoice and purchase-order identifiers and inspect whether the intended mismatch flag was actually set. Correction, 2026-10-06: an earlier version applied model-inference CLIENT span guidance to a GenAI tool-execution span and treated MCP session IDs as universal; the current source guidance distinguishes those span boundaries and makes session IDs conditional.

Sources: [semantic-conventions-genai/docs/gen-ai/gen-ai-spans.md at main · open-telemetry/semantic-conventions-genai](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-spans.md), [semantic-conventions-genai/docs/gen-ai/mcp.md at main · open-telemetry/semantic-conventions-genai](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/mcp.md).

## Primary sources

Sources checked 2026-10-06. Standards and product documentation can change; follow the linked version when implementing.

- [semantic-conventions-genai/docs/gen-ai/gen-ai-spans.md at main · open-telemetry/semantic-conventions-genai](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-spans.md) — OpenTelemetry
- [semantic-conventions-genai/docs/gen-ai/mcp.md at main · open-telemetry/semantic-conventions-genai](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/mcp.md) — OpenTelemetry

BotBento is in development. [Suggest a correction](/contact/).
