Canonical: https://botbento.com/blog/verify-ai-agent-task-completion/
Format: Markdown representation of the public HTML page.

[Home](/) / [Blog](/blog/)

FIELD NOTES / 3 MIN READ

# How Do You Verify an AI Agent Actually Finished the Task?

An agent's own 'task complete' message is generated text, not a fact. This piece explains why self-reported completion fails, what independent checks look like, and walks through a fictional calendar-booking bot to show where verification has to sit outside the agent's own output.

By BotBento Editorial · Published 2026-09-26 · Updated 2026-09-26

AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. [Editorial policy](/editorial/).

## In this article

- [Why 'the agent said it worked' isn't verification](#why-self-report-fails)
- [What independent verification actually checks](#build-independent-checks)
- [A worked example: a fictional meeting-scheduling bot](#worked-example)
- [What to log so verification is possible later](#what-to-log-for-evidence)[Read as Markdown ](/text/blog/verify-ai-agent-task-completion/index.md)

## Key takeaways

- An agent's success message is generated the same way as the rest of its output — it is not evidence the work happened.
- Verification has to check the actual side effect (a file, a calendar entry, a sent email), not the agent's description of that side effect.
- Anthropic's own engineering notes describe agents marking features complete without testing them end-to-end — this is a known, named failure mode, not a rare bug.

## Why 'the agent said it worked' isn't verification

An agent's completion message is a claim about the task, not an independent observation of the result. Even when it summarizes successful tool calls, the message may omit an error, a delayed failure or a missing end-to-end check. Treat 'done' as a signal to inspect evidence rather than as the evidence itself.

Anthropic's engineering write-up on long-running agent harnesses describes a concrete version of this failure: an agent marked a feature complete after making code changes and running limited checks, even though the feature did not work end-to-end. A passing unit test or a successful request to a development server did not establish that a person could use the finished feature.

In that experiment, explicit end-to-end testing with browser automation helped the agent find and fix problems that were not obvious from the code alone. The lesson is to choose an observation that matches the user's requested outcome. A prompt asking the agent to be careful cannot replace an actual check of the result.

Sources: [Effective harnesses for long-running agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents).

## What independent verification actually checks

Independent verification means checking evidence beyond the agent's narration. In coding tasks, an automated test can exercise behavior in the resulting program and report a pass or failure. Anthropic points to this verifiability as one reason coding tasks can suit agents: the agent can use test results as feedback and revise its work. Tests still have limits, so the check must cover the behavior the user actually requested.

Outside coding, define a check for each side effect. Read the calendar API to confirm the intended event and attendees; inspect the database row for the expected value; inspect a mail provider's status for the exact message. A provider's 'accepted' response means it took the request, not that a recipient received the message. Even a 'delivered' event generally indicates acceptance by a receiving server, not proof that a person saw it in an inbox.

Anthropic's agent-evaluation guidance separates an agent's work from the grading of its result. For a coding evaluation, the agent changes an environment and tests then grade the working result. The same separation is useful in an operational workflow: define the expected state first, let the agent act, then inspect that state with a read-only check whose output can be reviewed.

Sources: [Effective harnesses for long-running agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents), [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents), [Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents).

## A worked example: a fictional meeting-scheduling bot

Say a fictional bot, 'Scheduler,' is asked to book a 3pm meeting and send a confirmation email. Scheduler calls a calendar tool, calls an email tool, and returns: 'Done — meeting booked and confirmation sent.'

That sentence does not establish whether either tool call achieved the requested outcome. The calendar API could have returned a conflict that Scheduler misread as success. The email tool could have accepted the request but later reported a bounce. Neither problem is resolved by the bot repeating its own summary.

A verification step for this fictional example would read the calendar API for an event at 3pm with the expected attendees and use the email provider's message ID to inspect its status. A provider delivery event supports a narrower claim than 'the recipient read the email'; a bounce means that part of the task failed. If the calendar entry is absent or the mail status is unresolved, the run should remain incomplete or pending. This illustrates a check design, not the behavior of a real scheduling product.

## What to log so verification is possible later

Verification only works if there's something to check against. That means logging the tool call's actual arguments and the raw response from the external system — not just the agent's summary of that response. If a calendar tool returns an error code, the log needs the error code, not the agent's paraphrase of it.

Anthropic's guidance on building tools for agents recommends recording more than a top-level success score: tool errors, call and task duration, number of calls and token use can reveal what happened during a run. Keep the verification result and the external system's identifier alongside those operational details. That lets a reviewer distinguish an unchecked claim from a check that found a real problem.

BotBento is pre-release and does not currently offer a built completion-verification feature. The team is exploring how a run record might separate the agent's self-reported status from an independently-checked one, but nothing in that direction is available yet — the practical advice above holds regardless of what tooling a bot builder uses.

Sources: [Writing effective tools for AI agents—using AI agents](https://www.anthropic.com/engineering/writing-tools-for-agents).

## Primary sources

Sources checked 2026-09-26. Standards and product documentation can change; follow the linked version when implementing.

- [Effective harnesses for long-running agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents) — Anthropic
- [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) — Anthropic
- [Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) — Anthropic
- [Writing effective tools for AI agents—using AI agents](https://www.anthropic.com/engineering/writing-tools-for-agents) — Anthropic

BotBento is in development. [Suggest a correction](/contact/).
