FIELD NOTES / 5 MIN READ

What should an AI agent record after each run?

An AI agent run record should connect the requested outcome to what was actually verified. Record the run identity, task scope, outcome evidence, unresolved actions and the next safe step. Keep operational detail separate from the short result a person reads.

AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. Editorial policy.

Key takeaways

  • A completed tool call is evidence about one operation; the routine still needs an explicit check of the user's requested outcome.
  • Distinguish incomplete work from actions whose effects are unknown before deciding whether to retry.
  • Store references and useful diagnostics without copying credentials, private messages or unnecessary personal data into the run history.

Start with the result someone needs to check

Imagine a morning routine that reads three approved project sources, writes a short research note and posts a link to a team channel. A message saying the run succeeded leaves several questions unanswered. Were all three sources available? Where is the saved note? Did the channel receive the link, or did the routine only prepare a message? Those distinctions determine what a teammate should do next.

Our proposed result record starts with the task's acceptance conditions. For this example, success means that the three sources were checked, the note was saved in the intended workspace and the channel post was confirmed. A readable draft with one unavailable source is still useful, but it does not meet that complete definition. Record the useful output and the missing condition together.

This is a design proposal for a routine you control. It is not a claim that BotBento currently ships this record format, and the example below is illustrative rather than a measured production run.

Give the run and its attempts distinct identities

OpenTelemetry describes a trace as connected operations and a span as a unit of work with timing and identifying context. Child spans can represent sub-operations. These concepts help explain which operation happened inside a larger run, but a trace does not define your business acceptance conditions for you.

For a small routine, keep a stable logical run ID and identify individual attempts separately. Suppose the scheduled research note is run research-2026-09-08-am. A second attempt to retrieve a source should remain associated with that run. It should not look like a second successful morning publication in your daily count.

Record the routine version and the intended destination alongside the run ID. Use internal references where appropriate: a workspace identifier is usually more useful for diagnosing a wrong destination than the bot's display name. If you already collect traces, link the run record to its trace ID instead of embedding every diagnostic event in the human-facing summary.

Sources: Traces.

An illustrative record for a partially completed routine

Here is a deliberately small example. The field names are our suggested application contract, not OpenTelemetry standard attributes. In a real system, references would point to access-controlled records that the reviewer is permitted to open. None of the identifiers below represent a real person, workspace or delivery.

The useful distinction is between the known output and the unresolved action. Someone can read the saved note immediately. They can also see why replaying the entire routine would be a poor recovery instruction: it could create a second note and potentially a second channel post.

  • Run: research-2026-09-08-am; routine version: 3; attempt: 1.
  • Requested scope: sources A, B and C; save one note; post its link to channel R.
  • Observed result: sources A and B checked; C unavailable; draft note N saved and read back.
  • External action: channel post request timed out; whether a post exists is unresolved.
  • Outcome: partial; acceptance conditions not met; note N contains a visible missing-source notice.
  • Next step: inspect the channel operation before retrying it; recover source C separately; owner: the routine operator.

Treat an unknown external effect as a recovery decision

A timeout alone does not tell the caller whether a remote action happened. AWS's Builders' Library explains how a caller-provided request identifier can let a service recognize repeated intent and handle retries without duplicating the operation. The protection depends on the service's actual idempotency contract; writing a unique ID into your own log is not enough.

For the illustrative channel post, first look for a provider operation identifier or another supported way to inspect the result. If the provider supports idempotent retries, follow its rules about identifiers, request contents and retention windows. Do not generate a fresh action identity just because the original response was missing.

Your record should distinguish confirmed failure from an unresolved effect. We suggest separate fields for the observed error and the recovery decision: for example, timeout and inspect before retry. This makes an unattended worker's stopping point understandable. When the provider offers neither safe inspection nor a supported retry contract, leave the action for review instead of presenting a guessed delivery result.

Sources: Making retries safe with idempotent APIs.

Keep the record useful without copying private content

OpenTelemetry's sensitive-data guidance recommends collecting only what serves an observability purpose and reviewing instrumentation for information it may expose. That applies to the convenient fields in an agent run history as much as to a tracing backend. Authentication credentials, session tokens and private message bodies should not become routine diagnostic attachments.

For our research example, store a reference to note N and a short error category for source C. Avoid copying the complete source response into every run record. Keep detailed evidence in its appropriate storage location, with suitable access and retention, and make the summary useful even when a reader cannot open that evidence.

Decide which record fields are allowed before collecting them. A small list of operational fields is easier to inspect than an unrestricted dump of tool arguments and responses. Review any linked artifact too: a carefully redacted summary can still expose private content through an unrestricted evidence link.

Sources: Handling sensitive data.

Review the record before putting the routine on a schedule

Walk through the example with three deliberately different outcomes: everything completes, a source is unavailable, and a write request returns no clear result. For each one, ask another person to identify the usable output, the unverified condition and the next action from the result record alone. If they need the original chat to understand the outcome, the record is missing context.

Then consider a second attempt. Does it preserve the relationship to the original run? Can a reviewer tell whether the note was revised or a new copy was created? Does the record keep an earlier unknown action visible until evidence resolves it? These questions test the usefulness of the design without pretending that a tidy status label establishes completion.

Start with a short record that answers those questions. Add detail when a real recovery need justifies it. The goal is a routine whose result can be checked and whose next step can be chosen without reconstructing the entire conversation.

Primary sources

Sources checked 2026-09-08. Standards and product documentation can change; follow the linked version when implementing.

  1. Traces — OpenTelemetry
  2. Making retries safe with idempotent APIs — Amazon Web Services
  3. Handling sensitive data — OpenTelemetry

BotBento is in development. Suggest a correction.

Keep reading