AI agents vs workflows: choose who decides the next step
Use a workflow when you know the sequence. Give an agent discretion when discovering the sequence is part of the work. Then test the result, the cost and the boundaries separately.
AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. Editorial policy.
Key takeaways
- An agent chooses actions during a run; a workflow follows a sequence or routing rules you designed.
- A useful starting point is a fixed outer workflow with a small, bounded agent task inside it.
- Evaluate the delivered result and the permitted actions separately. A polished answer can still conceal an incomplete run.
Who decides what happens next?
An AI workflow is a sequence you arrange in advance: receive an input, classify it, retrieve relevant records, draft a response, and pass the draft to a reviewer. The model may perform one or several steps, but your program defines the permitted route. An AI agent has more discretion over the route. It can inspect a result, decide which tool would help next, and continue until a stopping condition is reached.
Anthropic’s engineering guide uses this distinction between predefined paths and model-directed processes. The practical question is how much choice a particular task needs. The term agent alone tells you little about the quality, reliability or permissions of a system. A workflow can involve sophisticated reasoning, and an agent can spend its time making quite ordinary decisions.
Sources: Building effective agents.
Example: a weekly research note
Consider an illustrative task for a small product team: prepare a Friday note about changes to three services the team depends on. The finished note should contain the change, its source, its relevance to the team, and a proposed next action. It should remain a draft until someone checks the recommendation. This example is a design exercise, not a claim about a deployed BotBento routine.
A workflow could fetch three known release-note pages, compare each page with the previous snapshot, extract new entries, ask a model to summarize them, and save the result. That is a reasonable design when the sources are stable and the question is narrow. An empty change set is also a valid result. The workflow should explicitly say that it found no changes rather than asking the model to invent something interesting.
An agent becomes more useful when a release note introduces an unfamiliar dependency or links to a migration guide. The next useful source may depend on the first finding. The agent could follow the migration link, look up the affected setting, and collect evidence for a proposed action. Give it a specific research question and allowed sources. Do not turn a request for a note into unrestricted permission to change production settings.
Keep the outer process predictable
For this research note, our suggested starting design is a fixed outer process with a bounded investigation inside it. The outer workflow owns the schedule, the list of approved services, the destination draft and the review requirement. The agent owns one question: does this particular change require attention, based on the available evidence?
Before each investigation, record the source URL and what has already been learned. Set a limit on the number of pages the agent can inspect and the time the investigation can take. Those limits are design choices for your budget, not universal values. If the limit is reached, preserve the useful evidence and label the question unresolved. A partial, accurately described finding is easier to review than a confident conclusion assembled after the budget has run out.
Keep collection and publication separate. The process can save a draft without having permission to send it to customers or alter a live account. If the agent recommends changing a setting, it should point to the relevant documentation and explain the expected effect. A different, explicitly authorized step can carry out the change after the recommendation has been checked.
Three questions before you add autonomy
The following questions are a practical decision aid we use for this example. They are not a benchmark or a maturity score. Answer them for the task you actually want to complete, rather than for the most impressive demonstration you can imagine.
- Is the next step known before the run starts? If it is, write that step into the workflow. For example, converting an approved document into two predetermined formats does not need a bot to invent a plan.
- Does evidence change what should be investigated? If it does, consider an agent for that investigation. For example, a confusing migration note may need a follow-up search that cannot be selected until the note has been read.
- Can you recognize a good stopping point? If you cannot, write the acceptance criteria first. For the Friday note, every included change needs a source and relevance explanation; an unresolved change needs an explicit unknown, not a guessed recommendation.
Test the result and the path
Anthropic’s evaluation guide distinguishes the record of a run from the outcome it produces and discusses combining different kinds of grading. That distinction is useful here: reading the final note is necessary, but it cannot by itself show whether the underlying sources were checked or whether the process performed an unauthorized action.
Build a small example set before you automate the weekly job. Include a week with no changes, a broken source page, two announcements that describe the same change, an announcement with an ambiguous date, and a migration guide that disagrees with an older overview. Keep the expected behavior beside each example. The expected answer can be “needs a person to resolve this contradiction.”
Inspect the saved note and the run record together. Did each source open successfully? Does a claim point to the exact page supporting it? Were failed sources reported? Was the output saved in the intended place? Were tool calls limited to the allowed actions? Compare the same cases using the simpler workflow before adopting a more autonomous approach. More steps are only helpful if they improve a result that matters to your team.
Sources: Demystifying evals for AI agents.
A useful first version
Start with one service and one draft destination. Write down the input, the permitted sources, the shape of the output and the conditions that require review. Run it manually on several contrasting examples. Keep the draft visible alongside its evidence so that a reviewer can check a claim without reconstructing the whole investigation.
Once the basic process is dependable, decide which repeated decisions are wasting time. A conditional branch may solve the problem. A narrowly scoped agent may solve it better. Add the smallest amount of discretion that helps, then rerun your examples. The goal is a research note someone can use every Friday, including the uneventful and awkward Fridays.
Primary sources
Sources checked 2026-09-07. Standards and product documentation can change; follow the linked version when implementing.
- Building effective agents — Anthropic
- Demystifying evals for AI agents — Anthropic
BotBento is in development. Suggest a correction.