Back to News & insightsEngineering

When does your product need an AI agent?

Decide whether a fixed workflow or a model-directed loop fits the task, then define tools, stopping conditions, and recovery paths.

Editorial guide · Updated September 20, 2026 · 3 min read
Floating metal and glass layers connected to a central core, illustrating application architecture.

The word agent often hides an important design choice: who decides the next step? In a fixed workflow, application code determines the sequence. In a model-directed system, the model can choose actions based on intermediate results. That flexibility can be useful, but it creates more states to test and recover from.

Anthropic's discussion of effective agents makes a similar distinction between predefined workflows and dynamically directed processes. For a product team, the practical goal is to give the model enough discretion to solve the task without making every part of the application unpredictable.

Start with a task that has a clear finish

Consider a supplier comparison assistant. A user wants three available replacements for a discontinued component, with supporting specifications. The task may require different searches depending on what the first results contain, so a dynamic loop could be useful.

Define the finish line: three candidates that meet named constraints, or a report explaining which constraints could not be verified. “Keep researching until confident” is not an operational stopping rule. Set limits on elapsed time, tool calls, and spending, alongside the quality criteria.

Keep deterministic decisions in code

If the process always reads a form, validates fields, and saves a record, ordinary application logic may be sufficient. A model can help interpret a free-text field without selecting the entire workflow. Adding autonomy to a predictable sequence can create work without improving the result.

For the supplier assistant, code can check whether a price is numeric and whether a candidate includes a source. The model can help interpret descriptions. Separate those responsibilities so a change in model behavior cannot silently bypass a hard requirement.

Make tools small and observable

Give each tool a narrow purpose, documented inputs, and a clear result. A search tool should report matches or an explicit failure. A draft-saving tool should not also email the draft unless that action is part of its visible contract.

Log tool names, status, timing, and appropriate identifiers. Avoid storing unnecessary sensitive content. When the assistant produces an unsupported claim, you should be able to determine whether the tool returned no evidence, the model misread it, or the application lost part of the response.

Plan for partial completion

Suppose the assistant finds two candidates and the third supplier site becomes unavailable. Save the verified work and describe the missing part. Do not turn a network failure into an invented third result or discard everything and start an unbounded retry loop.

For state-changing operations, define how retries behave. A search can often be repeated safely; a purchase cannot simply be repeated after an uncertain timeout. Keep a durable operation identifier and inspect the prior result before issuing another action where duplication matters.

Evaluate the outcome and the path

Review whether the final result meets the task and whether the actions were appropriate. A correct answer reached through an unauthorized data source is still a failure. Include tests where tools return empty results, conflicting information, and content that tries to instruct the assistant.

Begin with a small set of permitted tools and expand only when a demonstrated need justifies it. Keep a manual completion path for unresolved cases. The strongest case for an agent is a measured improvement on a task that benefits from adaptation, with costs and failure behavior that the team can explain.

Further reading

Anthropic: Engineering research and guidance

An original editorial guide. Provider capabilities and documentation can change. Follow the linked sources and test the exact model or service before relying on it.