Assistant, Workflow, or Agent: A Practical Way to Choose AI Tools
Most AI tool decisions are made backwards. Someone sees a product, likes it, and then looks for work it could do. The result is a subscription, a brief period of enthusiasm, and a tool that quietly stops being opened.
The alternative is not more careful product research. It is deciding, before looking at any product, what kind of help a given task actually needs — because the AI tool market currently contains at least three distinct categories of thing, and they fail in different ways when applied to the wrong work.
Three Categories, Not One
The labels vendors use are not reliable; nearly everything is described as an assistant, a copilot, or an agent regardless of what it does. The useful distinction is about where the work sits.
Assistants: you do the task, the tool helps
You remain in control of the sequence. The tool responds to what you ask, produces something, and you decide what happens next. Drafting help, summarization, code completion, and search fall here. The defining feature is that a bad output costs you the time to notice and redo it, and nothing else.
Workflows: the tool does a defined task, you approve
A bounded sequence runs without you, then hands back a result for review. Document processing, structured extraction, classification, routine transformation. The tool has autonomy over steps but not over goals, and there is a checkpoint before anything becomes consequential.
Agents: the tool decides the sequence
Given an objective, the system determines its own steps, uses tools, and proceeds without step-level supervision. This is where the market’s attention has been and where deployment has been hardest, for reasons we examined in our discussion of why long tasks break down even when each step is competent.
These are not points on a quality scale. They are different instruments, and the third is not an upgraded version of the first.
The Question That Sorts Tasks
The most useful diagnostic we know is this: what does it cost when this goes wrong, and how quickly would you notice?
Cheap and immediately visible failure — a badly drafted paragraph, an unhelpful summary — points toward assistants. You are the check, and being the check costs you nothing extra because you were reading the output anyway.
Expensive or delayed failure points toward workflows with explicit review, or away from automation entirely. If a wrong answer propagates into a decision, a customer communication, or a record that other work depends on, the review step is not overhead. It is the reason the system is safe to use.
The mistake we see most often is applying an agent to work with expensive, slow-to-surface failure modes because the demonstration was impressive. The demonstration ran on a clean task with a visible result. Production work frequently has neither property.
A Second Question: Is the Task Actually Stable?
Automation of any kind assumes the task looks roughly the same each time. A surprising number of tasks that feel routine are not.
A useful test: write down the steps as if instructing a new colleague. If the instructions require more than a handful of “unless” clauses, the task carries judgment that the person doing it may not have noticed they were exercising. Automating it will work until it meets one of the exceptions, which it will.
Tasks with genuine variability are candidates for assistance, not for workflows. The tool helps the person handle each case; it does not handle cases on its own.
What to Evaluate, Once the Category Is Settled
Only after the category is decided does product comparison become useful, and even then the relevant criteria are narrower than most reviews suggest.
Failure visibility. When the tool is wrong, is it obviously wrong? A tool that produces confident, plausible errors is more dangerous than one that fails loudly, and this is rarely mentioned in marketing material.
Correction cost. How long does fixing a bad output take relative to doing the task yourself? Tools that are fast when right and expensive when wrong can be net negative at surprisingly high accuracy rates.
Integration reality. Whether the tool sees the data it needs without manual staging. A tool requiring you to assemble its inputs has moved the work rather than removed it.
Exit cost. What happens to your data and process if you stop. Workflow tools in particular accumulate configuration that is difficult to reconstruct elsewhere.
Cost at real volume. Usage-based pricing behaves very differently at pilot scale and at full deployment.
Notably absent from this list: benchmark scores and model comparisons. At the tool level these are rarely the deciding factor, and the underlying models change often enough that a choice made on them is a choice made on a moving quantity.
Where the Market Actually Is
Our reading of the current landscape, offered as observation rather than prediction: assistants are mature and genuinely useful, and the productivity gains there are real if unspectacular. Workflow tools are the fastest-improving category and probably where most organizations get the best return right now. Agents are advancing quickly and remain difficult to deploy reliably outside narrow, well-bounded domains.
This is roughly the opposite of the attention distribution, which is worth keeping in mind when a product roadmap seems to be moving toward autonomy faster than the underlying reliability supports. We traced part of this trajectory in our look at the shift from plugins to agents, and the intervening months have mostly confirmed that the middle category was undersold.
A Practical Sequence
The approach we would suggest, in order:
List the tasks that consume time, before looking at any tool. Sort them by what a failure costs and how fast it would surface. Assign each to a category on that basis. Only then evaluate products, against the criteria above rather than against feature lists. Pilot on the tasks with the cheapest failure modes, because that is where you learn most safely. Then measure whether the time actually went down — which, in a striking number of cases, it did not.
The last step is the one most often skipped. Tools that feel faster are not always faster, particularly when the review burden they create is distributed across people who are not the ones evaluating the tool. This is a version of a broader difficulty that Executive Intelligence Society has explored in the context of what changes when machines begin exercising judgment: the cost of a delegated decision often lands somewhere other than where the delegation happened.
The Underlying Point
Tool selection reads as a technology question and is mostly a work-design question. What is the task, what does an error cost, who checks, and what is the fallback? Answer those and the product decision becomes narrow. Skip them and no amount of product research substitutes.
It also helps to be skeptical of the demonstration, which is engineered to be impressive and tells you little about the tool’s behavior on your least convenient input — a pattern we have written about before in the gap between interface quality and underlying capability.
If you are working through tool selection for a team and want to compare notes, I am on LinkedIn.

