From AI Assistants to AI Agents
An assistant answers a request; an agent pursues a goal by choosing steps, using tools and checking the results until it is done or stopped. Four checks tell the two apart, and the model is only one part of the system that passes them. Because small misses compound across steps, an agent is judged on the whole job, not on its best answer.
After this chapter you can
- Tell an AI assistant from an AI agent with four checks - goal, own next step, tools and feedback.
- Describe the agent loop, its state and the five conditions that should stop it.
- Explain why the model is only one layer of an agent and why exact rules belong in software.
- Show with arithmetic why small misses compound across steps, and ask for end-to-end results.
- Close Module 02 ready to ask where this capability creates business value.
Suppose, as an illustrative composite, that you run operations for a regional rail operator, and on Friday night a storm brings trees down across one line, so Saturday’s trains on that route are cancelled. You open a general-purpose AI assistant and ask it to write a notice telling passengers that the trains are not running. Seconds later you have a clear, polite draft. One request, one answer, and the rest of the work is still yours.
Now ask for something different: run replacement buses on Saturday. That is not one answer. Someone has to find coach companies with vehicles free, check which stations a coach can actually reach, confirm that the drivers’ hours are within the legal limits, load the new times into the journey planner, brief the staff at each station and tell passengers. Then someone has to read the replies, because one coach company will pull out on Friday evening, and start again on that part of the plan.
The first request is a writing task. The second is a job with a goal, several steps, several systems and a constant need to see whether each step worked. The difference between those two requests is the difference between an AI assistant and an AI agent.
Assistants answer; agents pursue goals
An AI assistant helps a person perform a task. It answers, drafts, summarizes and suggests, and the person decides what happens next. An AI agent pursues a goal through a sequence of actions, using tools and the results of those actions, until the goal is met or something stops it.
The idea is older than the products. A standard AI textbook has long defined an agent as anything that perceives its environment and acts on it1. A thermostat qualifies under that definition. What changed in the 2020s is that language models made the hard middle step, working out what to do next in an open-ended situation, general enough to use across ordinary business work.
Because the word is now attached to almost everything, it helps to have a test. Gartner has warned of “agent washing”, the relabeling of existing chatbots and automation as agents, and forecasts that over 40 percent of agentic AI projects will be canceled by the end of 20272. For now, four checks are enough.
Is it given an outcome rather than a single request? Does it choose its own next step? Does it act through tools that reach other systems? Does it read the result of each action and adjust? A product that passes all four is an agent, whatever its label. A product that fails any of them is an assistant, which may be exactly what you need.
Notice one question that is not on the list: how much the system may do without a person. That is a separate setting, not part of the definition. An agent can be confined to reading data and proposing a plan, or trusted to act on its own within limits. Agentic AI and Autonomous Actions in Module 06 sets out that autonomy ladder and how to choose a rung for each task. From Copilots to AI Agents in Module 01 gave the business test for the moment that matters most: when the AI, not a person, makes the change in the system of record.
The loop at the heart of every agent
However it is built, every agent runs some version of one loop. It starts from a goal and plans the next useful step. It acts, usually by calling a tool: a search, a database query, a booking system, a calculator. It observes the result. Then it decides whether to continue, change the plan, stop or ask a person, and goes round again.
The research that made this pattern popular is plain about why the observe step matters. In 2022, Shunyu Yao and colleagues at Princeton and Google described ReAct, short for reasoning and acting: the model writes out its reasoning, takes an action, reads what came back and reasons again. On two benchmarks of interactive tasks, a simulated household and an online shop, ReAct beat earlier methods by 34 and 10 percentage points in success rate. On question answering, consulting a simple Wikipedia lookup between steps reduced the invented facts and the errors that cascade when a model reasons from its own assumptions alone3. Looking at the world between steps is what keeps the plan honest.
Two ordinary details decide whether the loop works in a business. The first is state: a running record of where the job stands. In the rail example, the state says that two coach firms have confirmed, the stations have been briefed and the journey planner has not yet been updated. Without it, an agent repeats steps or skips them. The second is a set of stopping conditions. The loop must end when the goal is met, when a limit on steps, time or money is reached, when a person is needed, when a tool fails or when a policy blocks the action. A loop with no stopping rule is not autonomy. It is a system that does not know when to quit.
The model is not the agent
The model is the part you see in a demonstration, so it is easy to think it is the whole thing. It is not. The model interprets the goal, proposes the next step and makes sense of each result. On its own, it only produces text. The agent is the system built around it.
Around the model sit instructions, which say what the agent is for and what it must never do, and state, which records progress. Then come tools, the agent’s only way of touching anything outside itself, and guardrails: permissions, spending and step limits, approval points and a log of every action. Retrieval, taught in RAG and Enterprise Knowledge, becomes one tool among several: the agent looks something up when the plan needs it.
This anatomy explains a design pattern that many teams use. Let the model interpret, let software check, and let the model explain. When a coach company writes in to offer a driver, the model reads the message and extracts the offer. A rules engine checks the driver’s legal working hours exactly, the same way every time. The model then drafts a reply that explains the result. The probabilistic part handles language; the deterministic part handles the rule. Each does what it is good at.
Not every job needs a loop
Not every multi-step job needs an agent. Engineers at one model provider draw a useful line between workflows, in which models and tools follow steps fixed in advance by code, and agents, in which the model directs its own process and tool use. Their advice is to find the simplest solution that works and add complexity only when it is needed4. The operator’s ticket refunds follow the same path every time, so a fixed workflow is right, perhaps with a model reading free-text complaints at one step. Replacement buses are different: which systems to consult, and in what order, depends on what each coach firm, each station and the drivers’ hours allow. That is where the loop earns its cost. AI Agents and Intelligent Workflows in Module 05 turns this distinction into a test for choosing which workflows to give an agent.
The cost is real, because every turn of the loop is another model call, often with a longer context than the last. The same provider measured, by its own account, that its agents used about four times as many tokens as a chat interaction, and systems of several cooperating agents about fifteen times as many5. A single agent with a handful of tools is often enough. More agents add coordination, cost and failure points, and should follow a clear division of the work rather than an impressive diagram.
Small misses compound
An assistant is judged on one answer. An agent is judged on a chain of actions, and chains behave differently.
Think of a form with five fields, filled in by someone who gets each field right 95 percent of the time. The chance that the whole form is right is not 95 percent. It is 0.95 multiplied by itself five times: 0.95 to the fifth power is 0.774, or about 77 percent, so more than one form in five contains an error somewhere. Stretch it to ten fields and the figure falls to about 60 percent. Even at 99 percent per step, a twenty-step job comes out right only about 82 percent of the time.
The arithmetic assumes that any single miss spoils the job and that nothing catches it. Real systems can do better than that, and the whole point of good design is to make them do so. Shorten the chain where you can. Move exact rules out of the model and into software. Check each important result before the next step depends on it. Give the agent a way to recover, or to stop and ask. Benchmarks show why this matters: in the 2024 test that introduced a benchmark of agents serving simulated retail and airline customers under company policies, the best agent tested succeeded on fewer than half of the tasks, and in retail on fewer than a quarter when the same task had to succeed in each of eight repeated runs6. Later models have scored higher, but the gap between succeeding once and succeeding every time is the lesson that lasts.
The practical rule for a leader is simple. Ask for end-to-end results on real tasks: how often the whole job was done correctly, how many actions were wrong, how often it escalated and what each completed task cost. Accuracy on a single step is not the number that matters.
Story: from searching to supervising at Genentech
Genentech, the biotechnology company in the Roche group, offers a documented before-and-after of one organization making this shift. Its research scientists spend much of their time validating biomarkers, measurable signs in the body, such as a protein, that indicate a disease or a response to treatment. Before a team commits to a drug target, someone has to gather the evidence: what the published literature says, what public databases show and what the company’s own experimental data contain.
Before, that gathering was done by hand. Scientists sifted through scattered sources, including PubMed’s 38 million biomedical publications, public protein databases and internal repositories, and the company describes the work as taking weeks per question7. An assistant could help a scientist summarize a paper once it had been found. It could not do the searching.
After, Genentech built what it calls the gRED Research Agent. A scientist poses a research question. The agent breaks it into a multi-step plan, queries PubMed, the Human Protein Atlas and internal databases through their programming interfaces, adapts its approach as each result comes back, and returns a summary with citations. The company says tasks that took weeks now take minutes, and projects that the agent will help automate over 43,000 hours of manual effort in biomarker validation, a forecast rather than a measured result7.
Run the four checks and the system passes all of them. It is given a goal, chooses its own steps, acts through tools and adjusts to what it finds. Yet its autonomy is deliberately modest. It reads and assembles; it does not choose drug targets or change any record. Aviv Regev, who leads Genentech’s research organization, put the intended balance in one line: “Agents can’t replace our scientists, but they actually boost our scientists”7.
Two cautions keep the story in proportion. The figures come from an undated case study (2024 inferred) published by a technology vendor, and the 43,000 hours is the company’s projection, not an audited result. Nor does the case study report how often the agent’s summaries were incomplete or wrong, which is exactly the end-to-end question the previous section says to ask. The design lesson stands on its own, though: a large gain came from giving the agent a narrow job with read-only tools and keeping the decision with the expert.
What this means for leaders
The word agent will keep appearing in proposals, vendor pitches and board papers. The four checks let you sort real agents from relabeled assistants in a minute. Behind the label, what matters is the job the agent is given, the tools it can reach, the checks around it and the evidence that it completes the whole job. The aim is never to make an agent as autonomous as possible. It is to give it useful autonomy inside clearly defined boundaries.
Keep capability and outcome apart. Published evidence of agents’ business results is still thin, and much of it, like the Genentech case, is reported by the company or its vendor. Until a pilot in your own process measures the whole job, treat what an agent can do as a capability, not a result.
Check yourself
- A system that drafts excellent replies to customer emails is an AI agent.
- The model is the agent; everything else is packaging.
- An agent that is right 95 percent of the time at each step completes a five-step job about 77 percent of the time if nothing catches its mistakes.
- Every multi-step process should be rebuilt as an agent.
- An agent can pass all four checks and still be limited to reading data and proposing a plan.
- Letting the model check exact business rules itself is as reliable as using a rules engine.
Reflection: run the four checks
What comes next
You now have a sound mental model of modern AI, from what it is to how a loop of planning, tools and checks lets it act. Module 03, Business Value of AI, asks the question this module has set up: what would count as value from this capability, and how do you tell a convincing demonstration from a real outcome? It opens with What Does AI Value Actually Mean?, which defines what counts as value before anyone tries to measure it.
References
- Stuart Russell and Peter Norvig. Artificial Intelligence: A Modern Approach (4th edition). Pearson. 2020.
- Gartner. Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. Gartner Newsroom. 2025.
- Shunyu Yao et al. ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023 (arXiv:2210.03629). 2023.
- Erik Schluntz and Barry Zhang. Building effective agents. Anthropic Engineering. 2024.
- Anthropic. How we built our multi-agent research system. Anthropic Engineering blog. 2025.
- Shunyu Yao, Noah Shinn, Pedram Razavi and Karthik Narasimhan. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv (2406.12045). 2024.
- Amazon Web Services. Genentech Leverages Generative AI to Get Lifesaving Medicines to Patients Faster. AWS customer case study. 2024.
Further reading
- Shunyu Yao et al. ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023 (arXiv:2210.03629). 2023.
- Erik Schluntz and Barry Zhang. Building effective agents. Anthropic Engineering. 2024.
- Amazon Web Services. Genentech Leverages Generative AI to Get Lifesaving Medicines to Patients Faster. AWS customer case study. 2024.
Sources last verified 2026-10-08.