AI Academy · Book
Executives & Directors · Module 05 · Chapter 014

AI Agents and Intelligent Workflows

An agent's headline success rate hides the number that decides its value: how often it is wrong without telling anyone. Choose workflows an agent can verify, write it a mandate, design its handoffs, and score it on three outcomes, not one. Module 05's lesson holds here too: value comes from the work redesigned, not the tool added.

≈ 13 min read

After this chapter you can

  • Explain why an agent's silent-error rate matters more than its completion rate, using the cost of each outcome.
  • Apply five questions to decide whether a workflow suits an agent or ordinary automation.
  • Ask for repeat-run reliability rather than single-try success when an agent is evaluated.
  • Specify a good handoff and a five-field agent mandate with a named business owner.
  • Summarize Module 05's lesson and connect it to the risk questions of Module 06.

Take an illustrative case. Suppose your operations team pilots two AI agents on the same stream of routine requests for a month. Agent A completes nine requests in ten correctly, on its own. Agent B completes only seven in ten on its own and hands the other three to a person, with a note saying what it checked and why it stopped. Asked which one to scale, many leaders pick Agent A. The higher number looks like the better agent.

It may be the worse one. The comparison turns on a detail the headline number hides: what Agent A does on its tenth request. If it acts anyway, and the mistake surfaces weeks later in a complaint, a write-off or a rework queue, then its tenth case is a silent error. Agent B’s three cases are handoffs. A person sees each one on the day, with the context attached. Which agent is better depends on what each of those costs, and in many enterprise workflows a silent error costs far more than a handoff.

The questions that matter are about agents at work inside the enterprise: where they fit, how to delegate to them, and how to judge them. How an agent works as a technology, its loop of planning, tool use and checking, belongs to From AI Assistants to AI Agents in Module 02, and the shift from answering to acting in your systems of record belongs to From Copilots to AI Agents in Module 01. Here the question is managerial. You are about to give a piece of software a job.

Three outcomes, not one success rate

Every task an agent attempts ends in one of three ways. It is done and verified: the agent acted, and a check confirmed the result. It is handed off: the agent stopped and passed the case to a person. Or it is done wrong and unnoticed: the agent acted, the result was wrong, and nobody knows yet.

Every agent task ends one of three ways - done and verified, handed off, or wrong and unnoticed - and a success rate hides the third.Done and verifiedThe agent acted and a checkconfirmed itHanded offA person takes it today withthe contextWrong and unnoticedFound weeks later at a higher price
Figure 5.14.1 A single success rate merges the first and third outcomes. Ask for all three.

A single success rate merges the first and the third, because a wrong action and a right one both count as “completed” until someone notices. That is why the first question to ask about any agent is not how often it finishes, but how often it is wrong without saying so, and how you would know. The third number cannot come from the agent’s own log. By definition, nobody has seen those errors. It has to come from sampling: a person regularly checking a random set of completed tasks against what should have happened.

Illustrative arithmetic makes the point concrete. Take the two agents from the opening and give them 10,000 requests a month. Agent A acts on all of them: 9,000 are right and 1,000 are wrong without anyone noticing. Agent B acts on 7,000 and hands off 3,000, and it is not perfect either: suppose 100 of the 7,000 it handles are wrong. Now suppose a silent error costs 300 to find and fix later, in staff time, refunds and goodwill, and a handoff costs 20 of staff time on the day. Agent A’s 1,000 silent errors cost 300,000 a month. Agent B’s 100 silent errors cost 30,000 and its 3,000 handoffs 60,000, or 90,000 a month in all.

Illustrative monthly cost - Agent A's 1,000 silent errors cost 300,000, while Agent B's 100 silent errors and 3,000 handoffs cost 90,000.Silent errorsHandoffs to staffAgent A (9 in 10)300 k300 kAgent B (7 in 10)30 k60 k90 kILLUSTRATIVE NUMBERS
Figure 5.14.2 Illustrative. The agent with the lower completion rate costs less a month because its misses are visible.

The claim needs this framing to be honest. A lower completion rate is not a virtue in itself. Agent B pays 60,000 in handoffs to avoid 900 silent errors, so it comes out ahead only when a silent error costs more than 60,000 divided by 900, about 67. Below that price, Agent A would win as it stands. And if Agent A could hand off its own tenth case instead of acting on it, its 1,000 handoffs would cost 20,000 and it would beat Agent B easily. Agent B wins here because it knows where its competence ends and stops there. Accuracy, Hallucination and Reliability, in Module 06, explains why a system that can decline to answer is more trustworthy. The managerial version is simpler: put a price on a silent error before you compare agents, and the right choice often becomes clear.

Give agents the workflows that suit them

Agents are being tried widely, and mostly not yet at scale. McKinsey’s 2025 global survey, as reported in coverage of it, found 23 percent of respondents saying their organizations were scaling an agentic AI system somewhere, and another 39 percent experimenting. In any single business function, no more than 10 percent reportedly said they were scaling agents1. Gartner forecasts that over 40 percent of agentic AI projects will be canceled by the end of 2027, citing rising costs, unclear business value or inadequate risk controls; that is a forecast, not a measurement2. One common pattern behind both numbers, and the one a leader can most directly control, is an agent applied to a workflow it does not suit.

Five questions sort the candidates.

Five questions decide whether a workflow suits an agent - varied path, tool access, checkable results, contained actions and a team that owns it.QuestionIf yesIf noDoes the path vary from caseto case?An agent may add valueUse rules or ordinary automationCan it reach the systemsthrough approved tools?ProceedFix the integration firstCan a system check the result?The agent can verify itselfKeep a person approvingAre the actions containedor reversible?The agent may actThe agent recommends onlyDoes a team want it and own it?StartWait
Figure 5.14.3 An agent earns its place where the path varies, the result can be checked and the actions can be contained.

The first question often removes the most candidates. If the path is fixed, an agent is the expensive way to do it. Anthropic’s engineering guidance on building agents draws the same line between workflows, where models and tools follow predefined code paths, and agents, which direct their own process. It advises “finding the simplest solution possible, and only increasing complexity when needed”, and warns that agentic systems “often trade latency and cost for better task performance”3. A request that always takes the same five steps should be automated with rules. The agent belongs where the request arrives in messy language and the next step depends on what the last one found.

The third question decides how much the agent can be trusted. If a system can confirm the result, such as a booking that shows as accepted or a record that matches its source, the agent can check its own work before it reports success. If only a person can judge the result, a person has to approve it.

The fifth question is easy to skip. A Stanford study of 1,500 workers across 104 occupations found them positive about automating 46 percent of the 844 tasks studied, most often because it would free time for higher-value work4. An agent dropped on a team that did not ask for it is likely to be worked around.

Consistency: the same task a thousand times

A demonstration runs a task once. A workflow runs it thousands of times a month, and an agent can solve a task on Monday and fail the identical task on Tuesday. When a leading model of 2024 was tested on simulated retail service tasks, it succeeded at the first attempt on 61 percent, but on fewer than a quarter of tasks in all eight attempts5. Models have improved a great deal since, so treat the figures as a snapshot. The method is what lasts. From AI Assistants to AI Agents showed how small error rates compound across the steps of one task. Repetition is the second multiplier, across the many runs of the same task. When a team or vendor reports a success rate, ask how often the agent succeeds when the same case is run again, on your own cases, not theirs.

Design the handoff, not only the action

If handoffs are what make Agent B worth scaling, their quality matters as much as their number. A handoff that says only “unable to complete” sends the person back to the start. A good one hands over the work done so far.

A good handoff carries the request, what the agent checked, why it stopped and a suggested next step, routed to a named person.The requestIn the requester'sown wordsCheckedSystems andrecordsconsultedWhy it stoppedThe rule or doubtthat appliedNext stepSuggested, withthe evidenceRouted to a named person, with a clock
Figure 5.14.4 A good handoff passes on the work already done, so the person starts where the agent stopped.

Two design choices follow. First, the person receiving handoffs needs the skill to resolve them. As agents take the easy cases, the human queue fills with the hard ones, and the people doing the work see fewer routine cases to stay in practice on. Lisanne Bainbridge called this an irony of automation four decades ago: the more a system automates, the more its remaining human tasks depend on skills that automation stops exercising6. Second, the agent’s own rules for stopping must be written, not left to the model’s sense of doubt. The rules name the topics, amounts and kinds of customer that always go to a person, whatever the agent’s confidence. Autonomous Workflows and AI-Native Organizations, in Module 10, takes the exception path up to the scale of the whole organization.

Write the agent a mandate

Delegating to an agent is delegation, and good delegation is written down. A mandate has five fields, and each one answers a question an auditor, a regulator or an angry customer will eventually ask.

An agent mandate names its outcome, the systems it may touch, what it may do alone, what it must hand off and who owns it.OutcomeWhat done looks like, and how itis checkedSystemsWhat it may read and write, andnothing moreActs aloneActions, amounts and limits it maytake on its ownHands offTopics, amounts and cases thatalways go to a personOwnerA named business owner whoanswers for its results
Figure 5.14.5 Five fields turn “use AI here” into a job someone can audit.

The outcome must be checkable: “resolve supplier invoice queries” is a hope, while “every query answered with the matching purchase order and a confirmed payment date” is a goal an agent can be held to. Systems follows least privilege: the agent reaches only what the workflow needs. Acts alone and hands off together set its authority. How far up the scale that authority goes, from recommending to operating on its own, is the subject of Agentic AI and Autonomous Actions in Module 06, which also covers how authority is widened on evidence. The owner is the field most often left blank. It should be the business leader whose results the agent affects, not the team that built it.

Story: the fraud system that never handed off

One of the clearest public records of software given a job and the authority to finish it comes from before generative AI. In 2013, Michigan’s Unemployment Insurance Agency switched on a new system, MiDAS, built to detect benefit fraud automatically. In tens of thousands of cases it decided fraud on its own, with no person involved. Its fraud notices, IEEE Spectrum reported, were designed in a way that almost ensured some claimants would admit to fraud by mistake, and some findings rested on missing or corrupt data7. A fraud finding meant repaying the benefits plus a penalty of four times the amount, collected if necessary by seizing tax refunds and garnishing wages8. MiDAS was rule-based software, not an AI agent. That makes it a cleaner lesson, because nobody can blame a model: what failed was the mandate.

Michigan's MiDAS raised fraud findings about fivefold, but a later review found 85 percent of the findings it made alone were wrong, against 44 percent where a person was involved.Judged by its outputFraud findings up about fivefoldFund grew from about 3 to 69 millionEvery case closed; none handed offWhat a later review found85% of system-only findings wrong44% wrong with some human involvement20 million settlement in 2024
Figure 5.14.6 Every case ended as done. Most of the cases the system decided alone were wrong, and nobody saw it for years.

Judged by what it produced, the system was a success. Fraud findings rose about fivefold compared with the old system, and the fund that received the penalties reportedly grew from about 3 million to more than 69 million in little more than a year7. Every case it touched counted as finished. None was handed off.

The errors were silent until the people affected forced them into view through appeals and lawsuits. In September 2015, after a federal lawsuit was filed, the agency stopped letting the system decide fraud on its own. A later review, as IEEE Spectrum reported, found that of 40,195 cases the system had decided alone, 85 percent, or about 34,000, were wrong. Of a separate set of 22,589 cases with some human involvement, 44 percent were wrong7. In January 2024, a Michigan court gave final approval to a 20 million settlement for more than 3,000 people falsely accused between 2013 and 20159. Australia’s Robodebt scheme, told in AI vs Automation, followed the same pattern.

Read the case as a missing mandate. The outcome, “find fraud”, was counted but never checked. The system could act alone on everything, including penalties, and had to hand off nothing. Nobody sampled its finished cases until the courts did. And the 44 percent is a warning of its own: a person somewhere in the process is not the same as a designed handoff, with the evidence, to someone who has the time and skill to decide. Price the silent errors, and the fivefold rise in findings stops looking like success. An AI agent can fail the same way, faster. The defenses are the ones in this chapter: a written mandate, stopping rules, handoffs with context and a weekly sample of completed work.

What this means for leaders

Treat an agent as a new role in a workflow, not a feature in an application. That means choosing the workflow with care, writing down its authority, designing how it hands work back, and judging it on all three outcomes. The leadership decisions are the same ones you make when you delegate to a person, made explicitly because software will not ask when it is unsure unless you tell it to.

Agents are the strongest form of the lesson that runs through this module: value comes from AI built into real work and measured by outcomes. They do not only inform the work; they do it. That is why the controls named in this chapter, from handoff rules to audits, matter more here than anywhere else in the module.

Check yourself

  1. The agent with the higher completion rate is the better one to scale.
  2. An agent’s own logs can tell you its silent-error rate.
  3. If a request always follows the same steps, rules or ordinary automation are usually the better choice.
  4. An agent that succeeds on a task once will succeed on it every time.
  5. A good handoff includes what the agent checked and why it stopped.
  6. The team that built the agent should own its results.

Reflection: price the tenth case

What comes next

This module looked at where AI creates value; every chapter also touched on what happens when it is wrong. The next module makes that the subject. The AI Risk Landscape opens Module 06 by mapping where AI risk comes from and why it depends on how a system is used, not only on how accurate it is.

References

  1. McKinsey & Company (QuantumBlack). The state of AI in 2025: Agents, innovation, and transformation. McKinsey & Company. 2025.
  2. Gartner. Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. Gartner Newsroom. 2025.
  3. Erik Schluntz and Barry Zhang. Building effective agents. Anthropic Engineering. 2024.
  4. Yijia Shao, Humishka Zope, Yucheng Jiang, Jiaxin Pei, David Nguyen, Erik Brynjolfsson and Diyi Yang. Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce. arXiv (2506.06576), Stanford University. 2025.
  5. Shunyu Yao, Noah Shinn, Pedram Razavi and Karthik Narasimhan. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv (2406.12045). 2024.
  6. Lisanne Bainbridge. Ironies of Automation. Automatica 19(6). 1983.
  7. Robert N. Charette. Michigan's MiDAS Unemployment System: Algorithm Alchemy That Created Lead, Not Gold. IEEE Spectrum. 2018.
  8. Julia Angwin. The Seven-Year Struggle to Hold an Out-of-Control Algorithm to Account. The Markup (Hello World newsletter). 2022.
  9. Michigan Department of Attorney General. Class Action Settlement Approved by Court of Claims in Bauserman v. State of Michigan Unemployment Insurance Agency. State of Michigan. 2024.

Further reading

Sources last verified 2026-10-10.