AI Academy · Book
Executives & Directors · Module 03 · Chapter 006

From AI Capability to Business Outcome

A demonstration shows what AI can do; it says nothing about what will change in the business. Between the two sits a chain of changes, and every arrow in it is an assumption until someone tests it. Leaders who name the use case, test the weakest arrow first, redesign the work and give the outcome one owner are the ones who turn capability into results.

≈ 17 min read

After this chapter you can

  • Distinguish a capability, a use case and a business outcome, and test a proposal for a real use case.
  • Write each arrow of the chain as an assumption, with the person who controls it.
  • Choose the weakest arrow and test it cheaply before building.
  • Explain why workflow redesign, not the tool, converts capability into outcome.
  • Write an outcome hypothesis with a workflow change, a target, a guardrail and a business owner.

Suppose you sit on the steering committee of a seed company, and the demonstration is good. A dealer forwards a large grower’s request for next season’s planting plan: a forty-page soil-test report, five years of yield maps and a list of fields. In under a minute the AI produces a two-page agronomy summary, flags two fields where last year’s disease pressure rules out the variety the grower asked for and drafts three questions for the dealer. The committee is impressed. The technology lead asks for budget to give the tool to every agronomist in the commercial division.

Then someone asks what it will be worth, and the honest answer is that nobody knows yet. What the committee has seen is a capability. Nothing on the screen says whether agronomists will trust the summary, whether planting plans will go out sooner, whether sooner plans win more orders, or whether the disease history would have been caught anyway. Those questions sit below the waterline, and they decide the value. The seed company is an illustration, and it runs through this chapter; its people and numbers are invented.

The first chapter of this module defined AI value as a measured change in business results, net of cost, and drew the chain that runs from capability to value. That definition is the standard. The work of leadership happens one level down, where a specific capability meets a specific piece of work and someone has to say what will change.

The core idea

An AI capability becomes valuable only when it causes a measurable change in a business outcome. It does that through a sequence of changes in how people work, and each change is an assumption until it has been shown to hold.

Two consequences follow. First, the quality of the capability is never the whole case, because a perfect capability can sit at the start of a chain that breaks two links later. Second, the most useful thing an executive can do with an AI proposal is not to admire the demonstration but to ask what must be true, link by link, for the claimed outcome to appear, and who will make each link hold.

Capability, use case, outcome

Three words are used as if they were interchangeable, and the confusion is expensive.

Capabilities such as summarize and predict are verbs about the machine; outcomes such as plan turnaround and replant claims are measures of the business.CapabilitySummarizeClassifyPredictExtractGenerateOutcomePlan turnaroundReplant claimsUnplanned downtimeCost per orderA use case connects them by placing a capability in a specific piece of work.
Figure 3.6.1 Capabilities are verbs about the machine. Outcomes are measures of the business. None of the words on the left appear on the right.

A capability is something the technology can do: summarize, classify, predict, extract, generate. A business outcome is something the organization measures and cares about: how long a planting plan takes, how much replant claims cost, how many hours a rolling mill stands idle without warning. Notice that the two lists share no words. The gap between them is where many business cases go quiet.

A use case is what bridges the gap. It is a capability applied to a specific task, for specific people, at a specific point in a workflow, with a decision or action that will now be different. “Use generative AI in agronomy” is not a use case; it names a technology and a department. “Give the agronomists who serve large growers an AI summary of each field file so that they triage it within the first hour instead of reading it in full on the third day” is a use case. It names the users, the step that changes and the change itself, and that gives you something to measure.

A useful test for any proposal is to strike out the name of the technology and see what is left. If what remains is a clear statement of who will work differently and how, there is a use case. If nothing remains, there is only a purchase.

One capability, several outcomes

The same capability can produce very different outcomes depending on where it lands, and sometimes it produces none.

One summarization capability can lead to faster plans, fewer replant claims or no outcome at all, depending on whose work changes.Field-file summaryAgronomists triage on day oneFaster plansField scouts pick farm visitsFewer replant claimsNobody changes their routineNo outcome
Figure 3.6.2 The capability is identical on every branch. The outcome depends on whose work changes.

Put the seed company’s summary in front of agronomists who use it to sort requests on the first day, and the outcome might be faster planting plans and more orders booked. Put it in front of field scouts who use the disease flags to choose which farms to visit, and the outcome might be fewer replant claims. Put it into a routine that nobody changes, and the summaries pile up while the business carries on as before. Each branch would need a different owner, a different measure and a different business case.

The outcome also depends on who uses the capability. In a study of more than 5,000 customer-support agents, an AI assistant raised productivity about 15 percent on average, but the gains were uneven, and the most experienced agents gained little1. As AI Is Changing Everything showed with the jagged frontier, the same system can help on one task and hurt on the next2. A capability does not carry its outcome with it.

Every arrow is an assumption

The chain from the first chapter of this module can now be drawn for a single use case. It inserts a use case after the capability and shows the behavior change as changed work; value, net of cost, still follows the outcome. The lesson is in the arrows.

A chain from capability to use case, adoption, workflow change and outcome, in which every arrow is an assumption to test.CapabilityWhat canthe AI do?Use caseWhere isit applied?AdoptionDo peoplerely on it?WorkflowDoes theworkchange?OutcomeWhatmeasuremoves?Every arrow is an assumption to test
Figure 3.6.3 Business cases tend to describe the first box in detail and the last box with optimism. The arrows decide the result.

Each arrow can be written as a sentence that begins “this only works if”. For the seed company: the summaries capture what an agronomist needs, so the capability fits the use case. Agronomists read the summary before the full file and act on it, so the use case becomes adoption. The review rules let a summary-first triage route standard requests to a fast lane, so adoption becomes changed work. And growers book more orders with suppliers who answer sooner, so changed work becomes an outcome. Any one of those sentences can be false.

Adoption is an arrow often mistaken for proof, and it is subtler than a login count suggests. In the same support-agent study, agents followed only 38 percent of the AI’s suggestions on average, and those who followed the most in their first month gained far more than those who followed the least1. The authors call this an association, not proof of cause, but the point stands: two people on the same team, with the same tool, can use it in materially different ways. “Everyone has access” says almost nothing about whether the next arrow holds.

There is a picture worth keeping here. A relay team can field four fast runners and still lose in the exchange zones, because the race is decided at the handoffs as much as on the straights. The arrows in the chain are exchange zones, and the baton usually passes between different people: from the data team to the users, from the users to the managers who set the routine, from the routine to customers the organization does not control.

Writing the arrows down is an old discipline with a new application. Rita Gunther McGrath and Ian MacMillan argued that ventures facing real uncertainty should be planned on explicit assumptions, tested at milestones, with money released as the assumptions are confirmed rather than committed at the start3. An AI use case is exactly such a venture. Its assumptions are its arrows.

Some arrows run outside the AI

Not every arrow is under the same control, and that changes how you manage it. Some arrows depend on the technology: accuracy, speed, reliability. Some depend on the organization: the routine, the rules, the incentives, the training. And some depend on people the organization does not control at all: customers, dealers, regulators, the weather.

For each arrow the register names what must be true, who controls it and the cheapest early check; the last arrow depends on growers and dealers.ArrowWhat must be trueWho controls itCheapest early checkCapability to use caseSummaries capture whatagronomists needAI teamScore 50 past field filesUse case to adoptionAgronomists read thesummary firstAgronomy team leadObserve a week of triageAdoption to workflowRules allow a fast laneHead of agronomyRun the rule by handWorkflow to outcomeGrowers reward faster plansGrowers and dealers, not usCompare past orders withresponse speed
Figure 3.6.4 An illustrative arrow register for the seed company’s use case. The arrow nobody inside controls is often the one to check first.

An arrow register like this one does three jobs. It turns a business case’s optimism into checkable statements. It shows who has to act for each arrow to hold, which is usually not the team building the AI. And it shows which arrows depend on outsiders, where the organization can influence the result but not command it.

The last row is the one to notice. If growers buy on price, trial yields and their dealer relationship, and barely notice when a plan arrives, then faster plans will not win more orders, however good the summary is. That assumption can be checked against the company’s own history of plans and orders, before any model is built.

Test the weakest arrow first

Many AI programs test the arrow they are best equipped to test, which is the first one. The technology team owns the capability, so the pilot measures accuracy, and the result is often encouraging. The arrows further down, where the case is more likely to break, wait until after the money has been spent.

The better order is to find the arrow most likely to fail and test it first, as cheaply as possible. Often that test needs no AI at all. If the doubt is whether agronomists will triage from a summary, have an analyst write summaries by hand for a few weeks and watch what agronomists do with them. If the doubt is whether growers reward speed, look at the past. If the doubt is whether the routine can absorb a new input, run the new routine on paper. A failed check at this stage costs a few weeks. The same failure discovered after rollout costs the whole investment and, often, the credibility of the next proposal.

Redesign the work, not only the tool

The arrow from adoption to changed work is where much AI value is won or lost, and it is the one that depends least on the technology. Many organizations add AI to an existing process and change nothing else.

Adding a summary to the old process frees no time, while redesigning triage around the summary creates a path to faster plans.AI ADDED TO THE OLD PATHSummary added; every filestill read in full.No time freedWORK REDESIGNEDSummary-first triage;standard requests take afast lane.Faster plansvs
Figure 3.6.5 The capability is the same on both sides. Only the redesigned process has a path to the outcome.

In the bolt-on version, the seed company gives agronomists the summary but leaves the review rules alone. The rules still require every file to be read in full, so agronomists read the summary and then the file. The work has grown slightly. In the redesigned version, the rules change. The summary and its flags decide the route: standard requests go to a fast lane with a target of a plan within a day, complex farms go straight to senior agronomists, and field scouts visit the farms the flags pick out. A person still signs every recommendation. The capability is identical in both versions; only one has a path to an outcome.

The survey evidence points the same way. In McKinsey’s March 2025 global survey, the redesign of workflows had the biggest effect on whether organizations saw an impact on earnings from generative AI, out of 25 attributes tested. Yet only 21 percent of respondents whose organizations used generative AI said they had fundamentally redesigned even some of their workflows4.

Workflow redesign had the biggest effect on earnings impact of 25 attributes, yet only 21 percent of generative AI users had redesigned workflows.1 of 25Workflow redesignBiggest effect on earnings impact of theattributes tested21%Have redesigned workflowsShare of generative AI users, at least some workflowsSource: McKinsey, The state of AI · March 2025
Figure 3.6.6 The practice most linked to earnings impact is the one most organizations have not yet done.

Boston Consulting Group’s 10-20-70 rule, taught in The AI Transformation Challenge, points the same way: AI leaders put most of their effort into people and processes, not algorithms5. Both are surveys of self-reported practice, so treat them as a signal rather than a recipe. The direction is consistent with the long history of the electric motor told in AI Is Changing Everything: the gains came when the work was rebuilt around the technology, not when the technology arrived.

Write the outcome hypothesis, and name its owner

The first chapter of this module introduced the value hypothesis: if this capability, for this workflow, then this outcome improves, by this much, while a guardrail holds. Its by becomes the “how much, by when” of the outcome. For a single use case, add two things that the chain has shown to matter: the change in the work, stated explicitly, and the person who owns the result.

An outcome hypothesis names the people, the capability, the step of work that changes, the outcome and target, the guardrail and the owner.ForWhich people or processUsingWhich capabilityTo changeWhich step of the workSo thatWhich outcome, how much, by whenWhileWhich guardrail holdsOwned byWhich business leader
Figure 3.6.7 An outcome hypothesis is the value hypothesis plus the change in the work and a named owner.

Filled in for the seed company, with illustrative numbers, it reads like this: for the agronomists who serve large growers, using AI field-file summaries, to change triage from full reading to summary-first routing, so that median turnaround for standard planting plans falls from five working days to two within two quarters, while replant claims on those fields do not rise, owned by the head of agronomy services. Every part of that sentence can be checked, and every arrow in the register now has a home.

When an initiative has several goals, rank them: one primary outcome, a few secondary ones, and the guardrails that must not get worse. Leading vs Lagging AI Metrics, later in this module, develops that structure.

The owner deserves a moment. The technology team can own the capability arrow, because it controls accuracy, reliability and running cost. It cannot own the outcome, because it does not control the routine, the rules or the people. The outcome belongs with the business leader who controls the arrows that are most likely to break. AI Operating Model, in Module 4, sets out how that split works across an organization. For a single use case the rule is simple: if no business leader is willing to sign the hypothesis, the arrows below adoption have no one to make them hold.

Story: the accurate model and the unchanged stoppages

Predicting which equipment will fail is an established use of machine learning in asset-heavy industry. Pioneering research on New York City’s power grid described the models’ purpose as helping the power company prioritize its maintenance and repair work6. Prioritize is a verb about the work, not the model, and that is where this story turns.

A steelmaker wanted fewer unplanned stoppages at its rolling mills. Its data team built a model that ranked the mills’ motors, gearboxes and pumps by their risk of failing within ninety days. The back-test was excellent, and the steering committee was delighted. By the third month the rankings were live on a dashboard for maintenance planners, and usage reports looked healthy.

An accurate failure-prediction model went live and was used, but crews kept the old calendar and unplanned stoppages were unchanged after a year.Month 1Back-test looksexcellentCommittee delightedMonth 3Dashboard goesliveUsage looks healthyMonth 8Crews follow theold calendarRosters set sixweeks aheadMonth 12StoppagesunchangedWhere is the value?
Figure 3.6.8 In this illustrative composite, every report before month twelve measured the capability or its use. None measured the outcome.

In the eighth month a maintenance engineer pointed out something awkward. Crew rosters were still built six weeks ahead from the fixed inspection calendar. The dashboard was something planners looked at, not something they planned from. At twelve months, unplanned downtime hours were where they had started, and the board asked the obvious question: the model is accurate, so where is the value?

Nobody could answer, because nobody owned the outcome. The data team owned the model. Mill maintenance owned the crews. The outcome sat between them.

The post-mortem traced the chain arrow by arrow over the year.

Of 400 mill assets flagged as high-risk, 260 were opened, 90 became work orders and only 30 were inspected before failure.400Assets flagged high-risk260Opened by a planner90Became a work orderThe workflow arrow broke here30Inspected in time
Figure 3.6.9 Illustrative figures. The capability arrow held. The chain broke where a prediction had to become scheduled work.

Four hundred assets were flagged; the capability worked. Two hundred and sixty were opened by a planner, which was adoption of a kind. Only ninety became a work order, because the roster had no slot for a prediction. Thirty were inspected before they failed, and only those thirty could ever have touched the outcome.

The fix was not a better model. The head of mill maintenance became the outcome owner. Rankings fed straight into the weekly crew schedule, with a fixed share of crew hours held back for the highest-risk assets. Statutory inspections of cranes and pressure equipment became the guardrail: they had to stay on schedule. The measure became unplanned downtime hours, the figure the plant already reported every month, not model accuracy. In the language of this chapter, the steelmaker finally wrote its outcome hypothesis, a year after it had built its capability.

What this means for leaders

Four habits separate proposals that turn capability into results from those that stop at the demonstration. Insist on a use case, not a technology: a named group of people, a step of work and the change in it. Ask for the arrows, written as “this only works if” statements, with the person who controls each one. Fund the test of the weakest arrow before the build, especially when that arrow runs through customers or partners. And refuse to approve an outcome that has no business owner, because the arrows below adoption are management work, not technology work.

Check yourself

  1. A capability becomes a use case only when it is applied to a specific task, for specific people, in a specific workflow.
  2. If most staff use an AI tool every week, the adoption arrow has held and the outcome will follow.
  3. The first arrow to test is the capability, because everything else depends on it.
  4. In McKinsey’s March 2025 survey, workflow redesign had the biggest effect on earnings impact of the 25 attributes tested.
  5. The AI team should own the business outcome, because it understands the technology best.
  6. An accurate model can leave the business outcome unchanged.

Reflection: find the broken arrow

What comes next

An outcome hypothesis states what should change, by how much and against what starting point. It does not prove that the change happened, or that AI caused it. Every target in this chapter assumed that the starting point was known and that something else would not explain the movement. The next chapter, Baselines, Metrics and Measurement, shows how to establish the starting point and how to prove that the outcome actually moved.

References

  1. Erik Brynjolfsson, Danielle Li and Lindsey Raymond. Generative AI at Work. The Quarterly Journal of Economics 140(2). 2025.
  2. Fabrizio Dell'Acqua et al. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality, HBS Working Paper 24-013. Harvard Business School. 2023.
  3. Rita Gunther McGrath and Ian C. MacMillan. Discovery-Driven Planning. Harvard Business Review, July-August 1995. 1995.
  4. McKinsey & Company (QuantumBlack). The state of AI: How organizations are rewiring to capture value. McKinsey & Company. 2025.
  5. Boston Consulting Group. Where's the Value in AI?. Boston Consulting Group. 2024.
  6. Cynthia Rudin, David Waltz, Roger N. Anderson et al. Machine Learning for the New York City Power Grid. IEEE Transactions on Pattern Analysis and Machine Intelligence 34(2): 328-345. 2012.

Further reading

Sources last verified 2026-10-08.