AI Academy · Book
Executives & Directors · Module 03 · Chapter 001

What Does AI Value Actually Mean?

AI value is not how much a tool is used or how impressive it looks. It is the measurable change in business results that AI causes, compared with what would have happened without it, net of what it costs. Leaders who insist on that definition, and on a written hypothesis before anything is built, stop paying for activity and start paying for results.

≈ 16 min read

After this chapter you can

  • Define AI value as a measured, net difference against what would have happened without AI.
  • Distinguish capability and adoption from value using the capability-to-value chain.
  • Place any AI metric on the activity, output, outcome and value ladder.
  • Explain why a successful pilot proves feasibility, not return.
  • Write a testable value hypothesis before anything is built.

In the first half of 2025, the research group METR ran a randomized trial with sixteen experienced software developers working on large open-source projects they knew well. Each developer brought real tasks, 246 in all, and each task was randomly assigned to be done with or without AI tools. Before they started, the developers predicted that AI would cut their completion time by 24 percent. When the work was done, they estimated it had cut their time by 20 percent. The measured result went the other way: with AI allowed, tasks took 19 percent longer1.

Developers expected AI to save 24 percent of their time and believed it saved 20 percent, but measured tasks took 19 percent longer.24%Expected time savedDevelopers' forecast beforethe trial20%Believed time savedDevelopers' estimate afterwards19%Measured time addedTasks took longer with AI allowedSource: Becker et al., METR · 2025
Figure 3.1.1 Enthusiasm and experience said faster. The comparison with the world without AI said slower.

The study is narrow. It covers one kind of expert, one kind of codebase and the tools of early 2025, and other studies in other settings have found large gains. Its lesson for executives is not that AI fails. It is that the people using a tool, sincerely and with expertise, could not tell from the inside whether it was helping. Only a comparison with the work done without AI revealed the result.

Most AI programs report exactly what those developers could see: who uses the tool, how often, and how much they like it. The question that decides whether the money was well spent is harder, and it starts with a definition.

Value is a comparison, not a count

Here is the definition this module uses. AI value is the measurable business benefit created when AI changes a process, a decision, a product or a customer interaction, compared with what would have happened without it, net of what it costs to achieve.

Three phrases in that sentence do the work. Measurable rules out benefits that are only asserted. Changes means that value comes from work being done differently, not from a tool being available. And compared with what would have happened without it means that value is always a difference between two worlds, one of which you never get to see. A count of users, prompts or drafts is not a difference between anything.

The definition is deliberately broad about where the benefit lands. It can be revenue, cost, customer experience, risk, speed or strategic position. It is deliberately strict about proof.

Leaders have reason to be strict. In IBM’s 2025 survey of 2,000 chief executives, the CEOs themselves reported that only 25 percent of their AI initiatives had delivered the expected return over the previous few years, and only 16 percent had been scaled across the enterprise2. Those are self-reported figures, so the true picture could be better or worse. Either way, most initiatives were approved on promises that the organization did not later see in its results.

Value sits at the end of a chain

The simplest way to apply the definition is to see AI value as the last link in a chain. Each link asks its own question, and each can break.

AI value appears only at the end of a chain from capability to adoption, behavior change, outcome and net value.CapabilityWhat canit do?AdoptionAre peopleusing it?BehaviorDoes workchange?OutcomeWhatmeasurablychanged?ValueWhat is itworth, net?Most AI reporting stops at the second link
Figure 3.1.2 Capability and adoption are necessary. Value appears only if every link after them holds.

Capability is what the system can do: summarize a contract, flag a fraudulent payment, draft a design report. Demonstrations live here, and they are often genuinely impressive. Adoption is whether people use it. Most AI dashboards live here, because licenses, logins and prompts are easy to count. Behavior change is whether work is actually done differently: a step removed, a decision made earlier, a person freed for something else. Outcome is what measurably changed in the business as a result. Value is what that change is worth, after the cost of getting it.

Each link is necessary and none is sufficient. A capable tool nobody uses creates nothing. A widely used tool that leaves the workflow untouched usually creates little. Even a real outcome can be worth less than it cost. From AI Capability to Business Outcome, later in this module, examines each arrow in the chain as an assumption to be tested. For now the point is simpler: when someone reports AI success, ask which link the evidence comes from.

The same chain gives a quick test for any number in an AI update. Measured, its links show up as four rungs: activity and output are evidence from the adoption link, while outcome and value are the last two links.

Activity and output metrics arrive first, but the investment decision depends on outcome and value.Activity100,000 requests a monthARRIVES FIRSTOutput80,000 summaries producedOutcomeLess time preparing reportsValueFreed time wins revenue
Figure 3.1.3 The chain, as measured: activity and output come from the adoption link and arrive first. The investment decision lives on the top two rungs.

Take the claim that an assistant summarizes a hundred-page report in thirty seconds. A hundred thousand requests a month is activity; eighty thousand summaries is output; analysts spending less time on reports is an outcome; that time going into work that wins revenue, shown to have done so, is value. The bottom rungs dominate reporting because they arrive first, automatically and with false precision. As Measuring AI Business Value argued, activity is not value. For every number in an AI update, name its rung, and send anything below outcome back with the question, “and what did that change?”

The world you never see

The hardest phrase in the definition is the comparison: what would have happened without AI. Economists call this the counterfactual. It is hard because it never happens. You observe the quarter with AI. The same quarter without AI exists only as an estimate.

Value is the gap between the observed outcome with AI and an estimate of the outcome without it, after checking for other causes.Same teamand quarterWith AIWhat you observeWithout AIWhat you must estimateCHECK BEFORE YOU CLAIM THE GAPVolumeStaffingSeasonalityOther changes
Figure 3.1.4 Value is the gap between the two paths, not the number on the first one.

Suppose month-end close took eight working days before an AI tool was introduced and six days after. That looks like two days of value. Before claiming it, ask what else changed. Did transaction volume fall that quarter? Did the team add people, or a new finance system? Any of these could explain part of the gain, and a before-and-after comparison cannot tell them apart. The METR trial shows the opposite risk: the developers’ own before-and-after impression would have recorded a gain where the controlled comparison found a loss.

Good comparisons exist, and they are often cheaper than leaders assume. When Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied an AI assistant used by more than 5,000 agents at one company, they could measure its effect because access was rolled out to different teams at different times. Teams still waiting served as the comparison. Measured that way, productivity rose by 15 percent on average, with the largest gains among less experienced staff3. A staggered rollout costs almost nothing extra and turns a launch into evidence.

The deeper reason to insist on a comparison is that intuition about value is unreliable even among experts. At Microsoft, where product ideas were routinely tested in controlled experiments, only about one third of the ideas tested improved the metric they were designed to improve, according to a 2009 study4. In well-optimized products the success rate was lower still5. These were product and website ideas, not AI initiatives, but they were ideas that experienced teams believed in enough to build. Baselines, Metrics and Measurement covers how to choose the right comparison method. What every executive needs now is the reflex: compared with what?

Value arrives through several doors

Value is not one thing. AI can create it through several mechanisms, and a sound case names the one or two that matter for the use case at hand.

AI creates value through six mechanisms; each initiative should name the one or two that matter for its use case.RevenueNew, expanded or protected salesCostLower cost per unit of workCustomer experienceBetter and faster serviceWorkforce productivityMore or better output per personRiskFewer errors, losses and breachesSpeed and positionFaster learning, stronger strategy
Figure 3.1.5 Name the one or two mechanisms that matter for this use case. No initiative moves all six.

There is no universal AI return on investment. A tool that drafts tender documents and a model that predicts pump failures create value through different mechanisms, so they need different value models and different measures.

The mechanisms also differ in how certain their evidence can be. Some benefits are direct and countable: less downtime, fewer errors, more contracts won. Others are indirect: faster experimentation, better access to knowledge, a workforce more able to use the next tool. Both are real, but a credible case reports them with different degrees of confidence rather than adding them into one number.

Workforce productivity deserves one warning now, because it is the mechanism most often overstated. In a randomized trial by Shakked Noy and Whitney Zhang, professionals given ChatGPT for writing tasks finished them in about 40 percent less time, and their work was rated 18 percent higher in quality6. That is a genuine outcome. It becomes value only when the organization does something with the freed time.

Read the wording of any speed figure, because two different measures hide behind it. “Less time” is a time reduction: Noy and Zhang’s 40 percent less time frees 40 percent of the time, so a 100-minute task takes 60. “Faster” usually means more work per hour, which frees less: a task done 30 percent faster takes 100 divided by 1.3, about 77 minutes, freeing about 23 percent. The rule, X percent faster frees X divided by (100 + X), applies only to the second kind; applied to Noy and Zhang it would wrongly shrink their result to 29 percent. AI and Workforce Productivity and Productivity vs Realized Capacity develop both points. The next four chapters take revenue, cost, customer experience and productivity in turn.

A working pilot is a show home

The most expensive confusion in AI programs is between the technology works and the investment works. A pilot is designed to answer the first question. It is often reported as if it had answered the second.

Think of a property developer’s show home. It proves that the design is attractive and can be built. It does not prove that two hundred houses can be built on budget, on that ground, with that workforce, or that they will sell at a profit.

Like a show home, an AI pilot proves the design works but not that it works at scale or pays for itself.Does it proveShow homeAI pilotThe design worksIt works at full scaleIt pays for itself
Figure 3.1.6 A pilot proves feasibility. Scale and return are separate questions with separate evidence.

An AI pilot is a show home. High accuracy, happy users and strong technical performance are real achievements. Production can still fail on integration cost, data quality, security constraints, the effort of changing how people work, or simply scale. And a pilot’s benefit is gross. Value is net of the full cost of achieving it: integration, data preparation, security, monitoring, training and workflow redesign, not only the model or license fee. A valuable capability can therefore be a poor investment. AI ROI and Value Realization shows how much of a promised benefit typically survives the journey from approval to capture.

Write the value hypothesis first

If value is a measured difference, net of cost, then the way to manage it is to say in advance which difference you expect. That statement is a value hypothesis, and it is the most useful single habit in AI governance.

Compare two statements. The first: “We will implement an AI coding assistant.” The second: “If AI helps our developers generate unit tests, development cycle time will fall by at least 10 percent within two quarters, while escaped defects do not increase.” The first statement describes a project; it can only be completed. The second is a claim about the business; it can be tested, and it can fail. Only the second is a value hypothesis.

A value hypothesis has five parts - if this capability, for this workflow, then this outcome improves, by this target, while a constraint holds.IFwe introducethiscapabilityFORthis user orworkflowTHENthisoutcomeimprovesBYthis muchby this dateWHILEquality andrisk hold
Figure 3.1.7 A testable value hypothesis names the capability, the workflow, the outcome, the target and the constraint.

Each part does a job. If names the capability, so nobody confuses the tool with the result. For names the users or workflow, which keeps the claim small enough to test. Then names an outcome metric, from the top half of the ladder. By sets a target and a date, which turns hope into a commitment. While names the guardrail, so speed is not bought with quality or risk.

The hypothesis also tells you how to test it. A claim about one workflow can be checked with a controlled pilot, a staggered rollout or a comparison team. Given the Microsoft experience, expect some well-argued hypotheses to fail. That is the point of writing them: a failed hypothesis caught in month four is far cheaper than a failed program discovered in year two.

Story: the bid machine

An engineering consultancy of about seven hundred people designs roads, rail and water infrastructure. Most of its work comes through public tenders. In a typical year it submits about 100 bids and wins about 30, a win rate of 30 percent. Each bid takes weeks of engineers’ time, so bidding is one of the firm’s largest hidden costs.

Its leaders roll out an AI assistant that drafts tender responses from past submissions, project sheets and staff profiles. The board paper promises “faster bids and more wins.” It gives no number, no baseline beyond last year’s bid count, and no constraint.

An engineering consultancy's AI bid assistant went from launch to a green dashboard, yet after twelve months it won no more contracts.Month 1LaunchTraining for every bid teamMonth 4Drafts in halfthe timeBid managers delightedMonth 8Dashboard greenBids up 40 percentMonth 12Annual review140 bids, 30 wins
Figure 3.1.8 Every early signal was positive. The outcome that mattered did not move.

The early signals are all good. By month four, a first draft of a typical response takes half the time it used to. By month eight, the dashboard shows the firm on course to submit 40 percent more bids than the year before. At the annual review the numbers arrive: 140 bids submitted and 30 contracts won. The win rate has fallen from 30 percent to about 21 percent. Senior engineers, who review and price every bid, report more weekend work, not less.

The post-mortem finds four causes, and none of them is the technology. First, there was no value hypothesis. “More wins” was never written as “more profitable wins at no higher cost per win,” so the program optimized the number it could see, which was bids. Second, there was no comparison. Public tender volume in the firm’s markets had risen that year, so flat wins meant a falling share; no one had looked at the market or at a division that had not yet adopted the tool. Third, behavior did change, in the wrong direction. Because a bid now felt cheap, the bid/no-bid discipline loosened and the firm chased tenders it would once have declined. The hours saved on drafting went into more bids, not better ones. Fourth, the real constraint was senior review and pricing judgment, which the tool never touched.

The dashboard showed faster drafts, more bids and good feedback, while value was decided by a missing hypothesis, a missing comparison, looser bid discipline and an unchanged bottleneck.WHAT THE DASHBOARD SHOWEDDrafting time halved · Bids up 40percent · Glowing feedbackWHAT DECIDED VALUENo value hypothesisNo comparison with the marketBid discipline loosenedSenior review unchanged
Figure 3.1.9 The capability worked and adoption was excellent. The value was decided below the waterline.

When the firm relaunched, it started with a hypothesis: if AI drafts first versions of tender responses for the transport division, the win rate on bids submitted will rise from 30 to 34 percent within twelve months, while bid cost per contract won does not rise and the bid/no-bid gate stays unchanged. The water division would adopt six months later and serve as the comparison. The tool was the same. What changed was that the firm now knew what it was buying and how it would find out whether it had got it.

What this means for leaders

The definition of AI value changes the questions a leader asks, and the order in which they are asked. Before any discussion of models or vendors, the first question is what will be better because of AI, by how much, compared with what would have happened otherwise. If the answer is a capability or a usage target, the proposal is not yet a business proposal.

It also changes what counts as progress. Adoption is a leading signal worth watching, but reporting it as success trains the organization to produce adoption. Ask for at least one outcome metric from the top of the ladder in every AI update, and ask what the baseline was.

Finally, it changes how failure is treated. A value hypothesis that fails a fair test is useful information, bought cheaply. A program that never wrote one cannot fail, and for the same reason cannot succeed in any way that finance will recognize.

Check yourself

  1. High adoption of an AI tool shows that it is creating value.
  2. People using an AI tool every day can reliably judge whether it saves them time.
  3. “We will implement an AI coding assistant” is a testable value hypothesis.
  4. A successful pilot proves that an AI investment will pay back.
  5. A valuable AI capability can still be a poor investment.
  6. Many ideas that experienced teams believe in fail to improve their target metric when tested.

Reflection: find the missing hypothesis

What comes next

The definition says value can arrive through several doors. The one most boards ask about first is revenue, and it is also where attractive claims are hardest to prove. The next chapter, Where AI Creates Revenue, examines how AI can create, expand or protect revenue, and how to tell genuine incremental revenue from revenue that would have arrived anyway.

References

  1. Joel Becker, Nate Rush, Elizabeth Barnes and David Rein. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR (arXiv 2507.09089). 2025.
  2. IBM Institute for Business Value. 2025 CEO Study: 5 mindshifts to supercharge business growth. IBM Institute for Business Value, with Oxford Economics. 2025.
  3. Erik Brynjolfsson, Danielle Li and Lindsey Raymond. Generative AI at Work. The Quarterly Journal of Economics 140(2). 2025.
  4. Ron Kohavi, Thomas Crook and Roger Longbotham. Online Experimentation at Microsoft. Microsoft (Third Workshop on Data Mining Case Studies and Practice Prize). 2009.
  5. Ron Kohavi, Diane Tang and Ya Xu. Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press. 2020.
  6. Shakked Noy and Whitney Zhang. Experimental evidence on the productivity effects of generative artificial intelligence. Science 381(6654). 2023.

Further reading

Sources last verified 2026-10-08.