Measuring AI Business Value
AI is worth what the business measurably changes because of it, not how much it is used. Usage is easy to count and value is not, so many AI updates report activity. Because whatever gets reported gets funded and copied, the choice of number steers the whole program.
After this chapter you can
- Distinguish AI activity metrics from evidence of business value.
- Explain why measurement steers what an organization funds, copies and scales.
- Recognize how a single flattering number hides trade-offs.
- Apply the business-metric test to an AI proposal or update.
- Ask four value questions and name the Module 03 chapters that answer them.
In July 2024 the freelance marketplace Upwork published a survey of 2,500 people in the United States, the United Kingdom, Australia and Canada. Of the executives, 96 percent expected AI tools to raise their company’s overall productivity. Of the employees who used those tools, 77 percent said AI had added to their workload, and 39 percent said they were spending more time reviewing or moderating content the AI had produced1.
Treat the numbers with care. Upwork sells access to freelance talent, the employee sample was only 625 people, and the survey measured opinions rather than output. But the shape of the result is worth a leader’s attention, because both groups can be telling the truth. The executives see an AI program that is spreading: more licenses, more users, more pilots every quarter. The employees see the work itself, including the hours spent checking drafts that looked finished and were not. One group is looking at activity. The other is looking at what the activity does to the work.
Activity is not value
The gap has a simple explanation. Activity and value are two different axes, and much AI reporting measures only one of them.
Activity is how much AI is used. Value is whether a number the business already cares about, such as cycle time, cost to serve, quality, revenue or risk, has moved because of that use. The two are related, since AI that nobody uses changes nothing, but they are not the same thing. A program can sit in the bottom-right corner for a long time: broad use, growing dashboards, and no business number that has noticed. Call that corner adoption theater. Everyone is busy with AI, and the business is unchanged.
The top-left corner is the one leaders tend to forget. A single team using AI on a critical process can move a number that thousands of casual users never touch. Usage alone cannot tell you which corner you are in.
So the argument of this chapter is short. AI is worth what the business measurably changes because of it, not how much it is used. As AI Is Changing Everything showed, the distance between using AI and getting value from it is a central gap of this period2. Measurement is how a leader finds out which side of that gap a given initiative is on.
What popular AI metrics actually prove
Many AI updates are built from the same handful of numbers. Each proves something. None proves, on its own, what it is usually offered to prove.
Licenses prove the organization bought something. Weekly users prove people showed up, and prompts prove they tried it. Pilots prove teams experimented. Model accuracy proves the model performs well on a test set, which is not the same as performing well inside your process. Hours saved come closest to value and are the most often mistaken for it: they prove that a task got faster, if the clock was real, but not that the hour went anywhere useful.
None of these numbers is useless. Empty seats are a real warning, and a model that fails its test will not succeed in production. They are early signals, and they belong in an update. The mistake is to label them as results. How early signals relate to results is the subject of Leading vs Lagging AI Metrics, and what happens to a saved hour is the subject of Productivity vs Realized Capacity, both in Module 03.
Measurement steers the organization
If counting usage were merely incomplete, it would be a reporting problem. It is worse than that, because measurement does not just describe an organization. It steers it.
In 1975 the management scholar Steven Kerr published a paper with a title that has outlived most of its era’s research: “On the Folly of Rewarding A, While Hoping for B”. His examples, drawn from politics and war, medicine, universities and business, showed the same pattern again and again. Organizations hoped for one behavior and rewarded another, and then were surprised to get the behavior they rewarded3.
AI programs run Kerr’s loop every quarter. The update reports what is easy to count, which is usually activity. Leaders reward what the update shows, with funding, praise and the next rollout. Other teams notice what was rewarded and chase the same number. A few quarters later, that number has quietly become the organization’s definition of AI success. Run the loop on usage and you train the organization to produce usage. Run it on a business result and you train it to change the work.
The loop runs whether anyone designs it or not. The only choice is which number goes into it, and that choice belongs to the business owner of the process: not the vendor, and not the technology team alone, because only the owner answers for the number when it does not move. The AI Transformation Challenge placed ownership of AI outcomes with the business; measurement follows ownership.
There is a second reason to measure, and it is humbling. Even organizations that test everything find that most ideas do not work as hoped. Ron Kohavi and his co-authors report that at Microsoft, only about one third of the ideas tested in controlled experiments improved the metric they were designed to improve4. If two in three ideas from experienced product teams fail to improve their target, an AI initiative’s sponsor is unlikely to do better by instinct.
One number hides the one that fell
The third reason is that AI creates trade-offs, and a single number hides them. Speed can rise while quality falls. Volume can rise while customer effort rises with it. Cost can fall while incidents climb.
Suppose an assistant makes a team faster. Indexed to 100 in the first month, speed climbs to 130 by the sixth. That is the number that goes on the update. Over the same months, the share of work that is right the first time, and does not come back for correction, slides to 88. Nobody lies; the second line simply is not in the report. This is an illustration, but the shape is documented. As AI Is Changing Everything showed, consultants using AI on a task just outside its capabilities were 19 percentage points less likely to reach the correct answer than colleagues working without it5. The Upwork respondents who spend extra time reviewing AI output are living on the second line1.
A car dashboard makes the point. The speedometer is useful, but no one drives by it alone. You also watch the warning lights, because a fast car that is overheating is not doing well, and you check that you are on the right road at all, because a driver who watches only speed can arrive very quickly at the wrong place. An AI program needs the same set of instruments: adoption tells you the engine is running, productivity is the speedometer, quality and risk are the warning lights, and value is whether you are on the right road. One dial is still missing, fuel, which is what the journey costs; that is the next chapter’s subject.
The business-metric test
All of this reduces to one test that an executive can apply to every AI proposal and every AI update.
The test is less demanding than it sounds. It does not require every benefit to appear in the accounts this quarter. Some value is quality, resilience or faster learning, and it takes time to reach the profit and loss statement. But value that is real can still be observed. Intangible is not the same as unobservable: if nobody can name the number that should move, nobody will be able to tell you whether it did.
Notice what fails the test. Usage fails, because it proves adoption. Model accuracy fails, because a precise forecast that nobody acts on changes nothing. A vendor’s return-on-investment calculator fails too, because it cannot see your mix of work, your review effort or your hidden rework; until your own process and your own numbers say otherwise, it is marketing. How to build a case that passes is taught in Building the AI Business Case, in Module 03.
Four questions every value claim must survive
When someone tells you an AI initiative has created value, four questions will test the claim. Module 03 builds the tools behind each of them; here they are only questions.
Did the work change, and did the business notice? A capability must change how people work, then an operational number, then a business outcome; a claim that skips a link is a hope. Compared with what? Without a measurement from before, taken the same way, an improvement is a story. Where did the time go? A saved hour becomes value only when someone decides what it is for. What else moved? One number can hide another, so a claim is complete only when the neighboring numbers are on the page.
Story: the regulator that tested the summaries
The clearest demonstration of these questions comes from a regulator that asked them before scaling anything.
The Australian Securities and Investments Commission (ASIC) regularly deals with public submissions to parliamentary inquiries: long documents from firms, associations and individuals, each of which a staff member must read and summarize, noting what it says about ASIC and what it recommends. It is skilled, slow work, and it is exactly the kind of task that generative AI is sold to speed up. Between 15 January and 16 February 2024, ASIC ran a proof of concept with Amazon Web Services to find out whether a large language model, Meta’s Llama 2, could summarize a sample of submissions to a parliamentary inquiry into audit and consultancy firms6.
Stop here and decide. You run the trial. The obvious evidence is easy to get: how quickly the model produces a summary, and whether staff enjoy using it. Better evidence costs more: summaries judged on whether they could actually be used, against summaries written by your own people. Which would you measure before deciding?
ASIC chose the harder measure. Ten staff members of varying seniority wrote their own summaries of the same five submissions. Five assessors then scored all the summaries against the same criteria, including coherence, length, references to ASIC and to regulation, and whether the recommendations were identified, without being told that some had been written by AI, according to press accounts of the report78.
The human summaries scored 61 of 75 points, or 81 percent. The AI summaries scored 35, or 47 percent, and they scored lower than the human summaries on every criterion. The assessors found that the AI often missed emphasis, nuance and context, and sometimes included incorrect information. Most telling for a leader, they concluded that in its current state the tool could create more work rather than less, because its output had to be fact-checked against the original submissions, and because the originals sometimes presented the information better6. ASIC set out these results in its response to the Senate Select Committee on Adopting Artificial Intelligence, and the committee published them later that year68.
Two cautions keep the story honest. The model has since been superseded, and the team had only about a week to optimize the model and its prompts; ASIC itself warned that this limits how far the findings generalize, so the result says little about what AI summarization can do in 2026. ASIC concluded that AI should be positioned to augment people’s work rather than replace it6.
The lesson is not about the model. It is about the measurement. Had ASIC counted speed, or staff enthusiasm, the trial would have looked like a success, and the first number in the next update would have been the wrong one. By measuring usable output against its own people, judged blind, ASIC learned in five weeks what many organizations learn only after a rollout: a summary that has to be checked line by line is not a saved hour. Your answer to the dilemma is the measure your own organization would have used.
What this means for leaders
Measuring AI value is a leadership task before it is an analytical one. The analysts can produce any number you ask for. Only leaders decide which number counts, and the organization will learn to produce whatever that is. So ask for the business number first, and keep the activity numbers on the update as early signals rather than as the headline. Ask for a second number, the one that would fall if the first one rose for the wrong reasons. And name the business owner who will answer for both.
This is also where the AI program earns or loses its credibility with the board. An update that reports record usage and a flat business will be believed once. An update that reports a smaller, honest gain, with the trade-off beside it, is the one that gets the next round of funding.
Check yourself
- High weekly usage proves an AI initiative is working.
- Usage numbers are worth tracking as early signals.
- Time saved equals money saved.
- Whatever an AI update reports tends to become what the organization pursues.
- If the model scores well on accuracy, the business case is proven.
- A vendor’s return-on-investment calculator can serve as your business case.
Reflection: test your last AI update
What comes next
Measurement tells you whether value appeared. It does not tell you what the value cost. Every assistant, agent and dashboard costs something to build, run and change, and the price of the model is only one line of that bill. The next chapter, The Economics of AI, turns to the missing dial on the dashboard: what AI actually costs.
References
- Upwork Research Institute. From Burnout to Balance: AI-Enhanced Work Models for the Future. Upwork. 2024.
- McKinsey & Company (QuantumBlack). The state of AI in 2025: Agents, innovation, and transformation. McKinsey & Company. 2025.
- Steven Kerr. On the Folly of Rewarding A, While Hoping for B. Academy of Management Journal, 18(4), 769-783. 1975.
- Ron Kohavi, Diane Tang and Ya Xu. Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press. 2020.
- Fabrizio Dell'Acqua et al. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality, HBS Working Paper 24-013. Harvard Business School. 2023.
- Australian Securities and Investments Commission (ASIC). Response to Senate Select Committee on Adopting Artificial Intelligence: generative AI summarisation proof of concept. Parliament of Australia, Senate Select Committee on Adopting Artificial Intelligence. 2024.
- Crikey. AI worse than humans in every way at summarising information, government trial finds. Crikey. 2024.
- Information Age (Australian Computer Society). Humans outperform AI in Australian government trial. Information Age. 2024.
Further reading
- Steven Kerr. On the Folly of Rewarding A, While Hoping for B. Academy of Management Journal, 18(4), 769-783. 1975.
- Ron Kohavi, Diane Tang and Ya Xu. Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press. 2020.
- Australian Securities and Investments Commission (ASIC). Response to Senate Select Committee on Adopting Artificial Intelligence: generative AI summarisation proof of concept. Parliament of Australia, Senate Select Committee on Adopting Artificial Intelligence. 2024.
Sources last verified 2026-10-09.