AI Academy · Book
Executives & Directors · Module 03 · Chapter 014

Module 03 Synthesis — Choosing Value Over Hype

Hype is not a lie; it is a claim that skips links in the chain from capability to cash. Every tool in this module is a way of putting those links back: define the claim, find its baseline, check adoption and capture, subtract the cost, and read each piece of evidence for what it can and cannot prove. Leaders who run that test stay ambitious about value and become hard to fool.

≈ 14 min read

After this chapter you can

  • Recall the Module 03 chain in one pass, from the value hypothesis to the scorecard.
  • Walk a headline AI claim through definition, scope, adoption, capture and cost, and find the step that loses the most.
  • Identify where hype comes from, inside the organization as well as outside it, and what it stands in for.
  • Judge which link of the value chain a given piece of evidence actually covers.
  • Carry four questions into every AI review and let the evidence choose the next step.

Suppose a memo lands on your desk on Monday, with a request for your signature by Friday. It proposes a general-purpose AI assistant for all ten thousand employees. It says every employee will become 25 percent more productive. It does the arithmetic for you: ten thousand people at an average cost of 40,000 a year is a payroll of 400 million, and 25 percent of that is 100 million of value, every year.

A proposal memo claims 10,000 employees, 25 percent more productive, and 100 million a year in value.10,000EmployeesOne assistant for everyone+25%ProductivityClaimed for every employee100MValue per year25% of a 400M payrollILLUSTRATIVE NUMBERS
Figure 3.14.1 An illustrative memo. Each number is plausible on its own; the chain between them is missing.

Nothing in the memo is obviously false. The multiplication checks out. In controlled studies, assistants of this kind do make many tasks faster, and competitors are buying them. The memo is not a lie. It is something more common and more expensive: a claim that skips most of the links between what a technology can do and what a business banks.

This module has given you a tool for each of those links. This closing chapter puts them together into one test, runs the memo through it, and then applies it to a real, carefully run rollout, to show that even good evidence proves less than it seems to.

Value first: outcome, proof, capture

The difference between hype and value is where the thinking starts. Technology-first thinking asks what AI can do, finds somewhere to use it and hopes the value appears. Value-first thinking starts with a business result that matters and works backwards: which outcome must change, what evidence will show that it changed, and who will turn the change into money, capacity or avoided risk.

Technology-first thinking starts with capability and hopes for value; value-first thinking starts with the outcome, the proof and the capture.TECHNOLOGY FIRSTWhat can AI do?Hope for valueVALUE FIRSTWhat outcome must change,and how will we prove andcapture it?Evidence of valuevs
Figure 3.14.2 The same tool can sit in either column. What differs is the question asked before the money is spent.

Choosing value over hype is not caution. An organization that knows quickly which bets are working can afford to make more of them, and bigger ones. The discipline is a way of being ambitious about value while being exacting about evidence. Its opposite is not boldness; it is spending without learning.

Module 03 on one page

Each chapter in this module added one link to the chain, and read in order they form the test this chapter applies. Value is the change AI causes against what would have happened anyway, net of cost, and each source needs its own proof: incremental revenue, a budget line that falls, customer behavior that changes, time that is actually freed. Between capability and outcome, every arrow is an assumption. Measurement fixes the baseline and the decision rule before launch, treats freed time as potential until someone decides where it goes, and uses early signals only if they predict results. The business case shows a range; realization needs an owner because cost is fixed while benefit leaks; each benefit is counted once; and the scorecard ends in a verb, an owner and a date.

Module 03 as one chain - value sources, outcome, measurement, business case and decision - forming one test for every claim.ValueFoursourcesOutcomeEveryarrow is anassumptionMeasureBaseline,capacity,signalsCaseRange,return,impactDecideVerb,owner, dateOne test for every claim
Figure 3.14.3 Thirteen chapters, one chain. Each link is a place where a confident claim can lose most of its value.

Walk the memo before you sign

Run the memo through the chain, one link at a time. The numbers that follow are illustrative, but each step is a question you can send back to any author of any proposal.

Definition. “Twenty-five percent more productive” usually means a task gets done 25 percent faster. That does not free a quarter of the time. Work that took 125 units now takes 100, so the time freed is 25 out of 125, one fifth. The general rule, from AI and Workforce Productivity, is that X percent faster frees X/(100+X) of the time. The memo is already at 80 million, not 100.

Scope. The assistant does not touch every hour of the week. Suppose it helps with drafting, summarizing and searching, about 30 percent of a typical week. Twenty percent of 30 percent is 6 percent of total time, which is 24 million.

Adoption. Licenses are not use. Suppose half of the staff use the assistant every week in a way that changes how they work. The figure falls to 12 million.

Capture. Freed time is potential until someone decides what it becomes, as Productivity vs Realized Capacity showed. Suppose half of it turns into fewer hires, more output or overtime avoided, and the rest dissolves into the day. The figure is 6 million.

Cost. Licenses, integration, training and support come off the top. Say 4 million a year. What remains is about 2 million.

The memo's 100 million falls to about 2 million once definition, scope, adoption, capture and cost are applied.100 MMemo claim−20 MSpeed frees 20%−56 M30% of tasks−12 MHalf adopt−6 MHalf captured−4 MCost2 MNetILLUSTRATIVE NUMBERS
Figure 3.14.4 Illustrative. Definition, scope, adoption, capture and cost take a 100 million claim to about 2 million: 80, 24, 12, 6, then 2.

Two million a year may still be a good investment. But it is a different decision from one hundred million: it needs a smaller commitment, a pilot that tests the weakest of the five assumptions, and a named owner for the freed time. Notice also that the biggest loss came not from adoption or cost, which proposals often discuss, but from scope, which they often leave unstated. With other illustrative assumptions, a different step could lose the most. The test works because it makes each assumption visible, so that the right people can argue about the right number.

Where hype comes from

It is tempting to think of hype as something vendors do to buyers. Much of it is home-grown. Sponsors want their projects approved, and a large number travels further than a careful one. Building the AI Business Case showed that the UK Treasury treats optimism in appraisals as a proven, systematic tendency and corrects for it before a case is approved1. AI Value Scorecard showed that when federal auditors rechecked the ratings agencies gave their own IT investments, they found more risk than the rating showed in 24 of 53 cases2. None of this is peculiar to AI, and none of it requires anyone to lie. Each sponsor believed the number, and the approval process rewarded the most confident version of it. The defense is not suspicion of people. It is a process that asks for the links before it accepts the total.

Business cases are corrected for proven optimism, and self-ratings understated risk in 24 of 53 federal IT investments.Business casesOptimism treated as proven and corrected before approvalSelf-ratingsMore risk than the official rating in 24 of 53
Figure 3.14.5 Optimism is an internal problem before it is a vendor problem.

Hype is easiest to spot by what stands in for evidence. “Everyone is doing AI” stands in for a business problem. “The model is state of the art” stands in for an outcome. “People love it” stands in for a baseline. “The vendor says five times return” stands in for a counterfactual. None of these statements is wrong in itself. Each becomes a warning sign when it occupies the place where a measured link should be.

Hype phrases such as everyone is doing AI and people love it sit where questions about outcome, baseline, capture and cost should be.What hype offersEveryone is doing AIThe model is state of the artPeople love itUsage is up every weekThe vendor quotes 5x returnWhat evidence asksWhich outcome must change?Better than what we do now?Against which baseline?Who captures the freed time?Net of what cost?
Figure 3.14.6 Each hype phrase occupies the place of a question. Ask the question and the phrase either earns its place or falls away.

When hype leaves the building, it becomes a legal matter. As What Exactly Is Artificial Intelligence? described, US regulators have already penalized firms for overstating their use of AI to investors and customers34. An internal memo is not an investor document, but the numbers in it have a habit of travelling into one.

Read evidence for what it can prove

The second half of the test is reading the evidence itself. Every study, dashboard and survey answers a narrower question than the headline built on it. The way to stay honest is to ask which link of the chain a piece of evidence actually covers.

Software development is a useful place to see this, because the same kind of tool has been studied at every level. In a controlled experiment, developers with an AI pair programmer finished one well-defined programming task 55.8 percent faster than those without it5. Check the definition before applying the module’s rule: the paper’s figure is a cut in completion time, about 71 minutes against 161, so on that task the work ran more than twice as fast. That is strong evidence of capability on that task, and it says nothing yet about the rest of a developer’s week. Three randomized field experiments covering 4,867 developers found about 26 percent more completed tasks, with large uncertainty around the estimate6, which is evidence about output in real work. And a vendor’s analysis of about 800 developers, not peer reviewed, found no significant change in cycle time and a 41 percent higher bug rate among those with access7, which touches the link that matters to a customer: whether finished, working software arrives sooner.

A controlled experiment, field experiments and team metrics each prove a different link; only your own baseline shows whether the tool pays in your organization.EvidenceWhat it foundWhat it can proveControlled task experiment55.8% less time on one taskCapability on that taskField experiments inthree firmsAbout 26% more completed tasksOutput in real work in those firmsTeam delivery metricsNo change in cycle time; 41% more bugsWhether delivery and quality movedYour own baselineNot yet measuredWhether it pays here, net of cost
Figure 3.14.7 Each level of evidence covers a different link. None of them substitutes for a baseline in your own organization.

The lesson is not that one study is right and another wrong. They measure different links, in different firms, with different methods. A leader who quotes the 55.8 percent as the business case is using capability evidence to answer an outcome question. A leader who quotes the 41 percent to block the investment is making the same mistake in reverse. The honest position is that the external evidence makes the bet worth testing, and only your own baseline can show whether it pays.

Story: a careful rollout, read as a post-mortem

In January 2025, a team at ZoomInfo, a business-to-business data and software company, published an account of how the company rolled out an AI coding assistant to its developers8. It is worth reading as a post-mortem, not because anything went wrong, but because it is unusually candid about what a well-run deployment can and cannot show.

The rollout was staged. In July 2023, five engineers tried the tool for a week. The company then recruited 126 engineers, about a third of its developers, chosen to span specialties, seniority, locations and technology stacks, for a two-week trial in August. Company-wide rollout began in September, with licenses released in stages over several months, and eventually reached more than 400 engineers.

ZoomInfo moved from five engineers to a stratified trial of 126 to a staged rollout to more than 400 engineers, then published what it had measured.Jul 2023Five engineersOne-week first lookAug 2023126 engineersStratified two-week trialSep 2023Staged rolloutLicenses releasedover monthsJan 2025Results published400+ engineers; outcomesstill to come
Figure 3.14.8 A staged rollout with a stratified trial, as described by ZoomInfo’s engineers (Bakal et al., 2025).

The measurements were real and specific. Developers accepted 33 percent of the suggestions the tool showed them and 20 percent of the suggested lines of code. Satisfaction was 72 percent. Of the developers who answered the trial survey, 90 percent said the tool cut the time their tasks took, with a median reduction of 20 percent.

Now apply the test. A 20 percent cut in task time is the same thing as being 25 percent faster, the very number in the memo. Here it is a self-reported median, from people who chose to answer a survey, about the tasks they did with the tool. Acceptance rates and satisfaction are leading signals of adoption, and good ones. What the paper does not yet show is whether features shipped sooner, defects fell, or hiring plans changed. The authors say so plainly: they were watching their regular delivery metrics and planned to report on causality in a later paper.

That is the point of the story. ZoomInfo did nearly everything this module recommends at the start: it piloted before scaling, chose a representative trial group, measured more than log-ins and published its limits. Even so, the evidence it had reached the adoption and perception links of the chain, not the outcome and capture links. Those links do not arrive by themselves. They have to be designed in, with a baseline for the outcome and an owner for the freed time, before the rollout starts.

Evidence picks the verb

A test is useful only if it changes what happens next. AI Value Scorecard set out the verbs: scale, adjust, continue, pause or stop, each chosen by a rule agreed before the numbers arrived. Running the value test is how you know which rule applies. If the outcome has moved and the economics hold, scale. If the tool works and the outcome has not moved, the fix is usually in the workflow or the capture, not the model, so adjust. If there is no credible path to the outcome, stop, and treat the freed budget as a win. How to rank and stage many such bets at once is the subject of Module 09.

The test reduces to four questions that travel to every AI review, whoever is presenting and whatever the tool.

Four questions form a loop - what changes, how will we know, who captures it and what decides the next step.Outcome?What changes inthe businessProof?Baseline and metricCapture?Owner for the gainDecision?Forecast againstactualValue loopEvery AI review
Figure 3.14.9 Four questions form a loop. Comparing each forecast with its result makes the next forecast more honest.

The loop matters as much as the questions. Each time a forecast is compared with its result, the organization learns how optimistic its own proposals tend to be, and the next memo arrives with a more believable number.

What this means for leaders

Three habits carry this module into practice. First, when someone says “the value”, ask which number they mean: potential, expected or realized, in the terms of Building the AI Business Case. Many arguments about AI value are two people quoting different numbers. Second, walk every large claim through definition, scope, adoption, capture and cost before debating it, and find the step that loses the most. Third, read each piece of evidence for the link it covers, and insist that the outcome and capture links are designed in before launch, because they cannot be added afterwards.

None of this requires technical expertise. It requires knowing what evidence to demand, and continuing to demand it after the launch.

Check yourself

  1. Twenty-five percent faster frees a quarter of the time.
  2. Choosing value over hype means being cautious about AI.
  3. A high acceptance rate for AI suggestions shows the business has gained.
  4. Internal teams create inflated benefit claims as readily as vendors do.
  5. Overstating a firm’s use of AI has already drawn penalties from US regulators.
  6. If a tool raised output in another firm’s experiment, it will pay off in yours.

Reflection: find the missing link

What comes next

This module answered two questions: why invest in AI, and how to tell whether an investment creates value. A company can measure each bet well and still place its bets in the wrong direction. The next module asks what the bets are for. It begins with What Is an Enterprise AI Strategy?, which treats strategy not as a list of AI projects but as a short set of hard choices traced back to the business.

Laws referenced

US enforcement against false AI claims ("AI washing") · US - federal

Federal securities antifraud rules and Investment Advisers Act Marketing Rule (SEC); FTC Act Section 5 (unfair or deceptive practices)

There is no AI-specific federal statute, but existing law already applies to what companies say about AI. Regulators have penalized firms that overstated their use or capability of AI to investors (SEC) and to consumers (FTC). Claims about AI in marketing, investor materials and product descriptions need the same substantiation as any other claim.

  • 2024-03-18 — SEC's first AI-washing cases: Delphia and Global Predictions settle for USD 400,000 in total civil penalties
  • 2024-09-25 — FTC launches Operation AI Comply, with five actions over deceptive AI claims and uses

Last verified 2026-10-08

References

  1. HM Treasury. Supplementary Green Book Guidance: Optimism Bias. GOV.UK. 2013.
  2. U.S. Government Accountability Office. IT Dashboard: Selected Agencies' Investment Ratings Fail to Fully Consider Risks (GAO-27-108416). GAO. 2026.
  3. US Securities and Exchange Commission. SEC Charges Two Investment Advisers with Making False and Misleading Statements About Their Use of Artificial Intelligence (Press Release 2024-36). SEC. 2024.
  4. US Federal Trade Commission. FTC Announces Crackdown on Deceptive AI Claims and Schemes (Operation AI Comply). FTC. 2024.
  5. Sida Peng, Eirini Kalliamvakou, Peter Cihon and Mert Demirer. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv (2302.06590). 2023.
  6. Zheyuan (Kevin) Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng and Tobias Salz. The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. Management Science. 2026.
  7. Uplevel Data Labs. Does GenAI Improve Software Developer Productivity?. Uplevel. 2024.
  8. Gal Bakal, Ali Dasdan, Yaniv Katz, Michael Kaufman and Guy Levin. Experience with GitHub Copilot for Developer Productivity at Zoominfo. arXiv (2501.13282). 2025.

Further reading

Sources last verified 2026-10-08.