Module 03 Synthesis — Choosing Value Over Hype
Hype is not a lie; it is a claim that skips links in the chain from capability to cash. Every tool in this module is a way of putting those links back: define the claim, find its baseline, check adoption and capture, subtract the cost, and read each piece of evidence for what it can and cannot prove. Leaders who run that test stay ambitious about value and become hard to fool.
After this chapter you can
- Recall the Module 03 chain in one pass, from the value hypothesis to the scorecard.
- Walk a headline AI claim through definition, scope, adoption, capture and cost, and find the step that loses the most.
- Identify where hype comes from, inside the organization as well as outside it, and what it stands in for.
- Judge which link of the value chain a given piece of evidence actually covers.
- Carry four questions into every AI review and let the evidence choose the next step.
Suppose a memo lands on your desk on Monday, with a request for your signature by Friday. It proposes a general-purpose AI assistant for all ten thousand employees. It says every employee will become 25 percent more productive. It does the arithmetic for you: ten thousand people at an average cost of 40,000 a year is a payroll of 400 million, and 25 percent of that is 100 million of value, every year.
Nothing in the memo is obviously false. The multiplication checks out. In controlled studies, assistants of this kind do make many tasks faster, and competitors are buying them. The memo is not a lie. It is something more common and more expensive: a claim that skips most of the links between what a technology can do and what a business banks.
This module has given you a tool for each of those links. This closing chapter puts them together into one test, runs the memo through it, and then applies it to a real, carefully run rollout, to show that even good evidence proves less than it seems to.
Value first: outcome, proof, capture
The difference between hype and value is where the thinking starts. Technology-first thinking asks what AI can do, finds somewhere to use it and hopes the value appears. Value-first thinking starts with a business result that matters and works backwards: which outcome must change, what evidence will show that it changed, and who will turn the change into money, capacity or avoided risk.
Choosing value over hype is not caution. An organization that knows quickly which bets are working can afford to make more of them, and bigger ones. The discipline is a way of being ambitious about value while being exacting about evidence. Its opposite is not boldness; it is spending without learning.
Module 03 on one page
Each chapter in this module added one link to the chain, and read in order they form the test this chapter applies. Value is the change AI causes against what would have happened anyway, net of cost, and each source needs its own proof: incremental revenue, a budget line that falls, customer behavior that changes, time that is actually freed. Between capability and outcome, every arrow is an assumption. Measurement fixes the baseline and the decision rule before launch, treats freed time as potential until someone decides where it goes, and uses early signals only if they predict results. The business case shows a range; realization needs an owner because cost is fixed while benefit leaks; each benefit is counted once; and the scorecard ends in a verb, an owner and a date.
Walk the memo before you sign
Run the memo through the chain, one link at a time. The numbers that follow are illustrative, but each step is a question you can send back to any author of any proposal.
Definition. “Twenty-five percent more productive” usually means a task gets done 25 percent faster. That does not free a quarter of the time. Work that took 125 units now takes 100, so the time freed is 25 out of 125, one fifth. The general rule, from AI and Workforce Productivity, is that X percent faster frees X/(100+X) of the time. The memo is already at 80 million, not 100.
Scope. The assistant does not touch every hour of the week. Suppose it helps with drafting, summarizing and searching, about 30 percent of a typical week. Twenty percent of 30 percent is 6 percent of total time, which is 24 million.
Adoption. Licenses are not use. Suppose half of the staff use the assistant every week in a way that changes how they work. The figure falls to 12 million.
Capture. Freed time is potential until someone decides what it becomes, as Productivity vs Realized Capacity showed. Suppose half of it turns into fewer hires, more output or overtime avoided, and the rest dissolves into the day. The figure is 6 million.
Cost. Licenses, integration, training and support come off the top. Say 4 million a year. What remains is about 2 million.
Two million a year may still be a good investment. But it is a different decision from one hundred million: it needs a smaller commitment, a pilot that tests the weakest of the five assumptions, and a named owner for the freed time. Notice also that the biggest loss came not from adoption or cost, which proposals often discuss, but from scope, which they often leave unstated. With other illustrative assumptions, a different step could lose the most. The test works because it makes each assumption visible, so that the right people can argue about the right number.
Where hype comes from
It is tempting to think of hype as something vendors do to buyers. Much of it is home-grown. Sponsors want their projects approved, and a large number travels further than a careful one. Building the AI Business Case showed that the UK Treasury treats optimism in appraisals as a proven, systematic tendency and corrects for it before a case is approved1. AI Value Scorecard showed that when federal auditors rechecked the ratings agencies gave their own IT investments, they found more risk than the rating showed in 24 of 53 cases2. None of this is peculiar to AI, and none of it requires anyone to lie. Each sponsor believed the number, and the approval process rewarded the most confident version of it. The defense is not suspicion of people. It is a process that asks for the links before it accepts the total.
Hype is easiest to spot by what stands in for evidence. “Everyone is doing AI” stands in for a business problem. “The model is state of the art” stands in for an outcome. “People love it” stands in for a baseline. “The vendor says five times return” stands in for a counterfactual. None of these statements is wrong in itself. Each becomes a warning sign when it occupies the place where a measured link should be.
When hype leaves the building, it becomes a legal matter. As What Exactly Is Artificial Intelligence? described, US regulators have already penalized firms for overstating their use of AI to investors and customers34. An internal memo is not an investor document, but the numbers in it have a habit of travelling into one.
Read evidence for what it can prove
The second half of the test is reading the evidence itself. Every study, dashboard and survey answers a narrower question than the headline built on it. The way to stay honest is to ask which link of the chain a piece of evidence actually covers.
Software development is a useful place to see this, because the same kind of tool has been studied at every level. In a controlled experiment, developers with an AI pair programmer finished one well-defined programming task 55.8 percent faster than those without it5. Check the definition before applying the module’s rule: the paper’s figure is a cut in completion time, about 71 minutes against 161, so on that task the work ran more than twice as fast. That is strong evidence of capability on that task, and it says nothing yet about the rest of a developer’s week. Three randomized field experiments covering 4,867 developers found about 26 percent more completed tasks, with large uncertainty around the estimate6, which is evidence about output in real work. And a vendor’s analysis of about 800 developers, not peer reviewed, found no significant change in cycle time and a 41 percent higher bug rate among those with access7, which touches the link that matters to a customer: whether finished, working software arrives sooner.
The lesson is not that one study is right and another wrong. They measure different links, in different firms, with different methods. A leader who quotes the 55.8 percent as the business case is using capability evidence to answer an outcome question. A leader who quotes the 41 percent to block the investment is making the same mistake in reverse. The honest position is that the external evidence makes the bet worth testing, and only your own baseline can show whether it pays.
Story: a careful rollout, read as a post-mortem
In January 2025, a team at ZoomInfo, a business-to-business data and software company, published an account of how the company rolled out an AI coding assistant to its developers8. It is worth reading as a post-mortem, not because anything went wrong, but because it is unusually candid about what a well-run deployment can and cannot show.
The rollout was staged. In July 2023, five engineers tried the tool for a week. The company then recruited 126 engineers, about a third of its developers, chosen to span specialties, seniority, locations and technology stacks, for a two-week trial in August. Company-wide rollout began in September, with licenses released in stages over several months, and eventually reached more than 400 engineers.
The measurements were real and specific. Developers accepted 33 percent of the suggestions the tool showed them and 20 percent of the suggested lines of code. Satisfaction was 72 percent. Of the developers who answered the trial survey, 90 percent said the tool cut the time their tasks took, with a median reduction of 20 percent.
Now apply the test. A 20 percent cut in task time is the same thing as being 25 percent faster, the very number in the memo. Here it is a self-reported median, from people who chose to answer a survey, about the tasks they did with the tool. Acceptance rates and satisfaction are leading signals of adoption, and good ones. What the paper does not yet show is whether features shipped sooner, defects fell, or hiring plans changed. The authors say so plainly: they were watching their regular delivery metrics and planned to report on causality in a later paper.
That is the point of the story. ZoomInfo did nearly everything this module recommends at the start: it piloted before scaling, chose a representative trial group, measured more than log-ins and published its limits. Even so, the evidence it had reached the adoption and perception links of the chain, not the outcome and capture links. Those links do not arrive by themselves. They have to be designed in, with a baseline for the outcome and an owner for the freed time, before the rollout starts.
Evidence picks the verb
A test is useful only if it changes what happens next. AI Value Scorecard set out the verbs: scale, adjust, continue, pause or stop, each chosen by a rule agreed before the numbers arrived. Running the value test is how you know which rule applies. If the outcome has moved and the economics hold, scale. If the tool works and the outcome has not moved, the fix is usually in the workflow or the capture, not the model, so adjust. If there is no credible path to the outcome, stop, and treat the freed budget as a win. How to rank and stage many such bets at once is the subject of Module 09.
The test reduces to four questions that travel to every AI review, whoever is presenting and whatever the tool.
The loop matters as much as the questions. Each time a forecast is compared with its result, the organization learns how optimistic its own proposals tend to be, and the next memo arrives with a more believable number.
What this means for leaders
Three habits carry this module into practice. First, when someone says “the value”, ask which number they mean: potential, expected or realized, in the terms of Building the AI Business Case. Many arguments about AI value are two people quoting different numbers. Second, walk every large claim through definition, scope, adoption, capture and cost before debating it, and find the step that loses the most. Third, read each piece of evidence for the link it covers, and insist that the outcome and capture links are designed in before launch, because they cannot be added afterwards.
None of this requires technical expertise. It requires knowing what evidence to demand, and continuing to demand it after the launch.
Check yourself
- Twenty-five percent faster frees a quarter of the time.
- Choosing value over hype means being cautious about AI.
- A high acceptance rate for AI suggestions shows the business has gained.
- Internal teams create inflated benefit claims as readily as vendors do.
- Overstating a firm’s use of AI has already drawn penalties from US regulators.
- If a tool raised output in another firm’s experiment, it will pay off in yours.
Reflection: find the missing link
What comes next
This module answered two questions: why invest in AI, and how to tell whether an investment creates value. A company can measure each bet well and still place its bets in the wrong direction. The next module asks what the bets are for. It begins with What Is an Enterprise AI Strategy?, which treats strategy not as a list of AI projects but as a short set of hard choices traced back to the business.
Laws referenced
Not legal advice. Laws change; verify before relying on this, and consult counsel for decisions.
US enforcement against false AI claims ("AI washing") · US - federal
Federal securities antifraud rules and Investment Advisers Act Marketing Rule (SEC); FTC Act Section 5 (unfair or deceptive practices)
There is no AI-specific federal statute, but existing law already applies to what companies say about AI. Regulators have penalized firms that overstated their use or capability of AI to investors (SEC) and to consumers (FTC). Claims about AI in marketing, investor materials and product descriptions need the same substantiation as any other claim.
- 2024-03-18 — SEC's first AI-washing cases: Delphia and Global Predictions settle for USD 400,000 in total civil penalties
- 2024-09-25 — FTC launches Operation AI Comply, with five actions over deceptive AI claims and uses
Last verified 2026-10-08
References
- HM Treasury. Supplementary Green Book Guidance: Optimism Bias. GOV.UK. 2013.
- U.S. Government Accountability Office. IT Dashboard: Selected Agencies' Investment Ratings Fail to Fully Consider Risks (GAO-27-108416). GAO. 2026.
- US Securities and Exchange Commission. SEC Charges Two Investment Advisers with Making False and Misleading Statements About Their Use of Artificial Intelligence (Press Release 2024-36). SEC. 2024.
- US Federal Trade Commission. FTC Announces Crackdown on Deceptive AI Claims and Schemes (Operation AI Comply). FTC. 2024.
- Sida Peng, Eirini Kalliamvakou, Peter Cihon and Mert Demirer. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv (2302.06590). 2023.
- Zheyuan (Kevin) Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng and Tobias Salz. The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. Management Science. 2026.
- Uplevel Data Labs. Does GenAI Improve Software Developer Productivity?. Uplevel. 2024.
- Gal Bakal, Ali Dasdan, Yaniv Katz, Michael Kaufman and Guy Levin. Experience with GitHub Copilot for Developer Productivity at Zoominfo. arXiv (2501.13282). 2025.
Further reading
- HM Treasury. Supplementary Green Book Guidance: Optimism Bias. GOV.UK. 2013.
- Zheyuan (Kevin) Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng and Tobias Salz. The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. Management Science. 2026.
Sources last verified 2026-10-08.