AI Academy · Book
Executives & Directors · Module 03 · Chapter 007

Baselines, Metrics and Measurement

An AI result is the difference between two numbers. AI value is the difference between what happened and what would have happened without AI. Fix the baseline and the definitions before launch, estimate the unseen world with a comparison that fits the decision, and agree in advance which result changes the decision.

≈ 17 min read

After this chapter you can

  • Explain why an AI value claim needs a baseline measured before launch, by the same method as the result.
  • Write a metric definition with numerator, denominator, population, period, source and owner.
  • Name the forces besides AI that move a metric, including regression to the mean.
  • Match a comparison method to the size and reversibility of the decision.
  • Separate proxies from outcomes and agree a decision rule before results arrive.

From the early 2000s, safety cameras went up at more than 4,000 sites across Great Britain. Collisions and casualties at those sites then fell substantially, and the national four-year evaluation credited much of the fall to the cameras. In 2010 the road-safety researcher Richard Allsop reviewed the evidence for the RAC Foundation. He concluded that the cameras worked, and that the evaluation’s own percentages had “in all probability” exaggerated how well1.

Both conclusions are true, and the reason is the subject of this chapter. Many cameras had been installed at sites that had just recorded untypically high numbers of casualties. A bad few years on a stretch of road is partly chance, and chance does not repeat on schedule, so casualties at those sites would have fallen anyway. Casualties were also falling nationally, partly because cars protected their occupants better. After allowing for both, Allsop judged that the cameras prevented around 1,000 deaths and serious injuries in the year to March 2004: a large, real effect, but smaller than the raw before-and-after figures implied1.

Cameras at more than 4,000 British sites prevented an estimated 800 to 1,300 deaths and serious injuries a year, fewer than the raw figures implied.4,000+Camera sitesGreat Britain, four-year evaluation800 to 1,300Fewer deaths and serious injuriesYear to March 2004, after trend and regression tothe meanSource: Allsop, RAC Foundation · 2004
Figure 3.7.1 The cameras worked. The raw before-and-after figures still overstated how well, because sites were chosen after bad years.

Now replace the camera with an AI pilot. Pilots usually go where the pain is worst: the office with the backlog, the region with the bad quarter, the team with the complaints. A site chosen because its numbers are unusually bad will tend to improve whatever is installed there. A pilot that reports a 20 percent improvement at such a site has reported a difference. It has not yet reported what AI did.

Impact is a comparison, not a difference

Every claim that AI created value contains a hidden comparison: compared with what would have happened without it. That alternative is the counterfactual. It can never be observed directly, because the same team cannot live through the same quarter both with and without the tool. Measurement is the work of estimating it credibly, and of agreeing in advance what the estimate must show.

Five steps carry that work, and many AI value reports stop after the fourth.

Five steps - baseline, target, change, measure and compare - with impact defined as the observed outcome minus the counterfactual.BaselineWhere thework standsbefore launchTargetThe gainthat justifiesthe costChangeWhat AIalters inthe workMeasureWhat moved,measured thesame wayCompareWhat wouldhave movedanywayImpact = observed outcome minus the counterfactual
Figure 3.7.2 The measurement chain. A report that shows the after-state and calls it proof has skipped the last step.

The baseline says where the work stands before anything changes. The target says what improvement would justify the investment, agreed before launch. The change names exactly what AI alters in the work, so the measurement looks in the right place. The measurement records what actually moved, defined and collected the same way as the baseline. The comparison estimates what would have moved anyway, and subtracts it.

What Does AI Value Actually Mean? made “compared with what?” a reflex; this chapter turns the answer into evidence.

A baseline is a measurement, not a memory

A baseline is the measured starting point against which change will be judged. It is not the number someone remembers from last year’s board pack, and it is not a figure reconstructed after the results look good. It is measured before launch, over a period long enough to be representative, by the same method that will measure the result.

An illustration runs through the rest of the chapter. Suppose a shipping line’s export-documentation team prepares the paperwork for each container booking: the bill of lading, the customs filings, and the corrections when a shipper’s data is wrong. Over the last full quarter the team spent 30 minutes of handling per shipment file, at an all-in cost of 50 per file, and 2.0 percent of issued bills of lading needed a later amendment. Those three numbers are the baseline. The target, agreed before an AI drafting assistant goes live, is to cut cost per file by at least 15 percent, from 50 to 42.5, within two quarters, without the amendment rate rising.

Three properties make a baseline trustworthy. It is representative: a quarter that contains the peak export season will flatter any result measured in a quiet one. It is recent: a baseline from before a reorganization describes a different operation. And it is defined in writing, because measurement disputes often turn out to be arguments about definitions that nobody wrote down.

A one-page definition of cost per shipment file - numerator, denominator, population, period, source and owner - agreed before launch.FieldCost per shipment fileNumeratorActive handling cost, including reworkDenominatorShipment files completed in the periodPopulationAll export files, including the hard onesPeriodLast full quarter before launch, then each quarterSourceBooking system and time records, unchangedOwnerHead of documentation, signed off by finance
Figure 3.7.3 Illustrative, for the shipping line: a metric definition written before launch. Every field is a place where the result can drift if it is left blank.

If the pilot team quietly excludes the hardest files (“those go to the senior team anyway”), cost per file falls without any change in the work. If the source system changes in the middle of the pilot, the trend breaks for reasons that have nothing to do with AI. Writing the definition down before launch is what stops the goalposts from moving.

The denominator deserves particular care, because the same change can be reported in two honest ways that sound very different. If the amendment rate falls from 2.0 to 1.5 percent, that is a fall of 0.5 percentage points and also a fall of 25 percent relative to the baseline. Both are correct. A report that does not say which one it uses invites the reader to hear the larger. Speed needs the same discipline: as AI and Workforce Productivity showed, work that becomes 25 percent faster takes 20 percent less time, so the two figures are not interchangeable.

When there is no historical data, measure a baseline prospectively rather than skip it; the box later in this chapter shows how.

What else moves the number

The simplest measurement compares before with after. It is useful for monitoring and nearly useless for attribution, because many things besides AI move a business metric at the same time.

Season, trend, regression to the mean, other changes, novelty and learning, and self-selection can all move a pilot metric besides AI.-20%Cost per file: whatelse moved it?SeasonTrendRegressionNew projectsNoveltyVolunteers
Figure 3.7.4 Six things besides AI that move a business metric; the 20 percent fall is the illustrative shipping-line result. Each is a reason to compare.

Season and trend. Export volumes, claims and sales follow calendars, and many metrics drift for reasons that have nothing to do with the pilot, as British road casualties fell with safer cars. Regression to the mean. Sites chosen because their numbers were unusually bad tend to improve on their own, as the speed cameras showed. Other changes. A new customs portal, a pricing change or a staffing change in the same quarter acts on the same number. Novelty and learning. Early results can be inflated by enthusiasm or depressed by the learning curve, and experienced experimenters watch for effects that fade or grow over the first weeks2. Who volunteered. The first teams to adopt a tool are often the best-run teams, so comparing adopters with non-adopters measures the managers as much as the tool.

Each one is a reason to estimate what would have happened anyway.

Estimating the world without AI

Return to the documentation team. Two quarters after launch, cost per file in the pilot office has fallen from 50 to 40. The pilot team reports a 20 percent saving, comfortably past the 15 percent target. But the shipping line runs a second documentation office on the same systems, and it did not get the assistant. Its cost per file fell from 50 to 45 over the same months, because the customs authority’s new filing portal had cut rework everywhere.

From a baseline of 50 per file, the pilot reached 40 and the comparison office 45, so AI's estimated impact is about 5 per file, 10 percent rather than 20.Baseline: 50 per fileWith AI - what we see40 per fileWithout AI - estimated from the second office45 per fileESTIMATED AI IMPACTAbout 5 per fileAbout 10%, not 20%
Figure 3.7.5 Illustrative. The second office estimates the unseen world, and AI’s share of the saving is about half the raw figure.

Two figures now need separate names. The observed change is what moved in the pilot office: 20 percent. The incremental effect is what moved because of AI: the pilot’s change minus the comparison office’s. If the second office is a good estimate of the unseen world, AI’s incremental effect is about 5 per file, roughly 10 percent of the baseline, not 20. The assistant worked. It did about half of what the before-and-after figure claimed, and its incremental result falls short of the target that justified it.

The conclusion rests on how comparable the two offices are: the same systems, customs regimes, mix of shippers and pre-pilot trend make a credible comparison. If the second office handles mostly simple bookings, the comparison misleads in its own way. Choosing the comparison is therefore a design decision made before launch, not a search for a convenient benchmark afterward.

Match the method to the decision

There are four common ways to estimate the counterfactual. They trade cost for credibility.

Before-and-after cannot show cause; a comparison group partly can; a staggered rollout mostly can; a randomized test shows cause best and suits large bets.MethodShows causeKey assumptionExtra costUse forBefore and afterNothing else changedNoneMonitoringComparison grouppartlyThe groups would havemoved togetherLowMedium decisionsStaggered rolloutmostlyStart order is unrelatedto resultsLowPhased launchesRandomized testFewMediumLarge orhard-to-reverse bets
Figure 3.7.6 Four ways to estimate the world without AI. Credibility rises down the table; so does the need to design it in before launch.

Before and after assumes nothing else changed. It suits monitoring and small, reversible decisions.

A comparison group followed over the same period assumes that the two groups would have moved in parallel without the intervention. The estimate is the change in the treated group minus the change in the comparison group, which statisticians call a difference-in-differences. One of its most famous early uses comes from public health. In mid-nineteenth-century London, two companies piped water to the same south London districts, often down the same streets and sometimes to neighboring houses. In 1849 both drew water from the sewage-polluted tidal Thames. By the 1854 cholera epidemic, the Lambeth company had moved its intake upstream and the Southwark and Vauxhall company had not. The physician John Snow compared cholera deaths in the districts supplied by Southwark and Vauxhall alone with those where both companies competed, in both epidemics3.

Cholera deaths per 10,000 rose by 11.8 where only the polluted supply was used and fell by 45.2 where the clean supply competed, an estimated effect of about 57.Deaths per 10,000Southwark and Vauxhall onlyBoth companies1849, bothsupplies polluted134.9130.11854, Lambeth supply clean146.684.9Change+11.8-45.2
Figure 3.7.7 Snow’s comparison (1855), as reanalyzed by Coleman (2020). The supply change accounts for a fall of about 57, more than the raw 45.

In the jointly supplied districts, deaths fell by about 45 per 10,000 people. A before-and-after reading would stop there. But where no clean water was available, deaths rose by about 12 over the same period, because the 1854 epidemic was worse. The estimated effect of the cleaner supply is the difference between the two changes, a fall of about 57, larger than the raw fall3. A comparison corrects the estimate in whichever direction the world was moving. Sometimes, as with the cameras and the documentation office, it shrinks the claim. Sometimes it shows that the raw figure understated the effect.

A staggered rollout gives the tool to different teams at different times, so the teams still waiting serve as the comparison. It costs almost nothing extra, which is why some of the strongest evidence on AI at work comes from one4. Its assumption is that the order of the rollout is unrelated to how teams would have performed, which is easiest to defend when the order is set by lot or by logistics rather than by enthusiasm.

A randomized test assigns the tool by chance, to customers, cases or teams, so that the groups differ only in the tool. With a large enough sample, random assignment makes the comparison credible without assuming that two groups would have moved together. It is routine in online businesses, where randomizing customers is cheap, and its practical disciplines are well documented2.

Not every initiative needs a randomized test. The method should match the size and reversibility of the decision. A team trying a drafting assistant on its own work can rely on before-and-after and judgment. A rollout costing 5 million deserves at least a comparison group, and ideally a staggered or randomized start. Either must be designed in from the beginning, because it can rarely be added afterward.

Proxies, guardrails and the decision rule

Two more habits turn a credible estimate into a decision.

The first is to keep the proxy in its place. Many AI initiatives are tracked first through numbers that are easy to collect: usage, model accuracy, time per task, clicks. These are proxies for the outcome the business cares about, and proxies tend to separate from outcomes once people optimize them. The anthropologist Marilyn Strathern gave Goodhart’s law its best-known form: “When a measure becomes a target, it ceases to be a good measure”5. A proxy is valuable as an early signal and a health check. It should not be the number that decides.

The second is to agree the decision rule before the results arrive. A decision rule says which result leads to which action: scale, iterate, investigate or stop. It names the primary metric, the size of incremental effect that justifies the investment, and the guardrails, the measures that must not deteriorate while the primary metric improves. How guardrails pair with early and late indicators is the subject of Leading vs Lagging AI Metrics. The point here is that they are written down in advance.

Scale only if the incremental cut reaches 15 percent and amendments stay at or below 2 percent; investigate if amendments rise; otherwise iterate or stop.Incremental costcut of 15%or more?YesCheck the guardrailAmendments at or below2.0%: scaleAmendments up: investigateNoIterate or stopFix the weakest step,then retest
Figure 3.7.8 Illustrative. The decision rule for the documentation pilot, agreed before launch. The incremental result, not the raw one, enters the tree.

With that rule agreed, the documentation review is short. The observed cut of 20 percent would have cleared it. The incremental cut of about 10 percent does not. Same pilot, different decision, and nobody can accuse the reviewer of moving the goalposts, because the goalposts were set first.

A rule written in advance also protects the pilot team: a shortfall is information, not failure. And bringing finance in to write the rule, not to judge the result, settles the definitions while nobody has a stake in the answer.

Story: the better model

Suppose you run a product area at a large online travel platform. Your data scientists have built a new model to rank the hotels shown to each customer. On the standard offline test, scoring historical data it has never seen, it clearly beats the model in production. The team asks to switch it on for everyone. Do you approve?

Booking.com faced this question many times. Three of its data scientists described about 150 customer-facing machine learning applications, each validated in randomized controlled trials against business metrics such as conversion, customer-service tickets or cancellations6. They then examined 23 cases in which a new model was tested against a successful model already in production. They found no correlation between how much better the new model scored offline and how much it improved the business metric in the live trial. The Pearson correlation was -0.1, with a 90 percent confidence interval from -0.45 to 0.276.

Of about 150 models tested in live trials, 23 upgrades showed no correlation, about minus 0.1, between offline score gains and business gains.~150Models tested inlive trialsEach against a business metric23Model upgradescomparedOffline gain versus business gain-0.1CorrelationNo detectable link betweenthe twoSource: Bernardi et al., KDD · 2019
Figure 3.7.9 At Booking.com, a better offline score did not predict a better business result. Only the randomized trial could tell.

The authors offered several explanations. Some problems saturate: past a point, a better model adds nothing a customer notices. A model can also improve its proxy while the outcome suffers. One recommendation model learned to show hotels very similar to the one a customer was viewing; customers clicked, presumably to compare them, and the extra choice hurt conversion. Very accurate predictions can unsettle customers who feel the system knows too much. And a model carries costs the offline test cannot see: in a separate randomized test that slowed pages artificially, about 30 percent more latency cost more than 0.5 percent of conversion, which the authors called a relevant cost for the business6.

So the answer to the dilemma is neither yes nor no. It is: test it. Booking.com’s teams treat offline performance as “only a health check” that the model does what it should, and the decision to keep a model rests on the randomized comparison against the business metric6. The authors caution that the result comes from one company’s systems, with models that all improve on an earlier success, and should not be over-generalized. That is the right reading for an executive too. The lesson is not that accuracy never matters, as From AI Capability to Business Outcome argued in its own terms. It is that the only way to know whether it matters here is to measure the outcome against a comparison.

The case shows the whole chain at work. The baseline is the model already in production. The target is a business metric defined in advance. The comparison is randomized. The proxy, offline accuracy, is kept as a health check. Most organizations work at a smaller scale and with fewer customers to randomize, but the questions are the same.

What this means for leaders

Measurement is not a reporting task that starts after launch. Its most important decisions are made before launch: the baseline, the definition, the comparison and the rule. An executive who asks for them at the funding stage gets evidence at the review. One who asks at the review gets a before-and-after chart and an argument.

The practical posture is proportion. Small, reversible experiments need a baseline and honest judgment. Rollouts that commit real money, change many people’s work or are hard to undo need a comparison designed in from the start. In both cases, the number that decides should be the incremental outcome, not the usage, the accuracy or the raw change.

Check yourself

  1. A before-and-after comparison shows how much value AI created.
  2. A pilot placed at the worst-performing site will tend to look better than the tool deserves.
  3. A comparison group can only make an AI result look smaller.
  4. A fall from 2.0 to 1.5 percent is both 0.5 percentage points and 25 percent.
  5. A model that scores better offline will improve the business metric.
  6. Every AI initiative needs a randomized experiment.

Reflection: the last number you were shown

What comes next

The documentation office now has a credible estimate: the assistant saves about 5 in every 50 of cost per file, and handling time is down too. But a minute saved on a file is not yet a minute the business can use. The next chapter, Productivity vs Realized Capacity, asks what happens to saved time, and what it takes to turn it into capacity that shows up in the accounts.

References

  1. Richard Allsop. The Effectiveness of Speed Cameras: A Review of Evidence. RAC Foundation. 2010.
  2. Ron Kohavi, Diane Tang and Ya Xu. Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press. 2020.
  3. Thomas S. Coleman. John Snow and Cholera: Revisiting Difference-in-Differences and Randomized Trials in Research. University of Chicago, Harris School of Public Policy (working paper). 2020.
  4. Erik Brynjolfsson, Danielle Li and Lindsey Raymond. Generative AI at Work. The Quarterly Journal of Economics 140(2). 2025.
  5. Marilyn Strathern. 'Improving ratings': audit in the British University system. European Review 5(3): 305-321. 1997.
  6. Lucas Bernardi, Themistoklis Mavridis and Pablo Estevez. 150 Successful Machine Learning Models: 6 Lessons Learned at Booking.com. Proceedings of KDD '19, ACM, pages 1743-1751. 2019.

Further reading

Sources last verified 2026-10-08.