AI Academy · Book
Executives & Directors · Module 03 · Chapter 009

Leading vs Lagging AI Metrics

The numbers that prove AI value arrive last, often a year or more after the money is spent. Leaders need earlier signals that genuinely predict those results, a chain that connects the two, and guardrails that catch gains which quietly cost more than they save.

≈ 14 min read

After this chapter you can

  • Distinguish leading, intermediate and lagging AI metrics and the question each one answers.
  • Trace an AI initiative through the metric chain from adoption to net value.
  • Test whether a leading metric predicts the next link or is only a count that flatters.
  • Use disagreements between early and late signals to find the link that is breaking.
  • Pair every primary metric and guardrail with an early and a late reading, and agree the timing before launch.

In the late summer of 1937, with the US economy sliding back into recession, Treasury Secretary Henry Morgenthau Jr. put a practical question to the National Bureau of Economic Research. He asked for a list of statistical series that would best show when the recession would end. Wesley Mitchell and Arthur Burns answered with a list of the most reliable indicators of recovery, chosen for their record of consistent timing in past cycles and published in May 1938. It was the first set of leading, coincident and lagging indicators1. The idea outlived the crisis. The Commerce Department compiled the composite indexes for decades, and after a bidding process in 1995 the Conference Board became custodian of the official leading, coincident and lagging indexes, publishing them on its own from January 19962.

An AI program asks Morgenthau’s question in miniature. The numbers that prove value, such as unit cost, revenue and margin, arrive last, often a year or more after the money is spent. Picture the steering committee in the program’s second year. Looking only at the cost line, it would conclude that AI had failed. Looking only at how widely AI was being used, it would conclude the opposite. Both conclusions would be premature, because each rests on a single number, and the two numbers move at different times. Managing through that gap takes two kinds of measure and a way to connect them.

Where leading and lagging come from

Two details of the 1938 list matter for anyone borrowing the idea. The indicators earned their place by a track record of moving first, not by being easy to count. And economists kept testing them afterwards: most leaders continued to lead, but some changed their timing and had to be reclassified1.

Management borrowed the idea in the 1990s. Robert Kaplan and David Norton argued that every scorecard needs both outcome measures, which they called lag indicators, and performance drivers, the lead indicators. “Outcome measures without performance drivers do not communicate how the outcomes are to be achieved,” they wrote. The reverse fails too: drivers such as cycle times and defect rates, tracked without outcomes, can deliver operational improvements that never show up as business results3.

Manufacturers know a third version from the shop floor. The US Occupational Safety and Health Administration treats injury and illness rates as lagging indicators, because they record harm already done. Leading indicators, such as the time it takes to respond to a reported hazard, show whether the safety system is working before anyone is hurt. Its guidance puts the division of labor plainly: leading indicators drive change, and lagging indicators measure whether it worked4.

Leading and lagging indicators came from economics in 1938, management scorecards in 1996 and workplace safety in 2019; AI programs need the same discipline.1938EconomicsMitchell and Burns:indicators of recovery1996ManagementKaplan and Norton: driversand outcomes2019SafetyOSHA: drive change, thenmeasure itNowAI programsThe samediscipline applies
Figure 3.9.1 Three fields reached the same rule: steer by what moves first, judge by what moves last.

Three kinds of metric, three questions

The core idea of this chapter fits in one sentence: leading metrics steer the journey, lagging metrics judge the arrival, and intermediate metrics show the road in between. Each answers a different management question, and each moves at a different speed.

Leading metrics show what is happening now within days, intermediate metrics show whether the work improves within months, and lagging metrics show whether the result arrived within quarters or years.KindQuestion it answersAI examplesTypical paceLeadingWhat is happening now?Use on target work,acceptance, overridesDays to weeksIntermediateIs the work improving?Task time, cycle time, reworkWeeks to monthsLaggingDid the business result arrive?Unit cost, revenue, retention,net valueQuarters to years
Figure 3.9.2 Each kind answers its own question. The paces are a rough guide; your business sets the real cadence.

Leading metrics tell you whether the initiative is moving in the intended direction: whether people use the tool on the work it was built for, how often they accept its suggestions, how often they override them. You can act on them this week. Intermediate metrics are operational: task time, end-to-end cycle time, throughput, rework and first-time quality. They usually explain why the final result is or is not moving. Lagging metrics are the business results the investment was made for: unit cost, revenue, margin, retention and, finally, value net of what the initiative costs. They move last, and they are the proof.

A metric’s kind depends on the question you ask of it, not on the metric itself. Cycle time is intermediate when the goal is lower cost, and it is the lagging outcome for a team whose goal is speed. Label each metric by the role it plays in this initiative.

Read the metrics as one chain

The three kinds become useful when they are linked. An AI initiative rests on a chain of claims: people will use the tool, using it will improve the task, a better task will improve the workflow, a better workflow will move a business result, and the result will be worth more than it cost. From AI Capability to Business Outcome treated each arrow in that chain as an assumption to be tested. Metrics put a reading on each arrow.

The metric chain runs from adoption to task, workflow, business result and net value, and each link needs its own measurement.AdoptionUsed onthe targetwork?TaskIs eachtask better?WorkflowIs the wholeprocessbetter?BusinessDid theresultmove?ValueWorth morethan itcost?Every link needs its own reading
Figure 3.9.3 Leading readings sit at the left, lagging ones at the right. A break anywhere stops value reaching the end.

Reading the chain also disciplines the conversation in a review. A rise in adoption is good news about the first link and no news about the others. As Measuring AI Business Value argued, activity is not value; the chain shows exactly how far the activity has traveled. How to set the baseline for each reading, and how to attribute a change to AI rather than to something else, was the work of Baselines, Metrics and Measurement.

A leading metric must earn its place

Mitchell and Burns did not choose indicators because they were available. They chose them because their timing had held up across past cycles. AI programs often skip that test. The usual early numbers are whatever the tool’s admin console reports: licenses assigned, registered users, prompts sent, summaries generated.

Eric Ries called numbers like these vanity metrics. Cumulative counts can only climb, so they flatter the team whatever is happening underneath, and they do not tell anyone what to do next. He contrasted them with actionable metrics, which show cause and effect, and he asked that reports also be accessible, so people can understand them, and auditable, so people can trust them5.

Vanity counts such as licenses and prompts rise by themselves; leading metrics such as share of target work and suggestions that survive review predict the next link.Counts that flatterLicenses assignedRegistered usersPrompts sentSummaries generatedCandidate leading metricsShare of target work done with AISuggestions that survive reviewOverride rateRework on AI-assisted outputTest: if it doubled, what should move next, and when?
Figure 3.9.4 The left column rises by itself. The right column holds candidates: each must still show, over a quarter or two, that it leads the next link.

A practical test sorts the two. Ask: if this number doubled tomorrow, what should improve next, by roughly how much, and by when? If nobody can answer, the number is a count, not a leading metric. If someone can, write the answer down, because it is a prediction you can check. After a quarter or two, look back: did movements in the leading metric come before movements in the intermediate and lagging ones? A leading metric that never leads should be demoted, just as economists reclassified indicators that stopped behaving.

Find the link where the chain breaks

Some of the most instructive moments come when the leading and lagging readings point in different directions. Software delivery is the well-measured example, examined in *AI in Software Engineering* in Module 5: code-level measures improved with AI adoption while delivery stability slipped. A team that watched only its intermediate measures would have reported success; a team that watched only the outcome would have seen a decline and not known why. Watching both, linked, turns a disagreement into a diagnosis.
If adoption is up but task measures are flat, check fit and training; if tasks improved but the workflow did not, look downstream for queues, rework or quality problems.Adoption up. Didtask measuresimprove?NoTask unchangedCheck fit, training andtarget workYesTask improvedWorkflow flatLook downstream: queues,rework, qualityWorkflow betterCheck timing and thebusiness result
Figure 3.9.5 Walk the chain from the left. The first link that fails to move is where to spend management attention.

The decision rules are simple. If adoption is up but the task has not improved, people are opening the tool without it helping: look at fit, training and whether it is being used on the right work. If the task has improved but the workflow has not, the gain is being lost downstream, in a queue, in rework or in quality problems that send work back. If the workflow has improved but the business result has not yet moved, check whether enough time has passed before concluding anything. The point is not more metrics. It is knowing which link is breaking.

Guardrails: what must not get worse

Baselines, Metrics and Measurement introduced guardrails: the measures that must not deteriorate while the primary metric improves. Many guardrails share a gap: they are lagging. Delivery stability, customer complaints and incident counts move only after the damage has shipped. Workplace safety faced the same problem: injury rates record harm already done, so OSHA asks employers to track leading indicators alongside them4.

The remedy is to give guardrails the same treatment as primary metrics: a leading reading and a lagging one. Kaplan and Norton’s warning applies with extra force to AI, because a tool that speeds up one step makes it easy to improve a driver while the outcome quietly suffers3. Teams optimize what leaders put in front of them, so the guardrail has to be in front of them too, early enough to act on.

A two-by-two of leading and lagging metrics for primary outcomes and guardrails; the early-warning guardrail quadrant is the easiest to leave empty.PrimaryGuardrailRoleLeadingTiming · LaggingSteerUse on target work, task timeProveUnit cost, revenue, net valueWarn earlyOverrides, rework, first-time qualityConfirm no harmStability, complaints, incidents
Figure 3.9.6 Programs tend to fill the top row first. The bottom-left cell, early warnings on quality, is the easiest to leave empty.

Keep the set small. One primary outcome, a leading reading for each link of the chain, and two or three guardrails, each with an early and a late reading, is enough for most initiatives. How those lines fit on one page for a steering committee is the subject of AI Value Scorecard, later in this module.

Set the clock before you start

Lagging metrics are slow for a structural reason: the payoff waits on investments in new processes, skills and ways of working, which is why measured productivity can dip before it rises, the J-curve that Quick Wins, Strategic Bets and Transformation Initiatives examines in Module 9.

The practical consequence is to agree the timing in advance. For each link in the chain, write down the expected direction, the rough size and the date by which it should move. Then agree which leading readings would justify patience while the lagging ones are still flat, and which would trigger intervention. Without that agreement, a slow lagging metric invites two opposite mistakes: cancelling a sound program in its dip, or funding a broken one because its adoption chart keeps rising. And keep reading the leading and intermediate instruments through the dip, because that is when they show where the adjustment is failing.

Story: practice scores rose, exam scores fell

In the autumn of 2023, researchers from the University of Pennsylvania worked with a large high school in Turkey to test GPT-4 math tutors on nearly 1,000 students in about fifty ninth- to eleventh-grade classes. Each classroom was assigned to one of three arms. One practiced with GPT Base, a tutor built to mimic a standard chat assistant. A second practiced with GPT Tutor, which drew on teacher-written solutions and common mistakes and was prompted to give hints rather than answers. A control group practiced with course books and notes. Four 90-minute sessions covered about 15 percent of the semester’s math, and each session ended with an exam that students took on their own, without any AI6.

The study, published in 2025, took two readings that many AI programs never separate. The first was the work done with the tool: practice scores. They rose sharply, by 48 percent with GPT Base and 127 percent with GPT Tutor, compared with the control group. The second was the result the school exists for: what students could do on their own afterwards. On the unassisted exam, students who had practiced with GPT Base scored 17 percent lower than the control group. Students who had used GPT Tutor scored about the same as the control group, despite their far better practice scores6.

Practice scores rose 48 percent with GPT Base and 127 percent with GPT Tutor, but on the unassisted exam GPT Base students scored 17 percent worse and GPT Tutor students no better than the control group.ArmPractice with AIExam on their ownGPT Base48% better17% worseGPT Tutor (hints only)127% betterNo measurable difference
Figure 3.9.7 Both arms against a control group with books and notes. The reading taken during the work pointed the wrong way.

A program judged by the first reading would have reported a triumph, and the better the practice scores, the louder the triumph. Yet the arm with the biggest practice gain produced no learning gain at all, and the plain chat assistant did measurable harm.

The early warning was there to be read. The researchers logged every interaction. Students used GPT Base as a crutch: they asked for and copied solutions. In the first practice session, 67 percent of students’ first messages on a problem with GPT Base repeated the question or asked for the answer, against 37 percent with GPT Tutor. GPT Base also gave a correct answer only about half the time. And the students’ own views flattered both tools: those using GPT Base did not believe they had learned less, and those using GPT Tutor believed they had done significantly better on the exam, when they had not6.

The guardrail that limited the damage was a design choice, not a dashboard. GPT Tutor’s teacher-written hints and its refusal to hand over answers largely removed the harm, though it still did not produce a gain. Three things in this study are often missing from AI programs, and each maps onto this chapter. There was a lagging reading of the outcome that mattered, taken without the tool, rather than a reading of the work done with it. There was an early guardrail reading, the share of requests that simply asked for the answer, available in the first practice session. And there was a reason to distrust self-reports, which rose while the result fell.

What this means for leaders

Treat adoption as evidence about the first link of the chain and nothing more. Ask for a reading on every link, insist that each leading metric comes with a written prediction of what it should move and when, and look back to see whether it did. Give every guardrail an early reading as well as a late one. Agree the clock before launch, so that a slow lagging metric is judged against expectations rather than impatience. And keep the measurement routines running through the dip, because that is when they are worth the most.

Check yourself

  1. High adoption of an AI tool shows that it is creating business value.
  2. A metric can be leading for one initiative and lagging for another.
  3. Lagging metrics are enough, as long as you are patient.
  4. A leading metric should be checked against history to see whether it really moves first.
  5. A count that can only rise, such as registered users, makes a good leading metric.
  6. A guardrail needs an early reading as well as a late one.

Reflection: the last dashboard you saw

What comes next

A chain of leading and lagging metrics tells you whether an initiative is on its way and whether it arrived. Before the money is committed, leaders need something else: a credible argument that the arrival is worth the journey, with benefits, costs and risks stated in a form finance will accept. The next chapter, Building the AI Business Case, turns these measures into that argument.

References

  1. Geoffrey H. Moore. The Forty-second Anniversary of the Leading Indicators. NBER, in Business Cycles, Inflation, and Forecasting (Ballinger), chapter 24, pp. 369-400. 1983.
  2. The Conference Board. Business Cycle Indicators: Frequently Asked Questions. The Conference Board. 2026.
  3. Robert S. Kaplan and David P. Norton. Linking the Balanced Scorecard to Strategy. California Management Review, 39(1), 53-79. 1996.
  4. US Occupational Safety and Health Administration. Using Leading Indicators to Improve Safety and Health Outcomes (OSHA 3970). US Department of Labor. 2019.
  5. Eric Ries. The Lean Startup. Crown Business. 2011.
  6. Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı and Rei Mariman. Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences 122(26), e2422633122. 2025.

Further reading

Sources last verified 2026-10-10.