AI Value Scorecard
An AI value scorecard is one page that ends in a decision: scale, adjust, continue, pause or stop. It sets realized value against expected value, marks how sure each figure is, and applies rules agreed before the numbers arrived. A page that does not end in a verb, an owner and a date is a report, however green it looks.
After this chapter you can
- Distinguish an executive scorecard from the operational and business dashboards beneath it.
- Build a one-page scorecard whose eight lines each answer one of four questions, with a named owner.
- Calculate a realization rate on a countable, like-for-like basis and read it alongside ROI.
- Mark the confidence of each value figure and arrange an independent check of the page.
- Apply pre-agreed rules to choose scale, adjust, continue, pause or stop.
On 7 October 2026 the US Government Accountability Office published its latest check on a public scorecard. Every major federal IT investment carries a risk rating from its agency’s chief information officer, and the ratings are published on the government’s IT Dashboard so that anyone can see which projects are in trouble. GAO’s auditors assessed 53 of those investments themselves, from the same project records. Their view matched the official rating in 27 cases. In 24 they found more risk than the rating showed, and in only two did they find less. Twenty-one of the ratings had not been updated on the agencies’ own schedule1.
Ten years earlier, the same exercise on 95 investments had found more risk than the official rating in 60 of them, 63 percent, and less in 132. A decade of public scrutiny had not changed the direction of the error. The ratings were not lies. They were judgments made by the people closest to the work, without the evidence shown beside them and without anyone outside the project obliged to check them, on spending of more than 100 billion dollars a year.
Many AI programs report upward in the same way. Usage is up, accuracy is on target, response time is within limits, training is complete. Every line is green and every line is true, and a sponsor who asks “should we keep funding this?” still has nothing to decide with. The scorecard exists to close that gap.
The core idea
An AI value scorecard is a one-page executive view of a single AI initiative that ends in a decision. It is the artifact this module has been building toward. Building the AI Business Case produced the expected value, AI ROI and Value Realization measured what was delivered and explained the shortfall, and Total Business Impact decided which benefits may be counted. The scorecard puts the result of all three on one page, in an order that forces a choice.
The page answers four questions in sequence. Does the initiative matter: is it tied to a strategic priority and a material problem? Is it working: is the business outcome moving, not just usage? Is the value credible: how strong is the evidence behind the number? Is the risk acceptable: are quality and safety holding? Then it names a verb, an owner and a date.
The order matters. An initiative that does not serve a priority should not reach the value discussion, however well it performs, and a large value figure resting on thin evidence should not reach the decision without its confidence mark.
Above the dashboards, not instead of them
The idea of a short scorecard is older than AI. Kaplan and Norton’s 1992 balanced scorecard argued for restraint: one report tied to strategy, limiting information overload by limiting the number of measures3. Companies more often suffer from too many measures than too few, because one is added every time someone suggests it.
An AI initiative generates measures at three layers. The operational layer holds usage, response time, accuracy, errors and model cost, the numbers a delivery team watches daily. The business layer holds the workflow outcome: cost per unit of work, cycle time, revenue, customer effort. The executive layer holds strategic fit, value, confidence, risk and the decision. The scorecard is the top layer only.
A dashboard shows what is happening. A scorecard adds a target, a judgment against it and a decision. If the top page shows usage and model cost, it is a dashboard wearing a scorecard label. Those numbers still matter, and Leading vs Lagging AI Metrics explained how early operational signals predict later outcomes, but they belong one layer down.
Eight lines, one owner
The page has eight lines, and each earns its place by answering one of the four questions.
The first two lines keep the page anchored in the business rather than the technology. The outcome is what changes, such as cost per delivery, never “deploy the assistant”; the solution is not the outcome. Guardrails are the quality measures that must hold while the outcome improves, as Leading vs Lagging AI Metrics defined them. Cost is shown in full and against plan, because a benefit that arrives on time with a cost that has quietly doubled is a different decision.
Each line has one named owner, usually the person who can change it, and the executive sponsor owns the whole page. That last assignment is often missed. When everyone owns a line and nobody owns the outcome, the page describes the initiative accurately and never produces a decision.
Expected against realized
The most important pair on the page is expected value against realized value, and the number that joins them is the realization rate: realized benefit divided by expected benefit, for the same period and on the same basis. In an illustrative case, if a business case expected 1 million a year and the initiative is delivering 630,000 a year, the realization rate is 63 percent.
Three habits keep the rate honest. Compare like with like: a full-year expectation against a full-year run rate, not against a partial first year. Use only the benefit that Total Business Impact would let you count, so that the realized figure cannot be inflated by relabeled indicators. And never let the rate stand alone. A rate below expectation is a question, not a verdict, and AI ROI and Value Realization showed how to answer it by splitting the gap into adoption, benefit per use, conversion and cost variances, each with its own owner. The scorecard carries the result of that split in one line, so the discussion is about causes rather than about whether the number is bad.
The same comparison applies to cost and to every other forecast on the page. Recorded over many initiatives, forecast against actual becomes the organization’s own evidence about how optimistic its business cases tend to be, which is a better guide to the next case than any supplier’s figure.
The realization rate is not ROI. An initiative can realize 63 percent of its expected benefit and still earn a healthy return if its costs are low, or realize 90 percent and lose money if its costs have grown. The page shows both.
Show how sure you are
Not every figure on the page deserves the same trust, and a common failure of executive reporting is to print a forecast in the same font, color and certainty as a measured fact.
The fix is calibrated language. The Intergovernmental Panel on Climate Change attaches a confidence level to each key finding, judged from how much evidence there is, how good it is and how far independent lines of evidence agree4. The same two questions work for a business figure: how much good evidence stands behind it, and do independent sources agree?
A scorecard needs only three levels. Low confidence means the figure rests on an external benchmark or an assumption: good enough to justify exploring, not to justify scaling. Medium means the organization’s own pilot data supports it, but from a small, early or self-selected group. High means the effect has been confirmed internally, ideally against a comparison group that did not get the tool, using the methods Baselines, Metrics and Measurement described.
Confidence differs by line. A team can be highly confident that the system works and have only low confidence in what it is worth. The mark goes beside each figure, not once at the top of the page.
Who marks the page
The GAO findings point to a second design choice: who assigns the ratings. The people closest to an initiative know the most about it and are the least likely to see its risks, which is why their marks drift optimistic even when made in good faith. The United Kingdom’s major projects system builds in a correction: where the central authority has independently assured a project in the past six months, its rating replaces the owner’s, and a red rating is read as a call for action, not a prediction of failure5.
An AI scorecard can borrow both habits. The initiative’s owner drafts the page, and someone without a stake in the result, often finance, checks the value and confidence lines before the page reaches the sponsor. Each rating is then read as a prompt for a decision, not as a grade.
The same logic argues against a single overall score. It is tempting to average the eight lines into one number or one color, but the standard handbook on composite indicators warns that such indices can send misleading messages6. A strong value line and a weak risk line do not average out to amber; they describe a trade-off the sponsor has to see.
Rules before numbers
The page ends with a verb, and the rules that choose the verb should be written before the results arrive. Baselines, Metrics and Measurement made this point for pilots: agree the decision rule in advance and nobody can be accused of moving the goalposts. The scorecard applies the same discipline to a live initiative, with five verbs.
The thresholds are illustrative; a high-risk use case may need a higher bar for scaling, and a cheap experiment a lower one. What matters is that the rules exist, that finance has agreed them, and that they are applied as written. The stop rule rests on the arithmetic of AI ROI and Value Realization: money already spent does not count, so the test is whether the value still to come exceeds the cost still to be spent.
An early initiative uses a lighter version of the page: the problem, the hypothesis, the evidence so far and what it must learn before the next review. The page grows as the evidence grows. Comparing many initiatives side by side, and deciding which to fund, is portfolio work for Module 09.
Story: a sponsor’s day with the page
The case that follows is an illustrative composite. The company is unnamed and the figures are invented to make the rules visible.
The sponsor is head of operations at a regional parcel and pallet carrier with twelve depots. A year earlier she backed an AI assistant that helps planners build next-day delivery routes and replan them when vans break down, drivers call in sick or customers change a time slot. The business case expected 1 million a year in fuel and driver hours, from 5 percent fewer kilometers per stop on assisted routes, with 80 percent of routes planned with the assistant. Finance had agreed the rules in the table above before launch.
At half past seven, six reports arrive, all green. At ten she visits the largest depot. Planners tell her they use the assistant for the next day’s routes on quiet afternoons but switch it off on busy mornings, when vans are going out late and the old spreadsheet feels safer. At one, finance returns the value and confidence lines it has checked. At half past four the steering committee meets and, for the first time, receives one page instead of six reports.
The page shows 72 percent of routes assisted against a target of 80, and kilometers per stop down 3.5 percent on assisted routes against a target of 5. Together those give 0.9 times 0.7 of the expected benefit: 630,000 a year, a realization rate of 63 percent. The variance line splits the 370,000 gap into 100,000 of adoption and 270,000 of benefit per route. On-time delivery and failed first-attempt deliveries are unchanged. Total cost is 300,000 a year, about 7 percent above the plan of 280,000, so the initiative is already worth 330,000 a year net. Confidence is medium: the saving is measured on the organization’s own routes, but the assisted routes are the ones planners chose, and those exclude the hardest mornings.
The rule for scaling is 70 percent, and 63 is below it. The outcome is improving and the guardrails held, so this is not a pause or a stop; the value still to come is well above the cost still to be spent. The verb is adjust. The head of planning owns morning adoption, since reaching 80 percent of routes at today’s saving would on its own bring realization to exactly 70 percent. The data lead owns an eight-week test in which two depots plan every morning route with the assistant, matched against two that do not, to raise confidence and learn whether the saving holds on hard routes. The recheck is set for the next quarter’s committee. If realization reaches the bar with the guardrails holding and confidence at least medium, the next page will say scale, to the carrier’s second region. The sponsor leaves with a verb, two owners and a date.
What this means for leaders
The scorecard changes the conversation in the steering room from “is it going well?” to “what do we do next, and who does it?” That change depends less on the template than on three disciplines around it: realized value shown against expected value on a countable basis, a confidence mark beside every figure, and rules for the verb agreed before the numbers arrive. Without the rules, a page of honest numbers still ends in whichever verb the most enthusiastic person in the meeting prefers.
Leaders also shape whether the page is trusted. Ask for an independent check of the value and confidence lines, as the UK system does for its major projects. Reward the owner who reports 63 percent with a credible fix over the owner whose page is green everywhere and says nothing. And refuse a single overall score; it is easier to read and harder to act on.
Check yourself
- A good scorecard contains every KPI the delivery teams can produce.
- The realization rate and ROI are the same number expressed differently.
- Under the illustrative rules, an initiative at 63 percent realization with guardrails holding should be scaled.
- Confidence should be marked beside each value figure, not once for the whole initiative.
- Ratings made by the people closest to a project tend to be accurate because they know it best.
- Realized value below forecast can still justify continued investment.
What comes next
The scorecard completes the toolkit of this module: how to define value, find it, measure it, build the case, test the return, count impact once and put it all on one page that forces a decision. The final chapter, Module 03 Synthesis — Choosing Value Over Hype, brings these tools together into one test for telling real AI value from a confident claim.
References
- U.S. Government Accountability Office. IT Dashboard: Selected Agencies' Investment Ratings Fail to Fully Consider Risks (GAO-27-108416). GAO. 2026.
- U.S. Government Accountability Office. IT Dashboard: Agencies Need to Fully Consider Risks When Rating Their Major Investments (GAO-16-494). GAO. 2016.
- Robert S. Kaplan and David P. Norton. The Balanced Scorecard - Measures That Drive Performance. Harvard Business Review 70(1), January-February 1992, 71-79. 1992.
- Michael D. Mastrandrea, Christopher B. Field, Thomas F. Stocker and others. Guidance Note for Lead Authors of the IPCC Fifth Assessment Report on Consistent Treatment of Uncertainties. Intergovernmental Panel on Climate Change. 2010.
- National Infrastructure and Service Transformation Authority (UK). NISTA Major Projects Annual Report 2025-26. GOV.UK. 2026.
- OECD and European Commission Joint Research Centre. Handbook on Constructing Composite Indicators: Methodology and User Guide. OECD Publishing, Paris. 2008.
Further reading
- Robert S. Kaplan and David P. Norton. The Balanced Scorecard - Measures That Drive Performance. Harvard Business Review 70(1), January-February 1992, 71-79. 1992.
- Michael D. Mastrandrea, Christopher B. Field, Thomas F. Stocker and others. Guidance Note for Lead Authors of the IPCC Fifth Assessment Report on Consistent Treatment of Uncertainties. Intergovernmental Panel on Climate Change. 2010.
- U.S. Government Accountability Office. IT Dashboard: Selected Agencies' Investment Ratings Fail to Fully Consider Risks (GAO-27-108416). GAO. 2026.
Sources last verified 2026-10-08.