Building the AI Business Case
A business case is an argument, not a request for money: it shows why this use of resources beats the alternatives, including doing nothing, and what must be true for the value to appear. A credible AI case discounts potential value for adoption and realization, shows a range instead of one number, names the assumptions that could break it, and counts each risk once.
After this chapter you can
- Build a case from a measured problem, a baseline and business as usual, using the five questions of the Five Case Model.
- Separate potential, expected and realized value, and discount for adoption and realization.
- Show full cost by year and a range of scenarios instead of one number.
- Use sensitivity and switching values to find the assumptions that could break the case.
- Calculate risk-adjusted value without counting risk twice, and decide when to fund evidence before scale.
The British government has a standing instruction for anyone who writes a business case for an IT system. Before the details are examined, and in the absence of better evidence, add up to 200 percent to the estimated capital cost and up to 54 percent to the expected duration1. The figures are not a penalty. They come from a review of large public projects that measured how far earlier cases had been wrong at the outline stage, and the guidance is blunt about the reason: there is “a demonstrated, systematic, tendency for project appraisers to be overly optimistic”1.
The current edition of the Treasury’s appraisal manual, The Green Book, still defines optimism bias as “the proven tendency for appraisals to be over-optimistic about key assumptions,” and tells appraisers to correct it by raising costs and durations and lowering benefits2. An AI business case written today is exactly the kind of document that instruction was written for. It usually rests on a demo that worked, a supplier’s productivity figure from someone else’s workflow, and no history inside the organization. The question for an executive is not whether such a case is optimistic. It is whether the case has been built so that its optimism can be found.
The core idea
A business case answers one question: why should we put resources behind this initiative rather than use them elsewhere, including the option of changing nothing? A demo answers a different question, whether the technology can do the task. A successful demo is evidence about feasibility. It says nothing yet about whether the organization should invest in doing the task at scale, which is the only question an investment board is there to decide.
A credible case is a chain of claims, each of which can be checked. It starts with a problem that matters and can be measured, not with a technology. It records the baseline, the current state measured before anything changes, which Baselines, Metrics and Measurement treated in depth. It names the intervention precisely: which step of whose work the AI changes. It states the expected outcome with a target and a guardrail, so that a gain in speed cannot be bought with a loss in quality. Only then does it state the decision requested.
Every other part of the case hangs from this spine: the value, the cost, the range, the evidence and the risk. A case that opens with the model, the vendor or the demo has the spine upside down.
Five questions every case must answer
One widely used public-sector structure for a business case is the Treasury’s Five Case Model. It is worth borrowing in any organization, because it separates five questions that AI proposals tend to blur2.
Two rules from the same manual deserve a place in every AI investment review. First, the “business as usual” option, the outcome expected if current arrangements simply continue, “must always be taken forward” to the shortlist as the benchmark against which every other option is compared2. Doing nothing is rarely free. Demand grows, backlogs lengthen and hiring follows, and a case that never prices that path cannot show what the investment changes. Second, the effort should be proportionate: a case needs enough detail to support a sound decision and no more2. A three-month pilot needs a clear decision rule, not a five-year model.
Potential, expected and realized value
A common error in an AI business case is to present one kind of value as another. Three numbers need to be kept apart. Potential value is the theoretical maximum: every eligible person uses the tool from the first day, and every hour it frees becomes useful output. Expected value is what remains after honest discounts for adoption and realization. Realized value is what is measured after go-live, and it is the only one of the three that is a fact.
Suppose an initiative could, at full use, release capacity worth 10 million a year, or 30 million over three years. People do not all adopt at once. If use ramps from 40 percent of eligible staff in the first year to 70 percent in the second and 90 percent in the third, 20 million remains. Freed time does not all turn into output or avoided cost either; Productivity vs Realized Capacity showed how much of it is absorbed by the rest of the workflow. If half of it is realized, the expected value is 10 million. That is a third of the potential, and it is the number the case should present.
A second discount hides inside productivity claims. As AI and Workforce Productivity showed, work that is 30 percent faster frees about 23 percent of the time, not 30. A case that multiplies hours by the headline speed-up has overstated its benefit before any other assumption is examined.
Full cost, and when the money arrives
The cost side has its own version of potential value: the price list. A supplier’s quote shows licenses and model usage. The total also includes integration with existing systems, data preparation, training and change, evaluation and human review, monitoring, support and governance. Understanding AI Total Cost of Ownership, in Module 08, sets out the full method. The business case needs two things from it: the total, and the split between one-time and recurring costs, because that split shapes the years.
Continue the example. Suppose integration, data work and training cost 2 million once, and running the system costs 1 million a year. Over three years the full cost is 5 million against the expected benefit of 10 million, a net value of 5 million. The timing matters as much as the total.
A case should say when benefit starts and when it reaches a steady state, and it should show the negative first year rather than averaging it away. Turning these flows into return on investment, payback and present value is the work of the next chapter.
A range, and the assumptions that move it
A single net value of 5 million implies a precision no AI initiative has at approval. A case should show a range. In a conservative scenario, adoption ramps from 30 to 70 percent, only 40 percent of freed capacity is realized and integration overruns by half a million: benefit is 6 million, cost 5.5 million, and the net value nearly vanishes. In an upside scenario, faster adoption and 60 percent realization lift benefit to 13.5 million, but heavier use also raises running costs, to 6.5 million in total.
Scenarios show the spread. Sensitivity shows its cause. Move one assumption at a time by the same amount, here 20 percent in the wrong direction, and see how much net value is lost.
The Green Book adds a sharper test, the switching value: how far an assumption would have to move before the option stops being worth doing2. In the example, the case still breaks even if realization falls from half to a quarter, if adoption reaches only half of plan, or if running costs come in at up to about two and two-thirds times the estimate. Apply the Treasury’s upper-bound uplift of 200 percent to the 2 million of one-time cost and it becomes 6 million; the base case still clears, by 1 million. Switching values turn a debate about whose forecast is right into a more useful question: how likely is it that the assumption moves that far?
Sensitivity also tells you where to spend scrutiny, and that is often not where organizations spend it. Douglas Hubbard’s analysis of about 20 major IT investment cases found that the variables that mattered most to the decision were routinely the ones nobody had measured, while the most carefully measured variables mattered least3. In an AI case the pattern is easy to picture: the license price is negotiated to the last decimal, and the adoption rate is a guess.
Where the numbers come from
Every number in a case comes from one of two places. Daniel Kahneman called them the inside view, built from the details of the plan, and the outside view, built from what happened to similar efforts. His own textbook team learned the difference the hard way: their two-year estimate was an inside view, and the eight years the book actually took matched what comparable teams had experienced4. The outside view is less flattering and far more accurate.
A supplier’s productivity figure is an inside view borrowed from someone else’s workflow, and an average cannot simply be transplanted. In a large study of customer-support agents, an AI assistant lifted productivity by 15 percent on average, by far more for the least experienced agents and by little for the most experienced5. The same tool would produce a different number in a workforce with a different mix of experience. A case should therefore ask of every borrowed figure which reference class it came from, and whether the organization resembles it. The supplier’s number describes how that workforce performed, not how yours will.
When the most sensitive assumption rests on borrowed evidence, the right decision is usually to buy better evidence before buying scale. A pilot with its own business logic, a comparison group and a rule agreed in advance, the method Baselines, Metrics and Measurement set out, converts the weakest assumption into the strongest. How to score the confidence of each benefit line belongs to the AI Value Scorecard, later in this module.
Risk-adjusted value, counted once
Some risks do not shrink the benefit; they threaten to remove it. Integration may fail, the quality bar may not be met, a regulator or works council may object, or a supplier may withdraw the product. A simple way to bring delivery risk into the case is to multiply the value if the initiative is delivered by the probability that it is, and then subtract the full cost.
In the base case, value if delivered is the expected benefit of 10 million. Suppose the board judges a 70 percent chance of delivering it. Probability-weighted value is 7 million; minus the full cost of 5 million, the risk-adjusted value is 2 million. A case that looked like 5 million of net value is worth less than half that once delivery risk is priced. The calculation assumes the whole cost is spent even when delivery fails. Staging the investment, so that a failing initiative can be stopped before most of the money is gone, reduces that exposure; stage gates are part of the portfolio discipline in Module 09.
The formula is easy to misuse. Adoption and realization discounts describe how much value arrives if the system works; the probability of delivery describes whether it works. They are different discounts, and each belongs in the case once. A common error is to start from a value already labeled expected, already haircut for delivery risk, and multiply it by a probability again. That double-counts risk and can kill a sound case as surely as optimism can approve a weak one. Whoever builds the case should say which discount was applied where.
Story: the permits office asks for 1.5 million
The case that follows is an illustrative composite of a familiar public-sector proposal; the numbers are invented to make the arithmetic visible.
You sit on the investment board of a city government agency. The building-permits office brings a proposal. Applications take 60,000 reviewer hours a year, residents wait weeks for decisions, and complaints have reached the city council. Applications are forecast to rise by a tenth next year, about 6,000 more reviewer hours, or roughly four more reviewers if nothing changes. The proposal is an AI assistant that checks each application for completeness and drafts the reviewer’s notes. The supplier says reviewers work 30 percent faster. The team’s case is tidy: 30 percent of 60,000 hours is 18,000 hours saved, worth 900,000 a year at 50 per hour, against a three-year license of 1.5 million. The demo was impressive, and the office director is present. Before reading on, decide: approve, reject, or something else?
Many boards would split between approving, because the savings look large, and rejecting, because the number looks inflated. This board did neither. It rebuilt the benefit first. Thirty percent faster frees about 23 percent of the time, roughly 13,800 hours rather than 18,000, and with 60 percent of reviewers using the assistant in the first year, about 8,300. It then asked what the hours were worth. The city was not cutting posts, so the hours were not cash. The financial benefit was the hiring it could avoid as applications grew, and the larger public benefit was a shorter wait, measured in days to decision. Against business as usual, four extra reviewers and longer queues, the case looked quite different from the one in the proposal.
The board then turned to cost and evidence. The license was not the full cost: integration with the permit system, training, the reviewers’ time checking drafts and the governance a public body needs when software assists decisions about residents were all missing. Until a real estimate existed, the board applied the optimism-bias uplift to the one-time cost. And the 30 percent came from another city’s different workflow, a borrowed inside view. So the board funded a three-month pilot of 200,000 in two of the office’s six review teams, with the other four as the comparison and a written rule: scale if time per application falls at least 15 percent and errors in approved permits do not rise; redesign if speed improves but errors rise; stop if time falls by less than 5 percent. The case did not fail. It was converted from a promise into a test.
What this means for leaders
The leader’s job in a business case review is not to check the arithmetic, though the arithmetic must be right. It is to check the structure: that the case starts from a measured problem, compares against business as usual, separates potential from expected value, shows full cost by year, presents a range with the assumptions that drive it, and says where each important number came from. A case built that way can be wrong, but it will be wrong in ways the organization can see and correct.
Two habits make the biggest difference. Ask for switching values rather than arguing about forecasts, because they move the conversation to the assumptions that could actually break the case. And when the most sensitive assumption rests on someone else’s evidence, fund the evidence first. A small, well-designed pilot with a decision rule is usually a better use of money than a confident approval or a cautious rejection.
Check yourself
- A successful demo is strong evidence that an AI investment will pay off.
- The business-as-usual option belongs in the comparison even when nobody wants to keep the current process.
- If AI makes reviewers 30 percent faster, the case can count 30 percent of their hours as freed.
- The assumptions to scrutinize hardest are the ones that are easiest to measure.
- Multiplying an already risk-adjusted expected value by a probability of success makes the case more prudent.
- A supplier’s average productivity gain may not hold for a workforce with a different mix of experience.
Reflection: the last case you approved
What comes next
A well-built case gives the board a defensible forecast and a decision. It is still a forecast. Once the money is spent, the question becomes whether the expected 10 million is actually being captured, and what return the full cost is producing. The next chapter, AI ROI and Value Realization, sets the forecast against the result.
References
- HM Treasury. Supplementary Green Book Guidance: Optimism Bias. GOV.UK. 2013.
- HM Treasury. The Green Book (2026): Appraisal and Evaluation in Central Government. GOV.UK. 2026.
- Douglas W. Hubbard. How to Measure Anything: Finding the Value of Intangibles in Business, 3rd edition. Wiley. 2014.
- Daniel Kahneman. Thinking, Fast and Slow. Farrar, Straus and Giroux. 2011.
- Erik Brynjolfsson, Danielle Li and Lindsey Raymond. Generative AI at Work. The Quarterly Journal of Economics 140(2). 2025.
Further reading
- HM Treasury. The Green Book (2026): Appraisal and Evaluation in Central Government. GOV.UK. 2026.
- Daniel Kahneman. Thinking, Fast and Slow. Farrar, Straus and Giroux. 2011.
- Douglas W. Hubbard. How to Measure Anything: Finding the Value of Intangibles in Business, 3rd edition. Wiley. 2014.
Sources last verified 2026-10-08.