AI Academy · Book
Executives & Directors · Module 08 · Chapter 005

Experimentation vs Production Economics

An experiment buys information; a production system buys reliable delivery. The two are priced differently, so a cheap pilot says little about what production will cost, and an experiment that ends in a stop can be a bargain. Leaders who budget experiments by the decision they inform, and rebuild the case from scratch at the production gate, avoid both overbuilding and under-costing.

≈ 17 min read

After this chapter you can

  • Distinguish the economics of an experiment (cost of a good decision) from those of production (complete cost of reliable delivery).
  • Set a ceiling on an experiment's budget from the expected loss it lets the organization avoid.
  • Identify which pilot numbers carry over to a production case and which must be measured again.
  • Explain why quality and availability bars raise cost non-linearly, and why the bar must be agreed before the estimate.
  • Rebuild a production case from full TCO, real adoption, volume scenarios, an agreed bar and an owner.

Suppose an AI pilot handles 90 percent of its cases correctly, and the workflow it is meant to join needs 99 percent. How much more work is that? A common first answer is about ten percent more: nine points on a scale of a hundred. Count the errors instead. In 10,000 cases, a system that is right 90 percent of the time makes 1,000 errors. One that is right 99 percent of the time makes 100. The production bar is not nine points higher. It is a tenfold cut in errors.

In 10,000 cases, 90 percent correct means 1,000 errors and 99 percent means 100, a tenfold reduction.90% correct1,00095% correct50098% correct20099% correct100ILLUSTRATIVE NUMBERS
Figure 8.5.1 Errors in 10,000 cases. Going from 90 to 99 percent is a tenfold cut in errors, not a nine-point improvement.

Engineers who run large systems know this curve well. Google’s site reliability engineers put it plainly: cost does not rise in a straight line as reliability improves, and one more increment of reliability may cost a hundred times more than the one before1. The same is true of quality. That is one reason a pilot’s cost is such a poor guide to what production will cost. Data, Integration and Operational Costs ended with another: a pilot can prove an idea with clean sample data, one connection and a team willing to check every output by hand. Production can do none of those things.

The core idea

An experiment and a production system are different investments with different objectives, and they should be judged by different numbers. An experiment buys information. Its job is to reach a good decision, quickly and at a sensible cost: does the idea work, will people use it, is the value large enough? A production system buys reliable delivery. Its job is to produce business outcomes at the required quality, for as long as it runs, at an acceptable complete cost.

An experiment buys information and is judged by time to decision; production buys reliable delivery and is judged by complete cost per outcome.ExperimentProductionWhat it buysInformationReliable deliverySucceeds whenA good decision, made fastOutcomes at the required qualityMeasured byTime to decision, uncertainty removedComplete cost per outcomeCan tolerateManual steps, small scale, downtimeVery little of thatBudget set byThe decision it informsTotal cost of ownership
Figure 8.5.2 Two investments with two objectives. Judging either by the other’s numbers produces the classic mistakes.

Two mistakes follow from confusing them, and they point in opposite directions. The first is to run an experiment as if it were already production, spending on robustness before anyone knows the idea is worth it. The second is to approve production on an experiment’s numbers, as if the pilot’s invoice were a sample of the real bill. This module measures the real bill as total cost of ownership: build plus run plus operate plus change, over the system’s life, as Understanding AI Total Cost of Ownership defined it. An experiment’s cost is only the first few lines of that sum.

An experiment is priced by the decision it informs

How much should an experiment cost? The useful answer comes from decision analysis. In 1966 Ronald Howard showed that the probabilities and the economics of a decision can be considered together, so that reducing an uncertainty can be given a numerical value2. Information is worth what it would change about the decision, and no more.

An illustrative worked example makes the idea concrete; its figures are invented. Suppose a finance team is asked to approve 250,000 to test an AI tool that matches supplier invoices. If the tool works, the production system gains 3 million over its life, net of all costs. If it does not, the organization loses 2 million. Finance puts the chance of failure at 30 percent.

Illustrative - a test that reveals whether the tool works lets you stop in the 30 percent case and avoid 2 million, so it is worth at most 600,000.Test beforecommittingIt works: 70%Go and gain 3 millionIt fails: 30%Stop and avoid losing 2 millionMOST THE TEST CAN BE WORTH30% x 2 million= 600,000
Figure 8.5.3 Illustrative. An experiment is worth at most the expected loss it lets you avoid; here, 600,000.

Without a test, the expected value of going ahead is 0.7 times 3 million, minus 0.3 times 2 million: 2.1 million minus 600,000, or 1.5 million. It is worth doing, and three times in ten it loses 2 million. Now imagine a test that reliably showed in advance which case you were in. You would go only when the tool works, so the expected value rises to 2.1 million. The difference, 600,000, is the most any test of this decision can be worth. It equals the chance of failure times the loss it would let you avoid. A real experiment is sometimes wrong, so it is worth less than that ceiling, but the ceiling sets the scale. Here the 250,000 is under the ceiling, so it can be justified, but only if the test is quick and aimed at the riskiest assumption, such as the match rate on live invoices.

Three practical rules follow. First, budget an experiment against the decision it informs, not against zero. Spending 250,000 to protect a 2 million loss can be a bargain; the same 250,000 to protect a 200,000 loss is not. Second, design the experiment to answer the question that would change the decision, usually the assumption with the most money resting on it, rather than to produce a demonstration. Third, speed has value. A test that takes six months delays the decision and the gain with it, so time to decision is an economic measure in its own right.

The example also shows why an experiment that ends in a stop can be a success: in the 30 percent case, stopping is precisely what it was paid to make possible. The portfolio discipline that turns this into practice, with stop rules written before the money moves, belongs to Prioritizing the AI Portfolio in Module 09.

What a pilot proves, and what it does not

As What Does AI Value Actually Mean? showed, a pilot proves feasibility; scale and return are separate questions. For the production case the practical question is narrower: which of the pilot’s numbers can be carried across, and which must be measured again?

Pilot quality, time saved and exception rate carry over only partly, user enthusiasm rarely, and pilot cost per request not at all.Pilot numberCarries over?WhyQuality on test casesPartlyTest data is usually cleaner than live dataTime saved per taskPartlyOnly for the tasks and users testedUser enthusiasmRarelyVolunteers, novelty and close supportCost per requestNo integration, support or hand-checking pricedException ratePartlyCurated inputs hide the hard cases
Figure 8.5.4 Many pilot numbers are evidence about the idea. Few are evidence about the production system.

Quality measured on the pilot’s cases is real evidence, but those cases were often chosen, cleaned or both, and the Thai screening clinics in the previous chapter showed what live data does to a threshold. Time saved is real for the tasks and people tested; it says less about the tasks and people who were not. Enthusiasm is the weakest number of all, because pilot users are usually volunteers, enjoy close support from the team and are using something new.

Cost per request is the number most often carried across, and it should not be. A pilot rarely pays for integration, security, on-call support or monitoring, and many pilots contain a hidden worker: a team member who checks outputs by hand before anyone sees them. That person is not in the pilot’s budget, but the work they do is real, and in production it must either be priced as an exception queue or engineered away.

The pilot is a small share of the bill

The gap between a working prototype and a working system is well documented. In a widely cited 2015 paper, Google engineers observed that the model code is only a small fraction of a real machine learning system; the infrastructure around it is “vast and complex”3. Many experiments never cross the line: IDC research conducted with Lenovo, a hardware vendor, found that for every 33 AI proofs of concept a company launched, only four reached production4, and Gartner predicted in 2024 that at least 30 percent of generative AI projects would be abandoned after proof of concept, with escalating cost among its reasons5. A 2022 survey of published deployment case studies adds that problems arise at every stage, from data management to monitoring6.

Illustrative - a 200,000 pilot precedes production costs of 1.8 million in year one and 1.1 million in each later year, so the pilot is 5 percent of three-year cost.BuildRunOperateChangePilot150 k200 kYear 11,000 k300 k400 k1,800 kYear 2400 k500 k200 k1,100 kYear 3400 k500 k200 k1,100 kILLUSTRATIVE NUMBERS
Figure 8.5.5 Illustrative. A 200,000 pilot followed by 4 million of three-year production TCO: the pilot is one twentieth of the bill.

The illustrative figure shows the shape. A three-month pilot costs 200,000: mostly build, with a little cloud spend and some of the team’s time. Production costs 1.8 million in its first year, as integration, security, monitoring and support are built, and 1.1 million in each year after that, as the system runs, is operated and absorbs changes. Over three years that is 4 million. The pilot was 5 percent of it. None of the families were hidden; Understanding AI Total Cost of Ownership and Data, Integration and Operational Costs priced each of them. They simply arrive at the production gate, and nothing in the pilot’s invoice shows them coming.

The production bar is a cost decision

Quality is one bar production must clear. Availability is another, and it follows the same arithmetic. A system at 99 percent availability may be down for about 88 hours a year, which is acceptable for an internal report and unacceptable for a workflow that stops when the system does. Each extra nine cuts the permitted downtime tenfold, to about 9 hours at 99.9 percent, and requires more redundancy, more monitoring and someone on call. That is the non-linear cost curve the site reliability engineers describe1. A pilot that is down for an afternoon, by contrast, inconveniences a few volunteers.

The executive point is that the bar should be set by the workflow, not by engineering pride. What does one error cost the business, and what does one hour down cost? The expected-downtime test in Data, Integration and Operational Costs gives the method for the second question. Without an agreed bar, a production estimate is meaningless, because the same system can cost several times more at one bar than at another.

Two ways to get it wrong

The first error is to overbuild the experiment. A team builds a production-grade platform, with high availability, full security review and polished interfaces, for an idea nobody has yet shown to be valuable. Money and months are committed before the main uncertainty is reduced, and the time to decision stretches. Worse, the size of the investment makes stopping harder, which is exactly the escalation of commitment that Prioritizing the AI Portfolio warns against.

Overbuilding an experiment and underbuilding production are opposite errors; the right level of engineering fits the stage.Overbuilt experimentProduction-grade before value is provenUnderbuilt productionPrototype shortcuts carried into serviceFit to the stageCheap to learn, complete to run
Figure 8.5.6 Both errors are expensive. The right level of engineering depends on which investment you are making.

The second error is to underbuild production. A successful prototype is promoted with its shortcuts intact: the manual steps, the hand-checking, the single server, the absence of monitoring or anyone on call. Sculley and his colleagues called these shortcuts technical debt, and the interest is paid later as incidents, rework and hurried re-engineering3. Shortcuts are rational in an experiment because many experiments end; they are expensive in production because production systems do not.

A few things are not optional at either stage. Security for real customer data, and the law on personal data, apply from the first day an experiment touches them.

Rebuild the case at the gate

At the production gate, the pilot’s case should not be scaled up. It should be rebuilt, because almost every number in it was measured in conditions production will not have.

The production case needs full TCO, adoption among target users, volume scenarios, and an agreed quality bar and owner.Full TCOBuild, run,operate, changeReal adoptionTarget users,not volunteersScenariosBase, threetimes, one fifthBar and ownerQuality,availability,accountabilityA new case, not the pilot multiplied
Figure 8.5.7 Four things the production case needs that the pilot did not measure.

Four things belong in the rebuilt case. Full cost of ownership, including the hidden worker priced as an exception queue or removed by engineering. Adoption among the people who must use the system, rather than the volunteers who chose to. Scenarios for volume: what happens if usage is three times the forecast, and what happens if only a fifth of the intended users adopt it? And an agreed bar and an owner: the quality and availability the workflow needs, and a named person accountable for the outcome after launch. How those scenarios turn into cost per customer and break-even is the subject of the next chapter.

The case does not stop being an estimate on launch day. The FinOps discipline treats comparing forecast cost with actual cost, per unit of business value, as a continuing practice78. Usage, prices and behavior all drift, so the production case should be re-run with real figures each quarter. How to take a system across the line in stages, region by region, is covered in From AI Pilot to Production to Scale.

Story: Sweden’s screening trial and England’s decision

Breast screening programs in Sweden and England both rely on double reading: two specialists read every mammogram. By 2021 several retrospective studies, which ran AI over archives of past mammograms, had shown promising results, but no randomized trial had tested AI inside a working program9. That is the gap between a pilot and production. An archive is chosen data with known answers. A screening program is thousands of women a week, real radiologists, and outcomes that take years to show.

Sweden ran the test inside its national program. From April 2021 women at four screening sites in the southwest of the country were randomly assigned either to standard double reading or to AI-supported screening, in which an AI system scored each examination, sent most of them to one radiologist and the highest-risk ones to two, and marked suspicious areas. The first safety analysis, of 80,033 women, was published in August 2023. It found a similar cancer detection rate, no rise in false positives, and 44 percent less reading work. The main question, how many cancers would surface between screening rounds, needed two more years of follow-up9.

Suppose you run England’s program at the end of 2024. It carries out about 2.1 million screens a year, each read by two specialists10. Do you (A) adopt AI as one of the two readers nationally, on the strength of Sweden’s numbers; (B) wait for the Swedish follow-up and for others to move first; or © pay for a large trial in English screening conditions before deciding?

In February 2025 the UK government chose C. It put 11 million pounds, through the National Institute for Health and Care Research, into the EDITH trial: nearly 700,000 women at 30 NHS sites, testing whether AI lets one specialist safely do the screening work that takes two today10.

In the MASAI trial of 105,934 women, AI-supported screening found 29 percent more cancers with 44 percent fewer screen readings.105,934Women randomizedFour sites in Sweden+29%Cancers detected6.4 vs 5.0 per 1,000-44%Screen readings61,248 vs 109,692Source: MASAI trial, Lancet Digital Health · 2025
Figure 8.5.8 A randomized trial inside a working program, not a demo: more cancers found with almost half the reading work.

Sweden’s full results came out over the following year. Among the 105,934 women randomized, the AI group found 6.4 cancers per 1,000 women screened against 5.0, without significantly more recalls or false positives, and needed 61,248 screen readings against 109,692, a 44 percent cut11. The two figures measure different things: the 29 percent is cancers found at the screen itself, and the 44 percent is the number of human readings, not radiologists’ total time. In January 2026 the main result followed: cancers surfacing between screens were no more common with AI, 1.55 against 1.76 per 1,000, and sensitivity rose from 73.8 to 80.5 percent with the same specificity12. Option B would have bought that certainty, but only about Sweden.

Why not A? Because few pilot numbers carry over unchanged, even from a well-run randomized trial. The Swedish trial screened women aged 40 to 80, mostly every 18 to 24 months, with one AI product at four sites9. England invites women aged 50 to 71 every three years13, so cancers have longer to grow between rounds, and the result that matters most, cancers missed between screens, could come out differently. That reasoning is this chapter’s, not the government’s; its stated case was that a successful trial could free up hundreds of radiologists10.

The economics of option C follow the chapter’s rule: price the experiment by the decision it informs. A 2026 model of the UK program, built on prospective trial evidence, estimated that replacing one human reader with AI would cut lifetime costs by about 31 pounds for every woman invited while slightly improving outcomes14. It is a model, not a measurement, which is exactly why the measurement is worth buying. And if the trial finds that AI does not hold up in English conditions, the stop will be a success: it will have prevented a national change that would not have paid.

What this means for leaders

Treat the two investments differently and ask different questions of each. For an experiment, ask what decision it informs, how much money rests on that decision, and how soon the answer will come; a large experiment protecting a larger decision can be good value, and a stop can be a success. For production, ask for a rebuilt case: full cost of ownership, adoption among the real users, volume scenarios, an agreed quality and availability bar, and an owner. Never let the pilot’s invoice stand in for the production budget, and never let a pilot be built as if it were already production.

Check yourself

  1. A cheap pilot is good evidence that production will be cheap.
  2. The most an experiment can be worth is the expected loss it lets you avoid.
  3. Moving from 90 to 99 percent quality is roughly a ten percent improvement.
  4. Experiments should be built to production standards so they can scale quickly.
  5. An experiment that ends in a decision to stop can be an economic success.
  6. The production case should be rebuilt from full TCO, real adoption and volume scenarios, not scaled up from the pilot.

Reflection: one pilot, two budgets

What comes next

A rebuilt production case still has to answer the question finance will ask first: what does each customer, each transaction or each outcome cost, and does that number improve or worsen as use grows? England’s screening decision turns on exactly that: what each screen costs to read, and how value and cost behave across a whole national program. Measuring it properly is the subject of AI Unit Economics and Economics at Scale.

Laws referenced

General Data Protection Regulation · EU

Regulation (EU) 2016/679

Personal data is any information relating to an identified or identifiable person, directly or indirectly, including by an identifier such as an online ID (Art. 4(1)). Lawful basis and purpose limitation (Arts. 5-6); processing special-category data, including biometric data used to identify a person, health data and data revealing ethnicity, is prohibited unless a specific exception applies (Art. 9); data protection by design and by default (Art. 25); processors such as AI vendors may act only under a written contract with required terms and sufficient guarantees (Art. 28); transparency to data subjects (Arts. 13-14); right not to be subject to a decision based solely on automated processing with legal or similarly significant effects (Art. 22); breach notification to the supervisory authority within 72 hours (Art. 33) and to individuals without undue delay when the risk is high (Art. 34); data protection impact assessment for high-risk processing (Art. 35). Fines up to EUR 20 million or 4% of global turnover.

  • 2018-05-25 — Applies

Last verified 2026-10-08 · official text

EU AI Act · EU

Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744

Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.

  • 2024-08-01 — Entered into force
  • 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
  • 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
  • 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
  • 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
  • 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
  • 2028-08-02 — High-risk obligations for AI in products regulated under Annex I

Last verified 2026-10-06 · official text

References

  1. Marc Alvidrez (edited by Kavita Guliani). Embracing Risk (Chapter 3). Site Reliability Engineering: How Google Runs Production Systems, O'Reilly Media. 2016.
  2. Ronald A. Howard. Information Value Theory. IEEE Transactions on Systems Science and Cybernetics, 2(1), 22-26. 1966.
  3. D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Crespo and Dan Dennison. Hidden Technical Debt in Machine Learning Systems. Advances in Neural Information Processing Systems 28 (NIPS 2015). 2015.
  4. Evan Schuman. 88% of AI pilots fail to reach production - but that's not all on IT. CIO.com, reporting IDC research conducted with Lenovo. 2025.
  5. Gartner. Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept By End of 2025. Gartner Newsroom. 2024.
  6. Andrei Paleyes, Raoul-Gabriel Urma and Neil D. Lawrence. Challenges in Deploying Machine Learning: A Survey of Case Studies. ACM Computing Surveys, 55(6). 2022.
  7. FinOps Foundation. FinOps Framework. The Linux Foundation. 2026.
  8. J.R. Storment and Mike Fuller. Cloud FinOps: Collaborative, Real-Time Cloud Value Decision Making, 2nd edition. O'Reilly Media. 2023.
  9. Kristina Lång, Viktoria Josefsson, Anna-Maria Larsson, Stefan Larsson, Charlotte Högberg, Hanna Sartor, Solveig Hofvind, Ingvar Andersson and Aldana Rosso. Artificial intelligence-supported screen reading versus standard double reading in the Mammography Screening with Artificial Intelligence trial (MASAI): a clinical safety analysis of a randomised, controlled, non-inferiority, single-blinded, screening accuracy study. The Lancet Oncology, 24(8), 936-944. 2023.
  10. Department of Health and Social Care, Department for Science, Innovation and Technology and NHS England. World-leading AI trial to tackle breast cancer launched. GOV.UK. 2025.
  11. Veronica Hernström, Viktoria Josefsson, Hanna Sartor, David Schmidt, Anna-Maria Larsson, Solveig Hofvind, Ingvar Andersson, Aldana Rosso, Oskar Hagberg and Kristina Lång. Screening performance and characteristics of breast cancer detected in the Mammography Screening with Artificial Intelligence trial (MASAI). The Lancet Digital Health, 7(3), e175-e183. 2025.
  12. Jessie Gommers, Veronica Hernström, Viktoria Josefsson, Hanna Sartor, David Schmidt et al. Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study. The Lancet, 407(10527), 505-514. 2026.
  13. NHS. Breast screening (mammogram): when you'll be invited and who should go. NHS website. 2026.
  14. Harry Hill and C. Roadevin. Economic evaluation of artificial intelligence for cancer detection in the UK breast screening programme. British Journal of Cancer, 135(3), 453-460. 2026.

Further reading

Sources last verified 2026-10-08.