From AI Pilot to Production to Scale
A successful pilot proves that something can work under borrowed conditions. Validation proves it changes a business result, production proves the organization can run it every day without those conditions, and scale proves the result repeats, at a cost that still makes sense, in places the pilot never saw. These are four separate proofs, and each deserves its own evidence before the next commitment.
After this chapter you can
- Separate the four proofs (feasible, valuable, operable, repeatable) and the gate that tests each.
- Identify the conditions a pilot borrows and assign each an owner and a cost in production.
- Design a pilot for its production path, with a baseline, a production hypothesis and a stop rule.
- Name the six tests of operability, including monitoring for drift and a business owner.
- Plan a progressive rollout (silent run, canary, waves) and treat scale as a new proof, not a photocopy.
- Read pipeline measures honestly to spot both pilot purgatory and premature scale.
Here is a document most executives have seen in some form. It is the close-out summary of a six-week pilot at a facilities-management company that maintains offices, warehouses and hospitals for its clients, and it is an illustration, not a real company’s file. An AI assistant read incoming repair requests and suggested a triage route: urgent dispatch, a scheduled visit or a specialist contractor.
The status box is green and the recommendation is a rollout to all 4,000 coordinators. Now read the footnotes, which nobody presents. The request data was a hand-cleaned extract, refreshed weekly. A senior engineer checked every suggested route before it was used. Two data scientists were on call during business hours. The assistant received its requests through a nightly file, not through the work-order system. And the costs were calculated at pilot volume, on introductory pricing.
None of those footnotes means the pilot failed. Each one means the pilot ran on conditions that production will not provide. The memo answers one question, whether the assistant can triage repair requests well, and silently assumes the answers to the three others that follow it.
Four proofs, not one launch
It helps to treat the road from pilot to scale as a series of four separate proofs. A pilot proves the approach is feasible: under controlled conditions, it can work. Validation proves it is valuable: it changes a business result, not only the quality of the model’s answers. This is the proof pilots most often skip, because a model that produces good answers is easily mistaken for a process that produces a better business result. Production proves it is operable: the organization can run it every day, on live data, with real users, support and controls, without the team that built it standing by. Scale proves it is repeatable: the result holds in other teams, sites or markets, at a unit cost that still makes sense.
These gate outcomes judge one initiative’s evidence at a stage boundary; they are not the portfolio decisions of Prioritizing the AI Portfolio (fund, experiment, defer, stop), which place every initiative at the quarterly review.
The distinction matters because most organizations stall between the first gate and the third. The AI Maturity Model described that stall as pilot purgatory and showed how widespread it is. One more data point makes the scale of the problem concrete: in research with Lenovo reported in 2025, IDC found that for every 33 AI proofs of concept a company launched, only four reached production, about one in eight. IDC’s explanation was not weak models but low organizational readiness in data, processes and IT infrastructure1. The figure comes from a vendor-sponsored survey and should be read as a direction, not a benchmark. The direction is consistent with everything else we know. This chapter is about what it takes to cross the gap for a particular system.
What a pilot borrows
Pilots are easy to call successful because they borrow conditions they will later have to give back. The memo’s footnotes are a typical list: clean data, expert reviewers, engineers on call, a temporary integration, enthusiastic volunteers and a price that applies only at small volume.
Engineers have known this for a decade. In 2015 a team of Google engineers described “hidden technical debt” in machine learning: only a small fraction of a real-world machine learning system is the model code, and the surrounding infrastructure for data, configuration, monitoring and serving is vast, complex and costly to maintain. A RAND study based on interviews with 65 experienced data scientists and engineers found the same pattern from the failure side. Among its five root causes of failed AI projects were a lack of the data needed to train a useful model and inadequate infrastructure to manage data and deploy finished models3.
The practical test for an executive is simple. For every condition the pilot borrowed, ask who will supply it in production, and what it will cost. If nobody can answer, the pilot has proved less than the memo claims.
Design the pilot for the production path
The cheapest time to fix the borrowing is before the pilot starts. A pilot designed only to answer “can it work?” will answer that question and leave the rest to luck. A pilot designed with its production path in view collects the evidence the later gates will need.
Three disciplines do most of the work. First, write down the production hypothesis at the start: which system the assistant will live in, who will own it, what review it will need and what a single outcome should cost at full volume. Second, measure value against a baseline taken before the pilot, on the business result rather than the model’s output. Faster triage is only valuable if it changes something the company cares about, such as time to repair, repeat visits or penalties under its client contracts. And time saved is not money saved until the organization decides what to do with the freed capacity, as Productivity vs Realized Capacity showed. Third, time-box the pilot and decide the stop rule in advance, the discipline Prioritizing the AI Portfolio set out.
None of this requires production-grade controls during the experiment. It requires that the experiment be honest about which questions it is answering and which it is not.
Production: from prototype to service
Production is the point at which an AI system stops being a project and becomes a service that someone runs. It asks for a different set of capabilities, and few of them are about the model.
Monitoring deserves special emphasis, because AI systems change in ways that ordinary software does not. The data arriving in production drifts away from the data the model was built on; researchers call this concept drift, and it is the normal condition of a deployed model, not an exception4. Language models add further sources of change: a provider updates the model, a prompt is edited, users find new ways to ask. The NIST AI Risk Management Framework treats post-deployment monitoring, including mechanisms for user feedback, override, incident response and decommissioning, as a core part of managing an AI system rather than an optional extra5.
Ownership is the other test that pilots routinely fail. The team that built a pilot is rarely the team that should run it. Production needs a business owner accountable for outcome and adoption, with engineering accountable for reliability and security, as AI Roles, Ownership and Accountability describes. If the only person who can explain the system is the data scientist who built it, the system is not in production. It is a pilot with more users.
Roll out in waves, not in one jump
The memo’s proposal, from 200 coordinators to 4,000 in a quarter, is a twenty-fold jump in one step. Software engineering solved this problem long ago. Google’s site reliability engineers describe a canary as a partial and time-limited deployment of a change, evaluated to decide whether to proceed with the rollout6. The same logic applies to AI, with one addition: before the system acts at all, it can run silently.
A silent run, sometimes called shadow mode, connects the system to live data and records what it would have done, without anyone acting on it. It is the cheapest way to discover what the clean pilot extract was hiding. A canary puts the system in front of one team and compares its results with the baseline. Each wave after that has written entry criteria (quality, adoption, cost per outcome, incident rate) and a tested way back. The point is not caution for its own sake. It is that each wave produces the evidence that justifies the next one, and limits the damage if the evidence disappoints.
Scale is a new proof, not a photocopy
The most expensive mistake in this lifecycle is to treat scale as a photocopy of the pilot: the same system, more users. A photocopy reproduces everything, including the borrowed conditions, and multiplies their cost. A system that worked for one team of enthusiasts with an engineer down the corridor can fail for forty teams with none.
Scale changes three things. It changes economics: costs that were trivial at pilot volume, such as model usage, human review and support, grow with every user and transaction, while some benefits do not. Recompute cost per outcome at the target volume, the question The Economics of AI asked at ten times today’s use. It changes context: other regions bring different data, different processes and sometimes different law, so the result has to be shown again there. And it changes organization: training, support and governance must reach people who never volunteered.
The useful scaling question is therefore not “how fast can we roll it out?” but “what must be true at the next site for the result to repeat?” Reuse makes the answer cheaper. Shared connectors, evaluation sets and monitoring built once for the first production system can serve the next ten, which is why scale is also where an organization’s platform investment pays off.
Measure the pipeline honestly
Leaders who want to know whether their organization is crossing the gap need a few measures of the pipeline itself, not only of individual systems. Three are enough to start. The conversion rate is the share of qualified pilots that reach production. The time to production is the elapsed time from pilot start to live operation. And realized against expected value compares what each production system was forecast to deliver with what it delivered.
Each needs interpretation. A low conversion rate can mean good discipline, if pilots were designed to test uncertain ideas and most deserved to stop. It can also mean weak pilot design, if promising work keeps dying for production reasons nobody planned for. A long time to production usually points at a specific bottleneck: data access, security review, integration or a shortage of people who can run systems. The realized-against-expected comparison is the one most organizations skip, and it is the one that makes the next round of prioritization better.
The measures also guard against both traps on either side of the gap. Pilot purgatory shows up as a high pilot count, low conversion and long times to production. Premature scale shows up differently: systems that reached thousands of users quickly, with realized value well below forecast and costs above it. The first trap wastes learning. The second wastes money and trust, and is harder to reverse.
Story: Duke Health’s Sepsis Watch
One of the clearest published accounts of an AI system moving through these gates comes from a hospital, where the cost of getting it wrong is high and the evidence is unusually well documented. It documents three of the four proofs in detail, and is candid about the one it leaves open.
Before: in 2016 Duke University Hospital had a sepsis problem and a history. Its performance on the national sepsis treatment measure was poor, and earlier decision-support alerts for deteriorating patients had caused significant alarm fatigue without improving clinical outcomes. In April 2016 leadership launched a project to do better, with a multidisciplinary team of statisticians, data scientists, engineers and clinicians. The plan already named the production path: if the pilot at the flagship hospital succeeded, it would expand to two Duke Health community hospitals7.
The team did not treat the model as the product. The clinicians decided that alert-weary frontline staff were the wrong people to receive warnings and gave the job to rapid response team nurses, who helped design the application. The infrastructure was built for an emergency department of about 200 visits a day but designed to scale to 1,500 inpatient beds. Two monitoring systems watched the integration with the health record and the model’s daily inputs and outputs. A three-month silent period on live data preceded launch. A governance committee of nursing, physician and administrative leaders took charge of usage, training, reporting and, explicitly, sustainability after the pilot7.
After: Sepsis Watch launched on 5 November 2018 and stayed in continuous use by the nurses throughout the pilot. The authors were candid about cost. Significant investment went into aligning stakeholders, building trust, defining roles and training staff, and the product kept changing in its first months7. The report describes integration, not patient outcomes, so on the evidence it publishes the value proof is the one still open.
Scale came in two separate steps. The first was internal and was the expansion the original plan had named: in June 2019 Sepsis Watch was extended to the emergency departments of the two Duke Health community hospitals7. The second was external. In a study published in 2025, Duke researchers tested the model on 205,005 encounters at the four emergency departments of Summa Health, a separate health system in Ohio, and found it performed strongly at every site8. Two cautions apply: several authors have commercial interests in related software, and a retrospective validation is not yet a deployment. But that is the point. Even a system that worked at home had to prove itself again somewhere else.
What this means for leaders
The executive role in this lifecycle is not to approve rollouts. It is to ask, at each gate, what evidence would justify the next increment of money, risk and organizational attention, and to make sure that evidence is collected on purpose. Separate the four proofs in every review: a green pilot is evidence that the approach is feasible, not that it is valuable, operable or repeatable. Make the borrowed conditions visible and assign each one an owner and a cost. Insist on waves with entry criteria and a way back. And when a pilot fails a gate, treat redesign, deferral or a clean stop as decisions, not as failures.
Check yourself
- A pilot that met its targets is ready for production.
- A silent run lets a system meet live data before anyone acts on its outputs.
- Scaling a production system mostly means adding users.
- A low pilot-to-production conversion rate always signals failure.
- Only a small fraction of a real-world machine learning system is the model code.
- Once a model is in production, its performance stays stable unless someone changes it.
Reflection: read the footnotes
What comes next
A single system can now be walked from pilot to scale. But an organization runs many systems at once, with limited people and a calendar that does not wait. The next chapter, Building the 90-Day AI Plan, turns the portfolio and these gates into a concrete first quarter of action.
Laws referenced
Not legal advice. Laws change; verify before relying on this, and consult counsel for decisions.
EU AI Act · EU
Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744
Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.
- 2024-08-01 — Entered into force
- 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
- 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
- 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
- 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
- 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
- 2028-08-02 — High-risk obligations for AI in products regulated under Annex I
Last verified 2026-10-06 · official text
References
- Evan Schuman. 88% of AI pilots fail to reach production - but that's not all on IT. CIO.com, reporting IDC research conducted with Lenovo. 2025.
- D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Crespo and Dan Dennison. Hidden Technical Debt in Machine Learning Systems. Advances in Neural Information Processing Systems 28 (NIPS 2015). 2015.
- James Ryseff, Brandon F. De Bruhl and Sydne J. Newberry. The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed: Avoiding the Anti-Patterns of AI. RAND Corporation, RR-A2680-1. 2024.
- João Gama, Indrė Žliobaitė, Albert Bifet, Mykola Pechenizkiy and Abdelhamid Bouchachia. A Survey on Concept Drift Adaptation. ACM Computing Surveys 46(4). 2014.
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. NIST. 2023.
- Alec Warner and Štěpán Davidovič. Canarying Releases (Chapter 16). The Site Reliability Workbook, O'Reilly Media. 2018.
- Mark P. Sendak, William Ratliff, Dina Sarro and colleagues. Real-World Integration of a Sepsis Deep Learning Technology Into Routine Clinical Care: Implementation Study. JMIR Medical Informatics, 8(7), e15182. 2020.
- B. Valan, A. Prakash, W. Ratliff and colleagues (senior author M. Sendak). Evaluating sepsis watch generalizability through multisite external validation of a sepsis machine learning model. npj Digital Medicine, 8, 350. 2025.
Further reading
- Mark P. Sendak, William Ratliff, Dina Sarro and colleagues. Real-World Integration of a Sepsis Deep Learning Technology Into Routine Clinical Care: Implementation Study. JMIR Medical Informatics, 8(7), e15182. 2020.
- D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Crespo and Dan Dennison. Hidden Technical Debt in Machine Learning Systems. Advances in Neural Information Processing Systems 28 (NIPS 2015). 2015.
- Alec Warner and Štěpán Davidovič. Canarying Releases (Chapter 16). The Site Reliability Workbook, O'Reilly Media. 2018.
Sources last verified 2026-10-08.