AI Academy · Book
Executives & Directors · Module 08 · Chapter 008

Build vs Buy Economics and Investment Decisions

A vendor's annual price and an internal build estimate are not comparable numbers. A sound comparison counts both options on one boundary over the same service life, then adds what the quote leaves out: the value lost while waiting, the scarce people a build consumes, the risk that each option goes wrong and the cost of leaving. The output is not a winner but a decision with a stated point at which it would flip.

≈ 14 min read

After this chapter you can

  • Restate a build-versus-buy comparison on one four-family TCO boundary, at the same volume and service life.
  • Price the cost of delay and the opportunity cost of the team a build would use.
  • Risk-adjust both options, using IT overrun evidence for builds and repricing and exit costs for purchases.
  • Find the volume or price at which the answer flips and turn it into a review trigger.
  • Explain why buying first, with exit terms, can be a sound way to keep the option to build.

The memo below is illustrative, though its shape will be familiar to many investment committees. A regional water authority wants software that listens to acoustic sensors on its pipes and flags likely leaks. Two options reach the committee on a single page.

An illustrative memo compares a leak-detection subscription of 800,000 a year with a one-time build of 1.2 million and recommends building to save 1.2 million.OptionHeadline costMemo's three-year viewBuy: leak-detection service800,000 a year2.4 millionBuild: in-house model1.2 million, once1.2 millionRecommendationBuildSaves 1.2 million
Figure 8.8.1 Illustrative. A committee memo that compares a subscription with a build estimate and finds a saving that is not there.

The memo is tidy, confident and wrong in a way that is easy to miss. It sets a three-year subscription against the cost of reaching the first day of service. It says nothing about who runs the built system afterward, how long each option takes to start working, what the in-house team would otherwise be doing, or what happens if the vendor raises its price or the build runs late. Every one of those items has a number. This chapter puts them in.

The core idea

Build vs Buy vs Partner, in Module 04, settled the strategic question with two tests: would we still have an advantage if rivals had this capability, and how badly would a vendor’s change hurt us? This chapter takes the answer that survives those tests and prices it.

The idea is simple to state. Compare the decision, not the quote. Count both options on one boundary over the same service life, then add the costs that a price list never shows: time, scarce people, risk and exit. The output is not “build” or “buy”. It is a recommendation with the conditions under which it would change.

Five steps turn a build-versus-buy comparison into an investment decision - one cost boundary, the price of waiting, the team's best other use, risk adjustment and the point at which the answer flips.BoundaryFourfamilies foreach optionDelayValue lostwhile notlivePeopleBest otheruse ofthe teamRiskOverruns,repricing,exitFlip pointWhen theanswerchanges
Figure 8.8.2 Five steps from a headline price to an investment decision. Many memos stop after the first, or skip it.

The steps borrow nothing exotic. Management accountants have taught the make-or-buy decision this way for decades: compare only the costs that differ between the options, ask what the freed capacity could do instead, and weigh factors such as supplier reliability1. AI adds two twists. Its prices move faster than most components, and both options carry far more running cost than their headline suggests.

Buying is not zero engineering; building is not one-off

The two options hide their costs in opposite places. A purchase looks like a single run cost, but it still needs integration with the authority’s sensor platform and work-order system, data mapping, acceptance testing, a team to triage alerts and someone to manage the vendor. A build looks like a single project cost, but once it is live the authority owns everything the vendor would have carried: compute and storage, monitoring, security, on-call support, retraining and every model upgrade.

Buying still costs integration, testing, alert triage, vendor management and repricing; building keeps costing compute, monitoring and security, retraining, the team's best other use and overrun risk.What buying still costsIntegration and data mappingAcceptance testingAlert triageVendor managementRepricing at renewalWhat building keeps costingCompute and storageMonitoring, security, on-callRetraining and upgradesThe team's best other useDelivery overrun
Figure 8.8.3 Each option hides its costs where the other shows them. A fair comparison exposes both.

The cure is the module’s single definition of total cost of ownership, set out in Understanding AI Total Cost of Ownership: build plus run plus operate plus change, over the system’s whole life. Apply it to both options, on the same volume, for the same length of service. For the authority, whose figures here and below are invented for illustration, assume 4,000 sensors and three years of service from the day each option goes live.

The purchase costs 400,000 to integrate and test. The subscription is 200 per sensor a year, or 800,000, so run comes to 2.4 million over three years. Alert triage and vendor management add 350,000 a year, or 1.05 million, and configuration changes and renewal work add 50,000 a year, or 150,000. The three-year total is 4.0 million.

The build costs 1.2 million to reach its first day. Compute and storage run at 50 per sensor, or 200,000 a year: 600,000 over three years. Triage plus monitoring, security and on-call support come to 450,000 a year, or 1.35 million. Retraining and upgrades add 150,000 a year, or 450,000. The three-year total is 3.6 million.

Over three years of service, buying totals 4.0 million and building 3.6 million once all four cost families are counted; the memo's 1.2 million saving shrinks to 0.4 million.BuildRunOperateChangeBuy0.4 M2.4 M1.1 M4 MBuild1.2 M0.6 M1.4 M0.5 M3.6 MILLUSTRATIVE NUMBERS
Figure 8.8.4 Illustrative, three years of service at 4,000 sensors. On one boundary, building is 0.4 million cheaper, not 1.2 million.

On a like-for-like boundary the build still wins, but by 0.4 million, a tenth of the purchase cost, rather than the memo’s 1.2 million. That margin is small enough for the remaining steps to overturn.

The price of waiting

Options differ in when they start paying back, and the delay has a price. Donald Reinertsen, writing about product development, called it the cost of delay: the value lost for each unit of time that something useful is late. His advice was blunt: if you quantify only one thing, quantify the cost of delay2. Many investment memos quantify everything else and leave this one out.

The vendor’s service can be live in four months. The build needs twelve. Suppose that, once live, the system is worth about 100,000 a month in water not lost, bursts avoided and crews sent to the right street. That is 25 per sensor a month at 4,000 sensors. Eight months of extra waiting costs 800,000, which by itself wipes out the build’s 0.4 million advantage and lifts the build to 4.4 million, above the purchase.

The build's three-year cost of 3.6 million rises by 0.8 million for eight months of later service, to 4.4 million, above the purchase at 4.0 million.3.6 MBuild TCO+0.8 MEight months later4.4 MBuild with delayILLUSTRATIVE NUMBERS
Figure 8.8.5 Illustrative. Eight months of later service lift the build from 3.6 to 4.4 million, above the purchase at 4.0.

The estimate of monthly value will be uncertain, and that is not a reason to leave it out. A rough figure, stated with its range, is more honest than an implied value of zero. It also forces the useful question: is speed really worth that much here? For a capability that protects revenue or safety, it often is. For an internal convenience, it may be close to nothing, and then the build’s lower run cost deserves more weight.

Whose time does a build use?

A build estimate prices engineers at their salary. The real cost of using them is what they would otherwise have delivered. Accountants call this opportunity cost, and in make-or-buy decisions it works in both directions: buying frees capacity that can be applied to other work1.

Suppose the authority’s data team is five people and the build would take them for a year. If they would otherwise deliver the pressure-management program the regulator has asked for, the build’s true cost includes the value of that program arriving a year late, net of what the team costs. If the team has spare capacity, the opportunity cost is close to zero and the salary is the right number. The question for the committee is therefore not “what does the team cost?” but “what is the team’s best other use, and is that our bottleneck?”

This cost is easy to leave out, because it never appears on any invoice. Yet in organizations where skilled AI engineers are the scarcest resource, it can decide the matter. Price it, or at least name the work that will wait.

Risk-adjust both options

An investment decision compares expected outcomes, not best cases, so each option needs a risk adjustment, and the two options carry risks of different shapes. On the build side, Building the Multi-Year AI Roadmap sets out how IT projects overrun and why the tail matters more than the average; here the committee simply budgets the average overrun of about 27 percent3.

Applied to the authority’s 1.2 million build, that average adds about 0.32 million and lifts the build, already at 4.4 million with delay, to 4.72 million. It is the gentle case. The committee does not need a probability model to go further. It needs to ask what the build costs if it goes badly, how much later it would then arrive, and whether the organization could absorb both.

Buying carries risk too. Vendors reprice when their own costs or strategies change, as AI Unit Economics and Economics at Scale showed with documented cases from 2025. Suppose the authority expects a 20 percent increase at renewal, applied to the third year: that adds 160,000 and lifts the purchase to 4.16 million. Against the build’s 4.72 million, the purchase now leads by about 0.56 million.

Illustrative - with an expected overrun the build rises to 4.72 million; with a 20 percent renewal increase the purchase rises to 4.16 million and leads.4.72MBuild risk-adjustedPlus 0.32 million expected overrun4.16MBuy risk-adjustedPlus 0.16 million renewal increaseILLUSTRATIVE NUMBERS
Figure 8.8.6 Illustrative. Each option carries its own risk: the build overruns, the purchase reprices. The purchase now leads.

Vendors can also change products, degrade service or fail. Those risks are priced in the next step.

Price the exit before you sign

Every purchase is also a future migration. Carl Shapiro and Hal Varian showed that lock-in is an economic condition, not a moral failing: it exists whenever switching costs are high, and the time to evaluate them is before the first commitment, when the buyer still has leverage4. Build vs Buy vs Partner asked whether you could leave. The economic version asks what leaving would cost.

For the authority, a move to another provider or to its own model would mean reintegrating the sensor feeds, rebuilding the history of alerts and confirmed leaks, and running two systems in parallel for a few months. Suppose that is about 300,000. It is paid only if the authority leaves, so it sits beside the three-year comparison rather than inside it: any later move to building must save more than 300,000 to be worth making. The number also sets the contract terms worth negotiating: export of all sensor data, alert history and confirmed-leak labels in a usable format, a term short enough to revisit, and price protection at renewal. Each term lowers the cost of leaving, and a lower cost of leaving is worth paying for.

Find the point where the answer flips

The risk-adjusted comparison favors buying at 4,000 sensors. It is not a permanent verdict. The two options have different cost structures: the purchase is cheap to start and expensive per sensor, the build expensive to start and cheap per sensor. As AI Unit Economics and Economics at Scale showed with break-even volume, a higher fixed cost needs volume to pay for itself. So the useful output is the volume at which the answer flips.

The risk-adjusted cost of buying rises faster with sensor numbers than the cost of building; the lines cross near 5,900 sensors, about 50 percent above the planned 4,000.024573,0004,0005,0006,0007,000Buy (risk-adjusted) ·6.1 MBuild (risk-adjusted)· 5.8 MILLUSTRATIVE NUMBERS
Figure 8.8.7 Illustrative, three-year cost by number of sensors. The lines cross near 5,900 sensors, about half again the plan.

On the chapter’s assumptions the risk-adjusted purchase costs about 640 per sensor over three years plus 1.6 million of fixed cost, while the build costs about 350 per sensor, including its share of the delay, plus 3.32 million. The lines cross near 5,900 sensors. The figures are undiscounted for simplicity; whether the winning option is worth funding at all is a business-case question, answered with the method in Module 03.

The flip point turns a one-off argument into a managed decision. The authority can buy now, on a two-year term with the exit terms above, and write down the trigger: if the sensor network is funded beyond about 6,000 units, or the vendor’s renewal price rises more than expected, the build case reopens. The actual volumes and costs will come from the monthly FinOps rhythm described in Cost Optimization and AI FinOps, which treats a forecast as something to compare with actuals and correct56. Buying first is not a failure to decide. It buys time and evidence for the bigger commitment, at a price the committee can see.

Story: rent the model, or keep building?

Put yourself in a boardroom in Cupertino in March 2025. At its developer conference in June 2024, Apple had announced a more personalized Siri that would understand a user’s own context and act across apps. On 7 March 2025 it said publicly: “It’s going to take us longer than we thought to deliver on these features”7. The company had its own models, its own chips, its own private cloud and one of the largest cash flows in business. It also had a flagship feature a year late. Do you keep building the model yourself, or rent a stronger one from a rival?

Apple announced personalized Siri features in June 2024, delayed them in March 2025, was reported in November 2025 to be renting a Google model, and confirmed the partnership in January 2026.Jun 2024Siri featuresannouncedPersonal contextand actionsMar 2025Delay announced"Longer than we thought"Nov 2025Rental reportedAbout 1 billion ayear, unconfirmedJan 2026PartnershipconfirmedModels based on Gemini
Figure 8.8.8 A company that could afford to build anything chose to rent the model layer and keep the layers around it.

Apple tested models from three outside providers. In November 2025, Bloomberg reported, in an account summarized by other outlets, that Apple was nearing a deal to pay roughly 1 billion a year for a custom model from Google, and that Apple saw the arrangement as temporary until its own models were strong enough. Neither company confirmed the price8. On 12 January 2026 the two companies announced a multi-year collaboration under which the next generation of Apple’s foundation models would be based on Google’s Gemini models. Apple said it had chosen this “after careful evaluation”, and that Apple Intelligence would continue to run on Apple devices and its own Private Cloud Compute9.

Read the decision through this chapter’s steps. The cost of delay had become the dominant number: every further year of building was a year of a promised feature missing from the most important product in the company. The reported rental was large in absolute terms and small against that delay. Apple bought the layer where outside providers were ahead and kept the layers that carry its privacy promise, the devices and the private cloud. And by calling the arrangement temporary it kept a written route back to building, the flip point stated in public. Whether the bet pays off is not yet known. The shape of the decision is the lesson.

What this means for leaders

Refuse any build-versus-buy comparison that sets a subscription against a build estimate. Ask for both options on the module’s four-family boundary, at the same volume, for the same service life. Then ask for three numbers that memos often leave out: the monthly value lost while waiting, the work the build team would otherwise do, and the cost of leaving the vendor. Treat overrun and repricing as expected costs, not footnotes. And never approve a sourcing decision without its flip point, the volume, price or capability change at which you would decide differently, with a date to look again.

Check yourself

  1. A vendor’s annual price can be compared directly with an internal build estimate.
  2. Buying an AI capability removes the need for engineering work.
  3. The value lost while an option is not yet live is a real cost of that option.
  4. The cost of an internal team is its salary.
  5. Budgeting the average IT overrun covers the risk of a build.
  6. Buying first can be a sound way to keep the option to build later.

Reflection: reprice one decision

What comes next

This chapter priced one decision. An AI portfolio holds many, each with its own costs, delays, risks and flip points, and they share budgets, platforms and people. The final chapter of the module, Module 08 Synthesis — Making AI Economically Sustainable, brings the cost structure, scaling economics and investment decisions together into one way of running AI as a sustainable investment.

Laws referenced

EU AI Act · EU

Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744

Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.

  • 2024-08-01 — Entered into force
  • 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
  • 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
  • 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
  • 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
  • 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
  • 2028-08-02 — High-risk obligations for AI in products regulated under Annex I

Last verified 2026-10-06 · official text

Digital Operational Resilience Act (DORA) · EU (financial sector)

Regulation (EU) 2022/2554

Banks, insurers, investment firms and other financial entities must manage ICT risk, report major ICT incidents, test their resilience and manage ICT third-party risk: keep a register of ICT service contracts, include required contract terms (exit, audit, incident support) and plan for provider failure. Critical ICT third-party providers, which can include cloud and AI service providers, fall under direct EU oversight.

  • 2025-01-17 — Applies

Last verified 2026-10-08 · official text

References

  1. Mitchell Franklin, Patty Graybeal and Dixon Cooper. Principles of Accounting, Volume 2: Managerial Accounting, section 10.3 Evaluate and Determine Whether to Make or Buy a Component. OpenStax, Rice University. 2019.
  2. Donald G. Reinertsen. The Principles of Product Development Flow: Second Generation Lean Product Development. Celeritas Publishing. 2009.
  3. Bent Flyvbjerg and Alexander Budzier. Why Your IT Project May Be Riskier Than You Think. Harvard Business Review. 2011.
  4. Carl Shapiro and Hal R. Varian. Information Rules: A Strategic Guide to the Network Economy. Harvard Business School Press. 1999.
  5. FinOps Foundation. FinOps Framework. The Linux Foundation. 2026.
  6. J.R. Storment and Mike Fuller. Cloud FinOps: Collaborative, Real-Time Cloud Value Decision Making, 2nd edition. O'Reilly Media. 2023.
  7. TidBITS. Apple Delays "More Personalized" Siri. TidBITS. 2025.
  8. TechCrunch. Apple nears deal to pay Google $1B annually to power new Siri, report says. TechCrunch. 2025.
  9. Google and Apple. Joint statement from Google and Apple. Google (The Keyword). 2026.

Further reading

Sources last verified 2026-10-08.