AI Academy · Book
Executives & Directors · Module 08 · Chapter 006

AI Unit Economics and Economics at Scale

Scale does not improve AI economics by itself; it multiplies whatever each unit leaves behind, and spreads fixed cost only over units that are actually used. Leaders who know the contribution per unit, the break-even volume and the assumption that moves both can tell growth that creates value from growth that only creates activity.

≈ 15 min read

After this chapter you can

  • Compute contribution per unit from value per unit and complete variable cost, including expected review.
  • Compute break-even volume as fixed cost divided by contribution per unit, and explain why it is a threshold, not a verdict.
  • Explain fixed-cost absorption and why low adoption raises cost per unit.
  • Identify when scale raises unit cost through mix, review share, step costs and falling value per unit.
  • Test whether a price structure fits a skewed cost structure, and find the assumption that moves a case most.

Suppose a translation services firm sells AI-assisted document translation at a flat 40 per document. In its first quarter it translates 10,000 business documents. Each one costs it about 15 to deliver: model usage, file handling and a linguist’s quick check. The firm also carries 150,000 a quarter of platform and staff cost that does not move with volume. The quarter ends with 100,000 of profit, and the sales team is asked to grow. It does. In the second quarter volume doubles to 20,000 documents, and profit falls to 50,000.

Illustrative - a firm doubled its document volume and halved its profit because each new document cost 45 to deliver at a price of 40.2xDocuments translated10,000 to 20,000 a quarter-50%Quarterly profit100,000 to 50,000-5Left by eachnew documentPrice 40, cost 45ILLUSTRATIVE NUMBERS
Figure 8.6.1 Illustrative. Volume doubled and profit halved, because every new document cost more to deliver than it earned.

Nothing broke. The new clients sent long technical manuals rather than letters and contracts. Each manual needed many more model calls and much more of a linguist’s time, so it cost about 45 to deliver against a price of 40. The first 10,000 documents still left 25 each, or 250,000. The new 10,000 each lost 5, or 50,000. Contribution fell from 250,000 to 200,000, the fixed 150,000 did not change, and profit halved. The firm had grown, and every unit of growth had made it poorer.

The core idea

The Economics of AI asked leaders to pick a business unit, such as a case, a document or a customer, and to test whether the gap between value and cost per unit holds at ten times the use. This chapter does that calculation. It rests on one subtraction and one division, both borrowed from managerial accounting. Contribution per unit is the value of one unit minus the variable cost of delivering it1. Total contribution is that figure times volume, and what remains after fixed cost is the net result.

Unit economics runs from value per unit, less variable cost, to contribution, which volume multiplies before fixed cost is deducted.Value per unitPrice, orrealized valueVariable costModel, data,reviewContributionWhat eachunit leavesNet resultTimes volume,less fixed costScale multiplies contribution, whatever its sign
Figure 8.6.2 The whole method. Scale helps only while each unit leaves something behind.

The consequence is the argument of this chapter. Volume multiplies contribution, so growth is good news only while contribution is positive and stays positive as volume rises. Volume also spreads fixed cost, but only across units that are actually delivered. And all three terms move with scale: fixed cost per unit falls, variable cost per unit can rise, and value per unit can fall. An AI case that is sound at today’s volume can weaken at ten times, and one that loses money today can be sound at ten times. The arithmetic says which.

Contribution: what each unit leaves behind

Take a simple example. An AI service creates 50 of value per customer served each year. Delivering it costs 15 per customer in variable cost: 10 for model usage and data processing, and 5 for the expected cost of human review. Each customer therefore leaves 35 to cover fixed cost and return. The contribution margin, contribution divided by value, is 70 percent1.

Illustrative - value of 50 per customer, less 10 of model and data cost and 5 of expected review, leaves a contribution of 35.50Value per customer−10Model and data−5Expected review35ContributionILLUSTRATIVE NUMBERS
Figure 8.6.3 Illustrative. Each customer creates 50, costs 15 to serve and leaves 35: a 70 percent contribution margin.

Three disciplines keep this number honest. The variable cost must be complete. It includes the human review that a share of cases needs, which Data, Integration and Operational Costs showed can quietly outweigh the model, and the cost of failed attempts, which Model, Compute and Infrastructure Costs folded into the cost per successful outcome. The value must be realized value. For an internal system, minutes saved count only when someone turns them into avoided cost or extra output, as Productivity vs Realized Capacity in Module 03 showed. And the unit must be the one the business buys. A model request is a technical unit; one finished case may need many requests, so a low cost per request says almost nothing about contribution.

Break-even: how much volume the fixed cost needs

Contribution answers whether each unit pays. Break-even answers how many units are needed to pay for everything else. The standard formula is fixed cost divided by contribution per unit1. If the service above carries 350,000 a year of fixed cost, for its platform, its engineering team and its base capacity, then 350,000 divided by 35 is 10,000 customers.

Illustrative - total value rises at 50 per customer and total cost starts at 350,000 and rises at 15, so they cross at 10,000 customers.02505007501,00005,00010,00015,00020,000Total value · 1,000 kTotal cost · 650 kILLUSTRATIVE NUMBERS
Figure 8.6.4 Illustrative. Fixed cost of 350,000 and contribution of 35 per customer meet at 10,000 customers.

Below 10,000 customers the service loses money; above it, each extra customer adds 35. At 20,000 customers it earns 350,000 a year. This is a different threshold from the one in Model, Compute and Infrastructure Costs, which asked how busy owned capacity must be to beat renting it. Break-even volume asks how much business the whole case needs. Both are questions about the denominator.

Break-even is a threshold, not a verdict. A case that just clears it may still be a poor use of capital if an alternative earns more, if it carries risk the figures leave out, or if the volume forecast is the least certain number in it. Return on investment is the business of AI ROI and Value Realization in Module 03; break-even tells you how far the forecast volume sits above or below the line, which is often the first thing a finance committee wants to know.

Scale spreads fixed cost, if the units arrive

The friendliest effect of scale is fixed-cost absorption. A platform costing 1 million a year adds 100 to each unit at 10,000 units, 10 at 100,000 units and 1 at 1,000,000 units. That is why many AI cases look poor at pilot volume and attractive at full volume.

Illustrative - a fixed platform cost of 1 million a year falls from 100 per unit at 10,000 units to 1 per unit at one million units.10,000 units100100,000 units101,000,000 units1ILLUSTRATIVE NUMBERS
Figure 8.6.5 Illustrative. A 1 million platform adds 100 per unit at 10,000 units and 1 per unit at a million.

The same arithmetic punishes low adoption. A platform sized and paid for one million users that attracts 100,000 active users carries 10 of fixed cost per user, ten times the plan. The usual cause is not technical. Adoption is an economic variable, and a rollout that reaches a tenth of its intended users has a tenth of the intended denominator. Fixed costs also stay fixed only within a range. Understanding AI Total Cost of Ownership showed how review teams, capacity and support rise in steps once volume crosses a threshold2. A case that counts on absorption must show where the next step falls.

Scale can hurt as well as help

The less friendly effects arrive with the same growth. The first is mix: the earliest volume is usually the easiest, and later volume brings longer documents, harder cases and new languages, as the translation firm discovered. The second is review: when the share of cases needing a person rises, review cost grows faster than volume, the exception-rate trap of Data, Integration and Operational Costs. The third is value: the first uses of a system are usually the most valuable ones, and the next million units may be worth less each than the first hundred thousand.

Scale lowers unit cost when fixed cost dominates and new units resemble old ones; it raises unit cost when variable cost dominates or new units are harder.Scale helps whenFixed cost dominatesNew units resemble old onesValue per unit holdsScale hurts whenVariable cost dominatesNew units are harderReview share risesAsk how cost grows relative to value
Figure 8.6.6 Volume is not the test. The test is how cost per unit and value per unit move as volume grows.

AI businesses show this at market scale. In 2020, before today’s large language models, investors at Andreessen Horowitz reported that AI companies’ gross margins were often in the 50 to 60 percent range, below the 60 to 80 percent or more of comparable software companies, because compute and humans in the loop grow with every customer served3. The pattern survived the arrival of those models. In 2025, Bessemer Venture Partners put the average gross margin of its fastest-growing AI cohort, which it calls Supernovas, at about 25 percent, often negative early on, against about 60 percent for its steadier Shooting Stars4. Revenue grew quickly in both groups; where the variable cost of serving it grew as fast, margin did not follow.

The practical test compares growth rates. If volume rises 10 percent and total cost 5 percent, cost per unit falls. If volume rises 10 percent and cost 20 percent, cost per unit rises by about 9 percent. Ask the same of value: if the next 10 percent of volume adds 15 percent to value and 10 percent to cost, scale is working. Some economies appear only at volume, such as routing simple work to a smaller model, because their own fixed cost needs a large base to pay back. Choosing and running those levers belongs to the next chapter.

Averages hide the heavy users

An average contribution per unit can be healthy while a segment destroys value. That matters most when price is flat and cost is not. Two documented cases from 2025 show the pattern in AI products; the companies are named here as evidence of a pricing pattern, not as recommendations. In June 2025 Cursor, which sells an AI coding tool, moved its individual Pro plan from a fixed number of requests to a monthly pool of usage at API prices. Its explanation was that new models spend more tokens on longer tasks, and that the hardest requests cost an order of magnitude more than simple ones; it also admitted the change was poorly communicated and offered refunds5. A month later Anthropic announced weekly limits on its flat-fee subscription plans from 28 August, to curb users running its coding tool continuously in the background, and said the limits would affect fewer than 5 percent of subscribers6.

Flat seat pricing suits even usage, per-request pricing tracks cost but makes bills volatile, and per-outcome pricing needs a measurable outcome.Price perFits cost whenRiskSeat orsubscriptionUsage is evenHeavy users erode marginRequest or volumeCost follows usageVolatile bills for buyersCompletedoutcomeOutcomes can be measuredDisputes over what countsHybridA base plus usage tiersHarder to explain
Figure 8.6.7 A price structure should fit the cost structure. Where it does not, scale widens the gap.

In both cases a small share of users drove a large share of cost under a price that did not move with it, and the seller changed the structure rather than the price. The lesson applies in both directions. If your organization sells or charges back an AI service, break contribution down by segment, using the cost-to-serve view of Data, Integration and Operational Costs, and make sure the price structure fits the cost structure. If a service is priced at 10 per customer and costs 15 to serve, growth only enlarges the loss. If you buy one, expect the vendor’s price to move toward its own cost structure, as both did, and model your heaviest internal users before you sign.

Test the range, then find the assumption that moves it

A single forecast is a guess with decimals. A unit-economics case should be run across a range, for example a fifth of the forecast, the forecast itself, and two, five and ten times it, with value, variable cost, fixed cost and capacity recomputed at each point rather than scaled. Managerial accountants call the next step sensitivity analysis: asking what happens to the result if price, units, variable cost or fixed cost changes1.

Illustrative - at 20,000 customers a 20 percent adverse change cuts net from 350,000 to 290,000 for variable cost, 280,000 for fixed cost, 210,000 for volume and 150,000 for value.Base case350 kVariable cost +20%290 kFixed cost +20%280 kVolume -20%210 kValue per unit -20%150 kILLUSTRATIVE NUMBERS
Figure 8.6.8 Illustrative. The same 20 percent miss costs 60,000 on variable cost and 200,000 on value per unit.

Return to the service at 20,000 customers, earning 350,000 a year. A 20 percent rise in variable cost cuts that to 290,000; a 20 percent rise in fixed cost, to 280,000; 20 percent fewer customers, to 210,000. A 20 percent fall in value per customer cuts it to 150,000. Value per unit moves the result most because it is the largest number in the subtraction, and it is usually the least measured. Most review time goes to the model price; the assumption that deserves it is often the value.

The FinOps Foundation treats unit economics as a continuing practice, not a one-time calculation: the unit metrics are tracked as the system runs and compared with the forecast78. The most sensitive assumptions are the ones to measure first and watch most closely.

Story: the safety team that counted cases, not calls

This is an illustrative composite, not a documented case, told as a before and after inside one organization.

A mid-sized pharmaceutical company receives adverse-event reports about its medicines from doctors, patients and partners. Each report must be entered, coded and assessed as a safety case. Its drug-safety team deployed an AI system that reads incoming reports and drafts the case. The cases it cannot complete with confidence, and every serious case under the company’s own quality rules, go to a safety specialist.

Before. The team’s monthly report tracked model calls processed and cost per call, which had fallen by a third over the year. On that basis it proposed bringing every region onto the system, raising volume from 50,000 cases a year to 500,000. Nobody had asked what one completed case was worth or what the platform needed to break even.

After. Finance asked the question the dashboard could not answer: what is one completed case worth? Its value was 50, the price an outside provider charged the company for the same intake work. Variable cost was 18 per case for the model, data handling and quality checks, so each case left 32 before human handling. Ten percent of cases went to a specialist at 25 each, an expected 2.5 per case, which brought contribution to 29.5. Fixed cost was 1.6 million a year for the platform, validation and the engineering team. Break-even was 1.6 million divided by 29.5: about 54,200 cases. At 50,000 cases the system earned 1,475,000 of contribution, 125,000 short of its fixed cost. The dashboard had been reporting a success that lost money every year.

Finance then tested scale. At 500,000 cases the platform would need more capacity, regional validation and support, taking fixed cost in steps to 4 million a year. Contribution of 29.5 on 500,000 cases is 14.75 million, a net 10.75 million. Two assumptions were tested. Reports from new regions arrive in more languages and formats, so automation might fall from 90 to 75 percent: expected handling rises to 6.25 per case, contribution to 25.75, and the net at 500,000 to about 8.9 million. And the outside provider might cut its price, lowering the value of a completed case from 50 to 35: contribution falls to 14.5 and the net to 3.25 million. Even with both, the expansion nets about 1.4 million; it stops paying only if the value per case falls below about 32 with automation at 75 percent. The table keeps the two cost bases apart: today’s 1.6 million of fixed cost at 50,000 cases, and the 4 million the platform would need at 500,000.

Illustrative - today the system runs 125,000 short; at 500,000 cases on 4 million of fixed cost it nets 10.75 million in the base case, about 8.9 million at 75 percent automation, 3.25 million at a value of 35 and about 1.4 million with both.ScenarioContribution per caseBreak-even volumeNet a yearToday: 50,000 cases, 1.6million fixed29.5About 54,200Minus 125,000At 500,000 cases, 4 millionfixed: base29.5About 135,60010.75 millionAt 500,000: automationfalls to 75%25.75About 155,300About 8.9 millionAt 500,000: value fallsto 3514.5About 275,9003.25 millionAt 500,000: both10.75About 372,100About 1.4 million
Figure 8.6.9 Illustrative. The case fails at today’s volume and fixed cost, and holds at 500,000 in every scenario; value per case moves it most.

The decision changed shape. At 50,000 cases the problem was fixed-cost absorption, not the model price the team had been working to lower. The expansion was approved region by region, with cost and contribution per completed case reported every quarter in place of cost per call. Because value per case moved the result most, the company fixed the outside provider’s price for three years before migrating the second region.

What this means for leaders

Ask for the unit economics of every AI case that is meant to scale, and ask in the business’s own unit. A good answer gives the value per unit and where it comes from, the variable cost per unit including review and failures, the contribution, the fixed cost, and the break-even volume, with today’s and next year’s volume marked against it. It then shows the case at a range of volumes with step costs included, names the assumption that moves the result most, and breaks contribution down by segment wherever usage is uneven. A case that offers cost per request instead, or growth in usage as proof of success, has not yet been made.

Check yourself

  1. Cost per model request is a good measure of an AI system’s unit economics.
  2. Break-even volume is fixed cost divided by contribution per unit.
  3. Growing the volume of an AI service always lowers its cost per unit.
  4. A 1 million platform spread over one million units adds 1 to the cost of each unit.
  5. Under a flat price, a small share of heavy users can make a growing AI product less profitable.
  6. Once a case passes break-even volume, the investment is justified.

Reflection: one system, ten times the use

What comes next

Knowing the contribution per unit tells you where AI spending earns its keep and where it does not. Keeping it that way, month after month, as usage, prices and models change, needs an operating discipline: visibility of spend by unit, allocation to owners, budgets and alerts, and the optimization levers that scale makes worth pulling. That discipline is the subject of Cost Optimization and AI FinOps.

References

  1. Mitchell Franklin, Patty Graybeal and Dixon Cooper. Principles of Accounting, Volume 2: Managerial Accounting, chapter 3 Cost-Volume-Profit Analysis. OpenStax, Rice University. 2019.
  2. Mitchell Franklin, Patty Graybeal and Dixon Cooper. Principles of Accounting, Volume 2: Managerial Accounting, section 2.2 Identify and Apply Basic Cost Behavior Patterns. OpenStax, Rice University. 2019.
  3. Martin Casado and Matt Bornstein. The New Business of AI (and How It's Different From Traditional Software). Andreessen Horowitz. 2020.
  4. Bessemer Venture Partners. The State of AI 2025. Bessemer Venture Partners (Atlas). 2025.
  5. Cursor (Anysphere). Clarifying our pricing. Cursor blog. 2025.
  6. TechCrunch. Anthropic unveils new rate limits to curb Claude Code power users. TechCrunch. 2025.
  7. FinOps Foundation. FinOps Framework. The Linux Foundation. 2026.
  8. J.R. Storment and Mike Fuller. Cloud FinOps: Collaborative, Real-Time Cloud Value Decision Making, 2nd edition. O'Reilly Media. 2023.

Further reading

Sources last verified 2026-10-08.