AI Academy · Book
Executives & Directors · Module 08 · Chapter 001

The Economics of AI

The price of an AI answer keeps falling, and AI bills keep rising. Both are true because AI is metered: cost moves with every request, every page of context and every human check. Leaders judge AI by what each unit of value costs, and whether that still holds at ten times the use.

≈ 15 min read

After this chapter you can

  • Explain why the price of an AI answer can fall while the total AI bill rises.
  • Explain why AI cost behaves like a meter rather than a fixed software license.
  • Distinguish the six cost layers and map them to build, run and operate.
  • Treat human review as an economic variable that a business case must state.
  • Ask the unit question and the ten-times question before approving an AI case.

Between November 2022 and October 2024, the cost of getting an answer from a model that performs at the level of the original ChatGPT fell more than 280-fold1. Over roughly the same years, the amount of AI work being bought went the other way. In May 2024 Google was processing 9.7 trillion tokens a month across its products and programming interfaces. A year later it was processing more than 480 trillion2. By May 2026 the figure had passed 3.2 quadrillion, about 330 times the level of two years earlier3.

Google's monthly tokens processed rose from 9.7 trillion in May 2024 to 480 trillion in 2025 and 3.2 quadrillion in 2026, while the price per answer fell.9.7TTokens a monthMay 2024480TTokens a monthMay 20253.2QTokens a monthMay 2026, about 330xSource: Google I/O keynotes, 2025 and 2026 · May 2026
Figure 8.1.1 While the price of an answer fell, the volume of answers bought rose faster. Bills follow volume.

Finance teams have noticed. In the FinOps Foundation’s 2026 survey of cloud cost practitioners, 98 percent said they now manage AI spend, up from 31 percent two years earlier4. A line that was a rounding error in 2024 is now on almost every cloud cost review.

So the answer gets cheaper and the bill gets bigger. That is not a contradiction to be explained away. It is the central fact of AI economics, and an executive who understands why it happens can tell a sound AI business case from a fragile one.

The core idea

The Economics of AI in Module 01 made one point: the model price is one line of the bill. This chapter goes underneath that line and asks how the whole bill behaves.

AI economics is the relationship between three things: the value an AI capability creates, the full cost of delivering it, and how both change as use grows. Written as an equation, net AI value is total benefit minus total cost. The equation is simple. What makes it hard is that both sides move with scale, and they do not move at the same speed.

That is why the usual budget question, “what does the model cost?”, produces a price but not a decision. The better question has two parts: what does each unit of value cost us, and does that still hold when use grows ten times? The rest of this chapter builds the vocabulary to answer it.

Cheaper answers, bigger bills

In 1865 the English economist William Stanley Jevons noticed something odd about coal. James Watt’s engines had made steam power far more efficient, yet Britain was burning more coal than ever. “It is wholly a confusion of ideas,” he wrote, “to suppose that the economical use of fuel is equivalent to a diminished consumption. The very contrary is the truth”5. Cheaper steam power made new uses worth while, and the new uses outgrew the savings. Economists now call this the Jevons paradox, or the rebound effect. The figure below is an illustrative index, not measured data, but it has the shape finance teams see.

Illustrative index lines - the price per answer falls each year while the total AI bill rises, because usage grows faster than prices fall.0125250375500Year 1Year 2Year 3Year 4Price per answer · 18Total AI bill · 460ILLUSTRATIVE NUMBERS
Figure 8.1.2 Illustrative index, Year 1 = 100. When the price of an answer falls, demand for answers can grow faster.

AI follows the same pattern, for reasons any finance leader will recognize. When a request costs a fraction of a cent, teams find many more things worth asking. Each request also tends to grow: longer documents go in as context, longer answers come out, and systems that plan and check their own work call the model several times for a single task. A workflow that once made one call per customer may make ten. Falling prices invite more use, and more use moves the bill.

The rebound is not a reason to slow down. Cheaper intelligence is what makes many good use cases affordable. It is a reason to stop forecasting AI budgets from the price list. The price per request tells you what one unit costs. It does not tell you how many units the organization will consume once the capability is useful, and that number is usually the one that surprises.

A meter, not a license

Many executives learned software economics in an era when it worked like a license. Software, as Carl Shapiro and Hal Varian put it, is costly to produce but cheap to reproduce: the first copy is expensive, and each further copy costs almost nothing6. Once you had paid for access, much of the cost sat still whether people used the system a little or a lot.

AI behaves more like a meter. Every request uses computing power that someone has to pay for, every page of context adds to it, and every answer a person has to check adds labor. The more the organization uses intelligence, the faster the meter runs.

Traditional software behaves like a license whose cost sits still; AI behaves like a meter that runs faster with use, so the meter must stay slower than the value.LicensePay for accessCost sits stillScale barely moves itMeterPay per useMore context costs moreSpeeds up with adoptionKeep the meter slower than the value.
Figure 8.1.3 A license that was cheap at a hundred users is still a license at a million. A meter is a different object at a million.

The clearest evidence comes from the companies that sell AI. In 2020, the investors Martin Casado and Matt Bornstein compared the AI companies they knew with ordinary software businesses. AI companies’ gross margins were often 50 to 60 percent, against a benchmark of 60 to 80 percent or more for comparable software. Cloud computing often took a quarter or more of revenue, and the humans needed to label data and keep models accurate could take another 10 to 15 percent7. Five years later, a study of fast-growing AI start-ups found the fastest growers averaging gross margins of about 25 percent, often negative, while a steadier group averaged about 60 percent8.

AI companies report lower gross margins than comparable software businesses because every unit served carries compute and human cost.BusinessTypical gross marginSourceComparable software60 to 80% or moreCasado and Bornstein, 2020AI companiesOften 50 to 60%Casado and Bornstein, 2020Fastest-growing AI start-upsAbout 25%, often negativeBessemer, 2025
Figure 8.1.4 AI sellers carry real costs for every unit they serve. Those costs reach buyers as usage-based prices.

For a buyer, the lesson is direct. When your vendor has a real cost for every unit it serves, that cost reaches you as usage-based pricing, or as limits and tiers in a contract that looks fixed. Even a self-hosted model is a meter: the hardware is a fixed cost, but power, capacity and the people who run it grow with use. The job is not to stop the meter. It is to keep it running slower than the value.

Count the benefit honestly

The benefit side is easy to count too narrowly, and just as easy to count too generously. AI benefit comes in five kinds: revenue from new products, better conversion or customers who stay; cost savings, the same output for less spend; productivity, time released for people; quality, fewer errors and more consistent decisions; and strategic value, such as learning, market position or a data advantage.

AI benefit comes in five kinds - revenue, cost savings, productivity, quality and strategic value - and productivity counts only when released capacity is used.RevenueNew products, conversion, retentionCost savingsSame output, less spendProductivityCounts only when capacity is usedQualityFewer errorsStrategicLearning and position
Figure 8.1.5 Five kinds of benefit. Label each line financial, strategic, experimental or capability-building.

The most useful evidence for AI’s benefit is measured per unit of work. In a large field study of a generative AI assistant, covering more than 5,000 customer-support agents, those with access to it resolved 15 percent more issues per hour on average9. Notice the unit: issues resolved per hour. A per-unit gain like that is real, but it becomes money only when the organization decides what to do with the released hours: handle more volume, deliver faster or hire fewer people next year. Productivity vs Realized Capacity in Module 03 treats that decision. Here the rule is short: book only the share of released time that someone has committed to turn into output or avoided cost.

Strategic value is real, and it is not a blank check. Label every benefit line as financial, strategic, experimental or capability-building, and give the non-financial lines an objective, a budget limit and a date when they will be judged.

Six layers of cost

On the cost side, the model invoice is the most visible line and often not the largest. At executive level it helps to think in six layers, because each behaves differently as use grows.

Total AI cost sits in six layers - model, compute, data, integration, operations and governance - each tied to the build, run and operate families.ModelPer-request usageRunComputeServing and testingRunDataPreparation, pipelinesBuild and runIntegrationLinks to core systemsBuildOperationsReview, monitoring, supportOperateGovernanceTesting, auditOperate
Figure 8.1.6 Six layers of AI cost, each tagged with its main cost family. Model updates and replacement fall under change.

Model usage is usually priced per request and by how much text goes in and out. Compute is the infrastructure to test and serve the system, rented or owned. Data covers preparing, cleaning and moving the information the system relies on, and it is a layer that is often underestimated. Integration connects the system to the order, customer, identity and workflow systems it has to work with, and in many programs it becomes one of the largest lines. Operations keeps the system trustworthy once it is live: human review, monitoring, evaluation and support. Governance adds risk assessment, testing, audit and documentation; it creates cost, and that cost is weighed against the risk it manages.

This module uses one definition of total cost of ownership throughout: build plus run plus operate plus change, over the system’s whole life. The layers map onto those four families, as the tags in the figure show, and change touches every layer whenever a model is updated or replaced. The next chapter, Understanding AI Total Cost of Ownership, builds the full model, including the costs nobody budgets and the way costs jump in steps.

Human review is a dial

One cost inside operations often decides an AI business case more than any model price does: people checking the output. AI often removes one kind of work and creates another. The system drafts, a person reviews, and sometimes the person corrects. If review is heavy, the savings shrink. The sellers’ margins above show the same thing from the other side: humans in the loop are a recurring cost, not a launch expense7.

Human review is a dial from no review to full review; each rung up adds cost and removes risk, so a business case must state its rung.No reviewErrors reach customersSample reviewCheck a weekly shareException reviewPeople handle flagged casesFull reviewSavings can disappear
Figure 8.1.7 Each rung up adds cost and removes risk.

It helps to treat human involvement as a dial that is set on purpose. With no review, the system is fastest and cheapest, and its errors reach customers. With sample review, people check a share of outputs each week. With exception review, the system flags uncertain or sensitive cases and people handle only those. With full review, every output is checked: the safest setting, and one in which the savings case can disappear.

The right setting depends on what an error costs, not on what the demo showed. When a business case arrives, ask which setting it assumes, what the review rate is, and whether that setting would survive the first serious error. Data, Integration and Operational Costs later in this module prices the work of handling exceptions.

The unit question

An annual AI cost with no unit attached is hard to manage and impossible to compare. The FinOps Foundation calls the discipline of tying technology spend to the value it creates unit economics: metrics that show how technology use affects the value of the organization’s products, services or activities10. The practitioners’ standard text on cloud cost management makes the same case for cloud spend generally11.

The unit question - pick the unit, count full cost and realized value per unit, then test whether the gap holds at ten times the volume.Pick the unitCase, document,customerCost per unitAll six layersValue per unitRealized onlyTen timesDoes thegap hold?Value must outgrow cost.
Figure 8.1.8 Four steps that turn an annual AI cost into an economic case.

The method has four steps. Pick the unit that matches the business model: a resolved case, a processed document, a served customer, a delivered shipment. Count the cost per unit across all six layers, not only the model. Count the value per unit using realized value, not theoretical time saved. Then test the result at ten times today’s volume and ask whether the gap between value and cost per unit still holds.

Scale can push that gap either way. Costs that are mostly fixed, such as the platform, the integration and the core team, spread over more units as use grows, so unit cost falls. Costs that are mostly variable, such as model usage, data processing and review, rise with every request, so unit cost holds or climbs. Two systems with the same bill today can therefore have very different economics at ten times the volume. The test also changes how model choices look: a more expensive model can be the better buy if it cuts errors, retries and review, and a cheaper one can be worse if it adds them.

Skipping this test is expensive. When Gartner predicted in 2024 that at least 30 percent of generative AI projects would be abandoned after proof of concept by the end of 2025, escalating costs and unclear business value were two of the four reasons it gave12. Experimentation vs Production Economics and AI Unit Economics and Economics at Scale take the calculation further. For now, the question is enough.

Story: the assistant that looked cheap

Suppose a regional logistics company handles shipment inquiries all day: where is my parcel, why was the delivery missed, can I change the address? It pilots an AI assistant that drafts answers for its contact center in one region, about 100,000 inquiries a year. The pilot works. Average handling time falls by a quarter.

Before. The team writes the business case the way first cases are often written. It takes the pilot’s model bill and multiplies it by twenty for a national rollout of 2 million inquiries a year, which gives about 1 million a year for model and infrastructure. It takes the agent time released, multiplies it by salary, and books about 9 million a year in savings. Net value: 8 million a year. The case is tidy, and the pilot evidence is real.

After. The finance partner keeps the pilot evidence and adds what production needs. First, human review: the production design checks 30 percent of inquiries, the exceptions and the sensitive ones. That is 600,000 reviews a year at about 4 each, or 2.4 million. Second, integration and operations: links to the tracking and customer systems, monitoring, and a small team to support the assistant, about 1.2 million a year. Third, capacity not released: the operations lead expects 1.6 million of the 9 million in time savings never to show up in the budget.

The first case's net value of 8 million falls to 2.8 million after human review, integration and operations, and unreleased capacity - still positive at about a third.8 MFirst case net−2.4 MHuman review (30%)−1.2 MIntegration and operations−1.6 MCapacity not released2.8 MRebuilt netILLUSTRATIVE NUMBERS
Figure 8.1.9 The rebuilt case still pays, at about a third of the first estimate.

The rebuilt case is worth 2.8 million a year, about a third of the first. Per inquiry, the full running cost is 2.30 (4.6 million over 2 million inquiries) and the realized value is 3.70 (7.4 million over 2 million), so each inquiry earns 1.40. The review rate is the sensitive assumption. If review drifted to half of all inquiries, the net would fall to 1.2 million a year. At about 65 percent review, the case would break even.

The company did not cancel. It approved a staged rollout with two targets that the business owns: cost per resolved inquiry, and a falling review rate. The pilot was useful evidence. It was never the business case. The one-time build costs, which this story leaves aside, are where the next chapter begins.

What this means for leaders

Four habits follow from the economics. First, forecast AI budgets from expected use, not from the price list, because falling prices invite more use. Second, insist on a unit, so that cost and value can be compared and tracked as the system grows. Third, make the human review rate an explicit assumption with an owner, since it often decides whether a case pays. Fourth, test every case at ten times the volume before approving the rollout, and approve in stages with a cost-per-unit target rather than an open budget.

Check yourself

  1. Because the price of an AI answer keeps falling, AI budgets will fall too.
  2. AI economics is just cloud cost management.
  3. AI companies typically run lower gross margins than comparable software businesses.
  4. The model invoice is usually the largest line in an AI business case.
  5. Productivity gains automatically equal cash savings.
  6. Two AI systems with the same cost today can have very different economics at ten times the volume.

Reflection: read your own meter

What comes next

This chapter looked at how AI cost behaves: metered, layered, sensitive to review and to scale. The next step is to see the whole footprint over a system’s life, from the first build to retirement, including the costs nobody budgets. That is the subject of Understanding AI Total Cost of Ownership.

References

  1. Stanford Institute for Human-Centered AI (HAI). AI Index Report 2025, Chapter 1: Research and Development. Stanford University. 2025.
  2. Sundar Pichai. Google I/O 2025: From research to reality. Google (The Keyword blog). 2025.
  3. Sundar Pichai. I/O 2026: Welcome to the agentic Gemini era. Google (The Keyword blog). 2026.
  4. FinOps Foundation. State of FinOps 2026. The Linux Foundation. 2026.
  5. William Stanley Jevons. The Coal Question: An Inquiry Concerning the Progress of the Nation, and the Probable Exhaustion of Our Coal-Mines. Macmillan (2nd edition 1866, Library of Economics and Liberty). 1865.
  6. Carl Shapiro and Hal R. Varian. Information Rules: A Strategic Guide to the Network Economy. Harvard Business School Press. 1999.
  7. Martin Casado and Matt Bornstein. The New Business of AI (and How It's Different From Traditional Software). Andreessen Horowitz. 2020.
  8. Bessemer Venture Partners. The State of AI 2025. Bessemer Venture Partners (Atlas). 2025.
  9. Erik Brynjolfsson, Danielle Li and Lindsey Raymond. Generative AI at Work. The Quarterly Journal of Economics 140(2). 2025.
  10. FinOps Foundation. FinOps Framework. The Linux Foundation. 2026.
  11. J.R. Storment and Mike Fuller. Cloud FinOps: Collaborative, Real-Time Cloud Value Decision Making, 2nd edition. O'Reilly Media. 2023.
  12. Gartner. Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept By End of 2025. Gartner Newsroom. 2024.

Further reading

Sources last verified 2026-10-08.