The Economics of AI
The price of an AI answer keeps falling, and AI bills keep rising. Both are true because AI is metered: cost moves with every request, every page of context and every human check. Leaders judge AI by what each unit of value costs, and whether that still holds at ten times the use.
After this chapter you can
- Explain why the price of an AI answer can fall while the total AI bill rises.
- Explain why AI cost behaves like a meter rather than a fixed software license.
- Distinguish the six cost layers and map them to build, run and operate.
- Treat human review as an economic variable that a business case must state.
- Ask the unit question and the ten-times question before approving an AI case.
Between November 2022 and October 2024, the cost of getting an answer from a model that performs at the level of the original ChatGPT fell more than 280-fold1. Over roughly the same years, the amount of AI work being bought went the other way. In May 2024 Google was processing 9.7 trillion tokens a month across its products and programming interfaces. A year later it was processing more than 480 trillion2. By May 2026 the figure had passed 3.2 quadrillion, about 330 times the level of two years earlier3.
Finance teams have noticed. In the FinOps Foundation’s 2026 survey of cloud cost practitioners, 98 percent said they now manage AI spend, up from 31 percent two years earlier4. A line that was a rounding error in 2024 is now on almost every cloud cost review.
So the answer gets cheaper and the bill gets bigger. That is not a contradiction to be explained away. It is the central fact of AI economics, and an executive who understands why it happens can tell a sound AI business case from a fragile one.
The core idea
The Economics of AI in Module 01 made one point: the model price is one line of the bill. This chapter goes underneath that line and asks how the whole bill behaves.
AI economics is the relationship between three things: the value an AI capability creates, the full cost of delivering it, and how both change as use grows. Written as an equation, net AI value is total benefit minus total cost. The equation is simple. What makes it hard is that both sides move with scale, and they do not move at the same speed.
That is why the usual budget question, “what does the model cost?”, produces a price but not a decision. The better question has two parts: what does each unit of value cost us, and does that still hold when use grows ten times? The rest of this chapter builds the vocabulary to answer it.
Cheaper answers, bigger bills
In 1865 the English economist William Stanley Jevons noticed something odd about coal. James Watt’s engines had made steam power far more efficient, yet Britain was burning more coal than ever. “It is wholly a confusion of ideas,” he wrote, “to suppose that the economical use of fuel is equivalent to a diminished consumption. The very contrary is the truth”5. Cheaper steam power made new uses worth while, and the new uses outgrew the savings. Economists now call this the Jevons paradox, or the rebound effect. The figure below is an illustrative index, not measured data, but it has the shape finance teams see.
AI follows the same pattern, for reasons any finance leader will recognize. When a request costs a fraction of a cent, teams find many more things worth asking. Each request also tends to grow: longer documents go in as context, longer answers come out, and systems that plan and check their own work call the model several times for a single task. A workflow that once made one call per customer may make ten. Falling prices invite more use, and more use moves the bill.
The rebound is not a reason to slow down. Cheaper intelligence is what makes many good use cases affordable. It is a reason to stop forecasting AI budgets from the price list. The price per request tells you what one unit costs. It does not tell you how many units the organization will consume once the capability is useful, and that number is usually the one that surprises.
A meter, not a license
Many executives learned software economics in an era when it worked like a license. Software, as Carl Shapiro and Hal Varian put it, is costly to produce but cheap to reproduce: the first copy is expensive, and each further copy costs almost nothing6. Once you had paid for access, much of the cost sat still whether people used the system a little or a lot.
AI behaves more like a meter. Every request uses computing power that someone has to pay for, every page of context adds to it, and every answer a person has to check adds labor. The more the organization uses intelligence, the faster the meter runs.
The clearest evidence comes from the companies that sell AI. In 2020, the investors Martin Casado and Matt Bornstein compared the AI companies they knew with ordinary software businesses. AI companies’ gross margins were often 50 to 60 percent, against a benchmark of 60 to 80 percent or more for comparable software. Cloud computing often took a quarter or more of revenue, and the humans needed to label data and keep models accurate could take another 10 to 15 percent7. Five years later, a study of fast-growing AI start-ups found the fastest growers averaging gross margins of about 25 percent, often negative, while a steadier group averaged about 60 percent8.
For a buyer, the lesson is direct. When your vendor has a real cost for every unit it serves, that cost reaches you as usage-based pricing, or as limits and tiers in a contract that looks fixed. Even a self-hosted model is a meter: the hardware is a fixed cost, but power, capacity and the people who run it grow with use. The job is not to stop the meter. It is to keep it running slower than the value.
Count the benefit honestly
The benefit side is easy to count too narrowly, and just as easy to count too generously. AI benefit comes in five kinds: revenue from new products, better conversion or customers who stay; cost savings, the same output for less spend; productivity, time released for people; quality, fewer errors and more consistent decisions; and strategic value, such as learning, market position or a data advantage.
The most useful evidence for AI’s benefit is measured per unit of work. In a large field study of a generative AI assistant, covering more than 5,000 customer-support agents, those with access to it resolved 15 percent more issues per hour on average9. Notice the unit: issues resolved per hour. A per-unit gain like that is real, but it becomes money only when the organization decides what to do with the released hours: handle more volume, deliver faster or hire fewer people next year. Productivity vs Realized Capacity in Module 03 treats that decision. Here the rule is short: book only the share of released time that someone has committed to turn into output or avoided cost.
Strategic value is real, and it is not a blank check. Label every benefit line as financial, strategic, experimental or capability-building, and give the non-financial lines an objective, a budget limit and a date when they will be judged.
Six layers of cost
On the cost side, the model invoice is the most visible line and often not the largest. At executive level it helps to think in six layers, because each behaves differently as use grows.
Model usage is usually priced per request and by how much text goes in and out. Compute is the infrastructure to test and serve the system, rented or owned. Data covers preparing, cleaning and moving the information the system relies on, and it is a layer that is often underestimated. Integration connects the system to the order, customer, identity and workflow systems it has to work with, and in many programs it becomes one of the largest lines. Operations keeps the system trustworthy once it is live: human review, monitoring, evaluation and support. Governance adds risk assessment, testing, audit and documentation; it creates cost, and that cost is weighed against the risk it manages.
This module uses one definition of total cost of ownership throughout: build plus run plus operate plus change, over the system’s whole life. The layers map onto those four families, as the tags in the figure show, and change touches every layer whenever a model is updated or replaced. The next chapter, Understanding AI Total Cost of Ownership, builds the full model, including the costs nobody budgets and the way costs jump in steps.
Human review is a dial
One cost inside operations often decides an AI business case more than any model price does: people checking the output. AI often removes one kind of work and creates another. The system drafts, a person reviews, and sometimes the person corrects. If review is heavy, the savings shrink. The sellers’ margins above show the same thing from the other side: humans in the loop are a recurring cost, not a launch expense7.
It helps to treat human involvement as a dial that is set on purpose. With no review, the system is fastest and cheapest, and its errors reach customers. With sample review, people check a share of outputs each week. With exception review, the system flags uncertain or sensitive cases and people handle only those. With full review, every output is checked: the safest setting, and one in which the savings case can disappear.
The right setting depends on what an error costs, not on what the demo showed. When a business case arrives, ask which setting it assumes, what the review rate is, and whether that setting would survive the first serious error. Data, Integration and Operational Costs later in this module prices the work of handling exceptions.
The unit question
An annual AI cost with no unit attached is hard to manage and impossible to compare. The FinOps Foundation calls the discipline of tying technology spend to the value it creates unit economics: metrics that show how technology use affects the value of the organization’s products, services or activities10. The practitioners’ standard text on cloud cost management makes the same case for cloud spend generally11.
The method has four steps. Pick the unit that matches the business model: a resolved case, a processed document, a served customer, a delivered shipment. Count the cost per unit across all six layers, not only the model. Count the value per unit using realized value, not theoretical time saved. Then test the result at ten times today’s volume and ask whether the gap between value and cost per unit still holds.
Scale can push that gap either way. Costs that are mostly fixed, such as the platform, the integration and the core team, spread over more units as use grows, so unit cost falls. Costs that are mostly variable, such as model usage, data processing and review, rise with every request, so unit cost holds or climbs. Two systems with the same bill today can therefore have very different economics at ten times the volume. The test also changes how model choices look: a more expensive model can be the better buy if it cuts errors, retries and review, and a cheaper one can be worse if it adds them.
Skipping this test is expensive. When Gartner predicted in 2024 that at least 30 percent of generative AI projects would be abandoned after proof of concept by the end of 2025, escalating costs and unclear business value were two of the four reasons it gave12. Experimentation vs Production Economics and AI Unit Economics and Economics at Scale take the calculation further. For now, the question is enough.
Story: the assistant that looked cheap
Suppose a regional logistics company handles shipment inquiries all day: where is my parcel, why was the delivery missed, can I change the address? It pilots an AI assistant that drafts answers for its contact center in one region, about 100,000 inquiries a year. The pilot works. Average handling time falls by a quarter.
Before. The team writes the business case the way first cases are often written. It takes the pilot’s model bill and multiplies it by twenty for a national rollout of 2 million inquiries a year, which gives about 1 million a year for model and infrastructure. It takes the agent time released, multiplies it by salary, and books about 9 million a year in savings. Net value: 8 million a year. The case is tidy, and the pilot evidence is real.
After. The finance partner keeps the pilot evidence and adds what production needs. First, human review: the production design checks 30 percent of inquiries, the exceptions and the sensitive ones. That is 600,000 reviews a year at about 4 each, or 2.4 million. Second, integration and operations: links to the tracking and customer systems, monitoring, and a small team to support the assistant, about 1.2 million a year. Third, capacity not released: the operations lead expects 1.6 million of the 9 million in time savings never to show up in the budget.
The rebuilt case is worth 2.8 million a year, about a third of the first. Per inquiry, the full running cost is 2.30 (4.6 million over 2 million inquiries) and the realized value is 3.70 (7.4 million over 2 million), so each inquiry earns 1.40. The review rate is the sensitive assumption. If review drifted to half of all inquiries, the net would fall to 1.2 million a year. At about 65 percent review, the case would break even.
The company did not cancel. It approved a staged rollout with two targets that the business owns: cost per resolved inquiry, and a falling review rate. The pilot was useful evidence. It was never the business case. The one-time build costs, which this story leaves aside, are where the next chapter begins.
What this means for leaders
Four habits follow from the economics. First, forecast AI budgets from expected use, not from the price list, because falling prices invite more use. Second, insist on a unit, so that cost and value can be compared and tracked as the system grows. Third, make the human review rate an explicit assumption with an owner, since it often decides whether a case pays. Fourth, test every case at ten times the volume before approving the rollout, and approve in stages with a cost-per-unit target rather than an open budget.
Check yourself
- Because the price of an AI answer keeps falling, AI budgets will fall too.
- AI economics is just cloud cost management.
- AI companies typically run lower gross margins than comparable software businesses.
- The model invoice is usually the largest line in an AI business case.
- Productivity gains automatically equal cash savings.
- Two AI systems with the same cost today can have very different economics at ten times the volume.
Reflection: read your own meter
What comes next
This chapter looked at how AI cost behaves: metered, layered, sensitive to review and to scale. The next step is to see the whole footprint over a system’s life, from the first build to retirement, including the costs nobody budgets. That is the subject of Understanding AI Total Cost of Ownership.
References
- Stanford Institute for Human-Centered AI (HAI). AI Index Report 2025, Chapter 1: Research and Development. Stanford University. 2025.
- Sundar Pichai. Google I/O 2025: From research to reality. Google (The Keyword blog). 2025.
- Sundar Pichai. I/O 2026: Welcome to the agentic Gemini era. Google (The Keyword blog). 2026.
- FinOps Foundation. State of FinOps 2026. The Linux Foundation. 2026.
- William Stanley Jevons. The Coal Question: An Inquiry Concerning the Progress of the Nation, and the Probable Exhaustion of Our Coal-Mines. Macmillan (2nd edition 1866, Library of Economics and Liberty). 1865.
- Carl Shapiro and Hal R. Varian. Information Rules: A Strategic Guide to the Network Economy. Harvard Business School Press. 1999.
- Martin Casado and Matt Bornstein. The New Business of AI (and How It's Different From Traditional Software). Andreessen Horowitz. 2020.
- Bessemer Venture Partners. The State of AI 2025. Bessemer Venture Partners (Atlas). 2025.
- Erik Brynjolfsson, Danielle Li and Lindsey Raymond. Generative AI at Work. The Quarterly Journal of Economics 140(2). 2025.
- FinOps Foundation. FinOps Framework. The Linux Foundation. 2026.
- J.R. Storment and Mike Fuller. Cloud FinOps: Collaborative, Real-Time Cloud Value Decision Making, 2nd edition. O'Reilly Media. 2023.
- Gartner. Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept By End of 2025. Gartner Newsroom. 2024.
Further reading
- J.R. Storment and Mike Fuller. Cloud FinOps: Collaborative, Real-Time Cloud Value Decision Making, 2nd edition. O'Reilly Media. 2023.
- FinOps Foundation. FinOps Framework. The Linux Foundation. 2026.
- Martin Casado and Matt Bornstein. The New Business of AI (and How It's Different From Traditional Software). Andreessen Horowitz. 2020.
Sources last verified 2026-10-08.