Model, Compute and Infrastructure Costs
The advertised price of a model or a GPU is only the top half of a fraction. What decides the economics is the bottom half: how many requests succeed, and how much of the capacity you pay for does useful work. Leaders who ask for cost per successful outcome, and for utilization against a break-even, stop buying the cheapest line item and start buying the cheapest result.
After this chapter you can
- Calculate effective cost per successful outcome, including the cost of fixing failures, and find the fix cost at which two models break even.
- Explain why the cost of a task multiplies calls per task, text in, text out and price per unit, and why agents need cost per completed task.
- Calculate the break-even utilization at which owning or committing to capacity beats renting it.
- Match commitments, batch and premium options to confidence in demand and the business value of speed.
Here is a question worth putting to any leadership team that has just approved a GPU budget. A large technology company studied one of its own production clusters, the shared machines its researchers used to train AI models. Most of the GPUs in the cluster were allocated to a team at any moment, so the dashboard looked full. Of the GPUs that were in use, what share of the hardware was actually doing work? Make a guess before reading on.
The answer was about half. Researchers analyzed two months and about 100,000 jobs on Microsoft’s Philly cluster and found that the hardware utilization of GPUs in use averaged around 52 percent. They warned that the allocation figure, the one that suggested a full cluster, was “misleading” on its own1.
That gap, between what a dashboard or price list says and what the work actually costs, is the subject of this chapter. It applies to a GPU you own and to a model you rent by the request.
The core idea
You can pay for AI compute in two ways. You can pay per request, renting a model from a provider and paying for each question it answers. Or you can pay for capacity, owning or reserving machines by the hour whether or not they are busy. Most organizations do both.
Either way, the advertised number is only the top half of a fraction. Per request, the bottom half is the success rate: a request that fails still costs money, and so does the person who fixes it. Per hour of capacity, the bottom half is utilization: an idle hour costs as much as a busy one. Effective cost is what you pay divided by the useful share of what you paid for.
Understanding AI Total Cost of Ownership placed these costs in the run family of build, run, operate and change. This chapter explains why the run line moves so much, and what a leader should ask about it.
What one request really costs
For a system in use, most of the ongoing model bill comes from inference, using a trained model on each request, rather than from training it; Training vs Inference in Module 02 explains the difference. For a rented model, the cost of one business task is a product of four numbers: how many model calls the task needs, how much text goes in, how much comes out, and the price of each.
Each factor is set by design, not by the provider. Text is metered in tokens, small pieces of words. On one major provider’s published price list in October 2026, every current model charges five times as much for a token it writes as for a token it reads, and the price of output ranges from 0.50 to 50 US dollars per million tokens between the smallest and the largest model, a 100-fold spread2. A system that pastes a whole manual into every request, or writes three paragraphs where one would do, pays for it on every call.
Model choice moves the price as well as the quality: in the Stanford AI Index’s comparison, an early reasoning model that was far better at olympiad mathematics was also nearly six times more expensive and 30 times slower to run3. Sometimes that is a bargain. For a routine task, it is waste.
Often the largest multiplier is the number of calls. A chat assistant answers once. An agent plans, searches, calls tools, checks its own work and answers, and each step is another model call. One provider measured the difference in its own systems.
Agents typically used about four times the tokens of a chat interaction, and multi-agent systems about 15 times4. How agents work is the subject of From AI Assistants to AI Agents. The economic point is simple: an agent’s cost must be measured per completed task, because the price per call hides the number of calls.
Price per request is not cost per outcome
A common model-cost mistake is to compare two models by their price per request. A worked example shows why that comparison can point the wrong way.
Suppose a chemicals maker’s accounts team uses AI to read incoming supplier invoices and fill in the payment record. It tests two models. Model A costs 0.02 per document and gets 90 percent right. Model B costs 0.10 per document, five times as much, and gets 98 percent right. Every document the model gets wrong is caught and fixed by an accounts clerk, and each fix takes about five minutes of that person’s time, which costs about 3.
Look first at the model alone. Divide the price by the success rate and you get the model cost per correct document: about 0.022 for Model A and about 0.102 for Model B. On that view, A is still far cheaper. Now add the people. Per 1,000 documents, A produces 100 failures and B produces 20. At 3 a fix, A’s failures cost 300 and B’s cost 60. The full cost per resolved document is 0.32 for A and 0.16 for B. The model that is five times more expensive per request is half as expensive per outcome.
At 2 million documents a year, procurement sees a model invoice of 40,000 for A against 200,000 for B. The business pays 640,000 for A and 320,000 for B. That is why effective cost per successful outcome, not price per request, is the number to put in the business case:
Effective cost per successful outcome = (model cost + cost of handling the failures) ÷ outcomes delivered.
The comparison also has a break-even. Each failure that B avoids saves the cost of one fix, and B’s extra price is 0.08 a document. The two models cost the same when a fix costs 1. Above that, the more accurate model wins; below it, the cheaper one does. Here a fix costs 3, so the accurate model wins comfortably.
There is often a third option, illustrated with the same invented figures. Suppose half the documents are simple, where A is right 98 percent of the time, and half are complex, where A manages only 82 percent and B 97 percent. Sending simple documents to A and complex ones to B costs 135 per 1,000 documents, about 16 percent less than B alone, as long as the sorting step works and costs little. Choosing and routing between models is covered in The Evolution of AI Models; the economic test is the same one: cost per resolved document.
Renting or owning capacity
Behind every model call is compute: chips, memory, networks and the people who run them. An organization can rent that compute inside a provider’s per-request price, rent machines by the hour from a cloud, or own them. The choice changes the shape of the cost more than its size.
Per-request pricing turns compute into a variable cost. You pay nothing when nobody asks a question, the provider absorbs the peaks, and capacity risk is someone else’s problem. You pay a margin for that, and you accept the provider’s prices, models and schedule. Owned or reserved capacity turns compute into a fixed cost. You pay the same whether the machines are busy or idle, in exchange for a lower price per hour, more control and, for some data, fewer questions about where it is processed. Which is right is a sourcing decision, taken up in Build vs Buy Economics and Investment Decisions. Whichever you choose, one number decides whether fixed capacity was a good buy: utilization.
Utilization: the denominator you own
When capacity is fixed, its cost per useful hour is its cost per available hour divided by utilization. Suppose an owned GPU costs 2 an hour all-in, counting the hardware spread over its life, power, space and the staff who run it. Renting the same GPU on demand costs 6 an hour. At full use, owning is three times cheaper. At 50 percent utilization, each useful hour of the owned GPU costs 4. At 33 percent, it costs 6, the same as renting. Below that, owning is the more expensive choice.
The rule generalizes: the break-even utilization is the owned price per available hour divided by the rented price per hour. Every proposal to buy or reserve capacity should state both numbers and the utilization it expects.
Two things push utilization down. The first is the gap between allocation and use that the opening study found: capacity assigned to a team is not capacity doing work1. The second is peak demand. Capacity has to be sized for the busiest week, not the average one, so the more seasonal the demand, the more of the year the capacity sits idle. Capacity also comes in blocks, which is why compute is one of the largest sources of the step costs that Understanding AI Total Cost of Ownership described.
Paying for flexibility, speed and certainty
Cloud and model providers price the same compute differently depending on what you promise and what you are willing to wait for. The discounts are large, and each one comes with a risk.
One large cloud provider offers up to 72 percent off its on-demand prices for a one- or three-year commitment to a steady level of use5, and up to 90 percent off for spare capacity that it can take back when it needs it6. Model providers make a similar trade on time. At least two major providers charge half price for requests that can wait for an answer, with a 24-hour target72. In the other direction, one provider charges twice the standard rate for faster output2.
The arithmetic of a commitment is the same as the arithmetic of owning. If committed capacity costs 40 percent of the on-demand price, it pays off only if you would otherwise have used it more than 40 percent of the time. The safe rule is to commit for the steady base load you are confident of for the whole term, and to rent the peaks. Latency and availability follow the same logic: an answer in one second, or a system that never goes down, costs more in capacity and redundancy, and is worth it only where a slow or missing answer costs more. The practitioners’ standard text on cloud cost management frames all of this as two separate levers: using less, and paying less for what you use8. The FinOps Foundation’s framework treats rate optimization and usage optimization as distinct capabilities for the same reason9. Cost Optimization and AI FinOps, later in this module, turns these levers, along with caching and context limits, into an operating practice.
Story: the long tail on Alibaba Cloud’s model market
Alibaba Cloud Model Studio is a model marketplace. Customers call any of thousands of models through one service and pay per request. The simple way to run such a service is to give every model its own GPUs, so that each is ready the moment a request arrives. When the company’s engineers measured what that cost, they found demand spread very unevenly. In one sample of 779 models, 167.6 million requests and about 30,000 GPUs, 94.1 percent of the models received only 1.35 percent of the requests. Keeping capacity ready for that long tail of rarely used models tied up 17.7 percent of the GPUs, each handling fewer than 0.2 requests a second. The popular models had the opposite problem: bursts of demand overloaded the capacity reserved for them, and reserving extra dedicated GPUs for the bursts wasted capacity in the same way10.
Nothing on that bill was overpriced. Each GPU-hour cost what it cost. The waste sat in the denominator: capacity sized for each model’s own peak, which most models almost never reached. It is the pattern this chapter has followed from the start, a price per hour that looks reasonable and a cost per useful hour that does not.
The engineers’ answer was not a cheaper GPU. It was a larger share of useful work. Their system, Aegaeon, pools GPUs across models and switches a GPU from one model to another between individual tokens, so that one GPU can serve up to seven models instead of the two or three that earlier pooling methods managed. In a beta deployment lasting more than three months, 47 models, from 1.8 to 72 billion parameters, that had needed 1,192 GPUs ran on 213, a reduction of 82 percent. Over a 70-hour window, average GPU utilization rose from between 13.3 and 33.9 percent to 48.1 percent, with no observed breaches of the service targets. The paper adds that the production deployment keeps spare capacity above the minimum, for peaks and for fault tolerance, so utilization stays below the figures from the laboratory tests10. The trade press reported the 82 percent figure as the company’s claim11.
Put the measured figures into the break-even from earlier in this chapter, using its illustrative prices of 2 an hour to own and 6 to rent. A GPU busy 13 percent of the time costs about 15 per useful hour, far more than renting. At 48 percent it costs about 4, comfortably less. The hardware and its price did not change; the denominator did.
Three lessons travel well beyond a cloud provider. Capacity reserved separately for each workload, each sized for its own peak, is the expensive default, and it is often invisible because every reservation looks sensible on its own. Sharing one pool across many workloads is what keeps the base busy; the peaks are better rented, batched or shared than owned. And the number to ask for is the one these engineers measured: useful work per GPU, not GPUs allocated, which is the same warning the Philly study gave about a training cluster.
What this means for leaders
The model price and the GPU price are the easiest AI costs to find, which is why they dominate procurement conversations. They are the numerators. Leaders add value by asking for the denominators: how many requests succeed and what the failures cost, and how much of the capacity being bought will do useful work against a stated break-even.
Three habits make this routine. Compare models by cost per successful outcome, with the cost of human fixes included, and ask at what fix cost the decision would flip. Ask for any capacity proposal to state its expected utilization, its break-even and its peak. And match commitments to confidence: commit for the base load you are sure of, rent the peaks, and pay for speed or availability only where a slow answer costs more.
Check yourself
- The model with the lower price per request is the cheaper choice.
- An agent’s cost should be measured per completed task, not per model call.
- If almost every GPU in a cluster is allocated, the cluster is well utilized.
- Owning capacity beats renting it only above a break-even utilization.
- A three-year commitment at a large discount is always cheaper than paying on demand.
- Waiting longer for an answer can cut the model price.
Reflection: find your denominator
What comes next
Model, compute and infrastructure are the most visible run costs, and this chapter has shown how much their real cost depends on success and utilization. They are often not the largest costs. The data an AI system relies on, the connections to the systems it works with, and the people who keep it running often cost more. That is the subject of Data, Integration and Operational Costs.
References
- Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee, Junjie Qian, Wencong Xiao and Fan Yang. Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads. Proceedings of the 2019 USENIX Annual Technical Conference (USENIX ATC '19). 2019.
- Anthropic. Pricing (Claude Platform documentation). Anthropic. 2026.
- Stanford Institute for Human-Centered AI (HAI). AI Index Report 2025, Chapter 2: Technical Performance. Stanford University. 2025.
- Anthropic. How we built our multi-agent research system. Anthropic Engineering blog. 2025.
- Amazon Web Services. Compute Savings Plans pricing. Amazon Web Services. 2026.
- Amazon Web Services. Amazon EC2 Spot Instances. Amazon Web Services. 2026.
- Google. Batch API (Gemini API documentation). Google AI for Developers. 2026.
- J.R. Storment and Mike Fuller. Cloud FinOps: Collaborative, Real-Time Cloud Value Decision Making, 2nd edition. O'Reilly Media. 2023.
- FinOps Foundation. FinOps Framework. The Linux Foundation. 2026.
- Yuxing Xiang, Xue Li, Kun Qian, Yufan Yang, Diwen Zhu, Wenyuan Yu, Ennan Zhai, Xuanzhe Liu, Xin Jin and Jingren Zhou. Aegaeon: Effective GPU Pooling for Concurrent LLM Serving on the Market. Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles (SOSP '25), Seoul, 13-16 October 2025. 2025.
- Georgia Butler. Alibaba Cloud claims it can reduce GPU use by 82% with pooling system. Data Center Dynamics. 2025.
Further reading
- Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee, Junjie Qian, Wencong Xiao and Fan Yang. Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads. Proceedings of the 2019 USENIX Annual Technical Conference (USENIX ATC '19). 2019.
- J.R. Storment and Mike Fuller. Cloud FinOps: Collaborative, Real-Time Cloud Value Decision Making, 2nd edition. O'Reilly Media. 2023.
- Anthropic. How we built our multi-agent research system. Anthropic Engineering blog. 2025.
Sources last verified 2026-10-08.