AI Academy · Book
Executives & Directors · Module 02 · Chapter 007

Training vs Inference

Training creates a model's capability; inference delivers it, on every single request. The first phase changes the model and happens rarely. The second leaves the model untouched, repeats with every use and carries much of the lasting cost, the waiting time and the operational risk.

≈ 14 min read

After this chapter you can

  • Distinguish training, which changes a model's parameters, from inference, which uses them unchanged.
  • Explain why a model does not learn from use, and why learning from feedback must be a designed process.
  • Explain why inference becomes a major running cost as usage grows.
  • Recognize how reasoning models raise cost and latency per answer, and set that effort by request.
  • Decide whether a change belongs in the model or in what the model reads at answer time.

Training is the part of AI that makes headlines, with its vast clusters of chips and its eye-watering bills, so it is natural to assume that is where the energy goes. Google’s own measurements say otherwise. Between 2019 and 2021, machine learning accounted for 10 to 15 percent of all the energy Google used, and when its engineers measured one week of April in each of those three years, the answer was the same every time: about three-fifths of the machine learning energy went to inference, using trained models, and about two-fifths to training them. The authors gave the reason plainly: the many services with a billion or more users that run machine learning every time someone searches, types or takes a photo1.

About 60 percent of Google's machine learning energy in 2019 to 2021 went to inference and about 40 percent to training.Inference (using models)60%Training (building models)40%Source: Patterson et al., IEEE Computer · 2019-2021
Figure 2.7.1 At Google, using models took about three-fifths of machine learning energy in each year measured; building them took two-fifths.

That split is the whole subject of this chapter in one number. Training is what creates an AI model. Inference is what happens every time the model is used. They are different activities, with different costs, different risks and different people in charge, and confusing them leads to predictable mistakes in budgets, data policies and vendor negotiations.

Learning versus using

The idea fits in one line: training creates capability; inference delivers it.

Training is how a model learns. It is shown a very large number of examples, and its internal settings, called parameters, are adjusted again and again until its outputs improve. When training ends, the parameters are frozen and saved. That saved set of numbers is the trained model.

Inference is how a model is used. The trained model receives an input, such as a question, an image or a customer record, and produces an output: an answer, a label, a score. Its parameters do not change. The same model, with the same parameters, answers the first request and the millionth.

Training changes a model's parameters and creates capability; inference uses the trained model with fixed parameters and delivers that capability.TRAININGThe model learns - itsparameters change.Creates capabilityINFERENCEThe model is used - itsparameters stay fixed.Delivers capabilityvs
Figure 2.7.2 The test is simple: if the parameters change, it is training. If they do not, it is inference.

A manufacturing picture helps. Training is making a mold: the toolmakers shoot trial parts, measure them, adjust the tool and shoot again until the parts come out right. It is slow, skilled and expensive, and it happens before production starts. Inference is running the press. The finished mold makes part after part, and making a part does not reshape the mold. Every part costs material, energy and machine time, so over a long run the parts can cost far more than the mold did. The picture breaks down in one place: a model can be replaced by a new version far faster than a mold can be recut. But the core holds. You build the capability once, and you pay for every use.

Training: a loop that runs before launch

How Machine Learning Learns described the loop at the heart of training: take an example, make a prediction, measure the error against the right answer, nudge the parameters to make the error smaller, and repeat2. Training is the only phase in which the parameters change, and it ends with a fixed, saved model.

That model is a file of parameters that can be copied, versioned, tested and deployed. Fine-tuning runs the same loop again, smaller and on focused data, and produces a new version of that file. It is still training. Anything that changes the parameters is training; anything that does not is something else.

Inference: the model is used, not changed

Inference is a straight path rather than a loop. An input goes in, the trained model processes it, and an output comes out. The shape is the same whether the input is a question to a language model, a photograph to a vision model or a customer record to a churn model.

Inference is a straight path from input through a trained model with fixed parameters to an output; prompts and documents change the input, and only training changes the model.InputQuestion, instructions,documentsTrained modelParameters fixedOutputAnswer, label or scorePrompts change the input. Only training changes the model.
Figure 2.7.3 Everything an application adds at answer time changes the input, not the model.

This is where executives most often blur the two phases, because modern applications add a great deal at inference time. An instruction such as “answer in three bullet points” is part of the prompt, which Prompting and Context Engineering — Executive Mental Model covers. Today’s policy document, fetched from a store and placed beside the question, is retrieval, the subject of RAG and Enterprise Knowledge — Executive Mental Model. One customer’s account details, supplied for one answer, are runtime data, not training data. All three can change the answer dramatically. None of them changes the model. The moment the request is finished, the model is exactly as it was before.

Why the model does not learn while you use it

Many staff believe an AI assistant learns from every conversation, so correcting it today will make it better tomorrow. In standard use it does not. The conversation may be stored, and depending on the product’s settings and the contract it may later be used to improve a future model, but the model answering you is not changing as you type.

That design is deliberate, and a well-known exception shows why. In March 2016 Microsoft launched Tay, a chatbot for young adults that was meant to “get better and better” through interaction with the public. Within its first 24 hours, in Microsoft’s own words, “a coordinated attack by a subset of people exploited a vulnerability in Tay”, and it began posting offensive content. Microsoft took it offline and apologized3. A model that learns directly from its users learns from all of them, including the ones trying to break it.

There are two further reasons. First, training on new material can damage old skills. Neural networks are prone to what researchers call catastrophic forgetting: trained on something new, they can lose competence at what they did well before4. Every change to the parameters therefore needs testing against everything the model is supposed to do. Second, organizations need to know which model gave which answer. A model whose parameters shift with every conversation cannot be tested, versioned or audited.

So learning from feedback is a process someone has to design, not a side effect of use. Corrections are collected, reviewed and then acted on. In a well-run process, many become fixes to the documents or instructions the model reads at inference time, and only a few justify a new, tested version of the model itself.

Inference is where the running cost lives

Training is one large job before launch. Inference repeats after launch, with every user and every request. Ten users make ten requests; ten million users make ten million. Google’s energy split is what that arithmetic looks like at scale1. It measures energy rather than money, and it is one company’s figure, not an industry average, but spending follows the same arithmetic.

Illustrative cumulative cost - a one-time training cost of 10 stays flat while cumulative inference grows with adoption and passes it in the sixth quarter.05101520Q1Q2Q3Q4Q5Q6Q7Q8Training (one-time) ·10 MInference(cumulative) · 18 MILLUSTRATIVE NUMBERS
Figure 2.7.4 Illustrative: a one-time training cost stays flat, while cumulative inference cost grows with use and overtakes it.

Suppose building a model costs 10 million, once, and inference starts small and grows as people adopt the product. By the sixth quarter the cumulative inference bill, 10.5 million, has passed the training bill, and it keeps climbing. The numbers are invented; the shape is not. If you use a provider’s model, the split is even starker: the provider paid for training, and almost every bill you receive is for inference.

The price of each answer has also been falling fast. As Why AI, Why Now? showed, the cost of a query at a fixed level of performance fell more than 280-fold in under two years5. Cheaper answers do not make inference unimportant. They make it easy to use far more of it, which is why the whole bill, and not only the price per answer, belongs in every business case. The Economics of AI covers that bill.

Two things set the size of each request. For language models, cost usually tracks the amount of text going in and coming out, so long prompts and long answers cost more; Tokens, Context and Embeddings explains why. And one user request can trigger several model calls when a system checks its work or uses tools, which multiplies both cost and waiting time.

Reasoning models move more work into inference

A recent development shifts the balance further. Some models reason before they answer: instead of replying at once, they generate intermediate steps, check and revise them, and only then give a final answer. Products often call this a thinking mode. Researchers call it test-time compute, computation spent at inference, on each answer, rather than during training. Large Language Models explains how such models are trained; this chapter is concerned with what they do to the cost of every answer.

The idea grew from a 2022 finding that simply asking a model to show its intermediate steps improved its answers on arithmetic and logic problems6. When one developer released an early reasoning model in September 2024, it reported that performance improved both with more training and “with more time spent thinking (test-time compute)”7. Test-time compute had become a second lever, alongside training, for raising capability.

The lever has a price. Every intermediate step is generated, so each answer costs more and takes longer. The research also shows how to pull the lever well. Charlie Snell and colleagues found, on a competition-mathematics benchmark, that allocating extra computation according to how hard each question was proved more than four times as efficient as spending it uniformly. On problems a smaller model could already partly solve, extra thinking time let it beat a model 14 times its size for the same total computation8. These are benchmark results. Published evidence that reasoning modes improve business outcomes is still thin, so the gain has to be tested on your own requests.

A dial from answering at once to reasoning at length; reasoning spends more compute and time on each answer, so match the effort to the question.Answer at onceLower cost, secondsReason at lengthMore compute, slowerMatch effort to the questionReason where it pays
Figure 2.7.5 Reasoning effort is a dial set per request, not a free upgrade switched on for everything.

For leaders, that is the practical lesson. A routine look-up does not need deep reasoning; a hard diagnosis or a complex plan may. The question is which requests deserve the expensive setting, and who decides.

Serving: inference is an operation

A trained model does nothing until it is served: loaded onto hardware, connected to applications and scaled to meet demand. Serving is judged on different measures from training. Latency is how long one answer takes. Throughput is how many answers the system can produce per minute. Availability is whether it is there at all when people need it. An excellent model that cannot carry the traffic quickly and economically is not a viable product.

Production AI is a lifecycle - train, evaluate, serve, infer on every request, then monitor and update - and inference is the phase that repeats.Before launchTrainProvider or youBefore launchEvaluateTest before relyingon itLaunchServeScale, speed, uptimeEvery requestInferCost and timeper answerOngoingMonitor andupdateVersion everychange
Figure 2.7.6 Training and inference sit inside a longer lifecycle, and inference is the only phase that repeats with every use.

The two phases need different infrastructure. Training wants large clusters working on one long job. Inference wants speed, availability and a low cost per request, close to where users are. With a managed service, you do not run the hardware, but you pay for every answer. If you host a model yourself, you take on the hardware, scaling, security and uptime. Either way, version the model, the prompts and the documents together, because together they define what users experience.

Story: an assistant that should “learn every night”

This story is an illustrative composite, not a documented case.

A global manufacturer of industrial compressors runs an AI assistant for its field-service engineers. An engineer standing at a stopped machine types the fault code and what she sees; the assistant suggests likely causes and the right repair procedure from the manuals. Six months in, two complaints reach the head of service. Engineers correct the assistant in the chat and are annoyed when it repeats the same mistake a week later. And product engineering issues new service bulletins every week, while the list of superseded part numbers changes almost daily.

The head of service brings three options to the operations committee.

Three options for keeping a service assistant current - change nothing, fine-tune every night, or split facts from model changes - each resting on a different belief about what changes a model.How should theassistant keepup?AChange nothingBet: it learns from correctionsBFine-tune every nightBet: retraining keeps it currentCSplit the changeby typeBet: facts at answer time,model on a schedule
Figure 2.7.7 Three options, each resting on a different belief about what changes a model.

Option A assumes the assistant is learning from the corrections. Option B fine-tunes the model every night on the day’s corrections and the new bulletins. Option C keeps changing facts in the documents the assistant reads at answer time, sends corrections to a reviewed queue, and changes the model itself only in scheduled, tested releases. Before reading on, decide which you would approve.

Option A rests on the misconception this chapter has dismantled. The corrections were being stored, but the model’s parameters had not moved since launch; the same superseded part number would keep coming back until someone fixed the source.

Option B sounds diligent, but it turns the update problem into a training problem. Nightly fine-tuning means 365 new model versions a year, each of which should be tested against every fault type the assistant already handles, because narrow retraining can erode existing skills4. If a full test cycle takes a few days, the nightly versions either ship untested or pile up in a queue. Worse, unreviewed corrections become training data: one engineer’s mistaken “correction” becomes the model’s belief. And a bulletin issued on Monday morning still waits until the night’s run.

Illustrative comparison - nightly fine-tuning creates 365 model versions a year, while option C makes four tested releases and puts a new bulletin into answers the same day.365Model versions a yearUnder nightly fine-tuning4Tested releases a yearUnder option CSame dayNew bulletin in answersOnce the document is updatedILLUSTRATIVE NUMBERS
Figure 2.7.8 Illustrative: separating changing facts from model changes makes updates both faster and safer.

The committee chose C. A new bulletin reaches answers as soon as its document is updated, with no training run at all. Corrections go to a named owner, who decides whether each one is a wrong document, a missing instruction or a genuine gap in the model’s behavior. Only the last kind feeds a quarterly model release, tested before engineers see it. The lesson generalizes: before asking how to retrain a model, ask whether the model is what needs to change. That principle, retrieve rather than retrain, is taught in full in RAG and Enterprise Knowledge — Executive Mental Model.

What this means for leaders

Four lessons follow from the distinction.

  1. Budget for inference as a running cost. Training may have been paid for, by you or by a provider. Every answer after launch costs compute, and that cost grows with volume, with the length of each request and with reasoning effort.
  2. Set the reasoning dial deliberately. Reasoning modes can improve hard answers, but they raise cost and waiting time on every request that uses them. Decide which requests deserve them.
  3. Do not assume the model learns from use. Design the feedback loop, name its owner, and find out in writing where conversation data goes and whether it is used for training.
  4. Change the right thing. Facts that change belong in the documents and systems the model reads at answer time. Change the model only when its behavior must change, and test every new version.

Check yourself

  1. An AI assistant learns from every conversation it has.
  2. Once a model is trained, using it is nearly free.
  3. Putting today’s policy document into the prompt retrains the model.
  4. Fine-tuning is a form of training.
  5. Switching on a reasoning mode can raise both the cost and the waiting time of each answer.
  6. Retraining a model on new material can only add to what it knows.

Reflection: where does your AI really run?

What comes next

Training and inference are two phases of one model’s life, and both draw on the same three resources: the data a model learns from and reads, the model itself, and the compute that builds and runs it. The next chapter, Data, Models and Compute, brings the three together and explains why they have become the economic engine of modern AI.

References

  1. David Patterson, Joseph Gonzalez, Urs Hölzle, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier and Jeff Dean. The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink. IEEE Computer 55(7), 18-28; arXiv:2204.05149. 2022.
  2. Ian Goodfellow, Yoshua Bengio and Aaron Courville. Deep Learning. MIT Press. 2016.
  3. Peter Lee. Learning from Tay's introduction. The Official Microsoft Blog. 2016.
  4. James Kirkpatrick et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences 114(13), 3521-3526; arXiv:1612.00796. 2017.
  5. Stanford Institute for Human-Centered AI (HAI). AI Index Report 2025, Chapter 1: Research and Development. Stanford University. 2025.
  6. Jason Wei et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS 2022 (arXiv:2201.11903). 2022.
  7. OpenAI. Learning to reason with LLMs. OpenAI. 2024.
  8. Charlie Snell, Jaehoon Lee, Kelvin Xu and Aviral Kumar. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. arXiv:2408.03314 (ICLR 2025). 2024.

Further reading

Sources last verified 2026-10-08.