AI Academy · Book
Executives & Directors · Module 02 · Chapter 010

Large Language Models

A large language model is trained to do one thing, predict the next piece of text, and at scale that single skill yields drafting, summarizing, translating, coding and more. What it learned is a set of patterns, not a set of records. Leaders who keep that distinction clear can use the language capability fully and fetch the facts from where facts live.

≈ 15 min read

After this chapter you can

  • Explain in plain terms what a large language model is and where it sits among foundation models.
  • Describe next-token prediction, causal attention and the pretrain, post-train, prompt sequence at executive level.
  • Judge when a reasoning model is worth its extra time and cost, and why its visible steps are not proof.
  • Distinguish what a model learned in training from the facts supplied at run time.
  • Apply the rule that facts, figures and references come from a system of record, checked by an owner.

In 2011 the computer scientist Hector Levesque proposed a test that he believed only a machine with common sense could pass. Each question is a sentence pair that differs by one word: “The trophy doesn’t fit in the brown suitcase because it is too big”, and the same sentence ending “too small”. A person knows at once that it is the trophy in the first and the suitcase in the second. Nothing in the grammar says so; you need to know how objects and containers behave. At the first public contest, held in 2016, no system did much better than chance. By November 2019 a system built on a pretrained language model scored 90.1 percent on the standard set of 273 questions1.

In three years, the Winograd Schema Challenge went from beyond every machine to solved by language-model-based systems, without evidence that they had acquired reliable common sense.2011Test proposedSentence pairs that seemto need common sense2016First contestNo system much betterthan chance201990.1% on273 questionsBuilt on a pretrainedlanguage model2023Defeat declaredBut little evidence ofreliable common sense
Figure 2.10.1 A test designed to require common sense fell to models trained on text prediction (Kocijan et al., 2023).

The researchers who wrote the challenge’s post-mortem, including some of the field’s best-known skeptics, titled it The Defeat of the Winograd Schema Challenge. They drew two conclusions, and both matter to anyone who now buys or deploys these systems. First, they had been wrong to assume that common-sense reasoning was a prerequisite for the task: models trained to predict text did very well at it. Second, there was still little evidence that those models could reason with common sense consistently1. Large language models are both of those things at once: far more capable than their simple training goal suggests, and less reliable than their fluency suggests.

The core idea

A large language model, or LLM, is a large neural network trained on vast amounts of text and code to predict what comes next, given what came before. Large refers to the billions of learned values, called parameters, and to the data and computing power used to set them. Language covers human languages and, in most current models, programming code. Model means the behavior was learned from examples, not written as rules. In the family described in Foundation Models, an LLM is the foundation model specialized in language.

An LLM is a general language capability that predicts what comes next from learned patterns; it is not a database, not proof of truth and not a finished product.It isA general language capabilityA predictor of what comes nextPatterns learned from text and codeIt is notA database of factsProof that an answer is trueThe finished product
Figure 2.10.2 An LLM is a language capability. Facts, checks and accountability come from the system around it.

That one sentence explains most of what executives see. It explains why one model can answer a customer, summarize a contract and explain old code: each is a way of continuing text sensibly. It also explains the failures. A model built to produce the most plausible continuation will produce one even when the plausible answer and the true answer part company.

Prediction, one token at a time

Take a sentence from accounts payable: “Please pay this invoice within thirty …”. The chart below uses illustrative probabilities, not output from a real model. The model does not look anything up. It scores every possible next piece of text and turns the scores into probabilities. It chooses one, appends it and repeats, until the answer is complete. The pieces are called tokens, often fragments of words rather than whole words; Tokens, Context and Embeddings explains why they drive cost and limits. How the model chooses among likely candidates, and why the same question can get different answers, belongs to How Generative AI Works.

After the words please pay this invoice within thirty, the model rates days far above every other next token, then adds one token and repeats.days81%working11%calendar4%minutes1%ILLUSTRATIVE NUMBERS
Figure 2.10.3 Illustrative probabilities: after “within thirty”, the model rates every candidate token, picks one and predicts again.

The training goal sounds almost trivial, but doing it well across a large share of the world’s written text forces the model to pick up grammar, style, facts that appear often, the structure of code and patterns that look like reasoning. The Winograd result is an example: nobody taught the model about trophies and suitcases. That is the strength. The limit is just as important. A likely continuation is not checked against anything. The model has no separate step in which it asks whether a sentence is true.

Attention: how context steers the prediction

To predict well, the model has to work out which earlier words matter for the next one. Return to Levesque’s sentence and ask the model to continue it: “The trophy doesn’t fit in the brown suitcase because it is too big. The thing that is too big is the …”. To predict the next token, the model looks back over everything before it and weighs which words bear on the answer. “Big” and “doesn’t fit” pull toward trophy. Change “big” to “small” and the same mechanism pulls toward suitcase.

Change one word and the prediction switches from the trophy to the suitcase, because attention weighs the earlier words.The trophy doesn't fit inthe suitcase: it's too ...... bigit = the trophy... smallit = the suitcase
Figure 2.10.4 One word changes what the model attends to, and so what it predicts.

That mechanism is called attention, and it is the core of the Transformer, the design introduced by Vaswani and colleagues in 2017 and used by most major LLMs since2. Two details are worth knowing. In a generating model, each token can attend only to the tokens before it, never to words that have not been written yet; the model reads and writes strictly left to right. And during training, the Transformer processes all the positions in a passage at once rather than one after another, which made it practical to train on far more text with far more computing power. Data, Models and Compute showed how far that scaling went, and why bigger is not automatically better3.

Pretraining, post-training, prompting

An LLM is made in stages. In pretraining, it learns broad patterns by predicting the next token across a vast body of text and code. The result, a base model, is knowledgeable about language but is not yet an assistant. Give it a question and it continues the text in whatever way looks most like its training data; that might be an answer, or it might be more questions in the same style.

Pretraining gives broad patterns, post-training teaches the model to follow instructions, and prompting sets the task without changing the parameters, so one model serves many tasks.PretrainingBroad patterns from vasttext and codePost-trainingLearns to followinstructions helpfullyand safelyPromptingSets the task;parameters unchangedSame model. Different instructions. Different output.
Figure 2.10.5 Pretraining gives capability, post-training gives behavior, the prompt gives the task.

In post-training, developers shape that raw capability into useful behavior. People write examples of good answers and rate the model’s responses, and the model is trained toward the answers people prefer: following instructions, staying on task, declining harmful requests. AI vs Machine Learning described the best-known version of this method, in which people preferred the answers of a much smaller model trained this way over a far larger one that was not4. Post-training is why an assistant behaves like an assistant. The recipe differs between developers, and so does the behavior that results.

In prompting, the user sets the task: summarize this contract in five bullets, translate this product sheet, explain this code to a new engineer. The model’s parameters do not change. As Training vs Inference explained, the model keeps nothing from your request once the answer is delivered, unless the application stores it and supplies it again. This is why one model can do so many jobs without separate training for each: the task arrives with the request.

Reasoning models: what changes for leaders

The newest branch of the family is the reasoning model. It is still a language model, trained differently. In 2022, researchers showed that prompting a large model to write out intermediate steps before its answer markedly improved its results on multi-step arithmetic and logic problems5. Reasoning models are trained to work this way by default, largely with reinforcement learning that rewards answers a program can check, such as the result of a mathematics problem or code that passes its tests6. AI vs Machine Learning showed how far that training lifted one model’s scores. Training vs Inference explained the economics: the extra working is spent at inference, on every answer, which is what researchers call test-time compute.

Reasoning models are stronger on multi-step problems but slower and costlier per answer; standard models are usually enough for drafts and routine replies.ModelMulti-stepproblemsSpeedCost per answerBest forStandardGoodFastLowerDrafts, summaries,routine repliesReasoningStrongerSlowerHigherPlanning, analysis,hard code
Figure 2.10.6 Reasoning models trade speed and cost for strength on multi-step problems.

For a leader, three things change. The first is model choice. There is no longer one best model; there is a better model for each kind of task, and the reasoning model is worth its extra time and cost only where the work has several dependent steps. The second is how you read the visible working. A model’s written steps look like an audit trail, but they are generated text like everything else. In one study, researchers nudged models toward wrong answers with a hidden bias in the question; accuracy fell by as much as 36 percent across 13 tasks, and the step-by-step explanations systematically failed to mention the bias7. The steps can help a reviewer, but they are not proof of how the answer was reached.

The third change is the one most often missed: better reasoning is not better recall. When one developer tested its newer reasoning model on questions about people, it produced false statements on 33 percent of them, against 16 percent for its predecessor; the developer said the newer model made more claims overall, right and wrong, and that more research was needed8. Reasoning improves work on the problem in front of the model. It does not turn the model’s memory into a record.

What it learned is not a record

For enterprise use, this may be the most important distinction in the chapter. What a model learned in training is held in its parameters as patterns. It is not stored as documents, rows or references that can be looked up, and it stopped changing when training stopped. What you give the model at the moment of a request is context: today’s policy, the customer’s record, the question. Context is used for that request and then gone.

So an LLM can appear to know facts without holding any. Ask about something that is widely and consistently written about, and the patterns usually produce the right answer. Ask for something specific, rare, recent or private, such as an order status, a policy clause or a reference to a particular study, and the model will still produce text in the shape of the answer. Researchers call fluent but false or unsupported output hallucination, and it is a well-documented property of language generation, not a defect in one product9.

Part of the reason is the way models are trained and scored. Kalai and colleagues showed that when evaluations award points only for correct answers, a model that guesses outscores one that admits it does not know, so training and benchmarks reward confident guessing10. Their own comparison makes the point.

On a factual test, one model answered almost everything and was wrong 76.8 percent of the time, while another declined 63.2 percent of questions and was wrong 20.8 percent of the time, with similar correct rates.CorrectWrongDeclined to answerRarely declines20.6%76.8%100%Often declines16%20.8%63.2%100%Source: Kalai et al., Nature · 2026
Figure 2.10.7 Shares of all questions. On the same factual test, the model that almost never abstained was slightly more often right and far more often wrong.

The practical rule follows. Anything that must be exact, current or verifiable comes from a system of record at the moment of the question: a database, a document store, a catalog. The model’s job is to find the right words around those facts, not to supply them. Supplying your own documents at question time is called retrieval, and RAG and Enterprise Knowledge covers it. Why fluency so easily passes for accuracy is the subject of How Generative AI Works.

Story: the eighteen sources

Put yourself in the position of a senior official in Tromsø, a city in northern Norway, in March 2025. Your administration has prepared a 120-page report on restructuring the city’s schools and kindergartens, a proposal that would close some schools and move children. It goes out for public consultation, and the council is due to vote in June. The report is well written and its argument is supported by 18 sources, some of them by well-known Norwegian education researchers. Your team has used an AI assistant to help with the work. Would you have a librarian check every one of the 18 references before publication, or do the authors’ names and the quality of the writing satisfy you?

The report went out without that check. Soon after, the local newspaper iTromsø found that only 7 of the 18 sources could be verified. One was a book attributed to a real education professor; the book does not exist. The chief municipal executive called it embarrassing. The consultation was halted, and the council vote planned for 25 June was moved to December11.

In Tromsø, 11 of 18 cited sources in a school report did not exist, the council decision slipped six months and an external review cost 1.2 million kroner.11 of 18Sources that could notbe verifiedFound by a local newspaper6 monthsDecision delayedCouncil vote moved from Juneto December1.2 millionNorwegian kronerCost of the external reviewSource: NRK, 2025 · 2025
Figure 2.10.8 Plausible references cost a city half a year and an external review (NRK, 2025).

An external review by PwC, which cost 1.2 million kroner, found that the errors had not changed the direction of the proposal, which followed a 2021 council decision, but that they could damage trust in the municipality. Its broader finding was that the municipality had not been ready to use AI. Staff had received no AI training and had taught themselves, and PwC recommended an AI strategy, training, quality assurance and clear roles12. When the chat logs were released in September, they showed how the sources had been produced. The employee had asked the assistant for examples from Norway and for reports confirming the argument, with source citations13.

Read that request again with this chapter in mind. It asked a language model for records, and specifically for records that supported a conclusion already reached. The model did what it is trained to do. It produced text in the shape of an academic reference, with plausible titles and the names of researchers who really do write about schools. Nothing in the model checked whether those books existed, and nothing in the process around it did either. The failure was not the model’s alone, as the municipality itself acknowledged; it was the absence of a rule that facts, including references, come from a record and are checked by a person who owns the result.

What this means for leaders

The chapter’s argument reduces to one line: the model supplies the language, and your systems supply the facts. Four practical consequences follow.

Use the language capability, and measure it. Drafting, rewriting, summarizing, translating, explaining and changing tone are what next-token prediction does best, and they do not require the model to know anything you have not given it. The evidence of gains is real but narrow. In one controlled experiment, professionals writing short business documents took about 40 percent less time with an assistant, and graders rated their work about 18 percent higher14. Those were self-contained writing tasks, not whole jobs, so measure the gain in your own work.

Route every fact through a record. Prices, stock, order status, policies, legal clauses, figures and references must come from an authoritative source at the moment they are needed. If a design asks the model to remember them, the design is wrong, however good the demonstration.

Match the model to the task. Use reasoning models where the work has several dependent steps and the extra time and cost are worth it. Do not treat visible reasoning as proof, and do not expect it to fix recall.

Reward honesty about uncertainty. When you evaluate tools, count a confident wrong answer as worse than “I don’t know”, and design workflows in which the model can decline and hand over to a person.

Check yourself

  1. An LLM stores the documents it was trained on and looks answers up in them.
  2. In a generating model, each token attends only to the tokens before it.
  3. Writing a prompt retrains the model for future users.
  4. A reasoning model’s written steps show reliably how it reached its answer.
  5. Better reasoning ability guarantees fewer invented facts.
  6. In Tromsø, the invented references named real researchers.

Reflection: find the request for records

What comes next

You now have the model itself: what it predicts, how attention lets context steer it, how it is trained and where its knowledge ends. To work with it well, you need the units it operates on. What is a token, and why does it set the bill? How much can a model hold in view at once, and what happens to the material in the middle? How does it turn meaning into numbers? The next chapter, Tokens, Context and Embeddings, answers those questions.

References

  1. Vid Kocijan, Ernest Davis, Thomas Lukasiewicz, Gary Marcus and Leora Morgenstern. The defeat of the Winograd Schema Challenge. Artificial Intelligence 325, 103971. 2023.
  2. Ashish Vaswani et al. Attention Is All You Need. Advances in Neural Information Processing Systems 30 (NeurIPS 2017). 2017.
  3. Jared Kaplan et al. Scaling Laws for Neural Language Models. arXiv:2001.08361. 2020.
  4. Long Ouyang et al. Training language models to follow instructions with human feedback. NeurIPS 2022 (arXiv:2203.02155). 2022.
  5. Jason Wei et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS 2022 (arXiv:2201.11903). 2022.
  6. DeepSeek-AI. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv:2501.12948. 2025.
  7. Miles Turpin, Julian Michael, Ethan Perez and Samuel R. Bowman. Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. Advances in Neural Information Processing Systems 36 (NeurIPS 2023). 2023.
  8. OpenAI. OpenAI o3 and o4-mini System Card. OpenAI. 2025.
  9. Ziwei Ji et al. Survey of Hallucination in Natural Language Generation. ACM Computing Surveys 55(12). 2023.
  10. Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala and Edwin Zhang. Evaluating large language models for accuracy incentivizes hallucinations. Nature 653(8116), 1047-1051. 2026.
  11. NRK. Tromsø kommune har henvist til litteratur som ikke finnes i omstruktureringen av skoler. NRK Troms og Finnmark. 2025.
  12. NRK. KI-konklusjon klar: Tromsø kommune var ikke moden for å ta i bruk KI. NRK Troms og Finnmark. 2025.
  13. NRK. KI-skandalen i Tromsø: Ansatt ba om kilder som støtter skolenedleggelser. NRK Troms og Finnmark. 2025.
  14. Shakked Noy and Whitney Zhang. Experimental evidence on the productivity effects of generative artificial intelligence. Science 381(6654). 2023.

Further reading

Sources last verified 2026-10-08.