AI Academy · Book
Executives & Directors · Module 02 · Chapter 009

Foundation Models

A foundation model is trained once on broad data and then adapted to many tasks. That reuse changed the economics of AI, but it comes with two conditions leaders cannot delegate: the model is only the bottom layer of any product built on it, and every product built on it inherits its flaws along with its strengths.

≈ 15 min read

After this chapter you can

  • Explain what makes a model a foundation model - broad pretraining at scale, then adaptation to many tasks.
  • Explain why reuse changed the economics of building AI, and where the costly stages now sit.
  • Separate the foundation model from the context, application and workflow built around it.
  • Choose the lightest adaptation that does the job, and say why fine-tuning is a poor way to add facts.
  • Recognize homogenization - shared blind spots across applications - and what a model does not know.

In October 2018 a team at Google published a language model called BERT. They had trained it once, without any labeled examples, on English Wikipedia and a large collection of books. Then they tested it on the field’s standard benchmarks, the shared exams researchers use to compare systems: answering questions about a passage, judging whether one sentence follows from another, telling a positive review from a negative one. For each task they added one small output layer and trained briefly on that task’s examples.

Before reading on, make a guess. Researchers had spent years building a separate, specialized system for each of those tasks. On how many of them did the one model set a new record?

The answer was eleven. BERT set new state-of-the-art results on eleven language tasks and raised the headline score on the GLUE benchmark to 80.5 percent, 7.7 points above the previous best1. A year later the same model was inside one of the most used products in the world. In October 2019 Google said BERT would help its search engine understand one in ten English-language searches in the United States2. By October 2020 it was used in almost every query in English3.

BERT, trained once, set records on eleven language tasks, then helped with one in ten US English searches in 2019 and almost every English query by 2020.11Language tasks ledOne pretrained model, one smalllayer per task1 in 10US English searchesHelped by the same model,October 2019Almost allEnglish queriesBy October 2020Source: Devlin et al.; Google · 2018-2020
Figure 2.9.1 One model, trained once, beat the specialists on eleven tasks and then went to work inside a global product.

Nobody called BERT a foundation model at the time, because the term did not exist yet. But it showed the pattern that now shapes how many organizations get their AI: train one model broadly, then reuse it many times.

Train broadly once, adapt to many tasks

In August 2021 more than a hundred researchers at Stanford published a long report that gave the pattern its name. They defined foundation models as models “trained on broad data at scale and adaptable to a wide range of downstream tasks”4. The name was chosen for the role, not the technology. A foundation model is something other systems are built on.

One broadly trained foundation model at the center serves many tasks, from summarizing and translating to writing code and answering questions.Modeltrained broadlyonceSummarizeTranslateExtractClassifyDraftWrite codeAnswer
Figure 2.9.2 Reuse is the defining feature. One broadly trained model serves many tasks that once needed a model each.

The definition has two halves, and both matter to a leader. The first is reuse: one model, adapted many times. The second is easy to miss. The Stanford authors stressed that a foundation model is central but incomplete by design. It is a component waiting to be built on, not a finished product. Many of the costly mistakes organizations make with these models come from forgetting that second half.

Foundation models are not only about text. Depending on how it was trained, a model may work with images, audio, video or several of these at once, which Multimodal AI takes up later in this module. Large language models are the best-known kind, and the next chapter is about them.

The old way repeated the whole loop

To see why reuse matters, look at how most machine learning was built until recently. Each task ran its own loop.

Before foundation models, every task ran its own loop of collecting labeled data, training, testing and maintenance.Collect dataLabeled for this taskTrainFor this task onlyTestOn this task's casesRun and maintainRetrain asdata driftsEvery new taskrepeats the loop
Figure 2.9.3 In the task-specific approach the whole cost repeats for each new task.

Someone collected data and labeled it for one task. A team trained a model that did only that task, tested it on that task’s cases, and then ran, monitored and retrained it as the world changed. That approach works, and it still does. A demand forecast, a churn score or a maintenance alert is often best served by exactly this kind of focused model, which is cheaper to run and easier to validate.

The problem is repetition. A seventh task means a seventh loop, with new data, new training, new testing and new upkeep. Foundation models break the pattern by moving the most expensive step, broad training, to whoever builds the model, and by letting everyone else start from the result.

How a foundation model is made

A foundation model’s life has four stages, and only the last two usually involve your organization.

A foundation model is pretrained on broad data, post-trained to follow instructions, adapted to a use and run inside applications; providers pay for the first two stages.PretrainBroad data, atscale, no labelsPost-trainTaught to followinstructionsAdaptFitted to your useUseAnswers insideyour applicationProviders pay for the first two stages; many customers share them.
Figure 2.9.4 The expensive stages happen once, upstream. Your decisions start at adaptation.

Pretraining comes first. The model learns from an enormous mix of material, such as text, code, images or audio, and it learns from the structure of the data itself, so nobody has to label each example. BERT learned by filling in words hidden from sentences; the next chapter shows how today’s language models learn by predicting what comes next. As Data, Models and Compute showed, performance in this stage improves predictably as data, model size and compute grow together5.

Post-training then turns the raw model into one that follows instructions and behaves helpfully. Large Language Models explains how.

Adaptation fits the model to a particular use, and use is the model answering inside an application. These two stages are where your organization’s choices live.

The economics follow from the order. Pretraining at the frontier is the business of a few very large organizations, and, as Why AI, Why Now? showed, it gets more expensive every year while using a model gets cheaper. A provider pays for pretraining once and spreads the cost across many customers. That lowers the cost of trying an idea dramatically. It does not end engineering work. It moves it from training models to building what goes around them.

The model is not the application

The most important distinction in this chapter is between the model and the thing your people and customers actually use.

The foundation model is only the bottom layer; context and tools, the application and the workflow sit above it and become more specific to your organization.WorkflowPeople, process, measures, accountabilityApplicationInterface, rules, permissionsContext andtoolsYour documents, data and systemsFoundationmodelGeneral capability, shared by manyMORESPECIFICTOYOU
Figure 2.9.5 The model is the bottom layer. The higher the layer, the more specific it is to you, and the more it deserves your effort.

At the bottom sits the foundation model, general capability shared by many customers. Few organizations build this layer themselves. Above it are context and tools: your documents and data supply what the model does not know, and tools let the application reach your systems, for example to read a project schedule or check stock. Then comes the application: the interface, the business rules and the permissions that decide who may see what. At the top is the workflow: the people who use the output, the process it sits in, the measures that show whether it works and the person accountable when it does not.

The name itself is the right picture. A foundation is expensive to pour and engineered to carry a great deal, and similar foundations sit under very different buildings. But nobody lives in a foundation. People live in the floors built on top of it. Buying access to a foundation model is not the same as owning a finished application. The picture breaks in one useful place: a model can be swapped or upgraded under a running application, which a real foundation cannot, so every layer above it has to be tested again when it changes.

Adapt only as far as the job needs

How much should you change the model to fit your use? As little as the job needs. Adaptation is a ladder, and each rung adds cost, data work and upkeep.

Adaptation rises from instructing the model to grounding it in your documents, connecting tools and fine-tuning; each rung adds cost and upkeep.InstructInstructions and examplesSTART HEREGroundSupply your documentsConnectReach your systemsFine-tuneAdjust the model itself
Figure 2.9.6 Climb only as far as the job needs. Many uses never need the top rung.

The first rung is instructing the model: clear instructions, a few good examples and the right context in the request. That is often enough, and Prompting and Context Engineering treats it in depth. The second is grounding: supplying your own documents at the moment of the question, so answers rest on your sources rather than on the model’s memory. RAG and Enterprise Knowledge explains how. The third is connecting tools, so the application can read from and act in your systems.

Only then comes fine-tuning, further training on your own examples. It has become far cheaper than it was. A method called LoRA, for example, cut the number of parameters that have to be trained by a factor of 10,000, and the memory needed by a factor of three, compared with fully retraining a very large model, with similar or better quality6. But cheap is not the same as appropriate. Fine-tuning is good at changing behavior, format and style. It is poor at teaching a model new facts: a 2024 study found that models learn unfamiliar facts from fine-tuning slowly, and that as they do, they become more prone to making things up7. If the goal is for the model to know your information, the answer is usually grounding, not training.

Training your own foundation model from scratch sits far above this ladder and is rarely justified. Smaller models were the subject of Data, Models and Compute; hosted versus open-weight models belong to Build vs Buy vs Partner and AI Platform Strategy in Module Four.

Shared strengths, shared flaws

Reuse has a second side. When many applications are built on the same few models, they become alike underneath. The Stanford report called this homogenization. By 2021, it noted, nearly all of the best language systems were adapted from one of a handful of foundation models such as BERT4. That gives enormous leverage, because one improvement to the base model reaches every application built on it. It also means that the base model’s defects are inherited by every application adapted from it.

Applications built on the same foundation model inherit its broad skills and improvements, and also its blind spots, biases and failures on the same cases.What you inheritBroad language skillImprovements to the baseFast start on new tasksWhat you also inheritBlind spotsBiases in the training dataFailure on the same cases
Figure 2.9.7 Every application built on a model inherits its strengths and its flaws together.

This is not only a theoretical worry. In 2023 researchers examined commercial AI systems from different providers across text, images and speech. The systems failed on the same cases more often than chance would predict, and some people were misclassified by every system available to them8.

For a leader, the lesson is practical. Switching to another provider does not guarantee escape from a blind spot, because the alternatives may share it. A provider’s published scores describe the model on average, not on your customers or your cases. Test at your own layer, on your own tasks, and keep a route to a person for the cases the model gets wrong.

What the model does not know

A foundation model learned general concepts and public knowledge up to the date its training data ended. It may understand accounting, engineering and contract law very well. That is not the same as knowing your business.

A foundation model learned general, public knowledge up to a cutoff; the current, private information a business runs on sits below the waterline.WHAT THE MODEL LEARNEDGeneral concepts · Public knowledge· Up to a training cutoffWHAT YOUR BUSINESS RUNS ONThis quarter's numbersInternal policiesToday's inventoryCustomer recordsLast week's decisions
Figure 2.9.8 The information your business runs on sits below the waterline, out of the model’s sight unless you supply it.

Your business runs on what sits below the waterline: this quarter’s numbers, internal policies, today’s inventory, customer records, last week’s decisions. None of that is in the model unless your application supplies it. Nor is anything that happened after its training ended.

The model cannot go and look for today’s facts. Your application has to bring them to it.

Story: the summer section nobody checked

On Sunday, 18 May 2025, a big-city American daily newspaper delivered a 64-page summer guide to its print subscribers. The section had been licensed from a national syndicator and was produced outside the newsroom. No editor at the paper had reviewed it. A version also ran in a second large daily9.

A licensed summer section shipped unreviewed on 18 May 2025; readers found fake books the next day, and the paper's review on 29 May found errors in every story.18 MaySection ships64 pages,licensed, unreviewedNext dayReaders spot itTen recommended booksdo not exist20 MayPaper respondsPulled from e-paper,no charge29 MayFull reviewEvery story had errors
Figure 2.9.9 The model was used raw. The failure was every missing layer above it, and readers found it first.

Within about a day, readers on social media noticed that the summer reading list recommended books that did not exist. Of the fifteen titles, ten were inventions, credited to real and well-known novelists. The paper then fact-checked all ten stories in the section and found multiple errors or unverifiable claims in every one: experts who could not be found at the organizations named, a quotation that a well-known chef’s office said she never gave, and a gardening writer quoted at a 2024 event who had died in 202310. The freelance writer admitted he had used an AI tool and published what it produced without checking it, and the syndicator ended its relationship with him9.

The chief executive of the paper’s parent organization published her own post-mortem. She did not stop at the writer. She traced the failure through each place a person could have caught it: the writer, the syndicator’s editing, a circulation team that trusted licensed content and never sent the pages to editors, and her own approval of the practice without looking closely. The fix was not a better model. The paper removed the section from its digital edition, did not charge print subscribers for it, stopped buying editorial special sections from that syndicator, required third-party content to be labeled and reviewed by its standards editors, and drafted an AI policy requiring fact-checking, editorial review and disclosure of significant AI use, overseen by a committee from across the organization11.

Read through the lens of this chapter, the model did exactly what a general model does. Asked for summer books, it produced plausible titles in the style of real authors, because nothing in it knew which books exist. Researchers call this behavior hallucination: fluent, confident statements that are unsupported or simply wrong12. Why it happens and how to manage it is the subject of Accuracy, Hallucination and Reliability in Module Six. Fluency is a property of the model; accuracy is a property of the system built around it. What was missing was every layer above it: no grounding in a real catalog, no rule that sources be checked, no review step, and no named owner of the outcome. And because the flaw entered upstream, it traveled to every outlet that took the section. That is homogenization in miniature. One defect at the base reaches everything built on it.

What this means for leaders

Four lessons follow. The first is to start from an existing model and spend above it. Pretraining is the most expensive stage and providers already share its cost across many customers. Your budget belongs in the layers no provider can supply: your data, your rules and permissions, your workflow and the people who own the outcome.

The second is to adapt only as far as the job needs. Begin with instructions, ground the model in your documents before you consider changing it, and fine-tune only when behavior or format still falls short. Never fine-tune to teach it facts that you could simply supply.

The third is to treat the model’s flaws as yours. Every application built on it inherits them, and so do the alternatives that share its training. Test on your own cases, by people who know the work, before every release and whenever the model underneath changes.

The fourth is to name an owner above the model. In the newspaper’s case, every layer assumed another had checked. Accountability for an AI-assisted output has to sit with a person in the workflow, not with the model or its provider.

Check yourself

  1. A foundation model is defined by its size.
  2. One pretrained model can outperform specialized systems on many different tasks.
  3. Fine-tuning is the best way to teach a model our company’s facts.
  4. Applications built on the same model tend to share its blind spots.
  5. Switching to a different provider guarantees we escape a model’s blind spot.
  6. A strong model was the missing piece in the newspaper’s summer section.

Reflection: find the missing layer

What comes next

Foundation models give organizations broad, reusable capability, and the most influential kind today works with language. Why can one language model summarize, translate, draft and write code without separate training for each job? The next chapter, Large Language Models, opens the box.

References

  1. Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL-HLT 2019 (arXiv:1810.04805, first posted October 2018, Google AI Language). 2019.
  2. Pandu Nayak. Understanding searches better than ever before. The Keyword (Google blog). 2019.
  3. Prabhakar Raghavan. How AI is powering a more helpful Google. The Keyword (Google blog). 2020.
  4. Rishi Bommasani et al. On the Opportunities and Risks of Foundation Models. Stanford CRFM (arXiv:2108.07258). 2021.
  5. Jared Kaplan et al. Scaling Laws for Neural Language Models. arXiv:2001.08361. 2020.
  6. Edward J. Hu, Yelong Shen, Phillip Wallis et al. LoRA: Low-Rank Adaptation of Large Language Models. ICLR 2022 (arXiv:2106.09685, Microsoft). 2022.
  7. Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart and Jonathan Herzig. Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?. EMNLP 2024 (arXiv:2405.05904). 2024.
  8. Connor Toups, Rishi Bommasani, Kathleen A. Creel, Sarah H. Bana, Dan Jurafsky and Percy Liang. Ecosystem-level Analysis of Deployed Machine Learning Reveals Homogeneous Outcomes. NeurIPS 2023 (arXiv:2307.05862). 2023.
  9. Chicago Sun-Times. Syndicated content in Sun-Times special section included AI-generated misinformation. Chicago Sun-Times. 2025.
  10. Chicago Sun-Times. Special section with fake book list plagued with additional errors, Sun-Times review finds. Chicago Sun-Times. 2025.
  11. Melissa Bell. Lessons (and an apology) from the Sun-Times CEO on that AI-generated book list. Chicago Sun-Times. 2025.
  12. Ziwei Ji et al. Survey of Hallucination in Natural Language Generation. ACM Computing Surveys 55(12). 2023.

Further reading

Sources last verified 2026-10-08.