Data, Models and Compute
Every AI system is built from three things: the data it learns from and reads, the model that turns that data into behavior, and the compute that builds and runs the model. They work like factors in a product, so the weakest of the three sets the ceiling. Leaders who ask which corner is weakest make better investments than leaders who ask which model is biggest.
After this chapter you can
- Explain how data, models and compute combine to produce AI capability, and why the weakest one limits the rest.
- Distinguish training data from runtime data, and explain why data quality often matters more than volume.
- Explain why the right model is the smallest one that clears the quality bar, and how distillation and quantization help.
- Recognize compute as an operating capability that builds the model and then serves every request.
- Use the three building blocks to find the weakest point in an AI proposal before approving it.
In June 2023, a Microsoft research team published a coding model called phi-1. By the standards of the day it was tiny: 1.3 billion parameters, trained for four days on eight graphics processors on a 7-billion-token dataset (a token is a word fragment), about 51 billion tokens seen in training because the model read the dataset roughly eight times. On HumanEval, a widely used test of writing short Python programs, it solved 50.6 percent of the problems on the first try. StarCoder, a model with 15.5 billion parameters trained on about a trillion tokens, solved 33.6 percent1.
The researchers did not have a secret architecture. They had different data. Instead of feeding the model everything they could collect, they filtered public code for material of “textbook quality” and added about a billion tokens of textbooks and exercises written for the purpose. Two caveats keep the result honest. HumanEval measures one narrow skill, and those synthetic textbooks were written by a much larger model, so some of phi-1’s quality was bought with someone else’s compute. But that is precisely the lesson. Capability did not come from the model alone, or the data alone, or the compute alone. It came from how the three were combined.
Three building blocks, one capability
Any AI system, from a demand forecast to a chat assistant, can be examined with three questions. What does it learn from, and what does it see when it works? That is the data. How does it represent and learn patterns? That is the model. How much processing can be applied, first to build the model and then to run it? That is the compute.
The three do not add up; they behave more like factors in a product. A powerful model with poor data gives poor answers. Excellent data with an unsuitable model goes to waste. A strong model and good data without affordable compute fail economically, because the system costs too much to build or to run. A useful mental model, not a literal equation, is that capability is roughly data times model times compute: a near-zero in any one of them drags the whole result toward zero.
The research behind modern AI shows the same interdependence. As Why AI, Why Now? described, researchers found in 2020 that language models improved smoothly and predictably as model size, data and compute grew together. The same study carried a condition that matters more to a leader than the curve itself: grow the model while the data stays fixed, and the gains run into diminishing returns2. The three have to grow in balance.
Data enters a system twice
Data is the raw material: text, images, audio, sensor readings, transactions, documents, source code and business records. Different problems need different data, and it enters an AI system at two different moments, which map onto the two phases described in Training vs Inference.
Training data is used before go-live, to build or adapt the model. Years of past orders train a demand forecast; a large body of text trains a language model. Training data decides which patterns the model is able to learn, and which it never sees. Runtime data is supplied each time the system is used: a customer’s question, the current state of their account, this week’s price list. It gives the model information that did not exist when it was trained.
The distinction has practical consequences. Prices, stock levels and policies change daily, and retraining a model every time they change is rarely sensible. Well-designed applications look up current facts at the moment of use instead, an approach RAG and Enterprise Knowledge explains later in this module. When a team proposes an AI system, it is worth asking about both kinds of data separately, because they are usually owned, governed and refreshed by different people.
Quality decides more than volume
A common belief about data is that the organization with the most of it wins. Volume is what gets counted in a pitch: rows, terabytes, years of history. What decides the result is whether the data is accurate, relevant to the decision, complete, consistent, current and representative of the people and situations the system will meet.
Practitioners know this, and still find it hard to act on. In interviews with 53 AI practitioners working on high-stakes systems in India, East and West Africa and the United States, Google researchers found that 92 percent had experienced at least one data cascade: a data problem, often small and unnoticed at the start, that compounded into failures further down the line. Their title summed up the cause: everyone wants to do the model work, not the data work3.
Even carefully curated data contains errors. When researchers audited the test sets of ten of the most widely used benchmark datasets in machine learning, they estimated that at least 3.3 percent of labels were wrong on average. The authors showed that errors at that level can be enough to change which of two models looks better4. If the data used to judge a model is wrong, the choice of model can be wrong too.
Representativeness deserves particular attention because training data carries history. If past decisions were uneven, the records are uneven, and a model learns those patterns whether anyone intends it to or not. Before data is used, someone should be able to answer four questions: who is represented, who is missing, what assumptions shaped how the records were created, and whether the data fits the use now intended. How bias turns into harm, and how to govern it, is covered later in the program; the point here is that it begins as a property of the data.
The model: fit beats size
The model is the structure that learns patterns from data. In deep learning it is a neural network, designed for a job such as prediction, vision or language. The model decides what the data becomes, which is why two systems trained on the same data can give different results. Its architecture shapes what relationships it can represent, how efficiently it learns and how much compute it needs. For generative AI, one architecture, the transformer, has been especially important5; Large Language Models explains how it works.
Model size is usually measured in parameters, the internal settings adjusted during training. More parameters can mean more capability, and they nearly always mean more cost and slower answers. But as phi-1 showed, size is only one input. The target is not the largest model available. It is the smallest model that clears the quality bar for the task, tested on the task.
Two engineering techniques make that target reachable. Distillation uses a large model to train a smaller one to imitate it. DistilBERT, an early and well-documented example, was 40 percent smaller and 60 percent faster than the model it learned from while keeping 97 percent of its language-understanding performance6. Quantization stores a model’s numbers at lower precision. Researchers showed in 2022 that 8-bit quantization could halve the memory needed to run language models of up to 175 billion parameters without measurable loss in quality7.
Neither technique is free. A distilled or quantized model may lose abilities the task turns out to need, so the trade has to be measured against the quality bar, not assumed.
Compute is a capability, not a pile of chips
Compute is the processing capacity used to train and to run models. Training, as the two previous chapters showed, repeats one loop an enormous number of times, and every pass costs computation. Graphics processors became central because neural networks perform vast numbers of similar calculations that can run in parallel, which is exactly what those chips were designed to do8.
But compute is not a purchase order for chips. The chips have to be fed by memory, storage and networking; software has to schedule which job runs where; and serving systems have to answer requests as they arrive.
Compute at scale is an operation, not an asset in a rack: the team that trained the OPT model in 2022 logged at least 35 manual restarts from hardware failures across about two months on 992 GPUs9.
That operational burden is one reason the most advanced models are built by a small number of organizations. And compute does not stop when training stops. Every request a model serves consumes compute again, so the running bill grows with use, as Training vs Inference showed. Training scale determines how a model is built; inference scale determines how it is delivered.
Why one weak corner sinks the rest
Put the three back together and the product logic becomes concrete. The table below is an illustrative summary, not a measured result: when one corner is strong and another weak, the strong corner cannot rescue the result, and each mismatch fails in its own recognizable way.
Scale is not free, so progress also comes from efficiency: better data, better algorithms, better hardware and models sized to the job. phi-1 and DistilBERT are both examples of buying capability with quality and design rather than with more of everything.
Two related questions sit just beyond this chapter. Total cost of ownership, cloud versus on-premises and the economics of energy belong to Module 08, AI Economics, starting with Model, Compute and Infrastructure Costs. Whether your data can become a lasting competitive advantage belongs to Proprietary Data, AI Moats and Differentiation. Both start from the same triangle.
Story: New Orleans stops waiting to be asked
On 11 November 2014, five people, three of them children, died in a house fire in the Broadmoor neighborhood of New Orleans. There was no working smoke alarm in the house. Between 2010 and 2014 the city had recorded 22 deaths in structure fires10. The fire department had long given out free smoke alarms, but residents had to come to their local fire station to get one11. The alarms went to people who asked, and the people most at risk were often not the ones asking.
Before, the department’s data described demand, not need. A list of requests says who knows about the program and has time to act on it. It says nothing about which homes lack an alarm.
After, the city’s Office of Performance and Accountability worked with the fire department to answer a different question: which neighborhoods are least likely to have alarms and most likely to suffer a fatal fire? The hard part was the data. A national housing survey asked households directly whether they had a smoke alarm, but its results were only available for the whole parish. The census had detail for small neighborhood areas, called block groups, but never asked about alarms. The analysts used questions that appeared in both: how old the building was, how long the household had lived there, and household income relative to the poverty line. They fitted a logistic regression, one of the simplest statistical models, on the survey, applied it to census block groups, and combined the result with a fatality-risk score built from past fires and the share of residents under five or over 6510.
The city’s own estimate was that canvassing the tenth of the city ranked highest would find more than 20 percent of the homes needing alarms, twice what a random tenth would find. It flagged its own limitation too: because the data was by census block group rather than by house, real targeting would be less precise than the estimate10. Firefighters then went door to door. By mid-2016 the department had installed more than 8,000 alarms with the tool, and in October 2015 alarms in one targeted home sounded in time for all 11 residents to escape11. Those are outputs and one dramatic escape, not a measured outcome. Neither source reports a test of whether the targeting really doubled reach or reduced fire deaths, so read the doubling as the city’s forecast.
Look at the triangle. The data corner took the real work: finding a source that asked the right question and a second source at the right level of detail, and joining them. The model corner was deliberately modest. The compute corner was a set of scripts the city later published12; as the city put it, nothing it did “required big data or fancy machines”11. Each corner was sized to one decision: which doors to knock on first.
What this means for leaders
The triangle turns a vague question, “should we invest in this AI?”, into three specific ones. First, interrogate the data before the model. Many proposals describe the model in detail and the data in a sentence, yet the data cascades study suggests problems often start in the data. Ask what the system learns from, what it sees at runtime, who owns each, and whether anyone has measured its quality.
Second, treat model size as a cost to be justified. The default should be the smallest model that clears a quality bar the business has set, with evidence that smaller options were tested. Third, see compute as an operation that continues for as long as the system runs, not a one-off purchase. A pilot that serves fifty users hides the bill for fifty thousand. And finally, size all three to the decision. New Orleans did not need a frontier model, because its decision was which doors to knock on, made a few times a year.
Check yourself
- More data always makes an AI system better.
- Two models trained on the same data can give different results.
- The largest model available is usually the safest choice.
- Once a model is trained, its compute costs largely stop.
- Data problems in AI projects are rare when experienced teams are involved.
- New Orleans needed only a simple statistical model to target smoke alarms well.
Reflection: find your weakest corner
What comes next
Data, models and compute are the building blocks of every AI system. The next big shift came when a few organizations combined very broad data with very large amounts of compute to build models that many others could reuse, so that most organizations no longer had to build every model themselves. The next chapter, Foundation Models, explains what those models are and why they changed the economics of AI.
References
- Suriya Gunasekar, Yi Zhang, Jyoti Aneja et al. Textbooks Are All You Need. arXiv:2306.11644 (Microsoft Research). 2023.
- Jared Kaplan et al. Scaling Laws for Neural Language Models. arXiv:2001.08361. 2020.
- Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh and Lora Aroyo. "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (ACM). 2021.
- Curtis G. Northcutt, Anish Athalye and Jonas Mueller. Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks. NeurIPS 2021 Datasets and Benchmarks Track; arXiv:2103.14749. 2021.
- Ashish Vaswani et al. Attention Is All You Need. Advances in Neural Information Processing Systems 30 (NeurIPS 2017). 2017.
- Victor Sanh, Lysandre Debut, Julien Chaumond and Thomas Wolf. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. arXiv:1910.01108 (NeurIPS 2019 EMC2 Workshop). 2019.
- Tim Dettmers, Mike Lewis, Younes Belkada and Luke Zettlemoyer. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale. NeurIPS 2022; arXiv:2208.07339. 2022.
- Ian Goodfellow, Yoshua Bengio and Aaron Courville. Deep Learning. MIT Press. 2016.
- Susan Zhang, Stephen Roller, Naman Goyal et al. OPT: Open Pre-trained Transformer Language Models. arXiv:2205.01068 (Meta AI). 2022.
- City of New Orleans, Office of Performance and Accountability. Analytics-Informed Smoke Alarm Outreach Program (NOLAlytics report). City of New Orleans. 2015.
- Katherine Hillenbrand. New Orleans Develops Data-Intensive Predictive Fire Risk Model. Government Technology. 2016.
- City of New Orleans, Office of Performance and Accountability. smoke-alarm-outreach (source code). GitHub. 2015.
Further reading
- Ian Goodfellow, Yoshua Bengio and Aaron Courville. Deep Learning. MIT Press. 2016.
- Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh and Lora Aroyo. "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (ACM). 2021.
Sources last verified 2026-10-08.