AI Academy · Book
Executives & Directors · Module 02 · Chapter 005

How Machine Learning Learns

A machine-learning model learns by guessing on examples with known answers, measuring its error and adjusting, thousands of times over. What it learns is patterns, not answers, and those patterns are only worth something if they hold on cases the model has never seen and keep holding after the world moves. Capability comes from the loop around the model, and leaders own both ends of that loop.

≈ 16 min read

After this chapter you can

  • Describe the machine-learning loop from problem definition to monitoring and retraining.
  • Explain how a model learns by measuring and reducing its error on examples with known answers.
  • Distinguish generalization from memorization, and overfitting from underfitting.
  • Recognize leakage and shortcuts that make a held-back test easier than real use.
  • Explain data drift and concept drift, and why a strong test score can decay in production.
  • Separate the model from the business decision and from the wider system.

In February 2009 a team led by Google researchers published a paper in Nature describing a system that estimated flu activity across the United States from what people typed into Google1. Official figures from the Centers for Disease Control and Prevention (CDC), assembled from doctors’ and laboratories’ reports, typically arrived one to two weeks late. Google Flu Trends offered the same picture almost at once, and it was built on more data than any epidemiologist had ever worked with.

Four years later, in February 2013, Nature reported that the system was predicting more than double the share of doctor visits for influenza-like illness that the CDC recorded. A subsequent analysis in Science found that it had been too high in 100 of the 108 weeks from August 2011 to September 2013. Worse, a simple projection from CDC data that was already three weeks old estimated current flu levels better than Google’s model did2. In August 2015 Google stopped publishing the estimates and handed the data to research partners instead3.

That is the contradiction at the heart of machine learning. A model can be built by capable people, on enormous data, for a well-defined target, and still fail at the one thing it was built for. Nothing in Google’s code had broken. The failure came from how the model had learned and from what happened to the world after it learned. Those two mechanisms explain most of the machine-learning disappointments an executive will ever see, and both can be managed.

Data in, intelligence out is the wrong picture

It is tempting to picture machine learning as a pipe: pour in data and intelligence comes out. AI vs Machine Learning defined machine learning more carefully, as a system whose performance on a task, by some measure, improves with experience. The question here is how that improvement actually happens, and what it depends on.

The answer is a loop, not a pipe. Someone defines the problem and what counts as a better answer. Someone collects examples that show the answer. The model is trained on those examples, then tested on cases it has not seen. If it passes, it is deployed to predict new cases inside real work. Then it is monitored, and retrained or replaced when its predictions start to drift away from reality, which sends the loop around again.

Machine-learning capability comes from a loop of define, collect, train, test, deploy and monitor, not from a model alone.DefineProblem andobjectiveCollectExamples withanswersTrainLearn patternsTestOn unseen casesDeployPredict new casesMonitorWatch and retrainThe learning loopNever finished
Figure 2.5.2 The model is the output of this loop, never the whole capability. Leaders own the first and last stages.

The model is the product of the loop at one moment in time. Data scientists own the middle of it. The two stages that often decide success, defining what to learn and noticing when it stops working, belong to the business.

Define what the model must learn

A model cannot learn from an ambition. “Reduce churn with AI” is an ambition. A learning problem names an input, a target and a decision: each month, predict which subscribers are likely to cancel within ninety days, using their usage, complaints and tenure, so the retention team calls the right people first. Only the second version can be trained, tested and held to account.

The target is a choice, and the model will chase exactly what is chosen. Ask a model to optimize customer satisfaction and someone must decide whether that means a survey score, repeat purchases, complaint rates or retention. Each choice trains a different model with different blind spots. Google Flu Trends was asked to match the CDC’s curve, and it did what it was asked: it found search terms whose rise and fall matched that curve, whether or not they had anything to do with flu2. Machine-learning systems optimize what we define, not necessarily what we intended. Closing that gap is leadership work.

The examples matter as much as the target. Each training example carries a label, the recorded answer the model learns from: did this customer leave, was this invoice paid late, did this booked car turn up. Wrong labels teach the wrong lesson, and the model learns it efficiently. Bias and Fairness, later in the program, examines a well-known case in which a sensible-looking label quietly encoded the wrong goal, and Data, Models and Compute, later in this module, covers what makes data fit for the purpose.

How training works: guess, measure, adjust

Training sounds mysterious, but the principle is simple enough to describe in four steps. A model begins as a large set of adjustable numbers, called parameters, that turn inputs into an output. At first those numbers are close to arbitrary, and so are its predictions.

A model learns by guessing on examples with known answers, comparing, scoring the error and adjusting its parameters, repeated many times.GuessPredict a casewith aknown answerCompareCheck whatreally happenedMeasureScore the errorAdjustNudge theparametersRepeat across every example, many times over
Figure 2.5.3 Learning is error reduction. The model changes only because a measure says how wrong it was.

The model makes a guess for an example whose answer is already known. The guess is compared with what actually happened. A loss function, the measure of wrongness, turns the gap into a number: a subscriber who left but was given a 5 percent chance of leaving produces a large loss. The training procedure then works out which way to nudge each parameter so that the loss would have been a little smaller, and nudges it. Repeat across every example, often for many passes through the data, and the parameters settle into values that make good predictions on the training examples4.

Three consequences follow, and each matters to a leader. First, what comes out is a learned function, customer data in and probability of leaving out, not a rule anyone wrote. For complex models it may not be a rule anyone can read. Second, learning requires a measurable definition of “better”. If the loss function rewards the wrong thing, training will faithfully deliver the wrong thing. Third, the procedure has no idea which patterns are real. It reduces error on the examples it is given, by whatever patterns are available, including coincidences. How the adjustment is computed inside a deep network is the subject of the next chapter, and the difference between this learning phase and the phase in which the model is used has a chapter of its own, Training vs Inference.

The goal is generalization, not memory

Pedro Domingos, in a widely read essay on the subject, put the central point in four words: it is generalization that counts5. A model is valuable only if it performs well on cases it did not see during training. Doing well on the training examples is easy. A big enough model can simply memorize them.

A model can miss in two directions. It can underfit: be too simple to capture the real pattern, and so do poorly everywhere. Or it can overfit: fit its training examples so closely that it learns their noise and coincidences along with their signal. An overfitted model looks brilliant on the past and fails on anything new.

The goal sits between underfitting, which misses the pattern, and overfitting, which memorizes the examples - a model that generalizes to cases it never saw.UnderfittingToo simple - misses the real patternOverfittingToo specific - memorizes the examplesGeneralizesWorks on cases it never saw
Figure 2.5.4 The goal is not the most complex model, but the one that holds up on new cases.

Google Flu Trends is a textbook case of overfitting. Its first version screened 50 million search terms to find the best match for 1,152 weekly data points. With that many candidates and that few cases, some terms were bound to match the flu season by coincidence. The developers themselves had to weed out seasonal terms such as those about high school basketball. The authors of the Science analysis described the result as “part flu detector, part winter detector”, and it missed the 2009 H1N1 pandemic entirely, because that outbreak did not arrive in winter2.

A picture helps. One traveler prepares for a trip by memorizing a phrasebook: a few hundred perfect sentences. It works beautifully until a waiter replies with a sentence that is not in the book. Another traveler learns from many conversations, is corrected and adjusts, until she can understand sentences she has never heard. The first has memorized; the second has generalized. A good model is fluency in the patterns, not a phrasebook of past answers.

The discipline that protects against the phrasebook is a strict separation of data. Teams split their examples into three sets. The training set is what the model learns from. The validation set is used to compare designs and tune settings during development. The test set is locked away and used once, at the end, to estimate how the model will do on cases it has never seen. The rule is simple and often broken: never judge a model on the data it learned from.

When the test is not really unseen

A held-back test set is necessary, but it is not enough. The test cases must be unseen in the way that matters, which means they must not share hidden shortcuts with the training data.

Medical imaging offers a well-documented example. In 2018 researchers trained deep-learning models to detect pneumonia on chest X-rays from three hospital systems. The models scored well on held-back images from the hospitals they had trained on and markedly worse at a third hospital, because they had partly learned to recognize where an image was taken, and pneumonia was far more common in one hospital’s images than in the others’. The test set was held back, but it shared the shortcut, so it could not reveal the flaw6.

When information that would not be available in real use slips into training or testing, data scientists call it leakage. It is common: Sayash Kapoor and Arvind Narayanan found it in at least 294 published studies across 17 scientific fields7.

Production is not the test score

Even a model that generalizes well on launch day is generalizing from the past. When the world changes, its learned patterns can expire. This is drift, and the research literature distinguishes two kinds8.

Data drift means the inputs change; concept drift means the relationship between inputs and outcome changes; both appear only when predictions are compared with reality.DATA DRIFTThe inputs change: newcustomers, new products,new channels.Patterns meet unfamiliar casesCONCEPT DRIFTThe link between inputs andoutcome changes: the samesignal now meanssomething else.Learned patterns become wrongvs
Figure 2.5.5 Both kinds are normal. Neither shows up in the code, only in predictions compared with what really happened.

In data drift, the inputs change. A retailer opens a new channel and the model starts seeing customers unlike any it learned from. In concept drift, the relationship itself changes, so the same signal now means something else. Google Flu Trends suffered from both. The authors of the Science analysis pointed to what they called algorithm dynamics: Google kept changing its search engine, and people kept changing how they searched, so the link between searches and flu shifted under a model that was not watching for it2. AI vs Automation described a sharper version of the same thing, when the pandemic broke models across inventory management and fraud detection in a matter of weeks.

Drift is not a defect to be eliminated. It is a condition to be managed. The practical tools are unglamorous: compare predictions with actual outcomes on a schedule, watch for inputs that look unlike the training data, name an owner who receives those reports, and agree in advance what triggers a retrain. A test score describes the world the model learned from. It is not a permanent business result.

The model is neither the decision nor the system

A model produces a prediction. The organization decides what to do with it. Suppose the churn model gives a subscriber a 70 percent chance of leaving, an illustrative figure. Should she be offered a discount? That depends on her value, the cost of the offer, the experience the company wants to give and its appetite for risk. Turning a score into an action is a decision rule, and the business must own it.

Nor is the model the system. In a widely cited paper, engineers at Google showed that in a real-world machine-learning system the learning code is only a small fraction of the whole. Around it sit data collection and verification, configuration, serving infrastructure and monitoring, and that surrounding machinery is where much of the cost and much of the failure lives9.

The model and its score are the visible tip; data pipelines, monitoring, decision rules, human review, feedback design and ownership sit below and decide whether it works.WHAT THE DEMO SHOWSThe model · Its test scoreWHAT THE SYSTEM NEEDSData pipelinesMonitoringDecision rulesHuman reviewFeedback designOwnership
Figure 2.5.6 Evaluate the system, not only the model. Most of the work and most of the risk sit below the waterline.

The same paper warns about hidden feedback loops: systems whose own outputs shape the data they will later learn from9. A recommendation engine that learns only from clicks on what it chose to show keeps showing people more of the same, and next year’s training data is partly its own creation. Feedback can improve a system, but only if someone designs it.

Story: a day in ferry revenue management

Suppose a ferry operator that carries cars across a busy strait. What follows is an illustrative composite, not a single documented case, but every mechanism in it is standard revenue-management practice. On busy sailings the operator sells a few more car spaces than the deck holds, because some booked vehicles never turn up, and it sets those overbooking limits from forecasts of no-shows. Underestimate the no-shows and deck space sails empty; overestimate them and booked cars are left on the quay. The same forecasting problem is well documented in airline revenue management, where passenger-level models beat simple historical averages10.

At six in the morning, the operator’s revenue manager opens the overnight forecast. For every sailing in the next three days, the no-show model has predicted how many booked vehicles will not appear. It learned from two years of bookings, tested well on months it had never seen and has been in service for a year.

At half past eight she sets the limits. The model gives a number; she makes the decision, because the two errors cost different amounts. An empty lane on the deck is lost revenue. A booked car left on the quay is a refund, a missed appointment on the far side and a customer who may not come back. On the busiest crossings she sells slightly fewer extra spaces than the model suggests.

At eleven a message arrives from the port: the early sailing was oversold again, and booked cars had to wait for the next crossing. It is the third time this week.

Over one day a ferry revenue manager receives the no-show forecast, sets limits, hears of a third oversold sailing, finds the drift and puts containment and an owner in place.06:00The forecastarrivesNo-shows predictedfor every sailing08:30She setsthe limitsThe model advises;she decides11:00Oversold againThird time this week14:00Predictedmeets actualThe lines split inweek five18:00Contain andrepairManual limits, retrainplan, named owner
Figure 2.5.7 An illustrative composite day. The model ran perfectly; the problem was the world it was running in.

At two in the afternoon she sits down with the data team. Nothing is broken: the pipeline ran, the model ran, nobody has changed the code. Then someone puts predicted and actual no-shows on one chart, week by week; the chart below uses illustrative figures. For four weeks the lines move together. In week five they split.

Predicted and actual no-show rates move together until week five, when actual no-shows fall while the model keeps predicting the old rate.0%2%5%8%10%W1W2W3W4W5W6W7W8Predicted no-shows· 6%Actual no-shows· 2%ILLUSTRATIVE NUMBERS
Figure 2.5.8 Illustrative data. Week five is when a new fare launched; the model kept predicting last year’s travelers.

Week five was when the operator launched a new low fare that cannot be changed or refunded. Drivers who lose their money if they miss a sailing have every reason to turn up. The fare did not exist in the training data, so the model had never learned what its buyers do. It went on predicting no-shows that no longer happened, and the limits it informed sold space to cars that all arrived. That is concept drift, and the operator caused it itself: nobody in pricing knew that launching a fare also changed a model.

By six in the evening the team has done three things. They lower the limits by hand on the affected crossings. They plan a retrain once there are enough weeks of the new fare to learn from, tested on the most recent weeks rather than a random sample. And they add a daily comparison of predicted and actual no-shows, with a named owner and a rule that any new fare or schedule change triggers a model review. The original training was sound. What was missing was the last stage of the loop.

What this means for leaders

Machine learning rewards leaders who look past the model to the loop. Four habits make the difference.

Write the learning problem down. Insist on a named input, target, decision and purpose before anyone trains anything, and check that the target is what the business actually wants rather than a convenient proxy.

Ask what the test resembled. A score is evidence only about the data it was measured on. Ask whether the test cases came from a later period, a different site or a new segment, and whether anything in them could have leaked from training.

Separate the score from the decision. Ask who turns predictions into actions, what each kind of error costs and who can override the model.

Fund the last stage. Budget for monitoring and retraining from the start, give each production model a named owner, and treat your own launches, price changes and policy changes as events that can move a model.

Check yourself

  1. A model that is right on every training example will do well on new cases.
  2. A bigger dataset protects a model from learning coincidences.
  3. A model learns by measuring its error on examples with known answers and adjusting to reduce it.
  4. If a model was tested on held-back data, its score will hold in a new setting.
  5. A model can degrade in production even though nobody has changed its code.
  6. The model’s output is the business decision.

Reflection: find the week-five moment

What comes next

The model has so far been a closed box: parameters that are nudged until the error falls. Much of modern machine learning fills that box with something specific, many-layered neural networks that learn their own features from raw data. The next chapter, Deep Learning — The Executive Mental Model, opens the box at executive altitude.

References

  1. Jeremy Ginsberg et al. Detecting influenza epidemics using search engine query data. Nature 457, 1012-1014. 2009.
  2. David Lazer, Ryan Kennedy, Gary King and Alessandro Vespignani. The Parable of Google Flu: Traps in Big Data Analysis. Science 343(6176). 2014.
  3. Google Research. The Next Chapter for Flu Trends. Google Research Blog. 2015.
  4. Ian Goodfellow, Yoshua Bengio and Aaron Courville. Deep Learning. MIT Press. 2016.
  5. Pedro Domingos. A Few Useful Things to Know About Machine Learning. Communications of the ACM 55(10). 2012.
  6. John R. Zech, Marcus A. Badgeley, Manway Liu, Anthony B. Costa, Joseph J. Titano and Eric K. Oermann. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLOS Medicine 15(11), e1002683. 2018.
  7. Sayash Kapoor and Arvind Narayanan. Leakage and the reproducibility crisis in machine-learning-based science. Patterns 4(9), 100804. 2023.
  8. João Gama, Indrė Žliobaitė, Albert Bifet, Mykola Pechenizkiy and Abdelhamid Bouchachia. A Survey on Concept Drift Adaptation. ACM Computing Surveys 46(4). 2014.
  9. D. Sculley et al. Hidden Technical Debt in Machine Learning Systems. Advances in Neural Information Processing Systems 28 (NIPS 2015). 2015.
  10. Richard D. Lawrence, Se June Hong and Jacques Cherrier. Passenger-based predictive modeling of airline no-show rates. KDD '03: Proceedings of the Ninth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 397-406. 2003.

Further reading

Sources last verified 2026-10-08.