AI Academy · Book
Executives & Directors · Module 02 · Chapter 003

AI vs Machine Learning

Artificial intelligence is a field defined by its goal; machine learning is one way of reaching it, in which a system improves at a task from experience instead of being told every rule. Knowing which kind of learning sits behind a proposal, and where its answer key comes from, tells a leader more than the label AI ever will.

≈ 14 min read

After this chapter you can

  • Place AI, machine learning, deep learning and generative AI in one hierarchy, and judge "is it machine learning?" component by component.
  • Test any machine-learning proposal by its task, its experience and its measure.
  • Distinguish supervised, unsupervised, self-supervised and reinforcement learning by where each gets its answer key.
  • Explain why reinforcement learning is central to building today's AI assistants and reasoning models.
  • Choose an approach by naming the output the problem needs before naming a technique.

In May 1997, IBM’s Deep Blue beat the world chess champion Garry Kasparov over six games. Twenty years later, DeepMind’s AlphaZero taught itself chess in about nine hours and then beat Stockfish, the strongest chess program of its day. Both are milestones of artificial intelligence. Before reading on, predict the answer to a simpler question: which of them is machine learning?

Most people say both. The answer is one. Deep Blue’s chess knowledge was written by people. Its team, advised by grandmasters, built an evaluation function of more than 8,000 features that told the machine what to look for in a position, and created its opening book of about 4,000 positions by hand. Software tools helped tune some of the weights, but what mattered in chess was decided by engineers, and the machine’s strength came from brute search: on average, 126 million positions a second during the 1997 match1. AlphaZero was given nothing but the rules. It played against itself, kept what led to wins, and first outplayed Stockfish after four hours of training. In a 1,000-game match it won 155 games and lost 6. It examined about 60,000 positions a second, against Stockfish’s 60 million, because it had learned which positions were worth examining2.

Deep Blue's chess knowledge was written by engineers and grandmasters and it searched 126 million positions a second; AlphaZero learned from the rules and self-play and searched 60,000; both are AI but only AlphaZero is machine learning.SystemChess knowledge came fromPositions per secondMachine learning?Deep Blue (1997)Engineers and grandmasters126 millionAlphaZero (2017)The rules and self-play60,000
Figure 2.3.1 Both are AI. Only one learned its chess, and it needed to look at far fewer positions.

The story has a coda. Stockfish, the program AlphaZero beat, is open source, and its developers kept improving it. In 2020 they added a small neural network, trained on the evaluations of millions of positions, to judge positions; it tested at more than 80 Elo points stronger than the hand-written judgment it supplemented3. In 2024 they removed the hand-written evaluation altogether4. Its search, the part that explores moves, is still conventional code. So one of the strongest chess engines in the world today is partly machine learning and partly not. That is the first lesson of this chapter: “is it machine learning?” is a question about the parts of a system, and the answer is often “this part, not that one”.

The field, and one way of building it

Artificial intelligence names a goal. When the term was coined in 1955, as Why AI, Why Now? recounted, it described an ambition to make machines do what we would call intelligent, without prescribing any method5. Machine learning names a method. The term was popularized by Arthur Samuel at IBM, whose 1959 paper described a checkers program that improved by playing, and argued that teaching computers to learn from experience would one day spare programmers much of the detailed work of writing every instruction6.

The two words therefore sit at different levels. AI is the field. Machine learning is one major approach inside it, alongside approaches that encode knowledge by hand, such as expert rules, search and planning. Inside machine learning sits deep learning, the family of methods built on many-layered neural networks that powers most recent progress. The standard textbook on deep learning draws exactly this nesting: deep learning inside machine learning inside AI, with hand-built knowledge bases as AI that learns nothing from data7.

AI is the broad field; machine learning is one approach inside it, deep learning is a family inside machine learning, and generative AI is a capability mostly built with deep learning.ArtificialintelligenceThe goal, by any methodFieldMachinelearningImproves from experienceApproachDeeplearningMany-layered neural networksFamilyGenerative AICreates content; mostly deep learningCapability
Figure 2.3.2 AI is the field. Machine learning is one way to build it, and most generative AI is built with deep learning.

Generative AI fits this picture slightly differently. It is a description of what a system produces, new text, images or code, rather than of how it learned; today’s generative systems are almost all built with deep learning. The Major Types of AI, the next chapter, explains why labels that describe outputs and labels that describe methods should not be mixed. For now, two sentences do most of the work. Not every AI system learns from data: Deep Blue did not. And not every machine-learning system generates anything: most of them predict.

What makes it machine learning

The most useful definition for a leader is also one of the oldest. Tom Mitchell’s 1997 textbook says that a program learns when its performance at some task, as measured in some way, improves with experience. His worked example was Samuel’s game: the task is playing checkers, the measure is the share of games won, and the experience is games played against itself8.

The definition has three parts, and each becomes a question that any machine-learning proposal should answer in plain language. In the checkers example the task is playing checkers, the experience is games played against itself, and the measure is the share of games won. A business proposal needs the same three sentences: to flag deliveries likely to run more than a day late, for instance, the task is the flag at dispatch, the experience is past delivery records with actual arrival times, and the measure is late deliveries caught set against the false warnings someone must handle.

Machine learning means improving at a task, as judged by a measure, from experience, so every proposal must name its task, its data and its measure.TaskWhat exactly will it do?ExperienceWhat data will it learn from?MeasureHow will we know it improved?
Figure 2.3.3 If a team cannot name all three in a sentence each, it does not yet have a machine-learning project.

A team that cannot fill in all three does not yet have a machine-learning project; it has a hope. A vague task (“use AI on deliveries”) cannot be measured. Missing experience, such as no record of actual arrival times, means nothing to learn from. And the measure has to be the business’s measure, not just a technical score: a flag that cries wolf twenty times a day will be ignored, however clever it is. How the learning itself proceeds, and how it can go wrong, is the subject of How Machine Learning Learns.

Four ways a machine can learn

Most textbooks name three broad ways of learning. A fourth now matters as much as the other three, because it is how the largest models start. The difference between them comes down to one question: where does the answer key come from?

Supervised learning gets its answers from people's labels, unsupervised learning finds structure with no answers, self-supervised learning takes its answers from the data itself, and reinforcement learning follows a designed reward; each needs a different check.Way of learningThe answer key comes fromBusiness exampleAskSupervisedPeople who labeled past casesWhich loans defaultedWho labels, and how late?UnsupervisedNowhere: it finds structureCustomer segmentsDoes the pattern mean anything?Self-supervisedThe data itselfPredicting the next wordWhat was in the data?ReinforcementA reward someone designedSelf-play; human ratingsWhat exactly is rewarded?
Figure 2.3.4 Each way of learning gets its answer key from somewhere different, and each needs a different check.

In supervised learning, the system learns from examples paired with known answers: past loans and whether each one defaulted, past deliveries and whether each one was late. Most business machine learning works this way. Its answer key is made by people or by events, which means it costs money, arrives with a delay and carries the judgment of whoever made it. In unsupervised learning there is no answer key. The system finds structure on its own, such as clusters of similar customers or transactions unlike any others. That is useful for discovering what nobody thought to look for, but someone still has to decide whether a cluster means anything.

Self-supervised learning takes its answer key from the data itself. Hide the next word of a sentence, and the sentence already contains the right answer. Because no person has to label anything, a model can learn from a vast amount of text, which is how foundation models are first trained9. The price is that the model absorbs whatever the data contained, good and bad. In reinforcement learning, the system learns by acting: it tries something, receives a reward or a penalty, and adjusts to earn more reward over time10. AlphaZero learned chess this way, with winning as its reward. Real systems mix these freely, as the next section shows.

Reinforcement learning built today’s assistants

Reinforcement learning is often described as a technique for games, robots and control systems. That picture is out of date. Reinforcement learning is now central to how general-purpose AI assistants are made.

A modern AI assistant is first pretrained self-supervised on text, then tuned on example answers, then trained by reinforcement on people's rankings and on answers that can be checked.PretrainingSelf-supervisedon vast textExamplesSupervised onexample answersRankingsReinforcement onpeople'spreferencesChecksReinforcement onverified answers
Figure 2.3.5 A modern assistant is built in stages, and three different ways of learning take a turn.

A language model’s first stage is self-supervised: it learns to predict text. Such a model is fluent but not reliably helpful. In the method published by Ouyang and colleagues in 2022, people then wrote example answers for supervised tuning, and ranked alternative answers from the model; a reward model learned those preferences, and reinforcement learning trained the language model to earn them. People preferred the answers of the resulting 1.3-billion-parameter model to those of the untuned 175-billion-parameter base model of the same family, more than 100 times its size11. The newest step trains reasoning models with rewards that a program can check: a mathematics answer is compared with the correct result, and code is run against test cases. Work published in Nature in 2025 showed this kind of reinforcement learning lifting a model’s score on hard mathematics problems severalfold, without any human-written examples of reasoning12.

The executive lesson is not the mechanics, which Large Language Models covers. It is that a system trained by reward becomes good at whatever is rewarded. A model rewarded for answers people prefer learns to produce answers people prefer, which is not always the same as answers that are true. A model rewarded for checkable answers improves fastest where answers can be checked. So when a proposal says a model was “trained with reinforcement learning”, the right follow-up is a single question: rewarded for what, judged by whom?

Name the output before the technique

Most machine learning in business does not generate anything. It predicts: a credit score, a demand forecast, a label such as “likely duplicate”, a ranking of products to recommend. Organizations ran such systems for years before generative AI arrived, and they still carry much of the value. As Generative AI Changes the Game set out, generative systems add a different kind of output, new content, and with it different checks. Calling both “AI” hides the difference that matters most when choosing between them: what the system is supposed to produce.

That suggests the habit this chapter recommends. Before anyone names a technique, name the output the problem needs.

A fixed answer from known logic points to rules, a score or forecast to supervised learning, unnamed patterns to unsupervised learning, new drafts to generative AI, and consequential judgments to a person.What outputdoes theproblem need?Answer from known logicRulesA score, forecast or labelSupervised learningPatterns nobody has namedUnsupervised learningA new draft or summaryGenerative AIA consequential judgmentA person decides
Figure 2.3.6 Name the output first. The technique follows from it, and real systems usually combine several branches.

If the answer follows from known logic, such as a tax formula or an approval limit, write the rule; AI vs Automation showed why rules often win. If the problem needs a score, forecast or label, and history records the outcomes, supervised learning fits. If the problem is to find patterns nobody has named, unsupervised learning fits, with a person judging what they mean. If the problem needs a draft, a summary or an answer in plain language, generative AI fits, with someone checking the output. And where a decision has serious consequences for a person, a person makes it, whatever sits underneath. Teams that start from a technique tend to go looking for a problem that fits it. Teams that start from the output usually end up with several techniques, each in its place.

Story: the reviewers Facebook sent home

In March 2020, as the pandemic closed offices, Facebook, which renamed itself Meta the following year, sent its content reviewers home to protect their health and said it would rely more heavily on its technology to review content13. The decision turned one of the largest enforcement systems in the world into an unplanned test of this chapter’s question: which parts learn, and where does their answer key come from?

The system had several parts, and they were not all the same kind of AI. Written Community Standards, drafted and revised by people, define what is not allowed; those are rules. Matching technology acts on a new post that matches or comes very close to content already judged to break a rule. Learned classifiers flag likely violations before anyone reports them. And review teams make the final call on cases the technology sends them. In Meta’s own description, the technology “can learn from each human decision” and becomes more accurate after thousands of them14. The reviewers were a layer of judgment, and they were also the answer key.

Facebook's enforcement combined written policies, which are rules; matching against content already judged; learned classifiers trained on reviewers' decisions; and review teams who make the final call.PartKindWhere its answer comes fromWritten policiesRulesPeople write and revise themMatchingComparison, not learningContent people already judgedClassifiersMachine learningReviewers' past decisionsReview teamsPeoplePolicy and judgment
Figure 2.3.7 One enforcement system, four kinds of part. Two of them get their answers from the reviewers.

The report for the next quarter, published in August 2020, showed both halves. Where the learned parts were strong, they carried the load. Facebook had extended its automated hate-speech detection to Spanish, Arabic and Indonesian and improved it in English in the first quarter, then improved it again in English, Spanish and Burmese in the second. Its proactive detection rate, the share of actioned hate speech it found before anyone reported it, rose from 89 to 95 percent. The volume it acted on, a different measure, rose from 9.6 million pieces in the first quarter to 22.5 million in the second13. Where the work rested on people, it fell. Action on suicide and self-injury content on Facebook fell from about 1.7 million pieces to 911,000, as reported from the company’s data15. “With fewer content reviewers,” the company wrote, it acted on less of that content on Facebook and Instagram, and on less child sexual exploitation content on Instagram. It relies heavily on people to review those categories and to help improve the technology that finds copies of violating content, and it could not always offer appeals13. Facebook said it had prioritized the most harmful content within those categories, and that many reviewers were soon working again from home16.

Hate speech that Facebook acted on rose from 9.6 million to 22.5 million pieces between the first and second quarters of 2020, while action on suicide and self-injury content fell from about 1.7 million to 911,000.Hate speech, Q1 20209.6 millionHate speech, Q2 202022.5 millionSuicide and self-injury,Q1 20201.7 millionSuicide and self-injury,Q2 20200.9 millionSource: Meta, Community Standards Enforcement Report · August 2020
Figure 2.3.8 With fewer reviewers, the learned detector carried more of the load; the work that rested on people fell by almost half.

Nothing in the record says the technology failed. The lesson is about structure. The parts that had learned well kept working, and in one category improved. The parts that depended on people’s verdicts, either to decide cases or to supply the examples the technology learns from, shrank with the people. A leader asked to fund a model by cutting the people who judge cases should ask the same question: does the model learn from them? If it does, those people are not overhead around the system. They are where its experience comes from.

What this means for leaders

The word AI tells you almost nothing about what is inside a system. The words in this chapter tell you a great deal. Whether a component follows written rules or learned from experience decides how it fails and how you check it. Which way it learned decides where its answer key came from, which in turn decides its costs, its delays and its blind spots. And the output it produces, a rule’s verdict, a prediction, new content or a recommendation for a person, decides whether it fits the problem at all. Leaders do not need to build these systems, but they do need to ask about them in this vocabulary, and to expect real systems to combine several layers.

Check yourself

  1. Machine learning is another name for AI.
  2. Today’s strongest chess engines combine hand-written search with a learned evaluation.
  3. Most machine learning in business generates new content.
  4. Reinforcement learning is mainly a technique for games and robots.
  5. A large language model is first trained without anyone labeling its data.
  6. A model trained by reward becomes good at what the people who built it intended.

Reflection: find the answer key

What comes next

Machine learning is one major approach inside AI, sitting alongside rules, search and planning, with deep learning inside it and generative AI built on top. But AI can be sorted in other ways too: by what a system does, how it is built, how general it is and how autonomous it is. The next chapter, The Major Types of AI, maps that wider landscape and shows how to keep the different labels from blurring together.

References

  1. Murray Campbell, A. Joseph Hoane Jr. and Feng-hsiung Hsu. Deep Blue. Artificial Intelligence 134(1-2): 57-83. 2002.
  2. David Silver et al. A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play. Science 362(6419): 1140-1144. 2018.
  3. The Stockfish developers. Introducing NNUE Evaluation. Stockfish blog. 2020.
  4. The Stockfish developers. Stockfish 16.1. Stockfish blog. 2024.
  5. John McCarthy, Marvin L. Minsky, Nathaniel Rochester and Claude E. Shannon. A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence, August 31, 1955. AI Magazine 27(4), 2006 reprint. 1955.
  6. Arthur L. Samuel. Some Studies in Machine Learning Using the Game of Checkers. IBM Journal of Research and Development 3(3): 210-229. 1959.
  7. Ian Goodfellow, Yoshua Bengio and Aaron Courville. Deep Learning. MIT Press. 2016.
  8. Tom M. Mitchell. Machine Learning. McGraw-Hill. 1997.
  9. Rishi Bommasani et al. On the Opportunities and Risks of Foundation Models. Stanford CRFM (arXiv:2108.07258). 2021.
  10. Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction, second edition. MIT Press. 2018.
  11. Long Ouyang et al. Training language models to follow instructions with human feedback. NeurIPS 2022 (arXiv:2203.02155). 2022.
  12. Daya Guo et al. (DeepSeek-AI). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645: 633-638. 2025.
  13. Guy Rosen. Community Standards Enforcement Report, August 2020. Facebook Newsroom. 2020.
  14. Meta. How enforcement technology works. Meta Transparency Center. 2024.
  15. Tom McKay. Facebook Says Covid-19 Shutdowns Hurt Its Ability to Fight Suicide, Self-Injury, Child Exploitation Content. Gizmodo. 2020.
  16. Associated Press. Facebook: Pandemic hurt enforcement on suicide, child nudity. Associated Press (via WJXT News4JAX). 2020.

Further reading

Sources last verified 2026-10-10.