Deep Learning — The Executive Mental Model
Deep learning is machine learning built from many-layered neural networks that learn their own features from raw data. That one idea made photographs, sound and text usable by machines. It does not make depth, size or a test score a guarantee, and it does not make deep learning the right tool for every problem.
After this chapter you can
- Define deep learning and place it inside machine learning without treating it as a synonym for AI.
- Explain representation learning and what the layers of a deep network learn.
- Explain why more layers or parameters do not automatically make a better model.
- Describe backpropagation at executive level and separate explaining a model from testing it.
- Judge when rules or classic machine learning fit a problem better than deep learning.
In December 2015, four researchers at Microsoft Research posted a paper that opened with an awkward result. They had built two neural networks to recognize objects in photographs, identical except that one had 18 layers and the other 34. The deeper network was worse. On the ImageNet benchmark it got 28.5 percent of images wrong, against 27.9 percent for the shallower one. The authors stressed that this was not overfitting, the familiar failure of a model that memorizes its training examples. The deeper network was worse even on the examples it had been trained on1.
The same paper then showed the fix. With one change to how the layers were wired together, the 34-layer network’s error fell to 25.0 percent, and residual networks up to 152 layers deep went on to win that year’s ImageNet competition1.
Both halves of that result belong in an executive’s head. Depth is what gives deep learning its power. Depth on its own guarantees nothing. A useful mental model of deep learning has to hold both thoughts at once, and it has to be precise enough to let you question a proposal without becoming an engineer.
The core idea
As AI vs Machine Learning showed, deep learning is not another word for AI. It is one family inside machine learning, which is itself one approach inside AI. What sets the family apart is captured in the definition three of its pioneers gave in Nature: models composed of many processing layers that learn representations of data at multiple levels of abstraction2.
Four words carry that definition. A neural network is a large arrangement of simple units, organized in layers. Each unit takes numbers from the layer below, weighs them, and passes a result upward. The weights are the model’s parameters, and a modern network has millions or billions of them. The previous chapter, How Machine Learning Learns, showed how training nudges parameters to reduce error. What is new here is everything between input and output: layer after layer of intermediate descriptions, or representations, that the network shapes for itself from the data.
The network learns the features
The phrase “learns the features” is the heart of the chapter, so it is worth making concrete.
Think of a table of data: the route, the carrier, the season. Each of those columns is a feature, and a person chose every one. Classic machine learning then learns how much each feature matters. On tables of data like this, it works well, as a later section shows.
Now try the same approach with a photograph, a recorded voice or a paragraph of free text. Which features would you write down for every way a dog can appear in a photograph, or every way an annoyed customer can describe a late delivery? You cannot list them. For decades this was the wall that kept images, sound and language out of reach of ordinary software. Each new task needed careful engineering and deep domain expertise to design the feature extractor by hand2.
Deep learning removed much of that wall. Given raw pixels, sound or text and enough labeled examples, the network learns its own representations, the internal descriptions that make the final judgment possible. That is why unstructured data, the photographs, recordings, scanned forms and documents that make up much of what organizations hold, became usable by machines.
The trade carries a cost that leaders should name early. Features you did not choose are features you cannot easily read. A classic model can tell you that port congestion mattered most. A deep network’s representations are patterns spread across millions of numbers, and nobody wrote them down.
What makes a network deep
A network is “deep” when it has many layers between input and output. There is no official number at which it qualifies. What depth buys is the ability to build descriptions on top of descriptions.
The clearest evidence comes from a 2014 study that made the inside of an image network visible. Matthew Zeiler and Rob Fergus traced what made individual units in each layer respond. The second layer responded to corners and combinations of edges and colors. The third captured textures, such as mesh patterns. The fourth was more specific to the kinds of object in the data, responding to dog faces or birds’ legs. The fifth responded to whole objects, such as keyboards and dogs, across many poses3.
Two cautions keep this picture honest. First, it is a tendency, not a promise that every layer has a name a person would recognize; many internal patterns have no tidy human label. Second, the same idea applies well beyond photographs. Speech networks build up from sound patterns toward words. Language networks build up from words toward meaning in context, and at very large scale that path leads to the models covered in Large Language Models, later in this module.
Deeper is not automatically better
Return to the result that opened the chapter. If each layer adds a level of description, why did adding 16 layers make the network worse?
The answer is that very deep networks are harder to train. The process that adjusts the parameters struggles to find good settings for so many stacked layers, so the extra layers end up doing harm rather than adding detail. The Microsoft team’s fix, called residual connections, gave each block of layers a shortcut: pass the input forward unchanged and learn only a correction to it. If a layer had nothing useful to add, it could leave the signal alone. With that change, the deeper network beat the shallower one, as depth was supposed to1.
The lesson for executives is simple. Depth and size are capacity. Design, data and training decide whether that capacity turns into performance. A proposal that leads with layer counts or parameter counts is describing how big a model is, not whether it fits your problem.
How a deep network learns
Training a deep network uses the same loop that How Machine Learning Learns described: make a prediction, measure the error, adjust the parameters, repeat. Depth adds one hard question. When the network gets an answer wrong, which of its millions of parameters, some of them dozens of layers away from the output, deserve the blame?
The answer is backpropagation. Starting from the error at the output, it works backward through the layers and calculates, for every parameter, how much a small change would have reduced the error. Every parameter is then nudged in the helpful direction, and the loop runs again, millions of times. The method was popularized in a 1986 Nature paper by David Rumelhart, Geoffrey Hinton and Ronald Williams, whose title makes the point of this chapter: the network was learning representations by propagating errors backward4.
If the method dates from 1986, why did deep learning only break through in the 2010s? Because it needs far more data and computing power than existed then. As Why AI, Why Now? showed, the turning point came in 2012, when a deep network trained on 1.2 million labeled photographs, using two graphics cards built for video games, won the ImageNet competition by a wide margin5. How data, model size and compute trade off against each other is the subject of Data, Models and Compute.
Two consequences of this way of learning matter to anyone who has to rely on the result. First, what the network knows is spread across all of its parameters, not stored as rules anyone can read. That makes it hard to explain fully. Hard to explain is not the same as impossible to test, though. You can still measure accuracy, error rates across groups and behavior under difficult conditions. The NIST AI Risk Management Framework treats “valid and reliable” and “explainable and interpretable” as separate characteristics of trustworthy AI for exactly this reason6. Second, “learning” here means parameters changing to fit data. The word is borrowed from people; the process is not how people learn, and a proposal that says a model “learns like our best staff” is selling a metaphor.
When deep learning is the wrong tool
Deep learning is the leading approach for images, audio and language. It does not lead everywhere, and the evidence on this point is unusually clear.
For fixed policies, plain rules win. “Any expense above a set amount needs a second approval” needs no training data, costs almost nothing to run and can be explained in one sentence. For tables of data, classic machine learning remains hard to beat. A 2022 benchmark of 45 tabular datasets of around 10,000 rows found that tree-based methods, a classic family, were still state of the art over deep learning, even after extensive tuning of both7.
The frontier does move. In January 2025, Nature published a pretrained deep network for tables that outperformed earlier methods on datasets of up to 10,000 rows, beating a tuned ensemble of the strongest classic methods in 2.8 seconds against 4 hours8. The same paper notes that tree-based methods had dominated tabular data for twenty years. The lesson is not that one family wins. It is that the right choice is an empirical question: compare the proposed model with a simple baseline on your data, and choose the simplest approach that reliably solves the problem at a cost you accept.
Story: teaching a sprayer to tell weeds from crops
The weed-spraying system built by Blue River Technology and John Deere shows the whole arc unusually clearly, from the problem rules could not solve to a measured result in the field. Read it as a post-mortem of a success: what made it work, what it cost and what it promised compared with what it delivered.
The problem is old. Farmers usually spray herbicide across a whole field, although weeds cover only part of it. Simple optical sensors could already find a green plant on bare soil. That is a rule, “green means weed”, and it breaks the moment the crop comes up, because then everything is green. Telling a young weed from a young crop plant, from a moving machine, in changing light, is exactly the kind of judgment nobody can write down as features.
Blue River, founded in 2011, attacked it with deep learning. Its system photographed the ground ahead of the sprayer and used a neural network to separate weeds from crops plant by plant, then sprayed only the weeds9. The model was the visible part. The bulk of the work was data: by early 2018 the company had collected more than a million photographs of cotton plants, taken from every angle and in every lighting condition, and about 200 people had labeled them. Its 2017 trials ran in cotton fields in Arkansas, Missouri and Texas, at about 5 miles an hour, well short of the 12 to 15 the business needed10.
In September 2017, Deere bought the company for 305 million9. It then took its time. Deere said it wanted the machine to work “really well” for customers before going to market10. The commercial product arrived in March 2022 in limited numbers, for three crops only: corn, soybeans and cotton. One design choice is telling. The system identifies the crop and sprays every plant that is not crop, so it never needs a picture of every weed species that might appear11.
Then came the numbers. For the prototype, the company’s chief executive spoke of up to 90 percent savings in herbicide10. At launch, Deere claimed more than two-thirds11. After the 2024 season, with the system used on more than a million acres, Deere reported an average saving of 59 percent across corn, soybean and cotton fields, an estimated 8 million gallons of herbicide mix12. The three numbers do not measure the same thing. The first is a best case for a prototype, the second a launch claim, the third an average across a season of customers’ fields. All three are reported by the company, not by an independent study.
The project worked, and the lesson is about the numbers. The first was a hope and the last was a measurement. The number to plan on is the measured field average, not the prototype estimate: still large, and well below the first promise.
What this means for leaders
You do not need to know how to build a neural network to govern one. You need to recognize what deep learning is good at, what it costs and which claims about it are signals rather than noise.
Deep learning earns its place when the signal is too rich to describe with features people can list: images, sound, free text, sensor streams. For tables and fixed policies, ask why something simpler would not do, and expect the team to show you the comparison. When you see a model’s size used as an argument, treat it as a description, not evidence. Ask instead for performance on the conditions where errors cost most, measured on data the model has not seen. Plan for the work around the model, because in a deep learning project that is where most of the cost and time go: data collection, labeling, testing, the fallback when the model is unsure, and monitoring after launch. And discount early promises. The number that matters is the one measured in your own operations.
Check yourself
- Deep learning is another word for AI.
- Adding layers to a neural network always makes it more accurate.
- In a deep network, the people who build it choose the features the model uses.
- You can test a model even when you cannot fully explain how it reaches its answers.
- On typical tables of business data, deep learning reliably beats classic methods.
- In the weed-spraying case, the measured field saving was lower than the early estimates.
Reflection: where would features fail?
What comes next
Deep learning gives a model its architecture and its way of learning. A trained model is not yet useful, though: it has to be put to work, answering real requests one after another, at a cost that never stops. Learning and using are two different jobs, with different costs, risks and people involved. The next chapter, Training vs Inference, explains the difference.
References
- Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun. Deep Residual Learning for Image Recognition. IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016); arXiv:1512.03385, December 2015. 2016.
- Yann LeCun, Yoshua Bengio and Geoffrey Hinton. Deep learning. Nature 521, 436-444. 2015.
- Matthew D. Zeiler and Rob Fergus. Visualizing and Understanding Convolutional Networks. European Conference on Computer Vision (ECCV 2014); arXiv:1311.2901. 2014.
- David E. Rumelhart, Geoffrey E. Hinton and Ronald J. Williams. Learning representations by back-propagating errors. Nature 323, 533-536. 1986.
- Alex Krizhevsky, Ilya Sutskever and Geoffrey E. Hinton. ImageNet Classification with Deep Convolutional Neural Networks. Advances in Neural Information Processing Systems 25 (NIPS 2012). 2012.
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. NIST. 2023.
- Leo Grinsztajn, Edouard Oyallon and Gael Varoquaux. Why do tree-based models still outperform deep learning on typical tabular data?. NeurIPS 2022 Datasets and Benchmarks Track; arXiv:2207.08815. 2022.
- Noah Hollmann, Samuel Muller, ... and Frank Hutter. Accurate predictions on small data with a tabular foundation model. Nature 637, 319-326. 2025.
- TechCrunch. After scrapping Monsanto deal, Deere agrees to buy precision farming startup Blue River for 305M. TechCrunch. 2017.
- DTN/Progressive Farmer. Spray System Aims to do Search-and-Destroy Missions on Weeds. DTN/Progressive Farmer. 2018.
- John Deere. The Furrow, March 2022 - Tech@Work: See & Spray Ultimate. John Deere. 2022.
- John Deere. See & Spray customers see 59% average herbicide savings in 2024. John Deere news release, 18 September 2024. 2024.
Further reading
- Yann LeCun, Yoshua Bengio and Geoffrey Hinton. Deep learning. Nature 521, 436-444. 2015.
- Ian Goodfellow, Yoshua Bengio and Aaron Courville. Deep Learning. MIT Press. 2016.
Sources last verified 2026-10-08.