AI Academy · Book
Executives & Directors · Module 01 · Chapter 003

Why AI, Why Now?

The ideas behind AI are seventy years old; what is new is readiness. Four forces matured at roughly the same time: web-scale data, specialized chips and training at scale, general foundation models, and easy access. Because that readiness is shared by every competitor, the advantage now lies in applying AI, not in obtaining it.

≈ 15 min read

After this chapter you can

  • Explain why AI became an urgent business issue only now, after seventy years of research, without crediting a single invention or product.
  • Name the four forces that converged - web-scale data, specialized chips and training at scale, foundation models, and easy access.
  • Distinguish the rising cost of building frontier models from the falling cost of using them.
  • Explain why shared access moves advantage from obtaining AI to applying it.
  • Test whether your organization's reasons for waiting on AI still hold today.

In August 1955, four scientists asked for funding to bring ten researchers together for two months the following summer. Their proposal coined the term artificial intelligence, and it was confident about the timetable: a “significant advance” could be made if a carefully chosen group worked on the problems “together for a summer”1. The summer of 1956 came and went. So did the next six decades, with long stretches of disappointment between the bursts of progress.

Then the timetable inverted. A free conversational assistant released at the end of November 2022 reached an estimated 100 million monthly users by January 2023, two months after launch, which UBS analysts called the fastest growth of any consumer application to that date2. A field that had taken seventy years to deliver on its first promise reached a mass audience in eight weeks.

Both facts are true, and together they pose the question this chapter answers. If the idea is old, what changed? The honest answer is not one invention and not one product. The product was the moment the public noticed. The change had been building for years underneath it.

The world around AI became ready

AI did not change overnight. The world around it became ready for it. Four forces, each with its own history, matured at roughly the same time, and each made the others useful.

Four forces converged to make AI ready - web-scale data, specialized chips and training at scale, foundation models and easy access.Web-scale dataThe public web astraining materialChips and scaleSpecialized hardware fortraining at enormous scaleFoundation modelsOne general model,many tasksEasy accessCloud delivery andplain language
Figure 1.3.1 Four forces converged. None of them alone would have made AI a business issue for everyone.

The word for this is convergence. It is the same pattern the previous chapter described for every general-purpose technology: the breakthrough people remember sits on top of several slower developments that had to mature first. The four forces are distinct. Data is the raw material. Chips and training at scale are the engine. Foundation models are the product of that engine running on that material. Access is how the result reached ordinary organizations. The next sections take them one at a time.

Force one: data the size of the web

Learning systems learn from examples, and for most of AI’s history there were not enough of them in digital form. That changed with the public web. Since 2007 a nonprofit called Common Crawl has archived the open web and given the archive away. It now holds more than 300 billion pages and adds three to five billion each month3. Alongside it grew digitized books, open-source code, encyclopedias and billions of captioned images.

Two milestones show what this material made possible. ImageNet, a collection of more than a million photographs labeled by category, gave image recognition its proving ground in an annual competition4. In 2020 a language model called GPT-3 was trained on a mix in which filtered Common Crawl supplied 410 billion tokens, roughly words and word fragments, and 60 percent of the training material5. The datasets have kept growing: Stanford’s AI Index estimates that the data used to train large language models doubles about every eight months6.

General models learned from the public web and never saw your company's information, so they know the world's language but not your business.Learned fromPublic web pagesDigitized books and codeCaptioned imagesNever sawYour contractsYour customer historyHow your processes really runA general model knows the world's language, not your business.
Figure 1.3.2 Foundation models were trained on web-scale public data, not on company data.

One point here is often misunderstood, and it matters for every executive. These models were trained mostly on the public record and licensed material, not on your company’s data. They never saw your contracts, your customer history or how your processes actually run. That explains two things at once. It explains why the models were useful on the first day: nobody had to collect and label your data before you could try them. And it explains why your own information becomes important later, when a model is connected to it safely. How that connection works is the subject of RAG and Enterprise Knowledge in Module 02.

Force two: specialized chips and training at scale

For decades, AI was starved of computing power. The fix was not simply cheaper ordinary computers. It was hardware built for the kind of arithmetic that learning needs: millions of small calculations done in parallel.

The turning point is usually dated to 2012. A University of Toronto team trained a deep neural network on 1.2 million ImageNet photographs using two consumer graphics cards, chips designed for video games, in five to six days4. In that year’s ImageNet competition, their system’s top-5 error was 16.4 percent on the competition data alone, and 15.3 percent with extra training data; the next best entry managed 26.2 percent7. A gap that size ended an argument. Graphics chips became the standard tool for training, and companies soon designed chips for AI alone…

Researchers then found that capability improved predictably as models, data and computing power grew together, which gave the field a reason to scale up8. It did.

Training compute for machine learning doubled about every 20 months before 2010 and about every 6 months in the deep-learning era.Before 201020 monthsDeep-learning era6 monthsSource: Sevilla et al., IJCNN · 2022
Figure 1.3.3 Once specialized chips arrived, the computing power used to train leading models doubled about every six months instead of every 20.

Before 2010, the computing power used to train leading models doubled about every 20 months, roughly in line with ordinary chip progress. In the deep-learning era it doubled about every six months9, and the AI Index now puts the doubling time for notable models at about five months6. This is the second force: not cheap computing, but specialized chips that made training at enormous scale possible.

Costlier to build, cheaper to use

This force has a consequence that is often stated backward. People say AI “got cheap”. Using it did. Building it at the frontier did the opposite.

The chips themselves improved steadily. The AI Index estimates that the cost of machine-learning hardware has fallen about 30 percent a year while its energy efficiency has risen about 40 percent a year. But the scale of training grew much faster than the chips improved. Epoch AI estimates that the hardware and energy cost of the largest training runs has grown about 2.4 times a year since 2016, and that on that trend the largest runs would cost more than a billion dollars by 202710. Building at the frontier has become the business of a few very large organizations: nearly 90 percent of notable models in 2024 came from industry rather than universities6, and over 90 percent of notable frontier models in 202511.

Training compute doubles every five months and frontier costs rise, while using a model at equal performance became more than 280 times cheaper in under two years.5 monthsTraining computedoublesNotable models; frontier costskeep rising30%Yearly fall inhardware costChips improve, but scalegrows faster>280×Cheaper to useSame performance, Nov 2022 toOct 2024Source: Stanford AI Index · 2025
Figure 1.3.4 Building at the frontier got costlier; using a model got far cheaper. Most organizations rent the result.

These are three different measures: the price of the chips, the total bill for one frontier training run, and the price a user pays per query. Using a model went the other way. At the performance level of the first widely used chat models, the price of a query fell from 20 dollars per million tokens in November 2022 to 7 cents in October 2024, a drop of more than 280-fold6. For an executive, the split settles the first question that boards often ask. Your organization almost never needs to build its own frontier model. It rents the result of someone else’s very expensive training run, and that rent keeps falling. Where the full cost of using AI really lies is a separate question, taken up in The Economics of AI at the end of this module.

Force three: one general model for many tasks

For decades, organizations built narrow systems, one per task: a model to flag fraud, a model to translate, a model to rank search results. That work was real and valuable. But each new task meant new labeled data, a new team and a new project.

Training at scale on web-scale data produced something different. A 2017 design called the transformer made it practical to train language models on far larger amounts of text12. By 2020, GPT-3, with 175 billion parameters, could translate, answer questions and complete text from a handful of examples, without being retrained for each task5. In 2021 Stanford researchers gave the category its name: foundation models, trained once on broad data at scale and then adapted to a wide range of tasks13.

One general model, trained once, now serves many tasks that used to need a narrow model each.Onegeneral modelDraftSummarizeTranslateClassifyWrite codeAnswer
Figure 1.3.5 One general model, trained once, now serves tasks that used to need a narrow model each.

This is the force most people mean when they say that AI “got good”. It changes the economics of a first experiment: you no longer start from zero, you start from a foundation and adapt it. It does not make a model a strategy, and the brand names at the top of today’s rankings will keep changing. How these models work, and how they are adapted, is the subject of Foundation Models in Module 02.

Force four: access as a service and a conversation

The fourth force is access, and it arrived in two waves.

For most of AI’s history, using it meant having your own research team and your own specialized hardware. The first wave came from cloud computing. In August 2006 computing capacity became something a company could rent by the hour as a web service14. Over the following decade AI capabilities joined it: a model built and run by someone else became a service that any development team could call, paying only for what it used.

Access widened from research labs to developers through the cloud and then to everyone through plain language; with shared access, advantage moves to application.Research labsSpecialists with theirown hardwareDevelopersA cloud service any teamcan callEveryonePlain language,no programmingSame access for all. Advantage moves to application.
Figure 1.3.6 Access widened from research labs to developers to everyone. When access is shared, advantage moves to application.

The second wave was the conversational interface of late 2022. Plain language became the way to use advanced AI, so people could use it without programming or a data-science team, and they did so in their tens of millions within weeks2. Why language as an interface changes so much is the subject of Generative AI Changes the Game; why many employees started before their employers had a policy is the subject of AI Is Already Inside Your Organization.

Access also changed competition. A startup and a large enterprise can now reach broadly similar capability on the same day. The differentiator moves from who can obtain AI to who can apply it: to their own data, processes, people and controls.

Why not ten years ago?

The forces matured at different speeds, and that is why the moment came when it did.

The four forces matured at different speeds, and all four were in place only from about 2022, when adoption took off.EraDataChipsModelsAccess2000sGrowingGeneral-purposeNarrowSpecialistsEarly 2010sPlentifulGraphics chipsNarrowSpecialistsLate 2010sWeb-scaleAI chipsFirst general modelsDevelopersSince 2022Web-scaleAI chips at scaleFoundation modelsEveryone
Figure 1.3.7 A qualitative summary: all four forces were in place only from about 2022, and that is when adoption took off.

By the 2000s the web was filling with data, but chips were general-purpose, models were narrow and AI was for specialists. In the early 2010s graphics chips were being repurposed for learning, yet models were still one per task. By the late 2010s specialized chips, the transformer, the first general language models and cloud AI services were all in place, but almost nobody outside technology used them directly. Only from about 2022 did all four line up.

It is tempting to ask which force mattered most. There is no useful winner. Data without chips is an unread library. Chips without data are idle power. Both without a general model produce narrow tools. A general model without easy access remains a laboratory curiosity. Remove any one and today’s adoption becomes much harder.

The forces are also still moving, and not only upward. Researchers project that training datasets will reach the size of the entire stock of public human-written text sometime between 2026 and 203215. Almost every leading AI chip is fabricated by a single foundry11, and the power needed for training grows every year6. Any of these could slow a force. None is likely to reverse the convergence, but they are a reason to build capabilities that survive a change of tool rather than betting on one product.

Story: one shipping line, two verdicts

The container shipping line in this story is a composite, built from patterns common across the industry. It is an illustration of the argument, not evidence for it.

Around 2016, an international container line had a familiar problem. After every port strike or storm closure, its customer team was buried in emails from shippers, in more than a dozen languages: where is my cargo, and why am I being charged for returning containers late? Leaders asked whether AI could help. A small specialist team spent months labeling old emails by hand so that a model could learn to sort them. The model did one task in one language. It ran on servers the line had bought, and only the specialists could change it. When customers changed their wording, accuracy fell, and every new language or task meant a new project. A year later the steering committee gave its verdict: AI is not ready for us.

In 2016 the shipping line built one narrow model and concluded AI was not ready; in 2024 the same problem met a ready world and produced a year of evidence.2016One narrow model, onelanguage, months of labelingVerdict: not ready2024One general model, manylanguages, approvedcloud toolA year of evidence on real workvs
Figure 1.3.8 An illustrative composite: same shipping line, same problem, different world. The 2016 verdict was right for 2016.

In 2016 they were largely right. The forces had not yet converged.

In 2024 the problem was back, with a new twist: service staff were pasting customer emails into public AI assistants on their own phones to draft replies. That was a data risk, and also a signal. Leaders reopened the old verdict and tested it against the four forces. A general model could already read and draft in many languages with no labeling. Someone else had paid for the training. And the capability came as a cloud service inside the line’s own data boundary, used in plain language.

They chose one process, replies to disruption inquiries. Only the approved tool was allowed, and no customer data went into public assistants. The AI drafted; staff edited and sent every reply; a person still approved every waived charge or other commercial concession; and quality was sampled every week. A year later the line had not bought a revolution. It had a year of evidence about where the drafts helped, which languages needed closer review and which answers about charges and contract terms needed checking. The shipping line had not become smarter between the two attempts. The world around AI had become ready, and this time leadership tested its old verdict instead of repeating it.

What this means for leaders

The four forces lead to one practical conclusion. AI became ready for everyone at once. That makes it a present leadership issue, not a future research topic, and it means that waiting buys less than it used to, because competitors with the same access are already learning.

It also means that many of the reasons organizations give for waiting belong to an earlier world. “We would need to label years of data first” described narrow models. “We would need our own infrastructure” described the era before cloud AI. “Only specialists can use it” described the era before plain-language interfaces. Some objections remain valid, but they have moved from the global forces to local conditions: whether a model can reach the right information safely, whether people have an approved tool, and whether someone owns the process being changed.

Readiness is also a statement about capability, not results. It says the tools can now do this kind of work; whether they pay off in your organization is a separate question, and AI Is Changing Everything showed how rarely they yet do. Readiness is not a mandate for unfocused spending. It is a reason to start learning on one real problem, with people accountable for the output, and to judge progress by evidence. How organizations typically move from that first problem to wider change is the subject of the next chapter.

Check yourself

  1. AI appeared overnight with the launch of one consumer product.
  2. Today’s general models learned mostly from company data.
  3. Building frontier models got more expensive while using models got cheaper.
  4. The key compute breakthrough was cheaper ordinary computers.
  5. Shared access to AI moves the source of advantage from obtaining AI to applying it.
  6. Because the world is ready, every company should spend heavily on AI now.

Reflection: reopen an old verdict

What comes next

The world is ready, and it is ready for everyone at once. Yet organizations are responding at very different speeds, and some move from experiment to real change much faster than others. The next chapter, The AI Adoption Curve, describes that journey and how to tell where your organization stands on it.

References

  1. John McCarthy, Marvin L. Minsky, Nathaniel Rochester and Claude E. Shannon. A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence, August 31, 1955. AI Magazine 27(4), 2006 reprint. 1955.
  2. Krystal Hu. ChatGPT sets record for fastest-growing user base - analyst note. Reuters. 2023.
  3. Common Crawl Foundation. Common Crawl - free and open corpus of web crawl data. commoncrawl.org. 2026.
  4. Alex Krizhevsky, Ilya Sutskever and Geoffrey E. Hinton. ImageNet Classification with Deep Convolutional Neural Networks. Advances in Neural Information Processing Systems 25 (NIPS 2012). 2012.
  5. Tom B. Brown et al. Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems 33 (NeurIPS 2020). 2020.
  6. Stanford Institute for Human-Centered AI (HAI). AI Index Report 2025, Chapter 1: Research and Development. Stanford University. 2025.
  7. ImageNet (Stanford Vision Lab, Princeton and UNC). ImageNet Large Scale Visual Recognition Challenge 2012 (ILSVRC2012) results. image-net.org. 2012.
  8. Jared Kaplan et al. Scaling Laws for Neural Language Models. arXiv:2001.08361. 2020.
  9. Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu, Marius Hobbhahn and Pablo Villalobos. Compute Trends Across Three Eras of Machine Learning. 2022 International Joint Conference on Neural Networks (IJCNN). 2022.
  10. Ben Cottier, Robi Rahman, Loredana Fattorini, Nestor Maslej and David Owen. How Much Does It Cost to Train Frontier AI Models?. Epoch AI. 2024.
  11. Stanford Institute for Human-Centered AI (HAI). AI Index Report 2026. Stanford University. 2026.
  12. Ashish Vaswani et al. Attention Is All You Need. Advances in Neural Information Processing Systems 30 (NeurIPS 2017). 2017.
  13. Rishi Bommasani et al. On the Opportunities and Risks of Foundation Models. Stanford CRFM (arXiv:2108.07258). 2021.
  14. Amazon Web Services. Announcing Amazon Elastic Compute Cloud (Amazon EC2) - beta. Amazon Web Services. 2006.
  15. Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim and Marius Hobbhahn. Will We Run Out of Data? Limits of LLM Scaling Based on Human-Generated Data. Proceedings of the 41st International Conference on Machine Learning (ICML 2024). 2024.

Further reading

Sources last verified 2026-10-09.