AI Academy · Book
Executives & Directors · Module 10 · Chapter 002

The Evolution of AI Models

Leading AI capability is no longer scarce. Rival frontier models sit within a few percent of each other, freely downloadable models trail them by months, and small models now do what giant ones did two years ago. The strategic question has moved from choosing the best model to running a portfolio that is tested on your own work and built to be swapped.

≈ 15 min read

After this chapter you can

  • Explain the three engines of model improvement - scale, post-training and thinking time - and why size alone does not decide quality.
  • Describe how the frontier has converged and how capability moves into smaller, cheaper models, with evidence.
  • Distinguish what a public benchmark shows from what an organization's own evaluation set shows.
  • Design a model portfolio with routing, and judge when one model is enough.
  • Apply the rule that a model release triggers a test, not a migration, and recognize when adapting a model changes legal role.

Before reading on, make a guess. The most capable AI models are sold by a handful of companies as closed services: you can use them, but you cannot download them. Alongside them sit open-weight models, whose trained parameters anyone can download, run and adapt. How far behind the best closed model is the best open one? Three years? One year? A few months?

Many people guess a year or more. Epoch AI, which tracks model capabilities, found that from January to May 2026 the most capable open-weight models lagged the frontier closed models by an average of four months1. The 2026 AI Index adds a second view. On the Arena leaderboard, where people compare answers from two anonymous models and pick the better one, the top closed model led the top open model by 3.3 percent in early 2026, and the leading models from four different companies sat within 25 rating points of each other2. The two measures differ. Epoch’s is a lag in time, averaged across many benchmarks; the Arena’s is a gap in ratings that people give when they compare two answers.

In early 2026 the best closed model led the best open model by 3.3 percent, and the top four companies' models were within 25 rating points.3.3%Closed lead over openBest closed versus bestopen model25Rating pointsSpread of the top fourcompanies' best models2.7%US lead over ChinaTop model from eachSource: Stanford AI Index 2026 (Arena leaderboard) · March 2026
Figure 10.2.1 At the top, capability is crowded. The lead any one model holds is small, and it changes hands.

That answer changes the shape of every model decision. If the best model in January is matched by several rivals and an open alternative by May, then a strategy built on picking the winner is a strategy built on something that will not hold still.

Capability is abundant; design for the move

Here is the argument. Models have improved through three successive engines: more scale, better training after scale, and more thinking time at the moment of use. Each engine raised the frontier, and each was copied, compressed and cheapened within months. The result is that strong capability is now plentiful and short-lived at the top.

For an enterprise this turns model choice from a one-time purchase into an ongoing discipline. Organizations that handle this well tend not to bet on one model. They match each workload to the least expensive model that clears its bar, keep a set of their own real cases to test any new model against, and build their systems so that changing the model is a configuration change rather than a rebuild. The rest of this chapter explains why each of those habits follows from how the models have evolved.

Three engines of improvement

The modern story starts in 2017, when Google researchers introduced the transformer, the design that lets a model weigh every part of its input against every other part and that underlies today’s language models3. Large Language Models explains how it works; what matters here is what happened next.

Model capability grew through scale from 2020, post-training from 2022 and thinking time from 2024, with each advance quickly copied, including in open models.2017ThetransformerThe design undertoday's models2020Engine 1 - scaleBigger models anddata improvepredictably2022Engine 2- post-trainingA small tuned modelbeats a giantraw one2024Engine 3 -thinking timeReasoning modelswork beforeanswering2025Copied inthe openOpen-weightreasoning model
Figure 10.2.2 Each engine raised the frontier, and each was matched by others within months. Size alone was never the whole story.

Scale. In 2020, researchers showed that a language model’s errors fall smoothly and predictably as the model, its training data and its computing budget grow4. That finding turned model building into an investment race: the computing used to train notable models has doubled roughly every five months5. But scale was never only size. In 2022, DeepMind trained a 70-billion-parameter model on four times more data than its 280-billion-parameter predecessor, at the same computing cost, and the smaller model won across the board6. A model with more parameters is not automatically a better model.

Post-training. The second engine is what happens after the initial training. By showing a model examples of good answers and having people rate its outputs, developers taught it to follow instructions. The headline result: people preferred the answers of a 1.3-billion-parameter tuned model to those of the 175-billion-parameter original, a model more than 100 times larger7. Usefulness, not raw size, became the target.

Thinking time. The third engine arrived in September 2024 with the first widely available reasoning model, which spends extra computation working through a problem before it answers8. Four months later a developer published the weights of a reasoning model trained largely by reinforcement learning that performed comparably on reasoning tasks9. Training vs Inference explains why this engine moves cost from building the model to every single use of it.

Each engine changed which model is best for which job. Each was also matched quickly by competitors, which is the subject of the next section.

The frontier converges

Two years ago, the leading model held a visible lead. That lead has collapsed.

The rating gap between the top and tenth-ranked model fell from 11.9 to 5.4 percent, and between the top two from 4.9 to 0.7 percent.Top vs 10th model,early 202411.9%Top vs 10th model,early 20255.4%Top two models, 20234.9%Top two models, 20240.7%Source: Stanford AI Index 2025 (Chatbot Arena) · early 2025
Figure 10.2.3 The gap between the best and the tenth-best model halved in a year. The gap between the top two almost vanished.

On the Arena leaderboard, the rating gap between the top and the tenth-ranked model fell from 11.9 percent to 5.4 percent between early 2024 and early 2025, and the gap between the top two fell from 4.9 percent to 0.7 percent10. Convergence is not a straight line, though. The 2026 AI Index found that the closed lead over open models had widened again, to 3.3 percent from 0.5 percent in August 2024, and that six of the top ten models were closed2. Leads still open up. They just do not last long.

For an executive, the practical reading is twofold. First, a contract or architecture that locks you to one model locks you to a lead measured in months. Second, the choice between closed and open models is no longer a choice between capability and control. It is a trade-off among operating effort, data control, cost and support, which Model and Third-Party Risk and AI Platform Strategy examine in depth.

Capability moves downward

The frontier does not only converge. What it can do also migrates, quickly, into much smaller and cheaper models.

The smallest model to score above 60 percent on the MMLU knowledge test shrank from 540 billion parameters in 2022 to 3.8 billion in 2024.Smallest model above 60%on MMLU, 2022540 BSmallest model above 60%on MMLU, 20243.8 BSource: Stanford AI Index 2025 (MMLU) · 2024
Figure 10.2.4 The same test score that needed 540 billion parameters in 2022 needed 3.8 billion in 2024, a 142-fold reduction.

MMLU is a broad test of knowledge across 57 subjects. In 2022, the smallest model to score above 60 percent on it had 540 billion parameters. In 2024, a model with 3.8 billion parameters did the same, a 142-fold reduction in two years10. Much of this comes from distillation, in which a large model helps train a smaller one, and from better training data, both introduced in Data, Models and Compute. A model of that size is small enough to run on a laptop. These are test scores, a measure of capability; whether a small model handles a particular workload is something only a test on that workload shows.

The consequence is a moving boundary. A workload that needed a frontier model when you launched it may be handled by a small, fast and inexpensive model a year later, and the reverse holds for work you once ruled out. The question to ask each year is not “Is our model still good?” but “Which tier of model does this workload actually need now?”

Leaderboards saturate; your own cases do not

If capability moves this fast, how do you know which model is good enough? Public benchmarks are the usual answer, and they are a weak one.

Benchmarks are built to stay hard for years and are now often beaten within months. On SWE-bench, a test of fixing real software issues, the best score rose from 4.4 percent in 2023 to 71.7 percent in 202410. On Humanity’s Last Exam, designed as a very hard test of expert knowledge, frontier models gained 30 percentage points in a single year. The 2026 AI Index puts it bluntly: evaluations meant to stay challenging for years are saturated in months. It also cites a review that found invalid questions in widely used benchmarks at rates from 2 percent to 42 percent2.

Leaderboards rank models on public tests that age quickly; an evaluation set of your own cases shows whether a model clears your bar at an acceptable cost.A leaderboard showsRank on a public testA snapshot that ages in monthsQuestions that may be flawedYour evaluation set showsWhether a model clears your barCost per completed taskWhether a release helps you
Figure 10.2.5 A public benchmark ranks models. Only a set of your own real cases tells you which one your work needs.

The durable alternative is an evaluation set: a few hundred real, representative cases from your own work, with the right answer or an agreed standard for each, kept private and updated as the work changes. It is unglamorous, and it may be the most valuable model asset an organization can own, because it outlives every model it is used to judge. AI Evaluation and Approval Gates covers how evaluation feeds formal approval. Here the point is narrower: without your own cases, every model decision rests on someone else’s test.

From one model to a portfolio

Put the previous sections together and a different design emerges. Instead of one model for everything, the organization keeps a small portfolio: a frontier or reasoning model for the hardest work, a general model for most tasks, and small or specialized models for high-volume, routine or on-device work. A router, a thin layer of software in front of the models, decides which one handles each request.

Requests go to a router that sends simple work to a small model, typical work to a general model and hard cases to a frontier model, using rules set by the organization's own evaluation set.WORKRequestsMixed difficultyCONTROLRouterPicks the tierEvaluation setOur real casesMODELSSmall modelRoutine and high volumeGeneral modelMost tasksFrontier modelHardest casesevery requestsimpletypicalhardsets the rules
Figure 10.2.6 A portfolio separates the work from the model. The router and the evaluation set stay; the models behind them can change.

The research case for this design is promising, if early. Researchers at Stanford and Berkeley showed that a cascade, which tries an inexpensive model first and escalates only when the answer looks weak, matched the best single model on their test sets at up to 98 percent lower cost11. A Berkeley-led team trained routers that cut costs by more than half in some settings without lowering answer quality12. Those are laboratory results on public tasks, not promises for your workflow, and the cost mechanics belong to Cost Optimization and AI FinOps. But the direction is consistent: when most requests are easy, sending all of them to the most powerful model is the most expensive way to get the same result.

A portfolio has costs of its own. Each model added means more testing, more security review and more vendor management, and a simple organization with one dominant workload may rightly choose a single model. The aim is not the largest portfolio. It is the smallest one that fits the work, with the ability to change any part of it.

Models change under you

The last reason for a portfolio is that models do not stay put, even when their names do. Researchers at Stanford and Berkeley tested the same named hosted model in March and June 2023. Its accuracy at telling prime from composite numbers fell from 84 percent to 51 percent in three months13. Providers also retire model versions on their own schedules, so a model you have not chosen to replace will eventually be replaced for you.

A model release triggers a test on the organization's own cases; only workloads that improve are routed to the new model, and all models are monitored for drift.Release noticedNew modelor versionTest on our setQuality, cost, speedRoute or holdMove only theworkloads that gainMonitorWatch for driftand retirementAssessDo not auto-migrate
Figure 10.2.7 Every release starts a test, not a migration. Only the workloads that improve on your own cases move.

The right response is a standing routine. Every significant release triggers a test on your evaluation set, not a migration. Workloads move only where the new model is materially better or cheaper on your cases. Everything in use is monitored for drift. AI Lifecycle Governance covers reassessment and retirement, and Model and Third-Party Risk covers the exit plan for a critical provider. What this chapter adds is the reason those routines are no longer optional: the pace of model evolution guarantees that the model under any workflow will change during that workflow’s life.

Story: the sports league that was asked to migrate everything

This is a composite drawn from common patterns rather than a single organization. Its details and numbers are illustrative. Consider what you would decide before reading on.

A professional sports league has used one general-purpose model for two years across four workloads. A fan-services assistant answers questions about tickets, season memberships and refunds, tens of thousands a month. A records team uses the model to extract dates, scores and player names from scanned match reports going back decades. The analytics department uses a coding assistant to write and debug the code behind its performance statistics. And the accessibility team uses it to draft text descriptions of photos and charts on the league’s website for fans who use screen readers.

A new frontier model has just been released with markedly better reasoning. At the same time, small models have become far cheaper for routine work. The chief information officer proposes migrating all four workloads to the new model before the new season. Finance worries about the bill. Procurement worries about depending on one provider. The accessibility team worries that any change will alter descriptions fans rely on.

What would you decide? Migrate everything, freeze everything, or something else?

Testing each workload on its own cases led to four decisions - route one to a small model, keep one, move one to the new model and hold one for review.WorkloadWhat its own test showedDecisionFan servicesSmall model matched current answers on 96%of casesRoute to small model; escalate the restMatch-reportextractionCurrent model already above its barKeep; re-test each seasonAnalytics codeNew model fixed far more failing casesMove to the new modelImage descriptionsMixed results; some errors on chartsHold; specialists review samples
Figure 10.2.8 Illustrative composite. One release, four different answers. The evaluation set, not the announcement, decided each one.

The league did something else. Over two weeks, each workload owner assembled a few hundred past cases with known good answers. The tests gave four different answers. The fan-services questions were mostly simple, and a small model matched current answers on almost all of them, so that workload moved to the small model with harder questions escalated to the general one. Match-report extraction already met its bar; changing it would add risk for no gain. The analytics coding assistant improved sharply on the new model, so it moved. Image descriptions were mixed, so the accessibility team kept the current model and added a specialist review of samples.

The routing layer the IT team built for this made each change a configuration setting, and the next release a test rather than a project. The overall bill fell even though one workload moved to a more expensive model. The lesson is not that newer models are overrated. It is that “migrate everything” and “freeze everything” are both decisions taken without evidence, and the evidence was cheap to get.

What this means for leaders

The evolution of AI models has made capability abundant and impermanent. That should change what you ask your teams for. Stop asking which model is best. Ask which workloads need which tier of capability, what your own cases say about any new release, and how quickly you could change the model under a critical workflow if you had to. The model will change. The evaluation set, the routing layer and the workflow are the parts you keep.

Check yourself

  1. The best freely downloadable models are years behind the best closed models.
  2. A model with more parameters will outperform a smaller model trained on the same computing budget.
  3. Work that needed a very large model in 2022 can often be done by a far smaller model today.
  4. A strong public benchmark score is good evidence that a model suits your workflow.
  5. A hosted model can behave differently a few months later under the same name.
  6. When a better model is released, the safest course is to migrate every workload to it.

Reflection: your model map

What comes next

The engines described here did more than make models smarter and cheaper. They made them able to read images and speech, and to act through tools. The next chapter, Multimodal and Agentic AI, looks at those two shifts and what they mean for the work AI can take on.

Laws referenced

EU AI Act · EU

Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744

Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.

  • 2024-08-01 — Entered into force
  • 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
  • 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
  • 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
  • 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
  • 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
  • 2028-08-02 — High-risk obligations for AI in products regulated under Annex I

Last verified 2026-10-06 · official text

References

  1. Epoch AI. Open models lag state-of-the-art closed models by 4 months. Epoch AI (Data Insights). 2026.
  2. Stanford Institute for Human-Centered AI (HAI). AI Index Report 2026. Stanford University. 2026.
  3. Ashish Vaswani et al. Attention Is All You Need. Advances in Neural Information Processing Systems 30 (NeurIPS 2017). 2017.
  4. Jared Kaplan et al. Scaling Laws for Neural Language Models. arXiv:2001.08361. 2020.
  5. Stanford Institute for Human-Centered AI (HAI). AI Index Report 2025, Chapter 1: Research and Development. Stanford University. 2025.
  6. Jordan Hoffmann et al. Training Compute-Optimal Large Language Models. DeepMind (arXiv:2203.15556); NeurIPS 2022. 2022.
  7. Long Ouyang et al. Training language models to follow instructions with human feedback. NeurIPS 2022 (arXiv:2203.02155). 2022.
  8. OpenAI. Learning to reason with LLMs. OpenAI. 2024.
  9. DeepSeek-AI. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv:2501.12948. 2025.
  10. Stanford Institute for Human-Centered AI (HAI). AI Index Report 2025, Chapter 2: Technical Performance. Stanford University. 2025.
  11. Lingjiao Chen, Matei Zaharia and James Zou. FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. arXiv:2305.05176. 2023.
  12. Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous and Ion Stoica. RouteLLM: Learning to Route LLMs with Preference Data. arXiv:2406.18665. 2024.
  13. Lingjiao Chen, Matei Zaharia and James Zou. How Is ChatGPT's Behavior Changing Over Time?. Harvard Data Science Review 6(2). 2024.
  14. European Commission. Guidelines for providers of general-purpose AI models. European Commission, Shaping Europe's digital future. 2025.
  15. WilmerHale. European Commission Issues Guidelines for Providers of General-Purpose AI Models. WilmerHale Privacy and Cybersecurity Law blog. 2025.

Further reading

Sources last verified 2026-10-08.