AI Academy · Book
Executives & Directors · Module 10 · Chapter 001

Where AI Is Going Next

The capabilities many roadmaps still list as "next", reasoning, perception and action, are already in production. What changes from here is how reliably, how cheaply, how autonomously and how physically AI can do that work. Leaders who plan around thresholds rather than predictions are rarely surprised and rarely stampeded.

≈ 16 min read

After this chapter you can

  • Describe what AI can already do in production in 2026, and why reasoning, multimodal and agentic capability are no longer "next".
  • Explain the four dials that will change next - reliability, cost, autonomy and physical reach - with evidence for each.
  • Distinguish what a system can do in a demonstration from what it does reliably, and why that gap decides business value.
  • Write a threshold that turns a parked AI idea into a re-testable decision.
  • Use a simple Adopt / Trial / Watch / Park radar to avoid both technology fever and complacency.

Picture a slide from a strategy deck written in 2023. It is a staircase. At the bottom, Today: models and assistants. One step up, Next: reasoning and multimodal AI. Above that, Then: agents and tool use. Higher still, AI-native workflows, and at the top, AI in physical systems. This one is a reconstruction, not a quote from any company. It was careful, hype-free and, by the standards of its day, sensible.

Its middle steps came quickly. OpenAI’s first publicly available reasoning model, one that works through a problem before it answers, shipped in September 20241. A month later a model that could operate a computer by reading the screen, moving the cursor and typing was released, described by its own maker as “still experimental” and “error-prone”2. By mid-2026, 40 percent of respondents at organizations with more than 1 billion dollars in revenue told McKinsey they were scaling AI agents, up from 27 percent a year earlier3.

A reconstructed 2023 roadmap shows reasoning, multimodal AI and agents as future steps; by 2026 the first had shipped and agents were scaling, while AI-native workflows and physical AI were still ahead.TodayModels and assistantsNextReasoning and multimodal AIARRIVED 2024ThenAgents and tool useSCALING IN 2026FutureAI-native workflowsLonger termAI in physical systems
Figure 10.1.1 A typical 2023 roadmap staircase (reconstructed). Reasoning shipped in 2024 and agents were scaling by 2026; the top two steps were still ahead. The steps it named were never the hard part.

The slide was not wrong about direction. It was wrong about the unit. It staged the future by capability names, and names arrive quickly. What decides whether any of them changes your business is something the slide never measured: whether the capability works reliably enough, cheaply enough and safely enough for a particular piece of your work.

The capabilities are here; the thresholds are not

The argument of this chapter fits in two sentences. Today’s AI already reasons, perceives and acts. What changes next is how reliably, how cheaply, how autonomously and how far into the physical world it does those things.

Those four dials, not a list of product categories, are the trajectory to watch. Each one moves at its own speed, and each one matters only when it crosses a line that your workflow cares about: an error rate you can live with, a cost per task below what the work is worth, a degree of independence your controls can supervise, a physical setting where a mistake is survivable. The executive skill is not predicting when that happens. It is knowing in advance which line matters, watching it, and having decided what you will do when it is crossed.

What is already here

It is worth being exact about the starting point, because many plans still treat it as the future. Three capabilities are in routine production use in October 2026.

Reasoning, perception and tool use are in production today; reliability, cost, autonomy and physical reach are still moving.In production todayReasoning through multi-step problemsReading images, audio and documentsUsing tools and operating softwareStill movingDoing it every timeDoing it cheaply at scaleWorking longer without supervisionActing in the physical world
Figure 10.1.2 The capability list is settled. The four dials on the right are where the next few years of change will come from.

Reasoning. Models now spend extra computation working through a problem before answering, which is why their results on mathematics, science and code have jumped. How they do it is covered in Large Language Models and Training vs Inference. Perception. The same systems read images, charts, scanned documents and speech, as Multimodal AI explained. Action. Through tools and connections to other software, models can look things up, fill in forms, run code and complete multi-step tasks, the shift traced in From AI Assistants to AI Agents.

The 2026 AI Index summarizes the result plainly: models now meet or exceed human baselines on PhD-level science questions, multimodal reasoning and competition mathematics4. The same report notes that top models read an analog clock correctly only about half the time. That is the jagged frontier described in AI Is Changing Everything: brilliance and brittleness side by side. It is also the clue to where the trajectory goes next. Note what kind of evidence this is. Benchmark scores show capability on a test; they do not show that the same system delivers results in a real business process. That evidence comes only from deployments, and from tests on your own work.

Reliability: from “can” to “every time”

A demonstration shows that a system can do something. A business process needs it to do it every time, or to know when it has not. The distance between the two is the first and most important dial.

Success on real computer tasks rose from 12 percent in 2024 to about 66 percent in 2026, close to the human baseline of 72 percent but still failing a third of the time.AI agents in 202412%AI agents in 202666%Human baseline72%Source: Stanford AI Index 2026 (OSWorld) · 2026
Figure 10.1.3 Agents doing everyday computer tasks went from 12 to about 66 percent success in two years. They still fail about one task in three.

The progress is real. On OSWorld, a benchmark of everyday tasks across real operating systems and applications, AI agents went from about 12 percent success in 2024 to about 66 percent in 2026, against a human baseline of about 72 percent4. Read the same figure the other way, though, and agents still fail roughly one task in three. OSWorld is still a benchmark: it measures capability under test conditions, not results in a live process.

The research group METR makes the gap more precise. It measures the length of task, in human expert time, that a model can complete. When it first published the method in March 2025, the best model could finish tasks of about an hour with 50 percent success. Asked to succeed 80 percent of the time, the same model managed tasks of only about 15 minutes. Across models, the 80 percent horizon was roughly five times shorter than the 50 percent one5.

That ratio is the executive point. Few business processes can run on a coin flip, so the useful measure is not the most impressive thing a system has ever done but what it does on the thousandth ordinary case. Reliability is also a property of the whole system, including retrieval, checks and escalation, not only of the model, as Accuracy, Hallucination and Reliability showed. Expect the next few years to be spent closing this gap, and expect the closing to be uneven across tasks. The question to ask of any workflow is simple: what error rate can it tolerate, and what does one error cost?

Cost: the same capability keeps getting cheaper

The second dial is the price of capability. Here the trend has been steep and consistent.

The price of reaching a fixed level of AI performance fell between 9 and 900 times a year, with GPT-4-level science results falling about 40 times a year.9xSlowest annual declinePrice of a fixed levelof performance40xPhD-level sciencequestionsGPT-4-level results, per year900xFastest annual declineMost recent, may not lastSource: Epoch AI · March 2025
Figure 10.1.4 The price of a fixed level of AI performance has fallen by between 9 and 900 times a year, depending on the task.

Epoch AI tracked what it costs to reach a fixed level of performance on a range of tests. Depending on the task and the level, prices fell between 9 and 900 times a year. Reaching GPT-4’s level on PhD-level science questions became about 40 times cheaper each year6. The fastest declines were the most recent, so nobody should extrapolate them with confidence. Why AI, Why Now? gave a similar figure from another source: the cost of a GPT-3.5-level answer fell more than 280-fold in two years7.

Two cautions keep this from becoming a simple story. First, a falling price per unit of capability does not mean a falling bill. Reasoning models do more computation per answer, and agents call models many times per task, so the cost per completed task can rise even as the price per word falls, as Training vs Inference explained. Second, price is only one line of total cost; The Economics of AI covers the rest.

The practical consequence is still large. A use case you parked because it cost too much per transaction may be viable a year later without any new invention, simply because the price of the capability it needs has fallen past your line. Cost can turn yesterday’s “no” into today’s “yes”, and it is an easy dial to forget to re-check.

Autonomy: longer tasks, same governance question

The third dial is how long and how independently a system can work toward a goal. METR’s measure gives the clearest long-run view.

The length of software task AI completes with 50 percent success grew from about one hour in early 2025 to at least 16 hours by May 2026, doubling every few months.2019MeasurementbeginsTask length doubles aboutevery 7 months2024The pace quickensAbout every 3 to 4 monthsMar 2025About 1 hourBest model at50% successMay 2026At least 16 hoursAt the limit of the test
Figure 10.1.5 The length of software task AI can complete half the time has grown exponentially. The trend is a signal, not a schedule.

From 2019 to early 2025, the length of task frontier models could complete with 50 percent success doubled about every seven months; fitted only to models released in 2024 and early 2025, it was about 110 days, or three and a half months5. By May 2026, METR estimated that its best-performing model could complete tasks taking a skilled person at least 16 hours, and warned that results above 16 hours are unreliable with its current set of tasks8.

Read this carefully. METR itself stresses that the time horizon is not how long an AI can work on its own; its error bars span about a factor of two in each direction; its tasks are mostly software and research; and horizons on visual computer-use tasks are 40 to 100 times shorter9. The trend does not tell you when an agent will run your month-end close. It tells you that the ceiling on autonomy is rising fast, and that it is rising fastest in software work.

As the technical ceiling rises, autonomy becomes less a question of what the system can do and more a decision about what you will let it do. That decision is governed by the controls set out in Agentic AI and Autonomous Actions: least privilege, approval for consequential steps and actions that can be undone. Those controls do not loosen because the models improve; the case for them grows as agents take on longer chains of work.

Physical systems: slower, narrower and real

The fourth dial is AI acting in the physical world, through vehicles, robots and machines. It moves more slowly than the other three because errors there are physical, testing is expensive and the law is stricter. It is also no longer hypothetical.

Waymo reached 500,000 paid driverless rides a week in March 2026, ten times its weekly volume in May 2024.500,000Paid driverless rides a weekAcross 10 US cities10xGrowth in under two yearsFrom 50,000 a week in May 2024Source: Waymo figures via TechCrunch · March 2026
Figure 10.1.6 Driverless rides scaled tenfold in under two years, after more than a decade of development in a few carefully chosen cities.

Driverless taxis are the clearest case. Waymo reported 500,000 paid rides a week across ten US cities in March 2026, up from 50,000 a week in May 202410. A peer-reviewed comparison over 56.7 million driverless miles found 85 percent fewer crashes involving suspected serious or worse injuries than a human benchmark on the same roads11. In factories, the International Federation of Robotics counted more than 600,000 industrial robots installed in 2025, an 11 percent rise, taking the global working stock to a record 5 million. It credits AI, machine vision and easier programming with raising what robots can do and lowering what they cost to deploy12.

The pattern is the one to remember: physical AI scales only inside a well-defined operating domain, a set of cities, a factory cell, a warehouse, and widens that domain step by step. Its risk surface is different too. A wrong answer can be corrected; a wrong movement may not be reversible. AI + Robotics, later in this module, covers that ground in depth.

Plan for thresholds, not predictions

If the four dials are what moves, the planning discipline follows. Replace predictions such as “agents will handle our procurement by 2028” with conditional statements: if a capability crosses a stated threshold for a named workflow, then we will take a stated action. A threshold has four parts, quality, cost, speed and acceptable risk, and it belongs to a workflow, not to a technology. The card below is an illustration, not a real company’s thresholds.

A threshold card lists each workflow with the quality or cost line it must cross, the signal being watched, the action agreed in advance and its owner.WorkflowThresholdSignal we watchWhen crossedOwnerSupplier contractreview98% of key clausesfound on our test setQuarterly re-testMove to trialLegal operations leadInvoice exceptionsCost per case below aquarter of manualModel price changesRe-run business caseFinance controllerWarehouse pickingrobotsProven on our item mixPeer site resultsPilot one aisleOperations director
Figure 10.1.7 Illustrative. A threshold card makes a parked idea re-testable: a named workflow, a line to cross, a signal, a pre-agreed action and an owner.

Thresholds need a home. A simple radar, adapted from the rings that Thoughtworks has long used for its Technology Radar13, sorts every capability you track into one of four positions.

A four-level radar - adopt, trial, watch and park - where capabilities move up only when evidence crosses a written threshold.AdoptProven on our work; scale itBudget and ownerTrialPromising; needs our evidenceBounded experimentWatchRelevant; not yet over the lineThreshold written downParkNo material link to our value chainReview yearlyEVIDENCERISES
Figure 10.1.8 Every tracked capability sits in one of four positions, and moves only when evidence, not an announcement, crosses its threshold.

The radar protects against two opposite failures. The first is technology fever: every release enters the roadmap, and the organization reprioritizes constantly without learning anything. The second is complacency: the current stack is declared sufficient, and nobody notices when the economics change underneath it. Both fail for the same reason. Neither has written down what evidence would change its mind.

Two further habits make the radar work. Remember that model releases arrive in months while redesigned work takes years, a distinction drawn in The Speed of AI Adoption. And test the plan against more than one future, which Preparing for AI Uncertainty and Strategic Change develops into a full method.

Story: the feature that was parked correctly

This is a composite drawn from common patterns rather than a single company, and its details are illustrative.

A software company correctly parked an AI feature, set no trigger to revisit it, and learned from a customer that the economics had changed.ParkedEarly 2024:accuracy and costbelow the lineForgottenNo threshold,owner orre-test dateOvertakenPrices fell;quality roseDiscoveredCustomer builtit in-houseThe decision was right. The missing trigger was the failure.
Figure 10.1.9 An illustrative composite post-mortem in four steps: a sound decision to wait, with nothing in place to notice when waiting stopped being sound.

A mid-sized company sells procurement software to manufacturers. In early 2024 its product team tested AI to pull payment terms, renewal dates and penalty clauses out of supplier contracts, a feature customers had asked for. On a test set of scanned contracts, the model missed too many clauses, and the cost per contract was too high for the product’s price point. The team recommended waiting. The leadership team agreed. It was a good decision.

Eighteen months later, during a renewal negotiation, the company’s largest customer mentioned that its own operations team had built contract extraction into its workflow with an AI coding tool, and asked why it was still paying for manual uploads. The renewal closed at a discount. The customer was not unusual: in McKinsey’s 2026 survey, 32 percent of all respondents, not only those at large firms, said their organizations had decided against buying a software product or feature because they could build it with agentic coding tools3.

The post-mortem found no bad decision, only missing machinery. The 2024 test had no written threshold, so nobody could say what result would have changed the answer. The feature had no owner after the product manager moved on. The test set had been discarded, so re-testing would have taken weeks rather than an afternoon. And the company’s watch on AI was a newsletter of model announcements, which told it a great deal about what had been released and nothing about whether its own line had been crossed. The fix took a month: a threshold card for every parked AI idea, a named owner, a kept test set and a quarterly re-test.

What this means for leaders

The future of AI is best treated as four dials moving at different speeds, not as a staircase of product categories. That changes what you ask for. Ask your teams for the reliability their workflows need, the cost per completed task they can afford, the autonomy their controls can supervise and the physical settings where a mistake is survivable. Then ask them to watch the gap between those lines and today’s capability, rather than the stream of announcements.

Check yourself

  1. Reasoning models, multimodal AI and agents that use tools are still mostly future capabilities.
  2. An AI agent that succeeds on two-thirds of everyday computer tasks is ready to run a business process unattended.
  3. The price of reaching a fixed level of AI performance has been falling by many times a year.
  4. Falling prices per unit of AI capability guarantee a falling cost per completed task.
  5. METR’s time horizon measures how long an AI system can work on its own.
  6. Physical AI tends to scale inside a well-defined operating domain and widen it step by step.

Reflection: your parked list

What comes next

Each of the four dials depends, in the end, on the models underneath. Why have they become better at reasoning, cheaper to run and able to see and act, and what does that mean for the choices you make about them? The next chapter, The Evolution of AI Models, looks inside the technology itself.

Laws referenced

EU AI Act · EU

Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744

Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.

  • 2024-08-01 — Entered into force
  • 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
  • 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
  • 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
  • 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
  • 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
  • 2028-08-02 — High-risk obligations for AI in products regulated under Annex I

Last verified 2026-10-06 · official text

EU Machinery Regulation · EU

Regulation (EU) 2023/1230

Safety rules for machinery, covering self-evolving behaviour and safety components that use AI. Relevant to robots and autonomous equipment.

  • 2027-01-20 — Applies

Last verified 2026-10-06 · official text

EU Product Liability Directive (revised) · EU

Directive (EU) 2024/2853

No-fault liability now explicitly covers software, including AI systems and SaaS, and updates or the lack of security updates. Easier proof for claimants with complex products.

  • 2026-12-09 — Applies to products placed on the market from this date

Last verified 2026-10-06 · official text

References

  1. OpenAI. Learning to reason with LLMs. OpenAI. 2024.
  2. Anthropic. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. Anthropic. 2024.
  3. McKinsey & Company (QuantumBlack). The state of AI in 2026: On the road to ROI. McKinsey & Company. 2026.
  4. Stanford Institute for Human-Centered AI (HAI). AI Index Report 2026. Stanford University. 2026.
  5. Thomas Kwa, Ben West, Joel Becker and others (METR). Measuring AI Ability to Complete Long Software Tasks. arXiv 2503.14499. 2025.
  6. Ben Cottier, Ben Snodin, David Owen and Tom Adamczewski. LLM inference prices have fallen rapidly but unequally across tasks. Epoch AI. 2025.
  7. Stanford Institute for Human-Centered AI (HAI). AI Index Report 2025, Chapter 1: Research and Development. Stanford University. 2025.
  8. METR. Task-Completion Time Horizons of Frontier AI Models. METR. 2026.
  9. METR. Clarifying limitations of time horizon. METR. 2026.
  10. TechCrunch. Waymo's skyrocketing ridership in one chart. TechCrunch. 2026.
  11. Kristofer D. Kusano and others. Comparison of Waymo Rider-Only crash rates by crash type to human benchmarks at 56.7 million miles. Traffic Injury Prevention. 2025.
  12. International Federation of Robotics. Five Million Robots now Operate in Factories Globally. International Federation of Robotics (World Robotics 2026 press release). 2026.
  13. Thoughtworks. Technology Radar FAQ. Thoughtworks. 2026.

Further reading

Sources last verified 2026-10-08.