Building the Multi-Year AI Roadmap
A multi-year AI roadmap is a sequence of decisions made under falling certainty, not a three-year promise. Only the current quarter is a contract, the first year is a commitment, and later years are a direction held in place by named assumptions. The roadmap earns its keep through the dependencies it draws, the decisions it schedules and the signals that would change it.
After this chapter you can
- Distinguish the four levels of firmness on a roadmap - contract, commitment, direction and option - and explain why certainty falls with distance.
- Build the roadmap in three layers (outcomes, use cases, capabilities) with people and governance as capability lanes.
- Draw the dependencies between lanes so that no use case runs ahead of its controls and no capability runs ahead of demand.
- Replace activity milestones with decision, capability and outcome milestones that have owners and dates.
- Name load-bearing assumptions with signposts, and change the roadmap only at known moments.
Bent Flyvbjerg has spent three decades collecting data on large projects: bridges, IT systems, dams, rail lines, Olympic Games. His database now holds more than 16,000 of them, and in the book he wrote with Dan Gardner he reports that just under half, 47.9 percent, came in on budget or better. Only 8.5 percent came in on budget and on time, roughly one in twelve. And only 0.5 percent, one project in two hundred, came in on budget, on time and delivered the benefits that had been promised1.
A three-year AI roadmap is a bundle of large projects: platforms, data programs, use cases, changes to how thousands of people work. If it is written as a promise that every box will arrive on budget, on time and with its business case intact, the base rate says it will be broken, and probably early. That does not make multi-year planning pointless. It changes what the roadmap is for. Its job is not to predict the third year. Its job is to make the right decisions happen in the right order, and to show in advance what would change the plan.
A sequence of decisions under falling certainty
What Is an Enterprise AI Strategy? drew the boundary: the strategy says where the organization will compete and why, and the roadmap says in what order the work happens. Building the 90-Day AI Plan built the first slice of that order, a contract for one quarter. The multi-year roadmap is the structure that quarter sits inside, and the most useful thing it can show is how firm each part of it is.
Each level promises something different. The contract is the current quarter, exactly as the previous chapter defined it: a few outcomes, one owner each, and a decision on Day 90. The commitment is the first year: funded themes with named owners, whose order can change at each quarterly decision without anyone calling it a failure. The direction covers the second and third years. It names the capabilities the organization intends to have and the conditions under which it will invest in them, but not tasks and dates. Options are the bets that Quick Wins, Strategic Bets and Transformation Initiatives described, held open at low cost until evidence says whether to exercise them.
The gradient matters because uncertainty in technology work is not evenly spread. When Flyvbjerg and Alexander Budzier studied 1,471 IT projects, the average cost overrun was a manageable 27 percent. But one project in six was what they called a black swan, with an average cost overrun of 200 percent and a schedule overrun of almost 70 percent2. Averages hide the tail, and the tail is where multi-year programs die.
Flyvbjerg’s prescription is to build big things from small, repeatable modules, because modular projects are faster, cheaper and less likely to land in the tail1. For an AI roadmap, that means sequencing a series of production steps that each pay off and each teach something, rather than a single program that pays off only if everything arrives together.
Three layers: why, what and how
Roadmapping is older than AI. Motorola described its technology roadmap process in 1987, as a way to line up product plans with the technologies they would need3. Robert Phaal and colleagues at Cambridge later generalized the practice into a layered chart: the market and business layer says why, the product and service layer says what, and the technology and resources layer says how, all plotted against the same time axis4. Their point was that the value of the chart lies in the links between layers, not in any one layer.
The layers translate directly. Outcomes are the measured results the strategy promised. Use cases are the systems in production and the changed workflows around them. Capabilities are everything the use cases stand on: data access, the shared platform, evaluation and risk review, and the skills and roles of the people who will work differently. People and governance belong in the capability layer as lanes of their own. Redesigning a role, retraining a team or consulting employee representatives often takes longer than building the software, and a roadmap that leaves them out has hidden its longest lead times.
A quick test of any roadmap you are shown: cover the bottom layer with your hand. If what remains says nothing specific, the document is a technology plan with a business title.
Draw the arrows between lanes
The boxes on a roadmap are the easy part. The arrows are the work. Every use case depends on capabilities beneath it, and every capability should exist because a use case above it needs it. Suppose a mid-sized container port is drafting its roadmap. Its chart might look like this.
Three rules keep the arrows honest. First, no use case runs ahead of the capability it depends on. If the berth assistant needs risk review and the review process does not exist yet, the assistant waits, or the review is accelerated. Drawing the arrow forces that conversation before launch rather than after. Second, no capability runs ahead of the demand for it. Prioritizing the AI Portfolio gave the enabler test and AI Platform Strategy argued for a demand-led platform; on a roadmap, that shows up as a shared service that appears only in the year a second use case needs it. Third, the people lane has lead times too. If crane scheduling will change how yard planners work, the role redesign starts a year earlier, because it sets the earliest date the scheduling system can deliver value.
Milestones are decisions
Many roadmaps are studded with milestones that record activity: a program kicked off, a platform launched, a thousand people trained, a set of actions started. Activity milestones are easy to hit and say almost nothing about whether the organization is closer to its outcomes. A roadmap for executives should be built around a different kind.
A decision milestone names a date on which leadership will choose between defined options, on defined evidence. Two other kinds support it. A capability milestone counts only when the capability is in use by a use case, not when it is built. An outcome milestone counts only when the result is measured against a baseline. If a milestone can be reached without anyone deciding, using or measuring anything, it belongs in a project plan, not on the roadmap.
Decision milestones also make the roadmap governable. Each one has an owner who will make the call and a forum where it is made, which is exactly what the next chapter, AI Transformation Governance and Executive Sponsorship, builds.
Name the assumptions that hold it up
Every direction on a roadmap rests on assumptions: that the data will be accessible, that model costs will keep falling, that customers will accept an automated answer, that a regulation will apply from a certain date. The discipline for handling them comes from James Dewar’s assumption-based planning, developed at RAND for the US Army in the early 1990s5.
The method has five steps. Identify the load-bearing assumptions, the ones whose failure would force a real change to the plan. Find which of those are vulnerable within the roadmap’s horizon. For each vulnerable one, define a signpost: an observable signal that the assumption is failing, with a named person watching it. Then plan shaping actions that help keep the assumption true, and hedging actions that prepare for the case where it does not hold.
For the port, one load-bearing assumption behind the gate-booking forecast is that trucking companies will book slots through the port’s system. A signpost might be the share of trucks arriving with a booking by the end of year one. A shaping action is an incentive for booked arrivals; a hedge is a forecast that works from arrival history alone. If the signpost stays below its line, the year-two direction changes, and everyone knew in advance that it might.
Rebuild it on a schedule
Henry Mintzberg and James Waters showed that the strategy an organization actually follows is rarely purely the one it intended. Some intentions are realized, some are dropped, and new patterns emerge from what people learn along the way6. A roadmap that cannot absorb emergent learning ends up describing an organization that no longer exists. The answer is not to revise it constantly, which destroys the commitment layer, but to revise it at known moments.
There are three such moments. Every quarter, the Day 90 decisions settle the contract and may re-sequence the year-one commitment. Whenever a signpost fires, the lanes resting on that assumption are replanned, and only those. Every year, the whole roadmap is rebuilt from the strategy: the next slice of direction becomes commitment, and what has been learned reshapes the direction beyond it. The meeting rhythm that carries these reviews belongs to AI Change Management, KPIs and Operating Rhythm.
Story: the plan that was on schedule
This is a documented public case, told as a dilemma. Decide what you would have done before reading the outcome.
In October 2023, New York City published its AI Action Plan, which the city described as the first of its kind for a major American city. It was a serious document: 37 actions in seven initiatives, each with a timeframe. Action 1.6 would develop an AI risk assessment and project review process, to be initiated within twelve months. Action 1.8 would monitor AI tools in operation. Initiative 7 committed the city to refresh the plan and report on progress every year7. Announced alongside the plan was a live public service: the MyCity chatbot, which answered business owners’ questions about operating in the city.
In March 2024, reporters at The Markup tested the chatbot. It told users that employers could take a cut of their workers’ tips, that landlords did not have to accept housing vouchers and that businesses did not have to accept cash. Each answer contradicted the law8.
You lead the program. The plan is on schedule, and the risk review it promised has until October to get started. You have three options. Keep the chatbot live, add a warning and fix it in service. Take it offline until the risk review exists. Or narrow it to answers drawn from verified city pages while the review is built.
The city kept it live. A beta warning was added telling users not to rely on the answers as legal or professional advice, and the bot stayed online8. Two years after publication, the city’s progress report recorded 35 of the 37 actions initiated or completed. The risk review process and the monitoring of AI tools in operation were both still listed as in progress, with draft risk policies piloted on several agency projects9. In January 2026 the new mayor called the chatbot “functionally unusable”, said it was costing around half a million dollars a year, and announced its end as part of closing a budget gap; by early February it was offline10.
Read as a roadmap, the plan had several strengths: dated actions, a governance lane, a refresh mechanism and honest annual reporting. What it lacked were two things this chapter is about. There was no arrow between the live service in the value lane and the review in the governance lane that it depended on, so the service could run two years ahead of its control. And its milestones were mostly activities, actions initiated or completed, rather than a decision with a date: on what evidence would the chatbot be paused, narrowed or ended? The plan stayed on schedule while the decision waited for a new administration.
What this means for leaders
A multi-year roadmap is a management instrument, not a forecast. The executive’s job is to keep its levels of firmness honest, so that only the quarter is treated as a contract and the third year is not defended as if it were one. It is to insist on the arrows between lanes, especially between anything live and the controls it relies on. It is to replace activity milestones with decisions that have owners and dates, to name the assumptions that hold the direction up and to make sure someone is watching each signpost. And it is to rebuild the roadmap every year from the strategy, rather than renewing last year’s chart with the dates moved right.
Check yourself
- A good multi-year AI roadmap commits to the scope, budget and date of every item for all three years.
- In Flyvbjerg’s database of more than 16,000 large projects, about one in two hundred met budget, schedule and benefits.
- Because the average IT project overran its budget by only 27 percent, a 30 percent contingency covers the risk.
- A shared platform should usually appear on the roadmap in the year a second use case needs it.
- “Platform launched” and “1,000 staff trained” are strong executive milestones.
- A roadmap should be changed whenever new information arrives.
Reflection: the roadmap on your wall
What comes next
A roadmap full of decision milestones is only as good as the people who make the decisions. When functions disagree, when a signpost fires in someone else’s budget, or when a live system needs pausing, someone must have the authority to act. The next chapter, AI Transformation Governance and Executive Sponsorship, sets out who that is and how sponsorship keeps the sequence moving across organizational boundaries.
Laws referenced
Not legal advice. Laws change; verify before relying on this, and consult counsel for decisions.
EU AI Act · EU
Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744
Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.
- 2024-08-01 — Entered into force
- 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
- 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
- 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
- 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
- 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
- 2028-08-02 — High-risk obligations for AI in products regulated under Annex I
Last verified 2026-10-06 · official text
References
- Bent Flyvbjerg and Dan Gardner. How Big Things Get Done: The Surprising Factors That Determine the Fate of Every Project, from Home Renovations to Space Exploration and Everything in Between. Currency (Penguin Random House). 2023.
- Bent Flyvbjerg and Alexander Budzier. Why Your IT Project May Be Riskier Than You Think. Harvard Business Review. 2011.
- Charles H. Willyard and Cheryl W. McClees. Motorola's Technology Roadmap Process. Research Management, 30(5), 13-19. 1987.
- Robert Phaal, Clare J. P. Farrukh and David R. Probert. Technology roadmapping - A planning framework for evolution and revolution. Technological Forecasting and Social Change, 71(1-2), 5-26. 2004.
- James A. Dewar. Assumption-Based Planning: A Tool for Reducing Avoidable Surprises. Cambridge University Press (RAND Studies in Policy Analysis). 2002.
- Henry Mintzberg and James A. Waters. Of strategies, deliberate and emergent. Strategic Management Journal, 6(3), 257-272. 1985.
- New York City Office of Technology and Innovation. New York City Artificial Intelligence Action Plan. City of New York. 2023.
- Colin Lecher. NYC's AI Chatbot Tells Businesses to Break the Law. The Markup (with THE CITY). 2024.
- New York City Office of Technology and Innovation. New York City AI Action Plan Annual Progress Report. City of New York. 2025.
- The Markup. Mamdani to kill the NYC AI chatbot we caught telling businesses to break the law. The Markup. 2026.
Further reading
- Robert Phaal, Clare J. P. Farrukh and David R. Probert. Technology roadmapping - A planning framework for evolution and revolution. Technological Forecasting and Social Change, 71(1-2), 5-26. 2004.
- James A. Dewar. Assumption-Based Planning: A Tool for Reducing Avoidable Surprises. Cambridge University Press (RAND Studies in Policy Analysis). 2002.
- Bent Flyvbjerg and Dan Gardner. How Big Things Get Done: The Surprising Factors That Determine the Fate of Every Project, from Home Renovations to Space Exploration and Everything in Between. Currency (Penguin Random House). 2023.
Sources last verified 2026-10-08.