The AI Maturity Model
AI maturity is not how many pilots an organization runs or how much it spends. It is how reliably it can turn AI into governed, measured value, again and again. Five levels describe that climb; the right target is the level your strategy requires, assessed with evidence, unit by unit.
After this chapter you can
- Define AI maturity as repeatable, governed, measured AI value, and distinguish it from ambition and readiness.
- Describe the five levels of the course's maturity ladder, with the evidence and exit test for each.
- Explain why the step from experimentation to production is where most organizations stall.
- Set target maturity from strategy, risk and law rather than from the top of the ladder.
- Replace a single maturity score with an evidence-backed profile that produces owned actions.
Before reading on, make a prediction. In the last quarter of 2024 Gartner surveyed leaders at 432 organizations in six countries and placed each one on its AI maturity model, using a seven-question assessment. It then asked how long their AI initiatives stayed in production. What share of high-maturity organizations said their AI initiatives remain in production for three years or more, and what share of low-maturity ones?
The answers were 45 percent and 20 percent, a gap of more than two to one. On trust the gap was wider: business units were ready to use new AI solutions in 57 percent of high-maturity organizations and only 14 percent of low-maturity ones1.
These are self-reported figures from one analyst firm’s survey, and they show association, not cause. But they point at a capability gap. Almost every large organization can start an AI initiative. Far fewer can run one reliably, measure what it returns, and then do the same thing a second and a tenth time without starting from scratch. That ability is what this book means by AI maturity.
Maturity is repeatability, not activity
Ask a leadership team which company is more mature: one with 200 AI pilots, or one with 20 AI systems that reliably create measured business value. The question usually produces a pause, and the pause is useful. Pilot count, licenses bought and money spent are signs of activity. None of them shows that the organization can operate AI.
A working definition is this: AI maturity is an organization’s ability to identify, build, deploy, operate, govern, measure and scale AI value, repeatedly. The last word carries the weight. A single success can come from one talented team and a lucky choice of problem. Maturity is the system that makes the next success likely.
Three questions are easy to blur, and each has its own place in this course.
As Assessing Enterprise AI Readiness showed, readiness only means something against a named ambition. Maturity is the broader, slower-moving view: the habits, platforms and controls the organization brings to every AI ambition it attempts. An organization can be ready enough to run a careful experiment and still be immature as an enterprise.
Where maturity ladders come from
Maturity models are older than modern AI, and their history explains both their value and their limits. Richard Nolan described in 1979 how corporate computing grew in stages, with a contagion phase of fast, uncoordinated growth before a control phase2. Five years later John King and Kenneth Kraemer found many of his claims empirically weak: organizations did not march through the stages on a fixed clock. They still judged the model useful, which is the right attitude to any ladder, including this one3.
The most influential ladder came from software. In 1993 Carnegie Mellon’s Software Engineering Institute named the second level of its Capability Maturity Model Repeatable, the level at which basic management practices let an organization repeat earlier successes instead of depending on heroics. Its authors also warned that skipping levels is counterproductive4. Both lessons carry over to AI: the core of maturity is repeatability, and an organization cannot scale what it cannot yet run.
The five levels
The course uses one maturity ladder, with five levels. Each level is defined by what the organization can reliably do, not by the technology it owns.
Level 1, Awareness. AI is interesting. Executives are curious, individuals use public tools, and nobody can say with confidence where AI is used across the organization. The first objective is coordination: a sponsor, an inventory of use, basic rules for acceptable use and enough literacy for people to use the tools safely.
Level 2, Experimentation. The organization tests AI on purpose. Pilots have hypotheses, success criteria and a controlled environment. This is real progress, and it carries the best-known risk on the ladder, discussed in the next section.
Level 3, Production. Selected AI systems run inside real business processes. Each has a named owner, service expectations, monitoring, security, an incident route and a way to switch it off. The organization measures total cost, cost per unit of work and value against a baseline. The main risk at this level is that every system is a one-off, built in its own way by its own team.
Level 4, Scaled. The organization can deliver many AI systems consistently. Shared capabilities, such as model access, data services, evaluation, monitoring, security and cost management, mean the tenth system is faster and safer than the first. AI is managed as a portfolio in which initiatives are started, scaled, optimized and stopped. The risk here is the opposite of level 3: a central function that becomes a queue.
Level 5, AI-native. AI is built into strategy, products, decisions and operations. It is no longer mainly a technology program but an ordinary part of how the business creates value, and the organization keeps learning and reconfiguring as the technology changes. Organizational learning is the defining trait. An MIT Sloan Management Review and BCG study found that only about one company in ten reported significant financial benefits from AI, and that those which learned with AI, changing how people and systems work together, were far more likely to5.
Each level is easier to recognize by its evidence and its exit test than by its label.
There is no standard number of months per level. The pace depends on the starting point, the industry’s constraints, the investment and the size of the organization. Anyone who quotes a universal timeline is guessing.
The step that stalls: from experiment to production
The transition from level 2 to level 3 is where many organizations get stuck. Practitioners call it pilot purgatory: many experiments, few systems in production, and a steady stream of new pilots that hides the fact that nothing is being converted.
The evidence points the same way from several directions. In BCG’s 2024 survey of 1,000 senior executives in 59 countries, only 26 percent of companies had built the capabilities to move beyond proofs of concept and generate tangible value, and just 4 percent had cutting-edge capabilities across functions6. The measures differ: BCG classifies companies by capability, while McKinsey counts survey respondents. In McKinsey’s 2025 global survey, nearly two-thirds of respondents reported that their organizations had not yet begun to scale AI across the enterprise7. In 2024 Gartner predicted that at least 30 percent of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs and unclear business value8. That was a forecast, not a measurement, but its list of causes is a fair description of what level 3 demands.
The lesson for maturity is simple. The way out of purgatory is not more pilots. It is production engineering, measured value, security, governance and an owner who will still be accountable a year after launch. How to move a particular system across that boundary is the subject of From AI Pilot to Production to Scale, later in this module. For the maturity assessment, the question is whether the organization has done it repeatedly, or only once.
Level 5 is not a trophy
A common misuse of a maturity ladder is to treat the top rung as the goal everywhere. It is not. The right target is the maturity the strategy requires, and that differs by use, risk and scale.
An internal drafting assistant needs level 3 discipline: an owner, sensible controls and measured use. An AI system that informs credit decisions about customers needs production discipline plus full lifecycle governance, because the law treats it as high-risk. Agents running workflows across the enterprise need the shared platform, evaluation and portfolio control of level 4. Few organizations need level 5 everywhere, and none needs it before it has mastered level 3.
Target-setting also guards against two other mistakes. The first is copying a benchmark from another industry. A software company’s maturity profile is not a requirement for a utility or a bank, whose risk, regulation and business model differ. The second is waiting for the top before expecting value. Significant value is created on levels 2, 3 and 4, and organizations that wait for AI-native status delay both learning and returns. The level of ambition itself is a strategic choice, as Defining AI Ambition argued.
Uneven by design: a profile, not a score
No organization sits on one rung. It can be at level 4 in technology, level 3 in data, level 2 in governance and level 2 in change management. Business units sit on different rungs of the same ladder. A single enterprise score averages these into a number that describes nothing, and it hides the gap that is actually holding the organization back.
The useful view is a profile: the same dimensions used for readiness, rated for each major business unit, with evidence behind every rating.
The same ladder applies inside each dimension. Governance moves from a basic policy to experiment controls, to lifecycle governance, to portfolio governance and finally to continuous oversight. Economics moves from knowing what is spent, to experiment budgets, to total and unit cost, to portfolio economics. A leadership team does not need to recite every sub-ladder. It does need to stop presenting one number.
Evidence, not opinion
Maturity is easy to overstate, because the signals that are easiest to count are the least informative.
The last item on the right deserves attention. A mature organization stops things. It retires systems that no longer pay their way and ends pilots that failed their test. A portfolio that only ever grows is a sign of level 2 behavior, however large it is.
A maturity assessment that ends in a number has failed. For each dimension and unit, it should record five things: the current level, the evidence for it, the level the strategy requires, the gap that matters most, and the first action to close it. The gap is an input to the roadmap. It is not automatically the top priority, because a large gap in an area the strategy does not need may matter less than a small gap that blocks the next step. Choosing among gaps and opportunities is a portfolio decision. The assessment is also not a one-off. Repeat it when strategy, technology, regulation or the organization changes materially, so that it remains a management practice rather than a one-off report.
Story: a day of maturity at a glassmaker
The chief operating officer of a container-glass maker with four plants in the European Union had been asked for one thing by Friday: a maturity score for the board’s annual strategy session. A consultant’s survey had already produced one, 3.1 out of 5. She spent a day finding out what it meant.
At eight she sat with the quality team. Their camera-based models had inspected every bottle and jar on the lines for years. Each had an owner, monitoring for drift, a validation record from the quality engineers and a known cost per thousand containers inspected. When a new mold or glass color changed what the cameras saw, the team retrained and redeployed within days. That was level 4 behavior, built long before generative AI.
At ten she visited customer service. There were 41 generative AI pilots. Three teams had built three different tools for summarizing customer calls. None of the pilots had an owner beyond the original sponsor, and nobody could say what one summary cost. Staff liked the tools. The company could not run them. That was level 2, deep in pilot purgatory.
After lunch she met the workforce-planning team, who wanted generative AI to draft shift allocations and performance notes for plant staff. Their shift-scheduling models were mature, but the route for approving a generative system that touched decisions about workers did not exist yet. The EU AI Act’s high-risk rules for worker management, which apply to such systems from December 2027, made that route a requirement, not a nicety. At three the finance director showed her the total AI spend, which was rising, and admitted he could not report cost per transaction for any generative system.
At half past five she rewrote the board page. The score was gone. In its place were a profile by unit and dimension, with the evidence behind each rating, and a target for the next two years. Level 3 production discipline would cover generative AI in customer service, inspection would stay at level 4, and an approval route would be built for high-risk uses. Level 5 was not the target anywhere. Three gaps came first: a named owner for any pilot older than 90 days, one shared evaluation and approval route for generative systems, and unit-cost reporting. Which of the 41 pilots to keep she left for the next quarter’s portfolio review.
The board spent less time on the number, because there was none, and more on the three gaps.
What this means for leaders
Treat maturity as a measure of repeatability. Before you accept any claim of progress, ask what the organization can now do again, more cheaply and more safely, that it could not do before. Pilots, licenses and spending show activity. Owned production systems, measured value and falling time to production show maturity.
Set the target from the strategy, not from the top of the ladder. Decide what level each major use and unit actually needs, including where the law sets the floor, and resist both under-building and over-building. Insist on a profile instead of a score, and on evidence instead of opinion. Above all, make the assessment produce actions: each material gap should be given an owner and a date, and the next assessment should show whether it moved.
Check yourself
- An organization with more AI pilots is usually more mature.
- Every organization should aim for level 5, AI-native, across the business.
- The step from experimentation to production is where most organizations stall.
- A single enterprise maturity score is the best summary for the board.
- The first widely used maturity model treated repeatability as the step after the initial level.
- Stopping AI initiatives is a sign of low maturity.
Reflection: place yourself on the ladder
What comes next
An honest maturity profile shows where you stand and which gaps matter. It still leaves a choice: where to put limited money, people and executive attention. Not every gap and not every idea deserves the same investment. The next chapter, Prioritizing the AI Portfolio, shows how to decide which opportunities to fund, which to experiment with, which to defer and which to stop.
Laws referenced
Not legal advice. Laws change; verify before relying on this, and consult counsel for decisions.
EU AI Act · EU
Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744
Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.
- 2024-08-01 — Entered into force
- 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
- 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
- 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
- 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
- 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
- 2028-08-02 — High-risk obligations for AI in products regulated under Annex I
Last verified 2026-10-06 · official text
References
- Gartner, Inc. Gartner Survey Finds 45% of Organizations With High AI Maturity Keep AI Projects Operational for at Least Three Years. Gartner Newsroom (press release, Stamford, Conn.). 2025.
- Richard L. Nolan. Managing the Crises in Data Processing. Harvard Business Review, 57(2), 115-126. 1979.
- John Leslie King and Kenneth L. Kraemer. Evolution and Organizational Information Systems: An Assessment of Nolan's Stage Model. Communications of the ACM, 27(5), 466-475. 1984.
- Mark C. Paulk, Bill Curtis, Mary Beth Chrissis and Charles V. Weber. Capability Maturity Model for Software, Version 1.1 (CMU/SEI-93-TR-024). Software Engineering Institute, Carnegie Mellon University. 1993.
- Sam Ransbotham, Shervin Khodabandeh, David Kiron, François Candelon, Michael Chu and Burt LaFountain. Expanding AI's Impact With Organizational Learning. MIT Sloan Management Review and Boston Consulting Group. 2020.
- Boston Consulting Group. Where's the Value in AI?. Boston Consulting Group. 2024.
- McKinsey & Company (QuantumBlack). The state of AI in 2025: Agents, innovation, and transformation. McKinsey & Company. 2025.
- Gartner. Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept By End of 2025. Gartner press release, 29 July 2024. 2024.
- European Union. Regulation (EU) 2026/1744 (Digital Omnibus on AI) amending Regulation (EU) 2024/1689. Official Journal of the European Union. 2026.
Further reading
- Mark C. Paulk, Bill Curtis, Mary Beth Chrissis and Charles V. Weber. Capability Maturity Model for Software, Version 1.1 (CMU/SEI-93-TR-024). Software Engineering Institute, Carnegie Mellon University. 1993.
- John Leslie King and Kenneth L. Kraemer. Evolution and Organizational Information Systems: An Assessment of Nolan's Stage Model. Communications of the ACM, 27(5), 466-475. 1984.
- Boston Consulting Group. Where's the Value in AI?. Boston Consulting Group. 2024.
Sources last verified 2026-10-08.