Reading Time: 11 minutes

Quick Answer: An AI maturity model is a staged framework that rates how deeply and effectively an organisation uses AI, scored across dimensions such as strategy, data, governance, skills and measured value. Most published models use four or five stages, running from informal experimentation through to AI reshaping the operating model itself. The stage you land on matters far less than the gaps that stage number hides.

Deloitte’s State of AI in the Enterprise 2026, a survey of 3,235 senior leaders across 24 countries, found that 42% of organisations now rate their strategy as highly prepared for AI adoption. Those same organisations rated themselves considerably less prepared on infrastructure, data, risk and talent. Deloitte describes this as being strategically ready but operationally unsure.

That gap is the reason AI maturity models exist. It is also the reason most of them disappoint. A single overall grade averages away precisely the information a leader needs in order to act.

This guide covers what an AI maturity model is, what each stage looks like, how the leading frameworks compare, and what actually moves an organisation up a level.

What is an AI maturity model?

An AI maturity model is a structured framework that assesses how capable an organisation is at using AI, then places it at a defined stage on a progression. It scores that capability across several dimensions rather than one, typically covering strategy, data, technology, governance, talent and business value.

It has three practical uses: establishing a baseline, guiding investment by showing which dimension is furthest behind, and tracking movement over time so a claim of improvement can be evidenced rather than asserted.

One distinction is worth making early, because it causes constant confusion. An AI maturity model is the framework itself. An AI maturity assessment is the exercise of scoring your organisation against one. You can run many assessments against a single model, which is rather the point of having a model.

A credible AI maturity model looks beyond tooling. Buying more licences moves nothing on its own, because a licence is not a capability. What matters is whether the work itself has changed, which is closely tied to what agentic AI is and how it differs from earlier automation.

Where AI maturity models came from

The five-stage shape did not originate with AI. It comes from the Capability Maturity Model, developed at Carnegie Mellon University in the late 1980s to assess how rigorously software teams followed their own processes. Its levels ran from ad hoc through to optimised.

That structure was borrowed more or less wholesale when AI arrived. Microsoft states outright that its agentic AI adoption maturity model is based on the Capability Maturity Model, and most vendor frameworks follow the same shape whether or not they say so.

That lineage matters. The original model was built for slow-moving process audits, where an annual assessment was reasonable because the thing being measured moved slowly. AI capability does not, and that mismatch explains a great deal further down.

The five stages of AI maturity

Stage names differ between published frameworks, but the underlying progression is remarkably consistent. Informal use, then contained pilots, then something real in production, then spread across the organisation, then a genuinely changed operating model.

The five stages below use our own naming. Other published frameworks label a broadly similar progression differently, and the comparison table further down sets the main ones side by side.

StageWhat it looks likeTypical failure modeWhat moves you up
1. UnmappedInformal individual tool use, no coordination, no view of which work is worth automatingNothing compounds, the same solution is reinvented team by teamEstablish what the work actually is before choosing tools
2. ContainedFunded pilots and executive interest, little reaching productionPilots run indefinitely and never convertPick one workflow, give it an owner and a budget, ship it
3. EmbeddedAt least one workflow live with a named owner and measurementSuccess stays trapped in the team that built itA repeatable route from idea to production, plus broad enablement
4. CompoundingWins spread across functions, agents owned centrally, ROI measuredGovernance lags behind deployment speedCentral ownership of agents and continuous measurement
5. Self-improvingThe operating model adapts and the score drives reinvestmentDrift between what is deployed and what is actually governedContinuous re-scoring, with the score setting the next priority

Stage 1: Unmapped

Individuals use AI tools on their own initiative. Someone in marketing drafts copy with a chatbot, someone in operations has quietly built a forecasting spreadsheet, and none of it is coordinated, funded or visible to leadership. Almost all of it is generative rather than agentic AI, assisting a person rather than completing work. The defining feature is not an absence of tools. It is that nobody has established which work is genuinely worth automating, so the same solution gets reinvented team by team and nothing accumulates.

Stage 2: Contained

Pilots exist, they have budget, and executives are paying attention. What they lack is reach. Results stay inside the team that ran them and little converts into production. This is where most organisations stall, and they can stall for years, because commissioning another pilot always feels like forward motion.

Stage 3: Embedded

At least one workflow genuinely runs on AI in production, with a named owner, a budget and measurement attached to it. One workflow truly in production is worth more than ten pilots that never shipped. The risk is that success stays trapped in the team that built it, which is the containment problem in a more advanced form.

Stage 4: Compounding

Wins spread across functions. Each AI agent is owned centrally rather than living in an individual employee’s personal account, so capability stays with the organisation when people leave. Returns are measured rather than argued. The usual failure at this stage is governance lagging behind deployment speed, and Deloitte found only one in five companies has a mature governance model for autonomous agents.

Stage 5: Self-improving

The operating model itself adapts. Agents run material parts of the work, the score decides where the next investment goes, and improvement becomes continuous rather than project-based. Very few organisations are genuinely at this stage.

Most readers will place themselves at stage two or three, and the data supports that. McKinsey found that nearly two-thirds of organisations have not yet begun scaling AI across the enterprise, and just 39% report any EBIT impact at enterprise level. Deloitte’s split is similar: 34% are using AI to deeply transform, 30% are redesigning key processes around it, and 37% are using it at a surface level with little or no change to how the work is done.

How the leading AI maturity frameworks compare

Several organisations publish competing models. They differ more in vocabulary than in substance, and once you line them up the shared assumptions become difficult to miss.

FrameworkStagesStage namesDimensionsCadence
GartnerFiveFoundational, Emerging, Operational, Scaled, TransformationalSeven pillars, including strategy, data, governance and operating modelPeriodic assessment
MicrosoftFiveInitial, Repeatable, Defined, Capable, EfficientFive pillars, including governance, technology and culturePeriodic assessment
Capability Maturity ModelFiveInitial, Repeatable, Defined, Managed, OptimisingSoftware process rigour, the original 1980s lineagePeriodic audit
GrowthNationScore out of 100Scored per team and organisation wideDerived from interviews across every team, tied to hours savedContinuous, updated as work is mapped

One correction is worth making, because it is widely repeated. A great many published summaries still list Gartner’s stages using an earlier set of names. The current Gartner model runs Foundational, Emerging, Operational, Scaled and Transformational, assessed across seven pillars. If you intend to benchmark against Gartner, use the version currently published.

What these frameworks share is more revealing than what separates them. All assess multiple dimensions, all describe a progression from ad hoc to optimised, and all are scored periodically, usually by an external assessor. That last shared assumption is the one worth questioning.

Where traditional AI maturity models fall short

None of this makes the established frameworks useless. Any of them beats having no measurement at all. But three limitations surface the moment an organisation tries to act on its score.

A single score hides the variance that matters

An enterprise does not have one maturity level. Engineering might sit at stage four while finance sits at stage one, and the organisational average describes neither. This is what the Deloitte finding exposes. Rating your strategy as highly prepared while rating data, infrastructure, risk and talent as far less prepared is not a contradiction. It is what a real organisation looks like from the inside, and a single averaged grade conceals it.

An annual grade is stale before it circulates

This is the flaw inherited directly from the Capability Maturity Model lineage. Periodic assessment made sense for software process audits. Applied to a capability that shifts every quarter, an annual grade describes an organisation that no longer exists by the time anyone reads the report. It also makes progress invisible between assessments, which is precisely when budget decisions get made.

Knowing your stage does not tell you what to do on Monday

This is the largest gap. McKinsey found that fundamentally redesigning workflows has the strongest link to EBIT impact of any factor studied, yet only 21% of adopters had fundamentally redesigned any workflow at all. MIT’s NANDA study reported that 95% of enterprise generative AI pilots produced no measurable return, although that figure rests on a small sample of 52 interviews and 153 survey responses and is contested on how narrowly it defines success. Sample caveats aside, the pattern is consistent. Organisations can name their stage. What they cannot name is which work to hand over first, and no stage label will tell them. AI agents for business deliver returns when they are pointed at work that has been properly identified beforehand.

How to measure AI maturity continuously

If the limitations are cadence and granularity, the answer is not a better set of stage names. It is a different way of producing the score. Four things change when maturity is measured continuously rather than annually.

  • Score from the work itself. Interview every team rather than a handful of leaders, so the result reflects what people actually do rather than what leadership believes they do.
  • Score per team as well as organisation wide, so variance stays visible instead of being averaged away.
  • Update as new work is mapped, so the score behaves like a live trend rather than an annual snapshot.
  • Tie the score to hours saved and cost avoided, so progress is evidenced rather than argued at renewal time.

That is the reasoning behind our AI scorecard: a grade out of 100 with a written narrative, a trend over time and a benchmark against peers, produced from what teams tell us rather than from a questionnaire someone fills in once a year. You can see how it works in more detail.

If you want to know where your organisation genuinely sits, start with a single team. We will map how that team actually works, show you the score that comes out of it, and hand you back a working agent built from what we find. No deck, and no six-month assessment programme.

Frequently asked questions

What is an AI maturity model?

An AI maturity model is a staged framework that rates how deeply and effectively an organisation uses AI. It scores capability across dimensions such as strategy, data, governance, talent and measured business value, then places the organisation at a defined stage. Most published models use four or five stages.

What are the five stages of AI maturity?

Stage names vary between frameworks, but the progression is consistent. Informal individual use with no coordination, then funded pilots that stay contained within one team, then at least one workflow genuinely running in production, then AI spreading across functions with measured returns, and finally an operating model that improves continuously.

What is the difference between an AI maturity model and an AI readiness assessment?

A maturity model is the framework. An assessment is the exercise of scoring yourself against it. Maturity generally describes how much AI is already embedded in how you work, while readiness describes whether the foundations exist to adopt it at all. Readiness is usually the earlier question of the two.

How often should you reassess AI maturity?

Most published frameworks assume an annual or occasional cadence, inherited from software process audits. That is too slow for a capability that changes quarterly. Continuous scoring shows movement between formal reviews, which is when most budget and renewal decisions are actually taken.

Which AI maturity model should we use?

Any of the established frameworks will give you a defensible baseline, and the differences between them are largely vocabulary. The more useful question is how the score gets produced. A model scored from what teams actually do, broken down per team and updated as work is mapped, tells you more than a better set of stage labels.