A maturity model has five tiers. A maturity index has one number, and most G&A leaders cannot tell you theirs from last quarter. That gap is the whole subject. An AI maturity framework is a shared scoreboard, not a strategy or a roadmap, and its only job is to tell a Finance or People Ops leader two things they otherwise guess at: where the function sits today, and whether it has moved since the last review. The tiers, the radar charts, the named stages are presentation. The number is the point.
The number most G&A leaders cannot produce
Ask a product leader how their AI work is going and they point at shipped features. Ask a G&A leader the same question and you get anecdote, because Finance and People Ops rarely produce a visible artefact. Nothing rolls off a line. So the only signal of progress is whoever talks loudest in the leadership meeting, and loud is not the same as moved.
That is what a maturity index fixes, and why it matters more in G&A than anywhere else. It replaces the anecdote with a tracked figure on a fixed scale. Not to impress the board. To settle, in one number, an argument the function has every quarter about whether it is actually getting anywhere.
The confusion worth clearing first is readiness versus maturity, because the two get sold as the same product and are not. Readiness is a precondition: do we have the data, the skills, the callable tools to start. Maturity is an outcome: how much of the real work has moved. A team can score high on readiness and low on maturity, and often does, because the ingredients were bought and the meal was never cooked. The bridge between them is operating cadence, and cadence is the thing every framework underweights. If you want the capability side scored honestly, the five pillars of AI readiness is the precondition check that sits under everything here.
The five tiers every framework collapses to
The published frameworks differ in branding and stop at different places, but strip the diagrams and they share five tiers. The names below are the ones we use; the substance is common to all of them.
| Tier | What it looks like | Typical G&A signal |
|---|---|---|
| 1. Experimenting | Individuals using ChatGPT and Copilot ad hoc | Slack channel full of prompts, no shared library |
| 2. Piloting | Funded projects, one workflow at a time | A named "AI in Finance" pilot, owned by one person |
| 3. Scaling | The same workflow used by a whole team | Every recruiter uses the same JD assistant |
| 4. Integrating | The AI step is part of the official workflow, with governance | Month-end close has an AI review step in the SOP |
| 5. Operating | Workflows are measured, owned and improved on a cadence | Weekly review of three AI-assisted workflows, with metrics |
Tiers 1 and 2 produce the most internal noise: the demos, the town-hall slides, the Slack channel of clever prompts. Tiers 4 and 5 produce the actual operating value. The gap between them is where every framework quietly hides the difficult part, because the difficult part is unglamorous and slow. For the deeper version of this ladder and why each rung depends on the one below it, see the AI operating ladder.
The AI maturity frameworks you will actually be handed
You will be given one of these in the next twelve months, usually by someone with a reason to prefer it. Knowing what each is good for saves a quarter of arguing about the wrong axis.
Gartner-style tier models. Five tiers, useful as board-level vocabulary. Weakest on the integration tier; they treat scaling as the finish line when scaling is the middle. Best when the audience is a CFO or CHRO who wants a familiar diagram to nod at.
MIT Sloan stages. Stronger on the organisational side, especially leadership behaviour and data literacy. Less prescriptive about the operating cadence that holds maturity in place once you have it. Best for a culture-and-capability narrative.
Build-Operate-Transfer variants. Strong on the funding model and the hand-off from consultancy to in-house team. Weakest on what to actually do on a Tuesday morning. Best when the spend is large and the real question is governance.
Deepgrain's operating ladder. Built specifically for the integration tier. It names the bridge from pilot to operating as the work, not the gap you are meant to cross on faith. Best once the leader has done the readiness exercise and now needs to move the line.
Home-grown internal frameworks. Often the most accurate for a specific company, because they encode local constraints nobody outside would know. Almost always missing the operating-cadence section, because the people who design the framework are not the people who run the weekly review.
A working rule: use a public framework as your external vocabulary and a working ladder to run the work, and never put both in the same document. Board frameworks are about legitimacy. Team frameworks are about the next move. They read as contradictory to anyone forced to hold both at once.
An AI maturity index that survives a year
A maturity index is only worth running if the same number means the same thing in Q1 and Q4. That rules out most radar-chart versions, which quietly redefine their axes every time leadership changes and then wonder why the trend line is meaningless. The shape we use with G&A teams is five categories with hard caps, so a team cannot fake one by neglecting another.
- Workflow coverage (0 to 30). Percentage of named G&A workflows with at least one AI step running in production for ninety days. Not a demo. Production.
- Adoption depth (0 to 20). Percentage of staff with measured weekly use of AI in their core workflow. Self-report does not count.
- Time-back (0 to 20). Aggregate time saved per quarter against a documented baseline, normalised against headcount.
- Governance (0 to 15). Percentage of in-production AI steps with a current risk review and a named owner.
- Operating cadence (0 to 15). Whether a weekly review of the AI-assisted workflow stack actually happens and produces a change log.
Total: 0 to 100. The caps are the point. Coverage without governance plateaus at 30. Adoption without cadence plateaus at 50. The score forces the conversation back to the integration tier every time someone tries to buy their way out of it with another pilot.
The reason most indexes die in month four is not the categories. It is the compile. If assembling the number takes a week of chasing people, whoever owns it will stop after the second quarter, and a maturity index nobody updates is worse than none because it lies with authority. So pull the inputs from places that do not depend on goodwill: adoption depth from the tool vendor's usage logs, never a survey; coverage and governance from the workflow owner's change log; time-back from a baseline you wrote down before you started. Then wire the assembly into a scheduled job. A small n8n workflow, around twenty pounds per builder seat per month and self-hostable if the data is sensitive, can pull the usage logs, read the change logs and hand you a draft score every Monday. Where the change log is free text, summarise it with a model, not a regex, because a regex will miss every phrasing you did not anticipate and quietly undercount the work.
Where Finance and People Ops actually stall
The two functions fail in opposite directions, and the fix is structural in both, never motivational. Nobody stalls because they lack enthusiasm.
Where Finance stalls
Strong on governance, weak on operating cadence
Scores 40 to 55 for two years and calls it caution
Month-end close hoards all the pilot energy
The risk function formalises before anything ships
The fix: put the AI stack on the controls-review agenda
Where People Ops stalls
Strong on adoption, weak on workflow coverage
Scores 45 to 60 because everyone uses AI for something
No single workflow rewired end to end
A fourth pet workflow sneaks in every quarter
The fix: name three workflows, refuse to count a fourth
Opposite failure modes, the same remedy: give the AI stack the operating dignity of the numbers next to it.
Finance formalises early, which is healthy, but the operating side never gets the same airtime as the controls review. Put the AI-assisted workflow stack on that same weekly agenda and within a quarter the cadence score moves; within two, the time-back follows. The other Finance trap is treating month-end close as the only target worth having. It is the most visible workflow, so it absorbs all the pilot energy, and the hardest to integrate safely, so it absorbs all the governance overhead. Pick a quieter workflow first: procurement triage, accruals support, vendor onboarding. Maturity rises faster on a dull workflow that actually moves than on a glamorous one that never clears governance.
People Ops has the reverse problem. Everyone is using AI for something and no workflow has been rewired end to end. The discipline is not picking the right three workflows. Candidate sourcing, review summarisation and policy Q&A are the usual three because they have clean inputs, clean outputs and clear owners. The discipline is refusing to let a fourth in this quarter, however reasonable the fourth sounds in the moment.
Moving from pilot to integration
Every framework agrees the pilot-to-integration jump is the hard one. None of them are specific about how, because the how is boring and boring does not sell a deck. The demo worked. The rollout didn't. That is the whole failure in four words, and it repeats because the integration tier needs four things a pilot is never funded to have. Run any workflow you think is integrated through this before you write it down as such.
A team with those four things moves from tier 3 to tier 4 within two quarters, whichever published framework it cites. A team without them stays at tier 3 forever, however much pilot funding keeps arriving. The tooling is the easy part and the cheap part. What is scarce is a permanent owner whose performance review actually contains the workflow's numbers, and a funding line that treats operate-and-improve as the work rather than the afterthought. For the deeper anatomy of why demos clear and rollouts do not, see why AI pilots stall at production.
The honest limits of any maturity index
Three things worth knowing before you commission a maturity assessment, because a number wielded badly does more damage than no number.
First, the index is a lagging indicator. By the time it moves, the work has been happening for a quarter. Use it to confirm a direction, never to choose one. Second, it rewards only the workflows you can name. Anything invisible to the framework is invisible to the score, and over time invisible to the strategy and the budget, which is how the transit licence above survived for years. Audit the named-workflow list every six months and go looking for what is missing. Third, the index does not score judgement. A team can have an AI step in every workflow and still make worse calls than the team next to it. Maturity is necessary and not sufficient.
The last limit is the one that keeps me honest, so I will give it a real number rather than a reassurance.
Of the eleven functions we scored on a hundred-point readiness index, not one came in above seventy. I would rather you knew that than pretend the base rate is good.
That is not a reason to skip the exercise. It is the reason to run it and publish the number internally, ugly and all, because a function that knows it sits at sixty-something will do the boring integration work, and a function that assumes it is doing fine will keep buying pilots. The score is a mirror, not a trophy.
What to do with this on Monday
If you want a maturity exercise to be more than a slide, the entire programme fits in four moves.
- 01Day 1Pick one external framework
Choose a public tier model as your board vocabulary. Conservative board, use a familiar Gartner-style one; academic board, use the MIT stages.
- 02Week 1Baseline the index honestly
Score the five-category index once, privately. Do not publish the number outside the team for two quarters.
- 03Weeks 1-12Name three workflows
Pick three you will move from tier 3 to tier 4 this quarter. Put them on a fifteen-minute weekly review with an owner each.
- 04Day 90Score again at ninety days
Re-run the index. The number tells you whether the cadence was real or theatrical. Nothing else does.
That is it. Anyone offering more than this in the first quarter is selling transformation theatre, and G&A budgets have paid for enough of that. If you want a scored starting point before you build your own index, the free Readiness Assessment takes about ten minutes and gives you the honest baseline this whole exercise depends on. For a broader diagnostic run against the whole function, how to diagnose an organisation in 30 days is the fuller version, and the maturity work here sits inside the wider AI operating system picture rather than standing alone.
Common questions
- What is an AI maturity index?
- An AI maturity index is a number, not a description. The model tells you the tiers; the index tells you where you sit on a 0 to 100 scale and whether that number moved since last quarter. Most teams confuse the two: they can recite the five tiers from a slide but have never scored themselves. If you cannot produce last quarter's number from memory, you have a maturity model, not a maturity index.
- Which AI maturity framework is best for Finance and People Ops?
- Check who published it before you check what it says. Frameworks from companies selling AI tooling over-index on adoption, tiers 1 and 2, because that is what their product moves. Frameworks from advisers with nothing to sell are more honest about the pilot-to-integration gap. Use a public framework for board vocabulary and a working index to run the actual work. Never mix the two in one document.
- How is AI maturity different from AI readiness?
- Readiness is a precondition. Maturity is an outcome. Readiness asks whether you have the data, skills and tools to start. Maturity asks how much of the real work has actually moved. A team can be highly ready and barely mature. The bridge between the two is operating cadence: weekly reviews, owned workflows and a clear next bet.
- How do you measure organisational AI maturity?
- Pull the numbers from three places and no more: usage logs from the tool vendor, never self-report; a quarterly survey capped at five questions; and the workflow owner's own change log. Wire the assembly into a scheduled job so it takes an afternoon, not a week. Anything slower than that will not survive a second quarter, because whoever owns the measurement quietly stops doing it.
- Why do G&A teams stall between pilot and integration?
- Pilots are funded as projects and integration is funded as overhead, so the economics punish the work that actually moves maturity. The demo lands, the budget closes, the owner moves on. The fix is to fund three months of operate-and-improve at the same time as the pilot, with the same named owner. Without that, every pilot becomes a one-off.
Not sure where your function stands yet?Take the Readiness Assessment→
When reading turns into doing
The Grain Audit maps one People Ops process end to end, ranks the highest-return automations, and hands you a 90-day plan you keep whether or not we work together.
Two weeks. £2,000, credited in full against a programme. Three slots a month.
Book a Grain AuditIf this resonated, there's more.
Subscribe to receive new Intelligence pieces as they're published. No noise, just the work.
By subscribing you agree to our Privacy Policy. Unsubscribe any time.



