A company we talked to this summer runs 140 internal agents, built by people who genuinely know what they're doing. Nobody there could tell you which one helped. That conversation is the reason we drew a maturity matrix at all.
What we expected to find
Across dozens of companies now, we've asked the same questions about AI adoption: who owns it, what the plan is, who runs it, how far it has spread. We expected the answers to sort companies by sophistication, with the most enthusiastic teams at the top and everyone else below.
That's not what we found. As it turns out, the companies with the most impressive individual work often sat the most stuck.
What the research actually said
What lets a company improve, on AI adoption, sophistication, any of it, turns out to be less about how skilled individual employees are with AI and more about how the company manages AI ownership. A company with a mediocre AI toolset and a real owner keeps getting better. A company with brilliant individual work and no owner stays put, and the brilliant work retires with whoever did it. Once we saw that, the rest of the matrix fell out of the research on its own.
Other things we heard often enough to write down
Every one of these looks like a tooling problem. None of them are.
- Top performers are the top AI users, at every company we've asked. Whether AI is the cause is a fair question, and several operators have told us good hires are good with any tool.
- Skill lists go stale everywhere. At one company the prompt is on v5 and the team is still running v1. Nine people were solving the same problem, each slightly differently.
- Lunch-and-learns are hard to make stick, because the people who volunteer to present are the people already using AI.
- One rep ran a homegrown agent for two months before anyone found out about it. Self-reporting is most companies' entire detection system.
- People fork the shared skill to make it their own and never come back. The official version keeps aging while private copies multiply.
- Cost data will show you who spends the most on tokens. That isn't the same list as your best AI users, and most companies have no way to tell the two apart yet.
One more pattern: engineering teams are ahead of GTM teams on internal AI at nearly every company we've met. If you carry this for a GTM org, you sit earlier in a curve engineering started climbing first, which is not the same as behind.
Where the matrix lives
The framework itself sits on its own page rather than in this post: the AI Maturity Matrix. Five levels, the four questions that place you on them, and six dimensions read level by level.
L0 Wild West → L1 Prioritized → L2 Adopted → L3 Operationalized → L4 Instrumented
Ownership defines the level. The other five dimensions are the ones we measure, and they move at their own speed once someone owns the work.
Answer the four questions in order and the first no marks where you land. None of them ask how good your prompts are.
Behind the scenes: the agents that do the legwork
We didn't assemble any of this by hand. Two agents keep it current, and both draft while we decide.
The classifier runs every day against our recorded calls. It skips the internal ones, reads each call against the four breakpoints, and proposes a placement with the evidence quoted back at us. Every draft carries three verdicts: the level, the ICP fit, and the persona we're talking to.
It can also propose edits to the matrix itself, exact line out and exact line in, capped at three per call. Most days it proposes none. An agent that always finds something to say stops getting read.
The prep agent wakes up at 7am and reads the calendar. Today's external meetings get a pre-call profile: a provisional level, our best guess at who really owns AI adoption there, the public AI signals, and discovery questions aimed at that placement.
Tomorrow's meetings get a drafted email. It leaves companies we've already met alone, since the classifier owns those.
What keeps them honest
Both run as managed skills, each with an owner, a version, and a tested release. Changes get reviewed before they replace whatever runs today.
The guardrails do the real work here. Neither agent can touch a level definition or a breakpoint, and neither can put a customer's name on a public page. Those limits hold whether or not the model cooperates.
Those limits catch problems before they happen rather than after, which is the only reason we'd point a drafting agent at a public page.
And when something in the matrix lacks a call behind it, we go ask, because recorded calls only sample the world.
We keep revising this as the research accumulates, and if your company reads differently than the matrix predicts I'd genuinely like to hear about it at [email protected]. That's how the levels got their shape in the first place.
Run the four questions on your own company this week. The answer takes about twenty minutes and points at the single thing to work on next.
Then keep the answer somewhere you'll look again, because it moves. Ownership changes hands, plans get funded and quietly unfunded, and the six dimensions drift apart at different speeds. The companies that improve fastest watch that drift as it happens, instead of rediscovering it two quarters later.