Adoption
Most maturity models describe increasing sophistication. This one describes increasing ability to answer questions about your own system, because that is what separates organisations that scale from ones that stall.
Yaju Team · 7 May 2026
Maturity models are usually a ladder of adopted technology. We find a more useful version is a ladder of answerable questions, because capability without visibility produces the failures we see most often.
Here is the ladder as we observe it.
An agent exists and does something useful. Usually built by one enthusiastic person, evaluated by reading the output, running on credentials that live in a configuration file somewhere.
The question you cannot answer: what would happen if this person left.
Almost every organisation is here at some point, and there is nothing wrong with being here. The failure is staying here while the count grows.
There is an inventory. Every agent has a named owner. You can list what is running and say who is responsible for each.
The question you can now answer: what are we running, and who do I ask about it.
This is the single highest-value step on the ladder and the one most often skipped, because it produces no new capability. It is also the cheapest to do early and the most painful to retrofit.
Spend is attributed per agent, broken down by user and model, reported in currency. Budgets block rather than warn.
The question you can now answer: is this agent worth what it costs.
Before this level, cost conversations are about AI in general, which is a conversation that ends in slowing down. After it, they are about specific agents, which is a conversation that ends in decisions.
Evaluation criteria are written down and scored automatically on every run, not sampled. A failed check halts the run rather than producing a report someone reads later.
The question you can now answer: has this quietly got worse.
This is the level that makes scale possible, because it is what turns review from uniform suspicion into triage. Below it, every new agent adds a permanent human review cost.
Optimisation is measured against your own evals. Agents are shared through a catalogue with versioning and rollback. Policy is enforced at the point of action rather than described.
The question you can now answer: what happens if we change this, and can we undo it.
Organisations try to climb it in the order of perceived sophistication, which puts evaluation and optimisation early because they sound advanced, and inventory late because it sounds administrative.
That order does not work. Evaluation without ownership means nobody acts on a failing score. Optimisation without attribution means you cannot tell whether the saving was real. The unglamorous levels are load-bearing.
In our experience: capable of building agents at a level four standard, and operating them at level one or two. The gap between what teams can build and what they can account for is the defining problem of the current moment, and it is the reason this platform exists.
The Agent Orchestration System pages cover ownership, attribution, budgets, evaluation and the Agent Hub. If you want to work out which level you are at, book a demo and we will go through it against your actual deployment.