Adoption
A pilot that works and a programme that scales are different achievements, and the gap between them is almost never technical. Here are the five places we watch programmes stop.
Yaju Team · 28 April 2026
The pattern is common enough to be predictable. A pilot goes well, leadership is pleased, a rollout is announced, and eight months later the same three agents are running and nothing else has moved.
Nobody decides to stop. The programme simply stops progressing, and when you look at why, it is rarely the thing anyone expected.
Pilots are usually selected for demonstrability: a well-defined task, clean data, an enthusiastic team, a champion with time to spare.
Every one of those is absent in the second use case. The data is messy, the process is contested, the team is busy, and nobody has spare capacity to nurse it. The pilot proved the technology works in favourable conditions, which was never the open question.
The pilot had a champion. The rollout has a programme. A programme is not an owner, and agents without a named owner do not get maintained, so the second wave arrives with no one responsible for it in six months.
This is the single most reliable predictor we see. Not model choice, not budget, not sophistication: whether each agent has a person whose name is attached.
In a pilot, one person who understands the domain reads the output and judges it. That works for one agent and does not survive ten, because attention does not scale and disagreements about quality have no referee.
Without written criteria and automatic scoring, every new agent needs a human assessor indefinitely, and the programme is bounded by how many assessors you have. Teams that wrote evals during the pilot pass through this. Teams that did not discover that scaling means scaling review.
Pilot spend is small enough that nobody asks. Rollout spend is not, and the question arrives in the quarter when the programme has costs and few completed outcomes.
If spend is not attributed per agent, there is no answer. "AI cost us this much" invites a decision about AI in general, which is almost always a decision to slow down. "This agent costs this and produces this" invites a decision about that agent, which is the conversation you want.
The least technical reason and the most common. A pilot can succeed without anyone agreeing on the goal, because the goal was to see whether it worked. A programme cannot.
Teams that stalled generally could not answer what success would look like in a year. Teams that scaled had something specific, even if modest: this category of work, this much of it, at this quality, by then.
Attribution per agent and a named owner from the first day. Written evaluation criteria before building. Budgets that block rather than warn. A second use case chosen for representativeness rather than for demonstrability. And a stated goal specific enough to be wrong about.
None of that is sophisticated. All of it is easier to install before the sprawl than after.
The Agent Orchestration System pages cover ownership, attribution, budgets and evaluation. If you would rather discuss a specific programme than read about the pattern, book a demo and bring the constraints.