Models
The reflex when an agent underperforms is to reach for more capability. Sometimes that is right. More often the failure is somewhere the model cannot reach, and upgrading simply makes the same mistake more expensively.
Yaju Team · 2 April 2026
An agent produces a bad output. Someone proposes the larger model. It is the obvious move, it is easy to justify, and roughly half the time it changes nothing except the invoice.
The question worth asking first is what kind of failure this was.
Only one of these is a capability problem. The other three survive an upgrade intact.
Take the failing case and check, in order: was the correct source retrieved at all; how much irrelevant material came with it; does a written criterion exist that this output would have failed; and was the action within the boundary the agent should have had.
If the answer to the first is no, fix retrieval. If the second is large, fix chunking or trim tools. If no criterion exists, write one. If the boundary was wrong, that is a governance change, not a model change.
Only when all four check out is the model the remaining variable, and at that point the upgrade is a reasonable experiment with a clear hypothesis.
Because a model upgrade is a decision one person can make in a morning, and the alternatives are projects. Fixing retrieval means understanding a corpus. Writing criteria means agreeing what good means, which is an organisational conversation disguised as a technical one.
The upgrade is not chosen because it is likely to work. It is chosen because it is available.
A larger model changes the price of every call the agent makes, forever, including the calls that were already working. If it fixes ten percent of failures, you have paid for one hundred percent of traffic to improve a tenth of it.
This is exactly what evaluation-driven optimisation is for. When the Optimizer proposes a change, it is measured against your own evals, so the question stops being "does this feel better" and becomes "did the score move, and what did it cost".
They exist, and they have a recognisable shape: long multi-step reasoning where intermediate errors compound, tasks with unusual structure that smaller models mishandle consistently rather than occasionally, and work where the correct answer requires holding several constraints simultaneously.
If your failures are concentrated there and your retrieval, criteria and boundaries are sound, upgrade with confidence.
The versions page covers each Oran version and what it added. The Agent Orchestration System pages cover the evaluation and cost layers that turn a model decision into a measurable one.