Cost control
Model pricing is the part of agent spend that is easy to see and small to change. The expensive parts are the ones nobody attributes to anything: the retries, the oversized context, the agent that has been running since March and belongs to a team that dissolved in April.
Yaju Team · 14 May 2026
Ask an engineering leader what their agents cost and you will usually get a number from a provider dashboard. It is a real number. It is also the least interesting one available, because it answers a question nobody is actually asking.
The question is not "what did we spend on inference last month". It is "which of the things we are running is worth what it costs, and who decides". That question has a different shape, and a provider invoice cannot answer it.
None of these are exotic failures. They are the normal result of building quickly, and they compound quietly because no single one of them is large enough to trigger an investigation.
You cannot reduce a cost you cannot attribute. This sounds obvious and it is routinely skipped, because attribution is boring work and the dashboard already shows a number.
Oran 3.1 existed for exactly this reason. It attributed spend per agent, broke it down by user and by model, and reported it in currency rather than tokens. That is not a feature anyone demos well. It is the thing that makes every later conversation possible, because it turns "AI is expensive" into "this agent costs this much and produces this".
Until that shift happens, cost control is guesswork with a budget attached.
The second problem is timing. A monthly report tells you about an overrun after it has already happened, which is an accounting exercise rather than a control.
Circuit-breaker budgets close that gap: a limit per workspace and per team, a warning as the limit approaches, an automatic block when it is reached, and a forecast for the rest of the cycle. The difference between a warning and a block sounds procedural. It is not. One of them is information, and the other is a decision that gets taken whether or not anyone is looking at the dashboard on a Friday evening.
Here is where teams get hurt. Once spend is visible, the instinct is to cut it, and cutting it is easy: use a smaller model, shorten the prompt, remove a verification step. Spend falls immediately and quality falls slightly later, somewhere nobody is measuring.
This is why optimisation only works when it is measured against evaluation. Oran 3.4 reduces cost through model and prompt optimisation, tool trimming and hybrid agent conversion, and every one of those changes is scored against your own evals before it stands. Cheaper is only cheaper if the output still passes the checks you defined. Otherwise you have not saved money, you have deferred a cost into a place where it will be more expensive to find.
If you are starting from a provider invoice and nothing else, the order that has worked for the teams we see is fairly consistent.
Attribute spend per agent before anything else. Give every agent a named owner. Set a budget that blocks rather than warns. Write evals for your highest-volume agent, even crude ones. Only then start optimising, and check every change against those evals.
It is not a fast sequence. It is the one that ends with a number you can defend in a planning meeting, which is the actual goal.
The AI Spend Explorer will find overspend in a couple of minutes if you want a starting point rather than a project. If you want the whole picture, the Agent Orchestration System pages cover attribution, budgets, evaluation and optimisation as one system rather than four tools.