Engineering
There is a long list of techniques for making model inference faster and cheaper. Most of them are real. Very few of them are the reason your agent is slow, and picking from the list before measuring is how teams spend a quarter on a two percent gain.
Yaju Team · 2 June 2026
Optimisation work has a gravitational pull towards the technically interesting. Inference-level techniques are genuinely clever, well documented, and satisfying to implement.
They are also, in most deployments, somewhere below fifth on the list of things that would help.
First, remove work that does not need doing. Unused tool definitions occupying context on every call. Prompt sections added for an edge case two years ago. Retrieval returning ten passages where three would do. This is pure waste, and removing it costs nothing in quality because it was contributing nothing.
Second, fix retries. A run that fails and restarts pays twice for one outcome. Retries are usually the largest single addressable inefficiency and they are invisible unless measured apart from successes.
Third, parallelise what is sequential but independent. Most agent code runs steps in series because that is how it reads, not because it must.
Fourth, right-size the model per step. Classification and extraction steps rarely need frontier capability, and running it on them is a permanent cost for no benefit.
Fifth, and only now, inference-level techniques.
The first four are cheap, low-risk and frequently large. The fifth is sophisticated, and its impact is bounded by how much of your elapsed time was inference in the first place.
If inference is a third of your run time, a twenty percent inference improvement is a seven percent run improvement. That is worth having, and it is not worth having before the retries that are costing you a hundred percent on a share of your traffic.
Everything above assumes you know the breakdown. Most teams do not, because standard instrumentation reports per run rather than per step.
Per-step measurement is the prerequisite for every decision on this list. Without it, optimisation is a matter of taste, and taste reliably selects the interesting option.
The reason we insist on evaluation before optimisation is that all of these changes can reduce quality, and most of the reductions are not visible immediately.
A smaller model on a step that turned out to need capability. A trimmed prompt that removed the paragraph handling a real edge case. Fewer retrieved passages, which is fine until the question needs the fourth one.
Scoring each change against criteria you defined beforehand is what separates a saving from a deferral. This is exactly what the Optimizer does, and it is why it refuses to present a change as an improvement on cost alone.
The highest-return optimisation work in most agent systems is deleting things: unused tools, accumulated prompt text, unnecessary retrieval, steps that should not have been sequential.
None of that is interesting to write about, which is presumably why so much more is written about the alternatives.
The Optimizer and Token Monitoring pages cover per-step measurement and evaluation-measured optimisation. The AI Observability pages cover instrumentation inside a run.