Models
Once an agent system is broken into steps, most of those steps turn out to be classification, extraction and routing. Those do not need a frontier model, and running one on them is the most common avoidable cost we see.
Yaju Team · 10 June 2026
Look inside a single agent run and count what actually happens. Deciding which category a request falls into. Extracting three fields from a document. Judging whether a retrieved passage is relevant. Formatting a result. Choosing which tool to call next.
Then the reasoning step, which is usually one of them.
Most systems run the same large model on all of it, because that is how the system was built during the phase when getting it working at all was the goal.
A large model on a classification step costs several times what a small one would, produces the same answer, and does so on every single run forever.
The individual difference is trivial, which is exactly why it persists. Nobody reviews a step that works. The aggregate across a year of traffic is not trivial at all.
Classification into a known set of categories. Extraction of structured fields from semi-structured text. Relevance judgement over a shortlist. Reformatting. Short summarisation where the source is already focused.
These share a property: the output space is constrained. When the correct answer is one of five categories or a field from a document, capability beyond a threshold adds nothing because there is nowhere for it to go.
Multi-step reasoning where intermediate errors compound. Tasks requiring several constraints held simultaneously. Work with unusual structure that smaller models mishandle consistently rather than occasionally. Generation where quality is genuinely open-ended.
The distinction is not task difficulty in a human sense. It is whether the output space is bounded.
Instrument per step, not per run. Most systems report cost and latency for the whole agent, which averages an expensive reasoning step with a dozen trivial ones and hides both.
Once steps are visible, the experiment is direct: swap one step to a smaller model and rerun your evaluation set. If the scores hold, the saving is real and permanent. If they drop, you have learned something specific about that step rather than about model size in general.
This is exactly what the Optimizer does: model and prompt optimisation, tool trimming and hybrid conversion, each measured against your own evals before it stands.
The shape that tends to result is a small model handling the majority of steps, a larger one on the two or three that need it, and a routing decision between them that is itself cheap.
Teams arrive at this by measurement rather than design, which is the right order. Designing a hybrid system before knowing which steps need capability produces a complicated system optimised for a guess.
Cost reduction without evaluation is not a saving, it is a deferral. A cheaper step that is subtly worse produces failures further downstream, in a place where they are more expensive to diagnose than the money saved.
This is why every proposed optimisation should be scored against criteria you wrote before you started optimising.
The Optimizer and Token Monitoring pages cover per-step instrumentation and evaluation-measured optimisation. The versions page covers each Oran version and what it added.