Engineering
Teams set cost budgets and quality criteria and then let latency be whatever it turns out to be. It is the third axis, it trades against the other two constantly, and leaving it unspecified means it gets decided by accident.
Yaju Team · 25 March 2026
Ask a team what their agent should cost and you will get a number. Ask what good output looks like and you will usually get criteria. Ask how long it is allowed to take and you will get a pause.
That pause is where a lot of subsequent frustration originates, because latency is not a property that emerges. It is the result of decisions taken for other reasons.
Retrieve more candidates and rerank them: better grounding, more time. Add a verification step: fewer confident errors, more time. Give the agent another tool: more capability, more context to process on every call.
None of these are wrong. They are trades, and a trade made without a budget on one side is not a trade, it is a drift. Systems get slower one reasonable decision at a time until somebody notices that a thing which used to be interactive now is not.
Rarely where people assume. In the agent runs we have looked at, the distribution is fairly consistent: a minority in model inference, a substantial share in retrieval and tool calls, and a surprisingly large share in sequential steps that could have run in parallel but were written as a chain.
The last of those is usually the cheapest to fix and the least examined, because it looks like correct code rather than a performance problem.
Retries are the other quiet consumer. An agent that fails and restarts pays the whole latency twice for one outcome, and aggregate figures hide it unless successes and retries are measured separately.
A latency budget only means something in context. An agent a person is waiting on has a budget measured in seconds, and exceeding it makes the feature feel broken regardless of output quality.
An agent running overnight on a queue can take twenty minutes and nobody cares, which means it can afford verification steps, wider retrieval and slower models that an interactive one cannot.
Teams that classify their agents this way make much better decisions, because the same technique is correct in one context and wrong in the other.
This is why evals in Oran score latency alongside behavioural, structural and cost criteria rather than treating it as an operational metric read separately.
A change that improves accuracy and doubles response time is not automatically an improvement, and a system that only scores correctness will report it as one. Scoring all four together forces the trade to be explicit at the moment it is made, which is the only point where it is cheap to reconsider.
Write down the budget before building, even roughly, and classify the agent as interactive or background. Measure retrieval, tool calls and inference separately so that the budget can be attributed rather than merely exceeded. Count retries apart from successes. Then treat a latency regression the way you would treat a quality regression, which is to say as something that fails a check rather than something that is noticed in a meeting.
The Agent Orchestration System pages cover the evaluation layer and how cost, quality and latency are scored together. The observability pages cover where time is spent inside a run.