Engineering
Agent traffic does not behave like the request patterns most capacity planning assumes. It is spiky, it grows with confidence rather than with headcount, and a single run can consume what a hundred ordinary requests would.
Yaju Team · 29 April 2026
Infrastructure planning generally assumes something about the shape of demand: it correlates with users, it varies predictably across the day, and one request costs roughly what another does.
Agent workloads break all three assumptions, and teams tend to discover this in the quarter when the forecast stops being useful.
Traditional capacity scales with the number of people using a system. Agent usage scales with how much a team trusts the agents, which is a different variable entirely and moves in steps rather than smoothly.
A team that trusts one agent runs it occasionally. The same team, three months later, having watched it work, routes a whole category of work through it. Nothing about the organisation changed. The load multiplied.
This is why capacity comfortable in month one is frequently tight in month six, and why a forecast built on user counts will be wrong in a direction that is expensive.
A single agent run can involve retrieval, several model calls, multiple tool invocations and a verification step. The variance between a simple run and a complex one is enormous, and averages conceal it badly.
Planning on mean cost per run underestimates the tail, and the tail is where budgets are actually consumed. A small proportion of runs routinely account for a large proportion of spend.
An agent that fails partway and restarts pays for the whole run twice and produces one outcome. Unless successes and retries are measured separately, this appears as higher average cost rather than as a fixable failure.
In most deployments we have looked at, retries are the single largest addressable waste, and they are also the easiest to miss because nothing errors visibly.
Agent work arrives in bursts: a batch job kicks off, an incident triggers twenty parallel investigations, a scheduled run fires across teams simultaneously.
Capacity sized for the average will queue during the bursts, and queueing on an interactive agent is indistinguishable from the feature being broken.
Cost and latency per agent, not per workspace. Retries separately from successes. Percentiles rather than means, because the mean of a heavy-tailed distribution is not a useful planning number. Step counts per run, so wandering agents are visible before the budget catches them.
With those four, a forecast becomes possible. Without them, capacity planning is extrapolation from a number that conceals its own variance.
Circuit-breaker budgets are a capacity control as much as a financial one. A hard limit per workspace and team, with a warning as it approaches and a block when reached, bounds the worst case in a way that a monthly report cannot.
This matters most in self-hosted deployments, where exceeding capacity is not somebody else's scaling problem.
The AI Observability and Token Monitoring pages cover instrumentation, and the deployment pages cover self-hosted and air-gapped capacity considerations.