Use case
Cloud teams generate enormous volumes of signal and spend a surprising share of their week assembling context rather than acting on it. That gap is where agents belong, and it is not the part anyone demos.
Yaju Team · 20 May 2026
Cloud operations looks made for automation: huge event volume, repetitive investigation, expertise concentrated in few people, and a backlog of improvements nobody reaches.
The instinct is to point an agent at the alert stream and let it decide. That is the application most likely to fail, for a reason worth stating before anything else.
An alert means different things depending on what happened upstream, what is deploying, what maintenance is scheduled, and whether this pattern resembled something that turned out to be nothing last quarter.
Most of that context is not in the alert. It is in other systems and in engineers' heads. An agent deciding without it will be confidently wrong, and two confident mistakes are enough to end the trust that made the project possible.
There is also no natural review step. By the time a decision could be reviewed, it has already had consequences.
Context assembly before a human decides. Recent changes on the affected path, related past incidents, the current deployment state, the relevant runbook section. This removes the ten minutes of hunting that precedes every real diagnosis, on every incident, forever.
Post-incident timelines, which exist across several systems and get written late and badly.
Cost and capacity review, which is mechanical, continuous and permanently deprioritised.
Change validation against policy and past incidents, flagging rather than blocking.
An agent that can read infrastructure state and cannot change it is a fundamentally different risk from one that can do both. That boundary belongs at the point where the action occurs, not in the prompt.
Credentials resolve at runtime from the org vault, so no infrastructure credential lives in an agent's configuration and revocation is a single action.
Cloud teams are usually good at this and forget to apply it to the agent itself. Cost attributed per agent. Latency percentiles rather than means, because agent runs are heavily skewed. Retries counted separately from successes, since they are the same work billed twice.
Evals on every run rather than a periodic audit, because the failure you care about is the one that appeared gradually.
Context assembly on one class of incident. Named owner on the operations side. Evaluation criteria built from last quarter's real incidents. Read-only to begin with, and a budget that blocks.
Widen when the scores are boring, not when the demo is impressive.
The AI Observability and Token Monitoring pages cover instrumentation. The Agent Orchestration System pages cover policy, credentials and evaluation.