Skip to content

New — Introducing Yaju Software Factory — Try the tool for free

Explore now
Yaju AS
  • ProductProduct
  • SolutionsSolutions
  • ResourcesResources
  • BlogBlog
  • CompanyCompany
Sign in
Book demo

Platform

  • Agent Orchestration System

    Building agents is easy. Operating should be too.

  • Software Factory

    Turn your backlog into review-ready code

  • Agent Hub

    Browse, run, and share agents

Functionalities

  • AI Governance

    Policy enforced at the moment of action

  • AI Observability

    Observe and trust every agent

  • Token Monitoring

    Make every token count

  • Optimizer

    Same outcomes, lower cost

  • AI Spend Explorer

    Find overspend in two minutes

Product

  • MCP Gateway

  • CLI

  • Pricing

  • Versions

Featured

Choosing a model is an operations decision, not a benchmark decision

Use Cases

  • Cost Control

    Know what agents cost. Prove what they deliver.

  • AI Transformation

    Turn AI adoption into business transformation

Deployment

  • Credential Vault

    Org-level secrets, resolved at runtime

  • Self-hosted

    Run agents in your own environment

By Industry

  • IT & Developers

  • Financial Services

  • Public Sector

  • Engineering

  • Telecommunications

  • Healthcare and Life Sciences

  • Manufacturing

Featured

Running agents on hardware you own

Discover

  • Customer Stories

  • Partners

  • Yaju Labs

    Yaju Agent Systems research lab

For Learners

  • Versions

  • Agent Academy

Featured

Support triage is the best first agent most teams never build

Content

  • Blog

    The latest from Yaju, launches, and insights

Explorations

  • Future(s) of Work

    How will AI change the way we work?

  • Oran Models

    The generation teams run today

Initiatives

  • Scholars Program

    Finding the next generation of agent builders

  • Open Development Community

    Building agent tooling in the open

  • Catalyst Grants

    Backing ambitious work on agents

Featured

The future of work debate has an evidence problem
  • About

  • Careers

  • Newsroom

Yaju AS
Sign in
Book demo

Platform

  • Agent Orchestration System

    Building agents is easy. Operating should be too.

  • Software Factory

    Turn your backlog into review-ready code

  • Agent Hub

    Browse, run, and share agents


Functionalities

  • AI Governance

    Policy enforced at the moment of action

  • AI Observability

    Observe and trust every agent

  • Token Monitoring

    Make every token count

  • Optimizer

    Same outcomes, lower cost

  • AI Spend Explorer

    Find overspend in two minutes


Product

  • MCP Gateway

  • CLI

  • Pricing

  • Versions


Featured

Choosing a model is an operations decision, not a benchmark decision

Use Cases

  • Cost Control

    Know what agents cost. Prove what they deliver.

  • AI Transformation

    Turn AI adoption into business transformation


Deployment

  • Credential Vault

    Org-level secrets, resolved at runtime

  • Self-hosted

    Run agents in your own environment


By Industry

  • IT & Developers

  • Financial Services

  • Public Sector

  • Engineering

  • Telecommunications

  • Healthcare and Life Sciences

  • Manufacturing


Featured

Running agents on hardware you own

Discover

  • Customer Stories

  • Partners

  • Yaju Labs

    Yaju Agent Systems research lab


For Learners

  • Versions

  • Agent Academy


Featured

Support triage is the best first agent most teams never build

Content

  • Blog

    The latest from Yaju, launches, and insights


Explorations

  • Future(s) of Work

    How will AI change the way we work?

  • Oran Models

    The generation teams run today


Initiatives

  • Scholars Program

    Finding the next generation of agent builders

  • Open Development Community

    Building agent tooling in the open

  • Catalyst Grants

    Backing ambitious work on agents


Featured

The future of work debate has an evidence problem
  • About

  • Careers

  • Newsroom

Yaju AS
  • Agent Orchestration System

    Software Factory

    Agent Hub

  • MCP Gateway

    CLI

    Pricing

    Versions

  • AI Governance

    AI Observability

    Token Monitoring

    Optimizer

    AI Spend Explorer

  • Cost Control

    AI Transformation

    Solutions Overview

  • IT & Developers

    Financial Services

    Public Sector

    Engineering

    Telecommunications

    Healthcare and Life Sciences

    Manufacturing

  • Credential Vault

    Self-hosted

    Deployment Options

  • Blog

    Customer Stories

    Partners

    Agent Academy

    Yaju Labs

  • About

    Careers

    Newsroom

  • Legal Center

    Security

    Privacy Policy

    Terms of Use

Platform

  • Agent Orchestration System

  • Software Factory

  • Agent Hub

Product

  • MCP Gateway

  • CLI

  • Pricing

  • Versions

Functionalities

  • AI Governance

  • AI Observability

  • Token Monitoring

  • Optimizer

  • AI Spend Explorer

Solutions

  • Cost Control

  • AI Transformation

  • Solutions Overview

By Industry

  • IT & Developers

  • Financial Services

  • Public Sector

  • Engineering

  • Telecommunications

  • Healthcare and Life Sciences

  • Manufacturing

Deployment

  • Credential Vault

  • Self-hosted

  • Deployment Options

Resources

  • Blog

  • Customer Stories

  • Partners

  • Agent Academy

  • Yaju Labs

Company

  • About

  • Careers

  • Newsroom

Legal

  • Legal Center

  • Security

  • Privacy Policy

  • Terms of Use

Platform

  • Agent Orchestration System

  • Software Factory

  • Agent Hub

Product

  • MCP Gateway

  • CLI

  • Pricing

  • Versions

Functionalities

  • AI Governance

  • AI Observability

  • Token Monitoring

  • Optimizer

  • AI Spend Explorer

Solutions

  • Cost Control

  • AI Transformation

  • Solutions Overview

By Industry

  • IT & Developers

  • Financial Services

  • Public Sector

  • Engineering

  • Telecommunications

  • Healthcare and Life Sciences

  • Manufacturing

Deployment

  • Credential Vault

  • Self-hosted

  • Deployment Options

Resources

  • Blog

  • Customer Stories

  • Partners

  • Agent Academy

  • Yaju Labs

Company

  • About

  • Careers

  • Newsroom

Legal

  • Legal Center

  • Security

  • Privacy Policy

  • Terms of Use

LinkedInInstagramEmail

Yaju AS ©2026

  • English

Engineering

The optimisations worth doing, and the ones that only look impressive

There is a long list of techniques for making model inference faster and cheaper. Most of them are real. Very few of them are the reason your agent is slow, and picking from the list before measuring is how teams spend a quarter on a two percent gain.

Yaju Team · 2 June 2026

Optimisation work has a gravitational pull towards the technically interesting. Inference-level techniques are genuinely clever, well documented, and satisfying to implement.

They are also, in most deployments, somewhere below fifth on the list of things that would help.

The order we would actually work in

First, remove work that does not need doing. Unused tool definitions occupying context on every call. Prompt sections added for an edge case two years ago. Retrieval returning ten passages where three would do. This is pure waste, and removing it costs nothing in quality because it was contributing nothing.

Second, fix retries. A run that fails and restarts pays twice for one outcome. Retries are usually the largest single addressable inefficiency and they are invisible unless measured apart from successes.

Third, parallelise what is sequential but independent. Most agent code runs steps in series because that is how it reads, not because it must.

Fourth, right-size the model per step. Classification and extraction steps rarely need frontier capability, and running it on them is a permanent cost for no benefit.

Fifth, and only now, inference-level techniques.

Why the order matters more than the techniques

The first four are cheap, low-risk and frequently large. The fifth is sophisticated, and its impact is bounded by how much of your elapsed time was inference in the first place.

If inference is a third of your run time, a twenty percent inference improvement is a seven percent run improvement. That is worth having, and it is not worth having before the retries that are costing you a hundred percent on a share of your traffic.

Measure before choosing

Everything above assumes you know the breakdown. Most teams do not, because standard instrumentation reports per run rather than per step.

Per-step measurement is the prerequisite for every decision on this list. Without it, optimisation is a matter of taste, and taste reliably selects the interesting option.

Every optimisation is a hypothesis about quality

The reason we insist on evaluation before optimisation is that all of these changes can reduce quality, and most of the reductions are not visible immediately.

A smaller model on a step that turned out to need capability. A trimmed prompt that removed the paragraph handling a real edge case. Fewer retrieved passages, which is fine until the question needs the fourth one.

Scoring each change against criteria you defined beforehand is what separates a saving from a deferral. This is exactly what the Optimizer does, and it is why it refuses to present a change as an improvement on cost alone.

The unglamorous conclusion

The highest-return optimisation work in most agent systems is deleting things: unused tools, accumulated prompt text, unnecessary retrieval, steps that should not have been sequential.

None of that is interesting to write about, which is presumably why so much more is written about the alternatives.

Next

The Optimizer and Token Monitoring pages cover per-step measurement and evaluation-measured optimisation. The AI Observability pages cover instrumentation inside a run.

Loading...
The optimisations worth doing, and the ones that only look impressive | Yaju