Skip to content

New — Introducing Yaju Software Factory — Try the tool for free

Explore now
Yaju AS
  • ProductProduct
  • SolutionsSolutions
  • ResourcesResources
  • BlogBlog
  • CompanyCompany
Sign in
Book demo

Platform

  • Agent Orchestration System

    Building agents is easy. Operating should be too.

  • Software Factory

    Turn your backlog into review-ready code

  • Agent Hub

    Browse, run, and share agents

Functionalities

  • AI Governance

    Policy enforced at the moment of action

  • AI Observability

    Observe and trust every agent

  • Token Monitoring

    Make every token count

  • Optimizer

    Same outcomes, lower cost

  • AI Spend Explorer

    Find overspend in two minutes

Product

  • MCP Gateway

  • CLI

  • Pricing

  • Versions

Featured

Choosing a model is an operations decision, not a benchmark decision

Use Cases

  • Cost Control

    Know what agents cost. Prove what they deliver.

  • AI Transformation

    Turn AI adoption into business transformation

Deployment

  • Credential Vault

    Org-level secrets, resolved at runtime

  • Self-hosted

    Run agents in your own environment

By Industry

  • IT & Developers

  • Financial Services

  • Public Sector

  • Engineering

  • Telecommunications

  • Healthcare and Life Sciences

  • Manufacturing

Featured

Running agents on hardware you own

Discover

  • Customer Stories

  • Partners

  • Yaju Labs

    Yaju Agent Systems research lab

For Learners

  • Versions

  • Agent Academy

Featured

Support triage is the best first agent most teams never build

Content

  • Blog

    The latest from Yaju, launches, and insights

Explorations

  • Future(s) of Work

    How will AI change the way we work?

  • Oran Models

    The generation teams run today

Initiatives

  • Scholars Program

    Finding the next generation of agent builders

  • Open Development Community

    Building agent tooling in the open

  • Catalyst Grants

    Backing ambitious work on agents

Featured

The future of work debate has an evidence problem
  • About

  • Careers

  • Newsroom

Yaju AS
Sign in
Book demo

Platform

  • Agent Orchestration System

    Building agents is easy. Operating should be too.

  • Software Factory

    Turn your backlog into review-ready code

  • Agent Hub

    Browse, run, and share agents


Functionalities

  • AI Governance

    Policy enforced at the moment of action

  • AI Observability

    Observe and trust every agent

  • Token Monitoring

    Make every token count

  • Optimizer

    Same outcomes, lower cost

  • AI Spend Explorer

    Find overspend in two minutes


Product

  • MCP Gateway

  • CLI

  • Pricing

  • Versions


Featured

Choosing a model is an operations decision, not a benchmark decision

Use Cases

  • Cost Control

    Know what agents cost. Prove what they deliver.

  • AI Transformation

    Turn AI adoption into business transformation


Deployment

  • Credential Vault

    Org-level secrets, resolved at runtime

  • Self-hosted

    Run agents in your own environment


By Industry

  • IT & Developers

  • Financial Services

  • Public Sector

  • Engineering

  • Telecommunications

  • Healthcare and Life Sciences

  • Manufacturing


Featured

Running agents on hardware you own

Discover

  • Customer Stories

  • Partners

  • Yaju Labs

    Yaju Agent Systems research lab


For Learners

  • Versions

  • Agent Academy


Featured

Support triage is the best first agent most teams never build

Content

  • Blog

    The latest from Yaju, launches, and insights


Explorations

  • Future(s) of Work

    How will AI change the way we work?

  • Oran Models

    The generation teams run today


Initiatives

  • Scholars Program

    Finding the next generation of agent builders

  • Open Development Community

    Building agent tooling in the open

  • Catalyst Grants

    Backing ambitious work on agents


Featured

The future of work debate has an evidence problem
  • About

  • Careers

  • Newsroom

Yaju AS
  • Agent Orchestration System

    Software Factory

    Agent Hub

  • MCP Gateway

    CLI

    Pricing

    Versions

  • AI Governance

    AI Observability

    Token Monitoring

    Optimizer

    AI Spend Explorer

  • Cost Control

    AI Transformation

    Solutions Overview

  • IT & Developers

    Financial Services

    Public Sector

    Engineering

    Telecommunications

    Healthcare and Life Sciences

    Manufacturing

  • Credential Vault

    Self-hosted

    Deployment Options

  • Blog

    Customer Stories

    Partners

    Agent Academy

    Yaju Labs

  • About

    Careers

    Newsroom

  • Legal Center

    Security

    Privacy Policy

    Terms of Use

Platform

  • Agent Orchestration System

  • Software Factory

  • Agent Hub

Product

  • MCP Gateway

  • CLI

  • Pricing

  • Versions

Functionalities

  • AI Governance

  • AI Observability

  • Token Monitoring

  • Optimizer

  • AI Spend Explorer

Solutions

  • Cost Control

  • AI Transformation

  • Solutions Overview

By Industry

  • IT & Developers

  • Financial Services

  • Public Sector

  • Engineering

  • Telecommunications

  • Healthcare and Life Sciences

  • Manufacturing

Deployment

  • Credential Vault

  • Self-hosted

  • Deployment Options

Resources

  • Blog

  • Customer Stories

  • Partners

  • Agent Academy

  • Yaju Labs

Company

  • About

  • Careers

  • Newsroom

Legal

  • Legal Center

  • Security

  • Privacy Policy

  • Terms of Use

Platform

  • Agent Orchestration System

  • Software Factory

  • Agent Hub

Product

  • MCP Gateway

  • CLI

  • Pricing

  • Versions

Functionalities

  • AI Governance

  • AI Observability

  • Token Monitoring

  • Optimizer

  • AI Spend Explorer

Solutions

  • Cost Control

  • AI Transformation

  • Solutions Overview

By Industry

  • IT & Developers

  • Financial Services

  • Public Sector

  • Engineering

  • Telecommunications

  • Healthcare and Life Sciences

  • Manufacturing

Deployment

  • Credential Vault

  • Self-hosted

  • Deployment Options

Resources

  • Blog

  • Customer Stories

  • Partners

  • Agent Academy

  • Yaju Labs

Company

  • About

  • Careers

  • Newsroom

Legal

  • Legal Center

  • Security

  • Privacy Policy

  • Terms of Use

LinkedInInstagramEmail

Yaju AS ©2026

  • English

Models

Choosing a model is an operations decision, not a benchmark decision

Teams spend weeks comparing model scores and then discover that the thing determining their outcome was never the model. Here is how we think about model choice inside an agent system, and why it sits so far down our own list.

Yaju Team · 18 March 2026

There is a familiar sequence. A team compares models, picks one on published scores, builds around it, and finds that the agent behaves unpredictably in ways the comparison never surfaced. So they try another model. The behaviour changes; the unpredictability does not.

The diagnosis is usually the same: the model was never the variable that mattered.

What benchmarks can and cannot tell you

A published score tells you how a model performed on a fixed set of problems under conditions somebody else chose. That is genuinely useful information about capability in general, and nearly useless as a prediction about your specific work.

Your work has context they did not have. Your documents, your conventions, your idea of an acceptable answer, your tolerance for a confident wrong response versus a hedged right one. A model that scores marginally higher in the abstract can easily be the worse choice for a workflow where failure has a particular shape.

This is not an argument against benchmarks. It is an argument against treating them as a substitute for evaluating against your own work.

Build the evaluation before choosing

The practical advice is unglamorous: assemble thirty to fifty real examples from the work you actually intend to automate, with the output you would consider correct. That set is worth more than any comparison table, for three reasons.

It measures the thing you care about. It keeps measuring it when you change something else. And it makes model choice reversible, because swapping a model becomes an experiment with a result rather than a migration with a hope.

Teams that build this set early tend to make model decisions quickly and stop revisiting them. Teams that skip it revisit the decision indefinitely, because nothing ever settles it.

Why the same model behaves differently for two teams

Almost everything around the model is a variable. The retrieval strategy determines what the model sees. The tool set determines what it can attempt and how much context it carries before it starts. The prompt accumulates edge-case handling over months. The failure policy determines whether a shaky answer is returned or held.

Change any of those and behaviour changes measurably, usually by more than a model swap would. That is why we treat model choice as one setting inside an operational system rather than the foundation of it.

Where cost enters

Model pricing is the most visible cost and rarely the largest. Retries, oversized context and unused tool definitions routinely account for more, and none of them change when you switch model.

This is why the Optimizer measures model choice, prompt shape and tool set together, against your evals. Moving to a cheaper model is only a saving if quality holds on your own criteria. Sometimes it does, and the saving is real. Sometimes it does not, and you have traded a visible cost for an invisible one.

Our actual recommendation

Pick a capable model, build your evaluation set, instrument cost per agent, and then spend your attention on retrieval, tools, policy and review. Revisit the model when your evals tell you to, not when a new score is published.

That is a less exciting answer than a comparison table. It is the one that has produced stable systems in the deployments we have watched.

Next

The versions page covers every generation of the Yaju Agent Orchestration System and what each one added. The Agent Orchestration System pages cover the evaluation and optimisation layers that make model choice a reversible decision.

Loading...
Choosing a model is an operations decision, not a benchmark decision | Yaju