Skip to content

New — Introducing Yaju Software Factory — Try the tool for free

Explore now
Yaju AS
  • ProductProduct
  • SolutionsSolutions
  • ResourcesResources
  • BlogBlog
  • CompanyCompany
Sign in
Book demo

Platform

  • Agent Orchestration System

    Building agents is easy. Operating should be too.

  • Software Factory

    Turn your backlog into review-ready code

  • Agent Hub

    Browse, run, and share agents

Functionalities

  • AI Governance

    Policy enforced at the moment of action

  • AI Observability

    Observe and trust every agent

  • Token Monitoring

    Make every token count

  • Optimizer

    Same outcomes, lower cost

  • AI Spend Explorer

    Find overspend in two minutes

Product

  • MCP Gateway

  • CLI

  • Pricing

  • Versions

Featured

Choosing a model is an operations decision, not a benchmark decision

Use Cases

  • Cost Control

    Know what agents cost. Prove what they deliver.

  • AI Transformation

    Turn AI adoption into business transformation

Deployment

  • Credential Vault

    Org-level secrets, resolved at runtime

  • Self-hosted

    Run agents in your own environment

By Industry

  • IT & Developers

  • Financial Services

  • Public Sector

  • Engineering

  • Telecommunications

  • Healthcare and Life Sciences

  • Manufacturing

Featured

Running agents on hardware you own

Discover

  • Customer Stories

  • Partners

  • Yaju Labs

    Yaju Agent Systems research lab

For Learners

  • Versions

  • Agent Academy

Featured

Support triage is the best first agent most teams never build

Content

  • Blog

    The latest from Yaju, launches, and insights

Explorations

  • Future(s) of Work

    How will AI change the way we work?

  • Oran Models

    The generation teams run today

Initiatives

  • Scholars Program

    Finding the next generation of agent builders

  • Open Development Community

    Building agent tooling in the open

  • Catalyst Grants

    Backing ambitious work on agents

Featured

The future of work debate has an evidence problem
  • About

  • Careers

  • Newsroom

Yaju AS
Sign in
Book demo

Platform

  • Agent Orchestration System

    Building agents is easy. Operating should be too.

  • Software Factory

    Turn your backlog into review-ready code

  • Agent Hub

    Browse, run, and share agents


Functionalities

  • AI Governance

    Policy enforced at the moment of action

  • AI Observability

    Observe and trust every agent

  • Token Monitoring

    Make every token count

  • Optimizer

    Same outcomes, lower cost

  • AI Spend Explorer

    Find overspend in two minutes


Product

  • MCP Gateway

  • CLI

  • Pricing

  • Versions


Featured

Choosing a model is an operations decision, not a benchmark decision

Use Cases

  • Cost Control

    Know what agents cost. Prove what they deliver.

  • AI Transformation

    Turn AI adoption into business transformation


Deployment

  • Credential Vault

    Org-level secrets, resolved at runtime

  • Self-hosted

    Run agents in your own environment


By Industry

  • IT & Developers

  • Financial Services

  • Public Sector

  • Engineering

  • Telecommunications

  • Healthcare and Life Sciences

  • Manufacturing


Featured

Running agents on hardware you own

Discover

  • Customer Stories

  • Partners

  • Yaju Labs

    Yaju Agent Systems research lab


For Learners

  • Versions

  • Agent Academy


Featured

Support triage is the best first agent most teams never build

Content

  • Blog

    The latest from Yaju, launches, and insights


Explorations

  • Future(s) of Work

    How will AI change the way we work?

  • Oran Models

    The generation teams run today


Initiatives

  • Scholars Program

    Finding the next generation of agent builders

  • Open Development Community

    Building agent tooling in the open

  • Catalyst Grants

    Backing ambitious work on agents


Featured

The future of work debate has an evidence problem
  • About

  • Careers

  • Newsroom

Yaju AS
  • Agent Orchestration System

    Software Factory

    Agent Hub

  • MCP Gateway

    CLI

    Pricing

    Versions

  • AI Governance

    AI Observability

    Token Monitoring

    Optimizer

    AI Spend Explorer

  • Cost Control

    AI Transformation

    Solutions Overview

  • IT & Developers

    Financial Services

    Public Sector

    Engineering

    Telecommunications

    Healthcare and Life Sciences

    Manufacturing

  • Credential Vault

    Self-hosted

    Deployment Options

  • Blog

    Customer Stories

    Partners

    Agent Academy

    Yaju Labs

  • About

    Careers

    Newsroom

  • Legal Center

    Security

    Privacy Policy

    Terms of Use

Platform

  • Agent Orchestration System

  • Software Factory

  • Agent Hub

Product

  • MCP Gateway

  • CLI

  • Pricing

  • Versions

Functionalities

  • AI Governance

  • AI Observability

  • Token Monitoring

  • Optimizer

  • AI Spend Explorer

Solutions

  • Cost Control

  • AI Transformation

  • Solutions Overview

By Industry

  • IT & Developers

  • Financial Services

  • Public Sector

  • Engineering

  • Telecommunications

  • Healthcare and Life Sciences

  • Manufacturing

Deployment

  • Credential Vault

  • Self-hosted

  • Deployment Options

Resources

  • Blog

  • Customer Stories

  • Partners

  • Agent Academy

  • Yaju Labs

Company

  • About

  • Careers

  • Newsroom

Legal

  • Legal Center

  • Security

  • Privacy Policy

  • Terms of Use

Platform

  • Agent Orchestration System

  • Software Factory

  • Agent Hub

Product

  • MCP Gateway

  • CLI

  • Pricing

  • Versions

Functionalities

  • AI Governance

  • AI Observability

  • Token Monitoring

  • Optimizer

  • AI Spend Explorer

Solutions

  • Cost Control

  • AI Transformation

  • Solutions Overview

By Industry

  • IT & Developers

  • Financial Services

  • Public Sector

  • Engineering

  • Telecommunications

  • Healthcare and Life Sciences

  • Manufacturing

Deployment

  • Credential Vault

  • Self-hosted

  • Deployment Options

Resources

  • Blog

  • Customer Stories

  • Partners

  • Agent Academy

  • Yaju Labs

Company

  • About

  • Careers

  • Newsroom

Legal

  • Legal Center

  • Security

  • Privacy Policy

  • Terms of Use

LinkedInInstagramEmail

Yaju AS ©2026

  • English

Models

When a bigger model is the wrong answer

The reflex when an agent underperforms is to reach for more capability. Sometimes that is right. More often the failure is somewhere the model cannot reach, and upgrading simply makes the same mistake more expensively.

Yaju Team · 2 April 2026

An agent produces a bad output. Someone proposes the larger model. It is the obvious move, it is easy to justify, and roughly half the time it changes nothing except the invoice.

The question worth asking first is what kind of failure this was.

Four failures that look identical and are not

  • The agent did not have the information. Retrieval missed it, or it was never in the corpus. A larger model will produce a more fluent answer built on the same absence.
  • The agent had the information and too much else. The relevant passage was present, buried in noise that a longer context window makes worse rather than better.
  • The agent was not told what good looks like. Without criteria, output quality is whatever the model defaults to, which varies run to run regardless of size.
  • The agent did something it should not have been allowed to attempt. This is a policy failure wearing a capability costume.

Only one of these is a capability problem. The other three survive an upgrade intact.

How to tell the difference in an afternoon

Take the failing case and check, in order: was the correct source retrieved at all; how much irrelevant material came with it; does a written criterion exist that this output would have failed; and was the action within the boundary the agent should have had.

If the answer to the first is no, fix retrieval. If the second is large, fix chunking or trim tools. If no criterion exists, write one. If the boundary was wrong, that is a governance change, not a model change.

Only when all four check out is the model the remaining variable, and at that point the upgrade is a reasonable experiment with a clear hypothesis.

Why the reflex is so strong

Because a model upgrade is a decision one person can make in a morning, and the alternatives are projects. Fixing retrieval means understanding a corpus. Writing criteria means agreeing what good means, which is an organisational conversation disguised as a technical one.

The upgrade is not chosen because it is likely to work. It is chosen because it is available.

The cost shape

A larger model changes the price of every call the agent makes, forever, including the calls that were already working. If it fixes ten percent of failures, you have paid for one hundred percent of traffic to improve a tenth of it.

This is exactly what evaluation-driven optimisation is for. When the Optimizer proposes a change, it is measured against your own evals, so the question stops being "does this feel better" and becomes "did the score move, and what did it cost".

The cases where capability genuinely is the answer

They exist, and they have a recognisable shape: long multi-step reasoning where intermediate errors compound, tasks with unusual structure that smaller models mishandle consistently rather than occasionally, and work where the correct answer requires holding several constraints simultaneously.

If your failures are concentrated there and your retrieval, criteria and boundaries are sound, upgrade with confidence.

Next

The versions page covers each Oran version and what it added. The Agent Orchestration System pages cover the evaluation and cost layers that turn a model decision into a measurable one.

Loading...
When a bigger model is the wrong answer | Yaju