Skip to content

New — Introducing Yaju Software Factory — Try the tool for free

Explore now
Yaju AS
  • ProductProduct
  • SolutionsSolutions
  • ResourcesResources
  • BlogBlog
  • CompanyCompany
Sign in
Book demo

Platform

  • Agent Orchestration System

    Building agents is easy. Operating should be too.

  • Software Factory

    Turn your backlog into review-ready code

  • Agent Hub

    Browse, run, and share agents

Functionalities

  • AI Governance

    Policy enforced at the moment of action

  • AI Observability

    Observe and trust every agent

  • Token Monitoring

    Make every token count

  • Optimizer

    Same outcomes, lower cost

  • AI Spend Explorer

    Find overspend in two minutes

Product

  • MCP Gateway

  • CLI

  • Pricing

  • Versions

Featured

Choosing a model is an operations decision, not a benchmark decision

Use Cases

  • Cost Control

    Know what agents cost. Prove what they deliver.

  • AI Transformation

    Turn AI adoption into business transformation

Deployment

  • Credential Vault

    Org-level secrets, resolved at runtime

  • Self-hosted

    Run agents in your own environment

By Industry

  • IT & Developers

  • Financial Services

  • Public Sector

  • Engineering

  • Telecommunications

  • Healthcare and Life Sciences

  • Manufacturing

Featured

Running agents on hardware you own

Discover

  • Customer Stories

  • Partners

  • Yaju Labs

    Yaju Agent Systems research lab

For Learners

  • Versions

  • Agent Academy

Featured

Support triage is the best first agent most teams never build

Content

  • Blog

    The latest from Yaju, launches, and insights

Explorations

  • Future(s) of Work

    How will AI change the way we work?

  • Oran Models

    The generation teams run today

Initiatives

  • Scholars Program

    Finding the next generation of agent builders

  • Open Development Community

    Building agent tooling in the open

  • Catalyst Grants

    Backing ambitious work on agents

Featured

The future of work debate has an evidence problem
  • About

  • Careers

  • Newsroom

Yaju AS
Sign in
Book demo

Platform

  • Agent Orchestration System

    Building agents is easy. Operating should be too.

  • Software Factory

    Turn your backlog into review-ready code

  • Agent Hub

    Browse, run, and share agents


Functionalities

  • AI Governance

    Policy enforced at the moment of action

  • AI Observability

    Observe and trust every agent

  • Token Monitoring

    Make every token count

  • Optimizer

    Same outcomes, lower cost

  • AI Spend Explorer

    Find overspend in two minutes


Product

  • MCP Gateway

  • CLI

  • Pricing

  • Versions


Featured

Choosing a model is an operations decision, not a benchmark decision

Use Cases

  • Cost Control

    Know what agents cost. Prove what they deliver.

  • AI Transformation

    Turn AI adoption into business transformation


Deployment

  • Credential Vault

    Org-level secrets, resolved at runtime

  • Self-hosted

    Run agents in your own environment


By Industry

  • IT & Developers

  • Financial Services

  • Public Sector

  • Engineering

  • Telecommunications

  • Healthcare and Life Sciences

  • Manufacturing


Featured

Running agents on hardware you own

Discover

  • Customer Stories

  • Partners

  • Yaju Labs

    Yaju Agent Systems research lab


For Learners

  • Versions

  • Agent Academy


Featured

Support triage is the best first agent most teams never build

Content

  • Blog

    The latest from Yaju, launches, and insights


Explorations

  • Future(s) of Work

    How will AI change the way we work?

  • Oran Models

    The generation teams run today


Initiatives

  • Scholars Program

    Finding the next generation of agent builders

  • Open Development Community

    Building agent tooling in the open

  • Catalyst Grants

    Backing ambitious work on agents


Featured

The future of work debate has an evidence problem
  • About

  • Careers

  • Newsroom

Yaju AS
  • Agent Orchestration System

    Software Factory

    Agent Hub

  • MCP Gateway

    CLI

    Pricing

    Versions

  • AI Governance

    AI Observability

    Token Monitoring

    Optimizer

    AI Spend Explorer

  • Cost Control

    AI Transformation

    Solutions Overview

  • IT & Developers

    Financial Services

    Public Sector

    Engineering

    Telecommunications

    Healthcare and Life Sciences

    Manufacturing

  • Credential Vault

    Self-hosted

    Deployment Options

  • Blog

    Customer Stories

    Partners

    Agent Academy

    Yaju Labs

  • About

    Careers

    Newsroom

  • Legal Center

    Security

    Privacy Policy

    Terms of Use

Platform

  • Agent Orchestration System

  • Software Factory

  • Agent Hub

Product

  • MCP Gateway

  • CLI

  • Pricing

  • Versions

Functionalities

  • AI Governance

  • AI Observability

  • Token Monitoring

  • Optimizer

  • AI Spend Explorer

Solutions

  • Cost Control

  • AI Transformation

  • Solutions Overview

By Industry

  • IT & Developers

  • Financial Services

  • Public Sector

  • Engineering

  • Telecommunications

  • Healthcare and Life Sciences

  • Manufacturing

Deployment

  • Credential Vault

  • Self-hosted

  • Deployment Options

Resources

  • Blog

  • Customer Stories

  • Partners

  • Agent Academy

  • Yaju Labs

Company

  • About

  • Careers

  • Newsroom

Legal

  • Legal Center

  • Security

  • Privacy Policy

  • Terms of Use

Platform

  • Agent Orchestration System

  • Software Factory

  • Agent Hub

Product

  • MCP Gateway

  • CLI

  • Pricing

  • Versions

Functionalities

  • AI Governance

  • AI Observability

  • Token Monitoring

  • Optimizer

  • AI Spend Explorer

Solutions

  • Cost Control

  • AI Transformation

  • Solutions Overview

By Industry

  • IT & Developers

  • Financial Services

  • Public Sector

  • Engineering

  • Telecommunications

  • Healthcare and Life Sciences

  • Manufacturing

Deployment

  • Credential Vault

  • Self-hosted

  • Deployment Options

Resources

  • Blog

  • Customer Stories

  • Partners

  • Agent Academy

  • Yaju Labs

Company

  • About

  • Careers

  • Newsroom

Legal

  • Legal Center

  • Security

  • Privacy Policy

  • Terms of Use

LinkedInInstagramEmail

Yaju AS ©2026

  • English

Deployment

Running reranking inside your own cloud

Reranking is the retrieval step most worth adding and the one most often skipped in restricted environments, because it means another component to host. Here is what that actually costs and how to decide whether it earns its place.

Yaju Team · 8 April 2026

Teams operating in environments where data cannot leave face a recurring dilemma with retrieval quality. The techniques that help most are the ones that add components, and every additional component is something to deploy, monitor and upgrade inside the boundary.

Reranking is the clearest case. It is usually the largest single accuracy improvement available, and it is an extra service.

What you are adding

A reranking stage sits between first-stage retrieval and the model. It takes a shortlist of candidates and reorders them by judging each against the query directly, rather than comparing pre-computed vectors.

Operationally this is a service with its own resource profile: heavier per item than vector search, applied to a small set rather than a whole index. It needs capacity, monitoring and an upgrade path like anything else you run.

What you get back

Two things at once, which is unusual.

Better grounding, because the passage that answers the question arrives at the top rather than fifth.

Lower context cost on every call, because you can pass three good passages instead of ten mediocre ones. In a self-hosted deployment where you are paying for your own compute, that reduction applies to every single run, permanently.

The second benefit is the one that usually settles the decision in restricted environments, because it offsets the cost of the component you just added.

How to decide before deploying it

Measure retrieval position before you build anything. Take real questions with their correct passages and record whether the correct passage is retrieved and where it ranks.

If it is usually retrieved but ranked low, reranking will help substantially and you can estimate by how much before committing.

If it is usually not retrieved at all, reranking cannot help. The problem is upstream in chunking or embedding, and adding a component will cost capacity and change nothing.

This measurement takes an afternoon and prevents the most common wasted quarter in retrieval work.

Capacity planning

Reranking load is proportional to query volume rather than corpus size, which makes it easier to forecast than the index itself. The variable that matters is how many candidates you rerank per query.

Reranking fifty candidates instead of twenty costs more and helps less than most teams expect. Measure it rather than assuming more is better.

In air-gapped environments

The same considerations apply with one addition: model updates for the reranker arrive as artefacts on your schedule, like everything else. Build your evaluation set first so that accepting or declining an update is a measured decision rather than a leap.

Where it fits in evaluation

Adding a stage changes latency as well as accuracy and cost, and all three belong in the same scoring. Evals in Oran cover behavioural, structural, latency and cost criteria together for exactly this reason: an accuracy gain that breaks an interactive latency budget is a trade, not an improvement, and it should be visible at the moment it is made.

Next

The reranking and deployment pages cover mechanics and self-hosted options. The developer documentation covers building the evaluation set that makes this decision measurable.

Loading...
Running reranking inside your own cloud | Yaju