Deployment
Reranking is the retrieval step most worth adding and the one most often skipped in restricted environments, because it means another component to host. Here is what that actually costs and how to decide whether it earns its place.
Yaju Team · 8 April 2026
Teams operating in environments where data cannot leave face a recurring dilemma with retrieval quality. The techniques that help most are the ones that add components, and every additional component is something to deploy, monitor and upgrade inside the boundary.
Reranking is the clearest case. It is usually the largest single accuracy improvement available, and it is an extra service.
A reranking stage sits between first-stage retrieval and the model. It takes a shortlist of candidates and reorders them by judging each against the query directly, rather than comparing pre-computed vectors.
Operationally this is a service with its own resource profile: heavier per item than vector search, applied to a small set rather than a whole index. It needs capacity, monitoring and an upgrade path like anything else you run.
Two things at once, which is unusual.
Better grounding, because the passage that answers the question arrives at the top rather than fifth.
Lower context cost on every call, because you can pass three good passages instead of ten mediocre ones. In a self-hosted deployment where you are paying for your own compute, that reduction applies to every single run, permanently.
The second benefit is the one that usually settles the decision in restricted environments, because it offsets the cost of the component you just added.
Measure retrieval position before you build anything. Take real questions with their correct passages and record whether the correct passage is retrieved and where it ranks.
If it is usually retrieved but ranked low, reranking will help substantially and you can estimate by how much before committing.
If it is usually not retrieved at all, reranking cannot help. The problem is upstream in chunking or embedding, and adding a component will cost capacity and change nothing.
This measurement takes an afternoon and prevents the most common wasted quarter in retrieval work.
Reranking load is proportional to query volume rather than corpus size, which makes it easier to forecast than the index itself. The variable that matters is how many candidates you rerank per query.
Reranking fifty candidates instead of twenty costs more and helps less than most teams expect. Measure it rather than assuming more is better.
The same considerations apply with one addition: model updates for the reranker arrive as artefacts on your schedule, like everything else. Build your evaluation set first so that accepting or declining an update is a measured decision rather than a leap.
Adding a stage changes latency as well as accuracy and cost, and all three belong in the same scoring. Evals in Oran cover behavioural, structural, latency and cost criteria together for exactly this reason: an accuracy gain that breaks an interactive latency budget is a trade, not an improvement, and it should be visible at the moment it is made.
The reranking and deployment pages cover mechanics and self-hosted options. The developer documentation covers building the evaluation set that makes this decision measurable.