Engineering
Most retrieval systems return ten passages and hand all ten to the model. Adding a step that reorders them and keeps the best three is usually the largest single accuracy improvement available, and it reduces cost at the same time.
Yaju Team · 30 March 2026
First-stage retrieval is built for speed. It scans a large index quickly and returns candidates that are probably relevant. Probably is doing real work in that sentence.
The standard response is to widen the net: return more passages and let the model sort it out. This works, in the sense that the right passage is now usually somewhere in the set, and it fails in two ways nobody measures.
Every additional passage occupies context on every call, so the cost of the workaround is permanent and scales with traffic.
And the model has to locate the relevant material among the noise. Models are decent at this and not reliable at it, particularly when a plausible-looking irrelevant passage sits above the correct one. That is the failure that produces a confident answer built on the wrong paragraph.
A reranker looks at the query and each candidate together, rather than comparing pre-computed vectors. It is slower per item, which is why it cannot scan the whole index, and much better at judging relevance, which is why it is worth running over a shortlist.
The pattern is: retrieve twenty or fifty candidates cheaply, rerank them properly, keep the top three or five. The model receives less material and better material at the same time.
That combination is unusual. Most quality improvements cost money. This one typically reduces the context on every call while raising the proportion of answers grounded in the right source.
Measure retrieval separately from generation. Take real questions with their correct passages, and check whether the correct passage appears in the set the model receives, and at what position.
If the correct passage is usually retrieved but ranked fifth or eighth, a reranker will help substantially. If it is usually not retrieved at all, the problem is upstream in chunking or embedding, and reranking cannot fix an absence.
This distinction saves a lot of wasted effort, and it takes an afternoon to establish.
Small corpora where first-stage retrieval already ranks the right passage first. Latency-critical paths where the extra step does not fit the budget. Queries that are essentially lookups, where a metadata filter does the job more cheaply than any ranking.
A reranker is a good default, not a universal one.
Adding a stage changes both latency and cost, which means it belongs in the same evaluation loop as everything else. Evals in Oran score behavioural, structural, latency and cost criteria together for exactly this reason: a change that improves accuracy and breaks a latency budget is not automatically an improvement, and that trade-off should be visible rather than discovered.
The reranking pages cover the mechanics and deployment options. The developer documentation covers building the evaluation set that tells you whether any of this is worth doing for your corpus.