Engineering
The hardest part of improving retrieval is not technique, it is having something to measure against. If you cannot yet build an evaluation set from your own documents, building one from a public corpus first is a genuinely useful rehearsal.
Yaju Team · 13 April 2026
Almost every retrieval improvement is easy to implement and impossible to justify without measurement. Change the chunk size and something happens; whether it was an improvement is anybody's guess.
So the first piece of infrastructure worth building is not a better pipeline. It is a set of questions with known answers.
Building an evaluation set from internal documents runs into two obstacles immediately: access approval, and disagreement about what the correct answer is. Both are solvable and neither is fast.
A large public reference corpus sidesteps both. The documents are available, the facts are checkable, and nobody has to approve anything. You are not evaluating your own content, which means the results do not transfer directly, but the point of the exercise is to build the harness and learn to read it.
Treat it as a rehearsal, not a substitute.
Real questions, phrased the way people actually phrase them, including the short and ambiguous ones. A set made entirely of well-formed queries will flatter any system.
A known target passage for each question, so you can measure retrieval directly rather than inferring it from the final answer.
Enough variety to distinguish failure modes: some questions answerable from one passage, some requiring two, some whose answer is in a table, some where the obvious keyword appears in several irrelevant places.
Thirty to fifty questions is enough to start. Precision matters less than having anything at all.
Whether the target passage was retrieved at all, and at what position. That single number separates the two failure modes that otherwise look identical: material that was never found, and material that was found and buried.
The first sends you to chunking and embedding. The second sends you to reranking. Without the distinction, teams routinely spend weeks tuning the layer that was already working.
Once the harness exists and you can read its output, rebuilding the set against internal documents is a much smaller task, because the hard part was never the corpus. It was knowing what to record and what to compare.
Your own set will behave differently, and that is the point. Internal vocabulary, inconsistent formatting, superseded versions and documents nobody has opened in three years are exactly the conditions a public corpus does not reproduce.
An evaluation set is not a one-off exercise. It is the thing that makes every later change measurable: a new embedding model, a different chunking strategy, a reranker, a cheaper model proposed by the Optimizer.
This is why evals in Oran run on every agent execution rather than as a periodic audit. A set that is consulted once during a project tells you about that project. A set that runs continuously tells you when something drifted, which is the failure that would otherwise be found by a user.
The developer documentation covers retrieval and reranking mechanics. The Agent Orchestration System pages cover how evaluation, cost and optimisation work as one loop rather than three separate activities.