Research
Retrieval quality is usually treated as a ranking problem. A large share of the failures we investigate turn out to be absence problems: the answer was never written down, or was written down somewhere the index never reached.
Yaju Team · 24 June 2026
When an agent gives a wrong answer, the investigation usually starts with ranking. Was the right passage retrieved, and where did it sit.
A surprising share of the time the honest answer is that there was no right passage. The information the agent needed does not exist in written form, or exists in a place nothing indexed.
A retrieval system asked for something that is not there does not return nothing. It returns the closest thing, because that is what ranking does.
The agent then receives material that is topically adjacent and factually irrelevant, and produces an answer grounded in it. The output has citations. It looks better-supported than a hedge would.
That is the worst available failure mode: wrong, confident, and apparently sourced.
Knowledge that lives in people. The exception that is always made for one customer. The reason a process has a step that looks redundant. Nobody wrote it down because everybody knew.
Knowledge in formats nothing indexed. Images, spreadsheets, recordings, diagrams, the shared drive that was never connected.
Knowledge that is out of date rather than missing. Worse than absence, because it retrieves well and is wrong. A superseded policy with no marker saying so will be returned confidently for years.
Knowledge nobody is allowed to see. Correctly restricted, and the agent's answer is incomplete in a way the user cannot detect.
Build an evaluation set from questions people actually ask, not questions you can answer. The temptation is to write questions whose answers you know are in the corpus, which produces a set that reports everything is fine.
Then measure a separate thing from accuracy: for each question, does an adequate source exist at all. That splits failures into "not found" and "not there", which have completely different remedies. One is an indexing project. The other is a writing project.
An honest audit of what is missing is more useful than another tuning cycle, and organisations consistently underrate it.
It tells you what to document, what to connect and what to retire. It is also the only way to know whether your retrieval system is close to its ceiling or nowhere near it, which determines whether further tuning is worth anyone's time.
Where a gap cannot be closed, the system should say so. An agent that answers "I could not find a current source for this" is more useful than one that produces a fluent answer from adjacent material, even though it reads as less capable.
This is a configuration decision and a cultural one. Teams that reward apparent confidence get confident systems, and they get them in the places where confidence is least warranted.
Our research is about how organisations create, use, manage and improve agents, and corpus gaps sit squarely in the improve half. No model change fixes an absence, which makes it exactly the kind of operational problem that gets less attention than it deserves.
The Yaju Labs pages cover the research programme, and the Open Development Community is free to join. The developer documentation covers retrieval, chunking and provenance.