Engineering
Most teams pick a chunk size in the first hour of a project and never revisit it. It then quietly determines what an agent can and cannot find for the rest of that system's life. Here is how we think about it, and how to tell when yours is wrong.
Yaju Team · 3 March 2026
There is a moment early in every retrieval project where someone types a number. Five hundred tokens. A thousand. Whatever the tutorial used. The pipeline works, the demo answers questions, and that number is never examined again.
It deserves more attention than it gets, because chunking is not a preprocessing detail. It is the decision that determines which facts can ever be retrieved together, and no amount of clever ranking downstream can recover a relationship that chunking destroyed.
Chunks that are too small fragment a single idea across several pieces. The answer exists in your corpus, but no individual chunk contains enough of it to look relevant, so nothing scores highly and the agent answers from its own assumptions instead.
Chunks that are too large drag unrelated material along with the part that matched. The retrieval looks successful, the context window fills with noise, and the agent produces something that is subtly about the wrong paragraph. This failure is worse than the first, because it produces confident output rather than an obvious gap.
Both failures look identical from the outside: the agent gives an answer that is wrong in a way nobody can immediately explain.
The most useful change we see teams make is to stop thinking in token counts and start thinking in document structure.
A policy document has sections. A support archive has one problem per ticket. An API reference has one endpoint per entry. Those boundaries were put there by a person for the same reason you need them: they mark where one idea stops and another begins. Splitting on them, with a size cap as a fallback rather than the primary rule, consistently outperforms splitting on length alone.
Where documents have no usable structure, overlap is the cheap mitigation. Carrying a sentence or two across the boundary costs storage and very little else, and it rescues the specific case where the key sentence lands one token past a split.
A chunk is not just text. It is text plus everything you know about where it came from: the document it belongs to, the section heading above it, the date it was written, the team that owns it, whether it has been superseded.
That metadata is what lets an agent filter before it ranks, which is usually a larger accuracy win than any embedding change. "Search the current version of the compliance handbook" is a fundamentally easier problem than "search everything and hope the current version scores highest".
It also gives you an answer to the question that follows every wrong output: where did this come from. A chunk that carries its provenance turns debugging from archaeology into a lookup.
Chunking strategy is the kind of topic that generates long opinions and short evidence. The way out is the same as everywhere else in agent work: write the evaluation first.
Assemble thirty or forty real questions that people actually ask your system, with the passage that should answer each one. Then measure whether that passage is retrieved at all. Not whether the final answer reads well, which conflates retrieval and generation, but whether the right material reached the model.
With that in place, chunking becomes an experiment rather than a debate. Change the strategy, rerun the set, keep the version that retrieves more. Teams that build this harness early tend to stop having chunking arguments entirely, because there is now a cheap way to settle them.
Retrieval quality is not a separate concern from agent operation. A retrieval change that adds two hundred tokens to every call is a cost change, and a retrieval change that improves recall is a quality change. Both belong in the same evaluation loop as everything else the agent does.
That is why evals in Oran score behavioural, structural, latency and cost criteria together, and why the Optimizer measures its own changes against those same evals. Trimming context to save money is only a saving if retrieval still finds what it needs.
If you already have a retrieval system in production and no evaluation set, build the evaluation set first, even a crude one. It will tell you more in an afternoon than a month of tuning. The developer documentation covers the mechanics, and the Agent Orchestration System pages cover how the evaluation and cost layers fit around it.