Engineering
The diagram that explains the architecture, the screenshot in the incident report, the photographed whiteboard from the design review. All of it sits in systems an agent can reach and none of it is searchable, because the index only ever saw the words around it.
Yaju Team · 7 April 2026
Ask an engineer where the system architecture is documented and you will often be pointed at a diagram. Ask a retrieval system the same question and it will return the paragraph that says "see the diagram below".
That gap is not a small one. In most organisations a meaningful share of the actual knowledge lives in images, and the standard pipeline treats them as gaps between paragraphs.
When images are embedded in the same space as text, a question asked in words can retrieve a picture. "What talks to the payment service" can return the architecture diagram rather than the sentence introducing it. An incident description can retrieve the screenshot of the error rather than the ticket title.
The useful part is not novelty. It is that the material people actually rely on stops being invisible to the system that is supposed to find things for them.
Decorative images. Not every image carries information. Stock photography, logos and section dividers indexed alongside meaningful diagrams add noise to every query. A filter, even a crude one based on source or context, pays for itself.
Text inside images. Screenshots are mostly text that visual embedding handles poorly. These generally want recognition first, so the text becomes searchable text, with the image retained for reference.
Missing context. A chart is far more retrievable when it carries its caption, the heading above it and the document it came from. An image indexed alone is a picture with no anchor, and its retrieval quality suffers accordingly.
When a system returns a paragraph, a person can read it and judge. When it returns an image, judgement depends entirely on knowing where the image came from and when.
An outdated architecture diagram retrieved confidently is worse than no diagram, because it looks authoritative in a way prose does not. Carrying the source document, the date and the section with every indexed image is what keeps this safe.
The same discipline applies as everywhere else: build a set of real questions whose correct answer is an image, and measure whether that image is retrieved. Evaluating multimodal retrieval by reading the final answer conflates two layers and will send you tuning the wrong one.
Include the awkward cases deliberately. The near-duplicate diagram from last year. The screenshot that is mostly text. The photograph of a whiteboard. Those are where systems separate.
Images are not free to index or to carry in context. A retrieval change that routinely attaches an image to every answer is a cost change, and it belongs in the same evaluation as accuracy. Cheaper is only cheaper if quality holds, and richer is only better if someone is willing to pay for it on every call.
The developer documentation covers embedding and retrieval mechanics. The Agent Orchestration System pages cover the evaluation and cost layers that keep a change like this measurable rather than merely impressive.