Research
An agent that performs well in English and adequately elsewhere is not a multilingual agent, it is an English agent with a translation layer and an uneven failure rate. The difference matters most in the places it is hardest to see.
Yaju Team · 12 August 2026
Organisations that operate across several countries tend to discover language problems late, because the systems that surface them are themselves built around the dominant language.
The support queue in one language looks fine. The equivalent queue in another has a slightly higher escalation rate, which is attributed to staffing. Nobody connects it to retrieval.
Rarely in generation, which is the part people test. Usually upstream.
Retrieval. If your index was built with a strategy tuned on one language, chunk boundaries and embeddings can behave differently on another, particularly for languages with different word segmentation or morphology. The right passage exists and is not retrieved.
Parsing. Documents in other scripts, particularly mixed-direction text with Latin product names, can extract in the wrong order and produce text that is individually correct and collectively scrambled.
Vocabulary. The words that matter most in a business context, product names and internal terms, are exactly the ones a general system handles least reliably outside its dominant training language.
In English, a bad retrieval usually produces an answer that a reviewer can see is unsupported. In a language the reviewing team reads less fluently, the same output is harder to check, and the review that would have caught it is weaker precisely where it is needed more.
This is why aggregate quality metrics hide the problem. An overall accuracy figure averages a well-served majority with an underserved minority, and the average looks acceptable.
Per language, always. Never an aggregate. If you cannot break quality down by language, you cannot see the problem at all.
Retrieval separately from generation, per language. That distinguishes "the passage was never found" from "the passage was found and the answer was poor", and those have different fixes.
Your own vocabulary, deliberately. Build evaluation questions containing your product names and internal terms in each language you operate in, because that is where general systems diverge most.
Chunking on document structure rather than token counts, because structure is language-independent in a way that length is not.
Hybrid retrieval combining semantic and keyword search, which rescues exact terms that embeddings handle unevenly.
A vocabulary list for your organisation, where the system supports one.
And testing with real documents from each region rather than translations of one region's documents, which reproduce the source language's structure and flatter the system.
Yaju Labs studies how organisations create, use, manage and improve agents, and uneven performance across languages is one of the clearest cases where the operational layer matters more than model capability. A better model does not fix a retrieval strategy that was tuned on one language.
The Yaju Labs pages cover the research programme, and the Open Development Community is free to join if you want to contribute rather than read. The developer documentation covers retrieval and chunking.