Product
Translation inside a working system is not a general language problem. It is a narrow, repetitive job with a known input shape and a checkable output, which makes it one of the better first agents a team can build.
Yaju Team · 23 June 2026
Teams looking for a first agent often reach for something ambitious, on the theory that a small win is not worth the setup. We would argue the opposite, and translation inside an existing workflow is a good illustration.
The input shape is known. Support messages, product descriptions, internal notices, interface strings. Not arbitrary prose, but a recognisable kind of text arriving in a predictable form.
The output is checkable. Someone who reads both languages can judge it, which means an evaluation set is straightforward to build.
It is repetitive and continuous. The value does not depend on one impressive result, it accumulates across a year.
And it is recoverable. A poor translation in an internal notice is a correction, not an incident. First agents should be chosen with the failure in mind rather than the success.
Not grammar. Terminology.
Every organisation has words that must be rendered one specific way: product names, feature names, legal terms, status labels that map to something in a system. A translation that is linguistically excellent and uses the wrong term for a product is wrong in the way that matters.
This is why a glossary matters more than model capability here. An agent constrained to the organisation's approved terminology will outperform a more capable one that renders terms freshly each time and is inconsistent across a corpus.
Human translators working across a large body of text drift, because they are people doing careful work over months. An agent given a glossary does not drift, and the resulting consistency is genuinely valuable for anything users encounter repeatedly.
This is a case where the agent is not merely cheaper. It is better at one specific property that matters.
Per language, always. An aggregate quality figure averages a well-served language with an underserved one and reports that everything is fine.
On terminology specifically. Build a check that verifies approved terms were used, which is mechanical and catches the failure that matters most.
On your own material, including the awkward cases: the message with a product name mid-sentence, the string with placeholders, the notice with a legal phrase that has an approved rendering.
The temptation after it works is to widen it: add a language, add a document type, add a step. Resist for a while. The agents still running a year later are the ones whose boundary was set once and held.
If a second job appears, build a second agent. Ten agents doing one thing each are individually evaluable and individually replaceable. One agent doing ten things is neither.
Beyond the output, a first agent like this teaches a team the four habits that everything else depends on: give it an owner, write criteria before building, attribute its cost, and set a boundary that is enforced rather than described.
Learning those on something recoverable is considerably better than learning them on something that is not.
The Agent Hub pages cover publishing and reusing agents across teams. The Agent Orchestration System pages cover ownership, evaluation and attribution.