Company
Almost nobody has years of experience operating agents in production, because the job barely existed three years ago. So we stopped looking for it, and started looking for the things that transfer.
Yaju Team · 25 August 2026
If we required three years of agent operations experience, we would be choosing from a very small pool of people who mostly learned the same lessons at the same two or three companies. That is not a hiring strategy, it is a bottleneck.
So the question we ask instead is which existing experience actually transfers. The answer has been fairly consistent.
People who have run systems in production understand the thing that catches agent teams out: that the interesting part starts after it works.
They ask who gets paged. They ask what happens when this fails at three in the morning and the person who built it has left. They instrument before they optimise. They are suspicious of systems that cannot explain what they did.
Every one of those instincts is directly applicable to agents, and none of them are obvious to someone whose experience is building things that get demonstrated.
Agent quality is mostly a data problem wearing a model costume. Retrieval, chunking, provenance, staleness, the difference between a document that exists and one that can be found.
Engineers who have spent time on pipelines already think in terms of where data came from, what transformed it and whether the transformation lost something. That is exactly the lens that finds the parsing defect nobody was looking for.
We ask about a system someone ran that misbehaved in a way that was hard to explain, and we listen for the shape of the investigation. Did they form a hypothesis and test it, or did they change things until it stopped? Both happen, and only one of them scales.
We ask how they would know if a system had quietly got worse. Candidates who reach for measurement rather than vigilance tend to do well here.
And we ask what they would refuse to automate. It is a values question disguised as a technical one, and the answers are informative.
Recall of model names and benchmark scores. That knowledge has a half-life of about four months and is easy to acquire.
Whether someone has used our platform. Obviously not, and we would be a strange company if we expected it.
Speed at a whiteboard puzzle, which measures practice at whiteboard puzzles.
The work has a substantial unglamorous component. A lot of it is attribution, ownership, evaluation criteria and audit trails rather than model experimentation. People who find that boring are not wrong, they are just suited to something else, and it is better established in an interview than in month four.
We are remote by default with a preference for working in person where it is practical, and we say that rather than presenting distribution as universally better.
The Careers page carries open roles. The Scholars Program is a remote, part-time route with a mentor and no requirement for prior research experience. The Open Development Community is free to join and is a reasonable way to see how we work before deciding whether you want to.
Careers, the Scholars Program and the Open Development Community pages cover each route. Questions can go to contact@capconsultor.eu.