Adoption
Platform evaluations tend to compare capability, which is the dimension where every vendor looks similar after an hour. These are the questions that actually separate them, including the ones we would rather you did not ask us casually.
Yaju Team · 11 June 2026
Most evaluation processes we see spend eighty percent of their effort on the demo and twenty on everything else. That ratio is backwards, because the demo is the part every serious vendor has rehearsed.
Here is the list we would use.
Can you show spend per agent, not per workspace or per month? Broken down by user and by model? Reported in currency rather than tokens?
Can you set a limit that blocks rather than warns, per team, and see a forecast for the rest of the cycle?
If the answer to the first is no, cost control in that platform means reading a bill afterwards.
Does every agent have a named owner as a property of the agent, or is that a convention the customer is expected to maintain in a spreadsheet?
What happens to an agent whose owner leaves? Is there anything in the system that notices an agent nobody has touched in six months?
Orphaned agents are the quiet cost of every programme, and platforms differ enormously in whether they make them visible.
Where is a boundary enforced: in the prompt, in the orchestration layer, or at the point where the action occurs?
This is the question with the biggest gap between vendors and the one most easily answered vaguely. "The agent is instructed not to" is not enforcement. Ask what physically prevents the action.
How do credentials reach the agent? If the answer involves the agent holding a secret, ask what revocation looks like.
Can output be scored automatically against criteria you define, on every run rather than on a sample?
Can a run be halted when a check fails, or does scoring only produce a report someone reads later?
Do the criteria cover cost and latency alongside correctness, or only correctness? A platform that optimises cost without evaluating quality will make your system cheaper and worse.
Is every agent action recorded, or a sample? Can it be exported? Does it include what the agent drew on, not only what it did?
If a regulator asked why a specific output was produced eleven months ago, what would you hand them?
Can it run inside your infrastructure, in your own cloud, fully air-gapped? Is the governance identical in those modes or reduced?
And the one people skip: what does leaving look like. Can you export your data and your agent definitions in a form that remains usable? A platform you cannot leave is a dependency rather than a choice.
We would rather you asked these directly. We do not train on your data. Processing is in the EU by default and can be pinned. Governance is identical self-hosted and air-gapped. Our subprocessor list is available on request and we notify before adding one that processes personal data.
Where we are weaker: an isolated deployment cannot benefit from anything that depends on aggregate behaviour, and air-gapped operation changes how support and diagnosis work. Both are workable, and both are worth knowing in advance rather than discovering.
The Trust Center covers compliance, subprocessors and security reporting. Procurement questionnaires can go to contact@capconsultor.eu.