Company
Our community gathering ran for two days and the most useful sessions were not the ones we scheduled. A short account of what people actually wanted to discuss.
Yaju Team · 30 October 2025
We prepared an agenda for Connect. The agenda was fine. The sessions people stayed in afterwards were not the ones on it, which is worth recording.
Not how to build an agent. Nobody asked that. The recurring question was some version of: we have thirty of these now and I cannot tell you what they do.
That is the sprawl problem, and hearing it from teams in unrelated sectors on the same afternoon was clarifying. It is not a symptom of doing something wrong. It is what happens when building becomes cheap and operating has not caught up.
Someone asked how other people write evaluation criteria, and a scheduled thirty minutes became most of an afternoon.
The interesting disagreement was about when to write them. The textbook answer is before building. Almost everyone in the room had written their first eval after an agent produced something wrong, working backwards from the failure.
The consensus that emerged was that the backwards version is fine and possibly better, because criteria derived from a real failure are sharper than criteria imagined in advance. What matters is that they exist and run on every execution rather than being consulted once.
How much human review is enough. One group argued for review of everything until an agent has earned trust. Another argued that reviewing everything guarantees the programme never scales and produces inattentive review anyway.
Nobody changed their mind. Both positions are defensible and the answer clearly depends on what the agent touches, which is a less satisfying conclusion than either side wanted.
Three things went straight into how we talk about the platform. That ownership per agent is the first thing to install, not the last. That evals written after a failure are legitimate. That "what are we running" is the question underneath most others.
To everyone who came, and particularly to the people who described something that had not worked. Those were the useful sessions.
The Open Development Community is free to join and is where most of this continues between gatherings. Yaju Labs covers the research programme.