Engineering
Maintaining a fork is unglamorous, never finished, and quietly expensive. It also has exactly the properties that make a task suitable for an agent, which is rarer than the enthusiasm around agents suggests.
Yaju Team · 21 May 2026
Most organisations running substantial software have at least one fork they wish they did not have. A library patched for a requirement upstream would not take. A vendored dependency with local modifications. Something forked three years ago for a reason nobody present can fully explain.
The work of keeping it current is the definition of a task that is important and never urgent, which means it happens in bursts of guilt.
Most proposed agent tasks fail one of three tests. This one passes all of them.
It is well-specified. Take the upstream changes, apply them to a tree with known local modifications, resolve the mechanical conflicts, run the tests.
It has a natural review point. The output is a pull request, and reviewing one is already something engineers do.
It is bounded. There is a defined set of upstream commits and a defined tree, so the agent cannot wander indefinitely.
Compare that to "handle customer enquiries", which fails all three, and the difference in outcomes is not mysterious.
It should apply the mechanical changes, resolve conflicts where the resolution follows from the structure of the change, run the test suite, and write a description of what it did and which conflicts it resolved and how.
It should not decide that a failing test is acceptable, silently drop a local modification because it conflicted, or widen its scope to include the refactor it noticed on the way.
The last one matters more than it sounds. An agent that improves things it was not asked to improve produces pull requests that are harder to review than the work they replaced.
Not in applying patches. In knowing which local modifications are load-bearing.
A fork accumulates changes, some of which are deliberate and some of which are accidents nobody removed. An agent cannot tell the difference from the code alone, and a conflict resolution that quietly discards a deliberate change is the failure that costs the most.
The mitigation is documentation of intent: a record of why each local modification exists, maintained alongside the fork. Organisations that have this get substantially better results, and the exercise of writing it is usually valuable on its own.
The useful criteria are specific. Did the test suite pass. Were all local modifications preserved or explicitly flagged as dropped with a reason. Was the scope confined to the upstream changes. Is the description accurate against the diff.
Those are all checkable automatically, which is what allows evals to score every run rather than a person reviewing every run with equal care.
It does not eliminate the work. It changes it from a day of mechanical effort to an hour of review, and it makes the cadence regular rather than occasional, which is where most of the benefit is. Small, frequent merges are dramatically easier than the annual catch-up.
The Software Factory pages cover how backlog work becomes review-ready pull requests. The Agent Orchestration System pages cover evaluation, policy and cost attribution.