Build log / Operating method 002
The second agent needs an interface.
The moment a second AI enters the system, the problem changes. It is no longer only “can the model do the task?” It becomes “who owns the state, who decides, and how does the work stop?”
Agent count is not autonomy
Our first public role was the Publisher: observe verified activity, preserve the record and turn eligible work into something another person can inspect. The second role is the Orchestrator: decide what work should happen, route it, check the result and close the loop.
That sounds like progress. It can also create a new tax. Two capable agents can still duplicate work, wait on one another, overwrite state or escalate routine decisions to a human. The extra intelligence is useful only if the interface between the roles is explicit.
Observe
Verify
Publish
Decide
Route
Verify
The minimum contract
- STATE What has actually happened, with evidence? Plans and completed work cannot share the same status.
- OWNER Which agent has authority over the next decision? Shared responsibility usually means repeated work.
- HANDOFF What exact artifact or event moves the task forward? “Review this” is weaker than a defined input and expected output.
- STOP What closes the loop? A published URL, a verified event or a failed quality gate is a terminal state. Another summary is not.
Why we are starting small
OpenAI's practical guidance recommends keeping a single-agent system until added specialization justifies multi-agent complexity. Anthropic's account of its research system reaches the same boundary from the other direction: parallel agents can improve complex research, but coordination, evaluation and reliability become engineering problems of their own.
Our interpretation is narrow: add a role only when it owns a distinct decision or measurable piece of work. Then make the interface observable. A second agent should remove human coordination, not merely relocate it.
It does not claim that the two-agent system improved output, reduced cost or lowered Human Minutes per Transaction. Those are results. They require evidence and the experiment's observation window.
The next test
The architecture now has a public contract. The next question is measurable: when Publisher produces an eligible artifact, can Orchestrator validate, distribute and verify it without asking a human to repair the loop?
That result will be published only after it is old enough to qualify as evidence.
Sources that shaped the method
- OpenAI, A practical guide to building agents — incremental orchestration and single-agent vs multi-agent patterns.
- Anthropic, How we built our multi-agent research system — orchestrator-worker design and coordination challenges.
Audit your own agent graph: can you name the state, owner, handoff and stop condition for every edge? If one is missing, that edge is probably human work in disguise.