The leader as conductor of agent work
Assignment, result, evidence and decision. Three documented cases show how to direct agent work and recognise its actual completion.
2026-09-26 · 1.0 · PL / EN
01 / A task needs an ending
A leader asks for a publication to be corrected. One agent prepares the text, another checks the numbers and a third saves the material. All report completion. Does the reader see the correct version? We do not know yet. Someone needs to connect the assignment, the result and evidence of completion.
That is what we mean by conducting: defining what must be produced, who may act and how completion will be established. Our April 2026 draft Hermes handbook is the starting point. This public edition retains the method while updating its boundaries. Titles such as COO or CTO assigned to agents describe process functions. They do not make a model a company executive or transfer human accountability to it.
02 / Three records instead of a chorus of reports
The first record is the assignment: objective, versioned input, permitted actions, constraints and acceptance criteria. The second is the result: a specific file or change, with an account of what was and was not done. The third is the decision: acceptance within a stated scope, a correction or a stop with a named gap.
An example assignment might say: “Prepare an English version of this edition. Preserve the numbers, qualifications and links. Return the file and a list of ambiguities. Do not publish.” This is an editorial template, not an account of running an additional agent for this article. The reviewer knows what to inspect; the worker knows the boundary of the task.
A document helps preserve continuity, but formal storage does not make it true. A current instruction from the authorised owner may change an earlier agreement. Record the change and resolve the conflict instead of letting an agent choose the more convenient version. In particular, an older plan must not override a pause.
03 / Case: a reported successful save
In HTTP 200 is only the beginning, five local trials were performed on a fictional data store. One covered an ignored field: the response acknowledged the request, but a separate read revealed the old value. Another detected changes to categories that were meant to remain untouched. This was not a production CMS test.
The lesson for the conductor is to define acceptance by the expected state, rather than the worker's report alone. Assign changes to specific fields and preservation of the rest. Require a read after writing and a comparison. If an operation's outcome is unknown, establish what happened first. Creating the object again may add a duplicate instead of fixing a missing response.
04 / Case: a passing test with narrow scope
In What code checks and what people review, a local validator accepted a record containing an invented source identifier. It behaved as the test expected: it checked structure, not source existence. Seven passing trials did not establish seven true facts.
The conductor should therefore ask: “What exactly does this test check?” Valid syntax, numerical consistency, permission to act and factual accuracy are different scopes. The report should name them. Another agent may catch an error, but agreement between two answers does not become evidence by itself. If the author performed the review, call it an author review rather than implying an independent assessment.
05 / Case: a person wants to take over
Owner knowledge that changes code compares a conversation with a historical interface change. The owner needed to take a task over from running automation after learning the cost of interruption. Code inspection confirmed the change, but did not establish deployment or correctness of the entire execution path.
This is an important boundary in directing agents. A reserved task needs handover rules: who may take over, what happens to work in progress and how the final state will be confirmed. An on-screen warning is part of that route, not the whole acceptance process. The test plan must also cover cancellation and a failed handover. Avoid leaving two workers each believing they exclusively own the change.
06 / Delegate independent parts
Current Hermes delegation documentation describes a child's separate context and the parent passing its objective and necessary information. Do not assume a new worker knows the entire conversation. The same documentation warns that a directory shared by default can cause editing collisions. Available isolation mechanisms must be matched to the actual configuration.
Divide work so that one part does not depend on another's as-yet-unknown output. Research into two independent sources can be combined later; review of a completed translation starts after its version is available. Name the person or role that integrates results and resolves differences. This is our organisational rule, not a speed measurement.
Hermes tool documentation distinguishes toolsets for files, browsing, memory and automation. A documented feature is not proof that it is available in a particular session. Check tools and permissions before assigning work. Do not attribute every product capability to every agent.
07 / When to stop the orchestra
Set time, cost and scope limits before starting. Name the situations requiring a stop: conflicting sources, missing media rights, an unresolved write or a result outside the assignment. A gap report is useful when it states exactly what is needed for the next step.
For a small task, one worker and clear acceptance may be enough. More agents also mean context handoffs and integration; we promise no savings without measurement. Completion should identify the accepted result, remaining limitations and next decision. To apply this in your own workspace, we can start with one recurring assignment. A neutral example is enough for a conversation about roles and acceptance.
08 / Sources and AI contribution
The full April 2026 draft handbook was read. Vendor documentation was checked on 26 September 2026; no Hermes installation test or model comparison was performed. The three cases link to public editions dated 25 September and their recorded evidence. The historical tests and deployments described there were not repeated today.
The method is a proposed way to organise work, not evidence of agent-team performance. This public edition contains no private topology, sessions or client data. Codex prepared the Polish/English text, original diagram and author review. There was no independent second-model review. A person may withdraw the publication. Changes to tool documentation, corrections to a cited case or a discovered flaw in the method trigger another review.