Skip to main content

Deterministic Sub-Agent Orchestration using Claude Mods

A delegation rule written into the prompt is a recommendation: the model weighs it against everything else in the context and can override it. A mod enforces the same rule as code on every agent call. Every task goes to the subagent whose tools fit it and runs on the cheapest model that can solve it.

Part 1: The mod brings every task to the right specialist and model

Specialized subagents deliver better work at lower cost

Three things define a subagent: its prompt with instructions, skills and house rules, its tools, such as the Angular CLI’s MCP server and Chrome DevTools, and its model. A specialist with the right configuration works better and cheaper than a general-purpose agent.

Four cards compared: the orchestrator on Opus plans with the tools Agent, Read, Glob, Grep and Bash; the fact-gatherer on Haiku builds a read-only inventory, with Write, Edit and Bash struck out; angular-expert on Sonnet builds with angular-cli, chrome-devtools and the skills angular-conventions and ui-ux-pro-max; deployment-engineer on Sonnet deploys with the gh CLI, the secrets store, GitHub Actions and deploy-hetzner-box

The orchestrator plans on Opus, the fact-gatherer builds an inventory of what already exists on Haiku, and angular-expert and deployment-engineer work on Sonnet with the tools and skills of their stack. The cheap, fast Haiku 5.5 is enough for an inventory. Sonnet handles implementation and tests within one stack, Opus handles planning and creative work.

Three columns with the model a task typically needs: Haiku with one dollar sign for an inventory of what exists, with fact-gatherer; Sonnet with two dollar signs for implementation and tests in one stack, with angular-expert, dotnet-expert, hugo-expert and playwright-expert, angular-expert highlighted; Opus with three dollar signs for planning and creative work, with content-writer and media-creator

A simple two-tier app shows how this works, from the request to the checked result: “Add field X to list Y”. An inventory on Haiku collects the facts the contract depends on, one specialist on Sonnet per tier implements that tier’s half, and Opus plans and checks. Every step runs on the cheapest model that can handle it. The contract sits in a shared ledger, and every agent reads the same facts from it. No detail gets lost as summaries pass from agent to agent, and a process that breaks off, for instance because it runs out of tokens or its cache expires, picks up again from the state in the ledger.

Flow in four steps: the orchestrator on Opus sends the inventory to the fact-gatherer on Haiku, which reports every place the field occurs as path:line; the orchestrator records the contract in the shared ledger, owner as string or null, at most 100 characters, JSON name owner; then dotnet-expert for the API tier and angular-expert for the client tier work in parallel on Sonnet; their claims are checked on the main thread against diff and build

The flow has four steps. First the fact-gatherer on Haiku finds every place in the code where the field occurs and reports it with file and line. From that, the orchestrator fixes the contract in the ledger, in this example a field owner, text or empty, at most 100 characters. Then dotnet-expert in the API tier and angular-expert in the client tier implement it in parallel on Sonnet. Finally the main thread checks their claims against diff and build.

Choosing the agent means choosing its model, with its cost, and its expertise: the system prompt and tools, such as the Angular CLI for angular-expert or the gh CLI for deployment-engineer, decide how good the work turns out.

The model decides whether a delegation rule applies

The model decides, based on prompt text, whether Claude Code delegates a task and to whom. An agent description only raises the likelihood of a delegation. So Angular work that angular-expert would do on Sonnet ends up with a general-purpose agent. That agent inherits the orchestrator’s Opus model and lacks the right tools, such as the Angular CLI’s MCP server and Chrome DevTools. The result ignores the standards set in the specialist, is often wrong, and spends tokens inefficiently.

Terminal transcript: both agents, dotnet-expert and angular-expert, run on Sonnet, and the CLAUDE.md requires work to go to the responsible expert; the user asks for field X in list Y in the .NET API and the Angular client; Claude splits the request into one prompt per tier, gives the API to dotnet-expert on Sonnet and the client prompt to a general-purpose agent that inherits Opus; asked why angular-expert did not take the client, Claude answers that the rule in the CLAUDE.md is an instruction it weighs against the rest of the context, and nothing in the harness enforces it

In the example, the CLAUDE.md requires work to go to the responsible specialist. Claude gives the API to dotnet-expert on Sonnet and the client prompt to a general-purpose agent that inherits Opus; angular-expert never starts. Asked about it, Claude explains the reason itself: the rule in the CLAUDE.md is an instruction it weighs against the rest of the context, and nothing in the harness enforces it.

Hooks keep the CLAUDE.md short and repeat its rules before every prompt

Before mods, the delegation rules sat in the CLAUDE.md like a constitution, and hooks enforced them. claude-md-guard.sh made sure the Claude Constitution did not get bloated. delegation-reminder.sh brought the skill check and the delegation rules back into the context before every prompt, to counter context decay: in a long session, early instructions lose weight the more new text enters the context.

Left, the delegation rules in the CLAUDE.md, five enforceable and three that need judgement; right, the two hooks: claude-md-guard.sh runs after every change to the CLAUDE.md and blocks it at more than 120 lines, missing sections, inventories or src folders that do not exist; delegation-reminder.sh adds the rules again before every prompt; a red note says neither hook runs at the agent call, so the model still picks the subagent itself

Five of the rules in the CLAUDE.md can be checked and therefore enforced; three need judgement.

A mod enforces the delegation rules as code on every agent call

In Claude Code the flow runs from prompt through orchestrator and agent call to the subagent; the mod sits in the flow between agent call and subagent, holds the spawn and picks the subagent; agent files and routing.json feed the mod from above

The mod intercepts every agent call before the subagent starts. It holds the spawn and determines the target agent from the agent definitions and routing.json. Every potential delegation error lands in the log, and our learning loop makes sure it does not happen again. That way every session runs with the same rules. What else mods do in the harness is covered in the article Optimizing Agentic Software Engineering With Mods .

Before a subagent starts, the mod hands it four things: the prompt with your request, the part of the contract it is to implement, a boundary block with the actions it must not perform, and, where needed, the reroute to the responsible specialist. Every subagent starts under the same rules, however long the session has been running.

The card shows what the mod inserts at agent.spawn: Prompt, the orchestrator's request, traced back to your words; Contract slice, the shared contract with its path:line facts from the ledger; Boundary block, no checkout, stash, reset, commit or push, report side effects; Reroute, a generic agent with this work becomes angular-expert; adjustable through keywords and gates in .claude/routing.json

The mod enforces three rules as code: it holds a generic spawn when a specialist owns the work, it blocks gated agents you did not name, and it sends every inventory to the fact-gatherer on Haiku. It adds the agent roster to the system prompt and the boundary block to every prompt; both remain instructions to the model. It records each subagent’s request, prompt and claim in the ledger for the orchestrator.

Three columns with what the mod makes of each rule: as code at agent.spawn it holds a generic spawn that a specialist takes over, blocks a gated agent you did not name, and sends inventories to the fact-gatherer on Haiku, which no rule had asked for; at every prompt and spawn it adds the agent roster and the boundary block as instructions; for the coordinator it records request, prompt and claim in the session ledger

The same prompt reaches the same specialist on every spawn

Back to the example with the new field: the prompt for the Angular client that had ended up with the general-purpose agent now passes through the mod. The mod holds it and reroutes it to angular-expert on Sonnet.

The prompt 'Show field X in the list Y view of the Angular client, with Vitest specs' as Claude gives it to a general-purpose agent; the mod marks Angular and Vitest, reads the prompt and routes the agent call to angular-expert, which runs on Sonnet with the skills and house rules of its stack; the branches to the fact-gatherer on Haiku and to the general-purpose agent on Opus are greyed out; below, the matches per specialist, with angular-expert

The mod reads the client prompt, recognizes Angular and Vitest in it and routes the agent call to angular-expert, which works on Sonnet with the skills of its stack. You confirm the reroute.

The table shows five prompts and the agent the mod routes each one to. A prompt that changes code goes to the specialist for its stack, an inventory to the fact-gatherer on Haiku. When no specialist owns a task, it stays with the general-purpose agent.

ExamplePromptTarget agent
Client tier“Show field X in the list Y view of the Angular client, with Vitest specs”angular-expert
API tier“Add field X to list Y in the .NET API, with an EF Core delta script and an xUnit test”dotnet-expert
Inventory“Find out where list Y is defined and used, in the API and the client”fact-gatherer
List without changes“List every file that touches list Y, do not edit anything”fact-gatherer
No owner“Rename field X in the helper script”general-purpose

Because the mod routes by the same rules on every agent call, the same prompt always reaches the same specialist on the same model. That lets us assume deterministic orchestration.

Decision tree: a task from the orchestrator on Opus first meets the question whether it only searches, lists or locates; yes leads to the fact-gatherer on Haiku; no leads to the question whether a specialist's expertise fits the prompt; no leads to the general-purpose agent that inherits Opus, the expensive fallback; yes leads to the question whether the work is long texts, images or video; yes leads to a producer on Opus, no to a stack expert on Sonnet

Every task from the orchestrator goes through the same questions. If it only needs to find or list something, it goes as an inventory to the fact-gatherer on Haiku. If a specialist’s expertise fits, the kind of work decides: code goes to a stack expert on Sonnet, long texts, images and video go to a producer on Opus. If nobody owns the task, the general-purpose agent on Opus takes over, the most expensive path.

Part 2: Optimizations cut costs and delegation errors further

The optimization of our harness architecture goes further than this. Three more steps cut costs and delegation errors:

The graph ledger saves more with every additional consumer of a contract

Every session keeps a shared agent ledger, the file agent-ledger.md in the session’s scratchpad. Only the main thread may write to it. Per agent it records the prompt; the claim, meaning the agent’s report that the work is done; the proof that someone checked that report; and the side effects, meaning everything the agent changed outside the edited files, such as servers started, packages installed or new rows in the database. Subagents started later get their prompt from this ledger. What one agent has already checked, or what turned out to be a dead end, no other agent has to work out again, and you pay the tokens for it only once.

The simplest ledger is a Markdown file, a log-like list in which each entry sits below the previous one. It knows nothing of relationships between the entries. Which fact belongs to which contract, and which specialist needs which part of it, the orchestrator works out itself and repeats by hand in every prompt. A graph stores these relationships along with the entries: facts, contract and agents are linked, and each agent gets exactly the part it needs. GraphRAG and agent memory work the same way.

The owner field run in two lanes across five steps, inventory, contract, API tier, client tier and verify: in the default with keywords and a Markdown ledger, the inventory goes to Haiku through the words find out, the contract is typed into every prompt by hand, you confirm the reroute to dotnet-expert, the client prompt without a stack word runs as a general-purpose agent on Opus, and the claim stays unchecked until verify; with the graph ledger, 21 fact nodes are created, from them a contract recorded once, a contract slice goes to dotnet-expert and angular-expert, and the ledger check confirms the diff

With a Markdown ledger, the orchestrator types the contract into every prompt by hand, and a client prompt without a stack word ends up as a general-purpose agent on Opus. With the graph ledger, the inventory yields 21 facts, the contract is derived from them once, and dotnet-expert and angular-expert get their part of it at spawn. A claim counts as checked only when the ledger check confirms the diff.

Our question was whether a graph ledger could raise efficiency even further. We expected no gain for a contract with two consumers. The measurements confirmed that and also showed that the savings grow with every additional consumer. All figures in this section come from measured runs.

We measured three tasks with three runs per variant (Markdown ledger and graph ledger). Each run was its own Claude Code session without intervention, with the orchestrator on Opus and in a fresh copy of the repository. The cost per run is the cost Claude Code reports for the session; we added up the tokens per model from the transcripts of all agents. Equal quality means every layer contains the change and every build is green. Runtime tests were not part of the measurement, and one client whose build already failed before the runs was left out of the assessment.

We also tested Jev as a decision layer that picks the agent for a prompt before the LLM does. In our runs Jev brought no performance benefit, so we did not include those measurements in this article.

The graph ledger in four steps: the fact-gatherer on Haiku reports facts, every path:line becomes a fact node, anchored ones keep a hash; the orchestrator on Opus records a contract, derived from the facts, for each consumer and its folder, on the main thread only; at spawn, dotnet-expert and angular-expert on Sonnet get the contract slice, and the orchestrator no longer repeats the contract; the ledger check requires each folder's diff to add the field owner before a claim counts as checked

The graph ledger works in four steps. First the fact-gatherer on Haiku reports what it found in the code, and each finding becomes a fact of its own in the graph. From these, the orchestrator on Opus derives the contract once and decides which agent implements which part of it. At spawn, dotnet-expert and angular-expert on Sonnet get only their part, the contract slice, and the orchestrator no longer has to write the contract into every prompt. Finally the ledger check looks at each folder’s diff to see whether the new field was really added. Only then does an agent’s claim count as checked.

Measurement on the owner field in a two-tier app with API and client, orchestrator on Opus, three runs each, graph ledger against Markdown ledger: cost per run 0.85 against 0.87 dollars, minus 2 percent; wall time 170 against 139 seconds, plus 22 percent; Opus tokens 536k against 555k, minus 3 percent; orchestrator prompt length 5.5k against 5.8k characters, minus 6 percent; below, three tiles: equal quality, because in every run each layer received the change and both builds were green; 6 percent shorter prompts, because the hook appends contract and facts; 2 percent lower cost within the noise, so the graph ledger pays off only with more consumers, repeated spawns and long sessions

With two consumers, the 2% cost difference is within the noise, a run with the graph ledger takes 22% longer, and quality is the same in both variants. The fact-gatherer found 21 facts per run on average, meaning places in the code that the contract builds on.

Measurement on a larger task, an owner field through HTTP API, MCP tools, database schema and Angular client, three runs each, graph ledger against Markdown ledger: cost per run 0.92 against 1.00 dollars, minus 8 percent; Sonnet tokens 312k against 430k, minus 28 percent; Opus tokens 570k against 623k, minus 8 percent; wall time 167 against 142 seconds, plus 18 percent; below, three tiles: equal quality here too; with the graph ledger 3 of 3 runs kept the inventory on Haiku, with the Markdown ledger 1 of 3; a measurable effect with minus 8 percent cost and minus 28 percent Sonnet tokens at 25 seconds more wall time

The cost advantage grows with the number of consumers. With two agents on the same contract, both ledgers are almost level; with four, a run costs 1.49 dollars with the graph ledger and 1.65 dollars with the Markdown ledger.

Line chart of cost per run by number of consumers per contract. Measured: with two consumers 0.87 dollars with the Markdown ledger and 0.85 dollars with the graph ledger, with three 1.00 and 0.92 dollars, with four 1.65 and 1.49 dollars. Dashed, the linear projection: with six 2.34 and 2.05 dollars, with eight 3.12 and 2.69 dollars.
ConsumersMarkdown ledgerGraph ledger
2$0.87$0.85
3$1.00$0.92
4$1.65$1.49
6 (projection)$2.34$2.05
8 (projection)$3.12$2.69

The graph ledger saves 2% with two consumers, 8% with three and 10% with four. The linear projection reaches 13% with six and 14% with eight. Three tasks are few data points, which is why the dashed part is a projection. In every run of all three tasks, each layer received the change and both builds were green. For the task with four consumers, the costs of the individual runs overlap, so the 10% there is a trend.

Cost advantage of the graph ledger by number of consumers, meaning the agents that use the same contract: at 2, API and client, the Markdown ledger, because the graph ledger's minus 2 percent cost is within the noise and it needs 22 percent more time; at 3, API, MCP, schema and client, the graph ledger with minus 8 percent cost and minus 28 percent Sonnet tokens; at 4, two APIs and two clients, the graph ledger with minus 10 percent cost, minus 34 percent Opus tokens and minus 18 percent wall time; at 6 to 8, a projected saving of 13 to 14 percent; quality was the same in all 18 runs

With two consumers, the Markdown ledger is therefore enough; from three consumers on, the graph ledger pays off.

Delegation learning loop

A harness that learns from its own mistakes is state of the art today. During the session, the mod catches a delegation error at spawn and holds it for you. After the session, the error goes from the log to claude-learn. The skill corrects the agent’s trigger, routing.json or the mod itself, and the next session routes the same prompt correctly.

Three cards, each with one error in the session: Claude gives the client prompt to a generic agent; without the mod, Opus does Sonnet's work without Angular tools, with the mod the spawn is held and rerouted to angular-expert or run; Claude starts the Playwright expert for a page check; without the mod, browsers open and a long test suite runs, with the mod it first asks run or block; an agent reports done; without the mod, the answer scrolls away, with the mod there is a ledger row and a card in the /delegation pane, and the claim stays unverified until you press Verify

The mod catches three typical errors during the session. If Claude gives the client prompt to a generic agent, the mod holds the spawn and offers the reroute to angular-expert. If Claude starts an expensive agent like the Playwright expert that nobody asked for, the mod asks first. When an agent reports “done”, the claim lands in the ledger and stays open until you check it.

The learning loop in four stations: Error, the inventory went to dotnet-expert on Sonnet because a repo name matched MCP Servers, and the Angular half ran as a general-purpose agent on Opus; Captured, a staged error was held, rerouted to angular-expert and recorded as an open row; Learned, the user asked to learn from the delegation error, and claude-learn sorts every error by cause; Fixed, version 0.3.0 checks inventories first, a global angular-expert on Sonnet exists, and four tests replay the error; the next session routes the same prompt correctly

After the session, claude-learn fixes the cause. In the example, an inventory ended up with dotnet-expert because a project name matched the keyword MCP Servers, and the Angular half ran as a general-purpose agent on Opus. Since the fix, the mod checks inventories first, there is a dedicated angular-expert on Sonnet, and four tests replay the error so that it stays fixed.

An agent call passes the mod; when the mod detects an error, it writes an open row to delegation-errors.jsonl; after 20 minutes without input, learn-queue starts the claude-learn skill, which sorts the open rows by cause: a trigger that is too broad is narrowed in the description or in routing.json, one that is too narrow gets trigger phrases, a missing agent is created, an orchestrator error becomes a lesson when it recurs, and an error in the mod is fixed in register.ts with a test; the row is closed as fixed or dismissed, and the next session routes with the corrected files

Every detected error lands in the log as an open row. As soon as nobody has typed anything for 20 minutes, learn-queue starts the claude-learn skill. It sorts the errors by cause: a trigger that is too broad gets narrowed, one that is too narrow gets additional trigger phrases, a missing agent gets created, and an error in the mod gets fixed with a test. The next session already works with the corrected files.

The /delegation pane shows each subagent’s model, cost and open claims

Working with the orchestrator, you see the main thread. The subagents work in the background: which agent took which prompt, on which model, at what cost, and whether its “done” was checked is buried in the transcript. The mod records every spawn, every prompt, every claim and every model request anyway. The /delegation pane puts that record next to the conversation. In the terminal it is a plain text list. In the Code tab of the desktop app, the pane draws a header with the tiles Cost, Tokens, Time and Unverified, one card per agent, grouped by Running, Claimed and Verified, and a Decisions card with the reroutes.

Windows Terminal with the Dracula skin: on the left, the Claude Code session with the prompt Add an optional supplier field to the stock item list, in the .NET API and in the Angular client; the fact-gatherer collects the facts about the stock items, and /delegation opens the pane; on the right, the /delegation pane with the fact-gatherer on claude-haiku-5-5, running, with request and prompt; a red arrow leads from the command to the pane

Before any code changes, the fact-gatherer reads where stock items occur in the API and the Angular client. It works read-only and on the cheapest model; the experts start only once the contract built from its facts is settled. The band above the prompt shows one running agent and the specialists the mod found for the prompt. The arrow leads from the command to the pane: one line per agent with model, status, your request and the prompt it received.

In the Claude desktop app, the mod draws the same pane with its own UI. The desktop app can show elements the terminal cannot display, such as SVG graphics, so panes there can be far richer than in the terminal.

The Claude desktop app in the Code tab with the session Supplier field for stock items: on the left, the prompt, /delegation and the band with the specialists found, angular-expert, dotnet-expert and angular-engineer; on the right, the Delegation pane with the orchestrator, 1 agent and 0 decisions, the tiles Cost, Tokens 60k, Time 0:09 and Unverified 1, under Claimed the card stock-facts with a magnifying glass, Haiku in green, fact-gatherer, 60k tokens and 0:09, request, prompt and a Verify button

The header tiles add up cost, tokens and time of all subagents so far; Unverified counts the claims nobody has checked yet. The costs are an estimate: the mod prices the tokens of each model request at the rate of the model that handled it. The magnifying glass marks the fact-gatherer, and the model is color-coded: Haiku green, Sonnet blue, Opus amber. Each card shows its own tokens, cost and time, then your request, the prompt and, as soon as the agent answers, the first line of its answer. A finished agent stays under Claimed until you check its claim against the files and press Verify; then it moves to Verified.

How this article was made: Alexander Kastil selected and supervised the benchmarks, the text was written with AI assistance, and Alexander Kastil then reset the focus and corrected errors.

  • Mods overview : what a mod is, what it can change and where it runs; it draws panes in the terminal and in the Code tab of the desktop app.
  • Create custom subagents : agent files with model, tools and description, which set the price tier and expertise of each specialist.
  • How we built our multi-agent research system : Anthropic on a lead agent that coordinates parallel subagents, with prompts, tool selection and evaluation.
  • Building effective agents : Anthropic on orchestrator-workers and other composable patterns, and when a workflow fits better than an autonomous agent.
  • Hooks reference : settings hooks such as PostToolUse and UserPromptSubmit, which the delegation rules relied on before mods.
  • React to events : guarding or rewriting a tool call and following a turn, the mechanism behind holding and rerouting an agent call.
  • Draw in the interface : panes, the band above the prompt, buttons and state, as the /delegation pane and its Verify button use them.
  • Plugins overview : plugins, marketplaces and install scopes, through which a marketplace distributes the mod and its fixes to every repository.