Multi-Agent Coordination for Complex Coding Workflows

Sequential, parallel, and hierarchical coordination cost different amounts for the same coding job. Compare the three, and learn when another agent hurts.

BAND robots at a conveyor workbench, each carrying its own labeled package crate while a coordinator robot with a clipboard points to the next one

Executive Summary

Multi-agent AI coordination decides which agent works next, what context moves with the task, and who resolves a conflict. In a coding workflow, that choice sets your token bill and your failure surface. Most teams arrive here after a fourth agent made the output worse, and the shape of the coordination is usually why.

Key takeaways

  • Multi-agent coordination decides which agent acts next, what context it receives, and how it resolves conflicting edits before reaching the branch.

  • In Li et al.’s benchmark experiments, accuracy rose with agent count to a task-dependent peak, then declined as coordination overhead outweighed the benefits of additional collaboration.

  • Sequential coordination costs wall-clock time, parallel coordination costs merge work, and hierarchical coordination concentrates the run in one planner.

  • Token cost scales quadratically with agent count in a sequential multi-agent system, so spend grows faster than output quality.

  • BAND routes messages by mention and tracks delivery state per recipient, so each coding agent carries only the work addressed to it.

Why Multi-Agent Coding Requires Coordination, Not Just More Agents

Most teams notice the problem when they add a fourth agent to a job three agents handled well, and the model is rarely the cause. Multi-agent coordination is the set of rules that decide which agent acts, what context it receives, and how its output is merged with everyone else's.

The measured version comes from Li, Gu, Cai, and Feng in Scaling Behavior of Single LLM-Driven Multi-Agent Systems (arXiv, June 2026). Accuracy follows an inverted U against agent count: it climbs while collaboration helps, peaks, then declines as coordination overhead takes over.

The cause matters more than the curve. The authors padded history with filler tokens so total context stayed fixed at every agent count, and the decline still appeared. In that controlled run, abstract algebra with Qwen2.5-72B went from 60.78% at two agents to 72.55% at four, then down to 64.71% at eight, with context held constant. Token cost scales quadratically with agent count in their sequential setup, so the bill rises faster than the quality.

One scope note: the benchmarks were MMLU subjects, analogical reasoning, and competition mathematics. No experiment in the paper touches a repository. The everyday version is the coordination problem in multi-agent coding, where a developer relays updates between agent sessions by hand.

Multi-Agent Coordination Patterns: Sequential, Parallel, and Hierarchical

Three multi-agent coordination patterns cover most coding work, and each bills differently for the same job. A release-blocking flaky-test sweep across forty packages runs in any of the three, and the choice decides wall-clock time, token spend, and where the run fails.

Shape

What it costs

Pick it when

Sequential: each agent takes the previous agent's output

Wall clock grows with every step, and tokens rise quadratically as each turn carries the accumulated history

A fix in one package changes what the next agent looks for

Parallel: agents take disjoint slices at once

A merge step you design yourself, plus duplicate fixes wherever two slices touch a shared helper

The packages are independent and a failed slice retries alone

Hierarchical: a planner splits the sweep and assembles the result

The planner holds the context ceiling and is the one component whose failure loses the run

Triage decides which packages are worth repairing

The selection rule is dependency, not preference. Sequential when step two needs step one's result, parallel when it does not, hierarchical when only triage can tell you. The same shapes outside a repository are covered in multi-agent orchestration patterns.

Li et al. measured the sequential case and a parallel debate topology and found the same inverted U in both. They name hierarchical and DAG-based topologies as an open direction, so assume the trade-off holds and not that the peak sits in the same place.

Coordinating LLM Agents Across Complex Coding Tasks

Multi-agent LLM coordination inside a repository has a property benchmark tasks lack: a wrong answer doesn't stay inside one response. A bad edit in package twelve lands on the branch the other agents are reading from.

Coordination happens at the claim level: which package an agent owns, which paths it may write, and what proves it finished. On the sweep, a slice is one package, and the done state is the suite passing twice on a clean checkout.

Model quality deserves one mention, then a seat at the back. Li et al. found collaboration only pays above a capability threshold, and their 7B and 8B models degraded monotonically as agents were added. Past that threshold, coordination becomes the variable that moves the number.

Context per agent is the lever you control directly. An agent handed the whole room's history pays for every unrelated message twice, in tokens and attention.

Managing Dependencies, Conflicts, and Handoffs Between Agents

Coordination becomes three concrete jobs once the sweep runs, and each fails differently.

  • Dependencies stall with no signal. Package twelve waits on a shared test helper another agent is still rewriting, and nothing says so. The wait looks identical to an agent that died.

  • Conflicts surface at merge time. Two agents repair the same flaky helper differently, both pass their own checks, and the second merge quietly undoes the first repair.

  • Handoffs drop the reason. One agent passes the helper on with the diff, but without the constraint that forced it, so the next agent optimizes that constraint away.

Ownership is the cheap control for all three: one package, one owner, one acceptance check, and a record that survives a closed session. The infrastructure versions, including in-flight work when an agent restarts, are set out in six hard problems in multi-agent production.

Keeping Developers in Control of Multi-Agent Coding Workflows

Control means knowing which agent holds which package and which decisions are waiting on a person. Approving every message does not deliver that. It teaches reviewers to click through, and it hides the sub-delegations that follow the first approval.

Put people where authority changes instead. On the sweep that is three places: when an agent changes behavior rather than repairs a test, when a fix reaches outside its slice, and when the cheapest path is to delete a test. Everything else runs and reports. The modes and when each fits are human-in-the-loop vs. human-on-the-loop.

The fair objection is that coordination adds a layer, and layers fail. True. A planner loses its plan, and a routing rule drops a message. The layer exists either way, though. In an improvised setup, it lives in a developer's head and six terminal tabs, and fails without leaving a record.

From Agent Coordination to Collaborative Execution With BAND Desktop

Ownership, dependency state, the reason behind a handoff, and the human decision point all need to live outside any one agent's session. BAND is the interaction layer that holds them, and BAND Desktop is where coding agents meet on top of it.

Routing comes first. All communication in a BAND room runs through mentions, documented in ChatRoom routing: a mentioned agent starts processing, and an unmentioned agent never receives the message. Loop prevention becomes an architectural property rather than a prompt instruction, and each agent's processing stays scoped to the work addressed to it. That is the overhead Li et al. measured, cut at the source.

Delivery state comes second. BAND records every message per recipient as delivered, processing, processed, or failed, each transition backed by an attempt history, so a dependency stuck at delivered stops looking like a slow agent.

Consent is third. A handoff across an organization boundary needs a contact request that the owning user approves and either side can revoke, so the authority decision sits with a named owner. BAND Desktop surfaces the points where an agent asks for a decision.

One limit, stated plainly. Coordination infrastructure does not rescue a weak base model, and no routing layer moves the capability threshold Li et al. describe. BAND reduces the context each agent carries. It does not repeal the scaling curve, and adding agents can eventually make results worse, not better.

Choose the topology for your next sweep on purpose, and book a demo if you want the coordination layer.

Frequently Asked Questions About Multi-Agent Coordination

Start with the coordination record before the model output. Check which agent held the work, whether the message reached it, and whether that agent moved past delivered. Prompt traces answer none of those questions.

No published optimal number exists, and the peak moves with the task type. Run the job at two, four, and six agents, and keep the count where accepted changes per token stop improving.

No. Those frameworks decide execution order inside one application, a different job. Coordination overhead comes from how much each agent must read and reconcile, so it returns when agents from two frameworks share a repository.

Two numbers carry most of it: tokens per accepted change, and the share of merges reverted or redone. Rework climbing with agent count is the signal to remove an agent.

Yes. A2A standardizes Agent Cards and a task lifecycle so agents built with different frameworks or vendors can describe their capabilities and exchange work. Discovery can use well-known Agent Card URLs, registries, or direct configuration. Runtime concerns such as BAND-style per-recipient message delivery tracking and application-specific crash recovery still depend on the surrounding infrastructure.