Blind judging · position-debiased · multi-run averaged · living document

When do multi-agent pipelines actually beat a single prompt? We measured.

Blind evaluation results for Agency Orchestrator (AO), an open-source multi-agent workflow engine. We publish the losses along with the wins — including the tiers where multi-agent is worse. Reproduce it yourself: npm run eval in the repo.

The question

AO's core bet: is the output of a multi-role DAG pipeline actually better than what a user gets from one well-formed prompt? Marketing says yes. We wanted data.

Method

Results

Strong generator (Claude on both sides), high-confidence subset

Template Multi-agent One-shot baseline Winner
story-creation 9.0 8.0 ✅ multi-agent
tech-blog 8.0 9.0 ❌ baseline
ai-opinion 8.0 9.0 ❌ baseline
product-review 7.5 8.5 ❌ baseline

→ Multi-agent goes 1–3: roughly a tie, leaning negative.

Weak generator (ollama/llama3 8B on both sides), 3-run average

Template Multi-agent One-shot baseline Winner Stability (multi wins / total)
story-creation 5.0 6.3 ❌ baseline 0/3
tech-blog 3.0 4.5 ❌ baseline 0/3
ai-opinion 2.5 5.2 ❌ baseline 0/3
product-review 4.7 4.2 ✅ multi-agent 2/3

→ Multi-agent goes 1–3, and the three losses are 0/3 — it never won those in any run. A stable pattern, not noise.

Mid-tier generator (DeepSeek — AO's actual default), 3-run average

Template Multi-agent One-shot baseline Winner Stability
story-creation 8.0 7.0 ✅ multi-agent 2/3
tech-blog 8.2 5.7 ✅ multi-agent 3/3 (high confidence)
ai-opinion 8.7 8.2 ✅ multi-agent 2/3
product-review 7.5 7.7 ❌ baseline 1/3

→ Multi-agent goes 3–1 — the opposite of both extremes. On tech-blog it wins 3/3 decisively (8.2 vs 5.7: DeepSeek's one-shot blogs had bugs and truncation; the pipeline shipped complete, publishable posts).

2026-09-25 re-run: strong tier, with acceptance criteria written as data (claude-code both sides, story-creation, n=1)

Template Multi-agent One-shot baseline Verdict
story-creation 8.0 4.5 ✅ multi-agent (both directions agree)

Same template, same tier. Earlier it was 9.0 vs 8.0 (a squeaker); now it is 8.0 vs 4.5. Both blind judges spelled out why — they were counting acceptance items:

A has no title, opens on the action of pulling up the shutter, ends on the image of the glowing watch hands; B added a "# Seven Minutes Fast" title and spelled out the theme in the last line, violating acceptance items 1, 3 and 4.

What changed in between: (1) the template now declares deliverables plus a full acceptance on the final step ("no title / within ±30% of the target length / open on a concrete image / don't state the theme"); (2) the blind judge now anchors on the deliverable step's acceptance (it used to take the last completed step's — which is a different step whenever a review step runs at the end).

Read this number with its bias attached: the baseline never saw those acceptance criteria — it only got "goal + inputs + produce the final deliverable". The multi-agent side both saw them and was pushed to satisfy them by automatic verification. That is exactly how the product is used (you write the acceptance once, the pipeline enforces it), but it is not a prompt-to-prompt comparison — for that, the baseline prompt would have to carry the same criteria.

So we ran the control on the spot: same tier (claude-code on both sides), same day, but a template with no acceptance and no declared deliverables:

Template Multi-agent One-shot baseline Verdict
tech-blog (no acceptance) 8.5 8.5 ➖ tie (low confidence: the two directions disagree)

Neither judge could carry the argument — one preferred the baseline's coverage (profiling, the GIL, rayon, NumPy, CI all covered), the other preferred the multi-agent draft's "the first version was slower" narrative with reproducible timings. Lengths differed by 2× (5,517 vs 11,739 characters).

Read the pair together and the picture is cleaner than either half: at the strong-model tier, multi-agent on its own is roughly a tie — which matches this file's earlier finding; the gap opens only once the acceptance criteria are written as data. In other words, the 8.0 vs 4.5 number measures "is this way of working worth it", not "are multi-agent pipelines inherently better". Different questions, different answers — don't quote one for the other.

Conclusion: the relationship is non-monotonic (Goldilocks)

Generator tier Multi-agent vs one-shot Why
Very weak (llama3-8B class) ❌ decisively worse (1–3, losses at 0/3) each step is low quality; the hand-off chain amplifies drift and errors instead of correcting them
Mid-tier (DeepSeek — the default) ✅ better (3–1) the model is good enough to execute each specialized step, but one-shot isn't near its ceiling → division of labor genuinely lifts quality
Strong (Claude class) ≈ tie, leaning negative one-shot is already near the ceiling; orchestration overhead buys little

Mechanism: weak models accumulate drift across hand-offs (language mixing, topic drift, hallucinated APIs, dropped drafts) — the hand-off chain is an error amplifier. Mid-tier models handle each specialized sub-task competently, so "research first, then write, then review" produces real gains. Strong models are already good enough in one shot.

The practical takeaway: AO defaults to DeepSeek, which sits squarely in the "division-of-labor pays off" sweet spot. "A cheap-but-capable model + multi-agent = better output" is supported by the data — as long as the default never drops to the very-weak tier, where multi-agent actively makes things worse. If you bring a frontier model, the honest pitch is different: quality is a wash, and the value is structured reproducibility (deterministic DAGs, per-step acceptance criteria, resume/iterate) rather than a quality lift.

Real bugs this eval loop caught in our own product

Limits (read before quoting us)

Reproduce it

git clone https://github.com/jnMetaCode/agency-orchestrator
npm install && npm run eval

Agency Orchestrator is Apache-2.0. One sentence in, a team of AI specialists out — model-agnostic (11 providers incl. local Ollama), runs entirely on your machine.