Blind judging · position-debiased · multi-run averaged · living document
When do multi-agent pipelines actually beat a single prompt? We measured.
Blind evaluation results for Agency Orchestrator (AO), an open-source multi-agent workflow engine. We publish the losses along with the wins — including the tiers where multi-agent is worse. Reproduce it yourself:
npm run evalin the repo.
The question
AO's core bet: is the output of a multi-role DAG pipeline actually better than what a user gets from one well-formed prompt? Marketing says yes. We wanted data.
Method
- Comparison: same input, run twice — once through AO's multi-agent workflow, once as a single one-shot prompt (simulating a user not using AO at all).
- Blind, position-debiased judging: the judge model doesn't know which output came from AO. Each pair is judged twice with positions swapped (A=multi, B=baseline; then A=baseline, B=multi) and scores averaged — this cancels the largest known LLM-judge bias (position). If the two directions disagree, the result is flagged low-confidence.
- Generation/judging separation: outputs are generated by the model under test; judging is fixed to a strong model (Claude).
- Multiple runs averaged: each template runs N times — single runs are far too noisy (we watched one template swing 5.5 → 9.0 → 9.0 → 8.0 across four runs).
Results
Strong generator (Claude on both sides), high-confidence subset
| Template | Multi-agent | One-shot baseline | Winner |
|---|---|---|---|
| story-creation | 9.0 | 8.0 | ✅ multi-agent |
| tech-blog | 8.0 | 9.0 | ❌ baseline |
| ai-opinion | 8.0 | 9.0 | ❌ baseline |
| product-review | 7.5 | 8.5 | ❌ baseline |
→ Multi-agent goes 1–3: roughly a tie, leaning negative.
Weak generator (ollama/llama3 8B on both sides), 3-run average
| Template | Multi-agent | One-shot baseline | Winner | Stability (multi wins / total) |
|---|---|---|---|---|
| story-creation | 5.0 | 6.3 | ❌ baseline | 0/3 |
| tech-blog | 3.0 | 4.5 | ❌ baseline | 0/3 |
| ai-opinion | 2.5 | 5.2 | ❌ baseline | 0/3 |
| product-review | 4.7 | 4.2 | ✅ multi-agent | 2/3 |
→ Multi-agent goes 1–3, and the three losses are 0/3 — it never won those in any run. A stable pattern, not noise.
Mid-tier generator (DeepSeek — AO's actual default), 3-run average
| Template | Multi-agent | One-shot baseline | Winner | Stability |
|---|---|---|---|---|
| story-creation | 8.0 | 7.0 | ✅ multi-agent | 2/3 |
| tech-blog | 8.2 | 5.7 | ✅ multi-agent | 3/3 (high confidence) |
| ai-opinion | 8.7 | 8.2 | ✅ multi-agent | 2/3 |
| product-review | 7.5 | 7.7 | ❌ baseline | 1/3 |
→ Multi-agent goes 3–1 — the opposite of both extremes. On tech-blog it wins 3/3 decisively (8.2 vs 5.7: DeepSeek's one-shot blogs had bugs and truncation; the pipeline shipped complete, publishable posts).
2026-09-25 re-run: strong tier, with acceptance criteria written as data (claude-code both sides, story-creation, n=1)
| Template | Multi-agent | One-shot baseline | Verdict |
|---|---|---|---|
| story-creation | 8.0 | 4.5 | ✅ multi-agent (both directions agree) |
Same template, same tier. Earlier it was 9.0 vs 8.0 (a squeaker); now it is 8.0 vs 4.5. Both blind judges spelled out why — they were counting acceptance items:
A has no title, opens on the action of pulling up the shutter, ends on the image of the glowing watch hands; B added a "# Seven Minutes Fast" title and spelled out the theme in the last line, violating acceptance items 1, 3 and 4.
What changed in between: (1) the template now declares deliverables plus a full acceptance on the
final step ("no title / within ±30% of the target length / open on a concrete image / don't state the
theme"); (2) the blind judge now anchors on the deliverable step's acceptance (it used to take the
last completed step's — which is a different step whenever a review step runs at the end).
Read this number with its bias attached: the baseline never saw those acceptance criteria — it only got "goal + inputs + produce the final deliverable". The multi-agent side both saw them and was pushed to satisfy them by automatic verification. That is exactly how the product is used (you write the acceptance once, the pipeline enforces it), but it is not a prompt-to-prompt comparison — for that, the baseline prompt would have to carry the same criteria.
So we ran the control on the spot: same tier (claude-code on both sides), same day, but a template
with no acceptance and no declared deliverables:
| Template | Multi-agent | One-shot baseline | Verdict |
|---|---|---|---|
| tech-blog (no acceptance) | 8.5 | 8.5 | ➖ tie (low confidence: the two directions disagree) |
Neither judge could carry the argument — one preferred the baseline's coverage (profiling, the GIL, rayon, NumPy, CI all covered), the other preferred the multi-agent draft's "the first version was slower" narrative with reproducible timings. Lengths differed by 2× (5,517 vs 11,739 characters).
Read the pair together and the picture is cleaner than either half: at the strong-model tier, multi-agent on its own is roughly a tie — which matches this file's earlier finding; the gap opens only once the acceptance criteria are written as data. In other words, the 8.0 vs 4.5 number measures "is this way of working worth it", not "are multi-agent pipelines inherently better". Different questions, different answers — don't quote one for the other.
Conclusion: the relationship is non-monotonic (Goldilocks)
| Generator tier | Multi-agent vs one-shot | Why |
|---|---|---|
| Very weak (llama3-8B class) | ❌ decisively worse (1–3, losses at 0/3) | each step is low quality; the hand-off chain amplifies drift and errors instead of correcting them |
| Mid-tier (DeepSeek — the default) | ✅ better (3–1) | the model is good enough to execute each specialized step, but one-shot isn't near its ceiling → division of labor genuinely lifts quality |
| Strong (Claude class) | ≈ tie, leaning negative | one-shot is already near the ceiling; orchestration overhead buys little |
Mechanism: weak models accumulate drift across hand-offs (language mixing, topic drift, hallucinated APIs, dropped drafts) — the hand-off chain is an error amplifier. Mid-tier models handle each specialized sub-task competently, so "research first, then write, then review" produces real gains. Strong models are already good enough in one shot.
The practical takeaway: AO defaults to DeepSeek, which sits squarely in the "division-of-labor pays off" sweet spot. "A cheap-but-capable model + multi-agent = better output" is supported by the data — as long as the default never drops to the very-weak tier, where multi-agent actively makes things worse. If you bring a frontier model, the honest pitch is different: quality is a wash, and the value is structured reproducibility (deterministic DAGs, per-step acceptance criteria, resume/iterate) rather than a quality lift.
Real bugs this eval loop caught in our own product
- 4 templates referenced roles that didn't exist in the published role library (a shipping bug users would hit immediately) → dependency pinned.
- Final-step outputs were polluted with meta-commentary → anti-pollution constraints added; one template's score went 5.5 → 9.0.
- The judge harness truncated outputs at 3.5k chars before judging — systematically punishing longer outputs → raised to 20k.
- Judge JSON parse failures silently dropped runs → retry added.
Limits (read before quoting us)
- 4 templates × 3 tiers × 3 runs: a clear signal, not an exhaustive benchmark.
- The judge is Claude — strong, but still an LLM judge; low-confidence items are flagged, and position-debiasing + averaging only reduce bias.
- The sweet spot's boundaries are unmapped: DeepSeek wins, llama3-8B loses, but the wide space between (GPT-class minis, Qwen, larger open models) is untested.
- Baseline prompts were mechanically synthesized from the workflow goal + inputs — a skilled human prompt would make the baseline stronger.
- When a template declares
acceptance, the blind judge anchors on it — and the baseline never saw it (see the 2026-09-25 entry). That is a fair measure of "is AO less work to get what you wanted", and a biased one for "are multi-agent pipelines inherently better". Know which question you're asking.
Reproduce it
git clone https://github.com/jnMetaCode/agency-orchestrator
npm install && npm run eval
Agency Orchestrator is Apache-2.0. One sentence in, a team of AI specialists out — model-agnostic (11 providers incl. local Ollama), runs entirely on your machine.