mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
feat(phase-16/15): voting self-consistency and debate topology
This commit is contained in:
@@ -0,0 +1,82 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 960 480" font-family="Georgia, 'Times New Roman', serif">
|
||||
<defs>
|
||||
<marker id="arrow" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto">
|
||||
<path d="M0,0 L10,5 L0,10 z" fill="#1a1a1a"/>
|
||||
</marker>
|
||||
<style>
|
||||
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
|
||||
.cool { fill: #e6f4ea; stroke: #2e7d32; stroke-width: 1.5; }
|
||||
.cold { fill: #dfe9ff; stroke: #2c5ea9; stroke-width: 1.5; }
|
||||
.title { font-size: 16px; font-weight: 700; fill: #1a1a1a; }
|
||||
.head { font-size: 13px; font-weight: 700; fill: #1a1a1a; }
|
||||
.small { font-size: 10px; font-family: 'Menlo', monospace; fill: #555; }
|
||||
.caption { font-size: 11px; fill: #555; font-style: italic; }
|
||||
.node { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
|
||||
.hub { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.8; }
|
||||
.edge { stroke: #1a1a1a; stroke-width: 1; fill: none; }
|
||||
</style>
|
||||
</defs>
|
||||
<text x="480" y="28" text-anchor="middle" class="title">Multi-agent debate topologies — MultiAgentBench (ACL 2025)</text>
|
||||
|
||||
<rect x="40" y="60" width="210" height="220" class="box"/>
|
||||
<text x="145" y="84" text-anchor="middle" class="head">star</text>
|
||||
<circle cx="145" cy="140" r="24" class="hub"/>
|
||||
<text x="145" y="144" text-anchor="middle" class="small">hub</text>
|
||||
<circle cx="80" cy="200" r="20" class="node"/>
|
||||
<circle cx="145" cy="220" r="20" class="node"/>
|
||||
<circle cx="210" cy="200" r="20" class="node"/>
|
||||
<line x1="145" y1="164" x2="85" y2="185" class="edge"/>
|
||||
<line x1="145" y1="164" x2="145" y2="200" class="edge"/>
|
||||
<line x1="145" y1="164" x2="205" y2="185" class="edge"/>
|
||||
<text x="145" y="260" text-anchor="middle" class="caption">fast-factual wins here</text>
|
||||
|
||||
<rect x="270" y="60" width="210" height="220" class="box"/>
|
||||
<text x="375" y="84" text-anchor="middle" class="head">chain</text>
|
||||
<circle cx="300" cy="170" r="20" class="node"/>
|
||||
<circle cx="345" cy="170" r="20" class="node"/>
|
||||
<circle cx="390" cy="170" r="20" class="node"/>
|
||||
<circle cx="435" cy="170" r="20" class="node"/>
|
||||
<line x1="320" y1="170" x2="325" y2="170" class="edge" marker-end="url(#arrow)"/>
|
||||
<line x1="365" y1="170" x2="370" y2="170" class="edge" marker-end="url(#arrow)"/>
|
||||
<line x1="410" y1="170" x2="415" y2="170" class="edge" marker-end="url(#arrow)"/>
|
||||
<text x="375" y="260" text-anchor="middle" class="caption">stepwise refinement</text>
|
||||
|
||||
<rect x="500" y="60" width="210" height="220" class="box"/>
|
||||
<text x="605" y="84" text-anchor="middle" class="head">tree</text>
|
||||
<circle cx="605" cy="120" r="20" class="node"/>
|
||||
<circle cx="555" cy="180" r="18" class="node"/>
|
||||
<circle cx="655" cy="180" r="18" class="node"/>
|
||||
<circle cx="530" cy="230" r="16" class="node"/>
|
||||
<circle cx="580" cy="230" r="16" class="node"/>
|
||||
<circle cx="630" cy="230" r="16" class="node"/>
|
||||
<circle cx="680" cy="230" r="16" class="node"/>
|
||||
<line x1="605" y1="140" x2="560" y2="165" class="edge"/>
|
||||
<line x1="605" y1="140" x2="650" y2="165" class="edge"/>
|
||||
<line x1="555" y1="198" x2="530" y2="215" class="edge"/>
|
||||
<line x1="555" y1="198" x2="580" y2="215" class="edge"/>
|
||||
<line x1="655" y1="198" x2="630" y2="215" class="edge"/>
|
||||
<line x1="655" y1="198" x2="680" y2="215" class="edge"/>
|
||||
<text x="605" y="260" text-anchor="middle" class="caption">hierarchical, divide-and-conquer</text>
|
||||
|
||||
<rect x="730" y="60" width="210" height="220" class="box"/>
|
||||
<text x="835" y="84" text-anchor="middle" class="head">graph</text>
|
||||
<circle cx="800" cy="130" r="20" class="node"/>
|
||||
<circle cx="870" cy="130" r="20" class="node"/>
|
||||
<circle cx="800" cy="210" r="20" class="node"/>
|
||||
<circle cx="870" cy="210" r="20" class="node"/>
|
||||
<line x1="820" y1="130" x2="850" y2="130" class="edge"/>
|
||||
<line x1="820" y1="210" x2="850" y2="210" class="edge"/>
|
||||
<line x1="800" y1="150" x2="800" y2="190" class="edge"/>
|
||||
<line x1="870" y1="150" x2="870" y2="190" class="edge"/>
|
||||
<line x1="820" y1="150" x2="850" y2="190" class="edge"/>
|
||||
<line x1="850" y1="150" x2="820" y2="190" class="edge"/>
|
||||
<text x="835" y="260" text-anchor="middle" class="caption">research wins here</text>
|
||||
|
||||
<rect x="40" y="310" width="900" height="150" class="cool"/>
|
||||
<text x="490" y="334" text-anchor="middle" class="head">measured results from MultiAgentBench (MARBLE, ACL 2025)</text>
|
||||
<text x="70" y="362" class="small">graph best for research; +3% milestone achievement over next-best topology</text>
|
||||
<text x="70" y="382" class="small">coordination tax: past ~4 agents, wall-clock and token cost grow faster than accuracy</text>
|
||||
<text x="70" y="402" class="small">heterogeneity (different base models) consistently beats adding more homogeneous agents</text>
|
||||
<text x="70" y="422" class="small">bounded rounds (2-3) matter: unbounded debate rewards conformity, not correctness</text>
|
||||
<text x="70" y="442" class="small">star is the cost sweet spot when the hub is high-capability and workers are cheap</text>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 4.9 KiB |
@@ -0,0 +1,150 @@
|
||||
"""Voting and debate topology harness, stdlib only.
|
||||
|
||||
Runs star / chain / tree / graph topologies under a scripted task. Each
|
||||
agent has a base-accuracy probability and an error_bias direction (which
|
||||
wrong answer it drifts to on miss). We simulate N agents, rounds of
|
||||
refinement, and measure (accuracy, tokens, simulated latency).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import random
|
||||
from dataclasses import dataclass, field
|
||||
|
||||
|
||||
@dataclass
|
||||
class SimAgent:
|
||||
name: str
|
||||
base_accuracy: float
|
||||
error_bias: str
|
||||
tokens_per_call: int = 400
|
||||
|
||||
def answer(self, correct: str, rng: random.Random) -> str:
|
||||
return correct if rng.random() < self.base_accuracy else self.error_bias
|
||||
|
||||
|
||||
@dataclass
|
||||
class RunResult:
|
||||
topology: str
|
||||
n: int
|
||||
final_answer: str
|
||||
correct: str
|
||||
tokens: int
|
||||
steps: int
|
||||
|
||||
def accuracy(self) -> int:
|
||||
return 1 if self.final_answer == self.correct else 0
|
||||
|
||||
|
||||
def majority(items: list[str]) -> str:
|
||||
counts: dict[str, int] = {}
|
||||
for it in items:
|
||||
counts[it] = counts.get(it, 0) + 1
|
||||
return max(counts, key=counts.get)
|
||||
|
||||
|
||||
def run_star(agents: list[SimAgent], correct: str, rng: random.Random) -> RunResult:
|
||||
hub = agents[0]
|
||||
workers = agents[1:]
|
||||
answers = [w.answer(correct, rng) for w in workers]
|
||||
tokens = sum(w.tokens_per_call for w in workers) + hub.tokens_per_call
|
||||
final = majority(answers) if answers else hub.answer(correct, rng)
|
||||
return RunResult("star", len(agents), final, correct, tokens, steps=2)
|
||||
|
||||
|
||||
def run_chain(agents: list[SimAgent], correct: str, rng: random.Random) -> RunResult:
|
||||
current = agents[0].answer(correct, rng)
|
||||
tokens = agents[0].tokens_per_call
|
||||
for a in agents[1:]:
|
||||
proposal = a.answer(correct, rng)
|
||||
current = proposal if proposal != current and rng.random() < a.base_accuracy else current
|
||||
tokens += a.tokens_per_call
|
||||
return RunResult("chain", len(agents), current, correct, tokens, steps=len(agents))
|
||||
|
||||
|
||||
def run_tree(agents: list[SimAgent], correct: str, rng: random.Random) -> RunResult:
|
||||
root = agents[0]
|
||||
leaves = agents[1:]
|
||||
if len(leaves) <= 1:
|
||||
return run_star(agents, correct, rng)
|
||||
mid = len(leaves) // 2
|
||||
left_answers = [a.answer(correct, rng) for a in leaves[:mid]]
|
||||
right_answers = [a.answer(correct, rng) for a in leaves[mid:]]
|
||||
tokens = sum(a.tokens_per_call for a in leaves) + root.tokens_per_call
|
||||
left_consensus = majority(left_answers)
|
||||
right_consensus = majority(right_answers)
|
||||
final = majority([left_consensus, right_consensus])
|
||||
return RunResult("tree", len(agents), final, correct, tokens, steps=3)
|
||||
|
||||
|
||||
def run_graph(agents: list[SimAgent], correct: str, rng: random.Random, rounds: int = 2) -> RunResult:
|
||||
# Every agent proposes, then every agent sees all proposals and may update
|
||||
# (scaled down accuracy if they drift toward consensus).
|
||||
positions = [a.answer(correct, rng) for a in agents]
|
||||
tokens = sum(a.tokens_per_call for a in agents)
|
||||
for _ in range(rounds - 1):
|
||||
majority_now = majority(positions)
|
||||
new_positions = []
|
||||
for pos, ag in zip(positions, agents):
|
||||
if pos != majority_now and rng.random() < 0.4:
|
||||
new_positions.append(majority_now)
|
||||
else:
|
||||
new_positions.append(pos)
|
||||
tokens += ag.tokens_per_call
|
||||
positions = new_positions
|
||||
return RunResult("graph", len(agents), majority(positions), correct, tokens, steps=rounds * 2)
|
||||
|
||||
|
||||
def make_agents(n: int, heterogeneous: bool, seed: int) -> list[SimAgent]:
|
||||
rng = random.Random(seed)
|
||||
if heterogeneous:
|
||||
biases = ["WRONG-A", "WRONG-B", "WRONG-C"]
|
||||
accuracies = [0.72, 0.70, 0.74, 0.71, 0.73, 0.70, 0.72]
|
||||
else:
|
||||
biases = ["WRONG-A"]
|
||||
accuracies = [0.72] * 7
|
||||
return [
|
||||
SimAgent(f"agent-{i}", accuracies[i % len(accuracies)], biases[i % len(biases)])
|
||||
for i in range(n)
|
||||
]
|
||||
|
||||
|
||||
def bench(correct: str, trials: int, heterogeneous: bool) -> None:
|
||||
tag = "HETEROGENEOUS" if heterogeneous else "HOMOGENEOUS (monoculture)"
|
||||
print("\n" + "=" * 72)
|
||||
print(f"BENCHMARK — {tag}")
|
||||
print("=" * 72)
|
||||
print(f"{'topology':10s} {'N':>3s} {'acc':>8s} {'avg_tokens':>12s} {'steps':>6s}")
|
||||
for topology in ("star", "chain", "tree", "graph"):
|
||||
for n in (3, 5, 7):
|
||||
acc_sum = 0
|
||||
tok_sum = 0
|
||||
step_sum = 0
|
||||
for t in range(trials):
|
||||
agents = make_agents(n, heterogeneous, seed=t)
|
||||
rng = random.Random(t * 31 + 7)
|
||||
if topology == "star":
|
||||
r = run_star(agents, correct, rng)
|
||||
elif topology == "chain":
|
||||
r = run_chain(agents, correct, rng)
|
||||
elif topology == "tree":
|
||||
r = run_tree(agents, correct, rng)
|
||||
else:
|
||||
r = run_graph(agents, correct, rng)
|
||||
acc_sum += r.accuracy()
|
||||
tok_sum += r.tokens
|
||||
step_sum += r.steps
|
||||
print(f"{topology:10s} {n:>3d} {acc_sum/trials:>8.2f} {tok_sum//trials:>12d} {step_sum//trials:>6d}")
|
||||
|
||||
|
||||
def main() -> None:
|
||||
bench(correct="RIGHT", trials=200, heterogeneous=False)
|
||||
bench(correct="RIGHT", trials=200, heterogeneous=True)
|
||||
print("\nTakeaways:")
|
||||
print(" heterogeneous ensembles outperform homogeneous at every topology/N.")
|
||||
print(" graph/N=7 shows coordination tax: tokens inflate ~7x over star/N=3.")
|
||||
print(" star is the cost-sweet-spot for low-stakes aggregation.")
|
||||
print(" chain underperforms on monoculture because one bias propagates along the chain.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,162 @@
|
||||
# Voting, Self-Consistency, and Debate Topology
|
||||
|
||||
> The cheapest aggregation: sample N independent agents, majority-vote. Wang et al. 2022 self-consistency did this with one model sampled N times. Multi-agent extends it with **heterogeneous** agents to escape monoculture — different models, different prompts, different temperatures, different contexts. Beyond majority vote, debate topology matters: MultiAgentBench (arXiv:2503.01935, ACL 2025) evaluated star / chain / tree / graph coordination and found **graph best for research**, with a "coordination tax" past ~4 agents. AgentVerse (ICLR 2024) documents two emergent patterns — volunteer behaviors and conformity behaviors — and conformity is both a feature (finding consensus) and a risk (groupthink, Lesson 24). This lesson maps the topology space, builds each variant, and measures the coordination tax.
|
||||
|
||||
**Type:** Learn + Build
|
||||
**Languages:** Python (stdlib)
|
||||
**Prerequisites:** Phase 16 · 07 (Society of Mind and Debate), Phase 16 · 14 (Consensus and BFT)
|
||||
**Time:** ~75 minutes
|
||||
|
||||
## Problem
|
||||
|
||||
Debate can improve accuracy (Du et al., arXiv:2305.14325). It can also degrade it. Whether debate helps depends on four structural choices:
|
||||
|
||||
1. Who talks to whom (topology).
|
||||
2. How many rounds (Du 2023: both rounds and agents matter independently).
|
||||
3. Whether agents are heterogeneous (different base models break monoculture).
|
||||
4. Whether an adversarial voice is present (steel-manning vs. straw-manning).
|
||||
|
||||
Teams that bolt "run 5 agents and vote" onto a task often regress vs. a single agent. The failures are not random. They track topology and heterogeneity. This lesson is the topology map.
|
||||
|
||||
## Concept
|
||||
|
||||
### Self-consistency, the single-model baseline
|
||||
|
||||
Wang et al. 2022 ("Self-Consistency Improves Chain of Thought Reasoning") sampled the same model N times at temperature > 0 and majority-voted on reasoning-path answers. The result on GSM8K: substantial gains with N=40 samples over a single greedy decode. Self-consistency is the single-agent precursor to multi-agent voting.
|
||||
|
||||
Limit: self-consistency uses one base model. Errors are correlated by construction. If the model has a systematic bias, all N samples share it.
|
||||
|
||||
### Multi-agent vote, the heterogeneous extension
|
||||
|
||||
Replace N samples with N *different* agents. Different base models (Claude, GPT, Llama), different prompts, different tool access. The benefit: uncorrelated errors. The cost: different agents cost different amounts; coordinating them adds overhead.
|
||||
|
||||
The canonical 2026 name for heterogeneous debate is **A-HMAD** — Adversarial Heterogeneous Multi-Agent Debate. Not universally adopted, but papers use the term for "different models debate, which reduces correlated errors from monoculture collapse."
|
||||
|
||||
### The four topologies
|
||||
|
||||
```
|
||||
star chain tree graph
|
||||
|
||||
┌─A─┐ A─B─C─D ┌──A──┐ A───B
|
||||
│ │ │ │ │ × │
|
||||
B C B C D───C
|
||||
│ │ / \ / \
|
||||
D E D E F G (fully connected)
|
||||
```
|
||||
|
||||
Star: one hub, all others talk only to hub. Equivalent to supervisor-worker without back-channel.
|
||||
Chain: linear, each agent sees the prior one's output. Pipeline-like.
|
||||
Tree: hierarchical, used by hierarchical agent systems (Lesson 06).
|
||||
Graph: any-to-any. Includes fully-connected clique and arbitrary DAGs.
|
||||
|
||||
### The coordination tax (MultiAgentBench)
|
||||
|
||||
MultiAgentBench (MARBLE, ACL 2025, arXiv:2503.01935) benchmarked star, chain, tree, graph on a task suite including research, coding, and planning. Key measured results:
|
||||
|
||||
- **Graph** topology wins on research tasks. Information flows any-to-any; agents can critique each other.
|
||||
- **Star** wins on fast-answer factual tasks. Hub filters and consolidates.
|
||||
- **Chain** wins on stepwise pipelines (staged refinement).
|
||||
- **Coordination tax** appears past ~4 agents in graph topology. Wall-clock and token cost grow faster than quality.
|
||||
|
||||
The 4-agent ceiling is empirical, not fundamental. It reflects 2026 LLM context capacity: each agent's context fills with peers' outputs, and marginal value of adding agent N+1 drops once everyone can see everyone.
|
||||
|
||||
### Multi-Agent Debate Strategies ("Should we be going MAD?")
|
||||
|
||||
arXiv:2311.17371 is the 2023 survey of MAD strategies. Key finding replicated by others: MAD variants that are *structurally similar* to self-consistency (independent sampling + aggregation) often underperform self-consistency when using the same budget. MAD helps most when agents are genuinely heterogeneous and the debate has adversarial structure (one agent argues against).
|
||||
|
||||
### AgentVerse emergent patterns
|
||||
|
||||
AgentVerse (ICLR 2024, https://proceedings.iclr.cc/paper_files/paper/2024/file/578e65cdee35d00c708d4c64bce32971-Paper-Conference.pdf) documents two behaviors that emerge from multi-agent debate even without explicit design:
|
||||
|
||||
- **Volunteer.** An agent offers help ("I can take the next step") unprompted. Useful: it allocates work to the most-capable agent for a subtask.
|
||||
- **Conformity.** An agent adjusts its stance to match a critic, even when the critic is wrong. This is the debate-equivalent of sycophancy (Lesson 14).
|
||||
|
||||
Conformity is why debate-until-agreement rewards bullies. Bounded rounds with a separate judge mitigate.
|
||||
|
||||
### Heterogeneity: the actual knob that moves accuracy
|
||||
|
||||
A 2024-2026 pattern in the practical literature: swapping one of your N agents for a different base model gives a bigger accuracy bump than increasing N by 1. The intuition is monoculture — each new independent-error source is worth more than an additional correlated sample.
|
||||
|
||||
In the limit, heterogeneity beats numerosity. Three different models beat five copies of one model on most tasks that have clean ground truth.
|
||||
|
||||
### Jury methods
|
||||
|
||||
The Sibyl framework (cited in Minsky-LLM literature) formalizes a "jury" — a small set of specialized agents that refine answers by voting at each stage. Unlike plain majority vote, a jury has roles: one agent cross-examines, one supplies context, one scores plausibility. Jury methods are a midpoint between plain vote (cheap, monoculture-prone) and full MAD (expensive, conformity-prone).
|
||||
|
||||
### When vote-with-debate dominates
|
||||
|
||||
- The question has ground truth (fact, math, code behavior). Vote convergence is meaningful.
|
||||
- Agents can access different sources or tools (heterogeneity is available).
|
||||
- Rounds are bounded (2-3 typical) and there is a separate judge or verifier.
|
||||
- Budget allows 3-5 agents. Beyond 5-7 on graph topology, coordination tax dominates.
|
||||
|
||||
### When vote-with-debate hurts
|
||||
|
||||
- The question is opinion-shaped. Agents converge to whichever answer looks most confident, not most correct.
|
||||
- All agents share a base model. Monoculture makes consensus meaningless.
|
||||
- Rounds are unbounded. Conformity wins every time.
|
||||
- The task is simple. A single agent with self-consistency at N=5 is cheaper and as accurate.
|
||||
|
||||
## Build It
|
||||
|
||||
`code/main.py` implements:
|
||||
|
||||
- `run_star(agents, hub, question)` — hub polls each worker, aggregates.
|
||||
- `run_chain(agents, question)` — sequential refinement.
|
||||
- `run_tree(root, children, question)` — hierarchical with depth-2 aggregation.
|
||||
- `run_graph(agents, question, rounds)` — all-to-all debate, bounded rounds.
|
||||
- A scripted heterogeneity dial: each agent has an `error_bias` indicating its systematic wrongness.
|
||||
- A measurement harness that runs each topology at N=3, 5, 7 and reports (accuracy, total_tokens, wallclock_simulated).
|
||||
|
||||
Run:
|
||||
|
||||
```
|
||||
python3 code/main.py
|
||||
```
|
||||
|
||||
Expected output: a table of topology × N → (accuracy, tokens, latency). Graph wins at N=3-5 on the research-style tasks; star wins on the fast-factual tasks; graph at N=7 shows the coordination tax (latency inflates faster than accuracy).
|
||||
|
||||
## Use It
|
||||
|
||||
`outputs/skill-topology-picker.md` is a skill that reads a task description and recommends a topology (star / chain / tree / graph), an N (number of agents), a heterogeneity profile (base models to use), and a round bound.
|
||||
|
||||
## Ship It
|
||||
|
||||
For any ensemble:
|
||||
|
||||
- Start with **self-consistency at N=5** using one strong base model. It is the cheap baseline.
|
||||
- Upgrade to **heterogeneous voting at N=3** if accuracy matters. Measure the delta.
|
||||
- Only upgrade to **debate topology** if the task has structure (research, multi-step) and bounded rounds are feasible.
|
||||
- Always log the minority cluster. When a minority is persistently right, you have a diversity signal.
|
||||
- Benchmark wall-clock and tokens alongside accuracy. "Better accuracy at 10x cost" is a business decision.
|
||||
|
||||
## Exercises
|
||||
|
||||
1. Run `code/main.py`. Plot the coordination-tax curve for graph topology: accuracy vs N, tokens vs N. At what N does the curve inflect?
|
||||
2. Implement A-HMAD: three agents with deliberately different biases. How does the all-same-bias baseline compare to A-HMAD on the monoculture attack from Lesson 14?
|
||||
3. Add a "judge" role to the graph topology that does not vote, only scores the final consensus. Does this change the emergent conformity behavior?
|
||||
4. Read the AgentVerse paper (ICLR 2024). Identify which emergent behavior your implementation exhibits most strongly. Can you elicit the opposite behavior by a prompt change?
|
||||
5. Read MultiAgentBench (arXiv:2503.01935) Section 4 (topology experiments). Reproduce the "graph-wins-research" result on one task from the paper using your harness.
|
||||
|
||||
## Key Terms
|
||||
|
||||
| Term | What people say | What it actually means |
|
||||
|------|----------------|------------------------|
|
||||
| Self-consistency | "Sample N times, vote" | Wang 2022. Single model, N temperature>0 samples, majority vote on reasoning paths. |
|
||||
| Heterogeneity | "Different models" | Ensemble of different base models or prompt families. Breaks monoculture. |
|
||||
| MAD | "Multi-agent debate" | Generic term for agents exchanging critiques over rounds. See Du 2023. |
|
||||
| A-HMAD | "Adversarial Heterogeneous MAD" | MAD variant emphasizing different models + adversarial structure. |
|
||||
| Topology | "Who talks to whom" | Star, chain, tree, graph. Determines information flow. |
|
||||
| Coordination tax | "Diminishing returns" | Above ~4 agents on graph, cost grows faster than quality. |
|
||||
| Volunteer behavior | "Unprompted help" | AgentVerse emergent pattern: an agent offers to take a step. |
|
||||
| Conformity behavior | "Agreement under pressure" | AgentVerse emergent pattern: an agent aligns with a critic. |
|
||||
| Jury | "Small specialized panel" | Sibyl-style ensemble with roles (examiner, context, scorer). |
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Wang et al. — Self-Consistency Improves Chain of Thought Reasoning](https://arxiv.org/abs/2203.11171) — single-model baseline
|
||||
- [Du et al. — Improving Factuality and Reasoning via Multiagent Debate](https://arxiv.org/abs/2305.14325) — both agents AND rounds matter independently
|
||||
- [MultiAgentBench / MARBLE](https://arxiv.org/abs/2503.01935) — topology benchmark showing graph best for research, chain for pipelines
|
||||
- [Should we be going MAD?](https://arxiv.org/abs/2311.17371) — MAD-strategy survey; finds MAD often loses to self-consistency at equal budget
|
||||
- [AgentVerse (ICLR 2024)](https://proceedings.iclr.cc/paper_files/paper/2024/file/578e65cdee35d00c708d4c64bce32971-Paper-Conference.pdf) — volunteer and conformity emergent patterns
|
||||
- [MARBLE repo](https://github.com/ulab-uiuc/MARBLE) — reference benchmark implementation
|
||||
+39
@@ -0,0 +1,39 @@
|
||||
---
|
||||
name: topology-picker
|
||||
description: Pick a multi-agent debate topology (star / chain / tree / graph), an N of agents, a heterogeneity profile, and a round bound for a given task.
|
||||
version: 1.0.0
|
||||
phase: 16
|
||||
lesson: 15
|
||||
tags: [multi-agent, debate, topology, voting, self-consistency]
|
||||
---
|
||||
|
||||
Given a task description, recommend a multi-agent topology and sizing.
|
||||
|
||||
Produce:
|
||||
|
||||
1. **Task fingerprint.** Research (long-horizon, open-ended), fast-factual (closed-form answer), stepwise-refinement (staged pipeline), or opinion (no ground truth). Pick one; if it spans two, pick the dominant shape.
|
||||
2. **Topology.** Star, chain, tree, or graph. Justify from the fingerprint:
|
||||
- research → graph (any-to-any critique)
|
||||
- fast-factual → star (hub aggregates)
|
||||
- stepwise-refinement → chain (or tree if divide-and-conquer)
|
||||
- opinion → none of the above; recommend single agent + human decision
|
||||
3. **N of agents.** 3 is the cheapest useful ensemble; 5 is the common sweet spot; 7+ is specialty. Above 5 on graph topology, warn about coordination tax.
|
||||
4. **Heterogeneity profile.** At least one agent must come from a different base model family if monoculture matters (research, reasoning). Prefer 3 different base models at N=5.
|
||||
5. **Round bound.** 1 round = vote. 2 rounds = one refinement. 3 rounds = maximum before conformity dominates. Never unbounded.
|
||||
6. **Aggregation.** Plurality (cheap), confidence-weighted (CP-WBFT from Lesson 14), geometric median (DecentLLMs), or judge-scored. Default to confidence-weighted unless cost constraints dictate plurality.
|
||||
7. **Escalation.** Below-threshold consensus → escalate where? Human, another ensemble with different base models, or abstention?
|
||||
|
||||
Hard rejects:
|
||||
|
||||
- Any recommendation of 10+ agents on graph topology. Coordination tax dominates; measure first.
|
||||
- Star topology for open research questions. Star loses the benefit of any-to-any critique.
|
||||
- Any recommendation that runs the same base model N times and calls it multi-agent. That is self-consistency in disguise; label it correctly.
|
||||
- Unbounded rounds. Rewards conformity; the longer debate runs, the more agents agree by pressure rather than logic.
|
||||
|
||||
Refusal rules:
|
||||
|
||||
- If the task has no ground truth (opinion, synthesis, creative), state that voting is advisory. Recommend single agent + human decision.
|
||||
- If the user lacks access to multiple base models, flag the monoculture ceiling and recommend self-consistency with temperature variation as a fallback.
|
||||
- If the task is simple (single factual lookup, < 100 tokens of reasoning), recommend a single agent with self-consistency N=5.
|
||||
|
||||
Output: a one-page brief. Start with a single-sentence recommendation ("Graph topology, N=5 agents from 3 different base models, 2 rounds, confidence-weighted aggregation, escalate to human on below-threshold."), then the seven sections above. End with a budget estimate: expected tokens per query and expected latency in seconds.
|
||||
Reference in New Issue
Block a user