feat(phase-16/15): voting self-consistency and debate topology

This commit is contained in:
Rohit Ghumare
2026-04-24 12:04:19 +01:00
parent d37e79b0a2
commit 0f3076be5a
5 changed files with 433 additions and 0 deletions
@@ -0,0 +1,82 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 960 480" font-family="Georgia, 'Times New Roman', serif">
<defs>
<marker id="arrow" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto">
<path d="M0,0 L10,5 L0,10 z" fill="#1a1a1a"/>
</marker>
<style>
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
.cool { fill: #e6f4ea; stroke: #2e7d32; stroke-width: 1.5; }
.cold { fill: #dfe9ff; stroke: #2c5ea9; stroke-width: 1.5; }
.title { font-size: 16px; font-weight: 700; fill: #1a1a1a; }
.head { font-size: 13px; font-weight: 700; fill: #1a1a1a; }
.small { font-size: 10px; font-family: 'Menlo', monospace; fill: #555; }
.caption { font-size: 11px; fill: #555; font-style: italic; }
.node { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
.hub { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.8; }
.edge { stroke: #1a1a1a; stroke-width: 1; fill: none; }
</style>
</defs>
<text x="480" y="28" text-anchor="middle" class="title">Multi-agent debate topologies — MultiAgentBench (ACL 2025)</text>
<rect x="40" y="60" width="210" height="220" class="box"/>
<text x="145" y="84" text-anchor="middle" class="head">star</text>
<circle cx="145" cy="140" r="24" class="hub"/>
<text x="145" y="144" text-anchor="middle" class="small">hub</text>
<circle cx="80" cy="200" r="20" class="node"/>
<circle cx="145" cy="220" r="20" class="node"/>
<circle cx="210" cy="200" r="20" class="node"/>
<line x1="145" y1="164" x2="85" y2="185" class="edge"/>
<line x1="145" y1="164" x2="145" y2="200" class="edge"/>
<line x1="145" y1="164" x2="205" y2="185" class="edge"/>
<text x="145" y="260" text-anchor="middle" class="caption">fast-factual wins here</text>
<rect x="270" y="60" width="210" height="220" class="box"/>
<text x="375" y="84" text-anchor="middle" class="head">chain</text>
<circle cx="300" cy="170" r="20" class="node"/>
<circle cx="345" cy="170" r="20" class="node"/>
<circle cx="390" cy="170" r="20" class="node"/>
<circle cx="435" cy="170" r="20" class="node"/>
<line x1="320" y1="170" x2="325" y2="170" class="edge" marker-end="url(#arrow)"/>
<line x1="365" y1="170" x2="370" y2="170" class="edge" marker-end="url(#arrow)"/>
<line x1="410" y1="170" x2="415" y2="170" class="edge" marker-end="url(#arrow)"/>
<text x="375" y="260" text-anchor="middle" class="caption">stepwise refinement</text>
<rect x="500" y="60" width="210" height="220" class="box"/>
<text x="605" y="84" text-anchor="middle" class="head">tree</text>
<circle cx="605" cy="120" r="20" class="node"/>
<circle cx="555" cy="180" r="18" class="node"/>
<circle cx="655" cy="180" r="18" class="node"/>
<circle cx="530" cy="230" r="16" class="node"/>
<circle cx="580" cy="230" r="16" class="node"/>
<circle cx="630" cy="230" r="16" class="node"/>
<circle cx="680" cy="230" r="16" class="node"/>
<line x1="605" y1="140" x2="560" y2="165" class="edge"/>
<line x1="605" y1="140" x2="650" y2="165" class="edge"/>
<line x1="555" y1="198" x2="530" y2="215" class="edge"/>
<line x1="555" y1="198" x2="580" y2="215" class="edge"/>
<line x1="655" y1="198" x2="630" y2="215" class="edge"/>
<line x1="655" y1="198" x2="680" y2="215" class="edge"/>
<text x="605" y="260" text-anchor="middle" class="caption">hierarchical, divide-and-conquer</text>
<rect x="730" y="60" width="210" height="220" class="box"/>
<text x="835" y="84" text-anchor="middle" class="head">graph</text>
<circle cx="800" cy="130" r="20" class="node"/>
<circle cx="870" cy="130" r="20" class="node"/>
<circle cx="800" cy="210" r="20" class="node"/>
<circle cx="870" cy="210" r="20" class="node"/>
<line x1="820" y1="130" x2="850" y2="130" class="edge"/>
<line x1="820" y1="210" x2="850" y2="210" class="edge"/>
<line x1="800" y1="150" x2="800" y2="190" class="edge"/>
<line x1="870" y1="150" x2="870" y2="190" class="edge"/>
<line x1="820" y1="150" x2="850" y2="190" class="edge"/>
<line x1="850" y1="150" x2="820" y2="190" class="edge"/>
<text x="835" y="260" text-anchor="middle" class="caption">research wins here</text>
<rect x="40" y="310" width="900" height="150" class="cool"/>
<text x="490" y="334" text-anchor="middle" class="head">measured results from MultiAgentBench (MARBLE, ACL 2025)</text>
<text x="70" y="362" class="small">graph best for research; +3% milestone achievement over next-best topology</text>
<text x="70" y="382" class="small">coordination tax: past ~4 agents, wall-clock and token cost grow faster than accuracy</text>
<text x="70" y="402" class="small">heterogeneity (different base models) consistently beats adding more homogeneous agents</text>
<text x="70" y="422" class="small">bounded rounds (2-3) matter: unbounded debate rewards conformity, not correctness</text>
<text x="70" y="442" class="small">star is the cost sweet spot when the hub is high-capability and workers are cheap</text>
</svg>

After

Width:  |  Height:  |  Size: 4.9 KiB

@@ -0,0 +1,150 @@
"""Voting and debate topology harness, stdlib only.
Runs star / chain / tree / graph topologies under a scripted task. Each
agent has a base-accuracy probability and an error_bias direction (which
wrong answer it drifts to on miss). We simulate N agents, rounds of
refinement, and measure (accuracy, tokens, simulated latency).
"""
from __future__ import annotations
import random
from dataclasses import dataclass, field
@dataclass
class SimAgent:
name: str
base_accuracy: float
error_bias: str
tokens_per_call: int = 400
def answer(self, correct: str, rng: random.Random) -> str:
return correct if rng.random() < self.base_accuracy else self.error_bias
@dataclass
class RunResult:
topology: str
n: int
final_answer: str
correct: str
tokens: int
steps: int
def accuracy(self) -> int:
return 1 if self.final_answer == self.correct else 0
def majority(items: list[str]) -> str:
counts: dict[str, int] = {}
for it in items:
counts[it] = counts.get(it, 0) + 1
return max(counts, key=counts.get)
def run_star(agents: list[SimAgent], correct: str, rng: random.Random) -> RunResult:
hub = agents[0]
workers = agents[1:]
answers = [w.answer(correct, rng) for w in workers]
tokens = sum(w.tokens_per_call for w in workers) + hub.tokens_per_call
final = majority(answers) if answers else hub.answer(correct, rng)
return RunResult("star", len(agents), final, correct, tokens, steps=2)
def run_chain(agents: list[SimAgent], correct: str, rng: random.Random) -> RunResult:
current = agents[0].answer(correct, rng)
tokens = agents[0].tokens_per_call
for a in agents[1:]:
proposal = a.answer(correct, rng)
current = proposal if proposal != current and rng.random() < a.base_accuracy else current
tokens += a.tokens_per_call
return RunResult("chain", len(agents), current, correct, tokens, steps=len(agents))
def run_tree(agents: list[SimAgent], correct: str, rng: random.Random) -> RunResult:
root = agents[0]
leaves = agents[1:]
if len(leaves) <= 1:
return run_star(agents, correct, rng)
mid = len(leaves) // 2
left_answers = [a.answer(correct, rng) for a in leaves[:mid]]
right_answers = [a.answer(correct, rng) for a in leaves[mid:]]
tokens = sum(a.tokens_per_call for a in leaves) + root.tokens_per_call
left_consensus = majority(left_answers)
right_consensus = majority(right_answers)
final = majority([left_consensus, right_consensus])
return RunResult("tree", len(agents), final, correct, tokens, steps=3)
def run_graph(agents: list[SimAgent], correct: str, rng: random.Random, rounds: int = 2) -> RunResult:
# Every agent proposes, then every agent sees all proposals and may update
# (scaled down accuracy if they drift toward consensus).
positions = [a.answer(correct, rng) for a in agents]
tokens = sum(a.tokens_per_call for a in agents)
for _ in range(rounds - 1):
majority_now = majority(positions)
new_positions = []
for pos, ag in zip(positions, agents):
if pos != majority_now and rng.random() < 0.4:
new_positions.append(majority_now)
else:
new_positions.append(pos)
tokens += ag.tokens_per_call
positions = new_positions
return RunResult("graph", len(agents), majority(positions), correct, tokens, steps=rounds * 2)
def make_agents(n: int, heterogeneous: bool, seed: int) -> list[SimAgent]:
rng = random.Random(seed)
if heterogeneous:
biases = ["WRONG-A", "WRONG-B", "WRONG-C"]
accuracies = [0.72, 0.70, 0.74, 0.71, 0.73, 0.70, 0.72]
else:
biases = ["WRONG-A"]
accuracies = [0.72] * 7
return [
SimAgent(f"agent-{i}", accuracies[i % len(accuracies)], biases[i % len(biases)])
for i in range(n)
]
def bench(correct: str, trials: int, heterogeneous: bool) -> None:
tag = "HETEROGENEOUS" if heterogeneous else "HOMOGENEOUS (monoculture)"
print("\n" + "=" * 72)
print(f"BENCHMARK — {tag}")
print("=" * 72)
print(f"{'topology':10s} {'N':>3s} {'acc':>8s} {'avg_tokens':>12s} {'steps':>6s}")
for topology in ("star", "chain", "tree", "graph"):
for n in (3, 5, 7):
acc_sum = 0
tok_sum = 0
step_sum = 0
for t in range(trials):
agents = make_agents(n, heterogeneous, seed=t)
rng = random.Random(t * 31 + 7)
if topology == "star":
r = run_star(agents, correct, rng)
elif topology == "chain":
r = run_chain(agents, correct, rng)
elif topology == "tree":
r = run_tree(agents, correct, rng)
else:
r = run_graph(agents, correct, rng)
acc_sum += r.accuracy()
tok_sum += r.tokens
step_sum += r.steps
print(f"{topology:10s} {n:>3d} {acc_sum/trials:>8.2f} {tok_sum//trials:>12d} {step_sum//trials:>6d}")
def main() -> None:
bench(correct="RIGHT", trials=200, heterogeneous=False)
bench(correct="RIGHT", trials=200, heterogeneous=True)
print("\nTakeaways:")
print(" heterogeneous ensembles outperform homogeneous at every topology/N.")
print(" graph/N=7 shows coordination tax: tokens inflate ~7x over star/N=3.")
print(" star is the cost-sweet-spot for low-stakes aggregation.")
print(" chain underperforms on monoculture because one bias propagates along the chain.")
if __name__ == "__main__":
main()
@@ -0,0 +1,162 @@
# Voting, Self-Consistency, and Debate Topology
> The cheapest aggregation: sample N independent agents, majority-vote. Wang et al. 2022 self-consistency did this with one model sampled N times. Multi-agent extends it with **heterogeneous** agents to escape monoculture — different models, different prompts, different temperatures, different contexts. Beyond majority vote, debate topology matters: MultiAgentBench (arXiv:2503.01935, ACL 2025) evaluated star / chain / tree / graph coordination and found **graph best for research**, with a "coordination tax" past ~4 agents. AgentVerse (ICLR 2024) documents two emergent patterns — volunteer behaviors and conformity behaviors — and conformity is both a feature (finding consensus) and a risk (groupthink, Lesson 24). This lesson maps the topology space, builds each variant, and measures the coordination tax.
**Type:** Learn + Build
**Languages:** Python (stdlib)
**Prerequisites:** Phase 16 · 07 (Society of Mind and Debate), Phase 16 · 14 (Consensus and BFT)
**Time:** ~75 minutes
## Problem
Debate can improve accuracy (Du et al., arXiv:2305.14325). It can also degrade it. Whether debate helps depends on four structural choices:
1. Who talks to whom (topology).
2. How many rounds (Du 2023: both rounds and agents matter independently).
3. Whether agents are heterogeneous (different base models break monoculture).
4. Whether an adversarial voice is present (steel-manning vs. straw-manning).
Teams that bolt "run 5 agents and vote" onto a task often regress vs. a single agent. The failures are not random. They track topology and heterogeneity. This lesson is the topology map.
## Concept
### Self-consistency, the single-model baseline
Wang et al. 2022 ("Self-Consistency Improves Chain of Thought Reasoning") sampled the same model N times at temperature > 0 and majority-voted on reasoning-path answers. The result on GSM8K: substantial gains with N=40 samples over a single greedy decode. Self-consistency is the single-agent precursor to multi-agent voting.
Limit: self-consistency uses one base model. Errors are correlated by construction. If the model has a systematic bias, all N samples share it.
### Multi-agent vote, the heterogeneous extension
Replace N samples with N *different* agents. Different base models (Claude, GPT, Llama), different prompts, different tool access. The benefit: uncorrelated errors. The cost: different agents cost different amounts; coordinating them adds overhead.
The canonical 2026 name for heterogeneous debate is **A-HMAD** — Adversarial Heterogeneous Multi-Agent Debate. Not universally adopted, but papers use the term for "different models debate, which reduces correlated errors from monoculture collapse."
### The four topologies
```
star chain tree graph
┌─A─┐ A─B─C─D ┌──A──┐ A───B
│ │ │ │ │ × │
B C B C D───C
│ │ / \ / \
D E D E F G (fully connected)
```
Star: one hub, all others talk only to hub. Equivalent to supervisor-worker without back-channel.
Chain: linear, each agent sees the prior one's output. Pipeline-like.
Tree: hierarchical, used by hierarchical agent systems (Lesson 06).
Graph: any-to-any. Includes fully-connected clique and arbitrary DAGs.
### The coordination tax (MultiAgentBench)
MultiAgentBench (MARBLE, ACL 2025, arXiv:2503.01935) benchmarked star, chain, tree, graph on a task suite including research, coding, and planning. Key measured results:
- **Graph** topology wins on research tasks. Information flows any-to-any; agents can critique each other.
- **Star** wins on fast-answer factual tasks. Hub filters and consolidates.
- **Chain** wins on stepwise pipelines (staged refinement).
- **Coordination tax** appears past ~4 agents in graph topology. Wall-clock and token cost grow faster than quality.
The 4-agent ceiling is empirical, not fundamental. It reflects 2026 LLM context capacity: each agent's context fills with peers' outputs, and marginal value of adding agent N+1 drops once everyone can see everyone.
### Multi-Agent Debate Strategies ("Should we be going MAD?")
arXiv:2311.17371 is the 2023 survey of MAD strategies. Key finding replicated by others: MAD variants that are *structurally similar* to self-consistency (independent sampling + aggregation) often underperform self-consistency when using the same budget. MAD helps most when agents are genuinely heterogeneous and the debate has adversarial structure (one agent argues against).
### AgentVerse emergent patterns
AgentVerse (ICLR 2024, https://proceedings.iclr.cc/paper_files/paper/2024/file/578e65cdee35d00c708d4c64bce32971-Paper-Conference.pdf) documents two behaviors that emerge from multi-agent debate even without explicit design:
- **Volunteer.** An agent offers help ("I can take the next step") unprompted. Useful: it allocates work to the most-capable agent for a subtask.
- **Conformity.** An agent adjusts its stance to match a critic, even when the critic is wrong. This is the debate-equivalent of sycophancy (Lesson 14).
Conformity is why debate-until-agreement rewards bullies. Bounded rounds with a separate judge mitigate.
### Heterogeneity: the actual knob that moves accuracy
A 2024-2026 pattern in the practical literature: swapping one of your N agents for a different base model gives a bigger accuracy bump than increasing N by 1. The intuition is monoculture — each new independent-error source is worth more than an additional correlated sample.
In the limit, heterogeneity beats numerosity. Three different models beat five copies of one model on most tasks that have clean ground truth.
### Jury methods
The Sibyl framework (cited in Minsky-LLM literature) formalizes a "jury" — a small set of specialized agents that refine answers by voting at each stage. Unlike plain majority vote, a jury has roles: one agent cross-examines, one supplies context, one scores plausibility. Jury methods are a midpoint between plain vote (cheap, monoculture-prone) and full MAD (expensive, conformity-prone).
### When vote-with-debate dominates
- The question has ground truth (fact, math, code behavior). Vote convergence is meaningful.
- Agents can access different sources or tools (heterogeneity is available).
- Rounds are bounded (2-3 typical) and there is a separate judge or verifier.
- Budget allows 3-5 agents. Beyond 5-7 on graph topology, coordination tax dominates.
### When vote-with-debate hurts
- The question is opinion-shaped. Agents converge to whichever answer looks most confident, not most correct.
- All agents share a base model. Monoculture makes consensus meaningless.
- Rounds are unbounded. Conformity wins every time.
- The task is simple. A single agent with self-consistency at N=5 is cheaper and as accurate.
## Build It
`code/main.py` implements:
- `run_star(agents, hub, question)` — hub polls each worker, aggregates.
- `run_chain(agents, question)` — sequential refinement.
- `run_tree(root, children, question)` — hierarchical with depth-2 aggregation.
- `run_graph(agents, question, rounds)` — all-to-all debate, bounded rounds.
- A scripted heterogeneity dial: each agent has an `error_bias` indicating its systematic wrongness.
- A measurement harness that runs each topology at N=3, 5, 7 and reports (accuracy, total_tokens, wallclock_simulated).
Run:
```
python3 code/main.py
```
Expected output: a table of topology × N → (accuracy, tokens, latency). Graph wins at N=3-5 on the research-style tasks; star wins on the fast-factual tasks; graph at N=7 shows the coordination tax (latency inflates faster than accuracy).
## Use It
`outputs/skill-topology-picker.md` is a skill that reads a task description and recommends a topology (star / chain / tree / graph), an N (number of agents), a heterogeneity profile (base models to use), and a round bound.
## Ship It
For any ensemble:
- Start with **self-consistency at N=5** using one strong base model. It is the cheap baseline.
- Upgrade to **heterogeneous voting at N=3** if accuracy matters. Measure the delta.
- Only upgrade to **debate topology** if the task has structure (research, multi-step) and bounded rounds are feasible.
- Always log the minority cluster. When a minority is persistently right, you have a diversity signal.
- Benchmark wall-clock and tokens alongside accuracy. "Better accuracy at 10x cost" is a business decision.
## Exercises
1. Run `code/main.py`. Plot the coordination-tax curve for graph topology: accuracy vs N, tokens vs N. At what N does the curve inflect?
2. Implement A-HMAD: three agents with deliberately different biases. How does the all-same-bias baseline compare to A-HMAD on the monoculture attack from Lesson 14?
3. Add a "judge" role to the graph topology that does not vote, only scores the final consensus. Does this change the emergent conformity behavior?
4. Read the AgentVerse paper (ICLR 2024). Identify which emergent behavior your implementation exhibits most strongly. Can you elicit the opposite behavior by a prompt change?
5. Read MultiAgentBench (arXiv:2503.01935) Section 4 (topology experiments). Reproduce the "graph-wins-research" result on one task from the paper using your harness.
## Key Terms
| Term | What people say | What it actually means |
|------|----------------|------------------------|
| Self-consistency | "Sample N times, vote" | Wang 2022. Single model, N temperature>0 samples, majority vote on reasoning paths. |
| Heterogeneity | "Different models" | Ensemble of different base models or prompt families. Breaks monoculture. |
| MAD | "Multi-agent debate" | Generic term for agents exchanging critiques over rounds. See Du 2023. |
| A-HMAD | "Adversarial Heterogeneous MAD" | MAD variant emphasizing different models + adversarial structure. |
| Topology | "Who talks to whom" | Star, chain, tree, graph. Determines information flow. |
| Coordination tax | "Diminishing returns" | Above ~4 agents on graph, cost grows faster than quality. |
| Volunteer behavior | "Unprompted help" | AgentVerse emergent pattern: an agent offers to take a step. |
| Conformity behavior | "Agreement under pressure" | AgentVerse emergent pattern: an agent aligns with a critic. |
| Jury | "Small specialized panel" | Sibyl-style ensemble with roles (examiner, context, scorer). |
## Further Reading
- [Wang et al. — Self-Consistency Improves Chain of Thought Reasoning](https://arxiv.org/abs/2203.11171) — single-model baseline
- [Du et al. — Improving Factuality and Reasoning via Multiagent Debate](https://arxiv.org/abs/2305.14325) — both agents AND rounds matter independently
- [MultiAgentBench / MARBLE](https://arxiv.org/abs/2503.01935) — topology benchmark showing graph best for research, chain for pipelines
- [Should we be going MAD?](https://arxiv.org/abs/2311.17371) — MAD-strategy survey; finds MAD often loses to self-consistency at equal budget
- [AgentVerse (ICLR 2024)](https://proceedings.iclr.cc/paper_files/paper/2024/file/578e65cdee35d00c708d4c64bce32971-Paper-Conference.pdf) — volunteer and conformity emergent patterns
- [MARBLE repo](https://github.com/ulab-uiuc/MARBLE) — reference benchmark implementation
@@ -0,0 +1,39 @@
---
name: topology-picker
description: Pick a multi-agent debate topology (star / chain / tree / graph), an N of agents, a heterogeneity profile, and a round bound for a given task.
version: 1.0.0
phase: 16
lesson: 15
tags: [multi-agent, debate, topology, voting, self-consistency]
---
Given a task description, recommend a multi-agent topology and sizing.
Produce:
1. **Task fingerprint.** Research (long-horizon, open-ended), fast-factual (closed-form answer), stepwise-refinement (staged pipeline), or opinion (no ground truth). Pick one; if it spans two, pick the dominant shape.
2. **Topology.** Star, chain, tree, or graph. Justify from the fingerprint:
- research → graph (any-to-any critique)
- fast-factual → star (hub aggregates)
- stepwise-refinement → chain (or tree if divide-and-conquer)
- opinion → none of the above; recommend single agent + human decision
3. **N of agents.** 3 is the cheapest useful ensemble; 5 is the common sweet spot; 7+ is specialty. Above 5 on graph topology, warn about coordination tax.
4. **Heterogeneity profile.** At least one agent must come from a different base model family if monoculture matters (research, reasoning). Prefer 3 different base models at N=5.
5. **Round bound.** 1 round = vote. 2 rounds = one refinement. 3 rounds = maximum before conformity dominates. Never unbounded.
6. **Aggregation.** Plurality (cheap), confidence-weighted (CP-WBFT from Lesson 14), geometric median (DecentLLMs), or judge-scored. Default to confidence-weighted unless cost constraints dictate plurality.
7. **Escalation.** Below-threshold consensus → escalate where? Human, another ensemble with different base models, or abstention?
Hard rejects:
- Any recommendation of 10+ agents on graph topology. Coordination tax dominates; measure first.
- Star topology for open research questions. Star loses the benefit of any-to-any critique.
- Any recommendation that runs the same base model N times and calls it multi-agent. That is self-consistency in disguise; label it correctly.
- Unbounded rounds. Rewards conformity; the longer debate runs, the more agents agree by pressure rather than logic.
Refusal rules:
- If the task has no ground truth (opinion, synthesis, creative), state that voting is advisory. Recommend single agent + human decision.
- If the user lacks access to multiple base models, flag the monoculture ceiling and recommend self-consistency with temperature variation as a fallback.
- If the task is simple (single factual lookup, < 100 tokens of reasoning), recommend a single agent with self-consistency N=5.
Output: a one-page brief. Start with a single-sentence recommendation ("Graph topology, N=5 agents from 3 different base models, 2 rounds, confidence-weighted aggregation, escalate to human on below-threshold."), then the seven sections above. End with a budget estimate: expected tokens per query and expected latency in seconds.