mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
feat(phase-16/08): role specialization planner critic executor verifier
This commit is contained in:
@@ -0,0 +1,73 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 900 500" font-family="Georgia, 'Times New Roman', serif">
|
||||
<defs>
|
||||
<marker id="arrow" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto">
|
||||
<path d="M0,0 L10,5 L0,10 z" fill="#1a1a1a"/>
|
||||
</marker>
|
||||
<style>
|
||||
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
|
||||
.hot { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.5; }
|
||||
.cool { fill: #e6f4ea; stroke: #2e7d32; stroke-width: 1.5; }
|
||||
.cold { fill: #dfe9ff; stroke: #2c5ea9; stroke-width: 1.5; }
|
||||
.dsk { fill: #e9e6ff; stroke: #5a4fcf; stroke-width: 1.5; }
|
||||
.title { font-size: 16px; font-weight: 700; fill: #1a1a1a; }
|
||||
.head { font-size: 13px; font-weight: 700; fill: #1a1a1a; }
|
||||
.step { font-size: 12px; font-family: 'Menlo', monospace; fill: #222; }
|
||||
.small { font-size: 10px; font-family: 'Menlo', monospace; fill: #555; }
|
||||
.caption { font-size: 11px; fill: #555; font-style: italic; }
|
||||
.edge { stroke: #1a1a1a; stroke-width: 1.2; fill: none; }
|
||||
</style>
|
||||
</defs>
|
||||
<text x="450" y="28" text-anchor="middle" class="title">Role specialization — planner, executor, critic, verifier</text>
|
||||
|
||||
<rect x="40" y="60" width="170" height="100" class="cool"/>
|
||||
<text x="125" y="82" text-anchor="middle" class="head">Planner</text>
|
||||
<text x="125" y="104" text-anchor="middle" class="small">reads goal</text>
|
||||
<text x="125" y="120" text-anchor="middle" class="small">produces spec</text>
|
||||
<text x="125" y="138" text-anchor="middle" class="small">(signature, tests, I/O)</text>
|
||||
<text x="125" y="154" text-anchor="middle" class="small">MetaGPT: PM + Architect</text>
|
||||
|
||||
<line x1="210" y1="110" x2="250" y2="110" class="edge" marker-end="url(#arrow)"/>
|
||||
|
||||
<rect x="250" y="60" width="170" height="100" class="hot"/>
|
||||
<text x="335" y="82" text-anchor="middle" class="head">Executor</text>
|
||||
<text x="335" y="104" text-anchor="middle" class="small">reads one step</text>
|
||||
<text x="335" y="120" text-anchor="middle" class="small">produces artifact</text>
|
||||
<text x="335" y="138" text-anchor="middle" class="small">has the real tools</text>
|
||||
<text x="335" y="154" text-anchor="middle" class="small">ChatDev: asks if unclear</text>
|
||||
|
||||
<line x1="420" y1="110" x2="460" y2="110" class="edge" marker-end="url(#arrow)"/>
|
||||
|
||||
<rect x="460" y="60" width="170" height="100" class="dsk"/>
|
||||
<text x="545" y="82" text-anchor="middle" class="head">Critic</text>
|
||||
<text x="545" y="104" text-anchor="middle" class="small">subjective review</text>
|
||||
<text x="545" y="120" text-anchor="middle" class="small">catches taste issues</text>
|
||||
<text x="545" y="138" text-anchor="middle" class="small">fooled by plausible prose</text>
|
||||
<text x="545" y="154" text-anchor="middle" class="small">(LLM-based)</text>
|
||||
|
||||
<line x1="630" y1="110" x2="670" y2="110" class="edge" marker-end="url(#arrow)"/>
|
||||
|
||||
<rect x="670" y="60" width="200" height="100" class="cold"/>
|
||||
<text x="770" y="82" text-anchor="middle" class="head">Verifier</text>
|
||||
<text x="770" y="104" text-anchor="middle" class="small">deterministic check</text>
|
||||
<text x="770" y="120" text-anchor="middle" class="small">tests, type, schema</text>
|
||||
<text x="770" y="138" text-anchor="middle" class="small">cannot be fooled</text>
|
||||
<text x="770" y="154" text-anchor="middle" class="small">(code, not LLM)</text>
|
||||
|
||||
<line x1="335" y1="160" x2="335" y2="200" class="edge"/>
|
||||
<path d="M 335 200 Q 300 220 265 200" class="edge" fill="none" stroke-dasharray="3 3"/>
|
||||
<text x="295" y="195" class="small">communicative dehallucination (ChatDev)</text>
|
||||
|
||||
<rect x="40" y="240" width="830" height="110" class="box"/>
|
||||
<text x="450" y="262" text-anchor="middle" class="head">why verifier is load-bearing</text>
|
||||
<text x="450" y="284" text-anchor="middle" class="small">MAST taxonomy (Cemri et al. 2025): 21.3% of 1642 multi-agent failures are verification gaps</text>
|
||||
<text x="450" y="304" text-anchor="middle" class="small">PwC 2025: structured validation loop in CrewAI moved accuracy 10% -> 70% (7× gain)</text>
|
||||
<text x="450" y="324" text-anchor="middle" class="small">anti-pattern: every role is an LLM that says "looks good to me" — classic MAST failure</text>
|
||||
<text x="450" y="342" text-anchor="middle" class="small">always pair critic (subjective) with verifier (deterministic); they catch different bugs</text>
|
||||
|
||||
<rect x="40" y="370" width="830" height="110" class="box"/>
|
||||
<text x="450" y="392" text-anchor="middle" class="head">MetaGPT's SOP pattern</text>
|
||||
<text x="450" y="412" text-anchor="middle" class="small">Code = SOP(Team): Product Manager, Architect, Project Manager, Engineer, QA Engineer</text>
|
||||
<text x="450" y="430" text-anchor="middle" class="small">each role has strict input/output schema · pipeline becomes deterministic</text>
|
||||
<text x="450" y="450" text-anchor="middle" class="small">CrewAI Agent(role, goal, backstory) is the modern ergonomic surface</text>
|
||||
<text x="450" y="470" text-anchor="middle" class="small">LangGraph encodes roles as graph nodes with specialized prompts</text>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 5.0 KiB |
@@ -0,0 +1,132 @@
|
||||
"""Role specialization: planner, executor, critic, verifier.
|
||||
|
||||
Builds a small Python function. Critic (LLM-simulated) and verifier (code)
|
||||
together catch bugs that either alone would miss.
|
||||
|
||||
Run twice: once with correct executor output, once with off-spec output.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass, field
|
||||
|
||||
|
||||
@dataclass
|
||||
class Spec:
|
||||
task_name: str
|
||||
signature: str
|
||||
description: str
|
||||
tests: list[tuple[tuple, int]]
|
||||
|
||||
|
||||
@dataclass
|
||||
class Artifact:
|
||||
code: str
|
||||
|
||||
|
||||
@dataclass
|
||||
class CriticReport:
|
||||
approved: bool
|
||||
notes: list[str] = field(default_factory=list)
|
||||
|
||||
|
||||
@dataclass
|
||||
class VerifierReport:
|
||||
passed: bool
|
||||
failures: list[str] = field(default_factory=list)
|
||||
|
||||
|
||||
def planner(user_wish: str) -> Spec:
|
||||
"""Produces a structured spec from a high-level wish."""
|
||||
return Spec(
|
||||
task_name="add_two",
|
||||
signature="add_two(a: int, b: int) -> int",
|
||||
description=user_wish,
|
||||
tests=[((1, 2), 3), ((10, 20), 30), ((-5, 5), 0)],
|
||||
)
|
||||
|
||||
|
||||
def executor_correct(spec: Spec) -> Artifact:
|
||||
return Artifact(code="def add_two(a, b):\n return a + b\n")
|
||||
|
||||
|
||||
def executor_buggy(spec: Spec) -> Artifact:
|
||||
return Artifact(code="def add_two(a, b):\n return a * b\n")
|
||||
|
||||
|
||||
def critic(spec: Spec, art: Artifact) -> CriticReport:
|
||||
"""LLM-style review. Pattern-matches against common issues but can be fooled
|
||||
by plausible-looking code that is semantically wrong."""
|
||||
notes: list[str] = []
|
||||
if "def" not in art.code:
|
||||
notes.append("missing def statement")
|
||||
if "return" not in art.code:
|
||||
notes.append("missing return")
|
||||
if spec.task_name not in art.code:
|
||||
notes.append(f"function name does not match spec '{spec.task_name}'")
|
||||
approved = not notes
|
||||
return CriticReport(approved=approved, notes=notes)
|
||||
|
||||
|
||||
def verifier(spec: Spec, art: Artifact) -> VerifierReport:
|
||||
"""Run the code in a sandbox namespace and execute the tests. Deterministic."""
|
||||
ns: dict = {}
|
||||
try:
|
||||
exec(art.code, ns, ns)
|
||||
except Exception as e:
|
||||
return VerifierReport(passed=False, failures=[f"exec error: {e}"])
|
||||
fn = ns.get(spec.task_name)
|
||||
if not callable(fn):
|
||||
return VerifierReport(passed=False, failures=[f"no callable '{spec.task_name}' produced"])
|
||||
failures: list[str] = []
|
||||
for args, expected in spec.tests:
|
||||
try:
|
||||
got = fn(*args)
|
||||
except Exception as e:
|
||||
failures.append(f"call {args} raised {e}")
|
||||
continue
|
||||
if got != expected:
|
||||
failures.append(f"call {args}: expected {expected}, got {got}")
|
||||
return VerifierReport(passed=not failures, failures=failures)
|
||||
|
||||
|
||||
def run_pipeline(user_wish: str, executor, label: str) -> None:
|
||||
print(f"\n=== {label} ===")
|
||||
spec = planner(user_wish)
|
||||
print(f" [planner] spec: {spec.signature} with {len(spec.tests)} tests")
|
||||
art = executor(spec)
|
||||
print(f" [executor] produced:\n {art.code.replace(chr(10), chr(10)+' ')}")
|
||||
crep = critic(spec, art)
|
||||
print(f" [critic] approved={crep.approved}, notes={crep.notes}")
|
||||
vrep = verifier(spec, art)
|
||||
print(f" [verifier] passed={vrep.passed}, failures={vrep.failures}")
|
||||
if crep.approved and vrep.passed:
|
||||
print(" RESULT: ship it.")
|
||||
elif not vrep.passed:
|
||||
print(" RESULT: verifier blocked ship (deterministic catch).")
|
||||
elif not crep.approved:
|
||||
print(" RESULT: critic blocked ship (subjective catch).")
|
||||
|
||||
|
||||
def main() -> None:
|
||||
print("Role specialization pipeline — planner, executor, critic, verifier")
|
||||
print("-" * 70)
|
||||
|
||||
run_pipeline(
|
||||
"A function that returns the sum of two integers.",
|
||||
executor_correct,
|
||||
"Correct executor output",
|
||||
)
|
||||
|
||||
run_pipeline(
|
||||
"A function that returns the sum of two integers.",
|
||||
executor_buggy,
|
||||
"Buggy executor output (looks plausible; fails runtime)",
|
||||
)
|
||||
|
||||
print("\nKey insight: the critic passes the buggy code because it looks fine.")
|
||||
print("Only the verifier -- deterministic test execution -- catches the semantic bug.")
|
||||
print("All-LLM pipelines (no verifier) would ship the bug. Classic MAST failure mode.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,129 @@
|
||||
# Role Specialization — Planner, Critic, Executor, Verifier
|
||||
|
||||
> The most common multi-agent decomposition in 2026: one agent plans, one executes, one critiques or verifies. MetaGPT (arXiv:2308.00352) formalizes this as SOPs encoded into role prompts — Product Manager, Architect, Project Manager, Engineer, QA Engineer — following `Code = SOP(Team)`. ChatDev (arXiv:2307.07924) chains designer, programmer, reviewer, tester through a "chat chain" with "communicative dehallucination" (agents explicitly request missing details). The verifier is load-bearing: Cemri et al. (MAST, arXiv:2503.13657) show every multi-agent failure can be traced to missing or broken verification. PwC reported 7× accuracy gain (10% → 70%) from structured validation loops in CrewAI.
|
||||
|
||||
**Type:** Learn + Build
|
||||
**Languages:** Python (stdlib)
|
||||
**Prerequisites:** Phase 16 · 04 (Primitive Model), Phase 16 · 05 (Supervisor)
|
||||
**Time:** ~60 minutes
|
||||
|
||||
## Problem
|
||||
|
||||
Generic multi-agent systems produce generic output. Three coders in a group chat write three flavors of the same mediocre code. You can add more agents, add more rounds, and still not cross the quality threshold.
|
||||
|
||||
The fix is not more agents — it is *different* agents. Assign distinct roles. Give the critic tools the planner does not have. Give the verifier an objective test suite. Now the system has internal disagreement with grounded correction, not just parallel guessing.
|
||||
|
||||
## Concept
|
||||
|
||||
### The four canonical roles
|
||||
|
||||
**Planner.** Reads the goal, produces a step list or a spec. Tools: knowledge retrieval, docs. Output: structured plan.
|
||||
|
||||
**Executor.** Reads one plan step at a time, produces the artifact. Tools: the actual work tools (code compiler, shell, API client). Output: the artifact.
|
||||
|
||||
**Critic.** Reads the executor's output against the planner's intent. Tools: read-only access to the artifact, static analysis. Output: accept/reject with reasons.
|
||||
|
||||
**Verifier.** Reads the artifact and runs a deterministic check. Tools: test runner, type checker, schema validator. Output: pass/fail with evidence.
|
||||
|
||||
Critic is subjective, opinionated, often LLM-based. Verifier is objective, deterministic, often code-based. They are not the same role.
|
||||
|
||||
### MetaGPT's SOP pattern
|
||||
|
||||
MetaGPT (arXiv:2308.00352) encodes software engineering SOPs as role prompts:
|
||||
|
||||
- **Product Manager** writes the PRD.
|
||||
- **Architect** produces the system design.
|
||||
- **Project Manager** splits tasks.
|
||||
- **Engineer** implements.
|
||||
- **QA Engineer** runs tests.
|
||||
|
||||
Each role has a strict input/output schema. The role prompt says what the role *is* and what it *must produce*. The `Code = SOP(Team)` formulation — deterministic SOPs turn a team of LLMs into a predictable pipeline.
|
||||
|
||||
### ChatDev's communicative dehallucination
|
||||
|
||||
ChatDev adds a key move: when an executor needs a specific detail that was not in the plan, it explicitly asks the designer before continuing. This prevents the classic LLM failure of plausibly inventing the detail.
|
||||
|
||||
Implementation: the role prompt includes "when you need specific information you were not given, ask the relevant role by name before producing output."
|
||||
|
||||
### Why verifier matters most
|
||||
|
||||
Cemri et al. (MAST) traced 1642 multi-agent execution failures. 21.3% were verification gaps — the system shipped an answer no one had checked. The remaining 79% often trace back to "there was a check that failed silently or was never run." Verification is the load-bearing role.
|
||||
|
||||
PwC reported (CrewAI deployments, 2025) that adding a structured validation loop moved accuracy from 10% to 70%. 7× gain from one role.
|
||||
|
||||
### Critic vs verifier
|
||||
|
||||
- A critic is an LLM reviewing an artifact for quality. Subjective. Can be fooled by plausible prose.
|
||||
- A verifier is a deterministic program running on the artifact. Objective. Gives pass/fail with evidence.
|
||||
|
||||
Use both. Critic catches taste issues the verifier cannot articulate. Verifier catches bugs the critic cannot see because they show up only at runtime.
|
||||
|
||||
### The anti-pattern
|
||||
|
||||
Every role in your system is an LLM and every role's output is "looks good to me." Classic MAST failure mode. Add at least one verifier whose pass/fail is decided by code, not by an LLM.
|
||||
|
||||
### Framework mappings
|
||||
|
||||
- **CrewAI** — `Agent(role, goal, backstory)` is the textbook specialization surface.
|
||||
- **LangGraph** — nodes can have specialized prompts; edges enforce the pipeline.
|
||||
- **AutoGen** — role-specific ConversableAgents with one-word names in a GroupChat.
|
||||
- **OpenAI Agents SDK** — handoff tools between role-specialized Agents.
|
||||
|
||||
## Build It
|
||||
|
||||
`code/main.py` implements a 4-role pipeline building a simple Python function:
|
||||
|
||||
- **Planner** produces a spec.
|
||||
- **Executor** generates a code string.
|
||||
- **Critic** (LLM-simulated) flags obvious issues.
|
||||
- **Verifier** runs the generated code in a sandbox (`exec`) against a test case.
|
||||
|
||||
Demo runs twice: once where the executor produces correct code (critic + verifier both pass), once where the executor produces off-spec code (critic misses the bug because it looks plausible, verifier catches it because the test fails).
|
||||
|
||||
Run:
|
||||
|
||||
```
|
||||
python3 code/main.py
|
||||
```
|
||||
|
||||
## Use It
|
||||
|
||||
`outputs/skill-role-designer.md` takes a task and produces the role roster (3-5 roles), the input/output schema per role, and the verifier check. Use this before wiring agents into a framework.
|
||||
|
||||
## Ship It
|
||||
|
||||
Checklist:
|
||||
|
||||
- **At least one deterministic verifier.** Never all-LLM.
|
||||
- **Explicit I/O schema per role.** The planner returns a spec, not prose; the executor reads that schema.
|
||||
- **Communicative dehallucination.** Executor must ask the planner when info is missing; never invent it.
|
||||
- **Critic/verifier ordering.** Run critic first (cheap, catches design issues), verifier second (slow, catches bugs).
|
||||
- **Loop budget.** Max 2 critic-executor revision rounds before escalating to human.
|
||||
|
||||
## Exercises
|
||||
|
||||
1. Run `code/main.py` and observe how the verifier catches the bug the critic missed. Add a static-analysis check (count occurrences of `return`) as an additional verifier. What does it catch that the runtime test misses?
|
||||
2. Add a 5th role: "requirements analyst" that translates user wish into planner-ready spec. What communicative dehallucination requests should flow up to it?
|
||||
3. Read MetaGPT Section 3 ("Agents"). List the input/output schema of each of MetaGPT's 5 roles.
|
||||
4. Read ChatDev's chat-chain diagram (arXiv:2307.07924 Figure 3). Identify where communicative dehallucination breaks a loop that would otherwise be infinite.
|
||||
5. PwC's 7× accuracy gain came from verification loops. Hypothesize three tasks where adding a verifier would not help — where deterministic checking of correctness is impossible or prohibitively expensive.
|
||||
|
||||
## Key Terms
|
||||
|
||||
| Term | What people say | What it actually means |
|
||||
|------|----------------|------------------------|
|
||||
| Role specialization | "Different agents, different jobs" | Distinct system prompts tuned for planner/executor/critic/verifier roles. |
|
||||
| SOP pattern | "Encoded standard operating procedure" | MetaGPT's framing: strict I/O schemas per role turn a team into a pipeline. |
|
||||
| Communicative dehallucination | "Ask before inventing" | ChatDev pattern: executor asks planner when a detail is missing rather than making one up. |
|
||||
| Critic | "LLM reviewer" | Subjective, opinionated reviewer. Catches taste issues. Can be fooled by plausible prose. |
|
||||
| Verifier | "Deterministic check" | Code-based pass/fail. Test runner, type checker, schema validator. Cannot be fooled. |
|
||||
| Verification gap | "No one checked" | 21.3% of MAST failures. Answer shipped without a check that would have caught the bug. |
|
||||
| Revision loop | "Critic sends it back" | Critic rejection triggers executor re-run with feedback. Needs a budget. |
|
||||
| All-LLM anti-pattern | "Looks good to me" | Every role is an LLM, no deterministic check. Classic MAST failure. |
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Hong et al. — MetaGPT: Meta Programming for Multi-Agent Collaboration](https://arxiv.org/abs/2308.00352) — the SOP-as-role-prompt reference paper
|
||||
- [Qian et al. — Communicative Agents for Software Development (ChatDev)](https://arxiv.org/abs/2307.07924) — chat chain + communicative dehallucination
|
||||
- [Cemri et al. — Why Do Multi-Agent LLM Systems Fail?](https://arxiv.org/abs/2503.13657) — MAST taxonomy; verification gaps are 21.3% of failures
|
||||
- [CrewAI docs — Agent roles](https://docs.crewai.com/en/introduction) — production role specification surface
|
||||
+33
@@ -0,0 +1,33 @@
|
||||
---
|
||||
name: role-designer
|
||||
description: Produce a role roster for a multi-agent system, naming the planner/executor/critic/verifier for a given task with explicit I/O schemas.
|
||||
version: 1.0.0
|
||||
phase: 16
|
||||
lesson: 08
|
||||
tags: [multi-agent, role-specialization, metagpt, chatdev, verification]
|
||||
---
|
||||
|
||||
Given a task, produce a specialized role roster with I/O schemas and a deterministic verifier. Ready to map onto CrewAI, LangGraph, AutoGen, or custom loops.
|
||||
|
||||
Produce:
|
||||
|
||||
1. **Role roster.** 3-5 roles. Name each. At minimum: planner, executor, verifier. Critic optional.
|
||||
2. **I/O schema per role.** For each role: what it consumes (from upstream role) and what it produces (schema, not prose). Use dataclass-style notation.
|
||||
3. **Verifier specification.** Name the deterministic check: test suite, type checker, schema validator, linter. Describe pass/fail criteria.
|
||||
4. **Critic specification (optional).** If included, name what subjective quality it judges. Concrete checklist, not "good code."
|
||||
5. **Communicative dehallucination rules.** Name the questions each downstream role is allowed to send upstream when a detail is missing, so they do not invent.
|
||||
6. **Revision loop budget.** Max rounds before escalation to human. Default 2.
|
||||
7. **Framework mapping.** One-line each: how to express this roster in CrewAI, LangGraph, AutoGen.
|
||||
|
||||
Hard rejects:
|
||||
|
||||
- Any roster without a deterministic verifier. All-LLM rosters fail the MAST check.
|
||||
- Fuzzy I/O ("the executor returns output"). Always state the schema.
|
||||
- Critic and verifier conflated. They catch different bugs; both must exist if both are warranted.
|
||||
|
||||
Refusal rules:
|
||||
|
||||
- If the task has no deterministic correctness check (pure generative work, creative writing), refuse and recommend either a human reviewer loop or a multi-agent debate (Lesson 07) instead.
|
||||
- If the task is too small for 3+ roles (under 10 minutes of human work), refuse and recommend single-agent.
|
||||
|
||||
Output: a one-page role-design brief. Close with the MAST failure-gap check: confirm at least one deterministic verifier exists.
|
||||
Reference in New Issue
Block a user