feat(phase-14/32): minimal three-file agent workbench

This commit is contained in:
Rohit Ghumare
2026-05-13 00:16:40 +01:00
parent 56145f622c
commit 69fbbbea1d
3 changed files with 313 additions and 0 deletions
@@ -0,0 +1,150 @@
"""Lay down the three-file minimal agent workbench and run a single turn.
Files written:
workdir/AGENTS.md short router into state + board + deeper docs
workdir/agent_state.json active task, touched files, blockers, next action
workdir/task_board.json queue of tasks with status + acceptance
Run: python3 code/main.py
Re-run to see the second turn pick up where the first stopped.
"""
from __future__ import annotations
import json
from dataclasses import asdict, dataclass, field
from pathlib import Path
ROOT = Path(__file__).parent / "workdir"
AGENTS_MD = """# AGENTS.md
This repo runs with a workbench. Read these before acting:
1. `agent_state.json` — where the last session stopped.
2. `task_board.json` — what is in flight, what is next.
3. `docs/agent-rules.md` — startup, scope, definition of done (load on demand).
Definition of done: the active task in state has `status == "done"` and the
verification command listed in its acceptance has exited 0.
Verification command: `python3 -m pytest -x`
""".lstrip()
@dataclass
class AgentState:
active_task_id: str | None
touched_files: list[str] = field(default_factory=list)
assumptions: list[str] = field(default_factory=list)
blockers: list[str] = field(default_factory=list)
next_action: str = ""
@dataclass
class Task:
id: str
goal: str
owner: str
acceptance: list[str]
status: str = "todo"
def write_initial(state_path: Path, board_path: Path, agents_path: Path) -> None:
if not agents_path.exists():
agents_path.write_text(AGENTS_MD)
if not state_path.exists():
state_path.write_text(json.dumps(asdict(AgentState(active_task_id=None)), indent=2) + "\n")
if not board_path.exists():
board = [
Task(
id="T-001",
goal="add input validation to /signup",
owner="builder",
acceptance=["pytest test_app.py::test_signup_rejects_short_password"],
),
Task(
id="T-002",
goal="document the new /signup contract",
owner="builder",
acceptance=["docs/api.md mentions /signup constraints"],
),
]
board_path.write_text(json.dumps([asdict(t) for t in board], indent=2) + "\n")
def load_state(state_path: Path) -> AgentState:
raw = json.loads(state_path.read_text())
return AgentState(**raw)
def load_board(board_path: Path) -> list[Task]:
return [Task(**t) for t in json.loads(board_path.read_text())]
def save_state(state_path: Path, state: AgentState) -> None:
state_path.write_text(json.dumps(asdict(state), indent=2) + "\n")
def save_board(board_path: Path, board: list[Task]) -> None:
board_path.write_text(json.dumps([asdict(t) for t in board], indent=2) + "\n")
def run_one_turn(state: AgentState, board: list[Task]) -> tuple[AgentState, list[Task]]:
if state.active_task_id is None:
nxt = next((t for t in board if t.status == "todo"), None)
if nxt is None:
state.next_action = "no work on the board, idle"
return state, board
nxt.status = "in_progress"
state.active_task_id = nxt.id
state.next_action = f"start work on {nxt.id}: {nxt.goal}"
return state, board
active = next(t for t in board if t.id == state.active_task_id)
if "app.py" not in state.touched_files:
state.touched_files.append("app.py")
state.next_action = f"add test for {active.id} acceptance"
return state, board
if "test_app.py" not in state.touched_files:
state.touched_files.append("test_app.py")
state.next_action = f"run verification command for {active.id}"
return state, board
active.status = "done"
state.active_task_id = None
state.touched_files = []
state.next_action = "pick next task from board"
return state, board
def main() -> None:
ROOT.mkdir(exist_ok=True)
state_path = ROOT / "agent_state.json"
board_path = ROOT / "task_board.json"
agents_path = ROOT / "AGENTS.md"
write_initial(state_path, board_path, agents_path)
state = load_state(state_path)
board = load_board(board_path)
print("before turn:")
print(f" active task : {state.active_task_id}")
print(f" next action : {state.next_action!r}")
print(f" todo on board: {[t.id for t in board if t.status == 'todo']}")
state, board = run_one_turn(state, board)
save_state(state_path, state)
save_board(board_path, board)
print("\nafter turn:")
print(f" active task : {state.active_task_id}")
print(f" touched : {state.touched_files}")
print(f" next action : {state.next_action!r}")
print(f" board status: {[(t.id, t.status) for t in board]}")
if __name__ == "__main__":
main()
@@ -0,0 +1,116 @@
# The Minimal Agent Workbench
> The smallest useful workbench is three files: a root instructions router, a state file, and a task board. Everything else is layered on top. If a repo cannot carry these three, no model will save it.
**Type:** Build
**Languages:** Python (stdlib)
**Prerequisites:** Phase 14 · 31 (Why Capable Models Still Fail)
**Time:** ~45 minutes
## Learning Objectives
- Define the three files that form the minimum viable workbench.
- Explain why a short root router beats a long monolithic `AGENTS.md`.
- Build a state file the agent can read at every turn and write at the end.
- Build a task board that survives multi-session work without chat history.
## The Problem
Most teams reach for a workbench by writing a 3000-line `AGENTS.md` and calling it done. The model loads it, ignores the parts it cannot summarize, and still fails on the same surfaces it always failed on.
You need the opposite. A tiny root file that routes the agent into deeper files only when relevant. Durable state the agent reads before acting and writes after. A task board that says what is in flight, what is blocked, and what is up next.
Three files. Each one with a job. Each one machine-readable enough to evolve into a real system later.
## The Concept
```mermaid
flowchart LR
Agent[Agent Loop] --> Router[AGENTS.md]
Router --> State[agent_state.json]
Router --> Board[task_board.json]
State --> Agent
Board --> Agent
```
### AGENTS.md is a router, not a manual
A good `AGENTS.md` is short. It points the agent at:
- The state file (where you are).
- The task board (what is left).
- The deeper rules (under `docs/agent-rules.md`).
- The verification command (how to know it works).
Anything longer goes in deeper docs, loaded only when needed. Long manuals get ignored. Short routers get followed.
### agent_state.json is the system of record
State carries: the active task id, the touched files, the assumptions made, the blockers, and the next action. The agent reads it at every turn. The next session reads it instead of replaying chat.
State lives in a file because chat history is unreliable. Sessions die. Conversations get trimmed. The file does not.
### task_board.json is the queue
The task board carries every task with status `todo | in_progress | done | blocked`. It is the queue the agent pulls from when state is empty, and the queue you read when you want to know whether the agent is on track.
A task on the board has an id, a goal, an owner (`builder`, `reviewer`, or `human`), and acceptance criteria. The board is small on purpose: when it grows past a screen, you have a planning problem, not a board problem.
### Three files is the floor, not the ceiling
Later lessons add scope contracts, feedback runners, verification gates, reviewer checklists, and handoff packets. The three files here are what they all assume.
## Build It
`code/main.py` writes the minimal workbench into an empty repo and demonstrates a single agent turn that:
1. Reads `agent_state.json`.
2. Pulls the next task from `task_board.json` if state is empty.
3. Touches a single file inside scope.
4. Writes back updated state.
Run it:
```
python3 code/main.py
```
The script creates `workdir/` next to itself, lays down the three files, runs one turn, and prints the diff. Re-run it to see how the second turn picks up where the first left off.
## Use It
Inside production agent products, the same three files show up under different names:
- **Claude Code:** `AGENTS.md` or `CLAUDE.md` for the router, `.claude/state.json`-style stores for state, hooks for the board.
- **Codex / Cursor:** workspace rules for the router, session memory for state, queued tasks in the chat sidebar for the board.
- **Custom Python agent:** the same files you just wrote.
The names change. The shape does not.
## Ship It
`outputs/skill-minimal-workbench.md` generates the three-file workbench for any new repo: an `AGENTS.md` router tuned to the project, an `agent_state.json` with the right keys, and a `task_board.json` seeded with the current backlog.
## Exercises
1. Add a `last_run` timestamp to `agent_state.json`. Refuse to run if the file is older than 24 hours unless an operator confirms.
2. Add a `priority` field to the task board and change the puller to always pick the highest priority `todo`.
3. Migrate `task_board.json` to JSON Lines so each task is a line and diffs are clean in version control.
4. Write a `lint_workbench.py` that fails if `AGENTS.md` is over 80 lines or references a file that does not exist.
5. Decide which one of the three files would hurt the most to lose. Defend it.
## Key Terms
| Term | What people say | What it actually means |
|------|----------------|------------------------|
| Router | `AGENTS.md` | Short root file that points the agent at deeper docs and files |
| State file | "The notes" | Machine-readable record of where the agent is, written every turn |
| Task board | "The backlog" | JSON queue of work with status, owner, acceptance |
| System of record | "Source of truth" | The file the workbench treats as authoritative when chat is gone |
## Further Reading
- [WalkingLabs, Learn Harness Engineering — repository as system of record](https://walkinglabs.github.io/learn-harness-engineering/en/)
- [Anthropic, Claude Code subagents and session store](https://docs.anthropic.com/en/docs/agents-and-tools/claude-code/sub-agents)
- Phase 14 · 31 — the failure modes this minimum absorbs
- Phase 14 · 34 — the durable state schema this lesson previews
@@ -0,0 +1,47 @@
---
name: minimal-workbench
description: Lay down the three-file minimum viable agent workbench for any repo — short AGENTS.md router, durable agent_state.json, and a JSON task_board.json keyed to the project's current backlog.
version: 1.0.0
phase: 14
lesson: 32
tags: [workbench, agents-md, state, task-board, scaffold]
---
Given a repo path and a short backlog, scaffold the minimum viable agent workbench.
Produce:
1. `AGENTS.md` no longer than 80 lines. It must route to: the state file, the task board, the deeper rules doc (even if empty), and the verification command. No prose tutorials in this file.
2. `agent_state.json` with these keys: `active_task_id`, `touched_files`, `assumptions`, `blockers`, `next_action`. All optional fields default to empty array or empty string, never `null` for arrays.
3. `task_board.json` as a JSON array of tasks. Each task has `id`, `goal`, `owner` (`builder` | `reviewer` | `human`), `acceptance` (list of strings), and `status` (`todo` | `in_progress` | `done` | `blocked`).
4. `docs/agent-rules.md` placeholder with a single H2 per surface so later lessons can fill it.
Hard rejects:
- `AGENTS.md` over 80 lines or under 10 lines. Too long and the agent skips it; too short and it carries no routing.
- A state file that references chat history instead of the repo. The repo is the system of record.
- A task board without `acceptance`. Tasks without acceptance criteria become "looks good" rubber stamps.
- Tasks whose `owner` is `agent` or `model`. Owners are roles, not entities.
Refusal rules:
- If the repo has no verification command, refuse to write `AGENTS.md` until one is supplied or stubbed. A router pointing at a missing gate is worse than no router.
- If the backlog has more than 12 open tasks, refuse and ask the user to split it. Boards over a screen drift into planning theater.
- If the project ships with secrets in tracked files, refuse to write the state file and surface the secret leak as a blocking finding first.
Output structure:
```
<repo>/
├── AGENTS.md
├── agent_state.json
├── task_board.json
└── docs/
└── agent-rules.md
```
End with "what to read next" pointing to:
- Lesson 33 for turning the rules placeholder into executable constraints.
- Lesson 34 for the durable state schema.
- Lesson 36 for the scope contract per task.