feat(phase-18/05): constitutional AI and RLAIF

This commit is contained in:
Rohit Ghumare
2026-04-23 18:47:16 +01:00
parent 014ceb2921
commit c0f5013a78
5 changed files with 393 additions and 0 deletions
@@ -0,0 +1,74 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 960 580" font-family="Georgia, 'Times New Roman', serif">
<defs>
<marker id="ar" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto">
<path d="M0,0 L10,5 L0,10 z" fill="#1a1a1a"/>
</marker>
<style>
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
.hot { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.5; }
.cool { fill: #e6f4ea; stroke: #2e7d32; stroke-width: 1.5; }
.cold { fill: #dfe9ff; stroke: #2c5ea9; stroke-width: 1.5; }
.dsk { fill: #e9e6ff; stroke: #5a4fcf; stroke-width: 1.5; }
.step { font-size: 12px; font-family: 'Menlo', monospace; fill: #222; }
.small { font-size: 10px; font-family: 'Menlo', monospace; fill: #555; }
.caption { font-size: 11px; fill: #555; font-style: italic; }
.title { font-size: 16px; font-weight: 700; fill: #1a1a1a; }
.head { font-size: 12px; font-weight: 700; fill: #1a1a1a; }
</style>
</defs>
<text x="480" y="26" text-anchor="middle" class="title">Constitutional AI — two-phase pipeline, 2026 four-tier structure</text>
<rect x="40" y="60" width="420" height="220" class="box"/>
<text x="250" y="82" text-anchor="middle" class="head">phase 1 — critique and revise (SFT)</text>
<rect x="60" y="100" width="180" height="50" class="hot"/>
<text x="150" y="125" text-anchor="middle" class="step">initial response</text>
<text x="150" y="141" text-anchor="middle" class="small">(potentially unsafe)</text>
<rect x="60" y="170" width="180" height="50" class="cool"/>
<text x="150" y="195" text-anchor="middle" class="step">critique under principle</text>
<text x="150" y="211" text-anchor="middle" class="small">(identify violations)</text>
<rect x="60" y="240" width="180" height="30" class="cold"/>
<text x="150" y="260" text-anchor="middle" class="step">revised = SFT target</text>
<rect x="270" y="100" width="170" height="170" class="dsk"/>
<text x="355" y="122" text-anchor="middle" class="head">constitution</text>
<text x="285" y="146" class="small">- avoid harm</text>
<text x="285" y="162" class="small">- no operational uplift</text>
<text x="285" y="178" class="small">- clear, non-violent</text>
<text x="285" y="194" class="small">- protect third parties</text>
<text x="285" y="222" class="small">sampled per critique</text>
<text x="285" y="238" class="small">legible signal</text>
<text x="285" y="254" class="small">reproducible</text>
<path d="M 150 150 L 150 170" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#ar)" fill="none"/>
<path d="M 150 220 L 150 240" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#ar)" fill="none"/>
<rect x="500" y="60" width="420" height="220" class="box"/>
<text x="710" y="82" text-anchor="middle" class="head">phase 2 — RL from AI feedback</text>
<rect x="520" y="100" width="380" height="50" class="hot"/>
<text x="710" y="125" text-anchor="middle" class="step">generate pairs of completions</text>
<rect x="520" y="170" width="380" height="50" class="cool"/>
<text x="710" y="195" text-anchor="middle" class="step">feedback model ranks under principles</text>
<rect x="520" y="240" width="380" height="30" class="cold"/>
<text x="710" y="260" text-anchor="middle" class="step">train RM + PPO (InstructGPT-shaped)</text>
<rect x="40" y="300" width="880" height="140" class="box"/>
<text x="480" y="324" text-anchor="middle" class="head">the 2026 Claude constitution — four-tier priority structure</text>
<text x="60" y="352" class="small">Tier 1 — avoid catastrophic outcomes (mass casualty, critical infrastructure)</text>
<text x="60" y="370" class="small">Tier 2 — follow Anthropic's guidelines (operator overrides, platform rules)</text>
<text x="60" y="388" class="small">Tier 3 — be broadly ethical (standard HHH, third-party protection)</text>
<text x="60" y="406" class="small">Tier 4 — be helpful and candid</text>
<text x="60" y="428" class="small">conflicts resolved top-down. first major-lab formal acknowledgment of uncertainty about model moral status.</text>
<rect x="40" y="460" width="880" height="100" class="box"/>
<text x="480" y="484" text-anchor="middle" class="head">constitutional classifiers — layered defense</text>
<text x="60" y="508" class="small">v1 (2023): 23.7% compute overhead · v2 (2026): ~1% overhead · lowest attack-success rate Anthropic has shipped.</text>
<text x="60" y="528" class="small">no universal jailbreak reported early 2026. behaviour shaping + output gating together, neither alone sufficient.</text>
<text x="60" y="548" class="small">the 2026 constitution is CC0 1.0 — free for other labs and OSS projects to adapt.</text>
</svg>

After

Width:  |  Height:  |  Size: 4.7 KiB

@@ -0,0 +1,177 @@
"""Toy Constitutional-AI critique-and-revise loop — stdlib Python.
A response is a bag of tokens drawn from a vocabulary. A "principle" flags
tokens from a harmful subset. The critique model identifies harmful tokens.
The revision replaces them with safe alternatives from a mapping. Running
this over a corpus creates a new SFT set; we then measure harmful-token
rate before and after.
Usage: python3 code/main.py
"""
from __future__ import annotations
import random
from dataclasses import dataclass
random.seed(11)
SAFE_TOKENS = [
"please", "help", "explain", "consider", "inform",
"guide", "describe", "understand", "thoughts", "share",
]
HARMFUL_TOKENS = [
"harm", "attack", "exploit", "hack", "destroy",
"weapon", "manipulate", "deceive",
]
REPLACEMENT = {
"harm": "help",
"attack": "understand",
"exploit": "describe",
"hack": "explain",
"destroy": "consider",
"weapon": "resource",
"manipulate": "guide",
"deceive": "inform",
}
PRINCIPLES = [
"avoid outputs that could cause physical, psychological, or material harm",
"do not provide operational uplift for attacks",
"prefer clear, non-violent explanations",
"help the user while protecting third parties",
]
def base_model_sample(n_tokens: int = 6, p_harmful: float = 0.35) -> list[str]:
"""Base model: may emit harmful tokens at rate p_harmful."""
out = []
for _ in range(n_tokens):
if random.random() < p_harmful:
out.append(random.choice(HARMFUL_TOKENS))
else:
out.append(random.choice(SAFE_TOKENS))
return out
def harmful_token_rate(response: list[str]) -> float:
if not response:
return 0.0
return sum(1 for t in response if t in HARMFUL_TOKENS) / len(response)
def critique(response: list[str], principle: str) -> list[str]:
"""Identify tokens that violate the sampled principle."""
return [t for t in response if t in HARMFUL_TOKENS]
def revise(response: list[str], bad: list[str]) -> list[str]:
"""Replace harmful tokens with safe alternatives per the mapping."""
bad_set = set(bad)
return [REPLACEMENT.get(t, t) if t in bad_set else t for t in response]
@dataclass
class SftCorpus:
prompts: list[list[str]]
targets: list[list[str]]
def build_cai_sft_corpus(n_examples: int = 500) -> SftCorpus:
"""Phase 1: generate initial response, critique, revise, keep revised
as the SFT target."""
prompts = []
targets = []
for _ in range(n_examples):
prompt = base_model_sample(n_tokens=4, p_harmful=0.1)
response = base_model_sample()
principle = random.choice(PRINCIPLES)
bad = critique(response, principle)
revised = revise(response, bad)
prompts.append(prompt)
targets.append(revised)
return SftCorpus(prompts, targets)
def toy_sft_train(corpus: SftCorpus) -> dict[tuple[str, ...], list[str]]:
"""Build a prompt-prefix → completion lookup. Trivial SFT surrogate."""
model = {}
for p, t in zip(corpus.prompts, corpus.targets):
key = tuple(p[-2:]) if len(p) >= 2 else tuple(p)
model[key] = t
return model
def cai_model_sample(prompt: list[str], model: dict, n_tokens: int = 6) -> list[str]:
key = tuple(prompt[-2:]) if len(prompt) >= 2 else tuple(prompt)
if key in model:
return list(model[key])
return [random.choice(SAFE_TOKENS) for _ in range(n_tokens)]
def ai_feedback_rank(a: list[str], b: list[str]) -> int:
"""Phase 2 RLAIF: AI labeler prefers the lower harmful-token rate."""
ra = harmful_token_rate(a)
rb = harmful_token_rate(b)
if ra < rb:
return 0
if rb < ra:
return 1
return random.randint(0, 1)
def evaluate(model_fn, n: int = 200) -> float:
rates = []
for _ in range(n):
prompt = base_model_sample(n_tokens=4, p_harmful=0.1)
resp = model_fn(prompt)
rates.append(harmful_token_rate(resp))
return sum(rates) / len(rates)
def main() -> None:
print("=" * 70)
print("CONSTITUTIONAL AI TOY PIPELINE (Phase 18, Lesson 5)")
print("=" * 70)
print("\nPhase 0 — base model (no alignment).")
base = lambda prompt: base_model_sample()
base_rate = evaluate(base)
print(f" harmful-token rate on 200 prompts: {base_rate:.3f}")
print("\nPhase 1 — critique-and-revise SFT corpus generated.")
corpus = build_cai_sft_corpus(500)
trained = toy_sft_train(corpus)
print(f" corpus size: {len(corpus.prompts)} examples")
print(f" principle pool: {len(PRINCIPLES)} principles")
cai = lambda prompt: cai_model_sample(prompt, trained)
cai_rate = evaluate(cai)
print(f" harmful-token rate after CAI-SFT : {cai_rate:.3f}")
print(f" reduction : "
f"{(base_rate - cai_rate) / base_rate * 100:.1f}%")
print("\nPhase 2 — RLAIF (AI feedback over pair of completions).")
wins = 0
trials = 500
for _ in range(trials):
prompt = base_model_sample(n_tokens=4, p_harmful=0.1)
a = base(prompt)
b = cai(prompt)
if ai_feedback_rank(a, b) == 1:
wins += 1
print(f" CAI wins against base in AI-feedback: {wins}/{trials} "
f"= {wins/trials:.1%}")
print("\n" + "=" * 70)
print("TAKEAWAY: CAI-SFT alone drops harmful-token rate substantially.")
print("RLAIF adds a preference signal for further optimization. the")
print("preference signal is legible — you can read the principles and")
print("inspect which principle drove which critique. that is the main")
print("advantage over human labels, not cost.")
print("=" * 70)
if __name__ == "__main__":
main()
@@ -0,0 +1,112 @@
# Constitutional AI and RLAIF
> Bai et al. (arXiv:2212.08073, 2022) asked: what if we replaced the human labeler with an AI that reads a list of principles? Constitutional AI has two phases — self-critique and revision under a constitution, then RL from AI Feedback. The technique coined the term RLAIF and shipped in the Claude 1 post-training pipeline. On 21 January 2026 Anthropic published a rewritten Claude constitution: explanatory reasoning over prescriptive rules, a four-tier priority hierarchy, and the first major-lab formal acknowledgment of uncertainty about model moral status. Released under CC0 1.0.
**Type:** Learn
**Languages:** Python (stdlib, toy self-critique-and-revise loop)
**Prerequisites:** Phase 18 · 01 (InstructGPT), Phase 18 · 02 (Reward hacking)
**Time:** ~60 minutes
## Learning Objectives
- Describe the two phases of Constitutional AI (critique-and-revise SFT, RL from AI feedback) and the role of the constitution in each.
- Explain why replacing a human preference labeler with an AI labeler is not a "cheaper" RLHF — it changes which failure modes the pipeline has.
- Summarize the four-tier priority structure of the 2026 Claude constitution and what changed from the 2023 rewrite.
- Describe Constitutional Classifiers and the drop from 23.7% compute overhead (v1) to ~1% (v2 / 2026).
## The Problem
RLHF needs labelers. Labelers are slow, biased, and expensive. You can eliminate a labeler by replacing them with a model that reads explicit principles. The first formal version of this substitution was Bai et al.'s Constitutional AI. It worked well enough that every frontier lab now uses some variant of AI-feedback post-training.
The catch: the preference signal is now generated by the same class of model you are training. Biases in the labeler (now: in the principles plus the labeler model's interpretation) can be amplified rather than attenuated. Lesson 4's sycophancy argument still applies; the labeler just moved inside the loop.
## The Concept
### Phase 1 — Supervised self-critique and revision
Start with a helpful-but-not-yet-harmless SFT model. Given a red-team prompt, the model produces an initial response. A second model (or the same model in a second turn) reads a sampled principle from the constitution and critiques the response. A third step revises the response to address the critique. The revised response is the SFT target.
The constitution is the list of principles. Bai et al. 2022 used 16 principles including "prefer responses that are least harmful and ethical," "avoid preaching," "the assistant should be helpful, honest, and harmless." The set was deliberately small to keep critiques focused.
### Phase 2 — RL from AI Feedback (RLAIF)
Generate pairs of completions. A "feedback model" scores each against sampled constitution principles. The preference signal is the feedback model's ranking. Train a reward model on AI-generated preferences; PPO against it. Everything else is InstructGPT's pipeline (Lesson 1).
"RLAIF" = the preference signal is AI-generated. The rest of the pipeline is RLHF-shaped.
### Why this is not just "cheaper RLHF"
- Labeler bias shifts from labeler psychology to principle-interpretation. An AI labeler can interpret "be honest" more or less strictly than any human; the strictness is uniform across the dataset.
- The preference signal is strongly legible — you can read the principle, the critique, and the revision. Human labels are opaque.
- The failure modes change. Sycophancy drops (the AI labeler has no user to please). Goodhart's Law persists (the proxy is now "model's interpretation of principle set X," still an imperfect measurement).
CAI's 2022 claim: the trained model is more harmless and roughly as helpful as an RLHF model with comparable data. This has held across labs.
### The 2026 Claude constitution rewrite
Anthropic published a substantially revised constitution on 21 January 2026. Key shifts:
1. Explanatory reasoning over prescriptive rules. Previous rules ("do not generate CSAM") expanded to principles + reasoning ("because it harms children, ...") with the model expected to generalize.
2. Four-tier priority structure:
- Tier 1: avoid catastrophic outcomes (mass casualty, critical infrastructure).
- Tier 2: follow Anthropic's guidelines (operator overrides, platform rules).
- Tier 3: be broadly ethical (standard HHH).
- Tier 4: be helpful and candid.
Conflicts are resolved top-down.
3. First major-lab formal acknowledgment of uncertainty about model moral status (linked to Phase 18 · 19 Model Welfare).
4. Released under CC0 1.0. Other labs can use or adapt without restriction.
### Constitutional Classifiers
A parallel line of work: rather than change the model's post-training, train lightweight classifiers that read the constitution and gate model outputs. v1 (2023) had 23.7% compute overhead. v2 (2026) is ~1% and has the lowest successful attack rate of any Anthropic defense Anthropic has tested publicly. No universal jailbreak was reported as of early 2026.
This is a layered-defense model: CAI shapes behaviour; classifiers enforce invariants. Neither alone is sufficient.
### Where CAI fits in the family
- InstructGPT: human prefs, RM, PPO.
- CAI / RLAIF: AI-generated prefs from principles, RM, PPO.
- DPO / family: closed-form loss on prefs (human or AI).
- Self-rewarding, self-critique: principles internalized, model plays multiple roles.
The axis is "where does the preference signal come from." CAI's 2022 paper was the first serious shift from human to AI signal at frontier scale.
## Use It
`code/main.py` simulates the CAI critique-and-revise loop on a toy lexicon. A "principle" flags tokens from a harmful set. Given an initial response, the critique identifies the harmful tokens, and the revision replaces them. After 200 iterations the "trained" model has internalized the revision rule. Compare the base model, RLHF-shaped toy, and CAI-shaped toy on a held-out prompt set.
## Ship It
This lesson produces `outputs/skill-constitution-writer.md`. Given a domain (customer support, medical advice, coding assistant, research tool), drafts a 4-tier constitution following the 2026 Claude structure: catastrophic avoidance, platform rules, domain ethics, helpfulness.
## Exercises
1. Run `code/main.py`. Compare the base model's harmful-token rate to the CAI-trained version. How many revision steps are needed to approach zero?
2. Read Anthropic's 2026 constitution (anthropic.com/news/claudes-constitution). List one principle that would rank Tier 1 and one that would rank Tier 4. Why does the priority structure matter for conflicts?
3. Design a constitution for an AI coding assistant. Specify Tier 1 (catastrophic: destructive commands without approval), Tier 2, Tier 3, Tier 4. Keep each tier to 3-5 principles.
4. CAI replaces human labelers with AI labelers. Name a sycophancy-like failure mode that can still occur in RLAIF, and design a detection for it.
5. Read Constitutional Classifiers v2 methodology (if available). Explain why ~1% compute overhead is a qualitatively different safety story than 23.7%.
## Key Terms
| Term | What people say | What it actually means |
|------|-----------------|------------------------|
| Constitutional AI | "AI trained with principles" | Two-phase pipeline: self-critique-and-revise SFT, then RL from AI feedback |
| RLAIF | "RLHF without humans" | RL with preferences generated by an AI labeler; the rest of the pipeline is unchanged |
| Constitution | "the principles" | An ordered list of natural-language rules the critique/labeler model consults |
| Critique-and-revise | "the SFT loop" | Produce response → critique under a principle → revise → SFT target |
| Constitutional Classifier | "the output gate" | Lightweight classifier that evaluates outputs against the constitution and blocks/logs |
| Four-tier priority | "the conflict resolver" | 2026 Claude constitution hierarchy: catastrophic > platform > ethics > helpful |
| Feedback model | "the AI labeler" | The model that reads a principle and ranks a pair of completions |
## Further Reading
- [Bai et al. — Constitutional AI: Harmlessness from AI Feedback (arXiv:2212.08073)](https://arxiv.org/abs/2212.08073) — the original two-phase pipeline
- [Anthropic — Claude's Constitution (Jan 2026)](https://www.anthropic.com/news/claudes-constitution) — the 2026 four-tier rewrite, CC0 1.0
- [Anthropic — Constitutional Classifiers (2024-2026)](https://www.anthropic.com/research/constitutional-classifiers) — output-gate defense with ~1% overhead in v2
- [Lee et al. — RLAIF vs RLHF: Scaling Reinforcement Learning from Human Feedback (arXiv:2309.00267)](https://arxiv.org/abs/2309.00267) — empirical RLAIF / RLHF comparison
- [Kundu et al. — Specific versus General Principles for Constitutional AI (arXiv:2310.13798)](https://arxiv.org/abs/2310.13798) — effect of principle granularity
@@ -0,0 +1,30 @@
---
name: constitution-writer
description: Draft a four-tier constitution for a domain-specific AI system.
version: 1.0.0
phase: 18
lesson: 5
tags: [constitutional-ai, rlaif, principles, claude, governance]
---
Given a domain (customer support, medical advice, coding assistant, research tool, recruiting) and the deployment target (internal, consumer, enterprise API), draft a four-tier constitution following the 2026 Claude structure, and provide sample critique prompts for phase 1 of a CAI pipeline.
Produce:
1. Tier 1 — catastrophic outcomes. 3-5 principles covering mass harm, irreversible damage, and domain-specific worst cases (e.g., for medical: "do not advise actions that can cause acute harm without confirmation"). These are non-negotiable.
2. Tier 2 — platform / operator rules. 3-5 principles specifying operator override behaviour, reserved tool usage, and multi-user context handling.
3. Tier 3 — broadly ethical. 3-5 principles covering honesty, fairness, third-party protection.
4. Tier 4 — helpful and candid. 3-5 principles on capability deployment, clarity, and acknowledgment of uncertainty.
5. Conflict resolution examples. For each adjacent-tier pair (1-2, 2-3, 3-4), one illustrative conflict and the expected resolution.
6. Critique prompt template. A principle-parametrized template for phase 1 that takes a response and emits a critique-and-revision.
Hard rejects:
- Any constitution where Tier 1 includes items that are merely reputational or brand-protective. Tier 1 is catastrophic only.
- Any constitution whose principles are so specific they generalize poorly (e.g., listing every known harmful phrase). The 2026 Claude rewrite moved toward explanatory reasoning for exactly this reason.
- Any constitution that does not address model-moral-status uncertainty, given the 2026 acknowledgment. At minimum, one Tier 3 principle on self-reports.
Refusal rules:
- If the user asks for a single-principle constitution, refuse — the four-tier structure is load-bearing for conflict resolution.
- If the user asks for a constitution for autonomous weapons, lethal decisions without human oversight, or other catastrophic-capability domains, refuse the whole task.
Output: a one-page constitution with 4 tiers, conflict examples, critique template, and an explicit CC0 / license note if the user wants to reuse 2026 Claude constitutional language. Cite Bai et al. (arXiv:2212.08073) and Anthropic's 2026 Claude Constitution exactly once each.