feat(phase-15/12): durable execution for long-running agents

This commit is contained in:
Rohit Ghumare
2026-04-24 12:02:47 +01:00
parent e957ef9ec3
commit ba52aa6e5b
5 changed files with 398 additions and 0 deletions
@@ -0,0 +1,74 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 880 540" font-family="Georgia, 'Times New Roman', serif">
<defs>
<marker id="arrow" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto">
<path d="M0,0 L10,5 L0,10 z" fill="#1a1a1a"/>
</marker>
<style>
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
.hot { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.5; }
.cool { fill: #e6f4ea; stroke: #2e7d32; stroke-width: 1.5; }
.cold { fill: #dfe9ff; stroke: #2c5ea9; stroke-width: 1.5; }
.label { font-size: 13px; font-weight: 600; fill: #1a1a1a; }
.content { font-size: 11px; font-family: 'Menlo', monospace; fill: #333; }
.small { font-size: 10px; font-family: 'Menlo', monospace; fill: #555; }
.title { font-size: 16px; font-weight: 700; fill: #1a1a1a; }
</style>
</defs>
<text x="440" y="26" text-anchor="middle" class="title">Durable execution: replay re-runs workflow, logged activities skip</text>
<rect x="40" y="50" width="800" height="460" class="box"/>
<!-- Attempt 1 -->
<text x="60" y="80" class="label">attempt 1 — crash after call_llm</text>
<rect x="60" y="96" width="150" height="50" class="cool"/>
<text x="135" y="118" text-anchor="middle" class="label">fetch_docs</text>
<text x="135" y="136" text-anchor="middle" class="small">run, log done</text>
<rect x="230" y="96" width="150" height="50" class="cool"/>
<text x="305" y="118" text-anchor="middle" class="label">call_llm</text>
<text x="305" y="136" text-anchor="middle" class="small">run, log done</text>
<rect x="400" y="96" width="150" height="50" class="hot"/>
<text x="475" y="118" text-anchor="middle" class="label">CRASH</text>
<text x="475" y="136" text-anchor="middle" class="small">write_report not yet run</text>
<!-- Event log -->
<rect x="60" y="170" width="490" height="74" class="cold"/>
<text x="80" y="192" class="label">event log (durable)</text>
<text x="80" y="212" class="content">[fetch_docs started, fetch_docs done (result=X),</text>
<text x="80" y="228" class="content"> call_llm started, call_llm done (result=Y)]</text>
<!-- Attempt 2 -->
<text x="60" y="276" class="label">attempt 2 — replay</text>
<rect x="60" y="292" width="150" height="50" class="cold"/>
<text x="135" y="314" text-anchor="middle" class="label">fetch_docs</text>
<text x="135" y="332" text-anchor="middle" class="small">replay from log</text>
<rect x="230" y="292" width="150" height="50" class="cold"/>
<text x="305" y="314" text-anchor="middle" class="label">call_llm</text>
<text x="305" y="332" text-anchor="middle" class="small">replay from log</text>
<rect x="400" y="292" width="150" height="50" class="cool"/>
<text x="475" y="314" text-anchor="middle" class="label">write_report</text>
<text x="475" y="332" text-anchor="middle" class="small">run, log done</text>
<rect x="570" y="292" width="150" height="50" class="box"/>
<text x="645" y="314" text-anchor="middle" class="label">return report</text>
<text x="645" y="332" text-anchor="middle" class="small">done</text>
<path d="M210 317 L230 317" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
<path d="M380 317 L400 317" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
<path d="M550 317 L570 317" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
<!-- Why determinism -->
<rect x="60" y="370" width="760" height="120" class="hot"/>
<text x="440" y="394" text-anchor="middle" class="label">what makes replay work</text>
<text x="80" y="416" class="content">1. Workflow code is deterministic (no wall clock, no random, no hidden I/O).</text>
<text x="80" y="432" class="content">2. Every activity is logged with args + result; replay re-reads the log.</text>
<text x="80" y="448" class="content">3. Backend survives the crash (PostgreSQL / Durable Objects; not SQLite for prod).</text>
<text x="80" y="464" class="content">4. HITL pauses are a workflow signal, not a polling loop.</text>
<text x="80" y="480" class="content">LLM calls fit the activity shape: non-deterministic, expensive, potentially failing.</text>
</svg>

After

Width:  |  Height:  |  Size: 4.1 KiB

@@ -0,0 +1,171 @@
"""Minimal durable-execution engine — stdlib Python.
Models the workflow / activity / event-log pattern used by Temporal, LangGraph
checkpointing, Microsoft Agent Framework, and Claude Code Routines.
Activities are logged with inputs before execution and outputs after. A
replay of a workflow re-runs the workflow code but returns cached outputs
for activities whose event is already in the log. A crash mid-run loses
only the incomplete activity.
"""
from __future__ import annotations
import functools
import json
import os
import tempfile
from dataclasses import dataclass
# ---------- Event log ----------
@dataclass
class EventLog:
path: str
def __post_init__(self) -> None:
if not os.path.exists(self.path):
with open(self.path, "w") as f:
json.dump([], f)
def events(self) -> list[dict]:
with open(self.path) as f:
return json.load(f)
def append(self, ev: dict) -> None:
evs = self.events()
evs.append(ev)
with open(self.path, "w") as f:
json.dump(evs, f)
def lookup(self, name: str, args: tuple) -> dict | None:
for ev in self.events():
if ev["name"] == name and ev["args"] == list(args) and ev["status"] == "done":
return ev
return None
# ---------- Activity decorator ----------
def activity(name: str):
def deco(fn):
@functools.wraps(fn)
def wrapper(log: EventLog, *args):
hit = log.lookup(name, args)
if hit:
print(f" [replay] {name}({args}) -> {hit['result']} (from log)")
return hit["result"]
log.append({"name": name, "args": list(args), "status": "started"})
result = fn(*args)
log.append({"name": name, "args": list(args),
"status": "done", "result": result})
print(f" [run] {name}({args}) -> {result}")
return result
return wrapper
return deco
# ---------- Example activities ----------
@activity("fetch_docs")
def fetch_docs(query: str) -> int:
# Pretend to hit an API; return number of docs.
return len(query) * 3
@activity("call_llm")
def call_llm(doc_count: int) -> str:
# Pretend LLM call; deterministic here for pedagogy.
return f"summary({doc_count}_docs)"
@activity("write_report")
def write_report(summary: str) -> str:
# Pretend tool call with a side effect.
return f"report://{summary}"
# ---------- Workflow ----------
def workflow(log: EventLog, query: str, crash_after: int = -1) -> str:
"""Three-activity workflow with an optional crash for pedagogy."""
doc_count = fetch_docs(log, query)
if crash_after == 1:
raise RuntimeError("simulated crash after fetch_docs")
summary = call_llm(log, doc_count)
if crash_after == 2:
raise RuntimeError("simulated crash after call_llm")
report = write_report(log, summary)
return report
# ---------- Driver ----------
def reset_log(path: str) -> EventLog:
if os.path.exists(path):
os.remove(path)
return EventLog(path)
def count_runs(log: EventLog) -> int:
return sum(1 for ev in log.events() if ev["status"] == "started")
def main() -> None:
print("=" * 70)
print("DURABLE EXECUTION (Phase 15, Lesson 12)")
print("=" * 70)
tmpdir = tempfile.mkdtemp()
# Naive retry: lose the event log on crash. Every restart re-runs
# everything.
print("\nNaive retry (no event log persisted)")
print("-" * 70)
for attempt in range(1, 4):
log = reset_log(os.path.join(tmpdir, "naive.json"))
print(f" attempt {attempt}:")
try:
crash = 2 if attempt == 1 else -1
r = workflow(log, "hello", crash_after=crash)
print(f" -> result {r}")
print(f" -> {count_runs(log)} activity starts this attempt")
break
except RuntimeError as e:
print(f" -> crash: {e}; {count_runs(log)} activity starts wasted")
# Durable retry: keep the event log across attempts; replay does not
# re-execute completed activities.
print("\nDurable retry (event log preserved across attempts)")
print("-" * 70)
durable_path = os.path.join(tmpdir, "durable.json")
if os.path.exists(durable_path):
os.remove(durable_path)
for attempt in range(1, 4):
log = EventLog(durable_path)
print(f" attempt {attempt}:")
try:
crash = 2 if attempt == 1 else -1
r = workflow(log, "hello", crash_after=crash)
print(f" -> result {r}")
print(f" -> {count_runs(log)} total activity starts across attempts")
break
except RuntimeError as e:
print(f" -> crash: {e}")
print()
print("=" * 70)
print("HEADLINE: durability makes long-horizon runs affordable to fail")
print("-" * 70)
print(" Naive retry re-executes every activity on every attempt.")
print(" Durable retry replays completed activities from the log;")
print(" only the missing activity actually runs. Same design used")
print(" by Temporal, LangGraph checkpointing, Microsoft Agent")
print(" Framework, and Claude Code Routines. The LLM call is")
print(" just another non-deterministic activity in the log.")
if __name__ == "__main__":
main()
@@ -0,0 +1,112 @@
# Long-Running Background Agents: Durable Execution
> Production long-horizon agents do not run in `while True`. Every LLM call becomes an activity with checkpoint, retry, and replay. Temporal's OpenAI Agents SDK integration went GA March 2026. Claude Code Routines (Anthropic) runs scheduled Claude Code invocations without a persistent local process. Sessions pause on human-input, survive deploys, and resume from the latest checkpoint keyed by `thread_id`. Behind the new ergonomics sits an old pattern — workflow orchestration — with one new input: LLM calls as non-deterministic activities that must be deterministically replayed on recovery.
**Type:** Learn
**Languages:** Python (stdlib, minimal durable-execution state machine)
**Prerequisites:** Phase 15 · 10 (Permission modes), Phase 15 · 01 (Long-horizon agents)
**Time:** ~60 minutes
## The Problem
Consider an agent that runs for four hours. It calls three tools, prompts the user twice, and makes forty LLM calls. Halfway through, the host it is running on reboots. What happens?
- In a naive `while True` loop: everything is lost. The run restarts from scratch. The three tool calls (with real side effects) execute again. The user is prompted again for things they already approved. Forty LLM calls are re-billed.
- With durable execution: the run resumes from the most recent checkpoint. Already-completed activities are not re-executed; their results are replayed from the durable log. The user does not re-approve things they already approved. The LLM calls already made are not re-billed.
This is the same pattern workflow engines have shipped for a decade (Temporal, Cadence, Uber's Cherami). What's new is that LLM calls are now a kind of activity — non-deterministic, expensive, with side effects — and they fit this pattern cleanly.
The running theme of the lesson: long-horizon reliability decays (METR observes a "35-minute degradation" — success rate drops roughly quadratically with horizon). Durable execution enables runs that are longer than the reliability profile supports, which is a new way to fail safely if the design is right and unsafely if the design is wrong.
## The Concept
### Activities, workflows, and replay
- **Workflow**: deterministic orchestration code. Defines the sequence of activities, the branches, the waits. Must be deterministic so it can be replayed from the event log without surprising divergence.
- **Activity**: a non-deterministic, potentially failing unit of work. LLM call, tool call, file write, HTTP request. Each activity is logged with its inputs and (once complete) its outputs.
- **Event log**: the durable backing store. Every activity start, complete, fail, retry, and every workflow decision is recorded.
- **Replay**: on recovery, the workflow code re-runs from the start; every activity that already completed returns its logged result without re-executing. Only activities that had not completed are actually run.
This is the same shape as React re-rendering against a virtual DOM, or Git rebuilding a working tree from commits. Determinism in the orchestrator is what makes durability cheap.
### Why LLM calls fit the pattern
LLM calls are:
- Non-deterministic (temperature > 0; even temperature 0 drifts across model versions).
- Expensive (money and latency).
- Potentially failing (rate limits, timeouts).
- Side-effectful (if they invoke tools).
This is exactly the activity profile. Wrapping every LLM call as an activity gives you retry with exponential backoff, checkpointing across restarts, and a replayable trace for debugging.
### Checkpoints keyed by `thread_id`
LangGraph, Microsoft Agent Framework, Cloudflare Durable Objects, and Claude Code Routines all converged on the same API shape: a `thread_id` (or equivalent) identifies the session; each state transition persists to a backend (PostgreSQL default, SQLite for dev, Redis for cache); resume reads the latest checkpoint.
The backend choice matters:
- **PostgreSQL**: durable, queryable, survives deploys. Default for LangGraph.
- **SQLite**: local-dev only; loses data across hosts.
- **Redis**: fast but ephemeral unless AOF/snapshot configured.
- **Cloudflare Durable Objects**: transparently distributed; scoped by a unique key; survives for hours to weeks.
### Human-input as a first-class state
Propose-then-commit (Lesson 15) requires a durable "waiting on human" state. The workflow pauses, the external queue holds the pending request, and an approval resumes from exactly that point. Without durability this is best-effort; with it, an overnight approval arrives and the workflow picks up in the morning.
### The 35-minute degradation
METR observed that every agent class measured shows reliability decay beyond ~35 minutes of continuous operation. Doubling the task duration roughly quadruples the failure rate. Durable execution does not fix this; it lets you run longer than the reliability profile supports. The safe pattern is to combine durability with checkpoints that require fresh HITL on re-entry, and with budget kill switches (Lesson 13) that cap total compute regardless of wall-clock time.
### When durable execution is the wrong answer
- Runs shorter than a few minutes with no human input. Overhead > benefit.
- Strictly read-only information retrieval.
- Tasks where correctness requires end-to-end within one context window (some reasoning tasks; some one-shot generation).
## Use It
`code/main.py` implements a minimal durable-execution engine in stdlib Python. It supports:
- `@activity` decorator that logs inputs and outputs to a JSON event log.
- A workflow function that sequences activities.
- A `run_or_replay(workflow, event_log)` function that replays completed activities without re-executing them.
The driver simulates a three-activity workflow, crashes halfway through, and shows (a) a naive retry re-executing everything versus (b) a replay running only the missing activity.
## Ship It
`outputs/skill-durable-execution-review.md` reviews a proposed long-running agent deployment for correct durable-execution shape: activities, determinism, checkpoint backend, human-input state, and HITL-on-resume policy.
## Exercises
1. Run `code/main.py`. Observe the difference in activity-execution count between naive retry and replay. Change the crash point and show the replay count changes accordingly.
2. Convert the toy engine to use `thread_id` explicitly. Simulate two concurrent sessions sharing the engine and confirm their event logs do not collide.
3. Take one activity in the toy engine. Introduce a non-determinism (a wall-clock timestamp inside a workflow decision). Demonstrate the divergence on replay. Explain how real engines handle this (side-effect registration, `Workflow.now()` APIs).
4. Read the LangChain "Runtime behind production deep agents" post. List every state that the runtime persists and name which failure mode each covers.
5. Design a checkpoint policy for a 6-hour autonomous coding task. Where do you checkpoint? What does resume-on-crash look like? What requires fresh HITL?
## Key Terms
| Term | What people say | What it actually means |
|---|---|---|
| Workflow | "Agent's script" | Deterministic orchestration code; replayable from event log |
| Activity | "A step" | Non-deterministic unit (LLM call, tool call); logged before and after |
| Event log | "The backing store" | Durable record of every state transition |
| Replay | "Resume" | Re-run workflow; completed activities return logged results without re-execution |
| Checkpoint | "Save point" | Persisted state keyed by thread_id; latest-wins on resume |
| thread_id | "Session key" | Identifier that scopes durable state |
| 35-minute degradation | "Reliability decay" | METR: success rate drops ~quadratically with horizon |
| Non-determinism | "Drift on replay" | Wall clock, random, LLM output; must be registered as side effect |
## Further Reading
- [Anthropic — Claude Code Agent SDK: agent loop](https://code.claude.com/docs/en/agent-sdk/agent-loop) — budget, turns, and resume semantics.
- [Microsoft — Agent Framework: human-in-the-loop and checkpointing](https://learn.microsoft.com/en-us/agent-framework/workflows/human-in-the-loop) — RequestInfoEvent shape.
- [LangChain — The Runtime Behind Production Deep Agents](https://www.langchain.com/conceptual-guides/runtime-behind-production-deep-agents) — concrete runtime requirements.
- [OpenAI Agents SDK + Temporal integration (Trigger.dev announcement)](https://trigger.dev) — activity shape for LLM calls.
- [Anthropic — Measuring agent autonomy in practice](https://www.anthropic.com/research/measuring-agent-autonomy) — the 35-minute degradation reference.
@@ -0,0 +1,41 @@
---
name: durable-execution-review
description: Review a proposed long-running agent deployment for correct durable-execution shape (activities, determinism, checkpoint backend, human-input state, HITL-on-resume).
version: 1.0.0
phase: 15
lesson: 12
tags: [durable-execution, workflows, checkpointing, temporal, langgraph, agents-sdk]
---
Given a proposed long-running agent deployment (Temporal + OpenAI Agents SDK, LangGraph with PostgreSQL checkpointer, Microsoft Agent Framework, Claude Code Routines, Cloudflare Durable Objects, or an in-house equivalent), audit the design against the durable-execution pattern.
Produce:
1. **Activity inventory.** List every activity (LLM call, tool call, HTTP request, file write). For each, confirm it is wrapped as an activity with retry policy, timeout, and idempotency key. Raw LLM calls outside the activity envelope are a reliability hole.
2. **Workflow determinism.** Identify every non-deterministic read inside the workflow code (wall clock, random, external state). Each must be registered as a side-effect activity so replay returns the same value. Hidden non-determinism is the most common cause of replay drift.
3. **Checkpoint backend.** Name the backend (PostgreSQL, SQLite, Redis, Durable Objects). Confirm it survives deploys. SQLite is dev-only. Redis requires AOF or snapshot config. Cloudflare Durable Objects are transparent but require a unique key discipline.
4. **Human-input state.** Confirm pauses for HITL are a first-class workflow state, not a polling loop. The workflow should block on an external signal (approval queue, webhook, `interrupt()` primitive) that resumes exactly when the approval arrives.
5. **HITL-on-resume policy.** For any resume after a crash, state whether fresh HITL is required before executing the next activity. Without this, durable execution plus an approval granted before the crash may re-fire an approved action when the context has changed. Critical for long horizons.
Hard rejects:
- Agent SDK usage where LLM calls are not wrapped as activities.
- Checkpoint backends that do not survive a deploy.
- Workflows that embed wall clock or random without activity wrapping.
- Human-input modeled as a polling loop rather than a signal.
- Long-horizon runs (above one hour) with no HITL-on-resume policy.
- Runs with no budget kill switch (Lesson 13) layered on top of durability.
Refusal rules:
- If the user proposes a durable workflow with no explicit idempotency on side-effect activities, refuse and require idempotency keys first. Retries will double-execute otherwise.
- If the user cannot show a replay test (run workflow, crash mid-run, replay, assert no double side effects), refuse and require that test before production.
- If the user proposes a 24-hour unattended run with no HITL checkpoint, refuse. The 35-minute degradation (Lesson 12 notes) makes this a reliability problem even if durability is correct.
Output format:
Return a design-review memo with:
- **Activity table** (activity, retry policy, timeout, idempotency key)
- **Determinism audit** (non-deterministic reads and how each is handled)
- **Checkpoint backend** (name, survives-deploy y/n, replay-test status)
- **HITL state shape** (first-class state / polling / missing)
- **HITL-on-resume policy** (explicit, with rationale)
- **Readiness** (production / staging / research-only)