Files
Mario Zechner 49683a3647 feat(durable): lifecycle conformance and task graph view (Packages 22, 23)
Package 22: runtime.now() and report() throw after the invocation ended; a
migration-only commit is published; a failed Harness.open closes without the
caller's context and rethrows the open error; ConversationWatch alias. Tests
for close joining non-cooperative code, cancelled close, queued operations at
the seal, progress calls and viewers on a paused Harness, the Chord guide, and
the spec's usage examples.

Package 23: Harness.taskGraph() and watchTaskGraph() mount every live task with
its owner edge, committed status, flags, and owned conversations, advanced from
commit publications. Example 24 prints the checkout's tree.
2026-10-01 09:53:21 +02:00
..
2026-09-30 21:12:28 +02:00

@earendil-works/pi-durable

Experimental. The API changes without notice between releases.

A durable agent harness. Conversations, model turns, tool calls, and your own state are committed to storage before anything is shown. If the process dies mid-turn, reopening the storage picks the work up where it stopped.

Built on @earendil-works/pi-ai for model access and @earendil-works/chord for document state.

Table of Contents

Installation

npm install @earendil-works/pi-durable @earendil-works/pi-ai @earendil-works/chord

Quick Start

import { BACKGROUND_CONTEXT } from "@earendil-works/chord/context";
import { createModels } from "@earendil-works/pi-ai/models";
import { openaiProvider } from "@earendil-works/pi-ai/providers/openai";
import { AssistantEntry, createRegistry, Harness, MemoryStorage } from "@earendil-works/pi-durable";

const context = BACKGROUND_CONTEXT;

const models = createModels();
models.setProvider(openaiProvider()); // reads OPENAI_API_KEY

const harness = await Harness.open(new MemoryStorage(), { models, registry: createRegistry() }, context);
const root = await harness.root(context, { agent: { model: { provider: "openai", modelId: "gpt-6-sol" } } });

const submission = await root.submit({ type: "input", content: "What is the capital of France?" }, context);
const settled = await submission.wait(context);
if (settled.status === "done" && settled.type === "input") {
	const answer = await root.commit((tx) => tx.entry(AssistantEntry, settled.answer), context);
	console.log(answer?.model?.[0]);
}
await harness.close(context);

What happened:

  • Harness.open() opens a Session over a storage backend. MemoryStorage keeps everything in memory.
  • root() returns the root conversation, creating it on first use with the given agent choices. A conversation is a transcript of immutable entries.
  • submit() durably admits your input and returns a Submission. A built-in generation task calls the model and appends the answer.
  • wait() resolves once the input is answered (done) or has failed (unanswered, with a reason).

Every async call takes a Chord Context, which carries cancellation. BACKGROUND_CONTEXT never cancels. Cancelling a wait only cancels that wait, never the work.

Concepts

  • Harness: one open storage plus the machinery that runs agents on it. All changes go through one line of atomic commits, and nothing is shown before its commit is stored.
  • Conversation: a transcript. root() creates the root conversation on first use; you can create more and fork them. A Conversation handle holds no state; compare handles by id.
  • Entry: one immutable transcript record, such as a user message (pi.user), a model response (pi.assistant), a tool result (pi.tool-result), a system prompt change (pi.system), a reset (pi.reset), or your own kind. The model sees the entries from the newest reset onward.
  • Commit: an atomic write. conversation.commit((tx) => ...) can append entries, edit documents, and create tasks together; either all of it is stored or none of it.
  • Document: typed JSON state stored next to the transcript and changed in commits. Built-in ones hold each conversation's agent choices (pi.agent), the running generation and tools (pi.live), queued submissions (pi.inbox), and spend (pi.usage).
  • Task: a durable state machine that saves a checkpoint at every step, so a restarted process continues from the last one. Every task has an owner: its conversation, or another task. The Harness runs answers as built-in tasks: pi.generation calls the model and owns the pi.tool tasks of its tool calls, waits for them, and hands the run to the next generation.
  • Submission: something you hand to a conversation, either user input or an entry to write, which you can wait for.
  • Turn and run: a turn is one model response and its tool calls; a run is the turns from an input to its final answer. A conversation is busy while a run is going.
  • Extension: a named bundle of tools, system prompt sections, hooks, wrappers, and tasks.
  • Registry: the extensions this process installed. It can change while the Harness runs; new work uses the new state.
  • Agent: what a conversation runs with: model, thinking level, selected extensions, tools, instructions, and working directory. Stored per conversation as names in pi.agent, resolved against the registry at each use.

One answered input, as entries and tasks:

submit(input) → pi.user
  pi.generation → pi.system (only if the prompt or tools changed), pi.assistant (tool calls)
    pi.tool × n → pi.tool-result × n   (owned by the generation, which waits for them)
  pi.generation → pi.assistant (answer) → submission done

Persist and Resume

Use SQLite or JSONL storage to keep conversations across restarts:

import { openNodeSqliteStorage } from "@earendil-works/pi-durable/storage/sqlite/node";

const harness = await Harness.open(await openNodeSqliteStorage("./session.sqlite"), { models, registry }, context);
const root = await harness.root(context); // the same root as last time
harness.resume(); // continue any run the last process left unfinished

Work interrupted by a crash or close stays pending. resume() starts the task scheduler; submitting or waiting starts it too. A retried submission with the same requestId returns the existing submission instead of submitting twice:

const submission = await root.submit({ type: "input", content: "Hello", requestId: "greeting-1" }, context);
// After a restart: the same request ID finds the same submission.
const again = await root.submit({ type: "input", content: "Hello", requestId: "greeting-1" }, context);
// again.id === submission.id

harness.submission(id) reacquires a submission by ID, for example to wait for it after a restart.

Extensions

Code the Harness runs, other than its built-in tasks, comes in named extensions installed in a registry your process owns:

import { createRegistry, defineExtension, defineTool, hook, section, ToolTask } from "@earendil-works/pi-durable";
import { CodingTools } from "@earendil-works/pi-durable/tools";

const Coding = defineExtension({
	name: "coding",
	sections: [section("preamble", () => "You are a concise coding assistant.", { tag: false })],
	hooks: [hook(ToolTask, { beforeTool: (call) => (isDangerous(call) ? { block: "Needs approval" } : undefined) })],
});

const registry = createRegistry();
registry.install(CodingTools);
registry.install(Coding);

An extension may bring tools, sections, hooks, wraps (decorators of a tool or section by name), and tasks. By default every conversation selects every installed extension, in install order. Nothing in the registry is stored; conversations store extension names.

Tools

@earendil-works/pi-durable/tools provides read, write, edit, and bash, and the CodingTools extension with all four. They touch files and processes only through the call's environment (see Environment). Reading images is not supported yet.

Define your own tool with a TypeBox schema. defineTool() types args from parameters, which the Harness validates before execute(). api.output() streams running output, which becomes the result when execute() returns no content:

import { Type } from "@earendil-works/pi-ai";

const count = defineTool({
	name: "count",
	description: "Count from 1 to n",
	parameters: Type.Object({ n: Type.Number() }),
	execute: async (args, api) => {
		for (let i = 1; i <= args.n; i++) api.output(`${i}\n`);
		return {};
	},
});
registry.install(defineExtension({ name: "count", tools: [count] }));

Each call runs as its own durable task. Its intent is committed before execute() runs. If the process dies mid-call, the tool reruns on reopen only when it is declared replay: "safe"; otherwise the model gets an interrupted error result with the output committed so far. Throwing from execute() gives the model an error result. A result can also return usage, which is added to the conversation's usage. It can also return control: { terminate: true }: when every result of the round asks for it, the run ends without another model request.

A later extension's tool with the same name replaces an earlier one where both are selected, and wrapTool() decorates whichever tool won:

const Venv = defineExtension({ name: "venv", tools: [createBashTool({ commandPrefix: "source .venv/bin/activate" })] });
const Timing = defineExtension({
	name: "timing",
	wraps: [wrapTool(createBashTool(), (bash) => ({ ...bash, execute: (args, api, ctx) => timed(() => bash.execute(args, api, ctx)) }))],
});

System Prompt

The system prompt is built from the selected extensions' sections, rendered in order before each request. A section sees the resolved agent, the environment built for the request, and committed documents:

section("cwd", (input) => input.env?.cwd); // rendered as <cwd>\n...\n</cwd>; undefined omits it

A conversation's instructions render last, as the section instructions. Sections and tool changes are stored as positional system entries in the transcript. Only what changed is sent again, which keeps provider prompt caches warm. A section that returns something different every time, such as the current time, defeats that.

Per-Conversation Agent

Each conversation stores what it runs with in its pi.agent document. configure() changes it in one commit; unset fields follow the host:

await root.configure(
	{
		model: { provider: "openai", modelId: "gpt-6-sol" },
		thinkingLevel: "high",
		extensions: { remove: [Coding] }, // edits the host default; an array selects exactly these, in order
		tools: [readTool, bashTool], // an array offers exactly these; { remove: [...] } drops some
		instructions: "Only read; never edit files.",
		cwd: "/work/repo",
	},
	context,
);
await root.configure({ tools: null }, context); // null clears a field back to the host default
const agent = await root.agent(context); // resolved: model, extensions, tools, sections, cwd

Extensions and tools are passed as objects and stored by name, so a stored name outlives its code: after an extension is uninstalled, conversations that select it just stop getting it until it is installed again. createConversation(), fork(), and root() take the same change as agent. A task-owned conversation, such as a subagent's, starts as a copy of its owner's conversation's agent. A fork starts with the agent its parent had at the fork entry. The model, prompt, and offered tools of a request are fixed when it is prepared; a change applies from the next request. Tool calls and hooks use the agent as their task phase resolves it, and the environment is built from the current cwd at each use, so a cwd or extension change can reach calls the model already made.

Settings

Run policy shared by every conversation is passed as settings. It is read at every use and never stored, so getters make it live, for example backed by a settings file:

const harness = await Harness.open(storage, {
	models,
	registry,
	settings: {
		extensions: [CodingTools, Coding], // default selection; absent: every installed extension
		stream: { timeoutMs: 120_000 },
		retry: { maxRetries: 3 },
		compaction: { reserveTokens: 16384 },
		toolExecution: "parallel",
		get followUpMode() {
			return userSettings.followUpMode;
		},
	},
}, context);

Environment

env builds the execution environment for each tool call, section rendering, and runtime.env(). It receives the conversation's ID, its agent cwd, and committed reads, so one function serves a directory per conversation or a container per conversation:

import { NodeExecutionEnv } from "@earendil-works/pi-durable/env/node";

const harness = await Harness.open(storage, {
	models,
	registry,
	env: ({ cwd }) => new NodeExecutionEnv({ cwd: cwd ?? process.cwd() }),
}, context);

A throw from env becomes the call's error result. Without an environment, the built-in tools fail with an error result. A fresh environment object per call is fine: edit and write serialize changes to one file by the environment's id and path. A custom ExecutionEnv sets id so that equal ids see the same files at the same paths, for example one id per container.

Reload

Installing an extension with an installed name replaces it in place, in one step:

registry.install(await loadCodingExtension()); // same name "coding": replaces the installed one

registry.uninstall(extension) removes the installed extension with that name, whichever object it is.

Work that already started keeps the code it took: a running tool call finishes under its old implementation, and each task phase resolves hooks and the agent once, from the registry at the phase's start. The next phase, request, or call uses the new code. After a restart, install the same extensions again; pending tasks of an extension's tasks resume once it is installed.

Watching a Conversation

Everything a UI needs is committed state. viewState() returns the conversation's structural view as a read-only Chord state, updated after every commit that touches it:

const view = await root.viewState(context);
view.subscribe((value) => {
	// value.entries: the active transcript
	// value.docs["pi.live"]: the running generation (streamed partial, retry, deferred) and tool calls (output, details)
	// value.docs["pi.inbox"], value.docs["pi.usage"], value.docs["pi.agent"]
	render(value);
});
// later: view.dispose();

watch() delivers the same view with the exact Chord operations of each commit, one callback at a time:

const watch = await root.watch(context);
render(watch.value); // the state at attachment
watch.start(async (value, ops) => {
	await send(ops); // for example to a remote client that applies them
});
// later: await watch.stop();

A slow watch keeps at most 100 undelivered frames. After that, the pending frames are replaced by one frame holding the whole newest view. A client that joins late or reconnects starts from the current view; nothing is replayed.

Partial answers and tool output are committed at most every 100 ms, so a crash loses at most that window.

Busy Conversations

A conversation is busy while a run is working on an input. Submitting to a busy conversation queues the submission in the conversation's inbox, docs["pi.inbox"] in the view:

await root.submit({ type: "input", content: "Also run the tests" }, context); // follow-up (default)
await root.submit({ type: "input", content: "Use pnpm, not npm", whenBusy: "steer" }, context);
await root.submit({ type: "input", content: "Only if idle", whenBusy: "reject" }, context); // throws ConversationBusy
await root.submit({ type: "write", entry: { kind: "app.note", data: "user opened a file" } }, context);
  • Steers are placed after the current tool round and join the running work.
  • Follow-ups are placed when the run answers, and start the next run.
  • Writes append an entry without asking the model anything.
  • await submission.abort(context) withdraws a queued submission.
  • The settings steeringMode: "all" and followUpMode: "all" place every queued item at once instead of one per turn.

If a run fails, queued items stay in the inbox until the next submission places them, oldest first.

Reset and Handoff

reset() starts a new context. The model no longer sees older entries, but they stay in storage:

await root.reset(undefined, context);                                  // start from nothing
await root.reset("We were fixing the flaky login test. Continue.", context); // start from a handoff note

While busy, the reset is queued like a write. When it is placed during a tool round, the current run ends. A tool can request the same with control: { handoff: "..." }.

Compaction

Compaction shrinks what the model sees: it summarizes older entries and appends a pi.compaction entry that holds the summary and heads the first entry it keeps. Older entries stay in storage.

const id = await root.compact("Keep the failing test names", context); // manual, with optional instructions
const { outcome } = (await harness.waitForTask(id, context)).state;
if (outcome.status === "completed" && outcome.result.submissionId !== undefined) {
	const placed = await (await harness.submission(outcome.result.submissionId, context))!.wait(context);
	console.log(placed.status); // "done", or "unanswered" with reason "stale"
}

The conversation keeps working while the summary is made. The summary is placed at once when the conversation is idle, otherwise at the next turn boundary. Esc (abort()) cancels a manual compaction.

Generation also compacts on its own, controlled by the settings:

settings: {
	compaction: {
		enabled: true, // automatic compaction; manual compact() always works
		reserveTokens: 16384, // above contextWindow - reserveTokens, the next request waits for a compaction
		keepRecentTokens: 20000, // roughly how much recent context stays verbatim
		backgroundTokens: 32768, // this far below that, a compaction starts in the background; 0 disables it
	},
}

When a provider rejects a request because the context is too long, generation compacts and retries once. A summary that would cut before the start of the current context settles as stale when it is placed, so when several are in flight, the furthest cut stays in effect. Summarization spend counts in pi.usage. A beforeCompact hook on CompactionTask can decline or supply its own summary.

Running compactions are listed in docs["pi.live"].compactions with their reason, attempt, and retry backoff. The agent events add compaction_start and compaction_end, and a compactions field in the snapshot.

Agent Events (Experimental)

For consumers that want coding-agent style events (message_start, message_update, tool_execution_start, ...) instead of structural state:

import { watchEvents } from "@earendil-works/pi-durable";

const stream = await watchEvents(harness, root.id, context);
initialize(stream.snapshot); // entries, run, in-flight generation, tools, compactions, inbox, agent, usage
stream.start(async (events) => {
	for (const event of events) console.log(JSON.stringify(event));
});

Events are derived from commits, one batch per commit, and apply on top of the snapshot. Message and tool updates carry deltas: text and thinking appends, appended tool-call argument text, and output trims and appends. When a consumer falls more than 100 batches behind, it receives a fresh snapshot event instead. See test/examples/19-json.ts for the full stream of one run.

Hooks

Hooks let extensions observe or adjust the built-in tasks, in the conversations that select them:

import { GenerationTask, hook, ToolTask } from "@earendil-works/pi-durable";

const Guard = defineExtension({
	name: "guard",
	hooks: [
		hook(ToolTask, { beforeTool: (call) => (call.name === "bash" ? { block: "bash is disabled here" } : undefined) }),
		hook(GenerationTask, { onYield: (answer) => (needsMoreWork(answer) ? { continue: "Keep going." } : undefined) }),
	],
});
  • Generation: beforeRequest (replace the messages of one request), afterResponse, onYield (continue the run with another user message), and afterTools (runs once a round's tools are done).
  • Tools: beforeTool (block or rewrite arguments) and afterTool (replace the result).

To limit a hook to some conversations, select its extension only there, for example with configure({ extensions: { add: [Guard] } }).

More Conversations and Forks

const other = await harness.createConversation({ ownership: { kind: "ownerless" } }, context);
const fork = await root.fork(entryId, { ownership: { kind: "ownerless" } }, context);

A fork sees its parent's entries up to entryId and continues independently. It keeps the parent's agent as of that entry. Both take agent and init, applied in the creating commit.

Abort and Subagents

await root.abort(context) stops a conversation: queued inputs are withdrawn (queued writes stay), every task of its current work is aborted, and the call resolves once the conversation is idle.

A conversation can be owned by a task. A subagent tool creates its child inside api.commit() with ownership: { kind: "task", taskId: api.taskId }, then drives it through api.conversation(id):

const Subagent: Extension = defineExtension({
	name: "subagent",
	tools: [
		defineTool({
			name: "subagent",
			description: "Delegate a self-contained task to a subagent and get its answer back.",
			parameters: Type.Object({ task: Type.String() }),
			replay: "safe", // a rerun after a crash finds the same child and submission
			execute: async (args, api, context) => {
				const child = await api.commit(async (tx) => {
					// The ownership index remembers the child, so a rerun reuses it.
					const existing = (await tx.scanConversations({ ownerTaskId: api.taskId }, 1)).items[0];
					if (existing !== undefined) return existing.id;
					// Starts as a copy of this conversation's agent: model, extensions, tools, cwd.
					const created = await tx.createConversation({ ownership: { kind: "task", taskId: api.taskId } });
					// A cheaper model, and no subagents of its own.
					await configure(tx, created.id, { model: haiku, extensions: { remove: [Subagent] } });
					return created.id;
				}, context);
				await api.details({ conversationId: child }, context); // lets a UI attach to the child
				const request = { type: "input", content: args.task, requestId: `subagent:${api.taskId}` } as const;
				const settled = await (await (await api.conversation(child, context))!.submit(request, context)).wait(context);
				return { content: [{ type: "text", text: settled.status }] };
			},
		}),
	],
});

Owned work belongs to its owner:

  • Aborting the call aborts the child. So does the call failing: execute() throwing, or a crash that interrupts a call that is not replay-safe.
  • The parent is idle only once the child is.
  • A task created with { background: true } is a boundary: work it owns survives the parent's abort and does not keep the parent busy. root.abort(context, { background: true }) aborts it too.

The examples show both patterns as product code:

  • 22-subagent-foreground.ts: the tool above, returning the child's answer. The UI finds the child through the call's details and prints the child's events indented under the call.
  • 23-subagent-background.ts: persistent subagents behind one subagent tool that spawns, messages (steer or follow-up), waits for, stops, and lists them. Each child is owned by a background anchor task, so the parent's Esc and idle waits never reach it. Each message is delivered by a background reporter task that posts the answer back to the parent as a follow-up input once it arrives; request IDs keep a restart from sending a message or a report twice.

Child Tasks

A task can own child tasks, created with ownership: { kind: "task", taskId }, and wait for them by committing a waiting state:

pay: async (task, runtime, context) => {
	await runtime.commit(async (tx) => {
		const payments = [];
		for (const card of task.input.cards) {
			payments.push(await tx.createTask(Payment, { card }, { ownership: { kind: "task", taskId: task.id } }));
		}
		// Resume in `decide` once every payment is done; the first failure aborts the rest.
		return { status: "waiting", checkpoint: { phase: "decide", payments }, on: payments, policy: "failFast" };
	}, context);
},
decide: async (task, runtime, context) => {
	const outcomes = await runtime.outcomes(task.state.checkpoint.payments, context);
	// ...commit the checkout's own outcome
},
  • Waiting: the task runs no code while it waits. With allSettled it resumes once every task in on is done; with failFast the first failed child also aborts the others. on may name other tasks too, with allSettled.
  • Finishing: a task that finishes while work it owns is still running is completing: its outcome is decided, but it becomes terminal, and waitForTask() returns, only once that work is done. A failed or aborted outcome aborts that work first.
  • Aborting: abort runs bottom-up. Aborting a task aborts the work it owns first, and its own abort handler starts only once that work is done, so each task undoes its own effects.

24-child-tasks.ts runs a checkout with four payments: a declined card, a cancelled checkout, and a restart while the payments run.

Task Graph

harness.taskGraph(context) shows every live task of the Session as one Chord state, for a task panel or debugging. Each node has its owner edge (owner task, or none for a task its conversation owns), its status, whether it is background or abort-marked, and the conversations it owns. harness.watchTaskGraph(context) delivers the same value as a watch, like a conversation's watch().

const graph = await harness.taskGraph(context);
graph.subscribe((value) => {
	for (const node of Object.values(value.tasks)) {
		const status = node.state.status === "waiting" ? `waiting on ${node.state.on.join(", ")}` : node.state.status;
		console.log(`${node.id} ${node.kind} ${status}`, node.owner ?? `conversation ${node.conversationId}`);
	}
});

A task appears with the commit that creates it and leaves with the commit that makes it terminal. Statuses are the committed ones: pending, running, waiting (with on and policy), and completing (with the held outcome's status). After a restart, tasks that were running show as pending until they run again. Whether a pending task is blocked by a missing definition is not part of the graph; harness.inspect() reports that. The graph lists live tasks only: once a subagent's owner task is terminal, a later task in its conversation is a top-level node, and the conversation's ConversationRecord.owner (also in its view's conversation) links it to its parent. 24-child-tasks.ts prints the checkout's tree while its payments run.

Your Own State

Documents are typed JSON objects committed together with entries. Define one, and edit it in a commit:

import { defineDoc } from "@earendil-works/pi-durable";

const Todos = defineDoc<{ items: string[] }>({
	kind: "app.todos",
	version: 1,
	scope: "conversation",
	history: "latest", // or "rewindable" to read old values with snapshotAsOf()
	fork: "initial", // what a fork starts with: "initial", "current", or "asOf"
	initial: () => ({ items: [] }),
});

await root.commit(async (tx) => {
	(await tx.doc(Todos, root.id)).items.push("write docs");
}, context);
console.log(await harness.snapshot(Todos, root.id, context));

harness.watchDoc() and harness.documentState() observe one document like the view above. HarnessOptions.conversationCreated(tx, conversation) runs in every commit that creates or forks a conversation, including a tool's raw tx.createConversation(), so every conversation gets your documents; init in createConversation(), fork(), and root() writes per-call data in the same commit. An extension's tools, sections, and hooks read their own documents through api or input.read, and treat an absent one as its default (11-extension-state.ts).

Usage and Cost

Each conversation keeps token and cost totals in docs["pi.usage"]: per provider/model for model responses, and per tool name for tool results that report usage. Failed and aborted attempts count too. For the whole Session:

const usage = await harness.usage(context); // { models: { "openai/gpt-6-sol": Usage }, tools: {...} }

Storage

Backend Import Notes
Memory MemoryStorage from the package root Nothing is persisted.
SQLite openNodeSqliteStorage(file) from @earendil-works/pi-durable/storage/sqlite/node One database file. WAL mode with synchronous = NORMAL: commits survive process crashes; the newest may be lost on power or host failure.
JSONL openNodeJsonlStorage(directory, context) from @earendil-works/pi-durable/storage/jsonl/node Append-only files in one directory. Pass { fsync: true } to flush before each commit marker.

One process owns a storage at a time; there is no cross-process locking. The portable SQLite and JSONL cores (/storage/sqlite, /storage/jsonl) run without Node APIs, for example on Bun or in Cloudflare Durable Objects, given an asynchronous SqliteDatabase facade or a FileSystem from @earendil-works/pi-durable/env.

SQLite adapters implement promise-based exec, run, get, all, transaction, and close. run, get, and all take SQL text plus positional bindings; adapters may cache prepared statements by SQL text. A transaction callback receives a transaction handle; all work in the transaction must use it, and the handle expires when the callback settles. Adapters must queue unrelated operations and other transactions until the transaction finishes, so calling database itself inside the callback never settles:

await database.transaction(async (transaction) => {
	await transaction.exec("CREATE TABLE example (value TEXT)");
	await transaction.run("INSERT INTO example (value) VALUES (?)", "stored atomically");
});

Custom backends can run the shared conformance suite with any Vitest- or Jest-compatible runner:

import { registerStorageConformance } from "@earendil-works/pi-durable/testing";
import { describe, expect, it } from "vitest";

registerStorageConformance({ describe, expect, it }, "My Storage", async (use) => {
	const storage = await openMyStorage();
	try {
		await use(storage);
	} finally {
		await closeMyStorage(storage);
	}
});

The package root loads TypeBox, because the tool task validates arguments with pi-ai's validateToolArguments(). That costs about 23 MB of peak RSS unbundled, about 4 MB in a tree-shaken bundle.

Examples

Runnable examples live in test/examples. Run one from this package directory with:

node --conditions=source --experimental-strip-types test/examples/14-chat.ts
Example Shows
14-chat One question and answer
16-real-model Streaming an answer from OpenAI
17-coding-tools A tool-using turn on JSONL storage
18-print Print mode: submit a prompt, print the answer
19-json JSON mode: agent events or raw view operations, on SQLite, JSONL, or memory
20-inbox Steers, follow-ups, writes, and withdrawal while busy
21-late-join Attaching a view and an event stream mid-run
22-subagent-foreground A replay-safe subagent tool whose child the call owns, with the child's events under the call
23-subagent-background Persistent subagents: spawn, steer, stop, list, answers reported back, restart-safe
24-child-tasks A checkout that owns and waits for four payments: failFast, abort, restart
25-compaction A long chat compacted in the background, manually, and after a context overflow
26-coding-agent CodingTools, live settings from a settings object, an environment that follows the conversation's directory
27-plan-mode A read-only plan mode as an extension with its own document, switched with configure()
28-reviewer A reviewer conversation with its own model, extensions, tools, directory, and review loop
29-sandbox-per-conversation An environment per conversation, looked up from an app document
30-tool-override A same-name bash for some conversations, and a wrapper that times whichever bash won
31-reload-and-restart Reloading an extension mid-call, and stored choices surviving a restart
00–13 The layers underneath: sessions, documents, forks, watches, the Harness, agent configuration, reload, extension state, tasks, recovery

Examples that call OpenAI need OPENAI_API_KEY; most use the faux provider otherwise.

Design Documents

Benchmarks: npm run bench:storage, npm run bench:storage:memory, and npm run bench:tool-output.

License

MIT