Anthropic, OpenAI Codex, OpenRouter, and Radius now use one loopback callback server in pi-ai. MCP sign-in renders the same browser page through a new renderPage option on the pi-mcp callback server.
Built-in extensions (mcp, llama.cpp, codemode, tool-search) are now extension resources like files: the package manager resolves them as builtin:<name>, the extensions setting disables them with -builtin:<name> (per project too), pi config lists them, and --no-extensions turns them off unless loaded with -e builtin:<name>. They load after project trust is resolved and after file and package extensions.
Built-in extensions and tools are named builtin:<name> in errors, diagnostics, RPC source info, and bug reports.
Refreshes hold a per-server lock file from reading the tokens to saving new ones, and use tokens another process already refreshed. Rotating refresh tokens (Cloudflare) were lost when two pi processes refreshed at once, leaving the server needing sign-in. Closing a connection waits for a running refresh so shutdown does not drop rotated tokens.
pi.live turn control and generation presentation, built-in entry kinds,
pi.generation checkpoint and phases, typed entry tokens, built-in tasks in
the registry, TaskRuntime reads, stream options and retry policy config,
Harness hook for scheduler-written outcomes, and createConversation without
input.
Adds fullscreenWheelScrollLines ("auto" or 1-100). Auto keeps one line
per event in local macOS terminals and scales with wheel velocity
elsewhere, capped at six lines per event.
closes#9758
This adds codemode and MCP support to pi. Codemode (packages/codemode) runs
model-written JavaScript in a QuickJS wasm VM inside a worker. From there the
script calls pi's tools as async functions, uses a per-session store, and can
read the model catalog and run classifiers. The codemode tool in coding-agent
is compatible with Codex's exec tool, and codemode.mode works like Codex's
tool modes (on or only). Large results are written to a JSON temp file. MCP
(packages/mcp) is a standalone client for stdio and streamable HTTP servers,
with OAuth.
The core only gets general mechanisms; codemode, tool_search and MCP are
built-in extensions that use them. Other extensions can replace these
built-ins. Tools declare an exposure: direct, model-only, codemode, deferred
or hidden. prepareLoadout() lets an orchestrating tool change which tool
declarations the model sees. ctx.executeTool() runs nested calls through
validation and the tool_call/tool_result hooks. Those calls emit events with
parentToolCallId and are recorded as bounded nestedCalls on the parent
result, which compaction and HTML export use. Tools can also return
structuredContent with an outputSchema, and can report errors with isError
instead of throwing. MCP servers come from mcp.json (global, or per project
after the project is trusted) or from pi.registerMcpServer(). They are
managed with /mcp and pi mcp list|login|logout, and their runtimes load on
first use.
Closes#10040
Theme previews re-render the whole transcript, and streaming re-renders
every frame, so per-line and per-session work dominated in long sessions.
- visibleWidth: fast path for ANSI-styled ASCII, allocation-free escape scanning
- Box: compare unpadded child lines and skip a second width measurement
- Markdown: reuse parsed tokens across invalidation
- footer: cache usage totals and context usage per session state
- bash results: cache the collapsed preview including the expand hint
- sanitizeBinaryOutput: single regex replace
Together no longer lists Kimi K2.6, which broke type checks against the hydrated catalog. Tests use Kimi K3 and the Together default model now points at it.
Adds defineTask, registry-resolved phase handlers, and a scheduler that
decides every task transition in one callback on the Session line:
reservation with migration, dependencies, precedence rules in a step
before each phase, handover to replaced definitions, direct-task abort
(mark, signal, join, fresh abort invocation), orphaning of blocked tasks,
open reconciliation, and close that joins invocations without writing
outcomes. runtime.commit() callbacks return the typed next state; Tx no
longer exposes setTask. Adds Harness.resume, getTask, waitForTask,
abortTask, and task-aware waitForIdle; RegistryReader.subscribe wakes
the scheduler.
This adds experimental support for virtual models. A virtual model is a catalog entry that an extension registers with pi.registerVirtualModel(). It does not talk to a provider itself; for every request it picks a physical model and thinking level. The motivation is to make routing policies pluggable: plan on a strong model and implement on a cheap one, pick a model with a classifier, or move to a larger window when the context grows. Today this needs model switching by hand or hacks in prepareRequest.
The selection stays virtual: model_change records it, resume restores it, and each response records the physical model and thinking level that produced it. Router state can be stored on the session branch, so it follows /tree and forks and survives compaction. Context limits, compaction and image sizing use the physical model a request was routed to. examples/extensions/jev-router.ts shows a full router that uses the Jev classifier to pick Sol or Terra for planning and hands off to Luna after the first edit. docs/virtual-models.md describes the API.
The agent runs every tool call in the final assistant message, including
calls whose output_item.done never arrived. Those can carry cut-off or
mixed-up arguments. With llama.cpp, which omits output_index, two parallel
calls became three (echo a, echo a, echo b), two of them with the same id.
processResponsesStream now fails the stream when a completed response
still has an unfinished tool call.
fixes#9974
Read Finder file URLs before icon image data on macOS. Quote paths in
bash mode and reject terminal control characters before rendering.
Fixes#9999
Co-authored-by: 1dustycy <onedustycy@gmail.com>
Implements Pico5 Package 13: Harness.open over a Session with an application-owned
registry (tools, tool wrappers, hooks, tasks, system prompt sections; batched
publication, stable keyed order), lazy root creation with init, atomic
create/fork with init, conversation-bound commits, fork-aware entries, model
context derivation, and the built-in ConversationConfig document.
Fixes cached documents skipping migration when accessed with another definition
version. Rewrites the Pico5 spec and handoff to the registry/system-prompt
design and removes the superseded live-registries proposal.
Storage reports replayed deltas after the selected base; Session tracks the
count across adopted writes so checkpoint predicates can bound replay.
Also adds internal reserved-root creation and conversation-bound commit
defaults for the upcoming Harness.
Mistral models with effort levels now use reasoning_effort based on
their thinking level map from models.dev instead of a hardcoded id list.
GLM 5.3 no longer gets prompt_mode, GLM 5.2 can use max, and off sends
the model's off value when it has one.
fixes#9678
Add the llama-cpp-classify classifier API. Each question becomes one chat prompt with single-token answer labels; llama-server's pre-sampling next-token log-probabilities give the softmax over those labels. Missing labels are retried with deeper readouts and then reported as errors.
models.dev dropped Kimi K2.6 from OpenCode Go, which broke type checks
against the regenerated catalog. Use OpenCode Go Kimi K3 for
provider-level tests and OpenCode Zen Kimi K2.6 for thinking format tests.
* feat(coding-agent,tui): add system theme derived from terminal colors
The new default system theme builds pi's colors from the terminal's reported foreground, background, and ANSI palette. Tokens take their hue and relative saturation from a palette slot; lightness is placed by OKLab lightness difference and WCAG 2 contrast minimums. Terminals that report only a background get built-in hues, and terminals that report nothing get ANSI indices with faint secondary text.
The TUI queries OSC 10, 11, and 4 in one burst ended by DA1, so replies complete in one round trip, and reports replies that arrive after the timeout. Startup never waits: the system theme starts in grayscale and updates when colors arrive. Light/dark detection now uses the reported colors instead of the ?996 query; mode 2031 notifications trigger a new query.
* fix(coding-agent): wait for terminal colors before building the startup header
The header and startup notices bake theme colors into their text. When the color replies arrived after they were built, as in Windows Terminal, they stayed grayscale. Startup now waits for the first color query, which ends at the DA1 reply or after 100 ms.
* fix(coding-agent): rebuild header and notices on theme changes
The startup header, loaded resources, and chat notices baked theme colors into their text, so they kept the old colors after a theme switch or when the system theme received the terminal's colors. They now use ThemedText, which rebuilds its text from a function whenever the UI is invalidated.
* fix(tui): keep focus on components that forward mouse events to children
A component that forwards a mouse event to a child it hosts, such as SettingsList with an open submenu, left keyboard focus on the child. When the host removed the child, focus stayed on a detached component and all keys were lost. The forwarding component now keeps focus, like delegating containers do.
* feat(coding-agent): describe the system theme in the theme selector
* feat(coding-agent): list the system theme above automatic in the theme selector
* refactor(coding-agent,tui): remove leftovers of the light/dark startup detection
Theme detection only reports dark or light now; the source and confidence fields only served saving the detected theme. The first-time setup no longer shows the detected appearance. Removes the TUI's unused queryTerminalColorScheme(), queryTerminalBackgroundColor(), queryTerminalForegroundColor(), and parseOsc11BackgroundColor(); queryTerminalColors() covers them.
* feat(coding-agent): match the reviewed system and light/dark theme design
The system theme now follows the reviewed design: color families with OKHSL saturation curves, palette slots per family, contrast rules per token and panel, relaxation for mid-gray backgrounds, and the terminal foreground for body text. Each contrast level is a target-lightness curve fitted to the reviewed output. The built-in dark and light themes use the reviewed Pi dark and light colors.
* feat(coding-agent): show the pi logo in the startup header
A two-line pixel logo replaces the app name. Its first line carries the version, its second line the first line of key hints.
* feat(tui,coding-agent): support OKHSL colors in themes
parseColor() accepts okhsl(H S L), and okhslColor() and colorToOkhsl() construct and read OKHSL colors. OKHSL saturation is relative to the sRGB gamut, so every value is in gamut. Theme files accept okhsl() values; HTML exports convert them to hex. The OKHSL code moves from coding-agent into pi-tui, where colors.ts now shares its Oklab conversion instead of keeping a second copy.
* refactor(coding-agent): write the built-in themes in OKHSL
The built-in dark and light themes use okhsl() values, with variables for colors that several roles share, so they read as color families and serve as editable templates. Near-identical values collapsed into one; no color changes visibly.
* refactor(coding-agent,tui): simplify terminal color and theme code
Share the terminal color query between startup and the theme controller, merge the query's resolve/late-reply state into one callback, keep a single global for terminal colors, use oklab.ts arrays directly in colors.ts, and drop leftover helpers and exports.
* fix(coding-agent): use one terminal light/dark decision everywhere
Theme pairs, the system theme, and Theme.appearance now share getTerminalTheme(): the reported background decides, then the terminal's light/dark report, then COLORFGBG, then dark.
COLORFGBG is classified by palette index like Vim (0-6 and 8 dark) and only reads the last field, so rxvt's "default" no longer picks up the foreground index.
Simplify the tests added on this branch.
* fix(ai): use an OpenRouter model that still exists in beta test
OpenRouter removed anthropic/claude-3-haiku, which broke type checking in CI after model regeneration.
GPT-6 models return service_tier "fast" for Fast mode requests (the new
name for priority processing), which fell through to the 1x default.
closes#10034
* fix(ai): upgrade openai SDK to 7.19.0
Adds the "fast" service tier to the SDK types, needed to price
GPT-6 Fast mode requests correctly. Drops the local
prompt_cache_options type, which the SDK now defines.
* fix(coding-agent): check Cloudflare compat request via fetch instead of SDK mock
pi-ai now resolves its own nested openai 7.x while evals keeps 6.x at the
root, so vi.mock("openai") in coding-agent no longer reached pi-ai's client.
Capture the outgoing request with a fake fetch and assert the URL and headers.
* fix(ai): drop any casts for stream_options and max_tokens in openai-completions
stream_options is typed by the SDK, so assign it directly. max_tokens is
deprecated by OpenAI but still needed for OpenAI-compatible providers
that reject max_completion_tokens; use a narrow cast instead of any to
avoid the deprecation warning.
A new session was only written to disk once the first assistant message
existed. If pi exited during the first turn (Ctrl+C twice while streaming,
or the process was killed), the user's prompt was lost and the session
never appeared in /resume.
The file is now created as soon as the session has a user or assistant
message. A session with only setup entries (model, thinking level) still
stays in memory, so starting pi and quitting without chatting leaves no
file. Forking uses the same rule through a shared helper, so the two paths
cannot drift apart and write the header twice (#1672).
closes#10000
GLM models on Mistral send empty content deltas at the start of a
response, alongside tool call fragments, and sometimes mid-thinking.
Each one opened an empty text block, and a mid-thinking one split
thinking into two blocks, which Mistral rejects on replay with
"Expected at most one leading ThinkChunk".
closes#9674
Model-level samplingParams were only merged by streamSimple(). Direct
stream()/complete() calls, such as extensions using
modelRegistry.complete(), dropped them. The OpenAI-compatible adapters
now merge model defaults with request keys, and request keys take
precedence.
closes#9506
Successful RPC prompt, steer, and follow_up responses include
data.disposition ("handled", "queued", or "started" for prompt), so clients
know whether to wait for agent_settled.
fixes#9098
Closing the last overlay hid the terminal cursor. If an extension closed
an overlay during session shutdown, this ran after stop() had restored
the cursor, leaving the shell with an invisible cursor.
closes#10026
Replace the tsgo native preview with typescript@7.0.2 and port the check scripts to the TS 7 API. Target ES2024, drop useDefineForClassFields: false and unused decorator options, and enable verbatimModuleSyntax. Replace tsx with node plus the source resolver hook, passed to --import as a file URL.
closes#9965
Strict tool schemas make models send null for omitted optional fields.
The read renderer only treated undefined as missing, so full-file reads
rendered as "path:1".
fixes#9996
Temporary git extensions (-e git:...@ref) were cached in a folder keyed
only by host and repo path. After changing the pinned ref, pi reused the
old checkout without refreshing it and kept loading the first downloaded
commit. Include the ref in the temporary folder hash so each ref gets
its own checkout. Unpinned sources keep their existing folder.
closes#9982
Replace the alternate tracker implementations with the optimized immutable overlay, validate draft placements as strict JSON, and document the ownership contract. Consolidate tracker tests and benchmarks around the canonical implementation.