Move the codemode script reference, including the models API, into
docs/codemode.md and keep only one line per global in the tool
description. Declared tools get a one-line note on how scripts call them,
and the codemode system prompt guidance and MCP server section are shorter.
With the default tools a GPT-5.6 request drops from about 5,300 to 3,300
tokens.
Errors now point the way back: unknown tools and models members name close
matches, models.classify() and models.generateImages() validate their
arguments, unknown or mistyped models point to getAvailableOfType(),
store() overflow explains the store, and unshown generated images get a note.
Scripts run image models with the session's credentials, like
models.classify(). Results are base64 image blocks that image() attaches
to the codemode result; usage counts toward the session cost.
The top-level /login menu offers Sign in with Radius with its status,
with an animated shimmer on the Radius row and a short Radius intro.
After a Radius sign-in it offers to configure the Radius MCP server in
the global mcp.json with auth.provider and reloads. Cancelling a login
returns to the menu it was started from. OAuth sign-ins without a
subscription are labeled as accounts.
Co-authored-by: niloc <5925270+colindaymond@users.noreply.github.com>
The session restored its tool loadout before MCP servers reconnected, so deferred tools that tool_search had loaded were dropped. Restored tools that are not registered yet now stay pending and are activated when they register, until a setActiveTools() call deactivates a tool or the next prompt starts.
Runs the pinned @modelcontextprotocol/conformance client scenarios for 2025-03-26, 2025-06-18, and 2025-11-25 against pi's MCP connection and OAuth sign-in, and fails on regressions against a committed check-level baseline.
A step-up sign-in requested only the scopes of the insufficient_scope challenge, so the new token lost the scopes granted before and other requests asked for sign-in again. Tokens now record their granted scope and step-up requests the union (SEP-2350).
Empty or null optional fields in OAuth token and registration responses
(scope, refresh_token, id_token, client_secret) failed sign-in, and
expires_in: null marked tokens as expired. Also fall back to the server
origin when resource metadata has invalid URLs, skip empty scopes when
selecting the requested scope, and end pagination at an empty or null
nextCursor.
closes#10266
The configured metadata document replaces discovery and is trusted as
configured. Authorization codes are only exchanged when the response's
iss parameter names the flow's authorization server.
closes#10172
HTTP MCP servers can set "auth": { "provider": "<provider>" } to send the provider's current /login token as the bearer token instead of using MCP OAuth. The token is read on every request, so provider refreshes apply and nothing is copied to mcp-auth.json. Only allowed in the global mcp.json and from extensions, and requires https except on loopback hosts.
MCP tool and namespace names now replace - with _ like Codex, so the
pi tool name is also the codemode identifier. Tools of a server whose
names collide all get a hash suffix, and server names that differ only
in - and _ are rejected.
closes#10239
When the session's initial tools come from defaultTools, /reload activates
names added to the resolved selection. Removals do not deactivate tools, and
explicit --tools/--no-tools/--no-builtin-tools keep overriding the setting.
closes#10245
In codemode.mode "only", tools whose declarations codemode hides were still
listed in the <tools> section. Drop their snippets so the list matches the
declared tools.
closes#10192
The first prompt waits only for servers with direct tools. Other servers
connect in the background; codemode scripts, tool_search, and the
resource tools wait for the servers they need. codemode and tool_search
are activated from the config.
Tool declarations no longer depend on connected servers: the codemode
description leaves out deferred tools, and the tool_search description
no longer lists sources. Servers are listed in an mcp_servers system
prompt section instead, set in before_agent_start from the config and
connected state, so changes are appended to the conversation.
describeNamespace() and searchTools() accept mcp__dev-radius,
mcp__dev_radius, dev-radius, and dev_radius.
closes#10212
getBranchSelection looked up the model catalog once per assistant message,
which made prompt submission scale with session length. Walk the branch
backward instead: only the last model_change can hold against later
responses, so at most that model needs a lookup.
closes#10198
Adds oauth.clientName (and pi mcp add --oauth-client-name) to set the
client_name sent during dynamic client registration. Defaults to pi.
closes#10226
MCP codemode exposure no longer lists tools, tool counts, or server
instructions in the codemode description. Servers are listed by name and
an optional mcp.json description; scripts find tools with searchTools()
and read instructions with describeNamespace(). codemode-deferred is an
alias for codemode.
refs #10212
image() now rejects malformed base64 and data without a PNG, JPEG, GIF,
or WebP signature, and derives the MIME type from the signature. Invalid
image blocks were persisted and made every later provider request fail.
closes#10215
The lazy OAuth entry list in the bundle build was missing openai-chatgpt,
so /login with OpenAI failed with a missing module error. Also fail the
build if any importOAuthModule() target in load.ts has no lazy entry.
Add gpt-6.1-sol to the OpenAI, Azure OpenAI Responses, and OpenAI Codex
providers with official pricing ($2 input, $0.10 cached input, $2.50
cache writes, $10 output, long-context tier above 272k). The model
rejects reasoning.effort "none", so off is unsupported and clamps to
the lowest effort.
Addresses GHSA-82fw-gwwq-j7x9 (vitest mocker path traversal) and
GHSA-3wwx-pv8p-q78v (undici 6.x permessage-deflate DoS). Both are
dev/example-only; the shipped coding-agent shrinkwrap is unchanged.
Structured bash results now hold the full output up to 1 MiB instead of the
model-facing 2000 lines or 50KB, keeping the first and last 512 KiB of longer
output, and add truncated and full_output_path.
Codemode sums the usage of models.classify() calls into its tool result and shows each call's cost. Usage of tools called through ctx.executeTool() is added to the calling tool's result usage instead of being dropped.
System One services (TypeSafe, OpenRouter, OpenCode Zen, Cloudflare Workers AI) return usage.input_tokens/output_tokens. ClassifierResult now carries usage priced from the model catalog like chat usage.
Vercel evaluation models are generated from the AI Gateway catalog and
routed through its TypeSafe-compatible /typesafe/v1/systemone endpoint.
OpenCode Zen's jev-1.13 and jev-1.13-free are routed through
/zen/v1/systemone.
Tools without a custom call renderer now show their arguments as key=value
pairs when collapsed and one key: value line per argument when expanded.
MCP tools get a server/tool title and results that collapse to 5 lines.
Entries of only +name/-name modify the inherited tool selection, so
"defaultTools": ["+codemode"] enables codemode without repeating the
defaults. Project entries layer on top of user settings. Document how to
enable codemode without MCP and how to use classifier models from it.
Anthropic, OpenAI Codex, OpenRouter, and Radius now use one loopback callback server in pi-ai. MCP sign-in renders the same browser page through a new renderPage option on the pi-mcp callback server.
Built-in extensions (mcp, llama.cpp, codemode, tool-search) are now extension resources like files: the package manager resolves them as builtin:<name>, the extensions setting disables them with -builtin:<name> (per project too), pi config lists them, and --no-extensions turns them off unless loaded with -e builtin:<name>. They load after project trust is resolved and after file and package extensions.
Built-in extensions and tools are named builtin:<name> in errors, diagnostics, RPC source info, and bug reports.
Refreshes hold a per-server lock file from reading the tokens to saving new ones, and use tokens another process already refreshed. Rotating refresh tokens (Cloudflare) were lost when two pi processes refreshed at once, leaving the server needing sign-in. Closing a connection waits for a running refresh so shutdown does not drop rotated tokens.
Adds fullscreenWheelScrollLines ("auto" or 1-100). Auto keeps one line
per event in local macOS terminals and scales with wheel velocity
elsewhere, capped at six lines per event.
closes#9758
This adds codemode and MCP support to pi. Codemode (packages/codemode) runs
model-written JavaScript in a QuickJS wasm VM inside a worker. From there the
script calls pi's tools as async functions, uses a per-session store, and can
read the model catalog and run classifiers. The codemode tool in coding-agent
is compatible with Codex's exec tool, and codemode.mode works like Codex's
tool modes (on or only). Large results are written to a JSON temp file. MCP
(packages/mcp) is a standalone client for stdio and streamable HTTP servers,
with OAuth.
The core only gets general mechanisms; codemode, tool_search and MCP are
built-in extensions that use them. Other extensions can replace these
built-ins. Tools declare an exposure: direct, model-only, codemode, deferred
or hidden. prepareLoadout() lets an orchestrating tool change which tool
declarations the model sees. ctx.executeTool() runs nested calls through
validation and the tool_call/tool_result hooks. Those calls emit events with
parentToolCallId and are recorded as bounded nestedCalls on the parent
result, which compaction and HTML export use. Tools can also return
structuredContent with an outputSchema, and can report errors with isError
instead of throwing. MCP servers come from mcp.json (global, or per project
after the project is trusted) or from pi.registerMcpServer(). They are
managed with /mcp and pi mcp list|login|logout, and their runtimes load on
first use.
Closes#10040
Theme previews re-render the whole transcript, and streaming re-renders
every frame, so per-line and per-session work dominated in long sessions.
- visibleWidth: fast path for ANSI-styled ASCII, allocation-free escape scanning
- Box: compare unpadded child lines and skip a second width measurement
- Markdown: reuse parsed tokens across invalidation
- footer: cache usage totals and context usage per session state
- bash results: cache the collapsed preview including the expand hint
- sanitizeBinaryOutput: single regex replace
Together no longer lists Kimi K2.6, which broke type checks against the hydrated catalog. Tests use Kimi K3 and the Together default model now points at it.
This adds experimental support for virtual models. A virtual model is a catalog entry that an extension registers with pi.registerVirtualModel(). It does not talk to a provider itself; for every request it picks a physical model and thinking level. The motivation is to make routing policies pluggable: plan on a strong model and implement on a cheap one, pick a model with a classifier, or move to a larger window when the context grows. Today this needs model switching by hand or hacks in prepareRequest.
The selection stays virtual: model_change records it, resume restores it, and each response records the physical model and thinking level that produced it. Router state can be stored on the session branch, so it follows /tree and forks and survives compaction. Context limits, compaction and image sizing use the physical model a request was routed to. examples/extensions/jev-router.ts shows a full router that uses the Jev classifier to pick Sol or Terra for planning and hands off to Luna after the first edit. docs/virtual-models.md describes the API.
Add the llama-cpp-classify classifier API. Each question becomes one chat prompt with single-token answer labels; llama-server's pre-sampling next-token log-probabilities give the softmax over those labels. Missing labels are retried with deeper readouts and then reported as errors.
models.dev dropped Kimi K2.6 from OpenCode Go, which broke type checks
against the regenerated catalog. Use OpenCode Go Kimi K3 for
provider-level tests and OpenCode Zen Kimi K2.6 for thinking format tests.
* feat(coding-agent,tui): add system theme derived from terminal colors
The new default system theme builds pi's colors from the terminal's reported foreground, background, and ANSI palette. Tokens take their hue and relative saturation from a palette slot; lightness is placed by OKLab lightness difference and WCAG 2 contrast minimums. Terminals that report only a background get built-in hues, and terminals that report nothing get ANSI indices with faint secondary text.
The TUI queries OSC 10, 11, and 4 in one burst ended by DA1, so replies complete in one round trip, and reports replies that arrive after the timeout. Startup never waits: the system theme starts in grayscale and updates when colors arrive. Light/dark detection now uses the reported colors instead of the ?996 query; mode 2031 notifications trigger a new query.
* fix(coding-agent): wait for terminal colors before building the startup header
The header and startup notices bake theme colors into their text. When the color replies arrived after they were built, as in Windows Terminal, they stayed grayscale. Startup now waits for the first color query, which ends at the DA1 reply or after 100 ms.
* fix(coding-agent): rebuild header and notices on theme changes
The startup header, loaded resources, and chat notices baked theme colors into their text, so they kept the old colors after a theme switch or when the system theme received the terminal's colors. They now use ThemedText, which rebuilds its text from a function whenever the UI is invalidated.
* fix(tui): keep focus on components that forward mouse events to children
A component that forwards a mouse event to a child it hosts, such as SettingsList with an open submenu, left keyboard focus on the child. When the host removed the child, focus stayed on a detached component and all keys were lost. The forwarding component now keeps focus, like delegating containers do.
* feat(coding-agent): describe the system theme in the theme selector
* feat(coding-agent): list the system theme above automatic in the theme selector
* refactor(coding-agent,tui): remove leftovers of the light/dark startup detection
Theme detection only reports dark or light now; the source and confidence fields only served saving the detected theme. The first-time setup no longer shows the detected appearance. Removes the TUI's unused queryTerminalColorScheme(), queryTerminalBackgroundColor(), queryTerminalForegroundColor(), and parseOsc11BackgroundColor(); queryTerminalColors() covers them.
* feat(coding-agent): match the reviewed system and light/dark theme design
The system theme now follows the reviewed design: color families with OKHSL saturation curves, palette slots per family, contrast rules per token and panel, relaxation for mid-gray backgrounds, and the terminal foreground for body text. Each contrast level is a target-lightness curve fitted to the reviewed output. The built-in dark and light themes use the reviewed Pi dark and light colors.
* feat(coding-agent): show the pi logo in the startup header
A two-line pixel logo replaces the app name. Its first line carries the version, its second line the first line of key hints.
* feat(tui,coding-agent): support OKHSL colors in themes
parseColor() accepts okhsl(H S L), and okhslColor() and colorToOkhsl() construct and read OKHSL colors. OKHSL saturation is relative to the sRGB gamut, so every value is in gamut. Theme files accept okhsl() values; HTML exports convert them to hex. The OKHSL code moves from coding-agent into pi-tui, where colors.ts now shares its Oklab conversion instead of keeping a second copy.
* refactor(coding-agent): write the built-in themes in OKHSL
The built-in dark and light themes use okhsl() values, with variables for colors that several roles share, so they read as color families and serve as editable templates. Near-identical values collapsed into one; no color changes visibly.
* refactor(coding-agent,tui): simplify terminal color and theme code
Share the terminal color query between startup and the theme controller, merge the query's resolve/late-reply state into one callback, keep a single global for terminal colors, use oklab.ts arrays directly in colors.ts, and drop leftover helpers and exports.
* fix(coding-agent): use one terminal light/dark decision everywhere
Theme pairs, the system theme, and Theme.appearance now share getTerminalTheme(): the reported background decides, then the terminal's light/dark report, then COLORFGBG, then dark.
COLORFGBG is classified by palette index like Vim (0-6 and 8 dark) and only reads the last field, so rxvt's "default" no longer picks up the foreground index.
Simplify the tests added on this branch.
* fix(ai): use an OpenRouter model that still exists in beta test
OpenRouter removed anthropic/claude-3-haiku, which broke type checking in CI after model regeneration.
Replace the tsgo native preview with typescript@7.0.2 and port the check scripts to the TS 7 API. Target ES2024, drop useDefineForClassFields: false and unused decorator options, and enable verbatimModuleSyntax. Replace tsx with node plus the source resolver hook, passed to --import as a file URL.
closes#9965
Expose theme colors as values and add theme.style() for combined styling.
- @earendil-works/pi-tui gains a Color type (palette index, sRGB, OKLCH) with parseColor(), mixColors(), colorToHex(), colorToRgb(), styleText() and related helpers.
- Theme JSON accepts #rgb and oklch(...) values and an optional appearance ("dark" or "light"), detected from the theme colors when omitted.
- theme.style() combines tokens or concrete colors with text attributes; theme.colors exposes concrete colors, including the terminal's reported default colors (OSC 10/11) for tokens set to "".
- Terminals with TERM=*-direct are detected as truecolor.
OpenRouter decision models use its TypeSafe-compatible /api/v1/systemone endpoint. Cloudflare Workers AI gets a cloudflare-workers-ai-system-one classifier API for the /ai/run envelope. Request and answer parsing is shared in system-one-shared.ts.
Move catalog layout, index validation, version selection, and request
negotiation into scripts/model-catalog-protocol.ts. pi.dev keeps an
identical copy, so the publisher and a test running the current client
exercise the same selection pi.dev performs. Replaces the cross-repository
compatibility dispatch.
Related to #9099.
Since 0.86 the prompt and tool declarations live in transcript system
messages, so a context handler that filters or slices messages could drop
them. After extension-driven compaction this sent requests without built-in
tools and made Codex emit raw tool-call text.
context handlers now see the conversation only. An unchanged list keeps the
transcript as is; a changed list gets the replayed prompt sections and tool
declarations as one leading system message. Handler-added system messages
are kept after it.
Add context_with_system, which runs after every context handler on the full
transcript and sends its result verbatim, for extensions that need to edit
system messages per request. Dropping the leading system message is reported
as an extension error.
closes#9789, closes#9822
3349e1db1 (#9618) limited OSC 52 to SSH/mosh so a desktop terminal that ignores
it no longer reports a false success. That also disabled the only clipboard
route for containers and WSL without WSLg, which have no display to gate on.
Emit OSC 52 again on display-less Linux. On WSL, write the Windows clipboard
through PowerShell (verified, no OSC 52 size cap), preferring OSC 52 in
Windows Terminal where it is known to work. Desktop sessions with a display
still report the failure.
closes#9688