feat(coding-agent): add models.generateImages() to codemode

Scripts run image models with the session's credentials, like
models.classify(). Results are base64 image blocks that image() attaches
to the codemode result; usage counts toward the session cost.
This commit is contained in:
Armin Ronacher
2026-10-01 16:50:29 +02:00
parent ed8b3bcc19
commit aab34df655
7 changed files with 237 additions and 49 deletions
+1
View File
@@ -6,6 +6,7 @@
- Added an `oauth.authServerMetadataUrl` setting for MCP servers that advertise a wrong OAuth authorization server or none. Pi uses the configured metadata document instead of discovery ([#10172](https://github.com/earendil-works/pi/issues/10172)).
- Added `quietStartup: "header"`, which keeps the startup header with version and key hints but hides the model scope line and loaded-resource listing.
- Added `models.generateImages()` to codemode scripts. It runs image models such as OpenRouter's with the session's credentials and returns base64 image blocks that `image()` attaches to the result; usage counts toward the session cost like `models.classify()`. Extensions can call `ctx.modelRegistry.generateImages()`. See [Use image models](docs/models.md#use-image-models).
### Changed
+2 -2
View File
@@ -161,7 +161,7 @@ This keeps `read`, `bash`, `edit`, and `write` and adds `codemode`. For one invo
pi --tools read,bash,edit,write,codemode
```
Codemode is useful without MCP: scripts can run several tool calls in parallel, filter large output before it reaches the model, and call classifier models such as TypeSafe's Jev through `models.classify()` (see [Classifier models](models.md#use-classifier-models)).
Codemode is useful without MCP: scripts can run several tool calls in parallel, filter large output before it reaches the model, call classifier models such as TypeSafe's Jev through `models.classify()` (see [Classifier models](models.md#use-classifier-models)), and generate images through `models.generateImages()` (see [Image models](models.md#use-image-models)).
### How codemode works
@@ -175,7 +175,7 @@ The `codemode` description lists the callable tools with their TypeScript declar
Tools with an output schema resolve to structured values: `bash` to `{ output, truncated, full_output_path?, exit_code, wall_time_seconds }`, also for non-zero exit codes, and MCP tools to their `CallToolResult`. Other tools resolve to their text output. The `output` of `bash` is not limited to the 2000 lines or 50KB the model sees: it holds up to 1 MiB, and longer output keeps its first and last 512 KiB around an omission marker, with `truncated` set and the full output in `full_output_path`.
`store(key, value)` and `load(key)` keep JSON values across `codemode` calls: each successful script that stores values appends a `codemode-store` custom entry to the session, so resumed sessions keep the values and each branch sees only the values written on its path. Scripts can also use `models`: `getModelsOfType`, `getAvailableOfType`, and `getModelOfType` list the model catalog, and `classify(model, context)` runs a classifier model with the session's credentials, at most four at a time per script.
`store(key, value)` and `load(key)` keep JSON values across `codemode` calls: each successful script that stores values appends a `codemode-store` custom entry to the session, so resumed sessions keep the values and each branch sees only the values written on its path. Scripts can also use `models`: `getModelsOfType`, `getAvailableOfType`, and `getModelOfType` list the model catalog, `classify(model, context)` runs a classifier model, and `generateImages(model, context)` runs an image model. Both use the session's credentials, with at most four such calls in flight per script.
### Tool search
+19
View File
@@ -135,6 +135,25 @@ When the service reports token counts, as all System One services do, `result.us
Extensions call classifiers through `ctx.modelRegistry.classify()`, without codemode. [Virtual models](virtual-models.md#route-requests) can use them to route requests; see the `jev-router.ts` example.
## Use image models
Image models generate images from a prompt and optional input images. Pi lists OpenRouter's image models, such as `google/gemini-2.5-flash-image` and `black-forest-labs/flux.2-pro`, under the `openrouter` provider; they use the same `OPENROUTER_API_KEY` or `/login` credential as its chat models.
Like classifier models, image models do not appear in `/model`; the model reaches them through the [`codemode`](cli.md#enable-codemode) tool. Scripts list them with `models.getAvailableOfType("image")` and call `models.generateImages(model, { input })`. The result's `output` holds base64 image blocks, which `image()` attaches to the `codemode` result so the model sees them:
```js
const painter = await models.getModelOfType("image", "openrouter", "google/gemini-2.5-flash-image");
const result = await models.generateImages(painter, {
input: [{ type: "text", text: "A red fox in the snow, watercolor" }],
});
if (result.stopReason !== "stop") return result.errorMessage;
for (const block of result.output) if (block.type === "image") image(block);
```
`input` can also contain `{ type: "image", data, mimeType }` blocks to edit or use as references. Pi adds the usage of a script's image calls to the `codemode` tool result, like classifier calls. Generated images are not saved to disk.
Extensions generate images through `ctx.modelRegistry.generateImages()`, without codemode.
## Add a custom provider
Use an extension when the provider needs custom streaming, model discovery, or authentication behavior. See [Custom Providers](custom-provider.md) for the extension workflow.
@@ -1,5 +1,6 @@
import type {
Api,
AssistantImages,
AssistantMessage,
AssistantMessageEventStream,
AuthOperationOptions,
@@ -9,9 +10,13 @@ import type {
ClassifierModel,
ClassifierResult,
Context,
ImageApi,
ImageModel,
ImagesContext,
Model,
ModelsApiStreamOptions,
ModelsClassifierOptions,
ModelsImagesOptions,
ModelsRefreshOptions,
ModelsRefreshResult,
ModelsSimpleStreamOptions,
@@ -173,6 +178,15 @@ export class ModelRegistry {
return this.runtime.classify(model, context, options);
}
/** Generate images with request-time authentication. Never rejects. */
generateImages(
model: ImageModel<ImageApi>,
context: ImagesContext,
options?: ModelsImagesOptions,
): Promise<AssistantImages> {
return this.runtime.generateImages(model, context, options);
}
getProviderDisplayName(provider: string): string {
return this.runtime.getProvider(provider)?.name ?? provider;
}
@@ -8,7 +8,16 @@ import { writeFile } from "node:fs/promises";
import { tmpdir } from "node:os";
import { join } from "node:path";
import type { AgentTool, AgentToolCallOutcome, AgentToolResult } from "@earendil-works/pi-agent-core";
import type { AnyModel, ClassifierContext, ImageContent, ModelType, TextContent, Usage } from "@earendil-works/pi-ai";
import type {
AnyModel,
ClassifierContext,
ImageContent,
ImagesContext,
ModelType,
ModelTypeMap,
TextContent,
Usage,
} from "@earendil-works/pi-ai";
import {
type CodemodeResult,
CodemodeSandbox,
@@ -38,7 +47,7 @@ import {
const ARGS_PREVIEW_CHARS = 200;
const ERROR_PREVIEW_CHARS = 500;
/** Classifier calls one script may have in flight; `Promise.all` over many items queues the rest. */
/** `models.classify()` and `models.generateImages()` calls one script may have in flight; `Promise.all` over many items queues the rest. */
const MAX_CONCURRENT_MODEL_CALLS = 4;
/**
* Heap limit for the QuickJS VM. The worker shares pi's process, so without a limit a runaway
@@ -86,6 +95,13 @@ function toModelInfo(model: AnyModel): Record<string, unknown> {
return info;
}
/** The fields of `ClassifierResult` and `AssistantImages` that a nested call row reports. */
interface ModelCallResult {
stopReason: "stop" | "error" | "aborted";
errorMessage?: string;
usage?: Usage;
}
/** Runs at most `limit` calls at once, in call order. */
function createLimiter(limit: number): <T>(run: () => Promise<T>) => Promise<T> {
let active = 0;
@@ -404,8 +420,8 @@ function createDiscoveryGlobals(
/**
* `models.*` for scripts: the model registry methods declared in {@link MODEL_GLOBAL_DECLARATIONS}.
* Classifier calls appear as nested call rows so the renderer shows them, and their usage goes to
* `addUsage`.
* Classifier and image calls appear as nested call rows so the renderer shows them, and their usage
* goes to `addUsage`. Rows show only the model, never prompts or image data.
*/
function createModelGlobals(
models: CodemodeModelRuntime,
@@ -415,7 +431,45 @@ function createModelGlobals(
addUsage: (usage: Usage) => void,
): CodemodeTool[] {
const limit = createLimiter(MAX_CONCURRENT_MODEL_CALLS);
let classifyCount = 0;
let callCount = 0;
/**
* Resolve the script's model by provider and id only, then run the call as a nested call row. A
* script-supplied baseUrl or headers must never receive the credentials.
*/
const runModelCall = async <TType extends "classifier" | "image", TResult extends ModelCallResult>(
name: string,
type: TType,
model: unknown,
run: (resolved: ModelTypeMap[TType]) => Promise<TResult>,
): Promise<TResult> => {
const ref = model as { provider?: unknown; id?: unknown } | null;
if (typeof ref !== "object" || ref === null || typeof ref.provider !== "string" || typeof ref.id !== "string") {
throw new Error(`${name}() expects a model from models.getModelOfType() or models.getAvailableOfType()`);
}
const resolved = models.getModelOfType(type, ref.provider, ref.id);
if (!resolved) throw new Error(`Unknown ${type} model "${ref.provider}/${ref.id}"`);
const record: CodemodeNestedCall = {
id: `${toolCallId}/${name}/${++callCount}`,
name,
args: `${resolved.provider}/${resolved.id}`,
status: "running",
};
calls.push(record);
publish();
const startedAt = performance.now();
const result = await limit(() => run(resolved));
record.durationMs = performance.now() - startedAt;
record.status = result.stopReason === "stop" ? "ok" : result.stopReason === "aborted" ? "cancelled" : "error";
if (result.errorMessage) record.error = truncateText(result.errorMessage, ERROR_PREVIEW_CHARS);
if (result.usage) {
record.cost = result.usage.cost.total;
addUsage(result.usage);
}
publish();
return result;
};
const implementations: Record<string, CodemodeTool["execute"]> = {
"models.getModelsOfType": (args) => {
const [type, provider] = args as unknown[];
@@ -434,42 +488,17 @@ function createModelGlobals(
const model = models.getModelOfType(toModelType(type), provider, id);
return model === undefined ? undefined : toModelInfo(model);
},
"models.classify": async (args, { signal }) => {
"models.classify": (args, { signal }) => {
const [model, context] = args as unknown[];
const ref = model as { provider?: unknown; id?: unknown } | null;
if (
typeof ref !== "object" ||
ref === null ||
typeof ref.provider !== "string" ||
typeof ref.id !== "string"
) {
throw new Error(
"models.classify() expects a model from models.getModelOfType() or models.getAvailableOfType()",
);
}
// Only provider and id count. A script-supplied baseUrl or headers must never receive the credentials.
const resolved = models.getModelOfType("classifier", ref.provider, ref.id);
if (!resolved) throw new Error(`Unknown classifier model "${ref.provider}/${ref.id}"`);
const record: CodemodeNestedCall = {
id: `${toolCallId}/models.classify/${++classifyCount}`,
name: "models.classify",
args: `${resolved.provider}/${resolved.id}`,
status: "running",
};
calls.push(record);
publish();
const startedAt = performance.now();
const result = await limit(() => models.classify(resolved, context as ClassifierContext, { signal }));
record.durationMs = performance.now() - startedAt;
record.status = result.stopReason === "stop" ? "ok" : result.stopReason === "aborted" ? "cancelled" : "error";
if (result.errorMessage) record.error = truncateText(result.errorMessage, ERROR_PREVIEW_CHARS);
if (result.usage) {
record.cost = result.usage.cost.total;
addUsage(result.usage);
}
publish();
return result;
return runModelCall("models.classify", "classifier", model, (resolved) =>
models.classify(resolved, context as ClassifierContext, { signal }),
);
},
"models.generateImages": (args, { signal }) => {
const [model, context] = args as unknown[];
return runModelCall("models.generateImages", "image", model, (resolved) =>
models.generateImages(resolved, context as ImagesContext, { signal }),
);
},
};
return MODEL_GLOBAL_DECLARATIONS.map((declaration) => ({
@@ -1,8 +1,8 @@
/**
* The `codemode` tool: the model writes JavaScript that calls other tools. Scripts use `tools`,
* `ALL_TOOLS`, `text()`, `image()`, `exit()`, `store()`/`load()`, `console.*`, and `return <value>`,
* may start with a `// @options:` line, and reach the model catalog and classifiers through
* `models.*`. Results start with a "Script completed" or "Script failed" header.
* may start with a `// @options:` line, and reach the model catalog, classifiers, and image models
* through `models.*`. Results start with a "Script completed" or "Script failed" header.
*
* Scripts can call the agent loop's nested tools: active `direct` tools and every `codemode` or
* `deferred` tool. Nested calls run through the agent loop's tool pipeline (`ctx.executeTool`), so
@@ -58,7 +58,7 @@ export interface CodemodeStoreEntryData {
/** The part of the model registry that scripts reach through `models`. */
export type CodemodeModelRuntime = Pick<
ModelRegistry,
"getModelsOfType" | "getAvailableOfType" | "getModelOfType" | "classify"
"getModelsOfType" | "getAvailableOfType" | "getModelOfType" | "classify" | "generateImages"
>;
export interface CodemodeToolOptions {
@@ -180,13 +180,32 @@ interface ClassifierContext {
state: Record<string, unknown>;
questions: Record<string, ClassifierQuestion>;
}
/** Token counts reported by the service. Cost is in USD. */
type ModelUsage = { input: number; output: number; totalTokens: number; cost: { total: number } };
interface ClassifierResult {
api: string;
provider: string;
model: string;
answers: Record<string, ClassifierAnswer>;
/** Set when the service reports token counts. Cost is in USD. */
usage?: { input: number; output: number; totalTokens: number; cost: { total: number } };
usage?: ModelUsage;
stopReason: "stop" | "error" | "aborted";
errorMessage?: string;
timestamp: number;
}
type ModelTextBlock = { type: "text"; text: string };
/** \`data\` is base64. Show it with \`image(block)\`; never print \`data\` with \`text()\`, \`console\`, or \`return\`. */
type ModelImageBlock = { type: "image"; data: string; mimeType: string };
interface ImagesContext {
/** The prompt as text blocks, plus image blocks to edit or use as references. */
input: (ModelTextBlock | ModelImageBlock)[];
}
interface ImagesResult {
api: string;
provider: string;
model: string;
/** Generated images, and text blocks for models that also return text. */
output: (ModelTextBlock | ModelImageBlock)[];
usage?: ModelUsage;
stopReason: "stop" | "error" | "aborted";
errorMessage?: string;
timestamp: number;
@@ -215,6 +234,12 @@ export const MODEL_GLOBAL_DECLARATIONS: readonly Omit<CodemodeTool, "execute">[]
"Run a classifier model on one state. Only `provider` and `id` of `model` are used. Provider errors do not throw: check `stopReason` and `errorMessage`.",
signature: "(model: ModelInfo, context: ClassifierContext): Promise<ClassifierResult>",
},
{
name: "models.generateImages",
description:
'Generate images with an image model. Only `provider` and `id` of `model` are used. Provider errors do not throw: check `stopReason` and `errorMessage`. Generation can take minutes, so do not set a short `timeout_ms`. Show results with `for (const block of result.output) if (block.type === "image") image(block);`.',
signature: "(model: ModelInfo, context: ImagesContext): Promise<ImagesResult>",
},
];
const DEFERRED_TOOLS_GUIDANCE = `Some nested tools may be omitted from this description, such as deferred tools and MCP tools. They are still available on the global \`tools\` object and listed in \`ALL_TOOLS\`.
@@ -1,12 +1,15 @@
import { readFileSync, rmSync } from "node:fs";
import type { AgentTool } from "@earendil-works/pi-agent-core";
import {
type AssistantImages,
type ClassifierModel,
type ClassifierResult,
fauxAssistantMessage,
fauxToolCall,
getCurrentSystemPrompt,
getCurrentTools,
type ImageModel,
type ImagesContext,
type TranscriptContext,
} from "@earendil-works/pi-ai";
import type { ToolResultMessage, Usage } from "@earendil-works/pi-ai/compat";
@@ -548,12 +551,30 @@ describe("codemode models", () => {
headers: { "X-Secret": "hunter2" },
};
const painterModel: ImageModel<"test-images"> = {
type: "image",
id: "painter",
name: "Painter",
api: "test-images",
provider: "scorer",
baseUrl: "https://images.test/v1",
input: ["text", "image"],
output: ["text", "image"],
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
};
interface ClassifyObservation {
baseUrl: string;
apiKey: string | undefined;
text: unknown;
}
interface ImagesObservation {
baseUrl: string;
apiKey: string | undefined;
input: ImagesContext["input"];
}
async function setup() {
const harness = await createHarness({
initialActiveToolNames: ["codemode"],
@@ -561,11 +582,33 @@ describe("codemode models", () => {
});
harnesses.push(harness);
const observed: ClassifyObservation[] = [];
const imageRequests: ImagesObservation[] = [];
let active = 0;
let maxActive = 0;
harness.session.modelRuntime.registerProvider("scorer", {
apiKey: "secret-key",
models: [scorerModel],
models: [scorerModel, painterModel],
images: {
"test-images": {
generateImages: async (model, context, options): Promise<AssistantImages> => {
imageRequests.push({ baseUrl: model.baseUrl, apiKey: options?.apiKey, input: context.input });
const prompt = context.input.find((block) => block.type === "text")?.text;
const base = { api: model.api, provider: model.provider, model: model.id, timestamp: 0 };
if (prompt === "explode") {
return { ...base, output: [], stopReason: "error", errorMessage: "painter exploded" };
}
return {
...base,
output: [
{ type: "text", text: `painted ${prompt}` },
{ type: "image", data: TINY_PNG_BASE64, mimeType: "image/png" },
],
usage: usage(100, 0.04),
stopReason: "stop",
};
},
},
},
classifiers: {
"test-classifier": {
classify: async (model, context, options): Promise<ClassifierResult> => {
@@ -600,7 +643,7 @@ describe("codemode models", () => {
},
});
harness.session.setActiveToolsByName(["codemode"]);
return { harness, observed, maxActive: () => maxActive };
return { harness, observed, imageRequests, maxActive: () => maxActive };
}
async function run(harness: Harness, code: string): Promise<ToolResultMessage> {
@@ -622,6 +665,10 @@ describe("codemode models", () => {
"classify(model: ModelInfo, context: ClassifierContext): Promise<ClassifierResult>;",
);
expect(codemode?.description).toContain("interface ClassifierResult {");
expect(codemode?.description).toContain(
"generateImages(model: ModelInfo, context: ImagesContext): Promise<ImagesResult>;",
);
expect(codemode?.description).toContain("interface ImagesResult {");
const overridden = await createHarness({ tools: [createCodemodeTool() as AgentTool] });
harnesses.push(overridden);
@@ -678,6 +725,59 @@ describe("codemode models", () => {
expect(harness.session.getSessionStats().cost).toBeCloseTo(0.006, 10);
});
it("generates images with catalog auth and attaches them through image()", async () => {
const { harness, imageRequests } = await setup();
const result = await run(
harness,
`
const [model] = await models.getAvailableOfType("image", "scorer");
const reference = { type: "image", data: "${TINY_PNG_BASE64}", mimeType: "image/png" };
const generated = await models.generateImages(
{ ...model, baseUrl: "https://evil.test" },
{ input: [{ type: "text", text: "a fox" }, reference] },
);
for (const block of generated.output) {
if (block.type === "image") image(block);
else text(block.text);
}
const failed = await models.generateImages(model, { input: [{ type: "text", text: "explode" }] });
const attempt = async (fn) => { try { await fn(); return "ok"; } catch (error) { return error.message; } };
return {
id: model.id,
stopReason: generated.stopReason,
failed: [failed.stopReason, failed.errorMessage],
wrongType: await attempt(() => models.generateImages({ provider: "scorer", id: "judge" }, { input: [] })),
};
`,
);
expect(result.isError).toBe(false);
const [text, ...rest] = resultText(result).split("\n");
expect(text).toBe("painted a fox");
expect(rest[0]).toBe("<image>");
expect(JSON.parse(rest.slice(1).join("\n"))).toEqual({
id: "painter",
stopReason: "stop",
failed: ["error", "painter exploded"],
wrongType: 'Unknown image model "scorer/judge"',
});
expect(result.content[2]).toEqual({ type: "image", data: TINY_PNG_BASE64, mimeType: "image/png" });
expect(imageRequests.map((request) => [request.baseUrl, request.apiKey])).toEqual([
["https://images.test/v1", "secret-key"],
["https://images.test/v1", "secret-key"],
]);
expect(imageRequests[0].input).toEqual([
{ type: "text", text: "a fox" },
{ type: "image", data: TINY_PNG_BASE64, mimeType: "image/png" },
]);
const details = result.details as unknown as CodemodeToolDetails;
expect(details.calls.map((call) => [call.name, call.args, call.status, call.cost, call.error])).toEqual([
["models.generateImages", "scorer/painter", "ok", 0.04, undefined],
["models.generateImages", "scorer/painter", "error", undefined, "painter exploded"],
]);
expect(result.usage?.cost.total).toBeCloseTo(0.04, 10);
expect(harness.session.getSessionStats().cost).toBeCloseTo(0.04, 10);
});
it("reports provider errors as results and invalid arguments as exceptions", async () => {
const { harness } = await setup();
const result = await run(