mirror of
https://github.com/earendil-works/pi.git
synced 2026-10-02 08:44:38 +08:00
feat(coding-agent): add models.generateImages() to codemode
Scripts run image models with the session's credentials, like models.classify(). Results are base64 image blocks that image() attaches to the codemode result; usage counts toward the session cost.
This commit is contained in:
@@ -6,6 +6,7 @@
|
||||
|
||||
- Added an `oauth.authServerMetadataUrl` setting for MCP servers that advertise a wrong OAuth authorization server or none. Pi uses the configured metadata document instead of discovery ([#10172](https://github.com/earendil-works/pi/issues/10172)).
|
||||
- Added `quietStartup: "header"`, which keeps the startup header with version and key hints but hides the model scope line and loaded-resource listing.
|
||||
- Added `models.generateImages()` to codemode scripts. It runs image models such as OpenRouter's with the session's credentials and returns base64 image blocks that `image()` attaches to the result; usage counts toward the session cost like `models.classify()`. Extensions can call `ctx.modelRegistry.generateImages()`. See [Use image models](docs/models.md#use-image-models).
|
||||
|
||||
### Changed
|
||||
|
||||
|
||||
@@ -161,7 +161,7 @@ This keeps `read`, `bash`, `edit`, and `write` and adds `codemode`. For one invo
|
||||
pi --tools read,bash,edit,write,codemode
|
||||
```
|
||||
|
||||
Codemode is useful without MCP: scripts can run several tool calls in parallel, filter large output before it reaches the model, and call classifier models such as TypeSafe's Jev through `models.classify()` (see [Classifier models](models.md#use-classifier-models)).
|
||||
Codemode is useful without MCP: scripts can run several tool calls in parallel, filter large output before it reaches the model, call classifier models such as TypeSafe's Jev through `models.classify()` (see [Classifier models](models.md#use-classifier-models)), and generate images through `models.generateImages()` (see [Image models](models.md#use-image-models)).
|
||||
|
||||
### How codemode works
|
||||
|
||||
@@ -175,7 +175,7 @@ The `codemode` description lists the callable tools with their TypeScript declar
|
||||
|
||||
Tools with an output schema resolve to structured values: `bash` to `{ output, truncated, full_output_path?, exit_code, wall_time_seconds }`, also for non-zero exit codes, and MCP tools to their `CallToolResult`. Other tools resolve to their text output. The `output` of `bash` is not limited to the 2000 lines or 50KB the model sees: it holds up to 1 MiB, and longer output keeps its first and last 512 KiB around an omission marker, with `truncated` set and the full output in `full_output_path`.
|
||||
|
||||
`store(key, value)` and `load(key)` keep JSON values across `codemode` calls: each successful script that stores values appends a `codemode-store` custom entry to the session, so resumed sessions keep the values and each branch sees only the values written on its path. Scripts can also use `models`: `getModelsOfType`, `getAvailableOfType`, and `getModelOfType` list the model catalog, and `classify(model, context)` runs a classifier model with the session's credentials, at most four at a time per script.
|
||||
`store(key, value)` and `load(key)` keep JSON values across `codemode` calls: each successful script that stores values appends a `codemode-store` custom entry to the session, so resumed sessions keep the values and each branch sees only the values written on its path. Scripts can also use `models`: `getModelsOfType`, `getAvailableOfType`, and `getModelOfType` list the model catalog, `classify(model, context)` runs a classifier model, and `generateImages(model, context)` runs an image model. Both use the session's credentials, with at most four such calls in flight per script.
|
||||
|
||||
### Tool search
|
||||
|
||||
|
||||
@@ -135,6 +135,25 @@ When the service reports token counts, as all System One services do, `result.us
|
||||
|
||||
Extensions call classifiers through `ctx.modelRegistry.classify()`, without codemode. [Virtual models](virtual-models.md#route-requests) can use them to route requests; see the `jev-router.ts` example.
|
||||
|
||||
## Use image models
|
||||
|
||||
Image models generate images from a prompt and optional input images. Pi lists OpenRouter's image models, such as `google/gemini-2.5-flash-image` and `black-forest-labs/flux.2-pro`, under the `openrouter` provider; they use the same `OPENROUTER_API_KEY` or `/login` credential as its chat models.
|
||||
|
||||
Like classifier models, image models do not appear in `/model`; the model reaches them through the [`codemode`](cli.md#enable-codemode) tool. Scripts list them with `models.getAvailableOfType("image")` and call `models.generateImages(model, { input })`. The result's `output` holds base64 image blocks, which `image()` attaches to the `codemode` result so the model sees them:
|
||||
|
||||
```js
|
||||
const painter = await models.getModelOfType("image", "openrouter", "google/gemini-2.5-flash-image");
|
||||
const result = await models.generateImages(painter, {
|
||||
input: [{ type: "text", text: "A red fox in the snow, watercolor" }],
|
||||
});
|
||||
if (result.stopReason !== "stop") return result.errorMessage;
|
||||
for (const block of result.output) if (block.type === "image") image(block);
|
||||
```
|
||||
|
||||
`input` can also contain `{ type: "image", data, mimeType }` blocks to edit or use as references. Pi adds the usage of a script's image calls to the `codemode` tool result, like classifier calls. Generated images are not saved to disk.
|
||||
|
||||
Extensions generate images through `ctx.modelRegistry.generateImages()`, without codemode.
|
||||
|
||||
## Add a custom provider
|
||||
|
||||
Use an extension when the provider needs custom streaming, model discovery, or authentication behavior. See [Custom Providers](custom-provider.md) for the extension workflow.
|
||||
|
||||
@@ -1,5 +1,6 @@
|
||||
import type {
|
||||
Api,
|
||||
AssistantImages,
|
||||
AssistantMessage,
|
||||
AssistantMessageEventStream,
|
||||
AuthOperationOptions,
|
||||
@@ -9,9 +10,13 @@ import type {
|
||||
ClassifierModel,
|
||||
ClassifierResult,
|
||||
Context,
|
||||
ImageApi,
|
||||
ImageModel,
|
||||
ImagesContext,
|
||||
Model,
|
||||
ModelsApiStreamOptions,
|
||||
ModelsClassifierOptions,
|
||||
ModelsImagesOptions,
|
||||
ModelsRefreshOptions,
|
||||
ModelsRefreshResult,
|
||||
ModelsSimpleStreamOptions,
|
||||
@@ -173,6 +178,15 @@ export class ModelRegistry {
|
||||
return this.runtime.classify(model, context, options);
|
||||
}
|
||||
|
||||
/** Generate images with request-time authentication. Never rejects. */
|
||||
generateImages(
|
||||
model: ImageModel<ImageApi>,
|
||||
context: ImagesContext,
|
||||
options?: ModelsImagesOptions,
|
||||
): Promise<AssistantImages> {
|
||||
return this.runtime.generateImages(model, context, options);
|
||||
}
|
||||
|
||||
getProviderDisplayName(provider: string): string {
|
||||
return this.runtime.getProvider(provider)?.name ?? provider;
|
||||
}
|
||||
|
||||
@@ -8,7 +8,16 @@ import { writeFile } from "node:fs/promises";
|
||||
import { tmpdir } from "node:os";
|
||||
import { join } from "node:path";
|
||||
import type { AgentTool, AgentToolCallOutcome, AgentToolResult } from "@earendil-works/pi-agent-core";
|
||||
import type { AnyModel, ClassifierContext, ImageContent, ModelType, TextContent, Usage } from "@earendil-works/pi-ai";
|
||||
import type {
|
||||
AnyModel,
|
||||
ClassifierContext,
|
||||
ImageContent,
|
||||
ImagesContext,
|
||||
ModelType,
|
||||
ModelTypeMap,
|
||||
TextContent,
|
||||
Usage,
|
||||
} from "@earendil-works/pi-ai";
|
||||
import {
|
||||
type CodemodeResult,
|
||||
CodemodeSandbox,
|
||||
@@ -38,7 +47,7 @@ import {
|
||||
|
||||
const ARGS_PREVIEW_CHARS = 200;
|
||||
const ERROR_PREVIEW_CHARS = 500;
|
||||
/** Classifier calls one script may have in flight; `Promise.all` over many items queues the rest. */
|
||||
/** `models.classify()` and `models.generateImages()` calls one script may have in flight; `Promise.all` over many items queues the rest. */
|
||||
const MAX_CONCURRENT_MODEL_CALLS = 4;
|
||||
/**
|
||||
* Heap limit for the QuickJS VM. The worker shares pi's process, so without a limit a runaway
|
||||
@@ -86,6 +95,13 @@ function toModelInfo(model: AnyModel): Record<string, unknown> {
|
||||
return info;
|
||||
}
|
||||
|
||||
/** The fields of `ClassifierResult` and `AssistantImages` that a nested call row reports. */
|
||||
interface ModelCallResult {
|
||||
stopReason: "stop" | "error" | "aborted";
|
||||
errorMessage?: string;
|
||||
usage?: Usage;
|
||||
}
|
||||
|
||||
/** Runs at most `limit` calls at once, in call order. */
|
||||
function createLimiter(limit: number): <T>(run: () => Promise<T>) => Promise<T> {
|
||||
let active = 0;
|
||||
@@ -404,8 +420,8 @@ function createDiscoveryGlobals(
|
||||
|
||||
/**
|
||||
* `models.*` for scripts: the model registry methods declared in {@link MODEL_GLOBAL_DECLARATIONS}.
|
||||
* Classifier calls appear as nested call rows so the renderer shows them, and their usage goes to
|
||||
* `addUsage`.
|
||||
* Classifier and image calls appear as nested call rows so the renderer shows them, and their usage
|
||||
* goes to `addUsage`. Rows show only the model, never prompts or image data.
|
||||
*/
|
||||
function createModelGlobals(
|
||||
models: CodemodeModelRuntime,
|
||||
@@ -415,7 +431,45 @@ function createModelGlobals(
|
||||
addUsage: (usage: Usage) => void,
|
||||
): CodemodeTool[] {
|
||||
const limit = createLimiter(MAX_CONCURRENT_MODEL_CALLS);
|
||||
let classifyCount = 0;
|
||||
let callCount = 0;
|
||||
|
||||
/**
|
||||
* Resolve the script's model by provider and id only, then run the call as a nested call row. A
|
||||
* script-supplied baseUrl or headers must never receive the credentials.
|
||||
*/
|
||||
const runModelCall = async <TType extends "classifier" | "image", TResult extends ModelCallResult>(
|
||||
name: string,
|
||||
type: TType,
|
||||
model: unknown,
|
||||
run: (resolved: ModelTypeMap[TType]) => Promise<TResult>,
|
||||
): Promise<TResult> => {
|
||||
const ref = model as { provider?: unknown; id?: unknown } | null;
|
||||
if (typeof ref !== "object" || ref === null || typeof ref.provider !== "string" || typeof ref.id !== "string") {
|
||||
throw new Error(`${name}() expects a model from models.getModelOfType() or models.getAvailableOfType()`);
|
||||
}
|
||||
const resolved = models.getModelOfType(type, ref.provider, ref.id);
|
||||
if (!resolved) throw new Error(`Unknown ${type} model "${ref.provider}/${ref.id}"`);
|
||||
|
||||
const record: CodemodeNestedCall = {
|
||||
id: `${toolCallId}/${name}/${++callCount}`,
|
||||
name,
|
||||
args: `${resolved.provider}/${resolved.id}`,
|
||||
status: "running",
|
||||
};
|
||||
calls.push(record);
|
||||
publish();
|
||||
const startedAt = performance.now();
|
||||
const result = await limit(() => run(resolved));
|
||||
record.durationMs = performance.now() - startedAt;
|
||||
record.status = result.stopReason === "stop" ? "ok" : result.stopReason === "aborted" ? "cancelled" : "error";
|
||||
if (result.errorMessage) record.error = truncateText(result.errorMessage, ERROR_PREVIEW_CHARS);
|
||||
if (result.usage) {
|
||||
record.cost = result.usage.cost.total;
|
||||
addUsage(result.usage);
|
||||
}
|
||||
publish();
|
||||
return result;
|
||||
};
|
||||
const implementations: Record<string, CodemodeTool["execute"]> = {
|
||||
"models.getModelsOfType": (args) => {
|
||||
const [type, provider] = args as unknown[];
|
||||
@@ -434,42 +488,17 @@ function createModelGlobals(
|
||||
const model = models.getModelOfType(toModelType(type), provider, id);
|
||||
return model === undefined ? undefined : toModelInfo(model);
|
||||
},
|
||||
"models.classify": async (args, { signal }) => {
|
||||
"models.classify": (args, { signal }) => {
|
||||
const [model, context] = args as unknown[];
|
||||
const ref = model as { provider?: unknown; id?: unknown } | null;
|
||||
if (
|
||||
typeof ref !== "object" ||
|
||||
ref === null ||
|
||||
typeof ref.provider !== "string" ||
|
||||
typeof ref.id !== "string"
|
||||
) {
|
||||
throw new Error(
|
||||
"models.classify() expects a model from models.getModelOfType() or models.getAvailableOfType()",
|
||||
);
|
||||
}
|
||||
// Only provider and id count. A script-supplied baseUrl or headers must never receive the credentials.
|
||||
const resolved = models.getModelOfType("classifier", ref.provider, ref.id);
|
||||
if (!resolved) throw new Error(`Unknown classifier model "${ref.provider}/${ref.id}"`);
|
||||
|
||||
const record: CodemodeNestedCall = {
|
||||
id: `${toolCallId}/models.classify/${++classifyCount}`,
|
||||
name: "models.classify",
|
||||
args: `${resolved.provider}/${resolved.id}`,
|
||||
status: "running",
|
||||
};
|
||||
calls.push(record);
|
||||
publish();
|
||||
const startedAt = performance.now();
|
||||
const result = await limit(() => models.classify(resolved, context as ClassifierContext, { signal }));
|
||||
record.durationMs = performance.now() - startedAt;
|
||||
record.status = result.stopReason === "stop" ? "ok" : result.stopReason === "aborted" ? "cancelled" : "error";
|
||||
if (result.errorMessage) record.error = truncateText(result.errorMessage, ERROR_PREVIEW_CHARS);
|
||||
if (result.usage) {
|
||||
record.cost = result.usage.cost.total;
|
||||
addUsage(result.usage);
|
||||
}
|
||||
publish();
|
||||
return result;
|
||||
return runModelCall("models.classify", "classifier", model, (resolved) =>
|
||||
models.classify(resolved, context as ClassifierContext, { signal }),
|
||||
);
|
||||
},
|
||||
"models.generateImages": (args, { signal }) => {
|
||||
const [model, context] = args as unknown[];
|
||||
return runModelCall("models.generateImages", "image", model, (resolved) =>
|
||||
models.generateImages(resolved, context as ImagesContext, { signal }),
|
||||
);
|
||||
},
|
||||
};
|
||||
return MODEL_GLOBAL_DECLARATIONS.map((declaration) => ({
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
/**
|
||||
* The `codemode` tool: the model writes JavaScript that calls other tools. Scripts use `tools`,
|
||||
* `ALL_TOOLS`, `text()`, `image()`, `exit()`, `store()`/`load()`, `console.*`, and `return <value>`,
|
||||
* may start with a `// @options:` line, and reach the model catalog and classifiers through
|
||||
* `models.*`. Results start with a "Script completed" or "Script failed" header.
|
||||
* may start with a `// @options:` line, and reach the model catalog, classifiers, and image models
|
||||
* through `models.*`. Results start with a "Script completed" or "Script failed" header.
|
||||
*
|
||||
* Scripts can call the agent loop's nested tools: active `direct` tools and every `codemode` or
|
||||
* `deferred` tool. Nested calls run through the agent loop's tool pipeline (`ctx.executeTool`), so
|
||||
@@ -58,7 +58,7 @@ export interface CodemodeStoreEntryData {
|
||||
/** The part of the model registry that scripts reach through `models`. */
|
||||
export type CodemodeModelRuntime = Pick<
|
||||
ModelRegistry,
|
||||
"getModelsOfType" | "getAvailableOfType" | "getModelOfType" | "classify"
|
||||
"getModelsOfType" | "getAvailableOfType" | "getModelOfType" | "classify" | "generateImages"
|
||||
>;
|
||||
|
||||
export interface CodemodeToolOptions {
|
||||
@@ -180,13 +180,32 @@ interface ClassifierContext {
|
||||
state: Record<string, unknown>;
|
||||
questions: Record<string, ClassifierQuestion>;
|
||||
}
|
||||
/** Token counts reported by the service. Cost is in USD. */
|
||||
type ModelUsage = { input: number; output: number; totalTokens: number; cost: { total: number } };
|
||||
interface ClassifierResult {
|
||||
api: string;
|
||||
provider: string;
|
||||
model: string;
|
||||
answers: Record<string, ClassifierAnswer>;
|
||||
/** Set when the service reports token counts. Cost is in USD. */
|
||||
usage?: { input: number; output: number; totalTokens: number; cost: { total: number } };
|
||||
usage?: ModelUsage;
|
||||
stopReason: "stop" | "error" | "aborted";
|
||||
errorMessage?: string;
|
||||
timestamp: number;
|
||||
}
|
||||
type ModelTextBlock = { type: "text"; text: string };
|
||||
/** \`data\` is base64. Show it with \`image(block)\`; never print \`data\` with \`text()\`, \`console\`, or \`return\`. */
|
||||
type ModelImageBlock = { type: "image"; data: string; mimeType: string };
|
||||
interface ImagesContext {
|
||||
/** The prompt as text blocks, plus image blocks to edit or use as references. */
|
||||
input: (ModelTextBlock | ModelImageBlock)[];
|
||||
}
|
||||
interface ImagesResult {
|
||||
api: string;
|
||||
provider: string;
|
||||
model: string;
|
||||
/** Generated images, and text blocks for models that also return text. */
|
||||
output: (ModelTextBlock | ModelImageBlock)[];
|
||||
usage?: ModelUsage;
|
||||
stopReason: "stop" | "error" | "aborted";
|
||||
errorMessage?: string;
|
||||
timestamp: number;
|
||||
@@ -215,6 +234,12 @@ export const MODEL_GLOBAL_DECLARATIONS: readonly Omit<CodemodeTool, "execute">[]
|
||||
"Run a classifier model on one state. Only `provider` and `id` of `model` are used. Provider errors do not throw: check `stopReason` and `errorMessage`.",
|
||||
signature: "(model: ModelInfo, context: ClassifierContext): Promise<ClassifierResult>",
|
||||
},
|
||||
{
|
||||
name: "models.generateImages",
|
||||
description:
|
||||
'Generate images with an image model. Only `provider` and `id` of `model` are used. Provider errors do not throw: check `stopReason` and `errorMessage`. Generation can take minutes, so do not set a short `timeout_ms`. Show results with `for (const block of result.output) if (block.type === "image") image(block);`.',
|
||||
signature: "(model: ModelInfo, context: ImagesContext): Promise<ImagesResult>",
|
||||
},
|
||||
];
|
||||
|
||||
const DEFERRED_TOOLS_GUIDANCE = `Some nested tools may be omitted from this description, such as deferred tools and MCP tools. They are still available on the global \`tools\` object and listed in \`ALL_TOOLS\`.
|
||||
|
||||
@@ -1,12 +1,15 @@
|
||||
import { readFileSync, rmSync } from "node:fs";
|
||||
import type { AgentTool } from "@earendil-works/pi-agent-core";
|
||||
import {
|
||||
type AssistantImages,
|
||||
type ClassifierModel,
|
||||
type ClassifierResult,
|
||||
fauxAssistantMessage,
|
||||
fauxToolCall,
|
||||
getCurrentSystemPrompt,
|
||||
getCurrentTools,
|
||||
type ImageModel,
|
||||
type ImagesContext,
|
||||
type TranscriptContext,
|
||||
} from "@earendil-works/pi-ai";
|
||||
import type { ToolResultMessage, Usage } from "@earendil-works/pi-ai/compat";
|
||||
@@ -548,12 +551,30 @@ describe("codemode models", () => {
|
||||
headers: { "X-Secret": "hunter2" },
|
||||
};
|
||||
|
||||
const painterModel: ImageModel<"test-images"> = {
|
||||
type: "image",
|
||||
id: "painter",
|
||||
name: "Painter",
|
||||
api: "test-images",
|
||||
provider: "scorer",
|
||||
baseUrl: "https://images.test/v1",
|
||||
input: ["text", "image"],
|
||||
output: ["text", "image"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
};
|
||||
|
||||
interface ClassifyObservation {
|
||||
baseUrl: string;
|
||||
apiKey: string | undefined;
|
||||
text: unknown;
|
||||
}
|
||||
|
||||
interface ImagesObservation {
|
||||
baseUrl: string;
|
||||
apiKey: string | undefined;
|
||||
input: ImagesContext["input"];
|
||||
}
|
||||
|
||||
async function setup() {
|
||||
const harness = await createHarness({
|
||||
initialActiveToolNames: ["codemode"],
|
||||
@@ -561,11 +582,33 @@ describe("codemode models", () => {
|
||||
});
|
||||
harnesses.push(harness);
|
||||
const observed: ClassifyObservation[] = [];
|
||||
const imageRequests: ImagesObservation[] = [];
|
||||
let active = 0;
|
||||
let maxActive = 0;
|
||||
harness.session.modelRuntime.registerProvider("scorer", {
|
||||
apiKey: "secret-key",
|
||||
models: [scorerModel],
|
||||
models: [scorerModel, painterModel],
|
||||
images: {
|
||||
"test-images": {
|
||||
generateImages: async (model, context, options): Promise<AssistantImages> => {
|
||||
imageRequests.push({ baseUrl: model.baseUrl, apiKey: options?.apiKey, input: context.input });
|
||||
const prompt = context.input.find((block) => block.type === "text")?.text;
|
||||
const base = { api: model.api, provider: model.provider, model: model.id, timestamp: 0 };
|
||||
if (prompt === "explode") {
|
||||
return { ...base, output: [], stopReason: "error", errorMessage: "painter exploded" };
|
||||
}
|
||||
return {
|
||||
...base,
|
||||
output: [
|
||||
{ type: "text", text: `painted ${prompt}` },
|
||||
{ type: "image", data: TINY_PNG_BASE64, mimeType: "image/png" },
|
||||
],
|
||||
usage: usage(100, 0.04),
|
||||
stopReason: "stop",
|
||||
};
|
||||
},
|
||||
},
|
||||
},
|
||||
classifiers: {
|
||||
"test-classifier": {
|
||||
classify: async (model, context, options): Promise<ClassifierResult> => {
|
||||
@@ -600,7 +643,7 @@ describe("codemode models", () => {
|
||||
},
|
||||
});
|
||||
harness.session.setActiveToolsByName(["codemode"]);
|
||||
return { harness, observed, maxActive: () => maxActive };
|
||||
return { harness, observed, imageRequests, maxActive: () => maxActive };
|
||||
}
|
||||
|
||||
async function run(harness: Harness, code: string): Promise<ToolResultMessage> {
|
||||
@@ -622,6 +665,10 @@ describe("codemode models", () => {
|
||||
"classify(model: ModelInfo, context: ClassifierContext): Promise<ClassifierResult>;",
|
||||
);
|
||||
expect(codemode?.description).toContain("interface ClassifierResult {");
|
||||
expect(codemode?.description).toContain(
|
||||
"generateImages(model: ModelInfo, context: ImagesContext): Promise<ImagesResult>;",
|
||||
);
|
||||
expect(codemode?.description).toContain("interface ImagesResult {");
|
||||
|
||||
const overridden = await createHarness({ tools: [createCodemodeTool() as AgentTool] });
|
||||
harnesses.push(overridden);
|
||||
@@ -678,6 +725,59 @@ describe("codemode models", () => {
|
||||
expect(harness.session.getSessionStats().cost).toBeCloseTo(0.006, 10);
|
||||
});
|
||||
|
||||
it("generates images with catalog auth and attaches them through image()", async () => {
|
||||
const { harness, imageRequests } = await setup();
|
||||
const result = await run(
|
||||
harness,
|
||||
`
|
||||
const [model] = await models.getAvailableOfType("image", "scorer");
|
||||
const reference = { type: "image", data: "${TINY_PNG_BASE64}", mimeType: "image/png" };
|
||||
const generated = await models.generateImages(
|
||||
{ ...model, baseUrl: "https://evil.test" },
|
||||
{ input: [{ type: "text", text: "a fox" }, reference] },
|
||||
);
|
||||
for (const block of generated.output) {
|
||||
if (block.type === "image") image(block);
|
||||
else text(block.text);
|
||||
}
|
||||
const failed = await models.generateImages(model, { input: [{ type: "text", text: "explode" }] });
|
||||
const attempt = async (fn) => { try { await fn(); return "ok"; } catch (error) { return error.message; } };
|
||||
return {
|
||||
id: model.id,
|
||||
stopReason: generated.stopReason,
|
||||
failed: [failed.stopReason, failed.errorMessage],
|
||||
wrongType: await attempt(() => models.generateImages({ provider: "scorer", id: "judge" }, { input: [] })),
|
||||
};
|
||||
`,
|
||||
);
|
||||
expect(result.isError).toBe(false);
|
||||
const [text, ...rest] = resultText(result).split("\n");
|
||||
expect(text).toBe("painted a fox");
|
||||
expect(rest[0]).toBe("<image>");
|
||||
expect(JSON.parse(rest.slice(1).join("\n"))).toEqual({
|
||||
id: "painter",
|
||||
stopReason: "stop",
|
||||
failed: ["error", "painter exploded"],
|
||||
wrongType: 'Unknown image model "scorer/judge"',
|
||||
});
|
||||
expect(result.content[2]).toEqual({ type: "image", data: TINY_PNG_BASE64, mimeType: "image/png" });
|
||||
expect(imageRequests.map((request) => [request.baseUrl, request.apiKey])).toEqual([
|
||||
["https://images.test/v1", "secret-key"],
|
||||
["https://images.test/v1", "secret-key"],
|
||||
]);
|
||||
expect(imageRequests[0].input).toEqual([
|
||||
{ type: "text", text: "a fox" },
|
||||
{ type: "image", data: TINY_PNG_BASE64, mimeType: "image/png" },
|
||||
]);
|
||||
const details = result.details as unknown as CodemodeToolDetails;
|
||||
expect(details.calls.map((call) => [call.name, call.args, call.status, call.cost, call.error])).toEqual([
|
||||
["models.generateImages", "scorer/painter", "ok", 0.04, undefined],
|
||||
["models.generateImages", "scorer/painter", "error", undefined, "painter exploded"],
|
||||
]);
|
||||
expect(result.usage?.cost.total).toBeCloseTo(0.04, 10);
|
||||
expect(harness.session.getSessionStats().cost).toBeCloseTo(0.04, 10);
|
||||
});
|
||||
|
||||
it("reports provider errors as results and invalid arguments as exceptions", async () => {
|
||||
const { harness } = await setup();
|
||||
const result = await run(
|
||||
|
||||
Reference in New Issue
Block a user