PaperclipEvaluation directory
Evidence for better agents

Paperclip Evals

Two ways to test Paperclip. Follow each program from a test case to its evidence, results, and run history.

01 / RUNNER

Runner Evals

Does the runner use tools correctly, respect permissions, and preserve its execution contract?

Real runner + model provider
Controlled mock Paperclip control plane
Latest recorded measurement
$runner_counts
$runner_date · $runner_coverage
$runner_id
Runner operations guide
02 / PRODUCT

Product E2E Evals

Can a user complete work through Paperclip, from the first request to a usable result?

Real browser + server + database + runner + model provider
Local and Daytona environments, as selected
Latest recorded measurement
$product_counts
$product_date · $product_coverage
$product_id
Product E2E operations guide

Snapshot refreshed $refreshed. History pages may contain newer runs. A partial campaign covers only its selected cases. These programs measure different boundaries; their pass rates are not combined.

Choose a test. Make the result inspectable.

The eval guide explains which program to use, how to add a case, and how to distinguish a product defect from model behavior, a grading error, or a setup failure. Evalbook is a report framework used to present evidence; it is not an execution boundary.

Read the eval guide →