Files
Model-Optimizer/.agents/skills/launching-evals/references/benchmarks/terminal-bench-general-info.md
T
bdb793a42d [SKILL.md Chore] Make .agents/ the canonical agent-skills location (#1362)
### What does this PR do?

Type of change: documentation / repo housekeeping

Centralize agent-shared assets under **`.agents/`** as the single,
agent-agnostic source of truth, so the same `SKILL.md` (plus shared
scripts and cluster config) works across every coding agent without
maintaining N copies that drift out of sync. Claude Code discovers
skills only under `.claude/skills/`, so `.claude/` holds **relative
in-repo symlinks** into `.agents/` for back-compat.

```bash
repo-root/
├── .agents/                      ← canonical source of truth
│   ├── README.md
│   ├── clusters.yaml.example
│   ├── scripts/
│   │   └── sync-upstream-skills.sh
│   └── skills/
│       ├── accessing-mlflow/      compare-results/     debug/
│       ├── deployment/            eagle3-new-model/    eagle3-review-logs/
│       ├── eagle3-triage/         eagle3-validate/     evaluation/
│       ├── launching-evals/       monitor/             ptq/
│       ├── quant-recipe-search/   release-cherry-pick/ common/
│
├── .claude/                      ← back-compat (relative symlinks)
│   ├── clusters.yaml.example  →  ../.agents/clusters.yaml.example
│   ├── scripts                →  ../.agents/scripts
│   └── skills                 →  ../.agents/skills
│
└── (future agents — add a symlink/config, no copies)
    ├── .codex/skills           →  ../.agents/skills
    └── .cursor/skills          →  ../.agents/skills
```

### Why symlinks (and not "just point each agent's config at
`.agents/`")

Claude Code **only** auto-discovers project skills under
`.claude/skills/` — there is no setting/env var to redirect discovery to
an arbitrary path, and the plugin route would require committing a
`.claude/settings.json` + marketplace manifest, add a first-open
workspace-trust gate (breaks headless/CI runs), and namespace every
skill (`/ptq` → `/<plugin>:ptq`). A single relative in-repo symlink is
the smallest change that keeps `.agents/` canonical while satisfying
Claude Code's discovery requirement. This repo already commits relative
symlinks (`CLAUDE.md`, `tools/launcher/modules/Model-Optimizer`). See
the discussion thread for the full comparison.

### Changes

- Move `.claude/{skills,scripts,clusters.yaml.example}` → `.agents/`
(git renames preserve history).
- Add `.agents/README.md` documenting the convention and per-agent
wiring.
- Re-add `.claude/skills`, `.claude/scripts`,
`.claude/clusters.yaml.example` as relative symlinks into `.agents/`.
- Update internal path references and lint/sync config from `.claude/`
to `.agents/` (upstream provenance paths and `.claude/clusters.yaml`
back-compat lookups left intact).
- **Merged latest `main`** and folded in skills added there since branch
time — `compare-results`, `eagle3-new-model`, `eagle3-triage`,
`eagle3-review-logs`, `eagle3-validate`, `quant-recipe-search`, and new
`evaluation` recipes/tasks/references — into `.agents/skills/`.
- `main` converted `CLAUDE.md` into a symlink to a new agent-agnostic
`AGENTS.md`; this PR keeps that and moves the "skills live in
`.agents/`" guidance into `AGENTS.md`.

### Testing

- `.claude/skills` symlink resolves to all 15 skills; `ls
.claude/skills` and `ls .agents/skills` match.
- `pre-commit run check-symlinks --all-files` and `markdownlint-cli2
--all-files` pass.
- `bash -n` clean on `sync-upstream-skills.sh` and `remote_exec.sh`.
- `.claude/skills`, `.claude/scripts`, `.claude/clusters.yaml.example`,
and `CLAUDE.md` are all recorded as git symlinks (mode `120000`).

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).

- Is this change backward compatible?: ✅ `.claude/skills/`,
`.claude/scripts/`, and `.claude/clusters.yaml.example` continue to
resolve to the same content via symlinks; Claude Code auto-discovery is
unaffected; `remote_exec.sh` still accepts `.claude/clusters.yaml`.
- Did you write any new necessary tests?: N/A — directory move with
symlinks; verified as listed above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — repo housekeeping only, no API/feature/bugfix change.

### Additional Information

- A few vendored skill files still carry internal details (Slurm account
names, lustre paths, internal `:5005` GitLab registry advice in
`launching-evals/`) worth scrubbing in a follow-up.

---------

Signed-off-by: Seonghee Lee <seongheel@nvidia.com>
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Shengliang Xu <106840466+shengliangxu@users.noreply.github.com>
Co-authored-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 16:12:19 -07:00

7.7 KiB
Raw Blame History

Terminal Bench

Terminal Bench is an agentic benchmark where models interact with a terminal environment to solve tasks.

Key files

  • terminal_bench/agents/terminus_2/terminus_2.py — main agent implementation
  • terminal_bench/agents/failure_mode.py — failure mode definitions
  • terminal_bench/harness/harness.py — harness and result aggregation
  • core_evals/nvidia_terminal_bench/framework.yml — default config values

Key Facts

  • Task-first ordering: task1.1-of-N, task1.2-of-N, ..., task2.1-of-N, ... — mid-run results are biased toward early tasks.

Failure Modes

All failure modes (see failure_mode.py):

  • UNSET — no failure mode triggered (task ran to completion)
  • NONE — explicitly set: no failure (task solved)
  • UNSOLVED — task not completed within constraints
  • TOKEN_LIMIT_EXCEEDED — agent hit max_input_tokens_per_task (cumulative input tokens across all turns). Shows as outcome: token_limit_exceeded in task_status.json.
  • PARSE_ERROR — harness couldn't parse the test output (post-test.txt), e.g. pytest output missing short test summary info
  • FATAL_LLM_PARSE_ERROR — unrecoverable LLM/agent response parse error
  • CONTEXT_LENGTH_EXCEEDED — input exceeded model's context window (see Context Recovery)
  • OUTPUT_LENGTH_EXCEEDED — response truncated by max_completion_tokens; agent retries; recorded when all retries exhausted. Shows as finish_reason: length in eval_factory_metrics.json.
  • TEST_TIMEOUT — test verification timed out
  • AGENT_TIMEOUT — agent execution timed out (see Mitigating Agent Timeouts)
  • UNKNOWN_AGENT_ERROR — unexpected agent error (stops eval on default policy)
  • AGENT_INSTALLATION_FAILED — agent setup failed (stops eval on default policy)
  • UNKNOWN — unknown harness error (stops eval on default policy)

failed_samples_policy (default: default) — only stops on "no fair chance" failures: UNKNOWN, UNKNOWN_AGENT_ERROR, AGENT_INSTALLATION_FAILED. All other failures continue with score 0.

Artifacts

All paths relative to <output_dir>/<invocation>/terminal-bench-hard/.

Client logs

logs/client-*.log — contains rich/ANSI formatting (binary), always use grep -a. Shows live progress (Running tasks (X/Y, Accuracy: Z%)) and crash diagnostics.

Run-level artifacts

Path: artifacts/terminal-bench/

File Written Updated Content
tb.lock Run start Never Full resolved config: invocation args, agent kwargs (max_episodes, temperature, max_input_tokens_per_task), run config (n_concurrent_trials, global_agent_timeout_sec, failed_samples_policy), ECS/sandbox settings. Best for reproducing runs.
run_metadata.json Run start Once at end model_name, dataset_name/dataset_version, n_concurrent_trials, task_ids, start_time/end_time, accuracy, pass_at_k
task_status.json After 1st task After each task One entry per task (not per trial). status (success/failed), outcome, trial_name. "Success is sticky" — once a task succeeds, later failures don't overwrite. 48 entries total.
tb_results.json After 1st task After each task See below

Mid-run: task_status.json and tb_results.json grow incrementally. run_metadata.json exists but lacks final metrics.

tb_results.json details

The richest single artifact.

Per-trial fields:

  • is_resolved (bool) — ground truth for whether the task was solved. Use this, not passed or score.
  • failure_mode, parser_results (dict of test name → "passed"/"failed")
  • instruction — full task description given to the agent
  • Token usage: total_input_tokens, total_output_tokens
  • trajectory_length — number of agent episodes (turns)
  • Timestamps: trial_started_at, agent_started_at/ended_at, test_started_at/ended_at
  • recording_path — asciinema .cast file for replaying terminal sessions
  • error_type, error_message — populated on crashes

Aggregate fields:

  • pass_at_k, accuracy, n_resolved, n_unresolved
  • resolved_ids, unresolved_ids
  • failure_mode_counts, error_type_counts, token_limit_exceeded_count
  • total_input_tokens, total_output_tokens — run-wide totals

Per-trial artifacts/terminal-bench/<task>/<trial>/results.json files are the source — tb_results.json aggregates them (same schema).

Per-trial artifacts

Path: artifacts/terminal-bench/<task>/<trial>/

Agent logs (agent-logs/episode-N/, N = 0, 1, 2, ...):

  • prompt.txt — full prompt sent to the model (system instructions + task + terminal state)
  • response.txt — model's raw response (JSON with analysis, plan, commands, task_complete)
  • debug.json — LiteLLM trace: model, messages, optional_params, reasoning/reasoning_content (chain-of-thought), token usage, llm_api_duration_ms, response headers

Panes (panes/) — terminal screen snapshots:

  • pre-agent.txt — before agent starts (initial prompt)
  • post-agent.txt — after agent finishes (all commands and outputs)
  • post-test.txt — after test verification. If failure_mode: parse_error, check this first; for pytest tasks the summary block may be missing.

Panes are useful for quick triage without reading episode logs.

Troubleshooting

Mitigating Agent Timeouts

High AGENT_TIMEOUT rates (e.g. 85%+) are caused by inference contention: too many concurrent agent sessions competing for the same vLLM instance.

Two levers reduce contention: lower parallelism (fewer concurrent tasks) and scale inference (more deployment nodes / data-parallel replicas). Scaling inference has diminishing returns — requesting 32–64 nodes means long queue times and harder Slurm scheduling. The recommended approach combines both:

Split into independent single-sample runs with lower parallelism (8x1 pattern):

Instead of one run with n_samples: 8, parallelism: 100, submit 8 independent runs each with n_samples: 1 and reduced parallelism: 24. This scales horizontally with multiple smaller jobs.

Context Recovery

When the agent's input exceeds the model's context window, terminus_2 has two recovery paths. Both rely on litellm.get_max_tokens(model_name) to determine the context limit.

Proactive path (_check_proactive_summarization): Fires when free_tokens < 8000 before the API call. Summarizes while the full conversation history is still available. This is the healthier path.

Reactive path (on ContextLengthExceededError): Fires after the API rejects a request:

  1. Unwind (_unwind_messages_to_free_tokens): Drops the most recent user+assistant pairs until free_tokens >= 4000. Destructive — removed messages are permanently lost.
  2. Summarize (_summarize): Asks the model (using truncated history) to summarize, generates questions from summary + capture_pane(), answers from truncated history, resets chat._messages to just 3 messages (original instruction + Q&A).

Reactive path flaw: Unwind drops recent messages before summarize runs. The terminal reflects those actions but the summary doesn't contain them. Only capture_pane() partially compensates.

LiteLLM context limit is often wrong: litellm.get_max_tokens() returns the advertised context window, not the deployment limit. For unknown models it falls back to 1M tokens; for --max-model-len smaller than default, it reports the full spec. When the limit is too high, unwind removes nothing, summarize hits the same error, and recovery is a no-op — propagates as CONTEXT_LENGTH_EXCEEDED.

Agent Trace Analysis

See references/benchmarks/terminal-bench-trace-analysis.md for analyzing per-task agent traces, extracting behavior patterns, and categorizing failures.