### What does this PR do?
Type of change: documentation / repo housekeeping
Centralize agent-shared assets under **`.agents/`** as the single,
agent-agnostic source of truth, so the same `SKILL.md` (plus shared
scripts and cluster config) works across every coding agent without
maintaining N copies that drift out of sync. Claude Code discovers
skills only under `.claude/skills/`, so `.claude/` holds **relative
in-repo symlinks** into `.agents/` for back-compat.
```bash
repo-root/
├── .agents/ ← canonical source of truth
│ ├── README.md
│ ├── clusters.yaml.example
│ ├── scripts/
│ │ └── sync-upstream-skills.sh
│ └── skills/
│ ├── accessing-mlflow/ compare-results/ debug/
│ ├── deployment/ eagle3-new-model/ eagle3-review-logs/
│ ├── eagle3-triage/ eagle3-validate/ evaluation/
│ ├── launching-evals/ monitor/ ptq/
│ ├── quant-recipe-search/ release-cherry-pick/ common/
│
├── .claude/ ← back-compat (relative symlinks)
│ ├── clusters.yaml.example → ../.agents/clusters.yaml.example
│ ├── scripts → ../.agents/scripts
│ └── skills → ../.agents/skills
│
└── (future agents — add a symlink/config, no copies)
├── .codex/skills → ../.agents/skills
└── .cursor/skills → ../.agents/skills
```
### Why symlinks (and not "just point each agent's config at
`.agents/`")
Claude Code **only** auto-discovers project skills under
`.claude/skills/` — there is no setting/env var to redirect discovery to
an arbitrary path, and the plugin route would require committing a
`.claude/settings.json` + marketplace manifest, add a first-open
workspace-trust gate (breaks headless/CI runs), and namespace every
skill (`/ptq` → `/<plugin>:ptq`). A single relative in-repo symlink is
the smallest change that keeps `.agents/` canonical while satisfying
Claude Code's discovery requirement. This repo already commits relative
symlinks (`CLAUDE.md`, `tools/launcher/modules/Model-Optimizer`). See
the discussion thread for the full comparison.
### Changes
- Move `.claude/{skills,scripts,clusters.yaml.example}` → `.agents/`
(git renames preserve history).
- Add `.agents/README.md` documenting the convention and per-agent
wiring.
- Re-add `.claude/skills`, `.claude/scripts`,
`.claude/clusters.yaml.example` as relative symlinks into `.agents/`.
- Update internal path references and lint/sync config from `.claude/`
to `.agents/` (upstream provenance paths and `.claude/clusters.yaml`
back-compat lookups left intact).
- **Merged latest `main`** and folded in skills added there since branch
time — `compare-results`, `eagle3-new-model`, `eagle3-triage`,
`eagle3-review-logs`, `eagle3-validate`, `quant-recipe-search`, and new
`evaluation` recipes/tasks/references — into `.agents/skills/`.
- `main` converted `CLAUDE.md` into a symlink to a new agent-agnostic
`AGENTS.md`; this PR keeps that and moves the "skills live in
`.agents/`" guidance into `AGENTS.md`.
### Testing
- `.claude/skills` symlink resolves to all 15 skills; `ls
.claude/skills` and `ls .agents/skills` match.
- `pre-commit run check-symlinks --all-files` and `markdownlint-cli2
--all-files` pass.
- `bash -n` clean on `sync-upstream-skills.sh` and `remote_exec.sh`.
- `.claude/skills`, `.claude/scripts`, `.claude/clusters.yaml.example`,
and `CLAUDE.md` are all recorded as git symlinks (mode `120000`).
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).
- Is this change backward compatible?: ✅ `.claude/skills/`,
`.claude/scripts/`, and `.claude/clusters.yaml.example` continue to
resolve to the same content via symlinks; Claude Code auto-discovery is
unaffected; `remote_exec.sh` still accepts `.claude/clusters.yaml`.
- Did you write any new necessary tests?: N/A — directory move with
symlinks; verified as listed above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — repo housekeeping only, no API/feature/bugfix change.
### Additional Information
- A few vendored skill files still carry internal details (Slurm account
names, lustre paths, internal `:5005` GitLab registry advice in
`launching-evals/`) worth scrubbing in a follow-up.
---------
Signed-off-by: Seonghee Lee <seongheel@nvidia.com>
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Shengliang Xu <106840466+shengliangxu@users.noreply.github.com>
Co-authored-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
7.7 KiB
Terminal Bench
Terminal Bench is an agentic benchmark where models interact with a terminal environment to solve tasks.
Key files
terminal_bench/agents/terminus_2/terminus_2.py— main agent implementationterminal_bench/agents/failure_mode.py— failure mode definitionsterminal_bench/harness/harness.py— harness and result aggregationcore_evals/nvidia_terminal_bench/framework.yml— default config values
Key Facts
- Task-first ordering:
task1.1-of-N, task1.2-of-N, ..., task2.1-of-N, ...— mid-run results are biased toward early tasks.
Failure Modes
All failure modes (see failure_mode.py):
UNSET— no failure mode triggered (task ran to completion)NONE— explicitly set: no failure (task solved)UNSOLVED— task not completed within constraintsTOKEN_LIMIT_EXCEEDED— agent hitmax_input_tokens_per_task(cumulative input tokens across all turns). Shows asoutcome: token_limit_exceededintask_status.json.PARSE_ERROR— harness couldn't parse the test output (post-test.txt), e.g. pytest output missingshort test summary infoFATAL_LLM_PARSE_ERROR— unrecoverable LLM/agent response parse errorCONTEXT_LENGTH_EXCEEDED— input exceeded model's context window (see Context Recovery)OUTPUT_LENGTH_EXCEEDED— response truncated bymax_completion_tokens; agent retries; recorded when all retries exhausted. Shows asfinish_reason: lengthineval_factory_metrics.json.TEST_TIMEOUT— test verification timed outAGENT_TIMEOUT— agent execution timed out (see Mitigating Agent Timeouts)UNKNOWN_AGENT_ERROR— unexpected agent error (stops eval on default policy)AGENT_INSTALLATION_FAILED— agent setup failed (stops eval on default policy)UNKNOWN— unknown harness error (stops eval on default policy)
failed_samples_policy (default: default) — only stops on "no fair chance" failures: UNKNOWN, UNKNOWN_AGENT_ERROR, AGENT_INSTALLATION_FAILED. All other failures continue with score 0.
Artifacts
All paths relative to <output_dir>/<invocation>/terminal-bench-hard/.
Client logs
logs/client-*.log — contains rich/ANSI formatting (binary), always use grep -a. Shows live progress (Running tasks (X/Y, Accuracy: Z%)) and crash diagnostics.
Run-level artifacts
Path: artifacts/terminal-bench/
| File | Written | Updated | Content |
|---|---|---|---|
tb.lock |
Run start | Never | Full resolved config: invocation args, agent kwargs (max_episodes, temperature, max_input_tokens_per_task), run config (n_concurrent_trials, global_agent_timeout_sec, failed_samples_policy), ECS/sandbox settings. Best for reproducing runs. |
run_metadata.json |
Run start | Once at end | model_name, dataset_name/dataset_version, n_concurrent_trials, task_ids, start_time/end_time, accuracy, pass_at_k |
task_status.json |
After 1st task | After each task | One entry per task (not per trial). status (success/failed), outcome, trial_name. "Success is sticky" — once a task succeeds, later failures don't overwrite. 48 entries total. |
tb_results.json |
After 1st task | After each task | See below |
Mid-run: task_status.json and tb_results.json grow incrementally. run_metadata.json exists but lacks final metrics.
tb_results.json details
The richest single artifact.
Per-trial fields:
is_resolved(bool) — ground truth for whether the task was solved. Use this, notpassedorscore.failure_mode,parser_results(dict of test name → "passed"/"failed")instruction— full task description given to the agent- Token usage:
total_input_tokens,total_output_tokens trajectory_length— number of agent episodes (turns)- Timestamps:
trial_started_at,agent_started_at/ended_at,test_started_at/ended_at recording_path— asciinema.castfile for replaying terminal sessionserror_type,error_message— populated on crashes
Aggregate fields:
pass_at_k,accuracy,n_resolved,n_unresolvedresolved_ids,unresolved_idsfailure_mode_counts,error_type_counts,token_limit_exceeded_counttotal_input_tokens,total_output_tokens— run-wide totals
Per-trial artifacts/terminal-bench/<task>/<trial>/results.json files are the source — tb_results.json aggregates them (same schema).
Per-trial artifacts
Path: artifacts/terminal-bench/<task>/<trial>/
Agent logs (agent-logs/episode-N/, N = 0, 1, 2, ...):
prompt.txt— full prompt sent to the model (system instructions + task + terminal state)response.txt— model's raw response (JSON withanalysis,plan,commands,task_complete)debug.json— LiteLLM trace: model, messages, optional_params,reasoning/reasoning_content(chain-of-thought), token usage,llm_api_duration_ms, response headers
Panes (panes/) — terminal screen snapshots:
pre-agent.txt— before agent starts (initial prompt)post-agent.txt— after agent finishes (all commands and outputs)post-test.txt— after test verification. Iffailure_mode: parse_error, check this first; for pytest tasks the summary block may be missing.
Panes are useful for quick triage without reading episode logs.
Troubleshooting
Mitigating Agent Timeouts
High AGENT_TIMEOUT rates (e.g. 85%+) are caused by inference contention: too many concurrent agent sessions competing for the same vLLM instance.
Two levers reduce contention: lower parallelism (fewer concurrent tasks) and scale inference (more deployment nodes / data-parallel replicas). Scaling inference has diminishing returns — requesting 32–64 nodes means long queue times and harder Slurm scheduling. The recommended approach combines both:
Split into independent single-sample runs with lower parallelism (8x1 pattern):
Instead of one run with n_samples: 8, parallelism: 100, submit 8 independent runs each with n_samples: 1 and reduced parallelism: 24. This scales horizontally with multiple smaller jobs.
Context Recovery
When the agent's input exceeds the model's context window, terminus_2 has two recovery paths. Both rely on litellm.get_max_tokens(model_name) to determine the context limit.
Proactive path (_check_proactive_summarization): Fires when free_tokens < 8000 before the API call. Summarizes while the full conversation history is still available. This is the healthier path.
Reactive path (on ContextLengthExceededError): Fires after the API rejects a request:
- Unwind (
_unwind_messages_to_free_tokens): Drops the most recent user+assistant pairs untilfree_tokens >= 4000. Destructive — removed messages are permanently lost. - Summarize (
_summarize): Asks the model (using truncated history) to summarize, generates questions from summary +capture_pane(), answers from truncated history, resetschat._messagesto just 3 messages (original instruction + Q&A).
Reactive path flaw: Unwind drops recent messages before summarize runs. The terminal reflects those actions but the summary doesn't contain them. Only capture_pane() partially compensates.
LiteLLM context limit is often wrong: litellm.get_max_tokens() returns the advertised context window, not the deployment limit. For unknown models it falls back to 1M tokens; for --max-model-len smaller than default, it reports the full spec. When the limit is too high, unwind removes nothing, summarize hits the same error, and recovery is a no-op — propagates as CONTEXT_LENGTH_EXCEEDED.
Agent Trace Analysis
See references/benchmarks/terminal-bench-trace-analysis.md for analyzing per-task agent traces, extracting behavior patterns, and categorizing failures.