### What does this PR do? Type of change: new feature Packages the existing ModelOpt agent skills as installable Codex and Claude plugins: - Adds a repo-scoped Codex marketplace and Claude-compatible marketplace. - Adds the canonical `plugins/modelopt/` plugin tree and manifests. - Moves the skill tree into the plugin and keeps `.agents/skills` as a compatibility symlink. - Adds a minimal `common` placeholder skill required by Codex validation. - Documents installation from this repository. ### Usage ```bash codex plugin marketplace add NVIDIA/Model-Optimizer ``` Then open `/plugins`, select the `modelopt` marketplace, and install `modelopt`. For Claude Code: ```bash claude plugin marketplace add https://github.com/NVIDIA/Model-Optimizer.git claude plugin install modelopt@modelopt ``` ### Testing - Codex plugin validator - `claude plugin validate . --strict` - `claude plugin validate plugins/modelopt --strict` 1. Install the marketplace plugin with Codex and Claude from an unrelated temporary workspace. 2. Exercise packaged evaluation helpers, a day-0 gate, and the shared remote helper from that workspace. 3. Run `uv run --frozen --extra dev python -m pytest -q plugins/modelopt/skills/day0-release/tests/test_gates.py plugins/modelopt/skills/benchmark-model-kernels/tests`. 4. Run pre-commit hooks for all changed files. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — added a plugin-path validator; existing focused skill tests and installed-plugin smoke tests pass. - Did you update Changelog?: N/A — agent tooling and distribution only. - Did you get Claude approval on this PR?: N/A ### Additional Information Skills remain available through `.agents/skills`; bundled helpers are packaged under the plugin and resolved from `$SKILL_DIR` so installed workflows do not depend on the current workspace. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added installable ModelOpt plugins for Claude Code and Codex. * Added skills for PTQ, deployment, evaluation, monitoring, debugging, benchmarking, MLflow access, EAGLE3 workflows, and release management. * Added deployment helpers, evaluation recipes, checkpoint validation, and release-gating tools. * **Documentation** * Expanded setup, credential, SLURM, benchmarking, deployment, evaluation, troubleshooting, and workspace guidance. * Added installation instructions and updated agent-skill discovery guidance. * **Maintenance** * Updated skill references and compatibility links for reliable use across supported plugin environments. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
6.5 KiB
GDPVal (NeMo Gym "Stirrup" agent)
Task Details
- Reference:
references/gym-gdpval.md(SIF build, gym machinery, deploy sizing, scoring modes, failure modes) — read it before editing a GDPVal config. - Upstream README: https://github.com/NVIDIA-NeMo/Evaluator/blob/main/examples/nemotron/nemotron-3-ultra/v0.2/README.md
GDPVal is an agentic benchmark: the Stirrup agent produces office/PDF
deliverables inside a per-task Apptainer code-exec sandbox, then a pairwise/rubric
judge (Gemini 3.1 Pro) scores them. It is the most resource-intensive benchmark
in the suite — 220 tasks, num_repeats=1, 4 judge trials per rollout.
It runs on the 0.2.6 nel launcher as a nemo_gym task (NOT nel-next), so
Steps 1–9 apply — but with the branch differences below.
What makes GDPVal different (not a normal aa/ task)
- Standalone — one gym eval per config. Never add GDPVal to a multi-task
evaluation.taskslist, and never add other tasks to a GDPVal config. - Apptainer SIF sandbox — prefer a site-provided SIF; otherwise
$SKILL_DIR/scripts/gdpval-sif.shbuilds one into$GDPVAL_SIF_DIR(build-if-absent, never copied between clusters). Missing/misnamed → silent unsandboxed exec. - Thinking mode is mandatory — non-thinking loses ~86% of pairwise judgements.
Serve with the model's
--reasoning-parserand force it on via the adapter'schat_template_kwargs. - Scoring:
rubric(template default, no references, no ELO) vscomparison(the AA-comparablenormalized_elo; a conversion, not a flag flip). - Needs
INFERENCE_API_KEY,TAVILY_API_KEY,INFERENCE_JUDGE_URL,GDPVAL_SIF_DIRin.env, plusNEMO_EVALUATOR_TRUST_PRE_CMD=1(the config has apre_cmd).
All of the above — SIF handling, the SIF↔Gym-commit coupling, scoring modes, judge
panel, preflight and failure modes — is detailed in references/gym-gdpval.md.
Read it before editing a GDPVal config.
Config
Start from the self-contained example and edit it — do not copy a fragment into another config:
recipes/examples/gym_gdpval/example_gym_gdpval.yaml # SLURM + single-node vLLM,
# rubric mode, self-contained
num_repeats=1 — already set by the template via ++num_repeats=1; both
current goldens use it. A full 220-task run of a large MoE typically needs multi-node.
Canary — limit_samples does NOT work here
++…params.limit_samples=N is inert on the gym path. The gym does its own data
prep and rollout collection, so the launcher-level limiter is ignored: you get the
full 220-task run. Do not use it believing you launched a two-task smoke test — this
is the heaviest benchmark in the suite.
There is no cheap sample-limited canary. Instead, launch the real run and treat its first ~20–30 minutes as the canary, cancelling if any of these is wrong:
RD=<output_dir>/<run>/nemo_gym.0
grep -c "Using Apptainer container" $RD/logs/client-*.log # sandbox actually used
grep -c "falling back\|not a git repo" $RD/logs/client-*.log # unsandboxed / inert pin
grep -ciE " 401 | 403 |Internal Server Error" $RD/artifacts/nemo_gym_logs/gdpval_judge_model.log
wc -l $RD/artifacts/evaluator_rollouts.jsonl # rollouts flowing
In comparison mode stage 1 (45 tasks) is a natural early checkpoint — an ELO estimate appears before the full 220-task stage 2 starts.
Score Extraction
The GDPVal score is NOT in
artifacts/eval_factory_metrics.json. That file holds onlyresponse_stats/reasoning/evaluation(request-level telemetry). Looking there and finding no ELO does not mean the run failed to score.
The reported GDPVal score is normalized_elo — the AA 0–1 scale, comparable
across models and to the published AA index. eval_elo is the same fit on the raw
Elo axis (normalized_elo = (eval_elo - 500) / 2000); quote it as supporting
detail, not as the score.
The final numbers live in artifacts/results.yml (authoritative, local) and are
mirrored to MLflow. Read them by metric name:
| Mode | Metric (results.yml → groups.nemo_gym.metrics.<name>.scores.<name>.value) |
|---|---|
| comparison | gdpval_stirrup_agent/comparison/normalized_elo ← REPORT THIS (AA 0–1 scale) |
| comparison | gdpval_stirrup_agent/comparison/eval_elo (raw Elo; supporting detail) |
| comparison | gdpval_stirrup_agent/comparison/win_rate, /judged, /wins, /losses, /ties |
| comparison | per-reference: gdpval_stirrup_agent/comparison/ref/<ref_key>/{win_rate,wins,losses,ties,judged} |
| comparison | per-stage estimate: gdpval_stirrup_agent/comparison/stage_0/eval_elo (stage 1, all refs) — the final value is the top-level one, from the last stage |
| rubric | mean of reward across artifacts/evaluator_rollouts.jsonl (per-rollout 0–1) |
# COMPARISON mode — final score from the local results file (no MLflow needed)
python3 -c "
import yaml
m=yaml.safe_load(open('<output_dir>/<run>/nemo_gym.0/artifacts/results.yml'))['groups']['nemo_gym']['metrics']
for k in ('normalized_elo','eval_elo','win_rate'):
n=f'gdpval_stirrup_agent/comparison/{k}'
print(k, '=', m[n]['scores'][n]['value'])"
# RUBRIC mode (the template default) — there is no ELO; average the per-rollout reward
python3 -c "
import json
r=[json.loads(l).get('reward') for l in open('<output_dir>/<run>/nemo_gym.0/artifacts/evaluator_rollouts.jsonl')]
r=[x for x in r if isinstance(x,(int,float))]
print('mean reward =', sum(r)/len(r), 'over', len(r), 'rollouts')"
In MLflow the same values are prefixed nemo_gym_ and duplicated under a
key_metrics/ path — query these exact keys rather than browsing the UI, because a
comparison run logs ~200 metrics and most of them are per-reference, so the
headline is easy to miss:
nemo_gym_gdpval_stirrup_agent/key_metrics/comparison/normalized_elo <- report this
nemo_gym_gdpval_stirrup_agent/key_metrics/comparison/eval_elo
nemo_gym_gdpval_stirrup_agent/key_metrics/comparison/win_rate
Sanity checks before quoting a score: …/comparison/judged should be large (a few
hundred+), num_stages/num_references should match your multistage config, and the
unique task_id count in evaluator_rollouts.jsonl should be close to 220 — a short
count means tasks were lost (e.g. across a walltime resume) and the ELO is computed on
fewer tasks than the references were. Per-task detail is in
evaluator_rollouts.jsonl + nemo_gym_logs/; raw judge responses are under
PERSIST_DELIVERABLES_DIR.