[skill] evaluation: add GDPVal (NeMo Gym Stirrup agent) support (#2039)

### What does this PR do?

Type of change: new feature (agent skill)

Adds GDPVal support to the `evaluation` agent skill. GDPVal is an
agentic AA
benchmark: the NeMo Gym "Stirrup" agent produces office/PDF deliverables
inside a
per-task Apptainer code-exec sandbox, and a judge panel scores them. It
runs on the
0.2.6 launcher as a `nemo_gym` task, but it is **standalone** (one gym
eval per
config) and mechanically unlike the `aa/` nemo-skills tasks, so it gets
its own
branch in the skill rather than being merged into the `aa/` task list.

- `recipes/tasks/aa_gym/gdpval.md` — task recipe: standalone rule,
rubric-vs-comparison
  scoring, canary, score extraction.
- `references/gym-gdpval.md` — the machinery: Apptainer SIF sandbox, the
`_gym_prepare` venv-repair / process-group-reap workaround, deployment
sizing,
scoring modes, the MLflow deliverables trap, canary failure modes, and
the
  SIF ↔ Gym-version rebuild coupling.
- `recipes/examples/gym_gdpval/` — self-contained SLURM + vLLM template
plus the
co-located `_gym_prepare.yaml` Hydra include (it must travel with the
config).
- `scripts/gdpval-sif.sh` — build-if-absent / reuse-if-present Apptainer
SIF helper.
Builds on the target cluster only (never copies across clusters),
flock-guarded and
  atomic, driven by `$GDPVAL_SIF_DIR`.
- `SKILL.md` / `references/quantization-benchmarks.md` — GDPVal is part
of the AA
suite but a different harness, so it is generated as a companion
standalone config.
- `recipes/env.example` — `TAVILY_API_KEY` (agent web search) and
`GDPVAL_SIF_DIR`.

### Usage

```bash
# 1. Set GDPVAL_SIF_DIR in .env, then build the sandbox once on the target cluster
#    (build-if-absent, reuse-if-present):
srun -p cpu -t 01:00:00 --pty .agents/scripts/gdpval-sif.sh

# 2. Copy the whole example dir (the _gym_prepare.yaml include must travel with it),
#    fill in the ??? values, then dry-run -> canary -> full:
nel run --config gym_gdpval/example_gym_gdpval.yaml --dry-run
```

### Testing

Validated end-to-end on an aarch64 GB300 SLURM cluster with an NVFP4 MoE
checkpoint:
deploy → SIF build + sandboxed exec → gym head server → 220 rollouts +
deliverables →
judge scoring, producing a real rubric score with zero judge failures.
Several traps
found during that run are now documented in the reference (silent
unsandboxed
fallback, gym-commit/head-server hang, judge api-key value-vs-name, SIF
↔ Gym version
coupling, `limit_samples` not limiting gym rollouts).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (agent skill documentation
+ helper script; no library code)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (agent skill only, no API change)
- Did you get Claude approval on this PR?: ❌ (not yet run)

### Additional Information

Docs/skill-only change under `.agents/`; no `modelopt/` source is
touched.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added standalone GDPVal evaluation support through NeMo Gym, with
Slurm, vLLM, sandbox, judge, reasoning, sampling, and logging
configuration.
- Added automated Apptainer/Singularity image setup with reuse,
validation, locking, and reliable publishing.
- Added configurable settings for judge services, web search, and shared
image caching.
- Added GDPVal preparation and execution examples, including dry-run,
canary, and full-run guidance.

- **Documentation**
- Added GDPVal setup, troubleshooting, scoring, deployment, and
quantization guidance.
- Clarified separate GDPVal configuration, mandatory thinking mode, and
repeat-count requirements.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Chenjie Luo
2026-08-03 15:10:16 -07:00
committed by GitHub
co-authored by Claude Opus 5
parent 38826351eb
commit 14b20c0a12
7 changed files with 973 additions and 11 deletions
+176
View File
@@ -0,0 +1,176 @@
#!/usr/bin/env bash
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# gdpval-sif.sh — ensure the GDPVal Stirrup Apptainer SIF exists on THIS cluster.
#
# Build-if-absent, reuse-if-present. Self-contained: the SIF is built and reused
# on the TARGET cluster's own filesystem — this NEVER copies a SIF from another
# cluster. Idempotent, so it's safe to run before every `nel run`; a subsequent
# run reuses the built SIF instantly.
#
# Usage:
# .agents/scripts/gdpval-sif.sh [<sif-dir-or-file>] [--commit <sha>] [--force|--check]
# <sif-dir-or-file> Persistent path on the target cluster's shared FS.
# DEFAULTS to $GDPVAL_SIF_DIR (from .env) when omitted. A
# directory -> <dir>/$GDPVAL_SIF_NAME (default python-3.13.gdpval.sif,
# matching the example config); a *.sif path
# is used verbatim. Bind-mount this SAME dir into the eval
# container at /gdpval/sif (see recipes/examples/gym_gdpval/).
# --commit <sha> NeMo Gym commit whose gdpval.def to build. Keep in sync
# with the config's install_on_the_fly.commit.
# --force Rebuild even if the SIF already exists.
# --check Verify-only preflight: exit 0 if the expected SIF exists,
# is non-trivial in size, and is a readable SIF. NOTE: it
# inspects the filesystem it RUNS ON — run it on the cluster
# (srun/ssh), not the submitting box, or you validate the
# wrong filesystem.
# Never builds.
# Use before `nel run` — NEL's mount validation is `test -d`
# and cannot see a missing/misnamed SIF file.
#
# Requires `apptainer` (or `singularity`) on PATH with unprivileged/fakeroot
# build support, plus network egress to GitHub/base image. Run on a node that has
# it — a login node, or (preferred for the ~30-min build) the CPU partition:
# srun -p cpu -t 01:00:00 --pty \
# .agents/scripts/gdpval-sif.sh /lustre/<...>/gdpval/sif
#
# Env overrides: GDPVAL_GYM_COMMIT, GDPVAL_SIF_NAME, APPTAINER_BIN.
set -euo pipefail
# Keep GDPVAL_GYM_COMMIT in sync with install_on_the_fly.commit in the config.
GDPVAL_GYM_COMMIT="${GDPVAL_GYM_COMMIT:-dd41196f620f2af99947d776cbe5da9439d2a08d}" # pragma: allowlist secret
GDPVAL_SIF_NAME="${GDPVAL_SIF_NAME:-python-3.13.gdpval.sif}"
APPTAINER_BIN="${APPTAINER_BIN:-}"
_log() { printf '\033[2m %s\033[0m\n' "$*" >&2; }
_die() { printf '\033[31mgdpval-sif: %s\033[0m\n' "$*" >&2; exit 1; }
_usage() { sed -n '/^# gdpval-sif\.sh/,/^set -euo/p' "$0" | sed 's/^# \{0,1\}//; /^set -euo/d'; }
# --- parse args ---
target=""; force=0; check=0
while [[ $# -gt 0 ]]; do
case "$1" in
--commit) GDPVAL_GYM_COMMIT="${2:?--commit needs a value}"; shift 2 ;;
--force) force=1; shift ;;
--check) check=1; shift ;;
-h|--help) _usage; exit 0 ;;
-*) _die "unknown flag: $1 (see --help)" ;;
*) [[ -z "$target" ]] || _die "unexpected extra arg: $1"; target="$1"; shift ;;
esac
done
# Default to $GDPVAL_SIF_DIR (.env) when no path is given, so agents run it hands-free.
target="${target:-${GDPVAL_SIF_DIR:-}}"
[[ -n "$target" ]] || { _usage; _die "no path given and GDPVAL_SIF_DIR is unset — pass a dir or set GDPVAL_SIF_DIR (see recipes/env.example)"; }
# --- resolve dir vs *.sif ---
if [[ "$target" == *.sif ]]; then
sif="$target"; sif_dir="$(dirname "$target")"
else
sif_dir="$target"; sif="$sif_dir/$GDPVAL_SIF_NAME"
fi
# --- verify-only mode (preflight) ---
# NEL's submit-time mount validation runs `test -d`, so it only proves the SIF *dir*
# exists — a dir holding the WRONG sif name (e.g. python-3.12 after a gym bump to a
# 3.13 def) passes validation, and the Stirrup agent then SILENTLY falls back to
# non-sandboxed exec. Run this before submitting to fail loudly instead.
if [[ "$check" -eq 1 ]]; then
if [[ -f "$sif" ]]; then
# -f alone would pass on a truncated or 0-byte file (e.g. an interrupted copy).
# A real GDPVal SIF is ~1-4 GB; anything under 100 MB is not one.
_sz=$(stat -c %s "$sif" 2>/dev/null || echo 0)
if [[ "$_sz" -lt 104857600 ]]; then
printf '\033[31mgdpval-sif: %s exists but is only %s bytes — truncated/incomplete\033[0m\n' "$sif" "$_sz" >&2
echo " Rebuild with: $0 --force ${sif_dir}" >&2
exit 1
fi
if command -v apptainer >/dev/null 2>&1 && ! apptainer inspect "$sif" >/dev/null 2>&1; then
printf '\033[31mgdpval-sif: %s is not a readable SIF (apptainer inspect failed)\033[0m\n' "$sif" >&2
exit 1
fi
_log "SIF present: $sif ($(du -h "$sif" 2>/dev/null | cut -f1))"
echo "$sif"; exit 0
fi
printf '\033[31mgdpval-sif: MISSING expected SIF: %s\033[0m\n' "$sif" >&2
if [[ -d "$sif_dir" ]]; then
echo " dir exists but does not contain it; found:" >&2
if ls -1 "$sif_dir"/*.sif >/dev/null 2>&1; then ls -1 "$sif_dir"/*.sif | sed 's/^/ /' >&2
else echo " (no .sif files)" >&2; fi
fi
echo " Build it with: $0 ${sif_dir} (or --commit <gym-sha> for a different def)" >&2
exit 1
fi
mkdir -p "$sif_dir" || _die "cannot create SIF dir: $sif_dir"
# --- reuse if present ---
if [[ -f "$sif" && "$force" -eq 0 ]]; then
_log "reusing existing SIF (no rebuild): $sif"
echo "$sif"; exit 0
fi
# --- locate apptainer/singularity ---
if [[ -z "$APPTAINER_BIN" ]]; then
APPTAINER_BIN="$(command -v apptainer || command -v singularity || true)"
fi
[[ -n "$APPTAINER_BIN" ]] || _die "apptainer/singularity not found on PATH. Run on a node that has it \
(e.g. 'module load apptainer', or inside the eval image). This script does NOT copy a SIF from another cluster."
def_url="https://raw.githubusercontent.com/NVIDIA-NeMo/Gym/${GDPVAL_GYM_COMMIT}/responses_api_agents/stirrup_agent/containers/gdpval.def"
tmp="${sif_dir}/.build.$$.${GDPVAL_SIF_NAME}"
def_local="${sif_dir}/.gdpval.$$.def"
lock="${sif_dir}/.gdpval-sif.lock"
# --- build under a flock (double-checked) so concurrent runs don't double-build ---
exec 9>"$lock" || _die "cannot open lock file: $lock"
_log "acquiring build lock ($lock) ..."
flock 9
# Re-check inside the lock: another builder may have finished while we waited.
if [[ -f "$sif" && "$force" -eq 0 ]]; then
_log "another builder produced it: $sif"
echo "$sif"; exit 0
fi
_log "building GDPVal SIF (this can take ~20-40 min)"
_log " gym commit: ${GDPVAL_GYM_COMMIT}"
_log " def: ${def_url}"
_log " dest: ${sif}"
# Leave no temp artefacts if we are killed or exit early. $tmp is renamed on success,
# so this only ever removes leftovers.
trap 'rm -f "$tmp" "$def_local"' EXIT
rm -f "$tmp" "$def_local"
# apptainer build cannot take a remote def URL as its source — fetch the def to a
# local file first, then build from it.
if command -v curl >/dev/null 2>&1; then curl -fsSL "$def_url" -o "$def_local"
else wget -qO "$def_local" "$def_url"; fi
[ -s "$def_local" ] || { rm -f "$def_local"; _die "failed to download def from $def_url"; }
# Prefer --fakeroot (needs an /etc/subuid entry for the build user); fall back to an
# unprivileged build where fakeroot is unavailable.
if "$APPTAINER_BIN" build --fakeroot "$tmp" "$def_local"; then
:
# A failed --fakeroot attempt can leave a partial $tmp behind, and apptainer refuses an
# existing destination — clear it or the unprivileged fallback can never succeed.
elif rm -f "$tmp" && "$APPTAINER_BIN" build "$tmp" "$def_local"; then
_log "built without --fakeroot (unprivileged mode)"
else
rm -f "$tmp" "$def_local"
_die "apptainer build failed (see output above)."
fi
rm -f "$def_local"
# Atomic publish: a partial build never looks complete.
mv -f "$tmp" "$sif" || { rm -f "$tmp"; _die "failed to move built SIF into place: $sif"; }
_log "done: $sif"
echo "$sif"
+30 -1
View File
@@ -49,6 +49,34 @@ Steps 1–9 below are the 0.2.6 path — use them for everything else.
---
### GDPVal (NeMo Gym "Stirrup" agent) path — branch here too
GDPVal **does** run on the 0.2.6 `nel` launcher (as a `nemo_gym` task, not
nel-next), so Steps 1–9 apply — but it is mechanically special and **standalone**
(one gym eval per config; never mix it with `aa/` tasks). If the user asks for
GDPVal:
1. Read **`references/gym-gdpval.md`** (Apptainer SIF sandbox, gym prepare/reap
machinery, deploy sizing, rubric-vs-comparison scoring, MLflow deliverables trap,
failure modes) + **`recipes/tasks/aa_gym/gdpval.md`**.
2. Start from **`recipes/examples/gym_gdpval/example_gym_gdpval.yaml`** — a single
self-contained file.
3. Prerequisite — the Apptainer SIF. **If your site provides one, use it**
(NVIDIA-internal: `modelopttools:eval-config` Step 3c); otherwise set
`GDPVAL_SIF_DIR` in `.env` and build with `.agents/scripts/gdpval-sif.sh`
(build-if-absent, no cross-cluster copy). Either way the mounted dir must contain
the file `GDPVAL_CONTAINER_PATH` names (template: `python-3.13.gdpval.sif`) — a
name mismatch passes NEL's `test -d` check and the agent then silently runs
unsandboxed. Verify with `gdpval-sif.sh --check`. `.env` needs `HF_TOKEN`, `INFERENCE_API_KEY`, `TAVILY_API_KEY`,
`INFERENCE_JUDGE_URL`, `GDPVAL_SIF_DIR`, and `NEMO_EVALUATOR_TRUST_PRE_CMD=1` (the
config has a `pre_cmd`). Thinking mode is mandatory (non-thinking loses ~86%).
4. Dry-run → launch. **`limit_samples` is inert on the gym path** (the gym runs all
220 tasks regardless), so there is no cheap canary: watch the real run's first
~20–30 min for the SIF-sandbox line and judge auth, and cancel if wrong. See the
recipe's Canary section.
---
### Step 1 — Prerequisites
Run `nel --version`; if missing, instruct `pip install nemo-evaluator-launcher`. If user has an existing config, skip to Step 8 (optionally review for `???` and quantization flags first).
@@ -62,8 +90,9 @@ Run `nel --version`; if missing, instruct `pip install nemo-evaluator-launcher`.
- AA Index v2 suite (default for quantized-checkpoint validation, see `references/quantization-benchmarks.md`): `recipes/tasks/aa/{gpqa_diamond,hle,lcr,scicode,ifbench,mmmu_pro,tau2_bench_telecom,omniscience}.md`
- Optional: `recipes/tasks/mmlu_pro.md`, `recipes/tasks/aime_2025.md`, `recipes/tasks/livecodebench.md`
- **nel-next only** (different evaluator — see the nel-next section below, NOT the 0.2.6 steps): shared reference `references/nel-next.md` + per-benchmark recipes `recipes/tasks/aa_next/{terminal_bench_2_1,swebench_verified}.md` (agentic). The `aa_next/` dir holds tasks that require nemo-evaluator-next (0.3.x); `aa/` is the 0.2.6 suite.
- **GDPVal (NeMo Gym / agentic)** — **part of the AA suite** but a 0.2.6 `nemo_gym` task on a different harness, so it's **standalone** (see the GDPVal branch above): recipe `recipes/tasks/aa_gym/gdpval.md` + shared reference `references/gym-gdpval.md` + self-contained example `recipes/examples/gym_gdpval/`. Generated as its **own config** from the example, **never merged into the `aa/` multi-task `tasks` list**. The `aa_gym/` dir holds the NeMo Gym Stirrup-agent tasks.
**AA rule:** If the user mentions "AA" / "Artificial Analysis", generate **only** tasks under `recipes/tasks/aa/`. Do not add MMLU-Pro, AIME 2025, or LiveCodeBench unless explicitly asked.
**AA rule:** If the user mentions "AA" / "Artificial Analysis", generate the `recipes/tasks/aa/` tasks (one multi-task config) **plus a companion standalone GDPVal config** (`recipes/tasks/aa_gym/gdpval.md`, via the GDPVal branch) — GDPVal is part of the AA suite but a different harness, so it's its own config, never added to the `aa/` `tasks` list. Do not add MMLU-Pro, AIME 2025, or LiveCodeBench unless explicitly asked. GDPVal is the heaviest AA task (standalone, multi-hour, needs the SIF sandbox + judge) — surface it and let the user opt out per run.
**Shortcut path** (when task list is known up front, e.g. "run AA"):
@@ -37,8 +37,25 @@ NEMO_EVALUATOR_TRUST_PRE_CMD=1
# /v1 base; tau2-bench needs the full /v1/chat/completions.
# HLE + AA-LCR + AA-Omniscience judges (ns_hle_aa, ns_aa_lcr, ns_omniscience) — shared inference host
# GDPVal (nemo_gym) also reuses INFERENCE_JUDGE_URL for its pairwise judge.
# INFERENCE_JUDGE_URL=https://<your-inference-host>/v1
# GDPVal (nemo_gym Stirrup agent) — agent web search. Secret; exported and read
# by the harness. See recipes/tasks/aa_gym/gdpval.md + references/gym-gdpval.md.
# TAVILY_API_KEY=
# GDPVal (nemo_gym) — persistent Apptainer SIF cache dir on the TARGET cluster's
# shared FS (a path, not a secret). .agents/scripts/gdpval-sif.sh builds the SIF
# here if absent and reuses it otherwise; the config bind-mounts this dir at
# /gdpval/sif. Convention: a per-user .cache dir.
# GDPVAL_SIF_DIR=<shared-fs>/<user>/.cache/gdpval/sif
# GDPVal (nemo_gym) — Stirrup agent turn cap. Read at SUBMIT time from the
# launching shell (the config uses ${oc.env:GDPVAL_MAX_TURNS,250}), so it must be
# exported before `nel run`; setting it as a container env var has no effect.
# Default 250 (the golden value); lower it only to shorten a debugging run.
# GDPVAL_MAX_TURNS=250
# Tau2 (tau2_bench_telecom) — judger + user-simulator model_ids are hardcoded in
# the recipe; only the shared endpoint URL comes from here
# TAU2_ENDPOINT_URL=https://<your-inference-host>/v1/chat/completions # user + judger
@@ -0,0 +1,360 @@
# SPDX-FileCopyrightText: Copyright (c) 2025, NVIDIA CORPORATION. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
#
# =============================================================================
# GDPVal (NeMo Gym "Stirrup" agent) — STANDALONE gym eval, RUBRIC mode.
# Self-deploys a checkpoint via vLLM on one SLURM node. One gym eval per config.
#
# Read recipes/tasks/aa_gym/gdpval.md + references/gym-gdpval.md first — they
# cover the SIF sandbox, scoring modes, judge panel, preflight and failure modes.
#
# Before running: `.env` needs HF_TOKEN, INFERENCE_API_KEY, TAVILY_API_KEY,
# INFERENCE_JUDGE_URL, GDPVAL_SIF_DIR, NEMO_EVALUATOR_TRUST_PRE_CMD=1; and the
# SIF must exist (prefer a site-provided one, else
# `srun -p cpu -t 01:00:00 --pty .agents/scripts/gdpval-sif.sh`).
#
# nel run --config example_gym_gdpval.yaml --env-file .env
#
# `limit_samples` is INERT here — the gym always runs all 220 tasks. There is no
# cheap smoke test; see the recipe's Canary section.
# =============================================================================
defaults:
# slurm/default works anywhere; if your install ships a predefined
# internal/slurm/<cluster> config, prefer it (pre-fills hostname/partition/
# gres — see SKILL.md Step 4).
- execution: slurm/default
- deployment: vllm
- _self_
# GDPVal scoring mode. This template is RUBRIC-ONLY, deliberately: rubric needs no
# reference deliverables and runs on the public gym image, so it works standalone.
# rubric — standalone LLM-judge scoring, 0-1 reward per deliverable. (this template)
# comparison — pairwise vs anchored reference deliverables; the only mode that yields
# an AA-comparable ELO / win-rate. It needs a reference set, a newer gym
# image, and its own overrides — do NOT just flip reward_mode here.
# To run comparison, NVIDIA-internal users should follow `modelopttools:eval-config`
# Step 3c, which converts this config (container override + reference_models map +
# mounts + multistage). See references/gym-gdpval.md "Scoring modes".
gdpval:
reward_mode: rubric
# GDPVal pairwise judge. base_url is config (from .env), not a secret, so no
# export needed; only api_key (INFERENCE_API_KEY) is exported and read by the
# harness. Keep the judge fixed across comparable runs.
gdpval_judge:
base_url: <INFERENCE_JUDGE_URL> # from .env (/v1 base); shared inference host
model: gcp/google/gemini-3.1-pro-preview # Gemini 3.1 Pro; use an equivalent on your endpoint if needed
api_key: INFERENCE_API_KEY # doc only; common_params injects the VALUE ($INFERENCE_API_KEY).
# An env-var NAME reaches the proxy as a literal -> opaque judge 500.
cluster:
sbatch_comment: '{"OccupiedIdleGPUsJobReaper":{"exemptIdleTimeMins":"480","reason":"benchmarking","description":"Eval benchmark low GPU utilization"}}'
execution:
hostname: ???
username: ${oc.env:USER}
account: ???
output_dir: ??? # absolute host path; keeps deliverables + response cache across resumes
walltime: "04:00:00"
# gres: a predefined internal/slurm/<cluster> config sets this. On slurm/default
# it's gpu:8 — set to the node's GPU count (match --tensor/--data-parallel-size)
# or sbatch fails "Requested node configuration is not available".
mounts:
# mount_home ALWAYS false — see example_eval.yaml / SKILL Step 4 for why.
mount_home: false
deployment:
# Real HF cache -> /hf-cache (paired with HF_HOME below).
<HOST_HF_CACHE>: /hf-cache
evaluation:
<HOST_HF_CACHE>: /hf-cache
<HOST_UV_CACHE>: /cache/uv
# SIF dir -> /gdpval/sif, so the file lands at GDPVAL_CONTAINER_PATH below.
# Literal path only: mount KEYS are never interpolated.
<GDPVAL_SIF_DIR>: /gdpval/sif
# Writable shared-FS staging for ref files. Node-local /tmp breaks multi-node Ray.
<HOST_REF_FILES>: /gdpval_ref_files
# Comparison mode additionally mounts one dir per reference model at
# /gdpval/refs/<key> (both `deployment` and `evaluation`) — added by
# modelopttools:eval-config Step 3c, not here.
auto_export: # REQUIRED trigger for MLflow upload (see example_eval.yaml).
destinations:
- mlflow
# Auto-export is a separate CPU-only sbatch; GPU-only partitions reject it.
cpu_partition: ???
deployment:
env_vars:
HF_TOKEN: host:HF_TOKEN
HF_HOME: lit:/hf-cache
# vLLM backend toggles go HERE (not in command). Uncomment per model card —
# e.g. NVFP4 MoE on Blackwell needs FlashInfer FP4 kernels:
# VLLM_USE_FLASHINFER_MOE_FP4: lit:1
# VLLM_FLASHINFER_MOE_BACKEND: lit:throughput
checkpoint_path: ??? # prefer a path already on the cluster over hf_model_handle
hf_model_handle:
served_model_name: ???
image: vllm/vllm-openai:v0.19.1 # bump to the EXACT model's recipes.vllm.ai minimum (SKILL Step 3)
# GDPVal REQUIRES thinking mode from the policy (non-thinking loses ~86% of
# pairwise judgements). Serve with the model's --reasoning-parser so vLLM emits
# a separate reasoning channel; thinking is forced on via the adapter_config
# chat_template_kwargs below. Add --enable-expert-parallel for MoE, and
# --trust-remote-code for custom-code models.
# --max-num-seqs: on the GYM path derive it from the agent's concurrency, NOT from
# `parallelism` (which is gym-internal, not an in-flight request count — the generic
# SKILL rule would give an absurd cap). N = ceil(stirrup_agent.concurrency / DP).
command: >-
vllm serve /checkpoint
--served-model-name ${deployment.served_model_name}
--host 0.0.0.0
--port ${deployment.port}
--tensor-parallel-size 1
--data-parallel-size 1
--max-model-len 131072
--reasoning-parser <PARSER>
--model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 128}'
--max-num-batched-tokens 8192
--enable-chunked-prefill
evaluation:
env_vars:
HF_TOKEN: host:HF_TOKEN
HF_HOME: lit:/hf-cache
DUMMY_API_KEY: lit:dummy # the deployed vLLM endpoint's (policy) key
INFERENCE_API_KEY: host:INFERENCE_API_KEY # GDPVal judge auth (shared inference host)
TAVILY_API_KEY: host:TAVILY_API_KEY # Stirrup agent web search
UV_CACHE_DIR: lit:/cache/uv
# Apptainer SIF for the Stirrup per-task code-exec sandbox (must match the mount above).
GDPVAL_CONTAINER_PATH: lit:/gdpval/sif/python-3.13.gdpval.sif
# Shared-FS staging for ref files; node-local /tmp breaks multi-node Ray.
GDPVAL_REF_FILES_DIR: lit:/gdpval_ref_files
# Deliverables land under /results (already mounted) so they persist. The
# "*cache*" basename matches the mlflow exporter's exclusion, so the (large)
# deliverables are NOT auto-uploaded — they stay on disk for inspection.
# Drop "_cache" if you WANT them uploaded as artifacts.
PERSIST_DELIVERABLES_DIR: lit:/results/gdpval/deliverables_cache
NEL_INVOCATION_ID: runtime:NEL_INVOCATION_ID
# Installs apptainer + squashfuse into the eval container (needs
# NEMO_EVALUATOR_TRUST_PRE_CMD=1 in the launching shell). See references/gym-gdpval.md.
# Installs the apptainer RUNTIME into the eval container (it is not baked in) so the
# Stirrup agent can exec the prebuilt SIF. ARCH-AWARE: Blackwell/Grace clusters are
# aarch64, so never hardcode an amd64 .deb — the Ubuntu PPA builds both arches.
pre_cmd: |
set -ex
export DEBIAN_FRONTEND=noninteractive
apt-get update -qq
apt-get install -y -qq software-properties-common squashfuse fuse3 ca-certificates || true
add-apt-repository -y ppa:apptainer/ppa
apt-get update -qq
# Unpinned on purpose: the PPA is the only source with current arm64 builds. The
# resolved version is echoed below for the record.
apt-get install -y -qq apptainer
apptainer --version
mkdir -p /usr/local/var/apptainer/mnt/session
nemo_evaluator_config:
config:
params:
temperature: 1.0
top_p: 0.95
parallelism: 16384 # gym-internal concurrency, NOT a model-server cap
request_timeout: 36000 # gym rollouts are long-running
max_retries: 10
target:
api_endpoint:
api_key_name: DUMMY_API_KEY
adapter_config:
use_system_prompt: false
process_reasoning_traces: true
params_to_remove:
- max_tokens
- max_completion_tokens
# Thinking mode MUST be on — non-thinking loses ~86% of pairwise
# judgements. Keep the toggle key your model family uses (enable_thinking
# -> Qwen3.x/GLM; thinking -> Kimi/DeepSeek); see SKILL "Reasoning adapter config".
params_to_add:
chat_template_kwargs:
enable_thinking: true
skip_special_tokens: false
use_caching: false
use_progress_tracking: true
tracking_requests_stats: true
log_failed_requests: true
use_request_logging: true
max_logged_requests: 10
use_response_logging: true
max_logged_responses: 10
# STANDALONE: exactly one gym task. Do NOT add other tasks to this list.
tasks:
- name: nemo_gym
# Public image — fine for RUBRIC mode (this template's default).
# COMPARISON mode needs a NEWER Gym than this image ships and must OVERRIDE this
# line; see references/gym-gdpval.md "Scoring modes". NOTE: this image bakes Gym
# as a non-git directory, so `install_on_the_fly.commit` below is silently
# ignored here (the prepare step logs "/opt/Gym is not a git repo; using baked-in
# Gym version") — the pin only applies on images where /opt/Gym IS a git repo.
container: nvcr.io/nvidia/eval-factory/nemo-gym:26.05 # pin a verified tag
nemo_evaluator_config:
config:
params:
extra:
nemo_gym:
install_on_the_fly:
url: https://github.com/NVIDIA-NeMo/Gym
# Gym version. BUMPING THIS REQUIRES REBUILDING THE SIF from the matching
# commit — gdpval.def (the sandbox) is versioned with the gym repo, and an
# old SIF + new gym silently degrades deliverables. Rebuild with
# `gdpval-sif.sh --commit <sha>` to a new name + repoint GDPVAL_CONTAINER_PATH.
# See references/gym-gdpval.md "Rebuild the SIF when the Gym version changes".
# Current golden pin (updated GDPVal task-sampling algo, which drives
# multistage stage-1 selection). Its gdpval.def is byte-identical to
# 049b1fd0…, so a python-3.13 SIF built from either commit is valid —
# no rebuild when moving between them.
# NOTE: this pin is INERT on the public nemo-gym image (see the
# `container:` comment above) — it only takes effect on an image where
# /opt/Gym is a git repo. Confirm via "=== NeMo Gym commit ===" + SHA
# in the client log before crediting it with any behaviour change.
commit: dd41196f620f2af99947d776cbe5da9439d2a08d # pragma: allowlist secret
command: |
set -ex
cd /opt/Gym
# Honour a mounted uv cache if injected, else image-local. No ${VAR} braces.
[ -n "$UV_CACHE_DIR" ] || export UV_CACHE_DIR=/opt/cache/uv
source .venv/bin/activate
# install_on_the_fly: checkout the pin when /opt/Gym is a git repo; some
# images bake Gym at a fixed version (not a git repo) — use it as-is.
if [ -d .git ]; then
git remote add oss_pin "{{config.params.extra.nemo_gym.install_on_the_fly.url}}" 2>/dev/null || true
git fetch oss_pin
git checkout "{{config.params.extra.nemo_gym.install_on_the_fly.commit}}"
echo "=== NeMo Gym commit ===" && git rev-parse HEAD
else
echo "=== /opt/Gym is not a git repo; using baked-in Gym version ==="
fi
for r in /opt/Gym/responses_api_models/*/requirements.txt \
/opt/Gym/responses_api_agents/*/requirements.txt \
/opt/Gym/resources_servers/*/requirements.txt; do
[ -f "$r" ] || continue
grep -vE '^[[:space:]]*-e ' "$r" > "$r.fixed" || true
grep -qiE '^ray([<>=[]|$)' "$r.fixed" 2>/dev/null || echo 'ray==2.49.2' >> "$r.fixed"
grep -qiE '^tqdm' "$r.fixed" 2>/dev/null || echo 'tqdm' >> "$r.fixed"
mv "$r.fixed" "$r" 2>/dev/null || true
done
for v in /opt/Gym/responses_api_models/*/.venv \
/opt/Gym/responses_api_agents/*/.venv \
/opt/Gym/resources_servers/*/.venv; do
[ -d "$v" ] || continue
d="$(dirname "$v")"
[ -f "$d/requirements.txt" ] && uv pip install --python "$v/bin/python" -q -r "$d/requirements.txt" || true
done
# Preserve inherited PYTHONPATH. Plain $VAR only (OmegaConf parses ${...}).
_gym_sp="$(/opt/Gym/.venv/bin/python -c 'import site; print(site.getsitepackages()[0])')"
if [ -n "$PYTHONPATH" ]; then export PYTHONPATH="/opt/Gym:$_gym_sp:$PYTHONPATH"
else export PYTHONPATH="/opt/Gym:$_gym_sp"; fi
# Writable staging dir for ref files (bind-mounted, see execution).
mkdir -p /gdpval_ref_files
# SECRETS: `set -x` traces AFTER expansion, so without this the params'
# $HF_TOKEN / $INFERENCE_API_KEY / $TAVILY_API_KEY land in the log (and in
# MLflow via log_logs). The rollout call below is already safe via the heredoc.
set +x
ng_prepare_benchmark {{config.params.extra.nemo_gym.data_prep_params}} {{config.params.extra.nemo_gym.common_params}}
set -x
# The rollout command is written to a script and run from it, rather than passed
# to `bash -c '...'`. Gym params legitimately CONTAIN single quotes — comparison
# mode's ++multistage.stages='[{num_tasks: 45}, ...]' and ++...judge_panel='[{...}]'
# — and those would close a single-quoted `bash -c` wrapper early, so Hydra then
# receives the value split on spaces and dies with
# "no viable alternative at input '[{num_tasks:'".
# The heredoc delimiter is QUOTED ('GYM_RUN_EOF') so nothing expands while writing:
# $$ and $TAVILY_API_KEY / $INFERENCE_API_KEY land in the script literally and are
# expanded at run time by the shell that executes it, which is what we want.
cat > /tmp/gym_run.sh <<'GYM_RUN_EOF'
echo $$ > /tmp/gym_eval_pgid
exec ng_e2e_collect_rollouts {{config.params.extra.nemo_gym.collect_rollout_params}} {{config.params.extra.nemo_gym.common_params}}
GYM_RUN_EOF
setsid --wait bash /tmp/gym_run.sh &
__ev=$!
__rc=0; wait "$__ev" || __rc=$?
echo "Evaluator Gym finished!"
__pg=$(cat /tmp/gym_eval_pgid 2>/dev/null || echo)
[ -n "$__pg" ] && kill -9 -"$__pg" 2>/dev/null || true
timeout 60 ray stop --force >/dev/null 2>&1 || true
# uid-scoped: an unscoped pkill would kill other jobs' Ray daemons.
pkill -9 -u "$(id -u)" -f raylet >/dev/null 2>&1 || true
pkill -9 -u "$(id -u)" -f gcs_server >/dev/null 2>&1 || true
pkill -9 -u "$(id -u)" -f plasma_store >/dev/null 2>&1 || true
exit $__rc
data_prep_params: >-
"+config_paths=[responses_api_models/vllm_model/configs/vllm_model.yaml,benchmarks/gdpval/config.yaml]"
+hf_token=$HF_TOKEN
++use_cached_prepared_benchmarks=true
collect_rollout_params: >-
"+config_paths=[responses_api_models/vllm_model/configs/vllm_model.yaml,benchmarks/gdpval/config.yaml]"
++port_range_low=63000
++port_range_high=64000
++global_aiohttp_connector_limit_per_host={{config.params.parallelism}}
++uv_cache_dir=$UV_CACHE_DIR
++skip_venv_if_present=true
++output_jsonl_fpath={{config.output_dir}}/evaluator_rollouts.jsonl
++nemo_gym_log_dir={{config.output_dir}}/nemo_gym_logs
++overwrite_metrics_conflicts=true
++split=benchmark
++resume_from_cache=true
++reuse_existing_data_preparation=true
++upload_rollouts_to_wandb=false
++responses_create_params.temperature={{config.params.temperature}}
++responses_create_params.top_p={{config.params.top_p}}
# No judge_responses_create_params_overrides.model here: the judge model is
# already set by openai_model=${gdpval_judge.model}, and pinning it a second
# way silently collapses comparison mode's 3-member judge panel to one judge.
# num_repeats=1 — both current goldens. No reference_* overrides here:
# the single-reference keys conflict with comparison mode's reference_models
# map (added by modelopttools:eval-config Step 3c).
# NEVER put '#' inside this folded (>-) scalar — it folds to one line and
# comments out every override after it.
common_params: >-
++use_absolute_ip=true
++policy_base_url={{target.api_endpoint.url}}
++policy_api_key=$DUMMY_API_KEY
++policy_model_name={{target.api_endpoint.model_id}}
++gdpval_judge_model.responses_api_models.openai_model.openai_base_url=${gdpval_judge.base_url}
++gdpval_judge_model.responses_api_models.openai_model.openai_model=${gdpval_judge.model}
++gdpval_judge_model.responses_api_models.openai_model.openai_api_key=$INFERENCE_API_KEY
++gdpval_judge_model.responses_api_models.openai_model.max_concurrent_requests=10
++gdpval_stirrup_agent.responses_api_agents.stirrup_agent.concurrency=220
++gdpval_stirrup_agent.responses_api_agents.stirrup_agent.tavily_api_key=$TAVILY_API_KEY
++gdpval_stirrup_agent.responses_api_agents.stirrup_agent.agent_max_turns=${oc.env:GDPVAL_MAX_TURNS,250}
++num_repeats=1
++gdpval_resources_server.resources_servers.gdpval.reward_mode=${gdpval.reward_mode}
++gdpval_resources_server.resources_servers.gdpval.persist_raw_judge_responses=true
++gdpval_resources_server.resources_servers.gdpval.preconvert_max_concurrent=30
++gdpval_resources_server.resources_servers.gdpval.preconvert_office_to_pdf=true
export:
# LITERAL values only (auto_export resolves this block at submit time in a scope
# without deployment/evaluation nodes; ${oc.env:...} is fine). Keep the sampling
# tags EQUAL to evaluation params above — they're the only MLflow record of them.
mlflow:
tracking_uri: ${oc.env:MLFLOW_TRACKING_URI} # from modelopttools:eval-config
experiment_name: ${oc.env:USER}/CHANGEME-served-model-name
description: 'CHANGEME-served-model-name | GDPVal rubric | T=1.0, top_p=0.95, num_repeats=1'
log_logs: true
only_required: false
tags:
framework: vllm
model: CHANGEME-served-model-name
benchmark: nemo_gym.gdpval
temperature: '1.0'
top_p: '0.95'
@@ -0,0 +1,129 @@
# GDPVal (NeMo Gym "Stirrup" agent)
## Task Details
- Reference: `references/gym-gdpval.md` (SIF build, gym machinery, deploy sizing,
scoring modes, failure modes) — **read it before editing a GDPVal config.**
- Upstream README:
<https://github.com/NVIDIA-NeMo/Evaluator/blob/main/examples/nemotron/nemotron-3-ultra/v0.2/README.md>
GDPVal is an **agentic** benchmark: the Stirrup agent produces office/PDF
deliverables inside a per-task Apptainer code-exec sandbox, then a pairwise/rubric
judge (**Gemini 3.1 Pro**) scores them. It is the most resource-intensive benchmark
in the suite — **220 tasks**, `num_repeats=1`, 4 judge trials per rollout.
It runs on the **0.2.6 `nel` launcher** as a `nemo_gym` task (NOT nel-next), so
Steps 1–9 apply — but with the branch differences below.
## What makes GDPVal different (not a normal `aa/` task)
- **Standalone** — one gym eval per config. Never add GDPVal to a multi-task
`evaluation.tasks` list, and never add other tasks to a GDPVal config.
- **Apptainer SIF sandbox** — prefer a site-provided SIF; otherwise
`.agents/scripts/gdpval-sif.sh` builds one into `$GDPVAL_SIF_DIR` (build-if-absent,
never copied between clusters). Missing/misnamed → **silent** unsandboxed exec.
- **Thinking mode is mandatory** — non-thinking loses ~86% of pairwise judgements.
Serve with the model's `--reasoning-parser` and force it on via the adapter's
`chat_template_kwargs`.
- **Scoring:** `rubric` (template default, no references, no ELO) vs `comparison`
(the AA-comparable `normalized_elo`; a conversion, not a flag flip).
- Needs `INFERENCE_API_KEY`, `TAVILY_API_KEY`, `INFERENCE_JUDGE_URL`,
`GDPVAL_SIF_DIR` in `.env`, plus `NEMO_EVALUATOR_TRUST_PRE_CMD=1` (the config has a
`pre_cmd`).
All of the above — SIF handling, the SIF↔Gym-commit coupling, scoring modes, judge
panel, preflight and failure modes — is detailed in **`references/gym-gdpval.md`**.
Read it before editing a GDPVal config.
## Config
Start from the self-contained example and edit it — **do not** copy a fragment into
another config:
```text
recipes/examples/gym_gdpval/example_gym_gdpval.yaml # SLURM + single-node vLLM,
# rubric mode, self-contained
```
`num_repeats=1` — already set by the template via `++num_repeats=1`; both
current goldens use it. A full 220-task run of a large MoE typically needs multi-node.
## Canary — `limit_samples` does NOT work here
**`++…params.limit_samples=N` is inert on the gym path.** The gym does its own data
prep and rollout collection, so the launcher-level limiter is ignored: you get the
full 220-task run. Do not use it believing you launched a two-task smoke test — this
is the heaviest benchmark in the suite.
There is no cheap sample-limited canary. Instead, **launch the real run and treat its
first ~20–30 minutes as the canary**, cancelling if any of these is wrong:
```bash
RD=<output_dir>/<run>/nemo_gym.0
grep -c "Using Apptainer container" $RD/logs/client-*.log # sandbox actually used
grep -c "falling back\|not a git repo" $RD/logs/client-*.log # unsandboxed / inert pin
grep -ciE " 401 | 403 |Internal Server Error" $RD/artifacts/nemo_gym_logs/gdpval_judge_model.log
wc -l $RD/artifacts/evaluator_rollouts.jsonl # rollouts flowing
```
In comparison mode stage 1 (45 tasks) is a natural early checkpoint — an ELO estimate
appears before the full 220-task stage 2 starts.
## Score Extraction
> **The GDPVal score is NOT in `artifacts/eval_factory_metrics.json`.** That file
> holds only `response_stats` / `reasoning` / `evaluation` (request-level telemetry).
> Looking there and finding no ELO does not mean the run failed to score.
**The reported GDPVal score is `normalized_elo`** — the AA 0–1 scale, comparable
across models and to the published AA index. `eval_elo` is the same fit on the raw
Elo axis (`normalized_elo = (eval_elo - 500) / 2000`); quote it as supporting
detail, not as the score.
The final numbers live in **`artifacts/results.yml`** (authoritative, local) and are
mirrored to MLflow. Read them by metric name:
| Mode | Metric (results.yml → `groups.nemo_gym.metrics.<name>.scores.<name>.value`) |
| --- | --- |
| comparison | `gdpval_stirrup_agent/comparison/normalized_elo` ← **REPORT THIS** (AA 0–1 scale) |
| comparison | `gdpval_stirrup_agent/comparison/eval_elo` (raw Elo; supporting detail) |
| comparison | `gdpval_stirrup_agent/comparison/win_rate`, `/judged`, `/wins`, `/losses`, `/ties` |
| comparison | per-reference: `gdpval_stirrup_agent/comparison/ref/<ref_key>/{win_rate,wins,losses,ties,judged}` |
| comparison | per-stage estimate: `gdpval_stirrup_agent/comparison/stage_0/eval_elo` (stage 1, all refs) — the **final** value is the top-level one, from the last stage |
| rubric | mean of `reward` across `artifacts/evaluator_rollouts.jsonl` (per-rollout 0–1) |
```bash
# COMPARISON mode — final score from the local results file (no MLflow needed)
python3 -c "
import yaml
m=yaml.safe_load(open('<output_dir>/<run>/nemo_gym.0/artifacts/results.yml'))['groups']['nemo_gym']['metrics']
for k in ('normalized_elo','eval_elo','win_rate'):
n=f'gdpval_stirrup_agent/comparison/{k}'
print(k, '=', m[n]['scores'][n]['value'])"
# RUBRIC mode (the template default) — there is no ELO; average the per-rollout reward
python3 -c "
import json
r=[json.loads(l).get('reward') for l in open('<output_dir>/<run>/nemo_gym.0/artifacts/evaluator_rollouts.jsonl')]
r=[x for x in r if isinstance(x,(int,float))]
print('mean reward =', sum(r)/len(r), 'over', len(r), 'rollouts')"
```
In **MLflow** the same values are prefixed `nemo_gym_` and duplicated under a
`key_metrics/` path — query these exact keys rather than browsing the UI, because a
comparison run logs **~200 metrics and most of them are per-reference**, so the
headline is easy to miss:
```text
nemo_gym_gdpval_stirrup_agent/key_metrics/comparison/normalized_elo <- report this
nemo_gym_gdpval_stirrup_agent/key_metrics/comparison/eval_elo
nemo_gym_gdpval_stirrup_agent/key_metrics/comparison/win_rate
```
Sanity checks before quoting a score: `…/comparison/judged` should be large (a few
hundred+), `num_stages`/`num_references` should match your multistage config, and the
unique `task_id` count in `evaluator_rollouts.jsonl` should be close to 220 — a short
count means tasks were lost (e.g. across a walltime resume) and the ELO is computed on
fewer tasks than the references were. Per-task detail is in
`evaluator_rollouts.jsonl` + `nemo_gym_logs/`; raw judge responses are under
`PERSIST_DELIVERABLES_DIR`.
@@ -0,0 +1,236 @@
# GDPVal (NeMo Gym "Stirrup" agent) — reference for the gym / agentic path
GDPVal runs on the **0.2.6 `nel` launcher** as a `nemo_gym` task, but it is
mechanically unlike the `aa/` nemo-skills tasks: the Stirrup agent produces
office/PDF **deliverables** in a per-task **Apptainer** code-exec sandbox, a
pairwise/rubric **judge** (Gemini 3.1 Pro) scores them, and NeMo Gym is pulled and
run **inline in the eval container** (`install_on_the_fly`) via `ng_prepare_benchmark`
+ `ng_e2e_collect_rollouts`. This file is the shared machinery; the config template
is `recipes/examples/gym_gdpval/` and the per-task pointer is
`recipes/tasks/aa_gym/gdpval.md`.
## Where each piece runs
| Component | Where |
|---|---|
| Policy model (under test) | your self-deployed vLLM endpoint (SLURM GPU node) — or an external endpoint |
| NeMo Gym + Stirrup agent orchestration | inside the **eval** container (`nemo_gym` task), pulled via `install_on_the_fly` |
| Per-task code-exec | **Apptainer SIF** launched by the agent inside the eval container |
| Judge (pairwise/rubric) | external OpenAI-compatible endpoint (`gdpval_judge`, e.g. Gemini 3.1 Pro) |
| Agent web search | Tavily (`TAVILY_API_KEY`) |
## Apptainer SIF sandbox
The Stirrup agent runs each task's generated code in an Apptainer SIF, bind-mounted
at **exactly** the path `GDPVAL_CONTAINER_PATH` names (template:
`/gdpval/sif/python-3.13.gdpval.sif`). Missing or misnamed → the agent **silently**
runs code-exec unsandboxed; the run "succeeds" but the numbers aren't comparable.
**If your site provides a prebuilt SIF, use it** — a self-built one resolves its pip
stack at *your* build time and can drift from the sandbox a published reference set
was generated in. Mount the provided dir at `/gdpval/sif` and point
`GDPVAL_CONTAINER_PATH` at its filename. (NVIDIA-internal: `modelopttools:eval-config`
Step 3c has the path.) Otherwise build it on the target cluster — never copy a SIF
between clusters:
```bash
srun -p cpu -t 01:00:00 --pty .agents/scripts/gdpval-sif.sh # uses $GDPVAL_SIF_DIR
```
`gdpval-sif.sh` is idempotent (flock-guarded, atomic): it builds from `gdpval.def` at
the pinned commit if absent and is a no-op once present. It needs
apptainer/singularity with unprivileged-build support and network egress — run it on a
login or CPU node, **not** inside the eval job. The eval image doesn't ship apptainer,
so the config's `pre_cmd` installs the **runtime** (arch-aware: use the Ubuntu PPA, not
an amd64 `.deb` — most Blackwell/Grace clusters are aarch64), which needs
`NEMO_EVALUATOR_TRUST_PRE_CMD=1`.
**The SIF is versioned with the Gym repo.** `gdpval.def` changes across commits (e.g.
`2502893977` → `049b1fd0` moved python 3.12 → 3.13 and added TeX Live, playwright,
polars, geospatial), and the newer agent's prompt advertises that richer runtime. So
when you bump `install_on_the_fly.commit`, **diff the def at the two commits**
(`raw.githubusercontent.com/NVIDIA-NeMo/Gym/<sha>/responses_api_agents/stirrup_agent/containers/gdpval.def`);
if it changed, rebuild to a **new version-tagged filename** (`GDPVAL_SIF_NAME=… gdpval-sif.sh --commit <sha>`)
and repoint `GDPVAL_CONTAINER_PATH`. Running a new gym on an old SIF makes the
generated code fail its imports *inside the sandbox* — deliverables silently degrade
with no error in the eval.
**Exception — a site-provided SIF paired with a site-provided gym image.** Those
images typically bake Gym as a non-git dir, so `install_on_the_fly.commit` is inert
and the two provided artifacts are already matched to each other; that pairing is the
coherent one even when the SIF filename encodes a different SHA than your pin. You
cannot check the image's baked Gym from the config — confirm from the client log
(`=== NeMo Gym commit ===` + a SHA, or the "not a git repo" line).
## Gym prepare / reap (why the task `command:` is long)
The task `command:` carries two workaround blocks, inlined in the template:
1. **prepare** — activate the baked Gym venv, checkout the `install_on_the_fly` pin
(only if `/opt/Gym` is a git repo), repair the image's incomplete per-server venvs
(drop the editable `-e nemo-gym[dev]` line, which forces a ray re-resolve that
breaks venv-less servers; pin `ray==2.49.2` + `tqdm`), and front the main venv on
`PYTHONPATH`.
2. **run** — data prep, then `ng_e2e_collect_rollouts` executed from a script written
via a **quoted heredoc** and launched under `setsid`, so the whole server/Ray
process tree can be reaped by process group. Without that reap, orphaned Ray
workers hold the launcher's stdout open and the run **hangs in post-eval**; the
quoted heredoc keeps `$$` and `$*_API_KEY` unexpanded until run time and survives
params that contain single quotes (comparison mode's `stages='[{...}]'`).
Both compensate for the eval image's deployment-oriented packaging and Gym's
incomplete shutdown — remove them once the image ships complete ray-consistent venvs
and Gym reaps its own process groups. Avoid bash `${VAR}` inside these blocks:
OmegaConf parses `${...}`. `$(...)`, `$$` and `$VAR` are fine.
## Deployment sizing
GDPVal is heavy: 220 tasks × `num_repeats` rollouts, each a long multi-turn agent
episode with code-exec + judge calls (`request_timeout: 36000`). The example
self-deploys single-node vLLM, which is fine for a canary or a small policy. For the
**full run of a large MoE** (e.g. MiniMax-M2.7), the reviewed golden uses **multi-node
`vllm_ray`** (16 × 4-GPU HSG = 64 GPUs, `walltime 04:00:00`). To scale up:
+ Switch `defaults: - deployment: vllm_ray` and add nodes (`execution.num_nodes`);
see `references/multi-node.md` for the Ray TP/PP layout.
+ `parallelism` (`16384`) is **gym-internal concurrency**, not a server cap. The
real throttles are the agent's `stirrup_agent.concurrency` and the judge's
`max_concurrent_requests` — raise those only after the judge logs are clean of 429s.
+ **`--max-num-seqs`: derive it from `stirrup_agent.concurrency`, NOT `parallelism`.**
SKILL Step 3/4's `ceil(parallelism / DP)` rule assumes `parallelism` is the in-flight
request count; on the gym path it is not, and applying it literally gives an absurd
cap. Use `ceil(stirrup_agent.concurrency / DP)` — e.g. 220 / DP 4 → 55, round to 64.
+ **`max_new_tokens`:** the reviewed golden **does** set it alongside the adapter's
`params_to_remove: [max_tokens, max_completion_tokens]`, so do the Step 3 model-card
lookup as normal. The template omits it (five params) because the adapter strips the
per-request cap anyway; adding it back matches the golden and is harmless.
+ **Match `temperature` / `top_p` to whatever the reference deliverables were generated
with.** A pairwise ELO compares your deliverables against theirs, so a sampling
difference lands in the score as if it were a quality difference.
+ Long runs exceed 4h; rely on NEL's walltime dependency-chain resume
(`resume_from_cache=true` is already set). See SKILL Step 4 + `run-validation.md`.
## Scoring modes — rubric vs comparison
+ **`rubric`** (template default) — judge scores each deliverable against its rubric.
0–1 reward, **no ELO** (undefined without an opponent). Runs on the public gym image.
+ **`comparison`** — pairwise vs anchored reference deliverables; the **only** mode
yielding the AA-comparable `normalized_elo`. It is a conversion, not a flag flip:
it needs a reference set, a gym image whose Gym has the `reference_models` map, ref
mounts on **both** deployment and evaluation, and multistage overrides. Setting
`reward_mode=comparison` alone exits at startup with
`reward_mode=comparison requires reference_deliverables_dir to be set`, surfaced
only as `Process gdpval_resources_server finished unexpectedly!`.
NVIDIA-internal: `modelopttools:eval-config` Step 3c is the conversion checklist.
## Judge
Rubric mode uses a single judge. **Comparison mode uses a 3-member panel** —
`openai/gpt-5.5`, `gcp/google/gemini-3.1-pro-preview`,
`aws/anthropic/bedrock-claude-opus-4-8` — one **sampled per trial**, all routed
through the single `gdpval_judge_model` proxy (`<INFERENCE_JUDGE_URL>` from `.env`).
`++...judge_sampling_seed=42` makes that sampling reproducible.
+ **Inject the key's VALUE, not its name:** `openai_api_key=$INFERENCE_API_KEY`.
Passing an env-var *name* (e.g. via a `${...api_key}` interpolation that resolves to
the literal string `INFERENCE_API_KEY`) makes the proxy reply `LiteLLM Virtual Key
expected`, which the gym wraps as an opaque **500** — it looks like a judge outage,
not a config error. One key covers all three panel members.
+ **Do not set `judge_responses_create_params_overrides.model`.** Pinning a model
collapses the panel to a single judge, silently changing the scoring methodology.
+ **Throttles:** judge `max_concurrent_requests=10` and Stirrup `concurrency=220` are
the golden values — the judge rate-limits long before the served model does, so raise
these only after the judge logs are clean of 429s.
## Preflight — what NEL validates, and what it does NOT
NEL validates mount paths at **submit** time (`_collect_mount_paths` +
`_validate_remote_paths_exist`): it ssh's to the cluster, runs `test -d` on every
mount source, and `raise ValueError` listing the missing ones **before** any
`sbatch` — so a missing reference dir or cache costs you nothing. Three gaps to know:
| Artifact | Missing → | Loud? |
|---|---|---|
| mounted dirs (refs, caches, SIF **dir**, checkpoint) | `ValueError` at submit, no job queued | ✅ pre-allocation |
| **the SIF file inside that dir** | **agent silently runs code-exec unsandboxed** | ❌ **silent** |
| task `container:` (image / `.sqsh`) | not collected for validation → pyxis import failure | ⚠️ only after allocation |
1. **`test -d` proves the directory, not the SIF.** A `$GDPVAL_SIF_DIR` that exists but
holds the *wrong* filename (e.g. `python-3.12…` after bumping to a 3.13 def) passes
validation, and the run then silently degrades. Guard with the verify-only mode:
```bash
.agents/scripts/gdpval-sif.sh --check # uses $GDPVAL_SIF_DIR; exit 1 + lists what IS there
```
Keep `GDPVAL_SIF_NAME` / the helper's default in sync with the config's
`GDPVAL_CONTAINER_PATH`; they are the same string in two places.
2. **`--dry-run` skips remote validation entirely** (it never opens the ssh
connection). A clean dry-run says nothing about whether your mounts exist — run
the preflight separately.
3. **The container is never checked.** A wrong/rotated image path fails at pyxis
import, i.e. after the allocation is granted. Verify it with `ls -l` first
(comparison mode's internal image especially — see `modelopttools:eval-config`).
## Env vars
| Var | Prefix | Purpose |
|---|---|---|
| `HF_TOKEN` | host | model/dataset downloads |
| `INFERENCE_API_KEY` | host | **judge** auth (and policy if external) |
| `TAVILY_API_KEY` | host | Stirrup agent web search |
| `DUMMY_API_KEY` | lit:dummy | self-deployed vLLM policy key |
| `GDPVAL_CONTAINER_PATH` | lit | SIF path — must equal the SIF bind-mount target |
| `GDPVAL_REF_FILES_DIR` | lit:/gdpval_ref_files | shared-FS ref-file staging (node-local /tmp breaks multi-node Ray) |
| `PERSIST_DELIVERABLES_DIR` | lit | where deliverables persist (see MLflow note) |
| `GDPVAL_MAX_TURNS` | lit (optional) | Stirrup turn cap (default 100; golden uses 250) |
| `NEL_INVOCATION_ID` | runtime | run id |
`INFERENCE_JUDGE_URL` is the judge host — config (from `.env`), substituted as the
literal `<INFERENCE_JUDGE_URL>` placeholder in `gdpval_judge.base_url`, **not**
`${oc.env:...}`. Judge `model_id` is hardcoded in the config (swap for an equivalent
on your endpoint). The upstream OSS recipe uses a separate `GDPVAL_JUDGE_API_KEY`;
this template reuses the shared `INFERENCE_API_KEY` for the judge.
`GDPVAL_SIF_DIR` (`.env`) is host-side config, **not** a container env var: it's the
persistent SIF cache dir the helper builds into and the config bind-mounts at
`/gdpval/sif`. `gdpval-sif.sh` reads `$GDPVAL_SIF_DIR` directly (its default target);
the config mount is a **literal `<GDPVAL_SIF_DIR>` placeholder** you substitute with
that same path — mount KEYS aren't interpolated, so don't emit `${oc.env:...}` there
(same rule as the judge URLs). One `.env` value feeds both, so the build path and the
run path can't drift.
## MLflow export — the deliverables trap
Deliverables can be large. The mlflow exporter excludes any artifact dir whose
basename matches `*cache*`, so the template sets
`PERSIST_DELIVERABLES_DIR=/results/gdpval/deliverables_cache`: the deliverables stay
under `/results` for inspection but are **not** auto-uploaded. Drop the `_cache`
suffix only if you actually want them uploaded. Everything else about auto-export is
standard (SKILL Step 1 shortcut #4): `auto_export.destinations: [mlflow]` +
`cpu_partition` + a literal-valued `export.mlflow` block (tag `benchmark:
nemo_gym.gdpval`).
## num_repeats
**Use 1.** Both current goldens do, set with a top-level `++num_repeats=1` — it
works, and recent Gym pins already ship `num_repeats: 1` in
`benchmarks/gdpval/config.yaml`, so no `sed` patching is needed.
Historical only: pre-multistage single-reference configs used 2 (220 × 2 = 440
rollouts) and patched it with `sed` because the per-dataset key could not be set
via `++` on those pins. Do not carry a `=2` into a current run.
## Failure modes to check at canary
+ **Silent unsandboxed exec** — grep the eval log for the SIF fallback warning /
apptainer mount errors; confirm `GDPVAL_CONTAINER_PATH` == the mount target.
+ **Judge 401 / 429** — wrong `INFERENCE_JUDGE_URL` / key, or `max_concurrent_requests`
too high for the judge endpoint.
+ **Empty reasoning / low win-rate** — thinking mode off. Confirm
`chat_template_kwargs.enable_thinking: true` (right toggle key for the family) +
the policy's `--reasoning-parser`.
+ **Run hangs in post-eval** — orphaned Ray/gym processes holding stdout; that's what
the setsid + process-group reap in the task `command:` prevents.
+ **Multi-node ref-file errors** — `GDPVAL_REF_FILES_DIR` on node-local storage;
point it at a shared-FS staging dir.
@@ -3,15 +3,20 @@
When evaluating a quantized checkpoint, prioritize benchmarks that are sensitive
to precision loss. The Artificial Analysis (AA) Index v2 suite under
`recipes/tasks/aa/` is the default set for quantized-checkpoint validation.
**GDPVal** (`recipes/tasks/aa_gym/gdpval.md`) is also part of the AA suite, but a
different harness (NeMo Gym) — it runs as a **separate standalone config**, never
merged into the `aa/` multi-task list.
**Scope rule:**
- **Default quant validation** (when the user just says "evaluate this
quantized checkpoint"): use the AA suite plus the three always-include
benchmarks at `recipes/tasks/*.md` (MMLU-Pro, AIME 2025, LiveCodeBench).
quantized checkpoint"): use the AA suite — the `aa/` tasks **plus a standalone
GDPVal config** — plus the three always-include benchmarks at
`recipes/tasks/*.md` (MMLU-Pro, AIME 2025, LiveCodeBench).
- **Explicit AA request** ("AA" / "Artificial Analysis" / "AA Index v2"):
use **only** `recipes/tasks/aa/`. Do not add the three always-include
tasks unless the user asks. See the callout at the bottom of this file.
use the `aa/` tasks **and** a companion standalone GDPVal config. Do not add
the three always-include tasks unless the user asks. See the callout at the
bottom of this file.
## Available task recipes
@@ -28,6 +33,7 @@ to precision loss. The Artificial Analysis (AA) Index v2 suite under
| `tasks/aa/mmmu_pro.md` | MMMU-Pro | Multimodal reasoning | VLM-only; usually Low/Medium when only the LLM is quantized (vision encoder/adapter typically stay BF16) |
| `tasks/aa/tau2_bench_telecom.md` | Tau2-Bench Telecom | Agentic tool use (user-simulator + judge) | Medium-high — tool-call JSON is brittle, but user-sim + judge variance often dominates the signal |
| `tasks/aa/omniscience.md` | AA-Omniscience | Knowledge reliability (`ns_omniscience`, nemo-skills, `num_repeats: 10`) — correct vs hallucinate vs abstain on obscure facts, judge-scored | Medium — measures the hallucination/abstention balance; aggressive precision loss can erode factual recall and shift the omni-index |
| `tasks/aa_gym/gdpval.md` | GDPVal (`nemo_gym` Stirrup agent, **standalone config**) | Agentic office/PDF deliverables in an Apptainer code-exec sandbox, pairwise/rubric judge | High — long-horizon agentic reasoning + code + judge; precision loss compounds across many turns. **Heaviest task**: multi-hour, often multi-node, needs the SIF sandbox + judge. Runs as its own config, never in the `aa/` list |
## Recommended sets by use case
@@ -35,15 +41,17 @@ to precision loss. The Artificial Analysis (AA) Index v2 suite under
|----------|-----------|
| Quick sanity check | GPQA |
| Standard quant validation (text LLM) | GPQA, SciCode, LCR |
| AA / Artificial Analysis suite (text LLM) | All `tasks/aa/` text tasks: GPQA, HLE, LCR, SciCode, IFBench, Tau2-Bench Telecom, AA-Omniscience |
| AA / Artificial Analysis suite (multimodal) | AA text suite + MMMU-Pro |
| AA / Artificial Analysis suite (text LLM) | All `tasks/aa/` text tasks: GPQA, HLE, LCR, SciCode, IFBench, Tau2-Bench Telecom, AA-Omniscience — **plus GDPVal** (`tasks/aa_gym/`, a separate standalone config) |
| AA / Artificial Analysis suite (multimodal) | AA text suite (incl. GDPVal) + MMMU-Pro |
| Code-focused model | LiveCodeBench, SciCode |
| Reasoning model | AIME 2025, GPQA, HLE |
> If the user asks for "AA" or "Artificial Analysis", generate **only** tasks
> under `recipes/tasks/aa/`. Do not silently add MMLU-Pro, AIME 2025, or
> LiveCodeBench — they live at `recipes/tasks/*.md` and are a separate
> always-include set.
> If the user asks for "AA" or "Artificial Analysis", generate the
> `recipes/tasks/aa/` tasks **plus a companion standalone GDPVal config**
> (`recipes/tasks/aa_gym/gdpval.md`) — GDPVal is part of the AA suite but a
> different harness, so it's its own config, never in the `aa/` `tasks` list. Do
> not silently add MMLU-Pro, AIME 2025, or LiveCodeBench — they live at
> `recipes/tasks/*.md` and are a separate always-include set.
## Notes for quantized-checkpoint runs
@@ -62,6 +70,13 @@ to precision loss. The Artificial Analysis (AA) Index v2 suite under
apples-to-apples comparison.
- **IFBench** is the least quant-sensitive in the set but still useful as a
regression check for aggressive formats (NVFP4, INT4-AWQ).
- **GDPVal** is part of the AA suite but the heaviest task and a separate
harness: it runs as its **own standalone `aa_gym` config** (never in the `aa/`
`tasks` list), needs the Apptainer SIF sandbox + judge, and is multi-hour /
often multi-node. Generate it alongside the `aa/` config; see
`recipes/tasks/aa_gym/gdpval.md` + `references/gym-gdpval.md`. Thinking mode is
mandatory (non-thinking loses ~86% of pairwise judgements). `num_repeats` is **1** —
the value both current goldens use, already set by the template; do not raise it.
## How to use