mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Add AA-Omniscience eval recipe; harden judge/run conventions in the eval skill (#1834)
### What does this PR do? Type of change: documentation (agent `evaluation` skill) Updates the `evaluation` agent skill (`.agents/skills/evaluation/`) with an AA-Omniscience recipe plus several judge/run hardening conventions found while running the AA Index v2 suite on quantized checkpoints. - **AA-Omniscience recipe** (`recipes/tasks/aa/omniscience.md`): nemo-skills `ns_omniscience`, AA Index v2 params (`++parse_reasoning=False`, `num_repeats=10`), `gcp/google/gemini-3-flash-preview` judge; score `omniscience_pass_at_1_avg-of-N_judge_correct`. Added to the AA Index v2 suite list in `SKILL.md`. - **Judge model_ids hardcoded** in the HLE / AA-LCR / Tau2 recipes (with a "swap for an equivalent on your own endpoint" note); shared judge URL var renamed `NS_JUDGE_URL` -> `INFERENCE_JUDGE_URL`; judge API key folded into `INFERENCE_API_KEY` in `env.example`. - **Idle-reaper exemption** in `example_eval.yaml`: `cluster.sbatch_comment` exempts eval jobs from the `OccupiedIdleGPUsJobReaper`, which otherwise CANCELs jobs whose GPUs sit idle during model load / judge calls / aggregation. - **Resume-on-kill note** in `SKILL.md`: after a preemption / idle-reaper CANCEL (not a walltime timeout, which NEL auto-resumes), re-submit the job's `run.sub` to resume from the response cache (`skip_filled`) with cumulative progress. ### Usage \`\`\`bash nel run --config recipes/examples/example_eval.yaml # now ships the idle-reaper exemption # extend evaluation.tasks with the AA-Omniscience fragment from recipes/tasks/aa/omniscience.md \`\`\` ### Testing Validated end-to-end on gcp-nrt (B200): MiniMax-M2.7-NVFP4 AA-Omniscience full run (600 questions x 10 repeats) completed with \`omniscience_pass_at_1_avg-of-10_judge_correct = 18.77\`. The idle-reaper exemption and the \`sbatch run.sub\` resume path were both exercised - a reaped run resumed from cache and finished all 10 repeats. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: yes (skill docs/recipes only; no library API change) - If you copied code from any other sources or added a new PIP dependency: N/A - Did you write any new necessary tests?: N/A (agent skill recipes/docs) - Did you update Changelog?: N/A (agent skill, not a library feature) - Did you get Claude approval on this PR?: pending (/claude review) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an **AA-Omniscience** evaluation recipe. * Expanded the default **AA Index v2** quantized-checkpoint validation suite to include Omniscience. * **Documentation** * Clarified which judge/user-simulator **model identifiers** are fixed in recipes vs supplied via environment variables. * Updated HLE, LCR, and Tau2-Bench Telecom guidance for judge configuration and API-key/endpoint usage. * Added instructions for resuming evaluations after scheduler preemption/cancellation. * **Chores** * Updated the example evaluation config to include a GPU job reaper policy. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
This commit is contained in:
@@ -42,7 +42,7 @@ Run `nel --version`; if missing, instruct `pip install nemo-evaluator-launcher`.
|
||||
|
||||
**Task recipes** (always read before editing the relevant task in the config):
|
||||
|
||||
- AA Index v2 suite (default for quantized-checkpoint validation, see `references/quantization-benchmarks.md`): `recipes/tasks/aa/{gpqa_diamond,hle,lcr,scicode,ifbench,mmmu_pro,tau2_bench_telecom}.md`
|
||||
- AA Index v2 suite (default for quantized-checkpoint validation, see `references/quantization-benchmarks.md`): `recipes/tasks/aa/{gpqa_diamond,hle,lcr,scicode,ifbench,mmmu_pro,tau2_bench_telecom,omniscience}.md`
|
||||
- Optional: `recipes/tasks/mmlu_pro.md`, `recipes/tasks/aime_2025.md`, `recipes/tasks/livecodebench.md`
|
||||
|
||||
**AA rule:** If the user mentions "AA" / "Artificial Analysis", generate **only** tasks under `recipes/tasks/aa/`. Do not add MMLU-Pro, AIME 2025, or LiveCodeBench unless explicitly asked.
|
||||
@@ -240,6 +240,14 @@ On SLURM, several deploy/eval failures are invisible to `--dry-run` and only sur
|
||||
|
||||
Evals that exceed 4h of wall-clock time are handled by **NEL's built-in dependency-chain resume**, not by shrinking the eval. NEL submits the first SLURM job; if it hits walltime, a dependent follow-on job resumes from the response/result caches the first job wrote, then queues another follow-on. Long evals continue across walltime windows automatically. See `references/run-validation.md#nel-timeout-and-resume-behavior` for the full mechanism.
|
||||
|
||||
**Preemption / external kill — resume manually with `sbatch run.sub`.** On a preemptible account (common on busy internal clusters) the scheduler can **CANCEL** a run mid-eval for a higher-priority job — `sacct -j <id>` shows `CANCELLED by <uid>` (a `svc-*` service account) with `Elapsed` well under the 4h walltime. NEL does **not** auto-resume this (its dependency chain only fires on a genuine walltime timeout). But the `run.sub` that NEL generated for the job (in its run dir) is **re-submittable** and resumes from the same `output_dir` + response cache (`skip_filled`), continuing from the partial output rather than restarting:
|
||||
|
||||
```bash
|
||||
ssh <host> "cd <output_dir>/<timestamp>-<invocation>/<task>/ && sbatch run.sub"
|
||||
```
|
||||
|
||||
Re-submit again if it's preempted again — each resume re-deploys, then skips already-generated samples, so progress is **cumulative** across attempts until it completes. Always confirm via `sacct -j <id>` that the prior job was `CANCELLED` (not a real failure) before resuming.
|
||||
|
||||
Implications for the agent:
|
||||
|
||||
- Do **not** lower `num_repeats`, split heavy tasks (AA-LCR, SciCode) into separate configs, or otherwise carve up the eval to fit inside 4h. Let NEL chain.
|
||||
|
||||
@@ -33,27 +33,18 @@ NEMO_EVALUATOR_TRUST_PRE_CMD=1
|
||||
# JUDGE_API_KEY=
|
||||
# INFERENCE_API_KEY=
|
||||
|
||||
# --- Optional: judge / user-simulator endpoints (model_id + URL) ---
|
||||
# --- Optional: judge / user-simulator endpoints (URL only) ---
|
||||
#
|
||||
# External judge / user-simulator / scoring endpoints, for any task that needs one
|
||||
# (HLE, AA-LCR, Tau2 below — add more for other such benchmarks; auth via
|
||||
# INFERENCE_API_KEY above). These are config, not secrets: the values you set here are
|
||||
# substituted as literal model_id/url into the config (matching <VAR> placeholders in
|
||||
# the recipes) — they do NOT need to be exported; only INFERENCE_API_KEY is.
|
||||
# URL note: nemo-skills uses the /v1 base; tau2-bench needs the full /v1/chat/completions.
|
||||
# The `modelopttools:eval-config` skill fills the model_id/url values below (and
|
||||
# installs the optional launcher package) — invoke it for judge-scored runs;
|
||||
# otherwise point them at your own OpenAI-compatible judge host.
|
||||
# Judge / user-simulator `model_id`s are hardcoded in each task recipe (swap them
|
||||
# there if you serve an equivalent). Only the endpoint URLs come from here — config,
|
||||
# not secrets, so no export needed (only INFERENCE_API_KEY is). nemo-skills uses the
|
||||
# /v1 base; tau2-bench needs the full /v1/chat/completions.
|
||||
|
||||
# HLE judge (ns_hle_aa) — recommended GPT-4o
|
||||
# HLE_JUDGE_MODEL_ID=<judge-model-id>
|
||||
# AA-LCR judge (ns_aa_lcr) — recommended Qwen3 235B
|
||||
# LCR_JUDGE_MODEL_ID=<judge-model-id>
|
||||
# NS_JUDGE_URL=https://<your-inference-host>/v1 # shared by both judges above
|
||||
# HLE + AA-LCR + AA-Omniscience judges (ns_hle_aa, ns_aa_lcr, ns_omniscience) — shared inference host
|
||||
# INFERENCE_JUDGE_URL=https://<your-inference-host>/v1
|
||||
|
||||
# Tau2 (tau2_bench_telecom) — user-sim Qwen3 235B, judger gpt-oss-120B
|
||||
# TAU2_USER_MODEL_ID=<user-simulator-model-id>
|
||||
# TAU2_JUDGER_MODEL_ID=<judger-model-id>
|
||||
# Tau2 (tau2_bench_telecom) — judger + user-simulator model_ids are hardcoded in
|
||||
# the recipe; only the shared endpoint URL comes from here
|
||||
# TAU2_ENDPOINT_URL=https://<your-inference-host>/v1/chat/completions # user + judger
|
||||
|
||||
# terminal-bench-hard (AWS sandbox)
|
||||
|
||||
@@ -46,6 +46,8 @@ defaults:
|
||||
- execution: slurm/default
|
||||
- deployment: vllm
|
||||
- _self_
|
||||
cluster:
|
||||
sbatch_comment: '{"OccupiedIdleGPUsJobReaper":{"exemptIdleTimeMins":"480","reason":"benchmarking","description":"Eval benchmark low GPU utilization"}}'
|
||||
execution:
|
||||
hostname: ???
|
||||
username: ${oc.env:USER}
|
||||
|
||||
@@ -6,12 +6,12 @@
|
||||
|
||||
## Params
|
||||
|
||||
Text-only HLE, params aligned to Artificial Analysis Index v2; judge-scored.
|
||||
Substitute the judge `model_id`/`url` with the literal values you keep in `.env`
|
||||
(`HLE_JUDGE_MODEL_ID` rec. **GPT-4o**, `NS_JUDGE_URL`; see `recipes/env.example`) —
|
||||
they're config, not secrets, so they don't need exporting. Only `api_key`
|
||||
(`INFERENCE_API_KEY`) is exported and read by the harness. Keep the judge fixed
|
||||
across comparable runs.
|
||||
Text-only HLE, params aligned to Artificial Analysis Index v2; judge-scored. The
|
||||
judge `model_id` is hardcoded in the fragment below (**GPT-4o**) — swap it for an
|
||||
equivalent on your own endpoint if needed. The judge `url` still comes from `.env`
|
||||
(`INFERENCE_JUDGE_URL`; see `recipes/env.example`) — config, not a secret, so no export.
|
||||
Only `api_key` (`INFERENCE_API_KEY`) is exported and read by the harness. Keep the
|
||||
judge fixed across comparable runs.
|
||||
|
||||
`hle_strict_judge: true` (inside the `judge` block) enables strict judging.
|
||||
|
||||
@@ -35,8 +35,8 @@ Use this inside the top-level `evaluation.tasks` list:
|
||||
params:
|
||||
extra:
|
||||
judge:
|
||||
model_id: <HLE_JUDGE_MODEL_ID> # from .env; recommended GPT-4o
|
||||
url: <NS_JUDGE_URL> # from .env (/v1 base)
|
||||
model_id: azure/openai/gpt-4o # GPT-4o; use an equivalent on your own endpoint if needed
|
||||
url: <INFERENCE_JUDGE_URL> # from .env (/v1 base)
|
||||
api_key: INFERENCE_API_KEY # env-var name; exported, read by harness
|
||||
hle_strict_judge: true
|
||||
```
|
||||
|
||||
@@ -6,10 +6,11 @@
|
||||
|
||||
## Params
|
||||
|
||||
Judge-scored (equality checker). Substitute the judge `model_id`/`url` with the
|
||||
literal values you keep in `.env` (`LCR_JUDGE_MODEL_ID` rec. **Qwen3 235B**,
|
||||
`NS_JUDGE_URL`; see `recipes/env.example`) — config, not secrets, so no export
|
||||
needed; only `api_key` (`INFERENCE_API_KEY`) is exported. Keep the judge fixed.
|
||||
Judge-scored (equality checker). The judge `model_id` is hardcoded in the fragment
|
||||
below (**Qwen3 235B**) — swap it for an equivalent on your own endpoint if needed.
|
||||
The judge `url` still comes from `.env` (`INFERENCE_JUDGE_URL`; see `recipes/env.example`)
|
||||
— config, not a secret, so no export; only `api_key` (`INFERENCE_API_KEY`) is
|
||||
exported. Keep the judge fixed.
|
||||
|
||||
AA-LCR needs long context: plan for roughly 120K input tokens plus 16K
|
||||
generation tokens. Set deployment `--max-model-len` to at least `131072`, and
|
||||
@@ -55,8 +56,8 @@ block. Per SKILL.md Step 3, the deployment flag must live inside
|
||||
extra:
|
||||
num_repeats: 16
|
||||
judge:
|
||||
model_id: <LCR_JUDGE_MODEL_ID> # from .env; recommended Qwen3 235B
|
||||
url: <NS_JUDGE_URL> # from .env (/v1 base)
|
||||
model_id: nvidia/qwen/qwen-235b # Qwen3 235B; use an equivalent on your own endpoint if needed
|
||||
url: <INFERENCE_JUDGE_URL> # from .env (/v1 base)
|
||||
api_key: INFERENCE_API_KEY # env-var name; exported, read by harness
|
||||
```
|
||||
|
||||
|
||||
@@ -0,0 +1,49 @@
|
||||
# AA-Omniscience
|
||||
|
||||
## Task Details
|
||||
|
||||
- Reference: <https://docs.nvidia.com/nemo/evaluator/latest/evaluation/benchmarks/catalog/all/harnesses/nemo_skills.html#nemo-skills-ns-omniscience>
|
||||
|
||||
## Params
|
||||
|
||||
Knowledge / hallucination benchmark, params aligned to Artificial Analysis Index
|
||||
v2; judge-scored. The judge `model_id` is hardcoded in the fragment below
|
||||
(**gcp/google/gemini-3-flash-preview**) — swap it for an equivalent on your own
|
||||
endpoint if needed. The judge `url` comes from `.env` (`INFERENCE_JUDGE_URL`); only
|
||||
`api_key` (`INFERENCE_API_KEY`) is exported and read by the harness. Keep the judge
|
||||
fixed across comparable runs.
|
||||
|
||||
`++parse_reasoning=False` is required (golden knob — omniscience scores the final
|
||||
answer, not the reasoning trace).
|
||||
|
||||
## YAML Fragment
|
||||
|
||||
Use this inside the top-level `evaluation.tasks` list:
|
||||
|
||||
```yaml
|
||||
- name: nemo_skills.ns_omniscience
|
||||
container: nvcr.io/nvidia/eval-factory/nemo-skills:26.05.1
|
||||
env_vars:
|
||||
INFERENCE_API_KEY: host:INFERENCE_API_KEY
|
||||
nemo_evaluator_config:
|
||||
config:
|
||||
params:
|
||||
extra:
|
||||
num_repeats: 10
|
||||
args: "++parse_reasoning=False"
|
||||
judge:
|
||||
api_key: INFERENCE_API_KEY
|
||||
model_id: gcp/google/gemini-3-flash-preview # use an equivalent on your own endpoint if needed
|
||||
url: <INFERENCE_JUDGE_URL> # from .env (/v1 base)
|
||||
```
|
||||
|
||||
## Score Extraction from mlflow
|
||||
|
||||
**Primary result — Omniscience Index** (-100 to 100): `omniscience_pass_at_1_avg-of-N_judge_omni_index`. This is AA's headline metric: accuracy net of hallucinations, rewarding abstention over guessing wrong (so it can be negative). Report this one. See the AA methodology: <https://artificialanalysis.ai/methodology/intelligence-benchmarking#aa-omniscience>.
|
||||
|
||||
Also report (same `pass_at_1_avg-of-N` aggregation):
|
||||
|
||||
- **Accuracy** (0-100): `omniscience_pass_at_1_avg-of-N_judge_correct` — % of questions answered correctly.
|
||||
- **Non-hallucination rate** (0-100): `100 - omniscience_pass_at_1_avg-of-N_judge_omni_hallucination` — the `judge_omni_hallucination` key is the hallucination rate, so non-hallucination = `1 - hallucination`.
|
||||
|
||||
N is the repeat count (10). If the repeat count is unknown, use the highest available `avg-of-N`.
|
||||
@@ -7,12 +7,13 @@
|
||||
## Params
|
||||
|
||||
Tau2 uses the evaluated model as the agent plus a separate user-simulator endpoint;
|
||||
keep both fixed across runs. Substitute the user-sim & judger `model_id`/`url` with the
|
||||
literal values you keep in `.env` (`TAU2_USER_MODEL_ID` rec. **Qwen3 235B**,
|
||||
`TAU2_JUDGER_MODEL_ID` rec. **gpt-oss-120B**, `TAU2_ENDPOINT_URL`; see
|
||||
`recipes/env.example`) — config, not secrets, so no export needed; only `api_key`
|
||||
(`INFERENCE_API_KEY`) is exported. tau2-bench needs the full `/v1/chat/completions`
|
||||
URL (nemo-skills judges use the `/v1` base).
|
||||
keep both fixed across runs. The judger (**gpt-oss-120B**) and user-simulator
|
||||
(**Qwen3 235B**) `model_id`s are hardcoded in the fragment below — swap them for
|
||||
equivalents on your own endpoint if needed. Only the shared `url`
|
||||
(`TAU2_ENDPOINT_URL`) comes from `.env` (see `recipes/env.example`) — config, not a
|
||||
secret, so no export needed; only `api_key` (`INFERENCE_API_KEY`) is exported.
|
||||
tau2-bench needs the full `/v1/chat/completions` URL (nemo-skills judges use the
|
||||
`/v1` base).
|
||||
|
||||
For parallelism, we have to throttle to a smaller cap due to the test may be throttled by
|
||||
user and judger API rate limit. If frequent 429 errors are hit, the reported scores could be much lower.
|
||||
@@ -56,11 +57,11 @@ Use this inside the top-level `evaluation.tasks` list:
|
||||
skip_failed_samples: true
|
||||
n_samples: 8
|
||||
user:
|
||||
model_id: <TAU2_USER_MODEL_ID> # from .env; recommended Qwen3 235B
|
||||
model_id: nvidia/qwen/qwen-235b # Qwen3 235B; use an equivalent on your own endpoint if needed
|
||||
url: <TAU2_ENDPOINT_URL> # from .env (full /v1/chat/completions)
|
||||
api_key: INFERENCE_API_KEY # env-var name; exported, read by harness
|
||||
judger:
|
||||
model_id: <TAU2_JUDGER_MODEL_ID> # from .env; recommended gpt-oss-120B
|
||||
model_id: nvidia/openai/gpt-oss-120b # gpt-oss-120B; use an equivalent on your own endpoint if needed
|
||||
url: <TAU2_ENDPOINT_URL> # from .env (full /v1/chat/completions)
|
||||
api_key: INFERENCE_API_KEY # env-var name; exported, read by harness
|
||||
```
|
||||
|
||||
@@ -27,6 +27,7 @@ to precision loss. The Artificial Analysis (AA) Index v2 suite under
|
||||
| `tasks/aa/ifbench.md` | IFBench | Instruction following | Low — format-compliance is robust; even aggressive FP4 usually shows only small drops |
|
||||
| `tasks/aa/mmmu_pro.md` | MMMU-Pro | Multimodal reasoning | VLM-only; usually Low/Medium when only the LLM is quantized (vision encoder/adapter typically stay BF16) |
|
||||
| `tasks/aa/tau2_bench_telecom.md` | Tau2-Bench Telecom | Agentic tool use (user-simulator + judge) | Medium-high — tool-call JSON is brittle, but user-sim + judge variance often dominates the signal |
|
||||
| `tasks/aa/omniscience.md` | AA-Omniscience | Knowledge reliability (`ns_omniscience`, nemo-skills, `num_repeats: 10`) — correct vs hallucinate vs abstain on obscure facts, judge-scored | Medium — measures the hallucination/abstention balance; aggressive precision loss can erode factual recall and shift the omni-index |
|
||||
|
||||
## Recommended sets by use case
|
||||
|
||||
@@ -34,7 +35,7 @@ to precision loss. The Artificial Analysis (AA) Index v2 suite under
|
||||
|----------|-----------|
|
||||
| Quick sanity check | GPQA |
|
||||
| Standard quant validation (text LLM) | GPQA, SciCode, LCR |
|
||||
| AA / Artificial Analysis suite (text LLM) | All `tasks/aa/` text tasks: GPQA, HLE, LCR, SciCode, IFBench, Tau2-Bench Telecom |
|
||||
| AA / Artificial Analysis suite (text LLM) | All `tasks/aa/` text tasks: GPQA, HLE, LCR, SciCode, IFBench, Tau2-Bench Telecom, AA-Omniscience |
|
||||
| AA / Artificial Analysis suite (multimodal) | AA text suite + MMMU-Pro |
|
||||
| Code-focused model | LiveCodeBench, SciCode |
|
||||
| Reasoning model | AIME 2025, GPQA, HLE |
|
||||
@@ -53,10 +54,12 @@ to precision loss. The Artificial Analysis (AA) Index v2 suite under
|
||||
do **not** lower them for quant comparisons, or noise will mask real
|
||||
regressions. The field name differs by harness: `n_samples` for simple-evals
|
||||
(AIME `64`) and tau2-bench (Tau2 `8`); `num_repeats` for nemo-skills
|
||||
(AA-LCR/GPQA `16`, LiveCodeBench/SciCode `8`, IFBench `5`, MMLU-Pro `1`).
|
||||
- **Judge / user-simulator endpoints** are required by AA-LCR, HLE AA, and
|
||||
Tau2-Bench Telecom. Keep the judge and (for Tau2) user-simulator models
|
||||
fixed across baseline and quantized runs for apples-to-apples comparison.
|
||||
(AA-LCR/GPQA `16`, AA-Omniscience `10`, LiveCodeBench/SciCode `8`, IFBench `5`,
|
||||
MMLU-Pro `1`).
|
||||
- **Judge / user-simulator endpoints** are required by AA-LCR, HLE AA,
|
||||
AA-Omniscience, and Tau2-Bench Telecom. Keep the judge and (for Tau2)
|
||||
user-simulator models fixed across baseline and quantized runs for
|
||||
apples-to-apples comparison.
|
||||
- **IFBench** is the least quant-sensitive in the set but still useful as a
|
||||
regression check for aggressive formats (NVFP4, INT4-AWQ).
|
||||
|
||||
|
||||
Reference in New Issue
Block a user