Files
Model-Optimizer/plugins/modelopt
Zhiyu f2ee751089 Validate eval limit accounting and raise TB 2.1 timeout (#2499)
### What does this PR do?

Type of change: documentation

Successful evaluations can still contain timed-out trials or responses
stopped by output limits. Existing skills scan for errors but do not
require counts or rates, allowing serving-speed effects to be mistaken
for quantization accuracy changes.

Require per-task timeout and output-limit accounting before scores are
treated as validated, with explicit denominators, telemetry coverage,
effective limits, and artifact evidence. Separate request retries,
terminal trial failures, and resumable SLURM walltime events. Carry that
evidence into baseline/candidate comparisons; unknown accounting or
unresolved infrastructure effects prevent an acceptable verdict.
Preserve benchmark-defined failures and label diagnostic protocol
changes explicitly. Allow valid-with-warnings results for small
fractions of limit-hit responses/trials when coverage, protocol, and
other checks pass, without automatically retrying. Clarify that a BF16
acceptance gate cannot be satisfied with an FP8/INT4 baseline.

Correct Terminal-Bench 2.1 guidance that sharding cannot affect scores
and require accounting across the full trial set. Raise the ModelOpt TB
2.1 agent budget from 7200 to 14400 seconds in the recipe and set it
explicitly in the template. With timeout_strategy=max, this is a minimum
budget, not a hard ceiling. Set sandbox lifetime to six hours and
example SLURM walltime to eight hours to allow setup and verification;
partition limits and effective task budgets must still be checked.
Baseline and candidate must use the same policy, and old two-hour
results require remeasurement for a matched comparison. This changes
skill defaults, not the upstream benchmark protocol or harness
instrumentation. The upstream-vendored launching-evals skill is
unchanged per repository policy; its standalone workflow still needs an
upstream update.

Reduce the evaluation entrypoint from 6,156 to 888 words (86%) by moving
launcher Steps 1–8 into an on-demand reference without changing their
instructions. Preserve step headings/links for existing callers. Shorten
timeout accounting from 550 to 352 words (36%) while retaining the
checks and warning policy. This reduces initial context; full launcher
workflows still load the relevant detailed sections.

### Usage

TB 2.1 recipe/template defaults now include:

```yaml
solver:
  timeout_strategy: max
  run_timeout: 14400
sandbox:
  max_task_lifetime_sec: 21600
```

The example uses `cluster.walltime: "08:00:00"` where the partition
permits it. Request timeout remains 3600 seconds. For leaderboard
comparisons, use the benchmark protocol rather than assuming the
ModelOpt override is comparable.

### Testing

- Pre-commit passed on all six changed Markdown/YAML files.
- skill-creator quick_validate.py passed for evaluation and
compare-results.
- git diff --check passed.
- Validated recipe and template solver/sandbox settings against NEL
schemas with the pinned TB playbook; checked max and task timeout
resolution.
- Verified extracted launcher Steps 1–8 match the original text, apart
from a trailing blank line.
- Manually reviewed handling of recovered retries, resumed artifacts,
incomplete telemetry, benchmark-defined limits, and timeout-sensitive
comparisons. No evaluation jobs were launched; increased timeout
defaults have not been benchmarked.

### Before your PR is "*Ready for review*"

Contributor guidelines and security coding practices reviewed.

- Is this change backward compatible?: ✅ Existing configs remain valid.
New TB configs use longer budgets and may consume more runtime;
comparison requires matched timeout policies.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in CONTRIBUTING.md: N/A
- Did you write any new necessary tests?: N/A — skill/config changes
validated as listed above.
- Did you update Changelog?: N/A — internal skill guidance.
- Did you get Claude approval on this PR?: ❌ Not requested yet.

### Additional Information

A nonzero benchmark-defined timeout or output-cap rate does not
automatically invalidate a score. The change requires evidence and
protocol-aware interpretation, without choosing a universal acceptable
rate or silently excluding affected trials.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Expanded evaluation guidance with a centralized launcher workflow
covering setup, configuration, execution, monitoring, authentication,
and failure handling.
* Required timeout and output-limit accounting for every task, including
successful and non-reasoning runs, with rates, denominators, telemetry
coverage, effective limits, and recovery status.
* Clarified that unknown telemetry, mismatched limits, or unresolved
infrastructure effects prevent an acceptable verdict.
* Updated comparison guidance to require matched reruns, aligned
precision baselines, and limit-hit rate comparisons.
* Clarified distributed benchmark timeout handling and score extraction;
GDPVal support is no longer documented.
  * Updated example evaluation time limits and SLURM walltime.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
2026-09-23 00:32:06 +00:00
..
2026-09-10 17:30:37 +00:00