mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: documentation Successful evaluations can still contain timed-out trials or responses stopped by output limits. Existing skills scan for errors but do not require counts or rates, allowing serving-speed effects to be mistaken for quantization accuracy changes. Require per-task timeout and output-limit accounting before scores are treated as validated, with explicit denominators, telemetry coverage, effective limits, and artifact evidence. Separate request retries, terminal trial failures, and resumable SLURM walltime events. Carry that evidence into baseline/candidate comparisons; unknown accounting or unresolved infrastructure effects prevent an acceptable verdict. Preserve benchmark-defined failures and label diagnostic protocol changes explicitly. Allow valid-with-warnings results for small fractions of limit-hit responses/trials when coverage, protocol, and other checks pass, without automatically retrying. Clarify that a BF16 acceptance gate cannot be satisfied with an FP8/INT4 baseline. Correct Terminal-Bench 2.1 guidance that sharding cannot affect scores and require accounting across the full trial set. Raise the ModelOpt TB 2.1 agent budget from 7200 to 14400 seconds in the recipe and set it explicitly in the template. With timeout_strategy=max, this is a minimum budget, not a hard ceiling. Set sandbox lifetime to six hours and example SLURM walltime to eight hours to allow setup and verification; partition limits and effective task budgets must still be checked. Baseline and candidate must use the same policy, and old two-hour results require remeasurement for a matched comparison. This changes skill defaults, not the upstream benchmark protocol or harness instrumentation. The upstream-vendored launching-evals skill is unchanged per repository policy; its standalone workflow still needs an upstream update. Reduce the evaluation entrypoint from 6,156 to 888 words (86%) by moving launcher Steps 1–8 into an on-demand reference without changing their instructions. Preserve step headings/links for existing callers. Shorten timeout accounting from 550 to 352 words (36%) while retaining the checks and warning policy. This reduces initial context; full launcher workflows still load the relevant detailed sections. ### Usage TB 2.1 recipe/template defaults now include: ```yaml solver: timeout_strategy: max run_timeout: 14400 sandbox: max_task_lifetime_sec: 21600 ``` The example uses `cluster.walltime: "08:00:00"` where the partition permits it. Request timeout remains 3600 seconds. For leaderboard comparisons, use the benchmark protocol rather than assuming the ModelOpt override is comparable. ### Testing - Pre-commit passed on all six changed Markdown/YAML files. - skill-creator quick_validate.py passed for evaluation and compare-results. - git diff --check passed. - Validated recipe and template solver/sandbox settings against NEL schemas with the pinned TB playbook; checked max and task timeout resolution. - Verified extracted launcher Steps 1–8 match the original text, apart from a trailing blank line. - Manually reviewed handling of recovered retries, resumed artifacts, incomplete telemetry, benchmark-defined limits, and timeout-sensitive comparisons. No evaluation jobs were launched; increased timeout defaults have not been benchmarked. ### Before your PR is "*Ready for review*" Contributor guidelines and security coding practices reviewed. - Is this change backward compatible?: ✅ Existing configs remain valid. New TB configs use longer budgets and may consume more runtime; comparison requires matched timeout policies. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md: N/A - Did you write any new necessary tests?: N/A — skill/config changes validated as listed above. - Did you update Changelog?: N/A — internal skill guidance. - Did you get Claude approval on this PR?: ❌ Not requested yet. ### Additional Information A nonzero benchmark-defined timeout or output-cap rate does not automatically invalidate a score. The change requires evidence and protocol-aware interpretation, without choosing a universal acceptable rate or silently excluding affected trials. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Expanded evaluation guidance with a centralized launcher workflow covering setup, configuration, execution, monitoring, authentication, and failure handling. * Required timeout and output-limit accounting for every task, including successful and non-reasoning runs, with rates, denominators, telemetry coverage, effective limits, and recovery status. * Clarified that unknown telemetry, mismatched limits, or unresolved infrastructure effects prevent an acceptable verdict. * Updated comparison guidance to require matched reruns, aligned precision baselines, and limit-hit rate comparisons. * Clarified distributed benchmark timeout handling and score extraction; GDPVal support is no longer documented. * Updated example evaluation time limits and SLURM walltime. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>