mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
## Summary - fix launcher Slurm task annotation patching so nemo-run CLI resolves `slurm_factory` correctly for task slots - harden `modelopt-mcp` submit parsing/status resolution and add regression coverage for launcher false-positive success cases - add a minimal `nvidia-smi` smoke YAML/script and fix launcher packaging so source-backed Slurm jobs package required files recursively ## Validation - `uv run pytest tests/test_core.py -q` - `uv run pytest tests/test_bridge.py -q` - dry-run and live-submit validated through the patched local MCP server on `cw_dfw` - interactive smoke job succeeded end-to-end (`nvidia-smi` ran successfully in-container) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an NVIDIA SMI GPU smoke test (script and minimal Slurm YAML example) for launcher integration. * **Bug Fixes** * Improved detection of fatal launcher errors, including when the launcher exits with code 0. * Strengthened Slurm experiment/job identifier parsing and added early rejection of unsafe experiment IDs, with clearer “unparsed”/failure behavior. * Updated sandbox task Slurm config type handling and improved launcher packaging so examples/common are included consistently. * **Tests** * Expanded unit and filesystem-based coverage for parsing/validation, dry-run fatal stderr handling, and nested experiment directory layouts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Signed-off-by: Chenhan D. Yu <chenhany@nvidia.com>