mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
CI/Dev env bump (#1818)
### What does this PR do? Type of change: chore Bumps CI/dev tooling and test containers. **Container bumps** - NeMo test containers → 26.06 - TRT-LLM container → 1.3.0rc19 - transformers max version → 5.12 **Dev tooling bumps** - ruff bump 0.12.11 → 0.15.18 - mypy 1.17.1 → 2.1.0: enable new defaults (`local_partial_types`, `strict_bytes`); fix/narrow the errors newly surfaced by mypy 2.0 in 4 modules (rather than blanket-suppressing them); remove 2 stale `# type: ignore` comments - pre-commit 4.3.0 → 4.6.0 - sphinx 8.1 → 9.1 + sphinx-rtd-theme 3.0 → 3.1: add `suppress_warnings = ["ref.python"]` to fix cross-reference ambiguity error new in sphinx 9.x - trl fix for newly released 1.7 version **Bug fixes surfaced by the bumps** - sparsity (weight): make the weight mask DTensor-aware under FSDP. The transformers→5.12 bump routes the HF Trainer FSDP optimizer-state save through torch's DTensor-based `get_optimizer_state_dict`, which triggered `aten.mul.Tensor got mixed torch.Tensor and DTensor` in the dynamic `weight` getter. The mask is now distributed to the weight's mesh/placements before masking, cached, and rebuilt only when the sharding changes (invalidated on `set_mask`). Fixes the `llm_sparsity` example test. ### Testing - `pre-commit run --all-files` ✅ (including mypy 2.1.0) - `nox -s docs` ✅ - `tests/unit/torch/sparsity` + `tests/unit/torch/nas` ✅ - `llm_sparsity` GPU example test (FSDP path) verified in CI ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Summary * **Documentation** * Refreshed Docker pre-requisites across examples to recommend updated container image tags (and streamlined some instructions). * **Bug Fixes** * Improved sparse weight mask handling for DTensor/FSDP by aligning and caching distributed masks. * Made TensorRT engine byte retrieval return immutable `bytes`. * Reduced Sphinx cross-reference warnings and tuned Transformers compatibility warning thresholds. * **Tests** * Increased default unit test timeout on Windows runners. * **Chores** * Updated CI workflow container tags and refreshed linting/typing/docs version pins, plus related mypy configuration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 4.6
parent
55d6e75833
commit
33bfa8b1fe
+1
-1
@@ -9,5 +9,5 @@ coverage:
|
||||
project:
|
||||
default:
|
||||
target: auto
|
||||
threshold: 1% # Allow atmost 1% coverage drop from main branch.
|
||||
threshold: 2% # Allow atmost 2% coverage drop from main branch.
|
||||
patch: false
|
||||
|
||||
@@ -59,7 +59,7 @@ jobs:
|
||||
uses: ./.github/workflows/_example_tests_runner.yml
|
||||
secrets: inherit
|
||||
with:
|
||||
docker_image: "nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc17"
|
||||
docker_image: "nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc19"
|
||||
example: ${{ matrix.example }}
|
||||
pip_install_extras: "[hf,dev-test]"
|
||||
runner: linux-amd64-gpu-rtxpro6000-latest-1
|
||||
@@ -73,7 +73,7 @@ jobs:
|
||||
uses: ./.github/workflows/_example_tests_runner.yml
|
||||
secrets: inherit
|
||||
with:
|
||||
docker_image: "nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc17"
|
||||
docker_image: "nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc19"
|
||||
example: ${{ matrix.example }}
|
||||
pip_install_extras: "[hf,dev-test]"
|
||||
runner: linux-amd64-gpu-rtxpro6000-latest-2
|
||||
@@ -85,7 +85,7 @@ jobs:
|
||||
uses: ./.github/workflows/_example_tests_runner.yml
|
||||
secrets: inherit
|
||||
with:
|
||||
docker_image: "nvcr.io/nvidia/nemo:26.04"
|
||||
docker_image: "nvcr.io/nvidia/nemo:26.06"
|
||||
example: megatron_bridge
|
||||
timeout_minutes: 45
|
||||
pip_install_extras: "[hf,puzzletron,dev-test]"
|
||||
|
||||
@@ -43,10 +43,10 @@ jobs:
|
||||
container_image: nvcr.io/nvidia/pytorch:26.05-py3
|
||||
- example: gpu_megatron
|
||||
timeout: 60
|
||||
container_image: nvcr.io/nvidia/nemo:26.04
|
||||
container_image: nvcr.io/nvidia/nemo:26.06
|
||||
- example: gpu_trtllm
|
||||
timeout: 15
|
||||
container_image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc17
|
||||
container_image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc19
|
||||
- example: gpu_vllm
|
||||
timeout: 15
|
||||
container_image: docker.io/vllm/vllm-openai:v0.20.0
|
||||
|
||||
Reference in New Issue
Block a user