mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
76c04dfd990660c8a87ae700a89e040a03f278a5
15
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
51cc5dbade |
Make torch 2.14 the unit-test default and constrain it for the tensorrt example images (#2309)
### What does this PR do? Type of change: Bug fix (CI) + test coverage **Fixes `onnx (torch_onnx)` and `onnx (diffusers)`**, which have failed on every branch since `torch 2.14.0` was published to PyPI today (2026-09-02 13:42 UTC), and **adds torch 2.14 to the unit test matrix as the new default** so the next torch release is caught there rather than in an example job. ### Root cause Every test in those two jobs failed with: ``` RuntimeError: CUDNN_BACKEND_TENSOR_DESCRIPTOR cudnnFinalize failed ptrDesc->finalize() cudnn_status: CUDNN_STATUS_SUBLIBRARY_LOADING_FAILED ``` `nvcr.io/nvidia/tensorrt:26.05-py3` ships cuDNN **9.22** and has no preinstalled torch, so pip resolved the newest one — and torch 2.14 pins `nvidia-cudnn-cu13==9.24.0.43`. Loading 9.24 sublibraries against the image's 9.22 `libcudnn.so.9` is exactly what that status reports. | | last good run (08:55) | first failing run (13:34) | |---|---|---| | `torch` | 2.13.0 | **2.14.0** | | `nvidia-cudnn-cu13` | 9.20.0.48 | **9.24.0.43** | | image cuDNN | 9.22.0.52 | 9.22.0.52 | ### Why only these two jobs - The **nemo** and **pytorch** images have a preinstalled torch that already satisfies `torch>=2.8`, so pip never resolves a new one — confirmed from the megatron job log, where torch does not appear in `Successfully installed`. - **`tensorrt:26.05-py3` has no preinstalled torch**, so pip takes the newest from PyPI. - **`onnx (torch_trt)`** shares that image but passes throughout, because `torch-tensorrt<2.13` already holds torch below 2.14. ### The changes 1. **Constrain torch only where the incompatibility is.** `PIP_CONSTRAINT=torch<2.14` in the example runner, applied when the job's image is a `tensorrt` one. It also covers the `examples/*/requirements.txt` loop in the same shell, which matters because `nemo_automodel` pulls torch in too. Not pinned in `pyproject.toml`: torch 2.14 is fine anywhere its own bundled cuDNN is the one loaded, so that would constrain users to work around one pinned image. 2. **Test torch 2.14.** `torch_214` added to `TORCH_VERSIONS` (`torchvision~=0.29.0`) and promoted to the unit-test default across the supported Python versions, with 2.13 demoted to the back-compat row. `release.yml`'s basic unit test moves to the same default (it was still on 2.12). Nothing exercised 2.14 before — which is why a torch release reached us through an example job instead of a unit test. ### Testing - `actionlint` and YAML/TOML parse clean; pre-commit clean. - Verified by this PR's own jobs: `onnx (torch_onnx)` and `onnx (diffusers)` reproduce the failure on `main` right now, and the new `unit-3.12(torch_214, tf_latest)` job is the first run of ModelOpt against torch 2.14. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — CI-only; no source or package metadata change - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new dependency - Did you write any new necessary tests?: ✅ — torch 2.14 added to the unit test matrix - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — internal CI, not user-facing - Did you get Claude approval on this PR?: ❌ — not yet requested --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
d73278808b |
Bump nemo container requirement to 26.08 for MBridge examples (#2257)
### What does this PR do? Type of change: Bug fix Bumps the Megatron-Bridge examples, tests and launcher configs to `nemo:26.08` and removes the version-gated fallbacks they carried, plus the fixes needed to make the suites green on that container. **26.08 bump and shim removal** - Examples, CI workflows, `noxfile.py` and the `mbridge_*` launcher configs move to `nemo:26.08`. - `examples/megatron_bridge/_distillation_provider.py` is deleted — 26.08's Megatron-Bridge ships `convert_to_distillation_provider(..., distill_submodule=...)` natively, so `distill.py` imports it directly. - `prune_minitron.py` drops the `AutoBridge.from_hf_config` / config-only-export probing; `--no_moe_grouped_gemm` is no longer needed in the MoE pruning tests, and the Qwen3.5-MoE `skipif` is gone (native MoE expert mappings are in 26.08). - `_DynamicMambaMixer` targets only the raw `conv1d_weight` / `conv1d_bias` parameters that replaced the `conv1d` module in Megatron-Core. **MambaModel / MambaModelProvider removal** Megatron-Core has shipped `HybridModel` since 26.06 and `MambaModel` is a deprecated subclass that shares its `forward`, so `DMRegistry` resolves those instances to the `HybridModel` registration and the separate entry is redundant. Same for `MambaModelProvider` vs `HybridModelProvider` on the bridge side. `MambaMixer` / `MambaLayer` / `ExtendedRMSNorm` are untouched — the layers still exist. The deprecated `get_te_mamba_stack_spec` is removed; use `get_te_hybrid_stack_spec`. **Bug fix: compressed output_layer extra state** `mtq.compress` converts even a *disabled* `output_layer` into a `RealQuantLinear` (its weight is left uncompressed, since `pack_real_quantize_weight` skips disabled quantizers). The guard added in #2112 excluded `RealQuantLinear` by class, so the layer emitted extra state and every worker died in `GPTModel.sharded_state_dict`: ``` RuntimeError: Boolean value of Tensor with more than one value is ambiguous megatron/core/models/gpt/gpt_model.py:896, in sharded_state_dict output_extra_state and output_extra_state.data ``` The guard now keys off whether the weight was actually compressed (`QTensorWrapper`) instead of the class. This took out all 12 `test_homogeneous_compressed_sharded_state_dict` params, and the crashed workers poisoned the pool, which surfaced as unrelated timeouts and NCCL errors in `test_layer_sync_moe_local_experts_amax`, `test_kv_cache_quant`, `test_kv_cache_amax_sync`, `test_convert_mcore_te_gpt_model` and `test_homogeneous_sharded_state_dict_te_spec` — 21 tests in total. The e2e coverage is `skip_flaky_on_blackwell`, so CI never ran it; `test_output_layer_extra_state_empty_when_nothing_quantized` now asserts the contract directly and is not skipped. **Checkpoint import entry point** 26.08 replaced `examples/conversion/convert_checkpoints.py` with `scripts/conversion/convert.sh`, so `tools/launcher/common/megatron_bridge/import/import.sh` and the three README snippets are retargeted. `import.sh` uses the distributed GPU backend with `GPUS_PER_NODE` / `TP` / `PP` / `EP` knobs. **Megatron-LM on nemo:26.06** keeps working: `_get_mamba_conv1d` still dispatches between the `conv1d` module (26.06 and earlier) and the raw parameters (26.08+), so `import_mcore_gpt_from_hf` / `export_mcore_gpt_to_hf` handle NemotronH on both. Only the Megatron-Bridge examples and Minitron pruning of Mamba/hybrid models require 26.08. **Test consolidation** `test_export_distilled_megatron_to_hf.py` is merged into `test_distill.py`: `test_distill_llm` becomes `test_distill_llm_hf_export` and covers the standalone `--export_iterations all` run on the checkpoints it already produces, saving one full distillation (~185 s of CI time). The two mamba-named gpu test files are renamed to `hybrid`. ### Usage ```bash # HF -> Megatron import, via Megatron-Bridge's 26.08 conversion entry point bash /opt/Megatron-Bridge/scripts/conversion/convert.sh import \ --executor local \ --device gpu \ --gpus-per-node 8 \ --hf-model Qwen/Qwen3-8B \ --megatron-path /tmp/Qwen3-8B-megatron ``` ### Testing All on `nvcr.io/nvidia/nemo:26.08`, 2x RTX 6000 Ada, no timeout overrides: - `tests/examples/megatron_bridge`: 16 passed, 1 skipped (28m14s). The skip is the `gemma3vl` QAD param, now `@pytest.mark.manual` since `qwen3_5_moe_vl` covers the VLM QAD path. - `tests/gpu_megatron` (`_extensions`, `distill`, `export`, `opt`, `peft`, `sparsity`, `speculative`, `utils`): 61 passed, 5 xpassed. - `tests/gpu_megatron/torch/export` re-run after the conv1d dispatch change: 27 passed. - The 21 previously failing/hanging quantization tests: 21 passed (12 + 9). - `tests/gpu_megatron/torch/{nas,prune}`: verified separately. `import.sh` equivalence on a toy `qwen3_moe`, comparing all 12 weight tensors after flattening each dist checkpoint with `dcp_to_torch_save` — the GPU backend at 1 GPU, `--tp 2`, `--pp 2`, `--ep 2`, and `import.sh` end-to-end (`GPUS_PER_NODE=2 EP=2`) are all byte-identical to `--device cpu`. `nemo:26.06` compatibility was checked directly in that image: `megatron.core.models.hybrid.HybridModel`, the modelopt hybrid spec and `hybrid_layer_pattern` are all present, while `megatron.bridge.models.hybrid` and the bridge's `distill_submodule` are not. The NemotronH round-trip test failed there before the conv1d dispatch was restored and the dispatch is back in place; per project convention the suites themselves only run on 26.08. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ⚠️ Megatron-Bridge examples plus Minitron pruning of Mamba/hybrid models now require `nemo:26.08`. Megatron-LM quantization and checkpoint export still run on `nemo:26.06`. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — `test_output_layer_extra_state_empty_when_nothing_quantized` for the compress fix; existing tests extended for the merged export coverage. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added guidance for importing Hugging Face checkpoints into Megatron distributed format. * Expanded distillation workflows to export selected or all checkpoint iterations. * **Improvements** * Expanded Hybrid model support across Megatron workflows. * Updated distributed import tooling with GPU and parallelism options. * Updated supported environments and examples to NVIDIA NeMo 26.08. * **Bug Fixes** * Corrected output-layer quantization state handling when quantization is disabled. * **Documentation** * Added compatibility guidance for current and legacy NeMo containers. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a2fbac7bad |
Pin diffusers<0.40 for the tf_min unit test env (#2226)
### What does this PR do? **Type of change:** Bug fix (CI / test environment) Fixes the `tf_min` unit-test matrix, which has been failing on `main` and every open PR with diffusers export errors in `tests/unit/torch/export/test_export_diffusers.py`. **Root cause — a dependency conflict, not a code regression.** The `tf_min` matrix pins `transformers~=4.57`, which requires `huggingface_hub<1.0`. The project's diffusers dependency is unbounded (`diffusers>=0.32.2`), so a fresh env resolves diffusers **0.40**, which requires `huggingface_hub>=1.23` — unsatisfiable alongside transformers 4.57. The env keeps `huggingface_hub 0.36` (for transformers), so diffusers 0.40's `pipeline_utils` import fails (`get_cached_repo_tree` was added in hub 1.x). That makes `diffusers_utils._HAS_DIFFUSERS` False → `is_diffusers_object()` returns False → diffusers models silently misroute from `_export_diffusers_checkpoint` to the LLM export path, where the LLM dummy-forward feeds integer token inputs into diffusers conv/linear layers (Long/Float `RuntimeError`s) and the quantized-routing test fails (`assert 0 == 1`). It's a time-bomb from an external diffusers 0.40 release auto-pulled by the unbounded pin: the code paths involved were green when they merged, and required checks can't retroactively block already-merged code once a transitive dependency drifts. **Fix:** bound diffusers to `<0.40` for the `tf_min` env only (resolves to 0.39, which supports `huggingface_hub<1.0`), keeping that env internally consistent. `tf_latest` is unchanged. ### Usage N/A — CI / test-environment change. ### Testing Reproduced and validated in a local `tf_min` venv (transformers 4.57 / torch 2.13): - **Before** (diffusers 0.40): 4 failures in `test_export_diffusers.py` — `RuntimeError: Input type (long int) and bias type (float)` (dit), `mat1 and mat2 must have the same dtype, but got Long and Float` (flux), `AttributeError: 'NoneType'` (flux2), `assert 0 == 1` (quantized). - **After the pin** (diffusers resolves to 0.39): `from diffusers import DiffusionPipeline, ModelMixin` succeeds, `is_diffusers_object` routes correctly, and `tests/unit/torch/export/test_export_diffusers.py` passes (**11 passed**) with **unmodified source**. Full `tests/unit/torch/export/` = **165 passed, 10 skipped**. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — test-env only; `tf_latest` unchanged. - If you copied code from any other sources or added a new PIP dependency …: N/A - Did you write any new necessary tests?: N/A — the existing `test_export_diffusers.py` covers this once the env is consistent. - Did you update Changelog?: N/A — CI / test-env fix, not user-facing. - Did you get Claude approval on this PR?: ❌ <!-- run /claude review --> ### Additional Information Follow-ups a maintainer may want (out of scope here): - A **production** robustness fix so `is_diffusers_object` still detects `ModelMixin` components when `DiffusionPipeline` can't import — helps real users who install diffusers 0.40 with transformers 4.57. - A general upper bound / constraint reconciling `diffusers` with `huggingface_hub`. - A separate pre-existing `tf_min` ONNX flake (`test_autocast_quantize`) observed locally but green in CI. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved compatibility for the minimum Transformers environment by applying the appropriate Diffusers version constraint. * Updated test environment setup to install all required package versions for each Transformers configuration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
38826351eb |
Bump transformers dependency to >=4.57,<5.15 (#2050)
Transformers Dependency bump - Drop 4.56 (still support 4.57 with
deprecation note) and extend to 5.14 (nemo:26.08 ships with this
version)
- CI tests now use transformers 5.14
- Manually ran `tests/{gpu_megatron,examples/megatron_bridge}` in
nemo:26.08.rc3
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Compatibility
- Updated supported Transformers versions to 4.57 through 5.14.
- Qwen3-VL models are now available without version-based restrictions.
## Documentation
- Updated the changelog to reflect the new minimum Transformers version
and upcoming removal of Transformers 4.x support.
## Testing
- Updated validation and test configurations for the supported
Transformers versions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
42458def24 |
ci: fix torch_trt on torch 2.13; default unit tests to torch 2.13; add example allow-failure hatch (#1951)
### What does this PR do? Type of change: Bug fix + CI / infra torch 2.13 + torchvision 0.28 were published to PyPI on 2026-07-08 and broke the `onnx (torch_trt)` example job (which had passed the day before). This PR fixes that break and hardens CI against the next one: 1. **Fix `torch_trt` (torch/torchvision/torch-tensorrt trio pin).** The base install pulled torch 2.13 / torchvision 0.28, then `torch-tensorrt 2.12.1` downgraded torch back to 2.12 but left torchvision at 0.28 (which pins `torch==2.13`) — breaking `import`. `examples/torch_trt/requirements.txt` now caps `torch-tensorrt>=2.4.0,<2.13` + `torchvision<0.28` so the trio stays consistent (also protects direct `pip install -r` users). 2. **Unit tests default to torch 2.13.** `noxfile.py` gains `torch_213` (`torchvision~=0.28.0`); the required `linux`/`windows` jobs and the multi-version Python spread (3.10/3.11/3.13/3.14) now run torch 2.13, with torch 2.8–2.12 kept as back-compat legs on Python 3.12. 3. **Per-example allow-failure escape hatch.** `_example_tests_runner.yml` gains an `allow_failure` input; when set, a **test-run** failure is surfaced as a `::warning::` via `continue-on-error` instead of blocking the PR. `example_tests.yml` derives it per example from the repo variable **`ALLOW_FAILURE_EXAMPLE_TESTS`** (comma-separated example names, comma-wrapped so `onnx` ≠ `torch_onnx`). Future breakages can be quarantined by updating the variable — no code change / PR required. ### Testing - Ran the new default unit session locally in an isolated uv venv (torch **2.13.0**+cu130, torchvision **0.28.0**+cu130, transformers **5.12.1**): ``` nox -s "unit-3.12(torch_213, tf_latest)" => 2813 passed, 15 skipped, 1786 warnings in 250.37s ``` - Verified the allow-failure hatch: with `ALLOW_FAILURE_EXAMPLE_TESTS=torch_trt`, the (previously failing) `torch_trt` job reports success with a warning and does not block the required example check. `vars` is re-read on each job attempt, so "Re-run failed jobs" picks up the variable without a fresh trigger. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (CI-only; older torch versions still covered) - If you copied code from any other sources or added a new PIP dependency: N/A (no new dependency; only version caps) - Did you write any new necessary tests?: N/A (CI configuration change) - Did you update Changelog?: N/A (CI infra, no user-facing API change) - Did you get Claude approval on this PR?: ❌ (pending — will run `/claude review`) ### Additional Information The `ALLOW_FAILURE_EXAMPLE_TESTS` repo variable can be cleared for `torch_trt` now that the requirements pin lands the real fix; keep it as the standing escape hatch for future example breakages. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Example test jobs can now be configured to “allow failure” without failing the workflow; when enabled, a warning annotation is emitted. * **Bug Fixes** * CI unit-test and GPU-test configurations were refreshed (including a reduced timeout for the `gpu_megatron` job). * **Tests** * Updated unit-test coverage to use the newest Torch 2.13-based setup by default, with back-compat retained where applicable. * **Documentation** * Added inline guidance for how the allow-failure examples list is specified. * **Chores** * Refreshed `torch_trt` example dependency constraints to improve compatibility. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
795c589428 |
CI: CUDA build/test hygiene + fix Puzzletron Nemotron test failures (#1901)
### What does this PR do? Type of change: Bug fix (CI / tests) - **Fix Nemotron nightly failures:** install `mamba_ssm`/`causal-conv1d` from PyPI releases instead of git `main` (avoids the broken `apache-tvm-ffi 0.1.12` that crashes on import). - **Speed up CUDA builds:** set `TORCH_CUDA_ARCH_LIST=12.0` (runner's sm_120) in the GPU/example/regression workflow container env instead of the image's ~6 archs. - **Make unit tests CPU-only:** force CUDA off in the nox `unit` env and skip JIT-compiling CUDA extensions when no GPU is usable; move the two GPU-/`mamba_ssm`-requiring unit tests to `tests/gpu`. - **Harden example tests against HF flakes:** capture subprocess output and retry transient HuggingFace access errors (5xx / rate-limit / connection). - **Skip Blackwell-flaky sharded-state-dict tests:** `test_homogeneous_sharded_state_dict` and `test_regular_state_dict[320]` intermittently hit a CUDA illegal-memory-access on the sm_120 runner that poisons the CUDA context and cascades timeouts; gate them behind a reusable `skip_flaky_on_blackwell` marker (still run on non-Blackwell GPUs). - **Bump slow test timeout:** `test_prune_minitron_vlm` → 360s for the 2-GPU nightly. ### Testing - CI tests on this PR pass (1-gpu) - Manually triggerred 2-gpu test: - GPU: https://github.com/NVIDIA/Model-Optimizer/actions/runs/28774553356 - Examples: https://github.com/NVIDIA/Model-Optimizer/actions/runs/28774556693 - Regression: https://github.com/NVIDIA/Model-Optimizer/actions/runs/28771197586 ### Additional Information - Backward compatible: N/A (CI/tests only) - New dependency: N/A - Changelog: N/A (CI/test infra) 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
33bfa8b1fe |
CI/Dev env bump (#1818)
### What does this PR do? Type of change: chore Bumps CI/dev tooling and test containers. **Container bumps** - NeMo test containers → 26.06 - TRT-LLM container → 1.3.0rc19 - transformers max version → 5.12 **Dev tooling bumps** - ruff bump 0.12.11 → 0.15.18 - mypy 1.17.1 → 2.1.0: enable new defaults (`local_partial_types`, `strict_bytes`); fix/narrow the errors newly surfaced by mypy 2.0 in 4 modules (rather than blanket-suppressing them); remove 2 stale `# type: ignore` comments - pre-commit 4.3.0 → 4.6.0 - sphinx 8.1 → 9.1 + sphinx-rtd-theme 3.0 → 3.1: add `suppress_warnings = ["ref.python"]` to fix cross-reference ambiguity error new in sphinx 9.x - trl fix for newly released 1.7 version **Bug fixes surfaced by the bumps** - sparsity (weight): make the weight mask DTensor-aware under FSDP. The transformers→5.12 bump routes the HF Trainer FSDP optimizer-state save through torch's DTensor-based `get_optimizer_state_dict`, which triggered `aten.mul.Tensor got mixed torch.Tensor and DTensor` in the dynamic `weight` getter. The mask is now distributed to the weight's mesh/placements before masking, cached, and rebuilt only when the sharding changes (invalidated on `set_mask`). Fixes the `llm_sparsity` example test. ### Testing - `pre-commit run --all-files` ✅ (including mypy 2.1.0) - `nox -s docs` ✅ - `tests/unit/torch/sparsity` + `tests/unit/torch/nas` ✅ - `llm_sparsity` GPU example test (FSDP path) verified in CI ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Summary * **Documentation** * Refreshed Docker pre-requisites across examples to recommend updated container image tags (and streamlined some instructions). * **Bug Fixes** * Improved sparse weight mask handling for DTensor/FSDP by aligning and caching distributed masks. * Made TensorRT engine byte retrieval return immutable `bytes`. * Reduced Sphinx cross-reference warnings and tuned Transformers compatibility warning thresholds. * **Tests** * Increased default unit test timeout on Windows runners. * **Chores** * Updated CI workflow container tags and refreshed linting/typing/docs version pins, plus related mypy configuration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
b98a59557a |
Add vLLM-based runtime statistics for subblock latency measurement (#1358)
### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Runtime-based latency optimization: collect vLLM-measured inference latency to constrain optimization. * **Configuration** * New runtime config/template for Llama-3.1-8B pruning (runtime stats enabled, NCCL timeout templating, MIP target-latency). * Validation sample defaults adjusted (one flow: 128 → 8; runtime flow uses 128). * Human constraint key renamed to target_latency_seconds. * **Documentation** * README section describing runtime-based latency optimization setup and usage. * **Tests** * Added GPU end-to-end test for runtime stats collection. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1358?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Grzegorz Karch <gkarch@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
7aa0c95646 |
Add tests/gpu_vllm (#1517)
### What does this PR do? Type of change: new tests This PR adds unit tests for vLLM fakequant, specifically testing code in `modelopt/torch/quantization/plugins/vllm.py` ### Testing ``` pytest tests/gpu_vllm/torch/quantization/test_vllm_dynamic_modules.py -sv ``` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added comprehensive GPU vLLM test suite with end-to-end quantization checks and fixtures for TinyLlama, TinyQwen3-MoE, and DeepSeek V3; includes helpers to build tiny DeepSeek V3 models. * **Chores** * Updated GPU CI to use explicit container image references, added a GPU-focused test session, and adjusted test-run setup for vLLM. * **Documentation** * Documented new GPU test directory in contributing guide. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1517?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
eb5ed2df68 |
[CI] Bump torch, transformers and dev containers to latest (#1554)
- Transformers upper bound bumped from `<5.8` to `<5.10` - Enable torch 2.12 CICD testing - Bump TRT-LLM container to `1.3.0rc16` (transformers 5.5) - Use pytorch and tensorrt 26.04 containers in CICD <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated CI test container images and targeted Torch version across workflows; adjusted release CI job to use the newer torch config. * Broadened Transformers constraint in project metadata and test/dev pins. * Removed strict transformers pins from example requirements and lifted a compression dependency cap. * Raised the import-time Transformers version threshold for compatibility warnings. * **Tests** * Refactored a GPU test to collect and report validation errors and updated numeric expected baselines. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1554?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
4b270f0ea6 |
Support Mixed precision & Static MSE in MCore; Nemotron Super v3 NVFP4 recipe (#1521)
### What does this PR do?
Type of change: New features + Bug fixes
Mixed Precision and MSE support in MCore PTQ
- support mixed precision export in MCore by detecting mixed precision
layers in HF Quant Config
- Restore static quantizer in MCore checkpoint restore as `NVFP4QTensor`
(not TensorQuantizer which can call max calibrate. we want to skip max
calibrate for static quantizer during restore) --> fixes bug during
MCore export for MSE
- Fix dynamic block quantizer detection when `block_sizes` is
dict-backed.
- Add a YAML quantization recipe that roughly mirrors Nemotron 3 Super
NVFP4 `hf_quant_config.json`
Export bug fixes
- copy .py files properly from original HF ckpt (for reasoning parser
etc)
## Super recipe
Mirrors the published nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
hf_quant_config.json:
- MoE routed experts (mixer.experts.<N>.{up,down}_proj): NVFP4 W4A4
weight MSE, group_size 16
- MoE shared experts (mixer.shared_experts.{up,down}_proj): FP8
per-tensor
- Mamba mixer linears (mixer.{in,out}_proj): FP8 per-tensor
- KV cache: FP8
rest: not quantized
### Usage
```python
# Add a code snippet demonstrating how to use this
```
### Testing
Tested on Nemotron model
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added NVFP4 (4-bit) quantization checkpoint restore and export support
for Megatron-Core models
* Added tokenizer file export capability in model checkpoints
* Extended quantization support for expert-parallel distributed training
* Introduced new PTQ recipes for Nemotron-3-Super models with
mixed-precision quantization
* **Bug Fixes**
* Fixed FP8 and FP4 hardware compatibility detection on non-CUDA systems
* Improved offline Hugging Face Hub access handling with better error
messaging
* Enhanced calibration validation for mixture-of-experts models
* Fixed amplitude maximum validation for static block quantizers
* **Documentation**
* Updated expert weight quantization configuration documentation
<!-- review_stack_entry_start -->
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1521?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)
<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Jenny Chen <jennifchen@nvidia.com>
|
||
|
|
72799ab551 |
Pin transformers<5.8 to avoid CI failures (#1395)
Transformers now releases new versions weekly/biweekly causing frequest breaking changes in our tests (We always pin a max version for release branches so thats unaffected). Hence pinning max version to 5.7 in main branch. We will bump transformers monthly / with torch bump instead of always using latest one Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
70546bdd6a |
Enable Python 3.14 wheel support to unblock NGC PyTorch container testing on Ubuntu 26.04 + Python 3.14 (#1386)
Ubuntu 26.04 is here and very soon, NVIDIA PyTorch containers will ship with Python 3.14 requiring us to enable untested support to unblock them <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * DFlash offline speculative decoding training * MXFP4→NVFP4 weight conversion support * Shared hidden-state dump utilities * Updated DeepSeek PTQ calibration defaults * **Chores** * Added Python 3.14 support; updated Python requirement to <3.15 * **Documentation** * Updated installation documentation for Python version compatibility <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> |
||
|
|
c51c1762b3 |
fix: prevent gh-pages repo bloat from doc preview artifacts (#1309)
### What does this PR do? Type of change: Bug fix Fixes gh-pages branch bloat that grew from ~26 MB to ~441 MB in four weeks (nvbug 6099503). Three compounding causes were identified and addressed: 1. **Sphinx `.doctrees/` cache published to gh-pages** — `sphinx-build` was writing its build cache inside `build/html/` which was then uploaded verbatim. Accounts for ~3.3 GB uncompressed across history. 2. **`JamesIves/github-pages-deploy-action` appending a commit on every push** — main-site files accumulated forever with `single-commit: false` (default). 3. **PR preview deploying on every `synchronize` event for all PRs** — `rossjrw/pr-preview-action` re-deployed the full site for every push to any PR regardless of whether docs changed (e.g. PR #1128 triggered 64 preview deploys × ~11 MB each). Changes: - Pass `-d /tmp/doctrees` to `sphinx-build` so `.doctrees/` is never written into `build/html/` - Add `paths: [docs/**, modelopt/**]` filter to `pull_request` trigger so the docs workflow only runs on PRs that touch docs or source code - Set `single-commit: true` on the deploy action so main-site pushes squash into one commit - Deduplicate docs build: `deploy-preview` now downloads the artifact from `build-docs` instead of running a second `sphinx-build` - Set `retention-days: 1` on the artifact since it is only needed for the duration of the workflow run The one-time cleanup (force-push squashed orphan to gh-pages) was already applied separately — repo is now ~59 MB for a full clone vs ~441 MB before. ### Usage N/A — CI/workflow change only. ### Testing - Workflow logic reviewed manually. - The one-time cleanup was verified: `git rev-list --objects --disk-usage origin/gh-pages` now reports ~28 MB; full clone is ~59 MB. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information nvbug 6099503 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Optimized documentation build and deployment workflow in CI/CD pipeline. * Improved pull request documentation preview handling with faster build timeouts and refined artifact management. * Enhanced GitHub Pages deployment configuration for better consistency. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
3d0f0db49e |
[CI] Replace tox with nox, use nemo:26.04 for megatron tests, and simplify CI workflows (#1286)
### What does this PR do? Type of change: New feature / infrastructure improvement Follow-up to #1285 for correct CI test environment for megatron based tests Replaces `tox` + `tox-current-env` with `nox` for all test, lint, docs, and wheel build sessions. The primary motivation was that `tox-current-env` is incompatible with uv venvs in NGC containers (e.g. NeMo's `/opt/venv`) — it picks the system Python via `sys._base_executable` instead of the container's venv Python which has megatron packages pre-installed. Key changes: - **`noxfile.py`** replaces `tox.ini` with GPU, CPU unit, partial-install, pre-commit, docs, and wheel sessions - **GPU sessions** use `venv_backend="none"` (run directly in container env) and `python -m pip/pytest` to avoid PATH mismatches - **uv** is set as the default venv backend (if available) for CPU sessions (faster installs) Also includes CI workflow simplifications: - **`_pr_gate.yml`** new reusable workflow centralizing file-change detection + linux-check wait logic (was duplicated across 3 workflow files) - **Collapsed pr/non-pr job pairs** into single jobs with conditional `runs-on` in `gpu_tests.yml`, `example_tests.yml`, `regression_tests.yml` - **Collapsed `multi-py` / `multi-torch` / `multi-transformers`** into a single `multi-version` matrix job in `unit_tests.yml` - **PR path filtering** for unit test secondary jobs (multi-version, launcher, partial-install) — skipped if no relevant files changed - **Fixed schedule/workflow_dispatch skipping** — jobs with `needs: [pr-gate]` were incorrectly skipped when all pr-gate internal jobs were skipped; fixed by making the gate job always run - **multi-version, launcher, partial-install** now also run on `schedule` / `workflow_dispatch` ### Usage ```bash python -m pip install nox uv # install nox and uv (once) nox -l # list all sessions nox -s gpu_megatron # run a GPU session (inside container) nox -s "unit-3.12(torch_211, tf_latest)" # run a specific unit test combination nox -s "unit-3.12(torch_211, tf_latest)" -R # force-recreate venv (e.g. after dep changes) COVERAGE_PROCESS_START=pyproject.toml nox -s "unit-3.12(torch_211, tf_latest)" # with coverage ``` ### Testing - Ran `nox -l` to verify all session names - Ran `gpu_megatron` session locally inside NeMo container — confirmed it uses `/opt/venv/bin/python` correctly - Manually triggered nightly-runs: - Unit: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608013657 - GPU: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608018763 - Examples: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608017322 ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A — CI infrastructure only - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ (added `nox` and `uv` to `dev-test`, both Apache-2.0) - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — no user-facing changes ### Additional Information Supersedes the tox-current-env workaround in the parent branch. --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |