Commit Graph
15 Commits
Author SHA1 Message Date
Keval Morabia 51cc5dbade Make torch 2.14 the unit-test default and constrain it for the tensorrt example images (#2309)
### What does this PR do?

Type of change: Bug fix (CI) + test coverage

**Fixes `onnx (torch_onnx)` and `onnx (diffusers)`**, which have failed
on every branch since
`torch 2.14.0` was published to PyPI today (2026-09-02 13:42 UTC), and
**adds torch 2.14 to the unit
test matrix as the new default** so the next torch release is caught
there rather than in an example job.

### Root cause

Every test in those two jobs failed with:

```
RuntimeError: CUDNN_BACKEND_TENSOR_DESCRIPTOR cudnnFinalize failed
  ptrDesc->finalize() cudnn_status: CUDNN_STATUS_SUBLIBRARY_LOADING_FAILED
```

`nvcr.io/nvidia/tensorrt:26.05-py3` ships cuDNN **9.22** and has no
preinstalled torch, so pip
resolved the newest one — and torch 2.14 pins
`nvidia-cudnn-cu13==9.24.0.43`. Loading 9.24
sublibraries against the image's 9.22 `libcudnn.so.9` is exactly what
that status reports.

| | last good run (08:55) | first failing run (13:34) |
|---|---|---|
| `torch` | 2.13.0 | **2.14.0** |
| `nvidia-cudnn-cu13` | 9.20.0.48 | **9.24.0.43** |
| image cuDNN | 9.22.0.52 | 9.22.0.52 |

### Why only these two jobs

- The **nemo** and **pytorch** images have a preinstalled torch that
already satisfies `torch>=2.8`,
so pip never resolves a new one — confirmed from the megatron job log,
where torch does not appear
  in `Successfully installed`.
- **`tensorrt:26.05-py3` has no preinstalled torch**, so pip takes the
newest from PyPI.
- **`onnx (torch_trt)`** shares that image but passes throughout,
because `torch-tensorrt<2.13`
  already holds torch below 2.14.

### The changes

1. **Constrain torch only where the incompatibility is.**
`PIP_CONSTRAINT=torch<2.14` in the example
runner, applied when the job's image is a `tensorrt` one. It also covers
the
`examples/*/requirements.txt` loop in the same shell, which matters
because `nemo_automodel`
pulls torch in too. Not pinned in `pyproject.toml`: torch 2.14 is fine
anywhere its own bundled
cuDNN is the one loaded, so that would constrain users to work around
one pinned image.
2. **Test torch 2.14.** `torch_214` added to `TORCH_VERSIONS`
(`torchvision~=0.29.0`) and promoted to
the unit-test default across the supported Python versions, with 2.13
demoted to the back-compat
row. `release.yml`'s basic unit test moves to the same default (it was
still on 2.12).
Nothing exercised 2.14 before — which is why a torch release reached us
through an example
   job instead of a unit test.

### Testing

- `actionlint` and YAML/TOML parse clean; pre-commit clean.
- Verified by this PR's own jobs: `onnx (torch_onnx)` and `onnx
(diffusers)` reproduce the failure on
`main` right now, and the new `unit-3.12(torch_214, tf_latest)` job is
the first run of ModelOpt
  against torch 2.14.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — CI-only; no source or package
metadata change
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependency
- Did you write any new necessary tests?: ✅ — torch 2.14 added to the
unit test matrix
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — internal CI, not user-facing
- Did you get Claude approval on this PR?: ❌ — not yet requested

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-09-03 01:11:16 +05:30
Keval MorabiaandClaude Opus 5 d73278808b Bump nemo container requirement to 26.08 for MBridge examples (#2257)
### What does this PR do?

Type of change: Bug fix

Bumps the Megatron-Bridge examples, tests and launcher configs to
`nemo:26.08` and removes the version-gated fallbacks they carried, plus
the fixes needed to make the suites green on that container.

**26.08 bump and shim removal**

- Examples, CI workflows, `noxfile.py` and the `mbridge_*` launcher
configs move to `nemo:26.08`.
- `examples/megatron_bridge/_distillation_provider.py` is deleted —
26.08's Megatron-Bridge ships `convert_to_distillation_provider(...,
distill_submodule=...)` natively, so `distill.py` imports it directly.
- `prune_minitron.py` drops the `AutoBridge.from_hf_config` /
config-only-export probing; `--no_moe_grouped_gemm` is no longer needed
in the MoE pruning tests, and the Qwen3.5-MoE `skipif` is gone (native
MoE expert mappings are in 26.08).
- `_DynamicMambaMixer` targets only the raw `conv1d_weight` /
`conv1d_bias` parameters that replaced the `conv1d` module in
Megatron-Core.

**MambaModel / MambaModelProvider removal**

Megatron-Core has shipped `HybridModel` since 26.06 and `MambaModel` is
a deprecated subclass that shares its `forward`, so `DMRegistry`
resolves those instances to the `HybridModel` registration and the
separate entry is redundant. Same for `MambaModelProvider` vs
`HybridModelProvider` on the bridge side. `MambaMixer` / `MambaLayer` /
`ExtendedRMSNorm` are untouched — the layers still exist. The deprecated
`get_te_mamba_stack_spec` is removed; use `get_te_hybrid_stack_spec`.

**Bug fix: compressed output_layer extra state**

`mtq.compress` converts even a *disabled* `output_layer` into a
`RealQuantLinear` (its weight is left uncompressed, since
`pack_real_quantize_weight` skips disabled quantizers). The guard added
in #2112 excluded `RealQuantLinear` by class, so the layer emitted extra
state and every worker died in `GPTModel.sharded_state_dict`:

```
RuntimeError: Boolean value of Tensor with more than one value is ambiguous
  megatron/core/models/gpt/gpt_model.py:896, in sharded_state_dict
    output_extra_state and output_extra_state.data
```

The guard now keys off whether the weight was actually compressed
(`QTensorWrapper`) instead of the class. This took out all 12
`test_homogeneous_compressed_sharded_state_dict` params, and the crashed
workers poisoned the pool, which surfaced as unrelated timeouts and NCCL
errors in `test_layer_sync_moe_local_experts_amax`,
`test_kv_cache_quant`, `test_kv_cache_amax_sync`,
`test_convert_mcore_te_gpt_model` and
`test_homogeneous_sharded_state_dict_te_spec` — 21 tests in total. The
e2e coverage is `skip_flaky_on_blackwell`, so CI never ran it;
`test_output_layer_extra_state_empty_when_nothing_quantized` now asserts
the contract directly and is not skipped.

**Checkpoint import entry point**

26.08 replaced `examples/conversion/convert_checkpoints.py` with
`scripts/conversion/convert.sh`, so
`tools/launcher/common/megatron_bridge/import/import.sh` and the three
README snippets are retargeted. `import.sh` uses the distributed GPU
backend with `GPUS_PER_NODE` / `TP` / `PP` / `EP` knobs.

**Megatron-LM on nemo:26.06** keeps working: `_get_mamba_conv1d` still
dispatches between the `conv1d` module (26.06 and earlier) and the raw
parameters (26.08+), so `import_mcore_gpt_from_hf` /
`export_mcore_gpt_to_hf` handle NemotronH on both. Only the
Megatron-Bridge examples and Minitron pruning of Mamba/hybrid models
require 26.08.

**Test consolidation**

`test_export_distilled_megatron_to_hf.py` is merged into
`test_distill.py`: `test_distill_llm` becomes
`test_distill_llm_hf_export` and covers the standalone
`--export_iterations all` run on the checkpoints it already produces,
saving one full distillation (~185 s of CI time). The two mamba-named
gpu test files are renamed to `hybrid`.

### Usage

```bash
# HF -> Megatron import, via Megatron-Bridge's 26.08 conversion entry point
bash /opt/Megatron-Bridge/scripts/conversion/convert.sh import \
    --executor local \
    --device gpu \
    --gpus-per-node 8 \
    --hf-model Qwen/Qwen3-8B \
    --megatron-path /tmp/Qwen3-8B-megatron
```

### Testing

All on `nvcr.io/nvidia/nemo:26.08`, 2x RTX 6000 Ada, no timeout
overrides:

- `tests/examples/megatron_bridge`: 16 passed, 1 skipped (28m14s). The
skip is the `gemma3vl` QAD param, now `@pytest.mark.manual` since
`qwen3_5_moe_vl` covers the VLM QAD path.
- `tests/gpu_megatron` (`_extensions`, `distill`, `export`, `opt`,
`peft`, `sparsity`, `speculative`, `utils`): 61 passed, 5 xpassed.
- `tests/gpu_megatron/torch/export` re-run after the conv1d dispatch
change: 27 passed.
- The 21 previously failing/hanging quantization tests: 21 passed (12 +
9).
- `tests/gpu_megatron/torch/{nas,prune}`: verified separately.

`import.sh` equivalence on a toy `qwen3_moe`, comparing all 12 weight
tensors after flattening each dist checkpoint with `dcp_to_torch_save` —
the GPU backend at 1 GPU, `--tp 2`, `--pp 2`, `--ep 2`, and `import.sh`
end-to-end (`GPUS_PER_NODE=2 EP=2`) are all byte-identical to `--device
cpu`.

`nemo:26.06` compatibility was checked directly in that image:
`megatron.core.models.hybrid.HybridModel`, the modelopt hybrid spec and
`hybrid_layer_pattern` are all present, while
`megatron.bridge.models.hybrid` and the bridge's `distill_submodule` are
not. The NemotronH round-trip test failed there before the conv1d
dispatch was restored and the dispatch is back in place; per project
convention the suites themselves only run on 26.08.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ⚠️ Megatron-Bridge examples plus
Minitron pruning of Mamba/hybrid models now require `nemo:26.08`.
Megatron-LM quantization and checkpoint export still run on
`nemo:26.06`.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ —
`test_output_layer_extra_state_empty_when_nothing_quantized` for the
compress fix; existing tests extended for the merged export coverage.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added guidance for importing Hugging Face checkpoints into Megatron
distributed format.
* Expanded distillation workflows to export selected or all checkpoint
iterations.

* **Improvements**
  * Expanded Hybrid model support across Megatron workflows.
  * Updated distributed import tooling with GPU and parallelism options.
  * Updated supported environments and examples to NVIDIA NeMo 26.08.

* **Bug Fixes**
* Corrected output-layer quantization state handling when quantization
is disabled.

* **Documentation**
  * Added compatibility guidance for current and legacy NeMo containers.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 00:31:41 +05:30
Shengliang Xu a2fbac7bad Pin diffusers<0.40 for the tf_min unit test env (#2226)
### What does this PR do?

**Type of change:** Bug fix (CI / test environment)

Fixes the `tf_min` unit-test matrix, which has been failing on `main`
and every open PR with diffusers export errors in
`tests/unit/torch/export/test_export_diffusers.py`.

**Root cause — a dependency conflict, not a code regression.** The
`tf_min` matrix pins `transformers~=4.57`, which requires
`huggingface_hub<1.0`. The project's diffusers dependency is unbounded
(`diffusers>=0.32.2`), so a fresh env resolves diffusers **0.40**, which
requires `huggingface_hub>=1.23` — unsatisfiable alongside transformers
4.57. The env keeps `huggingface_hub 0.36` (for transformers), so
diffusers 0.40's `pipeline_utils` import fails (`get_cached_repo_tree`
was added in hub 1.x). That makes `diffusers_utils._HAS_DIFFUSERS` False
→ `is_diffusers_object()` returns False → diffusers models silently
misroute from `_export_diffusers_checkpoint` to the LLM export path,
where the LLM dummy-forward feeds integer token inputs into diffusers
conv/linear layers (Long/Float `RuntimeError`s) and the
quantized-routing test fails (`assert 0 == 1`).

It's a time-bomb from an external diffusers 0.40 release auto-pulled by
the unbounded pin: the code paths involved were green when they merged,
and required checks can't retroactively block already-merged code once a
transitive dependency drifts.

**Fix:** bound diffusers to `<0.40` for the `tf_min` env only (resolves
to 0.39, which supports `huggingface_hub<1.0`), keeping that env
internally consistent. `tf_latest` is unchanged.

### Usage

N/A — CI / test-environment change.

### Testing

Reproduced and validated in a local `tf_min` venv (transformers 4.57 /
torch 2.13):
- **Before** (diffusers 0.40): 4 failures in `test_export_diffusers.py`
— `RuntimeError: Input type (long int) and bias type (float)` (dit),
`mat1 and mat2 must have the same dtype, but got Long and Float` (flux),
`AttributeError: 'NoneType'` (flux2), `assert 0 == 1` (quantized).
- **After the pin** (diffusers resolves to 0.39): `from diffusers import
DiffusionPipeline, ModelMixin` succeeds, `is_diffusers_object` routes
correctly, and `tests/unit/torch/export/test_export_diffusers.py` passes
(**11 passed**) with **unmodified source**. Full
`tests/unit/torch/export/` = **165 passed, 10 skipped**.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — test-env only; `tf_latest`
unchanged.
- If you copied code from any other sources or added a new PIP
dependency …: N/A
- Did you write any new necessary tests?: N/A — the existing
`test_export_diffusers.py` covers this once the env is consistent.
- Did you update Changelog?: N/A — CI / test-env fix, not user-facing.
- Did you get Claude approval on this PR?: ❌ <!-- run /claude review -->

### Additional Information

Follow-ups a maintainer may want (out of scope here):
- A **production** robustness fix so `is_diffusers_object` still detects
`ModelMixin` components when `DiffusionPipeline` can't import — helps
real users who install diffusers 0.40 with transformers 4.57.
- A general upper bound / constraint reconciling `diffusers` with
`huggingface_hub`.
- A separate pre-existing `tf_min` ONNX flake (`test_autocast_quantize`)
observed locally but green in CI.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved compatibility for the minimum Transformers environment by
applying the appropriate Diffusers version constraint.
* Updated test environment setup to install all required package
versions for each Transformers configuration.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-08-21 16:21:46 +05:30
Keval Morabia 38826351eb Bump transformers dependency to >=4.57,<5.15 (#2050)
Transformers Dependency bump - Drop 4.56 (still support 4.57 with
deprecation note) and extend to 5.14 (nemo:26.08 ships with this
version)

- CI tests now use transformers 5.14
- Manually ran `tests/{gpu_megatron,examples/megatron_bridge}` in
nemo:26.08.rc3

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Compatibility
- Updated supported Transformers versions to 4.57 through 5.14.
- Qwen3-VL models are now available without version-based restrictions.

## Documentation
- Updated the changelog to reflect the new minimum Transformers version
and upcoming removal of Transformers 4.x support.

## Testing
- Updated validation and test configurations for the supported
Transformers versions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-08-04 01:43:32 +05:30
Keval MorabiaandClaude Opus 4.8 42458def24 ci: fix torch_trt on torch 2.13; default unit tests to torch 2.13; add example allow-failure hatch (#1951)
### What does this PR do?

Type of change: Bug fix + CI / infra

torch 2.13 + torchvision 0.28 were published to PyPI on 2026-07-08 and
broke the `onnx (torch_trt)` example job (which had passed the day
before). This PR fixes that break and hardens CI against the next one:

1. **Fix `torch_trt` (torch/torchvision/torch-tensorrt trio pin).** The
base install pulled torch 2.13 / torchvision 0.28, then `torch-tensorrt
2.12.1` downgraded torch back to 2.12 but left torchvision at 0.28
(which pins `torch==2.13`) — breaking `import`.
`examples/torch_trt/requirements.txt` now caps
`torch-tensorrt>=2.4.0,<2.13` + `torchvision<0.28` so the trio stays
consistent (also protects direct `pip install -r` users).
2. **Unit tests default to torch 2.13.** `noxfile.py` gains `torch_213`
(`torchvision~=0.28.0`); the required `linux`/`windows` jobs and the
multi-version Python spread (3.10/3.11/3.13/3.14) now run torch 2.13,
with torch 2.8–2.12 kept as back-compat legs on Python 3.12.
3. **Per-example allow-failure escape hatch.**
`_example_tests_runner.yml` gains an `allow_failure` input; when set, a
**test-run** failure is surfaced as a `::warning::` via
`continue-on-error` instead of blocking the PR. `example_tests.yml`
derives it per example from the repo variable
**`ALLOW_FAILURE_EXAMPLE_TESTS`** (comma-separated example names,
comma-wrapped so `onnx` ≠ `torch_onnx`). Future breakages can be
quarantined by updating the variable — no code change / PR required.

### Testing

- Ran the new default unit session locally in an isolated uv venv (torch
**2.13.0**+cu130, torchvision **0.28.0**+cu130, transformers
**5.12.1**):
  ```
  nox -s "unit-3.12(torch_213, tf_latest)"
  => 2813 passed, 15 skipped, 1786 warnings in 250.37s
  ```
- Verified the allow-failure hatch: with
`ALLOW_FAILURE_EXAMPLE_TESTS=torch_trt`, the (previously failing)
`torch_trt` job reports success with a warning and does not block the
required example check. `vars` is re-read on each job attempt, so
"Re-run failed jobs" picks up the variable without a fresh trigger.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (CI-only; older torch versions
still covered)
- If you copied code from any other sources or added a new PIP
dependency: N/A (no new dependency; only version caps)
- Did you write any new necessary tests?: N/A (CI configuration change)
- Did you update Changelog?: N/A (CI infra, no user-facing API change)
- Did you get Claude approval on this PR?: ❌ (pending — will run
`/claude review`)

### Additional Information

The `ALLOW_FAILURE_EXAMPLE_TESTS` repo variable can be cleared for
`torch_trt` now that the requirements pin lands the real fix; keep it as
the standing escape hatch for future example breakages.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Example test jobs can now be configured to “allow failure” without
failing the workflow; when enabled, a warning annotation is emitted.
* **Bug Fixes**
* CI unit-test and GPU-test configurations were refreshed (including a
reduced timeout for the `gpu_megatron` job).
* **Tests**
* Updated unit-test coverage to use the newest Torch 2.13-based setup by
default, with back-compat retained where applicable.
* **Documentation**
* Added inline guidance for how the allow-failure examples list is
specified.
* **Chores**
* Refreshed `torch_trt` example dependency constraints to improve
compatibility.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-09 19:01:37 +05:30
Keval MorabiaandClaude Opus 4.8 795c589428 CI: CUDA build/test hygiene + fix Puzzletron Nemotron test failures (#1901)
### What does this PR do?

Type of change: Bug fix (CI / tests)

- **Fix Nemotron nightly failures:** install `mamba_ssm`/`causal-conv1d`
from PyPI releases instead of git `main` (avoids the broken
`apache-tvm-ffi 0.1.12` that crashes on import).
- **Speed up CUDA builds:** set `TORCH_CUDA_ARCH_LIST=12.0` (runner's
sm_120) in the GPU/example/regression workflow container env instead of
the image's ~6 archs.
- **Make unit tests CPU-only:** force CUDA off in the nox `unit` env and
skip JIT-compiling CUDA extensions when no GPU is usable; move the two
GPU-/`mamba_ssm`-requiring unit tests to `tests/gpu`.
- **Harden example tests against HF flakes:** capture subprocess output
and retry transient HuggingFace access errors (5xx / rate-limit /
connection).
- **Skip Blackwell-flaky sharded-state-dict tests:**
`test_homogeneous_sharded_state_dict` and `test_regular_state_dict[320]`
intermittently hit a CUDA illegal-memory-access on the sm_120 runner
that poisons the CUDA context and cascades timeouts; gate them behind a
reusable `skip_flaky_on_blackwell` marker (still run on non-Blackwell
GPUs).
- **Bump slow test timeout:** `test_prune_minitron_vlm` → 360s for the
2-GPU nightly.

### Testing

- CI tests on this PR pass (1-gpu)
- Manually triggerred 2-gpu test:
- GPU:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/28774553356
- Examples:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/28774556693
- Regression:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/28771197586

### Additional Information

- Backward compatible: N/A (CI/tests only)
- New dependency: N/A
- Changelog: N/A (CI/test infra)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-06 13:48:29 +05:30
Keval MorabiaandClaude Sonnet 4.6 33bfa8b1fe CI/Dev env bump (#1818)
### What does this PR do?

Type of change: chore

Bumps CI/dev tooling and test containers.

**Container bumps**
- NeMo test containers → 26.06
- TRT-LLM container → 1.3.0rc19
- transformers max version → 5.12

**Dev tooling bumps**
- ruff bump 0.12.11 → 0.15.18
- mypy 1.17.1 → 2.1.0: enable new defaults (`local_partial_types`,
`strict_bytes`); fix/narrow the errors newly surfaced by mypy 2.0 in 4
modules (rather than blanket-suppressing them); remove 2 stale `# type:
ignore` comments
- pre-commit 4.3.0 → 4.6.0
- sphinx 8.1 → 9.1 + sphinx-rtd-theme 3.0 → 3.1: add `suppress_warnings
= ["ref.python"]` to fix cross-reference ambiguity error new in sphinx
9.x
- trl fix for newly released 1.7 version

**Bug fixes surfaced by the bumps**
- sparsity (weight): make the weight mask DTensor-aware under FSDP. The
transformers→5.12 bump routes the HF Trainer FSDP optimizer-state save
through torch's DTensor-based `get_optimizer_state_dict`, which
triggered `aten.mul.Tensor got mixed torch.Tensor and DTensor` in the
dynamic `weight` getter. The mask is now distributed to the weight's
mesh/placements before masking, cached, and rebuilt only when the
sharding changes (invalidated on `set_mask`). Fixes the `llm_sparsity`
example test.

### Testing

- `pre-commit run --all-files` ✅ (including mypy 2.1.0)
- `nox -s docs` ✅
- `tests/unit/torch/sparsity` + `tests/unit/torch/nas` ✅
- `llm_sparsity` GPU example test (FSDP path) verified in CI

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ✅
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Summary

* **Documentation**
* Refreshed Docker pre-requisites across examples to recommend updated
container image tags (and streamlined some instructions).
* **Bug Fixes**
* Improved sparse weight mask handling for DTensor/FSDP by aligning and
caching distributed masks.
  * Made TensorRT engine byte retrieval return immutable `bytes`.
* Reduced Sphinx cross-reference warnings and tuned Transformers
compatibility warning thresholds.
* **Tests**
  * Increased default unit test timeout on Windows runners.
* **Chores**
* Updated CI workflow container tags and refreshed linting/typing/docs
version pins, plus related mypy configuration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-27 01:00:25 +05:30
Grzegorz K. KarchandKeval Morabia b98a59557a Add vLLM-based runtime statistics for subblock latency measurement (#1358)
### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Runtime-based latency optimization: collect vLLM-measured inference
latency to constrain optimization.

* **Configuration**
* New runtime config/template for Llama-3.1-8B pruning (runtime stats
enabled, NCCL timeout templating, MIP target-latency).
* Validation sample defaults adjusted (one flow: 128 → 8; runtime flow
uses 128).
  * Human constraint key renamed to target_latency_seconds.

* **Documentation**
* README section describing runtime-based latency optimization setup and
usage.

* **Tests**
  * Added GPU end-to-end test for runtime stats collection.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1358?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Grzegorz Karch <gkarch@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-06-08 19:22:55 +00:00
kinjalpatel27andKeval Morabia 7aa0c95646 Add tests/gpu_vllm (#1517)
### What does this PR do?

Type of change: new tests

This PR adds unit tests for vLLM fakequant, specifically testing code in
`modelopt/torch/quantization/plugins/vllm.py`


### Testing

```
pytest tests/gpu_vllm/torch/quantization/test_vllm_dynamic_modules.py -sv
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ✅ 

### Additional Information


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Added comprehensive GPU vLLM test suite with end-to-end quantization
checks and fixtures for TinyLlama, TinyQwen3-MoE, and DeepSeek V3;
includes helpers to build tiny DeepSeek V3 models.

* **Chores**
* Updated GPU CI to use explicit container image references, added a
GPU-focused test session, and adjusted test-run setup for vLLM.

* **Documentation**
  * Documented new GPU test directory in contributing guide.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1517?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-29 20:47:23 +00:00
Keval Morabia eb5ed2df68 [CI] Bump torch, transformers and dev containers to latest (#1554)
- Transformers upper bound bumped from `<5.8` to `<5.10`
- Enable torch 2.12 CICD testing
- Bump TRT-LLM container to `1.3.0rc16` (transformers 5.5)
- Use pytorch and tensorrt 26.04 containers in CICD

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Updated CI test container images and targeted Torch version across
workflows; adjusted release CI job to use the newer torch config.
* Broadened Transformers constraint in project metadata and test/dev
pins.
* Removed strict transformers pins from example requirements and lifted
a compression dependency cap.
* Raised the import-time Transformers version threshold for
compatibility warnings.

* **Tests**
* Refactored a GPU test to collect and report validation errors and
updated numeric expected baselines.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1554?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-29 10:37:00 -07:00
Jenny Chen 4b270f0ea6 Support Mixed precision & Static MSE in MCore; Nemotron Super v3 NVFP4 recipe (#1521)
### What does this PR do?

Type of change: New features + Bug fixes

Mixed Precision and MSE support in MCore PTQ
- support mixed precision export in MCore by detecting mixed precision
layers in HF Quant Config
- Restore static quantizer in MCore checkpoint restore as `NVFP4QTensor`
(not TensorQuantizer which can call max calibrate. we want to skip max
calibrate for static quantizer during restore) --> fixes bug during
MCore export for MSE
- Fix dynamic block quantizer detection when `block_sizes` is
dict-backed.
- Add a YAML quantization recipe that roughly mirrors Nemotron 3 Super
NVFP4 `hf_quant_config.json`
  Export bug fixes     
- copy .py files properly from original HF ckpt (for reasoning parser
etc)
## Super recipe

Mirrors the published nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
hf_quant_config.json:
- MoE routed experts (mixer.experts.<N>.{up,down}_proj): NVFP4 W4A4
weight MSE, group_size 16
- MoE shared experts (mixer.shared_experts.{up,down}_proj): FP8
per-tensor
- Mamba mixer linears (mixer.{in,out}_proj): FP8 per-tensor
- KV cache:                                                    FP8
rest: not quantized

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
Tested on Nemotron model
### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added NVFP4 (4-bit) quantization checkpoint restore and export support
for Megatron-Core models
  * Added tokenizer file export capability in model checkpoints
* Extended quantization support for expert-parallel distributed training
* Introduced new PTQ recipes for Nemotron-3-Super models with
mixed-precision quantization

* **Bug Fixes**
* Fixed FP8 and FP4 hardware compatibility detection on non-CUDA systems
* Improved offline Hugging Face Hub access handling with better error
messaging
  * Enhanced calibration validation for mixture-of-experts models
  * Fixed amplitude maximum validation for static block quantizers

* **Documentation**
  * Updated expert weight quantization configuration documentation

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1521?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Jenny Chen <jennifchen@nvidia.com>
2026-05-29 16:23:09 +00:00
Keval Morabia 72799ab551 Pin transformers<5.8 to avoid CI failures (#1395)
Transformers now releases new versions weekly/biweekly causing frequest
breaking changes in our tests (We always pin a max version for release
branches so thats unaffected). Hence pinning max version to 5.7 in main
branch. We will bump transformers monthly / with torch bump instead of
always using latest one

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-06 00:37:48 +05:30
Keval Morabiaandcoderabbitai[bot] 70546bdd6a Enable Python 3.14 wheel support to unblock NGC PyTorch container testing on Ubuntu 26.04 + Python 3.14 (#1386)
Ubuntu 26.04 is here and very soon, NVIDIA PyTorch containers will ship
with Python 3.14 requiring us to enable untested support to unblock them

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
  * DFlash offline speculative decoding training
  * MXFP4→NVFP4 weight conversion support
  * Shared hidden-state dump utilities
  * Updated DeepSeek PTQ calibration defaults

* **Chores**
  * Added Python 3.14 support; updated Python requirement to <3.15

* **Documentation**
  * Updated installation documentation for Python version compatibility

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-05-04 21:56:04 +00:00
Keval Morabia c51c1762b3 fix: prevent gh-pages repo bloat from doc preview artifacts (#1309)
### What does this PR do?

Type of change: Bug fix

Fixes gh-pages branch bloat that grew from ~26 MB to ~441 MB in four
weeks (nvbug 6099503). Three compounding causes were identified and
addressed:

1. **Sphinx `.doctrees/` cache published to gh-pages** — `sphinx-build`
was writing its build cache inside `build/html/` which was then uploaded
verbatim. Accounts for ~3.3 GB uncompressed across history.
2. **`JamesIves/github-pages-deploy-action` appending a commit on every
push** — main-site files accumulated forever with `single-commit: false`
(default).
3. **PR preview deploying on every `synchronize` event for all PRs** —
`rossjrw/pr-preview-action` re-deployed the full site for every push to
any PR regardless of whether docs changed (e.g. PR #1128 triggered 64
preview deploys × ~11 MB each).

Changes:
- Pass `-d /tmp/doctrees` to `sphinx-build` so `.doctrees/` is never
written into `build/html/`
- Add `paths: [docs/**, modelopt/**]` filter to `pull_request` trigger
so the docs workflow only runs on PRs that touch docs or source code
- Set `single-commit: true` on the deploy action so main-site pushes
squash into one commit
- Deduplicate docs build: `deploy-preview` now downloads the artifact
from `build-docs` instead of running a second `sphinx-build`
- Set `retention-days: 1` on the artifact since it is only needed for
the duration of the workflow run

The one-time cleanup (force-push squashed orphan to gh-pages) was
already applied separately — repo is now ~59 MB for a full clone vs ~441
MB before.

### Usage

N/A — CI/workflow change only.

### Testing

- Workflow logic reviewed manually.
- The one-time cleanup was verified: `git rev-list --objects
--disk-usage origin/gh-pages` now reports ~28 MB; full clone is ~59 MB.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information

nvbug 6099503

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Optimized documentation build and deployment workflow in CI/CD
pipeline.
* Improved pull request documentation preview handling with faster build
timeouts and refined artifact management.
* Enhanced GitHub Pages deployment configuration for better consistency.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-21 20:19:08 +00:00
Keval MorabiaandClaude Sonnet 4.6 3d0f0db49e [CI] Replace tox with nox, use nemo:26.04 for megatron tests, and simplify CI workflows (#1286)
### What does this PR do?

Type of change: New feature / infrastructure improvement

Follow-up to #1285 for correct CI test environment for megatron based
tests

Replaces `tox` + `tox-current-env` with `nox` for all test, lint, docs,
and wheel build sessions. The primary motivation was that
`tox-current-env` is incompatible with uv venvs in NGC containers (e.g.
NeMo's `/opt/venv`) — it picks the system Python via
`sys._base_executable` instead of the container's venv Python which has
megatron packages pre-installed.

Key changes:
- **`noxfile.py`** replaces `tox.ini` with GPU, CPU unit,
partial-install, pre-commit, docs, and wheel sessions
- **GPU sessions** use `venv_backend="none"` (run directly in container
env) and `python -m pip/pytest` to avoid PATH mismatches
- **uv** is set as the default venv backend (if available) for CPU
sessions (faster installs)

Also includes CI workflow simplifications:
- **`_pr_gate.yml`** new reusable workflow centralizing file-change
detection + linux-check wait logic (was duplicated across 3 workflow
files)
- **Collapsed pr/non-pr job pairs** into single jobs with conditional
`runs-on` in `gpu_tests.yml`, `example_tests.yml`,
`regression_tests.yml`
- **Collapsed `multi-py` / `multi-torch` / `multi-transformers`** into a
single `multi-version` matrix job in `unit_tests.yml`
- **PR path filtering** for unit test secondary jobs (multi-version,
launcher, partial-install) — skipped if no relevant files changed
- **Fixed schedule/workflow_dispatch skipping** — jobs with `needs:
[pr-gate]` were incorrectly skipped when all pr-gate internal jobs were
skipped; fixed by making the gate job always run
- **multi-version, launcher, partial-install** now also run on
`schedule` / `workflow_dispatch`

### Usage

```bash
python -m pip install nox uv                                                    # install nox and uv (once)
nox -l                                                                          # list all sessions
nox -s gpu_megatron                                                             # run a GPU session (inside container)
nox -s "unit-3.12(torch_211, tf_latest)"                                        # run a specific unit test combination
nox -s "unit-3.12(torch_211, tf_latest)" -R                                     # force-recreate venv (e.g. after dep changes)
COVERAGE_PROCESS_START=pyproject.toml nox -s "unit-3.12(torch_211, tf_latest)"  # with coverage
```

### Testing
- Ran `nox -l` to verify all session names
- Ran `gpu_megatron` session locally inside NeMo container — confirmed
it uses `/opt/venv/bin/python` correctly
- Manually triggered nightly-runs:
- Unit:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608013657
- GPU:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608018763
- Examples:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608017322

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A — CI infrastructure only
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ (added `nox`
and `uv` to `dev-test`, both Apache-2.0)
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — no user-facing changes

### Additional Information
Supersedes the tox-current-env workaround in the parent branch.

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-18 16:56:02 +00:00