719 Commits
Author SHA1 Message Date
c78e654744 Skip Softmax diffusion export (#1269)
### What does this PR do?

Type of change: New Feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

Adds HuggingFace `config.json` export of skip-softmax sparse-attention
calibration for diffusion pipelines (e.g. Wan 2.2), on top of the base
skip-softmax work.

- **`_export_diffusers_checkpoint`** walks every `nn.Module` component
of a diffusers pipeline, calls `export_sparse_attention_config`, and
writes the result into that component's `config.json` under the
`sparse_attention_config` key. The sparse config lives **only** in
`config.json` — there is no standalone `sparse.yaml`.
- **`export_sparse_attention_config`** emits a `config_groups` schema
where each algorithm's parameters are nested inside its own group; only
`config_groups` and `producer` are top-level:
- skip-softmax group → `algorithm: "skip_softmax"`, `targets`, `ignore`
(layers kept dense — e.g. cross-attention + first/last blocks),
`initial_disabled_steps` (opt-in, user-set; emitted only when `> 0`),
`threshold_scale_factor` (`a * exp(b * target_sparsity)`), and
`target_sparsity`.
- N:M group → `algorithm: "sparse_softmax"` with
`sparsity_n`/`sparsity_m`, `dense_sink_tokens`, `dense_recent_tokens`
flattened into the group.
- **Deploy reader**
(`modelopt/torch/sparsity/attention_sparsity/plugins/sparse_attn_config.py`)
reads these per-group params back, keeping the export↔load round-trip
consistent.
- **Example wiring**:
`examples/diffusers/sparsity/wan22_skip_softmax.py` gains
`--export-dir`, `--skip-softmax-threshold`, and
`--initial-disabled-steps`. `--export-dir` runs
`export_hf_checkpoint(pipe, export_dir=...)` after calibration.
- Updated `CHANGELOG.rst`.

### Usage

```bash
python examples/diffusers/sparsity/wan22_skip_softmax.py \
    --model-path Wan-AI/Wan2.2-T2V-A14B-Diffusers \
    --calibrate --target-sparsity 0.5 --calib-size 4 \
    --initial-disabled-steps 5 \
    --export-dir ./wan22_skip_softmax_ckpt
```

Resulting layout — a `config.json` per component, **no `sparse.yaml`**:

```
wan22_skip_softmax_ckpt/
├── transformer/config.json        # carries sparse_attention_config
├── transformer_2/config.json      # carries sparse_attention_config
├── vae/ …  text_encoder/ …  tokenizer/ …  scheduler/ …
└── model_index.json
```

A representative `config.json` entry for a diffusion transformer:

```json
"sparse_attention_config": {
  "config_groups": {
    "group_0": {
      "algorithm": "skip_softmax",
      "targets": ["WanAttention"],
      "ignore": ["blocks.0.attn1", "blocks.0.attn2", "…"],
      "initial_disabled_steps": 5,
      "threshold_scale_factor": {
        "formula": "a * exp(b * target_sparsity)",
        "prefill": {"a": 1443.49, "b": 4.30}
      },
      "target_sparsity": {"prefill": 0.5}
    }
  },
  "producer": {"name": "modelopt", "version": "0.45.0..."}
}
```

The N:M variant adds a second group:

```json
"group_1": {
  "algorithm": "sparse_softmax",
  "targets": ["WanAttention"],
  "sparsity_n": 2, "sparsity_m": 4,
  "dense_sink_tokens": 0, "dense_recent_tokens": 64
}
```

### Testing

- `tests/examples/diffusers_sparsity/test_sparsity.py`: baseline /
triton-baseline / fixed-threshold runs of the Wan 2.2 example, plus a
Python-API calibrate → **export** test asserting the nested
`sparse_attention_config` (`threshold_scale_factor`, `target_sparsity`,
`ignore`, `initial_disabled_steps`) and the absence of any
`sparse.yaml`.
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attention_conversion.py`
and `test_sparse_attn_config.py`: unit coverage of the per-group export
schema and the deploy-reader round-trip (writer nests → reader reads
from groups → internal mtsa config unchanged).
- Validated end-to-end on Wan 2.2 T2V-A14B: full 4-prompt / 40-step /
81-frame calibration; the exported checkpoint carries the nested schema
in both `transformer` and `transformer_2` `config.json`, and runtime
measurement shows ~47–49% tile sparsity at a 0.5 target.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ The exported
`sparse_attention_config` schema was renamed and nested per-group during
0.45.x development, and the loader reads only the new layout —
checkpoints exported by earlier 0.45.x builds must be re-exported. No
released version is affected. <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 16:02:41 -07:00
Frida Hou 16d562a0bf Refactor local_hessian onto shared MSE flow + fused-MoE expert support (#1578)
### What does this PR do?

Type of change: Bug fix + new feature (fused-MoE coverage) <!-- Use one
of the following: Bug fix, new feature, new example, new tests,
documentation. -->

**Dusted off and refactored `local_hessian_calibrate`** to align with
the new MSE calibration flow, fix latent
drift, extend coverage to fused-MoE experts, and decouple
module-specific handling behind a clean extension point.

  Core changes (`modelopt/torch/quantization/model_calib.py`):
- Extracted a shared `_mse_calibrate_weights` helper now used by both
`mse_calibrate` and `local_hessian_calibrate`
- Replaced the monolithic `LocalHessianHelper` + bespoke per-weight loop
with a small `_LocalHessianAccumulator` (lazy fp32 buffer, freed after
building the error func), removing ~200 lines of duplicated scale-search
logic and all manual `cuda.synchronize`/`empty_cache` bookkeeping. The
`XᵀX` GEMM accumulates in fp32 to avoid bf16/fp16 precision loss.
- Removed dead NVFP4-static promotion (now handled inside
`max_calibrate`).
- **Fused-MoE expert support:** per-expert Hessians captured from each
expert's routed activations, keyed by
`id(weight_quantizer)` so dense and per-expert paths share one
calibration loop. Never-routed experts / non-eager kernels / `cin` not
divisible by `block_size` / registered backends fall back to plain MSE
(with an eager, module-named warning).

  Decoupling (zero module-type-specific code in `model_calib.py`):
- Added `QuantModule.register_calibration_input_hooks(callback)` — the
activation-side counterpart to `iter_weights_for_calibration`. Base
default is a no-op; `QuantLinearConvBase` pairs the weight quantizer
with the
forward input (linear only), and `_QuantFusedExperts` (in
`plugins/huggingface.py`) owns the per-expert pairing via
`_current_expert_idx`. Any future module type gains local-Hessian
support by implementing this one method.


### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
- new `tests/unit/torch/quantization/test_local_hessian.py` (accumulator
math/shape/dtype, dense end-to-end, backend-skip, block-size guard) and
a per-expert MoE test in `test_fused_experts.py`.
- Behavior-preserving check: refactored branch produces
**bit-identical** dense weight scales to `main` (216/216 tensors) on
Qwen3-8B with the fp32 accumulation neutralized to isolate the
structural change.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅  <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Introduced Hessian-weighted calibration method for enhanced weight
quantization refinement in mixed-precision models.
* Enhanced MSE calibration to support custom per-quantizer error
functions for specialized calibration workflows.

* **Tests**
* Added comprehensive unit tests for local Hessian calibration
validation across dense models and fused-expert architectures.
* Included tests for custom error function integration and fallback
behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
2026-06-08 12:29:00 -07:00
Grzegorz K. KarchandKeval Morabia b98a59557a Add vLLM-based runtime statistics for subblock latency measurement (#1358)
### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Runtime-based latency optimization: collect vLLM-measured inference
latency to constrain optimization.

* **Configuration**
* New runtime config/template for Llama-3.1-8B pruning (runtime stats
enabled, NCCL timeout templating, MIP target-latency).
* Validation sample defaults adjusted (one flow: 128 → 8; runtime flow
uses 128).
  * Human constraint key renamed to target_latency_seconds.

* **Documentation**
* README section describing runtime-based latency optimization setup and
usage.

* **Tests**
  * Added GPU end-to-end test for runtime stats collection.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1358?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Grzegorz Karch <gkarch@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-06-08 19:22:55 +00:00
Keval MorabiaandClaude Opus 4.8 01415c2788 fix(llm_eval): repair test_qwen3_eval_fp8 end-to-end (#1650)
### What does this PR do?

Type of change: Bug fix

`tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8` was
silently passing while its evals crashed, then began failing as a
timeout. This repairs the whole pipeline:

- **lm_eval `IndexError` (root cause):** TRT-LLM KV-cache prefix reuse
returns truncated `context_logits` for shared-prefix requests (e.g.
hellaswag's one-context / many-endings), which breaks `parse_logprobs`.
Add an `enable_kv_cache_reuse` flag to `modelopt.deploy.llm.LLM`
(default `True`, unchanged) and disable it for the eval deployment so
full-length context logits are returned.
- **Silent CI green:** `python eval.py | tee result.txt` returns `tee`'s
exit code, so a crashing eval was masked. Add `set -o pipefail` to
`huggingface_example.sh` so failures fail the test.
- **Long-prompt overflows:** with the tiny test model's toy tokenizer,
gsm8k/MMLU prompts exceed `max_seq_len`. Bump test
`max_position_embeddings` to 8192, skip MMLU prompts that don't fit even
at zero-shot, and add an MMLU sample limit (`--mmlu_limit`).
- **human-eval build failures:** install with `--no-build-isolation`
(`pkg_resources` is absent in pip's isolated build env), patch its
malformed `console_scripts` entry point, and pin the clone.
- **Cleanups:** gate the post-quant `run_tensorrt_llm.py` smoke test
behind the `quant` task (eval tasks deploy on their own; ~45s saved for
eval-only runs); replace the SIGPIPE-prone serve-readiness `tail -f |
while` with a poll loop (required under `pipefail`).

### Usage

N/A — example/test fix.

### Testing

All four eval tasks verified end-to-end in the CI container (TRT-LLM
1.3.0rc17, RTX 6000 Ada): lm_eval (hellaswag + gsm8k), MMLU, and
simple_eval (humaneval) all complete with exit 0 and no
`IndexError`/overflow. Cold full run ≈ 340s on this GPU.

CI test on 2-gpu:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/27154417497/job/80153551154

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (new `enable_kv_cache_reuse`
defaults to current behavior; new script flags are optional)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new
dependencies)
- Did you write any new necessary tests?: N/A (fixes and strengthens an
existing test)
- Did you update Changelog?: N/A (bug fix to examples/tests)
- Did you get Claude approval on this PR?: ❌ (pending)

### Additional Information

The full test runs ~340s on an RTX 6000 Ada; CI runners are historically
slower, while `@pytest.mark.timeout` is set to 600 — worth watching the
first CI run and bumping if it's close.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added an option to limit MMLU evaluation length.

* **Bug Fixes**
* Disabled KV-cache prefix reuse for evaluations needing per-token
context logits to prevent truncated/incorrect logprobs.
* Skip examples whose prompts remain too long; warn and report accuracy
as NaN if all examples are skipped.

* **Chores / Scripts**
* Improved example scripts for reproducible installs, patched entry
point handling, pipeline failure detection, conditional test invocation,
polling-based log wait, and a new CLI flag for MMLU limits.

* **Tests**
* Increased timeout and prompt headroom; capped MMLU smoke tests for
speed.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 18:11:51 +00:00
Wei-Ming Chen 1f4a489b0f Adds AutoQuant support for VLM / Qwen3.5-Qwen3.6 style models (#1381)
### What does this PR do?

Type of change: new feature, bug fix, new tests

### Details

- Enables AutoQuant search over fused MoE expert containers by
snapshotting/restoring their per-expert quantizers.
- Adds Qwen3.5/3.6 linear-attention grouping rules so fused deployment
layers keep compatible quant formats.
- Supports `w4a16_nvfp4` as an AutoQuant search format.
- Preserves disabled AutoQuant layer patterns in generated configs while
allowing selected modules like `lm_head` to override default disables.
- Keeps recipe-mode and AutoQuantize VLM paths on the outer CausalLM so
Qwen3.5/3.6-MoE `lm_head` remains visible.
- Skips `parent_class`-scoped quant config entries during AutoQuant bare
quantizer matching, preventing class-scoped global entries from
last-match overriding every selected module.
- Adds temporary hardcoded Qwen/VLM AutoQuant disabled-layer patterns in
`hf_ptq.py` with a TODO to refactor into the config system.

### Usage

```bash
python examples/llm_ptq/hf_ptq.py \
  --pyt_ckpt_path <model_path> \
  --qformat fp8,w4a16_nvfp4 \
  --auto_quantize_bits 5.0 \
  --auto_quantize_cost_model active_moe \
  --auto_quantize_checkpoint <autoquant_state.pt> \
  --export_path <output_dir>
```

### Testing

- `/Users/weimingc/miniconda3/envs/modelopt/bin/python -m pytest
tests/unit/torch/quantization/test_autoquant.py::test_get_auto_quantize_config_keeps_selected_lm_head_enabled
tests/unit/torch/quantization/test_config_validation.py::TestMatchQuantizerCfg::test_parent_class_scoped_entries_are_ignored_for_bare_autoquant_lookup`
- `/Users/weimingc/miniconda3/envs/modelopt/bin/python -m pytest
tests/unit/torch/quantization/test_autoquant.py
tests/unit/torch/quantization/test_config_validation.py -k "not
data_parallel"` (`120 passed, 1 deselected`)
- `/Users/weimingc/miniconda3/envs/modelopt/bin/python -m py_compile
examples/llm_ptq/hf_ptq.py modelopt/torch/quantization/algorithms.py
modelopt/torch/quantization/_auto_quantize_cost.py
tests/unit/torch/quantization/test_autoquant.py
tests/unit/torch/quantization/test_config_validation.py`
- Full local affected-file pytest without `-k "not data_parallel"` only
failed `test_data_parallel_auto_quantize` because this local sandbox
cannot bind a free socket (`PermissionError: Operation not permitted`).
- Ran Qwen3.6 35B AutoQuant e2e with `fp8,w4a16_nvfp4` and exported a
checkpoint.
- Verified exported checkpoint loads in vLLM nightly without local
patches.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added w4a16_nvfp4 quantization format and optional cost-exclusion
patterns for AutoQuantize.

* **Improvements**
* Safer multimodal/VLM handling and AutoQuantize now runs on the full
outer model when applicable.
* Better fused-MoE support, more accurate weight accounting, and refined
attention-grouping for improved quantization choices.
  * Dynamic layer-disabling support for targeted disables.

* **Tests**
* New unit tests covering cost-model exclusions, fused-MoE accounting,
and config selection.

* **Documentation**
  * Updated cost-constraint example to show exclusion-pattern usage.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-06-08 10:06:12 -07:00
Keval MorabiaandClaude Opus 4.8 54ce4e09d8 Add Quantization Aware Distillation (QAD) to Megatron-Bridge example (#1600)
### What does this PR do?

Type of change: new example

**Note:** This is **part 2 of 4** (builds on #1589):

- **Part 1 (#1589):** Megatron-Bridge `quantize.py` + `export.py`
support and tests.
- **Part 2 (this PR):** extend `distill.py` for quantization-aware
distillation (QAD) — load a quantized Megatron checkpoint as the
student.
- **Part 3:** https://github.com/NVIDIA/Model-Optimizer/pull/1601
- **Part 4:** repeat the NVFP4 + QAD experiments on a non-Nemotron
model.

Extends `examples/megatron_bridge/distill.py` to initialize the student
from a **Megatron checkpoint** (a quantized checkpoint from
`quantize.py`, or a pruned one) via `--student_megatron_path`, enabling
**Quantization Aware Distillation (QAD)**:

- `--student_hf_path` still builds the student architecture;
`--student_megatron_path` supplies the (optionally quantized) weights.
- For a quantized checkpoint, the ModelOpt quantize mode + base weights
are restored onto the **plain student before the knowledge-distillation
conversion** (`restore_sharded_modelopt_state` is a no-op once a model
is already converted), so the distilled checkpoint stays exportable as a
quantized model with `export.py`.

**Upstream dependency / workaround:** `DistillationProvider.provide()`
has no seam to transform the student before the KD conversion, so this
patches `provide()` at the class level (via an `id()`-keyed registry,
because the provider proxies instance-attribute assignment to its
teacher once the teacher is set). A companion Megatron-Bridge PR adds a
first-class `DistillationProvider.student_pre_conversion_hook`; from
nemo:26.06 onwards the workaround should be removed and replaced with
that hook (a removal note in `distill.py` documents exactly how).

### Usage

```bash
# 1) PTQ -> quantized Megatron checkpoint (part 1)
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-8B --quant_cfg fp8 --tp_size 2 \
    --export_megatron_path /tmp/Qwen3-8B-FP8-megatron

# 2) QAD: distill the quantized student from the unquantized teacher
torchrun --nproc_per_node 8 distill.py \
    --teacher_hf_path Qwen/Qwen3-8B \
    --student_hf_path Qwen/Qwen3-8B \
    --student_megatron_path /tmp/Qwen3-8B-FP8-megatron \
    --data_paths 1.0 tokenized/data_text_document \
    --train_iters 1000 --output_dir /output/qwen3_8b_qad

# 3) export the distilled quantized checkpoint (part 1)
torchrun --nproc_per_node 1 export.py \
    --hf_model_name_or_path Qwen/Qwen3-8B \
    --megatron_path /output/qwen3_8b_qad/checkpoints \
    --export_unified_hf_path /tmp/qwen3_8b_qad_fp8_hf
```

### Testing

`tests/examples/megatron_bridge/test_qad.py` (validated on a 2-GPU NeMo
`26.04` container): quantize a tiny Qwen3 at TP=2 → QAD distill from the
quantized student → `export.py` to a unified HF checkpoint, asserting
`hf_quant_config.json` is written (proves the quantize mode survived
QAD). Includes a commented-out vLLM deployment check, validated locally
(full flow passes; vLLM loads the export as `quantization=modelopt`).
Existing normal/Puzzletron distillation tests still pass.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A (new example feature; default
behavior unchanged when `--student_megatron_path` is not set)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new
dependencies)
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅

### Additional Information

Depends on a companion Megatron-Bridge PR adding
`DistillationProvider.student_pre_conversion_hook` (the upstream
replacement for the class-level `provide()` workaround). The Nemotron-3
tutorial NVFP4 + QAD experiments ship in part 3.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Quantization Aware Distillation (QAD) workflow to recover accuracy of
quantized Megatron students and distill from quantized checkpoints.
* CLI option to initialize a distillation student from a Megatron
checkpoint and a structure-only load path for bridging.

* **Documentation**
* Expanded runnable quantize → QAD → export guidance and best-practice
tips.

* **Tests**
  * End-to-end test validating quantize → QAD → export artifacts.

* **Chores / UX**
* Clearer rank-aware messages, improved tokenizer padding handling, and
more consistent export behavior (fixed export dtype).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-05 21:28:25 +00:00
Shengliang Xu dbdff11a7f Drive PTQ example qformat choices from preset YAMLs (hf_ptq, multinode_ptq, megatron_bridge) (#1525)
### What does this PR do?

Type of change: Refactor

Replace the hardcoded `QUANT_CFG_CHOICES` / `KV_QUANT_CFG_CHOICES` dicts
in the PTQ
example scripts with a small `_load_preset_cfg_choices()` helper that
discovers the
available qformat names by listing
`modelopt_recipes/configs/ptq/presets/{model,kv}/`
and **eagerly loads every preset YAML into a plain dict at import** via
the existing
`load_config(...,
schema_type=QuantizeConfig).model_dump(exclude_unset=True)` path.
The directory listing becomes the source of truth for the `--qformat` /
`--kv_cache_qformat` CLI vocabulary.

> Note: an earlier revision used a lazy, copy-on-access `Mapping`. That
was overkill
> for these example scripts — the previous `mtq.*_CFG` module constants
were
> themselves eagerly-loaded shared dicts, and every call site that
mutates a config
> already deepcopies first — so it is now a plain eager dict. A lazy
variant can be
> reintroduced later if import time ever matters.

**Scope.** Three scripts carried the same hardcoded tables and all three
are migrated:

- `examples/llm_ptq/hf_ptq.py` — `--qformat` / `--kv_cache_qformat`.
- `examples/llm_ptq/multinode_ptq.py` — `--qformat` /
`--kv_cache_qformat`.
- `examples/megatron_bridge/quantize.py` — `--quant_cfg` /
`--kv_cache_quant`
  (still also accepts any full `mtq.config.choices` name).

All three scripts share the discovery helper (`load_quant_cfg_choices`),
the canonical
alias table (`QFORMAT_ALIASES`), the KV disable sentinel
(`KV_CACHE_NONE`), and the
ready-built `QUANT_CFG_CHOICES` / `KV_QUANT_CFG_CHOICES` mappings via
the new
**`modelopt.recipe.presets`** module — no copy lives in the example
scripts anymore. The
module is a standalone import, so `import modelopt.recipe` stays cheap;
only an explicit
`import modelopt.recipe.presets` triggers the eager preset load.

A small alias table preserves previously-supported short CLI names
(`int8_sq`,
`nvfp4_awq`, `fp8_pb_wo`, …, plus the Megatron-Bridge `fp8_blockwise`)
as deprecation
shims. It is documented as not-for-extension — new formats land as
preset YAMLs, and
longer term, configurations should be authored as full recipes
(`--recipe`). The alias
logic is fail-fast: an alias pointing at a missing preset raises
`ValueError` at import.

Also adds `presets/kv/fp8_cast.yaml` and `presets/kv/nvfp4_cast.yaml`,
composed from the
existing `kv_fp8_cast` / `kv_nvfp4_cast` unit fragments. This promotes
`fp8_cast` /
`nvfp4_cast` to first-class KV presets and lets us delete the runtime
`_set_kv_cache_constant_amax` helper and all its call sites —
`use_constant_amax` is now
authoritative in the YAML. The KV-calibration-skip decision is derived
from the config
(`_kv_cfg_uses_constant_amax`), not from hardcoded format names.

**⚠️ CLI surface expansion (owner sign-off requested).** Because the
directory listing
is now the CLI vocabulary, each script accepts **every** preset under
`presets/{model,kv}/`, not just its previously curated subset. For
`hf_ptq.py` this is
the same surface the prior table covered; for `multinode_ptq.py` and the
Megatron-Bridge
script it is broader (e.g. KV `fp8_affine` / `fp8_cast` / `nvfp4_cast` /
`nvfp4_rotate`
are now selectable). This is intended ("the directory is the policy"),
but please confirm
those two scripts are meant to expose all presets — if a given path has
not validated a
format, it should be gated explicitly.

### Usage

```bash
# Old short names still work via the alias shim
python examples/llm_ptq/hf_ptq.py --pyt_ckpt_path <model> --qformat int8_sq --kv_cache_qformat fp8_cast --export_path out/

# Canonical preset basenames work directly
python examples/llm_ptq/hf_ptq.py --pyt_ckpt_path <model> --qformat int8_smoothquant --kv_cache_qformat fp8_cast --export_path out/

# A newly-added preset YAML is valid on the CLI of all three scripts with no code change
python examples/llm_ptq/hf_ptq.py --pyt_ckpt_path <model> --qformat nvfp4_awq_full --export_path out/
```

### Testing

- New: `tests/unit/recipe/test_presets.py` smoke tests for
`modelopt.recipe.presets` —
every discovered model/KV preset loads into a `quant_cfg` dict, the
directory listing is
fully covered, deprecation aliases resolve to their canonical preset,
the KV `none`
sentinel does not collide with a preset, and a stale alias raises. These
guard the eager
import-time load (one bad preset would otherwise break `import
modelopt.recipe.presets`
  and every PTQ example).
- Previously verified locally (uv `.venv` py3.13 + `dev-py310-modelopt`
conda):
all previously-supported `--qformat` / `--kv_cache_qformat` names
resolve to dicts
bit-equal to the corresponding `mtq.*_CFG` constants; `fp8_cast` /
`nvfp4_cast` carry
`use_constant_amax: true` while non-cast variants do not; argparse
accepts
`--kv_cache_qformat none` plus all variants; unknown qformats raise at
lookup / argparse.
- All pre-commit hooks pass (ruff, mypy, bandit, license, rst, yaml).
- Pre-merge manual checks recommended by review (run in an env with the
deps installed):
`python examples/llm_ptq/hf_ptq.py --help`, `… multinode_ptq.py --help`,
  `… megatron_bridge/quantize.py --help`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — all previously-valid CLI
values continue to work via the alias table; output configs are
bit-equivalent to the prior hardcoded path.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
deps.
- Did you write any new necessary tests?: ✅ —
`tests/examples/llm_ptq/test_example_utils.py` preset-discovery smoke
tests.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ☐

### Additional Information

Out of scope / follow-up: the `_AUTO_QUANTIZE_QFORMATS` table and
`_canonical_qformat`
helper in `hf_ptq.py` are intentionally left hardcoded — auto_quantize
is being
refactored/reimplemented and they are expected to be removed soon.

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-06-05 12:44:16 -07:00
realAsma 433b549cd8 [2/N] Simplify KDTrainer and enhance ModelOptHFTrainer (#1191)
## Summary

This PR simplifies the HuggingFace knowledge distillation trainer and
enhances the base `ModelOptHFTrainer` with Liger fused loss,
per-parameter learning rates, and training utilities.

### Model-agnostic Liger kernel fused loss

Adds custom Liger kernel integration in `ModelOptHFTrainer` that extends
HuggingFace's built-in support in three ways:

1. **Model-agnostic**: Works with any causal LM that has an `lm_head`,
unlike HF's Liger which only supports [a fixed set of model
architectures](https://github.com/linkedin/Liger-Kernel/blob/main/src/liger_kernel/transformers/monkey_patch.py).
2. **DeepSpeed ZeRO-3 support**: HF's Liger integration only works with
FSDP. ModelOpt adds distributed param gathering for DeepSpeed ZeRO-3 and
DDP as well.
3. **KD loss support**: `KDTrainer` extends fused loss to knowledge
distillation via `LigerFusedLinearJSD` for fused lm_head +
Jensen-Shannon divergence.

#### Liger kernel memory sweep (Qwen3-1.7B, 2×H100 FSDP2, NVFP4+FP8_KV)

Max per-GPU batch size before OOM at each sequence length:

**QAT (no teacher)**

| Seq Length | 512 | 1024 | 2048 | 4096 | 8192 | 16384 |
|------------|-----|------|------|------|------|-------|
| **Liger**  | 16  | 16   | 16   | 16   | 8    | 4     |
| **No Liger** | 16 | 16  | 8    | 4    | 2    | OOM   |

**QAD (with teacher)**

| Seq Length | 512 | 1024 | 2048 | 4096 | 8192 | 16384 |
|------------|-----|------|------|------|------|-------|
| **Liger**  | 16  | 16   | 8    | 4    | 2    | 1     |
| **No Liger** | 8 | 4   | 2    | 1    | OOM  | OOM   |

Liger fused loss enables **2-4× larger batch sizes** at long context
lengths by avoiding the materialization of the full logit tensor.

### ModelOptHFTrainer enhancements

- `ModelOptTrainerArguments` with `--trainable_params`,
`--frozen_params`, `--lr_config`, `--save_dtype`, and `--manual_gc`
flags
- Per-parameter learning rate support via YAML config (`lr_config`)
- `_prepare_model` and `_update_config_json_dtype` promoted to base
class

### KDTrainer simplification + fix

Removes `mtd.convert()` and the `DistillationModel` in-place class-swap
for the HF path. The teacher model now lives directly on the trainer and
is forwarded explicitly inside `compute_kd_loss_func`. This eliminates:

- `mtd.convert()` in-place class swap and DynamicModule wrapping
- Forward hooks for capturing intermediate outputs
- `hide_teacher_model` / `hide_loss_modules` context managers for
checkpointing
- Deferred initialization branching (FSDP2 vs DDP/DeepSpeed)
- `save_model` and `QADTrainer._quantize_model` overrides

**Bug fix**: The previous `DistillationModel`/`mtd.convert()` approach
did not support CPU RAM-efficient loading for QAD. The teacher model had
to be fully loaded on GPU before wrapping, which doubled peak memory
during initialization. The new approach loads the teacher lazily on the
trainer, enabling standard HF device-map and low-cpu-mem-usage loading.

Only logit-level distillation is supported for the HF path. The core
`DistillationModel`/`mtd.convert()` API remains for Megatron and
advanced intermediate-layer distillation use cases.

## Test plan

- [x] `pytest tests/unit/torch/distill/` (29 passed)
- [x] `pytest tests/unit/torch/opt/plugins/test_hf_patching.py` (2
passed)
- [x] `pytest tests/unit/torch/opt/plugins/test_lr_config.py`
- [x] Pre-commit hooks pass
- [ ] GPU example tests: `pytest tests/examples/llm_qat/` (QAT, QAD,
LoRA QAT, QLoRA)
- [ ] GPU distill example: `pytest tests/examples/llm_distill/`

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added Liger fused loss support in `ModelOptHFTrainer` for distributed
causal language models with JSD distillation loss support.
* Introduced `ModelOptTrainerArguments` with new training CLI flags:
per-parameter learning rates via YAML, parameter freezing, and manual
garbage collection.
* Simplified knowledge distillation trainer with logit-level
distillation support.

* **Documentation**
* Updated example configurations and documentation with new training
options and defaults.
  * Added learning rate configuration example guide.

* **Tests**
* Added test coverage for distillation training and per-parameter
optimizer configuration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-06-05 15:47:43 +00:00
Sepehr SameniandKeval Morabia 115cae2584 Puzzletron bypass distillation part 2 core (#1469)
## Summary

  This is PR 2 of 3 in the Puzzletron bypass/local-distillation stack.

This PR adds the bypass distillation core engine. It builds on PR 1’s
shared infrastructure, but it does not yet wire bypass into
  the full Puzzletron pipeline.

  Stack:
  1. `ssameni/puzzletron-bypass-1-prereqs`: shared prerequisites
  2. **This PR**: bypass distillation core
3. `ssameni/puzzletron-bypass-3-integration`: Puzzletron integration,
configs, docs, GPU coverage

  ## What Changed

  - Added `modelopt.torch.puzzletron.bypass_distillation`.
- Added bypass run identity/fingerprinting, experiment naming, state
manifests, and completion tracking.
- Added stitched teacher/student model construction for local blockwise
distillation.
  - Added bypass training loop with:
    - pipeline-parallel teacher activation stitching
    - per-block student losses
    - gradient accumulation
    - checkpoint/resume support
    - validation hooks
    - best/latest checkpoint realization
- Added bypass checkpoint helpers for saving optimizer/scaler state and
HF-format model checkpoints.
- Added scalable distributed checkpoint saving that gathers tensors per
safetensors file instead of materializing all shards on
  rank 0.
  - Restricted v1 `keys_to_learn` to subblock-level targets only:
    - `entire_block`
    - `subblock_attention`
    - `subblock_ffn`
    - `subblock_mamba`
    - lists of those keys
- Added pipeline ownership helper for deriving owned blocks and
neighboring PP ranks.
- Added robust `trust_remote_code` / `auto_map` handling for checkpoint
config saves.

  ## Why

Bypass distillation is the local-distillation stage used to train
pruned/reconfigured blocks before they are added to the
  replacement library.

Keeping the core engine separate from Puzzletron pipeline wiring makes
this PR reviewable on its own: reviewers can focus on
distributed training, checkpoint/resume semantics, and Sewing Kit
stitching behavior without also reviewing configs/docs/pipeline
  integration.

  ## Tests

  Added focused unit coverage for:

  - bypass checkpoint utilities
  - bypass run identity/fingerprinting and completion state
  - `keys_to_learn` subblock selection
  - LR scheduler behavior
  - launch/sweep dispatch behavior
  - HF checkpoint utility behavior
  - stitched model factory buffer ownership

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Full bypass distillation workflow: blockwise stitched student/teacher
distillation, deterministic experiment IDs/fingerprints, resume-capable
checkpointing with atomic symlink updates, optional asynchronous
checkpoint saves, and improved support for trust-remote-code model
loading.
  * Public bypass-distillation entrypoint exposed for easier launch.

* **Tests**
* Extensive unit tests covering bypass utilities, checkpoint behavior,
LR scheduler, stitched factory, and training orchestration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Sepehr Sameni <ssameni@nvidia.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-06-05 07:55:32 +00:00
Ajinkya RasaneandClaude Opus 4.8 01dec93580 [6058870] Fix ONNX AutoCast keep_io_types for empty I/O tensors (#1626)
### What does this PR do?

Type of change: Bug fix

ONNX AutoCast (used by `modelopt.onnx.quantization` with
`--high_precision_dtype fp16/bf16`) failed its internal sanity check
with `Unexpected type in I/O tensor <name>, keep_io_types=True, original
type: 1, converted type: 10` when a network input or output is an
**empty tensor** (a tensor with a dimension of size 0).

Root cause: `PrecisionConverter._add_cast` treats empty tensors
specially — instead of inserting a `Cast` node, it "fake-casts" them by
retyping their `value_info` in place (there is no data to convert).
However, `setup_mappings` stores the same `graph.input` / `graph.output`
`ValueInfoProto` objects in `value_info_map`, so retyping an empty
tensor that is a network input/output silently changed the model's
declared I/O type to the low-precision type — violating the
`keep_io_types=True` contract and tripping the sanity check.

Fix (`modelopt/onnx/autocast/precisionconverter.py`): skip the
metadata-only fake-cast for protected network I/O tensors (when
`keep_io_types=True` and the target type differs from the original I/O
type) and fall through to insert a real `Cast` node. This preserves the
declared I/O type while still bridging precision into the low-precision
consumer/producer. Intermediate empty tensors are unchanged and still
fake-cast.

### Usage

```bash
# Previously failed with "Sanity check failed: Unexpected type in I/O tensor ...";
# now completes and produces a valid mixed-precision model.
python -m modelopt.onnx.quantization \
    --quantize_mode int8 \
    --high_precision_dtype fp16 \
    --onnx_path model_with_empty_io_tensor.onnx \
    --output_path model_int8_fp16.onnx
```

### Testing

- Added `test_empty_tensor_network_input_keep_io_types` (fp16 + bf16,
both type-inference paths) in
`tests/unit/onnx/autocast/test_precisionconverter.py`, asserting network
I/O types stay FP32 and a real `Cast` bridges the empty input. Verified
it fails before the fix and passes after.
- Full autocast unit suite passes: `CUDA_VISIBLE_DEVICES="" pytest
tests/unit/onnx/autocast/` → 202 passed.
- Reproduced the original failure on the reported model and confirmed
the command now succeeds end-to-end; the resulting ONNX model passes
`onnx.checker`, runs in ONNX Runtime with correct FP32 I/O, and builds a
strongly-typed TensorRT engine.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Fixed an ONNX AutoCast issue where empty tensors with
keep_io_types=True could silently change model I/O types; now a proper
Cast is inserted for protected inputs/outputs to preserve declared I/O
types.

* **Tests**
* Added a regression test to ensure empty network-input tensors are
handled correctly in mixed-precision ONNX conversion when
keep_io_types=True.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 21:20:04 +00:00
yeyu-nvidiaandClaude Opus 4.6 a7b0a92047 EAGLE3 new model support: pipeline configs, triage docs, and Ministral-3 fixes (#1417)
## Summary

EAGLE3 automation triage work (OKR-30): testing the 4-step EAGLE3
offline pipeline against 12 new model architectures, documenting failure
modes, and fixing issues found.

### Code fixes (modelopt)

| File | Change |
|------|--------|
| `modelopt/torch/speculative/utils.py` | Extend VLM detection in
`load_vlm_or_llm` to check `text_config`/`llm_config` attrs (catches
`mistral3` models) |
| `modelopt/torch/speculative/plugins/modeling_fakebase.py` | Add
`consolidated.safetensors` fallback for checkpoints with incomplete HF
shards |
| `modelopt/torch/export/plugins/hf_spec_configs.py` | Set
`use_cache=True` in EAGLE export templates (fixes strict
`huggingface_hub` validation) |

### Pipeline infrastructure

- `examples/speculative_decoding/pipeline/eagle3/` — pipeline scripts
and configs:
- `offline_training.sh` — training + export with runtime patches for
older container modelopt
- `dump_offline_data_vllm.sh` — vLLM-based hidden state extraction (with
speculators compat patches)
- `dump_offline_data.sh`, `dump_offline_data_hf.sh` — alternative dump
paths
  - 18 quick-fail-check YAMLs for 12 models
  - 4 standalone task1 YAMLs

### Documentation

- `eagle3_triage_chart.md` — model test matrix, triage decision tree,
per-model results, failure catalog
- `eagle3_new_model_triage_guide.md` — step-by-step guide for triaging
new models

### Model test results (as of 2026-05-27)

| Model | task_0 | task_1 | task_2 | task_3 | Blocker |
|-------|--------|--------|--------|--------|---------|
| Qwen3-8B | - | - | - | - | Reference (existing) |
| Kimi-K2.5 | - | - | - | - | Existing (GB200) |
| **Ministral-3-8B** | SKIP | PASS | PASS | FAIL | `use_cache=null` in
export (fixed) |
| Ministral-3-14B | FAIL | - | - | FAIL | vLLM engine init fails |
| Qwen3.5-35B-A3B | TIMEOUT | - | - | - | Data synth too slow |
| gpt-oss-20b | FAIL | - | - | - | Tokenizer `HarmonyError` |
| Step-3.5-Flash | TIMEOUT | - | - | - | Data synth time limit |
| MiniMax-M2.5 | TIMEOUT | - | - | - | `trust_remote_code` needed |
| DeepSeek-V3.2 | no log | - | - | - | May not be mirrored |
| Qwen3.5-9B | - | - | - | - | Not yet run |
| Qwen3.5-27B | - | - | - | - | Not yet run |
| GLM-5 | - | - | - | - | Not yet run |

### Issues found and fixed

| # | Issue | Fix |
|---|-------|-----|
| 1 | `mistral3` model type not detected as VLM | Check
`text_config`/`llm_config` attrs in `load_vlm_or_llm` |
| 2 | Missing HF shard file (Ministral-3-8B) | Fallback to
`consolidated.safetensors` with Mistral native key aliases |
| 3 | `use_cache=null` in exported EAGLE config | Set `use_cache=True`
in export template configs |
| 4 | speculators incompatible with vLLM container | Runtime patches in
`dump_offline_data_vllm.sh` |
| 5 | `offline_training.sh` infra issues | Rewritten with runtime
patches for container modelopt |

## Test plan

- [x] Ministral-3-8B training passes (`cicd_1779829129`)
- [x] Ministral-3-8B export succeeds
- [ ] Ministral-3-8B benchmark passes (`cicd_1779901409` — pending with
all fixes)
- [ ] Dry-run remaining model configs

## Note

GitHub secret scanning alert #6 is a **false positive** —
`Mistral3ForConditionalGeneration` (a HuggingFace model class name in a
YAML comment) was flagged as a "Mistral AI API Key".

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-06-03 12:24:44 -07:00
yeyu-nvidia 651fd223e6 feat: EAGLE3 LoRA co-training improvements (#1607)
Layer-selective LoRA injection for EAGLE3 co-training, optimizer-stable
warmup (LoRA always in the optimizer; warmup gated by a flag), and an
export+merge+lm_eval evaluation script.

Review feedback addressed:
- trust_remote_code is caller-controlled (TRUST_REMOTE_CODE / --trust_remote_code), default False
- eagle_base_lora_start_layer raises ValueError instead of silently injecting zero adapters
- eval_lora.sh validates HF_MODEL_CKPT / EAGLE_CKPT up front
- added unit tests for start-layer injection and validation (8/8 passing, incl. on-cluster GPU run)

Signed-off-by: Ye Yu <yeyu@nvidia.com>
2026-06-03 17:33:12 +00:00
Keval MorabiaandClaude Opus 4.8 e40b4d69d6 Add Minitron pruning support for Gemma3 via Megatron-Bridge (#1604)
### What does this PR do?

Type of change: New feature (+ new tests)

Adds Minitron pruning support for **Gemma3** models loaded through
Megatron-Bridge (`Gemma3ForCausalLM` → `GPTModel`).

Gemma3's bridge implementation subclasses several megatron-core modules
and overrides `forward`/`__init__`, so `DMRegistry` (which matches by
`nn_cls.forward is parent.forward`) does not recognize them and
pruning's `convert_to_dynamic` raised:

```
KeyError: "<class '...Gemma3LanguageModelEmbedding'> is not registered for a dynamic module!"
```

There were actually **two** blockers, both fixed here:

1. **Layer spec was discarded.** `load_mbridge_model_from_hf`
unconditionally replaced `provider.transformer_layer_spec` with the
generic TE GPT spec (to disable grouped-GEMM), throwing away
`gemma3_layer_spec`. With the generic spec, plain
`TEDotProductAttention` received Gemma3's int `window_size` and crashed
at construction (`TypeError: 'int' object is not subscriptable`) before
pruning even started. Now the override is applied **only to MoE
models**; dense models keep the bridge's native spec.
2. **Custom layers weren't registered as dynamic modules.** Added
registrations for Gemma3's custom layers.

Changes:
- **`modelopt/torch/nas/plugins/mbridge.py`** (new): register dynamic
modules for `Gemma3LanguageModelEmbedding` (reuses the existing
embedding dynamic class), `Gemma3SelfAttention`, and the fused post-LN
`TERowParallelLinearLayerNorm` used by both `self_attention.linear_proj`
and `mlp.linear_fc2`. The fused `post_layernorm` is converted to a
dynamic module and sliced by the linear's `output_size` (==
`hidden_size`).
- **`modelopt/torch/nas/plugins/megatron.py`**: two small reuse seams in
`_DynamicSelfAttention` — build the core-attention dynamic class from
`type(self.core_attention)` (preserves `Gemma3TEDotProductAttention`)
and extract an overridable `_convert_linear_proj` hook. No behavior
change for megatron-core models.
- **`modelopt/torch/utils/plugins/mbridge.py`**: only override the layer
spec for MoE models; dense models keep the bridge's native spec.
- **`modelopt/torch/nas/plugins/__init__.py`**: import `mbridge` under
`import_plugin("megatron.bridge")`.
- **Tests**: new `get_tiny_gemma3` / `create_tiny_gemma3_dir` helpers;
parametrize `test_prune_minitron` over qwen3 and gemma3.

### Usage

```bash
# Prune a Gemma3 checkpoint with Minitron (same flow as other models)
torchrun --nproc_per_node 2 examples/megatron_bridge/prune_minitron.py \
    --hf_model_name_or_path google/gemma-3-1b-it \
    --prune_target_params 0.6e9 \
    --output_hf_path /tmp/gemma-3-pruned
```

### Testing

`tests/examples/megatron_bridge/test_prune_minitron.py` now runs for
both qwen3 and gemma3. Verified in the megatron-bridge container:

Also verified the gemma3 case standalone on 1 GPU (PP=1) and 2 GPUs
(PP=2).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (the spec-override change only
affects how dense Megatron-Bridge models are loaded — they now keep
their native, correct spec; MoE behavior is unchanged)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ (parametrized gemma3
coverage + tiny-model helper)
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ✅

### Additional Information

Other Megatron-Bridge models with custom layers were surveyed; only the
Gemma family needs this treatment. **Gemma1** (embedding scaling via
`EmbeddingScalingMixin`) and **Gemma2** (post-LN linear + non-TE
`Gemma2DotProductAttention` + custom output layer + mixin embedding) are
intentionally **left as outdated**. All other dense/MoE LLMs (llama,
qwen2/3, qwen3-moe, mistral, nemotron, deepseek, gpt-oss, etc.) use
standard megatron-core specs already covered by `megatron.py`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added Minitron pruning support for Megatron-Bridge Gemma3 models.
* Improved Megatron-Bridge layer configuration handling for
Mixture-of-Experts scenarios.

* **Tests**
  * Added Gemma3 model creation utilities for testing.
* Expanded pruning validation tests to cover both Qwen3 and Gemma3
models.

* **Examples**
* Updated pruning example to adjust tokenizer "use fast" detection based
on model architecture.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 12:00:41 +05:30
realAsmaandClaude Opus 4.8 196c091027 [1/N] Refactor llm_qat example: YAML configs + ModelOptArgParser (#1172)
### What does this PR do?

Type of change: new example

Refactors `examples/llm_qat` from a monolithic launch/script flow into a
modular, config-driven Hugging Face QAT/QAD workflow. The high-level
user flow is now:

1. Quantize a base model with a ModelOpt PTQ recipe.
2. Train or evaluate the quantized checkpoint with QAT, QAD, LoRA QAT,
QLoRA, or fine-tuning configs.
3. Export the trained checkpoint for deployment.

Highlights:

- Replaces the legacy `examples/llm_qat/launch.sh` +
`examples/llm_qat/main.py` path with separate `quantize.py` and
`train.py` entrypoints.
- Adds YAML-driven argument parsing through `ModelOptArgParser`,
including `--config <yaml>` defaults, CLI overrides, and generated
`examples/llm_qat/ARGUMENTS.md`.
- Adds declarative configs under `examples/llm_qat/configs/` for
training modes, dataset blends, and Accelerate backends.
- Adds `dataset_utils.py` for weighted multi-source dataset blending,
Hugging Face streaming, local dataset loading, distributed rank-aware
loading, tokenization caching, pre-tokenization, chat templating, and
assistant-token label masking.
- Moves the Hugging Face QAD flow into the shared trainer path through
`DistillArguments`, `QADTrainer`, and teacher-model distillation kwargs.
- Updates `QuantizationArguments` to prefer recipe paths via `--recipe`,
while keeping legacy `--quant_cfg` available with deprecation warnings
for in-trainer quantization.
- Updates `QATTrainer` handling for pre-quantized checkpoints, FSDP2
TensorQuantizer buffers, and recipe-resolved PTQ configs.
- Adds the `general/ptq/int4_blockwise_weight_only` recipe and refreshes
docs for NVFP4, FP8, INT4, FSDP2, DDP, DeepSpeed, QLoRA, and
LLaMA-Factory integration.
- Adds focused parser, dataset tokenization, assistant-mask, and example
workflow coverage.

### Usage

From the repo root:

```sh
cd examples/llm_qat

# 1. Quantize
python quantize.py \
  --model_name_or_path Qwen/Qwen3-8B \
  --dataset_config configs/dataset/blend.yaml \
  --recipe general/ptq/nvfp4_default-kv_fp8 \
  --output_dir qwen3-8b-quantized

# 2. QAT train
accelerate launch --config-file configs/accelerate/fsdp2.yaml train.py \
  --config configs/train/qat_nvfp4.yaml \
  --model_name_or_path qwen3-8b-quantized \
  --output_dir qwen3-8b-qat-nvfp4

# 3. QAD train
accelerate launch --config-file configs/accelerate/fsdp2.yaml train.py \
  --config configs/train/qad_nvfp4.yaml \
  --model_name_or_path qwen3-8b-quantized \
  --teacher_model Qwen/Qwen3-8B \
  --output_dir qwen3-8b-qad-nvfp4
```

Dataset blends can be pre-tokenized and cached before training:

```sh
cd examples/llm_qat
python dataset_utils.py \
  --dataset_config configs/dataset/blend.yaml \
  --model_name_or_path Qwen/Qwen3-8B
```

### Testing

Focused coverage added or updated:

- `tests/unit/torch/opt/plugins/test_modelopt_arg_parser.py`
- `tests/examples/llm_qat/test_dataset_tokenization.py`
- `tests/examples/llm_qat/test_assistant_mask.py`
- `tests/examples/llm_qat/test_llm_qat.py`

Recorded validation:

- [x] `pytest tests/unit/torch/opt/plugins/test_modelopt_arg_parser.py`
- [x] `pytest
tests/examples/llm_qat/test_llm_qat.py::test_dataset_utils_pretokenize`
- [x] `pre-commit run --all-files`
- [ ] `pytest tests/examples/llm_qat/test_llm_qat.py` full GPU/backend
suite

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: No for the legacy
`examples/llm_qat` CLI/file layout (`main.py`, `launch.sh`, and FSDP1
config are removed); library `quant_cfg` usage remains available but is
deprecated for this workflow.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: Yes.
`train.py`/`simple_qat_train.py` retain the upstream Alpaca attribution
where applicable; `examples/llm_qat/requirements.txt` switches
`tensorboardX` to `tensorboard`.
- Did you write any new necessary tests?: Yes. Parser, dataset
tokenization, assistant masking, pre-tokenization, and QAT/QAD workflow
tests were added or updated.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
Yes.
- Did you get Claude approval on this PR?: N/A in this description
update.

### Additional Information

This PR intentionally changes the `llm_qat` example surface. Existing
users should move from `examples/llm_qat/main.py` and
`examples/llm_qat/launch.sh` to `examples/llm_qat/quantize.py`,
`examples/llm_qat/train.py`, and the YAML configs under
`examples/llm_qat/configs/`.

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 01:25:02 +00:00
h-guo18 902d36921a [Feat]: Streaming Hidden-states Dataset (#1509)
### What does this PR do?

Type of change: new feature

**Design doc:**
https://gist.github.com/h-guo18/241c94968b0591324c361d97cf995dd0
Jira ticket: https://jirasw.nvidia.com/browse/OMNIML-4341
Sandbox CI:
https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/327911555#L2228

Streaming hidden-states dataset: per-sample activations pulled from a
live `vllm serve` over HTTP, replacing on-disk activation dumps.

Two axes for future extensions:
- **Backend** (`_fetch`): vLLM now; TRT-LLM / SGLang next.
- **Algorithm** (`_format`): Eagle now; distillation / probing next.
- **Sandbox CI**:
https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/merge_requests/169

Shared plumbing — async producer, token-level truncation to
`training_seq_len`, loss-mask alignment, DDP via Accelerate's dispatcher
(rank 0 fetches, broadcasts), circuit breaker, resume — lives in
`StreamingDataset`. First instance: **`EagleVllmStreamingDataset`**.

API: `data.mode ∈ {online, offline, streaming}`; legacy configs
auto-promote.

### Usage

```yaml
data:
  mode: streaming
  data_path: input_conversations/train.jsonl
  streaming_server_url: http://localhost:8000
  streaming_model_name: meta-llama/Llama-3.1-8B-Instruct
training:
  training_seq_len: 4096   # also caps the prompt sent to vllm
```

Requires `vllm serve` with `ExampleHiddenStatesConnector` and
`dataloader_num_workers=0`.

End-to-end Slurm pipeline:
`tools/launcher/examples/Qwen/Qwen3-8B/hf_streaming_eagle3.yaml`.

### Testing

- **Unit**: full-corpus invariant, rank-0-only iter, resume, circuit
breaker, mocked-httpx integration.
- **E2E**: `launch_train.sh` against a stdlib `HTTPServer` mimicking the
connector.
- **Smoke** (Qwen3-8B / 8×H100 / 4096 ultrachat samples, single epoch):
train_loss 32 → 18, MT-Bench AR 1.003 → 1.20.

### TODO before un-drafting

- [ ] Observability counters (filtered, fetch failures, queue depth,
latency).
- [ ] Changelog entry.
- [x] Add test in sandbox.

**Non-goals (v1):** multi-epoch streaming, cross-rank dynamic load
balancing.

### Before your PR is "*Ready for review*"

- Backward compatible: ✅ (legacy configs auto-promote)
- New PIP dep: N/A
- New tests: ✅
- Changelog: ❌ (TODO above)
- Claude approval: ❌


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Streaming training mode with server-backed hidden-state fetching,
deterministic seed control, resume support, and streaming-specific
dataset options (server, model, prefetch, shared storage).

* **Behavior / Bug Fixes**
* Stronger mode validation; offline behavior derived from data mode;
resume handling adjusted to avoid double-skip during streaming runs.

* **Tests**
* End-to-end CI streaming test and expanded unit tests covering
streaming, resume, DDP, determinism, and failure cases.

* **Infrastructure**
* Launcher script and pipeline config for end-to-end streaming training.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1509?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-06-02 14:56:23 -07:00
Keval MorabiaandClaude Opus 4.8 f21977a5fc Add Megatron-Bridge PTQ quantize + export example scripts (#1589)
### What does this PR do?

Type of change: new example

Adds a two-step post-training quantization (PTQ) flow for
**Megatron-Bridge** models under `examples/megatron_bridge/`, mirroring
the Megatron-LM `quantize.sh` / `export.sh` split:

- **`quantize.py`** — loads an HF model via Megatron-Bridge, applies
ModelOpt PTQ (via a `--quant_cfg` alias / full config name, or a
`--recipe` YAML), with optional KV-cache quant, weight-only,
compression, and MoE expert-ratio calibration, then saves a **Megatron
checkpoint** (with ModelOpt state). Tensor / pipeline / expert
parallelism are all supported, and the checkpoint can later be reloaded
for further training (QAT / distillation).
- **`export.py`** — loads the quantized Megatron checkpoint, **re-shards
to TP=1**, and exports a **HuggingFace (unified)** checkpoint deployable
with TensorRT-LLM / vLLM / SGLang.

**Why the split?** The unified HF exporter (`export_mcore_gpt_to_hf`)
does not gather tensor-parallel-sharded weights — Megatron-LM likewise
forces `TP=1` during its export step. Saving a TP-sharded Megatron
checkpoint first lets us calibrate at TP>1 (to fit large models) and
then reload re-sharded to TP=1 for the HF export. A combined
single-script flow silently produced corrupt HF checkpoints under TP>1
(collided per-rank shards), which this split avoids.

> **Note:** This is **part 1 of 4**:
> - **Part 1 (this PR):** Megatron-Bridge `quantize.py` + `export.py`
support and tests.
> - **Part 2:** extend `distill.py` for quantization-aware distillation
(QAD) — load a quantized Megatron checkpoint as the student.
> - **Part 3:** add NVFP4 + QAD-on-pruned-checkpoint experiments to the
Nemotron-3-Nano-30B-A3B tutorial.
> - **Part 4:** repeat the NVFP4 + QAD experiments on a non-Nemotron
model.

### Usage

```bash
# Step 1: quantize (TP/PP/EP supported) -> Megatron checkpoint
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-8B \
    --quant_cfg fp8 \
    --tp_size 2 \
    --export_megatron_path /tmp/Qwen3-8B-FP8-megatron

# Step 2: export -> deployable HuggingFace (unified) checkpoint (re-shards to TP=1)
torchrun --nproc_per_node 1 export.py \
    --hf_model_name_or_path Qwen/Qwen3-8B \
    --megatron_path /tmp/Qwen3-8B-FP8-megatron \
    --export_unified_hf_path /tmp/Qwen3-8B-FP8-hf
```

### Testing

`tests/examples/megatron_bridge/test_quantize.py` (validated on a 2-GPU
NeMo `26.04` container):

- `test_quantize_export_and_vllm_deployment` — quantize a tiny Qwen3 via
a recipe at TP=2 → `export.py` re-shards to TP=1 → load + generate with
**vLLM** (skipped if vLLM absent).
- `test_quantize_megatron_checkpoint_reload` — quantize at TP=2 → reload
the Megatron checkpoint via the bridge and assert ModelOpt quantizers
were restored.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A (new example)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new
dependencies)
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅

### Additional Information

The Nemotron-3 tutorial update to use these scripts is intentionally
**not** included here — it ships with the part 3 PR alongside the NVFP4
+ QAD experiments.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Expanded post-training quantization (PTQ) workflow documentation with
detailed step-by-step examples and configuration guidance for the
Megatron-Bridge framework.

* **New Features**
* Added quantization tool for applying PTQ to Megatron models with
calibration support.
* Added export tool for converting quantized models to a deployable
format.

* **Tests**
* Added integration tests validating the complete
quantization-export-deployment workflow, including inference validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 19:36:15 +00:00
sychen52 8f968322f6 [OMNIML-3994] Make sure all weight quantizers have _amax (#1560)
### What does this PR do?
- max_calibrate now always runs weight_only_quantize before the optional
forward_loop, so every weight quantizer has _amax populated regardless
of MoE routing.
- awq_lite search loop explicitly deletes _amax around the alpha sweep
to keep using the dynamic-amax path; no restore needed (postprocess
overwrites for enabled modules, disabled modules early-return).
  - Removed three band-aids:
    - _bootstrap_uncalibrated_weight_quantizers (mse_calibrate)
    - per-module max_calibrate fallback in awq_lite postprocess
    - _ensure_weight_quantizer_calibrated + helpers (export/quant_utils)
- Renamed test_bootstrap_populates_dead_expert_quantizers ->
test_max_calibrate_populates_dead_expert_quantizers; deleted the GPU
lazy-calibration test that no longer has a behavior to exercise.
Type of change:  refactor

<!-- Details about the change. -->

### Usage
same as before.

### Testing
run unittests locally

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Refactor**
* Simplified the weight quantization calibration process by
consolidating logic and removing redundant internal functions.

* **Tests**
  * Removed a redundant quantization format-specific export test.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1560?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
2026-06-02 15:47:15 +00:00
Wei-Ming Chen 72df833e0b Add active-MoE AutoQuant cost accounting (#1497)
### What does this PR do?

• Type of change: new feature

Adds an active_moe cost model for auto_quantize effective-bits search.
This lets AutoQuant account for routed MoE expert weights by active
decode weight traffic instead of total checkpoint weight
size, using active_moe_expert_ratio = num_experts_per_tok / num_experts.

The default behavior is unchanged: cost_model="weight" still counts all
quantizable weights equally.

  ### Usage

  import modelopt.torch.quantization as mtq

  model, search_state = mtq.auto_quantize(
      model,
      constraints={"effective_bits": 5.0},
      quantization_formats=[
          mtq.NVFP4_DEFAULT_CFG,
          mtq.FP8_DEFAULT_CFG,
      ],
      data_loader=calib_dataloader,
      forward_step=forward_step,
      loss_func=loss_func,
      cost_model="active_moe",
# Optional. If omitted, ModelOpt tries to infer this from model.config.
      active_moe_expert_ratio=2 / 64,
  )

  The HF PTQ example also exposes:

  --auto_quantize_cost_model active_moe \
  --auto_quantize_active_moe_expert_ratio 0.03125

  ### Testing

python -m pytest tests/unit/torch/quantization/test_autoquant.py -q -k
'active_moe or quant_recipe_hparam_cost_weight'
python -m pytest tests/unit/torch/quantization/test_autoquant.py -q -k
'not data_parallel_auto_quantize'

  Results:

  - 4 passed
  - 58 passed, 1 deselected

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added active-MoE cost model option for auto-quantization with
configurable expert ratio; API and CLI accept cost_model and
active_moe_expert_ratio
  * Unified auto-quantize supports new quant format w4a16_nvfp4

* **Bug Fixes**
* Ensure labels are moved to the logits device for base models without
an lm_head
* CLI enforces valid expert-ratio range and requires active-MoE mode
when a ratio is provided

* **Tests**
* Added unit tests for active-MoE behavior, cost-weighting, ratio
handling, and search budget selection

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1497?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-06-02 09:06:07 +00:00
hychiangandClaude Sonnet 4.6 f0d2237cbc Add Qwen3VL MCore Export support from PR 895 (#1482)
# [Megatron Export] Add Qwen3-VL mcore ↔ HF weight mapping

> This PR is duplicated from [PR
#895](https://github.com/NVIDIA/Model-Optimizer/pull/895).
> The original branch source is no longer available; this new branch
carries the same changes forward.

## What does this PR do?

**New feature:** Add Qwen3-VL (Vision-Language) model support to the
Megatron Core export/import
plugin, enabling HuggingFace-to-mcore weight conversion for PTQ/QAT/QAD
workflows.

### Overview

Qwen3-VL has a different weight structure from Qwen3 text-only models:

- Language model weights are under `model.language_model.` prefix (not
`model.`)
- Visual encoder weights are under `model.visual.` prefix
- `lm_head` is at root level, not nested under `language_model`

### What changed

| File | Change |
|---|---|
| `modelopt/torch/export/plugins/mcore_qwen3vl.py` | New plugin: derives
Qwen3-VL mcore↔HF mapping by rewriting `model.*` →
`model.language_model.*` on top of the existing Qwen3 dense rules;
`lm_head.` is intentionally left unchanged |
| `modelopt/torch/export/plugins/mcore_common.py` | Registers
`Qwen3VLForConditionalGeneration` in `all_mcore_hf_export_mapping` and
`all_mcore_hf_import_mapping` |
| `modelopt/torch/export/plugins/hf_checkpoint_utils.py` | Generalized
`load_multimodal_components` with a `prefixes` parameter; sharded
checkpoints now scan all shards (not just the first) |
| `modelopt/torch/export/unified_export_megatron.py` |
`save_pretrained`: added Qwen3-VL branch that copies `model.visual.*`
vision-encoder weights from the original HF checkpoint into the exported
directory, producing a complete, loadable checkpoint |
| `tests/_test_utils/torch/transformers_models.py` | Added
`get_tiny_qwen3vl` / `create_tiny_qwen3vl_dir` helpers; Qwen3VL classes
are lazy-imported inside the function to avoid collection failures on
older transformers builds |
| `tests/gpu_megatron/torch/export/test_unified_export_megatron.py` |
Integrated Qwen3-VL export/import tests into the existing
`test_unified_export_megatron` / `test_unified_import_megatron`
parametrized suites; removed standalone `test_mcore_qwen3vl.py` |
| `docs/source/deployment/3_unified_hf.rst` | Added Qwen3-VL (FP8 /
NVFP4) to the deployment support matrix for TensorRT-LLM |

### Workflow coverage

| Step | Status | Files |
|---|---|---|
| 1. Quantize Qwen3-VL with `hf_ptq` | ✅ existing | — |
| 2. Export quantized mcore → HF | ✅ this PR |
`plugins/mcore_qwen3vl.py` (weight name mapping),
`unified_export_megatron.py` (export path) |
| 3. Vision-encoder weights merged into export dir | ✅ this PR |
`plugins/hf_checkpoint_utils.py` (`load_multimodal_components` with
`prefixes`), `unified_export_megatron.py` (calls it when `arch ==
"Qwen3VLForConditionalGeneration"`) |
| 4. Import HF checkpoint back to mcore | ✅ this PR |
`plugins/mcore_qwen3vl.py` (same mapping, reverse direction),
`unified_export_megatron.py` (import path) |

### Design notes

- **MoE not supported**: `Qwen3VLMoeForConditionalGeneration` stores
expert weights as
3-D tensors (`mlp.experts.gate_up_proj`, `mlp.experts.down_proj`) that
require a
dedicated fused-expert mapping. A `NotImplementedError` comment in the
plugin
  documents this explicitly.
- **`copy.deepcopy` on `func_kwargs`**: each mapping entry gets its own
copy to
prevent shared-dict mutation when both Qwen3 and Qwen3-VL rules are
loaded.
- **`prefixes` parameter on `load_multimodal_components`**:
backward-compatible default
preserves existing LLaVA behaviour (`"multi_modal_projector"`,
`"vision_model"`);
  Qwen3-VL callers pass `("model.visual.",)`.
- **Sharded checkpoint scan**: the old code only looked in the first
shard. The
Qwen3-VL vision encoder can span multiple shards, so all shards are now
scanned.

## Usage

From the [Megatron-LM PR
comment](https://github.com/NVIDIA/Megatron-LM/pull/3444#issuecomment-3911271713):
> Qwen3VL is supported within
[Megatron-Bridge](https://github.com/NVIDIA-NeMo/Megatron-Bridge), and
pretraining and PEFT recipes for Qwen3VL are
[here](https://github.com/NVIDIA-NeMo/Megatron-Bridge/blob/main/src/megatron/bridge/recipes/qwen_vl/qwen3_vl.py)
and the core code logic
[here](https://github.com/NVIDIA-NeMo/Megatron-Bridge/tree/main/src/megatron/bridge/models/qwen_vl).

Create
`Megatron-LM/examples/post_training/modelopt/conf/Qwen/Qwen3-VL-8B-Instruct.sh`:

```bash
#!/bin/bash
# Qwen3-VL-8B-Instruct text-model config for Megatron-LM import/quantize.
#
# Text-model dimensions are identical to Qwen3-8B (4096 hidden, 36 layers,
# 32 heads, GQA=8).  Differences: rope_theta=5000000, checkpoint path uses
# model.language_model.* prefix (handled by mcore_qwen3vl plugin).

if [ -z ${HF_MODEL_CKPT} ]; then
    HF_MODEL_CKPT=Qwen/Qwen3-VL-8B-Instruct
    TOKENIZER_MODEL=Qwen/Qwen3-VL-8B-Instruct
else
    TOKENIZER_MODEL=${HF_MODEL_CKPT}
fi

MODEL_ARGS=" \
    --save-interval 100000 \
    --micro-batch-size 1 \
    --bf16 \
    --no-masked-softmax-fusion \
    --disable-bias-linear \
    --untie-embeddings-and-output-weights \
    --position-embedding-type rope \
    --no-rope-fusion \
    --normalization RMSNorm \
    --swiglu \
    --num-layers 36 \
    --hidden-size 4096 \
    --ffn-hidden-size 12288 \
    --num-attention-heads 32 \
    --group-query-attention \
    --num-query-groups 8 \
    --kv-channels 128 \
    --qk-layernorm \
    --seq-length 4096 \
    --max-position-embeddings 262144 \
    --tokenizer-type HuggingFaceTokenizer \
    --make-vocab-size-divisible-by 1187 \
    --use-mcore-models \
    --rotary-percent 1.0 \
    --rotary-base 5000000 \
    --no-bias-swiglu-fusion \
"
```

Import Qwen3-VL from HuggingFace to MCore (local, requires GPUs):

```bash
MLM_MODEL_CFG=Qwen/Qwen3-VL-8B-Instruct \
HF_MODEL_CKPT=Qwen/Qwen3-VL-8B-Instruct \
MLM_MODEL_SAVE=/tmp/qwen3vl_mcore \
TP=1 \
bash Megatron-LM/examples/post_training/modelopt/convert.sh Qwen/Qwen3-VL-8B-Instruct
```

Quantize (PTQ via Megatron-LM path):

```bash
MLM_MODEL_CFG=Qwen/Qwen3-VL-8B-Instruct \
HF_MODEL_CKPT=Qwen/Qwen3-VL-8B-Instruct \
QUANT_CFG=NVFP4_DEFAULT_CFG \
TP=4 \
bash Megatron-LM/examples/post_training/modelopt/quantize.sh Qwen/Qwen3-VL-8B-Instruct
```

## Testing

- Verified round-trip import/export with Qwen3-VL-8B-Instruct with the
example usage above
- Unit/GPU tests covering:
  - Registration in global export/import mappings
- Import mapping: dense keys, `model.language_model.` prefix, `lm_head.`
at root, `QKVMerging`, `GatedMLPMerging`, `REPLICATE` for layernorms, TP
sharding configs
- Export mapping: `QKVSlicing`, `GatedMLPSlicing`, no `parallel_config`
  - Import/export symmetry: same mcore keys, matching HF prefixes
- Qwen3-VL vs Qwen3 difference: same keys, VL adds `language_model.`
prefix, `lm_head` unchanged

## Before your PR is "Ready for review"

- Is this change backward compatible?: Yes, additive only
- Did you write any new necessary tests?: Yes,
`tests/gpu_megatron/torch/export/test_unified_export_megatron.py`
- Did you add or update any necessary documentation? Yes, see
`docs/source/deployment/3_unified_hf.rst`
- Did you update Changelog? Yes, see `CHANGELOG.rst`

## Additional Information

Companion Megatron-LM PR adds `Qwen3VLModel`, `Qwen3VLDataset`, and
`pretrain_qwenvl.py`.
See: https://github.com/NVIDIA/Megatron-LM/pull/3444

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: hychiang <hungyuehc@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-01 14:29:40 -07:00
h-guo18 40a4dd326d [Feat]: Eagle Dry Run Mode (#1566)
### What does this PR do?

Type of change: new feature

Adds `--dry_run` to `examples/speculative_decoding/main.py`: load →
`mtsp.convert` → save, then exit (no `trainer.train()`). With
`FakeBaseModel`, the convert→save→export chain runs in seconds and
produces an exportable EAGLE3 / Medusa / DFlash checkpoint with correct
structure but untrained draft-head weights — useful for end-to-end
plumbing smoke tests on downstream stacks (vLLM, TRT-LLM, SGLang)
without paying for a real training run.

Where it sits among existing EAGLE3 modes:

```
EAGLE3 modes
├── online      base model runs forward in-loop
├── offline     reads pre-dumped hidden states from disk
├── streaming   streams hidden states from a live server in-loop
└── dry-run ★  skip training entirely; convert + save + export   ← NEW (this PR)
```

Companion `FakeBaseModel` fixes so small base checkpoints work:
- Synthesize the weight_map from a single `model.safetensors` when no
sharded index is present (Llama-3.2-1B, Qwen3-0.6B, …).
- Honor `tie_word_embeddings`: reuse `embed_tokens` when `lm_head` is
absent from safetensors.

A new launcher YAML
(`tools/launcher/examples/Qwen/Qwen3-8B/hf_eagle3_dryrun.yaml`) wires
this together as a one-task pipeline.

### Usage

```bash
# Direct
python main.py --dry_run \
  --config modelopt_recipes/general/speculative_decoding/eagle3.yaml \
  model.model_name_or_path=meta-llama/Llama-3.1-8B-Instruct \
  model.use_fake_base_for_offline=true \
  data.offline_data_path=/tmp/dryrun-placeholder \
  training.output_dir=ckpts/dryrun
python scripts/export_hf_checkpoint.py --model_path ckpts/dryrun --export_path export/dryrun

# Launcher
uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_eagle3_dryrun.yaml --yes
```

### Testing

- `tests/unit/torch/speculative/plugins/test_fakebase.py`: 3 new cases —
single-file fallback, tied-embeddings fallback, and the negative case
(missing `lm_head` without tying).
-
`tests/examples/speculative_decoding/test_eagle.py::test_eagle3_dry_run`:
full `launch_train.sh --dry_run → export_hf_checkpoint.py` chain on
`tiny_llama`; asserts exported state_dict has all
`LLAMA_EAGLE_SINGLE_LAYER` required keys.
- Manually verified end-to-end on Llama-3.1-8B-Instruct (sharded),
Llama-3.2-1B-Instruct (single-file + tied), and Qwen3-0.6B
(single-file).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `--dry_run` is opt-in;
`FakeBaseModel` changes are additive fallbacks.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — examples-only addition.
- Did you get Claude approval on this PR?: ❌

### Additional Information

N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a `--dry_run` CLI flag to perform a fast early-exit execution
that saves model artifacts without training.
* Added support for single-file model checkpoint formats in the loader.

* **Bug Fixes**
  * Improved checkpoint-loading error messages and handling.
  * Enhanced tied-embeddings fallback when head weights are absent.

* **Tests**
* Added integration and unit tests covering dry-run behavior and
single-file/tied-embedding loading.

* **Documentation**
  * Added a launcher example for dry-run smoke testing.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1566?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-05-29 15:02:26 -07:00
realAsma d7e72f42ed Refine static NVFP4 MSE calibration (#1536)
### What does this PR do?

Type of change: Bug fix

Refines static NVFP4 MSE calibration and forces static NVFP4 amax state
to stay FP32 across calibration loading, quantizer promotion, dtype
casts, and restore paths.

Main changes:

- Tighten max/MSE calibration bootstrap and static NVFP4 quantizer
promotion.
- Keep static NVFP4 `_amax` and `_global_amax` in FP32.
- Update focused GPU/unit coverage for FP8 sweep calibration, promotion,
restore, and FP32 amax preservation.

### Usage

```yaml
algorithm:
  method: mse
  fp8_scale_sweep: true
```

### Testing

```bash
pre-commit run --files modelopt/torch/quantization/nn/modules/tensor_quantizer.py tests/gpu/torch/quantization/test_nvfp4_static_quantizer_cuda.py
pytest_pwd tests/gpu/torch/quantization/test_nvfp4_static_quantizer_cuda.py -q
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).

- Is this change backward compatible?: Yes
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: Yes
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: Yes

### Additional Information

N/A

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-05-29 21:04:04 +00:00
Gwena Cunhaandmodelopt-fix-agent-bot f99d83eeab [Auto-23/24][ONNX][Autocast] Clear stale Cast-output type metadata before ORT InferenceSession load (#1565)
### What does this PR do?

Type of change: Bug fix

**Error:**
```
onnxruntime.capi.onnxruntime_pybind11_state.Fail: [ONNXRuntimeError] : 1 : FAIL : Type Error: Type (tensor(float16)) of output arg (node_5bc985fa) of node (node_5bc985fa) does not match expected type (tensor(float)).
```

**Root cause:**
- Some ONNX exporters emit `graph.output` / `value_info` entries whose
dtype disagrees with the upstream `Cast` node's `to` attribute.
- ORT's type checker rejects such models on session load.

**Fix:**
- New helper `modelopt.onnx.utils.clear_stale_value_info()` reconciles
each `graph.output` elem_type to its producing Cast's `to`, then clears
`value_info` so ORT recomputes intermediate types.
- Called from `autocast/referencerunner.py` and
`quantization/quantize.py::_preprocess_onnx`.

### Usage

```python
# Internal fix; no new flag introduced. Generic CLI to exercise the affected Autocast path:
$ python -m modelopt.onnx.autocast --onnx=model.onnx
```

### Testing

```
pytest tests/unit/onnx/test_onnx_utils.py::test_clear_stale_value_info
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Automatic cleaning and reconciliation of stale ONNX type metadata
before runtime and quantization; reconciled model files are produced and
used when inconsistencies are found.
* **Tests**
* New unit tests covering metadata-cleaning behavior across cast/type
scenarios to ensure correctness and prevent regressions.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1565?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
Co-authored-by: modelopt-fix-agent-bot (Claude Opus 4.7) <noreply@anthropic.com>
2026-05-29 19:38:49 +00:00
Keval Morabia eb5ed2df68 [CI] Bump torch, transformers and dev containers to latest (#1554)
- Transformers upper bound bumped from `<5.8` to `<5.10`
- Enable torch 2.12 CICD testing
- Bump TRT-LLM container to `1.3.0rc16` (transformers 5.5)
- Use pytorch and tensorrt 26.04 containers in CICD

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Updated CI test container images and targeted Torch version across
workflows; adjusted release CI job to use the newer torch config.
* Broadened Transformers constraint in project metadata and test/dev
pins.
* Removed strict transformers pins from example requirements and lifted
a compression dependency cap.
* Raised the import-time Transformers version threshold for
compatibility warnings.

* **Tests**
* Refactored a GPU test to collect and report validation errors and
updated numeric expected baselines.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1554?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-29 10:37:00 -07:00
Jenny Chen 4b270f0ea6 Support Mixed precision & Static MSE in MCore; Nemotron Super v3 NVFP4 recipe (#1521)
### What does this PR do?

Type of change: New features + Bug fixes

Mixed Precision and MSE support in MCore PTQ
- support mixed precision export in MCore by detecting mixed precision
layers in HF Quant Config
- Restore static quantizer in MCore checkpoint restore as `NVFP4QTensor`
(not TensorQuantizer which can call max calibrate. we want to skip max
calibrate for static quantizer during restore) --> fixes bug during
MCore export for MSE
- Fix dynamic block quantizer detection when `block_sizes` is
dict-backed.
- Add a YAML quantization recipe that roughly mirrors Nemotron 3 Super
NVFP4 `hf_quant_config.json`
  Export bug fixes     
- copy .py files properly from original HF ckpt (for reasoning parser
etc)
## Super recipe

Mirrors the published nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
hf_quant_config.json:
- MoE routed experts (mixer.experts.<N>.{up,down}_proj): NVFP4 W4A4
weight MSE, group_size 16
- MoE shared experts (mixer.shared_experts.{up,down}_proj): FP8
per-tensor
- Mamba mixer linears (mixer.{in,out}_proj): FP8 per-tensor
- KV cache:                                                    FP8
rest: not quantized

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
Tested on Nemotron model
### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added NVFP4 (4-bit) quantization checkpoint restore and export support
for Megatron-Core models
  * Added tokenizer file export capability in model checkpoints
* Extended quantization support for expert-parallel distributed training
* Introduced new PTQ recipes for Nemotron-3-Super models with
mixed-precision quantization

* **Bug Fixes**
* Fixed FP8 and FP4 hardware compatibility detection on non-CUDA systems
* Improved offline Hugging Face Hub access handling with better error
messaging
  * Enhanced calibration validation for mixture-of-experts models
  * Fixed amplitude maximum validation for static block quantizers

* **Documentation**
  * Updated expert weight quantization configuration documentation

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1521?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Jenny Chen <jennifchen@nvidia.com>
2026-05-29 16:23:09 +00:00
Sepehr Sameni a9c156e24a Split bypass prerequisites (#1468)
## Summary

  This is PR 1 of 3 in the Puzzletron bypass/local-distillation stack.

This PR contains prerequisite infrastructure only. It does not wire
bypass distillation into the Puzzletron pipeline yet.

  Stack:
  1. **This PR**: shared prerequisites
  2. `ssameni/puzzletron-bypass-2-core`: bypass distillation core
3. `ssameni/puzzletron-bypass-3-integration`: Puzzletron integration,
configs, docs, GPU coverage

  ## What Changed

- Added `ModelDescriptor.pruning_mixins()` so model families can expose
pruning mixins needed by downstream bypass initialization.
- Added KV-head pruning mixin support for GPT-OSS, Nemotron-H,
Nemotron-H-v2, and Qwen3-VL descriptors.
- Improved pruning utilities for nested language-model configs and
missing attention bias config fields.
- Added `create_train_dataloader()` and streaming-safe shuffle handling.
- Added chat-template fallback for base models without
`tokenizer.chat_template`.
- Added Sewing Kit loss/helper exports needed by the later bypass core.
- Updated child-state initialization to support composing multiple
pruning mixins.
  - Updated warmup-step resolver to account for gradient accumulation.

  ## Why

The bypass distillation MR needs these reusable pieces, but they are
independently reviewable and useful without adding the bypass
  training stage itself.

Splitting them out keeps the bypass core PR focused on the actual
local-distillation engine.

  ## Tests

  Added focused unit coverage for:

  - Dataloader behavior
  - Bypass loss helpers
  - KV-head pruning utilities
  - Sewing Kit activity/input/function/needle behavior

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* KV-head pruning added for multiple model families; generic pruning
mixin hook available.
* New training dataloader factory for infinite, block-sized training
streams.
  * Vectorwise and batched normalized MSE loss utilities.

* **Improvements**
  * Loss reports now show Δ-from-initial and visual indicators.
* Chat-sample preprocessing tolerates tokenizers without chat templates.
* More robust head-dimension/bias handling and grad-accum-aware warmup
resolver.

* **Tests**
* Extensive unit tests added across dataloaders, losses, pruning, hydra
utils, and sewing-kit.

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1468)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Sepehr Sameni <ssameni@nvidia.com>
2026-05-29 12:01:51 +00:00
Keval Morabia 7ae18651a5 Fix puzzletron tests and make FFN pruning check robust (#1556)
### What does this PR do?

Type of change: Bug fix + test robustness

Fixes the puzzletron GPU tests and makes them robust to GPU vendor /
transformers version drift. The previous exact-equality checks failed in
CI because bf16 attention paths produce different operation orderings on
different GPUs (Ada vs Blackwell), which completely reshuffles
close-in-magnitude FFN importance scores on the small (hidden=256) test
models. Earlier set-membership relaxations weren't enough — the
orderings between GPUs were near-orthogonal.

**Changes:**

*Numerical determinism (root cause):*
- Switch test forward to **fp32 autocast** (`autocast_dtype:
torch.float32` in `validate_model_defaults.yaml`). Combined with TF32
disabled (below), this gives bit-identical results across GPU vendors.
- Reduce `block_size` from 8192 → **2048**: SDPA's flash attention only
supports bf16/fp16, so at fp32 it materializes the full attention
matrix; seq=8192 OOMs (32 GB), seq=2048 fits.
- Strengthen `set_seed`: also disable TF32 on cuda matmul + cudnn so
fp32 matmuls produce bit-identical results across GPU vendors.

*Test structure (orthogonal improvements):*
- Refactor `_assert_*` helpers to `_check_*` that return error lists;
the multiprocess job collects all failures and fails once at the end
(instead of fail-fast). Single CI run now surfaces every mismatch.
- Drop the brittle exact-equality `score[0]` (channel-rank) check. Use
**set-membership at both ends of the importance ranking**:
- Expected least-important channel must land in top-K least-important of
actual.
- Expected most-important channel must land in top-K most-important of
actual.
- `K = max(8, num_channels // 16)` gives slack for any residual drift
without defeating the check.

*Tolerances:*
- `_check_lm_loss` default tolerance: **0.001** for non-MoE (now
essentially bit-deterministic).
- MoE caller passes **0.01** to absorb expert-routing run-to-run noise
(~0.003).
- Nemotron-3-Nano-30B-A3B (hybrid mamba+MoE) remains excluded from
`EXPECTED_LM_LOSS` — its ~0.06 run-to-run noise would require a
tolerance that defeats the check.

*Values:*
- Update `EXPECTED_FFN_PRUNING_VALUES` (with `least_important` +
`most_important` per layer) and `EXPECTED_LM_LOSS` for the new fp32 +
shorter-seq baseline.

### Testing

Verified locally on RTX 6000 Ada across transformers `4.57.0`, `5.6.0`,
`5.9.0`, in both 1-GPU and 2-GPU modes:

|       | 4.57 | 5.6 | 5.9 |
|-------|-----|-----|-----|
| 1-GPU | 8/9 † | 9/9 | 9/9 |
| 2-GPU | — | — | 9/9 |

† Qwen3-VL-30B-A3B fails on 4.57: its lm_loss is 4.92 on 4.57 but
5.03125 on 5.6+. This is a pre-existing cross-major-version drift
documented in a code comment; not introduced by this PR. CI runs
transformers 5.9 (container 26.04).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (test-only change)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-29 07:33:11 +00:00
Eddy Shieh e012f064e1 [Quant] Fix padded NVFP4 MSE calibration (#1557)
### What does this PR do?

Type of change: Bug fix

Fixes static NVFP4 MSE calibration when the weight's last dimension
needs block padding. The reference MSE path now returns fake-quantized
tensors to the padded static block layout through the quantizer helper
before comparing losses, so padded weights such as `512x60` compare as
`2048x16` against `2048x16` instead of `2048x16` against `1920x16`.

The local-Hessian calibration path now reuses the same helper, and a
CUDA regression test covers the forced reference sweep for a padded
`nn.Linear(60, 512)` weight.

Fixes #1552.

### Usage

```python
import copy
import torch
from torch import nn
import modelopt.torch.quantization as mtq

model = nn.Linear(60, 512, bias=False).eval().cuda().bfloat16()
inputs = torch.randn(2, 60, device="cuda", dtype=torch.bfloat16)
mtq.quantize(
    model,
    copy.deepcopy(mtq.NVFP4_W4A4_WEIGHT_MSE_FP8_SWEEP_CFG),
    forward_loop=lambda model: model(inputs),
)
```

### Testing

- `git diff --check`
- `MODELOPT_NVFP4_TRITON_SWEEP=0 uv run --no-sync python - <<'PY' ...`
original padded `Linear(60, 512)` repro: passed
- `uv run --no-sync python -m pytest
tests/gpu/torch/quantization/test_nvfp4_static_quantizer_cuda.py`: 11
passed
- `uv run --no-sync python -m pytest
tests/unit/torch/quantization/test_mse_calibrator.py`: 14 passed
- `uv run --no-sync pre-commit run --files
modelopt/torch/quantization/model_calib.py
tests/gpu/torch/quantization/test_nvfp4_static_quantizer_cuda.py`:
passed

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

Related issue: #1552


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Strengthened quantization calibration robustness with improved state
management and explicit restoration mechanisms
* Unified quantization logic across different calibration methods to
eliminate code duplication and improve consistency
* Refined block-quantization postprocessing behavior for better accuracy

* **Tests**
* Added comprehensive test coverage validating quantization calibration
behavior with padded tensor dimensions and edge cases

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1557?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: eshieh <eshieh@nvidia.com>
2026-05-29 12:22:37 +05:30
Keval MorabiaandClaude Sonnet 4.6 999c99913e Add Nemotron-3-Nano-30B-A3B-BF16 e2e tutorial: Prune + Distill + Quantize + Nemo Evaluator + vLLM deployment (#1376)
### What does this PR do?

Type of change: example/tutorial <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

Add Nemotron-3-Nano-30B-A3B-BF16 e2e tutorial: Prune + Distill +
Quantize + Nemo Evaluator + vLLM deployment

<img width="2079" height="1613" alt="image"
src="https://github.com/user-attachments/assets/19b6ab82-7f01-45df-a0a5-d1c3282b384a"
/>



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* End-to-end Nemotron-3-Nano-30B tutorial: pruning, two‑phase
distillation, FP8 PTQ, evaluation, and vLLM deployment; new ablations
and long‑context analyses.
* Distillation CLI: configurable seed and activation‑recomputation
options.

* **Bug Fixes**
* Preprocessing hardened to skip malformed JSONL and normalize tool‑call
argument formats.

* **Documentation**
* Many README/examples/evaluator docs updated (news list, tokenization
guides, tutorials, configs, and deployment notes).

* **Tests**
* Added test verifying preprocessing handles stringified tool‑call
arguments.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1376?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-28 18:35:59 +00:00
Keval MorabiaandClaude Opus 4.7 5d0441ae3d Create shared Megatron calibration forward loop for prune / quantize with megatron pretraining data style sequence packing (#1501)
## Summary

Replaces the bespoke calibration loops in Megatron-LM and
Megatron-Bridge prune / quantize example scripts with a single shared
utility,
`modelopt.torch.utils.plugins.megatron_calibration.get_megatron_calibration_forward_loop`.

The shared loop iterates a packed calibration dataloader built via
`get_dataset_dataloader(pack=True)` and drives a logits-free prefill
pass through the model so activation hooks fire on every layer.
`pack=True` produces Megatron-LM pretraining-style **global-stream
packing**: all raw samples are concatenated into one EOS-separated token
stream and sliced into uniform-length rows. The trained model has seen
this distribution extensively during pretraining, so the activations
produced during calibration are representative of the model's natural
behavior.

Migrates four call sites:
- `examples/megatron_bridge/prune_minitron.py`
- `Megatron-LM/examples/post_training/modelopt/{prune,quantize}.py`
(separate PR:
[NVIDIA/Megatron-LM#4881](https://github.com/NVIDIA/Megatron-LM/pull/4881))
- `Megatron-Bridge/examples/quantization/quantize.py` (separate PR)

Each call site passes `pack=True` explicitly with an inline comment so
users see the option and know when to flip it. The function-level
defaults (`get_dataset_dataloader(pack=False)`,
`get_megatron_calibration_forward_loop(pack=False)`) remain
back-compat-safe.

Unified defaults across all four sites: `--calib-dataset
nemotron-post-training-dataset-v2`, `--calib-size 1024`,
`--calib-max-sequence-length 4096`, `--calib-batch-size 1`.

## Experimental results

Qwen3-8B on full 100% MMLU (n=14042; binomial 2σ noise floor ≈ ±0.78 pt
at acc ≈ 0.7), 0-shot, eval batch_size=4. Calibration on the default
workload: nemotron-post-training-dataset-v2, seq_length=4096,
calib_batch_size=8.

**Three calibration data shapes compared:**

- **Padded**: one doc per row, padded to `seq_length`, pad tokens flow
through the forward (legacy `get_calib_dataloader` pad+truncate
behavior).
- **Trimmed**: one doc per row, each row trimmed to its real content
length via `attention_mask`, with EOS forced at the last real position;
pad never enters the forward. This is no longer part of this PR.
- **Packed**: global-stream slicing — all docs concatenated
EOS-separated into one token stream, sliced into uniform `seq_length`
rows. Matches Megatron's `.bin`/`.idx` pretraining distribution. Enabled
via `pack=True` in `get_megatron_calibration_forward_loop`.

| Workload | Padded | Trimmed | **Packed** |
|---|---|---|---|
| M-LM NVFP4 quantize (`NVFP4_DEFAULT_CFG`) | 0.707 | 0.708 | 0.709 |
| M-Bridge Minitron prune (Qwen3-8B → 30L / 3584 / 11776 ≈ 6B params) |
0.576 | 0.573 | **0.589** |

### Key findings

- **M-LM quantize quality is calibration-mode-insensitive** for dense
Qwen3-8B NVFP4 — all three modes are nearly identical.
- **M-Bridge prune**: More sensitive to calibration data shape. Packed
wins over Padded on full MMLU.

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-28 18:34:37 +00:00
chochowski 48fbac0ed7 config for expert removal in nemotron3 (#1544)
### What does this PR do?

New example
This PR adds config and updates Nemotron descriptor to support expert
pruning for Nano3 30B-A3B

### Usage

```python
torchrun --nproc_per_node 2 examples/puzzletron/main.py --config examples/puzzletron/configs/nemotron-nano-30b-A3b-v3/nemotron_nano_v3_pruneexp.yaml 2>&1 | tee ./log.txt | grep "Puzzletron Progress"
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- Did you write any new necessary tests?:  ❌ 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for Nemotron Nano 30B model with multiple pruning and
compression strategies
* Enabled expert pruning to reduce model capacity, plus attention
optimization and feed-forward layer compression
* Introduced model validation and solution evaluation configurations for
efficient model optimization workflows

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1544?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: mchochowski <mchochowski@nvidia.com>
2026-05-27 09:07:13 +00:00
Shengliang Xu 9d0d97829a chore(lint): modernize typing (PEP 604/585) and enable UP032 (#1537)
### What does this PR do?

Type of change: chore / refactor (no behavior change)

Two small lint-cleanup commits:

**1. `chore(typing): modernize Union/Optional/List to PEP 604 / 585
syntax`** (8 files)

- Replace `X = Union[A, B] # noqa: UP007` with `X: TypeAlias = A | B`
for the six module-level type aliases (`ModelLike`, `Criterion`,
`NodeTarget`, `CalibrationDataType`, `Hparam.Importance` /
`ActiveSlice`). The `TypeAlias` annotation is required so mypy continues
to treat them as aliases under PEP 604.
- Modernize forward-ref unions in `modelopt/onnx/quantization/autotune/`
to full-string forward refs (e.g. `"RegionPattern | None"`).
- Update docstring type tags in
`examples/puzzletron/evaluation/hf_deployable_anymodel.py`.

**2. `chore(lint): remove UP032 ignore and convert .format() to
f-strings`** (10 files)

- Drop `UP032` from `extend-ignore` in `pyproject.toml`.
- Auto-convert 19 `"...".format(...)` calls to f-strings across export
plugins, examples, tests, and tools. One conversion in
`modelopt/torch/utils/plugins/megatron_generate.py` was wrapped manually
to stay under the 100-char limit.

**Intentionally left as-is:**

- `tools/launcher/slurm_config.py` keeps its `# ruff: noqa: UP045` —
nemo_run's CLI parser can't introspect PEP 604 optional annotations.
- `modelopt/torch/puzzletron/*` is **not** touched. The subtree disables
ruff's `UP` family entirely (per-file-ignore `"UP"`) while migration is
in progress, and converting `Optional[X]` to `X | None` there would
silently break runtime introspection in
`block_config._get_dataclass_type` that uses `get_origin(tp) is
typing.Union` (PEP 604 unions return `types.UnionType` from
`get_origin`, not `typing.Union`). Best revisited when puzzletron's lint
carve-out is narrowed.
- `UP038` (`isinstance(x, (int, float))` → `isinstance(x, int | float)`)
— ruff has officially deprecated this rule; PEP 604 in isinstance is
slightly slower and misleads readers about PEP 695 / `Optional`. Ignore
kept.

### Usage

No user-facing API changes.

### Testing

- Pre-commit hooks (ruff check, ruff format, mypy, bandit, license) pass
on both commits.
- Ruff status against `main`: 37 unrelated pre-existing findings
(W291/W293/E501/RUF005/PLR1704); zero new findings introduced by this
PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — Runtime behavior of the six
type aliases changes from a `typing.Union` instance to
`types.UnionType`. Downstream code introspecting via `get_origin(...) is
typing.Union` on these aliases would break, but no in-repo caller does
this on them. (The introspection in
`modelopt/torch/puzzletron/block_config.py` operates on user-supplied
dataclass field types, none of which are these aliases.)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (no behavior change)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — internal style refactor; happy to add a Misc note if reviewers want
one.
- Did you get Claude approval on this PR?: ❌ — not yet.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Style**
* Modernized type annotations across the codebase to use Python 3.10+
union syntax and TypeAlias where appropriate.
* Standardized string formatting to f-strings, improving clarity of
logs, errors, and validation messages.

* **Chores**
  * Updated linting configuration to reflect modern typing/style rules.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1537?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-05-26 16:47:31 -07:00
Ajinkya Rasane cd2ea46e1d [OMNIML-4730] Support quantized nn.Embedding (#1495)
### What does this PR do?

Type of change: new feature

Register `nn.Embedding` in `QuantModuleRegistry` so the embedding table
and lookup activations participate in quantization end-to-end:

- New `modelopt/torch/quantization/nn/modules/quant_embedding.py`
exposes `weight_quantizer` (embedding table), `output_quantizer` (lookup
activations, off by default), and an `input_quantizer` placeholder.
Embedding inputs are integer indices that cannot be fake-quantized, so
direct `enable()` / `enable_quant()` / `enable_calib()` calls on
`input_quantizer` raise; `forward()` never invokes `input_quantizer` at
all, so a back-door flip of `_disabled` is a no-op rather than a crash.
Wildcard configs (`*input_quantizer`) are accepted silently so the stock
deny-all → enable-wildcards → opt-out pattern in `NVFP4_DEFAULT_CFG` and
friends still works.
- `default_disabled_quantizers.yaml` installs `parent_class:
nn.Embedding, enable: false` so embedding quantization is opt-in and
existing model behavior is unchanged.
- `is_quantized_linear` in `core_utils.py` early-returns `False` for
`nn.Embedding` so AWQ / SmoothQuant / SVDQuant don't treat it as a GEMM
op.
- `_process_quantized_modules` in `unified_export_hf.py` routes
quantized `nn.Embedding` modules through `_export_quantized_weight`, so
the exported checkpoint contains the packed NVFP4 / FP8 / INT bytes plus
`weight_scale*` buffers, exactly like Linear layers.

### Usage

```python
import copy
import torch.nn as nn
import modelopt.torch.quantization as mtq
from modelopt.torch.export import export_hf_checkpoint

# Opt embeddings into the stock NVFP4 config — the YAML default is opt-out.
cfg = copy.deepcopy(mtq.NVFP4_DEFAULT_CFG)
cfg["quant_cfg"].append(
    {
        "parent_class": "nn.Embedding",
        "quantizer_name": "*weight_quantizer",
        "cfg": {"num_bits": (2, 1), "block_sizes": {-1: 16, "type": "dynamic", "scale_bits": (4, 3)}},
    }
)

model = mtq.quantize(model, cfg, forward_loop)
export_hf_checkpoint(model, export_dir="./out")
# out/model.safetensors contains: embedding.weight (uint8, NVFP4-packed),
# embedding.weight_scale (FP8 E4M3 per-block), embedding.weight_scale_2 (FP32).
```

### Testing

- New unit tests `tests/unit/torch/quantization/test_quant_embedding.py`
cover: default quantizer state, no-quant identity, per-tensor and
per-row weight fake quant against the manual
`tensor_quant.fake_tensor_quant` reference, output quantizer activation,
locked-mutator raises (parametrized over `enable` / `enable_quant` /
`enable_calib`), and the wildcard-then-opt-out pattern. All 8 cases pass
under `CUDA_VISIBLE_DEVICES=""`. Export-path coverage (weight packing +
tied-weight guard) lives in
`tests/gpu/torch/export/test_export_embedding.py` because the export
path hits `torch.cuda.empty_cache()`.
- Verified end-to-end on an embedding-only model: `mtq.quantize` with
`NVFP4_DEFAULT_CFG` + the embedding opt-in produces `embedding.weight
(uint8)`, `embedding.weight_scale (float8_e4m3fn)`,
`embedding.weight_scale_2 (float32)` in the exported safetensors, with
`"quant_algo": "NVFP4"` in `hf_quant_config.json`.
- End-to-end smoke test on TinyLlama-1.1B-Chat-v1.0 via `hf_ptq.py`
(32-sample calibration on the default cnn_dailymail +
nemotron-post-training datasets), sweeping the embedding's
`weight_quantizer` across FP8 (per-tensor E4M3) and NVFP4 (E2M1,
block_size=16, FP8 scales) while varying the Linear-layer format. KV
cache is FP8 in all runs. Greedy generation (`max_new_tokens=100`) from
the standard CNN/DailyMail preview prompt:

> `<s>` LONDON, England (Reuters) -- Harry Potter star Daniel Radcliffe
gains access to a reported £20 million ($41.1 million) fortune as he
turns 18 on Monday, but he insists the money won't cast a spell on him.
[…] The Londoner has filmed a TV movie called "My Boy Jack," about
author Rudyard Kipling and

- **Baseline FP16 (no quantization):** *"his son, who died at the age of
18. He also has a role in the upcoming "The Prestige" and is set to star
in the remake of "The League of Extraordinary Gentlemen." "I'm not going
to be a star for the rest of my life," he said. "I'm going to be a star
for a while, but I'm not going to be a star for the rest of my life."*
- **FP8 Linear + FP8 embedding** (`hf_quant_config.json`: `"quant_algo":
"FP8"`, `exclude_modules: ["lm_head"]`, so `model.embed_tokens` is FP8):
*"his son, who died at the age of 18. He also has a role in the upcoming
"The Prestige" and is set to star in the upcoming "The Golden Compass."
"I'm not going to be a star forever," he said. "I'm going to be a star
for a while, but I'm not going to be a star forever."* (matches FP16
token-for-token until a single lexical drift)
- **FP8 Linear + NVFP4 embedding** (`model.embed_tokens: W4A16_NVFP4,
group_size=16`; Linear: `FP8`): *"his son, who died in the First World
War. He also has a role in the upcoming "The Golden Compass" and is set
to star in the BBC's "The Bill." "I'm not going to be a star for the
rest of my life," he said. "I'm going to be a star for a while, but I'm
not going to be a star for the rest of my life."* (embedding format swap
is visible — first phrase diverges from baseline; rest of the sentence
still tracks)
- **NVFP4 Linear + FP8 embedding** (`model.embed_tokens: FP8`; Linear:
`NVFP4, group_size=16`): *"his son, and is also in talks to play the
title role in a film version of the stage play "The Importance of Being
Earnest." "I've been offered a lot of things," he said. "I've been
offered a lot of things, but I've also been offered a lot of things that
I've said no to." Radcliffe's first movie, "The Tailor of Gloucester,"
was released in the UK in"* (expected drift at W4A4)
- **NVFP4 Linear + NVFP4 embedding** (`model.embed_tokens: W4A16_NVFP4,
group_size=16`; Linear: `NVFP4, group_size=16`): *"his son, who died in
a car accident. Radcliffe will also star in the film, which is being
shot in London. "It's a very different kind of film," he said. "It's a
very personal story, and I've been very lucky to be able to do it."
Radcliffe's next film, "The Girl with the Dragon Tattoo," is due to be
released in the U.S. In December. He will also"* (full-NVFP4 path;
quantized embedding row is `W4A16_NVFP4` because the embedding's input
quantizer is permanently disabled by design)

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ — embedding quantizers are
opt-in via `parent_class: nn.Embedding, enable: false` in
`default_disabled_quantizers.yaml`, so existing model behavior is
unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — will run `/claude review`
after the PR is up.

### Additional Information

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Opt-in quantization for embedding layers: weight quantization enabled
by default, optional output quantization; input quantization remains
permanently disabled.

* **Bug Fixes**
* Preserve tied embedding weights during export by skipping packing
(emits warning) to avoid breaking weight sharing.

* **Tests**
* Added unit tests covering embedding quantization behavior, export
packing, calibration, and tied-weight scenarios.

* **Documentation**
  * Changelog updated with embedding quantization notes.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1495?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2026-05-26 09:22:24 -07:00
3ff15ccef3 Add support for postprocess exported model for block scale swizzling and support for different padding strategy (#1195)
### What does this PR do?

Type of change: ?  new feature

<!-- Details about the change. -->
Adds post-processing support for exported diffusion model checkpoints to
enable NVFP4 block scale swizzling and configurable padding strategies.
This allows exported quantized checkpoints to be directly consumed by
inference runtimes (e.g., ComfyUI with comfy_kitchen) that require
cuBLAS 2-D block-scaling-factors layout.
Changes:
1) Unified post-processing step (_postprocess_safetensors): Loads saved
safetensors files and applies merge, padding, swizzle, and quantization
metadata injection in a single pass.
2) NVFP4 scale swizzle (swizzle_nvfp4_scales): Rearranges block scales
from ModelOpt's flat [rows, cols // 16] layout to cuBLAS 2-D tiled
layout per the cuBLAS specification.
3) Configurable padding (pad_nvfp4_weights): Pads NVFP4 weight and scale
tensors to multiples of 16, with "row" (rows only) or "row_col" (both
dimensions) strategies.
4) Standalone quantization metadata (build_layerwise_quant_metadata):
Extracted from merge_diffusion_checkpoint so _quantization_metadata can
be injected independently of merging — works for both merged (LTX-2) and
standalone (Flux2) exports.
5) Bug fix (conversion.py): Wrapped yield in try/finally in
set_quantizer_by_cfg_context so quantizer states are always restored,
fixing an issue when yield fails.
### Usage

```python
# LTX-2 export with merge + swizzle + padding
export_hf_checkpoint(
  pipeline,
  export_dir="./output",
  merged_base_safetensor_path="./ltx-2-22b-dev.safetensors",
  enable_swizzle_layout=True,
  padding_strategy="row_col",
  enable_layerwise_quant_metadata=True,
)
# Flux2 standalone export with swizzle + padding (no merge needed)
export_hf_checkpoint(
  transformer,
  export_dir="./output",
  enable_swizzle_layout=True,
  padding_strategy="row_col",
)
# Via quantize.py CLI
python quantize.py \
  --model ltx-2 --format fp4 \
  --extra-param merged_base_safetensor_path=./ltx-2-22b-dev.safetensors \
  --extra-param enable_swizzle_layout=true \
  --extra-param padding_strategy=row_col \
  --hf-ckpt-dir ./output
```

### Testing
1) Exported LTX-2.3 NVFP4 with swizzle + padding + merged base
checkpoint. Verified checkpoint has correct uint8 weights, float8_e4m3fn
scales in swizzled layout, and _quantization_metadata . Ran the
checkpoint with ComfyUI
2) Exported Flux2 NVFP4 with swizzle + padding. Verified checkpoint has
correct uint8 weights, float8_e4m3fn scales in swizzled layout, and
_quantization_metadata . Ran the checkpoint with ComfyUI

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Diffusers export: optional NVFP4 support — swizzle layout, row/row_col
padding, and optional per-layer quantization metadata; exports are now
post-processed to apply these options.
* Export flow accepts new flags to enable swizzle, padding strategy, and
layerwise metadata.

* **Bug Fixes**
* Quantizer context manager now always restores state, including on
exceptions.

* **Tests**
* Added unit tests for NVFP4 padding, swizzling, metadata injection, and
post-processing.

* **Documentation**
  * README example updated to show swizzle and padding flags.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: YASH Nankani <ynankani@2u1g-x570-0073.ipp2a1.colossus.nvidia.com>
Signed-off-by: ynankani-nv <ynankani@nvidia.com>
Signed-off-by: YASH Nankani <ynankani@dl325g11-1979.ipp2a2.colossus.nvidia.com>
Signed-off-by: YASH Nankani <ynankani@dl325g11-0771.ipp4a1.colossus.nvidia.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: YASH Nankani <ynankani@2u1g-x570-0073.ipp2a1.colossus.nvidia.com>
Co-authored-by: YASH Nankani <ynankani@dl325g11-1979.ipp2a2.colossus.nvidia.com>
Co-authored-by: YASH Nankani <ynankani@dl325g11-0771.ipp4a1.colossus.nvidia.com>
2026-05-22 10:36:59 +00:00
realAsma b02e888550 fix(quantization): accept QuantizeAlgorithmConfig in get_modelike_from_algo_cfg (#201) (#1528)
Fixes #201

`get_modelike_from_algo_cfg` was typed to accept `QuantizeAlgoCfgType`
(which includes `QuantizeAlgorithmConfig` instances), but the body only
handled list / str / None / dict and raised `ValueError("Invalid config
type")` for `QuantizeAlgorithmConfig` instances. This broke the
documented pattern of passing typed config objects (`MaxCalibConfig`,
`AWQLiteCalibConfig`, etc.) via `quant_config['algorithm']`.

The fix pre-converts a `QuantizeAlgorithmConfig` instance to its
`model_dump()` dict at function entry, so the existing dict path handles
it unchanged. Two-line change; no new branch needed.

Verification: added a unit test in
`tests/unit/torch/quantization/test_mode.py` that builds a
`MaxCalibConfig`, calls `get_modelike_from_algo_cfg` on both the object
and its equivalent dict, and asserts the two paths produce the same
tuple. Also covers the list-of-object path.

Opened by OnCallBot on behalf of realAsma. Please review the diff before
merging.

cc Chenjie Luo (module owner for torch.quantization)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* `get_modelike_from_algo_cfg` now accepts configuration objects
directly alongside existing input formats.

* **Tests**
* Added regression test to verify configuration object handling works
correctly.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1528?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-05-21 22:39:36 +00:00
kaix-nv c9098b63fb [4/n] Add vLLM integration for modelopt sparse attention (#1127)
### What does this PR do?

Type of change: New feature, new example, new tests, documentation.

Adds vLLM integration for ModelOpt sparse attention with paged KV cache
support.

This PR extends the ModelOpt Triton flash attention path so K/V can be
read directly from vLLM's paged KV cache through `block_table` lookup.
This avoids gather-to-contiguous copies when serving exported
sparse-attention checkpoints with vLLM.

The vLLM integration swaps vLLM's `FlashAttentionImpl` with
`ModelOptSparseAttentionImpl` after model load. The sparse configuration
is read from the exported checkpoint's `config.json`
`sparse_attention_config` block, written by
`examples/llm_sparsity/attention_sparsity/hf_sa.py`.

The restored checkpoint metadata supports:

- calibrated skip-softmax metadata (`threshold_scale_factor`,
`target_sparse_ratio`)
- N:M sparse-softmax metadata (`sparsity_n`, `sparsity_m`)
- dense token preservation metadata (`dense_sink_tokens`,
`dense_recent_tokens`)

The vLLM path uses ModelOpt Triton for sparse prefill launches.
Decode-only launches, cascade/prefix-cache metadata, and launches
without active sparse work delegate back to vLLM FlashAttention.

### Limitations

- Sparse attention is enabled for sparse prefill only.
- Decode-only launches currently fall back to vLLM FlashAttention.
- Attention sinks from vLLM FlashAttention are rejected until the
ModelOpt Triton path supports them.
- CUDA graph capture is not validated with this sparse attention path
yet; use `--enforce-eager`.
- Quant-only serving remains covered by `vllm_serve_fakequant.py`.
- Combined sparse attention + quantization serving is not handled by
this launcher in this PR and is planned as follow-up work.

### Usage

Export a checkpoint with calibrated skip-softmax and sparse24 metadata:

```bash
python examples/llm_sparsity/attention_sparsity/hf_sa.py \
    --pyt_ckpt_path /path/to/hf-model \
    --sparse_attn skip_softmax_calib_sparse24 \
    --target_sparse_ratio 0.5 \
    --calib_samples 64 \
    --calib_max_seqlen 16384 \
    --calib_chunk_size 4096 \
    --seq_len 2048 \
    --export_dir /path/to/modelopt-skipsoftmax-sparse24-export
```

Serve the exported checkpoint with the vLLM sparse-attention launcher:

```bash
PYTHONPATH=$PWD python examples/vllm_serve/vllm_serve_sparse_attn.py \
    /path/to/modelopt-skipsoftmax-sparse24-export \
    --tensor-parallel-size 8 \
    --host 0.0.0.0 \
    --port 8000 \
    --trust-remote-code \
    --enforce-eager
```

Send a request through the OpenAI-compatible endpoint:

```bash
curl http://localhost:8000/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{
      "model": "/path/to/modelopt-skipsoftmax-sparse24-export",
      "messages": [{"role": "user", "content": "Explain sparse attention in one paragraph."}],
      "max_tokens": 128
    }'
```

### Testing

GitHub CI on the latest commit is green:

- DCO
- code-quality
- docs build / deploy preview
- unit tests, including Linux, Windows, multi-version, partial-install,
and launcher jobs
- example tests
- GPU tests, including required GPU gate
- regression tests, including required regression gate
- `codecov/project`

Focused test coverage added/updated for this PR includes:

-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_worker.py`
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_config.py`
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attention_conversion.py`
-
`tests/unit/torch/sparsity/attention_sparsity/test_triton_skip_softmax.py`
- `tests/gpu/torch/sparsity/attention_sparsity/test_vllm_plugin.py`
- `tests/gpu/torch/kernels/common/attention/test_triton_fa_paged.py`
-
`tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_skip_softmax.py`
-
`tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_sparse_nm.py`
-
`tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_calibrate.py`

Manual / NEL eval validation:

- Served a ModelOpt exported sparse-attention checkpoint through
`examples/vllm_serve/vllm_serve_sparse_attn.py`.
- Launched RULER64K NEL evals on DFW with `coreai_nvfm_llm`.
- Current partial RULER64K prediction scores, before final `results.yml`
is written:
  - `skipsoftmax-only`: 98.59% over 4500 flushed samples
  - `skipsoftmax-r0.7`: 99.70% over 1000 flushed samples
  - `skipsoftmax-r0.9`: 99.70% over 1000 flushed samples

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - no new
PIP dependency.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A - documentation, examples, and tests are updated for this vLLM
integration path.

### Additional Information

Follow-up work:

- Validate and enable CUDA graph capture for the sparse vLLM path.
- Add combined sparse attention + quantization serving once the combined
path is tested.
- Investigate whether skip-softmax should also be enabled during decode.

---------

Signed-off-by: Kai Xu <kaix@nvidia.com>
2026-05-20 14:11:25 -07:00
Shengliang Xu e4dc0205d1 [OMNIML-4775] Move built-in PTQ quantization configs to YAML (#1423)
### What does this PR do?

Type of change: refactor

This PR moves the built-in PTQ quantization config definitions out of
hard-coded Python dictionaries and into schema-backed YAML config files,
and factors shared blocks into reusable composable snippets.

- Adds reusable numeric config snippets under
`modelopt_recipes/configs/numerics/`.
- Adds YAML presets for the built-in model PTQ configs under
`modelopt_recipes/configs/ptq/presets/model/`.
- Adds YAML presets for KV-cache quantization configs under
`modelopt_recipes/configs/ptq/presets/kv/`.
- Adds YAML presets for the Diffusers-specific PTQ configs under
`modelopt_recipes/configs/ptq/presets/diffusers/` and re-points
`examples/diffusers/quantization/config.py` constants at them via
`load_config`.
- Adds reusable KV quantization units (`kv_fp8_affine`, `kv_nvfp4`,
`kv_nvfp4_affine`, `kv_nvfp4_rotate`, `kv_*_cast` variants) under
`modelopt_recipes/configs/ptq/units/`.
- Adds reusable model-side units following the
`component_numerics[_type]` convention:
- `attention_qkv_fp8` — FP8 E4M3 on attention q/k/v bmm and softmax
quantizers; shared by `model/` and `diffusers/` `nvfp4_fp8_mha` presets.
- `block_sparse_moe_nvfp4` — NVFP4 W4A4 on `*block_sparse_moe*`
weight/input quantizers; shared by `nvfp4_mlp_only`,
`nvfp4_experts_only`, `nvfp4_omlp_only`.
- `experts_nvfp4` — NVFP4 W4A4 on `*.experts.*` weight/input quantizers;
shared by `nvfp4_mlp_only` and `nvfp4_experts_only`.
- Switches the existing 5 NVFP4 presets (default + awq lite/clip/full +
svdquant) and 4 mamba_moe presets to `$import` the existing
`w4a4_nvfp4_nvfp4` / `w8a8_fp8_fp8` units instead of re-inlining the
same weight+input quantizer pairs.
- Moves the recently-added `W4A16_NVFP4_CFG` to YAML
(`presets/model/w4a16_nvfp4.yaml`) composed from the existing
`units/w4_nvfp4` snippet.
- Updates `modelopt.torch.quantization.config` built-in config constants
to load `QuantizeConfig` objects from YAML with `load_config(...,
schema_type=QuantizeConfig).model_dump(exclude_unset=True)` via a new
`_load_quantize_config_dict` helper; the constants remain plain
`dict[str, Any]` for backwards compatibility with consumers that do
mapping-style mutation (e.g. `entry["cfg"]` assignment).
- Simplifies the cfg-list loader (`_load_quantizer_cfg_dict_list`) down
to a 4-line list/single normalization now that the three call sites all
load schema-typed YAMLs.
- Adds/updates recipe loader coverage for built-in schema-backed config
snippets.

### Latent-bug fixes surfaced by the refactor

Two small correctness fixes are included alongside the mechanical
refactor; flagging them explicitly:

- **`examples/diffusers/quantization/quantize.py`** — adds an explicit
`base_cfg = copy.deepcopy(base_cfg)` before applying runtime overrides.
The existing `# Build a fresh config dict so we never mutate the global
constants` comment had been aspirational only; in practice
`reset_set_int8_config` accumulated `PercentileCalibrator` entries into
`mtq.INT8_SMOOTHQUANT_CFG`/`INT8_DEFAULT_CONFIG` across repeated calls,
and `set_quant_config_attr` added `trt_high_precision_dtype` keys into
globally-shared cfg dicts. The deepcopy makes the code match the
comment.
- **`choices` set in `modelopt/torch/quantization/config.py`** — adds
`MXFP6_DEFAULT_CFG` and `NVFP4_W4A4_WEIGHT_LOCAL_HESSIAN_CFG` to the
documented public set of valid `mtq.*_CFG` names. Both constants exist
on main but were missing from `choices`, so CLIs that gate on
`mtq.config.choices` (e.g., `hf_ptq.py --qformat`) couldn't reach them
even though the configs themselves were fully supported.

### Usage

Existing Python imports continue to work:

```python
import modelopt.torch.quantization as mtq

cfg = mtq.FP8_DEFAULT_CFG
model = mtq.quantize(model, cfg, forward_loop)
```

The built-in constants are plain `dict[str, Any]` (sparse — only
explicitly-set fields are present), but their definitions now come from
YAML snippets and presets composed through the existing `$import`
system.

Reusable YAML snippets can be composed through `$import`, for example:

```yaml
# modelopt-schema: modelopt.torch.quantization.config.QuantizeConfig
imports:
  base_disable_all: configs/ptq/units/base_disable_all
  w4a4_nvfp4_nvfp4: configs/ptq/units/w4a4_nvfp4_nvfp4
  default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers

algorithm: max
quant_cfg:
  - $import: base_disable_all
  - $import: w4a4_nvfp4_nvfp4
  - $import: default_disabled_quantizers
```

### Testing

Local checks run:

- `nox -s "unit-3.10(torch_211, tf_latest)"` — 2329 passed, 12 skipped.
- `nox -s pre_commit_all` — all hooks pass (ruff check / ruff format /
mypy / YAML format / license / bandit / markdownlint).
- YAML parse + `$import` resolution sanity check across all changed
config files.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ Existing built-in Python config
constants keep the same public names and dict semantics.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ Adds/updates recipe loader
coverage for schema-backed built-in snippets.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌

### Additional Information

This PR was previously stacked on #1405, which has since merged to
`main`. The branch has been rebased onto `main` and no longer depends on
any other open PR.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Many new quantization numeric configs and PTQ presets added
(INT4/INT8/MXFP4/MXFP6/MXFP8/MXINT8/NVFP4), plus Diffusers, KV-cache
(affine/cast/rotate) and MLP/MoE-targeted presets.

* **Refactor**
* Presets and shared snippets migrated to schema-backed YAML sources and
centralized loading; INT8 percentile calibration avoids mutating shared
base configs.

* **Tests**
* Tests now discover packaged config snippets at runtime and validate
import/append behaviors.

* **Documentation**
  * Presets README and numerous header descriptions updated.

* **Chores**
  * Minor typing and script improvements.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1423?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-05-20 08:21:25 -07:00
Chenjie LuoandClaude Opus 4.7 a5bc6f8123 Add DATASET_COMBOS for grouped calibration datasets (#1508)
## Summary

- Add ``DATASET_COMBOS`` to ``modelopt.torch.utils.dataset_utils`` —
single ``--dataset`` tokens that fan out to several entries in
``SUPPORTED_DATASET_CONFIG``. The per-entry ``num_samples`` is split
evenly across the members inside ``get_dataset_dataloader``.
- Two initial combos:
- ``cnn_nemotron_v2_mix`` → ``cnn_dailymail`` +
``nemotron-post-training-dataset-v2``. Replaces the hardcoded
two-element fallback list in ``hf_ptq.py`` when ``--dataset`` is
omitted.
- ``nemotron-post-training-v3`` → the seven ``nvidia/Nemotron-*`` SFT
datasets registered in #1498 (mirroring the upstream
[`nemotron-post-training-v3`
collection](https://huggingface.co/collections/nvidia/nemotron-post-training-v3)).
- ``get_supported_datasets()`` now appends combo names so they show up
in ``--dataset`` help.
- ``hf_ptq.py``'s default ``--calib_size`` bumped from ``512`` to
``1024`` so the ``cnn_nemotron_v2_mix`` combo's even split preserves the
previous total sample count (was 512 per-dataset × 2 datasets = 1024;
now 1024 split → 512 per-dataset × 2). ``--calib_size`` now denotes the
total calibration budget regardless of combo cardinality.
- Reject mixing a combo with one of its member datasets in the same
``--dataset`` list (e.g. ``cnn_dailymail,cnn_nemotron_v2_mix``) — combo
would otherwise double-sample the explicit member with a smaller
per-member quota.
- Reject combo names in ``get_dataset_samples``; combos are
dataloader-only. The error message points callers to
``get_dataset_dataloader``.
- Validate ``DATASET_COMBOS`` at import time: empty member lists, name
collisions with ``SUPPORTED_DATASET_CONFIG``, and references to unknown
datasets raise ``ValueError`` up front.

## Test plan

End-to-end validated against ``/hf-local/Qwen/Qwen3.5-0.8B`` via
``get_dataset_dataloader`` on the actual streamed data, plus 5 new unit
tests in ``TestDatasetCombosExpansion`` (all 44 tests in
``test_dataset_utils.py`` pass with no regressions).

- [x] ``python -c "from modelopt.torch.utils.dataset_utils import
DATASET_COMBOS, get_supported_datasets; assert 'cnn_nemotron_v2_mix' in
get_supported_datasets() and 'nemotron-post-training-v3' in
get_supported_datasets()"``
- [x] ``hf_ptq.py`` with no ``--dataset`` flag still calibrates on
cnn_dailymail + nemotron-post-training-dataset-v2 with the same total
sample count as before.
- [x] ``--dataset nemotron-post-training-v3 --calib_size 1024``
allocates 146 per member across the seven Nemotron datasets; full
1022-sample dataloader builds without error.
- [x] ``--dataset cnn_dailymail,nemotron-post-training-v3 --calib_size
256,1024`` composes correctly: 256 from cnn_dailymail (as a plain entry)
plus the 7-way split from the combo. (The earlier
``cnn_dailymail,cnn_nemotron_v2_mix`` example is rejected by design
since ``cnn_dailymail`` is a member of that combo.)
- [x] ``--dataset cnn_dailymail,cnn_nemotron_v2_mix`` raises
``ValueError`` with a clear message.
- [x] ``get_dataset_samples("cnn_nemotron_v2_mix", ...)`` raises
``ValueError`` pointing to ``get_dataset_dataloader``.
- [x] Unit tests: ``pytest
tests/unit/torch/utils/test_dataset_utils.py`` — 44 passed.

## Post-validation fix

End-to-end testing surfaced that the original
``nemotron-sft-agentic-v2`` entry kept the two splits
(``interactive_agent``, ``tool_calling``) that pyarrow's streaming JSON
reader cannot parse, and excluded ``search`` which is the only clean
split. Failures reproduce deterministically across cache wipes with
``force_redownload``, so they are content-level defects in the published
JSONL files at the pinned revision, not local artifacts:

- ``interactive_agent`` — heterogeneous schema
(``Column(.../member_id/type) changed from string to array``) at JSONL
row 4.
- ``tool_calling`` — malformed JSON row in a later shard, fails at
sample ~885 with ``Missing a closing quotation mark in string``.
- ``search`` — streams cleanly (verified to 2500 samples).

Commit ``10f3cfd`` corrects ``nemotron-sft-agentic-v2`` to use only
``search``, with an updated comment. The CHANGELOG calls this out as a
separate bullet so the behavior change on a previously-released dataset
entry from #1498 is discoverable.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Dataset combo support: a single dataset token can expand into multiple
registered datasets with even sample splitting; predefined combos added
(e.g., cnn_nemotron_v2_mix, nemotron-post-training-v3) and listed as
supported.

* **Updates**
  * Default dataset when none specified now uses cnn_nemotron_v2_mix.
  * Calibration size default increased from 512 to 1024.

* **Bug Fixes**
* nemotron-sft-agentic-v2 now uses only the deterministic "search" split
to avoid streaming JSON errors.

* **Tests**
* Added coverage for combo expansion, splitting, overlap validation, and
rejection behavior.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1508?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 03:08:43 +05:30
Keval Morabia d7733562dd fix(prune): Minitron HybridModel + GPT-family fused-TE-spec import/export (#1518)
## Summary

Split out of #1501 so the pack=True calibration packing change can land
independently. This PR carries the pruning + export-side fixes.

**Pruning bug fixes**

- Register `HybridModel` (parent of `MambaModel` in modern Megatron-LM)
under a new `HAS_HYBRID` flag so `mcore_minitron` actually prunes
Nemotron-H et al. Previously `HybridModel` instances fell through
`convert_to_dynamic`, got `freeze()`-ed (collapsing `hidden_size` /
`num_layers` to a single choice), and produced unloadable saved
checkpoints with mixed pruned/unpruned dims.
- Replace the `isinstance(MambaModel)` gate in `_get_hybrid_pattern_key`
with attribute-presence detection so both `MambaModel` (still using
`hybrid_override_pattern`) and plain `HybridModel`
(`hybrid_layer_pattern`) are handled uniformly.
- Track `in_features` as a dynamic attribute on
`_DynamicTEQKVLayerNormColumnParallelLinear` so TE's forward-time
`inp_shape[-1] == in_features` assertion holds when `hidden_size` is
pruned.
- Dedupe MambaModel / HybridModel divisor dict into `_HYBRID_DIVISORS`.

**Fused-TE-spec import/export for GPT-family**

- Importer: prefer per-context keys (`fused_input_layernorm`,
`fused_pre_mlp_layernorm`); fall back to legacy `fused_norm` for
Nemotron-H back-compat. **Raise `KeyError`** when a fused-TE model has
neither rule registered — the branch only fires when the model uses
fused `TELayerNormColumnParallelLinear`, so a missing rule is
unambiguously a plugin misconfig that would otherwise ship a
chance-accuracy checkpoint.
- Exporter: mirror the same fallback chain in `_get_fused_norm_weight`
so GPT-family models round-trip cleanly back to HF.
- Add the new rules to Qwen3, Qwen2.5, Llama, Llama4 (MoE-only, only
`fused_input_layernorm`), DeepSeek, GptOss (MoE-only, only
`fused_input_layernorm`) import and export mappings.
- Preserve TE `_extra_state` from the existing module state dict (don't
blank to `None`) at both call sites in the importer.

**Misc**

- `megatron_prefill`: `.contiguous()` on the logits slice before
`broadcast_from_last_pipeline_stage` — broadcast asserts contiguity
which fails when SP pads `seq_length` to a multiple of TP.
- `megatron_mmlu`: accept `mmlu_dataset` kwarg so callers can point at a
local copy of `cais/mmlu`.
- `warn_rank_0`: auto-bump `stacklevel` by 1 inside the wrapper so
callers' warnings point at user code, not at the wrapper frame.
- `tools/launcher/examples/Qwen/Qwen3-8B/megatron_lm_ptq.yaml`: bump
`mmlu_lower_bound` 0.68 → 0.75 (validated end-to-end with the fused-norm
import fix).
- CHANGELOG: bug-fix entry for the importer; date correction on the 0.44
entry.

## Consumer

Megatron-LM PR https://github.com/NVIDIA/Megatron-LM/pull/4807 —
`prune.py` / `mmlu.py` consume these APIs and currently ship inline WARs
against released 0.44. Once 0.45 ships and the modelopt pin is bumped,
those WARs collapse to one-liners.

Related: #1501 (calibration packing).

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Importer/exporter now correctly load fused LayerNorm weights for
GPT-family models, preferring context-specific fused keys with a legacy
fallback.

* **New Features**
  * Hybrid Mamba/HybridModel support added for pruning/NAS workflows.
* MMLU evaluation accepts a customizable dataset path (default:
"cais/mmlu").

* **Improvements**
* Extended export/import mappings and state handling across DeepSeek,
GPT, Llama, Qwen; ensured last-stage logits are contiguous.

* **Documentation**
  * Updated changelog entry and release date adjustment.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1518?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-20 01:14:12 +05:30
Shengliang Xu f5650bd02c Schematize config loading and quantizer config entries (#1405)
### What does this PR do?

This PR makes ModelOpt config loading schema-aware and moves `quant_cfg`
entries from `TypedDict` validation to Pydantic validation.

Key changes:

- `load_config()` now returns validated schema instances when a schema
is provided through `schema_type=...` or declared with `#
modelopt-schema:`.
  - Without a schema, it still returns the raw resolved dict/list.
- With a schema, it returns the validated schema object, such as
`QuantizeConfig` or `list[QuantizerCfgEntry]`.

- `QuantizerCfgEntry` is now a `ModeloptBaseConfig` Pydantic model.
- The “must specify `cfg`, `enable`, or both” rule is enforced wherever
entries are constructed.
- Enabled entries must provide non-empty `cfg` values when `cfg` is
present.
- Existing mapping-style access like `entry["cfg"]` and
`entry.get("enable")` continues to work.

- `RecipeMetadataConfig` is now a `ModeloptBaseConfig`, so recipe
metadata uses the same schema validation path.

- Recipe loading now delegates shape validation to `load_config()`
instead of manually checking loaded dicts.

- `normalize_quant_cfg_list()` now accepts:
  - `Sequence[QuantizerCfgEntry]`
  - `Sequence[Mapping[str, Any]]`
  - legacy flat mapping configs, with deprecation warnings

- Public quantization config constants remain plain dict/list structures
for backward compatibility, even when they are built from
schema-validated YAML snippets.

### Behavior changes

- `load_config()` may now return a schema instance instead of a raw
`dict`/`list` when a schema is available.
Callers that checked `isinstance(result, dict)` should use
`isinstance(result, Mapping)` or check the expected schema type.

- Normalized `quant_cfg` entries are now `QuantizerCfgEntry` objects
internally.
Mapping-style access is preserved, but direct equality checks against
literal dicts should use `entry.model_dump()`.

### Testing

- `python -m pytest tests/unit/recipe/test_loader.py`
- `python -m pytest
tests/unit/torch/quantization/test_config_validation.py`
- `ruff check` / `ruff format --check` on touched files
- Commit-time pre-commit hooks passed

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-05-18 08:40:33 -07:00
h-guo18 7038dec918 [1/2Refactor] speculative decoding: use mto config subsystem (#1328)
### What does this PR do?

Type of change: new feature

Port the speculative-decoding example to ModelOpt's recipe/config
subsystem: `model` / `data` / `training` / `<algo>` now load from a
single YAML with Pydantic validation and OmegaConf dotlist overrides.
Adds built-in `eagle3` / `dflash` recipes, drops the redundant
`training.mode` field (inferred from recipe class), and shrinks
`main.py` by ~145 lines (−208 / +63).

JIRA: OMNIML-3859

### Usage

```bash
python main.py --config general/speculative_decoding/eagle3 \
    model.model_name_or_path=meta-llama/Llama-3.2-1B \
    data.data_path=train.jsonl \
    training.output_dir=ckpts/test
```

### Testing

- `pytest tests/unit/recipe/test_loader.py` — new coverage for Eagle /
DFlash YAML loading, dotlist overrides, and field-level validation.
- Smoke-trained both built-in `eagle3` and `dflash` recipes end-to-end.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ — `main.py` CLI switched to
`--config <recipe>` (+ dotlist overrides); the old argparse flags are
removed.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
deps (`pydantic`, `omegaconf` already in core).
- Did you write any new necessary tests?: ✅ —
`tests/unit/recipe/test_loader.py`.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — to be added.

### Additional Information

Follow-up to the `modelopt.recipe` subsystem introduced for PTQ; this PR
extends the same declarative-YAML pattern to speculative decoding
(Eagle3 / DFlash / Medusa).


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added typed speculative-decoding recipe support for EAGLE, DFlash, and
Medusa; CLI dotlist overrides supported for single-file recipes.
* Trainer/config schema extended with speculative-training fields and
draft-vocab cache loading for Eagle.

* **Bug Fixes**
* Offline training no longer mutates model configs; loader enforces
required algorithm sections and prints recipe/config only on the primary
process.
* Reduced noisy per-rank logging by restricting status output to the
primary process.

* **Tests**
* Expanded tests for recipe loading, dotlist overrides, validation
strictness, and error cases.

* **Documentation**
  * Recipe YAMLs updated with metadata and usage notes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-05-16 17:29:24 -07:00
Frida HouandClaude Opus 4.7 81c4fb25ab Add 7 nvidia/Nemotron-* calibration datasets to SUPPORTED_DATASET_CONFIG (#1498)
Registers nemotron-{sft-instruction-following-chat-v2, science-v1,
competitive-programming-v1, sft-agentic-v2, math-v2, sft-swe-v2,
sft-multilingual-v1} so hf_ptq.py's --dataset flag (which enumerates
get_supported_datasets() automatically) can select these for PTQ
calibration. Splits with heterogeneous parquet schemas that crash
streaming CastError mid-iteration are excluded per inline comments.

Adds a parametrized smoke test that skips when HF_TOKEN is unset since
the nvidia/Nemotron-* datasets are gated.

### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Standardized chat-style preprocessing (joining message contents) and
normalized chat key to "messages" across seven Nemotron dataset
variants, with explicit dataset paths and curated splits.
* **Tests**
  * Added registry-shape unit checks for each new dataset entry.
* Added a gated integration smoke test that fetches sample data when a
Hugging Face Hub token is present.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1498?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-15 18:57:39 -07:00
Frida Hou 701934bf03 fixes for fused moe (qwen3.6, GLM5.1 + MSE calibration (#1382)
### What does this PR do?

Type of change: Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

Fixes several issues with NVFP4 MSE calibration and export for fused MoE
expert modules (_QuantFusedExperts — used by Qwen3.6, GLM-5.1, and other
HF transformers 5.0+ models that store expert weights as 3-D
nn.Parameters).

- Bug 1 — MSE weight calibration runs 0 iterations for fused experts
(model_calib.py)

The weight-quantizer discovery loop in mse_calibrate used the singular
attribute name gate_up_proj_weight_quantizer to look up quantizers, but
_QuantFusedExperts stores them in a plural nn.ModuleList named
gate_up_proj_weight_quantizers. All 20,480 expert quantizers were
silently skipped, resulting in "MSE weight calibration: 0it" and no
MSE-optimized scales.

Fix: add a second pass that detects plural {param}_weight_quantizers
ModuleLists and enqueues each per-expert quantizer with a (param_name,
expert_idx) tuple; step 3 unpacks the tuple to extract the per-expert
weight slice.

  - Bug 2 — Zero weight scales in exported checkpoint (nvfp4_tensor.py)

Per-block weight scales can silently underflow to 0 when cast to FP8
E4M3FN. The existing scale == 0 guard only catches exact float32 zeros;
values in (0, 2^-9) pass through and become 0 after the FP8 cast. This
affects both the dynamic recompute path (get_weights_scaling_factor) and
the static calibrated path (get_weights_scaling_factor_from_quantizer).
                                      
Fix: clamp per-block scales to 2^-9 (smallest positive FP8 E4M3FN
subnormal) before the FP8 cast in both paths.
   
- Bug 3 — Zero/corrupt amax for uncalibrated experts at export
(moe_utils.py)
                                      
Experts that receive no tokens during calibration have _amax = 0 or
uninitialized values. The existing scalar fallback used 1e-4 which
itself underflows to 0 in FP8 E4M3FN (1e-4 < 2^-9 ≈ 0.00195).
Additionally, the per-block fallback tensor had shape (H*W, 1) instead
of (H, W), causing a shape mismatch that silently bypassed the fallback
and fell through to the bad scalar. Finally, a stale zero global_amax
from an uncalibrated expert was not recomputed, causing division-by-zero
in the FP8 scale formula.

Fix: reshape the per-block fallback correctly; raise the clamp floor to
2e-3; always recompute global_amax from the current (possibly patched)
per-block _amax.
  Additional fixes:                   
- moe_utils.py: safe CPU extraction of _amax before deepcopy to avoid
async CUDA errors from corrupt bfloat16 amax storage on under-calibrated
experts.
- model_quant.py: print_quant_summary now calls os.makedirs(output_dir,
exist_ok=True) before writing .quant_summary.txt, preventing a
FileNotFoundError when the export directory doesn't exist yet.
- tensor_quantizer.py: change default format in _short_amax /
_short_tensor from ".4f" to ".2e" so small amax values (e.g. 3.5e-7)
display as 3.50e-07 instead of 0.0000.
- hf_ptq.py: strip leading pad tokens from the preview input and add
skip_special_tokens=True to input_decode, fixing degenerate pre/post-PTQ
output on models that use EOS as the pad token (e.g. Qwen3).
### Usage

```python
 # Quantize Qwen3.6-35B-A3B (or any compatible fused-expert MoE) with the new recipe:
  python examples/llm_ptq/hf_ptq.py \                                                                                                     
      --pyt_ckpt_path /path/to/Qwen3.6-35B-A3B \                                                                                          
      --recipe modelopt_recipes/general/ptq/nvfp4_experts_only_mse.yaml \                                                                 
      --export_path /path/to/output \                                                                                                     
      --calib_size 512 --calib_seq 2048   
```

### Testing
validated on Qwen3.6-35B-A3B (8× B200):
- 21,740 quantizers inserted; 20,480/20,480 MSE weight calibrations
completed (~11 min)
- 0 / 2,013,265,920 zero weight_scale entries in the exported checkpoint
(3 shards)
- Pre- and post-PTQ generation produce coherent, semantically consistent
output

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a new NVFP4 quantization recipe for expert layers with MSE-based
calibration.

* **Bug Fixes**
  * Fixed FP8 scale underflow handling to prevent zero scaling factors.
  * Fixed output directory creation for quantization summaries.

* **Improvements**
* Enhanced preview input handling for language models by removing
padding tokens.
  * Improved quantizer display precision for better readability.

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1382)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
2026-05-15 15:10:25 -07:00
a451a2baf8 Support NVFP4 W4A16 quantization (#1313)
### What does this PR do?

This PR supports NVFP4 W4A16 quantization.

### Usage

```python
  python examples/llm_ptq/scripts/huggingface_example.sh \
      --model Qwen/Qwen3-8B \
      --quant w4a16_nvfp4 \
      --calib 1 \
      --kv_cache_quant none \
      --tasks quant
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅ 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

### Additional Information

NVFP4 W4A16 is currently supported on vLLM only. TensorRT-LLM and SGLang
do not support this format.
   
vLLM loads NVFP4 W4A16 checkpoints via its `CompressedTensorsW4A16Fp4`
scheme, which requires the checkpoint to be in `compressed-tensors`
format. The raw ModelOpt export from this PR is still not directly
loadable by vLLM. After running hf_ptq.py, the checkpoint must be
converted by:
1. Renames tensors: .weight (uint8) → .weight_packed, .weight_scale_2
(fp32) → .weight_global_scale (inverted: 1 / weight_scale_2)
2. Rewrites config.json: sets quant_method: "compressed-tensors",
format: "nvfp4-pack-quantized", quantization_status: "compressed", and
derives the ignore list
  automatically from layers that lack a weight_packed tensor.  

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* NVFP4 W4A16 weight-only quantization added (FP4 weights,
group_size=16; BF16 activations; no calibration-forward pass).
Selectable via qformat/CLI and included in export flow.
* New --exclude_modules CLI option (and EXCLUDE_MODULES env support in
example script) to skip quantizing specific modules.

* **Documentation**
* Changelog entry describing NVFP4 W4A16 usage and vLLM deployment via
compressed-tensors conversion.

* **Tests**
  * Added test coverage for NVFP4 W4A16 export.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Co-authored-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-05-15 10:47:26 -07:00
Ajinkya Rasane 229ba61877 [OMNIML-3050] Enable torch.compile on _get_log_softmax_dist (#1479)
### What does this PR do?

Type of change: new feature

Wraps `_get_log_softmax_dist`
(`modelopt/torch/quantization/algorithms.py`) — the distributed
log-softmax helper used by `AutoQuantizeKLDivSearcher` under TP > 1 —
with `@torch.compile(dynamic=True)`. Inductor fuses the `amax /
all_reduce(MAX) / logsumexp / all_reduce(SUM) / log+sub+cast` pipeline
into one kernel. `dynamic=True` avoids recompiles across the varying
`[batch, seq]` shapes seen during calibration, matching the existing
pattern in `backends/fp8_per_tensor_gemm.py`. Stale TODOs are removed;
the prior ONNX-Windows concern no longer applies because the function is
only reachable when a TP group is initialized (never on the Windows CPU
unit-test job), and `@torch.compile` is import-time safe.

### Usage

```python
# Internal — invoked automatically by AutoQuantize KL-divergence search under TP > 1:
import modelopt.torch.quantization as mtq

model, _ = mtq.auto_quantize(
    model,
    constraints={"effective_bits": 6.0},
    quantization_formats=[mtq.INT4_AWQ_CFG, mtq.INT8_DEFAULT_CFG],
    data_loader=calib_loader,
    forward_step=lambda m, b: m(b),
    method="kl_div",
)
```

### Testing

Verified locally against the affected unit-test paths:

- `pytest tests/unit/torch/quantization/test_autoquant.py -k kl_div` →
22 passed (covers the call chain into `_get_log_prob`).
- `pytest tests/unit/onnx/` → 516 passed. The single failure
(`test_autocast_quantize.py::test_autocast_quantize_int8[False-True]` —
onnxruntime `CopyTensorAsync is not implemented`) is a pre-existing
failure on `main`, unrelated to this change (confirmed by stashing the
diff and re-running).
- `pytest tests/unit/torch/quantization/test_autoquant.py
tests/unit/torch/quantization/test_quantize_cpu.py
tests/unit/torch/quantization/test_config_validation.py` → 159 passed.
- Functional sanity in a single-process gloo group: compiled output
matches `torch.log_softmax` reference for fp32 and fp16, with no
recompiles across varying `[batch, seq]` shapes.
- `pre-commit run --files modelopt/torch/quantization/algorithms.py` →
ruff / mypy / bandit all pass.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (\`git commit -s -S\`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded \`trust_remote_code=True\`, \`torch.load(...,
weights_only=False)\`, \`pickle\`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in \`CONTRIBUTING.md\`: N/A
- Did you write any new necessary tests?: N/A — existing \`kl_div\`
autoquant tests already exercise the call chain.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ — will run \`/claude
review\` after marking ready.

### Additional Information

Single-line behavior change: \`_get_log_softmax_dist\` is now
\`@torch.compile(dynamic=True)\`-decorated. No API surface changes.

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2026-05-13 17:04:20 -04:00
Keval Morabia 94337ade75 fix(te-plugin): handle TE 2.15+ tuple return from _Linear / _GroupedLinear (#1481)
### What does this PR do?

Type of change: Bug fix

TE 2.15+ changed `_Linear.forward` and `_GroupedLinear.forward` to
return `(out, new_workspace)` tuples instead of a single tensor.
ModelOpt's patched `te_quantized_linear_fn` /
`te_grouped_quantized_linear_fn` still piped the whole tuple into
`self.output_quantizer`, crashing inside `TensorQuantizer.forward` on
`tuple.numel()`:

```
File ".../modelopt/torch/quantization/plugins/transformer_engine.py", line 184, in te_grouped_quantized_linear_fn
    return self.output_quantizer(output)
File ".../tensor_quantizer.py", line 1037, in forward
    if inputs.numel() == 0:
AttributeError: 'tuple' object has no attribute 'numel'
```

Mirror the existing pattern from `_QuantTELayerNormLinear.forward`: when
the underlying TE call returns a tuple, quantize only `output[0]` (the
activation tensor) and pass auxiliary workspace metadata through
unchanged. TE <= 2.14 returns a single tensor and falls through the
`isinstance` branch identically to before this change.

Already landed on `release/0.44.0` as commit `c897fbeaaf`; this brings
`main` in sync. Follow-up to
[#1473](https://github.com/NVIDIA/Model-Optimizer/pull/1473) (signature
introspection + `_forward` cache lookup), which fixed an earlier symptom
of the same TE 2.15 signature change but not this tuple-return path.

### Usage

No public API change. PTQ continues to work transparently across TE 2.x:

```python
import modelopt.torch.quantization as mtq
mtq.quantize(model, mtq.NVFP4_DEFAULT_CFG, forward_loop)  # now works on TE 2.15.x
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

Verified locally against **both TE 2.12** and **TE 2.15.0** using:

```bash
pytest tests/gpu_megatron/torch/quantization/plugins/test_transformer_engine.py
```

Without this fix on TE 2.15, the same test fails immediately with
`AttributeError: 'tuple' object has no attribute 'numel'`. With this
fix, both versions exercise the same code paths and pass — TE <= 2.14
skips the `isinstance(output, tuple)` branch and behaves identically to
before.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- Public API unchanged; TE
<= 2.14 path is identical (isinstance branch is false). -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!--- Existing
`test_transformer_engine.py` already exercises both paths; it would have
caught this on TE 2.15 had CI been running against that version. A
TE-version matrix is the right follow-up but is out of scope here. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information
<!-- E.g. related issue. -->
Triggered by Megatron-Bridge failing tests after their TE 2.15 bump. The
`release/0.44.0` cherry-pick was pushed directly (commit `c897fbeaaf`)
so Bridge could unblock; this PR carries the same fix forward to main.

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-14 02:04:01 +05:30
realAsma 62401e16df fix: layerwise calibration backward-compat, recipe split, batch-size guard (#1310)
## Summary

Follow-up to #1251 (which renamed `use_sequential` → `layerwise`). Three
related fixes bundled:

1. **Backward-compatible config loading.** PTQ checkpoints saved before
#1251 store the
legacy `use_sequential` key in the calibration-algorithm config, so
loading them now
raises `ValidationError: Extra inputs are not permitted
(use_sequential)` because
`QuantizeAlgorithmConfig` uses `extra='forbid'`. Accept `use_sequential`
as an alias
for `layerwise` via `AliasChoices`. The field still serializes as
`layerwise`, so
   round-trips through the current schema are clean.

2. **Recipe split.** `nvfp4_experts_only-fp8_kv` previously enabled
layerwise calibration
by default, which changes the calibration flow materially. Split into
two recipes:
   - `nvfp4_experts_only-fp8_kv.yaml` — default (no layerwise)
   - `nvfp4_experts_only-fp8_kv_layerwise.yaml` — layerwise variant

3. **`hf_ptq` batch-size guard.** Auto batch-size detection is not
supported together
with layerwise calibration. Default to `batch_size=1` when layerwise is
enabled and
   the user hasn't set a batch size explicitly.

Originally reported by Jenny Chen while resuming a PTQ checkpoint via
`restore_sharded_modelopt_state`:

```
pydantic_core._pydantic_core.ValidationError: 1 validation error for MaxCalibConfig
use_sequential
  Extra inputs are not permitted [type=extra_forbidden, input_value=False, input_type=bool]
```

## Test plan

- [x] `tests/unit/torch/quantization/test_config_validation.py` — legacy
alias accepted, current name accepted, dump serializes under current
name, `extra='forbid'` still rejects unknown keys.
- [x] `pre-commit run` — clean.

### Before your PR is *Ready for review*

- Is this change backward compatible?: ✅ (restores compatibility for
pre-#1251 checkpoints)
- New PIP dependency: N/A
- New necessary tests: ✅
- Changelog update: N/A (bug fix)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added new PTQ recipe for efficient layerwise calibration of large
models.
  * Automatic batch size optimization for layerwise calibration recipes.
  * Backward compatibility support for legacy input naming conventions.

* **Documentation**
* Updated recipe guides and changelog with new layerwise calibration
recipe.

* **Tests**
  * Added validation tests for configuration compatibility.

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1310)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-05-12 20:17:21 +00:00
Keval Morabia 5da4d636fc fix(te-plugin): make _Linear arg indexing robust to TE signature changes (#1473)
### What does this PR do?

Type of change: Bug fix

ModelOpt's `te_quantized_linear_fn` and `te_grouped_quantized_linear_fn`
read `weight` / `inp` from hard-coded positions in `args`. Two TE
signature changes broke this scheme:

- **TE 1.x → 2.0:** dropped the legacy `weight_fp8` slot between
`weight` and `inp`. ModelOpt handled this with an `if Version("2.0") <=
_TE_VERSION:` branch + a duplicate else branch.
- **TE 2.14 → 2.15:** inserted `weight_workspace` between `weight` and
`inp` at the `_Linear.forward` call site ([TE 2.15 linear.py
L1663](https://github.com/NVIDIA/TransformerEngine/blob/release_v2.15/transformer_engine/pytorch/module/linear.py#L1663)).
Unhandled by ModelOpt — `args[idx + 1]` resolved to `None` (workspace is
None outside FP8), which then crashed `TensorQuantizer.forward` on
`inputs.numel()` with `AttributeError: 'NoneType' object has no
attribute 'numel'`. Surfaced as a regression in Megatron-Bridge after
the TE 2.15 bump alongside ModelOpt 0.44.0rc3.
- **TE 2.10:** `_GroupedLinear.forward`'s second positional slot was
renamed `m_splits` → `non_tensor_args` (tuple wrapping). ModelOpt had a
separate `Version("2.10")` gate for this.

Replace all three version gates with **parameter-name introspection** of
the live `_Linear.forward` / `_GroupedLinear.forward` signature. The
parameter names (`weight`, `inp`, `m_splits`, `non_tensor_args`) have
been stable across TE 1.x, 2.x, and 2.15+; only their relative positions
shift. The new code reads the live signature via
`inspect.signature(...).parameters`, locates `weight`/`inp` by name, and
mutates only those positions in a list copy of `args` — everything
between (e.g. TE 2.15's `weight_workspace`) and after passes through
verbatim. The dual-branch code in `te_quantized_linear_fn` collapses to
a single path.

### Usage

No public API change. PTQ continues to work transparently across all
supported TE versions:

```python
import modelopt.torch.quantization as mtq
# Works on TE 1.x, 2.0-2.14, 2.15.x, and 2.16+ — no version flag needed.
mtq.quantize(model, mtq.NVFP4_DEFAULT_CFG, forward_loop)
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

Existing TE plugin tests
(`tests/gpu_megatron/torch/quantization/plugins/test_transformer_engine.py`)
exercise both the `_forward` (no-grad calibration) and `_apply`
(grad-enabled training) paths of `te_quantized_linear_fn` for
`te.pytorch.Linear` — they would have caught the original TE 2.15
regression on a CI matrix entry pinned to TE 2.15. Verified trace
correctness across:

| TE version | `_Linear.forward` signature | `_te_linear` weight→inp gap
| `_GroupedLinear.forward` second slot |
|---|---|---|---|
| 1.x | `(ctx, weight, weight_fp8, inp, …)` | 1 | n/a |
| 2.0–2.14 | `(ctx, weight, inp, bias, …)` | 0 | `m_splits` |
| 2.15.x | `(ctx, weight, weight_workspace, inp, …)` | 1 |
`non_tensor_args` |
| 2.16+ (main) | `(ctx, weight, inp, bias, fwd_args)` | 0 |
`non_tensor_args` |

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- Public API unchanged;
broadens the range of TE versions that work (TE 2.15.x now supported, TE
1.x still supported via the same introspection path). -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!--- Only
adds a stdlib `inspect` import. -->
- Did you write any new necessary tests?: Existing tests sufficient
<!--- Bug fix is covered by existing `test_transformer_engine.py` for
whatever single TE version CI exercises. A multi-version TE matrix is
the right next step but is out of scope for this PR. -->

### Additional Information
<!-- E.g. related issue. -->
Triggered by Megatron-Bridge
https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/3783 failing tests
after bumping ModelOpt 0.44.0rc2 → 0.44.0rc3 together with a Megatron-LM
bump that pulls TE 2.15. ModelOpt rc2 had the same latent bug — it just
wasn't exercised until TE 2.15 became the runtime version.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Improved Transformer Engine quantization plugin robustness by using
runtime parameter inspection instead of version-based branching,
ensuring compatibility across TE versions without requiring manual
updates.

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1473)

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-13 01:40:56 +05:30
d738995fe8 feat(opt): validate loaded modelopt state files (#1471)
Add validation to load_modelopt_state() to verify the loaded object is a
dict with the expected schema (modelopt_state_dict list and
modelopt_version str). Raises TypeError/ValueError with clear messages
when the file is malformed, and detects full checkpoints passed by
mistake, pointing users to mto.restore().

Closes #1041


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Added strict validation for model state files to surface format errors
with clear messages.
* Malformed or invalid state files now fail fast instead of being
returned silently.
* Improved detection to prevent accidental loading of full checkpoints
when only state dicts are expected.

* **Tests**
* New unit tests covering validation and loading behavior for various
malformed and valid state files.

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1471)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
Co-authored-by: Keval Morabia <kevalmorabia97@users.noreply.github.com>
2026-05-12 19:27:10 +00:00
Chenjie Luo 5508c327fb [Quantization] MSE-calibrate every per-expert weight in fused-experts MoE (#1421)
### What does this PR do?

Type of change: Bug fix

Two-part fix for transformers 5.x fused-experts containers (Qwen3-MoE /
Qwen3.5-MoE / Mixtral / DeepSeek / Kimi-K2.x ...) where weight
quantizers live in `nn.ModuleList`s (`gate_up_proj_weight_quantizers`,
`down_proj_weight_quantizers`):

1. **Per-expert weight iteration for calibration.** Add
`_QuantFusedExperts.iter_weights_for_calibration` that yields per-expert
`(weight_slice, quantizer)` pairs for both projections. The base impl
uses singular `*_weight_quantizer` and silently skips fused-experts
modules, so weight-only calibration paths never reached per-expert
quantizers.

2. **`mse_calibrate` refactor.**
- Add `_bootstrap_uncalibrated_weight_quantizers` after `max_calibrate`
to populate `_amax` on quantizers the forward pass didn't reach (dead
MoE experts that received no calibration tokens). Runs the existing
calibrator on the weight slice surfaced by
`iter_weights_for_calibration`.
- Replace the singular-only `weight_attr_names` discovery +
`getattr`-by-name walk with an `iter_weights_for_calibration` walk done
inside each parent module's `enable_weight_access_and_writeback`
context, so MSE processes every per-expert quantizer (active and dead)
and remains FSDP-safe.

Without this, the export-time fallback in `_export_fused_experts`
derived separate gate/up amaxes from each half of the fused weight,
breaking the `gate==up` `weight_scale_2` invariant on dead experts.

Also includes:
- `_sanitize_generation_config_for_save` in `unified_export_hf` —
coerces `do_sample=True` when an upstream `generation_config.json` has
`top_k`/`top_p` set, so newer transformers' strict validate doesn't
block `save_pretrained`.
- Small companion plumbing in `moe_utils.py`, `tensor_quantizer.py`, and
`core_utils.py` to support the per-expert iteration and bootstrap path.

### Usage

```python
import modelopt.torch.quantization as mtq
from modelopt.recipe import load_config

# Recipe `nvfp4_experts_only_mse-kv_fp8_cast` (already on main) now correctly
# MSE-calibrates every per-expert weight quantizer in fused-experts MoE models.
cfg = load_config("general/ptq/nvfp4_experts_only_mse-kv_fp8_cast")
mtq.quantize(model, cfg, forward_loop=calibration_forward_loop)
```

### Testing

**Original validation — Qwen3.5-122B-A10B with
`nvfp4_experts_only_mse-fp8_cast_kv`:**
- **Before:** 1/12288 (layer 38 expert 69) `gate \!= up`; 0 weights
MSE-calibrated.
- **After:** 0/12288 mismatches; 24576 weights MSE-calibrated; ~4.2 min.

**End-to-end pipeline validation — Qwen3.5-35B-A3B (40 layers × 256
experts × 2 projections = 20,480 per-expert weight quantizers), TRT-LLM
1.3.0rc13 + transformers 5.6 docker, single B200:**

| | Path A (4-sample calib, deliberately undercalibrated) | Path B (zero
forward-pass tokens) |
|---|---|---|
| Per-expert weight quantizers calibrated | 20,480 / 20,480 | 20,480 /
20,480 |
| Missing `_amax` | 0 | 0 |
| All-zero `_amax` | 0 | 0 |
| `mtq.quantize` time | 25–34 s | 23 s |

- **Cross-path diff:** every per-expert weight amax matches
**bit-for-bit** between the two paths (`n=20480 exact=20480 diff=0
max_rel=0`). With 8/256 experts routed per token and 4 calib samples,
almost all experts are "dead" in Path A. Bootstrap fills them from
`max(|weight|)`, MSE searches deterministically from there → identical
to Path B which bootstraps everything.
- **Export to HF NVFP4 checkpoint** succeeded (~95 s, 22 GB checkpoint).
Resulting `generation_config.json` has `do_sample: true` (upstream had
`top_k=20` + `top_p=0.95` which would have failed strict validate).
- **TRT-LLM inference loaded the checkpoint and generated text:** `"Born
in north-east France, Soyer trained as a"` → `" tailor. Demonstrating
his craft at a young age, at 20 he moved to Paris at the requests of the
noble people of Picardy."` (coherent grammar; factually wrong as
expected with 4-sample calib, but no NaN/Inf in logits, no
scale-mismatch crash). 92 GB GPU memory used.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ <\!-- relies on existing
recipe-level integration coverage; verified end-to-end on
Qwen3.5-122B-A10B and Qwen3.5-35B-A3B + TRT-LLM 1.3.0rc13 -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ <\!-- will run \`/claude
review\` -->

### Additional Information

Follow-up to PR #1407 (MSE+FP8-cast-KV recipes). The recipe YAML files
landed there; this PR fixes the calibration codepath so the MSE recipes
actually exercise per-expert weight quantizers in fused-experts MoE
containers.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed generation configuration validation for HuggingFace model
exports.
* Improved handling of quantization shape mismatches during expert
weight export.

* **New Features**
* Enhanced calibration process with automatic population of missing
expert quantizers.
* Added grouped quantizer synchronization for improved multi-expert
quantization.

* **Tests**
* Added regression tests for fused expert export and calibration
correctness.

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1421)

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-05-12 12:07:34 -07:00
Keval Morabiaandclaude[bot] 2ce745a92e Deprecate gradnas pruning and bert example (#1427)
### What does this PR do?

Type of change: Deprecation of dead code <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->

Deprecation warning already added in 0.44 as per 1-release deprecation
policy

GradNAS only works for Bert and GPT-J and we dont actively maintain it
or test it. Keeping it creates an expectation that it works plus it adds
one more option for user to choose from. We already have much better
pruning algorithms (Minitron and Puzzletron) for LLM pruning already
hence removing GradNas.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ No but we dont have any users
of this feature either <!--- If ❌, explain why. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Deprecation**
  * GradNAS pruning algorithm deprecated; related examples removed.
* **Documentation**
* Pruning and NAS guides and changelog updated to focus on Minitron and
FastNAS; GradNAS references removed.
* **Chores**
  * Chained-optimizations example and scripts removed.
* Ownership mappings updated for README and examples; license insertion
now applies to a previously excluded example file.
* **Tests**
* Multiple unit tests and test utilities related to GradNAS/transformer
NAS removed.

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1427)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
2026-05-12 22:32:22 +05:30