mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
main
719
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c78e654744 |
Skip Softmax diffusion export (#1269)
### What does this PR do?
Type of change: New Feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
Adds HuggingFace `config.json` export of skip-softmax sparse-attention
calibration for diffusion pipelines (e.g. Wan 2.2), on top of the base
skip-softmax work.
- **`_export_diffusers_checkpoint`** walks every `nn.Module` component
of a diffusers pipeline, calls `export_sparse_attention_config`, and
writes the result into that component's `config.json` under the
`sparse_attention_config` key. The sparse config lives **only** in
`config.json` — there is no standalone `sparse.yaml`.
- **`export_sparse_attention_config`** emits a `config_groups` schema
where each algorithm's parameters are nested inside its own group; only
`config_groups` and `producer` are top-level:
- skip-softmax group → `algorithm: "skip_softmax"`, `targets`, `ignore`
(layers kept dense — e.g. cross-attention + first/last blocks),
`initial_disabled_steps` (opt-in, user-set; emitted only when `> 0`),
`threshold_scale_factor` (`a * exp(b * target_sparsity)`), and
`target_sparsity`.
- N:M group → `algorithm: "sparse_softmax"` with
`sparsity_n`/`sparsity_m`, `dense_sink_tokens`, `dense_recent_tokens`
flattened into the group.
- **Deploy reader**
(`modelopt/torch/sparsity/attention_sparsity/plugins/sparse_attn_config.py`)
reads these per-group params back, keeping the export↔load round-trip
consistent.
- **Example wiring**:
`examples/diffusers/sparsity/wan22_skip_softmax.py` gains
`--export-dir`, `--skip-softmax-threshold`, and
`--initial-disabled-steps`. `--export-dir` runs
`export_hf_checkpoint(pipe, export_dir=...)` after calibration.
- Updated `CHANGELOG.rst`.
### Usage
```bash
python examples/diffusers/sparsity/wan22_skip_softmax.py \
--model-path Wan-AI/Wan2.2-T2V-A14B-Diffusers \
--calibrate --target-sparsity 0.5 --calib-size 4 \
--initial-disabled-steps 5 \
--export-dir ./wan22_skip_softmax_ckpt
```
Resulting layout — a `config.json` per component, **no `sparse.yaml`**:
```
wan22_skip_softmax_ckpt/
├── transformer/config.json # carries sparse_attention_config
├── transformer_2/config.json # carries sparse_attention_config
├── vae/ … text_encoder/ … tokenizer/ … scheduler/ …
└── model_index.json
```
A representative `config.json` entry for a diffusion transformer:
```json
"sparse_attention_config": {
"config_groups": {
"group_0": {
"algorithm": "skip_softmax",
"targets": ["WanAttention"],
"ignore": ["blocks.0.attn1", "blocks.0.attn2", "…"],
"initial_disabled_steps": 5,
"threshold_scale_factor": {
"formula": "a * exp(b * target_sparsity)",
"prefill": {"a": 1443.49, "b": 4.30}
},
"target_sparsity": {"prefill": 0.5}
}
},
"producer": {"name": "modelopt", "version": "0.45.0..."}
}
```
The N:M variant adds a second group:
```json
"group_1": {
"algorithm": "sparse_softmax",
"targets": ["WanAttention"],
"sparsity_n": 2, "sparsity_m": 4,
"dense_sink_tokens": 0, "dense_recent_tokens": 64
}
```
### Testing
- `tests/examples/diffusers_sparsity/test_sparsity.py`: baseline /
triton-baseline / fixed-threshold runs of the Wan 2.2 example, plus a
Python-API calibrate → **export** test asserting the nested
`sparse_attention_config` (`threshold_scale_factor`, `target_sparsity`,
`ignore`, `initial_disabled_steps`) and the absence of any
`sparse.yaml`.
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attention_conversion.py`
and `test_sparse_attn_config.py`: unit coverage of the per-group export
schema and the deploy-reader round-trip (writer nests → reader reads
from groups → internal mtsa config unchanged).
- Validated end-to-end on Wan 2.2 T2V-A14B: full 4-prompt / 40-step /
81-frame calibration; the exported checkpoint carries the nested schema
in both `transformer` and `transformer_2` `config.json`, and runtime
measurement shows ~47–49% tile sparsity at a 0.5 target.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ❌ The exported
`sparse_attention_config` schema was renamed and nested per-group during
0.45.x development, and the loader reads only the new layout —
checkpoints exported by earlier 0.45.x builds must be re-exported. No
released version is affected. <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
16d562a0bf |
Refactor local_hessian onto shared MSE flow + fused-MoE expert support (#1578)
### What does this PR do? Type of change: Bug fix + new feature (fused-MoE coverage) <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Dusted off and refactored `local_hessian_calibrate`** to align with the new MSE calibration flow, fix latent drift, extend coverage to fused-MoE experts, and decouple module-specific handling behind a clean extension point. Core changes (`modelopt/torch/quantization/model_calib.py`): - Extracted a shared `_mse_calibrate_weights` helper now used by both `mse_calibrate` and `local_hessian_calibrate` - Replaced the monolithic `LocalHessianHelper` + bespoke per-weight loop with a small `_LocalHessianAccumulator` (lazy fp32 buffer, freed after building the error func), removing ~200 lines of duplicated scale-search logic and all manual `cuda.synchronize`/`empty_cache` bookkeeping. The `XᵀX` GEMM accumulates in fp32 to avoid bf16/fp16 precision loss. - Removed dead NVFP4-static promotion (now handled inside `max_calibrate`). - **Fused-MoE expert support:** per-expert Hessians captured from each expert's routed activations, keyed by `id(weight_quantizer)` so dense and per-expert paths share one calibration loop. Never-routed experts / non-eager kernels / `cin` not divisible by `block_size` / registered backends fall back to plain MSE (with an eager, module-named warning). Decoupling (zero module-type-specific code in `model_calib.py`): - Added `QuantModule.register_calibration_input_hooks(callback)` — the activation-side counterpart to `iter_weights_for_calibration`. Base default is a no-op; `QuantLinearConvBase` pairs the weight quantizer with the forward input (linear only), and `_QuantFusedExperts` (in `plugins/huggingface.py`) owns the per-expert pairing via `_current_expert_idx`. Any future module type gains local-Hessian support by implementing this one method. ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing - new `tests/unit/torch/quantization/test_local_hessian.py` (accumulator math/shape/dtype, dense end-to-end, backend-skip, block-size guard) and a per-expert MoE test in `test_fused_experts.py`. - Behavior-preserving check: refactored branch produces **bit-identical** dense weight scales to `main` (216/216 tensors) on Qwen3-8B with the fp32 accumulation neutralized to isolate the structural change. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Introduced Hessian-weighted calibration method for enhanced weight quantization refinement in mixed-precision models. * Enhanced MSE calibration to support custom per-quantizer error functions for specialized calibration workflows. * **Tests** * Added comprehensive unit tests for local Hessian calibration validation across dense models and fused-expert architectures. * Included tests for custom error function integration and fallback behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> |
||
|
|
b98a59557a |
Add vLLM-based runtime statistics for subblock latency measurement (#1358)
### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Runtime-based latency optimization: collect vLLM-measured inference latency to constrain optimization. * **Configuration** * New runtime config/template for Llama-3.1-8B pruning (runtime stats enabled, NCCL timeout templating, MIP target-latency). * Validation sample defaults adjusted (one flow: 128 → 8; runtime flow uses 128). * Human constraint key renamed to target_latency_seconds. * **Documentation** * README section describing runtime-based latency optimization setup and usage. * **Tests** * Added GPU end-to-end test for runtime stats collection. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1358?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Grzegorz Karch <gkarch@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
01415c2788 |
fix(llm_eval): repair test_qwen3_eval_fp8 end-to-end (#1650)
### What does this PR do? Type of change: Bug fix `tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8` was silently passing while its evals crashed, then began failing as a timeout. This repairs the whole pipeline: - **lm_eval `IndexError` (root cause):** TRT-LLM KV-cache prefix reuse returns truncated `context_logits` for shared-prefix requests (e.g. hellaswag's one-context / many-endings), which breaks `parse_logprobs`. Add an `enable_kv_cache_reuse` flag to `modelopt.deploy.llm.LLM` (default `True`, unchanged) and disable it for the eval deployment so full-length context logits are returned. - **Silent CI green:** `python eval.py | tee result.txt` returns `tee`'s exit code, so a crashing eval was masked. Add `set -o pipefail` to `huggingface_example.sh` so failures fail the test. - **Long-prompt overflows:** with the tiny test model's toy tokenizer, gsm8k/MMLU prompts exceed `max_seq_len`. Bump test `max_position_embeddings` to 8192, skip MMLU prompts that don't fit even at zero-shot, and add an MMLU sample limit (`--mmlu_limit`). - **human-eval build failures:** install with `--no-build-isolation` (`pkg_resources` is absent in pip's isolated build env), patch its malformed `console_scripts` entry point, and pin the clone. - **Cleanups:** gate the post-quant `run_tensorrt_llm.py` smoke test behind the `quant` task (eval tasks deploy on their own; ~45s saved for eval-only runs); replace the SIGPIPE-prone serve-readiness `tail -f | while` with a poll loop (required under `pipefail`). ### Usage N/A — example/test fix. ### Testing All four eval tasks verified end-to-end in the CI container (TRT-LLM 1.3.0rc17, RTX 6000 Ada): lm_eval (hellaswag + gsm8k), MMLU, and simple_eval (humaneval) all complete with exit 0 and no `IndexError`/overflow. Cold full run ≈ 340s on this GPU. CI test on 2-gpu: https://github.com/NVIDIA/Model-Optimizer/actions/runs/27154417497/job/80153551154 ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (new `enable_kv_cache_reuse` defaults to current behavior; new script flags are optional) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new dependencies) - Did you write any new necessary tests?: N/A (fixes and strengthens an existing test) - Did you update Changelog?: N/A (bug fix to examples/tests) - Did you get Claude approval on this PR?: ❌ (pending) ### Additional Information The full test runs ~340s on an RTX 6000 Ada; CI runners are historically slower, while `@pytest.mark.timeout` is set to 600 — worth watching the first CI run and bumping if it's close. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an option to limit MMLU evaluation length. * **Bug Fixes** * Disabled KV-cache prefix reuse for evaluations needing per-token context logits to prevent truncated/incorrect logprobs. * Skip examples whose prompts remain too long; warn and report accuracy as NaN if all examples are skipped. * **Chores / Scripts** * Improved example scripts for reproducible installs, patched entry point handling, pipeline failure detection, conditional test invocation, polling-based log wait, and a new CLI flag for MMLU limits. * **Tests** * Increased timeout and prompt headroom; capped MMLU smoke tests for speed. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1f4a489b0f |
Adds AutoQuant support for VLM / Qwen3.5-Qwen3.6 style models (#1381)
### What does this PR do? Type of change: new feature, bug fix, new tests ### Details - Enables AutoQuant search over fused MoE expert containers by snapshotting/restoring their per-expert quantizers. - Adds Qwen3.5/3.6 linear-attention grouping rules so fused deployment layers keep compatible quant formats. - Supports `w4a16_nvfp4` as an AutoQuant search format. - Preserves disabled AutoQuant layer patterns in generated configs while allowing selected modules like `lm_head` to override default disables. - Keeps recipe-mode and AutoQuantize VLM paths on the outer CausalLM so Qwen3.5/3.6-MoE `lm_head` remains visible. - Skips `parent_class`-scoped quant config entries during AutoQuant bare quantizer matching, preventing class-scoped global entries from last-match overriding every selected module. - Adds temporary hardcoded Qwen/VLM AutoQuant disabled-layer patterns in `hf_ptq.py` with a TODO to refactor into the config system. ### Usage ```bash python examples/llm_ptq/hf_ptq.py \ --pyt_ckpt_path <model_path> \ --qformat fp8,w4a16_nvfp4 \ --auto_quantize_bits 5.0 \ --auto_quantize_cost_model active_moe \ --auto_quantize_checkpoint <autoquant_state.pt> \ --export_path <output_dir> ``` ### Testing - `/Users/weimingc/miniconda3/envs/modelopt/bin/python -m pytest tests/unit/torch/quantization/test_autoquant.py::test_get_auto_quantize_config_keeps_selected_lm_head_enabled tests/unit/torch/quantization/test_config_validation.py::TestMatchQuantizerCfg::test_parent_class_scoped_entries_are_ignored_for_bare_autoquant_lookup` - `/Users/weimingc/miniconda3/envs/modelopt/bin/python -m pytest tests/unit/torch/quantization/test_autoquant.py tests/unit/torch/quantization/test_config_validation.py -k "not data_parallel"` (`120 passed, 1 deselected`) - `/Users/weimingc/miniconda3/envs/modelopt/bin/python -m py_compile examples/llm_ptq/hf_ptq.py modelopt/torch/quantization/algorithms.py modelopt/torch/quantization/_auto_quantize_cost.py tests/unit/torch/quantization/test_autoquant.py tests/unit/torch/quantization/test_config_validation.py` - Full local affected-file pytest without `-k "not data_parallel"` only failed `test_data_parallel_auto_quantize` because this local sandbox cannot bind a free socket (`PermissionError: Operation not permitted`). - Ran Qwen3.6 35B AutoQuant e2e with `fp8,w4a16_nvfp4` and exported a checkpoint. - Verified exported checkpoint loads in vLLM nightly without local patches. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added w4a16_nvfp4 quantization format and optional cost-exclusion patterns for AutoQuantize. * **Improvements** * Safer multimodal/VLM handling and AutoQuantize now runs on the full outer model when applicable. * Better fused-MoE support, more accurate weight accounting, and refined attention-grouping for improved quantization choices. * Dynamic layer-disabling support for targeted disables. * **Tests** * New unit tests covering cost-model exclusions, fused-MoE accounting, and config selection. * **Documentation** * Updated cost-constraint example to show exclusion-pattern usage. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
54ce4e09d8 |
Add Quantization Aware Distillation (QAD) to Megatron-Bridge example (#1600)
### What does this PR do? Type of change: new example **Note:** This is **part 2 of 4** (builds on #1589): - **Part 1 (#1589):** Megatron-Bridge `quantize.py` + `export.py` support and tests. - **Part 2 (this PR):** extend `distill.py` for quantization-aware distillation (QAD) — load a quantized Megatron checkpoint as the student. - **Part 3:** https://github.com/NVIDIA/Model-Optimizer/pull/1601 - **Part 4:** repeat the NVFP4 + QAD experiments on a non-Nemotron model. Extends `examples/megatron_bridge/distill.py` to initialize the student from a **Megatron checkpoint** (a quantized checkpoint from `quantize.py`, or a pruned one) via `--student_megatron_path`, enabling **Quantization Aware Distillation (QAD)**: - `--student_hf_path` still builds the student architecture; `--student_megatron_path` supplies the (optionally quantized) weights. - For a quantized checkpoint, the ModelOpt quantize mode + base weights are restored onto the **plain student before the knowledge-distillation conversion** (`restore_sharded_modelopt_state` is a no-op once a model is already converted), so the distilled checkpoint stays exportable as a quantized model with `export.py`. **Upstream dependency / workaround:** `DistillationProvider.provide()` has no seam to transform the student before the KD conversion, so this patches `provide()` at the class level (via an `id()`-keyed registry, because the provider proxies instance-attribute assignment to its teacher once the teacher is set). A companion Megatron-Bridge PR adds a first-class `DistillationProvider.student_pre_conversion_hook`; from nemo:26.06 onwards the workaround should be removed and replaced with that hook (a removal note in `distill.py` documents exactly how). ### Usage ```bash # 1) PTQ -> quantized Megatron checkpoint (part 1) torchrun --nproc_per_node 2 quantize.py \ --hf_model_name_or_path Qwen/Qwen3-8B --quant_cfg fp8 --tp_size 2 \ --export_megatron_path /tmp/Qwen3-8B-FP8-megatron # 2) QAD: distill the quantized student from the unquantized teacher torchrun --nproc_per_node 8 distill.py \ --teacher_hf_path Qwen/Qwen3-8B \ --student_hf_path Qwen/Qwen3-8B \ --student_megatron_path /tmp/Qwen3-8B-FP8-megatron \ --data_paths 1.0 tokenized/data_text_document \ --train_iters 1000 --output_dir /output/qwen3_8b_qad # 3) export the distilled quantized checkpoint (part 1) torchrun --nproc_per_node 1 export.py \ --hf_model_name_or_path Qwen/Qwen3-8B \ --megatron_path /output/qwen3_8b_qad/checkpoints \ --export_unified_hf_path /tmp/qwen3_8b_qad_fp8_hf ``` ### Testing `tests/examples/megatron_bridge/test_qad.py` (validated on a 2-GPU NeMo `26.04` container): quantize a tiny Qwen3 at TP=2 → QAD distill from the quantized student → `export.py` to a unified HF checkpoint, asserting `hf_quant_config.json` is written (proves the quantize mode survived QAD). Includes a commented-out vLLM deployment check, validated locally (full flow passes; vLLM loads the export as `quantization=modelopt`). Existing normal/Puzzletron distillation tests still pass. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: N/A (new example feature; default behavior unchanged when `--student_megatron_path` is not set) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new dependencies) - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ✅ ### Additional Information Depends on a companion Megatron-Bridge PR adding `DistillationProvider.student_pre_conversion_hook` (the upstream replacement for the class-level `provide()` workaround). The Nemotron-3 tutorial NVFP4 + QAD experiments ship in part 3. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Quantization Aware Distillation (QAD) workflow to recover accuracy of quantized Megatron students and distill from quantized checkpoints. * CLI option to initialize a distillation student from a Megatron checkpoint and a structure-only load path for bridging. * **Documentation** * Expanded runnable quantize → QAD → export guidance and best-practice tips. * **Tests** * End-to-end test validating quantize → QAD → export artifacts. * **Chores / UX** * Clearer rank-aware messages, improved tokenizer padding handling, and more consistent export behavior (fixed export dtype). <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
dbdff11a7f |
Drive PTQ example qformat choices from preset YAMLs (hf_ptq, multinode_ptq, megatron_bridge) (#1525)
### What does this PR do?
Type of change: Refactor
Replace the hardcoded `QUANT_CFG_CHOICES` / `KV_QUANT_CFG_CHOICES` dicts
in the PTQ
example scripts with a small `_load_preset_cfg_choices()` helper that
discovers the
available qformat names by listing
`modelopt_recipes/configs/ptq/presets/{model,kv}/`
and **eagerly loads every preset YAML into a plain dict at import** via
the existing
`load_config(...,
schema_type=QuantizeConfig).model_dump(exclude_unset=True)` path.
The directory listing becomes the source of truth for the `--qformat` /
`--kv_cache_qformat` CLI vocabulary.
> Note: an earlier revision used a lazy, copy-on-access `Mapping`. That
was overkill
> for these example scripts — the previous `mtq.*_CFG` module constants
were
> themselves eagerly-loaded shared dicts, and every call site that
mutates a config
> already deepcopies first — so it is now a plain eager dict. A lazy
variant can be
> reintroduced later if import time ever matters.
**Scope.** Three scripts carried the same hardcoded tables and all three
are migrated:
- `examples/llm_ptq/hf_ptq.py` — `--qformat` / `--kv_cache_qformat`.
- `examples/llm_ptq/multinode_ptq.py` — `--qformat` /
`--kv_cache_qformat`.
- `examples/megatron_bridge/quantize.py` — `--quant_cfg` /
`--kv_cache_quant`
(still also accepts any full `mtq.config.choices` name).
All three scripts share the discovery helper (`load_quant_cfg_choices`),
the canonical
alias table (`QFORMAT_ALIASES`), the KV disable sentinel
(`KV_CACHE_NONE`), and the
ready-built `QUANT_CFG_CHOICES` / `KV_QUANT_CFG_CHOICES` mappings via
the new
**`modelopt.recipe.presets`** module — no copy lives in the example
scripts anymore. The
module is a standalone import, so `import modelopt.recipe` stays cheap;
only an explicit
`import modelopt.recipe.presets` triggers the eager preset load.
A small alias table preserves previously-supported short CLI names
(`int8_sq`,
`nvfp4_awq`, `fp8_pb_wo`, …, plus the Megatron-Bridge `fp8_blockwise`)
as deprecation
shims. It is documented as not-for-extension — new formats land as
preset YAMLs, and
longer term, configurations should be authored as full recipes
(`--recipe`). The alias
logic is fail-fast: an alias pointing at a missing preset raises
`ValueError` at import.
Also adds `presets/kv/fp8_cast.yaml` and `presets/kv/nvfp4_cast.yaml`,
composed from the
existing `kv_fp8_cast` / `kv_nvfp4_cast` unit fragments. This promotes
`fp8_cast` /
`nvfp4_cast` to first-class KV presets and lets us delete the runtime
`_set_kv_cache_constant_amax` helper and all its call sites —
`use_constant_amax` is now
authoritative in the YAML. The KV-calibration-skip decision is derived
from the config
(`_kv_cfg_uses_constant_amax`), not from hardcoded format names.
**⚠️ CLI surface expansion (owner sign-off requested).** Because the
directory listing
is now the CLI vocabulary, each script accepts **every** preset under
`presets/{model,kv}/`, not just its previously curated subset. For
`hf_ptq.py` this is
the same surface the prior table covered; for `multinode_ptq.py` and the
Megatron-Bridge
script it is broader (e.g. KV `fp8_affine` / `fp8_cast` / `nvfp4_cast` /
`nvfp4_rotate`
are now selectable). This is intended ("the directory is the policy"),
but please confirm
those two scripts are meant to expose all presets — if a given path has
not validated a
format, it should be gated explicitly.
### Usage
```bash
# Old short names still work via the alias shim
python examples/llm_ptq/hf_ptq.py --pyt_ckpt_path <model> --qformat int8_sq --kv_cache_qformat fp8_cast --export_path out/
# Canonical preset basenames work directly
python examples/llm_ptq/hf_ptq.py --pyt_ckpt_path <model> --qformat int8_smoothquant --kv_cache_qformat fp8_cast --export_path out/
# A newly-added preset YAML is valid on the CLI of all three scripts with no code change
python examples/llm_ptq/hf_ptq.py --pyt_ckpt_path <model> --qformat nvfp4_awq_full --export_path out/
```
### Testing
- New: `tests/unit/recipe/test_presets.py` smoke tests for
`modelopt.recipe.presets` —
every discovered model/KV preset loads into a `quant_cfg` dict, the
directory listing is
fully covered, deprecation aliases resolve to their canonical preset,
the KV `none`
sentinel does not collide with a preset, and a stale alias raises. These
guard the eager
import-time load (one bad preset would otherwise break `import
modelopt.recipe.presets`
and every PTQ example).
- Previously verified locally (uv `.venv` py3.13 + `dev-py310-modelopt`
conda):
all previously-supported `--qformat` / `--kv_cache_qformat` names
resolve to dicts
bit-equal to the corresponding `mtq.*_CFG` constants; `fp8_cast` /
`nvfp4_cast` carry
`use_constant_amax: true` while non-cast variants do not; argparse
accepts
`--kv_cache_qformat none` plus all variants; unknown qformats raise at
lookup / argparse.
- All pre-commit hooks pass (ruff, mypy, bandit, license, rst, yaml).
- Pre-merge manual checks recommended by review (run in an env with the
deps installed):
`python examples/llm_ptq/hf_ptq.py --help`, `… multinode_ptq.py --help`,
`… megatron_bridge/quantize.py --help`.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — all previously-valid CLI
values continue to work via the alias table; output configs are
bit-equivalent to the prior hardcoded path.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
deps.
- Did you write any new necessary tests?: ✅ —
`tests/examples/llm_ptq/test_example_utils.py` preset-discovery smoke
tests.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ☐
### Additional Information
Out of scope / follow-up: the `_AUTO_QUANTIZE_QFORMATS` table and
`_canonical_qformat`
helper in `hf_ptq.py` are intentionally left hardcoded — auto_quantize
is being
refactored/reimplemented and they are expected to be removed soon.
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
433b549cd8 |
[2/N] Simplify KDTrainer and enhance ModelOptHFTrainer (#1191)
## Summary This PR simplifies the HuggingFace knowledge distillation trainer and enhances the base `ModelOptHFTrainer` with Liger fused loss, per-parameter learning rates, and training utilities. ### Model-agnostic Liger kernel fused loss Adds custom Liger kernel integration in `ModelOptHFTrainer` that extends HuggingFace's built-in support in three ways: 1. **Model-agnostic**: Works with any causal LM that has an `lm_head`, unlike HF's Liger which only supports [a fixed set of model architectures](https://github.com/linkedin/Liger-Kernel/blob/main/src/liger_kernel/transformers/monkey_patch.py). 2. **DeepSpeed ZeRO-3 support**: HF's Liger integration only works with FSDP. ModelOpt adds distributed param gathering for DeepSpeed ZeRO-3 and DDP as well. 3. **KD loss support**: `KDTrainer` extends fused loss to knowledge distillation via `LigerFusedLinearJSD` for fused lm_head + Jensen-Shannon divergence. #### Liger kernel memory sweep (Qwen3-1.7B, 2×H100 FSDP2, NVFP4+FP8_KV) Max per-GPU batch size before OOM at each sequence length: **QAT (no teacher)** | Seq Length | 512 | 1024 | 2048 | 4096 | 8192 | 16384 | |------------|-----|------|------|------|------|-------| | **Liger** | 16 | 16 | 16 | 16 | 8 | 4 | | **No Liger** | 16 | 16 | 8 | 4 | 2 | OOM | **QAD (with teacher)** | Seq Length | 512 | 1024 | 2048 | 4096 | 8192 | 16384 | |------------|-----|------|------|------|------|-------| | **Liger** | 16 | 16 | 8 | 4 | 2 | 1 | | **No Liger** | 8 | 4 | 2 | 1 | OOM | OOM | Liger fused loss enables **2-4× larger batch sizes** at long context lengths by avoiding the materialization of the full logit tensor. ### ModelOptHFTrainer enhancements - `ModelOptTrainerArguments` with `--trainable_params`, `--frozen_params`, `--lr_config`, `--save_dtype`, and `--manual_gc` flags - Per-parameter learning rate support via YAML config (`lr_config`) - `_prepare_model` and `_update_config_json_dtype` promoted to base class ### KDTrainer simplification + fix Removes `mtd.convert()` and the `DistillationModel` in-place class-swap for the HF path. The teacher model now lives directly on the trainer and is forwarded explicitly inside `compute_kd_loss_func`. This eliminates: - `mtd.convert()` in-place class swap and DynamicModule wrapping - Forward hooks for capturing intermediate outputs - `hide_teacher_model` / `hide_loss_modules` context managers for checkpointing - Deferred initialization branching (FSDP2 vs DDP/DeepSpeed) - `save_model` and `QADTrainer._quantize_model` overrides **Bug fix**: The previous `DistillationModel`/`mtd.convert()` approach did not support CPU RAM-efficient loading for QAD. The teacher model had to be fully loaded on GPU before wrapping, which doubled peak memory during initialization. The new approach loads the teacher lazily on the trainer, enabling standard HF device-map and low-cpu-mem-usage loading. Only logit-level distillation is supported for the HF path. The core `DistillationModel`/`mtd.convert()` API remains for Megatron and advanced intermediate-layer distillation use cases. ## Test plan - [x] `pytest tests/unit/torch/distill/` (29 passed) - [x] `pytest tests/unit/torch/opt/plugins/test_hf_patching.py` (2 passed) - [x] `pytest tests/unit/torch/opt/plugins/test_lr_config.py` - [x] Pre-commit hooks pass - [ ] GPU example tests: `pytest tests/examples/llm_qat/` (QAT, QAD, LoRA QAT, QLoRA) - [ ] GPU distill example: `pytest tests/examples/llm_distill/` 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added Liger fused loss support in `ModelOptHFTrainer` for distributed causal language models with JSD distillation loss support. * Introduced `ModelOptTrainerArguments` with new training CLI flags: per-parameter learning rates via YAML, parameter freezing, and manual garbage collection. * Simplified knowledge distillation trainer with logit-level distillation support. * **Documentation** * Updated example configurations and documentation with new training options and defaults. * Added learning rate configuration example guide. * **Tests** * Added test coverage for distillation training and per-parameter optimizer configuration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
115cae2584 |
Puzzletron bypass distillation part 2 core (#1469)
## Summary
This is PR 2 of 3 in the Puzzletron bypass/local-distillation stack.
This PR adds the bypass distillation core engine. It builds on PR 1’s
shared infrastructure, but it does not yet wire bypass into
the full Puzzletron pipeline.
Stack:
1. `ssameni/puzzletron-bypass-1-prereqs`: shared prerequisites
2. **This PR**: bypass distillation core
3. `ssameni/puzzletron-bypass-3-integration`: Puzzletron integration,
configs, docs, GPU coverage
## What Changed
- Added `modelopt.torch.puzzletron.bypass_distillation`.
- Added bypass run identity/fingerprinting, experiment naming, state
manifests, and completion tracking.
- Added stitched teacher/student model construction for local blockwise
distillation.
- Added bypass training loop with:
- pipeline-parallel teacher activation stitching
- per-block student losses
- gradient accumulation
- checkpoint/resume support
- validation hooks
- best/latest checkpoint realization
- Added bypass checkpoint helpers for saving optimizer/scaler state and
HF-format model checkpoints.
- Added scalable distributed checkpoint saving that gathers tensors per
safetensors file instead of materializing all shards on
rank 0.
- Restricted v1 `keys_to_learn` to subblock-level targets only:
- `entire_block`
- `subblock_attention`
- `subblock_ffn`
- `subblock_mamba`
- lists of those keys
- Added pipeline ownership helper for deriving owned blocks and
neighboring PP ranks.
- Added robust `trust_remote_code` / `auto_map` handling for checkpoint
config saves.
## Why
Bypass distillation is the local-distillation stage used to train
pruned/reconfigured blocks before they are added to the
replacement library.
Keeping the core engine separate from Puzzletron pipeline wiring makes
this PR reviewable on its own: reviewers can focus on
distributed training, checkpoint/resume semantics, and Sewing Kit
stitching behavior without also reviewing configs/docs/pipeline
integration.
## Tests
Added focused unit coverage for:
- bypass checkpoint utilities
- bypass run identity/fingerprinting and completion state
- `keys_to_learn` subblock selection
- LR scheduler behavior
- launch/sweep dispatch behavior
- HF checkpoint utility behavior
- stitched model factory buffer ownership
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Full bypass distillation workflow: blockwise stitched student/teacher
distillation, deterministic experiment IDs/fingerprints, resume-capable
checkpointing with atomic symlink updates, optional asynchronous
checkpoint saves, and improved support for trust-remote-code model
loading.
* Public bypass-distillation entrypoint exposed for easier launch.
* **Tests**
* Extensive unit tests covering bypass utilities, checkpoint behavior,
LR scheduler, stitched factory, and training orchestration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Sepehr Sameni <ssameni@nvidia.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
01dec93580 |
[6058870] Fix ONNX AutoCast keep_io_types for empty I/O tensors (#1626)
### What does this PR do?
Type of change: Bug fix
ONNX AutoCast (used by `modelopt.onnx.quantization` with
`--high_precision_dtype fp16/bf16`) failed its internal sanity check
with `Unexpected type in I/O tensor <name>, keep_io_types=True, original
type: 1, converted type: 10` when a network input or output is an
**empty tensor** (a tensor with a dimension of size 0).
Root cause: `PrecisionConverter._add_cast` treats empty tensors
specially — instead of inserting a `Cast` node, it "fake-casts" them by
retyping their `value_info` in place (there is no data to convert).
However, `setup_mappings` stores the same `graph.input` / `graph.output`
`ValueInfoProto` objects in `value_info_map`, so retyping an empty
tensor that is a network input/output silently changed the model's
declared I/O type to the low-precision type — violating the
`keep_io_types=True` contract and tripping the sanity check.
Fix (`modelopt/onnx/autocast/precisionconverter.py`): skip the
metadata-only fake-cast for protected network I/O tensors (when
`keep_io_types=True` and the target type differs from the original I/O
type) and fall through to insert a real `Cast` node. This preserves the
declared I/O type while still bridging precision into the low-precision
consumer/producer. Intermediate empty tensors are unchanged and still
fake-cast.
### Usage
```bash
# Previously failed with "Sanity check failed: Unexpected type in I/O tensor ...";
# now completes and produces a valid mixed-precision model.
python -m modelopt.onnx.quantization \
--quantize_mode int8 \
--high_precision_dtype fp16 \
--onnx_path model_with_empty_io_tensor.onnx \
--output_path model_int8_fp16.onnx
```
### Testing
- Added `test_empty_tensor_network_input_keep_io_types` (fp16 + bf16,
both type-inference paths) in
`tests/unit/onnx/autocast/test_precisionconverter.py`, asserting network
I/O types stay FP32 and a real `Cast` bridges the empty input. Verified
it fails before the fix and passes after.
- Full autocast unit suite passes: `CUDA_VISIBLE_DEVICES="" pytest
tests/unit/onnx/autocast/` → 202 passed.
- Reproduced the original failure on the reported model and confirmed
the command now succeeds end-to-end; the resulting ONNX model passes
`onnx.checker`, runs in ONNX Runtime with correct FP32 I/O, and builds a
strongly-typed TensorRT engine.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Fixed an ONNX AutoCast issue where empty tensors with
keep_io_types=True could silently change model I/O types; now a proper
Cast is inserted for protected inputs/outputs to preserve declared I/O
types.
* **Tests**
* Added a regression test to ensure empty network-input tensors are
handled correctly in mixed-precision ONNX conversion when
keep_io_types=True.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
a7b0a92047 |
EAGLE3 new model support: pipeline configs, triage docs, and Ministral-3 fixes (#1417)
## Summary EAGLE3 automation triage work (OKR-30): testing the 4-step EAGLE3 offline pipeline against 12 new model architectures, documenting failure modes, and fixing issues found. ### Code fixes (modelopt) | File | Change | |------|--------| | `modelopt/torch/speculative/utils.py` | Extend VLM detection in `load_vlm_or_llm` to check `text_config`/`llm_config` attrs (catches `mistral3` models) | | `modelopt/torch/speculative/plugins/modeling_fakebase.py` | Add `consolidated.safetensors` fallback for checkpoints with incomplete HF shards | | `modelopt/torch/export/plugins/hf_spec_configs.py` | Set `use_cache=True` in EAGLE export templates (fixes strict `huggingface_hub` validation) | ### Pipeline infrastructure - `examples/speculative_decoding/pipeline/eagle3/` — pipeline scripts and configs: - `offline_training.sh` — training + export with runtime patches for older container modelopt - `dump_offline_data_vllm.sh` — vLLM-based hidden state extraction (with speculators compat patches) - `dump_offline_data.sh`, `dump_offline_data_hf.sh` — alternative dump paths - 18 quick-fail-check YAMLs for 12 models - 4 standalone task1 YAMLs ### Documentation - `eagle3_triage_chart.md` — model test matrix, triage decision tree, per-model results, failure catalog - `eagle3_new_model_triage_guide.md` — step-by-step guide for triaging new models ### Model test results (as of 2026-05-27) | Model | task_0 | task_1 | task_2 | task_3 | Blocker | |-------|--------|--------|--------|--------|---------| | Qwen3-8B | - | - | - | - | Reference (existing) | | Kimi-K2.5 | - | - | - | - | Existing (GB200) | | **Ministral-3-8B** | SKIP | PASS | PASS | FAIL | `use_cache=null` in export (fixed) | | Ministral-3-14B | FAIL | - | - | FAIL | vLLM engine init fails | | Qwen3.5-35B-A3B | TIMEOUT | - | - | - | Data synth too slow | | gpt-oss-20b | FAIL | - | - | - | Tokenizer `HarmonyError` | | Step-3.5-Flash | TIMEOUT | - | - | - | Data synth time limit | | MiniMax-M2.5 | TIMEOUT | - | - | - | `trust_remote_code` needed | | DeepSeek-V3.2 | no log | - | - | - | May not be mirrored | | Qwen3.5-9B | - | - | - | - | Not yet run | | Qwen3.5-27B | - | - | - | - | Not yet run | | GLM-5 | - | - | - | - | Not yet run | ### Issues found and fixed | # | Issue | Fix | |---|-------|-----| | 1 | `mistral3` model type not detected as VLM | Check `text_config`/`llm_config` attrs in `load_vlm_or_llm` | | 2 | Missing HF shard file (Ministral-3-8B) | Fallback to `consolidated.safetensors` with Mistral native key aliases | | 3 | `use_cache=null` in exported EAGLE config | Set `use_cache=True` in export template configs | | 4 | speculators incompatible with vLLM container | Runtime patches in `dump_offline_data_vllm.sh` | | 5 | `offline_training.sh` infra issues | Rewritten with runtime patches for container modelopt | ## Test plan - [x] Ministral-3-8B training passes (`cicd_1779829129`) - [x] Ministral-3-8B export succeeds - [ ] Ministral-3-8B benchmark passes (`cicd_1779901409` — pending with all fixes) - [ ] Dry-run remaining model configs ## Note GitHub secret scanning alert #6 is a **false positive** — `Mistral3ForConditionalGeneration` (a HuggingFace model class name in a YAML comment) was flagged as a "Mistral AI API Key". 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
651fd223e6 |
feat: EAGLE3 LoRA co-training improvements (#1607)
Layer-selective LoRA injection for EAGLE3 co-training, optimizer-stable warmup (LoRA always in the optimizer; warmup gated by a flag), and an export+merge+lm_eval evaluation script. Review feedback addressed: - trust_remote_code is caller-controlled (TRUST_REMOTE_CODE / --trust_remote_code), default False - eagle_base_lora_start_layer raises ValueError instead of silently injecting zero adapters - eval_lora.sh validates HF_MODEL_CKPT / EAGLE_CKPT up front - added unit tests for start-layer injection and validation (8/8 passing, incl. on-cluster GPU run) Signed-off-by: Ye Yu <yeyu@nvidia.com> |
||
|
|
e40b4d69d6 |
Add Minitron pruning support for Gemma3 via Megatron-Bridge (#1604)
### What does this PR do?
Type of change: New feature (+ new tests)
Adds Minitron pruning support for **Gemma3** models loaded through
Megatron-Bridge (`Gemma3ForCausalLM` → `GPTModel`).
Gemma3's bridge implementation subclasses several megatron-core modules
and overrides `forward`/`__init__`, so `DMRegistry` (which matches by
`nn_cls.forward is parent.forward`) does not recognize them and
pruning's `convert_to_dynamic` raised:
```
KeyError: "<class '...Gemma3LanguageModelEmbedding'> is not registered for a dynamic module!"
```
There were actually **two** blockers, both fixed here:
1. **Layer spec was discarded.** `load_mbridge_model_from_hf`
unconditionally replaced `provider.transformer_layer_spec` with the
generic TE GPT spec (to disable grouped-GEMM), throwing away
`gemma3_layer_spec`. With the generic spec, plain
`TEDotProductAttention` received Gemma3's int `window_size` and crashed
at construction (`TypeError: 'int' object is not subscriptable`) before
pruning even started. Now the override is applied **only to MoE
models**; dense models keep the bridge's native spec.
2. **Custom layers weren't registered as dynamic modules.** Added
registrations for Gemma3's custom layers.
Changes:
- **`modelopt/torch/nas/plugins/mbridge.py`** (new): register dynamic
modules for `Gemma3LanguageModelEmbedding` (reuses the existing
embedding dynamic class), `Gemma3SelfAttention`, and the fused post-LN
`TERowParallelLinearLayerNorm` used by both `self_attention.linear_proj`
and `mlp.linear_fc2`. The fused `post_layernorm` is converted to a
dynamic module and sliced by the linear's `output_size` (==
`hidden_size`).
- **`modelopt/torch/nas/plugins/megatron.py`**: two small reuse seams in
`_DynamicSelfAttention` — build the core-attention dynamic class from
`type(self.core_attention)` (preserves `Gemma3TEDotProductAttention`)
and extract an overridable `_convert_linear_proj` hook. No behavior
change for megatron-core models.
- **`modelopt/torch/utils/plugins/mbridge.py`**: only override the layer
spec for MoE models; dense models keep the bridge's native spec.
- **`modelopt/torch/nas/plugins/__init__.py`**: import `mbridge` under
`import_plugin("megatron.bridge")`.
- **Tests**: new `get_tiny_gemma3` / `create_tiny_gemma3_dir` helpers;
parametrize `test_prune_minitron` over qwen3 and gemma3.
### Usage
```bash
# Prune a Gemma3 checkpoint with Minitron (same flow as other models)
torchrun --nproc_per_node 2 examples/megatron_bridge/prune_minitron.py \
--hf_model_name_or_path google/gemma-3-1b-it \
--prune_target_params 0.6e9 \
--output_hf_path /tmp/gemma-3-pruned
```
### Testing
`tests/examples/megatron_bridge/test_prune_minitron.py` now runs for
both qwen3 and gemma3. Verified in the megatron-bridge container:
Also verified the gemma3 case standalone on 1 GPU (PP=1) and 2 GPUs
(PP=2).
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (the spec-override change only
affects how dense Megatron-Bridge models are loaded — they now keep
their native, correct spec; MoE behavior is unchanged)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ (parametrized gemma3
coverage + tiny-model helper)
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ✅
### Additional Information
Other Megatron-Bridge models with custom layers were surveyed; only the
Gemma family needs this treatment. **Gemma1** (embedding scaling via
`EmbeddingScalingMixin`) and **Gemma2** (post-LN linear + non-TE
`Gemma2DotProductAttention` + custom output layer + mixin embedding) are
intentionally **left as outdated**. All other dense/MoE LLMs (llama,
qwen2/3, qwen3-moe, mistral, nemotron, deepseek, gpt-oss, etc.) use
standard megatron-core specs already covered by `megatron.py`.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added Minitron pruning support for Megatron-Bridge Gemma3 models.
* Improved Megatron-Bridge layer configuration handling for
Mixture-of-Experts scenarios.
* **Tests**
* Added Gemma3 model creation utilities for testing.
* Expanded pruning validation tests to cover both Qwen3 and Gemma3
models.
* **Examples**
* Updated pruning example to adjust tokenizer "use fast" detection based
on model architecture.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
196c091027 |
[1/N] Refactor llm_qat example: YAML configs + ModelOptArgParser (#1172)
### What does this PR do? Type of change: new example Refactors `examples/llm_qat` from a monolithic launch/script flow into a modular, config-driven Hugging Face QAT/QAD workflow. The high-level user flow is now: 1. Quantize a base model with a ModelOpt PTQ recipe. 2. Train or evaluate the quantized checkpoint with QAT, QAD, LoRA QAT, QLoRA, or fine-tuning configs. 3. Export the trained checkpoint for deployment. Highlights: - Replaces the legacy `examples/llm_qat/launch.sh` + `examples/llm_qat/main.py` path with separate `quantize.py` and `train.py` entrypoints. - Adds YAML-driven argument parsing through `ModelOptArgParser`, including `--config <yaml>` defaults, CLI overrides, and generated `examples/llm_qat/ARGUMENTS.md`. - Adds declarative configs under `examples/llm_qat/configs/` for training modes, dataset blends, and Accelerate backends. - Adds `dataset_utils.py` for weighted multi-source dataset blending, Hugging Face streaming, local dataset loading, distributed rank-aware loading, tokenization caching, pre-tokenization, chat templating, and assistant-token label masking. - Moves the Hugging Face QAD flow into the shared trainer path through `DistillArguments`, `QADTrainer`, and teacher-model distillation kwargs. - Updates `QuantizationArguments` to prefer recipe paths via `--recipe`, while keeping legacy `--quant_cfg` available with deprecation warnings for in-trainer quantization. - Updates `QATTrainer` handling for pre-quantized checkpoints, FSDP2 TensorQuantizer buffers, and recipe-resolved PTQ configs. - Adds the `general/ptq/int4_blockwise_weight_only` recipe and refreshes docs for NVFP4, FP8, INT4, FSDP2, DDP, DeepSpeed, QLoRA, and LLaMA-Factory integration. - Adds focused parser, dataset tokenization, assistant-mask, and example workflow coverage. ### Usage From the repo root: ```sh cd examples/llm_qat # 1. Quantize python quantize.py \ --model_name_or_path Qwen/Qwen3-8B \ --dataset_config configs/dataset/blend.yaml \ --recipe general/ptq/nvfp4_default-kv_fp8 \ --output_dir qwen3-8b-quantized # 2. QAT train accelerate launch --config-file configs/accelerate/fsdp2.yaml train.py \ --config configs/train/qat_nvfp4.yaml \ --model_name_or_path qwen3-8b-quantized \ --output_dir qwen3-8b-qat-nvfp4 # 3. QAD train accelerate launch --config-file configs/accelerate/fsdp2.yaml train.py \ --config configs/train/qad_nvfp4.yaml \ --model_name_or_path qwen3-8b-quantized \ --teacher_model Qwen/Qwen3-8B \ --output_dir qwen3-8b-qad-nvfp4 ``` Dataset blends can be pre-tokenized and cached before training: ```sh cd examples/llm_qat python dataset_utils.py \ --dataset_config configs/dataset/blend.yaml \ --model_name_or_path Qwen/Qwen3-8B ``` ### Testing Focused coverage added or updated: - `tests/unit/torch/opt/plugins/test_modelopt_arg_parser.py` - `tests/examples/llm_qat/test_dataset_tokenization.py` - `tests/examples/llm_qat/test_assistant_mask.py` - `tests/examples/llm_qat/test_llm_qat.py` Recorded validation: - [x] `pytest tests/unit/torch/opt/plugins/test_modelopt_arg_parser.py` - [x] `pytest tests/examples/llm_qat/test_llm_qat.py::test_dataset_utils_pretokenize` - [x] `pre-commit run --all-files` - [ ] `pytest tests/examples/llm_qat/test_llm_qat.py` full GPU/backend suite ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: No for the legacy `examples/llm_qat` CLI/file layout (`main.py`, `launch.sh`, and FSDP1 config are removed); library `quant_cfg` usage remains available but is deprecated for this workflow. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: Yes. `train.py`/`simple_qat_train.py` retain the upstream Alpaca attribution where applicable; `examples/llm_qat/requirements.txt` switches `tensorboardX` to `tensorboard`. - Did you write any new necessary tests?: Yes. Parser, dataset tokenization, assistant masking, pre-tokenization, and QAT/QAD workflow tests were added or updated. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: Yes. - Did you get Claude approval on this PR?: N/A in this description update. ### Additional Information This PR intentionally changes the `llm_qat` example surface. Existing users should move from `examples/llm_qat/main.py` and `examples/llm_qat/launch.sh` to `examples/llm_qat/quantize.py`, `examples/llm_qat/train.py`, and the YAML configs under `examples/llm_qat/configs/`. --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
902d36921a |
[Feat]: Streaming Hidden-states Dataset (#1509)
### What does this PR do? Type of change: new feature **Design doc:** https://gist.github.com/h-guo18/241c94968b0591324c361d97cf995dd0 Jira ticket: https://jirasw.nvidia.com/browse/OMNIML-4341 Sandbox CI: https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/327911555#L2228 Streaming hidden-states dataset: per-sample activations pulled from a live `vllm serve` over HTTP, replacing on-disk activation dumps. Two axes for future extensions: - **Backend** (`_fetch`): vLLM now; TRT-LLM / SGLang next. - **Algorithm** (`_format`): Eagle now; distillation / probing next. - **Sandbox CI**: https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/merge_requests/169 Shared plumbing — async producer, token-level truncation to `training_seq_len`, loss-mask alignment, DDP via Accelerate's dispatcher (rank 0 fetches, broadcasts), circuit breaker, resume — lives in `StreamingDataset`. First instance: **`EagleVllmStreamingDataset`**. API: `data.mode ∈ {online, offline, streaming}`; legacy configs auto-promote. ### Usage ```yaml data: mode: streaming data_path: input_conversations/train.jsonl streaming_server_url: http://localhost:8000 streaming_model_name: meta-llama/Llama-3.1-8B-Instruct training: training_seq_len: 4096 # also caps the prompt sent to vllm ``` Requires `vllm serve` with `ExampleHiddenStatesConnector` and `dataloader_num_workers=0`. End-to-end Slurm pipeline: `tools/launcher/examples/Qwen/Qwen3-8B/hf_streaming_eagle3.yaml`. ### Testing - **Unit**: full-corpus invariant, rank-0-only iter, resume, circuit breaker, mocked-httpx integration. - **E2E**: `launch_train.sh` against a stdlib `HTTPServer` mimicking the connector. - **Smoke** (Qwen3-8B / 8×H100 / 4096 ultrachat samples, single epoch): train_loss 32 → 18, MT-Bench AR 1.003 → 1.20. ### TODO before un-drafting - [ ] Observability counters (filtered, fetch failures, queue depth, latency). - [ ] Changelog entry. - [x] Add test in sandbox. **Non-goals (v1):** multi-epoch streaming, cross-rank dynamic load balancing. ### Before your PR is "*Ready for review*" - Backward compatible: ✅ (legacy configs auto-promote) - New PIP dep: N/A - New tests: ✅ - Changelog: ❌ (TODO above) - Claude approval: ❌ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Streaming training mode with server-backed hidden-state fetching, deterministic seed control, resume support, and streaming-specific dataset options (server, model, prefetch, shared storage). * **Behavior / Bug Fixes** * Stronger mode validation; offline behavior derived from data mode; resume handling adjusted to avoid double-skip during streaming runs. * **Tests** * End-to-end CI streaming test and expanded unit tests covering streaming, resume, DDP, determinism, and failure cases. * **Infrastructure** * Launcher script and pipeline config for end-to-end streaming training. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1509?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
f21977a5fc |
Add Megatron-Bridge PTQ quantize + export example scripts (#1589)
### What does this PR do?
Type of change: new example
Adds a two-step post-training quantization (PTQ) flow for
**Megatron-Bridge** models under `examples/megatron_bridge/`, mirroring
the Megatron-LM `quantize.sh` / `export.sh` split:
- **`quantize.py`** — loads an HF model via Megatron-Bridge, applies
ModelOpt PTQ (via a `--quant_cfg` alias / full config name, or a
`--recipe` YAML), with optional KV-cache quant, weight-only,
compression, and MoE expert-ratio calibration, then saves a **Megatron
checkpoint** (with ModelOpt state). Tensor / pipeline / expert
parallelism are all supported, and the checkpoint can later be reloaded
for further training (QAT / distillation).
- **`export.py`** — loads the quantized Megatron checkpoint, **re-shards
to TP=1**, and exports a **HuggingFace (unified)** checkpoint deployable
with TensorRT-LLM / vLLM / SGLang.
**Why the split?** The unified HF exporter (`export_mcore_gpt_to_hf`)
does not gather tensor-parallel-sharded weights — Megatron-LM likewise
forces `TP=1` during its export step. Saving a TP-sharded Megatron
checkpoint first lets us calibrate at TP>1 (to fit large models) and
then reload re-sharded to TP=1 for the HF export. A combined
single-script flow silently produced corrupt HF checkpoints under TP>1
(collided per-rank shards), which this split avoids.
> **Note:** This is **part 1 of 4**:
> - **Part 1 (this PR):** Megatron-Bridge `quantize.py` + `export.py`
support and tests.
> - **Part 2:** extend `distill.py` for quantization-aware distillation
(QAD) — load a quantized Megatron checkpoint as the student.
> - **Part 3:** add NVFP4 + QAD-on-pruned-checkpoint experiments to the
Nemotron-3-Nano-30B-A3B tutorial.
> - **Part 4:** repeat the NVFP4 + QAD experiments on a non-Nemotron
model.
### Usage
```bash
# Step 1: quantize (TP/PP/EP supported) -> Megatron checkpoint
torchrun --nproc_per_node 2 quantize.py \
--hf_model_name_or_path Qwen/Qwen3-8B \
--quant_cfg fp8 \
--tp_size 2 \
--export_megatron_path /tmp/Qwen3-8B-FP8-megatron
# Step 2: export -> deployable HuggingFace (unified) checkpoint (re-shards to TP=1)
torchrun --nproc_per_node 1 export.py \
--hf_model_name_or_path Qwen/Qwen3-8B \
--megatron_path /tmp/Qwen3-8B-FP8-megatron \
--export_unified_hf_path /tmp/Qwen3-8B-FP8-hf
```
### Testing
`tests/examples/megatron_bridge/test_quantize.py` (validated on a 2-GPU
NeMo `26.04` container):
- `test_quantize_export_and_vllm_deployment` — quantize a tiny Qwen3 via
a recipe at TP=2 → `export.py` re-shards to TP=1 → load + generate with
**vLLM** (skipped if vLLM absent).
- `test_quantize_megatron_checkpoint_reload` — quantize at TP=2 → reload
the Megatron checkpoint via the bridge and assert ModelOpt quantizers
were restored.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: N/A (new example)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new
dependencies)
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅
### Additional Information
The Nemotron-3 tutorial update to use these scripts is intentionally
**not** included here — it ships with the part 3 PR alongside the NVFP4
+ QAD experiments.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Expanded post-training quantization (PTQ) workflow documentation with
detailed step-by-step examples and configuration guidance for the
Megatron-Bridge framework.
* **New Features**
* Added quantization tool for applying PTQ to Megatron models with
calibration support.
* Added export tool for converting quantized models to a deployable
format.
* **Tests**
* Added integration tests validating the complete
quantization-export-deployment workflow, including inference validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
8f968322f6 |
[OMNIML-3994] Make sure all weight quantizers have _amax (#1560)
### What does this PR do?
- max_calibrate now always runs weight_only_quantize before the optional
forward_loop, so every weight quantizer has _amax populated regardless
of MoE routing.
- awq_lite search loop explicitly deletes _amax around the alpha sweep
to keep using the dynamic-amax path; no restore needed (postprocess
overwrites for enabled modules, disabled modules early-return).
- Removed three band-aids:
- _bootstrap_uncalibrated_weight_quantizers (mse_calibrate)
- per-module max_calibrate fallback in awq_lite postprocess
- _ensure_weight_quantizer_calibrated + helpers (export/quant_utils)
- Renamed test_bootstrap_populates_dead_expert_quantizers ->
test_max_calibrate_populates_dead_expert_quantizers; deleted the GPU
lazy-calibration test that no longer has a behavior to exercise.
Type of change: refactor
<!-- Details about the change. -->
### Usage
same as before.
### Testing
run unittests locally
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Refactor**
* Simplified the weight quantization calibration process by
consolidating logic and removing redundant internal functions.
* **Tests**
* Removed a redundant quantization format-specific export test.
<!-- review_stack_entry_start -->
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1560?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)
<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
|
||
|
|
72df833e0b |
Add active-MoE AutoQuant cost accounting (#1497)
### What does this PR do?
• Type of change: new feature
Adds an active_moe cost model for auto_quantize effective-bits search.
This lets AutoQuant account for routed MoE expert weights by active
decode weight traffic instead of total checkpoint weight
size, using active_moe_expert_ratio = num_experts_per_tok / num_experts.
The default behavior is unchanged: cost_model="weight" still counts all
quantizable weights equally.
### Usage
import modelopt.torch.quantization as mtq
model, search_state = mtq.auto_quantize(
model,
constraints={"effective_bits": 5.0},
quantization_formats=[
mtq.NVFP4_DEFAULT_CFG,
mtq.FP8_DEFAULT_CFG,
],
data_loader=calib_dataloader,
forward_step=forward_step,
loss_func=loss_func,
cost_model="active_moe",
# Optional. If omitted, ModelOpt tries to infer this from model.config.
active_moe_expert_ratio=2 / 64,
)
The HF PTQ example also exposes:
--auto_quantize_cost_model active_moe \
--auto_quantize_active_moe_expert_ratio 0.03125
### Testing
python -m pytest tests/unit/torch/quantization/test_autoquant.py -q -k
'active_moe or quant_recipe_hparam_cost_weight'
python -m pytest tests/unit/torch/quantization/test_autoquant.py -q -k
'not data_parallel_auto_quantize'
Results:
- 4 passed
- 58 passed, 1 deselected
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added active-MoE cost model option for auto-quantization with
configurable expert ratio; API and CLI accept cost_model and
active_moe_expert_ratio
* Unified auto-quantize supports new quant format w4a16_nvfp4
* **Bug Fixes**
* Ensure labels are moved to the logits device for base models without
an lm_head
* CLI enforces valid expert-ratio range and requires active-MoE mode
when a ratio is provided
* **Tests**
* Added unit tests for active-MoE behavior, cost-weighting, ratio
handling, and search budget selection
<!-- review_stack_entry_start -->
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1497?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)
<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
|
||
|
|
f0d2237cbc |
Add Qwen3VL MCore Export support from PR 895 (#1482)
# [Megatron Export] Add Qwen3-VL mcore ↔ HF weight mapping > This PR is duplicated from [PR #895](https://github.com/NVIDIA/Model-Optimizer/pull/895). > The original branch source is no longer available; this new branch carries the same changes forward. ## What does this PR do? **New feature:** Add Qwen3-VL (Vision-Language) model support to the Megatron Core export/import plugin, enabling HuggingFace-to-mcore weight conversion for PTQ/QAT/QAD workflows. ### Overview Qwen3-VL has a different weight structure from Qwen3 text-only models: - Language model weights are under `model.language_model.` prefix (not `model.`) - Visual encoder weights are under `model.visual.` prefix - `lm_head` is at root level, not nested under `language_model` ### What changed | File | Change | |---|---| | `modelopt/torch/export/plugins/mcore_qwen3vl.py` | New plugin: derives Qwen3-VL mcore↔HF mapping by rewriting `model.*` → `model.language_model.*` on top of the existing Qwen3 dense rules; `lm_head.` is intentionally left unchanged | | `modelopt/torch/export/plugins/mcore_common.py` | Registers `Qwen3VLForConditionalGeneration` in `all_mcore_hf_export_mapping` and `all_mcore_hf_import_mapping` | | `modelopt/torch/export/plugins/hf_checkpoint_utils.py` | Generalized `load_multimodal_components` with a `prefixes` parameter; sharded checkpoints now scan all shards (not just the first) | | `modelopt/torch/export/unified_export_megatron.py` | `save_pretrained`: added Qwen3-VL branch that copies `model.visual.*` vision-encoder weights from the original HF checkpoint into the exported directory, producing a complete, loadable checkpoint | | `tests/_test_utils/torch/transformers_models.py` | Added `get_tiny_qwen3vl` / `create_tiny_qwen3vl_dir` helpers; Qwen3VL classes are lazy-imported inside the function to avoid collection failures on older transformers builds | | `tests/gpu_megatron/torch/export/test_unified_export_megatron.py` | Integrated Qwen3-VL export/import tests into the existing `test_unified_export_megatron` / `test_unified_import_megatron` parametrized suites; removed standalone `test_mcore_qwen3vl.py` | | `docs/source/deployment/3_unified_hf.rst` | Added Qwen3-VL (FP8 / NVFP4) to the deployment support matrix for TensorRT-LLM | ### Workflow coverage | Step | Status | Files | |---|---|---| | 1. Quantize Qwen3-VL with `hf_ptq` | ✅ existing | — | | 2. Export quantized mcore → HF | ✅ this PR | `plugins/mcore_qwen3vl.py` (weight name mapping), `unified_export_megatron.py` (export path) | | 3. Vision-encoder weights merged into export dir | ✅ this PR | `plugins/hf_checkpoint_utils.py` (`load_multimodal_components` with `prefixes`), `unified_export_megatron.py` (calls it when `arch == "Qwen3VLForConditionalGeneration"`) | | 4. Import HF checkpoint back to mcore | ✅ this PR | `plugins/mcore_qwen3vl.py` (same mapping, reverse direction), `unified_export_megatron.py` (import path) | ### Design notes - **MoE not supported**: `Qwen3VLMoeForConditionalGeneration` stores expert weights as 3-D tensors (`mlp.experts.gate_up_proj`, `mlp.experts.down_proj`) that require a dedicated fused-expert mapping. A `NotImplementedError` comment in the plugin documents this explicitly. - **`copy.deepcopy` on `func_kwargs`**: each mapping entry gets its own copy to prevent shared-dict mutation when both Qwen3 and Qwen3-VL rules are loaded. - **`prefixes` parameter on `load_multimodal_components`**: backward-compatible default preserves existing LLaVA behaviour (`"multi_modal_projector"`, `"vision_model"`); Qwen3-VL callers pass `("model.visual.",)`. - **Sharded checkpoint scan**: the old code only looked in the first shard. The Qwen3-VL vision encoder can span multiple shards, so all shards are now scanned. ## Usage From the [Megatron-LM PR comment](https://github.com/NVIDIA/Megatron-LM/pull/3444#issuecomment-3911271713): > Qwen3VL is supported within [Megatron-Bridge](https://github.com/NVIDIA-NeMo/Megatron-Bridge), and pretraining and PEFT recipes for Qwen3VL are [here](https://github.com/NVIDIA-NeMo/Megatron-Bridge/blob/main/src/megatron/bridge/recipes/qwen_vl/qwen3_vl.py) and the core code logic [here](https://github.com/NVIDIA-NeMo/Megatron-Bridge/tree/main/src/megatron/bridge/models/qwen_vl). Create `Megatron-LM/examples/post_training/modelopt/conf/Qwen/Qwen3-VL-8B-Instruct.sh`: ```bash #!/bin/bash # Qwen3-VL-8B-Instruct text-model config for Megatron-LM import/quantize. # # Text-model dimensions are identical to Qwen3-8B (4096 hidden, 36 layers, # 32 heads, GQA=8). Differences: rope_theta=5000000, checkpoint path uses # model.language_model.* prefix (handled by mcore_qwen3vl plugin). if [ -z ${HF_MODEL_CKPT} ]; then HF_MODEL_CKPT=Qwen/Qwen3-VL-8B-Instruct TOKENIZER_MODEL=Qwen/Qwen3-VL-8B-Instruct else TOKENIZER_MODEL=${HF_MODEL_CKPT} fi MODEL_ARGS=" \ --save-interval 100000 \ --micro-batch-size 1 \ --bf16 \ --no-masked-softmax-fusion \ --disable-bias-linear \ --untie-embeddings-and-output-weights \ --position-embedding-type rope \ --no-rope-fusion \ --normalization RMSNorm \ --swiglu \ --num-layers 36 \ --hidden-size 4096 \ --ffn-hidden-size 12288 \ --num-attention-heads 32 \ --group-query-attention \ --num-query-groups 8 \ --kv-channels 128 \ --qk-layernorm \ --seq-length 4096 \ --max-position-embeddings 262144 \ --tokenizer-type HuggingFaceTokenizer \ --make-vocab-size-divisible-by 1187 \ --use-mcore-models \ --rotary-percent 1.0 \ --rotary-base 5000000 \ --no-bias-swiglu-fusion \ " ``` Import Qwen3-VL from HuggingFace to MCore (local, requires GPUs): ```bash MLM_MODEL_CFG=Qwen/Qwen3-VL-8B-Instruct \ HF_MODEL_CKPT=Qwen/Qwen3-VL-8B-Instruct \ MLM_MODEL_SAVE=/tmp/qwen3vl_mcore \ TP=1 \ bash Megatron-LM/examples/post_training/modelopt/convert.sh Qwen/Qwen3-VL-8B-Instruct ``` Quantize (PTQ via Megatron-LM path): ```bash MLM_MODEL_CFG=Qwen/Qwen3-VL-8B-Instruct \ HF_MODEL_CKPT=Qwen/Qwen3-VL-8B-Instruct \ QUANT_CFG=NVFP4_DEFAULT_CFG \ TP=4 \ bash Megatron-LM/examples/post_training/modelopt/quantize.sh Qwen/Qwen3-VL-8B-Instruct ``` ## Testing - Verified round-trip import/export with Qwen3-VL-8B-Instruct with the example usage above - Unit/GPU tests covering: - Registration in global export/import mappings - Import mapping: dense keys, `model.language_model.` prefix, `lm_head.` at root, `QKVMerging`, `GatedMLPMerging`, `REPLICATE` for layernorms, TP sharding configs - Export mapping: `QKVSlicing`, `GatedMLPSlicing`, no `parallel_config` - Import/export symmetry: same mcore keys, matching HF prefixes - Qwen3-VL vs Qwen3 difference: same keys, VL adds `language_model.` prefix, `lm_head` unchanged ## Before your PR is "Ready for review" - Is this change backward compatible?: Yes, additive only - Did you write any new necessary tests?: Yes, `tests/gpu_megatron/torch/export/test_unified_export_megatron.py` - Did you add or update any necessary documentation? Yes, see `docs/source/deployment/3_unified_hf.rst` - Did you update Changelog? Yes, see `CHANGELOG.rst` ## Additional Information Companion Megatron-LM PR adds `Qwen3VLModel`, `Qwen3VLDataset`, and `pretrain_qwenvl.py`. See: https://github.com/NVIDIA/Megatron-LM/pull/3444 --------- Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com> Signed-off-by: hychiang <hungyuehc@nvidia.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
40a4dd326d |
[Feat]: Eagle Dry Run Mode (#1566)
### What does this PR do? Type of change: new feature Adds `--dry_run` to `examples/speculative_decoding/main.py`: load → `mtsp.convert` → save, then exit (no `trainer.train()`). With `FakeBaseModel`, the convert→save→export chain runs in seconds and produces an exportable EAGLE3 / Medusa / DFlash checkpoint with correct structure but untrained draft-head weights — useful for end-to-end plumbing smoke tests on downstream stacks (vLLM, TRT-LLM, SGLang) without paying for a real training run. Where it sits among existing EAGLE3 modes: ``` EAGLE3 modes ├── online base model runs forward in-loop ├── offline reads pre-dumped hidden states from disk ├── streaming streams hidden states from a live server in-loop └── dry-run ★ skip training entirely; convert + save + export ← NEW (this PR) ``` Companion `FakeBaseModel` fixes so small base checkpoints work: - Synthesize the weight_map from a single `model.safetensors` when no sharded index is present (Llama-3.2-1B, Qwen3-0.6B, …). - Honor `tie_word_embeddings`: reuse `embed_tokens` when `lm_head` is absent from safetensors. A new launcher YAML (`tools/launcher/examples/Qwen/Qwen3-8B/hf_eagle3_dryrun.yaml`) wires this together as a one-task pipeline. ### Usage ```bash # Direct python main.py --dry_run \ --config modelopt_recipes/general/speculative_decoding/eagle3.yaml \ model.model_name_or_path=meta-llama/Llama-3.1-8B-Instruct \ model.use_fake_base_for_offline=true \ data.offline_data_path=/tmp/dryrun-placeholder \ training.output_dir=ckpts/dryrun python scripts/export_hf_checkpoint.py --model_path ckpts/dryrun --export_path export/dryrun # Launcher uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_eagle3_dryrun.yaml --yes ``` ### Testing - `tests/unit/torch/speculative/plugins/test_fakebase.py`: 3 new cases — single-file fallback, tied-embeddings fallback, and the negative case (missing `lm_head` without tying). - `tests/examples/speculative_decoding/test_eagle.py::test_eagle3_dry_run`: full `launch_train.sh --dry_run → export_hf_checkpoint.py` chain on `tiny_llama`; asserts exported state_dict has all `LLAMA_EAGLE_SINGLE_LAYER` required keys. - Manually verified end-to-end on Llama-3.1-8B-Instruct (sharded), Llama-3.2-1B-Instruct (single-file + tied), and Qwen3-0.6B (single-file). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `--dry_run` is opt-in; `FakeBaseModel` changes are additive fallbacks. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — examples-only addition. - Did you get Claude approval on this PR?: ❌ ### Additional Information N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a `--dry_run` CLI flag to perform a fast early-exit execution that saves model artifacts without training. * Added support for single-file model checkpoint formats in the loader. * **Bug Fixes** * Improved checkpoint-loading error messages and handling. * Enhanced tied-embeddings fallback when head weights are absent. * **Tests** * Added integration and unit tests covering dry-run behavior and single-file/tied-embedding loading. * **Documentation** * Added a launcher example for dry-run smoke testing. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1566?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
d7e72f42ed |
Refine static NVFP4 MSE calibration (#1536)
### What does this PR do? Type of change: Bug fix Refines static NVFP4 MSE calibration and forces static NVFP4 amax state to stay FP32 across calibration loading, quantizer promotion, dtype casts, and restore paths. Main changes: - Tighten max/MSE calibration bootstrap and static NVFP4 quantizer promotion. - Keep static NVFP4 `_amax` and `_global_amax` in FP32. - Update focused GPU/unit coverage for FP8 sweep calibration, promotion, restore, and FP32 amax preservation. ### Usage ```yaml algorithm: method: mse fp8_scale_sweep: true ``` ### Testing ```bash pre-commit run --files modelopt/torch/quantization/nn/modules/tensor_quantizer.py tests/gpu/torch/quantization/test_nvfp4_static_quantizer_cuda.py pytest_pwd tests/gpu/torch/quantization/test_nvfp4_static_quantizer_cuda.py -q ``` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors). - Is this change backward compatible?: Yes - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: Yes - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: Yes ### Additional Information N/A Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
f99d83eeab |
[Auto-23/24][ONNX][Autocast] Clear stale Cast-output type metadata before ORT InferenceSession load (#1565)
### What does this PR do? Type of change: Bug fix **Error:** ``` onnxruntime.capi.onnxruntime_pybind11_state.Fail: [ONNXRuntimeError] : 1 : FAIL : Type Error: Type (tensor(float16)) of output arg (node_5bc985fa) of node (node_5bc985fa) does not match expected type (tensor(float)). ``` **Root cause:** - Some ONNX exporters emit `graph.output` / `value_info` entries whose dtype disagrees with the upstream `Cast` node's `to` attribute. - ORT's type checker rejects such models on session load. **Fix:** - New helper `modelopt.onnx.utils.clear_stale_value_info()` reconciles each `graph.output` elem_type to its producing Cast's `to`, then clears `value_info` so ORT recomputes intermediate types. - Called from `autocast/referencerunner.py` and `quantization/quantize.py::_preprocess_onnx`. ### Usage ```python # Internal fix; no new flag introduced. Generic CLI to exercise the affected Autocast path: $ python -m modelopt.onnx.autocast --onnx=model.onnx ``` ### Testing ``` pytest tests/unit/onnx/test_onnx_utils.py::test_clear_stale_value_info ``` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Automatic cleaning and reconciliation of stale ONNX type metadata before runtime and quantization; reconciled model files are produced and used when inconsistencies are found. * **Tests** * New unit tests covering metadata-cleaning behavior across cast/type scenarios to ensure correctness and prevent regressions. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1565?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com> Co-authored-by: modelopt-fix-agent-bot (Claude Opus 4.7) <noreply@anthropic.com> |
||
|
|
eb5ed2df68 |
[CI] Bump torch, transformers and dev containers to latest (#1554)
- Transformers upper bound bumped from `<5.8` to `<5.10` - Enable torch 2.12 CICD testing - Bump TRT-LLM container to `1.3.0rc16` (transformers 5.5) - Use pytorch and tensorrt 26.04 containers in CICD <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated CI test container images and targeted Torch version across workflows; adjusted release CI job to use the newer torch config. * Broadened Transformers constraint in project metadata and test/dev pins. * Removed strict transformers pins from example requirements and lifted a compression dependency cap. * Raised the import-time Transformers version threshold for compatibility warnings. * **Tests** * Refactored a GPU test to collect and report validation errors and updated numeric expected baselines. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1554?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
4b270f0ea6 |
Support Mixed precision & Static MSE in MCore; Nemotron Super v3 NVFP4 recipe (#1521)
### What does this PR do?
Type of change: New features + Bug fixes
Mixed Precision and MSE support in MCore PTQ
- support mixed precision export in MCore by detecting mixed precision
layers in HF Quant Config
- Restore static quantizer in MCore checkpoint restore as `NVFP4QTensor`
(not TensorQuantizer which can call max calibrate. we want to skip max
calibrate for static quantizer during restore) --> fixes bug during
MCore export for MSE
- Fix dynamic block quantizer detection when `block_sizes` is
dict-backed.
- Add a YAML quantization recipe that roughly mirrors Nemotron 3 Super
NVFP4 `hf_quant_config.json`
Export bug fixes
- copy .py files properly from original HF ckpt (for reasoning parser
etc)
## Super recipe
Mirrors the published nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
hf_quant_config.json:
- MoE routed experts (mixer.experts.<N>.{up,down}_proj): NVFP4 W4A4
weight MSE, group_size 16
- MoE shared experts (mixer.shared_experts.{up,down}_proj): FP8
per-tensor
- Mamba mixer linears (mixer.{in,out}_proj): FP8 per-tensor
- KV cache: FP8
rest: not quantized
### Usage
```python
# Add a code snippet demonstrating how to use this
```
### Testing
Tested on Nemotron model
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added NVFP4 (4-bit) quantization checkpoint restore and export support
for Megatron-Core models
* Added tokenizer file export capability in model checkpoints
* Extended quantization support for expert-parallel distributed training
* Introduced new PTQ recipes for Nemotron-3-Super models with
mixed-precision quantization
* **Bug Fixes**
* Fixed FP8 and FP4 hardware compatibility detection on non-CUDA systems
* Improved offline Hugging Face Hub access handling with better error
messaging
* Enhanced calibration validation for mixture-of-experts models
* Fixed amplitude maximum validation for static block quantizers
* **Documentation**
* Updated expert weight quantization configuration documentation
<!-- review_stack_entry_start -->
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1521?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)
<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Jenny Chen <jennifchen@nvidia.com>
|
||
|
|
a9c156e24a |
Split bypass prerequisites (#1468)
## Summary This is PR 1 of 3 in the Puzzletron bypass/local-distillation stack. This PR contains prerequisite infrastructure only. It does not wire bypass distillation into the Puzzletron pipeline yet. Stack: 1. **This PR**: shared prerequisites 2. `ssameni/puzzletron-bypass-2-core`: bypass distillation core 3. `ssameni/puzzletron-bypass-3-integration`: Puzzletron integration, configs, docs, GPU coverage ## What Changed - Added `ModelDescriptor.pruning_mixins()` so model families can expose pruning mixins needed by downstream bypass initialization. - Added KV-head pruning mixin support for GPT-OSS, Nemotron-H, Nemotron-H-v2, and Qwen3-VL descriptors. - Improved pruning utilities for nested language-model configs and missing attention bias config fields. - Added `create_train_dataloader()` and streaming-safe shuffle handling. - Added chat-template fallback for base models without `tokenizer.chat_template`. - Added Sewing Kit loss/helper exports needed by the later bypass core. - Updated child-state initialization to support composing multiple pruning mixins. - Updated warmup-step resolver to account for gradient accumulation. ## Why The bypass distillation MR needs these reusable pieces, but they are independently reviewable and useful without adding the bypass training stage itself. Splitting them out keeps the bypass core PR focused on the actual local-distillation engine. ## Tests Added focused unit coverage for: - Dataloader behavior - Bypass loss helpers - KV-head pruning utilities - Sewing Kit activity/input/function/needle behavior <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * KV-head pruning added for multiple model families; generic pruning mixin hook available. * New training dataloader factory for infinite, block-sized training streams. * Vectorwise and batched normalized MSE loss utilities. * **Improvements** * Loss reports now show Δ-from-initial and visual indicators. * Chat-sample preprocessing tolerates tokenizers without chat templates. * More robust head-dimension/bias handling and grad-accum-aware warmup resolver. * **Tests** * Extensive unit tests added across dataloaders, losses, pruning, hydra utils, and sewing-kit. [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1468) <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Sepehr Sameni <ssameni@nvidia.com> |
||
|
|
7ae18651a5 |
Fix puzzletron tests and make FFN pruning check robust (#1556)
### What does this PR do? Type of change: Bug fix + test robustness Fixes the puzzletron GPU tests and makes them robust to GPU vendor / transformers version drift. The previous exact-equality checks failed in CI because bf16 attention paths produce different operation orderings on different GPUs (Ada vs Blackwell), which completely reshuffles close-in-magnitude FFN importance scores on the small (hidden=256) test models. Earlier set-membership relaxations weren't enough — the orderings between GPUs were near-orthogonal. **Changes:** *Numerical determinism (root cause):* - Switch test forward to **fp32 autocast** (`autocast_dtype: torch.float32` in `validate_model_defaults.yaml`). Combined with TF32 disabled (below), this gives bit-identical results across GPU vendors. - Reduce `block_size` from 8192 → **2048**: SDPA's flash attention only supports bf16/fp16, so at fp32 it materializes the full attention matrix; seq=8192 OOMs (32 GB), seq=2048 fits. - Strengthen `set_seed`: also disable TF32 on cuda matmul + cudnn so fp32 matmuls produce bit-identical results across GPU vendors. *Test structure (orthogonal improvements):* - Refactor `_assert_*` helpers to `_check_*` that return error lists; the multiprocess job collects all failures and fails once at the end (instead of fail-fast). Single CI run now surfaces every mismatch. - Drop the brittle exact-equality `score[0]` (channel-rank) check. Use **set-membership at both ends of the importance ranking**: - Expected least-important channel must land in top-K least-important of actual. - Expected most-important channel must land in top-K most-important of actual. - `K = max(8, num_channels // 16)` gives slack for any residual drift without defeating the check. *Tolerances:* - `_check_lm_loss` default tolerance: **0.001** for non-MoE (now essentially bit-deterministic). - MoE caller passes **0.01** to absorb expert-routing run-to-run noise (~0.003). - Nemotron-3-Nano-30B-A3B (hybrid mamba+MoE) remains excluded from `EXPECTED_LM_LOSS` — its ~0.06 run-to-run noise would require a tolerance that defeats the check. *Values:* - Update `EXPECTED_FFN_PRUNING_VALUES` (with `least_important` + `most_important` per layer) and `EXPECTED_LM_LOSS` for the new fp32 + shorter-seq baseline. ### Testing Verified locally on RTX 6000 Ada across transformers `4.57.0`, `5.6.0`, `5.9.0`, in both 1-GPU and 2-GPU modes: | | 4.57 | 5.6 | 5.9 | |-------|-----|-----|-----| | 1-GPU | 8/9 † | 9/9 | 9/9 | | 2-GPU | — | — | 9/9 | † Qwen3-VL-30B-A3B fails on 4.57: its lm_loss is 4.92 on 4.57 but 5.03125 on 5.6+. This is a pre-existing cross-major-version drift documented in a code comment; not introduced by this PR. CI runs transformers 5.9 (container 26.04). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (test-only change) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
e012f064e1 |
[Quant] Fix padded NVFP4 MSE calibration (#1557)
### What does this PR do? Type of change: Bug fix Fixes static NVFP4 MSE calibration when the weight's last dimension needs block padding. The reference MSE path now returns fake-quantized tensors to the padded static block layout through the quantizer helper before comparing losses, so padded weights such as `512x60` compare as `2048x16` against `2048x16` instead of `2048x16` against `1920x16`. The local-Hessian calibration path now reuses the same helper, and a CUDA regression test covers the forced reference sweep for a padded `nn.Linear(60, 512)` weight. Fixes #1552. ### Usage ```python import copy import torch from torch import nn import modelopt.torch.quantization as mtq model = nn.Linear(60, 512, bias=False).eval().cuda().bfloat16() inputs = torch.randn(2, 60, device="cuda", dtype=torch.bfloat16) mtq.quantize( model, copy.deepcopy(mtq.NVFP4_W4A4_WEIGHT_MSE_FP8_SWEEP_CFG), forward_loop=lambda model: model(inputs), ) ``` ### Testing - `git diff --check` - `MODELOPT_NVFP4_TRITON_SWEEP=0 uv run --no-sync python - <<'PY' ...` original padded `Linear(60, 512)` repro: passed - `uv run --no-sync python -m pytest tests/gpu/torch/quantization/test_nvfp4_static_quantizer_cuda.py`: 11 passed - `uv run --no-sync python -m pytest tests/unit/torch/quantization/test_mse_calibrator.py`: 14 passed - `uv run --no-sync pre-commit run --files modelopt/torch/quantization/model_calib.py tests/gpu/torch/quantization/test_nvfp4_static_quantizer_cuda.py`: passed ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information Related issue: #1552 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Strengthened quantization calibration robustness with improved state management and explicit restoration mechanisms * Unified quantization logic across different calibration methods to eliminate code duplication and improve consistency * Refined block-quantization postprocessing behavior for better accuracy * **Tests** * Added comprehensive test coverage validating quantization calibration behavior with padded tensor dimensions and edge cases <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1557?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: eshieh <eshieh@nvidia.com> |
||
|
|
999c99913e |
Add Nemotron-3-Nano-30B-A3B-BF16 e2e tutorial: Prune + Distill + Quantize + Nemo Evaluator + vLLM deployment (#1376)
### What does this PR do? Type of change: example/tutorial <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Add Nemotron-3-Nano-30B-A3B-BF16 e2e tutorial: Prune + Distill + Quantize + Nemo Evaluator + vLLM deployment <img width="2079" height="1613" alt="image" src="https://github.com/user-attachments/assets/19b6ab82-7f01-45df-a0a5-d1c3282b384a" /> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * End-to-end Nemotron-3-Nano-30B tutorial: pruning, two‑phase distillation, FP8 PTQ, evaluation, and vLLM deployment; new ablations and long‑context analyses. * Distillation CLI: configurable seed and activation‑recomputation options. * **Bug Fixes** * Preprocessing hardened to skip malformed JSONL and normalize tool‑call argument formats. * **Documentation** * Many README/examples/evaluator docs updated (news list, tokenization guides, tutorials, configs, and deployment notes). * **Tests** * Added test verifying preprocessing handles stringified tool‑call arguments. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1376?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
5d0441ae3d |
Create shared Megatron calibration forward loop for prune / quantize with megatron pretraining data style sequence packing (#1501)
## Summary
Replaces the bespoke calibration loops in Megatron-LM and
Megatron-Bridge prune / quantize example scripts with a single shared
utility,
`modelopt.torch.utils.plugins.megatron_calibration.get_megatron_calibration_forward_loop`.
The shared loop iterates a packed calibration dataloader built via
`get_dataset_dataloader(pack=True)` and drives a logits-free prefill
pass through the model so activation hooks fire on every layer.
`pack=True` produces Megatron-LM pretraining-style **global-stream
packing**: all raw samples are concatenated into one EOS-separated token
stream and sliced into uniform-length rows. The trained model has seen
this distribution extensively during pretraining, so the activations
produced during calibration are representative of the model's natural
behavior.
Migrates four call sites:
- `examples/megatron_bridge/prune_minitron.py`
- `Megatron-LM/examples/post_training/modelopt/{prune,quantize}.py`
(separate PR:
[NVIDIA/Megatron-LM#4881](https://github.com/NVIDIA/Megatron-LM/pull/4881))
- `Megatron-Bridge/examples/quantization/quantize.py` (separate PR)
Each call site passes `pack=True` explicitly with an inline comment so
users see the option and know when to flip it. The function-level
defaults (`get_dataset_dataloader(pack=False)`,
`get_megatron_calibration_forward_loop(pack=False)`) remain
back-compat-safe.
Unified defaults across all four sites: `--calib-dataset
nemotron-post-training-dataset-v2`, `--calib-size 1024`,
`--calib-max-sequence-length 4096`, `--calib-batch-size 1`.
## Experimental results
Qwen3-8B on full 100% MMLU (n=14042; binomial 2σ noise floor ≈ ±0.78 pt
at acc ≈ 0.7), 0-shot, eval batch_size=4. Calibration on the default
workload: nemotron-post-training-dataset-v2, seq_length=4096,
calib_batch_size=8.
**Three calibration data shapes compared:**
- **Padded**: one doc per row, padded to `seq_length`, pad tokens flow
through the forward (legacy `get_calib_dataloader` pad+truncate
behavior).
- **Trimmed**: one doc per row, each row trimmed to its real content
length via `attention_mask`, with EOS forced at the last real position;
pad never enters the forward. This is no longer part of this PR.
- **Packed**: global-stream slicing — all docs concatenated
EOS-separated into one token stream, sliced into uniform `seq_length`
rows. Matches Megatron's `.bin`/`.idx` pretraining distribution. Enabled
via `pack=True` in `get_megatron_calibration_forward_loop`.
| Workload | Padded | Trimmed | **Packed** |
|---|---|---|---|
| M-LM NVFP4 quantize (`NVFP4_DEFAULT_CFG`) | 0.707 | 0.708 | 0.709 |
| M-Bridge Minitron prune (Qwen3-8B → 30L / 3584 / 11776 ≈ 6B params) |
0.576 | 0.573 | **0.589** |
### Key findings
- **M-LM quantize quality is calibration-mode-insensitive** for dense
Qwen3-8B NVFP4 — all three modes are nearly identical.
- **M-Bridge prune**: More sensitive to calibration data shape. Packed
wins over Padded on full MMLU.
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
48fbac0ed7 |
config for expert removal in nemotron3 (#1544)
### What does this PR do? New example This PR adds config and updates Nemotron descriptor to support expert pruning for Nano3 30B-A3B ### Usage ```python torchrun --nproc_per_node 2 examples/puzzletron/main.py --config examples/puzzletron/configs/nemotron-nano-30b-A3b-v3/nemotron_nano_v3_pruneexp.yaml 2>&1 | tee ./log.txt | grep "Puzzletron Progress" ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - Did you write any new necessary tests?: ❌ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for Nemotron Nano 30B model with multiple pruning and compression strategies * Enabled expert pruning to reduce model capacity, plus attention optimization and feed-forward layer compression * Introduced model validation and solution evaluation configurations for efficient model optimization workflows <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1544?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: mchochowski <mchochowski@nvidia.com> |
||
|
|
9d0d97829a |
chore(lint): modernize typing (PEP 604/585) and enable UP032 (#1537)
### What does this PR do? Type of change: chore / refactor (no behavior change) Two small lint-cleanup commits: **1. `chore(typing): modernize Union/Optional/List to PEP 604 / 585 syntax`** (8 files) - Replace `X = Union[A, B] # noqa: UP007` with `X: TypeAlias = A | B` for the six module-level type aliases (`ModelLike`, `Criterion`, `NodeTarget`, `CalibrationDataType`, `Hparam.Importance` / `ActiveSlice`). The `TypeAlias` annotation is required so mypy continues to treat them as aliases under PEP 604. - Modernize forward-ref unions in `modelopt/onnx/quantization/autotune/` to full-string forward refs (e.g. `"RegionPattern | None"`). - Update docstring type tags in `examples/puzzletron/evaluation/hf_deployable_anymodel.py`. **2. `chore(lint): remove UP032 ignore and convert .format() to f-strings`** (10 files) - Drop `UP032` from `extend-ignore` in `pyproject.toml`. - Auto-convert 19 `"...".format(...)` calls to f-strings across export plugins, examples, tests, and tools. One conversion in `modelopt/torch/utils/plugins/megatron_generate.py` was wrapped manually to stay under the 100-char limit. **Intentionally left as-is:** - `tools/launcher/slurm_config.py` keeps its `# ruff: noqa: UP045` — nemo_run's CLI parser can't introspect PEP 604 optional annotations. - `modelopt/torch/puzzletron/*` is **not** touched. The subtree disables ruff's `UP` family entirely (per-file-ignore `"UP"`) while migration is in progress, and converting `Optional[X]` to `X | None` there would silently break runtime introspection in `block_config._get_dataclass_type` that uses `get_origin(tp) is typing.Union` (PEP 604 unions return `types.UnionType` from `get_origin`, not `typing.Union`). Best revisited when puzzletron's lint carve-out is narrowed. - `UP038` (`isinstance(x, (int, float))` → `isinstance(x, int | float)`) — ruff has officially deprecated this rule; PEP 604 in isinstance is slightly slower and misleads readers about PEP 695 / `Optional`. Ignore kept. ### Usage No user-facing API changes. ### Testing - Pre-commit hooks (ruff check, ruff format, mypy, bandit, license) pass on both commits. - Ruff status against `main`: 37 unrelated pre-existing findings (W291/W293/E501/RUF005/PLR1704); zero new findings introduced by this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — Runtime behavior of the six type aliases changes from a `typing.Union` instance to `types.UnionType`. Downstream code introspecting via `get_origin(...) is typing.Union` on these aliases would break, but no in-repo caller does this on them. (The introspection in `modelopt/torch/puzzletron/block_config.py` operates on user-supplied dataclass field types, none of which are these aliases.) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (no behavior change) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ — internal style refactor; happy to add a Misc note if reviewers want one. - Did you get Claude approval on this PR?: ❌ — not yet. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Style** * Modernized type annotations across the codebase to use Python 3.10+ union syntax and TypeAlias where appropriate. * Standardized string formatting to f-strings, improving clarity of logs, errors, and validation messages. * **Chores** * Updated linting configuration to reflect modern typing/style rules. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1537?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
cd2ea46e1d |
[OMNIML-4730] Support quantized nn.Embedding (#1495)
### What does this PR do?
Type of change: new feature
Register `nn.Embedding` in `QuantModuleRegistry` so the embedding table
and lookup activations participate in quantization end-to-end:
- New `modelopt/torch/quantization/nn/modules/quant_embedding.py`
exposes `weight_quantizer` (embedding table), `output_quantizer` (lookup
activations, off by default), and an `input_quantizer` placeholder.
Embedding inputs are integer indices that cannot be fake-quantized, so
direct `enable()` / `enable_quant()` / `enable_calib()` calls on
`input_quantizer` raise; `forward()` never invokes `input_quantizer` at
all, so a back-door flip of `_disabled` is a no-op rather than a crash.
Wildcard configs (`*input_quantizer`) are accepted silently so the stock
deny-all → enable-wildcards → opt-out pattern in `NVFP4_DEFAULT_CFG` and
friends still works.
- `default_disabled_quantizers.yaml` installs `parent_class:
nn.Embedding, enable: false` so embedding quantization is opt-in and
existing model behavior is unchanged.
- `is_quantized_linear` in `core_utils.py` early-returns `False` for
`nn.Embedding` so AWQ / SmoothQuant / SVDQuant don't treat it as a GEMM
op.
- `_process_quantized_modules` in `unified_export_hf.py` routes
quantized `nn.Embedding` modules through `_export_quantized_weight`, so
the exported checkpoint contains the packed NVFP4 / FP8 / INT bytes plus
`weight_scale*` buffers, exactly like Linear layers.
### Usage
```python
import copy
import torch.nn as nn
import modelopt.torch.quantization as mtq
from modelopt.torch.export import export_hf_checkpoint
# Opt embeddings into the stock NVFP4 config — the YAML default is opt-out.
cfg = copy.deepcopy(mtq.NVFP4_DEFAULT_CFG)
cfg["quant_cfg"].append(
{
"parent_class": "nn.Embedding",
"quantizer_name": "*weight_quantizer",
"cfg": {"num_bits": (2, 1), "block_sizes": {-1: 16, "type": "dynamic", "scale_bits": (4, 3)}},
}
)
model = mtq.quantize(model, cfg, forward_loop)
export_hf_checkpoint(model, export_dir="./out")
# out/model.safetensors contains: embedding.weight (uint8, NVFP4-packed),
# embedding.weight_scale (FP8 E4M3 per-block), embedding.weight_scale_2 (FP32).
```
### Testing
- New unit tests `tests/unit/torch/quantization/test_quant_embedding.py`
cover: default quantizer state, no-quant identity, per-tensor and
per-row weight fake quant against the manual
`tensor_quant.fake_tensor_quant` reference, output quantizer activation,
locked-mutator raises (parametrized over `enable` / `enable_quant` /
`enable_calib`), and the wildcard-then-opt-out pattern. All 8 cases pass
under `CUDA_VISIBLE_DEVICES=""`. Export-path coverage (weight packing +
tied-weight guard) lives in
`tests/gpu/torch/export/test_export_embedding.py` because the export
path hits `torch.cuda.empty_cache()`.
- Verified end-to-end on an embedding-only model: `mtq.quantize` with
`NVFP4_DEFAULT_CFG` + the embedding opt-in produces `embedding.weight
(uint8)`, `embedding.weight_scale (float8_e4m3fn)`,
`embedding.weight_scale_2 (float32)` in the exported safetensors, with
`"quant_algo": "NVFP4"` in `hf_quant_config.json`.
- End-to-end smoke test on TinyLlama-1.1B-Chat-v1.0 via `hf_ptq.py`
(32-sample calibration on the default cnn_dailymail +
nemotron-post-training datasets), sweeping the embedding's
`weight_quantizer` across FP8 (per-tensor E4M3) and NVFP4 (E2M1,
block_size=16, FP8 scales) while varying the Linear-layer format. KV
cache is FP8 in all runs. Greedy generation (`max_new_tokens=100`) from
the standard CNN/DailyMail preview prompt:
> `<s>` LONDON, England (Reuters) -- Harry Potter star Daniel Radcliffe
gains access to a reported £20 million ($41.1 million) fortune as he
turns 18 on Monday, but he insists the money won't cast a spell on him.
[…] The Londoner has filmed a TV movie called "My Boy Jack," about
author Rudyard Kipling and
- **Baseline FP16 (no quantization):** *"his son, who died at the age of
18. He also has a role in the upcoming "The Prestige" and is set to star
in the remake of "The League of Extraordinary Gentlemen." "I'm not going
to be a star for the rest of my life," he said. "I'm going to be a star
for a while, but I'm not going to be a star for the rest of my life."*
- **FP8 Linear + FP8 embedding** (`hf_quant_config.json`: `"quant_algo":
"FP8"`, `exclude_modules: ["lm_head"]`, so `model.embed_tokens` is FP8):
*"his son, who died at the age of 18. He also has a role in the upcoming
"The Prestige" and is set to star in the upcoming "The Golden Compass."
"I'm not going to be a star forever," he said. "I'm going to be a star
for a while, but I'm not going to be a star forever."* (matches FP16
token-for-token until a single lexical drift)
- **FP8 Linear + NVFP4 embedding** (`model.embed_tokens: W4A16_NVFP4,
group_size=16`; Linear: `FP8`): *"his son, who died in the First World
War. He also has a role in the upcoming "The Golden Compass" and is set
to star in the BBC's "The Bill." "I'm not going to be a star for the
rest of my life," he said. "I'm going to be a star for a while, but I'm
not going to be a star for the rest of my life."* (embedding format swap
is visible — first phrase diverges from baseline; rest of the sentence
still tracks)
- **NVFP4 Linear + FP8 embedding** (`model.embed_tokens: FP8`; Linear:
`NVFP4, group_size=16`): *"his son, and is also in talks to play the
title role in a film version of the stage play "The Importance of Being
Earnest." "I've been offered a lot of things," he said. "I've been
offered a lot of things, but I've also been offered a lot of things that
I've said no to." Radcliffe's first movie, "The Tailor of Gloucester,"
was released in the UK in"* (expected drift at W4A4)
- **NVFP4 Linear + NVFP4 embedding** (`model.embed_tokens: W4A16_NVFP4,
group_size=16`; Linear: `NVFP4, group_size=16`): *"his son, who died in
a car accident. Radcliffe will also star in the film, which is being
shot in London. "It's a very different kind of film," he said. "It's a
very personal story, and I've been very lucky to be able to do it."
Radcliffe's next film, "The Girl with the Dragon Tattoo," is due to be
released in the U.S. In December. He will also"* (full-NVFP4 path;
quantized embedding row is `W4A16_NVFP4` because the embedding's input
quantizer is permanently disabled by design)
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ — embedding quantizers are
opt-in via `parent_class: nn.Embedding, enable: false` in
`default_disabled_quantizers.yaml`, so existing model behavior is
unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — will run `/claude review`
after the PR is up.
### Additional Information
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Opt-in quantization for embedding layers: weight quantization enabled
by default, optional output quantization; input quantization remains
permanently disabled.
* **Bug Fixes**
* Preserve tied embedding weights during export by skipping packing
(emits warning) to avoid breaking weight sharing.
* **Tests**
* Added unit tests covering embedding quantization behavior, export
packing, calibration, and tied-weight scenarios.
* **Documentation**
* Changelog updated with embedding quantization notes.
<!-- review_stack_entry_start -->
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1495?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)
<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
|
||
|
|
3ff15ccef3 |
Add support for postprocess exported model for block scale swizzling and support for different padding strategy (#1195)
### What does this PR do? Type of change: ? new feature <!-- Details about the change. --> Adds post-processing support for exported diffusion model checkpoints to enable NVFP4 block scale swizzling and configurable padding strategies. This allows exported quantized checkpoints to be directly consumed by inference runtimes (e.g., ComfyUI with comfy_kitchen) that require cuBLAS 2-D block-scaling-factors layout. Changes: 1) Unified post-processing step (_postprocess_safetensors): Loads saved safetensors files and applies merge, padding, swizzle, and quantization metadata injection in a single pass. 2) NVFP4 scale swizzle (swizzle_nvfp4_scales): Rearranges block scales from ModelOpt's flat [rows, cols // 16] layout to cuBLAS 2-D tiled layout per the cuBLAS specification. 3) Configurable padding (pad_nvfp4_weights): Pads NVFP4 weight and scale tensors to multiples of 16, with "row" (rows only) or "row_col" (both dimensions) strategies. 4) Standalone quantization metadata (build_layerwise_quant_metadata): Extracted from merge_diffusion_checkpoint so _quantization_metadata can be injected independently of merging — works for both merged (LTX-2) and standalone (Flux2) exports. 5) Bug fix (conversion.py): Wrapped yield in try/finally in set_quantizer_by_cfg_context so quantizer states are always restored, fixing an issue when yield fails. ### Usage ```python # LTX-2 export with merge + swizzle + padding export_hf_checkpoint( pipeline, export_dir="./output", merged_base_safetensor_path="./ltx-2-22b-dev.safetensors", enable_swizzle_layout=True, padding_strategy="row_col", enable_layerwise_quant_metadata=True, ) # Flux2 standalone export with swizzle + padding (no merge needed) export_hf_checkpoint( transformer, export_dir="./output", enable_swizzle_layout=True, padding_strategy="row_col", ) # Via quantize.py CLI python quantize.py \ --model ltx-2 --format fp4 \ --extra-param merged_base_safetensor_path=./ltx-2-22b-dev.safetensors \ --extra-param enable_swizzle_layout=true \ --extra-param padding_strategy=row_col \ --hf-ckpt-dir ./output ``` ### Testing 1) Exported LTX-2.3 NVFP4 with swizzle + padding + merged base checkpoint. Verified checkpoint has correct uint8 weights, float8_e4m3fn scales in swizzled layout, and _quantization_metadata . Ran the checkpoint with ComfyUI 2) Exported Flux2 NVFP4 with swizzle + padding. Verified checkpoint has correct uint8 weights, float8_e4m3fn scales in swizzled layout, and _quantization_metadata . Ran the checkpoint with ComfyUI ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Diffusers export: optional NVFP4 support — swizzle layout, row/row_col padding, and optional per-layer quantization metadata; exports are now post-processed to apply these options. * Export flow accepts new flags to enable swizzle, padding strategy, and layerwise metadata. * **Bug Fixes** * Quantizer context manager now always restores state, including on exceptions. * **Tests** * Added unit tests for NVFP4 padding, swizzling, metadata injection, and post-processing. * **Documentation** * README example updated to show swizzle and padding flags. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ynankani <ynankani@nvidia.com> Signed-off-by: YASH Nankani <ynankani@2u1g-x570-0073.ipp2a1.colossus.nvidia.com> Signed-off-by: ynankani-nv <ynankani@nvidia.com> Signed-off-by: YASH Nankani <ynankani@dl325g11-1979.ipp2a2.colossus.nvidia.com> Signed-off-by: YASH Nankani <ynankani@dl325g11-0771.ipp4a1.colossus.nvidia.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> Co-authored-by: YASH Nankani <ynankani@2u1g-x570-0073.ipp2a1.colossus.nvidia.com> Co-authored-by: YASH Nankani <ynankani@dl325g11-1979.ipp2a2.colossus.nvidia.com> Co-authored-by: YASH Nankani <ynankani@dl325g11-0771.ipp4a1.colossus.nvidia.com> |
||
|
|
b02e888550 |
fix(quantization): accept QuantizeAlgorithmConfig in get_modelike_from_algo_cfg (#201) (#1528)
Fixes #201 `get_modelike_from_algo_cfg` was typed to accept `QuantizeAlgoCfgType` (which includes `QuantizeAlgorithmConfig` instances), but the body only handled list / str / None / dict and raised `ValueError("Invalid config type")` for `QuantizeAlgorithmConfig` instances. This broke the documented pattern of passing typed config objects (`MaxCalibConfig`, `AWQLiteCalibConfig`, etc.) via `quant_config['algorithm']`. The fix pre-converts a `QuantizeAlgorithmConfig` instance to its `model_dump()` dict at function entry, so the existing dict path handles it unchanged. Two-line change; no new branch needed. Verification: added a unit test in `tests/unit/torch/quantization/test_mode.py` that builds a `MaxCalibConfig`, calls `get_modelike_from_algo_cfg` on both the object and its equivalent dict, and asserts the two paths produce the same tuple. Also covers the list-of-object path. Opened by OnCallBot on behalf of realAsma. Please review the diff before merging. cc Chenjie Luo (module owner for torch.quantization) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * `get_modelike_from_algo_cfg` now accepts configuration objects directly alongside existing input formats. * **Tests** * Added regression test to verify configuration object handling works correctly. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1528?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
c9098b63fb |
[4/n] Add vLLM integration for modelopt sparse attention (#1127)
### What does this PR do?
Type of change: New feature, new example, new tests, documentation.
Adds vLLM integration for ModelOpt sparse attention with paged KV cache
support.
This PR extends the ModelOpt Triton flash attention path so K/V can be
read directly from vLLM's paged KV cache through `block_table` lookup.
This avoids gather-to-contiguous copies when serving exported
sparse-attention checkpoints with vLLM.
The vLLM integration swaps vLLM's `FlashAttentionImpl` with
`ModelOptSparseAttentionImpl` after model load. The sparse configuration
is read from the exported checkpoint's `config.json`
`sparse_attention_config` block, written by
`examples/llm_sparsity/attention_sparsity/hf_sa.py`.
The restored checkpoint metadata supports:
- calibrated skip-softmax metadata (`threshold_scale_factor`,
`target_sparse_ratio`)
- N:M sparse-softmax metadata (`sparsity_n`, `sparsity_m`)
- dense token preservation metadata (`dense_sink_tokens`,
`dense_recent_tokens`)
The vLLM path uses ModelOpt Triton for sparse prefill launches.
Decode-only launches, cascade/prefix-cache metadata, and launches
without active sparse work delegate back to vLLM FlashAttention.
### Limitations
- Sparse attention is enabled for sparse prefill only.
- Decode-only launches currently fall back to vLLM FlashAttention.
- Attention sinks from vLLM FlashAttention are rejected until the
ModelOpt Triton path supports them.
- CUDA graph capture is not validated with this sparse attention path
yet; use `--enforce-eager`.
- Quant-only serving remains covered by `vllm_serve_fakequant.py`.
- Combined sparse attention + quantization serving is not handled by
this launcher in this PR and is planned as follow-up work.
### Usage
Export a checkpoint with calibrated skip-softmax and sparse24 metadata:
```bash
python examples/llm_sparsity/attention_sparsity/hf_sa.py \
--pyt_ckpt_path /path/to/hf-model \
--sparse_attn skip_softmax_calib_sparse24 \
--target_sparse_ratio 0.5 \
--calib_samples 64 \
--calib_max_seqlen 16384 \
--calib_chunk_size 4096 \
--seq_len 2048 \
--export_dir /path/to/modelopt-skipsoftmax-sparse24-export
```
Serve the exported checkpoint with the vLLM sparse-attention launcher:
```bash
PYTHONPATH=$PWD python examples/vllm_serve/vllm_serve_sparse_attn.py \
/path/to/modelopt-skipsoftmax-sparse24-export \
--tensor-parallel-size 8 \
--host 0.0.0.0 \
--port 8000 \
--trust-remote-code \
--enforce-eager
```
Send a request through the OpenAI-compatible endpoint:
```bash
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "/path/to/modelopt-skipsoftmax-sparse24-export",
"messages": [{"role": "user", "content": "Explain sparse attention in one paragraph."}],
"max_tokens": 128
}'
```
### Testing
GitHub CI on the latest commit is green:
- DCO
- code-quality
- docs build / deploy preview
- unit tests, including Linux, Windows, multi-version, partial-install,
and launcher jobs
- example tests
- GPU tests, including required GPU gate
- regression tests, including required regression gate
- `codecov/project`
Focused test coverage added/updated for this PR includes:
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_worker.py`
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_config.py`
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attention_conversion.py`
-
`tests/unit/torch/sparsity/attention_sparsity/test_triton_skip_softmax.py`
- `tests/gpu/torch/sparsity/attention_sparsity/test_vllm_plugin.py`
- `tests/gpu/torch/kernels/common/attention/test_triton_fa_paged.py`
-
`tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_skip_softmax.py`
-
`tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_sparse_nm.py`
-
`tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_calibrate.py`
Manual / NEL eval validation:
- Served a ModelOpt exported sparse-attention checkpoint through
`examples/vllm_serve/vllm_serve_sparse_attn.py`.
- Launched RULER64K NEL evals on DFW with `coreai_nvfm_llm`.
- Current partial RULER64K prediction scores, before final `results.yml`
is written:
- `skipsoftmax-only`: 98.59% over 4500 flushed samples
- `skipsoftmax-r0.7`: 99.70% over 1000 flushed samples
- `skipsoftmax-r0.9`: 99.70% over 1000 flushed samples
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - no new
PIP dependency.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A - documentation, examples, and tests are updated for this vLLM
integration path.
### Additional Information
Follow-up work:
- Validate and enable CUDA graph capture for the sparse vLLM path.
- Add combined sparse attention + quantization serving once the combined
path is tested.
- Investigate whether skip-softmax should also be enabled during decode.
---------
Signed-off-by: Kai Xu <kaix@nvidia.com>
|
||
|
|
e4dc0205d1 |
[OMNIML-4775] Move built-in PTQ quantization configs to YAML (#1423)
### What does this PR do? Type of change: refactor This PR moves the built-in PTQ quantization config definitions out of hard-coded Python dictionaries and into schema-backed YAML config files, and factors shared blocks into reusable composable snippets. - Adds reusable numeric config snippets under `modelopt_recipes/configs/numerics/`. - Adds YAML presets for the built-in model PTQ configs under `modelopt_recipes/configs/ptq/presets/model/`. - Adds YAML presets for KV-cache quantization configs under `modelopt_recipes/configs/ptq/presets/kv/`. - Adds YAML presets for the Diffusers-specific PTQ configs under `modelopt_recipes/configs/ptq/presets/diffusers/` and re-points `examples/diffusers/quantization/config.py` constants at them via `load_config`. - Adds reusable KV quantization units (`kv_fp8_affine`, `kv_nvfp4`, `kv_nvfp4_affine`, `kv_nvfp4_rotate`, `kv_*_cast` variants) under `modelopt_recipes/configs/ptq/units/`. - Adds reusable model-side units following the `component_numerics[_type]` convention: - `attention_qkv_fp8` — FP8 E4M3 on attention q/k/v bmm and softmax quantizers; shared by `model/` and `diffusers/` `nvfp4_fp8_mha` presets. - `block_sparse_moe_nvfp4` — NVFP4 W4A4 on `*block_sparse_moe*` weight/input quantizers; shared by `nvfp4_mlp_only`, `nvfp4_experts_only`, `nvfp4_omlp_only`. - `experts_nvfp4` — NVFP4 W4A4 on `*.experts.*` weight/input quantizers; shared by `nvfp4_mlp_only` and `nvfp4_experts_only`. - Switches the existing 5 NVFP4 presets (default + awq lite/clip/full + svdquant) and 4 mamba_moe presets to `$import` the existing `w4a4_nvfp4_nvfp4` / `w8a8_fp8_fp8` units instead of re-inlining the same weight+input quantizer pairs. - Moves the recently-added `W4A16_NVFP4_CFG` to YAML (`presets/model/w4a16_nvfp4.yaml`) composed from the existing `units/w4_nvfp4` snippet. - Updates `modelopt.torch.quantization.config` built-in config constants to load `QuantizeConfig` objects from YAML with `load_config(..., schema_type=QuantizeConfig).model_dump(exclude_unset=True)` via a new `_load_quantize_config_dict` helper; the constants remain plain `dict[str, Any]` for backwards compatibility with consumers that do mapping-style mutation (e.g. `entry["cfg"]` assignment). - Simplifies the cfg-list loader (`_load_quantizer_cfg_dict_list`) down to a 4-line list/single normalization now that the three call sites all load schema-typed YAMLs. - Adds/updates recipe loader coverage for built-in schema-backed config snippets. ### Latent-bug fixes surfaced by the refactor Two small correctness fixes are included alongside the mechanical refactor; flagging them explicitly: - **`examples/diffusers/quantization/quantize.py`** — adds an explicit `base_cfg = copy.deepcopy(base_cfg)` before applying runtime overrides. The existing `# Build a fresh config dict so we never mutate the global constants` comment had been aspirational only; in practice `reset_set_int8_config` accumulated `PercentileCalibrator` entries into `mtq.INT8_SMOOTHQUANT_CFG`/`INT8_DEFAULT_CONFIG` across repeated calls, and `set_quant_config_attr` added `trt_high_precision_dtype` keys into globally-shared cfg dicts. The deepcopy makes the code match the comment. - **`choices` set in `modelopt/torch/quantization/config.py`** — adds `MXFP6_DEFAULT_CFG` and `NVFP4_W4A4_WEIGHT_LOCAL_HESSIAN_CFG` to the documented public set of valid `mtq.*_CFG` names. Both constants exist on main but were missing from `choices`, so CLIs that gate on `mtq.config.choices` (e.g., `hf_ptq.py --qformat`) couldn't reach them even though the configs themselves were fully supported. ### Usage Existing Python imports continue to work: ```python import modelopt.torch.quantization as mtq cfg = mtq.FP8_DEFAULT_CFG model = mtq.quantize(model, cfg, forward_loop) ``` The built-in constants are plain `dict[str, Any]` (sparse — only explicitly-set fields are present), but their definitions now come from YAML snippets and presets composed through the existing `$import` system. Reusable YAML snippets can be composed through `$import`, for example: ```yaml # modelopt-schema: modelopt.torch.quantization.config.QuantizeConfig imports: base_disable_all: configs/ptq/units/base_disable_all w4a4_nvfp4_nvfp4: configs/ptq/units/w4a4_nvfp4_nvfp4 default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers algorithm: max quant_cfg: - $import: base_disable_all - $import: w4a4_nvfp4_nvfp4 - $import: default_disabled_quantizers ``` ### Testing Local checks run: - `nox -s "unit-3.10(torch_211, tf_latest)"` — 2329 passed, 12 skipped. - `nox -s pre_commit_all` — all hooks pass (ruff check / ruff format / mypy / YAML format / license / bandit / markdownlint). - YAML parse + `$import` resolution sanity check across all changed config files. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ Existing built-in Python config constants keep the same public names and dict semantics. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ Adds/updates recipe loader coverage for schema-backed built-in snippets. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ ### Additional Information This PR was previously stacked on #1405, which has since merged to `main`. The branch has been rebased onto `main` and no longer depends on any other open PR. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Many new quantization numeric configs and PTQ presets added (INT4/INT8/MXFP4/MXFP6/MXFP8/MXINT8/NVFP4), plus Diffusers, KV-cache (affine/cast/rotate) and MLP/MoE-targeted presets. * **Refactor** * Presets and shared snippets migrated to schema-backed YAML sources and centralized loading; INT8 percentile calibration avoids mutating shared base configs. * **Tests** * Tests now discover packaged config snippets at runtime and validate import/append behaviors. * **Documentation** * Presets README and numerous header descriptions updated. * **Chores** * Minor typing and script improvements. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1423?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
a5bc6f8123 |
Add DATASET_COMBOS for grouped calibration datasets (#1508)
## Summary - Add ``DATASET_COMBOS`` to ``modelopt.torch.utils.dataset_utils`` — single ``--dataset`` tokens that fan out to several entries in ``SUPPORTED_DATASET_CONFIG``. The per-entry ``num_samples`` is split evenly across the members inside ``get_dataset_dataloader``. - Two initial combos: - ``cnn_nemotron_v2_mix`` → ``cnn_dailymail`` + ``nemotron-post-training-dataset-v2``. Replaces the hardcoded two-element fallback list in ``hf_ptq.py`` when ``--dataset`` is omitted. - ``nemotron-post-training-v3`` → the seven ``nvidia/Nemotron-*`` SFT datasets registered in #1498 (mirroring the upstream [`nemotron-post-training-v3` collection](https://huggingface.co/collections/nvidia/nemotron-post-training-v3)). - ``get_supported_datasets()`` now appends combo names so they show up in ``--dataset`` help. - ``hf_ptq.py``'s default ``--calib_size`` bumped from ``512`` to ``1024`` so the ``cnn_nemotron_v2_mix`` combo's even split preserves the previous total sample count (was 512 per-dataset × 2 datasets = 1024; now 1024 split → 512 per-dataset × 2). ``--calib_size`` now denotes the total calibration budget regardless of combo cardinality. - Reject mixing a combo with one of its member datasets in the same ``--dataset`` list (e.g. ``cnn_dailymail,cnn_nemotron_v2_mix``) — combo would otherwise double-sample the explicit member with a smaller per-member quota. - Reject combo names in ``get_dataset_samples``; combos are dataloader-only. The error message points callers to ``get_dataset_dataloader``. - Validate ``DATASET_COMBOS`` at import time: empty member lists, name collisions with ``SUPPORTED_DATASET_CONFIG``, and references to unknown datasets raise ``ValueError`` up front. ## Test plan End-to-end validated against ``/hf-local/Qwen/Qwen3.5-0.8B`` via ``get_dataset_dataloader`` on the actual streamed data, plus 5 new unit tests in ``TestDatasetCombosExpansion`` (all 44 tests in ``test_dataset_utils.py`` pass with no regressions). - [x] ``python -c "from modelopt.torch.utils.dataset_utils import DATASET_COMBOS, get_supported_datasets; assert 'cnn_nemotron_v2_mix' in get_supported_datasets() and 'nemotron-post-training-v3' in get_supported_datasets()"`` - [x] ``hf_ptq.py`` with no ``--dataset`` flag still calibrates on cnn_dailymail + nemotron-post-training-dataset-v2 with the same total sample count as before. - [x] ``--dataset nemotron-post-training-v3 --calib_size 1024`` allocates 146 per member across the seven Nemotron datasets; full 1022-sample dataloader builds without error. - [x] ``--dataset cnn_dailymail,nemotron-post-training-v3 --calib_size 256,1024`` composes correctly: 256 from cnn_dailymail (as a plain entry) plus the 7-way split from the combo. (The earlier ``cnn_dailymail,cnn_nemotron_v2_mix`` example is rejected by design since ``cnn_dailymail`` is a member of that combo.) - [x] ``--dataset cnn_dailymail,cnn_nemotron_v2_mix`` raises ``ValueError`` with a clear message. - [x] ``get_dataset_samples("cnn_nemotron_v2_mix", ...)`` raises ``ValueError`` pointing to ``get_dataset_dataloader``. - [x] Unit tests: ``pytest tests/unit/torch/utils/test_dataset_utils.py`` — 44 passed. ## Post-validation fix End-to-end testing surfaced that the original ``nemotron-sft-agentic-v2`` entry kept the two splits (``interactive_agent``, ``tool_calling``) that pyarrow's streaming JSON reader cannot parse, and excluded ``search`` which is the only clean split. Failures reproduce deterministically across cache wipes with ``force_redownload``, so they are content-level defects in the published JSONL files at the pinned revision, not local artifacts: - ``interactive_agent`` — heterogeneous schema (``Column(.../member_id/type) changed from string to array``) at JSONL row 4. - ``tool_calling`` — malformed JSON row in a later shard, fails at sample ~885 with ``Missing a closing quotation mark in string``. - ``search`` — streams cleanly (verified to 2500 samples). Commit ``10f3cfd`` corrects ``nemotron-sft-agentic-v2`` to use only ``search``, with an updated comment. The CHANGELOG calls this out as a separate bullet so the behavior change on a previously-released dataset entry from #1498 is discoverable. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Dataset combo support: a single dataset token can expand into multiple registered datasets with even sample splitting; predefined combos added (e.g., cnn_nemotron_v2_mix, nemotron-post-training-v3) and listed as supported. * **Updates** * Default dataset when none specified now uses cnn_nemotron_v2_mix. * Calibration size default increased from 512 to 1024. * **Bug Fixes** * nemotron-sft-agentic-v2 now uses only the deterministic "search" split to avoid streaming JSON errors. * **Tests** * Added coverage for combo expansion, splitting, overlap validation, and rejection behavior. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1508?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
d7733562dd |
fix(prune): Minitron HybridModel + GPT-family fused-TE-spec import/export (#1518)
## Summary Split out of #1501 so the pack=True calibration packing change can land independently. This PR carries the pruning + export-side fixes. **Pruning bug fixes** - Register `HybridModel` (parent of `MambaModel` in modern Megatron-LM) under a new `HAS_HYBRID` flag so `mcore_minitron` actually prunes Nemotron-H et al. Previously `HybridModel` instances fell through `convert_to_dynamic`, got `freeze()`-ed (collapsing `hidden_size` / `num_layers` to a single choice), and produced unloadable saved checkpoints with mixed pruned/unpruned dims. - Replace the `isinstance(MambaModel)` gate in `_get_hybrid_pattern_key` with attribute-presence detection so both `MambaModel` (still using `hybrid_override_pattern`) and plain `HybridModel` (`hybrid_layer_pattern`) are handled uniformly. - Track `in_features` as a dynamic attribute on `_DynamicTEQKVLayerNormColumnParallelLinear` so TE's forward-time `inp_shape[-1] == in_features` assertion holds when `hidden_size` is pruned. - Dedupe MambaModel / HybridModel divisor dict into `_HYBRID_DIVISORS`. **Fused-TE-spec import/export for GPT-family** - Importer: prefer per-context keys (`fused_input_layernorm`, `fused_pre_mlp_layernorm`); fall back to legacy `fused_norm` for Nemotron-H back-compat. **Raise `KeyError`** when a fused-TE model has neither rule registered — the branch only fires when the model uses fused `TELayerNormColumnParallelLinear`, so a missing rule is unambiguously a plugin misconfig that would otherwise ship a chance-accuracy checkpoint. - Exporter: mirror the same fallback chain in `_get_fused_norm_weight` so GPT-family models round-trip cleanly back to HF. - Add the new rules to Qwen3, Qwen2.5, Llama, Llama4 (MoE-only, only `fused_input_layernorm`), DeepSeek, GptOss (MoE-only, only `fused_input_layernorm`) import and export mappings. - Preserve TE `_extra_state` from the existing module state dict (don't blank to `None`) at both call sites in the importer. **Misc** - `megatron_prefill`: `.contiguous()` on the logits slice before `broadcast_from_last_pipeline_stage` — broadcast asserts contiguity which fails when SP pads `seq_length` to a multiple of TP. - `megatron_mmlu`: accept `mmlu_dataset` kwarg so callers can point at a local copy of `cais/mmlu`. - `warn_rank_0`: auto-bump `stacklevel` by 1 inside the wrapper so callers' warnings point at user code, not at the wrapper frame. - `tools/launcher/examples/Qwen/Qwen3-8B/megatron_lm_ptq.yaml`: bump `mmlu_lower_bound` 0.68 → 0.75 (validated end-to-end with the fused-norm import fix). - CHANGELOG: bug-fix entry for the importer; date correction on the 0.44 entry. ## Consumer Megatron-LM PR https://github.com/NVIDIA/Megatron-LM/pull/4807 — `prune.py` / `mmlu.py` consume these APIs and currently ship inline WARs against released 0.44. Once 0.45 ships and the modelopt pin is bumped, those WARs collapse to one-liners. Related: #1501 (calibration packing). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Importer/exporter now correctly load fused LayerNorm weights for GPT-family models, preferring context-specific fused keys with a legacy fallback. * **New Features** * Hybrid Mamba/HybridModel support added for pruning/NAS workflows. * MMLU evaluation accepts a customizable dataset path (default: "cais/mmlu"). * **Improvements** * Extended export/import mappings and state handling across DeepSeek, GPT, Llama, Qwen; ensured last-stage logits are contiguous. * **Documentation** * Updated changelog entry and release date adjustment. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1518?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
f5650bd02c |
Schematize config loading and quantizer config entries (#1405)
### What does this PR do?
This PR makes ModelOpt config loading schema-aware and moves `quant_cfg`
entries from `TypedDict` validation to Pydantic validation.
Key changes:
- `load_config()` now returns validated schema instances when a schema
is provided through `schema_type=...` or declared with `#
modelopt-schema:`.
- Without a schema, it still returns the raw resolved dict/list.
- With a schema, it returns the validated schema object, such as
`QuantizeConfig` or `list[QuantizerCfgEntry]`.
- `QuantizerCfgEntry` is now a `ModeloptBaseConfig` Pydantic model.
- The “must specify `cfg`, `enable`, or both” rule is enforced wherever
entries are constructed.
- Enabled entries must provide non-empty `cfg` values when `cfg` is
present.
- Existing mapping-style access like `entry["cfg"]` and
`entry.get("enable")` continues to work.
- `RecipeMetadataConfig` is now a `ModeloptBaseConfig`, so recipe
metadata uses the same schema validation path.
- Recipe loading now delegates shape validation to `load_config()`
instead of manually checking loaded dicts.
- `normalize_quant_cfg_list()` now accepts:
- `Sequence[QuantizerCfgEntry]`
- `Sequence[Mapping[str, Any]]`
- legacy flat mapping configs, with deprecation warnings
- Public quantization config constants remain plain dict/list structures
for backward compatibility, even when they are built from
schema-validated YAML snippets.
### Behavior changes
- `load_config()` may now return a schema instance instead of a raw
`dict`/`list` when a schema is available.
Callers that checked `isinstance(result, dict)` should use
`isinstance(result, Mapping)` or check the expected schema type.
- Normalized `quant_cfg` entries are now `QuantizerCfgEntry` objects
internally.
Mapping-style access is preserved, but direct equality checks against
literal dicts should use `entry.model_dump()`.
### Testing
- `python -m pytest tests/unit/recipe/test_loader.py`
- `python -m pytest
tests/unit/torch/quantization/test_config_validation.py`
- `ruff check` / `ruff format --check` on touched files
- Commit-time pre-commit hooks passed
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
7038dec918 |
[1/2Refactor] speculative decoding: use mto config subsystem (#1328)
### What does this PR do?
Type of change: new feature
Port the speculative-decoding example to ModelOpt's recipe/config
subsystem: `model` / `data` / `training` / `<algo>` now load from a
single YAML with Pydantic validation and OmegaConf dotlist overrides.
Adds built-in `eagle3` / `dflash` recipes, drops the redundant
`training.mode` field (inferred from recipe class), and shrinks
`main.py` by ~145 lines (−208 / +63).
JIRA: OMNIML-3859
### Usage
```bash
python main.py --config general/speculative_decoding/eagle3 \
model.model_name_or_path=meta-llama/Llama-3.2-1B \
data.data_path=train.jsonl \
training.output_dir=ckpts/test
```
### Testing
- `pytest tests/unit/recipe/test_loader.py` — new coverage for Eagle /
DFlash YAML loading, dotlist overrides, and field-level validation.
- Smoke-trained both built-in `eagle3` and `dflash` recipes end-to-end.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ❌ — `main.py` CLI switched to
`--config <recipe>` (+ dotlist overrides); the old argparse flags are
removed.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
deps (`pydantic`, `omegaconf` already in core).
- Did you write any new necessary tests?: ✅ —
`tests/unit/recipe/test_loader.py`.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — to be added.
### Additional Information
Follow-up to the `modelopt.recipe` subsystem introduced for PTQ; this PR
extends the same declarative-YAML pattern to speculative decoding
(Eagle3 / DFlash / Medusa).
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added typed speculative-decoding recipe support for EAGLE, DFlash, and
Medusa; CLI dotlist overrides supported for single-file recipes.
* Trainer/config schema extended with speculative-training fields and
draft-vocab cache loading for Eagle.
* **Bug Fixes**
* Offline training no longer mutates model configs; loader enforces
required algorithm sections and prints recipe/config only on the primary
process.
* Reduced noisy per-rank logging by restricting status output to the
primary process.
* **Tests**
* Expanded tests for recipe loading, dotlist overrides, validation
strictness, and error cases.
* **Documentation**
* Recipe YAMLs updated with metadata and usage notes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
81c4fb25ab |
Add 7 nvidia/Nemotron-* calibration datasets to SUPPORTED_DATASET_CONFIG (#1498)
Registers nemotron-{sft-instruction-following-chat-v2, science-v1,
competitive-programming-v1, sft-agentic-v2, math-v2, sft-swe-v2,
sft-multilingual-v1} so hf_ptq.py's --dataset flag (which enumerates
get_supported_datasets() automatically) can select these for PTQ
calibration. Splits with heterogeneous parquet schemas that crash
streaming CastError mid-iteration are excluded per inline comments.
Adds a parametrized smoke test that skips when HF_TOKEN is unset since
the nvidia/Nemotron-* datasets are gated.
### What does this PR do?
Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->
<!-- Details about the change. -->
### Usage
```python
# Add a code snippet demonstrating how to use this
```
### Testing
<!-- Mention how have you tested your change if applicable. -->
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Standardized chat-style preprocessing (joining message contents) and
normalized chat key to "messages" across seven Nemotron dataset
variants, with explicit dataset paths and curated splits.
* **Tests**
* Added registry-shape unit checks for each new dataset entry.
* Added a gated integration smoke test that fetches sample data when a
Hugging Face Hub token is present.
<!-- review_stack_entry_start -->
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1498?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)
<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
701934bf03 |
fixes for fused moe (qwen3.6, GLM5.1 + MSE calibration (#1382)
### What does this PR do?
Type of change: Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
Fixes several issues with NVFP4 MSE calibration and export for fused MoE
expert modules (_QuantFusedExperts — used by Qwen3.6, GLM-5.1, and other
HF transformers 5.0+ models that store expert weights as 3-D
nn.Parameters).
- Bug 1 — MSE weight calibration runs 0 iterations for fused experts
(model_calib.py)
The weight-quantizer discovery loop in mse_calibrate used the singular
attribute name gate_up_proj_weight_quantizer to look up quantizers, but
_QuantFusedExperts stores them in a plural nn.ModuleList named
gate_up_proj_weight_quantizers. All 20,480 expert quantizers were
silently skipped, resulting in "MSE weight calibration: 0it" and no
MSE-optimized scales.
Fix: add a second pass that detects plural {param}_weight_quantizers
ModuleLists and enqueues each per-expert quantizer with a (param_name,
expert_idx) tuple; step 3 unpacks the tuple to extract the per-expert
weight slice.
- Bug 2 — Zero weight scales in exported checkpoint (nvfp4_tensor.py)
Per-block weight scales can silently underflow to 0 when cast to FP8
E4M3FN. The existing scale == 0 guard only catches exact float32 zeros;
values in (0, 2^-9) pass through and become 0 after the FP8 cast. This
affects both the dynamic recompute path (get_weights_scaling_factor) and
the static calibrated path (get_weights_scaling_factor_from_quantizer).
Fix: clamp per-block scales to 2^-9 (smallest positive FP8 E4M3FN
subnormal) before the FP8 cast in both paths.
- Bug 3 — Zero/corrupt amax for uncalibrated experts at export
(moe_utils.py)
Experts that receive no tokens during calibration have _amax = 0 or
uninitialized values. The existing scalar fallback used 1e-4 which
itself underflows to 0 in FP8 E4M3FN (1e-4 < 2^-9 ≈ 0.00195).
Additionally, the per-block fallback tensor had shape (H*W, 1) instead
of (H, W), causing a shape mismatch that silently bypassed the fallback
and fell through to the bad scalar. Finally, a stale zero global_amax
from an uncalibrated expert was not recomputed, causing division-by-zero
in the FP8 scale formula.
Fix: reshape the per-block fallback correctly; raise the clamp floor to
2e-3; always recompute global_amax from the current (possibly patched)
per-block _amax.
Additional fixes:
- moe_utils.py: safe CPU extraction of _amax before deepcopy to avoid
async CUDA errors from corrupt bfloat16 amax storage on under-calibrated
experts.
- model_quant.py: print_quant_summary now calls os.makedirs(output_dir,
exist_ok=True) before writing .quant_summary.txt, preventing a
FileNotFoundError when the export directory doesn't exist yet.
- tensor_quantizer.py: change default format in _short_amax /
_short_tensor from ".4f" to ".2e" so small amax values (e.g. 3.5e-7)
display as 3.50e-07 instead of 0.0000.
- hf_ptq.py: strip leading pad tokens from the preview input and add
skip_special_tokens=True to input_decode, fixing degenerate pre/post-PTQ
output on models that use EOS as the pad token (e.g. Qwen3).
### Usage
```python
# Quantize Qwen3.6-35B-A3B (or any compatible fused-expert MoE) with the new recipe:
python examples/llm_ptq/hf_ptq.py \
--pyt_ckpt_path /path/to/Qwen3.6-35B-A3B \
--recipe modelopt_recipes/general/ptq/nvfp4_experts_only_mse.yaml \
--export_path /path/to/output \
--calib_size 512 --calib_seq 2048
```
### Testing
validated on Qwen3.6-35B-A3B (8× B200):
- 21,740 quantizers inserted; 20,480/20,480 MSE weight calibrations
completed (~11 min)
- 0 / 2,013,265,920 zero weight_scale entries in the exported checkpoint
(3 shards)
- Pre- and post-PTQ generation produce coherent, semantically consistent
output
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added a new NVFP4 quantization recipe for expert layers with MSE-based
calibration.
* **Bug Fixes**
* Fixed FP8 scale underflow handling to prevent zero scaling factors.
* Fixed output directory creation for quantization summaries.
* **Improvements**
* Enhanced preview input handling for language models by removing
padding tokens.
* Improved quantizer display precision for better readability.
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1382)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
|
||
|
|
a451a2baf8 |
Support NVFP4 W4A16 quantization (#1313)
### What does this PR do?
This PR supports NVFP4 W4A16 quantization.
### Usage
```python
python examples/llm_ptq/scripts/huggingface_example.sh \
--model Qwen/Qwen3-8B \
--quant w4a16_nvfp4 \
--calib 1 \
--kv_cache_quant none \
--tasks quant
```
### Testing
<!-- Mention how have you tested your change if applicable. -->
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
### Additional Information
NVFP4 W4A16 is currently supported on vLLM only. TensorRT-LLM and SGLang
do not support this format.
vLLM loads NVFP4 W4A16 checkpoints via its `CompressedTensorsW4A16Fp4`
scheme, which requires the checkpoint to be in `compressed-tensors`
format. The raw ModelOpt export from this PR is still not directly
loadable by vLLM. After running hf_ptq.py, the checkpoint must be
converted by:
1. Renames tensors: .weight (uint8) → .weight_packed, .weight_scale_2
(fp32) → .weight_global_scale (inverted: 1 / weight_scale_2)
2. Rewrites config.json: sets quant_method: "compressed-tensors",
format: "nvfp4-pack-quantized", quantization_status: "compressed", and
derives the ignore list
automatically from layers that lack a weight_packed tensor.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* NVFP4 W4A16 weight-only quantization added (FP4 weights,
group_size=16; BF16 activations; no calibration-forward pass).
Selectable via qformat/CLI and included in export flow.
* New --exclude_modules CLI option (and EXCLUDE_MODULES env support in
example script) to skip quantizing specific modules.
* **Documentation**
* Changelog entry describing NVFP4 W4A16 usage and vLLM deployment via
compressed-tensors conversion.
* **Tests**
* Added test coverage for NVFP4 W4A16 export.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Co-authored-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
229ba61877 |
[OMNIML-3050] Enable torch.compile on _get_log_softmax_dist (#1479)
### What does this PR do?
Type of change: new feature
Wraps `_get_log_softmax_dist`
(`modelopt/torch/quantization/algorithms.py`) — the distributed
log-softmax helper used by `AutoQuantizeKLDivSearcher` under TP > 1 —
with `@torch.compile(dynamic=True)`. Inductor fuses the `amax /
all_reduce(MAX) / logsumexp / all_reduce(SUM) / log+sub+cast` pipeline
into one kernel. `dynamic=True` avoids recompiles across the varying
`[batch, seq]` shapes seen during calibration, matching the existing
pattern in `backends/fp8_per_tensor_gemm.py`. Stale TODOs are removed;
the prior ONNX-Windows concern no longer applies because the function is
only reachable when a TP group is initialized (never on the Windows CPU
unit-test job), and `@torch.compile` is import-time safe.
### Usage
```python
# Internal — invoked automatically by AutoQuantize KL-divergence search under TP > 1:
import modelopt.torch.quantization as mtq
model, _ = mtq.auto_quantize(
model,
constraints={"effective_bits": 6.0},
quantization_formats=[mtq.INT4_AWQ_CFG, mtq.INT8_DEFAULT_CFG],
data_loader=calib_loader,
forward_step=lambda m, b: m(b),
method="kl_div",
)
```
### Testing
Verified locally against the affected unit-test paths:
- `pytest tests/unit/torch/quantization/test_autoquant.py -k kl_div` →
22 passed (covers the call chain into `_get_log_prob`).
- `pytest tests/unit/onnx/` → 516 passed. The single failure
(`test_autocast_quantize.py::test_autocast_quantize_int8[False-True]` —
onnxruntime `CopyTensorAsync is not implemented`) is a pre-existing
failure on `main`, unrelated to this change (confirmed by stashing the
diff and re-running).
- `pytest tests/unit/torch/quantization/test_autoquant.py
tests/unit/torch/quantization/test_quantize_cpu.py
tests/unit/torch/quantization/test_config_validation.py` → 159 passed.
- Functional sanity in a single-process gloo group: compiled output
matches `torch.log_softmax` reference for fp32 and fp16, with no
recompiles across varying `[batch, seq]` shapes.
- `pre-commit run --files modelopt/torch/quantization/algorithms.py` →
ruff / mypy / bandit all pass.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (\`git commit -s -S\`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded \`trust_remote_code=True\`, \`torch.load(...,
weights_only=False)\`, \`pickle\`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in \`CONTRIBUTING.md\`: N/A
- Did you write any new necessary tests?: N/A — existing \`kl_div\`
autoquant tests already exercise the call chain.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ — will run \`/claude
review\` after marking ready.
### Additional Information
Single-line behavior change: \`_get_log_softmax_dist\` is now
\`@torch.compile(dynamic=True)\`-decorated. No API surface changes.
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
|
||
|
|
94337ade75 |
fix(te-plugin): handle TE 2.15+ tuple return from _Linear / _GroupedLinear (#1481)
### What does this PR do?
Type of change: Bug fix
TE 2.15+ changed `_Linear.forward` and `_GroupedLinear.forward` to
return `(out, new_workspace)` tuples instead of a single tensor.
ModelOpt's patched `te_quantized_linear_fn` /
`te_grouped_quantized_linear_fn` still piped the whole tuple into
`self.output_quantizer`, crashing inside `TensorQuantizer.forward` on
`tuple.numel()`:
```
File ".../modelopt/torch/quantization/plugins/transformer_engine.py", line 184, in te_grouped_quantized_linear_fn
return self.output_quantizer(output)
File ".../tensor_quantizer.py", line 1037, in forward
if inputs.numel() == 0:
AttributeError: 'tuple' object has no attribute 'numel'
```
Mirror the existing pattern from `_QuantTELayerNormLinear.forward`: when
the underlying TE call returns a tuple, quantize only `output[0]` (the
activation tensor) and pass auxiliary workspace metadata through
unchanged. TE <= 2.14 returns a single tensor and falls through the
`isinstance` branch identically to before this change.
Already landed on `release/0.44.0` as commit `c897fbeaaf`; this brings
`main` in sync. Follow-up to
[#1473](https://github.com/NVIDIA/Model-Optimizer/pull/1473) (signature
introspection + `_forward` cache lookup), which fixed an earlier symptom
of the same TE 2.15 signature change but not this tuple-return path.
### Usage
No public API change. PTQ continues to work transparently across TE 2.x:
```python
import modelopt.torch.quantization as mtq
mtq.quantize(model, mtq.NVFP4_DEFAULT_CFG, forward_loop) # now works on TE 2.15.x
```
### Testing
<!-- Mention how have you tested your change if applicable. -->
Verified locally against **both TE 2.12** and **TE 2.15.0** using:
```bash
pytest tests/gpu_megatron/torch/quantization/plugins/test_transformer_engine.py
```
Without this fix on TE 2.15, the same test fails immediately with
`AttributeError: 'tuple' object has no attribute 'numel'`. With this
fix, both versions exercise the same code paths and pass — TE <= 2.14
skips the `isinstance(output, tuple)` branch and behaves identically to
before.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ <!--- Public API unchanged; TE
<= 2.14 path is identical (isinstance branch is false). -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!--- Existing
`test_transformer_engine.py` already exercises both paths; it would have
caught this on TE 2.15 had CI been running against that version. A
TE-version matrix is the right follow-up but is out of scope here. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
### Additional Information
<!-- E.g. related issue. -->
Triggered by Megatron-Bridge failing tests after their TE 2.15 bump. The
`release/0.44.0` cherry-pick was pushed directly (commit `c897fbeaaf`)
so Bridge could unblock; this PR carries the same fix forward to main.
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
62401e16df |
fix: layerwise calibration backward-compat, recipe split, batch-size guard (#1310)
## Summary Follow-up to #1251 (which renamed `use_sequential` → `layerwise`). Three related fixes bundled: 1. **Backward-compatible config loading.** PTQ checkpoints saved before #1251 store the legacy `use_sequential` key in the calibration-algorithm config, so loading them now raises `ValidationError: Extra inputs are not permitted (use_sequential)` because `QuantizeAlgorithmConfig` uses `extra='forbid'`. Accept `use_sequential` as an alias for `layerwise` via `AliasChoices`. The field still serializes as `layerwise`, so round-trips through the current schema are clean. 2. **Recipe split.** `nvfp4_experts_only-fp8_kv` previously enabled layerwise calibration by default, which changes the calibration flow materially. Split into two recipes: - `nvfp4_experts_only-fp8_kv.yaml` — default (no layerwise) - `nvfp4_experts_only-fp8_kv_layerwise.yaml` — layerwise variant 3. **`hf_ptq` batch-size guard.** Auto batch-size detection is not supported together with layerwise calibration. Default to `batch_size=1` when layerwise is enabled and the user hasn't set a batch size explicitly. Originally reported by Jenny Chen while resuming a PTQ checkpoint via `restore_sharded_modelopt_state`: ``` pydantic_core._pydantic_core.ValidationError: 1 validation error for MaxCalibConfig use_sequential Extra inputs are not permitted [type=extra_forbidden, input_value=False, input_type=bool] ``` ## Test plan - [x] `tests/unit/torch/quantization/test_config_validation.py` — legacy alias accepted, current name accepted, dump serializes under current name, `extra='forbid'` still rejects unknown keys. - [x] `pre-commit run` — clean. ### Before your PR is *Ready for review* - Is this change backward compatible?: ✅ (restores compatibility for pre-#1251 checkpoints) - New PIP dependency: N/A - New necessary tests: ✅ - Changelog update: N/A (bug fix) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added new PTQ recipe for efficient layerwise calibration of large models. * Automatic batch size optimization for layerwise calibration recipes. * Backward compatibility support for legacy input naming conventions. * **Documentation** * Updated recipe guides and changelog with new layerwise calibration recipe. * **Tests** * Added validation tests for configuration compatibility. [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1310) <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
5da4d636fc |
fix(te-plugin): make _Linear arg indexing robust to TE signature changes (#1473)
### What does this PR do?
Type of change: Bug fix
ModelOpt's `te_quantized_linear_fn` and `te_grouped_quantized_linear_fn`
read `weight` / `inp` from hard-coded positions in `args`. Two TE
signature changes broke this scheme:
- **TE 1.x → 2.0:** dropped the legacy `weight_fp8` slot between
`weight` and `inp`. ModelOpt handled this with an `if Version("2.0") <=
_TE_VERSION:` branch + a duplicate else branch.
- **TE 2.14 → 2.15:** inserted `weight_workspace` between `weight` and
`inp` at the `_Linear.forward` call site ([TE 2.15 linear.py
L1663](https://github.com/NVIDIA/TransformerEngine/blob/release_v2.15/transformer_engine/pytorch/module/linear.py#L1663)).
Unhandled by ModelOpt — `args[idx + 1]` resolved to `None` (workspace is
None outside FP8), which then crashed `TensorQuantizer.forward` on
`inputs.numel()` with `AttributeError: 'NoneType' object has no
attribute 'numel'`. Surfaced as a regression in Megatron-Bridge after
the TE 2.15 bump alongside ModelOpt 0.44.0rc3.
- **TE 2.10:** `_GroupedLinear.forward`'s second positional slot was
renamed `m_splits` → `non_tensor_args` (tuple wrapping). ModelOpt had a
separate `Version("2.10")` gate for this.
Replace all three version gates with **parameter-name introspection** of
the live `_Linear.forward` / `_GroupedLinear.forward` signature. The
parameter names (`weight`, `inp`, `m_splits`, `non_tensor_args`) have
been stable across TE 1.x, 2.x, and 2.15+; only their relative positions
shift. The new code reads the live signature via
`inspect.signature(...).parameters`, locates `weight`/`inp` by name, and
mutates only those positions in a list copy of `args` — everything
between (e.g. TE 2.15's `weight_workspace`) and after passes through
verbatim. The dual-branch code in `te_quantized_linear_fn` collapses to
a single path.
### Usage
No public API change. PTQ continues to work transparently across all
supported TE versions:
```python
import modelopt.torch.quantization as mtq
# Works on TE 1.x, 2.0-2.14, 2.15.x, and 2.16+ — no version flag needed.
mtq.quantize(model, mtq.NVFP4_DEFAULT_CFG, forward_loop)
```
### Testing
<!-- Mention how have you tested your change if applicable. -->
Existing TE plugin tests
(`tests/gpu_megatron/torch/quantization/plugins/test_transformer_engine.py`)
exercise both the `_forward` (no-grad calibration) and `_apply`
(grad-enabled training) paths of `te_quantized_linear_fn` for
`te.pytorch.Linear` — they would have caught the original TE 2.15
regression on a CI matrix entry pinned to TE 2.15. Verified trace
correctness across:
| TE version | `_Linear.forward` signature | `_te_linear` weight→inp gap
| `_GroupedLinear.forward` second slot |
|---|---|---|---|
| 1.x | `(ctx, weight, weight_fp8, inp, …)` | 1 | n/a |
| 2.0–2.14 | `(ctx, weight, inp, bias, …)` | 0 | `m_splits` |
| 2.15.x | `(ctx, weight, weight_workspace, inp, …)` | 1 |
`non_tensor_args` |
| 2.16+ (main) | `(ctx, weight, inp, bias, fwd_args)` | 0 |
`non_tensor_args` |
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ <!--- Public API unchanged;
broadens the range of TE versions that work (TE 2.15.x now supported, TE
1.x still supported via the same introspection path). -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!--- Only
adds a stdlib `inspect` import. -->
- Did you write any new necessary tests?: Existing tests sufficient
<!--- Bug fix is covered by existing `test_transformer_engine.py` for
whatever single TE version CI exercises. A multi-version TE matrix is
the right next step but is out of scope for this PR. -->
### Additional Information
<!-- E.g. related issue. -->
Triggered by Megatron-Bridge
https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/3783 failing tests
after bumping ModelOpt 0.44.0rc2 → 0.44.0rc3 together with a Megatron-LM
bump that pulls TE 2.15. ModelOpt rc2 had the same latent bug — it just
wasn't exercised until TE 2.15 became the runtime version.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Refactor**
* Improved Transformer Engine quantization plugin robustness by using
runtime parameter inspection instead of version-based branching,
ensuring compatibility across TE versions without requiring manual
updates.
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1473)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
d738995fe8 |
feat(opt): validate loaded modelopt state files (#1471)
Add validation to load_modelopt_state() to verify the loaded object is a dict with the expected schema (modelopt_state_dict list and modelopt_version str). Raises TypeError/ValueError with clear messages when the file is malformed, and detects full checkpoints passed by mistake, pointing users to mto.restore(). Closes #1041 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Added strict validation for model state files to surface format errors with clear messages. * Malformed or invalid state files now fail fast instead of being returned silently. * Improved detection to prevent accidental loading of full checkpoints when only state dicts are expected. * **Tests** * New unit tests covering validation and loading behavior for various malformed and valid state files. [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1471) <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com> Co-authored-by: Keval Morabia <kevalmorabia97@users.noreply.github.com> |
||
|
|
5508c327fb |
[Quantization] MSE-calibrate every per-expert weight in fused-experts MoE (#1421)
### What does this PR do?
Type of change: Bug fix
Two-part fix for transformers 5.x fused-experts containers (Qwen3-MoE /
Qwen3.5-MoE / Mixtral / DeepSeek / Kimi-K2.x ...) where weight
quantizers live in `nn.ModuleList`s (`gate_up_proj_weight_quantizers`,
`down_proj_weight_quantizers`):
1. **Per-expert weight iteration for calibration.** Add
`_QuantFusedExperts.iter_weights_for_calibration` that yields per-expert
`(weight_slice, quantizer)` pairs for both projections. The base impl
uses singular `*_weight_quantizer` and silently skips fused-experts
modules, so weight-only calibration paths never reached per-expert
quantizers.
2. **`mse_calibrate` refactor.**
- Add `_bootstrap_uncalibrated_weight_quantizers` after `max_calibrate`
to populate `_amax` on quantizers the forward pass didn't reach (dead
MoE experts that received no calibration tokens). Runs the existing
calibrator on the weight slice surfaced by
`iter_weights_for_calibration`.
- Replace the singular-only `weight_attr_names` discovery +
`getattr`-by-name walk with an `iter_weights_for_calibration` walk done
inside each parent module's `enable_weight_access_and_writeback`
context, so MSE processes every per-expert quantizer (active and dead)
and remains FSDP-safe.
Without this, the export-time fallback in `_export_fused_experts`
derived separate gate/up amaxes from each half of the fused weight,
breaking the `gate==up` `weight_scale_2` invariant on dead experts.
Also includes:
- `_sanitize_generation_config_for_save` in `unified_export_hf` —
coerces `do_sample=True` when an upstream `generation_config.json` has
`top_k`/`top_p` set, so newer transformers' strict validate doesn't
block `save_pretrained`.
- Small companion plumbing in `moe_utils.py`, `tensor_quantizer.py`, and
`core_utils.py` to support the per-expert iteration and bootstrap path.
### Usage
```python
import modelopt.torch.quantization as mtq
from modelopt.recipe import load_config
# Recipe `nvfp4_experts_only_mse-kv_fp8_cast` (already on main) now correctly
# MSE-calibrates every per-expert weight quantizer in fused-experts MoE models.
cfg = load_config("general/ptq/nvfp4_experts_only_mse-kv_fp8_cast")
mtq.quantize(model, cfg, forward_loop=calibration_forward_loop)
```
### Testing
**Original validation — Qwen3.5-122B-A10B with
`nvfp4_experts_only_mse-fp8_cast_kv`:**
- **Before:** 1/12288 (layer 38 expert 69) `gate \!= up`; 0 weights
MSE-calibrated.
- **After:** 0/12288 mismatches; 24576 weights MSE-calibrated; ~4.2 min.
**End-to-end pipeline validation — Qwen3.5-35B-A3B (40 layers × 256
experts × 2 projections = 20,480 per-expert weight quantizers), TRT-LLM
1.3.0rc13 + transformers 5.6 docker, single B200:**
| | Path A (4-sample calib, deliberately undercalibrated) | Path B (zero
forward-pass tokens) |
|---|---|---|
| Per-expert weight quantizers calibrated | 20,480 / 20,480 | 20,480 /
20,480 |
| Missing `_amax` | 0 | 0 |
| All-zero `_amax` | 0 | 0 |
| `mtq.quantize` time | 25–34 s | 23 s |
- **Cross-path diff:** every per-expert weight amax matches
**bit-for-bit** between the two paths (`n=20480 exact=20480 diff=0
max_rel=0`). With 8/256 experts routed per token and 4 calib samples,
almost all experts are "dead" in Path A. Bootstrap fills them from
`max(|weight|)`, MSE searches deterministically from there → identical
to Path B which bootstraps everything.
- **Export to HF NVFP4 checkpoint** succeeded (~95 s, 22 GB checkpoint).
Resulting `generation_config.json` has `do_sample: true` (upstream had
`top_k=20` + `top_p=0.95` which would have failed strict validate).
- **TRT-LLM inference loaded the checkpoint and generated text:** `"Born
in north-east France, Soyer trained as a"` → `" tailor. Demonstrating
his craft at a young age, at 20 he moved to Paris at the requests of the
noble people of Picardy."` (coherent grammar; factually wrong as
expected with 4-sample calib, but no NaN/Inf in logits, no
scale-mismatch crash). 92 GB GPU memory used.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ <\!-- relies on existing
recipe-level integration coverage; verified end-to-end on
Qwen3.5-122B-A10B and Qwen3.5-35B-A3B + TRT-LLM 1.3.0rc13 -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ <\!-- will run \`/claude
review\` -->
### Additional Information
Follow-up to PR #1407 (MSE+FP8-cast-KV recipes). The recipe YAML files
landed there; this PR fixes the calibration codepath so the MSE recipes
actually exercise per-expert weight quantizers in fused-experts MoE
containers.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Fixed generation configuration validation for HuggingFace model
exports.
* Improved handling of quantization shape mismatches during expert
weight export.
* **New Features**
* Enhanced calibration process with automatic population of missing
expert quantizers.
* Added grouped quantizer synchronization for improved multi-expert
quantization.
* **Tests**
* Added regression tests for fused expert export and calibration
correctness.
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1421)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
|
||
|
|
2ce745a92e |
Deprecate gradnas pruning and bert example (#1427)
### What does this PR do? Type of change: Deprecation of dead code <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Deprecation warning already added in 0.44 as per 1-release deprecation policy GradNAS only works for Bert and GPT-J and we dont actively maintain it or test it. Keeping it creates an expectation that it works plus it adds one more option for user to choose from. We already have much better pruning algorithms (Minitron and Puzzletron) for LLM pruning already hence removing GradNas. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ No but we dont have any users of this feature either <!--- If ❌, explain why. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Deprecation** * GradNAS pruning algorithm deprecated; related examples removed. * **Documentation** * Pruning and NAS guides and changelog updated to focus on Minitron and FastNAS; GradNAS references removed. * **Chores** * Chained-optimizations example and scripts removed. * Ownership mappings updated for README and examples; license insertion now applies to a previously excluded example file. * **Tests** * Multiple unit tests and test utilities related to GradNAS/transformer NAS removed. [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1427) <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com> |