mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
168cd828c19cd8e150ed69aa6b6de6713a8a153c
741
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
168cd828c1 |
Add qwen3 moe experts only test (#1274)
## Summary - Add unit test for Qwen3 MoE HF export with `NVFP4_EXPERTS_ONLY_CFG` quantization config - Verifies that `hf_quant_config.json` correctly reports `quant_algo: NVFP4` and that non-expert modules (`self_attn`, `lm_head`) appear in `exclude_modules` while routed expert layers (`mlp.experts.*`) do not - Reference: https://huggingface.co/nvidia/Qwen3.5-397B-A17B-NVFP4/blob/main/hf_quant_config.json Type of change: New tests ### Known issue On `transformers>=5.0`, fused MoE experts (`_QuantFusedExperts`) are not recognized by `get_quant_config`, causing `quant_algo=None` in the exported config. This test currently **fails** on transformers 5.x and is intended to be fixed by a follow-up change. ## Testing - **transformers 4.57.6**: PASSED - **transformers 5.5.4**: FAILED (`quant_algo` is `None` due to fused expert export gap) ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added GPU test coverage for exporting Qwen3 Mixture-of-Experts models with NVFP4 quantization. * Verifies the exported checkpoint records the NVFP4 quantization algorithm and that module exclusion patterns correctly exclude attention and LM head components while not excluding routed expert paths. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
3ad4f4f093 |
[Fix] Re-expand target_input on OOM in get_max_batch_size (#1374)
## Summary - `get_max_batch_size` halved `target_data_batch` on `torch.cuda.OutOfMemoryError` but never rebuilt `target_input`, so each retry re-fed the same too-large tensor — the retry loop was effectively a no-op. - Refactor the expand logic into an `_expand_to(batch)` helper, rebuild `target_input` after halving, and call `torch.cuda.empty_cache()` between attempts. ## Test plan - [x] New unit test `test_get_max_batch_size_oom_retry_shrinks_input` mocks `torch.cuda.*` and asserts the second retry receives the halved tensor (shapes seen: `[1, 10, 5]`, regulated result `4`). - [x] `pytest tests/unit/torch/utils/test_dataset_utils.py` — 14/14 pass (skipping the network-only minipile test). - [x] `pre-commit` (ruff, mypy, bandit, license headers) clean on commit. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Enhanced GPU memory management during batch size detection. When out-of-memory errors occur during the initial probing phase, the system now properly adapts input tensors to smaller batch sizes and clears GPU cache before retry attempts, resulting in more reliable recovery and stable batch sizing across diverse hardware environments. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
bb08094ff1 |
Add Nemotron-Nano-9B-v2 → Pruned 7B e2e tutorial: Prune + Distill + Eval + Quantize + vLLM deployment (#1325)
## Summary End-to-end optimization walkthrough for Nemotron-Nano-9B-v2 showing how ModelOpt techniques stack: - **Pruning** — Minitron structured pruning 9B → 7B - **Distillation** — Megatron-Bridge knowledge distillation up to 80B tokens; near-parity with official 9B on MMLU Pro, GPQA, LCB, AIME, Math 500, IFEval, SciCode - **Evaluation** - using nemo-evaluator - **Quantization** — FP8 PTQ via \`hf_ptq.py\`; checkpoint deployable on vLLM/TRT-LLM/SGLang with no extra flags (quantization auto-detected from \`config.json\`) - **vLLM Throughput** — BF16 vs FP8 benchmark on single H100 <img width="2085" height="1740" alt="image" src="https://github.com/user-attachments/assets/8620a019-5c09-4a6b-a5d2-ca164aaa5d87" /> <img width="2085" height="810" alt="image" src="https://github.com/user-attachments/assets/742c8035-f1fb-4394-b11b-0c6c3ac4e843" /> ### Files changed - `examples/pruning/minitron/README.md` — index page for Minitron end-to-end tutorials - `examples/pruning/minitron/NVIDIA-Nemotron-Nano-9B-v2/README.md` — full repro doc with 6 sections: data prep, pruning, distillation, evaluation, FP8 quantization, vLLM benchmarking - `examples/pruning/minitron/NVIDIA-Nemotron-Nano-9B-v2/nemo_evaluator.yaml` — NeMo Evaluator config used for all benchmark numbers - `examples/pruning/puzzletron/README.md` — index page for Puzzletron distillation results - `examples/pruning/puzzletron/Llama-3.1-8B-Instruct.md` — Puzzletron distillation results (renamed from puzzletron.md) - `examples/pruning/README.md` — updated Results section with direct links to new locations - `examples/megatron_bridge/README.md` — updated results link to point to `examples/pruning/` - `examples/puzzletron/README.md` — updated distillation results link - `examples/dataset/MEGATRON_DATA_PREP.md` — tokenization commands for all datasets used in the data blend 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Documentation * **New end-to-end tutorial** for model optimization covering Minitron pruning, knowledge distillation, FP8 quantization, and vLLM deployment with reproducibility steps and benchmark results * **Dataset preparation guide** with ready-to-run tokenization templates for Nemotron HuggingFace datasets * **Evaluation configuration** and results documentation including ablation studies across multiple benchmarks * **Updated navigation** across pruning, distillation, and dataset examples to streamline user workflows <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
378366b8c9 |
[6034518] Remove return statement preventing remote auto tuning (#1361)
### What does this PR do?
Remove return statement from the code checking remote auto tuning config
arguments since that results in skipping adding the actual remote tuning
config to the trtexec cmd.
**Root cause**: The necessary flags do not get added to
`self._base_cmd.extend(trtexec_args)` when remote autotuning is enabled.
**Before fix**:
```
['trtexec', '--avgRuns=100', '--iterations=100', '--warmUp=50', '--stronglyTyped', \
'--saveEngine=engine.trt', '--timingCacheFile=trtexec_timing.cache', \
'--onnx=baseline.onnx']
```
**After fix**:
```
['trtexec', '--avgRuns=100', '--iterations=100', '--warmUp=50', '--stronglyTyped', \
'--saveEngine=engine.trt', '--timingCacheFile=trtexec_timing.cache', \
'--remoteAutoTuningConfig=$CONFIG', '--safe', '--skipInference', \
'--onnx=baseline.onnx']
```
Notice that the remote autotuning and related flags are now included in
the `trtexec` command.
**Related PR**: https://github.com/NVIDIA/Model-Optimizer/pull/1259
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Bug Fixes
* Fixed an issue where remote autotuning configuration arguments were
not being properly included in benchmark commands, ensuring all remote
autotuning settings are now correctly applied during execution.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: dmoodie <dmoodie@nvidia.com>
|
||
|
|
a492fa9a14 |
Ensure removal of temp files on error in ONNX INT4 quantization (#1359)
### What does this PR do? Type of change: Minor bug fix - Put quantization steps inside try-finally to ensure removal of temp files on error in ONNX INT4 quantization. - To avoid redundancy between awq_lite() and awq_clip() methods, created a utility _remove_augmented_onnx() for exception-handling based removal of augmented onnx file and its data file. ### Testing - Locally performed ONNX INT4 awq-lite and awq-clip quantization with Llama 1B model. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Improved reliability of the quantization pipeline by ensuring temporary conversion artifacts are always removed, making cleanup more robust. * Consolidated handling of external-data companions and added safer deletion behavior that logs failures instead of raising errors. * Ensured consistent session teardown and forced memory collection to reduce resource leakage and intermittent errors during model conversion. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: vipandya <vipandya@nvidia.com> |
||
|
|
c07ac215ef |
[Fix]: Relax Dflash Rregression Test Threshold fo 2GPUs (#1373)
### What does this PR do? Type of change: Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> The offline dflash regression test can be runned on 1 or 2 gpus. For 2 gpus, the total steps is half of 1 gpu. This PR relax the failing threshold for 2 gpu tests. ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
9bb917d57c |
[BUG6108338] Update windows documentation for onnxruntime quantization with Cuda13.x (#1368)
### What does this PR do? Type of change: ? documentation <!-- Details about the change. --> Update windows documentation for onnxruntime quantization with Cuda13.x ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!--- Mandatory --> - Did you write any new necessary tests?: N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated Windows installation guide with CUDA 13.x-specific setup instructions for GPU-accelerated dependencies, including CuPy and ONNX Runtime configuration with nightly builds. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: ynankani <ynankani@nvidia.com> |
||
|
|
077e29a48c |
[NVBug 6108145] Fix PTQ calibration and export for fused-experts MoE (Qwen3.5-MoE VLM) (#1340)
### What does this PR do?
Type of change: Bug fix
Fixes a 4-bug cascade that caused silent PTQ failure on Qwen3.5-MoE VLMs
(Qwen3.6-35B-A3B): calibration
appeared to succeed but produced token-salad at inference. Root cause:
HF's @use_experts_implementation
dispatches expert forward to torch._grouped_mm / torch.bmm, bypassing
the F.linear hook that captures
activations — so gate_up_proj_input_quantizer /
down_proj_input_quantizer never calibrated and no input_scale
tensors were emitted.
Changes:
- examples/llm_ptq/hf_ptq.py — force config._experts_implementation =
"eager" (recursing into text_config /
vision_config / …) so per-expert F.linear calls are visible to the
calibration hook.
- modelopt/torch/quantization/conversion.py — normalize plural
ModuleList quantizer names (weight_quantizers.N
→ weight_quantizer) before fnmatch, so wildcards like
*mlp.experts*weight_quantizer match fused-expert
quantizers.
- modelopt/torch/export/unified_export_hf.py — hoist the
_QuantFusedExperts export branch above the
get_quantization_format() gate so _export_fused_experts() runs even when
the top-level format query returns
QUANTIZATION_NONE (happens for experts-only recipes).
- modelopt_recipes/general/ptq/nvfp4_experts_only-fp8_kv.yaml —
layerwise: false (VLM nested layer structure
breaks the layerwise walker).
<!-- Details about the change. -->
### Usage
```python
python examples/llm_ptq/hf_ptq.py \
--pyt_ckpt_path Qwen/Qwen3.6-35B-A3B \
--qformat nvfp4 \
--kv_cache_qformat fp8 \
--calib_size 512 \
--export_path Qwen3.6-35B-A3B-NVFP4
```
### Testing
<!-- Mention how have you tested your change if applicable. -->
Testing
End-to-end PTQ → vLLM deploy → NEL eval on Qwen3.6-35B-A3B (256 experts
× 40 layers, 35B params):
Hook-call diagnostic: 0 → 6720 per-expert F.linear calls during
calibration after the fix; 0 → 30720
input_scale tensors emitted in the exported checkpoint.
FP8 fused-MoE path still produces gibberish — separate follow-up (vLLM
per-expert weight_scale handling).
* vLLM full-FP8: the FlashInfer TRTLLM Fp8MoE loader doesn't stack the
256 per-expert scalar weight_scale tensors
into a [num_experts] per-expert vector — it ends up applying one
expert's scale across all 256, so every
routed expert dequants with the wrong amplitude → coherent token stream
collapses into multilingual gibberish.
* SGLang full-FP8: qwen3_5.py::_make_packed_weight_loader rejects with
AssertionError: Unexpected scalar for
tuple shard load: loaded_shard_id=(0,1,2), split_sizes=[1,1,1] — its
packed-loader has no path for "N
independent per-tensor source scalars combining into one fused-shard
parameter," so the fused QKV (or
in_proj_qkvz) load is structurally refused and the model never finishes
loading.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Better fused-expert export flow, a plugin to force eager expert
execution during calibration/export, and a representative quantizer
discovery utility.
* **Bug Fixes**
* Reliable matching/discovery of per-expert indexed quantizers enabling
correct calibration and mixed-precision export; fixes for calibration in
nested decoder layouts.
* **Documentation**
* Clarified PTQ config guidance on layerwise calibration.
* **Tests**
* Added fused-experts calibration, export, and name-normalization tests.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
|
||
|
|
e5ce0ae83a |
[NVBug 6102977] Add _disable_use_cache context manager to fix PTQ AttributeError on custom configs (#1324)
### What does this PR do? Type of change: Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> - Summary: Running hf_ptq.py on stepfun-ai/Step-3.5-Flash (and any model whose custom HF config doesn't assign use_cache) crashed in get_max_batch_size() with AttributeError: 'Step3p5Config' object has no attribute 'use_cache' before calibration could start. - Extract the existing "disable KV cache during calibration" logic into a _disable_use_cache(model) context manager, apply it to both get_max_batch_size and _forward_loop. The CM sets config.use_cache = False unconditionally (not only when the attribute exists) and restores the prior value on exit if one was set. - Behavior unchanged for normal configs; the NemotronH hybrid-cache correctness guarantee from #1251 is preserved. ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> Step-3.5-Flash PTQ now passes get_max_batch_size ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Improved memory handling during model evaluation and calibration by consistently disabling KV cache for both single-batch probes and full dataloader runs, simplifying and stabilizing inference flow and ensuring cache state is managed reliably. * **Tests** * Added unit tests verifying cache-state handling across models with and without cache settings, including correct restoration behavior even when errors occur. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
8eec6d4459 |
Support EP mcore import for TE Spec and Fix mamba moe config (#1342)
### What does this PR do? Type of change: Bug fix - Enable EP (expert parallelism) import for HF to MCore when using TE Spec - Fix bug in mamba moe config which doesn't skip attention layers properly in MCore (Mcore uses different naming for attention layers than HF) - Add getter for Quant Config (used in MLM modelopt examples to get quant cfg fields) ### Usage ```python # In Megatron-LM/examples/post_training/modelopt MLM_EXTRA_ARGS="--export-default-te-spec --trust-remote-code --moe-router-dtype fp32" EP=4 HF_MODEL_CKPT=</path/to/hf> MLM_MODEL_SAVE=<save/path> ./convert.sh nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Corrected expert-slice assignment so each expert-parallel rank loads the proper expert slice. * Improved detection of pipeline-parallel layer indices in submodule names. * **Improvements** * Relaxed constraints between local and global expert counts for grouped-local-expert imports. * Added typed helpers for managing quantization configuration entries and expanded quantizer disable patterns. * Exporter now accepts an additional hybrid model type when available. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> |
||
|
|
6d330784a6 |
Add required keys to attention pruning config (#1360)
### What does this PR do? Type of change: ? Bug fix The config `examples/puzzletron/configs/llama-3_1-8B_pruneffn_memory/pruning/attn_pruning.yaml` didn't have required keys to use attention pruning in the example `examples/puzzletron/main.py` ### Usage ### Testing In `examples/puzzletron/configs/llama-3_1-8B_pruneffn_memory/Llama-3_1-8B.yaml` change `ffn_pruning` to `attn_pruning` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated pruning configuration for improved KV-head pruning support, including enhanced importance hook settings and attention output handling for memory optimization. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Grzegorz Karch <gkarch@nvidia.com> |
||
|
|
6e08b13fe6 |
Fix regex capture for Megatron KD PP layer renaming (#1355)
### What does this PR do? Type of change: Bug fix Previously the regex we had looked for a dot after the integer layer number, but it might not exist sometimes. ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved detection and handling of pipeline-parallel layer indices in model names to correctly support layer identifiers positioned at the end of submodule names, enhancing compatibility with various naming conventions in distillation workflows. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com> |
||
|
|
47a33db9b6 |
[NVBUG: 6103846] Fix nvfp4_awq export for uncalibrated MoE experts (#1354)
## Summary - NVBug: [6103846](https://nvbugspro.nvidia.com/bug/6103846) — `Qwen3-30B-A3B nvfp4_awq` quantization fails at export with `AssertionError: Modules have different quantization formats`. - Root cause: in `model_calib.awq_lite`, MoE experts that end up disabled (NaN in act/weight scales, or no search-pass tokens) get `max_calibrate`-d but no `pre_quant_scale`. `get_quantization_format` then returns `nvfp4` for those experts while siblings stay `nvfp4_awq`. `unified_export_hf.requantize_resmooth_fused_llm_layers` groups all 128 experts of each linear name (gate_proj/down_proj/up_proj) and calls `preprocess_linear_fusion(..., resmooth_only=True)`, which asserts uniform format → fires for any single mismatched expert. - Fix: unify the disabled-expert paths in the awq_lite postprocess loop so any expert with `is_enabled == False` (no cache hits, NaN scales, or no search-pass tokens) receives `max_calibrate` + a neutral all-ones `pre_quant_scale`, matching the existing behavior for `num_cache_steps == 0`. Emit a warning so users notice that calibration coverage is incomplete and accuracy may degrade. ## Test plan - [x] `pytest tests/unit/torch/quantization/test_calib.py -k 'awq'` → 5 passed - [x] End-to-end on `Qwen/Qwen3-30B-A3B` with `NVFP4_AWQ_LITE_CFG` and a small calib set that leaves many experts uncalibrated: - All 6144 gate_proj/up_proj/down_proj expert linears report `nvfp4_awq` (no mismatch) - `export_hf_checkpoint` succeeds with no `AssertionError` - The new "Forcing pre_quant_scale=1 ... may degrade accuracy" warning fires for each affected expert - [x] Re-run via `examples/llm_ptq/hf_ptq.py` with the bug-report CLI (cnn_dailymail, batch_size=8, calib_size=64 — scaled down from 512 to fit budget) on B200: - 36 "the second time did not forward data through ..experts.X.{gate,up,down}_proj" warnings — i.e. the exact bug-triggering condition from the original NVBug log naturally reproduces - 2058 "Forcing pre_quant_scale=1" warnings — fix path activates for uncalibrated/disabled experts - 0 `AssertionError`s — export completes - `Quantized model exported to: /tmp/test_plan_qwen3-30b-a3b-nvfp4_awq` and post-PTQ generation works --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
4f378ad275 |
Add release-cherry-pick Claude Code skill (#1352)
Adds `.claude/skills/release-cherry-pick/SKILL.md` — a Claude Code skill for cherry-picking labeled PRs to a release branch. Invoke with `/release-cherry-pick <version>`. See this PR created with the skill: https://github.com/NVIDIA/Model-Optimizer/pull/1350 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added automated release cherry-pick workflow to streamline selecting and applying multiple PRs into release branches. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> |
||
|
|
1306841989 |
Enable SonarQube Static Application Security Testing (SAST) (#1349)
Enable SonarQube as a Nvidia recommended and more comprehensive code scanning tools compared to Bandit we currently use in pre-commit hook (still left for now) Tested pipeline in internal gitlab and it works and results are uploaded in internal SonarQube website <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Added CI jobs to run SonarQube analysis and generate a vulnerability report, with scheduled and branch-triggered runs. * Configured scans to preserve full git history, use caching, and auto-cancel interruptible runs. * Added an ignore rule to exclude generated analysis artifacts from version control. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
fe3042b5bb |
fix incomplete mapping of safetensors in generated puzzletron checkpoint (#1330)
### What does this PR do? Type of change: ? Bug fix Fixes `https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/puzzletron/main.py` where multi-GPU run caused only part of the file `model.safetensors.index.json` to be written to disk. ### Usage does not apply ### Testing Follow [instructions, step 3](https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/puzzletron#compress-the-model) - run with `--nproc_per_node 2` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a public checkpoint-saving entry that consolidates distributed sharded model shards into a single filesystem checkpoint; retains direct saving for single-process runs. * **Refactor** * Validation/evaluation tooling now uses the consolidated checkpoint-saving flow when persisting realized model checkpoints during runs. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Grzegorz Karch <gkarch@nvidia.com> Signed-off-by: Grzegorz K. Karch <grzegorz-k-karch@users.noreply.github.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> Co-authored-by: CodeRabbit <noreply@coderabbit.ai> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
5b41ba4c34 |
chore: Move FP8 MHA quantization entry from 0.45 to 0.44 in CHANGELOG (#1351)
## Summary PR #1289 (FP8 MHA quantization for ViT) was merged to `main` after `0.44.0rc1` was tagged, so the rc1 wheel ships without the `nn.LayerNorm` registration that the example's `_FP8_MHA_OVERRIDE` now references — surfaced as nvbug 6114983 (`ValueError: parent_class 'nn.LayerNorm' not found in QuantModuleRegistry` when running `torch_quant_to_onnx.py --quantize_mode=fp8`). PR #1289 is labeled `cherry-pick-0.44.0` and will be cherry-picked to `release/0.44.0` for the next rc, so the feature ships in 0.44 — this PR moves the corresponding release-notes bullet from the `0.45 (Future)` section to `0.44 (2026-05-xx)` to match. Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> |
||
|
|
06860061c6 |
[chore]: weekly bump of uv.lock on main (2026-04-27) (#1347)
## Summary Automated weekly update of uv.lock file for nSpect Scanning: - `uv.lock` — upgraded all transitive dependencies to latest compatible versions <details> <summary>uv lock --upgrade output</summary> ``` Using CPython 3.12.13 interpreter at: /opt/hostedtoolcache/Python/3.12.13/x64/bin/python3 Resolved 180 packages in 7.12s Updated certifi v2026.2.25 -> v2026.4.22 Updated click v8.3.2 -> v8.3.3 Updated filelock v3.28.0 -> v3.29.0 Updated huggingface-hub v1.11.0 -> v1.12.0 Updated idna v3.11 -> v3.13 Updated onnx-ir v0.2.0 -> v0.2.1 Updated onnxscript v0.6.2 -> v0.7.0 Updated packaging v26.1 -> v26.2 Updated pathspec v1.0.4 -> v1.1.1 Updated pyarrow v23.0.1 -> v24.0.0 Updated pybind11 v3.0.3 -> v3.0.4 Updated pydantic v2.13.2 -> v2.13.3 Updated pydantic-core v2.46.2 -> v2.46.3 Updated pydantic-settings v2.13.1 -> v2.14.0 Updated transformers v5.5.4 -> v5.6.2 Updated typer v0.24.1 -> v0.25.0 Updated tzdata v2026.1 -> v2026.2 Updated uvicorn v0.44.0 -> v0.46.0 Updated wheel v0.46.3 -> v0.47.0 Updated xxhash v3.6.0 -> v3.7.0 ``` </details> Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> |
||
|
|
4b5549b663 |
[1/N] Polish evaluation skills and common skills based on an E2E workflow testing, vendor two Claude skills from NeMo Evaluator (#1239)
### What does this PR do? Type of change: Documentation / Skills Polishes evaluation and common skills based on end-to-end experience quantizing + deploying + evaluating LLMs. Vendors the two upstream Claude skills from NeMo Evaluator, splits shared credential setup into its own doc, and applies reviewer feedback. **Status:** ✅ Approved by @mxinO; all CI passing. ### Changes **New files** - `.claude/skills/launching-evals/` — vendored verbatim from [NVIDIA-NeMo/Evaluator](https://github.com/NVIDIA-NeMo/Evaluator) @ commit `8fa16b2` (latest). Covers run / check / debug / analyze flows for NEL evaluations. - `.claude/skills/accessing-mlflow/SKILL.md` — vendored verbatim from the same upstream. Queries MLflow runs via the `mlflow-mcp` MCP server. - `.claude/scripts/sync-upstream-skills.sh` — re-vendors the two skills above at a pinned SHA. Idempotent; re-applies our provenance frontmatter on each run. - `.claude/skills/common/credentials.md` — shared HF / NGC / Docker credential setup, referenced from slurm-setup.md. Generic (not NVIDIA-internal) — public NEL SLURM-executor users rely on the same NGC/HF setup. Includes a "check what's already set first" detection section so the agent skips already-configured credentials. **Updated files** - `.claude/skills/common/slurm-setup.md` — NGC credential block collapsed to a one-paragraph pointer at `credentials.md`. - `.claude/skills/common/remote-execution.md` — reframed "Checkpoint and storage availability" as "Staging checkpoints from your workstation". Drops the misleading login-vs-compute framing and the dlcluster-specific row. - `.claude/skills/common/workspace-management.md` — drops stale pointer to the deleted e2e doc. - `.claude/skills/evaluation/SKILL.md` — workspace-integration intro trimmed; NEL CI section removed (content moved to Model-Optimizer-Internal). Monitoring block replaced via #1252 merge with joint trim text that routes to `monitor`, `launching-evals`, and `accessing-mlflow`. - `.claude/skills/ptq/SKILL.md` — "Next steps" block refined. - `.markdownlint-cli2.yaml` — excludes vendored upstream skills from markdownlint so they stay byte-identical to upstream. **Deleted files** - `.claude/skills/common/end-to-end-workflow.md` — per @kaix-nv and @mxinO: redundant with skill descriptions that already handle cross-skill routing. - `.claude/skills/evaluation/references/nel-ci-guide.md` — per @shengliangxu: NVIDIA-internal. Moved to `Model-Optimizer-Internal:agent/evaluation_guide.md` (MR !57). ### Review status | Reviewer | Concern | How addressed | |---|---|---| | @shengliangxu | `nel-ci-guide.md` contains NVIDIA-internal content | Moved to Model-Optimizer-Internal MR !57 (renamed `evaluation_guide.md` since it covers both NEL SLURM executor and NEL CI). Refreshed against current upstream NEL / NEL-CI (current cluster list, Sybil/non-Sybil distinction, prerequisites checklist, `nel-ci-cli` preferred trigger). | | @kaix-nv, @mxinO | e2e workflow doc unnecessary | Deleted. Skill descriptions already route between ptq / deployment / evaluation. | | @kaix-nv (#1252) | Overlap on `evaluation/SKILL.md` Monitoring section | Coordinated via comment on #1252. @kaix-nv incorporated the joint trim text referencing `monitor` + `launching-evals` + `accessing-mlflow`. Landed via #1252 merge. | | @mxinO | `remote-execution.md` compute-node framing is misleading | Reframed as workstation→cluster staging. | | @mxinO, CodeRabbit, Copilot | slurm-setup NGC creds aren't SLURM-specific; `$oauthtoken` literal clarification; overwrite safety | New `credentials.md` with NGC / HF / Docker setup. Append-if-missing pattern (`grep -q … \|\| echo … >>`). Explicitly calls out `$oauthtoken` as literal, kept unexpanded via single quotes. | | @mxinO | `credentials.md`: check what's already set first; document `hf auth login` | Added detection section at the top covering `HF_TOKEN`, `~/.cache/huggingface/token`, Docker config, enroot creds. HF section now documents `hf auth login` as the recommended interactive path; env-var path kept as option 2 for scripts/CI. | | @kevalmorabia97 | Internal `gitlab-master.nvidia.com` container URL in `launching-evals/SKILL.md` | Auto-fixed by bumping pinned SHA from `01899f8` to `8fa16b2` — upstream PR [NVIDIA-NeMo/Evaluator#920](https://github.com/NVIDIA-NeMo/Evaluator/pull/920) already replaced it with the public `nvcr.io/nvidia/eval-factory/simple-evals:26.03`. SHA bump also picked up upstream's new `nemo-evaluator-launcher resume` command and a tighter "MANDATORY monitor after every `nel run`" directive. | | @kevalmorabia97 | Internal `PPP` terminology + `/lustre/fsw/portfolios/coreai/...` path | Still in upstream — vendoring verbatim means we can't sanitize locally without breaking the "verbatim" property. Filed upstream as [NVIDIA-NeMo/Evaluator#938](https://github.com/NVIDIA-NeMo/Evaluator/issues/938) to genericize. Next sync-script SHA bump will pick up the fix automatically. | ### Related - **Depends on:** #1236 (`deployment/references/unsupported-models.md`, merged) - **Coordinated with:** #1252 (monitor skill, merged — joint trim text incorporated) - **Internal counterpart:** [Model-Optimizer-Internal MR !57](https://gitlab-master.nvidia.com/omniml/Model-Optimizer-Internal/-/merge_requests/57) — `agent/evaluation_guide.md` - **Upstream coordination:** vendored skills synced from [NVIDIA-NeMo/Evaluator](https://github.com/NVIDIA-NeMo/Evaluator) @ `8fa16b2`. Follow-up issue: [NVIDIA-NeMo/Evaluator#938](https://github.com/NVIDIA-NeMo/Evaluator/issues/938). ### Motivation Learnings from running end-to-end PTQ → Deploy → Eval on Devstral-Small-2-24B (FP8 VLM → NVFP4 MLP-only) on dlcluster B100, plus prior NEL CI experience on oci-hsg. ### Testing Validated end-to-end: PTQ (6 min) → vLLM deployment (3 debug iterations) → NEL evaluation (MMLU 77.4%, GSM8K 80%, GPQA 40% on `limit_samples=10`). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (documentation / skills only) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - `.claude/skills/launching-evals/` and `.claude/skills/accessing-mlflow/` are vendored verbatim from [NVIDIA-NeMo/Evaluator](https://github.com/NVIDIA-NeMo/Evaluator) (Apache-2.0). Provenance SHA pinned in each `SKILL.md` frontmatter and in `.claude/scripts/sync-upstream-skills.sh`. - Did you write any new necessary tests?: N/A (skill documentation) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> |
||
|
|
1ec931c2c7 |
[2/3][Feat]: Offline DFlash training (#1343)
### What does this PR do? Type of change: new feature Part 2 of a 3-PR series splitting #1271: - **[1/3] #1296**: File reorg + deprecate `ParallelDraft` - **[2/3] this PR**: Offline DFlash training (depends on #1296) - **[3/3] #1297**: Extract `HFSpecDecMixin` Changes: - Add `dflash_offline` flag to `DFlashConfig` for training from pre-computed hidden states; deletes base model layers to save memory. - Add Pydantic validators on `DFlashConfig`: - `_derive_dflash_offline` — auto-derive `dflash_offline` from `data_args.offline_data_path` in validation context. Not user-configurable: any user-supplied value is overridden by the derived value. - `_resolve_mask_token_id` — auto-detect `dflash_mask_token_id` from `tokenizer.mask_token_id`. - `_check_mask_token_id` — fail fast if unset after resolution. - `HFDFlashModel.modify()`: select `num_orig_hidden_layers` when offline; pick `_base_model_lm_head` device when no base layers present; drop base-model `layers` module. - `HFDFlashModel.forward()`: add offline branch — consumes precomputed `base_model_outputs` via `DFlashBaseModelOutput.from_offline_dict`, and when `dflash_self_logit_distillation` is enabled with `base_model_logits` absent, recomputes logits from `base_model_hidden_states` via `_base_model_lm_head`. Raises a clear error from the non-training / `pseudo_speculative_generate` paths when `dflash_offline=True`, since base-model layers have been deleted. - `DFlashBaseModelOutput` dataclass in `modeling_dflash.py` (with `from_offline_dict` classmethod) to unify online/offline output shapes. `aux_hidden_states` is required in `from_offline_dict` so missing keys fail fast at the entry point rather than deeper in the forward. - `examples/speculative_decoding/main.py`: replace inline `mask_token_id` auto-detect with `DFlashConfig.model_validate(dflash_cfg, context={"tokenizer": tokenizer, "data_args": data_args})`. ### Silent bug fix — `add_generation_template` → `add_generation_prompt` The pre-refactor `compute_hidden_states_hf.py` passed `add_generation_template=False` to `tokenizer.apply_chat_template`. This kwarg does not exist on HF `apply_chat_template` and was being silently ignored, so the intended "don't append a generation prompt" behavior was never actually applied. The new `tokenize_with_loss_mask` helper in `examples/speculative_decoding/collect_hidden_states/common.py` uses the correct `add_generation_prompt=False`. **This is a real behavior change** for anyone re-dumping hidden states: trailing generation prompts that were previously appended to the tokenized sequences will no longer be included. ### Testing - New tests: - `tests/unit/torch/speculative/plugins/test_hf_dflash_offline.py` — CPU unit tests for convert path (online keeps base layers, offline deletes them; `num_orig_hidden_layers` drives `target_layer_ids` in offline mode) and `DFlashConfig._derive_dflash_offline` validator. - `TestDFlashOfflineForwardGPU` in `tests/gpu/torch/speculative/plugins/test_hf_dflash.py` — GPU forward smoke with precomputed `base_model_outputs`, plus the `dflash_self_logit_distillation` logit-recompute path. - training test: <img width="454" height="317" alt="image" src="https://github.com/user-attachments/assets/79b92790-4d15-4313-bb9b-f35665b012e6" /> <img width="456" height="310" alt="image" src="https://github.com/user-attachments/assets/4558559f-9c35-49ed-b36e-82fbc99eab23" /> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ — additive `dflash_offline` flag defaulting to `False`; validators fall through when context not provided. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — see Testing section above. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ ### TODO (follow-up) - [x] Update `examples/speculative_decoding/collect_hidden_states/compute_hidden_states_*.py` to support DFlash offline data. Current scripts are Eagle-specific — they hardcode the `[2, N/2, N-3]` aux-layer selection and emit `{input_ids, hidden_states, aux_hidden_states}`. DFlash offline needs: - Aux layer indices driven by `build_target_layer_ids(num_orig_hidden_layers, num_draft_layers)` (or a configurable list), not the Eagle triplet. - `base_model_hidden_states` key (last-layer hidden) so `DFlashBaseModelOutput.from_offline_dict` + the `dflash_self_logit_distillation` recompute path can consume it. - Optional `base_model_logits` dump so offline training can skip the self-distillation logit recomputation when logits are available. ### Additional Information Base branch is #1296 (file reorg). Retarget to `main` once #1296 merges. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Offline DFlash speculative-decoding training from precomputed base-model hidden states * Answer-only-loss training with persisted loss masks and optional chat-template support * Flexible auxiliary-layer selection via CLI and an exposed default aux-layer helper * Auto-derived offline flag in config and automatic memory optimization during offline conversion * **Documentation** * Updated guides for offline pipeline, aux-layer selection, and loss-masking options * **Tests** * New unit, GPU, and regression tests covering offline conversion, training, and config derivation <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
7c80d85751 |
[1/3][Refactor]: File reorg; deprecate ParallelDraft (#1296)
### What does this PR do? Type of change: refactoring Part 1 of a 3-PR series splitting #1271: - **[1/3] this PR**: File reorg + deprecate `ParallelDraft` - **[2/3] #1295**: Offline DFlash training - **[3/3] #1297**: Extract `HFSpecDecMixin` Changes: - **File reorg**: `transformers.py` → `hf_eagle.py`; extract `HFMedusaModel` → `hf_medusa.py`; extract `EagleModule` / `EagleBaseModelOutput` → `modeling_eagle.py`; extract `DFlashModule` / `DFlashAttention` / `DFlashDecoderLayer` / `build_target_layer_ids` / `apply_rotary_pos_emb` → `modeling_dflash.py`. - **Deprecate `ParallelDraft`**: remove `parallel_draft_step`, `parallel_draft_heads_num_layers`, and the `ParallelDraft` module from HF Eagle; remove the `EagleMedusaExporter` branch from `HFEagleModel.get_exporter()` (the `EagleMedusaExporter` class itself still lives in `hf_spec_export.py` for Megatron parity). - **Rename**: `_draft_model_config` → `eagle_config` in export plugin. - Update imports in `examples/speculative_decoding/` and `modelopt/torch/speculative/utils.py` to follow the module rename. ### Testing Validated with existing Eagle and DFlash training scripts (re-run after `9ae5302729 revert behavior change`). ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ — renames `modelopt.torch.speculative.plugins.transformers` → `.hf_eagle`; removes `parallel_draft_step` / `parallel_draft_heads_num_layers` from Eagle config; renames `_draft_model_config` → `eagle_config` in export plugin. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — pure refactor; existing tests updated for the rename. `test_hf_spec_rope_export.py` assertions were also corrected to reflect the actual production path (the old assertions were masked by `MagicMock` not invoking the `_draft_model_config` `@property`). - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ ### Additional Information Breaking changes: - `modelopt.torch.speculative.plugins.transformers` → `.hf_eagle` - `parallel_draft_step` / `parallel_draft_heads_num_layers` removed from Eagle config - `_draft_model_config` → `eagle_config` in export plugin <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactoring** * Reorganized speculative-decoding plugins into focused modules, converting the legacy "transformers" entry into a deprecated shim that re-exports the new plugin surface. * Consolidated DFlash implementation into a shared modeling component and introduced a dedicated EAGLE decoder module. * **New Features** * Added a Medusa speculative-decoding plugin with configurable heads and combined-loss training behavior. * **Chores** * Updated pre-commit license-hook exclusion and feature-flag wiring. * **Tests** * Updated export tests to expect rope-scaling fallback semantics. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
946639aa19 |
Fix PTQ for VLMs with image calibration (#1318)
### What does this PR do? This PR fixes PTQ with image claibration for VLMs. ### Usage ```python python3 examples/llm_ptq/hf_ptq.py --pyt_ckpt_path Qwen/Qwen3-VL-8B-Instruct --qformat fp8 --export_path Qwen3-VL-8B-Instruct-fp8 --trust_remote_code --kv_cache_qformat none --calib_with_images --calib_size 512 ``` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Image-text calibration now extends support to additional model architectures when image calibration is enabled. * Improved tokenizer truncation handling in multimodal dataset processing to prevent configuration conflicts when image inputs are present. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Liana Mikaelyan <lmikaelyan@nvidia.com> |
||
|
|
5887410a6d |
fix: EAGLE mix_hidden_states in-place op crash (#1088) (#1104)
### Type of change - [x] Bug fix (non-breaking change which fixes an issue) ### Description Fixes #1088 — `RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation: IndexPutBackward0` when training with `eagle_mix_hidden_states=True`. **Root cause:** In `HFEagleModel._eagle_training_forward`, the indexed assignment at line 991–994 modifies `eagle_input_hiddens` in-place while it is still part of the autograd computation graph. **Fix:** Clone the tensor before the in-place assignment. This is the same pattern already used in the Megatron backend at `megatron_eagle.py:1201-1202`: ```python # Clone to avoid inplace modification of view created in no_grad mode eagle_module_input_hidden_states = eagle_module_input_hidden_states.clone() ``` The HF backend was missing this clone. ### Usage ```python config["eagle_mix_hidden_states"] = True config["eagle_ttt_steps"] = 2 mtsp.convert(model, mode=[("eagle", config)]) model.train() outputs = model(input_ids=input_ids, labels=labels) outputs.loss.backward() # no longer crashes ``` ### Testing Added `test_eagle_mix_hidden_states_backward` parametrized over `eagle_ttt_steps` [1, 2] that: - Converts a tiny LLaMA to EAGLE with `eagle_mix_hidden_states=True` - Runs forward + backward pass - Asserts loss is not None and gradients flow to `eagle_module` ``` pytest tests/unit/torch/speculative/plugins/test_hf_speculative.py::test_eagle_mix_hidden_states_backward -v ``` ### Checklist - [x] I have read the [contributor guidelines](CONTRIBUTING.md) and signed my commits - [x] I have followed the [security best practices](SECURITY.md) - [x] This change is backward compatible - [x] I have followed third-party code and dependency guidelines - [x] I have added tests that prove my fix is effective <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed gradient computation issue in speculative decoding during model training to ensure proper autograd behavior. * **Tests** * Added regression test to validate gradient computation in speculative decoding scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: javierdejesusda <javier.dejesusj9@gmail.com> |
||
|
|
0a1ca5d48a |
Fix unified_export_megatron for transformers 5.6 (#1335)
### What does this PR do?
Type of change: Bug fix
Broaden the exception handler around `AutoTokenizer.from_pretrained` in
`GPTModelExporter.save_pretrained` to also catch `ValueError` and
`ImportError`.
In `transformers` 4.x, attempting to load a tokenizer from a directory
that contains only `config.json` (no `tokenizer.json` /
`tokenizer.model` / `tokenizer_config.json`) raised `OSError`, which was
already handled. In `transformers` 5.x the resolution path now reaches
`PreTrainedTokenizerFast.__init__` and raises a `ValueError` ("Couldn't
instantiate the backend tokenizer from one of: ...") when none of the
three backend sources are available. This caused
`export_mcore_gpt_to_hf` to hard-fail for checkpoint directories that
don't carry tokenizer files — including
`tests/gpu_megatron/torch/export/test_unified_export_megatron.py`, which
writes only a minimal `config.json`.
The broadened `except (OSError, TypeError, ValueError, ImportError)`
mirrors the pattern already used just below for
`AutoProcessor.from_pretrained` and keeps tokenizer export best-effort,
as originally intended.
### Usage
No API change. Existing call sites continue to work:
```python
from modelopt.torch.export import export_mcore_gpt_to_hf
export_mcore_gpt_to_hf(
model,
pretrained_model_name_or_path, # may or may not contain tokenizer files
dtype=torch.bfloat16,
export_dir=export_dir,
)
```
### Testing
- Reproduced the failure on `transformers==5.6` with:
```
pytest
tests/gpu_megatron/torch/export/test_unified_export_megatron.py::test_unified_export_megatron[llama-LlamaForCausalLM-medusa-None-None]
```
which failed with `ValueError: Couldn't instantiate the backend
tokenizer...` raised from `unified_export_megatron.py:299`.
- After the fix, the same parametrization passes, and the other `llama`
/ `nemotron` / `eagle` / `medusa` parametrizations in the same test file
remain green.
- No behavioral change on `transformers` 4.x: the `OSError` path is
still caught.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- Existing test
`test_unified_export_megatron` already covers this path; the fix makes
it pass on transformers 5.x. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
### Additional Information
Triggered by the upgrade to `transformers` 5.6 (used in
`nvcr.io/nvidia/nemo:26.04`, which is the container for the
`gpu_megatron` nox session). The error message from `transformers` —
*"You need to have sentencepiece or tiktoken installed..."* — is a
misleading generic fallback; `sentencepiece` and `tiktoken` are already
pulled in via the `[hf]` extras, and the real cause is the missing
tokenizer files in the export source directory.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Refactor**
* Improved error handling in the export process to gracefully manage
additional exception types.
* **Tests**
* Enhanced test validation for Megatron export by using actual tokenizer
artifacts, ensuring model vocabulary size matches test tokenizer
configuration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
fda0899e40 |
feat(recipes): add KV cache cast variants (fp8_cast / nvfp4_cast) (#1334)
## Summary
- Adds three built-in PTQ recipes that express the KV-cache *cast*
variants directly in YAML, using the existing `use_constant_amax: true`
quantizer field. These are recipe equivalents of
`--kv_cache_qformat=fp8_cast` / `nvfp4_cast`:
- `general/ptq/fp8_default-fp8_cast_kv`
- `general/ptq/nvfp4_default-fp8_cast_kv`
- `general/ptq/nvfp4_default-nvfp4_cast_kv`
- Makes `--recipe` authoritative in `examples/llm_ptq/hf_ptq.py`: the
post-hoc `_set_kv_cache_constant_amax` override now only runs when
`--recipe is None`, so a recipe YAML fully determines KV-cache config
instead of being silently overridden by the default
`--kv_cache_qformat=fp8_cast`. Updated help text on both flags.
- Extends the recipe loader smoke test to cover the three new recipes.
## Motivation
Before this change, the cast variants lived only in argparse
(`_KV_CAST_FORMATS = {"fp8_cast", "nvfp4_cast"}`) and were layered on
top of any recipe-loaded config. That meant `--recipe
nvfp4_default-fp8_kv` would silently become a cast recipe due to the
`--kv_cache_qformat` default. Now the recipe is self-contained: its YAML
either sets `use_constant_amax: true` on the `*[kv]_bmm_quantizer` entry
(cast) or doesn't (data-driven calibration).
## Test plan
- [x] `pytest tests/unit/recipe/test_loader.py` — all 24 tests pass,
including the three new parametrized recipes.
- [x] Verified each new recipe round-trips through `load_recipe()` with
`use_constant_amax: True` surviving Pydantic validation on the KV entry.
- [x] End-to-end run on `/models/Qwen/Qwen3-8B` (RTX 6000 Ada, 4
samples, seq_len=128) for all three new recipes:
- After `mtq.quantize(model, recipe.quantize.model_dump(),
forward_loop=...)`, all 72 `k_bmm_quantizer` / `v_bmm_quantizer` modules
have `_use_constant_amax=True` and `_get_amax()` returns `448.0` (FP8
E4M3 max).
- Weight quantizers still calibrate from data normally (sample amax
values: q_proj=0.5508, k_proj=0.6250, v_proj=0.1689, o_proj=0.7266).
- [x] Verified the `--recipe` authoritative behavior change:
- Non-cast recipe + default `--kv_cache_qformat=fp8_cast` → KV entry
does NOT get `use_constant_amax` (no silent override).
- Cast recipe + contradictory `--kv_cache_qformat=fp8` → KV entry keeps
`use_constant_amax=True` (recipe wins).
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Fixed CLI to respect KV cache quantization settings from recipe YAML
instead of overriding them.
* **New Features**
* Added three new post-training quantization recipe configurations for
FP8 and NVFP4 with optimized KV cache handling.
* **Documentation**
* Enhanced CLI help text for recipe and KV cache quantization options
with configuration examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
|
||
|
|
01788bb007 |
Deprecate Mllama support in llm_ptq/vlm_ptq examples (#1332)
## Summary - Removes Mllama (Llama 3.2 Vision) model-type branches from the `llm_ptq` example (`hf_ptq.py`, `example_utils.py`) and drops the now-unused `MllamaImageProcessor` wrapper from `modelopt/torch/utils/`. - Drops the legacy `MllamaImageProcessor` path in `modelopt/torch/utils/vlm_dataset_utils.py`; the generic HF ProcessorMixin path handles the remaining cases. - Adds a CHANGELOG entry under 0.44 Backward Breaking Changes. ## Test plan - [x] CI lint / unit tests pass - [x] Smoke-run ``examples/llm_ptq/scripts/huggingface_example.sh --model <llm> --quant fp8`` (text-only path, non-mllama) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Removed Mllama (Llama 3.2 Vision) support from quantization examples. This includes removal of dedicated image processor implementation, specialized model handling, and related calibration logic. * Updated VLM image-text calibration guidance to use `--calib_with_images` flag with other supported VLMs instead of Mllama-specific processing paths. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
8663678f12 |
fix: bug hf_ptq.py max_length setting ignored for LLMs (#1311)
### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage - fixes a bug in example script. We were trying why our models were not that strong at long context. Seems like a recent refactor did not implement max seq length. so 512 is used by default. ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Improvements** * Calibration data loading now enforces a maximum sequence/sample length during dataset preparation, ensuring calibration inputs adhere to configured length limits. This yields more predictable calibration behavior, reduces peak memory usage during calibration runs, and improves consistency of quantization preprocessing. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Michael Feil <63565275+michaelfeil@users.noreply.github.com> |
||
|
|
e4e3508f51 |
[OMNIML-3349] Add FP8 MHA quantization support for HuggingFace ViT (#1289)
## Summary Enables TensorRT attention-v2 fusion for vision transformers when exported to ONNX with FP8 Q/DQ. The core library changes are architecture-agnostic (drop-in for any FP8 ONNX export); coverage is exercised by the existing `examples/torch_onnx/torch_quant_to_onnx.py` pipeline. - **`modelopt/onnx/export/fp8_exporter.py`** — new post-processing passes: move attention-scaling `Mul` and K `Transpose` to the Q-side so DQ feeds MatMul directly, pre-transpose constant weights, and insert FP8 Q/DQ on Softmax outputs (fixed `1/448` scale, data-independent) for MHA-v2 fusion. Rewrites only fire when every downstream consumer is a MatMul so non-attention branches are never perturbed. - **`modelopt/onnx/utils.py`** — `fold_dq_fp32_to_fp16_casts` / `fold_q_fp16_to_fp32_casts` remove the Cast nodes `convert_float_to_float16` inserts around Q/DQ and rewrite scale initializers to FP16 so TRT fuses DQ into the downstream GEMM. Guarded behind opset >= 19 (FP16 Q/DQ scale requirement). Warns on FP16 overflow/underflow. - **`modelopt/torch/_deploy/utils/torch_onnx.py`** — calls the fold helpers for FP8-quantized models after `convert_float_to_float16`. - **`modelopt/torch/quantization/export_onnx.py`** — keeps FP8 Q/DQ scale in the native input dtype so no Cast is emitted between graph and Q/DQ. Removes the now-unused `trt_high_precision_dtype` parameter from `_fp8_quantize`/`_fp8_dequantize`. - **`modelopt/torch/quantization/nn/modules/quant_layernorm.py`** (new) — registers `nn.LayerNorm` in `QuantModuleRegistry` so LayerNorm output quantizers are honored. - **`modelopt/torch/quantization/plugins/huggingface.py`** — skips `*Attention` wrappers whose children are also `*Attention` per-instance (not per-class) to avoid double-patching `eager_attention_forward` (e.g. `ViTAttention` vs `ViTSelfAttention`). - **`examples/torch_onnx/torch_quant_to_onnx.py`** — adds a `_FP8_MHA_OVERRIDE` config block to FP8 mode that enables LayerNorm output quantizer + disables its input quantizer for TRT attention fusion. - **Unit tests** (12 CPU tests, ~1.2s total) — fp8_exporter rewrites + fanout safety, fold-cast helpers + opset guard, LayerNorm quant-wrapper identity, per-instance nested-attention detection. ## Benchmarks ViT-base-patch16-224, RTX 6000 Ada, strongly-typed FP8 via `trtexec`. Accuracy on 2 000 ImageNet-1k validation samples (streaming). **Batch = 1 (latency-bound)** | Model | Top-1 | Top-5 | TRT latency | Speedup | |---|---|---|---|---| | FP16 baseline | 80.96% | 95.80% | 0.722 ms | 1.00x | | Torch FP8 MHA | 80.66% | 95.75% | 0.657 ms | **1.10x** | | ONNX PTQ FP8 | — | — | 0.589 ms | **1.23x** | **Batch = 64 (throughput-bound, realistic inference)** | Model | TRT latency | Speedup | Images/s | |---|---|---|---| | FP16 baseline | 23.40 ms | 1.00x | 1152 | | Torch FP8 MHA | 15.89 ms | **1.47x** | 1152 | | ONNX PTQ FP8 | 15.89 ms | **1.47x** | 1216 | Top-1 accuracy stays within 0.30 pp of FP16; at batch=64 the Torch FP8 MHA path matches ONNX PTQ wall-time — attention is the bottleneck there and both paths achieve full FP8 attention fusion (36/36 attention MatMuls with QDQ in ViT-base). ## Test plan - [x] CPU unit tests (new): \`python -m pytest tests/unit/onnx/quantization/test_fp8_mha_exporter.py tests/unit/onnx/test_fold_casts.py tests/unit/torch/quantization/test_quant_layernorm.py tests/unit/torch/quantization/plugins/test_nested_attention_skip.py\` - [x] Existing ONNX / quantization unit suites unaffected: \`python -m pytest tests/unit/onnx tests/unit/torch/quantization\` - [x] End-to-end ViT FP8 export: \`python examples/torch_onnx/torch_quant_to_onnx.py --timm_model_name vit_base_patch16_224 --quantize_mode fp8 --onnx_save_path vit_base_fp8.onnx\` — expect log lines \`Folded 48 weight Transpose nodes\`, \`Inserted FP8 weight DequantizeLinear for 1 Conv nodes\`, and \`Attention QDQ rewrites: ... inserted QDQ on 12 Softmax outputs\` - [x] trtexec FP8 strongly-typed build: \`trtexec --onnx=vit_base_fp8.onnx --fp8 --stronglyTyped\` - [x] Accuracy within ~0.3 pp of FP16 baseline on ImageNet-1k subset --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> |
||
|
|
7265ca6793 |
Fix lm_eval version checking (#1321)
### What does this PR do? Type of change: bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> `lm_eval` does not have `__version__` attribute ### Additional Information <!-- E.g. related issue. --> NVBug 6102101 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Enhanced the package version detection system to improve overall reliability and stability of the application while reducing unnecessary external dependencies. All functionality, including version gating and system warnings, continues to operate exactly as expected with no impact on the user experience. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
2564da7324 |
Update vLLM deployment docs for heterogeneous models (#1317)
### What does this PR do? Type of change: ? documentation. This PR updates vLLM deployment instructions, taking into account heterogenous models created with AnyModel. ### Usage Does not apply. ### Testing Run the updated instructions in the documentation. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Replaced a benchmarking-focused section with a deployment guide for running compressed models on vLLM. * Added step-by-step setup for using an AnyModel-enabled vLLM fork, including checkout and install guidance and required model config edits (with optional architecture metadata). * Simplified runtime to a single vllm serve command, removing manual model rearrangement steps. * Restored inference benchmarking as a subsection, retaining vllm bench latency/throughput examples. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Grzegorz Karch <gkarch@nvidia.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
c7966119eb |
Reorg the sparse/quant/common kernel dir (#1303)
### What does this PR do? Type of change: re-org code <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ We changed the import path <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!--- Mandatory --> - Did you write any new necessary tests?: ❌ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Calibration support for skip-softmax multi-threshold measurement in sparse attention. * N:M sparse softmax masking and helpers for sparsity-aware attention. * **Chores** * Reorganized and consolidated kernel/backends for quantization and sparsity to a unified kernels layout, updating tests and examples to match. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
e56682e34a |
docs: update installation pages with legal-approved license notices (#1322)
## Summary
- Replaces the old pip license notice ("Please review the license terms
of ModelOpt and any dependencies before use") with the Legal-approved
wording: "Model Optimizer will download and install additional
third-party open source software projects. Review the license terms of
these open source projects before use."
- Adds a generic container license review notice ("Before pulling and
using the container images, please review their respective license
terms.") to the Linux installation doc (Docker tab) and README.
- Adds a `.. note::` with the pip notice to the Windows installation
page (covers both standalone and Olive child pages).
- Expands the README container section to explicitly list all four
recommended NVIDIA container images (`pytorch`, `nemo`, `tensorrt-llm`,
`tensorrt`).
## Test plan
- [x] Verify rendered docs look correct (`nox -s docs`)
- [x] Confirm legal notices appear in Linux, Windows, and README install
sections
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Updated installation guides with explicit references to supported
NVIDIA container images (PyTorch, NeMo, TensorRT-LLM and variants),
clarified pre-installed Model Optimizer in some images, and added notes
to review each container’s license terms; clarified conditional
environment setup wording and local install license guidance.
* **Chores**
* Updated project license header year.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
||
|
|
0678136335 |
Fix vLLM fakequant MoE megatron export bug (#1305)
### What does this PR do?
Type of change: Bug fix
Fixes two bugs in the vLLM + Megatron-Core MoE export path and cleans up
the related weight-collection helper:
1. **`_QuantFusedMoEBase` (vllm.py)**: The weight-quantizer path in
`_invoke_fused_moe_quantized_function` was temporarily mutating
`self.w13_weight` / `self.w2_weight` to the quantized tensor, then
restoring them via `finally`. This exposed a stale quantized tensor on
`self` between the mutation and the kernel call. Fixed by computing the
quantized weight directly into a local `B` without touching `self.*`
attributes.
2. **`GPTModelExporter` / `VllmFqGPTModelExporter`
(unified_export_megatron.py / vllm_fakequant_megatron.py)**:
`expert_bias` (present in grouped MoE layers) was silently dropped
during export because the bias collection ran after the early-return on
missing `weight`. Extracted a `_get_weight_bias` helper that collects
weight, bias, and expert_bias together, so bias/expert_bias are captured
even when weight is absent or zero-element.
### Usage
```python
# No API change; export pipelines pick this up automatically.
# export_mcore_gpt_to_hf_vllm_fq / export_mcore_gpt_to_hf now correctly
# export expert_bias for grouped-MoE checkpoints.
```
### Testing
Step 1 — Quantize (run from Megatron-LM
examples/post_training/modelopt):
```
HF_MODEL_CKPT=<path/to/hf/weights> MLM_MODEL_SAVE=<quant-ckpt-name> \
bash quantize.sh <hf-model-id> NVFP4_DEFAULT_CFG
```
Step 2 — Export for vLLM fakequant:
```
MLM_EXTRA_ARGS=--export-vllm-fq \
HF_MODEL_CKPT=<path/to/hf/weights> \
MLM_MODEL_CKPT=<quant-ckpt-name> \
EXPORT_DIR=<export-dir> \
bash export.sh <hf-model-id>
```
Step 3 — Serve (run from examples/vllm_serve):
```
QUANT_CFG=NVFP4_DEFAULT_CFG \
QUANT_FILE_PATH=<export-dir>/quantizer_state.pth \
python3 vllm_serve_fakequant.py <export-dir> \
-tp 1 --served-model-name <model-name> \
--host 0.0.0.0 --port 8000 \
--trust-remote-code --enforce-eager \
--disable-custom-all-reduce \
--gpu-memory-utilization 0.8
```
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Refactor**
* Centralized weight/bias/expert-bias extraction and export to a single
helper for consistent handling.
* Standardized quantized-weight flow to temporarily swap and restore
parameter tensors during computation.
* **Bug Fixes**
* Prevented missing or incorrect weight/bias exports by unifying
extraction logic.
* Broadened checkpoint key matching to preserve more quantizer state
during reloads.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
|
||
|
|
c417e6f4d9 |
Exclude small-k and small-n Matmul nodes from Int8 quantization (#1256)
### What does this PR do? Exclude small-dimension MatMul nodes from INT8 quantization. MatMuls with N or K < 16 cannot efficiently use INT8, causing performance regressions. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved quantization exclusions so MatMul/Gemm ops with derived K<16 or N<16 are skipped, honoring Gemm transB, using inferred and runtime-determined shapes, and avoiding duplicate outputs. * **Tests** * Expanded unit tests to cover constant, inferred, and runtime-derived shapes, Gemm transB behavior, small-dimension edge cases, and output deduplication. * **Documentation** * Added changelog entry documenting the new small-dimension exclusion thresholds and transB handling. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: samcheng <samcheng@nvidia.com> |
||
|
|
785d3a2df6 |
[CI] Bump test containers to latest (#1299)
- Use latest containers for testing in CICD <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Bumped TensorRT-LLM Docker images to 1.3.0rc12 in example and GPU test workflows. * Updated PyTorch container image from 26.01 to 26.03 for GPU tests. * Captured uv lock upgrade output to a temp file, inlined it into PR bodies, and adjusted workflow heredoc/templating and step behavior. * **Documentation** * Clarified an inline comment and simplified a warning message for an ONNX quantization extension. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
c51c1762b3 |
fix: prevent gh-pages repo bloat from doc preview artifacts (#1309)
### What does this PR do? Type of change: Bug fix Fixes gh-pages branch bloat that grew from ~26 MB to ~441 MB in four weeks (nvbug 6099503). Three compounding causes were identified and addressed: 1. **Sphinx `.doctrees/` cache published to gh-pages** — `sphinx-build` was writing its build cache inside `build/html/` which was then uploaded verbatim. Accounts for ~3.3 GB uncompressed across history. 2. **`JamesIves/github-pages-deploy-action` appending a commit on every push** — main-site files accumulated forever with `single-commit: false` (default). 3. **PR preview deploying on every `synchronize` event for all PRs** — `rossjrw/pr-preview-action` re-deployed the full site for every push to any PR regardless of whether docs changed (e.g. PR #1128 triggered 64 preview deploys × ~11 MB each). Changes: - Pass `-d /tmp/doctrees` to `sphinx-build` so `.doctrees/` is never written into `build/html/` - Add `paths: [docs/**, modelopt/**]` filter to `pull_request` trigger so the docs workflow only runs on PRs that touch docs or source code - Set `single-commit: true` on the deploy action so main-site pushes squash into one commit - Deduplicate docs build: `deploy-preview` now downloads the artifact from `build-docs` instead of running a second `sphinx-build` - Set `retention-days: 1` on the artifact since it is only needed for the duration of the workflow run The one-time cleanup (force-push squashed orphan to gh-pages) was already applied separately — repo is now ~59 MB for a full clone vs ~441 MB before. ### Usage N/A — CI/workflow change only. ### Testing - Workflow logic reviewed manually. - The one-time cleanup was verified: `git rev-list --objects --disk-usage origin/gh-pages` now reports ~28 MB; full clone is ~59 MB. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information nvbug 6099503 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Optimized documentation build and deployment workflow in CI/CD pipeline. * Improved pull request documentation preview handling with faster build timeouts and refined artifact management. * Enhanced GitHub Pages deployment configuration for better consistency. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
5ffb8487d9 |
add gptq fused kernel (#1291)
### What does this PR do? Add gptq fused kernel to improve speed. ### Usage check unittest ### Testing added a unittest ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Fused GPTQ backend for faster blockwise weight updates, toggleable via a new "fused" option. * Shared NVFP4 quantization primitives exposed for reuse. * **Refactor** * Consolidated FP4 scale/quantization logic into reusable utilities and centralized Hessian inversion handling. * **Tests** * Expanded GPU tests comparing fused vs unfused GPTQ, added Triton-availability gating and a local benchmark entrypoint. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shiyang Chen <shiychen@nvidia.com> |
||
|
|
2fef374ded |
fix: auto-compute dp_replicate_size from world_size (#1302)
## Summary - When `dp_shard_size < world_size` (e.g., `dp_shard_size=4` on 8 GPUs across 2 nodes), `ParallelismConfig` raises `total_size (4) does not match num_processes (8)` because `dp_replicate_size` defaults to 1 - Auto-compute `dp_replicate_size = world_size // (dp_shard_size * cp_size)` so intra-node FSDP2 sharding + inter-node data-parallel replication works without manual config - This enables `dp_shard_size` to be set to per-node GPU count (better NVLink utilization) while automatically creating replicas across nodes ## Test plan - [ ] Verify single-node training (dp_shard_size == world_size, dp_replicate_size == 1) unchanged - [ ] Verify multi-node with dp_shard_size < world_size creates correct replica groups - [ ] Verify existing EAGLE3/DFlash configs still work 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Enhanced parallelism configuration initialization in the speculative decoding example to better handle distributed training scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
355c6b7883 |
fix: PTQ 1GPU, export PP divisibility, hidden states conversations key (#1293)
## Summary - **megatron_lm_ptq.yaml**: Qwen3-8B PTQ to single GPU for L40 clusters (TP=1, all tasks) - **quantize.sh**: Auto-find largest PP dividing model's `num_hidden_layers` for export step. Qwen3-8B has 36 layers which isn't divisible by 8, causing `AssertionError` on 8-GPU nodes - **compute_hidden_states_trtllm.py**: Use `messages` with `conversations` fallback, matching the HF version. Fixes `KeyError: 'conversations'` when data uses OpenAI `messages` format ## Test plan - [x] Qwen3-8B PTQ runs on single L40 GPU - [x] Export PP auto-selects valid divisor (36 layers → PP=6 on 8 GPUs, PP=4 on 4 GPUs, PP=1 on 1 GPU) - [x] EAGLE3 offline pipeline reads data with `messages` field 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Dataset input handling now supports multiple field formats for enhanced compatibility. * **Bug Fixes** * Optimized GPU resource allocation during model quantization with improved pipeline parallelism computation. * Updated quantization configuration for more efficient resource utilization. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>0.45.0dev |
||
|
|
289a239ca5 |
fix: use data_dir for directory paths in ShardedDataset (#1301)
## Summary
- `datasets`' `resolve_pattern` only matches entries with
`type=="file"`, so passing a bare directory path as `data_files` to
`load_dataset` results in `FileNotFoundError` even when the directory
exists on disk
- Detect directory paths in `ShardedDataset._load_dataset()` and pass
them via `data_dir` instead of `data_files`
## Reproduction
```python
from datasets import load_dataset
# This fails with FileNotFoundError:
load_dataset("json", data_files="/path/to/data_directory")
# This works:
load_dataset("json", data_dir="/path/to/data_directory")
```
## Test plan
- [ ] Verify existing EAGLE3/DFlash training pipelines that pass
directory paths work
- [ ] Verify file path and glob patterns still work (falls through to
`data_files`)
- [ ] Verify `data_files=None` (no data_files arg) still works
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Bug Fixes
* Fixed an issue with dataset loading that prevented proper handling of
directory-based data sources. Directories are now correctly detected and
processed during dataset initialization.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
||
|
|
97d153118e |
[minor] Add custom calibration backend registry (#1281)
### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a public backend-specific calibrator registration API to support FP8 scale-sweep calibration, allowing backends to supply custom calibrators used during FP8 tuning. * **Tests** * Added unit tests confirming registry insertion/overwrite, that registered calibrators are invoked when FP8 scale-sweep is enabled, are not invoked when disabled, and that calibration falls back to defaults when no backend is registered. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
010b220dc0 |
vLLM fakequant export update for AWQ checkpoint (#1242)
### What does this PR do?
Type of change: Bug
Enables end-to-end AWQ checkpoint export and reload in the vLLM
fake-quant serving path (`MODELOPT_STATE_PATH`). Previously, the
`input_quantizer` was using incorrect `pre_quant_scale` especially with
grouped quantizers like `qkv_proj`, using simply the first
`input_quantizer.pre_quant_scale`. This MR adds
`_resmooth_experts_for_export` that non-mutatively averages
`pre_quant_scale` across MoE experts and unifies input `_amax`, required
because vLLM uses a single input quantizer per expert group. Adds
`merge_amax_tensors_for_group` (element-wise max for same-shape, `cat`
for GQA, scalar-max fallback) replacing the scalar-collapsing
`torch.stack().max()` that dropped per-channel `_amax` structure.
### Usage
```python
# Export AWQ checkpoint from HF model
from modelopt.torch.export.plugins.vllm_fakequant_hf import export_hf_vllm_fq_checkpoint
export_hf_vllm_fq_checkpoint(model, export_dir="./awq_vllm_checkpoint")
```
### Testing
**Step 1 — Export the quantized checkpoint:**
```bash
python examples/llm_ptq/hf_ptq.py \
--pyt_ckpt_path <MODEL_PATH> \
--recipe <AWQ_RECIPE> \
--calib_size 512 \
--export_path <EXPORT_DIR> \
--vllm_fakequant_export
```
This produces `<EXPORT_DIR>/vllm_fq_modelopt_state.pth` with the averaged per-expert
pre_quant_scale and unified _amax now included.
Step 2 — Serve via vLLM fakequant worker:
```bash
MODELOPT_STATE_PATH=<EXPORT_DIR>/vllm_fq_modelopt_state.pth \
python examples/vllm_serve/vllm_serve_fakequant.py \
<EXPORT_DIR> --tensor-parallel-size <TP>
```
Tested for quantization configurations:
```
FP8_DEFAULT_CFG
FP8_DEFAULT_CFG (input_q disabled)
INT8_SMOOTHQUANT_CFG
INT8_WEIGHT_ONLY_CFG
NVFP4_DEFAULT_CFG
NVFP4_AWQ_LITE_CFG
INT4_AWQ_CFG
NVFP4_AWQ_CFG
NVFP4_DEFAULT_CFG (input_q disabled)
```
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit
* **New Features**
* Added Nemotron-style MoE export support and group-aware AWQ resmoothing with optional requantization during export.
* Improved handling for shared-input / expert groups and tensor-parallel sharding of pre-quantization scales.
* **Bug Fixes**
* Removed AWQ reload limitation from known issues; improved checkpoint validation and safer save/load behavior.
* Better detection and handling of enabled weight-quantizers and clearer warnings for mismatched checkpoint keys.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
|
||
|
|
26ae8da517 |
[2/3] Implicit Gemm NVFP4 (#1227)
### What does this PR do? Type of change: new feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> - Add Conv3D implicit GEMM kernel with BF16 WMMA tensor cores and fused NVFP4 activation quantization for video diffusion VAE layers - Integrate into _QuantConv3d via QuantModuleRegistry — automatically dispatched when NVFP4 quantization is applied to nn.Conv3d - Move kernel from `experimental/conv/ to modelopt/torch/kernels/conv/`; move tests to `tests/gpu/torch/quantization/kernels/` ### Testing <!-- Mention how have you tested your change if applicable. --> - Added test cases to measure the difference between cuDNN and our CUDA implicit GEMM kernel - Added an NVFP4 fake quantization test using CUDA code ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Per-backbone quantization/export in a single run with per-backbone checkpoints and backbone-aware quant filters * Configurable NVFP4 block-size via CLI/config; improved NVFP4 Conv3D inference path and Wan 2.2 quantization support * **Bug Fixes** * Video-model calibration now respects extra params and forces video decoding during calibration * **Documentation** * Added comprehensive Conv3D implicit‑GEMM kernel documentation; removed experimental Conv3D prototype docs/benchmark * **Tests** * New Wan 2.2 quantization/export tests and expanded Conv3D/FP4 kernel test coverage <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com>0.44.0rc0 |
||
|
|
c20f9c411d |
Add a standalone monitor skill for persistent job tracking (#1252)
### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Add a standalone monitor skill for persistent job tracking across sessions, and integrate it with PTQ, evaluation, and deployment skills. Problem: Each skill had ad-hoc inline monitoring (squeue polling, nel status checks) that didn't survive session restarts and couldn't track multiple jobs. Users had to manually ask "check status" every time. Solution: A centralized monitor skill with: - Job registry (.claude/active_jobs.json): single source of truth for all active jobs - Durable recurring cron: polls every 15 min, survives session restarts, self-cleans when all jobs complete - User-initiated mode: works in new conversations by reading the registry - Aggregated reporting: "2 of 4 completed" instead of per-job noise ### Usage After any skill submits a job, the monitor skill automatically: 1. Registers the job in .claude/active_jobs.json 2. Sets up a durable cron to poll status every 15 minutes User can also trigger manually: User: "check my eval status" → reads registry, reports current state User: "is the PTQ done?" → finds job, checks status User: "what jobs are running?" → lists all registered jobs ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added monitor skill for tracking SLURM jobs, NEL evaluations, and launcher experiments with persistent job registry. * **Documentation** * Updated deployment, evaluation, and PTQ documentation to use the new monitor skill. * Simplified diagnostic and troubleshooting instructions. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kai Xu <kaix@nvidia.com> |
||
|
|
e9a49890f1 |
[chore]: weekly bump of uv.lock on main (2026-04-18) (#1292)
## Summary Automated weekly update of uv.lock file for nSpect Scanning: - `uv.lock` — upgraded all transitive dependencies to latest compatible versions Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> |
||
|
|
3d0f0db49e |
[CI] Replace tox with nox, use nemo:26.04 for megatron tests, and simplify CI workflows (#1286)
### What does this PR do? Type of change: New feature / infrastructure improvement Follow-up to #1285 for correct CI test environment for megatron based tests Replaces `tox` + `tox-current-env` with `nox` for all test, lint, docs, and wheel build sessions. The primary motivation was that `tox-current-env` is incompatible with uv venvs in NGC containers (e.g. NeMo's `/opt/venv`) — it picks the system Python via `sys._base_executable` instead of the container's venv Python which has megatron packages pre-installed. Key changes: - **`noxfile.py`** replaces `tox.ini` with GPU, CPU unit, partial-install, pre-commit, docs, and wheel sessions - **GPU sessions** use `venv_backend="none"` (run directly in container env) and `python -m pip/pytest` to avoid PATH mismatches - **uv** is set as the default venv backend (if available) for CPU sessions (faster installs) Also includes CI workflow simplifications: - **`_pr_gate.yml`** new reusable workflow centralizing file-change detection + linux-check wait logic (was duplicated across 3 workflow files) - **Collapsed pr/non-pr job pairs** into single jobs with conditional `runs-on` in `gpu_tests.yml`, `example_tests.yml`, `regression_tests.yml` - **Collapsed `multi-py` / `multi-torch` / `multi-transformers`** into a single `multi-version` matrix job in `unit_tests.yml` - **PR path filtering** for unit test secondary jobs (multi-version, launcher, partial-install) — skipped if no relevant files changed - **Fixed schedule/workflow_dispatch skipping** — jobs with `needs: [pr-gate]` were incorrectly skipped when all pr-gate internal jobs were skipped; fixed by making the gate job always run - **multi-version, launcher, partial-install** now also run on `schedule` / `workflow_dispatch` ### Usage ```bash python -m pip install nox uv # install nox and uv (once) nox -l # list all sessions nox -s gpu_megatron # run a GPU session (inside container) nox -s "unit-3.12(torch_211, tf_latest)" # run a specific unit test combination nox -s "unit-3.12(torch_211, tf_latest)" -R # force-recreate venv (e.g. after dep changes) COVERAGE_PROCESS_START=pyproject.toml nox -s "unit-3.12(torch_211, tf_latest)" # with coverage ``` ### Testing - Ran `nox -l` to verify all session names - Ran `gpu_megatron` session locally inside NeMo container — confirmed it uses `/opt/venv/bin/python` correctly - Manually triggered nightly-runs: - Unit: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608013657 - GPU: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608018763 - Examples: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608017322 ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A — CI infrastructure only - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ (added `nox` and `uv` to `dev-test`, both Apache-2.0) - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — no user-facing changes ### Additional Information Supersedes the tox-current-env workaround in the parent branch. --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
2b315eda47 |
Replace mip package with pulp (#663)
Replace mip package with more popular pulp package for puzzle mip solving. Both use the CBC solver under the hood ## Testing - Results very close for Qwen3-8B and Nemotron-Nano-12B-v2 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Simplified GPU test environment setup by removing unnecessary system dependency installation * Updated internal optimization solver dependencies in the puzzletron module <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
2004779a67 |
Update README.md for DMS (fix cd experimental/DMS to cd Model-Optimizer/experimental/DMS) (#879)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** ? ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated DMS installation instructions to reflect the repository structure and correct directory navigation during setup. * Clarified the setup steps so users follow the accurate directory change before running installation commands. * Small wording improvements to reduce confusion during the installation process. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Farid Adilazuarda <42537562+faridlazuarda@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
760c980727 |
Add ResNet50 support for torch_onnx quantization workflow (#1263)
## Summary - Add end-to-end ResNet50 support in the torch_onnx quantization → ONNX export → TRT engine pipeline - Fix multiple Conv2d-related export issues that blocked Conv2d-heavy models from working with FP8/INT8/MXFP8/NVFP4/auto quantization modes - Fix `configure_linear_module_onnx_quantizers` to handle all modules with block quantization (not just `nn.Linear`), fixing NVFP4/MXFP8 export for models with quantized non-Linear modules - Add `--trt_build` flag to `torch_quant_to_onnx.py` and simplify test infrastructure ### Files Changed - `modelopt/torch/_deploy/utils/torch_onnx.py` — Disable FP8 Conv2d weight quantizers and autocast during ONNX export - `modelopt/torch/quantization/export_onnx.py` — Fix `configure_linear_module_onnx_quantizers` for all module types with block quantization - `examples/torch_onnx/torch_quant_to_onnx.py` — Add `--trt_build` flag, calibration for FP8 override quantizers, Conv2d→FP8 override for auto mode, filter_func updates - `examples/torch_onnx/README.md` — Add ResNet50 to supported models table - `tests/examples/torch_onnx/test_torch_quant_to_onnx.py` — Add ResNet50 test entry, simplify using `--trt_build` - `tests/_test_utils/torch/vision_models.py` — Add ResNet50 to timm model registry ### Quantization modes passing - ✅ FP8, INT8, MXFP8, NVFP4, Auto (all 5 modes pass export + TRT build) - INT4_AWQ excluded (pre-existing limitation for all models) ## Test plan - [x] All 5 resnet50 test modes pass: `pytest tests/examples/torch_onnx/test_torch_quant_to_onnx.py -k resnet50` (5/5 passed) - [x] Full regression: 18 passed, 2 failed (pre-existing swinv2_tiny fp8/int8 failures) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added ResNet50 to supported ONNX export vision models with FP8, INT8, MXFP8, and NVFP4 support. * Optional TensorRT engine build after export via a new CLI flag. * **Improvements** * Enhanced quantization calibration and export flows for FP8/INT8 models, including broader block-quantization support across module types and safer export handling. * Tests updated to include ResNet50 in the model matrix. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> Signed-off-by: ajrasane <arasane@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
92622a9aa6 |
Add nvfp4_mse and nvfp4_local_hessian options to the ptq script. (#1113)
### What does this PR do? Type of change: Bugfix <!-- Details about the change. --> Add newly added quant configs to the example PTQ script. ### Testing I have locally run auto_quantize with these two quant_configs, and obtained successfully exported HF artifacts. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for three new quantization formats: nvfp4_mse, nvfp4_local_hessian, and nvfp4_experts_only, expanding available export options when using auto-quantize. * **Bug Fixes / UX** * Updated the invalid-quantization error message to include the newly accepted format identifiers. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Bilal Kartal <bkartal@nvidia.com> Signed-off-by: bkartal-dev <bkartal@nvidia.com> |