mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
a4fde491cc6c3c7c747c7c946416ae09d2c6fa5d
511
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a4fde491cc |
Update MOE block detection logic and enable in huggingface_script.sh (#962)
### What does this PR do? Type of change: Bug fix Add moe expert calib ratio in huggingface_script.sh Also fix minimax2.5 MOE detection which does not follow other HF MOE layer convention ### Usage scripts/huggingface_example.sh --model <MiniMax-M2.5> --quant nvfp4 --moe_calib_experts_ratio 1.0 --trust_remote_code ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, using `torch.load(..., weights_only=True)`, avoiding `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other source, did you follow IP policy in [CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Configure MOE calibration experts ratio for quantization via an environment/option, enabling finer control over calibration. * **Bug Fixes** * Improved detection of sparse MOE blocks to handle varying expert/topology layouts, inferring expert counts when needed for more reliable processing. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
78779f4a94 |
[OMNIML-3495] Fix megatron_generate to pass position_ids for MTP model support (#955)
### What does this PR do?
Type of change: Bug fix
`megatron_generate` and `megatron_prefill` set `position_ids = None` for
text-only LLM models. This crashes models with Multi-Token Prediction
(MTP) layers (e.g., Nemotron Super with `mtp_num_layers=2`) because
MTP's `roll_tensor` calls `torch.roll(position_ids, ...)` which requires
a Tensor, not None.
```
File "megatron/core/transformer/multi_token_prediction.py", line 161, in roll_tensor
rolled_tensor = torch.roll(tensor, shifts=shifts, dims=dims)
TypeError: roll(): argument 'input' (position 1) must be Tensor, not NoneType
```
This fix generates sequential `position_ids = [0, 1, 2, ..., seq_len-1]`
per batch element.
### Usage
No API change. Models with MTP now work with `megatron_generate` without
any caller-side changes:
```python
from modelopt.torch.utils.plugins.megatron_generate import megatron_generate
# This now works for MTP models (previously crashed with TypeError)
megatron_generate(model, tokens.input_ids.cuda(), osl=1)
```
### Testing
- Verified Nemotron Super (88-layer hybrid Mamba/Attention/MoE model
with `mtp_num_layers=2`) quantization + generation works end-to-end with
`megatron_generate` on 16 GPUs (2 nodes, TP=16, EP=16).
```
torchrun --nproc_per_node 8 examples/quantization/quantize.py \
--hf-model-id /models/nemotron-super-rl-021126/ \
--calib-size 1 \
--export-quant-cfg fp8 \
--megatron-save-path /models/nemotron-super-rl-021126-FP8-MLM \
--pp 1 \
--tp 8 \
--ep 8 \
--trust-remote-code
```
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).
- Is this change backward compatible?: ✅ All modern models (RoPE, YaRN,
Mamba) have `add_position_embedding = False`, so `position_ids` is
ignored by the embedding layer. RoPE generates identical sequential
positions internally whether `position_ids` is provided or `None`.
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
N/A
- Did you write any new necessary tests?: ❌ Existing quantization tests
cover the code path. The fix is a one-line behavioral change (None →
Tensor) with no new branches.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ Bug fix for a previously unsupported model type.
### Additional Information
The VLM code path in `_forward_step_func` already generates
`position_ids` correctly (lines 223-227). This fix applies the same
pattern to the text-only LLM code path for consistency.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved positional ID handling in the text generation pipeline,
ensuring proper position tracking across batch processing in both
prefill and generation stages.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: James Shen <yueshen@nvidia.com>
|
||
|
|
a076e6c218 |
Make GPU tests >2x faster by reusing spawn processes between tests (#958)
### What does this PR do? Type of change: Test infra improvement <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> * Instead of spawning multiprocesses per test case (adds ~10s overhead x 150 tests), we spawn once per file (gpu_megatron) or session (gpu) to make the tests much much faster * Remove some unused / unnecessary tests * Fix and enable Minitron NAS test on Hybrid models for Mcore 0.16 Test time (`pytest tests/<test_type>`) on 2x RTXPro 6000 Blackwell: https://github.com/NVIDIA/Model-Optimizer/actions/runs/22625069275 | | gpu | gpu_megatron | |------------|-----|--------------| | **Before** | 27m | 40m | | **After** | 16m | 19m | Full PR-merge CI/CD now only takes ~45mins ### Testing <!-- Mention how have you tested your change if applicable. --> Tested on 1,2,4,8 GPU setup ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, using `torch.load(..., weights_only=True)`, avoiding `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other source, did you follow IP policy in [CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?: N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Reworked GPU test infra to use a persistent distributed worker pool with improved per-worker teardown, aggregated error reporting, Megatron-specific cleanup, reduced GPU CI timeouts, and a small build-system dependency change. * **Tests** * Converted many multi-process GPU tests to a fixture-driven distributed runner, updating test entry points and orchestration while preserving test logic; adjusted imports and conditional guards for several GPU tests. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
d780fa593f |
Add uv.lock for experimental/dms (#964)
For nspect scanning, we need `uv.lock` with each `pyproject.toml` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated dependency management and lock configuration settings. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
860d0b4a70 |
Merge Linux and Windows Changelog (#954)
Since we no longer have any compiled packages, all releases are for all platforms so we dont have separate windows releases hence merging changelog as well. Going forward, windows can test on latest linux version and if fixes needed, they can go in next usual monthly linux release <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Reorganized changelog structure to consolidate Windows and Linux release information in a unified view. * Expanded Windows Support documentation across recent release versions. * **Chores** * Updated Windows example release badge to dynamically reflect the latest PyPI release version. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
edde087bbd |
Adding a special list to handle HF models that require (#950)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Bug fix **Overview:** ? 1. Some models check `layer_types` and our offline overwrites `num_hidden_layers=0` which will requires passing `layer_types=[]`. It is also possible that the model will raise unexpected arguments when passing `layer_types=[]`. So the current WAR is to use a list to handle by checking the architectures. 2. Fix relative path issue in `launch_train.sh` ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Enhanced model loading robustness for specific layer configurations. * **Chores** * Improved example script path handling and GPU initialization tracking. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
82f1d216d1 |
Add Security and IP related contributing guide and configure coderabbit to catch such issues (#935)
### What does this PR do? - Add Security related coding practices in `SECURITY.md` and merge with `2_security.rst` - Update `CONTRIBUTING.md` for instructions to follow if copying code from other repositories - Update PR template - Cleanup dependency files - New API `mto.load_modelopt_state` doing the insecure `torch.load(f, weights_only=False)` instead of doing it separately everywhere. This also allows us to later improve the input validation for `modelopt_state_path` or use safer alternatives to `torch.load` ### Testing <!-- Mention how have you tested your change if applicable. --> N/A ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=True)`, `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other source, did you follow IP policy in [CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?: ✅ <!--- Mandatory --> - Did you write any new necessary tests?: NA <!--- Mandatory for new features or examples. --> - Did you add or update any necessary documentation and update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: NA <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Expanded and reorganized security guidance and contributor procedures; updated PR template and several READMEs with clearer security, submission, and installation instructions * Replaced an older security document with an enhanced, centralized security guidance * **Chores** * Adjusted example dependency lists and optional extras (adds, removals, and version constraints) * Enabled automated incremental reviews, added pre-merge security checks, and introduced a knowledge-base of coding/security guidelines <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
f26b9c3a2f |
skip generate option for large models and mxfp8 (#942)
## What does this PR do?
**Type of change:** New feature
**Overview:** Adds a `--skip_generate` flag to `hf_ptq.py` that skips
the pre/post-quantization generation preview calls. These calls run
`model.generate()` which crashes for very large models (500B+) that are
split across GPU and CPU via `device_map="auto"` (e.g., models with
Mamba/Triton kernels that cannot handle CPU-offloaded tensors).
## Usage
```
python examples/llm_ptq/hf_ptq.py \
--pyt_ckpt_path /path/to/model \
--export_path /path/to/output \
--qformat mxfp8 \
--trust_remote_code \
--export_fmt hf \
--batch_size 1 \
--skip_generate \
--kv_cache_qformat none
```
## Testing
Tested with a 500B parameter NemotronH hybrid Mamba/attention model on
4x GB200 GPUs. Without --skip_generate, the script crashes at
model.generate() due to Mamba Triton kernels failing on CPU-offloaded
tensors. With --skip_generate, the generation preview is skipped and
quantization proceeds normally.
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
The --skip_generate flag sets generated_ids_before_ptq = None early,
which also causes the post-quantization generate to be skipped via the
existing if generated_ids_before_ptq is None: pass guard. Combined with
--batch_size 1 (to skip the get_max_batch_size forward-pass probe), this
eliminates all forward passes that can crash for device-map-split
models.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Introduced `--skip_generate` CLI option to skip pre-quantization text
and image generation, reducing processing time for very large models.
Useful when generation previews are computationally expensive.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: adithyare <adithyare@nvidia.com>
Signed-off-by: Adi Renduchintala <adithya.r@gmail.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
|
||
|
|
ba29ad706d |
Enable RTX Pro 6000 Blackwell runners for CI/CD (#944)
## What does this PR do? **Type of change:** CI/CD Improvement <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Updated CI/CD test matrix (new `cuda13-gpu-trtllm` dedicated job for gpu tests on trtllm container) | Workflow | Trigger | Test Matrix | GPU Runner | |---|---|---|---| | **GPU tests** | PR | `cuda13-gpu`, `cuda13-gpu-megatron`, `cuda13-gpu-trtllm` | 1x RTX Pro 6000 | | **GPU tests** | Nightly | `cuda13-gpu`, `cuda13-gpu-megatron`, `cuda13-gpu-trtllm` | 2x RTX Pro 6000 | | **Example tests (torch)** | PR | `llm_distill`, `llm_qat`, `llm_sparsity`, `speculative_decoding` | 1x H100 | | **Example tests (torch)** | Nightly | `llm_distill`, `llm_qat`, `llm_sparsity`, `speculative_decoding` | 2x RTX Pro 6000 | | **Example tests (trtllm)** | PR | `llm_ptq`, `vlm_ptq` | 1x RTX Pro 6000 | | **Example tests (trtllm)** | Nightly | `llm_autodeploy`, `llm_eval`, `llm_ptq`, `vlm_ptq` | 2x RTX Pro 6000 | | **Example tests (onnx)** | PR | `diffusers`, `torch_onnx` | 1x L4 | | **Example tests (onnx)** | Nightly | `diffusers`, `torch_onnx` | 2x RTX Pro 6000 | ## Testing <!-- Mention how have you tested your change if applicable. --> - [ ] Per PR tests pass in this PR - [ ] Nightly tests manually triggered: https://github.com/NVIDIA/Model-Optimizer/actions/runs/22495679199 (GPU tests), ? (Example tests) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Unified and simplified CI test matrices and reduced duplicate workflow configuration. * Updated GPU runner targets and container images for test jobs to newer GPU types. * **Tests** * Added new GPU test variants and several new test modules. * Simplified test gating: megatron auto-skip removed; tests now use a single dependency check for the mamba provider. * **Tooling** * Added a new tox environment for an additional GPU test suite. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
0f668a312f |
Replace setup.py with uv-compatible pyproject.toml and add uv.lock for nspect scanning (#945)
## What does this PR do? **Type of change:** Project config update <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Remove unnecessary setup.py as the same can be done with pyproject.toml. No change for users in terms of installation command. ## Testing <!-- Mention how have you tested your change if applicable. --> Per PR tests passing is sufficient validation <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Migrated project configuration from setup.py to pyproject.toml, adopting modern Python packaging standards. * Updated CI/CD workflows and development tools to reference the new configuration location. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
94b2cb272d |
Integrate Automated QDQ placement tool - part 3.3 (#839)
## What does this PR do?
This PR implements QDQ autotuner CLI. This is the initial version of
CLI, it will be integrated to modelopt.onnx.quantization.autotune.
Usage:
```
python -m modelopt.onnx.quantization.autotune
--onnx_path model.onnx --schemes_per_region 50
--pattern_cache cache.yaml --qdq_baseline baseline.onnx
--quant_type int8 --verbose
```
PR 3.1: https://github.com/NVIDIA/Model-Optimizer/pull/837
PR 3.2 https://github.com/NVIDIA/Model-Optimizer/pull/838
PR 3.3: https://github.com/NVIDIA/Model-Optimizer/pull/839
**Overview:** ?
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: Document will
be added in part 4.
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
CHANGE log will be added in part 4.
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added a command-line interface for ONNX quantization autotuning with
configurable parameters for models, output paths, quantization
strategies, and TensorRT benchmarking.
* Introduced an automated workflow for pattern-based region optimization
with state management, baseline comparison, and benchmarking
capabilities.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Will Guo <willg@nvidia.com>
|
||
|
|
fcdaf6519d |
Support decoder block-level sequential calibration (#924)
## What does this PR do?
**Type of change:** New feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:** Add support for sequential calibration of layers (at
decoder level granularity) in ModelOpt.
Calibration flow
1. Get list of decoder blocks
2. For current block call get input activations (considering weight and
activation QDQ from all other previous blocks) and call specified
calibration function.
functions added
1. get_decoder_layers() -> to detect and get list of blocks to iterate
over
2. LayerActivationCollector class -> to get input activations to the
layer
3. sequential_calibrate() -> to perform the described calibration flow
4. use_sequential field in QuantizeAlgorithmConfig
## Usage
<!-- You can potentially add a usage example below. -->
```python
# Sample config
NVFP4_DEFAULT_CFG = {
"quant_cfg": {
"*weight_quantizer": {
"num_bits": (2, 1),
"block_sizes": {-1: 16, "type": "dynamic", "scale_bits": (4, 3)},
"axis": None,
"enable": True,
},
"*input_quantizer": {
"num_bits": (2, 1),
"block_sizes": {-1: 16, "type": "dynamic", "scale_bits": (4, 3)},
"axis": None,
"enable": True,
},
**_default_disabled_quantizer_cfg,
},
"algorithm": {
"method": "max",
"use_sequential": True,
}
```
Set use_sequential=True in QUANT_CFG's "algorithm" section.
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Sequential layer-by-layer calibration: Quantization now supports
processing decoder layers sequentially to improve memory efficiency on
large models.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com>
|
||
|
|
2905cb0f2e |
Updated the diffusion config issue and more test cases (#937)
## What does this PR do?
**Type of change:** new tests, Bug fix <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->
**Overview:**
- **Fixed the INT8 config issue**
- **Add HF checkpoint export test coverage**
1. The `--hf-ckpt-dir` export path had zero test coverage. This MR adds
tests at two levels:
2. Unit tests (tests/unit/torch/export/test_export_diffusers.py):
- Extended test_export_diffusers_real_quantized to parametrize over
INT8, INT8 SmoothQuant, FP8, and FP4 configs
- (previously only FP8). This gives 3 models x 4 configs = 12 test
cases.
3. GPU integration tests
(tests/gpu/torch/export/test_export_diffusers_hf_ckpt.py)
- New file testing the full quantize.py --hf-ckpt-dir pipeline via
subprocess with 4 combos:
- SDXL INT8 smoothquant min-mean (the exact scenario that triggered the
bug)
- Flux INT8 smoothquant min-mean
- SDXL FP8
- Flux FP4
## Usage
<!-- You can potentially add a usage example below. -->
```python
# Add a code snippet demonstrating how to use this
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**:No <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**:No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No
<!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **Tests**
* Added test coverage for exporting Diffusers models with Hugging Face
checkpoints across multiple quantization formats (INT8, FP8, FP4)
* Extended quantization export testing to validate multiple
configuration scenarios
* **Chores**
* Refined INT8 quantization configuration with improved calibrator
support for convolution layers
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
|
||
|
|
56e97c8096 |
Bug fix 5875873 (#865)
## What does this PR do? **Type of change:** Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Newer version of trl uses dtype instead of torch_dtype. Modified code to set float32 as default for older versions of trl that you torch_dtype. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Enhanced error handling in model training examples to safely manage missing dtype attributes, preventing crashes during initialization when torch_dtype is not configured. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com> |
||
|
|
110a44c81e |
Add support for export ComfyUI compatible checkpoint for diffusion model(e.g., LTX-2) (#911)
## What does this PR do
Add support for export ComfyUI compatible checkpoint for diffusion
model(e.g., LTX-2)
**Type of change:** <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->
**Overview:**
Add support for export ComfyUI compatible checkpoint for diffusion
model(e.g., LTX-2)
1) Added a a parameter for merging the base vae, vocoder, connectors in
the quantized checkpoint
2) storing quantization metadata and export tool as modelopt , required
for ComfyUI compatibility.
3) Internally updating the transformer block prefixes to match the
expectation of ComfyUI
## Usage
<!-- You can potentially add a usage example below. -->
```python
export_hf_checkpoint(
pipeline,
export_dir=EXPORT_DIR,
merged_base_safetensor_path=BASE_CKPT, # merge VAE/vocoder from base
)
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
1) Tested with ltx-2 model
a) initializing a twoStagePipeline object
b) calling mtq.quantize on transformer with NVFP4_DEFAULT_CFG
c) then exporting with export_hf_checkpoint passing the param
merged_base_safetensor_path to generate merged
checkpoint
2) Ran the generated checkpoint with step1 on ComfyUI to validate
3) Ran step1 without merged_base_safetensor_path to check backward
compatibility.
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: NA
- **Did you add or update any necessary documentation?**: NA
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
NA <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added support for exporting LTX-2 diffusion models with merged base
checkpoint integration
* Enhanced export functionality to preserve and attach quantization
metadata during model export
* Extended model export capabilities with automatic model type detection
for improved export handling
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: ynankani <ynankani@nvidia.com>
|
||
|
|
6f094d7fa5 |
[For RL] Keep attrs after folding weight and fix empty extra state for Megatron (#779)
## What does this PR do? **Type of change:** improvement **Overview:** - For Quantization aware reinforcement learning, after folding weight of rollout, we want to keep the quantization attrs for next step. - Minor fix for empty extra state - Support getting dataloader from jsonl file, useful for using training data as calibration data. I can separate this to another PR if necessary. ## Usage `mtq.fold_weight(keep_attrs=True)` will keep quantizer attrs after folding weight, ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: NA - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for loading dataset samples directly from JSONL/JSONL.GZ files * Added optional parameter to skip logits return in generation prefill operations * Enhanced weight folding operations to optionally preserve quantization attributes during model optimization * **Bug Fixes** * Fixed handling of empty tensor states to prevent deserialization errors in Megatron module <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Meng Xin <mxin@nvidia.com> |
||
|
|
a538f2e256 |
Fix skip softmax calibration memory issue (#923)
Fix OOM issue when running skip softmax calibration
Test:
```
python examples/llm_sparsity/attention_sparsity/hf_sa.py \
--pyt_ckpt_path Qwen/Qwen3-30B-Instruct-A3B-2507 \
--sparse_attn skip_softmax_calib
```
works with >= 96GB GPU memory
---------
Signed-off-by: Rohan Joshi <rohjoshi@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
||
|
|
35e60991c7 |
Integrate Automated QDQ autotuner - part 3.2 (#838)
## What does this PR do? This PR implements QDQAutotuner class. This class is used to drive the main Autotuner workflow. The workflow is: 1. uses RegionSearch to build regions 2. generate QDQ ONNX models and evaluate perf 3. save best model This PR is part 2/4 of https://github.com/NVIDIA/Model-Optimizer/pull/703. PR 3.1: https://github.com/NVIDIA/Model-Optimizer/pull/837 PR 3.2 https://github.com/NVIDIA/Model-Optimizer/pull/838 PR 3.3: https://github.com/NVIDIA/Model-Optimizer/pull/839 **Overview:** ? ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Not in this part. - **Did you add or update any necessary documentation?**: No, document will be updated in part 4. - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No, change log will be updated when all changes are ready. ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Introduced ONNX Q/DQ autotuning framework with automatic region discovery and pattern-based optimization. * Added model profiling and quantization scheme generation capabilities. * Enabled state persistence and quantization model export functionality. * Introduced configuration management for quantization parameters and profiling workflows. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Will Guo <willg@nvidia.com> |
||
|
|
a415667992 |
Enable Qwen3.5-MoE PTQ (#897)
## What does this PR do? **Type of change:** New model support **Overview:** Add ModelOpt PTQ support for https://huggingface.co/Qwen/Qwen3.5-397B-A17B ## Usage <!-- You can potentially add a usage example below. --> ```python python3 hf_ptq.py --pyt_ckpt_path /home/omniml_data_3/models/Qwen3.5-397B-A17B --qformat nvfp4_mlp_only --export_path /home/omniml_data_3/zhiyuc/checkpoints/Qwen3.5-397B-A17B-NVFP4 --trust_remote_code ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Not yet <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Qwen3.5 Mixture-of-Experts model support in quantization workflows. * **Bug Fixes** * Enhanced error diagnostics during model export with detailed module information. * Improved dataset tokenizer processing with proper truncation and length handling. * Fixed model export stability issue related to framework integration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> |
||
|
|
a6cbcbad0a |
Support mixed-precision per layer quant config in config.json (#929)
## What does this PR do?
**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
**Overview:** Support mixed-precision per layer quant config in
config.json, since it's the first-class source of truth in deployment
FWs.
## Usage
<!-- You can potentially add a usage example below. -->
```python
python3 -c "
from modelopt.torch.export.convert_hf_config import convert_hf_quant_config_format
import json
# Test 1: Existing FP8 case still works
fp8_config = {
'producer': {'name': 'modelopt', 'version': '0.29.0'},
'quantization': {
'quant_algo': 'FP8',
'kv_cache_quant_algo': 'FP8',
'exclude_modules': ['lm_head'],
},
}
result = convert_hf_quant_config_format(fp8_config)
print('=== FP8 (existing) ===')
print(json.dumps(result, indent=2))
assert result['quant_algo'] == 'FP8'
assert 'group_0' in result['config_groups']
assert result['ignore'] == ['lm_head']
# Test 2: Mixed precision
mixed_config = {
'producer': {'name': 'modelopt', 'version': '0.29.0'},
'quantization': {
'quant_algo': 'MIXED_PRECISION',
'kv_cache_quant_algo': 'FP8',
'quantized_layers': {
'model.layers.0.self_attn.q_proj': {'quant_algo': 'FP8'},
'model.layers.0.self_attn.k_proj': {'quant_algo': 'FP8'},
'model.layers.0.mlp.gate_proj': {'quant_algo': 'NVFP4', 'group_size': 16},
'model.layers.0.mlp.up_proj': {'quant_algo': 'NVFP4', 'group_size': 16},
'model.layers.1.self_attn.q_proj': {'quant_algo': 'FP8'},
},
},
}
result = convert_hf_quant_config_format(mixed_config)
print()
print('=== MIXED_PRECISION ===')
print(json.dumps(result, indent=2))
assert result['quant_algo'] == 'MIXED_PRECISION'
assert 'config_groups' in result
assert 'quantized_layers' in result
# Should have 2 groups: one for FP8 layers, one for NVFP4 layers
assert len(result['config_groups']) == 2
# Verify per-layer detail is preserved
assert 'model.layers.0.self_attn.q_proj' in result['quantized_layers']
assert result['quantized_layers']['model.layers.0.mlp.gate_proj']['quant_algo'] == 'NVFP4'
# Check that FP8 group has correct targets
for gname, gcfg in result['config_groups'].items():
if gcfg.get('weights', {}).get('num_bits') == 8:
assert len(gcfg['targets']) == 3 # 3 FP8 layers
elif gcfg.get('weights', {}).get('num_bits') == 4:
assert len(gcfg['targets']) == 2 # 2 NVFP4 layers
print()
print('All tests passed!')
"
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added support for MIXED_PRECISION quantization configurations in model
export, enabling automatic aggregation of layers by their individual
quantization settings.
* Enhanced quantization config handling to dynamically manage multiple
configuration groups for complex quantization scenarios.
* Maintained backward compatibility with existing quantization
algorithms.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
|
||
|
|
dfe705aab5 |
support static NVFP4 HF export (#858)
## What does this PR do? **Type of change:** ? new feature **Overview:** ? Supports export `NVFP4StaticQantizer` in unified huggingface checkpoint, as a deployment path for PTQ algorithms such as MSE ## Usage <!-- You can potentially add a usage example below. --> ```python # checkpoint generation python examples/llm_ptq/hf_ptq.py --pyt_ckpt_path Qwen/Qwen3-8B --qformat nvfp4_mse --export_path test-Qwen3-8B-Instruct-MSE-FP8-sweep-FP4 --kv_cache_qformat none --trust_remote_code ``` ## Testing Tested generated Qwen3 8B checkpoint with trtllm serve and nv_eval example in `Model-Optimizer-Internal/examples/nv_eval`. NV eval results: ``` | Groups |Version|Filter|n-shot| Metric | |Value | |Stderr| |--------|-------|------|------|-----------|---|-----:|---|-----:| |mmlu_str| |none | |exact_match|↑ |0.7186|± |0.0036| ``` ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for static NVFP4 quantizers that utilize pre-computed calibration scales. * Introduced new NVFP4 W4A4 quantization configuration with optional FP8 scale sweep. * **Performance Improvements** * Static quantizers now skip unnecessary dynamic scaling factor recalculation. * Unified quantization handling for improved consistency and efficiency. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com> |
||
|
|
3fe7e65b70 |
[5615343,5597780,5371126] Upgrade ORT to 1.24 (#928)
## What does this PR do? **Type of change:** Bug fix **Overview:** Upgrade ORT to 1.24.x to fix various bugs (5615343, 5597780, 5371126). TODO: Verify no regressions by bumping ORT once @ajrasane is back ## Usage See each bug. ## Testing See each bug. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Bug Fixes * Upgraded ONNX Runtime to version 1.24.2, addressing multiple reported issues * Updated system requirements and installation documentation to reflect the new ONNX Runtime version <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com> Signed-off-by: Gwena Cunha <4861122+gcunhase@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
4a11486f6c |
Enable multinode training for HF speculative decoding (#922)
## What does this PR do? **Type of change:** New example **Overview:** Modify launch_train.sh script to enable multi-node training. Provide a slurm template script. ## Usage Add required fields in slurm.sh. Then use the command below to submit multi-node job: ```bash bash slurm.sh ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added multi-node training support with new configuration options for distributed setups. * Introduced Slurm batch script for streamlined job submission to cluster environments. * **Improvements** * Enhanced GPU resource management for multi-GPU and distributed training configurations. * Updated speculative decoding model support and validation handling. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> |
||
|
|
ef5a2dfc5d |
feat: Baseten contrib third-party-dataset support (#851)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** ? Background: the outage on `cnn_dailymail` has a couple of issues with out platform, with 20+ customers asking about breaking support for own datasets. - We still require trt-engines to support, which means e.g. the modelopt version cannot be upgraded to latest easily. - Applying the chat template on short context or out-of-distribution material leads to a stronger degradation on harder or generation tasks. This is visible when e.g. running bfcl benchmark, where the model is generating a lot of tokens, the degration with modelopt is large, e.g. 5%+ tokens. Going forward, we are looking to get stable support for customer-brought datasets into modelopt, because the existing options do not suffice. A hard validation is not desirable, many baseten users will see things like `abisee/cnn_dailymail` is not in supported datasets, because we use modelopt 0.35.x. ## Usage <!-- You can potentially add a usage example below. --> `llm_ptq --dataset baseten/quant_calibration_dataset_v1` ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Extended dataset loading to support third‑party and custom datasets, including message-style, prompt, and text formats. * Tokenizer-aware processing and an option to apply a chat template for message-style datasets. * Data-loading pipeline now propagates tokenizer context through sample retrieval. * **Bug Fixes / Improvements** * Improved warnings and validation for datasets that lack expected structures or chat-template support. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: michaelfeil <63565275+michaelfeil@users.noreply.github.com> |
||
|
|
6b0ea4d43d |
Fix MOE layer sync test (#936)
## What does this PR do? **Type of change:** Bug Fix **Overview:** ? Fix MOE layer sync test by initializing weights in MOE layer differently [Link to bug](https://github.com/NVIDIA/Model-Optimizer/actions/runs/22288772586/job/64472124733#step:7:851) ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> https://github.com/NVIDIA/Model-Optimizer/actions/runs/22414958311 ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Tests** * Enhanced quantization testing with improved expert weight initialization patterns for expert-based models. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> |
||
|
|
4eacb0da72 |
Add trust_remote_code cli option for mbridge distillation (#934)
## What does this PR do? Nemotron2/3 need `AutoBridge(..., trust_remote_code=True)` which was missing previously ## Testing <!-- Mention how have you tested your change if applicable. --> Nemotron-nano-v2 can be distilled using tokenized Nemotron-Pretraining-SFT-v1 data Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
e589ac80f6 |
Integrate Automated QDQ benchmark - part 3.1 (#837)
## What does this PR do? This PR integrates benchmark module to QDQ autotunner. This benchamrk module is used to evaluate ONNX model perf. This PR is 1/3 of https://github.com/NVIDIA/Model-Optimizer/pull/703. Once all small PRs are merged #703 could be closed. PR 3.1: https://github.com/NVIDIA/Model-Optimizer/pull/837 PR 3.2 https://github.com/NVIDIA/Model-Optimizer/pull/838 PR 3.3: https://github.com/NVIDIA/Model-Optimizer/pull/839 ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No, document will be added in part 4. - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No, change log will be updated when all changes are merged. ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added ONNX quantization autotuning capabilities with a consolidated module providing streamlined import paths for core components. * Introduced unified benchmarking framework supporting TensorRT-based model evaluation with both command-line and Python API implementations. * Added support for timing cache persistence, custom plugin libraries, shape validation, and dynamic input shape configuration for flexible model testing and optimization. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Will Guo <willg@nvidia.com> |
||
|
|
d78797b466 |
Update the LTX2 API calls during the calibration (#926)
## What does this PR do? **Type of change:** Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Update LTX-2 integration to match latest upstream API 1. The LTX-2 codebase removed/replaced several APIs. This MR updates all affected files: 2. Replace cfg_guidance_scale with MultiModalGuiderParams: The pipeline __call__ no longer accepts a single cfg_guidance_scale float. It now requires two MultiModalGuiderParams objects (video_guider_params and audio_guider_params) that control CFG, STG, rescale, cross-modality guidance, and skip-step settings. Updated in ltx-2.py, ltx-2-fp8.py, ltx-2-onestage.py, calibration.py, and models_utils.py. 3. Replace fp8transformer with QuantizationPolicy: The TI2VidTwoStagesPipeline constructor no longer accepts the fp8transformer boolean flag. FP8 quantization is now configured via quantization=QuantizationPolicy.fp8_cast(). Updated in ltx-2-fp8.py and pipeline_manager.py (with backwards-compatible support for the old --extra-param fp8transformer=true CLI flag). 4. Remove DEFAULT_CFG_GUIDANCE_SCALE constant: Replaced by DEFAULT_VIDEO_GUIDER_PARAMS and DEFAULT_AUDIO_GUIDER_PARAMS in all import sites. ## Usage <!-- You can potentially add a usage example below. --> ```bash python quantize.py --model ltx-2 --format fp4 --batch-size 1 --calib-size 1 --n-steps 40 --extra-param checkpoint_path=./ltx-2-19b-dev-fp8.safetensors --extra-param distilled_lora_path=./ltx-2-19b-distilled-lora-384.safetensors --extra-param spatial_upsampler_path=./ltx-2-spatial-upscaler-x2-1.0.safetensors --extra-param gemma_root=./gemma-3-12b-it-qat-q4_0-unquantized --extra-param fp8transformer=true --hf-ckpt-dir ./ltx2-nvfp4 ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Updates** * Default resolution for LTX2 models adjusted to 768x1280 * Guidance parameter configuration updated for video and audio pipelines * FP8 quantization parameter handling refined <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
03a1899dda |
Support force tokens to % of total experts during calibration (#910)
## What does this PR do?
**Type of change:** New feature
**Overview:** Adds a configurable `moe_calib_experts_ratio` parameter
that controls the percentage of experts to calibrate during the forward
pass in MoE (Mixture of Experts) models. Previously, the calibration
forward always routed tokens to **all** experts, which is expensive.
This PR allows the user to specify a ratio (default: still all experts
so no behavior change) to improve expert calibration coverage without
the cost of a full-expert forward. The token counting for the expert
coverage table now tracks the calibration routing and runs on CUDA for
efficiency.
**Changes include:**
- New `moe_calib_experts_ratio` field in `QuantizeAlgorithmConfig`
(`config.py`)
- Propagation of the ratio from the algorithm config to MoE modules
during calibration (`mode.py`)
- Updated `_QuantSparseMoe.forward` to use the configurable ratio
instead of hard-coding all experts (`huggingface.py`)
- New `--moe_calib_experts_ratio` CLI flag in `hf_ptq.py` (default
`0.25`)
- Moved `expert_token_count` tensor to CUDA and updated the HTML table
title in `moe_utils.py`
## Usage
Via hf_ptq.py CLI — calibrate 50% of experts during MoE calibration
python hf_ptq.py --model <model> --qformat int4_awq
--moe_calib_experts_ratio 0.5
Via Python API — pass the ratio through the algorithm config
import modelopt.torch.quantization as mtq
quant_cfg = {
"quant_cfg": { ... },
"algorithm": {
"method": "awq_lite",
"moe_calib_experts_ratio": 0.25, # calibrate 1/4 of experts
},
}
mtq.quantize(model, quant_cfg, forward_loop=calib_loop)
## Testing
Test with Qwen3 30B A3B calibration and check the tokens per expert.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added support for configurable expert calibration during Mixture of
Experts (MOE) model quantization. Users can now specify the percentage
of experts to include during calibration, enabling better expert
coverage and improved quantization accuracy for MOE models. Default: 25%
of all experts.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: realAsma <86726418+realAsma@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
|
||
|
|
75b5da9b83 |
Fix: quant config error on quantized offline eagle (#925)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** ? ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Refactor** * Enhanced quantization configuration handling for transformer models through improved type validation, ensuring more robust processing of quantized model configurations. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
2802302bf9 |
SpecDec Bench: February Update (#875)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Addition of SpecBench Dataset Addition of NVIDID SPEED-Bench dataset, preproc scripts, and custom metrics aggregator Addition of example of converting SpecBench Medusa to this FW Addition of Initial TRTLLM AutoDeploy Specdec support Updates to all frameworks for better perf (overlap/async scheduling etc) ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added SPEED-Bench dataset support with configurable throughput and qualitative configurations * Introduced SpecBench metrics with acceptance rate analysis and visualizations * Added progress bar during benchmark execution * New model implementations for auto-deployment and Medusa-style speculative decoding * Data preparation utility for benchmark datasets * Enhanced metrics with per-category analysis and performance charts * **Documentation** * Updated README with SPEED-Bench workflow and examples * New porting guide for integrating custom benchmark runners * **Refactor** * Streamlined model and runner interfaces for improved flexibility * Consolidated dataset implementations and removed deprecated base classes * **Chores** * Added required dependencies for data handling and visualizations <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Izzy Putterman <iputterman@nvidia.com> |
||
|
|
c689ea1176 |
flush print megatron tokenization stats and update readme (#927)
## What does this PR do? When running the script, I often see the print stats for tokenization (every `log_interval`) not showing up or showing up very very delayed. Hence using `print(..., flush=True)` to fix this. Also update README that the example shown for tokenization takes too long to run, split into multiple .jsonl files for efficiently running the tokenization; and try out a smaller dataset first to test the script ## Testing <!-- Mention how have you tested your change if applicable. --> Split Nemotron-pretraining-SFT-v1 dataset into multiple .jsonl splits and then tokenize them parallelly in different slurm jobs. Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
52e662dd1d |
Fix test_transformers_tp for torch 2.10 env (#915)
After bumping CICD dev containers to latest (with torch 2.10), `test_transformers_tp.py` is failing (was skpipped in PR-merge CICD as it requires 2-gpu) Failing test: https://github.com/NVIDIA/Model-Optimizer/actions/runs/22258743173/job/64393623736#step:7:617 Passing test after this fix: https://github.com/NVIDIA/Model-Optimizer/actions/runs/22259791793/job/64396179609 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved quantization calibration by converting outputs to local tensor representations and adding a normalization step before loss computation, ensuring more reliable and accurate model calibration results. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
f78385e70e |
Improve megatron dataset preprocessing script and update docs (#918)
## What does this PR do?
Improve megatron dataset preprocessing script and update docs
## Usage
<!-- You can potentially add a usage example below. -->
```python
python -m modelopt.torch.utils.plugins.megatron_preprocess_data \
--hf_dataset nvidia/Nemotron-Pretraining-SFT-v1 \
--hf_name Nemotron-SFT-General \
--hf_split train \
--hf_max_samples_per_split 10_000_000 \
--json_keys text \
--tokenizer Qwen/Qwen3-0.6B \
--output_dir /path/to/tokenized/data/qwen3 \
--workers 32 \
--max_sequence_length 256_000
```
```python
python -m modelopt.torch.utils.plugins.megatron_preprocess_data \
--jsonl_paths /path/to/data1.jsonl /path/to/data2.jsonl ... \
--json_keys text \
--tokenizer Qwen/Qwen3-0.6B \
--output_dir /path/to/tokenized/data/qwen3 \
--workers 32 \
--max_sequence_length 256_000
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
- Downloaded and tokenized Nemotron-Pretraining-SFT-v1 with
Nemotron-Nano-v2 tokenizer
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Updated data preparation guides with new CLI patterns and Hugging Face
Hub integration instructions.
* **New Features**
* Added batch tokenization via directory input and direct Hugging Face
dataset downloads with flexible subset/split filtering.
* **Configuration Updates**
* Optimized distillation settings: adjusted optimizer parameters and
increased checkpoint retention.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
df47e816a7 |
Added support to rotate in fp32 (optional) (#885)
## What does this PR do?
**Type of change:** New Feature
**Overview:**
This MR adds support to perform rotation for RHT in float32 if enabled
by quantization configuration. It also makes rotate argument in
quantization configuration of type bool (for backward compatibility) or
dict (added option for float32 rotation)
## Usage
```
python hf_ptq.py --pyt_ckpt_path meta-llama/Llama-3.2-3B-Instruct --qformat nvfp4 --export_fmt hf --dataset cnn_dailymail --export_path test --trust_remote_code --inference_pipeline_parallel 1 --batch_size 1 --calib_size 4 --kv_cache_qformat nvfp4_rotate
```
Updated `NVFP4_KV_ROTATE_CFG` locally with `"rotate": {"enable": True,
"rotate_fp32": True}`
```
...
model.layers.27.self_attn.k_bmm_quantizer TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=8.3750 rotated (fp32) calibrator
=MaxCalibrator quant)
...
```
Updated `NVFP4_KV_ROTATE_CFG` locally with `"rotate": {"enable": True,
"rotate_fp32": False}`
```
model.layers.27.self_attn.k_bmm_quantizer TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=8.3750 rotated calibrator=MaxCalibrator quant)
```
## Testing
Updated unit test in `tests/gpu/torch/quantization/test_hadamard.py`
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No (updated existing test)
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added rotational input capability prior to quantization for RHT
(Rotated Hyperplane Transform).
* Introduced granular rotation configuration options enabling FP32
casting for improved numerical stability during transforms.
* **Tests**
* Expanded test coverage for rotation functionality with parameterized
FP32 casting scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
|
||
|
|
02fa3623d8 |
Sync MOE layer input quantizer only (#903)
## What does this PR do? **Type of change:** Bug fix **Overview:** in MOE layer we currently sync both the weight and input quantizers so that all experts have the same weight amaxes and activation amaxes. VLLM/TRTLLM actually support non-uniform weight amaxes in MOE so we only need to sync the activation amaxes. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved input quantizer synchronization for Mixture of Experts models to ensure correct amax value handling across local experts. * **Documentation** * Fixed typos and clarified wording in quantization documentation. * **Tests** * Added test coverage for Mixture of Experts quantizer synchronization functionality. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> |
||
|
|
70ffb6f87c |
Remove test_llama_eval_sparse_attention (#914)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Bug fix **Overview:** ? Fix the example test. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Removed a sparse attention test variant to streamline test coverage. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Kai Xu <kaix@nvidia.com> |
||
|
|
9e23c6c312 |
Upgrade Dev containers for CICD to latest (#891)
## What does this PR do? - Upgrade CICD test containers to latest - Enable torch 2.10 testing in CICD ## Testing <!-- Mention how have you tested your change if applicable. --> CI/CD in this PR should pass <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for mixed-precision gradient handling with FSDP2. * **Documentation** * Updated Linux installation guide with CUDA 13.x support and cupy dependency guidance. * **Chores** * Updated CI/CD workflows and test infrastructure to support PyTorch 2.10 and CUDA 13. * Updated container image versions and test environment configurations. * Updated TensorRT-LLM version requirements. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
9975ba1065 |
Fix DeepSeek PTQ script (#912)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** ? Fix two bugs in the PTQ script ## Testing Run DeepseekV3.2 PTQ and export <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Enhanced data type handling in quantization examples for bf16 operations * Updated internal dependencies for quantization utilities to improve modularity <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
7c4c9fdbc9 |
Support multiple-batch input for autocast calibration. (#760)
## What does this PR do? Add multi-batch calibration data support for autocast precision conversion. This enhancement allows users to provide multiple batches of calibration data (via a directory of NPZ files or Polygraphy JSON with multiple batches) to aggregate tensor statistics across batches, resulting in more robust precision conversion decisions. ## Usage ### Single NPZ file (existing behavior) ``` python -m modelopt.onnx.autocast --onnx_path model.onnx --calibration_data calibration_data.npz --output_path model_fp16.onnx ``` ### Directory containing multiple NPZ files for multi-batch calibration (new) ``` python -m modelopt.onnx.autocast --onnx_path model.onnx --calibration_data calibration_data_dir/ --output_path model_fp16.onnx ``` ## Testing - Tested with single NPZ file to ensure backward compatibility - Tested with directory containing multiple NPZ files for multi-batch calibration - Verified that aggregated statistics (absmax, min, max) are correctly computed across batches ## Before your PR is "*Ready for review*" - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No ## Additional Information Key changes: - Added `TensorStats` dataclass to store aggregated tensor statistics (absmax, min_val, max_val, shape) - Updated `ReferenceRunner` to: - Load multiple NPZ files from a directory (`_load_inputs_from_npz`) - Aggregate statistics across batches (`_aggregate_tensor_stats`) - Process multi-batch inference in `run()` method - Updated `IORangeRule` and `DepthOfReductionRule` to handle both raw numpy arrays and `TensorStats` objects - Enhanced `--calibration_data` CLI help text to document multi-batch support <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added multi-batch calibration support via directories of NPZ files or Polygraphy JSON files. * Implemented cross-batch statistics aggregation for more robust precision conversion decisions. * **Documentation** * Expanded calibration_data CLI option guidance with detailed support for multiple input formats and batch processing benefits. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Tony Yin <toyin@nvidia.com> |
||
|
|
adcce614cb |
add local hessian calibration (#788)
## What does this PR do? **Type of change:** new feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Add a new calibration method for weight scale search. It considers activation information by weighing scale candidates with local hessian matrix. Initial experiments with Qwen3 8B NVFP4 shows improvements. ## Usage <!-- You can potentially add a usage example below. --> Use `NVFP4_W4A4_WEIGHT_LOCAL_HESSIAN_CFG` quantization config for quantization and evaluation. ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added local Hessian-weighted MSE calibration pathway for NVFP4 per-block quantization with configurable amax search parameters and FP8 scale sweep support. * **Tests** * Added test coverage for the new local Hessian weight-only quantization configuration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com> |
||
|
|
ac7c985d96 |
[NVBUG: 5804406] Auto detect MOE layers (#900)
## What does this PR do?
**Type of change:** New feature, new tests
**Overview:** Replace hardcoded per-model MoE class registrations
(Mixtral, Qwen2Moe, Qwen3Moe, Qwen3Next, Llama4TextMoe, Qwen3VLMoe,
MiniMaxM2, etc.) with a single generic auto-detection mechanism
(`register_sparse_moe_on_the_fly`) that walks the model tree and
identifies MoE blocks by their structural attributes (`gate` + `experts`
with `top_k`/`num_experts`). This makes MoE quantization
forward-compatible with new HuggingFace MoE architectures without
requiring explicit registration for each model family.
Additionally, this PR:
- Tracks per-expert token routing counts during calibration via a gate
forward hook, enabling visibility into expert utilization.
- Saves an HTML report of expert token counts during export
(`save_expert_token_count_table`), highlighting under-utilized experts.
- Fixes the `topk` -> `top_k` attribute name for transformers >= 5.0
compatibility.
- Also move the ptq summary prints to a file in hf_ptq.py to reduce the
prints
## Usage
Auto-detection is transparent -- no user-facing API changes are needed.
Any HuggingFace MoE model with the standard `gate`/`experts` pattern is
automatically detected and quantized:
import modelopt.torch.quantization as mtq
# Any HuggingFace MoE model (Mixtral, Qwen3Moe, DeepSeek, etc.)
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-30B-A3B")
mtq.quantize(model, mtq.INT8_DEFAULT_CFG, forward_loop)
# During export, an .moe.html report with per-expert token counts is
saved automatically
## Testing
unittest, also test exporting qwen MOE
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added expert token count visualization for Mixture of Experts models,
exported as HTML reports during model export.
* Enhanced sparse MoE quantization with improved calibration-aware
routing and automatic model block detection.
* **Tests**
* Added comprehensive test suite for sparse MoE quantization validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
|
||
|
|
c4b662fbc8 |
[Bug fix] Fake quantized model save after HF accelerate hooks are added (#906)
## What does this PR do? **Type of change:** Bug fix **Overview:** Fix `AttributeError: Can't get local object 'add_hook_to_module.<locals>.new_forward'` when saving a quantized model a second time after restoring it with `device_map="auto"`. When a model is loaded with `device_map="auto"`, accelerate's `add_hook_to_module` patches every submodule (including `TensorQuantizer` instances) and injects three instance attributes: `_hf_hook`, `_old_forward`, and `forward` (a `functools.partial` wrapping a local function). These are not picklable and were leaking into the modelopt state dict collected by `get_modelopt_state()`, causing `torch.save` to fail. This PR adds the three accelerate-injected attributes to `TensorQuantizer._skip_properties_for_save_restore` so they are excluded from the serialized state, matching the existing pattern used for HuggingFace and DeepSpeed attributes. ## Usage ```python mto.enable_huggingface_checkpointing() # Quantize and save model = AutoModelForCausalLM.from_pretrained(name, device_map="auto") model = mtq.quantize(model, mtq.FP8_DEFAULT_CFG, forward_loop=forward_loop) model.save_pretrained(save_dir) # Restore and save again (this previously failed) model2 = AutoModelForCausalLM.from_pretrained(save_dir, device_map="auto") model2.save_pretrained(save_dir_round2) # now works ``` ## Testing - Added unit test `test_tensor_quantizer_modelopt_state_with_accelerate_hook` in `tests/unit/torch/quantization/plugins/test_accelerate.py` that verifies accelerate hook attributes are excluded from modelopt state and the state dict remains picklable. ## Before your PR is "*Ready for review*" - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes — only adds entries to a skip set; existing saved checkpoints are unaffected. - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: No (internal fix, no API change) - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No ## Additional Information The root cause is in accelerate's `add_hook_to_module`, which defines `new_forward` as a local function and binds it via `functools.partial` onto `module.forward`. Since local functions cannot be pickled, any `TensorQuantizer` that has been hooked by accelerate becomes unserializable unless these attributes are excluded. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Enhanced compatibility with accelerate library by excluding framework-specific hooks and attributes from model state serialization, preventing issues during save/restore operations. * **Tests** * Added test to validate that accelerate-related attributes are properly excluded from model state and that the state remains picklable. * **Public API** * TensorQuantizer is now publicly exported. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
eb99488da1 |
Fix: restore requires_grad in transformers5 reloading (#907)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Patch transformers 5.x parameter loading to preserve original `requires_grad` settings. In transformers v5.x, loading a checkpoint forcibly sets parameters' requires_grad, which unintentionally unfreeze frozen parameters (e.g. Base model in eagle training). This leads to optimizer initialization error since the restored optimizer expected more parameter than the checkpoint. This monkey-patch restores the original`requires_grad` after loading parameters. Reference: https://github.com/huggingface/transformers/blob/v5.0.0.rc1-release/src/transformers/core_model_loading.py#L640 ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed model parameter loading in speculative decoding to properly preserve gradient requirements for each parameter when using HuggingFace Transformers 5.x, ensuring correct behavior during checkpoint resumption and model initialization. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
3dd52bf106 |
Diffusion export bug fixed for model_index.json (#901)
## What does this PR do? **Type of change:** Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Updated diffusers export to preserve the original model_index.json instead of always rebuilding a minimal one. The export now uses a simple fallback order: copy original model_index.json from source path if available, otherwise call `pipe.save_config(export_dir)`, and only then generate a minimal model_index.json as last resort. Non-diffusers export behavior is unchanged. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved Diffusers pipeline export with enhanced model configuration handling. The export process now better preserves original pipeline configurations and uses fallback strategies to ensure complete configuration files are generated. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
b8a4586702 |
Refactor: Eagle data loading (#668)
## What does this PR do? **Type of change:** Refactor <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Jira ticket: https://jirasw.nvidia.com/browse/OMNIML-2955 Main changes : - Consolidate Eagle data loading with @ChenhanYu 's implementation of `transformers_dataset.py` - Refactor: baked the following logics from `example/main.py` to `modelopt/torch` for cleaner example entrance: - default config selecting and merging with custom config - tokenizer post-processor (chat template and pad_tok_id) - d2t loading - Implementation refactor: In HF workflow, reuse base modfel's input hidden states as input_embedding, instead of calculating from input_ids. This has two main benefits: - Easier VLM support, which has various embedding processing logics. - Training effieicy. - Deprecating eagle1 from the example. It is still available by setting custom config. - Other minor fixes and readme updates. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> Tested that training curves after changes (both online&offline) is identical with original branch: <img width="1073" height="634" alt="image" src="https://github.com/user-attachments/assets/abfd7bea-c82c-48a7-8181-68c5a9e4da8d" /> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added draft vocabulary cache support for EAGLE model training, enabling runtime vocabulary customization via `--draft_vocab_cache` parameter * Introduced new data loading utilities with sharding, streaming, and tokenization support for large-scale training * Added optional `--log_steps` configuration to training launcher * **Documentation** * Updated EAGLE configuration guides with draft vocabulary cache setup instructions and examples * **Refactor** * Restructured data pipeline for offline training with improved dataset handling and batching * Updated command-line arguments across training scripts (`--input-data` replaces `--input-file`) <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
9e38041d34 |
[OMNIML-2850] [3/n] Adds sparse attention calibration (#538)
## What does this PR do?
**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
new feature
**Overview:** ?
- This PR adds the sparse attention calibration algorithm
- Chunked prefill to support long ctx_len
- Separated calibration for prefill and decode
## Usage
<!-- You can potentially add a usage example below. -->
```python
import modelopt.torch.sparsity.attention_sparsity as mtsa
# Apply sparse attention with calibration
model = mtsa.sparsify(model, config=SKIP_SOFTMAX_CALIB)
# Print summary - now shows actual thresholds
mtsa.print_sparse_attention_summary(model)
# Output:
# Method: flash_skip_softmax, Threshold: Dynamic (λ=437.395926)
# Or llm_eval integration
# HuggingFace sparse attention example
python examples/llm_sparsity/attention_sparsity/hf_sa.py \
--pyt_ckpt_path Qwen/Qwen3-4B \
--sparse_attn skip_softmax_calib
```
# The calibration method
## Calibration Algorithm
- Implemented the Inverse Power model: scale_factor = k / (1 -
sparsity)^p
- Fit model parameters (k, p) per phase using scipy.optimize.curve_fit
- At inference: threshold = k / (1 - target_sparsity)^p / seqlen
## Why Choosing the Inverse Power model?
The inverse power model better fits the relationship between sparsity
ratio and threshold_scale_factor.
<img width="2388" height="1082" alt="sparsity_model_analysis"
src="https://github.com/user-attachments/assets/4dfb45d4-8c16-4f15-a878-c8e08a9b6128"
/>
## Runtime Flexibility
- Target sparsity can be changed at inference time without recalibration
- Users can adjust module._sparse_method_instance.target_sparse_ratio
dynamically
- Threshold automatically adapts to sequence length
## Testing
<!-- Mention how have you tested your change if applicable. -->
The calibration results for `Qwen/Qwen3-30B-A3B-Thinking-2507` are shown
below and are mostly consistent with the ground-truth numbers collected
from the kernel side.
```
Prefill Calibration Results:
Model: scale_factor = k / (1 - sparsity)^p
Fitted k: 1003.3990
Fitted p: 1.2589
R-squared: 0.827549
Scale factors for different target sparsities:
Target Scale Factor
---------- ---------------
50% 2401.35
70% 4568.26
80% 7610.98
90% 18214.70
95% 43591.65
```
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: Kai Xu <kaix@nvidia.com>
|
||
|
|
3801923e9d |
Support MiniMax M2.1 (FP8 checkpoint) (#817)
## What does this PR do? **Type of change:** ? new feature **Overview:** ? Support loading the MiniMax M2.1 (FP8) checkpoint for PTQ. ## Usage scripts/huggingface_example.sh --model <minimax checkpoint> --quant nvfp4 --trust_remote_code ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added MiniMax M2.1 model quantization support with nvfp4 format. * Extended FP8 quantization capabilities with configurable dtype parameter for enhanced precision control. * **Improvements** * Enhanced detection of quantized linear module variants. * Improved weight unpacking for FP8-based linear modules. * **Documentation** * Updated supported models table to include MiniMax M2.1. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
590f9fc662 |
Mamba MOE Quant Configs + Fix Export Bug (#882)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** ? - Fix a bug in MCore export `exclude_modules` where the layers had an extra period at the end - Add custom quant configs for mamba moes ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added four new Mamba MOE quantization configurations: aggressive and conservative variants for both FP8 and NVFP4 quantization schemes, providing enhanced flexibility in quantization options for different use cases. * **Bug Fixes** * Improved quantization export module exclusion pattern handling to properly normalize trailing dots from exclude patterns during export. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> |
||
|
|
9763505981 |
[fix][5889686] AutoCast: Fix logger (#890)
## What does this PR do? **Type of change:** Bug fix **Overview:** Previously relied on quantization logger, which caused logs to be suppressed when onnx.autocast was used directly Instead: - Inherit format and level if called from onnx.quantization - Configure independent format and level if called from onnx.autocast ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved logging configuration to ensure consistent behavior across modules with enhanced file and console output management. * Fixed log file handling to automatically create required directories. * Enhanced logger propagation logic for more reliable output routing. * **Chores** * Refined logging initialization for better automatic configuration at startup. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Gal Hubara Agam <96368689+galagam@users.noreply.github.com> |