### What does this PR do?
- Add experimental support for transformers >=5.0 and remove deprecated
usages:
https://github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md
- ⚠️ For accelerate examples that used `--warmup-ratio: float`
(deprecated in 5.x), we now change it to `--warmup-steps: float | int`
which works as ratio if float but only for 5.x. For 4.x, it will error
out if float and prompt user to change back to `--warmup-ratio` or pass
an int absolute step count.
- ⚠️ Unified Hugging Face checkpoint export for quantized checkpoints
may not work for some models with transformers>=5.0 yet as it requires a
lot of fixes (e.g. change in how MoE experts are organized)
- ~Add Workaround for TRT-LLM's import of deprecated transformers
functions so trt-llm based gpu unit tests work fine. Still deployment
for models needs proper fixes directly in TRT-LLM hence llm/vlm ptq
example tests still run with transformers 4.57~
- Everything except PTQ and Export (mainly MoE) should work fine with
transformers>=5.0
- Bump min torch to 2.8 and enable 2.11 cicd testing
- NOTE: Upcoming Nemo:26.04 container comes with transformers 5.3
### Testing
<!-- Mention how have you tested your change if applicable. -->
- [x] CI/CD tests passing
- [x] Manually tested unit tests, gpu tests with transformers 4.56 and
5.4
- [x] Manually tested example tests (except trt-llm container tests)
with transformers 4.56 and 5.4
- [x] 2-gpu nightly CICD tests manually triggered and passing: [gpu
tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867257540),
[example
tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867260643)
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).
- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Make remote-code usage opt-in via a configurable --trust_remote_code
flag across examples and tools.
* **Bug Fixes**
* Improve checkpoint/resume detection and related training guidance to
avoid erroneous errors.
* **Refactor**
* Consolidate dtype/config naming, switch warmup settings from ratio →
steps, and unify tokenizer invocation patterns.
* **Documentation**
* Simplify changelog title and add misc notes for release 0.44.
* **Chores**
* Remove scheduled PR-branch cleanup workflow and relax/remove several
transformers version pins.
* **Tests**
* Adjust test gates, skips, and structures to align with updated deps
and behaviors.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
### What does this PR do?
- Add secure checkpoint loading support using
`torch.serialization.add_safe_globals([cls])`. This also removes 1
existing pickle usage.
- Remove hard-coded `trust_remote_code=True`
- Replaces https://github.com/NVIDIA/Model-Optimizer/pull/1056 by
@RinZ27
### Testing
<!-- Mention how have you tested your change if applicable. -->
CICD tests ran
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->
### Additional Information
NVBug: 5999336
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added safe checkpoint save/load helpers and a --trust_remote_code CLI
flag in examples to control remote-code loading.
* **Bug Fixes**
* Checkpoint loading now defaults to safer, weights-only semantics to
reduce arbitrary-code exposure.
* **Documentation**
* CHANGELOG updated with security guidance and opt-in procedure for
unsafe checkpoint loading.
* **Tests**
* New unit tests validating the safe-load behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: RinZ27 <222222878+RinZ27@users.noreply.github.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: RinZ27 <222222878+RinZ27@users.noreply.github.com>
### What does this PR do?
#### Summary
Redesigns the `quant_cfg` configuration format in ModelOpt's PyTorch
quantization stack, replacing the previous dict-based format with an
**ordered list of typed `QuantizerCfgEntry` dicts**.
##### Motivation
The old `quant_cfg` dict had several pain points:
- **Ambiguous precedence**: no explicit way to reason about which entry
wins when multiple keys match a quantizer
- **Mixed key namespaces**: wildcard paths and PyTorch class names lived
in the same dict level, requiring ad-hoc dispatch
- **Magic `"default"` key**: an implicit, undocumented catch-all that
was easy to misuse
- **Poor composability**: merging two configs required dict updates that
silently discarded keys
- **No YAML round-trip fidelity**: the nested structure couldn't be
expressed cleanly in YAML
##### New format
`quant_cfg` is now an ordered list of `QuantizerCfgEntry` TypedDicts.
Each entry has:
- `quantizer_name` *(required)*: `fnmatch` wildcard matched against
quantizer module names
- `cfg` *(optional)*: dict (or list of dicts) of
`QuantizerAttributeConfig` fields
- `enable` *(optional)*: toggles quantizer on/off independently of `cfg`
- `parent_class` *(optional)*: restricts match to quantizers whose
parent module is of the given PyTorch class (e.g. `"nn.BatchNorm2d"`)
Entries are applied in list order; later entries override earlier ones.
The canonical pattern is deny-all first (`_base_disable_all`), then
selectively re-enable and configure, then apply standard exclusions
(`_default_disabled_quantizer_cfg`).
##### Changes
**Core library (`modelopt/torch/quantization/`)**
- **`config.py`**:
- Added `QuantizerCfgEntry` TypedDict (line 163) and
`find_quant_cfg_entry_by_path()` helper for exact-match lookup of
entries by path.
- Added `normalize_quant_cfg_list()` (line 1539) that converts legacy
formats (flat dict, single-key dicts, `nn.*`-scoped dicts, `"default"`
key) to canonical `QuantizerCfgEntry` lists. After normalization every
entry is guaranteed to have explicit `quantizer_name`, `enable`, and
`cfg` keys.
- Converted `_default_disabled_quantizer_cfg` and
`_mamba_moe_disabled_quantizer_cfg` from dicts to lists of
`QuantizerCfgEntry`.
- Added `_base_disable_all` (line 205): canonical deny-all entry
(`[{"quantizer_name": "*", "enable": False}]`).
- Converted all ~30 built-in config constants (`INT8_DEFAULT_CFG`,
`FP8_DEFAULT_CFG`, `NVFP4_DEFAULT_CFG`, etc.) to list format using
`*_base_disable_all` and `*_default_disabled_quantizer_cfg` unpacking.
- KV-cache configs (`FP8_KV_CFG`, `NVFP4_KV_CFG`, etc.) are now minimal
lists designed to be concatenated with a primary config — they
intentionally omit `_base_disable_all` and `"algorithm"`.
- Added two `QuantizeConfig` Pydantic field validators: a
`mode="before"` validator that calls `normalize_quant_cfg_list()`, and a
`mode="after"` validator that validates `cfg` dicts against
`QuantizerAttributeConfig`.
- Updated `need_calibration()` to iterate the normalized list instead of
the old dict.
- Changed `QuantizeQuantCfgType` alias from `dict[str | Callable, ...]`
to `list[QuantizerCfgEntry]`.
- **`conversion.py`**:
- Rewrote `set_quantizer_by_cfg()` (line 217) to iterate the list
directly. Each entry's `parent_class` is resolved via
`QuantModuleRegistry[parent_class_name]` (the existing `_DMRegistryCls`
registry).
- Added `set_quantizer_attributes_full()` (line 314): full replacement
of quantizer attributes from a `QuantizerAttributeConfig`. Unspecified
fields revert to defaults, enforcing entry atomicity. Can also upgrade
`TensorQuantizer` → `SequentialQuantizer` or downgrade the reverse.
- Added `set_quantizer_attributes_partial()` (line 384): merges a
partial `dict` of attributes into existing quantizer state. Does NOT
change quantizer structure. Used for enable-only entries.
- Added `set_quantizer_by_cfg_context()` context manager (line 447) that
temporarily applies a `quant_cfg` list and restores original quantizer
state on exit.
- Deprecated `set_quantizer_attribute()` (line 525) with a
`DeprecationWarning` pointing to the new functions.
- **`tensor_quantizer.py`**:
- `TensorQuantizer.set_from_attribute_config()`: narrowed type hint from
`dict` to `dict[str, Any]`.
- Added `_axis_setter` and `_block_sizes_setter` custom setters so that
`axis` and `block_sizes` changes properly propagate to the calibrator
and maintain mutual exclusivity.
- `SequentialQuantizer.set_from_attribute_config()`: narrowed signature
to `list[QuantizerAttributeConfig] | list[dict[str, Any]]` (removed the
old union with single values).
- **`algorithms.py`**:
- Updated `_match_quantizer_cfg()` to iterate the list and return
`(matched_cfg, matched_enable)` tuple with last-match-wins.
- Updated `_cfg_to_dict()`, `estimate_quant_compression()`, and
`QuantRecipe` to work with the list-based format.
- Updated `get_auto_quantize_config()` to emit list-format `quant_cfg`.
- **`model_quant.py`**: `disable_quantizer()` / `enable_quantizer()` now
call `set_quantizer_attributes_partial()` directly instead of the
deprecated `set_quantizer_attribute()`. Updated docstrings and code
examples to show the list format.
- **`utils/core_utils.py`**: `disable_lora_quantizers_in_config()` and
`update_quant_cfg_with_kv_cache_quant()` updated to append
`QuantizerCfgEntry` dicts to the list.
- **Other**: minor updates to `backends/fp8_per_tensor_gemm.py`,
`backends/nvfp4_gemm.py`, `compress.py`, `model_calib.py`,
`export/unified_export_hf.py`, and
`sparsity/attention_sparsity/conversion.py` to use the list format.
- **`onnx/llm_export_utils/quantization_utils.py`**: Updated
quantization config construction to use list format.
**YAML recipes (`modelopt_recipes/`)**
- Converted all 5 general PTQ recipes to the new list format:
- `general/ptq/fp8_default-fp8_kv.yml`
- `general/ptq/nvfp4_default-fp8_kv.yml`
- `general/ptq/nvfp4_experts_only-fp8_kv.yml`
- `general/ptq/nvfp4_mlp_only-fp8_kv.yml`
- `general/ptq/nvfp4_omlp_only-fp8_kv.yml`
- Converted model-specific recipe:
`models/Step3.5-Flash/nvfp4-mlp-only.yaml`
**Documentation (`docs/`)**
- New guide: `docs/source/guides/_quant_cfg.rst` — comprehensive
reference covering entry format, ordering semantics, entry atomicity,
`enable` vs `cfg` independence, `parent_class` filtering, and common
patterns (deny-all-then-enable, customizing a built-in config, building
from scratch).
- Updated `_pytorch_quantization.rst` code examples to show the list
format with `copy.deepcopy` and `.append()`.
- Added `_quant_cfg.rst` to the quantization guide table of contents.
**Examples**
- Updated all quantization examples to use the list format:
`deepseek/ptq.py`, `diffusers/quantization/config.py`,
`llm_ptq/hf_ptq.py`, `llm_qat/main.py`, `vllm_serve/vllm_ptq_utils.py`,
`llm_autodeploy/run_auto_quantize.py`, `llm_eval/quantization_utils.py`,
`llm_ptq/example_utils.py`,
`windows/torch_onnx/diffusers/qad_example/sample_example_qad_diffusers.py`,
and 2 notebooks.
**Tests**
- New test file:
`tests/unit/torch/quantization/test_config_validation.py` — unit tests
for `need_calibration()`, `normalize_quant_cfg_list()` (new format,
legacy format conversions, error cases),
`find_quant_cfg_entry_by_path()`, `_match_quantizer_cfg()`, and
`QuantizeConfig` Pydantic validators.
- Extended `tests/unit/torch/quantization/test_quantize_cpu.py` with
tests for `set_quantizer_attributes_full()` (atomicity, parent_class
filtering, SequentialQuantizer creation), list ordering, enable-only
entry behavior, and end-to-end legacy dict format.
- Updated 20+ existing test files across `tests/unit/`, `tests/gpu/`,
`tests/gpu_megatron/`, and `tests/_test_utils/` to use the list format.
##### Backward compatibility
`normalize_quant_cfg_list()` is called automatically by the
`QuantizeConfig` Pydantic `mode="before"` validator, so existing code
passing the old dict-based format (flat dict like `{"*weight_quantizer":
{"num_bits": 8}}`, single-key dict lists, or `nn.*`-scoped dicts with
`parent_class` semantics) continues to work without modification. The
legacy `"default"` key is converted to `quantizer_name: "*"`.
`set_quantizer_attribute()` is preserved as a deprecated wrapper around
`set_quantizer_attributes_partial()`.
#### Test coverage
- **Unit tests**: new `test_config_validation.py` with tests for
normalization, validation, path lookup, and cfg matching. Extended
`test_quantize_cpu.py` with tests for full/partial attribute setting,
ordering, atomicity, and legacy backward compatibility.
- **System testing**:
```
python examples/llm_ptq/hf_ptq.py \
--model Qwen/Qwen3-8B \
--recipe general/ptq/fp8_default-fp8_kv \
--export_path=build/fp8_default-fp8_kv42 \
--calib_size=16 \
--batch_size=0 \
--trust_remote_code \
--export_fmt=hf
```
### Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
### What does this PR do?
Type of change: ? bugfix
<!-- Details about the change. -->
1. example_utils.py — _extract_layer_prefixes
Added extraction of the top-level MTP module prefix (e.g., mtp from
mtp.fc.weight). Previously it only captured mtp.layers.0 because it
looked exclusively for *.layers.<digit> patterns.
Now it returns {"mtp", "mtp.layers.0"} instead of just {"mtp.layers.0"}.
2. hf_ptq.py — quantization exclusion pattern
Updated pattern construction to handle single-component prefixes:
mtp.layers.0 (multi-component) → *layers.0* (same as before)
mtp (single-component) → *mtp* (new — covers mtp.fc, mtp.norm, etc.)
### Usage
```python
# Add a code snippet demonstrating how to use this
```
### Testing
<!-- Mention how have you tested your change if applicable. -->
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved quantization exclusion so all relevant model parameters
(including non-layer top-level components) are correctly identified and
skipped during quantization.
* Broadened exclusion matching to use full-prefix patterns, reducing
false positives/negatives when excluding parameters from optimization.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
### What does this PR do?
Type of change: Bug fix
Pass `include_buffers=True` to `init_empty_weights()` when computing the
device map for HuggingFace models. Gemma-4 registers its model
parameters as buffers rather than parameters, so without this flag they
are not accounted for during device map computation, causing incorrect
placement or OOM errors.
### Usage
```python
# No API change — fix is internal to get_model() in example_utils.py
# Simply run HF PTQ with a Gemma-4 model as usual:
python hf_ptq.py --pyt_ckpt_path google/gemma-4-... --quantize ...
```
### Testing
Manually tested with Gemma-4 model loading via `hf_ptq.py`.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
### Additional Information
Gemma-4 uses buffers instead of parameters for some model weights,
requiring `include_buffers=True` for correct device map estimation with
`accelerate`'s `init_empty_weights`.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved model initialization in quantization examples by ensuring
buffers are properly included during temporary model construction,
resulting in more accurate device mapping inference for model
optimization.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: James Shen <yueshen@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
### What does this PR do?
- Remove `examples/nemo_run` and other deprecated Nemo 2.0 references
- Add Megatron-Bridge example links where missing
<!-- Details about the change. -->
### Testing
<!-- Mention how have you tested your change if applicable. -->
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!---
Mandatory -->
- Did you write any new necessary tests?: N/A <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Deprecations**
* Removed deprecated NeMo 2.0 support and related example flows and
utilities.
* **Documentation**
* Updated docs and examples to emphasize Megatron-Bridge / Megatron-LM
and refreshed technique/deployment guidance and links.
* **New Features**
* Added CLI options for additional parallelism (context/expert
tensor/expert model) in Megatron-Bridge distillation.
* **Chores**
* Removed legacy CI configs and refreshed container image tags across
examples and docs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
### What does this PR do?
1. start a new config system using yaml/yml files.
2. add a new top level package: modelopt_recipes
I want it to be a top level package so we can make it clear that the
modelopt package holds the code, this new package holds recipes
3. implement some of the existing quantization recipes using the new
config system as model agnostic general recipes, but not actually in
use. these recipes sit inside modelopt_recipes/general/ptq/...
4. make sure the configs from the new config system match the exisiting
configs
5. extend the hf_ptq script to enable recipe based PTQ
8. testted hf_ptq using both builtin and extenal config file. example
script:
### Usage
```bash
python examples/llm_ptq/hf_ptq.py \
--model Qwen/Qwen3-8B \
--recipe general/ptq/fp8_default-fp8_kv \
...
```
### Testing
```bash
python examples/llm_ptq/hf_ptq.py \
--model Qwen/Qwen3-8B \
--recipe general/ptq/fp8_default-fp8_kv \
--export_path=fp8_default-fp8_kv \
--calib_size=16 \
--batch_size=0 \
--trust_remote_code \
--export_fmt=hf
```
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Recipe-driven PTQ workflows via YAML recipes and new recipe loader;
CLI gains a --recipe option and --pyt_ckpt_path renamed to --model.
* Many new PTQ recipe and config presets (FP8, INT4/INT8, NVFP4, MXFPx,
KV-cache variants) and improved runtime config loading/merging.
* **Documentation**
* Added READMEs describing recipe/config layout.
* **Tests**
* New unit tests covering config loading, inheritance and recipe
loading.
* **Chores**
* Added YAML/OmegaConf runtime support and packaging of recipe YAMLs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
## What does this PR do?
**Type of change:** New model support <!-- Use one of the following: Bug
fix, new feature, new example, new tests, documentation. -->
**Overview:** Support Kimi-K2.5 PTQ.
## Usage
<!-- You can potentially add a usage example below. -->
```python
python3 hf_ptq.py --pyt_ckpt_path moonshotai/Kimi-K2.5 --qformat nvfp4_mlp_only --export_path ./kimi-k2.5-nvfp4 --trust_remote_code
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
You may need `pip install transformers==4.57.1` and the model file here:
https://huggingface.co/nvidia/Kimi-K2.5-NVFP4/blob/main/modeling_kimi_k25.py
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Fixes model loading so mixed-precision (BF16) and expert weights are
correctly restored and stale placeholders removed.
* Adds error handling to avoid failures when optional decompression
components are missing.
* **New Features**
* Adds conditional patch/restore around pack‑quantized model loads,
final unpacking of weights after load, and on‑the‑fly decompression
during inference.
* Improves logging for load and decompression events.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Zhiyu <zhiyuc@nvidia.com>
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
### What does this PR do?
Type of change: ? new feature
1) Add vfp4_omlp_only config == nvfp4_mlp_only + o_proj quant
2) Add block sparse MOE to mlp only config
3) Simplfiy config.py
4) Update readme in llm_ptq mention these two configs for better
accuracy.
### Usage
huggingface_script.sh ... --quant nvfp4_omlp_only
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added nvfp4_omlp_only quantization format for NVFP4, enabling
selective quantization of MLP and output projection layers while
preserving attention QKV projection accuracy.
* **Changed**
* pass_through_bwd now defaults to True; set to False if using STE with
zeroed outlier gradients for better QAT accuracy.
* **Documentation**
* Updated post-training quantization guidance with NVFP4-specific
configuration recommendations and usage examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
## What does this PR do?
**Type of change:** New feature
**Overview:** Adds a configurable `moe_calib_experts_ratio` parameter
that controls the percentage of experts to calibrate during the forward
pass in MoE (Mixture of Experts) models. Previously, the calibration
forward always routed tokens to **all** experts, which is expensive.
This PR allows the user to specify a ratio (default: still all experts
so no behavior change) to improve expert calibration coverage without
the cost of a full-expert forward. The token counting for the expert
coverage table now tracks the calibration routing and runs on CUDA for
efficiency.
**Changes include:**
- New `moe_calib_experts_ratio` field in `QuantizeAlgorithmConfig`
(`config.py`)
- Propagation of the ratio from the algorithm config to MoE modules
during calibration (`mode.py`)
- Updated `_QuantSparseMoe.forward` to use the configurable ratio
instead of hard-coding all experts (`huggingface.py`)
- New `--moe_calib_experts_ratio` CLI flag in `hf_ptq.py` (default
`0.25`)
- Moved `expert_token_count` tensor to CUDA and updated the HTML table
title in `moe_utils.py`
## Usage
Via hf_ptq.py CLI — calibrate 50% of experts during MoE calibration
python hf_ptq.py --model <model> --qformat int4_awq
--moe_calib_experts_ratio 0.5
Via Python API — pass the ratio through the algorithm config
import modelopt.torch.quantization as mtq
quant_cfg = {
"quant_cfg": { ... },
"algorithm": {
"method": "awq_lite",
"moe_calib_experts_ratio": 0.25, # calibrate 1/4 of experts
},
}
mtq.quantize(model, quant_cfg, forward_loop=calib_loop)
## Testing
Test with Qwen3 30B A3B calibration and check the tokens per expert.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added support for configurable expert calibration during Mixture of
Experts (MOE) model quantization. Users can now specify the percentage
of experts to include during calibration, enabling better expert
coverage and improved quantization accuracy for MOE models. Default: 25%
of all experts.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: realAsma <86726418+realAsma@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
## What does this PR do?
**Type of change:** ? new feature
**Overview:** ?
Support loading the MiniMax M2.1 (FP8) checkpoint for PTQ.
## Usage
scripts/huggingface_example.sh --model <minimax checkpoint> --quant
nvfp4 --trust_remote_code
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added MiniMax M2.1 model quantization support with nvfp4 format.
* Extended FP8 quantization capabilities with configurable dtype
parameter for enhanced precision control.
* **Improvements**
* Enhanced detection of quantized linear module variants.
* Improved weight unpacking for FP8-based linear modules.
* **Documentation**
* Updated supported models table to include MiniMax M2.1.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
## What does this PR do?
**Type of change:** New model support <!-- Use one of the following: Bug
fix, new feature, new example, new tests, documentation. -->
**Overview:** Add PTQ support for
https://huggingface.co/nvidia/NVIDIA-Nemotron-Parse-v1.1
## Usage
<!-- You can potentially add a usage example below. -->
```python
python3 hf_ptq.py --pyt_ckpt_path /home/omniml_data_3/models/NVIDIA-Nemotron-Parse-v1.1 --qformat fp8 --export_path /home/omniml_data_3/zhiyuc/checkpoints/NVIDIA-Nemotron-Parse-v1.1-FP8 --trust_remote_code --kv_cache_qformat none --attn_implementation eager
```
By default, image-text data will be used in calibration for VLMs.
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Not yet <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added support for Nemotron-Parse multimodal models, including proper
device mapping, processor loading, and generation handling.
* **Improvements**
* Enhanced quantization robustness with safer handling of quantization
attributes and fallback logic.
* Improved model loading with better device placement and encoder buffer
management for vision-language models.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
## What does this PR do?
Fix MTP export for Qwen3 Next
**Overview:** ?
For Qwen3 next, the MTP weights are not stored separately in
safetensors. So we use "mtp" weights key to decide if the weights are
for MTP or not.
## Testing
Qwen3 Next PTQ and check if MTP is in the exported checkpoint.
scripts/huggingface_example.sh --model
<Qwen3-Next-80B-A3B-Instruct/Thinking> --quant nvfp4 --trust_remote_code
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Refactor**
* Optimized Multi-Token Prediction weight loading with improved layer
detection and handling.
* **Chores**
* Simplified status reporting to display total loaded weights and
detected layers.
* Removed verbose per-file warnings for cleaner console output.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Zhiyu <zhiyuc@nvidia.com>
## What does this PR do?
**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
**Overview:** Enable GLM-4.7 PTQ workflow, including loading the
standalone MTP modules and export as-is.
## Usage
<!-- You can potentially add a usage example below. -->
```python
python3 hf_ptq.py --pyt_ckpt_path /home/omniml_data_3/models/GLM-4.7 --qformat nvfp4_mlp_only --export_path /home/omniml_data_3/zhiyuc/checkpoints/GLM-4.7-NVFP4-0203 --trust_remote_code
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added quantization support for GLM-4.7 model with automatic handling
of specialized layer architecture.
* Added image-text data calibration capabilities for Nemotron VL model
quantization.
* **Documentation**
* Updated support matrix to reflect newly supported models and
quantization features.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
## What does this PR do?
**Type of change:** New feature
**Overview:**
The primary goal of this PR is to allow the model optimizer to use
image-text pair data during the calibration phase of quantization, which
is likely help improve accuracy of quantized VLMs like Nemotron VL on
visual understanding tasks particularly, compared to text-only
calibration data.
- New Feature: Adds support for VLM calibration specifically using
image-text data.
- Dataset Integration: Introduces support for sampling from the
`Nemotron-VLM-Dataset-v2`.
- Refactoring: Created a separate utility for VLM datasets to keep the
main Hugging Face PTQ script (`hf_ptq.py`) clean.
- Simplified logic for handling multimodal inputs.
- Addressed specific issues encountered when calibrating the
`Nemotron-Nano-VL-12B-V2` model with image data.
- Documentation: Updated the README to include instructions and examples
for VLM calibration.
This PR complements https://github.com/NVIDIA/Model-Optimizer/pull/347
and we will consolidate llm_ptq and vlm_ptq examples in follow-up PRs.
## Usage
<!-- You can potentially add a usage example below. -->
```python
python3 hf_ptq.py --pyt_ckpt_path /home/scratch.omniml_data_2/models/Nemotron-Nano-VL-12B-V2 --qformat nvfp4 --export_path /home/omniml_data_3/zhiyuc/checkpoints/Nemotron-Nano-VL-12B-V2-NVFP4-doccalib --trust_remote_code --kv_cache_qformat none --calib_with_images --vlm_dataset nemotron_vlm_dataset_v2 --vlm_subsets sparsetables,plotqa_cot --calib_size 512
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Not yet <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added Vision-Language Model (VLM) calibration support with image-text
pair data, specifically for Nemotron VL models.
* Added new `--calib_with_images` CLI flag to enable image-based
calibration workflows.
* Integrated Nemotron VLM dataset v2 for streaming multimodal
calibration data.
* **Documentation**
* Added VLM calibration guidance in the PTQ README with usage examples
and dataset information.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
## What does this PR do?
**Type of change:** ? new feature
**Overview:**
Support KIMI K2 Thinking PTQ from the original int4 checkpoint.
Tested with transformers 4.57.1, compressed-tensors 0.12.0
The model weights are dequantized on the fly to save GPU memory
## Usage
scripts/huggingface_example.sh --model <Kimi-K2-Thinking ckpt> --quant
nvfp4_mlp_only --trust_remote_code
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Support for nvfp4_mlp_only quantization format, enabling new
layer-wise quantization options
* Quantization support for CompressedLinear layers in quantized models
* **Improvements**
* Enhanced quantization for DeepSeek models with improved attention
configuration handling
* Optimized model loading with automatic precision configuration and
weight unpacking
* Better memory management during model export with automatic cache
cleanup
* Conditional sample generation output controlled via verbose mode
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
## What does this PR do?
**Type of change:** Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
**Overview:** ?
## Usage
<!-- You can potentially add a usage example below. -->
```python
# Add a code snippet demonstrating how to use this
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
## What does this PR do?
**Type of change:** ? Bug fix
**Overview:** ?
For Deepseek, let's force the user to apply trust_remote_code and use
AutoModelForCausalLM for loading the model.
## Testing
python hf_ptq.py --pyt_ckpt_path <Kimi-K2-Thinking_path> --qformat nvfp4
--export_path <quantized_ckpt> --kv_cache_qformat none --calib_size 64
--trust_remote_code --dataset cnn_dailymail
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
## What does this PR do?
**Type of change:** ? Recipe improvement
**Overview:** ?
Disable QKV NVFP4 quantization for Qwen3 MOE models following the Qwen3
Next recipe for accuracy recovery
## Testing
Model accuracy benchmarking
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
## What does this PR do?
Refactor and clean up hf_ptq.py
This script has several separate logic and the code of them are
entangled, making it really hard to add new features
Refactor them so that we separate these logics:
1. sparsity, all logic go to sparsity_main. TODO: we may actually move
this logic out to a separate script
2. quantize, all logic go to quantize_main.
2.1 plain quantization with a single quantization format
2.2 auto quantization
In the quantization pipeline, separate the pipeline to:
1. model loading
2. calibrate dataset loading
3. pre-quantize processing
4. actual quantize
5. post-quantize processing
6. quantized model export
## Testing
tested the mono quantization:
```
python examples/llm_ptq/hf_ptq.py \
--pyt_ckpt_path=Qwen/Qwen3-8B \
--export_path=qwen3-8B_fp8 \
--qformat=fp8 \
--kv_cache_qformat=fp8 \
--calib_size=16 \
--batch_size=0 \
--trust_remote_code \
--export_fmt=hf
```
and deployed to vLLM and TRTLLM, validated accuracy using lm_eval
tested auto quantize:
```
python examples/llm_ptq/hf_ptq.py \
--qformat=nvfp4,fp8 \
--auto_quantize_score_size 128 \
--auto_quantize_bits 5.0 \
--auto_quantize_checkpoint llama-8B-auto-quantize-checkpoint \
--pyt_ckpt_path=meta-llama/Meta-Llama-3-8B \
--export_path=llama-8B_auto_quantize
--kv_cache_qformat=fp8 \
--calib_size=16 \
--batch_size=0 \
--trust_remote_code \
--export_fmt=tensorrt_llm
```
and compared exported files
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
## What does this PR do?
**Type of change:** ? new feature
**Overview:**
Support Qwen3Next NVFP4 quantization. The QKV layers are not quantized
in PTQ to retain checkpoint accuracy across benchmarks.
## Testing
Checkpoint export and benchmarking using nemo-evaluator
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
## What does this PR do?
**Type of change:** Bugfix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
**Overview:**
- Update custom file name patterns when copy files in export.
- Remove problematic parameters in export for Nemotron Nano VL models
- Resolve nvbugs https://nvbugspro.nvidia.com/bug/5637594
## Usage
<!-- You can potentially add a usage example below. -->
```python
# Add a code snippet demonstrating how to use this
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
## What does this PR do?
**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
But fix
**Overview:** ?
Make VILA codebase importable and import configuration before load model
config
Fix https://nvbugspro.nvidia.com/bug/5616904
## Usage
<!-- You can potentially add a usage example below. -->
```python
# Add a code snippet demonstrating how to use this
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
```
CUDA_VISIBLE_DEVICES=0 bash -e scripts/huggingface_example.sh --model /models/vila1.5-3b --quant int4_awq --tp 1 --pp 1 --trust_remote_code --kv_cache_free_gpu_memory_fraction 0.5
```
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
Signed-off-by: Yue <yueshen@nvidia.com>