mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
main
522
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2ce745a92e |
Deprecate gradnas pruning and bert example (#1427)
### What does this PR do? Type of change: Deprecation of dead code <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Deprecation warning already added in 0.44 as per 1-release deprecation policy GradNAS only works for Bert and GPT-J and we dont actively maintain it or test it. Keeping it creates an expectation that it works plus it adds one more option for user to choose from. We already have much better pruning algorithms (Minitron and Puzzletron) for LLM pruning already hence removing GradNas. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ No but we dont have any users of this feature either <!--- If ❌, explain why. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Deprecation** * GradNAS pruning algorithm deprecated; related examples removed. * **Documentation** * Pruning and NAS guides and changelog updated to focus on Minitron and FastNAS; GradNAS references removed. * **Chores** * Chained-optimizations example and scripts removed. * Ownership mappings updated for README and examples; license insertion now applies to a previously excluded example file. * **Tests** * Multiple unit tests and test utilities related to GradNAS/transformer NAS removed. [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1427) <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com> |
||
|
|
d30ebbd455 |
fix(llm_eval): migrate lm_eval_hf.py to lm-eval >= 0.4.10 HarnessCLI (#1416)
### What does this PR do?
Type of change: Bug fix
Related: NVbug 6153721. To make `lm_eval_hf.py` and the `vllm_causallms`
adapter compatible with vLLM 0.20+, **lm_eval must be upgraded to
0.4.11**.
`examples/llm_eval/lm_eval_hf.py` failed to import on lm-eval >= 0.4.10:
```
ImportError: cannot import name 'parse_eval_args' from 'lm_eval.__main__'
```
lm-eval 0.4.10 replaced `lm_eval.__main__.{setup_parser,
parse_eval_args, cli_evaluate}` with a `HarnessCLI`-based interface in
`lm_eval._cli`. This PR drops the legacy code path and drives
`HarnessCLI` directly:
- Attach the ModelOpt arguments (`--quant_cfg`, `--auto_quantize_*`,
`--calib_*`, `--compress`, `--sparse_cfg`) to the new `run` subparser.
- After parsing, move those keys out of the argparse namespace and into
`args.model_args` (now a dict, courtesy of `MergeDictAction`), so
`EvaluatorConfig.from_cli` doesn't reject them as unknown kwargs and so
they reach our `HFLM.create_from_arg_obj` override.
- Hard-require lm-eval >= 0.4.10 via `packaging.version.Version` and
bump `examples/llm_eval/requirements.txt` and
`examples/puzzletron/requirements.txt` accordingly.
### Usage
```bash
# Same CLI as before — HarnessCLI auto-inserts the `run` subcommand for legacy-style invocations.
python examples/llm_eval/lm_eval_hf.py \
--model hf \
--model_args pretrained=<HF model> \
--tasks hellaswag \
--quant_cfg FP8_DEFAULT_CFG \
--batch_size 4
```
### Testing
Added `tests/examples/llm_eval/test_llm_eval.py::test_lm_eval_hf` — an
end-to-end test that builds a tiny qwen3 (no GPU required) and runs
`lm_eval_hf.py` against MMLU with `--limit 0.1`. Verifies both the new
HarnessCLI integration and the `HFLM.create_from_arg_obj` override
actually execute. Existing FP8 PTQ test (`test_qwen3_eval_fp8`) was
retargeted to the same tiny qwen3 helper so it no longer pulls down the
1.1B TinyLlama checkpoint.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ❌ — drops support for lm-eval <
0.4.10. Requirements pins are bumped in the same PR; existing CLI
invocations continue to work because HarnessCLI auto-inserts the `run`
subcommand for legacy-style argv.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌
### Additional Information
NVbug 6153721 — vLLM 0.20+ compatibility requires lm_eval 0.4.11. This
PR pins to >= 0.4.10 to fix the import error; bumping the floor to
0.4.11 (or letting users opt in via vLLM extras) unblocks vLLM 0.20+
usage.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Chores**
* Require lm_eval >= 0.4.10 for evaluation examples and remove an
lm-eval pin from a puzzletron example.
* Improve LM-eval CLI integration so model-related CLI options
(including trust_remote_code) are recognized and forwarded correctly.
* **Tests**
* Added an end-to-end smoke test for the LM evaluation example.
* Updated evaluation tests to use the new tiny model setup for more
robust validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
6a3b6b8329 |
[Recipes][LLM PTQ] Add nvfp4 MSE+FP8-cast-KV recipes (experts_only / mlp_only) + --recipe in example scripts (#1407)
## Summary
- Adds two PTQ recipes that combine **experts/MLP-only NVFP4 W4A4** with
**MSE FP8 scale-sweep weight calibration** and **FP8 KV cache with
`use_constant_amax: true`** (skips KV calibration; matches the
`nvfp4_default-fp8_cast_kv` contract):
- `modelopt_recipes/general/ptq/nvfp4_experts_only_mse-fp8_cast_kv.yaml`
— applies to `*mlp.experts*` / `*block_sparse_moe*` only.
- `modelopt_recipes/general/ptq/nvfp4_mlp_only_mse-fp8_cast_kv.yaml` —
applies to all `*mlp*` / `*block_sparse_moe*` (dense MLP + MoE).
- Threads a new `--recipe` flag through
`examples/llm_ptq/scripts/parser.sh` and `huggingface_example.sh`.
Either `--quant` or `--recipe` is required; passing **both errors out**.
Recipe names are not validated in the script — `hf_ptq.py` is the source
of truth.
- Drops the bash-side `qformat` whitelist case-statement in
`huggingface_example.sh` for the same reason.
## Files
**New recipes (`modelopt_recipes/general/ptq/`):**
- `nvfp4_experts_only_mse-fp8_cast_kv.yaml` — same patterns as
`nvfp4_experts_only-fp8_kv.yaml`.
- `nvfp4_mlp_only_mse-fp8_cast_kv.yaml` — same patterns as
`nvfp4_mlp_only-fp8_kv.yaml`.
Both differ from their `_kv` siblings by:
- `algorithm: max` → `{ method: mse, fp8_scale_sweep: true, layerwise:
false }`
- All targeted **weight quantizers** switch `type: dynamic` → `type:
static` (otherwise `mse_calibrate` skips them: only static block-quant
weight quantizers are recognized for the FP8 sweep — see
`model_calib.py:369-374`).
- Input quantizers stay dynamic.
- KV bmm adds `use_constant_amax: true` (the `_cast_kv` flavor).
**Scripts (`examples/llm_ptq/scripts/`):**
- `parser.sh` — adds `--recipe` long-option, default `RECIPE=""`,
validates one-of-{`--quant`, `--recipe`} and not-both.
- `huggingface_example.sh` — when `RECIPE` is set, derives `MODEL_NAME`
from the recipe basename, passes `--recipe=…` to `hf_ptq.py` instead of
`--qformat=…`, and exits after export with a TRT-LLM deployment hint
(recipes can produce arbitrary configs that the script's downstream
`run_tensorrt_llm.py` path doesn't know how to handle generically).
Drops the `qformat` whitelist; defers to `hf_ptq.py`.
## Behavior
```
# Errors with: "Cannot specify both --quant and --recipe; pick one."
bash huggingface_example.sh --model=... --quant=nvfp4 --recipe=... --tasks=quant
# Errors with usage if neither is given
bash huggingface_example.sh --model=... --tasks=quant
# Both of these are now accepted; --recipe is forwarded verbatim to hf_ptq.py
bash huggingface_example.sh --model=... --quant=nvfp4 --tasks=quant
bash huggingface_example.sh --model=... --recipe=general/ptq/nvfp4_experts_only_mse-fp8_cast_kv --tasks=quant
bash huggingface_example.sh --model=... --recipe=general/ptq/nvfp4_mlp_only_mse-fp8_cast_kv --tasks=quant
```
## Test plan
- [x] `experts_only_mse-fp8_cast_kv` loads via
`modelopt.recipe.load_recipe(...)` and produces the expected algorithm +
per-pattern `quant_cfg` (verified in a working env: `algorithm ==
{'method': 'mse', 'fp8_scale_sweep': True, 'layerwise': False}`; expert
weight quantizers `type: static`; KV bmm has `use_constant_amax: True`).
- [x] Parser sanity: 4 flag combinations (both, neither, only `--quant`,
only `--recipe`) all behave as designed.
## Note
Pre-commit hook `check-modelopt-recipes` was skipped on both commits
because the local conda env has a broken `torchvision` install
(`AttributeError: partially initialized module 'torchvision' has no
attribute 'extension'`) that prevents `from modelopt.recipe.loader
import load_recipe`. The `experts_only` recipe was validated
independently by running `tools/precommit/check_modelopt_recipes.py` in
a working environment (exits 0); the `mlp_only` one is the same shape
with a different glob.
Rebased onto `main` from #1391 (which targeted
`chenjiel/nvfp4-fp8-sweep-triton`). The diff is scoped to the recipes +
script wiring; no kernel/sweep changes are included here.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added recipe-based quantization as an alternative to format-based
quantization with a new `--recipe` CLI option.
* Added two new quantization recipes for targeted layer optimization:
one for expert-layer-only quantization and one for MLP-layer-only
quantization, both featuring NVFP4 and FP8 KV-cache optimization.
* **Configuration**
* `--quant` and `--recipe` options are now mutually exclusive; specify
one to configure quantization behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
|
||
|
|
fa05381b86 |
Add demo (Puzzletron vs Minitron guide) in examples/pruning/ with README and notebooks (#1320)
### What does this PR do? Type of change: new documentation/example (tutorial + notebooks) Adds an end-to-end pruning & distillation guide under `examples/pruning_demo/`, walking users through structural compression of Qwen3-8B with NVIDIA Model-Optimizer. The example compares two methods side-by-side on two concrete scenarios: - **Scenario 1 — Moderate compression (7B parameter target)**: homogeneous pruning with Minitron vs. heterogeneous NAS-based pruning with Puzzletron. - **Scenario 2 — Aggressive compression (78,000 MiB memory budget)**: same comparison under a hard memory constraint. Both scenarios are followed by knowledge distillation and evaluated on MMLU (end-to-end in the notebooks) plus HellaSwag and GSM8K (reported in the guide). Contents: - `README.md` — full guide (setup, two scenarios, head-to-head analysis, inference benchmarks with vLLM + AIPerf, decision rules, limitations, open questions). - `00_prerequisites.ipynb` — data prep (WikiText-103 → Megatron binary) and teacher baseline evaluation. - `scenario1_minitron.ipynb` / `scenario1_puzzletron.ipynb` — 7B-param target. - `scenario2_minitron.ipynb` / `scenario2_puzzletron.ipynb` — 78k-MiB target, including a Puzzletron memory-sweep bonus section. - `advanced_compression_experiments.md` — extended results (larger distillation budgets with Nemotron-Post-Training-Dataset-v2, BLD, chained Minitron→Puzzletron, Mamba-Transformer hybrid). - Companion plots (`summary_chart.png`, `distillation_curves.png`, `memory_sweep_combined.png`, `all_curves_throughput_vs_latency.png`, ...). ### Usage Follow setup instructions in README.md then run, in order: 1. 00_prerequisites.ipynb — prepare data + baseline eval (~15 min). 2. One (or more) of the scenario notebooks: - scenario1_minitron.ipynb (~1h45) - scenario1_puzzletron.ipynb (~6h first run) - scenario2_minitron.ipynb (~45 min) - scenario2_puzzletron.ipynb (~6h15 first run) ### Testing - All four scenario notebooks were executed and tested end-to-end on 2x H200 GPUs - Inference benchmarks were captured on 1x H200 NVL with vLLM (AnyModel backend for Puzzletron checkpoints) and AIPerf - No library code is modified, so no unit tests are affected ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ (documentation/examples-only addition) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new runtime dependencies; the notebooks use lm-eval==0.4.8 and the existing ModelOpt/NeMo stack. The vLLM serving appendix references an open PR (vllm-project/vllm#36512) for Puzzletron AnyModel support, clearly flagged as pre-release. - Did you write any new necessary tests?: N/A — tutorial / documentation example. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information - Base model: https://huggingface.co/Qwen/Qwen3-8B - Calibration dataset: nvidia/Nemotron-Post-Training-Dataset-v2 - Distillation dataset: WikiText-103 - Complements the existing examples/puzzletron/ and examples/megatron_bridge/ READMEs with a scenario-driven narrative and a direct Minitron↔Puzzletron comparison. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added comprehensive guides comparing Minitron (homogeneous pruning) and Puzzletron (heterogeneous NAS + MIP) for LLM compression. * Added step-by-step Jupyter notebooks demonstrating two end-to-end scenarios (prune → distill → evaluate) with expected MMLU baselines and memory budgeting. * Added advanced experiments doc with extended results, chaining strategies, benchmarking (including inference/vLLM notes), tips, limitations, and appendices. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Alexandre Chidiac <achidiac@nvidia.com> |
||
|
|
f34f488a83 |
Add a general composable $import system for YAML configs, and use it to implement composable recipes (#1253)
### What does this PR do?
Type of change: New feature
Adds a general composable YAML config loading layer for ModelOpt configs
and recipes. YAML remains the source of truth for configuration data,
while Python/Pydantic-compatible types provide schema validation at load
time. This PR uses that loader to de-duplicate PTQ recipes, introduce
reusable config snippets/presets, and start migrating selected hardcoded
quantization presets to YAML.
#### Problem
1. Built-in PTQ recipes duplicated numeric format definitions, KV-cache
entries, and the default quantizer exclusion list.
2. YAML snippets were reusable only by convention; they did not declare
or validate the schema they were meant to satisfy.
3. Loading YAML-backed quantization presets from
`modelopt.torch.quantization.config` could not depend on
`modelopt.recipe` without creating circular imports.
4. Directory-format recipes exposed import-resolution details in the
recipe loader and used `recipe.yaml` plus nested `metadata:` in a way
that made metadata handling inconsistent.
#### Solution
**Shared YAML config loader**
- Adds `modelopt.torch.opt.config_loader` as the low-level loader used
by both `modelopt.recipe` and `modelopt.torch.quantization.config`.
- Keeps the public `modelopt.recipe.load_config()` entry point, while
removing the private `modelopt/recipe/_config_loader.py` shim.
- Handles YAML loading, built-in/filesystem path resolution, suffix
probing, `ExMy` conversion for `num_bits` / `scale_bits`, `$import`
expansion, and schema validation.
- Lives below `modelopt.recipe` in the dependency graph to avoid
circular imports from quantization config code.
**Composable `$import` system**
Recipes and snippets can declare an `imports` mapping, then reference
entries with `{$import: name}`.
`$import` semantics:
- **Dict value**: replaced with the imported dict. Multiple imports are
supported with ordered precedence; inline keys override imported keys.
- **List entry**: schema-driven behavior for strongly typed lists. If
the snippet schema matches the containing list type, the imported list
is spliced. If the snippet schema matches the list element type, the
imported element is appended. Other schema combinations are rejected.
- **Multi-document YAML**: supports snippets that need an `imports`
header plus a list body.
- **Recursive and scoped**: snippets can import other snippets; import
names are scoped per file.
- **Cycle detection**: circular imports report a clear error.
**Snippet schema validation**
- Every reusable snippet referenced through `imports` must declare a `#
modelopt-schema: ...` preamble.
- Snippets are validated after nested imports are resolved.
- Schema paths are restricted to the `modelopt.` package and may be
Pydantic models, `TypedDict` classes, or explicitly typed container
aliases such as `list[QuantizerCfgEntry]`.
- Untyped list imports are rejected so list append/splice behavior stays
strongly typed.
**Recipe model and directory recipe cleanup**
- `ModelOptRecipeBase` now owns a `metadata: RecipeMetadataConfig`
field.
- `ModelOptPTQRecipe` is the PTQ recipe schema; the overlapping
YAML-specific PTQ config class was removed.
- Directory recipes now use `metadata.yaml` / `metadata.yml` for
top-level metadata fields, plus section files such as `quantize.yaml`.
- Directory recipe loading now delegates import resolution to
`load_config()` instead of manually using raw config loading.
**Config snippet and preset library**
Adds reusable snippets under `modelopt_recipes/configs/`:
- `numerics/`: `fp8`, `nvfp4`, `nvfp4_static`
- `ptq/units/`: `base_disable_all`, `default_disabled_quantizers`,
`w8a8_fp8_fp8`, `w4a4_nvfp4_nvfp4`, `kv_fp8`, `kv_fp8_cast`,
`kv_nvfp4_cast`
- `ptq/presets/`: YAML presets for `FP8_DEFAULT_CFG` and `FP8_KV_CFG`
`FP8_DEFAULT_CFG` and `FP8_KV_CFG` now load from YAML presets via
`load_config()`.
**Recipe migration and naming**
- General PTQ recipes now use shared imports instead of repeating the
same quantizer fragments inline.
- General PTQ recipe paths were renamed to KV-first naming, for example:
- `general/ptq/fp8_default-fp8_kv` -> `general/ptq/fp8_default-kv_fp8`
- `general/ptq/fp8_default-fp8_cast_kv` ->
`general/ptq/fp8_default-kv_fp8_cast`
- `general/ptq/nvfp4_default-none_kv_gptq` ->
`general/ptq/nvfp4_default-kv_none-gptq`
- `general/ptq/nvfp4_default-nvfp4_cast_kv` ->
`general/ptq/nvfp4_default-kv_nvfp4_cast`
- Example docs and `examples/llm_ptq/hf_ptq.py --recipe` help text were
updated to use the new paths.
**Pre-commit and documentation**
- Recipe validation accepts `$import` entries and handles directory
recipes using `metadata.yaml`.
- The recipe validation hook skips `modelopt_recipes/configs/` because
those files are reusable snippets, not full recipes.
- `docs/source/guides/10_recipes.rst` now documents imports, schema
modelines, list append/splice semantics, built-in snippets, built-in
recipe paths, directory recipes, and the current recipe data model.
#### Backward compatibility
- Existing inline YAML recipes without `$import` continue to load.
- `modelopt.recipe.load_config()` remains public.
- The built-in recipe path renames are user-visible; callers should
update recipe path strings to the KV-first names listed above.
#### Testing
- `pytest tests/unit/recipe/test_loader.py -q` - 90 passed
- `python tools/precommit/check_modelopt_recipes.py ...` for the renamed
built-in PTQ recipes
- `pre-commit run mypy --files
modelopt/onnx/llm_export_utils/quantization_utils.py`
- `python -m py_compile examples/llm_ptq/hf_ptq.py`
- `git diff --check`
### Before your PR is "Ready for review"
- Is this change backward compatible?: Partially. Loader/API behavior is
compatible for existing inline YAML recipes, but built-in recipe path
names were renamed to KV-first paths.
- Did you write any new necessary tests?: Yes.
- Did you update Changelog?: Yes.
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
7097a6910e |
[specdec_bench] Stratify --num_requests across categories (#1389)
When the dataset has a `category` column with >1 distinct categories and ``--num_requests N`` is below the dataset size, take ceil(N / num_categories) rows from each category and round-robin interleave them so any prefix is balanced. Falls back to the existing ``range(N)`` slice when category metadata is absent or there's only one category. Fixes a sampling bug where SPEED-Bench parquet files are sorted by category, so e.g. ``--num_requests 64`` on throughput_8k pulls 64 high_entropy prompts and zero from low_entropy / mixed. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * SPEEDBench dataset truncation now performs deterministic, balanced stratified sampling across categories (round-robin) when multiple categories exist and a smaller sample size is requested. * Falls back to simple prefix selection when no category or only one category is present, preserving prior behavior in that case. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Alexandre Milesi <milesial@users.noreply.github.com> |
||
|
|
d794595bf8 |
llm_sparsity: Set warmup_steps 0 instead of 0.0 for transformers 5.x compat (#1393)
Fix for NVBug 6120631 to fix ``` finetune.py: error: argument --warmup_steps/--warmup-steps: invalid int value: '0.0' ``` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Corrected parameter format in finetuning example script for consistency. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
b5df2a5a3c |
Update llm_ptq requirements.txt (#1394)
### What does this PR do? Type of change: Dependency update compressed_tensors 0.15 is not compatible with our current quant implementation for Kimi K2.5, K2.6 ### Testing Unittest ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Broadened the `compressed-tensors` dependency constraint to allow a wider range of compatible versions. * Removed the `rouge_score` dependency from the example requirements. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
ef326c8b9b |
Add Gemma4 MoE quantization support (#1219)
## Summary - Register `Gemma4TextExperts` with `_QuantQwen35MoeExperts` plugin to unfuse fused 3D expert tensors into per-expert `nn.Linear` layers for quantization - Add structural `is_moe()` detection for modules with `router` + `experts` attributes (Gemma4 has no dedicated `SparseMoeBlock` class — the decoder layer directly owns `router` and `experts`) - Add `Gemma4TextDecoderLayer` to `get_expert_linear_names()` returning `["gate_proj", "down_proj", "up_proj"]` - Add `"*.experts.*"` pattern to `NVFP4_MLP_ONLY_CFG` and `NVFP4_EXPERTS_ONLY_CFG` to match Gemma4's expert path (`model.layers.X.experts.*`, not nested under `mlp`) **Context:** Gemma4 MoE models (e.g. `google/gemma-4-26B-A4B-it`) store expert weights as fused 3D `nn.Parameter` tensors (`gate_up_proj`, `down_proj`) instead of `nn.ModuleList` of `nn.Linear`. Since ModelOpt's quantizer only discovers `nn.Linear` modules, it silently skips the expert weights — the bulk of the model remains unquantized. **Companion vLLM PR:** https://github.com/vllm-project/vllm/pull/39406 (robust quantized MoE weight loading for Gemma4) ## Test plan - [x] `hf_ptq.py --pyt_ckpt_path google/gemma-4-26B-A4B-it --qformat nvfp4_mlp_only` — 35k+ quantizers inserted, 17GB output (vs 49GB BF16) - [x] `vllm serve <path> --quantization modelopt` — loads and serves successfully - [x] Text generation: correct ("The capital of France is **Paris**.") - [x] Vision: correct (describes image content accurately) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Support quantizing models with separate base/full components (handles heads present only on the full model) * Enhanced Mixture-of-Experts detection and explicit support for Gemma4 expert layer layouts * Extended NVFP4 selective quantization presets and recipes to include expert-layer patterns and enable FP8 for expert modules * **Bug Fixes** * Improved loss/logit handling and clearer errors for unsupported quantization methods <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: James Shen <yueshen@nvidia.com> Signed-off-by: Yue Shen <yueshen@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
f0eaa198df |
Enable active-param and memory based Minitron pruning constraint (#1377)
### What does this PR do?
Type of change: New feature, new tests, documentation.
OMNIML-4108: Extends the Minitron NAS pruner to support pruning by
**active parameter count** (`active_params`) and **memory footprint**
(`memory_mb`) in addition to the existing total parameter count
(`params`) constraint. Also adds standalone utilities for analytical
model stats.
#### Changes
**New pruning constraint keys**
- `active_params`: prune to a target number of active (routed) params —
useful for MoE models where total ≫ active; when present,
`active_params` is the **primary sort/display metric** for candidates
(priority: `active_params` > `params` > `memory_mb`)
- `memory_mb`: prune to fit a memory budget (BF16 weights + KV-cache +
Mamba state at a given sequence length and batch size)
- Constraints can be combined (AND logic): e.g. `{"params": 6e9,
"memory_mb": 12288}`
**New standalone utilities**
(`modelopt.torch.nas.plugins.megatron_model_stats`)
- `mcore_param_count`: analytically computes total and active parameter
counts for GPT and Mamba/hybrid MCore models
- `mcore_memory_footprint_mb`: estimates memory in MB (weights +
KV-cache + Mamba state)
- `print_mcore_model_stats`: rich-formatted model stats panel
**Rich-formatted pruning logs** — search space, top-k candidate tables,
and best subnet panel printed on rank 0
**`prune_score_func` format update** — now `mmlu_<N>pct_bs<bs>` (e.g.
`mmlu_10pct_bs32`) to explicitly control batch size for MMLU evaluation;
old `mmlu_<N>pct` format removed
**Infrastructure**
- NeMo container bumped to `nvcr.io/nvidia/nemo:26.04` in CI and docs
- Added `examples/megatron_bridge/requirements.txt` with
`transformers<5.0` (required for saving some Nemotron-3-Nano models)
### Usage
```python
# Prune to 3B active params (MoE-aware) — active_params is the primary sort metric
mtp.prune(model, mode=[("mcore_minitron", ss_config)], constraints={"active_params": 3e9}, config=pruning_config)
# Prune to fit a 12 GB memory budget
mtp.prune(model, mode=[("mcore_minitron", ss_config)], constraints={"memory_mb": 12288}, config=pruning_config)
```
### Testing
Pruned Nemotron-3-Nano-30B-A3B (31.6B, A3.6B) --> A3.0B. Takes <1hr on
8x H100 (more details in #1376)
```bash
torchrun --nproc_per_node 8 examples/megatron_bridge/prune_minitron.py \
--pp_size 8 \
--hf_model_name_or_path nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 \
--trust_remote_code \
--prune_target_params 28e9 \
--prune_target_active_params 3e9 \
--hparams_to_skip num_attention_heads \
--seq_length 8192 \
--output_hf_path pruned/Nemotron-3-Nano-30B-A3B-Pruned-28B-A3B-top20-max15depth-max30width-mmlu_10pct_bs32 \
--top_k 20 \
--max_depth_pruning 0.15 \
--max_width_pruning 0.30 \
--prune_score_func mmlu_10pct_bs32 \
--num_layers_in_first_pipeline_stage 5 \
--num_layers_in_last_pipeline_stage 5
```
```
╭──────────────────────────────────────────────────── Original Model Stats ─────────────────────────────────────────────────────╮
│ Total Parameters 31.58B │
│ Active Parameters 3.58B │
│ Memory (BF16, seq_length=8192, batch_size=1) weights: 60230.1 MB, kv_cache: 48.0 MB, mamba_state: 23.8 MB, Total: 60301.9 MB │
╰───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
Top 20 Candidates with Scores
┏━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━┓
┃ # ┃ export_config ┃ active_params ┃ params ┃ score ┃
┡━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━┩
│ 1 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 64, 'num_moe_experts': 120, │ 3.00B │ 27.06B │ 0.3399 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 2 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 56, 'num_moe_experts': 112, │ 3.00B │ 25.37B │ 0.4650 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 3 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 64, 'mamba_head_dim': 56, 'num_moe_experts': 112, │ 3.00B │ 25.37B │ 0.2343 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 4 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 56, 'mamba_head_dim': 48, 'num_moe_experts': 96, │ 3.00B │ 20.09B │ 0.2552 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 5 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48, 'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 21.61B │ 0.2601 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 6 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 64, 'num_moe_experts': 96, │ 3.00B │ 19.28B │ 0.3762 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3712} │ │ │ │
│ 7 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 104, │ 3.00B │ 22.28B │ 0.4783 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 8 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 96, │ 3.00B │ 21.99B │ 0.2420 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3328} │ │ │ │
│ 9 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 112, │ 3.00B │ 25.37B │ 0.2399 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3712} │ │ │ │
│ 10 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 112, │ 3.00B │ 26.17B │ 0.2601 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3328} │ │ │ │
│ 11 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 64, 'num_moe_experts': 112, │ 3.00B │ 25.37B │ 0.2503 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 12 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 23.68B │ 0.4329 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 13 │ {'num_layers': 46, 'hidden_size': 2688, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 128, │ 3.00B │ 26.17B │ 0.2587 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 2816} │ │ │ │
│ 14 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 64, 'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 23.68B │ 0.2336 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 15 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48, 'mamba_head_dim': 56, 'num_moe_experts': 96, │ 3.00B │ 20.09B │ 0.2559 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 16 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 96, │ 3.00B │ 20.70B │ 0.4608 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 17 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 104, │ 3.00B │ 23.68B │ 0.2455 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3712} │ │ │ │
│ 18 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 104, │ 3.00B │ 24.42B │ 0.2503 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3328} │ │ │ │
│ 19 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 120, │ 3.00B │ 27.92B │ 0.2587 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3712} │ │ │ │
│ 20 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 64, 'num_moe_experts': 104, │ 3.00B │ 23.68B │ 0.2469 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
└────┴───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┴───────────────┴────────┴────────┘
╭──────────────────────────────────────────────────────────────────────── Best Subnet ─────────────────────────────────────────────────────────────────────────╮
│ export_config {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 104, 'moe_ffn_hidden_size': 1856, │
│ 'moe_shared_expert_intermediate_size': 3072} │
│ active_params 3.00B │
│ params 22.28B │
│ score 0.4783 │
╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
╭───────────────────────────────────────────────────── Pruned Model Stats ──────────────────────────────────────────────────────╮
│ Total Parameters 22.28B │
│ Active Parameters 3.00B │
│ Memory (BF16, seq_length=8192, batch_size=1) weights: 42489.7 MB, kv_cache: 48.0 MB, mamba_state: 23.8 MB, Total: 42561.6 MB │
╰───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
```
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
84fe91be5d |
Fix gpt-oss examples trl import error (#1390)
### What does this PR do?
Type of change: Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
<!-- Details about the change. -->
Cap kernels<0.13 and trackio<0.21 in examples/gpt-oss/requirements.txt.
Both newer versions require huggingface_hub>=1.x, but the example's
transformers pins huggingface_hub<1.0, so a fresh install breaks on
import (Unsupported type for field 'import_name': str | None from
kernels; cannot import name 'Volume' from trackio).
### Usage
No API change. On transformers<5.0, override the config's warmup_steps
with --warmup_ratio 0.03 --warmup_steps 0 (or edit the YAML), as already
noted by the comment in configs/sft_*.yaml.
```python
# Add a code snippet demonstrating how to use this
accelerate launch --config_file configs/zero3.yaml sft.py --config configs/sft_full.yaml --model_name_or_path openai/gpt-oss-20b --quant_cfg MXFP4_MLP_WEIGHT_ONLY_CFG --output_dir gpt-oss-20b-qat --warmup_steps 0 --warmup_ratio 0.03
```
### Testing
1. pip install -r examples/gpt-oss/requirements.txt
pip install transformers==4.57.3
```python
# Add a code snippet demonstrating how to use this
accelerate launch --config_file configs/zero3.yaml sft.py --config
configs/sft_full.yaml --model_name_or_path openai/gpt-oss-20b
--quant_cfg MXFP4_MLP_WEIGHT_ONLY_CFG --output_dir gpt-oss-20b-qat
--warmup_steps 0 --warmup_ratio 0.03
```
2. pip install -r examples/gpt-oss/requirements.txt
pip install --upgrade transformers
```python
# Add a code snippet demonstrating how to use this
accelerate launch --config_file configs/zero3.yaml sft.py --config
configs/sft_full.yaml --model_name_or_path openai/gpt-oss-20b
--quant_cfg MXFP4_MLP_WEIGHT_ONLY_CFG --output_dir gpt-oss-20b-qat
```
<!-- Mention how have you tested your change if applicable. -->
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
### Additional Information
<!-- E.g. related issue. -->
Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com>
|
||
|
|
1d21ab9e29 |
[DeepSeek] Default to top-k calibration with peer-max input amax sync (#1380)
## Summary - DeepSeek PTQ (`examples/deepseek/ptq.py`) now defaults to native top-k routing during MoE calibration. The previous all-tokens-to-all-experts path (`CalibMoe`) is preserved behind a new `--calib_all_experts` flag. - After `mtq.quantize`, `fixup_moe_expert_amax` syncs every expert's `input_quantizer.amax` (w1/w2/w3) to the per-layer global peer max via `dist.all_reduce(MAX)` across EP ranks. `weight_quantizer.amax` stays per-expert; any uncalibrated expert is filled by computing amax over the dequantized FP8 weight. - `mtq.print_quant_summary` is now also written to `<output_path>/.quant_summary.txt`, mirroring `llm_ptq/hf_ptq.py`. ## Why Forcing all tokens through every expert doubled calibration time and inflated `input_quantizer.amax` for cold-routing experts with outliers they never see at inference. The new flow matches the inference distribution, runs roughly 2x faster, and mirrors the `layer_sync_moe_local_experts_amax` semantics that mtq runs automatically for `QuantSequentialMLP`-derived MoEs. ## Validation (DeepSeek-V3.2-Exp, MP=8, NVFP4_DEFAULT_CFG) Compared `_amax_baseline` (CalibMoe) vs `_amax_synced` (new default): - All 44,544 expert weight amaxes bit-identical. - Attention, shared experts, gate: identical. - Expert `w1.input` and `w3.input` (shared MoE block input): identical. - Expert `w2.input` (post-SiLU gated, expert-specific): synced to layer-wide peer max — 99.3% are larger than baseline (median 11.4x) since peer-max captures the worst-case outlier from any expert in the layer; 0.7% are smaller. This is the same trade-off `set_expert_quantizer_amax` makes for HF MoEs in `unified_export_hf.py`. ## Test plan - [x] DeepSeek-V3.2-Exp MP8 PTQ with default flags — completes in ~7 min (vs ~27 min with CalibMoe), produces `_amax_synced/` consistent with the comparison above. - [x] DeepSeek-V3.2-Exp MP8 PTQ with `--calib_all_experts` — produces `_amax_baseline/` identical (other than rounding) to the prior `CalibMoe`-default behavior. - [x] `.quant_summary.txt` written under `output_path` on rank 0. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a `--calib_all_experts` option to enable an alternate PTQ calibration mode; default remains top-k routing with a post-calibration per-layer peer-max synchronization and a compute fallback for uncalibrated experts. * **Documentation** * Clarified default and alternate calibration behaviors and added note about generation of a `.quant_summary.txt` summary file. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
383ab4e224 |
fix: include medusa in data_module assignment in main.py (#1370)
## Problem
When `training.mode == "medusa"` is used in `main.py`, the `data_module`
variable is never assigned because line 344 only covered `eagle3` and
`dflash` modes. This causes an `UnboundLocalError` when the trainer is
constructed with `**data_module`.
Fixes OMNIML-4147
## Fix
Add `"medusa"` to the `training_args.mode in ("eagle3", "dflash")`
condition so `data_module` is correctly populated for medusa training.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Fixed speculative decoding example to properly handle "medusa" mode
alongside existing "eagle3" and "dflash" modes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
||
|
|
50706d1750 |
Add closed-form MXFP4 -> NVFP4 weight cast (--cast_mxfp4_to_nvfp4) (#1372)
## Summary - New `--cast_mxfp4_to_nvfp4` flag in `hf_ptq.py` (and `huggingface_example.sh`) that converts an MXFP4 source checkpoint (e.g. `openai/gpt-oss-20b`) into an NVFP4 export with **bit-exact** weight reconstruction for the in-range blocks. - The cast pins NVFP4's `scale_2 = 2^m` (where `m = k_max − 8`) and `_amax = 6·2^k_j` per NVFP4 block, both read from the source `*_scales`. The resulting per-block scale `2^(k_j − m)` is exactly representable in E4M3, so `round_to_E2M1(value / 2^k_j)` yields the original MXFP4 nibble verbatim. For out-of-range blocks (`k_max − k_j > 17`) the per-block amax falls back to data-derived `max(|w_block|)`, which keeps the post-E4M3-clamp scale close to the block's actual magnitude. ## Verification End-to-end on `openai/gpt-oss-20b` with `--qformat=nvfp4_mlp_only --cast_mxfp4_to_nvfp4`: ``` [cast_mxfp4_to_nvfp4] overrode 48/48 weight quantizers [cast_mxfp4_to_nvfp4] lossless layers: 48/48 (100.00%) [cast_mxfp4_to_nvfp4] lossless blocks: 597196800/597196800 (100.0000%) ``` End-to-end on `openai/gpt-oss-120b` with the same flags (4×B200, `--use_seq_device_map --gpu_max_mem_percentage 0.5 --calib_batch_size 4`): ``` [cast_mxfp4_to_nvfp4] overrode 72/72 weight quantizers [cast_mxfp4_to_nvfp4] lossless layers: 67/72 (93.06%) [cast_mxfp4_to_nvfp4] lossless blocks: 3583179586/3583180800 (100.0000%) ``` Five layers fall into the OOR regime (block-spread > 17); the remaining 1,214 OOR blocks use the data-derived per-block amax fallback. Block-level losslessness is **99.99996%** end-to-end. Per-tensor MSE between MXFP4 source dequant and NVFP4 export dequant (~19B elements): | Metric | Without cast | With cast | |---|---|---| | Per-tensor SNR | ~26.4 dB (FP4 noise floor) | **∞ (every tensor)** | | Total RMSE | 8.67e−02 | **0** | | max\|err\| | up to 8.0e+1 | **0** | ## Modelopt-side enablers - `max_calibrate` auto-promotes static-block NVFP4 weight quantizers to `NVFP4StaticQuantizer` at the end of calibration. - `static_blockwise_fp4_fake_quant` kernel accepts N-D inputs (was 2D-only), unblocking MoE expert weights of shape `(E, F, K)`. - BMM-experts NVFP4 export routes through `get_weights_scaling_factor_from_quantizer` for static-mode quantizers, so the pinned `_amax` is actually consumed. - `set_expert_quantizer_amax` scalar-reduces per-quantizer amax before stacking, supporting per-block (vs scalar) static-mode amax. ## Test plan - [x] Unit tests at `tests/examples/llm_ptq/test_cast_mxfp4_to_nvfp4.py` (15 tests, all passing) cover: scalar/global-amax math, per-block hybrid (in-range closed-form vs OOR data-derived), shape preservation, key collection, and end-to-end `build_amax_map` against a synthetic safetensors checkpoint. - [x] End-to-end PTQ → export on `openai/gpt-oss-20b` (`nvfp4_mlp_only` qformat) with `--cast_mxfp4_to_nvfp4` succeeds; export takes ~21 s. 100% lossless cast (48/48 layers, 597,196,800 / 597,196,800 blocks). - [x] End-to-end PTQ → export on `openai/gpt-oss-120b` (4×B200, `nvfp4_mlp_only`, `--use_seq_device_map --gpu_max_mem_percentage 0.5 --calib_batch_size 4`). 67/72 layers fully lossless; 99.99996% block-level losslessness (3,583,179,586 / 3,583,180,800). - [x] TRT-LLM serving validation (TRT-LLM 1.3.0rc11, B200) on both exported NVFP4 checkpoints via `examples/llm_ptq/run_tensorrt_llm.py`: - **20b** (TP=1): 18.3 GB GPU memory; coherent generation. Sample: *"Quantum computing is poised to revolutionize data analysis. However, its potential is currently limited by quantum hardware constraints, including error rates, qubit lifetimes, and lack of fault tolerance…"* - **120b** (TP=4): 36.4 GB / GPU; coherent generation. Sample: *"Quantum computing is poised to revolutionize data storage and processing. These rare earth-based systems could serve as robust qubits; resistant to environmental decoherence…"* - [x] MSE comparison script (run separately during development) confirms per-tensor SNR=∞ across all 48 MoE expert tensors. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a MXFP4→NVFP4 weight-format cast utility and a CLI flag to enable it; helper scripts updated to expose the option. * **Bug Fixes** * Fixed static NVFP4 export for expert weights. * Improved collection/handling of quantizer amax values to avoid shape issues. * Generalized FP4 kernel to accept flexible tensor dimensionality. * Ensured static-block NVFP4 promotion during calibration. * **Tests** * Added comprehensive tests for the conversion workflow, helpers, and end-to-end application. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
bb08094ff1 |
Add Nemotron-Nano-9B-v2 → Pruned 7B e2e tutorial: Prune + Distill + Eval + Quantize + vLLM deployment (#1325)
## Summary End-to-end optimization walkthrough for Nemotron-Nano-9B-v2 showing how ModelOpt techniques stack: - **Pruning** — Minitron structured pruning 9B → 7B - **Distillation** — Megatron-Bridge knowledge distillation up to 80B tokens; near-parity with official 9B on MMLU Pro, GPQA, LCB, AIME, Math 500, IFEval, SciCode - **Evaluation** - using nemo-evaluator - **Quantization** — FP8 PTQ via \`hf_ptq.py\`; checkpoint deployable on vLLM/TRT-LLM/SGLang with no extra flags (quantization auto-detected from \`config.json\`) - **vLLM Throughput** — BF16 vs FP8 benchmark on single H100 <img width="2085" height="1740" alt="image" src="https://github.com/user-attachments/assets/8620a019-5c09-4a6b-a5d2-ca164aaa5d87" /> <img width="2085" height="810" alt="image" src="https://github.com/user-attachments/assets/742c8035-f1fb-4394-b11b-0c6c3ac4e843" /> ### Files changed - `examples/pruning/minitron/README.md` — index page for Minitron end-to-end tutorials - `examples/pruning/minitron/NVIDIA-Nemotron-Nano-9B-v2/README.md` — full repro doc with 6 sections: data prep, pruning, distillation, evaluation, FP8 quantization, vLLM benchmarking - `examples/pruning/minitron/NVIDIA-Nemotron-Nano-9B-v2/nemo_evaluator.yaml` — NeMo Evaluator config used for all benchmark numbers - `examples/pruning/puzzletron/README.md` — index page for Puzzletron distillation results - `examples/pruning/puzzletron/Llama-3.1-8B-Instruct.md` — Puzzletron distillation results (renamed from puzzletron.md) - `examples/pruning/README.md` — updated Results section with direct links to new locations - `examples/megatron_bridge/README.md` — updated results link to point to `examples/pruning/` - `examples/puzzletron/README.md` — updated distillation results link - `examples/dataset/MEGATRON_DATA_PREP.md` — tokenization commands for all datasets used in the data blend 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Documentation * **New end-to-end tutorial** for model optimization covering Minitron pruning, knowledge distillation, FP8 quantization, and vLLM deployment with reproducibility steps and benchmark results * **Dataset preparation guide** with ready-to-run tokenization templates for Nemotron HuggingFace datasets * **Evaluation configuration** and results documentation including ablation studies across multiple benchmarks * **Updated navigation** across pruning, distillation, and dataset examples to streamline user workflows <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
6d330784a6 |
Add required keys to attention pruning config (#1360)
### What does this PR do? Type of change: ? Bug fix The config `examples/puzzletron/configs/llama-3_1-8B_pruneffn_memory/pruning/attn_pruning.yaml` didn't have required keys to use attention pruning in the example `examples/puzzletron/main.py` ### Usage ### Testing In `examples/puzzletron/configs/llama-3_1-8B_pruneffn_memory/Llama-3_1-8B.yaml` change `ffn_pruning` to `attn_pruning` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated pruning configuration for improved KV-head pruning support, including enhanced importance hook settings and attention output handling for memory optimization. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Grzegorz Karch <gkarch@nvidia.com> |
||
|
|
1ec931c2c7 |
[2/3][Feat]: Offline DFlash training (#1343)
### What does this PR do? Type of change: new feature Part 2 of a 3-PR series splitting #1271: - **[1/3] #1296**: File reorg + deprecate `ParallelDraft` - **[2/3] this PR**: Offline DFlash training (depends on #1296) - **[3/3] #1297**: Extract `HFSpecDecMixin` Changes: - Add `dflash_offline` flag to `DFlashConfig` for training from pre-computed hidden states; deletes base model layers to save memory. - Add Pydantic validators on `DFlashConfig`: - `_derive_dflash_offline` — auto-derive `dflash_offline` from `data_args.offline_data_path` in validation context. Not user-configurable: any user-supplied value is overridden by the derived value. - `_resolve_mask_token_id` — auto-detect `dflash_mask_token_id` from `tokenizer.mask_token_id`. - `_check_mask_token_id` — fail fast if unset after resolution. - `HFDFlashModel.modify()`: select `num_orig_hidden_layers` when offline; pick `_base_model_lm_head` device when no base layers present; drop base-model `layers` module. - `HFDFlashModel.forward()`: add offline branch — consumes precomputed `base_model_outputs` via `DFlashBaseModelOutput.from_offline_dict`, and when `dflash_self_logit_distillation` is enabled with `base_model_logits` absent, recomputes logits from `base_model_hidden_states` via `_base_model_lm_head`. Raises a clear error from the non-training / `pseudo_speculative_generate` paths when `dflash_offline=True`, since base-model layers have been deleted. - `DFlashBaseModelOutput` dataclass in `modeling_dflash.py` (with `from_offline_dict` classmethod) to unify online/offline output shapes. `aux_hidden_states` is required in `from_offline_dict` so missing keys fail fast at the entry point rather than deeper in the forward. - `examples/speculative_decoding/main.py`: replace inline `mask_token_id` auto-detect with `DFlashConfig.model_validate(dflash_cfg, context={"tokenizer": tokenizer, "data_args": data_args})`. ### Silent bug fix — `add_generation_template` → `add_generation_prompt` The pre-refactor `compute_hidden_states_hf.py` passed `add_generation_template=False` to `tokenizer.apply_chat_template`. This kwarg does not exist on HF `apply_chat_template` and was being silently ignored, so the intended "don't append a generation prompt" behavior was never actually applied. The new `tokenize_with_loss_mask` helper in `examples/speculative_decoding/collect_hidden_states/common.py` uses the correct `add_generation_prompt=False`. **This is a real behavior change** for anyone re-dumping hidden states: trailing generation prompts that were previously appended to the tokenized sequences will no longer be included. ### Testing - New tests: - `tests/unit/torch/speculative/plugins/test_hf_dflash_offline.py` — CPU unit tests for convert path (online keeps base layers, offline deletes them; `num_orig_hidden_layers` drives `target_layer_ids` in offline mode) and `DFlashConfig._derive_dflash_offline` validator. - `TestDFlashOfflineForwardGPU` in `tests/gpu/torch/speculative/plugins/test_hf_dflash.py` — GPU forward smoke with precomputed `base_model_outputs`, plus the `dflash_self_logit_distillation` logit-recompute path. - training test: <img width="454" height="317" alt="image" src="https://github.com/user-attachments/assets/79b92790-4d15-4313-bb9b-f35665b012e6" /> <img width="456" height="310" alt="image" src="https://github.com/user-attachments/assets/4558559f-9c35-49ed-b36e-82fbc99eab23" /> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ — additive `dflash_offline` flag defaulting to `False`; validators fall through when context not provided. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — see Testing section above. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ ### TODO (follow-up) - [x] Update `examples/speculative_decoding/collect_hidden_states/compute_hidden_states_*.py` to support DFlash offline data. Current scripts are Eagle-specific — they hardcode the `[2, N/2, N-3]` aux-layer selection and emit `{input_ids, hidden_states, aux_hidden_states}`. DFlash offline needs: - Aux layer indices driven by `build_target_layer_ids(num_orig_hidden_layers, num_draft_layers)` (or a configurable list), not the Eagle triplet. - `base_model_hidden_states` key (last-layer hidden) so `DFlashBaseModelOutput.from_offline_dict` + the `dflash_self_logit_distillation` recompute path can consume it. - Optional `base_model_logits` dump so offline training can skip the self-distillation logit recomputation when logits are available. ### Additional Information Base branch is #1296 (file reorg). Retarget to `main` once #1296 merges. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Offline DFlash speculative-decoding training from precomputed base-model hidden states * Answer-only-loss training with persisted loss masks and optional chat-template support * Flexible auxiliary-layer selection via CLI and an exposed default aux-layer helper * Auto-derived offline flag in config and automatic memory optimization during offline conversion * **Documentation** * Updated guides for offline pipeline, aux-layer selection, and loss-masking options * **Tests** * New unit, GPU, and regression tests covering offline conversion, training, and config derivation <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
7c80d85751 |
[1/3][Refactor]: File reorg; deprecate ParallelDraft (#1296)
### What does this PR do? Type of change: refactoring Part 1 of a 3-PR series splitting #1271: - **[1/3] this PR**: File reorg + deprecate `ParallelDraft` - **[2/3] #1295**: Offline DFlash training - **[3/3] #1297**: Extract `HFSpecDecMixin` Changes: - **File reorg**: `transformers.py` → `hf_eagle.py`; extract `HFMedusaModel` → `hf_medusa.py`; extract `EagleModule` / `EagleBaseModelOutput` → `modeling_eagle.py`; extract `DFlashModule` / `DFlashAttention` / `DFlashDecoderLayer` / `build_target_layer_ids` / `apply_rotary_pos_emb` → `modeling_dflash.py`. - **Deprecate `ParallelDraft`**: remove `parallel_draft_step`, `parallel_draft_heads_num_layers`, and the `ParallelDraft` module from HF Eagle; remove the `EagleMedusaExporter` branch from `HFEagleModel.get_exporter()` (the `EagleMedusaExporter` class itself still lives in `hf_spec_export.py` for Megatron parity). - **Rename**: `_draft_model_config` → `eagle_config` in export plugin. - Update imports in `examples/speculative_decoding/` and `modelopt/torch/speculative/utils.py` to follow the module rename. ### Testing Validated with existing Eagle and DFlash training scripts (re-run after `9ae5302729 revert behavior change`). ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ — renames `modelopt.torch.speculative.plugins.transformers` → `.hf_eagle`; removes `parallel_draft_step` / `parallel_draft_heads_num_layers` from Eagle config; renames `_draft_model_config` → `eagle_config` in export plugin. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — pure refactor; existing tests updated for the rename. `test_hf_spec_rope_export.py` assertions were also corrected to reflect the actual production path (the old assertions were masked by `MagicMock` not invoking the `_draft_model_config` `@property`). - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ ### Additional Information Breaking changes: - `modelopt.torch.speculative.plugins.transformers` → `.hf_eagle` - `parallel_draft_step` / `parallel_draft_heads_num_layers` removed from Eagle config - `_draft_model_config` → `eagle_config` in export plugin <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactoring** * Reorganized speculative-decoding plugins into focused modules, converting the legacy "transformers" entry into a deprecated shim that re-exports the new plugin surface. * Consolidated DFlash implementation into a shared modeling component and introduced a dedicated EAGLE decoder module. * **New Features** * Added a Medusa speculative-decoding plugin with configurable heads and combined-loss training behavior. * **Chores** * Updated pre-commit license-hook exclusion and feature-flag wiring. * **Tests** * Updated export tests to expect rope-scaling fallback semantics. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
946639aa19 |
Fix PTQ for VLMs with image calibration (#1318)
### What does this PR do? This PR fixes PTQ with image claibration for VLMs. ### Usage ```python python3 examples/llm_ptq/hf_ptq.py --pyt_ckpt_path Qwen/Qwen3-VL-8B-Instruct --qformat fp8 --export_path Qwen3-VL-8B-Instruct-fp8 --trust_remote_code --kv_cache_qformat none --calib_with_images --calib_size 512 ``` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Image-text calibration now extends support to additional model architectures when image calibration is enabled. * Improved tokenizer truncation handling in multimodal dataset processing to prevent configuration conflicts when image inputs are present. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Liana Mikaelyan <lmikaelyan@nvidia.com> |
||
|
|
fda0899e40 |
feat(recipes): add KV cache cast variants (fp8_cast / nvfp4_cast) (#1334)
## Summary
- Adds three built-in PTQ recipes that express the KV-cache *cast*
variants directly in YAML, using the existing `use_constant_amax: true`
quantizer field. These are recipe equivalents of
`--kv_cache_qformat=fp8_cast` / `nvfp4_cast`:
- `general/ptq/fp8_default-fp8_cast_kv`
- `general/ptq/nvfp4_default-fp8_cast_kv`
- `general/ptq/nvfp4_default-nvfp4_cast_kv`
- Makes `--recipe` authoritative in `examples/llm_ptq/hf_ptq.py`: the
post-hoc `_set_kv_cache_constant_amax` override now only runs when
`--recipe is None`, so a recipe YAML fully determines KV-cache config
instead of being silently overridden by the default
`--kv_cache_qformat=fp8_cast`. Updated help text on both flags.
- Extends the recipe loader smoke test to cover the three new recipes.
## Motivation
Before this change, the cast variants lived only in argparse
(`_KV_CAST_FORMATS = {"fp8_cast", "nvfp4_cast"}`) and were layered on
top of any recipe-loaded config. That meant `--recipe
nvfp4_default-fp8_kv` would silently become a cast recipe due to the
`--kv_cache_qformat` default. Now the recipe is self-contained: its YAML
either sets `use_constant_amax: true` on the `*[kv]_bmm_quantizer` entry
(cast) or doesn't (data-driven calibration).
## Test plan
- [x] `pytest tests/unit/recipe/test_loader.py` — all 24 tests pass,
including the three new parametrized recipes.
- [x] Verified each new recipe round-trips through `load_recipe()` with
`use_constant_amax: True` surviving Pydantic validation on the KV entry.
- [x] End-to-end run on `/models/Qwen/Qwen3-8B` (RTX 6000 Ada, 4
samples, seq_len=128) for all three new recipes:
- After `mtq.quantize(model, recipe.quantize.model_dump(),
forward_loop=...)`, all 72 `k_bmm_quantizer` / `v_bmm_quantizer` modules
have `_use_constant_amax=True` and `_get_amax()` returns `448.0` (FP8
E4M3 max).
- Weight quantizers still calibrate from data normally (sample amax
values: q_proj=0.5508, k_proj=0.6250, v_proj=0.1689, o_proj=0.7266).
- [x] Verified the `--recipe` authoritative behavior change:
- Non-cast recipe + default `--kv_cache_qformat=fp8_cast` → KV entry
does NOT get `use_constant_amax` (no silent override).
- Cast recipe + contradictory `--kv_cache_qformat=fp8` → KV entry keeps
`use_constant_amax=True` (recipe wins).
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Fixed CLI to respect KV cache quantization settings from recipe YAML
instead of overriding them.
* **New Features**
* Added three new post-training quantization recipe configurations for
FP8 and NVFP4 with optimized KV cache handling.
* **Documentation**
* Enhanced CLI help text for recipe and KV cache quantization options
with configuration examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
|
||
|
|
01788bb007 |
Deprecate Mllama support in llm_ptq/vlm_ptq examples (#1332)
## Summary - Removes Mllama (Llama 3.2 Vision) model-type branches from the `llm_ptq` example (`hf_ptq.py`, `example_utils.py`) and drops the now-unused `MllamaImageProcessor` wrapper from `modelopt/torch/utils/`. - Drops the legacy `MllamaImageProcessor` path in `modelopt/torch/utils/vlm_dataset_utils.py`; the generic HF ProcessorMixin path handles the remaining cases. - Adds a CHANGELOG entry under 0.44 Backward Breaking Changes. ## Test plan - [x] CI lint / unit tests pass - [x] Smoke-run ``examples/llm_ptq/scripts/huggingface_example.sh --model <llm> --quant fp8`` (text-only path, non-mllama) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Removed Mllama (Llama 3.2 Vision) support from quantization examples. This includes removal of dedicated image processor implementation, specialized model handling, and related calibration logic. * Updated VLM image-text calibration guidance to use `--calib_with_images` flag with other supported VLMs instead of Mllama-specific processing paths. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
8663678f12 |
fix: bug hf_ptq.py max_length setting ignored for LLMs (#1311)
### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage - fixes a bug in example script. We were trying why our models were not that strong at long context. Seems like a recent refactor did not implement max seq length. so 512 is used by default. ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Improvements** * Calibration data loading now enforces a maximum sequence/sample length during dataset preparation, ensuring calibration inputs adhere to configured length limits. This yields more predictable calibration behavior, reduces peak memory usage during calibration runs, and improves consistency of quantization preprocessing. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Michael Feil <63565275+michaelfeil@users.noreply.github.com> |
||
|
|
e4e3508f51 |
[OMNIML-3349] Add FP8 MHA quantization support for HuggingFace ViT (#1289)
## Summary Enables TensorRT attention-v2 fusion for vision transformers when exported to ONNX with FP8 Q/DQ. The core library changes are architecture-agnostic (drop-in for any FP8 ONNX export); coverage is exercised by the existing `examples/torch_onnx/torch_quant_to_onnx.py` pipeline. - **`modelopt/onnx/export/fp8_exporter.py`** — new post-processing passes: move attention-scaling `Mul` and K `Transpose` to the Q-side so DQ feeds MatMul directly, pre-transpose constant weights, and insert FP8 Q/DQ on Softmax outputs (fixed `1/448` scale, data-independent) for MHA-v2 fusion. Rewrites only fire when every downstream consumer is a MatMul so non-attention branches are never perturbed. - **`modelopt/onnx/utils.py`** — `fold_dq_fp32_to_fp16_casts` / `fold_q_fp16_to_fp32_casts` remove the Cast nodes `convert_float_to_float16` inserts around Q/DQ and rewrite scale initializers to FP16 so TRT fuses DQ into the downstream GEMM. Guarded behind opset >= 19 (FP16 Q/DQ scale requirement). Warns on FP16 overflow/underflow. - **`modelopt/torch/_deploy/utils/torch_onnx.py`** — calls the fold helpers for FP8-quantized models after `convert_float_to_float16`. - **`modelopt/torch/quantization/export_onnx.py`** — keeps FP8 Q/DQ scale in the native input dtype so no Cast is emitted between graph and Q/DQ. Removes the now-unused `trt_high_precision_dtype` parameter from `_fp8_quantize`/`_fp8_dequantize`. - **`modelopt/torch/quantization/nn/modules/quant_layernorm.py`** (new) — registers `nn.LayerNorm` in `QuantModuleRegistry` so LayerNorm output quantizers are honored. - **`modelopt/torch/quantization/plugins/huggingface.py`** — skips `*Attention` wrappers whose children are also `*Attention` per-instance (not per-class) to avoid double-patching `eager_attention_forward` (e.g. `ViTAttention` vs `ViTSelfAttention`). - **`examples/torch_onnx/torch_quant_to_onnx.py`** — adds a `_FP8_MHA_OVERRIDE` config block to FP8 mode that enables LayerNorm output quantizer + disables its input quantizer for TRT attention fusion. - **Unit tests** (12 CPU tests, ~1.2s total) — fp8_exporter rewrites + fanout safety, fold-cast helpers + opset guard, LayerNorm quant-wrapper identity, per-instance nested-attention detection. ## Benchmarks ViT-base-patch16-224, RTX 6000 Ada, strongly-typed FP8 via `trtexec`. Accuracy on 2 000 ImageNet-1k validation samples (streaming). **Batch = 1 (latency-bound)** | Model | Top-1 | Top-5 | TRT latency | Speedup | |---|---|---|---|---| | FP16 baseline | 80.96% | 95.80% | 0.722 ms | 1.00x | | Torch FP8 MHA | 80.66% | 95.75% | 0.657 ms | **1.10x** | | ONNX PTQ FP8 | — | — | 0.589 ms | **1.23x** | **Batch = 64 (throughput-bound, realistic inference)** | Model | TRT latency | Speedup | Images/s | |---|---|---|---| | FP16 baseline | 23.40 ms | 1.00x | 1152 | | Torch FP8 MHA | 15.89 ms | **1.47x** | 1152 | | ONNX PTQ FP8 | 15.89 ms | **1.47x** | 1216 | Top-1 accuracy stays within 0.30 pp of FP16; at batch=64 the Torch FP8 MHA path matches ONNX PTQ wall-time — attention is the bottleneck there and both paths achieve full FP8 attention fusion (36/36 attention MatMuls with QDQ in ViT-base). ## Test plan - [x] CPU unit tests (new): \`python -m pytest tests/unit/onnx/quantization/test_fp8_mha_exporter.py tests/unit/onnx/test_fold_casts.py tests/unit/torch/quantization/test_quant_layernorm.py tests/unit/torch/quantization/plugins/test_nested_attention_skip.py\` - [x] Existing ONNX / quantization unit suites unaffected: \`python -m pytest tests/unit/onnx tests/unit/torch/quantization\` - [x] End-to-end ViT FP8 export: \`python examples/torch_onnx/torch_quant_to_onnx.py --timm_model_name vit_base_patch16_224 --quantize_mode fp8 --onnx_save_path vit_base_fp8.onnx\` — expect log lines \`Folded 48 weight Transpose nodes\`, \`Inserted FP8 weight DequantizeLinear for 1 Conv nodes\`, and \`Attention QDQ rewrites: ... inserted QDQ on 12 Softmax outputs\` - [x] trtexec FP8 strongly-typed build: \`trtexec --onnx=vit_base_fp8.onnx --fp8 --stronglyTyped\` - [x] Accuracy within ~0.3 pp of FP16 baseline on ImageNet-1k subset --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> |
||
|
|
7265ca6793 |
Fix lm_eval version checking (#1321)
### What does this PR do? Type of change: bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> `lm_eval` does not have `__version__` attribute ### Additional Information <!-- E.g. related issue. --> NVBug 6102101 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Enhanced the package version detection system to improve overall reliability and stability of the application while reducing unnecessary external dependencies. All functionality, including version gating and system warnings, continues to operate exactly as expected with no impact on the user experience. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
2564da7324 |
Update vLLM deployment docs for heterogeneous models (#1317)
### What does this PR do? Type of change: ? documentation. This PR updates vLLM deployment instructions, taking into account heterogenous models created with AnyModel. ### Usage Does not apply. ### Testing Run the updated instructions in the documentation. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Replaced a benchmarking-focused section with a deployment guide for running compressed models on vLLM. * Added step-by-step setup for using an AnyModel-enabled vLLM fork, including checkout and install guidance and required model config edits (with optional architecture metadata). * Simplified runtime to a single vllm serve command, removing manual model rearrangement steps. * Restored inference benchmarking as a subsection, retaining vllm bench latency/throughput examples. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Grzegorz Karch <gkarch@nvidia.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
c7966119eb |
Reorg the sparse/quant/common kernel dir (#1303)
### What does this PR do? Type of change: re-org code <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ We changed the import path <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!--- Mandatory --> - Did you write any new necessary tests?: ❌ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Calibration support for skip-softmax multi-threshold measurement in sparse attention. * N:M sparse softmax masking and helpers for sparsity-aware attention. * **Chores** * Reorganized and consolidated kernel/backends for quantization and sparsity to a unified kernels layout, updating tests and examples to match. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
0678136335 |
Fix vLLM fakequant MoE megatron export bug (#1305)
### What does this PR do?
Type of change: Bug fix
Fixes two bugs in the vLLM + Megatron-Core MoE export path and cleans up
the related weight-collection helper:
1. **`_QuantFusedMoEBase` (vllm.py)**: The weight-quantizer path in
`_invoke_fused_moe_quantized_function` was temporarily mutating
`self.w13_weight` / `self.w2_weight` to the quantized tensor, then
restoring them via `finally`. This exposed a stale quantized tensor on
`self` between the mutation and the kernel call. Fixed by computing the
quantized weight directly into a local `B` without touching `self.*`
attributes.
2. **`GPTModelExporter` / `VllmFqGPTModelExporter`
(unified_export_megatron.py / vllm_fakequant_megatron.py)**:
`expert_bias` (present in grouped MoE layers) was silently dropped
during export because the bias collection ran after the early-return on
missing `weight`. Extracted a `_get_weight_bias` helper that collects
weight, bias, and expert_bias together, so bias/expert_bias are captured
even when weight is absent or zero-element.
### Usage
```python
# No API change; export pipelines pick this up automatically.
# export_mcore_gpt_to_hf_vllm_fq / export_mcore_gpt_to_hf now correctly
# export expert_bias for grouped-MoE checkpoints.
```
### Testing
Step 1 — Quantize (run from Megatron-LM
examples/post_training/modelopt):
```
HF_MODEL_CKPT=<path/to/hf/weights> MLM_MODEL_SAVE=<quant-ckpt-name> \
bash quantize.sh <hf-model-id> NVFP4_DEFAULT_CFG
```
Step 2 — Export for vLLM fakequant:
```
MLM_EXTRA_ARGS=--export-vllm-fq \
HF_MODEL_CKPT=<path/to/hf/weights> \
MLM_MODEL_CKPT=<quant-ckpt-name> \
EXPORT_DIR=<export-dir> \
bash export.sh <hf-model-id>
```
Step 3 — Serve (run from examples/vllm_serve):
```
QUANT_CFG=NVFP4_DEFAULT_CFG \
QUANT_FILE_PATH=<export-dir>/quantizer_state.pth \
python3 vllm_serve_fakequant.py <export-dir> \
-tp 1 --served-model-name <model-name> \
--host 0.0.0.0 --port 8000 \
--trust-remote-code --enforce-eager \
--disable-custom-all-reduce \
--gpu-memory-utilization 0.8
```
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Refactor**
* Centralized weight/bias/expert-bias extraction and export to a single
helper for consistent handling.
* Standardized quantized-weight flow to temporarily swap and restore
parameter tensors during computation.
* **Bug Fixes**
* Prevented missing or incorrect weight/bias exports by unifying
extraction logic.
* Broadened checkpoint key matching to preserve more quantizer state
during reloads.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
|
||
|
|
2fef374ded |
fix: auto-compute dp_replicate_size from world_size (#1302)
## Summary - When `dp_shard_size < world_size` (e.g., `dp_shard_size=4` on 8 GPUs across 2 nodes), `ParallelismConfig` raises `total_size (4) does not match num_processes (8)` because `dp_replicate_size` defaults to 1 - Auto-compute `dp_replicate_size = world_size // (dp_shard_size * cp_size)` so intra-node FSDP2 sharding + inter-node data-parallel replication works without manual config - This enables `dp_shard_size` to be set to per-node GPU count (better NVLink utilization) while automatically creating replicas across nodes ## Test plan - [ ] Verify single-node training (dp_shard_size == world_size, dp_replicate_size == 1) unchanged - [ ] Verify multi-node with dp_shard_size < world_size creates correct replica groups - [ ] Verify existing EAGLE3/DFlash configs still work 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Enhanced parallelism configuration initialization in the speculative decoding example to better handle distributed training scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
355c6b7883 |
fix: PTQ 1GPU, export PP divisibility, hidden states conversations key (#1293)
## Summary - **megatron_lm_ptq.yaml**: Qwen3-8B PTQ to single GPU for L40 clusters (TP=1, all tasks) - **quantize.sh**: Auto-find largest PP dividing model's `num_hidden_layers` for export step. Qwen3-8B has 36 layers which isn't divisible by 8, causing `AssertionError` on 8-GPU nodes - **compute_hidden_states_trtllm.py**: Use `messages` with `conversations` fallback, matching the HF version. Fixes `KeyError: 'conversations'` when data uses OpenAI `messages` format ## Test plan - [x] Qwen3-8B PTQ runs on single L40 GPU - [x] Export PP auto-selects valid divisor (36 layers → PP=6 on 8 GPUs, PP=4 on 4 GPUs, PP=1 on 1 GPU) - [x] EAGLE3 offline pipeline reads data with `messages` field 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Dataset input handling now supports multiple field formats for enhanced compatibility. * **Bug Fixes** * Optimized GPU resource allocation during model quantization with improved pipeline parallelism computation. * Updated quantization configuration for more efficient resource utilization. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
010b220dc0 |
vLLM fakequant export update for AWQ checkpoint (#1242)
### What does this PR do?
Type of change: Bug
Enables end-to-end AWQ checkpoint export and reload in the vLLM
fake-quant serving path (`MODELOPT_STATE_PATH`). Previously, the
`input_quantizer` was using incorrect `pre_quant_scale` especially with
grouped quantizers like `qkv_proj`, using simply the first
`input_quantizer.pre_quant_scale`. This MR adds
`_resmooth_experts_for_export` that non-mutatively averages
`pre_quant_scale` across MoE experts and unifies input `_amax`, required
because vLLM uses a single input quantizer per expert group. Adds
`merge_amax_tensors_for_group` (element-wise max for same-shape, `cat`
for GQA, scalar-max fallback) replacing the scalar-collapsing
`torch.stack().max()` that dropped per-channel `_amax` structure.
### Usage
```python
# Export AWQ checkpoint from HF model
from modelopt.torch.export.plugins.vllm_fakequant_hf import export_hf_vllm_fq_checkpoint
export_hf_vllm_fq_checkpoint(model, export_dir="./awq_vllm_checkpoint")
```
### Testing
**Step 1 — Export the quantized checkpoint:**
```bash
python examples/llm_ptq/hf_ptq.py \
--pyt_ckpt_path <MODEL_PATH> \
--recipe <AWQ_RECIPE> \
--calib_size 512 \
--export_path <EXPORT_DIR> \
--vllm_fakequant_export
```
This produces `<EXPORT_DIR>/vllm_fq_modelopt_state.pth` with the averaged per-expert
pre_quant_scale and unified _amax now included.
Step 2 — Serve via vLLM fakequant worker:
```bash
MODELOPT_STATE_PATH=<EXPORT_DIR>/vllm_fq_modelopt_state.pth \
python examples/vllm_serve/vllm_serve_fakequant.py \
<EXPORT_DIR> --tensor-parallel-size <TP>
```
Tested for quantization configurations:
```
FP8_DEFAULT_CFG
FP8_DEFAULT_CFG (input_q disabled)
INT8_SMOOTHQUANT_CFG
INT8_WEIGHT_ONLY_CFG
NVFP4_DEFAULT_CFG
NVFP4_AWQ_LITE_CFG
INT4_AWQ_CFG
NVFP4_AWQ_CFG
NVFP4_DEFAULT_CFG (input_q disabled)
```
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit
* **New Features**
* Added Nemotron-style MoE export support and group-aware AWQ resmoothing with optional requantization during export.
* Improved handling for shared-input / expert groups and tensor-parallel sharding of pre-quantization scales.
* **Bug Fixes**
* Removed AWQ reload limitation from known issues; improved checkpoint validation and safer save/load behavior.
* Better detection and handling of enabled weight-quantizers and clearer warnings for mismatched checkpoint keys.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
|
||
|
|
26ae8da517 |
[2/3] Implicit Gemm NVFP4 (#1227)
### What does this PR do? Type of change: new feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> - Add Conv3D implicit GEMM kernel with BF16 WMMA tensor cores and fused NVFP4 activation quantization for video diffusion VAE layers - Integrate into _QuantConv3d via QuantModuleRegistry — automatically dispatched when NVFP4 quantization is applied to nn.Conv3d - Move kernel from `experimental/conv/ to modelopt/torch/kernels/conv/`; move tests to `tests/gpu/torch/quantization/kernels/` ### Testing <!-- Mention how have you tested your change if applicable. --> - Added test cases to measure the difference between cuDNN and our CUDA implicit GEMM kernel - Added an NVFP4 fake quantization test using CUDA code ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Per-backbone quantization/export in a single run with per-backbone checkpoints and backbone-aware quant filters * Configurable NVFP4 block-size via CLI/config; improved NVFP4 Conv3D inference path and Wan 2.2 quantization support * **Bug Fixes** * Video-model calibration now respects extra params and forces video decoding during calibration * **Documentation** * Added comprehensive Conv3D implicit‑GEMM kernel documentation; removed experimental Conv3D prototype docs/benchmark * **Tests** * New Wan 2.2 quantization/export tests and expanded Conv3D/FP4 kernel test coverage <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
760c980727 |
Add ResNet50 support for torch_onnx quantization workflow (#1263)
## Summary - Add end-to-end ResNet50 support in the torch_onnx quantization → ONNX export → TRT engine pipeline - Fix multiple Conv2d-related export issues that blocked Conv2d-heavy models from working with FP8/INT8/MXFP8/NVFP4/auto quantization modes - Fix `configure_linear_module_onnx_quantizers` to handle all modules with block quantization (not just `nn.Linear`), fixing NVFP4/MXFP8 export for models with quantized non-Linear modules - Add `--trt_build` flag to `torch_quant_to_onnx.py` and simplify test infrastructure ### Files Changed - `modelopt/torch/_deploy/utils/torch_onnx.py` — Disable FP8 Conv2d weight quantizers and autocast during ONNX export - `modelopt/torch/quantization/export_onnx.py` — Fix `configure_linear_module_onnx_quantizers` for all module types with block quantization - `examples/torch_onnx/torch_quant_to_onnx.py` — Add `--trt_build` flag, calibration for FP8 override quantizers, Conv2d→FP8 override for auto mode, filter_func updates - `examples/torch_onnx/README.md` — Add ResNet50 to supported models table - `tests/examples/torch_onnx/test_torch_quant_to_onnx.py` — Add ResNet50 test entry, simplify using `--trt_build` - `tests/_test_utils/torch/vision_models.py` — Add ResNet50 to timm model registry ### Quantization modes passing - ✅ FP8, INT8, MXFP8, NVFP4, Auto (all 5 modes pass export + TRT build) - INT4_AWQ excluded (pre-existing limitation for all models) ## Test plan - [x] All 5 resnet50 test modes pass: `pytest tests/examples/torch_onnx/test_torch_quant_to_onnx.py -k resnet50` (5/5 passed) - [x] Full regression: 18 passed, 2 failed (pre-existing swinv2_tiny fp8/int8 failures) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added ResNet50 to supported ONNX export vision models with FP8, INT8, MXFP8, and NVFP4 support. * Optional TensorRT engine build after export via a new CLI flag. * **Improvements** * Enhanced quantization calibration and export flows for FP8/INT8 models, including broader block-quantization support across module types and safer export handling. * Tests updated to include ResNet50 in the model matrix. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> Signed-off-by: ajrasane <arasane@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
92622a9aa6 |
Add nvfp4_mse and nvfp4_local_hessian options to the ptq script. (#1113)
### What does this PR do? Type of change: Bugfix <!-- Details about the change. --> Add newly added quant configs to the example PTQ script. ### Testing I have locally run auto_quantize with these two quant_configs, and obtained successfully exported HF artifacts. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for three new quantization formats: nvfp4_mse, nvfp4_local_hessian, and nvfp4_experts_only, expanding available export options when using auto-quantize. * **Bug Fixes / UX** * Updated the invalid-quantization error message to include the newly accepted format identifiers. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Bilal Kartal <bkartal@nvidia.com> Signed-off-by: bkartal-dev <bkartal@nvidia.com> |
||
|
|
feec81ad2b |
Add the Skip softmax for diffusion (#1166)
### What does this PR do?
Type of change: new feature, new example <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->
<!-- Details about the change. -->
## Summary
- Add skip-softmax sparse attention (BLASST) for diffusion models via
dedicated Triton kernels — an inference kernel with tile skipping and a
calibration kernel with vectorized multi-threshold sparsity measurement
- Add `triton_skip_softmax` method with exponential model calibration
(`scale_factor = a * exp(b * sparsity)`) and log-space fitting for
diffusion models
- Add Triton kernel backends for diffusers and LTX attention dispatch
- Fix calibration to skip RULER dataset generation when user provides
their own `forward_loop` (required for non-LLM models)
## Changes
### Triton kernels (`modelopt/torch/kernels/triton_fa.py`)
- **`_attn_fwd`**: Forward kernel with optional tile skipping — tiles
whose max attention score is far below the running softmax max are
skipped entirely (no V load, no softmax, no accumulation). Runtime
sparsity measurement via atomic counters.
- **`_attn_fwd_calibrate`**: Calibration kernel that computes full
attention while measuring how many tiles would be skipped at each of N
thresholds simultaneously. Uses per-program output buffers (zero atomic
contention) and vectorized multi-threshold comparison.
- **`attention()`** / **`attention_calibrate()`**: Python wrappers for
inference and calibration kernels.
### Kernel backends
(`modelopt/torch/sparsity/attention_sparsity/kernels/`)
- **`diffusers_triton_attention.py`**: Registers `modelopt_triton`
backend in diffusers' attention dispatch. Handles [B, S, H, D] → varlen
layout conversion, calibration/inference mode switching, thread-local
configuration, and counter accumulation.
- **`ltx_triton_attention.py`**: Patches `ltx_core.Attention` modules
for Triton dispatch with the same calibration/inference modes.
### Method
(`modelopt/torch/sparsity/attention_sparsity/methods/triton_skip_softmax.py`)
- `TritonSkipSoftmaxMethod`: Context managers for calibration (→
calibration kernel) and inference (→ forward kernel with tile skipping).
Three threshold priority levels: raw threshold > calibrated scale_factor
> static threshold.
### Calibration
(`modelopt/torch/sparsity/attention_sparsity/calibration/`)
- **`calibrator.py`**: `DynamicThresholdCalibrator` with `fit_logspace`
option — fits exponential model in log space (minimizes relative error)
for diffusion models where scale_factors span many orders of magnitude.
Records observed sparsity range for extrapolation warnings.
- **`calibrate.py`**: Skips RULER dataset when `forward_loop` is
provided; passes `fit_logspace` through from config.
### Config & conversion
- **`config.py`**: `CalibrationConfig.fit_logspace` field (default
False, recommended True for diffusion models).
`skip_softmax_raw_threshold` field for direct threshold mode.
- **`conversion.py`**: Auto-registers diffusers/LTX Triton backends on
`sparsify()`. Updated summary display.
### Example
- **`wan22_skip_softmax.py`**: End-to-end example for WAN 2.2 5B/14B
with baseline, raw-threshold, and calibrated modes. Supports runtime
sparsity reporting.
## Threshold modes
| Mode | How it works | Use case |
|------|-------------|----------|
| **Raw threshold** (`--raw-threshold -0.7`) | Passed directly to kernel
as `skip_threshold_log2` | Quick testing, sweeps |
| **Calibrated** (`--calibrate --target-sparsity 0.5`) | `scale_factor =
a * exp(b * target)`, then `threshold = scale_factor / seq_k` at runtime
| Production use with seqlen adaptation |
| **Static** (default `skip_softmax_threshold=0.1`) | `log2(lambda) *
sm_scale` | Fallback |
## Usage
```bash
# Fixed raw threshold (no calibration)
python examples/diffusers/sparsity/wan22_skip_softmax.py \
--model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
--raw-threshold -0.7 \
--prompt "A cat playing piano" --output out.mp4
# With calibration (log-space fit for diffusion models)
python examples/diffusers/sparsity/wan22_skip_softmax.py \
--model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
--calibrate --target-sparsity 0.5 \
--prompt "A cat playing piano" --output out.mp4
# Dense baseline for comparison
python examples/diffusers/sparsity/wan22_skip_softmax.py \
--model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
--baseline \
--prompt "A cat playing piano" --output baseline.mp4
```
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added skip-softmax sparse attention support for Diffusers models,
enabling efficient video generation
* Added support for both eager and Triton attention backends for sparse
attention
* Added new example script for Wan 2.2 text-to-video generation with
sparse attention optimization
* **Documentation**
* Updated documentation with sparse attention configuration guide and
usage examples
* **Tests**
* Added comprehensive unit tests for kernel backend registration and
skip-softmax functionality
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
|
||
|
|
2d868d3f1f |
Performant layerwise calibration for large models (#1251)
## Summary
Adds **performant layerwise calibration** for quantizing large models
(e.g. DeepSeek-R1 671B) that don't fit entirely on GPU. ([Example
commands](#example-commands))
1. **Performant calibration for large models** — Each decoder layer is
moved from CPU/disk to GPU (accelerate) or unsharded (FSDP2) **only
once** and kept on GPU for the entire calibration step. Previously,
every calibration batch triggered weight transfer for every layer —
O(num_batches) weight movements per layer. Now it is O(1) per layer.
This also means you can **increase batch size** since only one layer's
weights occupy GPU at a time — e.g. DeepSeek-R1 on a single node
(8×80GB) with `batch_size=16` and `gpu_max_mem_percentage=0.5`.
2. **Checkpoint save/resume** — Saves progress after each layer, so jobs
that exceed cluster time limits (e.g. 4-hour Slurm windows for 100+
layer MoE models) can resume from the last completed layer.
3. **Rename** `sequential_calibrate` → `layerwise_calibrate` for
clarity.
### Design details
The existing layerwise state machine (skip/run/capture) already
processes one layer at a time, but skip-mode layers still kept their
parameters in the ModuleList — so frameworks transferred all weights
every forward pass. This PR adds:
- **`_SkipLayer`**: replaces fully-calibrated layers with a
parameter-free dummy in the ModuleList, so framework hooks have nothing
to transfer
- **`persistent_materialization`**: keeps the active layer on GPU for
the entire calibration step, avoiding repeated offload/reload cycles
Checkpoint save is per-layer; restore is bulk — quantizer state and
weights for layers 0..K-1 are restored once at the end of calibration,
keeping the hot path fast.
### Example commands
**Qwen3-8B** (NVFP4+GPTQ, single GPU):
```bash
python hf_ptq.py \
--pyt_ckpt_path Qwen/Qwen3-8B \
--recipe nvfp4_gptq_sequential.yaml \
--calib_size 64 \
--batch_size 16 \
--dataset cnn_dailymail \
--export_path outputs/qwen3_8b_nvfp4_gptq_seq \
--gpu_max_mem_percentage 0.5 \
--use_seq_device_map \
--vllm_fakequant_export
```
**DeepSeek-R1** (NVFP4 experts-only + FP8 KV, 8×80GB):
```bash
python hf_ptq.py \
--model unsloth/DeepSeek-R1-0528-BF16 \
--recipe ../../modelopt_recipes/general/ptq/nvfp4_experts_only-fp8_kv.yaml \
--dataset cnn_dailymail \
--batch_size 16 \
--calib_size 64 \
--calib_seq 512 \
--gpu_max_mem_percentage 0.5 \
--use_seq_device_map \
--trust_remote_code \
--export_path output/DeepSeek-R1-BF16-nvfp4-experts-only-fp8-kv \
--vllm_fakequant_export
```
### Example: NVFP4+GPTQ layerwise calibration on Qwen3-8B (36 layers,
single GPU — 20 GB peak)
**Initial run** (killed after layer 11):
```
Layerwise calibration: Found 36 transformer layers
Calibrating layer 1/36 | capture: [1]
Computing Hessians for 7 linear layers...
GPTQ time: 51.39s
Calibrating layer 2/36 | run: [1] | capture: [2]
Checkpoint: saved layer 0
GPTQ time: 50.06s
Calibrating layer 3/36 | skip: 1 | run: [2] | capture: [3]
Checkpoint: saved layer 1
...
Calibrating layer 12/36 | skip: 10 | run: [11] | capture: [12]
Checkpoint: saved layer 10
<killed>
```
**Resumed run** (picks up from layer 11, finishes all 36):
```
Layerwise calibration: Found 36 transformer layers
Checkpoint: resuming layerwise calibration from layer 11/36
Calibrating layer 12 (resumed)
GPTQ time: 51.45s
Calibrating layer 13/36 | skip: 11 | run: [12] | capture: [13]
Checkpoint: saved layer 11
...
Calibrating layer 36/36 | skip: 34 | run: [35] | capture: [36]
Checkpoint: saved layer 34
GPTQ time: 50.33s
Checkpoint: saved layer 35 (final)
Checkpoint: restored 11 previously calibrated layers
Layerwise calibration completed
Quantized model exported to: outputs/qwen3_8b_nvfp4_gptq_seq
GPU 0: Peak memory usage = 20.42 GB
```
## TODO
- [ ] Update CHANGELOG
## Test plan
- `tests/unit/torch/quantization/test_layerwise_calibrate.py` — unit
tests for skip/swap/restore
- `tests/unit/torch/quantization/test_sequential_checkpoint.py` —
checkpoint save/resume correctness
- `tests/gpu/torch/quantization/plugins/test_accelerate_gpu.py` —
CPU-offloaded layerwise + GPTQ + checkpoint resume
- `tests/gpu/torch/quantization/test_fsdp2.py` — FSDP2 layerwise
calibration
### Verified
- [x] Qwen3-8B: layerwise calibration + checkpoint save/restore +
fakequantized checkpoint export + vLLM serve
- [x] DeepSeek-R1: checkpoint resume tested
- [x] DeepSeek-R1: fakequantized checkpoint export verified
---------
Signed-off-by: realAsma <akuriparambi@nvidia.com>
|
||
|
|
dc7ad66b71 |
GPTQ vector (#1223)
### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added backend-specific GPTQ helper registration to allow backend-tailored GPTQ behavior. * **Bug Fixes** * Prevented KV-cache state from leaking across repeated per-layer forwards during calibration. * **Tests** * Added GPU-focused tests validating GPTQ combined with vector quantization, including accuracy and end-to-end comparisons. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com> |
||
|
|
e4b054bf32 |
Fix and Speedup megatron_mmlu by >10x via prefill scoring and global batching (#1280)
### What does this PR do?
Type of change: new feature + bug fix
Two improvements to Megatron inference utilities:
**1. Pipeline Parallel (PP) correctness fixes**
PP inference was producing garbage output (MMLU ~0.24, random chance).
Two root causes:
- `megatron_generate` / `megatron_prefill` used
`get_forward_backward_func()` (the training pipeline scheduler), which
is not designed for inference. Rewrote both functions to use explicit
P2P communication via `recv_from_prev_pipeline_rank_` /
`send_to_next_pipeline_rank`, matching the `run_mcore_inference`
pattern.
- `import_mcore_gpt_from_hf` loads HF weights into stage 0's embedding
but never updates the output_layer on the last PP stage when
`share_embeddings_and_output_weights=True`. At model init,
`setup_embeddings_and_output_layer()` all-reduces from stage 0 to sync
the output layer; after importing HF weights that all-reduce is stale.
Fix: call `model.setup_embeddings_and_output_layer()` again after
import.
**2. `megatron_mmlu` speedup (~6x)**
Replaces the `megatron_mmlu` implementation with a significantly faster
approach that matches how `lm-evaluation-harness` scores multiple-choice
questions.
**Before:** autoregressive generation (`megatron_generate`, `osl=2`) per
example, 114 separate `load_dataset` calls, batch_size=1 — 260s for 5%
data.
**After:** single prefill forward pass + argmax over {A,B,C,D} logits, 2
`load_dataset` calls, configurable batch_size — 18s for 5% data (~6x
faster).
### Changes
**PP fixes:**
- `megatron_generate` / `megatron_prefill`: replace
`get_forward_backward_func` with explicit P2P
(`recv_from_prev_pipeline_rank_` / `send_to_next_pipeline_rank`)
- `import_mcore_gpt_from_hf`: call
`model.setup_embeddings_and_output_layer()` after HF weight import when
PP>1 and `share_embeddings_and_output_weights=True`
- `megatron_prefill`: add `skip_return_logits` param and VLM support
(needed for PP non-last stages)
**MMLU speedup:**
- **Log-likelihood scoring**: replace `megatron_generate` with
`megatron_prefill` — one forward pass per batch, no autoregressive
decode loop
- **Global batching**: collect all examples across all subjects, sort by
descending sequence length, run in `batch_size` chunks
- **2 dataset loads** instead of 114: use `load_dataset("cais/mmlu",
"all")` with per-subject grouping; skip dev load when `few_shots=0`
- **`percentage` → `fraction`** parameter rename for clarity
- **tqdm progress bar** (rank-0 only)
### Testing
- `test_megatron_generate_and_mmlu` parametrized over `tp` and `pp`.
Accuracy assertion: `0.36 < score < 0.39`. Manually checked generated
text is coherent.
- Re-ran M-Bridge Minitron MMLU based pruning for Nano v2 9B -> 7B and
all top 10 candidate's MMLU numbers are ballpark similar as before
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ — `percentage` parameter
renamed to `fraction`; `enable_kv_cache` removed from `megatron_mmlu`
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — existing test updated and
parametrized for TP+PP
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
🤖 Generated with [Claude Code](https://claude.ai/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved pipeline-parallel generation and MMLU evaluation reliability;
fixed output-layer synchronization in shared-embedding + pipeline
setups.
* **New Features**
* MMLU scoring now uses batched prefill logit scoring for faster,
batched evaluation.
* **Behavior Changes**
* Default MMLU sampling increased from 5% to 10%; calibration batch
sizing adjusted and related CLI/help text updated.
* **Tests**
* Distributed tests cover tensor- and pipeline-parallel modes and
tighten MMLU validation ranges.
* **Documentation**
* Updated pruning example and benchmark timing to reflect new sampling
and speedup.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
||
|
|
fe8c5178c7 |
Removed version fixes for torch transformers in windows ptq example requirements (#1275)
### What does this PR do? Type of change: Bug fix Removed version fixes for torch and transformers ### Testing Tested quantization with a couple of models . Working as expected. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Relaxed dependency specs: removed strict pin for torch to allow latest compatible installs, and constrained transformers to <5.0.0 for broader compatibility and easier updates. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com> Signed-off-by: Hrishith Thadicherla <99313418+hthadicherla@users.noreply.github.com> |
||
|
|
f238d93d86 |
vLLM fakequant fold weight_quantizer for megatron export (#1246)
### What does this PR do? Type of change: Bug fix During Megatron→vLLM fakequant export (`export_mcore_gpt_to_hf_vllm_fq`), the `weight_quantizer` is now applied as fake-quantization (quantize + dequantize) directly into the exported weight tensor, and its amax is no longer saved to `quantizer_state.pth`. On reload, if `weight_quantizer` keys are absent from the checkpoint (because they were folded at export time), the corresponding quantizer modules are disabled. This change is useful especially when amax across experts are not synced for `weight_quantizer`, this allows the `weight_quantizer` to keep them different for better accuracy. ### Usage ```python # Unchanged — export API is the same export_mcore_gpt_to_hf_vllm_fq(model, pretrained_model_name_or_path=..., export_dir=...) ``` ### Testing Step 1 — Quantize (run from Megatron-LM `examples/post_training/modelopt`): ```bash HF_MODEL_CKPT=<path/to/hf/weights> MLM_MODEL_SAVE=<quant-ckpt-name> \ bash quantize.sh <hf-model-id> NVFP4_DEFAULT_CFG ``` Step 2 — Export for vLLM fakequant: ```bash MLM_EXTRA_ARGS=--export-vllm-fq \ HF_MODEL_CKPT=<path/to/hf/weights> \ MLM_MODEL_CKPT=<quant-ckpt-name> \ EXPORT_DIR=<export-dir> \ bash export.sh <hf-model-id> ``` Step 3 — Serve (run from examples/vllm_serve): ```bash QUANT_CFG=NVFP4_DEFAULT_CFG \ QUANT_FILE_PATH=<export-dir>/quantizer_state.pth \ python3 vllm_serve_fakequant.py <export-dir> \ -tp 1 --served-model-name <model-name> \ --host 0.0.0.0 --port 8000 \ --trust-remote-code --enforce-eager \ --disable-custom-all-reduce \ --gpu-memory-utilization 0.8 ``` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Better handling when loading checkpoints: missing weight-quantizer entries are validated and corresponding modules are disabled to avoid load failures. * **Improvements** * Export now folds enabled weight quantizers into exported weights when present and omits internal weight-quantizer tensors from the exported state to produce cleaner exports. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> |
||
|
|
07ae8e7128 |
Add LoRA co-training support for HF EAGLE speculative decoding (#1060)
### What does this PR do?
Type of change: New feature + bug fixes
Adds **LoRA co-training** support for HF EAGLE speculative decoding.
When `eagle_base_lora=True`, HF PEFT LoRA adapters are injected into the
base model and co-trained alongside the EAGLE draft module in a single
online training pass. A preservation loss (KL divergence between the
original frozen base model output and the LoRA-adapted output) prevents
base model drift. LoRA adapter weights are exported in standard peft
format alongside EAGLE draft artifacts.
### Key features
- **LoRA injection**: `peft.inject_adapter_in_model` applied in-place
(no wrapper), keeping the existing `HFEagleModel` structure intact.
- **Preservation loss**: Cross-entropy `H(ref, lora)` — equivalent
gradient to `KL(ref || lora)` since `H(ref)` is constant w.r.t. LoRA
params.
- **Warmup schedule**: `eagle_base_lora_warmup_steps` freezes LoRA for N
steps while the EAGLE head stabilizes, then enables co-training via a
`LoRAWarmupCallback`.
- **Logits detach regularization**: `eagle_base_lora_logits_detach_prob`
stochastically detaches base logits from the EAGLE loss path, preventing
LoRA from degenerating to maximize EAGLE accuracy at the cost of base
model quality.
- **Export**: Standard peft format (`adapter_model.safetensors` +
`adapter_config.json`) alongside EAGLE draft model.
- **Merge script**: `scripts/merge_lora.py` merges LoRA weights into the
base model and restores the original `config.json` (avoids transformers
5.x rewriting `rope_theta` → `rope_parameters` which breaks
vLLM/TRT-LLM).
- **Multinode fix**: `dp_shard_size` now uses `WORLD_SIZE` instead of
local GPU count.
### Config options
```python
mtsp.convert(model, mode=[("eagle", {
"eagle_base_lora": True, # enable LoRA co-training
"eagle_base_lora_rank": 64, # LoRA rank
"eagle_base_lora_alpha": 16.0, # LoRA scaling
"eagle_base_lora_target_modules": ["q_proj", "k_proj", "v_proj", "o_proj"],
"eagle_base_lora_preservation_loss_weight": 0.1, # preservation loss weight
"eagle_base_lora_warmup_steps": 0, # freeze LoRA for N steps
"eagle_base_lora_logits_detach_prob": 0.5, # detach prob (0=never, 1=always)
})])
```
### Experimental results (Qwen3-8B, checkpoint-60000)
Base model quality preserved across detach_prob sweep (lm_eval: IFEval,
ARC-C, Winogrande — results pending final collection).
**Acceptance rate** (mt_bench, draft_length=3, output_length=4096,
temperature=0):
| detach_prob | vLLM AR | TRT-LLM AR |
|---|---|---|
| baseline (no LoRA) | 2.14 | 2.15 |
| 0.5 | 1.45 | 1.44 |
| 0.8 | **3.06** | **3.01** |
| 0.85 | 2.90 | 2.90 |
| 0.9 | 2.76 | 2.77 |
| 0.95 | 2.51 | 2.58 |
| 0.99 | 2.37 | 2.37 |
| 0.999 | 2.30 | 2.27 |
| 0.9999 | 2.31 | 2.26 |
Best AR at `detach_prob=0.8`: ~40% improvement over baseline.
### Testing
`tests/unit/torch/speculative/plugins/test_hf_speculative_lora.py` (5
tests):
- `test_lora_layers_injected` — LoRA layers present after conversion
- `test_trainable_params` — only `lora_*` and `eagle_module` params are
trainable
- `test_forward_returns_loss` — forward returns non-zero scalar loss
- `test_eagle_offline_incompatible` — `eagle_base_lora=True` +
`eagle_offline=True` raises `ValueError`
- `test_export_lora_artifacts` — export produces standard peft adapter
files
### Bug fixes (included in this PR)
1. **`launch_train.sh` case pattern ordering**: glob
`--eagle_base_lora*` was before specific patterns
(`--eagle_base_lora_rank*`, etc.), silently swallowing LoRA args.
2. **LoRA optimizer exclusion during warmup**: warmup freezing excluded
LoRA from the optimizer entirely; fixed with `add_param_group` in the
callback.
3. **`merge_lora.py` config.json**: `save_pretrained()` with
transformers >=5.x rewrites `rope_theta` → `rope_parameters`, breaking
vLLM positional embeddings. Fixed by copying the original base model
config.
4. **Multinode `dp_shard_size`**: used local GPU count instead of
`WORLD_SIZE`.
### Checklist
- [x] Backward compatible (all new config fields have defaults)
- [x] Uses `peft` via lazy imports (no hard dependency)
- [x] Unit tests added
- [x] Online HF training only (`eagle_offline=True` blocked)
---------
Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
||
|
|
361f7e391b |
Merge puzzletron compression algorithm (#1121)
### What does this PR do? Implement puzzletron compression algorithm based on Puzzle paper (https://arxiv.org/abs/2411.19146) <details> <summary> Th list of reviewed and merged MRs that resulted in the feature/puzzletron branch</summary> Merging dkorzekwa/any_model to feature/puzzletron [Add anymodel directories to feature/puzzletron by danielkorzekwa · Pull Request #974 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/974) - merged [Draft: anymodel activation scoring by danielkorzekwa · Pull Request #989 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/989) - merged [Draft: Merge anymodel pruning by danielkorzekwa · Pull Request #990 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/990/) - merged [Draft: Merging anymodel:build_library_and_stats by danielkorzekwa · Pull Request #993 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/993) - merged [Dkorzekwa/any model calc one block scores by danielkorzekwa · Pull Request #994 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/994) - merged [Draft: merge any_model: mip_and_realize_models by danielkorzekwa · Pull Request #995 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/995) - merged [Dkorzekwa/any model other modeqls by danielkorztiekwa · Pull Request #1007 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1007/) - merged PR to 1007: https://github.com/NVIDIA/Model-Optimizer/pull/1039 - merged [Dkorzekwa/anymodel gptoss by danielkorzekwa · Pull Request #1020 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1020) - merged [Merge any_model tutorial by danielkorzekwa · Pull Request #1035 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1035) - merged [Merge mbridge distillation for any_model by danielkorzekwa · Pull Request #1036 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1036) - merged [MR branch for the remaining difference between dkorzekwa/any_model an… by danielkorzekwa · Pull Request #1047 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1047) - merged [Dkorzekwa/decilm hf code cleanup by danielkorzekwa · Pull Request #1071 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1071) - merged [Dkorzekwa/decilm hf code cleanup 2 by danielkorzekwa · Pull Request #1073 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1073) - merged [Dkorzekwa/anymodel subblock stats by danielkorzekwa · Pull Request #1085 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1085) - merged [Dkorzekwa/anymodel subblock stats nodecilm by danielkorzekwa · Pull Request #1102 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1102) - merged [Dkorzekwa/decilm cleanup post subblockstats by danielkorzekwa · Pull Request #1103 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1103) - merged [code clean up by danielkorzekwa · Pull Request #1110 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1110) - merged Merging into main: [Activation hooks redesign (reuse hooks component across both minitron and puzzletron) by danielkorzekwa · Pull Request #1022 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1022) - merged [Dkorzekwa/puzzletron use importance hooks from prune by danielkorzekwa · Pull Request #1115 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1115) - merged </details> <!-- Details about the change. --> ### Usage Puzzletron tutorial: https://github.com/NVIDIA/Model-Optimizer/tree/feature/puzzletron/examples/puzzletron ### Testing The main e2e test for compressing 9 models with Puzzletron: https://github.com/NVIDIA/Model-Optimizer/blob/feature/puzzletron/tests/gpu/torch/puzzletron/test_puzzletron.py 2-gpu nightly tests: - https://github.com/NVIDIA/Model-Optimizer/actions/runs/24468209205/job/71501061203 - https://github.com/NVIDIA/Model-Optimizer/actions/runs/24470214159/job/71508152952 ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Puzzletron: end-to-end heterogeneous pruning & NAS workflow with AnyModel support, example pipelines, deployment and evaluation utilities, and tools for converting/pruning and exporting compressed checkpoints. * **Documentation** * Comprehensive Puzzletron tutorials, model-specific guides, evaluator instructions, example configs, and changelog entry. * **Chores** * CI/workflow updates (extras installation, longer GPU test timeout), pre-commit hook exclusion updated, and CODEOWNERS entries added. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com> Signed-off-by: Liana Mikaelyan <lmikaelyan@nvidia.com> Signed-off-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com> Signed-off-by: Daniel Korzekwa <daniel.korzekwa@gmail.com> Signed-off-by: jrausch <jrausch@nvidia.com> Signed-off-by: root <root@pool0-00848.cm.cluster> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Liana Mikaelyan <lmikaelyan@nvidia.com> Co-authored-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com> Co-authored-by: J Rausch <38429553+j-rausch@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
dec2952992 |
[6034518] Downgrade TRT support for remote autotuning in Autotune from 10.16 to 10.15 (#1259)
### What does this PR do? Type of change: Bug fix Remote autotuning is supported in TensorRT from version 10.15, but fails with Autotune as it's checking for 10.16+. This PR fixes that check and updates documentation accordingly. ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing See bug 6034518. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!--- Mandatory --> - Did you write any new necessary tests?: N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added a Remote Autotuning guide for TensorRT 10.15+ with CLI examples; updated examples to require `--safe --skipInference`. * **Updates** * Lowered TensorRT minimum requirement for remote autotuning from 10.16 to 10.15. * Clarified CLI help text for trtexec/autotune arguments. * **Bug Fixes** * trtexec-based autotuning now verifies the trtexec executable version when checking compatibility. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com> Signed-off-by: dmoodie <dmoodie@nvidia.com> Co-authored-by: dmoodie <dmoodie@nvidia.com> |
||
|
|
1619421383 |
Added support for MoE for vllm >= 0.14.0rc1 (#1162)
### What does this PR do?
Type of change: Bug fix
`_QuantFusedMoEBase.forward()` previously replaced
`vllm_fused_moe_package.invoke_fused_moe_kernel`, which was replaced
starting in vLLM v0.14.0rc1,
There are two paths for FusedMoE forward:
```
Path 1 (Modular — standard CUDA path):
FusedMoE.forward()
→ self.runner.forward()
→ TritonExperts.apply()
→ invoke_fused_moe_triton_kernel() ← called twice (w1, w2)
Path 2 (legacy):
inplace_fused_experts / outplace_fused_experts
→ fused_experts_impl()
→ dispatch_fused_moe_kernel()
→ invoke_fused_moe_triton_kernel()
or invoke_fused_moe_wna16_triton_kernel()
or invoke_fused_moe_wna16_cuda_kernel()
```
This caused an `AttributeError` / assertion failure for any MoE model
quantized with vLLM ≥ v0.14.0rc1.
The fix refactors the kernel-patching logic into a `_patch_moe_kernel()`
context manager that probes for both attribute names (the two names are
mutually exclusive across vLLM versions — confirmed by inspecting every
release from v0.10.0 to v0.19.1).
### Usage
NA
### Testing
```
docker run --gpus all -it --shm-size=160GB --network host --rm -v <modelopt path>:/home/modelopt \
vllm/vllm-openai:v0.15.0 bash -c "cd /home/modelopt && pip install . && pip install datasets && \
QUANT_CFG=NVFP4_DEFAULT_CFG python3 /home/modelopt/examples/vllm_serve/vllm_serve_fakequant.py \
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 -tp 1 --served-model-name NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 \
--host 0.0.0.0 --port 8001 --trust-remote-code --disable-custom-all-reduce \
--gpu-memory-utilization 0.8"
```
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Refactor**
* Ensures quantized expert weights are correctly used by the fused-MoE
execution path so inference uses the intended quantized tensors.
* Replaces fragile manual swapping of the runtime kernel with a safer,
context-managed swap that reliably caches and restores the original.
* Adds runtime detection and selection among available fused-MoE kernel
entrypoints to support multiple variants.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
|
||
|
|
3131195241 |
add: DFlash block diffusion speculative decoding (#1211)
DFlash (Block Diffusion for Flash Speculative Decoding) predicts an entire block of tokens in a single forward pass using masked parallel prediction with KV injection from the target model's hidden states. Key features: - Feature fusion (multi-layer hidden states -> FC + RMSNorm) - KV injection (fused features as K/V in every draft layer with QK-norm) - Random anchor sampling with bidirectional intra-block attention - Logit distillation with exponential loss decay (gamma weighting) - Multi-node DDP training with checkpoint resume - Export to z-lab compatible HF format - Online validation (context-dependent ground truth) Training recipe: modelopt_recipes/general/speculative_decoding/dflash.yaml Results: examples/speculative_decoding/doc/dflash_results.md ### ModelOpt Eval (online validation, osl=512) | Dataset | z-lab | ModelOpt (306K) | Diff | |---------|-------|-----------------|------| | gsm8k | 4.10 | **5.19** | **+1.09** | | MT-Bench | 3.58 | **4.36** | **+0.78** | ### z-lab Official Eval (dflash.benchmark, osl=512) | Dataset | z-lab | ModelOpt (306K) | Diff | |---------|-------|-----------------|------| | gsm8k | **5.00** | 4.08 | -0.92 | | MT-Bench | **3.28** | 2.99 | -0.29 | > z-lab model trained with block_size=16. ModelOpt trained with block_size=8. ## Evaluation Method Impact (gsm8k) | Eval Method | z-lab checkpoint | ModelOpt (306K) | |-------------|-----------------|-----------------| | Fixed GT (ModelOpt eval) | 2.95 | 4.23 | | Online GT (ModelOpt eval) | 4.10 | **5.19** | | z-lab official eval | **5.00** | 4.08 | ### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added DFlash speculative decoding mode with parallel block prediction support. * Included training launchers and MT-Bench evaluation scripts for DFlash models. * Added online acceptance rate validation for improved inference verification. * **Documentation** * DFlash quick start guide with configuration parameters and training examples. * Performance results and benchmarks for DFlash-trained models. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
73be81037d |
vLLM fakequant: add recipe-based quantization support (#1233)
### What does this PR do? Type of change: example update This PR adds recipe-based quantization support to the vLLM fakequant example. ### Testing ``` docker run --gpus all -it --shm-size=160GB --network host --rm --entrypoint bash -v <modelopt>:/home/modelopt vllm/vllm-openai:v0.15.0 -c "cd /home/modelopt && pip install . && pip install datasets && RECIPE_PATH=/home/modelopt/modelopt_recipes/general/ptq/nvfp4_mlp_only-fp8_kv.yml python3 /home/modelopt/examples/vllm_serve/vllm_serve_fakequant.py Qwen/Qwen3-0.6B -tp 1 --served-model-name Qwen3-0.6B --host 0.0.0.0 --port 8001 --trust-remote-code --disable-custom-all-reduce --gpu-memory-utilization 0.8" ``` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added `RECIPE_PATH` environment variable support enabling users to specify ModelOpt PTQ recipe YAML files for quantization configuration in vLLM serving. * **Documentation** * Updated examples and documentation to support recipe-driven quantization configuration, aligning export workflow with recipe-based setup. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> |
||
|
|
b6c6ec342c |
use typed quantize config instead of a raw dict (#1249)
### What does this PR do? But fix: Use typed QuantizeConfig instead using raw dict for formal typed ModelOpt configs. The dict typing was accidental. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Quantization recipe configuration is now implemented with a strongly-typed, structured schema that enforces type safety and provides enhanced validation with comprehensive error detection capabilities. * **Tests** * Updated recipe loading tests to correctly validate quantization configurations when recipes are loaded from directories, fully supporting the new structured object-based configuration format. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
202c3d3894 |
Add SwinTransformer support for torch_onnx quantization workflow (#1235)
## Summary - Enable end-to-end quantize → ONNX export → TRT engine pipeline for SwinTransformer models (v1 and v2) across FP8, INT8, MXFP8, NVFP4, and auto precision modes - Add Conv2d quantization overrides for TRT compatibility (TRT only supports FP8/INT8 for convolutions) - Fix FP8 LayerNorm type mismatch in TRT stronglyTyped mode by adding `LayerNormalization` to `change_casts_to_fp16` - Fix `cast_initializer_to_dtype` crash when node has no initializer inputs - Simplify `download_example_onnx.py` to a single `--timm_model_name` (required) flag, removing redundant `--vit` and `--llama` flags - Add vision model support matrix to README (ViT, Swin, SwinV2) - Rewrite tests: parametrize over (ViT, Swin, SwinV2) × (fp8, int8, mxfp8, nvfp4, auto) with TRT engine build verification ## Test plan - [ ] `python -m pytest tests/examples/torch_onnx/test_torch_quant_to_onnx.py -v` — 15 tests (3 models × 5 modes), all pass - [ ] Verified Swin accuracy on ImageNet-1k across all precisions (FP8: 81.29%, INT8: 81.12%, MXFP8: 81.32%, NVFP4: 80.79%, Auto: 80.84% TRT top-1 vs 81.37% base) - [ ] INT4_AWQ deferred (TODO in test file) — requires INT4 exporter changes for non-MatMul/Gemm consumer patterns 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * ONNX export supports arbitrary timm vision models with auto device selection and new CLI options (--timm_model_name, --model_kwargs, --no_pretrained); batch-size/input sizing is now model-generic. * **Bug Fixes** * Expanded FP16/BF16 cast handling to additional ONNX ops. * Disabled inplace ReLU before auto-quantization to avoid incorrect transforms. * Conv2d quantization overrides added for improved TensorRT compatibility. * Safer handling when initializers are missing during dtype casting. * **Documentation** * README updated with supported models table, quantization mappings, and example CLI usage. * **Tests** * Tests expanded to multiple architectures/quant modes and now verify TensorRT engine build. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
6403389eb0 |
Feat: Configurable Eagle ROPE scaling during export (#1238)
### What does this PR do? JIRA ticket: https://jirasw.nvidia.com/browse/OMNIML-3469 Type of change: New feature Decouple EAGLE training rope configuration from export rope configuration, enabling separate YaRN rope scaling injection at export time for long-context inference. #### Changes **Configurable export rope scaling (`EagleConfig`)** - Add `eagle_export_rope_scaling` field to `EagleConfig` with default YaRN config (`factor=32.0`, `original_max_position_embeddings=2048`) - Set to `{}` to disable rope scaling injection at export **Simplified training defaults (`default_config.py`)** - Change default training rope from `llama3` (theta=500k) to `default` (theta=10k) — models now train with simple positional embeddings; rope scaling is applied only at export - Add `rope_theta` inside `rope_scaling` dict for transformers 5.x cross-version compatibility **Move config validation/rewriting into `EagleConfig` (`config.py`)** - `_derive_eagle_offline`: derives `eagle_offline` from `data_args.offline_data_path` via validation context, removing manual assignment in `main.py` - `_check_rope_scaling_consistency`: rejects configs where `eagle_export_rope_scaling` is set but training `rope_type` is not `"default"` - `_warn_rope_vs_training_seq_len`: warns when `original_max_position_embeddings` differs from `training_seq_len` **Export rope injection (`hf_spec_export.py`)** - Inject `eagle_export_rope_scaling` into the exported HF config when training rope_type is `"default"` - Fall back `rope_theta` from `rope_scaling` dict for transformers 5.x compatibility **Fix Megatron RotaryEmbedding crash (`megatron_eagle.py`)** - `dict_to_config()` set `rope_scaling=True` whenever the `rope_scaling` key existed, even without a `"factor"` — causing `RotaryEmbedding` to divide by `None` - Now only enables `rope_scaling` when the dict actually contains a `"factor"` key ### Usage Configure in YAML config (or use defaults from `eagle3.yaml`): ```yaml eagle: eagle_export_rope_scaling: rope_type: yarn factor: 32.0 original_max_position_embeddings: 2048 ``` Set to empty dict to disable export rope injection: ```yaml eagle: eagle_export_rope_scaling: {} ``` ### Testing - New unit tests: `tests/unit/torch/speculative/test_eagle_config.py` — rope consistency validator, seq_len warning, context-derived `eagle_offline` - New unit tests: `tests/unit/torch/export/test_hf_spec_rope_export.py` — export rope injection, fallback, and empty-config cases ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ (new field has sensible default; existing configs work unchanged) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ (should be added if merging as a feature) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Add export-time rope-scaling configuration for EAGLE models. * **Improvements** * Stronger validation and context-aware reconciliation between training and export configs. * Export now injects rope-scaling and rope-theta when appropriate. * Default rope-scaling values updated for EAGLE variants. * Model instances now expose export rope-scaling for downstream use. * **Tests** * Added unit tests covering rope-scaling export behavior and configuration validators. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
0b42c143dd |
Update LICENSE and SPDX-License-Identifier as per OSRB guidance (#1244)
Update LICENSE and SPDX-License-Identifier as per OSRB guidance <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated contribution guidelines with expanded license compliance instructions and SPDX identifier guidance. * Extended LICENSE file with new "Third-Party Software Notices" section documenting Apache 2.0, MIT, and BSD 3-Clause licensed components. * **Chores** * Updated SPDX license identifiers across multiple files to reflect dual and triple licensing (Apache 2.0 with MIT and/or BSD 3-Clause). <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
80a77d1cc0 |
Add LTX-2 third-party license notices for legal compliance (#1226)
## Summary LTX-2 (`ltx-core`, `ltx-pipelines`, `ltx-trainer`) is a third-party dependency developed and provided by Lightricks. It is governed by the [LTX Community License Agreement](https://github.com/Lightricks/LTX-2/blob/main/LICENSE), **not** the Apache 2.0 license that covers NVIDIA Model Optimizer. Per legal guidance, all integration points must clearly surface this to users. - Add `[!WARNING]` license notice blocks at the top of all LTX-2-related READMEs (`examples/diffusers`, `examples/diffusers/distillation`, `examples/windows/diffusers/qad_example`) - Add `warnings.warn(UserWarning)` at every LTX package import site in Python files, covering both top-level and lazy imports: - `examples/diffusers/distillation/distillation_trainer.py` - `examples/diffusers/quantization/calibration.py` - `examples/diffusers/quantization/pipeline_manager.py` - `examples/windows/diffusers/qad_example/sample_example_qad_diffusers.py` - `modelopt/torch/export/diffusers_utils.py` - `modelopt/torch/quantization/plugins/diffusion/ltx2.py` - Add license notice comment to `requirements.txt` files that list LTX packages, so the obligation is visible at install time - Update `.github/CODEOWNERS` so all `requirements*.txt` files (covering variants like `requirements-dev.txt`) are owned by `@NVIDIA/modelopt-setup-codeowners` regardless of location, via a last-match-wins rule **Design notes:** - For library files (`diffusers_utils.py`, `ltx2.py`), the warning is placed at the lazy import site inside functions — it fires only when LTX-2 code paths are actually invoked, not at module import time, to avoid polluting non-LTX users - For example entry-point scripts that are LTX-2-only, the warning fires at module load time (after all imports, to satisfy ruff E402) ## Test plan - [ ] Confirm `pre-commit run --all-files` passes (ruff, mypy, markdownlint, bandit all clean) - [ ] Verify warning appears at runtime when running an LTX-2 quantization or distillation example - [ ] Confirm non-LTX code paths (FLUX, SDXL, SD3) do not emit the warning 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added third-party license notices across documentation and requirements files clarifying LTX-2 packages are governed by the LTX Community License Agreement rather than NVIDIA Model Optimizer's Apache 2.0 license. * **Chores** * Updated code ownership configuration for requirements files. * Added runtime warnings to notify when LTX-2 dependencies are accessed. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |