522 Commits
Author SHA1 Message Date
Keval Morabiaandclaude[bot] 2ce745a92e Deprecate gradnas pruning and bert example (#1427)
### What does this PR do?

Type of change: Deprecation of dead code <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->

Deprecation warning already added in 0.44 as per 1-release deprecation
policy

GradNAS only works for Bert and GPT-J and we dont actively maintain it
or test it. Keeping it creates an expectation that it works plus it adds
one more option for user to choose from. We already have much better
pruning algorithms (Minitron and Puzzletron) for LLM pruning already
hence removing GradNas.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ No but we dont have any users
of this feature either <!--- If ❌, explain why. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Deprecation**
  * GradNAS pruning algorithm deprecated; related examples removed.
* **Documentation**
* Pruning and NAS guides and changelog updated to focus on Minitron and
FastNAS; GradNAS references removed.
* **Chores**
  * Chained-optimizations example and scripts removed.
* Ownership mappings updated for README and examples; license insertion
now applies to a previously excluded example file.
* **Tests**
* Multiple unit tests and test utilities related to GradNAS/transformer
NAS removed.

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1427)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
2026-05-12 22:32:22 +05:30
Keval Morabia d30ebbd455 fix(llm_eval): migrate lm_eval_hf.py to lm-eval >= 0.4.10 HarnessCLI (#1416)
### What does this PR do?

Type of change: Bug fix

Related: NVbug 6153721. To make `lm_eval_hf.py` and the `vllm_causallms`
adapter compatible with vLLM 0.20+, **lm_eval must be upgraded to
0.4.11**.

`examples/llm_eval/lm_eval_hf.py` failed to import on lm-eval >= 0.4.10:

```
ImportError: cannot import name 'parse_eval_args' from 'lm_eval.__main__'
```

lm-eval 0.4.10 replaced `lm_eval.__main__.{setup_parser,
parse_eval_args, cli_evaluate}` with a `HarnessCLI`-based interface in
`lm_eval._cli`. This PR drops the legacy code path and drives
`HarnessCLI` directly:

- Attach the ModelOpt arguments (`--quant_cfg`, `--auto_quantize_*`,
`--calib_*`, `--compress`, `--sparse_cfg`) to the new `run` subparser.
- After parsing, move those keys out of the argparse namespace and into
`args.model_args` (now a dict, courtesy of `MergeDictAction`), so
`EvaluatorConfig.from_cli` doesn't reject them as unknown kwargs and so
they reach our `HFLM.create_from_arg_obj` override.
- Hard-require lm-eval >= 0.4.10 via `packaging.version.Version` and
bump `examples/llm_eval/requirements.txt` and
`examples/puzzletron/requirements.txt` accordingly.

### Usage

```bash
# Same CLI as before — HarnessCLI auto-inserts the `run` subcommand for legacy-style invocations.
python examples/llm_eval/lm_eval_hf.py \
    --model hf \
    --model_args pretrained=<HF model> \
    --tasks hellaswag \
    --quant_cfg FP8_DEFAULT_CFG \
    --batch_size 4
```

### Testing

Added `tests/examples/llm_eval/test_llm_eval.py::test_lm_eval_hf` — an
end-to-end test that builds a tiny qwen3 (no GPU required) and runs
`lm_eval_hf.py` against MMLU with `--limit 0.1`. Verifies both the new
HarnessCLI integration and the `HFLM.create_from_arg_obj` override
actually execute. Existing FP8 PTQ test (`test_qwen3_eval_fp8`) was
retargeted to the same tiny qwen3 helper so it no longer pulls down the
1.1B TinyLlama checkpoint.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ — drops support for lm-eval <
0.4.10. Requirements pins are bumped in the same PR; existing CLI
invocations continue to work because HarnessCLI auto-inserts the `run`
subcommand for legacy-style argv.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌

### Additional Information

NVbug 6153721 — vLLM 0.20+ compatibility requires lm_eval 0.4.11. This
PR pins to >= 0.4.10 to fix the import error; bumping the floor to
0.4.11 (or letting users opt in via vLLM extras) unblocks vLLM 0.20+
usage.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Require lm_eval >= 0.4.10 for evaluation examples and remove an
lm-eval pin from a puzzletron example.
* Improve LM-eval CLI integration so model-related CLI options
(including trust_remote_code) are recognized and forwarded correctly.

* **Tests**
  * Added an end-to-end smoke test for the LM evaluation example.
* Updated evaluation tests to use the new tiny model setup for more
robust validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-11 16:44:25 +05:30
Chenjie Luo 6a3b6b8329 [Recipes][LLM PTQ] Add nvfp4 MSE+FP8-cast-KV recipes (experts_only / mlp_only) + --recipe in example scripts (#1407)
## Summary

- Adds two PTQ recipes that combine **experts/MLP-only NVFP4 W4A4** with
**MSE FP8 scale-sweep weight calibration** and **FP8 KV cache with
`use_constant_amax: true`** (skips KV calibration; matches the
`nvfp4_default-fp8_cast_kv` contract):
- `modelopt_recipes/general/ptq/nvfp4_experts_only_mse-fp8_cast_kv.yaml`
— applies to `*mlp.experts*` / `*block_sparse_moe*` only.
- `modelopt_recipes/general/ptq/nvfp4_mlp_only_mse-fp8_cast_kv.yaml` —
applies to all `*mlp*` / `*block_sparse_moe*` (dense MLP + MoE).
- Threads a new `--recipe` flag through
`examples/llm_ptq/scripts/parser.sh` and `huggingface_example.sh`.
Either `--quant` or `--recipe` is required; passing **both errors out**.
Recipe names are not validated in the script — `hf_ptq.py` is the source
of truth.
- Drops the bash-side `qformat` whitelist case-statement in
`huggingface_example.sh` for the same reason.

## Files

**New recipes (`modelopt_recipes/general/ptq/`):**
- `nvfp4_experts_only_mse-fp8_cast_kv.yaml` — same patterns as
`nvfp4_experts_only-fp8_kv.yaml`.
- `nvfp4_mlp_only_mse-fp8_cast_kv.yaml` — same patterns as
`nvfp4_mlp_only-fp8_kv.yaml`.

Both differ from their `_kv` siblings by:
- `algorithm: max` → `{ method: mse, fp8_scale_sweep: true, layerwise:
false }`
- All targeted **weight quantizers** switch `type: dynamic` → `type:
static` (otherwise `mse_calibrate` skips them: only static block-quant
weight quantizers are recognized for the FP8 sweep — see
`model_calib.py:369-374`).
- Input quantizers stay dynamic.
- KV bmm adds `use_constant_amax: true` (the `_cast_kv` flavor).

**Scripts (`examples/llm_ptq/scripts/`):**
- `parser.sh` — adds `--recipe` long-option, default `RECIPE=""`,
validates one-of-{`--quant`, `--recipe`} and not-both.
- `huggingface_example.sh` — when `RECIPE` is set, derives `MODEL_NAME`
from the recipe basename, passes `--recipe=…` to `hf_ptq.py` instead of
`--qformat=…`, and exits after export with a TRT-LLM deployment hint
(recipes can produce arbitrary configs that the script's downstream
`run_tensorrt_llm.py` path doesn't know how to handle generically).
Drops the `qformat` whitelist; defers to `hf_ptq.py`.

## Behavior

```
# Errors with: "Cannot specify both --quant and --recipe; pick one."
bash huggingface_example.sh --model=... --quant=nvfp4 --recipe=... --tasks=quant

# Errors with usage if neither is given
bash huggingface_example.sh --model=... --tasks=quant

# Both of these are now accepted; --recipe is forwarded verbatim to hf_ptq.py
bash huggingface_example.sh --model=... --quant=nvfp4 --tasks=quant
bash huggingface_example.sh --model=... --recipe=general/ptq/nvfp4_experts_only_mse-fp8_cast_kv --tasks=quant
bash huggingface_example.sh --model=... --recipe=general/ptq/nvfp4_mlp_only_mse-fp8_cast_kv  --tasks=quant
```

## Test plan

- [x] `experts_only_mse-fp8_cast_kv` loads via
`modelopt.recipe.load_recipe(...)` and produces the expected algorithm +
per-pattern `quant_cfg` (verified in a working env: `algorithm ==
{'method': 'mse', 'fp8_scale_sweep': True, 'layerwise': False}`; expert
weight quantizers `type: static`; KV bmm has `use_constant_amax: True`).
- [x] Parser sanity: 4 flag combinations (both, neither, only `--quant`,
only `--recipe`) all behave as designed.

## Note

Pre-commit hook `check-modelopt-recipes` was skipped on both commits
because the local conda env has a broken `torchvision` install
(`AttributeError: partially initialized module 'torchvision' has no
attribute 'extension'`) that prevents `from modelopt.recipe.loader
import load_recipe`. The `experts_only` recipe was validated
independently by running `tools/precommit/check_modelopt_recipes.py` in
a working environment (exits 0); the `mlp_only` one is the same shape
with a different glob.

Rebased onto `main` from #1391 (which targeted
`chenjiel/nvfp4-fp8-sweep-triton`). The diff is scoped to the recipes +
script wiring; no kernel/sweep changes are included here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added recipe-based quantization as an alternative to format-based
quantization with a new `--recipe` CLI option.
* Added two new quantization recipes for targeted layer optimization:
one for expert-layer-only quantization and one for MLP-layer-only
quantization, both featuring NVFP4 and FP8 KV-cache optimization.

* **Configuration**
* `--quant` and `--recipe` options are now mutually exclusive; specify
one to configure quantization behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-05-07 15:38:47 -07:00
achidiac-nv fa05381b86 Add demo (Puzzletron vs Minitron guide) in examples/pruning/ with README and notebooks (#1320)
### What does this PR do?

Type of change: new documentation/example (tutorial + notebooks)

Adds an end-to-end pruning & distillation guide under
`examples/pruning_demo/`, walking users through structural compression
of Qwen3-8B with NVIDIA Model-Optimizer.

The example compares two methods side-by-side on two concrete scenarios:

- **Scenario 1 — Moderate compression (7B parameter target)**:
homogeneous pruning with Minitron vs. heterogeneous NAS-based pruning
with Puzzletron.
- **Scenario 2 — Aggressive compression (78,000 MiB memory budget)**:
same comparison under a hard memory constraint.

Both scenarios are followed by knowledge distillation and evaluated on
MMLU (end-to-end in the notebooks) plus HellaSwag and GSM8K (reported in
the guide).

Contents:
- `README.md` — full guide (setup, two scenarios, head-to-head analysis,
inference benchmarks with vLLM + AIPerf, decision rules, limitations,
open questions).
- `00_prerequisites.ipynb` — data prep (WikiText-103 → Megatron binary)
and teacher baseline evaluation.
- `scenario1_minitron.ipynb` / `scenario1_puzzletron.ipynb` — 7B-param
target.
- `scenario2_minitron.ipynb` / `scenario2_puzzletron.ipynb` — 78k-MiB
target, including a Puzzletron memory-sweep bonus section.
- `advanced_compression_experiments.md` — extended results (larger
distillation budgets with Nemotron-Post-Training-Dataset-v2, BLD,
chained Minitron→Puzzletron, Mamba-Transformer hybrid).
- Companion plots (`summary_chart.png`, `distillation_curves.png`,
`memory_sweep_combined.png`, `all_curves_throughput_vs_latency.png`,
...).

### Usage

Follow setup instructions in README.md then run, in order:
1. 00_prerequisites.ipynb — prepare data + baseline eval (~15 min).
2. One (or more) of the scenario notebooks:
  - scenario1_minitron.ipynb (~1h45)
  - scenario1_puzzletron.ipynb (~6h first run)
  - scenario2_minitron.ipynb (~45 min)
  - scenario2_puzzletron.ipynb (~6h15 first run)

### Testing
- All four scenario notebooks were executed and tested end-to-end on 2x
H200 GPUs
- Inference benchmarks were captured on 1x H200 NVL with vLLM (AnyModel
backend for Puzzletron checkpoints) and AIPerf
- No library code is modified, so no unit tests are affected

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ (documentation/examples-only
addition)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
runtime dependencies; the notebooks use lm-eval==0.4.8 and the existing
ModelOpt/NeMo stack. The vLLM serving appendix references an open PR
(vllm-project/vllm#36512) for Puzzletron AnyModel support, clearly
flagged as pre-release.
- Did you write any new necessary tests?: N/A — tutorial / documentation
example.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information
- Base model: https://huggingface.co/Qwen/Qwen3-8B
- Calibration dataset: nvidia/Nemotron-Post-Training-Dataset-v2
- Distillation dataset: WikiText-103
- Complements the existing examples/puzzletron/ and
examples/megatron_bridge/ READMEs with a scenario-driven narrative and a
direct Minitron↔Puzzletron comparison.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added comprehensive guides comparing Minitron (homogeneous pruning)
and Puzzletron (heterogeneous NAS + MIP) for LLM compression.
* Added step-by-step Jupyter notebooks demonstrating two end-to-end
scenarios (prune → distill → evaluate) with expected MMLU baselines and
memory budgeting.
* Added advanced experiments doc with extended results, chaining
strategies, benchmarking (including inference/vLLM notes), tips,
limitations, and appendices.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Alexandre Chidiac <achidiac@nvidia.com>
2026-05-06 15:54:13 +02:00
Shengliang Xu f34f488a83 Add a general composable $import system for YAML configs, and use it to implement composable recipes (#1253)
### What does this PR do?

Type of change: New feature

Adds a general composable YAML config loading layer for ModelOpt configs
and recipes. YAML remains the source of truth for configuration data,
while Python/Pydantic-compatible types provide schema validation at load
time. This PR uses that loader to de-duplicate PTQ recipes, introduce
reusable config snippets/presets, and start migrating selected hardcoded
quantization presets to YAML.

#### Problem

1. Built-in PTQ recipes duplicated numeric format definitions, KV-cache
entries, and the default quantizer exclusion list.
2. YAML snippets were reusable only by convention; they did not declare
or validate the schema they were meant to satisfy.
3. Loading YAML-backed quantization presets from
`modelopt.torch.quantization.config` could not depend on
`modelopt.recipe` without creating circular imports.
4. Directory-format recipes exposed import-resolution details in the
recipe loader and used `recipe.yaml` plus nested `metadata:` in a way
that made metadata handling inconsistent.

#### Solution

**Shared YAML config loader**

- Adds `modelopt.torch.opt.config_loader` as the low-level loader used
by both `modelopt.recipe` and `modelopt.torch.quantization.config`.
- Keeps the public `modelopt.recipe.load_config()` entry point, while
removing the private `modelopt/recipe/_config_loader.py` shim.
- Handles YAML loading, built-in/filesystem path resolution, suffix
probing, `ExMy` conversion for `num_bits` / `scale_bits`, `$import`
expansion, and schema validation.
- Lives below `modelopt.recipe` in the dependency graph to avoid
circular imports from quantization config code.

**Composable `$import` system**

Recipes and snippets can declare an `imports` mapping, then reference
entries with `{$import: name}`.

`$import` semantics:

- **Dict value**: replaced with the imported dict. Multiple imports are
supported with ordered precedence; inline keys override imported keys.
- **List entry**: schema-driven behavior for strongly typed lists. If
the snippet schema matches the containing list type, the imported list
is spliced. If the snippet schema matches the list element type, the
imported element is appended. Other schema combinations are rejected.
- **Multi-document YAML**: supports snippets that need an `imports`
header plus a list body.
- **Recursive and scoped**: snippets can import other snippets; import
names are scoped per file.
- **Cycle detection**: circular imports report a clear error.

**Snippet schema validation**

- Every reusable snippet referenced through `imports` must declare a `#
modelopt-schema: ...` preamble.
- Snippets are validated after nested imports are resolved.
- Schema paths are restricted to the `modelopt.` package and may be
Pydantic models, `TypedDict` classes, or explicitly typed container
aliases such as `list[QuantizerCfgEntry]`.
- Untyped list imports are rejected so list append/splice behavior stays
strongly typed.

**Recipe model and directory recipe cleanup**

- `ModelOptRecipeBase` now owns a `metadata: RecipeMetadataConfig`
field.
- `ModelOptPTQRecipe` is the PTQ recipe schema; the overlapping
YAML-specific PTQ config class was removed.
- Directory recipes now use `metadata.yaml` / `metadata.yml` for
top-level metadata fields, plus section files such as `quantize.yaml`.
- Directory recipe loading now delegates import resolution to
`load_config()` instead of manually using raw config loading.

**Config snippet and preset library**

Adds reusable snippets under `modelopt_recipes/configs/`:

- `numerics/`: `fp8`, `nvfp4`, `nvfp4_static`
- `ptq/units/`: `base_disable_all`, `default_disabled_quantizers`,
`w8a8_fp8_fp8`, `w4a4_nvfp4_nvfp4`, `kv_fp8`, `kv_fp8_cast`,
`kv_nvfp4_cast`
- `ptq/presets/`: YAML presets for `FP8_DEFAULT_CFG` and `FP8_KV_CFG`

`FP8_DEFAULT_CFG` and `FP8_KV_CFG` now load from YAML presets via
`load_config()`.

**Recipe migration and naming**

- General PTQ recipes now use shared imports instead of repeating the
same quantizer fragments inline.
- General PTQ recipe paths were renamed to KV-first naming, for example:
  - `general/ptq/fp8_default-fp8_kv` -> `general/ptq/fp8_default-kv_fp8`
- `general/ptq/fp8_default-fp8_cast_kv` ->
`general/ptq/fp8_default-kv_fp8_cast`
- `general/ptq/nvfp4_default-none_kv_gptq` ->
`general/ptq/nvfp4_default-kv_none-gptq`
- `general/ptq/nvfp4_default-nvfp4_cast_kv` ->
`general/ptq/nvfp4_default-kv_nvfp4_cast`
- Example docs and `examples/llm_ptq/hf_ptq.py --recipe` help text were
updated to use the new paths.

**Pre-commit and documentation**

- Recipe validation accepts `$import` entries and handles directory
recipes using `metadata.yaml`.
- The recipe validation hook skips `modelopt_recipes/configs/` because
those files are reusable snippets, not full recipes.
- `docs/source/guides/10_recipes.rst` now documents imports, schema
modelines, list append/splice semantics, built-in snippets, built-in
recipe paths, directory recipes, and the current recipe data model.

#### Backward compatibility

- Existing inline YAML recipes without `$import` continue to load.
- `modelopt.recipe.load_config()` remains public.
- The built-in recipe path renames are user-visible; callers should
update recipe path strings to the KV-first names listed above.

#### Testing

- `pytest tests/unit/recipe/test_loader.py -q` - 90 passed
- `python tools/precommit/check_modelopt_recipes.py ...` for the renamed
built-in PTQ recipes
- `pre-commit run mypy --files
modelopt/onnx/llm_export_utils/quantization_utils.py`
- `python -m py_compile examples/llm_ptq/hf_ptq.py`
- `git diff --check`

### Before your PR is "Ready for review"

- Is this change backward compatible?: Partially. Loader/API behavior is
compatible for existing inline YAML recipes, but built-in recipe path
names were renamed to KV-first paths.
- Did you write any new necessary tests?: Yes.
- Did you update Changelog?: Yes.

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-05-05 17:01:28 -07:00
milesial 7097a6910e [specdec_bench] Stratify --num_requests across categories (#1389)
When the dataset has a `category` column with >1 distinct categories and
``--num_requests N`` is below the dataset size, take ceil(N /
num_categories) rows from each category and round-robin interleave them
so any prefix is balanced. Falls back to the existing ``range(N)`` slice
when category metadata is absent or there's only one category.

Fixes a sampling bug where SPEED-Bench parquet files are sorted by
category, so e.g. ``--num_requests 64`` on throughput_8k pulls 64
high_entropy prompts and zero from low_entropy / mixed.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* SPEEDBench dataset truncation now performs deterministic, balanced
stratified sampling across categories (round-robin) when multiple
categories exist and a smaller sample size is requested.
* Falls back to simple prefix selection when no category or only one
category is present, preserving prior behavior in that case.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Alexandre Milesi <milesial@users.noreply.github.com>
2026-05-05 11:31:42 -07:00
Keval Morabia d794595bf8 llm_sparsity: Set warmup_steps 0 instead of 0.0 for transformers 5.x compat (#1393)
Fix for NVBug 6120631 to fix

```
finetune.py: error: argument --warmup_steps/--warmup-steps: invalid int value: '0.0'
```

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Corrected parameter format in finetuning example script for
consistency.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-05 10:06:28 -07:00
Chenjie Luo b5df2a5a3c Update llm_ptq requirements.txt (#1394)
### What does this PR do?

Type of change: Dependency update

compressed_tensors 0.15 is not compatible with our current quant
implementation for Kimi K2.5, K2.6

### Testing
Unittest

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Broadened the `compressed-tensors` dependency constraint to allow a
wider range of compatible versions.
  * Removed the `rouge_score` dependency from the example requirements.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2026-05-05 17:00:57 +00:00
yueshen2016andClaude Opus 4.6 ef326c8b9b Add Gemma4 MoE quantization support (#1219)
## Summary
- Register `Gemma4TextExperts` with `_QuantQwen35MoeExperts` plugin to
unfuse fused 3D expert tensors into per-expert `nn.Linear` layers for
quantization
- Add structural `is_moe()` detection for modules with `router` +
`experts` attributes (Gemma4 has no dedicated `SparseMoeBlock` class —
the decoder layer directly owns `router` and `experts`)
- Add `Gemma4TextDecoderLayer` to `get_expert_linear_names()` returning
`["gate_proj", "down_proj", "up_proj"]`
- Add `"*.experts.*"` pattern to `NVFP4_MLP_ONLY_CFG` and
`NVFP4_EXPERTS_ONLY_CFG` to match Gemma4's expert path
(`model.layers.X.experts.*`, not nested under `mlp`)

**Context:** Gemma4 MoE models (e.g. `google/gemma-4-26B-A4B-it`) store
expert weights as fused 3D `nn.Parameter` tensors (`gate_up_proj`,
`down_proj`) instead of `nn.ModuleList` of `nn.Linear`. Since ModelOpt's
quantizer only discovers `nn.Linear` modules, it silently skips the
expert weights — the bulk of the model remains unquantized.

**Companion vLLM PR:** https://github.com/vllm-project/vllm/pull/39406
(robust quantized MoE weight loading for Gemma4)

## Test plan
- [x] `hf_ptq.py --pyt_ckpt_path google/gemma-4-26B-A4B-it --qformat
nvfp4_mlp_only` — 35k+ quantizers inserted, 17GB output (vs 49GB BF16)
- [x] `vllm serve <path> --quantization modelopt` — loads and serves
successfully
- [x] Text generation: correct ("The capital of France is **Paris**.")
- [x] Vision: correct (describes image content accurately)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Support quantizing models with separate base/full components (handles
heads present only on the full model)
* Enhanced Mixture-of-Experts detection and explicit support for Gemma4
expert layer layouts
* Extended NVFP4 selective quantization presets and recipes to include
expert-layer patterns and enable FP8 for expert modules

* **Bug Fixes**
* Improved loss/logit handling and clearer errors for unsupported
quantization methods
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: James Shen <yueshen@nvidia.com>
Signed-off-by: Yue Shen <yueshen@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-05-06 00:09:54 +08:00
Keval Morabia f0eaa198df Enable active-param and memory based Minitron pruning constraint (#1377)
### What does this PR do?

Type of change: New feature, new tests, documentation.

OMNIML-4108: Extends the Minitron NAS pruner to support pruning by
**active parameter count** (`active_params`) and **memory footprint**
(`memory_mb`) in addition to the existing total parameter count
(`params`) constraint. Also adds standalone utilities for analytical
model stats.

#### Changes

**New pruning constraint keys**
- `active_params`: prune to a target number of active (routed) params —
useful for MoE models where total ≫ active; when present,
`active_params` is the **primary sort/display metric** for candidates
(priority: `active_params` > `params` > `memory_mb`)
- `memory_mb`: prune to fit a memory budget (BF16 weights + KV-cache +
Mamba state at a given sequence length and batch size)
- Constraints can be combined (AND logic): e.g. `{"params": 6e9,
"memory_mb": 12288}`

**New standalone utilities**
(`modelopt.torch.nas.plugins.megatron_model_stats`)
- `mcore_param_count`: analytically computes total and active parameter
counts for GPT and Mamba/hybrid MCore models
- `mcore_memory_footprint_mb`: estimates memory in MB (weights +
KV-cache + Mamba state)
- `print_mcore_model_stats`: rich-formatted model stats panel

**Rich-formatted pruning logs** — search space, top-k candidate tables,
and best subnet panel printed on rank 0

**`prune_score_func` format update** — now `mmlu_<N>pct_bs<bs>` (e.g.
`mmlu_10pct_bs32`) to explicitly control batch size for MMLU evaluation;
old `mmlu_<N>pct` format removed

**Infrastructure**
- NeMo container bumped to `nvcr.io/nvidia/nemo:26.04` in CI and docs
- Added `examples/megatron_bridge/requirements.txt` with
`transformers<5.0` (required for saving some Nemotron-3-Nano models)

### Usage

```python
# Prune to 3B active params (MoE-aware) — active_params is the primary sort metric
mtp.prune(model, mode=[("mcore_minitron", ss_config)], constraints={"active_params": 3e9}, config=pruning_config)

# Prune to fit a 12 GB memory budget
mtp.prune(model, mode=[("mcore_minitron", ss_config)], constraints={"memory_mb": 12288}, config=pruning_config)
```

### Testing

Pruned Nemotron-3-Nano-30B-A3B (31.6B, A3.6B) --> A3.0B. Takes <1hr on
8x H100 (more details in #1376)

```bash
torchrun --nproc_per_node 8 examples/megatron_bridge/prune_minitron.py \
    --pp_size 8 \
    --hf_model_name_or_path nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 \
    --trust_remote_code \
    --prune_target_params 28e9 \
    --prune_target_active_params 3e9 \
    --hparams_to_skip num_attention_heads \
    --seq_length 8192 \
    --output_hf_path pruned/Nemotron-3-Nano-30B-A3B-Pruned-28B-A3B-top20-max15depth-max30width-mmlu_10pct_bs32 \
    --top_k 20 \
    --max_depth_pruning 0.15 \
    --max_width_pruning 0.30 \
    --prune_score_func mmlu_10pct_bs32 \
    --num_layers_in_first_pipeline_stage 5 \
    --num_layers_in_last_pipeline_stage 5
```

```
╭──────────────────────────────────────────────────── Original Model Stats ─────────────────────────────────────────────────────╮
│ Total Parameters                              31.58B                                                                          │
│ Active Parameters                             3.58B                                                                           │
│ Memory (BF16, seq_length=8192, batch_size=1)  weights: 60230.1 MB, kv_cache: 48.0 MB, mamba_state: 23.8 MB, Total: 60301.9 MB │
╰───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

                                                                 Top 20 Candidates with Scores
┏━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━┓
┃  # ┃ export_config                                                                                                         ┃ active_params ┃ params ┃  score ┃
┡━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━┩
│  1 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 64, 'num_moe_experts': 120,          │         3.00B │ 27.06B │ 0.3399 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│  2 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 56, 'num_moe_experts': 112,          │         3.00B │ 25.37B │ 0.4650 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│  3 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 64, 'mamba_head_dim': 56, 'num_moe_experts': 112,          │         3.00B │ 25.37B │ 0.2343 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│  4 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 56, 'mamba_head_dim': 48, 'num_moe_experts': 96,           │         3.00B │ 20.09B │ 0.2552 │
│    │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│  5 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48, 'mamba_head_dim': 56, 'num_moe_experts': 104,          │         3.00B │ 21.61B │ 0.2601 │
│    │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│  6 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 64, 'num_moe_experts': 96,           │         3.00B │ 19.28B │ 0.3762 │
│    │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3712}                                             │               │        │        │
│  7 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 104,          │         3.00B │ 22.28B │ 0.4783 │
│    │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│  8 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 96,           │         3.00B │ 21.99B │ 0.2420 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3328}                                             │               │        │        │
│  9 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 112,          │         3.00B │ 25.37B │ 0.2399 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3712}                                             │               │        │        │
│ 10 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 112,          │         3.00B │ 26.17B │ 0.2601 │
│    │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3328}                                             │               │        │        │
│ 11 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 64, 'num_moe_experts': 112,          │         3.00B │ 25.37B │ 0.2503 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│ 12 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 56, 'num_moe_experts': 104,          │         3.00B │ 23.68B │ 0.4329 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│ 13 │ {'num_layers': 46, 'hidden_size': 2688, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 128,          │         3.00B │ 26.17B │ 0.2587 │
│    │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 2816}                                             │               │        │        │
│ 14 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 64, 'mamba_head_dim': 56, 'num_moe_experts': 104,          │         3.00B │ 23.68B │ 0.2336 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│ 15 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48, 'mamba_head_dim': 56, 'num_moe_experts': 96,           │         3.00B │ 20.09B │ 0.2559 │
│    │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│ 16 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 96,           │         3.00B │ 20.70B │ 0.4608 │
│    │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│ 17 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 104,          │         3.00B │ 23.68B │ 0.2455 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3712}                                             │               │        │        │
│ 18 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 104,          │         3.00B │ 24.42B │ 0.2503 │
│    │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3328}                                             │               │        │        │
│ 19 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 120,          │         3.00B │ 27.92B │ 0.2587 │
│    │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3712}                                             │               │        │        │
│ 20 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 64, 'num_moe_experts': 104,          │         3.00B │ 23.68B │ 0.2469 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
└────┴───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┴───────────────┴────────┴────────┘

╭──────────────────────────────────────────────────────────────────────── Best Subnet ─────────────────────────────────────────────────────────────────────────╮
│ export_config  {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 104, 'moe_ffn_hidden_size': 1856,     │
│                'moe_shared_expert_intermediate_size': 3072}                                                                                                  │
│ active_params  3.00B                                                                                                                                         │
│ params         22.28B                                                                                                                                        │
│ score          0.4783                                                                                                                                        │
╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

╭───────────────────────────────────────────────────── Pruned Model Stats ──────────────────────────────────────────────────────╮
│ Total Parameters                              22.28B                                                                          │
│ Active Parameters                             3.00B                                                                           │
│ Memory (BF16, seq_length=8192, batch_size=1)  weights: 42489.7 MB, kv_cache: 48.0 MB, mamba_state: 23.8 MB, Total: 42561.6 MB │
╰───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
```

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-05 09:58:47 +05:30
sugunav14 84fe91be5d Fix gpt-oss examples trl import error (#1390)
### What does this PR do?

Type of change: Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

<!-- Details about the change. -->
Cap kernels<0.13 and trackio<0.21 in examples/gpt-oss/requirements.txt.
Both newer versions require huggingface_hub>=1.x, but the example's
transformers pins huggingface_hub<1.0, so a fresh install breaks on
import (Unsupported type for field 'import_name': str | None from
kernels; cannot import name 'Volume' from trackio).

### Usage
No API change. On transformers<5.0, override the config's warmup_steps
with --warmup_ratio 0.03 --warmup_steps 0 (or edit the YAML), as already
noted by the comment in configs/sft_*.yaml.

```python
# Add a code snippet demonstrating how to use this
accelerate launch --config_file configs/zero3.yaml sft.py --config configs/sft_full.yaml --model_name_or_path openai/gpt-oss-20b --quant_cfg MXFP4_MLP_WEIGHT_ONLY_CFG --output_dir gpt-oss-20b-qat --warmup_steps 0 --warmup_ratio 0.03
```

### Testing
1. pip install -r examples/gpt-oss/requirements.txt
    pip install transformers==4.57.3

    ```python
    # Add a code snippet demonstrating how to use this
accelerate launch --config_file configs/zero3.yaml sft.py --config
configs/sft_full.yaml --model_name_or_path openai/gpt-oss-20b
--quant_cfg MXFP4_MLP_WEIGHT_ONLY_CFG --output_dir gpt-oss-20b-qat
--warmup_steps 0 --warmup_ratio 0.03
     ```

2. pip install -r examples/gpt-oss/requirements.txt
    pip install --upgrade transformers

    ```python
    # Add a code snippet demonstrating how to use this
accelerate launch --config_file configs/zero3.yaml sft.py --config
configs/sft_full.yaml --model_name_or_path openai/gpt-oss-20b
--quant_cfg MXFP4_MLP_WEIGHT_ONLY_CFG --output_dir gpt-oss-20b-qat
     ```

<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information
<!-- E.g. related issue. -->

Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com>
2026-05-05 09:51:30 +05:30
Chenjie Luo 1d21ab9e29 [DeepSeek] Default to top-k calibration with peer-max input amax sync (#1380)
## Summary

- DeepSeek PTQ (`examples/deepseek/ptq.py`) now defaults to native top-k
routing during MoE calibration. The previous all-tokens-to-all-experts
path (`CalibMoe`) is preserved behind a new `--calib_all_experts` flag.
- After `mtq.quantize`, `fixup_moe_expert_amax` syncs every expert's
`input_quantizer.amax` (w1/w2/w3) to the per-layer global peer max via
`dist.all_reduce(MAX)` across EP ranks. `weight_quantizer.amax` stays
per-expert; any uncalibrated expert is filled by computing amax over the
dequantized FP8 weight.
- `mtq.print_quant_summary` is now also written to
`<output_path>/.quant_summary.txt`, mirroring `llm_ptq/hf_ptq.py`.

## Why

Forcing all tokens through every expert doubled calibration time and
inflated `input_quantizer.amax` for cold-routing experts with outliers
they never see at inference. The new flow matches the inference
distribution, runs roughly 2x faster, and mirrors the
`layer_sync_moe_local_experts_amax` semantics that mtq runs
automatically for `QuantSequentialMLP`-derived MoEs.

## Validation (DeepSeek-V3.2-Exp, MP=8, NVFP4_DEFAULT_CFG)

Compared `_amax_baseline` (CalibMoe) vs `_amax_synced` (new default):
- All 44,544 expert weight amaxes bit-identical.
- Attention, shared experts, gate: identical.
- Expert `w1.input` and `w3.input` (shared MoE block input): identical.
- Expert `w2.input` (post-SiLU gated, expert-specific): synced to
layer-wide peer max — 99.3% are larger than baseline (median 11.4x)
since peer-max captures the worst-case outlier from any expert in the
layer; 0.7% are smaller. This is the same trade-off
`set_expert_quantizer_amax` makes for HF MoEs in `unified_export_hf.py`.

## Test plan

- [x] DeepSeek-V3.2-Exp MP8 PTQ with default flags — completes in ~7 min
(vs ~27 min with CalibMoe), produces `_amax_synced/` consistent with the
comparison above.
- [x] DeepSeek-V3.2-Exp MP8 PTQ with `--calib_all_experts` — produces
`_amax_baseline/` identical (other than rounding) to the prior
`CalibMoe`-default behavior.
- [x] `.quant_summary.txt` written under `output_path` on rank 0.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a `--calib_all_experts` option to enable an alternate PTQ
calibration mode; default remains top-k routing with a post-calibration
per-layer peer-max synchronization and a compute fallback for
uncalibrated experts.
* **Documentation**
* Clarified default and alternate calibration behaviors and added note
about generation of a `.quant_summary.txt` summary file.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2026-05-04 11:08:01 -07:00
yeyu-nvidiaandClaude Sonnet 4.6 383ab4e224 fix: include medusa in data_module assignment in main.py (#1370)
## Problem
When `training.mode == "medusa"` is used in `main.py`, the `data_module`
variable is never assigned because line 344 only covered `eagle3` and
`dflash` modes. This causes an `UnboundLocalError` when the trainer is
constructed with `**data_module`.

Fixes OMNIML-4147

## Fix
Add `"medusa"` to the `training_args.mode in ("eagle3", "dflash")`
condition so `data_module` is correctly populated for medusa training.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed speculative decoding example to properly handle "medusa" mode
alongside existing "eagle3" and "dflash" modes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-04 15:17:50 +05:30
Chenjie LuoandClaude Opus 4.7 50706d1750 Add closed-form MXFP4 -> NVFP4 weight cast (--cast_mxfp4_to_nvfp4) (#1372)
## Summary

- New `--cast_mxfp4_to_nvfp4` flag in `hf_ptq.py` (and
`huggingface_example.sh`) that converts an MXFP4 source checkpoint (e.g.
`openai/gpt-oss-20b`) into an NVFP4 export with **bit-exact** weight
reconstruction for the in-range blocks.
- The cast pins NVFP4's `scale_2 = 2^m` (where `m = k_max − 8`) and
`_amax = 6·2^k_j` per NVFP4 block, both read from the source `*_scales`.
The resulting per-block scale `2^(k_j − m)` is exactly representable in
E4M3, so `round_to_E2M1(value / 2^k_j)` yields the original MXFP4 nibble
verbatim. For out-of-range blocks (`k_max − k_j > 17`) the per-block
amax falls back to data-derived `max(|w_block|)`, which keeps the
post-E4M3-clamp scale close to the block's actual magnitude.

## Verification

End-to-end on `openai/gpt-oss-20b` with `--qformat=nvfp4_mlp_only
--cast_mxfp4_to_nvfp4`:

```
[cast_mxfp4_to_nvfp4] overrode 48/48 weight quantizers
[cast_mxfp4_to_nvfp4] lossless layers: 48/48 (100.00%)
[cast_mxfp4_to_nvfp4] lossless blocks: 597196800/597196800 (100.0000%)
```

End-to-end on `openai/gpt-oss-120b` with the same flags (4×B200,
`--use_seq_device_map --gpu_max_mem_percentage 0.5 --calib_batch_size
4`):

```
[cast_mxfp4_to_nvfp4] overrode 72/72 weight quantizers
[cast_mxfp4_to_nvfp4] lossless layers: 67/72 (93.06%)
[cast_mxfp4_to_nvfp4] lossless blocks: 3583179586/3583180800 (100.0000%)
```

Five layers fall into the OOR regime (block-spread > 17); the remaining
1,214 OOR blocks use the data-derived per-block amax fallback.
Block-level losslessness is **99.99996%** end-to-end.

Per-tensor MSE between MXFP4 source dequant and NVFP4 export dequant
(~19B elements):

| Metric | Without cast | With cast |
|---|---|---|
| Per-tensor SNR | ~26.4 dB (FP4 noise floor) | **∞ (every tensor)** |
| Total RMSE | 8.67e−02 | **0** |
| max\|err\| | up to 8.0e+1 | **0** |

## Modelopt-side enablers

- `max_calibrate` auto-promotes static-block NVFP4 weight quantizers to
`NVFP4StaticQuantizer` at the end of calibration.
- `static_blockwise_fp4_fake_quant` kernel accepts N-D inputs (was
2D-only), unblocking MoE expert weights of shape `(E, F, K)`.
- BMM-experts NVFP4 export routes through
`get_weights_scaling_factor_from_quantizer` for static-mode quantizers,
so the pinned `_amax` is actually consumed.
- `set_expert_quantizer_amax` scalar-reduces per-quantizer amax before
stacking, supporting per-block (vs scalar) static-mode amax.

## Test plan

- [x] Unit tests at `tests/examples/llm_ptq/test_cast_mxfp4_to_nvfp4.py`
(15 tests, all passing) cover: scalar/global-amax math, per-block hybrid
(in-range closed-form vs OOR data-derived), shape preservation, key
collection, and end-to-end `build_amax_map` against a synthetic
safetensors checkpoint.
- [x] End-to-end PTQ → export on `openai/gpt-oss-20b` (`nvfp4_mlp_only`
qformat) with `--cast_mxfp4_to_nvfp4` succeeds; export takes ~21 s. 100%
lossless cast (48/48 layers, 597,196,800 / 597,196,800 blocks).
- [x] End-to-end PTQ → export on `openai/gpt-oss-120b` (4×B200,
`nvfp4_mlp_only`, `--use_seq_device_map --gpu_max_mem_percentage 0.5
--calib_batch_size 4`). 67/72 layers fully lossless; 99.99996%
block-level losslessness (3,583,179,586 / 3,583,180,800).
- [x] TRT-LLM serving validation (TRT-LLM 1.3.0rc11, B200) on both
exported NVFP4 checkpoints via `examples/llm_ptq/run_tensorrt_llm.py`:
- **20b** (TP=1): 18.3 GB GPU memory; coherent generation. Sample:
*"Quantum computing is poised to revolutionize data analysis. However,
its potential is currently limited by quantum hardware constraints,
including error rates, qubit lifetimes, and lack of fault tolerance…"*
- **120b** (TP=4): 36.4 GB / GPU; coherent generation. Sample: *"Quantum
computing is poised to revolutionize data storage and processing. These
rare earth-based systems could serve as robust qubits; resistant to
environmental decoherence…"*
- [x] MSE comparison script (run separately during development) confirms
per-tensor SNR=∞ across all 48 MoE expert tensors.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a MXFP4→NVFP4 weight-format cast utility and a CLI flag to
enable it; helper scripts updated to expose the option.

* **Bug Fixes**
  * Fixed static NVFP4 export for expert weights.
* Improved collection/handling of quantizer amax values to avoid shape
issues.
  * Generalized FP4 kernel to accept flexible tensor dimensionality.
  * Ensured static-block NVFP4 promotion during calibration.

* **Tests**
* Added comprehensive tests for the conversion workflow, helpers, and
end-to-end application.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-01 19:04:04 +00:00
Keval MorabiaandClaude Sonnet 4.6 bb08094ff1 Add Nemotron-Nano-9B-v2 → Pruned 7B e2e tutorial: Prune + Distill + Eval + Quantize + vLLM deployment (#1325)
## Summary

End-to-end optimization walkthrough for Nemotron-Nano-9B-v2 showing how
ModelOpt techniques stack:

- **Pruning** — Minitron structured pruning 9B → 7B
- **Distillation** — Megatron-Bridge knowledge distillation up to 80B
tokens; near-parity with official 9B on MMLU Pro, GPQA, LCB, AIME, Math
500, IFEval, SciCode
- **Evaluation** - using nemo-evaluator
- **Quantization** — FP8 PTQ via \`hf_ptq.py\`; checkpoint deployable on
vLLM/TRT-LLM/SGLang with no extra flags (quantization auto-detected from
\`config.json\`)
- **vLLM Throughput** — BF16 vs FP8 benchmark on single H100

<img width="2085" height="1740" alt="image"
src="https://github.com/user-attachments/assets/8620a019-5c09-4a6b-a5d2-ca164aaa5d87"
/>

<img width="2085" height="810" alt="image"
src="https://github.com/user-attachments/assets/742c8035-f1fb-4394-b11b-0c6c3ac4e843"
/>


### Files changed

- `examples/pruning/minitron/README.md` — index page for Minitron
end-to-end tutorials
- `examples/pruning/minitron/NVIDIA-Nemotron-Nano-9B-v2/README.md` —
full repro doc with 6 sections: data prep, pruning, distillation,
evaluation, FP8 quantization, vLLM benchmarking
-
`examples/pruning/minitron/NVIDIA-Nemotron-Nano-9B-v2/nemo_evaluator.yaml`
— NeMo Evaluator config used for all benchmark numbers
- `examples/pruning/puzzletron/README.md` — index page for Puzzletron
distillation results
- `examples/pruning/puzzletron/Llama-3.1-8B-Instruct.md` — Puzzletron
distillation results (renamed from puzzletron.md)
- `examples/pruning/README.md` — updated Results section with direct
links to new locations
- `examples/megatron_bridge/README.md` — updated results link to point
to `examples/pruning/`
- `examples/puzzletron/README.md` — updated distillation results link
- `examples/dataset/MEGATRON_DATA_PREP.md` — tokenization commands for
all datasets used in the data blend

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Documentation

* **New end-to-end tutorial** for model optimization covering Minitron
pruning, knowledge distillation, FP8 quantization, and vLLM deployment
with reproducibility steps and benchmark results
* **Dataset preparation guide** with ready-to-run tokenization templates
for Nemotron HuggingFace datasets
* **Evaluation configuration** and results documentation including
ablation studies across multiple benchmarks
* **Updated navigation** across pruning, distillation, and dataset
examples to streamline user workflows

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-30 14:47:16 +05:30
Grzegorz K. Karch 6d330784a6 Add required keys to attention pruning config (#1360)
### What does this PR do?

Type of change: ? Bug fix

The config
`examples/puzzletron/configs/llama-3_1-8B_pruneffn_memory/pruning/attn_pruning.yaml`
didn't have required keys to use attention pruning in the example
`examples/puzzletron/main.py`

### Usage


### Testing

In
`examples/puzzletron/configs/llama-3_1-8B_pruneffn_memory/Llama-3_1-8B.yaml`
change `ffn_pruning` to `attn_pruning`

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Updated pruning configuration for improved KV-head pruning support,
including enhanced importance hook settings and attention output
handling for memory optimization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Grzegorz Karch <gkarch@nvidia.com>
2026-04-28 13:04:51 +00:00
h-guo18 1ec931c2c7 [2/3][Feat]: Offline DFlash training (#1343)
### What does this PR do?

Type of change: new feature

Part 2 of a 3-PR series splitting #1271:
- **[1/3] #1296**: File reorg + deprecate `ParallelDraft`
- **[2/3] this PR**: Offline DFlash training (depends on #1296)
- **[3/3] #1297**: Extract `HFSpecDecMixin`

Changes:
- Add `dflash_offline` flag to `DFlashConfig` for training from
pre-computed hidden states; deletes base model layers to save memory.
- Add Pydantic validators on `DFlashConfig`:
- `_derive_dflash_offline` — auto-derive `dflash_offline` from
`data_args.offline_data_path` in validation context. Not
user-configurable: any user-supplied value is overridden by the derived
value.
- `_resolve_mask_token_id` — auto-detect `dflash_mask_token_id` from
`tokenizer.mask_token_id`.
  - `_check_mask_token_id` — fail fast if unset after resolution.
- `HFDFlashModel.modify()`: select `num_orig_hidden_layers` when
offline; pick `_base_model_lm_head` device when no base layers present;
drop base-model `layers` module.
- `HFDFlashModel.forward()`: add offline branch — consumes precomputed
`base_model_outputs` via `DFlashBaseModelOutput.from_offline_dict`, and
when `dflash_self_logit_distillation` is enabled with
`base_model_logits` absent, recomputes logits from
`base_model_hidden_states` via `_base_model_lm_head`. Raises a clear
error from the non-training / `pseudo_speculative_generate` paths when
`dflash_offline=True`, since base-model layers have been deleted.
- `DFlashBaseModelOutput` dataclass in `modeling_dflash.py` (with
`from_offline_dict` classmethod) to unify online/offline output shapes.
`aux_hidden_states` is required in `from_offline_dict` so missing keys
fail fast at the entry point rather than deeper in the forward.
- `examples/speculative_decoding/main.py`: replace inline
`mask_token_id` auto-detect with
`DFlashConfig.model_validate(dflash_cfg, context={"tokenizer":
tokenizer, "data_args": data_args})`.

### Silent bug fix — `add_generation_template` → `add_generation_prompt`

The pre-refactor `compute_hidden_states_hf.py` passed
`add_generation_template=False` to `tokenizer.apply_chat_template`. This
kwarg does not exist on HF `apply_chat_template` and was being silently
ignored, so the intended "don't append a generation prompt" behavior was
never actually applied. The new `tokenize_with_loss_mask` helper in
`examples/speculative_decoding/collect_hidden_states/common.py` uses the
correct `add_generation_prompt=False`. **This is a real behavior
change** for anyone re-dumping hidden states: trailing generation
prompts that were previously appended to the tokenized sequences will no
longer be included.


### Testing
- New tests:
- `tests/unit/torch/speculative/plugins/test_hf_dflash_offline.py` — CPU
unit tests for convert path (online keeps base layers, offline deletes
them; `num_orig_hidden_layers` drives `target_layer_ids` in offline
mode) and `DFlashConfig._derive_dflash_offline` validator.
- `TestDFlashOfflineForwardGPU` in
`tests/gpu/torch/speculative/plugins/test_hf_dflash.py` — GPU forward
smoke with precomputed `base_model_outputs`, plus the
`dflash_self_logit_distillation` logit-recompute path.

- training test:
<img width="454" height="317" alt="image"
src="https://github.com/user-attachments/assets/79b92790-4d15-4313-bb9b-f35665b012e6"
/> <img width="456" height="310" alt="image"
src="https://github.com/user-attachments/assets/4558559f-9c35-49ed-b36e-82fbc99eab23"
/>


### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ — additive `dflash_offline`
flag defaulting to `False`; validators fall through when context not
provided.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — see Testing section above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

### TODO (follow-up)

- [x] Update
`examples/speculative_decoding/collect_hidden_states/compute_hidden_states_*.py`
to support DFlash offline data. Current scripts are Eagle-specific —
they hardcode the `[2, N/2, N-3]` aux-layer selection and emit
`{input_ids, hidden_states, aux_hidden_states}`. DFlash offline needs:
- Aux layer indices driven by
`build_target_layer_ids(num_orig_hidden_layers, num_draft_layers)` (or a
configurable list), not the Eagle triplet.
- `base_model_hidden_states` key (last-layer hidden) so
`DFlashBaseModelOutput.from_offline_dict` + the
`dflash_self_logit_distillation` recompute path can consume it.
- Optional `base_model_logits` dump so offline training can skip the
self-distillation logit recomputation when logits are available.

### Additional Information

Base branch is #1296 (file reorg). Retarget to `main` once #1296 merges.



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Offline DFlash speculative-decoding training from precomputed
base-model hidden states
* Answer-only-loss training with persisted loss masks and optional
chat-template support
* Flexible auxiliary-layer selection via CLI and an exposed default
aux-layer helper
* Auto-derived offline flag in config and automatic memory optimization
during offline conversion

* **Documentation**
* Updated guides for offline pipeline, aux-layer selection, and
loss-masking options

* **Tests**
* New unit, GPU, and regression tests covering offline conversion,
training, and config derivation
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-04-25 18:38:35 -07:00
h-guo18 7c80d85751 [1/3][Refactor]: File reorg; deprecate ParallelDraft (#1296)
### What does this PR do?

Type of change: refactoring

Part 1 of a 3-PR series splitting #1271:
- **[1/3] this PR**: File reorg + deprecate `ParallelDraft`
- **[2/3] #1295**: Offline DFlash training
- **[3/3] #1297**: Extract `HFSpecDecMixin`

Changes:
- **File reorg**: `transformers.py` → `hf_eagle.py`; extract
`HFMedusaModel` → `hf_medusa.py`; extract `EagleModule` /
`EagleBaseModelOutput` → `modeling_eagle.py`; extract `DFlashModule` /
`DFlashAttention` / `DFlashDecoderLayer` / `build_target_layer_ids` /
`apply_rotary_pos_emb` → `modeling_dflash.py`.
- **Deprecate `ParallelDraft`**: remove `parallel_draft_step`,
`parallel_draft_heads_num_layers`, and the `ParallelDraft` module from
HF Eagle; remove the `EagleMedusaExporter` branch from
`HFEagleModel.get_exporter()` (the `EagleMedusaExporter` class itself
still lives in `hf_spec_export.py` for Megatron parity).
- **Rename**: `_draft_model_config` → `eagle_config` in export plugin.
- Update imports in `examples/speculative_decoding/` and
`modelopt/torch/speculative/utils.py` to follow the module rename.

### Testing

Validated with existing Eagle and DFlash training scripts (re-run after
`9ae5302729 revert behavior change`).

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ — renames
`modelopt.torch.speculative.plugins.transformers` → `.hf_eagle`; removes
`parallel_draft_step` / `parallel_draft_heads_num_layers` from Eagle
config; renames `_draft_model_config` → `eagle_config` in export plugin.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — pure refactor; existing
tests updated for the rename. `test_hf_spec_rope_export.py` assertions
were also corrected to reflect the actual production path (the old
assertions were masked by `MagicMock` not invoking the
`_draft_model_config` `@property`).
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌

### Additional Information

Breaking changes:
- `modelopt.torch.speculative.plugins.transformers` → `.hf_eagle`
- `parallel_draft_step` / `parallel_draft_heads_num_layers` removed from
Eagle config
- `_draft_model_config` → `eagle_config` in export plugin

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Refactoring**
* Reorganized speculative-decoding plugins into focused modules,
converting the legacy "transformers" entry into a deprecated shim that
re-exports the new plugin surface.
* Consolidated DFlash implementation into a shared modeling component
and introduced a dedicated EAGLE decoder module.

* **New Features**
* Added a Medusa speculative-decoding plugin with configurable heads and
combined-loss training behavior.

* **Chores**
  * Updated pre-commit license-hook exclusion and feature-flag wiring.

* **Tests**
  * Updated export tests to expect rope-scaling fallback semantics.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-04-24 14:46:10 -07:00
Liana Mikaelyan 946639aa19 Fix PTQ for VLMs with image calibration (#1318)
### What does this PR do?

This PR fixes PTQ with image claibration for VLMs. 

### Usage

```python
python3 examples/llm_ptq/hf_ptq.py --pyt_ckpt_path Qwen/Qwen3-VL-8B-Instruct --qformat fp8 --export_path Qwen3-VL-8B-Instruct-fp8 --trust_remote_code --kv_cache_qformat none --calib_with_images --calib_size 512
```

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Image-text calibration now extends support to additional model
architectures when image calibration is enabled.
* Improved tokenizer truncation handling in multimodal dataset
processing to prevent configuration conflicts when image inputs are
present.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Liana Mikaelyan <lmikaelyan@nvidia.com>
2026-04-24 13:29:59 -07:00
Chenjie Luo fda0899e40 feat(recipes): add KV cache cast variants (fp8_cast / nvfp4_cast) (#1334)
## Summary

- Adds three built-in PTQ recipes that express the KV-cache *cast*
variants directly in YAML, using the existing `use_constant_amax: true`
quantizer field. These are recipe equivalents of
`--kv_cache_qformat=fp8_cast` / `nvfp4_cast`:
  - `general/ptq/fp8_default-fp8_cast_kv`
  - `general/ptq/nvfp4_default-fp8_cast_kv`
  - `general/ptq/nvfp4_default-nvfp4_cast_kv`
- Makes `--recipe` authoritative in `examples/llm_ptq/hf_ptq.py`: the
post-hoc `_set_kv_cache_constant_amax` override now only runs when
`--recipe is None`, so a recipe YAML fully determines KV-cache config
instead of being silently overridden by the default
`--kv_cache_qformat=fp8_cast`. Updated help text on both flags.
- Extends the recipe loader smoke test to cover the three new recipes.

## Motivation

Before this change, the cast variants lived only in argparse
(`_KV_CAST_FORMATS = {"fp8_cast", "nvfp4_cast"}`) and were layered on
top of any recipe-loaded config. That meant `--recipe
nvfp4_default-fp8_kv` would silently become a cast recipe due to the
`--kv_cache_qformat` default. Now the recipe is self-contained: its YAML
either sets `use_constant_amax: true` on the `*[kv]_bmm_quantizer` entry
(cast) or doesn't (data-driven calibration).

## Test plan

- [x] `pytest tests/unit/recipe/test_loader.py` — all 24 tests pass,
including the three new parametrized recipes.
- [x] Verified each new recipe round-trips through `load_recipe()` with
`use_constant_amax: True` surviving Pydantic validation on the KV entry.
- [x] End-to-end run on `/models/Qwen/Qwen3-8B` (RTX 6000 Ada, 4
samples, seq_len=128) for all three new recipes:
- After `mtq.quantize(model, recipe.quantize.model_dump(),
forward_loop=...)`, all 72 `k_bmm_quantizer` / `v_bmm_quantizer` modules
have `_use_constant_amax=True` and `_get_amax()` returns `448.0` (FP8
E4M3 max).
- Weight quantizers still calibrate from data normally (sample amax
values: q_proj=0.5508, k_proj=0.6250, v_proj=0.1689, o_proj=0.7266).
- [x] Verified the `--recipe` authoritative behavior change:
- Non-cast recipe + default `--kv_cache_qformat=fp8_cast` → KV entry
does NOT get `use_constant_amax` (no silent override).
- Cast recipe + contradictory `--kv_cache_qformat=fp8` → KV entry keeps
`use_constant_amax=True` (recipe wins).

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed CLI to respect KV cache quantization settings from recipe YAML
instead of overriding them.

* **New Features**
* Added three new post-training quantization recipe configurations for
FP8 and NVFP4 with optimized KV cache handling.

* **Documentation**
* Enhanced CLI help text for recipe and KV cache quantization options
with configuration examples.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-23 18:45:16 +00:00
Chenjie Luo 01788bb007 Deprecate Mllama support in llm_ptq/vlm_ptq examples (#1332)
## Summary

- Removes Mllama (Llama 3.2 Vision) model-type branches from the
`llm_ptq` example (`hf_ptq.py`, `example_utils.py`) and drops the
now-unused `MllamaImageProcessor` wrapper from `modelopt/torch/utils/`.
- Drops the legacy `MllamaImageProcessor` path in
`modelopt/torch/utils/vlm_dataset_utils.py`; the generic HF
ProcessorMixin path handles the remaining cases.
- Adds a CHANGELOG entry under 0.44 Backward Breaking Changes.

## Test plan

- [x] CI lint / unit tests pass
- [x] Smoke-run ``examples/llm_ptq/scripts/huggingface_example.sh
--model <llm> --quant fp8`` (text-only path, non-mllama)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Removed Mllama (Llama 3.2 Vision) support from quantization examples.
This includes removal of dedicated image processor implementation,
specialized model handling, and related calibration logic.
* Updated VLM image-text calibration guidance to use
`--calib_with_images` flag with other supported VLMs instead of
Mllama-specific processing paths.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-23 23:24:42 +05:30
Michael Feil 8663678f12 fix: bug hf_ptq.py max_length setting ignored for LLMs (#1311)
### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage
- fixes a bug in example script. We were trying why our models were not
that strong at long context. Seems like a recent refactor did not
implement max seq length. so 512 is used by default.

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **Improvements**
* Calibration data loading now enforces a maximum sequence/sample length
during dataset preparation, ensuring calibration inputs adhere to
configured length limits. This yields more predictable calibration
behavior, reduces peak memory usage during calibration runs, and
improves consistency of quantization preprocessing.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Michael Feil <63565275+michaelfeil@users.noreply.github.com>
2026-04-23 09:02:31 -07:00
Ajinkya Rasane e4e3508f51 [OMNIML-3349] Add FP8 MHA quantization support for HuggingFace ViT (#1289)
## Summary
Enables TensorRT attention-v2 fusion for vision transformers when
exported to ONNX with FP8 Q/DQ. The core library changes are
architecture-agnostic (drop-in for any FP8 ONNX export); coverage is
exercised by the existing `examples/torch_onnx/torch_quant_to_onnx.py`
pipeline.

- **`modelopt/onnx/export/fp8_exporter.py`** — new post-processing
passes: move attention-scaling `Mul` and K `Transpose` to the Q-side so
DQ feeds MatMul directly, pre-transpose constant weights, and insert FP8
Q/DQ on Softmax outputs (fixed `1/448` scale, data-independent) for
MHA-v2 fusion. Rewrites only fire when every downstream consumer is a
MatMul so non-attention branches are never perturbed.
- **`modelopt/onnx/utils.py`** — `fold_dq_fp32_to_fp16_casts` /
`fold_q_fp16_to_fp32_casts` remove the Cast nodes
`convert_float_to_float16` inserts around Q/DQ and rewrite scale
initializers to FP16 so TRT fuses DQ into the downstream GEMM. Guarded
behind opset >= 19 (FP16 Q/DQ scale requirement). Warns on FP16
overflow/underflow.
- **`modelopt/torch/_deploy/utils/torch_onnx.py`** — calls the fold
helpers for FP8-quantized models after `convert_float_to_float16`.
- **`modelopt/torch/quantization/export_onnx.py`** — keeps FP8 Q/DQ
scale in the native input dtype so no Cast is emitted between graph and
Q/DQ. Removes the now-unused `trt_high_precision_dtype` parameter from
`_fp8_quantize`/`_fp8_dequantize`.
- **`modelopt/torch/quantization/nn/modules/quant_layernorm.py`** (new)
— registers `nn.LayerNorm` in `QuantModuleRegistry` so LayerNorm output
quantizers are honored.
- **`modelopt/torch/quantization/plugins/huggingface.py`** — skips
`*Attention` wrappers whose children are also `*Attention` per-instance
(not per-class) to avoid double-patching `eager_attention_forward` (e.g.
`ViTAttention` vs `ViTSelfAttention`).
- **`examples/torch_onnx/torch_quant_to_onnx.py`** — adds a
`_FP8_MHA_OVERRIDE` config block to FP8 mode that enables LayerNorm
output quantizer + disables its input quantizer for TRT attention
fusion.
- **Unit tests** (12 CPU tests, ~1.2s total) — fp8_exporter rewrites +
fanout safety, fold-cast helpers + opset guard, LayerNorm quant-wrapper
identity, per-instance nested-attention detection.

## Benchmarks
ViT-base-patch16-224, RTX 6000 Ada, strongly-typed FP8 via `trtexec`.
Accuracy on 2 000 ImageNet-1k validation samples (streaming).

**Batch = 1 (latency-bound)**
| Model | Top-1 | Top-5 | TRT latency | Speedup |
|---|---|---|---|---|
| FP16 baseline | 80.96% | 95.80% | 0.722 ms | 1.00x |
| Torch FP8 MHA | 80.66% | 95.75% | 0.657 ms | **1.10x** |
| ONNX PTQ FP8 | — | — | 0.589 ms | **1.23x** |

**Batch = 64 (throughput-bound, realistic inference)**
| Model | TRT latency | Speedup | Images/s |
|---|---|---|---|
| FP16 baseline | 23.40 ms | 1.00x | 1152 |
| Torch FP8 MHA | 15.89 ms | **1.47x** | 1152 |
| ONNX PTQ FP8 | 15.89 ms | **1.47x** | 1216 |

Top-1 accuracy stays within 0.30 pp of FP16; at batch=64 the Torch FP8
MHA path matches ONNX PTQ wall-time — attention is the bottleneck there
and both paths achieve full FP8 attention fusion (36/36 attention
MatMuls with QDQ in ViT-base).

## Test plan
- [x] CPU unit tests (new): \`python -m pytest
tests/unit/onnx/quantization/test_fp8_mha_exporter.py
tests/unit/onnx/test_fold_casts.py
tests/unit/torch/quantization/test_quant_layernorm.py
tests/unit/torch/quantization/plugins/test_nested_attention_skip.py\`
- [x] Existing ONNX / quantization unit suites unaffected: \`python -m
pytest tests/unit/onnx tests/unit/torch/quantization\`
- [x] End-to-end ViT FP8 export: \`python
examples/torch_onnx/torch_quant_to_onnx.py --timm_model_name
vit_base_patch16_224 --quantize_mode fp8 --onnx_save_path
vit_base_fp8.onnx\` — expect log lines \`Folded 48 weight Transpose
nodes\`, \`Inserted FP8 weight DequantizeLinear for 1 Conv nodes\`, and
\`Attention QDQ rewrites: ... inserted QDQ on 12 Softmax outputs\`
- [x] trtexec FP8 strongly-typed build: \`trtexec
--onnx=vit_base_fp8.onnx --fp8 --stronglyTyped\`
- [x] Accuracy within ~0.3 pp of FP16 baseline on ImageNet-1k subset

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2026-04-23 11:38:31 -04:00
Keval Morabia 7265ca6793 Fix lm_eval version checking (#1321)
### What does this PR do?

Type of change: bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

`lm_eval` does not have `__version__` attribute

### Additional Information
<!-- E.g. related issue. -->
NVBug 6102101


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Enhanced the package version detection system to improve overall
reliability and stability of the application while reducing unnecessary
external dependencies. All functionality, including version gating and
system warnings, continues to operate exactly as expected with no impact
on the user experience.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-23 16:57:54 +05:30
Grzegorz K. KarchandKeval Morabia 2564da7324 Update vLLM deployment docs for heterogeneous models (#1317)
### What does this PR do?

Type of change: ? documentation.

This PR updates vLLM deployment instructions, taking into account
heterogenous models created with AnyModel.

### Usage

Does not apply.

### Testing

Run the updated instructions in the documentation.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Replaced a benchmarking-focused section with a deployment guide for
running compressed models on vLLM.
* Added step-by-step setup for using an AnyModel-enabled vLLM fork,
including checkout and install guidance and required model config edits
(with optional architecture metadata).
* Simplified runtime to a single vllm serve command, removing manual
model rearrangement steps.
* Restored inference benchmarking as a subsection, retaining vllm bench
latency/throughput examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Grzegorz Karch <gkarch@nvidia.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-23 06:22:37 +00:00
jingyu-ml c7966119eb Reorg the sparse/quant/common kernel dir (#1303)
### What does this PR do?

Type of change: re-org code <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ We changed the import path
<!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ❌ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:✅
<!--- Only for new features, API changes, critical bug fixes or backward
incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Calibration support for skip-softmax multi-threshold measurement in
sparse attention.
  * N:M sparse softmax masking and helpers for sparsity-aware attention.

* **Chores**
* Reorganized and consolidated kernel/backends for quantization and
sparsity to a unified kernels layout, updating tests and examples to
match.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-04-22 23:34:56 +00:00
kinjalpatel27 0678136335 Fix vLLM fakequant MoE megatron export bug (#1305)
### What does this PR do?

Type of change: Bug fix                     
Fixes two bugs in the vLLM + Megatron-Core MoE export path and cleans up
the related weight-collection helper:
    
1. **`_QuantFusedMoEBase` (vllm.py)**: The weight-quantizer path in
`_invoke_fused_moe_quantized_function` was temporarily mutating
`self.w13_weight` / `self.w2_weight` to the quantized tensor, then
restoring them via `finally`. This exposed a stale quantized tensor on
`self` between the mutation and the kernel call. Fixed by computing the
quantized weight directly into a local `B` without touching `self.*`
attributes.
2. **`GPTModelExporter` / `VllmFqGPTModelExporter`
(unified_export_megatron.py / vllm_fakequant_megatron.py)**:
`expert_bias` (present in grouped MoE layers) was silently dropped
during export because the bias collection ran after the early-return on
missing `weight`. Extracted a `_get_weight_bias` helper that collects
weight, bias, and expert_bias together, so bias/expert_bias are captured
even when weight is absent or zero-element.

### Usage

```python
# No API change; export pipelines pick this up automatically.
# export_mcore_gpt_to_hf_vllm_fq / export_mcore_gpt_to_hf now correctly                   
# export expert_bias for grouped-MoE checkpoints.    
```

### Testing
Step 1 — Quantize (run from Megatron-LM
examples/post_training/modelopt):
```
HF_MODEL_CKPT=<path/to/hf/weights> MLM_MODEL_SAVE=<quant-ckpt-name> \
bash quantize.sh <hf-model-id> NVFP4_DEFAULT_CFG
```
Step 2 — Export for vLLM fakequant:
```
MLM_EXTRA_ARGS=--export-vllm-fq \ 
HF_MODEL_CKPT=<path/to/hf/weights> \ 
MLM_MODEL_CKPT=<quant-ckpt-name> \ 
EXPORT_DIR=<export-dir> \ 
bash export.sh <hf-model-id> 
```
Step 3 — Serve (run from examples/vllm_serve):
```
 QUANT_CFG=NVFP4_DEFAULT_CFG \ 
 QUANT_FILE_PATH=<export-dir>/quantizer_state.pth \ 
 python3 vllm_serve_fakequant.py <export-dir> \ 
 -tp 1 --served-model-name <model-name> \ 
 --host 0.0.0.0 --port 8000 \ 
--trust-remote-code --enforce-eager \ 
 --disable-custom-all-reduce \ 
--gpu-memory-utilization 0.8          
```
### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Refactor**
* Centralized weight/bias/expert-bias extraction and export to a single
helper for consistent handling.
* Standardized quantized-weight flow to temporarily swap and restore
parameter tensors during computation.

* **Bug Fixes**
* Prevented missing or incorrect weight/bias exports by unifying
extraction logic.
* Broadened checkpoint key matching to preserve more quantizer state
during reloads.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-04-22 09:24:25 -07:00
yeyu-nvidiaandClaude Opus 4.6 2fef374ded fix: auto-compute dp_replicate_size from world_size (#1302)
## Summary
- When `dp_shard_size < world_size` (e.g., `dp_shard_size=4` on 8 GPUs
across 2 nodes), `ParallelismConfig` raises `total_size (4) does not
match num_processes (8)` because `dp_replicate_size` defaults to 1
- Auto-compute `dp_replicate_size = world_size // (dp_shard_size *
cp_size)` so intra-node FSDP2 sharding + inter-node data-parallel
replication works without manual config
- This enables `dp_shard_size` to be set to per-node GPU count (better
NVLink utilization) while automatically creating replicas across nodes

## Test plan
- [ ] Verify single-node training (dp_shard_size == world_size,
dp_replicate_size == 1) unchanged
- [ ] Verify multi-node with dp_shard_size < world_size creates correct
replica groups
- [ ] Verify existing EAGLE3/DFlash configs still work

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Enhanced parallelism configuration initialization in the speculative
decoding example to better handle distributed training scenarios.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-20 20:39:36 +00:00
Chenhan D. YuandClaude Opus 4.6 355c6b7883 fix: PTQ 1GPU, export PP divisibility, hidden states conversations key (#1293)
## Summary
- **megatron_lm_ptq.yaml**: Qwen3-8B PTQ to single GPU for L40 clusters
(TP=1, all tasks)
- **quantize.sh**: Auto-find largest PP dividing model's
`num_hidden_layers` for export step. Qwen3-8B has 36 layers which isn't
divisible by 8, causing `AssertionError` on 8-GPU nodes
- **compute_hidden_states_trtllm.py**: Use `messages` with
`conversations` fallback, matching the HF version. Fixes `KeyError:
'conversations'` when data uses OpenAI `messages` format

## Test plan
- [x] Qwen3-8B PTQ runs on single L40 GPU
- [x] Export PP auto-selects valid divisor (36 layers → PP=6 on 8 GPUs,
PP=4 on 4 GPUs, PP=1 on 1 GPU)
- [x] EAGLE3 offline pipeline reads data with `messages` field

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Dataset input handling now supports multiple field formats for
enhanced compatibility.

* **Bug Fixes**
* Optimized GPU resource allocation during model quantization with
improved pipeline parallelism computation.
* Updated quantization configuration for more efficient resource
utilization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 12:43:58 -07:00
kinjalpatel27 010b220dc0 vLLM fakequant export update for AWQ checkpoint (#1242)
### What does this PR do?

Type of change: Bug

Enables end-to-end AWQ checkpoint export and reload in the vLLM
fake-quant serving path (`MODELOPT_STATE_PATH`). Previously, the
`input_quantizer` was using incorrect `pre_quant_scale` especially with
grouped quantizers like `qkv_proj`, using simply the first
`input_quantizer.pre_quant_scale`. This MR adds
`_resmooth_experts_for_export` that non-mutatively averages
`pre_quant_scale` across MoE experts and unifies input `_amax`, required
because vLLM uses a single input quantizer per expert group. Adds
`merge_amax_tensors_for_group` (element-wise max for same-shape, `cat`
for GQA, scalar-max fallback) replacing the scalar-collapsing
`torch.stack().max()` that dropped per-channel `_amax` structure.

### Usage

```python
# Export AWQ checkpoint from HF model
  from modelopt.torch.export.plugins.vllm_fakequant_hf import export_hf_vllm_fq_checkpoint
  export_hf_vllm_fq_checkpoint(model, export_dir="./awq_vllm_checkpoint")      
```

### Testing
**Step 1 — Export the quantized checkpoint:**
  ```bash                    
  python examples/llm_ptq/hf_ptq.py \
--pyt_ckpt_path <MODEL_PATH> \
    --recipe <AWQ_RECIPE> \
--calib_size 512 \
--export_path <EXPORT_DIR> \
    --vllm_fakequant_export
```
  This produces `<EXPORT_DIR>/vllm_fq_modelopt_state.pth` with the averaged per-expert                                                                                                                           
  pre_quant_scale and unified _amax now included.                                                                                                                                                              
   

 Step 2 — Serve via vLLM fakequant worker:                                                                                                                                                                    
```bash
  MODELOPT_STATE_PATH=<EXPORT_DIR>/vllm_fq_modelopt_state.pth \
python examples/vllm_serve/vllm_serve_fakequant.py \
      <EXPORT_DIR> --tensor-parallel-size <TP>   
```

Tested for quantization configurations:
```
FP8_DEFAULT_CFG
FP8_DEFAULT_CFG (input_q disabled)
INT8_SMOOTHQUANT_CFG
INT8_WEIGHT_ONLY_CFG
NVFP4_DEFAULT_CFG
NVFP4_AWQ_LITE_CFG
INT4_AWQ_CFG
NVFP4_AWQ_CFG
NVFP4_DEFAULT_CFG (input_q disabled)
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A 
- Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A 

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit

* **New Features**
  * Added Nemotron-style MoE export support and group-aware AWQ resmoothing with optional requantization during export.
  * Improved handling for shared-input / expert groups and tensor-parallel sharding of pre-quantization scales.

* **Bug Fixes**
  * Removed AWQ reload limitation from known issues; improved checkpoint validation and safer save/load behavior.
  * Better detection and handling of enabled weight-quantizers and clearer warnings for mismatched checkpoint keys.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-04-19 22:12:45 -07:00
jingyu-ml 26ae8da517 [2/3] Implicit Gemm NVFP4 (#1227)
### What does this PR do?

Type of change: new feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

- Add Conv3D implicit GEMM kernel with BF16 WMMA tensor cores and fused
NVFP4 activation quantization for video diffusion VAE layers
- Integrate into _QuantConv3d via QuantModuleRegistry — automatically
dispatched when NVFP4 quantization is applied to nn.Conv3d
- Move kernel from `experimental/conv/ to modelopt/torch/kernels/conv/`;
move tests to `tests/gpu/torch/quantization/kernels/`

### Testing
<!-- Mention how have you tested your change if applicable. -->

- Added test cases to measure the difference between cuDNN and our CUDA
implicit GEMM kernel
- Added an NVFP4 fake quantization test using CUDA code

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Per-backbone quantization/export in a single run with per-backbone
checkpoints and backbone-aware quant filters
* Configurable NVFP4 block-size via CLI/config; improved NVFP4 Conv3D
inference path and Wan 2.2 quantization support
* **Bug Fixes**
* Video-model calibration now respects extra params and forces video
decoding during calibration
* **Documentation**
* Added comprehensive Conv3D implicit‑GEMM kernel documentation; removed
experimental Conv3D prototype docs/benchmark
* **Tests**
* New Wan 2.2 quantization/export tests and expanded Conv3D/FP4 kernel
test coverage
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-04-19 12:20:14 +05:30
760c980727 Add ResNet50 support for torch_onnx quantization workflow (#1263)
## Summary
- Add end-to-end ResNet50 support in the torch_onnx quantization → ONNX
export → TRT engine pipeline
- Fix multiple Conv2d-related export issues that blocked Conv2d-heavy
models from working with FP8/INT8/MXFP8/NVFP4/auto quantization modes
- Fix `configure_linear_module_onnx_quantizers` to handle all modules
with block quantization (not just `nn.Linear`), fixing NVFP4/MXFP8
export for models with quantized non-Linear modules
- Add `--trt_build` flag to `torch_quant_to_onnx.py` and simplify test
infrastructure

### Files Changed
- `modelopt/torch/_deploy/utils/torch_onnx.py` — Disable FP8 Conv2d
weight quantizers and autocast during ONNX export
- `modelopt/torch/quantization/export_onnx.py` — Fix
`configure_linear_module_onnx_quantizers` for all module types with
block quantization
- `examples/torch_onnx/torch_quant_to_onnx.py` — Add `--trt_build` flag,
calibration for FP8 override quantizers, Conv2d→FP8 override for auto
mode, filter_func updates
- `examples/torch_onnx/README.md` — Add ResNet50 to supported models
table
- `tests/examples/torch_onnx/test_torch_quant_to_onnx.py` — Add ResNet50
test entry, simplify using `--trt_build`
- `tests/_test_utils/torch/vision_models.py` — Add ResNet50 to timm
model registry

### Quantization modes passing
- ✅ FP8, INT8, MXFP8, NVFP4, Auto (all 5 modes pass export + TRT build)
- INT4_AWQ excluded (pre-existing limitation for all models)

## Test plan
- [x] All 5 resnet50 test modes pass: `pytest
tests/examples/torch_onnx/test_torch_quant_to_onnx.py -k resnet50` (5/5
passed)
- [x] Full regression: 18 passed, 2 failed (pre-existing swinv2_tiny
fp8/int8 failures)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added ResNet50 to supported ONNX export vision models with FP8, INT8,
MXFP8, and NVFP4 support.
  * Optional TensorRT engine build after export via a new CLI flag.

* **Improvements**
* Enhanced quantization calibration and export flows for FP8/INT8
models, including broader block-quantization support across module types
and safer export handling.
  * Tests updated to include ResNet50 in the model matrix.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Signed-off-by: ajrasane <arasane@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-18 14:30:45 +00:00
bkartal-dev 92622a9aa6 Add nvfp4_mse and nvfp4_local_hessian options to the ptq script. (#1113)
### What does this PR do?

Type of change: Bugfix

<!-- Details about the change. -->

Add newly added quant configs to the example PTQ script.

### Testing

I have locally run auto_quantize with these two quant_configs, and
obtained successfully exported HF artifacts.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added support for three new quantization formats: nvfp4_mse,
nvfp4_local_hessian, and nvfp4_experts_only, expanding available export
options when using auto-quantize.

* **Bug Fixes / UX**
* Updated the invalid-quantization error message to include the newly
accepted format identifiers.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Bilal Kartal <bkartal@nvidia.com>
Signed-off-by: bkartal-dev <bkartal@nvidia.com>
2026-04-18 07:28:49 +00:00
jingyu-ml feec81ad2b Add the Skip softmax for diffusion (#1166)
### What does this PR do?

Type of change: new feature, new example <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->

<!-- Details about the change. -->

## Summary

- Add skip-softmax sparse attention (BLASST) for diffusion models via
dedicated Triton kernels — an inference kernel with tile skipping and a
calibration kernel with vectorized multi-threshold sparsity measurement
- Add `triton_skip_softmax` method with exponential model calibration
(`scale_factor = a * exp(b * sparsity)`) and log-space fitting for
diffusion models
- Add Triton kernel backends for diffusers and LTX attention dispatch
- Fix calibration to skip RULER dataset generation when user provides
their own `forward_loop` (required for non-LLM models)

## Changes

### Triton kernels (`modelopt/torch/kernels/triton_fa.py`)
- **`_attn_fwd`**: Forward kernel with optional tile skipping — tiles
whose max attention score is far below the running softmax max are
skipped entirely (no V load, no softmax, no accumulation). Runtime
sparsity measurement via atomic counters.
- **`_attn_fwd_calibrate`**: Calibration kernel that computes full
attention while measuring how many tiles would be skipped at each of N
thresholds simultaneously. Uses per-program output buffers (zero atomic
contention) and vectorized multi-threshold comparison.
- **`attention()`** / **`attention_calibrate()`**: Python wrappers for
inference and calibration kernels.

### Kernel backends
(`modelopt/torch/sparsity/attention_sparsity/kernels/`)
- **`diffusers_triton_attention.py`**: Registers `modelopt_triton`
backend in diffusers' attention dispatch. Handles [B, S, H, D] → varlen
layout conversion, calibration/inference mode switching, thread-local
configuration, and counter accumulation.
- **`ltx_triton_attention.py`**: Patches `ltx_core.Attention` modules
for Triton dispatch with the same calibration/inference modes.

### Method
(`modelopt/torch/sparsity/attention_sparsity/methods/triton_skip_softmax.py`)
- `TritonSkipSoftmaxMethod`: Context managers for calibration (→
calibration kernel) and inference (→ forward kernel with tile skipping).
Three threshold priority levels: raw threshold > calibrated scale_factor
> static threshold.

### Calibration
(`modelopt/torch/sparsity/attention_sparsity/calibration/`)
- **`calibrator.py`**: `DynamicThresholdCalibrator` with `fit_logspace`
option — fits exponential model in log space (minimizes relative error)
for diffusion models where scale_factors span many orders of magnitude.
Records observed sparsity range for extrapolation warnings.
- **`calibrate.py`**: Skips RULER dataset when `forward_loop` is
provided; passes `fit_logspace` through from config.

### Config & conversion
- **`config.py`**: `CalibrationConfig.fit_logspace` field (default
False, recommended True for diffusion models).
`skip_softmax_raw_threshold` field for direct threshold mode.
- **`conversion.py`**: Auto-registers diffusers/LTX Triton backends on
`sparsify()`. Updated summary display.

### Example
- **`wan22_skip_softmax.py`**: End-to-end example for WAN 2.2 5B/14B
with baseline, raw-threshold, and calibrated modes. Supports runtime
sparsity reporting.

## Threshold modes

| Mode | How it works | Use case |
|------|-------------|----------|
| **Raw threshold** (`--raw-threshold -0.7`) | Passed directly to kernel
as `skip_threshold_log2` | Quick testing, sweeps |
| **Calibrated** (`--calibrate --target-sparsity 0.5`) | `scale_factor =
a * exp(b * target)`, then `threshold = scale_factor / seq_k` at runtime
| Production use with seqlen adaptation |
| **Static** (default `skip_softmax_threshold=0.1`) | `log2(lambda) *
sm_scale` | Fallback |

## Usage

```bash
# Fixed raw threshold (no calibration)
python examples/diffusers/sparsity/wan22_skip_softmax.py \
    --model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
    --raw-threshold -0.7 \
    --prompt "A cat playing piano" --output out.mp4

# With calibration (log-space fit for diffusion models)
python examples/diffusers/sparsity/wan22_skip_softmax.py \
    --model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
    --calibrate --target-sparsity 0.5 \
    --prompt "A cat playing piano" --output out.mp4

# Dense baseline for comparison
python examples/diffusers/sparsity/wan22_skip_softmax.py \
    --model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
    --baseline \
    --prompt "A cat playing piano" --output baseline.mp4
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added skip-softmax sparse attention support for Diffusers models,
enabling efficient video generation
* Added support for both eager and Triton attention backends for sparse
attention
* Added new example script for Wan 2.2 text-to-video generation with
sparse attention optimization

* **Documentation**
* Updated documentation with sparse attention configuration guide and
usage examples

* **Tests**
* Added comprehensive unit tests for kernel backend registration and
skip-softmax functionality
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-04-18 06:54:19 +00:00
realAsma 2d868d3f1f Performant layerwise calibration for large models (#1251)
## Summary

Adds **performant layerwise calibration** for quantizing large models
(e.g. DeepSeek-R1 671B) that don't fit entirely on GPU. ([Example
commands](#example-commands))

1. **Performant calibration for large models** — Each decoder layer is
moved from CPU/disk to GPU (accelerate) or unsharded (FSDP2) **only
once** and kept on GPU for the entire calibration step. Previously,
every calibration batch triggered weight transfer for every layer —
O(num_batches) weight movements per layer. Now it is O(1) per layer.
This also means you can **increase batch size** since only one layer's
weights occupy GPU at a time — e.g. DeepSeek-R1 on a single node
(8×80GB) with `batch_size=16` and `gpu_max_mem_percentage=0.5`.
2. **Checkpoint save/resume** — Saves progress after each layer, so jobs
that exceed cluster time limits (e.g. 4-hour Slurm windows for 100+
layer MoE models) can resume from the last completed layer.
3. **Rename** `sequential_calibrate` → `layerwise_calibrate` for
clarity.

### Design details

The existing layerwise state machine (skip/run/capture) already
processes one layer at a time, but skip-mode layers still kept their
parameters in the ModuleList — so frameworks transferred all weights
every forward pass. This PR adds:
- **`_SkipLayer`**: replaces fully-calibrated layers with a
parameter-free dummy in the ModuleList, so framework hooks have nothing
to transfer
- **`persistent_materialization`**: keeps the active layer on GPU for
the entire calibration step, avoiding repeated offload/reload cycles

Checkpoint save is per-layer; restore is bulk — quantizer state and
weights for layers 0..K-1 are restored once at the end of calibration,
keeping the hot path fast.

### Example commands

**Qwen3-8B** (NVFP4+GPTQ, single GPU):
```bash
python hf_ptq.py \
    --pyt_ckpt_path Qwen/Qwen3-8B \
    --recipe nvfp4_gptq_sequential.yaml \
    --calib_size 64 \
    --batch_size 16 \
    --dataset cnn_dailymail \
    --export_path outputs/qwen3_8b_nvfp4_gptq_seq \
    --gpu_max_mem_percentage 0.5 \
    --use_seq_device_map \
    --vllm_fakequant_export
```

**DeepSeek-R1** (NVFP4 experts-only + FP8 KV, 8×80GB):
```bash
python hf_ptq.py \
    --model unsloth/DeepSeek-R1-0528-BF16 \
    --recipe ../../modelopt_recipes/general/ptq/nvfp4_experts_only-fp8_kv.yaml \
    --dataset cnn_dailymail \
    --batch_size 16 \
    --calib_size 64 \
    --calib_seq 512 \
    --gpu_max_mem_percentage 0.5 \
    --use_seq_device_map \
    --trust_remote_code \
    --export_path output/DeepSeek-R1-BF16-nvfp4-experts-only-fp8-kv \
    --vllm_fakequant_export
```

### Example: NVFP4+GPTQ layerwise calibration on Qwen3-8B (36 layers,
single GPU — 20 GB peak)

**Initial run** (killed after layer 11):
```
Layerwise calibration: Found 36 transformer layers
Calibrating layer 1/36 | capture: [1]
Computing Hessians for 7 linear layers...
GPTQ time: 51.39s
Calibrating layer 2/36 | run: [1] | capture: [2]
Checkpoint: saved layer 0
GPTQ time: 50.06s
Calibrating layer 3/36 | skip: 1 | run: [2] | capture: [3]
Checkpoint: saved layer 1
...
Calibrating layer 12/36 | skip: 10 | run: [11] | capture: [12]
Checkpoint: saved layer 10
<killed>
```

**Resumed run** (picks up from layer 11, finishes all 36):
```
Layerwise calibration: Found 36 transformer layers
Checkpoint: resuming layerwise calibration from layer 11/36
Calibrating layer 12 (resumed)
GPTQ time: 51.45s
Calibrating layer 13/36 | skip: 11 | run: [12] | capture: [13]
Checkpoint: saved layer 11
...
Calibrating layer 36/36 | skip: 34 | run: [35] | capture: [36]
Checkpoint: saved layer 34
GPTQ time: 50.33s
Checkpoint: saved layer 35 (final)
Checkpoint: restored 11 previously calibrated layers
Layerwise calibration completed
Quantized model exported to: outputs/qwen3_8b_nvfp4_gptq_seq
GPU 0: Peak memory usage = 20.42 GB
```

## TODO
- [ ] Update CHANGELOG

## Test plan
- `tests/unit/torch/quantization/test_layerwise_calibrate.py` — unit
tests for skip/swap/restore
- `tests/unit/torch/quantization/test_sequential_checkpoint.py` —
checkpoint save/resume correctness
- `tests/gpu/torch/quantization/plugins/test_accelerate_gpu.py` —
CPU-offloaded layerwise + GPTQ + checkpoint resume
- `tests/gpu/torch/quantization/test_fsdp2.py` — FSDP2 layerwise
calibration

### Verified
- [x] Qwen3-8B: layerwise calibration + checkpoint save/restore +
fakequantized checkpoint export + vLLM serve
- [x] DeepSeek-R1: checkpoint resume tested
- [x] DeepSeek-R1: fakequantized checkpoint export verified

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-04-18 00:32:34 +00:00
sugunav14 dc7ad66b71 GPTQ vector (#1223)
### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added backend-specific GPTQ helper registration to allow
backend-tailored GPTQ behavior.

* **Bug Fixes**
* Prevented KV-cache state from leaking across repeated per-layer
forwards during calibration.

* **Tests**
* Added GPU-focused tests validating GPTQ combined with vector
quantization, including accuracy and end-to-end comparisons.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com>
2026-04-17 15:54:39 -07:00
Keval MorabiaandClaude Sonnet 4.6 e4b054bf32 Fix and Speedup megatron_mmlu by >10x via prefill scoring and global batching (#1280)
### What does this PR do?

Type of change: new feature + bug fix

Two improvements to Megatron inference utilities:

**1. Pipeline Parallel (PP) correctness fixes**

PP inference was producing garbage output (MMLU ~0.24, random chance).
Two root causes:

- `megatron_generate` / `megatron_prefill` used
`get_forward_backward_func()` (the training pipeline scheduler), which
is not designed for inference. Rewrote both functions to use explicit
P2P communication via `recv_from_prev_pipeline_rank_` /
`send_to_next_pipeline_rank`, matching the `run_mcore_inference`
pattern.
- `import_mcore_gpt_from_hf` loads HF weights into stage 0's embedding
but never updates the output_layer on the last PP stage when
`share_embeddings_and_output_weights=True`. At model init,
`setup_embeddings_and_output_layer()` all-reduces from stage 0 to sync
the output layer; after importing HF weights that all-reduce is stale.
Fix: call `model.setup_embeddings_and_output_layer()` again after
import.

**2. `megatron_mmlu` speedup (~6x)**

Replaces the `megatron_mmlu` implementation with a significantly faster
approach that matches how `lm-evaluation-harness` scores multiple-choice
questions.

**Before:** autoregressive generation (`megatron_generate`, `osl=2`) per
example, 114 separate `load_dataset` calls, batch_size=1 — 260s for 5%
data.

**After:** single prefill forward pass + argmax over {A,B,C,D} logits, 2
`load_dataset` calls, configurable batch_size — 18s for 5% data (~6x
faster).

### Changes

**PP fixes:**
- `megatron_generate` / `megatron_prefill`: replace
`get_forward_backward_func` with explicit P2P
(`recv_from_prev_pipeline_rank_` / `send_to_next_pipeline_rank`)
- `import_mcore_gpt_from_hf`: call
`model.setup_embeddings_and_output_layer()` after HF weight import when
PP>1 and `share_embeddings_and_output_weights=True`
- `megatron_prefill`: add `skip_return_logits` param and VLM support
(needed for PP non-last stages)

**MMLU speedup:**
- **Log-likelihood scoring**: replace `megatron_generate` with
`megatron_prefill` — one forward pass per batch, no autoregressive
decode loop
- **Global batching**: collect all examples across all subjects, sort by
descending sequence length, run in `batch_size` chunks
- **2 dataset loads** instead of 114: use `load_dataset("cais/mmlu",
"all")` with per-subject grouping; skip dev load when `few_shots=0`
- **`percentage` → `fraction`** parameter rename for clarity
- **tqdm progress bar** (rank-0 only)

### Testing

- `test_megatron_generate_and_mmlu` parametrized over `tp` and `pp`.
Accuracy assertion: `0.36 < score < 0.39`. Manually checked generated
text is coherent.
- Re-ran M-Bridge Minitron MMLU based pruning for Nano v2 9B -> 7B and
all top 10 candidate's MMLU numbers are ballpark similar as before

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — `percentage` parameter
renamed to `fraction`; `enable_kv_cache` removed from `megatron_mmlu`
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — existing test updated and
parametrized for TP+PP
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

🤖 Generated with [Claude Code](https://claude.ai/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved pipeline-parallel generation and MMLU evaluation reliability;
fixed output-layer synchronization in shared-embedding + pipeline
setups.

* **New Features**
* MMLU scoring now uses batched prefill logit scoring for faster,
batched evaluation.

* **Behavior Changes**
* Default MMLU sampling increased from 5% to 10%; calibration batch
sizing adjusted and related CLI/help text updated.

* **Tests**
* Distributed tests cover tensor- and pipeline-parallel modes and
tighten MMLU validation ranges.

* **Documentation**
* Updated pruning example and benchmark timing to reflect new sampling
and speedup.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-17 19:20:54 +00:00
Hrishith Thadicherla fe8c5178c7 Removed version fixes for torch transformers in windows ptq example requirements (#1275)
### What does this PR do?

Type of change: Bug fix

Removed version fixes for torch and transformers



### Testing
Tested quantization with a couple of models . Working as expected.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Relaxed dependency specs: removed strict pin for torch to allow latest
compatible installs, and constrained transformers to <5.0.0 for broader
compatibility and easier updates.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com>
Signed-off-by: Hrishith Thadicherla <99313418+hthadicherla@users.noreply.github.com>
2026-04-17 09:16:48 +05:30
kinjalpatel27 f238d93d86 vLLM fakequant fold weight_quantizer for megatron export (#1246)
### What does this PR do?

Type of change: Bug fix
During Megatron→vLLM fakequant export
(`export_mcore_gpt_to_hf_vllm_fq`), the `weight_quantizer` is now
applied as fake-quantization (quantize + dequantize) directly into the
exported weight tensor, and its amax is no longer saved to
`quantizer_state.pth`. On reload, if `weight_quantizer` keys are absent
from the checkpoint (because they were folded at export time), the
corresponding quantizer modules are disabled.
This change is useful especially when amax across experts are not synced
for `weight_quantizer`, this allows the `weight_quantizer` to keep them
different for better accuracy.

### Usage
```python                                                                                                                                                                                                    
# Unchanged — export API is the same                                                                                                                                                                         
export_mcore_gpt_to_hf_vllm_fq(model, pretrained_model_name_or_path=..., export_dir=...)
```
 
### Testing
Step 1 — Quantize (run from Megatron-LM
`examples/post_training/modelopt`):
  ```bash
HF_MODEL_CKPT=<path/to/hf/weights> MLM_MODEL_SAVE=<quant-ckpt-name> \
bash quantize.sh <hf-model-id> NVFP4_DEFAULT_CFG
```  

Step 2 — Export for vLLM fakequant:                                                                                                                                                                          
```bash  
MLM_EXTRA_ARGS=--export-vllm-fq \ 
HF_MODEL_CKPT=<path/to/hf/weights> \ 
MLM_MODEL_CKPT=<quant-ckpt-name> \ 
EXPORT_DIR=<export-dir> \ 
bash export.sh <hf-model-id> 
```

Step 3 — Serve (run from examples/vllm_serve):                                                                                                                                                               
```bash
 QUANT_CFG=NVFP4_DEFAULT_CFG \ 
 QUANT_FILE_PATH=<export-dir>/quantizer_state.pth \ 
 python3 vllm_serve_fakequant.py <export-dir> \ 
 -tp 1 --served-model-name <model-name> \ 
 --host 0.0.0.0 --port 8000 \ 
--trust-remote-code --enforce-eager \ 
 --disable-custom-all-reduce \ 
--gpu-memory-utilization 0.8          
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A 
- Did you write any new necessary tests?: N/A 
- Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A 

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit

* **Bug Fixes**
  * Better handling when loading checkpoints: missing weight-quantizer entries are validated and corresponding modules are disabled to avoid load failures.

* **Improvements**
  * Export now folds enabled weight quantizers into exported weights when present and omits internal weight-quantizer tensors from the exported state to produce cleaner exports.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-04-16 09:42:36 -07:00
yeyu-nvidiaandClaude Opus 4.6 07ae8e7128 Add LoRA co-training support for HF EAGLE speculative decoding (#1060)
### What does this PR do?

Type of change: New feature + bug fixes

Adds **LoRA co-training** support for HF EAGLE speculative decoding.
When `eagle_base_lora=True`, HF PEFT LoRA adapters are injected into the
base model and co-trained alongside the EAGLE draft module in a single
online training pass. A preservation loss (KL divergence between the
original frozen base model output and the LoRA-adapted output) prevents
base model drift. LoRA adapter weights are exported in standard peft
format alongside EAGLE draft artifacts.

### Key features

- **LoRA injection**: `peft.inject_adapter_in_model` applied in-place
(no wrapper), keeping the existing `HFEagleModel` structure intact.
- **Preservation loss**: Cross-entropy `H(ref, lora)` — equivalent
gradient to `KL(ref || lora)` since `H(ref)` is constant w.r.t. LoRA
params.
- **Warmup schedule**: `eagle_base_lora_warmup_steps` freezes LoRA for N
steps while the EAGLE head stabilizes, then enables co-training via a
`LoRAWarmupCallback`.
- **Logits detach regularization**: `eagle_base_lora_logits_detach_prob`
stochastically detaches base logits from the EAGLE loss path, preventing
LoRA from degenerating to maximize EAGLE accuracy at the cost of base
model quality.
- **Export**: Standard peft format (`adapter_model.safetensors` +
`adapter_config.json`) alongside EAGLE draft model.
- **Merge script**: `scripts/merge_lora.py` merges LoRA weights into the
base model and restores the original `config.json` (avoids transformers
5.x rewriting `rope_theta` → `rope_parameters` which breaks
vLLM/TRT-LLM).
- **Multinode fix**: `dp_shard_size` now uses `WORLD_SIZE` instead of
local GPU count.

### Config options

```python
mtsp.convert(model, mode=[("eagle", {
    "eagle_base_lora": True,                          # enable LoRA co-training
    "eagle_base_lora_rank": 64,                       # LoRA rank
    "eagle_base_lora_alpha": 16.0,                    # LoRA scaling
    "eagle_base_lora_target_modules": ["q_proj", "k_proj", "v_proj", "o_proj"],
    "eagle_base_lora_preservation_loss_weight": 0.1,  # preservation loss weight
    "eagle_base_lora_warmup_steps": 0,                # freeze LoRA for N steps
    "eagle_base_lora_logits_detach_prob": 0.5,        # detach prob (0=never, 1=always)
})])
```

### Experimental results (Qwen3-8B, checkpoint-60000)

Base model quality preserved across detach_prob sweep (lm_eval: IFEval,
ARC-C, Winogrande — results pending final collection).

**Acceptance rate** (mt_bench, draft_length=3, output_length=4096,
temperature=0):

| detach_prob | vLLM AR | TRT-LLM AR |
|---|---|---|
| baseline (no LoRA) | 2.14 | 2.15 |
| 0.5 | 1.45 | 1.44 |
| 0.8 | **3.06** | **3.01** |
| 0.85 | 2.90 | 2.90 |
| 0.9 | 2.76 | 2.77 |
| 0.95 | 2.51 | 2.58 |
| 0.99 | 2.37 | 2.37 |
| 0.999 | 2.30 | 2.27 |
| 0.9999 | 2.31 | 2.26 |

Best AR at `detach_prob=0.8`: ~40% improvement over baseline.

### Testing

`tests/unit/torch/speculative/plugins/test_hf_speculative_lora.py` (5
tests):
- `test_lora_layers_injected` — LoRA layers present after conversion
- `test_trainable_params` — only `lora_*` and `eagle_module` params are
trainable
- `test_forward_returns_loss` — forward returns non-zero scalar loss
- `test_eagle_offline_incompatible` — `eagle_base_lora=True` +
`eagle_offline=True` raises `ValueError`
- `test_export_lora_artifacts` — export produces standard peft adapter
files

### Bug fixes (included in this PR)

1. **`launch_train.sh` case pattern ordering**: glob
`--eagle_base_lora*` was before specific patterns
(`--eagle_base_lora_rank*`, etc.), silently swallowing LoRA args.
2. **LoRA optimizer exclusion during warmup**: warmup freezing excluded
LoRA from the optimizer entirely; fixed with `add_param_group` in the
callback.
3. **`merge_lora.py` config.json**: `save_pretrained()` with
transformers >=5.x rewrites `rope_theta` → `rope_parameters`, breaking
vLLM positional embeddings. Fixed by copying the original base model
config.
4. **Multinode `dp_shard_size`**: used local GPU count instead of
`WORLD_SIZE`.

### Checklist

- [x] Backward compatible (all new config fields have defaults)
- [x] Uses `peft` via lazy imports (no hard dependency)
- [x] Unit tests added
- [x] Online HF training only (`eagle_offline=True` blocked)

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-16 01:35:15 +00:00
361f7e391b Merge puzzletron compression algorithm (#1121)
### What does this PR do?

Implement puzzletron compression algorithm based on Puzzle paper
(https://arxiv.org/abs/2411.19146)

<details>
<summary> Th list of reviewed and merged MRs that resulted in the
feature/puzzletron branch</summary>

Merging dkorzekwa/any_model to feature/puzzletron

[Add anymodel directories to feature/puzzletron by danielkorzekwa · Pull
Request #974 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/974)
- merged

[Draft: anymodel activation scoring by danielkorzekwa · Pull Request
#989 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/989)
- merged

[Draft: Merge anymodel pruning by danielkorzekwa · Pull Request #990 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/990/)
- merged

[Draft: Merging anymodel:build_library_and_stats by danielkorzekwa ·
Pull Request #993 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/993)
- merged

[Dkorzekwa/any model calc one block scores by danielkorzekwa · Pull
Request #994 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/994)
- merged

[Draft: merge any_model: mip_and_realize_models by danielkorzekwa · Pull
Request #995 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/995)
- merged

[Dkorzekwa/any model other modeqls by danielkorztiekwa · Pull Request
#1007 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1007/)
- merged

PR to 1007: https://github.com/NVIDIA/Model-Optimizer/pull/1039 - merged

[Dkorzekwa/anymodel gptoss by danielkorzekwa · Pull Request #1020 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1020)
- merged

[Merge any_model tutorial by danielkorzekwa · Pull Request #1035 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1035)
- merged

[Merge mbridge distillation for any_model by danielkorzekwa · Pull
Request #1036 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1036)
- merged

[MR branch for the remaining difference between dkorzekwa/any_model an…
by danielkorzekwa · Pull Request #1047 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1047)
- merged

[Dkorzekwa/decilm hf code cleanup by danielkorzekwa · Pull Request #1071
·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1071)
- merged

[Dkorzekwa/decilm hf code cleanup 2 by danielkorzekwa · Pull Request
#1073 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1073)
- merged

[Dkorzekwa/anymodel subblock stats by danielkorzekwa · Pull Request
#1085 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1085)
- merged

[Dkorzekwa/anymodel subblock stats nodecilm by danielkorzekwa · Pull
Request #1102 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1102)
- merged

[Dkorzekwa/decilm cleanup post subblockstats by danielkorzekwa · Pull
Request #1103 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1103)
- merged

[code clean up by danielkorzekwa · Pull Request #1110 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1110)
- merged

Merging into main:

[Activation hooks redesign (reuse hooks component across both minitron
and puzzletron) by danielkorzekwa · Pull Request #1022 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1022)
- merged

[Dkorzekwa/puzzletron use importance hooks from prune by danielkorzekwa
· Pull Request #1115 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1115)
- merged

</details>

<!-- Details about the change. -->

### Usage

Puzzletron tutorial:

https://github.com/NVIDIA/Model-Optimizer/tree/feature/puzzletron/examples/puzzletron

### Testing
The main e2e test for compressing 9 models with Puzzletron:

https://github.com/NVIDIA/Model-Optimizer/blob/feature/puzzletron/tests/gpu/torch/puzzletron/test_puzzletron.py

2-gpu nightly tests: 

-
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24468209205/job/71501061203
-
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24470214159/job/71508152952

### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added Puzzletron: end-to-end heterogeneous pruning & NAS workflow with
AnyModel support, example pipelines, deployment and evaluation
utilities, and tools for converting/pruning and exporting compressed
checkpoints.

* **Documentation**
* Comprehensive Puzzletron tutorials, model-specific guides, evaluator
instructions, example configs, and changelog entry.

* **Chores**
* CI/workflow updates (extras installation, longer GPU test timeout),
pre-commit hook exclusion updated, and CODEOWNERS entries added.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com>
Signed-off-by: Liana Mikaelyan <lmikaelyan@nvidia.com>
Signed-off-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com>
Signed-off-by: Daniel Korzekwa <daniel.korzekwa@gmail.com>
Signed-off-by: jrausch <jrausch@nvidia.com>
Signed-off-by: root <root@pool0-00848.cm.cluster>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Liana Mikaelyan <lmikaelyan@nvidia.com>
Co-authored-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com>
Co-authored-by: J Rausch <38429553+j-rausch@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-16 00:48:11 +05:30
Gwena Cunhaanddmoodie dec2952992 [6034518] Downgrade TRT support for remote autotuning in Autotune from 10.16 to 10.15 (#1259)
### What does this PR do?

Type of change: Bug fix

Remote autotuning is supported in TensorRT from version 10.15, but fails
with Autotune as it's checking for 10.16+. This PR fixes that check and
updates documentation accordingly.

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
See bug 6034518.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!---
Mandatory -->
- Did you write any new necessary tests?: N/A <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added a Remote Autotuning guide for TensorRT 10.15+ with CLI examples;
updated examples to require `--safe --skipInference`.

* **Updates**
* Lowered TensorRT minimum requirement for remote autotuning from 10.16
to 10.15.
  * Clarified CLI help text for trtexec/autotune arguments.

* **Bug Fixes**
* trtexec-based autotuning now verifies the trtexec executable version
when checking compatibility.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
Signed-off-by: dmoodie <dmoodie@nvidia.com>
Co-authored-by: dmoodie <dmoodie@nvidia.com>
2026-04-15 14:31:41 -04:00
kinjalpatel27 1619421383 Added support for MoE for vllm >= 0.14.0rc1 (#1162)
### What does this PR do?

Type of change: Bug fix
`_QuantFusedMoEBase.forward()` previously replaced
`vllm_fused_moe_package.invoke_fused_moe_kernel`, which was replaced
starting in vLLM v0.14.0rc1,

There are two paths for FusedMoE forward:
```
Path 1 (Modular — standard CUDA path):                                                                                                                                                                       
    FusedMoE.forward()                                                                                                                                                                                         
      → self.runner.forward()                                                                                                                                                                                  
        → TritonExperts.apply()                                                                                                                                                                                
          → invoke_fused_moe_triton_kernel()  ← called twice (w1, w2)                                                                                                                                          
                                                                                                                                                                                                               
  Path 2 (legacy):                      
    inplace_fused_experts / outplace_fused_experts                                                                                                                                                             
      → fused_experts_impl()                       
        → dispatch_fused_moe_kernel()              
          → invoke_fused_moe_triton_kernel()                                                           
            or invoke_fused_moe_wna16_triton_kernel()                                                  
            or invoke_fused_moe_wna16_cuda_kernel()
```
This caused an `AttributeError` / assertion failure for any MoE model
quantized with vLLM ≥ v0.14.0rc1.

The fix refactors the kernel-patching logic into a `_patch_moe_kernel()`
context manager that probes for both attribute names (the two names are
mutually exclusive across vLLM versions — confirmed by inspecting every
release from v0.10.0 to v0.19.1).
    
### Usage

NA

### Testing
```
docker run --gpus all -it --shm-size=160GB --network host --rm -v <modelopt path>:/home/modelopt \
vllm/vllm-openai:v0.15.0 bash -c "cd /home/modelopt && pip install . && pip install datasets && \
  QUANT_CFG=NVFP4_DEFAULT_CFG python3 /home/modelopt/examples/vllm_serve/vllm_serve_fakequant.py \
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 -tp 1 --served-model-name NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 \ 
  --host 0.0.0.0 --port 8001 --trust-remote-code --disable-custom-all-reduce \
--gpu-memory-utilization 0.8" 
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?:  N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Refactor**
* Ensures quantized expert weights are correctly used by the fused-MoE
execution path so inference uses the intended quantized tensors.
* Replaces fragile manual swapping of the runtime kernel with a safer,
context-managed swap that reliably caches and restores the original.
* Adds runtime detection and selection among available fused-MoE kernel
entrypoints to support multiple variants.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-04-14 16:40:35 -07:00
Chenhan D. YuandClaude Opus 4.6 3131195241 add: DFlash block diffusion speculative decoding (#1211)
DFlash (Block Diffusion for Flash Speculative Decoding) predicts an
entire block of tokens in a single forward pass using masked parallel
prediction with KV injection from the target model's hidden states.

Key features:
- Feature fusion (multi-layer hidden states -> FC + RMSNorm)
- KV injection (fused features as K/V in every draft layer with QK-norm)
- Random anchor sampling with bidirectional intra-block attention
- Logit distillation with exponential loss decay (gamma weighting)
- Multi-node DDP training with checkpoint resume
- Export to z-lab compatible HF format
- Online validation (context-dependent ground truth)

Training recipe:
modelopt_recipes/general/speculative_decoding/dflash.yaml
Results: examples/speculative_decoding/doc/dflash_results.md

### ModelOpt Eval (online validation, osl=512)

| Dataset | z-lab | ModelOpt (306K) | Diff |
|---------|-------|-----------------|------|
| gsm8k | 4.10 | **5.19** | **+1.09** |
| MT-Bench | 3.58 | **4.36** | **+0.78** |

### z-lab Official Eval (dflash.benchmark, osl=512)

| Dataset | z-lab | ModelOpt (306K) | Diff |
|---------|-------|-----------------|------|
| gsm8k | **5.00** | 4.08 | -0.92 |
| MT-Bench | **3.28** | 2.99 | -0.29 |

> z-lab model trained with block_size=16. ModelOpt trained with
block_size=8.

## Evaluation Method Impact (gsm8k)

| Eval Method | z-lab checkpoint | ModelOpt (306K) |
|-------------|-----------------|-----------------|
| Fixed GT (ModelOpt eval) | 2.95 | 4.23 |
| Online GT (ModelOpt eval) | 4.10 | **5.19** |
| z-lab official eval | **5.00** | 4.08 |

### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added DFlash speculative decoding mode with parallel block prediction
support.
* Included training launchers and MT-Bench evaluation scripts for DFlash
models.
* Added online acceptance rate validation for improved inference
verification.

* **Documentation**
* DFlash quick start guide with configuration parameters and training
examples.
  * Performance results and benchmarks for DFlash-trained models.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 15:58:39 -07:00
kinjalpatel27 73be81037d vLLM fakequant: add recipe-based quantization support (#1233)
### What does this PR do?

Type of change: example update

This PR adds recipe-based quantization support to the vLLM fakequant
example.


### Testing
```
docker run --gpus all -it --shm-size=160GB --network host --rm --entrypoint bash -v <modelopt>:/home/modelopt vllm/vllm-openai:v0.15.0 -c "cd /home/modelopt && pip install . && pip install datasets && RECIPE_PATH=/home/modelopt/modelopt_recipes/general/ptq/nvfp4_mlp_only-fp8_kv.yml python3 /home/modelopt/examples/vllm_serve/vllm_serve_fakequant.py Qwen/Qwen3-0.6B -tp 1 --served-model-name Qwen3-0.6B --host 0.0.0.0 --port 8001 --trust-remote-code --disable-custom-all-reduce --gpu-memory-utilization 0.8"
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added `RECIPE_PATH` environment variable support enabling users to
specify ModelOpt PTQ recipe YAML files for quantization configuration in
vLLM serving.

* **Documentation**
* Updated examples and documentation to support recipe-driven
quantization configuration, aligning export workflow with recipe-based
setup.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-04-14 13:37:02 -07:00
Shengliang Xu b6c6ec342c use typed quantize config instead of a raw dict (#1249)
### What does this PR do?

But fix:

Use typed QuantizeConfig instead using raw dict for formal typed
ModelOpt configs.

The dict typing was accidental.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Quantization recipe configuration is now implemented with a
strongly-typed, structured schema that enforces type safety and provides
enhanced validation with comprehensive error detection capabilities.

* **Tests**
* Updated recipe loading tests to correctly validate quantization
configurations when recipes are loaded from directories, fully
supporting the new structured object-based configuration format.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-04-13 18:24:22 -07:00
Ajinkya RasaneandClaude Opus 4.6 202c3d3894 Add SwinTransformer support for torch_onnx quantization workflow (#1235)
## Summary
- Enable end-to-end quantize → ONNX export → TRT engine pipeline for
SwinTransformer models (v1 and v2) across FP8, INT8, MXFP8, NVFP4, and
auto precision modes
- Add Conv2d quantization overrides for TRT compatibility (TRT only
supports FP8/INT8 for convolutions)
- Fix FP8 LayerNorm type mismatch in TRT stronglyTyped mode by adding
`LayerNormalization` to `change_casts_to_fp16`
- Fix `cast_initializer_to_dtype` crash when node has no initializer
inputs
- Simplify `download_example_onnx.py` to a single `--timm_model_name`
(required) flag, removing redundant `--vit` and `--llama` flags
- Add vision model support matrix to README (ViT, Swin, SwinV2)
- Rewrite tests: parametrize over (ViT, Swin, SwinV2) × (fp8, int8,
mxfp8, nvfp4, auto) with TRT engine build verification

## Test plan
- [ ] `python -m pytest
tests/examples/torch_onnx/test_torch_quant_to_onnx.py -v` — 15 tests (3
models × 5 modes), all pass
- [ ] Verified Swin accuracy on ImageNet-1k across all precisions (FP8:
81.29%, INT8: 81.12%, MXFP8: 81.32%, NVFP4: 80.79%, Auto: 80.84% TRT
top-1 vs 81.37% base)
- [ ] INT4_AWQ deferred (TODO in test file) — requires INT4 exporter
changes for non-MatMul/Gemm consumer patterns

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* ONNX export supports arbitrary timm vision models with auto device
selection and new CLI options (--timm_model_name, --model_kwargs,
--no_pretrained); batch-size/input sizing is now model-generic.

* **Bug Fixes**
  * Expanded FP16/BF16 cast handling to additional ONNX ops.
* Disabled inplace ReLU before auto-quantization to avoid incorrect
transforms.
* Conv2d quantization overrides added for improved TensorRT
compatibility.
  * Safer handling when initializers are missing during dtype casting.

* **Documentation**
* README updated with supported models table, quantization mappings, and
example CLI usage.

* **Tests**
* Tests expanded to multiple architectures/quant modes and now verify
TensorRT engine build.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 04:48:29 +05:30
h-guo18 6403389eb0 Feat: Configurable Eagle ROPE scaling during export (#1238)
### What does this PR do?

JIRA ticket: https://jirasw.nvidia.com/browse/OMNIML-3469

Type of change: New feature

Decouple EAGLE training rope configuration from export rope
configuration, enabling separate YaRN rope scaling injection at export
time for long-context inference.

#### Changes

**Configurable export rope scaling (`EagleConfig`)**
- Add `eagle_export_rope_scaling` field to `EagleConfig` with default
YaRN config (`factor=32.0`, `original_max_position_embeddings=2048`)
- Set to `{}` to disable rope scaling injection at export

**Simplified training defaults (`default_config.py`)**
- Change default training rope from `llama3` (theta=500k) to `default`
(theta=10k) — models now train with simple positional embeddings; rope
scaling is applied only at export
- Add `rope_theta` inside `rope_scaling` dict for transformers 5.x
cross-version compatibility

**Move config validation/rewriting into `EagleConfig` (`config.py`)**
- `_derive_eagle_offline`: derives `eagle_offline` from
`data_args.offline_data_path` via validation context, removing manual
assignment in `main.py`
- `_check_rope_scaling_consistency`: rejects configs where
`eagle_export_rope_scaling` is set but training `rope_type` is not
`"default"`
- `_warn_rope_vs_training_seq_len`: warns when
`original_max_position_embeddings` differs from `training_seq_len`

**Export rope injection (`hf_spec_export.py`)**
- Inject `eagle_export_rope_scaling` into the exported HF config when
training rope_type is `"default"`
- Fall back `rope_theta` from `rope_scaling` dict for transformers 5.x
compatibility

**Fix Megatron RotaryEmbedding crash (`megatron_eagle.py`)**
- `dict_to_config()` set `rope_scaling=True` whenever the `rope_scaling`
key existed, even without a `"factor"` — causing `RotaryEmbedding` to
divide by `None`
- Now only enables `rope_scaling` when the dict actually contains a
`"factor"` key

### Usage

Configure in YAML config (or use defaults from `eagle3.yaml`):
```yaml
eagle:
  eagle_export_rope_scaling:
    rope_type: yarn
    factor: 32.0
    original_max_position_embeddings: 2048
```

Set to empty dict to disable export rope injection:
```yaml
eagle:
  eagle_export_rope_scaling: {}
```

### Testing

- New unit tests: `tests/unit/torch/speculative/test_eagle_config.py` —
rope consistency validator, seq_len warning, context-derived
`eagle_offline`
- New unit tests: `tests/unit/torch/export/test_hf_spec_rope_export.py`
— export rope injection, fallback, and empty-config cases

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ (new field has sensible
default; existing configs work unchanged)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ (should be added if merging as a feature)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Add export-time rope-scaling configuration for EAGLE models.

* **Improvements**
* Stronger validation and context-aware reconciliation between training
and export configs.
  * Export now injects rope-scaling and rope-theta when appropriate.
  * Default rope-scaling values updated for EAGLE variants.
  * Model instances now expose export rope-scaling for downstream use.

* **Tests**
* Added unit tests covering rope-scaling export behavior and
configuration validators.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-04-13 16:06:26 -07:00
Keval Morabia 0b42c143dd Update LICENSE and SPDX-License-Identifier as per OSRB guidance (#1244)
Update LICENSE and SPDX-License-Identifier as per OSRB guidance

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated contribution guidelines with expanded license compliance
instructions and SPDX identifier guidance.
* Extended LICENSE file with new "Third-Party Software Notices" section
documenting Apache 2.0, MIT, and BSD 3-Clause licensed components.

* **Chores**
* Updated SPDX license identifiers across multiple files to reflect dual
and triple licensing (Apache 2.0 with MIT and/or BSD 3-Clause).

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-13 22:32:54 +05:30
Keval MorabiaandClaude Sonnet 4.6 80a77d1cc0 Add LTX-2 third-party license notices for legal compliance (#1226)
## Summary

LTX-2 (`ltx-core`, `ltx-pipelines`, `ltx-trainer`) is a third-party
dependency developed and provided by Lightricks. It is governed by the
[LTX Community License
Agreement](https://github.com/Lightricks/LTX-2/blob/main/LICENSE),
**not** the Apache 2.0 license that covers NVIDIA Model Optimizer. Per
legal guidance, all integration points must clearly surface this to
users.

- Add `[!WARNING]` license notice blocks at the top of all LTX-2-related
READMEs (`examples/diffusers`, `examples/diffusers/distillation`,
`examples/windows/diffusers/qad_example`)
- Add `warnings.warn(UserWarning)` at every LTX package import site in
Python files, covering both top-level and lazy imports:
  - `examples/diffusers/distillation/distillation_trainer.py`
  - `examples/diffusers/quantization/calibration.py`
  - `examples/diffusers/quantization/pipeline_manager.py`
-
`examples/windows/diffusers/qad_example/sample_example_qad_diffusers.py`
  - `modelopt/torch/export/diffusers_utils.py`
  - `modelopt/torch/quantization/plugins/diffusion/ltx2.py`
- Add license notice comment to `requirements.txt` files that list LTX
packages, so the obligation is visible at install time
- Update `.github/CODEOWNERS` so all `requirements*.txt` files (covering
variants like `requirements-dev.txt`) are owned by
`@NVIDIA/modelopt-setup-codeowners` regardless of location, via a
last-match-wins rule

**Design notes:**
- For library files (`diffusers_utils.py`, `ltx2.py`), the warning is
placed at the lazy import site inside functions — it fires only when
LTX-2 code paths are actually invoked, not at module import time, to
avoid polluting non-LTX users
- For example entry-point scripts that are LTX-2-only, the warning fires
at module load time (after all imports, to satisfy ruff E402)

## Test plan

- [ ] Confirm `pre-commit run --all-files` passes (ruff, mypy,
markdownlint, bandit all clean)
- [ ] Verify warning appears at runtime when running an LTX-2
quantization or distillation example
- [ ] Confirm non-LTX code paths (FLUX, SDXL, SD3) do not emit the
warning

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added third-party license notices across documentation and
requirements files clarifying LTX-2 packages are governed by the LTX
Community License Agreement rather than NVIDIA Model Optimizer's Apache
2.0 license.

* **Chores**
  * Updated code ownership configuration for requirements files.
* Added runtime warnings to notify when LTX-2 dependencies are accessed.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-13 11:04:29 +05:30