89 Commits
Author SHA1 Message Date
Daniel Korzekwa fadbf74d31 Make the QAT/QAD guide the central place for concepts, background, and framework selection (#2590)
### What does this PR do?

Make the QAT/QAD guide the central place for concepts, background, and
framework selection. Have the Hugging Face and Megatron Bridge tutorials
link back to it instead of repeating explanations of QAT and QAD,
keeping the tutorials focused on setup and execution. In main QAT/QAD
guide make links to all relevant blogposts.

Note: MBridge example doc is out of scope for this MR.

### Testing
Doc changes only, manual check.

### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ 
- Did you write any new necessary tests?: N/A docs changes only

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Expanded the QAT/QAD guide with workflows, use cases, and a
comparison, including QAD’s use of a frozen BF16 teacher and logit-level
loss to recover accuracy after quantization.
* Updated README and quick-start navigation to link to the combined
QAT/QAD guide; the previous standalone QAT guide now redirects readers
there.
* Reorganized the LLM QAT tutorial: recipe guidance is now part of the
end-to-end example, while trainer examples and Python
quantize-and-fine-tune guidance are in Advanced Topics. The tutorial
also notes Triton accelerated kernels.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com>
2026-10-01 20:56:57 +02:00
Shengliang XuandClaude Opus 5 d0142c9dca Reuse a whole recipe via $import, deprecate recipe_type, and start the published-checkpoint backfill with two aliases (#2376)
### What does this PR do?

**Type of change:** New feature (recipe loading) + one bug fix

Two things, the second built on the first:

1. **A recipe can now reuse another recipe whole.** A top-level
`$import` brings in the imported recipe's entire body; keys given
alongside it override the imported ones. `metadata.recipe_type` becomes
optional and is deprecated along the way.
2. **The deprecated `recipe_type` is swept out of every shipped recipe,
and the checkpoint backfill starts with two published checkpoints
recorded as aliases** that reuse a portable recipe wholesale — the first
users of the alias mechanism — plus a fix to two existing Nemotron NVFP4
recipes.

#### Declaring what kind of recipe a file is

`load_recipe` read `metadata.recipe_type` out of the raw YAML *before*
resolving imports, because it needs the schema class to hand to
`load_config`. That made the field impossible to inherit, so a recipe
reusing another had to restate a line it could only have copied.

It is now optional, and the loader takes the first of these that
answers:

1. a `# modelopt-schema:` comment naming the recipe's schema class,
2. `metadata.recipe_type` — **deprecated**; still read and still
honoured, so a recipe outside this repo keeps working unchanged,
3. the recipe it delegates to via a top-level `$import`.

Whatever a recipe *does* state must be true, in both directions. A
schema comment contradicting a `recipe_type` is rejected, and so is a
recipe importing a different kind of recipe — that used to surface as
whatever pydantic made of, say, an `eagle` section spliced into a PTQ
schema. The concrete recipe classes carry a `RECIPE_TYPE` ClassVar as
the single source of truth.

Only a recipe that another file **imports** needs the schema comment —
that is what `$import` resolution requires to validate the payload. The
sweep here drops `metadata.recipe_type` from all 78 shipped recipes that
carried it and gives the imported ones a `# modelopt-schema:` comment
instead, so nothing in-tree depends on the deprecated field.

A directory recipe's `metadata.yml` resolves its kind the same way —
schema comment first, `recipe_type` as the fallback — it just has no
`$import` to delegate through, since a directory recipe has no body of
its own to hand off. (Follow-up commit, after this PR's initial review:
it originally still required `recipe_type` unconditionally, the one
place the deprecation didn't reach.)

#### Checkpoint aliases

Two checkpoints NVIDIA has published in quantized form use a scheme a
portable recipe already produces, with no checkpoint-specific deviation,
so each is recorded as a thin **alias** (top-level `$import`, overriding
only `metadata`) at its own model-hub path -- the *source* checkpoint's
path, not the published quantized one's:

-
**`models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast`**
delegates to `general/ptq/nvfp4_experts_only_mse-kv_fp8_cast` —
expert-only NVFP4 (MSE static weights, dynamic inputs) with an FP8 KV
cache in cast mode — published as `nvidia/Kimi-K2.6-NVFP4`.
-
**`models/Qwen/Qwen3.5-397B-A17B/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8`**
delegates to the `qwen3_5_moe` architecture recipe
`model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4
(MSE static weights) on the routed experts, ModelOpt-default FP8
elsewhere, FP8 KV cache — published as
`nvidia/Qwen3.5-397B-A17B-NVFP4-V2`.

(Follow-up commit, after this PR's initial review: the Qwen entry
originally lived at `models/nvidia/Qwen3.5-397B-A17B/` -- nvidia is the
*published* checkpoint's org, not Qwen3.5-397B-A17B's own. Moved to
match the source model's actual hub path, same as the Kimi-K2.6 entry
above.)

Editing the base recipe changes every alias that points at it; nothing
is duplicated.

#### One fix

- **The Nemotron-3 Super and Ultra NVFP4 recipes** quantized the MTP
block on the **Megatron-Core** path, where it is a live `model.mtp`
submodule their broad `*mixer.*` patterns matched into, contrary to
their own descriptions. They now disable `mtp.*` explicitly. Hugging
Face runs were unaffected — `NemotronHPreTrainedModel` sets
`_keys_to_ignore_on_load_unexpected = [r"mtp.*"]` and builds no MTP
module.

### Usage

A checkpoint alias resolves through `--recipe` to the recipe it
delegates to:

```bash
python examples/hf_ptq/hf_ptq.py \
    --pyt_ckpt_path <checkpoint> \
    --recipe models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast \
    --export_path <output>
```

A recipe that reuses another whole — the shape the aliases use:

```yaml
imports:
  base: general/ptq/nvfp4_experts_only_mse-kv_fp8_cast

$import: base
metadata:
  description: What this checkpoint uses the base recipe for.
```

### Testing

- **`tests/unit/recipe/test_loader.py`** — 28 new cases covering
whole-recipe reuse with no `metadata` at all; kind resolution from each
of the three sources, from a delegation chain and from a `$import` list;
a delegation cycle failing with `ValueError` rather than recursing;
`peek_declared_schema` including a comment placed below the first YAML
line; `recipe_type` being optional, filled per class, and rejected when
it contradicts; a directory recipe resolving its kind from a schema
comment the same way, rejecting a comment/`recipe_type` disagreement,
and still requiring one or the other; and delegating across kinds being
an error.
- **`tests/unit/recipe/test_recipe_docs.py`** — the
model-specific-recipe check now also covers the two new alias folders,
which must be listed in `ptq.md` like every other
`models/<org>/<model_id>` entry.
- **Recipe validation** (`tools/precommit/check_modelopt_recipes.py`)
and **`pre-commit`** pass on the changed files. The full
`tests/unit/recipe/` suite is left to CI — a broken `transformer_engine`
in the local dev venv keeps the `mtq.quantize`-based cases from running
there.

Not covered: **numerics**. Nothing here asserts accuracy, or that
running one of these recipes reproduces a released checkpoint's weights.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `metadata.recipe_type` is
still read and honoured for recipes outside this repo, the schema
comments are inert for direct loads, and the loader change only relaxes
a check.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — two feature entries, one deprecation, and one bug fix under 0.48.0.
- Did you get Claude approval on this PR?: ❌ — not yet run.



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Recipes can delegate configurations, support checkpoint aliases, and
apply local metadata overrides.
* Recipe types can be inferred from schema declarations or delegated
recipes, with stronger consistency validation.
* Added unquantized KV-cache options, layerwise export, broader operator
calibration, and new PTQ examples.
  * Added checkpoint-specific recipes and MLflow experiment references.

* **Bug Fixes**
* Improved ONNX calibration, FSDP2 export, and fused-MoE quantization
handling.
  * Nemotron-3 recipes keep MTP blocks in BF16.

* **Documentation**
* Expanded guidance for aliases, delegation, schema declarations, and
recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 17:12:31 -07:00
ed5c5ed369 [OMNIML-5899] Export IQ checkpoints from HF and Megatron (#2447)
## Summary

- add IQ format metadata and packed-weight export
- support Hugging Face and TP=1 Megatron export paths
- reject fused-MoE IQ export until a deployment loader owns its packed
layout
- document the shaped `uint8` weight contract and the fused-expert
boundary
- add Hugging Face, Megatron, metadata, and fused-expert export tests

## PR split

This work is split into four focused PRs. Each PR targets `main` and
owns a disjoint file set:

1. **Kernel** — [#2448: Add CUDA kernels for IQ
packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448)
2. **Quantization** — [#2446: Add IQ quantization codecs and
backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446)
3. **Export** — [#2447: Export IQ checkpoints from HF and
Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447)
4. **Recipes** — [#2449: Add IQ post-training quantization
recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449)

The required merge order is #2448, #2446, #2447, then #2449.

## Scope

This PR owns only export code, deployment documentation, and export
tests. It targets `main` and should merge after #2448 and #2446. It does
not contain kernel, codec/backend, or recipe files.

## Deployment consumer boundary

Dense weights and individually named expert weights use the documented
shaped `uint8` contract. Megatron fused-MoE IQ export is intentionally
rejected with `NotImplementedError`: its payload would have shape
`[num_experts, out_features, in_features // 256, payload_bytes]`, and no
deployment loader in this stack currently owns that layout. Support
should be enabled only with a loader integration test.

## Test coverage

- [Hugging Face packed-weight
export](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/unit/torch/export/test_export_weight.py)
- [quantization
metadata](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/unit/torch/export/test_get_quantization.py)
- [Megatron unified export and fused-MoE
rejection](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/gpu_megatron/torch/export/test_unified_export_megatron.py)

## Validation

- all pre-commit hooks pass for the changed files
- 89 focused Hugging Face export, metadata, and fused-expert tests pass
locally
- direct checks cover both fused-MoE export entry points for IQ1_S and
IQ2_XS
- Megatron GPU execution remains delegated to GPU CI
- restricted-term scan passes


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for IQ1_S and IQ2_XS GGML quantization formats in
unified Hugging Face and Megatron exports.
* Added quantization metadata, tensor-shape recovery, packing details,
and IQ2_XS size documentation.
* Added validation for required block sizes and tensor parallelism
settings.

* **Limitations**
  * Fused-MoE and GPT-OSS IQ expert packing are not supported.
* IQ exports require standard `weight` attributes in Hugging Face
models.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 00:00:33 +00:00
didi 76c04dfd99 docs: clarify canonical pruning documentation source (#1871) (#2469)
Fixes #1871

The pruning documentation is split between
`docs/source/guides/3_pruning.rst`
and `examples/pruning/README.md`, causing confusion about which is
authoritative.

The examples/pruning/README.md is the comprehensive, up-to-date
reference
covering Minitron, Puzzletron, FastNAS, support matrix, guidelines, and
distillation hyperparameters.

This PR adds a note to the RST guide making clear:
- The README is canonical for Minitron and Puzzletron (LLM/VLM pruning)
- The guide covers FastNAS for Computer Vision models

Signed-off-by: Diya <diyaismahil7@gmail.com>

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Updated the pruning guide’s introductory content for clearer
separation of general guidance and the related Minitron/Puzzletron note.
- Clarified that the guide focuses on FastNAS pruning for computer
vision models.
- Added references to the Pruning README for Minitron and Puzzletron API
examples, support information, guidelines, and distillation
hyperparameters.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: didi <diyaismahil7@gmail.com>
2026-09-18 19:32:06 +00:00
realAsma b9cfdce8dc docs: add Local Hessian NVFP4 weight-scale announcement blog (#2417)
### What does this PR do?

Type of change: documentation

Adds a Local Hessian announcement blog at
`docs/source/announcements/local-hessian.rst`, covering the NVFP4
per-block
weight-scale rule that minimizes output error instead of weight error.

Contents:

- Derivation of the per-block output-error objective and its `16x16`
local
  Hessian, with numbered equations.
- Results on Qwen3.5-9B: scale-setting comparison against max, MSE, and
  Four-over-six, plus composition with GPTQ.
- Figure 1, a grouped bar chart of the Qwen3.8-27B W4A4 candidate scores
  (BF16 in gray, the two scale rules in NVIDIA greens).
- A "Using Local Hessian" section with the config example and the
  end-to-end `hf_ptq.py` command.

Two supporting changes outside the blog:

- `docs/source/_static/announcements.css`: the `shibuya` theme has no
`span.eqno` rule, so Sphinx's default `float: right` on equation numbers
cannot share a line with MathJax's full-width display block and the
number
renders *above* the equation. This anchors it to the right of the
equation
  instead, and shrinks the table-note class.
-
`docs/source/announcements/assets/qwen3-27b-w4a4-scale-rule-accuracy.png`:
  the Figure 1 asset.

### Usage

```python
import modelopt.torch.quantization as mtq

config = {
    "quant_cfg": [...],  # quantizer configuration
    "algorithm": {"method": "local_hessian", "fp8_scale_sweep": True},
}

model = mtq.quantize(model, config, forward_loop)
```

### Testing

Documentation only; no code paths change. The `.rst` parses cleanly
under
docutils. The rendered page has not been checked with a full
`sphinx-build`,
so the equation-number CSS fix and the figure placement are worth an
eyeball
on the built docs before merge.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌

### Additional Information

Two items to settle before this is ready to publish:

1. The `--recipe` example points at

`modelopt_recipes/models/Qwen/Qwen3.8-27B/ptq/nvfp4_local_hessian-fp8_attn-kv_fp8_cast.yaml`,
a placeholder path derived from the existing recipe naming convention.
It
   needs to match whatever lands in #2363.
2. The tables report single-run team measurements; the blog says so and
makes
   no significance claims.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Added guidance on NVFP4 Local-Hessian weight-scale selection,
including mathematical details, accuracy comparisons, runtime
considerations, limitations, configuration examples, and reproduction
steps.
- Updated announcement labels, headings, metadata, descriptions, and
filtering text to use “Local-Hessian.”

- **Style**
- Improved announcement formatting for equation labels, display-equation
spacing, Hessian results, table headers, and explanatory notes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-09-16 21:48:11 +00:00
Keval MorabiaandClaude Opus 5 a448ba9757 Add end-to-end W4A4 NVFP4 + QAD tutorial for Qwen3.6-35B-A3B (#2411)
### What does this PR do?

Type of change: new example + bug fix

<img width="2085" height="1239" alt="image"
src="https://github.com/user-attachments/assets/b9ced215-ce8c-4dbe-be74-a75c1c4714b3"
/>


Adds an end-to-end **W4A4 NVFP4 + Quantization-Aware Distillation**
tutorial for
[Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) at
`examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B/`.

It complements the existing Nemotron-3-Nano tutorial (pruning +
distillation + FP8). Here the model
is unpruned and the technique under test is **W4A4** — aggressive enough
that PTQ alone leaves a
measurable accuracy gap, which is what QAD exists to close.

**Why W4A4 rather than weight-only NVFP4:** W4A16 measured *slower than
BF16* in 10 of 12 shapes,
because a BF16 activation forces vLLM onto the Marlin dequant fallback
and never reaches the
Blackwell FP4 tensor cores. W4A4 beats BF16 in 9 of 12 shapes (up to
1.30x) and shrinks the
checkpoint 67 GiB -> 22 GiB (3.1x).

**What the study found:** only 2 of 6 benchmarks show a statistically
significant PTQ deficit, so
those are the only two QAD can recover. IFBench is recovered to parity
with BF16 (-2.62 pp ->
-0.29 pp, gain of +2.33 pp, p=0.036); MMMU-Pro recovers ~40% and retains
a significant gap. The
other four are lossless under W4A4 to begin with.

Also added:

-
`modelopt_recipes/model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore.yaml`
— the PTQ
  recipe used as the QAD student, usable via `--recipe`.
- `data_blend.yaml` — the token-budgeted blend config for the
distillation data.
- `eval_configs/*.yaml` — one NeMo Evaluator config per benchmark.
tau2-bench is separate because
it needs `--enable-auto-tool-choice --tool-call-parser qwen3_coder` and
`deployment.command` is
  global to a config.

**Two export fixes found while producing these checkpoints** (both
change library/example behaviour,
both have changelog entries under 0.48.0 Bug Fixes):

- `unified_export_megatron.py` — MCore builds `embedding` on the MTP
stage as well as the first, so
gating export on `hasattr(model, "embedding")` wrote a **second,
unreferenced copy of the vocab
embedding** whenever an MTP model was exported with PP > 1. The index
mapped the key to the later
shard, so the extra copy never loaded but still shipped — ~1 GB for this
model. Now gated on
`model.pre_process`, MCore's own "this rank owns the input embedding"
flag.
- `export_quantized_megatron_to_hf.py` — stopped passing Megatron's
`moe_router_dtype` as the
router's *storage* dtype. It is a routing *compute* dtype; the parameter
is bf16 in a bf16 model,
so the export was widening bf16 to fp32. All 21,495,808 router values in
the exported checkpoint
have their low 16 bits zero, and vLLM builds the gate at the model dtype
and rounds on load, so
the dropped bytes carried no information. `export_mcore_gpt_to_hf` still
accepts the override.

### Usage

```bash
# 1. PTQ (2 GB200 nodes for EP=8)
srun ... python examples/megatron_bridge/quantize.py \
    --hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \
    --recipe model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore \
    --tp_size 1 --ep_size 8 --pp_size 1 \
    --calib_dataset_name cnn_nemotron_v2_mix --calib_num_samples 1024 --calib_batch_size 1 \
    --seq_length 8192 --skip_generate \
    --export_megatron_path /path/to/qwen36_w4a4_megatron

# 2. QAD (32 nodes x 4 GB200)
python -u examples/megatron_bridge/distill.py \
    --teacher_hf_path Qwen/Qwen3.6-35B-A3B --student_hf_path Qwen/Qwen3.6-35B-A3B \
    --student_megatron_path /path/to/qwen36_w4a4_megatron \
    --tp_size 1 --pp_size 1 --cp_size 1 --ep_size 8 \
    --seq_length 32768 --mbs 1 --gbs 512 --train_iters 500 \
    --lr 1e-5 --min_lr 1e-6 --lr_warmup_iters 50 --logit_kl_topk 4096 \
    --recompute_granularity full --recompute_method uniform --recompute_num_layers 1 \
    --no_async_save --eval_iters 0 --save_interval 50 \
    --data_paths "${DATA_BLEND}" --output_dir /path/to/qad_output
```

### Testing

**Library changes.**
`tests/gpu_megatron/torch/export/test_unified_export_megatron.py` gains
`test_unified_export_megatron_pp2_mtp_no_duplicate_tensors`: it exports
a PP=2 model built with
`mtp_num_layers=1` and asserts no tensor lands in more than one shard.
Verified to **fail without
the fix**:

```
AssertionError: tensors written to more than one shard:
  {'model.embed_tokens.weight': ('model-00001-of-00002.safetensors',
                                 'model-00002-of-00002.safetensors')}
```

The pre-existing `..._pp2_mtp_metadata_matches_shards` test cannot catch
this — it fakes
`_get_mtp_state_dict` on a model with no real MTP, so the last stage
never builds an embedding.

Ran the whole `tests/gpu_megatron/torch/export/` suite with and without
the fix: identical failure
sets (3 failures both ways, all `qwen3_5_moe_vl_*` from a local
`ImportError: FLA is not installed`),
58 passed with vs 56 without — the +2 being the new test's two workers.
`tests/unit/recipe` passes
368/368 after the recipe path move. `model.pre_process` is always
present: `GPTModelExporter.__init__`
raises unless the model is `GPTModel` or `HybridModel`, and both set it
unconditionally.

Both export fixes were also applied to the real 23 GB checkpoints and
re-validated end to end: every
retained tensor md5-identical, index/shard integrity re-checked, and a
**full GPQA re-evaluation of
the fixed checkpoint** scored 83.49 vs 84.25 before (paired per-question
t-test over the same 198
questions x 16 repeats: -0.76 pp, p=0.21, not significant).

**Numbers in the tutorial** come from real runs, not estimates:

- **253 evaluation runs** across BF16, the published W4A16 checkpoint,
W4A4 PTQ, and QAD at
50 / 300 / 500 iterations — 8 repeats per benchmark (3 for tau2-bench;
GPQA is one
  `num_repeats: 16` run).
- The published `nvidia/Qwen3.6-35B-A3B-NVFP4` checkpoint was
re-evaluated under this same harness
(36 runs) rather than quoted from its card, so the W4A16 row is
same-harness.
- Every figure and results-table value is generated from the collected
`results.yml` files by a
script, and I verified the README table cell-by-cell against that data
after each edit.
- Throughput rows were cross-checked against the recorded AIPerf sweeps;
the QAD wall-clock figures
  against the two jobs' Slurm records (`03:34:49` + `02:09:49`).
- All CLI flags in the tutorial were verified to exist in `quantize.py`
/ `distill.py` /
`export_quantized_megatron_to_hf.py`, and `cnn_nemotron_v2_mix` against
`dataset_utils.py`.

The tutorial also records the non-obvious constraints found the hard
way: QAD on this model requires
`TP=PP=CP=1` (TP breaks quantizer `_amax` dist-checkpoint sharding, PP
starves Qwen3-VL's M-RoPE of
`position_ids`, CP hits a rope shard mismatch), EP must match the PTQ
checkpoint, and
`--logit_kl_topk` is mandatory at 32K because the dense `[seq, vocab]`
fp32 logits are 30.31 GiB per
tensor on a 248,320-token vocabulary.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ <!-- gpu_megatron PP=2+MTP
export dedup test -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!-- 0.48.0: Megatron Framework + two Bug Fixes -->
- Did you get Claude approval on this PR?: ✅ <!-- not yet run -->

### Additional Information

Changelog entries are filed under **0.48.0**; the `cherry-pick-0.47.0`
label has been removed.
Rebased onto `main` after #2328 renamed `modelopt_recipes/huggingface`
to `model_type` (it is now a
compatibility symlink), so the recipe moved to
`model_type/qwen3_6_moe/ptq/`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an end-to-end Qwen3.6-35B-A3B tutorial for W4A4 NVFP4
quantization and quantization-aware distillation.
* Added checkpoint export, accuracy evaluation, and vLLM throughput
benchmarking workflows.
* Added evaluation configurations for AA-LCR, GPQA, IFBench, MMMU-Pro,
SciCode, and tau2 Telecom.
  * Added a token-budgeted supervised fine-tuning data configuration.
  * Added a Megatron-Core NVFP4/FP8 quantization recipe for Qwen3.6-MoE.

* **Documentation**
* Added benchmark results, deployment guidance, hardware requirements,
reproduction steps, HTTPS endpoint guidance, announcement filters, and
tutorial links.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 02:10:12 +05:30
Chad Voegele c118d359c1 docs: replace legacy TensorRT-LLM engine deployment guidance (#2436)
Chad's Agent

### What does this PR do?

Type of change: documentation.

Replace legacy TensorRT checkpoint export, support matrix, and
engine-build instructions with `export_hf_checkpoint` and TensorRT-LLM's
PyTorch backend. Preserve the existing 0.48.0 deprecation / 0.49.0
removal notice and page URL. Update the customized-model guide and
deployment skill to match.

### Usage

Follow the linked unified HF export guide. No API changes.

### Testing

- `git diff --check` passed.
- `uvx pre-commit run --files docs/source/deployment/1_tensorrt_llm.rst
docs/source/guides/_customized_model_quantization.rst
plugins/modelopt/skills/deployment/references/trtllm.md` passed.
- Full Sphinx build delegated to the Docs workflow; preview expected
after deployment.

### Before your PR is "Ready for review"

- Is this change backward compatible?: ✅ Documentation only; page URL
retained.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — documentation only.
- Did you update Changelog?: N/A — guidance correction; no new API
deprecation.
- Did you get Claude approval on this PR?: ❌ Not requested yet.

### Additional Information

Removes instructions for the TensorRT backend that current TensorRT-LLM
releases no longer support.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Updated TensorRT-LLM deployment guidance to use `export_hf_checkpoint`
with the PyTorch backend.
- Clarified that this workflow does not require TensorRT engine
construction.
- Updated DBRX customization instructions for exporting and deploying
quantized models.
- Added TensorRT-LLM version requirements and links to unified Hugging
Face deployment guidance.
- Removed guidance for the legacy TensorRT-LLM checkpoint exporter and
outdated troubleshooting steps.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
2026-09-16 13:43:37 -05:00
Shengliang Xu c7ed23a103 Rename modelopt_recipes/huggingface to model_type with backward-compat alias (#2328)
### What does this PR do?

Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.

Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.

- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
  source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
  `$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
  notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
  failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.

### Usage

```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none

# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```

```python
from modelopt.recipe import load_recipe

load_recipe("model_type/vit/ptq/fp8")    # canonical
load_recipe("huggingface/vit/ptq/fp8")   # deprecated alias, resolves to the same recipe
```

### Testing

- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
  structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
  on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
  `model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
  the loader alias — including a recipe that pulls internal `$import`s.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.

### Additional Information

The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.

- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.

- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
  - Local recipe files now take precedence over built-in recipes.
  - Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-15 12:16:12 -07:00
Jenny Chen f70991f36e Docs: Add QAT and QAD guide [OMNIML-4859] (#2255)
### What does this PR do?

Type of change: Documentation

Add detailed quantization aware training (QAT and QAD) guide in our
docs, featuring examples on how to run in HF, Megatron-Bridge, and
Megatron-LM

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Very short summary of changes only for new features,
backward breaking changes, deprecations, or fixes for critical bugs
present in previous releases. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added a comprehensive guide for quantization-aware training (QAT) and
distillation (QAD), including workflows, framework guidance, setup,
training, and export steps.
  * Updated the Guides navigation to include the new QAT/QAD guide.
* Expanded quantization guidance with QAT/QAD workflows, scale handling,
and NVFP4 references.
* Improved save and restore guidance with clearer references and an
explanatory note.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
2026-09-14 07:19:00 -07:00
Ajinkya RasaneandCodex 5b1f7e86cc [6701308][OMNIML-5805] Correct ONNX PTQ documentation contracts (#2413)
### What does this PR do?

Type of change: documentation

Align the ONNX PTQ README, guide, and executable example with the
implemented contracts:

- use the canonical `--calibration_data_path` CLI option;
- load `.npy` calibration data before passing it to the Python API;
- document the supported Autotune modes and calibration methods;
- correct the minimum opsets to INT8 19, FP8 19, and INT4 21; and
- describe the no-data fallback as random calibration inputs.

This also removes an inaccurate source comment without changing runtime
behavior.

### Usage

```bash
python -m modelopt.onnx.quantization \
    --onnx_path=model.onnx \
    --quantize_mode=int8 \
    --calibration_data_path=calib.npy \
    --output_path=model.quant.onnx
```

### Testing

- `pre-commit run --files docs/source/guides/_onnx_quantization.rst
examples/onnx_ptq/README.md modelopt/onnx/quantization/quantize.py
tests/examples/test_onnx_ptq.sh`
- `bash -n tests/examples/test_onnx_ptq.sh`
- `CUDA_VISIBLE_DEVICES="" python -m pytest -o addopts="" -p
no:cacheprovider --confcutdir=tests/unit/onnx/quantization -q
tests/unit/onnx/quantization/test_autotune_quantization_integration.py`
(4 passed)
- `nox -s docs` (passed; Sphinx built 881 HTML files)
- Focused before/after contract probe covering the documented CLI
option, API data type, Autotune modes and methods, opset minimums, and
random-input wording

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

Tracking: [6701308]

> 🤖 _Generated by Codex (AI agent)._


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Documentation**
- Clarified that random calibration inputs are used when no calibration
dataset is provided.
- Updated ONNX post-training quantization examples with minimum opset
requirements and the `calibration_data_path` argument.
- Clarified Autotune support for FP8 and INT8 calibration methods using
`max` or `entropy`.

- **Tests**
- Updated quantization command examples to use the current calibration
data path option.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
2026-09-11 22:04:19 +00:00
Shengliang Xu d19925e446 simple refactor(export): split TensorRT-LLM-only code into modelopt/torch/export/trtllm (#2365)
### What does this PR do?

Type of change: refactor.

**The TensorRT-LLM checkpoint export format is deprecated.** Per
`docs/source/deployment/1_tensorrt_llm.rst`: *"The
`export_tensorrt_llm_checkpoint` API will be deprecated in future
releases. Users are encouraged to transition to the unified HF export
API, which provides enhanced functionality and flexibility for exporting
models to multiple inference frameworks including TensorRT-LLM, vLLM,
and SGLang."*

That deprecated code was not sitting off to one side — it was
**interleaved with the export path we actually want to grow.**
`modelopt/torch/export` mixed the deprecated TensorRT-LLM checkpoint
logic with the framework-agnostic HF/Megatron export code, in the same
modules:

- `layer_utils.py` was 1,986 lines, of which ~1,600 were TensorRT-LLM
`build_*_config` builders. The HF path imports this module for five
small predicates (`is_moe`, `is_quantlinear`, …) and dragged the whole
deprecated builder set in with them.
- `model_config.py` held the TensorRT-LLM `ModelConfig` dataclasses
*and* the `QUANTIZATION_*` / `KV_CACHE_*` constants that every backend
needs, so all of HF export imported the deprecated checkpoint schema to
get a format name string.
- `quant_utils.py` carried two helpers whose only caller is the
deprecated `postprocess.py`.

**This PR isolates the deprecated format so it stops polluting the
HuggingFace export path.** Everything reachable only from
`export_tensorrt_llm_checkpoint` now lives under
`modelopt/torch/export/trtllm/`, and the dependency is **one-way**:
`trtllm/` reaches into the parent through `quant_format`, `quant_utils`
and `layer_utils`, and **no implementation module in the parent imports
`trtllm/`.** The single exception is the deprecation re-export in
`modelopt/torch/export/__init__.py` described below, which is scheduled
for deletion in 0.49.0.

That one-way edge is the property worth protecting in review. It means
the deprecated format can be evolved, frozen, or eventually removed
without touching HF export, and HF export can no longer accidentally
grow a dependency on it.

### Deprecation handling

The format has carried a deprecation notice in the deployment docs since
`bc546943b4` (2025-10-08, first shipped in 0.39.0) — about 11 months.
But the deprecation policy in `README.md` also specifies *how* a
deprecation is communicated: a changelog entry, a source statement of
timing, and a runtime warning on use. **None of those existed**; only
one docs page ever said anything. So 0.48.0 is the first release that
gives users a signal they can act on, and this PR treats it as the
*start* of the migration period rather than the end:

- Both entry points now emit a `DeprecationWarning` naming 0.48.0 and
the 0.49.0 removal.
- `export_tensorrt_llm_checkpoint` and
`torch_to_tensorrt_llm_checkpoint` **remain importable from
`modelopt.torch.export`** for this release only, so existing callers
keep working *and* actually receive the warning. Removing the path in
the same release that first warns would mean callers hit `ImportError`
and never see it.
- The 0.49.0 removal date is stated in all four channels the policy
names: the runtime warning, the source (`.. deprecated:: 0.48.0` plus a
comment), the changelog, and the deployment doc.

The deeper module paths (`modelopt.torch.export.model_config_export`,
`modelopt.torch.export.model_config`) are **not** forwarded. Neither
appeared in a docs example, and `model_config.py` never declared
`__all__`, so by the `__all__` convention in `CONTRIBUTING.md` they were
never part of the public surface.

Eight modules had no non-TRT-LLM importer and moved whole:
`model_config_export`, `model_config_utils`, `postprocess`,
`distribute`, `tensorrt_llm_utils`, `tensorrt_llm_type`,
`hf_config_map`, `mcore_config_map`.

Three were genuinely mixed and were split by call-graph analysis rather
than by file:

| module | stayed shared (HF path) | moved to `trtllm/` (deprecated) |
|---|---|---|
| `model_config.py` | `QUANTIZATION_*`, `KV_CACHE_*`,
`FUSION_FREE_FORMATS` → new leaf module `quant_format.py` | the
`ModelConfig` dataclasses + `LINEAR_*`/`LAYERNORM_*` checkpoint-layout
constants |
| `layer_utils.py` | 9 module-shape predicates and MoE quantizer helpers
(`is_moe`, `is_quantlinear`, `get_experts_list`,
`sync_moe_gate_up_amax`, …) | the 39 `build_*_config` builders and
enc/dec helpers |
| `quant_utils.py` | everything else | `get_scaling_factor_from_weight`,
`resmooth_and_get_scale` (only caller is `trtllm/postprocess.py`) |

`adjust_attn_amax_values` was deliberately left in the shared
`quant_utils.py`: it has no production caller at all (only a test), so
"used only by TRT-LLM export" is not demonstrable for it.

Nothing was added or removed. `export_tensorrt_llm_checkpoint` behaves
exactly as before, just from a new import path and with a warning
attached.

### Usage

```python
# Deprecated TensorRT-LLM checkpoint export — new home, and warns on call
from modelopt.torch.export.trtllm import (
    export_tensorrt_llm_checkpoint,
    torch_to_tensorrt_llm_checkpoint,
)
from modelopt.torch.export.trtllm.model_config import ModelConfig

# The pre-0.48 path still works for one release, and warns — removed in 0.49.0
from modelopt.torch.export import export_tensorrt_llm_checkpoint

# Shared format constants — new home, still re-exported from the top level
from modelopt.torch.export.quant_format import QUANTIZATION_NVFP4, KV_CACHE_FP8
from modelopt.torch.export import QUANTIZATION_NVFP4  # still works

# The recommended path — unchanged
from modelopt.torch.export import export_hf_checkpoint, get_model_type
```

### Testing

- `pre-commit` on all changed files: passes (ruff, ruff-format,
**mypy**, bandit, markdownlint). mypy caught one implicit re-export of
`is_layernorm`, now imported from the shared module directly.
- `tests/unit/torch/export`: **189 passed**. With the new `trtllm/` test
dir: **193 passed**.
- Full `tests/unit/torch`: **2367 passed, 0 export failures**. The 45
failures are pre-existing environment issues — a deepspeed circular
import and a read-only HF cache — confirmed by reading their error text,
not assumed.
- `pytest tests/gpu/torch/export --collect-only`: 172 items, no
collection error.
- In-repo consumers updated and re-verified by an AST scan that imports
every `modelopt.torch.export*` module referenced anywhere in the tree
and checks each imported name still resolves: `hf_ptq.py`,
`export_trtllm_ckpt.py`, `deepseek_v3/ptq.py`, the AutoQuantize
notebook, `hf_ptq/README.md`, 2 docs pages, 4 tests.
- **Deprecation contract is covered by committed tests** (3 new, in the
`trtllm/` test dir): the pre-0.48 top-level import still resolves to the
same objects, `torch_to_tensorrt_llm_checkpoint` warns *at call time*
rather than on first `next()` (it returns a generator, so a naive
`warnings.warn` in the body would fire late or never), and one
`export_tensorrt_llm_checkpoint` call emits exactly one warning rather
than two. The first of these makes closing the migration window early a
test failure rather than a silent regression. `pyproject.toml` sets no
`filterwarnings = error`, so no suite fails on the new warning.
- **After merging `main`** (4 commits, incl. a 180-line rewrite of
`unified_export_megatron.py` that touches a file this PR also edits):
merged with no conflicts, then re-verified rather than trusted — import
scan clean across 24 export modules, `ruff check` clean repo-wide, 193
export tests passing, GPU collection still clean.

**Not run: the GPU suites** (`tests/gpu/torch/export`,
`tests/gpu_trtllm`) — no GPU in my environment.
`tests/gpu/torch/export/test_export.py` had its imports retargeted, so
it is the one most worth a GPU run before merge.

### Reviewer note: the deprecated path has no test coverage

Worth knowing before reviewing. **No test in the repo — including
`tests/examples/` — calls `export_tensorrt_llm_checkpoint`,
`torch_to_tensorrt_llm_checkpoint`, any `build_*_config`,
`convert_to_tensorrt_llm_config`, or `postprocess_model_config`.** So
~4,600 moved lines have no direct tests, and this refactor is validated
by import-graph reasoning, lint and mypy rather than by tests exercising
the moved code.

Given the format is deprecated and scheduled for removal in 0.49.0, **no
new coverage is planned for the conversion path itself** — writing fresh
tests for an API being removed next release isn't a good use of effort.
The gap is documented so reviewers can weigh the risk, not as a TODO.
(The deprecation *mechanism* is tested; see Testing.)

One caveat on how the gap was established: a runtime check showing all
12 `trtllm` modules in `sys.modules` after the export suites is *not*
evidence of coverage — importing any submodule runs
`trtllm/__init__.py`, which star-imports `model_config_export` and pulls
in the rest. Real line coverage could not be measured (`coverage`'s
tracer is incompatible with this venv's torch build: `ValueError: module
functions cannot set METH_CLASS or METH_STATIC`, on both the C tracer
and `sysmon`). The claim rests on a call-site audit generated from the
actual public symbols of those modules.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ for the public API —
`export_tensorrt_llm_checkpoint` and `torch_to_tensorrt_llm_checkpoint`
remain importable from `modelopt.torch.export` through the 0.49.0
migration period, now with a `DeprecationWarning`. The undocumented
submodule paths `modelopt.torch.export.model_config_export` and
`.model_config` did move; see **Usage**.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
code or dependencies; existing code relocated.
- Did you write any new necessary tests?: ✅ — 3 tests covering the
deprecation contract (old import path, call-time warning, exactly-one
warning). One existing test also moved to mirror the source split. No
new coverage for the deprecated conversion path itself; see the note
above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 **Deprecations**, covering both the runtime warning and
the new import location.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Git detected the moves, so the diff stays reviewable: 8 files show as
pure renames (100%), the three split files as rename/copy at 94–99%
similarity, and only `layer_utils.py` as a 79% rewrite — expected, since
it shed 1,616 lines to `trtllm/`.

`examples/hf_ptq/hf_ptq.py` and
`examples/llm_sparsity/weight_sparsity/export_trtllm_ckpt.py` still call
the deprecated API, so those examples now print the warning. That is the
intended nudge, but happy to silence or migrate them if preferred. They
import from the new `.trtllm` path already, so they need no change at
0.49.0.

Two incidental changes, easy to revert if unwanted:
- `modelopt/torch/export/layer_utils.py` mode `100755 → 100644` (it was
needlessly executable).
- The new test is named `test_trtllm_quant_utils.py`, not
`test_quant_utils.py`: these directories have no `__init__.py`, so
pytest derives the module name from the bare filename and the shorter
name fails collection with `import file mismatch` against the existing
`test_quant_utils.py` one level up.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added shared quantization and KV-cache format definitions for export
workflows.
- Added expanded TensorRT-LLM export support, including broader model
architecture and quantization handling.
- Added distributed export utilities for coordinating checkpoint data
across processes.

- **Deprecation**
- TensorRT-LLM checkpoint export now emits a warning and is scheduled
for removal in version 0.49.0.
- Use the documented export module and save optimized model state
explicitly when needed.

- **Documentation**
- Updated guides and examples with new import paths and deprecation
guidance.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-10 10:44:04 -07:00
haoxiz-nvidia acdf330414 Add TensorRT-RTX ABI EP support for ONNX quantization (#2262)
### What does this PR do?

Type of change: new feature

Adds opt-in support for using the standalone TensorRT-RTX ABI Execution
Provider during ModelOpt ONNX quantization.

Users select the ABI backend with:

`--calibration_eps=NvTensorRtRtx --trt_rtx_backend=abi`

When selected, ModelOpt imports and registers the installed TensorRT-RTX
ABI provider before creating the ONNX Runtime inference session. The
backend selection is propagated through INT8, FP8, and INT4 AWQ
calibration paths, including the Windows GenAI LLM quantization example.

The existing `--calibration_eps=NvTensorRtRtx` behavior remains backward
compatible. The `legacy` backend is still the default and continues to
use TensorRT-RTX libraries supplied through `PATH`.

For Windows x64 with Python 3.11 or newer, the ONNX dependencies now
include:

- `onnxruntime-gpu~=1.26.0`
- `onnxruntime-ep-nv-tensorrt-rtx-cu13==0.4.0`

Keeping `onnxruntime-gpu` allows users to select either CUDA EP or
TensorRT-RTX ABI EP for calibration. Windows-on-Arm source-build
instructions are intentionally out of scope and will be documented
separately.

### Usage

```powershell
python -m modelopt.onnx.quantization `
  --onnx_path="C:\path\to\Llama-3.2-3B-Instruct\model.onnx" `
  --model_id="C:\path\to\Llama-3.2-3B-Instruct\config.json" `
  --quantize_mode=int8 `
  --output_path="C:\path\to\int8_abi\model.onnx" `
  --calibration_eps=NvTensorRtRtx `
  --trt_rtx_backend=abi `
  --use_external_data_format `
  --high_precision_dtype=fp32 `
  --log_level=INFO

### Testing
unit test have been added

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅
- Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ 
- Did you get Claude approval on this PR?: pending



<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit

- **New Features**
  - Added optional TensorRT-RTX ABI backend support for ONNX calibration on Windows ARM64.
  - Added `legacy` and `abi` backend selection to quantization APIs and command-line tools; `legacy` remains the default.
  - Added validation for unsupported backends and incompatible TensorRT plugin configurations.
  - Updated Windows ARM64 installation support and platform-specific package configuration.

- **Documentation**
  - Updated Windows installation guidance, Python compatibility requirements, ARM64 setup, and verification steps.
  - Documented the new TensorRT-RTX backend command-line option.

- **Tests**
  - Added coverage for ABI provider registration, backend validation, and compatibility checks.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Haoxi Zhang <haoxiz@nvidia.com>
2026-09-09 04:04:46 +00:00
haoxiz-nvidia 4773f72f8a Docs: Add WOA documentation (#2264)
### What does this PR do?

Add WoA env setup guide. Includes build instruction of pyarrow, which
used by datatsets

### Usage

N/A

### Testing
N/A

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Added comprehensive Windows on Arm installation guidance, including
prerequisites, environment setup, dependency installation, verification,
and troubleshooting.
- Documented experimental Windows ARM64 support, supported quantization
formats, native dependency requirements, and Support Matrix details.
  - Expanded supported Windows Python versions through 3.13.
- Expanded TensorRT-RTX guidance for calibration, deployment, provider
setup, and standalone plugin usage.
- Clarified PyArrow requirements and local build instructions for
Windows ARM64.
- Added links to dedicated Windows on Arm installation resources and
shared TensorRT-RTX documentation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Haoxi Zhang <haoxiz@nvidia.com>
2026-09-08 15:04:04 +05:30
Ajinkya RasaneandCodex c49ce57d75 [5591371] Add performance guard for ONNX Autotune (#2318)
### What does this PR do?

Type of change: Bug fix

This PR prevents integrated ONNX Autotune from saving an INT8/FP8 result
that does not improve TensorRT latency. Autotune search models now use
the same FP16/BF16 conversion path as the delivered model. After
calibration and existing Q/DQ post-processing, the exact candidate is
benchmarked against its precision-matched no-Q/DQ baseline.

- Keep Q/DQ when the measured speedup meets
`Config.performance_threshold` (`1.02x` by default, inclusive).
- Otherwise save the high-precision no-Q/DQ fallback at the requested
output path and report `no_qdq`.
- Reject already-quantized Autotune inputs because they cannot produce a
true no-Q/DQ baseline.
- Leave quantization without Autotune, standalone uncalibrated Autotune,
pattern search, caches, and state schema v1 unchanged.

### Usage

```bash
python -m modelopt.onnx.quantization \
  --onnx_path=model.onnx \
  --quantize_mode=fp8 \
  --calibration_data_path=calibration.npz \
  --high_precision_dtype=fp16 \
  --autotune=default \
  --output_path=model.autotuned.onnx
```

The output contains either the accepted Q/DQ placement or the
high-precision fallback. The log reports `qdq` or `no_qdq`, the two
measured latencies, the speedup, and the threshold.

### Testing

- Ran the CPU-only ONNX Autotune, runtime-precision, and quantization
API suites with no GPU visible and CPU execution providers: 199 passed.
- Ran all applicable pre-commit hooks on the 12 changed files, including
Ruff, mypy, Bandit, license, and RST checks.
- On an RTX 6000 Ada GPU with TensorRT 10.8, ran explicit
`--autotune=default` on a synthetic `Conv(128→128) → Relu → MaxPool →
Gemm` graph. The calibrated guard retained two Q/DQ sites from its
paired measurement (`0.066 ms / 0.064 ms = 1.023x`, threshold `1.020x`).
The selected, baseline, and candidate models all built with `trtexec
--stronglyTyped` without an output-type error. Five alternating
follow-up trials also favored Q/DQ (`1.016x` median speedup).

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

Related to #439.

> 🤖 _Generated by Codex (AI agent)._


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- ONNX Autotune now benchmarks candidates at the requested runtime
precision and supports custom model transformations during export.
- Calibrated INT8/FP8 quantization is retained only when it meets the
configured performance threshold (default 1.02×); otherwise, the
high-precision model is saved without Q/DQ.

- **Bug Fixes**
  - Improved runtime-precision handling across INT8 and FP8 workflows.
- Improved validation of inputs, pre-quantized models, failures, and
temporary resources.

- **Documentation**
- Updated Autotune guidance and command-line help to explain runtime
precision, performance validation, and fallback behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
2026-09-03 12:20:27 -04:00
Shengliang Xu de3eda8a11 Restructure recipes: split per-model_type recipes from model-hub checkpoint recipes (#2219)
### What does this PR do?

**Type of change:** Refactor (recipe-library layout) + documentation —
backward-breaking for saved `--recipe` paths.

Separate the two kinds of built-in Hugging Face recipes that were
previously mixed under `modelopt_recipes/huggingface/`:

- **`huggingface/<model_type>/`** — architecture recipes keyed by the
transformers `model_type`; one recipe covers every checkpoint of that
architecture. **Unchanged.**
- **`models/<org>/<model_id>/`** — a *new top-level tier* for recipes
that mirror one specific published checkpoint, keyed by its **model-hub
path** (as on the Hugging Face Hub, ModelScope, etc.) so the on-disk
path equals the hub path.

Concretely, the model-instance recipes move out of `huggingface/` to the
top level:

- `huggingface/models/mistralai/…`, `huggingface/models/nvidia/…` →
`models/mistralai/…`, `models/nvidia/…`
- `huggingface/step3p5/Step3.5-Flash/…` →
`models/stepfun-ai/Step-3.5-Flash/…` (re-keyed to the canonical HF repo
id
[`stepfun-ai/Step-3.5-Flash`](https://huggingface.co/stepfun-ai/Step-3.5-Flash)
— org `step3p5`→`stepfun-ai`, id `Step3.5-Flash`→`Step-3.5-Flash`)

**Why:** `modelopt_recipes/README.md` already documented a top-level
`models/` tier, but the files lived under `huggingface/models/` and
instance-specific recipes were awkwardly nested under the
per-`model_type` tree. This aligns the filesystem with the documented
layout and makes the instance tier hub-addressable — given a checkpoint
id you can find (or place) its recipe with no lookup table.
`load_recipe` resolves paths directly under `modelopt_recipes/`, so a
top-level `models/` sibling of `general/` and `huggingface/` works
identically.

The move is metadata-only — all recipe YAML content is byte-identical
(`R100` renames). Everything else is updating references (nvidia
launcher YAMLs, `test_loader.py`) and docs: a new `models/README.md`,
plus `huggingface/README.md`, root `README.md`, `ptq.md`, and the
`10_recipes.rst` guide, which no longer describe instances under
`huggingface/`.

### Usage

Recipe paths for the moved checkpoint recipes lose the `huggingface/`
prefix (and Step 3.5 Flash is keyed by its hub id):

```python
from modelopt.recipe import load_recipe

# before
load_recipe("huggingface/models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16")
load_recipe("huggingface/step3p5/Step3.5-Flash/ptq/nvfp4-mlp-only")

# after
load_recipe("models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16")
load_recipe("models/stepfun-ai/Step-3.5-Flash/ptq/nvfp4-mlp-only")
```

The same rename applies to `--recipe …` CLI values and launcher
`QUANT_CFG:` entries. Architecture recipes under
`huggingface/<model_type>/` are unaffected.

### Testing

- **Recipe resolution (torch-free):** parsed every recipe under
`models/` and confirmed all `$import` targets resolve against the recipe
root — 0 dangling across the tier.
- **Docs consistency:** re-ran the
`tests/unit/recipe/test_recipe_docs.py` logic; it now globs both
`huggingface/` and `models/`, and every model dir (incl.
`Step-3.5-Flash`, `Nemotron-3-Nano-4B-BF16`, …) plus every `general/ptq`
recipe is still mentioned in `ptq.md`.
- **Reference sweep:** repo-wide grep confirms no remaining references
to the old paths outside the intentional historical CHANGELOG entries
(released 0.44 / 0.45).
- **pre-commit:** `markdownlint-cli2`, license-insert, and `bandit`
hooks pass on the changed files.
- Note: the full `pytest` suite was not run in my environment (no
`torch`), so `test_recipe_docs.py` / `test_loader.py` should be
exercised in CI.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — `--recipe` / `load_recipe`
paths for the checkpoint-mirror tier change (drop the `huggingface/`
prefix; `step3p5/Step3.5-Flash` → `stepfun-ai/Step-3.5-Flash`).
Documented as a Backward Breaking Change in `CHANGELOG.rst` (0.47); the
only *released* old paths affected shipped in 0.45. A clean break was
chosen over a symlink or loader-alias shim.
- If you copied code from any other sources or added a new PIP
dependency …: N/A
- Did you write any new necessary tests?: ✅ — updated
`test_recipe_docs.py` to also glob the top-level `models/` tier so
instance recipes stay covered by the doc-consistency check.
- Did you update Changelog?: ✅ — added a 0.47 **Backward Breaking
Changes** entry.
- Did you get Claude approval on this PR?: ❌ <!-- run /claude review -->

### Additional Information

Design note: an earlier iteration nested everything under
`huggingface/model_type/` + `huggingface/models/`; the final layout
keeps `huggingface/` flat (per-`model_type`) and lifts instances to a
top-level `models/` tier, matching what `modelopt_recipes/README.md`
already documented. The `Step3p5*` architecture class names (from the
model's `trust_remote_code` modeling code) are unrelated to the recipe
path and are left unchanged.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added checkpoint-specific PTQ recipes for Kimi-K3, Mistral Medium 3.5,
and NVIDIA Nemotron models.
  * Added a Nemotron speculative-decoding warm-start recipe.

* **Documentation**
  * Clarified recipe selection and directory organization.
  * Documented checkpoint naming conventions and updated usage examples.

* **Bug Fixes**
* Updated launcher configurations and examples to reference the new
recipe locations and corrected model names.

* **Tests**
* Improved automatic recipe discovery and validation of documented
recipe paths.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-01 10:22:27 -07:00
Keval MorabiaandClaude Opus 5 d73278808b Bump nemo container requirement to 26.08 for MBridge examples (#2257)
### What does this PR do?

Type of change: Bug fix

Bumps the Megatron-Bridge examples, tests and launcher configs to
`nemo:26.08` and removes the version-gated fallbacks they carried, plus
the fixes needed to make the suites green on that container.

**26.08 bump and shim removal**

- Examples, CI workflows, `noxfile.py` and the `mbridge_*` launcher
configs move to `nemo:26.08`.
- `examples/megatron_bridge/_distillation_provider.py` is deleted —
26.08's Megatron-Bridge ships `convert_to_distillation_provider(...,
distill_submodule=...)` natively, so `distill.py` imports it directly.
- `prune_minitron.py` drops the `AutoBridge.from_hf_config` /
config-only-export probing; `--no_moe_grouped_gemm` is no longer needed
in the MoE pruning tests, and the Qwen3.5-MoE `skipif` is gone (native
MoE expert mappings are in 26.08).
- `_DynamicMambaMixer` targets only the raw `conv1d_weight` /
`conv1d_bias` parameters that replaced the `conv1d` module in
Megatron-Core.

**MambaModel / MambaModelProvider removal**

Megatron-Core has shipped `HybridModel` since 26.06 and `MambaModel` is
a deprecated subclass that shares its `forward`, so `DMRegistry`
resolves those instances to the `HybridModel` registration and the
separate entry is redundant. Same for `MambaModelProvider` vs
`HybridModelProvider` on the bridge side. `MambaMixer` / `MambaLayer` /
`ExtendedRMSNorm` are untouched — the layers still exist. The deprecated
`get_te_mamba_stack_spec` is removed; use `get_te_hybrid_stack_spec`.

**Bug fix: compressed output_layer extra state**

`mtq.compress` converts even a *disabled* `output_layer` into a
`RealQuantLinear` (its weight is left uncompressed, since
`pack_real_quantize_weight` skips disabled quantizers). The guard added
in #2112 excluded `RealQuantLinear` by class, so the layer emitted extra
state and every worker died in `GPTModel.sharded_state_dict`:

```
RuntimeError: Boolean value of Tensor with more than one value is ambiguous
  megatron/core/models/gpt/gpt_model.py:896, in sharded_state_dict
    output_extra_state and output_extra_state.data
```

The guard now keys off whether the weight was actually compressed
(`QTensorWrapper`) instead of the class. This took out all 12
`test_homogeneous_compressed_sharded_state_dict` params, and the crashed
workers poisoned the pool, which surfaced as unrelated timeouts and NCCL
errors in `test_layer_sync_moe_local_experts_amax`,
`test_kv_cache_quant`, `test_kv_cache_amax_sync`,
`test_convert_mcore_te_gpt_model` and
`test_homogeneous_sharded_state_dict_te_spec` — 21 tests in total. The
e2e coverage is `skip_flaky_on_blackwell`, so CI never ran it;
`test_output_layer_extra_state_empty_when_nothing_quantized` now asserts
the contract directly and is not skipped.

**Checkpoint import entry point**

26.08 replaced `examples/conversion/convert_checkpoints.py` with
`scripts/conversion/convert.sh`, so
`tools/launcher/common/megatron_bridge/import/import.sh` and the three
README snippets are retargeted. `import.sh` uses the distributed GPU
backend with `GPUS_PER_NODE` / `TP` / `PP` / `EP` knobs.

**Megatron-LM on nemo:26.06** keeps working: `_get_mamba_conv1d` still
dispatches between the `conv1d` module (26.06 and earlier) and the raw
parameters (26.08+), so `import_mcore_gpt_from_hf` /
`export_mcore_gpt_to_hf` handle NemotronH on both. Only the
Megatron-Bridge examples and Minitron pruning of Mamba/hybrid models
require 26.08.

**Test consolidation**

`test_export_distilled_megatron_to_hf.py` is merged into
`test_distill.py`: `test_distill_llm` becomes
`test_distill_llm_hf_export` and covers the standalone
`--export_iterations all` run on the checkpoints it already produces,
saving one full distillation (~185 s of CI time). The two mamba-named
gpu test files are renamed to `hybrid`.

### Usage

```bash
# HF -> Megatron import, via Megatron-Bridge's 26.08 conversion entry point
bash /opt/Megatron-Bridge/scripts/conversion/convert.sh import \
    --executor local \
    --device gpu \
    --gpus-per-node 8 \
    --hf-model Qwen/Qwen3-8B \
    --megatron-path /tmp/Qwen3-8B-megatron
```

### Testing

All on `nvcr.io/nvidia/nemo:26.08`, 2x RTX 6000 Ada, no timeout
overrides:

- `tests/examples/megatron_bridge`: 16 passed, 1 skipped (28m14s). The
skip is the `gemma3vl` QAD param, now `@pytest.mark.manual` since
`qwen3_5_moe_vl` covers the VLM QAD path.
- `tests/gpu_megatron` (`_extensions`, `distill`, `export`, `opt`,
`peft`, `sparsity`, `speculative`, `utils`): 61 passed, 5 xpassed.
- `tests/gpu_megatron/torch/export` re-run after the conv1d dispatch
change: 27 passed.
- The 21 previously failing/hanging quantization tests: 21 passed (12 +
9).
- `tests/gpu_megatron/torch/{nas,prune}`: verified separately.

`import.sh` equivalence on a toy `qwen3_moe`, comparing all 12 weight
tensors after flattening each dist checkpoint with `dcp_to_torch_save` —
the GPU backend at 1 GPU, `--tp 2`, `--pp 2`, `--ep 2`, and `import.sh`
end-to-end (`GPUS_PER_NODE=2 EP=2`) are all byte-identical to `--device
cpu`.

`nemo:26.06` compatibility was checked directly in that image:
`megatron.core.models.hybrid.HybridModel`, the modelopt hybrid spec and
`hybrid_layer_pattern` are all present, while
`megatron.bridge.models.hybrid` and the bridge's `distill_submodule` are
not. The NemotronH round-trip test failed there before the conv1d
dispatch was restored and the dispatch is back in place; per project
convention the suites themselves only run on 26.08.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ⚠️ Megatron-Bridge examples plus
Minitron pruning of Mamba/hybrid models now require `nemo:26.08`.
Megatron-LM quantization and checkpoint export still run on
`nemo:26.06`.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ —
`test_output_layer_extra_state_empty_when_nothing_quantized` for the
compress fix; existing tests extended for the merged export coverage.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added guidance for importing Hugging Face checkpoints into Megatron
distributed format.
* Expanded distillation workflows to export selected or all checkpoint
iterations.

* **Improvements**
  * Expanded Hybrid model support across Megatron workflows.
  * Updated distributed import tooling with GPU and parallelism options.
  * Updated supported environments and examples to NVIDIA NeMo 26.08.

* **Bug Fixes**
* Corrected output-layer quantization state handling when quantization
is disabled.

* **Documentation**
  * Added compatibility guidance for current and legacy NeMo containers.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 00:31:41 +05:30
realAsma a0513f18bd Update documentation theme to Shibuya (#2242)
### What does this PR do?

Type of change: documentation.

Replaces the legacy Sphinx RTD theme with Shibuya and gives the ModelOpt
documentation a modern responsive light/dark presentation.

The configuration uses the Shibuya green palette with NVIDIA green
(#76b900) as the primary accent, follows the reader system color
preference by default, enables dark code blocks, and expands the first
level of global navigation. RTD-specific CSS is removed, while the
announcement page now uses Shibuya semantic color tokens in both light
and dark modes.

Shibuya 2026.7.12 is licensed under BSD-3-Clause.

### Usage

```python
html_theme = "shibuya"
html_theme_options = {
    "accent_color": "green",
    "color_mode": "auto",
    "dark_code": True,
    "globaltoc_expand_depth": 1,
}
```

Doc preview:
https://nvidia.github.io/Model-Optimizer/pr-preview/pr-2242/

### Testing

- `nox -N --envdir /tmp/modelopt-shibuya-nox -s docs`
- Sphinx 9.1 built all 411 pages successfully with `--fail-on-warning`.
- `pre-commit run --files docs/source/conf.py
docs/source/_static/custom.css docs/source/_static/announcements.css
pyproject.toml uv.lock --show-diff-on-failure`
  - All applicable hooks passed.
- Verified generated HTML loads Shibuya assets, declares the green
accent, includes automatic light/dark mode logic, and includes the
custom theme-token CSS.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: N/A — documentation theme
change covered by the full warning-as-error build.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — documentation presentation change.
- Did you get Claude approval on this PR?: N/A — focused documentation
theme migration.

### Additional Information

No source code or public API behavior changes.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
  * Updated the documentation site with the Shibuya theme.
* Added green theme accents, automatic light/dark mode, and dark code
blocks.
* Improved table-of-contents behavior and refreshed announcement
styling.
  * Removed outdated layout and table customization overrides.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-08-26 15:38:08 -07:00
realAsma 8c04ce6ee2 docs: add AutoQuantize mixed-precision search blog (#1979)
### What does this PR do?

Type of change: documentation.

Adds the AutoQuantize technical blog to the announcements system
introduced by #1971.

- Preserves the source derivation, deployment-aware search details,
results, usage example, and references in native Sphinx RST.
- Adds the MMLU accuracy-versus-effective-bits figure as a losslessly
optimized asset.
- Adds a newest-first landing-page card and AutoQuantize filter.
- Credits the authors in this order: Asma Beevi K T, Wei Ming, Frida
Hou, Juhi Mittal, Jenny Chen, Ajinkya Rasane, Meng Xin.

This is a stacked PR targeting the branch for #1971. After #1971 merges,
this PR can be retargeted to main.

### Usage

N/A; documentation only.

### Testing

- Focused pre-commit hooks on all three changed files.
- git diff --check.
- Focused Sphinx HTML build for the announcement and landing page.
- Rendered-output checks for equations, table, Python code block, image
and alt text, references, external links, card, filter, and exact author
order.
- Full fail-on-warning build was attempted; remaining warnings were
unrelated optional autodoc environment warnings.

### Before your PR is "Ready for review"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in CONTRIBUTING.md: N/A
- Did you write any new necessary tests?: N/A
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

Depends on #1971.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Added a comprehensive announcement introducing AutoQuantize for
gradient-based mixed-precision optimization.
- Documented sensitivity scoring, effective-bits cost modeling,
deployment-aware grouping, benchmark results, usage examples, future
plans, and references.
- Added the announcement to the documentation homepage with an August
24, 2026 release card from the Model Optimizer Team.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-08-26 21:49:45 +00:00
Chenhan D. Yu ddca53b4bf docs: add announcements landing page (#1971)
## Summary
- Add a JS/static announcements landing page for GitHub Pages
- Keep existing Sphinx docs available under the `api/` subpath
- Add PR-authored announcement posts with tags, search/filtering, image
support, and a DSpark vs Domino sample post

Jira: https://jirasw.nvidia.com/browse/OMNIML-5476

## Verification
- `python3 docs/build_site.py --output docs/build/html`
- Local preview checked at `http://127.0.0.1:8088/`

## Publishing approval
User explicitly approved publishing the `dspark-vs-domino` sample post
and copied image assets from `modelopt-site` to public GitHub in
`NVIDIA/Model-Optimizer`.

## Notes
- Full `uv run nox -s docs` was attempted locally, but dependency
setup/download did not complete in a reasonable time; CI should provide
the authoritative full docs build and PR Pages preview.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an interactive announcements hub to the documentation homepage
with date sorting, tag filtering, search, pagination, and empty-state
messaging.
* Improved announcement presentation with consistent cards, metadata,
typography, and interactive controls.

* **Documentation**
* Added announcements covering the GitHub Pages announcement hub and a
DSpark versus Domino comparison.
* Added guidance for authoring, reviewing, and discovering future
announcements.
* Updated documentation navigation to feature announcements while
keeping other sections accessible.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-08-13 21:04:54 +00:00
ZhiyuandClaude Opus 5 6261f854aa docs: rebuild the unified HF deployment support matrix from the deploy test suite (NVBug 6550792) (#2087)
### What does this PR do?

Type of change: documentation

Fixes [NVBug 6550792](https://nvbugspro.nvidia.com/bug/6550792) /
OMNIML-5693.

The **Unified HF Checkpoint Deployment Model Support Matrix** listed 9
model families and **no VLMs**, while
`tests/examples/hf_ptq/test_deploy.py` declares deployment cases for ~80
checkpoints across TRT-LLM, vLLM, and SGLang — including `Qwen2.5-VL`,
`Qwen3-VL-235B`, and `Nemotron-3-Nano-Omni`. QA (the filer) could not
use the doc to scope testing, and users could not tell what is actually
covered.

Filing also surfaced that the matrix lived in **three places that had
drifted apart**: only the `.rst` listed Qwen3-VL, only the README listed
Qwen3.5 MoE, and the skill reference had neither.

#### Changes

1. **Rebuilt the matrix in `docs/source/deployment/3_unified_hf.rst`**
from `test_deploy.py`, split into language models,
vision-language/multimodal, speculative decoding drafters, and
diffusion.

2. **Stated plainly what the matrix is and is not.** Review established
that the original "CI-validated" framing claimed more than the suite
substantiates, so a *What this matrix is based on* section now leads
with two limits:
- The suite is marked `release` and collects only under `--run-release`,
which **no workflow passes** — these are declared cases, not PR-gated
coverage.
- Each case is a **load-and-generate smoke check on the text path**: no
accuracy, no image/audio input, no diffusion output, no verification
that speculative decoding engages.

The legend follows from that: ✅ = declared in the suite, ⚠ = expected to
work but not a suite entry (or an entry that does not exercise the
feature the row names), `-` = not in the suite. Sections that would
otherwise over-read carry their own qualifiers — VLM rows are labelled
text-only smoke coverage, and Medusa and Wan 2.2 are ⚠ with the reason
stated.

3. **Removed the two duplicate copies**, replacing them with links, so
there is one table to maintain.

4. **Fixed stale prose**: the deployment tabs still claimed FP8-only
support on vLLM v0.6.5 and a source build of SGLang main from Jan 2025,
both contradicting the version table above them. The TRT-LLM floor moves
to v1.2.0, qualified as the oldest version stated rather than the oldest
that works.

5. **Dropped the Phi series** from the deployment matrix, following
#2115 (NVBug 6563509) and confirmation that Phi-4 is being deprecated.

### Usage

N/A — documentation only.

### Testing

- `docutils` parse of the modified `.rst`: no warnings or errors from
the new content; all 5 tables parse with every cell in the correct
column.
- Cell contents cross-checked against `test_deploy.py` by AST-parsing
the `ModelDeployerList(...)` calls rather than by eye; the scope caveats
were each verified against `tests/_test_utils/deploy_utils.py`.
- `pre-commit run --files …` passes; `build-docs` green.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — documentation only
- Did you get Claude approval on this PR?: ❌ — not yet run

### Additional Information

**Two known follow-ups, neither in scope here:**

1. **Nothing enforces that the doc matrix tracks `test_deploy.py`.**
Consolidating to one copy removes the three-way drift but not the
doc-vs-test drift; a generator plus a CI check would close it.
2. **The release deployment suite does not run in CI.** Wiring it into
per-backend release CI is what would let ✅ mean "verified to pass"
rather than "declared". That needs GPU capacity across three backends
and should be tracked on its own.

**For the filer (@Kenny Kang):** the ✅ cells are the scope the release
deploy suite declares, and `test_deploy.py` carries the checkpoint, TP
size, and minimum SM version per entry — but please read the legend
first, since those cases are not currently executed by CI.

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 01:37:07 +05:30
vishalpandya1990 87c9f8cf83 Update documentation guide for ONNX INT4 PTQ on Windows cuda13 host (#2022)
### What does this PR do?

Type of change: Documentation update

- Update documentation guide for ONNX INT4 PTQ on Windows cuda13 host -
mention about compatible onnxruntim-gpu and cupy-cuda13x packages.

### Testing

- Windows's onnx_ptq\genai_llm INT4 PTQ example with a 1B genai-cuda-ep
ONNX model + local doc building

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Clarified Windows CUDA prerequisites for calibration and
GPU-accelerated quantization.
* Added setup guidance for CUDA 12 and CUDA 13.x, including compatible
packages and cuDNN requirements.
* Expanded installation verification steps for CUDA, ONNX Runtime, and
CuPy.
* Updated the GenAI LLM example with CUDA version compatibility
guidance.

* **Enhancements**
* Added runtime logging of detected CUDA environment paths and version
details during quantization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: vipandya <vipandya@nvidia.com>
2026-07-28 11:42:47 +05:30
Gwena Cunha d984de3795 [6008361][ONNX][Quantization] Clarify autotune guidance (#1989)
### What does this PR do?

Type of change: documentation

This PR clarifies when users should use ONNX quantization with Autotune
enabled versus the direct Autotune entry point.

- Adds a warning to the Autotune guide explaining that direct Autotune
is a lower-level Q/DQ placement tool and does not replace calibrated
ONNX PTQ.
- Updates the ONNX quantization guide to show `autotune=True` in the
Python API and explain that it uses default Autotune settings.
- Updates ONNX PTQ example documentation to prefer `python -m
modelopt.onnx.quantization ... --autotune=<mode>` for accuracy-sensitive
PTQ from an unquantized model.
- Updates the direct Autotune CLI help text to point users back to the
full ONNX quantization workflow when calibration data and accuracy
validation are required.

### Usage

```python
N/A — documentation/help text change.
```

### Testing

- Ran `python -m py_compile
modelopt/onnx/quantization/autotune/__main__.py`.
- Built the Sphinx documentation with `python -m sphinx -b html
docs/source docs/build/html`; build succeeded. Remaining warnings are
from optional documentation imports and existing cross-reference labels.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Clarified that Direct Autotune is an advanced tool for Q/DQ placement
experiments, not a replacement for full calibrated ONNX quantization.
* Added guidance on when to use Direct Autotune versus the end-to-end
ONNX PTQ workflow.
* Documented the optional `autotune=True` setting, expected
calibration-time impact, and representative calibration data
requirements.
* Expanded links and guidance across ONNX PTQ examples and updated CLI
help text with the recommended workflow.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Gwenaelle Cunha Sergio <gcunhasergio@nvidia.com>
2026-07-23 21:17:49 +00:00
Chad Voegele 7d5d3f9046 MiniMax-M3 mixed MXFP8-base + NVFP4-experts PTQ export (#1806)
### What does this PR do?

Type of change: New example

Adds two workflows for producing MiniMax-M3 checkpoints with an MXFP8
language-model base and NVFP4 routed experts:

- A memory-bounded exporter that preserves the vendor MXFP8 base and
quantizes routed experts from BF16 one MoE layer at a time.
- A model-specific `hf_ptq.py` recipe that quantizes the complete BF16
model using MXFP8 for language-model linear layers and MSE-calibrated
NVFP4 for routed experts.

The routed-expert NVFP4 `input_scale` is fixed to 1.0. The vision
branch, routers, `lm_head`, and KV cache remain unquantized.

Supporting changes:

- Skip MSE `amax` calibration for MX formats, which do not use a global
scale.
- Discover decoder layers through the MiniMax-M3 VLM path
`model.model.language_model.layers`.
- Add focused tests for recipe precedence, MXFP8 MSE exclusion, and VLM
decoder discovery.
- Document the streaming exporter in `examples/minimax_m3/README.md` and
the model-specific recipe in the recipe guides.

### Usage

Quantize the complete BF16 model through `hf_ptq.py`:

```bash
python examples/hf_ptq/hf_ptq.py \
    --pyt_ckpt_path /models/minimax-m3-bf16 \
    --recipe huggingface/minimax_m3_vl/ptq/mxfp8_nvfp4_experts \
    --export_path /models/minimax-m3-mxfp8-nvfp4 \
    --use_seq_device_map \
    --gpu_max_mem_percentage 0.68 \
    --calib_size 1
```

Compose the vendor MXFP8 base with routed experts quantized from BF16:

```bash
python examples/minimax_m3/hf_ptq_mixed_mxfp8_nvfp4.py \
    --mxfp8_ckpt /models/minimax-m3-mxfp8 \
    --bf16_ckpt /models/minimax-m3-bf16 \
    --recipe huggingface/minimax_m3_vl/ptq/nvfp4_experts_only \
    --output_ckpt /models/minimax-m3-mxfp8-nvfp4 \
    --device cuda
```

### Testing

- Full unit suite: 2,939 passed, 15 skipped.
- All pre-commit hooks passed.
- The streaming exporter reproduced all 89,614 tensors in
`nvidia/MiniMax-M3-NVFP4` exactly.
- Full BF16 `hf_ptq.py` validation matched all 87,552 NVFP4 expert
tensors, 534 MXFP8 scales, and 21,888 expert input scales exactly.
- The remaining 92 reference differences are BF16 Q/K norm tensors whose
values already differ between the public BF16 and vendor MXFP8 source
checkpoints.
- Verified standard Hugging Face shard names and mixed-precision
metadata.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

The workflows were tested with PyTorch 26.05.

---------

Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
2026-07-20 13:50:52 -05:00
Keval Morabia 1b03381123 (Deps) Pin docs and test dependencies to avoid breaking CI on new packages (#1874)
Fix broken doc building CI and pin docs/test dependencies to avoid such
issues in future

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Improved several guide cross-references so links resolve more reliably
in the rendered docs.
* Updated references in the quantization guides for better navigation
between related topics.

* **Chores**
* Tightened versions for documentation and test tooling to improve
consistency in local and CI environments.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-07-01 21:53:46 +05:30
Keval MorabiaandClaude Opus 4.8 2fc352be2d Add VLM pruning and PTQ with image-text calibration (Megatron-Bridge) (#1792)
### What does this PR do?

Type of change: New feature

Adds **vision-language model (VLM) support** to the Megatron-Bridge
examples for both **Minitron pruning** (`prune_minitron.py`) and **PTQ**
(`quantize.py`). Only the **language model** is pruned/quantized — the
vision tower and vision→language projector are left in full precision —
and the full VLM is saved back. `hidden_size` is skipped for pruning
when it is shared with the vision→LM projector.

Supported VLMs (tested e2e): **Qwen{3,3.5}-VL** (dense; hybrid
GatedDeltaNet + gated attention) and **Gemma3-VL** (sliding/full
attention).

### Calibration (image-text)

Calibration is conditioned on real **image-text** data so the language
model's pruning importance / quantizer statistics see vision-conditioned
activations. The modality is inferred from `--calib_dataset_name`:

- an **image-text** dataset (default for VLMs,
`nemotron_vlm_dataset_v2`) drives the **full VLM forward**;
- a **text** dataset runs text-only calibration of the language model
(for text-vs-image ablations).

A shared `get_megatron_vlm_calibration_forward_loop` (built on
`megatron_prefill`) drives the full VLM forward over image-text pairs
from `vlm_dataset_utils` (`scienceqa`, `nemotron_vlm_dataset_v2`, with
config-driven subset/shard caps to bound downloads). It shards across
**data-parallel (DP)** ranks like the text loop (#1804); **context
parallelism (CP)** applies to text-only VLM calibration (the shared text
loop), not the multimodal forward — splitting the sequence would
misalign the merged vision embeddings.

### Results - Cosmos-Reason2-2B

Validated end-to-end on **Cosmos-Reason2-2B** (Qwen3-VL). Minitron NAS
prunes the language-model tower **1.72B → ~1.59B** (vision encoder +
projector frozen), top_k=1. Calibration data drives pruning importance;
image-text calibration runs the full VLM forward.

| Model | Calibration | MMLU | BLINK Rel-Depth | RealWorldQA |
|---|---|---|---|---|
| Baseline (1.72B) | — | 0.58 | 0.76 | 0.61 |
| Pruned (1.59B) | text (`nemotron-post-training-dataset-v2`) | 0.51\* |
~0.69 | ~0.57 |
| Pruned (1.59B) | image+text (`nemotron_vlm_dataset_v2`) | 0.49\* |
**0.77** | **0.61** |

\* Pruned MMLU on the 10% split (the pruning score function); baseline
MMLU is the full set. The VLM-benchmark numbers for the text row were
measured with a different text calibration set and are expected to be
similar for `nemotron-post-training-dataset-v2` (marked `~`).

> [!NOTE]
> These numbers come from short single runs on small eval splits — read
them for **high-level trends only**, not as exact values.

Takeaways: pruning the LM tower of a VLM works end-to-end. **Image-text
calibration** (this PR's feature) preserves the VLM benchmarks better
than text-only — BLINK Rel-Depth ~0.77 vs ~0.69 and RealWorldQA ~0.61 vs
~0.57, both close to the unpruned baseline (0.76 / 0.61) — which is the
motivation for calibrating on vision-conditioned activations.

### Results - Qwen3.5-9B

| Model                      | MMLU   | MMStar |
|----------------------------|:------:|:------:|
| Qwen3.5-9B      | 0.7003 | 0.6117 |
| Pruned-7B (text calib)        | 0.5527 | 0.4411 |
| Pruned-7B (image+text calib)  | 0.5107 | 0.3941 |

### Key changes

- `quantize.py`: quantizes the **root** model with non-LM (vision)
quantizers disabled, so the ModelOpt state lives on the root (required
by the Megatron save) while only the language model is quantized.
- `prune_minitron.py`: image-text (or text) calibration for VLM pruning
importance.
- Shared VLM calibration forward loop (`megatron_prefill`-based, unwraps
tuple outputs, DP-sharded) + `vlm_dataset_utils`.
- Tiny VLM test fixtures (Qwen3.5-VL, Gemma3-VL) with vision tokens
derived dynamically from the reference processor; VLM prune + quantize
example tests.
- README + CHANGELOG.

### Usage

```bash
# Prune the language model of a VLM (image-text calibration by default)
torchrun --nproc_per_node 2 prune_minitron.py \
    --pp_size 2 \
    --hf_model_name_or_path <vlm> \
    --prune_target_params 3e9 \
    --output_hf_path /tmp/vlm-pruned

# PTQ the language model of a VLM
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path <vlm> \
    --quant_cfg fp8 \
    --export_megatron_path /tmp/vlm-fp8-megatron
```

### Testing

- `test_prune_minitron.py::test_prune_minitron_vlm` — Gemma3-VL,
image-text (ScienceQA) calibration; full load → prune (depth + ffn) →
save → reload.
- `test_quantize_export.py::test_quantize_vlm` — Qwen3.5-VL, text
calibration; quantize LM → save Megatron checkpoint.
- LM regression tests (`test_prune_minitron`,
`test_quantize_and_export`) unchanged and passing.

### Not in scope

- **HF unified export of a quantized VLM** is not yet supported;
`export.py` saves the Megatron checkpoint only for VLMs (tracked by a
TODO in `export.py`). The recommended path is to route the megatron→HF
quant export through Megatron-Bridge's
`AutoBridge.export_hf_weights_quant(quantization_checker, quant_fn,
quant_block_size)`, which reuses the bridge's per-model mcore↔HF mapping
— covering Qwen3.5-VL / Gemma3-VL and the vision tower/projector (left
full precision) for free — so modelopt supplies only the checker +
pack/scale fn + `hf_quant_config` (KV-cache scales need a separate
path). This avoids re-authoring per-model mappings in modelopt (cf.
#1482's Qwen3-VL-only `mcore_qwen3vl.py`).

> [!NOTE]
> Qwen3.5-VL **MoE** is not tested e2e: the Megatron-Bridge weight
conversion expects packed (`gate_up_proj`) experts that transformers'
tiny checkpoint doesn't emit. MoE pruning itself is covered by
`test_mcore_qwen35_gdn_moe_pruning`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅

### Additional Information

Follow-up to the GatedDeltaNet/MLA/latent-MoE pruning PR (#1747).
Rebased on `main` to pick up CP/DP calibration (#1804); the VLM
calibration loop now shards across DP ranks the same way. `hidden_size`
pruning for VLMs (requires resizing the vision projector) is left for a
future PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added VLM-aware Minitron pruning and post-training quantization that
target only the language-model portion, keeping the vision
tower/projector in full precision.
* Calibration now auto-selects text vs image-text datasets based on
model type, with modality validation.
* Expanded Megatron-Core CP/DP guidance and introduced a `--cp_size`
flag in quantization examples.
* **Bug Fixes**
* Improved VLM generation/prefill output handling and made vocabulary
sizing more robust for VLM wrappers.
* **Tests / Documentation**
* Updated pruning/quantization docs and refreshed/added VLM-focused
tests.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-01 02:47:02 +05:30
ZhiyuandClaude Opus 4.8 f335459dc0 refactor(examples): rename llm_ptq → hf_ptq (symlink for back-compat) (#1759)
## What does this PR do?

**Type of change:** refactor / deprecation (examples)

Follow-up to #1705 (which consolidated `examples/vlm_ptq` into
`examples/llm_ptq`). Since that example now covers Hugging Face **LLM
and VLM** PTQ, the `llm_ptq` name is a misnomer. This renames the
directory to `examples/hf_ptq` and leaves a relative symlink
`examples/llm_ptq → hf_ptq` so existing paths/commands keep working
during a deprecation window.

Requested by @kevalmorabia97 on #1705 (with the symlink-for-back-compat
approach), targeted for the **same 0.46 release** as the consolidation.

### Changes
- `git mv examples/llm_ptq → examples/hf_ptq` and
`tests/examples/llm_ptq → tests/examples/hf_ptq` (the CI runner maps the
matrix name to both `examples/<name>` and `tests/examples/<name>`).
- Add a tracked back-compat symlink `examples/llm_ptq → hf_ptq`.
- Update CI matrices and all repo **path references** (docs, READMEs,
agent skills, launcher/debugger tools, tests) from `llm_ptq` to
`hf_ptq`.
- Keep Python identifiers / test-util module names
(`run_llm_ptq_command`, `llm_ptq_utils`) — they name the LLM-PTQ task,
not the directory.
- Preserve the CODEOWNERS team slug
(`modelopt-examples-llm_ptq-codeowners`) and historical CHANGELOG
entries; add a CHANGELOG deprecation note.

### Back-compat caveats (inherent to git directory symlinks)
- ✅ Linux/macOS CLI usage and Python `cwd`/pytest resolution work
through the symlink.
- ⚠️ Windows git checkouts don't materialize symlinks by default (low
impact — this example is Linux-only in practice).
- ⚠️ GitHub web doesn't follow directory symlinks, so legacy external
deep-links to `examples/llm_ptq/...` won't navigate in. All **internal**
references are repointed to `hf_ptq`, so the symlink is only for legacy
external/CLI use.

### Usage (unchanged via symlink)
```bash
# New canonical path
cd examples/hf_ptq
scripts/huggingface_example.sh --model <hf_model> --quant fp8

# Old path still works (forwards via symlink)
cd examples/llm_ptq && scripts/huggingface_example.sh --model <hf_model> --quant fp8
```

### Testing
- `bash -n` on moved/edited shell scripts (new path + via symlink).
- `py_compile` on moved/edited Python; test re-export shim repointed to
`examples/hf_ptq/example_utils`.
- Verified git tracks `examples/llm_ptq` as a single symlink (mode
120000), not a duplicated tree (no pre-commit / pytest
double-processing).
- `pre-commit run` on all changed files passes.

### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (relative symlink keeps
`examples/llm_ptq` paths valid; see caveats above)
- Did you write any new necessary tests?: N/A (pure rename; existing
tests moved with the dir)
- Did you update Changelog?: ✅

### Additional Information
Follow-up (later release): remove the `examples/llm_ptq` symlink once
external references have migrated.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* PTQ guidance now directs to the unified Hugging Face PTQ flow,
including VLM quantization via the shared `--vlm` entry point.
* **Documentation**
* Updated README and guide links, references, and command snippets to
use `hf_ptq` (replacing `llm_ptq`).
* Deprecated and consolidated `vlm_ptq` into `hf_ptq`; removed
VILA/NVILA coverage from the Hugging Face PTQ examples.
* **Bug Fixes**
* Improved detection and routing so local/manual setup uses the correct
PTQ source.
* **Tests / Chores**
  * CI and example tests updated to run the `hf_ptq` variants.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-27 08:48:48 +00:00
Keval MorabiaandClaude Sonnet 4.6 33bfa8b1fe CI/Dev env bump (#1818)
### What does this PR do?

Type of change: chore

Bumps CI/dev tooling and test containers.

**Container bumps**
- NeMo test containers → 26.06
- TRT-LLM container → 1.3.0rc19
- transformers max version → 5.12

**Dev tooling bumps**
- ruff bump 0.12.11 → 0.15.18
- mypy 1.17.1 → 2.1.0: enable new defaults (`local_partial_types`,
`strict_bytes`); fix/narrow the errors newly surfaced by mypy 2.0 in 4
modules (rather than blanket-suppressing them); remove 2 stale `# type:
ignore` comments
- pre-commit 4.3.0 → 4.6.0
- sphinx 8.1 → 9.1 + sphinx-rtd-theme 3.0 → 3.1: add `suppress_warnings
= ["ref.python"]` to fix cross-reference ambiguity error new in sphinx
9.x
- trl fix for newly released 1.7 version

**Bug fixes surfaced by the bumps**
- sparsity (weight): make the weight mask DTensor-aware under FSDP. The
transformers→5.12 bump routes the HF Trainer FSDP optimizer-state save
through torch's DTensor-based `get_optimizer_state_dict`, which
triggered `aten.mul.Tensor got mixed torch.Tensor and DTensor` in the
dynamic `weight` getter. The mask is now distributed to the weight's
mesh/placements before masking, cached, and rebuilt only when the
sharding changes (invalidated on `set_mask`). Fixes the `llm_sparsity`
example test.

### Testing

- `pre-commit run --all-files` ✅ (including mypy 2.1.0)
- `nox -s docs` ✅
- `tests/unit/torch/sparsity` + `tests/unit/torch/nas` ✅
- `llm_sparsity` GPU example test (FSDP path) verified in CI

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ✅
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Summary

* **Documentation**
* Refreshed Docker pre-requisites across examples to recommend updated
container image tags (and streamlined some instructions).
* **Bug Fixes**
* Improved sparse weight mask handling for DTensor/FSDP by aligning and
caching distributed masks.
  * Made TensorRT engine byte retrieval return immutable `bytes`.
* Reduced Sphinx cross-reference warnings and tuned Transformers
compatibility warning thresholds.
* **Tests**
  * Increased default unit test timeout on Windows runners.
* **Chores**
* Updated CI workflow container tags and refreshed linting/typing/docs
version pins, plus related mypy configuration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-27 01:00:25 +05:30
Chenjie LuoandClaude Opus 4.8 56c4af2333 feat(recipes): add kv_fp8_cast variants for partial-NVFP4 and weight-only PTQ recipes (#1652)
### What does this PR do?

Type of change: new feature (recipes)

Several `general/ptq` recipe families shipped a data-driven FP8 KV-cache
(`-kv_fp8`) variant but lacked the constant-amax `kv_fp8_cast` companion
that `fp8_default` and `nvfp4_default` already have. This PR adds the
missing cast variants so every KV-quantizing (and the weight-only)
family offers the calibration-free FP8 KV-cache option:

- `general/ptq/nvfp4_experts_only-kv_fp8_cast`
- `general/ptq/nvfp4_mlp_only-kv_fp8_cast`
- `general/ptq/nvfp4_omlp_only-kv_fp8_cast`
- `general/ptq/nvfp4_weight_only-kv_fp8_cast`

Each new recipe composes the exact same model-quant config as its
existing sibling and swaps the `kv_fp8` unit for the shared
`kv_fp8_cast` unit (constant-amax FP8 KV cache; no KV calibration
forward pass). The docs guide table/tree and the changelog are updated
to match.

### Usage

```bash
python examples/llm_ptq/hf_ptq.py \
    --pyt_ckpt_path <model> \
    --recipe general/ptq/nvfp4_mlp_only-kv_fp8_cast
```

### Testing

Extended the built-in PTQ smoke test
`tests/unit/recipe/test_loader.py::test_load_recipe_all_builtins` with
the four new recipe paths; all four load into a valid
`ModelOptPTQRecipe` with a populated `quantize` section.

```
$ python -m pytest tests/unit/recipe/test_loader.py tests/unit/recipe/test_presets.py -q
180 passed
```

`pre-commit` (including the `validate modelopt recipes` hook) passes on
all changed files.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (additive — only new recipe
files)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ (extended the builtin recipe
smoke test)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ (not yet)

### Additional Information

The two weight-only families were discussed for scope;
`nvfp4_weight_only` is included (it already names a KV mode, `kv_fp16`),
while `int4_blockwise_weight_only` is intentionally left untouched since
it carries no `-kv_` composition.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added four new NVFP4 PTQ (Post-Training Quantization) recipe variants:
experts-only, MLP-only, OMLP-only, and weight-only configurations.
* All new recipes include FP8 KV-cache cast mode support for improved
inference performance.

* **Documentation**
* Updated built-in recipes guide with new NVFP4 recipe options and
repository layout.

* **Tests**
  * Expanded recipe loader test coverage for new recipe configurations.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 16:13:25 -07:00
Grzegorz K. KarchandKeval Morabia b98a59557a Add vLLM-based runtime statistics for subblock latency measurement (#1358)
### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Runtime-based latency optimization: collect vLLM-measured inference
latency to constrain optimization.

* **Configuration**
* New runtime config/template for Llama-3.1-8B pruning (runtime stats
enabled, NCCL timeout templating, MIP target-latency).
* Validation sample defaults adjusted (one flow: 128 → 8; runtime flow
uses 128).
  * Human constraint key renamed to target_latency_seconds.

* **Documentation**
* README section describing runtime-based latency optimization setup and
usage.

* **Tests**
  * Added GPU end-to-end test for runtime stats collection.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1358?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Grzegorz Karch <gkarch@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-06-08 19:22:55 +00:00
hychiangandClaude Sonnet 4.6 f0d2237cbc Add Qwen3VL MCore Export support from PR 895 (#1482)
# [Megatron Export] Add Qwen3-VL mcore ↔ HF weight mapping

> This PR is duplicated from [PR
#895](https://github.com/NVIDIA/Model-Optimizer/pull/895).
> The original branch source is no longer available; this new branch
carries the same changes forward.

## What does this PR do?

**New feature:** Add Qwen3-VL (Vision-Language) model support to the
Megatron Core export/import
plugin, enabling HuggingFace-to-mcore weight conversion for PTQ/QAT/QAD
workflows.

### Overview

Qwen3-VL has a different weight structure from Qwen3 text-only models:

- Language model weights are under `model.language_model.` prefix (not
`model.`)
- Visual encoder weights are under `model.visual.` prefix
- `lm_head` is at root level, not nested under `language_model`

### What changed

| File | Change |
|---|---|
| `modelopt/torch/export/plugins/mcore_qwen3vl.py` | New plugin: derives
Qwen3-VL mcore↔HF mapping by rewriting `model.*` →
`model.language_model.*` on top of the existing Qwen3 dense rules;
`lm_head.` is intentionally left unchanged |
| `modelopt/torch/export/plugins/mcore_common.py` | Registers
`Qwen3VLForConditionalGeneration` in `all_mcore_hf_export_mapping` and
`all_mcore_hf_import_mapping` |
| `modelopt/torch/export/plugins/hf_checkpoint_utils.py` | Generalized
`load_multimodal_components` with a `prefixes` parameter; sharded
checkpoints now scan all shards (not just the first) |
| `modelopt/torch/export/unified_export_megatron.py` |
`save_pretrained`: added Qwen3-VL branch that copies `model.visual.*`
vision-encoder weights from the original HF checkpoint into the exported
directory, producing a complete, loadable checkpoint |
| `tests/_test_utils/torch/transformers_models.py` | Added
`get_tiny_qwen3vl` / `create_tiny_qwen3vl_dir` helpers; Qwen3VL classes
are lazy-imported inside the function to avoid collection failures on
older transformers builds |
| `tests/gpu_megatron/torch/export/test_unified_export_megatron.py` |
Integrated Qwen3-VL export/import tests into the existing
`test_unified_export_megatron` / `test_unified_import_megatron`
parametrized suites; removed standalone `test_mcore_qwen3vl.py` |
| `docs/source/deployment/3_unified_hf.rst` | Added Qwen3-VL (FP8 /
NVFP4) to the deployment support matrix for TensorRT-LLM |

### Workflow coverage

| Step | Status | Files |
|---|---|---|
| 1. Quantize Qwen3-VL with `hf_ptq` | ✅ existing | — |
| 2. Export quantized mcore → HF | ✅ this PR |
`plugins/mcore_qwen3vl.py` (weight name mapping),
`unified_export_megatron.py` (export path) |
| 3. Vision-encoder weights merged into export dir | ✅ this PR |
`plugins/hf_checkpoint_utils.py` (`load_multimodal_components` with
`prefixes`), `unified_export_megatron.py` (calls it when `arch ==
"Qwen3VLForConditionalGeneration"`) |
| 4. Import HF checkpoint back to mcore | ✅ this PR |
`plugins/mcore_qwen3vl.py` (same mapping, reverse direction),
`unified_export_megatron.py` (import path) |

### Design notes

- **MoE not supported**: `Qwen3VLMoeForConditionalGeneration` stores
expert weights as
3-D tensors (`mlp.experts.gate_up_proj`, `mlp.experts.down_proj`) that
require a
dedicated fused-expert mapping. A `NotImplementedError` comment in the
plugin
  documents this explicitly.
- **`copy.deepcopy` on `func_kwargs`**: each mapping entry gets its own
copy to
prevent shared-dict mutation when both Qwen3 and Qwen3-VL rules are
loaded.
- **`prefixes` parameter on `load_multimodal_components`**:
backward-compatible default
preserves existing LLaVA behaviour (`"multi_modal_projector"`,
`"vision_model"`);
  Qwen3-VL callers pass `("model.visual.",)`.
- **Sharded checkpoint scan**: the old code only looked in the first
shard. The
Qwen3-VL vision encoder can span multiple shards, so all shards are now
scanned.

## Usage

From the [Megatron-LM PR
comment](https://github.com/NVIDIA/Megatron-LM/pull/3444#issuecomment-3911271713):
> Qwen3VL is supported within
[Megatron-Bridge](https://github.com/NVIDIA-NeMo/Megatron-Bridge), and
pretraining and PEFT recipes for Qwen3VL are
[here](https://github.com/NVIDIA-NeMo/Megatron-Bridge/blob/main/src/megatron/bridge/recipes/qwen_vl/qwen3_vl.py)
and the core code logic
[here](https://github.com/NVIDIA-NeMo/Megatron-Bridge/tree/main/src/megatron/bridge/models/qwen_vl).

Create
`Megatron-LM/examples/post_training/modelopt/conf/Qwen/Qwen3-VL-8B-Instruct.sh`:

```bash
#!/bin/bash
# Qwen3-VL-8B-Instruct text-model config for Megatron-LM import/quantize.
#
# Text-model dimensions are identical to Qwen3-8B (4096 hidden, 36 layers,
# 32 heads, GQA=8).  Differences: rope_theta=5000000, checkpoint path uses
# model.language_model.* prefix (handled by mcore_qwen3vl plugin).

if [ -z ${HF_MODEL_CKPT} ]; then
    HF_MODEL_CKPT=Qwen/Qwen3-VL-8B-Instruct
    TOKENIZER_MODEL=Qwen/Qwen3-VL-8B-Instruct
else
    TOKENIZER_MODEL=${HF_MODEL_CKPT}
fi

MODEL_ARGS=" \
    --save-interval 100000 \
    --micro-batch-size 1 \
    --bf16 \
    --no-masked-softmax-fusion \
    --disable-bias-linear \
    --untie-embeddings-and-output-weights \
    --position-embedding-type rope \
    --no-rope-fusion \
    --normalization RMSNorm \
    --swiglu \
    --num-layers 36 \
    --hidden-size 4096 \
    --ffn-hidden-size 12288 \
    --num-attention-heads 32 \
    --group-query-attention \
    --num-query-groups 8 \
    --kv-channels 128 \
    --qk-layernorm \
    --seq-length 4096 \
    --max-position-embeddings 262144 \
    --tokenizer-type HuggingFaceTokenizer \
    --make-vocab-size-divisible-by 1187 \
    --use-mcore-models \
    --rotary-percent 1.0 \
    --rotary-base 5000000 \
    --no-bias-swiglu-fusion \
"
```

Import Qwen3-VL from HuggingFace to MCore (local, requires GPUs):

```bash
MLM_MODEL_CFG=Qwen/Qwen3-VL-8B-Instruct \
HF_MODEL_CKPT=Qwen/Qwen3-VL-8B-Instruct \
MLM_MODEL_SAVE=/tmp/qwen3vl_mcore \
TP=1 \
bash Megatron-LM/examples/post_training/modelopt/convert.sh Qwen/Qwen3-VL-8B-Instruct
```

Quantize (PTQ via Megatron-LM path):

```bash
MLM_MODEL_CFG=Qwen/Qwen3-VL-8B-Instruct \
HF_MODEL_CKPT=Qwen/Qwen3-VL-8B-Instruct \
QUANT_CFG=NVFP4_DEFAULT_CFG \
TP=4 \
bash Megatron-LM/examples/post_training/modelopt/quantize.sh Qwen/Qwen3-VL-8B-Instruct
```

## Testing

- Verified round-trip import/export with Qwen3-VL-8B-Instruct with the
example usage above
- Unit/GPU tests covering:
  - Registration in global export/import mappings
- Import mapping: dense keys, `model.language_model.` prefix, `lm_head.`
at root, `QKVMerging`, `GatedMLPMerging`, `REPLICATE` for layernorms, TP
sharding configs
- Export mapping: `QKVSlicing`, `GatedMLPSlicing`, no `parallel_config`
  - Import/export symmetry: same mcore keys, matching HF prefixes
- Qwen3-VL vs Qwen3 difference: same keys, VL adds `language_model.`
prefix, `lm_head` unchanged

## Before your PR is "Ready for review"

- Is this change backward compatible?: Yes, additive only
- Did you write any new necessary tests?: Yes,
`tests/gpu_megatron/torch/export/test_unified_export_megatron.py`
- Did you add or update any necessary documentation? Yes, see
`docs/source/deployment/3_unified_hf.rst`
- Did you update Changelog? Yes, see `CHANGELOG.rst`

## Additional Information

Companion Megatron-LM PR adds `Qwen3VLModel`, `Qwen3VLDataset`, and
`pretrain_qwenvl.py`.
See: https://github.com/NVIDIA/Megatron-LM/pull/3444

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: hychiang <hungyuehc@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-01 14:29:40 -07:00
Shengliang Xu 04f58166ab [OMNIML-3707] Model-specific PTQ recipes bootstrap (#1506)
### What does this PR do?

Type of change: new feature

Replaces the hardcoded model-type branches in `examples/llm_ptq/` with
opt-in declarative **model-specific recipes** under
`modelopt_recipes/huggingface/<model_type>/ptq/`. Any adjustment
specific to a model type or instance must live in that model's recipe —
there is no implicit model-specific path anymore. Users select a model's
recipe with `--recipe huggingface/<model_type>/ptq/<recipe>`; users on
the plain `--qformat` path get only the generic numerics.

What moved out of Python
(`examples/llm_ptq/example_utils.py::build_quant_cfg` and
`examples/llm_ptq/hf_ptq.py::mono_quantize`):

- **gemma / mpt** `w4a8_awq` → `awq_lite` with `alpha_step=1` (coarser
search to avoid TRT-LLM overflow).
- **gemma** `int8_sq` → SmoothQuant `alpha=0.5` (default `1.0` regresses
Gemma 7B).
- **phi4mm** → disable `*speech*`, `*audio*`, `*image*`, `*vision*`
(quantize only the language model).
- **Nemotron VL** → disable `*vision*`, `*image*`, `*radio*`,
`*visual*`, `*encoder*`, `*model_encoder*` (quantize only the decoder).

What stayed in Python:

- MTP dynamic layer exclusion in `hf_ptq.py` (depends on
runtime-detected layer indices).
- `is_nemotron_vl(full_model)` detection itself, which still drives the
VLM calibration loop and the post-quantize `full_model` update — only
the `quant_cfg` adjustment it triggered moved into the Nemotron VL
recipe.

`multinode_ptq.py` shares the same `build_quant_cfg` call site and was
updated to match the new 2/3-arg signature; multinode users on
`--qformat` get the generic numerics (no `--recipe` plumbing in
multinode yet, so model-specific recipes are only reachable via
`hf_ptq.py`).

Already-YAML recipes that were elsewhere in the tree are relocated into
the same `huggingface/<model_type>/ptq/` layout so all model-specific
recipes live under one convention:

- **Step3.5-Flash** — moved from
`modelopt_recipes/huggingface/step3p5/Step3.5-Flash/` to
`huggingface/step3p5/Step3.5-Flash/ptq/` to match the `<model>/ptq/`
convention.
- **Qwen3.5 / Qwen3.6** — moved from
`modelopt_recipes/models/Qwen3.5-Qwen3.6/w4a16.yaml` to per-model_type
folders, anchored on the HuggingFace `model_type` (verified against
transformers 5.8.1 + HF model hub `config.json` for `Qwen/Qwen3.6-27B`,
`Qwen/Qwen3.6-35B-A3B`, `nvidia/Qwen3.5-397B-A17B-NVFP4`):
- `huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml` —
dense `qwen3_5`
- `huggingface/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml` —
`qwen3_5_moe`
- Both wrappers `$import` the shared `quant_cfg` snippet
`huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg.yaml`
(one source of truth; the two model_types share the same hybrid
linear-attention + softmax-attention architecture so the rules apply
identically).

Full recipe layout (`modelopt_recipes/huggingface/`):

```
gemma/ptq/{w4a8_awq,int8_sq}-kv_fp8_cast.yaml
mpt/ptq/w4a8_awq-kv_fp8_cast.yaml
phi4mm/ptq/{disabled_quantizers,nvfp4-kv_fp8_cast}.yaml
nemotron_vl/ptq/{disabled_quantizers,nvfp4-kv_fp8_cast}.yaml
qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast{,.quant_cfg}.yaml
qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml
step3p5/Step3.5-Flash/ptq/nvfp4-mlp-only.yaml
```

All recipes ship with FP8 KV-cache cast (`kv_fp8_cast`). For phi4mm and
nemotron_vl, `disabled_quantizers.yaml` is a multi-document list unit
that `$import`s the standard `default_disabled_quantizers` exclusions
and appends the model-specific ones — so each recipe imports a single
disabled-quantizer slot instead of layering two, with no duplication in
YAML. Each `ptq/` folder has a `README.md` describing exactly what is
model-specific.

### Usage

```bash
# Gemma W4A8 AWQ with the Gemma-specific algorithm tuning + FP8 KV cache:
python examples/llm_ptq/hf_ptq.py \
  --pyt_ckpt_path google/gemma-7b \
  --recipe huggingface/gemma/ptq/w4a8_awq-kv_fp8_cast \
  --export_path ./out

# Nemotron VL with vision branches excluded automatically:
python examples/llm_ptq/hf_ptq.py \
  --pyt_ckpt_path nvidia/<nemotron-vl-model> \
  --recipe huggingface/nemotron_vl/ptq/nvfp4-kv_fp8_cast \
  --export_path ./out
```

### Testing

- Pre-commit recipe validator
(`tools/precommit/check_modelopt_recipes.py`) loads every new recipe via
`load_recipe()` — passes for all new YAMLs (gemma/mpt/phi4mm/nemotron_vl
recipes + phi4mm/nemotron_vl `disabled_quantizers` snippets + qwen3_5 /
qwen3_5_moe recipe wrappers + the shared
`w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg` snippet + Step3.5-Flash
relocation).
- For qwen3_5 / qwen3_5_moe specifically, `load_recipe(...)` on both
wrappers produces an identical 33-entry resolved `quant_cfg`, confirming
the shared snippet is the single source of truth.
- `yamlfmt` + `markdownlint` + `bandit` + license-insertion hooks all
pass.
- No tests reference the removed `build_quant_cfg(qformat, ...,
model_type, ...)` signature; the only call sites (`hf_ptq.py`,
`multinode_ptq.py`) were updated to the new 2/3-arg form.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — users who relied on
**automatic** model-specific quant_cfg behavior via `--qformat`
(gemma/mpt AWQ, gemma SmoothQuant, phi4mm exclusions, Nemotron VL
exclusions) now need to pass `--recipe
huggingface/<model_type>/ptq/<recipe>` to apply the model's recipe. The
flag itself is unchanged; only the implicit behavior was removed.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ — relies on the existing
pre-commit recipe validator that loads each new YAML.
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ❌


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added many model-specific PTQ recipes (Gemma, MPT, Nemotron VL,
Phi‑4‑Multimodal, Qwen3.5, Qwen3.5‑MoE) and support for AWQ block-size
and MoE calibration ratio in quantization options.

* **Documentation**
* Expanded READMEs and changelog to document recipe locations, layout,
and how to opt into model-specific PTQ recipes.

* **Refactor**
* Model-specific PTQ tweaks moved to opt‑in recipes; default behavior
uses generic numerics.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1506?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-05-22 13:39:32 -07:00
Shengliang Xu 2b8defc14b [OMNIML-4158] ModelOpt config system documentation (#1472)
### What does this PR do?

Type of change: documentation

Adds a new Sphinx guide for the ModelOpt config system and updates two
adjacent guides to match the current APIs.

New file:

- `docs/source/guides/11_config_system.rst` — covers the general config
contract and semantics rather than any single consumer:
- `ModeloptBaseConfig` schemas, validation boundaries, and
`MutableMapping` access semantics
  - YAML loading flow, effective-schema selection, and persistence
  - Checkpoint persistence
- Composable YAML via `imports` / `$import`, including the dict / list /
multi-document forms and the type-directed list splice-vs-append rule

The guide also explains *why* ModelOpt uses a small YAML DSL for
composition (instead of OmegaConf alone, Python factories, or hard-coded
registries): config files stay self-describing, reusable fragments can
be authored as YAML, and resolved values are still validated against the
Python schemas. Recipes are presented as one application of the shared
config system; recipe-specific authoring continues to live in the
existing recipes guide.

Updates to existing guides:

- `docs/source/guides/10_recipes.rst` — `RecipeMetadataConfig` and
`QuantizeConfig` are described as `ModeloptBaseConfig` subclasses;
`quant_cfg` entries are validated `QuantizerCfgEntry` instances after
loading; the schema-comment section now explains that validation runs
against the effective schema (in-file comment or `schema_type=`
argument, with the argument winning).
- `docs/source/guides/_quant_cfg.rst` — entries described as
`QuantizerCfgEntry` Pydantic instances (with dict input accepted and
normalized via `mode="before"` validator); `cfg:` field typed as
`QuantizerAttributeConfig` rather than a free-form dict.

### Usage

`load_config` resolves YAML composition (`imports` / `$import`) and,
when an effective schema is in scope (either via `schema_type=` or a `#
modelopt-schema:` comment), returns a validated schema instance
directly:

```python
from modelopt.recipe import load_config
from modelopt.torch.quantization.config import QuantizeConfig

cfg = load_config("configs/ptq/presets/model/fp8", schema_type=QuantizeConfig)
# cfg is a validated QuantizeConfig instance.
resolved = cfg.model_dump()
```

### Testing

- Docs-only change; no code paths touched.
- Pre-commit hooks pass (RST formatting checks).

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.). Docs-only change — no executable
code paths affected.

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (docs-only)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (docs-only)
- Did you get Claude approval on this PR?: ✅

### Additional Information

N/A

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-05-18 18:16:43 -07:00
realAsma 62401e16df fix: layerwise calibration backward-compat, recipe split, batch-size guard (#1310)
## Summary

Follow-up to #1251 (which renamed `use_sequential` → `layerwise`). Three
related fixes bundled:

1. **Backward-compatible config loading.** PTQ checkpoints saved before
#1251 store the
legacy `use_sequential` key in the calibration-algorithm config, so
loading them now
raises `ValidationError: Extra inputs are not permitted
(use_sequential)` because
`QuantizeAlgorithmConfig` uses `extra='forbid'`. Accept `use_sequential`
as an alias
for `layerwise` via `AliasChoices`. The field still serializes as
`layerwise`, so
   round-trips through the current schema are clean.

2. **Recipe split.** `nvfp4_experts_only-fp8_kv` previously enabled
layerwise calibration
by default, which changes the calibration flow materially. Split into
two recipes:
   - `nvfp4_experts_only-fp8_kv.yaml` — default (no layerwise)
   - `nvfp4_experts_only-fp8_kv_layerwise.yaml` — layerwise variant

3. **`hf_ptq` batch-size guard.** Auto batch-size detection is not
supported together
with layerwise calibration. Default to `batch_size=1` when layerwise is
enabled and
   the user hasn't set a batch size explicitly.

Originally reported by Jenny Chen while resuming a PTQ checkpoint via
`restore_sharded_modelopt_state`:

```
pydantic_core._pydantic_core.ValidationError: 1 validation error for MaxCalibConfig
use_sequential
  Extra inputs are not permitted [type=extra_forbidden, input_value=False, input_type=bool]
```

## Test plan

- [x] `tests/unit/torch/quantization/test_config_validation.py` — legacy
alias accepted, current name accepted, dump serializes under current
name, `extra='forbid'` still rejects unknown keys.
- [x] `pre-commit run` — clean.

### Before your PR is *Ready for review*

- Is this change backward compatible?: ✅ (restores compatibility for
pre-#1251 checkpoints)
- New PIP dependency: N/A
- New necessary tests: ✅
- Changelog update: N/A (bug fix)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added new PTQ recipe for efficient layerwise calibration of large
models.
  * Automatic batch size optimization for layerwise calibration recipes.
  * Backward compatibility support for legacy input naming conventions.

* **Documentation**
* Updated recipe guides and changelog with new layerwise calibration
recipe.

* **Tests**
  * Added validation tests for configuration compatibility.

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1310)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-05-12 20:17:21 +00:00
Keval Morabiaandclaude[bot] 2ce745a92e Deprecate gradnas pruning and bert example (#1427)
### What does this PR do?

Type of change: Deprecation of dead code <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->

Deprecation warning already added in 0.44 as per 1-release deprecation
policy

GradNAS only works for Bert and GPT-J and we dont actively maintain it
or test it. Keeping it creates an expectation that it works plus it adds
one more option for user to choose from. We already have much better
pruning algorithms (Minitron and Puzzletron) for LLM pruning already
hence removing GradNas.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ No but we dont have any users
of this feature either <!--- If ❌, explain why. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Deprecation**
  * GradNAS pruning algorithm deprecated; related examples removed.
* **Documentation**
* Pruning and NAS guides and changelog updated to focus on Minitron and
FastNAS; GradNAS references removed.
* **Chores**
  * Chained-optimizations example and scripts removed.
* Ownership mappings updated for README and examples; license insertion
now applies to a previously excluded example file.
* **Tests**
* Multiple unit tests and test utilities related to GradNAS/transformer
NAS removed.

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1427)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
2026-05-12 22:32:22 +05:30
Shengliang Xu f34f488a83 Add a general composable $import system for YAML configs, and use it to implement composable recipes (#1253)
### What does this PR do?

Type of change: New feature

Adds a general composable YAML config loading layer for ModelOpt configs
and recipes. YAML remains the source of truth for configuration data,
while Python/Pydantic-compatible types provide schema validation at load
time. This PR uses that loader to de-duplicate PTQ recipes, introduce
reusable config snippets/presets, and start migrating selected hardcoded
quantization presets to YAML.

#### Problem

1. Built-in PTQ recipes duplicated numeric format definitions, KV-cache
entries, and the default quantizer exclusion list.
2. YAML snippets were reusable only by convention; they did not declare
or validate the schema they were meant to satisfy.
3. Loading YAML-backed quantization presets from
`modelopt.torch.quantization.config` could not depend on
`modelopt.recipe` without creating circular imports.
4. Directory-format recipes exposed import-resolution details in the
recipe loader and used `recipe.yaml` plus nested `metadata:` in a way
that made metadata handling inconsistent.

#### Solution

**Shared YAML config loader**

- Adds `modelopt.torch.opt.config_loader` as the low-level loader used
by both `modelopt.recipe` and `modelopt.torch.quantization.config`.
- Keeps the public `modelopt.recipe.load_config()` entry point, while
removing the private `modelopt/recipe/_config_loader.py` shim.
- Handles YAML loading, built-in/filesystem path resolution, suffix
probing, `ExMy` conversion for `num_bits` / `scale_bits`, `$import`
expansion, and schema validation.
- Lives below `modelopt.recipe` in the dependency graph to avoid
circular imports from quantization config code.

**Composable `$import` system**

Recipes and snippets can declare an `imports` mapping, then reference
entries with `{$import: name}`.

`$import` semantics:

- **Dict value**: replaced with the imported dict. Multiple imports are
supported with ordered precedence; inline keys override imported keys.
- **List entry**: schema-driven behavior for strongly typed lists. If
the snippet schema matches the containing list type, the imported list
is spliced. If the snippet schema matches the list element type, the
imported element is appended. Other schema combinations are rejected.
- **Multi-document YAML**: supports snippets that need an `imports`
header plus a list body.
- **Recursive and scoped**: snippets can import other snippets; import
names are scoped per file.
- **Cycle detection**: circular imports report a clear error.

**Snippet schema validation**

- Every reusable snippet referenced through `imports` must declare a `#
modelopt-schema: ...` preamble.
- Snippets are validated after nested imports are resolved.
- Schema paths are restricted to the `modelopt.` package and may be
Pydantic models, `TypedDict` classes, or explicitly typed container
aliases such as `list[QuantizerCfgEntry]`.
- Untyped list imports are rejected so list append/splice behavior stays
strongly typed.

**Recipe model and directory recipe cleanup**

- `ModelOptRecipeBase` now owns a `metadata: RecipeMetadataConfig`
field.
- `ModelOptPTQRecipe` is the PTQ recipe schema; the overlapping
YAML-specific PTQ config class was removed.
- Directory recipes now use `metadata.yaml` / `metadata.yml` for
top-level metadata fields, plus section files such as `quantize.yaml`.
- Directory recipe loading now delegates import resolution to
`load_config()` instead of manually using raw config loading.

**Config snippet and preset library**

Adds reusable snippets under `modelopt_recipes/configs/`:

- `numerics/`: `fp8`, `nvfp4`, `nvfp4_static`
- `ptq/units/`: `base_disable_all`, `default_disabled_quantizers`,
`w8a8_fp8_fp8`, `w4a4_nvfp4_nvfp4`, `kv_fp8`, `kv_fp8_cast`,
`kv_nvfp4_cast`
- `ptq/presets/`: YAML presets for `FP8_DEFAULT_CFG` and `FP8_KV_CFG`

`FP8_DEFAULT_CFG` and `FP8_KV_CFG` now load from YAML presets via
`load_config()`.

**Recipe migration and naming**

- General PTQ recipes now use shared imports instead of repeating the
same quantizer fragments inline.
- General PTQ recipe paths were renamed to KV-first naming, for example:
  - `general/ptq/fp8_default-fp8_kv` -> `general/ptq/fp8_default-kv_fp8`
- `general/ptq/fp8_default-fp8_cast_kv` ->
`general/ptq/fp8_default-kv_fp8_cast`
- `general/ptq/nvfp4_default-none_kv_gptq` ->
`general/ptq/nvfp4_default-kv_none-gptq`
- `general/ptq/nvfp4_default-nvfp4_cast_kv` ->
`general/ptq/nvfp4_default-kv_nvfp4_cast`
- Example docs and `examples/llm_ptq/hf_ptq.py --recipe` help text were
updated to use the new paths.

**Pre-commit and documentation**

- Recipe validation accepts `$import` entries and handles directory
recipes using `metadata.yaml`.
- The recipe validation hook skips `modelopt_recipes/configs/` because
those files are reusable snippets, not full recipes.
- `docs/source/guides/10_recipes.rst` now documents imports, schema
modelines, list append/splice semantics, built-in snippets, built-in
recipe paths, directory recipes, and the current recipe data model.

#### Backward compatibility

- Existing inline YAML recipes without `$import` continue to load.
- `modelopt.recipe.load_config()` remains public.
- The built-in recipe path renames are user-visible; callers should
update recipe path strings to the KV-first names listed above.

#### Testing

- `pytest tests/unit/recipe/test_loader.py -q` - 90 passed
- `python tools/precommit/check_modelopt_recipes.py ...` for the renamed
built-in PTQ recipes
- `pre-commit run mypy --files
modelopt/onnx/llm_export_utils/quantization_utils.py`
- `python -m py_compile examples/llm_ptq/hf_ptq.py`
- `git diff --check`

### Before your PR is "Ready for review"

- Is this change backward compatible?: Partially. Loader/API behavior is
compatible for existing inline YAML recipes, but built-in recipe path
names were renamed to KV-first paths.
- Did you write any new necessary tests?: Yes.
- Did you update Changelog?: Yes.

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-05-05 17:01:28 -07:00
Keval Morabiaandcoderabbitai[bot] 70546bdd6a Enable Python 3.14 wheel support to unblock NGC PyTorch container testing on Ubuntu 26.04 + Python 3.14 (#1386)
Ubuntu 26.04 is here and very soon, NVIDIA PyTorch containers will ship
with Python 3.14 requiring us to enable untested support to unblock them

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
  * DFlash offline speculative decoding training
  * MXFP4→NVFP4 weight conversion support
  * Shared hidden-state dump utilities
  * Updated DeepSeek PTQ calibration defaults

* **Chores**
  * Added Python 3.14 support; updated Python requirement to <3.15

* **Documentation**
  * Updated installation documentation for Python version compatibility

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-05-04 21:56:04 +00:00
ynankani 9bb917d57c [BUG6108338] Update windows documentation for onnxruntime quantization with Cuda13.x (#1368)
### What does this PR do?

Type of change: ? documentation

<!-- Details about the change. -->
Update windows documentation for onnxruntime quantization with Cuda13.x




### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?:  N/A <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!---
Mandatory -->
- Did you write any new necessary tests?: N/A <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->




<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated Windows installation guide with CUDA 13.x-specific setup
instructions for GPU-accelerated dependencies, including CuPy and ONNX
Runtime configuration with nightly builds.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: ynankani <ynankani@nvidia.com>
2026-04-29 16:15:40 +05:30
Keval MorabiaandClaude Sonnet 4.6 e56682e34a docs: update installation pages with legal-approved license notices (#1322)
## Summary

- Replaces the old pip license notice ("Please review the license terms
of ModelOpt and any dependencies before use") with the Legal-approved
wording: "Model Optimizer will download and install additional
third-party open source software projects. Review the license terms of
these open source projects before use."
- Adds a generic container license review notice ("Before pulling and
using the container images, please review their respective license
terms.") to the Linux installation doc (Docker tab) and README.
- Adds a `.. note::` with the pip notice to the Windows installation
page (covers both standalone and Olive child pages).
- Expands the README container section to explicitly list all four
recommended NVIDIA container images (`pytorch`, `nemo`, `tensorrt-llm`,
`tensorrt`).

## Test plan

- [x] Verify rendered docs look correct (`nox -s docs`)
- [x] Confirm legal notices appear in Linux, Windows, and README install
sections

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated installation guides with explicit references to supported
NVIDIA container images (PyTorch, NeMo, TensorRT-LLM and variants),
clarified pre-installed Model Optimizer in some images, and added notes
to review each container’s license terms; clarified conditional
environment setup wording and local install license guidance.
* **Chores**
  * Updated project license header year.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-22 22:49:23 +05:30
361f7e391b Merge puzzletron compression algorithm (#1121)
### What does this PR do?

Implement puzzletron compression algorithm based on Puzzle paper
(https://arxiv.org/abs/2411.19146)

<details>
<summary> Th list of reviewed and merged MRs that resulted in the
feature/puzzletron branch</summary>

Merging dkorzekwa/any_model to feature/puzzletron

[Add anymodel directories to feature/puzzletron by danielkorzekwa · Pull
Request #974 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/974)
- merged

[Draft: anymodel activation scoring by danielkorzekwa · Pull Request
#989 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/989)
- merged

[Draft: Merge anymodel pruning by danielkorzekwa · Pull Request #990 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/990/)
- merged

[Draft: Merging anymodel:build_library_and_stats by danielkorzekwa ·
Pull Request #993 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/993)
- merged

[Dkorzekwa/any model calc one block scores by danielkorzekwa · Pull
Request #994 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/994)
- merged

[Draft: merge any_model: mip_and_realize_models by danielkorzekwa · Pull
Request #995 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/995)
- merged

[Dkorzekwa/any model other modeqls by danielkorztiekwa · Pull Request
#1007 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1007/)
- merged

PR to 1007: https://github.com/NVIDIA/Model-Optimizer/pull/1039 - merged

[Dkorzekwa/anymodel gptoss by danielkorzekwa · Pull Request #1020 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1020)
- merged

[Merge any_model tutorial by danielkorzekwa · Pull Request #1035 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1035)
- merged

[Merge mbridge distillation for any_model by danielkorzekwa · Pull
Request #1036 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1036)
- merged

[MR branch for the remaining difference between dkorzekwa/any_model an…
by danielkorzekwa · Pull Request #1047 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1047)
- merged

[Dkorzekwa/decilm hf code cleanup by danielkorzekwa · Pull Request #1071
·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1071)
- merged

[Dkorzekwa/decilm hf code cleanup 2 by danielkorzekwa · Pull Request
#1073 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1073)
- merged

[Dkorzekwa/anymodel subblock stats by danielkorzekwa · Pull Request
#1085 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1085)
- merged

[Dkorzekwa/anymodel subblock stats nodecilm by danielkorzekwa · Pull
Request #1102 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1102)
- merged

[Dkorzekwa/decilm cleanup post subblockstats by danielkorzekwa · Pull
Request #1103 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1103)
- merged

[code clean up by danielkorzekwa · Pull Request #1110 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1110)
- merged

Merging into main:

[Activation hooks redesign (reuse hooks component across both minitron
and puzzletron) by danielkorzekwa · Pull Request #1022 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1022)
- merged

[Dkorzekwa/puzzletron use importance hooks from prune by danielkorzekwa
· Pull Request #1115 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1115)
- merged

</details>

<!-- Details about the change. -->

### Usage

Puzzletron tutorial:

https://github.com/NVIDIA/Model-Optimizer/tree/feature/puzzletron/examples/puzzletron

### Testing
The main e2e test for compressing 9 models with Puzzletron:

https://github.com/NVIDIA/Model-Optimizer/blob/feature/puzzletron/tests/gpu/torch/puzzletron/test_puzzletron.py

2-gpu nightly tests: 

-
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24468209205/job/71501061203
-
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24470214159/job/71508152952

### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added Puzzletron: end-to-end heterogeneous pruning & NAS workflow with
AnyModel support, example pipelines, deployment and evaluation
utilities, and tools for converting/pruning and exporting compressed
checkpoints.

* **Documentation**
* Comprehensive Puzzletron tutorials, model-specific guides, evaluator
instructions, example configs, and changelog entry.

* **Chores**
* CI/workflow updates (extras installation, longer GPU test timeout),
pre-commit hook exclusion updated, and CODEOWNERS entries added.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com>
Signed-off-by: Liana Mikaelyan <lmikaelyan@nvidia.com>
Signed-off-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com>
Signed-off-by: Daniel Korzekwa <daniel.korzekwa@gmail.com>
Signed-off-by: jrausch <jrausch@nvidia.com>
Signed-off-by: root <root@pool0-00848.cm.cluster>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Liana Mikaelyan <lmikaelyan@nvidia.com>
Co-authored-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com>
Co-authored-by: J Rausch <38429553+j-rausch@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-16 00:48:11 +05:30
Gwena Cunhaanddmoodie dec2952992 [6034518] Downgrade TRT support for remote autotuning in Autotune from 10.16 to 10.15 (#1259)
### What does this PR do?

Type of change: Bug fix

Remote autotuning is supported in TensorRT from version 10.15, but fails
with Autotune as it's checking for 10.16+. This PR fixes that check and
updates documentation accordingly.

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
See bug 6034518.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!---
Mandatory -->
- Did you write any new necessary tests?: N/A <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added a Remote Autotuning guide for TensorRT 10.15+ with CLI examples;
updated examples to require `--safe --skipInference`.

* **Updates**
* Lowered TensorRT minimum requirement for remote autotuning from 10.16
to 10.15.
  * Clarified CLI help text for trtexec/autotune arguments.

* **Bug Fixes**
* trtexec-based autotuning now verifies the trtexec executable version
when checking compatibility.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
Signed-off-by: dmoodie <dmoodie@nvidia.com>
Co-authored-by: dmoodie <dmoodie@nvidia.com>
2026-04-15 14:31:41 -04:00
Shengliang Xu 14b78aed1f recipes doc (#1165)
### What does this PR do?

Added an extensive guide for the ModelOpt recipe system: recipe
structure, YAML schema (quantization-focused), built-in discovery and
path resolution, floating-point shorthand conversion, example usage
(Python/CLI), authoring guidance, repository layout, and future
directions.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added comprehensive guide for ModelOpt recipes, introducing
declarative YAML-based optimization specifications. Documents recipe
structure, configuration loading, path resolution, and the three-layer
system architecture. Includes built-in recipe discovery conventions and
detailed instructions for authoring custom recipes with practical
examples.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-04-13 13:56:34 -07:00
Keval Morabia 04cd596d79 Add experimental support for transformers>=5.0 + min torch 2.8 (#975)
### What does this PR do?

- Add experimental support for transformers >=5.0 and remove deprecated
usages:
https://github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md
- ⚠️ For accelerate examples that used `--warmup-ratio: float`
(deprecated in 5.x), we now change it to `--warmup-steps: float | int`
which works as ratio if float but only for 5.x. For 4.x, it will error
out if float and prompt user to change back to `--warmup-ratio` or pass
an int absolute step count.
- ⚠️ Unified Hugging Face checkpoint export for quantized checkpoints
may not work for some models with transformers>=5.0 yet as it requires a
lot of fixes (e.g. change in how MoE experts are organized)
- ~Add Workaround for TRT-LLM's import of deprecated transformers
functions so trt-llm based gpu unit tests work fine. Still deployment
for models needs proper fixes directly in TRT-LLM hence llm/vlm ptq
example tests still run with transformers 4.57~
- Everything except PTQ and Export (mainly MoE) should work fine with
transformers>=5.0
- Bump min torch to 2.8 and enable 2.11 cicd testing
- NOTE: Upcoming Nemo:26.04 container comes with transformers 5.3

### Testing
<!-- Mention how have you tested your change if applicable. -->

- [x] CI/CD tests passing
- [x] Manually tested unit tests, gpu tests with transformers 4.56 and
5.4
- [x] Manually tested example tests (except trt-llm container tests)
with transformers 4.56 and 5.4
- [x] 2-gpu nightly CICD tests manually triggered and passing: [gpu
tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867257540),
[example
tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867260643)

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Make remote-code usage opt-in via a configurable --trust_remote_code
flag across examples and tools.

* **Bug Fixes**
* Improve checkpoint/resume detection and related training guidance to
avoid erroneous errors.

* **Refactor**
* Consolidate dtype/config naming, switch warmup settings from ratio →
steps, and unify tokenizer invocation patterns.

* **Documentation**
  * Simplify changelog title and add misc notes for release 0.44.

* **Chores**
* Remove scheduled PR-branch cleanup workflow and relax/remove several
transformers version pins.

* **Tests**
* Adjust test gates, skips, and structures to align with updated deps
and behaviors.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-09 09:59:37 +05:30
Shengliang Xu 1cceb950d6 [OMNIML-3689] PTQ quant_cfg semantic correction. Design in doc _quant_cfg.rst (#1094)
### What does this PR do?

#### Summary

Redesigns the `quant_cfg` configuration format in ModelOpt's PyTorch
quantization stack, replacing the previous dict-based format with an
**ordered list of typed `QuantizerCfgEntry` dicts**.

##### Motivation

The old `quant_cfg` dict had several pain points:
- **Ambiguous precedence**: no explicit way to reason about which entry
wins when multiple keys match a quantizer
- **Mixed key namespaces**: wildcard paths and PyTorch class names lived
in the same dict level, requiring ad-hoc dispatch
- **Magic `"default"` key**: an implicit, undocumented catch-all that
was easy to misuse
- **Poor composability**: merging two configs required dict updates that
silently discarded keys
- **No YAML round-trip fidelity**: the nested structure couldn't be
expressed cleanly in YAML

##### New format

`quant_cfg` is now an ordered list of `QuantizerCfgEntry` TypedDicts.
Each entry has:
- `quantizer_name` *(required)*: `fnmatch` wildcard matched against
quantizer module names
- `cfg` *(optional)*: dict (or list of dicts) of
`QuantizerAttributeConfig` fields
- `enable` *(optional)*: toggles quantizer on/off independently of `cfg`
- `parent_class` *(optional)*: restricts match to quantizers whose
parent module is of the given PyTorch class (e.g. `"nn.BatchNorm2d"`)

Entries are applied in list order; later entries override earlier ones.
The canonical pattern is deny-all first (`_base_disable_all`), then
selectively re-enable and configure, then apply standard exclusions
(`_default_disabled_quantizer_cfg`).

##### Changes

**Core library (`modelopt/torch/quantization/`)**

- **`config.py`**:
- Added `QuantizerCfgEntry` TypedDict (line 163) and
`find_quant_cfg_entry_by_path()` helper for exact-match lookup of
entries by path.
- Added `normalize_quant_cfg_list()` (line 1539) that converts legacy
formats (flat dict, single-key dicts, `nn.*`-scoped dicts, `"default"`
key) to canonical `QuantizerCfgEntry` lists. After normalization every
entry is guaranteed to have explicit `quantizer_name`, `enable`, and
`cfg` keys.
- Converted `_default_disabled_quantizer_cfg` and
`_mamba_moe_disabled_quantizer_cfg` from dicts to lists of
`QuantizerCfgEntry`.
- Added `_base_disable_all` (line 205): canonical deny-all entry
(`[{"quantizer_name": "*", "enable": False}]`).
- Converted all ~30 built-in config constants (`INT8_DEFAULT_CFG`,
`FP8_DEFAULT_CFG`, `NVFP4_DEFAULT_CFG`, etc.) to list format using
`*_base_disable_all` and `*_default_disabled_quantizer_cfg` unpacking.
- KV-cache configs (`FP8_KV_CFG`, `NVFP4_KV_CFG`, etc.) are now minimal
lists designed to be concatenated with a primary config — they
intentionally omit `_base_disable_all` and `"algorithm"`.
- Added two `QuantizeConfig` Pydantic field validators: a
`mode="before"` validator that calls `normalize_quant_cfg_list()`, and a
`mode="after"` validator that validates `cfg` dicts against
`QuantizerAttributeConfig`.
- Updated `need_calibration()` to iterate the normalized list instead of
the old dict.
- Changed `QuantizeQuantCfgType` alias from `dict[str | Callable, ...]`
to `list[QuantizerCfgEntry]`.

- **`conversion.py`**:
- Rewrote `set_quantizer_by_cfg()` (line 217) to iterate the list
directly. Each entry's `parent_class` is resolved via
`QuantModuleRegistry[parent_class_name]` (the existing `_DMRegistryCls`
registry).
- Added `set_quantizer_attributes_full()` (line 314): full replacement
of quantizer attributes from a `QuantizerAttributeConfig`. Unspecified
fields revert to defaults, enforcing entry atomicity. Can also upgrade
`TensorQuantizer` → `SequentialQuantizer` or downgrade the reverse.
- Added `set_quantizer_attributes_partial()` (line 384): merges a
partial `dict` of attributes into existing quantizer state. Does NOT
change quantizer structure. Used for enable-only entries.
- Added `set_quantizer_by_cfg_context()` context manager (line 447) that
temporarily applies a `quant_cfg` list and restores original quantizer
state on exit.
- Deprecated `set_quantizer_attribute()` (line 525) with a
`DeprecationWarning` pointing to the new functions.

- **`tensor_quantizer.py`**:
- `TensorQuantizer.set_from_attribute_config()`: narrowed type hint from
`dict` to `dict[str, Any]`.
- Added `_axis_setter` and `_block_sizes_setter` custom setters so that
`axis` and `block_sizes` changes properly propagate to the calibrator
and maintain mutual exclusivity.
- `SequentialQuantizer.set_from_attribute_config()`: narrowed signature
to `list[QuantizerAttributeConfig] | list[dict[str, Any]]` (removed the
old union with single values).

- **`algorithms.py`**:
- Updated `_match_quantizer_cfg()` to iterate the list and return
`(matched_cfg, matched_enable)` tuple with last-match-wins.
- Updated `_cfg_to_dict()`, `estimate_quant_compression()`, and
`QuantRecipe` to work with the list-based format.
- Updated `get_auto_quantize_config()` to emit list-format `quant_cfg`.

- **`model_quant.py`**: `disable_quantizer()` / `enable_quantizer()` now
call `set_quantizer_attributes_partial()` directly instead of the
deprecated `set_quantizer_attribute()`. Updated docstrings and code
examples to show the list format.

- **`utils/core_utils.py`**: `disable_lora_quantizers_in_config()` and
`update_quant_cfg_with_kv_cache_quant()` updated to append
`QuantizerCfgEntry` dicts to the list.

- **Other**: minor updates to `backends/fp8_per_tensor_gemm.py`,
`backends/nvfp4_gemm.py`, `compress.py`, `model_calib.py`,
`export/unified_export_hf.py`, and
`sparsity/attention_sparsity/conversion.py` to use the list format.

- **`onnx/llm_export_utils/quantization_utils.py`**: Updated
quantization config construction to use list format.

**YAML recipes (`modelopt_recipes/`)**

- Converted all 5 general PTQ recipes to the new list format:
  - `general/ptq/fp8_default-fp8_kv.yml`
  - `general/ptq/nvfp4_default-fp8_kv.yml`
  - `general/ptq/nvfp4_experts_only-fp8_kv.yml`
  - `general/ptq/nvfp4_mlp_only-fp8_kv.yml`
  - `general/ptq/nvfp4_omlp_only-fp8_kv.yml`
- Converted model-specific recipe:
`models/Step3.5-Flash/nvfp4-mlp-only.yaml`

**Documentation (`docs/`)**

- New guide: `docs/source/guides/_quant_cfg.rst` — comprehensive
reference covering entry format, ordering semantics, entry atomicity,
`enable` vs `cfg` independence, `parent_class` filtering, and common
patterns (deny-all-then-enable, customizing a built-in config, building
from scratch).
- Updated `_pytorch_quantization.rst` code examples to show the list
format with `copy.deepcopy` and `.append()`.
- Added `_quant_cfg.rst` to the quantization guide table of contents.

**Examples**

- Updated all quantization examples to use the list format:
`deepseek/ptq.py`, `diffusers/quantization/config.py`,
`llm_ptq/hf_ptq.py`, `llm_qat/main.py`, `vllm_serve/vllm_ptq_utils.py`,
`llm_autodeploy/run_auto_quantize.py`, `llm_eval/quantization_utils.py`,
`llm_ptq/example_utils.py`,
`windows/torch_onnx/diffusers/qad_example/sample_example_qad_diffusers.py`,
and 2 notebooks.

**Tests**

- New test file:
`tests/unit/torch/quantization/test_config_validation.py` — unit tests
for `need_calibration()`, `normalize_quant_cfg_list()` (new format,
legacy format conversions, error cases),
`find_quant_cfg_entry_by_path()`, `_match_quantizer_cfg()`, and
`QuantizeConfig` Pydantic validators.
- Extended `tests/unit/torch/quantization/test_quantize_cpu.py` with
tests for `set_quantizer_attributes_full()` (atomicity, parent_class
filtering, SequentialQuantizer creation), list ordering, enable-only
entry behavior, and end-to-end legacy dict format.
- Updated 20+ existing test files across `tests/unit/`, `tests/gpu/`,
`tests/gpu_megatron/`, and `tests/_test_utils/` to use the list format.

##### Backward compatibility

`normalize_quant_cfg_list()` is called automatically by the
`QuantizeConfig` Pydantic `mode="before"` validator, so existing code
passing the old dict-based format (flat dict like `{"*weight_quantizer":
{"num_bits": 8}}`, single-key dict lists, or `nn.*`-scoped dicts with
`parent_class` semantics) continues to work without modification. The
legacy `"default"` key is converted to `quantizer_name: "*"`.

`set_quantizer_attribute()` is preserved as a deprecated wrapper around
`set_quantizer_attributes_partial()`.

#### Test coverage

- **Unit tests**: new `test_config_validation.py` with tests for
normalization, validation, path lookup, and cfg matching. Extended
`test_quantize_cpu.py` with tests for full/partial attribute setting,
ordering, atomicity, and legacy backward compatibility.
- **System testing**:

```
python examples/llm_ptq/hf_ptq.py \
      --model Qwen/Qwen3-8B  \
      --recipe general/ptq/fp8_default-fp8_kv \
      --export_path=build/fp8_default-fp8_kv42  \
      --calib_size=16 \
      --batch_size=0 \
      --trust_remote_code \
      --export_fmt=hf
```

### Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-04-06 15:38:44 -07:00
vishalpandya1990 20a46e04a7 Update ModelOpt-with-Olive documentation to mention CUDA EP commands (#1099)
### What does this PR do?

Type of change: Minor documentation update

- Update documentation (olive installation instructions) to mention
install commands for CUDA EP packages for ORT / ORT-genai.

### Testing

- Locally checked the readme, doc.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated Windows Olive installation guidance to recommend CUDA-based
ONNX Runtime packages instead of DirectML and added a link to ONNX
Runtime’s Execution-Provider docs for alternative EPs and requirements.
* Simplified Windows examples by removing explicit package install
commands and pointing users to the consolidated Olive setup
instructions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: vipandya <vipandya@nvidia.com>
2026-03-24 17:54:15 +05:30
Keval Morabia d0bf0bef96 Remove deprecated Nemo 2.0 references / examples (#1098)
### What does this PR do?

- Remove `examples/nemo_run` and other deprecated Nemo 2.0 references
- Add Megatron-Bridge example links where missing

<!-- Details about the change. -->

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!---
Mandatory -->
- Did you write any new necessary tests?: N/A <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Deprecations**
* Removed deprecated NeMo 2.0 support and related example flows and
utilities.

* **Documentation**
* Updated docs and examples to emphasize Megatron-Bridge / Megatron-LM
and refreshed technique/deployment guidance and links.

* **New Features**
* Added CLI options for additional parallelism (context/expert
tensor/expert model) in Megatron-Bridge distillation.

* **Chores**
* Removed legacy CI configs and refreshed container image tags across
examples and docs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-24 14:11:13 +05:30
cb1ff321ee Add Python 3.13 support (#1048)
Fixes https://github.com/NVIDIA/Model-Optimizer/issues/217

## Summary

- Bump `requires-python` from `>=3.10,<3.13` to `>=3.10,<3.14` to
formally include Python 3.13
- Add explicit Python 3.10–3.13 PyPI classifiers for better
discoverability
- Add `py313` to tox CPU unit test and partial-install environment
matrices
- Add Python 3.10–3.13 to the `multi-py` CI matrix in `unit_tests.yml`

## Background

Python 3.13 was previously excluded by the `<3.13` upper bound. Testing
in a related repo with `--ignore-requires-python` confirmed that the
library installs and runs correctly under Python 3.13. This PR lifts the
restriction and wires up CI to verify it going forward.

## Test plan

- [ ] CI `multi-py` job passes on `py313-torch210-tf_latest-unit`
- [ ] `tox -e py313-torch210-tf_latest-unit` passes locally (requires
Python 3.13 installed)
- [ ] `tox -e py313-partial-unit-torch` passes locally
- [ ] No regressions on existing Python 3.10/3.11/3.12 matrix jobs

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Extended Python support: minimum remains 3.10; added official support
up through 3.13 (upper bound advanced accordingly).
* **Tests**
* CI and test matrix expanded to include experimental Python 3.13
coverage.
* **Documentation**
* Installation docs and changelog updated to reflect Python 3.13
support.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ivan Basov <ibasov@nvidia.com>
Signed-off-by: Ivan Basov <5455484+ivanbasov@users.noreply.github.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-17 11:33:48 +05:30
Gwena Cunha 69c0d47946 [OMNIML-3252][ONNX] MOQ + Autotune moq integration docs (#1026)
### What does this PR do?

**Type of change**: documentation

**Overview**: This PR updates the documentation and does some folder
re-structuring and file re-naming related to
https://github.com/NVIDIA/Model-Optimizer/pull/951.

### Usage

Documentation

### Testing

Documentation

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: N/A <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ (renamed `AutoQDQ` to `Autotune`)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
  * Renamed AutoQDQ to Autotune across guides and changelog.
  * Updated Autotune guide descriptions and wording.
* Added a new section on optimizing Q/DQ node placement with Autotune,
including CLI usage and API links (appears twice in one README).
  * Applied minor grammar and capitalization corrections.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
2026-03-12 10:47:07 -07:00
695c8e8522 Integrate Automated QDQ placement tool - part 4.3 (#843)
## What does this PR do?

This PR upload user guide of Automated QDQ placement tool. This tool
automatically search QDQ insertion points with better performance.

**Overview:** ?

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added comprehensive guide for Automated Q/DQ Placement Optimization
workflow, including quick start instructions, advanced usage patterns,
configuration options, best practices, and troubleshooting.

* **New Features**
  * Exposed public API for CLI parser programmatic access.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Will Guo <willg@nvidia.com>
Signed-off-by: Gwena Cunha <4861122+gcunhase@users.noreply.github.com>
Co-authored-by: Gwena Cunha <4861122+gcunhase@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-10 18:42:51 +00:00
Keval Morabia 860d0b4a70 Merge Linux and Windows Changelog (#954)
Since we no longer have any compiled packages, all releases are for all
platforms so we dont have separate windows releases hence merging
changelog as well.

Going forward, windows can test on latest linux version and if fixes
needed, they can go in next usual monthly linux release

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Reorganized changelog structure to consolidate Windows and Linux
release information in a unified view.
* Expanded Windows Support documentation across recent release versions.

* **Chores**
* Updated Windows example release badge to dynamically reflect the
latest PyPI release version.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-03 11:31:59 +05:30
Keval Morabia 82f1d216d1 Add Security and IP related contributing guide and configure coderabbit to catch such issues (#935)
### What does this PR do?

- Add Security related coding practices in `SECURITY.md` and merge with
`2_security.rst`
- Update `CONTRIBUTING.md` for instructions to follow if copying code
from other repositories
- Update PR template
- Cleanup dependency files
- New API `mto.load_modelopt_state` doing the insecure `torch.load(f,
weights_only=False)` instead of doing it separately everywhere. This
also allows us to later improve the input validation for
`modelopt_state_path` or use safer alternatives to `torch.load`

### Testing
<!-- Mention how have you tested your change if applicable. -->

N/A

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=True)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ <!--- Mandatory -->
- Did you write any new necessary tests?: NA <!--- Mandatory for new
features or examples. -->
- Did you add or update any necessary documentation and update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
NA <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Expanded and reorganized security guidance and contributor procedures;
updated PR template and several READMEs with clearer security,
submission, and installation instructions
* Replaced an older security document with an enhanced, centralized
security guidance

* **Chores**
* Adjusted example dependency lists and optional extras (adds, removals,
and version constraints)
* Enabled automated incremental reviews, added pre-merge security
checks, and introduced a knowledge-base of coding/security guidelines
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-03 03:49:15 +05:30