mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
main
522
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5584ce4558 |
Migrate Nemotron-3-Nano tutorial PTQ to MBridge scripts and move under examples/megatron_bridge (#1601)
### What does this PR do? Type of change: documentation (+ minor test fixes) Migrates the Nemotron-3-Nano-30B-A3B-BF16 tutorial quantization step from `examples/llm_ptq/hf_ptq.py` to the Megatron-Bridge quantize + export, and relocates the tutorial next to the scripts it now uses. Now that the whole tutorial is Megatron-Bridge based, it lives under `examples/megatron_bridge/`. - **Quantization migration:** replace the single `hf_ptq.py` call with `examples/megatron_bridge/quantize.py` (calibrate + save a Megatron checkpoint) → `examples/megatron_bridge/export.py` (deployable unified HF checkpoint). The FP8 results table is refreshed with the `quantize.py` numbers (same defaults, slightly better on average). - **Relocation:** moved `examples/pruning/minitron/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/` → `examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/`. A **redirect-stub `README.md`** remains at the old path (a directory symlink isn't traversable in the GitHub web UI), and all in-repo references (root README, CHANGELOG, pruning READMEs, megatron_bridge README) plus the tutorial's own relative links are updated. - **Evaluation:** per-format vLLM benchmark commands (BF16 / FP8), FP8 deployment notes documented in `nemo_evaluator.yaml`, reduced LiveCodeBench/AIME `num_repeats` (were too slow), and bumped the `nemo-evaluator-launcher` pin. - **Misc:** drop the `examples/megatron_bridge/requirements.txt` `transformers<5` pin in favor of an inline "downgrade `transformers<5` to save pruned Nemotron checkpoints" note; guard the hybrid Mamba-MoE sharded-state-dict test behind `HAS_MAMBA` (requires `mamba_ssm`); shrink the tiny Gemma3 test fixture's attention heads. > **Note:** the **NVFP4 + QAD** experiments (formerly the focus of this PR) are split out — their accuracy/throughput results are still in progress — and will follow in a separate PR on top of this one. ### Testing Docs-only + test-guard changes. Pre-commit hooks (markdownlint, RST checks, ruff, mypy) pass. The tutorial's relative links and the old-path redirect stub were verified to resolve to real files. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (old tutorial path still resolves via a redirect-stub README; `quantize.py`/`export.py` already exist in `examples/megatron_bridge`) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (adjusts/guards existing tests only) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ (existing tutorial entry updated to the new path) - Did you get Claude approval on this PR?: ✅ ### Additional Information Supersedes the previous "Part 3 of 4 (NVFP4 + QAD docs)" scope of this PR; the NVFP4 + QAD tutorial additions will land in a follow-up. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Moved the Nemotron-3-Nano-30B-A3B tutorial into the Megatron-Bridge tutorials and replaced the old file with a pointer to the new location. * Updated vLLM throughput numbers to 2.6× and expanded results/throughput tables. * Reworked the FP8 quantization/export workflow and added a note to use transformers<5 when saving pruned models. * Added a tutorials index and adjusted evaluator launcher pin and repeat counts. * **Tests** * Tests now detect optional Mamba support and skip related tests when unavailable. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8b01ba4274 |
Skip softmax calibration via Triton kernel (#1597)
### What does this PR do? Adds skip softmax calibration for LLMs via Triton kernel (leveraging existing kernel used for diffusion) Type of change: New feature <!-- Details about the change. --> ### Usage ``` python hf_sa.py --pyt_ckpt_path Qwen/Qwen3-8B --sparse_attn skip_softmax_triton_calib ``` The Triton calibration equals PyTorch at every threshold, for both phases: | threshold | prefill triton/pytorch | decode triton/pytorch | |------|------------------------|-----------------------| | 0.30 | 0.0% / 0.0% | 12.5% / 12.5% | | 0.50 | 0.0% / 0.0% | 37.5% / 37.5% | | 0.70 | 10.0% / 10.0% | 62.5% / 62.5% | ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a Triton-based skip-softmax sparse-attention calibration option and a CLI flag to override the calibration data directory (defaults to adjacent RULER data). * **Bug Fixes** * Ensure calibration kernels run on the correct CUDA device; align measurement granularity and tile/block sizing; ignore padded query rows when counting skippable tiles. * **Tests** * Added GPU Triton calibration tests for end-to-end inference, multi-threshold stats, and decode-phase reporting. * **Documentation** * Updated changelog and example to expose the new option and flag. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Rohan Joshi <rohjoshi@nvidia.com> Signed-off-by: Kai Xu <kaix@nvidia.com> Co-authored-by: Kai Xu <kaix@nvidia.com> |
||
|
|
aec72ffa68 |
Add DMD2 distillation for Qwen-Image (fastgen) (#1326)
### What does this PR do?
**Type of change:** New example + new `modelopt.torch.fastgen` library
module.
Adds **DMD2 (Distribution Matching Distillation) for Qwen-Image** —
distilling the base model into a few-step (1–4) generator. Includes the
framework-agnostic `modelopt.torch.fastgen` loss library (DMD pipeline,
EMA, optional GAN discriminator) and a NeMo AutoModel–based training
example with a mock-data smoke config, a real-data config, and inference
/ export scripts.
**Noted**: the example script will be migrated to AutoModel repo
### Usage
```bash
# Mock-data wiring smoke — runs end-to-end with no dataset to prepare
torchrun --nproc-per-node=8 \
examples/diffusers/fastgen/dmd2_finetune.py \
--config examples/diffusers/fastgen/configs/dmd2_qwen_image_smoke.yaml
```
See `examples/diffusers/fastgen/README.md` for real-data training and
inference.
### Testing
Unit tests under `tests/unit/torch/fastgen/`; `pre-commit` /
code-quality clean.
### Before your PR is "*Ready for review*"
- Backward compatible?: ✅ (new, additive module)
- Followed `CONTRIBUTING.md` for any copied code / new deps: ✅
- New tests added?: ✅
- Updated Changelog?: N/A
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Adds a FastGen-based distillation framework (DMD2) with
student/fake-score training, EMA support, GAN discriminator branch,
inference pipeline, and export utilities.
* Qwen-Image integration with latent packing and feature-capture for
plugin-enabled pipelines.
* **Documentation**
* New README, example configs, and runnable example scripts for
Qwen-Image distillation and inference.
* **Tests**
* Comprehensive unit tests covering math parity, gradient routing,
plugins, hooks, EMA, and recipe setup.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
c78e654744 |
Skip Softmax diffusion export (#1269)
### What does this PR do?
Type of change: New Feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
Adds HuggingFace `config.json` export of skip-softmax sparse-attention
calibration for diffusion pipelines (e.g. Wan 2.2), on top of the base
skip-softmax work.
- **`_export_diffusers_checkpoint`** walks every `nn.Module` component
of a diffusers pipeline, calls `export_sparse_attention_config`, and
writes the result into that component's `config.json` under the
`sparse_attention_config` key. The sparse config lives **only** in
`config.json` — there is no standalone `sparse.yaml`.
- **`export_sparse_attention_config`** emits a `config_groups` schema
where each algorithm's parameters are nested inside its own group; only
`config_groups` and `producer` are top-level:
- skip-softmax group → `algorithm: "skip_softmax"`, `targets`, `ignore`
(layers kept dense — e.g. cross-attention + first/last blocks),
`initial_disabled_steps` (opt-in, user-set; emitted only when `> 0`),
`threshold_scale_factor` (`a * exp(b * target_sparsity)`), and
`target_sparsity`.
- N:M group → `algorithm: "sparse_softmax"` with
`sparsity_n`/`sparsity_m`, `dense_sink_tokens`, `dense_recent_tokens`
flattened into the group.
- **Deploy reader**
(`modelopt/torch/sparsity/attention_sparsity/plugins/sparse_attn_config.py`)
reads these per-group params back, keeping the export↔load round-trip
consistent.
- **Example wiring**:
`examples/diffusers/sparsity/wan22_skip_softmax.py` gains
`--export-dir`, `--skip-softmax-threshold`, and
`--initial-disabled-steps`. `--export-dir` runs
`export_hf_checkpoint(pipe, export_dir=...)` after calibration.
- Updated `CHANGELOG.rst`.
### Usage
```bash
python examples/diffusers/sparsity/wan22_skip_softmax.py \
--model-path Wan-AI/Wan2.2-T2V-A14B-Diffusers \
--calibrate --target-sparsity 0.5 --calib-size 4 \
--initial-disabled-steps 5 \
--export-dir ./wan22_skip_softmax_ckpt
```
Resulting layout — a `config.json` per component, **no `sparse.yaml`**:
```
wan22_skip_softmax_ckpt/
├── transformer/config.json # carries sparse_attention_config
├── transformer_2/config.json # carries sparse_attention_config
├── vae/ … text_encoder/ … tokenizer/ … scheduler/ …
└── model_index.json
```
A representative `config.json` entry for a diffusion transformer:
```json
"sparse_attention_config": {
"config_groups": {
"group_0": {
"algorithm": "skip_softmax",
"targets": ["WanAttention"],
"ignore": ["blocks.0.attn1", "blocks.0.attn2", "…"],
"initial_disabled_steps": 5,
"threshold_scale_factor": {
"formula": "a * exp(b * target_sparsity)",
"prefill": {"a": 1443.49, "b": 4.30}
},
"target_sparsity": {"prefill": 0.5}
}
},
"producer": {"name": "modelopt", "version": "0.45.0..."}
}
```
The N:M variant adds a second group:
```json
"group_1": {
"algorithm": "sparse_softmax",
"targets": ["WanAttention"],
"sparsity_n": 2, "sparsity_m": 4,
"dense_sink_tokens": 0, "dense_recent_tokens": 64
}
```
### Testing
- `tests/examples/diffusers_sparsity/test_sparsity.py`: baseline /
triton-baseline / fixed-threshold runs of the Wan 2.2 example, plus a
Python-API calibrate → **export** test asserting the nested
`sparse_attention_config` (`threshold_scale_factor`, `target_sparsity`,
`ignore`, `initial_disabled_steps`) and the absence of any
`sparse.yaml`.
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attention_conversion.py`
and `test_sparse_attn_config.py`: unit coverage of the per-group export
schema and the deploy-reader round-trip (writer nests → reader reads
from groups → internal mtsa config unchanged).
- Validated end-to-end on Wan 2.2 T2V-A14B: full 4-prompt / 40-step /
81-frame calibration; the exported checkpoint carries the nested schema
in both `transformer` and `transformer_2` `config.json`, and runtime
measurement shows ~47–49% tile sparsity at a 0.5 target.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ❌ The exported
`sparse_attention_config` schema was renamed and nested per-group during
0.45.x development, and the loader reads only the new layout —
checkpoints exported by earlier 0.45.x builds must be re-exported. No
released version is affected. <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
275d45dce2 |
Add Alpamayo-1 example (#1594)
### What does this PR do? Type of change: ? New example <!-- Details about the change. --> Adds example for Alpamayo-1 quantization with ModelOpt (FP8, NVFP4, AutoQuant) ### Usage ``` python quantize.py --ckpt nvidia/Alpamayo-R1-10B --output-dir ./alpamayo-r1-fp8 --quantize fp8 ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Alpamayo 1 vision-language-action model quantization example supporting FP8, NVFP4, and mixed-precision optimization modes * Introduced CLI quantization tool with calibration loop and checkpoint export capabilities for both fake-quantized and real-quantized formats * **Documentation** * Added comprehensive guide documenting the Alpamayo quantization example, model details, and usage instructions <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Rohan Joshi <rohjoshi@nvidia.com> |
||
|
|
b98a59557a |
Add vLLM-based runtime statistics for subblock latency measurement (#1358)
### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Runtime-based latency optimization: collect vLLM-measured inference latency to constrain optimization. * **Configuration** * New runtime config/template for Llama-3.1-8B pruning (runtime stats enabled, NCCL timeout templating, MIP target-latency). * Validation sample defaults adjusted (one flow: 128 → 8; runtime flow uses 128). * Human constraint key renamed to target_latency_seconds. * **Documentation** * README section describing runtime-based latency optimization setup and usage. * **Tests** * Added GPU end-to-end test for runtime stats collection. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1358?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Grzegorz Karch <gkarch@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
01415c2788 |
fix(llm_eval): repair test_qwen3_eval_fp8 end-to-end (#1650)
### What does this PR do? Type of change: Bug fix `tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8` was silently passing while its evals crashed, then began failing as a timeout. This repairs the whole pipeline: - **lm_eval `IndexError` (root cause):** TRT-LLM KV-cache prefix reuse returns truncated `context_logits` for shared-prefix requests (e.g. hellaswag's one-context / many-endings), which breaks `parse_logprobs`. Add an `enable_kv_cache_reuse` flag to `modelopt.deploy.llm.LLM` (default `True`, unchanged) and disable it for the eval deployment so full-length context logits are returned. - **Silent CI green:** `python eval.py | tee result.txt` returns `tee`'s exit code, so a crashing eval was masked. Add `set -o pipefail` to `huggingface_example.sh` so failures fail the test. - **Long-prompt overflows:** with the tiny test model's toy tokenizer, gsm8k/MMLU prompts exceed `max_seq_len`. Bump test `max_position_embeddings` to 8192, skip MMLU prompts that don't fit even at zero-shot, and add an MMLU sample limit (`--mmlu_limit`). - **human-eval build failures:** install with `--no-build-isolation` (`pkg_resources` is absent in pip's isolated build env), patch its malformed `console_scripts` entry point, and pin the clone. - **Cleanups:** gate the post-quant `run_tensorrt_llm.py` smoke test behind the `quant` task (eval tasks deploy on their own; ~45s saved for eval-only runs); replace the SIGPIPE-prone serve-readiness `tail -f | while` with a poll loop (required under `pipefail`). ### Usage N/A — example/test fix. ### Testing All four eval tasks verified end-to-end in the CI container (TRT-LLM 1.3.0rc17, RTX 6000 Ada): lm_eval (hellaswag + gsm8k), MMLU, and simple_eval (humaneval) all complete with exit 0 and no `IndexError`/overflow. Cold full run ≈ 340s on this GPU. CI test on 2-gpu: https://github.com/NVIDIA/Model-Optimizer/actions/runs/27154417497/job/80153551154 ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (new `enable_kv_cache_reuse` defaults to current behavior; new script flags are optional) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new dependencies) - Did you write any new necessary tests?: N/A (fixes and strengthens an existing test) - Did you update Changelog?: N/A (bug fix to examples/tests) - Did you get Claude approval on this PR?: ❌ (pending) ### Additional Information The full test runs ~340s on an RTX 6000 Ada; CI runners are historically slower, while `@pytest.mark.timeout` is set to 600 — worth watching the first CI run and bumping if it's close. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an option to limit MMLU evaluation length. * **Bug Fixes** * Disabled KV-cache prefix reuse for evaluations needing per-token context logits to prevent truncated/incorrect logprobs. * Skip examples whose prompts remain too long; warn and report accuracy as NaN if all examples are skipped. * **Chores / Scripts** * Improved example scripts for reproducible installs, patched entry point handling, pipeline failure detection, conditional test invocation, polling-based log wait, and a new CLI flag for MMLU limits. * **Tests** * Increased timeout and prompt headroom; capped MMLU smoke tests for speed. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1f4a489b0f |
Adds AutoQuant support for VLM / Qwen3.5-Qwen3.6 style models (#1381)
### What does this PR do? Type of change: new feature, bug fix, new tests ### Details - Enables AutoQuant search over fused MoE expert containers by snapshotting/restoring their per-expert quantizers. - Adds Qwen3.5/3.6 linear-attention grouping rules so fused deployment layers keep compatible quant formats. - Supports `w4a16_nvfp4` as an AutoQuant search format. - Preserves disabled AutoQuant layer patterns in generated configs while allowing selected modules like `lm_head` to override default disables. - Keeps recipe-mode and AutoQuantize VLM paths on the outer CausalLM so Qwen3.5/3.6-MoE `lm_head` remains visible. - Skips `parent_class`-scoped quant config entries during AutoQuant bare quantizer matching, preventing class-scoped global entries from last-match overriding every selected module. - Adds temporary hardcoded Qwen/VLM AutoQuant disabled-layer patterns in `hf_ptq.py` with a TODO to refactor into the config system. ### Usage ```bash python examples/llm_ptq/hf_ptq.py \ --pyt_ckpt_path <model_path> \ --qformat fp8,w4a16_nvfp4 \ --auto_quantize_bits 5.0 \ --auto_quantize_cost_model active_moe \ --auto_quantize_checkpoint <autoquant_state.pt> \ --export_path <output_dir> ``` ### Testing - `/Users/weimingc/miniconda3/envs/modelopt/bin/python -m pytest tests/unit/torch/quantization/test_autoquant.py::test_get_auto_quantize_config_keeps_selected_lm_head_enabled tests/unit/torch/quantization/test_config_validation.py::TestMatchQuantizerCfg::test_parent_class_scoped_entries_are_ignored_for_bare_autoquant_lookup` - `/Users/weimingc/miniconda3/envs/modelopt/bin/python -m pytest tests/unit/torch/quantization/test_autoquant.py tests/unit/torch/quantization/test_config_validation.py -k "not data_parallel"` (`120 passed, 1 deselected`) - `/Users/weimingc/miniconda3/envs/modelopt/bin/python -m py_compile examples/llm_ptq/hf_ptq.py modelopt/torch/quantization/algorithms.py modelopt/torch/quantization/_auto_quantize_cost.py tests/unit/torch/quantization/test_autoquant.py tests/unit/torch/quantization/test_config_validation.py` - Full local affected-file pytest without `-k "not data_parallel"` only failed `test_data_parallel_auto_quantize` because this local sandbox cannot bind a free socket (`PermissionError: Operation not permitted`). - Ran Qwen3.6 35B AutoQuant e2e with `fp8,w4a16_nvfp4` and exported a checkpoint. - Verified exported checkpoint loads in vLLM nightly without local patches. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added w4a16_nvfp4 quantization format and optional cost-exclusion patterns for AutoQuantize. * **Improvements** * Safer multimodal/VLM handling and AutoQuantize now runs on the full outer model when applicable. * Better fused-MoE support, more accurate weight accounting, and refined attention-grouping for improved quantization choices. * Dynamic layer-disabling support for targeted disables. * **Tests** * New unit tests covering cost-model exclusions, fused-MoE accounting, and config selection. * **Documentation** * Updated cost-constraint example to show exclusion-pattern usage. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
2c52e7bf4e |
[OMNIML-4788] specdec_bench/Qwen3.5-4B: throughput_32k benchmark + S3 upload step (#1564)
### What does this PR do? Type of change: enhancement (follow-up to [#1531](https://github.com/NVIDIA/Model-Optimizer/pull/1531)). Extends the merged Qwen3.5-4B SPEED-Bench launcher YAMLs from a single-task qualitative-only smoke into a **3-task pipeline** that also covers long-context throughput and verifies the S3-upload path end-to-end. Two commits, cleanly cherry-picked from #1531's late branch state — they were authored after the merge-commit was resolved against an earlier rebased head and so didn't ride along with that merge. ### Pipeline shape (both YAMLs) | Task | Split | Save dir | |---|---|---| | `task_0` | qualitative (existing quality / acceptance-rate signal) | `/scratchspace/specdec_bench{,_mtp}/qualitative` | | `task_1` | **throughput_32k** (new — long-context throughput) | `/scratchspace/specdec_bench{,_mtp}/throughput_32k` | | `task_2` | **upload to S3 in sweep layout** | `s3://team-specdec-workgroup/results/specdec_bench{,_mtp}/<split>/` | ### New artifacts * `tools/launcher/common/specdec_bench/upload_to_s3.sh` — thin wrapper around `examples/specdec_bench/upload_to_s3.py` so it can be invoked as a launcher task. Installs `boto3` from `requirements.txt` on cold containers; warm pipelines pick it up from the prior `run.sh`. * `tools/launcher/common/specdec_bench/runtime_params_throughput_32k.yaml` — pins `engine_args.max_model_len = 40,960` (32K input + 4K output + 4K headroom) so vLLM doesn't silently auto-cap `max_model_len` below the 36K minimum needed for `throughput_32k` prompts on single-GPU runs. ### Why max_model_len matters Without an explicit `max_model_len`, vLLM auto-derives it from the model config (Qwen3.5-4B = 128K) **and from the GPU-memory budget**. On a single GPU the second factor can cap effective `max_model_len` well below 36K, silently truncating 32K-token prompts and producing wrong throughput numbers. The qualitative split is not affected (its prompts top out around 8K, well under any auto-derivation floor) so only `task_1` carries the override. ### S3 credentials `upload_to_s3.sh` reads `S3_ENDPOINT` / `S3_KEY_ID` / `S3_SECRET` from the runtime environment (not hardcoded). `--skip-existing` + `--allow-incomplete-provenance` are passed by default so re-runs land alongside the prior upload, and runs lacking `CONTAINER_IMAGE` (Phase-2 harness work in OMNIML-4788 will populate it) still upload. ### Testing Cluster smoke on cw_dfw via: ``` uv run slurm.py --yaml modules/Model-Optimizer/tools/launcher/examples/Qwen/Qwen3.5-4B/specdec_bench.yaml --yes ``` is currently in-flight (jobs `12257378/79/80`, PD). Will update this PR with timing/AR numbers + S3 upload confirmation once it lands. ### Before your PR is "Ready for review" - Backward compatible: ✅ (additive — task_0 keeps the prior qualitative behavior, just with `/qualitative` suffix in `save_dir`) - New PIP dep: ✅ no (boto3 already in `examples/specdec_bench/requirements.txt` from #1531) - New tests: N/A (launcher YAML + shell wrapper; covered by cluster smoke) - Changelog: N/A (internal-facing tooling) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a 32K-context runtime configuration (higher max model length) to enable long-context throughput benchmarking and avoid silent prompt truncation. * Added a launcher helper to upload benchmark results to S3 with incremental/retry-friendly options and pass/fail reporting. * **Chores** * Split Qwen3.5-4B benchmark into separate qualitative and 32K throughput tasks and added coordinated S3 upload. * Applied the same multi-task pipeline layout and clearer output organization to the MTP speculative-decoding benchmark. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1564?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: chenhany <chenhany@nvidia.com> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
54ce4e09d8 |
Add Quantization Aware Distillation (QAD) to Megatron-Bridge example (#1600)
### What does this PR do? Type of change: new example **Note:** This is **part 2 of 4** (builds on #1589): - **Part 1 (#1589):** Megatron-Bridge `quantize.py` + `export.py` support and tests. - **Part 2 (this PR):** extend `distill.py` for quantization-aware distillation (QAD) — load a quantized Megatron checkpoint as the student. - **Part 3:** https://github.com/NVIDIA/Model-Optimizer/pull/1601 - **Part 4:** repeat the NVFP4 + QAD experiments on a non-Nemotron model. Extends `examples/megatron_bridge/distill.py` to initialize the student from a **Megatron checkpoint** (a quantized checkpoint from `quantize.py`, or a pruned one) via `--student_megatron_path`, enabling **Quantization Aware Distillation (QAD)**: - `--student_hf_path` still builds the student architecture; `--student_megatron_path` supplies the (optionally quantized) weights. - For a quantized checkpoint, the ModelOpt quantize mode + base weights are restored onto the **plain student before the knowledge-distillation conversion** (`restore_sharded_modelopt_state` is a no-op once a model is already converted), so the distilled checkpoint stays exportable as a quantized model with `export.py`. **Upstream dependency / workaround:** `DistillationProvider.provide()` has no seam to transform the student before the KD conversion, so this patches `provide()` at the class level (via an `id()`-keyed registry, because the provider proxies instance-attribute assignment to its teacher once the teacher is set). A companion Megatron-Bridge PR adds a first-class `DistillationProvider.student_pre_conversion_hook`; from nemo:26.06 onwards the workaround should be removed and replaced with that hook (a removal note in `distill.py` documents exactly how). ### Usage ```bash # 1) PTQ -> quantized Megatron checkpoint (part 1) torchrun --nproc_per_node 2 quantize.py \ --hf_model_name_or_path Qwen/Qwen3-8B --quant_cfg fp8 --tp_size 2 \ --export_megatron_path /tmp/Qwen3-8B-FP8-megatron # 2) QAD: distill the quantized student from the unquantized teacher torchrun --nproc_per_node 8 distill.py \ --teacher_hf_path Qwen/Qwen3-8B \ --student_hf_path Qwen/Qwen3-8B \ --student_megatron_path /tmp/Qwen3-8B-FP8-megatron \ --data_paths 1.0 tokenized/data_text_document \ --train_iters 1000 --output_dir /output/qwen3_8b_qad # 3) export the distilled quantized checkpoint (part 1) torchrun --nproc_per_node 1 export.py \ --hf_model_name_or_path Qwen/Qwen3-8B \ --megatron_path /output/qwen3_8b_qad/checkpoints \ --export_unified_hf_path /tmp/qwen3_8b_qad_fp8_hf ``` ### Testing `tests/examples/megatron_bridge/test_qad.py` (validated on a 2-GPU NeMo `26.04` container): quantize a tiny Qwen3 at TP=2 → QAD distill from the quantized student → `export.py` to a unified HF checkpoint, asserting `hf_quant_config.json` is written (proves the quantize mode survived QAD). Includes a commented-out vLLM deployment check, validated locally (full flow passes; vLLM loads the export as `quantization=modelopt`). Existing normal/Puzzletron distillation tests still pass. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: N/A (new example feature; default behavior unchanged when `--student_megatron_path` is not set) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new dependencies) - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ✅ ### Additional Information Depends on a companion Megatron-Bridge PR adding `DistillationProvider.student_pre_conversion_hook` (the upstream replacement for the class-level `provide()` workaround). The Nemotron-3 tutorial NVFP4 + QAD experiments ship in part 3. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Quantization Aware Distillation (QAD) workflow to recover accuracy of quantized Megatron students and distill from quantized checkpoints. * CLI option to initialize a distillation student from a Megatron checkpoint and a structure-only load path for bridging. * **Documentation** * Expanded runnable quantize → QAD → export guidance and best-practice tips. * **Tests** * End-to-end test validating quantize → QAD → export artifacts. * **Chores / UX** * Clearer rank-aware messages, improved tokenizer padding handling, and more consistent export behavior (fixed export dtype). <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
de525973cf |
[minor] fix for GLM4.7 mtp module in PTQ (#1630)
### What does this PR do? Type of change: Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> `_keys_to_prefixes` now drops a top-level `"model"` key fragment instead of emitting it as a prefix. Without this guard, an inlined-MTP key like `model.layers.92.eh_proj.weight` would emit `"model"` → exporter wraps it as `"model*"` in `quantization_config.exclude_modules` → `fnmatch` in TRT-LLM matches every `model.layers.X.*` module → entire backbone is treated as unquantized → FP8 weights get loaded into BF16 buffers via `.view(bf16)` (which halves the last dim of an FP8 tensor) → loader crashes with `tensor a (5120) must match tensor b (2560) at non-singleton dimension 1`. The main-branch caller in `load_mtp_weights` (PR #1532) already filters inlined keys before invoking `_keys_to_prefixes`, so this guard is defense-in-depth on `main` today. It freezes the invariant in code rather than as a docstring caveat — the same regression cannot return if a future caller forgets to filter. Adds a focused unit test that pins the behavior. ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing - New unit test `tests/examples/llm_ptq/test_example_utils.py::test_keys_to_prefixes_drops_model_top_level`: passes with this PR, would fail without the guard. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed prefix handling in LLM post-training quantization examples: the top-level "model" prefix is now omitted when deriving exclusion prefixes to avoid overly broad wildcard exclusions. * **Tests** * Added test coverage verifying correct conversion of tensor keys to exclusion prefixes, including cases with top-level "model" keys. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> |
||
|
|
dbdff11a7f |
Drive PTQ example qformat choices from preset YAMLs (hf_ptq, multinode_ptq, megatron_bridge) (#1525)
### What does this PR do?
Type of change: Refactor
Replace the hardcoded `QUANT_CFG_CHOICES` / `KV_QUANT_CFG_CHOICES` dicts
in the PTQ
example scripts with a small `_load_preset_cfg_choices()` helper that
discovers the
available qformat names by listing
`modelopt_recipes/configs/ptq/presets/{model,kv}/`
and **eagerly loads every preset YAML into a plain dict at import** via
the existing
`load_config(...,
schema_type=QuantizeConfig).model_dump(exclude_unset=True)` path.
The directory listing becomes the source of truth for the `--qformat` /
`--kv_cache_qformat` CLI vocabulary.
> Note: an earlier revision used a lazy, copy-on-access `Mapping`. That
was overkill
> for these example scripts — the previous `mtq.*_CFG` module constants
were
> themselves eagerly-loaded shared dicts, and every call site that
mutates a config
> already deepcopies first — so it is now a plain eager dict. A lazy
variant can be
> reintroduced later if import time ever matters.
**Scope.** Three scripts carried the same hardcoded tables and all three
are migrated:
- `examples/llm_ptq/hf_ptq.py` — `--qformat` / `--kv_cache_qformat`.
- `examples/llm_ptq/multinode_ptq.py` — `--qformat` /
`--kv_cache_qformat`.
- `examples/megatron_bridge/quantize.py` — `--quant_cfg` /
`--kv_cache_quant`
(still also accepts any full `mtq.config.choices` name).
All three scripts share the discovery helper (`load_quant_cfg_choices`),
the canonical
alias table (`QFORMAT_ALIASES`), the KV disable sentinel
(`KV_CACHE_NONE`), and the
ready-built `QUANT_CFG_CHOICES` / `KV_QUANT_CFG_CHOICES` mappings via
the new
**`modelopt.recipe.presets`** module — no copy lives in the example
scripts anymore. The
module is a standalone import, so `import modelopt.recipe` stays cheap;
only an explicit
`import modelopt.recipe.presets` triggers the eager preset load.
A small alias table preserves previously-supported short CLI names
(`int8_sq`,
`nvfp4_awq`, `fp8_pb_wo`, …, plus the Megatron-Bridge `fp8_blockwise`)
as deprecation
shims. It is documented as not-for-extension — new formats land as
preset YAMLs, and
longer term, configurations should be authored as full recipes
(`--recipe`). The alias
logic is fail-fast: an alias pointing at a missing preset raises
`ValueError` at import.
Also adds `presets/kv/fp8_cast.yaml` and `presets/kv/nvfp4_cast.yaml`,
composed from the
existing `kv_fp8_cast` / `kv_nvfp4_cast` unit fragments. This promotes
`fp8_cast` /
`nvfp4_cast` to first-class KV presets and lets us delete the runtime
`_set_kv_cache_constant_amax` helper and all its call sites —
`use_constant_amax` is now
authoritative in the YAML. The KV-calibration-skip decision is derived
from the config
(`_kv_cfg_uses_constant_amax`), not from hardcoded format names.
**⚠️ CLI surface expansion (owner sign-off requested).** Because the
directory listing
is now the CLI vocabulary, each script accepts **every** preset under
`presets/{model,kv}/`, not just its previously curated subset. For
`hf_ptq.py` this is
the same surface the prior table covered; for `multinode_ptq.py` and the
Megatron-Bridge
script it is broader (e.g. KV `fp8_affine` / `fp8_cast` / `nvfp4_cast` /
`nvfp4_rotate`
are now selectable). This is intended ("the directory is the policy"),
but please confirm
those two scripts are meant to expose all presets — if a given path has
not validated a
format, it should be gated explicitly.
### Usage
```bash
# Old short names still work via the alias shim
python examples/llm_ptq/hf_ptq.py --pyt_ckpt_path <model> --qformat int8_sq --kv_cache_qformat fp8_cast --export_path out/
# Canonical preset basenames work directly
python examples/llm_ptq/hf_ptq.py --pyt_ckpt_path <model> --qformat int8_smoothquant --kv_cache_qformat fp8_cast --export_path out/
# A newly-added preset YAML is valid on the CLI of all three scripts with no code change
python examples/llm_ptq/hf_ptq.py --pyt_ckpt_path <model> --qformat nvfp4_awq_full --export_path out/
```
### Testing
- New: `tests/unit/recipe/test_presets.py` smoke tests for
`modelopt.recipe.presets` —
every discovered model/KV preset loads into a `quant_cfg` dict, the
directory listing is
fully covered, deprecation aliases resolve to their canonical preset,
the KV `none`
sentinel does not collide with a preset, and a stale alias raises. These
guard the eager
import-time load (one bad preset would otherwise break `import
modelopt.recipe.presets`
and every PTQ example).
- Previously verified locally (uv `.venv` py3.13 + `dev-py310-modelopt`
conda):
all previously-supported `--qformat` / `--kv_cache_qformat` names
resolve to dicts
bit-equal to the corresponding `mtq.*_CFG` constants; `fp8_cast` /
`nvfp4_cast` carry
`use_constant_amax: true` while non-cast variants do not; argparse
accepts
`--kv_cache_qformat none` plus all variants; unknown qformats raise at
lookup / argparse.
- All pre-commit hooks pass (ruff, mypy, bandit, license, rst, yaml).
- Pre-merge manual checks recommended by review (run in an env with the
deps installed):
`python examples/llm_ptq/hf_ptq.py --help`, `… multinode_ptq.py --help`,
`… megatron_bridge/quantize.py --help`.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — all previously-valid CLI
values continue to work via the alias table; output configs are
bit-equivalent to the prior hardcoded path.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
deps.
- Did you write any new necessary tests?: ✅ —
`tests/examples/llm_ptq/test_example_utils.py` preset-discovery smoke
tests.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ☐
### Additional Information
Out of scope / follow-up: the `_AUTO_QUANTIZE_QFORMATS` table and
`_canonical_qformat`
helper in `hf_ptq.py` are intentionally left hardcoded — auto_quantize
is being
refactored/reimplemented and they are expected to be removed soon.
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
bcbe2b957e |
Fix non-deterministic T5 calibration NaN on multi-GPU (#1636)
## Summary Fixes known error failing nightly 2-gpu CI 4/5 times - `tests/examples/llm_ptq/test_llm_ptq.py::test_ptq_t5` intermittently fails on multi-GPU runners with `AssertionError: detected nan values in amax. nan in original tensor: True` during FP8 calibration (in T5 encoder self-attention `self.o(attn_output)`). - Root cause: `device_map="auto"` is memory-aware, so on a busy 2-GPU box accelerate sometimes shards the tiny `t5-small` across both GPUs. T5 ties the encoder/decoder `shared` embeddings and relies on relative position-bias buffers; splitting these across devices via naive model-parallel hooks produces NaN activations (HF transformers [#21093](https://github.com/huggingface/transformers/issues/21093)). The placement varies with free memory at load time, which is why it failed ~4/5 runs but passed ~1/5 on the same machine. - Fix: load T5 on a single device (`device_map=None`) in `examples/llm_ptq/example_utils.py::get_model`, mirroring the existing BART handling. The existing `model.to(device)` path then places it on a single GPU, making calibration deterministic. `t5-small` is tiny, so single-device placement is not a memory concern. ## Test plan - Merge and see if nightly CI is no longer flaky |
||
|
|
433b549cd8 |
[2/N] Simplify KDTrainer and enhance ModelOptHFTrainer (#1191)
## Summary This PR simplifies the HuggingFace knowledge distillation trainer and enhances the base `ModelOptHFTrainer` with Liger fused loss, per-parameter learning rates, and training utilities. ### Model-agnostic Liger kernel fused loss Adds custom Liger kernel integration in `ModelOptHFTrainer` that extends HuggingFace's built-in support in three ways: 1. **Model-agnostic**: Works with any causal LM that has an `lm_head`, unlike HF's Liger which only supports [a fixed set of model architectures](https://github.com/linkedin/Liger-Kernel/blob/main/src/liger_kernel/transformers/monkey_patch.py). 2. **DeepSpeed ZeRO-3 support**: HF's Liger integration only works with FSDP. ModelOpt adds distributed param gathering for DeepSpeed ZeRO-3 and DDP as well. 3. **KD loss support**: `KDTrainer` extends fused loss to knowledge distillation via `LigerFusedLinearJSD` for fused lm_head + Jensen-Shannon divergence. #### Liger kernel memory sweep (Qwen3-1.7B, 2×H100 FSDP2, NVFP4+FP8_KV) Max per-GPU batch size before OOM at each sequence length: **QAT (no teacher)** | Seq Length | 512 | 1024 | 2048 | 4096 | 8192 | 16384 | |------------|-----|------|------|------|------|-------| | **Liger** | 16 | 16 | 16 | 16 | 8 | 4 | | **No Liger** | 16 | 16 | 8 | 4 | 2 | OOM | **QAD (with teacher)** | Seq Length | 512 | 1024 | 2048 | 4096 | 8192 | 16384 | |------------|-----|------|------|------|------|-------| | **Liger** | 16 | 16 | 8 | 4 | 2 | 1 | | **No Liger** | 8 | 4 | 2 | 1 | OOM | OOM | Liger fused loss enables **2-4× larger batch sizes** at long context lengths by avoiding the materialization of the full logit tensor. ### ModelOptHFTrainer enhancements - `ModelOptTrainerArguments` with `--trainable_params`, `--frozen_params`, `--lr_config`, `--save_dtype`, and `--manual_gc` flags - Per-parameter learning rate support via YAML config (`lr_config`) - `_prepare_model` and `_update_config_json_dtype` promoted to base class ### KDTrainer simplification + fix Removes `mtd.convert()` and the `DistillationModel` in-place class-swap for the HF path. The teacher model now lives directly on the trainer and is forwarded explicitly inside `compute_kd_loss_func`. This eliminates: - `mtd.convert()` in-place class swap and DynamicModule wrapping - Forward hooks for capturing intermediate outputs - `hide_teacher_model` / `hide_loss_modules` context managers for checkpointing - Deferred initialization branching (FSDP2 vs DDP/DeepSpeed) - `save_model` and `QADTrainer._quantize_model` overrides **Bug fix**: The previous `DistillationModel`/`mtd.convert()` approach did not support CPU RAM-efficient loading for QAD. The teacher model had to be fully loaded on GPU before wrapping, which doubled peak memory during initialization. The new approach loads the teacher lazily on the trainer, enabling standard HF device-map and low-cpu-mem-usage loading. Only logit-level distillation is supported for the HF path. The core `DistillationModel`/`mtd.convert()` API remains for Megatron and advanced intermediate-layer distillation use cases. ## Test plan - [x] `pytest tests/unit/torch/distill/` (29 passed) - [x] `pytest tests/unit/torch/opt/plugins/test_hf_patching.py` (2 passed) - [x] `pytest tests/unit/torch/opt/plugins/test_lr_config.py` - [x] Pre-commit hooks pass - [ ] GPU example tests: `pytest tests/examples/llm_qat/` (QAT, QAD, LoRA QAT, QLoRA) - [ ] GPU distill example: `pytest tests/examples/llm_distill/` 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added Liger fused loss support in `ModelOptHFTrainer` for distributed causal language models with JSD distillation loss support. * Introduced `ModelOptTrainerArguments` with new training CLI flags: per-parameter learning rates via YAML, parameter freezing, and manual garbage collection. * Simplified knowledge distillation trainer with logit-level distillation support. * **Documentation** * Updated example configurations and documentation with new training options and defaults. * Added learning rate configuration example guide. * **Tests** * Added test coverage for distillation training and per-parameter optimizer configuration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
115cae2584 |
Puzzletron bypass distillation part 2 core (#1469)
## Summary
This is PR 2 of 3 in the Puzzletron bypass/local-distillation stack.
This PR adds the bypass distillation core engine. It builds on PR 1’s
shared infrastructure, but it does not yet wire bypass into
the full Puzzletron pipeline.
Stack:
1. `ssameni/puzzletron-bypass-1-prereqs`: shared prerequisites
2. **This PR**: bypass distillation core
3. `ssameni/puzzletron-bypass-3-integration`: Puzzletron integration,
configs, docs, GPU coverage
## What Changed
- Added `modelopt.torch.puzzletron.bypass_distillation`.
- Added bypass run identity/fingerprinting, experiment naming, state
manifests, and completion tracking.
- Added stitched teacher/student model construction for local blockwise
distillation.
- Added bypass training loop with:
- pipeline-parallel teacher activation stitching
- per-block student losses
- gradient accumulation
- checkpoint/resume support
- validation hooks
- best/latest checkpoint realization
- Added bypass checkpoint helpers for saving optimizer/scaler state and
HF-format model checkpoints.
- Added scalable distributed checkpoint saving that gathers tensors per
safetensors file instead of materializing all shards on
rank 0.
- Restricted v1 `keys_to_learn` to subblock-level targets only:
- `entire_block`
- `subblock_attention`
- `subblock_ffn`
- `subblock_mamba`
- lists of those keys
- Added pipeline ownership helper for deriving owned blocks and
neighboring PP ranks.
- Added robust `trust_remote_code` / `auto_map` handling for checkpoint
config saves.
## Why
Bypass distillation is the local-distillation stage used to train
pruned/reconfigured blocks before they are added to the
replacement library.
Keeping the core engine separate from Puzzletron pipeline wiring makes
this PR reviewable on its own: reviewers can focus on
distributed training, checkpoint/resume semantics, and Sewing Kit
stitching behavior without also reviewing configs/docs/pipeline
integration.
## Tests
Added focused unit coverage for:
- bypass checkpoint utilities
- bypass run identity/fingerprinting and completion state
- `keys_to_learn` subblock selection
- LR scheduler behavior
- launch/sweep dispatch behavior
- HF checkpoint utility behavior
- stitched model factory buffer ownership
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Full bypass distillation workflow: blockwise stitched student/teacher
distillation, deterministic experiment IDs/fingerprints, resume-capable
checkpointing with atomic symlink updates, optional asynchronous
checkpoint saves, and improved support for trust-remote-code model
loading.
* Public bypass-distillation entrypoint exposed for easier launch.
* **Tests**
* Extensive unit tests covering bypass utilities, checkpoint behavior,
LR scheduler, stitched factory, and training orchestration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Sepehr Sameni <ssameni@nvidia.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
0081473861 |
Speed up slow unit/gpu/example tests (#1616)
### What does this PR do?
Type of change: test infrastructure / test speedups + CI stabilization
Make the test suite faster, `tests/unit` hermetic, and the CI lanes
stable, without losing coverage. Most changes are mechanical test/infra
edits; the buckets below cover the diff broadly.
**Unit tests — hermetic (no HF Hub):** toy local datasets/configs + the
local tiny tokenizer (with a checked-in chat template) replace Hub
assets; `tests/unit/conftest.py` enforces offline mode. Genuinely-HF
tests moved to `tests/gpu*` (e.g. the new
`tests/gpu/torch/utils/test_dataset_utils.py`). `CONTRIBUTING.md`
documents the hermetic-unit-test expectation.
**Unit-test speedups (no coverage loss):** speculative (disable CPU
torch.compile), calibrator (fewer histogram bins), ONNX conv/dynamo
(smaller shapes + representative subset), Ruler/sparse-attention (local
tokenizer), data-parallel autoquant (world size 4→2). Shared
`tiny_tokenizer` fixture. The distributed test helper now uses a private
`spawn` context instead of mutating the global start method (avoids
cross-test contamination).
**Rarely-used autonas/fastnas tests:** heavy parametrize cases marked
`@pytest.mark.manual`, one representative kept per test (fastnas
preferred); lighter sibling tests still cover core behavior. The legacy
FSDP1 NAS distributed test is also dropped: FastNAS/AutoNAS aren't used
with either FSDP1 or FSDP2, and FSDP1 is superseded by the newer FSDP2
API — so we keep a single FSDP2 case as a sanity check and drop FSDP1,
leaving the suite leaner.
**gpu_megatron:** deduplicate distributed worker pools by world_size
within a module (saves a redundant pool spin-up in multi-pool files;
module-scoped, no cross-module reuse).
**Example tests:** reduce per-test work via args that default to current
behavior (tests pass the fast values) — torch_onnx TRT optimization
level, diffusers calibration/inference steps, eagle `sample_size`,
megatron_bridge iters/calib, llm_sparsity data slice, export
safetensors-structure `calib_size`. Also enable the recently added
`gpt-oss` example tests in CI.
**Per-test timeouts:** `pytest-timeout` with a default per-directory
timeout (60s unit / 300s gpu+example) enforced in `tests/conftest.py`
(`timeout_func_only` in `pyproject.toml`), so a new test cannot silently
exceed the budget — an unmapped test dir crashes collection. A few
inherently slow tests carry explicit higher per-test overrides
(CUDA-compile, autotune, dflash).
**CUDA kernel pre-compilation:** a dedicated `tests/gpu/_extensions`
test JIT-builds the conv3d implicit-GEMM kernel up front (collected
before the functional tests in the same process) so the one-time build
cost no longer lands on — and time out — the first functional test that
uses it. Mirrored into the `llm_ptq`/`vlm_ptq` example lanes.
**Test relocation & optional-dependency guards:** vLLM sparsity plugin
test moved to `tests/gpu_vllm` (drops the in-test `importorskip`);
diffusers-dependent unit test guarded with `importorskip("diffusers")`
for partial-install lanes; `gpt_oss` example test dir renamed to
`gpt-oss` to match the CI matrix.
**Diffusers test models:** shared model-path constants in
`tests/_test_utils/examples/models.py` consolidated/renamed and point at
tiny `hf-internal-testing` test pipes (SDXL/SD3/FLUX) so
cachify/quantize/export tests run on toy weights; `local_id`s
normalized.
**Shared dataset utils:** `examples/llm_sparsity/.../hf_pts.py` now uses
`get_dataset_dataloader` (drops the bespoke cnn_dailymail-only
`get_calib_dataloader`; supports any registered/HF/JSONL dataset,
includes attention_mask); `data_prep.py` gains `--max_samples`.
**CI workflows:** container image bumps (pytorch 26.04→26.05, TRT-LLM
rc16→rc17) and tightened lane timeouts (unit 30→15 min, gpu lanes
trimmed, onnx example lane 45 min).
**Imports at top of file:** in-function imports across the test suite
are moved to module top per the coding guideline, conservatively —
optional deps stay guarded (in-function or behind a module-level
`importorskip`) in `tests/unit` since the partial-install lane runs
without them, and build/hardware-availability imports (apex, triton,
megatron/transformer_engine, tensorrt_llm) plus `_test_utils` lazy
guards are left in place.
**Kernel warning filters:** the repeated `filterwarnings` blanket-ignore
in six `tests/gpu/torch/kernels/**` modules is consolidated into a
scoped hook in `tests/gpu/torch/kernels/conftest.py` (kernel tests only
— the rest of the suite keeps surfacing warnings).
**Eagle example speedups:** `torch.compile` (eagle recipe default) added
~2 min to every eagle training test; it's now disabled in the eagle
example tests except one smoke (`test_llama_eagle3[1-False]`), and the
downstream resume / AR-validate / export tests point at the compile-free
checkpoint. Measured: `test_ar_validate` 139s→17s, offline training
142s→22s, streaming 140s→23s — the compile path is still smoke-tested
once.
**Example lanes install editable (`-e`):** so example scripts launched
as subprocesses resolve `modelopt` to the same source path as the test
process and reuse the pre-compiled CUDA-extension cache instead of
recompiling (~2 min/test); verified in the TRT-LLM container.
**Tiny test tokenizer:** `get_tiny_tokenizer` defaults to left padding
(what decoder-LM calibration expects) and ships a terse
generation-tagged chat template — replacing a verbose ChatML one that
inflated tokenized length on the 128-vocab tokenizer and broke the
offline-PTQ example tests' `max-seq-len` filter.
**Restored Hub-download coverage:** the live (ungated) HF dataset
round-trips exercising `get_dataset_samples`' download branch now live
in `tests/gpu/torch/utils/test_dataset_utils.py` (they had been dropped
from the hermetic unit file without a counterpart).
Individual file changes not explicitly called out above fall under this
general test/CI cleanup.
### Testing
Unit + the touched gpu_megatron files validated locally; example/GPU
lanes validated in CI.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (tests + example CLI args
default to prior behavior)
- If you copied code from any other sources or added a new PIP
dependency: N/A
- Did you write any new necessary tests?: N/A (optimizes/relocates
existing tests)
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ❌ (pending)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **Chores**
* Updated container image versions for PyTorch (26.04→26.05),
TensorRT-LLM (1.3.0rc16→1.3.0rc17), and ONNX/TensorRT (26.04→26.05).
* **Tests**
* Enhanced test isolation: unit tests now run hermetically without
HuggingFace Hub access.
* Optimized test runtime via smaller model/dataset parameters and
parallel test caching.
* Added CUDA extension availability tests and extended dataset utility
coverage.
* **Documentation**
* Updated testing guidelines in `CONTRIBUTING.md` to emphasize offline
test design.
* **Chores**
* Added pytest timeout configuration and improved CI/CD workflow
efficiency with editable installs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
ca7eb64ad0 |
DSV4 PTQ example with dequant on the fly (#1341)
### What does this PR do? Type of change: new example <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Add deepseek v4 official modeling ptq example ### Usage See readme, and it requires the vllm PR: https://github.com/vllm-project/vllm/pull/42209 ### Testing Tested with ptq and export of dsv4 flash and served with vllm. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added DeepSeek‑V4 routed‑expert post‑training quantization and an NVFP4 checkpoint conversion utility. * **Documentation** * Expanded DeepSeek quantization guide with directory layout, updated V3/V3.2 workflows, and detailed V4 routed‑expert calibration, single/multi‑node examples, and export guidance. * **Chores** * Made example quantization scripts location‑independent. * Updated pre‑commit license hook to skip DeepSeek example quantization files. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1341?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Meng Xin <mxin@nvidia.com> |
||
|
|
5bd04c3876 |
Revert unverified EAGLE3 model examples; keep triage code + baseline (#1623)
### What does this PR do? Type of change: Revert / cleanup (follow-up to #1417) Per review feedback (@h-guo18): `main` should be production-ready and user-facing. Most of the EAGLE3 model example YAMLs added in #1417 are not yet verified to work end-to-end in modelopt (~80% fail at some pipeline stage), which is confusing to ship. The agreed plan is to **land the triage infrastructure now and re-add each model's launcher YAML in a dedicated follow-up PR once it is verified green**. **Removed** (unverified, to be re-added per-model once verified): - Per-model launcher configs (`hf_offline_eagle3.yaml` + `eagle3_quick_check.yaml`) for: DeepSeek-V3.2, GLM-5, MiniMax-M2.5, Ministral-3-8B, Ministral-3-14B, Kimi-K2.5, Kimi-K2.5-NVFP4, GPT-OSS-20B, Qwen3.5-9B, Qwen3.5-27B, Qwen3.5-35B-A3B, Step-3.5-Flash. - Per-model status docs: `tools/launcher/examples/EAGLE3_TRIAGE.md`, `examples/speculative_decoding/pipeline/eagle3/eagle3_triage_chart.md` (volatile status — tracked internally instead). **Kept** (the durable triage infrastructure from #1417): - Verified baseline example `tools/launcher/examples/Qwen/Qwen3-8B/eagle3_quick_check.yaml`. - Launcher common scripts (vLLM native-extractor dump, etc.) and `compute_hidden_states_vllm.py`. - modelopt code fixes: FakeBaseModel VLM detection, `consolidated.safetensors` load, `use_cache` export templates. - New-model triage guide (`eagle3_new_model_triage_guide.md`); its "document results" step now points at the internal tracker rather than the removed chart. ### Testing No code paths change — this only removes example YAMLs and two status docs and edits one doc reference. Pre-commit (ruff/markdownlint/yaml/license) passes on the kept/edited files. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (removes unverified examples only; kept infra unchanged) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: ❌ (pending) ### Additional Information Follow-up to #1417. Next step (tracked separately): verify each removed model end-to-end in modelopt, then re-add its YAML in a dedicated PR. Note: the nmm-sandbox weekly EAGLE3 CI is being trimmed to the Qwen3-8B baseline to match. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated EAGLE3 triage guide to streamline the verification workflow; contributors now record test outcomes (status, experiment IDs, errors, and applied fixes) in the team's internal triage tracker before submitting model launcher configurations. * **Chores** * Removed legacy EAGLE3 example pipeline configurations and deprecated triage documentation to reduce maintenance overhead. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
a7b0a92047 |
EAGLE3 new model support: pipeline configs, triage docs, and Ministral-3 fixes (#1417)
## Summary EAGLE3 automation triage work (OKR-30): testing the 4-step EAGLE3 offline pipeline against 12 new model architectures, documenting failure modes, and fixing issues found. ### Code fixes (modelopt) | File | Change | |------|--------| | `modelopt/torch/speculative/utils.py` | Extend VLM detection in `load_vlm_or_llm` to check `text_config`/`llm_config` attrs (catches `mistral3` models) | | `modelopt/torch/speculative/plugins/modeling_fakebase.py` | Add `consolidated.safetensors` fallback for checkpoints with incomplete HF shards | | `modelopt/torch/export/plugins/hf_spec_configs.py` | Set `use_cache=True` in EAGLE export templates (fixes strict `huggingface_hub` validation) | ### Pipeline infrastructure - `examples/speculative_decoding/pipeline/eagle3/` — pipeline scripts and configs: - `offline_training.sh` — training + export with runtime patches for older container modelopt - `dump_offline_data_vllm.sh` — vLLM-based hidden state extraction (with speculators compat patches) - `dump_offline_data.sh`, `dump_offline_data_hf.sh` — alternative dump paths - 18 quick-fail-check YAMLs for 12 models - 4 standalone task1 YAMLs ### Documentation - `eagle3_triage_chart.md` — model test matrix, triage decision tree, per-model results, failure catalog - `eagle3_new_model_triage_guide.md` — step-by-step guide for triaging new models ### Model test results (as of 2026-05-27) | Model | task_0 | task_1 | task_2 | task_3 | Blocker | |-------|--------|--------|--------|--------|---------| | Qwen3-8B | - | - | - | - | Reference (existing) | | Kimi-K2.5 | - | - | - | - | Existing (GB200) | | **Ministral-3-8B** | SKIP | PASS | PASS | FAIL | `use_cache=null` in export (fixed) | | Ministral-3-14B | FAIL | - | - | FAIL | vLLM engine init fails | | Qwen3.5-35B-A3B | TIMEOUT | - | - | - | Data synth too slow | | gpt-oss-20b | FAIL | - | - | - | Tokenizer `HarmonyError` | | Step-3.5-Flash | TIMEOUT | - | - | - | Data synth time limit | | MiniMax-M2.5 | TIMEOUT | - | - | - | `trust_remote_code` needed | | DeepSeek-V3.2 | no log | - | - | - | May not be mirrored | | Qwen3.5-9B | - | - | - | - | Not yet run | | Qwen3.5-27B | - | - | - | - | Not yet run | | GLM-5 | - | - | - | - | Not yet run | ### Issues found and fixed | # | Issue | Fix | |---|-------|-----| | 1 | `mistral3` model type not detected as VLM | Check `text_config`/`llm_config` attrs in `load_vlm_or_llm` | | 2 | Missing HF shard file (Ministral-3-8B) | Fallback to `consolidated.safetensors` with Mistral native key aliases | | 3 | `use_cache=null` in exported EAGLE config | Set `use_cache=True` in export template configs | | 4 | speculators incompatible with vLLM container | Runtime patches in `dump_offline_data_vllm.sh` | | 5 | `offline_training.sh` infra issues | Rewritten with runtime patches for container modelopt | ## Test plan - [x] Ministral-3-8B training passes (`cicd_1779829129`) - [x] Ministral-3-8B export succeeds - [ ] Ministral-3-8B benchmark passes (`cicd_1779901409` — pending with all fixes) - [ ] Dry-run remaining model configs ## Note GitHub secret scanning alert #6 is a **false positive** — `Mistral3ForConditionalGeneration` (a HuggingFace model class name in a YAML comment) was flagged as a "Mistral AI API Key". 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
651fd223e6 |
feat: EAGLE3 LoRA co-training improvements (#1607)
Layer-selective LoRA injection for EAGLE3 co-training, optimizer-stable warmup (LoRA always in the optimizer; warmup gated by a flag), and an export+merge+lm_eval evaluation script. Review feedback addressed: - trust_remote_code is caller-controlled (TRUST_REMOTE_CODE / --trust_remote_code), default False - eagle_base_lora_start_layer raises ValueError instead of silently injecting zero adapters - eval_lora.sh validates HF_MODEL_CKPT / EAGLE_CKPT up front - added unit tests for start-layer injection and validation (8/8 passing, incl. on-cluster GPU run) Signed-off-by: Ye Yu <yeyu@nvidia.com> |
||
|
|
e40b4d69d6 |
Add Minitron pruning support for Gemma3 via Megatron-Bridge (#1604)
### What does this PR do?
Type of change: New feature (+ new tests)
Adds Minitron pruning support for **Gemma3** models loaded through
Megatron-Bridge (`Gemma3ForCausalLM` → `GPTModel`).
Gemma3's bridge implementation subclasses several megatron-core modules
and overrides `forward`/`__init__`, so `DMRegistry` (which matches by
`nn_cls.forward is parent.forward`) does not recognize them and
pruning's `convert_to_dynamic` raised:
```
KeyError: "<class '...Gemma3LanguageModelEmbedding'> is not registered for a dynamic module!"
```
There were actually **two** blockers, both fixed here:
1. **Layer spec was discarded.** `load_mbridge_model_from_hf`
unconditionally replaced `provider.transformer_layer_spec` with the
generic TE GPT spec (to disable grouped-GEMM), throwing away
`gemma3_layer_spec`. With the generic spec, plain
`TEDotProductAttention` received Gemma3's int `window_size` and crashed
at construction (`TypeError: 'int' object is not subscriptable`) before
pruning even started. Now the override is applied **only to MoE
models**; dense models keep the bridge's native spec.
2. **Custom layers weren't registered as dynamic modules.** Added
registrations for Gemma3's custom layers.
Changes:
- **`modelopt/torch/nas/plugins/mbridge.py`** (new): register dynamic
modules for `Gemma3LanguageModelEmbedding` (reuses the existing
embedding dynamic class), `Gemma3SelfAttention`, and the fused post-LN
`TERowParallelLinearLayerNorm` used by both `self_attention.linear_proj`
and `mlp.linear_fc2`. The fused `post_layernorm` is converted to a
dynamic module and sliced by the linear's `output_size` (==
`hidden_size`).
- **`modelopt/torch/nas/plugins/megatron.py`**: two small reuse seams in
`_DynamicSelfAttention` — build the core-attention dynamic class from
`type(self.core_attention)` (preserves `Gemma3TEDotProductAttention`)
and extract an overridable `_convert_linear_proj` hook. No behavior
change for megatron-core models.
- **`modelopt/torch/utils/plugins/mbridge.py`**: only override the layer
spec for MoE models; dense models keep the bridge's native spec.
- **`modelopt/torch/nas/plugins/__init__.py`**: import `mbridge` under
`import_plugin("megatron.bridge")`.
- **Tests**: new `get_tiny_gemma3` / `create_tiny_gemma3_dir` helpers;
parametrize `test_prune_minitron` over qwen3 and gemma3.
### Usage
```bash
# Prune a Gemma3 checkpoint with Minitron (same flow as other models)
torchrun --nproc_per_node 2 examples/megatron_bridge/prune_minitron.py \
--hf_model_name_or_path google/gemma-3-1b-it \
--prune_target_params 0.6e9 \
--output_hf_path /tmp/gemma-3-pruned
```
### Testing
`tests/examples/megatron_bridge/test_prune_minitron.py` now runs for
both qwen3 and gemma3. Verified in the megatron-bridge container:
Also verified the gemma3 case standalone on 1 GPU (PP=1) and 2 GPUs
(PP=2).
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (the spec-override change only
affects how dense Megatron-Bridge models are loaded — they now keep
their native, correct spec; MoE behavior is unchanged)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ (parametrized gemma3
coverage + tiny-model helper)
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ✅
### Additional Information
Other Megatron-Bridge models with custom layers were surveyed; only the
Gemma family needs this treatment. **Gemma1** (embedding scaling via
`EmbeddingScalingMixin`) and **Gemma2** (post-LN linear + non-TE
`Gemma2DotProductAttention` + custom output layer + mixin embedding) are
intentionally **left as outdated**. All other dense/MoE LLMs (llama,
qwen2/3, qwen3-moe, mistral, nemotron, deepseek, gpt-oss, etc.) use
standard megatron-core specs already covered by `megatron.py`.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added Minitron pruning support for Megatron-Bridge Gemma3 models.
* Improved Megatron-Bridge layer configuration handling for
Mixture-of-Experts scenarios.
* **Tests**
* Added Gemma3 model creation utilities for testing.
* Expanded pruning validation tests to cover both Qwen3 and Gemma3
models.
* **Examples**
* Updated pruning example to adjust tokenizer "use fast" detection based
on model architecture.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
196c091027 |
[1/N] Refactor llm_qat example: YAML configs + ModelOptArgParser (#1172)
### What does this PR do? Type of change: new example Refactors `examples/llm_qat` from a monolithic launch/script flow into a modular, config-driven Hugging Face QAT/QAD workflow. The high-level user flow is now: 1. Quantize a base model with a ModelOpt PTQ recipe. 2. Train or evaluate the quantized checkpoint with QAT, QAD, LoRA QAT, QLoRA, or fine-tuning configs. 3. Export the trained checkpoint for deployment. Highlights: - Replaces the legacy `examples/llm_qat/launch.sh` + `examples/llm_qat/main.py` path with separate `quantize.py` and `train.py` entrypoints. - Adds YAML-driven argument parsing through `ModelOptArgParser`, including `--config <yaml>` defaults, CLI overrides, and generated `examples/llm_qat/ARGUMENTS.md`. - Adds declarative configs under `examples/llm_qat/configs/` for training modes, dataset blends, and Accelerate backends. - Adds `dataset_utils.py` for weighted multi-source dataset blending, Hugging Face streaming, local dataset loading, distributed rank-aware loading, tokenization caching, pre-tokenization, chat templating, and assistant-token label masking. - Moves the Hugging Face QAD flow into the shared trainer path through `DistillArguments`, `QADTrainer`, and teacher-model distillation kwargs. - Updates `QuantizationArguments` to prefer recipe paths via `--recipe`, while keeping legacy `--quant_cfg` available with deprecation warnings for in-trainer quantization. - Updates `QATTrainer` handling for pre-quantized checkpoints, FSDP2 TensorQuantizer buffers, and recipe-resolved PTQ configs. - Adds the `general/ptq/int4_blockwise_weight_only` recipe and refreshes docs for NVFP4, FP8, INT4, FSDP2, DDP, DeepSpeed, QLoRA, and LLaMA-Factory integration. - Adds focused parser, dataset tokenization, assistant-mask, and example workflow coverage. ### Usage From the repo root: ```sh cd examples/llm_qat # 1. Quantize python quantize.py \ --model_name_or_path Qwen/Qwen3-8B \ --dataset_config configs/dataset/blend.yaml \ --recipe general/ptq/nvfp4_default-kv_fp8 \ --output_dir qwen3-8b-quantized # 2. QAT train accelerate launch --config-file configs/accelerate/fsdp2.yaml train.py \ --config configs/train/qat_nvfp4.yaml \ --model_name_or_path qwen3-8b-quantized \ --output_dir qwen3-8b-qat-nvfp4 # 3. QAD train accelerate launch --config-file configs/accelerate/fsdp2.yaml train.py \ --config configs/train/qad_nvfp4.yaml \ --model_name_or_path qwen3-8b-quantized \ --teacher_model Qwen/Qwen3-8B \ --output_dir qwen3-8b-qad-nvfp4 ``` Dataset blends can be pre-tokenized and cached before training: ```sh cd examples/llm_qat python dataset_utils.py \ --dataset_config configs/dataset/blend.yaml \ --model_name_or_path Qwen/Qwen3-8B ``` ### Testing Focused coverage added or updated: - `tests/unit/torch/opt/plugins/test_modelopt_arg_parser.py` - `tests/examples/llm_qat/test_dataset_tokenization.py` - `tests/examples/llm_qat/test_assistant_mask.py` - `tests/examples/llm_qat/test_llm_qat.py` Recorded validation: - [x] `pytest tests/unit/torch/opt/plugins/test_modelopt_arg_parser.py` - [x] `pytest tests/examples/llm_qat/test_llm_qat.py::test_dataset_utils_pretokenize` - [x] `pre-commit run --all-files` - [ ] `pytest tests/examples/llm_qat/test_llm_qat.py` full GPU/backend suite ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: No for the legacy `examples/llm_qat` CLI/file layout (`main.py`, `launch.sh`, and FSDP1 config are removed); library `quant_cfg` usage remains available but is deprecated for this workflow. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: Yes. `train.py`/`simple_qat_train.py` retain the upstream Alpaca attribution where applicable; `examples/llm_qat/requirements.txt` switches `tensorboardX` to `tensorboard`. - Did you write any new necessary tests?: Yes. Parser, dataset tokenization, assistant masking, pre-tokenization, and QAT/QAD workflow tests were added or updated. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: Yes. - Did you get Claude approval on this PR?: N/A in this description update. ### Additional Information This PR intentionally changes the `llm_qat` example surface. Existing users should move from `examples/llm_qat/main.py` and `examples/llm_qat/launch.sh` to `examples/llm_qat/quantize.py`, `examples/llm_qat/train.py`, and the YAML configs under `examples/llm_qat/configs/`. --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
902d36921a |
[Feat]: Streaming Hidden-states Dataset (#1509)
### What does this PR do? Type of change: new feature **Design doc:** https://gist.github.com/h-guo18/241c94968b0591324c361d97cf995dd0 Jira ticket: https://jirasw.nvidia.com/browse/OMNIML-4341 Sandbox CI: https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/327911555#L2228 Streaming hidden-states dataset: per-sample activations pulled from a live `vllm serve` over HTTP, replacing on-disk activation dumps. Two axes for future extensions: - **Backend** (`_fetch`): vLLM now; TRT-LLM / SGLang next. - **Algorithm** (`_format`): Eagle now; distillation / probing next. - **Sandbox CI**: https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/merge_requests/169 Shared plumbing — async producer, token-level truncation to `training_seq_len`, loss-mask alignment, DDP via Accelerate's dispatcher (rank 0 fetches, broadcasts), circuit breaker, resume — lives in `StreamingDataset`. First instance: **`EagleVllmStreamingDataset`**. API: `data.mode ∈ {online, offline, streaming}`; legacy configs auto-promote. ### Usage ```yaml data: mode: streaming data_path: input_conversations/train.jsonl streaming_server_url: http://localhost:8000 streaming_model_name: meta-llama/Llama-3.1-8B-Instruct training: training_seq_len: 4096 # also caps the prompt sent to vllm ``` Requires `vllm serve` with `ExampleHiddenStatesConnector` and `dataloader_num_workers=0`. End-to-end Slurm pipeline: `tools/launcher/examples/Qwen/Qwen3-8B/hf_streaming_eagle3.yaml`. ### Testing - **Unit**: full-corpus invariant, rank-0-only iter, resume, circuit breaker, mocked-httpx integration. - **E2E**: `launch_train.sh` against a stdlib `HTTPServer` mimicking the connector. - **Smoke** (Qwen3-8B / 8×H100 / 4096 ultrachat samples, single epoch): train_loss 32 → 18, MT-Bench AR 1.003 → 1.20. ### TODO before un-drafting - [ ] Observability counters (filtered, fetch failures, queue depth, latency). - [ ] Changelog entry. - [x] Add test in sandbox. **Non-goals (v1):** multi-epoch streaming, cross-rank dynamic load balancing. ### Before your PR is "*Ready for review*" - Backward compatible: ✅ (legacy configs auto-promote) - New PIP dep: N/A - New tests: ✅ - Changelog: ❌ (TODO above) - Claude approval: ❌ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Streaming training mode with server-backed hidden-state fetching, deterministic seed control, resume support, and streaming-specific dataset options (server, model, prefetch, shared storage). * **Behavior / Bug Fixes** * Stronger mode validation; offline behavior derived from data mode; resume handling adjusted to avoid double-skip during streaming runs. * **Tests** * End-to-end CI streaming test and expanded unit tests covering streaming, resume, DDP, determinism, and failure cases. * **Infrastructure** * Launcher script and pipeline config for end-to-end streaming training. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1509?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
f21977a5fc |
Add Megatron-Bridge PTQ quantize + export example scripts (#1589)
### What does this PR do?
Type of change: new example
Adds a two-step post-training quantization (PTQ) flow for
**Megatron-Bridge** models under `examples/megatron_bridge/`, mirroring
the Megatron-LM `quantize.sh` / `export.sh` split:
- **`quantize.py`** — loads an HF model via Megatron-Bridge, applies
ModelOpt PTQ (via a `--quant_cfg` alias / full config name, or a
`--recipe` YAML), with optional KV-cache quant, weight-only,
compression, and MoE expert-ratio calibration, then saves a **Megatron
checkpoint** (with ModelOpt state). Tensor / pipeline / expert
parallelism are all supported, and the checkpoint can later be reloaded
for further training (QAT / distillation).
- **`export.py`** — loads the quantized Megatron checkpoint, **re-shards
to TP=1**, and exports a **HuggingFace (unified)** checkpoint deployable
with TensorRT-LLM / vLLM / SGLang.
**Why the split?** The unified HF exporter (`export_mcore_gpt_to_hf`)
does not gather tensor-parallel-sharded weights — Megatron-LM likewise
forces `TP=1` during its export step. Saving a TP-sharded Megatron
checkpoint first lets us calibrate at TP>1 (to fit large models) and
then reload re-sharded to TP=1 for the HF export. A combined
single-script flow silently produced corrupt HF checkpoints under TP>1
(collided per-rank shards), which this split avoids.
> **Note:** This is **part 1 of 4**:
> - **Part 1 (this PR):** Megatron-Bridge `quantize.py` + `export.py`
support and tests.
> - **Part 2:** extend `distill.py` for quantization-aware distillation
(QAD) — load a quantized Megatron checkpoint as the student.
> - **Part 3:** add NVFP4 + QAD-on-pruned-checkpoint experiments to the
Nemotron-3-Nano-30B-A3B tutorial.
> - **Part 4:** repeat the NVFP4 + QAD experiments on a non-Nemotron
model.
### Usage
```bash
# Step 1: quantize (TP/PP/EP supported) -> Megatron checkpoint
torchrun --nproc_per_node 2 quantize.py \
--hf_model_name_or_path Qwen/Qwen3-8B \
--quant_cfg fp8 \
--tp_size 2 \
--export_megatron_path /tmp/Qwen3-8B-FP8-megatron
# Step 2: export -> deployable HuggingFace (unified) checkpoint (re-shards to TP=1)
torchrun --nproc_per_node 1 export.py \
--hf_model_name_or_path Qwen/Qwen3-8B \
--megatron_path /tmp/Qwen3-8B-FP8-megatron \
--export_unified_hf_path /tmp/Qwen3-8B-FP8-hf
```
### Testing
`tests/examples/megatron_bridge/test_quantize.py` (validated on a 2-GPU
NeMo `26.04` container):
- `test_quantize_export_and_vllm_deployment` — quantize a tiny Qwen3 via
a recipe at TP=2 → `export.py` re-shards to TP=1 → load + generate with
**vLLM** (skipped if vLLM absent).
- `test_quantize_megatron_checkpoint_reload` — quantize at TP=2 → reload
the Megatron checkpoint via the bridge and assert ModelOpt quantizers
were restored.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: N/A (new example)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new
dependencies)
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅
### Additional Information
The Nemotron-3 tutorial update to use these scripts is intentionally
**not** included here — it ships with the part 3 PR alongside the NVFP4
+ QAD experiments.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Expanded post-training quantization (PTQ) workflow documentation with
detailed step-by-step examples and configuration guidance for the
Megatron-Bridge framework.
* **New Features**
* Added quantization tool for applying PTQ to Megatron models with
calibration support.
* Added export tool for converting quantized models to a deployable
format.
* **Tests**
* Added integration tests validating the complete
quantization-export-deployment workflow, including inference validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
72df833e0b |
Add active-MoE AutoQuant cost accounting (#1497)
### What does this PR do?
• Type of change: new feature
Adds an active_moe cost model for auto_quantize effective-bits search.
This lets AutoQuant account for routed MoE expert weights by active
decode weight traffic instead of total checkpoint weight
size, using active_moe_expert_ratio = num_experts_per_tok / num_experts.
The default behavior is unchanged: cost_model="weight" still counts all
quantizable weights equally.
### Usage
import modelopt.torch.quantization as mtq
model, search_state = mtq.auto_quantize(
model,
constraints={"effective_bits": 5.0},
quantization_formats=[
mtq.NVFP4_DEFAULT_CFG,
mtq.FP8_DEFAULT_CFG,
],
data_loader=calib_dataloader,
forward_step=forward_step,
loss_func=loss_func,
cost_model="active_moe",
# Optional. If omitted, ModelOpt tries to infer this from model.config.
active_moe_expert_ratio=2 / 64,
)
The HF PTQ example also exposes:
--auto_quantize_cost_model active_moe \
--auto_quantize_active_moe_expert_ratio 0.03125
### Testing
python -m pytest tests/unit/torch/quantization/test_autoquant.py -q -k
'active_moe or quant_recipe_hparam_cost_weight'
python -m pytest tests/unit/torch/quantization/test_autoquant.py -q -k
'not data_parallel_auto_quantize'
Results:
- 4 passed
- 58 passed, 1 deselected
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added active-MoE cost model option for auto-quantization with
configurable expert ratio; API and CLI accept cost_model and
active_moe_expert_ratio
* Unified auto-quantize supports new quant format w4a16_nvfp4
* **Bug Fixes**
* Ensure labels are moved to the logits device for base models without
an lm_head
* CLI enforces valid expert-ratio range and requires active-MoE mode
when a ratio is provided
* **Tests**
* Added unit tests for active-MoE behavior, cost-weighting, ratio
handling, and search budget selection
<!-- review_stack_entry_start -->
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1497?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)
<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
|
||
|
|
7ae4ee7afd |
Force vLLM non-gated MoE through Triton (#1572)
### What does this PR do? Type of change: bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> From v0.20.0, vLLM selects the FlashInfer CUTLASS unquantized MoE backend for Nano-style non-gated MoE layers. That fused backend hides the intermediate activation between the expert GEMMs, so the w2 input quantizer can not be inserted. Solution: Add the parameter to force use triton kernel. ### Usage ### Testing Tested on Nano3 ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated vLLM fakequant serve documentation with notes on moe-backend configuration defaults. * **Improvements** * Enhanced launcher logic to conditionally set moe-backend defaults only when the feature is available, improving configuration flexibility across different vLLM installations. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1572?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Meng Xin <mxin@nvidia.com> |
||
|
|
40a4dd326d |
[Feat]: Eagle Dry Run Mode (#1566)
### What does this PR do? Type of change: new feature Adds `--dry_run` to `examples/speculative_decoding/main.py`: load → `mtsp.convert` → save, then exit (no `trainer.train()`). With `FakeBaseModel`, the convert→save→export chain runs in seconds and produces an exportable EAGLE3 / Medusa / DFlash checkpoint with correct structure but untrained draft-head weights — useful for end-to-end plumbing smoke tests on downstream stacks (vLLM, TRT-LLM, SGLang) without paying for a real training run. Where it sits among existing EAGLE3 modes: ``` EAGLE3 modes ├── online base model runs forward in-loop ├── offline reads pre-dumped hidden states from disk ├── streaming streams hidden states from a live server in-loop └── dry-run ★ skip training entirely; convert + save + export ← NEW (this PR) ``` Companion `FakeBaseModel` fixes so small base checkpoints work: - Synthesize the weight_map from a single `model.safetensors` when no sharded index is present (Llama-3.2-1B, Qwen3-0.6B, …). - Honor `tie_word_embeddings`: reuse `embed_tokens` when `lm_head` is absent from safetensors. A new launcher YAML (`tools/launcher/examples/Qwen/Qwen3-8B/hf_eagle3_dryrun.yaml`) wires this together as a one-task pipeline. ### Usage ```bash # Direct python main.py --dry_run \ --config modelopt_recipes/general/speculative_decoding/eagle3.yaml \ model.model_name_or_path=meta-llama/Llama-3.1-8B-Instruct \ model.use_fake_base_for_offline=true \ data.offline_data_path=/tmp/dryrun-placeholder \ training.output_dir=ckpts/dryrun python scripts/export_hf_checkpoint.py --model_path ckpts/dryrun --export_path export/dryrun # Launcher uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_eagle3_dryrun.yaml --yes ``` ### Testing - `tests/unit/torch/speculative/plugins/test_fakebase.py`: 3 new cases — single-file fallback, tied-embeddings fallback, and the negative case (missing `lm_head` without tying). - `tests/examples/speculative_decoding/test_eagle.py::test_eagle3_dry_run`: full `launch_train.sh --dry_run → export_hf_checkpoint.py` chain on `tiny_llama`; asserts exported state_dict has all `LLAMA_EAGLE_SINGLE_LAYER` required keys. - Manually verified end-to-end on Llama-3.1-8B-Instruct (sharded), Llama-3.2-1B-Instruct (single-file + tied), and Qwen3-0.6B (single-file). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `--dry_run` is opt-in; `FakeBaseModel` changes are additive fallbacks. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — examples-only addition. - Did you get Claude approval on this PR?: ❌ ### Additional Information N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a `--dry_run` CLI flag to perform a fast early-exit execution that saves model artifacts without training. * Added support for single-file model checkpoint formats in the loader. * **Bug Fixes** * Improved checkpoint-loading error messages and handling. * Enhanced tied-embeddings fallback when head weights are absent. * **Tests** * Added integration and unit tests covering dry-run behavior and single-file/tied-embedding loading. * **Documentation** * Added a launcher example for dry-run smoke testing. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1566?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
d7e72f42ed |
Refine static NVFP4 MSE calibration (#1536)
### What does this PR do? Type of change: Bug fix Refines static NVFP4 MSE calibration and forces static NVFP4 amax state to stay FP32 across calibration loading, quantizer promotion, dtype casts, and restore paths. Main changes: - Tighten max/MSE calibration bootstrap and static NVFP4 quantizer promotion. - Keep static NVFP4 `_amax` and `_global_amax` in FP32. - Update focused GPU/unit coverage for FP8 sweep calibration, promotion, restore, and FP32 amax preservation. ### Usage ```yaml algorithm: method: mse fp8_scale_sweep: true ``` ### Testing ```bash pre-commit run --files modelopt/torch/quantization/nn/modules/tensor_quantizer.py tests/gpu/torch/quantization/test_nvfp4_static_quantizer_cuda.py pytest_pwd tests/gpu/torch/quantization/test_nvfp4_static_quantizer_cuda.py -q ``` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors). - Is this change backward compatible?: Yes - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: Yes - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: Yes ### Additional Information N/A Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
7aa0c95646 |
Add tests/gpu_vllm (#1517)
### What does this PR do? Type of change: new tests This PR adds unit tests for vLLM fakequant, specifically testing code in `modelopt/torch/quantization/plugins/vllm.py` ### Testing ``` pytest tests/gpu_vllm/torch/quantization/test_vllm_dynamic_modules.py -sv ``` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added comprehensive GPU vLLM test suite with end-to-end quantization checks and fixtures for TinyLlama, TinyQwen3-MoE, and DeepSeek V3; includes helpers to build tiny DeepSeek V3 models. * **Chores** * Updated GPU CI to use explicit container image references, added a GPU-focused test session, and adjusted test-run setup for vLLM. * **Documentation** * Documented new GPU test directory in contributing guide. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1517?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
eb5ed2df68 |
[CI] Bump torch, transformers and dev containers to latest (#1554)
- Transformers upper bound bumped from `<5.8` to `<5.10` - Enable torch 2.12 CICD testing - Bump TRT-LLM container to `1.3.0rc16` (transformers 5.5) - Use pytorch and tensorrt 26.04 containers in CICD <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated CI test container images and targeted Torch version across workflows; adjusted release CI job to use the newer torch config. * Broadened Transformers constraint in project metadata and test/dev pins. * Removed strict transformers pins from example requirements and lifted a compression dependency cap. * Raised the import-time Transformers version threshold for compatibility warnings. * **Tests** * Refactored a GPU test to collect and report validation errors and updated numeric expected baselines. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1554?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
4b270f0ea6 |
Support Mixed precision & Static MSE in MCore; Nemotron Super v3 NVFP4 recipe (#1521)
### What does this PR do?
Type of change: New features + Bug fixes
Mixed Precision and MSE support in MCore PTQ
- support mixed precision export in MCore by detecting mixed precision
layers in HF Quant Config
- Restore static quantizer in MCore checkpoint restore as `NVFP4QTensor`
(not TensorQuantizer which can call max calibrate. we want to skip max
calibrate for static quantizer during restore) --> fixes bug during
MCore export for MSE
- Fix dynamic block quantizer detection when `block_sizes` is
dict-backed.
- Add a YAML quantization recipe that roughly mirrors Nemotron 3 Super
NVFP4 `hf_quant_config.json`
Export bug fixes
- copy .py files properly from original HF ckpt (for reasoning parser
etc)
## Super recipe
Mirrors the published nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
hf_quant_config.json:
- MoE routed experts (mixer.experts.<N>.{up,down}_proj): NVFP4 W4A4
weight MSE, group_size 16
- MoE shared experts (mixer.shared_experts.{up,down}_proj): FP8
per-tensor
- Mamba mixer linears (mixer.{in,out}_proj): FP8 per-tensor
- KV cache: FP8
rest: not quantized
### Usage
```python
# Add a code snippet demonstrating how to use this
```
### Testing
Tested on Nemotron model
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added NVFP4 (4-bit) quantization checkpoint restore and export support
for Megatron-Core models
* Added tokenizer file export capability in model checkpoints
* Extended quantization support for expert-parallel distributed training
* Introduced new PTQ recipes for Nemotron-3-Super models with
mixed-precision quantization
* **Bug Fixes**
* Fixed FP8 and FP4 hardware compatibility detection on non-CUDA systems
* Improved offline Hugging Face Hub access handling with better error
messaging
* Enhanced calibration validation for mixture-of-experts models
* Fixed amplitude maximum validation for static block quantizers
* **Documentation**
* Updated expert weight quantization configuration documentation
<!-- review_stack_entry_start -->
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1521?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)
<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Jenny Chen <jennifchen@nvidia.com>
|
||
|
|
5eba879002 |
Launcher nvrx install from PyPI; specdec_bench guard modelopt import (#1567)
## Summary - **`tools/launcher/common/service_utils.sh`** — Megatron-LM (post PR #4522) asserts `nvrx >= 0.6.0` on import. The previous git-clone-HEAD install in `util_install_extra_dep` produced setuptools_scm versions like `0.0.0.dev1+hash` whenever upstream HEAD landed between tags, failing that assertion. Switch to a PyPI install pinned at `>= 0.6.0`. nemo containers ship nvrx in two Python envs; `pip` (system) and `python -m pip` (uv venv) target different ones, so uninstall from both and install into the venv where Python actually imports from. - **`examples/specdec_bench/specdec_bench/__init__.py`** — guard the `from modelopt import __version__` import so `specdec_bench` can be loaded inside a vllm container that doesn't have modelopt installed. ## Test plan - [x] Launcher YAMLs (e.g. Qwen3-8B PTQ on ComputeLab) install `nvidia-resiliency-ext>=0.6.0` and don't trip the Megatron-LM `has_nvrx_async_support` assertion. - [x] `specdec_bench` imports cleanly under a vllm container with no modelopt installed (prints the warning and uses `__version__ = "0.0.0"`). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Streamlined dependency installation: removed fragile repo-based install and now ensures a clean uninstall/upgrade flow for the external resiliency package to improve reliability across environments. * Improved version handling: added guarded fallback and warning when an optional package is missing, reporting a safe default version to avoid runtime errors. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1567?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
a9c156e24a |
Split bypass prerequisites (#1468)
## Summary This is PR 1 of 3 in the Puzzletron bypass/local-distillation stack. This PR contains prerequisite infrastructure only. It does not wire bypass distillation into the Puzzletron pipeline yet. Stack: 1. **This PR**: shared prerequisites 2. `ssameni/puzzletron-bypass-2-core`: bypass distillation core 3. `ssameni/puzzletron-bypass-3-integration`: Puzzletron integration, configs, docs, GPU coverage ## What Changed - Added `ModelDescriptor.pruning_mixins()` so model families can expose pruning mixins needed by downstream bypass initialization. - Added KV-head pruning mixin support for GPT-OSS, Nemotron-H, Nemotron-H-v2, and Qwen3-VL descriptors. - Improved pruning utilities for nested language-model configs and missing attention bias config fields. - Added `create_train_dataloader()` and streaming-safe shuffle handling. - Added chat-template fallback for base models without `tokenizer.chat_template`. - Added Sewing Kit loss/helper exports needed by the later bypass core. - Updated child-state initialization to support composing multiple pruning mixins. - Updated warmup-step resolver to account for gradient accumulation. ## Why The bypass distillation MR needs these reusable pieces, but they are independently reviewable and useful without adding the bypass training stage itself. Splitting them out keeps the bypass core PR focused on the actual local-distillation engine. ## Tests Added focused unit coverage for: - Dataloader behavior - Bypass loss helpers - KV-head pruning utilities - Sewing Kit activity/input/function/needle behavior <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * KV-head pruning added for multiple model families; generic pruning mixin hook available. * New training dataloader factory for infinite, block-sized training streams. * Vectorwise and batched normalized MSE loss utilities. * **Improvements** * Loss reports now show Δ-from-initial and visual indicators. * Chat-sample preprocessing tolerates tokenizers without chat templates. * More robust head-dimension/bias handling and grad-accum-aware warmup resolver. * **Tests** * Extensive unit tests added across dataloaders, losses, pruning, hydra utils, and sewing-kit. [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1468) <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Sepehr Sameni <ssameni@nvidia.com> |
||
|
|
a2c496af0d |
[OMNIML-4788] specdec_bench: configuration.json provenance + upload_to_s3 (#1531)
> [!WARNING] > **Breaking on-disk schema change (specdec_bench v1.0.0).** This PR renames the acceptance-rate metric fields across `AcceptanceRate` / `MTBench` / `SpecBench` writers: > > | Old (pre-1.0.0) | New (1.0.0) | > |---|---| > | `Request_AR` | `Request_AL` | > | `Category_AR` | `Category_AL` | > | `Average_AR` | `Average_AL` | > | — | `Joint_Acceptance_Rate` (new) | > > The renamed values were always **acceptance length** (mean tokens generated per inference step), not a rate, and the visualizer reads `*_AL`. Pre-1.0.0 runs in S3 have `*_AR` and no `Joint_AR`; they must be re-run or post-processed before comparing. The visualizer aggregates runs by `specdec_bench` major version so accidental cross-methodology comparison is blocked. ### What does this PR do? Type of change: new feature Adds reproducibility provenance to `specdec_bench/configuration.json` and ports `upload_to_s3.py` from `iputterman/specdec_bench@main` (personal-namespace fork) into upstream. This is the first PR in a multi-stage migration off Izzy's fork now that he's left the team. Tracked in [OMNIML-4788](https://jirasw.nvidia.com/browse/OMNIML-4788). **Provenance fields added to configuration.json** (alongside existing argv / engine_version / gpu / python_version): - `specdec_bench_version` — methodology semver declared in `specdec_bench/__init__.py`. Bump minor on additive metrics, major on changed metric *definitions*. The visualizer (Phase 4 of the migration) will aggregate runs by major version so plots don't accidentally compare across methodology changes. - `specdec_bench_sha`, `modelopt_sha`, `modelopt_version`, `nmm_sandbox_sha`, `container_image` — code/runtime provenance. Each prefers an env var set by the harness (`SPECDEC_BENCH_SHA`, `MODELOPT_SHA`, `MODELOPT_VERSION`, `NMM_SANDBOX_SHA`, `CONTAINER_IMAGE`) and falls back to `git rev-parse` / `modelopt.__version__` when running standalone. The env-var preference is necessary because the runtime container has no `.git/` (the launcher packager tarballs source without git metadata) and may not have `modelopt` installed. - `checkpoint.{path, size_bytes, index_sha256, index_source}` — cheap reproducibility fingerprint that hashes `model.safetensors.index.json` (or `config.json` fallback). Changes whenever any tensor changes. - `serving_config` — engine-level config dict captured after init via a new `Model.get_serving_config()` method. VLLM dumps `AsyncEngineArgs` + the live `vllm_config.to_dict()`; SGLANG dumps the `engine_kwargs` passed to `sgl.Engine`; TRTLLM left at the base default `{}` for a later iteration. - `timestamp` — UTC ISO 8601. **Other changes** - `upload_to_s3.py` + `specdec_bench/s3_utils.py` ported from iputterman/specdec_bench@main. Recognizes run dirs by sentinel files, refuses to overwrite existing S3 prefixes. - `_redact_config` allowlists `tokenizer`, `tokenizer_path`, `tokenizer_mode`, `tokenizer_revision` so the model path stops being redacted (latent bug from substring-matching `token` ⊂ `tokenizer`). - `requirements_speed.txt`: `boto3`, `botocore` added (used by `s3_utils`). **Out of scope** (deferred to Phase 1b / Phase 2): - `--sweep_config` driver that emits per-run-dir nesting `<sweep>/<NNN_dataset_c<conc>>/` - `--s3_upload` flag baked into `run.py` itself - Launcher auto-injection of the provenance env vars (currently the example YAML sets them statically) - `container_digest` (enroot integration) and full GPU/driver inventory - TRTLLM `get_serving_config()` ### Usage ```bash # Run a smoke benchmark (Qwen3.5-4B + vLLM + MTP draft=3) — example YAML included uv run launch.py --yaml examples/Qwen/Qwen3.5-4B/specdec_bench_mtp.yaml --yes # After it lands, upload the run directory to S3: S3_KEY_ID=team-specdec-workgroup \ S3_SECRET=... \ python upload_to_s3.py /path/to/sweep_dir s3://team-specdec-workgroup/results ``` ### Testing Cluster-tested end-to-end on cw-dfw (Slurm job 11978794, NeMo Run experiment `cicd_1779403623`, ~19 min wall): - Qwen3.5-4B + vLLM + MTP draft=3 + SPEED-Bench-Internal/qualitative (80 requests) - `configuration.json` (22 KB) populated all eight new provenance fields - `Request_AR` mean 3.327 (vs 3.330 on the pre-Phase-1a run — within noise; methodology unchanged) - `upload_to_s3.py` (real upload, not dry-run) landed [s3://team-specdec-workgroup/results/qwen35_4_mtp_smoke_2026-05-21/specdec_bench_mtp/](https://app.s8k.io/buckets/team-specdec-workgroup/?prefix=results%2Fqwen35_4_mtp_smoke_2026-05-21%2F) where the visualizer at http://10.131.132.205:8080 can pick it up. ### Before your PR is "Ready for review" - Is this change backward compatible?: ✅ - `configuration.json` only gains fields. `upload_to_s3.py` / `s3_utils.py` are new files. `Model.get_serving_config()` default = `{}` so existing subclasses without an override behave as before. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - `boto3` / `botocore` are Apache 2.0 (permissive); `upload_to_s3.py` + `s3_utils.py` are ported from a private NVIDIA repo with explicit copyright headers retained. - Did you write any new necessary tests?: ❌ - Validated by cluster smoke (see Testing). Will add unit-tests for `dump_env` provenance fields and `upload_to_s3._discover_runs` in a follow-up. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Internal-facing tooling. - Did you get Claude approval on this PR?: ❌ (triggering after open) ### Additional Information Tracked in JIRA [OMNIML-4788](https://jirasw.nvidia.com/browse/OMNIML-4788). The full multi-phase plan is on that ticket's SPEC block — this PR is Phase 1a. Cherry-picked alongside the harness change are two example YAMLs (`examples/Qwen/Qwen3.5-4B/specdec_bench.yaml` for the NONE autoregressive baseline, `..._mtp.yaml` for the MTP run) that gave us cluster-test evidence. Can be split out if preferred. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an S3 upload CLI for benchmark results with dry-run and skip-existing options * Automatic capture of run configuration, provenance and redacted environment into saved config * Models now export serving configuration for reproducible runs * New launcher entrypoint and example job configs for Qwen SPEED-Bench runs * **Documentation** * README section describing S3 upload usage and supported local layouts * **Bug Fixes / Changes** * Acceptance-rate metric keys renamed in output (AR -> AL) * **Tests / CI** * New tests for redaction and S3 utilities; CI now runs specdec_bench examples <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1531?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: chenhany <chenhany@nvidia.com> |
||
|
|
999c99913e |
Add Nemotron-3-Nano-30B-A3B-BF16 e2e tutorial: Prune + Distill + Quantize + Nemo Evaluator + vLLM deployment (#1376)
### What does this PR do? Type of change: example/tutorial <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Add Nemotron-3-Nano-30B-A3B-BF16 e2e tutorial: Prune + Distill + Quantize + Nemo Evaluator + vLLM deployment <img width="2079" height="1613" alt="image" src="https://github.com/user-attachments/assets/19b6ab82-7f01-45df-a0a5-d1c3282b384a" /> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * End-to-end Nemotron-3-Nano-30B tutorial: pruning, two‑phase distillation, FP8 PTQ, evaluation, and vLLM deployment; new ablations and long‑context analyses. * Distillation CLI: configurable seed and activation‑recomputation options. * **Bug Fixes** * Preprocessing hardened to skip malformed JSONL and normalize tool‑call argument formats. * **Documentation** * Many README/examples/evaluator docs updated (news list, tokenization guides, tutorials, configs, and deployment notes). * **Tests** * Added test verifying preprocessing handles stringified tool‑call arguments. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1376?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
5d0441ae3d |
Create shared Megatron calibration forward loop for prune / quantize with megatron pretraining data style sequence packing (#1501)
## Summary
Replaces the bespoke calibration loops in Megatron-LM and
Megatron-Bridge prune / quantize example scripts with a single shared
utility,
`modelopt.torch.utils.plugins.megatron_calibration.get_megatron_calibration_forward_loop`.
The shared loop iterates a packed calibration dataloader built via
`get_dataset_dataloader(pack=True)` and drives a logits-free prefill
pass through the model so activation hooks fire on every layer.
`pack=True` produces Megatron-LM pretraining-style **global-stream
packing**: all raw samples are concatenated into one EOS-separated token
stream and sliced into uniform-length rows. The trained model has seen
this distribution extensively during pretraining, so the activations
produced during calibration are representative of the model's natural
behavior.
Migrates four call sites:
- `examples/megatron_bridge/prune_minitron.py`
- `Megatron-LM/examples/post_training/modelopt/{prune,quantize}.py`
(separate PR:
[NVIDIA/Megatron-LM#4881](https://github.com/NVIDIA/Megatron-LM/pull/4881))
- `Megatron-Bridge/examples/quantization/quantize.py` (separate PR)
Each call site passes `pack=True` explicitly with an inline comment so
users see the option and know when to flip it. The function-level
defaults (`get_dataset_dataloader(pack=False)`,
`get_megatron_calibration_forward_loop(pack=False)`) remain
back-compat-safe.
Unified defaults across all four sites: `--calib-dataset
nemotron-post-training-dataset-v2`, `--calib-size 1024`,
`--calib-max-sequence-length 4096`, `--calib-batch-size 1`.
## Experimental results
Qwen3-8B on full 100% MMLU (n=14042; binomial 2σ noise floor ≈ ±0.78 pt
at acc ≈ 0.7), 0-shot, eval batch_size=4. Calibration on the default
workload: nemotron-post-training-dataset-v2, seq_length=4096,
calib_batch_size=8.
**Three calibration data shapes compared:**
- **Padded**: one doc per row, padded to `seq_length`, pad tokens flow
through the forward (legacy `get_calib_dataloader` pad+truncate
behavior).
- **Trimmed**: one doc per row, each row trimmed to its real content
length via `attention_mask`, with EOS forced at the last real position;
pad never enters the forward. This is no longer part of this PR.
- **Packed**: global-stream slicing — all docs concatenated
EOS-separated into one token stream, sliced into uniform `seq_length`
rows. Matches Megatron's `.bin`/`.idx` pretraining distribution. Enabled
via `pack=True` in `get_megatron_calibration_forward_loop`.
| Workload | Padded | Trimmed | **Packed** |
|---|---|---|---|
| M-LM NVFP4 quantize (`NVFP4_DEFAULT_CFG`) | 0.707 | 0.708 | 0.709 |
| M-Bridge Minitron prune (Qwen3-8B → 30L / 3584 / 11776 ≈ 6B params) |
0.576 | 0.573 | **0.589** |
### Key findings
- **M-LM quantize quality is calibration-mode-insensitive** for dense
Qwen3-8B NVFP4 — all three modes are nearly identical.
- **M-Bridge prune**: More sensitive to calibration data shape. Packed
wins over Padded on full MMLU.
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
48fbac0ed7 |
config for expert removal in nemotron3 (#1544)
### What does this PR do? New example This PR adds config and updates Nemotron descriptor to support expert pruning for Nano3 30B-A3B ### Usage ```python torchrun --nproc_per_node 2 examples/puzzletron/main.py --config examples/puzzletron/configs/nemotron-nano-30b-A3b-v3/nemotron_nano_v3_pruneexp.yaml 2>&1 | tee ./log.txt | grep "Puzzletron Progress" ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - Did you write any new necessary tests?: ❌ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for Nemotron Nano 30B model with multiple pruning and compression strategies * Enabled expert pruning to reduce model capacity, plus attention optimization and feed-forward layer compression * Introduced model validation and solution evaluation configurations for efficient model optimization workflows <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1544?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: mchochowski <mchochowski@nvidia.com> |
||
|
|
9d0d97829a |
chore(lint): modernize typing (PEP 604/585) and enable UP032 (#1537)
### What does this PR do? Type of change: chore / refactor (no behavior change) Two small lint-cleanup commits: **1. `chore(typing): modernize Union/Optional/List to PEP 604 / 585 syntax`** (8 files) - Replace `X = Union[A, B] # noqa: UP007` with `X: TypeAlias = A | B` for the six module-level type aliases (`ModelLike`, `Criterion`, `NodeTarget`, `CalibrationDataType`, `Hparam.Importance` / `ActiveSlice`). The `TypeAlias` annotation is required so mypy continues to treat them as aliases under PEP 604. - Modernize forward-ref unions in `modelopt/onnx/quantization/autotune/` to full-string forward refs (e.g. `"RegionPattern | None"`). - Update docstring type tags in `examples/puzzletron/evaluation/hf_deployable_anymodel.py`. **2. `chore(lint): remove UP032 ignore and convert .format() to f-strings`** (10 files) - Drop `UP032` from `extend-ignore` in `pyproject.toml`. - Auto-convert 19 `"...".format(...)` calls to f-strings across export plugins, examples, tests, and tools. One conversion in `modelopt/torch/utils/plugins/megatron_generate.py` was wrapped manually to stay under the 100-char limit. **Intentionally left as-is:** - `tools/launcher/slurm_config.py` keeps its `# ruff: noqa: UP045` — nemo_run's CLI parser can't introspect PEP 604 optional annotations. - `modelopt/torch/puzzletron/*` is **not** touched. The subtree disables ruff's `UP` family entirely (per-file-ignore `"UP"`) while migration is in progress, and converting `Optional[X]` to `X | None` there would silently break runtime introspection in `block_config._get_dataclass_type` that uses `get_origin(tp) is typing.Union` (PEP 604 unions return `types.UnionType` from `get_origin`, not `typing.Union`). Best revisited when puzzletron's lint carve-out is narrowed. - `UP038` (`isinstance(x, (int, float))` → `isinstance(x, int | float)`) — ruff has officially deprecated this rule; PEP 604 in isinstance is slightly slower and misleads readers about PEP 695 / `Optional`. Ignore kept. ### Usage No user-facing API changes. ### Testing - Pre-commit hooks (ruff check, ruff format, mypy, bandit, license) pass on both commits. - Ruff status against `main`: 37 unrelated pre-existing findings (W291/W293/E501/RUF005/PLR1704); zero new findings introduced by this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — Runtime behavior of the six type aliases changes from a `typing.Union` instance to `types.UnionType`. Downstream code introspecting via `get_origin(...) is typing.Union` on these aliases would break, but no in-repo caller does this on them. (The introspection in `modelopt/torch/puzzletron/block_config.py` operates on user-supplied dataclass field types, none of which are these aliases.) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (no behavior change) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ — internal style refactor; happy to add a Misc note if reviewers want one. - Did you get Claude approval on this PR?: ❌ — not yet. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Style** * Modernized type annotations across the codebase to use Python 3.10+ union syntax and TypeAlias where appropriate. * Standardized string formatting to f-strings, improving clarity of logs, errors, and validation messages. * **Chores** * Updated linting configuration to reflect modern typing/style rules. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1537?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
f59f3ae1c3 |
[Example]: Dflash-Offline Launcher Example (#1529)
### What does this PR do? Type of change: new example Offline DFlash training launcher example for Qwen3-0.6B. Two-task pipeline: dump base-model hidden states via HF forward, then train DFlash on the dump. - `tools/launcher/examples/Qwen/Qwen3-0.6B/hf_offline_dflash.yaml` — new launcher YAML - `tools/launcher/common/eagle3/dump_offline_data_hf.sh` — new HF-backed dump script (DFlash's `answer_only_loss=true` requires `loss_mask`, which the existing TRT-LLM dump backend does not produce) - `examples/dataset/synthetic_conversations_1k.jsonl` — `conversation_id` added at-source (asserted by `compute_hidden_states_*.py`; previously injected at test-fixture time) - `tests/regression/torch/speculative/test_dflash_offline.py` — `tagged_synth_data_path` fixture removed since the field is now at-source ### Usage ``` uv run launch.py --yaml examples/Qwen/Qwen3-0.6B/hf_offline_dflash.yaml --yes ``` ### Testing End-to-end on CoreWeave Slurm (1×1 GPU): task_0 dump succeeded; task_1 training loss 7.85 → 2.94 over 2 epochs, regression PASSED. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (`conversation_id` is purely additive on the dataset) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (existing `test_dflash_offline.py` covers the workflow) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (new example only) - Did you get Claude approval on this PR?: ❌ ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added launcher script for offline hidden-state processing with HuggingFace backend support and distributed execution capabilities * Added training configuration example for Qwen3-0.6B model with DFlash speculative decoding pipeline * **Tests** * Updated offline hidden-states regression test to improve data handling workflow <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1529?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
16a0130d5f |
fix: preserve inlined MTP layers for GLM5 (#1532)
### What does this PR do?
Type of change: Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
Extends `load_mtp_weights` to detect *inlined* MTP layers — keys
`model.layers.{i}.*` for `i in [num_hidden, num_hidden +
num_nextn_predict_layers)` — in addition to the existing `mtp.*`
separate-file convention.
**Bug.** `load_mtp_weights()` only matched the substring `"mtp"` in
safetensors keys. GLM-5.1 (`GlmMoeDsaForCausalLM`) stores MTP at
`model.layers.78.*` with no `mtp` substring, so detection returned `([],
{})`, `_mtp_layer_prefixes` was never set, and MTP tensors were silently
dropped from the exported safetensors (had to be re-added manually).
**Detection.**
1. **Detect** via `config.num_nextn_predict_layers` (the model's own
declaration of how many MTP layers exist).
2. **Compute** the inlined layer indices: `model.layers.{i}` for `i in
range(num_hidden, num_hidden + num_nextn)`.
3. **Load** matching tensors from the on-disk shards via `safe_open`
(walks `model.safetensors.index.json` if present,else falls back to the
single shard).
4. **Split** the loaded tensors by whether `model.state_dict()` has a
slot for them:
- keys present in `model.state_dict()` → `model.load_state_dict(...,
strict=False)` (DeepSeek-V3 case: HF instantiates the extra layers).
- keys absent from `model.state_dict()` → returned as
`not_in_state_dict` so the exporter routes them through
`extra_state_dict` (GLM-5.1
case: `GlmMoeDsaModel` in transformers ≥5.7 only builds `num_hidden`
decoders, leaving MTP keys orphaned at `from_pretrained` time).
The returned prefixes flow into the existing plumbing —
`_mtp_layer_prefixes` → `quant_cfg` disable
+`quantization_config.exclude_modules` — unchanged.
### Usage
```python
# Add a code snippet demonstrating how to use this
```
### Testing
Verified end-to-end on a mini GLM-5.1 fixture (4 hidden layers + 1
inlined MTP at `model.layers.4`, 7 synthesized MTP tensors mirroring the
full GLM-5.1 layout)
To be verified with full model
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Improvements**
* Quantization utilities now stream safetensors and unify loading of
multi-token-prediction (MTP) weights from both inline and
separate/sharded conventions, reporting detected MTP prefixes and counts
of loaded vs orphaned tensors.
* **Tests**
* Added unit tests and a test import helper covering MTP discovery,
loading behaviors (inlined vs standalone/indexed shards), orphan
reporting, and non‑MTP checkpoints.
<!-- review_stack_entry_start -->
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1532?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)
<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com>
|
||
|
|
04f58166ab |
[OMNIML-3707] Model-specific PTQ recipes bootstrap (#1506)
### What does this PR do?
Type of change: new feature
Replaces the hardcoded model-type branches in `examples/llm_ptq/` with
opt-in declarative **model-specific recipes** under
`modelopt_recipes/huggingface/<model_type>/ptq/`. Any adjustment
specific to a model type or instance must live in that model's recipe —
there is no implicit model-specific path anymore. Users select a model's
recipe with `--recipe huggingface/<model_type>/ptq/<recipe>`; users on
the plain `--qformat` path get only the generic numerics.
What moved out of Python
(`examples/llm_ptq/example_utils.py::build_quant_cfg` and
`examples/llm_ptq/hf_ptq.py::mono_quantize`):
- **gemma / mpt** `w4a8_awq` → `awq_lite` with `alpha_step=1` (coarser
search to avoid TRT-LLM overflow).
- **gemma** `int8_sq` → SmoothQuant `alpha=0.5` (default `1.0` regresses
Gemma 7B).
- **phi4mm** → disable `*speech*`, `*audio*`, `*image*`, `*vision*`
(quantize only the language model).
- **Nemotron VL** → disable `*vision*`, `*image*`, `*radio*`,
`*visual*`, `*encoder*`, `*model_encoder*` (quantize only the decoder).
What stayed in Python:
- MTP dynamic layer exclusion in `hf_ptq.py` (depends on
runtime-detected layer indices).
- `is_nemotron_vl(full_model)` detection itself, which still drives the
VLM calibration loop and the post-quantize `full_model` update — only
the `quant_cfg` adjustment it triggered moved into the Nemotron VL
recipe.
`multinode_ptq.py` shares the same `build_quant_cfg` call site and was
updated to match the new 2/3-arg signature; multinode users on
`--qformat` get the generic numerics (no `--recipe` plumbing in
multinode yet, so model-specific recipes are only reachable via
`hf_ptq.py`).
Already-YAML recipes that were elsewhere in the tree are relocated into
the same `huggingface/<model_type>/ptq/` layout so all model-specific
recipes live under one convention:
- **Step3.5-Flash** — moved from
`modelopt_recipes/huggingface/step3p5/Step3.5-Flash/` to
`huggingface/step3p5/Step3.5-Flash/ptq/` to match the `<model>/ptq/`
convention.
- **Qwen3.5 / Qwen3.6** — moved from
`modelopt_recipes/models/Qwen3.5-Qwen3.6/w4a16.yaml` to per-model_type
folders, anchored on the HuggingFace `model_type` (verified against
transformers 5.8.1 + HF model hub `config.json` for `Qwen/Qwen3.6-27B`,
`Qwen/Qwen3.6-35B-A3B`, `nvidia/Qwen3.5-397B-A17B-NVFP4`):
- `huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml` —
dense `qwen3_5`
- `huggingface/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml` —
`qwen3_5_moe`
- Both wrappers `$import` the shared `quant_cfg` snippet
`huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg.yaml`
(one source of truth; the two model_types share the same hybrid
linear-attention + softmax-attention architecture so the rules apply
identically).
Full recipe layout (`modelopt_recipes/huggingface/`):
```
gemma/ptq/{w4a8_awq,int8_sq}-kv_fp8_cast.yaml
mpt/ptq/w4a8_awq-kv_fp8_cast.yaml
phi4mm/ptq/{disabled_quantizers,nvfp4-kv_fp8_cast}.yaml
nemotron_vl/ptq/{disabled_quantizers,nvfp4-kv_fp8_cast}.yaml
qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast{,.quant_cfg}.yaml
qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml
step3p5/Step3.5-Flash/ptq/nvfp4-mlp-only.yaml
```
All recipes ship with FP8 KV-cache cast (`kv_fp8_cast`). For phi4mm and
nemotron_vl, `disabled_quantizers.yaml` is a multi-document list unit
that `$import`s the standard `default_disabled_quantizers` exclusions
and appends the model-specific ones — so each recipe imports a single
disabled-quantizer slot instead of layering two, with no duplication in
YAML. Each `ptq/` folder has a `README.md` describing exactly what is
model-specific.
### Usage
```bash
# Gemma W4A8 AWQ with the Gemma-specific algorithm tuning + FP8 KV cache:
python examples/llm_ptq/hf_ptq.py \
--pyt_ckpt_path google/gemma-7b \
--recipe huggingface/gemma/ptq/w4a8_awq-kv_fp8_cast \
--export_path ./out
# Nemotron VL with vision branches excluded automatically:
python examples/llm_ptq/hf_ptq.py \
--pyt_ckpt_path nvidia/<nemotron-vl-model> \
--recipe huggingface/nemotron_vl/ptq/nvfp4-kv_fp8_cast \
--export_path ./out
```
### Testing
- Pre-commit recipe validator
(`tools/precommit/check_modelopt_recipes.py`) loads every new recipe via
`load_recipe()` — passes for all new YAMLs (gemma/mpt/phi4mm/nemotron_vl
recipes + phi4mm/nemotron_vl `disabled_quantizers` snippets + qwen3_5 /
qwen3_5_moe recipe wrappers + the shared
`w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg` snippet + Step3.5-Flash
relocation).
- For qwen3_5 / qwen3_5_moe specifically, `load_recipe(...)` on both
wrappers produces an identical 33-entry resolved `quant_cfg`, confirming
the shared snippet is the single source of truth.
- `yamlfmt` + `markdownlint` + `bandit` + license-insertion hooks all
pass.
- No tests reference the removed `build_quant_cfg(qformat, ...,
model_type, ...)` signature; the only call sites (`hf_ptq.py`,
`multinode_ptq.py`) were updated to the new 2/3-arg form.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ — users who relied on
**automatic** model-specific quant_cfg behavior via `--qformat`
(gemma/mpt AWQ, gemma SmoothQuant, phi4mm exclusions, Nemotron VL
exclusions) now need to pass `--recipe
huggingface/<model_type>/ptq/<recipe>` to apply the model's recipe. The
flag itself is unchanged; only the implicit behavior was removed.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ — relies on the existing
pre-commit recipe validator that loads each new YAML.
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ❌
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added many model-specific PTQ recipes (Gemma, MPT, Nemotron VL,
Phi‑4‑Multimodal, Qwen3.5, Qwen3.5‑MoE) and support for AWQ block-size
and MoE calibration ratio in quantization options.
* **Documentation**
* Expanded READMEs and changelog to document recipe locations, layout,
and how to opt into model-specific PTQ recipes.
* **Refactor**
* Model-specific PTQ tweaks moved to opt‑in recipes; default behavior
uses generic numerics.
<!-- review_stack_entry_start -->
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1506?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)
<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
3ff15ccef3 |
Add support for postprocess exported model for block scale swizzling and support for different padding strategy (#1195)
### What does this PR do? Type of change: ? new feature <!-- Details about the change. --> Adds post-processing support for exported diffusion model checkpoints to enable NVFP4 block scale swizzling and configurable padding strategies. This allows exported quantized checkpoints to be directly consumed by inference runtimes (e.g., ComfyUI with comfy_kitchen) that require cuBLAS 2-D block-scaling-factors layout. Changes: 1) Unified post-processing step (_postprocess_safetensors): Loads saved safetensors files and applies merge, padding, swizzle, and quantization metadata injection in a single pass. 2) NVFP4 scale swizzle (swizzle_nvfp4_scales): Rearranges block scales from ModelOpt's flat [rows, cols // 16] layout to cuBLAS 2-D tiled layout per the cuBLAS specification. 3) Configurable padding (pad_nvfp4_weights): Pads NVFP4 weight and scale tensors to multiples of 16, with "row" (rows only) or "row_col" (both dimensions) strategies. 4) Standalone quantization metadata (build_layerwise_quant_metadata): Extracted from merge_diffusion_checkpoint so _quantization_metadata can be injected independently of merging — works for both merged (LTX-2) and standalone (Flux2) exports. 5) Bug fix (conversion.py): Wrapped yield in try/finally in set_quantizer_by_cfg_context so quantizer states are always restored, fixing an issue when yield fails. ### Usage ```python # LTX-2 export with merge + swizzle + padding export_hf_checkpoint( pipeline, export_dir="./output", merged_base_safetensor_path="./ltx-2-22b-dev.safetensors", enable_swizzle_layout=True, padding_strategy="row_col", enable_layerwise_quant_metadata=True, ) # Flux2 standalone export with swizzle + padding (no merge needed) export_hf_checkpoint( transformer, export_dir="./output", enable_swizzle_layout=True, padding_strategy="row_col", ) # Via quantize.py CLI python quantize.py \ --model ltx-2 --format fp4 \ --extra-param merged_base_safetensor_path=./ltx-2-22b-dev.safetensors \ --extra-param enable_swizzle_layout=true \ --extra-param padding_strategy=row_col \ --hf-ckpt-dir ./output ``` ### Testing 1) Exported LTX-2.3 NVFP4 with swizzle + padding + merged base checkpoint. Verified checkpoint has correct uint8 weights, float8_e4m3fn scales in swizzled layout, and _quantization_metadata . Ran the checkpoint with ComfyUI 2) Exported Flux2 NVFP4 with swizzle + padding. Verified checkpoint has correct uint8 weights, float8_e4m3fn scales in swizzled layout, and _quantization_metadata . Ran the checkpoint with ComfyUI ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Diffusers export: optional NVFP4 support — swizzle layout, row/row_col padding, and optional per-layer quantization metadata; exports are now post-processed to apply these options. * Export flow accepts new flags to enable swizzle, padding strategy, and layerwise metadata. * **Bug Fixes** * Quantizer context manager now always restores state, including on exceptions. * **Tests** * Added unit tests for NVFP4 padding, swizzling, metadata injection, and post-processing. * **Documentation** * README example updated to show swizzle and padding flags. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ynankani <ynankani@nvidia.com> Signed-off-by: YASH Nankani <ynankani@2u1g-x570-0073.ipp2a1.colossus.nvidia.com> Signed-off-by: ynankani-nv <ynankani@nvidia.com> Signed-off-by: YASH Nankani <ynankani@dl325g11-1979.ipp2a2.colossus.nvidia.com> Signed-off-by: YASH Nankani <ynankani@dl325g11-0771.ipp4a1.colossus.nvidia.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> Co-authored-by: YASH Nankani <ynankani@2u1g-x570-0073.ipp2a1.colossus.nvidia.com> Co-authored-by: YASH Nankani <ynankani@dl325g11-1979.ipp2a2.colossus.nvidia.com> Co-authored-by: YASH Nankani <ynankani@dl325g11-0771.ipp4a1.colossus.nvidia.com> |
||
|
|
c9098b63fb |
[4/n] Add vLLM integration for modelopt sparse attention (#1127)
### What does this PR do?
Type of change: New feature, new example, new tests, documentation.
Adds vLLM integration for ModelOpt sparse attention with paged KV cache
support.
This PR extends the ModelOpt Triton flash attention path so K/V can be
read directly from vLLM's paged KV cache through `block_table` lookup.
This avoids gather-to-contiguous copies when serving exported
sparse-attention checkpoints with vLLM.
The vLLM integration swaps vLLM's `FlashAttentionImpl` with
`ModelOptSparseAttentionImpl` after model load. The sparse configuration
is read from the exported checkpoint's `config.json`
`sparse_attention_config` block, written by
`examples/llm_sparsity/attention_sparsity/hf_sa.py`.
The restored checkpoint metadata supports:
- calibrated skip-softmax metadata (`threshold_scale_factor`,
`target_sparse_ratio`)
- N:M sparse-softmax metadata (`sparsity_n`, `sparsity_m`)
- dense token preservation metadata (`dense_sink_tokens`,
`dense_recent_tokens`)
The vLLM path uses ModelOpt Triton for sparse prefill launches.
Decode-only launches, cascade/prefix-cache metadata, and launches
without active sparse work delegate back to vLLM FlashAttention.
### Limitations
- Sparse attention is enabled for sparse prefill only.
- Decode-only launches currently fall back to vLLM FlashAttention.
- Attention sinks from vLLM FlashAttention are rejected until the
ModelOpt Triton path supports them.
- CUDA graph capture is not validated with this sparse attention path
yet; use `--enforce-eager`.
- Quant-only serving remains covered by `vllm_serve_fakequant.py`.
- Combined sparse attention + quantization serving is not handled by
this launcher in this PR and is planned as follow-up work.
### Usage
Export a checkpoint with calibrated skip-softmax and sparse24 metadata:
```bash
python examples/llm_sparsity/attention_sparsity/hf_sa.py \
--pyt_ckpt_path /path/to/hf-model \
--sparse_attn skip_softmax_calib_sparse24 \
--target_sparse_ratio 0.5 \
--calib_samples 64 \
--calib_max_seqlen 16384 \
--calib_chunk_size 4096 \
--seq_len 2048 \
--export_dir /path/to/modelopt-skipsoftmax-sparse24-export
```
Serve the exported checkpoint with the vLLM sparse-attention launcher:
```bash
PYTHONPATH=$PWD python examples/vllm_serve/vllm_serve_sparse_attn.py \
/path/to/modelopt-skipsoftmax-sparse24-export \
--tensor-parallel-size 8 \
--host 0.0.0.0 \
--port 8000 \
--trust-remote-code \
--enforce-eager
```
Send a request through the OpenAI-compatible endpoint:
```bash
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "/path/to/modelopt-skipsoftmax-sparse24-export",
"messages": [{"role": "user", "content": "Explain sparse attention in one paragraph."}],
"max_tokens": 128
}'
```
### Testing
GitHub CI on the latest commit is green:
- DCO
- code-quality
- docs build / deploy preview
- unit tests, including Linux, Windows, multi-version, partial-install,
and launcher jobs
- example tests
- GPU tests, including required GPU gate
- regression tests, including required regression gate
- `codecov/project`
Focused test coverage added/updated for this PR includes:
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_worker.py`
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_config.py`
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attention_conversion.py`
-
`tests/unit/torch/sparsity/attention_sparsity/test_triton_skip_softmax.py`
- `tests/gpu/torch/sparsity/attention_sparsity/test_vllm_plugin.py`
- `tests/gpu/torch/kernels/common/attention/test_triton_fa_paged.py`
-
`tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_skip_softmax.py`
-
`tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_sparse_nm.py`
-
`tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_calibrate.py`
Manual / NEL eval validation:
- Served a ModelOpt exported sparse-attention checkpoint through
`examples/vllm_serve/vllm_serve_sparse_attn.py`.
- Launched RULER64K NEL evals on DFW with `coreai_nvfm_llm`.
- Current partial RULER64K prediction scores, before final `results.yml`
is written:
- `skipsoftmax-only`: 98.59% over 4500 flushed samples
- `skipsoftmax-r0.7`: 99.70% over 1000 flushed samples
- `skipsoftmax-r0.9`: 99.70% over 1000 flushed samples
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - no new
PIP dependency.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A - documentation, examples, and tests are updated for this vLLM
integration path.
### Additional Information
Follow-up work:
- Validate and enable CUDA graph capture for the sparse vLLM path.
- Add combined sparse attention + quantization serving once the combined
path is tested.
- Investigate whether skip-softmax should also be enabled during decode.
---------
Signed-off-by: Kai Xu <kaix@nvidia.com>
|
||
|
|
910dc49a2c |
Add Qwen3.6 W4A16 PTQ recipe (#1503)
### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a W4A16 PTQ recipe for Qwen3.5/Qwen3.6 with mixed-precision rules (NVFP4 for MLP projections, FP8 for attention layers and KV cache). * **Updates** * PTQ workflow now respects supplied recipes and applies quantization rules to the full model (recipe-driven targeting enabled). <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1503?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
e4dc0205d1 |
[OMNIML-4775] Move built-in PTQ quantization configs to YAML (#1423)
### What does this PR do? Type of change: refactor This PR moves the built-in PTQ quantization config definitions out of hard-coded Python dictionaries and into schema-backed YAML config files, and factors shared blocks into reusable composable snippets. - Adds reusable numeric config snippets under `modelopt_recipes/configs/numerics/`. - Adds YAML presets for the built-in model PTQ configs under `modelopt_recipes/configs/ptq/presets/model/`. - Adds YAML presets for KV-cache quantization configs under `modelopt_recipes/configs/ptq/presets/kv/`. - Adds YAML presets for the Diffusers-specific PTQ configs under `modelopt_recipes/configs/ptq/presets/diffusers/` and re-points `examples/diffusers/quantization/config.py` constants at them via `load_config`. - Adds reusable KV quantization units (`kv_fp8_affine`, `kv_nvfp4`, `kv_nvfp4_affine`, `kv_nvfp4_rotate`, `kv_*_cast` variants) under `modelopt_recipes/configs/ptq/units/`. - Adds reusable model-side units following the `component_numerics[_type]` convention: - `attention_qkv_fp8` — FP8 E4M3 on attention q/k/v bmm and softmax quantizers; shared by `model/` and `diffusers/` `nvfp4_fp8_mha` presets. - `block_sparse_moe_nvfp4` — NVFP4 W4A4 on `*block_sparse_moe*` weight/input quantizers; shared by `nvfp4_mlp_only`, `nvfp4_experts_only`, `nvfp4_omlp_only`. - `experts_nvfp4` — NVFP4 W4A4 on `*.experts.*` weight/input quantizers; shared by `nvfp4_mlp_only` and `nvfp4_experts_only`. - Switches the existing 5 NVFP4 presets (default + awq lite/clip/full + svdquant) and 4 mamba_moe presets to `$import` the existing `w4a4_nvfp4_nvfp4` / `w8a8_fp8_fp8` units instead of re-inlining the same weight+input quantizer pairs. - Moves the recently-added `W4A16_NVFP4_CFG` to YAML (`presets/model/w4a16_nvfp4.yaml`) composed from the existing `units/w4_nvfp4` snippet. - Updates `modelopt.torch.quantization.config` built-in config constants to load `QuantizeConfig` objects from YAML with `load_config(..., schema_type=QuantizeConfig).model_dump(exclude_unset=True)` via a new `_load_quantize_config_dict` helper; the constants remain plain `dict[str, Any]` for backwards compatibility with consumers that do mapping-style mutation (e.g. `entry["cfg"]` assignment). - Simplifies the cfg-list loader (`_load_quantizer_cfg_dict_list`) down to a 4-line list/single normalization now that the three call sites all load schema-typed YAMLs. - Adds/updates recipe loader coverage for built-in schema-backed config snippets. ### Latent-bug fixes surfaced by the refactor Two small correctness fixes are included alongside the mechanical refactor; flagging them explicitly: - **`examples/diffusers/quantization/quantize.py`** — adds an explicit `base_cfg = copy.deepcopy(base_cfg)` before applying runtime overrides. The existing `# Build a fresh config dict so we never mutate the global constants` comment had been aspirational only; in practice `reset_set_int8_config` accumulated `PercentileCalibrator` entries into `mtq.INT8_SMOOTHQUANT_CFG`/`INT8_DEFAULT_CONFIG` across repeated calls, and `set_quant_config_attr` added `trt_high_precision_dtype` keys into globally-shared cfg dicts. The deepcopy makes the code match the comment. - **`choices` set in `modelopt/torch/quantization/config.py`** — adds `MXFP6_DEFAULT_CFG` and `NVFP4_W4A4_WEIGHT_LOCAL_HESSIAN_CFG` to the documented public set of valid `mtq.*_CFG` names. Both constants exist on main but were missing from `choices`, so CLIs that gate on `mtq.config.choices` (e.g., `hf_ptq.py --qformat`) couldn't reach them even though the configs themselves were fully supported. ### Usage Existing Python imports continue to work: ```python import modelopt.torch.quantization as mtq cfg = mtq.FP8_DEFAULT_CFG model = mtq.quantize(model, cfg, forward_loop) ``` The built-in constants are plain `dict[str, Any]` (sparse — only explicitly-set fields are present), but their definitions now come from YAML snippets and presets composed through the existing `$import` system. Reusable YAML snippets can be composed through `$import`, for example: ```yaml # modelopt-schema: modelopt.torch.quantization.config.QuantizeConfig imports: base_disable_all: configs/ptq/units/base_disable_all w4a4_nvfp4_nvfp4: configs/ptq/units/w4a4_nvfp4_nvfp4 default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers algorithm: max quant_cfg: - $import: base_disable_all - $import: w4a4_nvfp4_nvfp4 - $import: default_disabled_quantizers ``` ### Testing Local checks run: - `nox -s "unit-3.10(torch_211, tf_latest)"` — 2329 passed, 12 skipped. - `nox -s pre_commit_all` — all hooks pass (ruff check / ruff format / mypy / YAML format / license / bandit / markdownlint). - YAML parse + `$import` resolution sanity check across all changed config files. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ Existing built-in Python config constants keep the same public names and dict semantics. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ Adds/updates recipe loader coverage for schema-backed built-in snippets. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ ### Additional Information This PR was previously stacked on #1405, which has since merged to `main`. The branch has been rebased onto `main` and no longer depends on any other open PR. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Many new quantization numeric configs and PTQ presets added (INT4/INT8/MXFP4/MXFP6/MXFP8/MXINT8/NVFP4), plus Diffusers, KV-cache (affine/cast/rotate) and MLP/MoE-targeted presets. * **Refactor** * Presets and shared snippets migrated to schema-backed YAML sources and centralized loading; INT8 percentile calibration avoids mutating shared base configs. * **Tests** * Tests now discover packaged config snippets at runtime and validate import/append behaviors. * **Documentation** * Presets README and numerous header descriptions updated. * **Chores** * Minor typing and script improvements. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1423?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
a5bc6f8123 |
Add DATASET_COMBOS for grouped calibration datasets (#1508)
## Summary - Add ``DATASET_COMBOS`` to ``modelopt.torch.utils.dataset_utils`` — single ``--dataset`` tokens that fan out to several entries in ``SUPPORTED_DATASET_CONFIG``. The per-entry ``num_samples`` is split evenly across the members inside ``get_dataset_dataloader``. - Two initial combos: - ``cnn_nemotron_v2_mix`` → ``cnn_dailymail`` + ``nemotron-post-training-dataset-v2``. Replaces the hardcoded two-element fallback list in ``hf_ptq.py`` when ``--dataset`` is omitted. - ``nemotron-post-training-v3`` → the seven ``nvidia/Nemotron-*`` SFT datasets registered in #1498 (mirroring the upstream [`nemotron-post-training-v3` collection](https://huggingface.co/collections/nvidia/nemotron-post-training-v3)). - ``get_supported_datasets()`` now appends combo names so they show up in ``--dataset`` help. - ``hf_ptq.py``'s default ``--calib_size`` bumped from ``512`` to ``1024`` so the ``cnn_nemotron_v2_mix`` combo's even split preserves the previous total sample count (was 512 per-dataset × 2 datasets = 1024; now 1024 split → 512 per-dataset × 2). ``--calib_size`` now denotes the total calibration budget regardless of combo cardinality. - Reject mixing a combo with one of its member datasets in the same ``--dataset`` list (e.g. ``cnn_dailymail,cnn_nemotron_v2_mix``) — combo would otherwise double-sample the explicit member with a smaller per-member quota. - Reject combo names in ``get_dataset_samples``; combos are dataloader-only. The error message points callers to ``get_dataset_dataloader``. - Validate ``DATASET_COMBOS`` at import time: empty member lists, name collisions with ``SUPPORTED_DATASET_CONFIG``, and references to unknown datasets raise ``ValueError`` up front. ## Test plan End-to-end validated against ``/hf-local/Qwen/Qwen3.5-0.8B`` via ``get_dataset_dataloader`` on the actual streamed data, plus 5 new unit tests in ``TestDatasetCombosExpansion`` (all 44 tests in ``test_dataset_utils.py`` pass with no regressions). - [x] ``python -c "from modelopt.torch.utils.dataset_utils import DATASET_COMBOS, get_supported_datasets; assert 'cnn_nemotron_v2_mix' in get_supported_datasets() and 'nemotron-post-training-v3' in get_supported_datasets()"`` - [x] ``hf_ptq.py`` with no ``--dataset`` flag still calibrates on cnn_dailymail + nemotron-post-training-dataset-v2 with the same total sample count as before. - [x] ``--dataset nemotron-post-training-v3 --calib_size 1024`` allocates 146 per member across the seven Nemotron datasets; full 1022-sample dataloader builds without error. - [x] ``--dataset cnn_dailymail,nemotron-post-training-v3 --calib_size 256,1024`` composes correctly: 256 from cnn_dailymail (as a plain entry) plus the 7-way split from the combo. (The earlier ``cnn_dailymail,cnn_nemotron_v2_mix`` example is rejected by design since ``cnn_dailymail`` is a member of that combo.) - [x] ``--dataset cnn_dailymail,cnn_nemotron_v2_mix`` raises ``ValueError`` with a clear message. - [x] ``get_dataset_samples("cnn_nemotron_v2_mix", ...)`` raises ``ValueError`` pointing to ``get_dataset_dataloader``. - [x] Unit tests: ``pytest tests/unit/torch/utils/test_dataset_utils.py`` — 44 passed. ## Post-validation fix End-to-end testing surfaced that the original ``nemotron-sft-agentic-v2`` entry kept the two splits (``interactive_agent``, ``tool_calling``) that pyarrow's streaming JSON reader cannot parse, and excluded ``search`` which is the only clean split. Failures reproduce deterministically across cache wipes with ``force_redownload``, so they are content-level defects in the published JSONL files at the pinned revision, not local artifacts: - ``interactive_agent`` — heterogeneous schema (``Column(.../member_id/type) changed from string to array``) at JSONL row 4. - ``tool_calling`` — malformed JSON row in a later shard, fails at sample ~885 with ``Missing a closing quotation mark in string``. - ``search`` — streams cleanly (verified to 2500 samples). Commit ``10f3cfd`` corrects ``nemotron-sft-agentic-v2`` to use only ``search``, with an updated comment. The CHANGELOG calls this out as a separate bullet so the behavior change on a previously-released dataset entry from #1498 is discoverable. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Dataset combo support: a single dataset token can expand into multiple registered datasets with even sample splitting; predefined combos added (e.g., cnn_nemotron_v2_mix, nemotron-post-training-v3) and listed as supported. * **Updates** * Default dataset when none specified now uses cnn_nemotron_v2_mix. * Calibration size default increased from 512 to 1024. * **Bug Fixes** * nemotron-sft-agentic-v2 now uses only the deterministic "search" split to avoid streaming JSON errors. * **Tests** * Added coverage for combo expansion, splitting, overlap validation, and rejection behavior. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1508?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
7038dec918 |
[1/2Refactor] speculative decoding: use mto config subsystem (#1328)
### What does this PR do?
Type of change: new feature
Port the speculative-decoding example to ModelOpt's recipe/config
subsystem: `model` / `data` / `training` / `<algo>` now load from a
single YAML with Pydantic validation and OmegaConf dotlist overrides.
Adds built-in `eagle3` / `dflash` recipes, drops the redundant
`training.mode` field (inferred from recipe class), and shrinks
`main.py` by ~145 lines (−208 / +63).
JIRA: OMNIML-3859
### Usage
```bash
python main.py --config general/speculative_decoding/eagle3 \
model.model_name_or_path=meta-llama/Llama-3.2-1B \
data.data_path=train.jsonl \
training.output_dir=ckpts/test
```
### Testing
- `pytest tests/unit/recipe/test_loader.py` — new coverage for Eagle /
DFlash YAML loading, dotlist overrides, and field-level validation.
- Smoke-trained both built-in `eagle3` and `dflash` recipes end-to-end.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ❌ — `main.py` CLI switched to
`--config <recipe>` (+ dotlist overrides); the old argparse flags are
removed.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
deps (`pydantic`, `omegaconf` already in core).
- Did you write any new necessary tests?: ✅ —
`tests/unit/recipe/test_loader.py`.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — to be added.
### Additional Information
Follow-up to the `modelopt.recipe` subsystem introduced for PTQ; this PR
extends the same declarative-YAML pattern to speculative decoding
(Eagle3 / DFlash / Medusa).
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added typed speculative-decoding recipe support for EAGLE, DFlash, and
Medusa; CLI dotlist overrides supported for single-file recipes.
* Trainer/config schema extended with speculative-training fields and
draft-vocab cache loading for Eagle.
* **Bug Fixes**
* Offline training no longer mutates model configs; loader enforces
required algorithm sections and prints recipe/config only on the primary
process.
* Reduced noisy per-rank logging by restricting status output to the
primary process.
* **Tests**
* Expanded tests for recipe loading, dotlist overrides, validation
strictness, and error cases.
* **Documentation**
* Recipe YAMLs updated with metadata and usage notes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
7f1f223d8f |
[Feat]: Add dflash in specdec-bench (#1432)
### What does this PR do?
Type of change: new feature
Adds **DFLASH** speculative decoding support to the `specdec_bench`
example for both the SGLang and vLLM backends.
- `run.py`: register `DFLASH` in `--speculative_algorithm` choices.
- `models/sglang.py`: unify the SGLang engine setup so `engine_kwargs`
is built once and shared between the speculative and non-speculative
paths; add a DFLASH branch (default `speculative_num_draft_tokens=8`,
optional `speculative_dflash_draft_window_size`); set
`disable_cuda_graph_padding=True` and
`cuda_graph_max_bs=max_concurrent_requests` to avoid CUDA-graph
bucket-padding mismatches during DFLASH replay; expose
`mamba_scheduler_strategy` passthrough (needed for Qwen3.5).
- `models/vllm.py`: add a DFLASH branch wiring `method="dflash"` with
`speculative_num_draft_tokens` (default 8).
### Usage
```bash
python run.py \
--model_dir <target_model> \
--tokenizer <target_model> \
--draft_model_dir <dflash_draft_model> \
--mtbench mtbench.jsonl \
--engine SGLANG \
--speculative_algorithm DFLASH \
--concurrency 32 \
--tp_size 4 --ep_size 4
```
Use `--engine VLLM` for the vLLM backend.
### Testing
Manually validated end-to-end on Kimi-K2.5-NVFP4 + MTBench with both
SGLang and vLLM backends. `specdec_bench` has no existing test
infrastructure, so no automated tests are added (consistent with how
other algorithms in this example are handled).
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (purely additive — new
algorithm choice; existing algorithms unchanged)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ (specdec_bench is
example-only with no test harness; matches other algorithms)
- Did you update Changelog?: N/A (example-only change)
- Did you get Claude approval on this PR?: ❌
### Additional Information
Requires SGLang and vLLM versions that include DFLASH support.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added DFLASH as a supported speculative decoding algorithm option (CLI
selectable).
* DFLASH support extended to multiple inference backends with
configurable draft-token count and optional draft-window size; CLI now
emits an informational note when some options are ignored.
<!-- review_stack_entry_start -->
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1432)
<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
a451a2baf8 |
Support NVFP4 W4A16 quantization (#1313)
### What does this PR do?
This PR supports NVFP4 W4A16 quantization.
### Usage
```python
python examples/llm_ptq/scripts/huggingface_example.sh \
--model Qwen/Qwen3-8B \
--quant w4a16_nvfp4 \
--calib 1 \
--kv_cache_quant none \
--tasks quant
```
### Testing
<!-- Mention how have you tested your change if applicable. -->
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
### Additional Information
NVFP4 W4A16 is currently supported on vLLM only. TensorRT-LLM and SGLang
do not support this format.
vLLM loads NVFP4 W4A16 checkpoints via its `CompressedTensorsW4A16Fp4`
scheme, which requires the checkpoint to be in `compressed-tensors`
format. The raw ModelOpt export from this PR is still not directly
loadable by vLLM. After running hf_ptq.py, the checkpoint must be
converted by:
1. Renames tensors: .weight (uint8) → .weight_packed, .weight_scale_2
(fp32) → .weight_global_scale (inverted: 1 / weight_scale_2)
2. Rewrites config.json: sets quant_method: "compressed-tensors",
format: "nvfp4-pack-quantized", quantization_status: "compressed", and
derives the ignore list
automatically from layers that lack a weight_packed tensor.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* NVFP4 W4A16 weight-only quantization added (FP4 weights,
group_size=16; BF16 activations; no calibration-forward pass).
Selectable via qformat/CLI and included in export flow.
* New --exclude_modules CLI option (and EXCLUDE_MODULES env support in
example script) to skip quantizing specific modules.
* **Documentation**
* Changelog entry describing NVFP4 W4A16 usage and vLLM deployment via
compressed-tensors conversion.
* **Tests**
* Added test coverage for NVFP4 W4A16 export.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Co-authored-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
62401e16df |
fix: layerwise calibration backward-compat, recipe split, batch-size guard (#1310)
## Summary Follow-up to #1251 (which renamed `use_sequential` → `layerwise`). Three related fixes bundled: 1. **Backward-compatible config loading.** PTQ checkpoints saved before #1251 store the legacy `use_sequential` key in the calibration-algorithm config, so loading them now raises `ValidationError: Extra inputs are not permitted (use_sequential)` because `QuantizeAlgorithmConfig` uses `extra='forbid'`. Accept `use_sequential` as an alias for `layerwise` via `AliasChoices`. The field still serializes as `layerwise`, so round-trips through the current schema are clean. 2. **Recipe split.** `nvfp4_experts_only-fp8_kv` previously enabled layerwise calibration by default, which changes the calibration flow materially. Split into two recipes: - `nvfp4_experts_only-fp8_kv.yaml` — default (no layerwise) - `nvfp4_experts_only-fp8_kv_layerwise.yaml` — layerwise variant 3. **`hf_ptq` batch-size guard.** Auto batch-size detection is not supported together with layerwise calibration. Default to `batch_size=1` when layerwise is enabled and the user hasn't set a batch size explicitly. Originally reported by Jenny Chen while resuming a PTQ checkpoint via `restore_sharded_modelopt_state`: ``` pydantic_core._pydantic_core.ValidationError: 1 validation error for MaxCalibConfig use_sequential Extra inputs are not permitted [type=extra_forbidden, input_value=False, input_type=bool] ``` ## Test plan - [x] `tests/unit/torch/quantization/test_config_validation.py` — legacy alias accepted, current name accepted, dump serializes under current name, `extra='forbid'` still rejects unknown keys. - [x] `pre-commit run` — clean. ### Before your PR is *Ready for review* - Is this change backward compatible?: ✅ (restores compatibility for pre-#1251 checkpoints) - New PIP dependency: N/A - New necessary tests: ✅ - Changelog update: N/A (bug fix) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added new PTQ recipe for efficient layerwise calibration of large models. * Automatic batch size optimization for layerwise calibration recipes. * Backward compatibility support for legacy input naming conventions. * **Documentation** * Updated recipe guides and changelog with new layerwise calibration recipe. * **Tests** * Added validation tests for configuration compatibility. [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1310) <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> |