mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
main
522
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
42458def24 |
ci: fix torch_trt on torch 2.13; default unit tests to torch 2.13; add example allow-failure hatch (#1951)
### What does this PR do? Type of change: Bug fix + CI / infra torch 2.13 + torchvision 0.28 were published to PyPI on 2026-07-08 and broke the `onnx (torch_trt)` example job (which had passed the day before). This PR fixes that break and hardens CI against the next one: 1. **Fix `torch_trt` (torch/torchvision/torch-tensorrt trio pin).** The base install pulled torch 2.13 / torchvision 0.28, then `torch-tensorrt 2.12.1` downgraded torch back to 2.12 but left torchvision at 0.28 (which pins `torch==2.13`) — breaking `import`. `examples/torch_trt/requirements.txt` now caps `torch-tensorrt>=2.4.0,<2.13` + `torchvision<0.28` so the trio stays consistent (also protects direct `pip install -r` users). 2. **Unit tests default to torch 2.13.** `noxfile.py` gains `torch_213` (`torchvision~=0.28.0`); the required `linux`/`windows` jobs and the multi-version Python spread (3.10/3.11/3.13/3.14) now run torch 2.13, with torch 2.8–2.12 kept as back-compat legs on Python 3.12. 3. **Per-example allow-failure escape hatch.** `_example_tests_runner.yml` gains an `allow_failure` input; when set, a **test-run** failure is surfaced as a `::warning::` via `continue-on-error` instead of blocking the PR. `example_tests.yml` derives it per example from the repo variable **`ALLOW_FAILURE_EXAMPLE_TESTS`** (comma-separated example names, comma-wrapped so `onnx` ≠ `torch_onnx`). Future breakages can be quarantined by updating the variable — no code change / PR required. ### Testing - Ran the new default unit session locally in an isolated uv venv (torch **2.13.0**+cu130, torchvision **0.28.0**+cu130, transformers **5.12.1**): ``` nox -s "unit-3.12(torch_213, tf_latest)" => 2813 passed, 15 skipped, 1786 warnings in 250.37s ``` - Verified the allow-failure hatch: with `ALLOW_FAILURE_EXAMPLE_TESTS=torch_trt`, the (previously failing) `torch_trt` job reports success with a warning and does not block the required example check. `vars` is re-read on each job attempt, so "Re-run failed jobs" picks up the variable without a fresh trigger. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (CI-only; older torch versions still covered) - If you copied code from any other sources or added a new PIP dependency: N/A (no new dependency; only version caps) - Did you write any new necessary tests?: N/A (CI configuration change) - Did you update Changelog?: N/A (CI infra, no user-facing API change) - Did you get Claude approval on this PR?: ❌ (pending — will run `/claude review`) ### Additional Information The `ALLOW_FAILURE_EXAMPLE_TESTS` repo variable can be cleared for `torch_trt` now that the requirements pin lands the real fix; keep it as the standing escape hatch for future example breakages. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Example test jobs can now be configured to “allow failure” without failing the workflow; when enabled, a warning annotation is emitted. * **Bug Fixes** * CI unit-test and GPU-test configurations were refreshed (including a reduced timeout for the `gpu_megatron` job). * **Tests** * Updated unit-test coverage to use the newest Torch 2.13-based setup by default, with back-compat retained where applicable. * **Documentation** * Added inline guidance for how the allow-failure examples list is specified. * **Chores** * Refreshed `torch_trt` example dependency constraints to improve compatibility. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
6b4ad85849 |
Qwen-Image diffusers PTQ: FP8 / NVFP4 / NVFP4-SVDQuant HF checkpoints (#1706)
### What does this PR do? Type of change: New feature Adds **Qwen-Image** (`Qwen/Qwen-Image`, `QwenImageTransformer2DModel`) to the diffusers quantization example and exports HuggingFace checkpoints in three precisions — **FP8**, **NVFP4**, and **NVFP4 + SVDQuant** — through the unified HF export. - Registers `--model qwen-image` (lazy diffusers import; no `trust_remote_code`). - Transformer-block-range recipe: quantizes only the linears under `transformer_blocks`, keeping the **first 2 / last 2** blocks (and everything outside `transformer_blocks`) in original precision. Applied **before** calibration so SVDQuant never mutates the excluded blocks. Expressed with the top-level `enable` `QuantizerCfgEntry` field (disable-all → re-enable `transformer_blocks` → disable first/last-N). - SVDQuant export (AWQ-style): promotes quantizer-owned tensors to clean module-level safetensors keys at export time — `weight_quantizer.svdquant_lora_a/b → <module>.svdquant_lora_a/b` and `input_quantizer._pre_quant_scale → <module>.pre_quant_scale` — with a documented `NVFP4_SVD` `quantization_config` (`group_size`, `has_zero_point: false`, `pre_quant_scale: true`, `lora_rank`). **Core SVDQuant quantization code (`modelopt/torch/quantization`) is unchanged.** - Shared export-path change — **intentionally global** (applies to all diffusers exports — SDXL / Flux / Wan, not just Qwen; the full export suite was verified green on GB200): `hide_quantizers_from_state_dict` now strips quantizer state from *all* modules (not just quant-linears) so calibrated norm-layer input quantizers no longer leak `input_quantizer._amax`. (An earlier `max_shard_size` workaround was dropped after merging `main`: #1794 makes the ComfyUI layerwise-metadata post-processing a no-op unless explicitly opted in, so a default sharded export no longer hits the unsupported-sharded path.) ### Usage ```bash python examples/diffusers/quantization/quantize.py \ --model qwen-image --override-model-path <Qwen-Image> --model-dtype BFloat16 \ --format fp4 --quant-algo svdquant --lowrank 32 \ --calib-size 64 --n-steps 20 \ --hf-ckpt-dir <out> # FP8: --format fp8 --quant-algo max # NVFP4: --format fp4 --quant-algo max ``` ### Testing - Focused unit + example tests pass on GB200 (sm_100): block-range recipe, `NVFP4_SVD` config schema, SVDQuant forward/fold (LoRA stays on `weight_quantizer`), Qwen dummy-input / strict-QKV-fusion / promotion, pipeline loading, and the diffusers HF-export test for Qwen FP8 / NVFP4 / SVDQuant. - Full `tests/examples/diffusers/test_export_diffusers_hf_ckpt.py` is green (SDXL, Flux, Qwen, Wan2.2) — confirms the shared export changes do not regress other models. - End-to-end on the real `Qwen/Qwen-Image` (~20B): all three formats export valid HF checkpoints — only `transformer_blocks` 2..57 quantized, nothing outside, no quantizer/`_amax` leak, correct `weight_scale`(`_2`)/`input_scale`, promoted SVDQuant keys (rank-consistent shapes), and the expected `quantization_config`. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ <!-- live-model LoRA storage unchanged; existing exports unaffected --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ❌ <!-- new-example feature; add a CHANGELOG.rst entry if required --> - Did you get Claude approval on this PR?: ❌ <!-- run /claude review --> ### Additional Information All changes are confined to the diffusers example (`examples/diffusers/quantization`) plus the shared export path (`modelopt/torch/export`); the core quantization library is untouched. ### Follow-up (next step): fused-QKV SVDQuant for sglang / Nunchaku This export keeps attention `q/k/v` (and `add_q/k/v_proj`) as **separate** projections — the diffusers-native layout. That matches sglang's bf16 / FP8 / plain-NVFP4 paths (which also keep QKV separate) and ModelOpt/TRT-LLM consumers, so those load 1:1. sglang's **NVFP4-SVDQuant (Nunchaku)** path, however, builds a **fused** `to_qkv` with a *single* fused rank-r LoRA in Nunchaku-native format (`proj_down`/`proj_up`, `smooth_factor`, `wscales`/`wtscale`). Our per-projection tensors (`svdquant_lora_a/b` + `pre_quant_scale`; three independent rank-r decompositions) are not directly loadable there — and cannot be fused at load time, because the fp16 weight residual needed to derive a single fused rank-r is not preserved after export. **Planned next step:** an opt-in fused-QKV SVDQuant export mode that fuses q/k/v **before** SVDQuant calibration (yielding one rank-r over the fused weight) and emits a Nunchaku-compatible layout, enabling lower-latency fused-QKV inference in sglang. Tracked as a separate follow-up. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Qwen-Image (`QWEN_IMAGE`) model quantization and Diffusers export support * Added NVFP4_SVD (SVDQuant) export configuration support * Added transformer block-range quantization recipes (exclude first/last blocks) * **Bug Fixes** * Improved missing-pipeline error messaging for Qwen-Image * Prevented quantizer-related tensor/buffer leakage by promoting and cleaning quantizer outputs during export * **Tests** * Added Qwen-Image HF checkpoint export tests and offline fixtures * Added unit coverage for SVDQuant promotion/clean state-dict keys <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
bc5bc1ac5f |
[Feat]: Add Final Norm for vLLM Hidden Extractor (#1846)
### What does this PR do? **Type of change:** Bug fix vLLM captures the final-layer hidden state *before* the model's final norm, but the offline/streaming distillation path fed it straight into `lm_head`, so the reconstructed base logits (the KD target) were computed from un-normed hidden states. This PR re-applies the base model's final norm before `lm_head` when the producer declares a pre-norm capture (`base_hidden_prenorm`), for both DFlash and EAGLE: - Producer sets `base_hidden_prenorm` (streaming: `True`; offline: from the dump). - Consumer (`_maybe_apply_base_final_norm`) re-applies the base final norm, and **fails loud** if pre-norm is declared but the model's norm type isn't supported (no silent corruption). - `FakeBaseModel` now loads the base final norm (+ `rope_theta`/`rms_norm_eps`); norm type is gated by an explicit `model_type` allowlist (gpt_oss excluded pending a matching norm class). ### Testing `tests/unit/torch/speculative/plugins/test_modeling_final_norm.py`; DFlash/EAGLE streaming training verified end-to-end. - Backward compatible?: ✅ (post-norm captures declare `base_hidden_prenorm=False` → unchanged) - New tests: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Expanded speculative decoding/distillation to reconstruct missing base-model logits by optionally applying a base model’s final pre–LM-head normalization when pre-norm hidden states are provided. * Streaming dataset now emits `base_hidden_prenorm` and can configure RDMA backends from environment settings. * **Bug Fixes** * Rejects mixed `base_hidden_prenorm` values within a batch. * Fails fast on streaming token-length mismatches. * Streaming dataset loading supports directory inputs by expanding sorted JSONL shards (and errors if none are found). * **Tests** * Added/updated unit and dataset tests for final-norm behavior and the new batch field. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
43fee0cd70 |
feat(export): quant-aware reverse weight conversion for unified HF export (#1833)
## What
ModelOpt's unified HF export builds the state dict from the
**in-memory** (transformers post-conversion) module names and disables
transformers' save-side `revert_weight_conversion` (it raises
`IndexError` on ModelOpt's 0-d scalar scale tensors). So when
transformers applies a load-time `conversion_mapping` (renamed MoE
leaves, `block_sparse_moe`↔`mlp`, reordered `model`/`language_model`
prefix, fused dense `gate_up_proj`), the exported tensor names no longer
match the **original HF hub checkpoint**, breaking the
unified-checkpoint contract.
Observed concretely on MiniMax-M3: `nvidia/MiniMax-M3-NVFP4-v1` emitted
the converted names (`model.language_model.*`,
`mlp.experts.*.gate_proj`) instead of hub names
(`language_model.model.*`, `block_sparse_moe.experts.*.w{1,2,3}`).
## How
New `modelopt/torch/export/quant_aware_conversion.py` performs a
**quantization-aware reverse conversion**, derived from the model's
conversion mapping via transformers' own `reverse_transform()` (so
anchored regex renamings reverse correctly), carrying each weight's
companion scale tensors (`weight_scale`, `weight_scale_2`,
`input_scale`, `weight_scale_inv`, `bias`):
- **Rename** — key-level substitution; scale siblings follow the module
path automatically.
- **Split** — un-fuse a dense output-dim concatenation (`gate_up_proj` →
`gate_proj`+`up_proj`): split `weight`/`weight_scale`/`bias` on the
fused dim, duplicate the 0-d scalars.
Key insight from GPU validation: **ModelOpt's export already expands
fused, stacked in-memory experts** (`experts.gate_up_proj [E,2F,H]`)
into per-expert 2-D linears before save, so the reverse for experts is a
pure per-expert leaf **rename** (`gate_proj`→`w1`, `up_proj`→`w3`,
`down_proj`→`w2`), not a 3-D un-stack. Converters are classified by
their reversed ops: `SplitModulelist` present ⇒ expert rename;
`Chunk`-only ⇒ dense split.
`export_hf_checkpoint` applies this before `save_pretrained`; anything
not reversible raises `QuantConversionUnsupportedError` and the export
**falls back to prior behavior with a warning** (non-breaking).
## Tests
- **Unit** (`tests/unit/torch/export/test_quant_aware_conversion.py`,
CPU): rename carries scales; dense `gate_up_proj` un-fuse with scale
split + scalar duplication; 3-D / non-divisible guards; end-to-end
MiniMax-M3-like reversal.
- **GPU** (`tests/gpu/torch/export/test_quant_aware_conversion_gpu.py`):
quantize a tiny **Mixtral** to NVFP4 → `export_hf_checkpoint` → assert
exported tensor names **exactly equal** the canonical hub names from
transformers' own `revert_weight_conversion` on the reference model
(experts land as `block_sparse_moe.experts.N.w{1,2,3}`; no fused
in-memory names remain).
## Validation
- ✅ Unit tests pass (CPU). Existing `test_unified_export_hf.py` remains
green (no regression).
- ✅ **GPU end-to-end passes**: tiny Mixtral NVFP4 quantize→export yields
**0 missing / 0 extra** vs. canonical hub names (transformers 5.3.0, RTX
6000 Ada). Mixtral's conversion (`block_sparse_moe`↔`mlp`, expert
gate/up fuse) is the same machinery MoE VLMs like MiniMax-M3 use.
- Note: transformers 5.3.0 has no `minimax_m3_vl` entry, so M3 itself
isn't exercised in CI here; Mixtral is the structural surrogate. When a
transformers build with the M3 mapping is available, the same test can
be parametrized for M3.
## Follow-ups
- Extend the derivation if a model needs stacked-scalar-scale handling
that ModelOpt does not pre-expand (none known today).
- Upstream fix to transformers `Chunk.convert()` for 0-d tensors, then
drop the `_patch_revert_weight_conversion` no-op.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved Hugging Face export so quantized models save with canonical
hub tensor names instead of temporary in-memory names.
* Fixed export handling for scalar scale tensors and other quantized
companion tensors, helping checkpoints round-trip more reliably.
* Added a fallback path so export continues even when a conversion
pattern can’t be safely reversed.
* Exported Mixtral-style expert weights now preserve the expected hub
layout in saved checkpoints.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
b96a785b3c |
Add AutoQuantize recipe support (#1856)
### What does this PR do? **Type of change:** New feature (AutoQuantize recipes). The `--auto_quantize_*` CLI flags are **deprecated but still work** (kept as a thin backward-compat shim) — **not** removed. Makes **AutoQuantize recipe-driven**: `mtq.auto_quantize` is configured by a declarative YAML recipe (`--recipe`). The old `--auto_quantize_*` flags are converted into an `AutoQuantizeConfig` **on the fly** and run the exact same recipe path (emitting a `DeprecationWarning`), so old commands keep working. The recipe path is verified **byte-identical** to the CLI. - **Cost model** (`quantization/config.py`, `algorithms.py`): new `effective_bits` field on `QuantizeConfig` (recipe-level override) and `QuantizerAttributeConfig` (per-format default). `estimate_quant_compression` resolves recipe-level → per-entry → `num_bits` heuristic. `configs/numerics/nvfp4.yaml` ships `effective_bits: 4.5` (block-scale-accurate) as the single source of truth. - **Recipe schema** (`recipe/config.py`, `recipe/loader.py`): `RecipeType.AUTO_QUANTIZE` + `AutoQuantizeConfig` / `AutoQuantizeConstraints` / `AutoQuantizeCost`. Fields: `constraints` (`effective_bits`, `cost_model`, `cost.active_moe_expert_ratio`), `candidate_formats`, `auto_quantize_method` (`gradient`/`kl_div`), `score_size`, `disabled_layers`, `cost_excluded_layers` (e.g. VL vision towers), `kv_cache`. - **Dispatch** (`examples/hf_ptq/hf_ptq.py`): recipe → mtq inputs via `_mtq_inputs_from_auto_quantize_config`; `_match_candidate_to_preset` resolves candidates to shipped presets and **guards export-compatibility** (rejects export-unsafe presets before the search). - **Deprecated CLI shim:** `_auto_quantize_config_from_cli` builds an `AutoQuantizeConfig` from the old flags and appends the shared base `disabled_layers` / `cost_excluded_layers` (loaded once as module constants in `recipe/config.py`, mirroring `_default_disabled_quantizer_cfg`). No model introspection, no new user flags. - **Shipped recipes:** `general/auto_quantize/` (`nvfp4_fp8_at_5p4bits`, `nvfp4_fp8_kl_div_at_5p4bits`, `nvfp4_mse_fp8_at_6p0bits`, `w4a8_awq_beta_fp8_at_6p0bits`, `w4a16_nvfp4_fp8_at_6p0bits-active_moe`) and model-specific `huggingface/qwen3_6_moe/auto_quantize/...`. Shared `configs/auto_quantize/units/base_disabled_layers` + `base_cost_excluded_layers` spliced via `$import`. **Migration (deprecated flag → recipe field):** `--auto_quantize_bits` → `constraints.effective_bits` · `--auto_quantize_method` → `auto_quantize_method` · `--auto_quantize_score_size` → `score_size` · `--auto_quantize_cost_model` → `constraints.cost_model` · `--auto_quantize_active_moe_expert_ratio` → `constraints.cost.active_moe_expert_ratio` · `--qformat fp8,nvfp4` → `candidate_formats`. `--auto_quantize_checkpoint` unchanged. ### Usage ```sh # Recipe (preferred) python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/auto_quantize/nvfp4_fp8_at_5p4bits --export_path <out> # Deprecated CLI (converted to a recipe on the fly, still works) python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --qformat nvfp4,fp8 --auto_quantize_bits 5.4 --export_path <out> ``` ### Testing - **GPU-free unit tests:** recipe loader; recipe→`mtq.auto_quantize` mapping incl. `cost_excluded_layers`; export-compat guard (reject/warn/no-bypass); deprecated-CLI→`AutoQuantizeConfig` conversion; `effective_bits` resolver + validators. - **Byte-identical export smoke:** recipe path on Qwen3.6-35B-A3B (`fp8 + w4a16_nvfp4 @ 6.0`, `active_moe`) → identical `hf_quant_config.json` across CLI/recipe; also confirmed the deprecated **CLI shim ≡ recipe** on the same VL MoE. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ **Yes** — `--auto_quantize_*` flags are deprecated but still work (converted to a recipe on the fly + `DeprecationWarning`). Plain PTQ CLI unaffected. - New PIP dependency / copied code: N/A - New tests?: ✅ - Updated Changelog?: ✅ (Deprecations) - Claude approval?: pending `/claude review` 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Juhi Mittal <juhim@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
75b580379a |
Create a user guide: ModelOpt for Researchers: Fast Experimentation Workflows (#1872)
### What does this PR do? Create a user guide: ModelOpt for Researchers: Fast Experimentation Workflows. ### Usage examples/researcher_guide/README.md ### Testing Reviewed + run mmlu-pro experiments used in the guide. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Expanded LM-Eval-Harness guidance with a linked pointer for shortening research iteration cycles. * Clarified that `--limit 10` evaluates 10 samples **per task** (not 10 total), and refined “quick smoke test” wording. * Added a new researcher-focused guide for fast model experimentation workflows, including evaluation command tips and benchmark limit/time/error planning guidance. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com> |
||
|
|
fbbc5989ce |
[6078291][OMNIML-3716] Add ViT FP8 + Torch-TRT example, wire softmax_quantizer in _QuantAttention (#1569)
### What does this PR do?
Type of change: new feature + bug fix
Adds a Torch-TensorRT deployment path for HuggingFace ViT and closes the
modelopt-side gap that prevented `*softmax_quantizer` from being applied
on the standard attention forward path.
* **New ViT PTQ recipes** under `modelopt_recipes/huggingface/vit/ptq/`:
* `fp8.yaml` — W8A8 per-tensor FP8 E4M3 on encoder Linear
weights/inputs;
attention Q/K/V BMMs + softmax output at FP8; per-block LayerNorm output
at FP8 (one shared Q/DQ feeds Q/K/V + MLP); patch-embed `nn.Conv2d`,
`classifier`, and the final `vit.layernorm` left FP16. Uses max
calibration.
* The recipe is self-contained (no `$import` of shared snippets) and
use the "specific-enable" style: narrow `parent_class` + path scoping
on every enable rule, so no `enable: false` carve-outs are needed.
* **New example** under `examples/torch_trt/`:
* `torch_tensorrt_ptq.py` — single-model pipeline (load HF model,
calibrate from `zh-plus/tiny-imagenet`, `mtq.quantize`,
`torch_tensorrt.compile`, verify the compiled-model argmax matches the
fake-quant argmax). Defaults to `google/vit-large-patch16-224`; pass
`--model_id` and `--recipe` to target any model + recipe combination.
`--no_pretrained` + `--model_kwargs` shrink the model for fast tests.
* `README.md` documenting the flow, the shipped recipes, hardware
requirements, and CLI usage.
* `requirements.txt`.
* **Bug fix in `modelopt/torch/quantization/plugins/huggingface.py`** —
inside
`_QuantAttention._quantized_attention`, the non-kitchen branch now
temporarily replaces `torch.nn.functional.softmax` (via the existing
`replace_function` context manager) with a wrapper that pipes the
softmax
output through `self.softmax_quantizer`. Previously the slot was created
on every registered attention class but only consumed by the optional
Kitchen MXFP8 flash-attention path, so FP8 / NVFP4 recipes that enabled
`*softmax_quantizer` saw it stay uncalibrated (`amax=None`) and emitted
no Q/DQ around the softmax output during ONNX / Torch-TRT export. With
this fix the `softmax_quantizer` is calibrated alongside the rest of
the model, and both the modelopt ONNX exporter and
`torch_tensorrt.compile`
pick up the Q/DQ pair. The patch short-circuits to the unwrapped call
when the quantizer is disabled (zero-overhead) and has no effect on SDPA
paths that fuse softmax inside a C++ kernel.
* **New e2e integration test** at
`tests/examples/torch_trt/test_torch_tensorrt_ptq.py` — mirrors the
`torch_onnx` test pattern: invokes the example through
`run_example_command`, parametrizes over the two precision modes (fp8,
nvfp4), uses a 1-layer ViT config (`--no_pretrained` + `--model_kwargs`)
so each parametrized case completes in under a minute. `importorskip` on
`torch_tensorrt` so the test is automatically skipped on hosts without
the package.
### Usage
```bash
# FP8 (Hopper / Ada) — default model is google/vit-large-patch16-224
python examples/torch_trt/torch_tensorrt_ptq.py \
--precision fp8 \
--calib_samples 128 \
--batch_size 1
# Custom model + custom recipe
python examples/torch_trt/torch_tensorrt_ptq.py \
--model_id <huggingface/model-id> \
--recipe <recipe-path-relative-to-modelopt_recipes-or-absolute-yaml>
```
### Testing
* Recipes load via `modelopt.recipe.load_recipe()` and pass
`QuantizeConfig` schema validation.
* Run `pytest tests/examples/torch_trt/test_torch_tensorrt_ptq.py` →
1 parametrized case passes on RTX 6000 Ada (fp8).
* End-to-end on `google/vit-base-patch16-224`: `mtq.quantize` with the
new
FP8 recipe followed by `torch_tensorrt.compile(ir="dynamo")` produces a
TRT engine whose argmax matches the FP16 baseline.
* ONNX exported from the torch path now contains Q/DQ on **12 / 12**
softmax outputs (was 0 / 12 before this PR's `_QuantAttention` fix),
matching the ONNX-CLI output's quantization layout.
Both FP8 paths land within 0.13 pp Top-1 of the FP16 baseline; Top-5 is
within 0.02 pp across all three.
* ImageNet-1k validation accuracy via the new
`torch_tensorrt_accuracy.py`
(full 50000 samples, batch=1, **every model Torch-TensorRT-compiled —
including the baseline** — so the comparison is apples-to-apples) for
the
example's default `google/vit-large-patch16-224`:
| Model (Torch-TRT) | Top-1 | Top-5 | Δ Top-1 vs baseline |
|---|---:|---:|---:|
| Baseline (FP16) | 81.99% | 96.01% | — |
| FP8 | 82.01% | 96.05% | +0.02 pp |
FP8 is within noise of the FP16 TRT baseline and NVFP4 W4A4 costs only
−0.13 pp Top-1 / −0.05 pp Top-5. Absolute Top-1 sits below the model
card's
~85.5% because evaluation uses the HF `AutoImageProcessor` default
preprocessing (direct 224×224 resize, no resize-then-center-crop),
applied
identically to all three models — so the deltas are the comparison
signal.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — new e2e integration test
under `tests/examples/torch_trt/`.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added Torch‑TensorRT FP8/NVFP4 deployment examples and end‑to‑end
scripts for HuggingFace ViT, plus ViT-specific PTQ recipes and
ImageNet-1k vs FP16 accuracy reporting.
* **Bug Fixes**
* Fixed softmax quantization and export/compilation edge cases (softmax
calibration during export, IO casting for empty tensors, routed expert
weight syncing, importer key handling).
* **Documentation**
* Added comprehensive example README with setup, usage, recipes,
evaluation, and hardware guidance.
* **Requirements**
* Pinned minimum versions for example dependencies.
* **Tests**
* Added tests validating the Torch‑TensorRT quantization examples for
fp8.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
973cb09cbe |
Refine DeciLM dtype handling in HF PTQ (#1869)
## Summary - factor config dtype resolution into helpers for HF PTQ model loading - keep DeciLM empty-init and final-load kwargs on `torch_dtype` while avoiding unsupported `dtype` forwarding - update the DeciLM dtype unit assertion for the follow-up behavior Follow-up to #1857 for NVBug 6359821. ## Validation - `pytest_pwd tests/examples/hf_ptq/test_example_utils.py -q -x` (`15 passed`) - `git diff --check` - `pre-commit run --files examples/hf_ptq/example_utils.py tests/examples/hf_ptq/test_example_utils.py` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved model loading so precision (dtype) is applied more consistently across supported loading paths, including DeciLM models. * Updated initialization to derive dtype from model configuration and pass the expected precision into model loading kwargs. * **Tests** * Updated test expectations to reflect the new dtype kwarg behavior during `from_pretrained` for causal language model loading scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
838b2053df |
Improve lm_eval readme to clarify how accelerate launch works (#1867)
### What does this PR do? Improve lm_eval readme to clarify how accelerate launch works for lm_eval. ### Usage see examples/llm_eval/README.md ### Testing - Manually reviewed changed documentation. - No impact on the commands used in the tutorial, only language-based clarification. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Clarified LM-Eval-Harness baseline instructions for multi-GPU evaluation. * Updated model-sharding guidance to emphasize fitting a single model across multiple GPUs and enabling larger batches (which may speed up evaluation). * Expanded data-parallel details to explain how `--num_processes` controls concurrent model copies and GPU allocation, including an 8-GPU example. * Removed outdated shorthand wording and replaced it with clearer, step-by-step guidance. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com> |
||
|
|
2fc352be2d |
Add VLM pruning and PTQ with image-text calibration (Megatron-Bridge) (#1792)
### What does this PR do?
Type of change: New feature
Adds **vision-language model (VLM) support** to the Megatron-Bridge
examples for both **Minitron pruning** (`prune_minitron.py`) and **PTQ**
(`quantize.py`). Only the **language model** is pruned/quantized — the
vision tower and vision→language projector are left in full precision —
and the full VLM is saved back. `hidden_size` is skipped for pruning
when it is shared with the vision→LM projector.
Supported VLMs (tested e2e): **Qwen{3,3.5}-VL** (dense; hybrid
GatedDeltaNet + gated attention) and **Gemma3-VL** (sliding/full
attention).
### Calibration (image-text)
Calibration is conditioned on real **image-text** data so the language
model's pruning importance / quantizer statistics see vision-conditioned
activations. The modality is inferred from `--calib_dataset_name`:
- an **image-text** dataset (default for VLMs,
`nemotron_vlm_dataset_v2`) drives the **full VLM forward**;
- a **text** dataset runs text-only calibration of the language model
(for text-vs-image ablations).
A shared `get_megatron_vlm_calibration_forward_loop` (built on
`megatron_prefill`) drives the full VLM forward over image-text pairs
from `vlm_dataset_utils` (`scienceqa`, `nemotron_vlm_dataset_v2`, with
config-driven subset/shard caps to bound downloads). It shards across
**data-parallel (DP)** ranks like the text loop (#1804); **context
parallelism (CP)** applies to text-only VLM calibration (the shared text
loop), not the multimodal forward — splitting the sequence would
misalign the merged vision embeddings.
### Results - Cosmos-Reason2-2B
Validated end-to-end on **Cosmos-Reason2-2B** (Qwen3-VL). Minitron NAS
prunes the language-model tower **1.72B → ~1.59B** (vision encoder +
projector frozen), top_k=1. Calibration data drives pruning importance;
image-text calibration runs the full VLM forward.
| Model | Calibration | MMLU | BLINK Rel-Depth | RealWorldQA |
|---|---|---|---|---|
| Baseline (1.72B) | — | 0.58 | 0.76 | 0.61 |
| Pruned (1.59B) | text (`nemotron-post-training-dataset-v2`) | 0.51\* |
~0.69 | ~0.57 |
| Pruned (1.59B) | image+text (`nemotron_vlm_dataset_v2`) | 0.49\* |
**0.77** | **0.61** |
\* Pruned MMLU on the 10% split (the pruning score function); baseline
MMLU is the full set. The VLM-benchmark numbers for the text row were
measured with a different text calibration set and are expected to be
similar for `nemotron-post-training-dataset-v2` (marked `~`).
> [!NOTE]
> These numbers come from short single runs on small eval splits — read
them for **high-level trends only**, not as exact values.
Takeaways: pruning the LM tower of a VLM works end-to-end. **Image-text
calibration** (this PR's feature) preserves the VLM benchmarks better
than text-only — BLINK Rel-Depth ~0.77 vs ~0.69 and RealWorldQA ~0.61 vs
~0.57, both close to the unpruned baseline (0.76 / 0.61) — which is the
motivation for calibrating on vision-conditioned activations.
### Results - Qwen3.5-9B
| Model | MMLU | MMStar |
|----------------------------|:------:|:------:|
| Qwen3.5-9B | 0.7003 | 0.6117 |
| Pruned-7B (text calib) | 0.5527 | 0.4411 |
| Pruned-7B (image+text calib) | 0.5107 | 0.3941 |
### Key changes
- `quantize.py`: quantizes the **root** model with non-LM (vision)
quantizers disabled, so the ModelOpt state lives on the root (required
by the Megatron save) while only the language model is quantized.
- `prune_minitron.py`: image-text (or text) calibration for VLM pruning
importance.
- Shared VLM calibration forward loop (`megatron_prefill`-based, unwraps
tuple outputs, DP-sharded) + `vlm_dataset_utils`.
- Tiny VLM test fixtures (Qwen3.5-VL, Gemma3-VL) with vision tokens
derived dynamically from the reference processor; VLM prune + quantize
example tests.
- README + CHANGELOG.
### Usage
```bash
# Prune the language model of a VLM (image-text calibration by default)
torchrun --nproc_per_node 2 prune_minitron.py \
--pp_size 2 \
--hf_model_name_or_path <vlm> \
--prune_target_params 3e9 \
--output_hf_path /tmp/vlm-pruned
# PTQ the language model of a VLM
torchrun --nproc_per_node 2 quantize.py \
--hf_model_name_or_path <vlm> \
--quant_cfg fp8 \
--export_megatron_path /tmp/vlm-fp8-megatron
```
### Testing
- `test_prune_minitron.py::test_prune_minitron_vlm` — Gemma3-VL,
image-text (ScienceQA) calibration; full load → prune (depth + ffn) →
save → reload.
- `test_quantize_export.py::test_quantize_vlm` — Qwen3.5-VL, text
calibration; quantize LM → save Megatron checkpoint.
- LM regression tests (`test_prune_minitron`,
`test_quantize_and_export`) unchanged and passing.
### Not in scope
- **HF unified export of a quantized VLM** is not yet supported;
`export.py` saves the Megatron checkpoint only for VLMs (tracked by a
TODO in `export.py`). The recommended path is to route the megatron→HF
quant export through Megatron-Bridge's
`AutoBridge.export_hf_weights_quant(quantization_checker, quant_fn,
quant_block_size)`, which reuses the bridge's per-model mcore↔HF mapping
— covering Qwen3.5-VL / Gemma3-VL and the vision tower/projector (left
full precision) for free — so modelopt supplies only the checker +
pack/scale fn + `hf_quant_config` (KV-cache scales need a separate
path). This avoids re-authoring per-model mappings in modelopt (cf.
#1482's Qwen3-VL-only `mcore_qwen3vl.py`).
> [!NOTE]
> Qwen3.5-VL **MoE** is not tested e2e: the Megatron-Bridge weight
conversion expects packed (`gate_up_proj`) experts that transformers'
tiny checkpoint doesn't emit. MoE pruning itself is covered by
`test_mcore_qwen35_gdn_moe_pruning`.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅
### Additional Information
Follow-up to the GatedDeltaNet/MLA/latent-MoE pruning PR (#1747).
Rebased on `main` to pick up CP/DP calibration (#1804); the VLM
calibration loop now shards across DP ranks the same way. `hidden_size`
pruning for VLMs (requires resizing the vision projector) is left for a
future PR.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added VLM-aware Minitron pruning and post-training quantization that
target only the language-model portion, keeping the vision
tower/projector in full precision.
* Calibration now auto-selects text vs image-text datasets based on
model type, with modality validation.
* Expanded Megatron-Core CP/DP guidance and introduced a `--cp_size`
flag in quantization examples.
* **Bug Fixes**
* Improved VLM generation/prefill output handling and made vocabulary
sizing more robust for VLM wrappers.
* **Tests / Documentation**
* Updated pruning/quantization docs and refreshed/added VLM-focused
tests.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
d70c48c1ff |
Fix HF PTQ empty-init dtype kwargs (#1857)
## Summary Fixes NVBug 6359821: `hf_ptq.py` can fail for remote/custom architectures like `DeciLMForCausalLM` when dtype-related kwargs are forwarded into model construction paths that do not accept them. This change keeps the fix scoped to the observed DeciLM/Llama Nemotron path. It resolves the init config used for empty-weight construction, derives dtype consistently from the resolved config, forwards the supported dtype kwarg for the DeciLM empty-weight probe, and drops unsupported dtype forwarding from the DeciLM real `from_pretrained()` load. NVBug: https://nvbugspro.nvidia.com/bug/6359821 ## Validation - `pre-commit run --files examples/hf_ptq/example_utils.py tests/examples/hf_ptq/test_example_utils.py` - `pytest_pwd tests/examples/hf_ptq/test_example_utils.py -q -x` (15 passed) - Actual `Llama-3_3-Nemotron-Super-49B-v1` end-to-end `hf_ptq.py` export on one node with 6 GPUs, Transformers 4.48.3: https://github.com/NVIDIA/Model-Optimizer/pull/1857#issuecomment-4845730927 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved model loading for Hugging Face remote-code scenarios by safely re-deriving the initialization configuration when needed, with a warning-based fallback. * Ensured precision is derived consistently from the resolved config (including dtype name handling) with a safe default when unspecified. * Tightened forwarding of precision-related kwargs and `trust_remote_code`, and avoided passing `max_memory` during config loading. * **Tests** * Added unit coverage for initialization config resolution (including failure fallback). * Extended integration-style coverage to validate dtype/kwarg forwarding, `trust_remote_code` behavior, and eval-mode initialization. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
6e2efd749b |
Fix lm_eval_hf freezing issue on multi-gpu slurm interactive node (#1831)
### What does this PR do? Fix lm_eval_hf freezing issue on multi-gpu slurm interactive node. **Note (Slurm interactive nodes):** On Slurm interactive nodes, `WORLD_SIZE` is set to the number of available GPUs in the shell environment. Running `python` directly causes `lm_eval` to hang waiting for peer ranks that were never spawned. Prepend `WORLD_SIZE=1` to any of the above commands to fix this. ### Usage see examples/llm_eval/README.md ### Testing Tested manually on a slurm interactive node with 8 GPUs. - Is this change backward compatible?: ✅ - Did you write any new necessary tests?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the LLM evaluation baseline to explicitly support both standard Hugging Face models and heterogeneous pruned Puzzletron checkpoints. * Added a Slurm interactive-nodes note advising `WORLD_SIZE=1` for direct `python` runs, while noting distributed launch tooling handles `WORLD_SIZE`. * Recommended using `--limit 10` for quick smoke tests. * Simplified evaluation instructions by linking to the LM-Eval-Harness guide instead of repeating commands/snippets. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com> |
||
|
|
72651b29a8 |
Fix Nemotron-H PTQ failure on Transformers 5.x with --trust_remote_code (moe_latent_size AttributeError) (#1839)
### What does this PR do?
Type of change: Bug fix
Quantizing remote-code checkpoints such as
`nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16` with `--trust_remote_code`
fails on Transformers 5.x during model loading:
```
AttributeError: 'NemotronHConfig' object has no attribute 'moe_latent_size'
```
(It works on Transformers 4.57.x.)
**Root cause:** In `examples/llm_ptq/example_utils.py::get_model`,
`AutoConfig.from_pretrained(..., trust_remote_code=True)` loads the
checkpoint's bundled **remote** `NemotronHConfig` (authored for
Transformers 4.55.4, which has no `moe_latent_size`). But because
`NemotronHForCausalLM` is a **built-in** class in Transformers 5.x, the
empty-weights device-map build instantiates the built-in model class,
whose modeling code reads `config.moe_latent_size`. The remote (old)
config and the built-in (new) model are a mismatched pair. Transformers
4.57.x only worked by luck — its built-in model never accessed that
field.
**Fix:** When instantiating the built-in model class, feed it a config
from the **same version as the model definition**. If the loaded config
came from remote code (its class module lives under
`transformers_modules`), re-derive it with the built-in class
(`AutoConfig` without `trust_remote_code`) so required fields get their
defaults. Non-remote configs are untouched. The subsequent real model
load already resolves the config via the built-in `config_class`, so
only the device-map build needed aligning.
### Usage
No API change. The previously-failing command now works:
```python
python examples/llm_ptq/hf_ptq.py \
--pyt_ckpt_path nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 \
--qformat fp8,nvfp4_mse --calib_size 64 \
--export_path ./output/nemotron-nano-fp8-nvfp4_mse \
--trust_remote_code --dataset cnn_dailymail --auto_quantize_bits 4.75
```
### Testing
- Reproduced the original `AttributeError` on Transformers 5.7.0, then
confirmed the fix resolves it.
- End-to-end: `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16` nvfp4 PTQ +
export completes successfully on Transformers 5.7.0.
- `tests/examples/llm_ptq/` unit tests (`test_example_utils.py`,
`test_hf_ptq_args.py`, `test_cast_mxfp4_to_nvfp4.py`) all pass.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ <!-- Only changes the config
used for the device-map build when a remote-code config is paired with a
built-in model class; non-remote configs and all other code paths are
unchanged. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- Example-script bug
fix; a regression test would require a remote-code checkpoint to load.
Existing llm_ptq unit tests pass and the fix was validated end-to-end.
-->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ <!-- Run `/claude review`.
-->
### Additional Information
The fix is general: re-deriving the config with the built-in class
handles any field the built-in model adds in future Transformers
releases, not just `moe_latent_size`.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved model initialization when using remote-code configurations by
re-deriving the initialization config for built-in model classes when
possible.
* Added a safe fallback to keep using the original configuration if
re-derivation fails.
* **Tests**
* Added unit tests covering remote-code config re-derivation, ensuring
unchanged configs are not re-derived, and verifying fallback behavior on
errors.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
|
||
|
|
f335459dc0 |
refactor(examples): rename llm_ptq → hf_ptq (symlink for back-compat) (#1759)
## What does this PR do? **Type of change:** refactor / deprecation (examples) Follow-up to #1705 (which consolidated `examples/vlm_ptq` into `examples/llm_ptq`). Since that example now covers Hugging Face **LLM and VLM** PTQ, the `llm_ptq` name is a misnomer. This renames the directory to `examples/hf_ptq` and leaves a relative symlink `examples/llm_ptq → hf_ptq` so existing paths/commands keep working during a deprecation window. Requested by @kevalmorabia97 on #1705 (with the symlink-for-back-compat approach), targeted for the **same 0.46 release** as the consolidation. ### Changes - `git mv examples/llm_ptq → examples/hf_ptq` and `tests/examples/llm_ptq → tests/examples/hf_ptq` (the CI runner maps the matrix name to both `examples/<name>` and `tests/examples/<name>`). - Add a tracked back-compat symlink `examples/llm_ptq → hf_ptq`. - Update CI matrices and all repo **path references** (docs, READMEs, agent skills, launcher/debugger tools, tests) from `llm_ptq` to `hf_ptq`. - Keep Python identifiers / test-util module names (`run_llm_ptq_command`, `llm_ptq_utils`) — they name the LLM-PTQ task, not the directory. - Preserve the CODEOWNERS team slug (`modelopt-examples-llm_ptq-codeowners`) and historical CHANGELOG entries; add a CHANGELOG deprecation note. ### Back-compat caveats (inherent to git directory symlinks) - ✅ Linux/macOS CLI usage and Python `cwd`/pytest resolution work through the symlink. - ⚠️ Windows git checkouts don't materialize symlinks by default (low impact — this example is Linux-only in practice). - ⚠️ GitHub web doesn't follow directory symlinks, so legacy external deep-links to `examples/llm_ptq/...` won't navigate in. All **internal** references are repointed to `hf_ptq`, so the symlink is only for legacy external/CLI use. ### Usage (unchanged via symlink) ```bash # New canonical path cd examples/hf_ptq scripts/huggingface_example.sh --model <hf_model> --quant fp8 # Old path still works (forwards via symlink) cd examples/llm_ptq && scripts/huggingface_example.sh --model <hf_model> --quant fp8 ``` ### Testing - `bash -n` on moved/edited shell scripts (new path + via symlink). - `py_compile` on moved/edited Python; test re-export shim repointed to `examples/hf_ptq/example_utils`. - Verified git tracks `examples/llm_ptq` as a single symlink (mode 120000), not a duplicated tree (no pre-commit / pytest double-processing). - `pre-commit run` on all changed files passes. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (relative symlink keeps `examples/llm_ptq` paths valid; see caveats above) - Did you write any new necessary tests?: N/A (pure rename; existing tests moved with the dir) - Did you update Changelog?: ✅ ### Additional Information Follow-up (later release): remove the `examples/llm_ptq` symlink once external references have migrated. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * PTQ guidance now directs to the unified Hugging Face PTQ flow, including VLM quantization via the shared `--vlm` entry point. * **Documentation** * Updated README and guide links, references, and command snippets to use `hf_ptq` (replacing `llm_ptq`). * Deprecated and consolidated `vlm_ptq` into `hf_ptq`; removed VILA/NVILA coverage from the Hugging Face PTQ examples. * **Bug Fixes** * Improved detection and routing so local/manual setup uses the correct PTQ source. * **Tests / Chores** * CI and example tests updated to run the `hf_ptq` variants. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c248dd5434 |
[Feat]: Domino support (#1710)
### What does this PR do? Type of change: New feature Adds **Domino** speculative decoding: the parallel DFlash draft backbone plus a lightweight **GRU causal correction head**. The backbone produces *base* logits for a full draft block in one forward; a GRU over the block's teacher-forced tokens produces a causal state that is fused with the backbone hidden state and projected to a vocab-sized logit correction on the block suffix — injecting the intra-block causal dependency the parallel backbone lacks. Trained with a dual loss `(1-λ)*final + λ*base`, where `λ_base` decays linearly 1→0 (curriculum: learn the parallel backbone first, then the correction). Reuses the DFlash mode/config/recipe; selected via `dflash_architecture_config.projector_type=domino` and routed to its own registry so `HFDominoModel` does not shadow `HFDFlashModel`. Exports in the z-lab/SpecForge drafter format (`prefix_gru.*` / `embed_proj.*`). > Note: the inference side (vLLM / AR evaluation) is intentionally **not** wired up yet — the correction head is not applied in serving. To be added once the inference path lands. ### Usage ```bash # Online training (recipe: projector_type=domino) uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_online_domino.yaml --yes ``` ### Testing CPU unit tests in `tests/unit/torch/speculative/plugins/test_hf_domino.py` cover conversion routing, the training forward (dual loss + grads), the λ schedule, and the export format. Online Qwen3-8B training validated end-to-end (loss curve below). <img width="1803" height="809" alt="image" src="https://github.com/user-attachments/assets/7c9d2001-bd80-4dec-919b-443e61089cca" /> ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (opt-in via `projector_type=domino`; DFlash path unchanged) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new dependency) - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Reference: SpecForge PR #571 (z-lab); drafter format `huggingface.co/Huang2020/Qwen3-8B-Domino-b16`. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added Domino speculative-decoding training with a decaying base/final dual-loss curriculum and Domino-specific lambda scheduling (training-only; inference wiring not yet included). * Added Domino draft-head export support for training checkpoints. * **Documentation & Configuration** * Added a Domino speculative-decoding training recipe and an HF Online Domino launcher configuration for Qwen3-8B. * **Refactor** * Updated speculative model conversion/export to route to Domino variants based on the configured projector type. * **Tests** * Added unit tests for Domino conversion, training loss/metrics, lambda decay behavior, and exporter output layout. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
d5962c4f3b |
Remove deprecated examples/llm_autodeploy (#1797)
Remove the AutoQuant + TensorRT-LLM AutoDeploy example, deprecated in 0.45, after the migration period. Record the removal under the 0.46 Backward Breaking Changes section. Users should use TensorRT-LLM's AutoDeploy directly together with ModelOpt PTQ in examples/llm_ptq. ### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated AutoDeploy guidance to a workflow: quantize with ModelOpt PTQ (via `llm_ptq`) to produce a unified Hugging Face checkpoint, then deploy with TensorRT-LLM AutoDeploy. * Revised Hopper notes to recommend using FP8 (and refreshed related optimization guidance). * **Breaking Changes** * Removed the deprecated `examples/llm_autodeploy` example and documented the new recommended approach. * **Chores** * Dropped obsolete example docs, scripts, and coverage; adjusted example test workflow to exclude `llm_autodeploy`; updated ownership mapping for the removed example path. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
33bfa8b1fe |
CI/Dev env bump (#1818)
### What does this PR do? Type of change: chore Bumps CI/dev tooling and test containers. **Container bumps** - NeMo test containers → 26.06 - TRT-LLM container → 1.3.0rc19 - transformers max version → 5.12 **Dev tooling bumps** - ruff bump 0.12.11 → 0.15.18 - mypy 1.17.1 → 2.1.0: enable new defaults (`local_partial_types`, `strict_bytes`); fix/narrow the errors newly surfaced by mypy 2.0 in 4 modules (rather than blanket-suppressing them); remove 2 stale `# type: ignore` comments - pre-commit 4.3.0 → 4.6.0 - sphinx 8.1 → 9.1 + sphinx-rtd-theme 3.0 → 3.1: add `suppress_warnings = ["ref.python"]` to fix cross-reference ambiguity error new in sphinx 9.x - trl fix for newly released 1.7 version **Bug fixes surfaced by the bumps** - sparsity (weight): make the weight mask DTensor-aware under FSDP. The transformers→5.12 bump routes the HF Trainer FSDP optimizer-state save through torch's DTensor-based `get_optimizer_state_dict`, which triggered `aten.mul.Tensor got mixed torch.Tensor and DTensor` in the dynamic `weight` getter. The mask is now distributed to the weight's mesh/placements before masking, cached, and rebuilt only when the sharding changes (invalidated on `set_mask`). Fixes the `llm_sparsity` example test. ### Testing - `pre-commit run --all-files` ✅ (including mypy 2.1.0) - `nox -s docs` ✅ - `tests/unit/torch/sparsity` + `tests/unit/torch/nas` ✅ - `llm_sparsity` GPU example test (FSDP path) verified in CI ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Summary * **Documentation** * Refreshed Docker pre-requisites across examples to recommend updated container image tags (and streamlined some instructions). * **Bug Fixes** * Improved sparse weight mask handling for DTensor/FSDP by aligning and caching distributed masks. * Made TensorRT engine byte retrieval return immutable `bytes`. * Reduced Sphinx cross-reference warnings and tuned Transformers compatibility warning thresholds. * **Tests** * Increased default unit test timeout on Windows runners. * **Chores** * Updated CI workflow container tags and refreshed linting/typing/docs version pins, plus related mypy configuration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
51774473c6 |
using validation keyword both in puzzletron configs and in dataset pr… (#1830)
…eparation script ### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Standardized the validation dataset name from `valid` to `validation` across multiple example configurations and dataset preparation outputs. * Improved compatibility for validation workflows by aligning dataset naming used during training and evaluation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Grzegorz Karch <gkarch@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
d0c01a4e96 |
Add p quantization to our triton fa kernel (#1757)
### What does this PR do? Type of change: new feature - New P_QDQ feature in the Triton FA kernel: fake-quant softmax P before P·V, modes fp8/nvfp4; API attention(..., p_qdq, p_qdq_scale); denominator unquantized, backward is STE. - New quantization/attention/p_qdq.py + quantization/common/fp8_quant.py; reuses nvfp4_quant FP4 rounding. - _QuantAttention: softmax_quantizer→p_bmm_quantizer, dispatches FP8/NVFP4 to the kernel (no kitchen); adds TensorQuantizer.is_fp8/is_nvfp4_dynamic; envelope guards for unsupported cases. - Recipe wildcard + vLLM reload updated for the rename; nvfp4_tensor.py comment typo fixed; ruff ignores + tests added. ### Usage ### Testing added unit tests ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Summary * **New Features** * Added softmax probability quant-dequant (P_QDQ) support via `p_qdq` and `p_qdq_scale`. * When enabled, quantized attention can route through the Triton P_QDQ path. * Added stricter Triton attention “envelope” validation with `validate_triton_attention_envelope`. * **Bug Fixes** * Improved handling of quantizer keys to correctly skip softmax-P (`p_bmm_quantizer`) entries. * **Tests** * Added GPU forward/backward coverage for FP8 (E4M3) and NVFP4 (E2M1), including reference comparisons and invalid-parameter cases. * **Documentation** * Updated attention quantization configs to use `p_bmm` quantizers instead of `softmax_quantizer`. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shiyang Chen <shiychen@nvidia.com> |
||
|
|
aa2a6a1b5d |
Add context-parallel (CP) and data-parallel (DP) support to Megatron calibration, and MMLU (#1804)
### What does this PR do?
Type of change: new feature
Adds **context-parallel (CP)** and **data-parallel (DP)** support to the
shared Megatron-Core inference/calibration utilities so PTQ calibration
and the MMLU sanity check work across these parallelisms (in addition to
the existing TP/PP/SP/EP).
**Context parallelism (CP):**
- **`megatron_calibration` / `megatron_mmlu`** — partition each sequence
across CP ranks (zigzag load-balanced, via `get_batch_on_this_cp_rank`).
MMLU gathers the per-rank logits back to the full sequence for
last-token scoring.
- **`megatron_prefill`** — accepts a CP-partitioned `position_ids`, and
under CP passes `attention_mask=None` so the CP-aware causal attention
builds the mask itself (a local triu mask would be wrong for the
per-rank zigzag chunks). Also wrapped in `torch.no_grad()` (pure
inference; lets MMLU run at larger batch sizes without retaining the
autograd graph).
- **`examples/megatron_bridge/quantize.py`** — new `--cp_size` flag.
**Data parallelism (DP):**
- **`get_dataset_dataloader`** — new `distributed` / `sampler_kwargs` to
shard the dataset across ranks with a `DistributedSampler`.
- **`megatron_calibration`** — shards calibration data across the DP
group (amax is max-reduced across DP inside `mtq` calibration, so the
per-rank shards combine correctly).
- **`megatron_mmlu`** — shards whole batches across DP ranks and
all-reduces the per-subject counts back to full-dataset accuracy.
- DP is **implicit**: `DP size = world_size / (tp * pp * cp)` —
launching with more GPUs than `tp * pp * cp` engages it. (No `--dp_size`
flag.)
RoPE is applied by the model per CP rank, so `position_ids` are only
needed for models with absolute/learned position embeddings.
### Usage
```bash
# Context parallelism = 2
torchrun --nproc_per_node 2 examples/megatron_bridge/quantize.py \
--hf_model_name_or_path Qwen/Qwen3-8B --quant_cfg nvfp4 \
--cp_size 2 --export_megatron_path /tmp/Qwen3-8B-NVFP4-cp2
# Data parallelism = 2 (implicit: tp*pp*cp = 1, 2 GPUs)
torchrun --nproc_per_node 2 examples/megatron_bridge/quantize.py \
--hf_model_name_or_path Qwen/Qwen3-8B --quant_cfg nvfp4 \
--export_megatron_path /tmp/Qwen3-8B-NVFP4-dp2
```
### Testing
Added `cp` and `dp` cases to
`tests/gpu_megatron/torch/utils/plugins/test_utils_megatron.py::test_megatron_generate_and_mmlu`
(Qwen3-0.6B). All four parallelisms pass on 2 GPUs over the full
1430-example MMLU shard set (confirming the DP all-reduce reconstructs
the full count): tp=0.373, pp=0.375, cp=0.371, dp=0.375.
End-to-end PTQ on **Qwen3-8B → NVFP4**, comparing MMLU before and after
PTQ across parallelisms:
| Parallelism | MMLU before PTQ (bf16) | MMLU after PTQ (NVFP4) |
| :--- | :---: | :---: |
| TP=2 | 0.7294 | 0.7058 |
| CP=2 | 0.7292 | 0.7101 |
| DP=2 | 0.7292 | 0.7099 |
> **Common PTQ args used for the runs above:** `--hf_model_name_or_path
Qwen/Qwen3-8B --quant_cfg nvfp4 --seq_length 1024 --calib_num_samples
512 --calib_batch_size 16`, default calibration dataset
(`cnn_nemotron_v2_mix` = cnn_dailymail +
nemotron-post-training-dataset-v2 mix). MMLU evaluated at `fraction=1.0,
batch_size=16` on the full test set. Runs on 2× RTX 6000 Ada — TP=2 uses
`--tp_size 2`, CP=2 uses `--cp_size 2`, DP=2 uses all model-parallel
sizes = 1 (implicit DP over the 2 GPUs).
CP-, DP-, and TP-calibrated models all land within ~0.4% MMLU of each
other both before and after PTQ, confirming CP/DP calibration yields an
equivalently-quantized model.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — all CP/DP logic is gated on
`cp_size > 1` / `dp_size > 1`; non-CP/DP (TP/PP/SP/EP) behavior is
unchanged. `megatron_prefill` gains an optional `position_ids` arg
(defaults to the previous behavior); `get_dataset_dataloader` gains
optional `distributed`/`sampler_kwargs` (default off).
- If you copied code from any other sources or added a new PIP
dependency: N/A
- Did you write any new necessary tests?: ✅ — `cp` and `dp`
parametrizations added to the existing Megatron generate/MMLU test.
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ✅ (run `/claude review`)
### Additional Information
`torch.no_grad()` on `megatron_prefill` is a shared change (also
benefits the calibration / generate / PEFT-test callers) — pure
inference, so no behavioral change beyond lower memory.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added context-parallel (CP) and data-parallel (DP) support across
shared inference, calibration, and MMLU evaluation.
* Introduced CP-aware prefill with per-rank logits gathering for correct
last-token scoring.
* Added `--cp_size` to the quantization example (DP is derived
automatically).
* **Improvements**
* Extended dataset dataloader utilities with optional distributed
sampling controls.
* **Bug Fixes**
* Calibration and evaluation now work when CP is enabled (no longer
restricted to CP=1).
* **Tests**
* Expanded Megatron generate/MMLU coverage to include CP and DP modes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
1766d55a7b |
[6281412] docs: update TensorRT-Edge-LLM CLI commands in torch_onnx example (#1808)
### What does this PR do?
Type of change: documentation
TensorRT-Edge-LLM v0.8.0 consolidated its CLI entry points, leaving the
example commands in `examples/torch_onnx/README.md` referencing tools
that no longer exist (e.g. `tensorrt-edgellm-export-visual`). This
updates the README to the current interface:
- `tensorrt-edgellm-quantize-llm` / `tensorrt-edgellm-quantize-draft` →
`tensorrt-edgellm-quantize {llm,draft}` (subcommands)
- `tensorrt-edgellm-export-llm` / `-export-visual` / `-export-draft` →
unified `tensorrt-edgellm-export` with positional `model` / `output_dir`
args and automatic VLM/audio component detection
- `--is_eagle_base` → `--eagle-base`
- Updated the CLI Tools table and the LLM / VLM / EAGLE examples
accordingly
### Usage
N/A — documentation change.
### Testing
Verified against the live `main` branch of TensorRT-Edge-LLM by running
the actual entry-point code (`python -m
tensorrt_edgellm.scripts.quantize/export`):
- `--help` runs cleanly for `quantize`, `quantize llm`, `quantize
draft`, and `export`; all documented flags (`--model_dir`,
`--output_dir`, `--quantization`, `--base_model_dir`,
`--draft_model_dir`, positional `model`/`output_dir`, `--eagle-base`)
are present.
- Drove the parser with the exact README commands — they parse and
advance into the real quantize/export logic.
- Confirmed the old names are gone: `quantize-llm` subcommand rejected,
`--is_eagle_base` rejected, `scripts.export_visual` module not found.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: N/A (documentation only)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (documentation only)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (minor docs change)
> 🤖 _Generated by Claude (AI agent)._
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Updated TensorRT-Edge-LLM CLI documentation to reflect consolidated
command structure
* Updated command examples for LLM, VLM, and EAGLE speculative decoding
workflows
* Documented new unified CLI interfaces with updated subcommands and
flags
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
b6bf6b7997 |
Update Roadmap Issue link
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
c458ad36f1 |
[2/2] Remove examples/diffusers/eval image-quality evaluation example (#1694)
> **Part 2 of 2** — removal for the 0.46 release. Depends on **Part 1**: #1798 (the 0.45 deprecation). ### What does this PR do? Type of change: Backward breaking change (removal of a deprecated example) Removes the `examples/diffusers/eval` image-quality evaluation example (ImageReward / CLIP-IQA / CLIP metrics) and its references in `examples/diffusers/README.md`. The example was **deprecated in 0.45** (see #1798) and is removed here for **0.46**, per the [Deprecation Policy](https://github.com/NVIDIA/Model-Optimizer#deprecation-policy) (1-release migration before removal). Scope is limited to the diffusers example; the unrelated `examples/llm_ptq` "Evaluate Accuracy" section is untouched. ### Usage N/A — removes example scripts; no library API changes. ### Testing N/A — deletion of example scripts plus documentation cleanup. Verified `examples/diffusers/README.md` has no remaining references to the deleted `eval/` directory. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ❌ — removes the (previously deprecated) `examples/diffusers/eval` example. No public `modelopt` API is affected. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — removal only. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — 0.46 Backward Breaking Changes. - Did you get Claude approval on this PR?: N/A ### Additional Information Paired with #1798 (the 0.45 deprecation). This PR targets **0.46** and should **not** carry the `cherry-pick-0.45.0` label (that belongs on #1798). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Documentation** * Updated the Diffusers “Model Optimizations” README by shortening the overview, removing the “Evaluate Accuracy” guidance (links, inputs, commands, and example results), and refining notes about per-subsection `requirements.txt` usage. * **Breaking Changes / Deprecations** * Updated the 0.46 changelog to reflect that the Diffusers image-quality evaluation example (ImageReward/CLIP-IQA/CLIP) is no longer maintained. * **Chores** * Removed the Diffusers evaluation example, including its evaluation entrypoint, metrics, and shared utilities. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
f83a23cdce |
Deprecate examples/llm_autodeploy (#1796)
### What does this PR do? Type of change: deprecation <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Mark the AutoQuant + TensorRT-LLM AutoDeploy example as deprecated per the deprecation policy: add a deprecation banner to the example README and a note under the 0.45 Deprecations section of the changelog. The example will be removed in a future release; users should use TensorRT-LLM's AutoDeploy directly together with ModelOpt PTQ in examples/llm_ptq. ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!--- Mandatory --> - Did you write any new necessary tests?: N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Deprecations** * The `examples/llm_autodeploy` example is deprecated and will be removed in a future release. Users should migrate to TensorRT-LLM's AutoDeploy directly with ModelOpt PTQ instead. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c6f8f07d8e |
Add support for dLLM encoder-decoder models (DiffusionGemma) [tied-weight PTQ export support ] (#1707)
### What does this PR do? Type of change: new feature Adds end-to-end PTQ + HF-checkpoint export support for block-diffusion encoder-decoder LLMs (e.g. DiffusionGemma) whose encoder/decoder stacks share parameters via HF `_tied_weights_keys`. Six source commits + one test commit + one CHANGELOG entry — purely additive for existing modelopt users (non-tied models see no behavioral change). **Source commits:** 1. **Onboard DiffusionGemma to `hf_ptq.py`** — substring-list additions in `model_type_is_enc_dec` and `MODEL_NAME_TO_TYPE` (so calibration routes through `.generate()`), and a `.sequences` unwrap in the preview decode for `ModelOutput`-returning `.generate()`s. 2. **MoE experts dedup** in `_export_fused_experts` — when two fused-expert modules share their 3-D source params, alias the per-expert packed weight + scales on cache hit so downstream `postprocess_state_dict` dedup catches them. ~42% storage reduction on `nvfp4_experts_only` for tied 26B MoE checkpoints. 3. **Dense Linear dedup** in `_export_quantized_weight` — symmetric to (2) for dense Linears; no-op for `nvfp4_experts_only` (dense early-returns at `QUANTIZATION_NONE`). 4. **Opt-in canonical-side reorder** (`--canonical_tied_naming`, default off) — partitions state_dict so canonical-side tied keys iterate before alias-side, letting first-wins dedup keep canonical names. 5. **`sync_tied_input_amax`** — max-merges per-side `input_quantizer.amax` across tied modules BEFORE export, so single-backbone consumers (vLLM) that load one `input_scale` per parameter don't clip on either side. Extends commits 2+3's alias loops to include `input_scale`. 6. **Default `*self_conditioning*` exclude** — adds the diffusion-model self-conditioning wildcard to `default_disabled_quantizers.yaml`. Companion to PR #1691 which already added the vision-module excludes. **Test commit:** 7. New test fixture (`tests/_test_utils/torch/quantization/tied_modules.py`) with three factories (`make_tied_linear_pair`, `tie_fused_experts_3d_params`, `wrap_in_parent_with_tied_keys`) + 10 unit tests covering commits 2–5 across `tests/unit/torch/export/test_unified_export_hf.py` (new) and `tests/unit/torch/quantization/plugins/test_fused_experts.py` (extended). Pure-Python, CPU-only, ~1s wall total. **Docs commit:** 8. CHANGELOG entry under 0.46 New Features. ### Usage ```python # Standard usage — no API change for non-tied models. Existing recipes still work: python examples/llm_ptq/hf_ptq.py \ --pyt_ckpt_path <ckpt-dir> \ --qformat nvfp4_experts_only \ --calib_size 32 \ --trust_remote_code \ --export_path <export-dir> # For tied-weight models (e.g. DiffusionGemma), opt-in to canonical-side # naming so the exported state_dict uses the canonical (e.g. decoder-side) # names per HF's _tied_weights_keys declaration: --canonical_tied_naming true ``` ### Testing - **10 new unit tests** covering commits 2–5. Pure-Python, CPU-only, ~1s wall total. - **Full local unit suite**: 2588 passed, 17 intentional skips. 9 pre-existing failures observed in unrelated test files (PEFT/LoRA, ONNX, speculative-decoding) — verified pre-existing via static analysis (none import any file modified by this PR). - **End-to-end PTQ export** validated on DiffusionGemma v10 via `nvfp4_experts_only` recipe + `--canonical_tied_naming true` — produces a 2-shard, ~18 GB safetensors checkpoint with decoder-canonical naming. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `--canonical_tied_naming` is opt-in default off; the dedup/alias logic is cache-based on `data_ptr()` and no-ops for non-tied models; `sync_tied_input_amax` no-ops when no two modules share a weight `data_ptr`; the `*self_conditioning*` wildcard is a no-op for models without matching module names. - If you copied code from any other sources or added a new PIP dependency: N/A - Did you write any new necessary tests?: ✅ 10 unit tests across two files; fixture helpers shared via `_test_utils` - Did you update CHANGELOG?: ✅ — new bullet under 0.46 → New Features - Did you get Claude approval on this PR?: pending `/claude review` ### Additional Information Validated locally against DiffusionGemma v10 (`DiffusionGemmaForBlockDiffusion`, `model_type: diffusion_gemma`). Substring patterns are also chosen to match the older `DiffusionGemma4ModelForBlockDiffusion` / `diffusion_gemma4` spelling, so the same code path works on earlier and current model versions. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## New Features - Added tied-weight PTQ and Hugging Face checkpoint export with export-time deduplication for encoder-decoder LLMs (including DiffusionGemma). - Implemented canonical tied key ordering during export and automatic max-merging of shared `input_quantizer.amax` values. - Extended model-type detection and export handling for DiffusionGemma. ## Bug Fixes - Fixed quantized diffusion preview decoding when generation returns a `ModelOutput` wrapper. ## Documentation - Added DiffusionGemma PTQ guidance and new YAML recipes to disable `self_conditioning` quantizers. ## Tests - Added unit tests for tied export behavior, `amax` syncing, and fused-expert tied deduplication. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Juhi Mittal <juhim@nvidia.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
090b1c5114 |
fastgen DMD2: make the Qwen-Image example self-contained on stock nemo_automodel (#1688)
### What does this PR do?
**Type of change:** new example / refactor (example self-containment)
The published `examples/diffusers/fastgen` DMD2 Qwen-Image distillation
example previously relied on
**local, unpublished modifications to the sibling `nemo_automodel`
package** (data path, collate, a
partial-load checkpointer, and Qwen-Image preprocessing). An external
user running the published
Model-Optimizer example against **official stock `nemo_automodel`**
would hit import/attribute errors.
This PR makes the example **self-contained on stock
`nemo_automodel>=0.4.0`** — with the changes kept
**as small as possible**: the team's actual AutoModel delta is only ~430
lines, so wherever the
upstream module is importable, the delta is expressed as a thin subclass
/ small reimplementation rather than a
full-file copy.
- **`fastgen_data/`** — the DMD2 data path:
- `collate_fns.py` — **reuses the upstream `SequentialBucketSampler` and
reimplements the collate**:
it builds the DMD2 batch directly from the vendored dataset's per-item
output (`image_latents` /
`text_embeddings` / `text_embeddings_mask` + an optional broadcast
`negative_text_embeddings` for
CFG) and deliberately does **not** call upstream
`collate_fn_production`, which stacks
model-specific token keys (`clip_tokens` / `t5_tokens`) absent from the
Qwen-Image cache. The
builder loads an optional `negative_prompt_embedding_path`.
- `text_to_image_dataset.py` — a **faithful vendored copy** of the
upstream reader (its
`prompt_embeds_mask` emission is interleaved with cache loading, so
wrapping it would force a
redundant per-item `torch.load`; carried verbatim instead).
- **`fastgen_checkpoint.py`** — `PartialLoadCheckpointer(Checkpointer)`
that overrides only
`load_optimizer` (FSDP2 `DefaultLoadPlanner(allow_partial_load=True)`)
so optimizer resume works
without patching upstream. Injected via an in-place re-bless of
`self.checkpointer` in the
recipe's `load_checkpoint` (model-state load stays strict).
- **`preprocess/`** — Qwen-Image preprocessing
(`preprocessing_multiprocess.py` + `processors/`),
trimmed to the image path (drops the flux/wan/hunyuan processors and the
video base class).
It lives in AutoModel's top-level `tools/` tree, which is **not**
shipped in the pip package, so
it cannot be wrapped and is vendored; `MultiTierBucketCalculator` is
imported from stock upstream.
- **`make_negative_prompt_embedding.py`** — generates the optional CFG
negative-prompt embedding.
- All `configs/*.yaml` target `fastgen_data.build_*` (a test enumerates
every config).
- Licensing: the AutoModel-copied files are NVIDIA-authored Apache-2.0,
so they carry only the
standard NVIDIA SPDX header (managed by the `insert-license` hook) — no
per-file provenance note,
no duplicated license, no pre-commit exclusion, and no separate
`LICENSE` note.
`nemo_automodel[diffusion]` version bound in `requirements.txt`.
The DMD2 math in `modelopt/torch/fastgen/` is **unchanged** — only
example/training-time glue moved.
### Usage
```bash
# Install example deps (stock nemo_automodel) from a source checkout
pip install -r examples/diffusers/fastgen/requirements.txt
# Build the training cache from raw images (Qwen-Image VAE latents + text embeddings)
python examples/diffusers/fastgen/preprocess_qwen_image.py image \
--image_dir <raw images> --output_dir <cache dir> --processor qwen_image \
--caption_format meta_json
# Generate the CFG negative-prompt embedding once
python examples/diffusers/fastgen/make_negative_prompt_embedding.py \
--output <cache dir>/negative_prompt_embedding.pt
# Point the config's data.dataloader.cache_dir + negative_prompt_embedding_path at the cache, then train.
```
See `examples/diffusers/fastgen/README.md` → "Requirements &
self-contained data path".
### Testing
- New `tests/examples/diffusers/fastgen/test_vendored_migration.py`:
environment-independent
invariants (every config targets a vendored builder; no `tools.*`
imports; each former AutoModel
patch is vendored / wrapped / a documented exclusion; the
former-vendored files carry the standard NVIDIA SPDX header, no
provenance note or duplicate license) plus
dependency-guarded structural tests (the collate emits the batch
contract + broadcasts the
negative embedding; the builder accepts
`negative_prompt_embedding_path`; the checkpointer
overrides only `load_optimizer`; the Qwen-Image processor
self-registers).
- **Validated 9/9 against a pure stock `nemo_automodel` 0.4.0 worktree**
(none of the local patches
present) via SLURM — re-run after this slim-down.
- The migrated code path is exercised by a live multi-GPU DMD2 run that
resumed from a checkpoint
through the vendored `PartialLoadCheckpointer`.
- `ruff check` + `ruff format --check` clean on all changed files.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (additive; bundled configs
target the vendored builders)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
(NeMo-AutoModel @ `e42584e3`, Apache-2.0; per review these
NVIDIA-authored files carry the standard NVIDIA SPDX header, no separate
provenance / `LICENSE` note; `nemo_automodel` was already a dependency)
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A (example-only change)
- Did you get Claude approval on this PR?: ❌ (pending — opened as draft)
### Additional Information
Opened as a **draft** pending: OSRB review of the vendored
NeMo-AutoModel (Apache-2.0) code, CI green,
and (optional) a multi-GPU smoke-train + resume on stock upstream.
Vendored from
NVIDIA-NeMo/Automodel at commit `e42584e3`.
### Update (post-review)
- **Mid-run resume data-correctness fix** (`6ffbc52c9`): on resume the
`StatefulDataLoader`'s restored state did not advance past the resume
point, so each window re-served the same data slice and multi-window
(SLURM-windowed) runs under-covered the dataset. Fixed by rebuilding a
fresh loader and skipping the deterministic sampler to the position
implied by `global_step`; added a SLURM-free, GPU-free CPU regression
test (`tests/examples/diffusers/fastgen/test_resume_dataloader.py`).
- **Licensing review** (`be832ae95`): the AutoModel-copied files are
NVIDIA-authored Apache-2.0, so they now carry only the standard NVIDIA
SPDX header — dropped the per-file provenance note, the duplicated
original-license block, the `insert-license` pre-commit exclusion, and
the `LICENSE` note.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added vendored Qwen‑Image preprocessing (multi-process) with processor
registry support.
* Updated DMD2 data loading with a dedicated dataset/collation pipeline,
including negative-prompt embedding/mask handling.
* Improved training resume behavior by rebuilding dataloader state and
making checkpoint optimizer restore tolerant of partial FSDP2 optimizer
shards.
* **Documentation**
* Refreshed the fastgen README and config notes for real-data training;
removed the prior mock-data smoke workflow.
* **Tests**
* Added regression and migration tests covering vendored wiring,
collate/dataloader contracts, processor registration, and
resume/checkpoint behavior.
* **Chores**
* Updated licenses/attribution, vendoring/tooling guards, requirements
pinning, linting configuration, and repository ownership rules.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
87f1a4f496 |
Add Minitron hidden_size pruning support for GatedDeltaNet, MLA, and latent MoE (#1747)
### What does this PR do?
Type of change: New feature
Adds Minitron (`mcore_minitron`) pruning support for Megatron-Core
models with attention and MoE variants that were previously unsupported.
For all of these, only ``hidden_size`` is pruned (alongside the usual
`ffn_hidden_size` / `num_layers` / MoE dimensions); the variant-internal
dimensions are kept:
- **GatedDeltaNet** (linear attention) and **gated attention**
(`attention_output_gate`) — e.g. Qwen3.5 (hybrid GatedDeltaNet +
gated-attention), including MoE variants. Attention / linear-attention
heads are not pruned.
- **Multi-Latent Attention (MLA)** — e.g. DeepSeek. MLA latent ranks are
not pruned.
- **Latent MoE** (`moe_latent_size`) — e.g. Nemotron-3. `hidden_size`
pruning resizes the `fc1`/`fc2` latent projections while the experts
stay in the (static) latent dim.
Introduces a `_DynamicAttention` base class whose policy decides whether
only `hidden_size` is pruned or attention heads too —
`_DynamicSelfAttention`, `_DynamicGatedDeltaNet` and
`_DynamicMLASelfAttention` build on it (extensible for future attention
types). A TODO documents how to also prune `moe_latent_size` as a
follow-up.
Note that the new `megatron.core` imports are not guarded because they
are available in all containers we support (e.g. `nemo:26.02` onwards)
### Usage
```python
import modelopt.torch.prune as mtp
# Qwen3.5 (GatedDeltaNet), DeepSeek (MLA), or Nemotron-3 (latent MoE) GPTModel
model, _ = mtp.prune(
model,
mode="mcore_minitron",
constraints={"export_config": {"hidden_size": 2048, "ffn_hidden_size": 8192}},
dummy_input=None,
config={"forward_loop": forward_loop},
)
```
### Testing
New GPU tests in
`tests/gpu_megatron/.../test_mcore_gpt_minitron_pruning.py` —
`test_mcore_qwen35_gdn_moe_pruning`, `test_mcore_mla_pruning`,
`test_mcore_latent_moe_pruning` — run across available GPUs (covers
pipeline parallel). Existing GPT/MoE prune + parameter-sorting tests
still pass; the file's tests were refactored to share scaffolding
(`_build_and_prune_variant`, `_sort_and_capture`, `_make_forward_loop`,
`_assert_reprune_matches`). `pre-commit` (ruff, mypy) clean.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Expanded Minitron pruning for Megatron-Core to support GatedDeltaNet
(including gated-attention MoE variants) and Multi-Latent Attention
(MLA).
* Added optional latent MoE pruning behavior: pruning targets
hidden-size projections while keeping the MoE latent dimension
unchanged.
* **Documentation**
* Updated the pruning support matrix to clarify attention types that
don’t support `num_attention_heads` pruning.
* **Bug Fixes**
* Improved attention importance/activation capture so pruning hooks only
run when attention-head pruning is actually enabled.
* **Tests**
* Added/expanded GPU coverage for gated-delta-net MoE, MLA, and latent
MoE pruning, with shape/config and inference checks.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
48f8d8984f |
Fix conversation loading logic in UltraChat dataset (#1680)
### What does this PR do?
Type of change: Bug fix.
<!-- Details about the change. -->
Previously, only the first user prompt was extracted from each example,
discarding all subsequent turns. UltraChat stores full multi-turn
conversations in the "messages" field, so switching to that field
preserves the complete dialogue rather than truncating to a single user
message.
### Usage
```python
# Add a code snippet demonstrating how to use this
```
### Testing
python make_dataset.py -f test_cfg.yaml --full
python make_dataset.py -f test_cfg.yaml
```yaml
- name: "ultrachat"
splits:
train_gen: 100
train_sft: 100
```
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Breaking Changes**
* Removed UltraChat support from the example dataset mixer, so UltraChat
splits can no longer be loaded.
* **Documentation**
* Updated the dataset examples README and the example dataset
configuration to remove UltraChat and adjust split settings for other
datasets.
* Updated the speculative decoding fine-tuning documentation to
reference Daring-Anteater instead of UltraChat.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: jzh26 <226629529+jzh26@users.noreply.github.com>
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
d8eb5e5b07 |
fix(puzzletron): correct val_dataset_name from 'valid' to 'validation' (#1765)
The `validate_model_defaults.yaml` file has `val_dataset_name` set to `valid`, but it should be `validation` per the dataset split expected by the tutorial configuration. This causes the Puzzletron tutorial to fail at the subblock scoring step (6/8). ### What does this PR do? Type of change: Bug fix Updates the Puzzletron validation config to use the correct validation dataset split name: - `val_dataset_name: valid` → `val_dataset_name: validation` This fixes the tutorial configuration so the pipeline can proceed correctly during the validation/subblock scoring stage. ### Usage No user-facing API changes. This is a config fix for the existing Puzzletron tutorial. ### Testing Verified the YAML config was updated to use the correct dataset split name. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information This addresses the tutorial failure at step 6/8 caused by the incorrect dataset split name in: `examples/puzzletron/configs/llama-3_1-8B_pruneffn_memory/validate_model_defaults.yaml` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Updated default validation dataset configuration parameter to correct dataset identifier. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Sabari07 <sabursd18@gmail.com> |
||
|
|
9048d13b86 |
[Feat]:Support DPace (#1724)
### What does this PR do? Type of change: New feature Adds the **D-PACE** (Dynamic Position-Aware Cross-Entropy) loss objective for DFlash speculative-decoding training ([arXiv:2605.18810](https://arxiv.org/abs/2605.18810)). It replaces the static exponential position decay with per-position CE weights derived from the draft's own confidence `q_i = exp(-CE_i)`: smoothed `q̃_i = (1-α)q_i + α` (Eq.7) and weighted by the suffix-sum of prefix products `w_j = Σ_{m≥j} ∏_{i≤m} q̃_i` (Eq.8), which directly targets expected accepted block length and shifts signal toward whichever positions currently limit acceptance. Selected via `dflash_loss_objective` — **D-PACE is now the default** (`dpace`); set `dflash_loss_objective: decay` to restore the previous static schedule. Smoothing via `dflash_dpace_alpha` (default 0.5). Weights are detached from the gradient — training-only, ~2.3% overhead, no architecture or inference change. Mutually exclusive with `dflash_loss_decay_factor`. ### Usage ```yaml # DFlash recipe / training config dflash: dflash_loss_objective: dpace # default: decay dflash_dpace_alpha: 0.5 # smoothing in (0, 1]; stable in [0.3, 0.7] ``` ### Testing CPU unit tests in `tests/unit/torch/speculative/plugins/test_hf_dflash.py`: weights match the paper closed form, are detached and non-increasing, the α smoothing floor keeps later weights non-zero, and convert wires/validates the new fields (rejects bad objective and degenerate α). Training validated on Qwen3-8B (curve below). <img width="1803" height="809" alt="image" src="https://github.com/user-attachments/assets/d34dcd76-9e46-4051-94d4-c880b1987965" /> ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ⚠️ Behavior change — D-PACE is now the **default** objective, so DFlash training loss weighting changes unless you set `dflash_loss_objective=decay` (which reproduces the previous static-decay behavior). - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new dependency) - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Reference: D-PACE, [arXiv:2605.18810](https://arxiv.org/abs/2605.18810). See `examples/speculative_decoding/doc/dflash.md` for the math and tuning notes. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added a **D-PACE** training loss objective for DFlash speculative decoding (`dflash_loss_objective: dpace`), configurable via `dflash_dpace_alpha` (default `0.5`). * **Documentation** * Documented D-PACE’s confidence-derived, dynamically weighted per-position loss behavior (training-only) and noted that `dflash_loss_decay_factor` is ignored with D-PACE. * **Bug Fixes** * Updated DFlash loss to reuse the precomputed per-token cross-entropy in the non-KD path. * **Tests** * Added unit tests for D-PACE weight correctness, masking, gradient detachment, monotonicity, smoothing, and config/validation behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
6c32c37671 |
refactor(examples): consolidate vlm_ptq into llm_ptq (#1705)
### What does this PR do? Type of change: refactor / deprecation (examples) `examples/vlm_ptq` was effectively a thin wrapper over `examples/llm_ptq`: its `scripts/huggingface_example.sh` already sourced `llm_ptq/scripts/parser.sh` and called `llm_ptq/hf_ptq.py`, and all the actual VLM logic (vision-tower exclusion, `--calib_with_images`, Nemotron VL calibration, VILA loading, multimodal export) already lives under `llm_ptq`. The wrapper also referenced a `requirements-vila.txt` that did not exist in the repo. This PR makes `llm_ptq` the single source of truth for both LLM and VLM PTQ and deprecates `vlm_ptq`. **`llm_ptq` (canonical):** - Add `--vlm` and `--calib_with_images` flags to `scripts/parser.sh` and `scripts/huggingface_example.sh`. `--vlm` bootstraps VILA dependencies and runs the TensorRT-LLM multimodal quickstart as the deploy smoke test (instead of the text-only `run_tensorrt_llm.py`). - Add `examples/llm_ptq/requirements-vila.txt` (fixes the previously broken reference). - Document the VLM support matrix and the `--vlm` workflow in `README.md`. **`vlm_ptq` (deprecated):** - Replace `scripts/huggingface_example.sh` with a shim that prints a deprecation warning and forwards to the `llm_ptq` script with `--vlm`. - Convert `README.md` into a redirect/migration notice. - Repoint root `README.md` VLM links and add a `CHANGELOG.rst` deprecation entry. ### Usage ```bash cd examples/llm_ptq # VLM PTQ (was: examples/vlm_ptq/scripts/huggingface_example.sh) scripts/huggingface_example.sh --model <hf_model> --quant fp8 --vlm # VLM image-text calibration scripts/huggingface_example.sh --model <hf_model> --quant nvfp4 --vlm --calib_with_images --trust_remote_code ``` ### Testing - `bash -n` syntax check on the modified `parser.sh`, `llm_ptq` script, and the `vlm_ptq` shim. - `pre-commit run --files <changed files>` passes. - The existing VLM example test (`tests/examples/vlm_ptq/test_qwen_vl.py` via `run_vlm_ptq_command`) still exercises the path end-to-end through the deprecation shim, which forwards to the consolidated `llm_ptq` script. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (old `vlm_ptq` entry point still works via a forwarding shim) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (existing VLM test still covers the consolidated path) - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Follow-up (later release): remove the `examples/vlm_ptq` directory and its CI matrix entry once external references have migrated. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added VLM quantization support to `examples/llm_ptq` via a `--vlm` flag. * Enabled image-text pair calibration with `--calib_with_images`, including VLM multimodal smoke-test coverage. * **Deprecations** * `examples/vlm_ptq` is deprecated; it now forwards to the `examples/llm_ptq --vlm` flow with a warning. * VILA/NVILA VLM support was removed from `examples/llm_ptq` due to a model dependency compatibility conflict. * **Documentation** * Updated READMEs and the model support matrix with VLM quantization behavior and export limitations. * **Tests / CI** * Updated VLM PTQ tests and CI workflow matrices to stop running the deprecated `vlm_ptq` example. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
977d34dc3c |
[Fix](nvbug6304585): specdec README online base-model example should use Instruct model (#1755)
### What does this PR do?
Type of change: Bug fix
The **Training Draft Model with Online/Offline Base Model** examples in
`examples/speculative_decoding/README.md` used
`meta-llama/Llama-3.2-1B`, a
base / pretrained checkpoint that ships **no chat template**. The online
EAGLE3
flow tokenizes conversations through
`tokenizer.apply_chat_template(...)`, so
the data collator fails fast at startup:
```
ValueError: No valid chat template!
```
This PR:
- Switches both README example commands (online and offline) to
`meta-llama/Llama-3.2-1B-Instruct`, which carries a chat template.
- Makes the collator error message in
`modelopt/torch/utils/plugins/transformers_dataset.py` actionable — it
now
explains the cause (base checkpoints have no chat template) and points
users
at an Instruct model or a custom `chat_template`.
### Usage
```bash
./launch_train.sh \
--config ../../modelopt_recipes/general/speculative_decoding/eagle3.yaml \
model.model_name_or_path=meta-llama/Llama-3.2-1B-Instruct \
data.data_path=input_conversations/train.jsonl \
training.output_dir=ckpts/llama-3.2-1b-online
```
### Testing
- Reproduced the original `No valid chat template!` failure with the
base
`Llama-3.2-1B` and confirmed the Instruct variant carries a chat
template.
- Verified the new error message renders correctly when a template is
missing.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌
### Additional Information
Fixes nvbug 6304585: https://nvbugspro.nvidia.com/bug/6304585
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **Documentation**
* Updated speculative decoding example training commands to reference
the Llama-3.2-1B-Instruct model.
* **Bug Fixes**
* Enhanced error message when chat template configuration is missing,
providing actionable guidance for resolution.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
7545aef726 |
Mitigate CVE-2026-4372 transformers kernels RCE exposure (#1746)
### What does this PR do? Type of change: Bug fix (security) [CVE-2026-4372](https://nvd.nist.gov/vuln/detail/CVE-2026-4372): `transformers` versions `4.56.0`–`5.2.x` allow remote code execution when loading an untrusted model — a crafted `_attn_implementation_internal` field in `config.json` triggers an implicit kernel download/import from the Hub on a routine `from_pretrained()` call (bypassing `trust_remote_code=False`). It only triggers when the **optional `kernels` package** is installed, and is fixed in `transformers>=5.3`. ModelOpt's own code does not use the vulnerable path, and ModelOpt's core install never pulls in `kernels`. The only place that does is the `examples/gpt-oss` example (it needs the Hub MXFP4 kernels). So: - **Pin `transformers>=5.3` in `examples/gpt-oss/requirements.txt`** — the one environment that installs `kernels`, closing the CVE where both preconditions can coincide. - **Add a runtime warning in `modelopt.torch`**, gated on `kernels` being importable, for affected `transformers<5.3`. The gate keeps it quiet for the majority of users (who don't have `kernels`) and only nudges genuinely-exposed setups, wherever they obtained `kernels`. The global `transformers>=4.56,<5.10` constraint is intentionally **not** tightened, to avoid forcing all users (including the `transformers` 4.x line, which has no patched release) off their current version. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A <!-- import-time warning gated on optional dep; no new test infra added --> - Did you update Changelog?: N/A <!-- intentionally omitted --> - Did you get Claude approval on this PR?: TODO ### Additional Information ModelOpt source uses no vulnerable code path (no `kernels` / `hub_kernels` import; only sets `_attn_implementation` to safe local values). 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a21197c277 |
Remove unsafe torch.load from examples/diffusers/fastgen (#1740)
Follow-up to #1326 - Remove unsafe `torch.load(..., weights_only=False)` in `examples/diffusers/fastgen` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated checkpoint loading mechanisms in FastGen examples to improve compatibility and reliability during model restoration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
e6790ef7b4 |
[Examples]: GPT-oss, Qwen3Moe streaming specdec example (#1692)
### What does this PR do?
Type of change: new example
Adds **streaming speculative-decoding examples (EAGLE3 + DFlash)** for
**gpt-oss-20b** and **Qwen3-30B-A3B** to the ModelOpt launcher,
mirroring the existing Qwen3-8B/Kimi examples.
- New yamls:
`tools/launcher/examples/{openai/gpt-oss-20b,Qwen/Qwen3-30B-A3B}/hf_streaming_{eagle3,dflash}_multi_node.yaml`,
plus gpt-oss `chat_template_train.jinja` (generation-tagged, for
`answer_only_loss`).
- `eagle_utils.py`: the streaming path now installs a custom
`data.chat_template` on the tokenizer (the online path already did) —
needed for the tagged template.
### Usage
```bash
cd tools/launcher
export SLURM_HOST=... SLURM_ACCOUNT=... SLURM_HF_LOCAL=... SLURM_JOB_DIR=...
uv run launch.py --yaml examples/openai/gpt-oss-20b/hf_streaming_eagle3_multi_node.yaml --yes
```
### Testing
Pipeline sanity test on **unsynthesized** data (daring-anteater), 1×
H100-80GB, 12k steps. All four train and pass the vLLM acceptance-length
eval:
| Model | Method | Train speed | vLLM AL |
|---|---|---|---|
| Qwen3-30B-A3B | EAGLE3 | 7.12 it/s | **1.74** |
| Qwen3-30B-A3B | DFlash | 2.31 it/s | 1.29 |
| gpt-oss-20b | EAGLE3 | 5.07 it/s | 1.19 |
| gpt-oss-20b | DFlash | 2.01 it/s | 1.14 |
<img width="1300" height="780" alt="image"
src="https://github.com/user-attachments/assets/2be22562-5b77-4a9f-9dc0-6f936a059736"
/>
> Sanity test only, not a quality run. gpt-oss AL is low because it is a
reasoning model (CoT at inference) while daring-anteater has no
reasoning traces and `answer_only_loss` masks all but the final content
— quality runs need synthesized/reasoning data.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added multi-node speculative decoding pipeline configurations for
Qwen3-30B-A3B and gpt-oss-20b with DFlash and EAGLE3 support.
* Introduced chat template training support for improved model
instruction formatting.
* **Enhancements**
* Increased benchmark concurrency from 1 to 32 across Qwen3-8B
configurations for more realistic performance evaluation.
* Extended training runs from 500 to 2000 steps for Kimi-K2.5 models.
* Improved chat template handling in speculative decoding workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
e004d8d90e |
DFlash speculative decoding for MiniMax-M2.7 (FSDP2): auto mask-token, FSDP2 resume fixes, per-checkpoint draft export (#1621)
## What
Brings up DFlash block-diffusion speculative decoding for large MoE
targets (MiniMax-M2.7, 229B) trained under accelerate FSDP2, and fixes
the regressions that broke checkpoint resume and per-checkpoint draft
export.
## Commits
- **auto-add mask token for DFlash** when the tokenizer lacks one
(resize embeddings, restore dtype).
- **requeue support** in `build_slurm_executor` + **FSDP2
cpu_ram_efficient_loading** for 229B on multi-node.
- **FSDP2 buffer patch** (`fsdp2_buffer_patch.py`): handle non-DTensor
buffers in `fsdp2_load_full_state_dict`, broadcast dtype codes from rank
0, and an FSDP2-safe `clip_grad_norm_`. Required because MiniMax-M2.7
pins transformers 4.57.x (no native `ParallelismConfig`).
- **dtype fix**: use the broadcast dtype (rank 0) rather than the local
meta-device param dtype, so non-leader ranks don't cast bf16 back to
fp32 on resume.
- **restore `DFlashExportCallback`** (this PR's headline): the
Pydantic-recipe refactor (
|
||
|
|
fa94e6ea4d |
[6309094] updated readme to clarify simple_qat_train.py usecase (#1742)
### What does this PR do? Type of change: documentation Added clarification about simjple_qat_train.py which is a demonstration of QAT flow and not meant for multi-GPU training ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests? N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Expanded the QAT/QAD README note section with clearer, end-to-end minimal demo instructions, highlighting a single-GPU quantize+train+save workflow via a dedicated script. * Added guidance for multi-GPU training using `accelerate launch`, with a pointer to the appropriate training entry point. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> |
||
|
|
55a2101e2f |
Update Nemotron-3 Pruning, Distillation and PTQ results based on new shared calibration loop with seq packing and add tool-calling eval fix (#1660)
### What does this PR do? Type of change: documentation + minor example-script tweaks Follow-up to #1601. Originally scoped to add **NVFP4 + QAD**, this PR was **repurposed** to refresh the [Nemotron-3-Nano-30B-A3B-BF16 tutorial](examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md) results using the **new shared calibration loop (sequence packing)** and to **fix tool calling in evaluation**. - Refreshed the prune → distill → eval → **FP8** results (accuracy + vLLM throughput tables) with the new calibration loop. - **Tool-calling eval fix** (`nemo_evaluator.yaml`): GPQA and AIME now run the Python sandbox tool. The tutorial reports both **with-tools** and **no-tools** GPQA/AIME and shows `mean ± std_dev`. - Script tweaks: `quantize.py` calibration now uses sequence packing (`pack=True`) which leads to slight improvement in PTQ; `prune_minitron.py` defaults `inference_batch_size` to `calib_batch_size`. ### Testing Documentation + small example-script changes; tutorial relative links resolve and the results tables / figure were verified consistent. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: Yes - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ (will run `/claude review`) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the main guide and evaluator instructions for prune + distill + FP8/NVFP4 quantization, including refreshed vLLM deployment tips, benchmark/noise presentation, and long-context tool-calling attribution notes. * Refreshed README technique examples/links, reordered the model support matrix rows, and improved pruning overview/support-matrix text. * **Changes to Examples** * NAS pruning now documents higher GPU memory usage vs manual pruning; pruning batching defaults were improved. * Quantization PTQ calibration uses packed document packing; quantized checkpoint export messaging was streamlined. * Updated pruning/distillation/quantization tutorial guidance, metrics/tables, command parameters, and evaluator YAML settings (KV-cache dtype, generation defaults, task behavior). <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
4be7c7f8d8 |
fix memory leak issue during puzzletron scoring, #1681 (#1729)
fixes the oom (cpu ram) issue (reported in #1681) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Optimized memory management during model validation operations. Explicit resource cleanup procedures are now performed after each solution validation, preventing memory accumulation and eliminating out-of-memory errors during extended validation workflows. * **Configuration** * Updated default validation dataset configuration setting. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
95d4e12c70 |
fix(puzzletron): use prebuilt KD dataset to avoid 136GB download (#1726)
Fixes #1658 ### What does this PR do? Type of change: Bug fix, documentation This PR updates the Puzzletron dataset preparation flow to use the already published prebuilt dataset `nvidia/Puzzle-KD-Nemotron-Post-Training-Dataset-v2` by default, avoiding the need to download the full raw `nvidia/Nemotron-Post-Training-Dataset-v2` dataset (~136 GB) just to filter it down to the same ~2.6 GB result. Changes included: - Add `PREBUILT_KD_DATASET` constant in `prepare_dataset.py` - Short-circuit dataset preparation when `dataset_name` matches the prebuilt dataset, loading it directly and skipping the download + filtering pipeline - Update 8 Puzzletron example configs to use the prebuilt dataset path by default - Update the Puzzletron README to document the default ~3 GB path and clarify that the raw ~136 GB path is still available if users want to reproduce preprocessing ### Usage Default lightweight path: ```bash python -m modelopt.torch.puzzletron.dataset.prepare_dataset \ --dataset_name nvidia/Puzzle-KD-Nemotron-Post-Training-Dataset-v2 \ --output_dir path/to/Puzzle-KD-Nemotron-Post-Training-Dataset-v2 ``` Raw dataset path (existing behavior, still supported): ```bash python -m modelopt.torch.puzzletron.dataset.prepare_dataset \ --dataset_name nvidia/Nemotron-Post-Training-Dataset-v2 \ --output_dir path/to/Nemotron-Post-Training-Dataset-v2 ``` ### Testing - Ran `pre-commit run --all-files` - Most hooks passed successfully - Local pre-commit `mypy` reported unrelated existing errors in: - `modelopt/torch/opt/config_loader.py` - `modelopt/recipe/loader.py` - Verified this change separately with a local mock-based test: - prebuilt dataset path correctly loads and saves directly - original raw dataset path remains untouched ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information This change preserves the original raw-dataset workflow for users who explicitly want to regenerate the filtered dataset from scratch, while making the default example flow much lighter and easier to use. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Documentation** * Updated setup instructions to use a prebuilt, optimized dataset by default, simplifying the model compression workflow. * **Chores** * Updated model compression configurations across multiple examples to use the prebuilt dataset. * Enhanced dataset preparation to support prebuilt dataset handling for more efficient setup. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Sabari07 <sabursd18@gmail.com> |
||
|
|
2640551515 |
feat: Layerwise calibration: nested config + QDQ-from-prev-layer flag + checkpoint I/O knobs (#1571)
### What does this PR do?
Type of change: new feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
Groups all layerwise-calibration options under a nested
`LayerwiseConfig` and adds three new behavior knobs to it. All changes
are backward compatible.
#### 1. Nested `layerwise` config
`QuantizeAlgorithmConfig.layerwise` changes from `bool` to a Pydantic
submodel:
```python
class LayerwiseConfig(ModeloptBaseConfig):
enable: bool = False
get_qdq_activations_from_prev_layer: bool = False
checkpoint_dir: str | None = None
save_every: int = 1
save_quantizers_only: bool = False
```
Backward compatibility:
- `layerwise: True/False` still accepted (emits `DeprecationWarning`).
- Flat `layerwise_checkpoint_dir` silently migrated into
`layerwise.checkpoint_dir`.
- Legacy `use_sequential` alias preserved (and resolved during flat-key
migration so it can't be dropped).
- Conflicting flat+nested `checkpoint_dir` values raise.
- All 7 shipped PTQ recipes (`modelopt_recipes/general/ptq/*.yaml`,
`huggingface/qwen3_5*/ptq/*.yaml`) migrated to the
canonical nested shape — no semantic change.
#### 2. `get_qdq_activations_from_prev_layer` — correct GPTQ vs
max-calib semantics
Controls what layer N's calibration sees:
- **True** (GPTQ default): activations carry the quantize-dequantize
error of layers 0..N-1 — GPTQ's Hessian-compensation
goal.
- **False** (max/mse/local_hessian default): full-precision activations,
matching the non-layerwise pass exactly.
The False branch wraps the next-layer input capture forward with the
existing `set_quantizer_by_cfg_context` deny-all idiom
(`{"quantizer_name": "*", "enable": False}`).
GPTQ's per-algorithm default is enforced via
`@model_validator(mode="after")` that reads
`LayerwiseConfig.model_fields_set` —
works for every input shape (empty constructor, bool, partial dict, full
dict) and lets explicit user values override.
#### 3. `save_every` — gate the large activation-cache writes
`save_every: int = 1` (`ge=1`). With N > 1, the per-layer
`next_inputs.pt` (cached activation tensors, the largest checkpoint
artifact for most models) is only written for the boundary layer of each
N-layer window. Per-layer
weight/quantizer/output_meta files are still written every layer (resume
needs them to replay skip layers correctly).
Interrupting mid-window re-calibrates that window on resume.
#### **UPDATE: drop this in current PR and moved to
https://github.com/NVIDIA/Model-Optimizer/pull/1640**
#### 4. `save_quantizers_only` — algorithm-aware weight-blob skipping
`save_quantizers_only: bool = False`. When True, skip `weights.pt`
entirely and persist just the per-quantizer `state_dict`
slice (carries `_amax`) to a new `quantizer_buffers.pt`. On resume,
`full_restore` reloads only the quantizer slice and
trusts that algorithm semantics didn't mutate `layer.weight`.
Safety is enforced by a **whitelist**: `_supports_save_quantizers_only:
ClassVar[bool] = False` on `QuantizeAlgorithmConfig`,
overridden to `True` only on `MaxCalibConfig`, `MseCalibConfig`,
`LocalHessianCalibConfig` (audited — these only touch
`_amax`). Weight-mutating algorithms (GPTQ folds Hessian updates,
AWQ/SmoothQuant fold pre-quant scales) reject the flag at
config-construction time so in-place weight updates can't be silently
lost on resume.
### Usage
```python
import modelopt.torch.quantization as mtq
# GPTQ — `get_qdq_activations_from_prev_layer` defaults to True (Hessian
semantics).
# save_every reduces activation-cache I/O.
mtq.quantize(
model,
{
"quant_cfg": [...],
"algorithm": {
"method": "gptq",
"layerwise": {
"enable": True,
"checkpoint_dir": "/path/to/ckpts",
"save_every": 4,
},
},
},
forward_loop=forward_loop,
)
# Max-calibration — `get_qdq_activations_from_prev_layer` defaults to
False (FP from prior layers).
# save_quantizers_only skips the weights blob since max only updates
_amax.
mtq.quantize(
model,
{
"quant_cfg": [...],
"algorithm": {
"method": "max",
"layerwise": {
"enable": True,
"checkpoint_dir": "/path/to/ckpts",
"save_quantizers_only": True,
},
},
},
forward_loop=forward_loop,
)
```
### Testing
New / updated unit tests in `tests/unit/torch/quantization/`:
- **`test_config_validation.py`** — `TestLayerwiseNestedConfig` covers
nested-form acceptance, bool-form
`DeprecationWarning`, flat `layerwise_checkpoint_dir` migration,
conflicting flat+nested checkpoint_dir, `use_sequential`
alias survival under migration, per-algorithm qdq defaults (parametrized
Max/GPTQ), `save_every` `ge=1` validation, and the
`save_quantizers_only` whitelist — parametrized rejection on `[GPTQ,
AWQLite, SmoothQuant]` and acceptance on `[Max, Mse,
LocalHessian]`.
- **`test_layerwise_calibrate.py`** —
- `test_layerwise_no_qdq_matches_sequential_amax` — behavioral
equivalence: layerwise + `qdq=False` produces the same
per-quantizer `_amax` as the non-layerwise (sequential) max-calibration
flow (verified via `torch.testing.assert_close`).
-
`test_layerwise_save_every_writes_next_inputs_only_at_window_boundaries`
— window-save layout (all layer dirs present,
`next_inputs.pt` only at boundaries).
- `test_layerwise_save_quantizers_only_resume_matches_one_shot_amax` —
end-to-end resume: full run → manifest rewound →
fresh model resumes → final `_amax` matches the one-shot baseline; also
pins the on-disk shape (no `weights.pt`,
`quantizer_buffers.pt` present).
## End-to-end correctness verification
Ran 4 PTQ jobs on **Qwen3-8B** with NVFP4 W4A16 quant_cfg and
`--calib_size 16`, one GPU
each on 4 GPUs
| Run | Algorithm | Layerwise config | Purpose |
|---|---|---|---|
| **A** | GPTQ | `enable=true, get_qdq_activations_from_prev_layer=true,
save_every=5` | New nested form + the new`save_every` knob |
| **B** | MSE | `enable=true, get_qdq_activations_from_prev_layer=false,
save_quantizers_only=true` | New nested form + the new
`save_quantizers_only` knob |
| **C** | GPTQ | Legacy flat form: `layerwise: true,
layerwise_checkpoint_dir: ...` | Backward-compat baseline for A
(exercises the migration validator) |
| **D** | MSE | `enable=false` | Non-layerwise baseline for B |
Across two pairwise comparisons (A vs C ; B vs D), all 905 tensors in
the exported HF checkpoints are bit-identical with hf_quant_config.json
matching exactly — confirming the new layerwise knobs preserve
correctness and the flat-form backward-compat path is intact.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g.
avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ — `layerwise: True/False` still
accepted with a `DeprecationWarning`; flat
`layerwise_checkpoint_dir` silently migrated; `use_sequential` alias
preserved. The two new knobs default to no-op behavior
(`save_every=1`, `save_quantizers_only=False`).
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
— no new dependencies.
- Did you write any new necessary tests?: ✅ — see Testing section.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
ready for review.
- Did you get Claude approval on this PR?: ✅ — `/claude review`
consulted iteratively; review findings (GPTQ default
survival, save_quantizers_only whitelist scope, docstring accuracy, dead
`layer` param) addressed in-PR.
### Additional Information
Notes on design choices that came out of internal review:
- GPTQ's `qdq=True` default uses a `model_validator(mode="after")` +
`model_fields_set` check rather than a `default_factory`
— a `default_factory` is only fired when the field is absent, so any
user-supplied dict (the natural way to enable
layerwise) would silently lose the GPTQ default.
- `save_quantizers_only` is enforced as a whitelist
(`_supports_save_quantizers_only`) rather than a per-algorithm
blacklist,
which keeps future weight-mutating algorithms safe by default.
- `set_quantizer_by_cfg_context` is reused for the qdq-disable scope
instead of a bespoke helper, keeping `model_calib.py`
aligned with the existing "deny-all" idiom documented at
`conversion.py:240`.
- Pre-validation recipe helpers in `examples/llm_ptq/example_utils.py`
(`needs_checkpoint_path_update` /
`resolve_checkpoint_dir`) accept both flat and nested shapes since they
run before Pydantic validation;
`resolve_checkpoint_dir` now also returns the resolved path so the
caller doesn't re-derive it.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Layerwise calibration now uses a nested config (enable,
get_qdq_activations_from_prev_layer, checkpoint_dir, save_every,
save_quantizers_only); checkpointing supports quantizer-only saves and
windowed commits with resume support.
* **Behavior**
* GPTQ calibration defaults get_qdq_activations_from_prev_layer=True
when unspecified; save_every controls when next-inputs/manifests are
written and must be positive.
* **Deprecations**
* Legacy flat-style layerwise keys are still accepted but emit
DeprecationWarning and are auto-migrated (conflicts detected).
* **Examples/UX**
* Tools auto-detect legacy vs nested checkpoint layouts and report the
resolved path.
* **Tests**
* Expanded coverage for nested configs, validation, checkpoint/resume
semantics, and windowed saves.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com>
Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
d26c8af002 |
Fix DeepSeek V3 ptq.py inference-repo path resolution (nvbug 6311147) (#1702)
### What does this PR do? Type of change: Bug fix Fixes nvbug **6311147** (OMNIML-5103). `examples/deepseek/deepseek_v3/ptq.py` resolved the cloned DeepSeek-V3 / DeepSeek-V3.2-Exp inference repos relative to its own directory (`deepseek_v3/`) via `Path(__file__).resolve().parent`. But the [README](https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/deepseek) clones those repos into the parent `examples/deepseek/` directory and runs the script from there, so the lookup landed one level too deep and raised `ValueError: DeepSeek-V3 or DeepSeek-V3.2-Exp not found` (the error message also printed the wrong directory). The fix resolves from `parent.parent` via a single `DEEPSEEK_DIR` base shared by both repo paths and the error message. ### Usage ```bash # Run from examples/deepseek/ as documented in the README, after cloning # DeepSeek-V3 (or DeepSeek-V3.2-Exp) into that directory: torchrun --nproc-per-node 8 --master_port=12346 deepseek_v3/ptq.py \ --model_path $DS_CKPT \ --config DeepSeek-V3/inference/configs/config_671B.json \ --quant_cfg NVFP4_DEFAULT_CFG \ --output_path $FP4_QUANT_PATH ``` ### Testing - Confirmed against the repro path: with the file at `examples/deepseek/deepseek_v3/ptq.py` and the repos cloned into `examples/deepseek/`, `Path(__file__).resolve().parent.parent` now points at `examples/deepseek/` so `DeepSeek-V3/inference` resolves correctly. - Verified the sibling `examples/deepseek/deepseek_v4/` does not share the bug (it takes an explicit `--dsv4_inference_dir` argument instead). - `pre-commit` clean. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (one-line path fix in an example script that requires the DeepSeek repos + multi-GPU checkpoint to exercise) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (bug is in a 0.45-cycle example, not a regression from a released version) - Did you get Claude approval on this PR?: ❌ (not yet run) ### Additional Information nvbug 6311147 / OMNIML-5103. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved path resolution in the example script to more reliably locate the required inference repository. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
2201edebd9 |
Fix Qwen AutoQuant disabled layer test (#1698)
### What does this PR do? Type of change: Bug fix Fixes the `llm_ptq` example-test failure on current `main`: ```text tests/examples/llm_ptq/test_hf_ptq_args.py::test_qwen_autoquant_disabled_layers_are_scoped_to_qwen_models ``` `*linear_attn.in_proj_a*` and `*linear_attn.in_proj_b*` were promoted into the global default disabled quantizer list, so they are no longer Qwen-only AutoQuant exclusions. This PR removes the redundant entries from the example-specific Qwen AutoQuant exclusion tuple and narrows the test to the remaining Qwen-only pattern, `*shared_expert_gate*`. ### Usage N/A ### Testing - Reproduced the failure on `origin/main` (`cc17f2c45`) with: - `python -m pytest tests/examples/llm_ptq/test_hf_ptq_args.py::test_qwen_autoquant_disabled_layers_are_scoped_to_qwen_models -vv` - Verified the fix with: - `python -m pytest tests/examples/llm_ptq/test_hf_ptq_args.py::test_qwen_autoquant_disabled_layers_are_scoped_to_qwen_models tests/unit/recipe/test_presets.py::test_w4a16_nvfp4_preset_disables_vllm_marlin_incompatible_projections tests/unit/recipe/test_loader.py::test_nvfp4_weight_only_recipe_disables_vllm_marlin_incompatible_projections -vv` - `python -m py_compile examples/llm_ptq/example_utils.py tests/examples/llm_ptq/test_hf_ptq_args.py` - `git diff --check` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information Observed while checking the failing `trtllm-pr (llm_ptq) / run-test` job on PR #1697. The failure reproduces on `main`, so this PR is independent of #1697. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Adjusted quantization exclusion configuration for the Qwen model family to narrow which layers are excluded. * **Tests** * Updated test expectations to align with the revised quantization exclusion behavior for Qwen models. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
60b1af5fb2 |
Fix GPT-OSS MXFP4->NVFP4 PTQ load, export, and cast (nvbug 6295279, 6295242) (#1678)
### What does this PR do? Type of change: Bug fix Fixes the GPT-OSS MXFP4 → NVFP4 PTQ path (`examples/llm_ptq/hf_ptq.py` with `--cast_mxfp4_to_nvfp4`), which failed in three independent ways. The documented command now runs end-to-end and produces a bit-exact (100% lossless) NVFP4 checkpoint. Addresses **nvbug 6295279** (OMNIML-5046) and **nvbug 6295242** (OMNIML-5045). 1. **nvbug 6295242 — CUDA illegal memory access on load.** GPT-OSS ships native MXFP4 weights that Transformers dequantizes to BF16; the threaded weight loader trips an illegal-memory access when `device_map="auto"` shards the dequant across **multiple GPUs**. The missing optional `kernels` package only *forces* the dequant path — it is not the root cause. `get_model` now detects MXFP4 checkpoints and loads them with `Mxfp4Config(dequantize=True)` on a **sequential** device map so the dequant stays on a single device. `kernels` is no longer required. 2. **nvbug 6295279 #1 — `NotImplementedError: Mxfp4GptOssExperts` during unified HF export.** Forcing `dequantize=True` yields plain `GptOssExperts` (even when `kernels` is installed), which ModelOpt wraps and exports normally. 3. **nvbug 6295279 #2 — `FileNotFoundError` in the cast step.** `--cast_mxfp4_to_nvfp4` treated `--pyt_ckpt_path` as a local dir; a HF Hub ID now resolves to its cached snapshot dir via `_resolve_model_path`. Also fixes a **static-block NVFP4 regression** (surfaced by the cast's `force_weight_quantizers_static`, introduced by #1560's now-unconditional `weight_only_quantize`): `_QuantGptOssExperts` / `_QuantLlama4TextExperts` quantize their expert weights transposed in the forward (`_transposed_quantize`), but the inherited `iter_weights_for_calibration` fed the non-transposed weight, locking a mismatched block-quant `_original_shape` and raising `ValueError: Input shape has changed`. The override now calibrates on the transposed view, matching both the forward and the export's `_amax` orientation. ### Why this regressed (it worked when the cast was added) `get_model` never had explicit handling for a *natively pre-quantized MXFP4* checkpoint — GPT-OSS fell through the generic *unquantized-checkpoint* branch and relied on Transformers' **implicit** MXFP4 behavior, which is fragile across three axes. The cast was originally validated (#1372, 2026-05-01) in the "lucky" quadrant of each: - **GPU count:** `device_map="auto"` on a single GPU never shards, so the dequant stays on one device. On multiple GPUs `auto` balances the model and shards the MXFP4→BF16 dequant across devices → CUDA illegal-memory crash (6295242). - **`kernels` presence:** without `kernels`, Transformers auto-dequantizes to BF16 `GptOssExperts` (exportable). With `kernels` installed it keeps the packed `Mxfp4GptOssExperts` kernel path → export `NotImplementedError` (6295279 #1). - **Transformers version:** the kernel-backed experts wrapper and the threaded multi-GPU weight loader are newer-Transformers behavior (env here is 5.5.4). Earlier versions simply dequantized MXFP4 → BF16, which is what the old generic path happened to need. The QA env sat in the *breaking* quadrant (multi-GPU and/or `kernels` present, newer Transformers), so the implicit path failed. The new branch makes both decisions explicit and deterministic (`dequantize=True` + single-device load), regardless of environment — mirroring the existing `has_pack_quantized_config` branch for compressed-tensors checkpoints. The fourth issue (static-block `Input shape has changed`) is a separate regression: it was introduced by **#1560 (2026-06-02, "Make sure all weight quantizers have `_amax`")**, a month *after* the cast landed. #1560 made `weight_only_quantize` unconditional in `max_calibrate`; previously it ran only when no calibration `forward_loop` was supplied, and the cast always supplies one — so the non-transposed weight-quantizer call simply never happened before. The conflict only appears at the intersection of (a) transposed-quantize experts (GPT-OSS/Llama4), (b) static-block NVFP4 — which `--cast_mxfp4_to_nvfp4` forces via `force_weight_quantizers_static` — and (c) #1560. CI's GPT-OSS NVFP4 coverage uses the *dynamic*-block path, which never locks the block shape, so #1560 looked safe. ### Usage ```bash python hf_ptq.py \ --pyt_ckpt_path openai/gpt-oss-20b \ --qformat nvfp4_mlp_only \ --cast_mxfp4_to_nvfp4 \ --export_path ./gpt-oss-20b-nvfp4 ``` ### Testing - Ran the documented command end-to-end on 2xB200 (`openai/gpt-oss-20b`): cast overrode **48/48** expert weight quantizers, **100% lossless** layers/blocks, exported a valid packed-NVFP4 HF checkpoint (uint8 weights + FP8 per-block `weight_scale` + per-tensor `weight_scale_2` + `hf_quant_config.json`). - Verified plain `--qformat nvfp4_mlp_only` (no cast) still works end-to-end. - **Independently verified the export is bit-exact:** dequantized the exported NVFP4 weights (ModelOpt's E2M1 LUT + pack layout) and compared against Transformers' canonical MXFP4→BF16 dequant (`Mxfp4Config(dequantize=True)`) over all 24 layers × both expert weights — `max_abs_err = 0`, 100% bitwise-equal in bf16. So `dequant(exported NVFP4) == dequant(original MXFP4)` exactly. - New unit tests: `test_get_original_hf_quant_method_*` (load detection) and `test_gpt_oss_experts_iter_weights_for_calibration_transposed` (the transpose regression). Existing `test_cast_mxfp4_to_nvfp4.py` (8 tests) still pass. `pre-commit` clean. **Known limitation:** verified for gpt-oss-20b (fits one GPU). gpt-oss-120b dequantized does not fit a single GPU, so `sequential` would still span GPUs — that case would need a CPU-dequant-then-dispatch path and is left as a follow-up. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ (0.45 Bug Fixes) - Did you get Claude approval on this PR?: ❌ (not yet run) ### Additional Information nvbug 6295279, nvbug 6295242 / OMNIML-5046, OMNIML-5045. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Prevented CUDA illegal-memory access during MXFP4→NVFP4 casting. * Fixed expert-weight calibration orientation to avoid shape mismatches. * **New Features** * Support loading native MXFP4 checkpoints with automatic dequantization. * Resolve remote model identifiers to local checkpoints when casting MXFP4→NVFP4, improving reliability. * **Tests** * Added unit and GPU regression tests covering quant-method detection, casting, and expert-weight calibration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
dd49a460c8 |
[6294905] Fix --quant_cfg CLI parsing type in transformer trainer (#1676)
### What does this PR do? Type of change: Bug fix Fix `--quant_cfg` CLI parsing by typing `quant_cfg` as `str | None` instead of `str | QuantizeConfig | None` ### Testing ``` accelerate launch --config_file examples/gpt-oss/configs/zero3.yaml examples/gpt-oss/sft.py --config examples/gpt-oss/configs/sft_full.yaml --model_name_or_path openai/gpt-oss-20b --quant_cfg MXFP4_MLP_WEIGHT_ONLY_CFG --output_dir gpt-oss-20b-qa ``` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Quantization config parameter now accepts string identifiers or none; resolution behavior for named presets remains unchanged. * **Documentation** * Updated argument reference to reflect the parameter type change while preserving the deprecation note and usage guidance. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> |
||
|
|
c88b62beec |
Fix garbage generation preview in hf_ptq.py when pad_token == eos_token (#1673)
### What does this PR do? Type of change: Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Fixes the generation **preview** in `examples/llm_ptq/hf_ptq.py` producing garbage output (e.g. repeated `\u200b` zero-width-space tokens) for models whose tokenizer has `pad_token == eos_token` — most visibly GLM-5.1. The garbage appeared *before* quantization, so it was not a quantization issue. **Root cause:** `pre_quantize` / `post_quantize` take the first (left-padded) calibration sample and call `full_model.generate(preview_input_ids, ...)` **without an `attention_mask`**. HuggingFace only auto-infers the mask when `pad_token_id != eos_token_id` (`generation/utils.py:_prepare_attention_mask_for_generation`); when they are equal it falls back to an all-ones mask, so the model attends to the leading pad/eos tokens, ignores the real prompt, and (for GLM's MoE/DSA/MTP path) collapses to a single repeated token. Calibration itself was always correct — it already passes the mask; only the preview generation was missing it. **Fix:** thread the calibration batch's `attention_mask` through to both preview `generate()` calls. One file changed (`examples/llm_ptq/hf_ptq.py`, +20/-8). ### Usage No usage change — the same command now produces a coherent preview instead of `\u200b` repetition ### Testing Reproduced the exact mechanism (left padding + pad_token == eos_token + missing attention_mask) on a small model(GPT2): without the mask the model emits the same HF warning as the bug report and ignores the prompt; with the mask the output is byte-identical to the unpadded baseline. Verified no behavioral change for models where pad != eos (the explicit mask equals HF's inferred input_ids.ne(pad_id)) and for Whisper (its batch carries no attention_mask, so the path is unchanged). Pre-commit: ruff-check, ruff-format, and mypy (no new errors vs. main) all pass. Before your PR is "Ready for review" Make sure you read and follow Contributor guidelines and your commits are signed (git commit -s -S). Make sure you read and follow the Security Best Practices (e.g. avoiding hardcoded trust_remote_code=True, torch.load(..., weights_only=False), pickle, etc.). - Is this change backward compatible?: ✅ <!-- Only changes internal helper signatures within the example script; no public API affected. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md: N/A <!-- No copied code, no new dependency. --> - Did you write any new necessary tests?: N/A <!-- Preview path requires model loading; no existing unit-test harness covers it. Verified via a standalone repro of the root-cause mechanism. --> - Did you update Changelog?: N/A <!-- Bug fix confined to an example-script preview; not a library/API change. Happy to add a 0.46 bug-fix entry if preferred. --> - Did you get Claude approval on this PR?: ✅ <!-- Will run `/claude review` before requesting review. --> ### Additional Information Backward compatible across model familes: | Model class | Before (no mask passed) | After (mask passed) | Result | |---|---|---|---| | `pad != eos` (most: T5, BART, many LLMs) | HF infers mask = `input_ids.ne(pad_id)` | explicit calib mask = same tensor | **Identical output** — no change | | `pad == eos` (GLM-5.1, GPT-2-style) | all-ones fallback → attends to pad → garbage | correct mask | **Fixed** | | Whisper | no mask | batch has no `attention_mask` key → `None` → no mask | **Identical** — no change | | Nemotron-VL / DeepSeek / NemotronH / `--skip_generate` | `generate()` not called on this path | unchanged | No change | <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Bug Fixes** * Enhanced LLM post-quantization example to properly handle attention masks during preview generation. The quantization preview now correctly threads attention masks through generate() calls, ensuring accurate generation outputs are captured both before and after quantization steps. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> |
||
|
|
43b67a83c5 |
specdec_bench: keep method=mtp when adding model=<assistant> for Gemma 4 MTP (#1677)
### What does this PR do?
Type of change: Bug fix
Fixes the specdec_bench vLLM wrapper's MTP `speculative_config` emission
so Gemma 4 MTP no longer hits the wrong code path inside vLLM.
### Bug
vLLM's `SpeculativeConfig.__post_init__`
(`vllm/config/speculative.py:529-602`) auto-detects `method` ONLY when
it's unset. When `model` is provided and `method` is `None`, the default
branch sets `method = "draft_model"` — the generic same-architecture
draft path, NOT MTP. That path enforces equal num_heads between target
and draft and raises:
```
AssertionError: All layers in one attention group must share num_heads; got {8, 4}
```
on heterogeneous-head models. Gemma 4 has 8 target heads and 4 draft
heads by design.
### Where the previous fix went wrong
PR #1663 changed the MTP branch in the wrapper to emit `{model:
<assistant>, num_speculative_tokens: N}` WITHOUT `method` when
`draft_model_dir` was provided, based on a misread of vLLM PR #41745's
test plan that only showed the `{model, num_speculative_tokens}` shape.
That test plan was the direct `LLM(...)` constructor invocation; vLLM
had already defaulted method internally. Going through specdec_bench's
`AsyncEngineArgs(speculative_config=...)` path, the explicit `method`
key is required to avoid the auto-detect → draft_model fallback.
### Reference
vLLM's own test at
[`tests/v1/e2e/spec_decode/test_spec_decode.py:818-823`](https://github.com/vllm-project/vllm/blob/main/tests/v1/e2e/spec_decode/test_spec_decode.py#L818)
does exactly this for the gemma4-e4b parametrization:
```python
speculative_config = {
"method": method, # "mtp"
"num_speculative_tokens": ...,
}
if draft_model is not None: # Gemma 4 case
speculative_config["model"] = draft_model
```
### Fix
Restore `method="mtp"` as the unconditional MTP path. ADD `model` only
when `draft_model_dir` is set. Backward-compatible for Qwen 3.5 MTP /
DeepSeek MTP / other inline-MTP families (they keep the bare `{method:
"mtp"}` config).
### Validation
Field-tested via vLLM PR #41745's correctness test on `gemma-4-E4B-it` +
`gemma-4-E4B-it-assistant`: produced 304.7 output TPS at γ=4 vs 171.0
baseline (178% speedup) on H100. The same `speculative_config` shape
this fix emits.
### Surfaced on
[OMNIML-5024](https://jirasw.nvidia.com/browse/OMNIML-5024) pipeline
#54356795:
- Wrapper emitted `{model: assistant, num_speculative_tokens: 3}`
- vLLM auto-detected `method = "draft_model"`
- Loaded gemma-4-E4B-it-assistant (4 heads) as a generic draft for
gemma-4-E4B-it (8 heads)
- Attention-group num_heads check tripped → AssertionError, task_0
FAILED, task_1 CANCELLED
### Before your PR is "*Ready for review*"
- Backward compatible: ✅ (Qwen 3.5 / DeepSeek MTP unchanged; only the
MTP+`draft_model_dir` case changes).
- New tests: ❌ — the test exercising this codepath would need a GPU +
gemma-4 model checkout, which is cluster work, not unit-test scope.
JIRA-tracked validation via OMNIML-5024 dispatch after this lands.
- Changelog: ❌
### Additional Information
- vLLM PR #41745 (Gemma4 MTP support)
- Companion: NVIDIA/Model-Optimizer PR #1675 (launcher
`GlobalVariables.draft_model` schema fix)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Fixed speculative decoding configuration handling in the benchmark
example to ensure consistent method assignment and proper draft model
configuration.
* **Documentation**
* Updated configuration comments to reflect corrected behavior and
improved clarity.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
|
||
|
|
46eddab877 |
[Feat]: Specdec Streaming: RDMA + Multinode (#1611)
### What does this PR do? Type of change: New feature Multi-node **streaming** training for speculative decoding (EAGLE3 / DFlash): a live `vllm serve` captures the target model's hidden states and moves them straight to the trainer over **NIXL RDMA** — no disk round-trip. The streaming dataset is map-style — each rank fetches only its own `DistributedSampler` shard (concurrency from `dataloader_num_workers`), round-robins across multiple serve replicas (`server_urls`), and scales to multi-node DDP. Serve-side tensor parallelism (TP>1) is supported: hidden states are replicated across TP ranks, so rank 0 alone owns the pool + transfer. ### How - `RdmaHiddenStatesConnector` — out-of-tree vLLM connector (no vLLM source edits): one pre-registered pinned NIXL pool per serve, a ring slot per request, and a small HTTP sidecar serving transfer metadata. The trainer RDMA-READs the slot into a per-worker buffer. RDMA is the **only** transport (the earlier disk/safetensors path is removed). - Map-style dataset + multi-node accelerate launch (`--machine_rank`, optional Slurm `--segment` to keep nodes in one NVLink domain). ### Usage ```yaml data: mode: streaming streaming_server_url: "http://node0:8000,http://node1:8000" # round-robin ``` ### Validation (Qwen3-8B, oci-nrt H100) sandbox CI: https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/337489812 **1. End-to-end convergence — EAGLE3 & DFlash, 5000 steps.** Both algorithms converge and export a deployable draft; the DFlash drafts also serve under vLLM speculative decoding (8/8 smoke prompts pass). | algorithm | topology (nodes) | train loss (step 0 → 5000) | vLLM draft acc-len | |---|---|---|---| | EAGLE3 | 2 serve TP=2 + 2 trainer DDP (4) | 37.1 → 8.20 | — | | DFlash | 1 serve TP=1 + 1 trainer (2) | 11.7 → 5.56 | 1.11 | | DFlash | 2 serve TP=2 + 2 trainer DDP (4) | 10.9 → 5.26 | 1.19 | <!-- Drag these PNGs in here (GitHub turns them into asset URLs): eagle3_streaming_loss.png, dflash_streaming_loss_singlenode.png, dflash_streaming_loss_multinode.png --> **2. Scalability — 1 → 12 nodes (EAGLE3, 200 steps).** Throughput scales ~23× across the sweep below. The step-time growth is cross-node DDP all-reduce, not the streaming path — RDMA (~0.33 ms/req @ 2 MB, ~47 GB/s host-pinned READ) is never the bottleneck. Scale serve + trainer nodes together for near-linear speedup. | serve / trainer | nodes | step time | samples / step | samples / sec (global) | acc @ step 200 | |---|---|---|---|---|---| | 1 serve / 1 rank (co-located, 1 node 2 GPU) | 1 | 0.23 s | 1 | 4.4 | [0.141, 0.094, 0.072] | | 1 serve / 1 rank (cross-node) | 2 | 0.23 s | 1 | 4.3 | [0.137, 0.105, 0.074] | | 2 serve / 8 ranks | 3 | 0.26 s | 8 | 31.1 | [0.215, 0.126, 0.097] | | 4 serve / 16 ranks (2 trainer nodes) | 6 | 0.28 s | 16 | 56.5 | [0.217, 0.148, 0.110] | | 8 serve / 32 ranks (4 trainer nodes) | 12 | 0.31 s | 32 | 101.9 | [0.235, 0.165, 0.137] | **3. Serve-side TP correctness.** TP=1 vs TP=2 draft top-1 accuracy track step-for-step (hidden states are replicated across TP ranks). <img width="910" height="546" alt="serve-tp-acc" src="https://github.com/user-attachments/assets/73df9214-7ff0-4ab4-bf2f-95842b12cd5f" /> ### Before your PR is "*Ready for review*" - Backward compatible?: ❌ — streaming is now RDMA-only; `server_url` → `server_urls`; the disk transport (`HS_TRANSPORT`, `streaming_shared_storage_path`) is removed. - New tests?: ✅ `tests/unit/torch/speculative/plugins/test_hf_streaming_dataset.py` (map-style dataset + mocked RDMA fetch). - Updated Changelog?: ❌ - Claude approval?: ❌ --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
66b54ed040 |
[OMNIML-5024] specdec_bench cell t0_d3 — google/gemma-4-E4B-it / MTP / vllm (#1663)
### What does this PR do?
Type of change: Bug fix + new example
Wires SPEED-bench's MTP path to support **Gemma 4** (and any future MTP
variant that uses a separate assistant / draft model), and adds the
SPEED-bench MTP/vLLM example for `google/gemma-4-E4B-it`.
**Key difference: Gemma 4 MTP vs. generic MTP.** vLLM's
`speculative_config` accepts two different shapes for MTP:
| Variant | `speculative_config` shape | Models |
|---|---|---|
| **Generic MTP** | `{"method": "mtp", "num_speculative_tokens": N}` |
Models that carry their own MTP layer in-tree (e.g. Qwen 3.5 MTP
variants) — no separate draft / assistant model. |
| **Assistant-model MTP** | `{"model": "<assistant>",
"num_speculative_tokens": N}` (no `method` key — vLLM auto-detects from
the assistant) | Gemma 4 family (E2B / E4B / 26B-A4B / 31B); each target
model has a paired `<target>-assistant` checkpoint that acts as the MTP
draft. Landed in
[vllm-project/vllm#41745](https://github.com/vllm-project/vllm/pull/41745)
(2026-05-06). |
The specdec_bench vLLM wrapper at
`examples/specdec_bench/specdec_bench/models/vllm.py` previously emitted
only the generic shape for any `--speculative_algorithm MTP` invocation,
which produced `NotImplementedError: Unsupported speculative method:
'mtp'` on Gemma 4 even with a container that has the support
(`vllm/vllm-openai:v0.22.1`+). This PR teaches the wrapper to switch
shapes based on whether `--draft_model_dir` is provided.
**Concrete changes:**
1. **`examples/specdec_bench/specdec_bench/models/vllm.py`** — when
`speculative_algorithm == "MTP"` AND `draft_model_dir` is set, emit
`{"model": draft_model_dir, "num_speculative_tokens": N}`
(assistant-model shape). Otherwise emit the existing `{"method": "mtp",
...}` (generic shape). Backward-compatible — Qwen 3.5 MTP and other
callers that omit `--draft_model_dir` get the same config they got
before.
2. **`examples/specdec_bench/specdec_bench/utils.py`** — `get_tokenizer`
reads `extra_special_tokens` from the model's `tokenizer_config.json`
and passes them through to `AutoTokenizer.from_pretrained`. Gemma 4
tokenizers ship a list-shaped `extra_special_tokens` entry that the
constructor would otherwise reject. Necessary for any Gemma 4 cell.
3.
**`tools/launcher/examples/gemma-4/gemma-4-E4B-it/specdec_bench_mtp_vllm.yaml`**
— SPEED-bench parent YAML for `google/gemma-4-E4B-it`. Uses
`vllm/vllm-openai:v0.22.1` (has `gemma4_mtp.py` from #41745) and wires
`--draft_model_dir /hf-local/google/gemma-4-E4B-it-assistant` on both
task_0 (qualitative) and task_1 (throughput_32k).
4.
**`tools/launcher/common/specdec_bench/_cells/gemma-4-E4B-it_mtp_vllm_t0_d3.yaml`**
— runtime params for the `t0_d3` cell of OMNIML-5022 (`temperature=0`,
`max_model_len=40960`).
### Usage
```python
# Wrapper-level: same CLI as before, just pass --draft_model_dir for
# Gemma 4 MTP. The wrapper auto-routes to the assistant-model shape.
# python examples/specdec_bench/run.py \
# --engine VLLM \
# --speculative_algorithm MTP \
# --draft_model_dir /hf-local/google/gemma-4-E4B-it-assistant \
# --draft_length 3 \
# --tp_size 1 \
# ...other SPEED-bench knobs...
# Equivalent direct vLLM invocation (for reference, no wrapper):
from vllm import LLM, SamplingParams
llm = LLM(
model="google/gemma-4-E4B-it",
speculative_config={
"model": "google/gemma-4-E4B-it-assistant",
"num_speculative_tokens": 3,
},
trust_remote_code=True,
)
```
### Testing
- **Upstream existence checks**: verified the assistant models
`google/gemma-4-{E2B,E4B,26B-A4B}-it-assistant` exist, public, ungated
on HuggingFace; verified `vllm/model_executor/models/gemma4_mtp.py` is
in vLLM `v0.22.0`, `v0.22.1`, and `main`.
- **Backward compat**: `MTP` callers that don't pass `--draft_model_dir`
(e.g. the existing Qwen 3.5 MTP/vLLM cells under
`tools/launcher/examples/Qwen/Qwen3.5-4B/`) take the unchanged
`{"method": "mtp", ...}` branch. No diff for those.
- **End-to-end cluster validation**: pending. Will run via the
OMNIML-5022 cells (OMNIML-5024 / 5025 / 5026 / 5027) once the
nmm-sandbox submodule pin advances past this PR. Each cell exercises
`task_0` (SPEED-Bench qualitative, 880 samples) + `task_1`
(throughput_32k, 80 samples) on cw_dfw, single H100.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ — the wrapper only takes the
new branch when `--draft_model_dir` is provided alongside
`--speculative_algorithm MTP`. Existing MTP callers (Qwen 3.5 etc.) keep
the generic `method: "mtp"` config.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependencies.
- Did you write any new necessary tests?: ❌ — relying on the SPEED-bench
cluster cells (OMNIML-5024 …5027) for end-to-end validation; no unit
test fixture for the vLLM wrapper exists in `tests/` for me to extend
symmetrically. Happy to add one if reviewers want it.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — small fix + example addition. Can add if requested.
- Did you get Claude approval on this PR?: ❌ — will run `/claude review`
once the PR is marked Ready for review.
### Additional Information
- JIRA: [OMNIML-5024](https://jirasw.nvidia.com/browse/OMNIML-5024)
(cell_t0_d3); siblings OMNIML-5025/5026/5027 (cell_{t0_d7, t1_d3,
t1_d7}) of Epic OMNIML-5022 are blocked on this PR landing.
- Upstream reference: vllm-project/vllm#41745 — "[Spec Decode] Add
Gemma4 MTP speculative decoding support".
- Companion (pensieve-intern !91, internal): adds a
"Model-family-specific MTP invocation" table to the specdec_bench cell
SPEC so future agents pair `MTP` with the right `--draft_model_dir` from
SPEC-read time.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added a SPEED-bench pipeline for Gemma 4 using vLLM speculative
decoding (MTP) with qualitative and throughput tasks.
* **Improvements**
* Speculative-decoding logic updated to handle assistant-model and
generic MTP cases distinctly.
* Tokenizer loading now reads and normalizes extra special tokens from
tokenizer config when available.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Pensieve Intern <chenhany@nvidia.com>
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
|
||
|
|
bde162a701 |
feat(deepseek): add --cast_mxfp4_to_nvfp4 to deepseek_v4 quantize step (#1653)
### What does this PR do? Type of change: new feature Brings the GPT-OSS lossless MXFP4 → NVFP4 cast (#1372) to DeepSeek V4's routed-expert export by adding a `--cast_mxfp4_to_nvfp4` flag to `examples/deepseek/deepseek_v4/quantize_to_nvfp4.py`. To avoid duplicating the closed-form math, the shared numerics — `mxfp4_to_nvfp4_global_amax`, `mxfp4_to_nvfp4_per_block_amax`, and the E2M1/E4M3/E8M0 constants — are **hoisted out of the GPT-OSS example cast into the library** at `modelopt/torch/quantization/utils/numeric_utils.py`. Both the GPT-OSS cast (`examples/llm_ptq/cast_mxfp4_to_nvfp4.py`) and the new DeepSeek path now import them from there. DeepSeek V4's routed experts ship as MXFP4 (E2M1 nibbles + a power-of-two E8M0 scale per 32-element block). By default the export dequantizes them to BF16 and re-quantizes to NVFP4 using the calibrated per-tensor weight amax, which re-derives per-block scales from the data and is therefore lossy. With the flag, the cast pins `scale_2 = 2^(k_max-8)` and each per-block E4M3 scale to `2^(k_j-m)` straight from the source E8M0 scales, so `per_block_scale * scale_2 = 2^k_j` and the NVFP4 nibbles equal the source MXFP4 nibbles bit-for-bit (for every block whose `k_j` lands in E4M3's representable window; rare out-of-range blocks clamp). The one V4-specific addition is that w1/w3 share a single `scale_2` for the fused GEMM1, so `k_max` is taken over both projections. The flag only affects routed-expert **weights** — activation `input_scale` still comes from `--amax_path` calibration. ### Usage ```bash python deepseek_v4/quantize_to_nvfp4.py \ --amax_path ${AMAX} \ --source_ckpt ${DS_V4} \ --output_ckpt ${HF_NVFP4_PATH} \ --cast_mxfp4_to_nvfp4 ``` ### Testing - The hoisted numerics get unit tests in `tests/unit/torch/quantization/test_numeric_utils.py` (10 cases: per-tensor global_amax, per-block amax incl. out-of-range, magnitude-table cache) — 10/10 pass. The example test `tests/examples/llm_ptq/test_cast_mxfp4_to_nvfp4.py` keeps the cast-specific cases (quantizer naming, `build_amax_map`, `apply_to_model`). - Validated on real DeepSeek-V4-Flash expert tensors (incl. the on-disk `float8_e8m0fnu` scale dtype): 23.5M blocks, 100% lossless, 0 error. - Generated a full NVFP4 checkpoint for DeepSeek-V4-Flash (43 layers, 256 routed experts) end-to-end: `[cast] lossless MXFP4->NVFP4 blocks: 8,657,043,456/8,657,043,456 (100.0000%)`. Output weights match an independently-produced reference cast byte-for-byte (`weight_scale`, `weight_scale_2`, packed nibbles modulo the harmless sign-of-zero). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (new opt-in flag; default export behavior unchanged; hoist re-exports through the existing example module) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ N/A (no new deps; shared numerics moved into the library rather than duplicated) - Did you write any new necessary tests?: ✅ (library numerics covered by `tests/unit/torch/quantization/test_numeric_utils.py`; end-to-end validated on a real DeepSeek-V4 checkpoint) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ (will run `/claude review`) ### Additional Information Mirrors and reuses #1372 (GPT-OSS MXFP4 → NVFP4 cast); the closed-form numerics are now shared via `modelopt.torch.quantization.utils.numeric_utils`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added `--cast_mxfp4_to_nvfp4` flag to perform a closed-form, mostly lossless MXFP4→NVFP4 conversion for routed-expert weights with aggregated lossless/block statistics. * **Documentation** * Updated DeepSeek V4 export instructions and README to document the new flag and clarify calibration behavior for activation scales. * **Chores** * Exposed shared numeric quantization utilities for MXFP4→NVFP4 casting. * **Tests** * Added and updated tests to validate the new numeric helpers and conversion behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |