mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
b16356776b8dd329059985dea6fc98d26bf07d94
1251
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b16356776b |
Build the GGML IQ packing kernels as a single CUDA extension (#2462)
### What does this PR do? Type of change: Code refactoring `#2448` added the GGML IQ packing kernels as **two** torch extensions, `modelopt_cuda_ext_iq1_s` and `modelopt_cuda_ext_iq2_xs`. This merges them into one, `modelopt_cuda_ext_ggml`. The existing per-extension split in `extensions.py` exists for reasons that don't apply to the IQ formats: `get_cuda_ext` gates on CUDA `>=11` while `_fp8`/`_mx` gate on `>=11.8`, and `_mx` needs `--use_fast_math`, which must not reach the base `tensor_quant` kernels. `get_cuda_ext_iq1_s` and `get_cuda_ext_iq2_xs` differed in none of that — same `>=11.8` gate, same `-O3` flags, same `common.cuh` — so the split only compiled the shared header twice, ran nvcc twice, and grew the loader, `__getattr__`, and `precompile()` once per format. With IQ2_XXS / IQ3_S / IQ4_NL plausibly following, that scales badly. Changes: - New `ggml/ggml.cpp` holds both host-side validation wrappers and the single `PYBIND11_MODULE`, binding `iq1_s_pack` and `iq2_xs_pack` (previously each module exported a bare `pack`). Deletes `ggml/iq1_s.cpp` and `ggml/iq2_xs.cpp`; the validation logic and docstrings carry over unchanged. - `get_cuda_ext_iq1_s` + `get_cuda_ext_iq2_xs` → `get_cuda_ext_ggml`, which builds `ggml.cpp`, `iq1_s.cu`, and `iq2_xs.cu` together. The retry-on-`raise_if_failed` semantics of the old getters are preserved. - Each format keeps its kernels in its own translation unit, so adding a format is a new `.cu` plus one `module.def` — no new extension, loader, or `precompile()` line. No caller outside `extensions.py` and its tests referenced the old getters on `main`, so nothing else changes. **Note for the follow-up PRs in the `#2448` series (`#2446`/`#2447`/`#2449`): the codec layer should call `get_cuda_ext_ggml().iq1_s_pack(...)` / `.iq2_xs_pack(...)` instead of `get_cuda_ext_iq1_s().pack(...)` / `get_cuda_ext_iq2_xs().pack(...)`.** ### Usage ```python from modelopt.torch.quantization.extensions import get_cuda_ext_ggml ext = get_cuda_ext_ggml(raise_if_failed=True) iq1_s_payload = ext.iq1_s_pack(weight, iq1s_grid) # uint8 [numel / 256, 50] iq2_xs_payload = ext.iq2_xs_pack(weight, iq2xs_grid, scales) # uint8 [numel / 256, 74] ``` ### Testing Ran on a single H200 NVL (TRT-LLM `1.3.0rc27.dev202609170000` container), building the merged extension from scratch: - `pytest tests/gpu/_extensions/test_torch_extensions.py` — **24 passed** (6:44). This is the full existing IQ suite (zero-block layout, encode, dtype rejection, row-straddling rejection, invalid/negative-zero scales, byte-exact dtype equivalence, and the brute-force optimality round-trip) reparametrized onto the merged module, plus the untouched `modelopt_cuda_ext` / `_fp8` / `_mx` load tests. - Verified `precompile()` loads all four extensions and that the merged module exports exactly `iq1_s_pack` and `iq2_xs_pack` with the expected arities. - Off-GPU: compiled the three sources directly and linked them into one `.so` to confirm no duplicate-symbol collisions between the two `.cu` translation units. - `pre-commit run --files ...` passes on all changed files (ruff, mypy, clang-format, bandit, license headers). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — the removed getters were added in `#2448` (merged today, unreleased) and have no callers outside this file's own tests. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new code or dependencies; the moved wrappers keep their original attribution. - Did you write any new necessary tests?: ✅ — existing coverage reparametrized onto the merged module; no behavior change to test. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — internal refactor of an unreleased, not-yet-wired-up API. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information Follow-up to #2448. Merge before the remaining PRs in that series (#2446, #2447, #2449) land, so the codec layer is written against `get_cuda_ext_ggml` and no rename is needed afterwards. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added IQ1_S packing support through the GGML CUDA extension. - Added a unified GGML extension loader for IQ1_S and IQ2_XS packing. - Improved extension loading reliability when a cached extension is unavailable. - **Changes** - Renamed the IQ2_XS packing binding from `pack` to `iq2_xs_pack`. - Consolidated IQ1_S and IQ2_XS extension access under the shared GGML loader. - Updated GPU validation and coverage to use the unified extension interface. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
542012d4d5 |
Update CODEOWNERS (#2463)
Add modelopt-torch-kernels-codeowners <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Added code ownership coverage for the `modelopt/torch/kernels` directory. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
ad1bad7817 |
fix(specdec): gather sharded hidden states in DFlash/DSpark AR generation (#2458)
### What does this PR do? Type of change: Bug fix Fixes the nightly `tests/regression/torch/speculative/test_dflash.py::test_dflash_ar_validate`, red every night since 2026-09-10 with: ``` WARNING: sample 0 (writing) failed: Expected all tensors to be on the same device, but got tensors is on cuda:1, different from other tensors on cuda:0 (when checking argument in method wrapper_CUDA_cat) RuntimeError: AR validation produced no results: all 3 samples failed. ``` **No PR caused this.** `ar_validate.py` loads with `device_map="auto"`, and the regression runner (`…rtxpro6000-l-2-…`) has 2 GPUs, so even Qwen3-0.6B gets sharded. `pseudo_speculative_generate()` then concatenates hidden states from `target_layer_ids`, which spans early *and* late layers — different devices once the base model is split: ```python selected = [base_outputs.hidden_states[lid + hid_offset] for lid in self.target_layer_ids] target_hidden = torch.cat(selected, dim=-1) # cuda:0 ++ cuda:1 -> RuntimeError ``` The block tensors have the same problem from the other side: they follow `input_ids.device`, while the draft module sits on the *last* base layer's device (`_place_draft`). What changed on 2026-09-10 was detection, not behaviour. #2288 made `ar_validate.py` raise when every sample fails; before it, the report block was guarded by `if results and …`, so zero results printed nothing and exited **0**. The test had been green while validating nothing. Its runtime is the tell: 23–26s in every green nightly, 21.6s in the first red one, where it does no validation work at all. The same defect is in #2288's own description, on an 8-GPU sharded **EAGLE3** checkpoint: 80/80 samples dead on the identical error, job exiting 0. So this is not DFlash-specific in principle — but EAGLE3's generate path gathers separately (`pop_and_gather_aux_hiddens`) and needs its own look, which is why this PR stops at the DFlash family. ### Fix Gather everything the draft consumes onto the draft's device, and hand results back on the caller's device since `validate_online` cats them onto the running sequence. Every `.to()` is a no-op on a single device, so single-GPU behaviour is bit-identical. `HFDominoModel` inherits `HFDFlashModel.pseudo_speculative_generate`, so it is fixed by the same change; `HFDSparkModel` overrides it and is patched in parallel (including `prev_token` feeding `markov_step`). ### Usage No API change. ```bash python examples/speculative_decoding/scripts/ar_validate.py \ --model_path <dflash ckpt> --osl 10 --num_samples 3 --steps 7 ``` ### Testing - `tests/unit/torch/speculative/plugins/test_hf_dflash.py` and `test_hf_dspark.py`: **110 passed, 1 skipped**. - The skip is the new `TestShardedBaseGeneration`, which is gated on `torch.cuda.device_count() >= 2`. It stands in for accelerate's dispatch — the base forward stays whole, but its hidden states are spread across two devices the way a real `device_map="auto"` split spreads them — and asserts both returned tensors come back on the caller's device. **I have no 2-GPU box to hand, so this test is unverified locally; it will first execute in CI.** It reproduces the reported error against the unpatched code by construction, but please treat that claim as unconfirmed until the run goes green. - The real verification is the next nightly: `test_dflash_ar_validate` should pass *and* print an AR number, which it has not done in any log I can find. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — device moves only; no-ops when the base model is on one device. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — one multi-GPU regression test (skipped without 2 GPUs). - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ — bug in an unreleased-cycle path that only ever produced a silent no-op; no user-visible behaviour to describe. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information Worth a maintainer view: `device_map="auto"` is a poor default for AR validation, which is a single-process step-by-step loop. Every sharded run of it that I can find has failed. Pinning it to one device would be a smaller blast radius than making every drafting path device-safe — but it would also cap validation at models that fit on one GPU, so I have not changed it here. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved speculative generation with sharded Hugging Face models across multiple GPUs. * Ensured generated draft tokens and outputs are returned on the caller’s input device. * Preserved Markov sampling behavior while improving device coordination. * **Tests** * Added multi-GPU coverage for sharded model generation, including output placement and draft shape validation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
216f28a6e0 |
Consolidate speculative-decoding agent skills into one stage/algorithm tree (#2201)
### What does this PR do?
Type of change: documentation
Reorganizes the EAGLE3 agent skills into a single speculative-decoding
skill, then adds algorithm sheets for DFlash, DSpark, and Domino.
**The problem.** The four `eagle3-*` skills each baked the algorithm
into a *stage* of the same draft-model pipeline:
```
skills/eagle3-new-model/ skills/eagle3-review-logs/
skills/eagle3-triage/ skills/eagle3-validate/
```
Adding DFlash would have meant four more near-duplicate skills, since
the stages are shared and only the algorithm differs.
**The change.** One skill dir shaped like `ptq/` (SKILL.md +
references/), split along the two real axes:
```
plugins/modelopt/skills/speculative-decoding/
├── SKILL.md # router: stage table x algorithm table
└── references/
├── stages/ # the procedure — algorithm-independent
│ ├── configure.md # <- eagle3-new-model
│ ├── review-logs.md # <- eagle3-review-logs
│ ├── triage.md # <- eagle3-triage
│ └── validate.md # <- eagle3-validate
└── algorithms/ # the data sheet — per-algorithm
├── README.md # contract: 6 required sections
├── eagle3.md
├── dflash.md
├── dspark.md # DFlash variant — delta only
└── domino.md # DFlash variant — delta only
```
Stage docs cite algorithm-sheet sections by heading (*Pipeline tasks*,
*Success markers*, *Quality gate*, *Known failures*, ...), so a new
algorithm means one new file plus a table row — no stage edits. Every
recipe in `modelopt_recipes/general/speculative_decoding/` now has a
sheet.
DSpark and Domino are documented as **DFlash variants**, not separate
pipelines: same `recipe_type: speculative_dflash`, same training script,
same `dflash.*` config namespace, selected by
`dflash_architecture_config.projector_type`. Their sheets carry only the
delta.
Writing the sheets surfaced three things the old EAGLE3-only skills got
wrong or missed:
- **Task counts are not fixed.** The old skills hardcoded "4-step
pipeline, task_0 through task_3". DFlash offline is 2 tasks, DFlash
online is 3, Domino is 2. The stage docs no longer assume a count.
- **`--aux-layers` couples the dump to the draft.** For DFlash the
dump's layer count must equal the draft's `num_hidden_layers`; a
mismatch doesn't error, it silently captures the wrong layers. Recorded
under *Known failures*.
- **In-training AR is meaningless for DSpark and Domino.** Both recipes
pin `estimate_ar: false` / `ar_validate_steps: 0` because eval runs the
DFlash backbone with the new head bypassed. Each sheet says so under
*Quality gate* so nobody reads a backbone-only number as a result.
**Behavior change:** the four `/eagle3-*` slash commands are replaced by
one `/speculative-decoding`. This isn't optional —
`tools/precommit/sync_claude_skills.sh` iterates `.agents/skills/*/` one
level deep and plugin discovery is `skills/<name>/SKILL.md`, so a
directory is either one skill or a container of skills, not both.
`tools/launcher/docs/claude_code.md` is updated accordingly.
### Usage
```
/speculative-decoding
```
Or by description — the skill triggers on EAGLE3 / DFlash / DSpark /
draft model / acceptance rate. For a new model, follow the stages in
order:
```bash
# 1. Configure: copy the closest examples/<Org>/<Model>/hf_<mode>_<algo>.yaml and adapt
cd tools/launcher
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml --dryrun # preview
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml --yes # submit
# 2. review-logs -> 3. triage (if anything failed) -> 4. validate
```
### Testing
- `claude plugin validate . --strict` and `claude plugin validate
plugins/modelopt --strict` — both pass
- `pre-commit run --files ...` over all changed files — passes,
including `markdownlint-cli2` and the `sync-claude-skills` symlink hook
(it agrees with the new `.claude/skills/speculative-decoding` symlink)
- Verified the new skill is discovered and its description loads
- Every relative link across the skill tree resolves; every repo path
cited in the sheets exists; no dangling `eagle3-*` reference remains
anywhere in the repo
- Each factual claim in the sheets was checked against its source — the
launcher example YAMLs, the four recipes, `dflash_online_training.sh`,
`vllm_smoke_test.sh`, `check_regression.py`, and
`plugins/hf_{dflash,dspark,domino}.py` — rather than written from memory
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ — the four `/eagle3-*` slash
commands become `/speculative-decoding`. Agent tooling only; no library
or API surface is touched. The three YAML comment fixes are
comment-only, no behavior change.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — documentation; covered
by plugin validation and pre-commit
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — agent tooling only, matching how #2025 handled it
- Did you get Claude approval on this PR?: ✅ — `/claude review` run; 2
findings, both fixed in `9b3c568`
### Additional Information
Follows #2025, which moved the skill tree into the installable plugin.
Two stale in-repo comments were found while sourcing the sheets, and are
**fixed in this PR** (`c2696b1`, comment-only):
1. `modelopt_recipes/general/speculative_decoding/dflash.yaml` pointed
`chat_template` at a `chat_templates/` directory under
`modelopt_recipes` that does not exist — templates live per-model beside
each launcher example.
2. Both offline DFlash example YAMLs annotated `--aux-layers dflash`
with "Must match the draft model's num_hidden_layers". `--aux-layers` is
a preset keyword accepting only `eagle`, `dflash`, or an explicit id
list, so it carries no count. The constraint is real but belongs to the
draft depth the preset resolves to: `--num-draft-layers` on the vLLM
dump, and no override at all on the HF/TRT-LLM dumps, which hardcode 5
via `resolve_aux_layers`. This comment had already misled this PR's own
first draft, which is why it's fixed rather than just documented.
Because of (1), this PR now touches `modelopt_recipes/`, which adds
**@NVIDIA/modelopt-recipes-codeowners** to the required reviewers.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added comprehensive speculative-decoding guidance for configuration,
training, validation, troubleshooting, and supported algorithms.
* Added workflow references for DFlash, Domino, DSpark, and EAGLE3,
including quality checks and failure diagnosis.
* **Documentation**
* Generalized experiment-log review and pipeline triage across
algorithms.
* Clarified DFlash draft-depth configuration, resource sizing, task
recovery, and validation.
* Replaced the EAGLE3-specific workflow entry with the broader
speculative-decoding workflow.
* Removed standalone EAGLE3 skill documentation as guidance is now
consolidated under speculative decoding.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
2ff2e1bc80 |
[OMNIML-5899] Add CUDA kernels for IQ packing (#2448)
## Summary - add native CUDA packing kernels for IQ1_S and IQ2_XS - load both extensions through the quantization extension module - validate caller metadata and launch bounds before contiguous materialization - normalize non-finite input elements consistently with the Python reference path - accept caller-computed IQ2_XS FP16 superblock scales to avoid a duplicate reduction - share common packing helpers and add direct extension compilation and boundary tests ## PR split This work is split into four focused PRs. Each PR targets `main` and owns a disjoint file set: 1. **Kernel** — [#2448: Add CUDA kernels for IQ packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448) 2. **Quantization** — [#2446: Add IQ quantization codecs and backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446) 3. **Export** — [#2447: Export IQ checkpoints from HF and Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447) 4. **Recipes** — [#2449: Add IQ post-training quantization recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449) The required merge order is #2448, #2446, #2447, then #2449. ## Scope This PR owns only native kernel sources, shared packing helpers, extension loading and build registration, and direct extension tests. It does not contain Python codecs, export code, or recipes. ## GPU test coverage Direct kernel-boundary coverage is included in this PR: - [extension compilation, zero payloads, input validation, and row-alignment checks](https://github.com/NVIDIA/Model-Optimizer/blob/74e94db9601870e3569c7c8e73506a2f08c29da8/tests/gpu/_extensions/test_torch_extensions.py) Pack/dequantize numerical, native/reference byte-parity, and non-finite-policy tests are owned by the quantization PR: [IQ1_S](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/gpu/torch/quantization/test_iq1_s_cuda.py) and [IQ2_XS](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/gpu/torch/quantization/test_iq2_xs_cuda.py). ## Dependency behavior On `main`, this PR provides optional CUDA extension loaders and direct extension tests. The Python encoders and fallback dispatch land in #2446. Until #2446 lands, no quantization path calls these getters, so a load failure reports only that the extension is unavailable. The IQ2_XS packer accepts one caller-computed FP16 scale per 256-value block. #2446 owns that predictor and passes the same values to the native and reference encoders. ## Provenance - The CUDA kernels were independently written. - They implement the packed-format contract and sign-parity convention from the pinned [llama.cpp definition](https://github.com/ggml-org/llama.cpp/blob/9b05354ec6fb58b4e665e9a39ebc40285c015638/ggml/src/ggml-common.h). - `16.875 = 15 × (1 + 1/8)` is a derived IQ1_S constant. - `0.125` is part of the encoded format. - `0.61` is our empirical IQ1_S scale predictor, not copied from upstream code. - The IQ2_XS predictor constants are owned by #2446 and are not duplicated in this kernel. Human review is still required to confirm that the attribution and license treatment are sufficient. ## Validation - repository hooks, including native formatting, pass for all changed files - extension loader and test modules compile as Python - all 20 direct-extension and CUDA integration test cases collect locally - CUDA runtime execution is delegated to GPU CI because the local host is macOS - restricted-term scan passes --------- Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
835c041c58 |
fix(specdec): resolve the eagle aux-layer preset in the vLLM hidden-state dump (#2410)
### What does this PR do? Type of change: Bug fix Fixes `nvbugs/6753684`, filed against #2080 by the ModelOpt QA Sentinel. The vLLM offline hidden-state dump rejected `--aux-layers eagle` — **the flag's own default** — so the documented invocation aborted before writing any state: ``` File "collect_hidden_states/compute_hidden_states_vllm.py", line 76, in _resolve_aux_layers_standalone ids = sorted({int(t) for t in aux_layers.split(',') if t.strip()}) ValueError: invalid literal for int() with base 10: 'eagle' ``` **Root cause.** `compute_hidden_states_vllm.py` runs in a stock vLLM container, where importing `modelopt.torch` fails (the full init chain pulls in omegaconf and friends). It therefore carries `_resolve_aux_layers_standalone`, a local copy of the preset logic in `common.resolve_aux_layers`. That copy implemented the `dflash` preset and explicit id lists, but never `eagle` — while `add_aux_layers_args` defaults to `eagle`. The HF and TRT-LLM dumps call the shared helper and were unaffected; only the vLLM path forked, and nothing compared the fork against its source. This PR resolves `eagle` inline, mirroring `hf_eagle.default_eagle_aux_layer_ids`. It also fixes a second defect the bug exposes: the function already had a message naming the accepted values, but it was unreachable, because `int()` raised first. An unrecognised preset now reports what it accepts instead of surfacing the raw `int()` error — which is what made the original failure opaque. ### Usage The previously-broken documented invocation now works: ```bash cd examples/speculative_decoding python collect_hidden_states/compute_hidden_states_vllm.py \ --model Qwen/Qwen2.5-0.5B-Instruct \ --input-data ../dataset/synthetic_conversations_1k.jsonl \ --output-dir /tmp/hs_vllm \ --max-seq-len 512 --tp 1 ``` `--aux-layers dflash` and explicit lists such as `--aux-layers 2,5,8` are unchanged. ### Testing Added `tests/unit/examples/test_vllm_hidden_states_aux_layers.py`, which pins the standalone copy to the shared implementation it mirrors: - `eagle` matches `hf_eagle.default_eagle_aux_layer_ids` across layer counts 4, 6, 8, 12, 24, 28, 32, 36, 48, 52, 61, 80 — deliberately including counts small enough that the `max(0, ...)` clamps collapse ids together. - A named regression case for `nvbugs/6753684`. - `dflash` and explicit-list behaviour unchanged. - Unknown specs (`bogus`, `EAGLE3`, `eagle3`, empty) raise the actionable message. - Out-of-range ids still rejected. Divergence here is silent — the dump would write plausible-looking hidden states from the *wrong* layers, surfacing much later as a poor acceptance rate. Hence pinning to the reference rather than asserting hardcoded lists alone. All 20 assertions verified and every pre-commit hook passes (`ruff`, `mypy`, `bandit`, RST lint, license headers). One caveat worth stating plainly: **pytest could not be run locally.** `tests/unit/conftest.py` imports `modelopt.torch.utils.distributed`, which needs `CPUOffloadPolicy` from `torch.distributed.fsdp` — absent in this machine's torch. Each assertion was executed directly against the real module instead, but CI is the first genuine pytest run. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — strictly widens accepted input; `dflash` and explicit lists behave identically. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — bug fix for a defect present in a previous release. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information The underlying fragility is the duplicated implementation, not this one missing branch. The function's own `TODO: drop this once common.resolve_aux_layers is decoupled from the heavy modelopt.torch import chain` is the real fix; the new test narrows the gap but does not close it. Worth tracking separately if the vLLM dump is expected to keep pace with new presets. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed `--aux-layers eagle` for vLLM offline hidden-state collection. * Added support for the documented `eagle` preset alongside `dflash` and explicit layer IDs. * Improved invalid-option errors to clearly list accepted formats. * Rejects `dflash` configurations when the target model has too few layers. * Continues rejecting layer IDs outside the model’s available range. * **Documentation** * Added a v0.48.0 changelog entry for the fix. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
f377b77116 |
Fail fast on non-finite AutoQuantize output gradients (#2432)
### What does this PR do? Type of change: Bug fix Fail fast when AutoQuantize receives non-finite output gradients, before accumulating sensitivity scores. The error names the affected module and suggests checking the model, data, and loss; it also identifies cuDNN SDPA backward on fully masked rows as one possible cause and gives an explicit retry workaround. Unlike the earlier revision, this does not disable cuDNN or change any attention backend settings. Invalid gradients are not zeroed or ignored, and there is no automatic retry. ### Usage No API or recipe changes. For the reproduced cuDNN failure, the caller can explicitly set `torch.backends.cuda.enable_cudnn_sdp(False)` before a fresh AutoQuantize run. ### Testing - AutoQuantize unit suite: **110 passed**. Coverage includes NaN and positive/negative infinity, module diagnostics, preventing invalid score accumulation, model-state cleanup, and unchanged SDPA backend settings. - Real-model E2E on **four GB300 GPUs**, Qwen/Qwen3.6-35B-A3B, main `8025a3dc5481129aa21fef99cb13a879e1b5847e` plus this patch, batch size 8, 512 calibration samples, and `w4a16_nvfp4_fp8_at_6p0bits-active_moe.yaml`: - Default backend: the new diagnostic fired at `model.language_model.layers.39.self_attn.q_proj`; cuDNN remained enabled and no quantized model was exported. Expected-error check passed. - Explicit cuDNN-disabled fresh run: both 64-batch calibration passes, all 16 scoring batches, optimization at **5.99 effective bits**, and checkpoint export completed with exit code 0. Verified all three indexed safetensors shards and quantization configuration; the index contains 93,563 tensors. - All applicable pre-commit checks and `git diff --check` passed. ### Before your PR is "*Ready for review*" Contributor guidelines and security guidance followed; commits are signed and signed off. - Is this change backward compatible?: Yes; finite-gradient behavior and backend settings are unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A; neither added. - Did you write any new necessary tests?: Yes. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: Yes, 0.48.0 bug fixes. - Did you get Claude approval on this PR?: No; awaiting review. ### Additional Information This improves error reporting rather than fixing the upstream cuDNN kernel. Blackwell-specificity is not established. Checkpoint deployment/reload was not tested. A calibration-only checkpoint-resume attempt completed scoring but encountered a separate `candidate_stats` KeyError. That issue is outside this patch; the successful export validation above used a fresh run without search-state resume. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * AutoQuantize now fails fast when output gradients contain non-finite values, with an actionable error identifying the affected module. * Attention backend settings are preserved and restored after successful runs and failures. * Model state is restored when sensitivity scoring encounters an error during setup or execution. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
b9cfdce8dc |
docs: add Local Hessian NVFP4 weight-scale announcement blog (#2417)
### What does this PR do?
Type of change: documentation
Adds a Local Hessian announcement blog at
`docs/source/announcements/local-hessian.rst`, covering the NVFP4
per-block
weight-scale rule that minimizes output error instead of weight error.
Contents:
- Derivation of the per-block output-error objective and its `16x16`
local
Hessian, with numbered equations.
- Results on Qwen3.5-9B: scale-setting comparison against max, MSE, and
Four-over-six, plus composition with GPTQ.
- Figure 1, a grouped bar chart of the Qwen3.8-27B W4A4 candidate scores
(BF16 in gray, the two scale rules in NVIDIA greens).
- A "Using Local Hessian" section with the config example and the
end-to-end `hf_ptq.py` command.
Two supporting changes outside the blog:
- `docs/source/_static/announcements.css`: the `shibuya` theme has no
`span.eqno` rule, so Sphinx's default `float: right` on equation numbers
cannot share a line with MathJax's full-width display block and the
number
renders *above* the equation. This anchors it to the right of the
equation
instead, and shrinks the table-note class.
-
`docs/source/announcements/assets/qwen3-27b-w4a4-scale-rule-accuracy.png`:
the Figure 1 asset.
### Usage
```python
import modelopt.torch.quantization as mtq
config = {
"quant_cfg": [...], # quantizer configuration
"algorithm": {"method": "local_hessian", "fp8_scale_sweep": True},
}
model = mtq.quantize(model, config, forward_loop)
```
### Testing
Documentation only; no code paths change. The `.rst` parses cleanly
under
docutils. The rendered page has not been checked with a full
`sphinx-build`,
so the equation-number CSS fix and the figure placement are worth an
eyeball
on the built docs before merge.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌
### Additional Information
Two items to settle before this is ready to publish:
1. The `--recipe` example points at
`modelopt_recipes/models/Qwen/Qwen3.8-27B/ptq/nvfp4_local_hessian-fp8_attn-kv_fp8_cast.yaml`,
a placeholder path derived from the existing recipe naming convention.
It
needs to match whatever lands in #2363.
2. The tables report single-run team measurements; the blog says so and
makes
no significance claims.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Documentation**
- Added guidance on NVFP4 Local-Hessian weight-scale selection,
including mathematical details, accuracy comparisons, runtime
considerations, limitations, configuration examples, and reproduction
steps.
- Updated announcement labels, headings, metadata, descriptions, and
filtering text to use “Local-Hessian.”
- **Style**
- Improved announcement formatting for equation labels, display-equation
spacing, Hessian results, table headers, and explanatory notes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: realAsma <akuriparambi@nvidia.com>
|
||
|
|
a448ba9757 |
Add end-to-end W4A4 NVFP4 + QAD tutorial for Qwen3.6-35B-A3B (#2411)
### What does this PR do? Type of change: new example + bug fix <img width="2085" height="1239" alt="image" src="https://github.com/user-attachments/assets/b9ced215-ce8c-4dbe-be74-a75c1c4714b3" /> Adds an end-to-end **W4A4 NVFP4 + Quantization-Aware Distillation** tutorial for [Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) at `examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B/`. It complements the existing Nemotron-3-Nano tutorial (pruning + distillation + FP8). Here the model is unpruned and the technique under test is **W4A4** — aggressive enough that PTQ alone leaves a measurable accuracy gap, which is what QAD exists to close. **Why W4A4 rather than weight-only NVFP4:** W4A16 measured *slower than BF16* in 10 of 12 shapes, because a BF16 activation forces vLLM onto the Marlin dequant fallback and never reaches the Blackwell FP4 tensor cores. W4A4 beats BF16 in 9 of 12 shapes (up to 1.30x) and shrinks the checkpoint 67 GiB -> 22 GiB (3.1x). **What the study found:** only 2 of 6 benchmarks show a statistically significant PTQ deficit, so those are the only two QAD can recover. IFBench is recovered to parity with BF16 (-2.62 pp -> -0.29 pp, gain of +2.33 pp, p=0.036); MMMU-Pro recovers ~40% and retains a significant gap. The other four are lossless under W4A4 to begin with. Also added: - `modelopt_recipes/model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore.yaml` — the PTQ recipe used as the QAD student, usable via `--recipe`. - `data_blend.yaml` — the token-budgeted blend config for the distillation data. - `eval_configs/*.yaml` — one NeMo Evaluator config per benchmark. tau2-bench is separate because it needs `--enable-auto-tool-choice --tool-call-parser qwen3_coder` and `deployment.command` is global to a config. **Two export fixes found while producing these checkpoints** (both change library/example behaviour, both have changelog entries under 0.48.0 Bug Fixes): - `unified_export_megatron.py` — MCore builds `embedding` on the MTP stage as well as the first, so gating export on `hasattr(model, "embedding")` wrote a **second, unreferenced copy of the vocab embedding** whenever an MTP model was exported with PP > 1. The index mapped the key to the later shard, so the extra copy never loaded but still shipped — ~1 GB for this model. Now gated on `model.pre_process`, MCore's own "this rank owns the input embedding" flag. - `export_quantized_megatron_to_hf.py` — stopped passing Megatron's `moe_router_dtype` as the router's *storage* dtype. It is a routing *compute* dtype; the parameter is bf16 in a bf16 model, so the export was widening bf16 to fp32. All 21,495,808 router values in the exported checkpoint have their low 16 bits zero, and vLLM builds the gate at the model dtype and rounds on load, so the dropped bytes carried no information. `export_mcore_gpt_to_hf` still accepts the override. ### Usage ```bash # 1. PTQ (2 GB200 nodes for EP=8) srun ... python examples/megatron_bridge/quantize.py \ --hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \ --recipe model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore \ --tp_size 1 --ep_size 8 --pp_size 1 \ --calib_dataset_name cnn_nemotron_v2_mix --calib_num_samples 1024 --calib_batch_size 1 \ --seq_length 8192 --skip_generate \ --export_megatron_path /path/to/qwen36_w4a4_megatron # 2. QAD (32 nodes x 4 GB200) python -u examples/megatron_bridge/distill.py \ --teacher_hf_path Qwen/Qwen3.6-35B-A3B --student_hf_path Qwen/Qwen3.6-35B-A3B \ --student_megatron_path /path/to/qwen36_w4a4_megatron \ --tp_size 1 --pp_size 1 --cp_size 1 --ep_size 8 \ --seq_length 32768 --mbs 1 --gbs 512 --train_iters 500 \ --lr 1e-5 --min_lr 1e-6 --lr_warmup_iters 50 --logit_kl_topk 4096 \ --recompute_granularity full --recompute_method uniform --recompute_num_layers 1 \ --no_async_save --eval_iters 0 --save_interval 50 \ --data_paths "${DATA_BLEND}" --output_dir /path/to/qad_output ``` ### Testing **Library changes.** `tests/gpu_megatron/torch/export/test_unified_export_megatron.py` gains `test_unified_export_megatron_pp2_mtp_no_duplicate_tensors`: it exports a PP=2 model built with `mtp_num_layers=1` and asserts no tensor lands in more than one shard. Verified to **fail without the fix**: ``` AssertionError: tensors written to more than one shard: {'model.embed_tokens.weight': ('model-00001-of-00002.safetensors', 'model-00002-of-00002.safetensors')} ``` The pre-existing `..._pp2_mtp_metadata_matches_shards` test cannot catch this — it fakes `_get_mtp_state_dict` on a model with no real MTP, so the last stage never builds an embedding. Ran the whole `tests/gpu_megatron/torch/export/` suite with and without the fix: identical failure sets (3 failures both ways, all `qwen3_5_moe_vl_*` from a local `ImportError: FLA is not installed`), 58 passed with vs 56 without — the +2 being the new test's two workers. `tests/unit/recipe` passes 368/368 after the recipe path move. `model.pre_process` is always present: `GPTModelExporter.__init__` raises unless the model is `GPTModel` or `HybridModel`, and both set it unconditionally. Both export fixes were also applied to the real 23 GB checkpoints and re-validated end to end: every retained tensor md5-identical, index/shard integrity re-checked, and a **full GPQA re-evaluation of the fixed checkpoint** scored 83.49 vs 84.25 before (paired per-question t-test over the same 198 questions x 16 repeats: -0.76 pp, p=0.21, not significant). **Numbers in the tutorial** come from real runs, not estimates: - **253 evaluation runs** across BF16, the published W4A16 checkpoint, W4A4 PTQ, and QAD at 50 / 300 / 500 iterations — 8 repeats per benchmark (3 for tau2-bench; GPQA is one `num_repeats: 16` run). - The published `nvidia/Qwen3.6-35B-A3B-NVFP4` checkpoint was re-evaluated under this same harness (36 runs) rather than quoted from its card, so the W4A16 row is same-harness. - Every figure and results-table value is generated from the collected `results.yml` files by a script, and I verified the README table cell-by-cell against that data after each edit. - Throughput rows were cross-checked against the recorded AIPerf sweeps; the QAD wall-clock figures against the two jobs' Slurm records (`03:34:49` + `02:09:49`). - All CLI flags in the tutorial were verified to exist in `quantize.py` / `distill.py` / `export_quantized_megatron_to_hf.py`, and `cnn_nemotron_v2_mix` against `dataset_utils.py`. The tutorial also records the non-obvious constraints found the hard way: QAD on this model requires `TP=PP=CP=1` (TP breaks quantizer `_amax` dist-checkpoint sharding, PP starves Qwen3-VL's M-RoPE of `position_ids`, CP hits a rope shard mismatch), EP must match the PTQ checkpoint, and `--logit_kl_topk` is mandatory at 32K because the dense `[seq, vocab]` fp32 logits are 30.31 GiB per tensor on a 248,320-token vocabulary. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ <!-- gpu_megatron PP=2+MTP export dedup test --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!-- 0.48.0: Megatron Framework + two Bug Fixes --> - Did you get Claude approval on this PR?: ✅ <!-- not yet run --> ### Additional Information Changelog entries are filed under **0.48.0**; the `cherry-pick-0.47.0` label has been removed. Rebased onto `main` after #2328 renamed `modelopt_recipes/huggingface` to `model_type` (it is now a compatibility symlink), so the recipe moved to `model_type/qwen3_6_moe/ptq/`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an end-to-end Qwen3.6-35B-A3B tutorial for W4A4 NVFP4 quantization and quantization-aware distillation. * Added checkpoint export, accuracy evaluation, and vLLM throughput benchmarking workflows. * Added evaluation configurations for AA-LCR, GPQA, IFBench, MMMU-Pro, SciCode, and tau2 Telecom. * Added a token-budgeted supervised fine-tuning data configuration. * Added a Megatron-Core NVFP4/FP8 quantization recipe for Qwen3.6-MoE. * **Documentation** * Added benchmark results, deployment guidance, hardware requirements, reproduction steps, HTTPS endpoint guidance, announcement filters, and tutorial links. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f68bf83cfa |
Add published MiniMax M2.7 NVFP4 PTQ recipe (#2437)
Chad's Agent ### What does this PR do? Type of change: new example. Backfill the recipe for [nvidia/MiniMax-M2.7-NVFP4](https://huggingface.co/nvidia/MiniMax-M2.7-NVFP4): max-calibrated NVFP4 W4A4 experts (block size 16), BF16 attention/router/LM head, and FP8 KV-cache cast mode. The published metadata identifies ModelOpt `0.43.0rc2.dev96+g3baa2da62.d20260410`, matching the original run. At that revision, `nvfp4_experts_only` used max calibration and the CLI defaulted to `fp8_cast`. The upload record links that run's checkpoint to the published model. ### Usage ```bash python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path /path/to/MiniMax-M2.7 \ --recipe models/MiniMaxAI/MiniMax-M2.7/ptq/nvfp4_experts_only-kv_fp8_cast \ --moe_calib_experts_ratio 1.0 \ --export_path /path/to/MiniMax-M2.7-NVFP4 ``` ### Testing - Loaded the recipe and checked its resolved patterns/numerics against the published `config.json`: 47,927 Linear module names and 124 KV quantizers across all 62 layers. - Confirmed `hf_quant_config.json` converts exactly to the published `config.json` quantization config using ModelOpt's exporter conversion. - Pre-commit checks and `git diff --check` passed. - No unit tests or README added; full-model PTQ was not rerun. The published config confirms formats and exclusions, not the calibration algorithm or fixed KV amax; those come from the historical run. This is scheme verification, not a weight-hash comparison or bit-identical regeneration claim. ### Before your PR is "Ready for review" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — recipe-only backfill; validated directly. - Did you update Changelog?: N/A — recipe backfill. - Did you get Claude approval on this PR?: N/A ### Additional Information [Published config.json](https://huggingface.co/nvidia/MiniMax-M2.7-NVFP4/blob/main/config.json) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a post-training quantization recipe for the MiniMax-M2.7 model. * Supports NVFP4 quantization for expert weights and activations with max calibration. * Adds FP8 casting for the KV cache while retaining BF16 precision for attention, routing, and language-model output components. * **Documentation** * Documented the recipe and its corresponding published checkpoint. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chad Voegele <cvoegele@nvidia.com> |
||
|
|
8025a3dc54 |
specdec(recipe): add MiniMax-M2.7-DFlash streaming multi-node pipeline (#1835)
## Summary - Add `hf_streaming_dflash_multi_node.yaml` for MiniMax-M2.7 (229B MoE) streaming DFlash training - 2 serve replicas (TP=4, whole node) + 2 trainer nodes (4 GPU each) over NIXL RDMA hidden-state transport - Capture IDs `[2,17,32,47,62,64]` from `build_target_layer_ids(64, 5)` + final layer output - MiniMax-specific: trust_remote_code, FSDP2 via accelerate config, mask_token=200054, YaRN rope_scaling factor=48 - Topology matches Kimi-K2.5 large-MoE streaming recipe Resolves OMNIML-5221 ## Test plan - [ ] Dry-run validation (`uv run launch.py --yaml ... --dry-run`) - [ ] Server-only smoke on CW-DFW (task_1 with `training.max_steps=1`) - [ ] Full streaming training run Signed-off-by: Ye Yu <yeyu@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a multi-node launcher configuration for MiniMax-M2.7 streaming training with speculative decoding. * Added an end-to-end workflow for conversational data preparation, distributed training, checkpoint export, and vLLM smoke testing. * Supports configurable serving and training settings across multiple nodes. * **Enhancements** * Applies Transformers version overrides only to trainer or single-node environments. * Preserves the serving environment’s Transformers version on dedicated serving nodes. * Uses the launcher’s built-in bfloat16 mixed-precision configuration for training. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
c118d359c1 |
docs: replace legacy TensorRT-LLM engine deployment guidance (#2436)
Chad's Agent ### What does this PR do? Type of change: documentation. Replace legacy TensorRT checkpoint export, support matrix, and engine-build instructions with `export_hf_checkpoint` and TensorRT-LLM's PyTorch backend. Preserve the existing 0.48.0 deprecation / 0.49.0 removal notice and page URL. Update the customized-model guide and deployment skill to match. ### Usage Follow the linked unified HF export guide. No API changes. ### Testing - `git diff --check` passed. - `uvx pre-commit run --files docs/source/deployment/1_tensorrt_llm.rst docs/source/guides/_customized_model_quantization.rst plugins/modelopt/skills/deployment/references/trtllm.md` passed. - Full Sphinx build delegated to the Docs workflow; preview expected after deployment. ### Before your PR is "Ready for review" - Is this change backward compatible?: ✅ Documentation only; page URL retained. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — documentation only. - Did you update Changelog?: N/A — guidance correction; no new API deprecation. - Did you get Claude approval on this PR?: ❌ Not requested yet. ### Additional Information Removes instructions for the TensorRT backend that current TensorRT-LLM releases no longer support. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Documentation** - Updated TensorRT-LLM deployment guidance to use `export_hf_checkpoint` with the PyTorch backend. - Clarified that this workflow does not require TensorRT engine construction. - Updated DBRX customization instructions for exporting and deploying quantized models. - Added TensorRT-LLM version requirements and links to unified Hugging Face deployment guidance. - Removed guidance for the legacy TensorRT-LLM checkpoint exporter and outdated troubleshooting steps. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chad Voegele <cvoegele@nvidia.com> |
||
|
|
655f94c207 |
Document the nvfp4_act_headroom calibration variant in ptq.md (#2439)
### What does this PR do?
Type of change: documentation
`general/ptq/nvfp4_act_headroom-kv_fp8_cast.yaml` appears in the
shipped-recipes
table in `modelopt_recipes/ptq.md`, but the **Calibration variants**
section —
which documents `max`, `mse`, `input_scale1`, `gptq`, and the
`layerwise`
variants — had no entry for it. Someone scanning that section for "which
calibration do I pick when NVFP4 W4A4 regresses?" only found `mse`,
which
searches **weight** scales and so cannot help when the loss comes from
activation clipping.
This adds the missing entry: the scale formula
(`amax = max(rho * anchor, upper)`) and its defaults, the fact that it
costs one
calibration pass and exports a standard NVFP4 checkpoint with coverage
identical
to `nvfp4_default-kv_fp8_cast`, and the symptoms that should route you
here
rather than to a weight-side calibration — an A16 ablation clears the
regression
while `mse` does not, the symptom is behavioral (verbose or runaway
generations,
hitting the generation cap) rather than a flat score drop, inference
contexts run
longer than the calibration set, or a few rare blocks dominate the
activation
error. MoE experts-only scopes are called out as the common case.
It also extends step 3 of **Choosing a general recipe** so the
escalation path
reads `mse` first, then `nvfp4_act_headroom` when the evidence points at
activations rather than weights.
**Evidence.** The guidance comes from a GLM-5.3-Flash NVFP4 experts-only
W4A4
root-cause study on SciCode (temperature 1.0), which established
causally that
activation quantization at the routed-expert `down_proj` input drove a
large
generation-length blow-up. Swapping `max` for `nvfp4_act_headroom` cut
the median
generation-length regression versus source from +38% to +19% and the
mean from
+19% to +4%, with no capped generations. The entry states plainly that
this was
the best strict-W4A4 result in that study but still missed the p50/p75
near-lossless gate, so headroom is presented as a strong first lever for
activation-driven regressions rather than a guaranteed fix, with a note
that
`rho` should be swept.
### Usage
No API or recipe change; the recipe already ships. This PR only
documents when
to select it over plain `max`:
```python
from modelopt.recipe import load_recipe
cfg = load_recipe("general/ptq/nvfp4_act_headroom-kv_fp8_cast")
```
### Testing
Docs-only change; no code paths touched, so no new or updated tests.
- `pre-commit run --files modelopt_recipes/ptq.md` — all applicable
hooks pass,
including `markdownlint-cli2` and `check-modelopt-recipes`.
- Re-read the rendered section to confirm the new bullet nests correctly
in the
existing `Calibration variants` list and that surrounding entries are
unchanged.
- Cross-checked every claim against the implementation
(`modelopt/torch/quantization/calib/nvfp4_act_headroom.py`), the recipe
YAML,
and the existing `CHANGELOG.rst` entry, so the documented defaults
(`anchor_percentile=1`, `upper_percentile=99.99`, `rho=16384`) and the
NVFP4-input-quantizer-only scope match the code.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- docs-only; the
algorithm's tests already live in
tests/unit/torch/quantization/test_nvfp4_act_headroom.py -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- nvfp4_act_headroom already has a CHANGELOG entry from the PR
that added it; a docs-only follow-up is not changelog-worthy. -->
- Did you get Claude approval on this PR?: ❌ <!-- not run; docs-only
change -->
### Additional Information
The entry deliberately does not sell this on accuracy: in that study the
quantized subtask accuracy (56.80%) was *above* the source checkpoint
(51.18%),
so what headroom recovered was generation-length behavior, not accuracy.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added an NVFP4 activation headroom calibration option with
configurable percentile-based scaling.
* Supports standard NVFP4 checkpoint export with a single calibration
pass.
* Applies to dynamic-block NVFP4 activation quantizers while keeping
weight-scale configuration independent.
* Added guidance for addressing activation-related W4A4 accuracy
regressions, including calibration coverage and recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
6a4b3f147e |
[Fix] Calibrate non-decoder modules during layerwise quantization (#2339)
### What does this PR do? Type of change: Bug fix Layerwise calibration now calibrates enabled quantizers outside transformer layers, such as `lm_head`, while hiding decoder layers from the additional calibration traversal. It also fails early when these quantizers are combined with progressive layerwise export, whose in-place conversion makes the required model calibration unsafe. ### Usage No API changes. ### Testing - `pytest_pwd tests/unit/torch/quantization/test_calib.py tests/unit/torch/quantization/test_layerwise_calibrate.py -q` — 68 passed - Pre-commit on all six changed files — passed - `git diff --check` — passed ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Layerwise calibration now supports quantizers outside transformer decoder layers, including LM-head activation quantizers. - Calibration supports models with CPU-offloaded components while preserving model behavior. - The MSE calibration utility is now publicly available. - Added warnings for calibration passes involving offloaded components. - **Bug Fixes** - Prevented unsupported export-mode calibration when outside-layer quantizers are present. - Ensured outside-layer calibration runs after transformer-layer restoration. - Preserved original forwards while hiding decoder subtrees from traversal and state collection. - Ensured cleanup and alias restoration after errors or completion. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
c7ed23a103 |
Rename modelopt_recipes/huggingface to model_type with backward-compat alias (#2328)
### What does this PR do?
Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.
Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.
- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
`$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.
### Usage
```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
--recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none
# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
--recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```
```python
from modelopt.recipe import load_recipe
load_recipe("model_type/vit/ptq/fp8") # canonical
load_recipe("huggingface/vit/ptq/fp8") # deprecated alias, resolves to the same recipe
```
### Testing
- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
`model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
the loader alias — including a recipe that pulls internal `$import`s.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.
### Additional Information
The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.
- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.
- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
- Local recipe files now take precedence over built-in recipes.
- Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
30f89908f0 |
Audit missing labels before release cherry-picks (#2373)
Chad's Agent ## Summary - query NVBugs MCP for the exact `Committed_ModelOpt_<VERSION>` keyword and inspect linked PRs - audit PRs merged after the latest RC tag for missing cherry-pick labels - verify audit result pagination, report findings, and require confirmation before adding labels ## Testing - `npx --yes markdownlint-cli2@0.18.1 plugins/modelopt/skills/release-cherry-pick/SKILL.md` - exercised the 0.47.0 audit using `0.47.0rc1`; GitHub returned 13 complete results - verified the exact NVBug keyword query returned 33 bugs and inspected their comments - `git diff --check` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Documentation** - Updated the release cherry-pick workflow to exclude pull requests whose merge commits are already included in the target release branch. - Applied this audit to both NVBug-linked and recently merged pull request candidates. - Updated recent pull request queries to reference the selected version and release branch. - Renumbered the remaining workflow steps for consistency. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chad Voegele <cvoegele@nvidia.com> |
||
|
|
3c87751903 |
Deprecate the single-format quantization CLI flags in favour of --recipe (#2426)
### What does this PR do?
Type of change: deprecation
Deprecates the single-format quantization CLI flags in favour of
`--recipe`. Passing one now emits a `DeprecationWarning`; nothing else
changes.
| script | flags |
|---|---|
| `examples/hf_ptq` | `--qformat`, `--kv_cache_qformat` |
| `examples/megatron_bridge/quantize.py` | `--quant_cfg`,
`--kv_cache_quant`, `--weight_only` |
| `examples/torch_onnx` | `--qformat` |
`--recipe` was already authoritative over all six — silently on
`hf_ptq`, and with a runtime warning on `megatron_bridge` — and
`modelopt/recipe/presets.py` already records the intent in a comment:
*"the long-term direction is to retire `--qformat` /
`--kv_cache_qformat` in favour of `--recipe`"*. This makes that a real
deprecation.
A recipe carries the quantization config, the calibration algorithm and
the KV-cache setting in one file, so they cannot drift apart the way
separate flags can. That drift is not hypothetical: the preset path
applies no MTP exclusion while the recipe unit
`default_disabled_quantizers` disables `mtp.*`, so the same model
quantizes differently depending on which entry point was used.
#### The warning fires only when a flag is actually passed
`RecipeSupersededAction` is an `argparse.Action`, and argparse invokes
an action only for options present on the command line — never for a
default. That matters because several of these default to a *quantizing*
value (`--qformat fp8`, `--kv_cache_qformat fp8_cast`); warning on the
defaults would fire on every run, including runs that correctly use
`--recipe` and never mention the flag.
`examples/speculative_decoding/scripts/quantize_drafter.py` keeps
`--qformat` undeprecated: it has no `--recipe`, so there would be
nothing to migrate to.
### Usage
```bash
# deprecated
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> --qformat nvfp4 --kv_cache_qformat fp8_cast
# replacement
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> \
--recipe general/ptq/nvfp4_experts_only-kv_fp8_cast
```
### Testing
Three tests in `tests/examples/hf_ptq/test_hf_ptq_args.py`, all passing:
- passing `--qformat` / `--kv_cache_qformat` raises `DeprecationWarning`
and still parses the value;
- omitting them raises nothing and leaves the defaults (`fp8`,
`fp8_cast`) untouched;
- the action stays wired to both flags, so a future edit cannot drop it
while leaving the help text.
Defaults and parsed values were diffed against `main` and are unchanged
— the action stores exactly what `store` / `store_true` would have.
`ruff` findings are at parity with `main` on every changed file.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — the flags still work, they
only warn.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 Deprecations.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
Draft: the removal release for these flags is not decided here, only the
deprecation.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Deprecations**
* Legacy quantization CLI options now issue visible `FutureWarning`
messages only when explicitly provided.
* Use `--recipe` instead of deprecated options in Hugging Face PTQ,
Megatron-Bridge, and torch-to-ONNX workflows.
* Existing option values, defaults, and parsing behavior remain
unchanged.
* Weight AutoQuantize recipes without an explicit `kv_cache` setting
continue to use `--kv_cache_qformat` as a fallback.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
f70991f36e |
Docs: Add QAT and QAD guide [OMNIML-4859] (#2255)
### What does this PR do? Type of change: Documentation Add detailed quantization aware training (QAT and QAD) guide in our docs, featuring examples on how to run in HF, Megatron-Bridge, and Megatron-LM ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Very short summary of changes only for new features, backward breaking changes, deprecations, or fixes for critical bugs present in previous releases. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added a comprehensive guide for quantization-aware training (QAT) and distillation (QAD), including workflows, framework guidance, setup, training, and export steps. * Updated the Guides navigation to include the new QAT/QAD guide. * Expanded quantization guidance with QAT/QAD workflows, scale handling, and NVFP4 references. * Improved save and restore guidance with clearer references and an explanatory note. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> |
||
|
|
6878c0ebd2 |
[chore]: weekly bump of uv.lock on main (2026-09-14) (#2428)
## Summary Automated weekly update of uv.lock file for nSpect Scanning: - `uv.lock` — upgraded all transitive dependencies to latest compatible versions <details> <summary>uv lock --upgrade output</summary> ``` Using CPython 3.12.14 interpreter at: /opt/hostedtoolcache/Python/3.12.14/x64/bin/python3 Resolved 206 packages in 9.26s Updated accelerate v1.14.0 -> v1.15.0 Updated ast-serialize v0.9.0 -> v0.11.2 Updated databricks-sdk v0.136.0 -> v0.139.0 Updated filelock v3.32.5 -> v3.32.6 Updated google-auth v2.57.1 -> v2.58.0 Updated huggingface-hub v1.30.0 -> v1.31.0 Updated hydra-core v1.3.2 -> v1.3.6 Updated multidict v6.7.1 -> v6.8.0 Added onnxruntime-ep-nv-tensorrt-rtx-cu13 v0.4.0 Updated onnxscript v0.7.1 -> v0.7.2 Updated platformdirs v4.11.7 -> v4.11.8 Updated regex v2026.9.3 -> v2026.9.10 Updated tqdm v4.70.0 -> v4.70.1 Updated tzdata v2026.3 -> v2026.4 Updated uv v0.12.10 -> v0.12.13 Updated uvicorn v0.52.4 -> v0.53.0 Updated virtualenv v21.7.8 -> v21.7.9 ``` </details> Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> |
||
|
|
700e188ce5 |
Document copy-PR testing authorization (#2423)
### What does this PR do? Type of change: documentation. Explain how authorized vetters start NVIDIA-runner checks for pull requests from forks, including the SHA-qualified copy-PR command and links to the centralized contributor and vetter guidance. ### Usage N/A; documentation-only change. ### Testing - `pre-commit run --files CONTRIBUTING.md` - `git diff --check` - Verified both copy-PR documentation links resolve ### Before your PR is "*Ready for review*" - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update `CHANGELOG.rst`?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information Clarifies the missing required-check state encountered on #2231. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Documentation** - Updated fork pull-request guidance to require an authorized reviewer’s approval before NVIDIA-hosted CI runs. - Added instructions for contributors without write access to obtain GitHub workflow approval from a reviewer with write permission. - Added the `/ok to test <full-head-sha>` command for authorizing testing against a specific commit. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
1030791f53 |
Fix distributed AutoQuantize scoring and share backward setup (#2231)
### What does this PR do? Type of change: Bug fix AutoQuantize can measure a group of quantized expert layers at their enclosing MLP output. That enclosing module is often a plain PyTorch container and does not carry distributed-group information, so its sensitivity score was not combined across data- or expert-parallel workers. This PR obtains the distributed groups from the quantized layers when the scoring module does not provide them. It also preserves construction order for quantized modules, scoring modules, and their registered hyperparameters so every worker accumulates scores in the same order. The temporary state needed by backward-based scoring is now managed by one shared session. The session installs and removes forward patches and invocation-specific output-gradient hooks, controls parameter gradients, and restores the active quantization recipes even when scoring raises an exception. Scoring methods remain responsible for their own score calculation. ### Usage N/A — this fixes existing AutoQuantize behavior and does not add an API or flag. ### Testing - `pre-commit run --files modelopt/torch/quantization/algorithms.py tests/unit/torch/quantization/test_autoquant.py` - `pytest -q tests/unit/torch/quantization/test_autoquant.py` — 102 passed - Added a real two-rank gradient AutoQuantize test covering MoE experts scored at an enclosing MLP. - Added regressions for deterministic hyperparameter registration, per-invocation replay for reused score modules, exact `forward`-attribute restoration, partial setup rollback, and cleanup after a scoring failure. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors). - Is this change backward compatible?: ✅ — no API or checkpoint format changes; distributed sensitivity values now include the missing reduction. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — no new feature, deprecation, breaking change, or critical release-note item. - Did you get Claude approval on this PR?: ❌ — pending review. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved quantization scoring consistency through deterministic ordering and invocation handling. * Added more reliable distributed score aggregation, including support for mixture-of-experts models. * Improved gradient-based scoring for repeated evaluations, tuple outputs, and checkpoint-compatible workflows. * Ensured model behavior and scoring state are restored after successful or failed evaluations. * Avoided unnecessary output replay when gradients are not required. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Joshua Hill <joshua.hill@baseten.co> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
51de53e48c |
[6463897] Fix narrow FP16 histogram calibration (#2412)
### What does this PR do? Type of change: Bug fix FP16 entropy calibration can fail for sufficiently narrow activation ranges because NumPy may construct the histogram bin edges at FP16 precision. NumPy 2.2 and later reject the resulting collapsed bin spacing, while earlier versions can silently return invalid, non-monotonic edges. The same precision issue can recur when ONNX Runtime merges later calibration batches. Losslessly widen FP16 activation values to FP32 while calculating and merging histograms in both ONNX entropy calibration paths, then restore the source dtype at the calibration-to-quantization boundary. This keeps the histogram bins stable without changing FP16 Q/DQ scale or graph dtype semantics. FP32 inputs and public APIs are unchanged. For full-range FP16 activations, ONNX Runtime can overflow while subtracting FP16 calibration endpoints before it widens the result. Retry only a non-finite FP16 scale calculation with FP32 endpoints, then cast the finite scale back to FP16. Existing finite FP16 calculations and all non-FP16 calculations continue to use ONNX Runtime's original result. The fallback intentionally patches only the `qdq_quantizer` binding used for calibrated activation ranges. Initializer and weight quantization continue to use ONNX Runtime's existing `quant_utils` path unchanged; full-range FP16 weight scaling is outside this calibration fix. The regression tests exercise both collectors across initial collection, an equal-range merge, and an expanding-range merge. They also verify the internal FP32 histogram and external FP16 calibration-range contract, including finite saturation when restoring sanitized values. A real entropy calibration test covers full-range FP16 values and verifies finite FP16 Q/DQ scales and a loadable ONNX Runtime graph. The AutoCast integration verifies the same FP16 scale-type contract through the public quantization path. ### Usage N/A — no API or usage change. ### Testing All tests ran with CUDA hidden. - NumPy 1.26.4: affected histogram, AutoCast, and ORT-patching modules: `39 passed`. - NumPy 2.2.3: affected histogram, AutoCast, and ORT-patching modules: `39 passed`. - NumPy 2.3.5: affected histogram, AutoCast, and ORT-patching modules: `39 passed`. - Public INT8 entropy quantization with full-range FP16 calibration data: finite FP16 Q/DQ scales, full ONNX check passed, and the CPU ONNX Runtime session loaded. - Changed-file pre-commit hooks: passed. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Follow-up to #1558. > 🤖 _Generated by Codex (AI agent)._ --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com> |
||
|
|
5b1f7e86cc |
[6701308][OMNIML-5805] Correct ONNX PTQ documentation contracts (#2413)
### What does this PR do?
Type of change: documentation
Align the ONNX PTQ README, guide, and executable example with the
implemented contracts:
- use the canonical `--calibration_data_path` CLI option;
- load `.npy` calibration data before passing it to the Python API;
- document the supported Autotune modes and calibration methods;
- correct the minimum opsets to INT8 19, FP8 19, and INT4 21; and
- describe the no-data fallback as random calibration inputs.
This also removes an inaccurate source comment without changing runtime
behavior.
### Usage
```bash
python -m modelopt.onnx.quantization \
--onnx_path=model.onnx \
--quantize_mode=int8 \
--calibration_data_path=calib.npy \
--output_path=model.quant.onnx
```
### Testing
- `pre-commit run --files docs/source/guides/_onnx_quantization.rst
examples/onnx_ptq/README.md modelopt/onnx/quantization/quantize.py
tests/examples/test_onnx_ptq.sh`
- `bash -n tests/examples/test_onnx_ptq.sh`
- `CUDA_VISIBLE_DEVICES="" python -m pytest -o addopts="" -p
no:cacheprovider --confcutdir=tests/unit/onnx/quantization -q
tests/unit/onnx/quantization/test_autotune_quantization_integration.py`
(4 passed)
- `nox -s docs` (passed; Sphinx built 881 HTML files)
- Focused before/after contract probe covering the documented CLI
option, API data type, Autotune modes and methods, opset minimums, and
random-input wording
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A
### Additional Information
Tracking: [6701308]
> 🤖 _Generated by Codex (AI agent)._
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Documentation**
- Clarified that random calibration inputs are used when no calibration
dataset is provided.
- Updated ONNX post-training quantization examples with minimum opset
requirements and the `calibration_data_path` argument.
- Clarified Autotune support for FP8 and INT8 calibration methods using
`max` or `entropy`.
- **Tests**
- Updated quantization command examples to use the current calibration
data path option.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
|
||
|
|
c37a6948db |
feat(quantization): fail fast when a quant config matches no weight quantizer (#2203)
### What does this PR do?
Type of change: New feature (fail-fast guard; behavior change on a
previously silent path)
A `quant_cfg` whose module patterns don't match the model is not an
error to `set_quantizer_by_cfg` — every pattern simply matches nothing.
The run then calibrates, exports, and hands back a checkpoint that is
silently unquantized:
```json
{"quantization": {"quant_algo": null, "kv_cache_quant_algo": "FP8", "quantized_layers": {}}}
```
Nothing in the run says so. It has only ever been caught by someone
reading the exported `hf_quant_config.json` afterwards — most recently
on Step-3.7 ([NVBug 6518665](https://nvbugspro.nvidia.com/bug/6518665),
after a full 8×B200 calibration), and before that on MiniMax-M3, where
fused-expert detection skipped the experts and an experts-only recipe
matched nothing.
`mtq.quantize` now compares the config's intent against the outcome and
raises **before calibration**:
```
RuntimeError: The quantization config asks for weight quantization but no weight quantizer was
enabled, so nothing would be quantized (3 quantizer(s) inserted). These patterns matched no
weight quantizer:
*.experts.*weight_quantizer
Either the patterns do not match this architecture's module names (check the model-specific
recipes under modelopt_recipes/huggingface/<model_type>/), or the modules holding the weights
were never converted to quantized modules (an unsupported custom module, e.g. a
trust_remote_code MoE layout).
```
Scoped to avoid false positives:
- **Only configs that ask for weight quantization** are checked (an
entry with `enable` and `weight_quantizer` in its pattern), so
activation-only and KV-cache-only configs are unaffected.
- **Intent is read from each pattern's final entry**, since `quant_cfg`
entries apply in order: a pattern that is enabled and then disabled
later asks for nothing by the end.
- **Configs refining an already-quantized model** (weight quantizers
enabled by an earlier `mtq.quantize`) are left alone.
Matching goes through `conversion._match_quantizer` — the same matcher
`set_quantizer_by_cfg` used to apply the config — so "did this pattern
match anything?" is answered exactly as the applying code would. A local
`fnmatch` diverges on the two cases that matcher handles:
`SequentialQuantizer` modules (W4A8-style list-valued `cfg`) and
fused-experts names (`..._weight_quantizers.0` normalizing to
`..._weight_quantizer`).
### Usage
No API change. A config that would previously have produced an
unquantized checkpoint now raises:
```python
mtq.quantize(model, {"quant_cfg": [
{"quantizer_name": "*", "enable": False},
{"quantizer_name": "*.experts.*weight_quantizer", "cfg": {"num_bits": 8, "axis": 0}},
]}, forward_loop) # RuntimeError if the model has no `experts` modules
```
### Testing
Seven tests in `tests/unit/torch/quantization/test_quantize_cpu.py`, one
per branch of the guard: patterns matching nothing raise; an
activation-only config still runs; weight patterns disabled by a later
entry still run; enabled-then-retracted patterns still run;
`SequentialQuantizer` (list-valued `cfg`) and fused-experts quantizer
names count as matched; and the already-quantized refinement path is
exercised. Each was checked to be non-vacuous by removing the
corresponding branch and confirming exactly that test fails.
**One existing test changed.**
`tests/gpu/torch/export/test_fsdp2_export.py` parametrized over
`NVFP4_MLP_ONLY_CFG`, but its `SmallQKVModel` has no MLP — so that case
ran the FSDP2 paths against an *unquantized* model, and the new guard
reported it (4 GPU failures on the first CI run, all `quant_config6`;
`NVFP4_OMLP_ONLY_CFG` passed because that model does have `o_proj`). The
parametrization is dropped with a comment; `NVFP4_OMLP_ONLY_CFG` keeps
the scoped-recipe coverage. **If reviewers would rather not change that
test's meaning, the alternative is to downgrade the guard to a warning —
flagging it explicitly as a decision.**
I also swept every shipped `mtq.*_CFG` against `SmallQKVModel`: only the
four MLP/experts-scoped configs raise, and the other three
(`NVFP4_EXPERTS_ONLY_CFG`, `MXFP4_MLP_WEIGHT_ONLY_CFG`,
`NVFP4_MLP_WEIGHT_ONLY_CFG`) are used elsewhere only against real MoE
models (Qwen3-MoE, gpt-oss), so no other test is affected.
Ran locally after rebasing onto current `main` (torch 2.11, transformers
5.5.4): `tests/unit/torch/quantization` + `tests/unit/recipe` — 1271
passed, 7 skipped. Full `tests/unit` (minus onnx, and puzzletron which
needs `hydra`): 2676 passed, with 4 pre-existing
`test_quant_aware_conversion.py` failures that reproduce unchanged on
clean `main`. GPU tests were not run locally (no suitable GPU); the
FSDP2 change above is reasoned from the CI failure, not re-run.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ — deliberately. A config that
previously produced a `quant_algo: null` checkpoint now raises. Any such
run was already not doing what it claimed; the three scoping rules above
keep intentional non-weight quantization working.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ (Backward Breaking Changes)
- Did you get Claude approval on this PR?: ❌
### Additional Information
Pairs with #2202 (PTQ support for Step-3.7 MoE checkpoints), which fixes
the specific model that motivated this. Independent branches; either can
merge first.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Bug Fixes**
- Quantization now detects enabled weight-quantization patterns that do
not apply to any model weights and reports a clear validation error
before calibration.
- Broad wildcard patterns and nested quantizers are now handled
correctly.
- Overlapping patterns respect the final matching setting, including
later disabling rules.
- Existing quantized models can be refined using the parsed
configuration.
- Activation-only and explicitly disabled weight-quantization
configurations remain supported.
- Pipeline-parallel stages without targeted weights can bypass this
validation when configured to do so.
- **Documentation**
- Documented the process-wide override for bypassing unmatched
weight-quantizer validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
bd90a5ed51 |
[OMNIML-5774] Add BEVFormer ONNX PTQ and evaluation example (#2208)
### What does this PR do? Type of change: new example Adds an end-to-end BEVFormer-tiny ONNX PTQ example under `examples/onnx_ptq/bevformer` with: - The [BEVFormer Dockerfile](https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile) from pinned NVIDIA DL4AGX commit `9f7b291`, rather than a second container definition in this repository. - A mounted Model Optimizer checkout, the official DL4AGX TensorRT 10 source patch, and TensorRT plugin compilation at container runtime. - Ordered temporal calibration-data generation that propagates `prev_bev`, resets state at scene boundaries, computes CAN bus deltas, writes the exact requested sample count, and publishes output atomically. - INT8 and FP8 quantization with BEVFormer-specific calibration defaults, custom-plugin handling, `MatMul` exclusions, and FP16 fallback policy. - Strongly typed FP16, INT8, and FP8 TensorRT engine generation plus nuScenes evaluation instructions. - Shared temporary-ONNX-copy handling for BEVFormer and VoVNet quantization so shape inference cannot mutate the source model. - Focused CPU-only tests for temporal state, exact-count cleanup, quantization defaults, plugin configuration, and source-model preservation. The ONNX PTQ index continues to document the shared PETR/FAR3D containers separately and links to the BEVFormer guide. ### Usage Clone the pinned DL4AGX revision and build its BEVFormer image: ```bash git clone https://github.com/NVIDIA/DL4AGX.git /path/to/DL4AGX git -C /path/to/DL4AGX checkout --detach \ 9f7b29104c253d5bc68334e7b83b3eecb72d4572 docker build \ --build-arg TORCH_CUDA_ARCH_LIST=8.9 \ --file /path/to/DL4AGX/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile \ --tag modelopt-onnx-bevformer \ /path/to/DL4AGX/AV-Solutions/bevformer-int8-eq ``` The example command targets compute capability 8.9; use the deployment GPU's compute capability for another architecture. After exporting the model and generating temporal calibration data, quantize it with: ```bash python /opt/Model-Optimizer/examples/onnx_ptq/bevformer/quantize.py \ --onnx=/artifacts/bevformer_tiny_epoch_24_cp2_op13.onnx \ --calibration-dir=/artifacts/calibration \ --trt-plugins=/workspace/BEVFormer_tensorrt/TensorRT/lib/libtensorrt_ops.so \ --quantization-mode=fp8 \ --output=/artifacts/bevformer_tiny_epoch_24_cp2_op13.fp8.onnx ``` See `examples/onnx_ptq/bevformer/README.md` for dataset setup, plugin compilation, export, FP16 feedback-engine creation, temporal calibration, INT8/FP8 engine builds, and evaluation. ### Validation #### Current revision: three-sample smoke validation The Dockerfile at the pinned DL4AGX revision built successfully, and the current Model Optimizer checkout was mounted into it for the workflow below. The runtime audit confirmed TensorRT 10.14.1.48, CUDA 13.1, Torch 2.9, the TensorRT/CUDA/CPU execution providers, and plugin linkage. A fresh three-frame A/B/B smoke passed without calculating partial NDS or mAP: - Fresh Torch 2.9 export, AutoCast, and strongly typed FP16, INT8, and FP8 engine builds passed. - Temporal calibration published exactly three batches with `use_prev_bev=[0, 0, 1]`, zero state at both scene starts, recurrent feedback on the third frame, and the expected CAN bus deltas. - The source ONNX hash remained unchanged. INT8 contained 136 Q/DQ pairs with INT8 zero points; FP8 contained 127 Q/DQ pairs with FP8 zero points. - Every engine produced finite outputs and 300 non-empty decoded detections for each frame; maximum scores ranged from 0.9489 to 0.9608. CPU-only validation passed the 16-test focused example set and the full ONNX session (`649 passed, 1 skipped`). Full-diff pre-commit and the warning-as-error documentation build also passed. #### Historical full accuracy reference The following results were collected previously with TensorRT 10.14.1.48 on an NVIDIA RTX 6000 Ada Generation GPU, using 600 ordered calibration samples and all 6,019 nuScenes validation samples. They are reference results, not a completed full evaluation of the current revision. | Precision | NDS | mAP | | :-- | --: | --: | | FP16 | 0.3546 | 0.2515 | | INT8/FP16 | 0.3512 | 0.2505 | | FP8/FP16 | 0.3526 | 0.2489 | #### Performance Following the PETR and FAR3D convention, performance is reported only as speedup normalized to FP16. The archived full-validation engines were benchmarked with five interleaved TensorRT 10.14 trials on the same NVIDIA RTX 6000 Ada Generation GPU. Each trial used the `trtexec` GPU Compute Time median with data transfers disabled, CUDA Graphs enabled, spin-wait, a 1-second warmup, a 10-second measurement window, and one inference stream. | Precision | Speedup vs. FP16 | | :-- | --: | | FP16 | 1.00x | | INT8/FP16 | 1.87x | | FP8/FP16 | 1.18x | TODO: Investigate why FP8 delivers less speedup than INT8 for BEVFormer-tiny. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors). - Is this change backward compatible?: ✅ - If you copied code from another source or added a new PIP dependency, did you follow the guidance in `CONTRIBUTING.md`?: ✅ - Did you write the necessary tests?: ✅ - Did you update the [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ ### Additional information Reference workflow: [NVIDIA DL4AGX BEVFormer INT8 example](https://github.com/NVIDIA/DL4AGX/tree/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq). > 🤖 _Generated by Codex (AI agent)._ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an end-to-end BEVFormer 3D detection workflow for ONNX post-training quantization, including temporal calibration, INT8/FP8 quantization, TensorRT engine generation, and nuScenes evaluation. * Added command-line tools for preparing calibration data and quantizing BEVFormer models. * **Bug Fixes** * Prevented source ONNX models from being overwritten during quantization and improved temporary model cleanup. * Fixed FP8 export for BF16 models during real-weight compression. * **Documentation** * Added setup, usage, compatibility, and performance guidance for the BEVFormer workflow. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com> |
||
|
|
b80e164472 |
Carry a checkpoint's ModelOpt PTQ run into its eval's MLflow run (#2407)
### What does this PR do? Type of change: documentation (agent skill) Follow-up to #2374, which made a tracked `hf_ptq.py --mlflow` run leave `.experiment.json` in the checkpoint it writes, naming the MLflow run that quantized it. Nothing on the eval side read that file, so an eval of a quantized checkpoint recorded no link back to the quantization that produced it. The `evaluation` skill now reads it. **Step 3** `cat`s the file as soon as `checkpoint_path` is known; **Step 4** carries it into `export.mlflow`: | `.experiment.json` field | goes to | | --- | --- | | `experiment_name` | `export.mlflow.experiment_name`, verbatim | | `run_name` / `run_id` / `run_url` | tags `modelopt_run_name` / `modelopt_run_id` / `modelopt_run_url` | | `tracking_uri`, `experiment_id` | deliberately unmapped | The `modelopt_` prefix keeps them from reading as the eval's own run. Values are quoted, or an all-digit `run_id` (or a `run_name` like `20260910`) is YAML-coerced to an int or a date. **`tracking_uri` is deliberately not inherited**, and the consequence is documented rather than implied. Evals go to whatever `$MLFLOW_TRACKING_URI` names. When the PTQ tracked to a different server — the usual case, since `modelopttools:eval-config` points evals at `mlflow.frontier-evals` while `hf_ptq --mlflow` typically writes to `mlflow-modelopt` — the inherited name creates a *same-named, empty* experiment on the eval server, and `modelopt_run_url` is the only route back to the PTQ run. Evals of one checkpoint still group under a stable name. Both cases are spelled out in Step 4 so nobody goes looking for the PTQ run beside the eval. Also updated: `references/nel-next.md` (nel-next configures MLflow through its own `export_config.mlflow`) and `accessing-mlflow` (the `tags.modelopt_run_id` query that closes the loop — without it the tag is write-only). ### Usage ```bash cat "$CHECKPOINT_PATH"/.experiment.json # absent → name the experiment as usual ``` ```yaml export: mlflow: tracking_uri: ${oc.env:MLFLOW_TRACKING_URI} # NOT the file's experiment_name: alice/hf_ptq/Qwen3.8-27B-NVFP4 # verbatim from .experiment.json tags: modelopt_run_name: '20260910-175422' modelopt_run_id: '7bec239a3a154970b062f3024a5ff20e' modelopt_run_url: 'https://<modelopt-mlflow-server>/#/experiments/36/runs/7bec239a3a154970b062f3024a5ff20e' ``` Finding every eval of a checkpoint a given PTQ run produced: ```python MLflow:query_runs(experiment_id, "tags.modelopt_run_id = '<ptq_run_id>'") ``` ### Testing Docs-only, so it was tested by having an agent follow the new wording end to end and checking what NEL actually submitted. A GPQA Diamond config (single repeat) was generated against an NVFP4 checkpoint carrying a `.experiment.json`, dry-run clean, then submitted on a SLURM cluster. Verified in the `export_config.yml` heredoc **inside the submitted `run.sub`** — the file the export job actually consumes, not just the source YAML: ```yaml experiment_name: chenjiel/hf_ptq/Qwen3.8-27B-NVFP4 # inherited verbatim tags: modelopt_run_name: 20260910-000000-synthetic-test-fixture modelopt_run_id: 00000000fake0000fake0000fake0000 modelopt_run_url: https://<modelopt-mlflow-server>/#/experiments/999/runs/... tracking_uri: https://<eval-mlflow-server>/ # env var, not the file's ``` This confirmed the open question behind the change: NEL accepts arbitrary `modelopt_*` keys in `export.mlflow.tags`. The `.experiment.json` used was a synthetic fixture, since the checkpoints predate #2374. Repeated on a **second cluster with a second checkpoint** (different account, QoS-based scheduler, FP8 rather than NVFP4) and carried through to completion, so the result is confirmed as *stored by MLflow* rather than merely submitted. Canary scored `gpqa pass@1 symbolic_correct: 90.0` (card: 88.01) with `no_answer: 0.0`, then the auto-export wrote: ``` experiment chenjiel/hf_ptq/Qwen3.8-27B-FP8 (id 2106) run eval-<invocation>-ns_gpqa FINISHED modelopt_run_name 20260910-000000-synthetic-test-fixture modelopt_run_id 00000000fake0000fake0000fake0000 modelopt_run_url https://<modelopt-mlflow-server>/#/experiments/999/runs/... ``` Experiment 2106 did not exist on the eval server beforehand — the export created it. That is the shadow-experiment case this PR documents, observed rather than predicted. The same exercise corrected the first draft of the wording: it had claimed the rule makes a quantization and its evals "sit together on one server", which is false whenever the two servers differ. Confirmed by API — `chenjiel/hf_ptq/Qwen3.8-27B-NVFP4` returns `RESOURCE_DOES_NOT_EXIST` on `mlflow.frontier-evals`. Rewritten as described above. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — additive guidance; the absent-file path is unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — skill documentation; validated by a real submission, above. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ — agent-skill guidance, not a library/example feature. - Did you get Claude approval on this PR?: ❌ ### Additional Information Follow-up to #2374. Three adjacent problems were found while validating and deliberately **left out of scope** — happy to split them into their own PR: 1. `recipes/examples/example_eval.yaml` puts `sbatch_comment` under a top-level `cluster:` key, but the executor reads `cfg.execution.sbatch_comment` (`executors/slurm/executor.py:688`) and SKILL.md Step 1 says the same. As shipped the template silently drops the idle-GPU reaper exemption, so long evals get reaped. 2. Step 4's `execution.gres` bullet names the wrong key for `internal/slurm/<cluster>` configs, which supply `gpus_per_node` and no `gres`; following it literally emits a redundant flag. 3. Step 3's `max_new_tokens` rules can contradict each other — "take the card's highest" can equal `max_model_len`, leaving no room for the prompt. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Documentation** - Clarified MLflow evaluation configuration requirements, including literal experiment and sampling values, CPU partition settings, and ModelOpt provenance. - Documented validation and fallback behavior for unresolved `${...}` placeholders in provenance values. - Added guidance for querying cross-server evaluations by run ID and handling same-named local experiments. - Clarified that provenance files are optional, independent of quantization detection, and do not control deployment flags. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5f8c76e2aa |
Fix TEGroupedMLP quantizer checkpoint resharding (#2319)
### What does this PR do? Type of change: Bug fix for https://github.com/NVIDIA/Model-Optimizer/issues/2209 Fix TEGroupedMLP per-expert weight quantizer checkpoint resharding. `TEGroupedMLP` now saves its per-expert quantizer state as singleton local shards, allowing the distributed checkpoint format to retain each expert's global identity. Restore also initializes scalar `_amax` placeholders after ModelOpt extra-state restoration so distributed checkpoint loading can populate quantizer state for experts that move between ranks. This fixes restoring quantized TEGroupedMLP checkpoints across expert-parallel and tensor-parallel topology changes. Previously there was a bug that had two parts 1. TEGroupedMLP did not mark its per-expert quantizer state as singleton_local_shards. That meant the scalar weight_quantizer.<expert>._amax state was not saved with the same globally unique expert identity as the grouped-expert weights, so DCP could not reliably redistribute it across EP layouts. 2. During restore, ModelOpt’s extra-state restoration can leave _amax absent for experts that were not local on the checkpoint’s saving rank. The subsequent distributed checkpoint load then had no destination tensor to populate. ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing - `ruff format`, `ruff check`, `mypy`, `bandit`, and repository pre-commit hooks - Focused GPU regression: ```bash python3 -m pytest tests/gpu_megatron/torch/quantization/plugins/test_megatron.py \ -k te_grouped_sharded_state_dict_reshard -v Replaced the prior metadata-only TEGroupedMLP sharded-state test with an end-to-end distributed-checkpoint save/restore regression test. The new test: - Quantizes a TEGroupedMLP with per-expert NVFP4 weight quantizers. - Assigns each local expert a distinct, deterministic `_amax` based on its global expert index. - Saves both the model distributed checkpoint and sharded ModelOpt state. - Rebuilds the model under a different TP/EP topology. - Restores ModelOpt state, loads the distributed checkpoint, and verifies each target-local expert received the expected global-expert `_amax`. The parameterized test covers: - EP=2 -> EP=1 - EP=1 -> EP=2 - TP=1 -> TP=2 - TP=2 -> TP=1 ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Very short summary of changes only for new features, backward breaking changes, deprecations, or fixes for critical bugs present in previous releases. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Improved checkpoint restoration for grouped quantizers by initializing missing quantization statistics with compatible shapes. - Improved restoration across supported grouped quantizer configurations, including sequential groups and parallel checkpoint layouts. - Extra module state is now finalized through supported post-load callbacks when available. - Preserved populated quantized output-layer state during checkpoint operations while removing empty placeholders. - **Tests** - Expanded checkpoint resharding coverage across tensor- and expert-parallel configurations. - Added coverage for disabled, dynamic, and other grouped quantizer scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> Signed-off-by: Jenny Chen <jennifchen@nvidia.com> |
||
|
|
c5d1065331 |
Modeling Lib: per-model architecture kick off with spec registration (#1828)
### What does this PR do? **Type of change:** New feature — per-model infrastructure (with behavior changes, see below) Starts `modelopt/torch/models/`: somewhere for ModelOpt to keep what it knows about a model, keyed by HF model type (`config.model_type`), one package per type (`<model_type>/specs.py`). Importing the package registers every spec; consumers call `get_spec(model_type)` / `match_moe_block(module, model_type)` and read the fields. The point is the infrastructure, not the migration. Today ModelOpt's per-model knowledge is scattered across if/elif chains in whichever subsystem happened to need it, so the same fact gets restated per subsystem and drifts. A model's MoE block classes, its expert projection naming, which of its norms store `w - 1` — these are facts about the *model*, and several subsystems want them. This PR gives them one home and one lookup. It is a kick-off, so it is deliberately narrow: specs plus the first consumer. **Export is that first consumer**, which is why most of the changed lines are on the export side — not because this is an export refactor. Quantization, speculative decoding and the rest keep their own tables for now; the per-model directory is where their sections land later, and `specs.py` is just the first thing in it. | Data the first consumer moved in | Spec field | Consumers | |---|---|---| | MoE block classes and expert linear naming | `MoESpec` | `get_expert_linear_names`, `get_experts_list`, `is_moe`, `sync_moe_gate_up_amax` | | grouped expert-export support set | `ExportSpec.grouped_expert_export` | `get_experts_list` | | AWQ `pre_quant_scale` fusion rules | `ExportSpec.pqs_fuse_rules` | `fuse_prequant_to_linear` | | weight-plus-one norm class names | `ExportSpec.weight_plus_one_norm_names` | `_layernorm_uses_weight_plus_one` | `ModelSpec` holds each concern as a separate, optional section, and the split is what makes it extensible: **topic sections** hold architecture facts any subsystem can read (`MoESpec` — block classes, expert projection naming, how the experts are stored), **subsystem sections** hold one subsystem's own data and policy (`ExportSpec` — AWQ fusion rules, weight-plus-one norms, whether grouped expert export is validated for this model). A new subsystem adds a section rather than a table. A section is `None` when the model has nothing to say about it, so a dense model carries `moe_spec=None` rather than an empty one. Two further fields record where a model's classes come from at all: `modeling_source` (`transformers` or `remote_code`) and `min_transformers_version`. ### Behavior changes **1. Lookup key: `type(root_model).__name__.lower()` → `config.model_type`.** The old key changes after `quantize` wraps the model, forcing substring matching; `model_type` is stable, so `ExportContext` carries it and lookups are exact. Matching is also tightened to case-insensitive **exact** names against the module's **MRO**, so quantized subclasses still match via their base class without substring false positives. **2. ⚠️ `get_expert_linear_names` raises instead of guessing.** Unmatched MoE blocks previously fell through to `["w1", "w2", "w3"]` (Mixtral naming); they now raise `NotImplementedError` telling you to register a `ModelSpec`. The blast radius depends on the transformers release, because only *iterable* expert layouts consult the spec — a fused expert container is resolved by the structural first-projection check and never asks for naming. Measured against both ends of the support matrix: | transformers | Unregistered MoE families `is_moe` admits | Reach the spec lookup | Actually affected | |---|---|---|---| | 5.14.1 (`tf_latest`) | 19 | **0** — all fused | none | | 4.57.6 (`tf_min`) | 9 | 7 iterable | **2** | On 4.57, five of the seven (`ernie4_5_moe`, `flex_olmo`, `jamba`, `olmoe`, `qwen3_omni_moe`) name their experts `gate_proj`/`up_proj`/`down_proj`, so the `w1` fallback already failed with `AttributeError` on `main` — for them this trades an unhelpful error for one that names the fix. The other two, **`minimax` and `phimoe`**, are Mixtral-derived with `w1`/`w2`/`w3` experts and did export on `main`, so they are now **registered** rather than left to regress. Both only take the per-expert path on transformers 4; on 5 their experts are fused (`MiniMaxExperts`, `PhimoeExperts`) and the structural check resolves them, which their specs decline to answer for. **So no model that exported correctly before this change regresses.** This is backward-incompatible and has a `CHANGELOG` entry. **3. DeepSeek-V3 and V4 are now registered, and newly visible to the MoE path.** Neither was reachable before: `DeepseekV3MoE` (like `DeepseekV2Moe` and `DeepseekV32MoE`) is invisible to `is_moe` — its class name does not end in `SparseMoeBlock` and it calls its router `gate`, so neither the name test nor the structural `router`+`experts` test matches. `DeepseekV4SparseMoeBlock` was detected by name but had no spec to resolve expert naming from. Registering `block_names` is what puts V3 on the MoE path at all, so this is added coverage rather than a fix to existing behavior. `deepseek` stays registered alongside them: it describes the remote-code `DeepseekMoE` block of DeepSeek-MoE/V1, and its spec is the only thing making that block detectable. **4. Fused expert containers are skipped in `get_experts_list` instead of crashing.** transformers 5 replaced several iterable expert `ModuleList`s with a single module holding 3-D parameters, while the specs still describe the transformers 4 iterable layout — a spec cannot tell the two apart, only the module can. The AWQ/SVDQuant resmooth pass in `requantize_resmooth_fused_llm_layers` reached `len(module.experts)` on a module with no `__len__` and died with `TypeError` mid-export. **Mixtral already hits this on `main`**; registering `deepseek_v3` only made an existing bug easier to reach. The skip is scoped to specs that claim iterable experts, so a layout the spec calls unsupported (DBRX, whose per-expert linears live under `experts.mlp`) still fails loudly rather than silently dropping resmoothing. Everything else is behavior-preserving, pinned by tests: - **`sync_moe_gate_up_amax` keeps generic coverage.** Quantization admits MoE blocks *structurally*, so unregistered families (Olmoe, Jamba, MiniMax…) reach it. With no spec it falls back to every declared gate/up naming — what it did pre-registry. Skipping them would leave the fused `gate_up_proj` halves on inconsistent `weight_scale_2`. - **`ExportSpec.grouped_expert_export` matches the legacy `get_experts_list` support set**, except for `deepseek_v3`/`deepseek_v4`, which are new and had no legacy behavior to match. Note `qwen3_5_moe`: legacy keyed off the root class name and `"qwen3_5moeforcausallm"` matched none of its qwen substrings (the `_5` breaks `qwen3moeforcausallm`), so it raised — the spec keeps `False` to match. Enabling it belongs in its own PR. - **`is_moe` consults the model's own spec first**, before the generic name and structural fallbacks, so per-model data always wins. The three checks are or-ed, so this is a no-op; the same reordering is deliberately *not* applied to `get_expert_linear_names`, where the structural fused-experts check must stay first. - **Two values are corrected, not moved.** `gpt_oss` declared `block_names=("GptOssMoE",)`, which matches no real module (transformers names it `GptOssMLP`); naming still resolved via the single-naming shortcut, so this was invisible. Now `("GptOssMLP", "GptOssMoE")`. **5. ⚠️ DBRX expert input amax is now populated, changing its exported scales.** The second corrected value, called out separately because it has a numerical consequence. The legacy branch keyed on `DBRXMoeSparseMoeBlock`, a class transformers does not define — it names the block `DbrxFFN`. So DBRX fell through to the `w1`/`w2`/`w3` default, every `hasattr(experts_mlp, linear_name)` in `_prepare_dbrx_experts` evaluated `False`, and the handler wrote no expert input amax at all. `dbrx/specs.py` declares the real names (`w1_linear`, `w2_linear`, `v1_linear`), so the handler now does what it was written to do. Anyone diffing a re-exported DBRX checkpoint against an older one will see different expert activation scales. That is the fix working, not a regression from the refactor. No `CHANGELOG` entry: DBRX is not listed as a supported export target in `docs/source/deployment/` or `examples/hf_ptq/README.md`. ### Scope One consumer wired up: the unified HF export path. The TRT-LLM builders now live in `modelopt/torch/export/trtllm/` after #2365, and are deprecated. This PR changes one import line there — `is_moe` moved to `modelopt/torch/models/moe.py`, so importing it from `..layer_utils` no longer resolves — and nothing else. Their `is_moe()` calls still pass no `model_type` and fall back to the all-specs search, reproducing today's behavior; their hardcoded tables, including a second copy of the weight-plus-one norm names, stay as they are and migrate whenever that package does. Worth noting that #2365 landed the same boundary from the other direction: the slimmed `export/layer_utils.py` is now documented as "module-shape predicates and MoE quantizer helpers shared by every export backend", which is exactly the set of functions this PR rewires. The Megatron path has its own per-family registry and is untouched; folding it in is a later question, not this PR's. `is_moe` itself moved out of `modelopt/torch/export/layer_utils.py` into `modelopt/torch/models/moe.py` — whether a module is an MoE block is a modeling question, not an export one, and it is the first piece of shared modeling logic to follow the specs into the new package. ### Relationship to #1939 (why a second registry?) Orthogonal layers: #1939's `ExportModuleRegistry` dispatches on **module structure** (*which handler runs?*), this registry resolves **family data** (*what are its values?*). Exporting a `QuantMixtralSparseMoeBlock`, #1939 picks the shared iterable-experts handler (also serving Qwen/DeepSeek/Gemma4); inside, `get_expert_linear_names` resolves the `mixtral` spec to `("w1", "w2", "w3")`. One handler serves many families, one spec serves many handlers. Sharing the matcher machinery is a planned follow-up. ### Usage ```python # modelopt/torch/models/qwen3_moe/specs.py — adding a model needs no engine edits register( ModelSpec( model_type="qwen3_moe", min_transformers_version="4.57", moe_spec=MoESpec( block_names=("Qwen3MoeSparseMoeBlock",), expert_linear_names=("gate_proj", "down_proj", "up_proj"), gate_up_pair=("gate_proj", "up_proj"), ), export_spec=ExportSpec( grouped_expert_export=True, pqs_fuse_rules=( (("Qwen3MoeAttention",), "v_proj", "o_proj"), (("Qwen3MoeMLP",), "up_proj", "down_proj"), ), ), ) ) ``` One layout per model. `block_names` is a tuple, so a model whose MoE appears under several class names is covered as long as they share a layout (`gpt_oss`'s `GptOssMLP`/`GptOssMoE`). `gemma4_text` imports its section from `gemma4` rather than restating it. ### Testing `tests/unit/torch/export/` — **202 passed**; `tests/unit/torch/quantization/` — **903 passed**. The `test_export_diffusers.py` collection error and 6 `test_quant_aware_conversion.py` failures are pre-existing and reproduce on untouched `main`, as does the one failing pre-commit hook (`generate-arguments-md`, no torch in its venv). `tests/unit/torch/models/test_model_specs.py`: registry matching (MRO / quantized classes / model-type scoping); every registered MoE section's `(block_names, expert_linear_names, fused_expert_names, gate_up_pair)` as an exhaustive table that fails if a spec is added without a row; the exhaustive `grouped_expert_export` support set; the structural fused-expert shortcut and the spec's precedence over it; the fused-container skip in `get_experts_list`; the `sync_moe_gate_up_amax` fallback; and legacy-equivalence of the aggregated `pqs_fuse_rules` / `gate_up_pairs` / `weight_plus_one_norm_names`. Mutation-checked — flipping a flag, typo-ing a block name, or swapping the precedence each fail. `tests/unit/torch/models/test_specs_vs_transformers.py` validates the specs against the *installed* transformers rather than a mirrored table: every registered block class must name a real class from its `min_transformers_version` on, and every `remote_code` model must still be absent. It names no model — which specs to check, and which cannot be, both come from the registry. > The run counts above were measured before the follow-up commits in this branch (spec > precedence, the policy split, the `MoESpec`/`MoELayout` merge) and have not been > re-measured; CI on the current head is the authority. **Local per-model-type E2E sweep** (transformers 5.3.0, tiny models from config, run against this branch and base `main`; not checked in): | model_type | real block class | this branch | base `main` | |---|---|---|---| | `qwen3_moe` / `qwen2_moe` / `qwen3_next` | `Qwen*SparseMoeBlock` | `gate/down/up` | same | | `mixtral` | `MixtralSparseMoeBlock` | `w1/w2/w3` | same | | `nemotron_h` | `NemotronHMoE` | `up/down` | same | | `gpt_oss` | `GptOssMLP` | `gate_up/down` | ❌ `w1/w2/w3` | | `deepseek_v3` | `DeepseekV3MoE` | raises | ❌ `w1/w2/w3` | Five types identical to base; both differences are base bugs the registry surfaces. `dbrx` could not be built on transformers 5.3.0 (`DbrxAttentionConfig` missing `rope_theta`). `sync_moe_gate_up_amax` was separately checked against base across registered, unregistered, and no-config models — identical in all cases. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ❌ — `get_expert_linear_names` no longer falls back to `w1/w2/w3`; see "Behavior changes" #2. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no copied code, no new dependency (stdlib only). - Did you write any new necessary tests?: ✅ — `tests/unit/torch/export/test_model_specs.py`. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ — happy to add an entry given the backward-incompatible item. - Did you get Claude approval on this PR?: ✅ — run; its CRITICAL finding on `sync_moe_gate_up_amax` is fixed above. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added model-aware export support across a broad range of MoE and dense architectures, including newer DeepSeek, Gemma, Qwen, Nemotron, Arctic, DBRX, GPT-OSS, Llama, and Mixtral variants. - Improved quantization and export handling for model-specific expert layouts, projection fusion, normalization, and gate/up synchronization. - Added automatic architecture recognition using Hugging Face model metadata. - **Bug Fixes** - Unsupported or ambiguous model layouts now produce explicit errors instead of applying potentially incorrect defaults. - **Documentation** - Added documentation describing model specifications and supported architecture metadata. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> Co-authored-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
dbe28e1e05 |
Document the MSE calibration API (#2405)
### What does this PR do? Type of change: documentation Expose `mse_calibrate` through `model_calib.__all__` so Sphinx autosummary includes the existing MSE calibration API on the generated `model_calib` reference page. The documentation configuration honors each module's curated `__all__` surface. Although `mse_calibrate` was implemented and used by the calibration dispatcher, it was missing from that surface and was therefore filtered out during API generation. ### Usage ```python from modelopt.torch.quantization.model_calib import mse_calibrate ``` ### Testing - `pre-commit run --files modelopt/torch/quantization/model_calib.py` - `git diff --check -- modelopt/torch/quantization/model_calib.py` - Generated the recursive autosummary API tree using the repository Sphinx configuration and module template. - Verified the generated RST contains `mse_calibrate`. - Rendered the focused module page and verified its function-table link, anchor, signature, and docstring. The canonical full build was unavailable in the active environment because the configured `shibuya` theme is not installed. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — the existing documentation build directly exercises this declarative autosummary contract. - Did you update Changelog?: N/A — this is a documentation-visibility repair for an existing API. - Did you get Claude approval on this PR?: N/A ### Additional Information No source implementation or calibration behavior changed. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Made the MSE calibration capability publicly available for quantization workflows. - **Chores** - Increased documentation build time limits to improve reliability for longer-running builds. - Increased multi-version test time limits to better accommodate extended test runs. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
a1bcda4727 |
Fix protobuf size-check failures in ONNX deployment (#2403)
### What does this PR do? Type of change: Bug fix:6701737 The ONNX deployment path assumed that ModelProto.ByteSize() would always return a valid size. With newer protobuf versions, querying the size of a model exceeding the protobuf serialization limit can itself raise EncodeError: Failed to serialize proto. Replaced both direct size checks with the existing is_model_too_large_for_protobuf() helper. This helper handles size-query failures conservatively and checks the protobuf size limit: Shape inference now selects the external-data/file-based path when ByteSize() fails or the model is too large. Metadata creation uses the same safe check instead of raising another serialization error. The unused TWO_GB constant was also removed. ### Usage ``` python examples/diffusers/quantization/diffusion_trt.py --model flux-dev --benchmark --skip-image ``` ### Testing the above test case pass on B100 ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved ONNX model size detection during shape inference and export processing. * Ensured large models consistently use the appropriate external-data handling path. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Noey Yang <174223378+noeyy-mino@users.noreply.github.com> |
||
|
|
5d2d5a5d15 |
deprecate trtllm-build in weight_sparsity (#2371)
### What does this PR do?
Type of change: export PTS/finetuned model to Hugging Face checkpoint,
then replace trtllm-build with trtllm-serve
Renamed export_trtllm_ckpt.py to export_hf_ckpt.py.
Replaced the legacy export_tensorrt_llm_checkpoint() flow with
export_hf_checkpoint().
Fix bug: 5823190
<!-- Details about the change. -->
### Usage
```
python examples/llm_sparsity/weight_sparsity/hf_pts.py --model_name_or_path Llama-3.1-8B-Instruct --device cuda --model_max_length 1024 --dtype fp16 --sparsity_fmt sparsegpt --calib_size 128 --output_dir Llama-3.1-8B-Instruct_pts
python examples/llm_sparsity/weight_sparsity/export_hf_ckpt.py --model_name_or_path Llama-3.1-8B-Instruct --model_max_length 1024 --dtype fp16 --modelopt_restore_path Llama-3.1-8B-Instruct_pts/pts_modelopt_state.pth --output_dir Llama-3.1-8B-Instruct_pts/trtllm/ckpt_pts
trtllm-serve Llama-3.1-8B-Instruct_pts/trtllm/ckpt_pts \
--tp_size 1 \
--pp_size 1 \
--host 0.0.0.0 \
--port 8000
```
### Testing
PTS and SAT tested
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A
### Additional Information
N/A
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Updated sparsity example instructions to export Hugging Face
checkpoints and serve models with `trtllm-serve`.
* Documented tensor and pipeline parallelism, host and port settings,
and the OpenAI-compatible chat completions endpoint.
* Corrected the PTS model restoration path.
* **Bug Fixes**
* Model export now saves the tokenizer alongside the checkpoint.
* Model length configuration is interpreted as an integer.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Noey Yang <174223378+noeyy-mino@users.noreply.github.com>
|
||
|
|
59d93af064 |
[OMNIML-5570, OMNIML-5569] 1/2 Add layer-wise KV-cache AutoQuant with forward KL (#2272)
### What does this PR do?
Type of change: new feature.
Adds standalone layer-wise KV-cache AutoQuantize through the existing
public
`mtq.auto_quantize` API:
- dispatches KV search with
`constraints={"effective_bits": ..., "cost_model": "kv_cache"}` and
forward-KL
sensitivity;
- selects one supported K/V format for every eligible causal-attention
layer;
- supports persistent/exportable FP8 K/V, NVFP4 K/V, and FP8-K/NVFP4-V
candidates;
- solves a K/V-width- and scale-storage-aware additive recipe with the
existing
PuLP-backed constrained solver;
- uses `BaseSearcher` lifecycle and safe checkpoint restore/save
machinery;
- preserves existing non-KV execution while isolating K/V candidate
calibration;
- returns standard AutoQuantize state that can be re-solved at another
KV budget;
- produces a complete KV-only replay config that disables every non-KV
quantizer;
- saves JSON-safe sensitivity metadata and the exact selected layer
mapping; and
- invokes the public API from `examples/hf_ptq/hf_ptq.py` through a
standalone
calibration-free recipe.
The implementation is architecture-driven. Plain and
conditional-generation Qwen
causal attention is supported, VLM vision attention is excluded through
the existing
language-model extraction boundary, hybrid full-attention mixers are
discovered through
their paired K/V quantizers, and nonattention/Mamba modules remain
outside the search.
Ambiguous language-model roots, unsupported distributed execution,
structural
algorithms, invalid storage declarations, nonpersistent scales, and
unsupported K/V
pairs fail closed.
KV-only unified HF exports leave weight-quantization fields unset.
Uniform all-FP8 or
all-NVFP4 selections retain their legacy KV scheme while also carrying
the complete
`kv_cache_quantized_layers` map and schema version; genuinely
layer-mixed selections use
the KV-side `MIXED_PRECISION` marker plus the same map. This keeps
weight-loader metadata
accurate and prevents disabled vision attention from making uniform
language-model KV
quantization appear partially quantized.
GEMM PTQ/AutoQuantize followed by KV AutoQuantize is intentionally
excluded and proposed
separately in stacked PR #2273.
### Why KV search has a dedicated backend
The user-facing entry point remains `mtq.auto_quantize`; no separate
public KV search API
is introduced. `AutoQuantizeKVSearcher` extends `BaseSearcher` and
reuses its reset,
checkpoint load/save, and search lifecycle, along with existing Pydantic
configuration,
calibration, safe checkpoint I/O, and PuLP-backed selection utilities.
The backend remains KV-specific because a decision owns paired K/V
quantizers on one
attention layer, its cost depends on separate K/V widths and data/scale
storage, BF16 is a
scoring reference but not a deployable solver choice, and the
optimization objective is
additive isolated forward KL under a KV-storage constraint. These
contracts do not match
the weight-domain hparam grouping, parameter-count cost, or
threshold-selection behavior
of the existing weight AutoQuant searchers. Keeping the specialization
behind the shared
API avoids changing established weight-search solver and scoring
behavior.
### Usage
```bash
python examples/hf_ptq/hf_ptq.py \
--pyt_ckpt_path Qwen/Qwen3.8-27B \
--recipe general/auto_quantize/kv_fp8_nvfp4_cast_kl_div_at_5p4bits \
--auto_quantize_checkpoint /path/to/kv_autoquant.pth \
--export_path /path/to/qwen3.8-27b-mixed-kv
```
The search checkpoint is compatible only with the same model,
eligible-layer geometry,
candidate configurations, and scoring setup. Use a distinct checkpoint
path after any of
those inputs change.
KV-cache AutoQuantize rejects `--use_fsdp2` before model loading because
its sensitivity
scoring, selection, and checkpoint writes are single-process. Existing
weight
AutoQuantize retains its previous experimental FSDP2 warning and
behavior.
### Testing
- Focused coverage exercises candidate validation/calibration, paired
K/V scoring and
storage accounting, solving, checkpoint resume, failure atomicity,
disabled layers,
fresh-model replay, Qwen/VLM/hybrid boundaries, JSON-safe reports, and
unified export.
- Uniform FP8/NVFP4 KV-only exports retain the legacy KV scheme and
complete layer map
without claiming a weight algorithm; disabled VLM vision attention is
excluded from
causal-KV eligibility.
- The shipped recipe runs end to end on a tiny offline Qwen fixture and
preserves
exportable scale state.
- After merging current `main`: 432 focused recipe/KV/export/hf_ptq
tests passed, with one
unrelated optional-dependency skip; changed-file pre-commit hooks
passed.
### Deployment gate
The producer schema is covered here. Runtime consumption of
`kv_cache_quantized_layers` is tracked in vLLM PR
https://github.com/vllm-project/vllm/pull/52813. Do not treat a produced
checkpoint as
runtime-supported until that consumer lands and the target K/V kernels
are available.
### Before your PR is "Ready for review"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow
guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅
### Additional information
- This is split from the combined ground-truth implementation in draft
PR #2211 to reduce
review scope; composition is isolated in #2273.
- The standalone core tree contains no composed GEMM→KV recipe schema or
orchestration.
- No model-name checks, checkpoint-specific layer lists, campaign data
contracts, cluster
launch logic, or runtime-kernel implementations are included.
---------
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
|
||
|
|
d69e93a72b |
Record the MLflow run that produced a checkpoint in .experiment.json (#2374)
### What does this PR do?
Type of change: new feature
A tracked `hf_ptq` run already tags itself with the checkpoint it writes
(`checkpoint_path`), so a run can be followed to its output. The reverse
was missing: given a checkpoint on disk, there was no way to find the
run that quantized it without searching the tracking server by path.
A tracked run now writes `.experiment.json` into `--export_path` naming
the experiment, the MLflow run id and the run URL, and uploads the same
bytes as the `experiment.json` artifact so a downloaded artifact set is
self-describing. `MlflowRunLogger` gains a `run_info` property carrying
that identity, with the tracking URI credential-masked the way `run_url`
already was.
Two deliberate behaviours:
- **Written from a `finally`**, so a run that crashes after export still
leaves the pointer behind.
- **Skipped when the export directory is absent** — a run that exported
nothing has nowhere to put it, and creating the directory would suggest
a checkpoint that does not exist. The artifact is still uploaded in that
case, so a failed run is traceable from the server side.
A failed local write warns and continues rather than failing the job,
consistent with the rest of the MLflow path. Only the main rank writes,
since the logger is inert on other ranks.
### Usage
```bash
python hf_ptq.py --pyt_ckpt_path Qwen/Qwen3.5-0.8B --qformat fp8 \
--export_path /tmp/qwen35-fp8 --mlflow https://<your-mlflow-server>
```
```console
$ cat /tmp/qwen35-fp8/.experiment.json
{
"tracking_uri": "https://<your-mlflow-server>",
"experiment_name": "alice/hf_ptq/Qwen3.5-0.8B-fp8",
"experiment_id": "36",
"run_id": "7bec239a3a154970b062f3024a5ff20e",
"run_name": "20260910-175422",
"run_url": "https://<your-mlflow-server>/#/experiments/36/runs/7bec239a3a154970b062f3024a5ff20e"
}
```
```python
# checkpoint -> run
import json, mlflow
info = json.load(open("/tmp/qwen35-fp8/.experiment.json"))
mlflow.set_tracking_uri(info["tracking_uri"])
run = mlflow.get_run(info["run_id"])
```
### Testing
**Unit** — `tests/unit/torch/utils/test_mlflow.py` (61 passed):
`run_info` contents before/after the run opens, the defaulted run name
being reported rather than left blank, and credential masking of the
tracking URI.
**Example** — `tests/examples/hf_ptq/test_hf_ptq_args.py` (27 passed):
the file landing in the checkpoint and on the server with identical
content, the failed-run path, the no-export path, and untracked runs
writing nothing.
**Real runs**, 1x H200, `Qwen3.5-0.8B` FP8 PTQ,
`tensorrt-llm/release:1.3.0rc26`:
- Against a local MLflow server — checkpoint copy and uploaded artifact
byte-identical; artifacts on the run were `command.txt`,
`experiment.json`, `logs/hf_ptq.log`, `summary/quant_summary.txt`,
`version.txt`.
- Against the internal `mlflow-modelopt` server (experiment
`chenjiel/hf_ptq/Qwen3.5-0.8B-fp8`, run
`7bec239a3a154970b062f3024a5ff20e`) — same result, confirming artifact
upload against a real backend. Reading `.experiment.json` back and
calling `mlflow.get_run(run_id)` resolved to `FINISHED` with
`checkpoint_path` pointing at the export directory.
- Crash path exercised for real when a first attempt died on a gated
calibration dataset: no export directory created, `experiment.json`
still uploaded, run closed `FAILED`.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — new entry under `*Misc*` in the open 0.48.0 section, matching where
the MLflow entries sit in 0.47.0.
- Did you get Claude approval on this PR?: ❌
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Exported checkpoints now record experiment and run traceability
metadata in `.experiment.json`.
* Checkpoint metadata is uploaded with opened MLflow runs, including
runs where export fails.
* Active MLflow run details—including identifiers, resolved run name,
URL, and tracking server—are available with credentials redacted.
* **Bug Fixes**
* Improved handling of failed, untracked, and pre-existing exports to
prevent inherited metadata pointers.
* **Documentation**
* Updated MLflow integration guidance and changelog information for
checkpoint metadata.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
a74054ab2b |
Let callers add MLflow tags to a fakequant serve's run (#2364)
### What does this PR do?
Type of change: new feature
The quantization run records what this library can see — the model, the
checkpoint, the vLLM and ModelOpt versions — but nothing about the
harness that launched it. A downstream tool that wants its own revision,
a sweep id, or a ticket number on the run has no way to put it there
today:
- `_run_tags()` returns a fixed dict
- `quant_config` (which becomes the run's params) is a hardcoded set of
`QUANT_*` variables
- MLflow itself has no environment variable for arbitrary tags
`MODELOPT_MLFLOW_EXTRA_TAGS` takes comma-separated `key=value` pairs and
merges them into the run's tags.
Two details worth a reviewer's attention:
**It joins `MLFLOW_ENV_VARS`.** A Ray-backed serve receives only the
variables named there, and the tracker runs in the rank-0 worker —
omitting it would make the feature silently do nothing under Ray.
**Caller tags are merged first**, so the library's own keys (`tool`,
`model`, `checkpoint_path`, `vllm_version`) are written over them and
keep describing the run truthfully whatever a caller sends.
`key=value` rather than JSON, learned from a live run: the variable
reaches the worker through a shell `export VAR="..."`, and JSON's own
double quotes terminate that quoting —
```
export MODELOPT_MLFLOW_EXTRA_TAGS_732b_DEPLOYMENT="{"internal_version": "4d8c"}"
```
arrived as `{`. A quote-free format survives verbatim and needs no
`json` import or exception handling. Splitting on the first `=` keeps
values that contain one, such as a URL with a query string.
### Usage
```bash
export MODELOPT_MLFLOW_EXTRA_TAGS="modelopt_internal_version=49fa29d5,sweep=kv-study"
python3 vllm_serve_fakequant.py "$MODEL" --mlflow https://your-mlflow-server/ ...
```
### Testing
Unit-level, over the helper: unset and empty variable, one and several
pairs, surrounding whitespace, an empty value, an entry with no `=`, a
trailing comma, and a value containing `=`. None raise; malformed
entries warn and are skipped.
End to end on a real fakequant serve (Nemotron-3-Nano-30B-A3B BF16,
`NVFP4_DEFAULT_CFG`, TP=8, Ray executor, vLLM 0.15, SLURM):
```
modelopt_internal_version '49fa29d5'
modelopt_version '0.47.0rc0.post32+gd38ed5ead'
git_sha 'd38ed5ead'
quant_cfg 'NVFP4_DEFAULT_CFG'
```
The tag was written by the `RayWorkerWrapper` process, which exercises
the whole path — env var → shell export → `--container-env` → raylet →
Ray actor → `_run_tags` — and confirms the `MLFLOW_ENV_VARS` entry is
doing its job. Also verified that the emitted payload survives a shell
export round-trip unchanged.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ <!--- Additive; with the
variable unset the tags are exactly as before. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ <!--- Verified manually as
above; there is no existing test module for vllm_mlflow_utils. Happy to
add one if you would like it. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <!--- Small additive feature in an example; tell me if it warrants an
entry. -->
- Did you get Claude approval on this PR?: ❌
### Additional Information
Consumed by Model-Optimizer-Internal MR !141/!147, which sets the
variable so a fakequant eval records the same harness commit on both its
quantization run and its evaluation-score run.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
7f7c46d820 |
[6508436] Fix BF16 FP8 ONNX export (#2314)
### What does this PR do?
Type of change: Bug fix
Fix FP8 ONNX export for BF16 models during real-weight compression
without changing the public API or the `weights_dtype="fp32"` default.
The FP8 exporter preserves BF16 initializer bits when bridging
GraphSurgeon NumPy arrays to Torch, widens BF16 values exactly to FP32
for normalization, and leaves existing FP16/FP32 handling unchanged.
Conv scales and dequantized outputs retain the source dtype, and scales
round upward when needed so serialized values cannot cause FP8 overflow.
`weights_dtype="bf16"` is accepted as a no-op only for FP8-only models
whose floating parameters are all BF16. Registered buffers do not affect
this weight-focused decision and may preserve higher-precision regions
in the exported graph. Unsupported BF16 FP8-to-FP16 and FP32 or
mixed-parameter-to-BF16 conversions are rejected with `ValueError`
before temporary export paths are created. A narrow GraphSurgeon fix
preserves integer BF16 value-info dtypes.
### Usage
```python
onnx_bytes, metadata = get_onnx_bytes_and_metadata(
quantized_fp8_model,
(sample_input,),
weights_dtype="bf16",
onnx_opset=23,
)
```
### Testing
- Seven focused CPU regressions passed: BF16 QDQ compression and integer
dtype handling, BF16-to-BF16 and FP32-to-FP16 Conv/Linear export, and
four unsupported-conversion cases.
- QDQ utilities: 31 passed; pytest 2.25s, wall 29.88s.
- FP8 MHA exporter: 6 passed; pytest 2.05s, wall 32.71s.
- Torch deploy utilities: 51 passed; pytest 8.94s, wall 25.74s.
- Torch ONNX CPU export: 36 passed; pytest 4.85s, wall 32.38s.
- Changed-file pre-commit hooks: all passed; wall 8.07s.
- Exact-head FP8 BF16 GPU workflow at `f21d62a`: exit code 0; ONNX
checker passed; 6 FP8 initializers, 3 native `DequantizeLinear` nodes,
and 12 BF16 initializers.
- Refreshed GitHub CI at `f21d62a`: 50 passed and 1 skipped. Unit, GPU,
and regression required aggregates and Codecov passed. Two ONNX example
leaves failed because the runner could not load a cuDNN sublibrary;
their dependent example aggregate consequently failed.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
### Additional Information
- TODO: Deliver authoritative `native`/FP32/FP16/BF16 ONNX export across
all quantized formats in follow-up pull requests.
> 🤖 _Generated by Codex (AI agent)._
---------
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Codex <codex@openai.com>
|
||
|
|
d19925e446 |
simple refactor(export): split TensorRT-LLM-only code into modelopt/torch/export/trtllm (#2365)
### What does this PR do?
Type of change: refactor.
**The TensorRT-LLM checkpoint export format is deprecated.** Per
`docs/source/deployment/1_tensorrt_llm.rst`: *"The
`export_tensorrt_llm_checkpoint` API will be deprecated in future
releases. Users are encouraged to transition to the unified HF export
API, which provides enhanced functionality and flexibility for exporting
models to multiple inference frameworks including TensorRT-LLM, vLLM,
and SGLang."*
That deprecated code was not sitting off to one side — it was
**interleaved with the export path we actually want to grow.**
`modelopt/torch/export` mixed the deprecated TensorRT-LLM checkpoint
logic with the framework-agnostic HF/Megatron export code, in the same
modules:
- `layer_utils.py` was 1,986 lines, of which ~1,600 were TensorRT-LLM
`build_*_config` builders. The HF path imports this module for five
small predicates (`is_moe`, `is_quantlinear`, …) and dragged the whole
deprecated builder set in with them.
- `model_config.py` held the TensorRT-LLM `ModelConfig` dataclasses
*and* the `QUANTIZATION_*` / `KV_CACHE_*` constants that every backend
needs, so all of HF export imported the deprecated checkpoint schema to
get a format name string.
- `quant_utils.py` carried two helpers whose only caller is the
deprecated `postprocess.py`.
**This PR isolates the deprecated format so it stops polluting the
HuggingFace export path.** Everything reachable only from
`export_tensorrt_llm_checkpoint` now lives under
`modelopt/torch/export/trtllm/`, and the dependency is **one-way**:
`trtllm/` reaches into the parent through `quant_format`, `quant_utils`
and `layer_utils`, and **no implementation module in the parent imports
`trtllm/`.** The single exception is the deprecation re-export in
`modelopt/torch/export/__init__.py` described below, which is scheduled
for deletion in 0.49.0.
That one-way edge is the property worth protecting in review. It means
the deprecated format can be evolved, frozen, or eventually removed
without touching HF export, and HF export can no longer accidentally
grow a dependency on it.
### Deprecation handling
The format has carried a deprecation notice in the deployment docs since
`bc546943b4` (2025-10-08, first shipped in 0.39.0) — about 11 months.
But the deprecation policy in `README.md` also specifies *how* a
deprecation is communicated: a changelog entry, a source statement of
timing, and a runtime warning on use. **None of those existed**; only
one docs page ever said anything. So 0.48.0 is the first release that
gives users a signal they can act on, and this PR treats it as the
*start* of the migration period rather than the end:
- Both entry points now emit a `DeprecationWarning` naming 0.48.0 and
the 0.49.0 removal.
- `export_tensorrt_llm_checkpoint` and
`torch_to_tensorrt_llm_checkpoint` **remain importable from
`modelopt.torch.export`** for this release only, so existing callers
keep working *and* actually receive the warning. Removing the path in
the same release that first warns would mean callers hit `ImportError`
and never see it.
- The 0.49.0 removal date is stated in all four channels the policy
names: the runtime warning, the source (`.. deprecated:: 0.48.0` plus a
comment), the changelog, and the deployment doc.
The deeper module paths (`modelopt.torch.export.model_config_export`,
`modelopt.torch.export.model_config`) are **not** forwarded. Neither
appeared in a docs example, and `model_config.py` never declared
`__all__`, so by the `__all__` convention in `CONTRIBUTING.md` they were
never part of the public surface.
Eight modules had no non-TRT-LLM importer and moved whole:
`model_config_export`, `model_config_utils`, `postprocess`,
`distribute`, `tensorrt_llm_utils`, `tensorrt_llm_type`,
`hf_config_map`, `mcore_config_map`.
Three were genuinely mixed and were split by call-graph analysis rather
than by file:
| module | stayed shared (HF path) | moved to `trtllm/` (deprecated) |
|---|---|---|
| `model_config.py` | `QUANTIZATION_*`, `KV_CACHE_*`,
`FUSION_FREE_FORMATS` → new leaf module `quant_format.py` | the
`ModelConfig` dataclasses + `LINEAR_*`/`LAYERNORM_*` checkpoint-layout
constants |
| `layer_utils.py` | 9 module-shape predicates and MoE quantizer helpers
(`is_moe`, `is_quantlinear`, `get_experts_list`,
`sync_moe_gate_up_amax`, …) | the 39 `build_*_config` builders and
enc/dec helpers |
| `quant_utils.py` | everything else | `get_scaling_factor_from_weight`,
`resmooth_and_get_scale` (only caller is `trtllm/postprocess.py`) |
`adjust_attn_amax_values` was deliberately left in the shared
`quant_utils.py`: it has no production caller at all (only a test), so
"used only by TRT-LLM export" is not demonstrable for it.
Nothing was added or removed. `export_tensorrt_llm_checkpoint` behaves
exactly as before, just from a new import path and with a warning
attached.
### Usage
```python
# Deprecated TensorRT-LLM checkpoint export — new home, and warns on call
from modelopt.torch.export.trtllm import (
export_tensorrt_llm_checkpoint,
torch_to_tensorrt_llm_checkpoint,
)
from modelopt.torch.export.trtllm.model_config import ModelConfig
# The pre-0.48 path still works for one release, and warns — removed in 0.49.0
from modelopt.torch.export import export_tensorrt_llm_checkpoint
# Shared format constants — new home, still re-exported from the top level
from modelopt.torch.export.quant_format import QUANTIZATION_NVFP4, KV_CACHE_FP8
from modelopt.torch.export import QUANTIZATION_NVFP4 # still works
# The recommended path — unchanged
from modelopt.torch.export import export_hf_checkpoint, get_model_type
```
### Testing
- `pre-commit` on all changed files: passes (ruff, ruff-format,
**mypy**, bandit, markdownlint). mypy caught one implicit re-export of
`is_layernorm`, now imported from the shared module directly.
- `tests/unit/torch/export`: **189 passed**. With the new `trtllm/` test
dir: **193 passed**.
- Full `tests/unit/torch`: **2367 passed, 0 export failures**. The 45
failures are pre-existing environment issues — a deepspeed circular
import and a read-only HF cache — confirmed by reading their error text,
not assumed.
- `pytest tests/gpu/torch/export --collect-only`: 172 items, no
collection error.
- In-repo consumers updated and re-verified by an AST scan that imports
every `modelopt.torch.export*` module referenced anywhere in the tree
and checks each imported name still resolves: `hf_ptq.py`,
`export_trtllm_ckpt.py`, `deepseek_v3/ptq.py`, the AutoQuantize
notebook, `hf_ptq/README.md`, 2 docs pages, 4 tests.
- **Deprecation contract is covered by committed tests** (3 new, in the
`trtllm/` test dir): the pre-0.48 top-level import still resolves to the
same objects, `torch_to_tensorrt_llm_checkpoint` warns *at call time*
rather than on first `next()` (it returns a generator, so a naive
`warnings.warn` in the body would fire late or never), and one
`export_tensorrt_llm_checkpoint` call emits exactly one warning rather
than two. The first of these makes closing the migration window early a
test failure rather than a silent regression. `pyproject.toml` sets no
`filterwarnings = error`, so no suite fails on the new warning.
- **After merging `main`** (4 commits, incl. a 180-line rewrite of
`unified_export_megatron.py` that touches a file this PR also edits):
merged with no conflicts, then re-verified rather than trusted — import
scan clean across 24 export modules, `ruff check` clean repo-wide, 193
export tests passing, GPU collection still clean.
**Not run: the GPU suites** (`tests/gpu/torch/export`,
`tests/gpu_trtllm`) — no GPU in my environment.
`tests/gpu/torch/export/test_export.py` had its imports retargeted, so
it is the one most worth a GPU run before merge.
### Reviewer note: the deprecated path has no test coverage
Worth knowing before reviewing. **No test in the repo — including
`tests/examples/` — calls `export_tensorrt_llm_checkpoint`,
`torch_to_tensorrt_llm_checkpoint`, any `build_*_config`,
`convert_to_tensorrt_llm_config`, or `postprocess_model_config`.** So
~4,600 moved lines have no direct tests, and this refactor is validated
by import-graph reasoning, lint and mypy rather than by tests exercising
the moved code.
Given the format is deprecated and scheduled for removal in 0.49.0, **no
new coverage is planned for the conversion path itself** — writing fresh
tests for an API being removed next release isn't a good use of effort.
The gap is documented so reviewers can weigh the risk, not as a TODO.
(The deprecation *mechanism* is tested; see Testing.)
One caveat on how the gap was established: a runtime check showing all
12 `trtllm` modules in `sys.modules` after the export suites is *not*
evidence of coverage — importing any submodule runs
`trtllm/__init__.py`, which star-imports `model_config_export` and pulls
in the rest. Real line coverage could not be measured (`coverage`'s
tracer is incompatible with this venv's torch build: `ValueError: module
functions cannot set METH_CLASS or METH_STATIC`, on both the C tracer
and `sysmon`). The claim rests on a call-site audit generated from the
actual public symbols of those modules.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ for the public API —
`export_tensorrt_llm_checkpoint` and `torch_to_tensorrt_llm_checkpoint`
remain importable from `modelopt.torch.export` through the 0.49.0
migration period, now with a `DeprecationWarning`. The undocumented
submodule paths `modelopt.torch.export.model_config_export` and
`.model_config` did move; see **Usage**.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
code or dependencies; existing code relocated.
- Did you write any new necessary tests?: ✅ — 3 tests covering the
deprecation contract (old import path, call-time warning, exactly-one
warning). One existing test also moved to mirror the source split. No
new coverage for the deprecated conversion path itself; see the note
above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 **Deprecations**, covering both the runtime warning and
the new import location.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
Git detected the moves, so the diff stays reviewable: 8 files show as
pure renames (100%), the three split files as rename/copy at 94–99%
similarity, and only `layer_utils.py` as a 79% rewrite — expected, since
it shed 1,616 lines to `trtllm/`.
`examples/hf_ptq/hf_ptq.py` and
`examples/llm_sparsity/weight_sparsity/export_trtllm_ckpt.py` still call
the deprecated API, so those examples now print the warning. That is the
intended nudge, but happy to silence or migrate them if preferred. They
import from the new `.trtllm` path already, so they need no change at
0.49.0.
Two incidental changes, easy to revert if unwanted:
- `modelopt/torch/export/layer_utils.py` mode `100755 → 100644` (it was
needlessly executable).
- The new test is named `test_trtllm_quant_utils.py`, not
`test_quant_utils.py`: these directories have no `__init__.py`, so
pytest derives the module name from the bare filename and the shorter
name fails collection with `import file mismatch` against the existing
`test_quant_utils.py` one level up.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added shared quantization and KV-cache format definitions for export
workflows.
- Added expanded TensorRT-LLM export support, including broader model
architecture and quantization handling.
- Added distributed export utilities for coordinating checkpoint data
across processes.
- **Deprecation**
- TensorRT-LLM checkpoint export now emits a warning and is scheduled
for removal in version 0.49.0.
- Use the documented export module and save optimized model state
explicitly when needed.
- **Documentation**
- Updated guides and examples with new import paths and deprecation
guidance.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
28dc117594 |
Add Day 0 Sub-agent Roles (#2006)
### What does this PR do? Type of change: new feature Adding sub-agent definition for Day 0 workflows. I'm adding subagents as an alternative, while we evaluate which is better. They are duplicated because Claude & Codex have different formats & different locations. Deriving from a shared vendor-neutral format at build-time is too complicated. ### Usage Tell your agent "Use the modelopt_model_quantizer agent to quantize the model" ### Testing Ran a trial using Qwen-2.5, Muse Glimmer, Qwen-2.8, and GLM-5.3-Flash. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information See design doc. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## New Features - Added specialized Model Optimizer agents for downloading, quantization, recipe search, evaluation, deployment, and performance benchmarking. - Added coordinated Codex and Claude Code access with validation requirements, artifact preservation, and concise handoffs. - Improved skill discovery across plugin and repository installations. ## Documentation - Clarified agent discovery, layout, and canonical editing locations. - Updated configuration guidance for supported agent definitions. ## Tests - Added synchronization checks for agent definitions and links. - Improved test compatibility for Python versions below 3.11. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chad Voegele <cvoegele@nvidia.com> |
||
|
|
9d0df45849 |
specdec: config_overrides for nested text_config checkpoints + load VLM-capable bases in merge_lora (#2289)
### What does this PR do?
Type of change: New feature + bug fix
Two related gaps, both hit while enabling EAGLE3 on a checkpoint whose
config nests its text dims.
**1. `config_overrides` for checkpoints whose `text_config` dims don't
propagate.**
Some multimodal checkpoints carry the real text-tower dims only under
`config.text_config`, leaving the parent fields `None`.
`from_pretrained` then builds a text tower with the wrong shape.
`load_vlm_or_llm` gains an optional `config_overrides` dict applied to
*both* the parent config and its `text_config` before instantiation, and
the three entrypoints that load checkpoints — `ar_validate.py`,
`export_hf_checkpoint.py`, `merge_lora.py` — get a `--config_overrides`
passthrough. `main.py` threads it from `ModelArguments`.
**2. `merge_lora.py` could not merge into any VLM base.**
It loaded via `AutoModelForCausalLM`, which cannot load architectures
absent from the CausalLM Auto map — every VLM base failed. It now goes
through `load_vlm_or_llm`, which routes VLMs to
`AutoModelForVision2Seq`/`AutoModelForImageTextToText` and plain LLMs to
`AutoModelForCausalLM` with the same `dtype`/`device_map`, so LLM
behavior is byte-for-byte unchanged.
Also adds an optional `transformers_cosmos3` import so `cosmos3_omni` is
registered with `AutoConfig` before use, and dispatches that
`model_type` to its model class directly — that plugin registers only a
*config*, never a model under `Auto*`, so `AutoModelForCausalLM` raised
`KeyError('cosmos3_omni')` regardless of imports. The import is wrapped
in `contextlib.suppress(ImportError)`, so it is a no-op when the plugin
isn't installed.
### Usage
```bash
# Checkpoint whose real dims live under config.text_config
python examples/speculative_decoding/scripts/ar_validate.py \
--model_path <ckpt> --trust_remote_code \
--config_overrides '{"num_hidden_layers": 36, "intermediate_size": 12288, "num_key_value_heads": 8}'
# Same flag on export and merge
python examples/speculative_decoding/scripts/export_hf_checkpoint.py \
--model_path <ckpt> --export_path <out> --config_overrides '{"num_hidden_layers": 36}'
python examples/speculative_decoding/scripts/merge_lora.py \
--base_model_path <base> --exported_lora_dir <out> --output_path <merged> \
--config_overrides '{"num_hidden_layers": 36}'
```
```python
model = load_vlm_or_llm(path, config_overrides={"num_hidden_layers": 36}) # default None
```
### Testing
Exercised end-to-end on a Cosmos3-Nano (16B, 36-layer text tower) EAGLE3
LoRA run:
- **Training** — the base loads with all 36 text layers and correct
dims; two 4-epoch co-training runs completed (46,816 steps each).
- **Export + merge** — produced `adapter_model.safetensors` and a merged
base. Verified correct by per-layer weight diff: a `start_layer=18` run
changed **exactly** layers 18-35, with layers 0-17 bit-identical to the
base.
- **AR validation** — `--config_overrides` loads the trained checkpoint;
80/80 MT-Bench samples, AR 3.42.
- **Regression check** — `merge_lora` via `load_vlm_or_llm` produces a
base loadable by `lm_eval`; ifeval/arc_challenge/winogrande all ran to
completion.
No local unit-test run: `nvidia-modelopt` isn't installed in my
checkout, so `tests/conftest.py` fails to import. Relying on CI.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `config_overrides` defaults
to `None`; the `merge_lora` loader swap keeps the same class, dtype and
device_map for plain LLMs.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependency; `transformers_cosmos3` is an optional import guarded by
`contextlib.suppress`.
- Did you write any new necessary tests?: ❌ — exercising these paths
needs a checkpoint with a nested `text_config`, which the unit suite has
no fixture for. Happy to add one if a reviewer can point me at a small
suitable model.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — can add a *Speculative Decoding* entry for the `merge_lora` VLM fix
if you consider it changelog-worthy.
- Did you get Claude approval on this PR?: ❌ — not yet run.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added JSON-based model configuration overrides across speculative
decoding, training, validation, export, and LoRA workflows.
- Overrides can update primary model and text configuration settings.
- Expanded support for vision-language models and Cosmos3 Omni
checkpoints.
- **Bug Fixes**
- Improved configuration handling for offline loading and
checkpoint-based initialization.
- Restored draft-model precision during checkpoint loading and model
conversion.
- Added validation for malformed, unsupported, and non-finite override
values.
- Standardized configuration override guidance across command-line
workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
635688d26e |
[6410139] Fix ONNX AutoCast for large external initializers (#2317)
### What does this PR do?
Type of change: Bug fix
Fix ONNX AutoCast for models whose external initializers exceed the
in-memory protobuf limit.
- Keep external tensor payloads reference-only during graph sanitization
and type inference, then materialize them once before value-dependent
classification and conversion.
- Duplicate shared initializers directly in `GraphProto`, preserving
external-data metadata without reading tensor bytes.
- Use file-backed ONNX paths for validation, shape inference, reference
execution, and custom-operator inspection when required.
- Avoid a redundant sanitizer pass in the fully sanitized AutoCast path
while preserving existing behavior for direct `PrecisionConverter` and
`convert_to_f16()` callers.
No CLI flags, dependencies, or public return types change.
### Usage
```bash
python -m modelopt.onnx.autocast \
--onnx_path model.onnx \
--output_path model_bf16.onnx \
--low_precision_type bf16
```
### Testing
- Ran the complete CPU-only AutoCast unit suite with no GPU visible: 248
passed.
- Ran focused ONNX utility regressions covering shared initializer
duplication and file-backed protobuf routing: 8 passed.
- Ran pre-commit on all changed files.
- Ran CPU-only integration coverage with an exact 2,147,485,696-byte
external initializer using protobuf 7.35.1 and 6.33.6.
- Verified BF16, FP16, shared-initializer, and aggregate-external-data
runtime-cast cases.
- Verified every integration output with `onnx.checker.check_model(...,
full_check=True)`; outputs expected to remain external-data-backed did
so.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌
> 🤖 _Generated by Codex (AI agent)._
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Bug Fixes**
- Fixed ONNX AutoCast failures for models with external initializers
larger than 2 GiB.
- Improved handling of large or external-data models during shape
inference, conversion, and runtime validation.
- Preserved external initializer metadata while avoiding unnecessary
data materialization.
- Improved temporary-file cleanup when model loading or inference fails.
- Improved processing of shared initializers, nested graphs, and custom
nodes.
- **Tests**
- Added coverage for large models, external initializers, nested graphs,
custom nodes, and runtime cleanup.
- **Documentation**
- Updated the changelog with recent fixes and benchmarking information.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
|
||
|
|
079078de9d |
FSDP2 export optimizations (#2207)
### What does this PR do? **Type of change:** Performance enhancement FSDP2 checkpoint export gathered the whole model onto rank 0 and wrote it from a single process. On large models that made export the slowest phase of a PTQ run and could exhaust host memory. The model is now split into export units — one per decoder layer, plus one for everything else that owns state — and the units are dealt round-robin across ranks. Every rank enters the unshard window for every unit, because the gather is a collective; only the owner keeps the result, packs it, and streams it to its own shard files. Rank 0 names the shards and writes the index once every rank has finished. Two properties follow. Peak host memory per rank is ~model/world instead of the whole checkpoint on rank 0. And packing runs on a gathered, full weight, so it is byte-for-byte the single-process code path — no amax narrowing, no block-quantized fallback, no scale-buffer heuristics. Measured on B300 (bia), NVFP4, `hf_ptq.py --use_fsdp2`: | model | world | export | artifact | |---|---|---|---| | Qwen3-235B-A22B | 8 (1 node) | **168.6s** | 146,361 keys / 16 shards / 0.122 TiB | | oakhaven-max 4.5 TB | 64 (8 nodes) | **1359.7s** | 568,398 keys / 158 shards / 1.292 TiB | Checkpoints were validated, not just timed: every index entry present, no leftover `__shard_part*`, all ranks exit 0. A repeat of the oakhaven configuration landed at 1411.7s, so treat ~1.36–1.41 ks as its range. No speedup ratio is quoted: the old rank-0 path was last measured on different hardware on a different day, and this cluster drifts up to 14% run-to-run on an identical tree, so a cross-day ratio would not reproduce. A same-day baseline of `main` is the outstanding item. Two alternatives were considered and rejected. `torch.distributed.checkpoint`'s `get_model_state_dict` full-gathers to rank 0 only with no per-module owner knob, so it reproduces the exact host-memory bound this removes, and per layer it was not faster than `full_tensor()`. Packing each rank's shard in place was ~9% faster at Qwen3-235B w8 on the trees measured at the time, but rebuilt `FSDPParam` objects from torch's private FSDP internals and carried ~400 lines of format-specific shard handling; gathering to the owner deletes all of it. Configurations that would silently corrupt a checkpoint now raise instead: FSDP2 composed with another DTensor parallelism (FSDP2 + TP on a 2-D mesh; HSDP is fine), models with no discoverable decoder layers, a decoder layer object reused across layers, and a container that holds the decoder layers while owning direct state of its own. ### Usage No API change — `export_hf_checkpoint` selects the path automatically. ```python from modelopt.torch.export import export_hf_checkpoint # model is FSDP2-wrapped; every rank must call this (it unshards collectively) export_hf_checkpoint(model, export_dir="./exported") ``` Testing - `test_fsdp2_streaming_export_matches_reference` — world = GPU count under real FSDP2; merged checkpoint compared tensor-by-tensor against a single-process export of the same seeded model, index asserted to span more than one shard file. Covers the ownership split and the cross-rank index merge. - `test_fsdp2_gathered_pack_matches_reference` — gathered packing vs whole-weight packing over nvfp4 / fp8 / int8 / fp8-per-channel-per-token / int4-blockwise - World=1 CPU tests for unit enumeration, shard sub-splitting by `max_shard_size`, tied-alias dropping, and `extra_state_dict` - `tests/unit/torch/utils/test_perf.py` — sampling and threshold logic of `maybe_clear_cuda_cache`, which replaces an unconditional `torch.cuda.empty_cache()` after every packed weight Before your PR is "*Ready for review*" Make sure you read and follow Contributor guidelines and your commits are signed (git commit -s -S). Make sure you read and follow the Security Best Practices (e.g. avoiding hardcoded trust_remote_code=True, torch.load(..., weights_only=False), pickle, etc.). - Is this change backward compatible?: ✅ <!-- No API change; non-FSDP2 and offloaded export paths are unchanged. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ✅ ### Additional Information Pre-export time (load, quantizer insertion, calibration) is now the dominant phase of an FSDP2 PTQ run. It is unaffected by this PR and addressed separately in #2359. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Improved FSDP2 Hugging Face checkpoint export by streaming rank-owned layers to shard files, reducing export time and host-memory usage for large models. - Improved export handling for multimodal models with missing architecture metadata. - Standardized weight packing across supported parallelism configurations. - **Performance** - Reduced repeated module lookups and unnecessary CUDA cache cleanup. - Improved shard generation, merging, and index completeness. - **Compatibility** - Expanded support for FSDP2 quantized export workflows and shard-local packing. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com> Signed-off-by: sugunav14 <178320438+sugunav14@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
613e5e8b2e |
feat(quantization): PTQ support for Step-3.7 MoE checkpoints (#2202)
### What does this PR do? Type of change: New feature Adds PTQ support for **Step-3.7** (`stepfun-ai/Step-3.7-Flash`). Follow-up to [NVBug 6518665](https://nvbugspro.nvidia.com/bug/6518665) / [OMNIML-5583](https://jirasw.nvidia.com/browse/OMNIML-5583): with the export crash fixed in #2071 the run completes, but the checkpoint it writes is silently unquantized — ```json {"quantization": {"quant_algo": null, "kv_cache_quant_algo": "FP8", "quantized_layers": {}}} ``` Two independent causes, both from Step's `trust_remote_code` modeling code. **1. The expert weights were invisible to quantization.** Step-3.5 and Step-3.7 ship the same custom `MoELinear`: a plain `nn.Module` holding one 3-D `weight` of `[num_experts, out_features, in_features]`, whose `forward(x, expert_id)` runs `F.linear` against the selected slice. It is not an `nn.Linear`, and the weights sit on the projection submodule rather than on the expert container, so neither the plain-linear path nor `_fused_experts_wrapper_class` (which wants a 3-D `down_proj` *Parameter*) claims it. The `_QuantMoELinear` wrapper that handles exactly this layout has existed since #1063, but its registration was gated on the Step-3.5 class names: ```python if type(model).__name__ not in ("Step3p5ForCausalLM", "Step3p5Model"): return for module in model.modules(): if type(module).__name__ == "Step3p5MoEMLP": ``` Step-3.7's root is `Step3p7ForConditionalGeneration` and its container is `Step3p7MoEMLP`, so it returned immediately and no expert ever got a quantizer. Detection is now **structural** — a 3-D `weight` plus `num_experts` / `in_features` / `out_features` and a two-positional-argument forward — so any Step revision (or another model shipping this layout) is picked up without a third hardcoded name. `_reconstruct_fused_moe_linear` likewise matches the wrapper type instead of the generated `QuantMoELinear` class name; a model whose class is spelled differently would otherwise quantize fine but export unusable per-expert keys. **2. Step's module names don't match the general recipes.** The MoE block is `moe` and the dense sibling is `share_expert`, so `*.experts.*`, `*block_sparse_moe*` and `*mlp*` reach none of the routed experts. This PR ships `huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast` and `huggingface/step3p7/ptq/nvfp4_mlp_only-kv_fp8`, which select `*moe*` and disable the router (`moe.gate`) and the shared expert — mirroring the existing Step-3.5 recipe — and documents the naming trap in `modelopt_recipes/ptq.md`. ### Usage ```bash python examples/hf_ptq/hf_ptq.py --model /local/Step-3.7-Flash --trust_remote_code \ --recipe huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast \ --dataset /local/cnn_dailymail --calib_size 32 --export_path /local/Step-3.7-Flash-nvfp4 ``` ### Testing - `tests/unit/torch/quantization/plugins/test_moe_linear.py` — structural detection (positive plus 2-D-weight / wrong-forward negatives), registration on a Step-3.7-shaped model, per-expert quantizers with calibrated amax, and reconstruction back to the 3-D parameter. Plus two export-dispatch tests added from review: the registry resolves a `Quant_SyntheticMoELinear` to `_export_moe_linear` (this one fails against the old name-keyed registration), and the handler fills an unrouted expert's input amax. - `tests/unit/recipe/test_step3p7_recipes.py` — drives both shipped recipes over a model mirroring Step's real paths (`model.language_model.layers[i].{moe,share_expert,mlp}`): routed experts NVFP4-quantized per expert, router / shared expert / dense MLP / `lm_head` per recipe scope. Ran locally (torch 2.11, transformers 5.5.4 — the version in the bug report): the two new files (12 tests) plus `tests/unit/recipe`, `tests/unit/torch/quantization/plugins/`, `tests/unit/torch/export/test_export_weight.py` and `test_export_registry.py` — 324 passed. Full `tests/unit` (minus onnx/puzzletron): 2490 passed, with 4 pre-existing `test_quant_aware_conversion.py` failures that reproduce unchanged on clean `main`. Not run end-to-end on the real Step-3.7-Flash checkpoint (1.4 TB / 8×B200) — QA can re-run against this branch with the recipe above. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — Step-3.5 keeps working; the name gate is replaced by a superset. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ ### Review updates - **Export dispatch keyed on the class name** (caught in review): registration was made class-name-independent, but `_export_moe_linear` in `hf_export_handlers.py` was still registered for the literal string `"QuantMoELinear"`, so a compatible class under another name bypassed the input-amax fallback for unrouted experts. The predicate now matches `_QuantMoELinear` through the MRO (lazy import, since the wrapper lives in the optional transformers plugin), keeping the name check as a fallback so the synthetic stand-ins in `test_export_registry.py` / `test_export_weight.py` still match. This was the same class-name coupling already fixed in `_reconstruct_fused_moe_linear`, in the file I hadn't looked at. - Canonical 2026 license header on the new test file. - Recipe comments corrected: `*moe*` matches the router, but **not** `share_expert` (`layers.N.share_expert.*` contains no `moe` segment) — that entry is an explicit guard, not an override. - `ptq.md` recommendation scoped to Step-3.7, since Step-3.5 has its own recipe. ### Additional Information Pairs with #2203 (fail fast when a quant config matches no weight quantizer), which turns this class of silent no-op into an error for any model. Independent branches; either can merge first. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added post-training quantization (PTQ) support for Step-3.7 Flash models, including per-expert quantization for routed MoE layers. - Added NVFP4 recipes for expert-only and MLP-only quantization, with optional FP8 KV-cache support. - **Bug Fixes** - Improved detection, quantization, reconstruction, and export of expert-indexed MoE layers across Step model revisions. - Added safeguards to prevent quantization of unsupported offloaded expert weights. - **Documentation** - Clarified Step-3.7 recipe selection, calibration guidance, and quantization exclusions. - **Tests** - Added coverage for expert, MLP, router, attention, shared-expert, export, and calibration behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
74a1dce390 |
Speed up FSDP2 MoE calibration by dropping redundant expert-weight gathers (#2359)
### What does this PR do? Type of change: Perf enhancement `promote_static_block_weight_quantizers` runs at the end of `max_calibrate`. All it does is read quantizer state -- it takes the weight the iterator hands it and throws it away. For this it goes through `iter_weights_for_calibration`, and on a fused-MoE module that iterator yields `weight[idx]`, once per expert. Under FSDP2 the fused weight is a DTensor spread across ranks, so `weight[idx]` isn't a cheap view. Slicing it makes PyTorch pull the whole fused expert weight back from every rank, just to hand over one slice that the loop then drops. That happens once per expert, per projection, per layer -- tens of thousands of round-trips on a large MoE. The fix adds `iter_weight_quantizers_for_calibration`: - The base `QuantModule` implementation delegates to `iter_weights_for_calibration` and drops the weight, so every subclass gets a correct implementation for free and the two iterators cannot drift apart. - Only the fused-MoE class — the one where materializing the weight view is itself expensive — overrides it, walking the per-expert quantizer `ModuleList` directly. It keeps the same skip condition as the weight iterator, since fetching an attribute is free and only indexing collectives. `promote_static_block_weight_quantizers` then iterates quantizers instead of `(weight, quantizer)` pairs. No other caller changes: the other four call sites genuinely use the weight. Why not just wrap the promote loop in `enable_weight_access_and_writeback`, the way `weight_only_quantize` does? It would still gather per module for weights the loop never reads. `weight_only_quantize` needs the window because it actually computes amax from the weight; promote only needs the quantizer. Also included: a warn-once check if `_amax` is ever a `DTensor`. It should not be — `_amax` is a registered buffer, and `fully_shard` shards parameters, not buffers — which is why the `reduce_amax` in this loop stays local. If that assumption ever breaks, the reduction becomes a per-quantizer collective too, and the warning says so rather than letting it degrade silently. ### Results Measured on Qwen-3.8 2.4T at world 64 (8 nodes, B300), NVFP4: | | pre-export | |---|---| | before | 79.3 min | | after | **19.0 min** | 60.3 min saved, 4.2x. The block is gone rather than shortened -- the run logs 13 silent minutes in total, all of it model load. Calibration itself is unchanged at 2:23, so the saving comes from the promote loop rather than from work moving elsewhere. End-to-end projects to 41.7 min. ### Usage No API change. `mtq.quantize` picks this up automatically. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ Additive; `iter_weights_for_calibration` and all its callers that use the weight are untouched. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ Under 0.47 Bug Fixes — the call site dates to 0.46. - Did you get Claude approval on this PR?: ❌ Not yet. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Performance** - Reduced calibration overhead for FSDP2-sharded fused-MoE quantization by avoiding unnecessary expert-weight gathering when only quantizer state is needed. - Preserved quantizer ordering and projection filtering during calibration. - **Bug Fixes** - Added a warning when global amax reduction requires collective processing for individual quantizers. - **Tests** - Added coverage for gated and non-gated fused-expert calibration, including verification that fused weights are not unnecessarily indexed. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
279d510616 |
fix(specdec): correct resume and bound staging in the vLLM hidden-state dump (#2080)
### What does this PR do?
**Type of change:** Bug fix
Fixes two issues in the vLLM offline hidden-state dump
(`examples/speculative_decoding/collect_hidden_states/compute_hidden_states_vllm.py`).
Both are invisible on small dumps and only bite at scale, which is why
they survived until
now — they were found while dumping ~194k conversations for a MiniMax-M3
draft.
**1. Resume silently re-processed already-finished work.**
`keep_conversation` skips conversations whose `.pt` already exists, but
that predicate reads
**on-disk state**, which is not part of the fingerprint `datasets`
computes for `filter()`
(it hashes the function and the dataset). With a persistent HF cache
reused across a resumed
or requeued run, the cached *"keep everything"* result from an earlier
run — computed when
few or no `.pt` files existed — is replayed. The run then re-generates
and **overwrites**
conversations it had already completed, and reports `Removed 0
conversations due to existing
output files` while doing so.
Observed on a 194k-conversation dump: ~62k `.pt` rewritten over a
two-hour window with the
total output count completely flat.
Fix: pass `load_from_cache_file=False` so the filter re-checks the disk
on every run.
**2. Staging exhausted `/dev/shm` partway through large dumps.**
The script generated the **entire** dataset before saving anything. The
KV connector stages
each conversation's hidden states under its `shared_storage_path`
(`/dev/shm`, i.e. RAM, by
default) and they are only freed by `cleanup_hidden_states()` in the
save loop — so every
conversation stayed staged simultaneously. On a large dump this exhausts
the space and the
connector starts failing writes:
```
Hidden-states write failed for req_id=...:
SafetensorError('Error while serializing: I/O error: No space left on device (os error 28)')
```
Fix: generate and save in chunks of `--save-chunk-size` (default 256),
so at most one chunk
is staged at a time. As a side benefit the dump becomes **incrementally
durable** — an
interrupted run (walltime limit, node failure) keeps its finished
conversations and the
resume path above continues from them, instead of losing the whole run's
work.
### Testing
- Reproduced both failures on a 194k-conversation MiniMax-M3 dump (8-way
DP, TP8), and
confirmed both fixes on the same workload: after the change the output
count advanced
monotonically across requeues (123k → 194k) with no rewrites, and
`/dev/shm` stayed bounded
through completion.
- `pre-commit run --files ...` passes (ruff check/format, mypy, bandit,
license, rst checks).
- Behavior is unchanged for a fresh single-shot dump other than the
chunked generate calls;
the default `--save-chunk-size 256` is the only new knob.
### Additional Information
Extracted from #1749, which is otherwise superseded by the streaming
DFlash/DSpark path — these
two fixes are model-agnostic and apply to any offline dump, so they are
worth landing on their
own.
### Before your PR is "*Ready for review*"
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No — the failure modes are
multi-process/at-scale (datasets cache reuse across runs, connector RAM
staging) and are not reproducible in the unit-test harness.
- **Did you add or update any necessary documentation?**: Yes —
CHANGELOG entry.
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added chunked hidden-state generation for large vLLM offline runs.
* Added a configurable save-chunk size, defaulting to 256 conversations.
* Enabled incremental saving and resumption of hidden-state outputs.
* **Bug Fixes**
* Improved resume filtering to accurately detect existing output files.
* Reduced memory usage by saving and releasing each generated chunk.
* Ensured temporary files are cleaned up after interrupted or skipped
saves.
* Added atomic output-file replacement to prevent incomplete results.
* Added validation to prevent invalid conversation IDs from creating
unsafe paths.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
56af187565 |
ar_validate: fail loudly when every sample fails (#2288)
### What does this PR do?
Type of change: Bug fix
`validate_ar()` catches per-sample exceptions, prints a `WARNING`, and
returns whatever succeeded. When *every* sample failed it returned an
empty list, and the reporting block was guarded by `if results and
accelerator.is_main_process:` — so the script printed no results and
exited **0**. A run where 100% of samples failed was indistinguishable
from a successful one.
This bit us on a real run: an EAGLE3 checkpoint loaded with
`device_map="auto"` was sharded across 8 GPUs, every one of the 80
samples died with `Expected all tensors to be on the same device`, and
the job still exited 0 with no AR number anywhere in the log — the
wrapper stamped it PASS.
Now it raises, so the caller sees a non-zero exit. Any previously
"passing" run that printed no AR number was never meaningful.
### Usage
No API change. Existing invocations are unaffected when at least one
sample succeeds:
```bash
python examples/speculative_decoding/scripts/ar_validate.py \
--model_path <ckpt> --steps 3 --osl 1024 --num_samples 80
```
### Testing
Reproduced the silent-pass on a Cosmos3-Nano EAGLE3 checkpoint (80/80
samples failing): before this change the job exited 0 and stamped PASS;
after it, the job exits non-zero with the sample failures visible.
Confirmed the normal path is unchanged by a subsequent run that
completed 80/80 and printed AR 3.42.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — only affects the
all-samples-failed case, which previously produced no output and a
misleading exit 0.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ — the failure path requires
a model that errors during AR validation; the existing suite has no
harness for that.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — behavior fix in an example script, not a released-feature change.
- Did you get Claude approval on this PR?: ❌ — not yet run.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved validation error handling when all samples fail.
* Validation now rejects non-positive sample counts before processing.
* Empty validation results are clearly distinguished from cases where
all samples fail.
* Error messages report the actual number of validation samples
attempted, capped at the available dataset size.
* Empty validation results are no longer reported as successful.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Ye Yu <yeyu@nvidia.com>
|
||
|
|
4956213d67 |
Unblock Qwen3.5/3.6 QAD: Megatron export, calibration, and distillation fixes (#2334)
### What does this PR do? Type of change: Bug fix Everything that blocked running QAD on a quantized Qwen3.5 / Qwen3.6 MoE checkpoint: two Megatron-Core → HuggingFace export bugs that make it unservable (§1–2), the dead code the first leaves behind (§3), a no-op flag (§4), a multi-GPU calibration deadlock (§5), and four distillation / data-prep bugs that stopped QAD itself from running (§6). #### 1. Routed experts were exported packed, and vLLM cannot load that ``` AttributeError: Layer language_model.model.layers.23.mlp.experts has no parameter 'w2_weight_weight_scale_2' for checkpoint weight ...experts.down_proj_weight_scale_2 ``` `mcore_qwen35vl.py` used `GroupedMLPPacking`, mirroring the **BF16 upstream** checkpoint, which really is packed. But that mapping is only used for **quantized** export, and vLLM's quantized MoE loader needs per-expert scales — both released NVFP4 checkpoints (`Qwen3.6-35B-A3B-NVFP4` via hf_ptq, `Nemotron-3.5-Lightning-30B-A3B-NVFP4` via Megatron-LM) are per-expert. `_grouped_mlp_slicing` gains `gate_proj_name` / `up_proj_name` to split each expert's fused gate+up and slice its per-block `weight_scale`; `GroupedGatedMLPSlicing` wires it up. The Megatron checkpoint layout is unchanged, so affected checkpoints need only a **re-export**. `_verify_exported_keys` is relaxed to match: exported modules now contribute their ancestor prefixes, so expanding one source module into many is not reported as ~82 dropped tensors. A module genuinely absent still has nothing beneath its prefix and is still caught. #### 2. A quantized `output_layer` (`lm_head`) could not be checkpointed `GPTModel.sharded_state_dict` drops `output_layer._extra_state` and asserts it is empty. ModelOpt keeps quantizer state there, so saving raised and — since that method also backs the load plan — loading silently restored the layer **unquantized**. `keep_gpt_output_layer_extra_state()` retains it, applied from `megatron_replace_quant_module_hook` so **every** Megatron model gets it (Megatron-LM and NeMo users included, neither of whom can import `mbridge`, which needs `megatron.bridge`). It matches the upstream body by AST before replacing it and self-disables otherwise. [NVIDIA/Megatron-LM#7086](https://github.com/NVIDIA/Megatron-LM/pull/7086) is **closed, not merged**: nemo:26.10 migrates `GPTModel` to `HybridModel`, whose `sharded_state_dict` has no pop-and-assert, so this side keeps the workaround. Not cosmetic: `lm_head` is 248320×2048 = 509M params, **34.6% of per-token weight traffic** on a model with ~2.9B active params. #### 3. Cleanup Nothing maps `GroupedMLPPacking` once qwen3_5 is switched over; it is removed with `_grouped_mlp_packing` and the `quantize=` / `record_quant_config=` parameters that existed only to serve it. Llama-4's `PackNameRemapping` is unaffected. Two smaller review-driven fixes: the gated-split shape checks raise `ValueError` rather than `assert` (stripped under `-O`), and per-expert quant metadata is recorded for `local_expert_indices` rather than every global id, fixing non-contiguous EP. #### 4. Remove the no-op `--moe_calib_experts_ratio` from the Megatron quantize example `examples/megatron_bridge/quantize.py` accepted the flag and threaded it into the `mtq` config, but `_moe_calib_experts_ratio` exists only in `plugins/huggingface.py` (9 refs) and never in `plugins/megatron.py` (0); `mode.py:247` only assigns it to modules already exposing the attribute. On a Megatron MoE model it was accepted and silently ignored — a trap, since on a 256-expert model it reads like a major quality lever. `hf_ptq.py` keeps it, where it works. #### 5. Fix multi-GPU image-text (VLM) calibration deadlocking VLM calibration hung for 30 minutes and died on a gloo timeout whenever `world_size > 1`, with no error until the timeout fired. `NemotronTarPlusJsonlIterable` split its budget with truncating division, so the stream supplied fewer samples than requested (1024 over 3 subsets → 341×3 = **1023**). `_ShardedIterable` gives rank *r* items *r, r+W, r+2W…*, so a stream that is not a multiple of `world_size` leaves the trailing rank one short — it exits the forward loop early and the others block on the next collective. The arithmetic predicts both observed hangs exactly: 1024 → stall at **255/256**, 512 (yielding 510) → **127/128**. Fixed both ends: subset budgets are distributed with `divmod` so they sum exactly, and `_ShardedIterable` truncates every rank to `floor(len / world)` — which also covers `num_samples` not being divisible by `world_size`, as the first fix alone does not. Verified on Qwen3.6-35B-A3B (EP=4, `nemotron_vlm_dataset_v2`, 1024 samples): the configuration that hung twice now completes 256/256 and exports. Unit tests cover both fixes and fail without them. #### 6. Fix the distillation path so QAD can actually run Four independent bugs, all hit while running QAD end to end on Qwen3.6-35B-A3B. Each blocks a different configuration, and together they made every sequence length OOM or abort. - **Context parallel aborts.** The DDP config derived `average_in_collective` from `--sft` alone, but context parallel also needs per-token loss reduction, so any `--cp_size > 1` run died on `Cannot average in collective when calculating per-token loss`. - **`TopKLogitsKLLoss` was not memory-efficient.** Despite documenting "without gathering full logits", it cast the *whole* vocabulary to FP32 before selecting the top-k, allocating two `[seq, vocab]` tensors — 30.3 GiB each at seq 32768 on this model's 248k vocab. Reducing before the cast is equivalent: widening is exact and temperature scaling is monotonic, so the selected entries and the loss are unchanged. - **MTP cross-entropy ran when it had nothing to recover.** `skip_lm_loss` exempts the MTP heads unconditionally, so their CE materialised another FP32 `[seq, vocab]` tensor even when the MTP head is excluded from quantization — as it is in every recipe here (775 of 906 `exclude_modules`, zero MTP `weight_scale` tensors exported). It is now skipped **only** when the model is quantized and MTP is left out of it; plain distillation such as pruning recovery still trains the MTP head. `test_mtp_excluded_from_quantization` pins all four cases. - **One bad record deadlocked data prep.** `megatron_preprocess_data` re-raised chat-template failures out of a pool worker, stalling the whole job until it timed out — three malformed records cost a multi-hour tokenization run. They are now skipped with a warning, matching the existing handling of malformed JSONL a few lines above. Also exposes `--logit_kl_topk`, which `DistillationConfig` has supported for a while but the example never passed through; `test_qad` now exercises that path. §4, §5 and §6 are independent of §1–3; happy to split them out if reviewers prefer. ### Usage No API change. Exported names now match the released checkpoints: ``` model.language_model.layers.0.mlp.experts.<E>.{gate,up,down}_proj.{weight,weight_scale,weight_scale_2} lm_head.{weight,weight_scale,weight_scale_2} ``` ### Testing - `test_mcore_export_mappings.py` — qwen3_5 mappings emit per-expert rules. Verified these fail without the fix (2 failed / 11 passed), with `Qwen3MoeForCausalLM` / `NemotronHForCausalLM` as controls. - `test_unified_export_megatron.py` — the gate/up split, per-block scale slicing, the 0-dim scalar fallback, and both directions of the `_verify_exported_keys` relaxation. - `test_megatron.py::TestKeepGptOutputLayerExtraState` — 15 cases: payload detection, no-op second call, warn-and-skip on an unrecognised `sharded_state_dict`, and `test_patches_stock_megatron_core` which installs a replica of the real pre-fix upstream body (verified against `be08ce5b1~1`) so the patched path is exercised whichever megatron-core is installed. - `test_qad.py` — CI caught that its reference comparison still assumed packed experts; fixed. **End to end on `Qwen/Qwen3.6-35B-A3B` (35B MoE, 256 experts), 4×GB200, nemo:26.08:** | | before | after | | --- | --- | --- | | export self-check | `Export dropped 82 tensor(s)` | passes | | expert tensors | `mlp.experts.gate_up_proj` (packed) | `mlp.experts.<E>.{gate,up,down}_proj` | | **vLLM v0.28.0 load** | **`AttributeError`, engine never starts** | **`Loading weights took 25.61 s`** | | **NEL eval (GPQA-D, MMMU-Pro)** | **FAILED** | **SUCCESS** | ### Results these fixes unblocked The export fix is what made a Megatron-produced NVFP4 MoE checkpoint servable at all, so it enabled a full PTQ study on Qwen3.6-35B-A3B. Accuracy deltas are against a BF16 baseline measured on the same harness, from **paired** per-question tests: | recipe | throughput vs BF16 | GPQA-D | SciCode ×8 | MMMU-Pro | IFBench | | --- | --- | --- | --- | --- | --- | | **W4A16** (weight-only) | **0.64–0.86×** — *slower* | −0.06 | −0.15 | +0.48 | −0.44 | | **W4A4** | 8/12 shapes faster | −0.60 | −0.70 | −1.48 (p=0.019) | −0.53 | | **W4A4 + 4-bit `lm_head`** | **9/12 shapes**, up to **1.30×** | **+0.03** (p=0.96) | −0.81 | **−1.16** (p=0.016) | −1.65 (ns) | Repeats: GPQA-D is `pass@1[avg-of-16]`; SciCode is 8 pooled runs per recipe; MMMU-Pro is 3 runs per side and IFBench 2–3 for BF16 and the last row, 1 elsewhere. AA-LCR (68.33 → 71.33, p=0.25, 3 runs per side) and τ²-Telecom (94.25 → 94.25, 3 runs per side) are on par; at 100 questions and 114 tasks they cannot resolve below ~5 pp and ~3 pp, so they carry no claim either way. #### QAD status (what §6 unblocked) With the §6 fixes in place, QAD runs end to end on this model: 32 nodes, `TP=1 PP=1 CP=1 EP=8`, seq 32768, gbs 512, ~38 s/iter, 124 GB/GPU peak. First accuracy read, MMMU-Pro at iteration 50 (0.84 B tokens), 3 runs per side, paired per-question: | | MMMU-Pro | vs BF16 | | --- | --- | --- | | BF16 | 74.55 | — | | W4A4 + 4-bit `lm_head` (PTQ) | 73.39 | **−1.16, p=0.016** | | + QAD, iteration 50 | 73.78 | −0.77, p=0.089 (ns) | The PTQ deficit that motivated this work is no longer statistically significant after 50 QAD iterations. The improvement itself (+0.39 vs PTQ) is **not** significant at p=0.41, and 50 iterations is 10% of the planned budget, so this is a direction rather than a result. A full six-benchmark sweep at iterations 50 and 300 is running; these numbers will be superseded. Two findings worth flagging beyond this PR: - **Weight-only NVFP4 is slower than BF16 on Blackwell.** W4A16 leaves activations in BF16, so vLLM cannot use the FP4 tensor cores and falls back to `MarlinNvFp4LinearKernel` / `'MARLIN'` MoE. W4A4 selects `FLASHINFER_TRTLLM` + `FlashInferCuteDslNvFp4LinearKernel` and beats W4A16 in **12/12** shapes. The Marlin line count tracks the recipe exactly (one W4A16 layer ⇒ one Marlin line ⇒ zero once `lm_head` is W4A4). - **The only accuracy cost is multimodal**: **−1.2 pp on MMMU-Pro** for the fastest recipe, confirmed over 3 runs per side (p=0.016). GPQA-D, SciCode, IFBench, AA-LCR and τ²-Telecom show no significant regression. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ <!-- Megatron checkpoints unaffected; re-export to gain the loadable layout. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!-- 0.47.0 → Bug Fixes; includes the removed flag, since passing it now errors instead of being ignored --> - Did you get Claude approval on this PR?: ✅ <!-- Reviewed; all findings addressed, threads resolved. --> ### Additional Information Both export bugs were found while reproducing `nvidia/Qwen3.6-35B-A3B-NVFP4` through `examples/megatron_bridge/`. Follow-up to #2332. Upstream counterpart [NVIDIA/Megatron-LM#7086](https://github.com/NVIDIA/Megatron-LM/pull/7086) is closed — see §2. Labeled `cherry-pick-0.47.0`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2d35643452 |
LiLiCorr training (#2342)
### What does this PR do? Type of change: new feature Adds **LiLiCorr**, a candidate-lattice reranker for DFlash drafts, as a new `projector_type` on the existing `dflash` mode — plus three DFlash-wide improvements that apply to every variant, and an optional composition with DFlash2's grouped convolutions. A DFlash drafter is trained on per-position marginals rather than on the joint block distribution, so its drafted tokens are individually plausible yet jointly incoherent. LiLiCorr keeps the top-`k` candidates the backbone already produces at each block position, scores transitions between adjacent candidates with a small two-layer transformer, and commits a path through the lattice greedily. Serving is unchanged in kind: verify still checks every drafted token against the target, so the emitted distribution is untouched and only acceptance length moves. - Paper: [LiLiCorr: Lightweight Likelihood Correlation of Parallel Drafts for Speculative Decoding](https://arxiv.org/abs/2608.20530) (arXiv:2608.20530) - Blog: https://research.nvidia.com/labs/nemotron/lilicorr/ - **Companion PR — serving support:** [sgl-project/sglang#37462](https://github.com/sgl-project/sglang/pull/37462) This PR is the **training** half. It trains the drafters and exports them; the companion PR above is what serves the resulting checkpoints, and is what the comparison table below was measured through. **What is in the commits** | | | | --- | --- | | LiLiCorr draft variant | `hf_lilicorr.py`, `modeling_lilicorr.py`, conversion routing, config fields, export | | Three DFlash-wide features | fp32 master weights for the draft, draft activation checkpointing, and a DDP hang fix — all default-off or behaviour-preserving, all applying to `dflash`, `domino`, `dspark` and `dflash2` alike | | Optional grouped convolutions | composes LiLiCorr with DFlash2's `DFlashGroupedConv`; see the dependency note below | | Two recipes | `lilicorr.yaml` and `lilicorr_conv.yaml` | | CPU unit tests, CHANGELOG, one launcher example | | **⚠️ The convolutions depend on the DFlash2 branch, and cannot run until it merges.** `modeling_lilicorr.py` imports `DFlashGroupedConv` from `modeling_dflash2`, which today exists only on `haoguo/dflash2-support`. The class is **imported rather than copied on purpose** — it is the only way the two variants cannot drift apart arithmetically — but the consequence is that the convolutional recipe cannot run against `main` as it stands. So the import is **deferred into `_install_sublayer_convs`** rather than taken at module scope. Everything else in this PR, including the plain LiLiCorr reranker, has no DFlash2 dependency at all and works on `main` today; an eager import would have made the whole plugin unimportable for the sake of one optional feature. Requesting the convolutions without DFlash2 present raises an `ImportError` naming the two config keys to remove, rather than failing at import time. **This PR carries two of @h-guo18's commits, with authorship and sign-off preserved.** Both are independent of DFlash2 itself and both are needed here: - `1419d47e`, the no-op sublayer seam. Without it `DFlashDecoderLayer.forward` never calls the wrappers the convolutions install onto, so the modules would be built, counted and exported while computing nothing. It is arithmetically an identity on its own. - `ba377e7a`, the RoPE-θ fix. On Transformers 5 a config carries both a top-level `rope_theta` and a `rope_parameters` dict; the real base lives in the dict while the class default (10,000 for Qwen3) stays visible as the flat attribute. Reading the flat field first builds a draft whose RoPE base is 100× off a Qwen3-8B target's, which trains and exports without complaint. Both the training-side enforcement and the exporter's `_get_rope_theta` are affected on `main` today. Both are @h-guo18's work and belong to their branches; they are carried here only so that this PR stands on its own. **If those branches land first, this PR can be rebased onto them and the two commits dropped**, and they can equally be split out now if that is easier to review. The same applies to `dflash_fp32_master_weights`, which is also in flight on `haoguo/dflash-fp32-master-weights`. The field name is shared deliberately so that there is only ever one knob rather than two spellings of it, and both versions default to off. Whichever lands first, this PR can be rebased onto it. ### Usage Train with the shipped recipe: ```python from modelopt.recipe import load_recipe config = load_recipe("general/speculative_decoding/lilicorr.yaml") # Qwen3-8B target, 6 epochs, block size 16 (15 drafted slots, 16 verified), # DFlash decay objective at gamma 7.0, fp32 master weights for the draft. ``` Or convert directly: ```python import modelopt.torch.speculative as mtsp config = { "dflash_block_size": 16, "dflash_loss_objective": "decay", "dflash_loss_decay_factor": 7.0, "dflash_fp32_master_weights": True, "dflash_lilicorr_w_ce": 0.25, "dflash_lilicorr_w_margin": 0.0, "dflash_lilicorr_w_pen": 0.25, "dflash_architecture_config": { "num_hidden_layers": 5, "projector_type": "lilicorr", "lilicorr_candidate_topk": 8, # Optional, and all-or-nothing: adding these two keys wraps every draft # sublayer in DFlash2's grouped convolution. Requires the DFlash2 variant. # "conv_kernel_size": 2, # "conv_group_size": 16, }, } mtsp.convert(model, [("dflash", config)]) ``` ### Results Six drafters for a **Qwen3-8B** target, all trained **in ModelOpt on one matched contract** — the same corpus, schedule and block geometry for every arm, so no row carries a training advantage. Training data is NVIDIA's [Nemotron Post-Training Dataset v2](https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2) with the multilingual split excluded, generated from the target with **thinking disabled**; **6 epochs**; block size 16 (15 drafted slots, 16 verified); DFlash decay objective at gamma 7; **8 nodes × 8 H100, global batch size 64** (one sequence per device, no gradient accumulation). All six were then exported and served through SGLang on a **single H100 80GB**, `tp_size 1`, at concurrency 1, greedy, `fa3`, mean of two replicates, with the whole node held exclusive per benchmark. Speedup is output tokens/s against an autoregressive baseline measured in the same allocation. Cells are `acceptance length / speedup-vs-AR`; **★ fastest, ☆ second fastest**: | benchmark | LiLiCorr+conv | LiLiCorr | DSpark | DFlash2 | Domino | DFlash | |---|---|---|---|---|---|---| | gsm8k | ★ 7.715 / 5.26x | ☆ 7.557 / 5.22x | 7.375 / 4.86x | 7.252 / 5.06x | 7.225 / 4.87x | 6.341 / 4.59x | | math500 | ★ 9.241 / 6.54x | ☆ 9.064 / 6.52x | 9.012 / 6.15x | 8.999 / 6.49x | 8.976 / 6.25x | 7.909 / 5.88x | | aime25 | ★ 8.285 / 6.03x | ☆ 8.156 / 6.03x | 8.043 / 5.61x | 7.967 / 5.91x | 8.066 / 5.77x | 7.126 / 5.44x | | humaneval | ★ 7.393 / 4.01x | 7.077 / 3.93x | 7.163 / 3.72x | ☆ 7.081 / 3.95x | 6.864 / 3.73x | 6.156 / 3.68x | | mbpp_sanitized | ★ 5.999 / 4.18x | ☆ 5.849 / 4.13x | 5.888 / 3.95x | 5.685 / 4.05x | 5.679 / 3.91x | 5.027 / 3.70x | | livecodebench | ★ 7.975 / 5.40x | ☆ 7.754 / 5.33x | 7.775 / 5.10x | 7.601 / 5.26x | 7.553 / 5.04x | 6.808 / 4.88x | | alpaca_eval | ☆ 3.697 / 2.69x | ★ 3.656 / 2.70x | 3.588 / 2.52x | 3.467 / 2.58x | 3.627 / 2.59x | 3.222 / 2.46x | | mtbench | ★ 4.014 / 2.94x | ☆ 3.939 / 2.93x | 3.957 / 2.78x | 3.748 / 2.80x | 3.948 / 2.84x | 3.478 / 2.67x | **Against every other approach in the table, LiLiCorr with convolutions is the fastest on all eight benchmarks.** Plain LiLiCorr is the fastest on seven of the eight; the exception is humaneval, a 164-prompt slice, where DFlash2 is ahead by 0.5%. `DFlash` is the deliberately head-free control; every head clears it by +7.60% to +21.67% on acceptance, which is the check that a head actually loaded. Reproducing the `LiLiCorr+conv` column additionally needs the DFlash2 variant. Acceptance length is bit-reproducible under greedy decoding and its replicate spread here was 0.00% on every benchmark; throughput has a ~0.2% floor. ### What `dflash_fp32_master_weights` does, and what it is worth Today the draft is cast to the frozen base model's dtype — bf16 — before the optimizer is built. AdamW then allocates its moments with `zeros_like(p)`, so the **optimizer state becomes bf16 too**. That is the problem: bf16 has too few mantissa bits to represent the small updates Adam's second moment accumulates, so those updates round away and the effective step size decays on its own, independently of the learning-rate schedule. The flag is standard mixed precision instead: the draft's master weights stay in fp32 while the matmuls run in bf16. It requires a bf16 autocast around the forward, which HF `Trainer` supplies under `TrainingArguments.bf16`. Paths that do not go through the Trainer — evaluation, `pseudo_speculative_generate`, a plain `convert()` and forward — currently need the caller to supply it, and no shipped recipe exercises those (`estimate_ar: false`, `do_eval: false`). Making the draft supply its own autocast is a follow-up, held back from here on review because it touches every DFlash variant and wants e2e coverage of the existing recipes. Compute speed is unchanged. The cost is memory, about 12 bytes per parameter for the weight plus Adam's two moments instead of 6, plus a doubled gradient all-reduce under DDP, since fp32 parameters mean fp32 gradients. Under FSDP2 that second cost is what `MixedPrecisionPolicy(reduce_dtype=...)` exists to control. It is worth **7 to 14 percent of acceptance length**, measured at the end of training on gsm8k, and it helps every projector type: | arm | bf16 | fp32 | Δ acceptance length | | --- | ---: | ---: | ---: | | LiLiCorr | 6.8670 | 7.5573 | **+10.05%** | | DFlash2 | 6.7396 | 7.2518 | **+7.60%** | | Domino | 6.5854 | 7.2252 | **+9.71%** | | DSpark | 6.4621 | 7.3752 | **+14.13%** | | DFlash | 5.9030 | 6.3412 | **+7.42%** | Every arm in the comparison table above was trained with it on, and **both shipped recipes set it `true`**, so the documented path gets it. It defaults to **off**, so no existing DFlash, Domino or DSpark run changes behaviour. Both shipped LiLiCorr recipes set it `true`, which is the arithmetic their numbers were trained with. Flipping the default is a reasonable follow-up once the autocast above is in. The draft is drawn in fp32 and, under this flag, kept there; an unpromoted run rounds the same draw to the base model's dtype. So the bf16 and fp32 rows of the table above start from the same initialization at the precision each trains in, rather than from two different draws. A unit test pins that. The flag also survives a resume. `modify()` runs under `from_pretrained` with the base model still on meta and cannot place the draft at all, so `restore_draft_precision` re-applies the dtype, the device and the rotary buffer once the weights are loaded and before the Trainer builds the optimizer — the last point that can still decide the Adam moment dtype. It also reloads the draft's tensors at the dtype they were saved in, since checkpoints store the draft in fp32 while the base is bf16 and `dtype="auto"` gives every tensor one dtype. @h-guo18 has the same field in flight on `haoguo/dflash-fp32-master-weights`, plus an HF-format-resume fix this PR does not have. The name is shared deliberately so there is only ever one knob; whichever lands first, the other should be dropped rather than merged. ### Testing - **257 CPU unit tests pass** across `tests/unit/torch/speculative/`, including the existing DFlash, Domino, DSpark and Eagle suites. 48 of them are new and cover LiLiCorr specifically: conversion routing, head geometry, the required-field validation, the three-term objective and its absolute weights, gradient reach into both the head and the drafter body, and the export contract. - Both recipes load and validate through `modelopt.recipe.load_recipe`. - The three DFlash-wide changes are covered behaviourally: the fp32 flag is checked on the optimizer's moment dtypes rather than only on parameters, since the moments are the point of the change, and on the initialization described above; activation checkpointing is asserted to leave draft gradients bit-identical with the flag on and off; and the rotary buffer is asserted present after `modify()` on a real device while still deferred on meta, which is the case the laziness existed for. - The resume path has its own test: after a `save_pretrained` / `from_pretrained` round trip, `restore_draft_precision` is asserted to return the draft to fp32 with its stored weights intact and its Adam moments in fp32. Without it the draft comes back in the base dtype with the flag still set, which is the failure it exists to prevent. - `TestDFlashLazyRotaryEmb` was updated rather than left passing: it asserted the rotary buffer does *not* exist after convert, and the DDP fix deliberately changes that on non-meta devices. The replacement pins the refined invariant in both directions. - The published checkpoints were trained with this arithmetic, verified rather than assumed: a fingerprint over draft initialisation, loss and gradients is compared against the pre-review tree for both `dflash` and `lilicorr`. Loss and gradients are **bitwise identical**. Initialisation moves, by less than bf16 resolution, and that is the single-dtype change described above. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — every addition is opt-in. The new `projector_type` is selected only by config, `dflash_fp32_master_weights` defaults to off, and the activation-checkpointing and DDP fixes preserve behaviour. No existing default changes. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new dependencies. Four files carry `# Adapted from https://github.com/sgl-project/SpecForge/...` headers for the DFlash backbone and loss they derive from (Apache-2.0), matching the attribution already on `hf_dflash.py` in this repo. The two commits described above are @h-guo18's, cherry-picked with authorship and sign-off preserved. - Did you write any new necessary tests?: ✅ — 48 new CPU tests, plus the updated rotary test. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ — will run `/claude review` once opened. ### Additional Information The convolutional recipe is the memory worst case: at an 8B target, combined with fp32 master weights, it may need `training.gradient_checkpointing: true` to fit on 80 GiB, and it fits without at 4B. Checkpointing is mathematically neutral — same objective, same data order, same resulting model — but it trades step time for memory, so a run using it is not step-time-comparable with one that does not. The recipe header says so. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added LiLiCorr speculative decoding with candidate-lattice reranking, configurable objectives, metrics, export support, and optional grouped convolutions. * Added FP32 master-weight support with improved mixed-precision behavior and gradient checkpointing. * Added LiLiCorr training recipes and a Qwen3-8B launcher configuration. * **Bug Fixes** * Improved rotary-embedding configuration handling and corrected DFlash distributed-training hangs. * Added validation for invalid LiLiCorr configurations and improved exported reranking metadata. * **Documentation** * Expanded guidance for FP32 master weights, training workflows, and LiLiCorr configuration. * **Tests** * Expanded coverage across training, evaluation, generation, export, and checkpoint workflows. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: mrusanovsky <mrusanovsky@nvidia.com> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
acdf330414 |
Add TensorRT-RTX ABI EP support for ONNX quantization (#2262)
### What does this PR do? Type of change: new feature Adds opt-in support for using the standalone TensorRT-RTX ABI Execution Provider during ModelOpt ONNX quantization. Users select the ABI backend with: `--calibration_eps=NvTensorRtRtx --trt_rtx_backend=abi` When selected, ModelOpt imports and registers the installed TensorRT-RTX ABI provider before creating the ONNX Runtime inference session. The backend selection is propagated through INT8, FP8, and INT4 AWQ calibration paths, including the Windows GenAI LLM quantization example. The existing `--calibration_eps=NvTensorRtRtx` behavior remains backward compatible. The `legacy` backend is still the default and continues to use TensorRT-RTX libraries supplied through `PATH`. For Windows x64 with Python 3.11 or newer, the ONNX dependencies now include: - `onnxruntime-gpu~=1.26.0` - `onnxruntime-ep-nv-tensorrt-rtx-cu13==0.4.0` Keeping `onnxruntime-gpu` allows users to select either CUDA EP or TensorRT-RTX ABI EP for calibration. Windows-on-Arm source-build instructions are intentionally out of scope and will be documented separately. ### Usage ```powershell python -m modelopt.onnx.quantization ` --onnx_path="C:\path\to\Llama-3.2-3B-Instruct\model.onnx" ` --model_id="C:\path\to\Llama-3.2-3B-Instruct\config.json" ` --quantize_mode=int8 ` --output_path="C:\path\to\int8_abi\model.onnx" ` --calibration_eps=NvTensorRtRtx ` --trt_rtx_backend=abi ` --use_external_data_format ` --high_precision_dtype=fp32 ` --log_level=INFO ### Testing unit test have been added ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: pending <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added optional TensorRT-RTX ABI backend support for ONNX calibration on Windows ARM64. - Added `legacy` and `abi` backend selection to quantization APIs and command-line tools; `legacy` remains the default. - Added validation for unsupported backends and incompatible TensorRT plugin configurations. - Updated Windows ARM64 installation support and platform-specific package configuration. - **Documentation** - Updated Windows installation guidance, Python compatibility requirements, ARM64 setup, and verification steps. - Documented the new TensorRT-RTX backend command-line option. - **Tests** - Added coverage for ABI provider registration, backend validation, and compatibility checks. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Haoxi Zhang <haoxiz@nvidia.com> |
||
|
|
fbd5e942e2 |
Add NVFP4 PTQ recipe for zai-org/GLM-5.3-Flash (experts + dense MLP) (#2312)
### What does this PR do? Type of change: new feature (model recipe) Adds an NVFP4 PTQ recipe for **[zai-org/GLM-5.3-Flash](https://huggingface.co/zai-org/GLM-5.3-Flash)**. GLM-5.3-Flash is a `glm5_next` VLM MoE — 45 decoder layers, 288 routed experts, and hybrid attention: KDA (linear-attention) layers interleaved with NoPE sparse-MLA layers. It requires `transformers >= 5.16.1`; earlier releases cannot parse the config. **`nvfp4_experts_dense_mlp-kv_fp8_cast`** applies: | component | precision | |---|---| | routed experts (layers 3–44, 288 each) | NVFP4 W4A4 | | dense MLP (layers 0–2, 9 modules) | NVFP4 W4A4 | | KV cache | FP8 (cast mode, constant amax) | | shared experts, router gate, KDA + MLA attention, vision tower, embeddings, `lm_head` | BF16 | `mlp_layer_types` marks only layers 0–2 `dense` and 3–44 `sparse`, so the dense-MLP scope adds just 9 modules (`mlp.gate_proj` / `mlp.up_proj` / `mlp.down_proj`) on top of the routed experts. The recipe starts from `base_disable_all`, so only the listed globs re-enable anything. > **On scope / why only one recipe.** An earlier revision of this PR also shipped a model-specific `nvfp4_experts_only-kv_fp8_cast`. It was removed: on this model it enables the **identical** quantizer set as the general `general/ptq/nvfp4_experts_only-kv_fp8_cast` (its `*block_sparse_moe*` entries are no-ops here and `default_disabled_quantizers` is redundant in experts-only scope), so it wasn't a model-specific deviation. For plain experts-only NVFP4, use the general recipe. The genuine model-specific delta — shipped here — is the dense-MLP scope plus the vision-tower exclusion below. #### The load-bearing `*visual*` disable The vision tower reuses the language model's leaf names — `model.visual.blocks.<N>.mlp.gate_proj` and friends, across 24 blocks — so the dense-MLP patterns match **144 modules inside `model.visual.*`**. Entries apply in order, so a trailing `{quantizer_name: '*visual*', enable: false}` is what keeps them BF16, and it has to stay last. (`*.experts.*` needs a literal `.experts.`, so it never reaches the vision tower.) The shared `default_disabled_quantizers` unit is deliberately not imported: for this model only its `*visual*` pattern changes anything — every other pattern either matches no module here, or matches one that `base_disable_all` already left off (`lm_head`, the `mlp.gate.` routers) and that nothing re-enables. #### Two model-specific points, documented in the file header - **`layerwise.enable=false` is required, not incidental.** This is a VLM, so the decoder layers nest under `model.language_model.layers` and `layerwise_calibrate` cannot locate them. - **The MTP head is not built, so it is neither quantized nor exported.** The config declares `num_hidden_layers: 45` (with `num_nextn_predict_layers: 1`), so the HF model class instantiates decoder layers 0–44 only and never constructs the MTP layer. Filed under `modelopt_recipes/models/` per the split introduced in #2219, keyed by the source hub model — alongside `moonshotai/Kimi-K3` and `mistralai/Mistral-Medium-3.5-128B`. There is no published `nvidia/GLM-5.3-Flash-NVFP4` yet; the `models/` section explicitly covers "published **(or planned)**" checkpoints. ### Usage ```bash python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path <zai-org/GLM-5.3-Flash checkpoint> \ --recipe models/zai-org/GLM-5.3-Flash/ptq/nvfp4_experts_dense_mlp-kv_fp8_cast \ --export_path <output> ``` ### Testing - **`tests/unit/recipe/test_glm_5_3_recipe.py`** (new) — applies the recipe to a tiny `glm5_next`-like VLM MoE and asserts the enabled/disabled state per module: routed experts + dense MLP → NVFP4; vision tower, shared experts, router gate, KDA `conv1d`, MLA attention and `lm_head` → BF16. This pins the wildcard precedence — in particular that the trailing `*visual*` disable keeps the vision tower BF16 even though it reuses the dense-MLP leaf names, and that `*mlp.gate_proj*` doesn't catch the router `mlp.gate`. - **`tests/unit/recipe/test_recipe_docs.py`** — all checks pass, including `test_every_model_specific_ptq_dir_is_mentioned` (the `models/zai-org/GLM-5.3-Flash/ptq/` folder appears in `ptq.md`). The recipe's scope was also checked against the model's actual module names: the dense-MLP patterns match 144 modules under `model.visual.*`, which the trailing disable returns to BF16; an exported checkpoint carries `input_scale` / `weight_scale` / `weight_scale_2` on `layers.0–2.mlp.*_proj` while `visual.blocks.0.mlp.gate_proj` retains only `.weight` / `.bias`; and `kv_cache_quant_algo: FP8` survives the trailing disable. ### Before your PR is "*Ready for review*" - Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed. - Is this change backward compatible?: ✅ (one new recipe + one new unit test; `ptq.md` updated) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — `tests/unit/recipe/test_glm_5_3_recipe.py` pins the recipe's wildcard precedence - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — consistent with the recent recipe additions (#2219, #2269, #2287), which did not add entries - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information Source model: https://huggingface.co/zai-org/GLM-5.3-Flash <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added post-training quantization recipes for GLM-5.3-Flash. * Supports NVFP4 W4A4 quantization for routed experts, with an additional configuration covering dense MLP layers. * Enables FP8 key-value cache casting. * Uses maximum-based calibration with layerwise calibration disabled for the VLM layout. * Retains BF16 precision for shared experts, vision components, attention, embeddings, routing, language head, and MTP components. * **Documentation** * Documented the available GLM-5.3-Flash quantization configurations and precision assignments. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
3f4b95e8a4 |
Add the NVFP4 PTQ recipe for Qwen/Qwen3.8-2.4T-A95B (#2302)
### What does this PR do? Type of change: new feature (model recipe) Adds the NVFP4 PTQ recipe for **Qwen/Qwen3.8-2.4T-A95B** — the recipe used to produce [**nvidia/Qwen3.8-2.4T-A95B-NVFP4**](https://huggingface.co/nvidia/Qwen3.8-2.4T-A95B-NVFP4). Qwen/Qwen3.8-2.4T-A95B is a `qwen3_5_moe_text` MoE: 92 layers, 512 routed experts (top-10) plus a shared expert, with **hybrid attention** — gated-delta (linear-attention) layers interleaved with full-attention layers. It is transformers-native from >= 5.9 and its config ships `base_model_ep_plan`, so no ModelOpt plugin is required. The recipe applies: | component | precision | |---|---| | routed experts | NVFP4 (MSE-searched static weight scales, dynamic input scales) | | self-attention | FP8 (W8A8, all projections) | | linear-attention | FP8 (W8A8, the full gated-delta path — `conv1d` + all in/out projections) | | KV cache | FP8 (cast mode) | | everything else | BF16 — including MTP, left unquantized | Two things are documented in the file header because they affect how the recipe should be read: - **The full gated-delta path is FP8, and that was validated end-to-end.** The `conv1d` and the in/out projections (`in_proj_qkv` / `in_proj_z` / `in_proj_a` / `in_proj_b`, `out_proj`) are all FP8; only the norms stay BF16. `nn.Conv1d` is a registered ModelOpt quant module, so the recipe's broad `*linear_attn*` rules reach `linear_attn.conv1d` too — this is intentional and matches the published `nvidia/Qwen3.8-2.4T-A95B-NVFP4`, whose `hf_quant_config.json` lists `linear_attn.conv1d` as FP8 on every gated-delta layer (the interleaved full-attention layers have no `conv1d`). (An earlier revision of the file header / ptq.md wrongly stated the recurrent path is never quantized; corrected in this PR.) - **The source ships as native block-FP8** (`quant_method=fp8`, `weight_block_size [128,128]`, dynamic activations). The loader dequantizes it to BF16 before quantizers are inserted, so the calibrated scales are against BF16 weights, not against the shipped FP8. Filed under `modelopt_recipes/models/` per the split introduced in #2219 (per-`model_type` recipes vs model-hub checkpoint recipes); this one targets a published checkpoint, alongside `deepseek-ai/DeepSeek-V4-Pro-0813` and the Nemotron-3 entries. ### Usage ```bash # The recipe is consumed by the PTQ entrypoint the same way as the other # modelopt_recipes/models/ entries: python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path <Qwen/Qwen3.8-2.4T-A95B checkpoint> \ --recipe models/Qwen/Qwen3.8-2.4T-A95B/ptq/nvfp4_experts_mse-fp8_self_attn-fp8_linear_attn-kv_fp8_cast \ --export_path <output> ``` ### Testing The exported checkpoint was evaluated against the BF16 baseline on **GPQA, AA-LCR, SciCode, IFBench and Terminal-Bench 2.1**, with no meaningful accuracy regression on any of them. The published `nvidia/Qwen3.8-2.4T-A95B-NVFP4` checkpoint is the artifact this recipe produces — its `hf_quant_config.json` is the ground truth for which modules are quantized (routed experts NVFP4; self-attention, all linear-attention projections **and** `conv1d`, and KV cache FP8). No new unit tests: this is a declarative recipe composed entirely of existing units (`base_disable_all`, `nvfp4`, `nvfp4_static`, `fp8`, `kv_fp8_cast`), all already covered. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (new file only; no existing behaviour touched) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — declarative recipe over existing, tested units - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — consistent with the recent recipe additions (#2219, #2269, #2287), which did not add entries - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information Model card: https://huggingface.co/nvidia/Qwen3.8-2.4T-A95B-NVFP4 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a post-training quantization recipe for the Qwen3.8-2.4T-A95B model. * Supports MSE-searched NVFP4 quantization for routed expert layers and FP8 quantization across self-attention and gated-delta linear-attention paths. * Supports FP8 cast-mode key-value caching while retaining BF16 precision for multi-token prediction and gated-delta normalization layers. * **Documentation** * Clarified the model’s hybrid precision configuration, including FP8 treatment of the gated-delta convolution path and the scope of broad linear-attention patterns. * Documented source-checkpoint dequantization and validation behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
0688761ce9 |
feat(export): support multimodal and MTP models in layerwise export (#2303)
### What does this PR do? Type of change: New feature **Layerwise export now supports multimodal and MTP models.** Both were refused outright, and both were refused for the same reason: `finalize()` was called from inside `layerwise_calibrate`, which is the wrong scope for it. **1. Calibration does not know which model the checkpoint describes.** It only sees the module it was handed. A VLM calibrates its *language model*, so the shards, the exclusions and `config.json` all came out describing that submodel rather than the whole VLM. Moving the call out lets the caller root the exporter at the parent — and without the key prefixing, tower collection or ambient parent handle an earlier attempt needed, because the decoder layers are the same objects from either root. **2. Calibration runs before things the export needs exist.** Orphaned MTP weights are loaded *after* calibration, by which point every shard had already been written, so they could not be passed at all. After the move they are an ordinary argument to `finalize()`, with no staging attribute stashed on the model. ### How it works The exporter is created by whoever owns the export and **announced on the model** that `mtq.quantize` is given. Calibration picks it up, binds it, and drives it per layer; the export that follows reads it back and finishes the checkpoint: ```python LayerwiseExporter(full_model, export_path).announce(language_model) mtq.quantize(language_model, quant_cfg, forward_loop=loop) ... getattr(full_model, LAYERWISE_EXPORTER_ATTR).finalize(extra_state_dict=mtp_state_dict) ``` Calibration and export are handed *different* models, so `announce()` publishes the exporter on each end separately: the caller announces on the model being calibrated, and `bind()` announces on the export root. Neither side has to know where the other looked, and the lookup stays an O(1) `getattr` rather than a `named_modules()` scan — worth avoiding at roughly 1.65 µs/module, or ~500 ms on a Kimi-K3-sized model. For a non-VLM both roots are the same object and the second announcement is a no-op. `finalize()` clears every attachment it recorded, so the module graph does not retain a live exporter afterwards. `mtq.quantize` and `mtq.calibrate` are **unchanged** — a layerwise-only feature does not belong in the public quantization API. The attribute follows `_mtp_layer_prefixes`, which crosses the same calibration→export boundary the same way (`hf_ptq.py:538` sets it, `unified_export_hf.py:870` reads it back). Construction is inert: `__init__` records only the export root and the directory, because the caller builds it before `mtq.quantize`, when there are no quantizers yet to validate or read a config from. `bind()` does that, called from calibration after quantizer insertion and before any layer is converted — the only window where both hold, and the same instant the exporter used to be constructed, so unsupported models still fail in seconds rather than hours. Only the calibration pass that sets `export_dir` drives the exporter: a list-form algorithm runs one pass per entry, and an earlier one must not convert layers a later one still has to calibrate. ### Usage Nothing changes for a plain layerwise-export recipe: `layerwise.export_dir` still drives it. Pre-attaching an exporter is the opt-in for the two cases that need it — a checkpoint whose root is wider than the calibrated model, and orphaned tensors to merge at the end. The one behaviour change for a config-only caller is that `mtq.quantize` now writes the layer shards but no longer finishes the checkpoint. Both exit paths warn with what is still owed, and `LayerwiseConfig.export_dir`'s description has been corrected — it previously promised "a complete, loadable checkpoint when the last layer lands" and still listed multimodal and MTP as raising `NotImplementedError`. ### Testing `tests/gpu/torch/export/test_layerwise_export.py` — **29 passed**. Beyond the 24 inherited from #2136, five new ones, each with a negative control confirming it fails without its fix: - orphaned MTP tensors reach the tail shard *and* the index - an exporter rooted at the parent widens the checkpoint's namespace - the config-only path announces an exporter that can be finished, and finalize clears it - only the pass that sets `export_dir` drives the exporter - an exporter whose root holds a different number of layers is refused at `bind()` Full suites: `tests/gpu/torch/export` + `tests/gpu/torch/quantization` **1012 passed / 55 skipped**, `tests/unit` **3318 passed / 15 skipped**, pre-commit clean. Both suites also report failures in `test_implicit_gemm.py` (FP4 conv kernels), `test_triton_fa_p_qdq.py`, `test_autocast_quantize_int8` and `test_engine_builder.py` collection; all reproduce unchanged on `main` and none touch the paths in this diff. Measured against the whole-model exporter on a tiny Gemma3-VL, towers prepared exactly as `hf_ptq` does: ``` keys: baseline=80 layerwise=80 only-baseline=[] only-layerwise=[] differing values: 0 vision tower present: True VLM namespace: True config.json is the VLM: True hf_quant_config match: True exclude_modules: ['language_model.lm_head', 'vision_tower.vision_model*'] (both sides) ``` #### End-to-end through `hf_ptq.py` Same FP8 recipe both sides; the baseline drops `layerwise.export_dir` and is exported by `main`, so the diff isolates this PR. Every tensor matches in key, dtype, shape and value, and `config.json` / `hf_quant_config.json` match too. | Model | Covers | Keys | Differing | |---|---|---|---| | Qwen3-VL-8B-Instruct | multimodal | 1254 = 1254 | 0 | | GLM-4.7-Flash | MoE + MTP | 28119 = 28119 | 0 | The VLM checkpoint keeps the vision tower unquantized (351 `model.visual.*` keys, no `weight_scale` among them) while the language model is FP8. The MTP run reports 212 orphaned tensors; all 212 land in `model-tail.safetensors` and in the index, with `model.layers.47*` in `exclude_modules`. **Not yet validated:** an accelerate-offloaded run, and a serving canary on the exported checkpoints. ### Refusals `export_dir` without `enable`, and an exporting algorithm entry with no calibration method, are both refused before calibration starts — neither reaches the per-layer pass, so both would otherwise export nothing. The early gate is a heuristic on the recipe, so `hf_ptq` also raises a plain `RuntimeError` at export time if calibration turned out not to have run; that backstop, not the gate, is what makes the failure legible on paths the recipe check cannot predict. `bind()` requires the layers calibration will drive and refuses a root that discovers a different number of them. Only the count is checked here: `export_layer` already rejects a reordering or a substituted module on its first call, and a length difference is the one mismatch it structurally cannot catch — every call would pass and `_write_index` would then open a shard that was never written, at the very end of the run. Orphan tensors are merged into the tail with no collision check, matching the whole-model path (`unified_export_hf.py:1623`). `load_mtp_weights` returns exactly the keys absent from `model.state_dict()`, so a collision with an exported tensor is not reachable through the only producer, and a guard would only make the two export paths diverge. ### Why not reuse `export_hf_checkpoint` It was the first idea and it is the most expensive one. Its transformers path is whole-model at every step — `_prepare_moe_inputs`, `requantize_resmooth_fused_llm_layers` (which runs a dummy forward that would fail on already-converted layers), `_process_quantized_modules`, a full `model.state_dict()` in host RAM, then `save_pretrained` rewriting shards already on disk — and it raises outright under `has_accelerate_offload`. `save_pretrained(state_dict={})` is not an escape either: safetensors' shared-storage check fires on MoE even with an empty dict. The natural consolidation target is the **streaming** exporter, which is already most of `finalize()`: 122 lines vs 74, sharing `decoder_owned_ids`, `enable_weight_access_and_writeback`, `_dispatch_export_handler`, `_reconstruct_fused_moe_linear`, `_add_mtp_exclusions`, `_postprocess_single_tensor`, `requires_weight_materialization` and `save_non_weight_artifacts`. Folding them together needs roughly four knobs: skip the whole-model prep, skip layers already written, seed the index with the existing shards, and inject the quant config. That is a separate change and deliberately not in this one. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `mtq.quantize`/`mtq.calibrate` signatures are unchanged, and a recipe that only sets `layerwise.export_dir` behaves as before. The one behaviour change is that `mtq.quantize` no longer finishes the checkpoint on its own: callers must now call `finalize()` on the exporter, which calibration leaves on the model. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ❌ — pending. - Did you get Claude approval on this PR?: ❌ — the last review's findings are all addressed; needs a re-run. ### Additional Information Follow-ups this enables: #2259 (MTP) reduces to close to nothing, and the multimodal work in #2218 no longer needs `export_parent`, the key prefixing, or the tower collection. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Layerwise export now supports clearer control over export locations and calibrated layer handling. - Export workflows provide improved support for resuming, sharded checkpoints, mixture-of-experts models, and nested model namespaces. - **Bug Fixes** - Improved handling of exported checkpoint shards and extra tensors. - Added clearer warnings when exports require completion before loading. - **Documentation** - Clarified that layerwise exports write shards during calibration and require an explicit finalization step. - Documented that the in-memory model is not suitable for inference after layerwise export. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |