mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
main
556
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4f62d418f4 |
Use TensorRT optimization level 0 in Torch ONNX example tests (#2619)
### What does this PR do?
Type of change: Bug fix
The Torch ONNX example tests can exceed their 300-second deadline while
building a ResNet50 INT8 TensorRT engine at optimization level 4. Add
`--trt_builder_optimization_level` to the vision example and select
level 0 in the existing integration tests. The example and helper retain
level 4 by default. Quantization, ONNX export, residual Q/DQ assertions,
engine execution, and test timeout limits are unchanged.
Document the build-time versus inference-performance tradeoff in the
example README.
### Usage
```bash
cd examples/torch_onnx
python torch_quant_to_onnx.py \
--timm_model_name resnet50 \
--recipe timm/resnet/ptq/int8 \
--onnx_save_path resnet50.int8.onnx \
--calibration_data_size 1 --no_pretrained \
--trt_build --trt_builder_optimization_level 0
```
### Testing
Validation used the TensorRT 26.05 container, TensorRT 10.16.1.11, and
PyTorch 2.13.0, with the existing 300-second per-test deadline.
- RTX 6000 Ada: **8 passed**, covering FP8 and INT8 on ViT, Swin,
SwinV2, and ResNet50. ResNet50 INT8 passed in 96.65 seconds.
- RTX PRO 6000 Blackwell Max-Q: **20 passed, 3 existing skips**,
covering the complete test file. ResNet50 INT8 passed in 81.27 seconds;
the baseline timed out at 300 seconds.
```bash
# RTX 6000 Ada: supported FP8/INT8 cases
python -m pytest tests/examples/torch_onnx/test_torch_quant_to_onnx.py \
-k '(fp8 or int8) and not mxfp8' --cov
# RTX PRO 6000 Blackwell: complete example test file
python -m pytest tests/examples/torch_onnx/test_torch_quant_to_onnx.py --cov
```
On each GPU, levels 4 and 0 used the same exported ResNet50 INT8 graph
with TensorRT 10.16.1.11:
- RTX 6000 Ada: TensorRT-reported engine build time decreased from 122.6
seconds at level 4 to 27.7 seconds at level 0. Both builds and inference
runs succeeded.
- RTX PRO 6000 Blackwell Max-Q: the original level-4 test hit its
300-second deadline during the engine build; level 0 built that saved
graph in 10.9 seconds and completed inference successfully.
All pre-commit checks passed for the changed files. The existing
integration tests exercise the real engine build; no redundant mocked
tests were added.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — the existing integration
tests were updated to exercise level 0; no new test cases were needed.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — minor example/CI fix; library behavior and example defaults are
preserved.
- Did you get Claude approval on this PR?: N/A — not requested for this
focused change.
### Additional Information
Example timeout: [ResNet50 INT8 CI
failure](https://github.com/NVIDIA/Model-Optimizer/actions/runs/36560176214/job/109381161169).
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added a configurable TensorRT builder optimization level for engine
builds, with a default of 4 and support for values from 0 to 5.
* Documented that level 0 can speed up builds, while lower optimization
levels may reduce inference performance.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
|
||
|
|
9c1cf80f1b |
test(megatron_bridge): cover context parallelism in the VLM QAD test (#2592)
### What does this PR do? Type of change: new tests Runs the VLM case of `test_qad` under context parallelism, so QAD on a Qwen3-VL model is covered on the path that until now could not run at all. `Qwen3VLMultimodalRotaryEmbedding` CP-shards its own embedding, so the batch has to hand it full-length `position_ids`. Megatron-Bridge's `get_batch` was sharding them too, leaving the rotary embedding at `seq / cp**2` against hidden states at `seq / cp`. The fix is upstream in [NVIDIA-NeMo/Megatron-Bridge#6243](https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/6243); this PR is the coverage that would have caught it. The VLM case moves from tensor to context parallelism and the PTQ step is sized to the same TP, so QAD still loads a matching checkpoint. The LLM case is unchanged (`tp_size=num_gpus, cp_size=1`), and `test_distill_vlm` still covers TP for a VLM, so nothing loses coverage. ### Usage ```bash # Unchanged: --cp_size is already a distill.py flag. On a container carrying Megatron-Bridge#6243 # it now works for VLMs, where it previously died in the rotary embedding. python examples/megatron_bridge/distill.py --cp_size 2 --tp_size 1 ... ``` ### Testing On 2x RTX 6000 Ada, in `nemo:26.08` with Megatron-Bridge#6243 on `PYTHONPATH`: - `test_qad[qwen3_5_moe_vl]` at `--tp_size 1 --cp_size 2` — FP8 PTQ, QAD across 2 CP ranks, export; quantizers survive and the vision tower is byte-identical. **1 passed (183 s).** Without the upstream fix the same run dies with `AttributeError: 'NoneType' object has no attribute 'ndim'` in `rope.py:175`. - `test_qad[qwen3]`, the unchanged LLM path — **1 passed (194 s).** - Gate check: on today's `nemo:26.08` (no #6243) the probe resolves `False` and the VLM case stays at `cp_size=1`, byte-identical to current CI; with #6243 it resolves `True` and runs at `cp_size=num_gpus`. On a 1-GPU runner it degenerates to today's config either way. - `pre-commit run --files ...` clean (ruff check, ruff format, mypy, bandit, markdownlint). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — existing tests extended rather than new ones added. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — test coverage and one doc line; no feature, break, deprecation, or fix for a released bug. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information - Depends on [NVIDIA-NeMo/Megatron-Bridge#6243](https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/6243). Safe to merge before it lands: the gate keeps the VLM case at `cp_size=1` until a container ships the fix, at which point the coverage switches on by itself. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the Qwen3.6 QAD instructions to keep tensor and pipeline parallelism set to 1, while allowing context parallelism to increase for longer sequences with the `nemo:26.10` container. * **Tests** * QAD validation now selects parallelism settings based on whether the Megatron-Bridge context-parallel fix is available, and reports when multi-GPU VLM coverage is reduced. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
aa89722d38 |
[6/6] Add the IQ1_M CUDA encoder and register the format (#2595)
### What does this PR do? Type of change: new feature **Second of two PRs adding IQ1_M** (1.75 bits per weight). #2513 landed the PyTorch codec; this PR adds its **CUDA encoder** and makes the format reachable. With it, ModelOpt supports all five GGML IQ formats at one and two bits. - the CUDA encoder, its binding and extension build wiring, plus the CUDA path in `quantize_iq1_m` - an `IQFormat` record and **one `IQ_FORMAT_REGISTRY` entry**, so backend dispatch, both exporters and `convert_hf_config` take it from there - the `ggml` package export - the `general/ptq/iq1_m` recipe, its presets, `ptq.md` and a CHANGELOG entry The kernel lands with the registration so every registered format keeps a CUDA encoder. ### The kernel In the kernel the delta shift is free per group, so it sits above the entry index in the sort key: a tie still prefers the lower shift and then the lower entry, as the reference encoder does. The 2048-entry grid IQ1_M shares with IQ1_S is 64 KiB, past the 48 KiB static shared-memory limit, so both kernels read it from global memory and rely on the cache. | 5632×2048 weight | torch | CUDA | | |---|---|---|---| | IQ1_M encode | 5.6 M elem/s | **318 M elem/s** | **57×** | ### Shared with IQ1_S rather than copied The two IQ1 kernels load each vector, score it against a grid entry and apply the ±1/8 shift the same way. So those three steps move into `common.cuh` as `load_vector`, `grid_terms` and `shifted_error`, and IQ1_S uses them too. **IQ1_S's packed bytes are unchanged**: its CUDA output on a 5632×2048 weight hashes the same before and after, and so does IQ1_M's, compared against the pre-split version of this change. IQ1_S encodes at the same speed (306 M elem/s). ### Usage ```bash python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq1_m ``` ### Testing Registering the format brings it under every registry-driven test with no IQ1_M-specific test code: backend dispatch, weight caching, the `num_bits` guard, `convert_hf_config` metadata, Megatron export and the `TensorQuantizer` tests in the shared battery. The shared CUDA battery gains one row. - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py`, `tests/unit/recipe/test_presets.py`: **166 passed** - broader unit sweep (`-k 'ggml or iq or gguf or registry'` over quantization, export and recipe tests): **221 passed**. The one failure, `test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes`, is a `torchvision` import error in my environment, unrelated to IQ. - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **49 passed** on RTX PRO 6000 Blackwell (sm_120), 7 of them IQ1_M, including CUDA-vs-PyTorch encoder parity - `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k 'iq or ggml'`: **45 passed** (9 tests × 5 formats) in `nvcr.io/nvidia/nemo:26.08` - `tests/examples/hf_ptq/test_llm_ptq.py -k iq1_m`: **passed** - reconstruction error falls monotonically across all five formats, pinned by a test - `general/ptq` now holds 31 recipes; `ptq.md` is updated. Rebased onto `main` after #2513 merged. The resulting tree is identical to the one the runs above tested, and the unit set was rerun on it: 166 passed. On this GPU, two of #2515's Q8_0 tests in `tests/gpu/_extensions/test_torch_extensions.py` fail: `test_cuda_ext_q8_0_zero_and_roundf_layout` and `test_cuda_ext_q8_0_dequantizes_with_small_error`. They fail identically on a clean `main` checkout, so they are not from this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code sources or dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S codec) → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M codec), all merged → **this**. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added IQ1_M weight-only quantization at 1.75 bits per weight, with CUDA acceleration and a 256-value block size. * Added an IQ1_M post-training quantization recipe for eligible linear layers; calibration data is not required. * Added IQ1_M to the supported GGML-compatible formats and recipe listings. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
333ace1bc9 |
Fix multi-turn synthetic generation and mark incomplete outputs (#2571)
### What does this PR do?
Type of change: Bug fix
Fixes synthetic conversation generation that assumed alternating user
and assistant messages. That assumption skipped user turns in prompt
skeletons and mishandled
leading system messages.
• Preserve conversation history. Regenerate every user turn while
retaining system messages and generated reasoning for subsequent
requests.
• Expose generation controls. Support model-specific request parameters,
configurable timeouts, and server-managed response budgets.
• Handle failures explicitly. Reject empty final answers and unsupported
tool calls. Failed conversations remain retryable without duplicating
saved output.
• Identify incomplete outputs. Mark length- and repetition-stopped
conversations as truncated, preserve stop metadata, and stop generating
follow-up turns.
### Usage
Run from the repository root against a compatible Qwen server with
reasoning parsing enabled:
python examples/speculative_decoding/scripts/server_generate.py \
--data_path input_conversations/train.jsonl \
--output_path synthetic/train.jsonl \
--url http://localhost:8000/v1 \
--model model \
--max_tokens 0 \
--request_timeout 3600 \
--extra_body
'{"chat_template_kwargs":{"enable_thinking":true,"preserve_thinking":true}}'
The model name must match the server’s configured name. Filter truncated
conversations before training.
### Testing
Focused regression tests: 15 passed.
The tests execute the command-line entry point using the real OpenAI
client library with mocked HTTP transport.
Coverage includes multi-turn generation, system prompts, reasoning
preservation, request parameters, failure recovery, resume
deduplication, truncation, and invalid
responses.
python -m pytest \
--confcutdir=tests/examples/speculative_decoding \
tests/examples/speculative_decoding/test_server_generate.py -q
The isolated test configuration avoids an unrelated parent configuration
import failure. All applicable pre-commit checks passed for the
generator, tests, and
documentation.
### Before your PR is "Ready for review"
• Is this change backward compatible?: ✅ Existing valid inputs,
defaults, conversation output structure, and resume behavior remain
supported. Invalid inputs and
failed requests now raise errors instead of being silently accepted.
• Copied code or new PIP dependencies?: N/A. No new third-party code or
dependencies were added.
• Did you write any new necessary tests?: ✅ Added focused command-line
regression tests.
• Did you update Changelog?: N/A. These are example-script correctness
fixes, not critical released library fixes.
• Did you get Claude approval on this PR?: ❌ Not yet obtained.
### Additional Information
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Data generation supports `conversations` and `messages` inputs,
preserves reasoning content, and accepts additional chat settings and
configurable request timeouts.
* Failed conversations are recorded separately, with options to retry
failures or exit when errors occur. Resume behavior distinguishes
retryable failures from rejected inputs.
* Outputs identify conversations truncated by length or repetition
limits.
* **Documentation**
* Updated data preparation guides with generation setup, input formats,
failure handling, resuming, and training guidance.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
|
||
|
|
e5b63320ab |
[5/6] Add the IQ1_M codec (#2513)
### What does this PR do? Type of change: new feature (not yet user-reachable) **First of two PRs adding IQ1_M** at 1.75 bits per weight, just above IQ1_S. This one lands the **PyTorch codec**: encoder and decoder. It is deliberately **not registered**, so no quantizer dispatches to it and the `ggml` package does not export it. #2595 adds the CUDA encoder, registers the format and adds its recipe. With both, ModelOpt supports all five GGML IQ formats at one and two bits. On the mixed-precision checkpoint #2511 measured (`unsloth/Qwen3.8-27B-GGUF`), IQ1_M covers **25 tensors and 1.2 B parameters**. With all five formats we can read 89.0% of that file; the rest is k-quants and F32. ### What's distinctive about it **IQ1_M is the most irregular layout of the five.** There is no leading block scale field at all. The FP16 super-block scale is reassembled from the top nibble of each of four scale words: ```c scale.u16 = (sc[0] >> 12) | ((sc[1] >> 8) & 0x00f0) | ((sc[2] >> 4) & 0x0f00) | (sc[3] & 0xf000); ``` It is also finer grained than IQ1_S: a local scale per **two** groups rather than four, and a delta shift chosen **per group** rather than per sub-block. That is where its extra 0.1875 bits go. ### Shared with IQ1_S rather than copied IQ1_M searches exactly as IQ1_S does: the same 2048-entry grid, the same ±1/8 delta, every (shift, local scale) choice for every 8-value vector. It differs only in how it selects among those choices afterwards. So the search moves out of IQ1_S's encoder into `_search_shifted_grid`, which both call, and `iq1_m.py` keeps only its selection and packing. **IQ1_S's encoded bytes are unchanged**, checked by hashing its output before and after on a fixed input. ### A scale-anchor correction IQ1_M anchors its scale differently from IQ1_S: the ratio **rises with a block's peak-to-RMS** rather than being flat, and clamps higher. It uses `clamp(0.58 + 0.035 * peak_to_rms, 0.65, 0.95)` against IQ1_S's flat `0.61`. Measured over 15 Qwen3.8-27B MLP weights: | | flat 0.61 | correct anchor | | |---|---|---|---| | relative reconstruction MSE | 0.17372 | **0.17291** | **−0.47%** | It is consistent on every tensor, with no outliers. The anchor changes quality without touching layout, so neither round-trip nor conformance tests would catch it drifting. `test_scale_anchor_follows_peak_to_rms` now pins it, for all five formats; see Testing. ### Family parity Two surface asymmetries close here, so the five are uniform. `IQ1_S` now exposes `_predict_iq1_s_scales` like the other four, instead of computing its anchor inline. `IQ1_M` exposes `iq1_m_grid`, aliasing the IQ1_S table it shares. ### Testing **The decoder is validated against llama.cpp's own output, not just round-tripped:** ``` IQ1_M: 25 tensors, 4,730,880 blocks → 0 mismatched, max|diff| 0.0 ``` This mattered: **my first IQ1_M decoder had a real bug.** A `repeat_interleave` on the wrong axis produced `[h0,h1,h0,h1]` where llama.cpp needs `[h0,h0,h1,h1]`. A round-trip against our own encoder still passed, because the encoder made the matching mistake. Only comparison against bytes we did not produce caught it. Blocks from that checkpoint ship as conformance vectors, and mutation testing confirms they catch a mis-set scale nibble. The decoder unpacks every field in one vectorized pass, since fake quant decodes on every forward: 5.2 ms for a 5632×2048 weight (IQ1_S: 3.3). - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py`, `tests/unit/recipe/test_presets.py`: **153 passed**, 15 of them IQ1_M codec cases, including the llama.cpp conformance check - `test_scale_anchor_follows_peak_to_rms` pins every format's scale anchor. It predicts scales for blocks whose peak-to-RMS is exactly 1, 4, 8 and 16, reaching both clamps and two points on each slope, and compares them against anchors written out in the test. Mutations each fail exactly the mutated format: reverting IQ1_M to IQ1_S's flat 0.61, moving either IQ1_M clamp, changing its taper by 0.001, moving an IQ2_S or IQ2_XS clamp, and changing IQ1_S's anchor to 0.62. - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed**. IQ1_S's CUDA-vs-PyTorch parity still holds after its encoder refactor. - IQ1_S and IQ1_M PyTorch encoder output and IQ1_M decoder output hash identically to the pre-split version of this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ IQ1_M adds no codebook; it reuses the IQ1_S table already carried in `codebooks.py`. The new conformance vectors come from `unsloth/Qwen3.8-27B-GGUF`, which is Apache-2.0 like its base model `Qwen/Qwen3.8-27B`; the vectors' docstring now records that. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A. Nothing is user-reachable yet; #2595 carries the entry. - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S codec) → #2565 (IQ2_S CUDA encoder and registration), all merged → **this** → #2595 (IQ1_M CUDA encoder and registration). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added IQ1_M quantization and dequantization for compact, GGML-compatible blocks of 256 values. * Added access to the IQ1_M grid and configurable chunk sizes for processing data. * **Bug Fixes** * Improved IQ1_S scale prediction and grid-search organization while preserving its encoding behavior. * **Tests** * Added IQ1_M conformance data and included the format in shared IQ-format test coverage. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
ad9ea97a4b |
Fix vLLM compilation guard for models without marker (#2518)
### What does this PR do?
Type of change: Bug fix
Makes the vLLM `disable_compilation` context manager support inner model
implementations that do not predefine a `do_not_compile` attribute,
including GLM-5.3. The context manager now installs the marker
temporarily and removes it afterward, while preserving and restoring
existing marker values for other vLLM models.
Adds regression coverage for both supported wrapper layouts:
`model.model` and `model.language_model.model`.
### Usage
```python
with disable_compilation(model):
mtq.quantize(model, quant_cfg, forward_loop=calibrate_loop)
```
No caller changes are required.
### Testing
- Ran `tests/gpu_vllm/torch/quantization/test_vllm_dynamic_modules.py`:
24 passed with vLLM 0.28.
- Ran pre-commit on both changed files: all applicable hooks passed.
- Installed this branch into `vllm/vllm-openai:glm53-flash` on OCI-JHB
and served the GLM-5.3-Flash BF16 checkpoint with
`QUANT_CFG=NVFP4_DEFAULT_CFG`, TP=4, eager mode, and BF16 KV cache.
- GLM passed the previous `do_not_compile` failure point, inserted 1,700
quantizers, enabled 456 weight quantizers, reached a healthy API server,
and returned a relevant manual prompt response.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — integration compatibility fix; no user-facing API change.
- Did you get Claude approval on this PR?: N/A
### Additional Information
Validated against GLM-5.3-Flash using ModelOpt commit
`869b64fcee0b20be323663449b00e8c52940a289`.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Bug Fixes**
- Compilation settings are now handled across supported nested model
configurations and restored after calibration, including when errors
occur.
- Calibration inputs correctly exclude padding when an attention mask is
provided and reject empty sequences.
- vLLM warmup reserves the required cache space for supported tail-cache
configurations.
- Serving startup supports an alternate vLLM launcher import path when
the OpenAI entrypoint is unavailable.
- **Compatibility**
- The vLLM serving example now defaults to vLLM 0.30.0 and documents
tested support for Nemotron 3 Nano hybrid attention/Mamba serving on
vLLM 0.26.0 and 0.30.0.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
|
||
|
|
3091b8ff69 |
[4/5] Add the IQ2_S CUDA encoder and register the format (#2565)
### What does this PR do? Type of change: new feature **Second of two PRs adding IQ2_S** (2.5625 bits per weight). #2512 landed the PyTorch codec; this PR adds its **CUDA encoder** and makes the format reachable: - the CUDA encoder, its binding and extension build wiring, plus the CUDA path in `quantize_iq2_s` - an `IQFormat` record and **one `IQ_FORMAT_REGISTRY` entry**, so backend dispatch, both exporters and `convert_hf_config` take it from there - the `ggml` package export - the `general/ptq/iq2_s` recipe, its presets, `ptq.md` and a CHANGELOG entry The kernel lands with the registration so every registered format keeps a CUDA encoder. On the mixed-precision checkpoint #2511 measured (`unsloth/Qwen3.8-27B-GGUF`), IQ2_S covers **9 tensors and 0.6 B parameters**. ### The kernel IQ2_S's **1024-entry codebook is twice IQ2_XS's**, which makes its search the most expensive in the family. The codebook and its norms take 36 KiB of shared memory, the most of any IQ kernel but inside the 48 KiB static limit, so they are declared statically like the IQ2_XS and IQ2_XXS kernels. That cost is why the kernel matters more here than anywhere else: | | torch | CUDA | | |---|---|---|---| | IQ2_S, 5632×2048 weight | 0.8 M elem/s | **725.7 M elem/s** | **907×** | | extrapolated to a 27B model | ~9.8 hours | **~37 s** | | ### Usage ```bash python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq2_s ``` ### Testing Registering the format brings it under every registry-driven test with no IQ2_S-specific test code: backend dispatch and weight caching, the `num_bits` guard, `convert_hf_config` metadata (uniform and mixed precision), all 9 Megatron export tests, and the two `TensorQuantizer` tests in the shared battery. The shared CUDA battery gains one row. - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py`, `tests/unit/recipe/test_presets.py`: **134 passed** - broader unit sweep (`-k 'ggml or iq or gguf or registry'` over quantization, export and recipe tests): **192 passed**. The one failure, `test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes`, is a `torchvision` import error in my environment, unrelated to IQ. - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed** on RTX PRO 6000 Blackwell (sm_120). 7 of them are IQ2_S: CUDA-vs-PyTorch encoder parity, determinism, reconstruction at scale, zero and non-finite policy, float64 input and the fallback path. - `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k 'iq or ggml'`: **36 passed** (9 tests × 4 formats) in `nvcr.io/nvidia/nemo:26.08` - `tests/examples/hf_ptq/test_llm_ptq.py -k iq2_s`: **passed**. TinyLlama PTQ through unified HF export writes `quant_algo: IQ2_S`, `block_payload_bytes: 82`, and `down_proj` packed as `(2048, 22, 82)` uint8. - `general/ptq` now holds 30 recipes. - The shared-memory change in `b7739d5d0` leaves the packed bytes identical (same hash on a 5632×2048 weight), and packing runs at 849.1 M elem/s against 825.7 before on RTX PRO 6000. The GPU battery was rerun: 42 passed. All of the above was rerun after rebasing onto `main` at `c2aaa44f6`. That base adds a Q8_0 packer to the same GGML extension (#2515), and changes the hf_ptq example and the export code this format goes through. The packed IQ2_S bytes still hash the same. On this RTX PRO 6000 (sm_120), two of #2515's own Q8_0 tests in `tests/gpu/_extensions/test_torch_extensions.py` fail: `test_cuda_ext_q8_0_zero_and_roundf_layout` and `test_cuda_ext_q8_0_dequantizes_with_small_error`. They fail identically on a clean `main` checkout, so they are not from this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code sources or dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: #2511 (IQ2_XXS, merged) → #2525 (format registry, merged) → #2512 (IQ2_S codec, merged) → **this** → #2513 (IQ1_M). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added IQ2_S weight-only quantization for eligible linear layers, at 2.5625 bits per weight. * Added a PTQ recipe that requires no calibration data. Weights must meet the existing 256-value block-size constraint. * Added CUDA-accelerated packing for CUDA weights, with a Python fallback when the CUDA extension is unavailable. * **Documentation** * Updated the PTQ recipe catalog and IQ-format size tradeoffs. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
834c90d7a1 |
Fix NVFP4 fake quant zeroing blocks with small scales (#2549)
### What does this PR do? Type of change: Bug fix The dynamic NVFP4 Triton kernel (`fp4_fake_quant_block`, used on compute >= 8.9) and the Conv3D implicit-GEMM CUDA kernels replaced any FP8 block scale below 1e-5 with 1.0, so every block whose max |x| was below ~6e-5 was zeroed. The static Triton kernel, the CUDA extension fallback and NVFP4 export have no such floor; the floor was only guarding division by zero. The Triton kernel had its own copy of the scale/round code instead of the shared `nvfp4_scalar_quant`. - Triton: use the shared `nvfp4_scalar_quant` (zero only on a zero block scale). - Conv3D CUDA (fused kernel and standalone `fp4_fake_quant`): same rule. - Both: a zero, inf or NaN global amax uses a unit block scale, like the CUDA extension (the conv kernels returned NaN for inf/NaN before). Blocks with scale >= 1e-5 are unchanged. - New tests: power-of-two scaling of input and global amax (2^-10, 2^-20) scales the output by the same factor (Triton, standalone conv FP4, fused conv3d); invalid global amax gives unit-scale rounding. The conv test's Python reference drops the floor. - B200: 505 passed / 31 skipped (`tests/gpu/torch/quantization` NVFP4/FP4 files) and 180 passed (conv implicit GEMM + attention P-QDQ). - Negative control on `main`: the small-input tests fail on both kernels; conv also fails for inf/NaN global amax. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (only blocks with scale < 1e-5 or an inf/NaN global amax change, from zeros/NaN to correct values) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: ❌ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * FP4 quantization now preserves proportional output when inputs and valid global scales are reduced together, including at small scales. * Zero, infinite, or NaN global scales use a safe fallback, preserving inputs already representable in FP4. * Small positive block scales are no longer discarded by an absolute scale threshold; subnormal scale handling is covered across quantization paths. * **Tests** * Added coverage for scale consistency across input types and block sizes, invalid global scales, and subnormal FP8 block scales. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shiyang Chen <shiychen@nvidia.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c2aaa44f60 |
[Speculative Decoding] DFlash2 draft variant (grouped sublayer convolution + candidate selector) (#2216)
### What does this PR do? Type of change: new feature Adds **DFlash2** ([blog](https://inco.ai/blog/dflash2/)) as a draft variant of the existing DFlash mode, selected with `dflash_architecture_config.projector_type="dflash2"` alongside `domino`, `dspark` and `lilicorr`. DFlash2 keeps DFlash's one-pass parallel backbone and adds two components that recover the acceptance a purely parallel draft loses: - **Grouped dynamic depthwise convolution** around every attention and MLP sublayer, giving each block position a view of its predecessors *inside* the block. Taps do not cross the block boundary, so the draft stays one forward pass. - **Low-rank candidate selector** scoring transitions between adjacent positions' top-k candidates, so serving walks one coherent path instead of taking an independent argmax per position. Both start as exact no-ops — the convolution's `base_kernel` is an identity and `kernel_projection` is zeroed; the selector's `successor_codebook` is zeroed — so a freshly built DFlash2 draft *is* its DFlash backbone, and enabling the variant is an extension rather than a perturbation. This matches the reference implementation ([SpecForge#772](https://github.com/sgl-project/SpecForge/pull/772), merged) and the way `modeling_lilicorr` already installs this same convolution class. **This unblocks a recipe already shipped on `main`.** `modeling_lilicorr._install_sublayer_convs` imports `DFlashGroupedConv` from `modeling_dflash2`, so `modelopt_recipes/general/speculative_decoding/lilicorr_conv.yaml` raises at model build today and its CHANGELOG entry documents a feature that cannot run. Landing this makes it runnable. Module and parameter names match the SGLang/vLLM `DFlash2DraftModel` loaders. Verified against the released `z-lab/Qwen3.8-27B-DFlash2` checkpoint: **81 tensors, 21 name patterns, zero difference in either direction**. The serving side, [vllm-project/vllm#52816](https://github.com/vllm-project/vllm/pull/52816), has since merged (`b389ac294`) with no change to the checkpoint contract. ### Usage ```bash python examples/speculative_decoding/main.py \ --config modelopt_recipes/general/speculative_decoding/dflash2.yaml \ model.model_name_or_path=Qwen/Qwen3-8B \ data.data_path=<corpus>.jsonl \ training.output_dir=<out> ``` ```yaml # modelopt_recipes/general/speculative_decoding/dflash2.yaml dflash: dflash_selector_loss_alpha: 1.0 # weight of the candidate-selector CE term dflash_architecture_config: projector_type: dflash2 conv_kernel_size: 2 # taps; must not exceed the block size conv_group_size: 16 # must divide hidden_size selector_rank: 256 selector_top_k: 16 ``` ### Testing <img width="2000" height="1320" alt="image" src="https://github.com/user-attachments/assets/869004d1-92c6-41a4-a03d-fd824a06255c" /> **Unit** — 25 CPU tests in `tests/unit/torch/speculative/plugins/test_hf_dflash2.py`; the full `tests/unit/torch/speculative/` suite passes with no regressions. The ones worth keeping pin invariants that a decreasing loss does not catch: - the convolution is an exact identity on the **default** construction, and its taps stay inside the block while a position still sees its predecessors; - the block-offset contract shared by the training objective and `CandidateSelector.greedy_path` — a misaligned objective still converges; - the export fields the vLLM loader requires, including the top-level `block_size` that `DFlash2Exporter` derives the nested copy from; - which selector factors receive gradient on the first step. `successor_codebook` starts at zero, so `predecessor_codebook` and `hidden_projection` take one step to begin moving. That is a warm start, not a dead branch, and both sides are asserted. **End-to-end** — trained on Qwen3-8B against a plain DFlash control with every other argument identical (plot above). Monotonic convergence, no NaN/divergence, no DDP unused-parameter issues. Note the losses are **not comparable** across arms: DFlash2's includes the selector CE term. **Serving (vLLM)** — the exported drafter loads and drafts under the merged DFlash2 path (`RESOLVED draft architectures: ['DFlash2DraftModel']`). Two notes for anyone reproducing: vLLM sizes the convolution from `1 + num_speculative_tokens` at runtime rather than from the checkpoint, so a `block_size=16` drafter is only correct at `num_speculative_tokens=15`; and at that value the upstream path currently hits an illegal memory access in `_cache_draft_logits` ([vllm#55279](https://github.com/vllm-project/vllm/issues/55279)), independent of which checkpoint is used. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — additive. New `projector_type`, its own registry and exporter, one new config field; DFlash / Domino / DSpark / LiLiCorr numerics and `state_dict` contents are untouched. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — `modeling_dflash2.py` is adapted from [SpecForge#772](https://github.com/sgl-project/SpecForge/pull/772) and carries its MIT notice. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — under `0.48.0`. - Did you get Claude approval on this PR?: ✅ — run on 2026-08-20; all review threads addressed and resolved. ### Additional Information Rebased onto current `main`. Two commits from the original branch were dropped because [#2342](https://github.com/NVIDIA/Model-Optimizer/pull/2342) landed them first, with authorship preserved: the no-op sublayer seam in `modeling_dflash.py`, and the `rope_theta`/`rope_parameters` fix — `main`'s version of the latter is stricter, so this PR no longer touches `hf_dflash.py` at all. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added DFlash2 speculative decoding with grouped dynamic convolutions and low-rank candidate selection. * Added configurable selector-loss weighting, including an option to disable it. * Added DFlash2 model conversion and export support. * Added checkpoints compatible with SGLang and vLLM DFlash2 serving. * **Documentation** * Added training recipes and a Qwen3-8B online DFlash2 training configuration. * **Tests** * Added coverage for conversion, training, metrics, gradients, and export compatibility. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1d392999b4 |
[OMNIML-5570] 2/2 Compose GEMM and KV-cache AutoQuant workflows (#2273)
### What does this PR do? Type of change: new feature. Follow-up to merged #2272. Adds composition of existing GEMM quantization with KV-cache AutoQuantize: - fixed FP8 GEMM PTQ followed by mixed-KV AutoQuantize; - gradient-based NVFP4/FP8 GEMM AutoQuantize followed by independent mixed-KV AutoQuantize; - an optional `kv_auto_quantize` recipe stage with independent method, constraints, candidates, and checkpoint path; - ordered `hf_ptq.py` orchestration that keeps selected weight/activation QDQ active while its calibration state remains frozen during KV candidate calibration; - fail-closed validation when a preceding stage leaves actual K/V quantizers enabled; and - unified export of a uniform-weight or mixed-weight checkpoint together with the selected per-layer KV map. The KV search still uses the public `mtq.auto_quantize(..., constraints={"cost_model": "kv_cache", ...})` API from #2272. On a converted model, the API preserves existing non-KV quantizers and requires K/V to be disabled before search. Fresh-model behavior is unchanged and starts from a deny-all quantizer baseline. #### Why a follow-up field instead of a generic stage list? This PR deliberately supports the two composition forms required by `hf_ptq.py` without replacing the stable recipe schema. Existing recipes already express a fixed `quantize` baseline plus one primary `auto_quantize` search. A generic ordered `stages` list would require a broader recipe/API migration, indexed checkpoint semantics, and compatibility rules for arbitrary stage sequences. There is not yet a demonstrated third search stage that justifies that surface-area change. The two searches are not combined inside `mtq.auto_quantize`: each invocation owns one search domain, constraint model, scoring method, and resumable checkpoint. Their ordering and independent checkpoint paths are orchestration concerns, while candidate calibration, scoring, selection, and state application remain in the shared public API. A general stage pipeline can be considered separately if more than this one optional KV follow-up is needed. Both solvers and scoring protocols are unchanged. The KV checkpoint compatibility signature additionally fingerprints the preceding quantizer configuration and calibrated state. Unsupported uniform-weight plus mixed-KV exports record `kv_cache_deployment_supported: false` in both ModelOpt and converted HF metadata. ### Usage Fixed FP8 GEMM PTQ followed by KV AutoQuantize: ```bash python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path Qwen/Qwen3-8B \ --recipe general/auto_quantize/fp8_ptq_then_kv_fp8_nvfp4_cast_kl_div_at_5p4bits \ --kv_auto_quantize_checkpoint /path/to/kv_autoquant.pth \ --export_path /path/to/qwen3-8b-fp8-and-mixed-kv ``` Weight AutoQuantize followed by KV AutoQuantize: ```bash python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path Qwen/Qwen3-8B \ --recipe general/auto_quantize/nvfp4_fp8_gradient_then_kv_fp8_nvfp4_cast_kl_div_at_5p4bits \ --auto_quantize_checkpoint /path/to/weight_autoquant.pth \ --kv_auto_quantize_checkpoint /path/to/kv_autoquant.pth \ --export_path /path/to/qwen3-8b-autoquant-and-mixed-kv ``` KV checkpoint resume requires identical preceding non-K/V quantizer configuration and calibrated state. If rerunning the preceding stage changes that state, use a new KV checkpoint path to recompute sensitivities; configuration identity alone is insufficient to reuse the scores safely. ### Testing - Latest changed-area validation: 126 tests passed across `hf_ptq.py` orchestration, KV checkpoint compatibility, export metadata, and HF configuration conversion. - A broader local run had 604 passes, one skip, and six failures: two socket-binding failures under the sandbox and four local Transformers API incompatibilities. This is not a full-suite pass. - The fixed-PTQ→KV recipe executes end to end on a tiny offline Qwen fixture. - Public API coverage verifies that composed KV search preserves preceding weight quantization and rejects enabled K/V state. - Changed-file pre-commit hooks passed; the isolated recipe validator also passed after dependency bootstrap. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ (0.48.0 composition feature and KV checkpoint flag deprecation) - Did you get Claude approval on this PR?: ❌ ### Additional Information - This follow-up targets `main`, which contains merged #2272. - `--auto_quantize_checkpoint` and `--kv_auto_quantize_checkpoint` are intentionally separate because KV sensitivities depend on the preceding GEMM state. - Uniform-weight plus mixed-KV exports are for artifact inspection until the runtime's uniform-weight ModelOpt configuration consumes `kv_cache_quantized_layers`. Export emits an actionable warning and records `kv_cache_deployment_supported: false`; this marker does not itself add runtime support. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added staged post-training quantization workflows for weights and KV caches, including dedicated KV-cache checkpoints. - Added FP8/NVFP4 recipes with configurable bit constraints and scoring. - KV-cache quantization now supports pre-quantized models. - **Bug Fixes** - Mixed weight and KV-cache quantization now exports with a warning instead of failing. - Improved validation and checkpoint compatibility for staged configurations. - Added safeguards for configurations without enabled weight quantizers. - **Documentation** - Clarified staged KV-cache workflows, checkpoint options, configuration behavior, and unsupported deployment combinations. - Documented deprecated legacy quantization options and their replacement behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
de2e810219 |
[OMNIML-5899] Add Q8_0 CUDA packing kernel (#2515)
### What does this PR do? Type of change: new feature Adds the Q8_0 CUDA packing layer for the three-PR Q8_0 series: - packs 32-value blocks into the 34-byte GGML-compatible payload; - exposes `q8_0_pack` through the shared GGML extension; - validates device, dtype, row alignment, and CUDA launch bounds; - tests byte layout, accepted dtypes, non-finite handling, float64 narrowing, FP16 scale boundaries, reconstruction error, and invalid inputs. This PR contains only the kernel and extension boundary. The codec/backend and export/recipe layers remain in the later PRs. ### Usage ```python from modelopt.torch.quantization.extensions import get_cuda_ext_ggml extension = get_cuda_ext_ggml(raise_if_failed=True) packed = extension.q8_0_pack(weight) ``` ### Testing - Combined #2515 -> #2516 -> #2517 stack: 108 focused CPU codec, backend, export, and recipe tests passed. - Ruff, formatting, and whitespace checks passed for the changed Python test. - The shared-extension Q8_0 and existing IQ tests require CUDA CI; the latest run is pending on the current PR head. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: yes - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: yes; no dependency was added and no implementation code was copied - Did you write any new necessary tests?: yes - Did you update `CHANGELOG.rst`?: N/A; the user-facing entry is in #2517 - Did you get Claude approval on this PR?: pending ### Provenance The CUDA encoder was independently written for ModelOpt. It implements the packed-format contract and scalar quantization formula documented by the pinned llama.cpp definitions: - [Q8_0 packed structure](https://github.com/ggml-org/llama.cpp/blob/9b05354ec6fb58b4e665e9a39ebc40285c015638/ggml/src/ggml-common.h) - [Q8_0 scalar reference formula](https://github.com/ggml-org/llama.cpp/blob/9b05354ec6fb58b4e665e9a39ebc40285c015638/ggml/src/ggml-quants.c) No llama.cpp implementation code is incorporated into this CUDA source. Human code-owner confirmation of this provenance and attribution is requested before merge. ### Related PRs Merge order: 1. **Kernel - this PR** 2. [#2516 - Q8_0 quantization codec and backend](https://github.com/NVIDIA/Model-Optimizer/pull/2516) 3. [#2517 - Q8_0 checkpoint export and recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2517) All three PRs target `main`. --------- Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com> |
||
|
|
0058a15537 |
[2/2] Track every Megatron-Bridge script with MLflow (#2514)
### What does this PR do? Type of change: new feature **[2/2] of a split. Based on #2544 — merge that first; this PR's diff is only the Megatron-Bridge half.** #2477 added MLflow tracking to `examples/megatron_bridge/quantize.py`. It was one of five scripts in that directory that write a checkpoint; the other four recorded nothing, so the provenance chain stopped at the PTQ checkpoint and a deployed model could not be traced back to the run that produced it. All five now take the same `--mlflow` / `--mlflow_experiment` / `--mlflow_run_name` flags, and **each declares what it records as a `Tool` beside its own flags** — the shared `mlflow_utils.py` knows none of them: | Script | Records | | --- | --- | | `prune_minitron.py` | command, arguments, log, `prune_score` metric, pointer | | `quantize.py` (#2477, moved onto the shared `Tool` in #2544) | + resolved recipe, quantizer summary | | `distill.py` | + Megatron-Bridge's per-iteration metrics and resolved config | | `export_quantized_megatron_to_hf.py` | command, arguments, log, pointer | | `export_distilled_megatron_to_hf.py` | same, one pointer per exported checkpoint | Each writes `.experiment.json` into the checkpoint it produced, and each tags what it consumed, so `prune → quantize → distill → export` is walkable both from disk and by tag query on the server. **`distill.py` opens the run and Megatron-Bridge joins it.** Its `LoggerConfig` records per-iteration metrics and the full resolved config — which a wrapper around `main()` cannot see — but nothing of `distill.py`'s own arguments and no invocation. Megatron-Bridge takes `mlflow.active_run()` when one exists, applies the tags and logs into it, so `distill_run()` opens the run on the rank Megatron-Bridge looks at (the **last** one) and the two share it. Its early exit is handled explicitly: `train()` leaves through `sys.exit(0)` on `--exit_interval`, which a blanket handler would record as `FAILED`. **The library pieces that exist for that shared run land here with their first caller**, rather than in [1/2] where they would have none: `split_tracking_credentials`, so a URI handed to something which *records* it carries no credential; `log_active_run_experiment_json`, for pointing a checkpoint at a run this process did not open; and `MlflowRunLogger._reattach`, because a co-owner can end the run first — Megatron-Bridge does, as `KILLED`, when SIGTERM arrives mid-training. Two of Megatron-Bridge's defaults are deliberately not inherited: **checkpoint artifact upload stays off** unless `--mlflow_log_checkpoints` (it pushes the whole checkpoint over HTTP after every save), and **an untracked run passes no `mlflow_*` fields at all**, since they landed in Megatron-Bridge 0.6 and sending them unconditionally would break an untracked run on an older one. ### Usage ```bash # Any of the five, same flags: torchrun --nproc_per_node 8 prune_minitron.py ... --mlflow https://<server>/ torchrun --nproc_per_node 8 quantize.py ... --mlflow https://<server>/ torchrun --nproc_per_node 8 distill.py ... --mlflow https://<server>/ torchrun --nproc_per_node 8 export_quantized_megatron_to_hf.py ... --mlflow https://<server>/ # Each checkpoint names the run that wrote it: cat /output/qad/checkpoints/.experiment.json ``` Experiments default to `$USER/megatron_bridge_{prune,quantize,distill,export,distill_export}/<model basename>-<variant>`. ### Testing - Real runs on a toy Qwen3 in one MLflow experiment covering all five Megatron-Bridge scripts and `hf_ptq` — prune, quantize, QAD distillation, quantized export, BF16 distillation, distilled export, HF PTQ — each closing `FINISHED` with the invocation, its arguments as params, its log, and a matching `.experiment.json` on disk. The chain tags line up: each stage's `source_checkpoint_path` is the previous stage's `checkpoint_path`. - `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08`, the only lane that runs it: **76 passed**. Plus the three suites from #2544: **195 pass**. - `pre-commit run --files <changed>`: all hooks pass. - Each fix from the review rounds has a test that fails with the fix reverted: the resumed run, the foreign active run, the percent-decoded credential, the credential that cannot be moved, the rank-dependent `LoggerConfig`, the exit-callback guard, and the `iter_*` join. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: several rounds; re-requested on this head. ### Additional Information Split from a single ~1150-line PR at review's request; #2544 carries the library consolidation this builds on, and this branch is based on it. Earlier review threads here show as outdated after the rebases — they are all resolved and their fixes are in this branch. One known gap, stated in the README rather than implied: `distill.py --hf_export_path` writes a second HuggingFace checkpoint from rank 0, which is not the rank that owns the run, so it carries no pointer yet. For the same reason the uploaded `logs/distill.log` holds the last rank's output — `print_rank_0` keeps the script's own lines on rank 0 — which the README now says outright; carrying rank 0's log into a run owned by another rank needs cross-rank upload and is a follow-up. Two defects found on shared-run paths during review, both verified against the installed Megatron-Bridge 0.6 rather than its docs. Megatron-Bridge ends the run it shares with `distill.py` as `KILLED` from its SIGTERM handler (`train.py:1413`) and then leaves through `sys.exit()` (`train.py:805`), i.e. before `distill_run`'s `finally` — and MLflow's fluent calls resolve their target by *opening* a run when none is active, so a preempted distillation's log and metrics went to a second, empty run and its `KILLED` status was overwritten. Separately, an unreachable server disabled our logger but `logger_kwargs` still handed Megatron-Bridge the same URI, and `state.py` calls `set_experiment` unguarded from inside the training loop — so a best-effort `$MLFLOW_TRACKING_URI` aborted the training instead of degrading to untracked. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
767ef5533e |
[3/5] Add the IQ2_S codec (#2512)
### What does this PR do? Type of change: new feature (not yet user-reachable) **First of two PRs adding IQ2_S**, the widest of the GGML IQ formats at one and two bits (2.5625 bits per weight). This one lands the **PyTorch codec**: the encoder, the decoder and the 1024-entry codebook. It is deliberately **not registered**, so no quantizer dispatches to it and the `ggml` package does not export it. #2565 adds the CUDA encoder, registers the format and adds its recipe. ### What's distinctive about it **IQ2_S is the one format llama.cpp's own tooling gives no head start on**, so the search is written against the GGML layout directly. The interesting difference from IQ2_XS and IQ2_XXS is sign handling. IQ2_S stores a **full 8-bit sign mask** per group rather than a 7-bit parity-coded index. The encoder therefore takes the input signs as they are instead of flipping the weakest element to fix parity, and the search compares magnitudes directly, which is simpler than its siblings. ### Why the codec lands before the kernel The CUDA encoder's tests use this codec as their reference. They compare against the PyTorch encoder byte for byte and draw the grid and scale predictor from it. So the kernel cannot be tested before the codec exists, and it follows in #2565 together with the registration. Every registered format therefore keeps a CUDA encoder. ### Test changes that make the split possible A codec can now land before it is registered, so two test contracts in `test_iq_formats.py` are stated precisely: - The two tests that go through `TensorQuantizer` (pass-through gradient, error falls with bit width) iterate `IQ_FORMAT_REGISTRY`. Every other battery test calls the codec directly and covers IQ2_S here. - The coverage check now asserts `set(IQ_FORMAT_REGISTRY) <= set(FORMATS)` instead of equality. That is what its docstring already said: a registered format must be listed, or it escapes the contract. - `test_registry_lists_every_exported_encoder` is unchanged, and it is why this PR leaves the package exports alone: an exported encoder must be registered. The error-by-bit-width failure message also labels errors by the order they were measured in; it previously zipped them with alphabetical names. ### Testing **The decoder is validated against llama.cpp's own output, not just round-tripped:** ``` IQ2_S: 9 tensors, 2,355,200 blocks → 0 mismatched, max|diff| 0.0 ``` The new codebook matches the `ggml-common.h` table entry for entry. Blocks from `unsloth/Qwen3.8-27B-GGUF` ship as conformance vectors, so CI keeps checking bytes we did not produce. - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py`, `tests/unit/recipe/test_presets.py`: **121 passed**, 14 of them IQ2_S codec cases, including the llama.cpp conformance check - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **35 passed**, unchanged by this PR ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ The new codebook is a GGML table, carried in `codebooks.py` with the source revision recorded. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A. Nothing is user-reachable yet; #2565 carries the entry. - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: #2511 (IQ2_XXS, merged) → #2525 (format registry, merged) → **this** → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added GGML-compatible IQ2_S quantization and dequantization support, including access to its magnitude grid. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
57f929e358 |
Bound the stale-capture warning to once per capture and name its cause (#2492)
### What does this PR do?
Type of change: Bug fix
The activation-capture hooks latch `_intermediate_output` on every
forward, and only
`DistillationModel.compute_kd_loss()` ever clears it. Any training loop
that does not call
`compute_kd_loss()` therefore warns on every forward after the first,
once per hooked module:
```
plain transformers Trainer, 16 micro-batches
before: 30 warnings (Teacher 15 + Student 15)
after: 2 warnings (Teacher 1 + Student 1)
```
The count scales with forwards, not optimizer steps —
`gradient_accumulation_steps` of 1, 4 and
16 all produced 30 warnings for the same 16 forwards — so the
accumulator the report blamed is
not involved. A plain `transformers.Trainer` reproduces it because its
default `compute_loss`
reads the student's own CE from `outputs.loss` and never calls
`compute_kd_loss()`; the
warning is a symptom of that, but neither message said so. The student
message attributed the
re-forward to Activation Checkpointing only, and the teacher message
called the situation
"expected" while still raising `UserWarning`.
This PR:
- warns once per captured activation via
`_warn_once_about_stale_output`, re-armed wherever a
capture is cleared, so a consumed capture that is followed by another
unconsumed one is still
reported (Activation Checkpointing re-runs forwards and relies on that);
- states the cause and the consequence in both messages:
`compute_kd_loss()` did not run since
the previous forward, so no KD loss is applied;
- moves the three `_intermediate_output = None` reset sites onto one
`_clear_captured_output`
helper so the capture and its warning state cannot drift apart, and has
the layerwise teacher
hook share both helpers rather than keeping a second copy of the
warning.
### Usage
No API change. A loop that applies KD should call `compute_kd_loss()`
once per forward:
```python
class KDLossTrainer(Trainer):
def compute_loss(self, model, inputs, return_outputs=False, **kwargs):
outputs = model(**inputs)
loss = model.compute_kd_loss(student_loss=outputs.loss) # consumes the captures
return (loss, outputs) if return_outputs else loss
```
### Testing
`tests/unit/torch/distill/test_distill.py`:
- `test_duplicate_fwd_hook_call` now pins the bound under
`warnings.simplefilter("always")`
(3 forwards -> exactly 2 warnings) instead of relying on the
interpreter's per-message
deduplication, and asserts the message names `compute_kd_loss`;
- `test_stale_output_warning_rearms_after_consuming` covers the re-arm:
warn, consume, warn
again -> 4 warnings. This one passes before and after the change; it
guards against a future
"warn once ever" simplification silently muting the Activation
Checkpointing case.
`tests/unit/torch/distill/test_layerwise.py`:
- `test_layerwise_stale_output_warning_is_bounded` covers the layerwise
hooks (3 forwards ->
exactly 2 warnings).
The first and third tests fail on `main` (`assert 4 == 2`), verified in
a worktree of `main`
carrying these test files.
Measured with a `transformers.Trainer` over a 16-micro-batch run,
counting
`"already has an intermediate output stored"`:
| setup | before | after |
|---|---|---|
| plain `Trainer`, `gradient_accumulation_steps=1` | 30 | 2 |
| plain `Trainer`, `gradient_accumulation_steps=4` | 30 | 2 |
| plain `Trainer`, `gradient_accumulation_steps=16` | 30 | 2 |
| `Trainer` calling `compute_kd_loss()`, accum=4 | 0 | 0 |
```
$ python -m pytest tests/unit/torch/distill -q
34 passed
$ python -m pytest tests/unit/torch -q --ignore=tests/unit/torch/deploy
1 failed, 2673 passed, 18 skipped in 217.17s
```
The single failure is
`tests/unit/torch/quantization/plugins/test_huggingface.py::test_dbrx`,
which fails identically on `main` in this environment because
`transformers 5.17` is outside the
`transformers>=4.57,<5.15` range pinned in `pyproject.toml`. It is
unrelated to this change.
`pre-commit run --files <changed files>` passes every hook (ruff check,
ruff format, mypy,
bandit, insert-license, large files, line endings).
Note on severity, since it affects how the issue reads: Python already
deduplicates a warning
by (message, location), so with the default filters this shows up twice
rather than 30 times.
The repetition is user-visible under `-W always`,
`PYTHONWARNINGS=always`, pytest, or one
warning registry per DDP rank, and the count is what scales with epoch
length. The behavioural
fix that matters for all filters is the latch.
### Before your PR is "*Ready for review*"
Is this change backward compatible?: ✅
If you copied code from any other sources or added a new PIP dependency,
did you follow guidance in CONTRIBUTING.md: N/A
Did you write any new necessary tests?: ✅
Did you update Changelog?: N/A
Did you get Claude approval on this PR?: N/A
### Additional Information
Fixes #2487.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Bug Fixes**
- Improved handling of captured activations during distillation,
including clearing stale outputs and resetting warning behavior after
outputs are consumed.
- Stale-output warnings now appear once per affected activation and
clarify when knowledge-distillation loss was not applied.
- Warnings distinguish expected cases involving activation checkpointing
or teacher evaluation.
- **Tests**
- Expanded coverage for warning counts, warning reset behavior, and
stale-output handling in standard and layerwise distillation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Edwardssss <ed_129@qq.com>
|
||
|
|
4eb86524f0 |
[1/2] One MLflow tracking core behind a Tool record (#2544)
### What does this PR do? Type of change: refactor (no functional change) **[1/2] of a split. Merge this first; #2514 is [2/2] and is based on this branch.** Three example scripts had each reimplemented the same MLflow wiring: the flags, the `$USER/<tool>/<model>-<variant>` experiment convention, the params/tags/artifacts a run uploads, and the open/close dance with its status. The copies had already drifted — only `hf_ptq` wrote a provenance pointer, only `vllm_serve` republished the resolved URI — and every new tracked script meant another copy. What a script records is now one declarative `Tool` record, **declared in the script itself, beside the flags it reads**: ```python # examples/megatron_bridge/quantize.py QUANTIZE = Tool( name="megatron_bridge_quantize", tracks="Track this run on an MLflow server, uploading the command, the resolved recipe, ...", variant_help="recipe name, or --quant_cfg if no --recipe", variant=lambda args: Path(args.recipe).stem if args.recipe else (args.quant_cfg or "none"), model=lambda args: args.hf_model_name_or_path, checkpoint=lambda args: args.export_megatron_path, texts=lambda args: resolved_recipe_texts(args.recipe), outputs=lambda args: {"summary/quant_summary.txt": Path(args.export_megatron_path) / ".quant_summary.txt"}, ) ``` `tracked_run` takes that record and runs the whole thing, so a script adds tracking in three lines: `add_mlflow_args(parser, TOOL)`, `resolve_mlflow_args(args, parser, TOOL)`, and `with mlflow_run(args, TOOL):`. The shared module knows no script's flags. `examples/hf_ptq`, `examples/vllm_serve` and `examples/megatron_bridge/quantize.py` move onto it. Three helpers fall away as redundant (`track_run`, `checkpoint_run_tags`, and `hf_ptq`'s two flag pass-throughs). ### Usage No user-facing change. The flags, their spellings and the experiment naming are exactly as before; a script author now writes a `Tool` instead of four functions. ### Testing - `tests/unit/torch/utils/test_mlflow.py`, `tests/examples/hf_ptq/test_hf_ptq_args.py`, `tests/examples/vllm_serve/test_vllm_mlflow_utils.py` — **179 pass**. - `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08` (the only lane that runs it), which drives `quantize.py` for real: **34 passed**, locally and in this PR's `megatron` lane. - `pre-commit run --files <changed>`: all hooks pass. - The four suites shared four copies of a stand-in for the `mlflow` module, which had drifted — one recorded artifacts as a list, another as a dict, a third made `log_artifact` a no-op, so a test asserting on an upload asserted nothing. They now share one `tests/_test_utils/mlflow.py`, which also emulates the fluent API's habit of opening a run when none is active. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `track_run` and `checkpoint_run_tags` are removed, but neither shipped in a release (0.47.0's `__all__` is `MlflowRunLogger`, `command_text`, `current_user`, `default_experiment_name`, `validate_tracking_uri`, all unchanged here). Three deliberate behaviour changes, each in shared code and each tested: - The `source_checkpoint_path` tag resolves to an absolute path where it recorded the raw argument, which a chain of runs needs to join on the pair. `run_tags` is shared, so this applies to every script that writes the tag — `hf_ptq` **and** `megatron_bridge/quantize.py`, for a local `--hf_model_name_or_path`. A source that names no directory, such as a Hub `org/name` id, is still recorded as given. - `MlflowRunLogger.track()` — which *did* ship in 0.47.0 — records a block ending in `SystemExit(0)` as `FINISHED` where it recorded `FAILED`, since a script that ends by calling `sys.exit()` rather than returning has still finished. - `.experiment.json`'s `tracking_uri` and the `run_url` built from it drop a trailing `/` from the tracking URI, so the link is `https://host/#/...` rather than `https://host//#/...`. Only reachable by constructing `MlflowRunLogger` directly; every CLI path already stripped the slash in `resolve_tracking_uri`. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — no user-visible change; the entry is in [2/2]. - Did you get Claude approval on this PR?: several rounds; re-requested on this head. ### Additional Information Split out of #2514. This half is the enabling refactor with no behaviour change; #2514 is the feature it unlocks and is based on this branch. At ~605 changed lines of core logic it is over the ~500 guideline; the owner accepted a two-PR split rather than three, and everything #2514 alone consumes — `split_tracking_credentials`, `log_active_run_experiment_json`, `MlflowRunLogger._reattach` — lands there rather than here. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
23355eda90 |
fix(deps): declare httpx, unbreaking partial-install (torch) for every PR (#2547)
## Summary `partial-install (torch)` has been failing on **every** PR since 2026-09-24 — including PRs whose branches predate the breakage — and because it is a *collection* error rather than a test failure, it aborts the entire run: ``` ImportError while importing test module '.../tests/unit/torch/speculative/plugins/test_hf_streaming_dataset.py' E ModuleNotFoundError: No module named 'httpx' collected 2243 items / 1 error / 45 skipped !!!!!!!!!!!!!!!!!!!! Interrupted: 1 error during collection !!!!!!!!!!!!!!!!!!!! ``` `unit-pr-required-check` aggregates it, so nothing currently merges on a fresh run. ## What happened **No code changed.** `modelopt/torch/speculative/plugins/hf_streaming_dataset.py` has imported `httpx` at module scope since #1509 (2026-06-02), and `httpx` has never appeared in `pyproject.toml`. It arrived only transitively: `dev-test` → `timm` → `huggingface_hub` → `httpx`. **huggingface_hub 2.0.0**, published **2026-09-24T12:01:21Z**, replaced `httpx<1,>=0.23.0` with the separate **`httpx2<3,>=2.0.0`** distribution. Different package name, so `httpx` stopped being installed and the chain disappeared. The boundary is exact — every run *created* before that timestamp passes, every one after fails: | PR | run created | result | |---|---|---| | #2536 / #2535 | 09-23 22:02 | pass | | #2500 | 09-23 23:45 | pass — **merged 09-24 20:01 on this stale-green result** | | *hub 1.33.0 (still requires httpx)* | *09-24 09:49* | | | **hub 2.0.0 published** | **09-24 12:01** | ← | | #2539 | 09-24 16:57 | fail | | #2544 | 09-24 18:32 | fail | | #2216 | 09-25 11:58 | fail | #2500 merging afterwards is not a counterexample: GitHub does not re-run checks at merge time, so it merged on a result from ~20 hours earlier. That is also why this went unnoticed. ## The changes ### 1. Declare `httpx` in the `hf` extra `httpx` is not incidental to streaming — it is the only transport: - every fetch is HTTP: `POST /v1/completions` to the vLLM serve plus `GET /meta` and `/desc` against the connector's sidecar, all through `httpx.Client`; - there is no non-HTTP path — the base `StreamingDataset._fetch` is an abstract seam and `EagleVllmStreamingDataset._fetch` is its only implementation; - no other HTTP library appears in the module (`requests` / `urllib` / `aiohttp`: zero hits, and `requests` is not declared either); - even the retry predicate is built from it: `_TRANSIENT_FETCH_ERRORS = (httpx.HTTPError, OSError)`. It belongs in `hf` rather than in the core `dependencies`: the same module needs `transformers.trainer_pt_utils` at module scope, so one extra already gates the whole file, and a core install has no use for an HTTP client. The bound matches the 0.x API the code uses — `httpx` has no 1.0 release, and 2.x is a different distribution. This is the part that stops it recurring. `[hf]` currently gets `httpx` only because `datasets` happens to require it — the same accident with a different supplier, one release away from repeating. ### 2. Acquire `httpx` in the test through the existing skip guard The test file already intends to skip where the extra is absent — it has `pytest.importorskip("transformers")` and a comment explaining why, and `transformers` is absent in this job too. It broke only because `import httpx` sat **five lines above** that guard, where a missing module ends collection instead of skipping one file. ## Verification - With everything installed: **18 passed**, no behaviour change. - The import-order property is checked with an AST walk over the module's top-level statements: no `hf`-extra-only import precedes the first `importorskip` (which is now line 41). - A faithful local reproduction was attempted and abandoned honestly: hiding `httpx` locally also breaks `huggingface_hub` 1.28, which `modelopt.torch.opt.plugins.huggingface` imports, so the local failure is not the CI one. CI is the oracle for that half — this PR's own `partial-install (torch)` run is the check that matters. ## Scope Two files, five lines of declaration and four of test import order. Deliberately not folded into any feature PR: it blocks the whole repo, and burying a repo-wide fix inside unrelated work is how these stay invisible. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Optional Hugging Face installations now include `httpx`, supporting features that require HTTP communication without requiring it for all installations. * **Tests** * Hugging Face streaming dataset tests now skip when `httpx` is unavailable, allowing the remaining test suite to be collected and run without it. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ed7e87953c |
Fix grouped expert quantizer checkpoint replicas (#2500)
### What does this PR do?
Type of change: Bug fix
Fix distributed checkpoint saving for quantized Transformer Engine
grouped MoE experts when tensor parallelism and expert parallelism are
both greater than one.
#### Observed error
With NeMo 26.08 and TP=2, EP=4, ETP=1, quantization and calibration
complete, but the run crashes while saving the distributed checkpoint:
```text
megatron.core.dist_checkpointing.core.CheckpointingException:
Invalid sharding pattern validation.
Invalid access pattern for ShardedTensor(
key='decoder.layers.1.mlp.experts.experts.32.linear_fc1.weight_quantizer._amax',
...
)
```
#### Root cause
`_QuantMegatronTEGroupedLinear.sharded_state_dict` created each expert
quantizer's sharded tensors without passing the expert tensor-parallel
and expert data-parallel process groups to
`make_sharded_tensors_for_checkpoint`. It then manually replaced only
the expert-data-parallel component of `replica_id`.
That manual rewrite is insufficient when tensor and expert parallelism
are both enabled: replica ownership is derived using the wrong
process-group topology, so ranks can publish an inconsistent access
pattern for the same globally indexed expert quantizer key.
Megatron-Core correctly rejects that checkpoint during sharding
validation.
#### Fix
Resolve the expert model-, tensor-, and data-parallel groups from one
`_pg_collection`-aware helper, then pass the expert TP/DP groups to
`make_sharded_tensors_for_checkpoint` for per-expert quantizer buffers.
This avoids mixing process-group sources when models use a non-global
`ProcessGroupCollection` or grouped child modules convert without their
parent MLP, and removes the manual `replica_id` rewrite. Shared,
whole-linear quantizer buffers deliberately retain the default dense
TP/DP replica groups because their keys and offsets carry no expert
identity; using only expert TP/DP groups would make replica IDs collide
across EP ranks.
#### Relationship to PR #2319
[PR #2319](https://github.com/NVIDIA/Model-Optimizer/pull/2319) fixed
two earlier TEGroupedMLP checkpoint problems:
1. Per-expert quantizer state did not retain the same globally unique
expert identity as the grouped-expert weights, so it could not be
redistributed reliably when the EP layout changed.
2. After ModelOpt extra-state restoration, an expert that moved to a
different rank could be missing the `_amax` or `_global_amax`
destination buffer required by the subsequent distributed checkpoint
load.
That PR therefore fixed resharding and restore correctness across
topology changes. Its regression matrix changed one parallel dimension
at a time: EP changed while TP=1, or TP changed while EP=1.
The remaining issue was on the save path when TP and EP were
simultaneously greater than one. Although the expert keys and restore
buffers were correct after #2319, the per-expert quantizer shards were
still constructed without the expert TP/DP process groups and then had
only part of their `replica_id` rewritten manually. Under TP=2, EP=4,
ETP=1, this produced the invalid access pattern rejected by
Megatron-Core before the checkpoint could be saved.
This PR complements #2319 by fixing that process-group/replica mapping
and adding combined TP+EP coverage.
A four-rank save-and-restore regression test covers combined TP=2, EP=2,
and ETP=1.
### Usage
N/A. This fixes checkpoint behavior without changing the public API.
### Testing
- [x] Focused Ruff check and format validation
- [x] Focused mypy validation
- [x] `git diff --check origin/main..HEAD`
- [x] Four-rank grouped-expert checkpoint save/restore test with TP=2,
EP=2, ETP=1: 1 passed on the original fix
- [x] End-to-end NeMo 26.08 quantization and checkpoint save on 8 B200
GPUs with TP=2, EP=4, ETP=1 on the original fix
- [x] Review-response validation on `dbfa1cdb5`: Ruff 0.15.20
format/check, `git diff --check`, and Python compilation
- [x] Final four-rank grouped-expert checkpoint regression on
`c2be0f2c6`: TP=2, EP=2, ETP=1 with both ordinary and overridden
ModelOpt TP state; 2 passed in 459.38s on 4 B200 GPUs with NeMo 26.08
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌
### Additional Information
The end-to-end validation produced a complete eight-shard distributed
checkpoint and exited successfully.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Bug Fixes**
- Fixed checkpoint saving for quantized grouped MoE experts when tensor
and expert parallelism are enabled.
- Improved handling of shared and per-expert quantizer state across
parallel configurations to support more reliable checkpoint save and
restore.
- **Tests**
- Added NVFP4 grouped expert checkpoint round-trip coverage with tensor
and expert parallelism, including configurations with the
tensor-parallel group override enabled and disabled.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
|
||
|
|
63c4b660bd |
Add Aumann-Shapley sensitivity scoring method to auto_quantize (#2183)
[Paper](https://arxiv.org/abs/2607.12266) · [Overview](https://x.com/waterloo_intern/status/2076460984475263401) · [Implementation thread](https://x.com/the_joshua_hill/status/2076427869388255635) Depends on #2231, which provides the shared AutoQuantize backward-scoring infrastructure. Until that PR merges, the focused PR B diff is available [here](https://github.com/joshua-hill/Model-Optimizer/compare/fix/autoquant-scoring-infrastructure...feat/aumann-shapley-autoquant). ### What does this PR do? Type of change: new feature This PR adds `method="aumann_shapley"` to `mtq.auto_quantize`. It is a label-free scoring method that measures how each candidate quantization format affects the model across the path from full precision to quantized. The new method is opt-in; existing `gradient` and `kl_div` behavior is unchanged. For each calibration batch, the method: 1. Runs the baseline model and saves its next-token distribution. 2. Measures each candidate format at a configurable number of points along the quantization path. 3. Uses the KL-divergence gradients at those points to assign a damage contribution to every runtime group and candidate format. 4. Measures the most aggressive candidate configuration once and uses that value to calibrate the per-group damage model. The resulting scores use the existing AutoQuantize linear-program solver. The search can either: - choose the least damaging configuration that meets an `effective_bits` target; or - choose the smallest configuration whose predicted damage stays below `max_predicted_damage`. The selected recipe records `predicted_damage` in mean per-token KL units together with its validity and fit diagnostics. ### Public API `auto_quantize` gains an optional `method_options` dictionary. For `method="aumann_shapley"`, it accepts: | Option | Default | Meaning | |---|---:|---| | `num_path_nodes` | `2` | Number of points used to average gradients along the quantization path. | | `damage_link` | `"coverage"` | How per-group scores combine. `"coverage"` uses `damage = c * (1 - exp(-sum(b)))`; `"additive"` sums the path contributions. | | `max_predicted_damage` | `None` | Replaces the bit target with a maximum predicted mean per-token KL. | Method options are validated before the model is modified. Unknown options and incompatible targets fail early. ### Usage Select a configuration for a target effective bit width: ```python import modelopt.torch.quantization as mtq model, search_state = mtq.auto_quantize( model, constraints={"effective_bits": 4.8}, quantization_formats=["NVFP4_DEFAULT_CFG", "FP8_DEFAULT_CFG"], data_loader=calib_loader, forward_step=lambda model, batch: model(**batch), method="aumann_shapley", ) print(search_state["best"]["predicted_damage"]) print(search_state["best"]["predicted_damage_valid"]) ``` Or let the search choose the bit width for a predicted-damage target: ```python model, search_state = mtq.auto_quantize( model, constraints={}, quantization_formats=["NVFP4_DEFAULT_CFG", "FP8_DEFAULT_CFG"], data_loader=calib_loader, forward_step=forward_step, method="aumann_shapley", method_options={"max_predicted_damage": 0.05}, ) ``` ### Implementation - Reuses the candidate-replay and backward-scoring lifecycle introduced in #2231. - Reuses the existing AutoQuantize linear-program solver for both search directions. - Numerically integrates the coverage path when converting measured contributions into per-group damage costs. - Preserves deterministic runtime-group and candidate ordering. - Keeps raw measurements, solver scores, and damage-model diagnostics distinct in the search state. - Rejects incompatible checkpoint resumes while allowing the same scores to be re-solved for a new bit budget. - Retains the shared MoE score-module rules so routed experts are scored at their enclosing block. Recipe integration will follow separately. ### Testing Focused tests: ```text pytest -q \ tests/unit/torch/quantization/test_autoquant.py::test_backward_scoring_session_restores_partial_setup \ tests/unit/torch/quantization/test_autoquant_shapley.py ``` Result: **51 passed**. The tests cover: - end-to-end scoring and configuration generation; - agreement between path contributions and measured quantization damage; - exact allocation checks against exhaustive search; - effective-bits and predicted-damage search modes; - checkpoint resume and offline re-solving; - custom formats and heterogeneous candidate ladders; - distributed reductions and nested MoE score modules; - reused score modules and model-specific backward support; - non-finite measurements and invalid-fit reporting; and - input validation before model conversion. End-to-end checks with NVFP4 and FP8 candidates at a 6.0-bit target: - `Qwen/Qwen2.5-0.5B-Instruct` reaches 5.998 effective bits, with the summed path contributions reproducing 98% of the directly measured lowest-precision KL. - `Qwen/Qwen3-30B-A3B` reaches 6.000 effective bits and reproduces 99%, with all 48 MoE layers scored once at the sparse-MoE block rather than per expert. ### Production use We use this method in production for NVFP4 checkpoints of Kimi-K3, MiniMax-M3, and GLM-5.2. ### Before your PR is "Ready for review" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow the contributing guidance?: ✅ No copied code and no new dependencies. - Did you write the necessary tests?: ✅ - Did you update `CHANGELOG.rst`?: ✅ - Are the commits signed and signed off?: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added label-free Aumann–Shapley scoring for automatic quantization. - Added configurable path sampling, damage modeling, effective-bit targets, and predicted-damage bounds. - Added temporary weight-folding support with automatic state restoration. - Added method-specific search options, checkpoint resumption, and distributed scoring. - **Bug Fixes** - Improved cleanup and restoration of quantizer state, gradients, hooks, and forward behavior after scoring or failures. - Added validation and clearer handling for unsupported configurations and invalid measurements. - **Tests** - Expanded coverage for scoring, solver behavior, distributed execution, checkpointing, and custom quantization formats. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Joshua Hill <joshua.hill@baseten.co> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
400498d82d |
[2/4] Register each GGML IQ format once for dispatch and export (#2525)
### What does this PR do? Type of change: refactor (no behaviour change) Addresses review feedback on #2511. Backend dispatch and export each kept their own list of the GGML IQ formats: `_FAKE_QUANTS` in the backend, and `IQ_FORMATS`, `IQ_BLOCK_METADATA` and `IQ_PACKERS` in export. All four listed the same formats. Adding a format meant a row in each, and the lists could drift apart. That had already happened twice in #2511: `convert_hf_config.py` kept its own upper-case spelling of the family and dropped IQ2_XXS metadata, and the Megatron export tests were hard-wired to two formats. Each format module now declares **one `IQFormat` record** beside its encoder and decoder: name, block geometry, `quantize`, `dequantize`, and its encode and decode chunk defaults. **`IQ_FORMAT_REGISTRY`** lists them. - Backend dispatch looks formats up in the registry. - Both exporters take the packer and block geometry from it. - Export's `IQ_FORMATS` is derived from it instead of being written out again. - `_FAKE_QUANTS`, `IQ_BLOCK_METADATA` and `IQ_PACKERS` are removed. - The per-format fake-quant wrappers collapse into one `IQFormat.fake_quant`, which does the `num_bits` check and calls the existing cache helper. Codebooks, searches, payload layouts and CUDA encoders stay in each format's module. #### Series and merge order This is one slice of the IQ format series. It targets `main` so unit CI runs, and **its diff includes #2511's commits until #2511 merges**. 1. #2511 — IQ2_XXS format 2. **this PR** — one registration per format 3. #2512 — IQ2_S format 4. #2513 — IQ1_M format After this lands, #2512 and #2513 are restacked onto it, so each adds a format module and a single registry entry instead of rows in four tables. #### Design choices - **An explicit list, not self-registration at import.** If formats registered themselves when their module was imported, the registry's contents would depend on import order. - **Backward compatible, with one behaviour change.** `iq1_s_fake_quant` and `iq2_xs_fake_quant` are public on main, so each format keeps its `<fmt>_fake_quant` name as an alias of its record's method. The three removed tables were introduced by #2511 and never released. The behaviour change: on main, the alias looked the encoder up at call time, so patching `iq1_s.quantize_iq1_s` changed what it ran. Now the record captures the encoder and decoder when it's built, so patching those module functions reaches neither dispatch nor the alias. Substitute through `IQ_FORMAT_REGISTRY` instead. - **Registering a format declares it exportable, and that's intended.** Export's `IQ_FORMATS` is derived from the registry, so a format registered for dispatch is also claimed by both exporters and `convert_hf_config`. That can't be wrong for an IQ format: fake quant is `dequantize(quantize(w))`, so a format can't be dispatched without the packer and block geometry, and those are all export reads. A QAT-only IQ format can't exist. If one ever needs to land ahead of its export path, an `exportable` flag on the record is a one-line addition. - **The registry is the substitution seam.** Dispatch now reads the registry, so tests that swap an encoder or decoder swap the registry entry. Patching the format module's function would no longer reach dispatch. - **Test expectations stay independent of the registry.** Tests take the *list* of formats from the registry, but their expected values come from each format's own module (`quantize_<fmt>`, `<FMT>_BLOCK_BYTES`, …). A mis-wired registry entry therefore can't make both sides of an assertion agree. #### What it does not unify The CUDA side (`ggml.cpp` bindings, the `extensions.py` source list, codebook sizes in `common.cuh`) and the recipes and docs remain per format. "One registration" holds for the Python side, which is where all four tables lived. ### Usage Adding a format after this PR (for example IQ2_S in #2512) needs its module and one line in the registry: ```python # modelopt/torch/quantization/ggml/iq2_s.py IQ2_S_FORMAT = IQFormat( name="iq2_s", block_size=IQ2_S_BLOCK_SIZE, block_bytes=IQ2_S_BLOCK_BYTES, quantize=quantize_iq2_s, dequantize=dequantize_iq2_s, block_chunk_size=_DEFAULT_BLOCK_CHUNK_SIZE, decode_chunk_size=_DEFAULT_DECODE_CHUNK_SIZE, ) # modelopt/torch/quantization/ggml/registry.py IQ_FORMAT_REGISTRY = {fmt.name: fmt for fmt in (IQ1_S_FORMAT, IQ2_XXS_FORMAT, IQ2_XS_FORMAT, IQ2_S_FORMAT)} ``` Looking up a format: ```python from modelopt.torch.quantization.ggml import IQ_FORMAT_REGISTRY fmt = IQ_FORMAT_REGISTRY["iq2_xxs"] packed, shape = fmt.quantize(weight) # GGML blocks fmt.block_bytes, fmt.effective_bits # 66, 2.0625 ``` ### Testing - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py` — **94 passed** - `tests/gpu/torch/quantization/test_iq_formats_cuda.py` — **22 passed** (RTX PRO 6000) - `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k iq` — **27 passed** in `nvcr.io/nvidia/nemo:26.08`, the image CI uses for that suite - broader sweep of IQ, export and recipe unit tests — **153 passed**, none failed **New guards on the registry itself:** - every encoder the package exports is registered - each record points at its own format's codec, geometry and chunk defaults - the public `<fmt>_fake_quant` alias is the registered record's method - export's `IQ_FORMATS` and `QUANTIZATION_IQ*` constants match the registry - a format's `fake_quant` refuses a quantizer configured for another format. Dispatch picks the record by `num_bits`, so it never reaches this guard; the test covers direct callers of a record or alias. The three per-format guards it replaced were untested on main. - every registered format is listed in the shared test batteries Checked by mutation: leaving IQ2_XXS out of the registry, or registering it with the IQ2_XS encoder, each fails the guard written for that case. **Coverage gap closed along the way:** `test_ggml_backend.py` was hard-wired to IQ1_S and IQ2_XS, so IQ2_XXS had no backend, cache or packed-once coverage. Those tests now run over the registry. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — public per-format fake-quant names are kept as aliases; the removed tables were never released. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A — internal refactor with no user-visible change - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information Review feedback on #2511 that this addresses: *"`_FAKE_QUANTS`, `IQ_FORMATS`, `IQ_BLOCK_METADATA`, and `IQ_PACKERS` independently enumerate the same formats. A common pack/dequantize/fake_quant interface would let backend dispatch and export consume one registration."* 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * IQ quantization formats are available through a shared format registry, keeping format details and quantization behavior consistent across supported workflows. * IQ-format model exports use registered format information for quantization metadata and weight packing. * **Tests** * Expanded checks to cover registered IQ formats and verify consistent format support across quantization and export. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
1c4cde7788 |
Fix: Release per-layer expert weights in layerwise export under offload (#2466)
### What does this PR do?
Type of change: Bug fix
Layerwise export leaks one layer's worth of quantized tensors per layer
on offloaded models, and dies of OOM partway through a large MoE. Two
things accumulate, both because the offload window cannot reclaim what
the export pass adds inside it.
**1. Per-expert holder modules.** `_export_fused_experts` splits a fused
MoE experts module into per-expert holders and attaches them to the live
model:
```python
proj = nn.Module(); proj.weight = wrapper.weight # packed U8
expert.add_module(proj_name, proj)
module.add_module(str(idx), expert)
```
They are plain `nn.Module`s built *inside* the weight-access window, so
they carry no accelerate `_hf_hook`.
`weight_access_and_writeback_context` closes by iterating the modules it
collected at entry and calling `hook.post_forward()` through each one's
own offload hook — the holders satisfy neither condition, so nothing
returns them to meta.
**2. Scale buffers.** `_export_quantized_weight` registers
`weight_scale` / `weight_scale_2` / `input_scale` on the layer's
pre-existing, hooked sub-modules, and `AlignDevicesHook.post_forward`
runs with `offload_buffers=False`:
```
after post_forward: {'weight': 'meta', 'weight_scale': 'cpu'}
```
The packed weight goes back to meta; the scales do not. For
`_QuantMoELinear` models this half is the whole leak on its own —
`_reconstruct_fused_moe_linear` restacks every expert's scales into one
`register_buffer` on the hooked wrapper.
A whole-model export never notices: one pass, write the state dict,
exit. Layerwise runs the same pass once per decoder layer, so every
finished layer stays resident.
### Measurements
Through unmodified `examples/hf_ptq/hf_ptq.py` on a Qwen3.5-MoE-shaped
model (10 layers, 64 experts), printing `torch.cuda.memory_allocated()`
after each exported layer:
| placement | per layer | over 9 layers |
| --- | --- | --- |
| offload | +0.052 GiB | 0.579 → 1.047 GiB |
| resident | -0.135 GiB | falls, as designed |
+0.052 GiB is exactly one layer's quantized experts: packed U8 100.7M/2
= 0.047 GiB plus FP8 block scales 100.7M/16 = 0.006 GiB.
At scale it is fatal rather than wasteful. Qwen/Qwen3.8-2.4T-A95B (92
layers, 512 experts) leaks 12.9 GB packed + 1.6 GB scales per layer, so
92 layers want 1.33 TB that no budget on a 283 GB card or 952 GB host
absorbs. The run died of CUDA OOM at layer 16/92 with
`--max_gpu_memory_gb 240`, and at `--max_gpu_memory_gb 30` leaked the
same 15 GB/layer onto the host instead.
### The fix
`_release_exported_tensors` (`model_utils.py`) is a context manager that
snapshots each sub-module's child-module and buffer names on entry, and
on exit drops whatever appeared. Persisting happens *inside* the block,
so "release only once it is on disk" is structural rather than a
comment.
Both packing sites use it: `LayerwiseExporter.export_layer` and the
offload decoder loop in `_export_transformers_checkpoint_streaming`.
The streaming writer had solved the same leak inline with a heuristic —
null every CUDA buffer, and every CUDA parameter on a hook-less module —
and that block is deleted in favour of the shared helper. Keying on
*what the pass added* rather than on device and hook presence drops two
assumptions that only held for a terminal, offloaded export: it no
longer nulls buffers the layer already had, nor parameters of
sub-modules accelerate simply did not hook. That is also what makes it
safe for the layerwise path, where resident models are supported and the
model outlives the export.
Deliberately out of scope: the FSDP2 per-unit loop in
`collect_export_tensors` keeps its per-unit holders. That predates this
PR, this PR does not touch that loop, and closing it needs its own
change and its own test.
### Usage
No API change. Existing layerwise export under offload simply stops
growing:
```bash
python examples/hf_ptq/hf_ptq.py \
--pyt_ckpt_path Qwen/Qwen3.8-2.4T-A95B \
--qformat nvfp4 \
--export_path /path/to/export \
--max_gpu_memory_gb 240
```
### Testing
- `tests/unit/torch/export` and `tests/unit/torch/quantization` — 1246
passed
- `tests/gpu/torch/export/test_layerwise_export.py` +
`test_offload_export.py` — 35 passed
- `cuda_alloc` over 10 offloaded layers: 0.526 → 0.518 GiB (-0.008), was
+0.468
- Exported checkpoint byte-identical to the unfixed run: 7841 tensors, 0
mismatches, max abs diff 0.0; `hf_quant_config.json` / `config.json` /
index identical
- Full Qwen3.8-2.4T-A95B PTQ then completed all 92 layers: peak GPU 117
GB of 283, peak host RSS 71 GB of 952, flat across 50 consecutive layers
at 23-26 s/layer
Coverage added: the existing `test_export_creates_per_expert_submodules`
now runs the export inside the context manager and asserts the holders
are gone on exit, and a new test in `test_offload_export.py` pins the
`offload_buffers=False` behaviour the buffer half exists for —
export-registered scales dropped, pre-existing buffers untouched.
`tests/gpu/torch/export/test_fsdp2_export.py` reports 34 failures in my
environment. They are **pre-existing and unrelated**: the same 34 fail
identically on this branch and on the merge-base (`2b1f33d0ef`), with
byte-identical failure sets and runtimes within 3 s. All 34 are `Failed:
Timeout (>120.0s)` from hung NCCL collectives, with zero assertion
failures.
The GPU export suites and the whole-model measurements above were run at
`2999d7cfb7`. The two commits since — swapping the holder marker for a
child-name diff, and moving the helper to `model_utils` — are covered by
the unit suites; a GPU re-run before merge is worthwhile.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — one new offload test for
the buffer half, and the existing fused-experts export test now covers
holder release
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — `layerwise.export_dir` is new in the unreleased 0.48.0, so this
bug was introduced and fixed within the same cycle
- Did you get Claude approval on this PR?: ✅
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
a21411adde |
Add the IQ2_XXS weight-only quantization format (#2511)
### What does this PR do? Type of change: new feature llama.cpp defines five GGML IQ formats at one and two bits; we ship two. This adds **IQ2_XXS** at 2.0625 bits per weight, between IQ1_S and IQ2_XS, and is the **first of three**. On a real mixed-precision checkpoint (`unsloth/Qwen3.8-27B-GGUF`, `Qwen3.8-27B-UD-IQ1_S.gguf`) IQ2_XXS alone covers **59 tensors and 2.84 B parameters — 10.6% of the file**, which a reader limited to IQ1_S/IQ2_XS cannot consume. Across all three PRs the missing formats account for 17.3%. | format | bpw | bytes/256 | codebook | | |---|---|---|---|---| | `iq1_s` | 1.5625 | 50 | `iq1s_grid` (2048) | existing | | **`iq2_xxs`** | **2.0625** | **66** | **`iq2xxs_grid` (256)** | **this PR** | | `iq2_xs` | 2.3125 | 74 | `iq2xs_grid` (512) | existing | The encoder follows the existing single-pass grid search at a fixed anchored super-block scale, and the CUDA kernel the existing per-block structure. IQ2_XXS reuses IQ2_XS's even-parity sign rule but packs a 4-bit sub-block scale into the same 32-bit word as four 7-bit sign indices, and its 256-entry grid needs no high index bits. ### Groundwork the next two reuse Two things land here because IQ2_XXS is the first format to need them: - **Export registry.** The IQ family was spelled as a two-element tuple at **nine** sites across `quant_utils.py`, `unified_export_hf.py` and `unified_export_megatron.py`. Those become an `IQ_FORMATS` frozenset plus per-format packer and block-geometry tables, so a format is a row rather than a sweep through the exporters. - **Shared test contract.** The per-format test files had drifted apart — each of `iq1_s` and `iq2_xs` tested things the other did not. They become one parametrized module per layer (unit and CUDA), so every format is held to the same contract and a new one inherits it. ### Usage ```bash python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq2_xxs ``` ### Testing **The decoder is validated against llama.cpp's own output, not just round-tripped.** Every IQ2_XXS tensor in the checkpoint above, compared against `dequantize_row_iq2_xxs` from `ggml-quants.c`: ``` IQ2_XXS: 59 tensors, 11,100,160 blocks → 0 mismatched, max|diff| 0.0 ``` The new codebook matches the `ggml-common.h` table entry for entry, as does the `ksigns_iq2xs` sign table. Blocks lifted from that checkpoint ship as conformance vectors so CI keeps checking bytes we did not produce; mutation testing confirms they catch a wrong sign-field width. The CUDA encoder is byte-identical to the PyTorch reference on a fixed input and runs at **1047.9 M elem/s against the torch search's 10.7** on a 5632×2048 weight. - `tests/unit/torch/quantization/ -k 'ggml or iq1 or iq2 or iq_'` — 99 passed - `tests/gpu/torch/quantization/test_iq_formats_cuda.py` — 21 passed (7 checks × 3 formats) - `tests/unit/recipe/test_presets.py` — passing; `general/ptq` now holds 29 recipes, `ptq.md` updated - reconstruction error decreases monotonically with bit width, pinned by a test Pre-existing failures in `tests/unit/torch/export/` and `test_autoquant.py` are `transformers`/`torchvision` import problems in my environment — identical counts with and without this change. ### A finding about already-merged code Checking the new kernel against its PyTorch reference at 4096 blocks showed that **CUDA and torch encoders disagree on roughly 1 block in 6000 — including the already-merged `iq2_xs`**, at 0.0163% against IQ2_XXS's 0.0000%. Root cause: both compute `xnorm − 2·scale·dot + scale²·qnorm`, but CUDA fuses it with `fmaf` while torch uses separate ops; where two local scales fall within a float32 ULP the roundings pick different sides. Adjudicated against float64, neither path is better (5 to 6). Worst-case cost is **1.48e-08** relative reconstruction error, and run-to-run determinism on a given device holds. This is pre-existing, not introduced here — `test_iq2_xs_cuda.py` asserts exact byte parity but on a 16-block weight where ties essentially never arise. I have **not** changed that test; rewording a guarantee on merged code belongs in its own change. The new shared GPU tests assert exact parity on a small fixed input and compare reconstruction error at scale. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — the new codebook is a GGML table, carried in `codebooks.py` beside the existing ones so the MIT-licensed surface stays in that one file, with the source revision recorded. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information First of three; **IQ2_S** and **IQ1_M** follow and build on this branch. Replaces #2505, which carried all three at once. Follows #2446 / #2447 / #2448 / #2449, which landed IQ1_S and IQ2_XS. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added IQ2_XXS weight-only quantization, including CUDA acceleration and support for Hugging Face and Megatron exports. - Added the `general/ptq/iq2_xxs` recipe. It requires no calibration data and supports eligible layers with a weight dimension divisible by 256. - Updated the PTQ recipe catalog to list IQ1_S, IQ2_XXS, and IQ2_XS at approximately 1.56, 2.06, and 2.31 bits per weight, respectively. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f2f0d6958e |
Add MLflow tracking flags to megatron_bridge quantize.py (#2477)
### What does this PR do?
Type of change: new feature
`examples/megatron_bridge/quantize.py` gains the MLflow tracking flags
`examples/hf_ptq/hf_ptq.py` already has: `--mlflow <tracking-uri>`
(MLflow's own `$MLFLOW_TRACKING_URI` is honoured too),
`--mlflow_experiment` and `--mlflow_run_name`. Only the master rank
opens a run, so a `torchrun` launch produces one run carrying the
invocation, every command-line argument as a searchable param, the
resolved `--recipe` (with `$import`s expanded), that rank's log and the
quantizer summary. Once `bridge.save_megatron_model` returns,
`.experiment.json` is written into `--export_megatron_path`, so a
Megatron checkpoint found on disk names the run that produced it; a run
that fails is still recorded as `FAILED` with its traceback.
Rather than copy the wiring a third time, the part `hf_ptq` and
`vllm_serve` had each duplicated moves into
`modelopt.torch.utils.mlflow`:
- `add_mlflow_args(parser, tool, tracks=, variant_help=)` — the three
flags, registered under both the `--mlflow_x` and `--mlflow-x` spellings
(vLLM's `FlexibleArgumentParser` only matches the dashed one).
- `resolve_tracking_uri(uri, parser)` → `(uri, required)` — the flag
overrides the environment and is fatal when the URI is unusable; a URI
inferred from `$MLFLOW_TRACKING_URI` warns and continues untracked,
since that variable is commonly exported for unrelated tooling.
- `resolve_mlflow_args(args, parser, tool, model, variant)` — the same,
settled onto `args`, plus the default experiment name.
- `EXPERIMENT_JSON`, `MlflowRunLogger.log_experiment_json()` and
`drop_experiment_json()` — the checkpoint→run provenance pointer,
previously private to `hf_ptq`.
Both existing callers now delegate to those, keeping their own help
wording and variant naming, so the three scripts share one convention
instead of three copies (`example_utils.py` and `vllm_mlflow_utils.py`
each lose ~60 lines). Their flags and defaults are unchanged; the only
user-visible difference is that `hf_ptq`'s ignored-URI warning gains the
`$` the vLLM one already had (`Ignoring $MLFLOW_TRACKING_URI, continuing
untracked`), so one shared message serves both.
One behaviour change reaches `hf_ptq` through the shared helper, and it
is a fix: when tracking was inferred from `$MLFLOW_TRACKING_URI` and the
run never opened (unreachable server, or `mlflow` not installed), it
used to leave the previous run's `.experiment.json` beside a freshly
exported checkpoint. `log_experiment_json` now drops the pointer when it
has no run to record, so after a completed export the file is this run's
or absent.
The new example-side code lives in
`examples/megatron_bridge/mlflow_utils.py`, which deliberately imports
no Megatron, so the whole flag-to-artifact path is testable without the
Megatron container (the same split
`examples/vllm_serve/vllm_mlflow_utils.py` uses).
### Usage
```bash
torchrun --nproc_per_node 2 quantize.py \
--hf_model_name_or_path Qwen/Qwen3-8B \
--recipe general/ptq/nvfp4_default-kv_fp8 \
--tp_size 2 \
--export_megatron_path /tmp/Qwen3-8B-NVFP4-megatron \
--mlflow https://<your-mlflow-server>/
# The checkpoint then names the run that produced it:
cat /tmp/Qwen3-8B-NVFP4-megatron/.experiment.json
```
The experiment defaults to `$USER/megatron_bridge_quantize/<model
basename>-<recipe name, or --quant_cfg>`.
### Testing
- `tests/examples/megatron_bridge/test_mlflow_utils.py` — 20 new tests
covering the flags (both spellings, env-vs-flag precedence, the
fatal/best-effort split), the params/tags/artifacts a run records, rank
gating, and the `.experiment.json` lifecycle. The last one guards the
seam with `quantize.py` as text, since that script needs Megatron to
import.
- `tests/unit/torch/utils/test_mlflow.py` — 13 new tests for the
extracted library API; suite at **75 passed**.
- Full `tests/examples/megatron_bridge` suite in
`nvcr.io/nvidia/nemo:26.08` on an RTX 6000 Ada: **37 passed (26m)**,
including the three `test_quantize_export` cases that drive the real
`quantize.py`, plus QAD, distill and prune.
- Regression proof for the refactor:
`tests/examples/hf_ptq/test_hf_ptq_args.py` **47 passed** and
`tests/examples/vllm_serve/test_vllm_mlflow_utils.py` **32 passed**,
unchanged apart from one renamed constant reference.
- Both new guards were shown to fire: mutating the `checkpoint_exported`
gate and removing `with mlflow_run(args):` each failed exactly one test.
- End-to-end tracked run in `nvcr.io/nvidia/nemo:26.08` (tiny Qwen3-MoE,
`general/ptq/fp8_default-kv_fp8`, 1 GPU) against an internal MLflow
server: run `47d4ccd7cd9e48269e7248868347ccd0` under experiment
`$USER/megatron_bridge_quantize/mbridge-ptq-validation` closed
`FINISHED` carrying `command.txt`, `version.txt`, `experiment.json`,
`recipe/resolved_recipe.yaml`, `logs/quantize.log` and
`summary/quant_summary.txt`; all 19 CLI arguments plus `world_size`
logged as params with no `mlflow_*` leakage, the
`model`/`checkpoint_path`/`source_checkpoint_path` tags set, and
`.experiment.json` written into the Megatron checkpoint beside
`iter_0000000/`.
- `pre-commit run --files <changed>`: all hooks pass (ruff, mypy,
bandit, markdownlint).
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under *Megatron Framework (M-LM / M-Bridge)*.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
`mlflow` stays an optional dependency, imported only once tracking is
enabled, so an untracked run behaves exactly as before.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added optional MLflow tracking for Megatron-Bridge quantization runs.
- Configure tracking with `--mlflow` or `MLFLOW_TRACKING_URI`, with
customizable experiment and run names.
- Records searchable parameters, resolved recipes, quantization
summaries, logs, and checkpoint provenance.
- Captures successful and failed runs and cleans up stale checkpoint
metadata when appropriate.
- **Documentation**
- Added setup instructions and usage examples covering artifacts,
naming, checkpoint metadata, validation, and authentication.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
87f7d1432f |
fix(speculative): hold the DFlash draft's fp32 master weights in the optimizer (#2483)
### What does this PR do? Type of change: Bug fix **Follow-up to #2342**, which split this out on review (commit `c67784d9`), and a rethink of how the flag is implemented. `dflash_fp32_master_weights` exists because the DFlash draft is cast to the frozen bf16 target's dtype, so AdamW allocates its moments in bf16 — and bf16 is too coarse to hold them. At `beta2=0.999` a single step changes `v` by at most **0.100%**, while the smallest change bf16 can represent near `v` is **0.164% mean / 0.388% max** (measured): every decrease rounds away, `v` only grows, and the effective step size decays on its own from step 1. #2342 fixed that by **promoting the draft model to fp32**. Everything else followed from giving the model a dtype the rest of it does not have — a bf16 autocast at every entry point, two transformers loader hints so `from_pretrained(dtype="auto")` would not round the draft away, a post-condition check because those hints fail silently, and a doubled DDP gradient all-reduce. **This PR puts the fp32 in the optimizer instead**, where Megatron-LM, DeepSpeed and apex put it. `MasterWeightAdamW` holds an fp32 master copy of each non-fp32 parameter plus fp32 moments in `self.state[p]`, steps on the master, and copies back at the parameter's dtype. The model is never anything but the base dtype, so every one of those follow-on pieces is deleted, gradients stay bf16, and the exported drafter is unchanged. What the placement costs is that wiring the optimizer becomes the training loop's job: `EagleTrainerWithAccLog.create_optimizer` builds it, and `VerifyMasterWeightsCallback` raises at the end of step 1 if the moments are not fp32. **The default flips to `True`** — the flag now changes optimizer memory and optimizer arithmetic and nothing else. Flipping it on the model-promoted implementation turns **25 of 259** unit tests red; flipping it here is **259 passed**. Set it to `False` to reclaim the memory, about 12 bytes per draft parameter instead of 4. <details> <summary>Three drive-by fixes, independent of the above</summary> - `_place_draft` is folded back into `modify()` — it fused the draft's dtype, its device and an eager rotary buffer behind one meta guard. - The module docstring's claim that `DFlashModule` has an `_apply` meta-buffer fix is removed (`grep "def _apply"` matches nothing, and never did). - #2342's field description no longer lists `evaluation` as a broken path — `forward` short-circuits to the base model when `not self.training`, so the draft never runs there. </details> ### Usage No API change. `dflash_fp32_master_weights` now means the *optimizer* holds fp32 master weights rather than the draft model being fp32. ### Testing **1 · The refactor is arithmetically a no-op.** Both implementations run AdamW on an fp32 tensor, so given the same starting values and the same gradients the trajectories are identical — 1000 steps, `weight_decay=0.01`: ``` old fp32 parameter vs new fp32 master : bitwise equal = True (max |diff| 0.0e+00) exp_avg / exp_avg_sq : bitwise equal = True optimizer state dtypes : ['torch.float32'] model parameter dtype : torch.bfloat16 ``` Initial values have to be matched at bf16 first, or the bf16 arm's one-time rounding of the draw shows up as a 2e-4 "difference" that is not arithmetic. With that controlled, the two implementations differ only in their *inputs*: gradient precision (fp32 vs bf16 — torch 2.10 requires `grad.dtype == param.dtype`) and that one-time rounding. **2 · End to end on GPU: the effect survives the refactor.** Qwen3-1.7B base, real corpus, one GPU per arm, three arms — pure bf16 (flag off), the #2342 implementation, and this one — on two algorithms trained independently, sharing seed, data order and initialisation within an algorithm. <img width="2925" height="960" alt="image" src="https://github.com/user-attachments/assets/40d37059-ab8a-4924-b049-85f76b70b156" /> The two fp32 arms sit on top of each other for the whole run while bf16 stays above both, and the old-vs-new gap is 10–23× smaller than the fp32-vs-bf16 effect it has to be compared against. **Acceptance length says the same thing, and settles what the loss could not.** All six drafters at the end of those curves were exported and served under vLLM against the same base, and measured on MT-Bench (80 prompts, 8 categories, greedy, one request at a time, `num_speculative_tokens` = trained `block_size` − 1, every knob but the drafter held fixed): | | pure bf16 | fp32 in model (#2342) | fp32 in optimizer (this PR) | new − old | fp32 − bf16 | |---|---|---|---|---|---| | `dflash` | 1.3068 | 1.3708 | **1.3666** | −0.0042 `t=−0.90` | +0.0619 `t=+13.8` | | `lilicorr` | 1.2536 | 1.2814 | **1.2882** | +0.0068 `t=+1.42` | +0.0312 `t=+9.2` | Paired by prompt, n=80. On both algorithms the new-vs-old 95% CI straddles zero (`dflash` [−0.0134, +0.0051], `lilicorr` [−0.0028, +0.0164]) while fp32-vs-bf16 does not come close to it, and the sign of new-vs-old **flips between the two algorithms** — what a rounding difference looks like, not a bias. This is also the comparison the training loss could not give: all three arms are **exported and served in bf16**, so the old implementation's fp32 draft weights are rounded at export exactly as they would be for deployment, and the "its loss was computed on a more precise forward" caveat below does not apply. `lilicorr` needs [vllm-project/vllm#57934](https://github.com/vllm-project/vllm/pull/57934), applied as an overlay so that both algorithms are measured on one engine build. The right panel is the mechanism, and the one signal that depends on neither the seed nor the choice of loss statistic: Adam's updates to the draft's RMSNorm gains are smaller than the bf16 ULP at 1.0 (0.0078), so in the bf16 arm every one of them rounds away and the gains never move — not one of `dflash`'s 14 in 30000 steps, and two of `lilicorr`'s 20 by 3e-06. Both fp32 arms move all of them, by the same amount. Two results behind the figure rather than in it. **fp32-vs-bf16 grows with the horizon** while old-vs-new does not — on `dflash` −0.129 at 1500 steps → −0.262 at 15000 → −0.341 at 30000, and on `lilicorr` −0.191 → −0.220 → −0.285, against an old-vs-new difference that stays near 0.02 at every horizon and changes sign between them (−0.026 → +0.028 on `lilicorr`). That is what a compounding bias and a rounding difference respectively should look like, and it is the reason the longer runs were worth doing. And **across seeds**, the paired old-vs-new difference at 1500 steps is +0.0003 (n=6) on `dflash` and +0.0643 (n=10) on `lilicorr`, both with a 95% CI straddling zero. <details> <summary>Limits of the above, stated rather than smoothed over</summary> At 5 seeds the `lilicorr` paired difference read +0.1610 ± 0.0557 (t=+2.89, 4/5 seeds in the same direction) — nominally significant, suggesting the new implementation was genuinely worse there. Four further `lilicorr` seeds were run against that pre-declared question; two came back strongly negative and the estimate settled at +0.0643 (95% CI [−0.086, +0.214]). The earlier reading was small-sample noise. At 1500 steps on `lilicorr` that CI is *not* narrower than the fp32-vs-bf16 effect it is being compared against, so the 1500-step sweep alone cannot certify equivalence there — `lilicorr` is still at loss 9.3 and deep in its early transient, and it is the long runs that resolve it. On `dflash` the 1500-step CI (±0.031) is already 4× tighter than the effect (−0.129). One asymmetry the loss comparison cannot separate: the old implementation held the draft weights in fp32 *at forward time*, so its training loss was computed on a more precise forward, while both implementations export bf16. Any residual advantage it appears to have is therefore an upper bound. </details> **3 · Unit tests.** `tests/unit/torch/speculative/` — **259 passed** on **transformers 5.0.0** and **5.3.0**, both ends of the supported `>=5.0,<5.13` (CPU, torch 2.10). `TestDFlashFp32MasterWeights` is rewritten for the new mechanism; the two that would have caught the traps in this design are `test_resume_does_not_round_the_master_back_down` (`Optimizer.load_state_dict` casts float state to its parameter's dtype, so a naive subclass rounds the master and both moments to bf16 on *every* resume, silently, with the loss still falling) and `test_the_callback_refuses_a_loop_that_forgot_the_optimizer`. The rest cover the draft's dtype with the flag either way, that no forward path needs an autocast any more, that plain AdamW really does leave the moments in bf16, and that an fp32 model allocates no redundant master. A sharded FSDP2 `DTensor` keeps an fp32 master and fp32 moments through a step. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ for artifacts, with one intentional default change. The draft's stored dtype goes back to matching the base, as it was before #2342; existing checkpoints load unchanged and the exported drafter is unaffected. The flag now defaults to **`True`** — the measurements above are the reason, and the cost is fp32 master + fp32 moments for the draft only. A training loop that builds its own optimizer instead of using the shipped `create_optimizer` gets plain AdamW and none of this; `VerifyMasterWeightsCallback` makes that fail loudly at step 1 rather than skip the feature quietly. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new dependencies. - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ — draft; will run `/claude review` before marking ready. ### Additional Information **On the 7–14% acceptance-length gain quoted in #2342:** that was measured on the fp32-model arithmetic and is not re-derived here. What is measured above is the like-for-like comparison this PR has to answer — same corpus, same horizon, same serving path, one implementation swapped. **History:** commits 1–3 restore the autocast design as it was split out; commits 4–6 replace it. Happy to squash before review. <details> <summary>Alternatives measured and rejected, so they do not get re-proposed</summary> - **Swapping `p.data` to the master and calling `super().step()`** (reuses all of AdamW, ~20 lines instead of ~50): bit-identical on ordinary parameters over 25 steps, but silently wrong under FSDP2 — assigning `.data` on a `DTensor` parameter updates the wrapper's reported dtype while the local shard keeps the model's, so `p.dtype` reads fp32, `p.data.dtype` reads bf16, and `zeros_like(p)` allocates the moments in bf16 anyway. CPU tests pass either way. - **Narrowing the autocast from `__call__` to `forward`** (while it still existed): turns 10 Domino/DSpark tests red — the variants apply their heads in their own `forward` overrides, outside `DFlashModule.forward`. - **Building the rotary buffer on meta and letting the loader materialise it**: makes RoPE correctness depend on transformers selecting a branch by class-name substring (`"RotaryEmbedding" in module.__class__.__name__`), and the `if not hasattr` guard is then permanently satisfied, so a later `to_empty()` leaves garbage forever — measured `4.56e-41`, i.e. cos=1 / sin=0, no positional encoding at all. - **Building it eagerly in `DFlashModule.__init__`**: lands before the dtype cast, so `Module.to` rounds the RoPE frequencies to bf16 on the default path — measured `0.8659643530845642` → `0.8671875`, loss `3.47230935097` → `3.47114777565`. </details> 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - DFlash now uses FP32 optimizer master weights and Adam moments by default while keeping draft parameters in the base model’s dtype. - Master-weight training preserves optimizer precision when restoring checkpoints. - The feature can be disabled to reduce optimizer memory usage. - Draft models consistently follow the base model’s dtype and device. - **Bug Fixes** - DFlash workflows now support operation without autocast. - Added validation for compatible AdamW-family optimizers and master-weight precision, including resumed training runs. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1b4e7dfb14 |
[OMNIML-5899] Add IQ post-training quantization recipes (#2449)
## Summary - add numerics configs for IQ1_S and IQ2_XS - add model presets and general PTQ recipes - document the supported weight-shape requirement - reuse the packed IQ weight across forwards instead of re-encoding it every time - add end-to-end `hf_ptq` coverage for both formats - add the release-note entry ## PR split This work is split into four focused PRs. Each PR targets `main` and owns a disjoint file set: 1. **Kernel** — [#2448: Add CUDA kernels for IQ packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448) 2. **Quantization** — [#2446: Add IQ quantization codecs and backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446) 3. **Export** — [#2447: Export IQ checkpoints from HF and Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447) 4. **Recipes** — [#2449: Add IQ post-training quantization recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449) The required merge order is #2448, #2446, #2447, then #2449. ## Scope Sixteen files. The PR started as eight recipe config, documentation and changelog files; the end-to-end test added for them surfaced two performance bugs in the already-merged codec, and fixing those pulled in the codec files and their tests. Beyond the original recipe set it now touches four files owned by #2446 — `ggml/common.py`, `ggml/iq1_s.py`, `ggml/iq2_xs.py` and `ggml/backend.py` — plus three test files. It still contains no kernel or export files. Keeping those fixes here rather than moving them to #2446 is deliberate and confirmed with the stack owner: #2446 is already merged, and both bugs are only observable through the end-to-end test this PR adds, so splitting them would separate each fix from the test that demonstrates it. ## Packed-weight cache fix Adding the end-to-end test made the cost visible: a TinyLlama IQ1_S `hf_ptq` run spent **499 of its 537 seconds inside the IQ1_S encoder**, packing the same 154 weights 15400 times — 100 times each. The repacking is not calibration. These recipes set `algorithm: null` and `hf_ptq` logs `Dynamic quantization. Calibration skipped.`. The 100 passes are the sample `generate()` calls `hf_ptq` makes before and after quantization: one decode step re-runs weight fake-quant on every linear, and the packed payload was thrown away each time. `_PackedWeightCache` was already there to prevent exactly this, and it never hit. It keyed on the identity of the tensor the backend was handed, but `TensorQuantizer` passes a fresh *view* of the weight on every forward, so the identity check never matched twice. The fix anchors the entry to `inputs._base` — the parameter the view is taken from — held as a weakref. The parameter is stable across forwards, so the cache hits; the reference stays weak, so the payload is released with the weight and offloaded/meta-device flows are unaffected. (A strong reference does make the cache hit, but pins full-precision storage for the life of the quantizer, which is the opposite of what those flows need.) Measured on TinyLlama IQ1_S `hf_ptq`, 2×H100, same command before and after: | | packer calls | packing time | wall clock | |---|---|---|---| | before | 15400 | 499.0 s | 8m57s | | after | 154 (one per weight) | 6.2 s | 3m21s | `test_ggml_weight_is_packed_once_across_forwards` pins this: it counts encoder calls across five forwards under `torch.inference_mode()` (what `generate()` runs under) and asserts exactly one. ## Decode chunk fix With packing cached, the end-to-end cost moved entirely into the decode, and IQ2_XS was still 4x slower than IQ1_S (1081s vs 260s per case). Instrumenting both showed packing was no longer the cost at all — IQ2_XS packs *faster*: | | pack calls | packing time | wall clock | |---|---|---|---| | IQ1_S | 154 | 6.2 s | 3m21s | | IQ2_XS | 154 | 1.6 s | 17m59s | The cause was one constant serving two loops with opposite characteristics. `_DEFAULT_BLOCK_CHUNK_SIZE` bounds the torch encode fallback, which holds the large codebook-search temporaries and runs once per weight; IQ2_XS sets it to 256 rather than IQ1_S's 1024 because its search sweeps sixteen local scales per grid tile. But the *decode* shared it — and the decode has tiny temporaries, runs on every forward, and is never cached, so a small chunk only multiplies kernel launches. Decoding a 2048x5632 weight: | chunk | IQ2_XS decode | transient peak | |---|---|---| | 256 (was) | 91.4 ms | +24 MiB | | 1024 | 22.9 ms | +31 MiB | | 4096 (now) | 5.8 ms | +56 MiB | The decode now takes its own `_DEFAULT_DECODE_CHUNK_SIZE`, threaded through `fake_quantize_with_cache`. The encode bounds are untouched, so the memory ceiling stays where it was aimed. The seven-iteration sign-parity loop in `dequantize_iq2_xs` is also folded into three XOR steps, off the same per-forward path. Per case in `tests/examples/hf_ptq` on 2xH100: | | before | after | |---|---|---| | IQ1_S | 259.65 s | 94.64 s | | IQ2_XS | 1081.05 s | 102.77 s | Both now fit the 300s `tests/examples` default, so the cases carry no explicit timeout. ## Integration contracts - `block_sizes: {-1: 256}` records the native packed-block contract; it does not drive the GGML fake-quant scale search. The export path reads `TensorQuantizer.block_sizes[-1]` through `get_weight_block_size`, validates it against the format block size, and records `group_size: 256` in checkpoint metadata. - The model presets intentionally expose `--qformat iq1_s` and `--qformat iq2_xs` in `hf_ptq`. ## Testing - `tests/unit/torch/quantization/ -k 'ggml or iq1 or iq2'` — 52 passed, 1 skipped - `tests/unit/recipe/test_presets.py` — both shipped IQ recipes load with the expected backend and no unused search option - `tests/examples/hf_ptq/test_llm_ptq.py -k 'iq1_s or iq2_xs'` — 2 passed on 2×H100 (TinyLlama, both formats end to end through export), 3m17s for the pair - pre-commit hooks pass on all changed files - larger-model sanity check outside CI: Qwen3.8-27B (2256 quantizers) quantizes and exports with `general/ptq/iq1_s` on 2xH100 in 66m46s, export itself 192s --------- Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5bb7343592 |
[https://nvbugspro.nvidia.com/bug/6778095] Fix fused P-QDQ to respect disabled quantization during calibration (#2434)
### What does this PR do?
Type of change: Bug fix
During max calibration, `enable_stats_collection()` calls
`disable_quant()`, which sets `_if_quant=False`. The fused causal P-QDQ
attention paths bypass `TensorQuantizer.forward()` and previously
selected the Triton/Kitchen path from the configured enabled state
alone, so P quant-dequant could still execute while quantization was
inactive.
This change:
- enters the Triton P-QDQ path only when `p_bmm_quantizer._if_quant` is
true;
- bypasses all fused P-QDQ paths when the P quantizer is disabled or
quantization is inactive;
- adds focused parameterized regression tests covering Triton and
Kitchen dispatch across enabled, `disable()`, and `disable_quant()`
states.
### Usage
N/A. This restores the existing `disable_quant()` contract and does not
introduce a new API.
### Testing
- `pytest
tests/unit/torch/quantization/plugins/test_attention_quant.py`: 14
passed
- Targeted pre-commit checks passed
- Kitchen coverage verifies the full `{disable(), disable_quant()} ×
{lazy, already initialized}` dispatch matrix remains bypassed while
inactive and resumes after re-enabling
- Reproduced with the exact same ModelOpt 0.47.0rc1 wheel on both sides
on B300: [module regression build
#121](http://dlswqa-nas.nvidia.com:18880/view/yiguo/job/modelopt-quant-module/121/)
- Controlled B300 isolation passed when only the P-BMM quantizers were
hard-disabled during calibration, and also passed when the existing
single Triton attention configuration was forced
- Fixed-code B300 validation passed in three independent runs: [Jenkins
build
123](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/123/),
[Jenkins build
124](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/124/),
and [Jenkins build
125](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/125/).
Jenkins build 123 compared all 178 module outputs byte-identically and
found 0/377 `amax` and 0/377 `scale` changes.
The confirmed impact is incorrect calibration behavior plus unstable
quantizer state and module outputs; downstream benchmark accuracy impact
has not been established.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A
### Additional Information
The failure was isolated to fused P-QDQ runtime-state dispatch.
Quantizer topology and configuration were identical in the failing
same-wheel comparison.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Bug Fixes**
- Corrected attention dispatch so disabled or inactive quantization uses
the original attention implementation.
- Preserved optimized quantized attention when quantization is enabled.
- Ensured attention masks remain unchanged when using the original
attention implementation.
- Improved fallback behavior for fused attention paths, including
correct initialization and reuse when quantization is re-enabled.
- **Tests**
- Added coverage for enabled and disabled quantization states,
attention-mask handling, fallback selection, and result consistency.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: yingguo-trt <244492186+yingguo-trt@users.noreply.github.com>
|
||
|
|
d0142c9dca |
Reuse a whole recipe via $import, deprecate recipe_type, and start the published-checkpoint backfill with two aliases (#2376)
### What does this PR do?
**Type of change:** New feature (recipe loading) + one bug fix
Two things, the second built on the first:
1. **A recipe can now reuse another recipe whole.** A top-level
`$import` brings in the imported recipe's entire body; keys given
alongside it override the imported ones. `metadata.recipe_type` becomes
optional and is deprecated along the way.
2. **The deprecated `recipe_type` is swept out of every shipped recipe,
and the checkpoint backfill starts with two published checkpoints
recorded as aliases** that reuse a portable recipe wholesale — the first
users of the alias mechanism — plus a fix to two existing Nemotron NVFP4
recipes.
#### Declaring what kind of recipe a file is
`load_recipe` read `metadata.recipe_type` out of the raw YAML *before*
resolving imports, because it needs the schema class to hand to
`load_config`. That made the field impossible to inherit, so a recipe
reusing another had to restate a line it could only have copied.
It is now optional, and the loader takes the first of these that
answers:
1. a `# modelopt-schema:` comment naming the recipe's schema class,
2. `metadata.recipe_type` — **deprecated**; still read and still
honoured, so a recipe outside this repo keeps working unchanged,
3. the recipe it delegates to via a top-level `$import`.
Whatever a recipe *does* state must be true, in both directions. A
schema comment contradicting a `recipe_type` is rejected, and so is a
recipe importing a different kind of recipe — that used to surface as
whatever pydantic made of, say, an `eagle` section spliced into a PTQ
schema. The concrete recipe classes carry a `RECIPE_TYPE` ClassVar as
the single source of truth.
Only a recipe that another file **imports** needs the schema comment —
that is what `$import` resolution requires to validate the payload. The
sweep here drops `metadata.recipe_type` from all 78 shipped recipes that
carried it and gives the imported ones a `# modelopt-schema:` comment
instead, so nothing in-tree depends on the deprecated field.
A directory recipe's `metadata.yml` resolves its kind the same way —
schema comment first, `recipe_type` as the fallback — it just has no
`$import` to delegate through, since a directory recipe has no body of
its own to hand off. (Follow-up commit, after this PR's initial review:
it originally still required `recipe_type` unconditionally, the one
place the deprecation didn't reach.)
#### Checkpoint aliases
Two checkpoints NVIDIA has published in quantized form use a scheme a
portable recipe already produces, with no checkpoint-specific deviation,
so each is recorded as a thin **alias** (top-level `$import`, overriding
only `metadata`) at its own model-hub path -- the *source* checkpoint's
path, not the published quantized one's:
-
**`models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast`**
delegates to `general/ptq/nvfp4_experts_only_mse-kv_fp8_cast` —
expert-only NVFP4 (MSE static weights, dynamic inputs) with an FP8 KV
cache in cast mode — published as `nvidia/Kimi-K2.6-NVFP4`.
-
**`models/Qwen/Qwen3.5-397B-A17B/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8`**
delegates to the `qwen3_5_moe` architecture recipe
`model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4
(MSE static weights) on the routed experts, ModelOpt-default FP8
elsewhere, FP8 KV cache — published as
`nvidia/Qwen3.5-397B-A17B-NVFP4-V2`.
(Follow-up commit, after this PR's initial review: the Qwen entry
originally lived at `models/nvidia/Qwen3.5-397B-A17B/` -- nvidia is the
*published* checkpoint's org, not Qwen3.5-397B-A17B's own. Moved to
match the source model's actual hub path, same as the Kimi-K2.6 entry
above.)
Editing the base recipe changes every alias that points at it; nothing
is duplicated.
#### One fix
- **The Nemotron-3 Super and Ultra NVFP4 recipes** quantized the MTP
block on the **Megatron-Core** path, where it is a live `model.mtp`
submodule their broad `*mixer.*` patterns matched into, contrary to
their own descriptions. They now disable `mtp.*` explicitly. Hugging
Face runs were unaffected — `NemotronHPreTrainedModel` sets
`_keys_to_ignore_on_load_unexpected = [r"mtp.*"]` and builds no MTP
module.
### Usage
A checkpoint alias resolves through `--recipe` to the recipe it
delegates to:
```bash
python examples/hf_ptq/hf_ptq.py \
--pyt_ckpt_path <checkpoint> \
--recipe models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast \
--export_path <output>
```
A recipe that reuses another whole — the shape the aliases use:
```yaml
imports:
base: general/ptq/nvfp4_experts_only_mse-kv_fp8_cast
$import: base
metadata:
description: What this checkpoint uses the base recipe for.
```
### Testing
- **`tests/unit/recipe/test_loader.py`** — 28 new cases covering
whole-recipe reuse with no `metadata` at all; kind resolution from each
of the three sources, from a delegation chain and from a `$import` list;
a delegation cycle failing with `ValueError` rather than recursing;
`peek_declared_schema` including a comment placed below the first YAML
line; `recipe_type` being optional, filled per class, and rejected when
it contradicts; a directory recipe resolving its kind from a schema
comment the same way, rejecting a comment/`recipe_type` disagreement,
and still requiring one or the other; and delegating across kinds being
an error.
- **`tests/unit/recipe/test_recipe_docs.py`** — the
model-specific-recipe check now also covers the two new alias folders,
which must be listed in `ptq.md` like every other
`models/<org>/<model_id>` entry.
- **Recipe validation** (`tools/precommit/check_modelopt_recipes.py`)
and **`pre-commit`** pass on the changed files. The full
`tests/unit/recipe/` suite is left to CI — a broken `transformer_engine`
in the local dev venv keeps the `mtq.quantize`-based cases from running
there.
Not covered: **numerics**. Nothing here asserts accuracy, or that
running one of these recipes reproduces a released checkpoint's weights.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `metadata.recipe_type` is
still read and honoured for recipes outside this repo, the schema
comments are inert for direct loads, and the loader change only relaxes
a check.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — two feature entries, one deprecation, and one bug fix under 0.48.0.
- Did you get Claude approval on this PR?: ❌ — not yet run.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Recipes can delegate configurations, support checkpoint aliases, and
apply local metadata overrides.
* Recipe types can be inferred from schema declarations or delegated
recipes, with stronger consistency validation.
* Added unquantized KV-cache options, layerwise export, broader operator
calibration, and new PTQ examples.
* Added checkpoint-specific recipes and MLflow experiment references.
* **Bug Fixes**
* Improved ONNX calibration, FSDP2 export, and fused-MoE quantization
handling.
* Nemotron-3 recipes keep MTP blocks in BF16.
* **Documentation**
* Expanded guidance for aliases, delegation, schema declarations, and
recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
f1abc75626 |
Carry unplaced checkpoint weights using the loader's accounting, replacing MTP name-matching (#2427)
### What does this PR do?
Type of change: bug fix + new tests
Replaces `hf_ptq`'s name-based MTP detection with the Transformers
loader's own accounting of what it could not place.
#### The problem
`load_mtp_weights` found MTP weights by name — `"mtp" in key` plus
config-derived layer indices — backed by a support matrix of three
storage conventions (GLM-5.1 inlined, GLM-4.7 standalone file,
Qwen3-Next tail shard). Every architecture that spells it differently is
a silent miss, and a miss means a checkpoint exported without a
component its own config still advertises. The failure is quiet on both
sides: Transformers drops unexpected keys, and vLLM's weight loading is
pull-based, so a missing MTP produces no warning at all.
#### The fix
`from_pretrained(..., output_loading_info=True)` reports
`unexpected_keys` — *"keys that are found in the checkpoints, but not
expected in the model's architecture"* — which is exactly the carry-over
set, derived structurally rather than by naming, and already accounting
for on-the-fly key conversion that a set re-derived afterwards would
have to replay. Record those keys at load; carry them at export. Weights
the loader *did* place go through the normal export path unchanged.
#### Two mechanisms, disjoint by construction
Measured against a real `from_pretrained`, an off-index file reports
**nothing**: the loader opens only shards named in
`model.safetensors.index.json`, so it never saw those tensors to call
them unexpected. Those are sidecars, not weights — untouched by
quantization and absent from the export — so they are **copied
verbatim**, which costs no host memory, preserves the bytes and file
layout, and leaves the filename a consumer looks for where it was.
| on-disk layout | mechanism |
|---|---|
| inlined layer past `num_hidden_layers` (GLM-5.1, DeepSeek-V3) |
`extra_state_dict` |
| indexed `mtp.*` tail shard (Qwen3-Next) | `extra_state_dict` |
| standalone off-index file (GLM-4.7) | **copied whole** |
A tensor is never both copied and carried; that would export it twice,
and a test asserts it.
#### Removed as redundant
`load_mtp_weights`, `mtp_layer_prefixes_from_checkpoint`,
`get_inlined_mtp_prefixes`, `_load_tensors_matching`,
`_apply_to_model_state_dict`, `_keys_to_prefixes`, `_add_mtp_exclusions`
and its three call sites, the pre-quantization `enable: False` entries
`hf_ptq` appended to the recipe's `quant_cfg`, and the dead
`_mtp_layer_prefixes` fallback in `_get_num_nextn_predict_layers`.
#### Two deliberate behavioural changes
**MTP now follows the recipe** instead of being force-excluded by the
script — which is what `examples/megatron_bridge` already does; it has
no MTP-specific code at all. Recipes importing
`configs/ptq/units/default_disabled_quantizers` still disable `mtp.*`,
so their behaviour is unchanged; a recipe omitting that unit will now
quantize an MTP the model actually built.
**`quantization_config.ignore` can no longer claim a layer is
unquantized that the export in fact quantized.** That contradiction came
from `_add_mtp_exclusions` firing off a model attribute with no
cross-check against quantizer state.
### Usage
No API change for callers of `export_hf_checkpoint`. Within
`examples/hf_ptq`, model loading now goes through a wrapper that records
the loader's accounting:
```python
model, loading_info = auto_class.from_pretrained(ckpt_path, output_loading_info=True, **kwargs)
record_unplaced_source_keys(model, ckpt_path, loading_info.get("unexpected_keys"))
```
### Testing
`tests/examples/hf_ptq/test_carry_over_layouts.py` — 8 tests driving a
**real** `from_pretrained` against a tiny model, covering each of the
three conventions above plus an auxiliary (non-MTP) tower, two layouts
at once, a checkpoint with nothing stray, and that indexed shards are
never copied. CPU-only: the mechanism is bookkeeping during load, so a
GPU adds nothing; the export side already has GPU coverage in
`tests/gpu/torch/export/test_export_carry_over.py`.
The six `load_mtp_weights` tests are replaced by three on the recording
path, and the `get_model` test doubles now model `output_loading_info`
the way Transformers does.
All passing: 8 layout tests, 82 in the surrounding `examples/hf_ptq`
suite. `ruff` findings at parity with `main` on every changed file.
Files named like a main weight shard are excluded from the off-index set
whatever the index says — a fixture with an empty `weight_map` would
otherwise have made the source weights look like sidecars and copied
them into an export beside the quantized ones.
### The algorithm: which weights get carried, and how
Two disjoint sets of source weights reach the export without passing
through quantization. They are distinguished by **what the loader did
with the file**, and that difference decides both how each is found and
how each is moved.
**Set 1 — unplaced weights.** The loader opened the file and read the
tensor, but the model had no parameter for it, so Transformers reports
it in `unexpected_keys`. An MTP head the recipe did not quantize is the
common case. Moved as **tensors**: located in whichever shard holds
them, read, and merged into the exporter's `extra_state_dict`.
**Set 2 — off-index sidecars.** The index never names the file, so the
loader never opened it and never had the chance to call anything
unexpected. GLM-4.7 keeps its MTP head in a standalone `mtp.safetensors`
exactly this way. Moved as **files**: copied byte for byte, so no host
memory is spent re-serialising tensors the export does not otherwise
touch.
#### The index is not an inventory of the checkpoint
This is the part that is easy to get wrong, and it cost a silent
data-loss bug during review.
`model.safetensors.index.json` selects which **files** the loader opens
— not which **tensors** it sees. Within a file it opens, Transformers
enumerates every tensor present and reports the unexpected ones.
Verified by experiment against transformers 5.3.0:
| case | reported in `unexpected_keys`? |
|---|---|
| key absent from the index, in a shard the index names for *other*
tensors | **yes** |
| key in a file the index never names (`mtp.safetensors`) | **no** — the
file is never opened |
So a tensor missing from `weight_map` but sitting inside a main shard is
**set 1, not set 2**. An MTP head stored that way is reported, recorded
— and was then silently dropped, because resolution went through
`weight_map`, which by construction has no entry for it. The
`--vllm_fakequant_export` guard shared that lookup, so the check written
to refuse exports that drop weights stayed silent in exactly the case it
existed for.
`locate_source_keys` now resolves through the index first (free for
everything it lists) and header-scans the shards only for the leftovers
— names, never tensor data — warning when a key is in no file at all.
The carry and the guard share it, so they cannot disagree again.
#### Flow
1. **At load.** `record_unplaced_source_keys` stores Transformers' own
`unexpected_keys` on the model (`_modelopt_unplaced_source_keys`) plus
the resolved local checkpoint path. The question asked is "does the
model have a parameter for this key", never "is this an MTP head" — so
the mechanism is architecture-agnostic.
2. **At export, before dispatch.** `read_unplaced_weights` resolves each
recorded key to its shard and reads the tensors, merging them into
`extra_state_dict`. An explicitly passed `extra_state_dict` wins on a
name clash: a caller naming a tensor is more specific than our
inference.
3. **Rank behaviour.** Only the rank that writes `extra_state_dict`
reads the bytes — the FSDP2 writer emits it from rank 0 alone, so a full
read on every rank would be host memory spent and discarded (a
DeepSeek-V3-class MTP head is 10 GB+ in bf16). The **key list** is still
resolved on every rank, because `get_quant_config` runs per rank and the
configs must agree.
4. **Recording what was written.** `export_hf_checkpoint` records
`_modelopt_carried_over_names` — the union of carried tensors and the
off-index sidecars' tensor names — before `get_quant_config` runs,
because that is the first point that knows what was *written* rather
than what was merely unplaced.
5. **Exclusions.** Both sets must reach `quantization_config.ignore`, or
a deployment framework reads the top-level `quant_algo` and tries to
load an original-precision weight as a quantized one (the NVBug 5718750
class). `seed_carried_over_exclusions` is the single path for this,
called by `get_quant_config` after its per-layer pass and again by the
layerwise exporter from `finalize()` — which snapshots its config during
`bind()`, while calibration is still running, so it cannot see the
carried set any earlier.
#### What is deliberately excluded
- **Files that re-ship indexed weights.** Mistral's
`consolidated.safetensors` is a second full copy of the model, and
PEFT's `adapter_model.safetensors` is an adapter. Both are off-index,
and copying either would put unquantized weights beside the quantized
ones — vLLM's mistral load-format looks for `consolidated.safetensors`
by name, so it is not inert. Caught by name for the known conventions
and by tensor-name overlap for the rest.
- **Files named like a main weight shard**, whatever the index says, so
a broken or partial index cannot make the real weights look like
sidecars.
- **Symlinks are *not* excluded.** A Hugging Face snapshot stores every
file as a symlink into a sibling `blobs/`, so refusing links would drop
the sidecar of every hub-downloaded checkpoint.
`resolve_checkpoint_file` checks where the link *lands* — regular file,
inside the checkpoint dir or its blob root — rather than whether it is a
link.
- **Buffers Transformers recomputes.** `*.inv_freq` is skipped by the
fake-quant guard even when a shard provides it: older
Llama/Mistral-lineage conversions do list it in the index, and refusing
an export over it would reject checkpoints that export correctly today.
### How a carried, never-quantized MTP head reaches
`quantization_config.ignore`
Raised in review: `_add_mtp_exclusions` is gone, and an unplaced weight
has no module, so
`get_quant_config` walks right past it. That was a real gap, not just a
documentation one —
fixed here.
The export writes a carried MTP head in its original precision. If it is
absent from
`exclude_modules`, a deployment framework reads the top-level
`quant_algo` and tries to load
`eh_proj` as an FP8/NVFP4 weight — the same class of failure as NVBug
5718750, where a
`transformers>=5.0` MoE router was written in BF16 but never excluded.
The two cases now differ only in *why* the module is invisible to the
quantizer walk:
| | why invisible | handled by |
|---|---|---|
| MoE router (tf≥5.0) | module exists, never gets a quantizer |
`_get_unquantized_moe_router_names` |
| carried weight | no module at all in the live model |
`_get_carried_over_module_names` |
`_get_carried_over_module_names` reads the keys the loader recorded as
unplaced
(`_modelopt_unplaced_source_keys`), strips the trailing parameter name —
a state-dict key is
`<module path>.<parameter>` — and dedupes. Those names are seeded into
`layer_config_dict` as
`QUANTIZATION_NONE`, exactly as the router pass does, so they flow
through
`process_layer_quant_config` into `exclude_modules`, which
`convert_hf_config` emits as
`quantization_config.ignore`.
An MTP head the recipe *does* quantize is unaffected: it has a module,
is loaded normally, is
never in the unplaced set, and is reported as quantized. The
backward-breaking note above still
holds — MTP follows the recipe instead of being force-excluded — but a
head that ends up carried
rather than quantized is no longer silently missing from `ignore`.
Covered by `test_carried_over_weights_are_excluded_from_quantization`
and
`test_carried_over_module_names_strip_parameter_and_dedup`.
`--vllm_fakequant_export` does not carry unplaced weights; it now raises
rather than writing a
checkpoint quietly missing them.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ — MTP layers now follow the
recipe rather than being force-excluded, and
`quantization_config.ignore` no longer lists layers the export may have
quantized (carried, never-quantized weights are still listed -- see
below). Shipped recipes are unaffected; see the Changelog entry.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 Backward Breaking Changes.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
Draft: the behavioural change to MTP quantization is the part most worth
a second opinion — it aligns `hf_ptq` with `megatron_bridge`, which
special-cases nothing.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Quantization supports enabled operators outside transformer layers,
including language-model heads.
* Exports preserve checkpoint weights not loaded into the quantized
model.
* Additional safetensors sidecar files are copied unchanged into
exported checkpoints.
* Unquantized auxiliary components, such as vision layers, remain
available in exported models.
* **Behavior Changes**
* Local recipe files take precedence over built-in recipes.
* Legacy architecture-specific recipe paths remain supported with
warnings.
* MTP-specific export exclusions are no longer applied.
* **Bug Fixes**
* Preserved weights are no longer incorrectly reported as unquantized.
* Incomplete exports are rejected when source weights cannot be placed
or preserved.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
fc4c40fcbe |
Fix HF export crash when a dynamic-block quantizer has zero amax (#2438)
### What does this PR do? Type of change: Bug fix `TensorQuantizer.export_amax()` early-returns `self.amax` unsanitized for dynamic-block quantizers, while the static path immediately below it has always substituted `maxbound` for zero/NaN entries. The `nvfp4` numerics unit sets `type: dynamic`, so a recipe that applies it to an *activation* quantizer — e.g. `general/ptq/nvfp4_mlp_only-kv_fp8_cast`, which targets `*mlp*input_quantizer` — feeds a raw `0.0` into `NVFP4QTensor.get_activation_scaling_factor`, whose assert aborts the entire export: ``` AssertionError: Failed to export module 'model.language_model.layers.37.mlp.gate_proj' (type=QuantLinear): activation scaling factor 0.0 not positive. ``` Calibration leaves `amax` at 0 whenever a layer — or an unrouted MoE expert — saw only zeros, so one dead layer costs the whole run at the final export step. This factors the substitution into `_sanitize_export_amax()` and calls it from both branches. Two details beyond de-duplication: - **Branch-free, so it survives a meta `amax`.** `torch.where` + `nan_to_num` both have meta kernels; `bool()` on a meta tensor raises. The layerwise and streaming export flows carry meta `amax` — `validate_attr` short-circuits on `is_meta` for exactly that reason — so only the warning is gated on a materialized tensor. - **No longer mutates calibrated state.** The old in-place `amax[amax == 0] = ...` wrote through a view of `self._amax`; `torch.where` returns a fresh tensor, so that hazard disappears. - **Warns, with a count.** The fix turns a loud failure into a silent one, and a zero amax means calibration never activated that layer — worth surfacing rather than papering over. The message reports how many entries were substituted, since per-location dedup otherwise collapses many dead experts into one uninformative message. A healthy model emits none. Scope: only the activation path is data-dependent and reachable this way. Weight-side `_amax` uses are left alone, since a weight amax of 0 would require an all-zero weight matrix. **Knowingly left as follow-up:** `export/quant_utils.py::get_scaling_factor` discards the sanitized `amax` when `num_bits == (2, 1)` and recomputes via `get_weights_scaling_factor_2_from_quantizer`, which reads `weight_quantizer._amax` raw — so a dynamic-NVFP4 *input* quantizer on a module whose *weight* quantizer is a different format (or disabled) can still trip `assert torch.all(scaling_factor > 0)`. Format dispatch is weight-driven, so the reported recipe does not reach that branch; fixing it properly changes a signature shared with the weight-side callers and is out of scope here. Not a regression. The dynamic early return, the `type: dynamic` numerics unit, and the recipe that combines them all ship in released 0.46.0 / 0.46.1. ### Usage No new or changed API. Exports that previously aborted now complete and warn: ```python # Recipe applies dynamic NVFP4 to *mlp*input_quantizer; layer 37 never activated during calibration. mtq.quantize(model, quant_cfg, forward_loop) export_hf_checkpoint(model, export_dir=out) # before: AssertionError; now: exports + UserWarning ``` ### Testing - New `test_amax_export_unusable_amax`, parametrized over zero and NaN, covering the dynamic-NVFP4 and static per-tensor configs; asserts the exported scale is positive and that export leaves the calibrated `amax` untouched. Plus `test_amax_export_meta_amax`, pinning that a meta `amax` survives export rather than raising. Both run on CPU and CUDA via the shared tester. - `tests/unit/torch/quantization/test_tensor_quantizer_cpu.py` — 40 passed. `tests/gpu/torch/quantization/test_tensor_quantizer_cuda.py` — 40 passed (GB300). - End-to-end repro on GB300, small Llama with one MLP fed all-zero activations under `general/ptq/nvfp4_mlp_only-kv_fp8_cast`: dead layer `export_amax()` `0.0` → `6.0`, live layer unchanged at `3.921875`, and `export_hf_checkpoint` goes from the `AssertionError` above to writing `model.safetensors`. - Full `examples/hf_ptq/hf_ptq.py` with the reported recipe and flags on a healthy model (Qwen3-0.6B): exits 0 and writes the checkpoint, confirming the normal path is unaffected. - `pre-commit run` clean on all changed files. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ✅ — `/claude review` run; its one IMPORTANT finding (meta-tensor regression) and both SUGGESTIONs addressed or answered in 251f2e3d ### Additional Information Fixes NVBug 6768300, reported against 0.47.0rc1 on GB200. The reporter also notes it passed on 0.47.0rc0; that is not explained by code — `git diff 0.47.0rc0..0.47.0rc1` touches `export/quant_utils.py` only in `get_kv_cache_scaling_factor` (new `clamp_fp8_scales` argument whose default preserves the old behaviour) and the INT4-AWQ packing path, neither of which is on the dense-HF NVFP4 activation-scale path. Whether `amax` lands on exactly 0 is calibration/model-state dependent, which is what makes it look version-flaky. Worth flagging separately: in the reported log the **pre-PTQ** sample output is already gibberish, so that BF16 checkpoint looks broken independently of quantization. This change stops the crash, but such a run will now export a valid-but-garbage checkpoint — the new warning is the signal to investigate. Suggest the `cherry-pick-0.47.0` label so this lands in the ongoing release. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed Hugging Face checkpoint export when dynamic-block quantizers have zero or invalid calibration scales. * Exports now use a positive fallback scale and issue a warning instead of failing when applicable. * Export operations no longer modify the original calibrated quantizer state. * Meta-device exports remain non-erroring and preserve device placement. * **Tests** * Added coverage for zero- and invalid-scale exports across dynamic and static quantization modes. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Yue <yueshen@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b311c054de |
Fix hybrid stack spec serialization in Megatron-Bridge checkpoints (#2452)
### What does this PR do?
Type of change: Bug fix
Hybrid (e.g. Nemotron-H) checkpoints saved by the
`examples/megatron_bridge` scripts could not be reloaded or exported to
HuggingFace:
TypeError: MLPSubmodules.__init__() missing 2 required positional
arguments: 'linear_fc1' and 'linear_fc2'
**Root cause.** Megatron-LM's YAML writer represents a
`functools.partial` via `_partial_representer`, which passes each
keyword value through `represent_data`. A dataclass instance has no
representer, so it falls through to `_safe_object_representer`, which
emits only `{_target_, _call_}` and drops every field. The default
hybrid stack spec builds its dense-MLP and MoE layers as exactly such
partials, and `set_moe_expert_layout()` stored the *built* `ModuleSpec`
on the provider — which is serialized into every checkpoint's
`run_config.yaml`. So `MLPSubmodules` / `MoESubmodules` were written
with no fields at all. Both export paths hit it: `convert.sh` via
`from_auto_config`, and `export_distilled_megatron_to_hf.py` via
`export_ckpt → load_megatron_model`. Every hybrid provider is affected,
dense or MoE.
**Fix.** `set_moe_expert_layout()` stores a named, zero-argument
*factory* instead. The provider already calls a callable spec at build
time (`_resolve_hybrid_stack_spec`), so model construction is unchanged
— only the serialized form differs:
```yaml
hybrid_stack_spec:
_call_: false
_target_: megatron.bridge.models.hybrid.hybrid_provider.transformer_engine_hybrid_stack_spec
```
The grouped-GEMM factory is Megatron-Bridge's own
`transformer_engine_hybrid_stack_spec`, so stock tooling
(`scripts/conversion/convert.sh`) resolves it without importing
ModelOpt.
**Known limitation — the two layouts are not symmetric.** The
SequentialMLP layout has no bridge-side equivalent (the upstream TE
hybrid spec hardcodes `TEGroupedMLP`), so it serializes a ModelOpt
target, which `instantiate` only accepts in a process that has imported
`mbridge.py` and thereby run `register_allowed_target_prefix`. A
SequentialMLP hybrid checkpoint therefore converts through the ModelOpt
entrypoints but not through stock `convert.sh`, where it fails on the
disallowed prefix instead of on `MLPSubmodules` — no regression, but
that path stays broken for this one layout. The reach is narrow:
`use_moe_grouped_gemm()` returns True for any architecture with a
grouped-expert export rule, NemotronH included, so SequentialMLP
requires an explicit `--no_moe_grouped_gemm`. Closing it properly needs
an upstream `moe_grouped_gemm`-aware factory in Megatron-Bridge.
Both spec builders also move out of `nas/plugins/megatron.py`, which
never used them, into a new `utils/plugins/megatron_layer_specs.py`
beside the other Megatron-Core-only helpers. Not into `mbridge.py`: that
module needs `megatron.bridge`, while `get_te_hybrid_stack_spec` is
reached by 16 test files through
`tests/_test_utils/torch/megatron/models.py`, which is bridge-free.
The underlying defect is upstream in
`megatron/training/config/yaml_utils.py`; this only stops ModelOpt from
stepping on it, so it is worth a separate Megatron-LM issue.
### Usage
No API change — hybrid checkpoints saved after this fix convert with the
existing commands:
```bash
torchrun --nproc_per_node 1 examples/megatron_bridge/export_distilled_megatron_to_hf.py \
--student_hf_path <student_hf_model_or_path> \
--megatron_path <distill_out>/checkpoints \
--hf_export_path <hf_out> \
--export_iterations all
```
### Testing
Verified in `nemo:26.08` (megatron-core 0.19.0) against a 30B-A3B
Nemotron-3.5-Lightning pruned+distilled run:
- **Round trip, both MoE layouts.** Ran `set_moe_expert_layout` on a
real `HybridModelProvider`, dumped it through `dump_dataclass_to_yaml`
(the writer used for `run_config.yaml`), reloaded via `instantiate`,
resolved. `moe_grouped_gemm=True` →
`TELayerNormColumnParallelLinear`/`TERowParallelLinear` +
`TEGroupedMLP`; `False` → same MLP + `SequentialMLP`. The field stays
callable after `finalize()` and `_resolve_hybrid_stack_spec()`, so a
saved config cannot regress.
- Applying the equivalent `run_config.yaml` fix to 32 iteration
checkpoints: all 32 rebuild the provider (52 layers, hidden 2304, 104
experts) with populated `MLPSubmodules` / `MoESubmodules`.
- **End-to-end exports**, 6 iterations, all rc 0, each producing exactly
the source model's 5139 tensor keys (0 missing, 0 extra), 9 shards /
41.5 GiB, all weights finite, drift from the base rising monotonically
with iteration (lm_head 0.030 → 0.092). Covered `convert.sh` CPU,
`convert.sh` GPU (4×GB300, TP=4), and
`export_distilled_megatron_to_hf.py`. Same iteration and wrapper: CPU
123 s vs GPU 134 s — GPU is not faster, since with TP=4 each rank still
builds 20.9 B of 22.3 B params and the cost is I/O plus CPU-side
conversion.
- `ruff check` / `ruff format --check` passed on the source files before
the module move.
**Not yet run:**
`tests/gpu_megatron/torch/utils/plugins/test_mbridge.py` (added here) —
the GPU allocation expired. It asserts the round-trip property verified
manually above, but its `HybridModelProvider(num_layers=2,
hidden_size=64, num_attention_heads=4)` construction is unverified. The
module move is verified only by reference grep and syntax check, so
please also run one test that uses
`tests/_test_utils/torch/megatron/models.py`. `pre-commit` was not run
either (unavailable in the environment used).
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ behavior; note
`get_te_hybrid_stack_spec` moved module
(`modelopt.torch.nas.plugins.megatron` →
`modelopt.torch.utils.plugins.megatron_layer_specs`), and a checkpoint
from 0.46.1/0.47.0 needs the `run_config.yaml` edit described in the
changelog.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ (added, not yet executed —
see Testing)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — will run `/claude review`
before marking ready.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Hybrid checkpoints now record complete layer specifications in
`run_config.yaml`.
* Recorded specifications support conversion to Hugging Face format.
* Hybrid MoE configurations support grouped-GEMM and sequential-MLP
modes.
* Configuration-based reconstruction preserves the selected MoE layout.
* **Compatibility**
* Checkpoints from earlier releases may require manually setting the
hybrid layer specification before conversion.
* **Tests**
* Added coverage confirming hybrid specifications survive configuration
serialization and can be recreated successfully.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
0a8e70804a |
Fix hf_ptq.py discarding completed PTQ run on sanity-generate() failure (#2480)
### What does this PR do? Type of change: Bug fix `post_quantize()` in `examples/hf_ptq/hf_ptq.py` ran the optional post-quantization sanity-check `full_model.generate()` unguarded, directly before `export_quantized()`. Any exception raised there aborted the whole run and discarded a completed calibration without exporting a checkpoint. Root cause (traced from [NVBug 6752977](https://nvbugspro.nvidia.com/bug/6752977), DGX Spark GB10 / DeepSeek-R1-Distill-Llama-8B / NVFP4): `get_model()` loads with `device_map="auto"`, relying on `accelerate`'s `infer_auto_device_map`/`get_max_memory()` to decide GPU vs. CPU placement. On DGX Spark's unified-memory single-GPU host, that memory probe under-reports GPU capacity, so part of even an 8B model can land on CPU — and the existing fallback shrinks the GPU budget further (`* gpu_mem_percentage`), compounding it. Calibration survives this because it never invokes the real fake-quant kernel, but the post-PTQ sanity `generate()` does, and NVFP4's dynamic-block-quantize op (`modelopt/torch/quantization/tensor_quant.py`) hard-asserts `amax.is_cuda` with no CPU fallback, so any CPU-offloaded layer crashes there — after ~5.8 hours of calibration, before export. This PR does not attempt to fix the underlying `device_map`/memory-probing behavior (unverified without the actual hardware/logs, which weren't reachable from this environment). Instead it makes the failure mode safe: a failure in the optional sanity check now only skips that check and warns, and export always proceeds, regardless of why `generate()` failed. ### Usage No new API. Behavior change only: `examples/hf_ptq/hf_ptq.py` now completes export even if the post-quantization sanity `generate()` call raises. ### Testing - Added `tests/examples/hf_ptq/test_hf_ptq_args.py::test_post_quantize_export_survives_a_failed_sanity_generate`, which drives `post_quantize()` with a `full_model.generate()` that raises and asserts `export_quantized()` still runs. - Ran `pytest tests/examples/hf_ptq/test_hf_ptq_args.py` (48 passed). - Ran `pre-commit` on the changed files (`ruff-format` reformatted line wrapping only). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ <!-- pending: run `/claude review` --> ### Additional Information Fixes NVBug 6752977. Linked JIRA: OMNIML-5932. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Quantized checkpoint export now continues when the optional post-quantization generation check fails. - A warning is shown when the generation check cannot complete. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ed5c5ed369 |
[OMNIML-5899] Export IQ checkpoints from HF and Megatron (#2447)
## Summary - add IQ format metadata and packed-weight export - support Hugging Face and TP=1 Megatron export paths - reject fused-MoE IQ export until a deployment loader owns its packed layout - document the shaped `uint8` weight contract and the fused-expert boundary - add Hugging Face, Megatron, metadata, and fused-expert export tests ## PR split This work is split into four focused PRs. Each PR targets `main` and owns a disjoint file set: 1. **Kernel** — [#2448: Add CUDA kernels for IQ packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448) 2. **Quantization** — [#2446: Add IQ quantization codecs and backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446) 3. **Export** — [#2447: Export IQ checkpoints from HF and Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447) 4. **Recipes** — [#2449: Add IQ post-training quantization recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449) The required merge order is #2448, #2446, #2447, then #2449. ## Scope This PR owns only export code, deployment documentation, and export tests. It targets `main` and should merge after #2448 and #2446. It does not contain kernel, codec/backend, or recipe files. ## Deployment consumer boundary Dense weights and individually named expert weights use the documented shaped `uint8` contract. Megatron fused-MoE IQ export is intentionally rejected with `NotImplementedError`: its payload would have shape `[num_experts, out_features, in_features // 256, payload_bytes]`, and no deployment loader in this stack currently owns that layout. Support should be enabled only with a loader integration test. ## Test coverage - [Hugging Face packed-weight export](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/unit/torch/export/test_export_weight.py) - [quantization metadata](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/unit/torch/export/test_get_quantization.py) - [Megatron unified export and fused-MoE rejection](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/gpu_megatron/torch/export/test_unified_export_megatron.py) ## Validation - all pre-commit hooks pass for the changed files - 89 focused Hugging Face export, metadata, and fused-expert tests pass locally - direct checks cover both fused-MoE export entry points for IQ1_S and IQ2_XS - Megatron GPU execution remains delegated to GPU CI - restricted-term scan passes <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for IQ1_S and IQ2_XS GGML quantization formats in unified Hugging Face and Megatron exports. * Added quantization metadata, tensor-shape recovery, packing details, and IQ2_XS size documentation. * Added validation for required block sizes and tensor parallelism settings. * **Limitations** * Fused-MoE and GPT-OSS IQ expert packing are not supported. * IQ exports require standard `weight` attributes in Hugging Face models. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d23030f91d |
[1/n] Adds skip-softmax calibration through the vLLM serving path (#1992)
### What does this PR do? Type of change: new feature Calibrates skip-softmax thresholds through the vLLM V1 execution path for FlashAttention and FlashInfer. Calibration measures the paged KV-cache path used at serving time, aggregates raw skipped/total tile counts across tensor-parallel head shards, fits separate prefill and decode curves, and exports the existing `sparse_attention_config` checkpoint schema. This uses raw counts rather than averaging per-rank sparsity ratios because TP ranks can contribute different tile populations; summing numerators and denominators before division preserves the global tile-weighted result. The vLLM adapter lives in `plugins/sparse_attn_calibration.py` rather than `SparseAttentionStatsManager`: the latter records module-local ratios for the HF calibration flow and has no aligned cross-process merge contract, while this path must merge per-sample raw counts from vLLM workers. Fitting and export still reuse `DynamicThresholdCalibrator` and the canonical conversion helpers so the model and checkpoint schema do not fork. Skip decisions depend on tile geometry. The common Triton launch boundary fixes the KV tile at 128 tokens and the prefill query tile at 128 tokens, including for direct kernel callers. Single-query decode can use a 16x128 compute tile without changing its skip decision. Measurement bypasses autotuning; serving still tunes warp and pipeline-stage counts while keeping the decision geometry fixed. ### Usage ```bash python examples/vllm_serve/calibrate_sparse_attn.py <CKPT> \ --prompts_file prompts.txt \ --target_sparse_ratio 0.7 \ --fit_logspace \ --tensor_parallel_size 4 \ --decode_tokens 32 \ --update_checkpoint_config ``` Calibration supports tensor parallelism and requires pipeline-parallel and data-parallel sizes of 1. It always writes `sparse_attention_config.json`; `--update_checkpoint_config` also merges the result into `<CKPT>/config.json`. ### Testing Latest revision `1e969cb380` (rebased onto main `02b58eb146`, 2026-09-17): - Calibration/count-fitting unit tests: **33 passed** (`test_sparse_attn_calibration.py` and `test_calibrator_fitting.py`). - Paged and contiguous calibration GPU suite: **33 passed** (`test_paged_calibrate.py` and `test_triton_fa_calibrate.py`), including NHD/HND equivalence, partial query tiles, decode counts, and malformed-cache rejection. Run with `CUDA_VISIBLE_DEVICES=1` on an RTX A6000; local GPU 0 was unavailable. - Calibration CLI tests: **21 passed** (`tests/examples/vllm_serve/test_calibrate_sparse_attn.py`). - `pre-commit run --files <four changed files>`: passed, including Ruff, mypy, and Bandit. - The new regression tests reproduced the skipped-counter truncation and missing cache-boundary checks before the fix. Calibration arithmetic and the 20-point threshold grid are unchanged. Historical validation from earlier revisions (not rerun end-to-end for this update): - `PYTHONPATH="$PWD" pytest -q tests/examples/vllm_serve/test_calibrate_sparse_attn.py tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_calibration.py` — 37 passed. - `PYTHONPATH="$PWD" pytest -q tests/gpu_vllm/torch/sparsity/attention_sparsity/test_vllm_calibration.py tests/gpu_vllm/torch/sparsity/attention_sparsity/test_sparse_attn_worker.py` — 65 passed, including kv-first, blocks-first, and packed FlashAttention cache layouts. - `PYTHONPATH="$PWD" pytest -q tests/gpu_vllm/torch/sparsity/attention_sparsity/test_vllm_runtime.py tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_config.py` — 33 passed. - `PYTHONPATH="$PWD" pytest -q tests/gpu/torch/kernels/sparsity/attention/test_paged_calibrate.py tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_calibrate.py tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_skip_softmax.py` — 31 passed, 1 skipped because the GPU lacks enough shared memory for the fp32 tile. - `pre-commit run --files <changed files>` — passed. - Historical end-to-end Nemotron 3 Ultra (GCP job `558552`), TP4, FA4, 48 RULER prompts, and 20 threshold trials: completed `0:0` with prefill `(a, b) = (9.9104, 10.8881)`, respectively +0.147% and -0.066% versus the matching 20-point reference `(9.8958, 10.8953)`. The supplied legacy fit `(14.47, 10.91)` used a different threshold grid; its `b` differs by only -0.201%, while `a` retains the known grid-weighting shift. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ Active skip-softmax fixes the calibrated decision geometry (serving still tunes warp/stage counts), and sparse-only vLLM installs fail fast for unsupported DCP, DBO/ubatching, speculative decoding, and FULL mixed-batch graphs instead of installing silently. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no copied code or new dependency. - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Pipeline parallelism is rejected during calibration because the current count-merging contract aligns records across tensor-parallel head shards, not across pipeline stages with disjoint attention layers. The unrelated HF padded-query behavior change was removed from this PR so it can be reviewed independently with its own compatibility test. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added vLLM skip-softmax calibration for paged attention, including prefill/decode support and checkpoint configuration generation. - Added Muse Glimmer AutoQuantize, Alpamayo QAD, streaming Kimi-K3 conversion, and NVFP4 activation headroom calibration recipes. - Added calibration statistics aggregation, phase-specific fitting, threshold validation, and preservation of existing sparse-attention settings. - **Bug Fixes** - Improved NVFP4 CPU/ONNX scale validation and clamping. - Added clearer handling for unsupported quantization, cache, CUDA graph, and engine configurations. - Standardized serving and calibration tile behavior. - **Documentation** - Expanded vLLM serving guidance, calibration instructions, compatibility requirements, and sparse-attention limitations. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kai Xu <kaix@nvidia.com> |
||
|
|
cf1f48fa0f |
[5565357] Fix SDXL NVFP4 export and performance (#2336)
### What does this PR do?
Type of change: Bug fix
Adds a compact SDXL and SDXL-Turbo mixed-precision FP4 recipe:
- block-16 NVFP4 for non-QKV Linear/GEMM layers;
- FP8 for Conv2d layers;
- high-precision Q/K/V projection Linears to preserve TensorRT
horizontal fusion;
- optional FP8 MHA quantization.
For SDXL FP4 export, Conv2d quantizers export directly through the
shared FP8 custom-op path. The previous `generate_fp8_scales` plus
`convert_zp_fp8` INT8 zero-point workaround is removed. The graph then
uses the existing FP8 Q/DQ normalization and `NVFP4QuantExporter`
lowering, with opset 23 for FLOAT4 support. Flux FP8 export also saves
the graph returned by its RoPE weight conversion.
This PR also changes shared exporter behavior:
- `_fp8_quantize` refreshes ONNX shape/type inference after applying the
custom FP8 operator's uint8 output metadata, affecting all FP8 ONNX
exports through this symbolic.
- `_quantized_sdpa` derives `disable_fp8_mha` from the live Q/K/V
quantizer state instead of a restored private module flag.
Other model recipe configurations remain unchanged.
### Usage
```bash
python quantize.py \
--model sdxl-1.0 \
--model-dtype Half \
--trt-high-precision-dtype Half \
--format fp4 \
--block-size 16 \
--batch-size 2 \
--calib-size 128 \
--n-steps 20 \
--quantized-torch-ckpt-save-path ./sdxl-fp4 \
--onnx-dir ./onnx-sdxl-fp4
```
### Testing
- CPU-only focused and generic NVFP4 exporter tests: 44 passed in 4.35
seconds.
- Focused Flux returned-graph save test: 1 passed.
- Required Linux unit CI at `034fe23ec` passed with the `all` dependency
set, including `tests/unit/examples/test_diffusers_fp4.py`.
- Latest changed-file pre-commit checks: all passed.
- TensorRT 10.14 on a B200 GPU:
- 302 native block-scaled NVFP4 GEMM tactics;
- 38 native FP8 Conv tactics;
- no FP4 Q/K/V projections;
- all 11 FP16 Q/K/V projection-fusion groups preserved;
- three alternating batch-2 profiles measured 18.614 ms FP4 versus
20.028 ms FP16 median UNet latency, a 7.06% reduction.
- FP8 SDXL/SD3 ONNX-to-TensorRT end-to-end runs were not executed
because they require explicit approval. The existing end-to-end test
matrix now includes SD3 FP8 alongside SDXL FP8.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ — no public API or CLI flags
change; the shared changes preserve the intended FP8 export and
attention behavior.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — the shared NVFP4 opset, FP8 shape-inference, and Diffusers
attention-policy changes are recorded under bug fixes.
- Did you get Claude approval on this PR?: N/A
### Additional Information
Tracking: [5565357]
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added SDXL support for mixed NVFP4/FP8 quantization, including
convolution and softmax handling.
- Added an SDXL quantization preset for streamlined post-training
quantization workflows.
- Expanded FP4 ONNX export support to Flux and SDXL, with improved
FP4/FP8 graph processing and export reliability.
- Added automatic quantization policy and format restoration from
checkpoints.
- **Documentation**
- Documented SDXL layer behavior, optional FP8 attention quantization,
and Blackwell/TensorRT requirements for FP4 and FP8 deployment.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
> 🤖 _Generated by Codex (AI agent)._
---------
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
|
||
|
|
17ef5b6c39 |
Split calibrated ONNX graph capabilities (#2468)
### What does this PR do? Type of change: new feature Splits the calibrated ONNX graph utilities into four cohesive capability owners: - `graph_indexing.py` owns read-only graph indexing and pattern matching. - `graph_selection.py` owns calibrated placement decisions and the runtime probes needed to make those decisions; `get_extended_model_outputs` stays with selection for that reason. - `graph_rewrites.py` owns graph transformations that are independent of Q/DQ policy. - `qdq_graph.py` owns graph-level Q/DQ analysis and policy, including policy-specific rewrites such as `remove_partial_input_qdq`; `qdq_utils.py` remains the lower-level Q/DQ node, tensor, and format utility layer. This four-way split is the intended long-term layout. All production and test callers now import the owning modules directly, and the former `graph_utils.py` catch-all module is removed. The extraction preserves the existing graph-selection and rewrite algorithms; all 43 moved functions were verified to have identical ASTs before and after the split. ### Usage ```python # N/A: behavior-preserving internal refactor. ``` ### Testing - `pytest -q tests/unit/onnx/quantization` - 376 passed, 13 intentionally xfailed - Focused tests after rebasing onto latest `main` - 47 passed, 6 intentionally xfailed - GPU ONNX quantization suite, excluding the existing AutoTune integration test - 46 passed, 3 skipped - Broader ONNX unit suite - 658 passed, 1 skipped, 13 intentionally xfailed - One ONNX Runtime `CopyTensorAsync is not implemented` failure was reproduced unchanged on the base revision. - Review-feedback validation - Public `insert_matmul_casts` star-import assertion passed. - `pytest -q tests/unit/onnx/quantization/test_graph_selection.py tests/unit/onnx/test_partitioning.py`: 42 passed. - Changed-file pre-commit hooks passed, including Ruff, formatting, mypy, Bandit, RST checks, and license checks. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: No -- direct imports from `modelopt.onnx.quantization.graph_utils` must migrate to the new capability modules. No shim is retained because removing the catch-all path is the purpose of this ownership split; the package-level quantization entry points are unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: Yes - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: Yes -- the backward-breaking section contains a symbol-level migration map. - Did you get Claude approval on this PR?: N/A ### Additional Information Follows #2457, which added focused characterization coverage for the calibrated ONNX quantization paths. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added expanded ONNX graph analysis for tensor relationships, pattern detection, FP8 attention matching, and quantization candidate selection. - Added safer Q/DQ transformations, custom-operator casting, mixed-precision configuration, and redundant-cast cleanup. - Improved quantization support for INT4, INT8, FP8, convolution, attention, and weight-selection scenarios. - **Refactor** - Reorganized ONNX quantization utilities into specialized components. - Removed the legacy graph utility module; related functionality is now available through dedicated modules. - **Tests** - Updated coverage for graph selection, partitioning, insertion points, and legacy-module removal. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> |
||
|
|
9e3d555aa1 |
[OMNIML-5899] Add IQ quantization codecs and backend (#2446)
## Summary - add IQ1_S and IQ2_XS reference codecs and a weight-only fake-quant backend - register and export both formats from the quantization package - cache compact packed weights across unchanged forwards and invalidate on tensor or config changes - use one Python-side IQ2_XS FP16 scale predictor for both reference and CUDA packing - validate packed payload metadata, normalize CUDA cache keys, and define a shared non-finite policy ## PR split This work is split into four focused PRs. Each PR targets `main` and owns a disjoint file set: 1. **Kernel** — [#2448: Add CUDA kernels for IQ packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448) 2. **Quantization** — [#2446: Add IQ quantization codecs and backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446) 3. **Export** — [#2447: Export IQ checkpoints from HF and Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447) 4. **Recipes** — [#2449: Add IQ post-training quantization recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449) The required merge order is #2448, #2446, #2447, then #2449. ## Scope This PR owns the Python codecs, backend dispatch, package registration, license attribution, CPU codec/backend tests, and CUDA numerical/reference-path tests. The native CUDA layer and direct extension tests remain in #2448; export and recipes remain in their own PRs. ## Why the codecs are separate from `qtensor` The new `ggml/` package contains stateless reference codecs and fake-quant backend functions. They transform ordinary tensors into packed format payloads and reconstruct tensors for fake quantization; they do not define persistent runtime quantized-tensor objects. `BaseQuantizedTensor` subclasses under `qtensor/` own runtime tensor objects and execution dispatch. Keeping the codecs separate avoids claiming a runtime tensor contract that these formats do not yet provide. A `qtensor` type can be added later if a runtime execution path requires one. ## Compatibility boundary The Python encoders intentionally use fixed-scale, unweighted searches. They are not intended to reproduce another encoder's bytes for every input when that encoder performs iterative scale refinement or importance weighting. Compatibility is defined by the canonical codebooks, 50/74-byte payload layouts, and pinned dequantization formulas. IQ2_XS computes the FP16 superblock scale once in the Python predictor and passes it to the CUDA packer. This removes a duplicate floating-point reduction and makes native/reference byte parity use the same scale. Non-finite input elements are treated as zero during packing in both implementations. The unit tests construct nonzero payload fields independently and validate metadata, signs, local scales, and global scales. The CUDA tests compare native packed bytes with this Python reference encoder. ## Test coverage - [IQ1_S CPU codec tests](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/unit/torch/quantization/test_iq1_s.py) - [IQ2_XS CPU codec tests](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/unit/torch/quantization/test_iq2_xs.py) - [registered backend and cache tests](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/unit/torch/quantization/test_ggml_backend.py) - [IQ1_S CUDA byte-parity, numerical, non-finite, zero-payload, and fallback tests](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/gpu/torch/quantization/test_iq1_s_cuda.py) - [IQ2_XS CUDA byte-parity, numerical, non-finite, zero-payload, underflow, and fallback tests](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/gpu/torch/quantization/test_iq2_xs_cuda.py) ## Licensing The embedded codebook data cites the pinned upstream MIT source, carries its license notice, and uses the repository's third-party license mechanism. Human OSRB/code-owner confirmation is still required; this PR does not claim that approval. ## Validation - focused lint, format, and type checks pass for all changed Python files - 36 focused CPU codec and backend tests pass locally - all 20 direct-extension and CUDA integration test cases collect locally; runtime CUDA execution remains delegated to GPU CI - restricted-term scan passes <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added GGML quantization support for IQ1_S and IQ2_XS formats. - Added quantization, dequantization, and fake-quantization workflows with pass-through gradients. - Added CPU fallback when CUDA acceleration is unavailable. - Added validation for packed weights, tensor shapes, formats, and backend options. - Added configurable chunk processing and caching for repeated quantization. - **Tests** - Added comprehensive CPU and CUDA coverage for accuracy, validation, caching, fallback behavior, and edge cases. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9895d6f129 |
[6771663] Preserve ONNX API output types when wiring casts (#2451)
### What does this PR do? Type of change: Bug fix Preserves the public ONNX graph I/O types captured at the API boundary when `PrecisionConverter` wires output casts. Type inference can change the working graph's output declaration before conversion; consulting that mutated declaration caused the required cast back to the original public type to be discarded and metadata restoration to fail. The converter now derives its I/O type map from the preserved boundary metadata and uses that map when deciding whether a cast should become a public graph output. A regression test covers an FP32 output whose working declaration is inferred as FP16, and the changelog records the corrected behavior. ### Usage ```python # No API changes are required. Existing conversions now preserve the original # public I/O declarations when keep_io_types=True. converted = convert_to_f16(model, keep_io_types=True) ``` ### Testing - Ran `pytest tests/unit/onnx/autocast/test_precisionconverter.py` (186 passed). - Ran `pytest tests/unit/onnx/autocast` (249 passed). - Ran Ruff check and format validation on the changed Python files. - Verified the original minimal end-to-end reproduction with Python 3.12 and TensorRT 10.16.1.11; conversion now completes without the output-metadata restoration error. - Verified the full CLI quantization path proceeds through the formerly failing one-Q/DQ scheme and successfully benchmarks the generated TensorRT engine. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: N/A ### Additional Information Tracking: [6771663] <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Fixed ONNX FP16 conversion to preserve public graph output types when type inference changes internal declarations. - Ensured output casts are inserted correctly when preserving input/output types is enabled. - **Tests** - Added regression coverage confirming preserved output types, correct cast insertion, and valid ONNX model generation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Ajinkya Rasane <ajinkyaashwin@gmail.com> Co-authored-by: Ajinkya Rasane <ajinkyaashwin@gmail.com> |
||
|
|
2f4da27ba5 |
Add calibrated ONNX quantization characterization tests (#2457)
### What does this PR do? Type of change: new tests Adds focused ONNX quantization characterization coverage before the calibrated INT8 and FP8 execution paths are consolidated. The new test uses deterministic literal calibration batches that produce distinct entropy and max scales for both INT8 and FP8. It directly verifies Q/DQ placement, quantized types, scale and zero-point values, graph I/O types, and opset through the public quantization API. Strict expected-failure tests also record the intended future contracts for calibration defaults and source cardinality, exact mode tokens, calibration-cache removal, legacy import removal, retired exporter helpers, and the new AutoTune helper namespace. ### Usage ```python # N/A — test-only change. ``` ### Testing Run in `nvcr.io/nvidia/tensorrt:25.06-py3`: - `pytest -q tests/unit/onnx/quantization` - 375 passed, 14 intentionally xfailed - `pytest -q tests/gpu/onnx/quantization/test_quantize_fp8.py` - 1 passed - GPU quantization suite excluding the existing AutoTune integration test - 46 passed, 3 skipped - ONNX Runtime patching and simplification GPU tests - 31 passed - Focused Ruff, formatting, mypy, Bandit, and repository pre-commit hooks passed. The complete GPU quantization run reached the existing AutoTune integration test and encountered a native TensorRT engine-build segmentation fault. No Python assertion failed before the native crash. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Expanded ONNX quantization coverage for calibrated INT8 and FP8 workflows, including graph structure, calibration values, tensor types, axes, and opset metadata. * Added validation tests for calibration sources, duplicate inputs, mode handling, removed options and legacy symbols, and AutoTune export behavior. * Documented expected future behavior through explicitly marked pending tests. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> |
||
|
|
02b58eb146 |
Fix vLLM fakequant calibration for hybrid attention models (#2414)
### What does this PR do? Type of change: Bug fix Fix fakequant calibration for hybrid attention/Mamba models, including NVIDIA Nemotron-3-Nano, on vLLM 0.26 and 0.28. The manual calibration scheduler path previously submitted requests with empty KV-cache block tables. Hybrid models require scheduler-compatible cache state during prefill; on current vLLM releases the empty tables caused the Mamba state to use the reserved null block and calibration activations became NaN. Request cleanup also no longer matched the vLLM 0.28 execution lifecycle, which could leave request-scoped state in the persistent batch. This PR: - Allocates non-null scratch blocks for every KV-cache group using the vLLM warmup reservation policy. - Supports both the vLLM 0.28 reservation helper and the equivalent vLLM 0.26 calculation. - Passes newly allocated blocks through `new_block_ids_to_zero` when that scheduler field is available. - Validates that the calibration batch fits in the configured cache and reports how to reduce calibration demand if it does not. - Cleans up calibration requests through a zero-token scheduler step on current vLLM, with a direct cleanup fallback for older runners. - Updates the example Dockerfile to default to vLLM 0.28.0 while retaining vLLM 0.26.0 through `VLLM_VERSION`. - Documents the validated Nemotron-3-Nano NVFP4 KV-cache workflow and clarifies that reducing `--max-num-batched-tokens` is not required. ### Usage Build the default vLLM 0.28.0 image: ```bash docker build -f examples/vllm_serve/Dockerfile \ -t vllm-modelopt:v0.28.0 . ``` Build with vLLM 0.26.0: ```bash docker build --build-arg VLLM_VERSION=0.26.0 \ -f examples/vllm_serve/Dockerfile \ -t vllm-modelopt:v0.26.0 . ``` Calibrate and serve Nemotron-3-Nano with NVFP4 KV-cache fakequant: ```bash KV_QUANT_CFG=NVFP4_KV_CFG QUANT_CALIB_SIZE=512 \ python examples/vllm_serve/vllm_serve_fakequant.py \ <nemotron3_nano_model_path> \ --trust-remote-code --enforce-eager -tp 8 \ --max-model-len 8192 --host 0.0.0.0 --port 8000 ``` ### Testing Validated on omniml-a0 with `NVIDIA-Nemotron-3-Nano-30B-A3B-BF16`, tensor parallel size 8, `NVFP4_KV_CFG`, `QUANT_CALIB_SIZE=512`, and `--max-model-len 8192`. No `--max-num-batched-tokens` override was used. - vLLM 0.28.0: - All 512 calibration samples completed. - No NaNs or cache-cleanup warnings were observed. - The server started and `/health` passed. - An OpenAI-compatible completion request returned coherent generated text. - vLLM 0.26.0: - Repeated the same 512-sample TP8 calibration with the official `vllm/vllm-openai:v0.26.0` image. - No NaNs were observed. - The server started, passed `/health`, and returned coherent generated text. - Docker: - Built and verified the updated vLLM 0.28.0 image. - Focused tests: - `tests/examples/vllm_serve/test_vllm_mlflow_utils.py`: 32 passed. - Cleanup failure, missing legacy API, and legacy fallback tests: 5 passed on both vLLM 0.26.0 and 0.28.0. - Repository hooks: - Targeted pre-commit hooks for every changed Python, Markdown, and Docker file: passed. - `git diff --check`: passed. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — added focused coverage for fail-closed cleanup, exception chaining, and the legacy cleanup fallback; the full regression was also validated end to end. - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information The change is quantization-format agnostic. It corrects the calibration scheduler and cache lifecycle rather than special-casing `NVFP4_KV_CFG` or using an NVFP4 cast path. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added support for configuring the vLLM version through `VLLM_VERSION`, with vLLM 0.28.0 as the default. - Added calibration and serving guidance for hybrid attention/Mamba models, including Nemotron 3 Nano with NVFP4 KV-cache fake quantization. - **Bug Fixes** - Improved calibration block handling across supported vLLM versions. - Improved calibration cleanup to preserve original errors and provide reliable fallback behavior when standard cleanup is unavailable. - **Documentation** - Documented tested versions, direct installation commands, ModelOpt setup, serving options, and guidance to avoid NaNs during batched serving. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> |
||
|
|
b16356776b |
Build the GGML IQ packing kernels as a single CUDA extension (#2462)
### What does this PR do? Type of change: Code refactoring `#2448` added the GGML IQ packing kernels as **two** torch extensions, `modelopt_cuda_ext_iq1_s` and `modelopt_cuda_ext_iq2_xs`. This merges them into one, `modelopt_cuda_ext_ggml`. The existing per-extension split in `extensions.py` exists for reasons that don't apply to the IQ formats: `get_cuda_ext` gates on CUDA `>=11` while `_fp8`/`_mx` gate on `>=11.8`, and `_mx` needs `--use_fast_math`, which must not reach the base `tensor_quant` kernels. `get_cuda_ext_iq1_s` and `get_cuda_ext_iq2_xs` differed in none of that — same `>=11.8` gate, same `-O3` flags, same `common.cuh` — so the split only compiled the shared header twice, ran nvcc twice, and grew the loader, `__getattr__`, and `precompile()` once per format. With IQ2_XXS / IQ3_S / IQ4_NL plausibly following, that scales badly. Changes: - New `ggml/ggml.cpp` holds both host-side validation wrappers and the single `PYBIND11_MODULE`, binding `iq1_s_pack` and `iq2_xs_pack` (previously each module exported a bare `pack`). Deletes `ggml/iq1_s.cpp` and `ggml/iq2_xs.cpp`; the validation logic and docstrings carry over unchanged. - `get_cuda_ext_iq1_s` + `get_cuda_ext_iq2_xs` → `get_cuda_ext_ggml`, which builds `ggml.cpp`, `iq1_s.cu`, and `iq2_xs.cu` together. The retry-on-`raise_if_failed` semantics of the old getters are preserved. - Each format keeps its kernels in its own translation unit, so adding a format is a new `.cu` plus one `module.def` — no new extension, loader, or `precompile()` line. No caller outside `extensions.py` and its tests referenced the old getters on `main`, so nothing else changes. **Note for the follow-up PRs in the `#2448` series (`#2446`/`#2447`/`#2449`): the codec layer should call `get_cuda_ext_ggml().iq1_s_pack(...)` / `.iq2_xs_pack(...)` instead of `get_cuda_ext_iq1_s().pack(...)` / `get_cuda_ext_iq2_xs().pack(...)`.** ### Usage ```python from modelopt.torch.quantization.extensions import get_cuda_ext_ggml ext = get_cuda_ext_ggml(raise_if_failed=True) iq1_s_payload = ext.iq1_s_pack(weight, iq1s_grid) # uint8 [numel / 256, 50] iq2_xs_payload = ext.iq2_xs_pack(weight, iq2xs_grid, scales) # uint8 [numel / 256, 74] ``` ### Testing Ran on a single H200 NVL (TRT-LLM `1.3.0rc27.dev202609170000` container), building the merged extension from scratch: - `pytest tests/gpu/_extensions/test_torch_extensions.py` — **24 passed** (6:44). This is the full existing IQ suite (zero-block layout, encode, dtype rejection, row-straddling rejection, invalid/negative-zero scales, byte-exact dtype equivalence, and the brute-force optimality round-trip) reparametrized onto the merged module, plus the untouched `modelopt_cuda_ext` / `_fp8` / `_mx` load tests. - Verified `precompile()` loads all four extensions and that the merged module exports exactly `iq1_s_pack` and `iq2_xs_pack` with the expected arities. - Off-GPU: compiled the three sources directly and linked them into one `.so` to confirm no duplicate-symbol collisions between the two `.cu` translation units. - `pre-commit run --files ...` passes on all changed files (ruff, mypy, clang-format, bandit, license headers). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — the removed getters were added in `#2448` (merged today, unreleased) and have no callers outside this file's own tests. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new code or dependencies; the moved wrappers keep their original attribution. - Did you write any new necessary tests?: ✅ — existing coverage reparametrized onto the merged module; no behavior change to test. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — internal refactor of an unreleased, not-yet-wired-up API. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information Follow-up to #2448. Merge before the remaining PRs in that series (#2446, #2447, #2449) land, so the codec layer is written against `get_cuda_ext_ggml` and no rename is needed afterwards. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added IQ1_S packing support through the GGML CUDA extension. - Added a unified GGML extension loader for IQ1_S and IQ2_XS packing. - Improved extension loading reliability when a cached extension is unavailable. - **Changes** - Renamed the IQ2_XS packing binding from `pack` to `iq2_xs_pack`. - Consolidated IQ1_S and IQ2_XS extension access under the shared GGML loader. - Updated GPU validation and coverage to use the unified extension interface. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ad1bad7817 |
fix(specdec): gather sharded hidden states in DFlash/DSpark AR generation (#2458)
### What does this PR do? Type of change: Bug fix Fixes the nightly `tests/regression/torch/speculative/test_dflash.py::test_dflash_ar_validate`, red every night since 2026-09-10 with: ``` WARNING: sample 0 (writing) failed: Expected all tensors to be on the same device, but got tensors is on cuda:1, different from other tensors on cuda:0 (when checking argument in method wrapper_CUDA_cat) RuntimeError: AR validation produced no results: all 3 samples failed. ``` **No PR caused this.** `ar_validate.py` loads with `device_map="auto"`, and the regression runner (`…rtxpro6000-l-2-…`) has 2 GPUs, so even Qwen3-0.6B gets sharded. `pseudo_speculative_generate()` then concatenates hidden states from `target_layer_ids`, which spans early *and* late layers — different devices once the base model is split: ```python selected = [base_outputs.hidden_states[lid + hid_offset] for lid in self.target_layer_ids] target_hidden = torch.cat(selected, dim=-1) # cuda:0 ++ cuda:1 -> RuntimeError ``` The block tensors have the same problem from the other side: they follow `input_ids.device`, while the draft module sits on the *last* base layer's device (`_place_draft`). What changed on 2026-09-10 was detection, not behaviour. #2288 made `ar_validate.py` raise when every sample fails; before it, the report block was guarded by `if results and …`, so zero results printed nothing and exited **0**. The test had been green while validating nothing. Its runtime is the tell: 23–26s in every green nightly, 21.6s in the first red one, where it does no validation work at all. The same defect is in #2288's own description, on an 8-GPU sharded **EAGLE3** checkpoint: 80/80 samples dead on the identical error, job exiting 0. So this is not DFlash-specific in principle — but EAGLE3's generate path gathers separately (`pop_and_gather_aux_hiddens`) and needs its own look, which is why this PR stops at the DFlash family. ### Fix Gather everything the draft consumes onto the draft's device, and hand results back on the caller's device since `validate_online` cats them onto the running sequence. Every `.to()` is a no-op on a single device, so single-GPU behaviour is bit-identical. `HFDominoModel` inherits `HFDFlashModel.pseudo_speculative_generate`, so it is fixed by the same change; `HFDSparkModel` overrides it and is patched in parallel (including `prev_token` feeding `markov_step`). ### Usage No API change. ```bash python examples/speculative_decoding/scripts/ar_validate.py \ --model_path <dflash ckpt> --osl 10 --num_samples 3 --steps 7 ``` ### Testing - `tests/unit/torch/speculative/plugins/test_hf_dflash.py` and `test_hf_dspark.py`: **110 passed, 1 skipped**. - The skip is the new `TestShardedBaseGeneration`, which is gated on `torch.cuda.device_count() >= 2`. It stands in for accelerate's dispatch — the base forward stays whole, but its hidden states are spread across two devices the way a real `device_map="auto"` split spreads them — and asserts both returned tensors come back on the caller's device. **I have no 2-GPU box to hand, so this test is unverified locally; it will first execute in CI.** It reproduces the reported error against the unpatched code by construction, but please treat that claim as unconfirmed until the run goes green. - The real verification is the next nightly: `test_dflash_ar_validate` should pass *and* print an AR number, which it has not done in any log I can find. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — device moves only; no-ops when the base model is on one device. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — one multi-GPU regression test (skipped without 2 GPUs). - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ — bug in an unreleased-cycle path that only ever produced a silent no-op; no user-visible behaviour to describe. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information Worth a maintainer view: `device_map="auto"` is a poor default for AR validation, which is a single-process step-by-step loop. Every sharded run of it that I can find has failed. Pinning it to one device would be a smaller blast radius than making every drafting path device-safe — but it would also cap validation at models that fit on one GPU, so I have not changed it here. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved speculative generation with sharded Hugging Face models across multiple GPUs. * Ensured generated draft tokens and outputs are returned on the caller’s input device. * Preserved Markov sampling behavior while improving device coordination. * **Tests** * Added multi-GPU coverage for sharded model generation, including output placement and draft shape validation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2ff2e1bc80 |
[OMNIML-5899] Add CUDA kernels for IQ packing (#2448)
## Summary - add native CUDA packing kernels for IQ1_S and IQ2_XS - load both extensions through the quantization extension module - validate caller metadata and launch bounds before contiguous materialization - normalize non-finite input elements consistently with the Python reference path - accept caller-computed IQ2_XS FP16 superblock scales to avoid a duplicate reduction - share common packing helpers and add direct extension compilation and boundary tests ## PR split This work is split into four focused PRs. Each PR targets `main` and owns a disjoint file set: 1. **Kernel** — [#2448: Add CUDA kernels for IQ packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448) 2. **Quantization** — [#2446: Add IQ quantization codecs and backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446) 3. **Export** — [#2447: Export IQ checkpoints from HF and Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447) 4. **Recipes** — [#2449: Add IQ post-training quantization recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449) The required merge order is #2448, #2446, #2447, then #2449. ## Scope This PR owns only native kernel sources, shared packing helpers, extension loading and build registration, and direct extension tests. It does not contain Python codecs, export code, or recipes. ## GPU test coverage Direct kernel-boundary coverage is included in this PR: - [extension compilation, zero payloads, input validation, and row-alignment checks](https://github.com/NVIDIA/Model-Optimizer/blob/74e94db9601870e3569c7c8e73506a2f08c29da8/tests/gpu/_extensions/test_torch_extensions.py) Pack/dequantize numerical, native/reference byte-parity, and non-finite-policy tests are owned by the quantization PR: [IQ1_S](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/gpu/torch/quantization/test_iq1_s_cuda.py) and [IQ2_XS](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/gpu/torch/quantization/test_iq2_xs_cuda.py). ## Dependency behavior On `main`, this PR provides optional CUDA extension loaders and direct extension tests. The Python encoders and fallback dispatch land in #2446. Until #2446 lands, no quantization path calls these getters, so a load failure reports only that the extension is unavailable. The IQ2_XS packer accepts one caller-computed FP16 scale per 256-value block. #2446 owns that predictor and passes the same values to the native and reference encoders. ## Provenance - The CUDA kernels were independently written. - They implement the packed-format contract and sign-parity convention from the pinned [llama.cpp definition](https://github.com/ggml-org/llama.cpp/blob/9b05354ec6fb58b4e665e9a39ebc40285c015638/ggml/src/ggml-common.h). - `16.875 = 15 × (1 + 1/8)` is a derived IQ1_S constant. - `0.125` is part of the encoded format. - `0.61` is our empirical IQ1_S scale predictor, not copied from upstream code. - The IQ2_XS predictor constants are owned by #2446 and are not duplicated in this kernel. Human review is still required to confirm that the attribution and license treatment are sufficient. ## Validation - repository hooks, including native formatting, pass for all changed files - extension loader and test modules compile as Python - all 20 direct-extension and CUDA integration test cases collect locally - CUDA runtime execution is delegated to GPU CI because the local host is macOS - restricted-term scan passes --------- Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
835c041c58 |
fix(specdec): resolve the eagle aux-layer preset in the vLLM hidden-state dump (#2410)
### What does this PR do? Type of change: Bug fix Fixes `nvbugs/6753684`, filed against #2080 by the ModelOpt QA Sentinel. The vLLM offline hidden-state dump rejected `--aux-layers eagle` — **the flag's own default** — so the documented invocation aborted before writing any state: ``` File "collect_hidden_states/compute_hidden_states_vllm.py", line 76, in _resolve_aux_layers_standalone ids = sorted({int(t) for t in aux_layers.split(',') if t.strip()}) ValueError: invalid literal for int() with base 10: 'eagle' ``` **Root cause.** `compute_hidden_states_vllm.py` runs in a stock vLLM container, where importing `modelopt.torch` fails (the full init chain pulls in omegaconf and friends). It therefore carries `_resolve_aux_layers_standalone`, a local copy of the preset logic in `common.resolve_aux_layers`. That copy implemented the `dflash` preset and explicit id lists, but never `eagle` — while `add_aux_layers_args` defaults to `eagle`. The HF and TRT-LLM dumps call the shared helper and were unaffected; only the vLLM path forked, and nothing compared the fork against its source. This PR resolves `eagle` inline, mirroring `hf_eagle.default_eagle_aux_layer_ids`. It also fixes a second defect the bug exposes: the function already had a message naming the accepted values, but it was unreachable, because `int()` raised first. An unrecognised preset now reports what it accepts instead of surfacing the raw `int()` error — which is what made the original failure opaque. ### Usage The previously-broken documented invocation now works: ```bash cd examples/speculative_decoding python collect_hidden_states/compute_hidden_states_vllm.py \ --model Qwen/Qwen2.5-0.5B-Instruct \ --input-data ../dataset/synthetic_conversations_1k.jsonl \ --output-dir /tmp/hs_vllm \ --max-seq-len 512 --tp 1 ``` `--aux-layers dflash` and explicit lists such as `--aux-layers 2,5,8` are unchanged. ### Testing Added `tests/unit/examples/test_vllm_hidden_states_aux_layers.py`, which pins the standalone copy to the shared implementation it mirrors: - `eagle` matches `hf_eagle.default_eagle_aux_layer_ids` across layer counts 4, 6, 8, 12, 24, 28, 32, 36, 48, 52, 61, 80 — deliberately including counts small enough that the `max(0, ...)` clamps collapse ids together. - A named regression case for `nvbugs/6753684`. - `dflash` and explicit-list behaviour unchanged. - Unknown specs (`bogus`, `EAGLE3`, `eagle3`, empty) raise the actionable message. - Out-of-range ids still rejected. Divergence here is silent — the dump would write plausible-looking hidden states from the *wrong* layers, surfacing much later as a poor acceptance rate. Hence pinning to the reference rather than asserting hardcoded lists alone. All 20 assertions verified and every pre-commit hook passes (`ruff`, `mypy`, `bandit`, RST lint, license headers). One caveat worth stating plainly: **pytest could not be run locally.** `tests/unit/conftest.py` imports `modelopt.torch.utils.distributed`, which needs `CPUOffloadPolicy` from `torch.distributed.fsdp` — absent in this machine's torch. Each assertion was executed directly against the real module instead, but CI is the first genuine pytest run. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — strictly widens accepted input; `dflash` and explicit lists behave identically. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — bug fix for a defect present in a previous release. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information The underlying fragility is the duplicated implementation, not this one missing branch. The function's own `TODO: drop this once common.resolve_aux_layers is decoupled from the heavy modelopt.torch import chain` is the real fix; the new test narrows the gap but does not close it. Worth tracking separately if the vLLM dump is expected to keep pace with new presets. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed `--aux-layers eagle` for vLLM offline hidden-state collection. * Added support for the documented `eagle` preset alongside `dflash` and explicit layer IDs. * Improved invalid-option errors to clearly list accepted formats. * Rejects `dflash` configurations when the target model has too few layers. * Continues rejecting layer IDs outside the model’s available range. * **Documentation** * Added a v0.48.0 changelog entry for the fix. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
f377b77116 |
Fail fast on non-finite AutoQuantize output gradients (#2432)
### What does this PR do? Type of change: Bug fix Fail fast when AutoQuantize receives non-finite output gradients, before accumulating sensitivity scores. The error names the affected module and suggests checking the model, data, and loss; it also identifies cuDNN SDPA backward on fully masked rows as one possible cause and gives an explicit retry workaround. Unlike the earlier revision, this does not disable cuDNN or change any attention backend settings. Invalid gradients are not zeroed or ignored, and there is no automatic retry. ### Usage No API or recipe changes. For the reproduced cuDNN failure, the caller can explicitly set `torch.backends.cuda.enable_cudnn_sdp(False)` before a fresh AutoQuantize run. ### Testing - AutoQuantize unit suite: **110 passed**. Coverage includes NaN and positive/negative infinity, module diagnostics, preventing invalid score accumulation, model-state cleanup, and unchanged SDPA backend settings. - Real-model E2E on **four GB300 GPUs**, Qwen/Qwen3.6-35B-A3B, main `8025a3dc5481129aa21fef99cb13a879e1b5847e` plus this patch, batch size 8, 512 calibration samples, and `w4a16_nvfp4_fp8_at_6p0bits-active_moe.yaml`: - Default backend: the new diagnostic fired at `model.language_model.layers.39.self_attn.q_proj`; cuDNN remained enabled and no quantized model was exported. Expected-error check passed. - Explicit cuDNN-disabled fresh run: both 64-batch calibration passes, all 16 scoring batches, optimization at **5.99 effective bits**, and checkpoint export completed with exit code 0. Verified all three indexed safetensors shards and quantization configuration; the index contains 93,563 tensors. - All applicable pre-commit checks and `git diff --check` passed. ### Before your PR is "*Ready for review*" Contributor guidelines and security guidance followed; commits are signed and signed off. - Is this change backward compatible?: Yes; finite-gradient behavior and backend settings are unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A; neither added. - Did you write any new necessary tests?: Yes. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: Yes, 0.48.0 bug fixes. - Did you get Claude approval on this PR?: No; awaiting review. ### Additional Information This improves error reporting rather than fixing the upstream cuDNN kernel. Blackwell-specificity is not established. Checkpoint deployment/reload was not tested. A calibration-only checkpoint-resume attempt completed scoring but encountered a separate `candidate_stats` KeyError. That issue is outside this patch; the successful export validation above used a fresh run without search-state resume. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * AutoQuantize now fails fast when output gradients contain non-finite values, with an actionable error identifying the affected module. * Attention backend settings are preserved and restored after successful runs and failures. * Model state is restored when sensitivity scoring encounters an error during setup or execution. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
b9cfdce8dc |
docs: add Local Hessian NVFP4 weight-scale announcement blog (#2417)
### What does this PR do?
Type of change: documentation
Adds a Local Hessian announcement blog at
`docs/source/announcements/local-hessian.rst`, covering the NVFP4
per-block
weight-scale rule that minimizes output error instead of weight error.
Contents:
- Derivation of the per-block output-error objective and its `16x16`
local
Hessian, with numbered equations.
- Results on Qwen3.5-9B: scale-setting comparison against max, MSE, and
Four-over-six, plus composition with GPTQ.
- Figure 1, a grouped bar chart of the Qwen3.8-27B W4A4 candidate scores
(BF16 in gray, the two scale rules in NVIDIA greens).
- A "Using Local Hessian" section with the config example and the
end-to-end `hf_ptq.py` command.
Two supporting changes outside the blog:
- `docs/source/_static/announcements.css`: the `shibuya` theme has no
`span.eqno` rule, so Sphinx's default `float: right` on equation numbers
cannot share a line with MathJax's full-width display block and the
number
renders *above* the equation. This anchors it to the right of the
equation
instead, and shrinks the table-note class.
-
`docs/source/announcements/assets/qwen3-27b-w4a4-scale-rule-accuracy.png`:
the Figure 1 asset.
### Usage
```python
import modelopt.torch.quantization as mtq
config = {
"quant_cfg": [...], # quantizer configuration
"algorithm": {"method": "local_hessian", "fp8_scale_sweep": True},
}
model = mtq.quantize(model, config, forward_loop)
```
### Testing
Documentation only; no code paths change. The `.rst` parses cleanly
under
docutils. The rendered page has not been checked with a full
`sphinx-build`,
so the equation-number CSS fix and the figure placement are worth an
eyeball
on the built docs before merge.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌
### Additional Information
Two items to settle before this is ready to publish:
1. The `--recipe` example points at
`modelopt_recipes/models/Qwen/Qwen3.8-27B/ptq/nvfp4_local_hessian-fp8_attn-kv_fp8_cast.yaml`,
a placeholder path derived from the existing recipe naming convention.
It
needs to match whatever lands in #2363.
2. The tables report single-run team measurements; the blog says so and
makes
no significance claims.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Documentation**
- Added guidance on NVFP4 Local-Hessian weight-scale selection,
including mathematical details, accuracy comparisons, runtime
considerations, limitations, configuration examples, and reproduction
steps.
- Updated announcement labels, headings, metadata, descriptions, and
filtering text to use “Local-Hessian.”
- **Style**
- Improved announcement formatting for equation labels, display-equation
spacing, Hessian results, table headers, and explanatory notes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: realAsma <akuriparambi@nvidia.com>
|
||
|
|
a448ba9757 |
Add end-to-end W4A4 NVFP4 + QAD tutorial for Qwen3.6-35B-A3B (#2411)
### What does this PR do? Type of change: new example + bug fix <img width="2085" height="1239" alt="image" src="https://github.com/user-attachments/assets/b9ced215-ce8c-4dbe-be74-a75c1c4714b3" /> Adds an end-to-end **W4A4 NVFP4 + Quantization-Aware Distillation** tutorial for [Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) at `examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B/`. It complements the existing Nemotron-3-Nano tutorial (pruning + distillation + FP8). Here the model is unpruned and the technique under test is **W4A4** — aggressive enough that PTQ alone leaves a measurable accuracy gap, which is what QAD exists to close. **Why W4A4 rather than weight-only NVFP4:** W4A16 measured *slower than BF16* in 10 of 12 shapes, because a BF16 activation forces vLLM onto the Marlin dequant fallback and never reaches the Blackwell FP4 tensor cores. W4A4 beats BF16 in 9 of 12 shapes (up to 1.30x) and shrinks the checkpoint 67 GiB -> 22 GiB (3.1x). **What the study found:** only 2 of 6 benchmarks show a statistically significant PTQ deficit, so those are the only two QAD can recover. IFBench is recovered to parity with BF16 (-2.62 pp -> -0.29 pp, gain of +2.33 pp, p=0.036); MMMU-Pro recovers ~40% and retains a significant gap. The other four are lossless under W4A4 to begin with. Also added: - `modelopt_recipes/model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore.yaml` — the PTQ recipe used as the QAD student, usable via `--recipe`. - `data_blend.yaml` — the token-budgeted blend config for the distillation data. - `eval_configs/*.yaml` — one NeMo Evaluator config per benchmark. tau2-bench is separate because it needs `--enable-auto-tool-choice --tool-call-parser qwen3_coder` and `deployment.command` is global to a config. **Two export fixes found while producing these checkpoints** (both change library/example behaviour, both have changelog entries under 0.48.0 Bug Fixes): - `unified_export_megatron.py` — MCore builds `embedding` on the MTP stage as well as the first, so gating export on `hasattr(model, "embedding")` wrote a **second, unreferenced copy of the vocab embedding** whenever an MTP model was exported with PP > 1. The index mapped the key to the later shard, so the extra copy never loaded but still shipped — ~1 GB for this model. Now gated on `model.pre_process`, MCore's own "this rank owns the input embedding" flag. - `export_quantized_megatron_to_hf.py` — stopped passing Megatron's `moe_router_dtype` as the router's *storage* dtype. It is a routing *compute* dtype; the parameter is bf16 in a bf16 model, so the export was widening bf16 to fp32. All 21,495,808 router values in the exported checkpoint have their low 16 bits zero, and vLLM builds the gate at the model dtype and rounds on load, so the dropped bytes carried no information. `export_mcore_gpt_to_hf` still accepts the override. ### Usage ```bash # 1. PTQ (2 GB200 nodes for EP=8) srun ... python examples/megatron_bridge/quantize.py \ --hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \ --recipe model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore \ --tp_size 1 --ep_size 8 --pp_size 1 \ --calib_dataset_name cnn_nemotron_v2_mix --calib_num_samples 1024 --calib_batch_size 1 \ --seq_length 8192 --skip_generate \ --export_megatron_path /path/to/qwen36_w4a4_megatron # 2. QAD (32 nodes x 4 GB200) python -u examples/megatron_bridge/distill.py \ --teacher_hf_path Qwen/Qwen3.6-35B-A3B --student_hf_path Qwen/Qwen3.6-35B-A3B \ --student_megatron_path /path/to/qwen36_w4a4_megatron \ --tp_size 1 --pp_size 1 --cp_size 1 --ep_size 8 \ --seq_length 32768 --mbs 1 --gbs 512 --train_iters 500 \ --lr 1e-5 --min_lr 1e-6 --lr_warmup_iters 50 --logit_kl_topk 4096 \ --recompute_granularity full --recompute_method uniform --recompute_num_layers 1 \ --no_async_save --eval_iters 0 --save_interval 50 \ --data_paths "${DATA_BLEND}" --output_dir /path/to/qad_output ``` ### Testing **Library changes.** `tests/gpu_megatron/torch/export/test_unified_export_megatron.py` gains `test_unified_export_megatron_pp2_mtp_no_duplicate_tensors`: it exports a PP=2 model built with `mtp_num_layers=1` and asserts no tensor lands in more than one shard. Verified to **fail without the fix**: ``` AssertionError: tensors written to more than one shard: {'model.embed_tokens.weight': ('model-00001-of-00002.safetensors', 'model-00002-of-00002.safetensors')} ``` The pre-existing `..._pp2_mtp_metadata_matches_shards` test cannot catch this — it fakes `_get_mtp_state_dict` on a model with no real MTP, so the last stage never builds an embedding. Ran the whole `tests/gpu_megatron/torch/export/` suite with and without the fix: identical failure sets (3 failures both ways, all `qwen3_5_moe_vl_*` from a local `ImportError: FLA is not installed`), 58 passed with vs 56 without — the +2 being the new test's two workers. `tests/unit/recipe` passes 368/368 after the recipe path move. `model.pre_process` is always present: `GPTModelExporter.__init__` raises unless the model is `GPTModel` or `HybridModel`, and both set it unconditionally. Both export fixes were also applied to the real 23 GB checkpoints and re-validated end to end: every retained tensor md5-identical, index/shard integrity re-checked, and a **full GPQA re-evaluation of the fixed checkpoint** scored 83.49 vs 84.25 before (paired per-question t-test over the same 198 questions x 16 repeats: -0.76 pp, p=0.21, not significant). **Numbers in the tutorial** come from real runs, not estimates: - **253 evaluation runs** across BF16, the published W4A16 checkpoint, W4A4 PTQ, and QAD at 50 / 300 / 500 iterations — 8 repeats per benchmark (3 for tau2-bench; GPQA is one `num_repeats: 16` run). - The published `nvidia/Qwen3.6-35B-A3B-NVFP4` checkpoint was re-evaluated under this same harness (36 runs) rather than quoted from its card, so the W4A16 row is same-harness. - Every figure and results-table value is generated from the collected `results.yml` files by a script, and I verified the README table cell-by-cell against that data after each edit. - Throughput rows were cross-checked against the recorded AIPerf sweeps; the QAD wall-clock figures against the two jobs' Slurm records (`03:34:49` + `02:09:49`). - All CLI flags in the tutorial were verified to exist in `quantize.py` / `distill.py` / `export_quantized_megatron_to_hf.py`, and `cnn_nemotron_v2_mix` against `dataset_utils.py`. The tutorial also records the non-obvious constraints found the hard way: QAD on this model requires `TP=PP=CP=1` (TP breaks quantizer `_amax` dist-checkpoint sharding, PP starves Qwen3-VL's M-RoPE of `position_ids`, CP hits a rope shard mismatch), EP must match the PTQ checkpoint, and `--logit_kl_topk` is mandatory at 32K because the dense `[seq, vocab]` fp32 logits are 30.31 GiB per tensor on a 248,320-token vocabulary. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ <!-- gpu_megatron PP=2+MTP export dedup test --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!-- 0.48.0: Megatron Framework + two Bug Fixes --> - Did you get Claude approval on this PR?: ✅ <!-- not yet run --> ### Additional Information Changelog entries are filed under **0.48.0**; the `cherry-pick-0.47.0` label has been removed. Rebased onto `main` after #2328 renamed `modelopt_recipes/huggingface` to `model_type` (it is now a compatibility symlink), so the recipe moved to `model_type/qwen3_6_moe/ptq/`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an end-to-end Qwen3.6-35B-A3B tutorial for W4A4 NVFP4 quantization and quantization-aware distillation. * Added checkpoint export, accuracy evaluation, and vLLM throughput benchmarking workflows. * Added evaluation configurations for AA-LCR, GPQA, IFBench, MMMU-Pro, SciCode, and tau2 Telecom. * Added a token-budgeted supervised fine-tuning data configuration. * Added a Megatron-Core NVFP4/FP8 quantization recipe for Qwen3.6-MoE. * **Documentation** * Added benchmark results, deployment guidance, hardware requirements, reproduction steps, HTTPS endpoint guidance, announcement filters, and tutorial links. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6a4b3f147e |
[Fix] Calibrate non-decoder modules during layerwise quantization (#2339)
### What does this PR do? Type of change: Bug fix Layerwise calibration now calibrates enabled quantizers outside transformer layers, such as `lm_head`, while hiding decoder layers from the additional calibration traversal. It also fails early when these quantizers are combined with progressive layerwise export, whose in-place conversion makes the required model calibration unsafe. ### Usage No API changes. ### Testing - `pytest_pwd tests/unit/torch/quantization/test_calib.py tests/unit/torch/quantization/test_layerwise_calibrate.py -q` — 68 passed - Pre-commit on all six changed files — passed - `git diff --check` — passed ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Layerwise calibration now supports quantizers outside transformer decoder layers, including LM-head activation quantizers. - Calibration supports models with CPU-offloaded components while preserving model behavior. - The MSE calibration utility is now publicly available. - Added warnings for calibration passes involving offloaded components. - **Bug Fixes** - Prevented unsupported export-mode calibration when outside-layer quantizers are present. - Ensured outside-layer calibration runs after transformer-layer restoration. - Preserved original forwards while hiding decoder subtrees from traversal and state collection. - Ensured cleanup and alias restoration after errors or completion. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
c7ed23a103 |
Rename modelopt_recipes/huggingface to model_type with backward-compat alias (#2328)
### What does this PR do?
Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.
Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.
- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
`$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.
### Usage
```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
--recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none
# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
--recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```
```python
from modelopt.recipe import load_recipe
load_recipe("model_type/vit/ptq/fp8") # canonical
load_recipe("huggingface/vit/ptq/fp8") # deprecated alias, resolves to the same recipe
```
### Testing
- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
`model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
the loader alias — including a recipe that pulls internal `$import`s.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.
### Additional Information
The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.
- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.
- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
- Local recipe files now take precedence over built-in recipes.
- Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
3c87751903 |
Deprecate the single-format quantization CLI flags in favour of --recipe (#2426)
### What does this PR do?
Type of change: deprecation
Deprecates the single-format quantization CLI flags in favour of
`--recipe`. Passing one now emits a `DeprecationWarning`; nothing else
changes.
| script | flags |
|---|---|
| `examples/hf_ptq` | `--qformat`, `--kv_cache_qformat` |
| `examples/megatron_bridge/quantize.py` | `--quant_cfg`,
`--kv_cache_quant`, `--weight_only` |
| `examples/torch_onnx` | `--qformat` |
`--recipe` was already authoritative over all six — silently on
`hf_ptq`, and with a runtime warning on `megatron_bridge` — and
`modelopt/recipe/presets.py` already records the intent in a comment:
*"the long-term direction is to retire `--qformat` /
`--kv_cache_qformat` in favour of `--recipe`"*. This makes that a real
deprecation.
A recipe carries the quantization config, the calibration algorithm and
the KV-cache setting in one file, so they cannot drift apart the way
separate flags can. That drift is not hypothetical: the preset path
applies no MTP exclusion while the recipe unit
`default_disabled_quantizers` disables `mtp.*`, so the same model
quantizes differently depending on which entry point was used.
#### The warning fires only when a flag is actually passed
`RecipeSupersededAction` is an `argparse.Action`, and argparse invokes
an action only for options present on the command line — never for a
default. That matters because several of these default to a *quantizing*
value (`--qformat fp8`, `--kv_cache_qformat fp8_cast`); warning on the
defaults would fire on every run, including runs that correctly use
`--recipe` and never mention the flag.
`examples/speculative_decoding/scripts/quantize_drafter.py` keeps
`--qformat` undeprecated: it has no `--recipe`, so there would be
nothing to migrate to.
### Usage
```bash
# deprecated
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> --qformat nvfp4 --kv_cache_qformat fp8_cast
# replacement
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> \
--recipe general/ptq/nvfp4_experts_only-kv_fp8_cast
```
### Testing
Three tests in `tests/examples/hf_ptq/test_hf_ptq_args.py`, all passing:
- passing `--qformat` / `--kv_cache_qformat` raises `DeprecationWarning`
and still parses the value;
- omitting them raises nothing and leaves the defaults (`fp8`,
`fp8_cast`) untouched;
- the action stays wired to both flags, so a future edit cannot drop it
while leaving the help text.
Defaults and parsed values were diffed against `main` and are unchanged
— the action stores exactly what `store` / `store_true` would have.
`ruff` findings are at parity with `main` on every changed file.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — the flags still work, they
only warn.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 Deprecations.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
Draft: the removal release for these flags is not decided here, only the
deprecation.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Deprecations**
* Legacy quantization CLI options now issue visible `FutureWarning`
messages only when explicitly provided.
* Use `--recipe` instead of deprecated options in Hugging Face PTQ,
Megatron-Bridge, and torch-to-ONNX workflows.
* Existing option values, defaults, and parsing behavior remain
unchanged.
* Weight AutoQuantize recipes without an explicit `kv_cache` setting
continue to use `--kv_cache_qformat` as a fallback.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
1030791f53 |
Fix distributed AutoQuantize scoring and share backward setup (#2231)
### What does this PR do? Type of change: Bug fix AutoQuantize can measure a group of quantized expert layers at their enclosing MLP output. That enclosing module is often a plain PyTorch container and does not carry distributed-group information, so its sensitivity score was not combined across data- or expert-parallel workers. This PR obtains the distributed groups from the quantized layers when the scoring module does not provide them. It also preserves construction order for quantized modules, scoring modules, and their registered hyperparameters so every worker accumulates scores in the same order. The temporary state needed by backward-based scoring is now managed by one shared session. The session installs and removes forward patches and invocation-specific output-gradient hooks, controls parameter gradients, and restores the active quantization recipes even when scoring raises an exception. Scoring methods remain responsible for their own score calculation. ### Usage N/A — this fixes existing AutoQuantize behavior and does not add an API or flag. ### Testing - `pre-commit run --files modelopt/torch/quantization/algorithms.py tests/unit/torch/quantization/test_autoquant.py` - `pytest -q tests/unit/torch/quantization/test_autoquant.py` — 102 passed - Added a real two-rank gradient AutoQuantize test covering MoE experts scored at an enclosing MLP. - Added regressions for deterministic hyperparameter registration, per-invocation replay for reused score modules, exact `forward`-attribute restoration, partial setup rollback, and cleanup after a scoring failure. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors). - Is this change backward compatible?: ✅ — no API or checkpoint format changes; distributed sensitivity values now include the missing reduction. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — no new feature, deprecation, breaking change, or critical release-note item. - Did you get Claude approval on this PR?: ❌ — pending review. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved quantization scoring consistency through deterministic ordering and invocation handling. * Added more reliable distributed score aggregation, including support for mixture-of-experts models. * Improved gradient-based scoring for repeated evaluations, tuple outputs, and checkpoint-compatible workflows. * Ensured model behavior and scoring state are restored after successful or failed evaluations. * Avoided unnecessary output replay when gradients are not required. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Joshua Hill <joshua.hill@baseten.co> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
51de53e48c |
[6463897] Fix narrow FP16 histogram calibration (#2412)
### What does this PR do? Type of change: Bug fix FP16 entropy calibration can fail for sufficiently narrow activation ranges because NumPy may construct the histogram bin edges at FP16 precision. NumPy 2.2 and later reject the resulting collapsed bin spacing, while earlier versions can silently return invalid, non-monotonic edges. The same precision issue can recur when ONNX Runtime merges later calibration batches. Losslessly widen FP16 activation values to FP32 while calculating and merging histograms in both ONNX entropy calibration paths, then restore the source dtype at the calibration-to-quantization boundary. This keeps the histogram bins stable without changing FP16 Q/DQ scale or graph dtype semantics. FP32 inputs and public APIs are unchanged. For full-range FP16 activations, ONNX Runtime can overflow while subtracting FP16 calibration endpoints before it widens the result. Retry only a non-finite FP16 scale calculation with FP32 endpoints, then cast the finite scale back to FP16. Existing finite FP16 calculations and all non-FP16 calculations continue to use ONNX Runtime's original result. The fallback intentionally patches only the `qdq_quantizer` binding used for calibrated activation ranges. Initializer and weight quantization continue to use ONNX Runtime's existing `quant_utils` path unchanged; full-range FP16 weight scaling is outside this calibration fix. The regression tests exercise both collectors across initial collection, an equal-range merge, and an expanding-range merge. They also verify the internal FP32 histogram and external FP16 calibration-range contract, including finite saturation when restoring sanitized values. A real entropy calibration test covers full-range FP16 values and verifies finite FP16 Q/DQ scales and a loadable ONNX Runtime graph. The AutoCast integration verifies the same FP16 scale-type contract through the public quantization path. ### Usage N/A — no API or usage change. ### Testing All tests ran with CUDA hidden. - NumPy 1.26.4: affected histogram, AutoCast, and ORT-patching modules: `39 passed`. - NumPy 2.2.3: affected histogram, AutoCast, and ORT-patching modules: `39 passed`. - NumPy 2.3.5: affected histogram, AutoCast, and ORT-patching modules: `39 passed`. - Public INT8 entropy quantization with full-range FP16 calibration data: finite FP16 Q/DQ scales, full ONNX check passed, and the CPU ONNX Runtime session loaded. - Changed-file pre-commit hooks: passed. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Follow-up to #1558. > 🤖 _Generated by Codex (AI agent)._ --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com> |