mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
main
1310
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ad8cd63847 |
Share one CUDA encoder per IQ family (#2615)
### What does this PR do? Type of change: refactor (no behaviour change) The five GGML IQ CUDA encoders were five copies of the same search. IQ2_XS and IQ2_XXS shared 139 of their roughly 150 lines of encoder and launcher code, and IQ2_S 113 of them. IQ1_S and IQ1_M had the same structure with a different choice space. Every scaled packer also validated its scales twice, in the `ggml.cpp` pybind wrapper and again in the CUDA entry point. This PR keeps **one encoder per family**, as two templates: - **`iq2_family.cuh`** for IQ2_XS, IQ2_XXS and IQ2_S. The grid sits in shared memory, the 16 local scales are scored per group, and each vector then takes its best entry under the chosen scale. A format supplies its group shape, whether it stores seven sign bits and recovers the eighth from parity, and a `store()` that writes the chosen entries, sign masks and local scales into its layout. - **`iq1_family.cuh`** for IQ1_S and IQ1_M, over the shared ternary grid. Each group picks one of `kChoices` options. With `kSharedShift` the option also fixes the ±1/8 delta (IQ1_S: `shift * 8 + local`); otherwise each vector picks its own (IQ1_M). IQ1_S's scale kernel now writes FP16 scales, so both IQ1 formats take the same input. Each format file is now one `Format` struct, holding its layout constants and `store()`, plus its entry point: 58–100 lines each. Validation lives once in `common.cuh`, as `check_pack_inputs` and `check_scaled_pack_inputs`. `ggml.cpp` binds the CUDA entry points directly instead of through five wrappers. **The kernel sources shrink from 1,536 to 1,241 lines** (+665 / −960). This is the first of two PRs. #2604 builds on it: it adds CUDA decoders as a `decode()` next to each format's `store()`, and makes export reuse fake quant's packed payloads. ### Testing **Nothing changes in the output.** Before the refactor I hashed 40 outputs: 5 formats × float32/bfloat16/float16/float64 inputs × encode and decode, on a weight with zero, tiny, oversized and non-finite blocks. All 40 hash the same afterwards. **Encode speed is unchanged.** Old and new were timed alternately for four rounds, in both orders, on an idle RTX PRO 6000 with a 5632×2048 weight. They were within 1% for every format: IQ1_S 37.6 / 37.6 ms, IQ1_M 37.0 / 37.0, IQ2_XXS 10.9 / 10.9, IQ2_XS 11.9 / 12.0, IQ2_S 15.5 / 15.4. - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **49 passed** - **Validation reports the same errors in the same order.** Over 5 formats × 8 combinations of bad arguments (devices, dtype, width, grid shape, scales dtype, length and sign), every first error matches main's. - `tests/gpu/_extensions/test_torch_extensions.py`: the validation-message tests pass. #2515's two Q8_0 tests fail identically on a clean `main` on this GPU. - IQ unit tests (`test_ggml_backend.py`, `test_iq_formats.py`, `test_convert_hf_config.py`, `test_presets.py`, `test_export_weight.py`): **173 passed** ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ Same bindings, messages and bytes. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code sources or dependencies. - Did you write any new necessary tests?: N/A. A refactor with no behaviour change, verified by the hashes above and the existing GPU tests. - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: **this** → #2604 (pack each IQ weight once and decode on CUDA). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Quantization now checks that inputs and grids are CUDA tensors on the same device, with compatible shapes. Scaled formats also validate scale type, shape, and finite, non-negative values. * **Improvements** * IQ1 and IQ2 formats share common encoding paths while retaining their format-specific output layouts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
fadbf74d31 |
Make the QAT/QAD guide the central place for concepts, background, and framework selection (#2590)
### What does this PR do? Make the QAT/QAD guide the central place for concepts, background, and framework selection. Have the Hugging Face and Megatron Bridge tutorials link back to it instead of repeating explanations of QAT and QAD, keeping the tutorials focused on setup and execution. In main QAT/QAD guide make links to all relevant blogposts. Note: MBridge example doc is out of scope for this MR. ### Testing Doc changes only, manual check. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - Did you write any new necessary tests?: N/A docs changes only ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Expanded the QAT/QAD guide with workflows, use cases, and a comparison, including QAD’s use of a frozen BF16 teacher and logit-level loss to recover accuracy after quantization. * Updated README and quick-start navigation to link to the combined QAT/QAD guide; the previous standalone QAT guide now redirects readers there. * Reorganized the LLM QAT tutorial: recipe guidance is now part of the end-to-end example, while trainer examples and Python quantize-and-fine-tune guidance are in Advanced Topics. The tutorial also notes Triton accelerated kernels. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com> |
||
|
|
4f62d418f4 |
Use TensorRT optimization level 0 in Torch ONNX example tests (#2619)
### What does this PR do?
Type of change: Bug fix
The Torch ONNX example tests can exceed their 300-second deadline while
building a ResNet50 INT8 TensorRT engine at optimization level 4. Add
`--trt_builder_optimization_level` to the vision example and select
level 0 in the existing integration tests. The example and helper retain
level 4 by default. Quantization, ONNX export, residual Q/DQ assertions,
engine execution, and test timeout limits are unchanged.
Document the build-time versus inference-performance tradeoff in the
example README.
### Usage
```bash
cd examples/torch_onnx
python torch_quant_to_onnx.py \
--timm_model_name resnet50 \
--recipe timm/resnet/ptq/int8 \
--onnx_save_path resnet50.int8.onnx \
--calibration_data_size 1 --no_pretrained \
--trt_build --trt_builder_optimization_level 0
```
### Testing
Validation used the TensorRT 26.05 container, TensorRT 10.16.1.11, and
PyTorch 2.13.0, with the existing 300-second per-test deadline.
- RTX 6000 Ada: **8 passed**, covering FP8 and INT8 on ViT, Swin,
SwinV2, and ResNet50. ResNet50 INT8 passed in 96.65 seconds.
- RTX PRO 6000 Blackwell Max-Q: **20 passed, 3 existing skips**,
covering the complete test file. ResNet50 INT8 passed in 81.27 seconds;
the baseline timed out at 300 seconds.
```bash
# RTX 6000 Ada: supported FP8/INT8 cases
python -m pytest tests/examples/torch_onnx/test_torch_quant_to_onnx.py \
-k '(fp8 or int8) and not mxfp8' --cov
# RTX PRO 6000 Blackwell: complete example test file
python -m pytest tests/examples/torch_onnx/test_torch_quant_to_onnx.py --cov
```
On each GPU, levels 4 and 0 used the same exported ResNet50 INT8 graph
with TensorRT 10.16.1.11:
- RTX 6000 Ada: TensorRT-reported engine build time decreased from 122.6
seconds at level 4 to 27.7 seconds at level 0. Both builds and inference
runs succeeded.
- RTX PRO 6000 Blackwell Max-Q: the original level-4 test hit its
300-second deadline during the engine build; level 0 built that saved
graph in 10.9 seconds and completed inference successfully.
All pre-commit checks passed for the changed files. The existing
integration tests exercise the real engine build; no redundant mocked
tests were added.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — the existing integration
tests were updated to exercise level 0; no new test cases were needed.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — minor example/CI fix; library behavior and example defaults are
preserved.
- Did you get Claude approval on this PR?: N/A — not requested for this
focused change.
### Additional Information
Example timeout: [ResNet50 INT8 CI
failure](https://github.com/NVIDIA/Model-Optimizer/actions/runs/36560176214/job/109381161169).
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added a configurable TensorRT builder optimization level for engine
builds, with a default of 4 and support for values from 0 to 5.
* Documented that level 0 can speed up builds, while lower optimization
levels may reduce inference performance.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
|
||
|
|
9c1cf80f1b |
test(megatron_bridge): cover context parallelism in the VLM QAD test (#2592)
### What does this PR do? Type of change: new tests Runs the VLM case of `test_qad` under context parallelism, so QAD on a Qwen3-VL model is covered on the path that until now could not run at all. `Qwen3VLMultimodalRotaryEmbedding` CP-shards its own embedding, so the batch has to hand it full-length `position_ids`. Megatron-Bridge's `get_batch` was sharding them too, leaving the rotary embedding at `seq / cp**2` against hidden states at `seq / cp`. The fix is upstream in [NVIDIA-NeMo/Megatron-Bridge#6243](https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/6243); this PR is the coverage that would have caught it. The VLM case moves from tensor to context parallelism and the PTQ step is sized to the same TP, so QAD still loads a matching checkpoint. The LLM case is unchanged (`tp_size=num_gpus, cp_size=1`), and `test_distill_vlm` still covers TP for a VLM, so nothing loses coverage. ### Usage ```bash # Unchanged: --cp_size is already a distill.py flag. On a container carrying Megatron-Bridge#6243 # it now works for VLMs, where it previously died in the rotary embedding. python examples/megatron_bridge/distill.py --cp_size 2 --tp_size 1 ... ``` ### Testing On 2x RTX 6000 Ada, in `nemo:26.08` with Megatron-Bridge#6243 on `PYTHONPATH`: - `test_qad[qwen3_5_moe_vl]` at `--tp_size 1 --cp_size 2` — FP8 PTQ, QAD across 2 CP ranks, export; quantizers survive and the vision tower is byte-identical. **1 passed (183 s).** Without the upstream fix the same run dies with `AttributeError: 'NoneType' object has no attribute 'ndim'` in `rope.py:175`. - `test_qad[qwen3]`, the unchanged LLM path — **1 passed (194 s).** - Gate check: on today's `nemo:26.08` (no #6243) the probe resolves `False` and the VLM case stays at `cp_size=1`, byte-identical to current CI; with #6243 it resolves `True` and runs at `cp_size=num_gpus`. On a 1-GPU runner it degenerates to today's config either way. - `pre-commit run --files ...` clean (ruff check, ruff format, mypy, bandit, markdownlint). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — existing tests extended rather than new ones added. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — test coverage and one doc line; no feature, break, deprecation, or fix for a released bug. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information - Depends on [NVIDIA-NeMo/Megatron-Bridge#6243](https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/6243). Safe to merge before it lands: the gate keeps the VLM case at `cp_size=1` until a container ships the fix, at which point the coverage switches on by itself. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the Qwen3.6 QAD instructions to keep tensor and pipeline parallelism set to 1, while allowing context parallelism to increase for longer sequences with the `nemo:26.10` container. * **Tests** * QAD validation now selects parallelism settings based on whether the Megatron-Bridge context-parallel fix is available, and reports when multi-GPU VLM coverage is reduced. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7d9e07d14b |
Count only added lines toward the PR size budget in AGENTS.md (#2616)
### What does this PR do? Type of change: documentation Changes the PR sizing rule in `AGENTS.md` to count only **added source** lines toward the ~500-line budget, instead of total changed lines. Deletions are cheap to review, so a PR that mostly removes code shouldn't be pushed into a split. Tests and docs are excluded too, since every sub-PR has to carry its own tests. The check uses the insertions count from `git diff --shortstat` with a pathspec that excludes `tests/` and `docs/`. ### Usage ```bash git diff --shortstat origin/main...HEAD -- . ':!tests' ':!docs' # N files changed, X insertions(+), Y deletions(-) -> compare X against ~500 ``` ### Testing - `pre-commit run --files AGENTS.md` (markdownlint passes). - Ran the pathspec against recent commits (#2595, #2513) to confirm it drops test and doc lines from the count. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information Follow-up to #2494, which introduced the sizing guidance. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated review guidance to measure pull request size by added source lines, excluding deletions, tests, and documentation. The guidance retains the recommendation to check the size before opening a review. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
aa89722d38 |
[6/6] Add the IQ1_M CUDA encoder and register the format (#2595)
### What does this PR do? Type of change: new feature **Second of two PRs adding IQ1_M** (1.75 bits per weight). #2513 landed the PyTorch codec; this PR adds its **CUDA encoder** and makes the format reachable. With it, ModelOpt supports all five GGML IQ formats at one and two bits. - the CUDA encoder, its binding and extension build wiring, plus the CUDA path in `quantize_iq1_m` - an `IQFormat` record and **one `IQ_FORMAT_REGISTRY` entry**, so backend dispatch, both exporters and `convert_hf_config` take it from there - the `ggml` package export - the `general/ptq/iq1_m` recipe, its presets, `ptq.md` and a CHANGELOG entry The kernel lands with the registration so every registered format keeps a CUDA encoder. ### The kernel In the kernel the delta shift is free per group, so it sits above the entry index in the sort key: a tie still prefers the lower shift and then the lower entry, as the reference encoder does. The 2048-entry grid IQ1_M shares with IQ1_S is 64 KiB, past the 48 KiB static shared-memory limit, so both kernels read it from global memory and rely on the cache. | 5632×2048 weight | torch | CUDA | | |---|---|---|---| | IQ1_M encode | 5.6 M elem/s | **318 M elem/s** | **57×** | ### Shared with IQ1_S rather than copied The two IQ1 kernels load each vector, score it against a grid entry and apply the ±1/8 shift the same way. So those three steps move into `common.cuh` as `load_vector`, `grid_terms` and `shifted_error`, and IQ1_S uses them too. **IQ1_S's packed bytes are unchanged**: its CUDA output on a 5632×2048 weight hashes the same before and after, and so does IQ1_M's, compared against the pre-split version of this change. IQ1_S encodes at the same speed (306 M elem/s). ### Usage ```bash python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq1_m ``` ### Testing Registering the format brings it under every registry-driven test with no IQ1_M-specific test code: backend dispatch, weight caching, the `num_bits` guard, `convert_hf_config` metadata, Megatron export and the `TensorQuantizer` tests in the shared battery. The shared CUDA battery gains one row. - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py`, `tests/unit/recipe/test_presets.py`: **166 passed** - broader unit sweep (`-k 'ggml or iq or gguf or registry'` over quantization, export and recipe tests): **221 passed**. The one failure, `test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes`, is a `torchvision` import error in my environment, unrelated to IQ. - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **49 passed** on RTX PRO 6000 Blackwell (sm_120), 7 of them IQ1_M, including CUDA-vs-PyTorch encoder parity - `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k 'iq or ggml'`: **45 passed** (9 tests × 5 formats) in `nvcr.io/nvidia/nemo:26.08` - `tests/examples/hf_ptq/test_llm_ptq.py -k iq1_m`: **passed** - reconstruction error falls monotonically across all five formats, pinned by a test - `general/ptq` now holds 31 recipes; `ptq.md` is updated. Rebased onto `main` after #2513 merged. The resulting tree is identical to the one the runs above tested, and the unit set was rerun on it: 166 passed. On this GPU, two of #2515's Q8_0 tests in `tests/gpu/_extensions/test_torch_extensions.py` fail: `test_cuda_ext_q8_0_zero_and_roundf_layout` and `test_cuda_ext_q8_0_dequantizes_with_small_error`. They fail identically on a clean `main` checkout, so they are not from this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code sources or dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S codec) → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M codec), all merged → **this**. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added IQ1_M weight-only quantization at 1.75 bits per weight, with CUDA acceleration and a 256-value block size. * Added an IQ1_M post-training quantization recipe for eligible linear layers; calibration data is not required. * Added IQ1_M to the supported GGML-compatible formats and recipe listings. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
333ace1bc9 |
Fix multi-turn synthetic generation and mark incomplete outputs (#2571)
### What does this PR do?
Type of change: Bug fix
Fixes synthetic conversation generation that assumed alternating user
and assistant messages. That assumption skipped user turns in prompt
skeletons and mishandled
leading system messages.
• Preserve conversation history. Regenerate every user turn while
retaining system messages and generated reasoning for subsequent
requests.
• Expose generation controls. Support model-specific request parameters,
configurable timeouts, and server-managed response budgets.
• Handle failures explicitly. Reject empty final answers and unsupported
tool calls. Failed conversations remain retryable without duplicating
saved output.
• Identify incomplete outputs. Mark length- and repetition-stopped
conversations as truncated, preserve stop metadata, and stop generating
follow-up turns.
### Usage
Run from the repository root against a compatible Qwen server with
reasoning parsing enabled:
python examples/speculative_decoding/scripts/server_generate.py \
--data_path input_conversations/train.jsonl \
--output_path synthetic/train.jsonl \
--url http://localhost:8000/v1 \
--model model \
--max_tokens 0 \
--request_timeout 3600 \
--extra_body
'{"chat_template_kwargs":{"enable_thinking":true,"preserve_thinking":true}}'
The model name must match the server’s configured name. Filter truncated
conversations before training.
### Testing
Focused regression tests: 15 passed.
The tests execute the command-line entry point using the real OpenAI
client library with mocked HTTP transport.
Coverage includes multi-turn generation, system prompts, reasoning
preservation, request parameters, failure recovery, resume
deduplication, truncation, and invalid
responses.
python -m pytest \
--confcutdir=tests/examples/speculative_decoding \
tests/examples/speculative_decoding/test_server_generate.py -q
The isolated test configuration avoids an unrelated parent configuration
import failure. All applicable pre-commit checks passed for the
generator, tests, and
documentation.
### Before your PR is "Ready for review"
• Is this change backward compatible?: ✅ Existing valid inputs,
defaults, conversation output structure, and resume behavior remain
supported. Invalid inputs and
failed requests now raise errors instead of being silently accepted.
• Copied code or new PIP dependencies?: N/A. No new third-party code or
dependencies were added.
• Did you write any new necessary tests?: ✅ Added focused command-line
regression tests.
• Did you update Changelog?: N/A. These are example-script correctness
fixes, not critical released library fixes.
• Did you get Claude approval on this PR?: ❌ Not yet obtained.
### Additional Information
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Data generation supports `conversations` and `messages` inputs,
preserves reasoning content, and accepts additional chat settings and
configurable request timeouts.
* Failed conversations are recorded separately, with options to retry
failures or exit when errors occur. Resume behavior distinguishes
retryable failures from rejected inputs.
* Outputs identify conversations truncated by length or repetition
limits.
* **Documentation**
* Updated data preparation guides with generation setup, input formats,
failure handling, resuming, and training guidance.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
|
||
|
|
bc5d2c3610 |
Update roadmap link in README.md (#1700)
### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the Roadmap link in the README to point to the current roadmap reference. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Trenton Starkey - Product @ NVIDIA <trenton.starkey@outlook.com> |
||
|
|
e5b63320ab |
[5/6] Add the IQ1_M codec (#2513)
### What does this PR do? Type of change: new feature (not yet user-reachable) **First of two PRs adding IQ1_M** at 1.75 bits per weight, just above IQ1_S. This one lands the **PyTorch codec**: encoder and decoder. It is deliberately **not registered**, so no quantizer dispatches to it and the `ggml` package does not export it. #2595 adds the CUDA encoder, registers the format and adds its recipe. With both, ModelOpt supports all five GGML IQ formats at one and two bits. On the mixed-precision checkpoint #2511 measured (`unsloth/Qwen3.8-27B-GGUF`), IQ1_M covers **25 tensors and 1.2 B parameters**. With all five formats we can read 89.0% of that file; the rest is k-quants and F32. ### What's distinctive about it **IQ1_M is the most irregular layout of the five.** There is no leading block scale field at all. The FP16 super-block scale is reassembled from the top nibble of each of four scale words: ```c scale.u16 = (sc[0] >> 12) | ((sc[1] >> 8) & 0x00f0) | ((sc[2] >> 4) & 0x0f00) | (sc[3] & 0xf000); ``` It is also finer grained than IQ1_S: a local scale per **two** groups rather than four, and a delta shift chosen **per group** rather than per sub-block. That is where its extra 0.1875 bits go. ### Shared with IQ1_S rather than copied IQ1_M searches exactly as IQ1_S does: the same 2048-entry grid, the same ±1/8 delta, every (shift, local scale) choice for every 8-value vector. It differs only in how it selects among those choices afterwards. So the search moves out of IQ1_S's encoder into `_search_shifted_grid`, which both call, and `iq1_m.py` keeps only its selection and packing. **IQ1_S's encoded bytes are unchanged**, checked by hashing its output before and after on a fixed input. ### A scale-anchor correction IQ1_M anchors its scale differently from IQ1_S: the ratio **rises with a block's peak-to-RMS** rather than being flat, and clamps higher. It uses `clamp(0.58 + 0.035 * peak_to_rms, 0.65, 0.95)` against IQ1_S's flat `0.61`. Measured over 15 Qwen3.8-27B MLP weights: | | flat 0.61 | correct anchor | | |---|---|---|---| | relative reconstruction MSE | 0.17372 | **0.17291** | **−0.47%** | It is consistent on every tensor, with no outliers. The anchor changes quality without touching layout, so neither round-trip nor conformance tests would catch it drifting. `test_scale_anchor_follows_peak_to_rms` now pins it, for all five formats; see Testing. ### Family parity Two surface asymmetries close here, so the five are uniform. `IQ1_S` now exposes `_predict_iq1_s_scales` like the other four, instead of computing its anchor inline. `IQ1_M` exposes `iq1_m_grid`, aliasing the IQ1_S table it shares. ### Testing **The decoder is validated against llama.cpp's own output, not just round-tripped:** ``` IQ1_M: 25 tensors, 4,730,880 blocks → 0 mismatched, max|diff| 0.0 ``` This mattered: **my first IQ1_M decoder had a real bug.** A `repeat_interleave` on the wrong axis produced `[h0,h1,h0,h1]` where llama.cpp needs `[h0,h0,h1,h1]`. A round-trip against our own encoder still passed, because the encoder made the matching mistake. Only comparison against bytes we did not produce caught it. Blocks from that checkpoint ship as conformance vectors, and mutation testing confirms they catch a mis-set scale nibble. The decoder unpacks every field in one vectorized pass, since fake quant decodes on every forward: 5.2 ms for a 5632×2048 weight (IQ1_S: 3.3). - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py`, `tests/unit/recipe/test_presets.py`: **153 passed**, 15 of them IQ1_M codec cases, including the llama.cpp conformance check - `test_scale_anchor_follows_peak_to_rms` pins every format's scale anchor. It predicts scales for blocks whose peak-to-RMS is exactly 1, 4, 8 and 16, reaching both clamps and two points on each slope, and compares them against anchors written out in the test. Mutations each fail exactly the mutated format: reverting IQ1_M to IQ1_S's flat 0.61, moving either IQ1_M clamp, changing its taper by 0.001, moving an IQ2_S or IQ2_XS clamp, and changing IQ1_S's anchor to 0.62. - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed**. IQ1_S's CUDA-vs-PyTorch parity still holds after its encoder refactor. - IQ1_S and IQ1_M PyTorch encoder output and IQ1_M decoder output hash identically to the pre-split version of this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ IQ1_M adds no codebook; it reuses the IQ1_S table already carried in `codebooks.py`. The new conformance vectors come from `unsloth/Qwen3.8-27B-GGUF`, which is Apache-2.0 like its base model `Qwen/Qwen3.8-27B`; the vectors' docstring now records that. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A. Nothing is user-reachable yet; #2595 carries the entry. - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S codec) → #2565 (IQ2_S CUDA encoder and registration), all merged → **this** → #2595 (IQ1_M CUDA encoder and registration). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added IQ1_M quantization and dequantization for compact, GGML-compatible blocks of 256 values. * Added access to the IQ1_M grid and configurable chunk sizes for processing data. * **Bug Fixes** * Improved IQ1_S scale prediction and grid-search organization while preserving its encoding behavior. * **Tests** * Added IQ1_M conformance data and included the format in shared IQ-format test coverage. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
ad9ea97a4b |
Fix vLLM compilation guard for models without marker (#2518)
### What does this PR do?
Type of change: Bug fix
Makes the vLLM `disable_compilation` context manager support inner model
implementations that do not predefine a `do_not_compile` attribute,
including GLM-5.3. The context manager now installs the marker
temporarily and removes it afterward, while preserving and restoring
existing marker values for other vLLM models.
Adds regression coverage for both supported wrapper layouts:
`model.model` and `model.language_model.model`.
### Usage
```python
with disable_compilation(model):
mtq.quantize(model, quant_cfg, forward_loop=calibrate_loop)
```
No caller changes are required.
### Testing
- Ran `tests/gpu_vllm/torch/quantization/test_vllm_dynamic_modules.py`:
24 passed with vLLM 0.28.
- Ran pre-commit on both changed files: all applicable hooks passed.
- Installed this branch into `vllm/vllm-openai:glm53-flash` on OCI-JHB
and served the GLM-5.3-Flash BF16 checkpoint with
`QUANT_CFG=NVFP4_DEFAULT_CFG`, TP=4, eager mode, and BF16 KV cache.
- GLM passed the previous `do_not_compile` failure point, inserted 1,700
quantizers, enabled 456 weight quantizers, reached a healthy API server,
and returned a relevant manual prompt response.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — integration compatibility fix; no user-facing API change.
- Did you get Claude approval on this PR?: N/A
### Additional Information
Validated against GLM-5.3-Flash using ModelOpt commit
`869b64fcee0b20be323663449b00e8c52940a289`.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Bug Fixes**
- Compilation settings are now handled across supported nested model
configurations and restored after calibration, including when errors
occur.
- Calibration inputs correctly exclude padding when an attention mask is
provided and reject empty sequences.
- vLLM warmup reserves the required cache space for supported tail-cache
configurations.
- Serving startup supports an alternate vLLM launcher import path when
the OpenAI entrypoint is unavailable.
- **Compatibility**
- The vLLM serving example now defaults to vLLM 0.30.0 and documents
tested support for Nemotron 3 Nano hybrid attention/Mamba serving on
vLLM 0.26.0 and 0.30.0.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
|
||
|
|
3091b8ff69 |
[4/5] Add the IQ2_S CUDA encoder and register the format (#2565)
### What does this PR do? Type of change: new feature **Second of two PRs adding IQ2_S** (2.5625 bits per weight). #2512 landed the PyTorch codec; this PR adds its **CUDA encoder** and makes the format reachable: - the CUDA encoder, its binding and extension build wiring, plus the CUDA path in `quantize_iq2_s` - an `IQFormat` record and **one `IQ_FORMAT_REGISTRY` entry**, so backend dispatch, both exporters and `convert_hf_config` take it from there - the `ggml` package export - the `general/ptq/iq2_s` recipe, its presets, `ptq.md` and a CHANGELOG entry The kernel lands with the registration so every registered format keeps a CUDA encoder. On the mixed-precision checkpoint #2511 measured (`unsloth/Qwen3.8-27B-GGUF`), IQ2_S covers **9 tensors and 0.6 B parameters**. ### The kernel IQ2_S's **1024-entry codebook is twice IQ2_XS's**, which makes its search the most expensive in the family. The codebook and its norms take 36 KiB of shared memory, the most of any IQ kernel but inside the 48 KiB static limit, so they are declared statically like the IQ2_XS and IQ2_XXS kernels. That cost is why the kernel matters more here than anywhere else: | | torch | CUDA | | |---|---|---|---| | IQ2_S, 5632×2048 weight | 0.8 M elem/s | **725.7 M elem/s** | **907×** | | extrapolated to a 27B model | ~9.8 hours | **~37 s** | | ### Usage ```bash python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq2_s ``` ### Testing Registering the format brings it under every registry-driven test with no IQ2_S-specific test code: backend dispatch and weight caching, the `num_bits` guard, `convert_hf_config` metadata (uniform and mixed precision), all 9 Megatron export tests, and the two `TensorQuantizer` tests in the shared battery. The shared CUDA battery gains one row. - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py`, `tests/unit/recipe/test_presets.py`: **134 passed** - broader unit sweep (`-k 'ggml or iq or gguf or registry'` over quantization, export and recipe tests): **192 passed**. The one failure, `test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes`, is a `torchvision` import error in my environment, unrelated to IQ. - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed** on RTX PRO 6000 Blackwell (sm_120). 7 of them are IQ2_S: CUDA-vs-PyTorch encoder parity, determinism, reconstruction at scale, zero and non-finite policy, float64 input and the fallback path. - `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k 'iq or ggml'`: **36 passed** (9 tests × 4 formats) in `nvcr.io/nvidia/nemo:26.08` - `tests/examples/hf_ptq/test_llm_ptq.py -k iq2_s`: **passed**. TinyLlama PTQ through unified HF export writes `quant_algo: IQ2_S`, `block_payload_bytes: 82`, and `down_proj` packed as `(2048, 22, 82)` uint8. - `general/ptq` now holds 30 recipes. - The shared-memory change in `b7739d5d0` leaves the packed bytes identical (same hash on a 5632×2048 weight), and packing runs at 849.1 M elem/s against 825.7 before on RTX PRO 6000. The GPU battery was rerun: 42 passed. All of the above was rerun after rebasing onto `main` at `c2aaa44f6`. That base adds a Q8_0 packer to the same GGML extension (#2515), and changes the hf_ptq example and the export code this format goes through. The packed IQ2_S bytes still hash the same. On this RTX PRO 6000 (sm_120), two of #2515's own Q8_0 tests in `tests/gpu/_extensions/test_torch_extensions.py` fail: `test_cuda_ext_q8_0_zero_and_roundf_layout` and `test_cuda_ext_q8_0_dequantizes_with_small_error`. They fail identically on a clean `main` checkout, so they are not from this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code sources or dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: #2511 (IQ2_XXS, merged) → #2525 (format registry, merged) → #2512 (IQ2_S codec, merged) → **this** → #2513 (IQ1_M). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added IQ2_S weight-only quantization for eligible linear layers, at 2.5625 bits per weight. * Added a PTQ recipe that requires no calibration data. Weights must meet the existing 256-value block-size constraint. * Added CUDA-accelerated packing for CUDA weights, with a Python fallback when the CUDA extension is unavailable. * **Documentation** * Updated the PTQ recipe catalog and IQ-format size tradeoffs. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
1576f7d4ad |
Increase diffusers example-test timeout to 60 minutes (#2591)
### What does this PR do? Type of change: Bug fix The diffusers example job can exhaust its 45-minute job budget while tests are still progressing. On the same commit, an [initial attempt timed out](https://github.com/NVIDIA/Model-Optimizer/actions/runs/36456149060/job/109049102711), while a [retry passed all 47 tests in 44m50s](https://github.com/NVIDIA/Model-Optimizer/actions/runs/36456149060/job/109115221419), leaving only 10 seconds of headroom. Increase the diffusers timeout to 60 minutes in `.github/workflows/example_tests.yml` to accommodate the workload and observed runtime variation. The other ONNX matrix entries retain their 45-minute timeout. This applies to both PR and nightly diffusers jobs. ### Usage N/A — CI configuration change. ### Testing - `pre-commit run --files .github/workflows/example_tests.yml` — passed all applicable hooks. - Parsed the caller and reusable workflow with `yaml.safe_load` and inspected the timeout input and consumer. - `git diff --check` — passed; reviewed the one-line diff. - GPU tests were not rerun locally. CI validation of the increased timeout is pending. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no copied code or new dependencies. - Did you write any new necessary tests?: N/A — one-line CI configuration change; validation described above. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — CI-only change. - Did you get Claude approval on this PR?: ❌ Not run; opening as a draft. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated automated checks to allow ONNX example tests up to 60 minutes. Other example tests retain their existing 45-minute limit. This change affects test execution time limits only; it does not change application features or behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> |
||
|
|
834c90d7a1 |
Fix NVFP4 fake quant zeroing blocks with small scales (#2549)
### What does this PR do? Type of change: Bug fix The dynamic NVFP4 Triton kernel (`fp4_fake_quant_block`, used on compute >= 8.9) and the Conv3D implicit-GEMM CUDA kernels replaced any FP8 block scale below 1e-5 with 1.0, so every block whose max |x| was below ~6e-5 was zeroed. The static Triton kernel, the CUDA extension fallback and NVFP4 export have no such floor; the floor was only guarding division by zero. The Triton kernel had its own copy of the scale/round code instead of the shared `nvfp4_scalar_quant`. - Triton: use the shared `nvfp4_scalar_quant` (zero only on a zero block scale). - Conv3D CUDA (fused kernel and standalone `fp4_fake_quant`): same rule. - Both: a zero, inf or NaN global amax uses a unit block scale, like the CUDA extension (the conv kernels returned NaN for inf/NaN before). Blocks with scale >= 1e-5 are unchanged. - New tests: power-of-two scaling of input and global amax (2^-10, 2^-20) scales the output by the same factor (Triton, standalone conv FP4, fused conv3d); invalid global amax gives unit-scale rounding. The conv test's Python reference drops the floor. - B200: 505 passed / 31 skipped (`tests/gpu/torch/quantization` NVFP4/FP4 files) and 180 passed (conv implicit GEMM + attention P-QDQ). - Negative control on `main`: the small-input tests fail on both kernels; conv also fails for inf/NaN global amax. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (only blocks with scale < 1e-5 or an inf/NaN global amax change, from zeros/NaN to correct values) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: ❌ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * FP4 quantization now preserves proportional output when inputs and valid global scales are reduced together, including at small scales. * Zero, infinite, or NaN global scales use a safe fallback, preserving inputs already representable in FP4. * Small positive block scales are no longer discarded by an absolute scale threshold; subnormal scale handling is covered across quantization paths. * **Tests** * Added coverage for scale consistency across input types and block sizes, invalid global scales, and subnormal FP8 block scales. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shiyang Chen <shiychen@nvidia.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
94d7272e8b |
Improve technique-specific documentation links in the main README (OM… (#2584)
### What does this PR do? Improve technique-specific documentation links in the main README (OMNIML-5944): The Techniques table in ./README.md contains links that do not lead directly to the relevant documentation: - The Docs link for Quantization Aware Training / Distillation points to the general quantization documentation. It should point to ./guides/quantization_aware_training.html. - The Megatron Bridge links for Post Training Quantization, Quantization Aware Training / Distillation, Pruning, and Distillation all point to the same folder. Users must then find the relevant section themselves. Update the Megatron Bridge links to target the corresponding sections in ./examples/megatron_bridge/README.md: - In ./README.md’s Techniques table, rename Docs to Getting started and reverse the order of the two link columns: currently Examples → Docs, proposed Getting started → Examples. ### Usage just see the table of techniques in the main readme.me ### Testing tested manually ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - Did you write any new necessary tests?: ❌ , only manual testing, only docs changes <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the Techniques table with “Getting started” and “Examples” columns, replacing the former “Examples” and “Docs” columns. * Added technique-specific guide links, including a link to the README for pruning. * Updated some example destinations and anchors to point to relevant workflow sections. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
8990897c56 |
Increase Unit Test timeout
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
c2aaa44f60 |
[Speculative Decoding] DFlash2 draft variant (grouped sublayer convolution + candidate selector) (#2216)
### What does this PR do? Type of change: new feature Adds **DFlash2** ([blog](https://inco.ai/blog/dflash2/)) as a draft variant of the existing DFlash mode, selected with `dflash_architecture_config.projector_type="dflash2"` alongside `domino`, `dspark` and `lilicorr`. DFlash2 keeps DFlash's one-pass parallel backbone and adds two components that recover the acceptance a purely parallel draft loses: - **Grouped dynamic depthwise convolution** around every attention and MLP sublayer, giving each block position a view of its predecessors *inside* the block. Taps do not cross the block boundary, so the draft stays one forward pass. - **Low-rank candidate selector** scoring transitions between adjacent positions' top-k candidates, so serving walks one coherent path instead of taking an independent argmax per position. Both start as exact no-ops — the convolution's `base_kernel` is an identity and `kernel_projection` is zeroed; the selector's `successor_codebook` is zeroed — so a freshly built DFlash2 draft *is* its DFlash backbone, and enabling the variant is an extension rather than a perturbation. This matches the reference implementation ([SpecForge#772](https://github.com/sgl-project/SpecForge/pull/772), merged) and the way `modeling_lilicorr` already installs this same convolution class. **This unblocks a recipe already shipped on `main`.** `modeling_lilicorr._install_sublayer_convs` imports `DFlashGroupedConv` from `modeling_dflash2`, so `modelopt_recipes/general/speculative_decoding/lilicorr_conv.yaml` raises at model build today and its CHANGELOG entry documents a feature that cannot run. Landing this makes it runnable. Module and parameter names match the SGLang/vLLM `DFlash2DraftModel` loaders. Verified against the released `z-lab/Qwen3.8-27B-DFlash2` checkpoint: **81 tensors, 21 name patterns, zero difference in either direction**. The serving side, [vllm-project/vllm#52816](https://github.com/vllm-project/vllm/pull/52816), has since merged (`b389ac294`) with no change to the checkpoint contract. ### Usage ```bash python examples/speculative_decoding/main.py \ --config modelopt_recipes/general/speculative_decoding/dflash2.yaml \ model.model_name_or_path=Qwen/Qwen3-8B \ data.data_path=<corpus>.jsonl \ training.output_dir=<out> ``` ```yaml # modelopt_recipes/general/speculative_decoding/dflash2.yaml dflash: dflash_selector_loss_alpha: 1.0 # weight of the candidate-selector CE term dflash_architecture_config: projector_type: dflash2 conv_kernel_size: 2 # taps; must not exceed the block size conv_group_size: 16 # must divide hidden_size selector_rank: 256 selector_top_k: 16 ``` ### Testing <img width="2000" height="1320" alt="image" src="https://github.com/user-attachments/assets/869004d1-92c6-41a4-a03d-fd824a06255c" /> **Unit** — 25 CPU tests in `tests/unit/torch/speculative/plugins/test_hf_dflash2.py`; the full `tests/unit/torch/speculative/` suite passes with no regressions. The ones worth keeping pin invariants that a decreasing loss does not catch: - the convolution is an exact identity on the **default** construction, and its taps stay inside the block while a position still sees its predecessors; - the block-offset contract shared by the training objective and `CandidateSelector.greedy_path` — a misaligned objective still converges; - the export fields the vLLM loader requires, including the top-level `block_size` that `DFlash2Exporter` derives the nested copy from; - which selector factors receive gradient on the first step. `successor_codebook` starts at zero, so `predecessor_codebook` and `hidden_projection` take one step to begin moving. That is a warm start, not a dead branch, and both sides are asserted. **End-to-end** — trained on Qwen3-8B against a plain DFlash control with every other argument identical (plot above). Monotonic convergence, no NaN/divergence, no DDP unused-parameter issues. Note the losses are **not comparable** across arms: DFlash2's includes the selector CE term. **Serving (vLLM)** — the exported drafter loads and drafts under the merged DFlash2 path (`RESOLVED draft architectures: ['DFlash2DraftModel']`). Two notes for anyone reproducing: vLLM sizes the convolution from `1 + num_speculative_tokens` at runtime rather than from the checkpoint, so a `block_size=16` drafter is only correct at `num_speculative_tokens=15`; and at that value the upstream path currently hits an illegal memory access in `_cache_draft_logits` ([vllm#55279](https://github.com/vllm-project/vllm/issues/55279)), independent of which checkpoint is used. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — additive. New `projector_type`, its own registry and exporter, one new config field; DFlash / Domino / DSpark / LiLiCorr numerics and `state_dict` contents are untouched. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — `modeling_dflash2.py` is adapted from [SpecForge#772](https://github.com/sgl-project/SpecForge/pull/772) and carries its MIT notice. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — under `0.48.0`. - Did you get Claude approval on this PR?: ✅ — run on 2026-08-20; all review threads addressed and resolved. ### Additional Information Rebased onto current `main`. Two commits from the original branch were dropped because [#2342](https://github.com/NVIDIA/Model-Optimizer/pull/2342) landed them first, with authorship preserved: the no-op sublayer seam in `modeling_dflash.py`, and the `rope_theta`/`rope_parameters` fix — `main`'s version of the latter is stricter, so this PR no longer touches `hf_dflash.py` at all. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added DFlash2 speculative decoding with grouped dynamic convolutions and low-rank candidate selection. * Added configurable selector-loss weighting, including an option to disable it. * Added DFlash2 model conversion and export support. * Added checkpoints compatible with SGLang and vLLM DFlash2 serving. * **Documentation** * Added training recipes and a Qwen3-8B online DFlash2 training configuration. * **Tests** * Added coverage for conversion, training, metrics, gradients, and export compatibility. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
be74001256 |
Add concise AgentX benchmark skill (#2573)
### What does this PR do? Type of change: Documentation. Adds a concise AgentX skill covering harness installation, automatic dataset downloads, benchmark execution, and result reporting. Reuses the existing deployment skill and adds a Claude discovery link. ### Usage Use run-agentx to benchmark my deployed model with a concurrency sweep. ### Testing • Skill structure and metadata validation passed. • Shell syntax checks passed. • Benchmark arguments parsed and produced a valid configuration using the pinned harness. • All applicable pre-commit checks passed. • No GPU benchmark was launched. ### Before your PR is "Ready for review" Contributor guidelines and security practices were reviewed. The commit is signed and signed off. • Is this change backward compatible?: ✅ • If you copied code from other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md?: N/A. No copied implementation or project dependency changes. • Did you write any new necessary tests?: N/A. Documentation changes were validated as described above. • Did you update Changelog?: N/A. Skill documentation only. • Did you get Claude approval on this PR?: ❌ Not run. ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added guidance for configuring and running SemiAnalysis AgentX serving benchmarks, including endpoint, model, tokenizer, context limits, caching, and dataset setup. * Documented using a pinned benchmark harness in a separate client environment and running each concurrency level in a fresh artifact directory with a fixed seed. * Expanded reporting guidance to cover overlapping requests, cache and preemption metrics, errors, unfinished requests, warmup failures, and submission validity. * Clarified that missing server-reported usage makes cache-hit data unknown, invalid or missing submission validity should be flagged, and smoke runs are not benchmark results. * Directed AgentX agentic workloads from the optional AIPerf guidance to the AgentX benchmark instructions. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Shiyang Chen <shiychen@nvidia.com> |
||
|
|
1d392999b4 |
[OMNIML-5570] 2/2 Compose GEMM and KV-cache AutoQuant workflows (#2273)
### What does this PR do? Type of change: new feature. Follow-up to merged #2272. Adds composition of existing GEMM quantization with KV-cache AutoQuantize: - fixed FP8 GEMM PTQ followed by mixed-KV AutoQuantize; - gradient-based NVFP4/FP8 GEMM AutoQuantize followed by independent mixed-KV AutoQuantize; - an optional `kv_auto_quantize` recipe stage with independent method, constraints, candidates, and checkpoint path; - ordered `hf_ptq.py` orchestration that keeps selected weight/activation QDQ active while its calibration state remains frozen during KV candidate calibration; - fail-closed validation when a preceding stage leaves actual K/V quantizers enabled; and - unified export of a uniform-weight or mixed-weight checkpoint together with the selected per-layer KV map. The KV search still uses the public `mtq.auto_quantize(..., constraints={"cost_model": "kv_cache", ...})` API from #2272. On a converted model, the API preserves existing non-KV quantizers and requires K/V to be disabled before search. Fresh-model behavior is unchanged and starts from a deny-all quantizer baseline. #### Why a follow-up field instead of a generic stage list? This PR deliberately supports the two composition forms required by `hf_ptq.py` without replacing the stable recipe schema. Existing recipes already express a fixed `quantize` baseline plus one primary `auto_quantize` search. A generic ordered `stages` list would require a broader recipe/API migration, indexed checkpoint semantics, and compatibility rules for arbitrary stage sequences. There is not yet a demonstrated third search stage that justifies that surface-area change. The two searches are not combined inside `mtq.auto_quantize`: each invocation owns one search domain, constraint model, scoring method, and resumable checkpoint. Their ordering and independent checkpoint paths are orchestration concerns, while candidate calibration, scoring, selection, and state application remain in the shared public API. A general stage pipeline can be considered separately if more than this one optional KV follow-up is needed. Both solvers and scoring protocols are unchanged. The KV checkpoint compatibility signature additionally fingerprints the preceding quantizer configuration and calibrated state. Unsupported uniform-weight plus mixed-KV exports record `kv_cache_deployment_supported: false` in both ModelOpt and converted HF metadata. ### Usage Fixed FP8 GEMM PTQ followed by KV AutoQuantize: ```bash python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path Qwen/Qwen3-8B \ --recipe general/auto_quantize/fp8_ptq_then_kv_fp8_nvfp4_cast_kl_div_at_5p4bits \ --kv_auto_quantize_checkpoint /path/to/kv_autoquant.pth \ --export_path /path/to/qwen3-8b-fp8-and-mixed-kv ``` Weight AutoQuantize followed by KV AutoQuantize: ```bash python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path Qwen/Qwen3-8B \ --recipe general/auto_quantize/nvfp4_fp8_gradient_then_kv_fp8_nvfp4_cast_kl_div_at_5p4bits \ --auto_quantize_checkpoint /path/to/weight_autoquant.pth \ --kv_auto_quantize_checkpoint /path/to/kv_autoquant.pth \ --export_path /path/to/qwen3-8b-autoquant-and-mixed-kv ``` KV checkpoint resume requires identical preceding non-K/V quantizer configuration and calibrated state. If rerunning the preceding stage changes that state, use a new KV checkpoint path to recompute sensitivities; configuration identity alone is insufficient to reuse the scores safely. ### Testing - Latest changed-area validation: 126 tests passed across `hf_ptq.py` orchestration, KV checkpoint compatibility, export metadata, and HF configuration conversion. - A broader local run had 604 passes, one skip, and six failures: two socket-binding failures under the sandbox and four local Transformers API incompatibilities. This is not a full-suite pass. - The fixed-PTQ→KV recipe executes end to end on a tiny offline Qwen fixture. - Public API coverage verifies that composed KV search preserves preceding weight quantization and rejects enabled K/V state. - Changed-file pre-commit hooks passed; the isolated recipe validator also passed after dependency bootstrap. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ (0.48.0 composition feature and KV checkpoint flag deprecation) - Did you get Claude approval on this PR?: ❌ ### Additional Information - This follow-up targets `main`, which contains merged #2272. - `--auto_quantize_checkpoint` and `--kv_auto_quantize_checkpoint` are intentionally separate because KV sensitivities depend on the preceding GEMM state. - Uniform-weight plus mixed-KV exports are for artifact inspection until the runtime's uniform-weight ModelOpt configuration consumes `kv_cache_quantized_layers`. Export emits an actionable warning and records `kv_cache_deployment_supported: false`; this marker does not itself add runtime support. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added staged post-training quantization workflows for weights and KV caches, including dedicated KV-cache checkpoints. - Added FP8/NVFP4 recipes with configurable bit constraints and scoring. - KV-cache quantization now supports pre-quantized models. - **Bug Fixes** - Mixed weight and KV-cache quantization now exports with a warning instead of failing. - Improved validation and checkpoint compatibility for staged configurations. - Added safeguards for configurations without enabled weight quantizers. - **Documentation** - Clarified staged KV-cache workflows, checkpoint options, configuration behavior, and unsupported deployment combinations. - Documented deprecated legacy quantization options and their replacement behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
de2e810219 |
[OMNIML-5899] Add Q8_0 CUDA packing kernel (#2515)
### What does this PR do? Type of change: new feature Adds the Q8_0 CUDA packing layer for the three-PR Q8_0 series: - packs 32-value blocks into the 34-byte GGML-compatible payload; - exposes `q8_0_pack` through the shared GGML extension; - validates device, dtype, row alignment, and CUDA launch bounds; - tests byte layout, accepted dtypes, non-finite handling, float64 narrowing, FP16 scale boundaries, reconstruction error, and invalid inputs. This PR contains only the kernel and extension boundary. The codec/backend and export/recipe layers remain in the later PRs. ### Usage ```python from modelopt.torch.quantization.extensions import get_cuda_ext_ggml extension = get_cuda_ext_ggml(raise_if_failed=True) packed = extension.q8_0_pack(weight) ``` ### Testing - Combined #2515 -> #2516 -> #2517 stack: 108 focused CPU codec, backend, export, and recipe tests passed. - Ruff, formatting, and whitespace checks passed for the changed Python test. - The shared-extension Q8_0 and existing IQ tests require CUDA CI; the latest run is pending on the current PR head. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: yes - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: yes; no dependency was added and no implementation code was copied - Did you write any new necessary tests?: yes - Did you update `CHANGELOG.rst`?: N/A; the user-facing entry is in #2517 - Did you get Claude approval on this PR?: pending ### Provenance The CUDA encoder was independently written for ModelOpt. It implements the packed-format contract and scalar quantization formula documented by the pinned llama.cpp definitions: - [Q8_0 packed structure](https://github.com/ggml-org/llama.cpp/blob/9b05354ec6fb58b4e665e9a39ebc40285c015638/ggml/src/ggml-common.h) - [Q8_0 scalar reference formula](https://github.com/ggml-org/llama.cpp/blob/9b05354ec6fb58b4e665e9a39ebc40285c015638/ggml/src/ggml-quants.c) No llama.cpp implementation code is incorporated into this CUDA source. Human code-owner confirmation of this provenance and attribution is requested before merge. ### Related PRs Merge order: 1. **Kernel - this PR** 2. [#2516 - Q8_0 quantization codec and backend](https://github.com/NVIDIA/Model-Optimizer/pull/2516) 3. [#2517 - Q8_0 checkpoint export and recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2517) All three PRs target `main`. --------- Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com> |
||
|
|
0058a15537 |
[2/2] Track every Megatron-Bridge script with MLflow (#2514)
### What does this PR do? Type of change: new feature **[2/2] of a split. Based on #2544 — merge that first; this PR's diff is only the Megatron-Bridge half.** #2477 added MLflow tracking to `examples/megatron_bridge/quantize.py`. It was one of five scripts in that directory that write a checkpoint; the other four recorded nothing, so the provenance chain stopped at the PTQ checkpoint and a deployed model could not be traced back to the run that produced it. All five now take the same `--mlflow` / `--mlflow_experiment` / `--mlflow_run_name` flags, and **each declares what it records as a `Tool` beside its own flags** — the shared `mlflow_utils.py` knows none of them: | Script | Records | | --- | --- | | `prune_minitron.py` | command, arguments, log, `prune_score` metric, pointer | | `quantize.py` (#2477, moved onto the shared `Tool` in #2544) | + resolved recipe, quantizer summary | | `distill.py` | + Megatron-Bridge's per-iteration metrics and resolved config | | `export_quantized_megatron_to_hf.py` | command, arguments, log, pointer | | `export_distilled_megatron_to_hf.py` | same, one pointer per exported checkpoint | Each writes `.experiment.json` into the checkpoint it produced, and each tags what it consumed, so `prune → quantize → distill → export` is walkable both from disk and by tag query on the server. **`distill.py` opens the run and Megatron-Bridge joins it.** Its `LoggerConfig` records per-iteration metrics and the full resolved config — which a wrapper around `main()` cannot see — but nothing of `distill.py`'s own arguments and no invocation. Megatron-Bridge takes `mlflow.active_run()` when one exists, applies the tags and logs into it, so `distill_run()` opens the run on the rank Megatron-Bridge looks at (the **last** one) and the two share it. Its early exit is handled explicitly: `train()` leaves through `sys.exit(0)` on `--exit_interval`, which a blanket handler would record as `FAILED`. **The library pieces that exist for that shared run land here with their first caller**, rather than in [1/2] where they would have none: `split_tracking_credentials`, so a URI handed to something which *records* it carries no credential; `log_active_run_experiment_json`, for pointing a checkpoint at a run this process did not open; and `MlflowRunLogger._reattach`, because a co-owner can end the run first — Megatron-Bridge does, as `KILLED`, when SIGTERM arrives mid-training. Two of Megatron-Bridge's defaults are deliberately not inherited: **checkpoint artifact upload stays off** unless `--mlflow_log_checkpoints` (it pushes the whole checkpoint over HTTP after every save), and **an untracked run passes no `mlflow_*` fields at all**, since they landed in Megatron-Bridge 0.6 and sending them unconditionally would break an untracked run on an older one. ### Usage ```bash # Any of the five, same flags: torchrun --nproc_per_node 8 prune_minitron.py ... --mlflow https://<server>/ torchrun --nproc_per_node 8 quantize.py ... --mlflow https://<server>/ torchrun --nproc_per_node 8 distill.py ... --mlflow https://<server>/ torchrun --nproc_per_node 8 export_quantized_megatron_to_hf.py ... --mlflow https://<server>/ # Each checkpoint names the run that wrote it: cat /output/qad/checkpoints/.experiment.json ``` Experiments default to `$USER/megatron_bridge_{prune,quantize,distill,export,distill_export}/<model basename>-<variant>`. ### Testing - Real runs on a toy Qwen3 in one MLflow experiment covering all five Megatron-Bridge scripts and `hf_ptq` — prune, quantize, QAD distillation, quantized export, BF16 distillation, distilled export, HF PTQ — each closing `FINISHED` with the invocation, its arguments as params, its log, and a matching `.experiment.json` on disk. The chain tags line up: each stage's `source_checkpoint_path` is the previous stage's `checkpoint_path`. - `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08`, the only lane that runs it: **76 passed**. Plus the three suites from #2544: **195 pass**. - `pre-commit run --files <changed>`: all hooks pass. - Each fix from the review rounds has a test that fails with the fix reverted: the resumed run, the foreign active run, the percent-decoded credential, the credential that cannot be moved, the rank-dependent `LoggerConfig`, the exit-callback guard, and the `iter_*` join. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: several rounds; re-requested on this head. ### Additional Information Split from a single ~1150-line PR at review's request; #2544 carries the library consolidation this builds on, and this branch is based on it. Earlier review threads here show as outdated after the rebases — they are all resolved and their fixes are in this branch. One known gap, stated in the README rather than implied: `distill.py --hf_export_path` writes a second HuggingFace checkpoint from rank 0, which is not the rank that owns the run, so it carries no pointer yet. For the same reason the uploaded `logs/distill.log` holds the last rank's output — `print_rank_0` keeps the script's own lines on rank 0 — which the README now says outright; carrying rank 0's log into a run owned by another rank needs cross-rank upload and is a follow-up. Two defects found on shared-run paths during review, both verified against the installed Megatron-Bridge 0.6 rather than its docs. Megatron-Bridge ends the run it shares with `distill.py` as `KILLED` from its SIGTERM handler (`train.py:1413`) and then leaves through `sys.exit()` (`train.py:805`), i.e. before `distill_run`'s `finally` — and MLflow's fluent calls resolve their target by *opening* a run when none is active, so a preempted distillation's log and metrics went to a second, empty run and its `KILLED` status was overwritten. Separately, an unreachable server disabled our logger but `logger_kwargs` still handed Megatron-Bridge the same URI, and `state.py` calls `set_experiment` unguarded from inside the training loop — so a best-effort `$MLFLOW_TRACKING_URI` aborted the training instead of degrading to untracked. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
767ef5533e |
[3/5] Add the IQ2_S codec (#2512)
### What does this PR do? Type of change: new feature (not yet user-reachable) **First of two PRs adding IQ2_S**, the widest of the GGML IQ formats at one and two bits (2.5625 bits per weight). This one lands the **PyTorch codec**: the encoder, the decoder and the 1024-entry codebook. It is deliberately **not registered**, so no quantizer dispatches to it and the `ggml` package does not export it. #2565 adds the CUDA encoder, registers the format and adds its recipe. ### What's distinctive about it **IQ2_S is the one format llama.cpp's own tooling gives no head start on**, so the search is written against the GGML layout directly. The interesting difference from IQ2_XS and IQ2_XXS is sign handling. IQ2_S stores a **full 8-bit sign mask** per group rather than a 7-bit parity-coded index. The encoder therefore takes the input signs as they are instead of flipping the weakest element to fix parity, and the search compares magnitudes directly, which is simpler than its siblings. ### Why the codec lands before the kernel The CUDA encoder's tests use this codec as their reference. They compare against the PyTorch encoder byte for byte and draw the grid and scale predictor from it. So the kernel cannot be tested before the codec exists, and it follows in #2565 together with the registration. Every registered format therefore keeps a CUDA encoder. ### Test changes that make the split possible A codec can now land before it is registered, so two test contracts in `test_iq_formats.py` are stated precisely: - The two tests that go through `TensorQuantizer` (pass-through gradient, error falls with bit width) iterate `IQ_FORMAT_REGISTRY`. Every other battery test calls the codec directly and covers IQ2_S here. - The coverage check now asserts `set(IQ_FORMAT_REGISTRY) <= set(FORMATS)` instead of equality. That is what its docstring already said: a registered format must be listed, or it escapes the contract. - `test_registry_lists_every_exported_encoder` is unchanged, and it is why this PR leaves the package exports alone: an exported encoder must be registered. The error-by-bit-width failure message also labels errors by the order they were measured in; it previously zipped them with alphabetical names. ### Testing **The decoder is validated against llama.cpp's own output, not just round-tripped:** ``` IQ2_S: 9 tensors, 2,355,200 blocks → 0 mismatched, max|diff| 0.0 ``` The new codebook matches the `ggml-common.h` table entry for entry. Blocks from `unsloth/Qwen3.8-27B-GGUF` ship as conformance vectors, so CI keeps checking bytes we did not produce. - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py`, `tests/unit/recipe/test_presets.py`: **121 passed**, 14 of them IQ2_S codec cases, including the llama.cpp conformance check - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **35 passed**, unchanged by this PR ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ The new codebook is a GGML table, carried in `codebooks.py` with the source revision recorded. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A. Nothing is user-reachable yet; #2565 carries the entry. - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: #2511 (IQ2_XXS, merged) → #2525 (format registry, merged) → **this** → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added GGML-compatible IQ2_S quantization and dequantization support, including access to its magnitude grid. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
57f929e358 |
Bound the stale-capture warning to once per capture and name its cause (#2492)
### What does this PR do?
Type of change: Bug fix
The activation-capture hooks latch `_intermediate_output` on every
forward, and only
`DistillationModel.compute_kd_loss()` ever clears it. Any training loop
that does not call
`compute_kd_loss()` therefore warns on every forward after the first,
once per hooked module:
```
plain transformers Trainer, 16 micro-batches
before: 30 warnings (Teacher 15 + Student 15)
after: 2 warnings (Teacher 1 + Student 1)
```
The count scales with forwards, not optimizer steps —
`gradient_accumulation_steps` of 1, 4 and
16 all produced 30 warnings for the same 16 forwards — so the
accumulator the report blamed is
not involved. A plain `transformers.Trainer` reproduces it because its
default `compute_loss`
reads the student's own CE from `outputs.loss` and never calls
`compute_kd_loss()`; the
warning is a symptom of that, but neither message said so. The student
message attributed the
re-forward to Activation Checkpointing only, and the teacher message
called the situation
"expected" while still raising `UserWarning`.
This PR:
- warns once per captured activation via
`_warn_once_about_stale_output`, re-armed wherever a
capture is cleared, so a consumed capture that is followed by another
unconsumed one is still
reported (Activation Checkpointing re-runs forwards and relies on that);
- states the cause and the consequence in both messages:
`compute_kd_loss()` did not run since
the previous forward, so no KD loss is applied;
- moves the three `_intermediate_output = None` reset sites onto one
`_clear_captured_output`
helper so the capture and its warning state cannot drift apart, and has
the layerwise teacher
hook share both helpers rather than keeping a second copy of the
warning.
### Usage
No API change. A loop that applies KD should call `compute_kd_loss()`
once per forward:
```python
class KDLossTrainer(Trainer):
def compute_loss(self, model, inputs, return_outputs=False, **kwargs):
outputs = model(**inputs)
loss = model.compute_kd_loss(student_loss=outputs.loss) # consumes the captures
return (loss, outputs) if return_outputs else loss
```
### Testing
`tests/unit/torch/distill/test_distill.py`:
- `test_duplicate_fwd_hook_call` now pins the bound under
`warnings.simplefilter("always")`
(3 forwards -> exactly 2 warnings) instead of relying on the
interpreter's per-message
deduplication, and asserts the message names `compute_kd_loss`;
- `test_stale_output_warning_rearms_after_consuming` covers the re-arm:
warn, consume, warn
again -> 4 warnings. This one passes before and after the change; it
guards against a future
"warn once ever" simplification silently muting the Activation
Checkpointing case.
`tests/unit/torch/distill/test_layerwise.py`:
- `test_layerwise_stale_output_warning_is_bounded` covers the layerwise
hooks (3 forwards ->
exactly 2 warnings).
The first and third tests fail on `main` (`assert 4 == 2`), verified in
a worktree of `main`
carrying these test files.
Measured with a `transformers.Trainer` over a 16-micro-batch run,
counting
`"already has an intermediate output stored"`:
| setup | before | after |
|---|---|---|
| plain `Trainer`, `gradient_accumulation_steps=1` | 30 | 2 |
| plain `Trainer`, `gradient_accumulation_steps=4` | 30 | 2 |
| plain `Trainer`, `gradient_accumulation_steps=16` | 30 | 2 |
| `Trainer` calling `compute_kd_loss()`, accum=4 | 0 | 0 |
```
$ python -m pytest tests/unit/torch/distill -q
34 passed
$ python -m pytest tests/unit/torch -q --ignore=tests/unit/torch/deploy
1 failed, 2673 passed, 18 skipped in 217.17s
```
The single failure is
`tests/unit/torch/quantization/plugins/test_huggingface.py::test_dbrx`,
which fails identically on `main` in this environment because
`transformers 5.17` is outside the
`transformers>=4.57,<5.15` range pinned in `pyproject.toml`. It is
unrelated to this change.
`pre-commit run --files <changed files>` passes every hook (ruff check,
ruff format, mypy,
bandit, insert-license, large files, line endings).
Note on severity, since it affects how the issue reads: Python already
deduplicates a warning
by (message, location), so with the default filters this shows up twice
rather than 30 times.
The repetition is user-visible under `-W always`,
`PYTHONWARNINGS=always`, pytest, or one
warning registry per DDP rank, and the count is what scales with epoch
length. The behavioural
fix that matters for all filters is the latch.
### Before your PR is "*Ready for review*"
Is this change backward compatible?: ✅
If you copied code from any other sources or added a new PIP dependency,
did you follow guidance in CONTRIBUTING.md: N/A
Did you write any new necessary tests?: ✅
Did you update Changelog?: N/A
Did you get Claude approval on this PR?: N/A
### Additional Information
Fixes #2487.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Bug Fixes**
- Improved handling of captured activations during distillation,
including clearing stale outputs and resetting warning behavior after
outputs are consumed.
- Stale-output warnings now appear once per affected activation and
clarify when knowledge-distillation loss was not applied.
- Warnings distinguish expected cases involving activation checkpointing
or teacher evaluation.
- **Tests**
- Expanded coverage for warning counts, warning reset behavior, and
stale-output handling in standard and layerwise distillation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Edwardssss <ed_129@qq.com>
|
||
|
|
4eb86524f0 |
[1/2] One MLflow tracking core behind a Tool record (#2544)
### What does this PR do? Type of change: refactor (no functional change) **[1/2] of a split. Merge this first; #2514 is [2/2] and is based on this branch.** Three example scripts had each reimplemented the same MLflow wiring: the flags, the `$USER/<tool>/<model>-<variant>` experiment convention, the params/tags/artifacts a run uploads, and the open/close dance with its status. The copies had already drifted — only `hf_ptq` wrote a provenance pointer, only `vllm_serve` republished the resolved URI — and every new tracked script meant another copy. What a script records is now one declarative `Tool` record, **declared in the script itself, beside the flags it reads**: ```python # examples/megatron_bridge/quantize.py QUANTIZE = Tool( name="megatron_bridge_quantize", tracks="Track this run on an MLflow server, uploading the command, the resolved recipe, ...", variant_help="recipe name, or --quant_cfg if no --recipe", variant=lambda args: Path(args.recipe).stem if args.recipe else (args.quant_cfg or "none"), model=lambda args: args.hf_model_name_or_path, checkpoint=lambda args: args.export_megatron_path, texts=lambda args: resolved_recipe_texts(args.recipe), outputs=lambda args: {"summary/quant_summary.txt": Path(args.export_megatron_path) / ".quant_summary.txt"}, ) ``` `tracked_run` takes that record and runs the whole thing, so a script adds tracking in three lines: `add_mlflow_args(parser, TOOL)`, `resolve_mlflow_args(args, parser, TOOL)`, and `with mlflow_run(args, TOOL):`. The shared module knows no script's flags. `examples/hf_ptq`, `examples/vllm_serve` and `examples/megatron_bridge/quantize.py` move onto it. Three helpers fall away as redundant (`track_run`, `checkpoint_run_tags`, and `hf_ptq`'s two flag pass-throughs). ### Usage No user-facing change. The flags, their spellings and the experiment naming are exactly as before; a script author now writes a `Tool` instead of four functions. ### Testing - `tests/unit/torch/utils/test_mlflow.py`, `tests/examples/hf_ptq/test_hf_ptq_args.py`, `tests/examples/vllm_serve/test_vllm_mlflow_utils.py` — **179 pass**. - `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08` (the only lane that runs it), which drives `quantize.py` for real: **34 passed**, locally and in this PR's `megatron` lane. - `pre-commit run --files <changed>`: all hooks pass. - The four suites shared four copies of a stand-in for the `mlflow` module, which had drifted — one recorded artifacts as a list, another as a dict, a third made `log_artifact` a no-op, so a test asserting on an upload asserted nothing. They now share one `tests/_test_utils/mlflow.py`, which also emulates the fluent API's habit of opening a run when none is active. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `track_run` and `checkpoint_run_tags` are removed, but neither shipped in a release (0.47.0's `__all__` is `MlflowRunLogger`, `command_text`, `current_user`, `default_experiment_name`, `validate_tracking_uri`, all unchanged here). Three deliberate behaviour changes, each in shared code and each tested: - The `source_checkpoint_path` tag resolves to an absolute path where it recorded the raw argument, which a chain of runs needs to join on the pair. `run_tags` is shared, so this applies to every script that writes the tag — `hf_ptq` **and** `megatron_bridge/quantize.py`, for a local `--hf_model_name_or_path`. A source that names no directory, such as a Hub `org/name` id, is still recorded as given. - `MlflowRunLogger.track()` — which *did* ship in 0.47.0 — records a block ending in `SystemExit(0)` as `FINISHED` where it recorded `FAILED`, since a script that ends by calling `sys.exit()` rather than returning has still finished. - `.experiment.json`'s `tracking_uri` and the `run_url` built from it drop a trailing `/` from the tracking URI, so the link is `https://host/#/...` rather than `https://host//#/...`. Only reachable by constructing `MlflowRunLogger` directly; every CLI path already stripped the slash in `resolve_tracking_uri`. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — no user-visible change; the entry is in [2/2]. - Did you get Claude approval on this PR?: several rounds; re-requested on this head. ### Additional Information Split out of #2514. This half is the enabling refactor with no behaviour change; #2514 is the feature it unlocks and is based on this branch. At ~605 changed lines of core logic it is over the ~500 guideline; the owner accepted a two-PR split rather than three, and everything #2514 alone consumes — `split_tracking_credentials`, `log_active_run_experiment_json`, `MlflowRunLogger._reattach` — lands there rather than here. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
80e04b8816 |
docs(eval): align Terminal-Bench 2.1 / SWE-bench / MRCR with upstream configs (#2479)
### What does this PR do? Type of change: documentation (agent skill) + template bug fixes Aligns the `evaluation` skill's three upstream-tracked benchmarks (Terminal-Bench 2.1, SWE-bench Verified, MRCR) with the current `nvidia-eval-factory-benchmarking` configs, and fixes guidance that turned out to be wrong when a full three-benchmark campaign was run with the skill end to end. Rebased on #2499. The first commit is the alignment. The rest address review and a fresh config-generation test: MRCR serving scoped per variant, `limit_samples` canary guidance, the template's serve command passing `--gpu-memory-utilization` (replacing `command:` had silently dropped it, so the 1M golden's 0.95 never reached vLLM), upstream's 128K values, and regression tests for the `++limit` gate and the serve command. The skill text keeps only the rules; the evidence behind them is below. GDPVal is out of scope. It was removed from this skill in #2470, and this PR replaces #2464. **Alignment with upstream** - Sandbox region via `HARBOR_ECS_REGION`. TB2.1's ECR repo name tracks the region; SWE-bench's stays in us-west-2. - One interceptor order for both harbor benchmarks, with `http_pairs_dump` **last**. Upstream is split on its position, which changes only what the dump records, never the score. Interceptor lists replace wholesale on merge, so a leaf must restate the whole chain. - `capture_request_body` goes on the service. A shared block injects an alias-only entry with no `type`. - MLflow tags gain `task_name` and `nemo-evaluator-next-version`. - TB2.1 `max_concurrent`: 50 for nano-class models, 15 for larger ones (all upstream non-nano leaves override it to 15). - `proxy.request_timeout` must always be set explicitly. Otherwise it inherits the model fragment's serving value, which ranges from 3600 to 36000 upstream, and TB2.1 has no benchmark key for it. - MRCR: - `parallelism` is 512, deliberately above server capacity, so `--max-num-seqs` must no longer be derived from it. - `limit_samples` now reaches the gym through a gated `++limit`. - Observability capture is on. - 128K is its own upstream benchmark on the condensed gym schema. **Corrections found by running it** | what the skill said | what actually happens | |---|---| | `username: ${oc.env:USER}` | nel-next only expands `${VAR}` / `${VAR:-default}`. The config passes `--dry-run` and fails at `--submit` with *"remote username contains invalid characters"*. | | MRCR canary via `++limit` edited into `collect_rollout_params` | `limit_samples` is now gated through. Under the 0.2.6 launcher the `-o` path is `++evaluation.nemo_evaluator_config.config.params.limit_samples`. `++config.params…` creates a bogus top-level key. | | the condensed gym schema's bootstrap is in the runtime image | It lives in upstream `configs/models/gym_eval_command.yaml`, which is composed in. A standalone config must carry the `command:` block. | | `mean/prefix_matched ~0.55 is healthy` | That value is calibrated to the 1M golden. A 128K run at `pass@1` ≈ 95 measured ≈ 1.0. The signal is a collapse toward 0. | | NVFP4 MoE `VLLM_*` env vars as reliable knobs | They are build-dependent: one vLLM build logged them as unknown and ignored them. Check the server log once per image. | **Rules the skill lacked** - **MRCR variant.** Pick the largest variant within the checkpoint's trained context. On a 262K-context model, 1M needs `VLLM_ALLOW_LONG_MAX_MODEL_LEN` and measures extrapolation, which a quantization comparison would then entangle with quantization damage. The serving setup follows the variant: 128K serves at the trained context without the override. - **SWE-bench `reasoning_effort`.** openhands-sdk sends `reasoning_effort: high` on every call, and canonical `bench.yaml` doesn't strip it. A server whose accepted set excludes `high` returns HTTP 400 on the first call of every trial, so `pass@1` is 0. The fix is to overwrite it with the server default via `proxy.extra_body`. terminus-2 (TB2.1) and Gym's `simple_agent` (MRCR) never send the key (46/46 and 110/110 requests checked). - **Reasoning toggles** go in `extra_body.chat_template_kwargs`. Don't write out no-op sampling defaults; `top_k` is *not* one (vLLM's default is `-1`). - **Upstream model fragment.** Consult `configs/models/<model>/` when it exists, for serving flags, the thinking toggle and `reasoning_replay.mode`. ### Usage No API change. Regenerating a config from the skill now yields the aligned values: ```yaml # recipes/examples/example_eval_next.yaml services: model: proxy: request_timeout: 3600 # always explicit benchmarks: - max_concurrent: 50 # nano-class; larger models 15 sandbox: region: ${HARBOR_ECS_REGION:-us-east-1} cluster: username: ${USER} # NOT ${oc.env:USER} ``` ### Testing - `python -m pytest plugins/modelopt/skills/ -o addopts=""`: 7/7 pass, including the new `tests/test_example_mrcr.py`. It checks that `example_mrcr.yaml` emits `++limit=N` only when `limit_samples` is set, and that the folded vLLM serve command shell-parses with every flag intact, including `--gpu-memory-utilization`. Each check fails when its defect is reintroduced. - `markdownlint-cli2` on the changed Markdown: 0 errors. Both example YAMLs parse. The full `pre-commit` suite was not run after the squash, because the sandbox could not fetch hook repos. The pre-squash commits passed it. - **Exercised end to end.** Configs built from this skill ran a BF16 campaign for a 262K-context MoE reasoning model on an internal cluster to completion: - MRCR-128K: 1470/1470 rollouts - Terminal-Bench 2.1: 712/712 trials - SWE-bench Verified: 2500/2500 trials Each correction above is a defect that campaign surfaced. - **Differential check.** Terminal-Bench configs generated from the pre- and post-change skill, from the same brief and in isolation, differ on: - interceptor chain - concurrency - `top_k` - region interpolation - MLflow tags Not run: a scored evaluation of this PR itself. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ <!-- Docs + a template fix; no ModelOpt API surface touched. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A <!-- Documentation and config-template values. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- Agent skill docs, not a user-facing ModelOpt feature/breaking change/deprecation. --> - Did you get Claude approval on this PR?: ❌ <!-- Not run. --> ### Additional Information The companion internal `eval-config` change now carries only the internal values: the ECR URLs, the region default and the cluster image notes. It points here for the generic rules. **Known gaps not fixed here**, worth a follow-up: - `references/nel-next.md:136` and `references/launcher-workflow.md:207` derive `--max-num-seqs` from `parallelism / DP`. nel-next has no `parallelism` field (its analogue is `max_concurrent`), and MRCR's 512 is deliberately not a server cap. - `references/launcher-workflow.md:200-201` makes `--max-num-batched-tokens` and `--enable-chunked-prefill` always-include defaults, but `example_eval_next.yaml` omits both. - `references/launcher-workflow.md:27` says `sbatch_comment` belongs under `execution:` and is otherwise inert, yet all three shipped examples put it under `cluster:`. - MoE detection (`--enable-expert-parallel`) is unresolvable from the facts the skill asks for when the model handle has no `-A*B` suffix. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
23355eda90 |
fix(deps): declare httpx, unbreaking partial-install (torch) for every PR (#2547)
## Summary `partial-install (torch)` has been failing on **every** PR since 2026-09-24 — including PRs whose branches predate the breakage — and because it is a *collection* error rather than a test failure, it aborts the entire run: ``` ImportError while importing test module '.../tests/unit/torch/speculative/plugins/test_hf_streaming_dataset.py' E ModuleNotFoundError: No module named 'httpx' collected 2243 items / 1 error / 45 skipped !!!!!!!!!!!!!!!!!!!! Interrupted: 1 error during collection !!!!!!!!!!!!!!!!!!!! ``` `unit-pr-required-check` aggregates it, so nothing currently merges on a fresh run. ## What happened **No code changed.** `modelopt/torch/speculative/plugins/hf_streaming_dataset.py` has imported `httpx` at module scope since #1509 (2026-06-02), and `httpx` has never appeared in `pyproject.toml`. It arrived only transitively: `dev-test` → `timm` → `huggingface_hub` → `httpx`. **huggingface_hub 2.0.0**, published **2026-09-24T12:01:21Z**, replaced `httpx<1,>=0.23.0` with the separate **`httpx2<3,>=2.0.0`** distribution. Different package name, so `httpx` stopped being installed and the chain disappeared. The boundary is exact — every run *created* before that timestamp passes, every one after fails: | PR | run created | result | |---|---|---| | #2536 / #2535 | 09-23 22:02 | pass | | #2500 | 09-23 23:45 | pass — **merged 09-24 20:01 on this stale-green result** | | *hub 1.33.0 (still requires httpx)* | *09-24 09:49* | | | **hub 2.0.0 published** | **09-24 12:01** | ← | | #2539 | 09-24 16:57 | fail | | #2544 | 09-24 18:32 | fail | | #2216 | 09-25 11:58 | fail | #2500 merging afterwards is not a counterexample: GitHub does not re-run checks at merge time, so it merged on a result from ~20 hours earlier. That is also why this went unnoticed. ## The changes ### 1. Declare `httpx` in the `hf` extra `httpx` is not incidental to streaming — it is the only transport: - every fetch is HTTP: `POST /v1/completions` to the vLLM serve plus `GET /meta` and `/desc` against the connector's sidecar, all through `httpx.Client`; - there is no non-HTTP path — the base `StreamingDataset._fetch` is an abstract seam and `EagleVllmStreamingDataset._fetch` is its only implementation; - no other HTTP library appears in the module (`requests` / `urllib` / `aiohttp`: zero hits, and `requests` is not declared either); - even the retry predicate is built from it: `_TRANSIENT_FETCH_ERRORS = (httpx.HTTPError, OSError)`. It belongs in `hf` rather than in the core `dependencies`: the same module needs `transformers.trainer_pt_utils` at module scope, so one extra already gates the whole file, and a core install has no use for an HTTP client. The bound matches the 0.x API the code uses — `httpx` has no 1.0 release, and 2.x is a different distribution. This is the part that stops it recurring. `[hf]` currently gets `httpx` only because `datasets` happens to require it — the same accident with a different supplier, one release away from repeating. ### 2. Acquire `httpx` in the test through the existing skip guard The test file already intends to skip where the extra is absent — it has `pytest.importorskip("transformers")` and a comment explaining why, and `transformers` is absent in this job too. It broke only because `import httpx` sat **five lines above** that guard, where a missing module ends collection instead of skipping one file. ## Verification - With everything installed: **18 passed**, no behaviour change. - The import-order property is checked with an AST walk over the module's top-level statements: no `hf`-extra-only import precedes the first `importorskip` (which is now line 41). - A faithful local reproduction was attempted and abandoned honestly: hiding `httpx` locally also breaks `huggingface_hub` 1.28, which `modelopt.torch.opt.plugins.huggingface` imports, so the local failure is not the CI one. CI is the oracle for that half — this PR's own `partial-install (torch)` run is the check that matters. ## Scope Two files, five lines of declaration and four of test import order. Deliberately not folded into any feature PR: it blocks the whole repo, and burying a repo-wide fix inside unrelated work is how these stay invisible. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Optional Hugging Face installations now include `httpx`, supporting features that require HTTP communication without requiring it for all installations. * **Tests** * Hugging Face streaming dataset tests now skip when `httpx` is unavailable, allowing the remaining test suite to be collected and run without it. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ed7e87953c |
Fix grouped expert quantizer checkpoint replicas (#2500)
### What does this PR do?
Type of change: Bug fix
Fix distributed checkpoint saving for quantized Transformer Engine
grouped MoE experts when tensor parallelism and expert parallelism are
both greater than one.
#### Observed error
With NeMo 26.08 and TP=2, EP=4, ETP=1, quantization and calibration
complete, but the run crashes while saving the distributed checkpoint:
```text
megatron.core.dist_checkpointing.core.CheckpointingException:
Invalid sharding pattern validation.
Invalid access pattern for ShardedTensor(
key='decoder.layers.1.mlp.experts.experts.32.linear_fc1.weight_quantizer._amax',
...
)
```
#### Root cause
`_QuantMegatronTEGroupedLinear.sharded_state_dict` created each expert
quantizer's sharded tensors without passing the expert tensor-parallel
and expert data-parallel process groups to
`make_sharded_tensors_for_checkpoint`. It then manually replaced only
the expert-data-parallel component of `replica_id`.
That manual rewrite is insufficient when tensor and expert parallelism
are both enabled: replica ownership is derived using the wrong
process-group topology, so ranks can publish an inconsistent access
pattern for the same globally indexed expert quantizer key.
Megatron-Core correctly rejects that checkpoint during sharding
validation.
#### Fix
Resolve the expert model-, tensor-, and data-parallel groups from one
`_pg_collection`-aware helper, then pass the expert TP/DP groups to
`make_sharded_tensors_for_checkpoint` for per-expert quantizer buffers.
This avoids mixing process-group sources when models use a non-global
`ProcessGroupCollection` or grouped child modules convert without their
parent MLP, and removes the manual `replica_id` rewrite. Shared,
whole-linear quantizer buffers deliberately retain the default dense
TP/DP replica groups because their keys and offsets carry no expert
identity; using only expert TP/DP groups would make replica IDs collide
across EP ranks.
#### Relationship to PR #2319
[PR #2319](https://github.com/NVIDIA/Model-Optimizer/pull/2319) fixed
two earlier TEGroupedMLP checkpoint problems:
1. Per-expert quantizer state did not retain the same globally unique
expert identity as the grouped-expert weights, so it could not be
redistributed reliably when the EP layout changed.
2. After ModelOpt extra-state restoration, an expert that moved to a
different rank could be missing the `_amax` or `_global_amax`
destination buffer required by the subsequent distributed checkpoint
load.
That PR therefore fixed resharding and restore correctness across
topology changes. Its regression matrix changed one parallel dimension
at a time: EP changed while TP=1, or TP changed while EP=1.
The remaining issue was on the save path when TP and EP were
simultaneously greater than one. Although the expert keys and restore
buffers were correct after #2319, the per-expert quantizer shards were
still constructed without the expert TP/DP process groups and then had
only part of their `replica_id` rewritten manually. Under TP=2, EP=4,
ETP=1, this produced the invalid access pattern rejected by
Megatron-Core before the checkpoint could be saved.
This PR complements #2319 by fixing that process-group/replica mapping
and adding combined TP+EP coverage.
A four-rank save-and-restore regression test covers combined TP=2, EP=2,
and ETP=1.
### Usage
N/A. This fixes checkpoint behavior without changing the public API.
### Testing
- [x] Focused Ruff check and format validation
- [x] Focused mypy validation
- [x] `git diff --check origin/main..HEAD`
- [x] Four-rank grouped-expert checkpoint save/restore test with TP=2,
EP=2, ETP=1: 1 passed on the original fix
- [x] End-to-end NeMo 26.08 quantization and checkpoint save on 8 B200
GPUs with TP=2, EP=4, ETP=1 on the original fix
- [x] Review-response validation on `dbfa1cdb5`: Ruff 0.15.20
format/check, `git diff --check`, and Python compilation
- [x] Final four-rank grouped-expert checkpoint regression on
`c2be0f2c6`: TP=2, EP=2, ETP=1 with both ordinary and overridden
ModelOpt TP state; 2 passed in 459.38s on 4 B200 GPUs with NeMo 26.08
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌
### Additional Information
The end-to-end validation produced a complete eight-shard distributed
checkpoint and exited successfully.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Bug Fixes**
- Fixed checkpoint saving for quantized grouped MoE experts when tensor
and expert parallelism are enabled.
- Improved handling of shared and per-expert quantizer state across
parallel configurations to support more reliable checkpoint save and
restore.
- **Tests**
- Added NVFP4 grouped expert checkpoint round-trip coverage with tensor
and expert parallelism, including configurations with the
tensor-parallel group override enabled and disabled.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
|
||
|
|
63c4b660bd |
Add Aumann-Shapley sensitivity scoring method to auto_quantize (#2183)
[Paper](https://arxiv.org/abs/2607.12266) · [Overview](https://x.com/waterloo_intern/status/2076460984475263401) · [Implementation thread](https://x.com/the_joshua_hill/status/2076427869388255635) Depends on #2231, which provides the shared AutoQuantize backward-scoring infrastructure. Until that PR merges, the focused PR B diff is available [here](https://github.com/joshua-hill/Model-Optimizer/compare/fix/autoquant-scoring-infrastructure...feat/aumann-shapley-autoquant). ### What does this PR do? Type of change: new feature This PR adds `method="aumann_shapley"` to `mtq.auto_quantize`. It is a label-free scoring method that measures how each candidate quantization format affects the model across the path from full precision to quantized. The new method is opt-in; existing `gradient` and `kl_div` behavior is unchanged. For each calibration batch, the method: 1. Runs the baseline model and saves its next-token distribution. 2. Measures each candidate format at a configurable number of points along the quantization path. 3. Uses the KL-divergence gradients at those points to assign a damage contribution to every runtime group and candidate format. 4. Measures the most aggressive candidate configuration once and uses that value to calibrate the per-group damage model. The resulting scores use the existing AutoQuantize linear-program solver. The search can either: - choose the least damaging configuration that meets an `effective_bits` target; or - choose the smallest configuration whose predicted damage stays below `max_predicted_damage`. The selected recipe records `predicted_damage` in mean per-token KL units together with its validity and fit diagnostics. ### Public API `auto_quantize` gains an optional `method_options` dictionary. For `method="aumann_shapley"`, it accepts: | Option | Default | Meaning | |---|---:|---| | `num_path_nodes` | `2` | Number of points used to average gradients along the quantization path. | | `damage_link` | `"coverage"` | How per-group scores combine. `"coverage"` uses `damage = c * (1 - exp(-sum(b)))`; `"additive"` sums the path contributions. | | `max_predicted_damage` | `None` | Replaces the bit target with a maximum predicted mean per-token KL. | Method options are validated before the model is modified. Unknown options and incompatible targets fail early. ### Usage Select a configuration for a target effective bit width: ```python import modelopt.torch.quantization as mtq model, search_state = mtq.auto_quantize( model, constraints={"effective_bits": 4.8}, quantization_formats=["NVFP4_DEFAULT_CFG", "FP8_DEFAULT_CFG"], data_loader=calib_loader, forward_step=lambda model, batch: model(**batch), method="aumann_shapley", ) print(search_state["best"]["predicted_damage"]) print(search_state["best"]["predicted_damage_valid"]) ``` Or let the search choose the bit width for a predicted-damage target: ```python model, search_state = mtq.auto_quantize( model, constraints={}, quantization_formats=["NVFP4_DEFAULT_CFG", "FP8_DEFAULT_CFG"], data_loader=calib_loader, forward_step=forward_step, method="aumann_shapley", method_options={"max_predicted_damage": 0.05}, ) ``` ### Implementation - Reuses the candidate-replay and backward-scoring lifecycle introduced in #2231. - Reuses the existing AutoQuantize linear-program solver for both search directions. - Numerically integrates the coverage path when converting measured contributions into per-group damage costs. - Preserves deterministic runtime-group and candidate ordering. - Keeps raw measurements, solver scores, and damage-model diagnostics distinct in the search state. - Rejects incompatible checkpoint resumes while allowing the same scores to be re-solved for a new bit budget. - Retains the shared MoE score-module rules so routed experts are scored at their enclosing block. Recipe integration will follow separately. ### Testing Focused tests: ```text pytest -q \ tests/unit/torch/quantization/test_autoquant.py::test_backward_scoring_session_restores_partial_setup \ tests/unit/torch/quantization/test_autoquant_shapley.py ``` Result: **51 passed**. The tests cover: - end-to-end scoring and configuration generation; - agreement between path contributions and measured quantization damage; - exact allocation checks against exhaustive search; - effective-bits and predicted-damage search modes; - checkpoint resume and offline re-solving; - custom formats and heterogeneous candidate ladders; - distributed reductions and nested MoE score modules; - reused score modules and model-specific backward support; - non-finite measurements and invalid-fit reporting; and - input validation before model conversion. End-to-end checks with NVFP4 and FP8 candidates at a 6.0-bit target: - `Qwen/Qwen2.5-0.5B-Instruct` reaches 5.998 effective bits, with the summed path contributions reproducing 98% of the directly measured lowest-precision KL. - `Qwen/Qwen3-30B-A3B` reaches 6.000 effective bits and reproduces 99%, with all 48 MoE layers scored once at the sparse-MoE block rather than per expert. ### Production use We use this method in production for NVFP4 checkpoints of Kimi-K3, MiniMax-M3, and GLM-5.2. ### Before your PR is "Ready for review" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow the contributing guidance?: ✅ No copied code and no new dependencies. - Did you write the necessary tests?: ✅ - Did you update `CHANGELOG.rst`?: ✅ - Are the commits signed and signed off?: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added label-free Aumann–Shapley scoring for automatic quantization. - Added configurable path sampling, damage modeling, effective-bit targets, and predicted-damage bounds. - Added temporary weight-folding support with automatic state restoration. - Added method-specific search options, checkpoint resumption, and distributed scoring. - **Bug Fixes** - Improved cleanup and restoration of quantizer state, gradients, hooks, and forward behavior after scoring or failures. - Added validation and clearer handling for unsupported configurations and invalid measurements. - **Tests** - Expanded coverage for scoring, solver behavior, distributed execution, checkpointing, and custom quantization formats. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Joshua Hill <joshua.hill@baseten.co> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
400498d82d |
[2/4] Register each GGML IQ format once for dispatch and export (#2525)
### What does this PR do? Type of change: refactor (no behaviour change) Addresses review feedback on #2511. Backend dispatch and export each kept their own list of the GGML IQ formats: `_FAKE_QUANTS` in the backend, and `IQ_FORMATS`, `IQ_BLOCK_METADATA` and `IQ_PACKERS` in export. All four listed the same formats. Adding a format meant a row in each, and the lists could drift apart. That had already happened twice in #2511: `convert_hf_config.py` kept its own upper-case spelling of the family and dropped IQ2_XXS metadata, and the Megatron export tests were hard-wired to two formats. Each format module now declares **one `IQFormat` record** beside its encoder and decoder: name, block geometry, `quantize`, `dequantize`, and its encode and decode chunk defaults. **`IQ_FORMAT_REGISTRY`** lists them. - Backend dispatch looks formats up in the registry. - Both exporters take the packer and block geometry from it. - Export's `IQ_FORMATS` is derived from it instead of being written out again. - `_FAKE_QUANTS`, `IQ_BLOCK_METADATA` and `IQ_PACKERS` are removed. - The per-format fake-quant wrappers collapse into one `IQFormat.fake_quant`, which does the `num_bits` check and calls the existing cache helper. Codebooks, searches, payload layouts and CUDA encoders stay in each format's module. #### Series and merge order This is one slice of the IQ format series. It targets `main` so unit CI runs, and **its diff includes #2511's commits until #2511 merges**. 1. #2511 — IQ2_XXS format 2. **this PR** — one registration per format 3. #2512 — IQ2_S format 4. #2513 — IQ1_M format After this lands, #2512 and #2513 are restacked onto it, so each adds a format module and a single registry entry instead of rows in four tables. #### Design choices - **An explicit list, not self-registration at import.** If formats registered themselves when their module was imported, the registry's contents would depend on import order. - **Backward compatible, with one behaviour change.** `iq1_s_fake_quant` and `iq2_xs_fake_quant` are public on main, so each format keeps its `<fmt>_fake_quant` name as an alias of its record's method. The three removed tables were introduced by #2511 and never released. The behaviour change: on main, the alias looked the encoder up at call time, so patching `iq1_s.quantize_iq1_s` changed what it ran. Now the record captures the encoder and decoder when it's built, so patching those module functions reaches neither dispatch nor the alias. Substitute through `IQ_FORMAT_REGISTRY` instead. - **Registering a format declares it exportable, and that's intended.** Export's `IQ_FORMATS` is derived from the registry, so a format registered for dispatch is also claimed by both exporters and `convert_hf_config`. That can't be wrong for an IQ format: fake quant is `dequantize(quantize(w))`, so a format can't be dispatched without the packer and block geometry, and those are all export reads. A QAT-only IQ format can't exist. If one ever needs to land ahead of its export path, an `exportable` flag on the record is a one-line addition. - **The registry is the substitution seam.** Dispatch now reads the registry, so tests that swap an encoder or decoder swap the registry entry. Patching the format module's function would no longer reach dispatch. - **Test expectations stay independent of the registry.** Tests take the *list* of formats from the registry, but their expected values come from each format's own module (`quantize_<fmt>`, `<FMT>_BLOCK_BYTES`, …). A mis-wired registry entry therefore can't make both sides of an assertion agree. #### What it does not unify The CUDA side (`ggml.cpp` bindings, the `extensions.py` source list, codebook sizes in `common.cuh`) and the recipes and docs remain per format. "One registration" holds for the Python side, which is where all four tables lived. ### Usage Adding a format after this PR (for example IQ2_S in #2512) needs its module and one line in the registry: ```python # modelopt/torch/quantization/ggml/iq2_s.py IQ2_S_FORMAT = IQFormat( name="iq2_s", block_size=IQ2_S_BLOCK_SIZE, block_bytes=IQ2_S_BLOCK_BYTES, quantize=quantize_iq2_s, dequantize=dequantize_iq2_s, block_chunk_size=_DEFAULT_BLOCK_CHUNK_SIZE, decode_chunk_size=_DEFAULT_DECODE_CHUNK_SIZE, ) # modelopt/torch/quantization/ggml/registry.py IQ_FORMAT_REGISTRY = {fmt.name: fmt for fmt in (IQ1_S_FORMAT, IQ2_XXS_FORMAT, IQ2_XS_FORMAT, IQ2_S_FORMAT)} ``` Looking up a format: ```python from modelopt.torch.quantization.ggml import IQ_FORMAT_REGISTRY fmt = IQ_FORMAT_REGISTRY["iq2_xxs"] packed, shape = fmt.quantize(weight) # GGML blocks fmt.block_bytes, fmt.effective_bits # 66, 2.0625 ``` ### Testing - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py` — **94 passed** - `tests/gpu/torch/quantization/test_iq_formats_cuda.py` — **22 passed** (RTX PRO 6000) - `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k iq` — **27 passed** in `nvcr.io/nvidia/nemo:26.08`, the image CI uses for that suite - broader sweep of IQ, export and recipe unit tests — **153 passed**, none failed **New guards on the registry itself:** - every encoder the package exports is registered - each record points at its own format's codec, geometry and chunk defaults - the public `<fmt>_fake_quant` alias is the registered record's method - export's `IQ_FORMATS` and `QUANTIZATION_IQ*` constants match the registry - a format's `fake_quant` refuses a quantizer configured for another format. Dispatch picks the record by `num_bits`, so it never reaches this guard; the test covers direct callers of a record or alias. The three per-format guards it replaced were untested on main. - every registered format is listed in the shared test batteries Checked by mutation: leaving IQ2_XXS out of the registry, or registering it with the IQ2_XS encoder, each fails the guard written for that case. **Coverage gap closed along the way:** `test_ggml_backend.py` was hard-wired to IQ1_S and IQ2_XS, so IQ2_XXS had no backend, cache or packed-once coverage. Those tests now run over the registry. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — public per-format fake-quant names are kept as aliases; the removed tables were never released. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A — internal refactor with no user-visible change - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information Review feedback on #2511 that this addresses: *"`_FAKE_QUANTS`, `IQ_FORMATS`, `IQ_BLOCK_METADATA`, and `IQ_PACKERS` independently enumerate the same formats. A common pack/dequantize/fake_quant interface would let backend dispatch and export consume one registration."* 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * IQ quantization formats are available through a shared format registry, keeping format details and quantization behavior consistent across supported workflows. * IQ-format model exports use registered format information for quantization metadata and weight packing. * **Tests** * Expanded checks to cover registered IQ formats and verify consistent format support across quantization and export. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
1c4cde7788 |
Fix: Release per-layer expert weights in layerwise export under offload (#2466)
### What does this PR do?
Type of change: Bug fix
Layerwise export leaks one layer's worth of quantized tensors per layer
on offloaded models, and dies of OOM partway through a large MoE. Two
things accumulate, both because the offload window cannot reclaim what
the export pass adds inside it.
**1. Per-expert holder modules.** `_export_fused_experts` splits a fused
MoE experts module into per-expert holders and attaches them to the live
model:
```python
proj = nn.Module(); proj.weight = wrapper.weight # packed U8
expert.add_module(proj_name, proj)
module.add_module(str(idx), expert)
```
They are plain `nn.Module`s built *inside* the weight-access window, so
they carry no accelerate `_hf_hook`.
`weight_access_and_writeback_context` closes by iterating the modules it
collected at entry and calling `hook.post_forward()` through each one's
own offload hook — the holders satisfy neither condition, so nothing
returns them to meta.
**2. Scale buffers.** `_export_quantized_weight` registers
`weight_scale` / `weight_scale_2` / `input_scale` on the layer's
pre-existing, hooked sub-modules, and `AlignDevicesHook.post_forward`
runs with `offload_buffers=False`:
```
after post_forward: {'weight': 'meta', 'weight_scale': 'cpu'}
```
The packed weight goes back to meta; the scales do not. For
`_QuantMoELinear` models this half is the whole leak on its own —
`_reconstruct_fused_moe_linear` restacks every expert's scales into one
`register_buffer` on the hooked wrapper.
A whole-model export never notices: one pass, write the state dict,
exit. Layerwise runs the same pass once per decoder layer, so every
finished layer stays resident.
### Measurements
Through unmodified `examples/hf_ptq/hf_ptq.py` on a Qwen3.5-MoE-shaped
model (10 layers, 64 experts), printing `torch.cuda.memory_allocated()`
after each exported layer:
| placement | per layer | over 9 layers |
| --- | --- | --- |
| offload | +0.052 GiB | 0.579 → 1.047 GiB |
| resident | -0.135 GiB | falls, as designed |
+0.052 GiB is exactly one layer's quantized experts: packed U8 100.7M/2
= 0.047 GiB plus FP8 block scales 100.7M/16 = 0.006 GiB.
At scale it is fatal rather than wasteful. Qwen/Qwen3.8-2.4T-A95B (92
layers, 512 experts) leaks 12.9 GB packed + 1.6 GB scales per layer, so
92 layers want 1.33 TB that no budget on a 283 GB card or 952 GB host
absorbs. The run died of CUDA OOM at layer 16/92 with
`--max_gpu_memory_gb 240`, and at `--max_gpu_memory_gb 30` leaked the
same 15 GB/layer onto the host instead.
### The fix
`_release_exported_tensors` (`model_utils.py`) is a context manager that
snapshots each sub-module's child-module and buffer names on entry, and
on exit drops whatever appeared. Persisting happens *inside* the block,
so "release only once it is on disk" is structural rather than a
comment.
Both packing sites use it: `LayerwiseExporter.export_layer` and the
offload decoder loop in `_export_transformers_checkpoint_streaming`.
The streaming writer had solved the same leak inline with a heuristic —
null every CUDA buffer, and every CUDA parameter on a hook-less module —
and that block is deleted in favour of the shared helper. Keying on
*what the pass added* rather than on device and hook presence drops two
assumptions that only held for a terminal, offloaded export: it no
longer nulls buffers the layer already had, nor parameters of
sub-modules accelerate simply did not hook. That is also what makes it
safe for the layerwise path, where resident models are supported and the
model outlives the export.
Deliberately out of scope: the FSDP2 per-unit loop in
`collect_export_tensors` keeps its per-unit holders. That predates this
PR, this PR does not touch that loop, and closing it needs its own
change and its own test.
### Usage
No API change. Existing layerwise export under offload simply stops
growing:
```bash
python examples/hf_ptq/hf_ptq.py \
--pyt_ckpt_path Qwen/Qwen3.8-2.4T-A95B \
--qformat nvfp4 \
--export_path /path/to/export \
--max_gpu_memory_gb 240
```
### Testing
- `tests/unit/torch/export` and `tests/unit/torch/quantization` — 1246
passed
- `tests/gpu/torch/export/test_layerwise_export.py` +
`test_offload_export.py` — 35 passed
- `cuda_alloc` over 10 offloaded layers: 0.526 → 0.518 GiB (-0.008), was
+0.468
- Exported checkpoint byte-identical to the unfixed run: 7841 tensors, 0
mismatches, max abs diff 0.0; `hf_quant_config.json` / `config.json` /
index identical
- Full Qwen3.8-2.4T-A95B PTQ then completed all 92 layers: peak GPU 117
GB of 283, peak host RSS 71 GB of 952, flat across 50 consecutive layers
at 23-26 s/layer
Coverage added: the existing `test_export_creates_per_expert_submodules`
now runs the export inside the context manager and asserts the holders
are gone on exit, and a new test in `test_offload_export.py` pins the
`offload_buffers=False` behaviour the buffer half exists for —
export-registered scales dropped, pre-existing buffers untouched.
`tests/gpu/torch/export/test_fsdp2_export.py` reports 34 failures in my
environment. They are **pre-existing and unrelated**: the same 34 fail
identically on this branch and on the merge-base (`2b1f33d0ef`), with
byte-identical failure sets and runtimes within 3 s. All 34 are `Failed:
Timeout (>120.0s)` from hung NCCL collectives, with zero assertion
failures.
The GPU export suites and the whole-model measurements above were run at
`2999d7cfb7`. The two commits since — swapping the holder marker for a
child-name diff, and moving the helper to `model_utils` — are covered by
the unit suites; a GPU re-run before merge is worthwhile.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — one new offload test for
the buffer half, and the existing fused-experts export test now covers
holder release
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — `layerwise.export_dir` is new in the unreleased 0.48.0, so this
bug was introduced and fixed within the same cycle
- Did you get Claude approval on this PR?: ✅
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
d16dad1c20 |
docs(llm_distill): document KDTrainer-based example flow (#2524)
## Summary - Update `examples/llm_distill/README.md` to describe the current `main.py` flow, which uses `KDTrainer` (from `modelopt.torch.distill.plugins.huggingface`) instead of `mtd.convert()` / `DistillationModel` wrapping. - Document that `KDTrainer` only supports logit-level distillation today, and that hidden-state/intermediate-layer KD still requires `mtd.convert()` + `DistillationModel` until `KDTrainer` gains that support. ## Test plan - [x] Reviewed rendered README diff for accuracy against `modelopt/torch/distill/plugins/huggingface.py` and `examples/llm_distill/main.py` - [x] `pre-commit` hooks (markdownlint-cli2, etc.) passed on commit 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the Hugging Face getting-started example to use `KDTrainer` with the standard training and model-saving workflow, including an example of combining it with `SFTTrainer`. * Clarified that `KDTrainer` supports logit-level distillation; hidden-state distillation uses a separate approach. * Explained that KD loss and evaluation cross-entropy are reported separately, with weighted CE/KD loss combination unsupported. * Added guidance on distributed training options, including the FSDP2 requirement and alternatives to default DataParallel. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
0fdda7937b |
Date 0.47.0 changelog for official release (#2529)
Set the 0.47.0 changelog date to 2026-09-23. No code changes. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the 0.47.0 changelog entry to show September 23, 2026, as the release date. No user-facing product behavior changes are included in this update. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chad Voegele <cvoegele@nvidia.com> |
||
|
|
a21411adde |
Add the IQ2_XXS weight-only quantization format (#2511)
### What does this PR do? Type of change: new feature llama.cpp defines five GGML IQ formats at one and two bits; we ship two. This adds **IQ2_XXS** at 2.0625 bits per weight, between IQ1_S and IQ2_XS, and is the **first of three**. On a real mixed-precision checkpoint (`unsloth/Qwen3.8-27B-GGUF`, `Qwen3.8-27B-UD-IQ1_S.gguf`) IQ2_XXS alone covers **59 tensors and 2.84 B parameters — 10.6% of the file**, which a reader limited to IQ1_S/IQ2_XS cannot consume. Across all three PRs the missing formats account for 17.3%. | format | bpw | bytes/256 | codebook | | |---|---|---|---|---| | `iq1_s` | 1.5625 | 50 | `iq1s_grid` (2048) | existing | | **`iq2_xxs`** | **2.0625** | **66** | **`iq2xxs_grid` (256)** | **this PR** | | `iq2_xs` | 2.3125 | 74 | `iq2xs_grid` (512) | existing | The encoder follows the existing single-pass grid search at a fixed anchored super-block scale, and the CUDA kernel the existing per-block structure. IQ2_XXS reuses IQ2_XS's even-parity sign rule but packs a 4-bit sub-block scale into the same 32-bit word as four 7-bit sign indices, and its 256-entry grid needs no high index bits. ### Groundwork the next two reuse Two things land here because IQ2_XXS is the first format to need them: - **Export registry.** The IQ family was spelled as a two-element tuple at **nine** sites across `quant_utils.py`, `unified_export_hf.py` and `unified_export_megatron.py`. Those become an `IQ_FORMATS` frozenset plus per-format packer and block-geometry tables, so a format is a row rather than a sweep through the exporters. - **Shared test contract.** The per-format test files had drifted apart — each of `iq1_s` and `iq2_xs` tested things the other did not. They become one parametrized module per layer (unit and CUDA), so every format is held to the same contract and a new one inherits it. ### Usage ```bash python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq2_xxs ``` ### Testing **The decoder is validated against llama.cpp's own output, not just round-tripped.** Every IQ2_XXS tensor in the checkpoint above, compared against `dequantize_row_iq2_xxs` from `ggml-quants.c`: ``` IQ2_XXS: 59 tensors, 11,100,160 blocks → 0 mismatched, max|diff| 0.0 ``` The new codebook matches the `ggml-common.h` table entry for entry, as does the `ksigns_iq2xs` sign table. Blocks lifted from that checkpoint ship as conformance vectors so CI keeps checking bytes we did not produce; mutation testing confirms they catch a wrong sign-field width. The CUDA encoder is byte-identical to the PyTorch reference on a fixed input and runs at **1047.9 M elem/s against the torch search's 10.7** on a 5632×2048 weight. - `tests/unit/torch/quantization/ -k 'ggml or iq1 or iq2 or iq_'` — 99 passed - `tests/gpu/torch/quantization/test_iq_formats_cuda.py` — 21 passed (7 checks × 3 formats) - `tests/unit/recipe/test_presets.py` — passing; `general/ptq` now holds 29 recipes, `ptq.md` updated - reconstruction error decreases monotonically with bit width, pinned by a test Pre-existing failures in `tests/unit/torch/export/` and `test_autoquant.py` are `transformers`/`torchvision` import problems in my environment — identical counts with and without this change. ### A finding about already-merged code Checking the new kernel against its PyTorch reference at 4096 blocks showed that **CUDA and torch encoders disagree on roughly 1 block in 6000 — including the already-merged `iq2_xs`**, at 0.0163% against IQ2_XXS's 0.0000%. Root cause: both compute `xnorm − 2·scale·dot + scale²·qnorm`, but CUDA fuses it with `fmaf` while torch uses separate ops; where two local scales fall within a float32 ULP the roundings pick different sides. Adjudicated against float64, neither path is better (5 to 6). Worst-case cost is **1.48e-08** relative reconstruction error, and run-to-run determinism on a given device holds. This is pre-existing, not introduced here — `test_iq2_xs_cuda.py` asserts exact byte parity but on a 16-block weight where ties essentially never arise. I have **not** changed that test; rewording a guarantee on merged code belongs in its own change. The new shared GPU tests assert exact parity on a small fixed input and compare reconstruction error at scale. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — the new codebook is a GGML table, carried in `codebooks.py` beside the existing ones so the MIT-licensed surface stays in that one file, with the source revision recorded. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information First of three; **IQ2_S** and **IQ1_M** follow and build on this branch. Replaces #2505, which carried all three at once. Follows #2446 / #2447 / #2448 / #2449, which landed IQ1_S and IQ2_XS. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added IQ2_XXS weight-only quantization, including CUDA acceleration and support for Hugging Face and Megatron exports. - Added the `general/ptq/iq2_xxs` recipe. It requires no calibration data and supports eligible layers with a weight dimension divisible by 256. - Updated the PTQ recipe catalog to list IQ1_S, IQ2_XXS, and IQ2_XS at approximately 1.56, 2.06, and 2.31 bits per weight, respectively. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f2f0d6958e |
Add MLflow tracking flags to megatron_bridge quantize.py (#2477)
### What does this PR do?
Type of change: new feature
`examples/megatron_bridge/quantize.py` gains the MLflow tracking flags
`examples/hf_ptq/hf_ptq.py` already has: `--mlflow <tracking-uri>`
(MLflow's own `$MLFLOW_TRACKING_URI` is honoured too),
`--mlflow_experiment` and `--mlflow_run_name`. Only the master rank
opens a run, so a `torchrun` launch produces one run carrying the
invocation, every command-line argument as a searchable param, the
resolved `--recipe` (with `$import`s expanded), that rank's log and the
quantizer summary. Once `bridge.save_megatron_model` returns,
`.experiment.json` is written into `--export_megatron_path`, so a
Megatron checkpoint found on disk names the run that produced it; a run
that fails is still recorded as `FAILED` with its traceback.
Rather than copy the wiring a third time, the part `hf_ptq` and
`vllm_serve` had each duplicated moves into
`modelopt.torch.utils.mlflow`:
- `add_mlflow_args(parser, tool, tracks=, variant_help=)` — the three
flags, registered under both the `--mlflow_x` and `--mlflow-x` spellings
(vLLM's `FlexibleArgumentParser` only matches the dashed one).
- `resolve_tracking_uri(uri, parser)` → `(uri, required)` — the flag
overrides the environment and is fatal when the URI is unusable; a URI
inferred from `$MLFLOW_TRACKING_URI` warns and continues untracked,
since that variable is commonly exported for unrelated tooling.
- `resolve_mlflow_args(args, parser, tool, model, variant)` — the same,
settled onto `args`, plus the default experiment name.
- `EXPERIMENT_JSON`, `MlflowRunLogger.log_experiment_json()` and
`drop_experiment_json()` — the checkpoint→run provenance pointer,
previously private to `hf_ptq`.
Both existing callers now delegate to those, keeping their own help
wording and variant naming, so the three scripts share one convention
instead of three copies (`example_utils.py` and `vllm_mlflow_utils.py`
each lose ~60 lines). Their flags and defaults are unchanged; the only
user-visible difference is that `hf_ptq`'s ignored-URI warning gains the
`$` the vLLM one already had (`Ignoring $MLFLOW_TRACKING_URI, continuing
untracked`), so one shared message serves both.
One behaviour change reaches `hf_ptq` through the shared helper, and it
is a fix: when tracking was inferred from `$MLFLOW_TRACKING_URI` and the
run never opened (unreachable server, or `mlflow` not installed), it
used to leave the previous run's `.experiment.json` beside a freshly
exported checkpoint. `log_experiment_json` now drops the pointer when it
has no run to record, so after a completed export the file is this run's
or absent.
The new example-side code lives in
`examples/megatron_bridge/mlflow_utils.py`, which deliberately imports
no Megatron, so the whole flag-to-artifact path is testable without the
Megatron container (the same split
`examples/vllm_serve/vllm_mlflow_utils.py` uses).
### Usage
```bash
torchrun --nproc_per_node 2 quantize.py \
--hf_model_name_or_path Qwen/Qwen3-8B \
--recipe general/ptq/nvfp4_default-kv_fp8 \
--tp_size 2 \
--export_megatron_path /tmp/Qwen3-8B-NVFP4-megatron \
--mlflow https://<your-mlflow-server>/
# The checkpoint then names the run that produced it:
cat /tmp/Qwen3-8B-NVFP4-megatron/.experiment.json
```
The experiment defaults to `$USER/megatron_bridge_quantize/<model
basename>-<recipe name, or --quant_cfg>`.
### Testing
- `tests/examples/megatron_bridge/test_mlflow_utils.py` — 20 new tests
covering the flags (both spellings, env-vs-flag precedence, the
fatal/best-effort split), the params/tags/artifacts a run records, rank
gating, and the `.experiment.json` lifecycle. The last one guards the
seam with `quantize.py` as text, since that script needs Megatron to
import.
- `tests/unit/torch/utils/test_mlflow.py` — 13 new tests for the
extracted library API; suite at **75 passed**.
- Full `tests/examples/megatron_bridge` suite in
`nvcr.io/nvidia/nemo:26.08` on an RTX 6000 Ada: **37 passed (26m)**,
including the three `test_quantize_export` cases that drive the real
`quantize.py`, plus QAD, distill and prune.
- Regression proof for the refactor:
`tests/examples/hf_ptq/test_hf_ptq_args.py` **47 passed** and
`tests/examples/vllm_serve/test_vllm_mlflow_utils.py` **32 passed**,
unchanged apart from one renamed constant reference.
- Both new guards were shown to fire: mutating the `checkpoint_exported`
gate and removing `with mlflow_run(args):` each failed exactly one test.
- End-to-end tracked run in `nvcr.io/nvidia/nemo:26.08` (tiny Qwen3-MoE,
`general/ptq/fp8_default-kv_fp8`, 1 GPU) against an internal MLflow
server: run `47d4ccd7cd9e48269e7248868347ccd0` under experiment
`$USER/megatron_bridge_quantize/mbridge-ptq-validation` closed
`FINISHED` carrying `command.txt`, `version.txt`, `experiment.json`,
`recipe/resolved_recipe.yaml`, `logs/quantize.log` and
`summary/quant_summary.txt`; all 19 CLI arguments plus `world_size`
logged as params with no `mlflow_*` leakage, the
`model`/`checkpoint_path`/`source_checkpoint_path` tags set, and
`.experiment.json` written into the Megatron checkpoint beside
`iter_0000000/`.
- `pre-commit run --files <changed>`: all hooks pass (ruff, mypy,
bandit, markdownlint).
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under *Megatron Framework (M-LM / M-Bridge)*.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
`mlflow` stays an optional dependency, imported only once tracking is
enabled, so an untracked run behaves exactly as before.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added optional MLflow tracking for Megatron-Bridge quantization runs.
- Configure tracking with `--mlflow` or `MLFLOW_TRACKING_URI`, with
customizable experiment and run names.
- Records searchable parameters, resolved recipes, quantization
summaries, logs, and checkpoint provenance.
- Captures successful and failed runs and cleans up stale checkpoint
metadata when appropriate.
- **Documentation**
- Added setup instructions and usage examples covering artifacts,
naming, checkpoint metadata, validation, and authentication.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
25d8c91762 |
Add guidance on keeping skill updates concise to AGENTS.md (#2523)
### What does this PR do? Type of change: documentation Adds an `## Updating skills` section to `AGENTS.md` (symlinked as `CLAUDE.md`). Skills are loaded into agent context, so each extra line costs tokens every time the skill runs. The new guidance tells the agent to: - Keep skill edits concise: add only what changes agent behavior, and tighten existing text instead of appending more. - Do a final compression pass over the skill diff before opening a PR: drop unnecessary explanations and examples, cut redundancy, and merge overlapping guidance. ### Usage N/A — no API or flag change. ### Testing `pre-commit run --files AGENTS.md` (markdownlint and the other applicable hooks pass). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A <!-- documentation-only change --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- agent instructions only, not user-facing --> - Did you get Claude approval on this PR?: ❌ <!-- will run /claude review if reviewers want it --> ### Additional Information Follows #2494, which added the PR sizing guidance to the same file. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added guidance for keeping skill updates focused on behavior changes and reviewing edits for unnecessary detail. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
87f7d1432f |
fix(speculative): hold the DFlash draft's fp32 master weights in the optimizer (#2483)
### What does this PR do? Type of change: Bug fix **Follow-up to #2342**, which split this out on review (commit `c67784d9`), and a rethink of how the flag is implemented. `dflash_fp32_master_weights` exists because the DFlash draft is cast to the frozen bf16 target's dtype, so AdamW allocates its moments in bf16 — and bf16 is too coarse to hold them. At `beta2=0.999` a single step changes `v` by at most **0.100%**, while the smallest change bf16 can represent near `v` is **0.164% mean / 0.388% max** (measured): every decrease rounds away, `v` only grows, and the effective step size decays on its own from step 1. #2342 fixed that by **promoting the draft model to fp32**. Everything else followed from giving the model a dtype the rest of it does not have — a bf16 autocast at every entry point, two transformers loader hints so `from_pretrained(dtype="auto")` would not round the draft away, a post-condition check because those hints fail silently, and a doubled DDP gradient all-reduce. **This PR puts the fp32 in the optimizer instead**, where Megatron-LM, DeepSpeed and apex put it. `MasterWeightAdamW` holds an fp32 master copy of each non-fp32 parameter plus fp32 moments in `self.state[p]`, steps on the master, and copies back at the parameter's dtype. The model is never anything but the base dtype, so every one of those follow-on pieces is deleted, gradients stay bf16, and the exported drafter is unchanged. What the placement costs is that wiring the optimizer becomes the training loop's job: `EagleTrainerWithAccLog.create_optimizer` builds it, and `VerifyMasterWeightsCallback` raises at the end of step 1 if the moments are not fp32. **The default flips to `True`** — the flag now changes optimizer memory and optimizer arithmetic and nothing else. Flipping it on the model-promoted implementation turns **25 of 259** unit tests red; flipping it here is **259 passed**. Set it to `False` to reclaim the memory, about 12 bytes per draft parameter instead of 4. <details> <summary>Three drive-by fixes, independent of the above</summary> - `_place_draft` is folded back into `modify()` — it fused the draft's dtype, its device and an eager rotary buffer behind one meta guard. - The module docstring's claim that `DFlashModule` has an `_apply` meta-buffer fix is removed (`grep "def _apply"` matches nothing, and never did). - #2342's field description no longer lists `evaluation` as a broken path — `forward` short-circuits to the base model when `not self.training`, so the draft never runs there. </details> ### Usage No API change. `dflash_fp32_master_weights` now means the *optimizer* holds fp32 master weights rather than the draft model being fp32. ### Testing **1 · The refactor is arithmetically a no-op.** Both implementations run AdamW on an fp32 tensor, so given the same starting values and the same gradients the trajectories are identical — 1000 steps, `weight_decay=0.01`: ``` old fp32 parameter vs new fp32 master : bitwise equal = True (max |diff| 0.0e+00) exp_avg / exp_avg_sq : bitwise equal = True optimizer state dtypes : ['torch.float32'] model parameter dtype : torch.bfloat16 ``` Initial values have to be matched at bf16 first, or the bf16 arm's one-time rounding of the draw shows up as a 2e-4 "difference" that is not arithmetic. With that controlled, the two implementations differ only in their *inputs*: gradient precision (fp32 vs bf16 — torch 2.10 requires `grad.dtype == param.dtype`) and that one-time rounding. **2 · End to end on GPU: the effect survives the refactor.** Qwen3-1.7B base, real corpus, one GPU per arm, three arms — pure bf16 (flag off), the #2342 implementation, and this one — on two algorithms trained independently, sharing seed, data order and initialisation within an algorithm. <img width="2925" height="960" alt="image" src="https://github.com/user-attachments/assets/40d37059-ab8a-4924-b049-85f76b70b156" /> The two fp32 arms sit on top of each other for the whole run while bf16 stays above both, and the old-vs-new gap is 10–23× smaller than the fp32-vs-bf16 effect it has to be compared against. **Acceptance length says the same thing, and settles what the loss could not.** All six drafters at the end of those curves were exported and served under vLLM against the same base, and measured on MT-Bench (80 prompts, 8 categories, greedy, one request at a time, `num_speculative_tokens` = trained `block_size` − 1, every knob but the drafter held fixed): | | pure bf16 | fp32 in model (#2342) | fp32 in optimizer (this PR) | new − old | fp32 − bf16 | |---|---|---|---|---|---| | `dflash` | 1.3068 | 1.3708 | **1.3666** | −0.0042 `t=−0.90` | +0.0619 `t=+13.8` | | `lilicorr` | 1.2536 | 1.2814 | **1.2882** | +0.0068 `t=+1.42` | +0.0312 `t=+9.2` | Paired by prompt, n=80. On both algorithms the new-vs-old 95% CI straddles zero (`dflash` [−0.0134, +0.0051], `lilicorr` [−0.0028, +0.0164]) while fp32-vs-bf16 does not come close to it, and the sign of new-vs-old **flips between the two algorithms** — what a rounding difference looks like, not a bias. This is also the comparison the training loss could not give: all three arms are **exported and served in bf16**, so the old implementation's fp32 draft weights are rounded at export exactly as they would be for deployment, and the "its loss was computed on a more precise forward" caveat below does not apply. `lilicorr` needs [vllm-project/vllm#57934](https://github.com/vllm-project/vllm/pull/57934), applied as an overlay so that both algorithms are measured on one engine build. The right panel is the mechanism, and the one signal that depends on neither the seed nor the choice of loss statistic: Adam's updates to the draft's RMSNorm gains are smaller than the bf16 ULP at 1.0 (0.0078), so in the bf16 arm every one of them rounds away and the gains never move — not one of `dflash`'s 14 in 30000 steps, and two of `lilicorr`'s 20 by 3e-06. Both fp32 arms move all of them, by the same amount. Two results behind the figure rather than in it. **fp32-vs-bf16 grows with the horizon** while old-vs-new does not — on `dflash` −0.129 at 1500 steps → −0.262 at 15000 → −0.341 at 30000, and on `lilicorr` −0.191 → −0.220 → −0.285, against an old-vs-new difference that stays near 0.02 at every horizon and changes sign between them (−0.026 → +0.028 on `lilicorr`). That is what a compounding bias and a rounding difference respectively should look like, and it is the reason the longer runs were worth doing. And **across seeds**, the paired old-vs-new difference at 1500 steps is +0.0003 (n=6) on `dflash` and +0.0643 (n=10) on `lilicorr`, both with a 95% CI straddling zero. <details> <summary>Limits of the above, stated rather than smoothed over</summary> At 5 seeds the `lilicorr` paired difference read +0.1610 ± 0.0557 (t=+2.89, 4/5 seeds in the same direction) — nominally significant, suggesting the new implementation was genuinely worse there. Four further `lilicorr` seeds were run against that pre-declared question; two came back strongly negative and the estimate settled at +0.0643 (95% CI [−0.086, +0.214]). The earlier reading was small-sample noise. At 1500 steps on `lilicorr` that CI is *not* narrower than the fp32-vs-bf16 effect it is being compared against, so the 1500-step sweep alone cannot certify equivalence there — `lilicorr` is still at loss 9.3 and deep in its early transient, and it is the long runs that resolve it. On `dflash` the 1500-step CI (±0.031) is already 4× tighter than the effect (−0.129). One asymmetry the loss comparison cannot separate: the old implementation held the draft weights in fp32 *at forward time*, so its training loss was computed on a more precise forward, while both implementations export bf16. Any residual advantage it appears to have is therefore an upper bound. </details> **3 · Unit tests.** `tests/unit/torch/speculative/` — **259 passed** on **transformers 5.0.0** and **5.3.0**, both ends of the supported `>=5.0,<5.13` (CPU, torch 2.10). `TestDFlashFp32MasterWeights` is rewritten for the new mechanism; the two that would have caught the traps in this design are `test_resume_does_not_round_the_master_back_down` (`Optimizer.load_state_dict` casts float state to its parameter's dtype, so a naive subclass rounds the master and both moments to bf16 on *every* resume, silently, with the loss still falling) and `test_the_callback_refuses_a_loop_that_forgot_the_optimizer`. The rest cover the draft's dtype with the flag either way, that no forward path needs an autocast any more, that plain AdamW really does leave the moments in bf16, and that an fp32 model allocates no redundant master. A sharded FSDP2 `DTensor` keeps an fp32 master and fp32 moments through a step. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ for artifacts, with one intentional default change. The draft's stored dtype goes back to matching the base, as it was before #2342; existing checkpoints load unchanged and the exported drafter is unaffected. The flag now defaults to **`True`** — the measurements above are the reason, and the cost is fp32 master + fp32 moments for the draft only. A training loop that builds its own optimizer instead of using the shipped `create_optimizer` gets plain AdamW and none of this; `VerifyMasterWeightsCallback` makes that fail loudly at step 1 rather than skip the feature quietly. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new dependencies. - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ — draft; will run `/claude review` before marking ready. ### Additional Information **On the 7–14% acceptance-length gain quoted in #2342:** that was measured on the fp32-model arithmetic and is not re-derived here. What is measured above is the like-for-like comparison this PR has to answer — same corpus, same horizon, same serving path, one implementation swapped. **History:** commits 1–3 restore the autocast design as it was split out; commits 4–6 replace it. Happy to squash before review. <details> <summary>Alternatives measured and rejected, so they do not get re-proposed</summary> - **Swapping `p.data` to the master and calling `super().step()`** (reuses all of AdamW, ~20 lines instead of ~50): bit-identical on ordinary parameters over 25 steps, but silently wrong under FSDP2 — assigning `.data` on a `DTensor` parameter updates the wrapper's reported dtype while the local shard keeps the model's, so `p.dtype` reads fp32, `p.data.dtype` reads bf16, and `zeros_like(p)` allocates the moments in bf16 anyway. CPU tests pass either way. - **Narrowing the autocast from `__call__` to `forward`** (while it still existed): turns 10 Domino/DSpark tests red — the variants apply their heads in their own `forward` overrides, outside `DFlashModule.forward`. - **Building the rotary buffer on meta and letting the loader materialise it**: makes RoPE correctness depend on transformers selecting a branch by class-name substring (`"RotaryEmbedding" in module.__class__.__name__`), and the `if not hasattr` guard is then permanently satisfied, so a later `to_empty()` leaves garbage forever — measured `4.56e-41`, i.e. cos=1 / sin=0, no positional encoding at all. - **Building it eagerly in `DFlashModule.__init__`**: lands before the dtype cast, so `Module.to` rounds the RoPE frequencies to bf16 on the default path — measured `0.8659643530845642` → `0.8671875`, loss `3.47230935097` → `3.47114777565`. </details> 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - DFlash now uses FP32 optimizer master weights and Adam moments by default while keeping draft parameters in the base model’s dtype. - Master-weight training preserves optimizer precision when restoring checkpoints. - The feature can be disabled to reduce optimizer memory usage. - Draft models consistently follow the base model’s dtype and device. - **Bug Fixes** - DFlash workflows now support operation without autocast. - Added validation for compatible AdamW-family optimizers and master-weight precision, including resumed training runs. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f2ee751089 |
Validate eval limit accounting and raise TB 2.1 timeout (#2499)
### What does this PR do? Type of change: documentation Successful evaluations can still contain timed-out trials or responses stopped by output limits. Existing skills scan for errors but do not require counts or rates, allowing serving-speed effects to be mistaken for quantization accuracy changes. Require per-task timeout and output-limit accounting before scores are treated as validated, with explicit denominators, telemetry coverage, effective limits, and artifact evidence. Separate request retries, terminal trial failures, and resumable SLURM walltime events. Carry that evidence into baseline/candidate comparisons; unknown accounting or unresolved infrastructure effects prevent an acceptable verdict. Preserve benchmark-defined failures and label diagnostic protocol changes explicitly. Allow valid-with-warnings results for small fractions of limit-hit responses/trials when coverage, protocol, and other checks pass, without automatically retrying. Clarify that a BF16 acceptance gate cannot be satisfied with an FP8/INT4 baseline. Correct Terminal-Bench 2.1 guidance that sharding cannot affect scores and require accounting across the full trial set. Raise the ModelOpt TB 2.1 agent budget from 7200 to 14400 seconds in the recipe and set it explicitly in the template. With timeout_strategy=max, this is a minimum budget, not a hard ceiling. Set sandbox lifetime to six hours and example SLURM walltime to eight hours to allow setup and verification; partition limits and effective task budgets must still be checked. Baseline and candidate must use the same policy, and old two-hour results require remeasurement for a matched comparison. This changes skill defaults, not the upstream benchmark protocol or harness instrumentation. The upstream-vendored launching-evals skill is unchanged per repository policy; its standalone workflow still needs an upstream update. Reduce the evaluation entrypoint from 6,156 to 888 words (86%) by moving launcher Steps 1–8 into an on-demand reference without changing their instructions. Preserve step headings/links for existing callers. Shorten timeout accounting from 550 to 352 words (36%) while retaining the checks and warning policy. This reduces initial context; full launcher workflows still load the relevant detailed sections. ### Usage TB 2.1 recipe/template defaults now include: ```yaml solver: timeout_strategy: max run_timeout: 14400 sandbox: max_task_lifetime_sec: 21600 ``` The example uses `cluster.walltime: "08:00:00"` where the partition permits it. Request timeout remains 3600 seconds. For leaderboard comparisons, use the benchmark protocol rather than assuming the ModelOpt override is comparable. ### Testing - Pre-commit passed on all six changed Markdown/YAML files. - skill-creator quick_validate.py passed for evaluation and compare-results. - git diff --check passed. - Validated recipe and template solver/sandbox settings against NEL schemas with the pinned TB playbook; checked max and task timeout resolution. - Verified extracted launcher Steps 1–8 match the original text, apart from a trailing blank line. - Manually reviewed handling of recovered retries, resumed artifacts, incomplete telemetry, benchmark-defined limits, and timeout-sensitive comparisons. No evaluation jobs were launched; increased timeout defaults have not been benchmarked. ### Before your PR is "*Ready for review*" Contributor guidelines and security coding practices reviewed. - Is this change backward compatible?: ✅ Existing configs remain valid. New TB configs use longer budgets and may consume more runtime; comparison requires matched timeout policies. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md: N/A - Did you write any new necessary tests?: N/A — skill/config changes validated as listed above. - Did you update Changelog?: N/A — internal skill guidance. - Did you get Claude approval on this PR?: ❌ Not requested yet. ### Additional Information A nonzero benchmark-defined timeout or output-cap rate does not automatically invalidate a score. The change requires evidence and protocol-aware interpretation, without choosing a universal acceptable rate or silently excluding affected trials. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Expanded evaluation guidance with a centralized launcher workflow covering setup, configuration, execution, monitoring, authentication, and failure handling. * Required timeout and output-limit accounting for every task, including successful and non-reasoning runs, with rates, denominators, telemetry coverage, effective limits, and recovery status. * Clarified that unknown telemetry, mismatched limits, or unresolved infrastructure effects prevent an acceptable verdict. * Updated comparison guidance to require matched reruns, aligned precision baselines, and limit-hit rate comparisons. * Clarified distributed benchmark timeout handling and score extraction; GDPVal support is no longer documented. * Updated example evaluation time limits and SLURM walltime. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> |
||
|
|
7159c01d9d |
[chore]: weekly bump of uv.lock on main (2026-09-21) (#2490)
## Summary Automated weekly update of uv.lock file for nSpect Scanning: - `uv.lock` — upgraded all transitive dependencies to latest compatible versions <details> <summary>uv lock --upgrade output</summary> ``` Using CPython 3.12.14 interpreter at: /opt/hostedtoolcache/Python/3.12.14/x64/bin/python3 Resolved 206 packages in 9.16s Updated cachetools v7.1.8 -> v7.2.0 Updated cuda-pathfinder v1.8.1 -> v1.8.2 Updated databricks-sdk v0.139.0 -> v0.140.0 Updated deepspeed v0.19.6 -> v0.19.7 Updated filelock v3.32.6 -> v3.32.7 Updated huggingface-hub v1.31.0 -> v1.32.0 Updated hydra-core v1.3.6 -> v1.3.7 Updated idna v3.19 -> v3.20 Updated mlflow-skinny v3.16.0 -> v3.16.1 Updated multidict v6.8.0 -> v6.9.0 Updated pandas v2.3.3, v3.0.5 -> v2.3.3, v3.0.6 Updated peft v0.20.0 -> v0.21.0 Updated platformdirs v4.11.8 -> v4.11.11 Updated propcache v0.5.2 -> v0.5.4 Updated pyparsing v3.3.2 -> v3.3.3 Updated python-discovery v1.6.0 -> v1.6.1 Updated urllib3 v2.7.0 -> v2.8.0 Updated uv v0.12.13 -> v0.12.17 Updated virtualenv v21.7.9 -> v21.9.0 Updated watchfiles v1.2.0 -> v1.3.0 Updated yarl v1.24.5 -> v1.25.1 ``` </details> Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> |
||
|
|
1b4e7dfb14 |
[OMNIML-5899] Add IQ post-training quantization recipes (#2449)
## Summary - add numerics configs for IQ1_S and IQ2_XS - add model presets and general PTQ recipes - document the supported weight-shape requirement - reuse the packed IQ weight across forwards instead of re-encoding it every time - add end-to-end `hf_ptq` coverage for both formats - add the release-note entry ## PR split This work is split into four focused PRs. Each PR targets `main` and owns a disjoint file set: 1. **Kernel** — [#2448: Add CUDA kernels for IQ packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448) 2. **Quantization** — [#2446: Add IQ quantization codecs and backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446) 3. **Export** — [#2447: Export IQ checkpoints from HF and Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447) 4. **Recipes** — [#2449: Add IQ post-training quantization recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449) The required merge order is #2448, #2446, #2447, then #2449. ## Scope Sixteen files. The PR started as eight recipe config, documentation and changelog files; the end-to-end test added for them surfaced two performance bugs in the already-merged codec, and fixing those pulled in the codec files and their tests. Beyond the original recipe set it now touches four files owned by #2446 — `ggml/common.py`, `ggml/iq1_s.py`, `ggml/iq2_xs.py` and `ggml/backend.py` — plus three test files. It still contains no kernel or export files. Keeping those fixes here rather than moving them to #2446 is deliberate and confirmed with the stack owner: #2446 is already merged, and both bugs are only observable through the end-to-end test this PR adds, so splitting them would separate each fix from the test that demonstrates it. ## Packed-weight cache fix Adding the end-to-end test made the cost visible: a TinyLlama IQ1_S `hf_ptq` run spent **499 of its 537 seconds inside the IQ1_S encoder**, packing the same 154 weights 15400 times — 100 times each. The repacking is not calibration. These recipes set `algorithm: null` and `hf_ptq` logs `Dynamic quantization. Calibration skipped.`. The 100 passes are the sample `generate()` calls `hf_ptq` makes before and after quantization: one decode step re-runs weight fake-quant on every linear, and the packed payload was thrown away each time. `_PackedWeightCache` was already there to prevent exactly this, and it never hit. It keyed on the identity of the tensor the backend was handed, but `TensorQuantizer` passes a fresh *view* of the weight on every forward, so the identity check never matched twice. The fix anchors the entry to `inputs._base` — the parameter the view is taken from — held as a weakref. The parameter is stable across forwards, so the cache hits; the reference stays weak, so the payload is released with the weight and offloaded/meta-device flows are unaffected. (A strong reference does make the cache hit, but pins full-precision storage for the life of the quantizer, which is the opposite of what those flows need.) Measured on TinyLlama IQ1_S `hf_ptq`, 2×H100, same command before and after: | | packer calls | packing time | wall clock | |---|---|---|---| | before | 15400 | 499.0 s | 8m57s | | after | 154 (one per weight) | 6.2 s | 3m21s | `test_ggml_weight_is_packed_once_across_forwards` pins this: it counts encoder calls across five forwards under `torch.inference_mode()` (what `generate()` runs under) and asserts exactly one. ## Decode chunk fix With packing cached, the end-to-end cost moved entirely into the decode, and IQ2_XS was still 4x slower than IQ1_S (1081s vs 260s per case). Instrumenting both showed packing was no longer the cost at all — IQ2_XS packs *faster*: | | pack calls | packing time | wall clock | |---|---|---|---| | IQ1_S | 154 | 6.2 s | 3m21s | | IQ2_XS | 154 | 1.6 s | 17m59s | The cause was one constant serving two loops with opposite characteristics. `_DEFAULT_BLOCK_CHUNK_SIZE` bounds the torch encode fallback, which holds the large codebook-search temporaries and runs once per weight; IQ2_XS sets it to 256 rather than IQ1_S's 1024 because its search sweeps sixteen local scales per grid tile. But the *decode* shared it — and the decode has tiny temporaries, runs on every forward, and is never cached, so a small chunk only multiplies kernel launches. Decoding a 2048x5632 weight: | chunk | IQ2_XS decode | transient peak | |---|---|---| | 256 (was) | 91.4 ms | +24 MiB | | 1024 | 22.9 ms | +31 MiB | | 4096 (now) | 5.8 ms | +56 MiB | The decode now takes its own `_DEFAULT_DECODE_CHUNK_SIZE`, threaded through `fake_quantize_with_cache`. The encode bounds are untouched, so the memory ceiling stays where it was aimed. The seven-iteration sign-parity loop in `dequantize_iq2_xs` is also folded into three XOR steps, off the same per-forward path. Per case in `tests/examples/hf_ptq` on 2xH100: | | before | after | |---|---|---| | IQ1_S | 259.65 s | 94.64 s | | IQ2_XS | 1081.05 s | 102.77 s | Both now fit the 300s `tests/examples` default, so the cases carry no explicit timeout. ## Integration contracts - `block_sizes: {-1: 256}` records the native packed-block contract; it does not drive the GGML fake-quant scale search. The export path reads `TensorQuantizer.block_sizes[-1]` through `get_weight_block_size`, validates it against the format block size, and records `group_size: 256` in checkpoint metadata. - The model presets intentionally expose `--qformat iq1_s` and `--qformat iq2_xs` in `hf_ptq`. ## Testing - `tests/unit/torch/quantization/ -k 'ggml or iq1 or iq2'` — 52 passed, 1 skipped - `tests/unit/recipe/test_presets.py` — both shipped IQ recipes load with the expected backend and no unused search option - `tests/examples/hf_ptq/test_llm_ptq.py -k 'iq1_s or iq2_xs'` — 2 passed on 2×H100 (TinyLlama, both formats end to end through export), 3m17s for the pair - pre-commit hooks pass on all changed files - larger-model sanity check outside CI: Qwen3.8-27B (2256 quantizers) quantizes and exports with `general/ptq/iq1_s` on 2xH100 in 66m46s, export itself 192s --------- Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5bb7343592 |
[https://nvbugspro.nvidia.com/bug/6778095] Fix fused P-QDQ to respect disabled quantization during calibration (#2434)
### What does this PR do?
Type of change: Bug fix
During max calibration, `enable_stats_collection()` calls
`disable_quant()`, which sets `_if_quant=False`. The fused causal P-QDQ
attention paths bypass `TensorQuantizer.forward()` and previously
selected the Triton/Kitchen path from the configured enabled state
alone, so P quant-dequant could still execute while quantization was
inactive.
This change:
- enters the Triton P-QDQ path only when `p_bmm_quantizer._if_quant` is
true;
- bypasses all fused P-QDQ paths when the P quantizer is disabled or
quantization is inactive;
- adds focused parameterized regression tests covering Triton and
Kitchen dispatch across enabled, `disable()`, and `disable_quant()`
states.
### Usage
N/A. This restores the existing `disable_quant()` contract and does not
introduce a new API.
### Testing
- `pytest
tests/unit/torch/quantization/plugins/test_attention_quant.py`: 14
passed
- Targeted pre-commit checks passed
- Kitchen coverage verifies the full `{disable(), disable_quant()} ×
{lazy, already initialized}` dispatch matrix remains bypassed while
inactive and resumes after re-enabling
- Reproduced with the exact same ModelOpt 0.47.0rc1 wheel on both sides
on B300: [module regression build
#121](http://dlswqa-nas.nvidia.com:18880/view/yiguo/job/modelopt-quant-module/121/)
- Controlled B300 isolation passed when only the P-BMM quantizers were
hard-disabled during calibration, and also passed when the existing
single Triton attention configuration was forced
- Fixed-code B300 validation passed in three independent runs: [Jenkins
build
123](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/123/),
[Jenkins build
124](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/124/),
and [Jenkins build
125](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/125/).
Jenkins build 123 compared all 178 module outputs byte-identically and
found 0/377 `amax` and 0/377 `scale` changes.
The confirmed impact is incorrect calibration behavior plus unstable
quantizer state and module outputs; downstream benchmark accuracy impact
has not been established.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A
### Additional Information
The failure was isolated to fused P-QDQ runtime-state dispatch.
Quantizer topology and configuration were identical in the failing
same-wheel comparison.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Bug Fixes**
- Corrected attention dispatch so disabled or inactive quantization uses
the original attention implementation.
- Preserved optimized quantized attention when quantization is enabled.
- Ensured attention masks remain unchanged when using the original
attention implementation.
- Improved fallback behavior for fused attention paths, including
correct initialization and reuse when quantization is re-enabled.
- **Tests**
- Added coverage for enabled and disabled quantization states,
attention-mask handling, fallback selection, and result consistency.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: yingguo-trt <244492186+yingguo-trt@users.noreply.github.com>
|
||
|
|
7a35cada39 |
Add guidance on sizing and splitting PRs to AGENTS.md (#2494)
### What does this PR do? Type of change: documentation Adds a `## Sizing and splitting PRs` section to `AGENTS.md` (symlinked as `CLAUDE.md`) so AI-assisted work stops producing one giant PR that nobody wants to review. The new guidance tells the agent to: - Keep each PR that goes up for review under ~500 changed lines of source, and check the size before opening. - Propose the split *before* opening an oversized PR rather than after. - Split on file/directory/module boundaries first and fall back to feature boundaries (enabling refactor first, then one PR per behavior it unlocks). - Keep the series acyclic and linearly ordered — no circular dependencies between sub-PRs — and state the merge order. - Prefix sub-PR titles with `[x/N]` so reviewers know the PR is one slice of a planned split, and link the siblings. - Make every sub-PR stand on its own: it builds, carries unit tests for the code it introduces, and passes CI without the later PRs. - Optionally submit the whole change as a reference-only draft PR for the big picture, cross-linked with the sub-PRs. ### Usage N/A — no API or flag change. ### Testing `pre-commit run --files AGENTS.md` (markdownlint and the other applicable hooks pass). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A <!-- documentation-only change --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- agent instructions only, not user-facing --> - Did you get Claude approval on this PR?: ❌ <!-- will run /claude review if reviewers want it --> ### Additional Information This PR is itself well under the new budget (26 added lines in one file), so no split applies. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Documentation** - Updated pull request sizing guidance to allow oversized changes when they cannot be meaningfully split. - Clarified that draft aggregate pull requests are optional and intended for reference only when splitting would reduce clarity. - Added examples covering self-contained changes and new models or backends without a functional intermediate state. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ee5c256204 |
Noeyy/fix bug 6701777 (#2402)
### What does this PR do? Type of change: Bug fix: 6701777 Regression source: "[OMNIML-3349] Add FP8 MHA quantization support for HuggingFace ViT" (#1289), merged into 0.44.0rc3 via the batch cherry-pick #1350. This PR: 1. Registers nn.LayerNorm as a QuantModule for the first time (modelopt/torch/quantization/nn/modules/quant_layernorm.py), intended to let FP8_DEFAULT_CFG's BMM input / LayerNorm output quantizer rules apply to ViT. 2. Removes the prior forced Cast-alignment logic in export_onnx.py that used to normalize Q/DQ node dtypes to trt_high_precision_dtype. Root Cause: Once nn.LayerNorm became a registered QuantModule, these wildcards started unintentionally matching norm1.norm inside FLUX's AdaLayerNormZero block — an elementwise_affine=False LayerNorm with no learnable weight/bias. Its input got routed through NVFP4 Q/DQ (emitted as Float32) while its synthesized affine scale remained native BFloat16, producing the dtype mismatch. Chosen fix: Explicitly exclude nn.LayerNorm from the diffusers NVFP4 presets rather than touching the global QuantModuleRegistry (which ViT FP8 MHA still needs). Add, in both modelopt_recipes/configs/ptq/presets/diffusers/nvfp4.yaml and nvfp4_fp8_mha.yaml, after the existing weight/input wildcard rules (list order matters — later entries override earlier ones): - parent_class: 'nn.LayerNorm' quantizer_name: '*' enable: false ### Usage ``` python examples/diffusers/quantization/quantize.py --model flux-dev --format fp4 --batch-size 2 --percentile 1.0 --alpha 0.8 --quant-algo max --n-steps 20 --quantized-torch-ckpt-save-path /tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4.pt --onnx-dir /tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4 --collect-method default --calib-size 128 --model-dtype BFloat16 --trt-high-precision-dtype BFloat16 trtexec --onnx=/tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4/model.onnx --builderOptimizationLevel=4 --saveEngine=/tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4/model.plan --stronglyTyped --minShapes=hidden_states:1x1024x64,img_ids:1024x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1 --optShapes=hidden_states:1x4096x64,img_ids:4096x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1 --maxShapes=hidden_states:1x4096x64,img_ids:4096x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1 ``` ### Testing The above test commands. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved Diffusers NVFP4 and NVFP4/FP8 MHA quantization presets by excluding LayerNorm modules from quantization. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
d0142c9dca |
Reuse a whole recipe via $import, deprecate recipe_type, and start the published-checkpoint backfill with two aliases (#2376)
### What does this PR do?
**Type of change:** New feature (recipe loading) + one bug fix
Two things, the second built on the first:
1. **A recipe can now reuse another recipe whole.** A top-level
`$import` brings in the imported recipe's entire body; keys given
alongside it override the imported ones. `metadata.recipe_type` becomes
optional and is deprecated along the way.
2. **The deprecated `recipe_type` is swept out of every shipped recipe,
and the checkpoint backfill starts with two published checkpoints
recorded as aliases** that reuse a portable recipe wholesale — the first
users of the alias mechanism — plus a fix to two existing Nemotron NVFP4
recipes.
#### Declaring what kind of recipe a file is
`load_recipe` read `metadata.recipe_type` out of the raw YAML *before*
resolving imports, because it needs the schema class to hand to
`load_config`. That made the field impossible to inherit, so a recipe
reusing another had to restate a line it could only have copied.
It is now optional, and the loader takes the first of these that
answers:
1. a `# modelopt-schema:` comment naming the recipe's schema class,
2. `metadata.recipe_type` — **deprecated**; still read and still
honoured, so a recipe outside this repo keeps working unchanged,
3. the recipe it delegates to via a top-level `$import`.
Whatever a recipe *does* state must be true, in both directions. A
schema comment contradicting a `recipe_type` is rejected, and so is a
recipe importing a different kind of recipe — that used to surface as
whatever pydantic made of, say, an `eagle` section spliced into a PTQ
schema. The concrete recipe classes carry a `RECIPE_TYPE` ClassVar as
the single source of truth.
Only a recipe that another file **imports** needs the schema comment —
that is what `$import` resolution requires to validate the payload. The
sweep here drops `metadata.recipe_type` from all 78 shipped recipes that
carried it and gives the imported ones a `# modelopt-schema:` comment
instead, so nothing in-tree depends on the deprecated field.
A directory recipe's `metadata.yml` resolves its kind the same way —
schema comment first, `recipe_type` as the fallback — it just has no
`$import` to delegate through, since a directory recipe has no body of
its own to hand off. (Follow-up commit, after this PR's initial review:
it originally still required `recipe_type` unconditionally, the one
place the deprecation didn't reach.)
#### Checkpoint aliases
Two checkpoints NVIDIA has published in quantized form use a scheme a
portable recipe already produces, with no checkpoint-specific deviation,
so each is recorded as a thin **alias** (top-level `$import`, overriding
only `metadata`) at its own model-hub path -- the *source* checkpoint's
path, not the published quantized one's:
-
**`models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast`**
delegates to `general/ptq/nvfp4_experts_only_mse-kv_fp8_cast` —
expert-only NVFP4 (MSE static weights, dynamic inputs) with an FP8 KV
cache in cast mode — published as `nvidia/Kimi-K2.6-NVFP4`.
-
**`models/Qwen/Qwen3.5-397B-A17B/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8`**
delegates to the `qwen3_5_moe` architecture recipe
`model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4
(MSE static weights) on the routed experts, ModelOpt-default FP8
elsewhere, FP8 KV cache — published as
`nvidia/Qwen3.5-397B-A17B-NVFP4-V2`.
(Follow-up commit, after this PR's initial review: the Qwen entry
originally lived at `models/nvidia/Qwen3.5-397B-A17B/` -- nvidia is the
*published* checkpoint's org, not Qwen3.5-397B-A17B's own. Moved to
match the source model's actual hub path, same as the Kimi-K2.6 entry
above.)
Editing the base recipe changes every alias that points at it; nothing
is duplicated.
#### One fix
- **The Nemotron-3 Super and Ultra NVFP4 recipes** quantized the MTP
block on the **Megatron-Core** path, where it is a live `model.mtp`
submodule their broad `*mixer.*` patterns matched into, contrary to
their own descriptions. They now disable `mtp.*` explicitly. Hugging
Face runs were unaffected — `NemotronHPreTrainedModel` sets
`_keys_to_ignore_on_load_unexpected = [r"mtp.*"]` and builds no MTP
module.
### Usage
A checkpoint alias resolves through `--recipe` to the recipe it
delegates to:
```bash
python examples/hf_ptq/hf_ptq.py \
--pyt_ckpt_path <checkpoint> \
--recipe models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast \
--export_path <output>
```
A recipe that reuses another whole — the shape the aliases use:
```yaml
imports:
base: general/ptq/nvfp4_experts_only_mse-kv_fp8_cast
$import: base
metadata:
description: What this checkpoint uses the base recipe for.
```
### Testing
- **`tests/unit/recipe/test_loader.py`** — 28 new cases covering
whole-recipe reuse with no `metadata` at all; kind resolution from each
of the three sources, from a delegation chain and from a `$import` list;
a delegation cycle failing with `ValueError` rather than recursing;
`peek_declared_schema` including a comment placed below the first YAML
line; `recipe_type` being optional, filled per class, and rejected when
it contradicts; a directory recipe resolving its kind from a schema
comment the same way, rejecting a comment/`recipe_type` disagreement,
and still requiring one or the other; and delegating across kinds being
an error.
- **`tests/unit/recipe/test_recipe_docs.py`** — the
model-specific-recipe check now also covers the two new alias folders,
which must be listed in `ptq.md` like every other
`models/<org>/<model_id>` entry.
- **Recipe validation** (`tools/precommit/check_modelopt_recipes.py`)
and **`pre-commit`** pass on the changed files. The full
`tests/unit/recipe/` suite is left to CI — a broken `transformer_engine`
in the local dev venv keeps the `mtq.quantize`-based cases from running
there.
Not covered: **numerics**. Nothing here asserts accuracy, or that
running one of these recipes reproduces a released checkpoint's weights.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `metadata.recipe_type` is
still read and honoured for recipes outside this repo, the schema
comments are inert for direct loads, and the loader change only relaxes
a check.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — two feature entries, one deprecation, and one bug fix under 0.48.0.
- Did you get Claude approval on this PR?: ❌ — not yet run.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Recipes can delegate configurations, support checkpoint aliases, and
apply local metadata overrides.
* Recipe types can be inferred from schema declarations or delegated
recipes, with stronger consistency validation.
* Added unquantized KV-cache options, layerwise export, broader operator
calibration, and new PTQ examples.
* Added checkpoint-specific recipes and MLflow experiment references.
* **Bug Fixes**
* Improved ONNX calibration, FSDP2 export, and fused-MoE quantization
handling.
* Nemotron-3 recipes keep MTP blocks in BF16.
* **Documentation**
* Expanded guidance for aliases, delegation, schema declarations, and
recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
051d6adb20 |
Require PTQ recipe guidance before selecting quantization candidates (#2478)
### What does this PR do? Type of change: documentation Updates `quant-recipe-search` to require reading `modelopt_recipes/ptq.md` before selecting initial or subsequent quantization candidates. The skill uses the guide to inform quantization scope, KV-cache scheme, and calibration method. It requires inspecting candidate YAMLs and supporting configs, citing relevant guide sections, and explaining deviations. Existing coverage, runtime compatibility, and evaluation checks remain required. ### Usage Invoke `quant-recipe-search` as usual. The skill reads the guide from the ModelOpt source checkout. ### Testing - Passed a local Codex skill smoke test recommending an initial NVFP4 candidate for Qwen/Qwen3-8B targeting high-concurrency inference on Blackwell. - The test prompt named the skill without explicitly requesting that the agent read ptq.md. - Verified successful reads of the updated skill, ptq.md, the selected recipe YAML, and supporting configs before the final recommendation. - The agent selected general/ptq/nvfp4_mlp_only-kv_fp8_cast and cited relevant guide sections. It explicitly left coverage, accuracy, and throughput pending validation. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ ### Additional Information This change places the reading requirement in the shared recipe-search skill so consumers receive it directly. Consumers that pin the ModelOpt plugin to a commit must update their pin to adopt it. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Documentation** - Updated the quantization recipe search workflow to require reviewing the PTQ reference documentation before selecting a recipe. - Clarified that available quantization schemes and model-specific exceptions should be considered during recipe evaluation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Zhehao Hu <zhehaoh@nvidia.com> |
||
|
|
f1abc75626 |
Carry unplaced checkpoint weights using the loader's accounting, replacing MTP name-matching (#2427)
### What does this PR do?
Type of change: bug fix + new tests
Replaces `hf_ptq`'s name-based MTP detection with the Transformers
loader's own accounting of what it could not place.
#### The problem
`load_mtp_weights` found MTP weights by name — `"mtp" in key` plus
config-derived layer indices — backed by a support matrix of three
storage conventions (GLM-5.1 inlined, GLM-4.7 standalone file,
Qwen3-Next tail shard). Every architecture that spells it differently is
a silent miss, and a miss means a checkpoint exported without a
component its own config still advertises. The failure is quiet on both
sides: Transformers drops unexpected keys, and vLLM's weight loading is
pull-based, so a missing MTP produces no warning at all.
#### The fix
`from_pretrained(..., output_loading_info=True)` reports
`unexpected_keys` — *"keys that are found in the checkpoints, but not
expected in the model's architecture"* — which is exactly the carry-over
set, derived structurally rather than by naming, and already accounting
for on-the-fly key conversion that a set re-derived afterwards would
have to replay. Record those keys at load; carry them at export. Weights
the loader *did* place go through the normal export path unchanged.
#### Two mechanisms, disjoint by construction
Measured against a real `from_pretrained`, an off-index file reports
**nothing**: the loader opens only shards named in
`model.safetensors.index.json`, so it never saw those tensors to call
them unexpected. Those are sidecars, not weights — untouched by
quantization and absent from the export — so they are **copied
verbatim**, which costs no host memory, preserves the bytes and file
layout, and leaves the filename a consumer looks for where it was.
| on-disk layout | mechanism |
|---|---|
| inlined layer past `num_hidden_layers` (GLM-5.1, DeepSeek-V3) |
`extra_state_dict` |
| indexed `mtp.*` tail shard (Qwen3-Next) | `extra_state_dict` |
| standalone off-index file (GLM-4.7) | **copied whole** |
A tensor is never both copied and carried; that would export it twice,
and a test asserts it.
#### Removed as redundant
`load_mtp_weights`, `mtp_layer_prefixes_from_checkpoint`,
`get_inlined_mtp_prefixes`, `_load_tensors_matching`,
`_apply_to_model_state_dict`, `_keys_to_prefixes`, `_add_mtp_exclusions`
and its three call sites, the pre-quantization `enable: False` entries
`hf_ptq` appended to the recipe's `quant_cfg`, and the dead
`_mtp_layer_prefixes` fallback in `_get_num_nextn_predict_layers`.
#### Two deliberate behavioural changes
**MTP now follows the recipe** instead of being force-excluded by the
script — which is what `examples/megatron_bridge` already does; it has
no MTP-specific code at all. Recipes importing
`configs/ptq/units/default_disabled_quantizers` still disable `mtp.*`,
so their behaviour is unchanged; a recipe omitting that unit will now
quantize an MTP the model actually built.
**`quantization_config.ignore` can no longer claim a layer is
unquantized that the export in fact quantized.** That contradiction came
from `_add_mtp_exclusions` firing off a model attribute with no
cross-check against quantizer state.
### Usage
No API change for callers of `export_hf_checkpoint`. Within
`examples/hf_ptq`, model loading now goes through a wrapper that records
the loader's accounting:
```python
model, loading_info = auto_class.from_pretrained(ckpt_path, output_loading_info=True, **kwargs)
record_unplaced_source_keys(model, ckpt_path, loading_info.get("unexpected_keys"))
```
### Testing
`tests/examples/hf_ptq/test_carry_over_layouts.py` — 8 tests driving a
**real** `from_pretrained` against a tiny model, covering each of the
three conventions above plus an auxiliary (non-MTP) tower, two layouts
at once, a checkpoint with nothing stray, and that indexed shards are
never copied. CPU-only: the mechanism is bookkeeping during load, so a
GPU adds nothing; the export side already has GPU coverage in
`tests/gpu/torch/export/test_export_carry_over.py`.
The six `load_mtp_weights` tests are replaced by three on the recording
path, and the `get_model` test doubles now model `output_loading_info`
the way Transformers does.
All passing: 8 layout tests, 82 in the surrounding `examples/hf_ptq`
suite. `ruff` findings at parity with `main` on every changed file.
Files named like a main weight shard are excluded from the off-index set
whatever the index says — a fixture with an empty `weight_map` would
otherwise have made the source weights look like sidecars and copied
them into an export beside the quantized ones.
### The algorithm: which weights get carried, and how
Two disjoint sets of source weights reach the export without passing
through quantization. They are distinguished by **what the loader did
with the file**, and that difference decides both how each is found and
how each is moved.
**Set 1 — unplaced weights.** The loader opened the file and read the
tensor, but the model had no parameter for it, so Transformers reports
it in `unexpected_keys`. An MTP head the recipe did not quantize is the
common case. Moved as **tensors**: located in whichever shard holds
them, read, and merged into the exporter's `extra_state_dict`.
**Set 2 — off-index sidecars.** The index never names the file, so the
loader never opened it and never had the chance to call anything
unexpected. GLM-4.7 keeps its MTP head in a standalone `mtp.safetensors`
exactly this way. Moved as **files**: copied byte for byte, so no host
memory is spent re-serialising tensors the export does not otherwise
touch.
#### The index is not an inventory of the checkpoint
This is the part that is easy to get wrong, and it cost a silent
data-loss bug during review.
`model.safetensors.index.json` selects which **files** the loader opens
— not which **tensors** it sees. Within a file it opens, Transformers
enumerates every tensor present and reports the unexpected ones.
Verified by experiment against transformers 5.3.0:
| case | reported in `unexpected_keys`? |
|---|---|
| key absent from the index, in a shard the index names for *other*
tensors | **yes** |
| key in a file the index never names (`mtp.safetensors`) | **no** — the
file is never opened |
So a tensor missing from `weight_map` but sitting inside a main shard is
**set 1, not set 2**. An MTP head stored that way is reported, recorded
— and was then silently dropped, because resolution went through
`weight_map`, which by construction has no entry for it. The
`--vllm_fakequant_export` guard shared that lookup, so the check written
to refuse exports that drop weights stayed silent in exactly the case it
existed for.
`locate_source_keys` now resolves through the index first (free for
everything it lists) and header-scans the shards only for the leftovers
— names, never tensor data — warning when a key is in no file at all.
The carry and the guard share it, so they cannot disagree again.
#### Flow
1. **At load.** `record_unplaced_source_keys` stores Transformers' own
`unexpected_keys` on the model (`_modelopt_unplaced_source_keys`) plus
the resolved local checkpoint path. The question asked is "does the
model have a parameter for this key", never "is this an MTP head" — so
the mechanism is architecture-agnostic.
2. **At export, before dispatch.** `read_unplaced_weights` resolves each
recorded key to its shard and reads the tensors, merging them into
`extra_state_dict`. An explicitly passed `extra_state_dict` wins on a
name clash: a caller naming a tensor is more specific than our
inference.
3. **Rank behaviour.** Only the rank that writes `extra_state_dict`
reads the bytes — the FSDP2 writer emits it from rank 0 alone, so a full
read on every rank would be host memory spent and discarded (a
DeepSeek-V3-class MTP head is 10 GB+ in bf16). The **key list** is still
resolved on every rank, because `get_quant_config` runs per rank and the
configs must agree.
4. **Recording what was written.** `export_hf_checkpoint` records
`_modelopt_carried_over_names` — the union of carried tensors and the
off-index sidecars' tensor names — before `get_quant_config` runs,
because that is the first point that knows what was *written* rather
than what was merely unplaced.
5. **Exclusions.** Both sets must reach `quantization_config.ignore`, or
a deployment framework reads the top-level `quant_algo` and tries to
load an original-precision weight as a quantized one (the NVBug 5718750
class). `seed_carried_over_exclusions` is the single path for this,
called by `get_quant_config` after its per-layer pass and again by the
layerwise exporter from `finalize()` — which snapshots its config during
`bind()`, while calibration is still running, so it cannot see the
carried set any earlier.
#### What is deliberately excluded
- **Files that re-ship indexed weights.** Mistral's
`consolidated.safetensors` is a second full copy of the model, and
PEFT's `adapter_model.safetensors` is an adapter. Both are off-index,
and copying either would put unquantized weights beside the quantized
ones — vLLM's mistral load-format looks for `consolidated.safetensors`
by name, so it is not inert. Caught by name for the known conventions
and by tensor-name overlap for the rest.
- **Files named like a main weight shard**, whatever the index says, so
a broken or partial index cannot make the real weights look like
sidecars.
- **Symlinks are *not* excluded.** A Hugging Face snapshot stores every
file as a symlink into a sibling `blobs/`, so refusing links would drop
the sidecar of every hub-downloaded checkpoint.
`resolve_checkpoint_file` checks where the link *lands* — regular file,
inside the checkpoint dir or its blob root — rather than whether it is a
link.
- **Buffers Transformers recomputes.** `*.inv_freq` is skipped by the
fake-quant guard even when a shard provides it: older
Llama/Mistral-lineage conversions do list it in the index, and refusing
an export over it would reject checkpoints that export correctly today.
### How a carried, never-quantized MTP head reaches
`quantization_config.ignore`
Raised in review: `_add_mtp_exclusions` is gone, and an unplaced weight
has no module, so
`get_quant_config` walks right past it. That was a real gap, not just a
documentation one —
fixed here.
The export writes a carried MTP head in its original precision. If it is
absent from
`exclude_modules`, a deployment framework reads the top-level
`quant_algo` and tries to load
`eh_proj` as an FP8/NVFP4 weight — the same class of failure as NVBug
5718750, where a
`transformers>=5.0` MoE router was written in BF16 but never excluded.
The two cases now differ only in *why* the module is invisible to the
quantizer walk:
| | why invisible | handled by |
|---|---|---|
| MoE router (tf≥5.0) | module exists, never gets a quantizer |
`_get_unquantized_moe_router_names` |
| carried weight | no module at all in the live model |
`_get_carried_over_module_names` |
`_get_carried_over_module_names` reads the keys the loader recorded as
unplaced
(`_modelopt_unplaced_source_keys`), strips the trailing parameter name —
a state-dict key is
`<module path>.<parameter>` — and dedupes. Those names are seeded into
`layer_config_dict` as
`QUANTIZATION_NONE`, exactly as the router pass does, so they flow
through
`process_layer_quant_config` into `exclude_modules`, which
`convert_hf_config` emits as
`quantization_config.ignore`.
An MTP head the recipe *does* quantize is unaffected: it has a module,
is loaded normally, is
never in the unplaced set, and is reported as quantized. The
backward-breaking note above still
holds — MTP follows the recipe instead of being force-excluded — but a
head that ends up carried
rather than quantized is no longer silently missing from `ignore`.
Covered by `test_carried_over_weights_are_excluded_from_quantization`
and
`test_carried_over_module_names_strip_parameter_and_dedup`.
`--vllm_fakequant_export` does not carry unplaced weights; it now raises
rather than writing a
checkpoint quietly missing them.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ — MTP layers now follow the
recipe rather than being force-excluded, and
`quantization_config.ignore` no longer lists layers the export may have
quantized (carried, never-quantized weights are still listed -- see
below). Shipped recipes are unaffected; see the Changelog entry.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 Backward Breaking Changes.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
Draft: the behavioural change to MTP quantization is the part most worth
a second opinion — it aligns `hf_ptq` with `megatron_bridge`, which
special-cases nothing.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Quantization supports enabled operators outside transformer layers,
including language-model heads.
* Exports preserve checkpoint weights not loaded into the quantized
model.
* Additional safetensors sidecar files are copied unchanged into
exported checkpoints.
* Unquantized auxiliary components, such as vision layers, remain
available in exported models.
* **Behavior Changes**
* Local recipe files take precedence over built-in recipes.
* Legacy architecture-specific recipe paths remain supported with
warnings.
* MTP-specific export exclusions are no longer applied.
* **Bug Fixes**
* Preserved weights are no longer incorrectly reported as unquantized.
* Incomplete exports are rejected when source weights cannot be placed
or preserved.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
fc4c40fcbe |
Fix HF export crash when a dynamic-block quantizer has zero amax (#2438)
### What does this PR do? Type of change: Bug fix `TensorQuantizer.export_amax()` early-returns `self.amax` unsanitized for dynamic-block quantizers, while the static path immediately below it has always substituted `maxbound` for zero/NaN entries. The `nvfp4` numerics unit sets `type: dynamic`, so a recipe that applies it to an *activation* quantizer — e.g. `general/ptq/nvfp4_mlp_only-kv_fp8_cast`, which targets `*mlp*input_quantizer` — feeds a raw `0.0` into `NVFP4QTensor.get_activation_scaling_factor`, whose assert aborts the entire export: ``` AssertionError: Failed to export module 'model.language_model.layers.37.mlp.gate_proj' (type=QuantLinear): activation scaling factor 0.0 not positive. ``` Calibration leaves `amax` at 0 whenever a layer — or an unrouted MoE expert — saw only zeros, so one dead layer costs the whole run at the final export step. This factors the substitution into `_sanitize_export_amax()` and calls it from both branches. Two details beyond de-duplication: - **Branch-free, so it survives a meta `amax`.** `torch.where` + `nan_to_num` both have meta kernels; `bool()` on a meta tensor raises. The layerwise and streaming export flows carry meta `amax` — `validate_attr` short-circuits on `is_meta` for exactly that reason — so only the warning is gated on a materialized tensor. - **No longer mutates calibrated state.** The old in-place `amax[amax == 0] = ...` wrote through a view of `self._amax`; `torch.where` returns a fresh tensor, so that hazard disappears. - **Warns, with a count.** The fix turns a loud failure into a silent one, and a zero amax means calibration never activated that layer — worth surfacing rather than papering over. The message reports how many entries were substituted, since per-location dedup otherwise collapses many dead experts into one uninformative message. A healthy model emits none. Scope: only the activation path is data-dependent and reachable this way. Weight-side `_amax` uses are left alone, since a weight amax of 0 would require an all-zero weight matrix. **Knowingly left as follow-up:** `export/quant_utils.py::get_scaling_factor` discards the sanitized `amax` when `num_bits == (2, 1)` and recomputes via `get_weights_scaling_factor_2_from_quantizer`, which reads `weight_quantizer._amax` raw — so a dynamic-NVFP4 *input* quantizer on a module whose *weight* quantizer is a different format (or disabled) can still trip `assert torch.all(scaling_factor > 0)`. Format dispatch is weight-driven, so the reported recipe does not reach that branch; fixing it properly changes a signature shared with the weight-side callers and is out of scope here. Not a regression. The dynamic early return, the `type: dynamic` numerics unit, and the recipe that combines them all ship in released 0.46.0 / 0.46.1. ### Usage No new or changed API. Exports that previously aborted now complete and warn: ```python # Recipe applies dynamic NVFP4 to *mlp*input_quantizer; layer 37 never activated during calibration. mtq.quantize(model, quant_cfg, forward_loop) export_hf_checkpoint(model, export_dir=out) # before: AssertionError; now: exports + UserWarning ``` ### Testing - New `test_amax_export_unusable_amax`, parametrized over zero and NaN, covering the dynamic-NVFP4 and static per-tensor configs; asserts the exported scale is positive and that export leaves the calibrated `amax` untouched. Plus `test_amax_export_meta_amax`, pinning that a meta `amax` survives export rather than raising. Both run on CPU and CUDA via the shared tester. - `tests/unit/torch/quantization/test_tensor_quantizer_cpu.py` — 40 passed. `tests/gpu/torch/quantization/test_tensor_quantizer_cuda.py` — 40 passed (GB300). - End-to-end repro on GB300, small Llama with one MLP fed all-zero activations under `general/ptq/nvfp4_mlp_only-kv_fp8_cast`: dead layer `export_amax()` `0.0` → `6.0`, live layer unchanged at `3.921875`, and `export_hf_checkpoint` goes from the `AssertionError` above to writing `model.safetensors`. - Full `examples/hf_ptq/hf_ptq.py` with the reported recipe and flags on a healthy model (Qwen3-0.6B): exits 0 and writes the checkpoint, confirming the normal path is unaffected. - `pre-commit run` clean on all changed files. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ✅ — `/claude review` run; its one IMPORTANT finding (meta-tensor regression) and both SUGGESTIONs addressed or answered in 251f2e3d ### Additional Information Fixes NVBug 6768300, reported against 0.47.0rc1 on GB200. The reporter also notes it passed on 0.47.0rc0; that is not explained by code — `git diff 0.47.0rc0..0.47.0rc1` touches `export/quant_utils.py` only in `get_kv_cache_scaling_factor` (new `clamp_fp8_scales` argument whose default preserves the old behaviour) and the INT4-AWQ packing path, neither of which is on the dense-HF NVFP4 activation-scale path. Whether `amax` lands on exactly 0 is calibration/model-state dependent, which is what makes it look version-flaky. Worth flagging separately: in the reported log the **pre-PTQ** sample output is already gibberish, so that BF16 checkpoint looks broken independently of quantization. This change stops the crash, but such a run will now export a valid-but-garbage checkpoint — the new warning is the signal to investigate. Suggest the `cherry-pick-0.47.0` label so this lands in the ongoing release. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed Hugging Face checkpoint export when dynamic-block quantizers have zero or invalid calibration scales. * Exports now use a positive fallback scale and issue a warning instead of failing when applicable. * Export operations no longer modify the original calibrated quantizer state. * Meta-device exports remain non-erroring and preserve device placement. * **Tests** * Added coverage for zero- and invalid-scale exports across dynamic and static quantization modes. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Yue <yueshen@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b311c054de |
Fix hybrid stack spec serialization in Megatron-Bridge checkpoints (#2452)
### What does this PR do?
Type of change: Bug fix
Hybrid (e.g. Nemotron-H) checkpoints saved by the
`examples/megatron_bridge` scripts could not be reloaded or exported to
HuggingFace:
TypeError: MLPSubmodules.__init__() missing 2 required positional
arguments: 'linear_fc1' and 'linear_fc2'
**Root cause.** Megatron-LM's YAML writer represents a
`functools.partial` via `_partial_representer`, which passes each
keyword value through `represent_data`. A dataclass instance has no
representer, so it falls through to `_safe_object_representer`, which
emits only `{_target_, _call_}` and drops every field. The default
hybrid stack spec builds its dense-MLP and MoE layers as exactly such
partials, and `set_moe_expert_layout()` stored the *built* `ModuleSpec`
on the provider — which is serialized into every checkpoint's
`run_config.yaml`. So `MLPSubmodules` / `MoESubmodules` were written
with no fields at all. Both export paths hit it: `convert.sh` via
`from_auto_config`, and `export_distilled_megatron_to_hf.py` via
`export_ckpt → load_megatron_model`. Every hybrid provider is affected,
dense or MoE.
**Fix.** `set_moe_expert_layout()` stores a named, zero-argument
*factory* instead. The provider already calls a callable spec at build
time (`_resolve_hybrid_stack_spec`), so model construction is unchanged
— only the serialized form differs:
```yaml
hybrid_stack_spec:
_call_: false
_target_: megatron.bridge.models.hybrid.hybrid_provider.transformer_engine_hybrid_stack_spec
```
The grouped-GEMM factory is Megatron-Bridge's own
`transformer_engine_hybrid_stack_spec`, so stock tooling
(`scripts/conversion/convert.sh`) resolves it without importing
ModelOpt.
**Known limitation — the two layouts are not symmetric.** The
SequentialMLP layout has no bridge-side equivalent (the upstream TE
hybrid spec hardcodes `TEGroupedMLP`), so it serializes a ModelOpt
target, which `instantiate` only accepts in a process that has imported
`mbridge.py` and thereby run `register_allowed_target_prefix`. A
SequentialMLP hybrid checkpoint therefore converts through the ModelOpt
entrypoints but not through stock `convert.sh`, where it fails on the
disallowed prefix instead of on `MLPSubmodules` — no regression, but
that path stays broken for this one layout. The reach is narrow:
`use_moe_grouped_gemm()` returns True for any architecture with a
grouped-expert export rule, NemotronH included, so SequentialMLP
requires an explicit `--no_moe_grouped_gemm`. Closing it properly needs
an upstream `moe_grouped_gemm`-aware factory in Megatron-Bridge.
Both spec builders also move out of `nas/plugins/megatron.py`, which
never used them, into a new `utils/plugins/megatron_layer_specs.py`
beside the other Megatron-Core-only helpers. Not into `mbridge.py`: that
module needs `megatron.bridge`, while `get_te_hybrid_stack_spec` is
reached by 16 test files through
`tests/_test_utils/torch/megatron/models.py`, which is bridge-free.
The underlying defect is upstream in
`megatron/training/config/yaml_utils.py`; this only stops ModelOpt from
stepping on it, so it is worth a separate Megatron-LM issue.
### Usage
No API change — hybrid checkpoints saved after this fix convert with the
existing commands:
```bash
torchrun --nproc_per_node 1 examples/megatron_bridge/export_distilled_megatron_to_hf.py \
--student_hf_path <student_hf_model_or_path> \
--megatron_path <distill_out>/checkpoints \
--hf_export_path <hf_out> \
--export_iterations all
```
### Testing
Verified in `nemo:26.08` (megatron-core 0.19.0) against a 30B-A3B
Nemotron-3.5-Lightning pruned+distilled run:
- **Round trip, both MoE layouts.** Ran `set_moe_expert_layout` on a
real `HybridModelProvider`, dumped it through `dump_dataclass_to_yaml`
(the writer used for `run_config.yaml`), reloaded via `instantiate`,
resolved. `moe_grouped_gemm=True` →
`TELayerNormColumnParallelLinear`/`TERowParallelLinear` +
`TEGroupedMLP`; `False` → same MLP + `SequentialMLP`. The field stays
callable after `finalize()` and `_resolve_hybrid_stack_spec()`, so a
saved config cannot regress.
- Applying the equivalent `run_config.yaml` fix to 32 iteration
checkpoints: all 32 rebuild the provider (52 layers, hidden 2304, 104
experts) with populated `MLPSubmodules` / `MoESubmodules`.
- **End-to-end exports**, 6 iterations, all rc 0, each producing exactly
the source model's 5139 tensor keys (0 missing, 0 extra), 9 shards /
41.5 GiB, all weights finite, drift from the base rising monotonically
with iteration (lm_head 0.030 → 0.092). Covered `convert.sh` CPU,
`convert.sh` GPU (4×GB300, TP=4), and
`export_distilled_megatron_to_hf.py`. Same iteration and wrapper: CPU
123 s vs GPU 134 s — GPU is not faster, since with TP=4 each rank still
builds 20.9 B of 22.3 B params and the cost is I/O plus CPU-side
conversion.
- `ruff check` / `ruff format --check` passed on the source files before
the module move.
**Not yet run:**
`tests/gpu_megatron/torch/utils/plugins/test_mbridge.py` (added here) —
the GPU allocation expired. It asserts the round-trip property verified
manually above, but its `HybridModelProvider(num_layers=2,
hidden_size=64, num_attention_heads=4)` construction is unverified. The
module move is verified only by reference grep and syntax check, so
please also run one test that uses
`tests/_test_utils/torch/megatron/models.py`. `pre-commit` was not run
either (unavailable in the environment used).
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ behavior; note
`get_te_hybrid_stack_spec` moved module
(`modelopt.torch.nas.plugins.megatron` →
`modelopt.torch.utils.plugins.megatron_layer_specs`), and a checkpoint
from 0.46.1/0.47.0 needs the `run_config.yaml` edit described in the
changelog.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ (added, not yet executed —
see Testing)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — will run `/claude review`
before marking ready.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Hybrid checkpoints now record complete layer specifications in
`run_config.yaml`.
* Recorded specifications support conversion to Hugging Face format.
* Hybrid MoE configurations support grouped-GEMM and sequential-MLP
modes.
* Configuration-based reconstruction preserves the selected MoE layout.
* **Compatibility**
* Checkpoints from earlier releases may require manually setting the
hybrid layer specification before conversion.
* **Tests**
* Added coverage confirming hybrid specifications survive configuration
serialization and can be recreated successfully.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
0a8e70804a |
Fix hf_ptq.py discarding completed PTQ run on sanity-generate() failure (#2480)
### What does this PR do? Type of change: Bug fix `post_quantize()` in `examples/hf_ptq/hf_ptq.py` ran the optional post-quantization sanity-check `full_model.generate()` unguarded, directly before `export_quantized()`. Any exception raised there aborted the whole run and discarded a completed calibration without exporting a checkpoint. Root cause (traced from [NVBug 6752977](https://nvbugspro.nvidia.com/bug/6752977), DGX Spark GB10 / DeepSeek-R1-Distill-Llama-8B / NVFP4): `get_model()` loads with `device_map="auto"`, relying on `accelerate`'s `infer_auto_device_map`/`get_max_memory()` to decide GPU vs. CPU placement. On DGX Spark's unified-memory single-GPU host, that memory probe under-reports GPU capacity, so part of even an 8B model can land on CPU — and the existing fallback shrinks the GPU budget further (`* gpu_mem_percentage`), compounding it. Calibration survives this because it never invokes the real fake-quant kernel, but the post-PTQ sanity `generate()` does, and NVFP4's dynamic-block-quantize op (`modelopt/torch/quantization/tensor_quant.py`) hard-asserts `amax.is_cuda` with no CPU fallback, so any CPU-offloaded layer crashes there — after ~5.8 hours of calibration, before export. This PR does not attempt to fix the underlying `device_map`/memory-probing behavior (unverified without the actual hardware/logs, which weren't reachable from this environment). Instead it makes the failure mode safe: a failure in the optional sanity check now only skips that check and warns, and export always proceeds, regardless of why `generate()` failed. ### Usage No new API. Behavior change only: `examples/hf_ptq/hf_ptq.py` now completes export even if the post-quantization sanity `generate()` call raises. ### Testing - Added `tests/examples/hf_ptq/test_hf_ptq_args.py::test_post_quantize_export_survives_a_failed_sanity_generate`, which drives `post_quantize()` with a `full_model.generate()` that raises and asserts `export_quantized()` still runs. - Ran `pytest tests/examples/hf_ptq/test_hf_ptq_args.py` (48 passed). - Ran `pre-commit` on the changed files (`ruff-format` reformatted line wrapping only). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ <!-- pending: run `/claude review` --> ### Additional Information Fixes NVBug 6752977. Linked JIRA: OMNIML-5932. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Quantized checkpoint export now continues when the optional post-quantization generation check fails. - A warning is shown when the generation check cannot complete. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ed5c5ed369 |
[OMNIML-5899] Export IQ checkpoints from HF and Megatron (#2447)
## Summary - add IQ format metadata and packed-weight export - support Hugging Face and TP=1 Megatron export paths - reject fused-MoE IQ export until a deployment loader owns its packed layout - document the shaped `uint8` weight contract and the fused-expert boundary - add Hugging Face, Megatron, metadata, and fused-expert export tests ## PR split This work is split into four focused PRs. Each PR targets `main` and owns a disjoint file set: 1. **Kernel** — [#2448: Add CUDA kernels for IQ packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448) 2. **Quantization** — [#2446: Add IQ quantization codecs and backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446) 3. **Export** — [#2447: Export IQ checkpoints from HF and Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447) 4. **Recipes** — [#2449: Add IQ post-training quantization recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449) The required merge order is #2448, #2446, #2447, then #2449. ## Scope This PR owns only export code, deployment documentation, and export tests. It targets `main` and should merge after #2448 and #2446. It does not contain kernel, codec/backend, or recipe files. ## Deployment consumer boundary Dense weights and individually named expert weights use the documented shaped `uint8` contract. Megatron fused-MoE IQ export is intentionally rejected with `NotImplementedError`: its payload would have shape `[num_experts, out_features, in_features // 256, payload_bytes]`, and no deployment loader in this stack currently owns that layout. Support should be enabled only with a loader integration test. ## Test coverage - [Hugging Face packed-weight export](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/unit/torch/export/test_export_weight.py) - [quantization metadata](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/unit/torch/export/test_get_quantization.py) - [Megatron unified export and fused-MoE rejection](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/gpu_megatron/torch/export/test_unified_export_megatron.py) ## Validation - all pre-commit hooks pass for the changed files - 89 focused Hugging Face export, metadata, and fused-expert tests pass locally - direct checks cover both fused-MoE export entry points for IQ1_S and IQ2_XS - Megatron GPU execution remains delegated to GPU CI - restricted-term scan passes <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for IQ1_S and IQ2_XS GGML quantization formats in unified Hugging Face and Megatron exports. * Added quantization metadata, tensor-shape recovery, packing details, and IQ2_XS size documentation. * Added validation for required block sizes and tensor parallelism settings. * **Limitations** * Fused-MoE and GPT-OSS IQ expert packing are not supported. * IQ exports require standard `weight` attributes in Hugging Face models. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a7166965e3 |
Drop GDPVal support from the evaluation skill (#2470)
### What does this PR do? Type of change: deprecation (agent skill) Removes GDPVal support from the `evaluation` agent skill. Its task recipe, example config and Apptainer SIF build helper are deleted. GDPVal was not cleanly separable, so this is not just a delete: - **MRCR depended on GDPVal's infrastructure.** `scripts/nel-gdpval.sh` was a generic pinned-0.2.6 `nel` launcher that only happened to be GDPVal-named, and MRCR ran through it; `references/gym-gdpval.md` documented the gym bootstrap machinery (prepare/reap, `install_on_the_fly` pin↔container coupling, the trust env vars) that both examples share. So the shared parts are kept and renamed rather than dropped: `scripts/nel-gym.sh` and `references/gym.md`. The launcher pin itself is unchanged (0.2.6), as are its env-override semantics. - **GDPVal is an AA-suite member**, so the "AA rule" in `SKILL.md` and `references/quantization-benchmarks.md` had to change. An "AA" / "Artificial Analysis" request now generates the `aa/` tasks as one multi-task config and nothing else. Both files now state that the resulting set omits GDPVal and is therefore not directly comparable to a published AA Index — report per-task scores rather than an aggregate. `SKILL.md` also tells the agent to say GDPVal is unsupported rather than reconstruct a config from an older copy of the skill. Everything GDPVal-specific is gone from `references/gym.md`: the SIF sandbox and its silent-unsandboxed-exec failure mode, the 3-member judge panel, rubric vs. comparison scoring, the deliverables/MLflow `*cache*` trap as a GDPVal concern, and Stirrup-agent deploy sizing. `recipes/env.example` loses `GDPVAL_SIF_DIR`, `GDPVAL_MAX_TURNS` and `TAVILY_API_KEY`, and gains a documented `NEMO_EVALUATOR_TRUST_UNLISTED_TASKS` (required by every gym task, previously only mentioned in prose). One correction carried along: the old reference said the bootstrap pins `ray==2.49.2`, but `example_mrcr.yaml` actually pins to whatever ray version the image already carries. `references/gym.md` now describes what the template does. ### Usage Not an API change. The skill-facing surface that moved: ```text scripts/nel-gdpval.sh -> scripts/nel-gym.sh (NEL_GDPVAL_* -> NEL_GYM_*) references/gym-gdpval.md -> references/gym.md tests/test_nel_gdpval.py -> tests/test_nel_gym.py ``` ### Testing - `pytest plugins/modelopt/skills/evaluation/tests/` — 1 passed. The launcher test was renamed rather than deleted: it is the only coverage for the pinned launcher, which MRCR still depends on. - `pre-commit run --files <changed>` — clean. `sync-claude-skills` fails in my working copy because `.claude/skills/` is a read-only harness mount there; it is unrelated to this change (it trips on `speculative-decoding`) and was skipped for the commit. - Grepped the repo for `gdpval` (case-insensitive): the only remaining hits are the deliberate "no longer supported" notes in `SKILL.md` and `references/quantization-benchmarks.md`. No dangling pointers to the deleted files, and no other skill referenced GDPVal. - Checked `tests/evals.json` for both `evaluation` and `day0-release`: no eval case expected a GDPVal companion config, so no expectations needed updating. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ❌ — anyone with a saved GDPVal config keeps it, but the skill no longer generates one and the SIF helper is gone. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — existing launcher test renamed and kept passing. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — under **Deprecations**. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information A follow-up is needed in the modelopt-internal repo: `modelopttools:eval-config` Step 3c is the GDPVal SIF / comparison-mode conversion checklist, and Step 3d names the gym image. Step 3c is now dead and the pointers this skill used to make into it are gone. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Evaluation Updates** * Added standalone NeMo Gym support for MRCR tasks with pinned launcher and Gym configuration requirements. * Added shared guidance for Gym setup, validation, execution, and recovery. * AA requests now generate only `aa/` tasks and report per-task scores. * **Removed Support** * Removed GDPVal evaluation recipes, documentation, task guidance, Apptainer helper tooling, and AA-suite inclusion. * Updated default quantized-checkpoint validation recommendations to exclude GDPVal. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
76c04dfd99 |
docs: clarify canonical pruning documentation source (#1871) (#2469)
Fixes #1871 The pruning documentation is split between `docs/source/guides/3_pruning.rst` and `examples/pruning/README.md`, causing confusion about which is authoritative. The examples/pruning/README.md is the comprehensive, up-to-date reference covering Minitron, Puzzletron, FastNAS, support matrix, guidelines, and distillation hyperparameters. This PR adds a note to the RST guide making clear: - The README is canonical for Minitron and Puzzletron (LLM/VLM pruning) - The guide covers FastNAS for Computer Vision models Signed-off-by: Diya <diyaismahil7@gmail.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Documentation** - Updated the pruning guide’s introductory content for clearer separation of general guidance and the related Minitron/Puzzletron note. - Clarified that the guide focuses on FastNAS pruning for computer vision models. - Added references to the Pruning README for Minitron and Puzzletron API examples, support information, guidelines, and distillation hyperparameters. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: didi <diyaismahil7@gmail.com> |