mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
main
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ad8cd63847 |
Share one CUDA encoder per IQ family (#2615)
### What does this PR do? Type of change: refactor (no behaviour change) The five GGML IQ CUDA encoders were five copies of the same search. IQ2_XS and IQ2_XXS shared 139 of their roughly 150 lines of encoder and launcher code, and IQ2_S 113 of them. IQ1_S and IQ1_M had the same structure with a different choice space. Every scaled packer also validated its scales twice, in the `ggml.cpp` pybind wrapper and again in the CUDA entry point. This PR keeps **one encoder per family**, as two templates: - **`iq2_family.cuh`** for IQ2_XS, IQ2_XXS and IQ2_S. The grid sits in shared memory, the 16 local scales are scored per group, and each vector then takes its best entry under the chosen scale. A format supplies its group shape, whether it stores seven sign bits and recovers the eighth from parity, and a `store()` that writes the chosen entries, sign masks and local scales into its layout. - **`iq1_family.cuh`** for IQ1_S and IQ1_M, over the shared ternary grid. Each group picks one of `kChoices` options. With `kSharedShift` the option also fixes the ±1/8 delta (IQ1_S: `shift * 8 + local`); otherwise each vector picks its own (IQ1_M). IQ1_S's scale kernel now writes FP16 scales, so both IQ1 formats take the same input. Each format file is now one `Format` struct, holding its layout constants and `store()`, plus its entry point: 58–100 lines each. Validation lives once in `common.cuh`, as `check_pack_inputs` and `check_scaled_pack_inputs`. `ggml.cpp` binds the CUDA entry points directly instead of through five wrappers. **The kernel sources shrink from 1,536 to 1,241 lines** (+665 / −960). This is the first of two PRs. #2604 builds on it: it adds CUDA decoders as a `decode()` next to each format's `store()`, and makes export reuse fake quant's packed payloads. ### Testing **Nothing changes in the output.** Before the refactor I hashed 40 outputs: 5 formats × float32/bfloat16/float16/float64 inputs × encode and decode, on a weight with zero, tiny, oversized and non-finite blocks. All 40 hash the same afterwards. **Encode speed is unchanged.** Old and new were timed alternately for four rounds, in both orders, on an idle RTX PRO 6000 with a 5632×2048 weight. They were within 1% for every format: IQ1_S 37.6 / 37.6 ms, IQ1_M 37.0 / 37.0, IQ2_XXS 10.9 / 10.9, IQ2_XS 11.9 / 12.0, IQ2_S 15.5 / 15.4. - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **49 passed** - **Validation reports the same errors in the same order.** Over 5 formats × 8 combinations of bad arguments (devices, dtype, width, grid shape, scales dtype, length and sign), every first error matches main's. - `tests/gpu/_extensions/test_torch_extensions.py`: the validation-message tests pass. #2515's two Q8_0 tests fail identically on a clean `main` on this GPU. - IQ unit tests (`test_ggml_backend.py`, `test_iq_formats.py`, `test_convert_hf_config.py`, `test_presets.py`, `test_export_weight.py`): **173 passed** ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ Same bindings, messages and bytes. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code sources or dependencies. - Did you write any new necessary tests?: N/A. A refactor with no behaviour change, verified by the hashes above and the existing GPU tests. - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: **this** → #2604 (pack each IQ weight once and decode on CUDA). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Quantization now checks that inputs and grids are CUDA tensors on the same device, with compatible shapes. Scaled formats also validate scale type, shape, and finite, non-negative values. * **Improvements** * IQ1 and IQ2 formats share common encoding paths while retaining their format-specific output layouts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
7d9e07d14b |
Count only added lines toward the PR size budget in AGENTS.md (#2616)
### What does this PR do? Type of change: documentation Changes the PR sizing rule in `AGENTS.md` to count only **added source** lines toward the ~500-line budget, instead of total changed lines. Deletions are cheap to review, so a PR that mostly removes code shouldn't be pushed into a split. Tests and docs are excluded too, since every sub-PR has to carry its own tests. The check uses the insertions count from `git diff --shortstat` with a pathspec that excludes `tests/` and `docs/`. ### Usage ```bash git diff --shortstat origin/main...HEAD -- . ':!tests' ':!docs' # N files changed, X insertions(+), Y deletions(-) -> compare X against ~500 ``` ### Testing - `pre-commit run --files AGENTS.md` (markdownlint passes). - Ran the pathspec against recent commits (#2595, #2513) to confirm it drops test and doc lines from the count. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information Follow-up to #2494, which introduced the sizing guidance. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated review guidance to measure pull request size by added source lines, excluding deletions, tests, and documentation. The guidance retains the recommendation to check the size before opening a review. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
aa89722d38 |
[6/6] Add the IQ1_M CUDA encoder and register the format (#2595)
### What does this PR do? Type of change: new feature **Second of two PRs adding IQ1_M** (1.75 bits per weight). #2513 landed the PyTorch codec; this PR adds its **CUDA encoder** and makes the format reachable. With it, ModelOpt supports all five GGML IQ formats at one and two bits. - the CUDA encoder, its binding and extension build wiring, plus the CUDA path in `quantize_iq1_m` - an `IQFormat` record and **one `IQ_FORMAT_REGISTRY` entry**, so backend dispatch, both exporters and `convert_hf_config` take it from there - the `ggml` package export - the `general/ptq/iq1_m` recipe, its presets, `ptq.md` and a CHANGELOG entry The kernel lands with the registration so every registered format keeps a CUDA encoder. ### The kernel In the kernel the delta shift is free per group, so it sits above the entry index in the sort key: a tie still prefers the lower shift and then the lower entry, as the reference encoder does. The 2048-entry grid IQ1_M shares with IQ1_S is 64 KiB, past the 48 KiB static shared-memory limit, so both kernels read it from global memory and rely on the cache. | 5632×2048 weight | torch | CUDA | | |---|---|---|---| | IQ1_M encode | 5.6 M elem/s | **318 M elem/s** | **57×** | ### Shared with IQ1_S rather than copied The two IQ1 kernels load each vector, score it against a grid entry and apply the ±1/8 shift the same way. So those three steps move into `common.cuh` as `load_vector`, `grid_terms` and `shifted_error`, and IQ1_S uses them too. **IQ1_S's packed bytes are unchanged**: its CUDA output on a 5632×2048 weight hashes the same before and after, and so does IQ1_M's, compared against the pre-split version of this change. IQ1_S encodes at the same speed (306 M elem/s). ### Usage ```bash python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq1_m ``` ### Testing Registering the format brings it under every registry-driven test with no IQ1_M-specific test code: backend dispatch, weight caching, the `num_bits` guard, `convert_hf_config` metadata, Megatron export and the `TensorQuantizer` tests in the shared battery. The shared CUDA battery gains one row. - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py`, `tests/unit/recipe/test_presets.py`: **166 passed** - broader unit sweep (`-k 'ggml or iq or gguf or registry'` over quantization, export and recipe tests): **221 passed**. The one failure, `test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes`, is a `torchvision` import error in my environment, unrelated to IQ. - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **49 passed** on RTX PRO 6000 Blackwell (sm_120), 7 of them IQ1_M, including CUDA-vs-PyTorch encoder parity - `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k 'iq or ggml'`: **45 passed** (9 tests × 5 formats) in `nvcr.io/nvidia/nemo:26.08` - `tests/examples/hf_ptq/test_llm_ptq.py -k iq1_m`: **passed** - reconstruction error falls monotonically across all five formats, pinned by a test - `general/ptq` now holds 31 recipes; `ptq.md` is updated. Rebased onto `main` after #2513 merged. The resulting tree is identical to the one the runs above tested, and the unit set was rerun on it: 166 passed. On this GPU, two of #2515's Q8_0 tests in `tests/gpu/_extensions/test_torch_extensions.py` fail: `test_cuda_ext_q8_0_zero_and_roundf_layout` and `test_cuda_ext_q8_0_dequantizes_with_small_error`. They fail identically on a clean `main` checkout, so they are not from this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code sources or dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S codec) → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M codec), all merged → **this**. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added IQ1_M weight-only quantization at 1.75 bits per weight, with CUDA acceleration and a 256-value block size. * Added an IQ1_M post-training quantization recipe for eligible linear layers; calibration data is not required. * Added IQ1_M to the supported GGML-compatible formats and recipe listings. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
e5b63320ab |
[5/6] Add the IQ1_M codec (#2513)
### What does this PR do? Type of change: new feature (not yet user-reachable) **First of two PRs adding IQ1_M** at 1.75 bits per weight, just above IQ1_S. This one lands the **PyTorch codec**: encoder and decoder. It is deliberately **not registered**, so no quantizer dispatches to it and the `ggml` package does not export it. #2595 adds the CUDA encoder, registers the format and adds its recipe. With both, ModelOpt supports all five GGML IQ formats at one and two bits. On the mixed-precision checkpoint #2511 measured (`unsloth/Qwen3.8-27B-GGUF`), IQ1_M covers **25 tensors and 1.2 B parameters**. With all five formats we can read 89.0% of that file; the rest is k-quants and F32. ### What's distinctive about it **IQ1_M is the most irregular layout of the five.** There is no leading block scale field at all. The FP16 super-block scale is reassembled from the top nibble of each of four scale words: ```c scale.u16 = (sc[0] >> 12) | ((sc[1] >> 8) & 0x00f0) | ((sc[2] >> 4) & 0x0f00) | (sc[3] & 0xf000); ``` It is also finer grained than IQ1_S: a local scale per **two** groups rather than four, and a delta shift chosen **per group** rather than per sub-block. That is where its extra 0.1875 bits go. ### Shared with IQ1_S rather than copied IQ1_M searches exactly as IQ1_S does: the same 2048-entry grid, the same ±1/8 delta, every (shift, local scale) choice for every 8-value vector. It differs only in how it selects among those choices afterwards. So the search moves out of IQ1_S's encoder into `_search_shifted_grid`, which both call, and `iq1_m.py` keeps only its selection and packing. **IQ1_S's encoded bytes are unchanged**, checked by hashing its output before and after on a fixed input. ### A scale-anchor correction IQ1_M anchors its scale differently from IQ1_S: the ratio **rises with a block's peak-to-RMS** rather than being flat, and clamps higher. It uses `clamp(0.58 + 0.035 * peak_to_rms, 0.65, 0.95)` against IQ1_S's flat `0.61`. Measured over 15 Qwen3.8-27B MLP weights: | | flat 0.61 | correct anchor | | |---|---|---|---| | relative reconstruction MSE | 0.17372 | **0.17291** | **−0.47%** | It is consistent on every tensor, with no outliers. The anchor changes quality without touching layout, so neither round-trip nor conformance tests would catch it drifting. `test_scale_anchor_follows_peak_to_rms` now pins it, for all five formats; see Testing. ### Family parity Two surface asymmetries close here, so the five are uniform. `IQ1_S` now exposes `_predict_iq1_s_scales` like the other four, instead of computing its anchor inline. `IQ1_M` exposes `iq1_m_grid`, aliasing the IQ1_S table it shares. ### Testing **The decoder is validated against llama.cpp's own output, not just round-tripped:** ``` IQ1_M: 25 tensors, 4,730,880 blocks → 0 mismatched, max|diff| 0.0 ``` This mattered: **my first IQ1_M decoder had a real bug.** A `repeat_interleave` on the wrong axis produced `[h0,h1,h0,h1]` where llama.cpp needs `[h0,h0,h1,h1]`. A round-trip against our own encoder still passed, because the encoder made the matching mistake. Only comparison against bytes we did not produce caught it. Blocks from that checkpoint ship as conformance vectors, and mutation testing confirms they catch a mis-set scale nibble. The decoder unpacks every field in one vectorized pass, since fake quant decodes on every forward: 5.2 ms for a 5632×2048 weight (IQ1_S: 3.3). - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py`, `tests/unit/recipe/test_presets.py`: **153 passed**, 15 of them IQ1_M codec cases, including the llama.cpp conformance check - `test_scale_anchor_follows_peak_to_rms` pins every format's scale anchor. It predicts scales for blocks whose peak-to-RMS is exactly 1, 4, 8 and 16, reaching both clamps and two points on each slope, and compares them against anchors written out in the test. Mutations each fail exactly the mutated format: reverting IQ1_M to IQ1_S's flat 0.61, moving either IQ1_M clamp, changing its taper by 0.001, moving an IQ2_S or IQ2_XS clamp, and changing IQ1_S's anchor to 0.62. - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed**. IQ1_S's CUDA-vs-PyTorch parity still holds after its encoder refactor. - IQ1_S and IQ1_M PyTorch encoder output and IQ1_M decoder output hash identically to the pre-split version of this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ IQ1_M adds no codebook; it reuses the IQ1_S table already carried in `codebooks.py`. The new conformance vectors come from `unsloth/Qwen3.8-27B-GGUF`, which is Apache-2.0 like its base model `Qwen/Qwen3.8-27B`; the vectors' docstring now records that. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A. Nothing is user-reachable yet; #2595 carries the entry. - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S codec) → #2565 (IQ2_S CUDA encoder and registration), all merged → **this** → #2595 (IQ1_M CUDA encoder and registration). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added IQ1_M quantization and dequantization for compact, GGML-compatible blocks of 256 values. * Added access to the IQ1_M grid and configurable chunk sizes for processing data. * **Bug Fixes** * Improved IQ1_S scale prediction and grid-search organization while preserving its encoding behavior. * **Tests** * Added IQ1_M conformance data and included the format in shared IQ-format test coverage. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
3091b8ff69 |
[4/5] Add the IQ2_S CUDA encoder and register the format (#2565)
### What does this PR do? Type of change: new feature **Second of two PRs adding IQ2_S** (2.5625 bits per weight). #2512 landed the PyTorch codec; this PR adds its **CUDA encoder** and makes the format reachable: - the CUDA encoder, its binding and extension build wiring, plus the CUDA path in `quantize_iq2_s` - an `IQFormat` record and **one `IQ_FORMAT_REGISTRY` entry**, so backend dispatch, both exporters and `convert_hf_config` take it from there - the `ggml` package export - the `general/ptq/iq2_s` recipe, its presets, `ptq.md` and a CHANGELOG entry The kernel lands with the registration so every registered format keeps a CUDA encoder. On the mixed-precision checkpoint #2511 measured (`unsloth/Qwen3.8-27B-GGUF`), IQ2_S covers **9 tensors and 0.6 B parameters**. ### The kernel IQ2_S's **1024-entry codebook is twice IQ2_XS's**, which makes its search the most expensive in the family. The codebook and its norms take 36 KiB of shared memory, the most of any IQ kernel but inside the 48 KiB static limit, so they are declared statically like the IQ2_XS and IQ2_XXS kernels. That cost is why the kernel matters more here than anywhere else: | | torch | CUDA | | |---|---|---|---| | IQ2_S, 5632×2048 weight | 0.8 M elem/s | **725.7 M elem/s** | **907×** | | extrapolated to a 27B model | ~9.8 hours | **~37 s** | | ### Usage ```bash python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq2_s ``` ### Testing Registering the format brings it under every registry-driven test with no IQ2_S-specific test code: backend dispatch and weight caching, the `num_bits` guard, `convert_hf_config` metadata (uniform and mixed precision), all 9 Megatron export tests, and the two `TensorQuantizer` tests in the shared battery. The shared CUDA battery gains one row. - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py`, `tests/unit/recipe/test_presets.py`: **134 passed** - broader unit sweep (`-k 'ggml or iq or gguf or registry'` over quantization, export and recipe tests): **192 passed**. The one failure, `test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes`, is a `torchvision` import error in my environment, unrelated to IQ. - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed** on RTX PRO 6000 Blackwell (sm_120). 7 of them are IQ2_S: CUDA-vs-PyTorch encoder parity, determinism, reconstruction at scale, zero and non-finite policy, float64 input and the fallback path. - `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k 'iq or ggml'`: **36 passed** (9 tests × 4 formats) in `nvcr.io/nvidia/nemo:26.08` - `tests/examples/hf_ptq/test_llm_ptq.py -k iq2_s`: **passed**. TinyLlama PTQ through unified HF export writes `quant_algo: IQ2_S`, `block_payload_bytes: 82`, and `down_proj` packed as `(2048, 22, 82)` uint8. - `general/ptq` now holds 30 recipes. - The shared-memory change in `b7739d5d0` leaves the packed bytes identical (same hash on a 5632×2048 weight), and packing runs at 849.1 M elem/s against 825.7 before on RTX PRO 6000. The GPU battery was rerun: 42 passed. All of the above was rerun after rebasing onto `main` at `c2aaa44f6`. That base adds a Q8_0 packer to the same GGML extension (#2515), and changes the hf_ptq example and the export code this format goes through. The packed IQ2_S bytes still hash the same. On this RTX PRO 6000 (sm_120), two of #2515's own Q8_0 tests in `tests/gpu/_extensions/test_torch_extensions.py` fail: `test_cuda_ext_q8_0_zero_and_roundf_layout` and `test_cuda_ext_q8_0_dequantizes_with_small_error`. They fail identically on a clean `main` checkout, so they are not from this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code sources or dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: #2511 (IQ2_XXS, merged) → #2525 (format registry, merged) → #2512 (IQ2_S codec, merged) → **this** → #2513 (IQ1_M). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added IQ2_S weight-only quantization for eligible linear layers, at 2.5625 bits per weight. * Added a PTQ recipe that requires no calibration data. Weights must meet the existing 256-value block-size constraint. * Added CUDA-accelerated packing for CUDA weights, with a Python fallback when the CUDA extension is unavailable. * **Documentation** * Updated the PTQ recipe catalog and IQ-format size tradeoffs. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
767ef5533e |
[3/5] Add the IQ2_S codec (#2512)
### What does this PR do? Type of change: new feature (not yet user-reachable) **First of two PRs adding IQ2_S**, the widest of the GGML IQ formats at one and two bits (2.5625 bits per weight). This one lands the **PyTorch codec**: the encoder, the decoder and the 1024-entry codebook. It is deliberately **not registered**, so no quantizer dispatches to it and the `ggml` package does not export it. #2565 adds the CUDA encoder, registers the format and adds its recipe. ### What's distinctive about it **IQ2_S is the one format llama.cpp's own tooling gives no head start on**, so the search is written against the GGML layout directly. The interesting difference from IQ2_XS and IQ2_XXS is sign handling. IQ2_S stores a **full 8-bit sign mask** per group rather than a 7-bit parity-coded index. The encoder therefore takes the input signs as they are instead of flipping the weakest element to fix parity, and the search compares magnitudes directly, which is simpler than its siblings. ### Why the codec lands before the kernel The CUDA encoder's tests use this codec as their reference. They compare against the PyTorch encoder byte for byte and draw the grid and scale predictor from it. So the kernel cannot be tested before the codec exists, and it follows in #2565 together with the registration. Every registered format therefore keeps a CUDA encoder. ### Test changes that make the split possible A codec can now land before it is registered, so two test contracts in `test_iq_formats.py` are stated precisely: - The two tests that go through `TensorQuantizer` (pass-through gradient, error falls with bit width) iterate `IQ_FORMAT_REGISTRY`. Every other battery test calls the codec directly and covers IQ2_S here. - The coverage check now asserts `set(IQ_FORMAT_REGISTRY) <= set(FORMATS)` instead of equality. That is what its docstring already said: a registered format must be listed, or it escapes the contract. - `test_registry_lists_every_exported_encoder` is unchanged, and it is why this PR leaves the package exports alone: an exported encoder must be registered. The error-by-bit-width failure message also labels errors by the order they were measured in; it previously zipped them with alphabetical names. ### Testing **The decoder is validated against llama.cpp's own output, not just round-tripped:** ``` IQ2_S: 9 tensors, 2,355,200 blocks → 0 mismatched, max|diff| 0.0 ``` The new codebook matches the `ggml-common.h` table entry for entry. Blocks from `unsloth/Qwen3.8-27B-GGUF` ship as conformance vectors, so CI keeps checking bytes we did not produce. - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py`, `tests/unit/recipe/test_presets.py`: **121 passed**, 14 of them IQ2_S codec cases, including the llama.cpp conformance check - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **35 passed**, unchanged by this PR ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ The new codebook is a GGML table, carried in `codebooks.py` with the source revision recorded. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A. Nothing is user-reachable yet; #2565 carries the entry. - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: #2511 (IQ2_XXS, merged) → #2525 (format registry, merged) → **this** → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added GGML-compatible IQ2_S quantization and dequantization support, including access to its magnitude grid. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
80e04b8816 |
docs(eval): align Terminal-Bench 2.1 / SWE-bench / MRCR with upstream configs (#2479)
### What does this PR do? Type of change: documentation (agent skill) + template bug fixes Aligns the `evaluation` skill's three upstream-tracked benchmarks (Terminal-Bench 2.1, SWE-bench Verified, MRCR) with the current `nvidia-eval-factory-benchmarking` configs, and fixes guidance that turned out to be wrong when a full three-benchmark campaign was run with the skill end to end. Rebased on #2499. The first commit is the alignment. The rest address review and a fresh config-generation test: MRCR serving scoped per variant, `limit_samples` canary guidance, the template's serve command passing `--gpu-memory-utilization` (replacing `command:` had silently dropped it, so the 1M golden's 0.95 never reached vLLM), upstream's 128K values, and regression tests for the `++limit` gate and the serve command. The skill text keeps only the rules; the evidence behind them is below. GDPVal is out of scope. It was removed from this skill in #2470, and this PR replaces #2464. **Alignment with upstream** - Sandbox region via `HARBOR_ECS_REGION`. TB2.1's ECR repo name tracks the region; SWE-bench's stays in us-west-2. - One interceptor order for both harbor benchmarks, with `http_pairs_dump` **last**. Upstream is split on its position, which changes only what the dump records, never the score. Interceptor lists replace wholesale on merge, so a leaf must restate the whole chain. - `capture_request_body` goes on the service. A shared block injects an alias-only entry with no `type`. - MLflow tags gain `task_name` and `nemo-evaluator-next-version`. - TB2.1 `max_concurrent`: 50 for nano-class models, 15 for larger ones (all upstream non-nano leaves override it to 15). - `proxy.request_timeout` must always be set explicitly. Otherwise it inherits the model fragment's serving value, which ranges from 3600 to 36000 upstream, and TB2.1 has no benchmark key for it. - MRCR: - `parallelism` is 512, deliberately above server capacity, so `--max-num-seqs` must no longer be derived from it. - `limit_samples` now reaches the gym through a gated `++limit`. - Observability capture is on. - 128K is its own upstream benchmark on the condensed gym schema. **Corrections found by running it** | what the skill said | what actually happens | |---|---| | `username: ${oc.env:USER}` | nel-next only expands `${VAR}` / `${VAR:-default}`. The config passes `--dry-run` and fails at `--submit` with *"remote username contains invalid characters"*. | | MRCR canary via `++limit` edited into `collect_rollout_params` | `limit_samples` is now gated through. Under the 0.2.6 launcher the `-o` path is `++evaluation.nemo_evaluator_config.config.params.limit_samples`. `++config.params…` creates a bogus top-level key. | | the condensed gym schema's bootstrap is in the runtime image | It lives in upstream `configs/models/gym_eval_command.yaml`, which is composed in. A standalone config must carry the `command:` block. | | `mean/prefix_matched ~0.55 is healthy` | That value is calibrated to the 1M golden. A 128K run at `pass@1` ≈ 95 measured ≈ 1.0. The signal is a collapse toward 0. | | NVFP4 MoE `VLLM_*` env vars as reliable knobs | They are build-dependent: one vLLM build logged them as unknown and ignored them. Check the server log once per image. | **Rules the skill lacked** - **MRCR variant.** Pick the largest variant within the checkpoint's trained context. On a 262K-context model, 1M needs `VLLM_ALLOW_LONG_MAX_MODEL_LEN` and measures extrapolation, which a quantization comparison would then entangle with quantization damage. The serving setup follows the variant: 128K serves at the trained context without the override. - **SWE-bench `reasoning_effort`.** openhands-sdk sends `reasoning_effort: high` on every call, and canonical `bench.yaml` doesn't strip it. A server whose accepted set excludes `high` returns HTTP 400 on the first call of every trial, so `pass@1` is 0. The fix is to overwrite it with the server default via `proxy.extra_body`. terminus-2 (TB2.1) and Gym's `simple_agent` (MRCR) never send the key (46/46 and 110/110 requests checked). - **Reasoning toggles** go in `extra_body.chat_template_kwargs`. Don't write out no-op sampling defaults; `top_k` is *not* one (vLLM's default is `-1`). - **Upstream model fragment.** Consult `configs/models/<model>/` when it exists, for serving flags, the thinking toggle and `reasoning_replay.mode`. ### Usage No API change. Regenerating a config from the skill now yields the aligned values: ```yaml # recipes/examples/example_eval_next.yaml services: model: proxy: request_timeout: 3600 # always explicit benchmarks: - max_concurrent: 50 # nano-class; larger models 15 sandbox: region: ${HARBOR_ECS_REGION:-us-east-1} cluster: username: ${USER} # NOT ${oc.env:USER} ``` ### Testing - `python -m pytest plugins/modelopt/skills/ -o addopts=""`: 7/7 pass, including the new `tests/test_example_mrcr.py`. It checks that `example_mrcr.yaml` emits `++limit=N` only when `limit_samples` is set, and that the folded vLLM serve command shell-parses with every flag intact, including `--gpu-memory-utilization`. Each check fails when its defect is reintroduced. - `markdownlint-cli2` on the changed Markdown: 0 errors. Both example YAMLs parse. The full `pre-commit` suite was not run after the squash, because the sandbox could not fetch hook repos. The pre-squash commits passed it. - **Exercised end to end.** Configs built from this skill ran a BF16 campaign for a 262K-context MoE reasoning model on an internal cluster to completion: - MRCR-128K: 1470/1470 rollouts - Terminal-Bench 2.1: 712/712 trials - SWE-bench Verified: 2500/2500 trials Each correction above is a defect that campaign surfaced. - **Differential check.** Terminal-Bench configs generated from the pre- and post-change skill, from the same brief and in isolation, differ on: - interceptor chain - concurrency - `top_k` - region interpolation - MLflow tags Not run: a scored evaluation of this PR itself. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ <!-- Docs + a template fix; no ModelOpt API surface touched. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A <!-- Documentation and config-template values. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- Agent skill docs, not a user-facing ModelOpt feature/breaking change/deprecation. --> - Did you get Claude approval on this PR?: ❌ <!-- Not run. --> ### Additional Information The companion internal `eval-config` change now carries only the internal values: the ECR URLs, the region default and the cluster image notes. It points here for the generic rules. **Known gaps not fixed here**, worth a follow-up: - `references/nel-next.md:136` and `references/launcher-workflow.md:207` derive `--max-num-seqs` from `parallelism / DP`. nel-next has no `parallelism` field (its analogue is `max_concurrent`), and MRCR's 512 is deliberately not a server cap. - `references/launcher-workflow.md:200-201` makes `--max-num-batched-tokens` and `--enable-chunked-prefill` always-include defaults, but `example_eval_next.yaml` omits both. - `references/launcher-workflow.md:27` says `sbatch_comment` belongs under `execution:` and is otherwise inert, yet all three shipped examples put it under `cluster:`. - MoE detection (`--enable-expert-parallel`) is unresolvable from the facts the skill asks for when the model handle has no `-A*B` suffix. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
400498d82d |
[2/4] Register each GGML IQ format once for dispatch and export (#2525)
### What does this PR do? Type of change: refactor (no behaviour change) Addresses review feedback on #2511. Backend dispatch and export each kept their own list of the GGML IQ formats: `_FAKE_QUANTS` in the backend, and `IQ_FORMATS`, `IQ_BLOCK_METADATA` and `IQ_PACKERS` in export. All four listed the same formats. Adding a format meant a row in each, and the lists could drift apart. That had already happened twice in #2511: `convert_hf_config.py` kept its own upper-case spelling of the family and dropped IQ2_XXS metadata, and the Megatron export tests were hard-wired to two formats. Each format module now declares **one `IQFormat` record** beside its encoder and decoder: name, block geometry, `quantize`, `dequantize`, and its encode and decode chunk defaults. **`IQ_FORMAT_REGISTRY`** lists them. - Backend dispatch looks formats up in the registry. - Both exporters take the packer and block geometry from it. - Export's `IQ_FORMATS` is derived from it instead of being written out again. - `_FAKE_QUANTS`, `IQ_BLOCK_METADATA` and `IQ_PACKERS` are removed. - The per-format fake-quant wrappers collapse into one `IQFormat.fake_quant`, which does the `num_bits` check and calls the existing cache helper. Codebooks, searches, payload layouts and CUDA encoders stay in each format's module. #### Series and merge order This is one slice of the IQ format series. It targets `main` so unit CI runs, and **its diff includes #2511's commits until #2511 merges**. 1. #2511 — IQ2_XXS format 2. **this PR** — one registration per format 3. #2512 — IQ2_S format 4. #2513 — IQ1_M format After this lands, #2512 and #2513 are restacked onto it, so each adds a format module and a single registry entry instead of rows in four tables. #### Design choices - **An explicit list, not self-registration at import.** If formats registered themselves when their module was imported, the registry's contents would depend on import order. - **Backward compatible, with one behaviour change.** `iq1_s_fake_quant` and `iq2_xs_fake_quant` are public on main, so each format keeps its `<fmt>_fake_quant` name as an alias of its record's method. The three removed tables were introduced by #2511 and never released. The behaviour change: on main, the alias looked the encoder up at call time, so patching `iq1_s.quantize_iq1_s` changed what it ran. Now the record captures the encoder and decoder when it's built, so patching those module functions reaches neither dispatch nor the alias. Substitute through `IQ_FORMAT_REGISTRY` instead. - **Registering a format declares it exportable, and that's intended.** Export's `IQ_FORMATS` is derived from the registry, so a format registered for dispatch is also claimed by both exporters and `convert_hf_config`. That can't be wrong for an IQ format: fake quant is `dequantize(quantize(w))`, so a format can't be dispatched without the packer and block geometry, and those are all export reads. A QAT-only IQ format can't exist. If one ever needs to land ahead of its export path, an `exportable` flag on the record is a one-line addition. - **The registry is the substitution seam.** Dispatch now reads the registry, so tests that swap an encoder or decoder swap the registry entry. Patching the format module's function would no longer reach dispatch. - **Test expectations stay independent of the registry.** Tests take the *list* of formats from the registry, but their expected values come from each format's own module (`quantize_<fmt>`, `<FMT>_BLOCK_BYTES`, …). A mis-wired registry entry therefore can't make both sides of an assertion agree. #### What it does not unify The CUDA side (`ggml.cpp` bindings, the `extensions.py` source list, codebook sizes in `common.cuh`) and the recipes and docs remain per format. "One registration" holds for the Python side, which is where all four tables lived. ### Usage Adding a format after this PR (for example IQ2_S in #2512) needs its module and one line in the registry: ```python # modelopt/torch/quantization/ggml/iq2_s.py IQ2_S_FORMAT = IQFormat( name="iq2_s", block_size=IQ2_S_BLOCK_SIZE, block_bytes=IQ2_S_BLOCK_BYTES, quantize=quantize_iq2_s, dequantize=dequantize_iq2_s, block_chunk_size=_DEFAULT_BLOCK_CHUNK_SIZE, decode_chunk_size=_DEFAULT_DECODE_CHUNK_SIZE, ) # modelopt/torch/quantization/ggml/registry.py IQ_FORMAT_REGISTRY = {fmt.name: fmt for fmt in (IQ1_S_FORMAT, IQ2_XXS_FORMAT, IQ2_XS_FORMAT, IQ2_S_FORMAT)} ``` Looking up a format: ```python from modelopt.torch.quantization.ggml import IQ_FORMAT_REGISTRY fmt = IQ_FORMAT_REGISTRY["iq2_xxs"] packed, shape = fmt.quantize(weight) # GGML blocks fmt.block_bytes, fmt.effective_bits # 66, 2.0625 ``` ### Testing - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py` — **94 passed** - `tests/gpu/torch/quantization/test_iq_formats_cuda.py` — **22 passed** (RTX PRO 6000) - `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k iq` — **27 passed** in `nvcr.io/nvidia/nemo:26.08`, the image CI uses for that suite - broader sweep of IQ, export and recipe unit tests — **153 passed**, none failed **New guards on the registry itself:** - every encoder the package exports is registered - each record points at its own format's codec, geometry and chunk defaults - the public `<fmt>_fake_quant` alias is the registered record's method - export's `IQ_FORMATS` and `QUANTIZATION_IQ*` constants match the registry - a format's `fake_quant` refuses a quantizer configured for another format. Dispatch picks the record by `num_bits`, so it never reaches this guard; the test covers direct callers of a record or alias. The three per-format guards it replaced were untested on main. - every registered format is listed in the shared test batteries Checked by mutation: leaving IQ2_XXS out of the registry, or registering it with the IQ2_XS encoder, each fails the guard written for that case. **Coverage gap closed along the way:** `test_ggml_backend.py` was hard-wired to IQ1_S and IQ2_XS, so IQ2_XXS had no backend, cache or packed-once coverage. Those tests now run over the registry. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — public per-format fake-quant names are kept as aliases; the removed tables were never released. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A — internal refactor with no user-visible change - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information Review feedback on #2511 that this addresses: *"`_FAKE_QUANTS`, `IQ_FORMATS`, `IQ_BLOCK_METADATA`, and `IQ_PACKERS` independently enumerate the same formats. A common pack/dequantize/fake_quant interface would let backend dispatch and export consume one registration."* 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * IQ quantization formats are available through a shared format registry, keeping format details and quantization behavior consistent across supported workflows. * IQ-format model exports use registered format information for quantization metadata and weight packing. * **Tests** * Expanded checks to cover registered IQ formats and verify consistent format support across quantization and export. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
a21411adde |
Add the IQ2_XXS weight-only quantization format (#2511)
### What does this PR do? Type of change: new feature llama.cpp defines five GGML IQ formats at one and two bits; we ship two. This adds **IQ2_XXS** at 2.0625 bits per weight, between IQ1_S and IQ2_XS, and is the **first of three**. On a real mixed-precision checkpoint (`unsloth/Qwen3.8-27B-GGUF`, `Qwen3.8-27B-UD-IQ1_S.gguf`) IQ2_XXS alone covers **59 tensors and 2.84 B parameters — 10.6% of the file**, which a reader limited to IQ1_S/IQ2_XS cannot consume. Across all three PRs the missing formats account for 17.3%. | format | bpw | bytes/256 | codebook | | |---|---|---|---|---| | `iq1_s` | 1.5625 | 50 | `iq1s_grid` (2048) | existing | | **`iq2_xxs`** | **2.0625** | **66** | **`iq2xxs_grid` (256)** | **this PR** | | `iq2_xs` | 2.3125 | 74 | `iq2xs_grid` (512) | existing | The encoder follows the existing single-pass grid search at a fixed anchored super-block scale, and the CUDA kernel the existing per-block structure. IQ2_XXS reuses IQ2_XS's even-parity sign rule but packs a 4-bit sub-block scale into the same 32-bit word as four 7-bit sign indices, and its 256-entry grid needs no high index bits. ### Groundwork the next two reuse Two things land here because IQ2_XXS is the first format to need them: - **Export registry.** The IQ family was spelled as a two-element tuple at **nine** sites across `quant_utils.py`, `unified_export_hf.py` and `unified_export_megatron.py`. Those become an `IQ_FORMATS` frozenset plus per-format packer and block-geometry tables, so a format is a row rather than a sweep through the exporters. - **Shared test contract.** The per-format test files had drifted apart — each of `iq1_s` and `iq2_xs` tested things the other did not. They become one parametrized module per layer (unit and CUDA), so every format is held to the same contract and a new one inherits it. ### Usage ```bash python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq2_xxs ``` ### Testing **The decoder is validated against llama.cpp's own output, not just round-tripped.** Every IQ2_XXS tensor in the checkpoint above, compared against `dequantize_row_iq2_xxs` from `ggml-quants.c`: ``` IQ2_XXS: 59 tensors, 11,100,160 blocks → 0 mismatched, max|diff| 0.0 ``` The new codebook matches the `ggml-common.h` table entry for entry, as does the `ksigns_iq2xs` sign table. Blocks lifted from that checkpoint ship as conformance vectors so CI keeps checking bytes we did not produce; mutation testing confirms they catch a wrong sign-field width. The CUDA encoder is byte-identical to the PyTorch reference on a fixed input and runs at **1047.9 M elem/s against the torch search's 10.7** on a 5632×2048 weight. - `tests/unit/torch/quantization/ -k 'ggml or iq1 or iq2 or iq_'` — 99 passed - `tests/gpu/torch/quantization/test_iq_formats_cuda.py` — 21 passed (7 checks × 3 formats) - `tests/unit/recipe/test_presets.py` — passing; `general/ptq` now holds 29 recipes, `ptq.md` updated - reconstruction error decreases monotonically with bit width, pinned by a test Pre-existing failures in `tests/unit/torch/export/` and `test_autoquant.py` are `transformers`/`torchvision` import problems in my environment — identical counts with and without this change. ### A finding about already-merged code Checking the new kernel against its PyTorch reference at 4096 blocks showed that **CUDA and torch encoders disagree on roughly 1 block in 6000 — including the already-merged `iq2_xs`**, at 0.0163% against IQ2_XXS's 0.0000%. Root cause: both compute `xnorm − 2·scale·dot + scale²·qnorm`, but CUDA fuses it with `fmaf` while torch uses separate ops; where two local scales fall within a float32 ULP the roundings pick different sides. Adjudicated against float64, neither path is better (5 to 6). Worst-case cost is **1.48e-08** relative reconstruction error, and run-to-run determinism on a given device holds. This is pre-existing, not introduced here — `test_iq2_xs_cuda.py` asserts exact byte parity but on a 16-block weight where ties essentially never arise. I have **not** changed that test; rewording a guarantee on merged code belongs in its own change. The new shared GPU tests assert exact parity on a small fixed input and compare reconstruction error at scale. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — the new codebook is a GGML table, carried in `codebooks.py` beside the existing ones so the MIT-licensed surface stays in that one file, with the source revision recorded. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information First of three; **IQ2_S** and **IQ1_M** follow and build on this branch. Replaces #2505, which carried all three at once. Follows #2446 / #2447 / #2448 / #2449, which landed IQ1_S and IQ2_XS. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added IQ2_XXS weight-only quantization, including CUDA acceleration and support for Hugging Face and Megatron exports. - Added the `general/ptq/iq2_xxs` recipe. It requires no calibration data and supports eligible layers with a weight dimension divisible by 256. - Updated the PTQ recipe catalog to list IQ1_S, IQ2_XXS, and IQ2_XS at approximately 1.56, 2.06, and 2.31 bits per weight, respectively. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
25d8c91762 |
Add guidance on keeping skill updates concise to AGENTS.md (#2523)
### What does this PR do? Type of change: documentation Adds an `## Updating skills` section to `AGENTS.md` (symlinked as `CLAUDE.md`). Skills are loaded into agent context, so each extra line costs tokens every time the skill runs. The new guidance tells the agent to: - Keep skill edits concise: add only what changes agent behavior, and tighten existing text instead of appending more. - Do a final compression pass over the skill diff before opening a PR: drop unnecessary explanations and examples, cut redundancy, and merge overlapping guidance. ### Usage N/A — no API or flag change. ### Testing `pre-commit run --files AGENTS.md` (markdownlint and the other applicable hooks pass). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A <!-- documentation-only change --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- agent instructions only, not user-facing --> - Did you get Claude approval on this PR?: ❌ <!-- will run /claude review if reviewers want it --> ### Additional Information Follows #2494, which added the PR sizing guidance to the same file. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added guidance for keeping skill updates focused on behavior changes and reviewing edits for unnecessary detail. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
7a35cada39 |
Add guidance on sizing and splitting PRs to AGENTS.md (#2494)
### What does this PR do? Type of change: documentation Adds a `## Sizing and splitting PRs` section to `AGENTS.md` (symlinked as `CLAUDE.md`) so AI-assisted work stops producing one giant PR that nobody wants to review. The new guidance tells the agent to: - Keep each PR that goes up for review under ~500 changed lines of source, and check the size before opening. - Propose the split *before* opening an oversized PR rather than after. - Split on file/directory/module boundaries first and fall back to feature boundaries (enabling refactor first, then one PR per behavior it unlocks). - Keep the series acyclic and linearly ordered — no circular dependencies between sub-PRs — and state the merge order. - Prefix sub-PR titles with `[x/N]` so reviewers know the PR is one slice of a planned split, and link the siblings. - Make every sub-PR stand on its own: it builds, carries unit tests for the code it introduces, and passes CI without the later PRs. - Optionally submit the whole change as a reference-only draft PR for the big picture, cross-linked with the sub-PRs. ### Usage N/A — no API or flag change. ### Testing `pre-commit run --files AGENTS.md` (markdownlint and the other applicable hooks pass). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A <!-- documentation-only change --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- agent instructions only, not user-facing --> - Did you get Claude approval on this PR?: ❌ <!-- will run /claude review if reviewers want it --> ### Additional Information This PR is itself well under the new budget (26 added lines in one file), so no split applies. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Documentation** - Updated pull request sizing guidance to allow oversized changes when they cannot be meaningfully split. - Clarified that draft aggregate pull requests are optional and intended for reference only when splitting would reduce clarity. - Added examples covering self-contained changes and new models or backends without a functional intermediate state. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a7166965e3 |
Drop GDPVal support from the evaluation skill (#2470)
### What does this PR do? Type of change: deprecation (agent skill) Removes GDPVal support from the `evaluation` agent skill. Its task recipe, example config and Apptainer SIF build helper are deleted. GDPVal was not cleanly separable, so this is not just a delete: - **MRCR depended on GDPVal's infrastructure.** `scripts/nel-gdpval.sh` was a generic pinned-0.2.6 `nel` launcher that only happened to be GDPVal-named, and MRCR ran through it; `references/gym-gdpval.md` documented the gym bootstrap machinery (prepare/reap, `install_on_the_fly` pin↔container coupling, the trust env vars) that both examples share. So the shared parts are kept and renamed rather than dropped: `scripts/nel-gym.sh` and `references/gym.md`. The launcher pin itself is unchanged (0.2.6), as are its env-override semantics. - **GDPVal is an AA-suite member**, so the "AA rule" in `SKILL.md` and `references/quantization-benchmarks.md` had to change. An "AA" / "Artificial Analysis" request now generates the `aa/` tasks as one multi-task config and nothing else. Both files now state that the resulting set omits GDPVal and is therefore not directly comparable to a published AA Index — report per-task scores rather than an aggregate. `SKILL.md` also tells the agent to say GDPVal is unsupported rather than reconstruct a config from an older copy of the skill. Everything GDPVal-specific is gone from `references/gym.md`: the SIF sandbox and its silent-unsandboxed-exec failure mode, the 3-member judge panel, rubric vs. comparison scoring, the deliverables/MLflow `*cache*` trap as a GDPVal concern, and Stirrup-agent deploy sizing. `recipes/env.example` loses `GDPVAL_SIF_DIR`, `GDPVAL_MAX_TURNS` and `TAVILY_API_KEY`, and gains a documented `NEMO_EVALUATOR_TRUST_UNLISTED_TASKS` (required by every gym task, previously only mentioned in prose). One correction carried along: the old reference said the bootstrap pins `ray==2.49.2`, but `example_mrcr.yaml` actually pins to whatever ray version the image already carries. `references/gym.md` now describes what the template does. ### Usage Not an API change. The skill-facing surface that moved: ```text scripts/nel-gdpval.sh -> scripts/nel-gym.sh (NEL_GDPVAL_* -> NEL_GYM_*) references/gym-gdpval.md -> references/gym.md tests/test_nel_gdpval.py -> tests/test_nel_gym.py ``` ### Testing - `pytest plugins/modelopt/skills/evaluation/tests/` — 1 passed. The launcher test was renamed rather than deleted: it is the only coverage for the pinned launcher, which MRCR still depends on. - `pre-commit run --files <changed>` — clean. `sync-claude-skills` fails in my working copy because `.claude/skills/` is a read-only harness mount there; it is unrelated to this change (it trips on `speculative-decoding`) and was skipped for the commit. - Grepped the repo for `gdpval` (case-insensitive): the only remaining hits are the deliberate "no longer supported" notes in `SKILL.md` and `references/quantization-benchmarks.md`. No dangling pointers to the deleted files, and no other skill referenced GDPVal. - Checked `tests/evals.json` for both `evaluation` and `day0-release`: no eval case expected a GDPVal companion config, so no expectations needed updating. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ❌ — anyone with a saved GDPVal config keeps it, but the skill no longer generates one and the SIF helper is gone. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — existing launcher test renamed and kept passing. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — under **Deprecations**. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information A follow-up is needed in the modelopt-internal repo: `modelopttools:eval-config` Step 3c is the GDPVal SIF / comparison-mode conversion checklist, and Step 3d names the gym image. Step 3c is now dead and the pointers this skill used to make into it are gone. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Evaluation Updates** * Added standalone NeMo Gym support for MRCR tasks with pinned launcher and Gym configuration requirements. * Added shared guidance for Gym setup, validation, execution, and recovery. * AA requests now generate only `aa/` tasks and report per-task scores. * **Removed Support** * Removed GDPVal evaluation recipes, documentation, task guidance, Apptainer helper tooling, and AA-suite inclusion. * Updated default quantized-checkpoint validation recommendations to exclude GDPVal. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b16356776b |
Build the GGML IQ packing kernels as a single CUDA extension (#2462)
### What does this PR do? Type of change: Code refactoring `#2448` added the GGML IQ packing kernels as **two** torch extensions, `modelopt_cuda_ext_iq1_s` and `modelopt_cuda_ext_iq2_xs`. This merges them into one, `modelopt_cuda_ext_ggml`. The existing per-extension split in `extensions.py` exists for reasons that don't apply to the IQ formats: `get_cuda_ext` gates on CUDA `>=11` while `_fp8`/`_mx` gate on `>=11.8`, and `_mx` needs `--use_fast_math`, which must not reach the base `tensor_quant` kernels. `get_cuda_ext_iq1_s` and `get_cuda_ext_iq2_xs` differed in none of that — same `>=11.8` gate, same `-O3` flags, same `common.cuh` — so the split only compiled the shared header twice, ran nvcc twice, and grew the loader, `__getattr__`, and `precompile()` once per format. With IQ2_XXS / IQ3_S / IQ4_NL plausibly following, that scales badly. Changes: - New `ggml/ggml.cpp` holds both host-side validation wrappers and the single `PYBIND11_MODULE`, binding `iq1_s_pack` and `iq2_xs_pack` (previously each module exported a bare `pack`). Deletes `ggml/iq1_s.cpp` and `ggml/iq2_xs.cpp`; the validation logic and docstrings carry over unchanged. - `get_cuda_ext_iq1_s` + `get_cuda_ext_iq2_xs` → `get_cuda_ext_ggml`, which builds `ggml.cpp`, `iq1_s.cu`, and `iq2_xs.cu` together. The retry-on-`raise_if_failed` semantics of the old getters are preserved. - Each format keeps its kernels in its own translation unit, so adding a format is a new `.cu` plus one `module.def` — no new extension, loader, or `precompile()` line. No caller outside `extensions.py` and its tests referenced the old getters on `main`, so nothing else changes. **Note for the follow-up PRs in the `#2448` series (`#2446`/`#2447`/`#2449`): the codec layer should call `get_cuda_ext_ggml().iq1_s_pack(...)` / `.iq2_xs_pack(...)` instead of `get_cuda_ext_iq1_s().pack(...)` / `get_cuda_ext_iq2_xs().pack(...)`.** ### Usage ```python from modelopt.torch.quantization.extensions import get_cuda_ext_ggml ext = get_cuda_ext_ggml(raise_if_failed=True) iq1_s_payload = ext.iq1_s_pack(weight, iq1s_grid) # uint8 [numel / 256, 50] iq2_xs_payload = ext.iq2_xs_pack(weight, iq2xs_grid, scales) # uint8 [numel / 256, 74] ``` ### Testing Ran on a single H200 NVL (TRT-LLM `1.3.0rc27.dev202609170000` container), building the merged extension from scratch: - `pytest tests/gpu/_extensions/test_torch_extensions.py` — **24 passed** (6:44). This is the full existing IQ suite (zero-block layout, encode, dtype rejection, row-straddling rejection, invalid/negative-zero scales, byte-exact dtype equivalence, and the brute-force optimality round-trip) reparametrized onto the merged module, plus the untouched `modelopt_cuda_ext` / `_fp8` / `_mx` load tests. - Verified `precompile()` loads all four extensions and that the merged module exports exactly `iq1_s_pack` and `iq2_xs_pack` with the expected arities. - Off-GPU: compiled the three sources directly and linked them into one `.so` to confirm no duplicate-symbol collisions between the two `.cu` translation units. - `pre-commit run --files ...` passes on all changed files (ruff, mypy, clang-format, bandit, license headers). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — the removed getters were added in `#2448` (merged today, unreleased) and have no callers outside this file's own tests. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new code or dependencies; the moved wrappers keep their original attribution. - Did you write any new necessary tests?: ✅ — existing coverage reparametrized onto the merged module; no behavior change to test. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — internal refactor of an unreleased, not-yet-wired-up API. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information Follow-up to #2448. Merge before the remaining PRs in that series (#2446, #2447, #2449) land, so the codec layer is written against `get_cuda_ext_ggml` and no rename is needed afterwards. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added IQ1_S packing support through the GGML CUDA extension. - Added a unified GGML extension loader for IQ1_S and IQ2_XS packing. - Improved extension loading reliability when a cached extension is unavailable. - **Changes** - Renamed the IQ2_XS packing binding from `pack` to `iq2_xs_pack`. - Consolidated IQ1_S and IQ2_XS extension access under the shared GGML loader. - Updated GPU validation and coverage to use the unified extension interface. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
542012d4d5 |
Update CODEOWNERS (#2463)
Add modelopt-torch-kernels-codeowners <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Added code ownership coverage for the `modelopt/torch/kernels` directory. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
655f94c207 |
Document the nvfp4_act_headroom calibration variant in ptq.md (#2439)
### What does this PR do?
Type of change: documentation
`general/ptq/nvfp4_act_headroom-kv_fp8_cast.yaml` appears in the
shipped-recipes
table in `modelopt_recipes/ptq.md`, but the **Calibration variants**
section —
which documents `max`, `mse`, `input_scale1`, `gptq`, and the
`layerwise`
variants — had no entry for it. Someone scanning that section for "which
calibration do I pick when NVFP4 W4A4 regresses?" only found `mse`,
which
searches **weight** scales and so cannot help when the loss comes from
activation clipping.
This adds the missing entry: the scale formula
(`amax = max(rho * anchor, upper)`) and its defaults, the fact that it
costs one
calibration pass and exports a standard NVFP4 checkpoint with coverage
identical
to `nvfp4_default-kv_fp8_cast`, and the symptoms that should route you
here
rather than to a weight-side calibration — an A16 ablation clears the
regression
while `mse` does not, the symptom is behavioral (verbose or runaway
generations,
hitting the generation cap) rather than a flat score drop, inference
contexts run
longer than the calibration set, or a few rare blocks dominate the
activation
error. MoE experts-only scopes are called out as the common case.
It also extends step 3 of **Choosing a general recipe** so the
escalation path
reads `mse` first, then `nvfp4_act_headroom` when the evidence points at
activations rather than weights.
**Evidence.** The guidance comes from a GLM-5.3-Flash NVFP4 experts-only
W4A4
root-cause study on SciCode (temperature 1.0), which established
causally that
activation quantization at the routed-expert `down_proj` input drove a
large
generation-length blow-up. Swapping `max` for `nvfp4_act_headroom` cut
the median
generation-length regression versus source from +38% to +19% and the
mean from
+19% to +4%, with no capped generations. The entry states plainly that
this was
the best strict-W4A4 result in that study but still missed the p50/p75
near-lossless gate, so headroom is presented as a strong first lever for
activation-driven regressions rather than a guaranteed fix, with a note
that
`rho` should be swept.
### Usage
No API or recipe change; the recipe already ships. This PR only
documents when
to select it over plain `max`:
```python
from modelopt.recipe import load_recipe
cfg = load_recipe("general/ptq/nvfp4_act_headroom-kv_fp8_cast")
```
### Testing
Docs-only change; no code paths touched, so no new or updated tests.
- `pre-commit run --files modelopt_recipes/ptq.md` — all applicable
hooks pass,
including `markdownlint-cli2` and `check-modelopt-recipes`.
- Re-read the rendered section to confirm the new bullet nests correctly
in the
existing `Calibration variants` list and that surrounding entries are
unchanged.
- Cross-checked every claim against the implementation
(`modelopt/torch/quantization/calib/nvfp4_act_headroom.py`), the recipe
YAML,
and the existing `CHANGELOG.rst` entry, so the documented defaults
(`anchor_percentile=1`, `upper_percentile=99.99`, `rho=16384`) and the
NVFP4-input-quantizer-only scope match the code.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- docs-only; the
algorithm's tests already live in
tests/unit/torch/quantization/test_nvfp4_act_headroom.py -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- nvfp4_act_headroom already has a CHANGELOG entry from the PR
that added it; a docs-only follow-up is not changelog-worthy. -->
- Did you get Claude approval on this PR?: ❌ <!-- not run; docs-only
change -->
### Additional Information
The entry deliberately does not sell this on accuracy: in that study the
quantized subtask accuracy (56.80%) was *above* the source checkpoint
(51.18%),
so what headroom recovered was generation-length behavior, not accuracy.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added an NVFP4 activation headroom calibration option with
configurable percentile-based scaling.
* Supports standard NVFP4 checkpoint export with a single calibration
pass.
* Applies to dynamic-block NVFP4 activation quantizers while keeping
weight-scale configuration independent.
* Added guidance for addressing activation-related W4A4 accuracy
regressions, including calibration coverage and recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
b80e164472 |
Carry a checkpoint's ModelOpt PTQ run into its eval's MLflow run (#2407)
### What does this PR do? Type of change: documentation (agent skill) Follow-up to #2374, which made a tracked `hf_ptq.py --mlflow` run leave `.experiment.json` in the checkpoint it writes, naming the MLflow run that quantized it. Nothing on the eval side read that file, so an eval of a quantized checkpoint recorded no link back to the quantization that produced it. The `evaluation` skill now reads it. **Step 3** `cat`s the file as soon as `checkpoint_path` is known; **Step 4** carries it into `export.mlflow`: | `.experiment.json` field | goes to | | --- | --- | | `experiment_name` | `export.mlflow.experiment_name`, verbatim | | `run_name` / `run_id` / `run_url` | tags `modelopt_run_name` / `modelopt_run_id` / `modelopt_run_url` | | `tracking_uri`, `experiment_id` | deliberately unmapped | The `modelopt_` prefix keeps them from reading as the eval's own run. Values are quoted, or an all-digit `run_id` (or a `run_name` like `20260910`) is YAML-coerced to an int or a date. **`tracking_uri` is deliberately not inherited**, and the consequence is documented rather than implied. Evals go to whatever `$MLFLOW_TRACKING_URI` names. When the PTQ tracked to a different server — the usual case, since `modelopttools:eval-config` points evals at `mlflow.frontier-evals` while `hf_ptq --mlflow` typically writes to `mlflow-modelopt` — the inherited name creates a *same-named, empty* experiment on the eval server, and `modelopt_run_url` is the only route back to the PTQ run. Evals of one checkpoint still group under a stable name. Both cases are spelled out in Step 4 so nobody goes looking for the PTQ run beside the eval. Also updated: `references/nel-next.md` (nel-next configures MLflow through its own `export_config.mlflow`) and `accessing-mlflow` (the `tags.modelopt_run_id` query that closes the loop — without it the tag is write-only). ### Usage ```bash cat "$CHECKPOINT_PATH"/.experiment.json # absent → name the experiment as usual ``` ```yaml export: mlflow: tracking_uri: ${oc.env:MLFLOW_TRACKING_URI} # NOT the file's experiment_name: alice/hf_ptq/Qwen3.8-27B-NVFP4 # verbatim from .experiment.json tags: modelopt_run_name: '20260910-175422' modelopt_run_id: '7bec239a3a154970b062f3024a5ff20e' modelopt_run_url: 'https://<modelopt-mlflow-server>/#/experiments/36/runs/7bec239a3a154970b062f3024a5ff20e' ``` Finding every eval of a checkpoint a given PTQ run produced: ```python MLflow:query_runs(experiment_id, "tags.modelopt_run_id = '<ptq_run_id>'") ``` ### Testing Docs-only, so it was tested by having an agent follow the new wording end to end and checking what NEL actually submitted. A GPQA Diamond config (single repeat) was generated against an NVFP4 checkpoint carrying a `.experiment.json`, dry-run clean, then submitted on a SLURM cluster. Verified in the `export_config.yml` heredoc **inside the submitted `run.sub`** — the file the export job actually consumes, not just the source YAML: ```yaml experiment_name: chenjiel/hf_ptq/Qwen3.8-27B-NVFP4 # inherited verbatim tags: modelopt_run_name: 20260910-000000-synthetic-test-fixture modelopt_run_id: 00000000fake0000fake0000fake0000 modelopt_run_url: https://<modelopt-mlflow-server>/#/experiments/999/runs/... tracking_uri: https://<eval-mlflow-server>/ # env var, not the file's ``` This confirmed the open question behind the change: NEL accepts arbitrary `modelopt_*` keys in `export.mlflow.tags`. The `.experiment.json` used was a synthetic fixture, since the checkpoints predate #2374. Repeated on a **second cluster with a second checkpoint** (different account, QoS-based scheduler, FP8 rather than NVFP4) and carried through to completion, so the result is confirmed as *stored by MLflow* rather than merely submitted. Canary scored `gpqa pass@1 symbolic_correct: 90.0` (card: 88.01) with `no_answer: 0.0`, then the auto-export wrote: ``` experiment chenjiel/hf_ptq/Qwen3.8-27B-FP8 (id 2106) run eval-<invocation>-ns_gpqa FINISHED modelopt_run_name 20260910-000000-synthetic-test-fixture modelopt_run_id 00000000fake0000fake0000fake0000 modelopt_run_url https://<modelopt-mlflow-server>/#/experiments/999/runs/... ``` Experiment 2106 did not exist on the eval server beforehand — the export created it. That is the shadow-experiment case this PR documents, observed rather than predicted. The same exercise corrected the first draft of the wording: it had claimed the rule makes a quantization and its evals "sit together on one server", which is false whenever the two servers differ. Confirmed by API — `chenjiel/hf_ptq/Qwen3.8-27B-NVFP4` returns `RESOURCE_DOES_NOT_EXIST` on `mlflow.frontier-evals`. Rewritten as described above. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — additive guidance; the absent-file path is unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — skill documentation; validated by a real submission, above. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ — agent-skill guidance, not a library/example feature. - Did you get Claude approval on this PR?: ❌ ### Additional Information Follow-up to #2374. Three adjacent problems were found while validating and deliberately **left out of scope** — happy to split them into their own PR: 1. `recipes/examples/example_eval.yaml` puts `sbatch_comment` under a top-level `cluster:` key, but the executor reads `cfg.execution.sbatch_comment` (`executors/slurm/executor.py:688`) and SKILL.md Step 1 says the same. As shipped the template silently drops the idle-GPU reaper exemption, so long evals get reaped. 2. Step 4's `execution.gres` bullet names the wrong key for `internal/slurm/<cluster>` configs, which supply `gpus_per_node` and no `gres`; following it literally emits a redundant flag. 3. Step 3's `max_new_tokens` rules can contradict each other — "take the card's highest" can equal `max_model_len`, leaving no room for the prompt. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Documentation** - Clarified MLflow evaluation configuration requirements, including literal experiment and sampling values, CPU partition settings, and ModelOpt provenance. - Documented validation and fallback behavior for unresolved `${...}` placeholders in provenance values. - Added guidance for querying cross-server evaluations by run ID and handling same-named local experiments. - Clarified that provenance files are optional, independent of quantization detection, and do not control deployment flags. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d69e93a72b |
Record the MLflow run that produced a checkpoint in .experiment.json (#2374)
### What does this PR do?
Type of change: new feature
A tracked `hf_ptq` run already tags itself with the checkpoint it writes
(`checkpoint_path`), so a run can be followed to its output. The reverse
was missing: given a checkpoint on disk, there was no way to find the
run that quantized it without searching the tracking server by path.
A tracked run now writes `.experiment.json` into `--export_path` naming
the experiment, the MLflow run id and the run URL, and uploads the same
bytes as the `experiment.json` artifact so a downloaded artifact set is
self-describing. `MlflowRunLogger` gains a `run_info` property carrying
that identity, with the tracking URI credential-masked the way `run_url`
already was.
Two deliberate behaviours:
- **Written from a `finally`**, so a run that crashes after export still
leaves the pointer behind.
- **Skipped when the export directory is absent** — a run that exported
nothing has nowhere to put it, and creating the directory would suggest
a checkpoint that does not exist. The artifact is still uploaded in that
case, so a failed run is traceable from the server side.
A failed local write warns and continues rather than failing the job,
consistent with the rest of the MLflow path. Only the main rank writes,
since the logger is inert on other ranks.
### Usage
```bash
python hf_ptq.py --pyt_ckpt_path Qwen/Qwen3.5-0.8B --qformat fp8 \
--export_path /tmp/qwen35-fp8 --mlflow https://<your-mlflow-server>
```
```console
$ cat /tmp/qwen35-fp8/.experiment.json
{
"tracking_uri": "https://<your-mlflow-server>",
"experiment_name": "alice/hf_ptq/Qwen3.5-0.8B-fp8",
"experiment_id": "36",
"run_id": "7bec239a3a154970b062f3024a5ff20e",
"run_name": "20260910-175422",
"run_url": "https://<your-mlflow-server>/#/experiments/36/runs/7bec239a3a154970b062f3024a5ff20e"
}
```
```python
# checkpoint -> run
import json, mlflow
info = json.load(open("/tmp/qwen35-fp8/.experiment.json"))
mlflow.set_tracking_uri(info["tracking_uri"])
run = mlflow.get_run(info["run_id"])
```
### Testing
**Unit** — `tests/unit/torch/utils/test_mlflow.py` (61 passed):
`run_info` contents before/after the run opens, the defaulted run name
being reported rather than left blank, and credential masking of the
tracking URI.
**Example** — `tests/examples/hf_ptq/test_hf_ptq_args.py` (27 passed):
the file landing in the checkpoint and on the server with identical
content, the failed-run path, the no-export path, and untracked runs
writing nothing.
**Real runs**, 1x H200, `Qwen3.5-0.8B` FP8 PTQ,
`tensorrt-llm/release:1.3.0rc26`:
- Against a local MLflow server — checkpoint copy and uploaded artifact
byte-identical; artifacts on the run were `command.txt`,
`experiment.json`, `logs/hf_ptq.log`, `summary/quant_summary.txt`,
`version.txt`.
- Against the internal `mlflow-modelopt` server (experiment
`chenjiel/hf_ptq/Qwen3.5-0.8B-fp8`, run
`7bec239a3a154970b062f3024a5ff20e`) — same result, confirming artifact
upload against a real backend. Reading `.experiment.json` back and
calling `mlflow.get_run(run_id)` resolved to `FINISHED` with
`checkpoint_path` pointing at the export directory.
- Crash path exercised for real when a first attempt died on a gated
calibration dataset: no export directory created, `experiment.json`
still uploaded, run closed `FAILED`.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — new entry under `*Misc*` in the open 0.48.0 section, matching where
the MLflow entries sit in 0.47.0.
- Did you get Claude approval on this PR?: ❌
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Exported checkpoints now record experiment and run traceability
metadata in `.experiment.json`.
* Checkpoint metadata is uploaded with opened MLflow runs, including
runs where export fails.
* Active MLflow run details—including identifiers, resolved run name,
URL, and tracking server—are available with credentials redacted.
* **Bug Fixes**
* Improved handling of failed, untracked, and pre-existing exports to
prevent inherited metadata pointers.
* **Documentation**
* Updated MLflow integration guidance and changelog information for
checkpoint metadata.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
a74054ab2b |
Let callers add MLflow tags to a fakequant serve's run (#2364)
### What does this PR do?
Type of change: new feature
The quantization run records what this library can see — the model, the
checkpoint, the vLLM and ModelOpt versions — but nothing about the
harness that launched it. A downstream tool that wants its own revision,
a sweep id, or a ticket number on the run has no way to put it there
today:
- `_run_tags()` returns a fixed dict
- `quant_config` (which becomes the run's params) is a hardcoded set of
`QUANT_*` variables
- MLflow itself has no environment variable for arbitrary tags
`MODELOPT_MLFLOW_EXTRA_TAGS` takes comma-separated `key=value` pairs and
merges them into the run's tags.
Two details worth a reviewer's attention:
**It joins `MLFLOW_ENV_VARS`.** A Ray-backed serve receives only the
variables named there, and the tracker runs in the rank-0 worker —
omitting it would make the feature silently do nothing under Ray.
**Caller tags are merged first**, so the library's own keys (`tool`,
`model`, `checkpoint_path`, `vllm_version`) are written over them and
keep describing the run truthfully whatever a caller sends.
`key=value` rather than JSON, learned from a live run: the variable
reaches the worker through a shell `export VAR="..."`, and JSON's own
double quotes terminate that quoting —
```
export MODELOPT_MLFLOW_EXTRA_TAGS_732b_DEPLOYMENT="{"internal_version": "4d8c"}"
```
arrived as `{`. A quote-free format survives verbatim and needs no
`json` import or exception handling. Splitting on the first `=` keeps
values that contain one, such as a URL with a query string.
### Usage
```bash
export MODELOPT_MLFLOW_EXTRA_TAGS="modelopt_internal_version=49fa29d5,sweep=kv-study"
python3 vllm_serve_fakequant.py "$MODEL" --mlflow https://your-mlflow-server/ ...
```
### Testing
Unit-level, over the helper: unset and empty variable, one and several
pairs, surrounding whitespace, an empty value, an entry with no `=`, a
trailing comma, and a value containing `=`. None raise; malformed
entries warn and are skipped.
End to end on a real fakequant serve (Nemotron-3-Nano-30B-A3B BF16,
`NVFP4_DEFAULT_CFG`, TP=8, Ray executor, vLLM 0.15, SLURM):
```
modelopt_internal_version '49fa29d5'
modelopt_version '0.47.0rc0.post32+gd38ed5ead'
git_sha 'd38ed5ead'
quant_cfg 'NVFP4_DEFAULT_CFG'
```
The tag was written by the `RayWorkerWrapper` process, which exercises
the whole path — env var → shell export → `--container-env` → raylet →
Ray actor → `_run_tags` — and confirms the `MLFLOW_ENV_VARS` entry is
doing its job. Also verified that the emitted payload survives a shell
export round-trip unchanged.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ <!--- Additive; with the
variable unset the tags are exactly as before. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ <!--- Verified manually as
above; there is no existing test module for vllm_mlflow_utils. Happy to
add one if you would like it. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <!--- Small additive feature in an example; tell me if it warrants an
entry. -->
- Did you get Claude approval on this PR?: ❌
### Additional Information
Consumed by Model-Optimizer-Internal MR !141/!147, which sets the
variable so a fakequant eval records the same harness commit on both its
quantization run and its evaluation-score run.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
19de0075cb |
Forward kv_cache_free_gpu_memory_fraction to the lm_eval TensorRT-LLM engine (NVBug 6701763) (#2300)
### What does this PR do? Type of change: Bug fix `scripts/huggingface_example.sh --kv_cache_free_gpu_memory_fraction` has no effect on the `lm_eval` task: the value is parsed by `parser.sh`, printed, and then dropped. lm-eval's built-in `trtllm` backend (`lm_eval.models.trtllm_causallms.TRTLLM.__init__`, which this example switched to in #2066) accepts `**kwargs`, but builds `KvCacheConfig(enable_block_reuse=False)` and passes `LLM(...)` a fixed set of keys — `kwargs` is never merged in. So an extra `--model_args` entry is accepted by the CLI and silently discarded, and the KV cache is sized from TensorRT-LLM's default `free_gpu_memory_fraction=0.9`. There is no way to fix this from the caller: `--model_args` only yields scalars, so a `KvCacheConfig` object cannot be passed in either. On a GH200 that means ~119.6 GiB of KV cache (`119.55 / 0.9 ≈ 132.8 GiB free`), leaving 87.8 MiB free, and `prompt_logprobs` deserialization then OOMs asking for 2.82 GiB. `examples/llm_eval/lm_eval_trtllm.py` already exists to patch this backend (its `_parse_logprobs` misaligns TensorRT-LLM's `prompt_logprobs` by one). It now also injects the fraction into the `KvCacheConfig` the backend builds, defaulting to 0.8 — the same default `parser.sh` declares, and below TensorRT-LLM's 0.9. `huggingface_example.sh` passes the parsed value through in `--model_args`. Scoped deliberately to the `lm_eval` path: the `quant` smoke test and `mmlu` go through `modelopt.deploy.llm.LLM` (0.7, hardcoded) and `simple_eval`/`livecodebench` through `trtllm-serve` (0.9); those are left as they are. ### Usage ```bash # Via the example script (parser.sh default 0.8) scripts/huggingface_example.sh --model $HF_PATH --quant fp8 --tp 1 \ --tasks quant,lm_eval --lm_eval_tasks mmlu --lm_eval_limit 50 \ --kv_cache_free_gpu_memory_fraction 0.5 ``` ```bash # Standalone, via lm-eval's --model_args python lm_eval_trtllm.py --model trtllm \ --model_args model=<ckpt>,tokenizer=<tok>,max_input_len=4096,kv_cache_free_gpu_memory_fraction=0.5 \ --tasks mmlu --batch_size 8 ``` ### Testing - `pytest tests/examples/llm_eval/test_lm_eval_trtllm.py` — 21 passed (lm-eval 0.4.12, no GPU). - The new tests instantiate the **real** upstream `TRTLLM.__init__` through `create_from_arg_obj`, with `tensorrt_llm` and the tokenizer stubbed, and assert the engine receives `KvCacheConfig(enable_block_reuse=False, free_gpu_memory_fraction=0.5)`; that an unset key still yields 0.8 rather than 0.9; and that the patch does not outlive the constructor. Reverting the fix fails 3 of them. - Tripwire test asserts upstream still neither declares nor forwards the argument, so this shim gets deleted rather than silently kept once lm-eval fixes it. - `pre-commit run --files <changed>` clean (ruff, mypy, bandit, markdownlint); `bash -n` on the modified script. - Not run: the GPU end-to-end `tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8`, which exercises `lm_eval` through the modified script — no GPU in this environment. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — the `lm_eval` KV cache goes from TensorRT-LLM's 0.9 to 0.8, which is strictly more conservative; `parser.sh`'s declared default is unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information NVBug 6701763. The 0.9 default on this path arrived with #2066 and was documented as a known limitation in `examples/llm_eval/README.md` ("the KV cache uses 90% of free GPU memory rather than 70%"); that note is replaced by the working knob. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Fixed the TensorRT-LLM evaluation workflow so `kv_cache_free_gpu_memory_fraction` is correctly passed to the backend. - The setting now defaults to `0.8`, providing more predictable GPU memory allocation for KV-cache usage. - **Documentation** - Updated the TensorRT-LLM evaluation example and usage guidance to describe the KV-cache memory setting and its default behavior. - Updated the Hugging Face example to pass the configured KV-cache memory fraction. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
19ce447d62 |
[skill] evaluation: mandate 8 SciCode runs and report the mean (#2327)
### What does this PR do? Type of change: documentation SciCode's repeat count was pinned to `num_repeats: 1` in #1945, which dropped its effective sample count from the original `8` (set when the recipe was written in #1561) to `1`. #2254 later documented that a single run cannot gate on — scored single-shot at `temperature 1.0`, a paired comparison moved **3.92 pp and changed sign** once repeated, and the day-0 skill's own table records a DeepSeek-V4-Pro drop reading 2.96 pp (`REGRESSION`) at 1 run versus **-0.96 pp** (`PASS`) at 8. But #2254 left the remedy as a judgment call — *"pool until the standard error is below the threshold, or report the task `INDETERMINATE`"* — so a one-run SciCode number was still reportable. This restores the original avg-of-8 statistics as a hard requirement, without reintroducing the in-run repeats that #1945 removed for a reason (repeating inside one run multiplies exposure to code-execution sandbox errors). The task keeps `num_repeats: 1` per run; the repeat budget of 8 is spent as **8 independent submissions, reported as their mean** — 8 per side for a comparison, `INDETERMINATE` below 8. Changes: - **`recipes/tasks/aa/scicode.md`** — mandatory *8 runs* section: how to submit the 8 inside a multi-task AA config, fresh-run requirement (no `run.sub` replay off a warm cache), duplicate-score check, per-run validation before averaging, and mean + `stdev/sqrt(8)` + run-count reporting. - **`compare-results` / `day0-release`** — 8 runs per side is a floor, not a variance-dependent choice. - **`references/quantization-benchmarks.md`** — the repeat-count table still listed SciCode at `num_repeats: 8`, stale since #1945. - **`evaluation/SKILL.md`** — AA rule points at the 8 submissions; the walltime section distinguishes them from the forbidden practice of splitting a heavy task across configs to dodge the 4h cap. - **`evaluation/tests/evals.json`** — behavioral eval case for the rule. ### Usage ```bash # Full AA suite (= SciCode run 1), then SciCode alone for runs 2-8. nel run --config <cfg>.yaml for _ in $(seq 7); do nel run --config <cfg>.yaml -t ns_scicode; done # Report the mean of scicode_pass_at_1_avg-of-1_subtask_accuracy over the 8 runs, # plus stdev/sqrt(8) and the run count. ``` ### Testing - `pre-commit run --files <changed files>` — all hooks pass (`markdownlint-cli2`, `check json`, symlink sync), no hook-applied modifications. - `json.load` on `evaluation/tests/evals.json` parses; 4 cases, existing three unchanged (append-only diff, original formatting preserved). - Cross-checked every remaining SciCode reference in the plugin (`grep -rn -i scicode plugins/modelopt/`) so no doc still claims `num_repeats: 8` or a variance-dependent pool size. **Ran the protocol end-to-end** on Qwen3.8-27B-FP8 (gcp-nrt, 8xB200, `temperature 1.0`), 8 independent `nel run` submissions, all `COMPLETED`, each scoring the full 80 problems / 338 subtasks. MLflow experiment 2017: | passed/338 | 168 | 167 | 166 | 164 | 163 | 161 | 157 | 155 | |---|---|---|---|---|---|---|---|---| | score | 49.70 | 49.41 | 49.11 | 48.52 | 48.22 | 47.63 | 46.45 | 45.86 | **Result 48.11, stdev 1.39, stderr 0.49.** Pooled 1301/2704 subtasks reproduces 48.1139 exactly. This is the evidence for the change: **the spread is 3.85 pp and the worst single run sits 2.26 pp from the mean**, so against a 1% gate any one of these eight, reported alone, would have been defensible and wrong. Previously this PR rested on the variance figures inherited from #2254; it now rests on a measured pool. Also validated by that campaign: an agent given only "run a full SciCode eval, follow the evaluation skill" — with no mention of the protocol — read the recipe, submitted 8 runs and reported the mean with a standard error. And the corrected score key is what made the numbers harvestable at all; the `avg-of-1` name returns nothing. Two operational findings from the run, one of which is folded into the recipe: - An MLflow export job timed out (`nel-export-ns_scicode.0`, elapsed `00:30:11` against a `00:30:00` limit) because all 8 exports hit the CPU partition at once and each reinstalls the launcher. Caused by the fan-out this PR mandates, so the recipe now warns about it and says to re-submit `export.sbatch` rather than treat the run as lost. - Out of scope, filed for separate fix: `.claude/agents` is a 0-byte read-only placeholder, so the `monitor` skill's instruction to create a session registry under it fails with `Not a directory`. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — added `scicode-eight-run-average` to `evaluation/tests/evals.json` - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — agent-skill guidance, not a user-facing library change - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information Restores the sampling behavior of #1561 while keeping the sandbox-load fix from #1945 and honoring the variance evidence from #2254. Note: `origin/feature/puzzletron_v2` still carries the pre-#1945 version of this file (`num_repeats: 8`, single submission) and rewrites it to source the repeat count from `examples/llm_eval/task_contracts.yaml`. If that branch lands it will need to be reconciled with this protocol. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Documentation** - Clarified SciCode evaluation requirements: eight independent, valid runs per comparison side with one repeat each. - Standardized score reporting using per-run metrics, mean, standard error, and run count. - Added provenance checks and replacement runs for invalid or replayed results, including timeout recovery. - Fewer than eight valid runs are reported as **INDETERMINATE**. - Updated benchmarking guidance to distinguish separate submissions from repeated runs. - **Tests** - Added coverage for eight-run averaging, validation, replacement runs, and **INDETERMINATE** outcomes. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
73d7784223 |
docs(eval-skill): add NVFP4 model-card sampling reference (#2224)
### What does this PR do? Type of change: documentation Adds `plugins/modelopt/skills/evaluation/references/nvfp4-modelcard-sampling.md` — the published `temperature` / `top_p` / max generation length for the **2026 NVFP4 checkpoints under [huggingface.co/nvidia](https://huggingface.co/nvidia/models) that disclose them** — and points `SKILL.md` Step 3 and `model-card-research.md` at it. **Why.** Config generation currently re-derives sampling params per run from one model card, which is slow and silently wrong in two ways: 1. **Quickstart boilerplate reads as an eval setting.** Most NVIDIA NVFP4 cards paste a TensorRT-LLM snippet containing `SamplingParams(temperature=0.8, top_p=0.95)`. That string is byte-identical across `Llama-3.1-8B`, `Llama-3.3-70B`, `Llama-4-Scout`, `Phi-4-reasoning-plus`, `Phi-4-multimodal`, `Qwen2.5-VL-7B` and six `Qwen3-*` repos. It is template text, not what the accuracy table was measured with — but it is the most prominent `temperature=` in the card. 2. **A generic fallback beats an available answer.** When a card is silent, Step 3 falls back to 65536/16384 even where a sibling in the same family publishes an exact value. **Scope — 25 rows.** All 69 NVFP4 checkpoints in the org were read. A row exists only for a 2026 **target** checkpoint whose card discloses usable settings, so absence means "read the card", not "not yet checked". Excluded by construction: pre-2026 releases, cards that publish nothing, and — via the curation regex `-NVFP4(-V\d+|-QAD)?$` — `-DSpark`/`-DFlash` speculative-decoding variants (verified against the target, so identical accuracy; they would only duplicate the base row), `-Eagle3` draft heads, and `-MLPerf-Inference-Closed-*` snapshots. | provenance | rows | meaning | | --- | --- | --- | | `eval` | 20 | card ties the values to its accuracy table — authoritative | | `rec` | 5 | recommended inference sampling, not tied to the eval | A `—` in one value column means that field specifically is unpublished — four rows give `temperature`/`top_p` but state no generation cap (`Mistral-Medium-3.5`, `Nemotron-3.5-Lightning`, `Nemotron-3-Super-120B`, `Nemotron-Labs-3-Elastic-30B`). Per-task exceptions are recorded where cards state them: GLM-5.2 GPQA Diamond `100000` vs `64000`; Qwen3.5-397B-V2 τ²-Bench Telecom `128000`; Qwen3.6 SciCode `temperature=0.6`; Kimi-K3 uncapped for Terminal-Bench. **Posture: the card is the source of truth; the table is a reference, not a constraint.** It is there to confirm a value you read, fill a gap when the card is silent, and catch a misreading — never to override what a card states. Listed and in agreement → proceed; listed and different → the card wins, re-read, surface the discrepancy. The `max_num_tokens` column records the card's *headline* cap, so Step 3's existing take-the-highest rule still governs when a card names more than one. Per-task `temperature`/`top_p` in the notes is precedent rather than mandate — engineers do tune sampling per benchmark — so the guidance is to follow the card and escalate only on a regime change (greedy vs sampled), not a nudge (`0.95` vs `1.0`). ### Usage Not an API change; the reference is consumed by the `evaluation` skill when generating a NEL config. ```yaml # nvidia/GLM-5.2-NVFP4 -> references/nvfp4-modelcard-sampling.md (provenance: eval) nemo_evaluator_config: config: params: max_new_tokens: 100000 # card: GPQA Diamond 100000, others 64000 -> take the highest temperature: 1.0 top_p: 0.95 ``` ### Testing **Coverage — enumeration cross-verified two ways.** Rather than paging `https://huggingface.co/nvidia/models?p=N` by hand, the candidate list came from the HF API (`author=nvidia&limit=1000` → 917 repos, one page), then verified against the website pagination: 918 unique model links across `p=0..31`, and the 68 NVFP4-named repos were **identical in both** (`api-only: []`, `web-only: []`). The committed curation regex was then re-run against that full set and confirmed to select every one of the 25 table rows plus 7 further 2026 target checkpoints whose cards disclose nothing — i.e. the documented filter reproduces the table. **Accuracy — every `eval` row diffed against its source sentence.** A script re-parsed the table and printed each row beside the matching card line. To be precise about what that covers: the diff was run when the table was larger, and all 28 machine-checkable *Benchmarked with* rows matched exactly (`Qwen3.6-35B-A3B` states its settings on the second line of a blockquote and was checked by hand). Every edit since has been a row **removal** — the scope cut to 2026 and the spec-decode removal — each verified by re-parsing the table before and after and confirming no surviving row's `temperature`/`top_p`/`max_num_tokens`/`provenance` changed. Extraction was restricted to *"Benchmarked with…"* / *"…were evaluated with…"* / *"We evaluate the model using…"* sentences, "Recommended Sampling" rows, and accuracy-table footnotes — never the quickstart snippets. **Two extraction traps found while building it**, both now in the file's refresh recipe as a second mandatory grep: some cards publish the cap **only** as a footnote beneath the accuracy table (`*Max OSL for evals can be as high as 64K`), and DeepSeek states sampling in its `## Input:` usage block rather than a *Benchmarked with* sentence. Neither is reachable from the obvious grep. **Hygiene.** `pre-commit run --files <the 3 files>` passes, including `markdownlint-cli2` and the `sync .claude/skills/ symlinks` hook; the reference resolves through both `.claude/skills/evaluation/references/` and `.agents/skills/evaluation/references/`. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A <!-- documentation-only; verification described under Testing --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- agent-skill documentation; no user-facing API or behavior change --> - Did you get Claude approval on this PR?: ✅ <!-- reviewed; findings addressed or answered below --> ### Additional Information Data collected 2026-08-20; the file is an explicitly dated snapshot and says to trust the card for anything newer. It ends with a refresh recipe (API enumeration → `-NVFP4(-V\d+|-QAD)?$` filter → `createdAt >= 2026-01-01` → card fetch → the grep phrasings that carry eval settings) so the table can be regenerated as new checkpoints ship — roughly four NVFP4 target checkpoints per month over 2026 so far. Review findings addressed: the mandatory-match framing was softened to reference-only, per-task sampling reframed as precedent with a human-escalation trigger, `mkdir -p cards` added to the recipe, the HF token no longer interpolated into a `curl` argument, and explicit uncapped generation (`max_new_tokens: null`) distinguished from an unpublished cap. The provenance question on the Nemotron-3.5-Lightning rows is answered in a review reply — the differing labels were correct (the base card says "Recommended Sampling", the sibling cards say "Benchmarked with"), and those sibling rows have since been removed as spec-decode duplicates. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d32c2c2a56 |
docs(eval-skill): add MRCR (NeMo Gym) benchmark (#2192)
### What does this PR do?
Type of change: documentation (agent skill)
Adds **MRCR** — OpenAI's Multi-Round Co-reference Resolution, a
long-context
retrieval benchmark — to the `evaluation` skill as a standalone NeMo Gym
task,
derived from the reviewed `nemotron_nano_v35_nvfp4_mrcr_gym` golden;
regroups the
gym tasks/examples under one `gym/` dir; and fixes several latent bugs
in the
shared gym command block that were found by running the benchmark
end-to-end.
MRCR tasks are long multi-turn conversations containing N near-identical
"needle"
responses; the model must reproduce the Nth verbatim behind a random
prefix.
Grading is deterministic (`SequenceMatcher.ratio()`, gated on the
prefix). Unlike
GDPVal it uses the `simple_agent`: no SIF, no judge, no Tavily —
`HF_TOKEN` is the
only secret, and the cost is context length (up to 1M tokens), not agent
turns.
**It is not an AA benchmark** and is never generated for an "AA"
request.
### Layout
| File | |
| --- | --- |
| `recipes/tasks/gym/mrcr.md` | new recipe — variants, 1M serving
envelope, canary, score extraction |
| `recipes/examples/gym/example_mrcr.yaml` | new self-contained SLURM +
vLLM config |
| `SKILL.md` | MRCR branch + gym index table |
| `references/quantization-benchmarks.md` | table row + comparability
notes |
`examples/gym_{gdpval,mrcr}/` → `examples/gym/example_<task>.yaml` and
`tasks/aa_gym/gdpval.md` → `tasks/gym/gdpval.md`, all tracked as git
renames.
`aa_gym` encoded "in the AA suite" in the *path*; since GDPVal is AA and
MRCR is
not, membership is now stated explicitly in both recipe headers and an
"In AA
suite?" column in the SKILL.md index. `gym/` groups by harness, not
suite.
### Bugs fixed (found by running it, not by reading it)
The first three are in the **shared** gym command block, so they also
affect
`example_gdpval.yaml` — GDPVal was broken independently of MRCR.
1. **Invalid OmegaConf interpolation.** A literal `${...}` inside a
comment in the
gym `command:` block is parsed as an interpolation and rejected, so
*every*
`nel run --dry-run` of either gym template died with
`hydra.errors.ConfigCompositionException`.
2. **Hardcoded `ray==2.49.2`** injected into each sub-server's
requirements. Against
an image carrying `ray[default]==2.55.1` this makes `uv` unsatisfiable
and the
gym resources server exits at startup. Now derived from the image at
runtime.
3. **`--max-num-seqs` sized against the wrong topology.** The template
shipped DP1
→ 4 replicas at 64 concurrent while citing a golden that is TP2×DP2 → 8
replicas
at 32. Following it ran double the reference's per-replica load, which
on
1M-token prompts is what decides whether the KV cache fits.
4. **Score extraction pointed at an empty map.** `results.yml` →
`groups.nemo_gym.metrics` holds only `key_metrics/mean/*` telemetry; the
scores
are in `artifacts/evaluator_rollouts_aggregate_metrics.json` →
`[0].agent_metrics`.
Also documents that `pass@1/accuracy` is already 0-100 while
`mean/reward` is the
same number as a 0-1 fraction, and records `mean/prefix_matched ≈ 0.55`
as the
healthy calibration.
5. **The template shipped a container its own bootstrap rejects.** After
adding the
hard-fail on an unpinned Gym, `container:` still defaulted to public
`nemo-gym:26.05` — the image the docs say has a non-git `/opt/Gym`. Now
`???`,
so `--dry-run`'s mandatory-value check catches it instead of the job
dying at
startup.
### Review feedback
All items from @meenchen and CodeRabbit addressed or answered inline.
Two worth
surfacing here:
- **`process_reasoning_traces` vs `use_reasoning`** — verified against
`nemo_evaluator/adapters/adapter_config.py`: both exist, `use_reasoning`
is the
**deprecated** one and they are bidirectionally aliased. Documented
rather than
switched to the deprecated name.
- **`tiktoken` / `transformers` left unpinned** — deliberate. The n3
prepare path
uses `transformers.AutoTokenizer` to decide which samples exceed the
cap, so a
bump can shift dataset membership; but pinning would diverge from the
golden's
`pre_cmd` and therefore from the run that produced the reference number.
Risk is
now documented under "Deferred, know the risk" instead of being
implicit.
Topology, `gres` and TP/DP were aligned to the **existing** sibling
templates
rather than to a new convention: `gres` stays a comment (already the
convention in
`example_eval.yaml` and `example_gdpval.yaml`), TP is a concrete `1`
like both
siblings, and `num_nodes`/`num_instances`/`--max-num-seqs` are guidance
rather than
baked values. `--max-num-seqs` sizing follows **AA-LCR**, since MRCR is
the same
KV-bound problem at ~1M tokens vs LCR's ~120K: the formula gives a
ceiling, not a
target, because oversubscribing causes preemption and recomputing a
1M-token
prefill makes the run slower.
### Usage
```bash
cp plugins/modelopt/skills/evaluation/recipes/examples/gym/example_mrcr.yaml mrcr.yaml
# fill checkpoint_path / served_model_name / container / SLURM ??? values
export NEMO_EVALUATOR_TRUST_PRE_CMD=1 NEMO_EVALUATOR_TRUST_UNLISTED_TASKS=1
nel run --config mrcr.yaml --dry-run && nel run --config mrcr.yaml
```
### Testing
Docs/config-only; no library code touched. `pre-commit` clean on every
commit.
- Both gym YAMLs parse; asserted the folded-scalar rule (no `#` inside
`>-`), that
every string survives `OmegaConf.create` (bug 1), and that the variant
is
identical in `data_prep_params` and `collect_rollout_params`.
- After the rename: zero stale `aa_gym`/`gym_gdpval`/`gym_mrcr` refs and
every
`recipes/**` path referenced across both skill trees resolves on disk —
this
caught `env.example`, which an extension-filtered grep missed. The
`nemo_gym_gdpval_stirrup_agent` metric names contain the substring
`gym_gdpval`
and were deliberately left untouched.
- **Ran the benchmark end-to-end.** A full run on gcp-nrt (B200)
completed
2363/2363 rollouts and scored; that run produced bugs 2 and 4 and the
corrected
metric paths. A second attempt on aws-cmh was preempted and later
cancelled.
`--dry-run` now passes (it did not before bug 1 was fixed), and the bare
template
correctly fails validation on unresolved mandatory values.
- Regenerated the config from scratch with a fresh agent against the
fixed
templates as a regression check; its dry-run passed and its findings
drove the
topology/`gres`/TP-DP alignment above.
Model-specific scores are deliberately **not** recorded in the recipe —
the
reference there stays the golden's BF16 Nano 3.5 shape, since quoting a
different
model's number beside it invites exactly the false comparison the recipe
warns
about.
### Before your PR is "*Ready for review*"
- Backward compatible?: ✅ — additive plus a doc-tree rename; all
referencing files
updated and verified. The template fixes change only broken behaviour.
- Copied code / new PIP dependency: N/A.
- New tests: N/A — agent-skill documentation.
- Changelog: N/A — skill/docs change; consistent with prior skill-only
commits.
- Claude approval: ❌ not yet — will trigger `/claude review`.
### Additional Information
Upstream drift flagged to reviewers (**no change made to that repo**):
in
`nvidia-eval-factory-benchmarking`, `configs/benchmarks/mrcr/bench.yaml`
is still
old-style `ng_*` while `configs/models/nemotron_nano_v35/gym.yaml` moved
to the
`gym eval` CLI on 2026-07-29 — RULER migrated, MRCR did not, so the
layers no
longer compose. This template follows `ng_*`, which is what MRCR's own
bench.yaml
specifies and what the sign-off run executed.
NVIDIA-internal companion (container specifics kept out of this public
tree):
Model-Optimizer-Internal MR !116.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
195a5af430 |
Add modelopt-agents-codeowners as owner of plugins/ (#2212)
### What does this PR do? Type of change: Documentation `plugins/` holds the installable ModelOpt agent plugin — the canonical skill tree at `plugins/modelopt/skills/`, which `.agents/skills` and `.claude/skills` expose via relative symlinks. It currently has no CODEOWNERS rule, so it falls through to the default `* @NVIDIA/modelopt-devs` owner. This routes it to `@NVIDIA/modelopt-agents-codeowners` instead. ``` # Agent plugin (skills, agent config) /plugins @NVIDIA/modelopt-agents-codeowners ``` The path is root-anchored (`/plugins`) to match the style of the `/examples` block above it. ### Usage N/A — no API or flag change. ### Testing No test surface. Verified the rule resolves as intended against the existing file: `/plugins` is the last (and only) match for paths under `plugins/`, so it wins over the default `*` rule; it does not overlap any other entry. Effective ownership of `plugins/` on this branch is `@NVIDIA/modelopt-agents-codeowners`. GitHub validates the team handle when the PR is opened — if the team does not exist or lacks write access to this repo, the Settings > CODEOWNERS errors page will flag the line. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- Repo-internal review routing, not user-facing. --> - Did you get Claude approval on this PR?: N/A ### Additional Information Note for reviewers: shared agent config and scripts under `.agents/` (everything that is not a symlink into `plugins/`) are still owned by the default `@NVIDIA/modelopt-devs` rule. Happy to add `/.agents` here too if the agents team should own that as well. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Added repository ownership rules for the plugins directory and its agent configuration files. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
b96841db3e |
Add optional MLflow tracking to the vLLM fake-quant server (#2120)
### What does this PR do? Type of change: new feature Wires `examples/vllm_serve/vllm_serve_fakequant.py` up to `modelopt.torch.utils.mlflow` via `--mlflow <tracking-uri>`, the same way #2023 did for `hf_ptq.py`, so a fake-quant serve records **what it actually quantized** and an evaluation of that endpoint can be traced back to a recipe. Without the flag, behavior is unchanged — every hook is gated on it. Three design points worth review: 1. **The run is recorded in the vLLM worker, not the launcher.** `vllm_serve_fakequant.py` is the API-server frontend; the engine and its workers are separate processes whose stdout it never sees, so a run opened there would capture none of the calibration. The launcher instead only settles the tracking configuration — validating the URI, naming the experiment, recording the command the user actually typed — and publishes it through the environment, which is how every other setting in this example (`QUANT_CFG`, `RECIPE_PATH`, …) already reaches the workers. Global rank 0 opens the run, so a TP-8 serve produces one run. 2. **The run covers load-through-warm-up, not the server's lifetime.** It opens *before the weights load*, so an unreachable server or a missing token fails in seconds rather than after a load and a full calibration, and it closes `FINISHED` once the model is quantized and warmed up. A run that stayed open for the serving lifetime would never close cleanly on SIGTERM. 3. **`recipe/quant_cfg.yaml` is only written on the preset path.** With `RECIPE_PATH`, `get_quant_config` returns the recipe's `quantize` section unchanged and `resolved_recipe.yaml` already carries it. With `QUANT_CFG`/`KV_QUANT_CFG` it is the *only* record of what ran: the params carry the preset names, while the config reaching `mtq.quantize` is those two deep-copied, merged, and — for an MLA model — extended at runtime with `*kv_c_bmm_quantizer` / `*k_pe_bmm_quantizer` by inspecting the loaded model. Uploaded artifacts: | Artifact | Contents | | --- | --- | | `command.txt` | The launcher's invocation, copy-pasteable, credentials masked | | `version.txt` | The ModelOpt version that ran | | `recipe/resolved_recipe.yaml` | `RECIPE_PATH` with its `$import`s expanded | | `recipe/quant_cfg.yaml` | Merged `QUANT_CFG`/`KV_QUANT_CFG` + MLA fixup (preset path only) | | `logs/<script>.log` | The rank-0 worker's stdout/stderr, including a crash traceback | | `summary/quant_summary.txt` | The per-quantizer summary | Plus the quantization *and* serving settings as searchable params, and `user` / `hostname` / `modelopt_version` / `git_sha` / `vllm_version` tags. The `checkpoint_path` tag matches the one `hf_ptq.py` sets, so a checkpoint's PTQ run and every serve of it join up. Two small library additions, both consumed by the new example module: - `command_text(argv=None)` — records another process's invocation, since a spawned worker's own `sys.argv` is vLLM plumbing rather than anything a user typed. - `MlflowRunLogger.log_text()` — uploads a value settled midway through a run, so a crash during calibration still keeps the config that caused it. The example `Dockerfile` installs the `mlflow` extra; the client remains optional and is imported only once tracking is enabled. ### Usage ```bash RECIPE_PATH=<recipe.yaml> python vllm_serve_fakequant.py <model_path> -tp 8 \ --host 0.0.0.0 --port 8000 \ --mlflow https://<your-mlflow-server>/ ``` ``` [mlflow] tracking to https://<your-mlflow-server>, experiment $USER/vllm_serve_fakequant/<model>-<recipe> (Worker_TP0) [mlflow] run: https://<your-mlflow-server>/#/experiments/19/runs/1c6679448f25... ``` `--mlflow-experiment` / `--mlflow-run-name` override the defaults. `$MLFLOW_TRACKING_URI` enables tracking on its own and is best-effort; an explicit `--mlflow` overrides it and fails loudly. > This is the **quantization** tracking server. It is unrelated to any server an evaluation harness exports its scores to — NeMo Evaluator Launcher has its own `export.mlflow.tracking_uri`. The README calls this out. ### Testing **Unit — 87 passing** (`tests/examples/vllm_serve/test_vllm_mlflow_utils.py`, 33 new; `tests/unit/torch/utils/test_mlflow.py`, +5). `vllm_mlflow_utils` deliberately imports no vLLM, so the whole launcher→worker handover is covered without a GPU, a server, or the mlflow client. **End to end on aws-cmh** (4× GB300, `simple_evals.gpqa_diamond`, Nemotron-3.5-Lightning-30B-A3B-BF16 fake-quantized with `general/ptq/nvfp4_mlp_only-kv_fp8_cast`): run `FINISHED` in 261.5 s, opened by `Worker_TP0` only, all artifacts present and verified by content — `command.txt` held the launcher's invocation rather than the worker's spawn argv, and `resolved_recipe.yaml` was 6797 B against 1845 B of source. 104 quantizers enabled (92 NVFP4 dynamic block-16 expert weight/input with calibrated amax, 12 FP8 KV bmm). The eval then ran to completion against the served endpoint, 22/22 requests HTTP 200. Two bugs the hardware run caught, both fixed here with regression tests: - `--mlflow_run_name` was rejected. vLLM's `FlexibleArgumentParser.parse_args` rewrites **every** `--foo_bar` to `--foo-bar` before matching, so a flag registered only under the underscored spelling is unreachable from its CLI. Both spellings are now registered. A unit test on a plain `ArgumentParser` could not have caught this. - `recipe/quant_cfg.yaml` uploaded a Python `repr` blob under a `.yaml` name: a recipe's `quantize` is a `QuantizeConfig`, `yaml.safe_dump` raises `RepresenterError` on it, and the old JSON fallback stringified the object. `_dump_yaml` now unwraps pydantic via `model_dump(mode="json")` and raises otherwise, with the caller downgrading that to a warning so a bad config cannot take down a serve. **Known coverage gap:** the preset (`QUANT_CFG`/`KV_QUANT_CFG`) path — the only one that now writes `recipe/quant_cfg.yaml` — is covered by unit test but has not been exercised on hardware; the canary used `RECIPE_PATH`. Likewise the case where `$MLFLOW_TRACKING_URI` is present *inside* the deployment container and `--mlflow` overrides it is unit-tested only: NeMo Evaluator Launcher forwards only declared env vars, so the eval server's URI never entered the container in the canary. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — new optional flags only; no `--mlflow` means no behavior change. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new dependency. Uses the existing optional `nvidia-modelopt[mlflow]` extra (`mlflow-skinny`, Apache-2.0) added in #2023; the example `Dockerfile` now installs it. No code copied from other sources. - Did you write any new necessary tests?: ✅ — 38 new tests. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — 0.47 Misc. - Did you get Claude approval on this PR?: ❌ — `/claude review` not yet run. ### Additional Information Follows #2023, which added `MlflowRunLogger` and the `hf_ptq.py` integration. Note for anyone tracking from an OCI cluster: `mlflow-modelopt.nvidia.com` is unreachable from oci-nrt and oci-hsg. TCP 443 completes and the connection is then reset on the first application byte, regardless of SNI or protocol, one RTT away — the PDX PaaS ingress appears to apply a source-IP policy, and the OCI clusters egress from Oracle-owned addresses (`155.248.190.0`, `168.110.199.1`) rather than NVIDIA's. gcp-nrt, aws-cmh and cw-dfw all reach it. This is an infrastructure matter, not a property of this change, but it determines where the feature is usable today. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added optional MLflow tracking for vLLM fake-quantization serving runs. * Records serving, quantization, worker, and invocation metadata, including configuration and summary artifacts. * Supports tracking URI, credentials, environment, and command-line configuration. * Added command and text artifact logging for active MLflow runs. * **Documentation** * Documented setup, configuration, recorded artifacts, lifecycle, and fallback behavior. * Updated the example container to include MLflow support. * **Tests** * Added comprehensive coverage for tracking configuration, logging, failures, and disabled tracking. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
06480249fd |
[skill] evaluation: align nel-next TB2.1/SWE-bench with golden toolchain (#2063)
### What does this PR do? Type of change: Documentation / tooling (agent skill) Terminal-Bench 2.1 configs generated from the `evaluation` skill had drifted from the canonical eval-factory config (`configs/benchmarks/terminal-bench-2.1/bench.yaml`). The **scoring contract already matched** the reference configs exactly — playbook, `repeats: 8`, `timeout_strategy: max`, `run_timeout: 7200`, `llm_kwargs.timeout: 3600`, concurrency. What had drifted was the toolchain and a few proxy-level defaults. The drift was in the **skill**, not in individual configs: a config generated fresh from the skill reproduced every stale value, so patching configs alone would not have held. - **`nel-next.sh` installs from the public upstream repo** (`github.com/NVIDIA-NeMo/Evaluator`, default branch → `0.4.0`) instead of PyPI. PyPI `nemo-evaluator` tops out at `0.3.0` and cannot reach the 0.4.x toolchain the reference runs use. `NEL_NEXT_SPEC` becomes the PyPI escape hatch and now takes precedence when explicitly set; `NEL_NEXT_ORIGIN` stays overridable from `.env` so internal mirrors stay out of this repo. - **`eval_image`**: document the pinned `0.5.0.1-harbor` (single source of truth: `configs/shared/nel_next_containers.yaml`) rather than `0.3.1.1-harbor` as a floor. - **`proxy.request_timeout` 1800 → 3600** — must be `>=` the solver's `llm_kwargs.timeout`, else the proxy truncates long agent turns the harness is still awaiting. - **`drop_params`**: add `max_input_tokens_per_task`, `no_rebuild` — sent by the 0.5.x harbor eval image; vLLM returns 400 unless stripped. - **`exclude_patterns`**: add `model_traffic.jsonl` so captured request bodies stay in the run dir and never reach MLflow. - **`http_pairs_dump`** interceptor (last in chain) for HTTP diagnostics. - **Sharding documented**: `max_concurrent`/`sandbox.concurrency` are *per shard*, so `shards: N` multiplies both serving capacity and live sandboxes (`N x concurrency`). - **`.gitignore`**: broaden `.env` / `.env-*` to `.env*` so secret backups such as `.env.bak-tb21` cannot be staged. **This does not move the benchmark.** The TB2.1 task set is pinned by a vendored registry override that has not changed since 2026-06-03, and both `0.3.1.1-harbor` and `0.5.0.1-harbor` score 89 samples — so the image bump is a toolchain fix and scores stay comparable across it. ### Usage ```bash set -a && source .env && set +a .agents/scripts/nel-next.sh --version # 0.4.0 (public upstream build) .agents/scripts/nel-next.sh eval run <tb21-config>.yaml --dry-run ``` ### Testing **End-to-end parity runs.** Configs generated from these changes were run to completion on aws-cmh (4x GB300 aarch64, sm_103) with an NVFP4 checkpoint of Qwen3.6-35B-A3B, and compared against the reference BF16 results for the same base model: | Benchmark | Reference (BF16) | This run (NVFP4) | Delta | |---|---|---|---| | Terminal-Bench 2.1 | 0.4438 `[0.4215, 0.4661]` | **0.4438** `[0.4233, 0.4644]` | **0.0000** | | SWE-bench Verified | 0.7012 `[0.6920, 0.7104]` | **0.7040** `[0.6950, 0.7130]` | **+0.0028 (+0.40%)** | Both `pass@1` over the full task sets (89 x r8 = 712 trials; 500 x r5 = 2500 trials), with overlapping 95% CIs in both cases. This exercises the changes in this PR directly: the `0.5.0.1-harbor` eval image, the `proxy.request_timeout >= llm_kwargs.timeout` fix, the four-param `drop_params`, the per-benchmark interceptor ordering, and the SWE-bench `system_message` / `instruction_template` requirements. A wrong tool-call parser, sampling preset, or instruction template would each have moved these numbers well outside the intervals. Also verified: - `nel-next.sh --version` -> `0.4.0`, built from `Evaluator.git@4d081325` — the commit the scored runs above actually executed, confirmed from their recorded uv archive (`direct_url.json` -> `commit_id 4d081325170aababd0c8f27c58bed31a81ce82ac`). This is the SHA `NEL_NEXT_REF` now defaults to. - Both configs pass `eval run --dry-run` with no schema errors. These schemas are `extra="forbid"`, so `http_pairs_dump` and the new `drop_params` entries would hard-fail if unsupported by the pinned toolchain. - `0.5.0.1-harbor` resolves in the generated `nel_eval.sbatch`; the tag is multi-arch (linux/amd64 + linux/arm64) and imported successfully on aarch64 compute nodes. - `pre-commit run --files <changed>` - all hooks pass, no file modifications. - Values cross-checked against the canonical benchmark configs in `dl/JoC/competitive_evaluation/nvidia-eval-factory-benchmarking` @ `main`. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `NEL_NEXT_SPEC` restores the previous PyPI install. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — agent-skill docs/config; validated via `--dry-run` + pre-commit. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — no library API change. - Did you get Claude approval on this PR?: ❌ — pending. ### Additional Information Personal run configs under `.agents/skills/evaluation/runs/` are deliberately **not** included: they carry internal cluster hostnames, lustre paths, account names and an AWS account id, which do not belong in this public repo. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated evaluation guidance for `nemo-evaluator` 0.4.x and the `0.5.0.1-harbor` image. * Added canonical benchmark configurations, longer proxy timeouts, request filtering, diagnostics, MLflow exclusions, sharding, capacity planning, concurrency, replay behavior, and deployment verification guidance. * Documented separate virtual-environment setup and version tracking. * **Configuration** * Evaluation tooling now defaults to a pinned Git-based installation with configurable sources and version references. * Environment files with any `.env`-prefixed name are now ignored. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9220fac053 |
[NVBug: 6563509] Drop Phi-3-vision / Phi-4-multimodal PTQ support (#2115)
### What does this PR do?
Type of change: Deprecation
Resolves [NVBug 6563509](https://nvbugspro.nvidia.com/bug/6563509),
where
`hf_ptq.py` on Phi-4-multimodal-instruct died with
`RuntimeError: Tensor.item() cannot be called on meta tensors`.
The crash is real but not fixable on our side, and it is not the reason
the model
is unusable. Phi-4-multimodal's bundled remote code predates
Transformers v5 and
does not load on **any** version in our supported range
(`transformers>=4.57,<5.15`):
| Blocker | Where |
|---|---|
| `peft.get_peft_model` reads `prepare_inputs_for_generation`, gone
since transformers 4.52 dropped `GenerationMixin` from `PreTrainedModel`
| `modeling_phi4mm.py:1959` |
| `_tied_weights_keys` declared as a list; Transformers 5.x calls
`.keys()` on it in `post_init` | `modeling_phi4mm.py:1937` |
| `int(torch.tensor(...))` in `__init__`, which cannot run on a meta
device — the reported crash | `speech_conformer_encoder.py:1435` |
The model card pins `transformers==4.48.2` / `peft==0.13.2`, so there is
no
overlap with our floor and nothing on our side can bridge it. The model
is
therefore dropped rather than worked around.
**Phi-3-vision is dropped alongside it because it is the older,
superseded model
in the same family** — with its successor unsupportable there is no
reason to
keep carrying the predecessor. This is a product-scope call, not a
separate
compatibility finding: Phi-3-vision shares the list-valued
`_tied_weights_keys`
defect (`modeling_phi3_v.py:1214`) and so is likewise broken on
Transformers 5.x,
but it does **not** hit the `peft` blocker, and it was not re-verified
on 4.57.
Per the 0.46 changelog we have already bumped the floor to 4.57 and
noted that
"Transformers 4.x support will be dropped in a future release", so any
remaining
window closes on its own. Same reasoning already applied to VILA / NVILA
in this
release.
**Removed**
- the support-matrix row in `examples/hf_ptq/README.md`
- `"Phi4MMForCausalLM": "phi4mm"` from `MODEL_NAME_TO_TYPE`
- the multimodal-detection heuristics that only ever matched these two —
`vision_lora`, `audio_processor`, `embd_layer.image_embd_layer`, and the
`phi4mm` model-type check — in both `is_multimodal_model` and
`_is_multimodal_config`
- the `Phi3Image` / `PhiImage` exclusions in `is_embedding`
- the phi4mm input-mode warning in `hf_ptq.py`
- `modelopt_recipes/huggingface/phi4mm/` and its references in
`modelopt_recipes/ptq.md`
**Not changed:** the device-map sizing path (meta-device skeleton,
`infer_auto_device_map`, and the `--gpu_max_mem_percentage` cap) keeps
its
original behavior. That cap is wanted exactly where it already fires —
when the
model is already offloading to CPU, where it costs little and the
headroom is
required. With the affected checkpoints removed, there is no supported
model
that trips the meta-device build, so there is nothing to work around
here.
Text-only **Phi-3/Phi-4** and **Phi-3.5-MoE** are natively supported by
transformers and are untouched.
### Testing
On H200, `nvcr.io/nvidia/tensorrt-llm/release` (torch 2.12, transformers
5.5.4),
against the real checkpoint:
- **Version matrix** (vanilla transformers, no modelopt) — Phi-4-MM
loads at
4.48.2 / 4.49.0 / 4.50.0 / 4.51.3 and fails at 4.53.3 / 4.56.2 / 4.57.1
(`AttributeError: 'Phi4MMModel' object has no attribute
'prepare_inputs_for_generation'`) and at 5.5.4 (meta-init, then
tied-keys).
This is what establishes that no supported version works.
- `tests/examples/hf_ptq/test_example_utils.py` — 28 passed.
- **Sweep**: `tests/examples/hf_ptq` + `tests/unit/torch/export` —
failure set
identical to the pre-change tree (GPU/model-dependent `test_vlm_ptq`,
plus
`test_quant_aware_conversion` scoped-mapping tests), so none are
introduced
here.
- `pre-commit` clean on all changed files, including recipe validation.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ — PTQ for Phi-3-vision and
Phi-4-multimodal is removed, along with the `huggingface/phi4mm/ptq/*`
recipes. Phi-4-multimodal is already unloadable on every supported
transformers
version, so no working workflow regresses; Phi-3-vision is a deliberate
scope
removal as its superseded predecessor.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — this is a deletion; the
existing
`test_get_model_*` / `test_resolve_init_config_*` tests are unchanged
and still
pass.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
Two related references were left in place deliberately; say the word and
I'll
fold them in:
- `tests/examples/hf_ptq/test_deploy.py` still deploys the
already-published
`nvidia/Phi-4-multimodal-instruct-{NVFP4,FP8}` checkpoints. Those
artifacts
exist and serve fine; this PR only removes the ability to *produce*
them.
- `examples/torch_onnx/README.md` still lists Phi-4-multimodal-instruct.
That
is a separate ONNX pipeline that does not go through `get_model()` and
was not
tested here.
Earlier revisions of this branch also reworked the device-map sizing so
the
meta-tensor crash could not occur. That was reverted in
|
||
|
|
9b8caf623a |
Use lm-eval 0.4.12's built-in trtllm backend, deprecate lm_eval_tensorrt_llm.py (#2066)
### What does this PR do?
Type of change: documentation / example update (with a behaviour fix)
lm-evaluation-harness **0.4.12** is the first release that ships a
TensorRT-LLM backend
(`lm_eval.models.trtllm_causallms`, registered as `trtllm`) — it is
absent in 0.4.10 and
0.4.11. This example no longer maintains its own, so:
- Pin `lm_eval[api,ifeval]>=0.4.12,<0.5` (the 0.5.0.dev line drops the
file) and bump
`lm_eval_hf.py`'s version guard to match.
- **Delete** `examples/llm_eval/lm_eval_tensorrt_llm.py` (the `trt-llm`
model). Replace
`python lm_eval_tensorrt_llm.py --model trt-llm --model_args
tokenizer=<tok>,checkpoint_dir=<ckpt>`
with `python lm_eval_trtllm.py --model trtllm --model_args
model=<ckpt>,tokenizer=<tok>`.
- Add `examples/llm_eval/lm_eval_trtllm.py`, whose entire content is one
corrected
`_parse_logprobs` plus `cli_evaluate()` (see below). `lm_eval_hf.py`
stays HF-only.
- `examples/hf_ptq/scripts/huggingface_example.sh` and the docs use the
upstream backend.
`parser.sh` gains `--input` (`BUILD_MAX_INPUT_LEN`, default 4096) — it
already *echoed*
that variable but never parsed or defaulted it, so it printed empty on
every run.
#### Why `lm_eval_trtllm.py` exists: an upstream off-by-one
TensorRT-LLM aligns `prompt_logprobs` to the *next* token.
`executor/base_worker.py`:
```python
# Pass prompt_token_ids with an offset of 1 for correct mapping to the context logits
prompt_token_ids = generation_result._generation_request.prompt_token_ids[1:] + first_generation_token
```
So entry `i` is the distribution that predicted `tokens[i + 1]`, and
`_topk_logprobs`
appends that token's id when it is not in the top-k. lm-eval's
`_parse_logprobs` instead
reads `prompt_logprobs[i][tokens[i]]` and applies its own shift on top,
which raises
`KeyError` on the **first request of every loglikelihood task**
(hellaswag, mmlu, arc, ...):
```
File ".../lm_eval/models/trtllm_causallms.py", line 324, in _parse_logprobs
current_token_logprob = prompt_logprob[tokens[i]]
KeyError: 6503
```
Probed against TRT-LLM 1.3.0rc23 with a 14-token prompt for
`prompt_logprobs` 0, 1 and 2:
`tokens[i]` is missing at **every** position, `tokens[i+1]` is present
at every position.
Only `generate_until` tasks work unpatched. **This wants an upstream
issue against
EleutherAI/lm-evaluation-harness.**
The override also fails loudly rather than quietly: it checks
`prompt_logprobs` covers
every prompt token and raises on a missing token, instead of skipping
the term and
silently inflating the reported accuracy.
#### Defaults that must be set explicitly
`TRTLLM.__init__` accepts `**kwargs` but forwards only a fixed set to
the **TensorRT-LLM
`LLM` API**, so extra `--model_args` aimed at the engine are silently
dropped. (lm-eval's
own named parameters — `max_gen_toks`, `batch_size`, `truncation_side`,
... — are honored
normally.) Two engine defaults are unsafe for few-shot eval:
- `tensor_parallel_size` defaults to **1** (the deleted wrapper used
every visible GPU).
- `max_input_len` defaults to **2048**, and longer prompts are silently
left-truncated —
5-shot MMLU/gsm8k prompts exceed that.
### Usage
```bash
python lm_eval_trtllm.py --model trtllm \
--model_args model=<quantized checkpoint dir>,tokenizer=<HF model folder>,tensor_parallel_size=<tp>,max_batch_size=<bs>,max_input_len=4096,max_output_len=512 \
--tasks hellaswag,gsm8k \
--batch_size <bs>
```
Flat arguments (no `run` subcommand) are what 0.4.12's
`HarnessCLI.parse_args` inserts
`run` for automatically (`_cli/harness.py:48-51`); this is the exact
command form used for
the results below.
### Testing
**Unit** — `tests/examples/llm_eval/test_lm_eval_trtllm.py`, no GPU and
no `tensorrt_llm`
install: stubs the response object and pins the `i-1` alignment, the
`rank != 1` →
`is_greedy` rule, the `ctxlen=0` edge, and both `RuntimeError` paths.
Mutation-checked —
dropping the `-1` shift is caught by 5/5 cases, ignoring `ctxlen` by
4/5. A sixth test is a
**tripwire**: it asserts lm-eval's own implementation is still
misaligned, so a future
0.4.x that fixes the bug fails the test and says to delete this file
rather than being
silently re-broken by the override.
**End to end** — `nvidia/Qwen3.5-122B-A10B-NVFP4` (NVFP4 MoE, 256
experts) on **4x B300**,
TRT-LLM 1.3.0rc23, lm-eval 0.4.12, `--limit 32`:
| run | hellaswag acc | hellaswag acc_norm | gsm8k flexible | gsm8k
strict |
|---|---|---|---|---|
| deleted impl (`trt-llm`), tp=4 | 0.7188 | 0.7812 | 0.8438 | 0.7812 |
| `lm_eval_trtllm.py`, tp=1 | 0.7188 | 0.7812 | 0.8438 | 0.8125 |
| `lm_eval_trtllm.py`, tp=2 | 0.7188 | 0.7812 | 0.9062 | 0.8125 |
| `lm_eval_trtllm.py`, tp=4 | 0.7188 | 0.7812 | 0.8750 | 0.8438 |
- hellaswag (the loglikelihood path this PR fixes) is **identical at
every tp and identical
to the deleted implementation** — the alignment fix is exact, not
approximate.
- gsm8k varies by 1–2 samples out of 32 (generation path: upstream uses
native `stop=`
sequences and per-request `SamplingParams`; the old wrapper used
beam-search-of-1 with
post-hoc string truncation).
- Without the override, every hellaswag run above dies with the
`KeyError`.
- Re-verified at tp=4 after the code moved out of `lm_eval_hf.py` into
`lm_eval_trtllm.py`.
Note: NVFP4 fused-MoE has no CUTLASS tactic on Hopper (`No supported MoE
GEMM tactic
remains after replacing unsupported NO_SMEM epilogues.`), so this had to
be validated on
Blackwell.
### Feature parity notes
Gained from upstream: `loglikelihood_rolling` (was
`NotImplementedError`), pipeline
parallelism, `add_bos_token` auto-detection, prompt truncation,
per-request sampling params,
`prompt_logprobs` instead of full-vocab context logits (much lower
memory), thinking-tag
handling, `batch_size=auto`.
Not reachable through the upstream backend (were set by
`modelopt.deploy.llm.LLM`):
`enable_attention_dp` for MoE, `CudaGraphConfig`,
`enable_chunked_prefill`,
`moe_expert_parallel_size=1`, and `free_gpu_memory_fraction=0.7` with a
capped
`kv_cache.max_tokens` — upstream uses the TRT-LLM default 0.9 (observed
allocating 218 GiB
of paged KV cache on B300), so OOM risk is higher on smaller GPUs. This
is documented in
`examples/llm_eval/README.md`, and `huggingface_example.sh` honours a
preset `LM_EVAL_TP`
so users can lower the tensor-parallel size without editing the script.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ — `lm_eval_tensorrt_llm.py` is
removed and the CLI changes (`--model trt-llm` → `trtllm`,
`checkpoint_dir=` → `model=`). Migration command is in the README and
CHANGELOG.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependency; existing `lm_eval` pin tightened.
- Did you write any new necessary tests?: ✅ —
`tests/examples/llm_eval/test_lm_eval_trtllm.py` (6 cases, no GPU).
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.47 *Deprecations*.
- Did you get Claude approval on this PR?: ✅ — reviewed, feedback
addressed in `dcedd37b4` and `622b97c26`.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added TensorRT-LLM evaluation through lm-evaluation-harness’s `trtllm`
backend.
* Added configurable input/output lengths, batching, tensor parallelism,
and build input length.
* Improved prompt log-probability alignment for more accurate evaluation
results.
* **Documentation**
* Updated evaluation instructions, truncation guidance, backend
limitations, and configuration examples.
* **Deprecations**
* Removed the legacy TensorRT-LLM evaluation script and entry point.
* **Updates**
* lm-evaluation-harness now requires versions 0.4.12 through 0.4.x.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
77dbeb1872 |
Add optional MLflow tracking to hf_ptq.py (#2023)
### What does this PR do? Type of change: new feature Adds `modelopt.torch.utils.mlflow.MlflowRunLogger`, a reusable helper for recording a script run on an MLflow tracking server, and wires `examples/hf_ptq/hf_ptq.py` up to it via `--mlflow <tracking-uri>` so a PTQ run can be reproduced from its MLflow entry alone. Without the flag, behavior is unchanged — every hook is gated on it. The logger lives in the library rather than the example so other scripts can record runs the same way: it takes a tracking URI, an experiment name and an explicit `enabled` flag, with params, tags and artifacts passed in. `hf_ptq.py` supplies only the PTQ-specific pieces (its params, the resolved recipe, the quantization summaries). `mlflow` is an optional dependency, imported only once tracking is enabled, so it is not a new requirement for the library. The run is opened **before the model loads**, so a bad URI or an unreachable server fails in seconds rather than after hours of calibration. The invocation and the recipe are uploaded at that point too, which keeps a crashed run useful: it is still recorded, with status `FAILED` and its log attached. Uploaded artifacts: | Artifact | Contents | | --- | --- | | `command.txt` | The full invocation, copy-pasteable | | `version.txt` | The ModelOpt version that ran (also a searchable tag) | | `recipe/resolved_recipe.yaml` | The `--recipe` with `$import`s expanded | | `logs/hf_ptq.log` | Everything the run printed, including a crash traceback | | `summary/quant_summary.txt` | Per-quantizer summary (unless `--no-verbose`) | | `summary/moe.html` | Per-expert calibration token counts, when the run produces them | Plus model / format / calibration settings as searchable params, and `user` / `hostname` / `modelopt_version` / `git_sha` tags. Three design points worth review: 1. **The recipe is uploaded resolved, not verbatim.** A recipe may be a directory or use `$import`s, so the source file is not self-contained. For `huggingface/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_at_6p0bits-active_moe` the source is 2,230 B / 58 lines against 7,563 B / 308 lines resolved — the raw file records under 30% of what actually ran. 2. **`hf_ptq.py` has no logging framework** (bare `print()`), so the log is produced by teeing stdout/stderr. Handlers that libraries bound to `sys.stderr` at import time are re-pointed at the tee for the run's duration and handed back afterwards; without that, `transformers` / `huggingface_hub` warnings reach the console but never the log. Native (C-level) output is still not captured — documented in the README. 3. **The recipe upload lives in the caller, not the library.** That keeps `modelopt.recipe` out of `modelopt.torch.utils`, which would otherwise risk a `modelopt.torch.utils` → `modelopt.recipe` → `modelopt.torch.quantization` → `modelopt.torch.utils` import cycle. 4. **MLflow failures never fail the quantization.** Startup validation is fatal by design (it is before any GPU work); the end-of-run upload is best-effort. Only the main rank uploads, so `--use_fsdp2` runs produce a single run. ### Usage ```bash python hf_ptq.py \ --pyt_ckpt_path <huggingface_model_card> \ --recipe general/ptq/nvfp4_default-kv_fp8_cast \ --export_path <quantized_ckpt_path> \ --mlflow https://<your-mlflow-server>/ ``` ``` [mlflow] experiment: $USER/hf_ptq/<checkpoint basename>-<recipe name> [mlflow] run: https://<your-mlflow-server>/#/experiments/13/runs/c243352e... ``` `--mlflow_experiment` and `--mlflow_run_name` override the defaults (`$USER/hf_ptq/<basename>-<recipe name or --qformat>`, and the UTC start time). Passing `--mlflow` with no value uses `$MLFLOW_TRACKING_URI`. Authentication uses MLflow's own env vars. ### Testing **Unit** — 51 tests in `tests/unit/torch/utils/test_mlflow.py` for the library, plus 13 in `tests/examples/hf_ptq/test_hf_ptq_args.py` for the hf_ptq wiring. CPU-only, no network and no `mlflow` dependency (driven against a stub module). Covers experiment-name derivation and sanitization, URI accept/reject, tee pass-through, the pre-bound-handler redirect, artifact renaming, skipping absent optional outputs, the disabled path, and `version.txt`. 85 tests pass together with the existing `test_hf_ptq_args.py` / `test_example_utils.py`. **Hardware** — real PTQ runs against a live MLflow server: | Run | Result | | --- | --- | | Qwen3-0.6B, NVFP4 PTQ, 1×B200 | `FINISHED`, all artifacts, sane post-quant generations | | Qwen3.6-35B-A3B MoE, AutoQuantize `w4a16_nvfp4_fp8_at_6p0bits-active_moe`, 2×B200 | `FINISHED` in 63 min, search hit `effective bits: 6.00`; 106 KB log capturing every per-layer decision, 4.4 MB quant summary | | Qwen3.6-35B-A3B, plain NVFP4 PTQ, 2×B200 | `FINISHED` | | Qwen3-0.6B re-run after the library move, 1×H200 | `FINISHED`, all five artifacts including `version.txt` | | Run **without** `--mlflow` after the review fixes | exactly 1 `[load_recipe]` line and 0 `[mlflow]` lines, confirming the untracked path is untouched | | Two runs sharing one `--export_path`, second crashed early | second run uploads **no** summary — the first run's 124 KB file on disk is correctly not attributed to it, and its traceback is in the log | | Crash mid-run (gated HF dataset) | `FAILED` recorded with log + traceback attached, summaries correctly absent | | Malformed URI | Rejected by `argparse` with a `Did you mean https://…?` hint | | Unreachable host | Fails in 9.9 s total, before any model load | | No `--mlflow` | Exit 0, no MLflow output, unchanged export | **Coverage gap, stated plainly:** `summary/moe.html` is verified only against a synthetic file (unit test + a real upload). It could not be produced naturally — `expert_token_count` buffers live on `_QuantSparseSequentialMoe`, while Qwen3.5/3.6 experts take the fused `_QuantFusedExperts` path, so no such file is written for these models regardless of `--moe_calib_experts_ratio`. The uploader's conditional is correct; the branch simply had no natural input available here. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — new optional flags only; no `--mlflow` means no behavior change. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — adds `mlflow` as an optional extra in `pyproject.toml` (`nvidia-modelopt[mlflow]`, folded into `all`) and to `examples/hf_ptq/requirements.txt`. Apache-2.0 (permissive). Imported lazily, so it is not required to install or import ModelOpt. No code copied from other sources. - Did you write any new necessary tests?: ✅ — 29 new unit tests. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — 0.47 New Features. - Did you get Claude approval on this PR?: ❌ — `/claude review` not yet run. A self-review was done first and its six findings are fixed in the third commit (the notable one: gathering the MLflow inputs re-read the recipe on *every* run, including without `--mlflow`). ### Additional Information The one deliberate coverage gap is `summary/moe.html`, described under Testing: no model available here takes the sparse-sequential MoE path that writes it, so it is covered by unit test and a synthetic upload rather than a natural one. The uploader treats it as an optional output and skips it when absent, which is exercised by test. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5e019d882f |
Add PR review-feedback guidance to AGENTS.md (#2057)
### What does this PR do? Type of change: documentation Adds a `## Responding to PR review feedback` section to `AGENTS.md` (which `CLAUDE.md` symlinks to, so it covers both Claude Code and Codex). Today the agent instructions stop at "open the PR" — nothing says what to do with the review comments that come back, so agents either apply every suggestion reflexively (including stale or wrong bot findings) or fix things silently and leave reviewers guessing whether their comment landed. The new section says to: - **Triage before acting** — check each comment against the current code; CODEOWNERS reviewers outweigh bot reviewers (CodeRabbit, Claude), whose findings are claims to verify rather than instructions. A reviewer reaffirming after pushback settles it. - **Pick one outcome per thread** — address in a commit, push back citing the code that shows the comment is wrong, or postpone as out of scope, and report which threads got which when asking for push approval. - **Reply in the thread after pushing** — a sentence on what changed and where. Those replies ride on the approval to push; pushback and postpone replies need their own approval, since no commit backs them. The agent never resolves threads — that stays the reviewer's call. ### Usage N/A — agent instructions only, no code change. ### Testing `pre-commit run --files AGENTS.md` passes (markdownlint included). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added guidance for responding to pull request review feedback. * Clarified how to validate comments, prioritize reviewers, report outcomes, and respond to addressed threads. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
14b20c0a12 |
[skill] evaluation: add GDPVal (NeMo Gym Stirrup agent) support (#2039)
### What does this PR do? Type of change: new feature (agent skill) Adds GDPVal support to the `evaluation` agent skill. GDPVal is an agentic AA benchmark: the NeMo Gym "Stirrup" agent produces office/PDF deliverables inside a per-task Apptainer code-exec sandbox, and a judge panel scores them. It runs on the 0.2.6 launcher as a `nemo_gym` task, but it is **standalone** (one gym eval per config) and mechanically unlike the `aa/` nemo-skills tasks, so it gets its own branch in the skill rather than being merged into the `aa/` task list. - `recipes/tasks/aa_gym/gdpval.md` — task recipe: standalone rule, rubric-vs-comparison scoring, canary, score extraction. - `references/gym-gdpval.md` — the machinery: Apptainer SIF sandbox, the `_gym_prepare` venv-repair / process-group-reap workaround, deployment sizing, scoring modes, the MLflow deliverables trap, canary failure modes, and the SIF ↔ Gym-version rebuild coupling. - `recipes/examples/gym_gdpval/` — self-contained SLURM + vLLM template plus the co-located `_gym_prepare.yaml` Hydra include (it must travel with the config). - `scripts/gdpval-sif.sh` — build-if-absent / reuse-if-present Apptainer SIF helper. Builds on the target cluster only (never copies across clusters), flock-guarded and atomic, driven by `$GDPVAL_SIF_DIR`. - `SKILL.md` / `references/quantization-benchmarks.md` — GDPVal is part of the AA suite but a different harness, so it is generated as a companion standalone config. - `recipes/env.example` — `TAVILY_API_KEY` (agent web search) and `GDPVAL_SIF_DIR`. ### Usage ```bash # 1. Set GDPVAL_SIF_DIR in .env, then build the sandbox once on the target cluster # (build-if-absent, reuse-if-present): srun -p cpu -t 01:00:00 --pty .agents/scripts/gdpval-sif.sh # 2. Copy the whole example dir (the _gym_prepare.yaml include must travel with it), # fill in the ??? values, then dry-run -> canary -> full: nel run --config gym_gdpval/example_gym_gdpval.yaml --dry-run ``` ### Testing Validated end-to-end on an aarch64 GB300 SLURM cluster with an NVFP4 MoE checkpoint: deploy → SIF build + sandboxed exec → gym head server → 220 rollouts + deliverables → judge scoring, producing a real rubric score with zero judge failures. Several traps found during that run are now documented in the reference (silent unsandboxed fallback, gym-commit/head-server hang, judge api-key value-vs-name, SIF ↔ Gym version coupling, `limit_samples` not limiting gym rollouts). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (agent skill documentation + helper script; no library code) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (agent skill only, no API change) - Did you get Claude approval on this PR?: ❌ (not yet run) ### Additional Information Docs/skill-only change under `.agents/`; no `modelopt/` source is touched. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added standalone GDPVal evaluation support through NeMo Gym, with Slurm, vLLM, sandbox, judge, reasoning, sampling, and logging configuration. - Added automated Apptainer/Singularity image setup with reuse, validation, locking, and reliable publishing. - Added configurable settings for judge services, web search, and shared image caching. - Added GDPVal preparation and execution examples, including dry-run, canary, and full-run guidance. - **Documentation** - Added GDPVal setup, troubleshooting, scoring, deployment, and quantization guidance. - Clarified separate GDPVal configuration, mandatory thinking mode, and repeat-count requirements. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7ed91540fa |
[NVBug: 6538278] Pass trust_remote_code to the TRT-LLM model load in deploy/eval examples (#2056)
### What does this PR do? Type of change: Bug fix Fixes [nvbug 6538278](https://nvbugspro.nvidia.com/bug/6538278). `examples/hf_ptq/run_tensorrt_llm.py` loaded the **tokenizer** with `args.trust_remote_code` but constructed `LLM()` without it, so it defaulted to `False`. Deploying any checkpoint that ships custom modeling code (`auto_map`) — e.g. `Llama-3.3-Nemotron-Super-49B-v1` (DeciLM) — failed at executor init: ``` Failed to initialize executor on rank 0: The repository ... contains custom code which must be executed to correctly load the model ... Please pass the argument trust_remote_code=True. -> ValueError -> Executor worker returned error -> run_tensorrt_llm.py subprocess exit 1 ``` The `modelopt.deploy.llm.LLM` wrapper already accepts and forwards `trust_remote_code` (`modelopt/deploy/llm/generate.py:153`), and `scripts/huggingface_example.sh` already forwards `--trust_remote_code` to the script — only the call site dropped it. The same omission exists in the sibling TRT-LLM deploy paths driven by the same launcher and the same checkpoints, so they are fixed together: | File | Fix | | --- | --- | | `examples/hf_ptq/run_tensorrt_llm.py` | the reported bug | | `examples/llm_eval/lm_eval_tensorrt_llm.py` | `LLM()` ignored the `trust_remote_code` that lm-eval injects into `model_args` | | `examples/llm_eval/mmlu.py` | `LLM()` ignored it (the tokenizer already used it) | | `examples/hf_ptq/scripts/huggingface_example.sh` | the `mmlu` stage never forwarded the flag at all | All four propagate the **user-provided** flag; none hardcode `trust_remote_code=True` (per SECURITY.md). ### Usage No API change. Existing flag now takes effect on the model load: ```bash scripts/huggingface_example.sh --model <Llama-3.3-Nemotron-Super-49B-v1> \ --quant fp8 --tasks quant --trust_remote_code ``` ### Testing - New CPU regression test `tests/examples/hf_ptq/test_run_tensorrt_llm.py` stubs the `tensorrt_llm`-dependent import and asserts both the tokenizer **and** the model load receive the flag, parametrized over `True`/`False`. Confirmed it fails without the fix (`KeyError: 'trust_remote_code'`) and passes with it. - Verified against the real `fire` package that a bare `--trust_remote_code` maps to `True` in `mmlu.py`'s `**kwargs`, and is absent (defaulting to `False`) when not passed. - Simulated the launcher's argument assembly both ways: with `TRUST_REMOTE_CODE=false` the emitted commands are byte-identical to before this change. - `bash -n` on the launcher; full pre-commit clean. - GPU deploy of the `fp8` DeciLM checkpoint was verified by the bug reporter on GB10 with this one-line change (model loaded, TRT engine built, generation succeeded, deploy exit 0). The `mmlu` / `lm_eval` stages are not GPU-verified here. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- examples-only bug fix --> - Did you get Claude approval on this PR?: ❌ <!-- not yet run --> ### Additional Information nvbug 6538278 / OMNIML-5659. Reported against modelopt 0.46.0rc0 on GB10 / DGX Spark (`tensorrt-llm/release:1.3.0rc22`). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Added consistent handling of the `trust_remote_code` setting across TensorRT-LLM inference and MMLU evaluation workflows. * Ensured tokenizer and model loading receive the configured remote-code behavior. * **Tests** * Added coverage validating remote-code settings and preserving KV-cache behavior when context logits are requested. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b75227ce1b |
[NVBug: 5987078] Fix unified HF export of compressed NVFP4 weights (--low_memory_mode) (#2038)
### What does this PR do? Type of change: Bug fix Fixes unified HF export of already-compressed NVFP4 weights — `mtq.compress` and, through it, `examples/llm_ptq/hf_ptq.py --low_memory_mode` (NVBug 5987078). Two defects, one root cause: the NVFP4 export branch has no handling for weights that were already real-quantized, unlike the `FP8_PB_REAL` branch which consumes `weight_quantizer._scale`. **1. `weight_scale` was recomputed from packed data.** After compression the weight is a `QTensorWrapper` of packed NVFP4 nibbles, and `QTensorWrapper.__new__` builds the Parameter from `_quantized_data`, so `.shape` reports the *packed* shape (the logical shape survives only in `metadata["shape"]`). The export derived the block count from `weight.shape[-1]` and took amax over nibble-pair bytes, so it wrote a scale of half the required size with meaningless values: | | `weight` | `weight_scale` written | expected | | --- | --- | --- | --- | | TinyLlama-1.1B `q_proj` | `[2048, 1024]` U8 | `[2048, 64]` | `[2048, 128]` | | DeepSeek-R1-Distill-Llama-70B `q_proj` | `[8192, 4096]` U8 | `[8192, 256]` | `[8192, 512]` | **2. An internal quantizer buffer leaked into the checkpoint.** `postprocess_state_dict` strips `weight_quantizer.<name>` for every name in `RealQuantLinear.list_of_scale_tensors`, but that list carried `"double_scale"` where the buffer is `_double_scale` — a missing underscore. So `_scale` was stripped and `_double_scale` was not, and it reached the checkpoint as `*.weight_quantizer._double_scale` (560 entries in the 70B checkpoint). Downstream loaders reject it before loading any weight: ``` KeyError: 'layers.0.mlp.down_proj.weight_quantizer._double_scale' RuntimeError: Engine core initialization failed. ``` This is why only the TensorRT backend appeared usable in the bug report — its converter tolerates the stray key, then produces `!!!!!!` output from the broken scales, while the PyTorch backend (vLLM / TensorRT-LLM) fails to load outright. The fix reuses the per-block scale captured at compression time, rescaled into the exported `weight_scale_2` convention, and corrects the typo above. The rescale matters: compression normalizes per-block FP8 scales against the global scale it captured at that moment, which is not the post-calibration `weight_scale_2` the export writes. Exporting the stored scale as-is loads fine but leaves every block off by a constant factor (~1.96x measured), so the `weight_scale * weight_scale_2` product that dequantization consumes must be preserved. Note the typo fix also affects `modelopt/torch/quantization/plugins/megatron.py:505,511`, which filter on the same list — these are internal buffers so excluding them looks correct, but calling it out since it is a behavior change outside the export path. **Compression-time scale layout.** `TensorQuantizer._real_quantize` calls `NVFP4QTensor.quantize(..., try_tensorrt=True)`, so on an FP4-capable device with TensorRT-LLM importable the stored `_scale` is the **cutlass-swizzled 1-D uint8** scale rather than the modelopt 2-D E4M3 layout. Confirmed on GB10 in a TRT-LLM container: ``` logical weight (512, 256) -> modelopt scale should be (512, 16) e4m3 _scale : (8192,) torch.uint8 (ndim=1) <- cutlass-swizzled after cutlass_fp4_scale_to_modelopt_fp4_scale: (512, 16) torch.float8_e4m3fn ``` The export therefore normalizes it the same way `NVFP4QTensor.dequantize` does, and raises if `tensorrt_llm` cannot be imported to convert, rather than writing raw byte values. ### Usage No API change. The previously broken path now works: ```bash python hf_ptq.py --pyt_ckpt_path <local_ckpt_dir> --qformat nvfp4 \ --low_memory_mode --export_path <out> ``` ### Testing All runs on DGX Spark (GB10, sm121, aarch64). Both compression-time scale layouts are covered, since the layout depends on whether TensorRT-LLM is importable in the process: **New tests** - `tests/gpu/torch/export/test_export_weight_gpu.py::test_export_compressed_nvfp4_weight` — dense E4M3 path. Asserts the per-block scale covers the logical input dim, that `weight_scale * weight_scale_2` matches an uncompressed export of the same model, and that `postprocess_state_dict` strips both internal buffers. - `tests/gpu_trtllm/torch/export/test_export_compressed_nvfp4.py::test_export_compressed_nvfp4_weight_trtllm_scale` — cutlass-swizzled path. Asserts as a *precondition* that the environment really produced a 1-D uint8 scale, so it cannot silently degrade into the dense case when TensorRT-LLM is absent. **Results** | suite | TensorRT-LLM 1.3.0rc17 container | vLLM 26.05 container | | --- | --- | --- | | both files above | 3 passed | 2 passed, 1 skipped | **Negative controls** (each test fails without the code it guards) - On `main` with only the test file applied: `test_export_compressed_nvfp4_weight` fails. - With only the un-swizzle conversion neutered, the rest of the fix intact: `test_export_compressed_nvfp4_weight_trtllm_scale` fails. **End-to-end, TensorRT-LLM container** (the environment from the bug report, where `_scale` is swizzled) — TinyLlama-1.1B, `--qformat nvfp4 --low_memory_mode`, plus a normal export as control. Exported checkpoints are structurally identical (0 stray `_double_scale` keys vs. 560 on `main`; 663 keys each; `q_proj.weight_scale` `[2048, 128]` FP8 in both; QKV `weight_scale_2` unified in both). Loaded on the **TensorRT-LLM PyTorch backend** — the backend reported as unusable: ``` ##### trt_normal ##### GEN: France and is the most populous city in the country. It is located on the Seine River... ##### trt_lowmem ##### GEN: France and is the most popular tourist destination in the country. It is a city of art, history... ``` On `main` that same load fails with `KeyError: '...weight_quantizer._double_scale'` before a single weight is read. **End-to-end, vLLM container** (dense scale path) — TinyLlama-1.1B, NVFP4 + `--low_memory_mode`: - before: `KeyError: '...weight_quantizer._double_scale'`, engine fails to start - after: loads and generates coherently (`"Paris is the capital of"` → `" France and is best known for the awe inspiring Notre-Dame C"`) - The fixed checkpoint is structurally equivalent to a normal (non-low-memory) export: identical key set (663 tensors), `input_scale` identical, and `weight_scale_2` identical — including the values unified across fused groups, so the fused-GEMM contract (one shared `weight_scale_2` per QKV / gate-up group) still holds. - DeepSeek-R1-Distill-Llama-70B (the model in the bug) reproduces the same signature on `main` and is the source of the numbers in the table above. ### Known limitation: fused groups are slightly less accurate than a full-memory PTQ This restores a working checkpoint, but `--low_memory_mode` NVFP4 is **not numerically identical** to a normal PTQ, and cannot be made so at export time. The nibbles are packed at *load* time against the layer's own global scale, so the effective per-block scale baked into them is `fp8_own * ws2_own`. `preprocess_linear_fusion` later unifies `weight_scale_2` across a fused group to the group max, and the format requires the per-block scale be E4M3, so the best the export can write is `round_fp8(fp8_own * ws2_own / ws2_unified)`. For the group member owning the max amax that ratio is exactly 1 and the round trip is bit-exact; for the others it costs one extra E4M3 rounding, bounded by a half-ULP (6.25%). The uncompressed path never pays this because its weights are still high precision at export, so `to_quantized_weight` re-quantizes the nibbles *after* unification. Measured on TinyLlama-1.1B (weight relative error vs. the source BF16 weights, 154 quantized tensors): | | mean rel. error | | --- | --- | | normal PTQ | 0.090248 | | `--low_memory_mode` (this PR) | 0.091443 | The degradation is confined to exactly 66 of 154 tensors = 22 layers x 3, i.e. the non-max members of each `q/k/v` and `gate/up` group (worst observed +0.0066, e.g. `layers.11.self_attn.k_proj` 0.0898 -> 0.0964). The 88 remaining tensors — group winners plus the unfused `o_proj` / `down_proj` — are bit-exact. Removing this requires compressing a fusion group against one shared scale in the compress-on-load path (`RealQuantParameterDict`), where the group max first becomes known; that is a larger change and is left as a follow-up. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information NVBug 5987078. Not addressed here: at 70B scale on DGX Spark, `--low_memory_mode` can also crash *before* export when the device map offloads, because `QTensorWrapper.to()` cannot represent a `meta` tensor: ``` accelerate/hooks.py: set_module_tensor_to_device(module, name, "meta") RuntimeError: Attempted to call `variable.set_data(tensor)`, but `variable` and `tensor` have incompatible tensor type. ``` It is reproducible on demand under low free GPU memory (crashes at 41.2 GB and 58.5 GB free; succeeds at 88.1 GB) and is easy to hit on Spark's unified memory, where page cache from reading the checkpoint counts against `torch.cuda.mem_get_info()`. That is an independent defect and will be filed separately. Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
23be82655b |
Add nvfp4_act_headroom activation calibration for NVFP4 (#2028)
### What does this PR do?
Type of change: new feature
Adds `nvfp4_act_headroom`, a calibration algorithm for NVFP4
**activation** global scales.
NVFP4 scales a tensor in two levels: an FP8-E4M3 scale per 16-element
block, plus one per-tensor global scale. Plain `max` calibration sets
that global scale from the largest per-block amax seen during
calibration, which leaves no room above it — any activation larger than
the calibration max saturates.
`nvfp4_act_headroom` instead anchors the global scale to a low
percentile of the per-block amax distribution:
```
amax = max(rho * anchor, floor)
```
`anchor` is the per-block amax at `anchor_percentile`; `upper` is the
per-block amax at `upper_percentile` and is the top of the range the
scale commits to representing. Multiplying the low anchor by `rho`
places the calibrated blocks in the lower part of the FP8 scale range
and leaves the rest as headroom.
`upper_percentile` defaults to 99.99 rather than the literal maximum on
purpose. Flooring at the literal max means one freak block drags the
global scale up until every other block's FP8 block scale falls below
subnormal: on a tensor with a single block seven orders of magnitude
above the rest, that flushes 99.998% of elements to zero, versus 6.7%
when the rare blocks are clipped instead. On a benign wide-range tensor
the two choices differ by 0.4% relative MSE. Set `upper_percentile=100`
to floor at the literal observed max, which guarantees no calibration
data is clipped, at that exposure. The calibrator warns when the
per-block range is too wide for `rho` to clear any headroom.
The algorithm applies only to NVFP4 dynamic-block **input** quantizers;
weight quantizers and everything else keep plain `max`.
Files:
- `calib/nvfp4_act_headroom.py` — `NVFP4ActHeadroomCalibrator` (log2
histogram of per-block amaxes, bounded memory)
- `config.py` — `NVFP4ActHeadroomCalibConfig` (`anchor_percentile`
default 1, `rho` default 16384)
- `model_calib.py` — `nvfp4_act_headroom_calibrate`
- `mode.py` — mode registration
- `modelopt_recipes/general/ptq/nvfp4_act_headroom-kv_fp8_cast.yaml` —
mirrors `nvfp4_default-kv_fp8_cast` (dynamic NVFP4 W4A4 + FP8 KV-cache
cast, same module coverage) with only the calibration algorithm swapped:
```diff
- algorithm: max
+ algorithm:
+ method: nvfp4_act_headroom
+ anchor_percentile: 1
```
### Why a dedicated collector rather than reusing `HistogramCalibrator`
`HistogramCalibrator` histograms the tensor values themselves into
linear, dynamically re-ranged bins, and its percentile mode returns an
amax that *clips* a high-tail fraction. This algorithm needs a different
statistic (per-block amaxes, a derived quantity), different bin spacing
(log2, because block amaxes span many decades and the anchor is read
from the **low** tail, where linear bins have almost no resolution), and
a different reduction (a low-percentile anchor scaled up, not an
upper-tail clip). Reusing it would mean changing its collection
semantics for every existing caller; composing it would still leave the
log-spacing and the derived statistic unaddressed. The collector here is
a fixed 512-bin int64 histogram per quantizer, so the memory argument
for a histogram (rather than retaining values) is preserved.
### Usage
```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> \
--recipe general/ptq/nvfp4_act_headroom-kv_fp8_cast --dataset <calib.jsonl>
```
Or directly:
```python
import modelopt.torch.quantization as mtq
NVFP4 = {"num_bits": (2, 1), "block_sizes": {-1: 16, "type": "dynamic", "scale_bits": (4, 3)}}
config = {
"quant_cfg": [
{"quantizer_name": "*", "enable": False},
{"quantizer_name": "*weight_quantizer", "cfg": NVFP4},
{"quantizer_name": "*input_quantizer", "cfg": NVFP4},
],
"algorithm": {"method": "nvfp4_act_headroom", "anchor_percentile": 1, "rho": 16384},
}
model = mtq.quantize(model, config, forward_loop)
```
### Testing
**Unit tests** —
`tests/unit/torch/quantization/test_nvfp4_act_headroom.py`, 26 tests
covering the anchor and range terms, monotonicity in
`anchor_percentile`, headroom above plain max, the too-wide-range
fallback and its warning, the rare-outlier case (the default clips it;
`upper_percentile=100` chases it), the `upper_percentile=100`
no-clipping guarantee across six distributions, input validation (`rho`
/ percentile bounds, NaN/Inf, all-zero, a last dim that is not a
multiple of the block size), quantizer selection, calibrator restoration
after calibration, `shared_states` / `sync_expert_weight_amax`
forwarding, and W4A4 end-to-end confirming weights fall back to plain
max while only activations get the headroom scale.
Full `tests/unit/torch/quantization/` and `tests/unit/recipe/` suites
pass with no regressions (the remaining failures in this environment are
a pre-existing broken-`torchvision` import and
`test_data_parallel_auto_quantize`, both of which fail identically on
`main`).
**End-to-end on a real model** — ran the shipped recipe on a 30B hybrid
Mamba/attention MoE (1024 calibration samples @ seq 4096) and verified
the export is a standard NVFP4 checkpoint: 6268 activation quantizers
calibrated; 19G vs the 62G BF16 source; 6260 `input_scale`, 6267
`weight_scale` / `weight_scale_2`; `quant_algo: NVFP4`,
`kv_cache_quant_algo: FP8`; 53 excluded modules; no module carrying an
`input_scale` without a `weight_scale`. 32 of 6268 quantizers (0.5%) had
a per-block range too wide for `rho` to clear and warned.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — purely additive: a new opt-in
algorithm, config class, and recipe. No existing behavior changes.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no
copied code, no new dependencies.
- Did you write any new necessary tests?: ✅ — 12 new unit tests.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — entry under 0.47 → New Features → Quantization.
- Did you get Claude approval on this PR?: ❌ — not yet run.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
55f1880f70 |
[skill] evaluation: correct the stale vLLM CUDA-13 image tag rule (#2042)
### What does this PR do? Type of change: Bug fix (agent skill documentation) The `evaluation` skill instructed the agent to **"append `-cu130` to the image tag"** for NVFP4 checkpoints on Blackwell B300/GB300 (sm_103). That was correct for v0.19.x, but **vLLM inverted its tag convention at v0.20.0**: the *unsuffixed* tag is now the CUDA-13 build, and `-cu129` is the CUDA-12 opt-out. Consequences of the stale rule: - `v0.20.1-cu130` / `v0.24.0-cu130` / `v0.26.0-cu130` **do not exist** — following the rule literally asks for a missing tag. - The documented fallback (`cu130-nightly-<arch>`) points at ~v0.20-era builds that are *older* than several models' documented minimum vLLM version, so it can't serve as an escape hatch either. This replaces the "append a suffix" instruction with a version-keyed table plus the durable check: **select a tag whose config blob reports `CUDA_VERSION` ≥ 13**, resolving the child manifest for the platform you actually deploy on. While the PR was open the default image was also bumped, and review surfaced two follow-on corrections. Full contents: 1. **Tag-convention fix** — version-keyed table, `-cu130` fallback removed. 2. **Default image `v0.19.1` → `v0.26.0`** (latest vLLM release) everywhere it was pinned: `SKILL.md` Step 3 and the Step 7.5 table, `example_eval.yaml`, `example_eval_next.yaml`. Version specifics that the bump made stale or self-contradictory were dropped (the `e.g. v0.20.0` bump example and the MiniMax-M2.7 `≥0.20.0` anecdote, both now *below* the default; the failure-mode lesson is kept). 3. **Convention boundary corrected to v0.20.0** — the first draft said `≤ v0.20.x` suffixed / `≥ ~v0.21` unsuffixed. Off by a minor release in both rows; see Testing. 4. **Config blob resolved per deployment platform** — the check said "arm64 child". GB300/Grace is arm64, but plenty of B300 deployments are `linux/amd64`. 5. **Same corrections applied to the `deployment` skill**, which carried the original append rule untouched: its NVFP4 note, `references/support-matrix.md`, `references/benchmarking.md`, and the `:latest` pins in `references/setup.md` (now `v0.26.0`, matching the evaluation skill's never-`:latest` stance). An earlier revision of this branch also carried a `.claude/skills/benchmark-model-kernels` symlink, added automatically by `tools/precommit/sync_claude_skills.sh` — it repairs missing symlinks repo-wide on any touch of `.agents/skills/`, and #1980 landed that skill without its link. It has been dropped from this branch to keep the scope on the vLLM image guidance. Worth its own one-line PR: without the symlink, Claude Code doesn't load that skill at all. ### Usage ```bash # Durable check — resolve the child manifest for YOUR platform (arm64 for # Grace/GB300, amd64 for x86) and read CUDA_VERSION from its config blob: # v0.19.1 -> CUDA_VERSION=12.9.1 (unsuffixed = CUDA 12, old convention) # v0.19.1-cu130 -> CUDA_VERSION=13.0.1 (suffixed = CUDA 13, old convention) # v0.20.0 -> CUDA_VERSION=13.0.2 (transition release: ships both suffixes) # v0.26.0 -> CUDA_VERSION=13.0.2 (unsuffixed = CUDA 13, new convention) # v0.26.0-cu129 -> CUDA_VERSION=12.9.1 (suffixed = CUDA 12, new convention) ``` ### Testing Verified empirically against the Docker registry API for `vllm/vllm-openai` — resolved each tag's child manifests and read `CUDA_VERSION` / `TORCH_CUDA_ARCH_LIST` from the config blob. **Where the convention flips (arm64):** | release | unsuffixed | `-cu130` | `-cu129` | |---|---|---|---| | v0.18.0 | 12.9.1 | 13.0.1 | absent | | v0.19.0 / v0.19.1 | 12.9.1 | 13.0.1 | absent | | **v0.20.0** | **13.0.2** | 13.0.2 | 12.9.1 | | v0.20.1 | 13.0.2 | **absent** | 12.9.1 | | v0.20.2 | 13.0.2 | **absent** | 12.9.1 | | v0.21.0 … v0.26.0 | 13.0.2 | absent | 12.9.1 | v0.20.0 is the transition release — it publishes both suffixes *and* its unsuffixed tag is already CUDA 13. That duplication is what hid the boundary: confirming `v0.20.0-cu130` exists reads as "old convention still applies at 0.20", while `v0.20.1-cu130` and `v0.20.2-cu130` don't exist at all. **Why the platform matters.** `CUDA_VERSION` is identical across children on every tag checked (v0.26.0, v0.26.0-cu129, v0.20.0, v0.19.1, v0.19.1-cu130, kimi-k3), but `TORCH_CUDA_ARCH_LIST` is not: ``` v0.26.0 amd64 7.5 8.0 8.6 8.9 9.0 10.0 12.0 v0.26.0 arm64 8.0 8.7 8.9 9.0 10.0 11.0 12.0 v0.19.1 amd64 7.0 7.5 8.0 8.9 9.0 10.0 12.0 v0.19.1 arm64 8.7 8.9 9.0 10.0+PTX 12.0 ``` `11.0` appears only on arm64, `7.5` / `8.6` only on amd64 — so the arch check has to read the child you'll actually run. **Default bump.** The `0.26.0` family is `{,-aarch64,-x86_64} × {,-cu129} × {,-ubuntu2404}` — 12 tags, `cu129` the only CUDA axis, no `-cu130`. The new default is therefore already a CUDA-13 build, and NVFP4 on B300/GB300 needs no suffix at all. Docs-only change; no runtime code touched. `pre-commit` clean. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (agent skill documentation) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ (not yet run) ### Additional Information Split out of the GDPVal skill work (#2039) because it is independent of GDPVal and applies to every NVFP4-on-Blackwell deployment the skill generates. Now spans both the `evaluation` and `deployment` skills, all under `.agents/`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated deployment and evaluation guidance to use vLLM v0.26.0. * Clarified CUDA image-tag conventions, including CUDA 13 defaults and CUDA 12 opt-outs. * Added validation guidance for resolved CUDA versions, platform architecture settings, and image compatibility. * Updated NVFP4 Blackwell B300/GB300 support notes, including required `sm_103` kernel availability. * Refreshed example recipes, setup commands, benchmarking notes, and serving-image requirements. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7b8da80205 |
docs(recipes): sync ptq.md with shipped recipes and enforce via unit tests (#1970)
### What does this PR do? **Type of change:** documentation (+ new tests) `modelopt_recipes/ptq.md` is meant to track every shipped PTQ recipe, but recent recipe PRs did not update it. This PR brings the doc back in sync and adds unit tests so it can't drift again. **Doc sync — general recipes:** - Add `nvfp4_experts_only_input_scale1-kv_fp8_cast` (#1947) to the shipped-recipes table (now 20 recipes) and document the `input_scale1` variant (constant amax 2688 → exported NVFP4 `input_scale == 1.0`, expert activations uncalibrated). - Add a pointer that the PTQ recipes' `quantize` sections also drive QAT/QAD, with the LSQ / Dual-LSQ QAD recipes under `general/qad/` (#1884). **Doc sync — model-specific recipes:** - Add `gemma4` (algorithm override, #1690), `diffusion_gemma` (extra `*self_conditioning*` exclusion, #1707), `vit` (FP8 with attention BMM quantizers for Torch-TRT, #1569), and the qwen3_5 / qwen3_5_moe `w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast` twins (#1620). - Rewrite the checkpoint-mirror section for the `huggingface/models/nvidia/` tier: Super recipe rename (`super-nvfp4.yaml` → `nvfp4-mse.yaml` / `nvfp4-max-calib.yaml`), Nano-4B GGUF Q4_K_M mirror (#1327/#1606), Ultra-550B `nvfp4-4o6` (#1684); fix stale `models/<checkpoint>` paths. **Enforcement — new `tests/unit/recipe/test_recipe_docs.py`:** 1. Every `general/ptq/*.yaml` stem must appear (backticked) in `ptq.md`. 2. Every recipe row in the shipped-recipes table must exist on disk (catches renames/removals). 3. The "All N `general/ptq/` recipes" count must match the file count. 4. Every model folder under `huggingface/` containing `ptq/*.yaml` must be mentioned in the doc. ### Usage ```bash pytest tests/unit/recipe/test_recipe_docs.py -q ``` ### Testing - `pytest tests/unit/recipe/ -q` → 218 passed (includes the 4 new tests). - Verified enforcement bites: against the pre-update `ptq.md`, 3 of the 4 new tests fail with actionable messages. - `pre-commit run --files modelopt_recipes/ptq.md tests/unit/recipe/test_recipe_docs.py` → all hooks pass. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ (docs + tests only; no runtime code changed) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — `tests/unit/recipe/test_recipe_docs.py` - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (documentation/test change) - Did you get Claude approval on this PR?: ❌ (pending) ### Additional Information Follow-up to the recipe guide added in #1662. The doc-consistency tests could also be wired into `tools/precommit/check_modelopt_recipes.py` for commit-time feedback if reviewers prefer both layers. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added the `nvfp4_experts_only_input_scale1-kv_fp8_cast` recipe to the general PTQ catalog. * Documented the new `input_scale1` calibration behavior and related quantization/caching details. * Expanded and reworked PTQ documentation, including updated navigation, “Beyond PTQ” guidance, new/updated model-specific recipes, and expanded checkpoint mirror entries. * Updated the general PTQ shipped-recipe count to 20. * **Tests** * Added automated consistency checks between on-disk shipped recipe YAML files and the PTQ documentation (coverage, existence, and counts). <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
a05850bffa |
Add nel-next (0.3.x) agentic AA benchmark support to eval skill (#1861)
### What does this PR do?
Type of change: new feature (agent skill — nel-next agentic eval
support)
Adds support for the **nel-next** evaluator (`nemo-evaluator[harbor]`
0.3.x) to the `evaluation` agent skill, for the **agentic** AA
benchmarks (Terminal-Bench 2.1, SWE-bench Verified) that do not run on
the default `nemo-evaluator-launcher` 0.2.6 (different package, CLI,
override syntax, and config schema).
- `.agents/scripts/nel-next.sh` — bootstrap + run nel-next from an
**isolated venv**, leaving the 0.2.6 install untouched.
- `references/nel-next.md` — shared reference (venv,
`services`/`benchmarks`/`cluster`/`output` schema, harbor/ECS-Fargate
architecture, timeout strategy, MLflow export, run flow, gotchas).
- `recipes/tasks/aa_next/{terminal_bench_2_1,swebench_verified}.md` +
`recipes/examples/example_eval_next.yaml` — per-benchmark recipes +
self-contained template.
- Internal harbor infra (`eval_image`, ECR repos) is **not hardcoded** —
configs read `${NEL_NEXT_EVAL_IMAGE}` / `${HARBOR_*_ECR_REPOSITORY}`
from `.env`, written by the companion `modelopttools:eval-config` skill
(internal repo, separate change). `SKILL.md` + `recipes/env.example`
updated.
### Usage
```bash
.agents/scripts/nel-next.sh --setup-only # one-time isolated 0.3.x venv
set -a && source .env && set +a # HF_TOKEN, AWS_*, harbor infra vars
.agents/scripts/nel-next.sh eval run <config>.yaml --dry-run # then --submit
```
### Testing
- Dry-run validated against `nemo-evaluator` 0.3.0 (`extra="forbid"`
schema) for both the Terminal-Bench 2.1 and SWE-bench Verified configs.
- Live runs on gcp-nrt (MiniMax-M2.7 NVFP4, 8×B200): Terminal-Bench 2.0
completed `pass@1=0.469` (712 trials, auto-resume across walltime
windows); SWE-bench Verified canary passed end-to-end (vLLM deploy →
OpenHands agent → AWS ECS Fargate sandbox → scoring/merge/report).
- `pre-commit run` clean (insert-license, markdownlint, YAML format,
symlink sync).
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (additive — new files + one
branch section in the eval skill; the 0.2.6 path is unchanged)
- If you copied code or added a new PIP dependency, did you follow
`CONTRIBUTING.md`: N/A (the helper installs `nemo-evaluator` into a
runtime venv; no new repo dependency)
- Did you write any new necessary tests?: N/A (agent-skill docs/tooling;
validated via dry-run + live cluster runs)
- Did you update Changelog?: N/A (agent skill content, not a
Model-Optimizer feature/API)
- Did you get Claude approval on this PR?: N/A
### Additional Information
Companion change (separate internal repo, `modelopt-internal`): the
`modelopttools:eval-config` skill gains a Step 3b that writes the harbor
infra rows (`NEL_NEXT_EVAL_IMAGE`, `HARBOR_ECR_REPOSITORY`,
`HARBOR_SWEBENCH_ECR_REPOSITORY`) into `.env` and points at the
canonical per-benchmark `bench.yaml` source of truth.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added a `nel-next` CLI wrapper to run NeMo Evaluator 0.3.x from an
isolated venv, including setup/version helpers and safer install/refresh
behavior.
* Introduced nel-next (harbor) templates and benchmark runbooks for
Terminal-Bench 2.1 and SWE-bench Verified.
* **Documentation**
* Added a dedicated nel-next reference guide covering configuration,
required workflow, reporting, and common gotchas.
* Updated evaluation skill docs and the `.env` example with
nel-next/harbor credential and output variables, plus new example
evaluation configuration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
138564f443 |
Add AA-Omniscience eval recipe; harden judge/run conventions in the eval skill (#1834)
### What does this PR do? Type of change: documentation (agent `evaluation` skill) Updates the `evaluation` agent skill (`.agents/skills/evaluation/`) with an AA-Omniscience recipe plus several judge/run hardening conventions found while running the AA Index v2 suite on quantized checkpoints. - **AA-Omniscience recipe** (`recipes/tasks/aa/omniscience.md`): nemo-skills `ns_omniscience`, AA Index v2 params (`++parse_reasoning=False`, `num_repeats=10`), `gcp/google/gemini-3-flash-preview` judge; score `omniscience_pass_at_1_avg-of-N_judge_correct`. Added to the AA Index v2 suite list in `SKILL.md`. - **Judge model_ids hardcoded** in the HLE / AA-LCR / Tau2 recipes (with a "swap for an equivalent on your own endpoint" note); shared judge URL var renamed `NS_JUDGE_URL` -> `INFERENCE_JUDGE_URL`; judge API key folded into `INFERENCE_API_KEY` in `env.example`. - **Idle-reaper exemption** in `example_eval.yaml`: `cluster.sbatch_comment` exempts eval jobs from the `OccupiedIdleGPUsJobReaper`, which otherwise CANCELs jobs whose GPUs sit idle during model load / judge calls / aggregation. - **Resume-on-kill note** in `SKILL.md`: after a preemption / idle-reaper CANCEL (not a walltime timeout, which NEL auto-resumes), re-submit the job's `run.sub` to resume from the response cache (`skip_filled`) with cumulative progress. ### Usage \`\`\`bash nel run --config recipes/examples/example_eval.yaml # now ships the idle-reaper exemption # extend evaluation.tasks with the AA-Omniscience fragment from recipes/tasks/aa/omniscience.md \`\`\` ### Testing Validated end-to-end on gcp-nrt (B200): MiniMax-M2.7-NVFP4 AA-Omniscience full run (600 questions x 10 repeats) completed with \`omniscience_pass_at_1_avg-of-10_judge_correct = 18.77\`. The idle-reaper exemption and the \`sbatch run.sub\` resume path were both exercised - a reaped run resumed from cache and finished all 10 repeats. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: yes (skill docs/recipes only; no library API change) - If you copied code from any other sources or added a new PIP dependency: N/A - Did you write any new necessary tests?: N/A (agent skill recipes/docs) - Did you update Changelog?: N/A (agent skill, not a library feature) - Did you get Claude approval on this PR?: pending (/claude review) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an **AA-Omniscience** evaluation recipe. * Expanded the default **AA Index v2** quantized-checkpoint validation suite to include Omniscience. * **Documentation** * Clarified which judge/user-simulator **model identifiers** are fixed in recipes vs supplied via environment variables. * Updated HLE, LCR, and Tau2-Bench Telecom guidance for judge configuration and API-key/endpoint usage. * Added instructions for resuming evaluations after scheduler preemption/cancellation. * **Chores** * Updated the example evaluation config to include a GPU job reaper policy. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
1c6bdb3021 |
Fix reduce_amax NotImplementedError on FP8 weights (NVBug 6360175) (#1824)
### What does this PR do? Type of change: Bug fix Fixes [NVBug 6360175](https://nvbugspro.nvidia.com/bug/6360175) / OMNIML-5265: quantizing a model whose weights are stored natively in FP8 (e.g. DeepSeek-V3 in `float8_e4m3fn`) crashes during `mtq.quantize` calibration with: ``` File ".../modelopt/torch/quantization/utils/core_utils.py", line 162, in reduce_amax max_val = torch.max(input) NotImplementedError: "max_all_cuda" not implemented for 'Float8_e4m3fn' ``` **Root cause:** FP8 dtypes (`float8_e4m3fn` / `float8_e5m2`) implement no full-tensor reduction kernel (`max_all_cuda`/`min_all_cuda`), nor `amax`/`amin`, `abs`, or elementwise `maximum`. `reduce_amax` called these directly on the FP8 weight tensor. **Fix:** Upcast FP8 inputs to the default float dtype (`torch.get_default_dtype()`) at the top of `reduce_amax`, before any reduction. The upcast is **lossless** (any default float dtype represents every FP8 value exactly) and only affects the FP8 path — the common (fp16/bf16/fp32) path is untouched. Placing the upcast at the top covers all branches (`torch.max`/`min`, `torch.amax`/`amin`, `torch.abs`), not just the line in the traceback. ### Usage No API change. Quantization of natively-FP8 checkpoints (e.g. DeepSeek-V3 NVFP4 PTQ) now runs through calibration instead of raising. ### Testing - New CPU regression test `test_reduce_amax_fp8` in `tests/unit/torch/quantization/test_utils.py` covering both FP8 dtypes (`float8_e4m3fn`, `float8_e5m2`) across all axis modes (`None`, `0`, `1`, `(0, 1)`); asserts results equal the float reference and the output dtype is the default float dtype. CPU reproduces the original error (no FP8 reduction kernel there either), so the test is GPU-free. - `pre-commit run --files ...` passes (ruff, mypy, bandit, license, rst checks). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ (0.45 Bug Fixes) - Did you get Claude approval on this PR?: ❌ (not yet) ### Additional Information NVBug 6360175 is tagged `Committed_ModelOpt_0.45.0` (regression); the changelog entry is under 0.45 and this will be cherry-picked to `release/0.45` after merge. Supersedes #1823, which got a stuck head ref (frozen at the original commit, no sync on force-push) after the repo move `TensorRT-Model-Optimizer` → `Model-Optimizer`; it could not be re-synced or reopened, so this PR replaces it from the same branch. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9cfd7dd3e8 |
Align eval skill AA benchmarks to golden NeMo configs + harden skill (#1790)
### What does this PR do?
Type of change: Documentation (agent skill under
`.agents/skills/evaluation/`)
Aligns the `evaluation` agent skill's AA benchmark recipes with NeMo
Evaluator's Nemotron-3-Ultra golden reproducibility configs, and fixes
several robustness gaps surfaced while validating the changes end-to-end
on SLURM (MiniMax-M2.7, FP8 + ModelOpt NVFP4).
**AA benchmark alignment**
- GPQA: simple-evals `gpqa_diamond_aa_v3` → nemo-skills `ns_gpqa`
(`++prompt_config=eval/aai/mcq-4choices`).
- MMLU-Pro: `mmlu_pro_aa_v3` → `ns_mmlu_pro` (`mcq-10choices-boxed`).
- Repeat counts unchanged (GPQA 16, MMLU-Pro 1); score-extraction metric
keys updated to the nemo-skills names
(`gpqa_pass_at_1_avg-of-N_symbolic_correct`,
`mmlu-pro_pass_at_1_symbolic_correct`, verified against MLflow run
data).
- HLE: add golden knobs `hle_strict_judge: true` +
`++server.enable_soft_fail=True`.
- Updated `references/{quantization-benchmarks,parallelism}.md` and the
example config's default task to match.
**Skill robustness fixes**
- Invoke `modelopttools:eval-config` at Step 1 for judge-scored runs;
create/populate `.env` before Step 5 (it was only created at Step 8) —
fixes a fresh-environment ordering gap.
- Secret-safety rule: never open `.env` with file tools (the harness
mirrors later external edits back into context, leaking keys) — interact
via shell only.
- Require fetching `recipes.vllm.ai` for the **exact model variant** and
bumping the image to its minimum vLLM — variant minimums differ (e.g.
MiniMax-M2 ≥0.11.0 vs M2.7 ≥0.20.0), and running below crashes
mid-inference with `CUDA illegal memory access`. Note that `--dry-run`
does not validate the image/version.
- Remove internal cluster names from the public skill docs.
- gitignore the local eval run-config dir.
### Usage
N/A — agent skill docs/recipes; no library/API change.
### Testing
Validated on SLURM with MiniMax-M2.7 (16-sample canaries):
- FP8: GPQA + MMLU-Pro ran through the new nemo-skills harness and
produced valid scores (e.g. MMLU-Pro `symbolic_correct=75.0`),
confirming the harness switch + metric keys.
- NVFP4: confirmed the vLLM-min-version finding — deploy crashed on
0.19.1 (`CUDA illegal memory access` mid-inference) and succeeded on
0.20.2 (the recipe minimum).
- (HLE tokenizer gap was reproduced deterministically; its fix is a
separate follow-up, not in this PR.)
### Before your PR is "*Ready for review*"
Commits are signed (`git commit -s -S`). ✅
- Is this change backward compatible?: N/A — agent skill/recipe docs.
Note: the GPQA/MMLU-Pro harness switch changes the MLflow metric keys,
so new runs are not directly comparable to prior simple-evals baselines
(re-baseline).
- If you copied code from any other sources or added a new PIP
dependency: N/A
- Did you write any new necessary tests?: N/A (docs/recipes only)
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ❌ (run `/claude review`)
### Additional Information
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Strengthened evaluation setup guidance: `.env` is now required earlier
with safer creation/sourcing rules, and endpoint placeholders are
clarified.
* Updated judge/user-simulator and vLLM deployment instructions,
including exact minimum `recipes.vllm.ai` variant/image requirements,
expected failure symptoms, and a warning about `--dry-run`.
* Refreshed AA benchmark/task recipes to the `nemo-skills` harness
(e.g., `ns_gpqa`, `ns_mmlu_pro`) with updated parameters and
score-metric guidance, plus stricter HLE judging options.
* **Bug Fixes**
* Improved score-reporting guidance for GPQA and MMLU-Pro to match
current outputs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
106781659e |
Add W4A16 NVFP4-MSE Qwen3.5 dense/MoE PTQ recipes (#1620)
### What does this PR do? Type of change: new feature (PTQ recipe) Adds an MSE-calibrated counterpart of the existing `w4a16_nvfp4-fp8_attn-kv_fp8_cast` PTQ recipe for the Qwen3.5 family (dense `qwen3_5` and MoE `qwen3_5_moe`). New files: - `modelopt_recipes/huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg.yaml` — shared `quant_cfg` snippet - `modelopt_recipes/huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml` — dense recipe - `modelopt_recipes/huggingface/qwen3_5_moe/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml` — MoE recipe The only difference from the `max` variant: NVFP4 MLP / `lm_head` weight scales come from an MSE FP8-scale sweep (`method: mse`, `fp8_scale_sweep: true`, `nvfp4_static`) instead of max calibration. FP8 attention / linear-attention projections and the FP8 KV cast are unchanged. The dense and MoE families share a single `quant_cfg` snippet under `qwen3_5/ptq`, matching the existing recipe's layout. ### Usage ```bash # Dense python hf_ptq.py --recipe huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast ... # MoE python hf_ptq.py --recipe huggingface/qwen3_5_moe/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast ... ``` ### Testing - Both recipes load and resolve their `$import`s via `modelopt.recipe.loader.load_recipe`. - The `check-modelopt-recipes` pre-commit validator passes on all three files. - Verified against the source that weight-only MSE (`method: mse` + `fp8_scale_sweep: true`) is supported for W4A16: `mse_calibrate` refines only weight quantizers (`iter_weights_for_calibration`) and needs no input/activation quantizers, and the `nvfp4_static` numeric satisfies `is_nvfp4_static` so the FP8 scale sweep engages on the MLP/lm_head weights. No code changes were required. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (config-only; covered by existing recipe-loader validation) - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: ❌ (not yet) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## New Features * Added PTQ recipe configurations for Qwen3.5 and Qwen3.5-MoE model families * Supports W4A16 quantization with NVFP4 static weights and MSE-based calibration * Enables FP8 precision for self-attention and KV-cache optimization for improved model performance and reduced memory footprint <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d7df14d12a |
[tools/debugger] Enforce a single relay owner across hosts (#1735)
### What does this PR do? Type of change: Bug fix The `tools/debugger` file-based relay assumes a single server but never enforced it. Because the relay lives on shared NFS (the repo is often the same checkout mounted on multiple hosts), a forgotten `server.sh` on another host kept polling the same `.relay/` and could **steal commands** (executing them on the wrong host), and killing one server's cleanup could **wipe the active server's markers**. This adds a `.relay/owner` ownership token (`host:pid:nanos`): - Each server writes `owner` atomically at startup and **takes over** instead of refusing when a stale `server.ready` exists (the old `kill -0 <pid>` guard was host-local and meaningless across hosts). - The handshake and main loops exit cleanly if `owner` changes (`[server] Superseded by <id> — exiting.`), so a freshly started server **evicts** any stale one — even on another host. - `cleanup()` only clears shared markers if we still own them, so a stepping-down server never clobbers its successor's `server.ready`/`owner`. Also gitignores `tools/debugger/logs/` and documents the `owner` file in the README. ### Usage ```bash # Inside the container; a previously-running server elsewhere that shares this # NFS .relay/ steps down automatically once this one claims ownership: bash tools/debugger/server.sh # [server] Note: existing server.ready found (<host:pid:ts>); taking over. # (the stale server logs: "[server] Superseded by <id> — exiting.") ``` ### Testing Verified live on computelab: a forgotten `server.sh` on another host was evicted when a new server started, after which `client.sh run` executed on the correct (new) host; confirmed the stepping-down server's cleanup does not remove the successor's `server.ready`/`owner`. `server.sh` passes `bash -n`. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ <!-- additive: new .relay/owner file; client.sh and the wire protocol are unchanged --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A <!-- the file-based relay tool has no test harness; behavior verified manually --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- internal dev tooling, not a shipped feature/API --> - Did you get Claude approval on this PR?: N/A <!-- can run /claude review --> ### Additional Information Scope is limited to `tools/debugger/` (`server.sh`, `README.md`, `.gitignore`). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated relay protocol documentation to clarify ownership-based server coordination. * **Bug Fixes** * Improved reliability of multi-server coordination in shared relay environments. * **Chores** * Updated ignore patterns for logging files. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d26c8af002 |
Fix DeepSeek V3 ptq.py inference-repo path resolution (nvbug 6311147) (#1702)
### What does this PR do? Type of change: Bug fix Fixes nvbug **6311147** (OMNIML-5103). `examples/deepseek/deepseek_v3/ptq.py` resolved the cloned DeepSeek-V3 / DeepSeek-V3.2-Exp inference repos relative to its own directory (`deepseek_v3/`) via `Path(__file__).resolve().parent`. But the [README](https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/deepseek) clones those repos into the parent `examples/deepseek/` directory and runs the script from there, so the lookup landed one level too deep and raised `ValueError: DeepSeek-V3 or DeepSeek-V3.2-Exp not found` (the error message also printed the wrong directory). The fix resolves from `parent.parent` via a single `DEEPSEEK_DIR` base shared by both repo paths and the error message. ### Usage ```bash # Run from examples/deepseek/ as documented in the README, after cloning # DeepSeek-V3 (or DeepSeek-V3.2-Exp) into that directory: torchrun --nproc-per-node 8 --master_port=12346 deepseek_v3/ptq.py \ --model_path $DS_CKPT \ --config DeepSeek-V3/inference/configs/config_671B.json \ --quant_cfg NVFP4_DEFAULT_CFG \ --output_path $FP4_QUANT_PATH ``` ### Testing - Confirmed against the repro path: with the file at `examples/deepseek/deepseek_v3/ptq.py` and the repos cloned into `examples/deepseek/`, `Path(__file__).resolve().parent.parent` now points at `examples/deepseek/` so `DeepSeek-V3/inference` resolves correctly. - Verified the sibling `examples/deepseek/deepseek_v4/` does not share the bug (it takes an explicit `--dsv4_inference_dir` argument instead). - `pre-commit` clean. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (one-line path fix in an example script that requires the DeepSeek repos + multi-GPU checkpoint to exercise) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (bug is in a 0.45-cycle example, not a regression from a released version) - Did you get Claude approval on this PR?: ❌ (not yet run) ### Additional Information nvbug 6311147 / OMNIML-5103. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved path resolution in the example script to more reliably locate the required inference repository. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
60b1af5fb2 |
Fix GPT-OSS MXFP4->NVFP4 PTQ load, export, and cast (nvbug 6295279, 6295242) (#1678)
### What does this PR do? Type of change: Bug fix Fixes the GPT-OSS MXFP4 → NVFP4 PTQ path (`examples/llm_ptq/hf_ptq.py` with `--cast_mxfp4_to_nvfp4`), which failed in three independent ways. The documented command now runs end-to-end and produces a bit-exact (100% lossless) NVFP4 checkpoint. Addresses **nvbug 6295279** (OMNIML-5046) and **nvbug 6295242** (OMNIML-5045). 1. **nvbug 6295242 — CUDA illegal memory access on load.** GPT-OSS ships native MXFP4 weights that Transformers dequantizes to BF16; the threaded weight loader trips an illegal-memory access when `device_map="auto"` shards the dequant across **multiple GPUs**. The missing optional `kernels` package only *forces* the dequant path — it is not the root cause. `get_model` now detects MXFP4 checkpoints and loads them with `Mxfp4Config(dequantize=True)` on a **sequential** device map so the dequant stays on a single device. `kernels` is no longer required. 2. **nvbug 6295279 #1 — `NotImplementedError: Mxfp4GptOssExperts` during unified HF export.** Forcing `dequantize=True` yields plain `GptOssExperts` (even when `kernels` is installed), which ModelOpt wraps and exports normally. 3. **nvbug 6295279 #2 — `FileNotFoundError` in the cast step.** `--cast_mxfp4_to_nvfp4` treated `--pyt_ckpt_path` as a local dir; a HF Hub ID now resolves to its cached snapshot dir via `_resolve_model_path`. Also fixes a **static-block NVFP4 regression** (surfaced by the cast's `force_weight_quantizers_static`, introduced by #1560's now-unconditional `weight_only_quantize`): `_QuantGptOssExperts` / `_QuantLlama4TextExperts` quantize their expert weights transposed in the forward (`_transposed_quantize`), but the inherited `iter_weights_for_calibration` fed the non-transposed weight, locking a mismatched block-quant `_original_shape` and raising `ValueError: Input shape has changed`. The override now calibrates on the transposed view, matching both the forward and the export's `_amax` orientation. ### Why this regressed (it worked when the cast was added) `get_model` never had explicit handling for a *natively pre-quantized MXFP4* checkpoint — GPT-OSS fell through the generic *unquantized-checkpoint* branch and relied on Transformers' **implicit** MXFP4 behavior, which is fragile across three axes. The cast was originally validated (#1372, 2026-05-01) in the "lucky" quadrant of each: - **GPU count:** `device_map="auto"` on a single GPU never shards, so the dequant stays on one device. On multiple GPUs `auto` balances the model and shards the MXFP4→BF16 dequant across devices → CUDA illegal-memory crash (6295242). - **`kernels` presence:** without `kernels`, Transformers auto-dequantizes to BF16 `GptOssExperts` (exportable). With `kernels` installed it keeps the packed `Mxfp4GptOssExperts` kernel path → export `NotImplementedError` (6295279 #1). - **Transformers version:** the kernel-backed experts wrapper and the threaded multi-GPU weight loader are newer-Transformers behavior (env here is 5.5.4). Earlier versions simply dequantized MXFP4 → BF16, which is what the old generic path happened to need. The QA env sat in the *breaking* quadrant (multi-GPU and/or `kernels` present, newer Transformers), so the implicit path failed. The new branch makes both decisions explicit and deterministic (`dequantize=True` + single-device load), regardless of environment — mirroring the existing `has_pack_quantized_config` branch for compressed-tensors checkpoints. The fourth issue (static-block `Input shape has changed`) is a separate regression: it was introduced by **#1560 (2026-06-02, "Make sure all weight quantizers have `_amax`")**, a month *after* the cast landed. #1560 made `weight_only_quantize` unconditional in `max_calibrate`; previously it ran only when no calibration `forward_loop` was supplied, and the cast always supplies one — so the non-transposed weight-quantizer call simply never happened before. The conflict only appears at the intersection of (a) transposed-quantize experts (GPT-OSS/Llama4), (b) static-block NVFP4 — which `--cast_mxfp4_to_nvfp4` forces via `force_weight_quantizers_static` — and (c) #1560. CI's GPT-OSS NVFP4 coverage uses the *dynamic*-block path, which never locks the block shape, so #1560 looked safe. ### Usage ```bash python hf_ptq.py \ --pyt_ckpt_path openai/gpt-oss-20b \ --qformat nvfp4_mlp_only \ --cast_mxfp4_to_nvfp4 \ --export_path ./gpt-oss-20b-nvfp4 ``` ### Testing - Ran the documented command end-to-end on 2xB200 (`openai/gpt-oss-20b`): cast overrode **48/48** expert weight quantizers, **100% lossless** layers/blocks, exported a valid packed-NVFP4 HF checkpoint (uint8 weights + FP8 per-block `weight_scale` + per-tensor `weight_scale_2` + `hf_quant_config.json`). - Verified plain `--qformat nvfp4_mlp_only` (no cast) still works end-to-end. - **Independently verified the export is bit-exact:** dequantized the exported NVFP4 weights (ModelOpt's E2M1 LUT + pack layout) and compared against Transformers' canonical MXFP4→BF16 dequant (`Mxfp4Config(dequantize=True)`) over all 24 layers × both expert weights — `max_abs_err = 0`, 100% bitwise-equal in bf16. So `dequant(exported NVFP4) == dequant(original MXFP4)` exactly. - New unit tests: `test_get_original_hf_quant_method_*` (load detection) and `test_gpt_oss_experts_iter_weights_for_calibration_transposed` (the transpose regression). Existing `test_cast_mxfp4_to_nvfp4.py` (8 tests) still pass. `pre-commit` clean. **Known limitation:** verified for gpt-oss-20b (fits one GPU). gpt-oss-120b dequantized does not fit a single GPU, so `sequential` would still span GPUs — that case would need a CPU-dequant-then-dispatch path and is left as a follow-up. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ (0.45 Bug Fixes) - Did you get Claude approval on this PR?: ❌ (not yet run) ### Additional Information nvbug 6295279, nvbug 6295242 / OMNIML-5046, OMNIML-5045. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Prevented CUDA illegal-memory access during MXFP4→NVFP4 casting. * Fixed expert-weight calibration orientation to avoid shape mismatches. * **New Features** * Support loading native MXFP4 checkpoints with automatic dequantization. * Resolve remote model identifiers to local checkpoints when casting MXFP4→NVFP4, improving reliability. * **Tests** * Added unit and GPU regression tests covering quant-method detection, casting, and expert-weight calibration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
9cb00ce43a |
docs: add modelopt_recipes README and PTQ recipe/scheme guide (#1662)
### What does this PR do?
Type of change: documentation
Adds two docs under `modelopt_recipes/` (no code or behavior changes):
- **`README.md`** — catalog of the recipe library: its purpose (a recipe
is the
single, version-controlled source of truth for *how* a model is
optimized), the
directory layout (`general/`, `huggingface/`, `models/`, `configs/`),
how to
load/select recipes (`load_recipe`, `--recipe`), and a high-level map of
the
general PTQ combos, speculative-decoding, and distillation recipes.
- **`recipe.md`** — a focused guide to the PTQ schemes: the general
`general/ptq/`
body scopes (full-model FP8/NVFP4, scoped experts-only / mlp-only /
omlp-only,
weight-only), KV-cache modes (`kv_fp8_cast` / `kv_nvfp4_cast` /
`kv_fp8`),
calibration variants (max / mse / gptq / layerwise), low- vs
high-concurrency
deployment guidance, and the model-specific recipes under `huggingface/`
and
`models/` — each compared to its general baseline.
### Usage
```python
# Documentation only. The recipes themselves load as before, e.g.:
from modelopt.recipe import load_recipe
cfg = load_recipe("general/ptq/nvfp4_experts_only-kv_fp8_cast")
```
### Testing
`pre-commit run --files modelopt_recipes/README.md
modelopt_recipes/recipe.md`
passes (markdownlint, modelopt recipe validation, license/format hooks).
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: N/A <!-- docs only -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- docs only -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- docs only -->
- Did you get Claude approval on this PR?: ❌ <!-- not yet -->
### Additional Information
Documentation for the `modelopt_recipes/` library; content verified
against the
recipe YAMLs and the `modelopt.recipe` / config-loader source.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Added comprehensive ModelOpt recipes guide describing YAML-based,
composable optimization workflows, directory/lookup layout, reuse via
imports, and how to add or share recipes.
* Added PTQ quantization guide covering recipe naming/structure,
quantization scopes and KV-cache options, calibration variant guidance,
model-specific overrides, multimodal considerations, and a
checkpoint-mirroring example.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
bde162a701 |
feat(deepseek): add --cast_mxfp4_to_nvfp4 to deepseek_v4 quantize step (#1653)
### What does this PR do? Type of change: new feature Brings the GPT-OSS lossless MXFP4 → NVFP4 cast (#1372) to DeepSeek V4's routed-expert export by adding a `--cast_mxfp4_to_nvfp4` flag to `examples/deepseek/deepseek_v4/quantize_to_nvfp4.py`. To avoid duplicating the closed-form math, the shared numerics — `mxfp4_to_nvfp4_global_amax`, `mxfp4_to_nvfp4_per_block_amax`, and the E2M1/E4M3/E8M0 constants — are **hoisted out of the GPT-OSS example cast into the library** at `modelopt/torch/quantization/utils/numeric_utils.py`. Both the GPT-OSS cast (`examples/llm_ptq/cast_mxfp4_to_nvfp4.py`) and the new DeepSeek path now import them from there. DeepSeek V4's routed experts ship as MXFP4 (E2M1 nibbles + a power-of-two E8M0 scale per 32-element block). By default the export dequantizes them to BF16 and re-quantizes to NVFP4 using the calibrated per-tensor weight amax, which re-derives per-block scales from the data and is therefore lossy. With the flag, the cast pins `scale_2 = 2^(k_max-8)` and each per-block E4M3 scale to `2^(k_j-m)` straight from the source E8M0 scales, so `per_block_scale * scale_2 = 2^k_j` and the NVFP4 nibbles equal the source MXFP4 nibbles bit-for-bit (for every block whose `k_j` lands in E4M3's representable window; rare out-of-range blocks clamp). The one V4-specific addition is that w1/w3 share a single `scale_2` for the fused GEMM1, so `k_max` is taken over both projections. The flag only affects routed-expert **weights** — activation `input_scale` still comes from `--amax_path` calibration. ### Usage ```bash python deepseek_v4/quantize_to_nvfp4.py \ --amax_path ${AMAX} \ --source_ckpt ${DS_V4} \ --output_ckpt ${HF_NVFP4_PATH} \ --cast_mxfp4_to_nvfp4 ``` ### Testing - The hoisted numerics get unit tests in `tests/unit/torch/quantization/test_numeric_utils.py` (10 cases: per-tensor global_amax, per-block amax incl. out-of-range, magnitude-table cache) — 10/10 pass. The example test `tests/examples/llm_ptq/test_cast_mxfp4_to_nvfp4.py` keeps the cast-specific cases (quantizer naming, `build_amax_map`, `apply_to_model`). - Validated on real DeepSeek-V4-Flash expert tensors (incl. the on-disk `float8_e8m0fnu` scale dtype): 23.5M blocks, 100% lossless, 0 error. - Generated a full NVFP4 checkpoint for DeepSeek-V4-Flash (43 layers, 256 routed experts) end-to-end: `[cast] lossless MXFP4->NVFP4 blocks: 8,657,043,456/8,657,043,456 (100.0000%)`. Output weights match an independently-produced reference cast byte-for-byte (`weight_scale`, `weight_scale_2`, packed nibbles modulo the harmless sign-of-zero). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (new opt-in flag; default export behavior unchanged; hoist re-exports through the existing example module) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ N/A (no new deps; shared numerics moved into the library rather than duplicated) - Did you write any new necessary tests?: ✅ (library numerics covered by `tests/unit/torch/quantization/test_numeric_utils.py`; end-to-end validated on a real DeepSeek-V4 checkpoint) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ (will run `/claude review`) ### Additional Information Mirrors and reuses #1372 (GPT-OSS MXFP4 → NVFP4 cast); the closed-form numerics are now shared via `modelopt.torch.quantization.utils.numeric_utils`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added `--cast_mxfp4_to_nvfp4` flag to perform a closed-form, mostly lossless MXFP4→NVFP4 conversion for routed-expert weights with aggregated lossless/block statistics. * **Documentation** * Updated DeepSeek V4 export instructions and README to document the new flag and clarify calibration behavior for activation scales. * **Chores** * Exposed shared numeric quantization utilities for MXFP4→NVFP4 casting. * **Tests** * Added and updated tests to validate the new numeric helpers and conversion behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
56c4af2333 |
feat(recipes): add kv_fp8_cast variants for partial-NVFP4 and weight-only PTQ recipes (#1652)
### What does this PR do?
Type of change: new feature (recipes)
Several `general/ptq` recipe families shipped a data-driven FP8 KV-cache
(`-kv_fp8`) variant but lacked the constant-amax `kv_fp8_cast` companion
that `fp8_default` and `nvfp4_default` already have. This PR adds the
missing cast variants so every KV-quantizing (and the weight-only)
family offers the calibration-free FP8 KV-cache option:
- `general/ptq/nvfp4_experts_only-kv_fp8_cast`
- `general/ptq/nvfp4_mlp_only-kv_fp8_cast`
- `general/ptq/nvfp4_omlp_only-kv_fp8_cast`
- `general/ptq/nvfp4_weight_only-kv_fp8_cast`
Each new recipe composes the exact same model-quant config as its
existing sibling and swaps the `kv_fp8` unit for the shared
`kv_fp8_cast` unit (constant-amax FP8 KV cache; no KV calibration
forward pass). The docs guide table/tree and the changelog are updated
to match.
### Usage
```bash
python examples/llm_ptq/hf_ptq.py \
--pyt_ckpt_path <model> \
--recipe general/ptq/nvfp4_mlp_only-kv_fp8_cast
```
### Testing
Extended the built-in PTQ smoke test
`tests/unit/recipe/test_loader.py::test_load_recipe_all_builtins` with
the four new recipe paths; all four load into a valid
`ModelOptPTQRecipe` with a populated `quantize` section.
```
$ python -m pytest tests/unit/recipe/test_loader.py tests/unit/recipe/test_presets.py -q
180 passed
```
`pre-commit` (including the `validate modelopt recipes` hook) passes on
all changed files.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (additive — only new recipe
files)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ (extended the builtin recipe
smoke test)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ (not yet)
### Additional Information
The two weight-only families were discussed for scope;
`nvfp4_weight_only` is included (it already names a KV mode, `kv_fp16`),
while `int4_blockwise_weight_only` is intentionally left untouched since
it carries no `-kv_` composition.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added four new NVFP4 PTQ (Post-Training Quantization) recipe variants:
experts-only, MLP-only, OMLP-only, and weight-only configurations.
* All new recipes include FP8 KV-cache cast mode support for improved
inference performance.
* **Documentation**
* Updated built-in recipes guide with new NVFP4 recipe options and
repository layout.
* **Tests**
* Expanded recipe loader test coverage for new recipe configurations.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
6f08731fe6 |
docs(eval skill): vLLM backend env vars + SLURM HF-cache/cpu_partition guidance (#1625)
### What does this PR do? Type of change: documentation Hardens the **evaluation** skill with operational guidance discovered while running AA-Index evals (NVFP4 Nemotron-3-Nano) on SLURM. Two themes: **1. vLLM backend env vars (original commits).** Some models need a model-card backend toggle that is an env var, not a CLI flag — e.g. NVFP4 MoE models like [NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4) need `VLLM_USE_FLASHINFER_MOE_FP4=1` + `VLLM_FLASHINFER_MOE_BACKEND=throughput`. These go in `deployment.env_vars` (with the `lit:` prefix), not the `vllm serve` command. **2. SLURM deploy/eval operational lessons (new commits).** - **`mount_home: false` always.** Some internal cluster templates default it `true`, which mounts the host `~/.cache`; where that is a symlink into a shared/networked filesystem, it dangles in the container and the vLLM `trust-remote-code` deploy dies with `FileNotFoundError: /root/.cache/huggingface` — **invisible to `--dry-run`**. - **HF cache:** mount the realpath of `~/.cache/huggingface` to `/hf-cache` and set `HF_HOME: lit:/hf-cache` for both stages. - **`execution.cpu_partition`:** on split GPU/CPU-partition clusters, the CPU-only MLflow auto-export job is rejected by the GPU partition and marks the whole task FAILED despite `EVAL_EXIT_CODE=0`. Set `cpu_partition` to route it correctly. - **Top-level `env_vars:`** for vars both stages need (`HF_TOKEN`, `HF_HOME`); `execution.env_vars` is unsupported and hard-errors. - **`lcr.md`:** `LOG_LEVEL=WARNING` to skip logging AA-LCR's ~120K-token inputs. ### Usage ```yaml execution: cpu_partition: <cpu-partition> # CPU-only auto-export job mounts: mount_home: false deployment: { <shared-fs>/<user>/.cache/huggingface: /hf-cache } evaluation: { <shared-fs>/<user>/.cache/huggingface: /hf-cache } env_vars: # shared by both stages HF_TOKEN: host:HF_TOKEN HF_HOME: lit:/hf-cache deployment: env_vars: VLLM_USE_FLASHINFER_MOE_FP4: lit:1 VLLM_FLASHINFER_MOE_BACKEND: lit:throughput ``` ### Testing Docs/skill-only. Validated end-to-end by running the full AA suite (GPQA + SciCode + AA-LCR) across 5 NVFP4 checkpoints on SLURM: dry-run + canary + full runs all SUCCESS with these settings. Pre-commit (yamlfmt, markdownlint) passed. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: ❌ 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Clarified where backend-related environment variables must be configured (deployment-level) and not embedded in execution commands. * Added guidance to mount the real HuggingFace cache path and set home-mounting to false for containerized SLURM runs. * Documented routing CPU-only jobs via CPU partitions and added a task-level LOG_LEVEL setting to reduce verbose long-context logging. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a548be5da6 |
[Eval skill] Add reasoning adapter_config to eval template (#1610)
### What does this PR do?
Type of change: documentation
Adds a reasoning-aware adapter (interceptor) config to the `evaluation`
skill's
default eval template and documents how to set `use_reasoning` per model
type.
- **`recipes/examples/example_eval.yaml`** — adds an `adapter_config`
block under
`evaluation.nemo_evaluator_config.target.api_endpoint` with
request/response
logging, `use_reasoning`, `log_failed_requests`, and a
`params_to_add.chat_template_kwargs` thinking block, with inline
comments on
which lines to flip for instruct vs. reasoning vs. hybrid models.
- **`SKILL.md`** — new "Reasoning adapter config (`use_reasoning`)"
subsection at
the end of Step 3:
- Instruct (non-reasoning) models → `use_reasoning: false` (no CoT trace
to
strip; drop the `chat_template_kwargs` thinking block).
- Reasoning models → `use_reasoning: true`, especially when the
deployment
command sets `--reasoning-parser`.
- Hybrid models (can run reasoning on or off) → always force thinking on
via
`chat_template_kwargs` and keep `use_reasoning: true` (highest scores).
### Usage
```bash
nel run --config recipes/examples/example_eval.yaml \
-o deployment.checkpoint_path=/path/to/checkpoint \
-o deployment.served_model_name=my-model
```
### Testing
Docs/template-only change. Ran `pre-commit` on both files (yamlfmt +
markdownlint
pass).
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- docs/template only
-->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ <!-- not yet run -->
### Additional Information
Note: there is a pre-existing inconsistency between Step 7 (which
references
`nemo_evaluator_config.config.target.api_endpoint.adapter_config`) and
the
template's `nemo_evaluator_config.target.api_endpoint` placement. This
PR follows
the template's existing path; reconciling Step 7 is left as a follow-up.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Added a guide for configuring adapter reasoning/chain-of-thought
behavior by model type (instruct, reasoning, hybrid), how reasoning
traces can be stripped before scoring, model-family template keyword
variants, and guidance to verify template kwargs and select appropriate
reasoning-effort settings.
* **Chores**
* Added an example configuration demonstrating enabling the reasoning
adapter with caching, limited request/response logging, failed-request
logging, forced “thinking” flags for hybrid models, and a
reasoning-effort option.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
88fd7ff958 |
eval skill: auto-detect predefined per-cluster execution configs (#1599)
### What does this PR do? Type of change: documentation Some NEL installs ship ready-made per-cluster execution configs as an `internal/slurm/<cluster>` group (via the optional `nemo_evaluator_launcher_internal` package). When present, the matching one pre-fills the cluster's `hostname` / `partition` / `gres` (and node-exclusivity), so the user only sets `account` / `output_dir` / `walltime` instead of hand-entering the hostname. - **`SKILL.md` Step 4** — adds a "check FIRST" step: run a discovery snippet that lists the available `cluster → hostname` pairs **from the installed package at runtime**; on a hostname match, use `defaults: - execution: internal/slurm/<cluster>` (replacing `slurm/default`) and drop the now-redundant `execution.hostname`. If the package isn't installed or nothing matches, fall back to `slurm/default` and fill the fields manually. - **`example_eval.yaml`** — a short, name-free comment on the `defaults` block pointing to that check. **Discovery-based by design (no internal data committed):** cluster names, hostnames, and accounts are read from the install at runtime and are **not** hardcoded here. The only internal references in the repo are the package name `nemo_evaluator_launcher_internal` and the generic `internal/slurm/<cluster>` group pattern. External users without the package degrade gracefully (import fails → `slurm/default`). ### Usage N/A — documentation / skill guidance only. ### Testing Ran the discovery snippet verbatim (lists `cluster → hostname` from the installed package) and confirmed `internal/slurm/<cluster>` resolves via a `--dry-run`. `pre-commit run` passes (markdownlint + YAML format). Grep-verified no cluster names/hostnames are hardcoded in the changed files. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (falls back to `slurm/default`) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (documentation) - Did you update Changelog?: N/A (skill docs) - Did you get Claude approval on this PR?: ❌ (pending) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added Step 4 guidance for discovering and using optional internal per-cluster SLURM execution configs: how to list available internal configs, select the matching cluster by switching the default execution, remove redundant hostname entries, and verify with a --dry-run. * Clarified that slurm/default is cluster-agnostic, recommend using internal per-cluster configs when present, and explained fallback behavior to manual hostname/account/output_dir entry. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
db9ea8f17e |
eval skill: parameterize external judge/user-sim endpoints via .env (#1591)
### What does this PR do?
Type of change: documentation
Several AA tasks call an external **judge / user-simulator / scoring
endpoint** whose `model_id` + `url` vary per user/site (HLE, AA-LCR,
Tau2 today — and the guidance is written as a general pattern for any
future such benchmark). Previously each recipe hardcoded `<...>`
placeholders that a user had to hand-edit in every config. This makes
them reusable **without committing any internal infrastructure**:
- **`recipes/env.example`**: add placeholders — `NS_JUDGE_URL`,
`HLE_JUDGE_MODEL_ID`, `LCR_JUDGE_MODEL_ID`, `TAU2_USER_MODEL_ID`,
`TAU2_JUDGER_MODEL_ID`, `TAU2_ENDPOINT_URL` — with the **recommended
model named** (GPT-4o / Qwen3 235B / gpt-oss-120B) but only **generic
hosts** (`https://<your-inference-host>/v1`). No internal hostnames or
gateway model-routing strings are committed; real values live in the
user's gitignored `.env`.
- **`recipes/tasks/aa/{hle,lcr,tau2_bench_telecom}.md`**: carry `<VAR>`
literal placeholders (named after the `.env` keys) that the skill
**substitutes as literal values** from the user's `.env`. These are
*config, not secrets*, so they are **not** exported — which avoids the
`${oc.env:...}` footgun (it silently fails unless the var was exported
with `set -a`). Only `api_key` (`INFERENCE_API_KEY`) stays an exported
env var read by the harness.
- **`SKILL.md`** (Step 5): instructs literal substitution from `.env`,
framed as a general pattern for any external-endpoint task.
The `/v1`-base (nemo-skills) vs full `/v1/chat/completions` (tau2-bench)
URL distinction is documented.
### Usage
N/A — documentation / skill-template only.
### Testing
`pre-commit run` passes (markdownlint). Verified no internal hostnames /
gateway model IDs are present in any committed file.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (documentation)
- Did you update Changelog?: N/A (skill docs)
- Did you get Claude approval on this PR?: ❌ (pending)
### Additional Information
Branched off latest `main` (includes #1583). Touches `lcr.md`, which
#1583 also edited (parallelism field) — different lines, rebased
cleanly.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Added comprehensive guidance for configuring external judge and
user-simulator endpoints in evaluation tasks
* Clarified best practices for substituting configuration values from
environment configuration files while keeping API keys as environment
variables
* Updated task documentation for HLE, LCR, and Tau2-Bench with improved
instructions for credential and endpoint handling
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
905259fbf5 |
Fix: use python3 in debugger server.sh (#1590)
### What does this PR do? Type of change: Bug fix `tools/debugger/server.sh` invokes `python -c "..."` inside `check_modelopt_local` (and `pip install -e .[dev]` in the fallback path). Containers that only ship `python3` (no `python` shim) cause the check to fail with `python: command not found`, exit 127. The server interprets this as "modelopt not editable-installed", runs `pip install`, the second check fails the same way, and the server aborts before it ever listens for commands. Switches both the inline import check and the install fallback to `python3` / `python3 -m pip`. ### Usage ```bash bash tools/debugger/server.sh # now starts cleanly on python3-only images ``` ### Testing Verified inside a container where `which python` returns nothing and `python3` is `/usr/bin/python3` (Python 3.12). With the old script, `server.sh` aborted at `check_modelopt_local`. With this fix, the check passes against an existing editable install and the server proceeds to wait for the client handshake. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `python3` is present on every image that previously had `python`. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — internal tooling shell script. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — internal debugger tooling, not user-facing. - Did you get Claude approval on this PR?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated debugger tooling to ensure consistent use of Python 3 for package validation and installation processes. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
38c78439f2 |
Add parallelism sizing reference to evaluation skill (#1583)
### What does this PR do? Type of change: documentation Adds `.claude/skills/evaluation/references/parallelism.md`, a reference for sizing both the **GPU topology** and the **request concurrency** of NEL evaluation runs, and wires the main `SKILL.md` to it. Pure docs for the agent-facing evaluation skill — no code or runtime behavior changes. The new reference covers: - **GPU topology (TP / DP / PP):** decision procedure (smallest-TP-then-max-DP), the TP-up triggers, and the TP/DP-split tradeoff for a fixed world size (e.g. all `TP×DP=8` factorizations on an 8-GPU node). - **Expert parallelism (EP):** `--enable-expert-parallel` is a boolean and `EP = TP × DP` (no direct EP-size flag); the DP-attention + EP-MoE dataflow (per-layer dispatch/combine all-to-all); when to enable vs not. - **Concurrency (`parallelism` / `--max-num-seqs`):** the request-count-vs-serving-capacity ceiling, KV-driven `--max-num-seqs`, and empirical tuning from vLLM startup logs + preemption. - **Gotcha + worked examples:** bit-width read from `config.json` (not the model name) sets the topology; includes the FP8-vs-4-bit Kimi example. `SKILL.md` gains pointers to the reference from the deployment-command, expert-parallel, evaluation-params, Step 4, and canary sections. ### Usage N/A — documentation only (agent skill guidance). ### Testing `pre-commit run` passes on both files (markdownlint included). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (documentation) - Did you update Changelog?: N/A (skill docs, not a library feature) - Did you get Claude approval on this PR?: ❌ (pending) ### Additional Information Commits are signed (`git commit -s -S`). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Expanded deployment guidance with clearer GPU topology, MoE/EP semantics, and bit‑width caveats. * Refined concurrency guidance: explicit rules for sizing top-level parallelism and computing max concurrent sequences, plus canary-run tuning to watch preemption and KV utilization. * Added task-level guidance for long‑context/judge‑bound jobs with recomputation advice and worked examples. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
2b7668d227 |
Fix MLflow auto-export in evaluation skill template (#1586)
### What does this PR do?
Type of change: documentation
Fixes two issues in the evaluation skill that caused completed eval runs
to silently **not** upload to MLflow (found while running a real GPQA
eval — the run finished successfully but never appeared in MLflow):
1. **Missing auto-export trigger.** NEL only auto-exports when
`execution.auto_export.destinations` is set; the `export.mlflow` block
alone just *configures* the exporter. The skill's shortcut said to "copy
the `export.mlflow` block" without calling out the trigger, so a config
could end up with the block but no upload.
2. **Interpolated export values break at submit time.** With auto-export
enabled, NEL resolves the `export.mlflow` block at submit time in a
scope that does **not** include `deployment` / `evaluation`, so
`${deployment.served_model_name}` / `${evaluation.*}` cross-references
fail hard with `Interpolation key '...' not found`. (`example_eval.yaml`
shipped this interpolated pattern *with* auto-export, so it would hit
the error.)
Changes:
- **`example_eval.yaml`**: mark the `auto_export.destinations` trigger
as REQUIRED; switch the `export.mlflow` block to **literal** values
(`CHANGEME-served-model-name` placeholders) with a comment explaining
why cross-refs break. `${oc.env:USER}` (env interpolation) is retained
since it resolves fine.
- **`SKILL.md`**: rewrite the MLflow shortcut step to require **both**
the trigger and literal export values, and warn against
`${deployment.*}` / `${evaluation.*}` cross-references.
### Usage
N/A — documentation/skill-template only.
### Testing
`pre-commit run` passes on both files (markdownlint + YAML format).
Verified the literal-export approach uploads correctly by exporting a
real run to MLflow.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (documentation)
- Did you update Changelog?: N/A (skill docs)
- Did you get Claude approval on this PR?: ❌ (pending)
### Additional Information
Separate from #1583 (parallelism reference); both touch the evaluation
skill but are independent.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Clarified MLflow auto-export: auto-export must be explicitly enabled
for uploads and the MLflow export block is ignored otherwise.
* Example config now uses literal experiment name, description, and tags
(no cross-scope interpolation) and shows concrete sampling values.
* Warned to keep sampling params consistent between tags and runtime
params and to set the tracking URI.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
ed0a4b175d |
Refactor evaluation skill: vLLM cross-check, MLflow defaults, walltime cap (#1561)
### What does this PR do?
Type of change: documentation
Refactor of the `evaluation` skill (`.claude/skills/evaluation/`) with
several substantive rule additions plus a significant compression pass.
**Skill rules added / tightened:**
- **Cross-check `recipes.vllm.ai` + HF model card before composing the
vLLM command.** Both sources matter; conflicts get surfaced to the user
instead of silently picked. WebFetch caveat triage (3 cases) for
JS-rendered variant tabs (this caveat bit the agent multiple times
during testing; the triage rules name the failure modes).
- **Single `deployment.command:` field** replaces separate
`tensor_parallel_size` / `data_parallel_size` / `extra_args` YAML
fields. NEL mounts the model at `/checkpoint`; Hydra interpolates
`${deployment.port}`.
- **vLLM defaults always included unless a recipe contradicts them**
(silence ≠ contradiction):
- `--max-num-batched-tokens 8192`
- `--enable-chunked-prefill`
- `--enable-expert-parallel` (MoE-only, detected via active-param suffix
or `num_experts`-like config field)
- `--max-num-seqs N` where `N = ceil(max_parallelism /
data_parallel_size)`, computed after Step 4 fills in `parallelism`.
- **Six-field evaluation params template:** `parallelism`,
`request_timeout`, `max_retries`, `max_new_tokens`, `temperature`,
`top_p`. No `top_k` / `presence_penalty` / `repetition_penalty` /
`min_p` at top level (task harnesses have their own defaults that
conflict). No per-task `max_new_tokens` overrides — one ceiling
everywhere.
- **`max_new_tokens` mandatory model-card lookup:** highest
card-recommended value wins; "card not yet checked + use generic
default" is explicitly forbidden (this was a real bug the user caught).
- **AA Index v2 (`recipes/tasks/aa/`)** is the default benchmark set for
quantized-checkpoint validation. "AA" / "Artificial Analysis" triggers
AA-only mode (no MMLU-Pro / AIME / LiveCodeBench unless asked).
- **MLflow auto-export on by default in shortcut path**, with
Hydra-interpolated `experiment_name` (`${USER}/${served_model_name}`),
`description` (embeds T / top_p / max_new_tokens), and `tags`
(string-coerced via single quotes for MLflow's tag type requirement).
Only `tracking_uri` needs user input.
- **Walltime capped at 4h** in generated configs (longer walltimes lower
scheduler priority → longer queue). Skill suggests three alternatives
for over-4h runs.
- **Example template** updated to the new conventions; `--max-model-len`
fallback bumped 32K → 131K to cover AA-LCR.
- **`tau2_bench_telecom` recipe:** `parallelism` left as `???` with
rate-limit guidance (canary 32–128, cap 512).
- **`model-card-research.md`:** output-length extraction is now a
mandatory, top-level checklist item with a cross-reference to the
SKILL.md rule.
- **`env.example`:** added `JUDGE_API_KEY` entry for AIME.
**Compression pass:** SKILL.md compressed ~670 → ~330 lines while
preserving every rule. Long prose collapsed into tables and bullets;
duplicated workflow checklist at the bottom removed.
### Usage
The skill is invoked when users ask to evaluate a model. The shortcut
path now produces a config like:
```yaml
deployment:
command: >-
vllm serve /checkpoint
--host 0.0.0.0
--port ${deployment.port}
--tensor-parallel-size 8
...
evaluation:
nemo_evaluator_config:
config:
params:
parallelism: ??? # Required — ask user
request_timeout: 3600
max_retries: 10
max_new_tokens: 81920 # from model card (highest)
temperature: 1.0
top_p: 0.95
export:
mlflow:
tracking_uri: ???
experiment_name: ${oc.env:USER}/${deployment.served_model_name}
description: '...'
tags:
framework: vllm
...
```
### Testing
Manually exercised the skill end-to-end on four models during refactor
(Qwen3.5-122B-A10B-FP8, Qwen3.6-35B-A3B-FP8, Kimi-K2.6, GLM-5.1-NVFP4)
to validate that the cross-check rules surface conflicts, the MLflow
defaults populate correctly, the AA suite excludes the right tasks, and
the walltime cap holds. Test configs are not included in this PR
(session artifacts at repo root, gitignored locally).
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — the skill is
editorial/operational; no runtime API changes. Existing configs continue
to work.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — skill documentation;
manually exercised on four models.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — skill documentation only.
- Did you get Claude approval on this PR?: ❌ — happy to run `/claude
review` if maintainers want it.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
- Rewrote the evaluation workflow with a stricter end-to-end checklist,
workspace-reuse guidance, AA-only shortcut, mandatory validation steps,
iterative finalization loop, and a hard 04:00:00 walltime cap
- Enforced vLLM as a single deployment command and stricter
generation/config constraints, including exact top-level params and
model-card–derived max_new_tokens
- Updated env var guidance (JUDGE_API_KEY / INFERENCE_API_KEY), MLflow
export metadata, and registry/auth preflight flow
- Added several AA-task recipes, added new benchmark recipes, and
removed or consolidated legacy task docs
<!-- review_stack_entry_start -->
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1561?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)
<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
a5bc6f8123 |
Add DATASET_COMBOS for grouped calibration datasets (#1508)
## Summary - Add ``DATASET_COMBOS`` to ``modelopt.torch.utils.dataset_utils`` — single ``--dataset`` tokens that fan out to several entries in ``SUPPORTED_DATASET_CONFIG``. The per-entry ``num_samples`` is split evenly across the members inside ``get_dataset_dataloader``. - Two initial combos: - ``cnn_nemotron_v2_mix`` → ``cnn_dailymail`` + ``nemotron-post-training-dataset-v2``. Replaces the hardcoded two-element fallback list in ``hf_ptq.py`` when ``--dataset`` is omitted. - ``nemotron-post-training-v3`` → the seven ``nvidia/Nemotron-*`` SFT datasets registered in #1498 (mirroring the upstream [`nemotron-post-training-v3` collection](https://huggingface.co/collections/nvidia/nemotron-post-training-v3)). - ``get_supported_datasets()`` now appends combo names so they show up in ``--dataset`` help. - ``hf_ptq.py``'s default ``--calib_size`` bumped from ``512`` to ``1024`` so the ``cnn_nemotron_v2_mix`` combo's even split preserves the previous total sample count (was 512 per-dataset × 2 datasets = 1024; now 1024 split → 512 per-dataset × 2). ``--calib_size`` now denotes the total calibration budget regardless of combo cardinality. - Reject mixing a combo with one of its member datasets in the same ``--dataset`` list (e.g. ``cnn_dailymail,cnn_nemotron_v2_mix``) — combo would otherwise double-sample the explicit member with a smaller per-member quota. - Reject combo names in ``get_dataset_samples``; combos are dataloader-only. The error message points callers to ``get_dataset_dataloader``. - Validate ``DATASET_COMBOS`` at import time: empty member lists, name collisions with ``SUPPORTED_DATASET_CONFIG``, and references to unknown datasets raise ``ValueError`` up front. ## Test plan End-to-end validated against ``/hf-local/Qwen/Qwen3.5-0.8B`` via ``get_dataset_dataloader`` on the actual streamed data, plus 5 new unit tests in ``TestDatasetCombosExpansion`` (all 44 tests in ``test_dataset_utils.py`` pass with no regressions). - [x] ``python -c "from modelopt.torch.utils.dataset_utils import DATASET_COMBOS, get_supported_datasets; assert 'cnn_nemotron_v2_mix' in get_supported_datasets() and 'nemotron-post-training-v3' in get_supported_datasets()"`` - [x] ``hf_ptq.py`` with no ``--dataset`` flag still calibrates on cnn_dailymail + nemotron-post-training-dataset-v2 with the same total sample count as before. - [x] ``--dataset nemotron-post-training-v3 --calib_size 1024`` allocates 146 per member across the seven Nemotron datasets; full 1022-sample dataloader builds without error. - [x] ``--dataset cnn_dailymail,nemotron-post-training-v3 --calib_size 256,1024`` composes correctly: 256 from cnn_dailymail (as a plain entry) plus the 7-way split from the combo. (The earlier ``cnn_dailymail,cnn_nemotron_v2_mix`` example is rejected by design since ``cnn_dailymail`` is a member of that combo.) - [x] ``--dataset cnn_dailymail,cnn_nemotron_v2_mix`` raises ``ValueError`` with a clear message. - [x] ``get_dataset_samples("cnn_nemotron_v2_mix", ...)`` raises ``ValueError`` pointing to ``get_dataset_dataloader``. - [x] Unit tests: ``pytest tests/unit/torch/utils/test_dataset_utils.py`` — 44 passed. ## Post-validation fix End-to-end testing surfaced that the original ``nemotron-sft-agentic-v2`` entry kept the two splits (``interactive_agent``, ``tool_calling``) that pyarrow's streaming JSON reader cannot parse, and excluded ``search`` which is the only clean split. Failures reproduce deterministically across cache wipes with ``force_redownload``, so they are content-level defects in the published JSONL files at the pinned revision, not local artifacts: - ``interactive_agent`` — heterogeneous schema (``Column(.../member_id/type) changed from string to array``) at JSONL row 4. - ``tool_calling`` — malformed JSON row in a later shard, fails at sample ~885 with ``Missing a closing quotation mark in string``. - ``search`` — streams cleanly (verified to 2500 samples). Commit ``10f3cfd`` corrects ``nemotron-sft-agentic-v2`` to use only ``search``, with an updated comment. The CHANGELOG calls this out as a separate bullet so the behavior change on a previously-released dataset entry from #1498 is discoverable. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Dataset combo support: a single dataset token can expand into multiple registered datasets with even sample splitting; predefined combos added (e.g., cnn_nemotron_v2_mix, nemotron-post-training-v3) and listed as supported. * **Updates** * Default dataset when none specified now uses cnn_nemotron_v2_mix. * Calibration size default increased from 512 to 1024. * **Bug Fixes** * nemotron-sft-agentic-v2 now uses only the deterministic "search" split to avoid streaming JSON errors. * **Tests** * Added coverage for combo expansion, splitting, overlap validation, and rejection behavior. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1508?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
5508c327fb |
[Quantization] MSE-calibrate every per-expert weight in fused-experts MoE (#1421)
### What does this PR do?
Type of change: Bug fix
Two-part fix for transformers 5.x fused-experts containers (Qwen3-MoE /
Qwen3.5-MoE / Mixtral / DeepSeek / Kimi-K2.x ...) where weight
quantizers live in `nn.ModuleList`s (`gate_up_proj_weight_quantizers`,
`down_proj_weight_quantizers`):
1. **Per-expert weight iteration for calibration.** Add
`_QuantFusedExperts.iter_weights_for_calibration` that yields per-expert
`(weight_slice, quantizer)` pairs for both projections. The base impl
uses singular `*_weight_quantizer` and silently skips fused-experts
modules, so weight-only calibration paths never reached per-expert
quantizers.
2. **`mse_calibrate` refactor.**
- Add `_bootstrap_uncalibrated_weight_quantizers` after `max_calibrate`
to populate `_amax` on quantizers the forward pass didn't reach (dead
MoE experts that received no calibration tokens). Runs the existing
calibrator on the weight slice surfaced by
`iter_weights_for_calibration`.
- Replace the singular-only `weight_attr_names` discovery +
`getattr`-by-name walk with an `iter_weights_for_calibration` walk done
inside each parent module's `enable_weight_access_and_writeback`
context, so MSE processes every per-expert quantizer (active and dead)
and remains FSDP-safe.
Without this, the export-time fallback in `_export_fused_experts`
derived separate gate/up amaxes from each half of the fused weight,
breaking the `gate==up` `weight_scale_2` invariant on dead experts.
Also includes:
- `_sanitize_generation_config_for_save` in `unified_export_hf` —
coerces `do_sample=True` when an upstream `generation_config.json` has
`top_k`/`top_p` set, so newer transformers' strict validate doesn't
block `save_pretrained`.
- Small companion plumbing in `moe_utils.py`, `tensor_quantizer.py`, and
`core_utils.py` to support the per-expert iteration and bootstrap path.
### Usage
```python
import modelopt.torch.quantization as mtq
from modelopt.recipe import load_config
# Recipe `nvfp4_experts_only_mse-kv_fp8_cast` (already on main) now correctly
# MSE-calibrates every per-expert weight quantizer in fused-experts MoE models.
cfg = load_config("general/ptq/nvfp4_experts_only_mse-kv_fp8_cast")
mtq.quantize(model, cfg, forward_loop=calibration_forward_loop)
```
### Testing
**Original validation — Qwen3.5-122B-A10B with
`nvfp4_experts_only_mse-fp8_cast_kv`:**
- **Before:** 1/12288 (layer 38 expert 69) `gate \!= up`; 0 weights
MSE-calibrated.
- **After:** 0/12288 mismatches; 24576 weights MSE-calibrated; ~4.2 min.
**End-to-end pipeline validation — Qwen3.5-35B-A3B (40 layers × 256
experts × 2 projections = 20,480 per-expert weight quantizers), TRT-LLM
1.3.0rc13 + transformers 5.6 docker, single B200:**
| | Path A (4-sample calib, deliberately undercalibrated) | Path B (zero
forward-pass tokens) |
|---|---|---|
| Per-expert weight quantizers calibrated | 20,480 / 20,480 | 20,480 /
20,480 |
| Missing `_amax` | 0 | 0 |
| All-zero `_amax` | 0 | 0 |
| `mtq.quantize` time | 25–34 s | 23 s |
- **Cross-path diff:** every per-expert weight amax matches
**bit-for-bit** between the two paths (`n=20480 exact=20480 diff=0
max_rel=0`). With 8/256 experts routed per token and 4 calib samples,
almost all experts are "dead" in Path A. Bootstrap fills them from
`max(|weight|)`, MSE searches deterministically from there → identical
to Path B which bootstraps everything.
- **Export to HF NVFP4 checkpoint** succeeded (~95 s, 22 GB checkpoint).
Resulting `generation_config.json` has `do_sample: true` (upstream had
`top_k=20` + `top_p=0.95` which would have failed strict validate).
- **TRT-LLM inference loaded the checkpoint and generated text:** `"Born
in north-east France, Soyer trained as a"` → `" tailor. Demonstrating
his craft at a young age, at 20 he moved to Paris at the requests of the
noble people of Picardy."` (coherent grammar; factually wrong as
expected with 4-sample calib, but no NaN/Inf in logits, no
scale-mismatch crash). 92 GB GPU memory used.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ <\!-- relies on existing
recipe-level integration coverage; verified end-to-end on
Qwen3.5-122B-A10B and Qwen3.5-35B-A3B + TRT-LLM 1.3.0rc13 -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ <\!-- will run \`/claude
review\` -->
### Additional Information
Follow-up to PR #1407 (MSE+FP8-cast-KV recipes). The recipe YAML files
landed there; this PR fixes the calibration codepath so the MSE recipes
actually exercise per-expert weight quantizers in fused-experts MoE
containers.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Fixed generation configuration validation for HuggingFace model
exports.
* Improved handling of quantization shape mismatches during expert
weight export.
* **New Features**
* Enhanced calibration process with automatic population of missing
expert quantizers.
* Added grouped quantizer synchronization for improved multi-expert
quantization.
* **Tests**
* Added regression tests for fused expert export and calibration
correctness.
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1421)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
|
||
|
|
1d796f9714 |
[Quantization] Fused Triton kernel for NVFP4 FP8 scale sweep search (#1387)
## Summary
- Replaces the 126-iteration Python sweep in `NVFP4MSECalibrator` with a
fused Triton kernel that loads each NVFP4 block once, evaluates all 126
valid FP8 E4M3 scale candidates in registers, and emits the per-block
``best_amax`` directly.
- Triton-backed `TritonNVFP4MSECalibrator` is the **default** for
`mse_calibrate(..., fp8_scale_sweep=True)`. Set
`MODELOPT_NVFP4_TRITON_SWEEP=0` to fall back to the reference for
debugging or numerics comparison.
- Microbenchmark: **~42x speedup** on a B300 over the reference
(`NVFP4MSECalibrator`) on a representative LLM weight (`8192x4096`, ~2M
NVFP4 blocks): `176.68 ms -> 4.23 ms`.
- End-to-end on Qwen3-8B PTQ (`hf_ptq.py --qformat nvfp4_mse`,
calib=128): **~6.7x faster mtq.quantize** with **identical global
weight-MSE** as the reference.
## End-to-end Qwen3-8B PTQ (B300)
> Note: ran on **Qwen3-8B** instead of Qwen3.5-9B because the docker's
transformers (4.57.3) doesn't yet recognize the `qwen3_5` (multimodal)
architecture and the model dir doesn't ship `modeling_*.py` for
`trust_remote_code`. Qwen3-8B is the same family / similar size and
gives a representative comparison.
Settings: `--calib_size 128 --calib_seq 512`, default `nvfp4_mse` config
(`NVFP4_W4A4_WEIGHT_MSE_FP8_SWEEP_CFG`). The first `nvfp4` run was
discarded as a warm-up for HF weight loading.
| qformat | `mtq.quantize` time | global weight MSE vs FP16 orig | notes
|
|---|---:|---:|---|
| `nvfp4` (no MSE search) | **3.46 s** | 6.363e-6 | max-calibration
baseline |
| `nvfp4_mse` (reference, slow) | **43.42 s** | 4.788e-6 |
`MODELOPT_NVFP4_TRITON_SWEEP=0` |
| `nvfp4_mse` (Triton, fast — this PR)| **6.51 s** | 4.788e-6 | default
in this PR |
- The Triton path's quantized weights produce **bit-identical global
weight MSE** to the reference (4.788e-6 vs 4.788e-6), validating the
kernel on a real model.
- Triton `nvfp4_mse` is **6.67x faster** than the reference `nvfp4_mse`,
and adds only **~3 s** over the no-MSE baseline (vs ~40 s for the
reference) — making the MSE-search option nearly free in practice.
- The MSE search itself reduces weight quantization error by **~25%** vs
plain max-calibration (6.363e-6 → 4.788e-6).
Per-layer MSE CSVs are produced by `tools/debugger/compare_mse_qwen.py`
(one row per Linear weight) for closer inspection if needed.
## Microbenchmark (B300, 8192x4096 weight, ~2M NVFP4 blocks)
```
reference NVFP4MSECalibrator: 176.68 ms
triton TritonNVFP4MSECalibrator: 4.23 ms
speedup: 41.8x
```
## Why this works
Each candidate is constructed as `valid_fp8_e4m3_value / 448`. With
`block_amax = global_amax * candidate`, the FP8 round-trip on the
per-block scale `block_amax / 6` (using `global_amax / 6` as the FP8
amax) is the **identity** — so the kernel can compute `scale = candidate
* global_amax / 6.0` inline and skip the FP8 cast. This keeps the kernel
runnable on any CUDA + Triton (no `tl.float8e4nv` requirement).
Because every candidate's per-block scale is just a rescaling of the
same input block, all 126 candidates can be evaluated against a single
`[BLOCKS_PER_PROGRAM, BLOCK_SIZE]` tile held in registers — replacing
126 weight-bandwidth passes with 1.
Two follow-on optimizations close the gap to the compute ceiling:
- `@triton.autotune` over `(BLOCKS_PER_PROGRAM, num_warps)` — the
original hand-picked default (BPP=4, num_warps=4) left ~4x on the table;
the best B300 config is `BPP=64, num_warps=8`.
- Drop the sign-handling `tl.where`: FP4 quant preserves sign, so `(w -
w_q)^2 == (|w| - |w_q|)^2` and the kernel works on `|w|` throughout (one
fewer where + negation per element per candidate).
## Files
- `modelopt/torch/kernels/quantization/gemm/nvfp4_fp8_sweep.py` — new
kernel + `nvfp4_fp8_scale_sweep` wrapper, autotuned.
- `modelopt/torch/kernels/quantization/gemm/__init__.py` — wire-in.
- `modelopt/torch/quantization/calib/mse.py` — new
`TritonNVFP4MSECalibrator(NVFP4MSECalibrator)`.
- `modelopt/torch/quantization/model_calib.py` — opt-out env var
(`MODELOPT_NVFP4_TRITON_SWEEP=0`); Triton path is default.
- `tests/gpu/torch/quantization/test_nvfp4_fp8_sweep_kernel.py` — 16 GPU
tests covering parity (across seeds, block counts, dtypes), input
validation, output round-trip, reset, and a wall-clock speedup report.
## Numerics
Bit-identical to the reference for typical block counts (`{4, 64, 1024}`
blocks, 3 seeds, fp32/fp16/bf16 — 14/15 microbenchmark tests bit-exact).
On multi-million-block weights an occasional adjacent-candidate
tie-break can differ at the fp32-noise level (observed **2 / 2,097,152
blocks** in the speedup test, per-block MSE within ~1e-7 relative). The
reference's CUDA `fake_e4m3fy` and the Triton inline math have slightly
different op ordering, which lets nearly-tied candidates flip. The
speedup test asserts the worst per-block MSE gap is `< 1e-5` relative on
differing blocks — both choices are valid argmins; the resulting
quantized weights are equally good. The Qwen3-8B end-to-end run confirms
this: aggregate weight MSE matches the reference exactly at the
displayed precision.
## Test plan
- [x] `pytest
tests/gpu/torch/quantization/test_nvfp4_fp8_sweep_kernel.py -v` — 16/16
pass on B300
- [x] `pytest
tests/gpu/torch/quantization/test_nvfp4_static_quantizer_cuda.py -v` —
existing NVFP4 tests still pass (9/9)
- [x] End-to-end PTQ on Qwen3-8B with `--qformat nvfp4_mse`: 6.67x
speedup, identical weight MSE to reference
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
- Added a Triton-based fused NVFP4 FP8 scale-sweep for faster per-block
scale selection.
- Exposed a device-aware FP8 scale candidate generator and a GPU sweep
API used by the calibrator.
- Calibrator now uses the Triton fast-path by default with env var
opt-out (MODELOPT_NVFP4_TRITON_SWEEP).
* **Bug Fixes / Behavior**
- One-shot fast-path enforcement in the NVFP4 calibrator; reset restores
reuse.
* **Tests**
- Added comprehensive GPU tests validating parity, input validation,
dispatch behavior, one-shot semantics, and speed benchmarks.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
|
||
|
|
e2d29c869b |
[NVBug 6143871] Fix awq_lite uncalibrated branch leaving input_quantizer disabled (#1410)
## Summary `awq_lite.setup()` disables `module.input_quantizer` at the start of search. The calibrated branch re-enables it inside `postprocess()`, but the uncalibrated branch (no cache-pass tokens, e.g. an MoE expert that never gets routed) never did. Worse, for experts that had cache hits but missed the search pass, the per-channel `_amax` left over from `max_calibrate` during cache mode tripped `preprocess_linear_fusion`'s `numel == 1` assertion and prevented `_export_quantized_weight` from emitting a per-tensor `input_scale`. Result: per-expert `input_scale` was missing in the exported HF checkpoint, and TRT-LLM `CutlassFusedMoE` crashed on load with `KeyError: '<idx>.w1.input_scale'` for any expert that did not see enough calibration tokens (e.g. Qwen3-30B-A3B + `nvfp4_awq` from the bug report). ## Fix Mirror the calibrated `postprocess()` path in `modelopt/torch/quantization/model_calib.py`: collapse any per-channel `_amax` to scalar (axis=None) and re-enable the `input_quantizer`. ## Test plan - [x] Added regression test `test_awq_lite_uncalibrated_linear_keeps_input_quantizer_enabled` using `NVFP4_AWQ_LITE_CFG` with a two-branch model where only one branch is exercised; verifies the uncalibrated linear's `input_quantizer` remains enabled after `mtq.quantize`. - [x] End-to-end pipeline test on a tiny synthetic Qwen3-MoE (8 experts, top-1 routing) confirms all 48 expert `input_scale` keys are present in the exported state_dict (vs. multiple missing pre-fix). - [x] Manual repro of the bug command (Qwen3-30B-A3B, `--qformat nvfp4_awq`) confirmed the missing `input_scale` keys before the fix. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed AWQ-Lite quantization for uncalibrated modules so export/preprocessing invariants are preserved even when calibration/parameter updates are skipped. * **Tests** * Added regression tests: one verifies uncalibrated linear modules keep their input quantizer enabled after quantization; another verifies weight-only AWQ-Lite cases keep the input quantizer disabled as expected. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
6a3b6b8329 |
[Recipes][LLM PTQ] Add nvfp4 MSE+FP8-cast-KV recipes (experts_only / mlp_only) + --recipe in example scripts (#1407)
## Summary
- Adds two PTQ recipes that combine **experts/MLP-only NVFP4 W4A4** with
**MSE FP8 scale-sweep weight calibration** and **FP8 KV cache with
`use_constant_amax: true`** (skips KV calibration; matches the
`nvfp4_default-fp8_cast_kv` contract):
- `modelopt_recipes/general/ptq/nvfp4_experts_only_mse-fp8_cast_kv.yaml`
— applies to `*mlp.experts*` / `*block_sparse_moe*` only.
- `modelopt_recipes/general/ptq/nvfp4_mlp_only_mse-fp8_cast_kv.yaml` —
applies to all `*mlp*` / `*block_sparse_moe*` (dense MLP + MoE).
- Threads a new `--recipe` flag through
`examples/llm_ptq/scripts/parser.sh` and `huggingface_example.sh`.
Either `--quant` or `--recipe` is required; passing **both errors out**.
Recipe names are not validated in the script — `hf_ptq.py` is the source
of truth.
- Drops the bash-side `qformat` whitelist case-statement in
`huggingface_example.sh` for the same reason.
## Files
**New recipes (`modelopt_recipes/general/ptq/`):**
- `nvfp4_experts_only_mse-fp8_cast_kv.yaml` — same patterns as
`nvfp4_experts_only-fp8_kv.yaml`.
- `nvfp4_mlp_only_mse-fp8_cast_kv.yaml` — same patterns as
`nvfp4_mlp_only-fp8_kv.yaml`.
Both differ from their `_kv` siblings by:
- `algorithm: max` → `{ method: mse, fp8_scale_sweep: true, layerwise:
false }`
- All targeted **weight quantizers** switch `type: dynamic` → `type:
static` (otherwise `mse_calibrate` skips them: only static block-quant
weight quantizers are recognized for the FP8 sweep — see
`model_calib.py:369-374`).
- Input quantizers stay dynamic.
- KV bmm adds `use_constant_amax: true` (the `_cast_kv` flavor).
**Scripts (`examples/llm_ptq/scripts/`):**
- `parser.sh` — adds `--recipe` long-option, default `RECIPE=""`,
validates one-of-{`--quant`, `--recipe`} and not-both.
- `huggingface_example.sh` — when `RECIPE` is set, derives `MODEL_NAME`
from the recipe basename, passes `--recipe=…` to `hf_ptq.py` instead of
`--qformat=…`, and exits after export with a TRT-LLM deployment hint
(recipes can produce arbitrary configs that the script's downstream
`run_tensorrt_llm.py` path doesn't know how to handle generically).
Drops the `qformat` whitelist; defers to `hf_ptq.py`.
## Behavior
```
# Errors with: "Cannot specify both --quant and --recipe; pick one."
bash huggingface_example.sh --model=... --quant=nvfp4 --recipe=... --tasks=quant
# Errors with usage if neither is given
bash huggingface_example.sh --model=... --tasks=quant
# Both of these are now accepted; --recipe is forwarded verbatim to hf_ptq.py
bash huggingface_example.sh --model=... --quant=nvfp4 --tasks=quant
bash huggingface_example.sh --model=... --recipe=general/ptq/nvfp4_experts_only_mse-fp8_cast_kv --tasks=quant
bash huggingface_example.sh --model=... --recipe=general/ptq/nvfp4_mlp_only_mse-fp8_cast_kv --tasks=quant
```
## Test plan
- [x] `experts_only_mse-fp8_cast_kv` loads via
`modelopt.recipe.load_recipe(...)` and produces the expected algorithm +
per-pattern `quant_cfg` (verified in a working env: `algorithm ==
{'method': 'mse', 'fp8_scale_sweep': True, 'layerwise': False}`; expert
weight quantizers `type: static`; KV bmm has `use_constant_amax: True`).
- [x] Parser sanity: 4 flag combinations (both, neither, only `--quant`,
only `--recipe`) all behave as designed.
## Note
Pre-commit hook `check-modelopt-recipes` was skipped on both commits
because the local conda env has a broken `torchvision` install
(`AttributeError: partially initialized module 'torchvision' has no
attribute 'extension'`) that prevents `from modelopt.recipe.loader
import load_recipe`. The `experts_only` recipe was validated
independently by running `tools/precommit/check_modelopt_recipes.py` in
a working environment (exits 0); the `mlp_only` one is the same shape
with a different glob.
Rebased onto `main` from #1391 (which targeted
`chenjiel/nvfp4-fp8-sweep-triton`). The diff is scoped to the recipes +
script wiring; no kernel/sweep changes are included here.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added recipe-based quantization as an alternative to format-based
quantization with a new `--recipe` CLI option.
* Added two new quantization recipes for targeted layer optimization:
one for expert-layer-only quantization and one for MLP-layer-only
quantization, both featuring NVFP4 and FP8 KV-cache optimization.
* **Configuration**
* `--quant` and `--recipe` options are now mutually exclusive; specify
one to configure quantization behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
|
||
|
|
097293b44d |
[Quantization] Saturate NVFP4 export FP8 scale cast to avoid NaN (#1397)
## Summary - Saturates `per_block_scale * 448 / per_block_scale_max` to ≤ 448 before the `to(torch.float8_e4m3fn)` cast in `NVFP4QTensor.get_weights_scaling_factor_from_quantizer`. - Adds a regression test that reproduces the NaN byte without the clamp. ## Why When `_amax` contains a zero entry (e.g. an all-zero weight block left untouched by max calibration), the existing `per_block_scale[per_block_scale == 0] = 1.0` safety net drives the pre-cast value to `1.0 * 448 / (global_amax / 6)`. `fp8_e4m3fn` has no Inf — anything `≥ 480` rounds to NaN — so a 0x7F byte slips into the exported `weight_scale`. This was observed in a saved Kimi-K2.6-NVFP4-MSE checkpoint at `language_model.model.layers.1.mlp.experts.21.down_proj.weight_scale[4001, 18]`. The MSE FP8 sweep itself never produces zero per-block amax (it always emits at least `c[0] * global_amax`), but any export path where `_amax` ends up zero — including pure max calibration — hits the bug. With the clamp the byte saturates to `0x7E` (= 448, fp8 max finite) and dequantization is unaffected: the FP4 nibbles for an all-zero block are all 0, so `0 × 448 × weight_scale_2 = 0` regardless of the stored fp8 scale. For non-degenerate blocks the clamp is a no-op since `per_block_amax ≤ global_amax` already bounds the pre-cast value at 448. ## Test plan - [x] New regression test `test_export_fp8_scale_no_nan_for_zero_amax_block` fails on `main`'s export code (reproduces the 0x7F NaN byte) and passes with the clamp. - [x] Existing tests in `tests/gpu/torch/quantization/test_nvfp4_static_quantizer_cuda.py` still pass (10/10). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved numerical stability in FP8 quantization scaling by preventing overflow and NaN conditions * Enhanced handling of edge cases in quantization processing for zero-weight blocks <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
b5df2a5a3c |
Update llm_ptq requirements.txt (#1394)
### What does this PR do? Type of change: Dependency update compressed_tensors 0.15 is not compatible with our current quant implementation for Kimi K2.5, K2.6 ### Testing Unittest ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Broadened the `compressed-tensors` dependency constraint to allow a wider range of compatible versions. * Removed the `rouge_score` dependency from the example requirements. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
1d21ab9e29 |
[DeepSeek] Default to top-k calibration with peer-max input amax sync (#1380)
## Summary - DeepSeek PTQ (`examples/deepseek/ptq.py`) now defaults to native top-k routing during MoE calibration. The previous all-tokens-to-all-experts path (`CalibMoe`) is preserved behind a new `--calib_all_experts` flag. - After `mtq.quantize`, `fixup_moe_expert_amax` syncs every expert's `input_quantizer.amax` (w1/w2/w3) to the per-layer global peer max via `dist.all_reduce(MAX)` across EP ranks. `weight_quantizer.amax` stays per-expert; any uncalibrated expert is filled by computing amax over the dequantized FP8 weight. - `mtq.print_quant_summary` is now also written to `<output_path>/.quant_summary.txt`, mirroring `llm_ptq/hf_ptq.py`. ## Why Forcing all tokens through every expert doubled calibration time and inflated `input_quantizer.amax` for cold-routing experts with outliers they never see at inference. The new flow matches the inference distribution, runs roughly 2x faster, and mirrors the `layer_sync_moe_local_experts_amax` semantics that mtq runs automatically for `QuantSequentialMLP`-derived MoEs. ## Validation (DeepSeek-V3.2-Exp, MP=8, NVFP4_DEFAULT_CFG) Compared `_amax_baseline` (CalibMoe) vs `_amax_synced` (new default): - All 44,544 expert weight amaxes bit-identical. - Attention, shared experts, gate: identical. - Expert `w1.input` and `w3.input` (shared MoE block input): identical. - Expert `w2.input` (post-SiLU gated, expert-specific): synced to layer-wide peer max — 99.3% are larger than baseline (median 11.4x) since peer-max captures the worst-case outlier from any expert in the layer; 0.7% are smaller. This is the same trade-off `set_expert_quantizer_amax` makes for HF MoEs in `unified_export_hf.py`. ## Test plan - [x] DeepSeek-V3.2-Exp MP8 PTQ with default flags — completes in ~7 min (vs ~27 min with CalibMoe), produces `_amax_synced/` consistent with the comparison above. - [x] DeepSeek-V3.2-Exp MP8 PTQ with `--calib_all_experts` — produces `_amax_baseline/` identical (other than rounding) to the prior `CalibMoe`-default behavior. - [x] `.quant_summary.txt` written under `output_path` on rank 0. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a `--calib_all_experts` option to enable an alternate PTQ calibration mode; default remains top-k routing with a post-calibration per-layer peer-max synchronization and a compute fallback for uncalibrated experts. * **Documentation** * Clarified default and alternate calibration behaviors and added note about generation of a `.quant_summary.txt` summary file. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
50706d1750 |
Add closed-form MXFP4 -> NVFP4 weight cast (--cast_mxfp4_to_nvfp4) (#1372)
## Summary - New `--cast_mxfp4_to_nvfp4` flag in `hf_ptq.py` (and `huggingface_example.sh`) that converts an MXFP4 source checkpoint (e.g. `openai/gpt-oss-20b`) into an NVFP4 export with **bit-exact** weight reconstruction for the in-range blocks. - The cast pins NVFP4's `scale_2 = 2^m` (where `m = k_max − 8`) and `_amax = 6·2^k_j` per NVFP4 block, both read from the source `*_scales`. The resulting per-block scale `2^(k_j − m)` is exactly representable in E4M3, so `round_to_E2M1(value / 2^k_j)` yields the original MXFP4 nibble verbatim. For out-of-range blocks (`k_max − k_j > 17`) the per-block amax falls back to data-derived `max(|w_block|)`, which keeps the post-E4M3-clamp scale close to the block's actual magnitude. ## Verification End-to-end on `openai/gpt-oss-20b` with `--qformat=nvfp4_mlp_only --cast_mxfp4_to_nvfp4`: ``` [cast_mxfp4_to_nvfp4] overrode 48/48 weight quantizers [cast_mxfp4_to_nvfp4] lossless layers: 48/48 (100.00%) [cast_mxfp4_to_nvfp4] lossless blocks: 597196800/597196800 (100.0000%) ``` End-to-end on `openai/gpt-oss-120b` with the same flags (4×B200, `--use_seq_device_map --gpu_max_mem_percentage 0.5 --calib_batch_size 4`): ``` [cast_mxfp4_to_nvfp4] overrode 72/72 weight quantizers [cast_mxfp4_to_nvfp4] lossless layers: 67/72 (93.06%) [cast_mxfp4_to_nvfp4] lossless blocks: 3583179586/3583180800 (100.0000%) ``` Five layers fall into the OOR regime (block-spread > 17); the remaining 1,214 OOR blocks use the data-derived per-block amax fallback. Block-level losslessness is **99.99996%** end-to-end. Per-tensor MSE between MXFP4 source dequant and NVFP4 export dequant (~19B elements): | Metric | Without cast | With cast | |---|---|---| | Per-tensor SNR | ~26.4 dB (FP4 noise floor) | **∞ (every tensor)** | | Total RMSE | 8.67e−02 | **0** | | max\|err\| | up to 8.0e+1 | **0** | ## Modelopt-side enablers - `max_calibrate` auto-promotes static-block NVFP4 weight quantizers to `NVFP4StaticQuantizer` at the end of calibration. - `static_blockwise_fp4_fake_quant` kernel accepts N-D inputs (was 2D-only), unblocking MoE expert weights of shape `(E, F, K)`. - BMM-experts NVFP4 export routes through `get_weights_scaling_factor_from_quantizer` for static-mode quantizers, so the pinned `_amax` is actually consumed. - `set_expert_quantizer_amax` scalar-reduces per-quantizer amax before stacking, supporting per-block (vs scalar) static-mode amax. ## Test plan - [x] Unit tests at `tests/examples/llm_ptq/test_cast_mxfp4_to_nvfp4.py` (15 tests, all passing) cover: scalar/global-amax math, per-block hybrid (in-range closed-form vs OOR data-derived), shape preservation, key collection, and end-to-end `build_amax_map` against a synthetic safetensors checkpoint. - [x] End-to-end PTQ → export on `openai/gpt-oss-20b` (`nvfp4_mlp_only` qformat) with `--cast_mxfp4_to_nvfp4` succeeds; export takes ~21 s. 100% lossless cast (48/48 layers, 597,196,800 / 597,196,800 blocks). - [x] End-to-end PTQ → export on `openai/gpt-oss-120b` (4×B200, `nvfp4_mlp_only`, `--use_seq_device_map --gpu_max_mem_percentage 0.5 --calib_batch_size 4`). 67/72 layers fully lossless; 99.99996% block-level losslessness (3,583,179,586 / 3,583,180,800). - [x] TRT-LLM serving validation (TRT-LLM 1.3.0rc11, B200) on both exported NVFP4 checkpoints via `examples/llm_ptq/run_tensorrt_llm.py`: - **20b** (TP=1): 18.3 GB GPU memory; coherent generation. Sample: *"Quantum computing is poised to revolutionize data analysis. However, its potential is currently limited by quantum hardware constraints, including error rates, qubit lifetimes, and lack of fault tolerance…"* - **120b** (TP=4): 36.4 GB / GPU; coherent generation. Sample: *"Quantum computing is poised to revolutionize data storage and processing. These rare earth-based systems could serve as robust qubits; resistant to environmental decoherence…"* - [x] MSE comparison script (run separately during development) confirms per-tensor SNR=∞ across all 48 MoE expert tensors. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a MXFP4→NVFP4 weight-format cast utility and a CLI flag to enable it; helper scripts updated to expose the option. * **Bug Fixes** * Fixed static NVFP4 export for expert weights. * Improved collection/handling of quantizer amax values to avoid shape issues. * Generalized FP4 kernel to accept flexible tensor dimensionality. * Ensured static-block NVFP4 promotion during calibration. * **Tests** * Added comprehensive tests for the conversion workflow, helpers, and end-to-end application. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
168cd828c1 |
Add qwen3 moe experts only test (#1274)
## Summary - Add unit test for Qwen3 MoE HF export with `NVFP4_EXPERTS_ONLY_CFG` quantization config - Verifies that `hf_quant_config.json` correctly reports `quant_algo: NVFP4` and that non-expert modules (`self_attn`, `lm_head`) appear in `exclude_modules` while routed expert layers (`mlp.experts.*`) do not - Reference: https://huggingface.co/nvidia/Qwen3.5-397B-A17B-NVFP4/blob/main/hf_quant_config.json Type of change: New tests ### Known issue On `transformers>=5.0`, fused MoE experts (`_QuantFusedExperts`) are not recognized by `get_quant_config`, causing `quant_algo=None` in the exported config. This test currently **fails** on transformers 5.x and is intended to be fixed by a follow-up change. ## Testing - **transformers 4.57.6**: PASSED - **transformers 5.5.4**: FAILED (`quant_algo` is `None` due to fused expert export gap) ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added GPU test coverage for exporting Qwen3 Mixture-of-Experts models with NVFP4 quantization. * Verifies the exported checkpoint records the NVFP4 quantization algorithm and that module exclusion patterns correctly exclude attention and LM head components while not excluding routed expert paths. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
3ad4f4f093 |
[Fix] Re-expand target_input on OOM in get_max_batch_size (#1374)
## Summary - `get_max_batch_size` halved `target_data_batch` on `torch.cuda.OutOfMemoryError` but never rebuilt `target_input`, so each retry re-fed the same too-large tensor — the retry loop was effectively a no-op. - Refactor the expand logic into an `_expand_to(batch)` helper, rebuild `target_input` after halving, and call `torch.cuda.empty_cache()` between attempts. ## Test plan - [x] New unit test `test_get_max_batch_size_oom_retry_shrinks_input` mocks `torch.cuda.*` and asserts the second retry receives the halved tensor (shapes seen: `[1, 10, 5]`, regulated result `4`). - [x] `pytest tests/unit/torch/utils/test_dataset_utils.py` — 14/14 pass (skipping the network-only minipile test). - [x] `pre-commit` (ruff, mypy, bandit, license headers) clean on commit. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Enhanced GPU memory management during batch size detection. When out-of-memory errors occur during the initial probing phase, the system now properly adapts input tensors to smaller batch sizes and clears GPU cache before retry attempts, resulting in more reliable recovery and stable batch sizing across diverse hardware environments. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
47a33db9b6 |
[NVBUG: 6103846] Fix nvfp4_awq export for uncalibrated MoE experts (#1354)
## Summary - NVBug: [6103846](https://nvbugspro.nvidia.com/bug/6103846) — `Qwen3-30B-A3B nvfp4_awq` quantization fails at export with `AssertionError: Modules have different quantization formats`. - Root cause: in `model_calib.awq_lite`, MoE experts that end up disabled (NaN in act/weight scales, or no search-pass tokens) get `max_calibrate`-d but no `pre_quant_scale`. `get_quantization_format` then returns `nvfp4` for those experts while siblings stay `nvfp4_awq`. `unified_export_hf.requantize_resmooth_fused_llm_layers` groups all 128 experts of each linear name (gate_proj/down_proj/up_proj) and calls `preprocess_linear_fusion(..., resmooth_only=True)`, which asserts uniform format → fires for any single mismatched expert. - Fix: unify the disabled-expert paths in the awq_lite postprocess loop so any expert with `is_enabled == False` (no cache hits, NaN scales, or no search-pass tokens) receives `max_calibrate` + a neutral all-ones `pre_quant_scale`, matching the existing behavior for `num_cache_steps == 0`. Emit a warning so users notice that calibration coverage is incomplete and accuracy may degrade. ## Test plan - [x] `pytest tests/unit/torch/quantization/test_calib.py -k 'awq'` → 5 passed - [x] End-to-end on `Qwen/Qwen3-30B-A3B` with `NVFP4_AWQ_LITE_CFG` and a small calib set that leaves many experts uncalibrated: - All 6144 gate_proj/up_proj/down_proj expert linears report `nvfp4_awq` (no mismatch) - `export_hf_checkpoint` succeeds with no `AssertionError` - The new "Forcing pre_quant_scale=1 ... may degrade accuracy" warning fires for each affected expert - [x] Re-run via `examples/llm_ptq/hf_ptq.py` with the bug-report CLI (cnn_dailymail, batch_size=8, calib_size=64 — scaled down from 512 to fit budget) on B200: - 36 "the second time did not forward data through ..experts.X.{gate,up,down}_proj" warnings — i.e. the exact bug-triggering condition from the original NVBug log naturally reproduces - 2058 "Forcing pre_quant_scale=1" warnings — fix path activates for uncalibrated/disabled experts - 0 `AssertionError`s — export completes - `Quantized model exported to: /tmp/test_plan_qwen3-30b-a3b-nvfp4_awq` and post-PTQ generation works --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
fda0899e40 |
feat(recipes): add KV cache cast variants (fp8_cast / nvfp4_cast) (#1334)
## Summary
- Adds three built-in PTQ recipes that express the KV-cache *cast*
variants directly in YAML, using the existing `use_constant_amax: true`
quantizer field. These are recipe equivalents of
`--kv_cache_qformat=fp8_cast` / `nvfp4_cast`:
- `general/ptq/fp8_default-fp8_cast_kv`
- `general/ptq/nvfp4_default-fp8_cast_kv`
- `general/ptq/nvfp4_default-nvfp4_cast_kv`
- Makes `--recipe` authoritative in `examples/llm_ptq/hf_ptq.py`: the
post-hoc `_set_kv_cache_constant_amax` override now only runs when
`--recipe is None`, so a recipe YAML fully determines KV-cache config
instead of being silently overridden by the default
`--kv_cache_qformat=fp8_cast`. Updated help text on both flags.
- Extends the recipe loader smoke test to cover the three new recipes.
## Motivation
Before this change, the cast variants lived only in argparse
(`_KV_CAST_FORMATS = {"fp8_cast", "nvfp4_cast"}`) and were layered on
top of any recipe-loaded config. That meant `--recipe
nvfp4_default-fp8_kv` would silently become a cast recipe due to the
`--kv_cache_qformat` default. Now the recipe is self-contained: its YAML
either sets `use_constant_amax: true` on the `*[kv]_bmm_quantizer` entry
(cast) or doesn't (data-driven calibration).
## Test plan
- [x] `pytest tests/unit/recipe/test_loader.py` — all 24 tests pass,
including the three new parametrized recipes.
- [x] Verified each new recipe round-trips through `load_recipe()` with
`use_constant_amax: True` surviving Pydantic validation on the KV entry.
- [x] End-to-end run on `/models/Qwen/Qwen3-8B` (RTX 6000 Ada, 4
samples, seq_len=128) for all three new recipes:
- After `mtq.quantize(model, recipe.quantize.model_dump(),
forward_loop=...)`, all 72 `k_bmm_quantizer` / `v_bmm_quantizer` modules
have `_use_constant_amax=True` and `_get_amax()` returns `448.0` (FP8
E4M3 max).
- Weight quantizers still calibrate from data normally (sample amax
values: q_proj=0.5508, k_proj=0.6250, v_proj=0.1689, o_proj=0.7266).
- [x] Verified the `--recipe` authoritative behavior change:
- Non-cast recipe + default `--kv_cache_qformat=fp8_cast` → KV entry
does NOT get `use_constant_amax` (no silent override).
- Cast recipe + contradictory `--kv_cache_qformat=fp8` → KV entry keeps
`use_constant_amax=True` (recipe wins).
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Fixed CLI to respect KV cache quantization settings from recipe YAML
instead of overriding them.
* **New Features**
* Added three new post-training quantization recipe configurations for
FP8 and NVFP4 with optimized KV cache handling.
* **Documentation**
* Enhanced CLI help text for recipe and KV cache quantization options
with configuration examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
|
||
|
|
01788bb007 |
Deprecate Mllama support in llm_ptq/vlm_ptq examples (#1332)
## Summary - Removes Mllama (Llama 3.2 Vision) model-type branches from the `llm_ptq` example (`hf_ptq.py`, `example_utils.py`) and drops the now-unused `MllamaImageProcessor` wrapper from `modelopt/torch/utils/`. - Drops the legacy `MllamaImageProcessor` path in `modelopt/torch/utils/vlm_dataset_utils.py`; the generic HF ProcessorMixin path handles the remaining cases. - Adds a CHANGELOG entry under 0.44 Backward Breaking Changes. ## Test plan - [x] CI lint / unit tests pass - [x] Smoke-run ``examples/llm_ptq/scripts/huggingface_example.sh --model <llm> --quant fp8`` (text-only path, non-mllama) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Removed Mllama (Llama 3.2 Vision) support from quantization examples. This includes removal of dedicated image processor implementation, specialized model handling, and related calibration logic. * Updated VLM image-text calibration guidance to use `--calib_with_images` flag with other supported VLMs instead of Mllama-specific processing paths. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
04fcf24227 |
Fix LLM deploy test failure by defaulting expert parallelism to 1 (#1273)
### What does this PR do? Type of change: Bug fix Fixes TRT-LLM DeepEP kernel failures during LLM deployment on unsupported GPUs (e.g. Blackwell SM 12.0) by defaulting expert parallelism (`ep`) to 1 instead of auto-setting it to the GPU count for MoE models. Previously, when the model config contained expert-related keys, `ep` was automatically set to `torch.cuda.device_count()`, which triggered DeepEP kernel failures on GPUs that don't support it. Now `ep` defaults to 1 while still enabling attention data parallelism for MoE models. Expert parallelism can be enabled explicitly by the caller when the environment is known to support it. ### Testing - [x] Verified that the `llm_ptq` test passes with this fix on Blackwell GPUs. - [x] 2-gpu CI test triggered: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24495054531/job/71588037727 ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
d45219b390 |
Fix debugger server failing to detect editable-installed modelopt (#1270)
## Summary - Removed `PYTHONPATH="" python -I` override in `check_modelopt_local()` so the PYTHONPATH validation uses the actual environment instead of an isolated one - Moved the workdir log line earlier in `server.sh` for better debugging visibility ## Test plan - [x] Start `server.sh` inside a Docker container and verify it correctly detects editable-installed modelopt - [x] Confirm the workdir is logged before the modelopt check runs 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Server startup now displays the configured work directory earlier in the initialization process, providing improved visibility of the active directory during server launch. * Simplified the modelopt validation check during server initialization while maintaining the same validation behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
7c8557158d |
Add job cancellation support to the debugger command relay (#1262)
## Summary - Add a `cancel` subcommand to the client that terminates the currently running command on the server - Server now runs commands in the background with PID tracking, enabling cancellation mid-execution - Client-side timeouts automatically cancel the running command on the server (previously the server process was left running) - Hardened against race conditions through 4 rounds of adversarial review (15 fixes total) ### Key changes **server.sh:** - Commands run in background with PID tracked in `$RELAY_DIR/running` (atomic tmp+mv write) - Cancel detection loop checks for `$RELAY_DIR/cancel` file with cmd_id verification - SIGTERM with 5s grace period, then SIGKILL escalation for stuck processes - `.exit` file written before `running` marker removed (ordering guarantee) - `set -e`-safe: `wait` uses `|| exit_code=$?` pattern; cleanup trap fully guarded - Stale cancel files cleared at command start; mismatched/empty signals rejected - Command file read into memory and removed before execution (eliminates TOCTOU with client timeout) **client.sh:** - New `cancel` subcommand: writes target cmd_id to cancel file, waits for server acknowledgment (30s timeout) - `run` timeout now sends targeted cancel signal (verifies cmd_id match to avoid killing wrong command) - `run` timeout cleans up orphaned result files - `status` shows currently running command - `flush` rejects if a command is currently running (prevents state corruption) - Exit code validated as numeric before use ### Protocol additions ``` .relay/ ├── running # server writes cmd_id:pid while executing (atomic) ├── cancel # client writes target cmd_id to request cancellation ``` ## Test plan - [ ] Start server in Docker, handshake from host - [ ] Run a command (`client.sh run "sleep 30"`), cancel it (`client.sh cancel`), verify exit code 130 - [ ] Run a command with short timeout (`--timeout 5 run "sleep 30"`), verify auto-cancel - [ ] Run a command that exits non-zero, verify server stays alive - [ ] Run `status` during execution, verify it shows the running command - [ ] Attempt `flush` during execution, verify it is rejected - [ ] Cancel when nothing is running, verify clean message 🤖 Generated with [Claude Code](https://claude.com/claude-code) ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (bash scripts for dev tooling, tested manually) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (internal tooling) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added a comprehensive debug skill guide and protocol reference with quick CLI examples and a new “Cancelling Commands” section. * **New Features** * Client-side `cancel` command to terminate the currently running remote command. * Status now reports active command (or `(idle)`). * **Improvements** * Stronger startup/validation guidance, safer shutdown/cleanup, deterministic cancel exit semantics (130), and auto-cancel on client-side timeout. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
952a62bf65 |
Fix missing attention_mask in calibration dataloader (#1261)
## Summary - When `include_labels=False` (the default for PTQ calibration), `get_dataset_dataloader` was discarding the `attention_mask` produced by the tokenizer and only returning `input_ids`. - Without `attention_mask`, HuggingFace models create a full causal mask, causing padding tokens to participate in attention during calibration and skewing quantization statistics. - This fix includes `attention_mask` alongside `input_ids` so the model correctly ignores padding tokens during calibration forward passes. ## Details In `modelopt/torch/utils/dataset_utils.py`, the tokenizer call at line 387 with `padding=True` produces both `input_ids` and `attention_mask`. The `include_labels=True` path (line 406) already preserves the full `batch_encoded` dict including `attention_mask`. However, the `include_labels=False` path was only keeping `input_ids` "for backward compatibility." During the calibration forward loop (`_forward_loop` → `_process_batch`), the batch dict is unpacked as `**kwargs` into `model.forward()`. Without `attention_mask`, HF models default to attending to all positions including padding, which pollutes calibration statistics. **Practical impact**: With `batch_size=1` there is no padding so the bug is invisible. With larger batch sizes and variable-length samples, shorter sequences get padded and the effect grows. ## Test plan - [x] Existing unit tests pass (`tests/unit/torch/utils/test_dataset_utils.py`) - [x] Pre-commit hooks pass - [ ] Verify PTQ accuracy with batch_size > 1 on a padded calibration dataset (GPU required) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
f7557221e3 |
Add file-based command relay for remote Docker testing (#1174)
## Summary - Adds a lightweight file-based client/server relay (`tools/debugger/`) that enables Claude Code (or any host-side automation) to execute commands inside a remote Docker container using only a shared filesystem — no networking setup required. - The server auto-detects the repo root, installs modelopt (`pip install -e .[dev]`), sets `PYTHONPATH`, and listens for commands. - The client supports `handshake`, `run`, `status`, and `flush` subcommands. - Includes `README.md` (full protocol docs) and `CLAUDE.md` (quick reference for Claude Code). ## Tested - Ran Qwen3.5-35B-A3B MoE PTQ with `nvfp4_experts_only` quantization via the relay: ``` bash examples/llm_ptq/scripts/huggingface_example.sh --model /hf-local/Qwen/Qwen3.5-35B-A3B/ --quant nvfp4_experts_only ``` - 42,140 quantizers inserted, MTP layers correctly excluded - Quantized checkpoint exported successfully (~208s, 147.61 GB peak GPU memory) ## Test plan - [x] Start `server.sh` inside a Docker container with the repo mounted - [x] Run `client.sh handshake` from the host - [x] Run `client.sh run "echo hello"` and verify output - [x] Run `client.sh flush` and verify `.relay/` is cleared - [x] Run a real PTQ workload end-to-end 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a file-based command relay system with host↔container client and server CLIs, supporting handshake, run, status and flush workflows to execute commands inside containers. * **Documentation** * Added guides describing the relay protocol, usage examples, CLI options, lifecycle, and operational notes (workdir, timeouts, sequential execution). * **Chores** * Updated ignore rules to exclude ephemeral relay artifacts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
18ce04f1ce |
Update the hf_ptq.yaml (#1175)
### What does this PR do? Type of change: Bug fix Fix the CLI override example comment in `tools/launcher/examples/Qwen/Qwen3-8B/hf_ptq.yaml` by adding a missing `--` (double-dash separator) to the `task_0.args` override. Without the `--` separator, the `--quant` flag would be parsed as an argument to the launcher/download script rather than being passed through to the PTQ script (`huggingface_example.sh`). This aligns the comment example with the actual `task_0.args` definition in the YAML (line 42), which already correctly includes the `--` separator. ### Testing Verified the comment now matches the actual args format used in the YAML config. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated example configuration to correct command-line parameter formatting in commented usage example. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
87ea8babe1 |
Add HuggingFace PTQ pipeline to launcher (#1100)
### What does this PR do? Type of change: New feature Adds a HuggingFace PTQ pipeline to the launcher, replacing the old `hf_ptq.sh`/`hf_ptq_local.yaml` approach with a cleaner wrapper around `huggingface_example.sh`. **Key changes:** - **New `common/hf/ptq.sh`** — wrapper script that downloads the model via `huggingface-cli` if needed, then delegates to `examples/llm_ptq/scripts/huggingface_example.sh` - **New `examples/Qwen/Qwen3-8B/hf_ptq.yaml`** — example config for Qwen3-8B nvfp4 quantization, supports both Slurm and local Docker - **Removed `common/hf_ptq/hf_ptq.sh`** and **`examples/Qwen/Qwen3-8B/hf_ptq_local.yaml`** — replaced by the new unified pipeline - **Configurable Slurm time limit** — `SlurmConfig.time` field replaces the hardcoded `"04:00:00"` in `build_slurm_executor` - **Configurable Slurm partition** — `slurm_factory` now reads `SLURM_PARTITION` env var (default: `batch`) - **`--clean` flag** — new `launch.py` option to `git clean -xdf` the examples directory before job submission - **Package `modelopt_recipes/`** — added to the nemo_run packager include list ### Testing - Tested HF PTQ pipeline on Slurm with Qwen3-8B ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ❌ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Hugging Face PTQ workflow configuration for Qwen model quantization. * Added `clean` parameter to launcher for clearing directories before job execution. * **Improvements** * Made Slurm execution time configurable per job instead of hardcoded values. * Slurm partition configuration now respects environment variables. * **Deprecated** * Removed legacy PTQ launcher scripts, replaced with unified wrapper for improved maintainability. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
fcb09bf11d |
[NVBug: 6038899] Fix MoE export crash on meta tensors with CPU offload (#1155)
## Summary Fixes `NotImplementedError` in `sync_moe_gate_up_amax` when quantizing MoE models (e.g. Qwen3-30B-A3B) on a single GPU with insufficient VRAM. When GPU memory is insufficient, ModelOpt enables CPU offload via accelerate, leaving uncalibrated expert parameters on the `meta` device. During export, `sync_moe_gate_up_amax` calls `torch.equal()` on these meta tensors, which raises `NotImplementedError` because `aten::equal` does not support meta tensors — even though calibration itself completed successfully. ## Changes - Add a guard in `sync_moe_gate_up_amax` to skip amax sync for meta tensors (which have no real data to sync) and emit a warning explaining the root cause. Bug: https://nvbugspro.nvidia.com/bug/6038899 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Added warning messages for unsupported tensor configurations in quantization workflows. * Improved edge case detection to gracefully skip processing in incompatible scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
ada1e26ba0 |
[NVBug: 6000530] Fix AWQ crash for uncalibrated MoE experts (#1142)
## Summary - Fixes NVBugs 6000530: `AttributeError: 'float' object has no attribute 'pow'` when running AWQ lite with `moe_calib_experts_ratio < 1.0` on MoE models (e.g. Qwen3-30B-A3B). - **Root cause**: When `moe_calib_experts_ratio=0.5`, some MoE experts receive zero tokens during the AWQ cache phase, leaving `act_scale` as a Python float `0.0` instead of a tensor. This causes two failures: 1. **Search phase crash**: Uncalibrated experts crash in `get_scale()` because `float.pow()` doesn't exist. 2. **Export crash**: Calibrated experts have `pre_quant_scale` but uncalibrated ones don't, causing `torch.stack()` to fail on mixed `None`/tensor values in `preprocess_linear_fusion()`. - **Fix**: Handle uncalibrated experts (`num_cache_steps == 0`) in two stages: 1. **Before search**: Disable AWQ search (`is_enabled = False`) to prevent `get_scale()` crash on float `act_scale`. 2. **During postprocessing**: Max calibrate weights and apply a neutral (all-ones) `pre_quant_scale` so export can stack scaling factors consistently across all experts. The `pre_quant_scale` buffer must be registered outside `enable_weight_access_and_writeback` because HF accelerate's `post_forward` hook drops newly-registered submodule buffers. ## Test plan - [x] Reproduce with `Qwen/Qwen3-30B-A3B`, `--qformat int4_awq`, `--moe_calib_experts_ratio 0.5` — verify no crash during calibration and export 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
16203d676a |
[NVBug 6007314] Deprecate MT-Bench support, remove openai pin, and add NeMo Evaluator reference (#1116)
### What does this PR do? Type of change: Deprecation, Bug fix, Documentation Removes MT-Bench (FastChat) evaluation support from `examples/llm_eval` and `examples/llm_ptq`. Also removes the stale `openai>=0.28.1` pin from `requirements.txt` that caused dependency conflicts with TRT-LLM (see [NVBug 6007314](https://nvbugspro.nvidia.com/bug/6007314)). Adds a NeMo Evaluator section to the llm_eval README as the recommended evaluation workflow for quantized checkpoints. **Changes:** - Delete `examples/llm_eval/run_fastchat.sh` and `examples/llm_eval/gen_model_answer.py` - Remove `mtbench` task from `examples/llm_ptq/scripts/parser.sh` and `huggingface_example.sh` - Remove `openai` dependency from `examples/llm_eval/requirements.txt` - Add NeMo Evaluator section to `examples/llm_eval/README.md` as the recommended way to evaluate quantized checkpoints from llm_ptq via TensorRT-LLM, vLLM, or SGLang - Update README docs in both `llm_eval` and `llm_ptq` - Add deprecation note to CHANGELOG.rst for 0.43 ### Usage N/A — this is a removal and documentation update. ### Testing N/A — removed code paths; no new functionality introduced. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ — MT-Bench evaluation via `--tasks mtbench` is no longer supported. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ ### Additional Information Related: [NVBug 6007314](https://nvbugspro.nvidia.com/bug/6007314) — openai dependency conflict caused by FastChat's `llm_judge` extra pinning `openai<1`. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Deprecations** * Removed MT-Bench (FastChat) evaluation support. NeMo Evaluator is now the recommended approach for evaluating quantized model checkpoints across multiple benchmarks. * **Documentation** * Updated evaluation guides to reflect NeMo Evaluator as the primary evaluation method, with support for TensorRT-LLM, vLLM, and SGLang serving backends. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
07bc4852a8 |
Default limit max hf quantzed safetensors file to 10GB (#1087)
### What does this PR do? Type of change: New feature Adds a `max_shard_size` parameter to `export_hf_checkpoint()` (and the internal `_export_diffusers_checkpoint()`) that controls the maximum size of each exported safetensors shard file. Defaults to `"10GB"`. Previously, the shard size was not explicitly controlled, which could result in very large single-file checkpoints [E.g. Qwen3.5] that are difficult to handle (e.g., slow uploads, memory issues, or exceeding file size limits on model hubs). This change ensures exported checkpoints are automatically sharded into ≤10GB files by default, while allowing users to customize the threshold. ### Usage ```python import modelopt.torch.quantization as mtq from modelopt.torch.export import export_hf_checkpoint # Default: shards capped at 10GB export_hf_checkpoint(model, dtype=torch.float16, export_dir="./output") # Custom shard size export_hf_checkpoint(model, dtype=torch.float16, export_dir="./output", max_shard_size="5GB") ``` ### Testing Verified that the `max_shard_size` parameter is correctly propagated to both the HF `save_pretrained()` (for transformers models) and diffusers component `save_pretrained()` calls. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ (new optional parameter with default matching previous behavior of HF's `save_pretrained`) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ❌ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ ### Additional Information The 10GB default was chosen to keep exported safetensors files within common file size limits while minimizing unnecessary sharding for most models. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added configurable maximum shard size parameter for Hugging Face checkpoint exports. When exporting both transformer and diffusion models, users can now specify the maximum size of each safetensors shard, controlling how checkpoint files are divided. The default shard size is set to 10GB. This feature works with both quantized and non-quantized models. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
acce79ffa3 |
Add NVFP4_EXPERTS_ONLY_CFG quantization config and YAML recipe (#1030)
### What does this PR do?
Type of change: New feature
Add `NVFP4_EXPERTS_ONLY_CFG` quantization config that targets only MoE
expert layers (`*mlp.experts*` and `*block_sparse_moe*`) with NVFP4
(W4A4) quantization, leaving all other layers (including non-expert MLP)
unquantized. This is useful for MoE models where selectively quantizing
only expert layers provides a good accuracy-performance tradeoff.
Changes:
- Refactored `_nvfp4_experts_only_quant_cfg` as a reusable building
block in `config.py`, with `_nvfp4_mlp_only_quant_cfg` now composing on
top of it
- Added `NVFP4_EXPERTS_ONLY_CFG` to the Python config choices
- Added corresponding `nvfp4_experts_only-fp8_kv.yml` YAML recipe to the
new recipe system (`modelopt_recipes/general/ptq/`)
- Updated `hf_ptq.py`, `multinode_ptq.py`, example scripts, and README
to include the new config
### Usage
```python
import modelopt.torch.quantization as mtq
model = mtq.quantize(model, mtq.NVFP4_EXPERTS_ONLY_CFG, forward_loop)
```
Or via the YAML recipe system:
```python
from modelopt.recipe import load_recipe
recipe = load_recipe("general/ptq/nvfp4_experts_only-fp8_kv")
```
### Testing
- Verified the YAML recipe matches the Python config definition
- Existing unit tests cover the quantization config infrastructure
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <\!-- Config is exercised
by existing quantization tests -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <\!-- Minor config addition -->
### Additional Information
The `experts_only` config is a subset of `mlp_only`: it quantizes
`*mlp.experts*` and `*block_sparse_moe*` patterns but not the broader
`*mlp*` pattern. The Python config was refactored so
`_nvfp4_mlp_only_quant_cfg` composes on top of
`_nvfp4_experts_only_quant_cfg`.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added an "experts-only" NVFP4 quantization option that selectively
quantizes MoE expert layers (preserving dense MLP/attention) for
improved PTQ accuracy.
* Added a corresponding PTQ recipe enabling expert-only W4A4
quantization with FP8 KV cache support.
* **Documentation**
* Updated README, examples, scripts, and changelog to document and
surface the new experts-only quantization choice.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
|
||
|
|
1dc890d971 |
Remove _moe_count_expert_calib_tokens flag; tie token counting to moe_calib_experts_ratio (#1062)
Cherry-pick for 0.43.0 ## Summary - **Remove `moe_count_expert_calib_tokens`** config field and the `_moe_count_expert_calib_tokens` internal flag. Token counting is now implicitly enabled when `moe_calib_experts_ratio` is set, removing a redundant knob. - **Change `--moe_calib_experts_ratio` default to `None`** in `hf_ptq.py` (was `1.0`). Previously all experts were force-calibrated by default; now the feature is opt-in and non-MoE models are unaffected without any flag. - **Disable `layer_sync_moe_local_experts_amax`** when `moe_calib_experts_ratio` is set, since each expert is calibrated independently with sufficient token coverage in that mode. - **Simplify `_QuantSparseMoe.forward`**: remove redundant truthy checks on `_moe_calib_experts_ratio` inside the branch that already assumes it is set. ## Changed files | File | Change | |------|--------| | `modelopt/torch/quantization/config.py` | Remove `moe_count_expert_calib_tokens` field; update `moe_calib_experts_ratio` description to document amax sync behavior | | `modelopt/torch/quantization/mode.py` | Remove `moe_count_expert_calib_tokens` propagation in `wrapped_calib_func` | | `modelopt/torch/quantization/plugins/huggingface.py` | Remove `_moe_count_expert_calib_tokens` from `_QuantSparseMoe`; simplify `forward`; skip `layer_sync_moe_local_experts_amax` when ratio is set | | `examples/llm_ptq/hf_ptq.py` | Default `--moe_calib_experts_ratio` to `None`; guard validation | | `tests/unit/.../test_sparse_moe.py` | Update tests to use `_moe_calib_experts_ratio` instead of removed flag | ## Test plan - [x] Verify `hf_ptq.py` works without `--moe_calib_experts_ratio` (non-MoE model, default `None`) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Configuration Changes** * moe_calib_experts_ratio now defaults to None (disabled) instead of 1.0; validation only occurs when a value is provided. * **Refactor** * Simplified MoE calibration flow and token-counting behavior; removed a deprecated expert-calibration configuration field. * **Documentation** * Changelog and docstrings updated to reflect the new default and calibration behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
42482b1b0f |
Add nvfp4_omlp_only config and simplify the config.py (#973)
### What does this PR do? Type of change: ? new feature 1) Add vfp4_omlp_only config == nvfp4_mlp_only + o_proj quant 2) Add block sparse MOE to mlp only config 3) Simplfiy config.py 4) Update readme in llm_ptq mention these two configs for better accuracy. ### Usage huggingface_script.sh ... --quant nvfp4_omlp_only ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, using `torch.load(..., weights_only=True)`, avoiding `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other source, did you follow IP policy in [CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added nvfp4_omlp_only quantization format for NVFP4, enabling selective quantization of MLP and output projection layers while preserving attention QKV projection accuracy. * **Changed** * pass_through_bwd now defaults to True; set to False if using STE with zeroed outlier gradients for better QAT accuracy. * **Documentation** * Updated post-training quantization guidance with NVFP4-specific configuration recommendations and usage examples. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
a4fde491cc |
Update MOE block detection logic and enable in huggingface_script.sh (#962)
### What does this PR do? Type of change: Bug fix Add moe expert calib ratio in huggingface_script.sh Also fix minimax2.5 MOE detection which does not follow other HF MOE layer convention ### Usage scripts/huggingface_example.sh --model <MiniMax-M2.5> --quant nvfp4 --moe_calib_experts_ratio 1.0 --trust_remote_code ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, using `torch.load(..., weights_only=True)`, avoiding `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other source, did you follow IP policy in [CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Configure MOE calibration experts ratio for quantization via an environment/option, enabling finer control over calibration. * **Bug Fixes** * Improved detection of sparse MOE blocks to handle varying expert/topology layouts, inferring expert counts when needed for more reliable processing. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
03a1899dda |
Support force tokens to % of total experts during calibration (#910)
## What does this PR do?
**Type of change:** New feature
**Overview:** Adds a configurable `moe_calib_experts_ratio` parameter
that controls the percentage of experts to calibrate during the forward
pass in MoE (Mixture of Experts) models. Previously, the calibration
forward always routed tokens to **all** experts, which is expensive.
This PR allows the user to specify a ratio (default: still all experts
so no behavior change) to improve expert calibration coverage without
the cost of a full-expert forward. The token counting for the expert
coverage table now tracks the calibration routing and runs on CUDA for
efficiency.
**Changes include:**
- New `moe_calib_experts_ratio` field in `QuantizeAlgorithmConfig`
(`config.py`)
- Propagation of the ratio from the algorithm config to MoE modules
during calibration (`mode.py`)
- Updated `_QuantSparseMoe.forward` to use the configurable ratio
instead of hard-coding all experts (`huggingface.py`)
- New `--moe_calib_experts_ratio` CLI flag in `hf_ptq.py` (default
`0.25`)
- Moved `expert_token_count` tensor to CUDA and updated the HTML table
title in `moe_utils.py`
## Usage
Via hf_ptq.py CLI — calibrate 50% of experts during MoE calibration
python hf_ptq.py --model <model> --qformat int4_awq
--moe_calib_experts_ratio 0.5
Via Python API — pass the ratio through the algorithm config
import modelopt.torch.quantization as mtq
quant_cfg = {
"quant_cfg": { ... },
"algorithm": {
"method": "awq_lite",
"moe_calib_experts_ratio": 0.25, # calibrate 1/4 of experts
},
}
mtq.quantize(model, quant_cfg, forward_loop=calib_loop)
## Testing
Test with Qwen3 30B A3B calibration and check the tokens per expert.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added support for configurable expert calibration during Mixture of
Experts (MOE) model quantization. Users can now specify the percentage
of experts to include during calibration, enabling better expert
coverage and improved quantization accuracy for MOE models. Default: 25%
of all experts.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: realAsma <86726418+realAsma@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
|
||
|
|
9975ba1065 |
Fix DeepSeek PTQ script (#912)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** ? Fix two bugs in the PTQ script ## Testing Run DeepseekV3.2 PTQ and export <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Enhanced data type handling in quantization examples for bf16 operations * Updated internal dependencies for quantization utilities to improve modularity <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
ac7c985d96 |
[NVBUG: 5804406] Auto detect MOE layers (#900)
## What does this PR do?
**Type of change:** New feature, new tests
**Overview:** Replace hardcoded per-model MoE class registrations
(Mixtral, Qwen2Moe, Qwen3Moe, Qwen3Next, Llama4TextMoe, Qwen3VLMoe,
MiniMaxM2, etc.) with a single generic auto-detection mechanism
(`register_sparse_moe_on_the_fly`) that walks the model tree and
identifies MoE blocks by their structural attributes (`gate` + `experts`
with `top_k`/`num_experts`). This makes MoE quantization
forward-compatible with new HuggingFace MoE architectures without
requiring explicit registration for each model family.
Additionally, this PR:
- Tracks per-expert token routing counts during calibration via a gate
forward hook, enabling visibility into expert utilization.
- Saves an HTML report of expert token counts during export
(`save_expert_token_count_table`), highlighting under-utilized experts.
- Fixes the `topk` -> `top_k` attribute name for transformers >= 5.0
compatibility.
- Also move the ptq summary prints to a file in hf_ptq.py to reduce the
prints
## Usage
Auto-detection is transparent -- no user-facing API changes are needed.
Any HuggingFace MoE model with the standard `gate`/`experts` pattern is
automatically detected and quantized:
import modelopt.torch.quantization as mtq
# Any HuggingFace MoE model (Mixtral, Qwen3Moe, DeepSeek, etc.)
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-30B-A3B")
mtq.quantize(model, mtq.INT8_DEFAULT_CFG, forward_loop)
# During export, an .moe.html report with per-expert token counts is
saved automatically
## Testing
unittest, also test exporting qwen MOE
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added expert token count visualization for Mixture of Experts models,
exported as HTML reports during model export.
* Enhanced sparse MoE quantization with improved calibration-aware
routing and automatic model block detection.
* **Tests**
* Added comprehensive test suite for sparse MoE quantization validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
|
||
|
|
3801923e9d |
Support MiniMax M2.1 (FP8 checkpoint) (#817)
## What does this PR do? **Type of change:** ? new feature **Overview:** ? Support loading the MiniMax M2.1 (FP8) checkpoint for PTQ. ## Usage scripts/huggingface_example.sh --model <minimax checkpoint> --quant nvfp4 --trust_remote_code ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added MiniMax M2.1 model quantization support with nvfp4 format. * Extended FP8 quantization capabilities with configurable dtype parameter for enhanced precision control. * **Improvements** * Enhanced detection of quantized linear module variants. * Improved weight unpacking for FP8-based linear modules. * **Documentation** * Updated supported models table to include MiniMax M2.1. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
5e43b2a5f5 |
Support Qwen3 Next MTP load and export (#860)
## What does this PR do? Fix MTP export for Qwen3 Next **Overview:** ? For Qwen3 next, the MTP weights are not stored separately in safetensors. So we use "mtp" weights key to decide if the weights are for MTP or not. ## Testing Qwen3 Next PTQ and check if MTP is in the exported checkpoint. scripts/huggingface_example.sh --model <Qwen3-Next-80B-A3B-Instruct/Thinking> --quant nvfp4 --trust_remote_code ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Optimized Multi-Token Prediction weight loading with improved layer detection and handling. * **Chores** * Simplified status reporting to display total loaded weights and detected layers. * Removed verbose per-file warnings for cleaner console output. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Zhiyu <zhiyuc@nvidia.com> |
||
|
|
615f99e746 |
Support KIMI K2 Thinking int4 checkpoint PTQ (#669)
## What does this PR do? **Type of change:** ? new feature **Overview:** Support KIMI K2 Thinking PTQ from the original int4 checkpoint. Tested with transformers 4.57.1, compressed-tensors 0.12.0 The model weights are dequantized on the fly to save GPU memory ## Usage scripts/huggingface_example.sh --model <Kimi-K2-Thinking ckpt> --quant nvfp4_mlp_only --trust_remote_code ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Support for nvfp4_mlp_only quantization format, enabling new layer-wise quantization options * Quantization support for CompressedLinear layers in quantized models * **Improvements** * Enhanced quantization for DeepSeek models with improved attention configuration handling * Optimized model loading with automatic precision configuration and weight unpacking * Better memory management during model export with automatic cache cleanup * Conditional sample generation output controlled via verbose mode <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
b0e7d9fd96 |
Define kv cache scaling factor as amax / 448 (#790)
## What does this PR do? **Overview:** ? Unified the FP8 and NVFP4 kv cache scaling factor definition so the same checkpoint can be used for both FP8 and NVFP4 kv cache quantization deployment ## Testing Unit test ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Refactor** * Fixed KV cache maximum bound to 448 for FP8 and NVFP4 quantization, simplifying configuration logic. * **Chores** * Removed internal constants from public exports. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
9c24e2c08e |
Fix Deepseek transformers model loading (#740)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** ? For Deepseek, let's force the user to apply trust_remote_code and use AutoModelForCausalLM for loading the model. ## Testing python hf_ptq.py --pyt_ckpt_path <Kimi-K2-Thinking_path> --qformat nvfp4 --export_path <quantized_ckpt> --kv_cache_qformat none --calib_size 64 --trust_remote_code --dataset cnn_dailymail ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
d541324e84 |
Disable QKV NVFP4 quantization for Qwen3 MOE (#735)
## What does this PR do? **Type of change:** ? Recipe improvement **Overview:** ? Disable QKV NVFP4 quantization for Qwen3 MOE models following the Qwen3 Next recipe for accuracy recovery ## Testing Model accuracy benchmarking Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
b1b9321877 |
Update llm_ptq doc (#685)
## What does this PR do? **Type of change:** ? documentation **Overview:** Update example doc Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
bfec1b86fd |
Support Qwen3Next NVFP4 quantization (#681)
## What does this PR do? **Type of change:** ? new feature **Overview:** Support Qwen3Next NVFP4 quantization. The QKV layers are not quantized in PTQ to retain checkpoint accuracy across benchmarks. ## Testing Checkpoint export and benchmarking using nemo-evaluator ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
e4c5a68bd8 |
[NVBug: 5707914] Update DS V3 repo commit id (#635)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** DS V3 repo commit id update so the DS R1/V3.1 models can be correctly loaded. Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
5ade7b03f8 |
[NNBUG: 5701866] Update DS V3.2 PTQ code (#630)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** 1) Update the DS V3.2 repo code reference to the latest version 2) The new DS V3.2 model now includes fp32 layers. We cast it down to match the checkpoint format during loading 3) Fix get_quant_config API change. ## Testing Generate the deepseek-ai/DeepSeek-V3.2 checkpoint ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
e20d218b43 |
[OMNIML-2857] [Experimental] Support the DeepSeek V3.2 model (#435)
## What does this PR do? **Type of change:** ? New model support **Overview:** ? ## Usage Please see examples/deepseek/README.md <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * New Features * Support for DeepSeek V3.2 quantization and automatic detection of available DeepSeek versions. * Triton-backed weight dequantization utility and MoE-aware calibration mode to improve calibration fidelity. * Documentation * DeepSeek examples README expanded with setup, conversion, calibration, and FP8→FP4 quantization workflows for R1, V3, and V3.2. * Bug Fixes * More robust, failure-tolerant copying of auxiliary files/assets during quantization. * Chores * Updated changelog and lint/ignore rules for example artifacts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
ce8ce2229c |
[NVBUG: 5617733] Update LLM generate API for modelopt LLM eval (#498)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** ? 1) Remove kv_cache_config in the generate API. It's no longer used in the code as well. We just estimate KV cache usage from other parameters 2) Add max_seq_len in the generate API to better estimate the real KV cache usage. 3) Assume default lm_eval max input sequence length to be 4096 Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
9a85e4922b |
[NVBUG: 5612606] Clear GPU cache for large models layer quantization during export (#497)
## What does this PR do? **Type of change:** Bug fix **Overview:** ? For large models like llama4 maverick, the stacked weights to fp8 conversion might hit OOM. This change aim to fix that. --------- Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
5f0ef3b310 |
[NVBUG: 5608888] Update link in vlm_ptq README for support matrix details (#499)
## What does this PR do? **Type of change:** ? documentation --------- Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |