mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: refactor (no behaviour change) The five GGML IQ CUDA encoders were five copies of the same search. IQ2_XS and IQ2_XXS shared 139 of their roughly 150 lines of encoder and launcher code, and IQ2_S 113 of them. IQ1_S and IQ1_M had the same structure with a different choice space. Every scaled packer also validated its scales twice, in the `ggml.cpp` pybind wrapper and again in the CUDA entry point. This PR keeps **one encoder per family**, as two templates: - **`iq2_family.cuh`** for IQ2_XS, IQ2_XXS and IQ2_S. The grid sits in shared memory, the 16 local scales are scored per group, and each vector then takes its best entry under the chosen scale. A format supplies its group shape, whether it stores seven sign bits and recovers the eighth from parity, and a `store()` that writes the chosen entries, sign masks and local scales into its layout. - **`iq1_family.cuh`** for IQ1_S and IQ1_M, over the shared ternary grid. Each group picks one of `kChoices` options. With `kSharedShift` the option also fixes the ±1/8 delta (IQ1_S: `shift * 8 + local`); otherwise each vector picks its own (IQ1_M). IQ1_S's scale kernel now writes FP16 scales, so both IQ1 formats take the same input. Each format file is now one `Format` struct, holding its layout constants and `store()`, plus its entry point: 58–100 lines each. Validation lives once in `common.cuh`, as `check_pack_inputs` and `check_scaled_pack_inputs`. `ggml.cpp` binds the CUDA entry points directly instead of through five wrappers. **The kernel sources shrink from 1,536 to 1,241 lines** (+665 / −960). This is the first of two PRs. #2604 builds on it: it adds CUDA decoders as a `decode()` next to each format's `store()`, and makes export reuse fake quant's packed payloads. ### Testing **Nothing changes in the output.** Before the refactor I hashed 40 outputs: 5 formats × float32/bfloat16/float16/float64 inputs × encode and decode, on a weight with zero, tiny, oversized and non-finite blocks. All 40 hash the same afterwards. **Encode speed is unchanged.** Old and new were timed alternately for four rounds, in both orders, on an idle RTX PRO 6000 with a 5632×2048 weight. They were within 1% for every format: IQ1_S 37.6 / 37.6 ms, IQ1_M 37.0 / 37.0, IQ2_XXS 10.9 / 10.9, IQ2_XS 11.9 / 12.0, IQ2_S 15.5 / 15.4. - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **49 passed** - **Validation reports the same errors in the same order.** Over 5 formats × 8 combinations of bad arguments (devices, dtype, width, grid shape, scales dtype, length and sign), every first error matches main's. - `tests/gpu/_extensions/test_torch_extensions.py`: the validation-message tests pass. #2515's two Q8_0 tests fail identically on a clean `main` on this GPU. - IQ unit tests (`test_ggml_backend.py`, `test_iq_formats.py`, `test_convert_hf_config.py`, `test_presets.py`, `test_export_weight.py`): **173 passed** ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ Same bindings, messages and bytes. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code sources or dependencies. - Did you write any new necessary tests?: N/A. A refactor with no behaviour change, verified by the hashes above and the existing GPU tests. - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: **this** → #2604 (pack each IQ weight once and decode on CUDA). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Quantization now checks that inputs and grids are CUDA tensors on the same device, with compatible shapes. Scaled formats also validate scale type, shape, and finite, non-negative values. * **Improvements** * IQ1 and IQ2 formats share common encoding paths while retaining their format-specific output layouts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>