mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
## Summary This PR adds two scale-learning algorithms to Model Optimizer—**LSQ** and **Dual-LSQ**—and provides focused NVFP4 Quantization-Aware Distillation (QAD) recipes for them. - **LSQ (learnt scale quantization)** corresponds to the original [Learned Step Size Quantization paper](https://arxiv.org/pdf/1902.08153). ModelOpt uses the name *learnt scale quantization* and learns one scale shared before and after quantization. - **Dual-LSQ** learns separate pre-quantization and post-quantization scales. ModelOpt implements both algorithms with an `amax` reparameterization: it learns `amax` rather than the scale directly, with `scale = amax / max_bound`. This allows scale parameters to be optimized during QAD while the quantized model weights remain fixed. ## What changed - Add an `LSQConfig` algorithm and calibration mode with configurable learned pre/post amax values, tied-scale support, and optional pre-scale quantization. - Extend `TensorQuantizer`, tensor quantization, and the FP4 GEMM path so gradients propagate through LSQ scale parameters. - Preserve LSQ parameter dtypes when training with FSDP2. - Add two modular NVFP4 recipes: - **LSQ:** one shared learned pre/post amax value. - **Dual-LSQ:** independent learned pre/post amax values; the pre-scale is not quantized. - Add a scale-only QAD training config that trains only LSQ amax parameters. - Add focused CPU/GPU LSQ behavior, recipe, and FSDP2 coverage. ## QAD GPU memory Measured on the same four-GPU Qwen3-1.7B QAD setup: | Training mode | Peak allocated / GPU | Peak reserved / GPU | Reduction vs full-parameter QAD | |---|---:|---:|---:| | Full-parameter QAD | 28.3 GiB | 30.2 GiB | — | | Scale-only QAD | 24.1 GiB | 24.8 GiB | 4.2 GiB allocated (15%); 5.4 GiB reserved (18%) | Scale-only QAD learns only a small subset of parameters—typically about the model size divided by 16 or 8, depending on scale granularity. Optimizer states and gradients are therefore maintained only for those learned scale/`amax` parameters, which accounts for the lower GPU-memory footprint compared with full-parameter QAD. ## Validation - Focused LSQ and recipe unit tests: **38 passed** - GPU tests: `test_lsq_cuda.py` 6 passed; FSDP2 LSQ test passed. - Pre-commit checks passed for all changed files. - `git diff --check` passed. ## Qwen3-1.7B NVFP4 PTQ + QAD results All four variants are compared at the same 600-step cutoff. The full-parameter Dual-LSQ run had already reached step 757 when it was stopped for the requested cap, so its curves and summary below discard steps after 600. | Variant | Mean train loss (steps 1–600) | Eval loss (step 600) | |---|---:|---:| | Full-parameter Dual-LSQ | 0.2289485 | 0.1064966 | | Scale-only Dual-LSQ | 0.2604457 | 0.1264785 | | Scale-only LSQ | 0.3045209 | 0.1495215 | | Dynamic-max full-parameter | 0.2504419 | 0.1215034 |  <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added LSQ quantization with learnable/tied `amax`, optional pre-scale quantization, and configurable scale-calibration. * Added FP4/INT4 straight-through casting utilities for training-time gradients. * Added unified QAT/QAD-capable quantization recipes and new LSQ-based recipe examples. * **Bug Fixes** * Improved quantization recipe loading/validation and broadened accepted recipe types for recipe directories and scripts. * Enhanced FSDP2 quantization by aligning LSQ `amax` parameter dtypes. * **Documentation** * Updated guides and examples to use the new quantization recipe name, including deprecation guidance for the old PTQ alias. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> Signed-off-by: realAsma <86726418+realAsma@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>