Files
Model-Optimizer/modelopt_recipes/general/qad
realAsmaandClaude Fable 5 d641b4a524 Add LSQ (Learned Scale Quantization) support and recipes (#1884)
## Summary

This PR adds two scale-learning algorithms to Model Optimizer—**LSQ**
and **Dual-LSQ**—and provides focused NVFP4 Quantization-Aware
Distillation (QAD) recipes for them.

- **LSQ (learnt scale quantization)** corresponds to the original
[Learned Step Size Quantization
paper](https://arxiv.org/pdf/1902.08153). ModelOpt uses the name *learnt
scale quantization* and learns one scale shared before and after
quantization.
- **Dual-LSQ** learns separate pre-quantization and post-quantization
scales.

ModelOpt implements both algorithms with an `amax` reparameterization:
it learns `amax` rather than the scale directly, with `scale = amax /
max_bound`. This allows scale parameters to be optimized during QAD
while the quantized model weights remain fixed.

## What changed

- Add an `LSQConfig` algorithm and calibration mode with configurable
learned pre/post amax values, tied-scale support, and optional pre-scale
quantization.
- Extend `TensorQuantizer`, tensor quantization, and the FP4 GEMM path
so gradients propagate through LSQ scale parameters.
- Preserve LSQ parameter dtypes when training with FSDP2.
- Add two modular NVFP4 recipes:
  - **LSQ:** one shared learned pre/post amax value.
- **Dual-LSQ:** independent learned pre/post amax values; the pre-scale
is not quantized.
- Add a scale-only QAD training config that trains only LSQ amax
parameters.
- Add focused CPU/GPU LSQ behavior, recipe, and FSDP2 coverage.

## QAD GPU memory

Measured on the same four-GPU Qwen3-1.7B QAD setup:

| Training mode | Peak allocated / GPU | Peak reserved / GPU | Reduction
vs full-parameter QAD |
|---|---:|---:|---:|
| Full-parameter QAD | 28.3 GiB | 30.2 GiB | — |
| Scale-only QAD | 24.1 GiB | 24.8 GiB | 4.2 GiB allocated (15%); 5.4
GiB reserved (18%) |

Scale-only QAD learns only a small subset of parameters—typically about
the model size divided by 16 or 8, depending on scale granularity.
Optimizer states and gradients are therefore maintained only for those
learned scale/`amax` parameters, which accounts for the lower GPU-memory
footprint compared with full-parameter QAD.

## Validation

- Focused LSQ and recipe unit tests: **38 passed**
- GPU tests: `test_lsq_cuda.py` 6 passed; FSDP2 LSQ test passed.
- Pre-commit checks passed for all changed files.
- `git diff --check` passed.


## Qwen3-1.7B NVFP4 PTQ + QAD results

All four variants are compared at the same 600-step cutoff. The
full-parameter Dual-LSQ run had already reached step 757 when it was
stopped for the requested cap, so its curves and summary below discard
steps after 600.

| Variant | Mean train loss (steps 1–600) | Eval loss (step 600) |
|---|---:|---:|
| Full-parameter Dual-LSQ | 0.2289485 | 0.1064966 |
| Scale-only Dual-LSQ | 0.2604457 | 0.1264785 |
| Scale-only LSQ | 0.3045209 | 0.1495215 |
| Dynamic-max full-parameter | 0.2504419 | 0.1215034 |

![Qwen3-1.7B NVFP4 QAD training and evaluation loss
curves](https://raw.githubusercontent.com/NVIDIA/Model-Optimizer/14c403c2e3aa833ef025c835d24ce3292556d6e4/docs/source/assets/pr1884-qwen3-1.7b-qad-loss-curves.png)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added LSQ quantization with learnable/tied `amax`, optional pre-scale
quantization, and configurable scale-calibration.
* Added FP4/INT4 straight-through casting utilities for training-time
gradients.
* Added unified QAT/QAD-capable quantization recipes and new LSQ-based
recipe examples.
* **Bug Fixes**
* Improved quantization recipe loading/validation and broadened accepted
recipe types for recipe directories and scripts.
  * Enhanced FSDP2 quantization by aligning LSQ `amax` parameter dtypes.
* **Documentation**
* Updated guides and examples to use the new quantization recipe name,
including deprecation guidance for the old PTQ alias.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
Signed-off-by: realAsma <86726418+realAsma@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 17:18:41 -07:00
..