mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: new feature. Follow-up to merged #2272. Adds composition of existing GEMM quantization with KV-cache AutoQuantize: - fixed FP8 GEMM PTQ followed by mixed-KV AutoQuantize; - gradient-based NVFP4/FP8 GEMM AutoQuantize followed by independent mixed-KV AutoQuantize; - an optional `kv_auto_quantize` recipe stage with independent method, constraints, candidates, and checkpoint path; - ordered `hf_ptq.py` orchestration that keeps selected weight/activation QDQ active while its calibration state remains frozen during KV candidate calibration; - fail-closed validation when a preceding stage leaves actual K/V quantizers enabled; and - unified export of a uniform-weight or mixed-weight checkpoint together with the selected per-layer KV map. The KV search still uses the public `mtq.auto_quantize(..., constraints={"cost_model": "kv_cache", ...})` API from #2272. On a converted model, the API preserves existing non-KV quantizers and requires K/V to be disabled before search. Fresh-model behavior is unchanged and starts from a deny-all quantizer baseline. #### Why a follow-up field instead of a generic stage list? This PR deliberately supports the two composition forms required by `hf_ptq.py` without replacing the stable recipe schema. Existing recipes already express a fixed `quantize` baseline plus one primary `auto_quantize` search. A generic ordered `stages` list would require a broader recipe/API migration, indexed checkpoint semantics, and compatibility rules for arbitrary stage sequences. There is not yet a demonstrated third search stage that justifies that surface-area change. The two searches are not combined inside `mtq.auto_quantize`: each invocation owns one search domain, constraint model, scoring method, and resumable checkpoint. Their ordering and independent checkpoint paths are orchestration concerns, while candidate calibration, scoring, selection, and state application remain in the shared public API. A general stage pipeline can be considered separately if more than this one optional KV follow-up is needed. Both solvers and scoring protocols are unchanged. The KV checkpoint compatibility signature additionally fingerprints the preceding quantizer configuration and calibrated state. Unsupported uniform-weight plus mixed-KV exports record `kv_cache_deployment_supported: false` in both ModelOpt and converted HF metadata. ### Usage Fixed FP8 GEMM PTQ followed by KV AutoQuantize: ```bash python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path Qwen/Qwen3-8B \ --recipe general/auto_quantize/fp8_ptq_then_kv_fp8_nvfp4_cast_kl_div_at_5p4bits \ --kv_auto_quantize_checkpoint /path/to/kv_autoquant.pth \ --export_path /path/to/qwen3-8b-fp8-and-mixed-kv ``` Weight AutoQuantize followed by KV AutoQuantize: ```bash python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path Qwen/Qwen3-8B \ --recipe general/auto_quantize/nvfp4_fp8_gradient_then_kv_fp8_nvfp4_cast_kl_div_at_5p4bits \ --auto_quantize_checkpoint /path/to/weight_autoquant.pth \ --kv_auto_quantize_checkpoint /path/to/kv_autoquant.pth \ --export_path /path/to/qwen3-8b-autoquant-and-mixed-kv ``` KV checkpoint resume requires identical preceding non-K/V quantizer configuration and calibrated state. If rerunning the preceding stage changes that state, use a new KV checkpoint path to recompute sensitivities; configuration identity alone is insufficient to reuse the scores safely. ### Testing - Latest changed-area validation: 126 tests passed across `hf_ptq.py` orchestration, KV checkpoint compatibility, export metadata, and HF configuration conversion. - A broader local run had 604 passes, one skip, and six failures: two socket-binding failures under the sandbox and four local Transformers API incompatibilities. This is not a full-suite pass. - The fixed-PTQ→KV recipe executes end to end on a tiny offline Qwen fixture. - Public API coverage verifies that composed KV search preserves preceding weight quantization and rejects enabled K/V state. - Changed-file pre-commit hooks passed; the isolated recipe validator also passed after dependency bootstrap. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ (0.48.0 composition feature and KV checkpoint flag deprecation) - Did you get Claude approval on this PR?: ❌ ### Additional Information - This follow-up targets `main`, which contains merged #2272. - `--auto_quantize_checkpoint` and `--kv_auto_quantize_checkpoint` are intentionally separate because KV sensitivities depend on the preceding GEMM state. - Uniform-weight plus mixed-KV exports are for artifact inspection until the runtime's uniform-weight ModelOpt configuration consumes `kv_cache_quantized_layers`. Export emits an actionable warning and records `kv_cache_deployment_supported: false`; this marker does not itself add runtime support. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added staged post-training quantization workflows for weights and KV caches, including dedicated KV-cache checkpoints. - Added FP8/NVFP4 recipes with configurable bit constraints and scoring. - KV-cache quantization now supports pre-quantized models. - **Bug Fixes** - Mixed weight and KV-cache quantization now exports with a warning instead of failing. - Improved validation and checkpoint compatibility for staged configurations. - Added safeguards for configurations without enabled weight quantizers. - **Documentation** - Clarified staged KV-cache workflows, checkpoint options, configuration behavior, and unsupported deployment combinations. - Documented deprecated legacy quantization options and their replacement behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>