mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: Bug fix Layerwise calibration now calibrates enabled quantizers outside transformer layers, such as `lm_head`, while hiding decoder layers from the additional calibration traversal. It also fails early when these quantizers are combined with progressive layerwise export, whose in-place conversion makes the required model calibration unsafe. ### Usage No API changes. ### Testing - `pytest_pwd tests/unit/torch/quantization/test_calib.py tests/unit/torch/quantization/test_layerwise_calibrate.py -q` — 68 passed - Pre-commit on all six changed files — passed - `git diff --check` — passed ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Layerwise calibration now supports quantizers outside transformer decoder layers, including LM-head activation quantizers. - Calibration supports models with CPU-offloaded components while preserving model behavior. - The MSE calibration utility is now publicly available. - Added warnings for calibration passes involving offloaded components. - **Bug Fixes** - Prevented unsupported export-mode calibration when outside-layer quantizers are present. - Ensured outside-layer calibration runs after transformer-layer restoration. - Preserved original forwards while hiding decoder subtrees from traversal and state collection. - Ensured cleanup and alias restoration after errors or completion. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: realAsma <akuriparambi@nvidia.com>