mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: Bug fix Reject INT4 and W4A8 AWQ export when a weight's input dimension is not divisible by the configured block size. Nemotron-3-Nano-4B has weights with input dimension 3136, which is not divisible by the configured block size 128. Partial INT4/W4A8 blocks are not supported. The export path previously inferred an incorrect block size and reached an out-of-bounds CUDA scale index. It now raises a clear `NotImplementedError` before GPU indexing. ### Usage No API change. Unsupported partial-block INT4/W4A8 AWQ exports now fail early with a clear error instead of a CUDA device-side assertion. ### Testing - `pre-commit run --files modelopt/torch/export/quant_utils.py tests/gpu/torch/export/test_export.py` — passed. - Focused GPU test for supported and partial-block INT4/W4A8 AWQ packing — 2 cases passed. - The reported Nemotron-3-Nano-4B failure was reproduced and localized with `CUDA_LAUNCH_BLOCKING=1` before applying the guard. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A — this fixes behavior in the current unreleased 0.47 development line. - Did you get Claude approval on this PR?: N/A — the focused candidate passed independent senior code review and test review. ### Additional Information The check is format-generic and does not special-case Nemotron or any architecture. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * INT4/AWQ and W4A8 AWQ checkpoint exports now validate block sizes before processing weights. * Exports reject block sizes that are non-positive, non-integer, or do not evenly divide the input dimension. * Improved validation for compressed INT4/AWQ weights, including partially filled quantization blocks. * **Tests** * Added GPU coverage for valid and invalid INT4/AWQ and W4A8 packing scenarios. * Expanded quantized-model export coverage using larger test dimensions. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: realAsma <akuriparambi@nvidia.com>