Files
Model-Optimizer/modelopt
realAsma a7f339ed08 Reject unsupported partial-block INT4/W4A8 AWQ export (#2320)
### What does this PR do?

Type of change: Bug fix

Reject INT4 and W4A8 AWQ export when a weight's input dimension is not
divisible by
the configured block size.

Nemotron-3-Nano-4B has weights with input dimension 3136, which is not
divisible by
the configured block size 128. Partial INT4/W4A8 blocks are not
supported. The
export path previously inferred an incorrect block size and reached an
out-of-bounds
CUDA scale index. It now raises a clear `NotImplementedError` before GPU
indexing.

### Usage

No API change. Unsupported partial-block INT4/W4A8 AWQ exports now fail
early with a
clear error instead of a CUDA device-side assertion.

### Testing

- `pre-commit run --files modelopt/torch/export/quant_utils.py
tests/gpu/torch/export/test_export.py` — passed.
- Focused GPU test for supported and partial-block INT4/W4A8 AWQ packing
— 2 cases passed.
- The reported Nemotron-3-Nano-4B failure was reproduced and localized
with
  `CUDA_LAUNCH_BLOCKING=1` before applying the guard.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A — this fixes behavior in the current
unreleased 0.47 development line.
- Did you get Claude approval on this PR?: N/A — the focused candidate
passed independent senior code review and test review.

### Additional Information

The check is format-generic and does not special-case Nemotron or any
architecture.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* INT4/AWQ and W4A8 AWQ checkpoint exports now validate block sizes
before processing weights.
* Exports reject block sizes that are non-positive, non-integer, or do
not evenly divide the input dimension.
* Improved validation for compressed INT4/AWQ weights, including
partially filled quantization blocks.

* **Tests**
* Added GPU coverage for valid and invalid INT4/AWQ and W4A8 packing
scenarios.
* Expanded quantized-model export coverage using larger test dimensions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-09-04 15:45:45 -07:00
..