Autoquant and GPTQ in support in Megatron-Core [OMNIML-3151] (#1562)

### What does this PR do?

Type of change: New Feature

Autoquant and GPTQ in support in Megatron-Core

- Add EP support to AutoQuantize
- Register MCore support in AutoQuantize
- Add decoder `output_layer` (lm head) to layerwise hook so that GPTQ
can register all decoder layers & lm head
- Split dataloader helper function out of megatron calibration utils so
that AutoQuantize in Megatron-LM can reuse the same dataloader

### Usage
See https://github.com/NVIDIA/Megatron-LM/pull/4821 for Autoquant usage
in Megatron
```python
# For GPTQ pick a recipe that uses gptq algorithm and run mtq.quantize
# e.g. general/ptq/nvfp4_default-kv_none-gptq
```

### Testing
Tested AutoQuant on Nemotron Nano and Ultra.
Tested GPTQ on Nano 3.
Added unit tests for both AutoQuant and GPTQ

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Summary by CodeRabbit

* **New Features**
* Added lazy Megatron-Core AutoQuant integration with Megatron-specific
quantization hooks and better decoder-layer discovery for layerwise
calibration.
* Improved AutoQuantize for expert-parallel (EP) models, including
consistent per-layer recipe selection across DP/TP/EP.
* Extended quant-layer grouping for NemotronH MCore fused
“local_experts” linear layers.

* **Bug Fixes**
* Prevented division-by-zero when calibration inputs are empty during
Hessian updates.

* **Tests**
* Added/extended unit and GPU coverage for EP AutoQuant, decoder-layer
calibration discovery behavior, and zero-token Hessian no-op.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Jenny Chen <jennifchen@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Jenny Chen
2026-07-01 19:12:06 -07:00
committed by GitHub
co-authored by Claude Opus 4.8
parent 892d27aa76
commit 9038b71f09
16 changed files with 637 additions and 183 deletions
+1 -1
View File
@@ -42,7 +42,7 @@ jobs:
timeout: 60
container_image: nvcr.io/nvidia/pytorch:26.05-py3
- example: gpu_megatron
timeout: 60
timeout: 75
container_image: nvcr.io/nvidia/nemo:26.06
- example: gpu_trtllm
timeout: 15