Files
Model-Optimizer/modelopt
ed5c5ed369 [OMNIML-5899] Export IQ checkpoints from HF and Megatron (#2447)
## Summary

- add IQ format metadata and packed-weight export
- support Hugging Face and TP=1 Megatron export paths
- reject fused-MoE IQ export until a deployment loader owns its packed
layout
- document the shaped `uint8` weight contract and the fused-expert
boundary
- add Hugging Face, Megatron, metadata, and fused-expert export tests

## PR split

This work is split into four focused PRs. Each PR targets `main` and
owns a disjoint file set:

1. **Kernel** — [#2448: Add CUDA kernels for IQ
packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448)
2. **Quantization** — [#2446: Add IQ quantization codecs and
backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446)
3. **Export** — [#2447: Export IQ checkpoints from HF and
Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447)
4. **Recipes** — [#2449: Add IQ post-training quantization
recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449)

The required merge order is #2448, #2446, #2447, then #2449.

## Scope

This PR owns only export code, deployment documentation, and export
tests. It targets `main` and should merge after #2448 and #2446. It does
not contain kernel, codec/backend, or recipe files.

## Deployment consumer boundary

Dense weights and individually named expert weights use the documented
shaped `uint8` contract. Megatron fused-MoE IQ export is intentionally
rejected with `NotImplementedError`: its payload would have shape
`[num_experts, out_features, in_features // 256, payload_bytes]`, and no
deployment loader in this stack currently owns that layout. Support
should be enabled only with a loader integration test.

## Test coverage

- [Hugging Face packed-weight
export](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/unit/torch/export/test_export_weight.py)
- [quantization
metadata](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/unit/torch/export/test_get_quantization.py)
- [Megatron unified export and fused-MoE
rejection](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/gpu_megatron/torch/export/test_unified_export_megatron.py)

## Validation

- all pre-commit hooks pass for the changed files
- 89 focused Hugging Face export, metadata, and fused-expert tests pass
locally
- direct checks cover both fused-MoE export entry points for IQ1_S and
IQ2_XS
- Megatron GPU execution remains delegated to GPU CI
- restricted-term scan passes


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for IQ1_S and IQ2_XS GGML quantization formats in
unified Hugging Face and Megatron exports.
* Added quantization metadata, tensor-shape recovery, packing details,
and IQ2_XS size documentation.
* Added validation for required block sizes and tensor parallelism
settings.

* **Limitations**
  * Fused-MoE and GPT-OSS IQ expert packing are not supported.
* IQ exports require standard `weight` attributes in Hugging Face
models.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 00:00:33 +00:00
..