mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
## Summary - add IQ format metadata and packed-weight export - support Hugging Face and TP=1 Megatron export paths - reject fused-MoE IQ export until a deployment loader owns its packed layout - document the shaped `uint8` weight contract and the fused-expert boundary - add Hugging Face, Megatron, metadata, and fused-expert export tests ## PR split This work is split into four focused PRs. Each PR targets `main` and owns a disjoint file set: 1. **Kernel** — [#2448: Add CUDA kernels for IQ packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448) 2. **Quantization** — [#2446: Add IQ quantization codecs and backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446) 3. **Export** — [#2447: Export IQ checkpoints from HF and Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447) 4. **Recipes** — [#2449: Add IQ post-training quantization recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449) The required merge order is #2448, #2446, #2447, then #2449. ## Scope This PR owns only export code, deployment documentation, and export tests. It targets `main` and should merge after #2448 and #2446. It does not contain kernel, codec/backend, or recipe files. ## Deployment consumer boundary Dense weights and individually named expert weights use the documented shaped `uint8` contract. Megatron fused-MoE IQ export is intentionally rejected with `NotImplementedError`: its payload would have shape `[num_experts, out_features, in_features // 256, payload_bytes]`, and no deployment loader in this stack currently owns that layout. Support should be enabled only with a loader integration test. ## Test coverage - [Hugging Face packed-weight export](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/unit/torch/export/test_export_weight.py) - [quantization metadata](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/unit/torch/export/test_get_quantization.py) - [Megatron unified export and fused-MoE rejection](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/gpu_megatron/torch/export/test_unified_export_megatron.py) ## Validation - all pre-commit hooks pass for the changed files - 89 focused Hugging Face export, metadata, and fused-expert tests pass locally - direct checks cover both fused-MoE export entry points for IQ1_S and IQ2_XS - Megatron GPU execution remains delegated to GPU CI - restricted-term scan passes <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for IQ1_S and IQ2_XS GGML quantization formats in unified Hugging Face and Megatron exports. * Added quantization metadata, tensor-shape recovery, packing details, and IQ2_XS size documentation. * Added validation for required block sizes and tensor parallelism settings. * **Limitations** * Fused-MoE and GPT-OSS IQ expert packing are not supported. * IQ exports require standard `weight` attributes in Hugging Face models. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>