mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Main changes: - Refactored speculative decoding export logics into `class EagleExporter` to improve cohesion; - Separated speculative decoding export entrance with quantization export (`export_hf_checkpoint()`) due to their fundamental differences: - Quantization export base model's state_dict and config, while speculative decoding only export drafter's. - Most of the model-specific logics of quantization export (e.g. diffusers, vlms) are not needed for speculative decoding export. - Quantization export produce different format than speculative decoding checkpoint. (The former produce tokenizer config, generation config, e.t.c, while the later does not need. ) ## Usage <!-- You can potentially add a usage example below. --> To export an regular bf16 eagle checkpoint without quantization, the commands are the same: ```python python scripts/export_hf_checkpoint.py --model_path <x> --export_path <x> ``` To run PTQ on online-trained eagle checkpoint and export it: ```python python hf_ptq.py --pyt_ckpt_path <x> --qformat fp8 --export_path <x> ``` The above two commands will produce drafter ckpt for deployment, in the same foramt. ## Testing <!-- Mention how have you tested your change if applicable. --> Tested setting: - Base model: llama3.1-8b - Algorithms: eagle - Export path tested: - (Unquantized online ckpt) `python scripts/export_hf_checkpoint.py --model_path <x> --export_path <x>` - (PTQ) export `python hf_ptq.py --pyt_ckpt_path <x> --qformat fp8 --export_path <x>` - Tested deployment on vllm. Got normal AR. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added export functionality for speculative decoding-optimized models * Support for multiple speculative decoding architectures with pre-configured deployment templates * Enhanced model export detection and automatic routing for optimized models * **Tests** * Updated export validation tests for speculative decoding models <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>