Files
Model-Optimizer/examples/speculative_decoding/scripts
h-guo18 a34d613d3c Feat: Speculatice Decoding export with quantization support (#913)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** 


Main changes:  
- Refactored speculative decoding export logics into `class
EagleExporter` to improve cohesion;

 
- Separated speculative decoding export entrance with quantization
export (`export_hf_checkpoint()`) due to their fundamental differences:
- Quantization export base model's state_dict and config, while
speculative decoding only export drafter's.
- Most of the model-specific logics of quantization export (e.g.
diffusers, vlms) are not needed for speculative decoding export.
- Quantization export produce different format than speculative decoding
checkpoint. (The former produce tokenizer config, generation config,
e.t.c, while the later does not need. )

## Usage
<!-- You can potentially add a usage example below. -->

To export an regular bf16 eagle checkpoint without quantization, the
commands are the same:
```python
python scripts/export_hf_checkpoint.py --model_path <x> --export_path <x>
```

To run PTQ on online-trained eagle checkpoint and export it:
```python
python hf_ptq.py --pyt_ckpt_path <x> --qformat fp8 --export_path <x>
```

The above two commands will produce drafter ckpt for deployment, in the
same foramt.

## Testing
<!-- Mention how have you tested your change if applicable. -->

Tested setting:
- Base model: llama3.1-8b
- Algorithms: eagle
- Export path tested: 
- (Unquantized online ckpt) `python scripts/export_hf_checkpoint.py
--model_path <x> --export_path <x>`
- (PTQ) export `python hf_ptq.py --pyt_ckpt_path <x> --qformat fp8
--export_path <x>`
- Tested deployment on vllm. Got normal AR. 

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
  * Added export functionality for speculative decoding-optimized models
* Support for multiple speculative decoding architectures with
pre-configured deployment templates
* Enhanced model export detection and automatic routing for optimized
models

* **Tests**
  * Updated export validation tests for speculative decoding models

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-03-04 01:46:02 +00:00
..