Files
Model-Optimizer/examples/llm_autodeploy
Frida Hou eb3e6edb50 [fix][5875912] Fix autoquant-autodeploy example (#878)
## What does this PR do?

**Type of change:** Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** ?
Please check Bug ticket

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
Tested with
```
./scripts/run_auto_quant_and_deploy.sh     --hf_ckpt ./models/Qwen/Qwen3-8B     --save_quantized_ckpt ./qwen3_8B_autoquant     --quant fp8     --effective_bits 10.0
```

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Simplified LLM initialization by removing intermediate configuration
layer
  * Updated attention backend from triton to flashinfer

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
2026-02-11 20:33:08 +00:00
..
2025-03-03 22:54:22 +05:30
2025-03-03 22:54:22 +05:30
2025-06-05 13:24:07 -07:00

Deploy AutoQuant Models with AutoDeploy

This guide demonstrates how to deploy mixed-precision models using ModelOpt's AutoQuant and TRT-LLM's AutoDeploy.

ModelOpt's AutoQuant is a post-training quantization (PTQ) algorithm that optimizes model quantization by selecting the best quantization format for each layer while adhering to user-defined compression constraints. This approach allows users to balance model accuracy and performance effectively.

TRT-LLM's AutoDeploy is designed to simplify and accelerate the deployment of PyTorch models, including off-the-shelf models like those from Hugging Face, to optimized inference environments with TRT-LLM. It automates graph transformations to integrate inference optimizations such as tensor parallelism, KV-caching and quantization. AutoDeploy supports optimized in-framework deployment, minimizing the amount of manual modification needed.

Prerequisites

AutoDeploy is available in TensorRT-LLM docker images. Please refer to our Installation Guide for more details.

1. Quantize and Deploy Model

Run the following command to quantize your model and launch an OpenAI-compatible endpoint:

./scripts/run_auto_quant_and_deploy.sh \
    --hf_ckpt <path_to_HF_model> \
    --save_quantized_ckpt <path_to_save_quantized_checkpoint> \
    --quant fp8,nvfp4 \
    --effective_bits 4.5

Parameters:

  • --hf_ckpt: Path to the unquantized Hugging Face checkpoint
  • --save_quantized_ckpt: Output path for the quantized checkpoint
  • --quant: Quantization formats to use (e.g., fp8,nvfp4)
  • --effective_bits: Target overall precision (higher values preserve accuracy for sensitive layers)
  • --calib_batch_size: (Optional, default=8) Calibration batch size. Reduce if encountering OOM issues

Note

:

  • NVFP4 is only available on Blackwell GPUs. For Hopper GPUs:
    • Remove nvfp4 from the --quant parameter
    • Increase --effective_bits above 8.0 for FP8-only AutoQuant
  • For tensor parallelism, add --world_size <gpu_num>
  • Additional generation and sampling configurations can be found in api_server.py

2. Test the Deployment

Send test prompts to the server:

python api_client.py --prompt "What is AI?" "What is golf?"

This will return generated responses for both prompts from your deployed model.