## What does this PR do? **Type of change:** Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** ? Please check Bug ticket ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing Tested with ``` ./scripts/run_auto_quant_and_deploy.sh --hf_ckpt ./models/Qwen/Qwen3-8B --save_quantized_ckpt ./qwen3_8B_autoquant --quant fp8 --effective_bits 10.0 ``` ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Simplified LLM initialization by removing intermediate configuration layer * Updated attention backend from triton to flashinfer <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Deploy AutoQuant Models with AutoDeploy
This guide demonstrates how to deploy mixed-precision models using ModelOpt's AutoQuant and TRT-LLM's AutoDeploy.
ModelOpt's AutoQuant is a post-training quantization (PTQ) algorithm that optimizes model quantization by selecting the best quantization format for each layer while adhering to user-defined compression constraints. This approach allows users to balance model accuracy and performance effectively.
TRT-LLM's AutoDeploy is designed to simplify and accelerate the deployment of PyTorch models, including off-the-shelf models like those from Hugging Face, to optimized inference environments with TRT-LLM. It automates graph transformations to integrate inference optimizations such as tensor parallelism, KV-caching and quantization. AutoDeploy supports optimized in-framework deployment, minimizing the amount of manual modification needed.
Prerequisites
AutoDeploy is available in TensorRT-LLM docker images. Please refer to our Installation Guide for more details.
1. Quantize and Deploy Model
Run the following command to quantize your model and launch an OpenAI-compatible endpoint:
./scripts/run_auto_quant_and_deploy.sh \
--hf_ckpt <path_to_HF_model> \
--save_quantized_ckpt <path_to_save_quantized_checkpoint> \
--quant fp8,nvfp4 \
--effective_bits 4.5
Parameters:
--hf_ckpt: Path to the unquantized Hugging Face checkpoint--save_quantized_ckpt: Output path for the quantized checkpoint--quant: Quantization formats to use (e.g.,fp8,nvfp4)--effective_bits: Target overall precision (higher values preserve accuracy for sensitive layers)--calib_batch_size: (Optional, default=8) Calibration batch size. Reduce if encountering OOM issues
Note
:
- NVFP4 is only available on Blackwell GPUs. For Hopper GPUs:
- Remove
nvfp4from the--quantparameter- Increase
--effective_bitsabove 8.0 for FP8-only AutoQuant- For tensor parallelism, add
--world_size <gpu_num>- Additional generation and sampling configurations can be found in
api_server.py
2. Test the Deployment
Send test prompts to the server:
python api_client.py --prompt "What is AI?" "What is golf?"
This will return generated responses for both prompts from your deployed model.