mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Make the QAT/QAD guide the central place for concepts, background, and framework selection (#2590)
### What does this PR do? Make the QAT/QAD guide the central place for concepts, background, and framework selection. Have the Hugging Face and Megatron Bridge tutorials link back to it instead of repeating explanations of QAT and QAD, keeping the tutorials focused on setup and execution. In main QAT/QAD guide make links to all relevant blogposts. Note: MBridge example doc is out of scope for this MR. ### Testing Doc changes only, manual check. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - Did you write any new necessary tests?: N/A docs changes only ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Expanded the QAT/QAD guide with workflows, use cases, and a comparison, including QAD’s use of a frozen BF16 teacher and logit-level loss to recover accuracy after quantization. * Updated README and quick-start navigation to link to the combined QAT/QAD guide; the previous standalone QAT guide now redirects readers there. * Reorganized the LLM QAT tutorial: recipe guidance is now part of the end-to-end example, while trainer examples and Python quantize-and-fine-tune guidance are in Advanced Topics. The tutorial also notes Triton accelerated kernels. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com>
This commit is contained in:
@@ -109,7 +109,7 @@ more fine-grained control on installed dependencies or for alternative docker im
|
||||
| **Technique** | **Description** | **Getting started** | **Examples** |
|
||||
| :------------: | :------------: | :------------: | :------------: |
|
||||
| Post Training Quantization | Compress model size by 2x-4x, speeding up inference while preserving model quality! | \[[Start here](https://nvidia.github.io/Model-Optimizer/guides/1_quantization.html)\] | \[[HF LLMs / VLMs](./examples/hf_ptq/)\] \[[Megatron-Bridge LLMs / VLMs](./examples/megatron_bridge/README.md#post-training-quantization)\] \[[Diffusers](./examples/diffusers/)\] \[[ONNX](./examples/onnx_ptq/)\] \[[Windows](./examples/windows/)\] |
|
||||
| Quantization Aware Training / Distillation | Refine accuracy of quantized models even further with a few training steps! | \[[Start here](https://nvidia.github.io/Model-Optimizer/guides/quantization_aware_training.html)\] | \[[Hugging Face](./examples/llm_qat/)\] \[[Megatron-Bridge](./examples/megatron_bridge/README.md#quantization-aware-distillation-qad)\] |
|
||||
| Quantization Aware Training / Distillation | Refine accuracy of quantized models even further with a few training steps! | \[[Start here](https://nvidia.github.io/Model-Optimizer/guides/quantization_aware_training_and_distillation.html)\] | \[[Hugging Face](./examples/llm_qat/)\] \[[Megatron-Bridge](./examples/megatron_bridge/README.md#quantization-aware-distillation-qad)\] |
|
||||
| Pruning | Reduce your model parameters or memory footprint and accelerate inference by removing unnecessary weights! | \[[Start here](./examples/pruning/README.md)\] | \[[General](./examples/pruning/)\] \[[Megatron-Bridge](./examples/megatron_bridge/README.md#pruning)\] |
|
||||
| Distillation | Reduce deployment model size by teaching small models to behave like larger models! | \[[Start here](https://nvidia.github.io/Model-Optimizer/guides/4_distillation.html)\] | \[[Hugging Face](./examples/llm_distill/)\] \[[Megatron-Bridge](./examples/megatron_bridge/README.md#distillation)\] \[[Megatron-LM](./examples/llm_distill/README.md#knowledge-distillation-kd-in-nvidia-megatron-lm-framework)\] |
|
||||
| Speculative Decoding | Train draft modules to predict extra tokens during inference! | \[[Start here](https://nvidia.github.io/Model-Optimizer/guides/5_speculative_decoding.html)\] | \[[Hugging Face](./examples/speculative_decoding/)\] \[[Megatron-LM](./examples/speculative_decoding#mlm-example)\] |
|
||||
|
||||
@@ -1,125 +1,7 @@
|
||||
.. _quantization-aware-training:
|
||||
:orphan:
|
||||
|
||||
===============================================
|
||||
Quantization-Aware Training and Distillation
|
||||
===============================================
|
||||
============================================
|
||||
|
||||
Quantization-aware training (QAT) and quantization-aware distillation (QAD) recover
|
||||
quality lost when a model is quantized. Both train with simulated quantization
|
||||
enabled, so the resulting checkpoint retains its ModelOpt quantization state for
|
||||
deployment.
|
||||
|
||||
QAT versus QAD
|
||||
==============
|
||||
|
||||
QAT is standard supervised fine-tuning of a quantized model. It uses the usual
|
||||
cross-entropy (CE) loss against labeled data while the quantized forward pass lets
|
||||
the weights adapt to quantization error. Use QAT to adapt a quantized model to a
|
||||
task or dataset.
|
||||
|
||||
QAD uses knowledge distillation: a frozen BF16 teacher guides the quantized student
|
||||
with a logit-level KL-divergence loss. Use QAD after PTQ to recover accuracy lost
|
||||
specifically to quantization. It requires the teacher during training and therefore
|
||||
uses more memory and compute than QAT.
|
||||
|
||||
In both cases, start from a PTQ checkpoint and retain its quantization configuration
|
||||
during training. Choose QAT for task adaptation; choose QAD for quantization-accuracy
|
||||
recovery.
|
||||
|
||||
QAD rationale
|
||||
=============
|
||||
|
||||
`Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
|
||||
<https://arxiv.org/abs/2601.20088>`_ recommends QAD for recovery after aggressive
|
||||
quantization, especially for models that have passed through multi-stage
|
||||
post-training such as SFT, RL, or model merging. The teacher signal makes recovery
|
||||
more robust when training-data quality or coverage is limited.
|
||||
|
||||
For broader background, see `How Quantization-Aware Training Enables Low-Precision
|
||||
Accuracy Recovery
|
||||
<https://developer.nvidia.com/blog/how-quantization-aware-training-enables-low-precision-accuracy-recovery/>`_. The
|
||||
`Nemotron 3.5 Lightning QAD blog
|
||||
<https://developer.nvidia.com/blog/developing-nemotron-3-5-lightning-nvfp4-with-qad-using-nvidia-model-optimizer/>`_
|
||||
discusses the PTQ-to-QAD-to-export workflow and its scale-handling considerations.
|
||||
|
||||
Choose a framework
|
||||
==================
|
||||
|
||||
ModelOpt supports QAT with Hugging Face, Megatron-Bridge, and Megatron-LM.
|
||||
|
||||
.. list-table::
|
||||
:header-rows: 1
|
||||
:widths: 20 40 40
|
||||
|
||||
* - Framework
|
||||
- Advantages
|
||||
- Trade-offs
|
||||
* - Hugging Face
|
||||
- Starts directly from Hugging Face checkpoints and is the simplest path for
|
||||
small to medium models. It supports FSDP2, DDP, and DeepSpeed through
|
||||
Accelerate, with no conversion to Megatron-Core.
|
||||
- Its parallelism is less efficient for large-scale training, so it is better
|
||||
suited to smaller models than the Megatron-based options.
|
||||
* - Megatron-Bridge
|
||||
- Automatically converts Hugging Face models to Megatron-Core and uses
|
||||
Megatron-LM's distributed training stack. It is a convenient scalable
|
||||
workflow without a separate conversion step.
|
||||
- The high-level workflow exposes fewer customization points than working
|
||||
directly in Megatron-LM.
|
||||
* - Megatron-LM
|
||||
- Provides the most control over model configuration, data, parallelism, and
|
||||
training behavior, making it the most customizable option for large-model
|
||||
training.
|
||||
- Requires manually converting the model to Megatron-Core and managing that
|
||||
checkpoint workflow.
|
||||
|
||||
Implementation guides
|
||||
=====================
|
||||
|
||||
Use the framework README as the executable source of truth. Each guide owns its
|
||||
prerequisites, commands, data preparation, distributed topology, and export options.
|
||||
|
||||
* `Hugging Face QAT/QAD Quick Start
|
||||
<https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/llm_qat#quick-start>`_
|
||||
* `Megatron-Bridge README
|
||||
<https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/megatron_bridge>`_
|
||||
* `Megatron-LM ModelOpt post-training documentation
|
||||
<https://github.com/NVIDIA/Megatron-LM/tree/main/examples/post_training/modelopt>`_
|
||||
|
||||
QAD launcher examples
|
||||
===================================
|
||||
|
||||
Model Optimizer includes a `launcher
|
||||
<https://github.com/NVIDIA/Model-Optimizer/tree/main/tools/launcher>`_ for
|
||||
running supported QAD pipelines as one-click commands. Follow the launcher's
|
||||
`Quick Start <https://github.com/NVIDIA/Model-Optimizer/tree/main/tools/launcher#quick-start>`_
|
||||
to set it up.
|
||||
|
||||
The `NVIDIA Nemotron 3.5 Lightning QAD launcher examples
|
||||
<https://github.com/NVIDIA/Model-Optimizer/tree/main/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16>`_
|
||||
provide Slurm pipelines for Megatron-Bridge and Megatron-LM. Both create an NVFP4
|
||||
student with PTQ, distill it from the BF16 teacher, and export a deployable Hugging
|
||||
Face checkpoint.
|
||||
|
||||
Review the example directory's README and the selected YAML before running. Adapt
|
||||
the model and data locations, output paths, and Slurm topology for your environment.
|
||||
|
||||
Run the Megatron-Bridge pipeline:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
cd tools/launcher
|
||||
source .env-slurm
|
||||
uv run launch.py \
|
||||
--yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_qad.yaml \
|
||||
--yes
|
||||
|
||||
Run the Megatron-LM pipeline:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
cd tools/launcher
|
||||
source .env-slurm
|
||||
uv run launch.py \
|
||||
--yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/megatron_lm_qad.yaml \
|
||||
--yes
|
||||
This guide has moved. See
|
||||
:doc:`quantization_aware_training_and_distillation`.
|
||||
|
||||
@@ -0,0 +1,159 @@
|
||||
.. _quantization-aware-training:
|
||||
|
||||
===============================================
|
||||
Quantization-Aware Training and Distillation
|
||||
===============================================
|
||||
|
||||
Quantization-aware training (QAT) and quantization-aware distillation (QAD) recover
|
||||
quality lost when a model is quantized. Both train with simulated quantization
|
||||
enabled, so the resulting checkpoint retains its ModelOpt quantization state for
|
||||
deployment.
|
||||
|
||||
**Quantization Aware Training (QAT)** inserts simulated quantization operations
|
||||
into the model graph and then fine-tunes the model so its weights learn to
|
||||
compensate for quantization error. During training, quantization scales are frozen
|
||||
while weights are updated. QAT is a general technique — it learns from labeled
|
||||
data on a quantized model using the usual cross-entropy (CE) loss.
|
||||
|
||||
**Quantization Aware Distillation (QAD)** is a special case of QAT that uses a
|
||||
frozen BF16 teacher (typically the original unquantized model) to guide the
|
||||
quantized student via a logit-level KL-divergence loss. QAD is a **pure accuracy
|
||||
recovery technique** — its goal is to recover accuracy lost from quantization, not
|
||||
to teach the model a new task. It requires the teacher during training and
|
||||
therefore uses more memory and compute than QAT.
|
||||
|
||||
In both cases, start from a PTQ checkpoint and retain its quantization configuration
|
||||
during training.
|
||||
|
||||
When to Use QAT vs QAD
|
||||
======================
|
||||
|
||||
.. list-table::
|
||||
:header-rows: 1
|
||||
:stub-columns: 1
|
||||
:widths: 20 40 40
|
||||
|
||||
* -
|
||||
- QAT (without distillation)
|
||||
- QAD (with distillation)
|
||||
* - What it does
|
||||
- Fine-tunes a quantized model on labeled data
|
||||
- Recovers quantization accuracy using the original model as teacher
|
||||
* - When to use
|
||||
- The model is already quantized and you want to fine-tune it for a
|
||||
**new task** (e.g., fine-tuning a `GPT-OSS
|
||||
<https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/gpt-oss>`_
|
||||
quantized checkpoint)
|
||||
- You want the **best possible accuracy recovery** after quantization
|
||||
* - Recommended workflow
|
||||
- Start from a quantized checkpoint, fine-tune with task-specific data
|
||||
- Full-precision fine-tuning first, then QAD to recover quantization loss
|
||||
|
||||
QAD rationale
|
||||
=============
|
||||
|
||||
**QAD is Model Optimizer's recommended strategy for accuracy recovery after
|
||||
quantization.**
|
||||
|
||||
In our experiments, full-precision fine-tuning followed by QAD
|
||||
delivers the best accuracy, especially at aggressive quantization levels (e.g.,
|
||||
NVFP4).
|
||||
|
||||
`Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
|
||||
<https://arxiv.org/abs/2601.20088>`_ recommends QAD for recovery after aggressive
|
||||
quantization, especially for models that have passed through multi-stage
|
||||
post-training such as SFT, RL, or model merging. The teacher signal makes recovery
|
||||
more robust when training-data quality or coverage is limited.
|
||||
|
||||
The optimal balance between QAT and QAD for a given model and task is an
|
||||
active area of research.
|
||||
|
||||
To learn more, read the `QAT/QAD blog post
|
||||
<https://developer.nvidia.com/blog/how-quantization-aware-training-enables-low-precision-accuracy-recovery/>`_.
|
||||
|
||||
The
|
||||
`Nemotron 3.5 Lightning QAD blog
|
||||
<https://developer.nvidia.com/blog/developing-nemotron-3-5-lightning-nvfp4-with-qad-using-nvidia-model-optimizer/>`_
|
||||
discusses the PTQ-to-QAD-to-export workflow and its scale-handling considerations.
|
||||
|
||||
Choose a framework
|
||||
==================
|
||||
|
||||
ModelOpt supports QAT with Hugging Face, Megatron-Bridge, and Megatron-LM.
|
||||
|
||||
.. list-table::
|
||||
:header-rows: 1
|
||||
:widths: 20 40 40
|
||||
|
||||
* - Framework
|
||||
- Advantages
|
||||
- Trade-offs
|
||||
* - Hugging Face
|
||||
- Starts directly from Hugging Face checkpoints and is the simplest path for
|
||||
small to medium models. It supports FSDP2, DDP, and DeepSpeed through
|
||||
Accelerate, with no conversion to Megatron-Core.
|
||||
- Its parallelism is less efficient for large-scale training, so it is better
|
||||
suited to smaller models than the Megatron-based options.
|
||||
* - Megatron-Bridge
|
||||
- Automatically converts Hugging Face models to Megatron-Core and uses
|
||||
Megatron-LM's distributed training stack. It is a convenient scalable
|
||||
workflow without a separate conversion step.
|
||||
- The high-level workflow exposes fewer customization points than working
|
||||
directly in Megatron-LM.
|
||||
* - Megatron-LM
|
||||
- Provides the most control over model configuration, data, parallelism, and
|
||||
training behavior, making it the most customizable option for large-model
|
||||
training.
|
||||
- Requires manually converting the model to Megatron-Core and managing that
|
||||
checkpoint workflow.
|
||||
|
||||
Implementation guides
|
||||
=====================
|
||||
|
||||
Use the framework README as the executable source of truth. Each guide owns its
|
||||
prerequisites, commands, data preparation, distributed topology, and export options.
|
||||
|
||||
* `Hugging Face QAT/QAD Quick Start
|
||||
<https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/llm_qat#quick-start>`_
|
||||
* `Megatron-Bridge README
|
||||
<https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/megatron_bridge>`_
|
||||
* `Megatron-LM ModelOpt post-training documentation
|
||||
<https://github.com/NVIDIA/Megatron-LM/tree/main/examples/post_training/modelopt>`_
|
||||
|
||||
QAD launcher examples
|
||||
===================================
|
||||
|
||||
Model Optimizer includes a `launcher
|
||||
<https://github.com/NVIDIA/Model-Optimizer/tree/main/tools/launcher>`_ for
|
||||
running supported QAD pipelines as one-click commands. Follow the launcher's
|
||||
`Quick Start <https://github.com/NVIDIA/Model-Optimizer/tree/main/tools/launcher#quick-start>`_
|
||||
to set it up.
|
||||
|
||||
The `NVIDIA Nemotron 3.5 Lightning QAD launcher examples
|
||||
<https://github.com/NVIDIA/Model-Optimizer/tree/main/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16>`_
|
||||
provide Slurm pipelines for Megatron-Bridge and Megatron-LM. Both create an NVFP4
|
||||
student with PTQ, distill it from the BF16 teacher, and export a deployable Hugging
|
||||
Face checkpoint.
|
||||
|
||||
Review the example directory's README and the selected YAML before running. Adapt
|
||||
the model and data locations, output paths, and Slurm topology for your environment.
|
||||
|
||||
Run the Megatron-Bridge pipeline:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
cd tools/launcher
|
||||
source .env-slurm
|
||||
uv run launch.py \
|
||||
--yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_qad.yaml \
|
||||
--yes
|
||||
|
||||
Run the Megatron-LM pipeline:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
cd tools/launcher
|
||||
source .env-slurm
|
||||
uv run launch.py \
|
||||
--yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/megatron_lm_qad.yaml \
|
||||
--yes
|
||||
@@ -86,7 +86,7 @@ Release notes, technical updates, examples, and deployment stories from the Mode
|
||||
Quick Start: PTQ - ONNX <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/onnx_ptq>
|
||||
Quick Start: PTQ - PyTorch to ONNX <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/torch_onnx>
|
||||
Quick Start: PTQ - Windows <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/windows>
|
||||
Quick Start: QAT and QAD <guides/quantization_aware_training>
|
||||
Quick Start: QAT and QAD <guides/quantization_aware_training_and_distillation>
|
||||
Quick Start: Pruning <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/pruning>
|
||||
Quick Start: Distillation <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/llm_distill>
|
||||
Quick Start: Speculative Decoding <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/speculative_decoding>
|
||||
@@ -99,7 +99,7 @@ Release notes, technical updates, examples, and deployment stories from the Mode
|
||||
|
||||
guides/0_support_matrix
|
||||
guides/1_quantization
|
||||
guides/quantization_aware_training
|
||||
guides/quantization_aware_training_and_distillation
|
||||
guides/2_save_load
|
||||
guides/3_pruning
|
||||
guides/4_distillation
|
||||
|
||||
+82
-103
@@ -1,22 +1,21 @@
|
||||
# Quantization Aware Training (QAT) and Distillation (QAD)
|
||||
|
||||
Quantization Aware Training (QAT) improves model accuracy beyond post-training quantization (PTQ) at low precisions (e.g., INT4, FP4 on [NVIDIA Blackwell](https://www.nvidia.com/en-us/data-center/technologies/blackwell-architecture/)). Quantization Aware Distillation (QAD) further improves accuracy by using the original full-precision model as a teacher.
|
||||
This tutorial shows how to run QAT and QAD with Hugging Face Transformers: set up the environment, quantize a model, train it, evaluate the checkpoint, and export it for deployment.
|
||||
|
||||
For background on how QAT enables low-precision accuracy recovery, see the [QAT/QAD blog post](https://developer.nvidia.com/blog/how-quantization-aware-training-enables-low-precision-accuracy-recovery/).
|
||||
For background on QAT and QAD and help choosing between Hugging Face, Megatron Bridge, and Megatron-LM, start with the [QAT/QAD guide](https://nvidia.github.io/Model-Optimizer/guides/quantization_aware_training_and_distillation.html).
|
||||
|
||||
<div align="center">
|
||||
|
||||
| **Section** | **Description** | **Link** | **Docs** |
|
||||
| :---: | :---: | :---: | :---: |
|
||||
| Quick Start | Prerequisites and setup | \[[Link](#quick-start)\] | |
|
||||
| End-to-End Example | Run QAT/QAD in 3 steps: quantize, train, export | \[[Link](#run-end-to-end-qatqad-example)\] | |
|
||||
| Arguments | Full CLI/YAML argument reference | \[[Link](ARGUMENTS.md)\] | |
|
||||
| Background | How QAT/QAD work and when to use each | \[[Link](#background)\] | \[[docs](https://nvidia.github.io/Model-Optimizer/guides/1_quantization.html)\] |
|
||||
| Support Matrix | Supported models, quantization formats, and backends | \[[Link](#support-matrix)\] | |
|
||||
| QLoRA | Model training with reduced GPU memory | \[[Link](#qlora-real-quantization)\] | |
|
||||
| Advanced Topics | FSDP2 config, YAML options | \[[Link](#advanced-topics)\] | |
|
||||
| Results | Accuracy benchmarks | \[[Link](#results)\] | |
|
||||
| Resources | Extra links and references | \[[Link](#resources)\] | |
|
||||
| **Section** | **Description** | **Link** |
|
||||
| :---: | :---: | :---: |
|
||||
| Quick Start | Prerequisites and setup | \[[Link](#quick-start)\] |
|
||||
| End-to-End Example | Run QAT/QAD in 3 steps: quantize, train, export | \[[Link](#run-end-to-end-qatqad-example)\] |
|
||||
| Arguments | Full CLI/YAML argument reference | \[[Link](ARGUMENTS.md)\] |
|
||||
| Support Matrix | Supported models, quantization formats, and backends | \[[Link](#support-matrix)\] |
|
||||
| QLoRA | Model training with reduced GPU memory | \[[Link](#qlora-real-quantization)\] |
|
||||
| Advanced Topics | Trainer APIs, FSDP2 config, YAML options | \[[Link](#advanced-topics)\] |
|
||||
| Results | Accuracy benchmarks | \[[Link](#results)\] |
|
||||
| Resources | Extra links and references | \[[Link](#resources)\] |
|
||||
|
||||
</div>
|
||||
|
||||
@@ -35,11 +34,24 @@ pip install -r examples/llm_qat/requirements.txt
|
||||
|
||||
The Qwen3-8B example below requires a minimum of **2 x 80GB GPUs**.
|
||||
|
||||
> ModelOpt provides accelerated quantization kernels using Triton for NVFP4 QAT. See the [installation guide](https://nvidia.github.io/Model-Optimizer/getting_started/_installation_for_Linux.html#accelerated-quantization-with-triton-kernels).
|
||||
|
||||
## Run End-to-End QAT/QAD Example
|
||||
|
||||
All arguments can be set via YAML, CLI, or both (CLI overrides YAML). See
|
||||
[ARGUMENTS.md](ARGUMENTS.md), `--help`, and [Configuration](#advanced-configuration).
|
||||
|
||||
### Quantization Recipes
|
||||
|
||||
Recipes are declarative YAML files that specify the quantization configuration. Built-in recipes are available in [`modelopt_recipes/`](../../modelopt_recipes/):
|
||||
|
||||
```sh
|
||||
# From the Model-Optimizer repository root, list available built-in recipes
|
||||
ls modelopt_recipes/general/ptq/
|
||||
```
|
||||
|
||||
See [custom calibration](https://nvidia.github.io/Model-Optimizer/guides/_pytorch_quantization.html#advanced-configuration-creation) for creating your own recipe.
|
||||
|
||||
### QAT
|
||||
|
||||
Quantize, fine-tune on labeled data, and export:
|
||||
@@ -99,96 +111,6 @@ Exported checkpoints can be deployed on [TensorRT-LLM](https://github.com/NVIDIA
|
||||
> [!TIP]
|
||||
> For more performant QAD, please refer to [examples/megatron_bridge/README.md](../megatron_bridge/README.md) for example scripts for PTQ / QAD with Megatron-Bridge which is generally more performant than the Hugging Face scripts.
|
||||
|
||||
## Background
|
||||
|
||||
### What is QAT?
|
||||
|
||||
**Quantization Aware Training (QAT)** inserts simulated quantization operations into the model graph and then fine-tunes the model so its weights learn to compensate for quantization error. During training, quantization scales are frozen while weights are updated. QAT is a general technique — it learns from labeled data on a quantized model.
|
||||
|
||||
```python
|
||||
import modelopt.torch.quantization as mtq
|
||||
from modelopt.recipe import load_recipe
|
||||
|
||||
# 1. Load a quantization recipe
|
||||
recipe = load_recipe("general/ptq/nvfp4_default-kv_fp8")
|
||||
|
||||
# 2. Quantize the model in-place
|
||||
model = mtq.quantize(model, recipe.quantize, forward_loop)
|
||||
|
||||
# 3. Fine-tune the quantized model
|
||||
trainer.train()
|
||||
trainer.save_model()
|
||||
```
|
||||
|
||||
> ModelOpt provides accelerated quantization kernels using Triton for NVFP4 QAT. See the [installation guide](https://nvidia.github.io/Model-Optimizer/getting_started/_installation_for_Linux.html#accelerated-quantization-with-triton-kernels).
|
||||
|
||||
### What is QAD?
|
||||
|
||||
**Quantization Aware Distillation (QAD)** is a special case of QAT that uses a teacher model (typically the original unquantized model) to guide the quantized student via a distillation loss. QAD is a **pure accuracy recovery technique** — its goal is to recover accuracy lost from quantization, not to teach the model a new task.
|
||||
|
||||
To learn more, read the [QAT/QAD blog post](https://developer.nvidia.com/blog/how-quantization-aware-training-enables-low-precision-accuracy-recovery/).
|
||||
|
||||
### When to Use QAT vs QAD
|
||||
|
||||
| | **QAT** (without distillation) | **QAD** (with distillation) |
|
||||
|-|---------|----------------------|
|
||||
| **What it does** | Fine-tunes a quantized model on labeled data | Recovers quantization accuracy using the original model as teacher |
|
||||
| **When to use** | The model is already quantized and you want to fine-tune it for a **new task** (e.g., fine-tuning a [GPT-OSS](../gpt-oss/) quantized checkpoint) | You want the **best possible accuracy recovery** after quantization |
|
||||
| **Recommended workflow** | Start from a quantized checkpoint, fine-tune with task-specific data | Full-precision fine-tuning first, then QAD to recover quantization loss |
|
||||
|
||||
**QAD is Model Optimizer's recommended strategy for accuracy recovery after quantization.** In our experiments, full-precision fine-tuning followed by QAD delivers the best accuracy, especially at aggressive quantization levels (e.g., NVFP4). The optimal balance between QAT and QAD for a given model and task is an active area of research.
|
||||
|
||||
### Using `QATTrainer` and `QADTrainer`
|
||||
|
||||
`QATTrainer` is a drop-in replacement for HuggingFace's `Trainer` that handles quantization-aware training seamlessly with various distributed backends (FSDP2, DeepSpeed, DDP):
|
||||
|
||||
```python
|
||||
from modelopt.torch.quantization.plugins.transformers_trainer import QATTrainer
|
||||
|
||||
trainer = QATTrainer(
|
||||
model=model, # pre-quantized model
|
||||
processing_class=tokenizer,
|
||||
args=training_args,
|
||||
**data_module,
|
||||
)
|
||||
trainer.train()
|
||||
trainer.save_model()
|
||||
```
|
||||
|
||||
`QADTrainer` extends `QATTrainer` with distillation. Pass the teacher model and a `DistillArguments` instance:
|
||||
|
||||
```python
|
||||
from modelopt.torch.distill.plugins.huggingface import DistillArguments
|
||||
from modelopt.torch.quantization.plugins.transformers_trainer import QADTrainer
|
||||
|
||||
distill_args = DistillArguments(
|
||||
distill=True,
|
||||
teacher_model="Qwen/Qwen3-8B",
|
||||
criterion="logits_loss",
|
||||
)
|
||||
|
||||
trainer = QADTrainer(
|
||||
model=model, # pre-quantized model
|
||||
processing_class=tokenizer,
|
||||
args=training_args,
|
||||
distill_args=distill_args,
|
||||
**data_module,
|
||||
)
|
||||
trainer.train()
|
||||
trainer.save_model()
|
||||
```
|
||||
|
||||
### Quantization Recipes
|
||||
|
||||
Recipes are declarative YAML files that specify the quantization configuration. Built-in recipes are available in [`modelopt_recipes/`](../../modelopt_recipes/):
|
||||
|
||||
```sh
|
||||
# List available built-in recipes
|
||||
ls modelopt_recipes/general/ptq/
|
||||
```
|
||||
|
||||
See [custom calibration](https://nvidia.github.io/Model-Optimizer/guides/_pytorch_quantization.html#advanced-configuration-creation) for creating your own recipe.
|
||||
|
||||
## Support Matrix
|
||||
|
||||
### Supported Models
|
||||
@@ -262,6 +184,63 @@ vllm serve qwen3-8b-fp4-qlora-hf/base_model --enable-lora \
|
||||
|
||||
## Advanced Topics
|
||||
|
||||
### Quantize and Fine-Tune with Python
|
||||
|
||||
```python
|
||||
import modelopt.torch.quantization as mtq
|
||||
from modelopt.recipe import load_recipe
|
||||
|
||||
# 1. Load a quantization recipe
|
||||
recipe = load_recipe("general/ptq/nvfp4_default-kv_fp8")
|
||||
|
||||
# 2. Quantize the model in-place
|
||||
model = mtq.quantize(model, recipe.quantize, forward_loop)
|
||||
|
||||
# 3. Fine-tune the quantized model
|
||||
trainer.train()
|
||||
trainer.save_model()
|
||||
```
|
||||
|
||||
### Using `QATTrainer` and `QADTrainer`
|
||||
|
||||
`QATTrainer` is a drop-in replacement for HuggingFace's `Trainer` that handles quantization-aware training seamlessly with various distributed backends (FSDP2, DeepSpeed, DDP):
|
||||
|
||||
```python
|
||||
from modelopt.torch.quantization.plugins.transformers_trainer import QATTrainer
|
||||
|
||||
trainer = QATTrainer(
|
||||
model=model, # pre-quantized model
|
||||
processing_class=tokenizer,
|
||||
args=training_args,
|
||||
**data_module,
|
||||
)
|
||||
trainer.train()
|
||||
trainer.save_model()
|
||||
```
|
||||
|
||||
`QADTrainer` extends `QATTrainer` with distillation. Pass the teacher model and a `DistillArguments` instance:
|
||||
|
||||
```python
|
||||
from modelopt.torch.distill.plugins.huggingface import DistillArguments
|
||||
from modelopt.torch.quantization.plugins.transformers_trainer import QADTrainer
|
||||
|
||||
distill_args = DistillArguments(
|
||||
distill=True,
|
||||
teacher_model="Qwen/Qwen3-8B",
|
||||
criterion="logits_loss",
|
||||
)
|
||||
|
||||
trainer = QADTrainer(
|
||||
model=model, # pre-quantized model
|
||||
processing_class=tokenizer,
|
||||
args=training_args,
|
||||
distill_args=distill_args,
|
||||
**data_module,
|
||||
)
|
||||
trainer.train()
|
||||
trainer.save_model()
|
||||
```
|
||||
|
||||
<details>
|
||||
<summary><b>FSDP2 and Model-Specific Layer Wrapping</b></summary>
|
||||
|
||||
|
||||
Reference in New Issue
Block a user