mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
refactor(examples): rename llm_ptq → hf_ptq (symlink for back-compat) (#1759)
## What does this PR do? **Type of change:** refactor / deprecation (examples) Follow-up to #1705 (which consolidated `examples/vlm_ptq` into `examples/llm_ptq`). Since that example now covers Hugging Face **LLM and VLM** PTQ, the `llm_ptq` name is a misnomer. This renames the directory to `examples/hf_ptq` and leaves a relative symlink `examples/llm_ptq → hf_ptq` so existing paths/commands keep working during a deprecation window. Requested by @kevalmorabia97 on #1705 (with the symlink-for-back-compat approach), targeted for the **same 0.46 release** as the consolidation. ### Changes - `git mv examples/llm_ptq → examples/hf_ptq` and `tests/examples/llm_ptq → tests/examples/hf_ptq` (the CI runner maps the matrix name to both `examples/<name>` and `tests/examples/<name>`). - Add a tracked back-compat symlink `examples/llm_ptq → hf_ptq`. - Update CI matrices and all repo **path references** (docs, READMEs, agent skills, launcher/debugger tools, tests) from `llm_ptq` to `hf_ptq`. - Keep Python identifiers / test-util module names (`run_llm_ptq_command`, `llm_ptq_utils`) — they name the LLM-PTQ task, not the directory. - Preserve the CODEOWNERS team slug (`modelopt-examples-llm_ptq-codeowners`) and historical CHANGELOG entries; add a CHANGELOG deprecation note. ### Back-compat caveats (inherent to git directory symlinks) - ✅ Linux/macOS CLI usage and Python `cwd`/pytest resolution work through the symlink. - ⚠️ Windows git checkouts don't materialize symlinks by default (low impact — this example is Linux-only in practice). - ⚠️ GitHub web doesn't follow directory symlinks, so legacy external deep-links to `examples/llm_ptq/...` won't navigate in. All **internal** references are repointed to `hf_ptq`, so the symlink is only for legacy external/CLI use. ### Usage (unchanged via symlink) ```bash # New canonical path cd examples/hf_ptq scripts/huggingface_example.sh --model <hf_model> --quant fp8 # Old path still works (forwards via symlink) cd examples/llm_ptq && scripts/huggingface_example.sh --model <hf_model> --quant fp8 ``` ### Testing - `bash -n` on moved/edited shell scripts (new path + via symlink). - `py_compile` on moved/edited Python; test re-export shim repointed to `examples/hf_ptq/example_utils`. - Verified git tracks `examples/llm_ptq` as a single symlink (mode 120000), not a duplicated tree (no pre-commit / pytest double-processing). - `pre-commit run` on all changed files passes. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (relative symlink keeps `examples/llm_ptq` paths valid; see caveats above) - Did you write any new necessary tests?: N/A (pure rename; existing tests moved with the dir) - Did you update Changelog?: ✅ ### Additional Information Follow-up (later release): remove the `examples/llm_ptq` symlink once external references have migrated. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * PTQ guidance now directs to the unified Hugging Face PTQ flow, including VLM quantization via the shared `--vlm` entry point. * **Documentation** * Updated README and guide links, references, and command snippets to use `hf_ptq` (replacing `llm_ptq`). * Deprecated and consolidated `vlm_ptq` into `hf_ptq`; removed VILA/NVILA coverage from the Hugging Face PTQ examples. * **Bug Fixes** * Improved detection and routing so local/manual setup uses the correct PTQ source. * **Tests / Chores** * CI and example tests updated to run the `hf_ptq` variants. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -5,7 +5,7 @@ Common detection for all ModelOpt skills. After this, you know what's available.
|
||||
## Env-1. Get ModelOpt source
|
||||
|
||||
```bash
|
||||
ls examples/llm_ptq/hf_ptq.py 2>/dev/null && echo "Source found"
|
||||
ls examples/hf_ptq/hf_ptq.py 2>/dev/null && echo "Source found"
|
||||
```
|
||||
|
||||
If not found: `git clone https://github.com/NVIDIA/Model-Optimizer.git && cd Model-Optimizer`
|
||||
|
||||
@@ -62,4 +62,4 @@ This matrix covers officially validated combinations. For unlisted models:
|
||||
|
||||
- **NVFP4 inference requires Blackwell GPUs** (B100, B200, GB200). Hopper can run FP4 calibration but not inference.
|
||||
- INT4_AWQ and W4A8_AWQ are only supported by TRT-LLM (not vLLM or SGLang).
|
||||
- Source: `examples/llm_ptq/README.md` and `docs/source/deployment/3_unified_hf.rst`
|
||||
- Source: `examples/hf_ptq/README.md` and `docs/source/deployment/3_unified_hf.rst`
|
||||
|
||||
@@ -5,7 +5,7 @@ description: This skill should be used when the user asks to "quantize a model",
|
||||
|
||||
# ModelOpt Post-Training Quantization
|
||||
|
||||
Produce a quantized checkpoint from a pretrained model. **Read `examples/llm_ptq/README.md` first** — it has the support matrix, CLI flags, and accuracy guidance.
|
||||
Produce a quantized checkpoint from a pretrained model. **Read `examples/hf_ptq/README.md` first** — it has the support matrix, CLI flags, and accuracy guidance.
|
||||
|
||||
## Step 1 — Environment
|
||||
|
||||
@@ -19,7 +19,7 @@ Read `skills/common/environment-setup.md` and `skills/common/workspace-managemen
|
||||
|
||||
## Step 2 — Is the model supported?
|
||||
|
||||
Check the support table in `examples/llm_ptq/README.md` for verified HF models.
|
||||
Check the support table in `examples/hf_ptq/README.md` for verified HF models.
|
||||
|
||||
- **Listed** → supported, use `hf_ptq.py` (step 4A/4B)
|
||||
- **Not listed** → read `references/unsupported-models.md` to determine if `hf_ptq.py` can still work or if a custom script is needed (step 4C)
|
||||
@@ -53,7 +53,7 @@ ls modelopt_recipes/huggingface/<model_type>/ptq/ 2>/dev/null # per-arch; <mode
|
||||
|
||||
If a model-specific recipe exists, prefer `--recipe <path>` — but **inspect its include/exclude patterns** rather than assuming (e.g. for VLMs, confirm the vision tower is actually excluded).
|
||||
|
||||
**If no model-specific recipe**, choose a format based on GPU (details in `examples/llm_ptq/README.md`):
|
||||
**If no model-specific recipe**, choose a format based on GPU (details in `examples/hf_ptq/README.md`):
|
||||
|
||||
- **Blackwell** (B100/B200/GB200): `nvfp4` variants
|
||||
- **Hopper** (H100/H200) or older: `fp8` or `int4_awq`
|
||||
@@ -90,9 +90,9 @@ In README table? ─→ YES ──→ SLURM (local or remote)? ──→ LAUNCHE
|
||||
|
||||
```bash
|
||||
pip install --no-build-isolation "nvidia-modelopt[hf]"
|
||||
pip install -r examples/llm_ptq/requirements.txt
|
||||
pip install -r examples/hf_ptq/requirements.txt
|
||||
|
||||
python examples/llm_ptq/hf_ptq.py \
|
||||
python examples/hf_ptq/hf_ptq.py \
|
||||
--pyt_ckpt_path <model> \
|
||||
--qformat <format> \
|
||||
--calib_size 512 \
|
||||
@@ -105,7 +105,7 @@ For remote: use `remote_run` from `remote_exec.sh` (see `skills/common/remote-ex
|
||||
|
||||
### 4B — Launcher: supported model on SLURM or local Docker
|
||||
|
||||
Write a YAML config using `common/hf_ptq/hf_ptq.sh`. See `references/launcher-guide.md` for the full template.
|
||||
Write a YAML config using `common/hf/ptq.sh`. See `references/launcher-guide.md` for the full template.
|
||||
|
||||
```bash
|
||||
cd tools/launcher
|
||||
@@ -179,7 +179,7 @@ Report the gate result before moving on. The report must include source size, ou
|
||||
| `skills/common/remote-execution.md` | Step 4A/4C only, if target is remote |
|
||||
| `skills/common/slurm-setup.md` | Step 4A/4C only, if using SLURM manually (not launcher) |
|
||||
| `references/slurm-setup-ptq.md` | Step 4A/4C only, PTQ-specific SLURM (container, GPU sizing, FSDP2) |
|
||||
| `examples/llm_ptq/README.md` | Step 3: support matrix, CLI flags, accuracy |
|
||||
| `examples/hf_ptq/README.md` | Step 3: support matrix, CLI flags, accuracy |
|
||||
| `modelopt/torch/quantization/config.py` | Step 3: format definitions |
|
||||
| `modelopt/torch/export/model_utils.py` | Step 4C: TRT-LLM export type mapping |
|
||||
| `modelopt_recipes/` | Step 3: pre-built recipes |
|
||||
|
||||
@@ -7,7 +7,7 @@ monitoring), see `skills/common/slurm-setup.md`.
|
||||
|
||||
## 1. Container
|
||||
|
||||
Get the recommended image version from `examples/llm_ptq/README.md`, then look for an existing `.sqsh` file:
|
||||
Get the recommended image version from `examples/hf_ptq/README.md`, then look for an existing `.sqsh` file:
|
||||
|
||||
```bash
|
||||
ls *.sqsh ../*.sqsh ~/containers/*.sqsh 2>/dev/null
|
||||
@@ -63,17 +63,17 @@ pip install -U transformers --no-deps
|
||||
|
||||
Estimate GPU count from model size and available GPU memory. `hf_ptq.py` uses `device_map="auto"` so it fills GPUs automatically — request only as many as needed.
|
||||
|
||||
For multi-node PTQ (200B+ params), use `examples/llm_ptq/multinode_ptq.py` with FSDP2 and accelerate:
|
||||
For multi-node PTQ (200B+ params), use `examples/hf_ptq/multinode_ptq.py` with FSDP2 and accelerate:
|
||||
|
||||
```bash
|
||||
accelerate launch \
|
||||
--config_file examples/llm_ptq/fsdp2.yaml \
|
||||
--config_file examples/hf_ptq/fsdp2.yaml \
|
||||
--num_machines $NUM_NODES \
|
||||
--num_processes $((NUM_NODES * GPUS_PER_NODE)) \
|
||||
--main_process_ip $MASTER_ADDR \
|
||||
--main_process_port $MASTER_PORT \
|
||||
--machine_rank $SLURM_PROCID \
|
||||
examples/llm_ptq/multinode_ptq.py \
|
||||
examples/hf_ptq/multinode_ptq.py \
|
||||
--pyt_ckpt_path <model> \
|
||||
--qformat <format> \
|
||||
--export_path <output>
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Handling Unlisted Models
|
||||
|
||||
The model is not in the verified support table (`examples/llm_ptq/README.md`). This does NOT mean it won't work — ModelOpt auto-detects standard HF modules (linear layers, attention, MoE blocks with `gate`+`experts`). Many unlisted models work with `hf_ptq.py` out of the box.
|
||||
The model is not in the verified support table (`examples/hf_ptq/README.md`). This does NOT mean it won't work — ModelOpt auto-detects standard HF modules (linear layers, attention, MoE blocks with `gate`+`experts`). Many unlisted models work with `hf_ptq.py` out of the box.
|
||||
|
||||
Follow the investigation steps below to determine if `hf_ptq.py` works or if patches are needed.
|
||||
|
||||
@@ -147,7 +147,7 @@ class QuantCustomModule(OriginalModule):
|
||||
| Fused 2D weights (experts stacked in rows) | Two-level expansion | `_QuantDbrxExpertGLU` |
|
||||
| Fused weights + `forward(x, expert_id)` | Expand + reconstruct on export | `_QuantMoELinear` (Step3.5) |
|
||||
|
||||
For the full guide, see `examples/llm_ptq/moe.md`.
|
||||
For the full guide, see `examples/hf_ptq/README.md`.
|
||||
|
||||
**Critical: always check the weight layout.** `nn.Linear` expects `(out_features, in_features)` — the last dimension must be `in_features`. If the fused tensor is `(num_experts, in_dim, out_dim)`, you must transpose (`.T`) when copying. Getting this wrong silently corrupts quantization scales. Inspect the original forward pass to determine which dimension is which.
|
||||
|
||||
|
||||
+1
-2
@@ -48,7 +48,7 @@ modelopt_recipes @NVIDIA/modelopt-recipes-codeowners
|
||||
/examples/gpt-oss @NVIDIA/modelopt-examples-gpt-oss-codeowners
|
||||
/examples/llm_distill @NVIDIA/modelopt-torch-distill-codeowners
|
||||
/examples/llm_eval @NVIDIA/modelopt-examples-llm_ptq-codeowners
|
||||
/examples/llm_ptq @NVIDIA/modelopt-examples-llm_ptq-codeowners
|
||||
/examples/hf_ptq @NVIDIA/modelopt-examples-llm_ptq-codeowners
|
||||
/examples/llm_qat @NVIDIA/modelopt-examples-llm_qat-codeowners
|
||||
/examples/llm_sparsity @NVIDIA/modelopt-torch-sparsity-codeowners
|
||||
/examples/megatron_bridge @NVIDIA/modelopt-examples-megatron-codeowners
|
||||
@@ -59,7 +59,6 @@ modelopt_recipes @NVIDIA/modelopt-recipes-codeowners
|
||||
/examples/specdec_bench @NVIDIA/modelopt-torch-speculative-codeowners
|
||||
/examples/speculative_decoding @NVIDIA/modelopt-torch-speculative-codeowners
|
||||
/examples/torch_onnx @NVIDIA/modelopt-onnx-codeowners
|
||||
/examples/vlm_ptq @NVIDIA/modelopt-examples-vlm-codeowners
|
||||
/examples/vllm_serve @NVIDIA/modelopt-examples-llm_ptq-codeowners
|
||||
/examples/windows @NVIDIA/modelopt-windows-codeowners
|
||||
|
||||
|
||||
@@ -9,7 +9,7 @@ on:
|
||||
required: true
|
||||
type: string
|
||||
example:
|
||||
description: "Example name to test (e.g. 'llm_ptq')"
|
||||
description: "Example name to test (e.g. 'hf_ptq')"
|
||||
required: true
|
||||
type: string
|
||||
timeout_minutes:
|
||||
|
||||
@@ -55,7 +55,7 @@ jobs:
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
example: [llm_ptq]
|
||||
example: [hf_ptq]
|
||||
uses: ./.github/workflows/_example_tests_runner.yml
|
||||
secrets: inherit
|
||||
with:
|
||||
@@ -69,7 +69,7 @@ jobs:
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
example: [llm_eval, llm_ptq]
|
||||
example: [llm_eval, hf_ptq]
|
||||
uses: ./.github/workflows/_example_tests_runner.yml
|
||||
secrets: inherit
|
||||
with:
|
||||
|
||||
+3
-2
@@ -11,8 +11,9 @@ Changelog
|
||||
|
||||
**Deprecations**
|
||||
|
||||
- Consolidated ``examples/vlm_ptq`` into ``examples/llm_ptq``. Vision-language model PTQ now shares the ``hf_ptq.py`` entry point and ``scripts/huggingface_example.sh``; pass ``--vlm`` to run the TensorRT-LLM multimodal quickstart smoke test. The ``examples/vlm_ptq/scripts/huggingface_example.sh`` entry point is deprecated: it now prints a warning and forwards to the ``llm_ptq`` script with ``--vlm``, and will be removed in a future release. See `examples/llm_ptq/README.md <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/llm_ptq#vlm-quantization>`__.
|
||||
- Dropped VILA / NVILA vision-language model support in ``examples/llm_ptq``. VILA's modeling code requires ``transformers<=4.50.0``, which conflicts with ModelOpt's minimum supported ``transformers`` version. The VILA-specific bootstrap (repo clone, ``requirements-vila.txt``) and loading paths in ``example_utils.py`` have been removed.
|
||||
- Renamed ``examples/llm_ptq`` to ``examples/hf_ptq`` to reflect that it covers Hugging Face LLM **and** VLM PTQ. A relative symlink ``examples/llm_ptq`` -> ``hf_ptq`` keeps existing paths and commands working; it will be removed in a future release. Please update references to the new ``examples/hf_ptq`` path.
|
||||
- Consolidated ``examples/vlm_ptq`` into ``examples/hf_ptq``. Vision-language model PTQ now shares the ``hf_ptq.py`` entry point and ``scripts/huggingface_example.sh``; pass ``--vlm`` to run the TensorRT-LLM multimodal quickstart smoke test. The ``examples/vlm_ptq/scripts/huggingface_example.sh`` entry point is deprecated: it now prints a warning and forwards to the ``hf_ptq`` script with ``--vlm``, and will be removed in a future release. See `examples/hf_ptq/README.md <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/hf_ptq#vlm-quantization>`__.
|
||||
- Dropped VILA / NVILA vision-language model support in ``examples/hf_ptq``. VILA's modeling code requires ``transformers<=4.50.0``, which conflicts with ModelOpt's minimum supported ``transformers`` version. The VILA-specific bootstrap (repo clone, ``requirements-vila.txt``) and loading paths in ``example_utils.py`` have been removed.
|
||||
|
||||
**New Features**
|
||||
|
||||
|
||||
@@ -30,7 +30,7 @@ Model Optimizer is also integrated with [NVIDIA Megatron-Bridge](https://github.
|
||||
- [2026/05/13] [**Puzzletron**](./examples/puzzletron): A new algorithm for heterogeneous pruning & NAS of LLM and VLM models.
|
||||
- [2026/04/15] Customer story: [Domyn compresses Colosseum-355B → 260B using ModelOpt's Minitron pruning + distillation](https://www.domyn.com/blog/domyn-large-the-journey-of-a-european-sovereign-ai-model-for-regulated-industries)
|
||||
- [2026/03/17] Customer story: [Bielik.AI builds Bielik Minitron 7B (33% smaller, 50% faster, 90% quality retained) using ModelOpt's Minitron pruning + distillation](https://bielik.ai/en/nvidia-gtc-bielik-minitron-premiere/)
|
||||
- [2026/03/11] Model Optimizer quantized Nemotron-3-Super checkpoints are available on Hugging Face for download: [FP8](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8), [NVFP4](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4). Learn more in the [Nemotron 3 Super release blog](https://blogs.nvidia.com/blog/nemotron-3-super-agentic-ai/). Check out how to quantize Nemotron 3 models for deployment acceleration [here](./examples/llm_ptq/README.md)
|
||||
- [2026/03/11] Model Optimizer quantized Nemotron-3-Super checkpoints are available on Hugging Face for download: [FP8](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8), [NVFP4](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4). Learn more in the [Nemotron 3 Super release blog](https://blogs.nvidia.com/blog/nemotron-3-super-agentic-ai/). Check out how to quantize Nemotron 3 models for deployment acceleration [here](./examples/hf_ptq/README.md)
|
||||
- [2026/03/11] [NeMo Megatron Bridge](https://github.com/NVIDIA-NeMo/Megatron-Bridge) now supports Nemotron-3-Super quantization (PTQ and QAT) and export workflows using the Model Optimizer library. See the [Quantization (PTQ and QAT) guide](https://github.com/NVIDIA-NeMo/Megatron-Bridge/blob/super-v3/docs/models/llm/nemotron3-super.md#quantization-ptq-and-qat) for FP8/NVFP4 quantization and HF export instructions.
|
||||
- [2025/12/11] [BLOG: Top 5 AI Model Optimization Techniques for Faster, Smarter Inference](https://developer.nvidia.com/blog/top-5-ai-model-optimization-techniques-for-faster-smarter-inference/)
|
||||
- [2025/12/08] NVIDIA TensorRT Model Optimizer is now officially rebranded as NVIDIA Model Optimizer.
|
||||
@@ -42,10 +42,10 @@ Model Optimizer is also integrated with [NVIDIA Megatron-Bridge](https://github.
|
||||
- [2025/06/24] [BLOG: Introducing NVFP4 for Efficient and Accurate Low-Precision Inference](https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/)
|
||||
- [2025/05/14] [NVIDIA TensorRT Unlocks FP4 Image Generation for NVIDIA Blackwell GeForce RTX 50 Series GPUs](https://developer.nvidia.com/blog/nvidia-tensorrt-unlocks-fp4-image-generation-for-nvidia-blackwell-geforce-rtx-50-series-gpus/)
|
||||
- [2025/04/21] [Adobe optimized deployment using Model-Optimizer + TensorRT leading to a 60% reduction in diffusion latency, a 40% reduction in total cost of ownership](https://developer.nvidia.com/blog/optimizing-transformer-based-diffusion-models-for-video-generation-with-nvidia-tensorrt/)
|
||||
- [2025/04/05] [NVIDIA Accelerates Inference on Meta Llama 4 Scout and Maverick](https://developer.nvidia.com/blog/nvidia-accelerates-inference-on-meta-llama-4-scout-and-maverick/). Check out how to quantize Llama4 for deployment acceleration [here](./examples/llm_ptq/README.md#llama-4)
|
||||
- [2025/04/05] [NVIDIA Accelerates Inference on Meta Llama 4 Scout and Maverick](https://developer.nvidia.com/blog/nvidia-accelerates-inference-on-meta-llama-4-scout-and-maverick/). Check out how to quantize Llama4 for deployment acceleration [here](./examples/hf_ptq/README.md#support-matrix)
|
||||
- [2025/03/18] [World's Fastest DeepSeek-R1 Inference with Blackwell FP4 & Increasing Image Generation Efficiency on Blackwell](https://developer.nvidia.com/blog/nvidia-blackwell-delivers-world-record-deepseek-r1-inference-performance/)
|
||||
- [2025/02/25] Model Optimizer quantized NVFP4 models available on Hugging Face for download: [DeepSeek-R1-FP4](https://huggingface.co/nvidia/DeepSeek-R1-FP4), [Llama-3.3-70B-Instruct-FP4](https://huggingface.co/nvidia/Llama-3.3-70B-Instruct-FP4), [Llama-3.1-405B-Instruct-FP4](https://huggingface.co/nvidia/Llama-3.1-405B-Instruct-FP4)
|
||||
- [2025/01/28] Model Optimizer has added support for NVFP4. Check out an example of NVFP4 PTQ [here](./examples/llm_ptq/README.md#model-quantization-and-trt-llm-conversion).
|
||||
- [2025/01/28] Model Optimizer has added support for NVFP4. Check out an example of NVFP4 PTQ [here](./examples/hf_ptq/README.md#getting-started).
|
||||
- [2025/01/28] Model Optimizer is now open source!
|
||||
|
||||
<details close>
|
||||
@@ -56,7 +56,7 @@ Model Optimizer is also integrated with [NVIDIA Megatron-Bridge](https://github.
|
||||
- [2024/08/28] [Boosting Llama 3.1 405B Performance up to 44% with Model Optimizer on NVIDIA H200 GPUs](https://developer.nvidia.com/blog/boosting-llama-3-1-405b-performance-by-up-to-44-with-nvidia-tensorrt-model-optimizer-on-nvidia-h200-gpus/)
|
||||
- [2024/08/28] [Up to 1.9X Higher Llama 3.1 Performance with Medusa](https://developer.nvidia.com/blog/low-latency-inference-chapter-1-up-to-1-9x-higher-llama-3-1-performance-with-medusa-on-nvidia-hgx-h200-with-nvlink-switch/)
|
||||
- [2024/08/15] New features in recent releases: [Cache Diffusion](./examples/diffusers/cache_diffusion), [QLoRA workflow with NVIDIA NeMo](https://docs.nvidia.com/nemo-framework/user-guide/24.09/sft_peft/qlora.html), and more. Check out [our blog](https://developer.nvidia.com/blog/nvidia-tensorrt-model-optimizer-v0-15-boosts-inference-performance-and-expands-model-support/) for details.
|
||||
- [2024/06/03] Model Optimizer now has an experimental feature to deploy to vLLM as part of our effort to support popular deployment frameworks. Check out the workflow [here](./examples/llm_ptq/README.md#deploy-fp8-quantized-model-using-vllm)
|
||||
- [2024/06/03] Model Optimizer now has an experimental feature to deploy to vLLM as part of our effort to support popular deployment frameworks. Check out the workflow [here](./examples/hf_ptq/README.md#vllm)
|
||||
- [2024/05/08] [Announcement: Model Optimizer Now Formally Available to Further Accelerate GenAI Inference Performance](https://developer.nvidia.com/blog/accelerate-generative-ai-inference-performance-with-nvidia-tensorrt-model-optimizer-now-publicly-available/)
|
||||
- [2024/03/27] [Model Optimizer supercharges TensorRT-LLM to set MLPerf LLM inference records](https://developer.nvidia.com/blog/nvidia-h200-tensor-core-gpus-and-nvidia-tensorrt-llm-set-mlperf-llm-inference-records/)
|
||||
- [2024/03/18] [GTC Session: Optimize Generative AI Inference with Quantization in TensorRT-LLM and TensorRT](https://www.nvidia.com/en-us/on-demand/session/gtc24-s63213/)
|
||||
@@ -102,7 +102,7 @@ more fine-grained control on installed dependencies or for alternative docker im
|
||||
|
||||
| **Technique** | **Description** | **Examples** | **Docs** |
|
||||
| :------------: | :------------: | :------------: | :------------: |
|
||||
| Post Training Quantization | Compress model size by 2x-4x, speeding up inference while preserving model quality! | \[[HF LLMs / VLMs](./examples/llm_ptq/)\] \[[Megatron-Bridge LLMs / VLMs](./examples/megatron_bridge/)\] \[[Diffusers](./examples/diffusers/)\] \[[ONNX](./examples/onnx_ptq/)\] \[[Windows](./examples/windows/)\] | \[[docs](https://nvidia.github.io/Model-Optimizer/guides/1_quantization.html)\] |
|
||||
| Post Training Quantization | Compress model size by 2x-4x, speeding up inference while preserving model quality! | \[[HF LLMs / VLMs](./examples/hf_ptq/)\] \[[Megatron-Bridge LLMs / VLMs](./examples/megatron_bridge/)\] \[[Diffusers](./examples/diffusers/)\] \[[ONNX](./examples/onnx_ptq/)\] \[[Windows](./examples/windows/)\] | \[[docs](https://nvidia.github.io/Model-Optimizer/guides/1_quantization.html)\] |
|
||||
| Quantization Aware Training / Distillation | Refine accuracy of quantized models even further with a few training steps! | \[[Hugging Face](./examples/llm_qat/)\] \[[Megatron-Bridge](./examples/megatron_bridge)\] | \[[docs](https://nvidia.github.io/Model-Optimizer/guides/1_quantization.html)\] |
|
||||
| Pruning | Reduce your model parameters or memory footprint and accelerate inference by removing unnecessary weights! | \[[General](./examples/pruning/)\] \[[Megatron-Bridge](./examples/megatron_bridge/)\] | |
|
||||
| Distillation | Reduce deployment model size by teaching small models to behave like larger models! | \[[Hugging Face](./examples/llm_distill/)\] \[[Megatron-Bridge](./examples/megatron_bridge/)\] \[[Megatron-LM](./examples/llm_distill/README.md#knowledge-distillation-kd-in-nvidia-megatron-lm-framework)\] | \[[docs](https://nvidia.github.io/Model-Optimizer/guides/4_distillation.html)\] |
|
||||
@@ -130,8 +130,7 @@ more fine-grained control on installed dependencies or for alternative docker im
|
||||
|
||||
| Model Type | Support Matrix |
|
||||
|------------|----------------|
|
||||
| LLM Quantization | [View Support Matrix](./examples/llm_ptq/README.md#support-matrix) |
|
||||
| VLM Quantization | [View Support Matrix](./examples/llm_ptq/README.md#hugging-face-supported-models) |
|
||||
| LLM / VLM Quantization | [View Support Matrix](./examples/hf_ptq/README.md#support-matrix) |
|
||||
| Diffusers Quantization | [View Support Matrix](./examples/diffusers/README.md#support-matrix) |
|
||||
| ONNX Quantization | [View Support Matrix](./examples/torch_onnx/README.md#onnx-export-supported-llm-models) |
|
||||
| Windows Quantization | [View Support Matrix](./examples/windows/README.md#support-matrix) |
|
||||
|
||||
@@ -6,7 +6,7 @@ We support exporting modelopt-optimized Hugging Face models (transformers and di
|
||||
|
||||
The workflow is as follows:
|
||||
|
||||
#. Load the Huggingface models or Megatron Core models, `quantize with modelopt <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/llm_ptq#ptq-post-training-quantization>`_ , and export to the unified checkpoint format, where the layer structures and tensor names are aligned with the original checkpoint.
|
||||
#. Load the Huggingface models or Megatron Core models, `quantize with modelopt <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/hf_ptq#ptq-post-training-quantization>`_ , and export to the unified checkpoint format, where the layer structures and tensor names are aligned with the original checkpoint.
|
||||
#. Load the unified checkpoint in the supported inference framework for accelerated inference.
|
||||
|
||||
|
||||
|
||||
@@ -570,7 +570,7 @@ Some example scripts accept a ``--recipe`` flag. For instance, the PTQ example:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
python examples/llm_ptq/hf_ptq.py \
|
||||
python examples/hf_ptq/hf_ptq.py \
|
||||
--model Qwen/Qwen3-8B \
|
||||
--recipe general/ptq/fp8_default-kv_fp8_cast \
|
||||
--export_path build/fp8 \
|
||||
|
||||
@@ -58,4 +58,4 @@ For quantized formats like NVFP4, you can reduce memory usage by up to 4x compar
|
||||
.. note::
|
||||
|
||||
An example implementation of this workflow can be found in:
|
||||
``examples/llm_ptq/hf_ptq.py``, which reduces the memory requirements of model calibration.
|
||||
``examples/hf_ptq/hf_ptq.py``, which reduces the memory requirements of model calibration.
|
||||
|
||||
@@ -15,7 +15,7 @@ As ModelOpt cannot detect these linear ops out-of-the-box, a HugggingFace plugin
|
||||
#. Define a customized ``_QuantDbrxExpertGLU`` as a ``DynamicModule`` with the same ``forward`` signature.
|
||||
#. Rewrite the linear ops (w1, v1 and v2) as a standard ``nn.Linear`` op, and re-implement the ``forward`` method.
|
||||
#. Register the new dynamic ``_QuantDbrxExperts`` to replace the ``DbrxExperts`` from the modeling_dbrx.py in the ``transformers`` library
|
||||
#. Try quantize the DBRX model after the plugin is implemented, feel free to follow the `llm_ptq example <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/llm_ptq>`_.
|
||||
#. Try quantize the DBRX model after the plugin is implemented, feel free to follow the `hf_ptq example <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/hf_ptq>`_.
|
||||
#. TensorRT-LLM is open-sourced. If this customized model is not supported by TensorRT-LLM yet, please modify :meth:`export_tensorrt_llm_checkpoint <modelopt.torch.export.export_tensorrt_llm_checkpoint>` or :meth:`export_hf_checkpoint <modelopt.torch.export.export_hf_checkpoint>` to export the quantized model for deployment with a customized TensorRT-LLM modeling implementation. Feel free to :doc:`contact us <../support/1_contact>` if further support is needed.
|
||||
|
||||
The following code snippet is excerpted from ``modelopt/torch/quantization/plugins/huggingface.py``
|
||||
|
||||
@@ -7,7 +7,7 @@ Welcome to Model Optimizer (ModelOpt) documentation!
|
||||
:caption: Getting Started
|
||||
|
||||
getting_started/[0-9]*
|
||||
Quick Start: PTQ - PyTorch <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/llm_ptq>
|
||||
Quick Start: PTQ - PyTorch <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/hf_ptq>
|
||||
Quick Start: PTQ - ONNX <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/onnx_ptq>
|
||||
Quick Start: PTQ - PyTorch to ONNX <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/torch_onnx>
|
||||
Quick Start: PTQ - Windows <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/windows>
|
||||
|
||||
@@ -193,6 +193,6 @@ lands in E4M3's representable window; the rare out-of-range block falls back to
|
||||
data-derived scale). The flag only affects routed-expert **weights** — activation
|
||||
`input_scale` still comes from `${AMAX}` calibration — and the run prints a
|
||||
`[cast] lossless MXFP4->NVFP4 blocks: …` summary. This mirrors the GPTOSS cast in
|
||||
[`examples/llm_ptq/cast_mxfp4_to_nvfp4.py`](../llm_ptq/cast_mxfp4_to_nvfp4.py); the
|
||||
[`examples/hf_ptq/cast_mxfp4_to_nvfp4.py`](../hf_ptq/cast_mxfp4_to_nvfp4.py); the
|
||||
V4 twist is that w1/w3 share one `scale_2` (fused GEMM1), so `k_max` is taken over
|
||||
both projections.
|
||||
|
||||
@@ -74,7 +74,7 @@ and the NVFP4 nibbles equal the source MXFP4 nibbles bit-for-bit (for every
|
||||
block whose ``k_j`` lands in E4M3's representable window). The flag only affects
|
||||
routed-expert *weights*; activation ``input_scale`` still comes from
|
||||
``--amax_path`` calibration. This mirrors the GPTOSS cast in
|
||||
``examples/llm_ptq/cast_mxfp4_to_nvfp4.py`` (PR #1372); the V4 twist is that
|
||||
``examples/hf_ptq/cast_mxfp4_to_nvfp4.py`` (PR #1372); the V4 twist is that
|
||||
w1/w3 share one ``scale_2`` (fused GEMM1), so ``k_max`` is taken over both.
|
||||
|
||||
Usage (single compute node, CPU-default; dequant+requant math is cheap
|
||||
|
||||
@@ -184,4 +184,4 @@ ModelOpt provides easy end to end QAT via [LLaMA-Factory](https://github.com/hiy
|
||||
|
||||
### Deployment of ModelOpt QAT/PTQ models beyond GPT-OSS
|
||||
|
||||
ModelOpt supports exporting a wide variety of models after QAT/PTQ to TensorRT-LLM, vLLM, SGLang etc. Please refer to [llm_ptq](../llm_ptq).
|
||||
ModelOpt supports exporting a wide variety of models after QAT/PTQ to TensorRT-LLM, vLLM, SGLang etc. Please refer to [hf_ptq](../hf_ptq).
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
# =============================================================================
|
||||
# FSDP Configuration for running LLM PTQ on multinode setup. This file is consumed by examples/llm_ptq/multinode_ptq.py
|
||||
# FSDP Configuration for running LLM PTQ on multinode setup. This file is consumed by examples/hf_ptq/multinode_ptq.py
|
||||
# =============================================================================
|
||||
|
||||
compute_environment: LOCAL_MACHINE
|
||||
@@ -6,7 +6,7 @@ The following instructions show how to evaluate the Model Optimizer quantized LL
|
||||
|
||||
## NeMo Evaluator
|
||||
|
||||
[NeMo Evaluator](https://docs.nvidia.com/nemo/evaluator/latest/get-started/quickstart/index.html#self-hosted-options) is the recommended way to evaluate a large choice of benchmarks on quantized checkpoints generated from [llm_ptq](../llm_ptq). Quantized checkpoints can be served with [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM), [vLLM](https://github.com/vllm-project/vllm), or [SGLang](https://github.com/sgl-project/sglang) and then evaluated using NeMo Evaluator.
|
||||
[NeMo Evaluator](https://docs.nvidia.com/nemo/evaluator/latest/get-started/quickstart/index.html#self-hosted-options) is the recommended way to evaluate a large choice of benchmarks on quantized checkpoints generated from [hf_ptq](../hf_ptq). Quantized checkpoints can be served with [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM), [vLLM](https://github.com/vllm-project/vllm), or [SGLang](https://github.com/sgl-project/sglang) and then evaluated using NeMo Evaluator.
|
||||
|
||||
## LM-Eval-Harness
|
||||
|
||||
@@ -233,7 +233,7 @@ This is useful for evaluating quantized models deployed with vLLM or any model s
|
||||
--tensor-parallel-size <tp_size> # Adjust as needed
|
||||
```
|
||||
|
||||
To generate the quantized model such as `nvidia/Llama-3.1-8B-Instruct-FP8`, please refer to instructions [here](https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/llm_ptq#deploy-fp8-quantized-model-using-vllm-and-sglang). Note currently modelopt quantized model support in vLLM is limited, we are working on expanding the model and quant formats support.
|
||||
To generate the quantized model such as `nvidia/Llama-3.1-8B-Instruct-FP8`, please refer to instructions [here](https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/hf_ptq#deploy-fp8-quantized-model-using-vllm-and-sglang). Note currently modelopt quantized model support in vLLM is limited, we are working on expanding the model and quant formats support.
|
||||
|
||||
1. **Make the script executable (if not already):**
|
||||
|
||||
|
||||
Symlink
+1
@@ -0,0 +1 @@
|
||||
hf_ptq
|
||||
@@ -24,7 +24,7 @@ For background on how QAT enables low-precision accuracy recovery, see the [QAT/
|
||||
|
||||
### Prerequisites
|
||||
|
||||
Please refer to [llm_ptq/README.md](../llm_ptq/README.md#pre-requisites) for container
|
||||
Please refer to [hf_ptq/README.md](../hf_ptq/README.md#pre-requisites) for container
|
||||
recommendations and base ModelOpt installation guidance. For this QAT/QAD example,
|
||||
install the Hugging Face dependencies and the example-specific requirements:
|
||||
|
||||
@@ -85,7 +85,7 @@ accelerate launch --config-file configs/accelerate/fsdp2.yaml train.py \
|
||||
python export.py --pyt_ckpt_path qwen3-8b-qad-nvfp4 --export_path qwen3-8b-qad-deploy
|
||||
```
|
||||
|
||||
Exported checkpoints can be deployed on [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM), [vLLM](https://github.com/vllm-project/vllm), or [SGLang](https://github.com/sgl-project/sglang). See [llm_ptq/README.md](../llm_ptq/README.md#deployment) for deployment instructions. For quick accuracy evaluation without exporting, see [Native Fake-Quantized Evaluation](#native-fake-quantized-evaluation).
|
||||
Exported checkpoints can be deployed on [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM), [vLLM](https://github.com/vllm-project/vllm), or [SGLang](https://github.com/sgl-project/sglang). See [hf_ptq/README.md](../hf_ptq/README.md#deployment) for deployment instructions. For quick accuracy evaluation without exporting, see [Native Fake-Quantized Evaluation](#native-fake-quantized-evaluation).
|
||||
|
||||
> [!NOTE]
|
||||
> For a minimal end-to-end demo (quantize + train + save in one script), see [simple_qat_train.py](simple_qat_train.py). It runs on a **single GPU** only and is intended as a quick introduction to the QAT flow (without transformer trainer)—not for distributed training.
|
||||
|
||||
@@ -94,9 +94,9 @@ The final QAT/QAD model after training is similar in architecture to that of PTQ
|
||||
To run QAT/QAD model with TRTLLM, run:
|
||||
|
||||
```sh
|
||||
cd ../../llm_ptq
|
||||
cd ../../hf_ptq
|
||||
|
||||
./scripts/huggingface_example.sh --model <path-to-QAT/QAD-model> --quant nvfp4
|
||||
```
|
||||
|
||||
See more details on deployment of quantized model [here](../../llm_ptq/README.md).
|
||||
See more details on deployment of quantized model [here](../../hf_ptq/README.md).
|
||||
|
||||
@@ -544,7 +544,7 @@
|
||||
"cell_type": "markdown",
|
||||
"id": "10acc50c-c876-41d5-8f7e-00dab8842ccd",
|
||||
"metadata": {},
|
||||
"source": "**Note:** The QAT checkpoint for `nvfp4` config can also be created using the CLI scripts. See the [QAT README](../README.md) for the full end-to-end workflow using `quantize.py`, `train.py`, and `export.py`.\n\nSee more details on deployment of quantized model [here](https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/llm_ptq/README.md)."
|
||||
"source": "**Note:** The QAT checkpoint for `nvfp4` config can also be created using the CLI scripts. See the [QAT README](../README.md) for the full end-to-end workflow using `quantize.py`, `train.py`, and `export.py`.\n\nSee more details on deployment of quantized model [here](https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/hf_ptq/README.md)."
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
@@ -603,7 +603,7 @@
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Exporting Quantized Model for deployment\n",
|
||||
"Before deploying the model with TensorRT-LLM you will need to export the model checkpoint files. This is similar to the step you take for a quantized PTQ Model. To export the unified Hugging Face checkpoints, which can be deployed on TensorRT-LLM Pytorch, vLLM and SGLang you will need to run the [huggingface_example.sh](https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/llm_ptq/scripts/huggingface_example.sh) script found in the Model Optimizer repo. "
|
||||
"Before deploying the model with TensorRT-LLM you will need to export the model checkpoint files. This is similar to the step you take for a quantized PTQ Model. To export the unified Hugging Face checkpoints, which can be deployed on TensorRT-LLM Pytorch, vLLM and SGLang you will need to run the [huggingface_example.sh](https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/hf_ptq/scripts/huggingface_example.sh) script found in the Model Optimizer repo. "
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -671,7 +671,7 @@
|
||||
"\n",
|
||||
"# run conversion script\n",
|
||||
"cd ..\n",
|
||||
"bash Model-Optimizer/examples/llm_ptq/scripts/huggingface_example.sh --model $(pwd)/qat/checkpoint-450/ --quant nvfp4"
|
||||
"bash Model-Optimizer/examples/hf_ptq/scripts/huggingface_example.sh --model $(pwd)/qat/checkpoint-450/ --quant nvfp4"
|
||||
]
|
||||
},
|
||||
{
|
||||
|
||||
@@ -73,7 +73,7 @@ from modelopt.torch.utils.plugins.megatron_calibration import get_megatron_calib
|
||||
from modelopt.torch.utils.plugins.megatron_generate import megatron_generate
|
||||
|
||||
# The --quant_cfg / --kv_cache_quant CLI vocabularies are discovered from the preset
|
||||
# YAMLs (shared with the llm_ptq examples via modelopt.recipe.presets). --quant_cfg
|
||||
# YAMLs (shared with the hf_ptq examples via modelopt.recipe.presets). --quant_cfg
|
||||
# additionally accepts any full config name from ``mtq.config.choices`` (e.g.
|
||||
# ``FP8_DEFAULT_CFG``); see get_quant_config below.
|
||||
|
||||
|
||||
@@ -27,4 +27,4 @@ To deploy and run on SGLang:
|
||||
python run_llama_fp8_sglang.py
|
||||
```
|
||||
|
||||
If you want to run post-training quantization with Model Optimizer for your selected models, check [here](../llm_ptq/README.md).
|
||||
If you want to run post-training quantization with Model Optimizer for your selected models, check [here](../hf_ptq/README.md).
|
||||
|
||||
@@ -9,7 +9,7 @@ End-to-end optimization of [Nemotron-Nano-9B-v2](https://huggingface.co/nvidia/N
|
||||
2. **[Pruning](#2-pruning)** — Minitron structured pruning from 9B to 7B
|
||||
3. **[Distillation](#3-distillation)** — recovering accuracy via Megatron-Bridge knowledge distillation (up to 80B tokens)
|
||||
4. **[Evaluation](#4-evaluation)** — benchmarking with NeMo Evaluator across MMLU Pro, GPQA Diamond, AIME, and more
|
||||
5. **[Quantization](#5-quantization)** — FP8 PTQ on the distilled checkpoint using ModelOpt's `examples/llm_ptq/hf_ptq.py` script
|
||||
5. **[Quantization](#5-quantization)** — FP8 PTQ on the distilled checkpoint using ModelOpt's `examples/hf_ptq/hf_ptq.py` script
|
||||
6. **[vLLM Inference Benchmarking](#6-vllm-inference-benchmarking)** — throughput comparison of BF16 vs FP8 on a single H100
|
||||
|
||||
## Results
|
||||
@@ -317,11 +317,11 @@ For more details on NeMo Evaluator, see the [GitHub repo](https://github.com/NVI
|
||||
|
||||
### 5. Quantization
|
||||
|
||||
ModelOpt allows stacking multiple optimization techniques. Here we stack FP8 quantization on top of the pruned and distilled model to get an even more optimized model. See [examples/llm_ptq/README.md](../../../llm_ptq/README.md) for the full PTQ documentation.
|
||||
ModelOpt allows stacking multiple optimization techniques. Here we stack FP8 quantization on top of the pruned and distilled model to get an even more optimized model. See [examples/hf_ptq/README.md](../../../hf_ptq/README.md) for the full PTQ documentation.
|
||||
|
||||
Similar to the official [Nemotron-Nano-9B-v2-FP8](https://huggingface.co/nvidia/NVIDIA-Nemotron-Nano-9B-v2-FP8) model, if you want to quantize the pruned 7B model to FP8, the Mamba and MLP layers are quantized to FP8, while all 4 attention layers and the Conv1d components within the Mamba layers are kept in BF16 to avoid accuracy degradation.
|
||||
|
||||
This is done with the `mtq.MAMBA_MOE_FP8_CONSERVATIVE_CFG` config defined in [`modelopt/torch/quantization/config.py`](../../../../modelopt/torch/quantization/config.py). To apply this, you need to modify `QUANT_CFG_CHOICES["fp8"]` in [`examples/llm_ptq/hf_ptq.py`](../../../llm_ptq/hf_ptq.py) to use `mtq.MAMBA_MOE_FP8_CONSERVATIVE_CFG`. For a faster model at the cost of a larger accuracy drop, you can use `mtq.MAMBA_MOE_FP8_AGGRESSIVE_CFG` instead.
|
||||
This is done with the `mtq.MAMBA_MOE_FP8_CONSERVATIVE_CFG` config defined in [`modelopt/torch/quantization/config.py`](../../../../modelopt/torch/quantization/config.py). To apply this, you need to modify `QUANT_CFG_CHOICES["fp8"]` in [`examples/hf_ptq/hf_ptq.py`](../../../hf_ptq/hf_ptq.py) to use `mtq.MAMBA_MOE_FP8_CONSERVATIVE_CFG`. For a faster model at the cost of a larger accuracy drop, you can use `mtq.MAMBA_MOE_FP8_AGGRESSIVE_CFG` instead.
|
||||
|
||||
> [!NOTE]
|
||||
> You can also quantize to NVFP4 using `mtq.MAMBA_MOE_NVFP4_CONSERVATIVE_CFG` (default) or `mtq.MAMBA_MOE_NVFP4_AGGRESSIVE_CFG` (faster, more accuracy drop), which may require further distillation (QAD) to recover accuracy and Blackwell GPU for deployment.
|
||||
@@ -329,7 +329,7 @@ This is done with the `mtq.MAMBA_MOE_FP8_CONSERVATIVE_CFG` config defined in [`m
|
||||
Calibrate and export the HF checkpoint from iteration 12800 to FP8 (takes 1-2 mins on 8x H100):
|
||||
|
||||
```bash
|
||||
python /opt/Model-Optimizer/examples/llm_ptq/hf_ptq.py \
|
||||
python /opt/Model-Optimizer/examples/hf_ptq/hf_ptq.py \
|
||||
--pyt_ckpt_path <output_dir>/checkpoints/hf_iter_12800 \
|
||||
--export_path <output_dir>/checkpoints/hf_iter_12800_fp8 \
|
||||
--qformat fp8 \
|
||||
|
||||
@@ -203,7 +203,7 @@ One can also use [examples/specdec_bench](../specdec_bench) to validate the trai
|
||||
|
||||
### Deploying Quantized model
|
||||
|
||||
See more details on deployment of quantized model to TRTLLM [here](../llm_ptq/README.md).
|
||||
See more details on deployment of quantized model to TRTLLM [here](../hf_ptq/README.md).
|
||||
|
||||
## Advanced Usage
|
||||
|
||||
|
||||
@@ -65,10 +65,10 @@ lm_eval --model local-completions --tasks gsm8k --model_args model=<model_name>,
|
||||
|
||||
Step 1: export the model with bf16 weights and quantizer state. To export the model:
|
||||
|
||||
- For **HF** models, use `examples/llm_ptq/hf_ptq.py` with `--vllm_fakequant_export`:
|
||||
- For **HF** models, use `examples/hf_ptq/hf_ptq.py` with `--vllm_fakequant_export`:
|
||||
|
||||
```bash
|
||||
python ../llm_ptq/hf_ptq.py \
|
||||
python ../hf_ptq/hf_ptq.py \
|
||||
--pyt_ckpt_path <MODEL_PATH> \
|
||||
--recipe <PATH_TO_RECIPE> \
|
||||
--calib_size 512 \
|
||||
|
||||
@@ -1,15 +1,15 @@
|
||||
# [Deprecated] Post-training quantization (PTQ) for Vision Language Models
|
||||
|
||||
> **This example has been consolidated into [`examples/llm_ptq`](../llm_ptq/README.md) and is
|
||||
> **This example has been consolidated into [`examples/hf_ptq`](../hf_ptq/README.md) and is
|
||||
> deprecated.** It will be removed in a future release. VLM PTQ now shares the same entry point
|
||||
> (`hf_ptq.py`) and shell script as LLM PTQ.
|
||||
|
||||
## Migration
|
||||
|
||||
Use the `llm_ptq` script with the `--vlm` flag:
|
||||
Use the `hf_ptq` script with the `--vlm` flag:
|
||||
|
||||
```bash
|
||||
cd examples/llm_ptq
|
||||
cd examples/hf_ptq
|
||||
scripts/huggingface_example.sh --model <Hugging Face model card or checkpoint> --quant [fp8|nvfp4|int8_sq|int4_awq|w4a8_awq] --vlm
|
||||
```
|
||||
|
||||
@@ -20,9 +20,9 @@ prints a deprecation warning and forwards to the command above.
|
||||
|
||||
| Topic | New location |
|
||||
| :--- | :--- |
|
||||
| Supported VLMs / support matrix | [llm_ptq/README.md#hugging-face-supported-models](../llm_ptq/README.md#hugging-face-supported-models) |
|
||||
| VLM quantization workflow (`--vlm`) | [llm_ptq/README.md#vlm-quantization](../llm_ptq/README.md#vlm-quantization) |
|
||||
| Image-text calibration (`--calib_with_images`) | [llm_ptq/README.md#vlm-calibration-with-image-text-pairs-eg-nemotron-vl](../llm_ptq/README.md#vlm-calibration-with-image-text-pairs-eg-nemotron-vl) |
|
||||
| Supported VLMs / support matrix | [hf_ptq/README.md#hugging-face-supported-models](../hf_ptq/README.md#hugging-face-supported-models) |
|
||||
| VLM quantization workflow (`--vlm`) | [hf_ptq/README.md#vlm-quantization](../hf_ptq/README.md#vlm-quantization) |
|
||||
| Image-text calibration (`--calib_with_images`) | [hf_ptq/README.md#vlm-calibration-with-image-text-pairs-eg-nemotron-vl](../hf_ptq/README.md#vlm-calibration-with-image-text-pairs-eg-nemotron-vl) |
|
||||
| Megatron-Bridge VLM PTQ | [examples/megatron_bridge/](../megatron_bridge/README.md) |
|
||||
|
||||
## Resources
|
||||
|
||||
@@ -14,21 +14,21 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# DEPRECATED: examples/vlm_ptq has been consolidated into examples/llm_ptq.
|
||||
# This shim forwards all arguments to the llm_ptq script with the --vlm flag so existing
|
||||
# DEPRECATED: examples/vlm_ptq has been consolidated into examples/hf_ptq.
|
||||
# This shim forwards all arguments to the hf_ptq script with the --vlm flag so existing
|
||||
# commands keep working. Please migrate to:
|
||||
#
|
||||
# cd examples/llm_ptq
|
||||
# cd examples/hf_ptq
|
||||
# scripts/huggingface_example.sh --model <model> --quant <qformat> --vlm
|
||||
#
|
||||
# See examples/llm_ptq/README.md#vlm-quantization for details.
|
||||
# See examples/hf_ptq/README.md#vlm-quantization for details.
|
||||
|
||||
set -e
|
||||
|
||||
echo "WARNING: examples/vlm_ptq is deprecated and will be removed in a future release." >&2
|
||||
echo " Forwarding to examples/llm_ptq/scripts/huggingface_example.sh --vlm" >&2
|
||||
echo " See examples/llm_ptq/README.md#vlm-quantization" >&2
|
||||
echo " Forwarding to examples/hf_ptq/scripts/huggingface_example.sh --vlm" >&2
|
||||
echo " See examples/hf_ptq/README.md#vlm-quantization" >&2
|
||||
|
||||
script_dir="$(dirname "$(readlink -f "$0")")"
|
||||
|
||||
exec "$script_dir/../../llm_ptq/scripts/huggingface_example.sh" --vlm "$@"
|
||||
exec "$script_dir/../../hf_ptq/scripts/huggingface_example.sh" --vlm "$@"
|
||||
|
||||
@@ -15,8 +15,8 @@
|
||||
|
||||
"""PTQ quant-config preset discovery shared by the PTQ example scripts.
|
||||
|
||||
The example PTQ entry points (``examples/llm_ptq/hf_ptq.py``,
|
||||
``examples/llm_ptq/multinode_ptq.py``, ``examples/megatron_bridge/quantize.py``)
|
||||
The example PTQ entry points (``examples/hf_ptq/hf_ptq.py``,
|
||||
``examples/hf_ptq/multinode_ptq.py``, ``examples/megatron_bridge/quantize.py``)
|
||||
expose a ``--qformat`` / ``--kv_cache_qformat`` (``--quant_cfg`` /
|
||||
``--kv_cache_quant`` for Megatron-Bridge) CLI vocabulary. Rather than hardcoding a
|
||||
name → config table in each script, the vocabulary is discovered by listing the
|
||||
|
||||
@@ -19,7 +19,7 @@ These helpers turn an MXFP4 source layer's E8M0 block scales into the per-tensor
|
||||
``global_amax`` and per-NVFP4-block ``amax`` that pin NVFP4's two-level scale so
|
||||
the cast reproduces the source MXFP4 weights bit-for-bit (see PR #1372 for the
|
||||
derivation). They are pure tensor math with no model or checkpoint dependencies,
|
||||
shared by the GPT-OSS (``examples/llm_ptq``) and DeepSeek-V4
|
||||
shared by the GPT-OSS (``examples/hf_ptq``) and DeepSeek-V4
|
||||
(``examples/deepseek``) PTQ cast paths.
|
||||
"""
|
||||
|
||||
|
||||
+5
-5
@@ -12,8 +12,8 @@
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
"""Re-export ``examples/llm_ptq/example_utils`` so tests can import it via
|
||||
``from _test_utils.examples.llm_ptq_example_utils import example_utils``
|
||||
"""Re-export ``examples/hf_ptq/example_utils`` so tests can import it via
|
||||
``from _test_utils.examples.hf_ptq_example_utils import example_utils``
|
||||
without per-file ``sys.path`` shims.
|
||||
"""
|
||||
|
||||
@@ -21,9 +21,9 @@ import sys
|
||||
|
||||
from _test_utils.examples.run_command import MODELOPT_ROOT
|
||||
|
||||
_LLM_PTQ_DIR = MODELOPT_ROOT / "examples" / "llm_ptq"
|
||||
if str(_LLM_PTQ_DIR) not in sys.path:
|
||||
sys.path.insert(0, str(_LLM_PTQ_DIR))
|
||||
_HF_PTQ_DIR = MODELOPT_ROOT / "examples" / "hf_ptq"
|
||||
if str(_HF_PTQ_DIR) not in sys.path:
|
||||
sys.path.insert(0, str(_HF_PTQ_DIR))
|
||||
|
||||
import example_utils
|
||||
|
||||
+2
-2
@@ -19,7 +19,7 @@ from dataclasses import asdict, dataclass
|
||||
|
||||
import pytest
|
||||
import torch
|
||||
from _test_utils.examples.run_command import run_llm_ptq_command
|
||||
from _test_utils.examples.run_command import run_hf_ptq_command
|
||||
|
||||
|
||||
@dataclass
|
||||
@@ -62,7 +62,7 @@ class PTQCommand:
|
||||
param_dict.pop("min_gpu", None)
|
||||
|
||||
quant = param_dict.pop("quant")
|
||||
run_llm_ptq_command(model=model_path, quant=quant, **param_dict)
|
||||
run_hf_ptq_command(model=model_path, quant=quant, **param_dict)
|
||||
|
||||
def param_str(self):
|
||||
param_dict = asdict(self)
|
||||
@@ -62,14 +62,14 @@ def run_command_in_background(
|
||||
return process
|
||||
|
||||
|
||||
def run_llm_ptq_command(*, model: str, quant: str, vlm: bool = False, **kwargs):
|
||||
def run_hf_ptq_command(*, model: str, quant: str, vlm: bool = False, **kwargs):
|
||||
kwargs.update({"model": model, "quant": quant})
|
||||
kwargs.setdefault("tasks", "quant")
|
||||
kwargs.setdefault("calib", 16)
|
||||
|
||||
cmd_parts = ["scripts/huggingface_example.sh", "--no-verbose"]
|
||||
if vlm:
|
||||
# VLM PTQ shares the llm_ptq entry point; --vlm runs the multimodal deploy smoke test.
|
||||
# VLM PTQ shares the hf_ptq entry point; --vlm runs the multimodal deploy smoke test.
|
||||
cmd_parts.append("--vlm")
|
||||
cmd_parts = extend_cmd_parts(cmd_parts, **kwargs)
|
||||
run_example_command(cmd_parts, "llm_ptq")
|
||||
run_example_command(cmd_parts, "hf_ptq")
|
||||
|
||||
+5
-5
@@ -12,10 +12,10 @@
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
"""Unit tests for ``examples/llm_ptq/cast_mxfp4_to_nvfp4.py``.
|
||||
"""Unit tests for ``examples/hf_ptq/cast_mxfp4_to_nvfp4.py``.
|
||||
|
||||
The module lives next to the example script (not inside the ``modelopt`` package),
|
||||
so we add ``examples/llm_ptq/`` to ``sys.path`` before importing it.
|
||||
so we add ``examples/hf_ptq/`` to ``sys.path`` before importing it.
|
||||
"""
|
||||
|
||||
import json
|
||||
@@ -26,9 +26,9 @@ import pytest
|
||||
import torch
|
||||
from safetensors.torch import save_file
|
||||
|
||||
_LLM_PTQ_DIR = Path(__file__).resolve().parents[3] / "examples" / "llm_ptq"
|
||||
if str(_LLM_PTQ_DIR) not in sys.path:
|
||||
sys.path.insert(0, str(_LLM_PTQ_DIR))
|
||||
_HF_PTQ_DIR = Path(__file__).resolve().parents[3] / "examples" / "hf_ptq"
|
||||
if str(_HF_PTQ_DIR) not in sys.path:
|
||||
sys.path.insert(0, str(_HF_PTQ_DIR))
|
||||
|
||||
import cast_mxfp4_to_nvfp4 as cast
|
||||
|
||||
+2
-2
@@ -12,7 +12,7 @@
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
"""End-to-end unit tests for ``examples/llm_ptq/example_utils.load_mtp_weights``.
|
||||
"""End-to-end unit tests for ``examples/hf_ptq/example_utils.load_mtp_weights``.
|
||||
|
||||
One test per supported on-disk MTP convention (inlined-orphaned, inlined-in-state-dict,
|
||||
separate-file-standalone, separate-file-indexed) plus a negative case.
|
||||
@@ -22,7 +22,7 @@ import json
|
||||
from types import SimpleNamespace
|
||||
|
||||
import torch
|
||||
from _test_utils.examples.llm_ptq_example_utils import example_utils
|
||||
from _test_utils.examples.hf_ptq_example_utils import example_utils
|
||||
from safetensors.torch import save_file
|
||||
|
||||
|
||||
+1
-1
@@ -20,7 +20,7 @@ from types import SimpleNamespace
|
||||
|
||||
import pytest
|
||||
|
||||
_EXAMPLES_DIR = Path(__file__).resolve().parents[3] / "examples" / "llm_ptq"
|
||||
_EXAMPLES_DIR = Path(__file__).resolve().parents[3] / "examples" / "hf_ptq"
|
||||
|
||||
|
||||
def _import_hf_ptq(monkeypatch):
|
||||
@@ -14,7 +14,7 @@
|
||||
# limitations under the License.
|
||||
import pytest
|
||||
import transformers
|
||||
from _test_utils.examples.llm_ptq_utils import PTQCommand
|
||||
from _test_utils.examples.hf_ptq_utils import PTQCommand
|
||||
from _test_utils.examples.models import (
|
||||
BART_PATH,
|
||||
MIXTRAL_PATH,
|
||||
@@ -15,9 +15,9 @@
|
||||
|
||||
import pytest
|
||||
from _test_utils.examples.models import QWEN_VL_PATH
|
||||
from _test_utils.examples.run_command import run_llm_ptq_command
|
||||
from _test_utils.examples.run_command import run_hf_ptq_command
|
||||
|
||||
|
||||
@pytest.mark.parametrize("quant", ["fp8", "int8_sq", "nvfp4"])
|
||||
def test_qwen_vl(quant):
|
||||
run_llm_ptq_command(model=QWEN_VL_PATH, quant=quant, vlm=True)
|
||||
run_hf_ptq_command(model=QWEN_VL_PATH, quant=quant, vlm=True)
|
||||
@@ -19,7 +19,7 @@ import pytest
|
||||
from _test_utils.examples.run_command import (
|
||||
extend_cmd_parts,
|
||||
run_example_command,
|
||||
run_llm_ptq_command,
|
||||
run_hf_ptq_command,
|
||||
)
|
||||
from _test_utils.torch.misc import minimum_sm
|
||||
from _test_utils.torch.transformers_models import create_tiny_qwen3_dir
|
||||
@@ -49,7 +49,7 @@ def test_qwen3_eval_fp8(tmp_path):
|
||||
# so 8192 leaves headroom for the longest prompts we evaluate.
|
||||
model_dir = create_tiny_qwen3_dir(tmp_path, with_tokenizer=True, max_position_embeddings=8192)
|
||||
try:
|
||||
run_llm_ptq_command(
|
||||
run_hf_ptq_command(
|
||||
model=str(model_dir),
|
||||
quant="fp8",
|
||||
tasks="mmlu,lm_eval,simple_eval",
|
||||
|
||||
@@ -132,7 +132,7 @@ def test_offline_ptq(offline_ptq_dirs):
|
||||
"--export_path",
|
||||
str(offline_ptq_dirs["ptq_export"]),
|
||||
],
|
||||
"llm_ptq",
|
||||
"hf_ptq",
|
||||
)
|
||||
|
||||
# Verify the exported checkpoint exists and has the expected EAGLE keys
|
||||
|
||||
@@ -108,7 +108,7 @@ def test_unified_hf_export_and_check_safetensors(
|
||||
env = os.environ.copy()
|
||||
if expected_suffix.startswith("t5_tiny"):
|
||||
env["CUDA_VISIBLE_DEVICES"] = "0"
|
||||
run_example_command(cmd_parts, "llm_ptq", env=env)
|
||||
run_example_command(cmd_parts, "hf_ptq", env=env)
|
||||
|
||||
# Now we expect a file named model.safetensors in output_dir
|
||||
generated_file = output_dir / "model.safetensors"
|
||||
|
||||
@@ -17,7 +17,7 @@
|
||||
|
||||
This exercises the exact combination that silently regressed (nvbug 6295279 / 6295242):
|
||||
a transposed-quantize MoE (``GptOssExperts``) with **static-block** NVFP4 weight quantizers
|
||||
-- which is what ``examples/llm_ptq --cast_mxfp4_to_nvfp4`` produces via
|
||||
-- which is what ``examples/hf_ptq --cast_mxfp4_to_nvfp4`` produces via
|
||||
``force_weight_quantizers_static`` -- calibrated with a forward loop.
|
||||
|
||||
The regression (the unconditional ``weight_only_quantize`` from #1560 feeding the
|
||||
@@ -43,9 +43,9 @@ from modelopt.torch.quantization.config import NVFP4_DEFAULT_CFG
|
||||
from modelopt.torch.quantization.nn import NVFP4StaticQuantizer
|
||||
|
||||
# The cast helpers live next to the example script, not in the ``modelopt`` package.
|
||||
_LLM_PTQ_DIR = Path(__file__).resolve().parents[4] / "examples" / "llm_ptq"
|
||||
if str(_LLM_PTQ_DIR) not in sys.path:
|
||||
sys.path.insert(0, str(_LLM_PTQ_DIR))
|
||||
_HF_PTQ_DIR = Path(__file__).resolve().parents[4] / "examples" / "hf_ptq"
|
||||
if str(_HF_PTQ_DIR) not in sys.path:
|
||||
sys.path.insert(0, str(_HF_PTQ_DIR))
|
||||
|
||||
from cast_mxfp4_to_nvfp4 import apply_to_model, force_weight_quantizers_static
|
||||
|
||||
|
||||
@@ -41,7 +41,7 @@ bash tools/debugger/client.sh --timeout 1800 run "<command>"
|
||||
|
||||
```bash
|
||||
# Run PTQ test
|
||||
bash tools/debugger/client.sh run "bash llm_ptq/scripts/huggingface_example.sh"
|
||||
bash tools/debugger/client.sh run "bash hf_ptq/scripts/huggingface_example.sh"
|
||||
|
||||
# Run pytest
|
||||
bash tools/debugger/client.sh run "python -m pytest tests/gpu -k test_quantize"
|
||||
|
||||
@@ -47,7 +47,7 @@ bash tools/debugger/client.sh handshake
|
||||
bash tools/debugger/client.sh run "echo hello"
|
||||
|
||||
# Run a test script
|
||||
bash tools/debugger/client.sh run "bash llm_ptq/scripts/huggingface_example.sh"
|
||||
bash tools/debugger/client.sh run "bash hf_ptq/scripts/huggingface_example.sh"
|
||||
|
||||
# Run with a long timeout (default is 600s)
|
||||
bash tools/debugger/client.sh --timeout 1800 run "python my_long_test.py"
|
||||
|
||||
@@ -22,6 +22,6 @@ trap 'error_handler $0 $LINENO' ERR
|
||||
|
||||
###################################################################################################
|
||||
|
||||
python modules/Model-Optimizer/examples/llm_ptq/hf_ptq.py \
|
||||
python modules/Model-Optimizer/examples/hf_ptq/hf_ptq.py \
|
||||
--model ${HF_MODEL_CKPT} \
|
||||
${@}
|
||||
|
||||
@@ -64,7 +64,7 @@ fi
|
||||
|
||||
# --- Step 2: Run huggingface_example.sh ---
|
||||
script_dir="$(dirname "$(readlink -f "$0")")"
|
||||
HF_EXAMPLE="${script_dir}/../../modules/Model-Optimizer/examples/llm_ptq/scripts/huggingface_example.sh"
|
||||
HF_EXAMPLE="${script_dir}/../../modules/Model-Optimizer/examples/hf_ptq/scripts/huggingface_example.sh"
|
||||
|
||||
echo "Running huggingface_example.sh --model $LOCAL_DIR --trust_remote_code ${PTQ_ARGS[*]}"
|
||||
exec bash "$HF_EXAMPLE" --model "$LOCAL_DIR" --trust_remote_code "${PTQ_ARGS[@]}"
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# HuggingFace PTQ via huggingface_example.sh
|
||||
#
|
||||
# Quantizes a HuggingFace model using examples/llm_ptq/scripts/huggingface_example.sh.
|
||||
# Quantizes a HuggingFace model using examples/hf_ptq/scripts/huggingface_example.sh.
|
||||
# Default: Qwen/Qwen3.5-9B with nvfp4_mlp_only on 8xH200.
|
||||
#
|
||||
# Usage (Slurm):
|
||||
@@ -11,14 +11,14 @@
|
||||
# export SLURM_HF_LOCAL=/home/scratch.<user>/hf-local
|
||||
# export HF_TOKEN=<your-hf-token> # for gated models; auto-injected into all tasks
|
||||
# cd tools/launcher
|
||||
# uv run launch.py --yaml examples/llm_ptq/hf_ptq.yaml --yes
|
||||
# uv run launch.py --yaml examples/hf_ptq/hf_ptq.yaml --yes
|
||||
#
|
||||
# Usage (local Docker):
|
||||
# cd tools/launcher
|
||||
# uv run launch.py --yaml examples/llm_ptq/hf_ptq.yaml hf_local=/mnt/hf-local --yes
|
||||
# uv run launch.py --yaml examples/hf_ptq/hf_ptq.yaml hf_local=/mnt/hf-local --yes
|
||||
#
|
||||
# Override model/quant via CLI:
|
||||
# uv run launch.py --yaml examples/llm_ptq/hf_ptq.yaml \
|
||||
# uv run launch.py --yaml examples/hf_ptq/hf_ptq.yaml \
|
||||
# pipeline.global_vars.hf_model=Qwen/Qwen3-8B \
|
||||
# pipeline.task_0.args='[--model,<<global_vars.hf_local>>Qwen/Qwen3-8B,--,--quant,nvfp4]' \
|
||||
# --yes
|
||||
|
||||
Reference in New Issue
Block a user