refactor(examples): rename llm_ptq → hf_ptq (symlink for back-compat) (#1759)

## What does this PR do?

**Type of change:** refactor / deprecation (examples)

Follow-up to #1705 (which consolidated `examples/vlm_ptq` into
`examples/llm_ptq`). Since that example now covers Hugging Face **LLM
and VLM** PTQ, the `llm_ptq` name is a misnomer. This renames the
directory to `examples/hf_ptq` and leaves a relative symlink
`examples/llm_ptq → hf_ptq` so existing paths/commands keep working
during a deprecation window.

Requested by @kevalmorabia97 on #1705 (with the symlink-for-back-compat
approach), targeted for the **same 0.46 release** as the consolidation.

### Changes
- `git mv examples/llm_ptq → examples/hf_ptq` and
`tests/examples/llm_ptq → tests/examples/hf_ptq` (the CI runner maps the
matrix name to both `examples/<name>` and `tests/examples/<name>`).
- Add a tracked back-compat symlink `examples/llm_ptq → hf_ptq`.
- Update CI matrices and all repo **path references** (docs, READMEs,
agent skills, launcher/debugger tools, tests) from `llm_ptq` to
`hf_ptq`.
- Keep Python identifiers / test-util module names
(`run_llm_ptq_command`, `llm_ptq_utils`) — they name the LLM-PTQ task,
not the directory.
- Preserve the CODEOWNERS team slug
(`modelopt-examples-llm_ptq-codeowners`) and historical CHANGELOG
entries; add a CHANGELOG deprecation note.

### Back-compat caveats (inherent to git directory symlinks)
- ✅ Linux/macOS CLI usage and Python `cwd`/pytest resolution work
through the symlink.
- ⚠️ Windows git checkouts don't materialize symlinks by default (low
impact — this example is Linux-only in practice).
- ⚠️ GitHub web doesn't follow directory symlinks, so legacy external
deep-links to `examples/llm_ptq/...` won't navigate in. All **internal**
references are repointed to `hf_ptq`, so the symlink is only for legacy
external/CLI use.

### Usage (unchanged via symlink)
```bash
# New canonical path
cd examples/hf_ptq
scripts/huggingface_example.sh --model <hf_model> --quant fp8

# Old path still works (forwards via symlink)
cd examples/llm_ptq && scripts/huggingface_example.sh --model <hf_model> --quant fp8
```

### Testing
- `bash -n` on moved/edited shell scripts (new path + via symlink).
- `py_compile` on moved/edited Python; test re-export shim repointed to
`examples/hf_ptq/example_utils`.
- Verified git tracks `examples/llm_ptq` as a single symlink (mode
120000), not a duplicated tree (no pre-commit / pytest
double-processing).
- `pre-commit run` on all changed files passes.

### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (relative symlink keeps
`examples/llm_ptq` paths valid; see caveats above)
- Did you write any new necessary tests?: N/A (pure rename; existing
tests moved with the dir)
- Did you update Changelog?: ✅

### Additional Information
Follow-up (later release): remove the `examples/llm_ptq` symlink once
external references have migrated.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* PTQ guidance now directs to the unified Hugging Face PTQ flow,
including VLM quantization via the shared `--vlm` entry point.
* **Documentation**
* Updated README and guide links, references, and command snippets to
use `hf_ptq` (replacing `llm_ptq`).
* Deprecated and consolidated `vlm_ptq` into `hf_ptq`; removed
VILA/NVILA coverage from the Hugging Face PTQ examples.
* **Bug Fixes**
* Improved detection and routing so local/manual setup uses the correct
PTQ source.
* **Tests / Chores**
  * CI and example tests updated to run the `hf_ptq` variants.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Zhiyu
2026-06-27 08:48:48 +00:00
committed by GitHub
co-authored by Claude Opus 4.8
parent c248dd5434
commit f335459dc0
67 changed files with 109 additions and 109 deletions
+1 -1
View File
@@ -5,7 +5,7 @@ Common detection for all ModelOpt skills. After this, you know what's available.
## Env-1. Get ModelOpt source
```bash
ls examples/llm_ptq/hf_ptq.py 2>/dev/null && echo "Source found"
ls examples/hf_ptq/hf_ptq.py 2>/dev/null && echo "Source found"
```
If not found: `git clone https://github.com/NVIDIA/Model-Optimizer.git && cd Model-Optimizer`
@@ -62,4 +62,4 @@ This matrix covers officially validated combinations. For unlisted models:
- **NVFP4 inference requires Blackwell GPUs** (B100, B200, GB200). Hopper can run FP4 calibration but not inference.
- INT4_AWQ and W4A8_AWQ are only supported by TRT-LLM (not vLLM or SGLang).
- Source: `examples/llm_ptq/README.md` and `docs/source/deployment/3_unified_hf.rst`
- Source: `examples/hf_ptq/README.md` and `docs/source/deployment/3_unified_hf.rst`
+7 -7
View File
@@ -5,7 +5,7 @@ description: This skill should be used when the user asks to "quantize a model",
# ModelOpt Post-Training Quantization
Produce a quantized checkpoint from a pretrained model. **Read `examples/llm_ptq/README.md` first** — it has the support matrix, CLI flags, and accuracy guidance.
Produce a quantized checkpoint from a pretrained model. **Read `examples/hf_ptq/README.md` first** — it has the support matrix, CLI flags, and accuracy guidance.
## Step 1 — Environment
@@ -19,7 +19,7 @@ Read `skills/common/environment-setup.md` and `skills/common/workspace-managemen
## Step 2 — Is the model supported?
Check the support table in `examples/llm_ptq/README.md` for verified HF models.
Check the support table in `examples/hf_ptq/README.md` for verified HF models.
- **Listed** → supported, use `hf_ptq.py` (step 4A/4B)
- **Not listed** → read `references/unsupported-models.md` to determine if `hf_ptq.py` can still work or if a custom script is needed (step 4C)
@@ -53,7 +53,7 @@ ls modelopt_recipes/huggingface/<model_type>/ptq/ 2>/dev/null # per-arch; <mode
If a model-specific recipe exists, prefer `--recipe <path>` — but **inspect its include/exclude patterns** rather than assuming (e.g. for VLMs, confirm the vision tower is actually excluded).
**If no model-specific recipe**, choose a format based on GPU (details in `examples/llm_ptq/README.md`):
**If no model-specific recipe**, choose a format based on GPU (details in `examples/hf_ptq/README.md`):
- **Blackwell** (B100/B200/GB200): `nvfp4` variants
- **Hopper** (H100/H200) or older: `fp8` or `int4_awq`
@@ -90,9 +90,9 @@ In README table? ─→ YES ──→ SLURM (local or remote)? ──→ LAUNCHE
```bash
pip install --no-build-isolation "nvidia-modelopt[hf]"
pip install -r examples/llm_ptq/requirements.txt
pip install -r examples/hf_ptq/requirements.txt
python examples/llm_ptq/hf_ptq.py \
python examples/hf_ptq/hf_ptq.py \
--pyt_ckpt_path <model> \
--qformat <format> \
--calib_size 512 \
@@ -105,7 +105,7 @@ For remote: use `remote_run` from `remote_exec.sh` (see `skills/common/remote-ex
### 4B — Launcher: supported model on SLURM or local Docker
Write a YAML config using `common/hf_ptq/hf_ptq.sh`. See `references/launcher-guide.md` for the full template.
Write a YAML config using `common/hf/ptq.sh`. See `references/launcher-guide.md` for the full template.
```bash
cd tools/launcher
@@ -179,7 +179,7 @@ Report the gate result before moving on. The report must include source size, ou
| `skills/common/remote-execution.md` | Step 4A/4C only, if target is remote |
| `skills/common/slurm-setup.md` | Step 4A/4C only, if using SLURM manually (not launcher) |
| `references/slurm-setup-ptq.md` | Step 4A/4C only, PTQ-specific SLURM (container, GPU sizing, FSDP2) |
| `examples/llm_ptq/README.md` | Step 3: support matrix, CLI flags, accuracy |
| `examples/hf_ptq/README.md` | Step 3: support matrix, CLI flags, accuracy |
| `modelopt/torch/quantization/config.py` | Step 3: format definitions |
| `modelopt/torch/export/model_utils.py` | Step 4C: TRT-LLM export type mapping |
| `modelopt_recipes/` | Step 3: pre-built recipes |
@@ -7,7 +7,7 @@ monitoring), see `skills/common/slurm-setup.md`.
## 1. Container
Get the recommended image version from `examples/llm_ptq/README.md`, then look for an existing `.sqsh` file:
Get the recommended image version from `examples/hf_ptq/README.md`, then look for an existing `.sqsh` file:
```bash
ls *.sqsh ../*.sqsh ~/containers/*.sqsh 2>/dev/null
@@ -63,17 +63,17 @@ pip install -U transformers --no-deps
Estimate GPU count from model size and available GPU memory. `hf_ptq.py` uses `device_map="auto"` so it fills GPUs automatically — request only as many as needed.
For multi-node PTQ (200B+ params), use `examples/llm_ptq/multinode_ptq.py` with FSDP2 and accelerate:
For multi-node PTQ (200B+ params), use `examples/hf_ptq/multinode_ptq.py` with FSDP2 and accelerate:
```bash
accelerate launch \
--config_file examples/llm_ptq/fsdp2.yaml \
--config_file examples/hf_ptq/fsdp2.yaml \
--num_machines $NUM_NODES \
--num_processes $((NUM_NODES * GPUS_PER_NODE)) \
--main_process_ip $MASTER_ADDR \
--main_process_port $MASTER_PORT \
--machine_rank $SLURM_PROCID \
examples/llm_ptq/multinode_ptq.py \
examples/hf_ptq/multinode_ptq.py \
--pyt_ckpt_path <model> \
--qformat <format> \
--export_path <output>
@@ -1,6 +1,6 @@
# Handling Unlisted Models
The model is not in the verified support table (`examples/llm_ptq/README.md`). This does NOT mean it won't work — ModelOpt auto-detects standard HF modules (linear layers, attention, MoE blocks with `gate`+`experts`). Many unlisted models work with `hf_ptq.py` out of the box.
The model is not in the verified support table (`examples/hf_ptq/README.md`). This does NOT mean it won't work — ModelOpt auto-detects standard HF modules (linear layers, attention, MoE blocks with `gate`+`experts`). Many unlisted models work with `hf_ptq.py` out of the box.
Follow the investigation steps below to determine if `hf_ptq.py` works or if patches are needed.
@@ -147,7 +147,7 @@ class QuantCustomModule(OriginalModule):
| Fused 2D weights (experts stacked in rows) | Two-level expansion | `_QuantDbrxExpertGLU` |
| Fused weights + `forward(x, expert_id)` | Expand + reconstruct on export | `_QuantMoELinear` (Step3.5) |
For the full guide, see `examples/llm_ptq/moe.md`.
For the full guide, see `examples/hf_ptq/README.md`.
**Critical: always check the weight layout.** `nn.Linear` expects `(out_features, in_features)` — the last dimension must be `in_features`. If the fused tensor is `(num_experts, in_dim, out_dim)`, you must transpose (`.T`) when copying. Getting this wrong silently corrupts quantization scales. Inspect the original forward pass to determine which dimension is which.
+1 -2
View File
@@ -48,7 +48,7 @@ modelopt_recipes @NVIDIA/modelopt-recipes-codeowners
/examples/gpt-oss @NVIDIA/modelopt-examples-gpt-oss-codeowners
/examples/llm_distill @NVIDIA/modelopt-torch-distill-codeowners
/examples/llm_eval @NVIDIA/modelopt-examples-llm_ptq-codeowners
/examples/llm_ptq @NVIDIA/modelopt-examples-llm_ptq-codeowners
/examples/hf_ptq @NVIDIA/modelopt-examples-llm_ptq-codeowners
/examples/llm_qat @NVIDIA/modelopt-examples-llm_qat-codeowners
/examples/llm_sparsity @NVIDIA/modelopt-torch-sparsity-codeowners
/examples/megatron_bridge @NVIDIA/modelopt-examples-megatron-codeowners
@@ -59,7 +59,6 @@ modelopt_recipes @NVIDIA/modelopt-recipes-codeowners
/examples/specdec_bench @NVIDIA/modelopt-torch-speculative-codeowners
/examples/speculative_decoding @NVIDIA/modelopt-torch-speculative-codeowners
/examples/torch_onnx @NVIDIA/modelopt-onnx-codeowners
/examples/vlm_ptq @NVIDIA/modelopt-examples-vlm-codeowners
/examples/vllm_serve @NVIDIA/modelopt-examples-llm_ptq-codeowners
/examples/windows @NVIDIA/modelopt-windows-codeowners
+1 -1
View File
@@ -9,7 +9,7 @@ on:
required: true
type: string
example:
description: "Example name to test (e.g. 'llm_ptq')"
description: "Example name to test (e.g. 'hf_ptq')"
required: true
type: string
timeout_minutes:
+2 -2
View File
@@ -55,7 +55,7 @@ jobs:
strategy:
fail-fast: false
matrix:
example: [llm_ptq]
example: [hf_ptq]
uses: ./.github/workflows/_example_tests_runner.yml
secrets: inherit
with:
@@ -69,7 +69,7 @@ jobs:
strategy:
fail-fast: false
matrix:
example: [llm_eval, llm_ptq]
example: [llm_eval, hf_ptq]
uses: ./.github/workflows/_example_tests_runner.yml
secrets: inherit
with:
+3 -2
View File
@@ -11,8 +11,9 @@ Changelog
**Deprecations**
- Consolidated ``examples/vlm_ptq`` into ``examples/llm_ptq``. Vision-language model PTQ now shares the ``hf_ptq.py`` entry point and ``scripts/huggingface_example.sh``; pass ``--vlm`` to run the TensorRT-LLM multimodal quickstart smoke test. The ``examples/vlm_ptq/scripts/huggingface_example.sh`` entry point is deprecated: it now prints a warning and forwards to the ``llm_ptq`` script with ``--vlm``, and will be removed in a future release. See `examples/llm_ptq/README.md <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/llm_ptq#vlm-quantization>`__.
- Dropped VILA / NVILA vision-language model support in ``examples/llm_ptq``. VILA's modeling code requires ``transformers<=4.50.0``, which conflicts with ModelOpt's minimum supported ``transformers`` version. The VILA-specific bootstrap (repo clone, ``requirements-vila.txt``) and loading paths in ``example_utils.py`` have been removed.
- Renamed ``examples/llm_ptq`` to ``examples/hf_ptq`` to reflect that it covers Hugging Face LLM **and** VLM PTQ. A relative symlink ``examples/llm_ptq`` -> ``hf_ptq`` keeps existing paths and commands working; it will be removed in a future release. Please update references to the new ``examples/hf_ptq`` path.
- Consolidated ``examples/vlm_ptq`` into ``examples/hf_ptq``. Vision-language model PTQ now shares the ``hf_ptq.py`` entry point and ``scripts/huggingface_example.sh``; pass ``--vlm`` to run the TensorRT-LLM multimodal quickstart smoke test. The ``examples/vlm_ptq/scripts/huggingface_example.sh`` entry point is deprecated: it now prints a warning and forwards to the ``hf_ptq`` script with ``--vlm``, and will be removed in a future release. See `examples/hf_ptq/README.md <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/hf_ptq#vlm-quantization>`__.
- Dropped VILA / NVILA vision-language model support in ``examples/hf_ptq``. VILA's modeling code requires ``transformers<=4.50.0``, which conflicts with ModelOpt's minimum supported ``transformers`` version. The VILA-specific bootstrap (repo clone, ``requirements-vila.txt``) and loading paths in ``example_utils.py`` have been removed.
**New Features**
+6 -7
View File
@@ -30,7 +30,7 @@ Model Optimizer is also integrated with [NVIDIA Megatron-Bridge](https://github.
- [2026/05/13] [**Puzzletron**](./examples/puzzletron): A new algorithm for heterogeneous pruning & NAS of LLM and VLM models.
- [2026/04/15] Customer story: [Domyn compresses Colosseum-355B → 260B using ModelOpt's Minitron pruning + distillation](https://www.domyn.com/blog/domyn-large-the-journey-of-a-european-sovereign-ai-model-for-regulated-industries)
- [2026/03/17] Customer story: [Bielik.AI builds Bielik Minitron 7B (33% smaller, 50% faster, 90% quality retained) using ModelOpt's Minitron pruning + distillation](https://bielik.ai/en/nvidia-gtc-bielik-minitron-premiere/)
- [2026/03/11] Model Optimizer quantized Nemotron-3-Super checkpoints are available on Hugging Face for download: [FP8](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8), [NVFP4](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4). Learn more in the [Nemotron 3 Super release blog](https://blogs.nvidia.com/blog/nemotron-3-super-agentic-ai/). Check out how to quantize Nemotron 3 models for deployment acceleration [here](./examples/llm_ptq/README.md)
- [2026/03/11] Model Optimizer quantized Nemotron-3-Super checkpoints are available on Hugging Face for download: [FP8](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8), [NVFP4](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4). Learn more in the [Nemotron 3 Super release blog](https://blogs.nvidia.com/blog/nemotron-3-super-agentic-ai/). Check out how to quantize Nemotron 3 models for deployment acceleration [here](./examples/hf_ptq/README.md)
- [2026/03/11] [NeMo Megatron Bridge](https://github.com/NVIDIA-NeMo/Megatron-Bridge) now supports Nemotron-3-Super quantization (PTQ and QAT) and export workflows using the Model Optimizer library. See the [Quantization (PTQ and QAT) guide](https://github.com/NVIDIA-NeMo/Megatron-Bridge/blob/super-v3/docs/models/llm/nemotron3-super.md#quantization-ptq-and-qat) for FP8/NVFP4 quantization and HF export instructions.
- [2025/12/11] [BLOG: Top 5 AI Model Optimization Techniques for Faster, Smarter Inference](https://developer.nvidia.com/blog/top-5-ai-model-optimization-techniques-for-faster-smarter-inference/)
- [2025/12/08] NVIDIA TensorRT Model Optimizer is now officially rebranded as NVIDIA Model Optimizer.
@@ -42,10 +42,10 @@ Model Optimizer is also integrated with [NVIDIA Megatron-Bridge](https://github.
- [2025/06/24] [BLOG: Introducing NVFP4 for Efficient and Accurate Low-Precision Inference](https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/)
- [2025/05/14] [NVIDIA TensorRT Unlocks FP4 Image Generation for NVIDIA Blackwell GeForce RTX 50 Series GPUs](https://developer.nvidia.com/blog/nvidia-tensorrt-unlocks-fp4-image-generation-for-nvidia-blackwell-geforce-rtx-50-series-gpus/)
- [2025/04/21] [Adobe optimized deployment using Model-Optimizer + TensorRT leading to a 60% reduction in diffusion latency, a 40% reduction in total cost of ownership](https://developer.nvidia.com/blog/optimizing-transformer-based-diffusion-models-for-video-generation-with-nvidia-tensorrt/)
- [2025/04/05] [NVIDIA Accelerates Inference on Meta Llama 4 Scout and Maverick](https://developer.nvidia.com/blog/nvidia-accelerates-inference-on-meta-llama-4-scout-and-maverick/). Check out how to quantize Llama4 for deployment acceleration [here](./examples/llm_ptq/README.md#llama-4)
- [2025/04/05] [NVIDIA Accelerates Inference on Meta Llama 4 Scout and Maverick](https://developer.nvidia.com/blog/nvidia-accelerates-inference-on-meta-llama-4-scout-and-maverick/). Check out how to quantize Llama4 for deployment acceleration [here](./examples/hf_ptq/README.md#support-matrix)
- [2025/03/18] [World's Fastest DeepSeek-R1 Inference with Blackwell FP4 & Increasing Image Generation Efficiency on Blackwell](https://developer.nvidia.com/blog/nvidia-blackwell-delivers-world-record-deepseek-r1-inference-performance/)
- [2025/02/25] Model Optimizer quantized NVFP4 models available on Hugging Face for download: [DeepSeek-R1-FP4](https://huggingface.co/nvidia/DeepSeek-R1-FP4), [Llama-3.3-70B-Instruct-FP4](https://huggingface.co/nvidia/Llama-3.3-70B-Instruct-FP4), [Llama-3.1-405B-Instruct-FP4](https://huggingface.co/nvidia/Llama-3.1-405B-Instruct-FP4)
- [2025/01/28] Model Optimizer has added support for NVFP4. Check out an example of NVFP4 PTQ [here](./examples/llm_ptq/README.md#model-quantization-and-trt-llm-conversion).
- [2025/01/28] Model Optimizer has added support for NVFP4. Check out an example of NVFP4 PTQ [here](./examples/hf_ptq/README.md#getting-started).
- [2025/01/28] Model Optimizer is now open source!
<details close>
@@ -56,7 +56,7 @@ Model Optimizer is also integrated with [NVIDIA Megatron-Bridge](https://github.
- [2024/08/28] [Boosting Llama 3.1 405B Performance up to 44% with Model Optimizer on NVIDIA H200 GPUs](https://developer.nvidia.com/blog/boosting-llama-3-1-405b-performance-by-up-to-44-with-nvidia-tensorrt-model-optimizer-on-nvidia-h200-gpus/)
- [2024/08/28] [Up to 1.9X Higher Llama 3.1 Performance with Medusa](https://developer.nvidia.com/blog/low-latency-inference-chapter-1-up-to-1-9x-higher-llama-3-1-performance-with-medusa-on-nvidia-hgx-h200-with-nvlink-switch/)
- [2024/08/15] New features in recent releases: [Cache Diffusion](./examples/diffusers/cache_diffusion), [QLoRA workflow with NVIDIA NeMo](https://docs.nvidia.com/nemo-framework/user-guide/24.09/sft_peft/qlora.html), and more. Check out [our blog](https://developer.nvidia.com/blog/nvidia-tensorrt-model-optimizer-v0-15-boosts-inference-performance-and-expands-model-support/) for details.
- [2024/06/03] Model Optimizer now has an experimental feature to deploy to vLLM as part of our effort to support popular deployment frameworks. Check out the workflow [here](./examples/llm_ptq/README.md#deploy-fp8-quantized-model-using-vllm)
- [2024/06/03] Model Optimizer now has an experimental feature to deploy to vLLM as part of our effort to support popular deployment frameworks. Check out the workflow [here](./examples/hf_ptq/README.md#vllm)
- [2024/05/08] [Announcement: Model Optimizer Now Formally Available to Further Accelerate GenAI Inference Performance](https://developer.nvidia.com/blog/accelerate-generative-ai-inference-performance-with-nvidia-tensorrt-model-optimizer-now-publicly-available/)
- [2024/03/27] [Model Optimizer supercharges TensorRT-LLM to set MLPerf LLM inference records](https://developer.nvidia.com/blog/nvidia-h200-tensor-core-gpus-and-nvidia-tensorrt-llm-set-mlperf-llm-inference-records/)
- [2024/03/18] [GTC Session: Optimize Generative AI Inference with Quantization in TensorRT-LLM and TensorRT](https://www.nvidia.com/en-us/on-demand/session/gtc24-s63213/)
@@ -102,7 +102,7 @@ more fine-grained control on installed dependencies or for alternative docker im
| **Technique** | **Description** | **Examples** | **Docs** |
| :------------: | :------------: | :------------: | :------------: |
| Post Training Quantization | Compress model size by 2x-4x, speeding up inference while preserving model quality! | \[[HF LLMs / VLMs](./examples/llm_ptq/)\] \[[Megatron-Bridge LLMs / VLMs](./examples/megatron_bridge/)\] \[[Diffusers](./examples/diffusers/)\] \[[ONNX](./examples/onnx_ptq/)\] \[[Windows](./examples/windows/)\] | \[[docs](https://nvidia.github.io/Model-Optimizer/guides/1_quantization.html)\] |
| Post Training Quantization | Compress model size by 2x-4x, speeding up inference while preserving model quality! | \[[HF LLMs / VLMs](./examples/hf_ptq/)\] \[[Megatron-Bridge LLMs / VLMs](./examples/megatron_bridge/)\] \[[Diffusers](./examples/diffusers/)\] \[[ONNX](./examples/onnx_ptq/)\] \[[Windows](./examples/windows/)\] | \[[docs](https://nvidia.github.io/Model-Optimizer/guides/1_quantization.html)\] |
| Quantization Aware Training / Distillation | Refine accuracy of quantized models even further with a few training steps! | \[[Hugging Face](./examples/llm_qat/)\] \[[Megatron-Bridge](./examples/megatron_bridge)\] | \[[docs](https://nvidia.github.io/Model-Optimizer/guides/1_quantization.html)\] |
| Pruning | Reduce your model parameters or memory footprint and accelerate inference by removing unnecessary weights! | \[[General](./examples/pruning/)\] \[[Megatron-Bridge](./examples/megatron_bridge/)\] | |
| Distillation | Reduce deployment model size by teaching small models to behave like larger models! | \[[Hugging Face](./examples/llm_distill/)\] \[[Megatron-Bridge](./examples/megatron_bridge/)\] \[[Megatron-LM](./examples/llm_distill/README.md#knowledge-distillation-kd-in-nvidia-megatron-lm-framework)\] | \[[docs](https://nvidia.github.io/Model-Optimizer/guides/4_distillation.html)\] |
@@ -130,8 +130,7 @@ more fine-grained control on installed dependencies or for alternative docker im
| Model Type | Support Matrix |
|------------|----------------|
| LLM Quantization | [View Support Matrix](./examples/llm_ptq/README.md#support-matrix) |
| VLM Quantization | [View Support Matrix](./examples/llm_ptq/README.md#hugging-face-supported-models) |
| LLM / VLM Quantization | [View Support Matrix](./examples/hf_ptq/README.md#support-matrix) |
| Diffusers Quantization | [View Support Matrix](./examples/diffusers/README.md#support-matrix) |
| ONNX Quantization | [View Support Matrix](./examples/torch_onnx/README.md#onnx-export-supported-llm-models) |
| Windows Quantization | [View Support Matrix](./examples/windows/README.md#support-matrix) |
+1 -1
View File
@@ -6,7 +6,7 @@ We support exporting modelopt-optimized Hugging Face models (transformers and di
The workflow is as follows:
#. Load the Huggingface models or Megatron Core models, `quantize with modelopt <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/llm_ptq#ptq-post-training-quantization>`_ , and export to the unified checkpoint format, where the layer structures and tensor names are aligned with the original checkpoint.
#. Load the Huggingface models or Megatron Core models, `quantize with modelopt <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/hf_ptq#ptq-post-training-quantization>`_ , and export to the unified checkpoint format, where the layer structures and tensor names are aligned with the original checkpoint.
#. Load the unified checkpoint in the supported inference framework for accelerated inference.
+1 -1
View File
@@ -570,7 +570,7 @@ Some example scripts accept a ``--recipe`` flag. For instance, the PTQ example:
.. code-block:: bash
python examples/llm_ptq/hf_ptq.py \
python examples/hf_ptq/hf_ptq.py \
--model Qwen/Qwen3-8B \
--recipe general/ptq/fp8_default-kv_fp8_cast \
--export_path build/fp8 \
@@ -58,4 +58,4 @@ For quantized formats like NVFP4, you can reduce memory usage by up to 4x compar
.. note::
An example implementation of this workflow can be found in:
``examples/llm_ptq/hf_ptq.py``, which reduces the memory requirements of model calibration.
``examples/hf_ptq/hf_ptq.py``, which reduces the memory requirements of model calibration.
@@ -15,7 +15,7 @@ As ModelOpt cannot detect these linear ops out-of-the-box, a HugggingFace plugin
#. Define a customized ``_QuantDbrxExpertGLU`` as a ``DynamicModule`` with the same ``forward`` signature.
#. Rewrite the linear ops (w1, v1 and v2) as a standard ``nn.Linear`` op, and re-implement the ``forward`` method.
#. Register the new dynamic ``_QuantDbrxExperts`` to replace the ``DbrxExperts`` from the modeling_dbrx.py in the ``transformers`` library
#. Try quantize the DBRX model after the plugin is implemented, feel free to follow the `llm_ptq example <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/llm_ptq>`_.
#. Try quantize the DBRX model after the plugin is implemented, feel free to follow the `hf_ptq example <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/hf_ptq>`_.
#. TensorRT-LLM is open-sourced. If this customized model is not supported by TensorRT-LLM yet, please modify :meth:`export_tensorrt_llm_checkpoint <modelopt.torch.export.export_tensorrt_llm_checkpoint>` or :meth:`export_hf_checkpoint <modelopt.torch.export.export_hf_checkpoint>` to export the quantized model for deployment with a customized TensorRT-LLM modeling implementation. Feel free to :doc:`contact us <../support/1_contact>` if further support is needed.
The following code snippet is excerpted from ``modelopt/torch/quantization/plugins/huggingface.py``
+1 -1
View File
@@ -7,7 +7,7 @@ Welcome to Model Optimizer (ModelOpt) documentation!
:caption: Getting Started
getting_started/[0-9]*
Quick Start: PTQ - PyTorch <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/llm_ptq>
Quick Start: PTQ - PyTorch <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/hf_ptq>
Quick Start: PTQ - ONNX <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/onnx_ptq>
Quick Start: PTQ - PyTorch to ONNX <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/torch_onnx>
Quick Start: PTQ - Windows <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/windows>
+1 -1
View File
@@ -193,6 +193,6 @@ lands in E4M3's representable window; the rare out-of-range block falls back to
data-derived scale). The flag only affects routed-expert **weights** — activation
`input_scale` still comes from `${AMAX}` calibration — and the run prints a
`[cast] lossless MXFP4->NVFP4 blocks: …` summary. This mirrors the GPTOSS cast in
[`examples/llm_ptq/cast_mxfp4_to_nvfp4.py`](../llm_ptq/cast_mxfp4_to_nvfp4.py); the
[`examples/hf_ptq/cast_mxfp4_to_nvfp4.py`](../hf_ptq/cast_mxfp4_to_nvfp4.py); the
V4 twist is that w1/w3 share one `scale_2` (fused GEMM1), so `k_max` is taken over
both projections.
@@ -74,7 +74,7 @@ and the NVFP4 nibbles equal the source MXFP4 nibbles bit-for-bit (for every
block whose ``k_j`` lands in E4M3's representable window). The flag only affects
routed-expert *weights*; activation ``input_scale`` still comes from
``--amax_path`` calibration. This mirrors the GPTOSS cast in
``examples/llm_ptq/cast_mxfp4_to_nvfp4.py`` (PR #1372); the V4 twist is that
``examples/hf_ptq/cast_mxfp4_to_nvfp4.py`` (PR #1372); the V4 twist is that
w1/w3 share one ``scale_2`` (fused GEMM1), so ``k_max`` is taken over both.
Usage (single compute node, CPU-default; dequant+requant math is cheap
+1 -1
View File
@@ -184,4 +184,4 @@ ModelOpt provides easy end to end QAT via [LLaMA-Factory](https://github.com/hiy
### Deployment of ModelOpt QAT/PTQ models beyond GPT-OSS
ModelOpt supports exporting a wide variety of models after QAT/PTQ to TensorRT-LLM, vLLM, SGLang etc. Please refer to [llm_ptq](../llm_ptq).
ModelOpt supports exporting a wide variety of models after QAT/PTQ to TensorRT-LLM, vLLM, SGLang etc. Please refer to [hf_ptq](../hf_ptq).
@@ -1,5 +1,5 @@
# =============================================================================
# FSDP Configuration for running LLM PTQ on multinode setup. This file is consumed by examples/llm_ptq/multinode_ptq.py
# FSDP Configuration for running LLM PTQ on multinode setup. This file is consumed by examples/hf_ptq/multinode_ptq.py
# =============================================================================
compute_environment: LOCAL_MACHINE
+2 -2
View File
@@ -6,7 +6,7 @@ The following instructions show how to evaluate the Model Optimizer quantized LL
## NeMo Evaluator
[NeMo Evaluator](https://docs.nvidia.com/nemo/evaluator/latest/get-started/quickstart/index.html#self-hosted-options) is the recommended way to evaluate a large choice of benchmarks on quantized checkpoints generated from [llm_ptq](../llm_ptq). Quantized checkpoints can be served with [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM), [vLLM](https://github.com/vllm-project/vllm), or [SGLang](https://github.com/sgl-project/sglang) and then evaluated using NeMo Evaluator.
[NeMo Evaluator](https://docs.nvidia.com/nemo/evaluator/latest/get-started/quickstart/index.html#self-hosted-options) is the recommended way to evaluate a large choice of benchmarks on quantized checkpoints generated from [hf_ptq](../hf_ptq). Quantized checkpoints can be served with [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM), [vLLM](https://github.com/vllm-project/vllm), or [SGLang](https://github.com/sgl-project/sglang) and then evaluated using NeMo Evaluator.
## LM-Eval-Harness
@@ -233,7 +233,7 @@ This is useful for evaluating quantized models deployed with vLLM or any model s
--tensor-parallel-size <tp_size> # Adjust as needed
```
To generate the quantized model such as `nvidia/Llama-3.1-8B-Instruct-FP8`, please refer to instructions [here](https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/llm_ptq#deploy-fp8-quantized-model-using-vllm-and-sglang). Note currently modelopt quantized model support in vLLM is limited, we are working on expanding the model and quant formats support.
To generate the quantized model such as `nvidia/Llama-3.1-8B-Instruct-FP8`, please refer to instructions [here](https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/hf_ptq#deploy-fp8-quantized-model-using-vllm-and-sglang). Note currently modelopt quantized model support in vLLM is limited, we are working on expanding the model and quant formats support.
1. **Make the script executable (if not already):**
+1
View File
@@ -0,0 +1 @@
hf_ptq
+2 -2
View File
@@ -24,7 +24,7 @@ For background on how QAT enables low-precision accuracy recovery, see the [QAT/
### Prerequisites
Please refer to [llm_ptq/README.md](../llm_ptq/README.md#pre-requisites) for container
Please refer to [hf_ptq/README.md](../hf_ptq/README.md#pre-requisites) for container
recommendations and base ModelOpt installation guidance. For this QAT/QAD example,
install the Hugging Face dependencies and the example-specific requirements:
@@ -85,7 +85,7 @@ accelerate launch --config-file configs/accelerate/fsdp2.yaml train.py \
python export.py --pyt_ckpt_path qwen3-8b-qad-nvfp4 --export_path qwen3-8b-qad-deploy
```
Exported checkpoints can be deployed on [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM), [vLLM](https://github.com/vllm-project/vllm), or [SGLang](https://github.com/sgl-project/sglang). See [llm_ptq/README.md](../llm_ptq/README.md#deployment) for deployment instructions. For quick accuracy evaluation without exporting, see [Native Fake-Quantized Evaluation](#native-fake-quantized-evaluation).
Exported checkpoints can be deployed on [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM), [vLLM](https://github.com/vllm-project/vllm), or [SGLang](https://github.com/sgl-project/sglang). See [hf_ptq/README.md](../hf_ptq/README.md#deployment) for deployment instructions. For quick accuracy evaluation without exporting, see [Native Fake-Quantized Evaluation](#native-fake-quantized-evaluation).
> [!NOTE]
> For a minimal end-to-end demo (quantize + train + save in one script), see [simple_qat_train.py](simple_qat_train.py). It runs on a **single GPU** only and is intended as a quick introduction to the QAT flow (without transformer trainer)—not for distributed training.
+2 -2
View File
@@ -94,9 +94,9 @@ The final QAT/QAD model after training is similar in architecture to that of PTQ
To run QAT/QAD model with TRTLLM, run:
```sh
cd ../../llm_ptq
cd ../../hf_ptq
./scripts/huggingface_example.sh --model <path-to-QAT/QAD-model> --quant nvfp4
```
See more details on deployment of quantized model [here](../../llm_ptq/README.md).
See more details on deployment of quantized model [here](../../hf_ptq/README.md).
@@ -544,7 +544,7 @@
"cell_type": "markdown",
"id": "10acc50c-c876-41d5-8f7e-00dab8842ccd",
"metadata": {},
"source": "**Note:** The QAT checkpoint for `nvfp4` config can also be created using the CLI scripts. See the [QAT README](../README.md) for the full end-to-end workflow using `quantize.py`, `train.py`, and `export.py`.\n\nSee more details on deployment of quantized model [here](https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/llm_ptq/README.md)."
"source": "**Note:** The QAT checkpoint for `nvfp4` config can also be created using the CLI scripts. See the [QAT README](../README.md) for the full end-to-end workflow using `quantize.py`, `train.py`, and `export.py`.\n\nSee more details on deployment of quantized model [here](https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/hf_ptq/README.md)."
},
{
"cell_type": "markdown",
@@ -603,7 +603,7 @@
"metadata": {},
"source": [
"## Exporting Quantized Model for deployment\n",
"Before deploying the model with TensorRT-LLM you will need to export the model checkpoint files. This is similar to the step you take for a quantized PTQ Model. To export the unified Hugging Face checkpoints, which can be deployed on TensorRT-LLM Pytorch, vLLM and SGLang you will need to run the [huggingface_example.sh](https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/llm_ptq/scripts/huggingface_example.sh) script found in the Model Optimizer repo. "
"Before deploying the model with TensorRT-LLM you will need to export the model checkpoint files. This is similar to the step you take for a quantized PTQ Model. To export the unified Hugging Face checkpoints, which can be deployed on TensorRT-LLM Pytorch, vLLM and SGLang you will need to run the [huggingface_example.sh](https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/hf_ptq/scripts/huggingface_example.sh) script found in the Model Optimizer repo. "
]
},
{
@@ -671,7 +671,7 @@
"\n",
"# run conversion script\n",
"cd ..\n",
"bash Model-Optimizer/examples/llm_ptq/scripts/huggingface_example.sh --model $(pwd)/qat/checkpoint-450/ --quant nvfp4"
"bash Model-Optimizer/examples/hf_ptq/scripts/huggingface_example.sh --model $(pwd)/qat/checkpoint-450/ --quant nvfp4"
]
},
{
+1 -1
View File
@@ -73,7 +73,7 @@ from modelopt.torch.utils.plugins.megatron_calibration import get_megatron_calib
from modelopt.torch.utils.plugins.megatron_generate import megatron_generate
# The --quant_cfg / --kv_cache_quant CLI vocabularies are discovered from the preset
# YAMLs (shared with the llm_ptq examples via modelopt.recipe.presets). --quant_cfg
# YAMLs (shared with the hf_ptq examples via modelopt.recipe.presets). --quant_cfg
# additionally accepts any full config name from ``mtq.config.choices`` (e.g.
# ``FP8_DEFAULT_CFG``); see get_quant_config below.
+1 -1
View File
@@ -27,4 +27,4 @@ To deploy and run on SGLang:
python run_llama_fp8_sglang.py
```
If you want to run post-training quantization with Model Optimizer for your selected models, check [here](../llm_ptq/README.md).
If you want to run post-training quantization with Model Optimizer for your selected models, check [here](../hf_ptq/README.md).
@@ -9,7 +9,7 @@ End-to-end optimization of [Nemotron-Nano-9B-v2](https://huggingface.co/nvidia/N
2. **[Pruning](#2-pruning)** — Minitron structured pruning from 9B to 7B
3. **[Distillation](#3-distillation)** — recovering accuracy via Megatron-Bridge knowledge distillation (up to 80B tokens)
4. **[Evaluation](#4-evaluation)** — benchmarking with NeMo Evaluator across MMLU Pro, GPQA Diamond, AIME, and more
5. **[Quantization](#5-quantization)** — FP8 PTQ on the distilled checkpoint using ModelOpt's `examples/llm_ptq/hf_ptq.py` script
5. **[Quantization](#5-quantization)** — FP8 PTQ on the distilled checkpoint using ModelOpt's `examples/hf_ptq/hf_ptq.py` script
6. **[vLLM Inference Benchmarking](#6-vllm-inference-benchmarking)** — throughput comparison of BF16 vs FP8 on a single H100
## Results
@@ -317,11 +317,11 @@ For more details on NeMo Evaluator, see the [GitHub repo](https://github.com/NVI
### 5. Quantization
ModelOpt allows stacking multiple optimization techniques. Here we stack FP8 quantization on top of the pruned and distilled model to get an even more optimized model. See [examples/llm_ptq/README.md](../../../llm_ptq/README.md) for the full PTQ documentation.
ModelOpt allows stacking multiple optimization techniques. Here we stack FP8 quantization on top of the pruned and distilled model to get an even more optimized model. See [examples/hf_ptq/README.md](../../../hf_ptq/README.md) for the full PTQ documentation.
Similar to the official [Nemotron-Nano-9B-v2-FP8](https://huggingface.co/nvidia/NVIDIA-Nemotron-Nano-9B-v2-FP8) model, if you want to quantize the pruned 7B model to FP8, the Mamba and MLP layers are quantized to FP8, while all 4 attention layers and the Conv1d components within the Mamba layers are kept in BF16 to avoid accuracy degradation.
This is done with the `mtq.MAMBA_MOE_FP8_CONSERVATIVE_CFG` config defined in [`modelopt/torch/quantization/config.py`](../../../../modelopt/torch/quantization/config.py). To apply this, you need to modify `QUANT_CFG_CHOICES["fp8"]` in [`examples/llm_ptq/hf_ptq.py`](../../../llm_ptq/hf_ptq.py) to use `mtq.MAMBA_MOE_FP8_CONSERVATIVE_CFG`. For a faster model at the cost of a larger accuracy drop, you can use `mtq.MAMBA_MOE_FP8_AGGRESSIVE_CFG` instead.
This is done with the `mtq.MAMBA_MOE_FP8_CONSERVATIVE_CFG` config defined in [`modelopt/torch/quantization/config.py`](../../../../modelopt/torch/quantization/config.py). To apply this, you need to modify `QUANT_CFG_CHOICES["fp8"]` in [`examples/hf_ptq/hf_ptq.py`](../../../hf_ptq/hf_ptq.py) to use `mtq.MAMBA_MOE_FP8_CONSERVATIVE_CFG`. For a faster model at the cost of a larger accuracy drop, you can use `mtq.MAMBA_MOE_FP8_AGGRESSIVE_CFG` instead.
> [!NOTE]
> You can also quantize to NVFP4 using `mtq.MAMBA_MOE_NVFP4_CONSERVATIVE_CFG` (default) or `mtq.MAMBA_MOE_NVFP4_AGGRESSIVE_CFG` (faster, more accuracy drop), which may require further distillation (QAD) to recover accuracy and Blackwell GPU for deployment.
@@ -329,7 +329,7 @@ This is done with the `mtq.MAMBA_MOE_FP8_CONSERVATIVE_CFG` config defined in [`m
Calibrate and export the HF checkpoint from iteration 12800 to FP8 (takes 1-2 mins on 8x H100):
```bash
python /opt/Model-Optimizer/examples/llm_ptq/hf_ptq.py \
python /opt/Model-Optimizer/examples/hf_ptq/hf_ptq.py \
--pyt_ckpt_path <output_dir>/checkpoints/hf_iter_12800 \
--export_path <output_dir>/checkpoints/hf_iter_12800_fp8 \
--qformat fp8 \
+1 -1
View File
@@ -203,7 +203,7 @@ One can also use [examples/specdec_bench](../specdec_bench) to validate the trai
### Deploying Quantized model
See more details on deployment of quantized model to TRTLLM [here](../llm_ptq/README.md).
See more details on deployment of quantized model to TRTLLM [here](../hf_ptq/README.md).
## Advanced Usage
+2 -2
View File
@@ -65,10 +65,10 @@ lm_eval --model local-completions --tasks gsm8k --model_args model=<model_name>,
Step 1: export the model with bf16 weights and quantizer state. To export the model:
- For **HF** models, use `examples/llm_ptq/hf_ptq.py` with `--vllm_fakequant_export`:
- For **HF** models, use `examples/hf_ptq/hf_ptq.py` with `--vllm_fakequant_export`:
```bash
python ../llm_ptq/hf_ptq.py \
python ../hf_ptq/hf_ptq.py \
--pyt_ckpt_path <MODEL_PATH> \
--recipe <PATH_TO_RECIPE> \
--calib_size 512 \
+6 -6
View File
@@ -1,15 +1,15 @@
# [Deprecated] Post-training quantization (PTQ) for Vision Language Models
> **This example has been consolidated into [`examples/llm_ptq`](../llm_ptq/README.md) and is
> **This example has been consolidated into [`examples/hf_ptq`](../hf_ptq/README.md) and is
> deprecated.** It will be removed in a future release. VLM PTQ now shares the same entry point
> (`hf_ptq.py`) and shell script as LLM PTQ.
## Migration
Use the `llm_ptq` script with the `--vlm` flag:
Use the `hf_ptq` script with the `--vlm` flag:
```bash
cd examples/llm_ptq
cd examples/hf_ptq
scripts/huggingface_example.sh --model <Hugging Face model card or checkpoint> --quant [fp8|nvfp4|int8_sq|int4_awq|w4a8_awq] --vlm
```
@@ -20,9 +20,9 @@ prints a deprecation warning and forwards to the command above.
| Topic | New location |
| :--- | :--- |
| Supported VLMs / support matrix | [llm_ptq/README.md#hugging-face-supported-models](../llm_ptq/README.md#hugging-face-supported-models) |
| VLM quantization workflow (`--vlm`) | [llm_ptq/README.md#vlm-quantization](../llm_ptq/README.md#vlm-quantization) |
| Image-text calibration (`--calib_with_images`) | [llm_ptq/README.md#vlm-calibration-with-image-text-pairs-eg-nemotron-vl](../llm_ptq/README.md#vlm-calibration-with-image-text-pairs-eg-nemotron-vl) |
| Supported VLMs / support matrix | [hf_ptq/README.md#hugging-face-supported-models](../hf_ptq/README.md#hugging-face-supported-models) |
| VLM quantization workflow (`--vlm`) | [hf_ptq/README.md#vlm-quantization](../hf_ptq/README.md#vlm-quantization) |
| Image-text calibration (`--calib_with_images`) | [hf_ptq/README.md#vlm-calibration-with-image-text-pairs-eg-nemotron-vl](../hf_ptq/README.md#vlm-calibration-with-image-text-pairs-eg-nemotron-vl) |
| Megatron-Bridge VLM PTQ | [examples/megatron_bridge/](../megatron_bridge/README.md) |
## Resources
@@ -14,21 +14,21 @@
# See the License for the specific language governing permissions and
# limitations under the License.
# DEPRECATED: examples/vlm_ptq has been consolidated into examples/llm_ptq.
# This shim forwards all arguments to the llm_ptq script with the --vlm flag so existing
# DEPRECATED: examples/vlm_ptq has been consolidated into examples/hf_ptq.
# This shim forwards all arguments to the hf_ptq script with the --vlm flag so existing
# commands keep working. Please migrate to:
#
# cd examples/llm_ptq
# cd examples/hf_ptq
# scripts/huggingface_example.sh --model <model> --quant <qformat> --vlm
#
# See examples/llm_ptq/README.md#vlm-quantization for details.
# See examples/hf_ptq/README.md#vlm-quantization for details.
set -e
echo "WARNING: examples/vlm_ptq is deprecated and will be removed in a future release." >&2
echo " Forwarding to examples/llm_ptq/scripts/huggingface_example.sh --vlm" >&2
echo " See examples/llm_ptq/README.md#vlm-quantization" >&2
echo " Forwarding to examples/hf_ptq/scripts/huggingface_example.sh --vlm" >&2
echo " See examples/hf_ptq/README.md#vlm-quantization" >&2
script_dir="$(dirname "$(readlink -f "$0")")"
exec "$script_dir/../../llm_ptq/scripts/huggingface_example.sh" --vlm "$@"
exec "$script_dir/../../hf_ptq/scripts/huggingface_example.sh" --vlm "$@"
+2 -2
View File
@@ -15,8 +15,8 @@
"""PTQ quant-config preset discovery shared by the PTQ example scripts.
The example PTQ entry points (``examples/llm_ptq/hf_ptq.py``,
``examples/llm_ptq/multinode_ptq.py``, ``examples/megatron_bridge/quantize.py``)
The example PTQ entry points (``examples/hf_ptq/hf_ptq.py``,
``examples/hf_ptq/multinode_ptq.py``, ``examples/megatron_bridge/quantize.py``)
expose a ``--qformat`` / ``--kv_cache_qformat`` (``--quant_cfg`` /
``--kv_cache_quant`` for Megatron-Bridge) CLI vocabulary. Rather than hardcoding a
name → config table in each script, the vocabulary is discovered by listing the
@@ -19,7 +19,7 @@ These helpers turn an MXFP4 source layer's E8M0 block scales into the per-tensor
``global_amax`` and per-NVFP4-block ``amax`` that pin NVFP4's two-level scale so
the cast reproduces the source MXFP4 weights bit-for-bit (see PR #1372 for the
derivation). They are pure tensor math with no model or checkpoint dependencies,
shared by the GPT-OSS (``examples/llm_ptq``) and DeepSeek-V4
shared by the GPT-OSS (``examples/hf_ptq``) and DeepSeek-V4
(``examples/deepseek``) PTQ cast paths.
"""
@@ -12,8 +12,8 @@
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
"""Re-export ``examples/llm_ptq/example_utils`` so tests can import it via
``from _test_utils.examples.llm_ptq_example_utils import example_utils``
"""Re-export ``examples/hf_ptq/example_utils`` so tests can import it via
``from _test_utils.examples.hf_ptq_example_utils import example_utils``
without per-file ``sys.path`` shims.
"""
@@ -21,9 +21,9 @@ import sys
from _test_utils.examples.run_command import MODELOPT_ROOT
_LLM_PTQ_DIR = MODELOPT_ROOT / "examples" / "llm_ptq"
if str(_LLM_PTQ_DIR) not in sys.path:
sys.path.insert(0, str(_LLM_PTQ_DIR))
_HF_PTQ_DIR = MODELOPT_ROOT / "examples" / "hf_ptq"
if str(_HF_PTQ_DIR) not in sys.path:
sys.path.insert(0, str(_HF_PTQ_DIR))
import example_utils
@@ -19,7 +19,7 @@ from dataclasses import asdict, dataclass
import pytest
import torch
from _test_utils.examples.run_command import run_llm_ptq_command
from _test_utils.examples.run_command import run_hf_ptq_command
@dataclass
@@ -62,7 +62,7 @@ class PTQCommand:
param_dict.pop("min_gpu", None)
quant = param_dict.pop("quant")
run_llm_ptq_command(model=model_path, quant=quant, **param_dict)
run_hf_ptq_command(model=model_path, quant=quant, **param_dict)
def param_str(self):
param_dict = asdict(self)
+3 -3
View File
@@ -62,14 +62,14 @@ def run_command_in_background(
return process
def run_llm_ptq_command(*, model: str, quant: str, vlm: bool = False, **kwargs):
def run_hf_ptq_command(*, model: str, quant: str, vlm: bool = False, **kwargs):
kwargs.update({"model": model, "quant": quant})
kwargs.setdefault("tasks", "quant")
kwargs.setdefault("calib", 16)
cmd_parts = ["scripts/huggingface_example.sh", "--no-verbose"]
if vlm:
# VLM PTQ shares the llm_ptq entry point; --vlm runs the multimodal deploy smoke test.
# VLM PTQ shares the hf_ptq entry point; --vlm runs the multimodal deploy smoke test.
cmd_parts.append("--vlm")
cmd_parts = extend_cmd_parts(cmd_parts, **kwargs)
run_example_command(cmd_parts, "llm_ptq")
run_example_command(cmd_parts, "hf_ptq")
@@ -12,10 +12,10 @@
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
"""Unit tests for ``examples/llm_ptq/cast_mxfp4_to_nvfp4.py``.
"""Unit tests for ``examples/hf_ptq/cast_mxfp4_to_nvfp4.py``.
The module lives next to the example script (not inside the ``modelopt`` package),
so we add ``examples/llm_ptq/`` to ``sys.path`` before importing it.
so we add ``examples/hf_ptq/`` to ``sys.path`` before importing it.
"""
import json
@@ -26,9 +26,9 @@ import pytest
import torch
from safetensors.torch import save_file
_LLM_PTQ_DIR = Path(__file__).resolve().parents[3] / "examples" / "llm_ptq"
if str(_LLM_PTQ_DIR) not in sys.path:
sys.path.insert(0, str(_LLM_PTQ_DIR))
_HF_PTQ_DIR = Path(__file__).resolve().parents[3] / "examples" / "hf_ptq"
if str(_HF_PTQ_DIR) not in sys.path:
sys.path.insert(0, str(_HF_PTQ_DIR))
import cast_mxfp4_to_nvfp4 as cast
@@ -12,7 +12,7 @@
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
"""End-to-end unit tests for ``examples/llm_ptq/example_utils.load_mtp_weights``.
"""End-to-end unit tests for ``examples/hf_ptq/example_utils.load_mtp_weights``.
One test per supported on-disk MTP convention (inlined-orphaned, inlined-in-state-dict,
separate-file-standalone, separate-file-indexed) plus a negative case.
@@ -22,7 +22,7 @@ import json
from types import SimpleNamespace
import torch
from _test_utils.examples.llm_ptq_example_utils import example_utils
from _test_utils.examples.hf_ptq_example_utils import example_utils
from safetensors.torch import save_file
@@ -20,7 +20,7 @@ from types import SimpleNamespace
import pytest
_EXAMPLES_DIR = Path(__file__).resolve().parents[3] / "examples" / "llm_ptq"
_EXAMPLES_DIR = Path(__file__).resolve().parents[3] / "examples" / "hf_ptq"
def _import_hf_ptq(monkeypatch):
@@ -14,7 +14,7 @@
# limitations under the License.
import pytest
import transformers
from _test_utils.examples.llm_ptq_utils import PTQCommand
from _test_utils.examples.hf_ptq_utils import PTQCommand
from _test_utils.examples.models import (
BART_PATH,
MIXTRAL_PATH,
@@ -15,9 +15,9 @@
import pytest
from _test_utils.examples.models import QWEN_VL_PATH
from _test_utils.examples.run_command import run_llm_ptq_command
from _test_utils.examples.run_command import run_hf_ptq_command
@pytest.mark.parametrize("quant", ["fp8", "int8_sq", "nvfp4"])
def test_qwen_vl(quant):
run_llm_ptq_command(model=QWEN_VL_PATH, quant=quant, vlm=True)
run_hf_ptq_command(model=QWEN_VL_PATH, quant=quant, vlm=True)
+2 -2
View File
@@ -19,7 +19,7 @@ import pytest
from _test_utils.examples.run_command import (
extend_cmd_parts,
run_example_command,
run_llm_ptq_command,
run_hf_ptq_command,
)
from _test_utils.torch.misc import minimum_sm
from _test_utils.torch.transformers_models import create_tiny_qwen3_dir
@@ -49,7 +49,7 @@ def test_qwen3_eval_fp8(tmp_path):
# so 8192 leaves headroom for the longest prompts we evaluate.
model_dir = create_tiny_qwen3_dir(tmp_path, with_tokenizer=True, max_position_embeddings=8192)
try:
run_llm_ptq_command(
run_hf_ptq_command(
model=str(model_dir),
quant="fp8",
tasks="mmlu,lm_eval,simple_eval",
@@ -132,7 +132,7 @@ def test_offline_ptq(offline_ptq_dirs):
"--export_path",
str(offline_ptq_dirs["ptq_export"]),
],
"llm_ptq",
"hf_ptq",
)
# Verify the exported checkpoint exists and has the expected EAGLE keys
@@ -108,7 +108,7 @@ def test_unified_hf_export_and_check_safetensors(
env = os.environ.copy()
if expected_suffix.startswith("t5_tiny"):
env["CUDA_VISIBLE_DEVICES"] = "0"
run_example_command(cmd_parts, "llm_ptq", env=env)
run_example_command(cmd_parts, "hf_ptq", env=env)
# Now we expect a file named model.safetensors in output_dir
generated_file = output_dir / "model.safetensors"
@@ -17,7 +17,7 @@
This exercises the exact combination that silently regressed (nvbug 6295279 / 6295242):
a transposed-quantize MoE (``GptOssExperts``) with **static-block** NVFP4 weight quantizers
-- which is what ``examples/llm_ptq --cast_mxfp4_to_nvfp4`` produces via
-- which is what ``examples/hf_ptq --cast_mxfp4_to_nvfp4`` produces via
``force_weight_quantizers_static`` -- calibrated with a forward loop.
The regression (the unconditional ``weight_only_quantize`` from #1560 feeding the
@@ -43,9 +43,9 @@ from modelopt.torch.quantization.config import NVFP4_DEFAULT_CFG
from modelopt.torch.quantization.nn import NVFP4StaticQuantizer
# The cast helpers live next to the example script, not in the ``modelopt`` package.
_LLM_PTQ_DIR = Path(__file__).resolve().parents[4] / "examples" / "llm_ptq"
if str(_LLM_PTQ_DIR) not in sys.path:
sys.path.insert(0, str(_LLM_PTQ_DIR))
_HF_PTQ_DIR = Path(__file__).resolve().parents[4] / "examples" / "hf_ptq"
if str(_HF_PTQ_DIR) not in sys.path:
sys.path.insert(0, str(_HF_PTQ_DIR))
from cast_mxfp4_to_nvfp4 import apply_to_model, force_weight_quantizers_static
+1 -1
View File
@@ -41,7 +41,7 @@ bash tools/debugger/client.sh --timeout 1800 run "<command>"
```bash
# Run PTQ test
bash tools/debugger/client.sh run "bash llm_ptq/scripts/huggingface_example.sh"
bash tools/debugger/client.sh run "bash hf_ptq/scripts/huggingface_example.sh"
# Run pytest
bash tools/debugger/client.sh run "python -m pytest tests/gpu -k test_quantize"
+1 -1
View File
@@ -47,7 +47,7 @@ bash tools/debugger/client.sh handshake
bash tools/debugger/client.sh run "echo hello"
# Run a test script
bash tools/debugger/client.sh run "bash llm_ptq/scripts/huggingface_example.sh"
bash tools/debugger/client.sh run "bash hf_ptq/scripts/huggingface_example.sh"
# Run with a long timeout (default is 600s)
bash tools/debugger/client.sh --timeout 1800 run "python my_long_test.py"
+1 -1
View File
@@ -22,6 +22,6 @@ trap 'error_handler $0 $LINENO' ERR
###################################################################################################
python modules/Model-Optimizer/examples/llm_ptq/hf_ptq.py \
python modules/Model-Optimizer/examples/hf_ptq/hf_ptq.py \
--model ${HF_MODEL_CKPT} \
${@}
+1 -1
View File
@@ -64,7 +64,7 @@ fi
# --- Step 2: Run huggingface_example.sh ---
script_dir="$(dirname "$(readlink -f "$0")")"
HF_EXAMPLE="${script_dir}/../../modules/Model-Optimizer/examples/llm_ptq/scripts/huggingface_example.sh"
HF_EXAMPLE="${script_dir}/../../modules/Model-Optimizer/examples/hf_ptq/scripts/huggingface_example.sh"
echo "Running huggingface_example.sh --model $LOCAL_DIR --trust_remote_code ${PTQ_ARGS[*]}"
exec bash "$HF_EXAMPLE" --model "$LOCAL_DIR" --trust_remote_code "${PTQ_ARGS[@]}"
@@ -1,6 +1,6 @@
# HuggingFace PTQ via huggingface_example.sh
#
# Quantizes a HuggingFace model using examples/llm_ptq/scripts/huggingface_example.sh.
# Quantizes a HuggingFace model using examples/hf_ptq/scripts/huggingface_example.sh.
# Default: Qwen/Qwen3.5-9B with nvfp4_mlp_only on 8xH200.
#
# Usage (Slurm):
@@ -11,14 +11,14 @@
# export SLURM_HF_LOCAL=/home/scratch.<user>/hf-local
# export HF_TOKEN=<your-hf-token> # for gated models; auto-injected into all tasks
# cd tools/launcher
# uv run launch.py --yaml examples/llm_ptq/hf_ptq.yaml --yes
# uv run launch.py --yaml examples/hf_ptq/hf_ptq.yaml --yes
#
# Usage (local Docker):
# cd tools/launcher
# uv run launch.py --yaml examples/llm_ptq/hf_ptq.yaml hf_local=/mnt/hf-local --yes
# uv run launch.py --yaml examples/hf_ptq/hf_ptq.yaml hf_local=/mnt/hf-local --yes
#
# Override model/quant via CLI:
# uv run launch.py --yaml examples/llm_ptq/hf_ptq.yaml \
# uv run launch.py --yaml examples/hf_ptq/hf_ptq.yaml \
# pipeline.global_vars.hf_model=Qwen/Qwen3-8B \
# pipeline.task_0.args='[--model,<<global_vars.hf_local>>Qwen/Qwen3-8B,--,--quant,nvfp4]' \
# --yes