## What does this PR do? **Type of change:** refactor / deprecation (examples) Follow-up to #1705 (which consolidated `examples/vlm_ptq` into `examples/llm_ptq`). Since that example now covers Hugging Face **LLM and VLM** PTQ, the `llm_ptq` name is a misnomer. This renames the directory to `examples/hf_ptq` and leaves a relative symlink `examples/llm_ptq → hf_ptq` so existing paths/commands keep working during a deprecation window. Requested by @kevalmorabia97 on #1705 (with the symlink-for-back-compat approach), targeted for the **same 0.46 release** as the consolidation. ### Changes - `git mv examples/llm_ptq → examples/hf_ptq` and `tests/examples/llm_ptq → tests/examples/hf_ptq` (the CI runner maps the matrix name to both `examples/<name>` and `tests/examples/<name>`). - Add a tracked back-compat symlink `examples/llm_ptq → hf_ptq`. - Update CI matrices and all repo **path references** (docs, READMEs, agent skills, launcher/debugger tools, tests) from `llm_ptq` to `hf_ptq`. - Keep Python identifiers / test-util module names (`run_llm_ptq_command`, `llm_ptq_utils`) — they name the LLM-PTQ task, not the directory. - Preserve the CODEOWNERS team slug (`modelopt-examples-llm_ptq-codeowners`) and historical CHANGELOG entries; add a CHANGELOG deprecation note. ### Back-compat caveats (inherent to git directory symlinks) - ✅ Linux/macOS CLI usage and Python `cwd`/pytest resolution work through the symlink. - ⚠️ Windows git checkouts don't materialize symlinks by default (low impact — this example is Linux-only in practice). - ⚠️ GitHub web doesn't follow directory symlinks, so legacy external deep-links to `examples/llm_ptq/...` won't navigate in. All **internal** references are repointed to `hf_ptq`, so the symlink is only for legacy external/CLI use. ### Usage (unchanged via symlink) ```bash # New canonical path cd examples/hf_ptq scripts/huggingface_example.sh --model <hf_model> --quant fp8 # Old path still works (forwards via symlink) cd examples/llm_ptq && scripts/huggingface_example.sh --model <hf_model> --quant fp8 ``` ### Testing - `bash -n` on moved/edited shell scripts (new path + via symlink). - `py_compile` on moved/edited Python; test re-export shim repointed to `examples/hf_ptq/example_utils`. - Verified git tracks `examples/llm_ptq` as a single symlink (mode 120000), not a duplicated tree (no pre-commit / pytest double-processing). - `pre-commit run` on all changed files passes. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (relative symlink keeps `examples/llm_ptq` paths valid; see caveats above) - Did you write any new necessary tests?: N/A (pure rename; existing tests moved with the dir) - Did you update Changelog?: ✅ ### Additional Information Follow-up (later release): remove the `examples/llm_ptq` symlink once external references have migrated. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * PTQ guidance now directs to the unified Hugging Face PTQ flow, including VLM quantization via the shared `--vlm` entry point. * **Documentation** * Updated README and guide links, references, and command snippets to use `hf_ptq` (replacing `llm_ptq`). * Deprecated and consolidated `vlm_ptq` into `hf_ptq`; removed VILA/NVILA coverage from the Hugging Face PTQ examples. * **Bug Fixes** * Improved detection and routing so local/manual setup uses the correct PTQ source. * **Tests / Chores** * CI and example tests updated to run the `hf_ptq` variants. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Quantization Aware Training and Distillation with LLaMA-Factory
This directory provides integration between LLaMA-Factory and ModelOpt for Quantization-Aware Training (QAT) and Quantization-Aware Distillation (QAD). This enables efficient training of large language models with quantization techniques while maintaining model quality. This README only covers QAT/QAD training with LLaMA-Factory. For more information on setting up other datasets, training different models, or customizing training configurations, please refer to LLaMA-Factory.
Quick Start
Basic QAT/QAD Training
./launch_llamafactory.sh llama_config.yaml
NOTE: The
launch_llamafactory.shscript automatically installs LLaMA Factory if it's not already present in your environment.
By default, the script uses fsdp2.yaml for distributed training.
Use Custom FSDP Arguments:
Pass additional FSDP parameters using the FSDP_ARGS environment variable:
FSDP_ARGS="--fsdp_transformer_layer_cls_to_wrap Qwen3DecoderLayer" \
./launch_llamafactory.sh llama_config.yaml
NOTE: The default
fsdp2.yamlusesQwen3DecoderLayeras the transformer layer class. If your model uses a different layer class (e.g.,LlamaDecoderLayerfor Llama models), pass--fsdp_transformer_layer_cls_to_wrap <your_layer_class>to thelaunch_llamafactory.shscript or provide a custom FSDP configuration file. See Supported Backends for details.
Custom Config File:
Specify your own FSDP configuration:
./launch_llamafactory.sh llama_config.yaml \
--accelerate_config /path/to/custom_fsdp.yaml
Training using CLI
For QAT/QAD training using llamafactory_cli, run
./launch_llamafactory.sh train llama_config.yaml
Configuration Guide
This section explains how to configure quantization parameters in your training setup.
YAML Configuration Structure
Your configuration file should follow same structure as sft example from llama3_full_sft.yaml with addition of modelopt configuration:
# Model Configuration
model:
model_name_or_path: /path/to/your/model
trust_remote_code: true
### method
stage: sft # Supervised Fine-Tuning
do_train: true
finetuning_type: full # Full parameter fine-tuning
### dataset
dataset: your_dataset
cutoff_len: 4096 # Maximum sequence length
max_samples: 100000 # Maximum training samples
### train
per_device_train_batch_size: 1
gradient_accumulation_steps: 1
learning_rate: 1.0e-5
num_train_epochs: 1
bf16: true # Use bfloat16 precision
val_size: 0.1 # Validation split ratio
### ModelOpt Configuration
modelopt:
recipe: general/ptq/nvfp4_default-kv_fp8 # Quantization recipe (built-in or custom path)
calib_size: 1024 # Calibration dataset size
compress: false # Enable weight compression
distill: false # Modify distill to true for QAD
teacher_model: /path/to/teacher/model # For QAD (optional)
NOTE:
compress: trueenables weight compression and will by default use ddp.yaml. NOTE: When training without cli, avoid using deepspeed option in the YAML configuration file.
Deployment
The final QAT/QAD model after training is similar in architecture to that of PTQ model. It simply has updated weights as compared to the PTQ model. It can be deployed to TensorRT-LLM (TRTLLM) or to TensorRT just like a regular ModelOpt PTQ model if the quantization format is supported for deployment.
To run QAT/QAD model with TRTLLM, run:
cd ../../hf_ptq
./scripts/huggingface_example.sh --model <path-to-QAT/QAD-model> --quant nvfp4
See more details on deployment of quantized model here.