Files
Model-Optimizer/examples/llm_qat/llama_factory
ZhiyuandClaude Opus 4.8 f335459dc0 refactor(examples): rename llm_ptq → hf_ptq (symlink for back-compat) (#1759)
## What does this PR do?

**Type of change:** refactor / deprecation (examples)

Follow-up to #1705 (which consolidated `examples/vlm_ptq` into
`examples/llm_ptq`). Since that example now covers Hugging Face **LLM
and VLM** PTQ, the `llm_ptq` name is a misnomer. This renames the
directory to `examples/hf_ptq` and leaves a relative symlink
`examples/llm_ptq → hf_ptq` so existing paths/commands keep working
during a deprecation window.

Requested by @kevalmorabia97 on #1705 (with the symlink-for-back-compat
approach), targeted for the **same 0.46 release** as the consolidation.

### Changes
- `git mv examples/llm_ptq → examples/hf_ptq` and
`tests/examples/llm_ptq → tests/examples/hf_ptq` (the CI runner maps the
matrix name to both `examples/<name>` and `tests/examples/<name>`).
- Add a tracked back-compat symlink `examples/llm_ptq → hf_ptq`.
- Update CI matrices and all repo **path references** (docs, READMEs,
agent skills, launcher/debugger tools, tests) from `llm_ptq` to
`hf_ptq`.
- Keep Python identifiers / test-util module names
(`run_llm_ptq_command`, `llm_ptq_utils`) — they name the LLM-PTQ task,
not the directory.
- Preserve the CODEOWNERS team slug
(`modelopt-examples-llm_ptq-codeowners`) and historical CHANGELOG
entries; add a CHANGELOG deprecation note.

### Back-compat caveats (inherent to git directory symlinks)
- ✅ Linux/macOS CLI usage and Python `cwd`/pytest resolution work
through the symlink.
- ⚠️ Windows git checkouts don't materialize symlinks by default (low
impact — this example is Linux-only in practice).
- ⚠️ GitHub web doesn't follow directory symlinks, so legacy external
deep-links to `examples/llm_ptq/...` won't navigate in. All **internal**
references are repointed to `hf_ptq`, so the symlink is only for legacy
external/CLI use.

### Usage (unchanged via symlink)
```bash
# New canonical path
cd examples/hf_ptq
scripts/huggingface_example.sh --model <hf_model> --quant fp8

# Old path still works (forwards via symlink)
cd examples/llm_ptq && scripts/huggingface_example.sh --model <hf_model> --quant fp8
```

### Testing
- `bash -n` on moved/edited shell scripts (new path + via symlink).
- `py_compile` on moved/edited Python; test re-export shim repointed to
`examples/hf_ptq/example_utils`.
- Verified git tracks `examples/llm_ptq` as a single symlink (mode
120000), not a duplicated tree (no pre-commit / pytest
double-processing).
- `pre-commit run` on all changed files passes.

### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (relative symlink keeps
`examples/llm_ptq` paths valid; see caveats above)
- Did you write any new necessary tests?: N/A (pure rename; existing
tests moved with the dir)
- Did you update Changelog?: ✅

### Additional Information
Follow-up (later release): remove the `examples/llm_ptq` symlink once
external references have migrated.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* PTQ guidance now directs to the unified Hugging Face PTQ flow,
including VLM quantization via the shared `--vlm` entry point.
* **Documentation**
* Updated README and guide links, references, and command snippets to
use `hf_ptq` (replacing `llm_ptq`).
* Deprecated and consolidated `vlm_ptq` into `hf_ptq`; removed
VILA/NVILA coverage from the Hugging Face PTQ examples.
* **Bug Fixes**
* Improved detection and routing so local/manual setup uses the correct
PTQ source.
* **Tests / Chores**
  * CI and example tests updated to run the `hf_ptq` variants.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-27 08:48:48 +00:00
..
…

Quantization Aware Training and Distillation with LLaMA-Factory

This directory provides integration between LLaMA-Factory and ModelOpt for Quantization-Aware Training (QAT) and Quantization-Aware Distillation (QAD). This enables efficient training of large language models with quantization techniques while maintaining model quality. This README only covers QAT/QAD training with LLaMA-Factory. For more information on setting up other datasets, training different models, or customizing training configurations, please refer to LLaMA-Factory.

Quick Start

Basic QAT/QAD Training

./launch_llamafactory.sh llama_config.yaml

NOTE: The launch_llamafactory.sh script automatically installs LLaMA Factory if it's not already present in your environment.

By default, the script uses fsdp2.yaml for distributed training.

Use Custom FSDP Arguments:

Pass additional FSDP parameters using the FSDP_ARGS environment variable:

FSDP_ARGS="--fsdp_transformer_layer_cls_to_wrap Qwen3DecoderLayer" \
./launch_llamafactory.sh llama_config.yaml

NOTE: The default fsdp2.yaml uses Qwen3DecoderLayer as the transformer layer class. If your model uses a different layer class (e.g., LlamaDecoderLayer for Llama models), pass --fsdp_transformer_layer_cls_to_wrap <your_layer_class> to the launch_llamafactory.sh script or provide a custom FSDP configuration file. See Supported Backends for details.

Custom Config File:

Specify your own FSDP configuration:

./launch_llamafactory.sh llama_config.yaml \
  --accelerate_config /path/to/custom_fsdp.yaml

Training using CLI

For QAT/QAD training using llamafactory_cli, run

./launch_llamafactory.sh train llama_config.yaml

Configuration Guide

This section explains how to configure quantization parameters in your training setup.

YAML Configuration Structure

Your configuration file should follow same structure as sft example from llama3_full_sft.yaml with addition of modelopt configuration:

# Model Configuration
model:
  model_name_or_path: /path/to/your/model
  trust_remote_code: true

### method
stage: sft                    # Supervised Fine-Tuning
do_train: true
finetuning_type: full         # Full parameter fine-tuning

### dataset
dataset: your_dataset
cutoff_len: 4096             # Maximum sequence length
max_samples: 100000          # Maximum training samples

### train
per_device_train_batch_size: 1
gradient_accumulation_steps: 1
learning_rate: 1.0e-5
num_train_epochs: 1
bf16: true                   # Use bfloat16 precision

val_size: 0.1               # Validation split ratio

### ModelOpt Configuration
modelopt:
  recipe: general/ptq/nvfp4_default-kv_fp8  # Quantization recipe (built-in or custom path)
  calib_size: 1024                          # Calibration dataset size
  compress: false                           # Enable weight compression
  distill: false                            # Modify distill to true for QAD
  teacher_model: /path/to/teacher/model     # For QAD (optional)

NOTE: compress: true enables weight compression and will by default use ddp.yaml. NOTE: When training without cli, avoid using deepspeed option in the YAML configuration file.

Deployment

The final QAT/QAD model after training is similar in architecture to that of PTQ model. It simply has updated weights as compared to the PTQ model. It can be deployed to TensorRT-LLM (TRTLLM) or to TensorRT just like a regular ModelOpt PTQ model if the quantization format is supported for deployment.

To run QAT/QAD model with TRTLLM, run:

cd ../../hf_ptq

./scripts/huggingface_example.sh --model <path-to-QAT/QAD-model> --quant nvfp4

See more details on deployment of quantized model here.