refactor(examples): consolidate vlm_ptq into llm_ptq (#1705)

### What does this PR do?

Type of change: refactor / deprecation (examples)

`examples/vlm_ptq` was effectively a thin wrapper over
`examples/llm_ptq`: its
`scripts/huggingface_example.sh` already sourced
`llm_ptq/scripts/parser.sh` and
called `llm_ptq/hf_ptq.py`, and all the actual VLM logic (vision-tower
exclusion,
`--calib_with_images`, Nemotron VL calibration, VILA loading, multimodal
export)
already lives under `llm_ptq`. The wrapper also referenced a
`requirements-vila.txt` that did not exist in the repo.

This PR makes `llm_ptq` the single source of truth for both LLM and VLM
PTQ and
deprecates `vlm_ptq`.

**`llm_ptq` (canonical):**
- Add `--vlm` and `--calib_with_images` flags to `scripts/parser.sh` and
`scripts/huggingface_example.sh`. `--vlm` bootstraps VILA dependencies
and runs
the TensorRT-LLM multimodal quickstart as the deploy smoke test (instead
of the
  text-only `run_tensorrt_llm.py`).
- Add `examples/llm_ptq/requirements-vila.txt` (fixes the previously
broken
  reference).
- Document the VLM support matrix and the `--vlm` workflow in
`README.md`.

**`vlm_ptq` (deprecated):**
- Replace `scripts/huggingface_example.sh` with a shim that prints a
deprecation
  warning and forwards to the `llm_ptq` script with `--vlm`.
- Convert `README.md` into a redirect/migration notice.
- Repoint root `README.md` VLM links and add a `CHANGELOG.rst`
deprecation entry.

### Usage

```bash
cd examples/llm_ptq
# VLM PTQ (was: examples/vlm_ptq/scripts/huggingface_example.sh)
scripts/huggingface_example.sh --model <hf_model> --quant fp8 --vlm

# VLM image-text calibration
scripts/huggingface_example.sh --model <hf_model> --quant nvfp4 --vlm --calib_with_images --trust_remote_code
```

### Testing

- `bash -n` syntax check on the modified `parser.sh`, `llm_ptq` script,
and the
  `vlm_ptq` shim.
- `pre-commit run --files <changed files>` passes.
- The existing VLM example test
(`tests/examples/vlm_ptq/test_qwen_vl.py` via
  `run_vlm_ptq_command`) still exercises the path end-to-end through the
  deprecation shim, which forwards to the consolidated `llm_ptq` script.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (old `vlm_ptq` entry point
still works via a forwarding shim)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (existing VLM test still
covers the consolidated path)
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

Follow-up (later release): remove the `examples/vlm_ptq` directory and
its CI
matrix entry once external references have migrated.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added VLM quantization support to `examples/llm_ptq` via a `--vlm`
flag.
* Enabled image-text pair calibration with `--calib_with_images`,
including VLM multimodal smoke-test coverage.

* **Deprecations**
* `examples/vlm_ptq` is deprecated; it now forwards to the
`examples/llm_ptq --vlm` flow with a warning.
* VILA/NVILA VLM support was removed from `examples/llm_ptq` due to a
model dependency compatibility conflict.

* **Documentation**
* Updated READMEs and the model support matrix with VLM quantization
behavior and export limitations.

* **Tests / CI**
* Updated VLM PTQ tests and CI workflow matrices to stop running the
deprecated `vlm_ptq` example.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Zhiyu
2026-06-16 18:12:34 -07:00
committed by GitHub
co-authored by Claude Opus 4.8
parent 977d34dc3c
commit 6c32c37671
12 changed files with 231 additions and 353 deletions
+2 -2
View File
@@ -55,7 +55,7 @@ jobs:
strategy:
fail-fast: false
matrix:
example: [llm_ptq, vlm_ptq]
example: [llm_ptq]
uses: ./.github/workflows/_example_tests_runner.yml
secrets: inherit
with:
@@ -69,7 +69,7 @@ jobs:
strategy:
fail-fast: false
matrix:
example: [llm_autodeploy, llm_eval, llm_ptq, vlm_ptq]
example: [llm_autodeploy, llm_eval, llm_ptq]
uses: ./.github/workflows/_example_tests_runner.yml
secrets: inherit
with: