mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
refactor(examples): consolidate vlm_ptq into llm_ptq (#1705)
### What does this PR do? Type of change: refactor / deprecation (examples) `examples/vlm_ptq` was effectively a thin wrapper over `examples/llm_ptq`: its `scripts/huggingface_example.sh` already sourced `llm_ptq/scripts/parser.sh` and called `llm_ptq/hf_ptq.py`, and all the actual VLM logic (vision-tower exclusion, `--calib_with_images`, Nemotron VL calibration, VILA loading, multimodal export) already lives under `llm_ptq`. The wrapper also referenced a `requirements-vila.txt` that did not exist in the repo. This PR makes `llm_ptq` the single source of truth for both LLM and VLM PTQ and deprecates `vlm_ptq`. **`llm_ptq` (canonical):** - Add `--vlm` and `--calib_with_images` flags to `scripts/parser.sh` and `scripts/huggingface_example.sh`. `--vlm` bootstraps VILA dependencies and runs the TensorRT-LLM multimodal quickstart as the deploy smoke test (instead of the text-only `run_tensorrt_llm.py`). - Add `examples/llm_ptq/requirements-vila.txt` (fixes the previously broken reference). - Document the VLM support matrix and the `--vlm` workflow in `README.md`. **`vlm_ptq` (deprecated):** - Replace `scripts/huggingface_example.sh` with a shim that prints a deprecation warning and forwards to the `llm_ptq` script with `--vlm`. - Convert `README.md` into a redirect/migration notice. - Repoint root `README.md` VLM links and add a `CHANGELOG.rst` deprecation entry. ### Usage ```bash cd examples/llm_ptq # VLM PTQ (was: examples/vlm_ptq/scripts/huggingface_example.sh) scripts/huggingface_example.sh --model <hf_model> --quant fp8 --vlm # VLM image-text calibration scripts/huggingface_example.sh --model <hf_model> --quant nvfp4 --vlm --calib_with_images --trust_remote_code ``` ### Testing - `bash -n` syntax check on the modified `parser.sh`, `llm_ptq` script, and the `vlm_ptq` shim. - `pre-commit run --files <changed files>` passes. - The existing VLM example test (`tests/examples/vlm_ptq/test_qwen_vl.py` via `run_vlm_ptq_command`) still exercises the path end-to-end through the deprecation shim, which forwards to the consolidated `llm_ptq` script. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (old `vlm_ptq` entry point still works via a forwarding shim) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (existing VLM test still covers the consolidated path) - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Follow-up (later release): remove the `examples/vlm_ptq` directory and its CI matrix entry once external references have migrated. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added VLM quantization support to `examples/llm_ptq` via a `--vlm` flag. * Enabled image-text pair calibration with `--calib_with_images`, including VLM multimodal smoke-test coverage. * **Deprecations** * `examples/vlm_ptq` is deprecated; it now forwards to the `examples/llm_ptq --vlm` flow with a warning. * VILA/NVILA VLM support was removed from `examples/llm_ptq` due to a model dependency compatibility conflict. * **Documentation** * Updated READMEs and the model support matrix with VLM quantization behavior and export limitations. * **Tests / CI** * Updated VLM PTQ tests and CI workflow matrices to stop running the deprecated `vlm_ptq` example. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -55,7 +55,7 @@ jobs:
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
example: [llm_ptq, vlm_ptq]
|
||||
example: [llm_ptq]
|
||||
uses: ./.github/workflows/_example_tests_runner.yml
|
||||
secrets: inherit
|
||||
with:
|
||||
@@ -69,7 +69,7 @@ jobs:
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
example: [llm_autodeploy, llm_eval, llm_ptq, vlm_ptq]
|
||||
example: [llm_autodeploy, llm_eval, llm_ptq]
|
||||
uses: ./.github/workflows/_example_tests_runner.yml
|
||||
secrets: inherit
|
||||
with:
|
||||
|
||||
Reference in New Issue
Block a user