mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: documentation Fixes [NVBug 6550792](https://nvbugspro.nvidia.com/bug/6550792) / OMNIML-5693. The **Unified HF Checkpoint Deployment Model Support Matrix** listed 9 model families and **no VLMs**, while `tests/examples/hf_ptq/test_deploy.py` declares deployment cases for ~80 checkpoints across TRT-LLM, vLLM, and SGLang — including `Qwen2.5-VL`, `Qwen3-VL-235B`, and `Nemotron-3-Nano-Omni`. QA (the filer) could not use the doc to scope testing, and users could not tell what is actually covered. Filing also surfaced that the matrix lived in **three places that had drifted apart**: only the `.rst` listed Qwen3-VL, only the README listed Qwen3.5 MoE, and the skill reference had neither. #### Changes 1. **Rebuilt the matrix in `docs/source/deployment/3_unified_hf.rst`** from `test_deploy.py`, split into language models, vision-language/multimodal, speculative decoding drafters, and diffusion. 2. **Stated plainly what the matrix is and is not.** Review established that the original "CI-validated" framing claimed more than the suite substantiates, so a *What this matrix is based on* section now leads with two limits: - The suite is marked `release` and collects only under `--run-release`, which **no workflow passes** — these are declared cases, not PR-gated coverage. - Each case is a **load-and-generate smoke check on the text path**: no accuracy, no image/audio input, no diffusion output, no verification that speculative decoding engages. The legend follows from that: ✅ = declared in the suite, ⚠ = expected to work but not a suite entry (or an entry that does not exercise the feature the row names), `-` = not in the suite. Sections that would otherwise over-read carry their own qualifiers — VLM rows are labelled text-only smoke coverage, and Medusa and Wan 2.2 are ⚠ with the reason stated. 3. **Removed the two duplicate copies**, replacing them with links, so there is one table to maintain. 4. **Fixed stale prose**: the deployment tabs still claimed FP8-only support on vLLM v0.6.5 and a source build of SGLang main from Jan 2025, both contradicting the version table above them. The TRT-LLM floor moves to v1.2.0, qualified as the oldest version stated rather than the oldest that works. 5. **Dropped the Phi series** from the deployment matrix, following #2115 (NVBug 6563509) and confirmation that Phi-4 is being deprecated. ### Usage N/A — documentation only. ### Testing - `docutils` parse of the modified `.rst`: no warnings or errors from the new content; all 5 tables parse with every cell in the correct column. - Cell contents cross-checked against `test_deploy.py` by AST-parsing the `ModelDeployerList(...)` calls rather than by eye; the scope caveats were each verified against `tests/_test_utils/deploy_utils.py`. - `pre-commit run --files …` passes; `build-docs` green. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — documentation only - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information **Two known follow-ups, neither in scope here:** 1. **Nothing enforces that the doc matrix tracks `test_deploy.py`.** Consolidating to one copy removes the three-way drift but not the doc-vs-test drift; a generator plus a CI check would close it. 2. **The release deployment suite does not run in CI.** Wiring it into per-backend release CI is what would let ✅ mean "verified to pass" rather than "declared". That needs GPU capacity across three backends and should be tracked on its own. **For the filer (@Kenny Kang):** the ✅ cells are the scope the release deploy suite declares, and `test_deploy.py` carries the checkpoint, TP size, and minimum SM version per entry — but please read the legend first, since those cases are not currently executed by CI. --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>