Remove deprecated examples/llm_autodeploy (#1797)

Remove the AutoQuant + TensorRT-LLM AutoDeploy example, deprecated in
0.45, after the migration period. Record the removal under the 0.46
Backward Breaking Changes section. Users should use TensorRT-LLM's
AutoDeploy directly together with ModelOpt PTQ in examples/llm_ptq.

### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated AutoDeploy guidance to a workflow: quantize with ModelOpt PTQ
(via `llm_ptq`) to produce a unified Hugging Face checkpoint, then
deploy with TensorRT-LLM AutoDeploy.
* Revised Hopper notes to recommend using FP8 (and refreshed related
optimization guidance).
* **Breaking Changes**
* Removed the deprecated `examples/llm_autodeploy` example and
documented the new recommended approach.
* **Chores**
* Dropped obsolete example docs, scripts, and coverage; adjusted example
test workflow to exclude `llm_autodeploy`; updated ownership mapping for
the removed example path.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Frida Hou
2026-06-27 09:46:50 +05:30
committed by GitHub
co-authored by Claude Opus 4.8
parent 4b04e732ed
commit d5962c4f3b
11 changed files with 20 additions and 788 deletions
+2 -2
View File
@@ -48,7 +48,7 @@ jobs:
pip_install_extras: "[hf,dev-test]"
runner: ${{ startsWith(github.ref, 'refs/heads/pull-request/') && 'linux-amd64-gpu-rtxpro6000-latest-1' || 'linux-amd64-gpu-rtxpro6000-latest-2' }}
##### TensorRT-LLM Example Tests (pr/non-pr split: non-pr runs extra autodeploy+eval examples) #####
##### TensorRT-LLM Example Tests (pr/non-pr split: non-pr runs extra eval examples) #####
trtllm-pr:
needs: [pr-gate]
if: startsWith(github.ref, 'refs/heads/pull-request/') && needs.pr-gate.outputs.any_changed == 'true'
@@ -69,7 +69,7 @@ jobs:
strategy:
fail-fast: false
matrix:
example: [llm_autodeploy, llm_eval, llm_ptq]
example: [llm_eval, llm_ptq]
uses: ./.github/workflows/_example_tests_runner.yml
secrets: inherit
with: