Add tests/gpu_vllm (#1517)

### What does this PR do?

Type of change: new tests

This PR adds unit tests for vLLM fakequant, specifically testing code in
`modelopt/torch/quantization/plugins/vllm.py`


### Testing

```
pytest tests/gpu_vllm/torch/quantization/test_vllm_dynamic_modules.py -sv
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ✅ 

### Additional Information


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Added comprehensive GPU vLLM test suite with end-to-end quantization
checks and fixtures for TinyLlama, TinyQwen3-MoE, and DeepSeek V3;
includes helpers to build tiny DeepSeek V3 models.

* **Chores**
* Updated GPU CI to use explicit container image references, added a
GPU-focused test session, and adjusted test-run setup for vLLM.

* **Documentation**
  * Documented new GPU test directory in contributing guide.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1517?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
This commit is contained in:
kinjalpatel27
2026-05-29 20:47:23 +00:00
committed by GitHub
co-authored by Keval Morabia
parent ed0a4b175d
commit 7aa0c95646
8 changed files with 369 additions and 19 deletions
@@ -34,9 +34,6 @@ jobs:
timeout-minutes: ${{ inputs.timeout_minutes }}
container:
image: ${{ inputs.docker_image }}
credentials:
username: $oauthtoken
password: ${{ secrets.NGC_API_KEY }}
options: --shm-size=2gb # TRT-LLM tests on 2-GPU runner needs more shared memory
env:
PIP_CONSTRAINT: "" # Disable pip constraint for upgrading packages
+15 -8
View File
@@ -29,6 +29,7 @@ jobs:
tests/gpu/**
tests/gpu_megatron/**
tests/gpu_trtllm/**
tests/gpu_vllm/**
gpu-tests:
needs: [pr-gate]
@@ -39,25 +40,30 @@ jobs:
include:
- example: gpu
timeout: 75
container_image: pytorch:26.04-py3
container_image: nvcr.io/nvidia/pytorch:26.04-py3
- example: gpu_megatron
timeout: 45
container_image: nemo:26.04
container_image: nvcr.io/nvidia/nemo:26.04
- example: gpu_trtllm
timeout: 30
container_image: tensorrt-llm/release:1.3.0rc16
container_image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc16
- example: gpu_vllm
timeout: 30
container_image: docker.io/vllm/vllm-openai:v0.20.0
runs-on: ${{ startsWith(github.ref, 'refs/heads/pull-request/') && 'linux-amd64-gpu-rtxpro6000-latest-1' || 'linux-amd64-gpu-rtxpro6000-latest-2' }}
timeout-minutes: ${{ matrix.timeout }}
container:
image: nvcr.io/nvidia/${{ matrix.container_image }}
credentials:
username: $oauthtoken
password: ${{ secrets.NGC_API_KEY }}
image: ${{ matrix.container_image }}
env:
GIT_DEPTH: 1000 # For correct version
PIP_CONSTRAINT: "" # Disable pip constraint for upgrading packages
HF_TOKEN: ${{ secrets.HF_TOKEN }}
steps:
- name: Install git
# The vllm container ships without git; needed for a real checkout (correct
# setuptools-scm version) and for the Codecov upload below.
if: matrix.example == 'gpu_vllm'
run: apt-get update && apt-get install -y git
- uses: actions/checkout@v6
- uses: nv-gha-runners/setup-proxy-cache@main
- name: Setup environment variables
@@ -68,7 +74,8 @@ jobs:
COVERAGE_PROCESS_START: ${{ github.workspace }}/pyproject.toml
COVERAGE_FILE: ${{ github.workspace }}/.coverage
run: |
python -m pip install nox && nox -s ${{ matrix.example }}
# Use `python3` (the vllm image has no `python` on PATH)
python3 -m pip install nox && nox -s ${{ matrix.example }}
- name: Upload GPU coverage to Codecov
uses: codecov/codecov-action@v5
with: