mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Add tests/gpu_vllm (#1517)
### What does this PR do? Type of change: new tests This PR adds unit tests for vLLM fakequant, specifically testing code in `modelopt/torch/quantization/plugins/vllm.py` ### Testing ``` pytest tests/gpu_vllm/torch/quantization/test_vllm_dynamic_modules.py -sv ``` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added comprehensive GPU vLLM test suite with end-to-end quantization checks and fixtures for TinyLlama, TinyQwen3-MoE, and DeepSeek V3; includes helpers to build tiny DeepSeek V3 models. * **Chores** * Updated GPU CI to use explicit container image references, added a GPU-focused test session, and adjusted test-run setup for vLLM. * **Documentation** * Documented new GPU test directory in contributing guide. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1517?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
This commit is contained in:
co-authored by
Keval Morabia
parent
ed0a4b175d
commit
7aa0c95646
@@ -34,9 +34,6 @@ jobs:
|
||||
timeout-minutes: ${{ inputs.timeout_minutes }}
|
||||
container:
|
||||
image: ${{ inputs.docker_image }}
|
||||
credentials:
|
||||
username: $oauthtoken
|
||||
password: ${{ secrets.NGC_API_KEY }}
|
||||
options: --shm-size=2gb # TRT-LLM tests on 2-GPU runner needs more shared memory
|
||||
env:
|
||||
PIP_CONSTRAINT: "" # Disable pip constraint for upgrading packages
|
||||
|
||||
@@ -29,6 +29,7 @@ jobs:
|
||||
tests/gpu/**
|
||||
tests/gpu_megatron/**
|
||||
tests/gpu_trtllm/**
|
||||
tests/gpu_vllm/**
|
||||
|
||||
gpu-tests:
|
||||
needs: [pr-gate]
|
||||
@@ -39,25 +40,30 @@ jobs:
|
||||
include:
|
||||
- example: gpu
|
||||
timeout: 75
|
||||
container_image: pytorch:26.04-py3
|
||||
container_image: nvcr.io/nvidia/pytorch:26.04-py3
|
||||
- example: gpu_megatron
|
||||
timeout: 45
|
||||
container_image: nemo:26.04
|
||||
container_image: nvcr.io/nvidia/nemo:26.04
|
||||
- example: gpu_trtllm
|
||||
timeout: 30
|
||||
container_image: tensorrt-llm/release:1.3.0rc16
|
||||
container_image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc16
|
||||
- example: gpu_vllm
|
||||
timeout: 30
|
||||
container_image: docker.io/vllm/vllm-openai:v0.20.0
|
||||
runs-on: ${{ startsWith(github.ref, 'refs/heads/pull-request/') && 'linux-amd64-gpu-rtxpro6000-latest-1' || 'linux-amd64-gpu-rtxpro6000-latest-2' }}
|
||||
timeout-minutes: ${{ matrix.timeout }}
|
||||
container:
|
||||
image: nvcr.io/nvidia/${{ matrix.container_image }}
|
||||
credentials:
|
||||
username: $oauthtoken
|
||||
password: ${{ secrets.NGC_API_KEY }}
|
||||
image: ${{ matrix.container_image }}
|
||||
env:
|
||||
GIT_DEPTH: 1000 # For correct version
|
||||
PIP_CONSTRAINT: "" # Disable pip constraint for upgrading packages
|
||||
HF_TOKEN: ${{ secrets.HF_TOKEN }}
|
||||
steps:
|
||||
- name: Install git
|
||||
# The vllm container ships without git; needed for a real checkout (correct
|
||||
# setuptools-scm version) and for the Codecov upload below.
|
||||
if: matrix.example == 'gpu_vllm'
|
||||
run: apt-get update && apt-get install -y git
|
||||
- uses: actions/checkout@v6
|
||||
- uses: nv-gha-runners/setup-proxy-cache@main
|
||||
- name: Setup environment variables
|
||||
@@ -68,7 +74,8 @@ jobs:
|
||||
COVERAGE_PROCESS_START: ${{ github.workspace }}/pyproject.toml
|
||||
COVERAGE_FILE: ${{ github.workspace }}/.coverage
|
||||
run: |
|
||||
python -m pip install nox && nox -s ${{ matrix.example }}
|
||||
# Use `python3` (the vllm image has no `python` on PATH)
|
||||
python3 -m pip install nox && nox -s ${{ matrix.example }}
|
||||
- name: Upload GPU coverage to Codecov
|
||||
uses: codecov/codecov-action@v5
|
||||
with:
|
||||
|
||||
Reference in New Issue
Block a user