mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: Bug fix (CI) + test coverage **Fixes `onnx (torch_onnx)` and `onnx (diffusers)`**, which have failed on every branch since `torch 2.14.0` was published to PyPI today (2026-09-02 13:42 UTC), and **adds torch 2.14 to the unit test matrix as the new default** so the next torch release is caught there rather than in an example job. ### Root cause Every test in those two jobs failed with: ``` RuntimeError: CUDNN_BACKEND_TENSOR_DESCRIPTOR cudnnFinalize failed ptrDesc->finalize() cudnn_status: CUDNN_STATUS_SUBLIBRARY_LOADING_FAILED ``` `nvcr.io/nvidia/tensorrt:26.05-py3` ships cuDNN **9.22** and has no preinstalled torch, so pip resolved the newest one — and torch 2.14 pins `nvidia-cudnn-cu13==9.24.0.43`. Loading 9.24 sublibraries against the image's 9.22 `libcudnn.so.9` is exactly what that status reports. | | last good run (08:55) | first failing run (13:34) | |---|---|---| | `torch` | 2.13.0 | **2.14.0** | | `nvidia-cudnn-cu13` | 9.20.0.48 | **9.24.0.43** | | image cuDNN | 9.22.0.52 | 9.22.0.52 | ### Why only these two jobs - The **nemo** and **pytorch** images have a preinstalled torch that already satisfies `torch>=2.8`, so pip never resolves a new one — confirmed from the megatron job log, where torch does not appear in `Successfully installed`. - **`tensorrt:26.05-py3` has no preinstalled torch**, so pip takes the newest from PyPI. - **`onnx (torch_trt)`** shares that image but passes throughout, because `torch-tensorrt<2.13` already holds torch below 2.14. ### The changes 1. **Constrain torch only where the incompatibility is.** `PIP_CONSTRAINT=torch<2.14` in the example runner, applied when the job's image is a `tensorrt` one. It also covers the `examples/*/requirements.txt` loop in the same shell, which matters because `nemo_automodel` pulls torch in too. Not pinned in `pyproject.toml`: torch 2.14 is fine anywhere its own bundled cuDNN is the one loaded, so that would constrain users to work around one pinned image. 2. **Test torch 2.14.** `torch_214` added to `TORCH_VERSIONS` (`torchvision~=0.29.0`) and promoted to the unit-test default across the supported Python versions, with 2.13 demoted to the back-compat row. `release.yml`'s basic unit test moves to the same default (it was still on 2.12). Nothing exercised 2.14 before — which is why a torch release reached us through an example job instead of a unit test. ### Testing - `actionlint` and YAML/TOML parse clean; pre-commit clean. - Verified by this PR's own jobs: `onnx (torch_onnx)` and `onnx (diffusers)` reproduce the failure on `main` right now, and the new `unit-3.12(torch_214, tf_latest)` job is the first run of ModelOpt against torch 2.14. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — CI-only; no source or package metadata change - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new dependency - Did you write any new necessary tests?: ✅ — torch 2.14 added to the unit test matrix - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — internal CI, not user-facing - Did you get Claude approval on this PR?: ❌ — not yet requested --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
108 lines
4.8 KiB
YAML
108 lines
4.8 KiB
YAML
# Reusable workflow for running example tests
|
|
name: Example Tests Runner
|
|
|
|
on:
|
|
workflow_call:
|
|
inputs:
|
|
docker_image:
|
|
description: "Docker image to use for tests"
|
|
required: true
|
|
type: string
|
|
example:
|
|
description: "Example name to test (e.g. 'hf_ptq')"
|
|
required: true
|
|
type: string
|
|
timeout_minutes:
|
|
description: "Timeout in minutes for the job"
|
|
required: false
|
|
type: number
|
|
default: 60
|
|
pip_install_extras:
|
|
description: "Pip install extras (e.g. '[hf,dev-test]' or '[all,dev-test]')"
|
|
required: false
|
|
type: string
|
|
default: "[all,dev-test]"
|
|
runner:
|
|
description: "GitHub runner to use"
|
|
required: false
|
|
type: string
|
|
default: "linux-amd64-gpu-rtxpro6000-latest-1"
|
|
allow_failure:
|
|
description: "If true, test failures are reported as a warning and do not fail the job (used to keep a known-broken example non-blocking)"
|
|
required: false
|
|
type: boolean
|
|
default: false
|
|
|
|
jobs:
|
|
run-test:
|
|
runs-on: ${{ inputs.runner }}
|
|
timeout-minutes: ${{ inputs.timeout_minutes }}
|
|
permissions:
|
|
contents: read
|
|
container:
|
|
image: ${{ inputs.docker_image }}
|
|
options: --shm-size=2gb # TRT-LLM tests on 2-GPU runner needs more shared memory
|
|
env:
|
|
PIP_CONSTRAINT: "" # Disable pip constraint for upgrading packages
|
|
HF_TOKEN: ${{ secrets.HF_TOKEN }}
|
|
# Build CUDA kernels only for the runner's RTX PRO 6000 (sm_120), not the image's ~6 archs.
|
|
TORCH_CUDA_ARCH_LIST: "12.0"
|
|
steps:
|
|
- uses: actions/checkout@v6
|
|
- uses: nv-gha-runners/setup-proxy-cache@main
|
|
- uses: ./.github/actions/cache-extensions
|
|
with:
|
|
cache-key: rtxpro6000-${{ inputs.docker_image }}
|
|
- name: Setup environment variables
|
|
run: |
|
|
echo "LD_LIBRARY_PATH=${LD_LIBRARY_PATH}:/usr/include:/usr/lib/x86_64-linux-gnu:/usr/local/tensorrt/targets/x86_64-linux-gnu/lib" >> $GITHUB_ENV
|
|
echo "PATH=${PATH}:/usr/local/tensorrt/targets/x86_64-linux-gnu/bin" >> $GITHUB_ENV
|
|
- name: Install dependencies
|
|
run: |
|
|
# Uninstall conflicting system-wide installed modelopt in nemo containers
|
|
pip uninstall -y nvidia-modelopt || true
|
|
|
|
# nvcr.io/nvidia/tensorrt:26.05-py3 ships cuDNN 9.22 with no preinstalled torch, and
|
|
# torch 2.14 pins cuDNN 9.24: mixing them fails with CUDNN_SUBLIBRARY_LOADING_FAILED.
|
|
if [[ "${{ inputs.docker_image }}" == *"/tensorrt:"* ]]; then
|
|
echo "torch<2.14" > /tmp/pip-constraints.txt
|
|
export PIP_CONSTRAINT=/tmp/pip-constraints.txt
|
|
fi
|
|
|
|
# Use `python -m pip` instead of `pip` to avoid conflicts with system pip for nemo containers
|
|
# Editable install so example scripts launched as subprocesses resolve modelopt to the same source path as the test process
|
|
python -m pip install -e ".${{ inputs.pip_install_extras }}"
|
|
|
|
if [[ "${{ inputs.example }}" == *"diffusers"* ]]; then
|
|
echo "Uninstalling apex for diffusers: T5 Int8 (PixArt) + Apex is not supported as per https://github.com/huggingface/transformers/issues/21391"
|
|
python -m pip uninstall -y apex || true
|
|
fi
|
|
|
|
find examples/${{ inputs.example }} -name "requirements.txt" | while read req_file; do python -m pip install -r "$req_file" || exit 1; done
|
|
- name: Run tests
|
|
id: run_tests
|
|
continue-on-error: ${{ inputs.allow_failure }}
|
|
env:
|
|
# Absolute paths so subprocesses running from different working directories
|
|
# all find the config and write .coverage.* files to the same location.
|
|
COVERAGE_PROCESS_START: ${{ github.workspace }}/pyproject.toml
|
|
COVERAGE_FILE: ${{ github.workspace }}/.coverage
|
|
run: |
|
|
echo "Running tests for: ${{ inputs.example }}"
|
|
python -m pytest tests/examples/${{ inputs.example }} --cov
|
|
- name: Flag allowed failure
|
|
if: ${{ inputs.allow_failure && steps.run_tests.outcome == 'failure' }}
|
|
run: |
|
|
echo "::warning title=Allowed example failure::'${{ inputs.example }}' failed but is in the allow-failure list (vars.ALLOW_FAILURE_EXAMPLE_TESTS); not blocking. Remove it from the variable once fixed."
|
|
- name: Upload coverage to Codecov
|
|
uses: codecov/codecov-action@v7
|
|
with:
|
|
token: ${{ secrets.CODECOV_TOKEN }}
|
|
files: coverage.xml
|
|
# One flag per example, not a shared `examples`: carryforward only applies to a flag
|
|
# with no upload, so a shared flag would replace every lane's coverage with the subset
|
|
# that ran once lanes are gated independently.
|
|
flags: examples-${{ inputs.example }}
|
|
fail_ci_if_error: false # test may be skipped if relevant file changes are not detected
|
|
verbose: true
|