Files
Keval Morabia 51cc5dbade Make torch 2.14 the unit-test default and constrain it for the tensorrt example images (#2309)
### What does this PR do?

Type of change: Bug fix (CI) + test coverage

**Fixes `onnx (torch_onnx)` and `onnx (diffusers)`**, which have failed
on every branch since
`torch 2.14.0` was published to PyPI today (2026-09-02 13:42 UTC), and
**adds torch 2.14 to the unit
test matrix as the new default** so the next torch release is caught
there rather than in an example job.

### Root cause

Every test in those two jobs failed with:

```
RuntimeError: CUDNN_BACKEND_TENSOR_DESCRIPTOR cudnnFinalize failed
  ptrDesc->finalize() cudnn_status: CUDNN_STATUS_SUBLIBRARY_LOADING_FAILED
```

`nvcr.io/nvidia/tensorrt:26.05-py3` ships cuDNN **9.22** and has no
preinstalled torch, so pip
resolved the newest one — and torch 2.14 pins
`nvidia-cudnn-cu13==9.24.0.43`. Loading 9.24
sublibraries against the image's 9.22 `libcudnn.so.9` is exactly what
that status reports.

| | last good run (08:55) | first failing run (13:34) |
|---|---|---|
| `torch` | 2.13.0 | **2.14.0** |
| `nvidia-cudnn-cu13` | 9.20.0.48 | **9.24.0.43** |
| image cuDNN | 9.22.0.52 | 9.22.0.52 |

### Why only these two jobs

- The **nemo** and **pytorch** images have a preinstalled torch that
already satisfies `torch>=2.8`,
so pip never resolves a new one — confirmed from the megatron job log,
where torch does not appear
  in `Successfully installed`.
- **`tensorrt:26.05-py3` has no preinstalled torch**, so pip takes the
newest from PyPI.
- **`onnx (torch_trt)`** shares that image but passes throughout,
because `torch-tensorrt<2.13`
  already holds torch below 2.14.

### The changes

1. **Constrain torch only where the incompatibility is.**
`PIP_CONSTRAINT=torch<2.14` in the example
runner, applied when the job's image is a `tensorrt` one. It also covers
the
`examples/*/requirements.txt` loop in the same shell, which matters
because `nemo_automodel`
pulls torch in too. Not pinned in `pyproject.toml`: torch 2.14 is fine
anywhere its own bundled
cuDNN is the one loaded, so that would constrain users to work around
one pinned image.
2. **Test torch 2.14.** `torch_214` added to `TORCH_VERSIONS`
(`torchvision~=0.29.0`) and promoted to
the unit-test default across the supported Python versions, with 2.13
demoted to the back-compat
row. `release.yml`'s basic unit test moves to the same default (it was
still on 2.12).
Nothing exercised 2.14 before — which is why a torch release reached us
through an example
   job instead of a unit test.

### Testing

- `actionlint` and YAML/TOML parse clean; pre-commit clean.
- Verified by this PR's own jobs: `onnx (torch_onnx)` and `onnx
(diffusers)` reproduce the failure on
`main` right now, and the new `unit-3.12(torch_214, tf_latest)` job is
the first run of ModelOpt
  against torch 2.14.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — CI-only; no source or package
metadata change
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependency
- Did you write any new necessary tests?: ✅ — torch 2.14 added to the
unit test matrix
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — internal CI, not user-facing
- Did you get Claude approval on this PR?: ❌ — not yet requested

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-09-03 01:11:16 +05:30

108 lines
4.8 KiB
YAML

# Reusable workflow for running example tests
name: Example Tests Runner
on:
workflow_call:
inputs:
docker_image:
description: "Docker image to use for tests"
required: true
type: string
example:
description: "Example name to test (e.g. 'hf_ptq')"
required: true
type: string
timeout_minutes:
description: "Timeout in minutes for the job"
required: false
type: number
default: 60
pip_install_extras:
description: "Pip install extras (e.g. '[hf,dev-test]' or '[all,dev-test]')"
required: false
type: string
default: "[all,dev-test]"
runner:
description: "GitHub runner to use"
required: false
type: string
default: "linux-amd64-gpu-rtxpro6000-latest-1"
allow_failure:
description: "If true, test failures are reported as a warning and do not fail the job (used to keep a known-broken example non-blocking)"
required: false
type: boolean
default: false
jobs:
run-test:
runs-on: ${{ inputs.runner }}
timeout-minutes: ${{ inputs.timeout_minutes }}
permissions:
contents: read
container:
image: ${{ inputs.docker_image }}
options: --shm-size=2gb # TRT-LLM tests on 2-GPU runner needs more shared memory
env:
PIP_CONSTRAINT: "" # Disable pip constraint for upgrading packages
HF_TOKEN: ${{ secrets.HF_TOKEN }}
# Build CUDA kernels only for the runner's RTX PRO 6000 (sm_120), not the image's ~6 archs.
TORCH_CUDA_ARCH_LIST: "12.0"
steps:
- uses: actions/checkout@v6
- uses: nv-gha-runners/setup-proxy-cache@main
- uses: ./.github/actions/cache-extensions
with:
cache-key: rtxpro6000-${{ inputs.docker_image }}
- name: Setup environment variables
run: |
echo "LD_LIBRARY_PATH=${LD_LIBRARY_PATH}:/usr/include:/usr/lib/x86_64-linux-gnu:/usr/local/tensorrt/targets/x86_64-linux-gnu/lib" >> $GITHUB_ENV
echo "PATH=${PATH}:/usr/local/tensorrt/targets/x86_64-linux-gnu/bin" >> $GITHUB_ENV
- name: Install dependencies
run: |
# Uninstall conflicting system-wide installed modelopt in nemo containers
pip uninstall -y nvidia-modelopt || true
# nvcr.io/nvidia/tensorrt:26.05-py3 ships cuDNN 9.22 with no preinstalled torch, and
# torch 2.14 pins cuDNN 9.24: mixing them fails with CUDNN_SUBLIBRARY_LOADING_FAILED.
if [[ "${{ inputs.docker_image }}" == *"/tensorrt:"* ]]; then
echo "torch<2.14" > /tmp/pip-constraints.txt
export PIP_CONSTRAINT=/tmp/pip-constraints.txt
fi
# Use `python -m pip` instead of `pip` to avoid conflicts with system pip for nemo containers
# Editable install so example scripts launched as subprocesses resolve modelopt to the same source path as the test process
python -m pip install -e ".${{ inputs.pip_install_extras }}"
if [[ "${{ inputs.example }}" == *"diffusers"* ]]; then
echo "Uninstalling apex for diffusers: T5 Int8 (PixArt) + Apex is not supported as per https://github.com/huggingface/transformers/issues/21391"
python -m pip uninstall -y apex || true
fi
find examples/${{ inputs.example }} -name "requirements.txt" | while read req_file; do python -m pip install -r "$req_file" || exit 1; done
- name: Run tests
id: run_tests
continue-on-error: ${{ inputs.allow_failure }}
env:
# Absolute paths so subprocesses running from different working directories
# all find the config and write .coverage.* files to the same location.
COVERAGE_PROCESS_START: ${{ github.workspace }}/pyproject.toml
COVERAGE_FILE: ${{ github.workspace }}/.coverage
run: |
echo "Running tests for: ${{ inputs.example }}"
python -m pytest tests/examples/${{ inputs.example }} --cov
- name: Flag allowed failure
if: ${{ inputs.allow_failure && steps.run_tests.outcome == 'failure' }}
run: |
echo "::warning title=Allowed example failure::'${{ inputs.example }}' failed but is in the allow-failure list (vars.ALLOW_FAILURE_EXAMPLE_TESTS); not blocking. Remove it from the variable once fixed."
- name: Upload coverage to Codecov
uses: codecov/codecov-action@v7
with:
token: ${{ secrets.CODECOV_TOKEN }}
files: coverage.xml
# One flag per example, not a shared `examples`: carryforward only applies to a flag
# with no upload, so a shared flag would replace every lane's coverage with the subset
# that ran once lanes are gated independently.
flags: examples-${{ inputs.example }}
fail_ci_if_error: false # test may be skipped if relevant file changes are not detected
verbose: true