14 Commits
Author SHA1 Message Date
Keval Morabia 51cc5dbade Make torch 2.14 the unit-test default and constrain it for the tensorrt example images (#2309)
### What does this PR do?

Type of change: Bug fix (CI) + test coverage

**Fixes `onnx (torch_onnx)` and `onnx (diffusers)`**, which have failed
on every branch since
`torch 2.14.0` was published to PyPI today (2026-09-02 13:42 UTC), and
**adds torch 2.14 to the unit
test matrix as the new default** so the next torch release is caught
there rather than in an example job.

### Root cause

Every test in those two jobs failed with:

```
RuntimeError: CUDNN_BACKEND_TENSOR_DESCRIPTOR cudnnFinalize failed
  ptrDesc->finalize() cudnn_status: CUDNN_STATUS_SUBLIBRARY_LOADING_FAILED
```

`nvcr.io/nvidia/tensorrt:26.05-py3` ships cuDNN **9.22** and has no
preinstalled torch, so pip
resolved the newest one — and torch 2.14 pins
`nvidia-cudnn-cu13==9.24.0.43`. Loading 9.24
sublibraries against the image's 9.22 `libcudnn.so.9` is exactly what
that status reports.

| | last good run (08:55) | first failing run (13:34) |
|---|---|---|
| `torch` | 2.13.0 | **2.14.0** |
| `nvidia-cudnn-cu13` | 9.20.0.48 | **9.24.0.43** |
| image cuDNN | 9.22.0.52 | 9.22.0.52 |

### Why only these two jobs

- The **nemo** and **pytorch** images have a preinstalled torch that
already satisfies `torch>=2.8`,
so pip never resolves a new one — confirmed from the megatron job log,
where torch does not appear
  in `Successfully installed`.
- **`tensorrt:26.05-py3` has no preinstalled torch**, so pip takes the
newest from PyPI.
- **`onnx (torch_trt)`** shares that image but passes throughout,
because `torch-tensorrt<2.13`
  already holds torch below 2.14.

### The changes

1. **Constrain torch only where the incompatibility is.**
`PIP_CONSTRAINT=torch<2.14` in the example
runner, applied when the job's image is a `tensorrt` one. It also covers
the
`examples/*/requirements.txt` loop in the same shell, which matters
because `nemo_automodel`
pulls torch in too. Not pinned in `pyproject.toml`: torch 2.14 is fine
anywhere its own bundled
cuDNN is the one loaded, so that would constrain users to work around
one pinned image.
2. **Test torch 2.14.** `torch_214` added to `TORCH_VERSIONS`
(`torchvision~=0.29.0`) and promoted to
the unit-test default across the supported Python versions, with 2.13
demoted to the back-compat
row. `release.yml`'s basic unit test moves to the same default (it was
still on 2.12).
Nothing exercised 2.14 before — which is why a torch release reached us
through an example
   job instead of a unit test.

### Testing

- `actionlint` and YAML/TOML parse clean; pre-commit clean.
- Verified by this PR's own jobs: `onnx (torch_onnx)` and `onnx
(diffusers)` reproduce the failure on
`main` right now, and the new `unit-3.12(torch_214, tf_latest)` job is
the first run of ModelOpt
  against torch 2.14.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — CI-only; no source or package
metadata change
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependency
- Did you write any new necessary tests?: ✅ — torch 2.14 added to the
unit test matrix
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — internal CI, not user-facing
- Did you get Claude approval on this PR?: ❌ — not yet requested

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-09-03 01:11:16 +05:30
Keval MorabiaandClaude Opus 5 3d4d9249f4 Gate example and GPU test lanes on the files they cover, and consolidate the CI gate (#2090)
### What does this PR do?

Type of change: CI/CD improvement

Follow-up to #2086, which carried the speculative-decoding fix; this PR
is the CI half.

**Lanes now run only when their own files change.** A one-line edit to
any example started all 12 example lanes, and any `modelopt/**` change
started every GPU suite.

- Each lane gates itself inside `_example_tests_runner.yml` /
`_gpu_tests_runner.yml`, deriving its watch list from the example or
suite name it was already given (`examples/<name>/**`,
`tests/examples/<name>/**`, `tests/<suite>/**`). **Adding a new example
stays a one-line matrix entry** — no central mapping to update.
- The five cross-example dependencies are declared as `watch_extra` next
to the example that needs them: `hf_ptq` → `llm_eval`
(`huggingface_example.sh` runs lm_eval from `../llm_eval`), `torch_trt`
→ `onnx_ptq`, `speculative_decoding` → `hf_ptq` + `dataset`, `llm_qat` /
`gpt-oss` → `dataset`.
- `gpu_tests.yml` is split into caller + runner to match. This shape is
forced, not stylistic: job-level `if:` cannot read `matrix`, and
`container:` images are pulled before any step runs, so gating inside
the job would still pull 10–20 GB and hold a GPU runner for every
skipped suite.
- `modelopt/**`, `modelopt_recipes/**`, `pyproject.toml` and
`tests/_test_utils/**` still run every lane. The gate now watches the
last two, which example tests depend on but it previously ignored.

**Docs-only changes no longer start GPU jobs.** The gate ignores
`**.md`, `**.rst`, `**.png` and `**.ipynb` by default, so a README edit
short-circuits the whole workflow. Nothing executes notebooks (no
`nbmake`/`nbval` in the repo), and `.sh`/`.yaml`/`.txt` stay watched
since examples run them.

**One file holds the gate logic.** `.github/actions/changed-files-gate`
is a composite action doing the merge-base + changed-files comparison
and, optionally, the `^linux$` wait. `_pr_gate.yml` and
`_wait_for_checks.yml` are both deleted: each top-level workflow keeps a
12-line `pr-gate` job that is pure wiring, and the runners use the
action as steps so a lane shows `gate` + `run-test` rather than three
checks. Calling a reusable workflow always materializes all of its jobs,
including skipped ones, which is what made the per-lane check list
noisy.

`unit_tests.yml` also drops its DCO wait: DCO can be marked passing
manually, so blocking the matrix on it only delayed feedback. The
`^linux$` wait remains, which is the gate that actually protects GPU
runners.

**Also, from the original lane consolidation:**

- **One TensorRT-LLM lane.** `trtllm-pr` and `trtllm-non-pr` merge into
a single `trtllm` job gated like the others, so `llm_eval` now runs on
PRs, where it was nightly-only.
- **`gpt-oss` moves to the TensorRT-LLM image.** Its deploy step needs
`tensorrt_llm`, which `pytorch` doesn't have, so `deploy_gpt_oss_trtllm`
silently skipped in CI. Its `importorskip` is dropped now that the lane
guarantees the dependency.
- **Containers bumped where no reason was documented:** pytorch
`26.06`/`26.01` → `26.07` (torch example lane, regression). Left pinned
with their existing in-file reasons: pytorch `26.05` for the gpu lane
(`EXPLICIT_BATCH` removed in TensorRT 11), tensorrt `26.05` for the onnx
lane (`torch-tensorrt` needs `libnvinfer.so.10`), vllm `v0.20.0` (legacy
FusedMoE coverage). TensorRT-LLM stays on `1.3.0rc20`: rc21–rc23 ship a
`quickstart_multimodal.py` importing `MultimodalConfig` before
`tensorrt_llm.llmapi` exported it (fixed upstream in
NVIDIA/TensorRT-LLM#17112, one day after rc23 was cut), which fails the
`hf_ptq` VLM deploy smoke test.

Also clarifies the changelog line in the PR template to spell out when
an entry is expected.

**Two silent-failure fixes found in review:**

- Every gate used `any_changed`, which is ACMR and excludes deletions,
so a delete-only PR (removing an example, a test, or library code) ran
nothing. Now `any_modified` (ACMRD), fixed in `example_tests.yml`,
`_pr_gate.yml` and `unit_tests.yml`.
- `git merge-base` was piped into `tee`, so a failure returned `tee`'s
exit status and emitted an empty base, quietly changing which lanes run
instead of failing.

**Nightly secret scanning is fixed.** `code_quality.yml` excludes
trufflehog's `lob` detector: its `(live|test)_[a-zA-Z0-9_]{35}` pattern
matches any pytest function whose name is exactly 35 characters after
`test_` — 60 of them in this repo, e.g.
`test_all_zero_activation_yields_no_scale` — and reports them as
*verified* secrets. This only ever failed nightly because the action
scans all history on `schedule` and only the PR diff on `pull_request`.

### Testing

Workflow YAML validated locally and the selection logic checked by hand
across scenarios (`examples/diffusers/**` → onnx lane only;
`examples/dataset/**` → `llm_qat`, `speculative_decoding`, `gpt-oss`;
`examples/llm_eval/**` → `hf_ptq` + `llm_eval`; `modelopt/**` and
nightly → everything). Gating behavior itself can only be exercised by a
real PR run — the failure mode to watch for is a lane skipping when it
should have run.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — CI configuration
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ — not yet run


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **CI Improvements**
- Improved change detection for model recipes, test utilities, workflow
actions, and documentation-only updates.
  - Streamlined pull request checks and status monitoring.
- Consolidated TensorRT-LLM example validation and refined conditional
test execution.
  - Updated test environments to newer PyTorch releases.
  - Refined GPU, regression, and unit test triggers.
- Deployment tests no longer automatically skip when TensorRT-LLM is
unavailable.
  - Updated secret scanning configuration.

- **Documentation**
- Updated pull request checklist guidance to include deprecations and
critical bug fixes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:46:46 +05:30
Md Shahrier Islam Arham 0c40c374e5 ci: give non-PR code quality runs distinct concurrency groups (#1894)
### What does this PR do?

Type of change: Bug fix

The concurrency group used `github.event.pull_request.number` with no
fallback, so every nightly and manually dispatched run shared the
literal group `Code Quality-` with `cancel-in-progress: true` — a manual
dispatch cancels an in-flight nightly and vice versa. Adds the `||
github.sha` fallback already used by `unit_tests.yml`.

### Usage

N/A — CI workflow configuration change.

### Testing

Matches the existing pattern in `unit_tests.yml` line-for-line. YAML
validated.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (workflow config;
validated as described above)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (CI-only change)
- Did you get Claude approval on this PR?: N/A (external contributor;
cannot trigger `/claude review`)

### Additional Information

Part of #1890 (item 4).


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Improved workflow concurrency handling so automated runs are grouped
and canceled more reliably across pull request, manual, and scheduled
triggers.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: arham766 <arhamislam766@yahoo.com>
2026-07-05 11:14:20 +00:00
Keval MorabiaandClaude Sonnet 4.6 3d0f0db49e [CI] Replace tox with nox, use nemo:26.04 for megatron tests, and simplify CI workflows (#1286)
### What does this PR do?

Type of change: New feature / infrastructure improvement

Follow-up to #1285 for correct CI test environment for megatron based
tests

Replaces `tox` + `tox-current-env` with `nox` for all test, lint, docs,
and wheel build sessions. The primary motivation was that
`tox-current-env` is incompatible with uv venvs in NGC containers (e.g.
NeMo's `/opt/venv`) — it picks the system Python via
`sys._base_executable` instead of the container's venv Python which has
megatron packages pre-installed.

Key changes:
- **`noxfile.py`** replaces `tox.ini` with GPU, CPU unit,
partial-install, pre-commit, docs, and wheel sessions
- **GPU sessions** use `venv_backend="none"` (run directly in container
env) and `python -m pip/pytest` to avoid PATH mismatches
- **uv** is set as the default venv backend (if available) for CPU
sessions (faster installs)

Also includes CI workflow simplifications:
- **`_pr_gate.yml`** new reusable workflow centralizing file-change
detection + linux-check wait logic (was duplicated across 3 workflow
files)
- **Collapsed pr/non-pr job pairs** into single jobs with conditional
`runs-on` in `gpu_tests.yml`, `example_tests.yml`,
`regression_tests.yml`
- **Collapsed `multi-py` / `multi-torch` / `multi-transformers`** into a
single `multi-version` matrix job in `unit_tests.yml`
- **PR path filtering** for unit test secondary jobs (multi-version,
launcher, partial-install) — skipped if no relevant files changed
- **Fixed schedule/workflow_dispatch skipping** — jobs with `needs:
[pr-gate]` were incorrectly skipped when all pr-gate internal jobs were
skipped; fixed by making the gate job always run
- **multi-version, launcher, partial-install** now also run on
`schedule` / `workflow_dispatch`

### Usage

```bash
python -m pip install nox uv                                                    # install nox and uv (once)
nox -l                                                                          # list all sessions
nox -s gpu_megatron                                                             # run a GPU session (inside container)
nox -s "unit-3.12(torch_211, tf_latest)"                                        # run a specific unit test combination
nox -s "unit-3.12(torch_211, tf_latest)" -R                                     # force-recreate venv (e.g. after dep changes)
COVERAGE_PROCESS_START=pyproject.toml nox -s "unit-3.12(torch_211, tf_latest)"  # with coverage
```

### Testing
- Ran `nox -l` to verify all session names
- Ran `gpu_megatron` session locally inside NeMo container — confirmed
it uses `/opt/venv/bin/python` correctly
- Manually triggered nightly-runs:
- Unit:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608013657
- GPU:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608018763
- Examples:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608017322

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A — CI infrastructure only
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ (added `nox`
and `uv` to `dev-test`, both Apache-2.0)
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — no user-facing changes

### Additional Information
Supersedes the tox-current-env workaround in the parent branch.

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-18 16:56:02 +00:00
Keval Morabia b484efb84e [CI] Cleanup ubuntu-runner disk storage before installing deps (#765)
We started seeing this issue in GitHub's free ubuntu-latest runners:
`ERROR: Could not install packages due to an OSError: [Errno 28] No
space left on device`.

Suggested by other GH runners users to remove unnecessary android /
dotnet files to avoid the storage issue.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Consolidated GitHub Actions workflow setup into a reusable custom
action for improved maintainability and consistency across CI/CD
pipelines.
* Enhanced release workflow with automated unit testing and artifact
upload capabilities.
* Streamlined runner initialization by reducing redundant configuration
steps.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-01-13 02:18:41 +05:30
Salman Chishti d8d5a2947f Upgrade GitHub Actions for Node 24 compatibility (#698)
## Summary

Upgrade GitHub Actions to their latest versions to ensure compatibility
with Node 24, as Node 20 will reach end-of-life in April 2026.

## Changes

| Action | Old Version(s) | New Version | Release | Files |
|--------|---------------|-------------|---------|-------|
| `actions/checkout` |
[`v4`](https://github.com/actions/checkout/releases/tag/v4) |
[`v6`](https://github.com/actions/checkout/releases/tag/v6) |
[Release](https://github.com/actions/checkout/releases/tag/v6) |
_example_tests_runner.yml, code_quality.yml, example_tests.yml,
gpu_tests.yml, pages.yml, unit_tests.yml |
| `actions/setup-python` |
[`v5`](https://github.com/actions/setup-python/releases/tag/v5) |
[`v6`](https://github.com/actions/setup-python/releases/tag/v6) |
[Release](https://github.com/actions/setup-python/releases/tag/v6) |
code_quality.yml, pages.yml, unit_tests.yml |

## Context

Per [GitHub's
announcement](https://github.blog/changelog/2025-09-19-deprecation-of-node-20-on-github-actions-runners/),
Node 20 is being deprecated and runners will begin using Node 24 by
default starting March 4th, 2026.

### Why this matters

- **Node 20 EOL**: April 2026
- **Node 24 default**: March 4th, 2026
- **Action**: Update to latest action versions that support Node 24

### Security Note

Actions that were previously pinned to commit SHAs remain pinned to SHAs
(updated to the latest release SHA) to maintain the security benefits of
immutable references.

### Testing

These changes only affect CI/CD workflow configurations and should not
impact application functionality. The workflows should be tested by
running them on a branch before merging.

Signed-off-by: Salman Muin Kayser Chishti <13schishti@gmail.com>
2025-12-16 22:00:34 +05:30
Keval Morabia 37c4974311 Run CICD on feature/* breanches as well
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-10-27 02:42:16 -07:00
Keval Morabia 4c36abe536 Add llm_ptq PR test, Cleanup dependency installation in modelopt docker (#338)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-09-19 10:48:44 +05:30
Keval Morabia 90bcf3aafa Fix GPU test CI condition to run nightly or on demand (#269)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-08-27 02:55:53 +05:30
Keval Morabia 4b2847256e Add trufflehog secret scanning CI action (#267)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-08-26 00:57:52 +05:30
Keval Morabia 5efc1729de Schedule nightly unit/gpu tests
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-08-12 17:02:35 +05:30
Keval Morabia de20a6a054 Enable cpu unit tests in Github CI (#210) 2025-06-17 03:27:57 +05:30
Keval Morabia 11b3eb6c78 Add tox.ini and fix code_quality workflow 2025-06-11 16:15:59 -07:00
Keval Morabia d6e32e9968 Add code quality checks for pull requests 2025-06-10 15:56:35 -07:00