mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: CI/CD improvement Follow-up to #2086, which carried the speculative-decoding fix; this PR is the CI half. **Lanes now run only when their own files change.** A one-line edit to any example started all 12 example lanes, and any `modelopt/**` change started every GPU suite. - Each lane gates itself inside `_example_tests_runner.yml` / `_gpu_tests_runner.yml`, deriving its watch list from the example or suite name it was already given (`examples/<name>/**`, `tests/examples/<name>/**`, `tests/<suite>/**`). **Adding a new example stays a one-line matrix entry** — no central mapping to update. - The five cross-example dependencies are declared as `watch_extra` next to the example that needs them: `hf_ptq` → `llm_eval` (`huggingface_example.sh` runs lm_eval from `../llm_eval`), `torch_trt` → `onnx_ptq`, `speculative_decoding` → `hf_ptq` + `dataset`, `llm_qat` / `gpt-oss` → `dataset`. - `gpu_tests.yml` is split into caller + runner to match. This shape is forced, not stylistic: job-level `if:` cannot read `matrix`, and `container:` images are pulled before any step runs, so gating inside the job would still pull 10–20 GB and hold a GPU runner for every skipped suite. - `modelopt/**`, `modelopt_recipes/**`, `pyproject.toml` and `tests/_test_utils/**` still run every lane. The gate now watches the last two, which example tests depend on but it previously ignored. **Docs-only changes no longer start GPU jobs.** The gate ignores `**.md`, `**.rst`, `**.png` and `**.ipynb` by default, so a README edit short-circuits the whole workflow. Nothing executes notebooks (no `nbmake`/`nbval` in the repo), and `.sh`/`.yaml`/`.txt` stay watched since examples run them. **One file holds the gate logic.** `.github/actions/changed-files-gate` is a composite action doing the merge-base + changed-files comparison and, optionally, the `^linux$` wait. `_pr_gate.yml` and `_wait_for_checks.yml` are both deleted: each top-level workflow keeps a 12-line `pr-gate` job that is pure wiring, and the runners use the action as steps so a lane shows `gate` + `run-test` rather than three checks. Calling a reusable workflow always materializes all of its jobs, including skipped ones, which is what made the per-lane check list noisy. `unit_tests.yml` also drops its DCO wait: DCO can be marked passing manually, so blocking the matrix on it only delayed feedback. The `^linux$` wait remains, which is the gate that actually protects GPU runners. **Also, from the original lane consolidation:** - **One TensorRT-LLM lane.** `trtllm-pr` and `trtllm-non-pr` merge into a single `trtllm` job gated like the others, so `llm_eval` now runs on PRs, where it was nightly-only. - **`gpt-oss` moves to the TensorRT-LLM image.** Its deploy step needs `tensorrt_llm`, which `pytorch` doesn't have, so `deploy_gpt_oss_trtllm` silently skipped in CI. Its `importorskip` is dropped now that the lane guarantees the dependency. - **Containers bumped where no reason was documented:** pytorch `26.06`/`26.01` → `26.07` (torch example lane, regression). Left pinned with their existing in-file reasons: pytorch `26.05` for the gpu lane (`EXPLICIT_BATCH` removed in TensorRT 11), tensorrt `26.05` for the onnx lane (`torch-tensorrt` needs `libnvinfer.so.10`), vllm `v0.20.0` (legacy FusedMoE coverage). TensorRT-LLM stays on `1.3.0rc20`: rc21–rc23 ship a `quickstart_multimodal.py` importing `MultimodalConfig` before `tensorrt_llm.llmapi` exported it (fixed upstream in NVIDIA/TensorRT-LLM#17112, one day after rc23 was cut), which fails the `hf_ptq` VLM deploy smoke test. Also clarifies the changelog line in the PR template to spell out when an entry is expected. **Two silent-failure fixes found in review:** - Every gate used `any_changed`, which is ACMR and excludes deletions, so a delete-only PR (removing an example, a test, or library code) ran nothing. Now `any_modified` (ACMRD), fixed in `example_tests.yml`, `_pr_gate.yml` and `unit_tests.yml`. - `git merge-base` was piped into `tee`, so a failure returned `tee`'s exit status and emitted an empty base, quietly changing which lanes run instead of failing. **Nightly secret scanning is fixed.** `code_quality.yml` excludes trufflehog's `lob` detector: its `(live|test)_[a-zA-Z0-9_]{35}` pattern matches any pytest function whose name is exactly 35 characters after `test_` — 60 of them in this repo, e.g. `test_all_zero_activation_yields_no_scale` — and reports them as *verified* secrets. This only ever failed nightly because the action scans all history on `schedule` and only the PR diff on `pull_request`. ### Testing Workflow YAML validated locally and the selection logic checked by hand across scenarios (`examples/diffusers/**` → onnx lane only; `examples/dataset/**` → `llm_qat`, `speculative_decoding`, `gpt-oss`; `examples/llm_eval/**` → `hf_ptq` + `llm_eval`; `modelopt/**` and nightly → everything). Gating behavior itself can only be exercised by a real PR run — the failure mode to watch for is a lane skipping when it should have run. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — CI configuration - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ — not yet run <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **CI Improvements** - Improved change detection for model recipes, test utilities, workflow actions, and documentation-only updates. - Streamlined pull request checks and status monitoring. - Consolidated TensorRT-LLM example validation and refined conditional test execution. - Updated test environments to newer PyTorch releases. - Refined GPU, regression, and unit test triggers. - Deployment tests no longer automatically skip when TensorRT-LLM is unavailable. - Updated secret scanning configuration. - **Documentation** - Updated pull request checklist guidance to include deprecations and critical bug fixes. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>