mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: CI/CD bug fix Every example lane uploaded coverage under one shared `examples` flag. That was correct while all lanes ran on every PR, but #2090 gates them independently, and **Codecov carryforward only applies to a flag with no upload on the commit**. | scenario | `examples` flag | outcome | |---|---|---| | no lanes run (docs-only) | absent | ✅ carried forward | | all lanes run | complete | ✅ correct | | **one lane runs** (now common) | **present but partial** | ❌ carryforward skipped; full coverage replaced by that lane's subset | The third row is what gating made routine: a PR touching only `examples/diffusers/**` runs the onnx lane, uploads `examples` containing onnx coverage alone, and Codecov reports a drop for code the PR never touched. One flag per example (`examples-<name>`, 12 flags) restores the intent already documented in `.github/codecov.yml`: a skipped lane has no upload for its flag and is carried forward; a lane that ran replaces only its own slice. The config comment is updated to explain why a shared flag defeats carryforward, so this isn't re-introduced. `gpu_tests` deliberately keeps a single `gpu` flag — its five suites are gated at workflow level, so they upload together or not at all. It would need the same change if per-suite gating is ever added. ### Testing Not directly observable on this PR: it changes workflow files, which are in the gate's `common` group, so **all twelve lanes run** and every flag is uploaded — the healthy case either way. The behavior it fixes appears on the next PR that touches a single example, where `codecov/project` should now stay accurate instead of reporting a drop. Worth noting the symptom was never blocking: `codecov/project` is not a required check and the threshold allows a 2% drop. This is about the coverage data being right. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — flags are new names; historical data under `examples` is unaffected - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — CI configuration - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ — not yet run <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Improved coverage reporting by tracking results separately for each test example. * Clarified coverage configuration and documented how skipped uploads are carried forward. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
101 lines
4.5 KiB
YAML
101 lines
4.5 KiB
YAML
# Reusable workflow for running example tests
|
|
name: Example Tests Runner
|
|
|
|
on:
|
|
workflow_call:
|
|
inputs:
|
|
docker_image:
|
|
description: "Docker image to use for tests"
|
|
required: true
|
|
type: string
|
|
example:
|
|
description: "Example name to test (e.g. 'hf_ptq')"
|
|
required: true
|
|
type: string
|
|
timeout_minutes:
|
|
description: "Timeout in minutes for the job"
|
|
required: false
|
|
type: number
|
|
default: 60
|
|
pip_install_extras:
|
|
description: "Pip install extras (e.g. '[hf,dev-test]' or '[all,dev-test]')"
|
|
required: false
|
|
type: string
|
|
default: "[all,dev-test]"
|
|
runner:
|
|
description: "GitHub runner to use"
|
|
required: false
|
|
type: string
|
|
default: "linux-amd64-gpu-rtxpro6000-latest-1"
|
|
allow_failure:
|
|
description: "If true, test failures are reported as a warning and do not fail the job (used to keep a known-broken example non-blocking)"
|
|
required: false
|
|
type: boolean
|
|
default: false
|
|
|
|
jobs:
|
|
run-test:
|
|
runs-on: ${{ inputs.runner }}
|
|
timeout-minutes: ${{ inputs.timeout_minutes }}
|
|
permissions:
|
|
contents: read
|
|
container:
|
|
image: ${{ inputs.docker_image }}
|
|
options: --shm-size=2gb # TRT-LLM tests on 2-GPU runner needs more shared memory
|
|
env:
|
|
PIP_CONSTRAINT: "" # Disable pip constraint for upgrading packages
|
|
HF_TOKEN: ${{ secrets.HF_TOKEN }}
|
|
# Build CUDA kernels only for the runner's RTX PRO 6000 (sm_120), not the image's ~6 archs.
|
|
TORCH_CUDA_ARCH_LIST: "12.0"
|
|
steps:
|
|
- uses: actions/checkout@v6
|
|
- uses: nv-gha-runners/setup-proxy-cache@main
|
|
- uses: ./.github/actions/cache-extensions
|
|
with:
|
|
cache-key: rtxpro6000-${{ inputs.docker_image }}
|
|
- name: Setup environment variables
|
|
run: |
|
|
echo "LD_LIBRARY_PATH=${LD_LIBRARY_PATH}:/usr/include:/usr/lib/x86_64-linux-gnu:/usr/local/tensorrt/targets/x86_64-linux-gnu/lib" >> $GITHUB_ENV
|
|
echo "PATH=${PATH}:/usr/local/tensorrt/targets/x86_64-linux-gnu/bin" >> $GITHUB_ENV
|
|
- name: Install dependencies
|
|
run: |
|
|
# Uninstall conflicting system-wide installed modelopt in nemo containers
|
|
pip uninstall -y nvidia-modelopt || true
|
|
|
|
# Use `python -m pip` instead of `pip` to avoid conflicts with system pip for nemo containers
|
|
# Editable install so example scripts launched as subprocesses resolve modelopt to the same source path as the test process
|
|
python -m pip install -e ".${{ inputs.pip_install_extras }}"
|
|
|
|
if [[ "${{ inputs.example }}" == *"diffusers"* ]]; then
|
|
echo "Uninstalling apex for diffusers: T5 Int8 (PixArt) + Apex is not supported as per https://github.com/huggingface/transformers/issues/21391"
|
|
python -m pip uninstall -y apex || true
|
|
fi
|
|
|
|
find examples/${{ inputs.example }} -name "requirements.txt" | while read req_file; do python -m pip install -r "$req_file" || exit 1; done
|
|
- name: Run tests
|
|
id: run_tests
|
|
continue-on-error: ${{ inputs.allow_failure }}
|
|
env:
|
|
# Absolute paths so subprocesses running from different working directories
|
|
# all find the config and write .coverage.* files to the same location.
|
|
COVERAGE_PROCESS_START: ${{ github.workspace }}/pyproject.toml
|
|
COVERAGE_FILE: ${{ github.workspace }}/.coverage
|
|
run: |
|
|
echo "Running tests for: ${{ inputs.example }}"
|
|
python -m pytest tests/examples/${{ inputs.example }} --cov
|
|
- name: Flag allowed failure
|
|
if: ${{ inputs.allow_failure && steps.run_tests.outcome == 'failure' }}
|
|
run: |
|
|
echo "::warning title=Allowed example failure::'${{ inputs.example }}' failed but is in the allow-failure list (vars.ALLOW_FAILURE_EXAMPLE_TESTS); not blocking. Remove it from the variable once fixed."
|
|
- name: Upload coverage to Codecov
|
|
uses: codecov/codecov-action@v7
|
|
with:
|
|
token: ${{ secrets.CODECOV_TOKEN }}
|
|
files: coverage.xml
|
|
# One flag per example, not a shared `examples`: carryforward only applies to a flag
|
|
# with no upload, so a shared flag would replace every lane's coverage with the subset
|
|
# that ran once lanes are gated independently.
|
|
flags: examples-${{ inputs.example }}
|
|
fail_ci_if_error: false # test may be skipped if relevant file changes are not detected
|
|
verbose: true
|