mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? - Add experimental support for transformers >=5.0 and remove deprecated usages: https://github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md - ⚠️ For accelerate examples that used `--warmup-ratio: float` (deprecated in 5.x), we now change it to `--warmup-steps: float | int` which works as ratio if float but only for 5.x. For 4.x, it will error out if float and prompt user to change back to `--warmup-ratio` or pass an int absolute step count. - ⚠️ Unified Hugging Face checkpoint export for quantized checkpoints may not work for some models with transformers>=5.0 yet as it requires a lot of fixes (e.g. change in how MoE experts are organized) - ~Add Workaround for TRT-LLM's import of deprecated transformers functions so trt-llm based gpu unit tests work fine. Still deployment for models needs proper fixes directly in TRT-LLM hence llm/vlm ptq example tests still run with transformers 4.57~ - Everything except PTQ and Export (mainly MoE) should work fine with transformers>=5.0 - Bump min torch to 2.8 and enable 2.11 cicd testing - NOTE: Upcoming Nemo:26.04 container comes with transformers 5.3 ### Testing <!-- Mention how have you tested your change if applicable. --> - [x] CI/CD tests passing - [x] Manually tested unit tests, gpu tests with transformers 4.56 and 5.4 - [x] Manually tested example tests (except trt-llm container tests) with transformers 4.56 and 5.4 - [x] 2-gpu nightly CICD tests manually triggered and passing: [gpu tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867257540), [example tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867260643) ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, using `torch.load(..., weights_only=True)`, avoiding `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other source, did you follow IP policy in [CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?: N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Make remote-code usage opt-in via a configurable --trust_remote_code flag across examples and tools. * **Bug Fixes** * Improve checkpoint/resume detection and related training guidance to avoid erroneous errors. * **Refactor** * Consolidate dtype/config naming, switch warmup settings from ratio → steps, and unify tokenizer invocation patterns. * **Documentation** * Simplify changelog title and add misc notes for release 0.44. * **Chores** * Remove scheduled PR-branch cleanup workflow and relax/remove several transformers version pins. * **Tests** * Adjust test gates, skips, and structures to align with updated deps and behaviors. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
120 lines
4.6 KiB
YAML
120 lines
4.6 KiB
YAML
name: GPU tests
|
|
|
|
on:
|
|
push:
|
|
branches: ["pull-request/[0-9]+"]
|
|
# NOTE: paths cannot be used since push happens to copied PR and only latest commit to PR is used
|
|
schedule:
|
|
- cron: "0 0 * * *" # Nightly
|
|
workflow_dispatch: # On-demand
|
|
|
|
# Cancel previous runs if new commit is pushed to the same PR
|
|
concurrency:
|
|
group: ${{ github.workflow }}-${{ startsWith(github.ref, 'refs/heads/pull-request/') && github.ref || github.sha }}
|
|
cancel-in-progress: true
|
|
|
|
jobs:
|
|
check-file-changes:
|
|
if: startsWith(github.ref, 'refs/heads/pull-request/')
|
|
runs-on: ubuntu-latest
|
|
outputs:
|
|
any_changed: ${{ steps.changed-tests.outputs.any_changed }}
|
|
steps:
|
|
- uses: actions/checkout@v6
|
|
with:
|
|
fetch-depth: 0
|
|
- id: get-pr-info
|
|
uses: nv-gha-runners/get-pr-info@main
|
|
# Get commit from main branch that is present in the PR to use as base for changed files
|
|
- id: calculate-merge-base
|
|
env:
|
|
PR_SHA: ${{ fromJSON(steps.get-pr-info.outputs.pr-info).head.sha }}
|
|
BASE_SHA: ${{ fromJSON(steps.get-pr-info.outputs.pr-info).base.sha }}
|
|
run: |
|
|
(echo -n "merge-base="; git merge-base "$BASE_SHA" "$PR_SHA") | tee --append "${GITHUB_OUTPUT}"
|
|
- name: Check for changes in test-relevant directories
|
|
id: changed-tests
|
|
uses: step-security/changed-files@v46.0.5
|
|
with:
|
|
base_sha: ${{ steps.calculate-merge-base.outputs.merge-base }}
|
|
sha: ${{ fromJSON(steps.get-pr-info.outputs.pr-info).head.sha }}
|
|
files: |
|
|
.github/workflows/gpu_tests.yml
|
|
modelopt/**
|
|
tests/gpu/**
|
|
pyproject.toml
|
|
tox.ini
|
|
fail_on_initial_diff_error: true
|
|
wait-checks:
|
|
needs: [check-file-changes]
|
|
if: needs.check-file-changes.outputs.any_changed == 'true'
|
|
uses: ./.github/workflows/_wait_for_checks.yml
|
|
permissions:
|
|
checks: read
|
|
secrets: inherit
|
|
with:
|
|
match_pattern: "^DCO$|^linux$" # Wait for DCO and Unit tests / linux to pass
|
|
delay: 300s
|
|
gpu-tests-pr:
|
|
needs: [check-file-changes, wait-checks]
|
|
if: needs.check-file-changes.outputs.any_changed == 'true'
|
|
strategy: &gpu_strategy
|
|
fail-fast: false
|
|
matrix:
|
|
include:
|
|
- example: gpu
|
|
timeout: 45
|
|
container_image: pytorch:26.01-py3
|
|
# tests/gpu/_extensions/test_onnx_extensions.py fails for newer containers until https://github.com/tbenthompson/cppimport/pull/98
|
|
- example: gpu-megatron
|
|
timeout: 45
|
|
container_image: pytorch:26.01-py3
|
|
- example: gpu-trtllm
|
|
timeout: 30
|
|
container_image: tensorrt-llm/release:1.3.0rc10
|
|
runs-on: linux-amd64-gpu-rtxpro6000-latest-1
|
|
timeout-minutes: ${{ matrix.timeout }}
|
|
container: &gpu_container
|
|
image: nvcr.io/nvidia/${{ matrix.container_image }}
|
|
env:
|
|
GIT_DEPTH: 1000 # For correct version
|
|
PIP_CONSTRAINT: "" # Disable pip constraint for upgrading packages
|
|
HF_TOKEN: ${{ secrets.HF_TOKEN }}
|
|
steps: &gpu_steps
|
|
- uses: actions/checkout@v6
|
|
- uses: nv-gha-runners/setup-proxy-cache@main
|
|
- name: Setup environment variables
|
|
run: |
|
|
echo "LD_LIBRARY_PATH=${LD_LIBRARY_PATH}:/usr/include:/usr/lib/x86_64-linux-gnu" >> $GITHUB_ENV
|
|
- name: Run gpu tests
|
|
env:
|
|
COVERAGE_PROCESS_START: ${{ github.workspace }}/pyproject.toml
|
|
COVERAGE_FILE: ${{ github.workspace }}/.coverage
|
|
run: |
|
|
pip install tox-current-env
|
|
COV_ARGS="--cov" tox -e cuda13-${{ matrix.example }} --current-env
|
|
- name: Upload GPU coverage to Codecov
|
|
uses: codecov/codecov-action@v5
|
|
with:
|
|
token: ${{ secrets.CODECOV_TOKEN }}
|
|
files: coverage.xml
|
|
flags: gpu
|
|
fail_ci_if_error: false # test may be skipped if relevant file changes are not detected
|
|
verbose: true
|
|
gpu-tests-non-pr:
|
|
if: ${{ !startsWith(github.ref, 'refs/heads/pull-request/') }}
|
|
strategy: *gpu_strategy
|
|
runs-on: linux-amd64-gpu-rtxpro6000-latest-2
|
|
timeout-minutes: ${{ matrix.timeout }}
|
|
container: *gpu_container
|
|
steps: *gpu_steps
|
|
gpu-pr-required-check:
|
|
# Run even if gpu-tests-pr is skipped
|
|
if: ${{ startsWith(github.ref, 'refs/heads/pull-request/') && always() }}
|
|
needs: [check-file-changes, gpu-tests-pr]
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- name: Required GPU tests did not succeed
|
|
if: ${{ needs.check-file-changes.result != 'success' || (needs.check-file-changes.outputs.any_changed == 'true' && needs.gpu-tests-pr.result != 'success') }}
|
|
run: exit 1
|