mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do?
Type of change: CI / infrastructure (build-time speedup)
ModelOpt's CUDA quantization extensions (`modelopt_cuda_ext`, `_fp8`,
`_mx`) JIT-compile via `torch.utils.cpp_extension.load()` on first use —
~110–140s **each** in a fresh container, which is the dominant cost of
the `gpu_trtllm` job and the TRT-LLM example jobs. This caches them
across runs.
The logic lives in a reusable composite action,
**`.github/actions/cache-extensions`**, used by both `gpu_tests.yml` and
`_example_tests_runner.yml`:
- Sets a **literal in-container `TORCH_EXTENSIONS_DIR`**
(`/root/.cache/torch_extensions`). `${{ github.workspace }}` can't be
used — for `container:` jobs it resolves to the *host* path, which is
mounted elsewhere (`/__w`) inside the container, so torch and the cache
step would disagree on the location.
- Caches that dir with `actions/cache`, keyed on a caller-supplied **env
discriminator** (`rtxpro6000` + container image) plus a `hashFiles` of
the kernel/loader sources — so the cache busts on any kernel change and
is scoped per arch+image.
- On an **exact hit**, **backdates the kernel sources** below the cached
objects so ninja reuses them. (Touching the *objects* instead desyncs
ninja's `.ninja_deps`, which records each output's build-time mtime →
`stored deps info out of date` → rebuild.)
Also fixes the unused `runner` default in `_example_tests_runner.yml`
(`h100` → `rtxpro6000`) so it can't seed a wrong-arch cache.
### Usage
N/A — CI only. To reuse from another job:
```yaml
- uses: ./.github/actions/cache-extensions
with:
cache-key: rtxpro6000-${{ matrix.container_image }} # GPU arch + image
```
### Testing
Validated on `gpu_trtllm`: cache hit → `ninja: no work to do` →
`test_cuda_ext*` dropped from **113s / 108s / 139s → 2.8s / 0.03s /
0.03s** (~360s saved per run). Jobs that build no extension (e.g.
`gpu_vllm`) simply skip the save.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (CI-only; key busts on
source/image change)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update Changelog?: N/A (CI infrastructure)
- Did you get Claude approval on this PR?: ❌ (pending)
### Additional Information
- Single-arch assumption: callers pass `rtxpro6000` in `cache-key`; if
the runner fleet ever mixes GPU archs, update that prefix (the cache
path is not arch-specific).
- No explicit TTL: the key is content-addressed, and GitHub auto-evicts
caches unused for 7 days (+ 10 GB/repo LRU).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>