mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
main
159
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1576f7d4ad |
Increase diffusers example-test timeout to 60 minutes (#2591)
### What does this PR do? Type of change: Bug fix The diffusers example job can exhaust its 45-minute job budget while tests are still progressing. On the same commit, an [initial attempt timed out](https://github.com/NVIDIA/Model-Optimizer/actions/runs/36456149060/job/109049102711), while a [retry passed all 47 tests in 44m50s](https://github.com/NVIDIA/Model-Optimizer/actions/runs/36456149060/job/109115221419), leaving only 10 seconds of headroom. Increase the diffusers timeout to 60 minutes in `.github/workflows/example_tests.yml` to accommodate the workload and observed runtime variation. The other ONNX matrix entries retain their 45-minute timeout. This applies to both PR and nightly diffusers jobs. ### Usage N/A — CI configuration change. ### Testing - `pre-commit run --files .github/workflows/example_tests.yml` — passed all applicable hooks. - Parsed the caller and reusable workflow with `yaml.safe_load` and inspected the timeout input and consumer. - `git diff --check` — passed; reviewed the one-line diff. - GPU tests were not rerun locally. CI validation of the increased timeout is pending. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no copied code or new dependencies. - Did you write any new necessary tests?: N/A — one-line CI configuration change; validation described above. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — CI-only change. - Did you get Claude approval on this PR?: ❌ Not run; opening as a draft. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated automated checks to allow ONNX example tests up to 60 minutes. Other example tests retain their existing 45-minute limit. This change affects test execution time limits only; it does not change application features or behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> |
||
|
|
8990897c56 |
Increase Unit Test timeout
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
d0142c9dca |
Reuse a whole recipe via $import, deprecate recipe_type, and start the published-checkpoint backfill with two aliases (#2376)
### What does this PR do?
**Type of change:** New feature (recipe loading) + one bug fix
Two things, the second built on the first:
1. **A recipe can now reuse another recipe whole.** A top-level
`$import` brings in the imported recipe's entire body; keys given
alongside it override the imported ones. `metadata.recipe_type` becomes
optional and is deprecated along the way.
2. **The deprecated `recipe_type` is swept out of every shipped recipe,
and the checkpoint backfill starts with two published checkpoints
recorded as aliases** that reuse a portable recipe wholesale — the first
users of the alias mechanism — plus a fix to two existing Nemotron NVFP4
recipes.
#### Declaring what kind of recipe a file is
`load_recipe` read `metadata.recipe_type` out of the raw YAML *before*
resolving imports, because it needs the schema class to hand to
`load_config`. That made the field impossible to inherit, so a recipe
reusing another had to restate a line it could only have copied.
It is now optional, and the loader takes the first of these that
answers:
1. a `# modelopt-schema:` comment naming the recipe's schema class,
2. `metadata.recipe_type` — **deprecated**; still read and still
honoured, so a recipe outside this repo keeps working unchanged,
3. the recipe it delegates to via a top-level `$import`.
Whatever a recipe *does* state must be true, in both directions. A
schema comment contradicting a `recipe_type` is rejected, and so is a
recipe importing a different kind of recipe — that used to surface as
whatever pydantic made of, say, an `eagle` section spliced into a PTQ
schema. The concrete recipe classes carry a `RECIPE_TYPE` ClassVar as
the single source of truth.
Only a recipe that another file **imports** needs the schema comment —
that is what `$import` resolution requires to validate the payload. The
sweep here drops `metadata.recipe_type` from all 78 shipped recipes that
carried it and gives the imported ones a `# modelopt-schema:` comment
instead, so nothing in-tree depends on the deprecated field.
A directory recipe's `metadata.yml` resolves its kind the same way —
schema comment first, `recipe_type` as the fallback — it just has no
`$import` to delegate through, since a directory recipe has no body of
its own to hand off. (Follow-up commit, after this PR's initial review:
it originally still required `recipe_type` unconditionally, the one
place the deprecation didn't reach.)
#### Checkpoint aliases
Two checkpoints NVIDIA has published in quantized form use a scheme a
portable recipe already produces, with no checkpoint-specific deviation,
so each is recorded as a thin **alias** (top-level `$import`, overriding
only `metadata`) at its own model-hub path -- the *source* checkpoint's
path, not the published quantized one's:
-
**`models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast`**
delegates to `general/ptq/nvfp4_experts_only_mse-kv_fp8_cast` —
expert-only NVFP4 (MSE static weights, dynamic inputs) with an FP8 KV
cache in cast mode — published as `nvidia/Kimi-K2.6-NVFP4`.
-
**`models/Qwen/Qwen3.5-397B-A17B/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8`**
delegates to the `qwen3_5_moe` architecture recipe
`model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4
(MSE static weights) on the routed experts, ModelOpt-default FP8
elsewhere, FP8 KV cache — published as
`nvidia/Qwen3.5-397B-A17B-NVFP4-V2`.
(Follow-up commit, after this PR's initial review: the Qwen entry
originally lived at `models/nvidia/Qwen3.5-397B-A17B/` -- nvidia is the
*published* checkpoint's org, not Qwen3.5-397B-A17B's own. Moved to
match the source model's actual hub path, same as the Kimi-K2.6 entry
above.)
Editing the base recipe changes every alias that points at it; nothing
is duplicated.
#### One fix
- **The Nemotron-3 Super and Ultra NVFP4 recipes** quantized the MTP
block on the **Megatron-Core** path, where it is a live `model.mtp`
submodule their broad `*mixer.*` patterns matched into, contrary to
their own descriptions. They now disable `mtp.*` explicitly. Hugging
Face runs were unaffected — `NemotronHPreTrainedModel` sets
`_keys_to_ignore_on_load_unexpected = [r"mtp.*"]` and builds no MTP
module.
### Usage
A checkpoint alias resolves through `--recipe` to the recipe it
delegates to:
```bash
python examples/hf_ptq/hf_ptq.py \
--pyt_ckpt_path <checkpoint> \
--recipe models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast \
--export_path <output>
```
A recipe that reuses another whole — the shape the aliases use:
```yaml
imports:
base: general/ptq/nvfp4_experts_only_mse-kv_fp8_cast
$import: base
metadata:
description: What this checkpoint uses the base recipe for.
```
### Testing
- **`tests/unit/recipe/test_loader.py`** — 28 new cases covering
whole-recipe reuse with no `metadata` at all; kind resolution from each
of the three sources, from a delegation chain and from a `$import` list;
a delegation cycle failing with `ValueError` rather than recursing;
`peek_declared_schema` including a comment placed below the first YAML
line; `recipe_type` being optional, filled per class, and rejected when
it contradicts; a directory recipe resolving its kind from a schema
comment the same way, rejecting a comment/`recipe_type` disagreement,
and still requiring one or the other; and delegating across kinds being
an error.
- **`tests/unit/recipe/test_recipe_docs.py`** — the
model-specific-recipe check now also covers the two new alias folders,
which must be listed in `ptq.md` like every other
`models/<org>/<model_id>` entry.
- **Recipe validation** (`tools/precommit/check_modelopt_recipes.py`)
and **`pre-commit`** pass on the changed files. The full
`tests/unit/recipe/` suite is left to CI — a broken `transformer_engine`
in the local dev venv keeps the `mtq.quantize`-based cases from running
there.
Not covered: **numerics**. Nothing here asserts accuracy, or that
running one of these recipes reproduces a released checkpoint's weights.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `metadata.recipe_type` is
still read and honoured for recipes outside this repo, the schema
comments are inert for direct loads, and the loader change only relaxes
a check.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — two feature entries, one deprecation, and one bug fix under 0.48.0.
- Did you get Claude approval on this PR?: ❌ — not yet run.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Recipes can delegate configurations, support checkpoint aliases, and
apply local metadata overrides.
* Recipe types can be inferred from schema declarations or delegated
recipes, with stronger consistency validation.
* Added unquantized KV-cache options, layerwise export, broader operator
calibration, and new PTQ examples.
* Added checkpoint-specific recipes and MLflow experiment references.
* **Bug Fixes**
* Improved ONNX calibration, FSDP2 export, and fused-MoE quantization
handling.
* Nemotron-3 recipes keep MTP blocks in BF16.
* **Documentation**
* Expanded guidance for aliases, delegation, schema declarations, and
recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
d23030f91d |
[1/n] Adds skip-softmax calibration through the vLLM serving path (#1992)
### What does this PR do? Type of change: new feature Calibrates skip-softmax thresholds through the vLLM V1 execution path for FlashAttention and FlashInfer. Calibration measures the paged KV-cache path used at serving time, aggregates raw skipped/total tile counts across tensor-parallel head shards, fits separate prefill and decode curves, and exports the existing `sparse_attention_config` checkpoint schema. This uses raw counts rather than averaging per-rank sparsity ratios because TP ranks can contribute different tile populations; summing numerators and denominators before division preserves the global tile-weighted result. The vLLM adapter lives in `plugins/sparse_attn_calibration.py` rather than `SparseAttentionStatsManager`: the latter records module-local ratios for the HF calibration flow and has no aligned cross-process merge contract, while this path must merge per-sample raw counts from vLLM workers. Fitting and export still reuse `DynamicThresholdCalibrator` and the canonical conversion helpers so the model and checkpoint schema do not fork. Skip decisions depend on tile geometry. The common Triton launch boundary fixes the KV tile at 128 tokens and the prefill query tile at 128 tokens, including for direct kernel callers. Single-query decode can use a 16x128 compute tile without changing its skip decision. Measurement bypasses autotuning; serving still tunes warp and pipeline-stage counts while keeping the decision geometry fixed. ### Usage ```bash python examples/vllm_serve/calibrate_sparse_attn.py <CKPT> \ --prompts_file prompts.txt \ --target_sparse_ratio 0.7 \ --fit_logspace \ --tensor_parallel_size 4 \ --decode_tokens 32 \ --update_checkpoint_config ``` Calibration supports tensor parallelism and requires pipeline-parallel and data-parallel sizes of 1. It always writes `sparse_attention_config.json`; `--update_checkpoint_config` also merges the result into `<CKPT>/config.json`. ### Testing Latest revision `1e969cb380` (rebased onto main `02b58eb146`, 2026-09-17): - Calibration/count-fitting unit tests: **33 passed** (`test_sparse_attn_calibration.py` and `test_calibrator_fitting.py`). - Paged and contiguous calibration GPU suite: **33 passed** (`test_paged_calibrate.py` and `test_triton_fa_calibrate.py`), including NHD/HND equivalence, partial query tiles, decode counts, and malformed-cache rejection. Run with `CUDA_VISIBLE_DEVICES=1` on an RTX A6000; local GPU 0 was unavailable. - Calibration CLI tests: **21 passed** (`tests/examples/vllm_serve/test_calibrate_sparse_attn.py`). - `pre-commit run --files <four changed files>`: passed, including Ruff, mypy, and Bandit. - The new regression tests reproduced the skipped-counter truncation and missing cache-boundary checks before the fix. Calibration arithmetic and the 20-point threshold grid are unchanged. Historical validation from earlier revisions (not rerun end-to-end for this update): - `PYTHONPATH="$PWD" pytest -q tests/examples/vllm_serve/test_calibrate_sparse_attn.py tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_calibration.py` — 37 passed. - `PYTHONPATH="$PWD" pytest -q tests/gpu_vllm/torch/sparsity/attention_sparsity/test_vllm_calibration.py tests/gpu_vllm/torch/sparsity/attention_sparsity/test_sparse_attn_worker.py` — 65 passed, including kv-first, blocks-first, and packed FlashAttention cache layouts. - `PYTHONPATH="$PWD" pytest -q tests/gpu_vllm/torch/sparsity/attention_sparsity/test_vllm_runtime.py tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_config.py` — 33 passed. - `PYTHONPATH="$PWD" pytest -q tests/gpu/torch/kernels/sparsity/attention/test_paged_calibrate.py tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_calibrate.py tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_skip_softmax.py` — 31 passed, 1 skipped because the GPU lacks enough shared memory for the fp32 tile. - `pre-commit run --files <changed files>` — passed. - Historical end-to-end Nemotron 3 Ultra (GCP job `558552`), TP4, FA4, 48 RULER prompts, and 20 threshold trials: completed `0:0` with prefill `(a, b) = (9.9104, 10.8881)`, respectively +0.147% and -0.066% versus the matching 20-point reference `(9.8958, 10.8953)`. The supplied legacy fit `(14.47, 10.91)` used a different threshold grid; its `b` differs by only -0.201%, while `a` retains the known grid-weighting shift. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ Active skip-softmax fixes the calibrated decision geometry (serving still tunes warp/stage counts), and sparse-only vLLM installs fail fast for unsupported DCP, DBO/ubatching, speculative decoding, and FULL mixed-batch graphs instead of installing silently. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no copied code or new dependency. - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Pipeline parallelism is rejected during calibration because the current count-merging contract aligns records across tensor-parallel head shards, not across pipeline stages with disjoint attention layers. The unrelated HF padded-query behavior change was removed from this PR so it can be reviewed independently with its own compatibility test. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added vLLM skip-softmax calibration for paged attention, including prefill/decode support and checkpoint configuration generation. - Added Muse Glimmer AutoQuantize, Alpamayo QAD, streaming Kimi-K3 conversion, and NVFP4 activation headroom calibration recipes. - Added calibration statistics aggregation, phase-specific fitting, threshold validation, and preservation of existing sparse-attention settings. - **Bug Fixes** - Improved NVFP4 CPU/ONNX scale validation and clamping. - Added clearer handling for unsupported quantization, cache, CUDA graph, and engine configurations. - Standardized serving and calibration tile behavior. - **Documentation** - Expanded vLLM serving guidance, calibration instructions, compatibility requirements, and sparse-attention limitations. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kai Xu <kaix@nvidia.com> |
||
|
|
542012d4d5 |
Update CODEOWNERS (#2463)
Add modelopt-torch-kernels-codeowners <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Added code ownership coverage for the `modelopt/torch/kernels` directory. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
28dc117594 |
Add Day 0 Sub-agent Roles (#2006)
### What does this PR do? Type of change: new feature Adding sub-agent definition for Day 0 workflows. I'm adding subagents as an alternative, while we evaluate which is better. They are duplicated because Claude & Codex have different formats & different locations. Deriving from a shared vendor-neutral format at build-time is too complicated. ### Usage Tell your agent "Use the modelopt_model_quantizer agent to quantize the model" ### Testing Ran a trial using Qwen-2.5, Muse Glimmer, Qwen-2.8, and GLM-5.3-Flash. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information See design doc. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## New Features - Added specialized Model Optimizer agents for downloading, quantization, recipe search, evaluation, deployment, and performance benchmarking. - Added coordinated Codex and Claude Code access with validation requirements, artifact preservation, and concise handoffs. - Improved skill discovery across plugin and repository installations. ## Documentation - Clarified agent discovery, layout, and canonical editing locations. - Updated configuration guidance for supported agent definitions. ## Tests - Added synchronization checks for agent definitions and links. - Improved test compatibility for Python versions below 3.11. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chad Voegele <cvoegele@nvidia.com> |
||
|
|
40bed19b2d |
[chore]: release-branch housekeeping — auto-close stale uv.lock bump PRs, document cherry-pick labels (#2335)
### What does this PR do?
Type of change: Bug fix, documentation
Two pieces of release/CI housekeeping.
**1. Close stale uv.lock bump PRs
(`.github/workflows/bump_uv_lock.yml`)**
The weekly `Bump uv.lock` workflow opens a new PR every Monday but never
cleans up
the previous one, so un-merged bump PRs (and their `auto/bump-uv-lock-*`
branches)
pile up in the repo.
A new step runs just before the new PR is created: it lists open PRs
against the
same base whose head branch matches this workflow's own
`auto/bump-uv-lock-<base>-`
prefix, closes each with `gh pr close --delete-branch`, and leaves a
comment saying
it was superseded.
- Only PRs whose head branch carries the workflow's own prefix are
touched, so
unrelated PRs can't be closed.
- The step is gated on `steps.changes.outputs.changed == 'true'`, so a
pending bump
PR is only dropped when a replacement is actually about to be opened. If
`uv lock --upgrade` produces no changes, nothing is closed.
- It runs before the new branch is pushed, so the newly created PR can
never close
itself.
- No new permissions needed: the job already has `contents: write` and
`pull-requests: write`.
**2. Document the `cherry-pick-<X.Y.Z>` label (`CONTRIBUTING.md`)**
The "Submitting your code" section now tells contributors that a bug fix
which
should also land in an ongoing release branch needs the matching
`cherry-pick-<X.Y.Z>` label (e.g. `cherry-pick-0.47.0`), so the fix gets
picked into
the release branch and tested in the next release candidate. This
convention was
already in use but wasn't written down anywhere.
### Usage
N/A — CI workflow and contributor-docs change, no user-facing API.
### Testing
- `python -c "import yaml;
yaml.safe_load(open('.github/workflows/bump_uv_lock.yml'))"` passes.
- `pre-commit run` on both changed files passes (markdownlint included).
- The cleanup path itself only exercises on the next scheduled or
manually
dispatched run of the workflow, which has not happened yet.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — CI and docs only, not user facing.
- Did you get Claude approval on this PR?: ❌
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Chores**
- Automated dependency update pull requests are now superseded after a
replacement is created, keeping updates aligned with the latest lockfile
state.
- Superseded updates are cleaned up more safely, while active and
externally sourced pull requests are preserved.
- **Documentation**
- Added guidance for labeling bug-fix pull requests that should be
included in an ongoing release branch.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
51cc5dbade |
Make torch 2.14 the unit-test default and constrain it for the tensorrt example images (#2309)
### What does this PR do? Type of change: Bug fix (CI) + test coverage **Fixes `onnx (torch_onnx)` and `onnx (diffusers)`**, which have failed on every branch since `torch 2.14.0` was published to PyPI today (2026-09-02 13:42 UTC), and **adds torch 2.14 to the unit test matrix as the new default** so the next torch release is caught there rather than in an example job. ### Root cause Every test in those two jobs failed with: ``` RuntimeError: CUDNN_BACKEND_TENSOR_DESCRIPTOR cudnnFinalize failed ptrDesc->finalize() cudnn_status: CUDNN_STATUS_SUBLIBRARY_LOADING_FAILED ``` `nvcr.io/nvidia/tensorrt:26.05-py3` ships cuDNN **9.22** and has no preinstalled torch, so pip resolved the newest one — and torch 2.14 pins `nvidia-cudnn-cu13==9.24.0.43`. Loading 9.24 sublibraries against the image's 9.22 `libcudnn.so.9` is exactly what that status reports. | | last good run (08:55) | first failing run (13:34) | |---|---|---| | `torch` | 2.13.0 | **2.14.0** | | `nvidia-cudnn-cu13` | 9.20.0.48 | **9.24.0.43** | | image cuDNN | 9.22.0.52 | 9.22.0.52 | ### Why only these two jobs - The **nemo** and **pytorch** images have a preinstalled torch that already satisfies `torch>=2.8`, so pip never resolves a new one — confirmed from the megatron job log, where torch does not appear in `Successfully installed`. - **`tensorrt:26.05-py3` has no preinstalled torch**, so pip takes the newest from PyPI. - **`onnx (torch_trt)`** shares that image but passes throughout, because `torch-tensorrt<2.13` already holds torch below 2.14. ### The changes 1. **Constrain torch only where the incompatibility is.** `PIP_CONSTRAINT=torch<2.14` in the example runner, applied when the job's image is a `tensorrt` one. It also covers the `examples/*/requirements.txt` loop in the same shell, which matters because `nemo_automodel` pulls torch in too. Not pinned in `pyproject.toml`: torch 2.14 is fine anywhere its own bundled cuDNN is the one loaded, so that would constrain users to work around one pinned image. 2. **Test torch 2.14.** `torch_214` added to `TORCH_VERSIONS` (`torchvision~=0.29.0`) and promoted to the unit-test default across the supported Python versions, with 2.13 demoted to the back-compat row. `release.yml`'s basic unit test moves to the same default (it was still on 2.12). Nothing exercised 2.14 before — which is why a torch release reached us through an example job instead of a unit test. ### Testing - `actionlint` and YAML/TOML parse clean; pre-commit clean. - Verified by this PR's own jobs: `onnx (torch_onnx)` and `onnx (diffusers)` reproduce the failure on `main` right now, and the new `unit-3.12(torch_214, tf_latest)` job is the first run of ModelOpt against torch 2.14. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — CI-only; no source or package metadata change - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new dependency - Did you write any new necessary tests?: ✅ — torch 2.14 added to the unit test matrix - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — internal CI, not user-facing - Did you get Claude approval on this PR?: ❌ — not yet requested --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
411d072a2e |
Speed up megatron_bridge example tests by ~6x on a single GPU (#2296)
### What does this PR do? Type of change: Test infrastructure / CI time `tests/examples/megatron_bridge` spends most of its time importing Python, not testing. Each step of a test spawns `torchrun`, and the new process spends **~25s importing torch/megatron/modelopt** before doing any work. A single `test_qad` run pays that **six** times — three steps, plus a spawned child per distributed checkpoint save, because Megatron-Core's async writer uses `mp_mode="spawn"` and spawn re-imports `__main__`. Profiled with phase timers in the example scripts: | | share of `test_qad[qwen3]` | |---|---| | Python imports (6 process launches × ~26s) | **~76%** | | actual compute (`mtq.quantize` 4.2s, model build 0.24s, export 0.06s) | ~5s | `run_example_command` now dispatches each step internally instead of shelling out: - **single-rank steps** run directly in the pytest process, driving the script's own `get_args()` + `main()` — no new interpreter, no re-import; - **multi-rank steps** drive `torch.distributed.run.main()` in-process with patched `sys.argv`, the same pattern Megatron-Bridge uses in its own functional tests. ### Results | suite | before | after | |---|---|---| | `tests/examples/megatron_bridge`, **1 GPU** | **21m27** | **4m03 - 6m17** | | `tests/examples/megatron_bridge`, **2 GPU** | 26m10 | 25m03 | Both figures are on current `main` (17 tests). The 2-GPU number is up from 20m41 before merging #2276, which added a test and made one previously single-rank step multi-rank. The 1-GPU figure is a range, not a best case. Across ten runs on a verified-idle box the suite lands at ~4m most of the time and at ~6m otherwise, always with the same result (17 passed). Per-test durations show the entire spread is one test: `test_qad[qwen3]` runs at ~15s or at ~149s. It is 25s in isolation, 15s after `test_distill.py` and ~25s after `test_prune_minitron.py`, so it needs the full sequence and does not reproduce on demand — three attempts to catch it under instrumentation all landed on fast runs. In those, CUDA state immediately before it is 46 MiB allocated / 68 MiB reserved / 7 segments / 22 MiB inactive-split, and the preceding test's 498/984 MiB is fully reclaimed, so a fragmented allocator is measured *not* to be the cause in the fast path at least. Left documented rather than guessed at: correctness is unaffected across every run, and the worst case sits inside the 30-minute PR budget (CI `Run tests` 902s). The single-GPU path is the big win, and it is the one the per-PR runner uses — that job now finishes in **8 minutes** in CI. Multi-rank steps still launch worker processes that re-import, so the 2-GPU nightly improves far less. This also fixes the timeouts under coverage. With `--cov` (how CI runs it), on the same three tests: in-process **3 passed in 1m15**, subprocess **3 failed on `Timeout (>360.0s)` in 18m57**. ### CI timeout The 2-GPU nightly runs every test multi-GPU and measured **58 minutes against a 60-minute cap** — too close to be reliable. `timeout_minutes` is now ref-conditional, mirroring the `runner` line directly below it: **30 minutes on PRs** (single-GPU, ~8 min) and **75 on nightly**. Keeping the nightly at full multi-GPU coverage is deliberate. Making individual tests single-rank cut it to ~7 minutes, but it gives up the parallel-path coverage that is the whole point of the 2-GPU job, and it surfaced a real fragility: `test_prune_minitron[nemotron_h]` fails with *"No scores collected for importance estimation"* when it runs single-rank after the full distill file. It passes alone and after any single preceding test — multi-rank tests are immune because `torchrun` gives them fresh worker processes. Nightly is the right place to spend the wall-clock. ### What is and isn't covered Each script's real `get_args()` still runs, so CLI flags, defaults and recipe-string resolution stay covered. Not covered for single-rank steps: the `torchrun` invocation itself and the `__main__` block (`dist.setup()` / `dist.abort()`). Multi-rank steps still go through the real launcher. **No test file changes.** The tests still read as "launch this torchrun command" and their assertions are untouched. ### Keeping it that way There is no toggle and no fallback. Every step in this suite must be `torchrun --nproc_per_node=<int> <script>.py` with the script exposing `get_args()` + `main()`, and `run_example_step` raises otherwise — it returns `str`, not `str | None`, so a step cannot quietly become a subprocess. That matters because a silent fallback still *passes*, just ~6x slower, so a new script or test could cost the suite its speed-up with nothing to show for it. Both guards verified by breaking them on purpose, each failing in ~1.4s rather than burning a run: | broken convention | result | |---|---| | `--nproc_per_node=gpu` | `AssertionError: --nproc_per_node must be a plain integer: [...]` | | step invoking `generate_vllm.py` (no `get_args`) | `AssertionError: generate_vllm.py must define get_args() and main(args)` | ### Layout The runner lives in `tests/_test_utils/examples/megatron_example_runner.py`, next to the `run_command.py` it plugs into and the other per-example helpers. It is deliberately not under `tests/_test_utils/torch/megatron/`: both files there import megatron at module top, whereas this one must not, since importing `megatron.bridge` would initialise CUDA in the pytest process and hold a context on device 0 for the whole session. ### Isolation Sharing one interpreter means anything global has to be put back between steps, or one failing test cascades into the next. Each of these was previously cleaned up by `torchrun` simply exiting: - **`NVTE_*`** — Transformer-Engine records its chosen attention backend in the environment, so a Mamba hybrid failed after an attention model ran. The environment is restored wholesale rather than by naming variables. - **Allocator** — `empty_cache()` frees nothing while a finished step's model is still reachable; a later test ran **9x slower** (162s vs 18s) against a fragmented allocator until `gc.collect()` was added first. - **Parallel state and the rerun state machine** — two separate singletons; `destroy_model_parallel()` does not touch the latter. - **Signal handlers** — `PContext.start()` installs its own `SIGTERM`/`SIGINT`/`SIGHUP`/`SIGQUIT` handlers and never restores them. With a subprocess launcher, process exit did that for us; in-process they are saved and put back, or pytest's Ctrl-C and CI cancellation would break for the rest of the session. Verified rather than assumed: injecting a failure mid-test (after a model was built and parallel state left live) gives **1 failed, 2 passed**, with the surviving tests at full speed. ### Coverage Coverage of the exercised code **improves**. In subprocess mode the child imports modelopt as `site-packages/modelopt/...` while pytest measures `modelopt/...`, so the data never merges — which is also why the subprocess report showed exactly double the statement count. | module | subprocess | in-process | |---|---|---| | `unified_export_megatron.py` | 8% | **43%** | | `mcore_custom.py` | 34% | **44%** | ### Testing All in `nvcr.io/nvidia/nemo:26.08` on 2x RTX 6000 Ada, with per-test caps enforced. | run | result | |---|---| | `tests/examples/megatron_bridge`, 1 GPU | 17 passed — 4m03 (6m17 worst of 10 runs) | | same, subprocess baseline | 15 passed, 1 skipped — 21m27 | | `tests/examples/megatron_bridge`, 2 GPU | 17 passed — 25m03 | | CI 1-GPU example job (`megatron / run-test`) | passed — `Run tests` 902s, job ~20m (30m cap) | | CI 2-GPU nightly | passed — `Run tests` 3162s, job 58m (75m cap) | | cascade check (injected mid-test failure) | 1 failed, 2 passed, survivors at full speed | | pre-commit | clean | ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — test-only; no source or public API changes - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new dependencies (`pytest-forked` was evaluated and rejected: `import megatron.bridge` initialises CUDA, and CUDA cannot be re-initialised in a forked child) - Did you write any new necessary tests?: N/A — this changes how existing tests are executed - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — internal test infrastructure, not user-facing - Did you get Claude approval on this PR?: ✅ — reviewed by Claude and CodeRabbit, all threads addressed --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
61757c9781 |
Support quantized Qwen3-VL / Qwen3.5-VL (dense + MoE) export from Megatron-Bridge and verify exported checkpoints (#2276)
### What does this PR do?
Type of change: Bug fix + new feature
Enables quantized **Qwen3-VL** and **Qwen3.5-VL** (dense and MoE) →
unified HuggingFace export from Megatron-Bridge, and fixes the bugs
found along the way (ten from testing, plus a further round from
review). Most of them produced a valid-looking checkpoint and a green
test run, so the PR also makes the export path verify its own output.
Review is easiest commit-by-commit — each of the eleven commits is
self-contained and independently green.
#### Two blockers
1. **The exporter rejected the Megatron-Bridge VLM wrapper.**
`GPTModelExporter` only unwrapped MCore's `LLaVAModel`, so
`Qwen3VLModel` raised `ValueError: Input to GPTModelExport must be a
megatron.core.models.GPTModel!`. It now unwraps any wrapper exposing
`.language_model`.
2. **A VLM QAD checkpoint couldn't be loaded back.** `distill.py` passes
`distill_submodule="language_model"`, so the checkpoint holds only the
language model and the load died on `KeyError:
vision_model.patch_embed.proj.weight`. The loader now reads the
checkpoint metadata and targets `.language_model` when there are no
vision weights.
#### Four silent-corruption bugs
3. **VLM QAD discarded all ModelOpt state** (shipped in 0.46).
`ModeloptStateManager` requires state on the **root** of whatever gets
checkpointed. `quantize.py` quantizes the VLM root, so PTQ anchors it
there — but QAD checkpoints only `language_model`, orphaning it. The
saved `modelopt_state_dict` was literally `[]`; the `*_quantizer._amax`
tensors were still present but got dropped on load
(`dist_ckpt_strictness="assume_ok_unexpected"`), and the export came out
plain BF16 with no `hf_quant_config.json`.
4. **Fused grouped-GEMM MoE experts were omitted entirely.** The MoE
dispatch had no `else`, so an architecture without an
`experts.linear_fc1` rule exported *zero routed experts*. This hit
**`Qwen3MoeForCausalLM`** — a registered, supported architecture with no
export test — not just VLMs. A tiny Qwen3-MoE exported 37 of 45 tensors,
exit 0, no warning.
5. **Qwen3.5's GatedDeltaNet output norm was off by exactly 1.0.**
Megatron stores that gamma zero-centered, HF centers it on 1. Correct
names, correct shapes, wrong values — invisible to any structural check.
Megatron-Bridge's importer confirms the convention
(`RMSNorm2ZeroCenteredRMSNormMapping`).
6. **The disabled-quantizer patterns silently no-op on Megatron paths.**
They are written against HuggingFace module names. `*mixer.conv1d*`
matches only because MCore and HF happen to agree on "mixer" for Mamba;
`*linear_attn.conv1d*` never matched (Megatron calls it
`self_attention.conv1d`), so the conv1d was calibrated.
`*linear_attn.in_proj_a/b*` **cannot** match at all — Megatron fuses all
six GDN sections behind one quantizer — so the alpha/beta gates the
recipe wants in BF16 were exported in FP8.
#### Four more bugs, found only by running real checkpoints
The tiny fixtures could not reach these; each came from a real model or
a real quant format.
7. **Routed experts were written in a layout no real Qwen3.5 checkpoint
uses.** Real Qwen3.5 stores experts packed as `[num_experts, out, in]`;
the mapping emitted per-expert names, so every routed expert was
dropped. The fixture actively hid this: transformers *unpacks* experts
on `save_pretrained`, so the saved reference agreed with the wrong
output. Fixed with a `transpose` kwarg on `_pack_name_remapping` plus a
`GroupedMLPPacking` rule, so fused `TEGroupedMLP` reaches the same
packed tensors — which is also what lets Qwen3.5 keep grouped GEMM
(**22.1 GB/GPU vs 38.9 GB/GPU** on a 20-layer, 256-expert model).
8. **`_grouped_mlp_packing` was broken for NVFP4.** It max-merged
`weight_scale`, but NVFP4 needs each expert's per-block scales *stacked*
with only the global `weight_scale_2` merged; it also dequantized packed
`uint8` against per-block scales, and passed `block_size=None`.
`weight_scale_2` is never populated in an FP8 run, so the whole branch
was dead code under FP8-only testing. `_grouped_mlp_slicing` gained
`quantize=False` so packing can quantize once over the stack, matching
`_pack_name_remapping`.
9. **`_mtp_prefix` corrupted every VLM's MTP tensor names.** It did
`prefix.replace("model", "mtp")` uncounted, so
`model.language_model.layers.{}` became `mtp.language_mtp.layers.0.*` —
tensors present and correctly valued, under names nothing loads.
LLM-only prefixes contain one occurrence, so this was invisible until a
VLM with MTP was exported.
10. **`load_multimodal_components` rejected HF repo ids.** `quantize.py
--hf_model_name_or_path Qwen/Qwen3.5-0.8B` worked, but the documented
export step failed with *"It should be a directory"*. Its sibling in the
same file already resolved repo ids via `snapshot_download`; now it does
too. This affected **every** VLM export.
`Qwen3_5ForConditionalGeneration` (dense Qwen3.5-VL) is now registered
for export and vision passthrough, which bugs 9 and 10 were blocking.
#### New: Qwen3.5-VL
`GatedDeltaNetSlicing` splits Megatron's fused `in_proj` (`[query, key,
value, z, beta, alpha]`) into HF's `in_proj_qkv` / `_z` / `_b` / `_a`,
taking sizes from the module's own `in_proj_split_sections` so TP
sharding falls out. Widening coverage to Qwen3.5's *gated
full-attention* layers then exposed a further split bug: gated attention
packs a per-head output gate beside each query head, so `_qkv_slicing`
split 192 rows as 96/48/48 instead of 128/32/32. It now derives the
group stride from `config.attention_output_gate`, matching
Megatron-Bridge's `split_qkv_weights`. The non-gated path is unchanged.
#### New: the export path verifies itself
- `assert_exported_checkpoint_matches` compares an exported checkpoint
against the model it came from — key set, shapes (accounting for NVFP4
`uint8` packing), safetensors index consistency, and values — replacing
existence-only assertions in all three export tests.
- `GPTModelExporter.save_pretrained` now raises if the export dropped
tensors the source checkpoint has, so *user* runs on architectures CI
never sees are protected too, not just tiny models.
- Loading a checkpoint whose quantizer tensors have no restorable state
now raises instead of silently loading unquantized.
- `assert_has_modelopt_state` replaces `rglob("modelopt_state")`, which
passes on an empty state; `assert_no_quantizers_matching` fails on
future HF↔Megatron name drift.
The mapping is also table-driven now: vision-tower prefixes live in
`all_mcore_hf_vision_passthrough_mapping` and
`with_language_model_prefix` is shared, so adding a VLM no longer means
editing `unified_export_megatron.py`. Five call sites that answered "is
this a VLM" three different ways now share `get_language_model` /
`is_vlm_config`.
### Usage
```bash
# Dense VLM (Qwen3-VL) -- no extra flags
torchrun --nproc_per_node 2 quantize.py \
--hf_model_name_or_path Qwen/Qwen3-VL-8B-Instruct \
--quant_cfg nvfp4 --tp_size 2 \
--export_megatron_path /tmp/Qwen3-VL-8B-NVFP4-megatron
torchrun --nproc_per_node 2 export_quantized_megatron_to_hf.py \
--hf_model_name_or_path Qwen/Qwen3-VL-8B-Instruct \
--megatron_path /tmp/Qwen3-VL-8B-NVFP4-megatron \
--pp_size 2 --export_unified_hf_path /tmp/Qwen3-VL-8B-NVFP4-hf
# Gated MoE (Qwen3.5-VL, Qwen3-MoE) -- no extra flags either. The scripts derive the
# expert layout from the model config, so quantize / distill / export all agree.
# --no_moe_grouped_gemm forces SequentialMLP if you want it explicitly.
```
### Testing
All in `nvcr.io/nvidia/nemo:26.08` on 2x RTX 6000 Ada.
| Suite | Result | Time |
|---|---|---|
| `tests/examples/megatron_bridge/` (full) | 18 passed | 27m58 |
| `tests/gpu_megatron/torch/export/` | 38 passed | 2m13 |
| `tests/unit/torch/export/` | 186 passed | 1.5s |
| pre-commit (ruff, ruff format, mypy, bandit) | clean | — |
| `tests/examples/megatron_bridge/test_quantize_export.py` on **2 GPUs**
(`pp_size=2`) | 3 passed | 5m |
The export leg of `test_quantize_and_export` now scales with `num_gpus`
like its quantize leg
already did. Previously it was hardcoded to one process, so the
collective checkpoint load ran at
PP=1 on both the 1-GPU PR runner and the 2-GPU nightly — which is how a
guard that raised on only
some pipeline stages (and therefore hung the job) reached review. The
dense `qwen3` case was dropped
in exchange: `qwen3_moe` already covers the non-VLM script path,
`qwen3vl` covers a dense decoder,
and that case was the one exceeding the 300s cap in CI.
#### Model coverage
`tests/gpu_megatron` runs in-process and is cheap, so it owns
per-architecture **mapping**
correctness. The example tests spawn `torchrun` per step and are ~50x
slower per case, so they
cover **script wiring** only — CLI flags, recipe resolution, and
checkpoint hand-off between steps.
| Suite | Models |
|---|---|
| `test_unified_export_megatron` | llama, nemotron, nemotron_h, qwen3vl,
qwen3_moe, qwen3_5_moe_vl x {none, FP8, NVFP4, +/-KV} x {grouped GEMM,
SequentialMLP} + eagle / medusa / MTP (29 params) |
| `test_megatron_importer` | nemotron_h, llama export->import round-trip
|
| `test_moe_layout_choice` | per-architecture grouped-GEMM exportability
(6 architectures) |
| `test_distill_megatron` | KD loss mechanics |
| Model | prune | quantize+export | QAD | distill+export |
|---|:--:|:--:|:--:|:--:|
| qwen3 | Y | Y | Y | Y |
| qwen3_moe | - | **Y (new)** | - | - |
| qwen3vl | - | **Y (moved from QAD)** | - | - |
| nemotron_h | Y | **Y (new)** | - | - |
| qwen3_5_vl | - | - | - | Y |
| qwen3_5_moe_vl | Y | **Y (new, both expert layouts)** | Y | - |
| deepseek_v3 | Y | - | - | - |
| gemma3vl | Y | - | ~~manual~~ removed | - |
QAD's unique property is that ModelOpt state survives distillation,
which needs one LLM and one
VLM rather than one case per architecture. Moving the rest to
quantize+export drops a `torchrun`
launch each: QAD went from 3 CI cases to 2 while quantize+export went
from 1 to 4, adding two
architectures for about a minute.
#### Real-model validation
Tiny fixtures cannot catch layout or scale bugs that only appear at real
dimensions, so the export
path was run end-to-end on released checkpoints. This is where bugs 7-10
came from.
| Model | Run | Result |
|---|---|---|
| Nemotron-3.5-Lightning-30B-A3B | NVFP4 4o6 PTQ → export → MMLU |
**0.7825 ± 0.0105** (gate 0.75) |
| Nemotron-3.5-Lightning-30B-A3B | Minitron pruning | 22.28B/3.00B
active, **0.5944** (gate 0.58) |
| Qwen3.5-0.8B (dense VLM) | FP8 PTQ → export → MMLU | BF16 0.4895 →
**0.4832** (±0.0127) |
| Qwen3.5-35B-A3B, half-depth (20 layers, 256 experts) | FP8 + NVFP4 PTQ
→ export | keys + shapes + **values** match reference |
| Qwen3.5-35B-A3B, full | FP8 PTQ | OOM on 2x48GB (see below) |
The half-depth model keeps real weights, real dims and all 256 experts.
Both expert layouts produce
identical key sets, and all exports pass
`assert_exported_checkpoint_matches(..., check_values=True)`
— every tensor, including all 20 x 256 experts, dequantizes to within
tolerance of the BF16
reference, so a transposed or mis-ordered expert stack would fail. NVFP4
lands in the correct packed
layout (`gate_up_proj [256, 1024, 1024]` U8, `weight_scale [256, 1024,
128]` E4M3,
`weight_scale_2 []` F32). Its *accuracy* is not meaningful — truncating
to 20 of 40 layers leaves a
chance-level model (BF16 0.2322, FP8 0.2538) — so it validates
correctness, not quality.
**Re-validated on the final code.** The numbers above were first taken
mid-review; since then the
NVFP4 block-scale merge changed on both packed paths, the vision-tower
download became two-stage,
and an expert-layout load guard was added. Both gating runs were
therefore repeated end to end:
Nemotron went 0.7748 → **0.7825 ± 0.0105** and Qwen3.5-0.8B went 0.4678
→ **0.4832 ± 0.0127**, with
the rest of the Nemotron pipeline reproducing exactly (3519 quantizers,
69GB checkpoint, 21GB
export). Both deltas are inside their own stderr, so the claim is that
the rework costs no accuracy
— not that it improved it. The Nemotron export also runs at `--pp_size
2`, exercising the new
collective layout guard on a real 30B MoE across pipeline stages.
Two limitations worth stating plainly:
- **No quantized accuracy number for a full-size MoE.** The full 35B
OOMs at 47.37 GiB while
*constructing* the model on 2x48GB, with grouped GEMM already enabled,
so no calibration knob
helps. Needs more GPUs than this setup has.
- **vLLM cannot yet serve packed FP8 Qwen3.5 experts.** `vllm
0.24.1.dev0` builds its fused expert
mapping weight-only, rewriting `experts.down_proj_input_scale` to
`w2_weight_input_scale` while the
parameter it registers is `w2_input_scale`. This is upstream and
independent of how the checkpoint
is produced — both of our export paths fail it identically. The 0.8B
numbers above are unaffected
(dense), and the packed exports are verified against the reference
checkpoint instead.
#### Guard verification
Each new guard was made to fire, not just to compile:
| Guard | Verification |
|---|---|
| Export self-check | Disabled the MoE guard, re-exported Qwen3-MoE -
independently reported all 24 dropped tensors. No false positives across
llama, nemotron, qwen3, qwen3-moe, qwen3vl, qwen3.5-vl, deepseek_v3
incl. eagle / medusa / MTP |
| Dropped-state raise | Deleted `modelopt_state` from a checkpoint with
50 quantizer tensors - raised instead of loading unquantized |
| NVFP4 value check | Flipped a `q_proj` - failed at `max_rel_err=1.74`
against a 0.3 threshold |
| Zero-centered gamma | Reproduced the off-by-1.0 on a good export -
caught as "not bit-exact" |
| Exclusion guard | Asserts no calibrated quantizer matches `conv1d` /
`mlp.router` / `output_layer` |
Exported artifacts are validated, not just their existence: 0 missing
keys vs reference, vision
tower bitwise-identical, dequantized weights within FP8 E4M3 error
(<=4.6%). The
`in_proj_a`/`in_proj_b` check is load-bearing - swapped alpha/beta would
still match on shape but
show ~100% error.
Also ran a tiny-Qwen3 **LLM** control through both steps to confirm the
exporter changes are a
no-op off the VLM path.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — the scripts now derive the
MoE expert layout from the model config, building SequentialMLP only for
architectures with no `experts.linear_fc1` rule, and the exporter raises
rather than dropping experts it has no rule for. Those runs previously
"succeeded" while writing a checkpoint containing no expert weights, so
no working behaviour is removed. `--no_moe_grouped_gemm` forces
SequentialMLP explicitly.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅ — approved (round 10: 0
CRITICAL, 0 IMPORTANT, 0 new suggestions); CodeRabbit approved earlier
### Additional Information
**MoE expert layout is now chosen automatically.** Only Nemotron-H can
export fused grouped-GEMM experts, so every other MoE architecture would
otherwise need `--no_moe_grouped_gemm` on all four scripts or hit a wall
at export. The scripts derive the layout from the model config — grouped
GEMM unless it would not be exportable — so they agree without threading
a flag. This changes MoE activation scales from one shared scale to
per-expert for the affected architectures.
Known gaps, unchanged by this PR:
- **Gated MoE still cannot use fused grouped GEMM.**
`_grouped_mlp_slicing` emits one weight per expert with no gate/up split
— its only prior caller, Nemotron-H, is non-gated, so every other MoE
architecture is built as `SequentialMLP` (see below). Adding that split
would restore the faster layout, but it needs a deliberate call on
activation-scale semantics: grouped GEMM keeps **one shared** activation
scale across experts while `SequentialMLP` has **per-expert** scales, so
the two are not numerically equivalent. It also needs EP>1 coverage.
- **Qwen3.5's alpha/beta gates share Megatron's fused `in_proj`
quantizer,** so they can only be kept in BF16 at export, not excluded by
name. Full fidelity needs per-section quantizers on the fused
projection.
- **Anchoring ModelOpt state on `.language_model`** (which would let
`quantize.py` quantize the language model directly and drop its
name-based non-LM disabling) needs a coordinated Megatron-Bridge change:
`save_sharded_modelopt_state` is ModelOpt code, but the restore the
Bridge path uses is Bridge's own and unconditionally restores onto the
root.
- **Gemma3-VL** remains Megatron-checkpoint only (`OMNIML-5366`).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added Muse Glimmer AutoQuantize and Alpamayo QAD workflows.
* Added streaming Kimi-K3 conversion and NVFP4 activation headroom
calibration.
* Added SFT-masked distillation for Megatron-Bridge.
* Added unified Hugging Face export for quantized Qwen3-VL and
Qwen3.5-VL checkpoints.
* MoE expert layouts are selected automatically, with an option to force
sequential experts.
* **Bug Fixes**
* Improved export validation for tensor coverage, MoE mappings,
quantizer state, and NVFP4 scales.
* Fixed Qwen3.5-VL GatedDeltaNet export handling.
* Preserved visual-model weights exactly during export.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
d73278808b |
Bump nemo container requirement to 26.08 for MBridge examples (#2257)
### What does this PR do? Type of change: Bug fix Bumps the Megatron-Bridge examples, tests and launcher configs to `nemo:26.08` and removes the version-gated fallbacks they carried, plus the fixes needed to make the suites green on that container. **26.08 bump and shim removal** - Examples, CI workflows, `noxfile.py` and the `mbridge_*` launcher configs move to `nemo:26.08`. - `examples/megatron_bridge/_distillation_provider.py` is deleted — 26.08's Megatron-Bridge ships `convert_to_distillation_provider(..., distill_submodule=...)` natively, so `distill.py` imports it directly. - `prune_minitron.py` drops the `AutoBridge.from_hf_config` / config-only-export probing; `--no_moe_grouped_gemm` is no longer needed in the MoE pruning tests, and the Qwen3.5-MoE `skipif` is gone (native MoE expert mappings are in 26.08). - `_DynamicMambaMixer` targets only the raw `conv1d_weight` / `conv1d_bias` parameters that replaced the `conv1d` module in Megatron-Core. **MambaModel / MambaModelProvider removal** Megatron-Core has shipped `HybridModel` since 26.06 and `MambaModel` is a deprecated subclass that shares its `forward`, so `DMRegistry` resolves those instances to the `HybridModel` registration and the separate entry is redundant. Same for `MambaModelProvider` vs `HybridModelProvider` on the bridge side. `MambaMixer` / `MambaLayer` / `ExtendedRMSNorm` are untouched — the layers still exist. The deprecated `get_te_mamba_stack_spec` is removed; use `get_te_hybrid_stack_spec`. **Bug fix: compressed output_layer extra state** `mtq.compress` converts even a *disabled* `output_layer` into a `RealQuantLinear` (its weight is left uncompressed, since `pack_real_quantize_weight` skips disabled quantizers). The guard added in #2112 excluded `RealQuantLinear` by class, so the layer emitted extra state and every worker died in `GPTModel.sharded_state_dict`: ``` RuntimeError: Boolean value of Tensor with more than one value is ambiguous megatron/core/models/gpt/gpt_model.py:896, in sharded_state_dict output_extra_state and output_extra_state.data ``` The guard now keys off whether the weight was actually compressed (`QTensorWrapper`) instead of the class. This took out all 12 `test_homogeneous_compressed_sharded_state_dict` params, and the crashed workers poisoned the pool, which surfaced as unrelated timeouts and NCCL errors in `test_layer_sync_moe_local_experts_amax`, `test_kv_cache_quant`, `test_kv_cache_amax_sync`, `test_convert_mcore_te_gpt_model` and `test_homogeneous_sharded_state_dict_te_spec` — 21 tests in total. The e2e coverage is `skip_flaky_on_blackwell`, so CI never ran it; `test_output_layer_extra_state_empty_when_nothing_quantized` now asserts the contract directly and is not skipped. **Checkpoint import entry point** 26.08 replaced `examples/conversion/convert_checkpoints.py` with `scripts/conversion/convert.sh`, so `tools/launcher/common/megatron_bridge/import/import.sh` and the three README snippets are retargeted. `import.sh` uses the distributed GPU backend with `GPUS_PER_NODE` / `TP` / `PP` / `EP` knobs. **Megatron-LM on nemo:26.06** keeps working: `_get_mamba_conv1d` still dispatches between the `conv1d` module (26.06 and earlier) and the raw parameters (26.08+), so `import_mcore_gpt_from_hf` / `export_mcore_gpt_to_hf` handle NemotronH on both. Only the Megatron-Bridge examples and Minitron pruning of Mamba/hybrid models require 26.08. **Test consolidation** `test_export_distilled_megatron_to_hf.py` is merged into `test_distill.py`: `test_distill_llm` becomes `test_distill_llm_hf_export` and covers the standalone `--export_iterations all` run on the checkpoints it already produces, saving one full distillation (~185 s of CI time). The two mamba-named gpu test files are renamed to `hybrid`. ### Usage ```bash # HF -> Megatron import, via Megatron-Bridge's 26.08 conversion entry point bash /opt/Megatron-Bridge/scripts/conversion/convert.sh import \ --executor local \ --device gpu \ --gpus-per-node 8 \ --hf-model Qwen/Qwen3-8B \ --megatron-path /tmp/Qwen3-8B-megatron ``` ### Testing All on `nvcr.io/nvidia/nemo:26.08`, 2x RTX 6000 Ada, no timeout overrides: - `tests/examples/megatron_bridge`: 16 passed, 1 skipped (28m14s). The skip is the `gemma3vl` QAD param, now `@pytest.mark.manual` since `qwen3_5_moe_vl` covers the VLM QAD path. - `tests/gpu_megatron` (`_extensions`, `distill`, `export`, `opt`, `peft`, `sparsity`, `speculative`, `utils`): 61 passed, 5 xpassed. - `tests/gpu_megatron/torch/export` re-run after the conv1d dispatch change: 27 passed. - The 21 previously failing/hanging quantization tests: 21 passed (12 + 9). - `tests/gpu_megatron/torch/{nas,prune}`: verified separately. `import.sh` equivalence on a toy `qwen3_moe`, comparing all 12 weight tensors after flattening each dist checkpoint with `dcp_to_torch_save` — the GPU backend at 1 GPU, `--tp 2`, `--pp 2`, `--ep 2`, and `import.sh` end-to-end (`GPUS_PER_NODE=2 EP=2`) are all byte-identical to `--device cpu`. `nemo:26.06` compatibility was checked directly in that image: `megatron.core.models.hybrid.HybridModel`, the modelopt hybrid spec and `hybrid_layer_pattern` are all present, while `megatron.bridge.models.hybrid` and the bridge's `distill_submodule` are not. The NemotronH round-trip test failed there before the conv1d dispatch was restored and the dispatch is back in place; per project convention the suites themselves only run on 26.08. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ⚠️ Megatron-Bridge examples plus Minitron pruning of Mamba/hybrid models now require `nemo:26.08`. Megatron-LM quantization and checkpoint export still run on `nemo:26.06`. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — `test_output_layer_extra_state_empty_when_nothing_quantized` for the compress fix; existing tests extended for the merged export coverage. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added guidance for importing Hugging Face checkpoints into Megatron distributed format. * Expanded distillation workflows to export selected or all checkpoint iterations. * **Improvements** * Expanded Hybrid model support across Megatron workflows. * Updated distributed import tooling with GPU and parallelism options. * Updated supported environments and examples to NVIDIA NeMo 26.08. * **Bug Fixes** * Corrected output-layer quantization state handling when quantization is disabled. * **Documentation** * Added compatibility guidance for current and legacy NeMo containers. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
195a5af430 |
Add modelopt-agents-codeowners as owner of plugins/ (#2212)
### What does this PR do? Type of change: Documentation `plugins/` holds the installable ModelOpt agent plugin — the canonical skill tree at `plugins/modelopt/skills/`, which `.agents/skills` and `.claude/skills` expose via relative symlinks. It currently has no CODEOWNERS rule, so it falls through to the default `* @NVIDIA/modelopt-devs` owner. This routes it to `@NVIDIA/modelopt-agents-codeowners` instead. ``` # Agent plugin (skills, agent config) /plugins @NVIDIA/modelopt-agents-codeowners ``` The path is root-anchored (`/plugins`) to match the style of the `/examples` block above it. ### Usage N/A — no API or flag change. ### Testing No test surface. Verified the rule resolves as intended against the existing file: `/plugins` is the last (and only) match for paths under `plugins/`, so it wins over the default `*` rule; it does not overlap any other entry. Effective ownership of `plugins/` on this branch is `@NVIDIA/modelopt-agents-codeowners`. GitHub validates the team handle when the PR is opened — if the team does not exist or lacks write access to this repo, the Settings > CODEOWNERS errors page will flag the line. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- Repo-internal review routing, not user-facing. --> - Did you get Claude approval on this PR?: N/A ### Additional Information Note for reviewers: shared agent config and scripts under `.agents/` (everything that is not a symlink into `plugins/`) are still owned by the default `@NVIDIA/modelopt-devs` rule. Happy to add `/.agents` here too if the agents team should own that as well. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Added repository ownership rules for the plugins directory and its agent configuration files. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
96b4aac90a |
Condense changelog entries and tighten conciseness guidelines (#2181)
### What does this PR do? Type of change: documentation Trims the 0.46 and 0.47 changelog entries and adds the guidelines that keep future entries (and comments, docstrings, tests) from growing the same way. **`CHANGELOG.rst`** - 0.47: `**New Features**` split into `*Quantization*` / `*Megatron Framework (M-LM / M-Bridge)*` / `*Misc*`, matching 0.46's organization. - 0.46 bug fixes: all 14 entries condensed to symptom → cause → fix. - Removed internal `NVBug` IDs (0.45, 0.46) and root-cause forensics that only matter to maintainers — exact tensor shapes, upstream-bug analysis, internal helper names. - Trimmed the most verbose feature entries (dLLM, Torch-TensorRT ViT, `local_hessian` Triton, `module_search_spaces`, `constant_amax`, CP/DP, `day0-release`) and the Phi-4-multimodal breaking-change note. **`AGENTS.md`** - What earns a changelog entry, and that entries are one or two sentences written for external users, with features filed under the right `**New Features**` sub-section. - What belongs in a PR description, so detail redirected out of the changelog and docstrings has a defined home. **`CONTRIBUTING.md`** - *Comment cautiously*: comments and docstrings capped at one or two lines; rationale, benchmarks, and root cause go in the PR description. - *Test design principles*: new lead bullet preferring the highest-level test that runs the real code path, with GPU/framework behavior going to `tests/gpu*` instead of monkeypatched CPU approximations in `tests/unit`; one test per behavior, `@pytest.mark.parametrize` over near-duplicate test functions. **`.github/PULL_REQUEST_TEMPLATE.md`** - Changelog checklist item now asks for a very short summary. ### Usage N/A — documentation only. ### Testing `pre-commit run --files CHANGELOG.rst AGENTS.md CONTRIBUTING.md .github/PULL_REQUEST_TEMPLATE.md` passes, including the RST and markdownlint hooks. No code changes, so no test suite applies. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- this PR only edits existing entries --> - Did you get Claude approval on this PR?: ❌ ### Additional Information Only the wording of released 0.45/0.46 entries changed; no entry was added or removed. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d3fe8117ff |
Add ModelOpt agent plugin marketplace (#2025)
### What does this PR do? Type of change: new feature Packages the existing ModelOpt agent skills as installable Codex and Claude plugins: - Adds a repo-scoped Codex marketplace and Claude-compatible marketplace. - Adds the canonical `plugins/modelopt/` plugin tree and manifests. - Moves the skill tree into the plugin and keeps `.agents/skills` as a compatibility symlink. - Adds a minimal `common` placeholder skill required by Codex validation. - Documents installation from this repository. ### Usage ```bash codex plugin marketplace add NVIDIA/Model-Optimizer ``` Then open `/plugins`, select the `modelopt` marketplace, and install `modelopt`. For Claude Code: ```bash claude plugin marketplace add https://github.com/NVIDIA/Model-Optimizer.git claude plugin install modelopt@modelopt ``` ### Testing - Codex plugin validator - `claude plugin validate . --strict` - `claude plugin validate plugins/modelopt --strict` 1. Install the marketplace plugin with Codex and Claude from an unrelated temporary workspace. 2. Exercise packaged evaluation helpers, a day-0 gate, and the shared remote helper from that workspace. 3. Run `uv run --frozen --extra dev python -m pytest -q plugins/modelopt/skills/day0-release/tests/test_gates.py plugins/modelopt/skills/benchmark-model-kernels/tests`. 4. Run pre-commit hooks for all changed files. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — added a plugin-path validator; existing focused skill tests and installed-plugin smoke tests pass. - Did you update Changelog?: N/A — agent tooling and distribution only. - Did you get Claude approval on this PR?: N/A ### Additional Information Skills remain available through `.agents/skills`; bundled helpers are packaged under the plugin and resolved from `$SKILL_DIR` so installed workflows do not depend on the current workspace. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added installable ModelOpt plugins for Claude Code and Codex. * Added skills for PTQ, deployment, evaluation, monitoring, debugging, benchmarking, MLflow access, EAGLE3 workflows, and release management. * Added deployment helpers, evaluation recipes, checkpoint validation, and release-gating tools. * **Documentation** * Expanded setup, credential, SLURM, benchmarking, deployment, evaluation, troubleshooting, and workspace guidance. * Added installation instructions and updated agent-skill discovery guidance. * **Maintenance** * Updated skill references and compatibility links for reliable use across supported plugin environments. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chad Voegele <cvoegele@nvidia.com> |
||
|
|
43e9d15e0b |
Give each example lane its own Codecov flag (#2111)
### What does this PR do? Type of change: CI/CD bug fix Every example lane uploaded coverage under one shared `examples` flag. That was correct while all lanes ran on every PR, but #2090 gates them independently, and **Codecov carryforward only applies to a flag with no upload on the commit**. | scenario | `examples` flag | outcome | |---|---|---| | no lanes run (docs-only) | absent | ✅ carried forward | | all lanes run | complete | ✅ correct | | **one lane runs** (now common) | **present but partial** | ❌ carryforward skipped; full coverage replaced by that lane's subset | The third row is what gating made routine: a PR touching only `examples/diffusers/**` runs the onnx lane, uploads `examples` containing onnx coverage alone, and Codecov reports a drop for code the PR never touched. One flag per example (`examples-<name>`, 12 flags) restores the intent already documented in `.github/codecov.yml`: a skipped lane has no upload for its flag and is carried forward; a lane that ran replaces only its own slice. The config comment is updated to explain why a shared flag defeats carryforward, so this isn't re-introduced. `gpu_tests` deliberately keeps a single `gpu` flag — its five suites are gated at workflow level, so they upload together or not at all. It would need the same change if per-suite gating is ever added. ### Testing Not directly observable on this PR: it changes workflow files, which are in the gate's `common` group, so **all twelve lanes run** and every flag is uploaded — the healthy case either way. The behavior it fixes appears on the next PR that touches a single example, where `codecov/project` should now stay accurate instead of reporting a drop. Worth noting the symptom was never blocking: `codecov/project` is not a required check and the threshold allows a 2% drop. This is about the coverage data being right. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — flags are new names; historical data under `examples` is unaffected - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — CI configuration - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ — not yet run <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Improved coverage reporting by tracking results separately for each test example. * Clarified coverage configuration and documented how skipped uploads are carried forward. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
99116c3fb9 |
Bump codecov-action to v7, the last Node 20 warning (#2110)
### What does this PR do? Type of change: CI/CD maintenance #2102 and #2103 bumped every action this repo references *directly*, but Node 20 annotations still appear — e.g. [this run on #2103](https://github.com/NVIDIA/Model-Optimizer/actions/runs/31170982088?pr=2103): > The following actions target Node.js 20 but are being forced to run on Node.js 24: `actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea` We never reference `github-script`. It comes in transitively through `codecov/codecov-action`: | codecov-action | bundled github-script | runtime | |---|---|---| | `@v5` (= v5.5.5) | `60a0d830…` | **node20** | | `@v6.0.2` / `@v7.0.0` | `ed597411…` | node24 | Bumps all four call sites (`unit_tests`, `gpu_tests`, `regression_tests`, `_example_tests_runner`) to `@v7`. Checked: v6's notes call out node24 support as the only breaking aspect — the same pattern as the bumps in #2102 — and our runners report 2.336.0. The inputs used here (`token`, `files`, `flags`, `fail_ci_if_error`, `verbose`) all still exist in v7. A scan of every referenced action, direct and one level transitive, now finds no node20 runtimes left. ### Testing Coverage upload runs in every unit, gpu, regression and example job, so CI exercises this broadly. Worth checking that coverage still lands in Codecov rather than only that the step is green — `fail_ci_if_error: false` means an upload failure would not turn the job red. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — CI configuration - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ — not yet run <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated automated test workflows to use the latest coverage reporting action. * Improved compatibility and reliability of coverage report uploads across example, GPU, regression, and unit tests. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
02b64bcce7 |
Bump step-security/changed-files to v47 (#2103)
### What does this PR do? Type of change: CI/CD maintenance Last of the [Node 20](https://github.blog/changelog/2025-09-19-deprecation-of-node-20-on-github-actions-runners/) actions: `step-security/changed-files` v46.0.5 → v47.0.5, in `_pr_gate.yml`, `example_tests.yml` and `unit_tests.yml`. Deliberately separate from #2102. Every gate in the repo runs through this action, and the example lanes depend on its `files_yaml` per-group outputs plus `any_modified` semantics. v47 has no release notes describing output behavior, and the upstream v47 notes are dependency bumps only — so this is the one bump I could not clear from a changelog. On its own, any gating regression is unambiguous. What I did verify at `v47.0.5`: - `micromatch` is still `^4.0.5` — the matcher the lane patterns were validated against - `files`, `files_ignore`, `files_yaml`, `files_ignore_yaml` are all still inputs - `any_modified`, `any_changed`, `changed_keys` are all still documented outputs ### Testing Static checks above. The behavior that matters cannot be proven from this PR: it changes workflow files, which are in the `common` group, so **every lane runs regardless** of whether gating still works. I plan to confirm with a throwaway probe PR against this branch — a docs-only change must run nothing, and a single-example change must run exactly one lane — the same method that caught the two ignore bugs fixed in #2101. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — CI configuration - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ — not yet run <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated pull-request file-change checks to use the latest available file-detection action. * Applied the update consistently across example-test and unit-test workflows. * Improved consistency and reliability across automated pull-request validation checks without changing application behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
30b68c38c7 |
Bump the Node 20 actions that have no behavior change (#2102)
### What does this PR do? Type of change: CI/CD maintenance Clears the [Node 20 deprecation](https://github.blog/changelog/2025-09-19-deprecation-of-node-20-on-github-actions-runners/) warnings. The runner already forces these onto Node 24, so this changes what the actions *declare*, not how they run. | action | bump | why it is safe | |---|---|---| | `actions/cache` | v4 → v6 | v5 requires runner >= 2.327.1; our self-hosted GPU runners report **2.336.0** | | `actions/download-artifact` | v4 → v8 | v5's breaking change affects downloads **by artifact ID**; `pages.yml` downloads by `name: docs-html` | | `actions/upload-artifact` | v4 → v7 | node24 runtime only | | `dorny/paths-filter` | v3 → v4 | node24 runtime only | | `poseidon/wait-for-status-checks` | v0.6.0 → v0.7.0 | node24 runtime only | `step-security/changed-files` stays at v46.0.5 and is bumped in a separate PR: the lane gating depends on its `files_yaml` group outputs and `any_modified` semantics, and v47 ships no release notes covering them. Splitting keeps a gating regression attributable. ### Testing `actions/cache` is exercised by every GPU job through `cache-extensions`; the `pages.yml` three are exercised by the docs build. Worth watching on this PR: cache **hit rate** rather than just success, since a silent cache-key change costs build time without failing. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — CI configuration - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ — not yet run <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated automated caching and build workflow actions to newer versions. * Improved pull request status-check handling. * Updated artifact upload, download, and path-filtering actions used by page deployments and tests. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3d4430f367 |
Apply the docs-only ignore to the example lane groups (#2101)
### What does this PR do? Type of change: Bug fix (CI) Follow-up to #2090. `files_ignore` does not apply to `files_yaml` groups — the action keys ignores separately through `files_ignore_yaml`. Docs-only PRs therefore still matched a lane: a change to `examples/diffusers/README.md` alone started the onnx lane's three GPU jobs. ```yaml files_ignore_yaml: | common: &docs - "**.ipynb" - "**.md" - "**.png" - "**.rst" torch: *docs trtllm: *docs megatron: *docs onnx: *docs ``` Only `example_tests.yml` is affected. `gpu_tests`, `regression_tests` and `unit_tests` go through `_pr_gate.yml`, which uses the plain `files` + `files_ignore` pair where the ignore does apply. ### Testing Found by a probe PR opened against merged main (a README-only change), which showed `onnx` running when nothing should have. The same probe is re-run against this branch to confirm the fix — see the linked draft PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — CI configuration - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ — not yet run <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated example-test workflow filtering to consistently ignore documentation, image, and notebook-only changes across all test lanes. * Improved pull request gate file matching by using recursive patterns for Markdown, reStructuredText, PNG, and notebook files. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3d4d9249f4 |
Gate example and GPU test lanes on the files they cover, and consolidate the CI gate (#2090)
### What does this PR do? Type of change: CI/CD improvement Follow-up to #2086, which carried the speculative-decoding fix; this PR is the CI half. **Lanes now run only when their own files change.** A one-line edit to any example started all 12 example lanes, and any `modelopt/**` change started every GPU suite. - Each lane gates itself inside `_example_tests_runner.yml` / `_gpu_tests_runner.yml`, deriving its watch list from the example or suite name it was already given (`examples/<name>/**`, `tests/examples/<name>/**`, `tests/<suite>/**`). **Adding a new example stays a one-line matrix entry** — no central mapping to update. - The five cross-example dependencies are declared as `watch_extra` next to the example that needs them: `hf_ptq` → `llm_eval` (`huggingface_example.sh` runs lm_eval from `../llm_eval`), `torch_trt` → `onnx_ptq`, `speculative_decoding` → `hf_ptq` + `dataset`, `llm_qat` / `gpt-oss` → `dataset`. - `gpu_tests.yml` is split into caller + runner to match. This shape is forced, not stylistic: job-level `if:` cannot read `matrix`, and `container:` images are pulled before any step runs, so gating inside the job would still pull 10–20 GB and hold a GPU runner for every skipped suite. - `modelopt/**`, `modelopt_recipes/**`, `pyproject.toml` and `tests/_test_utils/**` still run every lane. The gate now watches the last two, which example tests depend on but it previously ignored. **Docs-only changes no longer start GPU jobs.** The gate ignores `**.md`, `**.rst`, `**.png` and `**.ipynb` by default, so a README edit short-circuits the whole workflow. Nothing executes notebooks (no `nbmake`/`nbval` in the repo), and `.sh`/`.yaml`/`.txt` stay watched since examples run them. **One file holds the gate logic.** `.github/actions/changed-files-gate` is a composite action doing the merge-base + changed-files comparison and, optionally, the `^linux$` wait. `_pr_gate.yml` and `_wait_for_checks.yml` are both deleted: each top-level workflow keeps a 12-line `pr-gate` job that is pure wiring, and the runners use the action as steps so a lane shows `gate` + `run-test` rather than three checks. Calling a reusable workflow always materializes all of its jobs, including skipped ones, which is what made the per-lane check list noisy. `unit_tests.yml` also drops its DCO wait: DCO can be marked passing manually, so blocking the matrix on it only delayed feedback. The `^linux$` wait remains, which is the gate that actually protects GPU runners. **Also, from the original lane consolidation:** - **One TensorRT-LLM lane.** `trtllm-pr` and `trtllm-non-pr` merge into a single `trtllm` job gated like the others, so `llm_eval` now runs on PRs, where it was nightly-only. - **`gpt-oss` moves to the TensorRT-LLM image.** Its deploy step needs `tensorrt_llm`, which `pytorch` doesn't have, so `deploy_gpt_oss_trtllm` silently skipped in CI. Its `importorskip` is dropped now that the lane guarantees the dependency. - **Containers bumped where no reason was documented:** pytorch `26.06`/`26.01` → `26.07` (torch example lane, regression). Left pinned with their existing in-file reasons: pytorch `26.05` for the gpu lane (`EXPLICIT_BATCH` removed in TensorRT 11), tensorrt `26.05` for the onnx lane (`torch-tensorrt` needs `libnvinfer.so.10`), vllm `v0.20.0` (legacy FusedMoE coverage). TensorRT-LLM stays on `1.3.0rc20`: rc21–rc23 ship a `quickstart_multimodal.py` importing `MultimodalConfig` before `tensorrt_llm.llmapi` exported it (fixed upstream in NVIDIA/TensorRT-LLM#17112, one day after rc23 was cut), which fails the `hf_ptq` VLM deploy smoke test. Also clarifies the changelog line in the PR template to spell out when an entry is expected. **Two silent-failure fixes found in review:** - Every gate used `any_changed`, which is ACMR and excludes deletions, so a delete-only PR (removing an example, a test, or library code) ran nothing. Now `any_modified` (ACMRD), fixed in `example_tests.yml`, `_pr_gate.yml` and `unit_tests.yml`. - `git merge-base` was piped into `tee`, so a failure returned `tee`'s exit status and emitted an empty base, quietly changing which lanes run instead of failing. **Nightly secret scanning is fixed.** `code_quality.yml` excludes trufflehog's `lob` detector: its `(live|test)_[a-zA-Z0-9_]{35}` pattern matches any pytest function whose name is exactly 35 characters after `test_` — 60 of them in this repo, e.g. `test_all_zero_activation_yields_no_scale` — and reports them as *verified* secrets. This only ever failed nightly because the action scans all history on `schedule` and only the PR diff on `pull_request`. ### Testing Workflow YAML validated locally and the selection logic checked by hand across scenarios (`examples/diffusers/**` → onnx lane only; `examples/dataset/**` → `llm_qat`, `speculative_decoding`, `gpt-oss`; `examples/llm_eval/**` → `hf_ptq` + `llm_eval`; `modelopt/**` and nightly → everything). Gating behavior itself can only be exercised by a real PR run — the failure mode to watch for is a lane skipping when it should have run. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — CI configuration - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ — not yet run <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **CI Improvements** - Improved change detection for model recipes, test utilities, workflow actions, and documentation-only updates. - Streamlined pull request checks and status monitoring. - Consolidated TensorRT-LLM example validation and refined conditional test execution. - Updated test environments to newer PyTorch releases. - Refined GPU, regression, and unit test triggers. - Deployment tests no longer automatically skip when TensorRT-LLM is unavailable. - Updated secret scanning configuration. - **Documentation** - Updated pull request checklist guidance to include deprecations and critical bug fixes. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
22b6a148b0 |
Fix EAGLE-3 context-parallel training and re-enable its tests (#2086)
### What does this PR do?
Type of change: Bug fix
**EAGLE-3 context-parallel training (`--cp_size > 1`) is fixed, and its
tests run again.** CP has been broken since `accelerate` 1.13, and the
tests never caught it: the guard compared `Version("2.10.0a0")` against
`Version("2.10.0")`, which is False on every NGC alpha torch build, so
`test_llama_eagle3[cp_size=2]` has never actually run in CI.
Five fixes:
- **`main.py`** — rebuild the FSDP2 plugin accelerate requires for
`cp_size > 1`. The `--fsdp full_shard --fsdp_config` launcher flags that
used to supply it were dropped from `launch_train.sh`, so CP could not
start at all. Also pass the CP degree to the draft model.
- **`modeling_eagle.py`** — apply the draft model's first input norm
inside `layers[0]`'s own forward, where FSDP2 has actually unsharded its
weights, and only stash the input embeds on the path whose pre-hook
consumes them.
- **`hf_eagle.py`** — skip the dense eagle attention mask under CP
(causal masking comes from `is_causal`, TTT masking from the
ring-attention patch), and warn that padded positions are therefore
unmasked. Also stop `(eagle_loss or 0)` replacing a `0.0` loss tensor
with a plain `int`, which detached the graph.
- **`eagle_utils.py`** — key TTT-mask injection off the backward call's
`grad_out` kwarg, since newer torch omits `attn_bias` on the forward
call, silently disabling TTT masking.
- **`utils.py`** — CUDNN-only SDPA under CP; the `MATH` backend
decomposes SDPA and breaks on DTensors. Scoped to `cp_size > 1`, since
this context manager wraps every training forward and CPU has no cudnn
backend.
**Drops the `speculative_decoding` 26.01 container override.** It was
added when the lane ran 25.06 and spec-dec needed something *newer* — a
floor. Later bumps moved the default past it, so it had silently become
a ceiling holding spec-dec on a 6-month-old image.
### Testing
Ran `tests/examples/speculative_decoding` in
`nvcr.io/nvidia/pytorch:26.07-py3` on 2 GPUs, reproducing the CI install
steps (`pip uninstall -y nvidia-modelopt`, `pip install -e
".[hf,dev-test]"`, example requirements): **16 passed, 2 skipped** — the
2 skipped being pre-existing `--run-manual` tests. All four
`test_llama_eagle3` cases pass, including both `cp_size=2` ones.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — the existing `cp_size=2`
tests are re-enabled
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — not yet run
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
5dde396bdf |
Fix vLLM 0.24+ compatibility: registry TypeError and MoE RoutedExperts port (#2054)
### What does this PR do? Type of change: Bug fix vLLM 0.24 (as shipped in `nemo:26.08`) reworked the fused-MoE layer, which broke ModelOpt in two ways: 1. **`FusedMoE` became a factory function** returning a `MoERunner` pipeline. Registering it put a plain function into `QuantModuleRegistry`, so *every* later registry lookup raised `TypeError: issubclass() arg 2 must be a class, ...` — taking down large parts of `tests/gpu_megatron` (TE, Megatron chaining, MSE calibrator) that have nothing to do with vLLM. 2. **Expert weights moved onto a `RoutedExperts` submodule** of `MoERunner`, and `UnquantizedFusedMoEMethod` moved modules, so the MoE fakequant path had no valid registration target. Changes: - `_DMRegistryCls.register` now asserts keys are `nn.Module` subclasses, so a future upstream change fails at the registration site instead of as a confusing `TypeError` at lookup. - Register `RoutedExperts` (vLLM >= 0.24) while keeping the `FusedMoE` / `SharedFusedMoE` registrations for older releases; both go through the same `_QuantFusedMoEBase`. `MoERunner` calls `forward_modular` / `forward_monolithic` directly (`RoutedExperts.forward` raises by design), so those are hooked instead of `forward`. - The fused-MoE kernel patch now covers `experts.triton_moe` in addition to `fused_moe`. 0.24's launcher binds the kernel names at import time, so patching only the defining module would leave fakequant **silently inactive**. - `UnquantizedFusedMoEMethod` is resolved from either module layout. - `examples/vllm_serve/vllm_reload_utils.py`: quantizer module paths are now `mlp.experts.routed_experts.*`, so HF→vLLM expert key mapping inserts a matching `.routed_experts` infix when that layout is present. - CI `gpu_vllm` now runs on **two** containers: `v0.24.0` (first release with the FusedMoE-factory / `RoutedExperts` layout — `v0.24.1` was never released, `nemo:26.08` ships a `0.24.1.dev0` build of the same layout) and `v0.20.0`, which keeps the legacy `FusedMoE`/`SharedFusedMoE` branches covered. Test fixes for the newer vLLM (not product bugs): - The tiny Llama fixture used `hidden_size=32 / 16 heads` → `head_dim=2`, which `FLEX_ATTENTION` (the only backend available in this image) rejects with `NYI: embedding dimension ... must be at least 16`, killing the engine core at warmup. Now `head_dim=64`, matching the Qwen3-MoE fixture. - The FlashInfer metadata-builder stub used `causal=False`, which in 0.24 forces the FI-native path (`all_uses_trtllm = causal and ...`) requiring workspace buffers and real wrapper planning. Keep it on the all-TRTLLM path it was originally exercising; the stashed `_modelopt_*` fields are path-independent. ### Usage No API change — existing `mtq.quantize` / `examples/vllm_serve` flows work unmodified on both old and new vLLM. ### Testing All runs in the `nemo:26.08.rc3` container (vLLM `0.24.1.dev0+gee0da84ab`, the same 0.24.1 the CI job now pins). **Suites** - `tests/gpu_vllm`: **73 passed, 1 skipped** (was 70 passed / 3 failed). - `tests/gpu_megatron`: all pass (previously ~120 failures, all from the registry `TypeError` — TE, Megatron chaining and MSE-calibrator tests that never touch vLLM). - `tests/unit/torch/opt/test_dynamic.py`: 2 passed, including the new `test_register_rejects_non_module_classes` (rejects a factory function and a non-`nn.Module` class, and asserts no partial registration). - `pre-commit` clean on all touched files. **MoE fakequant verified by module-tree probe, not just by test assertions** After `mtq.quantize(..., NVFP4_DEFAULT_CFG)` inside the vLLM worker, every weight-owning module was enumerated on tiny DeepSeek-V3 (MLA + routed MoE + shared experts) and tiny Qwen3-MoE: ``` model.layers.0.mlp.experts.routed_experts [QuantRoutedExperts] w13_input_quantizer=3.484 w13_weight_quantizer=0.0840 w2_input_quantizer=0.1060 w2_weight_quantizer=0.0845 model.layers.0.mlp.shared_experts.gate_up_proj [QuantMergedColumnParallelLinear] ✅ model.layers.0.mlp.shared_experts.down_proj [QuantRowParallelLinear] ✅ ``` Weight amax being populated (not just input amax) means the `B is self.w13_weight` identity check and the Parameter-swap weight-fakequant branch actually execute through 0.24's kernel path — i.e. `forward_modular`/`forward_monolithic` really are the live entry points and the `experts.triton_moe` patch target is the one that fires. Unquantized modules were only the expected ones: embeddings, RMSNorms, MoE router `gate`, `lm_head`. Registration parity vs. older vLLM: Row/Column/MergedColumn/QKV `ParallelLinear` and all four attention types (`Attention`, `CrossAttention`, `EncoderOnlyAttention`, `MLAAttention`) register unchanged; `FusedMoE` → `RoutedExperts`; `SharedFusedMoE` has no counterpart because the `shared_fused_moe` module no longer exists in 0.24 — shared experts are now a plain MLP whose linears we already quantize (confirmed above). **Known gaps (pre-existing, not regressions from this PR)** - `DeepSeekV2FusedQkvAProjLinear` is not quantized: it subclasses `MergedColumnParallelLinear` but overrides `forward`, so the registry's shared-forward rule declines it. Pre-0.24 the equivalent (`q_a_proj` / `kv_a_proj_with_mqa`) were `ReplicatedLinear`, which ModelOpt never quantized — effective coverage is unchanged. - MoE fakequant hooks only the Triton expert kernels; FlashInfer/CUTLASS/DeepGEMM MoE backends bypass them (why the fixtures pin `moe_backend="triton"`). - `_setup` still requires a plain `UnquantizedFusedMoEMethod`; a `FusedMoEModularMethod` swap (some DP/all2all configs) still asserts. **Not covered by tests** - `examples/vllm_serve/vllm_reload_utils.py` — the expert key mapping is now asserted in `test_tiny_qwen3_moe_quantize` against the quantizer module paths of a booted MoE model, so a stale infix fails loudly instead of silently serving uncalibrated experts. The rest of the reload path is still inspection-only. Note the registry key moved `vllm_FusedMoE` → `vllm_RoutedExperts` and quantizer paths gained `.routed_experts`, so a `modelopt_state` saved under an older vLLM will not restore onto 0.24 as-is. - The legacy `FusedMoE`/`SharedFusedMoE` branches are covered by the second CI entry; the `v0.20.0` job is green on this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — all new paths are feature-detected; older vLLM keeps the `FusedMoE` registration. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — existing `tests/gpu_vllm` coverage exercises the new registration path. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ ### Additional Information Found while bumping the Megatron test environment from `nemo:26.06` to `nemo:26.08.rc3`. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added compatibility for newer vLLM MoE implementations, module layouts, and routed-expert configurations. * Improved model reload support for models using routed-expert submodules. * **Bug Fixes** * Improved detection and patching of vLLM MoE execution paths across supported configurations. * Registry validation now rejects invalid module registrations without partially applying changes. * **Tests** * Expanded GPU coverage for vLLM 0.24.0, dynamic module validation, and causal attention metadata paths. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
c062b3f829 |
Add sidecar GPU/CPU memory+utilization monitor for HF PTQ (#2000)
### What does this PR do?
**Type of change:** New feature (developer tool) + new tests
Adds a **standalone, cross-process resource monitor** —
`tools/resource_monitor.py` —
that samples a workload's GPU and CPU usage *from the outside* while it
runs, so
we can verify per-run memory budgets (e.g. the single-GPU layerwise PTQ
target of
≤80 GB GPU / ≤80 GB CPU for OMNIML-4947) without instrumenting
`hf_ptq.py` itself.
The monitor wraps any command, samples at a fixed interval, and on exit
writes a
CSV timeseries plus a peak/mean/min summary:
- **GPU** (per device, via NVML with an `nvidia-smi` fallback): used
memory,
utilization %, power draw (W), temperature (°C).
- **CPU** (via `psutil`): system total/used/free memory + utilization %,
and the
monitored process tree's RSS + CPU %.
It is opt-in from `examples/hf_ptq/scripts/huggingface_example.sh` via
`MODELOPT_MEM_MONITOR=1` (off by default → byte-for-byte identical
behavior). When
enabled it wraps the `hf_ptq.py` run and writes the trace/summary to a
**sibling**
`${SAVE_PATH}_mem_monitor/` directory, kept out of the exported
checkpoint that is
uploaded and consumed downstream.
**Files:**
- `tools/resource_monitor.py` — the sidecar (NVML + `nvidia-smi`
fallback; `psutil`).
- `tests/unit/tools/test_resource_monitor.py` — CPU-only unit tests (run
in the `unit` nox lane).
- `examples/hf_ptq/scripts/huggingface_example.sh` — opt-in
`MODELOPT_MEM_MONITOR=1` wrapper.
- `examples/hf_ptq/requirements.txt` — adds `psutil`.
- `pyproject.toml` — adds `psutil` to the `dev-test` extra
(deterministic import in the unit lane).
- `.github/workflows/unit_tests.yml` — adds `tools/resource_monitor.py`
to the unit-test path filters.
#### Why a new tool instead of extending
`modelopt/torch/utils/memory_monitor.py`?
The existing `GPUMemoryMonitor` is a fundamentally different tool and
cannot serve
this use case by extension:
| | `modelopt.torch.utils.memory_monitor.GPUMemoryMonitor` |
`tools/resource_monitor.py` (this PR) |
|---|---|---|
| Scope | **In-process** thread inside the workload | **Cross-process**
— wraps an external command |
| Survives workload OOM/SIGKILL | ❌ dies with the process | ✅ keeps
sampling, still writes the summary |
| Import cost | Pulls `torch` (~19 s) — lives in the workload |
Torch-free (`psutil`+`pynvml`, ~0.03 s) |
| Metrics | GPU device memory only | GPU mem/util/**power/temp** +
**CPU** mem/util + process-tree RSS |
| Output | In-memory / logs | CSV timeseries + peak/mean/min summary |
Merging the two would force `torch` into a standalone sidecar (defeating
the point)
or split the in-process monitor's threading model. A future refactor may
factor out
a **shared torch-free sampling core with two thin frontends**
(in-process + sidecar);
that is tracked as a follow-up rather than blocking this monitoring
harness, which
PR #2008 (single-GPU disk-offload PTQ) depends on.
### Usage
```bash
# Wrap mode (preferred): monitor exits with the workload's return code
python tools/resource_monitor.py --gpus 2,3 --out mem.csv --summary peak.txt -- \
python hf_ptq.py --pyt_ckpt_path=<model> --qformat=nvfp4 ...
# Opt-in from the HF PTQ example (off by default):
MODELOPT_MEM_MONITOR=1 CUDA_VISIBLE_DEVICES=2,3 CUDA_DEVICE_ORDER=PCI_BUS_ID \
bash examples/hf_ptq/scripts/huggingface_example.sh <args>
# -> writes ${SAVE_PATH}_mem_monitor/mem_trace.csv and mem_peak.txt
```
### Testing
- **Unit (CPU-only, in the `unit` nox lane):** `pytest
tests/unit/tools/test_resource_monitor.py`
— 11 tests covering `--gpus` parsing (CSV + space-separated, UUID/MIG
rejection),
the disabled/`nvidia-smi` sampling paths (including `[N/A]` → `None` and
the
smi-failure-yields-empty guard), CPU sampling, the accumulator, and
end-to-end
CSV/summary + exit-code propagation in wrap mode.
- **GPU-validated** on a B200 node (GPUs 2,3): confirmed the `gpu{i}_*`
memory /
utilization / power / temperature columns populate and the summary is
written.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — new tool; the example wrapper
is off unless `MODELOPT_MEM_MONITOR=1`.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — `psutil`
added to `dev-test` + the hf_ptq example requirements;
`pynvml`/`nvidia-ml-py` is an optional runtime dep (graceful
`nvidia-smi` fallback).
- Did you write any new necessary tests?: ✅ —
`tests/unit/tools/test_resource_monitor.py`.
- Did you update Changelog?: N/A — repo-level `tools/` script, not
shipped in the wheel.
- Did you get Claude approval on this PR?: ❌ <!-- run /claude review -->
### Additional Information
Part of **OMNIML-4947** (single-GPU disk-offload PTQ). This is PR 1 of
the stack —
the monitoring harness that PR #2008 (disk-offload layerwise PTQ +
offload-aware
export) builds on.
---------
Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
c2070cfd7a |
ci: skip docs preview deploy for fork PRs (#2029)
### What does this PR do? Type of change: Bug fix (CI) The `deploy-preview` job in the `Docs` workflow fails on **every pull request opened from a fork**, which blocks merging for all external contributors. **Root cause.** `deploy-preview` runs `rossjrw/pr-preview-action@v1`, which pushes the built HTML to the `gh-pages` branch. The workflow declares `permissions: contents: write`, but for a `pull_request` event originating from a forked repository GitHub caps the `GITHUB_TOKEN` at **read-only** — the `permissions:` block cannot elevate above that cap. The push is therefore rejected: ``` remote: Permission to NVIDIA/Model-Optimizer.git denied to github-actions[bot]. fatal: unable to access 'https://github.com/NVIDIA/Model-Optimizer.git/': The requested URL returned error: 403 ``` The job's `if:` condition gated on `github.event_name`, `github.event.action` and the `changes` path filter, but never on whether the PR came from a fork — so it always ran and always failed. Because the `changes` filter matches `docs/**`, `modelopt/**` and `.github/workflows/pages.yml`, essentially any substantive fork PR trips this. **Fix.** Restrict `deploy-preview` to PRs whose head branch lives in this repository: ```yaml github.event.pull_request.head.repo.full_name == github.repository ``` A skipped job is not a failed job, so fork PRs are no longer blocked by it. ### Usage N/A — CI-only change. ### Testing Behaviour by scenario: | Scenario | Before | After | | --- | --- | --- | | PR from a branch in this repo | preview deployed | preview deployed (**unchanged**) | | PR from a fork | ❌ fails with 403 | ⏭️ skipped | | Fork deleted (`head.repo` is `null`) | ❌ fails | ⏭️ skipped | - Confirmed against workflow history: recent `Docs` runs on in-repo branches (`main`, `chenjiel/nvfp4-act-headroom`, `mxin/qad-skill`, `haoguo/dspark-ptq-script`) all succeed, while fork-branch runs fail with the 403 above. - `build-docs` was already passing on the affected PRs — only the deploy step failed, so documentation builds are unaffected either way. - YAML parses; `pre-commit run --files .github/workflows/pages.yml` passes. (`yamlfmt` excludes `^.github/workflows/`, so this file is not auto-formatted.) - This PR edits `.github/workflows/pages.yml`, which is itself in the `changes` filter, so it exercises `deploy-preview` on the in-repo path — the preview deploy on this PR passing is a self-check that the unchanged path still works. The `closed` cleanup path is gated by the same condition. That is intentional: a fork PR never deployed a preview directory, so there is nothing to remove. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — workflow-condition change; not unit-testable - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — CI infrastructure, not user-facing - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information Currently blocking #1975 (`add Qwen3-VL support for DFlash training`), which is approved with every other check green and sits at `mergeStateStatus: BLOCKED` solely because of this job. Note that a PR only picks up this fix once its branch contains it, since workflows run from the PR branch's own definitions. A follow-up option, if doc previews for external contributors are wanted: build in the `pull_request` workflow and deploy from a separate `workflow_run`-triggered workflow, which executes in the base-repo context and does get a write token. Deliberately not using `pull_request_target` here — that would run unreviewed PR code with write permissions. Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
42458def24 |
ci: fix torch_trt on torch 2.13; default unit tests to torch 2.13; add example allow-failure hatch (#1951)
### What does this PR do? Type of change: Bug fix + CI / infra torch 2.13 + torchvision 0.28 were published to PyPI on 2026-07-08 and broke the `onnx (torch_trt)` example job (which had passed the day before). This PR fixes that break and hardens CI against the next one: 1. **Fix `torch_trt` (torch/torchvision/torch-tensorrt trio pin).** The base install pulled torch 2.13 / torchvision 0.28, then `torch-tensorrt 2.12.1` downgraded torch back to 2.12 but left torchvision at 0.28 (which pins `torch==2.13`) — breaking `import`. `examples/torch_trt/requirements.txt` now caps `torch-tensorrt>=2.4.0,<2.13` + `torchvision<0.28` so the trio stays consistent (also protects direct `pip install -r` users). 2. **Unit tests default to torch 2.13.** `noxfile.py` gains `torch_213` (`torchvision~=0.28.0`); the required `linux`/`windows` jobs and the multi-version Python spread (3.10/3.11/3.13/3.14) now run torch 2.13, with torch 2.8–2.12 kept as back-compat legs on Python 3.12. 3. **Per-example allow-failure escape hatch.** `_example_tests_runner.yml` gains an `allow_failure` input; when set, a **test-run** failure is surfaced as a `::warning::` via `continue-on-error` instead of blocking the PR. `example_tests.yml` derives it per example from the repo variable **`ALLOW_FAILURE_EXAMPLE_TESTS`** (comma-separated example names, comma-wrapped so `onnx` ≠ `torch_onnx`). Future breakages can be quarantined by updating the variable — no code change / PR required. ### Testing - Ran the new default unit session locally in an isolated uv venv (torch **2.13.0**+cu130, torchvision **0.28.0**+cu130, transformers **5.12.1**): ``` nox -s "unit-3.12(torch_213, tf_latest)" => 2813 passed, 15 skipped, 1786 warnings in 250.37s ``` - Verified the allow-failure hatch: with `ALLOW_FAILURE_EXAMPLE_TESTS=torch_trt`, the (previously failing) `torch_trt` job reports success with a warning and does not block the required example check. `vars` is re-read on each job attempt, so "Re-run failed jobs" picks up the variable without a fresh trigger. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (CI-only; older torch versions still covered) - If you copied code from any other sources or added a new PIP dependency: N/A (no new dependency; only version caps) - Did you write any new necessary tests?: N/A (CI configuration change) - Did you update Changelog?: N/A (CI infra, no user-facing API change) - Did you get Claude approval on this PR?: ❌ (pending — will run `/claude review`) ### Additional Information The `ALLOW_FAILURE_EXAMPLE_TESTS` repo variable can be cleared for `torch_trt` now that the requirements pin lands the real fix; keep it as the standing escape hatch for future example breakages. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Example test jobs can now be configured to “allow failure” without failing the workflow; when enabled, a warning annotation is emitted. * **Bug Fixes** * CI unit-test and GPU-test configurations were refreshed (including a reduced timeout for the `gpu_megatron` job). * **Tests** * Updated unit-test coverage to use the newest Torch 2.13-based setup by default, with back-compat retained where applicable. * **Documentation** * Added inline guidance for how the allow-failure examples list is specified. * **Chores** * Refreshed `torch_trt` example dependency constraints to improve compatibility. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
795c589428 |
CI: CUDA build/test hygiene + fix Puzzletron Nemotron test failures (#1901)
### What does this PR do? Type of change: Bug fix (CI / tests) - **Fix Nemotron nightly failures:** install `mamba_ssm`/`causal-conv1d` from PyPI releases instead of git `main` (avoids the broken `apache-tvm-ffi 0.1.12` that crashes on import). - **Speed up CUDA builds:** set `TORCH_CUDA_ARCH_LIST=12.0` (runner's sm_120) in the GPU/example/regression workflow container env instead of the image's ~6 archs. - **Make unit tests CPU-only:** force CUDA off in the nox `unit` env and skip JIT-compiling CUDA extensions when no GPU is usable; move the two GPU-/`mamba_ssm`-requiring unit tests to `tests/gpu`. - **Harden example tests against HF flakes:** capture subprocess output and retry transient HuggingFace access errors (5xx / rate-limit / connection). - **Skip Blackwell-flaky sharded-state-dict tests:** `test_homogeneous_sharded_state_dict` and `test_regular_state_dict[320]` intermittently hit a CUDA illegal-memory-access on the sm_120 runner that poisons the CUDA context and cascades timeouts; gate them behind a reusable `skip_flaky_on_blackwell` marker (still run on non-Blackwell GPUs). - **Bump slow test timeout:** `test_prune_minitron_vlm` → 360s for the 2-GPU nightly. ### Testing - CI tests on this PR pass (1-gpu) - Manually triggerred 2-gpu test: - GPU: https://github.com/NVIDIA/Model-Optimizer/actions/runs/28774553356 - Examples: https://github.com/NVIDIA/Model-Optimizer/actions/runs/28774556693 - Regression: https://github.com/NVIDIA/Model-Optimizer/actions/runs/28771197586 ### Additional Information - Backward compatible: N/A (CI/tests only) - New dependency: N/A - Changelog: N/A (CI/test infra) 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
d1c8ea95eb |
ci: gate example and GPU tests on changes to their own runner and cache action (#1891)
### What does this PR do? Type of change: Bug fix Adds `.github/workflows/_example_tests_runner.yml` and `.github/actions/cache-extensions/**` to the example_tests pr-gate `files:` list, and `cache-extensions/**` to the gpu_tests gate. Every example-test job executes through the reusable runner (which is taken from the PR's ref), so a PR changing only the runner previously merged with all example-test jobs skipped — observed on PR #1651, whose example-tests run shows pr-gate success with torch/trtllm-pr/megatron/onnx all skipped. Gating on the workflow's own machinery matches the existing intent (the gate already lists `example_tests.yml` itself). ### Usage N/A — CI workflow configuration change. ### Testing YAML validated; the change is additive to the gate lists only. The observed-failure evidence is PR #1651's run history (linked in #1890). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (workflow config; validated as described above) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (CI-only change) - Did you get Claude approval on this PR?: N/A (external contributor; cannot trigger `/claude review`) ### Additional Information Part of #1890 (item 1). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Expanded automated test-triggering rules to cover additional workflow and action changes. * Updated gating for example and GPU test runs so more relevant changes are validated automatically. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: arham766 <arhamislam766@yahoo.com> |
||
|
|
0c40c374e5 |
ci: give non-PR code quality runs distinct concurrency groups (#1894)
### What does this PR do? Type of change: Bug fix The concurrency group used `github.event.pull_request.number` with no fallback, so every nightly and manually dispatched run shared the literal group `Code Quality-` with `cancel-in-progress: true` — a manual dispatch cancels an in-flight nightly and vice versa. Adds the `|| github.sha` fallback already used by `unit_tests.yml`. ### Usage N/A — CI workflow configuration change. ### Testing Matches the existing pattern in `unit_tests.yml` line-for-line. YAML validated. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (workflow config; validated as described above) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (CI-only change) - Did you get Claude approval on this PR?: N/A (external contributor; cannot trigger `/claude review`) ### Additional Information Part of #1890 (item 4). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Improved workflow concurrency handling so automated runs are grouped and canceled more reliably across pull request, manual, and scheduled triggers. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: arham766 <arhamislam766@yahoo.com> |
||
|
|
ecc1e4f15f |
ci: scope bump_uv_lock change detection to uv.lock (#1892)
### What does this PR do? Type of change: Bug fix The torch-override step rewrites `pyproject.toml` via `toml.load`/`toml.dump`, which drops comments and reformats the file, so the unscoped `git diff --quiet` was always dirty and `changed=false` unreachable. In a week where `uv lock --upgrade` produces no changes, the create-pull-request step stages nothing and `git commit` exits 1, failing the scheduled run instead of no-oping. Scoping the diff to `uv.lock` restores the intended detection. ### Usage N/A — CI workflow configuration change. ### Testing Reproduced the failure locally by executing the workflow's steps in a clean checkout: after the toml mutation, unscoped `git diff --quiet` exits 1 with `uv.lock` byte-identical, and `git add uv.lock && git commit -s -m x` exits 1. With `git diff --quiet -- uv.lock`, the no-update path correctly yields `changed=false`. YAML validated. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (workflow config; validated as described above) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (CI-only change) - Did you get Claude approval on this PR?: N/A (external contributor; cannot trigger `/claude review`) ### Additional Information Part of #1890 (item 2). Signed-off-by: arham766 <arhamislam766@yahoo.com> |
||
|
|
9038b71f09 |
Autoquant and GPTQ in support in Megatron-Core [OMNIML-3151] (#1562)
### What does this PR do? Type of change: New Feature Autoquant and GPTQ in support in Megatron-Core - Add EP support to AutoQuantize - Register MCore support in AutoQuantize - Add decoder `output_layer` (lm head) to layerwise hook so that GPTQ can register all decoder layers & lm head - Split dataloader helper function out of megatron calibration utils so that AutoQuantize in Megatron-LM can reuse the same dataloader ### Usage See https://github.com/NVIDIA/Megatron-LM/pull/4821 for Autoquant usage in Megatron ```python # For GPTQ pick a recipe that uses gptq algorithm and run mtq.quantize # e.g. general/ptq/nvfp4_default-kv_none-gptq ``` ### Testing Tested AutoQuant on Nemotron Nano and Ultra. Tested GPTQ on Nano 3. Added unit tests for both AutoQuant and GPTQ ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Summary by CodeRabbit * **New Features** * Added lazy Megatron-Core AutoQuant integration with Megatron-specific quantization hooks and better decoder-layer discovery for layerwise calibration. * Improved AutoQuantize for expert-parallel (EP) models, including consistent per-layer recipe selection across DP/TP/EP. * Extended quant-layer grouping for NemotronH MCore fused “local_experts” linear layers. * **Bug Fixes** * Prevented division-by-zero when calibration inputs are empty during Hessian updates. * **Tests** * Added/extended unit and GPU coverage for EP AutoQuant, decoder-layer calibration discovery behavior, and zero-token Hessian no-op. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> Signed-off-by: Jenny Chen <jennifchen@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
fbbc5989ce |
[6078291][OMNIML-3716] Add ViT FP8 + Torch-TRT example, wire softmax_quantizer in _QuantAttention (#1569)
### What does this PR do?
Type of change: new feature + bug fix
Adds a Torch-TensorRT deployment path for HuggingFace ViT and closes the
modelopt-side gap that prevented `*softmax_quantizer` from being applied
on the standard attention forward path.
* **New ViT PTQ recipes** under `modelopt_recipes/huggingface/vit/ptq/`:
* `fp8.yaml` — W8A8 per-tensor FP8 E4M3 on encoder Linear
weights/inputs;
attention Q/K/V BMMs + softmax output at FP8; per-block LayerNorm output
at FP8 (one shared Q/DQ feeds Q/K/V + MLP); patch-embed `nn.Conv2d`,
`classifier`, and the final `vit.layernorm` left FP16. Uses max
calibration.
* The recipe is self-contained (no `$import` of shared snippets) and
use the "specific-enable" style: narrow `parent_class` + path scoping
on every enable rule, so no `enable: false` carve-outs are needed.
* **New example** under `examples/torch_trt/`:
* `torch_tensorrt_ptq.py` — single-model pipeline (load HF model,
calibrate from `zh-plus/tiny-imagenet`, `mtq.quantize`,
`torch_tensorrt.compile`, verify the compiled-model argmax matches the
fake-quant argmax). Defaults to `google/vit-large-patch16-224`; pass
`--model_id` and `--recipe` to target any model + recipe combination.
`--no_pretrained` + `--model_kwargs` shrink the model for fast tests.
* `README.md` documenting the flow, the shipped recipes, hardware
requirements, and CLI usage.
* `requirements.txt`.
* **Bug fix in `modelopt/torch/quantization/plugins/huggingface.py`** —
inside
`_QuantAttention._quantized_attention`, the non-kitchen branch now
temporarily replaces `torch.nn.functional.softmax` (via the existing
`replace_function` context manager) with a wrapper that pipes the
softmax
output through `self.softmax_quantizer`. Previously the slot was created
on every registered attention class but only consumed by the optional
Kitchen MXFP8 flash-attention path, so FP8 / NVFP4 recipes that enabled
`*softmax_quantizer` saw it stay uncalibrated (`amax=None`) and emitted
no Q/DQ around the softmax output during ONNX / Torch-TRT export. With
this fix the `softmax_quantizer` is calibrated alongside the rest of
the model, and both the modelopt ONNX exporter and
`torch_tensorrt.compile`
pick up the Q/DQ pair. The patch short-circuits to the unwrapped call
when the quantizer is disabled (zero-overhead) and has no effect on SDPA
paths that fuse softmax inside a C++ kernel.
* **New e2e integration test** at
`tests/examples/torch_trt/test_torch_tensorrt_ptq.py` — mirrors the
`torch_onnx` test pattern: invokes the example through
`run_example_command`, parametrizes over the two precision modes (fp8,
nvfp4), uses a 1-layer ViT config (`--no_pretrained` + `--model_kwargs`)
so each parametrized case completes in under a minute. `importorskip` on
`torch_tensorrt` so the test is automatically skipped on hosts without
the package.
### Usage
```bash
# FP8 (Hopper / Ada) — default model is google/vit-large-patch16-224
python examples/torch_trt/torch_tensorrt_ptq.py \
--precision fp8 \
--calib_samples 128 \
--batch_size 1
# Custom model + custom recipe
python examples/torch_trt/torch_tensorrt_ptq.py \
--model_id <huggingface/model-id> \
--recipe <recipe-path-relative-to-modelopt_recipes-or-absolute-yaml>
```
### Testing
* Recipes load via `modelopt.recipe.load_recipe()` and pass
`QuantizeConfig` schema validation.
* Run `pytest tests/examples/torch_trt/test_torch_tensorrt_ptq.py` →
1 parametrized case passes on RTX 6000 Ada (fp8).
* End-to-end on `google/vit-base-patch16-224`: `mtq.quantize` with the
new
FP8 recipe followed by `torch_tensorrt.compile(ir="dynamo")` produces a
TRT engine whose argmax matches the FP16 baseline.
* ONNX exported from the torch path now contains Q/DQ on **12 / 12**
softmax outputs (was 0 / 12 before this PR's `_QuantAttention` fix),
matching the ONNX-CLI output's quantization layout.
Both FP8 paths land within 0.13 pp Top-1 of the FP16 baseline; Top-5 is
within 0.02 pp across all three.
* ImageNet-1k validation accuracy via the new
`torch_tensorrt_accuracy.py`
(full 50000 samples, batch=1, **every model Torch-TensorRT-compiled —
including the baseline** — so the comparison is apples-to-apples) for
the
example's default `google/vit-large-patch16-224`:
| Model (Torch-TRT) | Top-1 | Top-5 | Δ Top-1 vs baseline |
|---|---:|---:|---:|
| Baseline (FP16) | 81.99% | 96.01% | — |
| FP8 | 82.01% | 96.05% | +0.02 pp |
FP8 is within noise of the FP16 TRT baseline and NVFP4 W4A4 costs only
−0.13 pp Top-1 / −0.05 pp Top-5. Absolute Top-1 sits below the model
card's
~85.5% because evaluation uses the HF `AutoImageProcessor` default
preprocessing (direct 224×224 resize, no resize-then-center-crop),
applied
identically to all three models — so the deltas are the comparison
signal.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — new e2e integration test
under `tests/examples/torch_trt/`.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added Torch‑TensorRT FP8/NVFP4 deployment examples and end‑to‑end
scripts for HuggingFace ViT, plus ViT-specific PTQ recipes and
ImageNet-1k vs FP16 accuracy reporting.
* **Bug Fixes**
* Fixed softmax quantization and export/compilation edge cases (softmax
calibration during export, IO casting for empty tensors, routed expert
weight syncing, importer key handling).
* **Documentation**
* Added comprehensive example README with setup, usage, recipes,
evaluation, and hardware guidance.
* **Requirements**
* Pinned minimum versions for example dependencies.
* **Tests**
* Added tests validating the Torch‑TensorRT quantization examples for
fp8.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
43c20342bf |
Update specdec_bench codeowner group
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
f335459dc0 |
refactor(examples): rename llm_ptq → hf_ptq (symlink for back-compat) (#1759)
## What does this PR do? **Type of change:** refactor / deprecation (examples) Follow-up to #1705 (which consolidated `examples/vlm_ptq` into `examples/llm_ptq`). Since that example now covers Hugging Face **LLM and VLM** PTQ, the `llm_ptq` name is a misnomer. This renames the directory to `examples/hf_ptq` and leaves a relative symlink `examples/llm_ptq → hf_ptq` so existing paths/commands keep working during a deprecation window. Requested by @kevalmorabia97 on #1705 (with the symlink-for-back-compat approach), targeted for the **same 0.46 release** as the consolidation. ### Changes - `git mv examples/llm_ptq → examples/hf_ptq` and `tests/examples/llm_ptq → tests/examples/hf_ptq` (the CI runner maps the matrix name to both `examples/<name>` and `tests/examples/<name>`). - Add a tracked back-compat symlink `examples/llm_ptq → hf_ptq`. - Update CI matrices and all repo **path references** (docs, READMEs, agent skills, launcher/debugger tools, tests) from `llm_ptq` to `hf_ptq`. - Keep Python identifiers / test-util module names (`run_llm_ptq_command`, `llm_ptq_utils`) — they name the LLM-PTQ task, not the directory. - Preserve the CODEOWNERS team slug (`modelopt-examples-llm_ptq-codeowners`) and historical CHANGELOG entries; add a CHANGELOG deprecation note. ### Back-compat caveats (inherent to git directory symlinks) - ✅ Linux/macOS CLI usage and Python `cwd`/pytest resolution work through the symlink. - ⚠️ Windows git checkouts don't materialize symlinks by default (low impact — this example is Linux-only in practice). - ⚠️ GitHub web doesn't follow directory symlinks, so legacy external deep-links to `examples/llm_ptq/...` won't navigate in. All **internal** references are repointed to `hf_ptq`, so the symlink is only for legacy external/CLI use. ### Usage (unchanged via symlink) ```bash # New canonical path cd examples/hf_ptq scripts/huggingface_example.sh --model <hf_model> --quant fp8 # Old path still works (forwards via symlink) cd examples/llm_ptq && scripts/huggingface_example.sh --model <hf_model> --quant fp8 ``` ### Testing - `bash -n` on moved/edited shell scripts (new path + via symlink). - `py_compile` on moved/edited Python; test re-export shim repointed to `examples/hf_ptq/example_utils`. - Verified git tracks `examples/llm_ptq` as a single symlink (mode 120000), not a duplicated tree (no pre-commit / pytest double-processing). - `pre-commit run` on all changed files passes. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (relative symlink keeps `examples/llm_ptq` paths valid; see caveats above) - Did you write any new necessary tests?: N/A (pure rename; existing tests moved with the dir) - Did you update Changelog?: ✅ ### Additional Information Follow-up (later release): remove the `examples/llm_ptq` symlink once external references have migrated. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * PTQ guidance now directs to the unified Hugging Face PTQ flow, including VLM quantization via the shared `--vlm` entry point. * **Documentation** * Updated README and guide links, references, and command snippets to use `hf_ptq` (replacing `llm_ptq`). * Deprecated and consolidated `vlm_ptq` into `hf_ptq`; removed VILA/NVILA coverage from the Hugging Face PTQ examples. * **Bug Fixes** * Improved detection and routing so local/manual setup uses the correct PTQ source. * **Tests / Chores** * CI and example tests updated to run the `hf_ptq` variants. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
d5962c4f3b |
Remove deprecated examples/llm_autodeploy (#1797)
Remove the AutoQuant + TensorRT-LLM AutoDeploy example, deprecated in 0.45, after the migration period. Record the removal under the 0.46 Backward Breaking Changes section. Users should use TensorRT-LLM's AutoDeploy directly together with ModelOpt PTQ in examples/llm_ptq. ### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated AutoDeploy guidance to a workflow: quantize with ModelOpt PTQ (via `llm_ptq`) to produce a unified Hugging Face checkpoint, then deploy with TensorRT-LLM AutoDeploy. * Revised Hopper notes to recommend using FP8 (and refreshed related optimization guidance). * **Breaking Changes** * Removed the deprecated `examples/llm_autodeploy` example and documented the new recommended approach. * **Chores** * Dropped obsolete example docs, scripts, and coverage; adjusted example test workflow to exclude `llm_autodeploy`; updated ownership mapping for the removed example path. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
33bfa8b1fe |
CI/Dev env bump (#1818)
### What does this PR do? Type of change: chore Bumps CI/dev tooling and test containers. **Container bumps** - NeMo test containers → 26.06 - TRT-LLM container → 1.3.0rc19 - transformers max version → 5.12 **Dev tooling bumps** - ruff bump 0.12.11 → 0.15.18 - mypy 1.17.1 → 2.1.0: enable new defaults (`local_partial_types`, `strict_bytes`); fix/narrow the errors newly surfaced by mypy 2.0 in 4 modules (rather than blanket-suppressing them); remove 2 stale `# type: ignore` comments - pre-commit 4.3.0 → 4.6.0 - sphinx 8.1 → 9.1 + sphinx-rtd-theme 3.0 → 3.1: add `suppress_warnings = ["ref.python"]` to fix cross-reference ambiguity error new in sphinx 9.x - trl fix for newly released 1.7 version **Bug fixes surfaced by the bumps** - sparsity (weight): make the weight mask DTensor-aware under FSDP. The transformers→5.12 bump routes the HF Trainer FSDP optimizer-state save through torch's DTensor-based `get_optimizer_state_dict`, which triggered `aten.mul.Tensor got mixed torch.Tensor and DTensor` in the dynamic `weight` getter. The mask is now distributed to the weight's mesh/placements before masking, cached, and rebuilt only when the sharding changes (invalidated on `set_mask`). Fixes the `llm_sparsity` example test. ### Testing - `pre-commit run --all-files` ✅ (including mypy 2.1.0) - `nox -s docs` ✅ - `tests/unit/torch/sparsity` + `tests/unit/torch/nas` ✅ - `llm_sparsity` GPU example test (FSDP path) verified in CI ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Summary * **Documentation** * Refreshed Docker pre-requisites across examples to recommend updated container image tags (and streamlined some instructions). * **Bug Fixes** * Improved sparse weight mask handling for DTensor/FSDP by aligning and caching distributed masks. * Made TensorRT engine byte retrieval return immutable `bytes`. * Reduced Sphinx cross-reference warnings and tuned Transformers compatibility warning thresholds. * **Tests** * Increased default unit test timeout on Windows runners. * **Chores** * Updated CI workflow container tags and refreshed linting/typing/docs version pins, plus related mypy configuration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
6c32c37671 |
refactor(examples): consolidate vlm_ptq into llm_ptq (#1705)
### What does this PR do? Type of change: refactor / deprecation (examples) `examples/vlm_ptq` was effectively a thin wrapper over `examples/llm_ptq`: its `scripts/huggingface_example.sh` already sourced `llm_ptq/scripts/parser.sh` and called `llm_ptq/hf_ptq.py`, and all the actual VLM logic (vision-tower exclusion, `--calib_with_images`, Nemotron VL calibration, VILA loading, multimodal export) already lives under `llm_ptq`. The wrapper also referenced a `requirements-vila.txt` that did not exist in the repo. This PR makes `llm_ptq` the single source of truth for both LLM and VLM PTQ and deprecates `vlm_ptq`. **`llm_ptq` (canonical):** - Add `--vlm` and `--calib_with_images` flags to `scripts/parser.sh` and `scripts/huggingface_example.sh`. `--vlm` bootstraps VILA dependencies and runs the TensorRT-LLM multimodal quickstart as the deploy smoke test (instead of the text-only `run_tensorrt_llm.py`). - Add `examples/llm_ptq/requirements-vila.txt` (fixes the previously broken reference). - Document the VLM support matrix and the `--vlm` workflow in `README.md`. **`vlm_ptq` (deprecated):** - Replace `scripts/huggingface_example.sh` with a shim that prints a deprecation warning and forwards to the `llm_ptq` script with `--vlm`. - Convert `README.md` into a redirect/migration notice. - Repoint root `README.md` VLM links and add a `CHANGELOG.rst` deprecation entry. ### Usage ```bash cd examples/llm_ptq # VLM PTQ (was: examples/vlm_ptq/scripts/huggingface_example.sh) scripts/huggingface_example.sh --model <hf_model> --quant fp8 --vlm # VLM image-text calibration scripts/huggingface_example.sh --model <hf_model> --quant nvfp4 --vlm --calib_with_images --trust_remote_code ``` ### Testing - `bash -n` syntax check on the modified `parser.sh`, `llm_ptq` script, and the `vlm_ptq` shim. - `pre-commit run --files <changed files>` passes. - The existing VLM example test (`tests/examples/vlm_ptq/test_qwen_vl.py` via `run_vlm_ptq_command`) still exercises the path end-to-end through the deprecation shim, which forwards to the consolidated `llm_ptq` script. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (old `vlm_ptq` entry point still works via a forwarding shim) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (existing VLM test still covers the consolidated path) - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Follow-up (later release): remove the `examples/vlm_ptq` directory and its CI matrix entry once external references have migrated. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added VLM quantization support to `examples/llm_ptq` via a `--vlm` flag. * Enabled image-text pair calibration with `--calib_with_images`, including VLM multimodal smoke-test coverage. * **Deprecations** * `examples/vlm_ptq` is deprecated; it now forwards to the `examples/llm_ptq --vlm` flow with a warning. * VILA/NVILA VLM support was removed from `examples/llm_ptq` due to a model dependency compatibility conflict. * **Documentation** * Updated READMEs and the model support matrix with VLM quantization behavior and export limitations. * **Tests / CI** * Updated VLM PTQ tests and CI workflow matrices to stop running the deprecated `vlm_ptq` example. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
07ce8e5817 |
ci: speed up Claude PR review to cut timeouts (#1753)
### What does this PR do? Type of change: Bug fix (CI/workflow) The `Claude Code Review` job frequently hits the 30-min timeout. I analyzed several recent runs to find why and tuned the workflow to reduce it. **Root cause (from log analysis):** the dominant cost is the inference proxy **throttling the job** — `429`/`503`/`529` responses trigger SDK exponential backoff, which ate **~22 of 30 minutes** in the timeout run. It is *not* PR size: the same 6-file PR ranged **12.5m → 23.7m → 30.5m timeout**. Turn count is also not the constraint — a **71-turn** run finished in **11.5m** while a **25-turn** run timed out at 28m (slow turns scattered from the start, consistent with throttling, not context growth). The lever within the workflow's control is **total request volume** (each tool round-trip is a separate uncached, rate-limited request). This PR reduces it: - **Batch independent tool calls** into single turns - **Conditional pre-reads** — only read a sub-package's `mode.py`/`config.py`/`__init__.py` when the diff touches registration/config/public API (was unconditional) - **Scoped diffs** instead of one giant `gh pr diff`; **hunks-only reads**; don't re-read lines the diff already shows - **Prioritize** `modelopt/` > `examples/` > `tests/`; cap large (>50-file) PRs at ~15 files - **Restrict cross-file symbol tracing** to new/renamed public symbols - **Post findings as you go** so partial runs still deliver value - **Block subagent fan-out** (`--disallowedTools "Task"`); **drop `--max-turns`** (the data showed it was the wrong lever) ### Testing Workflow-only change; validated by the run-log analysis described above. Real effect will be visible on the next `/claude review` invocations. ### Out of scope The dominant remaining fix is **infra-side** and not addressable here: raise the proxy rate-limit/quota and re-enable prompt caching (`DISABLE_PROMPT_CACHING`). Tracking separately with the proxy team. - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (CI workflow change) - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: ❌ (pending) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Chores** * Improved the automated code review workflow’s handling of review requests and review scope, including safer treatment of user-provided comment text. * Updated the review strategy to use tighter, prioritized investigation and reduce redundant reads, with safeguards for large pull requests. **Note:** Internal infrastructure change only—no direct impact on end-user features or functionality. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0f61c984a2 |
Add missing CODEOWNER fields
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
9f37fe1969 |
feat(tools/mcp): MCP server for ModelOpt launcher (OMNIML-5123) (#1701)
## Summary `tools/mcp/` — a new MCP server exposing the existing `tools/launcher/core.py` orchestration as **typed MCP tools** that codex / Claude Code agents can call directly, instead of shelling out to `uv run launch.py --yaml ...` and parsing prose output. Tracked under [OMNIML-5123](https://jirasw.nvidia.com/browse/OMNIML-5123) (Epic). Ships **Phase 1 + Phase 1.5** together: the core launcher surface plus the four highest-leverage helpers from the `cell.md` simplification loop ([OMNIML-5128](https://jirasw.nvidia.com/browse/OMNIML-5128) partial, [OMNIML-5132](https://jirasw.nvidia.com/browse/OMNIML-5132) full). ## Nine tools **Phase 1 — core launcher surface:** | Tool | Description | |---|---| | `list_examples` | Enumerate `tools/launcher/examples/` with model + description metadata extracted from each YAML | | `verify_setup` | Fail-fast probe for the named executor. Docker: `docker info` (daemon up) + `docker info --format` runtime-registry check for the `nvidia` runtime — no image pull, daemon-fast. Slurm: `ssh -o BatchMode=yes -o ConnectTimeout=5` to the cluster login node. ~1 s probe saves 30+ s of wasted submission on bad config | | `submit_job` | Submit a launcher YAML. Mode is determined by mutually-exclusive args: `hf_local` → Docker (local GPU), `cluster_host` → Slurm (remote SSH). Returns experiment_id immediately; the actual job runs detached | | `job_status` | Filesystem-based status from nemo_run's experiment dir (`_DONE`, `status_*.out`) — no in-memory registry, survives MCP server restarts | | `job_logs` | Read `log_<task>.out` from experiment dir; per-task filtering + optional tail | **Phase 1.5 — `cell.md` simplification (OMNIML-5128 / 5132):** | Tool | Description | |---|---| | `wait_for_experiment` | Replaces the agent's `while True: status; sleep` poll loop with one tool call. Reuses `job_status_impl` so terminal-state semantics stay identical. Returns final status plus `waited_seconds`; on timeout returns structured `{ok: False, reason: "wait_timeout", last_status: …}` | | `provision_passwordless_ssh_dry_run` | No-side-effect inspection of `~/.ssh/` that emits the exact `ssh-keygen` / `ssh-copy-id` commands the operator should run to make `verify_setup(executor='slurm')` pass. Closes the verify_setup "ssh_auth_failed → now what?" gap | | `read_cluster_artifact` | Uses nemo_run's tunnel primitives, not a reinvented SSH layer. `path=None` wraps `nemo experiment logs <id> <job_idx>` (built-in log fetch); with a `path`, uses the experiment's `Tunnel` to read the file. Structured failure on subprocess error / timeout | | `open_draft_pr` | `git push -u origin HEAD` + `gh pr create --draft …`. Validates cwd is a git repo first; on gh failure after push succeeds, reports `branch_pushed=True` so the operator can retry just the PR-open step | ## Design constants 1. **Single `submit_job` with mode by args** (not separate `submit_docker` / `submit_slurm` tools). Keeps the LLM tool catalog compact; mutual-exclusion is a runtime check. 2. **Filesystem is the source of truth** for status + logs. No in-memory registry. Survives MCP server restarts cleanly — important because operators / agents kill + restart their hosts often. 3. **`verify_setup` is auto-called by `submit_job`** by default (skippable when caller just probed). The probe is ~1 s; the cost of a misconfigured submission is 30+ s of cluster timeout or container-pull. Always-on verify pays back immediately. 4. **Delegate to nemo_run for tunnels.** `read_cluster_artifact` and `wait_for_experiment` use nemo_run's existing `Experiment` / `Tunnel` / `nemo experiment logs` primitives rather than reinventing SSH/rsync. One source of truth for cluster I/O. ## Layout ``` tools/mcp/ ├── pyproject.toml # name: modelopt-mcp, console_script ├── modelopt_mcp/ │ ├── __init__.py │ ├── server.py # FastMCP entry; 9 tool definitions │ └── bridge.py # thin wrapper over launcher's core.py │ # + filesystem status/log helpers │ # + tunnel/PR helpers (Phase 1.5) └── tests/ └── test_bridge.py # 34 unit tests, fully hermetic # (mocked subprocess + tmp_path fixtures) ``` ## Install Two paths, both **from source via uv**. No PyPI wheel exists; OMNIML-5123 opted for the uvx-from-git pattern to skip publication overhead. ### End-user install (recommended) `uvx` from the git subdirectory — single command, no manual clone: ```bash # Claude Code claude mcp add modelopt -- uvx --from \ "git+https://github.com/NVIDIA/Model-Optimizer.git#subdirectory=tools/mcp" \ modelopt-mcp # Codex codex mcp add modelopt -- uvx --from \ "git+https://github.com/NVIDIA/Model-Optimizer.git#subdirectory=tools/mcp" \ modelopt-mcp ``` Under the hood `uvx` clones the whole repo to its cache, installs `tools/mcp/` as the entry, and resolves the sibling `modelopt-launcher` dep via `[tool.uv.sources]` (`path = "../launcher"`) inside the cloned tree. ### Dev install (local checkout) ```bash uv pip install -e tools/launcher # sibling dep first uv pip install -e tools/mcp # then this package modelopt-mcp # entry on PATH ``` ### Why no plain `pip install` today Two specific reasons, worth flagging so reviewers know what's intentional vs missing: 1. **Nothing on PyPI yet.** Neither `modelopt-mcp` nor `modelopt-launcher` are published — this PR introduces the package but doesn't add release machinery. 2. **`pip` doesn't read `[tool.uv.sources]`.** Even from a local checkout, plain `pip install -e tools/mcp` fails because `modelopt-launcher` is a bare name (no URL) and pip can't find it. Sticking with `uv` / `uvx` is the practical path while we're git-only. If we later want plain-pip support: publish to PyPI, or switch to a PEP-440 direct URL (`"modelopt-launcher @ git+…#subdirectory=tools/launcher"`). Out of scope for this PR. ## Post-review changes Addressed all CodeRabbit + claude[bot] review findings on the original Phase-1 surface. See the inline replies for details; the substantive bug-fix highlights: * **Slurm `cluster_host`** — propagate via `env=child_env` (launch.py reads SLURM_HOST, not a CLI arg) * **`shlex.quote`** removed from nemo-run k=v overrides (subprocess list-form doesn't shell-quote) * **Docker `Popen`** now uses `stdout=DEVNULL, stderr=DEVNULL, start_new_session=True` to avoid pipe-buffer blocking * **`NEMORUN_HOME`** pinned in subprocess env so submit + status sides agree * **GPU verify** swapped from `docker run --gpus all` image-pull (slow + flaky) to `docker info --format` runtime-registry check (daemon-fast) * **Task-status word match** anchors on first word against a fixed failure-word set (no more `"fail" in "succeeded after retry; previous attempt failed"` false-positive) * **`experiment_id` regex** generalized for non-NVIDIA cluster paths * **`pyproject.toml`** dropped the unsatisfiable `modelopt-launcher` bare-name dep (launcher is a file-layout sibling, not a Python import dep) * **`Field(ge=1)`** on `job_logs.tail` * **Docstring contract** clarified (Docker returns `pid`, Slurm returns `experiment_id`) ## Validation - [x] `uv pip install -e .` succeeds (modelopt-launcher resolved transitively) - [x] 34/34 unit tests pass (`uv run python -m pytest tests/`) - [x] stdio handshake works end-to-end; `tools/list` returns all 9 with full schemas + descriptions - [x] Mode-resolution: `submit_job` correctly rejects no-executor + both-executors with structured `reason` - [x] Filesystem status: correctly classifies `done` / `failed` / `running` from `_DONE` + `status_*.out` - [x] `wait_for_experiment` short-circuits on already-terminal experiments; honors timeout without raising - [x] `provision_passwordless_ssh_dry_run` distinguishes no-key / key-only / key+pubkey cases - [x] `read_cluster_artifact` handles subprocess timeout + non-zero exit with structured reasons - [x] `open_draft_pr` reports `branch_pushed=True` on gh-failure-after-push so retries are cheap - [x] Pre-commit clean: ruff, ruff-format, mypy, bandit, license-headers ## Acceptance criteria **OMNIML-5123 (Phase 1):** - [x] `list_examples` returns all bundled YAMLs with path and model name - [x] `submit_job` with `hf_local` runs via Docker executor and returns immediately (Phase 1: returns PID; experiment_id capture in Phase 2) - [x] `submit_job` with `cluster_host`/`user` runs via Slurm executor (`detach=True`) and returns experiment_id - [x] `job_status` correctly reflects running / done / failed from nemo_run filesystem - [x] `job_logs` returns stdout for a completed job - [x] `uvx --from git+...#subdirectory=tools/mcp modelopt-mcp --help` resolves and starts - [x] Existing launcher tests unaffected (no changes to `tools/launcher/`) **OMNIML-5128 (Phase 1.5, partial):** - [x] `wait_for_experiment` blocks until terminal or timeout - [x] `read_cluster_artifact` pulls remote artifacts via nemo_run tunnel - [x] `open_draft_pr` opens a draft PR against a target repo - [ ] Capture `experiment_id` from Docker subprocess output — deferred to Phase 2 **OMNIML-5132 (Phase 1.5, full):** - [x] `provision_passwordless_ssh_dry_run` emits operator-facing commands without side effects ## Phase 2 (separate PR) * Capture `experiment_id` from Docker subprocess output (tail until nemo_run logs the id). * Extract the verify + submit helpers into a shared lib that [`nmm-sandbox-mcp`](https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/tree/main/tools/mcp) (companion server, separate repo) can consume for internal-ergonomics tools — cluster short-name → factory lookup + GitLab CI dispatch. * NEL integration ([OMNIML-5133](https://jirasw.nvidia.com/browse/OMNIML-5133)) + checkpoint introspection ([OMNIML-5134](https://jirasw.nvidia.com/browse/OMNIML-5134)). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added ModelOpt MCP server and console entrypoint; tools: list_examples, verify_setup, submit_job, job_status, job_logs, wait_for_experiment, provision_passwordless_ssh_dry_run, read_cluster_artifact, open_draft_pr; Docker and Slurm support. * **Documentation** * Expanded README with install steps, design notes, end-to-end agent example, roadmap, and repo layout. * **Tests** * Expanded unit tests covering bridge helpers, polling, SSH flows, artifact reads, and PR automation. * **Chores** * CI updated to run MCP tests; package/meta config added. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
d87f810953 |
ci: cache JIT-compiled CUDA torch extensions in GPU/example tests (#1651)
### What does this PR do?
Type of change: CI / infrastructure (build-time speedup)
ModelOpt's CUDA quantization extensions (`modelopt_cuda_ext`, `_fp8`,
`_mx`) JIT-compile via `torch.utils.cpp_extension.load()` on first use —
~110–140s **each** in a fresh container, which is the dominant cost of
the `gpu_trtllm` job and the TRT-LLM example jobs. This caches them
across runs.
The logic lives in a reusable composite action,
**`.github/actions/cache-extensions`**, used by both `gpu_tests.yml` and
`_example_tests_runner.yml`:
- Sets a **literal in-container `TORCH_EXTENSIONS_DIR`**
(`/root/.cache/torch_extensions`). `${{ github.workspace }}` can't be
used — for `container:` jobs it resolves to the *host* path, which is
mounted elsewhere (`/__w`) inside the container, so torch and the cache
step would disagree on the location.
- Caches that dir with `actions/cache`, keyed on a caller-supplied **env
discriminator** (`rtxpro6000` + container image) plus a `hashFiles` of
the kernel/loader sources — so the cache busts on any kernel change and
is scoped per arch+image.
- On an **exact hit**, **backdates the kernel sources** below the cached
objects so ninja reuses them. (Touching the *objects* instead desyncs
ninja's `.ninja_deps`, which records each output's build-time mtime →
`stored deps info out of date` → rebuild.)
Also fixes the unused `runner` default in `_example_tests_runner.yml`
(`h100` → `rtxpro6000`) so it can't seed a wrong-arch cache.
### Usage
N/A — CI only. To reuse from another job:
```yaml
- uses: ./.github/actions/cache-extensions
with:
cache-key: rtxpro6000-${{ matrix.container_image }} # GPU arch + image
```
### Testing
Validated on `gpu_trtllm`: cache hit → `ninja: no work to do` →
`test_cuda_ext*` dropped from **113s / 108s / 139s → 2.8s / 0.03s /
0.03s** (~360s saved per run). Jobs that build no extension (e.g.
`gpu_vllm`) simply skip the save.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (CI-only; key busts on
source/image change)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update Changelog?: N/A (CI infrastructure)
- Did you get Claude approval on this PR?: ❌ (pending)
### Additional Information
- Single-arch assumption: callers pass `rtxpro6000` in `cache-key`; if
the runner fleet ever mixes GPU archs, update that prefix (the cache
path is not arch-specific).
- No explicit TTL: the key is content-addressed, and GitHub auto-evicts
caches unused for 7 days (+ 10 GB/repo LRU).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
e2c3e036a3 |
Add day0-release orchestration skill with enforced gates (#1596)
### What does this PR do?
Type of change: new feature (Claude agent skill)
Adds a `day0-release` orchestration skill that chains the existing
domain skills (`ptq` → `deployment` → `evaluation` → `compare-results`)
into the day-0 release happy path, with **code-enforced gates** between
stages.
**The problem it solves.** Today the day-0 release sequence is
*model-driven* — the agent re-decides the PTQ → deploy → eval → compare
order each session, and can silently skip a gate (e.g. report scores
from an incomplete eval, or hand off a checkpoint that missed
quantization coverage). These are failure modes we hit repeatedly in
live trials. `day0-release` makes the sequence and the gates
deterministic so the common case is repeatable.
**What it is — and isn't.** It's a *conductor*, not a new instrument. It
owns the **sequence, the gates, and the accept/retry decision**; the
domain skills still own the actual work. It is **not** for single-stage
asks ("just quantize X" → `ptq`; "run MMLU on this endpoint" →
`evaluation`) — its negative trigger excludes those. It fires only for
the full goal-driven release.
**Goal it drives toward** (the documented day-0 criterion): a quantized
checkpoint smaller than the original, with <1% accuracy drop on the
standard set vs the matching baseline, plus a publish recommendation.
**Contents:**
- `.claude/skills/day0-release/SKILL.md` — the chain + gate calls +
decision logic.
- `.claude/skills/day0-release/scripts/gate_ptq.py`, `gate_run.py`,
`gate_compare.py` — deterministic gate scripts. Each is a pure decision
function (`evaluate_*`) plus a thin file-reading `main`, returning JSON
`{pass, failure_class, detail, ...}`. The `failure_class` values match
the `modelopt-agent-protocols` strawman (`MODEL_UNSUPPORTED`,
`QUANT_COVERAGE_FAILURE`, `EVAL_JUDGE_FAILED`,
`SAMPLE_ACCOUNTING_FAILED`, …).
- `.claude/skills/day0-release/tests/test_gates.py` — 21 unit tests for
the gate decision functions (GPU-free, no cluster).
- `.claude/skills/day0-release/tests/evals.json` — 5 routing + behavior
assertions.
**As-built note — `gate_ptq.py` input.** v1 takes a `--summary
<validation.json>` (size scan + hf_ptq quant summary: `source_bytes`,
`output_bytes`, `recipe`, `layer_precision_counts`, `metadata_diffs`)
rather than reading the safetensors checkpoint directly. The
`--checkpoint/--source/--recipe` args are reserved stubs; wiring the
gate to build that summary from the exported checkpoint is a follow-up.
`gate_run.py` and `gate_compare.py` likewise read small JSON summaries
the agent produces from the run artifacts.
### Usage
Ask Claude Code:
```
Release `<org>/<model>` at day-0: quantize to NVFP4, validate accuracy is within
1% of the BF16 baseline on the AA suite, and tell me if it's publishable.
Run on <cluster>.
```
The skill then runs the chain, enforcing a gate after each stage:
```text
setup ─▶ PTQ ─▶ deploy ─▶ baseline-eval ─▶ quantized-eval ─▶ compare ─▶ closeout
│ │ │ │ │
gate_ptq health gate_run gate_run gate_compare
```
| After stage | Gate | Pass condition | On fail |
|---|---|---|---|
| Setup | reachability | creds present, cluster SSH ok | SYSTEMIC →
abort |
| PTQ | `gate_ptq.py` | size ratio <1, layer coverage matches recipe,
metadata consistent | triage → PATCH / skip recipe / abort |
| Deploy | health | endpoint up + 1 successful generation | triage →
PATCH flags/TP / skip |
| Each eval | `gate_run.py` | complete, all samples scored, no
judge/parse failure | retry / triage |
| Compare | `gate_compare.py` | every task within <1% drop | REGRESSION
→ report; ANOMALOUS → human |
It returns a **decision**, not a raw artifact: `ACCEPT` (publishable,
with report) / `REGRESSION` (which tasks failed the threshold) /
`ANOMALOUS` / `INFEASIBLE` — plus the workspace path and MLflow run IDs.
### Testing
Tested in two layers — deterministic control flow (CI-able) separate
from the agentic stages (integration-level):
- **Gate-script unit tests** — `tests/test_gates.py`, **21 cases, all
passing** (`python -m pytest
.claude/skills/day0-release/tests/test_gates.py`). Covers the pass path
plus each `failure_class` branch for all three gates: `gate_compare`
(ACCEPT / REGRESSION / ANOMALOUS-on-implausible-gain /
ANOMALOUS-out-of-range / mismatched-task-sets / relative-threshold /
non-numeric-score), `gate_run` (valid / dropped-samples / judge-error /
missing-score / non-numeric-score / timeout-non-terminal /
running-non-terminal / no-tasks), `gate_ptq` (pass / not-smaller /
zero-coverage→MODEL_UNSUPPORTED / unexpected-unquantized / metadata-diff
/ unknown-recipe). No GPU or cluster needed.
- **Routing assertions** (`tests/evals.json`, 5 cases): documents
expected routing — fires on "release model X at day-0", does **not**
fire on "just quantize X" / "run MMLU on this endpoint", and the
gate-blocking / regression-reports-and-stops behaviors. These are
behavior specs for manual QA; there is no automated skill-routing
harness yet (tracked as "Remaining Work" in the design doc).
- **End-to-end integration (run on aws-pdx / B300, 2026-06-04)** — ✅ the
full chain and **all five gate types** were exercised on **real cluster
artifacts** (Qwen3-0.6B). Results below.
#### End-to-end integration results
Every gate fired correctly on real PTQ / serve / eval outputs, and
**both** failure-routing branches were exercised:
| Stage | Gate | Real artifact | Verdict | ✓ |
|---|---|---|---|---|
| Setup | reachability | aws-pdx SSH + SLURM | PASS | ✅ |
| PTQ (happy) | `gate_ptq` | FP8 ckpt, 0.50×, FP8=196 layers,
unexpected=0 | `pass:true` | ✅ |
| **PTQ (failure fixture)** | `gate_ptq` | `nvfp4_experts_only` on a
**dense** model → NVFP4=0, 196 silently-unquantized |
**`MODEL_UNSUPPORTED`** → chain **STOPS** before deploy | ✅ |
| Deploy | health | 2 vLLM endpoints, `/health`=200 + generation | PASS
| ✅ |
| Eval ×2 | `gate_run` | real run-summaries, 300/300 scored, SUCCESS |
`pass:true` | ✅ |
| Compare | `gate_compare` | arc_easy BF16=57.0 vs FP8=54.0 |
**`REGRESSION`** (drop 3.0 > 1pt) | ✅ |
- **Failure fixture (the key control-flow proof):** an experts-only
recipe on a dense model matched 0 modules, so PTQ "succeeded" in 3.5s
and exported a checkpoint with `quant_algo:null` / `quantized_layers:{}`
— every linear silently BF16. `gate_ptq` caught the zero-coverage and
routed to `MODEL_UNSUPPORTED`, so the chain **did not** deploy/eval the
bad checkpoint. This is exactly the silent coverage-miss the gate exists
to stop.
- **Caveat — ACCEPT not reached on real data:** the happy path returned
`REGRESSION`, which is the **correct** verdict — Qwen3-0.6B FP8
genuinely loses ~3pt on arc_easy (real, ~2 SE at 300 samples;
weight-only FP8 = 54.0 was no better than KV-FP8 = 54.67, so the
KV-cache wasn't the cause). Tiny models are quantization-fragile and
don't meet the <1% criterion; the gate correctly refused to rubber-stamp
it. The threshold was **not** loosened to manufacture a pass. The
`ACCEPT` terminal + publish-recommendation path therefore remains
validated by unit test (`test_compare_accept_within_threshold`) only,
not end-to-end — demonstrating it on real data needs a larger model
(e.g. Qwen3-4B/8B) where FP8 is near-lossless, deferred as optional
follow-up.
- **Infra bugs shaken out (incidental, not in this skill's code):** (1)
the ModelOpt launcher's `/hf-local` bind mount has no host dir on
aws-pdx → set `SLURM_HF_LOCAL=<lustre dir>`; (2) `lm-eval` is unusable
inside `vllm/vllm-openai:cu130-nightly` (ships transformers 5.x, which
removed `AutoModelForVision2Seq` that lm-eval imports at load) → used a
dependency-free `/v1/completions` log-likelihood scorer.
### Scope
**In v1:** the linear chain + gate scripts + `ACCEPT`/fail-with-report
outcomes. On `REGRESSION`, v1 *reports* "recipe R regressed on tasks
[...]" and stops.
**Deferred (follow-up PR):** the evaluator-optimizer recipe loop
(compare → pick next recipe → re-PTQ), which needs the `bigpareto`
integration and the shared `modelopt-agent-protocols` schema adopted on
both sides.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (new skill; no change to
existing skills)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new
deps)
- Did you write any new necessary tests?: ✅ (21 gate-script unit tests +
5 routing evals; end-to-end integration run on aws-pdx — see Testing)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ⬜ (run `/claude review`
before marking ready)
### Additional Information
Design rationale: ModelOpt Agent Skills Design doc — this skill
implements the "deterministic day-0 chain driver"
(prompt-chaining-with-gates, the code-driven orchestration pattern from
Anthropic's *Building Effective Agents*). The gate scripts double as the
data source for the Observability stage metrics, and `gate_compare.py`'s
verdict is the entry point for the deferred evaluator-optimizer recipe
loop.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added a deterministic day0-release skill: linear gated flow (setup →
PTQ → deploy → eval → compare) that halts on failed gates.
* **New Tools**
* Added CLI gates for PTQ checkpoint validation, run/evaluation
validation, and baseline-vs-candidate comparison (scale-aware
thresholds, anomaly detection, clear verdicts and exit codes).
* **Tests**
* Added unit tests covering pass/regress/anomaly outcomes, failure-class
triage, and gate scenarios.
* **Documentation**
* Added spec and changelog entry describing inputs, gating rules, triage
table, and closeout reporting.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
c78e654744 |
Skip Softmax diffusion export (#1269)
### What does this PR do?
Type of change: New Feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
Adds HuggingFace `config.json` export of skip-softmax sparse-attention
calibration for diffusion pipelines (e.g. Wan 2.2), on top of the base
skip-softmax work.
- **`_export_diffusers_checkpoint`** walks every `nn.Module` component
of a diffusers pipeline, calls `export_sparse_attention_config`, and
writes the result into that component's `config.json` under the
`sparse_attention_config` key. The sparse config lives **only** in
`config.json` — there is no standalone `sparse.yaml`.
- **`export_sparse_attention_config`** emits a `config_groups` schema
where each algorithm's parameters are nested inside its own group; only
`config_groups` and `producer` are top-level:
- skip-softmax group → `algorithm: "skip_softmax"`, `targets`, `ignore`
(layers kept dense — e.g. cross-attention + first/last blocks),
`initial_disabled_steps` (opt-in, user-set; emitted only when `> 0`),
`threshold_scale_factor` (`a * exp(b * target_sparsity)`), and
`target_sparsity`.
- N:M group → `algorithm: "sparse_softmax"` with
`sparsity_n`/`sparsity_m`, `dense_sink_tokens`, `dense_recent_tokens`
flattened into the group.
- **Deploy reader**
(`modelopt/torch/sparsity/attention_sparsity/plugins/sparse_attn_config.py`)
reads these per-group params back, keeping the export↔load round-trip
consistent.
- **Example wiring**:
`examples/diffusers/sparsity/wan22_skip_softmax.py` gains
`--export-dir`, `--skip-softmax-threshold`, and
`--initial-disabled-steps`. `--export-dir` runs
`export_hf_checkpoint(pipe, export_dir=...)` after calibration.
- Updated `CHANGELOG.rst`.
### Usage
```bash
python examples/diffusers/sparsity/wan22_skip_softmax.py \
--model-path Wan-AI/Wan2.2-T2V-A14B-Diffusers \
--calibrate --target-sparsity 0.5 --calib-size 4 \
--initial-disabled-steps 5 \
--export-dir ./wan22_skip_softmax_ckpt
```
Resulting layout — a `config.json` per component, **no `sparse.yaml`**:
```
wan22_skip_softmax_ckpt/
├── transformer/config.json # carries sparse_attention_config
├── transformer_2/config.json # carries sparse_attention_config
├── vae/ … text_encoder/ … tokenizer/ … scheduler/ …
└── model_index.json
```
A representative `config.json` entry for a diffusion transformer:
```json
"sparse_attention_config": {
"config_groups": {
"group_0": {
"algorithm": "skip_softmax",
"targets": ["WanAttention"],
"ignore": ["blocks.0.attn1", "blocks.0.attn2", "…"],
"initial_disabled_steps": 5,
"threshold_scale_factor": {
"formula": "a * exp(b * target_sparsity)",
"prefill": {"a": 1443.49, "b": 4.30}
},
"target_sparsity": {"prefill": 0.5}
}
},
"producer": {"name": "modelopt", "version": "0.45.0..."}
}
```
The N:M variant adds a second group:
```json
"group_1": {
"algorithm": "sparse_softmax",
"targets": ["WanAttention"],
"sparsity_n": 2, "sparsity_m": 4,
"dense_sink_tokens": 0, "dense_recent_tokens": 64
}
```
### Testing
- `tests/examples/diffusers_sparsity/test_sparsity.py`: baseline /
triton-baseline / fixed-threshold runs of the Wan 2.2 example, plus a
Python-API calibrate → **export** test asserting the nested
`sparse_attention_config` (`threshold_scale_factor`, `target_sparsity`,
`ignore`, `initial_disabled_steps`) and the absence of any
`sparse.yaml`.
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attention_conversion.py`
and `test_sparse_attn_config.py`: unit coverage of the per-group export
schema and the deploy-reader round-trip (writer nests → reader reads
from groups → internal mtsa config unchanged).
- Validated end-to-end on Wan 2.2 T2V-A14B: full 4-prompt / 40-step /
81-frame calibration; the exported checkpoint carries the nested schema
in both `transformer` and `transformer_2` `config.json`, and runtime
measurement shows ~47–49% tile sparsity at a 0.5 target.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ❌ The exported
`sparse_attention_config` schema was renamed and nested per-group during
0.45.x development, and the loader reads only the new layout —
checkpoints exported by earlier 0.45.x builds must be re-exported. No
released version is affected. <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
1555e6de20 |
Fix CodeCov upload issues (#1648)
Disable codecov binary validation which seems to be constantly failing
```
gpg: Signature made Tue Apr 21 19:28:03 2026 UTC
gpg: using RSA key 27034E7FDB850E0BBC2C62FF806BB28AED779869
gpg: Can't check signature: No public key
==> Could not verify signature. Please contact Codecov if problem continues
Exiting...
```
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Chores**
* Updated CI workflow notes and removed an outdated header comment.
* Added explanatory comments to the Linux job and adjusted the code
coverage upload step to use a relaxed validation mode (no other upload
settings changed).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
54ce4e09d8 |
Add Quantization Aware Distillation (QAD) to Megatron-Bridge example (#1600)
### What does this PR do? Type of change: new example **Note:** This is **part 2 of 4** (builds on #1589): - **Part 1 (#1589):** Megatron-Bridge `quantize.py` + `export.py` support and tests. - **Part 2 (this PR):** extend `distill.py` for quantization-aware distillation (QAD) — load a quantized Megatron checkpoint as the student. - **Part 3:** https://github.com/NVIDIA/Model-Optimizer/pull/1601 - **Part 4:** repeat the NVFP4 + QAD experiments on a non-Nemotron model. Extends `examples/megatron_bridge/distill.py` to initialize the student from a **Megatron checkpoint** (a quantized checkpoint from `quantize.py`, or a pruned one) via `--student_megatron_path`, enabling **Quantization Aware Distillation (QAD)**: - `--student_hf_path` still builds the student architecture; `--student_megatron_path` supplies the (optionally quantized) weights. - For a quantized checkpoint, the ModelOpt quantize mode + base weights are restored onto the **plain student before the knowledge-distillation conversion** (`restore_sharded_modelopt_state` is a no-op once a model is already converted), so the distilled checkpoint stays exportable as a quantized model with `export.py`. **Upstream dependency / workaround:** `DistillationProvider.provide()` has no seam to transform the student before the KD conversion, so this patches `provide()` at the class level (via an `id()`-keyed registry, because the provider proxies instance-attribute assignment to its teacher once the teacher is set). A companion Megatron-Bridge PR adds a first-class `DistillationProvider.student_pre_conversion_hook`; from nemo:26.06 onwards the workaround should be removed and replaced with that hook (a removal note in `distill.py` documents exactly how). ### Usage ```bash # 1) PTQ -> quantized Megatron checkpoint (part 1) torchrun --nproc_per_node 2 quantize.py \ --hf_model_name_or_path Qwen/Qwen3-8B --quant_cfg fp8 --tp_size 2 \ --export_megatron_path /tmp/Qwen3-8B-FP8-megatron # 2) QAD: distill the quantized student from the unquantized teacher torchrun --nproc_per_node 8 distill.py \ --teacher_hf_path Qwen/Qwen3-8B \ --student_hf_path Qwen/Qwen3-8B \ --student_megatron_path /tmp/Qwen3-8B-FP8-megatron \ --data_paths 1.0 tokenized/data_text_document \ --train_iters 1000 --output_dir /output/qwen3_8b_qad # 3) export the distilled quantized checkpoint (part 1) torchrun --nproc_per_node 1 export.py \ --hf_model_name_or_path Qwen/Qwen3-8B \ --megatron_path /output/qwen3_8b_qad/checkpoints \ --export_unified_hf_path /tmp/qwen3_8b_qad_fp8_hf ``` ### Testing `tests/examples/megatron_bridge/test_qad.py` (validated on a 2-GPU NeMo `26.04` container): quantize a tiny Qwen3 at TP=2 → QAD distill from the quantized student → `export.py` to a unified HF checkpoint, asserting `hf_quant_config.json` is written (proves the quantize mode survived QAD). Includes a commented-out vLLM deployment check, validated locally (full flow passes; vLLM loads the export as `quantization=modelopt`). Existing normal/Puzzletron distillation tests still pass. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: N/A (new example feature; default behavior unchanged when `--student_megatron_path` is not set) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new dependencies) - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ✅ ### Additional Information Depends on a companion Megatron-Bridge PR adding `DistillationProvider.student_pre_conversion_hook` (the upstream replacement for the class-level `provide()` workaround). The Nemotron-3 tutorial NVFP4 + QAD experiments ship in part 3. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Quantization Aware Distillation (QAD) workflow to recover accuracy of quantized Megatron students and distill from quantized checkpoints. * CLI option to initialize a distillation student from a Megatron checkpoint and a structure-only load path for bridging. * **Documentation** * Expanded runnable quantize → QAD → export guidance and best-practice tips. * **Tests** * End-to-end test validating quantize → QAD → export artifacts. * **Chores / UX** * Clearer rank-aware messages, improved tokenizer padding handling, and more consistent export behavior (fixed export dtype). <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0081473861 |
Speed up slow unit/gpu/example tests (#1616)
### What does this PR do?
Type of change: test infrastructure / test speedups + CI stabilization
Make the test suite faster, `tests/unit` hermetic, and the CI lanes
stable, without losing coverage. Most changes are mechanical test/infra
edits; the buckets below cover the diff broadly.
**Unit tests — hermetic (no HF Hub):** toy local datasets/configs + the
local tiny tokenizer (with a checked-in chat template) replace Hub
assets; `tests/unit/conftest.py` enforces offline mode. Genuinely-HF
tests moved to `tests/gpu*` (e.g. the new
`tests/gpu/torch/utils/test_dataset_utils.py`). `CONTRIBUTING.md`
documents the hermetic-unit-test expectation.
**Unit-test speedups (no coverage loss):** speculative (disable CPU
torch.compile), calibrator (fewer histogram bins), ONNX conv/dynamo
(smaller shapes + representative subset), Ruler/sparse-attention (local
tokenizer), data-parallel autoquant (world size 4→2). Shared
`tiny_tokenizer` fixture. The distributed test helper now uses a private
`spawn` context instead of mutating the global start method (avoids
cross-test contamination).
**Rarely-used autonas/fastnas tests:** heavy parametrize cases marked
`@pytest.mark.manual`, one representative kept per test (fastnas
preferred); lighter sibling tests still cover core behavior. The legacy
FSDP1 NAS distributed test is also dropped: FastNAS/AutoNAS aren't used
with either FSDP1 or FSDP2, and FSDP1 is superseded by the newer FSDP2
API — so we keep a single FSDP2 case as a sanity check and drop FSDP1,
leaving the suite leaner.
**gpu_megatron:** deduplicate distributed worker pools by world_size
within a module (saves a redundant pool spin-up in multi-pool files;
module-scoped, no cross-module reuse).
**Example tests:** reduce per-test work via args that default to current
behavior (tests pass the fast values) — torch_onnx TRT optimization
level, diffusers calibration/inference steps, eagle `sample_size`,
megatron_bridge iters/calib, llm_sparsity data slice, export
safetensors-structure `calib_size`. Also enable the recently added
`gpt-oss` example tests in CI.
**Per-test timeouts:** `pytest-timeout` with a default per-directory
timeout (60s unit / 300s gpu+example) enforced in `tests/conftest.py`
(`timeout_func_only` in `pyproject.toml`), so a new test cannot silently
exceed the budget — an unmapped test dir crashes collection. A few
inherently slow tests carry explicit higher per-test overrides
(CUDA-compile, autotune, dflash).
**CUDA kernel pre-compilation:** a dedicated `tests/gpu/_extensions`
test JIT-builds the conv3d implicit-GEMM kernel up front (collected
before the functional tests in the same process) so the one-time build
cost no longer lands on — and time out — the first functional test that
uses it. Mirrored into the `llm_ptq`/`vlm_ptq` example lanes.
**Test relocation & optional-dependency guards:** vLLM sparsity plugin
test moved to `tests/gpu_vllm` (drops the in-test `importorskip`);
diffusers-dependent unit test guarded with `importorskip("diffusers")`
for partial-install lanes; `gpt_oss` example test dir renamed to
`gpt-oss` to match the CI matrix.
**Diffusers test models:** shared model-path constants in
`tests/_test_utils/examples/models.py` consolidated/renamed and point at
tiny `hf-internal-testing` test pipes (SDXL/SD3/FLUX) so
cachify/quantize/export tests run on toy weights; `local_id`s
normalized.
**Shared dataset utils:** `examples/llm_sparsity/.../hf_pts.py` now uses
`get_dataset_dataloader` (drops the bespoke cnn_dailymail-only
`get_calib_dataloader`; supports any registered/HF/JSONL dataset,
includes attention_mask); `data_prep.py` gains `--max_samples`.
**CI workflows:** container image bumps (pytorch 26.04→26.05, TRT-LLM
rc16→rc17) and tightened lane timeouts (unit 30→15 min, gpu lanes
trimmed, onnx example lane 45 min).
**Imports at top of file:** in-function imports across the test suite
are moved to module top per the coding guideline, conservatively —
optional deps stay guarded (in-function or behind a module-level
`importorskip`) in `tests/unit` since the partial-install lane runs
without them, and build/hardware-availability imports (apex, triton,
megatron/transformer_engine, tensorrt_llm) plus `_test_utils` lazy
guards are left in place.
**Kernel warning filters:** the repeated `filterwarnings` blanket-ignore
in six `tests/gpu/torch/kernels/**` modules is consolidated into a
scoped hook in `tests/gpu/torch/kernels/conftest.py` (kernel tests only
— the rest of the suite keeps surfacing warnings).
**Eagle example speedups:** `torch.compile` (eagle recipe default) added
~2 min to every eagle training test; it's now disabled in the eagle
example tests except one smoke (`test_llama_eagle3[1-False]`), and the
downstream resume / AR-validate / export tests point at the compile-free
checkpoint. Measured: `test_ar_validate` 139s→17s, offline training
142s→22s, streaming 140s→23s — the compile path is still smoke-tested
once.
**Example lanes install editable (`-e`):** so example scripts launched
as subprocesses resolve `modelopt` to the same source path as the test
process and reuse the pre-compiled CUDA-extension cache instead of
recompiling (~2 min/test); verified in the TRT-LLM container.
**Tiny test tokenizer:** `get_tiny_tokenizer` defaults to left padding
(what decoder-LM calibration expects) and ships a terse
generation-tagged chat template — replacing a verbose ChatML one that
inflated tokenized length on the 128-vocab tokenizer and broke the
offline-PTQ example tests' `max-seq-len` filter.
**Restored Hub-download coverage:** the live (ungated) HF dataset
round-trips exercising `get_dataset_samples`' download branch now live
in `tests/gpu/torch/utils/test_dataset_utils.py` (they had been dropped
from the hermetic unit file without a counterpart).
Individual file changes not explicitly called out above fall under this
general test/CI cleanup.
### Testing
Unit + the touched gpu_megatron files validated locally; example/GPU
lanes validated in CI.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (tests + example CLI args
default to prior behavior)
- If you copied code from any other sources or added a new PIP
dependency: N/A
- Did you write any new necessary tests?: N/A (optimizes/relocates
existing tests)
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ❌ (pending)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **Chores**
* Updated container image versions for PyTorch (26.04→26.05),
TensorRT-LLM (1.3.0rc16→1.3.0rc17), and ONNX/TensorRT (26.04→26.05).
* **Tests**
* Enhanced test isolation: unit tests now run hermetically without
HuggingFace Hub access.
* Optimized test runtime via smaller model/dataset parameters and
parallel test caching.
* Added CUDA extension availability tests and extended dataset utility
coverage.
* **Documentation**
* Updated testing guidelines in `CONTRIBUTING.md` to emphasize offline
test design.
* **Chores**
* Added pytest timeout configuration and improved CI/CD workflow
efficiency with editable installs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
72df833e0b |
Add active-MoE AutoQuant cost accounting (#1497)
### What does this PR do?
• Type of change: new feature
Adds an active_moe cost model for auto_quantize effective-bits search.
This lets AutoQuant account for routed MoE expert weights by active
decode weight traffic instead of total checkpoint weight
size, using active_moe_expert_ratio = num_experts_per_tok / num_experts.
The default behavior is unchanged: cost_model="weight" still counts all
quantizable weights equally.
### Usage
import modelopt.torch.quantization as mtq
model, search_state = mtq.auto_quantize(
model,
constraints={"effective_bits": 5.0},
quantization_formats=[
mtq.NVFP4_DEFAULT_CFG,
mtq.FP8_DEFAULT_CFG,
],
data_loader=calib_dataloader,
forward_step=forward_step,
loss_func=loss_func,
cost_model="active_moe",
# Optional. If omitted, ModelOpt tries to infer this from model.config.
active_moe_expert_ratio=2 / 64,
)
The HF PTQ example also exposes:
--auto_quantize_cost_model active_moe \
--auto_quantize_active_moe_expert_ratio 0.03125
### Testing
python -m pytest tests/unit/torch/quantization/test_autoquant.py -q -k
'active_moe or quant_recipe_hparam_cost_weight'
python -m pytest tests/unit/torch/quantization/test_autoquant.py -q -k
'not data_parallel_auto_quantize'
Results:
- 4 passed
- 58 passed, 1 deselected
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added active-MoE cost model option for auto-quantization with
configurable expert ratio; API and CLI accept cost_model and
active_moe_expert_ratio
* Unified auto-quantize supports new quant format w4a16_nvfp4
* **Bug Fixes**
* Ensure labels are moved to the logits device for base models without
an lm_head
* CLI enforces valid expert-ratio range and requires active-MoE mode
when a ratio is provided
* **Tests**
* Added unit tests for active-MoE behavior, cost-weighting, ratio
handling, and search budget selection
<!-- review_stack_entry_start -->
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1497?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)
<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
|
||
|
|
b3d83ba6f0 |
Increase Claude review timeout
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
7aa0c95646 |
Add tests/gpu_vllm (#1517)
### What does this PR do? Type of change: new tests This PR adds unit tests for vLLM fakequant, specifically testing code in `modelopt/torch/quantization/plugins/vllm.py` ### Testing ``` pytest tests/gpu_vllm/torch/quantization/test_vllm_dynamic_modules.py -sv ``` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added comprehensive GPU vLLM test suite with end-to-end quantization checks and fixtures for TinyLlama, TinyQwen3-MoE, and DeepSeek V3; includes helpers to build tiny DeepSeek V3 models. * **Chores** * Updated GPU CI to use explicit container image references, added a GPU-focused test session, and adjusted test-run setup for vLLM. * **Documentation** * Documented new GPU test directory in contributing guide. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1517?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
eb5ed2df68 |
[CI] Bump torch, transformers and dev containers to latest (#1554)
- Transformers upper bound bumped from `<5.8` to `<5.10` - Enable torch 2.12 CICD testing - Bump TRT-LLM container to `1.3.0rc16` (transformers 5.5) - Use pytorch and tensorrt 26.04 containers in CICD <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated CI test container images and targeted Torch version across workflows; adjusted release CI job to use the newer torch config. * Broadened Transformers constraint in project metadata and test/dev pins. * Removed strict transformers pins from example requirements and lifted a compression dependency cap. * Raised the import-time Transformers version threshold for compatibility warnings. * **Tests** * Refactored a GPU test to collect and report validation errors and updated numeric expected baselines. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1554?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
a2c496af0d |
[OMNIML-4788] specdec_bench: configuration.json provenance + upload_to_s3 (#1531)
> [!WARNING] > **Breaking on-disk schema change (specdec_bench v1.0.0).** This PR renames the acceptance-rate metric fields across `AcceptanceRate` / `MTBench` / `SpecBench` writers: > > | Old (pre-1.0.0) | New (1.0.0) | > |---|---| > | `Request_AR` | `Request_AL` | > | `Category_AR` | `Category_AL` | > | `Average_AR` | `Average_AL` | > | — | `Joint_Acceptance_Rate` (new) | > > The renamed values were always **acceptance length** (mean tokens generated per inference step), not a rate, and the visualizer reads `*_AL`. Pre-1.0.0 runs in S3 have `*_AR` and no `Joint_AR`; they must be re-run or post-processed before comparing. The visualizer aggregates runs by `specdec_bench` major version so accidental cross-methodology comparison is blocked. ### What does this PR do? Type of change: new feature Adds reproducibility provenance to `specdec_bench/configuration.json` and ports `upload_to_s3.py` from `iputterman/specdec_bench@main` (personal-namespace fork) into upstream. This is the first PR in a multi-stage migration off Izzy's fork now that he's left the team. Tracked in [OMNIML-4788](https://jirasw.nvidia.com/browse/OMNIML-4788). **Provenance fields added to configuration.json** (alongside existing argv / engine_version / gpu / python_version): - `specdec_bench_version` — methodology semver declared in `specdec_bench/__init__.py`. Bump minor on additive metrics, major on changed metric *definitions*. The visualizer (Phase 4 of the migration) will aggregate runs by major version so plots don't accidentally compare across methodology changes. - `specdec_bench_sha`, `modelopt_sha`, `modelopt_version`, `nmm_sandbox_sha`, `container_image` — code/runtime provenance. Each prefers an env var set by the harness (`SPECDEC_BENCH_SHA`, `MODELOPT_SHA`, `MODELOPT_VERSION`, `NMM_SANDBOX_SHA`, `CONTAINER_IMAGE`) and falls back to `git rev-parse` / `modelopt.__version__` when running standalone. The env-var preference is necessary because the runtime container has no `.git/` (the launcher packager tarballs source without git metadata) and may not have `modelopt` installed. - `checkpoint.{path, size_bytes, index_sha256, index_source}` — cheap reproducibility fingerprint that hashes `model.safetensors.index.json` (or `config.json` fallback). Changes whenever any tensor changes. - `serving_config` — engine-level config dict captured after init via a new `Model.get_serving_config()` method. VLLM dumps `AsyncEngineArgs` + the live `vllm_config.to_dict()`; SGLANG dumps the `engine_kwargs` passed to `sgl.Engine`; TRTLLM left at the base default `{}` for a later iteration. - `timestamp` — UTC ISO 8601. **Other changes** - `upload_to_s3.py` + `specdec_bench/s3_utils.py` ported from iputterman/specdec_bench@main. Recognizes run dirs by sentinel files, refuses to overwrite existing S3 prefixes. - `_redact_config` allowlists `tokenizer`, `tokenizer_path`, `tokenizer_mode`, `tokenizer_revision` so the model path stops being redacted (latent bug from substring-matching `token` ⊂ `tokenizer`). - `requirements_speed.txt`: `boto3`, `botocore` added (used by `s3_utils`). **Out of scope** (deferred to Phase 1b / Phase 2): - `--sweep_config` driver that emits per-run-dir nesting `<sweep>/<NNN_dataset_c<conc>>/` - `--s3_upload` flag baked into `run.py` itself - Launcher auto-injection of the provenance env vars (currently the example YAML sets them statically) - `container_digest` (enroot integration) and full GPU/driver inventory - TRTLLM `get_serving_config()` ### Usage ```bash # Run a smoke benchmark (Qwen3.5-4B + vLLM + MTP draft=3) — example YAML included uv run launch.py --yaml examples/Qwen/Qwen3.5-4B/specdec_bench_mtp.yaml --yes # After it lands, upload the run directory to S3: S3_KEY_ID=team-specdec-workgroup \ S3_SECRET=... \ python upload_to_s3.py /path/to/sweep_dir s3://team-specdec-workgroup/results ``` ### Testing Cluster-tested end-to-end on cw-dfw (Slurm job 11978794, NeMo Run experiment `cicd_1779403623`, ~19 min wall): - Qwen3.5-4B + vLLM + MTP draft=3 + SPEED-Bench-Internal/qualitative (80 requests) - `configuration.json` (22 KB) populated all eight new provenance fields - `Request_AR` mean 3.327 (vs 3.330 on the pre-Phase-1a run — within noise; methodology unchanged) - `upload_to_s3.py` (real upload, not dry-run) landed [s3://team-specdec-workgroup/results/qwen35_4_mtp_smoke_2026-05-21/specdec_bench_mtp/](https://app.s8k.io/buckets/team-specdec-workgroup/?prefix=results%2Fqwen35_4_mtp_smoke_2026-05-21%2F) where the visualizer at http://10.131.132.205:8080 can pick it up. ### Before your PR is "Ready for review" - Is this change backward compatible?: ✅ - `configuration.json` only gains fields. `upload_to_s3.py` / `s3_utils.py` are new files. `Model.get_serving_config()` default = `{}` so existing subclasses without an override behave as before. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - `boto3` / `botocore` are Apache 2.0 (permissive); `upload_to_s3.py` + `s3_utils.py` are ported from a private NVIDIA repo with explicit copyright headers retained. - Did you write any new necessary tests?: ❌ - Validated by cluster smoke (see Testing). Will add unit-tests for `dump_env` provenance fields and `upload_to_s3._discover_runs` in a follow-up. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Internal-facing tooling. - Did you get Claude approval on this PR?: ❌ (triggering after open) ### Additional Information Tracked in JIRA [OMNIML-4788](https://jirasw.nvidia.com/browse/OMNIML-4788). The full multi-phase plan is on that ticket's SPEC block — this PR is Phase 1a. Cherry-picked alongside the harness change are two example YAMLs (`examples/Qwen/Qwen3.5-4B/specdec_bench.yaml` for the NONE autoregressive baseline, `..._mtp.yaml` for the MTP run) that gave us cluster-test evidence. Can be split out if preferred. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an S3 upload CLI for benchmark results with dry-run and skip-existing options * Automatic capture of run configuration, provenance and redacted environment into saved config * Models now export serving configuration for reproducible runs * New launcher entrypoint and example job configs for Qwen SPEED-Bench runs * **Documentation** * README section describing S3 upload usage and supported local layouts * **Bug Fixes / Changes** * Acceptance-rate metric keys renamed in output (AR -> AL) * **Tests / CI** * New tests for redaction and S3 utilities; CI now runs specdec_bench examples <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1531?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: chenhany <chenhany@nvidia.com> |