mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
3d4d9249f4a3333f782e24fb9a830ca7a0dc5d5d
109
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3d4d9249f4 |
Gate example and GPU test lanes on the files they cover, and consolidate the CI gate (#2090)
### What does this PR do? Type of change: CI/CD improvement Follow-up to #2086, which carried the speculative-decoding fix; this PR is the CI half. **Lanes now run only when their own files change.** A one-line edit to any example started all 12 example lanes, and any `modelopt/**` change started every GPU suite. - Each lane gates itself inside `_example_tests_runner.yml` / `_gpu_tests_runner.yml`, deriving its watch list from the example or suite name it was already given (`examples/<name>/**`, `tests/examples/<name>/**`, `tests/<suite>/**`). **Adding a new example stays a one-line matrix entry** — no central mapping to update. - The five cross-example dependencies are declared as `watch_extra` next to the example that needs them: `hf_ptq` → `llm_eval` (`huggingface_example.sh` runs lm_eval from `../llm_eval`), `torch_trt` → `onnx_ptq`, `speculative_decoding` → `hf_ptq` + `dataset`, `llm_qat` / `gpt-oss` → `dataset`. - `gpu_tests.yml` is split into caller + runner to match. This shape is forced, not stylistic: job-level `if:` cannot read `matrix`, and `container:` images are pulled before any step runs, so gating inside the job would still pull 10–20 GB and hold a GPU runner for every skipped suite. - `modelopt/**`, `modelopt_recipes/**`, `pyproject.toml` and `tests/_test_utils/**` still run every lane. The gate now watches the last two, which example tests depend on but it previously ignored. **Docs-only changes no longer start GPU jobs.** The gate ignores `**.md`, `**.rst`, `**.png` and `**.ipynb` by default, so a README edit short-circuits the whole workflow. Nothing executes notebooks (no `nbmake`/`nbval` in the repo), and `.sh`/`.yaml`/`.txt` stay watched since examples run them. **One file holds the gate logic.** `.github/actions/changed-files-gate` is a composite action doing the merge-base + changed-files comparison and, optionally, the `^linux$` wait. `_pr_gate.yml` and `_wait_for_checks.yml` are both deleted: each top-level workflow keeps a 12-line `pr-gate` job that is pure wiring, and the runners use the action as steps so a lane shows `gate` + `run-test` rather than three checks. Calling a reusable workflow always materializes all of its jobs, including skipped ones, which is what made the per-lane check list noisy. `unit_tests.yml` also drops its DCO wait: DCO can be marked passing manually, so blocking the matrix on it only delayed feedback. The `^linux$` wait remains, which is the gate that actually protects GPU runners. **Also, from the original lane consolidation:** - **One TensorRT-LLM lane.** `trtllm-pr` and `trtllm-non-pr` merge into a single `trtllm` job gated like the others, so `llm_eval` now runs on PRs, where it was nightly-only. - **`gpt-oss` moves to the TensorRT-LLM image.** Its deploy step needs `tensorrt_llm`, which `pytorch` doesn't have, so `deploy_gpt_oss_trtllm` silently skipped in CI. Its `importorskip` is dropped now that the lane guarantees the dependency. - **Containers bumped where no reason was documented:** pytorch `26.06`/`26.01` → `26.07` (torch example lane, regression). Left pinned with their existing in-file reasons: pytorch `26.05` for the gpu lane (`EXPLICIT_BATCH` removed in TensorRT 11), tensorrt `26.05` for the onnx lane (`torch-tensorrt` needs `libnvinfer.so.10`), vllm `v0.20.0` (legacy FusedMoE coverage). TensorRT-LLM stays on `1.3.0rc20`: rc21–rc23 ship a `quickstart_multimodal.py` importing `MultimodalConfig` before `tensorrt_llm.llmapi` exported it (fixed upstream in NVIDIA/TensorRT-LLM#17112, one day after rc23 was cut), which fails the `hf_ptq` VLM deploy smoke test. Also clarifies the changelog line in the PR template to spell out when an entry is expected. **Two silent-failure fixes found in review:** - Every gate used `any_changed`, which is ACMR and excludes deletions, so a delete-only PR (removing an example, a test, or library code) ran nothing. Now `any_modified` (ACMRD), fixed in `example_tests.yml`, `_pr_gate.yml` and `unit_tests.yml`. - `git merge-base` was piped into `tee`, so a failure returned `tee`'s exit status and emitted an empty base, quietly changing which lanes run instead of failing. **Nightly secret scanning is fixed.** `code_quality.yml` excludes trufflehog's `lob` detector: its `(live|test)_[a-zA-Z0-9_]{35}` pattern matches any pytest function whose name is exactly 35 characters after `test_` — 60 of them in this repo, e.g. `test_all_zero_activation_yields_no_scale` — and reports them as *verified* secrets. This only ever failed nightly because the action scans all history on `schedule` and only the PR diff on `pull_request`. ### Testing Workflow YAML validated locally and the selection logic checked by hand across scenarios (`examples/diffusers/**` → onnx lane only; `examples/dataset/**` → `llm_qat`, `speculative_decoding`, `gpt-oss`; `examples/llm_eval/**` → `hf_ptq` + `llm_eval`; `modelopt/**` and nightly → everything). Gating behavior itself can only be exercised by a real PR run — the failure mode to watch for is a lane skipping when it should have run. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — CI configuration - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ — not yet run <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **CI Improvements** - Improved change detection for model recipes, test utilities, workflow actions, and documentation-only updates. - Streamlined pull request checks and status monitoring. - Consolidated TensorRT-LLM example validation and refined conditional test execution. - Updated test environments to newer PyTorch releases. - Refined GPU, regression, and unit test triggers. - Deployment tests no longer automatically skip when TensorRT-LLM is unavailable. - Updated secret scanning configuration. - **Documentation** - Updated pull request checklist guidance to include deprecations and critical bug fixes. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
22b6a148b0 |
Fix EAGLE-3 context-parallel training and re-enable its tests (#2086)
### What does this PR do?
Type of change: Bug fix
**EAGLE-3 context-parallel training (`--cp_size > 1`) is fixed, and its
tests run again.** CP has been broken since `accelerate` 1.13, and the
tests never caught it: the guard compared `Version("2.10.0a0")` against
`Version("2.10.0")`, which is False on every NGC alpha torch build, so
`test_llama_eagle3[cp_size=2]` has never actually run in CI.
Five fixes:
- **`main.py`** — rebuild the FSDP2 plugin accelerate requires for
`cp_size > 1`. The `--fsdp full_shard --fsdp_config` launcher flags that
used to supply it were dropped from `launch_train.sh`, so CP could not
start at all. Also pass the CP degree to the draft model.
- **`modeling_eagle.py`** — apply the draft model's first input norm
inside `layers[0]`'s own forward, where FSDP2 has actually unsharded its
weights, and only stash the input embeds on the path whose pre-hook
consumes them.
- **`hf_eagle.py`** — skip the dense eagle attention mask under CP
(causal masking comes from `is_causal`, TTT masking from the
ring-attention patch), and warn that padded positions are therefore
unmasked. Also stop `(eagle_loss or 0)` replacing a `0.0` loss tensor
with a plain `int`, which detached the graph.
- **`eagle_utils.py`** — key TTT-mask injection off the backward call's
`grad_out` kwarg, since newer torch omits `attn_bias` on the forward
call, silently disabling TTT masking.
- **`utils.py`** — CUDNN-only SDPA under CP; the `MATH` backend
decomposes SDPA and breaks on DTensors. Scoped to `cp_size > 1`, since
this context manager wraps every training forward and CPU has no cudnn
backend.
**Drops the `speculative_decoding` 26.01 container override.** It was
added when the lane ran 25.06 and spec-dec needed something *newer* — a
floor. Later bumps moved the default past it, so it had silently become
a ceiling holding spec-dec on a 6-month-old image.
### Testing
Ran `tests/examples/speculative_decoding` in
`nvcr.io/nvidia/pytorch:26.07-py3` on 2 GPUs, reproducing the CI install
steps (`pip uninstall -y nvidia-modelopt`, `pip install -e
".[hf,dev-test]"`, example requirements): **16 passed, 2 skipped** — the
2 skipped being pre-existing `--run-manual` tests. All four
`test_llama_eagle3` cases pass, including both `cp_size=2` ones.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — the existing `cp_size=2`
tests are re-enabled
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — not yet run
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
5dde396bdf |
Fix vLLM 0.24+ compatibility: registry TypeError and MoE RoutedExperts port (#2054)
### What does this PR do? Type of change: Bug fix vLLM 0.24 (as shipped in `nemo:26.08`) reworked the fused-MoE layer, which broke ModelOpt in two ways: 1. **`FusedMoE` became a factory function** returning a `MoERunner` pipeline. Registering it put a plain function into `QuantModuleRegistry`, so *every* later registry lookup raised `TypeError: issubclass() arg 2 must be a class, ...` — taking down large parts of `tests/gpu_megatron` (TE, Megatron chaining, MSE calibrator) that have nothing to do with vLLM. 2. **Expert weights moved onto a `RoutedExperts` submodule** of `MoERunner`, and `UnquantizedFusedMoEMethod` moved modules, so the MoE fakequant path had no valid registration target. Changes: - `_DMRegistryCls.register` now asserts keys are `nn.Module` subclasses, so a future upstream change fails at the registration site instead of as a confusing `TypeError` at lookup. - Register `RoutedExperts` (vLLM >= 0.24) while keeping the `FusedMoE` / `SharedFusedMoE` registrations for older releases; both go through the same `_QuantFusedMoEBase`. `MoERunner` calls `forward_modular` / `forward_monolithic` directly (`RoutedExperts.forward` raises by design), so those are hooked instead of `forward`. - The fused-MoE kernel patch now covers `experts.triton_moe` in addition to `fused_moe`. 0.24's launcher binds the kernel names at import time, so patching only the defining module would leave fakequant **silently inactive**. - `UnquantizedFusedMoEMethod` is resolved from either module layout. - `examples/vllm_serve/vllm_reload_utils.py`: quantizer module paths are now `mlp.experts.routed_experts.*`, so HF→vLLM expert key mapping inserts a matching `.routed_experts` infix when that layout is present. - CI `gpu_vllm` now runs on **two** containers: `v0.24.0` (first release with the FusedMoE-factory / `RoutedExperts` layout — `v0.24.1` was never released, `nemo:26.08` ships a `0.24.1.dev0` build of the same layout) and `v0.20.0`, which keeps the legacy `FusedMoE`/`SharedFusedMoE` branches covered. Test fixes for the newer vLLM (not product bugs): - The tiny Llama fixture used `hidden_size=32 / 16 heads` → `head_dim=2`, which `FLEX_ATTENTION` (the only backend available in this image) rejects with `NYI: embedding dimension ... must be at least 16`, killing the engine core at warmup. Now `head_dim=64`, matching the Qwen3-MoE fixture. - The FlashInfer metadata-builder stub used `causal=False`, which in 0.24 forces the FI-native path (`all_uses_trtllm = causal and ...`) requiring workspace buffers and real wrapper planning. Keep it on the all-TRTLLM path it was originally exercising; the stashed `_modelopt_*` fields are path-independent. ### Usage No API change — existing `mtq.quantize` / `examples/vllm_serve` flows work unmodified on both old and new vLLM. ### Testing All runs in the `nemo:26.08.rc3` container (vLLM `0.24.1.dev0+gee0da84ab`, the same 0.24.1 the CI job now pins). **Suites** - `tests/gpu_vllm`: **73 passed, 1 skipped** (was 70 passed / 3 failed). - `tests/gpu_megatron`: all pass (previously ~120 failures, all from the registry `TypeError` — TE, Megatron chaining and MSE-calibrator tests that never touch vLLM). - `tests/unit/torch/opt/test_dynamic.py`: 2 passed, including the new `test_register_rejects_non_module_classes` (rejects a factory function and a non-`nn.Module` class, and asserts no partial registration). - `pre-commit` clean on all touched files. **MoE fakequant verified by module-tree probe, not just by test assertions** After `mtq.quantize(..., NVFP4_DEFAULT_CFG)` inside the vLLM worker, every weight-owning module was enumerated on tiny DeepSeek-V3 (MLA + routed MoE + shared experts) and tiny Qwen3-MoE: ``` model.layers.0.mlp.experts.routed_experts [QuantRoutedExperts] w13_input_quantizer=3.484 w13_weight_quantizer=0.0840 w2_input_quantizer=0.1060 w2_weight_quantizer=0.0845 model.layers.0.mlp.shared_experts.gate_up_proj [QuantMergedColumnParallelLinear] ✅ model.layers.0.mlp.shared_experts.down_proj [QuantRowParallelLinear] ✅ ``` Weight amax being populated (not just input amax) means the `B is self.w13_weight` identity check and the Parameter-swap weight-fakequant branch actually execute through 0.24's kernel path — i.e. `forward_modular`/`forward_monolithic` really are the live entry points and the `experts.triton_moe` patch target is the one that fires. Unquantized modules were only the expected ones: embeddings, RMSNorms, MoE router `gate`, `lm_head`. Registration parity vs. older vLLM: Row/Column/MergedColumn/QKV `ParallelLinear` and all four attention types (`Attention`, `CrossAttention`, `EncoderOnlyAttention`, `MLAAttention`) register unchanged; `FusedMoE` → `RoutedExperts`; `SharedFusedMoE` has no counterpart because the `shared_fused_moe` module no longer exists in 0.24 — shared experts are now a plain MLP whose linears we already quantize (confirmed above). **Known gaps (pre-existing, not regressions from this PR)** - `DeepSeekV2FusedQkvAProjLinear` is not quantized: it subclasses `MergedColumnParallelLinear` but overrides `forward`, so the registry's shared-forward rule declines it. Pre-0.24 the equivalent (`q_a_proj` / `kv_a_proj_with_mqa`) were `ReplicatedLinear`, which ModelOpt never quantized — effective coverage is unchanged. - MoE fakequant hooks only the Triton expert kernels; FlashInfer/CUTLASS/DeepGEMM MoE backends bypass them (why the fixtures pin `moe_backend="triton"`). - `_setup` still requires a plain `UnquantizedFusedMoEMethod`; a `FusedMoEModularMethod` swap (some DP/all2all configs) still asserts. **Not covered by tests** - `examples/vllm_serve/vllm_reload_utils.py` — the expert key mapping is now asserted in `test_tiny_qwen3_moe_quantize` against the quantizer module paths of a booted MoE model, so a stale infix fails loudly instead of silently serving uncalibrated experts. The rest of the reload path is still inspection-only. Note the registry key moved `vllm_FusedMoE` → `vllm_RoutedExperts` and quantizer paths gained `.routed_experts`, so a `modelopt_state` saved under an older vLLM will not restore onto 0.24 as-is. - The legacy `FusedMoE`/`SharedFusedMoE` branches are covered by the second CI entry; the `v0.20.0` job is green on this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — all new paths are feature-detected; older vLLM keeps the `FusedMoE` registration. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — existing `tests/gpu_vllm` coverage exercises the new registration path. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ ### Additional Information Found while bumping the Megatron test environment from `nemo:26.06` to `nemo:26.08.rc3`. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added compatibility for newer vLLM MoE implementations, module layouts, and routed-expert configurations. * Improved model reload support for models using routed-expert submodules. * **Bug Fixes** * Improved detection and patching of vLLM MoE execution paths across supported configurations. * Registry validation now rejects invalid module registrations without partially applying changes. * **Tests** * Expanded GPU coverage for vLLM 0.24.0, dynamic module validation, and causal attention metadata paths. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
c062b3f829 |
Add sidecar GPU/CPU memory+utilization monitor for HF PTQ (#2000)
### What does this PR do?
**Type of change:** New feature (developer tool) + new tests
Adds a **standalone, cross-process resource monitor** —
`tools/resource_monitor.py` —
that samples a workload's GPU and CPU usage *from the outside* while it
runs, so
we can verify per-run memory budgets (e.g. the single-GPU layerwise PTQ
target of
≤80 GB GPU / ≤80 GB CPU for OMNIML-4947) without instrumenting
`hf_ptq.py` itself.
The monitor wraps any command, samples at a fixed interval, and on exit
writes a
CSV timeseries plus a peak/mean/min summary:
- **GPU** (per device, via NVML with an `nvidia-smi` fallback): used
memory,
utilization %, power draw (W), temperature (°C).
- **CPU** (via `psutil`): system total/used/free memory + utilization %,
and the
monitored process tree's RSS + CPU %.
It is opt-in from `examples/hf_ptq/scripts/huggingface_example.sh` via
`MODELOPT_MEM_MONITOR=1` (off by default → byte-for-byte identical
behavior). When
enabled it wraps the `hf_ptq.py` run and writes the trace/summary to a
**sibling**
`${SAVE_PATH}_mem_monitor/` directory, kept out of the exported
checkpoint that is
uploaded and consumed downstream.
**Files:**
- `tools/resource_monitor.py` — the sidecar (NVML + `nvidia-smi`
fallback; `psutil`).
- `tests/unit/tools/test_resource_monitor.py` — CPU-only unit tests (run
in the `unit` nox lane).
- `examples/hf_ptq/scripts/huggingface_example.sh` — opt-in
`MODELOPT_MEM_MONITOR=1` wrapper.
- `examples/hf_ptq/requirements.txt` — adds `psutil`.
- `pyproject.toml` — adds `psutil` to the `dev-test` extra
(deterministic import in the unit lane).
- `.github/workflows/unit_tests.yml` — adds `tools/resource_monitor.py`
to the unit-test path filters.
#### Why a new tool instead of extending
`modelopt/torch/utils/memory_monitor.py`?
The existing `GPUMemoryMonitor` is a fundamentally different tool and
cannot serve
this use case by extension:
| | `modelopt.torch.utils.memory_monitor.GPUMemoryMonitor` |
`tools/resource_monitor.py` (this PR) |
|---|---|---|
| Scope | **In-process** thread inside the workload | **Cross-process**
— wraps an external command |
| Survives workload OOM/SIGKILL | ❌ dies with the process | ✅ keeps
sampling, still writes the summary |
| Import cost | Pulls `torch` (~19 s) — lives in the workload |
Torch-free (`psutil`+`pynvml`, ~0.03 s) |
| Metrics | GPU device memory only | GPU mem/util/**power/temp** +
**CPU** mem/util + process-tree RSS |
| Output | In-memory / logs | CSV timeseries + peak/mean/min summary |
Merging the two would force `torch` into a standalone sidecar (defeating
the point)
or split the in-process monitor's threading model. A future refactor may
factor out
a **shared torch-free sampling core with two thin frontends**
(in-process + sidecar);
that is tracked as a follow-up rather than blocking this monitoring
harness, which
PR #2008 (single-GPU disk-offload PTQ) depends on.
### Usage
```bash
# Wrap mode (preferred): monitor exits with the workload's return code
python tools/resource_monitor.py --gpus 2,3 --out mem.csv --summary peak.txt -- \
python hf_ptq.py --pyt_ckpt_path=<model> --qformat=nvfp4 ...
# Opt-in from the HF PTQ example (off by default):
MODELOPT_MEM_MONITOR=1 CUDA_VISIBLE_DEVICES=2,3 CUDA_DEVICE_ORDER=PCI_BUS_ID \
bash examples/hf_ptq/scripts/huggingface_example.sh <args>
# -> writes ${SAVE_PATH}_mem_monitor/mem_trace.csv and mem_peak.txt
```
### Testing
- **Unit (CPU-only, in the `unit` nox lane):** `pytest
tests/unit/tools/test_resource_monitor.py`
— 11 tests covering `--gpus` parsing (CSV + space-separated, UUID/MIG
rejection),
the disabled/`nvidia-smi` sampling paths (including `[N/A]` → `None` and
the
smi-failure-yields-empty guard), CPU sampling, the accumulator, and
end-to-end
CSV/summary + exit-code propagation in wrap mode.
- **GPU-validated** on a B200 node (GPUs 2,3): confirmed the `gpu{i}_*`
memory /
utilization / power / temperature columns populate and the summary is
written.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — new tool; the example wrapper
is off unless `MODELOPT_MEM_MONITOR=1`.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — `psutil`
added to `dev-test` + the hf_ptq example requirements;
`pynvml`/`nvidia-ml-py` is an optional runtime dep (graceful
`nvidia-smi` fallback).
- Did you write any new necessary tests?: ✅ —
`tests/unit/tools/test_resource_monitor.py`.
- Did you update Changelog?: N/A — repo-level `tools/` script, not
shipped in the wheel.
- Did you get Claude approval on this PR?: ❌ <!-- run /claude review -->
### Additional Information
Part of **OMNIML-4947** (single-GPU disk-offload PTQ). This is PR 1 of
the stack —
the monitoring harness that PR #2008 (disk-offload layerwise PTQ +
offload-aware
export) builds on.
---------
Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
c2070cfd7a |
ci: skip docs preview deploy for fork PRs (#2029)
### What does this PR do? Type of change: Bug fix (CI) The `deploy-preview` job in the `Docs` workflow fails on **every pull request opened from a fork**, which blocks merging for all external contributors. **Root cause.** `deploy-preview` runs `rossjrw/pr-preview-action@v1`, which pushes the built HTML to the `gh-pages` branch. The workflow declares `permissions: contents: write`, but for a `pull_request` event originating from a forked repository GitHub caps the `GITHUB_TOKEN` at **read-only** — the `permissions:` block cannot elevate above that cap. The push is therefore rejected: ``` remote: Permission to NVIDIA/Model-Optimizer.git denied to github-actions[bot]. fatal: unable to access 'https://github.com/NVIDIA/Model-Optimizer.git/': The requested URL returned error: 403 ``` The job's `if:` condition gated on `github.event_name`, `github.event.action` and the `changes` path filter, but never on whether the PR came from a fork — so it always ran and always failed. Because the `changes` filter matches `docs/**`, `modelopt/**` and `.github/workflows/pages.yml`, essentially any substantive fork PR trips this. **Fix.** Restrict `deploy-preview` to PRs whose head branch lives in this repository: ```yaml github.event.pull_request.head.repo.full_name == github.repository ``` A skipped job is not a failed job, so fork PRs are no longer blocked by it. ### Usage N/A — CI-only change. ### Testing Behaviour by scenario: | Scenario | Before | After | | --- | --- | --- | | PR from a branch in this repo | preview deployed | preview deployed (**unchanged**) | | PR from a fork | ❌ fails with 403 | ⏭️ skipped | | Fork deleted (`head.repo` is `null`) | ❌ fails | ⏭️ skipped | - Confirmed against workflow history: recent `Docs` runs on in-repo branches (`main`, `chenjiel/nvfp4-act-headroom`, `mxin/qad-skill`, `haoguo/dspark-ptq-script`) all succeed, while fork-branch runs fail with the 403 above. - `build-docs` was already passing on the affected PRs — only the deploy step failed, so documentation builds are unaffected either way. - YAML parses; `pre-commit run --files .github/workflows/pages.yml` passes. (`yamlfmt` excludes `^.github/workflows/`, so this file is not auto-formatted.) - This PR edits `.github/workflows/pages.yml`, which is itself in the `changes` filter, so it exercises `deploy-preview` on the in-repo path — the preview deploy on this PR passing is a self-check that the unchanged path still works. The `closed` cleanup path is gated by the same condition. That is intentional: a fork PR never deployed a preview directory, so there is nothing to remove. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — workflow-condition change; not unit-testable - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — CI infrastructure, not user-facing - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information Currently blocking #1975 (`add Qwen3-VL support for DFlash training`), which is approved with every other check green and sits at `mergeStateStatus: BLOCKED` solely because of this job. Note that a PR only picks up this fix once its branch contains it, since workflows run from the PR branch's own definitions. A follow-up option, if doc previews for external contributors are wanted: build in the `pull_request` workflow and deploy from a separate `workflow_run`-triggered workflow, which executes in the base-repo context and does get a write token. Deliberately not using `pull_request_target` here — that would run unreviewed PR code with write permissions. Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
42458def24 |
ci: fix torch_trt on torch 2.13; default unit tests to torch 2.13; add example allow-failure hatch (#1951)
### What does this PR do? Type of change: Bug fix + CI / infra torch 2.13 + torchvision 0.28 were published to PyPI on 2026-07-08 and broke the `onnx (torch_trt)` example job (which had passed the day before). This PR fixes that break and hardens CI against the next one: 1. **Fix `torch_trt` (torch/torchvision/torch-tensorrt trio pin).** The base install pulled torch 2.13 / torchvision 0.28, then `torch-tensorrt 2.12.1` downgraded torch back to 2.12 but left torchvision at 0.28 (which pins `torch==2.13`) — breaking `import`. `examples/torch_trt/requirements.txt` now caps `torch-tensorrt>=2.4.0,<2.13` + `torchvision<0.28` so the trio stays consistent (also protects direct `pip install -r` users). 2. **Unit tests default to torch 2.13.** `noxfile.py` gains `torch_213` (`torchvision~=0.28.0`); the required `linux`/`windows` jobs and the multi-version Python spread (3.10/3.11/3.13/3.14) now run torch 2.13, with torch 2.8–2.12 kept as back-compat legs on Python 3.12. 3. **Per-example allow-failure escape hatch.** `_example_tests_runner.yml` gains an `allow_failure` input; when set, a **test-run** failure is surfaced as a `::warning::` via `continue-on-error` instead of blocking the PR. `example_tests.yml` derives it per example from the repo variable **`ALLOW_FAILURE_EXAMPLE_TESTS`** (comma-separated example names, comma-wrapped so `onnx` ≠ `torch_onnx`). Future breakages can be quarantined by updating the variable — no code change / PR required. ### Testing - Ran the new default unit session locally in an isolated uv venv (torch **2.13.0**+cu130, torchvision **0.28.0**+cu130, transformers **5.12.1**): ``` nox -s "unit-3.12(torch_213, tf_latest)" => 2813 passed, 15 skipped, 1786 warnings in 250.37s ``` - Verified the allow-failure hatch: with `ALLOW_FAILURE_EXAMPLE_TESTS=torch_trt`, the (previously failing) `torch_trt` job reports success with a warning and does not block the required example check. `vars` is re-read on each job attempt, so "Re-run failed jobs" picks up the variable without a fresh trigger. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (CI-only; older torch versions still covered) - If you copied code from any other sources or added a new PIP dependency: N/A (no new dependency; only version caps) - Did you write any new necessary tests?: N/A (CI configuration change) - Did you update Changelog?: N/A (CI infra, no user-facing API change) - Did you get Claude approval on this PR?: ❌ (pending — will run `/claude review`) ### Additional Information The `ALLOW_FAILURE_EXAMPLE_TESTS` repo variable can be cleared for `torch_trt` now that the requirements pin lands the real fix; keep it as the standing escape hatch for future example breakages. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Example test jobs can now be configured to “allow failure” without failing the workflow; when enabled, a warning annotation is emitted. * **Bug Fixes** * CI unit-test and GPU-test configurations were refreshed (including a reduced timeout for the `gpu_megatron` job). * **Tests** * Updated unit-test coverage to use the newest Torch 2.13-based setup by default, with back-compat retained where applicable. * **Documentation** * Added inline guidance for how the allow-failure examples list is specified. * **Chores** * Refreshed `torch_trt` example dependency constraints to improve compatibility. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
795c589428 |
CI: CUDA build/test hygiene + fix Puzzletron Nemotron test failures (#1901)
### What does this PR do? Type of change: Bug fix (CI / tests) - **Fix Nemotron nightly failures:** install `mamba_ssm`/`causal-conv1d` from PyPI releases instead of git `main` (avoids the broken `apache-tvm-ffi 0.1.12` that crashes on import). - **Speed up CUDA builds:** set `TORCH_CUDA_ARCH_LIST=12.0` (runner's sm_120) in the GPU/example/regression workflow container env instead of the image's ~6 archs. - **Make unit tests CPU-only:** force CUDA off in the nox `unit` env and skip JIT-compiling CUDA extensions when no GPU is usable; move the two GPU-/`mamba_ssm`-requiring unit tests to `tests/gpu`. - **Harden example tests against HF flakes:** capture subprocess output and retry transient HuggingFace access errors (5xx / rate-limit / connection). - **Skip Blackwell-flaky sharded-state-dict tests:** `test_homogeneous_sharded_state_dict` and `test_regular_state_dict[320]` intermittently hit a CUDA illegal-memory-access on the sm_120 runner that poisons the CUDA context and cascades timeouts; gate them behind a reusable `skip_flaky_on_blackwell` marker (still run on non-Blackwell GPUs). - **Bump slow test timeout:** `test_prune_minitron_vlm` → 360s for the 2-GPU nightly. ### Testing - CI tests on this PR pass (1-gpu) - Manually triggerred 2-gpu test: - GPU: https://github.com/NVIDIA/Model-Optimizer/actions/runs/28774553356 - Examples: https://github.com/NVIDIA/Model-Optimizer/actions/runs/28774556693 - Regression: https://github.com/NVIDIA/Model-Optimizer/actions/runs/28771197586 ### Additional Information - Backward compatible: N/A (CI/tests only) - New dependency: N/A - Changelog: N/A (CI/test infra) 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
d1c8ea95eb |
ci: gate example and GPU tests on changes to their own runner and cache action (#1891)
### What does this PR do? Type of change: Bug fix Adds `.github/workflows/_example_tests_runner.yml` and `.github/actions/cache-extensions/**` to the example_tests pr-gate `files:` list, and `cache-extensions/**` to the gpu_tests gate. Every example-test job executes through the reusable runner (which is taken from the PR's ref), so a PR changing only the runner previously merged with all example-test jobs skipped — observed on PR #1651, whose example-tests run shows pr-gate success with torch/trtllm-pr/megatron/onnx all skipped. Gating on the workflow's own machinery matches the existing intent (the gate already lists `example_tests.yml` itself). ### Usage N/A — CI workflow configuration change. ### Testing YAML validated; the change is additive to the gate lists only. The observed-failure evidence is PR #1651's run history (linked in #1890). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (workflow config; validated as described above) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (CI-only change) - Did you get Claude approval on this PR?: N/A (external contributor; cannot trigger `/claude review`) ### Additional Information Part of #1890 (item 1). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Expanded automated test-triggering rules to cover additional workflow and action changes. * Updated gating for example and GPU test runs so more relevant changes are validated automatically. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: arham766 <arhamislam766@yahoo.com> |
||
|
|
0c40c374e5 |
ci: give non-PR code quality runs distinct concurrency groups (#1894)
### What does this PR do? Type of change: Bug fix The concurrency group used `github.event.pull_request.number` with no fallback, so every nightly and manually dispatched run shared the literal group `Code Quality-` with `cancel-in-progress: true` — a manual dispatch cancels an in-flight nightly and vice versa. Adds the `|| github.sha` fallback already used by `unit_tests.yml`. ### Usage N/A — CI workflow configuration change. ### Testing Matches the existing pattern in `unit_tests.yml` line-for-line. YAML validated. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (workflow config; validated as described above) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (CI-only change) - Did you get Claude approval on this PR?: N/A (external contributor; cannot trigger `/claude review`) ### Additional Information Part of #1890 (item 4). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Improved workflow concurrency handling so automated runs are grouped and canceled more reliably across pull request, manual, and scheduled triggers. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: arham766 <arhamislam766@yahoo.com> |
||
|
|
ecc1e4f15f |
ci: scope bump_uv_lock change detection to uv.lock (#1892)
### What does this PR do? Type of change: Bug fix The torch-override step rewrites `pyproject.toml` via `toml.load`/`toml.dump`, which drops comments and reformats the file, so the unscoped `git diff --quiet` was always dirty and `changed=false` unreachable. In a week where `uv lock --upgrade` produces no changes, the create-pull-request step stages nothing and `git commit` exits 1, failing the scheduled run instead of no-oping. Scoping the diff to `uv.lock` restores the intended detection. ### Usage N/A — CI workflow configuration change. ### Testing Reproduced the failure locally by executing the workflow's steps in a clean checkout: after the toml mutation, unscoped `git diff --quiet` exits 1 with `uv.lock` byte-identical, and `git add uv.lock && git commit -s -m x` exits 1. With `git diff --quiet -- uv.lock`, the no-update path correctly yields `changed=false`. YAML validated. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (workflow config; validated as described above) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (CI-only change) - Did you get Claude approval on this PR?: N/A (external contributor; cannot trigger `/claude review`) ### Additional Information Part of #1890 (item 2). Signed-off-by: arham766 <arhamislam766@yahoo.com> |
||
|
|
9038b71f09 |
Autoquant and GPTQ in support in Megatron-Core [OMNIML-3151] (#1562)
### What does this PR do? Type of change: New Feature Autoquant and GPTQ in support in Megatron-Core - Add EP support to AutoQuantize - Register MCore support in AutoQuantize - Add decoder `output_layer` (lm head) to layerwise hook so that GPTQ can register all decoder layers & lm head - Split dataloader helper function out of megatron calibration utils so that AutoQuantize in Megatron-LM can reuse the same dataloader ### Usage See https://github.com/NVIDIA/Megatron-LM/pull/4821 for Autoquant usage in Megatron ```python # For GPTQ pick a recipe that uses gptq algorithm and run mtq.quantize # e.g. general/ptq/nvfp4_default-kv_none-gptq ``` ### Testing Tested AutoQuant on Nemotron Nano and Ultra. Tested GPTQ on Nano 3. Added unit tests for both AutoQuant and GPTQ ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Summary by CodeRabbit * **New Features** * Added lazy Megatron-Core AutoQuant integration with Megatron-specific quantization hooks and better decoder-layer discovery for layerwise calibration. * Improved AutoQuantize for expert-parallel (EP) models, including consistent per-layer recipe selection across DP/TP/EP. * Extended quant-layer grouping for NemotronH MCore fused “local_experts” linear layers. * **Bug Fixes** * Prevented division-by-zero when calibration inputs are empty during Hessian updates. * **Tests** * Added/extended unit and GPU coverage for EP AutoQuant, decoder-layer calibration discovery behavior, and zero-token Hessian no-op. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> Signed-off-by: Jenny Chen <jennifchen@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
fbbc5989ce |
[6078291][OMNIML-3716] Add ViT FP8 + Torch-TRT example, wire softmax_quantizer in _QuantAttention (#1569)
### What does this PR do?
Type of change: new feature + bug fix
Adds a Torch-TensorRT deployment path for HuggingFace ViT and closes the
modelopt-side gap that prevented `*softmax_quantizer` from being applied
on the standard attention forward path.
* **New ViT PTQ recipes** under `modelopt_recipes/huggingface/vit/ptq/`:
* `fp8.yaml` — W8A8 per-tensor FP8 E4M3 on encoder Linear
weights/inputs;
attention Q/K/V BMMs + softmax output at FP8; per-block LayerNorm output
at FP8 (one shared Q/DQ feeds Q/K/V + MLP); patch-embed `nn.Conv2d`,
`classifier`, and the final `vit.layernorm` left FP16. Uses max
calibration.
* The recipe is self-contained (no `$import` of shared snippets) and
use the "specific-enable" style: narrow `parent_class` + path scoping
on every enable rule, so no `enable: false` carve-outs are needed.
* **New example** under `examples/torch_trt/`:
* `torch_tensorrt_ptq.py` — single-model pipeline (load HF model,
calibrate from `zh-plus/tiny-imagenet`, `mtq.quantize`,
`torch_tensorrt.compile`, verify the compiled-model argmax matches the
fake-quant argmax). Defaults to `google/vit-large-patch16-224`; pass
`--model_id` and `--recipe` to target any model + recipe combination.
`--no_pretrained` + `--model_kwargs` shrink the model for fast tests.
* `README.md` documenting the flow, the shipped recipes, hardware
requirements, and CLI usage.
* `requirements.txt`.
* **Bug fix in `modelopt/torch/quantization/plugins/huggingface.py`** —
inside
`_QuantAttention._quantized_attention`, the non-kitchen branch now
temporarily replaces `torch.nn.functional.softmax` (via the existing
`replace_function` context manager) with a wrapper that pipes the
softmax
output through `self.softmax_quantizer`. Previously the slot was created
on every registered attention class but only consumed by the optional
Kitchen MXFP8 flash-attention path, so FP8 / NVFP4 recipes that enabled
`*softmax_quantizer` saw it stay uncalibrated (`amax=None`) and emitted
no Q/DQ around the softmax output during ONNX / Torch-TRT export. With
this fix the `softmax_quantizer` is calibrated alongside the rest of
the model, and both the modelopt ONNX exporter and
`torch_tensorrt.compile`
pick up the Q/DQ pair. The patch short-circuits to the unwrapped call
when the quantizer is disabled (zero-overhead) and has no effect on SDPA
paths that fuse softmax inside a C++ kernel.
* **New e2e integration test** at
`tests/examples/torch_trt/test_torch_tensorrt_ptq.py` — mirrors the
`torch_onnx` test pattern: invokes the example through
`run_example_command`, parametrizes over the two precision modes (fp8,
nvfp4), uses a 1-layer ViT config (`--no_pretrained` + `--model_kwargs`)
so each parametrized case completes in under a minute. `importorskip` on
`torch_tensorrt` so the test is automatically skipped on hosts without
the package.
### Usage
```bash
# FP8 (Hopper / Ada) — default model is google/vit-large-patch16-224
python examples/torch_trt/torch_tensorrt_ptq.py \
--precision fp8 \
--calib_samples 128 \
--batch_size 1
# Custom model + custom recipe
python examples/torch_trt/torch_tensorrt_ptq.py \
--model_id <huggingface/model-id> \
--recipe <recipe-path-relative-to-modelopt_recipes-or-absolute-yaml>
```
### Testing
* Recipes load via `modelopt.recipe.load_recipe()` and pass
`QuantizeConfig` schema validation.
* Run `pytest tests/examples/torch_trt/test_torch_tensorrt_ptq.py` →
1 parametrized case passes on RTX 6000 Ada (fp8).
* End-to-end on `google/vit-base-patch16-224`: `mtq.quantize` with the
new
FP8 recipe followed by `torch_tensorrt.compile(ir="dynamo")` produces a
TRT engine whose argmax matches the FP16 baseline.
* ONNX exported from the torch path now contains Q/DQ on **12 / 12**
softmax outputs (was 0 / 12 before this PR's `_QuantAttention` fix),
matching the ONNX-CLI output's quantization layout.
Both FP8 paths land within 0.13 pp Top-1 of the FP16 baseline; Top-5 is
within 0.02 pp across all three.
* ImageNet-1k validation accuracy via the new
`torch_tensorrt_accuracy.py`
(full 50000 samples, batch=1, **every model Torch-TensorRT-compiled —
including the baseline** — so the comparison is apples-to-apples) for
the
example's default `google/vit-large-patch16-224`:
| Model (Torch-TRT) | Top-1 | Top-5 | Δ Top-1 vs baseline |
|---|---:|---:|---:|
| Baseline (FP16) | 81.99% | 96.01% | — |
| FP8 | 82.01% | 96.05% | +0.02 pp |
FP8 is within noise of the FP16 TRT baseline and NVFP4 W4A4 costs only
−0.13 pp Top-1 / −0.05 pp Top-5. Absolute Top-1 sits below the model
card's
~85.5% because evaluation uses the HF `AutoImageProcessor` default
preprocessing (direct 224×224 resize, no resize-then-center-crop),
applied
identically to all three models — so the deltas are the comparison
signal.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — new e2e integration test
under `tests/examples/torch_trt/`.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added Torch‑TensorRT FP8/NVFP4 deployment examples and end‑to‑end
scripts for HuggingFace ViT, plus ViT-specific PTQ recipes and
ImageNet-1k vs FP16 accuracy reporting.
* **Bug Fixes**
* Fixed softmax quantization and export/compilation edge cases (softmax
calibration during export, IO casting for empty tensors, routed expert
weight syncing, importer key handling).
* **Documentation**
* Added comprehensive example README with setup, usage, recipes,
evaluation, and hardware guidance.
* **Requirements**
* Pinned minimum versions for example dependencies.
* **Tests**
* Added tests validating the Torch‑TensorRT quantization examples for
fp8.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
f335459dc0 |
refactor(examples): rename llm_ptq → hf_ptq (symlink for back-compat) (#1759)
## What does this PR do? **Type of change:** refactor / deprecation (examples) Follow-up to #1705 (which consolidated `examples/vlm_ptq` into `examples/llm_ptq`). Since that example now covers Hugging Face **LLM and VLM** PTQ, the `llm_ptq` name is a misnomer. This renames the directory to `examples/hf_ptq` and leaves a relative symlink `examples/llm_ptq → hf_ptq` so existing paths/commands keep working during a deprecation window. Requested by @kevalmorabia97 on #1705 (with the symlink-for-back-compat approach), targeted for the **same 0.46 release** as the consolidation. ### Changes - `git mv examples/llm_ptq → examples/hf_ptq` and `tests/examples/llm_ptq → tests/examples/hf_ptq` (the CI runner maps the matrix name to both `examples/<name>` and `tests/examples/<name>`). - Add a tracked back-compat symlink `examples/llm_ptq → hf_ptq`. - Update CI matrices and all repo **path references** (docs, READMEs, agent skills, launcher/debugger tools, tests) from `llm_ptq` to `hf_ptq`. - Keep Python identifiers / test-util module names (`run_llm_ptq_command`, `llm_ptq_utils`) — they name the LLM-PTQ task, not the directory. - Preserve the CODEOWNERS team slug (`modelopt-examples-llm_ptq-codeowners`) and historical CHANGELOG entries; add a CHANGELOG deprecation note. ### Back-compat caveats (inherent to git directory symlinks) - ✅ Linux/macOS CLI usage and Python `cwd`/pytest resolution work through the symlink. - ⚠️ Windows git checkouts don't materialize symlinks by default (low impact — this example is Linux-only in practice). - ⚠️ GitHub web doesn't follow directory symlinks, so legacy external deep-links to `examples/llm_ptq/...` won't navigate in. All **internal** references are repointed to `hf_ptq`, so the symlink is only for legacy external/CLI use. ### Usage (unchanged via symlink) ```bash # New canonical path cd examples/hf_ptq scripts/huggingface_example.sh --model <hf_model> --quant fp8 # Old path still works (forwards via symlink) cd examples/llm_ptq && scripts/huggingface_example.sh --model <hf_model> --quant fp8 ``` ### Testing - `bash -n` on moved/edited shell scripts (new path + via symlink). - `py_compile` on moved/edited Python; test re-export shim repointed to `examples/hf_ptq/example_utils`. - Verified git tracks `examples/llm_ptq` as a single symlink (mode 120000), not a duplicated tree (no pre-commit / pytest double-processing). - `pre-commit run` on all changed files passes. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (relative symlink keeps `examples/llm_ptq` paths valid; see caveats above) - Did you write any new necessary tests?: N/A (pure rename; existing tests moved with the dir) - Did you update Changelog?: ✅ ### Additional Information Follow-up (later release): remove the `examples/llm_ptq` symlink once external references have migrated. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * PTQ guidance now directs to the unified Hugging Face PTQ flow, including VLM quantization via the shared `--vlm` entry point. * **Documentation** * Updated README and guide links, references, and command snippets to use `hf_ptq` (replacing `llm_ptq`). * Deprecated and consolidated `vlm_ptq` into `hf_ptq`; removed VILA/NVILA coverage from the Hugging Face PTQ examples. * **Bug Fixes** * Improved detection and routing so local/manual setup uses the correct PTQ source. * **Tests / Chores** * CI and example tests updated to run the `hf_ptq` variants. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
d5962c4f3b |
Remove deprecated examples/llm_autodeploy (#1797)
Remove the AutoQuant + TensorRT-LLM AutoDeploy example, deprecated in 0.45, after the migration period. Record the removal under the 0.46 Backward Breaking Changes section. Users should use TensorRT-LLM's AutoDeploy directly together with ModelOpt PTQ in examples/llm_ptq. ### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated AutoDeploy guidance to a workflow: quantize with ModelOpt PTQ (via `llm_ptq`) to produce a unified Hugging Face checkpoint, then deploy with TensorRT-LLM AutoDeploy. * Revised Hopper notes to recommend using FP8 (and refreshed related optimization guidance). * **Breaking Changes** * Removed the deprecated `examples/llm_autodeploy` example and documented the new recommended approach. * **Chores** * Dropped obsolete example docs, scripts, and coverage; adjusted example test workflow to exclude `llm_autodeploy`; updated ownership mapping for the removed example path. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
33bfa8b1fe |
CI/Dev env bump (#1818)
### What does this PR do? Type of change: chore Bumps CI/dev tooling and test containers. **Container bumps** - NeMo test containers → 26.06 - TRT-LLM container → 1.3.0rc19 - transformers max version → 5.12 **Dev tooling bumps** - ruff bump 0.12.11 → 0.15.18 - mypy 1.17.1 → 2.1.0: enable new defaults (`local_partial_types`, `strict_bytes`); fix/narrow the errors newly surfaced by mypy 2.0 in 4 modules (rather than blanket-suppressing them); remove 2 stale `# type: ignore` comments - pre-commit 4.3.0 → 4.6.0 - sphinx 8.1 → 9.1 + sphinx-rtd-theme 3.0 → 3.1: add `suppress_warnings = ["ref.python"]` to fix cross-reference ambiguity error new in sphinx 9.x - trl fix for newly released 1.7 version **Bug fixes surfaced by the bumps** - sparsity (weight): make the weight mask DTensor-aware under FSDP. The transformers→5.12 bump routes the HF Trainer FSDP optimizer-state save through torch's DTensor-based `get_optimizer_state_dict`, which triggered `aten.mul.Tensor got mixed torch.Tensor and DTensor` in the dynamic `weight` getter. The mask is now distributed to the weight's mesh/placements before masking, cached, and rebuilt only when the sharding changes (invalidated on `set_mask`). Fixes the `llm_sparsity` example test. ### Testing - `pre-commit run --all-files` ✅ (including mypy 2.1.0) - `nox -s docs` ✅ - `tests/unit/torch/sparsity` + `tests/unit/torch/nas` ✅ - `llm_sparsity` GPU example test (FSDP path) verified in CI ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Summary * **Documentation** * Refreshed Docker pre-requisites across examples to recommend updated container image tags (and streamlined some instructions). * **Bug Fixes** * Improved sparse weight mask handling for DTensor/FSDP by aligning and caching distributed masks. * Made TensorRT engine byte retrieval return immutable `bytes`. * Reduced Sphinx cross-reference warnings and tuned Transformers compatibility warning thresholds. * **Tests** * Increased default unit test timeout on Windows runners. * **Chores** * Updated CI workflow container tags and refreshed linting/typing/docs version pins, plus related mypy configuration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
6c32c37671 |
refactor(examples): consolidate vlm_ptq into llm_ptq (#1705)
### What does this PR do? Type of change: refactor / deprecation (examples) `examples/vlm_ptq` was effectively a thin wrapper over `examples/llm_ptq`: its `scripts/huggingface_example.sh` already sourced `llm_ptq/scripts/parser.sh` and called `llm_ptq/hf_ptq.py`, and all the actual VLM logic (vision-tower exclusion, `--calib_with_images`, Nemotron VL calibration, VILA loading, multimodal export) already lives under `llm_ptq`. The wrapper also referenced a `requirements-vila.txt` that did not exist in the repo. This PR makes `llm_ptq` the single source of truth for both LLM and VLM PTQ and deprecates `vlm_ptq`. **`llm_ptq` (canonical):** - Add `--vlm` and `--calib_with_images` flags to `scripts/parser.sh` and `scripts/huggingface_example.sh`. `--vlm` bootstraps VILA dependencies and runs the TensorRT-LLM multimodal quickstart as the deploy smoke test (instead of the text-only `run_tensorrt_llm.py`). - Add `examples/llm_ptq/requirements-vila.txt` (fixes the previously broken reference). - Document the VLM support matrix and the `--vlm` workflow in `README.md`. **`vlm_ptq` (deprecated):** - Replace `scripts/huggingface_example.sh` with a shim that prints a deprecation warning and forwards to the `llm_ptq` script with `--vlm`. - Convert `README.md` into a redirect/migration notice. - Repoint root `README.md` VLM links and add a `CHANGELOG.rst` deprecation entry. ### Usage ```bash cd examples/llm_ptq # VLM PTQ (was: examples/vlm_ptq/scripts/huggingface_example.sh) scripts/huggingface_example.sh --model <hf_model> --quant fp8 --vlm # VLM image-text calibration scripts/huggingface_example.sh --model <hf_model> --quant nvfp4 --vlm --calib_with_images --trust_remote_code ``` ### Testing - `bash -n` syntax check on the modified `parser.sh`, `llm_ptq` script, and the `vlm_ptq` shim. - `pre-commit run --files <changed files>` passes. - The existing VLM example test (`tests/examples/vlm_ptq/test_qwen_vl.py` via `run_vlm_ptq_command`) still exercises the path end-to-end through the deprecation shim, which forwards to the consolidated `llm_ptq` script. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (old `vlm_ptq` entry point still works via a forwarding shim) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (existing VLM test still covers the consolidated path) - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Follow-up (later release): remove the `examples/vlm_ptq` directory and its CI matrix entry once external references have migrated. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added VLM quantization support to `examples/llm_ptq` via a `--vlm` flag. * Enabled image-text pair calibration with `--calib_with_images`, including VLM multimodal smoke-test coverage. * **Deprecations** * `examples/vlm_ptq` is deprecated; it now forwards to the `examples/llm_ptq --vlm` flow with a warning. * VILA/NVILA VLM support was removed from `examples/llm_ptq` due to a model dependency compatibility conflict. * **Documentation** * Updated READMEs and the model support matrix with VLM quantization behavior and export limitations. * **Tests / CI** * Updated VLM PTQ tests and CI workflow matrices to stop running the deprecated `vlm_ptq` example. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
07ce8e5817 |
ci: speed up Claude PR review to cut timeouts (#1753)
### What does this PR do? Type of change: Bug fix (CI/workflow) The `Claude Code Review` job frequently hits the 30-min timeout. I analyzed several recent runs to find why and tuned the workflow to reduce it. **Root cause (from log analysis):** the dominant cost is the inference proxy **throttling the job** — `429`/`503`/`529` responses trigger SDK exponential backoff, which ate **~22 of 30 minutes** in the timeout run. It is *not* PR size: the same 6-file PR ranged **12.5m → 23.7m → 30.5m timeout**. Turn count is also not the constraint — a **71-turn** run finished in **11.5m** while a **25-turn** run timed out at 28m (slow turns scattered from the start, consistent with throttling, not context growth). The lever within the workflow's control is **total request volume** (each tool round-trip is a separate uncached, rate-limited request). This PR reduces it: - **Batch independent tool calls** into single turns - **Conditional pre-reads** — only read a sub-package's `mode.py`/`config.py`/`__init__.py` when the diff touches registration/config/public API (was unconditional) - **Scoped diffs** instead of one giant `gh pr diff`; **hunks-only reads**; don't re-read lines the diff already shows - **Prioritize** `modelopt/` > `examples/` > `tests/`; cap large (>50-file) PRs at ~15 files - **Restrict cross-file symbol tracing** to new/renamed public symbols - **Post findings as you go** so partial runs still deliver value - **Block subagent fan-out** (`--disallowedTools "Task"`); **drop `--max-turns`** (the data showed it was the wrong lever) ### Testing Workflow-only change; validated by the run-log analysis described above. Real effect will be visible on the next `/claude review` invocations. ### Out of scope The dominant remaining fix is **infra-side** and not addressable here: raise the proxy rate-limit/quota and re-enable prompt caching (`DISABLE_PROMPT_CACHING`). Tracking separately with the proxy team. - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (CI workflow change) - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: ❌ (pending) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Chores** * Improved the automated code review workflow’s handling of review requests and review scope, including safer treatment of user-provided comment text. * Updated the review strategy to use tighter, prioritized investigation and reduce redundant reads, with safeguards for large pull requests. **Note:** Internal infrastructure change only—no direct impact on end-user features or functionality. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
9f37fe1969 |
feat(tools/mcp): MCP server for ModelOpt launcher (OMNIML-5123) (#1701)
## Summary `tools/mcp/` — a new MCP server exposing the existing `tools/launcher/core.py` orchestration as **typed MCP tools** that codex / Claude Code agents can call directly, instead of shelling out to `uv run launch.py --yaml ...` and parsing prose output. Tracked under [OMNIML-5123](https://jirasw.nvidia.com/browse/OMNIML-5123) (Epic). Ships **Phase 1 + Phase 1.5** together: the core launcher surface plus the four highest-leverage helpers from the `cell.md` simplification loop ([OMNIML-5128](https://jirasw.nvidia.com/browse/OMNIML-5128) partial, [OMNIML-5132](https://jirasw.nvidia.com/browse/OMNIML-5132) full). ## Nine tools **Phase 1 — core launcher surface:** | Tool | Description | |---|---| | `list_examples` | Enumerate `tools/launcher/examples/` with model + description metadata extracted from each YAML | | `verify_setup` | Fail-fast probe for the named executor. Docker: `docker info` (daemon up) + `docker info --format` runtime-registry check for the `nvidia` runtime — no image pull, daemon-fast. Slurm: `ssh -o BatchMode=yes -o ConnectTimeout=5` to the cluster login node. ~1 s probe saves 30+ s of wasted submission on bad config | | `submit_job` | Submit a launcher YAML. Mode is determined by mutually-exclusive args: `hf_local` → Docker (local GPU), `cluster_host` → Slurm (remote SSH). Returns experiment_id immediately; the actual job runs detached | | `job_status` | Filesystem-based status from nemo_run's experiment dir (`_DONE`, `status_*.out`) — no in-memory registry, survives MCP server restarts | | `job_logs` | Read `log_<task>.out` from experiment dir; per-task filtering + optional tail | **Phase 1.5 — `cell.md` simplification (OMNIML-5128 / 5132):** | Tool | Description | |---|---| | `wait_for_experiment` | Replaces the agent's `while True: status; sleep` poll loop with one tool call. Reuses `job_status_impl` so terminal-state semantics stay identical. Returns final status plus `waited_seconds`; on timeout returns structured `{ok: False, reason: "wait_timeout", last_status: …}` | | `provision_passwordless_ssh_dry_run` | No-side-effect inspection of `~/.ssh/` that emits the exact `ssh-keygen` / `ssh-copy-id` commands the operator should run to make `verify_setup(executor='slurm')` pass. Closes the verify_setup "ssh_auth_failed → now what?" gap | | `read_cluster_artifact` | Uses nemo_run's tunnel primitives, not a reinvented SSH layer. `path=None` wraps `nemo experiment logs <id> <job_idx>` (built-in log fetch); with a `path`, uses the experiment's `Tunnel` to read the file. Structured failure on subprocess error / timeout | | `open_draft_pr` | `git push -u origin HEAD` + `gh pr create --draft …`. Validates cwd is a git repo first; on gh failure after push succeeds, reports `branch_pushed=True` so the operator can retry just the PR-open step | ## Design constants 1. **Single `submit_job` with mode by args** (not separate `submit_docker` / `submit_slurm` tools). Keeps the LLM tool catalog compact; mutual-exclusion is a runtime check. 2. **Filesystem is the source of truth** for status + logs. No in-memory registry. Survives MCP server restarts cleanly — important because operators / agents kill + restart their hosts often. 3. **`verify_setup` is auto-called by `submit_job`** by default (skippable when caller just probed). The probe is ~1 s; the cost of a misconfigured submission is 30+ s of cluster timeout or container-pull. Always-on verify pays back immediately. 4. **Delegate to nemo_run for tunnels.** `read_cluster_artifact` and `wait_for_experiment` use nemo_run's existing `Experiment` / `Tunnel` / `nemo experiment logs` primitives rather than reinventing SSH/rsync. One source of truth for cluster I/O. ## Layout ``` tools/mcp/ ├── pyproject.toml # name: modelopt-mcp, console_script ├── modelopt_mcp/ │ ├── __init__.py │ ├── server.py # FastMCP entry; 9 tool definitions │ └── bridge.py # thin wrapper over launcher's core.py │ # + filesystem status/log helpers │ # + tunnel/PR helpers (Phase 1.5) └── tests/ └── test_bridge.py # 34 unit tests, fully hermetic # (mocked subprocess + tmp_path fixtures) ``` ## Install Two paths, both **from source via uv**. No PyPI wheel exists; OMNIML-5123 opted for the uvx-from-git pattern to skip publication overhead. ### End-user install (recommended) `uvx` from the git subdirectory — single command, no manual clone: ```bash # Claude Code claude mcp add modelopt -- uvx --from \ "git+https://github.com/NVIDIA/Model-Optimizer.git#subdirectory=tools/mcp" \ modelopt-mcp # Codex codex mcp add modelopt -- uvx --from \ "git+https://github.com/NVIDIA/Model-Optimizer.git#subdirectory=tools/mcp" \ modelopt-mcp ``` Under the hood `uvx` clones the whole repo to its cache, installs `tools/mcp/` as the entry, and resolves the sibling `modelopt-launcher` dep via `[tool.uv.sources]` (`path = "../launcher"`) inside the cloned tree. ### Dev install (local checkout) ```bash uv pip install -e tools/launcher # sibling dep first uv pip install -e tools/mcp # then this package modelopt-mcp # entry on PATH ``` ### Why no plain `pip install` today Two specific reasons, worth flagging so reviewers know what's intentional vs missing: 1. **Nothing on PyPI yet.** Neither `modelopt-mcp` nor `modelopt-launcher` are published — this PR introduces the package but doesn't add release machinery. 2. **`pip` doesn't read `[tool.uv.sources]`.** Even from a local checkout, plain `pip install -e tools/mcp` fails because `modelopt-launcher` is a bare name (no URL) and pip can't find it. Sticking with `uv` / `uvx` is the practical path while we're git-only. If we later want plain-pip support: publish to PyPI, or switch to a PEP-440 direct URL (`"modelopt-launcher @ git+…#subdirectory=tools/launcher"`). Out of scope for this PR. ## Post-review changes Addressed all CodeRabbit + claude[bot] review findings on the original Phase-1 surface. See the inline replies for details; the substantive bug-fix highlights: * **Slurm `cluster_host`** — propagate via `env=child_env` (launch.py reads SLURM_HOST, not a CLI arg) * **`shlex.quote`** removed from nemo-run k=v overrides (subprocess list-form doesn't shell-quote) * **Docker `Popen`** now uses `stdout=DEVNULL, stderr=DEVNULL, start_new_session=True` to avoid pipe-buffer blocking * **`NEMORUN_HOME`** pinned in subprocess env so submit + status sides agree * **GPU verify** swapped from `docker run --gpus all` image-pull (slow + flaky) to `docker info --format` runtime-registry check (daemon-fast) * **Task-status word match** anchors on first word against a fixed failure-word set (no more `"fail" in "succeeded after retry; previous attempt failed"` false-positive) * **`experiment_id` regex** generalized for non-NVIDIA cluster paths * **`pyproject.toml`** dropped the unsatisfiable `modelopt-launcher` bare-name dep (launcher is a file-layout sibling, not a Python import dep) * **`Field(ge=1)`** on `job_logs.tail` * **Docstring contract** clarified (Docker returns `pid`, Slurm returns `experiment_id`) ## Validation - [x] `uv pip install -e .` succeeds (modelopt-launcher resolved transitively) - [x] 34/34 unit tests pass (`uv run python -m pytest tests/`) - [x] stdio handshake works end-to-end; `tools/list` returns all 9 with full schemas + descriptions - [x] Mode-resolution: `submit_job` correctly rejects no-executor + both-executors with structured `reason` - [x] Filesystem status: correctly classifies `done` / `failed` / `running` from `_DONE` + `status_*.out` - [x] `wait_for_experiment` short-circuits on already-terminal experiments; honors timeout without raising - [x] `provision_passwordless_ssh_dry_run` distinguishes no-key / key-only / key+pubkey cases - [x] `read_cluster_artifact` handles subprocess timeout + non-zero exit with structured reasons - [x] `open_draft_pr` reports `branch_pushed=True` on gh-failure-after-push so retries are cheap - [x] Pre-commit clean: ruff, ruff-format, mypy, bandit, license-headers ## Acceptance criteria **OMNIML-5123 (Phase 1):** - [x] `list_examples` returns all bundled YAMLs with path and model name - [x] `submit_job` with `hf_local` runs via Docker executor and returns immediately (Phase 1: returns PID; experiment_id capture in Phase 2) - [x] `submit_job` with `cluster_host`/`user` runs via Slurm executor (`detach=True`) and returns experiment_id - [x] `job_status` correctly reflects running / done / failed from nemo_run filesystem - [x] `job_logs` returns stdout for a completed job - [x] `uvx --from git+...#subdirectory=tools/mcp modelopt-mcp --help` resolves and starts - [x] Existing launcher tests unaffected (no changes to `tools/launcher/`) **OMNIML-5128 (Phase 1.5, partial):** - [x] `wait_for_experiment` blocks until terminal or timeout - [x] `read_cluster_artifact` pulls remote artifacts via nemo_run tunnel - [x] `open_draft_pr` opens a draft PR against a target repo - [ ] Capture `experiment_id` from Docker subprocess output — deferred to Phase 2 **OMNIML-5132 (Phase 1.5, full):** - [x] `provision_passwordless_ssh_dry_run` emits operator-facing commands without side effects ## Phase 2 (separate PR) * Capture `experiment_id` from Docker subprocess output (tail until nemo_run logs the id). * Extract the verify + submit helpers into a shared lib that [`nmm-sandbox-mcp`](https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/tree/main/tools/mcp) (companion server, separate repo) can consume for internal-ergonomics tools — cluster short-name → factory lookup + GitLab CI dispatch. * NEL integration ([OMNIML-5133](https://jirasw.nvidia.com/browse/OMNIML-5133)) + checkpoint introspection ([OMNIML-5134](https://jirasw.nvidia.com/browse/OMNIML-5134)). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added ModelOpt MCP server and console entrypoint; tools: list_examples, verify_setup, submit_job, job_status, job_logs, wait_for_experiment, provision_passwordless_ssh_dry_run, read_cluster_artifact, open_draft_pr; Docker and Slurm support. * **Documentation** * Expanded README with install steps, design notes, end-to-end agent example, roadmap, and repo layout. * **Tests** * Expanded unit tests covering bridge helpers, polling, SSH flows, artifact reads, and PR automation. * **Chores** * CI updated to run MCP tests; package/meta config added. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
d87f810953 |
ci: cache JIT-compiled CUDA torch extensions in GPU/example tests (#1651)
### What does this PR do?
Type of change: CI / infrastructure (build-time speedup)
ModelOpt's CUDA quantization extensions (`modelopt_cuda_ext`, `_fp8`,
`_mx`) JIT-compile via `torch.utils.cpp_extension.load()` on first use —
~110–140s **each** in a fresh container, which is the dominant cost of
the `gpu_trtllm` job and the TRT-LLM example jobs. This caches them
across runs.
The logic lives in a reusable composite action,
**`.github/actions/cache-extensions`**, used by both `gpu_tests.yml` and
`_example_tests_runner.yml`:
- Sets a **literal in-container `TORCH_EXTENSIONS_DIR`**
(`/root/.cache/torch_extensions`). `${{ github.workspace }}` can't be
used — for `container:` jobs it resolves to the *host* path, which is
mounted elsewhere (`/__w`) inside the container, so torch and the cache
step would disagree on the location.
- Caches that dir with `actions/cache`, keyed on a caller-supplied **env
discriminator** (`rtxpro6000` + container image) plus a `hashFiles` of
the kernel/loader sources — so the cache busts on any kernel change and
is scoped per arch+image.
- On an **exact hit**, **backdates the kernel sources** below the cached
objects so ninja reuses them. (Touching the *objects* instead desyncs
ninja's `.ninja_deps`, which records each output's build-time mtime →
`stored deps info out of date` → rebuild.)
Also fixes the unused `runner` default in `_example_tests_runner.yml`
(`h100` → `rtxpro6000`) so it can't seed a wrong-arch cache.
### Usage
N/A — CI only. To reuse from another job:
```yaml
- uses: ./.github/actions/cache-extensions
with:
cache-key: rtxpro6000-${{ matrix.container_image }} # GPU arch + image
```
### Testing
Validated on `gpu_trtllm`: cache hit → `ninja: no work to do` →
`test_cuda_ext*` dropped from **113s / 108s / 139s → 2.8s / 0.03s /
0.03s** (~360s saved per run). Jobs that build no extension (e.g.
`gpu_vllm`) simply skip the save.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (CI-only; key busts on
source/image change)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update Changelog?: N/A (CI infrastructure)
- Did you get Claude approval on this PR?: ❌ (pending)
### Additional Information
- Single-arch assumption: callers pass `rtxpro6000` in `cache-key`; if
the runner fleet ever mixes GPU archs, update that prefix (the cache
path is not arch-specific).
- No explicit TTL: the key is content-addressed, and GitHub auto-evicts
caches unused for 7 days (+ 10 GB/repo LRU).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
e2c3e036a3 |
Add day0-release orchestration skill with enforced gates (#1596)
### What does this PR do?
Type of change: new feature (Claude agent skill)
Adds a `day0-release` orchestration skill that chains the existing
domain skills (`ptq` → `deployment` → `evaluation` → `compare-results`)
into the day-0 release happy path, with **code-enforced gates** between
stages.
**The problem it solves.** Today the day-0 release sequence is
*model-driven* — the agent re-decides the PTQ → deploy → eval → compare
order each session, and can silently skip a gate (e.g. report scores
from an incomplete eval, or hand off a checkpoint that missed
quantization coverage). These are failure modes we hit repeatedly in
live trials. `day0-release` makes the sequence and the gates
deterministic so the common case is repeatable.
**What it is — and isn't.** It's a *conductor*, not a new instrument. It
owns the **sequence, the gates, and the accept/retry decision**; the
domain skills still own the actual work. It is **not** for single-stage
asks ("just quantize X" → `ptq`; "run MMLU on this endpoint" →
`evaluation`) — its negative trigger excludes those. It fires only for
the full goal-driven release.
**Goal it drives toward** (the documented day-0 criterion): a quantized
checkpoint smaller than the original, with <1% accuracy drop on the
standard set vs the matching baseline, plus a publish recommendation.
**Contents:**
- `.claude/skills/day0-release/SKILL.md` — the chain + gate calls +
decision logic.
- `.claude/skills/day0-release/scripts/gate_ptq.py`, `gate_run.py`,
`gate_compare.py` — deterministic gate scripts. Each is a pure decision
function (`evaluate_*`) plus a thin file-reading `main`, returning JSON
`{pass, failure_class, detail, ...}`. The `failure_class` values match
the `modelopt-agent-protocols` strawman (`MODEL_UNSUPPORTED`,
`QUANT_COVERAGE_FAILURE`, `EVAL_JUDGE_FAILED`,
`SAMPLE_ACCOUNTING_FAILED`, …).
- `.claude/skills/day0-release/tests/test_gates.py` — 21 unit tests for
the gate decision functions (GPU-free, no cluster).
- `.claude/skills/day0-release/tests/evals.json` — 5 routing + behavior
assertions.
**As-built note — `gate_ptq.py` input.** v1 takes a `--summary
<validation.json>` (size scan + hf_ptq quant summary: `source_bytes`,
`output_bytes`, `recipe`, `layer_precision_counts`, `metadata_diffs`)
rather than reading the safetensors checkpoint directly. The
`--checkpoint/--source/--recipe` args are reserved stubs; wiring the
gate to build that summary from the exported checkpoint is a follow-up.
`gate_run.py` and `gate_compare.py` likewise read small JSON summaries
the agent produces from the run artifacts.
### Usage
Ask Claude Code:
```
Release `<org>/<model>` at day-0: quantize to NVFP4, validate accuracy is within
1% of the BF16 baseline on the AA suite, and tell me if it's publishable.
Run on <cluster>.
```
The skill then runs the chain, enforcing a gate after each stage:
```text
setup ─▶ PTQ ─▶ deploy ─▶ baseline-eval ─▶ quantized-eval ─▶ compare ─▶ closeout
│ │ │ │ │
gate_ptq health gate_run gate_run gate_compare
```
| After stage | Gate | Pass condition | On fail |
|---|---|---|---|
| Setup | reachability | creds present, cluster SSH ok | SYSTEMIC →
abort |
| PTQ | `gate_ptq.py` | size ratio <1, layer coverage matches recipe,
metadata consistent | triage → PATCH / skip recipe / abort |
| Deploy | health | endpoint up + 1 successful generation | triage →
PATCH flags/TP / skip |
| Each eval | `gate_run.py` | complete, all samples scored, no
judge/parse failure | retry / triage |
| Compare | `gate_compare.py` | every task within <1% drop | REGRESSION
→ report; ANOMALOUS → human |
It returns a **decision**, not a raw artifact: `ACCEPT` (publishable,
with report) / `REGRESSION` (which tasks failed the threshold) /
`ANOMALOUS` / `INFEASIBLE` — plus the workspace path and MLflow run IDs.
### Testing
Tested in two layers — deterministic control flow (CI-able) separate
from the agentic stages (integration-level):
- **Gate-script unit tests** — `tests/test_gates.py`, **21 cases, all
passing** (`python -m pytest
.claude/skills/day0-release/tests/test_gates.py`). Covers the pass path
plus each `failure_class` branch for all three gates: `gate_compare`
(ACCEPT / REGRESSION / ANOMALOUS-on-implausible-gain /
ANOMALOUS-out-of-range / mismatched-task-sets / relative-threshold /
non-numeric-score), `gate_run` (valid / dropped-samples / judge-error /
missing-score / non-numeric-score / timeout-non-terminal /
running-non-terminal / no-tasks), `gate_ptq` (pass / not-smaller /
zero-coverage→MODEL_UNSUPPORTED / unexpected-unquantized / metadata-diff
/ unknown-recipe). No GPU or cluster needed.
- **Routing assertions** (`tests/evals.json`, 5 cases): documents
expected routing — fires on "release model X at day-0", does **not**
fire on "just quantize X" / "run MMLU on this endpoint", and the
gate-blocking / regression-reports-and-stops behaviors. These are
behavior specs for manual QA; there is no automated skill-routing
harness yet (tracked as "Remaining Work" in the design doc).
- **End-to-end integration (run on aws-pdx / B300, 2026-06-04)** — ✅ the
full chain and **all five gate types** were exercised on **real cluster
artifacts** (Qwen3-0.6B). Results below.
#### End-to-end integration results
Every gate fired correctly on real PTQ / serve / eval outputs, and
**both** failure-routing branches were exercised:
| Stage | Gate | Real artifact | Verdict | ✓ |
|---|---|---|---|---|
| Setup | reachability | aws-pdx SSH + SLURM | PASS | ✅ |
| PTQ (happy) | `gate_ptq` | FP8 ckpt, 0.50×, FP8=196 layers,
unexpected=0 | `pass:true` | ✅ |
| **PTQ (failure fixture)** | `gate_ptq` | `nvfp4_experts_only` on a
**dense** model → NVFP4=0, 196 silently-unquantized |
**`MODEL_UNSUPPORTED`** → chain **STOPS** before deploy | ✅ |
| Deploy | health | 2 vLLM endpoints, `/health`=200 + generation | PASS
| ✅ |
| Eval ×2 | `gate_run` | real run-summaries, 300/300 scored, SUCCESS |
`pass:true` | ✅ |
| Compare | `gate_compare` | arc_easy BF16=57.0 vs FP8=54.0 |
**`REGRESSION`** (drop 3.0 > 1pt) | ✅ |
- **Failure fixture (the key control-flow proof):** an experts-only
recipe on a dense model matched 0 modules, so PTQ "succeeded" in 3.5s
and exported a checkpoint with `quant_algo:null` / `quantized_layers:{}`
— every linear silently BF16. `gate_ptq` caught the zero-coverage and
routed to `MODEL_UNSUPPORTED`, so the chain **did not** deploy/eval the
bad checkpoint. This is exactly the silent coverage-miss the gate exists
to stop.
- **Caveat — ACCEPT not reached on real data:** the happy path returned
`REGRESSION`, which is the **correct** verdict — Qwen3-0.6B FP8
genuinely loses ~3pt on arc_easy (real, ~2 SE at 300 samples;
weight-only FP8 = 54.0 was no better than KV-FP8 = 54.67, so the
KV-cache wasn't the cause). Tiny models are quantization-fragile and
don't meet the <1% criterion; the gate correctly refused to rubber-stamp
it. The threshold was **not** loosened to manufacture a pass. The
`ACCEPT` terminal + publish-recommendation path therefore remains
validated by unit test (`test_compare_accept_within_threshold`) only,
not end-to-end — demonstrating it on real data needs a larger model
(e.g. Qwen3-4B/8B) where FP8 is near-lossless, deferred as optional
follow-up.
- **Infra bugs shaken out (incidental, not in this skill's code):** (1)
the ModelOpt launcher's `/hf-local` bind mount has no host dir on
aws-pdx → set `SLURM_HF_LOCAL=<lustre dir>`; (2) `lm-eval` is unusable
inside `vllm/vllm-openai:cu130-nightly` (ships transformers 5.x, which
removed `AutoModelForVision2Seq` that lm-eval imports at load) → used a
dependency-free `/v1/completions` log-likelihood scorer.
### Scope
**In v1:** the linear chain + gate scripts + `ACCEPT`/fail-with-report
outcomes. On `REGRESSION`, v1 *reports* "recipe R regressed on tasks
[...]" and stops.
**Deferred (follow-up PR):** the evaluator-optimizer recipe loop
(compare → pick next recipe → re-PTQ), which needs the `bigpareto`
integration and the shared `modelopt-agent-protocols` schema adopted on
both sides.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (new skill; no change to
existing skills)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new
deps)
- Did you write any new necessary tests?: ✅ (21 gate-script unit tests +
5 routing evals; end-to-end integration run on aws-pdx — see Testing)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ⬜ (run `/claude review`
before marking ready)
### Additional Information
Design rationale: ModelOpt Agent Skills Design doc — this skill
implements the "deterministic day-0 chain driver"
(prompt-chaining-with-gates, the code-driven orchestration pattern from
Anthropic's *Building Effective Agents*). The gate scripts double as the
data source for the Observability stage metrics, and `gate_compare.py`'s
verdict is the entry point for the deferred evaluator-optimizer recipe
loop.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added a deterministic day0-release skill: linear gated flow (setup →
PTQ → deploy → eval → compare) that halts on failed gates.
* **New Tools**
* Added CLI gates for PTQ checkpoint validation, run/evaluation
validation, and baseline-vs-candidate comparison (scale-aware
thresholds, anomaly detection, clear verdicts and exit codes).
* **Tests**
* Added unit tests covering pass/regress/anomaly outcomes, failure-class
triage, and gate scenarios.
* **Documentation**
* Added spec and changelog entry describing inputs, gating rules, triage
table, and closeout reporting.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
c78e654744 |
Skip Softmax diffusion export (#1269)
### What does this PR do?
Type of change: New Feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
Adds HuggingFace `config.json` export of skip-softmax sparse-attention
calibration for diffusion pipelines (e.g. Wan 2.2), on top of the base
skip-softmax work.
- **`_export_diffusers_checkpoint`** walks every `nn.Module` component
of a diffusers pipeline, calls `export_sparse_attention_config`, and
writes the result into that component's `config.json` under the
`sparse_attention_config` key. The sparse config lives **only** in
`config.json` — there is no standalone `sparse.yaml`.
- **`export_sparse_attention_config`** emits a `config_groups` schema
where each algorithm's parameters are nested inside its own group; only
`config_groups` and `producer` are top-level:
- skip-softmax group → `algorithm: "skip_softmax"`, `targets`, `ignore`
(layers kept dense — e.g. cross-attention + first/last blocks),
`initial_disabled_steps` (opt-in, user-set; emitted only when `> 0`),
`threshold_scale_factor` (`a * exp(b * target_sparsity)`), and
`target_sparsity`.
- N:M group → `algorithm: "sparse_softmax"` with
`sparsity_n`/`sparsity_m`, `dense_sink_tokens`, `dense_recent_tokens`
flattened into the group.
- **Deploy reader**
(`modelopt/torch/sparsity/attention_sparsity/plugins/sparse_attn_config.py`)
reads these per-group params back, keeping the export↔load round-trip
consistent.
- **Example wiring**:
`examples/diffusers/sparsity/wan22_skip_softmax.py` gains
`--export-dir`, `--skip-softmax-threshold`, and
`--initial-disabled-steps`. `--export-dir` runs
`export_hf_checkpoint(pipe, export_dir=...)` after calibration.
- Updated `CHANGELOG.rst`.
### Usage
```bash
python examples/diffusers/sparsity/wan22_skip_softmax.py \
--model-path Wan-AI/Wan2.2-T2V-A14B-Diffusers \
--calibrate --target-sparsity 0.5 --calib-size 4 \
--initial-disabled-steps 5 \
--export-dir ./wan22_skip_softmax_ckpt
```
Resulting layout — a `config.json` per component, **no `sparse.yaml`**:
```
wan22_skip_softmax_ckpt/
├── transformer/config.json # carries sparse_attention_config
├── transformer_2/config.json # carries sparse_attention_config
├── vae/ … text_encoder/ … tokenizer/ … scheduler/ …
└── model_index.json
```
A representative `config.json` entry for a diffusion transformer:
```json
"sparse_attention_config": {
"config_groups": {
"group_0": {
"algorithm": "skip_softmax",
"targets": ["WanAttention"],
"ignore": ["blocks.0.attn1", "blocks.0.attn2", "…"],
"initial_disabled_steps": 5,
"threshold_scale_factor": {
"formula": "a * exp(b * target_sparsity)",
"prefill": {"a": 1443.49, "b": 4.30}
},
"target_sparsity": {"prefill": 0.5}
}
},
"producer": {"name": "modelopt", "version": "0.45.0..."}
}
```
The N:M variant adds a second group:
```json
"group_1": {
"algorithm": "sparse_softmax",
"targets": ["WanAttention"],
"sparsity_n": 2, "sparsity_m": 4,
"dense_sink_tokens": 0, "dense_recent_tokens": 64
}
```
### Testing
- `tests/examples/diffusers_sparsity/test_sparsity.py`: baseline /
triton-baseline / fixed-threshold runs of the Wan 2.2 example, plus a
Python-API calibrate → **export** test asserting the nested
`sparse_attention_config` (`threshold_scale_factor`, `target_sparsity`,
`ignore`, `initial_disabled_steps`) and the absence of any
`sparse.yaml`.
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attention_conversion.py`
and `test_sparse_attn_config.py`: unit coverage of the per-group export
schema and the deploy-reader round-trip (writer nests → reader reads
from groups → internal mtsa config unchanged).
- Validated end-to-end on Wan 2.2 T2V-A14B: full 4-prompt / 40-step /
81-frame calibration; the exported checkpoint carries the nested schema
in both `transformer` and `transformer_2` `config.json`, and runtime
measurement shows ~47–49% tile sparsity at a 0.5 target.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ❌ The exported
`sparse_attention_config` schema was renamed and nested per-group during
0.45.x development, and the loader reads only the new layout —
checkpoints exported by earlier 0.45.x builds must be re-exported. No
released version is affected. <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
1555e6de20 |
Fix CodeCov upload issues (#1648)
Disable codecov binary validation which seems to be constantly failing
```
gpg: Signature made Tue Apr 21 19:28:03 2026 UTC
gpg: using RSA key 27034E7FDB850E0BBC2C62FF806BB28AED779869
gpg: Can't check signature: No public key
==> Could not verify signature. Please contact Codecov if problem continues
Exiting...
```
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Chores**
* Updated CI workflow notes and removed an outdated header comment.
* Added explanatory comments to the Linux job and adjusted the code
coverage upload step to use a relaxed validation mode (no other upload
settings changed).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
54ce4e09d8 |
Add Quantization Aware Distillation (QAD) to Megatron-Bridge example (#1600)
### What does this PR do? Type of change: new example **Note:** This is **part 2 of 4** (builds on #1589): - **Part 1 (#1589):** Megatron-Bridge `quantize.py` + `export.py` support and tests. - **Part 2 (this PR):** extend `distill.py` for quantization-aware distillation (QAD) — load a quantized Megatron checkpoint as the student. - **Part 3:** https://github.com/NVIDIA/Model-Optimizer/pull/1601 - **Part 4:** repeat the NVFP4 + QAD experiments on a non-Nemotron model. Extends `examples/megatron_bridge/distill.py` to initialize the student from a **Megatron checkpoint** (a quantized checkpoint from `quantize.py`, or a pruned one) via `--student_megatron_path`, enabling **Quantization Aware Distillation (QAD)**: - `--student_hf_path` still builds the student architecture; `--student_megatron_path` supplies the (optionally quantized) weights. - For a quantized checkpoint, the ModelOpt quantize mode + base weights are restored onto the **plain student before the knowledge-distillation conversion** (`restore_sharded_modelopt_state` is a no-op once a model is already converted), so the distilled checkpoint stays exportable as a quantized model with `export.py`. **Upstream dependency / workaround:** `DistillationProvider.provide()` has no seam to transform the student before the KD conversion, so this patches `provide()` at the class level (via an `id()`-keyed registry, because the provider proxies instance-attribute assignment to its teacher once the teacher is set). A companion Megatron-Bridge PR adds a first-class `DistillationProvider.student_pre_conversion_hook`; from nemo:26.06 onwards the workaround should be removed and replaced with that hook (a removal note in `distill.py` documents exactly how). ### Usage ```bash # 1) PTQ -> quantized Megatron checkpoint (part 1) torchrun --nproc_per_node 2 quantize.py \ --hf_model_name_or_path Qwen/Qwen3-8B --quant_cfg fp8 --tp_size 2 \ --export_megatron_path /tmp/Qwen3-8B-FP8-megatron # 2) QAD: distill the quantized student from the unquantized teacher torchrun --nproc_per_node 8 distill.py \ --teacher_hf_path Qwen/Qwen3-8B \ --student_hf_path Qwen/Qwen3-8B \ --student_megatron_path /tmp/Qwen3-8B-FP8-megatron \ --data_paths 1.0 tokenized/data_text_document \ --train_iters 1000 --output_dir /output/qwen3_8b_qad # 3) export the distilled quantized checkpoint (part 1) torchrun --nproc_per_node 1 export.py \ --hf_model_name_or_path Qwen/Qwen3-8B \ --megatron_path /output/qwen3_8b_qad/checkpoints \ --export_unified_hf_path /tmp/qwen3_8b_qad_fp8_hf ``` ### Testing `tests/examples/megatron_bridge/test_qad.py` (validated on a 2-GPU NeMo `26.04` container): quantize a tiny Qwen3 at TP=2 → QAD distill from the quantized student → `export.py` to a unified HF checkpoint, asserting `hf_quant_config.json` is written (proves the quantize mode survived QAD). Includes a commented-out vLLM deployment check, validated locally (full flow passes; vLLM loads the export as `quantization=modelopt`). Existing normal/Puzzletron distillation tests still pass. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: N/A (new example feature; default behavior unchanged when `--student_megatron_path` is not set) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new dependencies) - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ✅ ### Additional Information Depends on a companion Megatron-Bridge PR adding `DistillationProvider.student_pre_conversion_hook` (the upstream replacement for the class-level `provide()` workaround). The Nemotron-3 tutorial NVFP4 + QAD experiments ship in part 3. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Quantization Aware Distillation (QAD) workflow to recover accuracy of quantized Megatron students and distill from quantized checkpoints. * CLI option to initialize a distillation student from a Megatron checkpoint and a structure-only load path for bridging. * **Documentation** * Expanded runnable quantize → QAD → export guidance and best-practice tips. * **Tests** * End-to-end test validating quantize → QAD → export artifacts. * **Chores / UX** * Clearer rank-aware messages, improved tokenizer padding handling, and more consistent export behavior (fixed export dtype). <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0081473861 |
Speed up slow unit/gpu/example tests (#1616)
### What does this PR do?
Type of change: test infrastructure / test speedups + CI stabilization
Make the test suite faster, `tests/unit` hermetic, and the CI lanes
stable, without losing coverage. Most changes are mechanical test/infra
edits; the buckets below cover the diff broadly.
**Unit tests — hermetic (no HF Hub):** toy local datasets/configs + the
local tiny tokenizer (with a checked-in chat template) replace Hub
assets; `tests/unit/conftest.py` enforces offline mode. Genuinely-HF
tests moved to `tests/gpu*` (e.g. the new
`tests/gpu/torch/utils/test_dataset_utils.py`). `CONTRIBUTING.md`
documents the hermetic-unit-test expectation.
**Unit-test speedups (no coverage loss):** speculative (disable CPU
torch.compile), calibrator (fewer histogram bins), ONNX conv/dynamo
(smaller shapes + representative subset), Ruler/sparse-attention (local
tokenizer), data-parallel autoquant (world size 4→2). Shared
`tiny_tokenizer` fixture. The distributed test helper now uses a private
`spawn` context instead of mutating the global start method (avoids
cross-test contamination).
**Rarely-used autonas/fastnas tests:** heavy parametrize cases marked
`@pytest.mark.manual`, one representative kept per test (fastnas
preferred); lighter sibling tests still cover core behavior. The legacy
FSDP1 NAS distributed test is also dropped: FastNAS/AutoNAS aren't used
with either FSDP1 or FSDP2, and FSDP1 is superseded by the newer FSDP2
API — so we keep a single FSDP2 case as a sanity check and drop FSDP1,
leaving the suite leaner.
**gpu_megatron:** deduplicate distributed worker pools by world_size
within a module (saves a redundant pool spin-up in multi-pool files;
module-scoped, no cross-module reuse).
**Example tests:** reduce per-test work via args that default to current
behavior (tests pass the fast values) — torch_onnx TRT optimization
level, diffusers calibration/inference steps, eagle `sample_size`,
megatron_bridge iters/calib, llm_sparsity data slice, export
safetensors-structure `calib_size`. Also enable the recently added
`gpt-oss` example tests in CI.
**Per-test timeouts:** `pytest-timeout` with a default per-directory
timeout (60s unit / 300s gpu+example) enforced in `tests/conftest.py`
(`timeout_func_only` in `pyproject.toml`), so a new test cannot silently
exceed the budget — an unmapped test dir crashes collection. A few
inherently slow tests carry explicit higher per-test overrides
(CUDA-compile, autotune, dflash).
**CUDA kernel pre-compilation:** a dedicated `tests/gpu/_extensions`
test JIT-builds the conv3d implicit-GEMM kernel up front (collected
before the functional tests in the same process) so the one-time build
cost no longer lands on — and time out — the first functional test that
uses it. Mirrored into the `llm_ptq`/`vlm_ptq` example lanes.
**Test relocation & optional-dependency guards:** vLLM sparsity plugin
test moved to `tests/gpu_vllm` (drops the in-test `importorskip`);
diffusers-dependent unit test guarded with `importorskip("diffusers")`
for partial-install lanes; `gpt_oss` example test dir renamed to
`gpt-oss` to match the CI matrix.
**Diffusers test models:** shared model-path constants in
`tests/_test_utils/examples/models.py` consolidated/renamed and point at
tiny `hf-internal-testing` test pipes (SDXL/SD3/FLUX) so
cachify/quantize/export tests run on toy weights; `local_id`s
normalized.
**Shared dataset utils:** `examples/llm_sparsity/.../hf_pts.py` now uses
`get_dataset_dataloader` (drops the bespoke cnn_dailymail-only
`get_calib_dataloader`; supports any registered/HF/JSONL dataset,
includes attention_mask); `data_prep.py` gains `--max_samples`.
**CI workflows:** container image bumps (pytorch 26.04→26.05, TRT-LLM
rc16→rc17) and tightened lane timeouts (unit 30→15 min, gpu lanes
trimmed, onnx example lane 45 min).
**Imports at top of file:** in-function imports across the test suite
are moved to module top per the coding guideline, conservatively —
optional deps stay guarded (in-function or behind a module-level
`importorskip`) in `tests/unit` since the partial-install lane runs
without them, and build/hardware-availability imports (apex, triton,
megatron/transformer_engine, tensorrt_llm) plus `_test_utils` lazy
guards are left in place.
**Kernel warning filters:** the repeated `filterwarnings` blanket-ignore
in six `tests/gpu/torch/kernels/**` modules is consolidated into a
scoped hook in `tests/gpu/torch/kernels/conftest.py` (kernel tests only
— the rest of the suite keeps surfacing warnings).
**Eagle example speedups:** `torch.compile` (eagle recipe default) added
~2 min to every eagle training test; it's now disabled in the eagle
example tests except one smoke (`test_llama_eagle3[1-False]`), and the
downstream resume / AR-validate / export tests point at the compile-free
checkpoint. Measured: `test_ar_validate` 139s→17s, offline training
142s→22s, streaming 140s→23s — the compile path is still smoke-tested
once.
**Example lanes install editable (`-e`):** so example scripts launched
as subprocesses resolve `modelopt` to the same source path as the test
process and reuse the pre-compiled CUDA-extension cache instead of
recompiling (~2 min/test); verified in the TRT-LLM container.
**Tiny test tokenizer:** `get_tiny_tokenizer` defaults to left padding
(what decoder-LM calibration expects) and ships a terse
generation-tagged chat template — replacing a verbose ChatML one that
inflated tokenized length on the 128-vocab tokenizer and broke the
offline-PTQ example tests' `max-seq-len` filter.
**Restored Hub-download coverage:** the live (ungated) HF dataset
round-trips exercising `get_dataset_samples`' download branch now live
in `tests/gpu/torch/utils/test_dataset_utils.py` (they had been dropped
from the hermetic unit file without a counterpart).
Individual file changes not explicitly called out above fall under this
general test/CI cleanup.
### Testing
Unit + the touched gpu_megatron files validated locally; example/GPU
lanes validated in CI.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (tests + example CLI args
default to prior behavior)
- If you copied code from any other sources or added a new PIP
dependency: N/A
- Did you write any new necessary tests?: N/A (optimizes/relocates
existing tests)
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ❌ (pending)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **Chores**
* Updated container image versions for PyTorch (26.04→26.05),
TensorRT-LLM (1.3.0rc16→1.3.0rc17), and ONNX/TensorRT (26.04→26.05).
* **Tests**
* Enhanced test isolation: unit tests now run hermetically without
HuggingFace Hub access.
* Optimized test runtime via smaller model/dataset parameters and
parallel test caching.
* Added CUDA extension availability tests and extended dataset utility
coverage.
* **Documentation**
* Updated testing guidelines in `CONTRIBUTING.md` to emphasize offline
test design.
* **Chores**
* Added pytest timeout configuration and improved CI/CD workflow
efficiency with editable installs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
72df833e0b |
Add active-MoE AutoQuant cost accounting (#1497)
### What does this PR do?
• Type of change: new feature
Adds an active_moe cost model for auto_quantize effective-bits search.
This lets AutoQuant account for routed MoE expert weights by active
decode weight traffic instead of total checkpoint weight
size, using active_moe_expert_ratio = num_experts_per_tok / num_experts.
The default behavior is unchanged: cost_model="weight" still counts all
quantizable weights equally.
### Usage
import modelopt.torch.quantization as mtq
model, search_state = mtq.auto_quantize(
model,
constraints={"effective_bits": 5.0},
quantization_formats=[
mtq.NVFP4_DEFAULT_CFG,
mtq.FP8_DEFAULT_CFG,
],
data_loader=calib_dataloader,
forward_step=forward_step,
loss_func=loss_func,
cost_model="active_moe",
# Optional. If omitted, ModelOpt tries to infer this from model.config.
active_moe_expert_ratio=2 / 64,
)
The HF PTQ example also exposes:
--auto_quantize_cost_model active_moe \
--auto_quantize_active_moe_expert_ratio 0.03125
### Testing
python -m pytest tests/unit/torch/quantization/test_autoquant.py -q -k
'active_moe or quant_recipe_hparam_cost_weight'
python -m pytest tests/unit/torch/quantization/test_autoquant.py -q -k
'not data_parallel_auto_quantize'
Results:
- 4 passed
- 58 passed, 1 deselected
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added active-MoE cost model option for auto-quantization with
configurable expert ratio; API and CLI accept cost_model and
active_moe_expert_ratio
* Unified auto-quantize supports new quant format w4a16_nvfp4
* **Bug Fixes**
* Ensure labels are moved to the logits device for base models without
an lm_head
* CLI enforces valid expert-ratio range and requires active-MoE mode
when a ratio is provided
* **Tests**
* Added unit tests for active-MoE behavior, cost-weighting, ratio
handling, and search budget selection
<!-- review_stack_entry_start -->
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1497?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)
<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
|
||
|
|
b3d83ba6f0 |
Increase Claude review timeout
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
7aa0c95646 |
Add tests/gpu_vllm (#1517)
### What does this PR do? Type of change: new tests This PR adds unit tests for vLLM fakequant, specifically testing code in `modelopt/torch/quantization/plugins/vllm.py` ### Testing ``` pytest tests/gpu_vllm/torch/quantization/test_vllm_dynamic_modules.py -sv ``` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added comprehensive GPU vLLM test suite with end-to-end quantization checks and fixtures for TinyLlama, TinyQwen3-MoE, and DeepSeek V3; includes helpers to build tiny DeepSeek V3 models. * **Chores** * Updated GPU CI to use explicit container image references, added a GPU-focused test session, and adjusted test-run setup for vLLM. * **Documentation** * Documented new GPU test directory in contributing guide. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1517?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
eb5ed2df68 |
[CI] Bump torch, transformers and dev containers to latest (#1554)
- Transformers upper bound bumped from `<5.8` to `<5.10` - Enable torch 2.12 CICD testing - Bump TRT-LLM container to `1.3.0rc16` (transformers 5.5) - Use pytorch and tensorrt 26.04 containers in CICD <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated CI test container images and targeted Torch version across workflows; adjusted release CI job to use the newer torch config. * Broadened Transformers constraint in project metadata and test/dev pins. * Removed strict transformers pins from example requirements and lifted a compression dependency cap. * Raised the import-time Transformers version threshold for compatibility warnings. * **Tests** * Refactored a GPU test to collect and report validation errors and updated numeric expected baselines. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1554?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
a2c496af0d |
[OMNIML-4788] specdec_bench: configuration.json provenance + upload_to_s3 (#1531)
> [!WARNING] > **Breaking on-disk schema change (specdec_bench v1.0.0).** This PR renames the acceptance-rate metric fields across `AcceptanceRate` / `MTBench` / `SpecBench` writers: > > | Old (pre-1.0.0) | New (1.0.0) | > |---|---| > | `Request_AR` | `Request_AL` | > | `Category_AR` | `Category_AL` | > | `Average_AR` | `Average_AL` | > | — | `Joint_Acceptance_Rate` (new) | > > The renamed values were always **acceptance length** (mean tokens generated per inference step), not a rate, and the visualizer reads `*_AL`. Pre-1.0.0 runs in S3 have `*_AR` and no `Joint_AR`; they must be re-run or post-processed before comparing. The visualizer aggregates runs by `specdec_bench` major version so accidental cross-methodology comparison is blocked. ### What does this PR do? Type of change: new feature Adds reproducibility provenance to `specdec_bench/configuration.json` and ports `upload_to_s3.py` from `iputterman/specdec_bench@main` (personal-namespace fork) into upstream. This is the first PR in a multi-stage migration off Izzy's fork now that he's left the team. Tracked in [OMNIML-4788](https://jirasw.nvidia.com/browse/OMNIML-4788). **Provenance fields added to configuration.json** (alongside existing argv / engine_version / gpu / python_version): - `specdec_bench_version` — methodology semver declared in `specdec_bench/__init__.py`. Bump minor on additive metrics, major on changed metric *definitions*. The visualizer (Phase 4 of the migration) will aggregate runs by major version so plots don't accidentally compare across methodology changes. - `specdec_bench_sha`, `modelopt_sha`, `modelopt_version`, `nmm_sandbox_sha`, `container_image` — code/runtime provenance. Each prefers an env var set by the harness (`SPECDEC_BENCH_SHA`, `MODELOPT_SHA`, `MODELOPT_VERSION`, `NMM_SANDBOX_SHA`, `CONTAINER_IMAGE`) and falls back to `git rev-parse` / `modelopt.__version__` when running standalone. The env-var preference is necessary because the runtime container has no `.git/` (the launcher packager tarballs source without git metadata) and may not have `modelopt` installed. - `checkpoint.{path, size_bytes, index_sha256, index_source}` — cheap reproducibility fingerprint that hashes `model.safetensors.index.json` (or `config.json` fallback). Changes whenever any tensor changes. - `serving_config` — engine-level config dict captured after init via a new `Model.get_serving_config()` method. VLLM dumps `AsyncEngineArgs` + the live `vllm_config.to_dict()`; SGLANG dumps the `engine_kwargs` passed to `sgl.Engine`; TRTLLM left at the base default `{}` for a later iteration. - `timestamp` — UTC ISO 8601. **Other changes** - `upload_to_s3.py` + `specdec_bench/s3_utils.py` ported from iputterman/specdec_bench@main. Recognizes run dirs by sentinel files, refuses to overwrite existing S3 prefixes. - `_redact_config` allowlists `tokenizer`, `tokenizer_path`, `tokenizer_mode`, `tokenizer_revision` so the model path stops being redacted (latent bug from substring-matching `token` ⊂ `tokenizer`). - `requirements_speed.txt`: `boto3`, `botocore` added (used by `s3_utils`). **Out of scope** (deferred to Phase 1b / Phase 2): - `--sweep_config` driver that emits per-run-dir nesting `<sweep>/<NNN_dataset_c<conc>>/` - `--s3_upload` flag baked into `run.py` itself - Launcher auto-injection of the provenance env vars (currently the example YAML sets them statically) - `container_digest` (enroot integration) and full GPU/driver inventory - TRTLLM `get_serving_config()` ### Usage ```bash # Run a smoke benchmark (Qwen3.5-4B + vLLM + MTP draft=3) — example YAML included uv run launch.py --yaml examples/Qwen/Qwen3.5-4B/specdec_bench_mtp.yaml --yes # After it lands, upload the run directory to S3: S3_KEY_ID=team-specdec-workgroup \ S3_SECRET=... \ python upload_to_s3.py /path/to/sweep_dir s3://team-specdec-workgroup/results ``` ### Testing Cluster-tested end-to-end on cw-dfw (Slurm job 11978794, NeMo Run experiment `cicd_1779403623`, ~19 min wall): - Qwen3.5-4B + vLLM + MTP draft=3 + SPEED-Bench-Internal/qualitative (80 requests) - `configuration.json` (22 KB) populated all eight new provenance fields - `Request_AR` mean 3.327 (vs 3.330 on the pre-Phase-1a run — within noise; methodology unchanged) - `upload_to_s3.py` (real upload, not dry-run) landed [s3://team-specdec-workgroup/results/qwen35_4_mtp_smoke_2026-05-21/specdec_bench_mtp/](https://app.s8k.io/buckets/team-specdec-workgroup/?prefix=results%2Fqwen35_4_mtp_smoke_2026-05-21%2F) where the visualizer at http://10.131.132.205:8080 can pick it up. ### Before your PR is "Ready for review" - Is this change backward compatible?: ✅ - `configuration.json` only gains fields. `upload_to_s3.py` / `s3_utils.py` are new files. `Model.get_serving_config()` default = `{}` so existing subclasses without an override behave as before. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - `boto3` / `botocore` are Apache 2.0 (permissive); `upload_to_s3.py` + `s3_utils.py` are ported from a private NVIDIA repo with explicit copyright headers retained. - Did you write any new necessary tests?: ❌ - Validated by cluster smoke (see Testing). Will add unit-tests for `dump_env` provenance fields and `upload_to_s3._discover_runs` in a follow-up. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Internal-facing tooling. - Did you get Claude approval on this PR?: ❌ (triggering after open) ### Additional Information Tracked in JIRA [OMNIML-4788](https://jirasw.nvidia.com/browse/OMNIML-4788). The full multi-phase plan is on that ticket's SPEC block — this PR is Phase 1a. Cherry-picked alongside the harness change are two example YAMLs (`examples/Qwen/Qwen3.5-4B/specdec_bench.yaml` for the NONE autoregressive baseline, `..._mtp.yaml` for the MTP run) that gave us cluster-test evidence. Can be split out if preferred. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an S3 upload CLI for benchmark results with dry-run and skip-existing options * Automatic capture of run configuration, provenance and redacted environment into saved config * Models now export serving configuration for reproducible runs * New launcher entrypoint and example job configs for Qwen SPEED-Bench runs * **Documentation** * README section describing S3 upload usage and supported local layouts * **Bug Fixes / Changes** * Acceptance-rate metric keys renamed in output (AR -> AL) * **Tests / CI** * New tests for redaction and S3 utilities; CI now runs specdec_bench examples <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1531?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: chenhany <chenhany@nvidia.com> |
||
|
|
a5b7d3e824 |
ci: authenticate nvcr.io container pulls with NGC API key (#1559)
### What does this PR do? Type of change: Bug fix (CI infrastructure) NGC now requires authenticated pulls for `nvcr.io/nvidia/nemo:*` (and likely other) containers after an EULA acceptance gate. Unauthenticated pulls fail with: ``` Error response from daemon: Head "https://nvcr.io/v2/nvidia/nemo/manifests/26.04": denied: Please accept license on the browser to be able to download ``` Add a `credentials:` block to the `container:` spec in `gpu_tests.yml` and `_example_tests_runner.yml` so GitHub Actions logs in to nvcr.io with an NGC API key before pulling. ### Usage ```yaml container: image: nvcr.io/nvidia/${{ matrix.container_image }} credentials: username: $oauthtoken password: ${{ secrets.NGC_API_KEY }} ``` ### Required follow-up (cannot be done in code) 1. Add a repo secret `NGC_API_KEY` (Settings → Secrets and variables → Actions). 2. An NGC org admin must log into ngc.nvidia.com once and accept the EULA on the relevant containers (NeMo, PyTorch, TensorRT-LLM). The API key alone does not bypass the EULA gate. ### Testing - Workflow YAML validates locally (no syntax errors). - Full verification requires the secret to be set and the EULA accepted — will be confirmed on the next workflow run. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information The lint warning `Context access might be invalid: NGC_API_KEY` shown in IDE will resolve once the secret is added to the repository. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated GitHub Actions workflows to improve container registry authentication for test infrastructure. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1559?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
8f1529abd3 |
Consolidate coding standards in CONTRIBUTING.md; tune review automation (#1519)
### What does this PR do? Type of change: documentation + tooling Consolidates the previously agent-only coding guidelines into `CONTRIBUTING.md` so both human contributors and AI agents work from the same source of truth, and tunes the `/claude review` and CodeRabbit automation so they are advisory (approve-only, never blocking). **Docs reorg** - Move the content of `.agents/developer-guidelines.md` into `CONTRIBUTING.md` as a new `📐 Coding standards` section plus a `Test design principles` subsection under the existing test section. Delete the now-redundant file. - Add two new coding principles: - **Imports at top of file** in both source and test files. In-function imports are reserved for resolving circular imports or guarding optional dependencies (e.g., TRT-LLM, Megatron-Core), with a brief comment naming the reason. - **`__all__` + `from .module import *`** in package `__init__.py` to make the public API surface explicit at the definition site and keep star-imports safe. - Update `AGENTS.md` to point at the merged location and move the agent-specific "use relative paths" rule into it. Drop the dangling reference in `claude_review.yml`. **AI Review automation** - `/claude review` (`.github/workflows/claude_review.yml`): - Step 1 of the mandatory workflow now reads prior Claude comments/reviews so the bot doesn't duplicate already-raised findings. - Add brief thoroughness directives (per-file category coverage, cross-file dataflow trace for new/modified public symbols) to push more issues into the first review pass. - Replace the ambiguous "approve if no significant issues" line with an explicit decision rule: `0 CRITICAL AND 0 IMPORTANT → approve`, regardless of SUGGESTION count. SUGGESTIONs alone never block approval. - **Never submits a formal "Request changes" review.** When issues are found, posts a `--comment` review instead. Claude review is advisory and must not block merges. - CodeRabbit (`.coderabbit.yaml`): - Enable `reviews.request_changes_workflow: true` so CodeRabbit can formally approve PRs once its comments are resolved and pre-merge checks pass. Despite the name, this setting never submits a blocking "Request changes" review. - Add a `tests/**/*.py` path_instruction encoding the imports-at-top rule, lean-tests guidance, and correct test-directory placement. ### Usage N/A — documentation and automation configuration. ### Testing - `pre-commit run` ran clean on both commits (markdownlint, YAML format, etc.). - `/claude review` workflow changes will be validated on the next PR that triggers it. - CodeRabbit `tests/**/*.py` rule and approval behavior will be exercised on the next PR touching tests. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (docs + workflow config only) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ (will run `/claude review` after PR opens) ### Additional Information The previous `.agents/developer-guidelines.md` content was not agent-specific — it described universal coding standards. Hiding it under `.agents/` discouraged human contributors from reading it. The merge into `CONTRIBUTING.md` makes it discoverable on the standard GitHub PR-creation flow. The shift to approve-only automated reviews keeps the signal (approvals when clean, inline comments when issues are found) without giving the bots formal merge-blocking authority — the existing pre-merge `custom_checks` security gates remain the hard blockers. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Enhanced contribution guidelines with comprehensive coding standards and new test design principles. * Updated agents guidance to remove the legacy developer-guidelines reference. * **Chores** * Enabled formal "request changes" workflow and added test-path review rules in CI config. * Refined GitHub Actions review workflow (timeout, review prompt, single-pass requirements, and approval criteria). * Removed legacy developer-guidelines file. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1519?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
e27f76fbfe |
Refine agent contribution guidance (#1488)
### What does this PR do? Type of change: documentation. This PR centralizes repository agent guidance around `AGENTS.md` as the shared entrypoint. `CLAUDE.md` now points to `AGENTS.md`, while detailed coding principles and tool-specific setup notes live under `.agents/`. Key changes: - Add `AGENTS.md` as the shared repository agent instructions file. - Point `CLAUDE.md` at `AGENTS.md` so Claude Code reads the same entrypoint. - Add `.agents/developer-guidelines.md` for production code and review principles, including minimal changes, extension points, testing, performance, and compatibility expectations. - Add `.agents/TOOLING.md` for human-maintained notes about local agent overrides and shared instruction maintenance. - Update the Claude review workflow to read `AGENTS.md` and `.agents/developer-guidelines.md`. - Update contributor and README guidance with focused local validation examples and an AI-agent pointer. - Ignore local agent override files. ### Usage N/A. Documentation-only change. ### Testing - `git diff --check` - Commit pre-commit hooks passed, including `markdownlint-cli2`. - GitHub PR checks passed, including code quality, docs, unit Linux, required gate checks, DCO, CodeRabbit, and Codecov. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information Docs-only agent guidance update. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added comprehensive AI-agent instructions, developer guidelines, and tooling notes for AI-assisted workflows. * Updated contribution docs with clearer test-running guidance and linked agent resources from the main README. * Adjusted the code-review workflow to reference the new guidance materials. * **Chores** * Updated ignore rules to exclude local agent override files. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
be85c7fc52 |
Setup pre-commit hooks for claude PR creation
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
794a4e3705 |
Let @claude to create PRs from GH Issues
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
fe9644eb5c |
[CI] Fix Claude GitHub Actions setup (#1401)
## Summary Fix Claude GitHub Actions setup added in #1400 ## Test plan - [ ] Comment `/claude review` on this PR and verify the deep-review job runs (this is also serving as the smoke-test for the workflows merged in #1400) - [ ] Comment `@claude any thoughts on this change?` and verify the interactive job runs - [ ] Check that both jobs respect the new 10-minute cap if they ever stall 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated CI/CD workflow timeouts for job execution management. * Disabled experimental AI beta features in CI runs to ensure consistent behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
1c2085f7df |
[CI] Add Claude Code GitHub Action workflows (#1400)
## Summary Adds two GitHub Actions workflows for code review and interactive Q&A on PRs/issues, routed through NVIDIA's internal Anthropic inference proxy: - **`claude_review.yml`** — on-demand deep review triggered by `/claude review` comment. Focused on **algorithm correctness**, **mode/state composability** (`apply_mode` / `restore` / `modelopt_state` round-trip), **export compatibility** (HF / TRT-LLM / ONNX), and **backward compatibility** for saved checkpoints and recipes. Explicitly scoped to be **orthogonal to CodeRabbit** (which already handles style, typos, and security anti-patterns via `.coderabbit.yaml`). - **`claude.yml`** — interactive `@claude` assistant in PR/issue/review comments. Read-only (`contents: read`) — answers questions and posts suggestions but cannot push commits. Both workflows are gated on `author_association in [OWNER, MEMBER, COLLABORATOR]` to limit usage to NVIDIA org members and prevent abuse from external contributors on this public repo. ## Configuration Repo secrets / variables already configured (values omitted): | Name | Type | Purpose | |---|---|---| | `ANTHROPIC_API_KEY` | secret | Inference proxy key | | `ANTHROPIC_BASE_URL` | secret | Anthropic-compatible endpoint | | `CLAUDE_MODEL` | variable | Model identifier | ## Test plan `issue_comment` / `pull_request_review` workflows can only fire from the default branch — testing requires merging first. Plan: - [ ] Merge this PR - [ ] On a follow-up PR, comment `/claude review` and verify the deep-review job runs - [ ] On the same PR, post a comment with `@claude what do you think of this change?` and verify the interactive job runs - [ ] Confirm that the same triggers from a non-NVIDIA author are skipped (no run) - [ ] Iterate on prompts / fix forward as needed Pattern adapted from `NVIDIA/Megatron-LM/.github/workflows/claude_review.yml` (self-contained, no external workflow templates). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Introduced GitHub Actions workflow for AI-assisted interactions on issues and pull requests, triggered by mentions from authorized users. * Added GitHub Actions workflow for automated code review with AI assistance, triggered by specific commands in pull requests. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
f0eaa198df |
Enable active-param and memory based Minitron pruning constraint (#1377)
### What does this PR do?
Type of change: New feature, new tests, documentation.
OMNIML-4108: Extends the Minitron NAS pruner to support pruning by
**active parameter count** (`active_params`) and **memory footprint**
(`memory_mb`) in addition to the existing total parameter count
(`params`) constraint. Also adds standalone utilities for analytical
model stats.
#### Changes
**New pruning constraint keys**
- `active_params`: prune to a target number of active (routed) params —
useful for MoE models where total ≫ active; when present,
`active_params` is the **primary sort/display metric** for candidates
(priority: `active_params` > `params` > `memory_mb`)
- `memory_mb`: prune to fit a memory budget (BF16 weights + KV-cache +
Mamba state at a given sequence length and batch size)
- Constraints can be combined (AND logic): e.g. `{"params": 6e9,
"memory_mb": 12288}`
**New standalone utilities**
(`modelopt.torch.nas.plugins.megatron_model_stats`)
- `mcore_param_count`: analytically computes total and active parameter
counts for GPT and Mamba/hybrid MCore models
- `mcore_memory_footprint_mb`: estimates memory in MB (weights +
KV-cache + Mamba state)
- `print_mcore_model_stats`: rich-formatted model stats panel
**Rich-formatted pruning logs** — search space, top-k candidate tables,
and best subnet panel printed on rank 0
**`prune_score_func` format update** — now `mmlu_<N>pct_bs<bs>` (e.g.
`mmlu_10pct_bs32`) to explicitly control batch size for MMLU evaluation;
old `mmlu_<N>pct` format removed
**Infrastructure**
- NeMo container bumped to `nvcr.io/nvidia/nemo:26.04` in CI and docs
- Added `examples/megatron_bridge/requirements.txt` with
`transformers<5.0` (required for saving some Nemotron-3-Nano models)
### Usage
```python
# Prune to 3B active params (MoE-aware) — active_params is the primary sort metric
mtp.prune(model, mode=[("mcore_minitron", ss_config)], constraints={"active_params": 3e9}, config=pruning_config)
# Prune to fit a 12 GB memory budget
mtp.prune(model, mode=[("mcore_minitron", ss_config)], constraints={"memory_mb": 12288}, config=pruning_config)
```
### Testing
Pruned Nemotron-3-Nano-30B-A3B (31.6B, A3.6B) --> A3.0B. Takes <1hr on
8x H100 (more details in #1376)
```bash
torchrun --nproc_per_node 8 examples/megatron_bridge/prune_minitron.py \
--pp_size 8 \
--hf_model_name_or_path nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 \
--trust_remote_code \
--prune_target_params 28e9 \
--prune_target_active_params 3e9 \
--hparams_to_skip num_attention_heads \
--seq_length 8192 \
--output_hf_path pruned/Nemotron-3-Nano-30B-A3B-Pruned-28B-A3B-top20-max15depth-max30width-mmlu_10pct_bs32 \
--top_k 20 \
--max_depth_pruning 0.15 \
--max_width_pruning 0.30 \
--prune_score_func mmlu_10pct_bs32 \
--num_layers_in_first_pipeline_stage 5 \
--num_layers_in_last_pipeline_stage 5
```
```
╭──────────────────────────────────────────────────── Original Model Stats ─────────────────────────────────────────────────────╮
│ Total Parameters 31.58B │
│ Active Parameters 3.58B │
│ Memory (BF16, seq_length=8192, batch_size=1) weights: 60230.1 MB, kv_cache: 48.0 MB, mamba_state: 23.8 MB, Total: 60301.9 MB │
╰───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
Top 20 Candidates with Scores
┏━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━┓
┃ # ┃ export_config ┃ active_params ┃ params ┃ score ┃
┡━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━┩
│ 1 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 64, 'num_moe_experts': 120, │ 3.00B │ 27.06B │ 0.3399 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 2 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 56, 'num_moe_experts': 112, │ 3.00B │ 25.37B │ 0.4650 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 3 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 64, 'mamba_head_dim': 56, 'num_moe_experts': 112, │ 3.00B │ 25.37B │ 0.2343 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 4 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 56, 'mamba_head_dim': 48, 'num_moe_experts': 96, │ 3.00B │ 20.09B │ 0.2552 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 5 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48, 'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 21.61B │ 0.2601 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 6 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 64, 'num_moe_experts': 96, │ 3.00B │ 19.28B │ 0.3762 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3712} │ │ │ │
│ 7 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 104, │ 3.00B │ 22.28B │ 0.4783 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 8 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 96, │ 3.00B │ 21.99B │ 0.2420 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3328} │ │ │ │
│ 9 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 112, │ 3.00B │ 25.37B │ 0.2399 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3712} │ │ │ │
│ 10 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 112, │ 3.00B │ 26.17B │ 0.2601 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3328} │ │ │ │
│ 11 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 64, 'num_moe_experts': 112, │ 3.00B │ 25.37B │ 0.2503 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 12 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 23.68B │ 0.4329 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 13 │ {'num_layers': 46, 'hidden_size': 2688, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 128, │ 3.00B │ 26.17B │ 0.2587 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 2816} │ │ │ │
│ 14 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 64, 'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 23.68B │ 0.2336 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 15 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48, 'mamba_head_dim': 56, 'num_moe_experts': 96, │ 3.00B │ 20.09B │ 0.2559 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 16 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 96, │ 3.00B │ 20.70B │ 0.4608 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 17 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 104, │ 3.00B │ 23.68B │ 0.2455 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3712} │ │ │ │
│ 18 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 104, │ 3.00B │ 24.42B │ 0.2503 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3328} │ │ │ │
│ 19 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 120, │ 3.00B │ 27.92B │ 0.2587 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3712} │ │ │ │
│ 20 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 64, 'num_moe_experts': 104, │ 3.00B │ 23.68B │ 0.2469 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
└────┴───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┴───────────────┴────────┴────────┘
╭──────────────────────────────────────────────────────────────────────── Best Subnet ─────────────────────────────────────────────────────────────────────────╮
│ export_config {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 104, 'moe_ffn_hidden_size': 1856, │
│ 'moe_shared_expert_intermediate_size': 3072} │
│ active_params 3.00B │
│ params 22.28B │
│ score 0.4783 │
╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
╭───────────────────────────────────────────────────── Pruned Model Stats ──────────────────────────────────────────────────────╮
│ Total Parameters 22.28B │
│ Active Parameters 3.00B │
│ Memory (BF16, seq_length=8192, batch_size=1) weights: 42489.7 MB, kv_cache: 48.0 MB, mamba_state: 23.8 MB, Total: 42561.6 MB │
╰───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
```
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
70546bdd6a |
Enable Python 3.14 wheel support to unblock NGC PyTorch container testing on Ubuntu 26.04 + Python 3.14 (#1386)
Ubuntu 26.04 is here and very soon, NVIDIA PyTorch containers will ship with Python 3.14 requiring us to enable untested support to unblock them <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * DFlash offline speculative decoding training * MXFP4→NVFP4 weight conversion support * Shared hidden-state dump utilities * Updated DeepSeek PTQ calibration defaults * **Chores** * Added Python 3.14 support; updated Python requirement to <3.15 * **Documentation** * Updated installation documentation for Python version compatibility <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> |
||
|
|
b5597546b6 |
Increase gpu_tests CI timeout from 60 to 75 mins
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
785d3a2df6 |
[CI] Bump test containers to latest (#1299)
- Use latest containers for testing in CICD <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Bumped TensorRT-LLM Docker images to 1.3.0rc12 in example and GPU test workflows. * Updated PyTorch container image from 26.01 to 26.03 for GPU tests. * Captured uv lock upgrade output to a temp file, inlined it into PR bodies, and adjusted workflow heredoc/templating and step behavior. * **Documentation** * Clarified an inline comment and simplified a warning message for an ONNX quantization extension. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
c51c1762b3 |
fix: prevent gh-pages repo bloat from doc preview artifacts (#1309)
### What does this PR do? Type of change: Bug fix Fixes gh-pages branch bloat that grew from ~26 MB to ~441 MB in four weeks (nvbug 6099503). Three compounding causes were identified and addressed: 1. **Sphinx `.doctrees/` cache published to gh-pages** — `sphinx-build` was writing its build cache inside `build/html/` which was then uploaded verbatim. Accounts for ~3.3 GB uncompressed across history. 2. **`JamesIves/github-pages-deploy-action` appending a commit on every push** — main-site files accumulated forever with `single-commit: false` (default). 3. **PR preview deploying on every `synchronize` event for all PRs** — `rossjrw/pr-preview-action` re-deployed the full site for every push to any PR regardless of whether docs changed (e.g. PR #1128 triggered 64 preview deploys × ~11 MB each). Changes: - Pass `-d /tmp/doctrees` to `sphinx-build` so `.doctrees/` is never written into `build/html/` - Add `paths: [docs/**, modelopt/**]` filter to `pull_request` trigger so the docs workflow only runs on PRs that touch docs or source code - Set `single-commit: true` on the deploy action so main-site pushes squash into one commit - Deduplicate docs build: `deploy-preview` now downloads the artifact from `build-docs` instead of running a second `sphinx-build` - Set `retention-days: 1` on the artifact since it is only needed for the duration of the workflow run The one-time cleanup (force-push squashed orphan to gh-pages) was already applied separately — repo is now ~59 MB for a full clone vs ~441 MB before. ### Usage N/A — CI/workflow change only. ### Testing - Workflow logic reviewed manually. - The one-time cleanup was verified: `git rev-list --objects --disk-usage origin/gh-pages` now reports ~28 MB; full clone is ~59 MB. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information nvbug 6099503 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Optimized documentation build and deployment workflow in CI/CD pipeline. * Improved pull request documentation preview handling with faster build timeouts and refined artifact management. * Enhanced GitHub Pages deployment configuration for better consistency. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
3d0f0db49e |
[CI] Replace tox with nox, use nemo:26.04 for megatron tests, and simplify CI workflows (#1286)
### What does this PR do? Type of change: New feature / infrastructure improvement Follow-up to #1285 for correct CI test environment for megatron based tests Replaces `tox` + `tox-current-env` with `nox` for all test, lint, docs, and wheel build sessions. The primary motivation was that `tox-current-env` is incompatible with uv venvs in NGC containers (e.g. NeMo's `/opt/venv`) — it picks the system Python via `sys._base_executable` instead of the container's venv Python which has megatron packages pre-installed. Key changes: - **`noxfile.py`** replaces `tox.ini` with GPU, CPU unit, partial-install, pre-commit, docs, and wheel sessions - **GPU sessions** use `venv_backend="none"` (run directly in container env) and `python -m pip/pytest` to avoid PATH mismatches - **uv** is set as the default venv backend (if available) for CPU sessions (faster installs) Also includes CI workflow simplifications: - **`_pr_gate.yml`** new reusable workflow centralizing file-change detection + linux-check wait logic (was duplicated across 3 workflow files) - **Collapsed pr/non-pr job pairs** into single jobs with conditional `runs-on` in `gpu_tests.yml`, `example_tests.yml`, `regression_tests.yml` - **Collapsed `multi-py` / `multi-torch` / `multi-transformers`** into a single `multi-version` matrix job in `unit_tests.yml` - **PR path filtering** for unit test secondary jobs (multi-version, launcher, partial-install) — skipped if no relevant files changed - **Fixed schedule/workflow_dispatch skipping** — jobs with `needs: [pr-gate]` were incorrectly skipped when all pr-gate internal jobs were skipped; fixed by making the gate job always run - **multi-version, launcher, partial-install** now also run on `schedule` / `workflow_dispatch` ### Usage ```bash python -m pip install nox uv # install nox and uv (once) nox -l # list all sessions nox -s gpu_megatron # run a GPU session (inside container) nox -s "unit-3.12(torch_211, tf_latest)" # run a specific unit test combination nox -s "unit-3.12(torch_211, tf_latest)" -R # force-recreate venv (e.g. after dep changes) COVERAGE_PROCESS_START=pyproject.toml nox -s "unit-3.12(torch_211, tf_latest)" # with coverage ``` ### Testing - Ran `nox -l` to verify all session names - Ran `gpu_megatron` session locally inside NeMo container — confirmed it uses `/opt/venv/bin/python` correctly - Manually triggered nightly-runs: - Unit: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608013657 - GPU: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608018763 - Examples: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608017322 ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A — CI infrastructure only - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ (added `nox` and `uv` to `dev-test`, both Apache-2.0) - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — no user-facing changes ### Additional Information Supersedes the tox-current-env workaround in the parent branch. --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
feec81ad2b |
Add the Skip softmax for diffusion (#1166)
### What does this PR do?
Type of change: new feature, new example <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->
<!-- Details about the change. -->
## Summary
- Add skip-softmax sparse attention (BLASST) for diffusion models via
dedicated Triton kernels — an inference kernel with tile skipping and a
calibration kernel with vectorized multi-threshold sparsity measurement
- Add `triton_skip_softmax` method with exponential model calibration
(`scale_factor = a * exp(b * sparsity)`) and log-space fitting for
diffusion models
- Add Triton kernel backends for diffusers and LTX attention dispatch
- Fix calibration to skip RULER dataset generation when user provides
their own `forward_loop` (required for non-LLM models)
## Changes
### Triton kernels (`modelopt/torch/kernels/triton_fa.py`)
- **`_attn_fwd`**: Forward kernel with optional tile skipping — tiles
whose max attention score is far below the running softmax max are
skipped entirely (no V load, no softmax, no accumulation). Runtime
sparsity measurement via atomic counters.
- **`_attn_fwd_calibrate`**: Calibration kernel that computes full
attention while measuring how many tiles would be skipped at each of N
thresholds simultaneously. Uses per-program output buffers (zero atomic
contention) and vectorized multi-threshold comparison.
- **`attention()`** / **`attention_calibrate()`**: Python wrappers for
inference and calibration kernels.
### Kernel backends
(`modelopt/torch/sparsity/attention_sparsity/kernels/`)
- **`diffusers_triton_attention.py`**: Registers `modelopt_triton`
backend in diffusers' attention dispatch. Handles [B, S, H, D] → varlen
layout conversion, calibration/inference mode switching, thread-local
configuration, and counter accumulation.
- **`ltx_triton_attention.py`**: Patches `ltx_core.Attention` modules
for Triton dispatch with the same calibration/inference modes.
### Method
(`modelopt/torch/sparsity/attention_sparsity/methods/triton_skip_softmax.py`)
- `TritonSkipSoftmaxMethod`: Context managers for calibration (→
calibration kernel) and inference (→ forward kernel with tile skipping).
Three threshold priority levels: raw threshold > calibrated scale_factor
> static threshold.
### Calibration
(`modelopt/torch/sparsity/attention_sparsity/calibration/`)
- **`calibrator.py`**: `DynamicThresholdCalibrator` with `fit_logspace`
option — fits exponential model in log space (minimizes relative error)
for diffusion models where scale_factors span many orders of magnitude.
Records observed sparsity range for extrapolation warnings.
- **`calibrate.py`**: Skips RULER dataset when `forward_loop` is
provided; passes `fit_logspace` through from config.
### Config & conversion
- **`config.py`**: `CalibrationConfig.fit_logspace` field (default
False, recommended True for diffusion models).
`skip_softmax_raw_threshold` field for direct threshold mode.
- **`conversion.py`**: Auto-registers diffusers/LTX Triton backends on
`sparsify()`. Updated summary display.
### Example
- **`wan22_skip_softmax.py`**: End-to-end example for WAN 2.2 5B/14B
with baseline, raw-threshold, and calibrated modes. Supports runtime
sparsity reporting.
## Threshold modes
| Mode | How it works | Use case |
|------|-------------|----------|
| **Raw threshold** (`--raw-threshold -0.7`) | Passed directly to kernel
as `skip_threshold_log2` | Quick testing, sweeps |
| **Calibrated** (`--calibrate --target-sparsity 0.5`) | `scale_factor =
a * exp(b * target)`, then `threshold = scale_factor / seq_k` at runtime
| Production use with seqlen adaptation |
| **Static** (default `skip_softmax_threshold=0.1`) | `log2(lambda) *
sm_scale` | Fallback |
## Usage
```bash
# Fixed raw threshold (no calibration)
python examples/diffusers/sparsity/wan22_skip_softmax.py \
--model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
--raw-threshold -0.7 \
--prompt "A cat playing piano" --output out.mp4
# With calibration (log-space fit for diffusion models)
python examples/diffusers/sparsity/wan22_skip_softmax.py \
--model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
--calibrate --target-sparsity 0.5 \
--prompt "A cat playing piano" --output out.mp4
# Dense baseline for comparison
python examples/diffusers/sparsity/wan22_skip_softmax.py \
--model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
--baseline \
--prompt "A cat playing piano" --output baseline.mp4
```
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added skip-softmax sparse attention support for Diffusers models,
enabling efficient video generation
* Added support for both eager and Triton attention backends for sparse
attention
* Added new example script for Wan 2.2 text-to-video generation with
sparse attention optimization
* **Documentation**
* Updated documentation with sparse attention configuration guide and
usage examples
* **Tests**
* Added comprehensive unit tests for kernel backend registration and
skip-softmax functionality
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
|
||
|
|
361f7e391b |
Merge puzzletron compression algorithm (#1121)
### What does this PR do? Implement puzzletron compression algorithm based on Puzzle paper (https://arxiv.org/abs/2411.19146) <details> <summary> Th list of reviewed and merged MRs that resulted in the feature/puzzletron branch</summary> Merging dkorzekwa/any_model to feature/puzzletron [Add anymodel directories to feature/puzzletron by danielkorzekwa · Pull Request #974 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/974) - merged [Draft: anymodel activation scoring by danielkorzekwa · Pull Request #989 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/989) - merged [Draft: Merge anymodel pruning by danielkorzekwa · Pull Request #990 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/990/) - merged [Draft: Merging anymodel:build_library_and_stats by danielkorzekwa · Pull Request #993 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/993) - merged [Dkorzekwa/any model calc one block scores by danielkorzekwa · Pull Request #994 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/994) - merged [Draft: merge any_model: mip_and_realize_models by danielkorzekwa · Pull Request #995 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/995) - merged [Dkorzekwa/any model other modeqls by danielkorztiekwa · Pull Request #1007 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1007/) - merged PR to 1007: https://github.com/NVIDIA/Model-Optimizer/pull/1039 - merged [Dkorzekwa/anymodel gptoss by danielkorzekwa · Pull Request #1020 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1020) - merged [Merge any_model tutorial by danielkorzekwa · Pull Request #1035 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1035) - merged [Merge mbridge distillation for any_model by danielkorzekwa · Pull Request #1036 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1036) - merged [MR branch for the remaining difference between dkorzekwa/any_model an… by danielkorzekwa · Pull Request #1047 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1047) - merged [Dkorzekwa/decilm hf code cleanup by danielkorzekwa · Pull Request #1071 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1071) - merged [Dkorzekwa/decilm hf code cleanup 2 by danielkorzekwa · Pull Request #1073 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1073) - merged [Dkorzekwa/anymodel subblock stats by danielkorzekwa · Pull Request #1085 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1085) - merged [Dkorzekwa/anymodel subblock stats nodecilm by danielkorzekwa · Pull Request #1102 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1102) - merged [Dkorzekwa/decilm cleanup post subblockstats by danielkorzekwa · Pull Request #1103 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1103) - merged [code clean up by danielkorzekwa · Pull Request #1110 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1110) - merged Merging into main: [Activation hooks redesign (reuse hooks component across both minitron and puzzletron) by danielkorzekwa · Pull Request #1022 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1022) - merged [Dkorzekwa/puzzletron use importance hooks from prune by danielkorzekwa · Pull Request #1115 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1115) - merged </details> <!-- Details about the change. --> ### Usage Puzzletron tutorial: https://github.com/NVIDIA/Model-Optimizer/tree/feature/puzzletron/examples/puzzletron ### Testing The main e2e test for compressing 9 models with Puzzletron: https://github.com/NVIDIA/Model-Optimizer/blob/feature/puzzletron/tests/gpu/torch/puzzletron/test_puzzletron.py 2-gpu nightly tests: - https://github.com/NVIDIA/Model-Optimizer/actions/runs/24468209205/job/71501061203 - https://github.com/NVIDIA/Model-Optimizer/actions/runs/24470214159/job/71508152952 ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Puzzletron: end-to-end heterogeneous pruning & NAS workflow with AnyModel support, example pipelines, deployment and evaluation utilities, and tools for converting/pruning and exporting compressed checkpoints. * **Documentation** * Comprehensive Puzzletron tutorials, model-specific guides, evaluator instructions, example configs, and changelog entry. * **Chores** * CI/workflow updates (extras installation, longer GPU test timeout), pre-commit hook exclusion updated, and CODEOWNERS entries added. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com> Signed-off-by: Liana Mikaelyan <lmikaelyan@nvidia.com> Signed-off-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com> Signed-off-by: Daniel Korzekwa <daniel.korzekwa@gmail.com> Signed-off-by: jrausch <jrausch@nvidia.com> Signed-off-by: root <root@pool0-00848.cm.cluster> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Liana Mikaelyan <lmikaelyan@nvidia.com> Co-authored-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com> Co-authored-by: J Rausch <38429553+j-rausch@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
3131195241 |
add: DFlash block diffusion speculative decoding (#1211)
DFlash (Block Diffusion for Flash Speculative Decoding) predicts an entire block of tokens in a single forward pass using masked parallel prediction with KV injection from the target model's hidden states. Key features: - Feature fusion (multi-layer hidden states -> FC + RMSNorm) - KV injection (fused features as K/V in every draft layer with QK-norm) - Random anchor sampling with bidirectional intra-block attention - Logit distillation with exponential loss decay (gamma weighting) - Multi-node DDP training with checkpoint resume - Export to z-lab compatible HF format - Online validation (context-dependent ground truth) Training recipe: modelopt_recipes/general/speculative_decoding/dflash.yaml Results: examples/speculative_decoding/doc/dflash_results.md ### ModelOpt Eval (online validation, osl=512) | Dataset | z-lab | ModelOpt (306K) | Diff | |---------|-------|-----------------|------| | gsm8k | 4.10 | **5.19** | **+1.09** | | MT-Bench | 3.58 | **4.36** | **+0.78** | ### z-lab Official Eval (dflash.benchmark, osl=512) | Dataset | z-lab | ModelOpt (306K) | Diff | |---------|-------|-----------------|------| | gsm8k | **5.00** | 4.08 | -0.92 | | MT-Bench | **3.28** | 2.99 | -0.29 | > z-lab model trained with block_size=16. ModelOpt trained with block_size=8. ## Evaluation Method Impact (gsm8k) | Eval Method | z-lab checkpoint | ModelOpt (306K) | |-------------|-----------------|-----------------| | Fixed GT (ModelOpt eval) | 2.95 | 4.23 | | Online GT (ModelOpt eval) | 4.10 | **5.19** | | z-lab official eval | **5.00** | 4.08 | ### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added DFlash speculative decoding mode with parallel block prediction support. * Included training launchers and MT-Bench evaluation scripts for DFlash models. * Added online acceptance rate validation for improved inference verification. * **Documentation** * DFlash quick start guide with configuration parameters and training examples. * Performance results and benchmarks for DFlash-trained models. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
04cd596d79 |
Add experimental support for transformers>=5.0 + min torch 2.8 (#975)
### What does this PR do? - Add experimental support for transformers >=5.0 and remove deprecated usages: https://github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md - ⚠️ For accelerate examples that used `--warmup-ratio: float` (deprecated in 5.x), we now change it to `--warmup-steps: float | int` which works as ratio if float but only for 5.x. For 4.x, it will error out if float and prompt user to change back to `--warmup-ratio` or pass an int absolute step count. - ⚠️ Unified Hugging Face checkpoint export for quantized checkpoints may not work for some models with transformers>=5.0 yet as it requires a lot of fixes (e.g. change in how MoE experts are organized) - ~Add Workaround for TRT-LLM's import of deprecated transformers functions so trt-llm based gpu unit tests work fine. Still deployment for models needs proper fixes directly in TRT-LLM hence llm/vlm ptq example tests still run with transformers 4.57~ - Everything except PTQ and Export (mainly MoE) should work fine with transformers>=5.0 - Bump min torch to 2.8 and enable 2.11 cicd testing - NOTE: Upcoming Nemo:26.04 container comes with transformers 5.3 ### Testing <!-- Mention how have you tested your change if applicable. --> - [x] CI/CD tests passing - [x] Manually tested unit tests, gpu tests with transformers 4.56 and 5.4 - [x] Manually tested example tests (except trt-llm container tests) with transformers 4.56 and 5.4 - [x] 2-gpu nightly CICD tests manually triggered and passing: [gpu tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867257540), [example tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867260643) ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, using `torch.load(..., weights_only=True)`, avoiding `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other source, did you follow IP policy in [CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?: N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Make remote-code usage opt-in via a configurable --trust_remote_code flag across examples and tools. * **Bug Fixes** * Improve checkpoint/resume detection and related training guidance to avoid erroneous errors. * **Refactor** * Consolidate dtype/config naming, switch warmup settings from ratio → steps, and unify tokenizer invocation patterns. * **Documentation** * Simplify changelog title and add misc notes for release 0.44. * **Chores** * Remove scheduled PR-branch cleanup workflow and relax/remove several transformers version pins. * **Tests** * Adjust test gates, skips, and structures to align with updated deps and behaviors. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
ba4f42df1c |
Minor fix for example tests
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
80d2f02a2d |
Fix spec dec example tests (#1183)
### What does this PR do? Type of change: Test fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> - Fix `tests/examples/speculative_decoding` - previously silently skipped - Avoid pulling nemotron-post-training-dataset-v2 in tests to reduce chances of HF loading timeout in CICD - Make slow and redundant tests manual to speed up CICD ### Testing <!-- Mention how have you tested your change if applicable. --> - Tests passing ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Removed git‑LFS install step from CI and deleted an automated branch‑cleanup workflow * Trimmed example environment dependencies and relaxed transformers compatibility; added an optional tokenization dependency * **Tests** * Switched tests to generate datasets dynamically and improved fixture handling * Standardized PTQ test parameters (explicit calibration dataset) and refined GPU/test selection * **Bug Fixes** * Improved device-awareness and numeric handling in speculative decoding attention paths <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
2ae407c559 |
Include gpu and example tests also in codecov coverage reporting and enable omitted folder coverage (#1154)
So far, we only measured unit test coverage but we also have gpu test and example tests which needed to be setup differently to track in overall codecov coverage so we get accurate coverage reporting <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Strengthened code coverage measurement with parallelized collection across unit, GPU, and example test suites * Enhanced continuous integration workflow configuration with improved coverage reporting and threshold management * Updated testing infrastructure dependencies and settings to support more robust quality assurance processes <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
8f4c11aa44 |
Upgrade the TensorRT container version (#1112)
### What does this PR do? Type of change: Container version update - Upgraded the TensorRT container version to 26.02 - This supports TensorRT 10.15.1 ### Testing Unit and integrations tests pass ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated example guides to reference the latest TensorRT Docker image version (26.02) * Added compatibility note for onnxruntime-gpu usage * **Chores** * Updated CI/CD workflows to use the latest TensorRT Docker image version <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> |