mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
e2c4d083d40976c38bf7efd9a05628ef5eed80a5
65
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e2c4d083d4 |
[OMNIML-4922] Four over Six PTQ & Updating Nemotron Ultra Example (#1684)
### What does this PR do? [Four Over Six](https://arxiv.org/pdf/2512.02010) PTQ implementation for weight-only quantization. Four Over Six was used to produce the Nemotron 3 Ultra NVFP4 checkpoint. Also updates the Ultra PTQ example in the launcher to use this new 4/6 config `huggingface/nvidia/Nemotron-3-Ultra-550B-A55B/ptq/ultra-nvfp4-46-max` ### Usage ```bash uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/megatron_lm_ptq.yaml --yes ``` ### Testing - [x] Unit tests pass - [x] Run launcher example ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added NVFP4 Four‑Over‑Six (4/6) adaptive per‑block weight scaling and a configurable FP8 normalization option for FP4/FP8 quantization. * **Documentation** * Added PTQ recipes/presets and updated config docs to document FP8 max variants and the Four‑Over‑Six option. * **Tests** * Added unit and GPU tests validating 4/6 selection, normalization threading, scaling behavior, and reconstruction error checks. * **Chores** * Updated a PTQ pipeline example to use the NVFP4‑46‑max recipe and bumped a container image; minor project config tweak. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> Signed-off-by: Jenny Chen <jennifchen@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
37dbbdac5a |
Fix ModelOpt MCP Slurm launcher submit (#1799)
## Summary - fix launcher Slurm task annotation patching so nemo-run CLI resolves `slurm_factory` correctly for task slots - harden `modelopt-mcp` submit parsing/status resolution and add regression coverage for launcher false-positive success cases - add a minimal `nvidia-smi` smoke YAML/script and fix launcher packaging so source-backed Slurm jobs package required files recursively ## Validation - `uv run pytest tests/test_core.py -q` - `uv run pytest tests/test_bridge.py -q` - dry-run and live-submit validated through the patched local MCP server on `cw_dfw` - interactive smoke job succeeded end-to-end (`nvidia-smi` ran successfully in-container) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an NVIDIA SMI GPU smoke test (script and minimal Slurm YAML example) for launcher integration. * **Bug Fixes** * Improved detection of fatal launcher errors, including when the launcher exits with code 0. * Strengthened Slurm experiment/job identifier parsing and added early rejection of unsafe experiment IDs, with clearer “unparsed”/failure behavior. * Updated sandbox task Slurm config type handling and improved launcher packaging so examples/common are included consistently. * **Tests** * Expanded unit and filesystem-based coverage for parsing/validation, dry-run fatal stderr handling, and nested experiment directory layouts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Signed-off-by: Chenhan D. Yu <chenhany@nvidia.com> |
||
|
|
28b5e26fdb |
Fix torch import error to remove circular dependency & move Nemotron configs (#1606)
### What does this PR do? Type of change: Bug fix when running Megatron-LM modelopt example `generate.py` a circular import causes an error. the cause was because it import modelopt.torch.quantization which imports modelopt/torch --> ``` modelopt/torch/__init__.py:26. The chain is: import modelopt.torch.opt → runs modelopt/torch/__init__.py (parent package) → which pulls in distill → distill/mode.py:25 → back into opt before opt.utils is bound. ``` ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Clarified configuration comments in Nemotron-3-Super-120B quantization recipes for improved clarity. * **Chores** * Internal package initialization updates. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> |
||
|
|
985ad4361c |
launcher: make Slurm memory and user defaults configurable (#1791)
## Summary - Adds `SlurmConfig.mem` and a `SLURM_MEM` env override for launcher Slurm jobs. - Passes configured memory through to `nemo_run.SlurmExecutor` instead of always using `"0"`. - Lets `SLURM_USER` provide the launcher default user when local and cluster usernames differ. - Adds focused tests for Slurm memory defaults and overrides. ## Test plan - [x] `uv run pytest tools/launcher/tests/test_slurm_config.py tools/launcher/tests/test_slurm_executor.py` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Slurm job memory can now be configured via `SLURM_MEM`, with a default of `0` when unset. * Job launch username can now be set via `SLURM_USER`; if omitted, it falls back to the local login name. * **Bug Fixes** * Slurm executor now correctly uses the configured memory value from Slurm settings, falling back to `0` only when missing/empty. * **Tests** * Expanded unit tests to cover default, environment-driven, and executor parameter memory behavior (including missing/empty/`None` cases). <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
2a88b6056c |
launcher: package as modelopt_launcher; mcp: call console script directly (#1766)
## Summary
- `tools/launcher/__init__.py` gains `PACKAGE_DIR`, making the directory
an importable package named `modelopt_launcher` via a `package-dir`
mapping in `pyproject.toml`. **No file moves** — `common/` and
`examples/` stay exactly where they are.
- `pyproject.toml` adds the `packages`/`package-dir`/`package-data`
declarations and a `modelopt-launcher` console script entry point.
- `launch.py` switches imports to
`modelopt_launcher.{core,slurm_config}` (works in both `uv run
launch.py` via editable install and the installed console script), adds
`_has_modelopt_src` to skip packaging modelopt source when running
installed (cluster container already has it), and adds `main()`.
- `bridge.py` deletes the 75-line `_find_launcher_dir()` filesystem
walker and `_launcher_dir_not_found_response()`; simplifies
`_find_launcher_examples_dir()` to 2 strategies (env override → `import
modelopt_launcher`); switches `submit_job` subprocesses from `["uv",
"run", "launch.py"]` with `cwd=launcher_dir` to `["modelopt-launcher"]`
with no `cwd`.
- `tools/mcp/pyproject.toml` declares `modelopt-launcher` as a proper
dependency (dev: editable `../launcher`; published: PyPI), replacing a
lengthy comment explaining why it could not be declared.
- 7 tests for the deleted functions removed; all 38 remaining tests
pass.
## Test plan
- [ ] `cd tools/launcher && uv run python3 -m pytest tests/ -v` — all 65
tests pass
- [ ] `cd tools/mcp && uv run python3 -m pytest tests/ -v` — 38 tests
pass
- [ ] `uv run launch.py --yaml
examples/Qwen/Qwen3-8B/megatron_lm_ptq.yaml --dryrun --yes -v` — dry-run
resolves correctly
- [ ] `modelopt-launcher --yaml
examples/Qwen/Qwen3-8B/megatron_lm_ptq.yaml --dryrun --yes` — console
script works after `pip install -e tools/launcher`
- [ ] `cd tools/mcp && uvx modelopt-mcp` — MCP server starts without
FileNotFoundError
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added a `modelopt-launcher` CLI command as the standardized entry
point for launching optimization jobs.
* **Bug Fixes**
* Improved launcher detection and error reporting when the launcher
isn’t installed.
* Simplified experiment-directory discovery and standardized environment
handling for job submission and log retrieval.
* **Chores**
* Updated launcher packaging and example/resource discovery to work
reliably from installed distributions.
* Added dev-mode symlink and cleanup safeguards.
* **Tests**
* Adjusted expectations for draft PR failure behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
||
|
|
12ae5fbb35 |
[OMNIML-5233] hf_synth.yaml: relative-leaf output_dir contract (#1773)
## What does this PR do? **Type of change:** Bug fix **Overview:** Fix the `output_dir` shape for `tools/launcher/examples/Qwen/Qwen3-8B/hf_synth.yaml` so the downstream synth_run stage (OMNIML-5233) can submit + resume + publish correctly. Replaces the prior `/scratchspace/modelopt/qwen3-8b-synth-v1` (which fails resume + publish) with the canonical **relative-leaf** shape: `hf-local/modelopt/qwen3-8b-synth-v1`. The absolute team-folder prefix is **deliberately not committed** — it's an NVIDIA-internal cluster mount path, and `pensieve-intern`'s `_scan_for_internal_path_leak` guard refuses any YAML diff that bakes it (rightly, since this repo is public). The downstream stage that has the cluster context (synth_run) resolves the prefix via `mcp__nmm-sandbox__resolve_team_folder` at submit time and injects the absolute path through `extra_overrides`. **Why this shape:** the relative leaf `hf-local/modelopt/<id>` carries the contract between authoring (this stage), submitting (synth_run), and publishing (`modelopt-storage publish`). Publish promotes from the team's `hf-local/` tree as a metadata-only op; the leaf shape tells the operator + downstream what category the artifact lands in, without leaking the cluster mount. **Companion MR:** [pensieve-intern !126](https://gitlab-master.nvidia.com/omniml/integration/pensieve-intern/-/merge_requests/126) lands the matching SPEC contract in synth_support.md + synth_run.md, AND elevates the path-contract rule to the engine preamble (`_ENGINE_AGENT_RULES_PREAMBLE`) so every agent dispatch — across agent/subprocess/pensieve-artifact tasks — reads it as workflow-wide common knowledge. ## Usage Same as before. The shard write location is now resolved at submit time by synth_run. ## Testing - [x] One-line YAML change. - [ ] End-to-end: re-fire OMNIML-5233 synth_run after this + the pensieve-intern MR merge; expect the agent to submit the slurm array and stamp success. ## Before your PR is "Ready for review" - [x] Make sure you read and follow Contributor guidelines and your commits are signed. - [x] Is this change backward compatible? **Yes** (Qwen3-8B is the only consumer; no published checkpoints reference the old path). - [x] Did you write any new necessary tests? **N/A** — example YAML; behavior covered by the downstream synth_run agent run. - [x] Did you add or update any necessary documentation? Covered in the comment block on the changed lines. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated the configured output directory to a more stable, persist-friendly path so generated artifacts remain available across repeated runs and re-dispatches. * Added clearer inline guidance explaining how the resolved absolute path under the new directory is used to reuse prior shards and to handle publishing consistently. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
fa1d13f85d |
launcher: add Nemotron-3-Super-120B-A12B-BF16 MTP vLLM specdec bench config (#1714)
Adds a SPEED-bench MTP speculative-decoding YAML for `NVIDIA-Nemotron-3-Super-120B-A12B-BF16` via vLLM. Covers two splits: - `qualitative` — 32 concurrent, 4096 output tokens - `throughput_32k` — 8 concurrent, 80 requests, 4096 output tokens Both tasks run `tp_size=4` on a single 4×H100/A100 node. Part of OMNIML-5095 / OMNIML-5098. ## Test plan - [ ] `uv run slurm.py --yaml modules/Model-Optimizer/tools/launcher/examples/Nemotron-h/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/specdec_bench_mtp_vllm.yaml --dryrun --yes -v` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Chores** * Added a new benchmark configuration for evaluating NVIDIA Nemotron-3-Super-120B model performance using speculative decoding with vLLM. Configuration includes qualitative and throughput benchmark tasks. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
93dd08f429 |
[OMNIML-4760] synth_support (#1696)
Draft PR opened by **pensieve-intern** for [OMNIML-4760](https://jirasw.nvidia.com/browse/OMNIML-4760). Stage `synth_support` of Epic `OMNIML-4755`. The agent ran from the SPEC on the ticket description; review every change before marking ready. _Always-draft is enforced — the bot never auto-merges._ --- **Agent's self-narration** (stripped from PR diff; surfaced here for context): `VERIFICATION_COMMENT.txt`: ``` OMNIML-4760 synth_support verification for qwen3-8b found and fixed one task_0 issue in `tools/launcher/examples/Qwen/Qwen3-8B/hf_offline_eagle3.yaml`. What changed: - Updated `task_0` data input from missing `/hf-local/modelopt/Speculative-Decoding-Prompt-Samples` to existing `/hf-local/modelopt/Speculative-Decoding-Dataset-v1-Qwen3-8B/default-openai.jsonl`. - Left the rest of the 4-task monolithic `hf_offline_eagle3.yaml` unchanged. Verification: - `task_0.script` is `common/tensorrt_llm/query.sh`. - `--model <<global_vars.hf_model>>` resolves to `/hf-local/Qwen/Qwen3-8B`. - `task_0.slurm_config.container` is `nvcr.io/nvidia/tensorrt-llm/release:1.2.0`. - Cluster validation on cw_dfw succeeded for the fixed task_0 data path: experiment `cicd_1781221901`, Slurm job `12739762`, remote directory `/lustre/fsw/portfolios/coreai/users/chenhany/experiments/cicd/cicd_1781221901/Qwen3-8B_EAGLE3_offline_task0_verify_jsonl_0`. - Log evidence included successful SSH tunnel/authentication, TensorRT-LLM 1.2.0 container import, Qwen3-8B server health checks, a successful `/v1/chat/completions` request, and loading `1393367` train examples from the Qwen3-8B JSONL data file. Next: - Runner should open the single-file PR for human review because this was a task_0 config fix, not verification-only. ``` _Pollution-strip removed `VERIFICATION_COMMENT.txt` from this commit (sidecar narration and/or incidental lockfile regeneration are never part of the agent's intended deliverable)._ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Enhanced TE grouped MoE weight quantization with optional per-GEMM/per-expert quantizers (environment-controlled) and updated calibration flow * Expanded launcher CLI with `--shard-id` and per-run sample limiting via `--num-samples` * Added a Qwen3-8B standalone vLLM synthesis job example (`hf_synth.yaml`) * Added Slurm job requeue support * **Bug Fixes** * Improved sharded dataset synthesis with idempotent `.done` markers and smarter shard sizing/capping * **Tests** * Added coverage validating TEGrouped vs sequential MoE default amax behavior and output divergence/accuracy <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Pensieve Intern <pensieve-intern@nvidia.com> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Pensieve Intern <pensieve-intern@nvidia.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
e0125294f9 |
launcher: add Qwen3-8B/specdec_bench_dflash_vllm.yaml parent (OMNIML-5057) (#1764)
## Summary Adds the parent YAML for the Qwen3-8B / DFlash / vLLM SPEED-bench sweep (Epic OMNIML-5057, 4 cells: t0_d3 / t0_d7 / t1_d3 / t1_d7). Mirror of [`Qwen/Qwen3.5-4B/specdec_bench_dflash_vllm.yaml`](https://github.com/NVIDIA/Model-Optimizer/blob/main/tools/launcher/examples/Qwen/Qwen3.5-4B/specdec_bench_dflash_vllm.yaml). ### Differences vs Qwen3.5-4B template - `hf_model`: `/hf-local/Qwen/Qwen3-8B` - `draft_model`: `/hf-local/modelopt/Qwen3-8B-DFlash-bs8-seq4096-150000` — an internal NVIDIA checkpoint, not on HF Hub. `cell.md` Step 1's `PRIVATE_DRAFT_ORGS` branch already covers the expected 404; the cluster has the file staged by the Epic's `prep_inputs` stage. - `block_size`: 8 (was 4) — matches the canonical `cell_t0_d7` draft_length=7. ### Known limitation (acceptable per Shape (2)) Qwen3-8B's `max_position_embeddings = 40960`. The `throughput_32k` split contains rows whose prompts exceed 40960 input tokens; those rows fail with the vLLM context-limit assert. Per `cell.md` "Shape (2)" recovery contract, cells ship qualitative metrics + `null` throughput_32k AL. See OMNIML-5060 Notes (2026-06-15 correction) for the empirical evidence. ### Reference run `cicd_1781655951` (OMNIML-5060 pipeline #55054230, cw_dfw): - qualitative `Average_AL` = **3.4849** - throughput_32k = `null` (Shape (2) — overlong rows skipped) ### Cascade 3 subprocess cells (OMNIML-5059 / 5061 / 5062) re-run slurm against this parent YAML at invoke time with per-cell CLI overrides for temperature / block_size / save_dir / max_seq_len. ## Test plan - [x] `tools/precommit/check_launcher_yaml.py` passes locally - [x] YAML successfully drives a real slurm submission (`cicd_1781655951`) producing valid qualitative SPEED-bench output - [ ] Once merged: `intern_advance OMNIML-5057` fires the 3 subprocess cells against this YAML 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Added benchmark configuration file for evaluating speculative decoding performance with Qwen3-8B model using DFlash acceleration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
e6790ef7b4 |
[Examples]: GPT-oss, Qwen3Moe streaming specdec example (#1692)
### What does this PR do?
Type of change: new example
Adds **streaming speculative-decoding examples (EAGLE3 + DFlash)** for
**gpt-oss-20b** and **Qwen3-30B-A3B** to the ModelOpt launcher,
mirroring the existing Qwen3-8B/Kimi examples.
- New yamls:
`tools/launcher/examples/{openai/gpt-oss-20b,Qwen/Qwen3-30B-A3B}/hf_streaming_{eagle3,dflash}_multi_node.yaml`,
plus gpt-oss `chat_template_train.jinja` (generation-tagged, for
`answer_only_loss`).
- `eagle_utils.py`: the streaming path now installs a custom
`data.chat_template` on the tokenizer (the online path already did) —
needed for the tagged template.
### Usage
```bash
cd tools/launcher
export SLURM_HOST=... SLURM_ACCOUNT=... SLURM_HF_LOCAL=... SLURM_JOB_DIR=...
uv run launch.py --yaml examples/openai/gpt-oss-20b/hf_streaming_eagle3_multi_node.yaml --yes
```
### Testing
Pipeline sanity test on **unsynthesized** data (daring-anteater), 1×
H100-80GB, 12k steps. All four train and pass the vLLM acceptance-length
eval:
| Model | Method | Train speed | vLLM AL |
|---|---|---|---|
| Qwen3-30B-A3B | EAGLE3 | 7.12 it/s | **1.74** |
| Qwen3-30B-A3B | DFlash | 2.31 it/s | 1.29 |
| gpt-oss-20b | EAGLE3 | 5.07 it/s | 1.19 |
| gpt-oss-20b | DFlash | 2.01 it/s | 1.14 |
<img width="1300" height="780" alt="image"
src="https://github.com/user-attachments/assets/2be22562-5b77-4a9f-9dc0-6f936a059736"
/>
> Sanity test only, not a quality run. gpt-oss AL is low because it is a
reasoning model (CoT at inference) while daring-anteater has no
reasoning traces and `answer_only_loss` masks all but the final content
— quality runs need synthesized/reasoning data.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added multi-node speculative decoding pipeline configurations for
Qwen3-30B-A3B and gpt-oss-20b with DFlash and EAGLE3 support.
* Introduced chat template training support for improved model
instruction formatting.
* **Enhancements**
* Increased benchmark concurrency from 1 to 32 across Qwen3-8B
configurations for more realistic performance evaluation.
* Extended training runs from 500 to 2000 steps for Kimi-K2.5 models.
* Improved chat template handling in speculative decoding workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
e004d8d90e |
DFlash speculative decoding for MiniMax-M2.7 (FSDP2): auto mask-token, FSDP2 resume fixes, per-checkpoint draft export (#1621)
## What
Brings up DFlash block-diffusion speculative decoding for large MoE
targets (MiniMax-M2.7, 229B) trained under accelerate FSDP2, and fixes
the regressions that broke checkpoint resume and per-checkpoint draft
export.
## Commits
- **auto-add mask token for DFlash** when the tokenizer lacks one
(resize embeddings, restore dtype).
- **requeue support** in `build_slurm_executor` + **FSDP2
cpu_ram_efficient_loading** for 229B on multi-node.
- **FSDP2 buffer patch** (`fsdp2_buffer_patch.py`): handle non-DTensor
buffers in `fsdp2_load_full_state_dict`, broadcast dtype codes from rank
0, and an FSDP2-safe `clip_grad_norm_`. Required because MiniMax-M2.7
pins transformers 4.57.x (no native `ParallelismConfig`).
- **dtype fix**: use the broadcast dtype (rank 0) rather than the local
meta-device param dtype, so non-leader ranks don't cast bf16 back to
fp32 on resume.
- **restore `DFlashExportCallback`** (this PR's headline): the
Pydantic-recipe refactor (
|
||
|
|
601401b134 |
[OMNIML-5025] cell_t0_d7 (#1738)
Draft PR opened by **pensieve-intern** for [OMNIML-5025](https://jirasw.nvidia.com/browse/OMNIML-5025). Stage `cell_t0_d7` of Epic `OMNIML-5022`. The agent ran from the SPEC on the ticket description; review every change before marking ready. _Always-draft is enforced — the bot never auto-merges._ --- **Agent's self-narration** (stripped from PR diff; surfaced here for context): `INTERN_ARTIFACTS.json`: ``` { "sweep_name": "gemma-4-E4B-it_mtp_vllm_t0_d7", "experiment_id": "cicd_1781548540", "experiment_dir": "/lustre/fsw/portfolios/coreai/projects/coreai_dlalgo_modelopt/cicd/cicd/cicd_1781548540/", "AL_qualitative_overall": 3.2945, "AL_qualitative_categories": { "coding": 4.8883, "humanities": 2.703, "math": 4.1139, "multilingual": 4.1401, "qa": 2.59, "rag": 3.9329, "reasoning": 3.6251, "roleplay": 1.8026, "stem": 3.2041, "summarization": 2.8275, "writing": 2.4118 }, "AL_throughput_32k_overall": 3.3803, "AL_throughput_32k_categories": { "high_entropy": 2.0839, "low_entropy": 4.3797, "mixed": 3.7143 } } ``` `VERIFICATION_COMMENT.txt`: ``` Completed OMNIML-5025 cell_t0_d7 for `google/gemma-4-E4B-it` / MTP / vLLM. What was done: - Authored/kept `tools/launcher/common/specdec_bench/_cells/gemma-4-E4B-it_mtp_vllm_t0_d7.yaml` for sweep `gemma-4-E4B-it_mtp_vllm_t0_d7`. - Updated/kept `tools/launcher/examples/google/gemma-4-E4B-it/specdec_bench_mtp_vllm.yaml` with the Gemma-specific `vllm/vllm-openai:gemma` container after generic v0.22.1 failed startup. - Verified HF Hub model endpoints for `google/gemma-4-E4B-it` and `google/gemma-4-E4B-it-assistant` returned HTTP 200. - Wrote `INTERN_ARTIFACTS.json` with parsed AL metrics from the successful cluster run. Metrics extracted: - experiment_id: `cicd_1781548540` - experiment_dir: `/lustre/fsw/portfolios/coreai/projects/coreai_dlalgo_modelopt/cicd/cicd/cicd_1781548540/` - qualitative Average_AL: `3.2945` - throughput_32k Average_AL: `3.3803` PR status: - PR opened: NONE — this runner prompt says not to commit, push, or create PRs because the runner handles that. What's next: - Engine should consume `INTERN_ARTIFACTS.json` and advance the downstream wrap-up/aggregation stage. Trigger pipeline URL: - `https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/pipelines/54841359` Slurm job status: - submission: completed successfully - experiment_id: `cicd_1781548540` - experiment_dir: `/lustre/fsw/portfolios/coreai/projects/coreai_dlalgo_modelopt/cicd/cicd/cicd_1781548540/` ``` _Pollution-strip removed `INTERN_ARTIFACTS.json`, `INTERN_LEARNING_TICKET.md`, `VERIFICATION_COMMENT.txt` from this commit (sidecar narration and/or incidental lockfile regeneration are never part of the agent's intended deliverable)._ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated benchmark configuration to use an optimized container image for Gemma 4 MTP speculative-decoding benchmarks. * Revised documentation in the configuration to reflect the current image and its capabilities. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: pensieve-intern agent <noreply@nvidia.com> |
||
|
|
c661366352 |
launcher: Nemotron-120B specdec_bench — kill stale _cells/ reference (post-PR-#1564) (#1741)
## What Comment-only change to the Nemotron-3-Super-120B-A12B-BF16 DFlash parent YAML doc-comment header. Rewrites the example invocation from the pre-PR-#1564 pattern (with `--runtime_params common/specdec_bench/_cells/<sweep_name>.yaml`) to the current post-#1564 pattern (CLI-flag-only overrides, no per-cell file). No yaml content / config change — only the prose example in the header. ## Why PR #1564 removed the `tools/launcher/common/specdec_bench/_cells/` and `_runtime_params/` directories. Cell-specific knobs are now CLI overrides at slurm-invoke time, not committed files. The cell SPEC in pensieve-intern (`specdec_bench/specs/cell.md`) is explicit about this — five times in prose. But this parent YAML's doc-comment kept advertising the OLD pattern. When pensieve-intern's agent runner scans `tools/launcher/examples/` for a reference invocation (good practice — agents should imitate working examples), it lands on this Nemotron parent and copies the stale pattern. **Five recent agent dispatches on the gemma-4 Epic OMNIML-5022 (cells OMNIML-5024 / 5025 / 5026 / 5027) authored new `_cells/<sweep>.yaml` files for this reason, despite the SPEC telling them not to.** Prose loses to a concrete checked-in counter-example. The Qwen3.5-4B reference template that cell.md officially points at (`tools/launcher/examples/Qwen/Qwen3.5-4B/specdec_bench_mtp_vllm.yaml`) is clean and shows only the bare `uv run slurm.py --yaml ...` form. This PR makes Nemotron-120B consistent with that template. ## How surfaced Diagnosed 2026-06-15 on OMNIML-5025 cell_t0_d7 ([intern-agent job 341631795](https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/341631795)). The cell agent authored `_cells/gemma-4-E4B-it_mtp_vllm_t0_d7.yaml` — the engine's diff-shape classifier should have rejected it, but didn't (tracked separately as OMNIML-5170). Root-causing the agent's behavior surfaced this stale doc-comment as the source of the pattern. ## Verification `grep -rn '_cells\|runtime_params common'` across the entire launcher tree returned only this file. After this PR, the launcher tree carries zero stale references. Pairs with NVIDIA/Model-Optimizer#1738 (gemma-4-E4B-it container fix) and OMNIML-5170 (engine-side classifier hard-reject for `_cells/` paths — defense in depth). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated example documentation to clarify how to override per-cell parameters using CLI flags in Slurm configuration runs. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
55a2101e2f |
Update Nemotron-3 Pruning, Distillation and PTQ results based on new shared calibration loop with seq packing and add tool-calling eval fix (#1660)
### What does this PR do? Type of change: documentation + minor example-script tweaks Follow-up to #1601. Originally scoped to add **NVFP4 + QAD**, this PR was **repurposed** to refresh the [Nemotron-3-Nano-30B-A3B-BF16 tutorial](examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md) results using the **new shared calibration loop (sequence packing)** and to **fix tool calling in evaluation**. - Refreshed the prune → distill → eval → **FP8** results (accuracy + vLLM throughput tables) with the new calibration loop. - **Tool-calling eval fix** (`nemo_evaluator.yaml`): GPQA and AIME now run the Python sandbox tool. The tutorial reports both **with-tools** and **no-tools** GPQA/AIME and shows `mean ± std_dev`. - Script tweaks: `quantize.py` calibration now uses sequence packing (`pack=True`) which leads to slight improvement in PTQ; `prune_minitron.py` defaults `inference_batch_size` to `calib_batch_size`. ### Testing Documentation + small example-script changes; tutorial relative links resolve and the results tables / figure were verified consistent. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: Yes - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ (will run `/claude review`) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the main guide and evaluator instructions for prune + distill + FP8/NVFP4 quantization, including refreshed vLLM deployment tips, benchmark/noise presentation, and long-context tool-calling attribution notes. * Refreshed README technique examples/links, reordered the model support matrix rows, and improved pruning overview/support-matrix text. * **Changes to Examples** * NAS pruning now documents higher GPU memory usage vs manual pruning; pruning batching defaults were improved. * Quantization PTQ calibration uses packed document packing; quantized checkpoint export messaging was streamlined. * Updated pruning/distillation/quantization tutorial guidance, metrics/tables, command parameters, and evaluator YAML settings (KV-cache dtype, generation defaults, task behavior). <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
d7df14d12a |
[tools/debugger] Enforce a single relay owner across hosts (#1735)
### What does this PR do? Type of change: Bug fix The `tools/debugger` file-based relay assumes a single server but never enforced it. Because the relay lives on shared NFS (the repo is often the same checkout mounted on multiple hosts), a forgotten `server.sh` on another host kept polling the same `.relay/` and could **steal commands** (executing them on the wrong host), and killing one server's cleanup could **wipe the active server's markers**. This adds a `.relay/owner` ownership token (`host:pid:nanos`): - Each server writes `owner` atomically at startup and **takes over** instead of refusing when a stale `server.ready` exists (the old `kill -0 <pid>` guard was host-local and meaningless across hosts). - The handshake and main loops exit cleanly if `owner` changes (`[server] Superseded by <id> — exiting.`), so a freshly started server **evicts** any stale one — even on another host. - `cleanup()` only clears shared markers if we still own them, so a stepping-down server never clobbers its successor's `server.ready`/`owner`. Also gitignores `tools/debugger/logs/` and documents the `owner` file in the README. ### Usage ```bash # Inside the container; a previously-running server elsewhere that shares this # NFS .relay/ steps down automatically once this one claims ownership: bash tools/debugger/server.sh # [server] Note: existing server.ready found (<host:pid:ts>); taking over. # (the stale server logs: "[server] Superseded by <id> — exiting.") ``` ### Testing Verified live on computelab: a forgotten `server.sh` on another host was evicted when a new server started, after which `client.sh run` executed on the correct (new) host; confirmed the stepping-down server's cleanup does not remove the successor's `server.ready`/`owner`. `server.sh` passes `bash -n`. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ <!-- additive: new .relay/owner file; client.sh and the wire protocol are unchanged --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A <!-- the file-based relay tool has no test harness; behavior verified manually --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- internal dev tooling, not a shipped feature/API --> - Did you get Claude approval on this PR?: N/A <!-- can run /claude review --> ### Additional Information Scope is limited to `tools/debugger/` (`server.sh`, `README.md`, `.gitignore`). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated relay protocol documentation to clarify ownership-based server coordination. * **Bug Fixes** * Improved reliability of multi-server coordination in shared relay environments. * **Chores** * Updated ignore patterns for logging files. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6968fe78fd |
feat(tools/mcp): resolve launcher dir via env override + cwd walk-up (#1731)
## Summary Fixes Symptom A of [OMNIML-5151](https://jirasw.nvidia.com/browse/OMNIML-5151): the bridge's launcher-dir resolution doesn't work in `uv tool install` layouts (which is how the intern-agent CI installs modelopt-mcp). Empirically validated by nmm-sandbox pipelines [54755376](https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/pipelines/54755376) and [54829964](https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/pipelines/54829964) — the agent reached for `mcp__modelopt__submit_job`, got `launcher_dir_not_found`, fell back to shell. ## Root cause 5 sites used: ```python launcher_dir = _THIS_DIR.parent.parent / "launcher" ``` Works in: - `pip install -e tools/mcp` (dev) — `_THIS_DIR` is `<repo>/tools/mcp/modelopt_mcp/` - Direct `Model-Optimizer` clone — same path Doesn't work in: - `uv tool install --from "git+...#subdirectory=tools/mcp" modelopt-mcp` — `_THIS_DIR` is `~/.local/share/uv/tools/modelopt-mcp/lib/.../site-packages/modelopt_mcp/` and `parent.parent` doesn't reach a launcher ## How New `_find_launcher_dir()` helper resolves in this order: 1. `$MODELOPT_LAUNCHER_DIR` env override (deterministic) 2. `_THIS_DIR.parent.parent / "launcher"` (in-repo layout) 3. Walk up from `os.getcwd()` looking for `modules/Model-Optimizer/tools/launcher` (agent workspace) or `tools/launcher` (direct clone) Step 3 specifically unblocks intern-agent: the agent's cwd is inside its cloned nmm-sandbox workspace where `modules/Model-Optimizer/tools/launcher/` does exist — just needs the walk-up. Centralizes the structured-failure response too — five callsites had slightly different `launcher_dir_not_found` shapes; new `_launcher_dir_not_found_response()` helper produces a consistent dict listing the searched paths so the next failure is obvious. ## Sites updated - `submit_job_impl` — live submission - `_submit_job_dry_run` — dry-run validation - `_resolve_experiment_dir` — soft fallback (was crashing on `None.exists()` before) - `read_cluster_artifact_impl` — cwd for `nemo experiment logs` ## Tests 7 new hermetic tests in `tests/test_bridge.py`: - env override wins - env-points-at-ghost falls through gracefully - walk-up via `modules/Model-Optimizer/tools/launcher` - walk-up via plain `tools/launcher` - returns `None` when nothing is found - structured-failure response shape (with and without `dry_run` flag) All 45 tests pass (38 + 7 new). Pre-commit clean (ruff / mypy / bandit / license-headers). ## Out of scope Symptom B of OMNIML-5151 — `NEMORUN_HOME` env not propagated to MCP subprocess, affecting `job_status` / `wait_for_experiment` / `read_cluster_artifact(path=None)`. That's a cross-repo change in pensieve-harness's `mcp_config` writer + pensieve-modelopt's `intern_runner`. Follow-up PR. ## Anchors - [OMNIML-5151](https://jirasw.nvidia.com/browse/OMNIML-5151) — bug ticket - [OMNIML-5142](https://jirasw.nvidia.com/browse/OMNIML-5142) (install + register), [OMNIML-5144](https://jirasw.nvidia.com/browse/OMNIML-5144) (SPEC migration), [OMNIML-5123](https://jirasw.nvidia.com/browse/OMNIML-5123) (Epic) - nmm-sandbox MR !226 (CI install), !232 (PATH fix) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Improvements** * Enhanced launcher directory discovery mechanism with support for environment variable overrides and multiple search paths * Improved error diagnostics when launcher directory cannot be resolved * **Tests** * Added comprehensive test coverage for launcher directory discovery and failure scenarios <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
c4f39bd260 |
feat(tools/mcp): add dry_run flag to submit_job (#1718)
## Summary `submit_job(yaml_path, ..., dry_run=True)` validates a launcher YAML via `launch.py --dry-run` — exercises the YAML loader, factory resolution, and arg parser, with no cluster contact / no container spawn / no sbatch. Used by verify-task workflow stages (`pensieve-intern/workflows/spec_decoding_release/deployment_support.md`, `hidden_state_dump_support.md`, `training_support.md`, `synth_support.md`; `nemotron_quant_release/mlm_eval.md`, `mlm_ptq.md`, `mlm_qad.md`) that today shell out to `uv run launch.py --yaml X --dry-run` because the MCP catalog has no equivalent. ## Why Without this, [OMNIML-5144](https://jirasw.nvidia.com/browse/OMNIML-5144) (SPEC migration) can't finish — the verify-task SPECs all stay on shell prose for their YAML-validation step. Slice 1 of 5144 (specdec_bench's `cell.md`, pensieve-intern MR !100) didn't need `dry_run` because cell.md submits real cluster jobs, not validates configs. The other workflows do. ## Tool surface | Arg | Type | Notes | |---|---|---| | `dry_run` | `bool = False` | New. When True, skips cluster contact entirely. | | `hf_local` / `cluster_host` | `str?` | Now optional when `dry_run=True` (pass one to validate executor-specific config, omit both for shape-only validation). Still mutually exclusive in live submission. | | `skip_verify` | `bool = False` | Auto-bypassed when `dry_run=True` (verify_setup is meaningless without cluster contact). | Returns `{ok, dry_run: True, validated: bool, exit_code, stdout_tail, stderr_tail, argv}` instead of `experiment_id`. **Note on the ok/validated split:** when the launcher rejects the YAML (exit code non-zero), the tool still returns `ok: True` because the TOOL ran cleanly — only `validated: False`. Same shape `verify_setup` already uses: `ok: True` regardless of probe outcome; the probe outcome is the field the caller reads. ## Implementation Clean fork at the top of `submit_job_impl` — the dry_run branch lives in a dedicated `_submit_job_dry_run` helper to keep the live-submission path uncluttered. ~120 lines added, no changes to existing behavior. ## Tests 4 new hermetic tests in `tests/test_bridge.py` (total: 38 passing): * `test_submit_job_dry_run_yaml_validates` — happy path: ok+validated, `--dry-run` in argv, no `--yes` * `test_submit_job_dry_run_yaml_invalid` — launcher rejects YAML → `ok=True, validated=False`, `stderr_tail` surfaces the error * `test_submit_job_dry_run_yaml_not_found` — yaml_path missing → `yaml_not_found` carrying `dry_run: True` for caller introspection * `test_submit_job_dry_run_skips_verify` — bypasses `verify_setup` even with `skip_verify=False` ## Scope policy `dry_run` is universal launcher environment tooling, the [`tools/mcp/SCOPE.md`](https://github.com/NVIDIA/Model-Optimizer/blob/main/tools/mcp/SCOPE.md) test ("would this tool exist whether or not workflow X existed?") passes trivially — every workflow that validates a YAML wants this. No workload-specific knobs added. ## Anchors * [OMNIML-5145](https://jirasw.nvidia.com/browse/OMNIML-5145) — this ticket * [OMNIML-5144](https://jirasw.nvidia.com/browse/OMNIML-5144) — parent SPEC migration * [OMNIML-5123](https://jirasw.nvidia.com/browse/OMNIML-5123) — Epic * pensieve-intern MR !100 — slice 1 of the SPEC migration (the live-submit path) ## Validation * 38/38 unit tests pass (`uv run pytest tests/`) * Pre-commit clean (ruff, ruff-format, mypy, bandit, markdownlint, license-header) * No changes to the existing live-submission path — `dry_run=False` (default) is byte-for-byte equivalent to today's behavior <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a `dry_run` option to the `submit_job` tool to validate launcher YAML and referenced artifacts without contacting the cluster, spawning containers, or running `sbatch`. * Dry-run responses now report validation status and include details such as `exit_code`, stdout/stderr tails, and the launcher argv. * **Documentation** * Updated `submit_job` tool docs to describe the new `dry_run` behavior and the dry-run return shape. * **Tests** * Added unit tests for successful validation, validation failures, missing YAML, and ensuring setup verification is bypassed during dry-run. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
463229ad34 |
docs(tools/mcp): scope policy — environment tooling, not workflow policy (#1712)
## Summary
Adds `tools/mcp/SCOPE.md` documenting the design boundary for the MCP
server family (`modelopt-mcp` + `nmm-sandbox-mcp` +
`pensieve-intern-mcp`):
- **In scope:** universal verb-shaped operations on the cluster,
launcher, agent engine — environment tooling that every workflow can
rely on as pre-knowledge.
- **Out of scope:** workflow-specific logic ("run an EAGLE3 cell",
"publish a specdec release"). That belongs in SPEC text + agent
reasoning, composed out of these primitives.
## Why
Surfaced during the OMNIML-5123 follow-up discussion. The 14 tools
currently in scope across the three servers are deliberately a small,
closed set. Letting workflow-specific tools sneak in would sprawl the
catalog, force per-workflow opt-in, and break the "MCP-as-pre-knowledge"
promise that makes the catalog useful as a stable baseline for
pensieve-intern's agents.
SCOPE.md documents:
- The test: would this tool be useful across *any* workflow that uses
the same environment?
- A side-by-side table of in-scope vs out-of-scope tool shapes
- Symptoms that a tool is misclassified
- Why the line matters (interface-drift blast radius collapses, agent
tool catalog stays learnable)
- Four practical questions to ask before adding a tool
- A reference list of the 14 currently-in-scope tools
## Anchor
[OMNIML-5123](https://jirasw.nvidia.com/browse/OMNIML-5123) (Epic).
Docs-only, no code changes.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Added scope documentation clarifying that MCP environment tooling
covers universal verb-shaped operations (job submission, verification,
artifact reading) while workflow-specific policies belong elsewhere.
* Documented inclusion criteria and tool inventory guidance for MCP
servers.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
9f37fe1969 |
feat(tools/mcp): MCP server for ModelOpt launcher (OMNIML-5123) (#1701)
## Summary `tools/mcp/` — a new MCP server exposing the existing `tools/launcher/core.py` orchestration as **typed MCP tools** that codex / Claude Code agents can call directly, instead of shelling out to `uv run launch.py --yaml ...` and parsing prose output. Tracked under [OMNIML-5123](https://jirasw.nvidia.com/browse/OMNIML-5123) (Epic). Ships **Phase 1 + Phase 1.5** together: the core launcher surface plus the four highest-leverage helpers from the `cell.md` simplification loop ([OMNIML-5128](https://jirasw.nvidia.com/browse/OMNIML-5128) partial, [OMNIML-5132](https://jirasw.nvidia.com/browse/OMNIML-5132) full). ## Nine tools **Phase 1 — core launcher surface:** | Tool | Description | |---|---| | `list_examples` | Enumerate `tools/launcher/examples/` with model + description metadata extracted from each YAML | | `verify_setup` | Fail-fast probe for the named executor. Docker: `docker info` (daemon up) + `docker info --format` runtime-registry check for the `nvidia` runtime — no image pull, daemon-fast. Slurm: `ssh -o BatchMode=yes -o ConnectTimeout=5` to the cluster login node. ~1 s probe saves 30+ s of wasted submission on bad config | | `submit_job` | Submit a launcher YAML. Mode is determined by mutually-exclusive args: `hf_local` → Docker (local GPU), `cluster_host` → Slurm (remote SSH). Returns experiment_id immediately; the actual job runs detached | | `job_status` | Filesystem-based status from nemo_run's experiment dir (`_DONE`, `status_*.out`) — no in-memory registry, survives MCP server restarts | | `job_logs` | Read `log_<task>.out` from experiment dir; per-task filtering + optional tail | **Phase 1.5 — `cell.md` simplification (OMNIML-5128 / 5132):** | Tool | Description | |---|---| | `wait_for_experiment` | Replaces the agent's `while True: status; sleep` poll loop with one tool call. Reuses `job_status_impl` so terminal-state semantics stay identical. Returns final status plus `waited_seconds`; on timeout returns structured `{ok: False, reason: "wait_timeout", last_status: …}` | | `provision_passwordless_ssh_dry_run` | No-side-effect inspection of `~/.ssh/` that emits the exact `ssh-keygen` / `ssh-copy-id` commands the operator should run to make `verify_setup(executor='slurm')` pass. Closes the verify_setup "ssh_auth_failed → now what?" gap | | `read_cluster_artifact` | Uses nemo_run's tunnel primitives, not a reinvented SSH layer. `path=None` wraps `nemo experiment logs <id> <job_idx>` (built-in log fetch); with a `path`, uses the experiment's `Tunnel` to read the file. Structured failure on subprocess error / timeout | | `open_draft_pr` | `git push -u origin HEAD` + `gh pr create --draft …`. Validates cwd is a git repo first; on gh failure after push succeeds, reports `branch_pushed=True` so the operator can retry just the PR-open step | ## Design constants 1. **Single `submit_job` with mode by args** (not separate `submit_docker` / `submit_slurm` tools). Keeps the LLM tool catalog compact; mutual-exclusion is a runtime check. 2. **Filesystem is the source of truth** for status + logs. No in-memory registry. Survives MCP server restarts cleanly — important because operators / agents kill + restart their hosts often. 3. **`verify_setup` is auto-called by `submit_job`** by default (skippable when caller just probed). The probe is ~1 s; the cost of a misconfigured submission is 30+ s of cluster timeout or container-pull. Always-on verify pays back immediately. 4. **Delegate to nemo_run for tunnels.** `read_cluster_artifact` and `wait_for_experiment` use nemo_run's existing `Experiment` / `Tunnel` / `nemo experiment logs` primitives rather than reinventing SSH/rsync. One source of truth for cluster I/O. ## Layout ``` tools/mcp/ ├── pyproject.toml # name: modelopt-mcp, console_script ├── modelopt_mcp/ │ ├── __init__.py │ ├── server.py # FastMCP entry; 9 tool definitions │ └── bridge.py # thin wrapper over launcher's core.py │ # + filesystem status/log helpers │ # + tunnel/PR helpers (Phase 1.5) └── tests/ └── test_bridge.py # 34 unit tests, fully hermetic # (mocked subprocess + tmp_path fixtures) ``` ## Install Two paths, both **from source via uv**. No PyPI wheel exists; OMNIML-5123 opted for the uvx-from-git pattern to skip publication overhead. ### End-user install (recommended) `uvx` from the git subdirectory — single command, no manual clone: ```bash # Claude Code claude mcp add modelopt -- uvx --from \ "git+https://github.com/NVIDIA/Model-Optimizer.git#subdirectory=tools/mcp" \ modelopt-mcp # Codex codex mcp add modelopt -- uvx --from \ "git+https://github.com/NVIDIA/Model-Optimizer.git#subdirectory=tools/mcp" \ modelopt-mcp ``` Under the hood `uvx` clones the whole repo to its cache, installs `tools/mcp/` as the entry, and resolves the sibling `modelopt-launcher` dep via `[tool.uv.sources]` (`path = "../launcher"`) inside the cloned tree. ### Dev install (local checkout) ```bash uv pip install -e tools/launcher # sibling dep first uv pip install -e tools/mcp # then this package modelopt-mcp # entry on PATH ``` ### Why no plain `pip install` today Two specific reasons, worth flagging so reviewers know what's intentional vs missing: 1. **Nothing on PyPI yet.** Neither `modelopt-mcp` nor `modelopt-launcher` are published — this PR introduces the package but doesn't add release machinery. 2. **`pip` doesn't read `[tool.uv.sources]`.** Even from a local checkout, plain `pip install -e tools/mcp` fails because `modelopt-launcher` is a bare name (no URL) and pip can't find it. Sticking with `uv` / `uvx` is the practical path while we're git-only. If we later want plain-pip support: publish to PyPI, or switch to a PEP-440 direct URL (`"modelopt-launcher @ git+…#subdirectory=tools/launcher"`). Out of scope for this PR. ## Post-review changes Addressed all CodeRabbit + claude[bot] review findings on the original Phase-1 surface. See the inline replies for details; the substantive bug-fix highlights: * **Slurm `cluster_host`** — propagate via `env=child_env` (launch.py reads SLURM_HOST, not a CLI arg) * **`shlex.quote`** removed from nemo-run k=v overrides (subprocess list-form doesn't shell-quote) * **Docker `Popen`** now uses `stdout=DEVNULL, stderr=DEVNULL, start_new_session=True` to avoid pipe-buffer blocking * **`NEMORUN_HOME`** pinned in subprocess env so submit + status sides agree * **GPU verify** swapped from `docker run --gpus all` image-pull (slow + flaky) to `docker info --format` runtime-registry check (daemon-fast) * **Task-status word match** anchors on first word against a fixed failure-word set (no more `"fail" in "succeeded after retry; previous attempt failed"` false-positive) * **`experiment_id` regex** generalized for non-NVIDIA cluster paths * **`pyproject.toml`** dropped the unsatisfiable `modelopt-launcher` bare-name dep (launcher is a file-layout sibling, not a Python import dep) * **`Field(ge=1)`** on `job_logs.tail` * **Docstring contract** clarified (Docker returns `pid`, Slurm returns `experiment_id`) ## Validation - [x] `uv pip install -e .` succeeds (modelopt-launcher resolved transitively) - [x] 34/34 unit tests pass (`uv run python -m pytest tests/`) - [x] stdio handshake works end-to-end; `tools/list` returns all 9 with full schemas + descriptions - [x] Mode-resolution: `submit_job` correctly rejects no-executor + both-executors with structured `reason` - [x] Filesystem status: correctly classifies `done` / `failed` / `running` from `_DONE` + `status_*.out` - [x] `wait_for_experiment` short-circuits on already-terminal experiments; honors timeout without raising - [x] `provision_passwordless_ssh_dry_run` distinguishes no-key / key-only / key+pubkey cases - [x] `read_cluster_artifact` handles subprocess timeout + non-zero exit with structured reasons - [x] `open_draft_pr` reports `branch_pushed=True` on gh-failure-after-push so retries are cheap - [x] Pre-commit clean: ruff, ruff-format, mypy, bandit, license-headers ## Acceptance criteria **OMNIML-5123 (Phase 1):** - [x] `list_examples` returns all bundled YAMLs with path and model name - [x] `submit_job` with `hf_local` runs via Docker executor and returns immediately (Phase 1: returns PID; experiment_id capture in Phase 2) - [x] `submit_job` with `cluster_host`/`user` runs via Slurm executor (`detach=True`) and returns experiment_id - [x] `job_status` correctly reflects running / done / failed from nemo_run filesystem - [x] `job_logs` returns stdout for a completed job - [x] `uvx --from git+...#subdirectory=tools/mcp modelopt-mcp --help` resolves and starts - [x] Existing launcher tests unaffected (no changes to `tools/launcher/`) **OMNIML-5128 (Phase 1.5, partial):** - [x] `wait_for_experiment` blocks until terminal or timeout - [x] `read_cluster_artifact` pulls remote artifacts via nemo_run tunnel - [x] `open_draft_pr` opens a draft PR against a target repo - [ ] Capture `experiment_id` from Docker subprocess output — deferred to Phase 2 **OMNIML-5132 (Phase 1.5, full):** - [x] `provision_passwordless_ssh_dry_run` emits operator-facing commands without side effects ## Phase 2 (separate PR) * Capture `experiment_id` from Docker subprocess output (tail until nemo_run logs the id). * Extract the verify + submit helpers into a shared lib that [`nmm-sandbox-mcp`](https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/tree/main/tools/mcp) (companion server, separate repo) can consume for internal-ergonomics tools — cluster short-name → factory lookup + GitLab CI dispatch. * NEL integration ([OMNIML-5133](https://jirasw.nvidia.com/browse/OMNIML-5133)) + checkpoint introspection ([OMNIML-5134](https://jirasw.nvidia.com/browse/OMNIML-5134)). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added ModelOpt MCP server and console entrypoint; tools: list_examples, verify_setup, submit_job, job_status, job_logs, wait_for_experiment, provision_passwordless_ssh_dry_run, read_cluster_artifact, open_draft_pr; Docker and Slurm support. * **Documentation** * Expanded README with install steps, design notes, end-to-end agent example, roadmap, and repo layout. * **Tests** * Expanded unit tests covering bridge helpers, polling, SSH flows, artifact reads, and PR automation. * **Chores** * CI updated to run MCP tests; package/meta config added. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
cfc823d127 |
[Tests]: Precommit Check for Spec-Dec Recipes (#1527)
### What does this PR do? Type of change: new tests / tooling Adds pre-commit validation for speculative-decoding recipes (the existing `check-modelopt-recipes` hook only ran on PTQ) and for launcher YAML references into the recipe library. - `tools/precommit/check_modelopt_recipes.py`: accept `speculative_eagle` / `speculative_dflash` / `speculative_medusa` in addition to `ptq`, so per-model spec-dec recipes (e.g. `modelopt_recipes/models/Qwen3-8B/dflash.yaml`) get full Pydantic validation via `load_recipe()` at commit time. - `tools/precommit/check_launcher_yaml.py` (new): scans every `tools/launcher/examples/**/*.yaml` for `--config <path>` and `data.chat_template=<path>` references, verifies the resolved files exist, and runs `load_recipe()` on any path under `modelopt_recipes/`. Skips `<<global_vars.x>>` interpolation. `pass_filenames: false` so recipe-side edits also re-validate all launcher references. ### Usage ```bash pre-commit run check-modelopt-recipes --all-files pre-commit run check-launcher-yaml --all-files ``` ### Testing Smoke-tested both hooks manually: | Scenario | Result | |---|---| | spec-dec recipe with `dflash_block_size: not_an_int` | exit 1, Pydantic int_parsing error | | launcher YAML with non-existent `--config` path | exit 1, source file + resolved path reported | | launcher YAML with non-existent `data.chat_template` path | exit 1 | | Current repo state (all valid) | exit 0 | ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ (hooks themselves are the tests) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ pending ### Additional Information Motivated by the per-model recipe migration in #TBD — without these hooks, broken `--config` paths and recipe schema typos surface only at CI or runtime. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Added an automated pre-commit check that validates launcher example YAMLs, reporting parse errors and missing or invalid references. * Expanded recipe validation to cover additional recipe types beyond PTQ, improving detection of invalid recipe formats and metadata. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
cb5be1f92b |
launcher: move gemma-4-E4B-it yaml to google/ subdir (#1689)
## Summary - Moves `tools/launcher/examples/gemma-4/gemma-4-E4B-it/specdec_bench_mtp_vllm.yaml` to `tools/launcher/examples/google/gemma-4-E4B-it/specdec_bench_mtp_vllm.yaml` to align with the `google/` vendor-prefix convention used by other models in the launcher examples directory. - Updates the embedded `--yaml` path in the YAML comment to match the new location. ## Test plan - [ ] Confirm the file appears at the new path - [ ] Confirm the old path is gone <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated an example benchmark invocation reference to reflect the project’s reorganized directory layout. This edits only the example comment in the benchmark YAML; no benchmark settings, model paths, task arguments, or runtime configuration were modified. Low-risk documentation alignment. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan D. Yu <5185878+ChenhanYu@users.noreply.github.com> |
||
|
|
46eddab877 |
[Feat]: Specdec Streaming: RDMA + Multinode (#1611)
### What does this PR do? Type of change: New feature Multi-node **streaming** training for speculative decoding (EAGLE3 / DFlash): a live `vllm serve` captures the target model's hidden states and moves them straight to the trainer over **NIXL RDMA** — no disk round-trip. The streaming dataset is map-style — each rank fetches only its own `DistributedSampler` shard (concurrency from `dataloader_num_workers`), round-robins across multiple serve replicas (`server_urls`), and scales to multi-node DDP. Serve-side tensor parallelism (TP>1) is supported: hidden states are replicated across TP ranks, so rank 0 alone owns the pool + transfer. ### How - `RdmaHiddenStatesConnector` — out-of-tree vLLM connector (no vLLM source edits): one pre-registered pinned NIXL pool per serve, a ring slot per request, and a small HTTP sidecar serving transfer metadata. The trainer RDMA-READs the slot into a per-worker buffer. RDMA is the **only** transport (the earlier disk/safetensors path is removed). - Map-style dataset + multi-node accelerate launch (`--machine_rank`, optional Slurm `--segment` to keep nodes in one NVLink domain). ### Usage ```yaml data: mode: streaming streaming_server_url: "http://node0:8000,http://node1:8000" # round-robin ``` ### Validation (Qwen3-8B, oci-nrt H100) sandbox CI: https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/337489812 **1. End-to-end convergence — EAGLE3 & DFlash, 5000 steps.** Both algorithms converge and export a deployable draft; the DFlash drafts also serve under vLLM speculative decoding (8/8 smoke prompts pass). | algorithm | topology (nodes) | train loss (step 0 → 5000) | vLLM draft acc-len | |---|---|---|---| | EAGLE3 | 2 serve TP=2 + 2 trainer DDP (4) | 37.1 → 8.20 | — | | DFlash | 1 serve TP=1 + 1 trainer (2) | 11.7 → 5.56 | 1.11 | | DFlash | 2 serve TP=2 + 2 trainer DDP (4) | 10.9 → 5.26 | 1.19 | <!-- Drag these PNGs in here (GitHub turns them into asset URLs): eagle3_streaming_loss.png, dflash_streaming_loss_singlenode.png, dflash_streaming_loss_multinode.png --> **2. Scalability — 1 → 12 nodes (EAGLE3, 200 steps).** Throughput scales ~23× across the sweep below. The step-time growth is cross-node DDP all-reduce, not the streaming path — RDMA (~0.33 ms/req @ 2 MB, ~47 GB/s host-pinned READ) is never the bottleneck. Scale serve + trainer nodes together for near-linear speedup. | serve / trainer | nodes | step time | samples / step | samples / sec (global) | acc @ step 200 | |---|---|---|---|---|---| | 1 serve / 1 rank (co-located, 1 node 2 GPU) | 1 | 0.23 s | 1 | 4.4 | [0.141, 0.094, 0.072] | | 1 serve / 1 rank (cross-node) | 2 | 0.23 s | 1 | 4.3 | [0.137, 0.105, 0.074] | | 2 serve / 8 ranks | 3 | 0.26 s | 8 | 31.1 | [0.215, 0.126, 0.097] | | 4 serve / 16 ranks (2 trainer nodes) | 6 | 0.28 s | 16 | 56.5 | [0.217, 0.148, 0.110] | | 8 serve / 32 ranks (4 trainer nodes) | 12 | 0.31 s | 32 | 101.9 | [0.235, 0.165, 0.137] | **3. Serve-side TP correctness.** TP=1 vs TP=2 draft top-1 accuracy track step-for-step (hidden states are replicated across TP ranks). <img width="910" height="546" alt="serve-tp-acc" src="https://github.com/user-attachments/assets/73df9214-7ff0-4ab4-bf2f-95842b12cd5f" /> ### Before your PR is "*Ready for review*" - Backward compatible?: ❌ — streaming is now RDMA-only; `server_url` → `server_urls`; the disk transport (`HS_TRANSPORT`, `streaming_shared_storage_path`) is removed. - New tests?: ✅ `tests/unit/torch/speculative/plugins/test_hf_streaming_dataset.py` (map-style dataset + mocked RDMA fetch). - Updated Changelog?: ❌ - Claude approval?: ❌ --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
13e540f0b3 |
launcher: add draft_model to GlobalVariables for spec-decode parents (#1675)
SPEED-bench MTP / EAGLE3 / DRAFT_TARGET / DFLASH parent YAMLs reference the speculative draft/assistant model path on BOTH the qualitative and throughput_32k tasks via `<<global_vars.draft_model>>` indirection so the path lives in one place. The launcher's `GlobalVariables` dataclass had a strict whitelist of keys and rejected `draft_model` with `ValueError: No parameter named 'draft_model' exists`. Surfaced on OMNIML-5024 pipeline #54356795: the gemma-4-E4B-it / MTP / vLLM parent used the indirection but failed at launcher load. The agent worked around it inline. This adds `draft_model: str = None` as the fifth top-level field on `GlobalVariables`. The existing `<<global_vars.X>>` resolution in `SandboxPipeline.__post_init__` (uses `dataclasses.asdict`) picks the new field up automatically. Backward-compatible: existing parents that don't use `draft_model` keep working. Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
808e8b3320 |
Make .claude/skills a real folder for claude sandbox to work (#1674)
### What does this PR do? Type of change: ? Bug fix replace .claude/skills dir-symlink with real dir + per-skill symlinks Claude Code hardcodes .claude/skills in its sandbox denyWithinAllow list. bwrap enforces this by creating a mount point at that path, but fails with "Can't create file: Is a directory" when the path is a symlink to a directory — breaking all sandboxed commands, not just writes. Fix by making .claude/skills a real directory containing per-skill symlinks into .agents/skills/. Add a pre-commit hook that automatically creates and stages a new symlink whenever a skill directory is added to .agents/skills/, so authors need no extra steps. ### Usage claude with sandbox ### Testing tried on local. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated internal development tooling to streamline workflow automation and improve developer processes. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Shiyang Chen <shiychen@nvidia.com> |
||
|
|
06af68cbdb |
[OMNIML-4962] specdec_bench parent — Qwen/Qwen3.5-4B / DFlash / vLLM (#1638)
## Summary\n- Add specdec_bench DFlash vLLM cell t0_d3 for Qwen3.5-4B\n\n## Testing\n- Not run (cell YAML only)\n <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Added new benchmark configuration for Qwen3.5-4B model performance evaluation with speculative decoding optimization across qualitative and throughput testing scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: chenhany <chenhany@nvidia.com> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
66b54ed040 |
[OMNIML-5024] specdec_bench cell t0_d3 — google/gemma-4-E4B-it / MTP / vllm (#1663)
### What does this PR do?
Type of change: Bug fix + new example
Wires SPEED-bench's MTP path to support **Gemma 4** (and any future MTP
variant that uses a separate assistant / draft model), and adds the
SPEED-bench MTP/vLLM example for `google/gemma-4-E4B-it`.
**Key difference: Gemma 4 MTP vs. generic MTP.** vLLM's
`speculative_config` accepts two different shapes for MTP:
| Variant | `speculative_config` shape | Models |
|---|---|---|
| **Generic MTP** | `{"method": "mtp", "num_speculative_tokens": N}` |
Models that carry their own MTP layer in-tree (e.g. Qwen 3.5 MTP
variants) — no separate draft / assistant model. |
| **Assistant-model MTP** | `{"model": "<assistant>",
"num_speculative_tokens": N}` (no `method` key — vLLM auto-detects from
the assistant) | Gemma 4 family (E2B / E4B / 26B-A4B / 31B); each target
model has a paired `<target>-assistant` checkpoint that acts as the MTP
draft. Landed in
[vllm-project/vllm#41745](https://github.com/vllm-project/vllm/pull/41745)
(2026-05-06). |
The specdec_bench vLLM wrapper at
`examples/specdec_bench/specdec_bench/models/vllm.py` previously emitted
only the generic shape for any `--speculative_algorithm MTP` invocation,
which produced `NotImplementedError: Unsupported speculative method:
'mtp'` on Gemma 4 even with a container that has the support
(`vllm/vllm-openai:v0.22.1`+). This PR teaches the wrapper to switch
shapes based on whether `--draft_model_dir` is provided.
**Concrete changes:**
1. **`examples/specdec_bench/specdec_bench/models/vllm.py`** — when
`speculative_algorithm == "MTP"` AND `draft_model_dir` is set, emit
`{"model": draft_model_dir, "num_speculative_tokens": N}`
(assistant-model shape). Otherwise emit the existing `{"method": "mtp",
...}` (generic shape). Backward-compatible — Qwen 3.5 MTP and other
callers that omit `--draft_model_dir` get the same config they got
before.
2. **`examples/specdec_bench/specdec_bench/utils.py`** — `get_tokenizer`
reads `extra_special_tokens` from the model's `tokenizer_config.json`
and passes them through to `AutoTokenizer.from_pretrained`. Gemma 4
tokenizers ship a list-shaped `extra_special_tokens` entry that the
constructor would otherwise reject. Necessary for any Gemma 4 cell.
3.
**`tools/launcher/examples/gemma-4/gemma-4-E4B-it/specdec_bench_mtp_vllm.yaml`**
— SPEED-bench parent YAML for `google/gemma-4-E4B-it`. Uses
`vllm/vllm-openai:v0.22.1` (has `gemma4_mtp.py` from #41745) and wires
`--draft_model_dir /hf-local/google/gemma-4-E4B-it-assistant` on both
task_0 (qualitative) and task_1 (throughput_32k).
4.
**`tools/launcher/common/specdec_bench/_cells/gemma-4-E4B-it_mtp_vllm_t0_d3.yaml`**
— runtime params for the `t0_d3` cell of OMNIML-5022 (`temperature=0`,
`max_model_len=40960`).
### Usage
```python
# Wrapper-level: same CLI as before, just pass --draft_model_dir for
# Gemma 4 MTP. The wrapper auto-routes to the assistant-model shape.
# python examples/specdec_bench/run.py \
# --engine VLLM \
# --speculative_algorithm MTP \
# --draft_model_dir /hf-local/google/gemma-4-E4B-it-assistant \
# --draft_length 3 \
# --tp_size 1 \
# ...other SPEED-bench knobs...
# Equivalent direct vLLM invocation (for reference, no wrapper):
from vllm import LLM, SamplingParams
llm = LLM(
model="google/gemma-4-E4B-it",
speculative_config={
"model": "google/gemma-4-E4B-it-assistant",
"num_speculative_tokens": 3,
},
trust_remote_code=True,
)
```
### Testing
- **Upstream existence checks**: verified the assistant models
`google/gemma-4-{E2B,E4B,26B-A4B}-it-assistant` exist, public, ungated
on HuggingFace; verified `vllm/model_executor/models/gemma4_mtp.py` is
in vLLM `v0.22.0`, `v0.22.1`, and `main`.
- **Backward compat**: `MTP` callers that don't pass `--draft_model_dir`
(e.g. the existing Qwen 3.5 MTP/vLLM cells under
`tools/launcher/examples/Qwen/Qwen3.5-4B/`) take the unchanged
`{"method": "mtp", ...}` branch. No diff for those.
- **End-to-end cluster validation**: pending. Will run via the
OMNIML-5022 cells (OMNIML-5024 / 5025 / 5026 / 5027) once the
nmm-sandbox submodule pin advances past this PR. Each cell exercises
`task_0` (SPEED-Bench qualitative, 880 samples) + `task_1`
(throughput_32k, 80 samples) on cw_dfw, single H100.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ — the wrapper only takes the
new branch when `--draft_model_dir` is provided alongside
`--speculative_algorithm MTP`. Existing MTP callers (Qwen 3.5 etc.) keep
the generic `method: "mtp"` config.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependencies.
- Did you write any new necessary tests?: ❌ — relying on the SPEED-bench
cluster cells (OMNIML-5024 …5027) for end-to-end validation; no unit
test fixture for the vLLM wrapper exists in `tests/` for me to extend
symmetrically. Happy to add one if reviewers want it.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — small fix + example addition. Can add if requested.
- Did you get Claude approval on this PR?: ❌ — will run `/claude review`
once the PR is marked Ready for review.
### Additional Information
- JIRA: [OMNIML-5024](https://jirasw.nvidia.com/browse/OMNIML-5024)
(cell_t0_d3); siblings OMNIML-5025/5026/5027 (cell_{t0_d7, t1_d3,
t1_d7}) of Epic OMNIML-5022 are blocked on this PR landing.
- Upstream reference: vllm-project/vllm#41745 — "[Spec Decode] Add
Gemma4 MTP speculative decoding support".
- Companion (pensieve-intern !91, internal): adds a
"Model-family-specific MTP invocation" table to the specdec_bench cell
SPEC so future agents pair `MTP` with the right `--draft_model_dir` from
SPEC-read time.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added a SPEED-bench pipeline for Gemma 4 using vLLM speculative
decoding (MTP) with qualitative and throughput tasks.
* **Improvements**
* Speculative-decoding logic updated to handle assistant-model and
generic MTP cases distinctly.
* Tokenizer loading now reads and normalizes extra special tokens from
tokenizer config when available.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Pensieve Intern <chenhany@nvidia.com>
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
|
||
|
|
111b7ebd29 |
[OMNIML-4969] specdec_bench cell t0_d3 — nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 / dflash / vllm (#1656)
## Summary\n- Add specdec_bench cell t0_d3 for NVIDIA-Nemotron-3-Super-120B-A12B-BF16 dflash/vLLM\n\n## Testing\n- Not run (cell YAML only)\n <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a new benchmark configuration to run Nemotron-3-Super-120B on NVIDIA with vLLM using the DFlash speculative-decoding method. * Provides both qualitative and high-throughput benchmarking scenarios, tunable concurrency and request counts, and automated job execution with Slurm-compatible orchestration and output saving. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: chenhany <chenhany@nvidia.com> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Signed-off-by: Chenhan D. Yu <5185878+ChenhanYu@users.noreply.github.com> |
||
|
|
2c52e7bf4e |
[OMNIML-4788] specdec_bench/Qwen3.5-4B: throughput_32k benchmark + S3 upload step (#1564)
### What does this PR do? Type of change: enhancement (follow-up to [#1531](https://github.com/NVIDIA/Model-Optimizer/pull/1531)). Extends the merged Qwen3.5-4B SPEED-Bench launcher YAMLs from a single-task qualitative-only smoke into a **3-task pipeline** that also covers long-context throughput and verifies the S3-upload path end-to-end. Two commits, cleanly cherry-picked from #1531's late branch state — they were authored after the merge-commit was resolved against an earlier rebased head and so didn't ride along with that merge. ### Pipeline shape (both YAMLs) | Task | Split | Save dir | |---|---|---| | `task_0` | qualitative (existing quality / acceptance-rate signal) | `/scratchspace/specdec_bench{,_mtp}/qualitative` | | `task_1` | **throughput_32k** (new — long-context throughput) | `/scratchspace/specdec_bench{,_mtp}/throughput_32k` | | `task_2` | **upload to S3 in sweep layout** | `s3://team-specdec-workgroup/results/specdec_bench{,_mtp}/<split>/` | ### New artifacts * `tools/launcher/common/specdec_bench/upload_to_s3.sh` — thin wrapper around `examples/specdec_bench/upload_to_s3.py` so it can be invoked as a launcher task. Installs `boto3` from `requirements.txt` on cold containers; warm pipelines pick it up from the prior `run.sh`. * `tools/launcher/common/specdec_bench/runtime_params_throughput_32k.yaml` — pins `engine_args.max_model_len = 40,960` (32K input + 4K output + 4K headroom) so vLLM doesn't silently auto-cap `max_model_len` below the 36K minimum needed for `throughput_32k` prompts on single-GPU runs. ### Why max_model_len matters Without an explicit `max_model_len`, vLLM auto-derives it from the model config (Qwen3.5-4B = 128K) **and from the GPU-memory budget**. On a single GPU the second factor can cap effective `max_model_len` well below 36K, silently truncating 32K-token prompts and producing wrong throughput numbers. The qualitative split is not affected (its prompts top out around 8K, well under any auto-derivation floor) so only `task_1` carries the override. ### S3 credentials `upload_to_s3.sh` reads `S3_ENDPOINT` / `S3_KEY_ID` / `S3_SECRET` from the runtime environment (not hardcoded). `--skip-existing` + `--allow-incomplete-provenance` are passed by default so re-runs land alongside the prior upload, and runs lacking `CONTAINER_IMAGE` (Phase-2 harness work in OMNIML-4788 will populate it) still upload. ### Testing Cluster smoke on cw_dfw via: ``` uv run slurm.py --yaml modules/Model-Optimizer/tools/launcher/examples/Qwen/Qwen3.5-4B/specdec_bench.yaml --yes ``` is currently in-flight (jobs `12257378/79/80`, PD). Will update this PR with timing/AR numbers + S3 upload confirmation once it lands. ### Before your PR is "Ready for review" - Backward compatible: ✅ (additive — task_0 keeps the prior qualitative behavior, just with `/qualitative` suffix in `save_dir`) - New PIP dep: ✅ no (boto3 already in `examples/specdec_bench/requirements.txt` from #1531) - New tests: N/A (launcher YAML + shell wrapper; covered by cluster smoke) - Changelog: N/A (internal-facing tooling) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a 32K-context runtime configuration (higher max model length) to enable long-context throughput benchmarking and avoid silent prompt truncation. * Added a launcher helper to upload benchmark results to S3 with incremental/retry-friendly options and pass/fail reporting. * **Chores** * Split Qwen3.5-4B benchmark into separate qualitative and 32K throughput tasks and added coordinated S3 upload. * Applied the same multi-task pipeline layout and clearer output organization to the MTP speculative-decoding benchmark. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1564?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: chenhany <chenhany@nvidia.com> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
6b73e933bf |
Nemotron Ultra & Super launcher examples (#1609)
### What does this PR do? Type of change: New example New launcher example for Nemotron Super with PTQ + Export + VLLM smoke test on small GPQA-style dataset ### Usage ```python # Usage: # source .env-slurm # cd tools/launcher # uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/megatron_lm_ptq.yaml --yes ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added checkpoint export capability for quantized models to Hugging Face format. * Introduced complete quantization pipelines with conditional MMLU evaluation and model export stages. * **Bug Fixes** * Fixed num_shards calculation to prevent invalid minimum values. * **Documentation** * Updated vLLM version requirements for optimal NVFP4 model performance. * Enhanced quantization pipeline documentation with improved output paths and conditional execution details. * **Chores** * Updated Megatron-LM module to latest version. * Added sample dataset for model evaluation testing. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> |
||
|
|
5bd04c3876 |
Revert unverified EAGLE3 model examples; keep triage code + baseline (#1623)
### What does this PR do? Type of change: Revert / cleanup (follow-up to #1417) Per review feedback (@h-guo18): `main` should be production-ready and user-facing. Most of the EAGLE3 model example YAMLs added in #1417 are not yet verified to work end-to-end in modelopt (~80% fail at some pipeline stage), which is confusing to ship. The agreed plan is to **land the triage infrastructure now and re-add each model's launcher YAML in a dedicated follow-up PR once it is verified green**. **Removed** (unverified, to be re-added per-model once verified): - Per-model launcher configs (`hf_offline_eagle3.yaml` + `eagle3_quick_check.yaml`) for: DeepSeek-V3.2, GLM-5, MiniMax-M2.5, Ministral-3-8B, Ministral-3-14B, Kimi-K2.5, Kimi-K2.5-NVFP4, GPT-OSS-20B, Qwen3.5-9B, Qwen3.5-27B, Qwen3.5-35B-A3B, Step-3.5-Flash. - Per-model status docs: `tools/launcher/examples/EAGLE3_TRIAGE.md`, `examples/speculative_decoding/pipeline/eagle3/eagle3_triage_chart.md` (volatile status — tracked internally instead). **Kept** (the durable triage infrastructure from #1417): - Verified baseline example `tools/launcher/examples/Qwen/Qwen3-8B/eagle3_quick_check.yaml`. - Launcher common scripts (vLLM native-extractor dump, etc.) and `compute_hidden_states_vllm.py`. - modelopt code fixes: FakeBaseModel VLM detection, `consolidated.safetensors` load, `use_cache` export templates. - New-model triage guide (`eagle3_new_model_triage_guide.md`); its "document results" step now points at the internal tracker rather than the removed chart. ### Testing No code paths change — this only removes example YAMLs and two status docs and edits one doc reference. Pre-commit (ruff/markdownlint/yaml/license) passes on the kept/edited files. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (removes unverified examples only; kept infra unchanged) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: ❌ (pending) ### Additional Information Follow-up to #1417. Next step (tracked separately): verify each removed model end-to-end in modelopt, then re-add its YAML in a dedicated PR. Note: the nmm-sandbox weekly EAGLE3 CI is being trimmed to the Qwen3-8B baseline to match. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated EAGLE3 triage guide to streamline the verification workflow; contributors now record test outcomes (status, experiment IDs, errors, and applied fixes) in the team's internal triage tracker before submitting model launcher configurations. * **Chores** * Removed legacy EAGLE3 example pipeline configurations and deprecated triage documentation to reduce maintenance overhead. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
a7b0a92047 |
EAGLE3 new model support: pipeline configs, triage docs, and Ministral-3 fixes (#1417)
## Summary EAGLE3 automation triage work (OKR-30): testing the 4-step EAGLE3 offline pipeline against 12 new model architectures, documenting failure modes, and fixing issues found. ### Code fixes (modelopt) | File | Change | |------|--------| | `modelopt/torch/speculative/utils.py` | Extend VLM detection in `load_vlm_or_llm` to check `text_config`/`llm_config` attrs (catches `mistral3` models) | | `modelopt/torch/speculative/plugins/modeling_fakebase.py` | Add `consolidated.safetensors` fallback for checkpoints with incomplete HF shards | | `modelopt/torch/export/plugins/hf_spec_configs.py` | Set `use_cache=True` in EAGLE export templates (fixes strict `huggingface_hub` validation) | ### Pipeline infrastructure - `examples/speculative_decoding/pipeline/eagle3/` — pipeline scripts and configs: - `offline_training.sh` — training + export with runtime patches for older container modelopt - `dump_offline_data_vllm.sh` — vLLM-based hidden state extraction (with speculators compat patches) - `dump_offline_data.sh`, `dump_offline_data_hf.sh` — alternative dump paths - 18 quick-fail-check YAMLs for 12 models - 4 standalone task1 YAMLs ### Documentation - `eagle3_triage_chart.md` — model test matrix, triage decision tree, per-model results, failure catalog - `eagle3_new_model_triage_guide.md` — step-by-step guide for triaging new models ### Model test results (as of 2026-05-27) | Model | task_0 | task_1 | task_2 | task_3 | Blocker | |-------|--------|--------|--------|--------|---------| | Qwen3-8B | - | - | - | - | Reference (existing) | | Kimi-K2.5 | - | - | - | - | Existing (GB200) | | **Ministral-3-8B** | SKIP | PASS | PASS | FAIL | `use_cache=null` in export (fixed) | | Ministral-3-14B | FAIL | - | - | FAIL | vLLM engine init fails | | Qwen3.5-35B-A3B | TIMEOUT | - | - | - | Data synth too slow | | gpt-oss-20b | FAIL | - | - | - | Tokenizer `HarmonyError` | | Step-3.5-Flash | TIMEOUT | - | - | - | Data synth time limit | | MiniMax-M2.5 | TIMEOUT | - | - | - | `trust_remote_code` needed | | DeepSeek-V3.2 | no log | - | - | - | May not be mirrored | | Qwen3.5-9B | - | - | - | - | Not yet run | | Qwen3.5-27B | - | - | - | - | Not yet run | | GLM-5 | - | - | - | - | Not yet run | ### Issues found and fixed | # | Issue | Fix | |---|-------|-----| | 1 | `mistral3` model type not detected as VLM | Check `text_config`/`llm_config` attrs in `load_vlm_or_llm` | | 2 | Missing HF shard file (Ministral-3-8B) | Fallback to `consolidated.safetensors` with Mistral native key aliases | | 3 | `use_cache=null` in exported EAGLE config | Set `use_cache=True` in export template configs | | 4 | speculators incompatible with vLLM container | Runtime patches in `dump_offline_data_vllm.sh` | | 5 | `offline_training.sh` infra issues | Rewritten with runtime patches for container modelopt | ## Test plan - [x] Ministral-3-8B training passes (`cicd_1779829129`) - [x] Ministral-3-8B export succeeds - [ ] Ministral-3-8B benchmark passes (`cicd_1779901409` — pending with all fixes) - [ ] Dry-run remaining model configs ## Note GitHub secret scanning alert #6 is a **false positive** — `Mistral3ForConditionalGeneration` (a HuggingFace model class name in a YAML comment) was flagged as a "Mistral AI API Key". 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
902d36921a |
[Feat]: Streaming Hidden-states Dataset (#1509)
### What does this PR do? Type of change: new feature **Design doc:** https://gist.github.com/h-guo18/241c94968b0591324c361d97cf995dd0 Jira ticket: https://jirasw.nvidia.com/browse/OMNIML-4341 Sandbox CI: https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/327911555#L2228 Streaming hidden-states dataset: per-sample activations pulled from a live `vllm serve` over HTTP, replacing on-disk activation dumps. Two axes for future extensions: - **Backend** (`_fetch`): vLLM now; TRT-LLM / SGLang next. - **Algorithm** (`_format`): Eagle now; distillation / probing next. - **Sandbox CI**: https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/merge_requests/169 Shared plumbing — async producer, token-level truncation to `training_seq_len`, loss-mask alignment, DDP via Accelerate's dispatcher (rank 0 fetches, broadcasts), circuit breaker, resume — lives in `StreamingDataset`. First instance: **`EagleVllmStreamingDataset`**. API: `data.mode ∈ {online, offline, streaming}`; legacy configs auto-promote. ### Usage ```yaml data: mode: streaming data_path: input_conversations/train.jsonl streaming_server_url: http://localhost:8000 streaming_model_name: meta-llama/Llama-3.1-8B-Instruct training: training_seq_len: 4096 # also caps the prompt sent to vllm ``` Requires `vllm serve` with `ExampleHiddenStatesConnector` and `dataloader_num_workers=0`. End-to-end Slurm pipeline: `tools/launcher/examples/Qwen/Qwen3-8B/hf_streaming_eagle3.yaml`. ### Testing - **Unit**: full-corpus invariant, rank-0-only iter, resume, circuit breaker, mocked-httpx integration. - **E2E**: `launch_train.sh` against a stdlib `HTTPServer` mimicking the connector. - **Smoke** (Qwen3-8B / 8×H100 / 4096 ultrachat samples, single epoch): train_loss 32 → 18, MT-Bench AR 1.003 → 1.20. ### TODO before un-drafting - [ ] Observability counters (filtered, fetch failures, queue depth, latency). - [ ] Changelog entry. - [x] Add test in sandbox. **Non-goals (v1):** multi-epoch streaming, cross-rank dynamic load balancing. ### Before your PR is "*Ready for review*" - Backward compatible: ✅ (legacy configs auto-promote) - New PIP dep: N/A - New tests: ✅ - Changelog: ❌ (TODO above) - Claude approval: ❌ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Streaming training mode with server-backed hidden-state fetching, deterministic seed control, resume support, and streaming-specific dataset options (server, model, prefetch, shared storage). * **Behavior / Bug Fixes** * Stronger mode validation; offline behavior derived from data mode; resume handling adjusted to avoid double-skip during streaming runs. * **Tests** * End-to-end CI streaming test and expanded unit tests covering streaming, resume, DDP, determinism, and failure cases. * **Infrastructure** * Launcher script and pipeline config for end-to-end streaming training. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1509?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
905259fbf5 |
Fix: use python3 in debugger server.sh (#1590)
### What does this PR do? Type of change: Bug fix `tools/debugger/server.sh` invokes `python -c "..."` inside `check_modelopt_local` (and `pip install -e .[dev]` in the fallback path). Containers that only ship `python3` (no `python` shim) cause the check to fail with `python: command not found`, exit 127. The server interprets this as "modelopt not editable-installed", runs `pip install`, the second check fails the same way, and the server aborts before it ever listens for commands. Switches both the inline import check and the install fallback to `python3` / `python3 -m pip`. ### Usage ```bash bash tools/debugger/server.sh # now starts cleanly on python3-only images ``` ### Testing Verified inside a container where `which python` returns nothing and `python3` is `/usr/bin/python3` (Python 3.12). With the old script, `server.sh` aborted at `check_modelopt_local`. With this fix, the check passes against an existing editable install and the server proceeds to wait for the client handshake. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `python3` is present on every image that previously had `python`. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — internal tooling shell script. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — internal debugger tooling, not user-facing. - Did you get Claude approval on this PR?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated debugger tooling to ensure consistent use of Python 3 for package validation and installation processes. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
54fb87e972 |
[OMNIML-4869] author_yaml (#1574)
Draft PR opened by **pensieve-intern** for [OMNIML-4869](https://jirasw.nvidia.com/browse/OMNIML-4869). Stage `author_yaml` of Epic `OMNIML-4868`. The agent ran from the SPEC on the ticket description; review every change before marking ready. _Always-draft is enforced — the bot never auto-merges._ --- **Agent's self-narration** (stripped from PR diff; surfaced here for context): `VERIFICATION_COMMENT.txt`: ``` OMNIML-4869 status: model not staged on cw_dfw. Spec checks: - HF Hub model exists: Qwen/Qwen3.5-4B (public, not gated; model_type=qwen3_5). - Cluster stage check failed: /lustre/fsw/portfolios/coreai/projects/coreai_dlalgo_modelopt/hf-local/Qwen/Qwen3.5-4B not found. Requested action on Epic: please stage the model via `/manage-assets cw_dfw get Qwen/Qwen3.5-4B`. Notes: - YAML already exists at tools/launcher/examples/Qwen/Qwen3.5-4B/specdec_bench_mtp_vllm.yaml with tp_size=1, gpus_per_node=1, and container vllm/vllm-openai:qwen3_5-cu130. - Once staged, I can run the required dry-run: `uv run launch.py --yaml tools/launcher/examples/Qwen/Qwen3.5-4B/specdec_bench_mtp_vllm.yaml --dryrun --yes -v`. ``` _Pollution-strip removed `VERIFICATION_COMMENT.txt` from this commit (sidecar narration and/or incidental lockfile regeneration are never part of the agent's intended deliverable)._ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Added a benchmark configuration to evaluate speculative decoding (MTP) performance for the Qwen3.5-4B model, enabling separate speed and high-throughput (32k) runs with adjustable decoding and runtime parameters. * Configured execution settings for single-node GPU benchmarking and standardized container/runtime invocation for reproducible performance tests. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
40a4dd326d |
[Feat]: Eagle Dry Run Mode (#1566)
### What does this PR do? Type of change: new feature Adds `--dry_run` to `examples/speculative_decoding/main.py`: load → `mtsp.convert` → save, then exit (no `trainer.train()`). With `FakeBaseModel`, the convert→save→export chain runs in seconds and produces an exportable EAGLE3 / Medusa / DFlash checkpoint with correct structure but untrained draft-head weights — useful for end-to-end plumbing smoke tests on downstream stacks (vLLM, TRT-LLM, SGLang) without paying for a real training run. Where it sits among existing EAGLE3 modes: ``` EAGLE3 modes ├── online base model runs forward in-loop ├── offline reads pre-dumped hidden states from disk ├── streaming streams hidden states from a live server in-loop └── dry-run ★ skip training entirely; convert + save + export ← NEW (this PR) ``` Companion `FakeBaseModel` fixes so small base checkpoints work: - Synthesize the weight_map from a single `model.safetensors` when no sharded index is present (Llama-3.2-1B, Qwen3-0.6B, …). - Honor `tie_word_embeddings`: reuse `embed_tokens` when `lm_head` is absent from safetensors. A new launcher YAML (`tools/launcher/examples/Qwen/Qwen3-8B/hf_eagle3_dryrun.yaml`) wires this together as a one-task pipeline. ### Usage ```bash # Direct python main.py --dry_run \ --config modelopt_recipes/general/speculative_decoding/eagle3.yaml \ model.model_name_or_path=meta-llama/Llama-3.1-8B-Instruct \ model.use_fake_base_for_offline=true \ data.offline_data_path=/tmp/dryrun-placeholder \ training.output_dir=ckpts/dryrun python scripts/export_hf_checkpoint.py --model_path ckpts/dryrun --export_path export/dryrun # Launcher uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_eagle3_dryrun.yaml --yes ``` ### Testing - `tests/unit/torch/speculative/plugins/test_fakebase.py`: 3 new cases — single-file fallback, tied-embeddings fallback, and the negative case (missing `lm_head` without tying). - `tests/examples/speculative_decoding/test_eagle.py::test_eagle3_dry_run`: full `launch_train.sh --dry_run → export_hf_checkpoint.py` chain on `tiny_llama`; asserts exported state_dict has all `LLAMA_EAGLE_SINGLE_LAYER` required keys. - Manually verified end-to-end on Llama-3.1-8B-Instruct (sharded), Llama-3.2-1B-Instruct (single-file + tied), and Qwen3-0.6B (single-file). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `--dry_run` is opt-in; `FakeBaseModel` changes are additive fallbacks. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — examples-only addition. - Did you get Claude approval on this PR?: ❌ ### Additional Information N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a `--dry_run` CLI flag to perform a fast early-exit execution that saves model artifacts without training. * Added support for single-file model checkpoint formats in the loader. * **Bug Fixes** * Improved checkpoint-loading error messages and handling. * Enhanced tied-embeddings fallback when head weights are absent. * **Tests** * Added integration and unit tests covering dry-run behavior and single-file/tied-embedding loading. * **Documentation** * Added a launcher example for dry-run smoke testing. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1566?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
5eba879002 |
Launcher nvrx install from PyPI; specdec_bench guard modelopt import (#1567)
## Summary - **`tools/launcher/common/service_utils.sh`** — Megatron-LM (post PR #4522) asserts `nvrx >= 0.6.0` on import. The previous git-clone-HEAD install in `util_install_extra_dep` produced setuptools_scm versions like `0.0.0.dev1+hash` whenever upstream HEAD landed between tags, failing that assertion. Switch to a PyPI install pinned at `>= 0.6.0`. nemo containers ship nvrx in two Python envs; `pip` (system) and `python -m pip` (uv venv) target different ones, so uninstall from both and install into the venv where Python actually imports from. - **`examples/specdec_bench/specdec_bench/__init__.py`** — guard the `from modelopt import __version__` import so `specdec_bench` can be loaded inside a vllm container that doesn't have modelopt installed. ## Test plan - [x] Launcher YAMLs (e.g. Qwen3-8B PTQ on ComputeLab) install `nvidia-resiliency-ext>=0.6.0` and don't trip the Megatron-LM `has_nvrx_async_support` assertion. - [x] `specdec_bench` imports cleanly under a vllm container with no modelopt installed (prints the warning and uses `__version__ = "0.0.0"`). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Streamlined dependency installation: removed fragile repo-based install and now ensures a clean uninstall/upgrade flow for the external resiliency package to improve reliability across environments. * Improved version handling: added guarded fallback and warning when an optional package is missing, reporting a safe default version to avoid runtime errors. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1567?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
a2c496af0d |
[OMNIML-4788] specdec_bench: configuration.json provenance + upload_to_s3 (#1531)
> [!WARNING] > **Breaking on-disk schema change (specdec_bench v1.0.0).** This PR renames the acceptance-rate metric fields across `AcceptanceRate` / `MTBench` / `SpecBench` writers: > > | Old (pre-1.0.0) | New (1.0.0) | > |---|---| > | `Request_AR` | `Request_AL` | > | `Category_AR` | `Category_AL` | > | `Average_AR` | `Average_AL` | > | — | `Joint_Acceptance_Rate` (new) | > > The renamed values were always **acceptance length** (mean tokens generated per inference step), not a rate, and the visualizer reads `*_AL`. Pre-1.0.0 runs in S3 have `*_AR` and no `Joint_AR`; they must be re-run or post-processed before comparing. The visualizer aggregates runs by `specdec_bench` major version so accidental cross-methodology comparison is blocked. ### What does this PR do? Type of change: new feature Adds reproducibility provenance to `specdec_bench/configuration.json` and ports `upload_to_s3.py` from `iputterman/specdec_bench@main` (personal-namespace fork) into upstream. This is the first PR in a multi-stage migration off Izzy's fork now that he's left the team. Tracked in [OMNIML-4788](https://jirasw.nvidia.com/browse/OMNIML-4788). **Provenance fields added to configuration.json** (alongside existing argv / engine_version / gpu / python_version): - `specdec_bench_version` — methodology semver declared in `specdec_bench/__init__.py`. Bump minor on additive metrics, major on changed metric *definitions*. The visualizer (Phase 4 of the migration) will aggregate runs by major version so plots don't accidentally compare across methodology changes. - `specdec_bench_sha`, `modelopt_sha`, `modelopt_version`, `nmm_sandbox_sha`, `container_image` — code/runtime provenance. Each prefers an env var set by the harness (`SPECDEC_BENCH_SHA`, `MODELOPT_SHA`, `MODELOPT_VERSION`, `NMM_SANDBOX_SHA`, `CONTAINER_IMAGE`) and falls back to `git rev-parse` / `modelopt.__version__` when running standalone. The env-var preference is necessary because the runtime container has no `.git/` (the launcher packager tarballs source without git metadata) and may not have `modelopt` installed. - `checkpoint.{path, size_bytes, index_sha256, index_source}` — cheap reproducibility fingerprint that hashes `model.safetensors.index.json` (or `config.json` fallback). Changes whenever any tensor changes. - `serving_config` — engine-level config dict captured after init via a new `Model.get_serving_config()` method. VLLM dumps `AsyncEngineArgs` + the live `vllm_config.to_dict()`; SGLANG dumps the `engine_kwargs` passed to `sgl.Engine`; TRTLLM left at the base default `{}` for a later iteration. - `timestamp` — UTC ISO 8601. **Other changes** - `upload_to_s3.py` + `specdec_bench/s3_utils.py` ported from iputterman/specdec_bench@main. Recognizes run dirs by sentinel files, refuses to overwrite existing S3 prefixes. - `_redact_config` allowlists `tokenizer`, `tokenizer_path`, `tokenizer_mode`, `tokenizer_revision` so the model path stops being redacted (latent bug from substring-matching `token` ⊂ `tokenizer`). - `requirements_speed.txt`: `boto3`, `botocore` added (used by `s3_utils`). **Out of scope** (deferred to Phase 1b / Phase 2): - `--sweep_config` driver that emits per-run-dir nesting `<sweep>/<NNN_dataset_c<conc>>/` - `--s3_upload` flag baked into `run.py` itself - Launcher auto-injection of the provenance env vars (currently the example YAML sets them statically) - `container_digest` (enroot integration) and full GPU/driver inventory - TRTLLM `get_serving_config()` ### Usage ```bash # Run a smoke benchmark (Qwen3.5-4B + vLLM + MTP draft=3) — example YAML included uv run launch.py --yaml examples/Qwen/Qwen3.5-4B/specdec_bench_mtp.yaml --yes # After it lands, upload the run directory to S3: S3_KEY_ID=team-specdec-workgroup \ S3_SECRET=... \ python upload_to_s3.py /path/to/sweep_dir s3://team-specdec-workgroup/results ``` ### Testing Cluster-tested end-to-end on cw-dfw (Slurm job 11978794, NeMo Run experiment `cicd_1779403623`, ~19 min wall): - Qwen3.5-4B + vLLM + MTP draft=3 + SPEED-Bench-Internal/qualitative (80 requests) - `configuration.json` (22 KB) populated all eight new provenance fields - `Request_AR` mean 3.327 (vs 3.330 on the pre-Phase-1a run — within noise; methodology unchanged) - `upload_to_s3.py` (real upload, not dry-run) landed [s3://team-specdec-workgroup/results/qwen35_4_mtp_smoke_2026-05-21/specdec_bench_mtp/](https://app.s8k.io/buckets/team-specdec-workgroup/?prefix=results%2Fqwen35_4_mtp_smoke_2026-05-21%2F) where the visualizer at http://10.131.132.205:8080 can pick it up. ### Before your PR is "Ready for review" - Is this change backward compatible?: ✅ - `configuration.json` only gains fields. `upload_to_s3.py` / `s3_utils.py` are new files. `Model.get_serving_config()` default = `{}` so existing subclasses without an override behave as before. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - `boto3` / `botocore` are Apache 2.0 (permissive); `upload_to_s3.py` + `s3_utils.py` are ported from a private NVIDIA repo with explicit copyright headers retained. - Did you write any new necessary tests?: ❌ - Validated by cluster smoke (see Testing). Will add unit-tests for `dump_env` provenance fields and `upload_to_s3._discover_runs` in a follow-up. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Internal-facing tooling. - Did you get Claude approval on this PR?: ❌ (triggering after open) ### Additional Information Tracked in JIRA [OMNIML-4788](https://jirasw.nvidia.com/browse/OMNIML-4788). The full multi-phase plan is on that ticket's SPEC block — this PR is Phase 1a. Cherry-picked alongside the harness change are two example YAMLs (`examples/Qwen/Qwen3.5-4B/specdec_bench.yaml` for the NONE autoregressive baseline, `..._mtp.yaml` for the MTP run) that gave us cluster-test evidence. Can be split out if preferred. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an S3 upload CLI for benchmark results with dry-run and skip-existing options * Automatic capture of run configuration, provenance and redacted environment into saved config * Models now export serving configuration for reproducible runs * New launcher entrypoint and example job configs for Qwen SPEED-Bench runs * **Documentation** * README section describing S3 upload usage and supported local layouts * **Bug Fixes / Changes** * Acceptance-rate metric keys renamed in output (AR -> AL) * **Tests / CI** * New tests for redaction and S3 utilities; CI now runs specdec_bench examples <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1531?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: chenhany <chenhany@nvidia.com> |
||
|
|
5d0441ae3d |
Create shared Megatron calibration forward loop for prune / quantize with megatron pretraining data style sequence packing (#1501)
## Summary
Replaces the bespoke calibration loops in Megatron-LM and
Megatron-Bridge prune / quantize example scripts with a single shared
utility,
`modelopt.torch.utils.plugins.megatron_calibration.get_megatron_calibration_forward_loop`.
The shared loop iterates a packed calibration dataloader built via
`get_dataset_dataloader(pack=True)` and drives a logits-free prefill
pass through the model so activation hooks fire on every layer.
`pack=True` produces Megatron-LM pretraining-style **global-stream
packing**: all raw samples are concatenated into one EOS-separated token
stream and sliced into uniform-length rows. The trained model has seen
this distribution extensively during pretraining, so the activations
produced during calibration are representative of the model's natural
behavior.
Migrates four call sites:
- `examples/megatron_bridge/prune_minitron.py`
- `Megatron-LM/examples/post_training/modelopt/{prune,quantize}.py`
(separate PR:
[NVIDIA/Megatron-LM#4881](https://github.com/NVIDIA/Megatron-LM/pull/4881))
- `Megatron-Bridge/examples/quantization/quantize.py` (separate PR)
Each call site passes `pack=True` explicitly with an inline comment so
users see the option and know when to flip it. The function-level
defaults (`get_dataset_dataloader(pack=False)`,
`get_megatron_calibration_forward_loop(pack=False)`) remain
back-compat-safe.
Unified defaults across all four sites: `--calib-dataset
nemotron-post-training-dataset-v2`, `--calib-size 1024`,
`--calib-max-sequence-length 4096`, `--calib-batch-size 1`.
## Experimental results
Qwen3-8B on full 100% MMLU (n=14042; binomial 2σ noise floor ≈ ±0.78 pt
at acc ≈ 0.7), 0-shot, eval batch_size=4. Calibration on the default
workload: nemotron-post-training-dataset-v2, seq_length=4096,
calib_batch_size=8.
**Three calibration data shapes compared:**
- **Padded**: one doc per row, padded to `seq_length`, pad tokens flow
through the forward (legacy `get_calib_dataloader` pad+truncate
behavior).
- **Trimmed**: one doc per row, each row trimmed to its real content
length via `attention_mask`, with EOS forced at the last real position;
pad never enters the forward. This is no longer part of this PR.
- **Packed**: global-stream slicing — all docs concatenated
EOS-separated into one token stream, sliced into uniform `seq_length`
rows. Matches Megatron's `.bin`/`.idx` pretraining distribution. Enabled
via `pack=True` in `get_megatron_calibration_forward_loop`.
| Workload | Padded | Trimmed | **Packed** |
|---|---|---|---|
| M-LM NVFP4 quantize (`NVFP4_DEFAULT_CFG`) | 0.707 | 0.708 | 0.709 |
| M-Bridge Minitron prune (Qwen3-8B → 30L / 3584 / 11776 ≈ 6B params) |
0.576 | 0.573 | **0.589** |
### Key findings
- **M-LM quantize quality is calibration-mode-insensitive** for dense
Qwen3-8B NVFP4 — all three modes are nearly identical.
- **M-Bridge prune**: More sensitive to calibration data shape. Packed
wins over Padded on full MMLU.
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
b49f9b9e2d |
Megatron-Bridge import example in launcher for Nemotron Super V3 (#1516)
### What does this PR do? Type of change: New feature Add Megatron-Bridge import example in launcher for Nemotron Super V3 ### Usage ```python # Usage: # update .env-slurm with environment variables # cd tools/launcher # uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/megatron_bridge_import.yaml --yes ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Megatron-Bridge Hugging Face→Megatron checkpoint import (CPU-capable entrypoint) and an example pipeline for importing the Nvidia Nemotron-3 Super 120B model. * Exposed an `output_dir` global variable for pipeline interpolation. * **Chores** * Updated pinned Megatron-LM module. * Extended .gitignore to cover `.env*` variants. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1516?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> |
||
|
|
9d0d97829a |
chore(lint): modernize typing (PEP 604/585) and enable UP032 (#1537)
### What does this PR do? Type of change: chore / refactor (no behavior change) Two small lint-cleanup commits: **1. `chore(typing): modernize Union/Optional/List to PEP 604 / 585 syntax`** (8 files) - Replace `X = Union[A, B] # noqa: UP007` with `X: TypeAlias = A | B` for the six module-level type aliases (`ModelLike`, `Criterion`, `NodeTarget`, `CalibrationDataType`, `Hparam.Importance` / `ActiveSlice`). The `TypeAlias` annotation is required so mypy continues to treat them as aliases under PEP 604. - Modernize forward-ref unions in `modelopt/onnx/quantization/autotune/` to full-string forward refs (e.g. `"RegionPattern | None"`). - Update docstring type tags in `examples/puzzletron/evaluation/hf_deployable_anymodel.py`. **2. `chore(lint): remove UP032 ignore and convert .format() to f-strings`** (10 files) - Drop `UP032` from `extend-ignore` in `pyproject.toml`. - Auto-convert 19 `"...".format(...)` calls to f-strings across export plugins, examples, tests, and tools. One conversion in `modelopt/torch/utils/plugins/megatron_generate.py` was wrapped manually to stay under the 100-char limit. **Intentionally left as-is:** - `tools/launcher/slurm_config.py` keeps its `# ruff: noqa: UP045` — nemo_run's CLI parser can't introspect PEP 604 optional annotations. - `modelopt/torch/puzzletron/*` is **not** touched. The subtree disables ruff's `UP` family entirely (per-file-ignore `"UP"`) while migration is in progress, and converting `Optional[X]` to `X | None` there would silently break runtime introspection in `block_config._get_dataclass_type` that uses `get_origin(tp) is typing.Union` (PEP 604 unions return `types.UnionType` from `get_origin`, not `typing.Union`). Best revisited when puzzletron's lint carve-out is narrowed. - `UP038` (`isinstance(x, (int, float))` → `isinstance(x, int | float)`) — ruff has officially deprecated this rule; PEP 604 in isinstance is slightly slower and misleads readers about PEP 695 / `Optional`. Ignore kept. ### Usage No user-facing API changes. ### Testing - Pre-commit hooks (ruff check, ruff format, mypy, bandit, license) pass on both commits. - Ruff status against `main`: 37 unrelated pre-existing findings (W291/W293/E501/RUF005/PLR1704); zero new findings introduced by this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — Runtime behavior of the six type aliases changes from a `typing.Union` instance to `types.UnionType`. Downstream code introspecting via `get_origin(...) is typing.Union` on these aliases would break, but no in-repo caller does this on them. (The introspection in `modelopt/torch/puzzletron/block_config.py` operates on user-supplied dataclass field types, none of which are these aliases.) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (no behavior change) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ — internal style refactor; happy to add a Misc note if reviewers want one. - Did you get Claude approval on this PR?: ❌ — not yet. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Style** * Modernized type annotations across the codebase to use Python 3.10+ union syntax and TypeAlias where appropriate. * Standardized string formatting to f-strings, improving clarity of logs, errors, and validation messages. * **Chores** * Updated linting configuration to reflect modern typing/style rules. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1537?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
df68ccdcbd |
[OMNIML-4740] synth_support (#1496)
## Summary
Adds the EAGLE3 offline pipeline YAML for `moonshotai/Kimi-K2.5-DFlash`,
adapted from `Qwen/Qwen3-8B`'s offline YAML. **`task_0` cluster-tested
green on cw-dfw** (Slurm 11782946, experiment `cicd_1778864959`, elapsed
1:02:11).
Driven by /babysit-jira on OMNIML-4740. Replaces the original
pensieve-intern synth_support agent draft (1b021024) which had three
structural issues:
1. Set `global_vars.hf_model` to the *output* path (`Kimi-K2.5 DFlash`)
instead of the *input* checkpoint.
2. Used a TRT-LLM container (`release:1.2.0`) that doesn't register
`KimiK25ForConditionalGeneration` as of 2026-05-14.
3. Committed `VERIFICATION_COMMENT.txt` (a runner sidecar, not artifact
code) and `uv.lock` regen (+1037/-81, the source of the prior
`mergeable_state: dirty`).
This PR is the cleaned + cluster-validated replacement.
## Changes
- Directory `Kimi-K2.5 DFlash` → `Kimi-K2.5-DFlash` (Slurm tar packaging
breaks on spaces in `job_name`/path).
- `global_vars.hf_model: /hf-local/moonshotai/Kimi-K2.6` — the canonical
Kimi-K2.5 input stand-in staged by the operator on cw-dfw.
- `task_0` migrated from TRT-LLM to vLLM:
- `script: common/vllm/query.sh`
- `container: vllm/vllm-openai:latest`
- `ntasks_per_node: 1` (vLLM is single-process)
- `--tensor-parallel-size 8`, `--trust-remote-code`
- `--enforce-eager` (vllm-openai:latest is missing `torch/bin/ptxas` for
inductor autotuning)
- `--gpu-memory-utilization 0.95` + `--max-model-len 4096` (Kimi weights
are 595 GB bf16 on 8×80 GB = 93% weight occupancy; default 0.9 leaves
-1.1 GiB for KV cache)
- `VLLM_STARTUP_TIMEOUT=1800` env (Kimi load is ~7.7 min, default 600s
in `query.sh` is not enough)
- `--data` switched to the in-repo `synthetic_conversations_1k.jsonl` —
the canonical `Speculative-Decoding-Prompt-Samples` isn't on cw-dfw; the
in-repo dataset is the portable, packager-shipped input for smoke
testing.
- Sidecars + uv.lock removed from the diff.
## Cluster-test evidence (mandatory per the refined synth_support spec)
```
$ SLURM_CLUSTER=cw_dfw uv run slurm.py \
--yaml '.../moonshotai/Kimi-K2.5-DFlash/hf_offline_eagle3.yaml' \
pipeline.task_1.skip=true pipeline.task_2.skip=true pipeline.task_3.skip=true \
--yes --detach
Slurm 11782946: COMPLETED, elapsed 1:02:11
Loading weights took 461.45 seconds
Model loading took 71.44 GiB memory and 465.10 seconds
Map (num_proc=32): 100%|██████████| 100/100 [06:42<00:00, 4.02s/example]
Saved 10 shards: /scratchspace/data/train-{1..10}-00010.jsonl
```
Real assistant response verified end-to-end — Kimi correctly answered
the "bat and ball" CRT problem:
> The ball costs **$0.05** (5 cents). Here's why: If the ball costs
$0.05, then the bat costs $1.05 (which is $1.00 more). Together they
cost $1.10.
## Test plan
- [x] Dry-run validation: `slurm.py --yaml ... --dry-run` exits 0
- [x] Cluster test on cw-dfw: `task_0` only (`task_1/2/3.skip=true`) —
green, 1000 prompts processed, 10 synth-data shards written
- [ ] Downstream `run_pipeline` stage of
[OMNIML-4735](https://jirasw.nvidia.com/browse/OMNIML-4735) will
exercise task_1..3 end-to-end with the produced synth data
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Improvements**
* Enhanced rank initialization with extended fallback support for
various resource scheduling systems, improving compatibility across
distributed environments.
* Improved cross-rank coordination mechanism for dependency
installation, ensuring reliable and efficient setup in multi-node
deployments.
* **New Features**
* Added configuration example for Kimi-K2.5 offline speculative decoding
pipeline with vLLM integration.
<!-- review_stack_entry_start -->
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1496?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)
<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Signed-off-by: chenhany <chenhany@nvidia.com>
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
f59f3ae1c3 |
[Example]: Dflash-Offline Launcher Example (#1529)
### What does this PR do? Type of change: new example Offline DFlash training launcher example for Qwen3-0.6B. Two-task pipeline: dump base-model hidden states via HF forward, then train DFlash on the dump. - `tools/launcher/examples/Qwen/Qwen3-0.6B/hf_offline_dflash.yaml` — new launcher YAML - `tools/launcher/common/eagle3/dump_offline_data_hf.sh` — new HF-backed dump script (DFlash's `answer_only_loss=true` requires `loss_mask`, which the existing TRT-LLM dump backend does not produce) - `examples/dataset/synthetic_conversations_1k.jsonl` — `conversation_id` added at-source (asserted by `compute_hidden_states_*.py`; previously injected at test-fixture time) - `tests/regression/torch/speculative/test_dflash_offline.py` — `tagged_synth_data_path` fixture removed since the field is now at-source ### Usage ``` uv run launch.py --yaml examples/Qwen/Qwen3-0.6B/hf_offline_dflash.yaml --yes ``` ### Testing End-to-end on CoreWeave Slurm (1×1 GPU): task_0 dump succeeded; task_1 training loss 7.85 → 2.94 over 2 epochs, regression PASSED. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (`conversation_id` is purely additive on the dataset) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (existing `test_dflash_offline.py` covers the workflow) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (new example only) - Did you get Claude approval on this PR?: ❌ ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added launcher script for offline hidden-state processing with HuggingFace backend support and distributed execution capabilities * Added training configuration example for Qwen3-0.6B model with DFlash speculative decoding pipeline * **Tests** * Updated offline hidden-states regression test to improve data handling workflow <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1529?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
e2d4d73920 |
Agent Skills Updates From Live Trials (#1493)
### What does this PR do? Type of change: bug fix <!-- Details about the change. --> ### Usage Ask Claude Code: ``` Quantize `mistralai/Mistral-Medium-3.5-128B` to NVFP4 using the ModelOpt NVFP4 experts-only recipe. Run on $cluster Evaluate the resulting quantized checkpoint on: - GPQA Diamond AA v3 - SciCode AA v2 Complete the quantization and evaluation workflow end to end. Prompt when you require user input, otherwise keep going. ``` ### Testing I'm running the full loop with the above prompt, and iterating on skills to resolve undesired agent behavior. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: TODO ### Additional Information See trials log for details. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added SLURM Quality of Service (QoS) configuration support for job submission * Introduced 8 new evaluation task recipes (AIME 2025, GPQA, IFBench, LiveCodeBench, SciCode, AA-LCR, HLE-AA, MMMU-Pro, tau2_bench) * Enhanced job monitoring with continuous polling-based tracking * **Documentation** * Restructured evaluation workflow with explicit dry-run, canary, and full-run validation stages * Expanded PTQ validation with mandatory pre-deployment verification gates * Updated remote cluster selection and quantization detection guidance * **Tests** * Updated evaluation test expectations to reflect refined workflow stages <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1493?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chad Voegele <cvoegele@nvidia.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
d7733562dd |
fix(prune): Minitron HybridModel + GPT-family fused-TE-spec import/export (#1518)
## Summary Split out of #1501 so the pack=True calibration packing change can land independently. This PR carries the pruning + export-side fixes. **Pruning bug fixes** - Register `HybridModel` (parent of `MambaModel` in modern Megatron-LM) under a new `HAS_HYBRID` flag so `mcore_minitron` actually prunes Nemotron-H et al. Previously `HybridModel` instances fell through `convert_to_dynamic`, got `freeze()`-ed (collapsing `hidden_size` / `num_layers` to a single choice), and produced unloadable saved checkpoints with mixed pruned/unpruned dims. - Replace the `isinstance(MambaModel)` gate in `_get_hybrid_pattern_key` with attribute-presence detection so both `MambaModel` (still using `hybrid_override_pattern`) and plain `HybridModel` (`hybrid_layer_pattern`) are handled uniformly. - Track `in_features` as a dynamic attribute on `_DynamicTEQKVLayerNormColumnParallelLinear` so TE's forward-time `inp_shape[-1] == in_features` assertion holds when `hidden_size` is pruned. - Dedupe MambaModel / HybridModel divisor dict into `_HYBRID_DIVISORS`. **Fused-TE-spec import/export for GPT-family** - Importer: prefer per-context keys (`fused_input_layernorm`, `fused_pre_mlp_layernorm`); fall back to legacy `fused_norm` for Nemotron-H back-compat. **Raise `KeyError`** when a fused-TE model has neither rule registered — the branch only fires when the model uses fused `TELayerNormColumnParallelLinear`, so a missing rule is unambiguously a plugin misconfig that would otherwise ship a chance-accuracy checkpoint. - Exporter: mirror the same fallback chain in `_get_fused_norm_weight` so GPT-family models round-trip cleanly back to HF. - Add the new rules to Qwen3, Qwen2.5, Llama, Llama4 (MoE-only, only `fused_input_layernorm`), DeepSeek, GptOss (MoE-only, only `fused_input_layernorm`) import and export mappings. - Preserve TE `_extra_state` from the existing module state dict (don't blank to `None`) at both call sites in the importer. **Misc** - `megatron_prefill`: `.contiguous()` on the logits slice before `broadcast_from_last_pipeline_stage` — broadcast asserts contiguity which fails when SP pads `seq_length` to a multiple of TP. - `megatron_mmlu`: accept `mmlu_dataset` kwarg so callers can point at a local copy of `cais/mmlu`. - `warn_rank_0`: auto-bump `stacklevel` by 1 inside the wrapper so callers' warnings point at user code, not at the wrapper frame. - `tools/launcher/examples/Qwen/Qwen3-8B/megatron_lm_ptq.yaml`: bump `mmlu_lower_bound` 0.68 → 0.75 (validated end-to-end with the fused-norm import fix). - CHANGELOG: bug-fix entry for the importer; date correction on the 0.44 entry. ## Consumer Megatron-LM PR https://github.com/NVIDIA/Megatron-LM/pull/4807 — `prune.py` / `mmlu.py` consume these APIs and currently ship inline WARs against released 0.44. Once 0.45 ships and the modelopt pin is bumped, those WARs collapse to one-liners. Related: #1501 (calibration packing). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Importer/exporter now correctly load fused LayerNorm weights for GPT-family models, preferring context-specific fused keys with a legacy fallback. * **New Features** * Hybrid Mamba/HybridModel support added for pruning/NAS workflows. * MMLU evaluation accepts a customizable dataset path (default: "cais/mmlu"). * **Improvements** * Extended export/import mappings and state handling across DeepSeek, GPT, Llama, Qwen; ensured last-stage logits are contiguous. * **Documentation** * Updated changelog entry and release date adjustment. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1518?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
1f9c0bfae2 |
fix(launcher): propagate dflash training failures and fast-fail vllm smoke test on missing draft (#1409)
## Summary Two related fixes to the speculative-decoding launcher scripts that surfaced as silent failures in nmm-sandbox CI pipeline [50509971](https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/pipelines/50509971): - **`dflash_online_training.sh`** — install an `EXIT` trap so `error_handler`'s `FAIL_EXIT=1` actually surfaces as a non-zero shell exit. Previously, training failures (e.g. CUDA OOM) emitted `<REPORT>` tags but the script exited 0, and Slurm marked the task `SUCCEEDED`. The orchestrator then proceeded to run dependent tasks against a non-existent checkpoint. - **`vllm_smoke_test.sh`** — when `DRAFT_CKPT_DIR` is set but contains no `exported-checkpoint-*` (typically because the upstream training task crashed), exit 1 with a clear message instead of falling through to the self-draft branch. The self-draft fallback produced a confusing `num_speculative_tokens was provided but without speculative model` error from vLLM's `SpeculativeConfig` validation that obscured the real cause (the upstream training failure). ## Test plan - [ ] Re-run `Qwen3-8B_DFlash_online` on Computelab with insufficient GPU memory; verify task_0 surfaces as `FAILED` (not `SUCCEEDED`) and task_1 doesn't bother trying the self-draft path. - [ ] Re-run on adequate hardware (H200/B200) and verify the trap doesn't fire for the success path (script exits 0 normally). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Bug Fixes * Improved error handling in training script to properly detect and report training failures (e.g., out-of-memory errors) instead of incorrectly marking them as successful * Enhanced validation for missing draft checkpoints with clearer error messaging to prevent confusing downstream configuration errors <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
f34f488a83 |
Add a general composable $import system for YAML configs, and use it to implement composable recipes (#1253)
### What does this PR do?
Type of change: New feature
Adds a general composable YAML config loading layer for ModelOpt configs
and recipes. YAML remains the source of truth for configuration data,
while Python/Pydantic-compatible types provide schema validation at load
time. This PR uses that loader to de-duplicate PTQ recipes, introduce
reusable config snippets/presets, and start migrating selected hardcoded
quantization presets to YAML.
#### Problem
1. Built-in PTQ recipes duplicated numeric format definitions, KV-cache
entries, and the default quantizer exclusion list.
2. YAML snippets were reusable only by convention; they did not declare
or validate the schema they were meant to satisfy.
3. Loading YAML-backed quantization presets from
`modelopt.torch.quantization.config` could not depend on
`modelopt.recipe` without creating circular imports.
4. Directory-format recipes exposed import-resolution details in the
recipe loader and used `recipe.yaml` plus nested `metadata:` in a way
that made metadata handling inconsistent.
#### Solution
**Shared YAML config loader**
- Adds `modelopt.torch.opt.config_loader` as the low-level loader used
by both `modelopt.recipe` and `modelopt.torch.quantization.config`.
- Keeps the public `modelopt.recipe.load_config()` entry point, while
removing the private `modelopt/recipe/_config_loader.py` shim.
- Handles YAML loading, built-in/filesystem path resolution, suffix
probing, `ExMy` conversion for `num_bits` / `scale_bits`, `$import`
expansion, and schema validation.
- Lives below `modelopt.recipe` in the dependency graph to avoid
circular imports from quantization config code.
**Composable `$import` system**
Recipes and snippets can declare an `imports` mapping, then reference
entries with `{$import: name}`.
`$import` semantics:
- **Dict value**: replaced with the imported dict. Multiple imports are
supported with ordered precedence; inline keys override imported keys.
- **List entry**: schema-driven behavior for strongly typed lists. If
the snippet schema matches the containing list type, the imported list
is spliced. If the snippet schema matches the list element type, the
imported element is appended. Other schema combinations are rejected.
- **Multi-document YAML**: supports snippets that need an `imports`
header plus a list body.
- **Recursive and scoped**: snippets can import other snippets; import
names are scoped per file.
- **Cycle detection**: circular imports report a clear error.
**Snippet schema validation**
- Every reusable snippet referenced through `imports` must declare a `#
modelopt-schema: ...` preamble.
- Snippets are validated after nested imports are resolved.
- Schema paths are restricted to the `modelopt.` package and may be
Pydantic models, `TypedDict` classes, or explicitly typed container
aliases such as `list[QuantizerCfgEntry]`.
- Untyped list imports are rejected so list append/splice behavior stays
strongly typed.
**Recipe model and directory recipe cleanup**
- `ModelOptRecipeBase` now owns a `metadata: RecipeMetadataConfig`
field.
- `ModelOptPTQRecipe` is the PTQ recipe schema; the overlapping
YAML-specific PTQ config class was removed.
- Directory recipes now use `metadata.yaml` / `metadata.yml` for
top-level metadata fields, plus section files such as `quantize.yaml`.
- Directory recipe loading now delegates import resolution to
`load_config()` instead of manually using raw config loading.
**Config snippet and preset library**
Adds reusable snippets under `modelopt_recipes/configs/`:
- `numerics/`: `fp8`, `nvfp4`, `nvfp4_static`
- `ptq/units/`: `base_disable_all`, `default_disabled_quantizers`,
`w8a8_fp8_fp8`, `w4a4_nvfp4_nvfp4`, `kv_fp8`, `kv_fp8_cast`,
`kv_nvfp4_cast`
- `ptq/presets/`: YAML presets for `FP8_DEFAULT_CFG` and `FP8_KV_CFG`
`FP8_DEFAULT_CFG` and `FP8_KV_CFG` now load from YAML presets via
`load_config()`.
**Recipe migration and naming**
- General PTQ recipes now use shared imports instead of repeating the
same quantizer fragments inline.
- General PTQ recipe paths were renamed to KV-first naming, for example:
- `general/ptq/fp8_default-fp8_kv` -> `general/ptq/fp8_default-kv_fp8`
- `general/ptq/fp8_default-fp8_cast_kv` ->
`general/ptq/fp8_default-kv_fp8_cast`
- `general/ptq/nvfp4_default-none_kv_gptq` ->
`general/ptq/nvfp4_default-kv_none-gptq`
- `general/ptq/nvfp4_default-nvfp4_cast_kv` ->
`general/ptq/nvfp4_default-kv_nvfp4_cast`
- Example docs and `examples/llm_ptq/hf_ptq.py --recipe` help text were
updated to use the new paths.
**Pre-commit and documentation**
- Recipe validation accepts `$import` entries and handles directory
recipes using `metadata.yaml`.
- The recipe validation hook skips `modelopt_recipes/configs/` because
those files are reusable snippets, not full recipes.
- `docs/source/guides/10_recipes.rst` now documents imports, schema
modelines, list append/splice semantics, built-in snippets, built-in
recipe paths, directory recipes, and the current recipe data model.
#### Backward compatibility
- Existing inline YAML recipes without `$import` continue to load.
- `modelopt.recipe.load_config()` remains public.
- The built-in recipe path renames are user-visible; callers should
update recipe path strings to the KV-first names listed above.
#### Testing
- `pytest tests/unit/recipe/test_loader.py -q` - 90 passed
- `python tools/precommit/check_modelopt_recipes.py ...` for the renamed
built-in PTQ recipes
- `pre-commit run mypy --files
modelopt/onnx/llm_export_utils/quantization_utils.py`
- `python -m py_compile examples/llm_ptq/hf_ptq.py`
- `git diff --check`
### Before your PR is "Ready for review"
- Is this change backward compatible?: Partially. Loader/API behavior is
compatible for existing inline YAML recipes, but built-in recipe path
names were renamed to KV-first paths.
- Did you write any new necessary tests?: Yes.
- Did you update Changelog?: Yes.
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
9d2e6087d1 |
[Fix]: $HOME in launcher eagle example (#1365)
### What does this PR do? Type of change: Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> Launcher example bug raised by @cjluo-nv Before fix: task1 in tools/launcher/examples/Qwen/Qwen3-8B/hf_online_eagle3.yaml fails Reason: due to `HOME: /tmp` set in container, enroot credentials in `$HOME/.config/enroot/.crendential` not found ``` GpuFreq=control_disabled pyxis: importing docker image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc10 Apr 28 13:35:59.491365 2515157 slurmstepd 0x155552c3b780: error: pyxis: child 2515158 failed with error code: 1 Apr 28 13:35:59.491415 2515157 slurmstepd 0x155552c3b780: error: pyxis: failed to import docker image Apr 28 13:35:59.491433 2515157 slurmstepd 0x155552c3b780: error: pyxis: printing enroot log file: Apr 28 13:35:59.491453 2515157 slurmstepd 0x155552c3b780: error: pyxis: [INFO] Querying registry for permission grant Apr 28 13:35:59.491469 2515157 slurmstepd 0x155552c3b780: error: pyxis: [INFO] Authenticating with user: <anonymous> Apr 28 13:35:59.491483 2515157 slurmstepd 0x155552c3b780: error: pyxis: [INFO] Authentication succeeded Apr 28 13:35:59.491499 2515157 slurmstepd 0x155552c3b780: error: pyxis: [INFO] Fetching image manifest list Apr 28 13:35:59.491512 2515157 slurmstepd 0x155552c3b780: error: pyxis: [INFO] Fetching image manifest Apr 28 13:35:59.491524 2515157 slurmstepd 0x155552c3b780: error: pyxis: [ERROR] URL https://registry-1.docker.io/v2/nvcr.io/nvidia/tensorrt-llm/release/manifests/1.3.0rc10 returned error code: 401 Unauthorized Apr 28 13:35:59.491564 2515157 slurmstepd 0x155552c3b780: error: pyxis: couldn't start container Apr 28 13:35:59.491579 2515157 slurmstepd 0x155552c3b780: error: spank: required plugin spank_pyxis.so: task_init() failed with rc=-1 Apr 28 13:35:59.491593 2515157 slurmstepd 0x155552c3b780: error: Failed to invoke spank plugin stack Apr 28 13:35:59.515523 2515146 slurmstepd 0x155552c3b780: error: pyxis: child 2515240 failed with error code: 1 ``` After fix: ``` GpuFreq=control_disabled pyxis: importing docker image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc10 pyxis: imported docker image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc10 ``` ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated example pipeline to use the standardized dataset example path. * Removed unnecessary per-task overrides of the process home and cache directory to simplify environment setup. * Preserved required model checkpoint environment setting for the relevant task so model resolution continues to work. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
355c6b7883 |
fix: PTQ 1GPU, export PP divisibility, hidden states conversations key (#1293)
## Summary - **megatron_lm_ptq.yaml**: Qwen3-8B PTQ to single GPU for L40 clusters (TP=1, all tasks) - **quantize.sh**: Auto-find largest PP dividing model's `num_hidden_layers` for export step. Qwen3-8B has 36 layers which isn't divisible by 8, causing `AssertionError` on 8-GPU nodes - **compute_hidden_states_trtllm.py**: Use `messages` with `conversations` fallback, matching the HF version. Fixes `KeyError: 'conversations'` when data uses OpenAI `messages` format ## Test plan - [x] Qwen3-8B PTQ runs on single L40 GPU - [x] Export PP auto-selects valid divisor (36 layers → PP=6 on 8 GPUs, PP=4 on 4 GPUs, PP=1 on 1 GPU) - [x] EAGLE3 offline pipeline reads data with `messages` field 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Dataset input handling now supports multiple field formats for enhanced compatibility. * **Bug Fixes** * Optimized GPU resource allocation during model quantization with improved pipeline parallelism computation. * Updated quantization configuration for more efficient resource utilization. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
3d0f0db49e |
[CI] Replace tox with nox, use nemo:26.04 for megatron tests, and simplify CI workflows (#1286)
### What does this PR do? Type of change: New feature / infrastructure improvement Follow-up to #1285 for correct CI test environment for megatron based tests Replaces `tox` + `tox-current-env` with `nox` for all test, lint, docs, and wheel build sessions. The primary motivation was that `tox-current-env` is incompatible with uv venvs in NGC containers (e.g. NeMo's `/opt/venv`) — it picks the system Python via `sys._base_executable` instead of the container's venv Python which has megatron packages pre-installed. Key changes: - **`noxfile.py`** replaces `tox.ini` with GPU, CPU unit, partial-install, pre-commit, docs, and wheel sessions - **GPU sessions** use `venv_backend="none"` (run directly in container env) and `python -m pip/pytest` to avoid PATH mismatches - **uv** is set as the default venv backend (if available) for CPU sessions (faster installs) Also includes CI workflow simplifications: - **`_pr_gate.yml`** new reusable workflow centralizing file-change detection + linux-check wait logic (was duplicated across 3 workflow files) - **Collapsed pr/non-pr job pairs** into single jobs with conditional `runs-on` in `gpu_tests.yml`, `example_tests.yml`, `regression_tests.yml` - **Collapsed `multi-py` / `multi-torch` / `multi-transformers`** into a single `multi-version` matrix job in `unit_tests.yml` - **PR path filtering** for unit test secondary jobs (multi-version, launcher, partial-install) — skipped if no relevant files changed - **Fixed schedule/workflow_dispatch skipping** — jobs with `needs: [pr-gate]` were incorrectly skipped when all pr-gate internal jobs were skipped; fixed by making the gate job always run - **multi-version, launcher, partial-install** now also run on `schedule` / `workflow_dispatch` ### Usage ```bash python -m pip install nox uv # install nox and uv (once) nox -l # list all sessions nox -s gpu_megatron # run a GPU session (inside container) nox -s "unit-3.12(torch_211, tf_latest)" # run a specific unit test combination nox -s "unit-3.12(torch_211, tf_latest)" -R # force-recreate venv (e.g. after dep changes) COVERAGE_PROCESS_START=pyproject.toml nox -s "unit-3.12(torch_211, tf_latest)" # with coverage ``` ### Testing - Ran `nox -l` to verify all session names - Ran `gpu_megatron` session locally inside NeMo container — confirmed it uses `/opt/venv/bin/python` correctly - Manually triggered nightly-runs: - Unit: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608013657 - GPU: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608018763 - Examples: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608017322 ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A — CI infrastructure only - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ (added `nox` and `uv` to `dev-test`, both Apache-2.0) - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — no user-facing changes ### Additional Information Supersedes the tox-current-env workaround in the parent branch. --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
76b6fd51a5 |
fix: DFlash regression tests and vLLM server liveness (#1288)
## Summary - **hf_online_dflash.yaml**: Add 100K-sample training config with regression baselines (B200 loss curve), `MAX_FINAL_LOSS`/`MIN_FINAL_ACC`/`MIN_ACCEPTANCE_LENGTH` thresholds, vLLM nightly container for DFlash support - **vllm_smoke_test.sh**: Parse acceptance length from vLLM server log for regression check; `pip install pandas` workaround for broken nightly container; capture server output to temp file - **query.sh**: Detect vLLM server death during startup (PID liveness check) + 600s timeout to prevent infinite polling that wastes GPU hours; `pip install pandas` workaround - Fix empty `environment:` key in DFlash YAML causing nemo_run `ListParseError` ## Test plan - [x] E2E pipeline passed on 8x B200 (training + vLLM smoke test + AR eval) - [x] Training regression: final loss 3.82 < 5.0, acc 0.20 > 0.15 - [x] vLLM acceptance length: 1.79 >= 1.4 threshold - [x] AR evaluation: 2.02 overall on MT-Bench (8 categories) - [x] Server liveness check prevents GPU waste on vLLM crash 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added optional regression validation for vLLM acceptance metrics * Introduced configurable vLLM server startup timeout (default 600 seconds) * **Improvements** * Enhanced logging for vLLM server startup with progress tracking and waited time reporting * Faster detection of vLLM server process failures during initialization * **Configuration Updates** * Increased training dataset size and logging granularity * Scaled tensor parallelism from 4 to 8 across multiple pipelines * Expanded PTQ quantization to multi-step pipeline * Added configurable training metric thresholds <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |