Commit Graph
65 Commits
Author SHA1 Message Date
Jenny ChenandClaude Opus 4.8 e2c4d083d4 [OMNIML-4922] Four over Six PTQ & Updating Nemotron Ultra Example (#1684)
### What does this PR do?

[Four Over Six](https://arxiv.org/pdf/2512.02010) PTQ implementation for
weight-only quantization. Four Over Six was used to produce the Nemotron
3 Ultra NVFP4 checkpoint. Also updates the Ultra PTQ example in the
launcher to use this new 4/6 config
`huggingface/nvidia/Nemotron-3-Ultra-550B-A55B/ptq/ultra-nvfp4-46-max`

### Usage

```bash
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/megatron_lm_ptq.yaml --yes
```

### Testing

- [x] Unit tests pass
- [x] Run launcher example

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added NVFP4 Four‑Over‑Six (4/6) adaptive per‑block weight scaling and
a configurable FP8 normalization option for FP4/FP8 quantization.

* **Documentation**
* Added PTQ recipes/presets and updated config docs to document FP8 max
variants and the Four‑Over‑Six option.

* **Tests**
* Added unit and GPU tests validating 4/6 selection, normalization
threading, scaling behavior, and reconstruction error checks.

* **Chores**
* Updated a PTQ pipeline example to use the NVFP4‑46‑max recipe and
bumped a container image; minor project config tweak.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Jenny Chen <jennifchen@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 15:00:25 +00:00
Chenhan D. Yu 37dbbdac5a Fix ModelOpt MCP Slurm launcher submit (#1799)
## Summary
- fix launcher Slurm task annotation patching so nemo-run CLI resolves
`slurm_factory` correctly for task slots
- harden `modelopt-mcp` submit parsing/status resolution and add
regression coverage for launcher false-positive success cases
- add a minimal `nvidia-smi` smoke YAML/script and fix launcher
packaging so source-backed Slurm jobs package required files recursively

## Validation
- `uv run pytest tests/test_core.py -q`
- `uv run pytest tests/test_bridge.py -q`
- dry-run and live-submit validated through the patched local MCP server
on `cw_dfw`
- interactive smoke job succeeded end-to-end (`nvidia-smi` ran
successfully in-container)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an NVIDIA SMI GPU smoke test (script and minimal Slurm YAML
example) for launcher integration.
* **Bug Fixes**
* Improved detection of fatal launcher errors, including when the
launcher exits with code 0.
* Strengthened Slurm experiment/job identifier parsing and added early
rejection of unsafe experiment IDs, with clearer “unparsed”/failure
behavior.
* Updated sandbox task Slurm config type handling and improved launcher
packaging so examples/common are included consistently.
* **Tests**
* Expanded unit and filesystem-based coverage for parsing/validation,
dry-run fatal stderr handling, and nested experiment directory layouts.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Signed-off-by: Chenhan D. Yu <chenhany@nvidia.com>
2026-06-24 01:05:59 +05:30
Jenny Chen 28b5e26fdb Fix torch import error to remove circular dependency & move Nemotron configs (#1606)
### What does this PR do?

Type of change:  Bug fix
when running Megatron-LM modelopt example `generate.py` a circular
import causes an error. the cause was because it import
modelopt.torch.quantization which imports modelopt/torch -->
```
modelopt/torch/__init__.py:26. The chain is: import modelopt.torch.opt → runs modelopt/torch/__init__.py (parent package) → which pulls in distill
  → distill/mode.py:25 → back into opt before opt.utils is bound.
```

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Clarified configuration comments in Nemotron-3-Super-120B quantization
recipes for improved clarity.

* **Chores**
  * Internal package initialization updates.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
2026-06-23 18:21:17 +00:00
Chenhan D. Yu 985ad4361c launcher: make Slurm memory and user defaults configurable (#1791)
## Summary

- Adds `SlurmConfig.mem` and a `SLURM_MEM` env override for launcher
Slurm jobs.
- Passes configured memory through to `nemo_run.SlurmExecutor` instead
of always using `"0"`.
- Lets `SLURM_USER` provide the launcher default user when local and
cluster usernames differ.
- Adds focused tests for Slurm memory defaults and overrides.

## Test plan

- [x] `uv run pytest tools/launcher/tests/test_slurm_config.py
tools/launcher/tests/test_slurm_executor.py`


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Slurm job memory can now be configured via `SLURM_MEM`, with a default
of `0` when unset.
* Job launch username can now be set via `SLURM_USER`; if omitted, it
falls back to the local login name.

* **Bug Fixes**
* Slurm executor now correctly uses the configured memory value from
Slurm settings, falling back to `0` only when missing/empty.

* **Tests**
* Expanded unit tests to cover default, environment-driven, and executor
parameter memory behavior (including missing/empty/`None` cases).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-06-22 14:39:31 -07:00
Chenhan D. YuandClaude Sonnet 4.6 2a88b6056c launcher: package as modelopt_launcher; mcp: call console script directly (#1766)
## Summary

- `tools/launcher/__init__.py` gains `PACKAGE_DIR`, making the directory
an importable package named `modelopt_launcher` via a `package-dir`
mapping in `pyproject.toml`. **No file moves** — `common/` and
`examples/` stay exactly where they are.
- `pyproject.toml` adds the `packages`/`package-dir`/`package-data`
declarations and a `modelopt-launcher` console script entry point.
- `launch.py` switches imports to
`modelopt_launcher.{core,slurm_config}` (works in both `uv run
launch.py` via editable install and the installed console script), adds
`_has_modelopt_src` to skip packaging modelopt source when running
installed (cluster container already has it), and adds `main()`.
- `bridge.py` deletes the 75-line `_find_launcher_dir()` filesystem
walker and `_launcher_dir_not_found_response()`; simplifies
`_find_launcher_examples_dir()` to 2 strategies (env override → `import
modelopt_launcher`); switches `submit_job` subprocesses from `["uv",
"run", "launch.py"]` with `cwd=launcher_dir` to `["modelopt-launcher"]`
with no `cwd`.
- `tools/mcp/pyproject.toml` declares `modelopt-launcher` as a proper
dependency (dev: editable `../launcher`; published: PyPI), replacing a
lengthy comment explaining why it could not be declared.
- 7 tests for the deleted functions removed; all 38 remaining tests
pass.

## Test plan

- [ ] `cd tools/launcher && uv run python3 -m pytest tests/ -v` — all 65
tests pass
- [ ] `cd tools/mcp && uv run python3 -m pytest tests/ -v` — 38 tests
pass
- [ ] `uv run launch.py --yaml
examples/Qwen/Qwen3-8B/megatron_lm_ptq.yaml --dryrun --yes -v` — dry-run
resolves correctly
- [ ] `modelopt-launcher --yaml
examples/Qwen/Qwen3-8B/megatron_lm_ptq.yaml --dryrun --yes` — console
script works after `pip install -e tools/launcher`
- [ ] `cd tools/mcp && uvx modelopt-mcp` — MCP server starts without
FileNotFoundError

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added a `modelopt-launcher` CLI command as the standardized entry
point for launching optimization jobs.

* **Bug Fixes**
* Improved launcher detection and error reporting when the launcher
isn’t installed.
* Simplified experiment-directory discovery and standardized environment
handling for job submission and log retrieval.

* **Chores**
* Updated launcher packaging and example/resource discovery to work
reliably from installed distributions.
  * Added dev-mode symlink and cleanup safeguards.

* **Tests**
  * Adjusted expectations for draft PR failure behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-22 12:07:44 -07:00
Chenhan D. Yu 12ae5fbb35 [OMNIML-5233] hf_synth.yaml: relative-leaf output_dir contract (#1773)
## What does this PR do?

**Type of change:** Bug fix

**Overview:** Fix the `output_dir` shape for
`tools/launcher/examples/Qwen/Qwen3-8B/hf_synth.yaml` so the downstream
synth_run stage (OMNIML-5233) can submit + resume + publish correctly.
Replaces the prior `/scratchspace/modelopt/qwen3-8b-synth-v1` (which
fails resume + publish) with the canonical **relative-leaf** shape:
`hf-local/modelopt/qwen3-8b-synth-v1`.

The absolute team-folder prefix is **deliberately not committed** — it's
an NVIDIA-internal cluster mount path, and `pensieve-intern`'s
`_scan_for_internal_path_leak` guard refuses any YAML diff that bakes it
(rightly, since this repo is public). The downstream stage that has the
cluster context (synth_run) resolves the prefix via
`mcp__nmm-sandbox__resolve_team_folder` at submit time and injects the
absolute path through `extra_overrides`.

**Why this shape:** the relative leaf `hf-local/modelopt/<id>` carries
the contract between authoring (this stage), submitting (synth_run), and
publishing (`modelopt-storage publish`). Publish promotes from the
team's `hf-local/` tree as a metadata-only op; the leaf shape tells the
operator + downstream what category the artifact lands in, without
leaking the cluster mount.

**Companion MR:** [pensieve-intern
!126](https://gitlab-master.nvidia.com/omniml/integration/pensieve-intern/-/merge_requests/126)
lands the matching SPEC contract in synth_support.md + synth_run.md, AND
elevates the path-contract rule to the engine preamble
(`_ENGINE_AGENT_RULES_PREAMBLE`) so every agent dispatch — across
agent/subprocess/pensieve-artifact tasks — reads it as workflow-wide
common knowledge.

## Usage
Same as before. The shard write location is now resolved at submit time
by synth_run.

## Testing
- [x] One-line YAML change.
- [ ] End-to-end: re-fire OMNIML-5233 synth_run after this + the
pensieve-intern MR merge; expect the agent to submit the slurm array and
stamp success.

## Before your PR is "Ready for review"

- [x] Make sure you read and follow Contributor guidelines and your
commits are signed.
- [x] Is this change backward compatible? **Yes** (Qwen3-8B is the only
consumer; no published checkpoints reference the old path).
- [x] Did you write any new necessary tests? **N/A** — example YAML;
behavior covered by the downstream synth_run agent run.
- [x] Did you add or update any necessary documentation? Covered in the
comment block on the changed lines.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Updated the configured output directory to a more stable,
persist-friendly path so generated artifacts remain available across
repeated runs and re-dispatches.
* Added clearer inline guidance explaining how the resolved absolute
path under the new directory is used to reuse prior shards and to handle
publishing consistently.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-06-19 16:18:46 -07:00
Chenhan D. YuandClaude Sonnet 4.6 fa1d13f85d launcher: add Nemotron-3-Super-120B-A12B-BF16 MTP vLLM specdec bench config (#1714)
Adds a SPEED-bench MTP speculative-decoding YAML for
`NVIDIA-Nemotron-3-Super-120B-A12B-BF16` via vLLM.

Covers two splits:
- `qualitative` — 32 concurrent, 4096 output tokens
- `throughput_32k` — 8 concurrent, 80 requests, 4096 output tokens

Both tasks run `tp_size=4` on a single 4×H100/A100 node.

Part of OMNIML-5095 / OMNIML-5098.

## Test plan
- [ ] `uv run slurm.py --yaml
modules/Model-Optimizer/tools/launcher/examples/Nemotron-h/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/specdec_bench_mtp_vllm.yaml
--dryrun --yes -v`

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Chores**
* Added a new benchmark configuration for evaluating NVIDIA
Nemotron-3-Super-120B model performance using speculative decoding with
vLLM. Configuration includes qualitative and throughput benchmark tasks.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-19 16:18:28 -07:00
93dd08f429 [OMNIML-4760] synth_support (#1696)
Draft PR opened by **pensieve-intern** for
[OMNIML-4760](https://jirasw.nvidia.com/browse/OMNIML-4760).

Stage `synth_support` of Epic `OMNIML-4755`. The agent ran from the SPEC
on the ticket description; review every change before marking ready.

_Always-draft is enforced — the bot never auto-merges._

---

**Agent's self-narration** (stripped from PR diff; surfaced here for
context):

`VERIFICATION_COMMENT.txt`:
```
OMNIML-4760 synth_support verification for qwen3-8b found and fixed one task_0 issue in `tools/launcher/examples/Qwen/Qwen3-8B/hf_offline_eagle3.yaml`.

What changed:
- Updated `task_0` data input from missing `/hf-local/modelopt/Speculative-Decoding-Prompt-Samples` to existing `/hf-local/modelopt/Speculative-Decoding-Dataset-v1-Qwen3-8B/default-openai.jsonl`.
- Left the rest of the 4-task monolithic `hf_offline_eagle3.yaml` unchanged.

Verification:
- `task_0.script` is `common/tensorrt_llm/query.sh`.
- `--model <<global_vars.hf_model>>` resolves to `/hf-local/Qwen/Qwen3-8B`.
- `task_0.slurm_config.container` is `nvcr.io/nvidia/tensorrt-llm/release:1.2.0`.
- Cluster validation on cw_dfw succeeded for the fixed task_0 data path: experiment `cicd_1781221901`, Slurm job `12739762`, remote directory `/lustre/fsw/portfolios/coreai/users/chenhany/experiments/cicd/cicd_1781221901/Qwen3-8B_EAGLE3_offline_task0_verify_jsonl_0`.
- Log evidence included successful SSH tunnel/authentication, TensorRT-LLM 1.2.0 container import, Qwen3-8B server health checks, a successful `/v1/chat/completions` request, and loading `1393367` train examples from the Qwen3-8B JSONL data file.

Next:
- Runner should open the single-file PR for human review because this was a task_0 config fix, not verification-only.
```



_Pollution-strip removed `VERIFICATION_COMMENT.txt` from this commit
(sidecar narration and/or incidental lockfile regeneration are never
part of the agent's intended deliverable)._

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Enhanced TE grouped MoE weight quantization with optional
per-GEMM/per-expert quantizers (environment-controlled) and updated
calibration flow
* Expanded launcher CLI with `--shard-id` and per-run sample limiting
via `--num-samples`
* Added a Qwen3-8B standalone vLLM synthesis job example
(`hf_synth.yaml`)
  * Added Slurm job requeue support
* **Bug Fixes**
* Improved sharded dataset synthesis with idempotent `.done` markers and
smarter shard sizing/capping
* **Tests**
* Added coverage validating TEGrouped vs sequential MoE default amax
behavior and output divergence/accuracy
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Pensieve Intern <pensieve-intern@nvidia.com>
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Pensieve Intern <pensieve-intern@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-18 20:35:10 -07:00
Chenhan D. Yu e0125294f9 launcher: add Qwen3-8B/specdec_bench_dflash_vllm.yaml parent (OMNIML-5057) (#1764)
## Summary

Adds the parent YAML for the Qwen3-8B / DFlash / vLLM SPEED-bench sweep
(Epic OMNIML-5057, 4 cells: t0_d3 / t0_d7 / t1_d3 / t1_d7). Mirror of
[`Qwen/Qwen3.5-4B/specdec_bench_dflash_vllm.yaml`](https://github.com/NVIDIA/Model-Optimizer/blob/main/tools/launcher/examples/Qwen/Qwen3.5-4B/specdec_bench_dflash_vllm.yaml).

### Differences vs Qwen3.5-4B template

- `hf_model`: `/hf-local/Qwen/Qwen3-8B`
- `draft_model`: `/hf-local/modelopt/Qwen3-8B-DFlash-bs8-seq4096-150000`
— an internal NVIDIA checkpoint, not on HF Hub. `cell.md` Step 1's
`PRIVATE_DRAFT_ORGS` branch already covers the expected 404; the cluster
has the file staged by the Epic's `prep_inputs` stage.
- `block_size`: 8 (was 4) — matches the canonical `cell_t0_d7`
draft_length=7.

### Known limitation (acceptable per Shape (2))

Qwen3-8B's `max_position_embeddings = 40960`. The `throughput_32k` split
contains rows whose prompts exceed 40960 input tokens; those rows fail
with the vLLM context-limit assert. Per `cell.md` "Shape (2)" recovery
contract, cells ship qualitative metrics + `null` throughput_32k AL. See
OMNIML-5060 Notes (2026-06-15 correction) for the empirical evidence.

### Reference run

`cicd_1781655951` (OMNIML-5060 pipeline #55054230, cw_dfw):
- qualitative `Average_AL` = **3.4849**
- throughput_32k = `null` (Shape (2) — overlong rows skipped)

### Cascade

3 subprocess cells (OMNIML-5059 / 5061 / 5062) re-run slurm against this
parent YAML at invoke time with per-cell CLI overrides for temperature /
block_size / save_dir / max_seq_len.

## Test plan

- [x] `tools/precommit/check_launcher_yaml.py` passes locally
- [x] YAML successfully drives a real slurm submission
(`cicd_1781655951`) producing valid qualitative SPEED-bench output
- [ ] Once merged: `intern_advance OMNIML-5057` fires the 3 subprocess
cells against this YAML

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Added benchmark configuration file for evaluating speculative decoding
performance with Qwen3-8B model using DFlash acceleration.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-06-18 11:40:00 -07:00
h-guo18 e6790ef7b4 [Examples]: GPT-oss, Qwen3Moe streaming specdec example (#1692)
### What does this PR do?

Type of change: new example

Adds **streaming speculative-decoding examples (EAGLE3 + DFlash)** for
**gpt-oss-20b** and **Qwen3-30B-A3B** to the ModelOpt launcher,
mirroring the existing Qwen3-8B/Kimi examples.

- New yamls:
`tools/launcher/examples/{openai/gpt-oss-20b,Qwen/Qwen3-30B-A3B}/hf_streaming_{eagle3,dflash}_multi_node.yaml`,
plus gpt-oss `chat_template_train.jinja` (generation-tagged, for
`answer_only_loss`).
- `eagle_utils.py`: the streaming path now installs a custom
`data.chat_template` on the tokenizer (the online path already did) —
needed for the tagged template.

### Usage

```bash
cd tools/launcher
export SLURM_HOST=... SLURM_ACCOUNT=... SLURM_HF_LOCAL=... SLURM_JOB_DIR=...
uv run launch.py --yaml examples/openai/gpt-oss-20b/hf_streaming_eagle3_multi_node.yaml --yes
```

### Testing

Pipeline sanity test on **unsynthesized** data (daring-anteater), 1×
H100-80GB, 12k steps. All four train and pass the vLLM acceptance-length
eval:

| Model | Method | Train speed | vLLM AL |
|---|---|---|---|
| Qwen3-30B-A3B | EAGLE3 | 7.12 it/s | **1.74** |
| Qwen3-30B-A3B | DFlash | 2.31 it/s | 1.29 |
| gpt-oss-20b | EAGLE3 | 5.07 it/s | 1.19 |
| gpt-oss-20b | DFlash | 2.01 it/s | 1.14 |

<img width="1300" height="780" alt="image"
src="https://github.com/user-attachments/assets/2be22562-5b77-4a9f-9dc0-6f936a059736"
/>

> Sanity test only, not a quality run. gpt-oss AL is low because it is a
reasoning model (CoT at inference) while daring-anteater has no
reasoning traces and `answer_only_loss` masks all but the final content
— quality runs need synthesized/reasoning data.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added multi-node speculative decoding pipeline configurations for
Qwen3-30B-A3B and gpt-oss-20b with DFlash and EAGLE3 support.
* Introduced chat template training support for improved model
instruction formatting.

* **Enhancements**
* Increased benchmark concurrency from 1 to 32 across Qwen3-8B
configurations for more realistic performance evaluation.
  * Extended training runs from 500 to 2000 steps for Kimi-K2.5 models.
  * Improved chat template handling in speculative decoding workflows.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-06-15 16:49:39 -07:00
yeyu-nvidiaandClaude Opus 4.6 e004d8d90e DFlash speculative decoding for MiniMax-M2.7 (FSDP2): auto mask-token, FSDP2 resume fixes, per-checkpoint draft export (#1621)
## What

Brings up DFlash block-diffusion speculative decoding for large MoE
targets (MiniMax-M2.7, 229B) trained under accelerate FSDP2, and fixes
the regressions that broke checkpoint resume and per-checkpoint draft
export.

## Commits
- **auto-add mask token for DFlash** when the tokenizer lacks one
(resize embeddings, restore dtype).
- **requeue support** in `build_slurm_executor` + **FSDP2
cpu_ram_efficient_loading** for 229B on multi-node.
- **FSDP2 buffer patch** (`fsdp2_buffer_patch.py`): handle non-DTensor
buffers in `fsdp2_load_full_state_dict`, broadcast dtype codes from rank
0, and an FSDP2-safe `clip_grad_norm_`. Required because MiniMax-M2.7
pins transformers 4.57.x (no native `ParallelismConfig`).
- **dtype fix**: use the broadcast dtype (rank 0) rather than the local
meta-device param dtype, so non-leader ranks don't cast bf16 back to
fp32 on resume.
- **restore `DFlashExportCallback`** (this PR's headline): the
Pydantic-recipe refactor (7038dec918) dropped the callback that exported
the draft submodule after each checkpoint save, leaving a stale "export
happens during training via DFlashExportCallback" comment with no
callback. FSDP2 SHARDED_STATE_DICT checkpoints carry no
`model.safetensors`, so without it there is nothing for vLLM /
acceptance-length eval to load. The callback gathers only the ~328 MB
draft submodule across shards via `get_model_state_dict(...,
submodules={dflash_module}, full_state_dict=True, cpu_offload=True)` —
works under SHARDED_STATE_DICT without materializing the 229B base — and
writes `exported-checkpoint-{step}/`.

## Testing
- Resume from FSDP2 sharded checkpoints verified end-to-end (loss/AR
continuity).
- Draft export validated against vLLM: exported drafts load and produce
acceptance-length metrics on MT-Bench across the full checkpoint sweep.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Export draft-submodule weights to dedicated exported checkpoints
during training.
* FSDP2 buffer compatibility and DTensor-aware gradient clipping for
safer distributed loading/training.
  * Detect HF-format checkpoints for smarter resume/load behavior.
  * Auto-add and handle a mask special token for draft workflows.
  * vLLM: disable prefix caching to preserve full prompt hidden states.
  * Add CLI option for answer-only loss and save aligned loss masks.
  * Add a SPEED-Bench config for DFLASH/vLLM benchmarking.

* **Chores**
  * Robust package version fallback to avoid import failures.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-06-15 16:47:50 -07:00
Chenhan D. Yuandpensieve-intern agent 601401b134 [OMNIML-5025] cell_t0_d7 (#1738)
Draft PR opened by **pensieve-intern** for
[OMNIML-5025](https://jirasw.nvidia.com/browse/OMNIML-5025).

Stage `cell_t0_d7` of Epic `OMNIML-5022`. The agent ran from the SPEC on
the ticket description; review every change before marking ready.

_Always-draft is enforced — the bot never auto-merges._

---

**Agent's self-narration** (stripped from PR diff; surfaced here for
context):

`INTERN_ARTIFACTS.json`:
```
{
  "sweep_name": "gemma-4-E4B-it_mtp_vllm_t0_d7",
  "experiment_id": "cicd_1781548540",
  "experiment_dir": "/lustre/fsw/portfolios/coreai/projects/coreai_dlalgo_modelopt/cicd/cicd/cicd_1781548540/",
  "AL_qualitative_overall": 3.2945,
  "AL_qualitative_categories": {
    "coding": 4.8883,
    "humanities": 2.703,
    "math": 4.1139,
    "multilingual": 4.1401,
    "qa": 2.59,
    "rag": 3.9329,
    "reasoning": 3.6251,
    "roleplay": 1.8026,
    "stem": 3.2041,
    "summarization": 2.8275,
    "writing": 2.4118
  },
  "AL_throughput_32k_overall": 3.3803,
  "AL_throughput_32k_categories": {
    "high_entropy": 2.0839,
    "low_entropy": 4.3797,
    "mixed": 3.7143
  }
}
```

`VERIFICATION_COMMENT.txt`:
```
Completed OMNIML-5025 cell_t0_d7 for `google/gemma-4-E4B-it` / MTP / vLLM.

What was done:
- Authored/kept `tools/launcher/common/specdec_bench/_cells/gemma-4-E4B-it_mtp_vllm_t0_d7.yaml` for sweep `gemma-4-E4B-it_mtp_vllm_t0_d7`.
- Updated/kept `tools/launcher/examples/google/gemma-4-E4B-it/specdec_bench_mtp_vllm.yaml` with the Gemma-specific `vllm/vllm-openai:gemma` container after generic v0.22.1 failed startup.
- Verified HF Hub model endpoints for `google/gemma-4-E4B-it` and `google/gemma-4-E4B-it-assistant` returned HTTP 200.
- Wrote `INTERN_ARTIFACTS.json` with parsed AL metrics from the successful cluster run.

Metrics extracted:
- experiment_id: `cicd_1781548540`
- experiment_dir: `/lustre/fsw/portfolios/coreai/projects/coreai_dlalgo_modelopt/cicd/cicd/cicd_1781548540/`
- qualitative Average_AL: `3.2945`
- throughput_32k Average_AL: `3.3803`

PR status:
- PR opened: NONE — this runner prompt says not to commit, push, or create PRs because the runner handles that.

What's next:
- Engine should consume `INTERN_ARTIFACTS.json` and advance the downstream wrap-up/aggregation stage.

Trigger pipeline URL:
- `https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/pipelines/54841359`

Slurm job status:
- submission: completed successfully
- experiment_id: `cicd_1781548540`
- experiment_dir: `/lustre/fsw/portfolios/coreai/projects/coreai_dlalgo_modelopt/cicd/cicd/cicd_1781548540/`
```



_Pollution-strip removed `INTERN_ARTIFACTS.json`,
`INTERN_LEARNING_TICKET.md`, `VERIFICATION_COMMENT.txt` from this commit
(sidecar narration and/or incidental lockfile regeneration are never
part of the agent's intended deliverable)._

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Updated benchmark configuration to use an optimized container image
for Gemma 4 MTP speculative-decoding benchmarks.
* Revised documentation in the configuration to reflect the current
image and its capabilities.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: pensieve-intern agent <noreply@nvidia.com>
2026-06-15 20:09:36 +00:00
Chenhan D. Yu c661366352 launcher: Nemotron-120B specdec_bench — kill stale _cells/ reference (post-PR-#1564) (#1741)
## What

Comment-only change to the Nemotron-3-Super-120B-A12B-BF16 DFlash parent
YAML doc-comment header. Rewrites the example invocation from the
pre-PR-#1564 pattern (with `--runtime_params
common/specdec_bench/_cells/<sweep_name>.yaml`) to the current
post-#1564 pattern (CLI-flag-only overrides, no per-cell file).

No yaml content / config change — only the prose example in the header.

## Why

PR #1564 removed the `tools/launcher/common/specdec_bench/_cells/` and
`_runtime_params/` directories. Cell-specific knobs are now CLI
overrides at slurm-invoke time, not committed files. The cell SPEC in
pensieve-intern (`specdec_bench/specs/cell.md`) is explicit about this —
five times in prose.

But this parent YAML's doc-comment kept advertising the OLD pattern.
When pensieve-intern's agent runner scans `tools/launcher/examples/` for
a reference invocation (good practice — agents should imitate working
examples), it lands on this Nemotron parent and copies the stale
pattern. **Five recent agent dispatches on the gemma-4 Epic OMNIML-5022
(cells OMNIML-5024 / 5025 / 5026 / 5027) authored new
`_cells/<sweep>.yaml` files for this reason, despite the SPEC telling
them not to.**

Prose loses to a concrete checked-in counter-example. The Qwen3.5-4B
reference template that cell.md officially points at
(`tools/launcher/examples/Qwen/Qwen3.5-4B/specdec_bench_mtp_vllm.yaml`)
is clean and shows only the bare `uv run slurm.py --yaml ...` form. This
PR makes Nemotron-120B consistent with that template.

## How surfaced

Diagnosed 2026-06-15 on OMNIML-5025 cell_t0_d7 ([intern-agent job
341631795](https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/341631795)).
The cell agent authored `_cells/gemma-4-E4B-it_mtp_vllm_t0_d7.yaml` —
the engine's diff-shape classifier should have rejected it, but didn't
(tracked separately as OMNIML-5170). Root-causing the agent's behavior
surfaced this stale doc-comment as the source of the pattern.

## Verification

`grep -rn '_cells\|runtime_params common'` across the entire launcher
tree returned only this file. After this PR, the launcher tree carries
zero stale references.

Pairs with NVIDIA/Model-Optimizer#1738 (gemma-4-E4B-it container fix)
and OMNIML-5170 (engine-side classifier hard-reject for `_cells/` paths
— defense in depth).

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated example documentation to clarify how to override per-cell
parameters using CLI flags in Slurm configuration runs.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-06-15 12:58:14 -07:00
Keval MorabiaandClaude Opus 4.8 55a2101e2f Update Nemotron-3 Pruning, Distillation and PTQ results based on new shared calibration loop with seq packing and add tool-calling eval fix (#1660)
### What does this PR do?

Type of change: documentation + minor example-script tweaks

Follow-up to #1601. Originally scoped to add **NVFP4 + QAD**, this PR
was **repurposed** to refresh the [Nemotron-3-Nano-30B-A3B-BF16
tutorial](examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md)
results using the **new shared calibration loop (sequence packing)** and
to **fix tool calling in evaluation**.

- Refreshed the prune → distill → eval → **FP8** results (accuracy +
vLLM throughput tables) with the new calibration loop.
- **Tool-calling eval fix** (`nemo_evaluator.yaml`): GPQA and AIME now
run the Python sandbox tool. The tutorial reports both **with-tools**
and **no-tools** GPQA/AIME and shows `mean ± std_dev`.
- Script tweaks: `quantize.py` calibration now uses sequence packing
(`pack=True`) which leads to slight improvement in PTQ;
`prune_minitron.py` defaults `inference_batch_size` to
`calib_batch_size`.

### Testing

Documentation + small example-script changes; tutorial relative links
resolve and the results tables / figure were verified consistent.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: Yes
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ (will run `/claude review`)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated the main guide and evaluator instructions for prune + distill
+ FP8/NVFP4 quantization, including refreshed vLLM deployment tips,
benchmark/noise presentation, and long-context tool-calling attribution
notes.
* Refreshed README technique examples/links, reordered the model support
matrix rows, and improved pruning overview/support-matrix text.

* **Changes to Examples**
* NAS pruning now documents higher GPU memory usage vs manual pruning;
pruning batching defaults were improved.
* Quantization PTQ calibration uses packed document packing; quantized
checkpoint export messaging was streamlined.
* Updated pruning/distillation/quantization tutorial guidance,
metrics/tables, command parameters, and evaluator YAML settings
(KV-cache dtype, generation defaults, task behavior).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 18:32:13 +00:00
Chenjie LuoandClaude Opus 4.8 d7df14d12a [tools/debugger] Enforce a single relay owner across hosts (#1735)
### What does this PR do?

Type of change: Bug fix

The `tools/debugger` file-based relay assumes a single server but never
enforced it.
Because the relay lives on shared NFS (the repo is often the same
checkout mounted on
multiple hosts), a forgotten `server.sh` on another host kept polling
the same
`.relay/` and could **steal commands** (executing them on the wrong
host), and killing
one server's cleanup could **wipe the active server's markers**.

This adds a `.relay/owner` ownership token (`host:pid:nanos`):

- Each server writes `owner` atomically at startup and **takes over**
instead of
refusing when a stale `server.ready` exists (the old `kill -0 <pid>`
guard was
  host-local and meaningless across hosts).
- The handshake and main loops exit cleanly if `owner` changes
(`[server] Superseded by <id> — exiting.`), so a freshly started server
**evicts**
  any stale one — even on another host.
- `cleanup()` only clears shared markers if we still own them, so a
stepping-down
  server never clobbers its successor's `server.ready`/`owner`.

Also gitignores `tools/debugger/logs/` and documents the `owner` file in
the README.

### Usage

```bash
# Inside the container; a previously-running server elsewhere that shares this
# NFS .relay/ steps down automatically once this one claims ownership:
bash tools/debugger/server.sh
# [server] Note: existing server.ready found (<host:pid:ts>); taking over.
#   (the stale server logs: "[server] Superseded by <id> — exiting.")
```

### Testing

Verified live on computelab: a forgotten `server.sh` on another host was
evicted when
a new server started, after which `client.sh run` executed on the
correct (new) host;
confirmed the stepping-down server's cleanup does not remove the
successor's
`server.ready`/`owner`. `server.sh` passes `bash -n`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ <!-- additive: new .relay/owner
file; client.sh and the wire protocol are unchanged -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- the file-based relay
tool has no test harness; behavior verified manually -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- internal dev tooling, not a shipped feature/API -->
- Did you get Claude approval on this PR?: N/A <!-- can run /claude
review -->

### Additional Information

Scope is limited to `tools/debugger/` (`server.sh`, `README.md`,
`.gitignore`).


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated relay protocol documentation to clarify ownership-based server
coordination.

* **Bug Fixes**
* Improved reliability of multi-server coordination in shared relay
environments.

* **Chores**
  * Updated ignore patterns for logging files.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 18:27:52 +00:00
Chenhan D. YuandClaude Opus 4.7 6968fe78fd feat(tools/mcp): resolve launcher dir via env override + cwd walk-up (#1731)
## Summary

Fixes Symptom A of
[OMNIML-5151](https://jirasw.nvidia.com/browse/OMNIML-5151): the
bridge's launcher-dir resolution doesn't work in `uv tool install`
layouts (which is how the intern-agent CI installs modelopt-mcp).
Empirically validated by nmm-sandbox pipelines
[54755376](https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/pipelines/54755376)
and
[54829964](https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/pipelines/54829964)
— the agent reached for `mcp__modelopt__submit_job`, got
`launcher_dir_not_found`, fell back to shell.

## Root cause

5 sites used:

```python
launcher_dir = _THIS_DIR.parent.parent / "launcher"
```

Works in:
- `pip install -e tools/mcp` (dev) — `_THIS_DIR` is
`<repo>/tools/mcp/modelopt_mcp/`
- Direct `Model-Optimizer` clone — same path

Doesn't work in:
- `uv tool install --from "git+...#subdirectory=tools/mcp" modelopt-mcp`
— `_THIS_DIR` is
`~/.local/share/uv/tools/modelopt-mcp/lib/.../site-packages/modelopt_mcp/`
and `parent.parent` doesn't reach a launcher

## How

New `_find_launcher_dir()` helper resolves in this order:

1. `$MODELOPT_LAUNCHER_DIR` env override (deterministic)
2. `_THIS_DIR.parent.parent / "launcher"` (in-repo layout)
3. Walk up from `os.getcwd()` looking for
`modules/Model-Optimizer/tools/launcher` (agent workspace) or
`tools/launcher` (direct clone)

Step 3 specifically unblocks intern-agent: the agent's cwd is inside its
cloned nmm-sandbox workspace where
`modules/Model-Optimizer/tools/launcher/` does exist — just needs the
walk-up.

Centralizes the structured-failure response too — five callsites had
slightly different `launcher_dir_not_found` shapes; new
`_launcher_dir_not_found_response()` helper produces a consistent dict
listing the searched paths so the next failure is obvious.

## Sites updated

- `submit_job_impl` — live submission
- `_submit_job_dry_run` — dry-run validation
- `_resolve_experiment_dir` — soft fallback (was crashing on
`None.exists()` before)
- `read_cluster_artifact_impl` — cwd for `nemo experiment logs`

## Tests

7 new hermetic tests in `tests/test_bridge.py`:

- env override wins
- env-points-at-ghost falls through gracefully
- walk-up via `modules/Model-Optimizer/tools/launcher`
- walk-up via plain `tools/launcher`
- returns `None` when nothing is found
- structured-failure response shape (with and without `dry_run` flag)

All 45 tests pass (38 + 7 new). Pre-commit clean (ruff / mypy / bandit /
license-headers).

## Out of scope

Symptom B of OMNIML-5151 — `NEMORUN_HOME` env not propagated to MCP
subprocess, affecting `job_status` / `wait_for_experiment` /
`read_cluster_artifact(path=None)`. That's a cross-repo change in
pensieve-harness's `mcp_config` writer + pensieve-modelopt's
`intern_runner`. Follow-up PR.

## Anchors

- [OMNIML-5151](https://jirasw.nvidia.com/browse/OMNIML-5151) — bug
ticket
- [OMNIML-5142](https://jirasw.nvidia.com/browse/OMNIML-5142) (install +
register), [OMNIML-5144](https://jirasw.nvidia.com/browse/OMNIML-5144)
(SPEC migration),
[OMNIML-5123](https://jirasw.nvidia.com/browse/OMNIML-5123) (Epic)
- nmm-sandbox MR !226 (CI install), !232 (PATH fix)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Improvements**
* Enhanced launcher directory discovery mechanism with support for
environment variable overrides and multiple search paths
* Improved error diagnostics when launcher directory cannot be resolved

* **Tests**
* Added comprehensive test coverage for launcher directory discovery and
failure scenarios

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-15 10:32:08 -07:00
Chenhan D. YuandClaude Opus 4.7 c4f39bd260 feat(tools/mcp): add dry_run flag to submit_job (#1718)
## Summary

`submit_job(yaml_path, ..., dry_run=True)` validates a launcher YAML via
`launch.py --dry-run` — exercises the YAML loader, factory resolution,
and arg parser, with no cluster contact / no container spawn / no
sbatch.

Used by verify-task workflow stages
(`pensieve-intern/workflows/spec_decoding_release/deployment_support.md`,
`hidden_state_dump_support.md`, `training_support.md`,
`synth_support.md`; `nemotron_quant_release/mlm_eval.md`, `mlm_ptq.md`,
`mlm_qad.md`) that today shell out to `uv run launch.py --yaml X
--dry-run` because the MCP catalog has no equivalent.

## Why

Without this,
[OMNIML-5144](https://jirasw.nvidia.com/browse/OMNIML-5144) (SPEC
migration) can't finish — the verify-task SPECs all stay on shell prose
for their YAML-validation step. Slice 1 of 5144 (specdec_bench's
`cell.md`, pensieve-intern MR !100) didn't need `dry_run` because
cell.md submits real cluster jobs, not validates configs. The other
workflows do.

## Tool surface

| Arg | Type | Notes |
|---|---|---|
| `dry_run` | `bool = False` | New. When True, skips cluster contact
entirely. |
| `hf_local` / `cluster_host` | `str?` | Now optional when
`dry_run=True` (pass one to validate executor-specific config, omit both
for shape-only validation). Still mutually exclusive in live submission.
|
| `skip_verify` | `bool = False` | Auto-bypassed when `dry_run=True`
(verify_setup is meaningless without cluster contact). |

Returns `{ok, dry_run: True, validated: bool, exit_code, stdout_tail,
stderr_tail, argv}` instead of `experiment_id`.

**Note on the ok/validated split:** when the launcher rejects the YAML
(exit code non-zero), the tool still returns `ok: True` because the TOOL
ran cleanly — only `validated: False`. Same shape `verify_setup` already
uses: `ok: True` regardless of probe outcome; the probe outcome is the
field the caller reads.

## Implementation

Clean fork at the top of `submit_job_impl` — the dry_run branch lives in
a dedicated `_submit_job_dry_run` helper to keep the live-submission
path uncluttered. ~120 lines added, no changes to existing behavior.

## Tests

4 new hermetic tests in `tests/test_bridge.py` (total: 38 passing):

* `test_submit_job_dry_run_yaml_validates` — happy path: ok+validated,
`--dry-run` in argv, no `--yes`
* `test_submit_job_dry_run_yaml_invalid` — launcher rejects YAML →
`ok=True, validated=False`, `stderr_tail` surfaces the error
* `test_submit_job_dry_run_yaml_not_found` — yaml_path missing →
`yaml_not_found` carrying `dry_run: True` for caller introspection
* `test_submit_job_dry_run_skips_verify` — bypasses `verify_setup` even
with `skip_verify=False`

## Scope policy

`dry_run` is universal launcher environment tooling, the
[`tools/mcp/SCOPE.md`](https://github.com/NVIDIA/Model-Optimizer/blob/main/tools/mcp/SCOPE.md)
test ("would this tool exist whether or not workflow X existed?") passes
trivially — every workflow that validates a YAML wants this. No
workload-specific knobs added.

## Anchors

* [OMNIML-5145](https://jirasw.nvidia.com/browse/OMNIML-5145) — this
ticket
* [OMNIML-5144](https://jirasw.nvidia.com/browse/OMNIML-5144) — parent
SPEC migration
* [OMNIML-5123](https://jirasw.nvidia.com/browse/OMNIML-5123) — Epic
* pensieve-intern MR !100 — slice 1 of the SPEC migration (the
live-submit path)

## Validation

* 38/38 unit tests pass (`uv run pytest tests/`)
* Pre-commit clean (ruff, ruff-format, mypy, bandit, markdownlint,
license-header)
* No changes to the existing live-submission path — `dry_run=False`
(default) is byte-for-byte equivalent to today's behavior

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a `dry_run` option to the `submit_job` tool to validate launcher
YAML and referenced artifacts without contacting the cluster, spawning
containers, or running `sbatch`.
* Dry-run responses now report validation status and include details
such as `exit_code`, stdout/stderr tails, and the launcher argv.

* **Documentation**
* Updated `submit_job` tool docs to describe the new `dry_run` behavior
and the dry-run return shape.

* **Tests**
* Added unit tests for successful validation, validation failures,
missing YAML, and ensuring setup verification is bypassed during
dry-run.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-14 13:40:20 -07:00
Chenhan D. YuandClaude Opus 4.7 463229ad34 docs(tools/mcp): scope policy — environment tooling, not workflow policy (#1712)
## Summary

Adds `tools/mcp/SCOPE.md` documenting the design boundary for the MCP
server family (`modelopt-mcp` + `nmm-sandbox-mcp` +
`pensieve-intern-mcp`):

- **In scope:** universal verb-shaped operations on the cluster,
launcher, agent engine — environment tooling that every workflow can
rely on as pre-knowledge.
- **Out of scope:** workflow-specific logic ("run an EAGLE3 cell",
"publish a specdec release"). That belongs in SPEC text + agent
reasoning, composed out of these primitives.

## Why

Surfaced during the OMNIML-5123 follow-up discussion. The 14 tools
currently in scope across the three servers are deliberately a small,
closed set. Letting workflow-specific tools sneak in would sprawl the
catalog, force per-workflow opt-in, and break the "MCP-as-pre-knowledge"
promise that makes the catalog useful as a stable baseline for
pensieve-intern's agents.

SCOPE.md documents:

- The test: would this tool be useful across *any* workflow that uses
the same environment?
- A side-by-side table of in-scope vs out-of-scope tool shapes
- Symptoms that a tool is misclassified
- Why the line matters (interface-drift blast radius collapses, agent
tool catalog stays learnable)
- Four practical questions to ask before adding a tool
- A reference list of the 14 currently-in-scope tools

## Anchor

[OMNIML-5123](https://jirasw.nvidia.com/browse/OMNIML-5123) (Epic).

Docs-only, no code changes.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added scope documentation clarifying that MCP environment tooling
covers universal verb-shaped operations (job submission, verification,
artifact reading) while workflow-specific policies belong elsewhere.
* Documented inclusion criteria and tool inventory guidance for MCP
servers.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-14 04:15:04 +00:00
Chenhan D. YuandClaude Opus 4.7 9f37fe1969 feat(tools/mcp): MCP server for ModelOpt launcher (OMNIML-5123) (#1701)
## Summary

`tools/mcp/` — a new MCP server exposing the existing
`tools/launcher/core.py` orchestration as **typed MCP tools** that codex
/ Claude Code agents can call directly, instead of shelling out to `uv
run launch.py --yaml ...` and parsing prose output.

Tracked under
[OMNIML-5123](https://jirasw.nvidia.com/browse/OMNIML-5123) (Epic).
Ships **Phase 1 + Phase 1.5** together: the core launcher surface plus
the four highest-leverage helpers from the `cell.md` simplification loop
([OMNIML-5128](https://jirasw.nvidia.com/browse/OMNIML-5128) partial,
[OMNIML-5132](https://jirasw.nvidia.com/browse/OMNIML-5132) full).

## Nine tools

**Phase 1 — core launcher surface:**

| Tool | Description |
|---|---|
| `list_examples` | Enumerate `tools/launcher/examples/` with model +
description metadata extracted from each YAML |
| `verify_setup` | Fail-fast probe for the named executor. Docker:
`docker info` (daemon up) + `docker info --format` runtime-registry
check for the `nvidia` runtime — no image pull, daemon-fast. Slurm: `ssh
-o BatchMode=yes -o ConnectTimeout=5` to the cluster login node. ~1 s
probe saves 30+ s of wasted submission on bad config |
| `submit_job` | Submit a launcher YAML. Mode is determined by
mutually-exclusive args: `hf_local` → Docker (local GPU), `cluster_host`
→ Slurm (remote SSH). Returns experiment_id immediately; the actual job
runs detached |
| `job_status` | Filesystem-based status from nemo_run's experiment dir
(`_DONE`, `status_*.out`) — no in-memory registry, survives MCP server
restarts |
| `job_logs` | Read `log_<task>.out` from experiment dir; per-task
filtering + optional tail |

**Phase 1.5 — `cell.md` simplification (OMNIML-5128 / 5132):**

| Tool | Description |
|---|---|
| `wait_for_experiment` | Replaces the agent's `while True: status;
sleep` poll loop with one tool call. Reuses `job_status_impl` so
terminal-state semantics stay identical. Returns final status plus
`waited_seconds`; on timeout returns structured `{ok: False, reason:
"wait_timeout", last_status: …}` |
| `provision_passwordless_ssh_dry_run` | No-side-effect inspection of
`~/.ssh/` that emits the exact `ssh-keygen` / `ssh-copy-id` commands the
operator should run to make `verify_setup(executor='slurm')` pass.
Closes the verify_setup "ssh_auth_failed → now what?" gap |
| `read_cluster_artifact` | Uses nemo_run's tunnel primitives, not a
reinvented SSH layer. `path=None` wraps `nemo experiment logs <id>
<job_idx>` (built-in log fetch); with a `path`, uses the experiment's
`Tunnel` to read the file. Structured failure on subprocess error /
timeout |
| `open_draft_pr` | `git push -u origin HEAD` + `gh pr create --draft
…`. Validates cwd is a git repo first; on gh failure after push
succeeds, reports `branch_pushed=True` so the operator can retry just
the PR-open step |

## Design constants

1. **Single `submit_job` with mode by args** (not separate
`submit_docker` / `submit_slurm` tools). Keeps the LLM tool catalog
compact; mutual-exclusion is a runtime check.
2. **Filesystem is the source of truth** for status + logs. No in-memory
registry. Survives MCP server restarts cleanly — important because
operators / agents kill + restart their hosts often.
3. **`verify_setup` is auto-called by `submit_job`** by default
(skippable when caller just probed). The probe is ~1 s; the cost of a
misconfigured submission is 30+ s of cluster timeout or container-pull.
Always-on verify pays back immediately.
4. **Delegate to nemo_run for tunnels.** `read_cluster_artifact` and
`wait_for_experiment` use nemo_run's existing `Experiment` / `Tunnel` /
`nemo experiment logs` primitives rather than reinventing SSH/rsync. One
source of truth for cluster I/O.

## Layout

```
tools/mcp/
├── pyproject.toml          # name: modelopt-mcp, console_script
├── modelopt_mcp/
│   ├── __init__.py
│   ├── server.py           # FastMCP entry; 9 tool definitions
│   └── bridge.py           # thin wrapper over launcher's core.py
│                           #   + filesystem status/log helpers
│                           #   + tunnel/PR helpers (Phase 1.5)
└── tests/
    └── test_bridge.py      # 34 unit tests, fully hermetic
                            # (mocked subprocess + tmp_path fixtures)
```

## Install

Two paths, both **from source via uv**. No PyPI wheel exists;
OMNIML-5123 opted for the uvx-from-git pattern to skip publication
overhead.

### End-user install (recommended)

`uvx` from the git subdirectory — single command, no manual clone:

```bash
# Claude Code
claude mcp add modelopt -- uvx --from \
  "git+https://github.com/NVIDIA/Model-Optimizer.git#subdirectory=tools/mcp" \
  modelopt-mcp

# Codex
codex mcp add modelopt -- uvx --from \
  "git+https://github.com/NVIDIA/Model-Optimizer.git#subdirectory=tools/mcp" \
  modelopt-mcp
```

Under the hood `uvx` clones the whole repo to its cache, installs
`tools/mcp/` as the entry, and resolves the sibling `modelopt-launcher`
dep via `[tool.uv.sources]` (`path = "../launcher"`) inside the cloned
tree.

### Dev install (local checkout)

```bash
uv pip install -e tools/launcher    # sibling dep first
uv pip install -e tools/mcp         # then this package
modelopt-mcp                         # entry on PATH
```

### Why no plain `pip install` today

Two specific reasons, worth flagging so reviewers know what's
intentional vs missing:

1. **Nothing on PyPI yet.** Neither `modelopt-mcp` nor
`modelopt-launcher` are published — this PR introduces the package but
doesn't add release machinery.
2. **`pip` doesn't read `[tool.uv.sources]`.** Even from a local
checkout, plain `pip install -e tools/mcp` fails because
`modelopt-launcher` is a bare name (no URL) and pip can't find it.
Sticking with `uv` / `uvx` is the practical path while we're git-only.

If we later want plain-pip support: publish to PyPI, or switch to a
PEP-440 direct URL (`"modelopt-launcher @
git+…#subdirectory=tools/launcher"`). Out of scope for this PR.

## Post-review changes

Addressed all CodeRabbit + claude[bot] review findings on the original
Phase-1 surface. See the inline replies for details; the substantive
bug-fix highlights:

* **Slurm `cluster_host`** — propagate via `env=child_env` (launch.py
reads SLURM_HOST, not a CLI arg)
* **`shlex.quote`** removed from nemo-run k=v overrides (subprocess
list-form doesn't shell-quote)
* **Docker `Popen`** now uses `stdout=DEVNULL, stderr=DEVNULL,
start_new_session=True` to avoid pipe-buffer blocking
* **`NEMORUN_HOME`** pinned in subprocess env so submit + status sides
agree
* **GPU verify** swapped from `docker run --gpus all` image-pull (slow +
flaky) to `docker info --format` runtime-registry check (daemon-fast)
* **Task-status word match** anchors on first word against a fixed
failure-word set (no more `"fail" in "succeeded after retry; previous
attempt failed"` false-positive)
* **`experiment_id` regex** generalized for non-NVIDIA cluster paths
* **`pyproject.toml`** dropped the unsatisfiable `modelopt-launcher`
bare-name dep (launcher is a file-layout sibling, not a Python import
dep)
* **`Field(ge=1)`** on `job_logs.tail`
* **Docstring contract** clarified (Docker returns `pid`, Slurm returns
`experiment_id`)

## Validation

- [x] `uv pip install -e .` succeeds (modelopt-launcher resolved
transitively)
- [x] 34/34 unit tests pass (`uv run python -m pytest tests/`)
- [x] stdio handshake works end-to-end; `tools/list` returns all 9 with
full schemas + descriptions
- [x] Mode-resolution: `submit_job` correctly rejects no-executor +
both-executors with structured `reason`
- [x] Filesystem status: correctly classifies `done` / `failed` /
`running` from `_DONE` + `status_*.out`
- [x] `wait_for_experiment` short-circuits on already-terminal
experiments; honors timeout without raising
- [x] `provision_passwordless_ssh_dry_run` distinguishes no-key /
key-only / key+pubkey cases
- [x] `read_cluster_artifact` handles subprocess timeout + non-zero exit
with structured reasons
- [x] `open_draft_pr` reports `branch_pushed=True` on
gh-failure-after-push so retries are cheap
- [x] Pre-commit clean: ruff, ruff-format, mypy, bandit, license-headers

## Acceptance criteria

**OMNIML-5123 (Phase 1):**

- [x] `list_examples` returns all bundled YAMLs with path and model name
- [x] `submit_job` with `hf_local` runs via Docker executor and returns
immediately (Phase 1: returns PID; experiment_id capture in Phase 2)
- [x] `submit_job` with `cluster_host`/`user` runs via Slurm executor
(`detach=True`) and returns experiment_id
- [x] `job_status` correctly reflects running / done / failed from
nemo_run filesystem
- [x] `job_logs` returns stdout for a completed job
- [x] `uvx --from git+...#subdirectory=tools/mcp modelopt-mcp --help`
resolves and starts
- [x] Existing launcher tests unaffected (no changes to
`tools/launcher/`)

**OMNIML-5128 (Phase 1.5, partial):**

- [x] `wait_for_experiment` blocks until terminal or timeout
- [x] `read_cluster_artifact` pulls remote artifacts via nemo_run tunnel
- [x] `open_draft_pr` opens a draft PR against a target repo
- [ ] Capture `experiment_id` from Docker subprocess output — deferred
to Phase 2

**OMNIML-5132 (Phase 1.5, full):**

- [x] `provision_passwordless_ssh_dry_run` emits operator-facing
commands without side effects

## Phase 2 (separate PR)

* Capture `experiment_id` from Docker subprocess output (tail until
nemo_run logs the id).
* Extract the verify + submit helpers into a shared lib that
[`nmm-sandbox-mcp`](https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/tree/main/tools/mcp)
(companion server, separate repo) can consume for internal-ergonomics
tools — cluster short-name → factory lookup + GitLab CI dispatch.
* NEL integration
([OMNIML-5133](https://jirasw.nvidia.com/browse/OMNIML-5133)) +
checkpoint introspection
([OMNIML-5134](https://jirasw.nvidia.com/browse/OMNIML-5134)).


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added ModelOpt MCP server and console entrypoint; tools:
list_examples, verify_setup, submit_job, job_status, job_logs,
wait_for_experiment, provision_passwordless_ssh_dry_run,
read_cluster_artifact, open_draft_pr; Docker and Slurm support.

* **Documentation**
* Expanded README with install steps, design notes, end-to-end agent
example, roadmap, and repo layout.

* **Tests**
* Expanded unit tests covering bridge helpers, polling, SSH flows,
artifact reads, and PR automation.

* **Chores**
  * CI updated to run MCP tests; package/meta config added.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-12 18:56:18 -07:00
h-guo18 cfc823d127 [Tests]: Precommit Check for Spec-Dec Recipes (#1527)
### What does this PR do?

Type of change: new tests / tooling

Adds pre-commit validation for speculative-decoding recipes (the
existing `check-modelopt-recipes` hook only ran on PTQ) and for launcher
YAML references into the recipe library.

- `tools/precommit/check_modelopt_recipes.py`: accept
`speculative_eagle` / `speculative_dflash` / `speculative_medusa` in
addition to `ptq`, so per-model spec-dec recipes (e.g.
`modelopt_recipes/models/Qwen3-8B/dflash.yaml`) get full Pydantic
validation via `load_recipe()` at commit time.
- `tools/precommit/check_launcher_yaml.py` (new): scans every
`tools/launcher/examples/**/*.yaml` for `--config <path>` and
`data.chat_template=<path>` references, verifies the resolved files
exist, and runs `load_recipe()` on any path under `modelopt_recipes/`.
Skips `<<global_vars.x>>` interpolation. `pass_filenames: false` so
recipe-side edits also re-validate all launcher references.

### Usage

```bash
pre-commit run check-modelopt-recipes --all-files
pre-commit run check-launcher-yaml --all-files
```

### Testing

Smoke-tested both hooks manually:

| Scenario | Result |
|---|---|
| spec-dec recipe with `dflash_block_size: not_an_int` | exit 1,
Pydantic int_parsing error |
| launcher YAML with non-existent `--config` path | exit 1, source file
+ resolved path reported |
| launcher YAML with non-existent `data.chat_template` path | exit 1 |
| Current repo state (all valid) | exit 0 |

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ (hooks themselves are the
tests)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ pending

### Additional Information

Motivated by the per-model recipe migration in #TBD — without these
hooks, broken `--config` paths and recipe schema typos surface only at
CI or runtime.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Added an automated pre-commit check that validates launcher example
YAMLs, reporting parse errors and missing or invalid references.
* Expanded recipe validation to cover additional recipe types beyond
PTQ, improving detection of invalid recipe formats and metadata.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-06-11 17:40:15 -07:00
Chenhan D. Yu cb5be1f92b launcher: move gemma-4-E4B-it yaml to google/ subdir (#1689)
## Summary

- Moves
`tools/launcher/examples/gemma-4/gemma-4-E4B-it/specdec_bench_mtp_vllm.yaml`
to
`tools/launcher/examples/google/gemma-4-E4B-it/specdec_bench_mtp_vllm.yaml`
to align with the `google/` vendor-prefix convention used by other
models in the launcher examples directory.
- Updates the embedded `--yaml` path in the YAML comment to match the
new location.

## Test plan
- [ ] Confirm the file appears at the new path
- [ ] Confirm the old path is gone

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Updated an example benchmark invocation reference to reflect the
project’s reorganized directory layout. This edits only the example
comment in the benchmark YAML; no benchmark settings, model paths, task
arguments, or runtime configuration were modified. Low-risk
documentation alignment.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan D. Yu <5185878+ChenhanYu@users.noreply.github.com>
2026-06-11 23:19:15 +00:00
h-guo18 46eddab877 [Feat]: Specdec Streaming: RDMA + Multinode (#1611)
### What does this PR do?

Type of change: New feature

Multi-node **streaming** training for speculative decoding (EAGLE3 /
DFlash):
a live `vllm serve` captures the target model's hidden states and moves
them
straight to the trainer over **NIXL RDMA** — no disk round-trip. The
streaming
dataset is map-style — each rank fetches only its own
`DistributedSampler` shard
(concurrency from `dataloader_num_workers`), round-robins across
multiple serve
replicas (`server_urls`), and scales to multi-node DDP. Serve-side
tensor
parallelism (TP>1) is supported: hidden states are replicated across TP
ranks, so
rank 0 alone owns the pool + transfer.

### How

- `RdmaHiddenStatesConnector` — out-of-tree vLLM connector (no vLLM
source edits):
one pre-registered pinned NIXL pool per serve, a ring slot per request,
and a
small HTTP sidecar serving transfer metadata. The trainer RDMA-READs the
slot
  into a per-worker buffer. RDMA is the **only** transport (the earlier
  disk/safetensors path is removed).
- Map-style dataset + multi-node accelerate launch (`--machine_rank`,
optional
  Slurm `--segment` to keep nodes in one NVLink domain).

### Usage

```yaml
data:
  mode: streaming
  streaming_server_url: "http://node0:8000,http://node1:8000"  # round-robin
```

### Validation (Qwen3-8B, oci-nrt H100)
sandbox CI:
https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/337489812

**1. End-to-end convergence — EAGLE3 & DFlash, 5000 steps.** Both
algorithms
converge and export a deployable draft; the DFlash drafts also serve
under vLLM
speculative decoding (8/8 smoke prompts pass).

| algorithm | topology (nodes) | train loss (step 0 → 5000) | vLLM draft
acc-len |
|---|---|---|---|
| EAGLE3 | 2 serve TP=2 + 2 trainer DDP (4) | 37.1 → 8.20 | — |
| DFlash | 1 serve TP=1 + 1 trainer (2)      | 11.7 → 5.56 | 1.11 |
| DFlash | 2 serve TP=2 + 2 trainer DDP (4)  | 10.9 → 5.26 | 1.19 |

<!-- Drag these PNGs in here (GitHub turns them into asset URLs):
eagle3_streaming_loss.png, dflash_streaming_loss_singlenode.png,
dflash_streaming_loss_multinode.png -->

**2. Scalability — 1 → 12 nodes (EAGLE3, 200 steps).** Throughput scales
~23× across the
sweep below. The step-time growth is cross-node DDP all-reduce, not the
streaming path —
RDMA (~0.33 ms/req @ 2 MB, ~47 GB/s host-pinned READ) is never the
bottleneck. Scale
serve + trainer nodes together for near-linear speedup.

| serve / trainer | nodes | step time | samples / step | samples / sec
(global) | acc @ step 200 |
|---|---|---|---|---|---|
| 1 serve / 1 rank (co-located, 1 node 2 GPU) | 1 | 0.23 s | 1 | 4.4 |
[0.141, 0.094, 0.072] |
| 1 serve / 1 rank (cross-node) | 2 | 0.23 s | 1 | 4.3 | [0.137, 0.105,
0.074] |
| 2 serve / 8 ranks | 3 | 0.26 s | 8 | 31.1 | [0.215, 0.126, 0.097] |
| 4 serve / 16 ranks (2 trainer nodes) | 6 | 0.28 s | 16 | 56.5 |
[0.217, 0.148, 0.110] |
| 8 serve / 32 ranks (4 trainer nodes) | 12 | 0.31 s | 32 | 101.9 |
[0.235, 0.165, 0.137] |

**3. Serve-side TP correctness.** TP=1 vs TP=2 draft top-1 accuracy
track
step-for-step (hidden states are replicated across TP ranks).

<img width="910" height="546" alt="serve-tp-acc"
src="https://github.com/user-attachments/assets/73df9214-7ff0-4ab4-bf2f-95842b12cd5f"
/>

### Before your PR is "*Ready for review*"

- Backward compatible?: ❌ — streaming is now RDMA-only; `server_url` →
`server_urls`;
the disk transport (`HS_TRANSPORT`, `streaming_shared_storage_path`) is
removed.
- New tests?: ✅
`tests/unit/torch/speculative/plugins/test_hf_streaming_dataset.py`
  (map-style dataset + mocked RDMA fetch).
- Updated Changelog?: ❌
- Claude approval?: ❌

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-06-10 18:35:02 -07:00
Chenhan D. Yu 13e540f0b3 launcher: add draft_model to GlobalVariables for spec-decode parents (#1675)
SPEED-bench MTP / EAGLE3 / DRAFT_TARGET / DFLASH parent YAMLs reference the speculative draft/assistant model path on BOTH the qualitative and throughput_32k tasks via `<<global_vars.draft_model>>` indirection so the path lives in one place. The launcher's `GlobalVariables` dataclass had a strict whitelist of keys and rejected `draft_model` with `ValueError: No parameter named 'draft_model' exists`.

Surfaced on OMNIML-5024 pipeline #54356795: the gemma-4-E4B-it / MTP / vLLM parent used the indirection but failed at launcher load. The agent worked around it inline. This adds `draft_model: str = None` as the fifth top-level field on `GlobalVariables`.

The existing `<<global_vars.X>>` resolution in `SandboxPipeline.__post_init__` (uses `dataclasses.asdict`) picks the new field up automatically.

Backward-compatible: existing parents that don't use `draft_model` keep working.

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-06-10 16:09:24 -07:00
sychen52 808e8b3320 Make .claude/skills a real folder for claude sandbox to work (#1674)
### What does this PR do?

Type of change: ?  Bug fix

replace .claude/skills dir-symlink with real dir + per-skill symlinks

Claude Code hardcodes .claude/skills in its sandbox denyWithinAllow
list. bwrap enforces this by creating a mount point at that path, but
fails with "Can't create file: Is a directory" when the path is a
symlink to a directory — breaking all sandboxed commands, not just
writes.

Fix by making .claude/skills a real directory containing per-skill
symlinks into .agents/skills/. Add a pre-commit hook that automatically
creates and stages a new symlink whenever a skill directory is added to
.agents/skills/, so authors need no extra steps.

### Usage

claude with sandbox

### Testing
tried on local.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Updated internal development tooling to streamline workflow automation
and improve developer processes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
2026-06-10 15:00:48 -07:00
Chenhan D. Yu 06af68cbdb [OMNIML-4962] specdec_bench parent — Qwen/Qwen3.5-4B / DFlash / vLLM (#1638)
## Summary\n- Add specdec_bench DFlash vLLM cell t0_d3 for
Qwen3.5-4B\n\n## Testing\n- Not run (cell YAML only)\n

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Added new benchmark configuration for Qwen3.5-4B model performance
evaluation with speculative decoding optimization across qualitative and
throughput testing scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: chenhany <chenhany@nvidia.com>
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-06-10 21:41:00 +00:00
Chenhan D. Yu 66b54ed040 [OMNIML-5024] specdec_bench cell t0_d3 — google/gemma-4-E4B-it / MTP / vllm (#1663)
### What does this PR do?

Type of change: Bug fix + new example

Wires SPEED-bench's MTP path to support **Gemma 4** (and any future MTP
variant that uses a separate assistant / draft model), and adds the
SPEED-bench MTP/vLLM example for `google/gemma-4-E4B-it`.

**Key difference: Gemma 4 MTP vs. generic MTP.** vLLM's
`speculative_config` accepts two different shapes for MTP:

| Variant | `speculative_config` shape | Models |
|---|---|---|
| **Generic MTP** | `{"method": "mtp", "num_speculative_tokens": N}` |
Models that carry their own MTP layer in-tree (e.g. Qwen 3.5 MTP
variants) — no separate draft / assistant model. |
| **Assistant-model MTP** | `{"model": "<assistant>",
"num_speculative_tokens": N}` (no `method` key — vLLM auto-detects from
the assistant) | Gemma 4 family (E2B / E4B / 26B-A4B / 31B); each target
model has a paired `<target>-assistant` checkpoint that acts as the MTP
draft. Landed in
[vllm-project/vllm#41745](https://github.com/vllm-project/vllm/pull/41745)
(2026-05-06). |

The specdec_bench vLLM wrapper at
`examples/specdec_bench/specdec_bench/models/vllm.py` previously emitted
only the generic shape for any `--speculative_algorithm MTP` invocation,
which produced `NotImplementedError: Unsupported speculative method:
'mtp'` on Gemma 4 even with a container that has the support
(`vllm/vllm-openai:v0.22.1`+). This PR teaches the wrapper to switch
shapes based on whether `--draft_model_dir` is provided.

**Concrete changes:**

1. **`examples/specdec_bench/specdec_bench/models/vllm.py`** — when
`speculative_algorithm == "MTP"` AND `draft_model_dir` is set, emit
`{"model": draft_model_dir, "num_speculative_tokens": N}`
(assistant-model shape). Otherwise emit the existing `{"method": "mtp",
...}` (generic shape). Backward-compatible — Qwen 3.5 MTP and other
callers that omit `--draft_model_dir` get the same config they got
before.

2. **`examples/specdec_bench/specdec_bench/utils.py`** — `get_tokenizer`
reads `extra_special_tokens` from the model's `tokenizer_config.json`
and passes them through to `AutoTokenizer.from_pretrained`. Gemma 4
tokenizers ship a list-shaped `extra_special_tokens` entry that the
constructor would otherwise reject. Necessary for any Gemma 4 cell.

3.
**`tools/launcher/examples/gemma-4/gemma-4-E4B-it/specdec_bench_mtp_vllm.yaml`**
— SPEED-bench parent YAML for `google/gemma-4-E4B-it`. Uses
`vllm/vllm-openai:v0.22.1` (has `gemma4_mtp.py` from #41745) and wires
`--draft_model_dir /hf-local/google/gemma-4-E4B-it-assistant` on both
task_0 (qualitative) and task_1 (throughput_32k).

4.
**`tools/launcher/common/specdec_bench/_cells/gemma-4-E4B-it_mtp_vllm_t0_d3.yaml`**
— runtime params for the `t0_d3` cell of OMNIML-5022 (`temperature=0`,
`max_model_len=40960`).

### Usage

```python
# Wrapper-level: same CLI as before, just pass --draft_model_dir for
# Gemma 4 MTP. The wrapper auto-routes to the assistant-model shape.

# python examples/specdec_bench/run.py \
#     --engine VLLM \
#     --speculative_algorithm MTP \
#     --draft_model_dir /hf-local/google/gemma-4-E4B-it-assistant \
#     --draft_length 3 \
#     --tp_size 1 \
#     ...other SPEED-bench knobs...

# Equivalent direct vLLM invocation (for reference, no wrapper):
from vllm import LLM, SamplingParams
llm = LLM(
    model="google/gemma-4-E4B-it",
    speculative_config={
        "model": "google/gemma-4-E4B-it-assistant",
        "num_speculative_tokens": 3,
    },
    trust_remote_code=True,
)
```

### Testing

- **Upstream existence checks**: verified the assistant models
`google/gemma-4-{E2B,E4B,26B-A4B}-it-assistant` exist, public, ungated
on HuggingFace; verified `vllm/model_executor/models/gemma4_mtp.py` is
in vLLM `v0.22.0`, `v0.22.1`, and `main`.
- **Backward compat**: `MTP` callers that don't pass `--draft_model_dir`
(e.g. the existing Qwen 3.5 MTP/vLLM cells under
`tools/launcher/examples/Qwen/Qwen3.5-4B/`) take the unchanged
`{"method": "mtp", ...}` branch. No diff for those.
- **End-to-end cluster validation**: pending. Will run via the
OMNIML-5022 cells (OMNIML-5024 / 5025 / 5026 / 5027) once the
nmm-sandbox submodule pin advances past this PR. Each cell exercises
`task_0` (SPEED-Bench qualitative, 880 samples) + `task_1`
(throughput_32k, 80 samples) on cw_dfw, single H100.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ — the wrapper only takes the
new branch when `--draft_model_dir` is provided alongside
`--speculative_algorithm MTP`. Existing MTP callers (Qwen 3.5 etc.) keep
the generic `method: "mtp"` config.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependencies.
- Did you write any new necessary tests?: ❌ — relying on the SPEED-bench
cluster cells (OMNIML-5024 …5027) for end-to-end validation; no unit
test fixture for the vLLM wrapper exists in `tests/` for me to extend
symmetrically. Happy to add one if reviewers want it.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — small fix + example addition. Can add if requested.
- Did you get Claude approval on this PR?: ❌ — will run `/claude review`
once the PR is marked Ready for review.

### Additional Information

- JIRA: [OMNIML-5024](https://jirasw.nvidia.com/browse/OMNIML-5024)
(cell_t0_d3); siblings OMNIML-5025/5026/5027 (cell_{t0_d7, t1_d3,
t1_d7}) of Epic OMNIML-5022 are blocked on this PR landing.
- Upstream reference: vllm-project/vllm#41745 — "[Spec Decode] Add
Gemma4 MTP speculative decoding support".
- Companion (pensieve-intern !91, internal): adds a
"Model-family-specific MTP invocation" table to the specdec_bench cell
SPEC so future agents pair `MTP` with the right `--draft_model_dir` from
SPEC-read time.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a SPEED-bench pipeline for Gemma 4 using vLLM speculative
decoding (MTP) with qualitative and throughput tasks.

* **Improvements**
* Speculative-decoding logic updated to handle assistant-model and
generic MTP cases distinctly.
* Tokenizer loading now reads and normalizes extra special tokens from
tokenizer config when available.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Pensieve Intern <chenhany@nvidia.com>
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-06-10 12:46:13 -07:00
Chenhan D. Yu 111b7ebd29 [OMNIML-4969] specdec_bench cell t0_d3 — nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 / dflash / vllm (#1656)
## Summary\n- Add specdec_bench cell t0_d3 for
NVIDIA-Nemotron-3-Super-120B-A12B-BF16 dflash/vLLM\n\n## Testing\n- Not
run (cell YAML only)\n

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a new benchmark configuration to run Nemotron-3-Super-120B on
NVIDIA with vLLM using the DFlash speculative-decoding method.
* Provides both qualitative and high-throughput benchmarking scenarios,
tunable concurrency and request counts, and automated job execution with
Slurm-compatible orchestration and output saving.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: chenhany <chenhany@nvidia.com>
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Signed-off-by: Chenhan D. Yu <5185878+ChenhanYu@users.noreply.github.com>
2026-06-09 23:59:50 +00:00
Chenhan D. Yu 2c52e7bf4e [OMNIML-4788] specdec_bench/Qwen3.5-4B: throughput_32k benchmark + S3 upload step (#1564)
### What does this PR do?

Type of change: enhancement (follow-up to
[#1531](https://github.com/NVIDIA/Model-Optimizer/pull/1531)).

Extends the merged Qwen3.5-4B SPEED-Bench launcher YAMLs from a
single-task qualitative-only smoke into a **3-task pipeline** that also
covers long-context throughput and verifies the S3-upload path
end-to-end. Two commits, cleanly cherry-picked from #1531's late branch
state — they were authored after the merge-commit was resolved against
an earlier rebased head and so didn't ride along with that merge.

### Pipeline shape (both YAMLs)

| Task | Split | Save dir |
|---|---|---|
| `task_0` | qualitative (existing quality / acceptance-rate signal) |
`/scratchspace/specdec_bench{,_mtp}/qualitative` |
| `task_1` | **throughput_32k** (new — long-context throughput) |
`/scratchspace/specdec_bench{,_mtp}/throughput_32k` |
| `task_2` | **upload to S3 in sweep layout** |
`s3://team-specdec-workgroup/results/specdec_bench{,_mtp}/<split>/` |

### New artifacts

* `tools/launcher/common/specdec_bench/upload_to_s3.sh` — thin wrapper
around `examples/specdec_bench/upload_to_s3.py` so it can be invoked as
a launcher task. Installs `boto3` from `requirements.txt` on cold
containers; warm pipelines pick it up from the prior `run.sh`.
*
`tools/launcher/common/specdec_bench/runtime_params_throughput_32k.yaml`
— pins `engine_args.max_model_len = 40,960` (32K input + 4K output + 4K
headroom) so vLLM doesn't silently auto-cap `max_model_len` below the
36K minimum needed for `throughput_32k` prompts on single-GPU runs.

### Why max_model_len matters

Without an explicit `max_model_len`, vLLM auto-derives it from the model
config (Qwen3.5-4B = 128K) **and from the GPU-memory budget**. On a
single GPU the second factor can cap effective `max_model_len` well
below 36K, silently truncating 32K-token prompts and producing wrong
throughput numbers. The qualitative split is not affected (its prompts
top out around 8K, well under any auto-derivation floor) so only
`task_1` carries the override.

### S3 credentials

`upload_to_s3.sh` reads `S3_ENDPOINT` / `S3_KEY_ID` / `S3_SECRET` from
the runtime environment (not hardcoded). `--skip-existing` +
`--allow-incomplete-provenance` are passed by default so re-runs land
alongside the prior upload, and runs lacking `CONTAINER_IMAGE` (Phase-2
harness work in OMNIML-4788 will populate it) still upload.

### Testing

Cluster smoke on cw_dfw via:

```
uv run slurm.py --yaml modules/Model-Optimizer/tools/launcher/examples/Qwen/Qwen3.5-4B/specdec_bench.yaml --yes
```

is currently in-flight (jobs `12257378/79/80`, PD). Will update this PR
with timing/AR numbers + S3 upload confirmation once it lands.

### Before your PR is "Ready for review"

- Backward compatible: ✅ (additive — task_0 keeps the prior qualitative
behavior, just with `/qualitative` suffix in `save_dir`)
- New PIP dep: ✅ no (boto3 already in
`examples/specdec_bench/requirements.txt` from #1531)
- New tests: N/A (launcher YAML + shell wrapper; covered by cluster
smoke)
- Changelog: N/A (internal-facing tooling)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a 32K-context runtime configuration (higher max model length) to
enable long-context throughput benchmarking and avoid silent prompt
truncation.
* Added a launcher helper to upload benchmark results to S3 with
incremental/retry-friendly options and pass/fail reporting.

* **Chores**
* Split Qwen3.5-4B benchmark into separate qualitative and 32K
throughput tasks and added coordinated S3 upload.
* Applied the same multi-task pipeline layout and clearer output
organization to the MTP speculative-decoding benchmark.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1564?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: chenhany <chenhany@nvidia.com>
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-06-07 03:04:34 +00:00
Jenny Chen 6b73e933bf Nemotron Ultra & Super launcher examples (#1609)
### What does this PR do?

Type of change: New example

New launcher example for Nemotron Super with PTQ + Export + VLLM smoke
test on small GPQA-style dataset

### Usage

```python
# Usage:
#   source .env-slurm
#   cd tools/launcher
#   uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/megatron_lm_ptq.yaml --yes
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added checkpoint export capability for quantized models to Hugging
Face format.
* Introduced complete quantization pipelines with conditional MMLU
evaluation and model export stages.

* **Bug Fixes**
  * Fixed num_shards calculation to prevent invalid minimum values.

* **Documentation**
* Updated vLLM version requirements for optimal NVFP4 model performance.
* Enhanced quantization pipeline documentation with improved output
paths and conditional execution details.

* **Chores**
  * Updated Megatron-LM module to latest version.
  * Added sample dataset for model evaluation testing.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
2026-06-04 17:46:58 +00:00
yeyu-nvidiaandClaude Opus 4.6 5bd04c3876 Revert unverified EAGLE3 model examples; keep triage code + baseline (#1623)
### What does this PR do?

Type of change: Revert / cleanup (follow-up to #1417)

Per review feedback (@h-guo18): `main` should be production-ready and
user-facing. Most of the EAGLE3 model example YAMLs added in #1417 are
not yet verified to work end-to-end in modelopt (~80% fail at some
pipeline stage), which is confusing to ship. The agreed plan is to
**land the triage infrastructure now and re-add each model's launcher
YAML in a dedicated follow-up PR once it is verified green**.

**Removed** (unverified, to be re-added per-model once verified):
- Per-model launcher configs (`hf_offline_eagle3.yaml` +
`eagle3_quick_check.yaml`) for: DeepSeek-V3.2, GLM-5, MiniMax-M2.5,
Ministral-3-8B, Ministral-3-14B, Kimi-K2.5, Kimi-K2.5-NVFP4,
GPT-OSS-20B, Qwen3.5-9B, Qwen3.5-27B, Qwen3.5-35B-A3B, Step-3.5-Flash.
- Per-model status docs: `tools/launcher/examples/EAGLE3_TRIAGE.md`,
`examples/speculative_decoding/pipeline/eagle3/eagle3_triage_chart.md`
(volatile status — tracked internally instead).

**Kept** (the durable triage infrastructure from #1417):
- Verified baseline example
`tools/launcher/examples/Qwen/Qwen3-8B/eagle3_quick_check.yaml`.
- Launcher common scripts (vLLM native-extractor dump, etc.) and
`compute_hidden_states_vllm.py`.
- modelopt code fixes: FakeBaseModel VLM detection,
`consolidated.safetensors` load, `use_cache` export templates.
- New-model triage guide (`eagle3_new_model_triage_guide.md`); its
"document results" step now points at the internal tracker rather than
the removed chart.

### Testing

No code paths change — this only removes example YAMLs and two status
docs and edits one doc reference. Pre-commit
(ruff/markdownlint/yaml/license) passes on the kept/edited files.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (removes unverified examples
only; kept infra unchanged)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ❌ (pending)

### Additional Information

Follow-up to #1417. Next step (tracked separately): verify each removed
model end-to-end in modelopt, then re-add its YAML in a dedicated PR.
Note: the nmm-sandbox weekly EAGLE3 CI is being trimmed to the Qwen3-8B
baseline to match.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated EAGLE3 triage guide to streamline the verification workflow;
contributors now record test outcomes (status, experiment IDs, errors,
and applied fixes) in the team's internal triage tracker before
submitting model launcher configurations.

* **Chores**
* Removed legacy EAGLE3 example pipeline configurations and deprecated
triage documentation to reduce maintenance overhead.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-06-03 20:56:45 +00:00
yeyu-nvidiaandClaude Opus 4.6 a7b0a92047 EAGLE3 new model support: pipeline configs, triage docs, and Ministral-3 fixes (#1417)
## Summary

EAGLE3 automation triage work (OKR-30): testing the 4-step EAGLE3
offline pipeline against 12 new model architectures, documenting failure
modes, and fixing issues found.

### Code fixes (modelopt)

| File | Change |
|------|--------|
| `modelopt/torch/speculative/utils.py` | Extend VLM detection in
`load_vlm_or_llm` to check `text_config`/`llm_config` attrs (catches
`mistral3` models) |
| `modelopt/torch/speculative/plugins/modeling_fakebase.py` | Add
`consolidated.safetensors` fallback for checkpoints with incomplete HF
shards |
| `modelopt/torch/export/plugins/hf_spec_configs.py` | Set
`use_cache=True` in EAGLE export templates (fixes strict
`huggingface_hub` validation) |

### Pipeline infrastructure

- `examples/speculative_decoding/pipeline/eagle3/` — pipeline scripts
and configs:
- `offline_training.sh` — training + export with runtime patches for
older container modelopt
- `dump_offline_data_vllm.sh` — vLLM-based hidden state extraction (with
speculators compat patches)
- `dump_offline_data.sh`, `dump_offline_data_hf.sh` — alternative dump
paths
  - 18 quick-fail-check YAMLs for 12 models
  - 4 standalone task1 YAMLs

### Documentation

- `eagle3_triage_chart.md` — model test matrix, triage decision tree,
per-model results, failure catalog
- `eagle3_new_model_triage_guide.md` — step-by-step guide for triaging
new models

### Model test results (as of 2026-05-27)

| Model | task_0 | task_1 | task_2 | task_3 | Blocker |
|-------|--------|--------|--------|--------|---------|
| Qwen3-8B | - | - | - | - | Reference (existing) |
| Kimi-K2.5 | - | - | - | - | Existing (GB200) |
| **Ministral-3-8B** | SKIP | PASS | PASS | FAIL | `use_cache=null` in
export (fixed) |
| Ministral-3-14B | FAIL | - | - | FAIL | vLLM engine init fails |
| Qwen3.5-35B-A3B | TIMEOUT | - | - | - | Data synth too slow |
| gpt-oss-20b | FAIL | - | - | - | Tokenizer `HarmonyError` |
| Step-3.5-Flash | TIMEOUT | - | - | - | Data synth time limit |
| MiniMax-M2.5 | TIMEOUT | - | - | - | `trust_remote_code` needed |
| DeepSeek-V3.2 | no log | - | - | - | May not be mirrored |
| Qwen3.5-9B | - | - | - | - | Not yet run |
| Qwen3.5-27B | - | - | - | - | Not yet run |
| GLM-5 | - | - | - | - | Not yet run |

### Issues found and fixed

| # | Issue | Fix |
|---|-------|-----|
| 1 | `mistral3` model type not detected as VLM | Check
`text_config`/`llm_config` attrs in `load_vlm_or_llm` |
| 2 | Missing HF shard file (Ministral-3-8B) | Fallback to
`consolidated.safetensors` with Mistral native key aliases |
| 3 | `use_cache=null` in exported EAGLE config | Set `use_cache=True`
in export template configs |
| 4 | speculators incompatible with vLLM container | Runtime patches in
`dump_offline_data_vllm.sh` |
| 5 | `offline_training.sh` infra issues | Rewritten with runtime
patches for container modelopt |

## Test plan

- [x] Ministral-3-8B training passes (`cicd_1779829129`)
- [x] Ministral-3-8B export succeeds
- [ ] Ministral-3-8B benchmark passes (`cicd_1779901409` — pending with
all fixes)
- [ ] Dry-run remaining model configs

## Note

GitHub secret scanning alert #6 is a **false positive** —
`Mistral3ForConditionalGeneration` (a HuggingFace model class name in a
YAML comment) was flagged as a "Mistral AI API Key".

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-06-03 12:24:44 -07:00
h-guo18 902d36921a [Feat]: Streaming Hidden-states Dataset (#1509)
### What does this PR do?

Type of change: new feature

**Design doc:**
https://gist.github.com/h-guo18/241c94968b0591324c361d97cf995dd0
Jira ticket: https://jirasw.nvidia.com/browse/OMNIML-4341
Sandbox CI:
https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/327911555#L2228

Streaming hidden-states dataset: per-sample activations pulled from a
live `vllm serve` over HTTP, replacing on-disk activation dumps.

Two axes for future extensions:
- **Backend** (`_fetch`): vLLM now; TRT-LLM / SGLang next.
- **Algorithm** (`_format`): Eagle now; distillation / probing next.
- **Sandbox CI**:
https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/merge_requests/169

Shared plumbing — async producer, token-level truncation to
`training_seq_len`, loss-mask alignment, DDP via Accelerate's dispatcher
(rank 0 fetches, broadcasts), circuit breaker, resume — lives in
`StreamingDataset`. First instance: **`EagleVllmStreamingDataset`**.

API: `data.mode ∈ {online, offline, streaming}`; legacy configs
auto-promote.

### Usage

```yaml
data:
  mode: streaming
  data_path: input_conversations/train.jsonl
  streaming_server_url: http://localhost:8000
  streaming_model_name: meta-llama/Llama-3.1-8B-Instruct
training:
  training_seq_len: 4096   # also caps the prompt sent to vllm
```

Requires `vllm serve` with `ExampleHiddenStatesConnector` and
`dataloader_num_workers=0`.

End-to-end Slurm pipeline:
`tools/launcher/examples/Qwen/Qwen3-8B/hf_streaming_eagle3.yaml`.

### Testing

- **Unit**: full-corpus invariant, rank-0-only iter, resume, circuit
breaker, mocked-httpx integration.
- **E2E**: `launch_train.sh` against a stdlib `HTTPServer` mimicking the
connector.
- **Smoke** (Qwen3-8B / 8×H100 / 4096 ultrachat samples, single epoch):
train_loss 32 → 18, MT-Bench AR 1.003 → 1.20.

### TODO before un-drafting

- [ ] Observability counters (filtered, fetch failures, queue depth,
latency).
- [ ] Changelog entry.
- [x] Add test in sandbox.

**Non-goals (v1):** multi-epoch streaming, cross-rank dynamic load
balancing.

### Before your PR is "*Ready for review*"

- Backward compatible: ✅ (legacy configs auto-promote)
- New PIP dep: N/A
- New tests: ✅
- Changelog: ❌ (TODO above)
- Claude approval: ❌


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Streaming training mode with server-backed hidden-state fetching,
deterministic seed control, resume support, and streaming-specific
dataset options (server, model, prefetch, shared storage).

* **Behavior / Bug Fixes**
* Stronger mode validation; offline behavior derived from data mode;
resume handling adjusted to avoid double-skip during streaming runs.

* **Tests**
* End-to-end CI streaming test and expanded unit tests covering
streaming, resume, DDP, determinism, and failure cases.

* **Infrastructure**
* Launcher script and pipeline config for end-to-end streaming training.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1509?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-06-02 14:56:23 -07:00
Chenjie Luo 905259fbf5 Fix: use python3 in debugger server.sh (#1590)
### What does this PR do?

Type of change: Bug fix

`tools/debugger/server.sh` invokes `python -c "..."` inside
`check_modelopt_local` (and `pip install -e .[dev]` in the fallback
path). Containers that only ship `python3` (no `python` shim) cause the
check to fail with `python: command not found`, exit 127. The server
interprets this as "modelopt not editable-installed", runs `pip
install`, the second check fails the same way, and the server aborts
before it ever listens for commands.

Switches both the inline import check and the install fallback to
`python3` / `python3 -m pip`.

### Usage

```bash
bash tools/debugger/server.sh
# now starts cleanly on python3-only images
```

### Testing

Verified inside a container where `which python` returns nothing and
`python3` is `/usr/bin/python3` (Python 3.12). With the old script,
`server.sh` aborted at `check_modelopt_local`. With this fix, the check
passes against an existing editable install and the server proceeds to
wait for the client handshake.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `python3` is present on every
image that previously had `python`.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — internal tooling shell
script.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — internal debugger tooling, not user-facing.
- Did you get Claude approval on this PR?: N/A

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Updated debugger tooling to ensure consistent use of Python 3 for
package validation and installation processes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-06-01 21:47:21 +00:00
Chenhan D. Yu 54fb87e972 [OMNIML-4869] author_yaml (#1574)
Draft PR opened by **pensieve-intern** for
[OMNIML-4869](https://jirasw.nvidia.com/browse/OMNIML-4869).

Stage `author_yaml` of Epic `OMNIML-4868`. The agent ran from the SPEC
on the ticket description; review every change before marking ready.

_Always-draft is enforced — the bot never auto-merges._

---

**Agent's self-narration** (stripped from PR diff; surfaced here for
context):

`VERIFICATION_COMMENT.txt`:
```
OMNIML-4869 status: model not staged on cw_dfw.

Spec checks:
- HF Hub model exists: Qwen/Qwen3.5-4B (public, not gated; model_type=qwen3_5).
- Cluster stage check failed: /lustre/fsw/portfolios/coreai/projects/coreai_dlalgo_modelopt/hf-local/Qwen/Qwen3.5-4B not found.

Requested action on Epic: please stage the model via `/manage-assets cw_dfw get Qwen/Qwen3.5-4B`.

Notes:
- YAML already exists at tools/launcher/examples/Qwen/Qwen3.5-4B/specdec_bench_mtp_vllm.yaml with tp_size=1, gpus_per_node=1, and container vllm/vllm-openai:qwen3_5-cu130.
- Once staged, I can run the required dry-run: `uv run launch.py --yaml tools/launcher/examples/Qwen/Qwen3.5-4B/specdec_bench_mtp_vllm.yaml --dryrun --yes -v`.
```



_Pollution-strip removed `VERIFICATION_COMMENT.txt` from this commit
(sidecar narration and/or incidental lockfile regeneration are never
part of the agent's intended deliverable)._

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Added a benchmark configuration to evaluate speculative decoding (MTP)
performance for the Qwen3.5-4B model, enabling separate speed and
high-throughput (32k) runs with adjustable decoding and runtime
parameters.
* Configured execution settings for single-node GPU benchmarking and
standardized container/runtime invocation for reproducible performance
tests.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-05-31 11:53:00 +05:30
h-guo18 40a4dd326d [Feat]: Eagle Dry Run Mode (#1566)
### What does this PR do?

Type of change: new feature

Adds `--dry_run` to `examples/speculative_decoding/main.py`: load →
`mtsp.convert` → save, then exit (no `trainer.train()`). With
`FakeBaseModel`, the convert→save→export chain runs in seconds and
produces an exportable EAGLE3 / Medusa / DFlash checkpoint with correct
structure but untrained draft-head weights — useful for end-to-end
plumbing smoke tests on downstream stacks (vLLM, TRT-LLM, SGLang)
without paying for a real training run.

Where it sits among existing EAGLE3 modes:

```
EAGLE3 modes
├── online      base model runs forward in-loop
├── offline     reads pre-dumped hidden states from disk
├── streaming   streams hidden states from a live server in-loop
└── dry-run ★  skip training entirely; convert + save + export   ← NEW (this PR)
```

Companion `FakeBaseModel` fixes so small base checkpoints work:
- Synthesize the weight_map from a single `model.safetensors` when no
sharded index is present (Llama-3.2-1B, Qwen3-0.6B, …).
- Honor `tie_word_embeddings`: reuse `embed_tokens` when `lm_head` is
absent from safetensors.

A new launcher YAML
(`tools/launcher/examples/Qwen/Qwen3-8B/hf_eagle3_dryrun.yaml`) wires
this together as a one-task pipeline.

### Usage

```bash
# Direct
python main.py --dry_run \
  --config modelopt_recipes/general/speculative_decoding/eagle3.yaml \
  model.model_name_or_path=meta-llama/Llama-3.1-8B-Instruct \
  model.use_fake_base_for_offline=true \
  data.offline_data_path=/tmp/dryrun-placeholder \
  training.output_dir=ckpts/dryrun
python scripts/export_hf_checkpoint.py --model_path ckpts/dryrun --export_path export/dryrun

# Launcher
uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_eagle3_dryrun.yaml --yes
```

### Testing

- `tests/unit/torch/speculative/plugins/test_fakebase.py`: 3 new cases —
single-file fallback, tied-embeddings fallback, and the negative case
(missing `lm_head` without tying).
-
`tests/examples/speculative_decoding/test_eagle.py::test_eagle3_dry_run`:
full `launch_train.sh --dry_run → export_hf_checkpoint.py` chain on
`tiny_llama`; asserts exported state_dict has all
`LLAMA_EAGLE_SINGLE_LAYER` required keys.
- Manually verified end-to-end on Llama-3.1-8B-Instruct (sharded),
Llama-3.2-1B-Instruct (single-file + tied), and Qwen3-0.6B
(single-file).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `--dry_run` is opt-in;
`FakeBaseModel` changes are additive fallbacks.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — examples-only addition.
- Did you get Claude approval on this PR?: ❌

### Additional Information

N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a `--dry_run` CLI flag to perform a fast early-exit execution
that saves model artifacts without training.
* Added support for single-file model checkpoint formats in the loader.

* **Bug Fixes**
  * Improved checkpoint-loading error messages and handling.
  * Enhanced tied-embeddings fallback when head weights are absent.

* **Tests**
* Added integration and unit tests covering dry-run behavior and
single-file/tied-embedding loading.

* **Documentation**
  * Added a launcher example for dry-run smoke testing.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1566?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-05-29 15:02:26 -07:00
Keval MorabiaandClaude Opus 4.7 5eba879002 Launcher nvrx install from PyPI; specdec_bench guard modelopt import (#1567)
## Summary

- **`tools/launcher/common/service_utils.sh`** — Megatron-LM (post PR
#4522) asserts `nvrx >= 0.6.0` on import. The previous git-clone-HEAD
install in `util_install_extra_dep` produced setuptools_scm versions
like `0.0.0.dev1+hash` whenever upstream HEAD landed between tags,
failing that assertion. Switch to a PyPI install pinned at `>= 0.6.0`.
nemo containers ship nvrx in two Python envs; `pip` (system) and `python
-m pip` (uv venv) target different ones, so uninstall from both and
install into the venv where Python actually imports from.
- **`examples/specdec_bench/specdec_bench/__init__.py`** — guard the
`from modelopt import __version__` import so `specdec_bench` can be
loaded inside a vllm container that doesn't have modelopt installed.

## Test plan

- [x] Launcher YAMLs (e.g. Qwen3-8B PTQ on ComputeLab) install
`nvidia-resiliency-ext>=0.6.0` and don't trip the Megatron-LM
`has_nvrx_async_support` assertion.
- [x] `specdec_bench` imports cleanly under a vllm container with no
modelopt installed (prints the warning and uses `__version__ =
"0.0.0"`).

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Streamlined dependency installation: removed fragile repo-based
install and now ensures a clean uninstall/upgrade flow for the external
resiliency package to improve reliability across environments.
* Improved version handling: added guarded fallback and warning when an
optional package is missing, reporting a safe default version to avoid
runtime errors.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1567?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-29 19:58:35 +05:30
Chenhan D. Yu a2c496af0d [OMNIML-4788] specdec_bench: configuration.json provenance + upload_to_s3 (#1531)
> [!WARNING]
> **Breaking on-disk schema change (specdec_bench v1.0.0).** This PR
renames the acceptance-rate metric fields across `AcceptanceRate` /
`MTBench` / `SpecBench` writers:
>
> | Old (pre-1.0.0) | New (1.0.0) |
> |---|---|
> | `Request_AR` | `Request_AL` |
> | `Category_AR` | `Category_AL` |
> | `Average_AR` | `Average_AL` |
> | — | `Joint_Acceptance_Rate` (new) |
>
> The renamed values were always **acceptance length** (mean tokens
generated per inference step), not a rate, and the visualizer reads
`*_AL`. Pre-1.0.0 runs in S3 have `*_AR` and no `Joint_AR`; they must be
re-run or post-processed before comparing. The visualizer aggregates
runs by `specdec_bench` major version so accidental cross-methodology
comparison is blocked.

### What does this PR do?

Type of change: new feature

Adds reproducibility provenance to `specdec_bench/configuration.json`
and ports `upload_to_s3.py` from `iputterman/specdec_bench@main`
(personal-namespace fork) into upstream. This is the first PR in a
multi-stage migration off Izzy's fork now that he's left the team.
Tracked in [OMNIML-4788](https://jirasw.nvidia.com/browse/OMNIML-4788).

**Provenance fields added to configuration.json** (alongside existing
argv / engine_version / gpu / python_version):

- `specdec_bench_version` — methodology semver declared in
`specdec_bench/__init__.py`. Bump minor on additive metrics, major on
changed metric *definitions*. The visualizer (Phase 4 of the migration)
will aggregate runs by major version so plots don't accidentally compare
across methodology changes.
- `specdec_bench_sha`, `modelopt_sha`, `modelopt_version`,
`nmm_sandbox_sha`, `container_image` — code/runtime provenance. Each
prefers an env var set by the harness (`SPECDEC_BENCH_SHA`,
`MODELOPT_SHA`, `MODELOPT_VERSION`, `NMM_SANDBOX_SHA`,
`CONTAINER_IMAGE`) and falls back to `git rev-parse` /
`modelopt.__version__` when running standalone. The env-var preference
is necessary because the runtime container has no `.git/` (the launcher
packager tarballs source without git metadata) and may not have
`modelopt` installed.
- `checkpoint.{path, size_bytes, index_sha256, index_source}` — cheap
reproducibility fingerprint that hashes `model.safetensors.index.json`
(or `config.json` fallback). Changes whenever any tensor changes.
- `serving_config` — engine-level config dict captured after init via a
new `Model.get_serving_config()` method. VLLM dumps `AsyncEngineArgs` +
the live `vllm_config.to_dict()`; SGLANG dumps the `engine_kwargs`
passed to `sgl.Engine`; TRTLLM left at the base default `{}` for a later
iteration.
- `timestamp` — UTC ISO 8601.

**Other changes**
- `upload_to_s3.py` + `specdec_bench/s3_utils.py` ported from
iputterman/specdec_bench@main. Recognizes run dirs by sentinel files,
refuses to overwrite existing S3 prefixes.
- `_redact_config` allowlists `tokenizer`, `tokenizer_path`,
`tokenizer_mode`, `tokenizer_revision` so the model path stops being
redacted (latent bug from substring-matching `token` ⊂ `tokenizer`).
- `requirements_speed.txt`: `boto3`, `botocore` added (used by
`s3_utils`).

**Out of scope** (deferred to Phase 1b / Phase 2):
- `--sweep_config` driver that emits per-run-dir nesting
`<sweep>/<NNN_dataset_c<conc>>/`
- `--s3_upload` flag baked into `run.py` itself
- Launcher auto-injection of the provenance env vars (currently the
example YAML sets them statically)
- `container_digest` (enroot integration) and full GPU/driver inventory
- TRTLLM `get_serving_config()`

### Usage

```bash
# Run a smoke benchmark (Qwen3.5-4B + vLLM + MTP draft=3) — example YAML included
uv run launch.py --yaml examples/Qwen/Qwen3.5-4B/specdec_bench_mtp.yaml --yes

# After it lands, upload the run directory to S3:
S3_KEY_ID=team-specdec-workgroup \
S3_SECRET=... \
python upload_to_s3.py /path/to/sweep_dir s3://team-specdec-workgroup/results
```

### Testing

Cluster-tested end-to-end on cw-dfw (Slurm job 11978794, NeMo Run
experiment `cicd_1779403623`, ~19 min wall):

- Qwen3.5-4B + vLLM + MTP draft=3 + SPEED-Bench-Internal/qualitative (80
requests)
- `configuration.json` (22 KB) populated all eight new provenance fields
- `Request_AR` mean 3.327 (vs 3.330 on the pre-Phase-1a run — within
noise; methodology unchanged)
- `upload_to_s3.py` (real upload, not dry-run) landed
[s3://team-specdec-workgroup/results/qwen35_4_mtp_smoke_2026-05-21/specdec_bench_mtp/](https://app.s8k.io/buckets/team-specdec-workgroup/?prefix=results%2Fqwen35_4_mtp_smoke_2026-05-21%2F)
where the visualizer at http://10.131.132.205:8080 can pick it up.

### Before your PR is "Ready for review"

- Is this change backward compatible?: ✅
- `configuration.json` only gains fields. `upload_to_s3.py` /
`s3_utils.py` are new files. `Model.get_serving_config()` default = `{}`
so existing subclasses without an override behave as before.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- `boto3` / `botocore` are Apache 2.0 (permissive); `upload_to_s3.py` +
`s3_utils.py` are ported from a private NVIDIA repo with explicit
copyright headers retained.
- Did you write any new necessary tests?: ❌
- Validated by cluster smoke (see Testing). Will add unit-tests for
`dump_env` provenance fields and `upload_to_s3._discover_runs` in a
follow-up.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
  - Internal-facing tooling.
- Did you get Claude approval on this PR?: ❌ (triggering after open)

### Additional Information

Tracked in JIRA
[OMNIML-4788](https://jirasw.nvidia.com/browse/OMNIML-4788). The full
multi-phase plan is on that ticket's SPEC block — this PR is Phase 1a.

Cherry-picked alongside the harness change are two example YAMLs
(`examples/Qwen/Qwen3.5-4B/specdec_bench.yaml` for the NONE
autoregressive baseline, `..._mtp.yaml` for the MTP run) that gave us
cluster-test evidence. Can be split out if preferred.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an S3 upload CLI for benchmark results with dry-run and
skip-existing options
* Automatic capture of run configuration, provenance and redacted
environment into saved config
  * Models now export serving configuration for reproducible runs
* New launcher entrypoint and example job configs for Qwen SPEED-Bench
runs

* **Documentation**
* README section describing S3 upload usage and supported local layouts

* **Bug Fixes / Changes**
  * Acceptance-rate metric keys renamed in output (AR -> AL)

* **Tests / CI**
* New tests for redaction and S3 utilities; CI now runs specdec_bench
examples

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1531?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: chenhany <chenhany@nvidia.com>
2026-05-28 19:12:30 +00:00
Keval MorabiaandClaude Opus 4.7 5d0441ae3d Create shared Megatron calibration forward loop for prune / quantize with megatron pretraining data style sequence packing (#1501)
## Summary

Replaces the bespoke calibration loops in Megatron-LM and
Megatron-Bridge prune / quantize example scripts with a single shared
utility,
`modelopt.torch.utils.plugins.megatron_calibration.get_megatron_calibration_forward_loop`.

The shared loop iterates a packed calibration dataloader built via
`get_dataset_dataloader(pack=True)` and drives a logits-free prefill
pass through the model so activation hooks fire on every layer.
`pack=True` produces Megatron-LM pretraining-style **global-stream
packing**: all raw samples are concatenated into one EOS-separated token
stream and sliced into uniform-length rows. The trained model has seen
this distribution extensively during pretraining, so the activations
produced during calibration are representative of the model's natural
behavior.

Migrates four call sites:
- `examples/megatron_bridge/prune_minitron.py`
- `Megatron-LM/examples/post_training/modelopt/{prune,quantize}.py`
(separate PR:
[NVIDIA/Megatron-LM#4881](https://github.com/NVIDIA/Megatron-LM/pull/4881))
- `Megatron-Bridge/examples/quantization/quantize.py` (separate PR)

Each call site passes `pack=True` explicitly with an inline comment so
users see the option and know when to flip it. The function-level
defaults (`get_dataset_dataloader(pack=False)`,
`get_megatron_calibration_forward_loop(pack=False)`) remain
back-compat-safe.

Unified defaults across all four sites: `--calib-dataset
nemotron-post-training-dataset-v2`, `--calib-size 1024`,
`--calib-max-sequence-length 4096`, `--calib-batch-size 1`.

## Experimental results

Qwen3-8B on full 100% MMLU (n=14042; binomial 2σ noise floor ≈ ±0.78 pt
at acc ≈ 0.7), 0-shot, eval batch_size=4. Calibration on the default
workload: nemotron-post-training-dataset-v2, seq_length=4096,
calib_batch_size=8.

**Three calibration data shapes compared:**

- **Padded**: one doc per row, padded to `seq_length`, pad tokens flow
through the forward (legacy `get_calib_dataloader` pad+truncate
behavior).
- **Trimmed**: one doc per row, each row trimmed to its real content
length via `attention_mask`, with EOS forced at the last real position;
pad never enters the forward. This is no longer part of this PR.
- **Packed**: global-stream slicing — all docs concatenated
EOS-separated into one token stream, sliced into uniform `seq_length`
rows. Matches Megatron's `.bin`/`.idx` pretraining distribution. Enabled
via `pack=True` in `get_megatron_calibration_forward_loop`.

| Workload | Padded | Trimmed | **Packed** |
|---|---|---|---|
| M-LM NVFP4 quantize (`NVFP4_DEFAULT_CFG`) | 0.707 | 0.708 | 0.709 |
| M-Bridge Minitron prune (Qwen3-8B → 30L / 3584 / 11776 ≈ 6B params) |
0.576 | 0.573 | **0.589** |

### Key findings

- **M-LM quantize quality is calibration-mode-insensitive** for dense
Qwen3-8B NVFP4 — all three modes are nearly identical.
- **M-Bridge prune**: More sensitive to calibration data shape. Packed
wins over Padded on full MMLU.

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-28 18:34:37 +00:00
Jenny Chen b49f9b9e2d Megatron-Bridge import example in launcher for Nemotron Super V3 (#1516)
### What does this PR do?

Type of change: New feature

Add Megatron-Bridge import example in launcher for Nemotron Super V3

### Usage

```python
# Usage:
#   update .env-slurm with environment variables
#   cd tools/launcher
#   uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/megatron_bridge_import.yaml --yes
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added Megatron-Bridge Hugging Face→Megatron checkpoint import
(CPU-capable entrypoint) and an example pipeline for importing the
Nvidia Nemotron-3 Super 120B model.
  * Exposed an `output_dir` global variable for pipeline interpolation.

* **Chores**
  * Updated pinned Megatron-LM module.
  * Extended .gitignore to cover `.env*` variants.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1516?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
2026-05-27 15:36:56 +00:00
Shengliang Xu 9d0d97829a chore(lint): modernize typing (PEP 604/585) and enable UP032 (#1537)
### What does this PR do?

Type of change: chore / refactor (no behavior change)

Two small lint-cleanup commits:

**1. `chore(typing): modernize Union/Optional/List to PEP 604 / 585
syntax`** (8 files)

- Replace `X = Union[A, B] # noqa: UP007` with `X: TypeAlias = A | B`
for the six module-level type aliases (`ModelLike`, `Criterion`,
`NodeTarget`, `CalibrationDataType`, `Hparam.Importance` /
`ActiveSlice`). The `TypeAlias` annotation is required so mypy continues
to treat them as aliases under PEP 604.
- Modernize forward-ref unions in `modelopt/onnx/quantization/autotune/`
to full-string forward refs (e.g. `"RegionPattern | None"`).
- Update docstring type tags in
`examples/puzzletron/evaluation/hf_deployable_anymodel.py`.

**2. `chore(lint): remove UP032 ignore and convert .format() to
f-strings`** (10 files)

- Drop `UP032` from `extend-ignore` in `pyproject.toml`.
- Auto-convert 19 `"...".format(...)` calls to f-strings across export
plugins, examples, tests, and tools. One conversion in
`modelopt/torch/utils/plugins/megatron_generate.py` was wrapped manually
to stay under the 100-char limit.

**Intentionally left as-is:**

- `tools/launcher/slurm_config.py` keeps its `# ruff: noqa: UP045` —
nemo_run's CLI parser can't introspect PEP 604 optional annotations.
- `modelopt/torch/puzzletron/*` is **not** touched. The subtree disables
ruff's `UP` family entirely (per-file-ignore `"UP"`) while migration is
in progress, and converting `Optional[X]` to `X | None` there would
silently break runtime introspection in
`block_config._get_dataclass_type` that uses `get_origin(tp) is
typing.Union` (PEP 604 unions return `types.UnionType` from
`get_origin`, not `typing.Union`). Best revisited when puzzletron's lint
carve-out is narrowed.
- `UP038` (`isinstance(x, (int, float))` → `isinstance(x, int | float)`)
— ruff has officially deprecated this rule; PEP 604 in isinstance is
slightly slower and misleads readers about PEP 695 / `Optional`. Ignore
kept.

### Usage

No user-facing API changes.

### Testing

- Pre-commit hooks (ruff check, ruff format, mypy, bandit, license) pass
on both commits.
- Ruff status against `main`: 37 unrelated pre-existing findings
(W291/W293/E501/RUF005/PLR1704); zero new findings introduced by this
PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — Runtime behavior of the six
type aliases changes from a `typing.Union` instance to
`types.UnionType`. Downstream code introspecting via `get_origin(...) is
typing.Union` on these aliases would break, but no in-repo caller does
this on them. (The introspection in
`modelopt/torch/puzzletron/block_config.py` operates on user-supplied
dataclass field types, none of which are these aliases.)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (no behavior change)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — internal style refactor; happy to add a Misc note if reviewers want
one.
- Did you get Claude approval on this PR?: ❌ — not yet.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Style**
* Modernized type annotations across the codebase to use Python 3.10+
union syntax and TypeAlias where appropriate.
* Standardized string formatting to f-strings, improving clarity of
logs, errors, and validation messages.

* **Chores**
  * Updated linting configuration to reflect modern typing/style rules.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1537?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-05-26 16:47:31 -07:00
Chenhan D. Yuandh-guo18 df68ccdcbd [OMNIML-4740] synth_support (#1496)
## Summary

Adds the EAGLE3 offline pipeline YAML for `moonshotai/Kimi-K2.5-DFlash`,
adapted from `Qwen/Qwen3-8B`'s offline YAML. **`task_0` cluster-tested
green on cw-dfw** (Slurm 11782946, experiment `cicd_1778864959`, elapsed
1:02:11).

Driven by /babysit-jira on OMNIML-4740. Replaces the original
pensieve-intern synth_support agent draft (1b021024) which had three
structural issues:

1. Set `global_vars.hf_model` to the *output* path (`Kimi-K2.5 DFlash`)
instead of the *input* checkpoint.
2. Used a TRT-LLM container (`release:1.2.0`) that doesn't register
`KimiK25ForConditionalGeneration` as of 2026-05-14.
3. Committed `VERIFICATION_COMMENT.txt` (a runner sidecar, not artifact
code) and `uv.lock` regen (+1037/-81, the source of the prior
`mergeable_state: dirty`).

This PR is the cleaned + cluster-validated replacement.

## Changes

- Directory `Kimi-K2.5 DFlash` → `Kimi-K2.5-DFlash` (Slurm tar packaging
breaks on spaces in `job_name`/path).
- `global_vars.hf_model: /hf-local/moonshotai/Kimi-K2.6` — the canonical
Kimi-K2.5 input stand-in staged by the operator on cw-dfw.
- `task_0` migrated from TRT-LLM to vLLM:
  - `script: common/vllm/query.sh`
  - `container: vllm/vllm-openai:latest`
  - `ntasks_per_node: 1` (vLLM is single-process)
  - `--tensor-parallel-size 8`, `--trust-remote-code`
- `--enforce-eager` (vllm-openai:latest is missing `torch/bin/ptxas` for
inductor autotuning)
- `--gpu-memory-utilization 0.95` + `--max-model-len 4096` (Kimi weights
are 595 GB bf16 on 8×80 GB = 93% weight occupancy; default 0.9 leaves
-1.1 GiB for KV cache)
- `VLLM_STARTUP_TIMEOUT=1800` env (Kimi load is ~7.7 min, default 600s
in `query.sh` is not enough)
- `--data` switched to the in-repo `synthetic_conversations_1k.jsonl` —
the canonical `Speculative-Decoding-Prompt-Samples` isn't on cw-dfw; the
in-repo dataset is the portable, packager-shipped input for smoke
testing.
- Sidecars + uv.lock removed from the diff.

## Cluster-test evidence (mandatory per the refined synth_support spec)

```
$ SLURM_CLUSTER=cw_dfw uv run slurm.py \
    --yaml '.../moonshotai/Kimi-K2.5-DFlash/hf_offline_eagle3.yaml' \
    pipeline.task_1.skip=true pipeline.task_2.skip=true pipeline.task_3.skip=true \
    --yes --detach

Slurm 11782946: COMPLETED, elapsed 1:02:11
Loading weights took 461.45 seconds
Model loading took 71.44 GiB memory and 465.10 seconds
Map (num_proc=32): 100%|██████████| 100/100 [06:42<00:00, 4.02s/example]
Saved 10 shards: /scratchspace/data/train-{1..10}-00010.jsonl
```

Real assistant response verified end-to-end — Kimi correctly answered
the "bat and ball" CRT problem:

> The ball costs **$0.05** (5 cents). Here's why: If the ball costs
$0.05, then the bat costs $1.05 (which is $1.00 more). Together they
cost $1.10.

## Test plan
- [x] Dry-run validation: `slurm.py --yaml ... --dry-run` exits 0
- [x] Cluster test on cw-dfw: `task_0` only (`task_1/2/3.skip=true`) —
green, 1000 prompts processed, 10 synth-data shards written
- [ ] Downstream `run_pipeline` stage of
[OMNIML-4735](https://jirasw.nvidia.com/browse/OMNIML-4735) will
exercise task_1..3 end-to-end with the produced synth data

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Improvements**
* Enhanced rank initialization with extended fallback support for
various resource scheduling systems, improving compatibility across
distributed environments.
* Improved cross-rank coordination mechanism for dependency
installation, ensuring reliable and efficient setup in multi-node
deployments.

* **New Features**
* Added configuration example for Kimi-K2.5 offline speculative decoding
pipeline with vLLM integration.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1496?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Signed-off-by: chenhany <chenhany@nvidia.com>
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-05-25 01:12:31 +00:00
h-guo18 f59f3ae1c3 [Example]: Dflash-Offline Launcher Example (#1529)
### What does this PR do?

Type of change: new example

Offline DFlash training launcher example for Qwen3-0.6B. Two-task
pipeline: dump base-model hidden states via HF forward, then train
DFlash on the dump.

- `tools/launcher/examples/Qwen/Qwen3-0.6B/hf_offline_dflash.yaml` — new
launcher YAML
- `tools/launcher/common/eagle3/dump_offline_data_hf.sh` — new HF-backed
dump script (DFlash's `answer_only_loss=true` requires `loss_mask`,
which the existing TRT-LLM dump backend does not produce)
- `examples/dataset/synthetic_conversations_1k.jsonl` —
`conversation_id` added at-source (asserted by
`compute_hidden_states_*.py`; previously injected at test-fixture time)
- `tests/regression/torch/speculative/test_dflash_offline.py` —
`tagged_synth_data_path` fixture removed since the field is now
at-source

### Usage

```
uv run launch.py --yaml examples/Qwen/Qwen3-0.6B/hf_offline_dflash.yaml --yes
```

### Testing

End-to-end on CoreWeave Slurm (1×1 GPU): task_0 dump succeeded; task_1
training loss 7.85 → 2.94 over 2 epochs, regression PASSED.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (`conversation_id` is purely
additive on the dataset)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (existing
`test_dflash_offline.py` covers the workflow)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (new example only)
- Did you get Claude approval on this PR?: ❌

### Additional Information

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added launcher script for offline hidden-state processing with
HuggingFace backend support and distributed execution capabilities
* Added training configuration example for Qwen3-0.6B model with DFlash
speculative decoding pipeline

* **Tests**
* Updated offline hidden-states regression test to improve data handling
workflow

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1529?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-05-24 17:45:13 -07:00
Chad VoegeleandClaude Opus 4.7 e2d4d73920 Agent Skills Updates From Live Trials (#1493)
### What does this PR do?

Type of change: bug fix

<!-- Details about the change. -->

### Usage

Ask Claude Code:
```
Quantize `mistralai/Mistral-Medium-3.5-128B` to NVFP4 using the ModelOpt NVFP4 experts-only recipe.

Run on $cluster

Evaluate the resulting quantized checkpoint on:
- GPQA Diamond AA v3
- SciCode AA v2

Complete the quantization and evaluation workflow end to end. Prompt when you require user input, otherwise keep going.
```

### Testing
I'm running the full loop with the above prompt, and iterating on skills
to resolve undesired agent behavior.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: TODO

### Additional Information
See trials log for details.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added SLURM Quality of Service (QoS) configuration support for job
submission
* Introduced 8 new evaluation task recipes (AIME 2025, GPQA, IFBench,
LiveCodeBench, SciCode, AA-LCR, HLE-AA, MMMU-Pro, tau2_bench)
  * Enhanced job monitoring with continuous polling-based tracking

* **Documentation**
* Restructured evaluation workflow with explicit dry-run, canary, and
full-run validation stages
* Expanded PTQ validation with mandatory pre-deployment verification
gates
  * Updated remote cluster selection and quantization detection guidance

* **Tests**
* Updated evaluation test expectations to reflect refined workflow
stages

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1493?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 17:16:51 -05:00
Keval Morabia d7733562dd fix(prune): Minitron HybridModel + GPT-family fused-TE-spec import/export (#1518)
## Summary

Split out of #1501 so the pack=True calibration packing change can land
independently. This PR carries the pruning + export-side fixes.

**Pruning bug fixes**

- Register `HybridModel` (parent of `MambaModel` in modern Megatron-LM)
under a new `HAS_HYBRID` flag so `mcore_minitron` actually prunes
Nemotron-H et al. Previously `HybridModel` instances fell through
`convert_to_dynamic`, got `freeze()`-ed (collapsing `hidden_size` /
`num_layers` to a single choice), and produced unloadable saved
checkpoints with mixed pruned/unpruned dims.
- Replace the `isinstance(MambaModel)` gate in `_get_hybrid_pattern_key`
with attribute-presence detection so both `MambaModel` (still using
`hybrid_override_pattern`) and plain `HybridModel`
(`hybrid_layer_pattern`) are handled uniformly.
- Track `in_features` as a dynamic attribute on
`_DynamicTEQKVLayerNormColumnParallelLinear` so TE's forward-time
`inp_shape[-1] == in_features` assertion holds when `hidden_size` is
pruned.
- Dedupe MambaModel / HybridModel divisor dict into `_HYBRID_DIVISORS`.

**Fused-TE-spec import/export for GPT-family**

- Importer: prefer per-context keys (`fused_input_layernorm`,
`fused_pre_mlp_layernorm`); fall back to legacy `fused_norm` for
Nemotron-H back-compat. **Raise `KeyError`** when a fused-TE model has
neither rule registered — the branch only fires when the model uses
fused `TELayerNormColumnParallelLinear`, so a missing rule is
unambiguously a plugin misconfig that would otherwise ship a
chance-accuracy checkpoint.
- Exporter: mirror the same fallback chain in `_get_fused_norm_weight`
so GPT-family models round-trip cleanly back to HF.
- Add the new rules to Qwen3, Qwen2.5, Llama, Llama4 (MoE-only, only
`fused_input_layernorm`), DeepSeek, GptOss (MoE-only, only
`fused_input_layernorm`) import and export mappings.
- Preserve TE `_extra_state` from the existing module state dict (don't
blank to `None`) at both call sites in the importer.

**Misc**

- `megatron_prefill`: `.contiguous()` on the logits slice before
`broadcast_from_last_pipeline_stage` — broadcast asserts contiguity
which fails when SP pads `seq_length` to a multiple of TP.
- `megatron_mmlu`: accept `mmlu_dataset` kwarg so callers can point at a
local copy of `cais/mmlu`.
- `warn_rank_0`: auto-bump `stacklevel` by 1 inside the wrapper so
callers' warnings point at user code, not at the wrapper frame.
- `tools/launcher/examples/Qwen/Qwen3-8B/megatron_lm_ptq.yaml`: bump
`mmlu_lower_bound` 0.68 → 0.75 (validated end-to-end with the fused-norm
import fix).
- CHANGELOG: bug-fix entry for the importer; date correction on the 0.44
entry.

## Consumer

Megatron-LM PR https://github.com/NVIDIA/Megatron-LM/pull/4807 —
`prune.py` / `mmlu.py` consume these APIs and currently ship inline WARs
against released 0.44. Once 0.45 ships and the modelopt pin is bumped,
those WARs collapse to one-liners.

Related: #1501 (calibration packing).

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Importer/exporter now correctly load fused LayerNorm weights for
GPT-family models, preferring context-specific fused keys with a legacy
fallback.

* **New Features**
  * Hybrid Mamba/HybridModel support added for pruning/NAS workflows.
* MMLU evaluation accepts a customizable dataset path (default:
"cais/mmlu").

* **Improvements**
* Extended export/import mappings and state handling across DeepSeek,
GPT, Llama, Qwen; ensured last-stage logits are contiguous.

* **Documentation**
  * Updated changelog entry and release date adjustment.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1518?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-20 01:14:12 +05:30
Chenhan D. Yu 1f9c0bfae2 fix(launcher): propagate dflash training failures and fast-fail vllm smoke test on missing draft (#1409)
## Summary

Two related fixes to the speculative-decoding launcher scripts that
surfaced as silent failures in nmm-sandbox CI pipeline
[50509971](https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/pipelines/50509971):

- **`dflash_online_training.sh`** — install an `EXIT` trap so
`error_handler`'s `FAIL_EXIT=1` actually surfaces as a non-zero shell
exit. Previously, training failures (e.g. CUDA OOM) emitted `<REPORT>`
tags but the script exited 0, and Slurm marked the task `SUCCEEDED`. The
orchestrator then proceeded to run dependent tasks against a
non-existent checkpoint.

- **`vllm_smoke_test.sh`** — when `DRAFT_CKPT_DIR` is set but contains
no `exported-checkpoint-*` (typically because the upstream training task
crashed), exit 1 with a clear message instead of falling through to the
self-draft branch. The self-draft fallback produced a confusing
`num_speculative_tokens was provided but without speculative model`
error from vLLM's `SpeculativeConfig` validation that obscured the real
cause (the upstream training failure).

## Test plan

- [ ] Re-run `Qwen3-8B_DFlash_online` on Computelab with insufficient
GPU memory; verify task_0 surfaces as `FAILED` (not `SUCCEEDED`) and
task_1 doesn't bother trying the self-draft path.
- [ ] Re-run on adequate hardware (H200/B200) and verify the trap
doesn't fire for the success path (script exits 0 normally).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Bug Fixes

* Improved error handling in training script to properly detect and
report training failures (e.g., out-of-memory errors) instead of
incorrectly marking them as successful
* Enhanced validation for missing draft checkpoints with clearer error
messaging to prevent confusing downstream configuration errors

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-05-07 20:15:42 +00:00
Shengliang Xu f34f488a83 Add a general composable $import system for YAML configs, and use it to implement composable recipes (#1253)
### What does this PR do?

Type of change: New feature

Adds a general composable YAML config loading layer for ModelOpt configs
and recipes. YAML remains the source of truth for configuration data,
while Python/Pydantic-compatible types provide schema validation at load
time. This PR uses that loader to de-duplicate PTQ recipes, introduce
reusable config snippets/presets, and start migrating selected hardcoded
quantization presets to YAML.

#### Problem

1. Built-in PTQ recipes duplicated numeric format definitions, KV-cache
entries, and the default quantizer exclusion list.
2. YAML snippets were reusable only by convention; they did not declare
or validate the schema they were meant to satisfy.
3. Loading YAML-backed quantization presets from
`modelopt.torch.quantization.config` could not depend on
`modelopt.recipe` without creating circular imports.
4. Directory-format recipes exposed import-resolution details in the
recipe loader and used `recipe.yaml` plus nested `metadata:` in a way
that made metadata handling inconsistent.

#### Solution

**Shared YAML config loader**

- Adds `modelopt.torch.opt.config_loader` as the low-level loader used
by both `modelopt.recipe` and `modelopt.torch.quantization.config`.
- Keeps the public `modelopt.recipe.load_config()` entry point, while
removing the private `modelopt/recipe/_config_loader.py` shim.
- Handles YAML loading, built-in/filesystem path resolution, suffix
probing, `ExMy` conversion for `num_bits` / `scale_bits`, `$import`
expansion, and schema validation.
- Lives below `modelopt.recipe` in the dependency graph to avoid
circular imports from quantization config code.

**Composable `$import` system**

Recipes and snippets can declare an `imports` mapping, then reference
entries with `{$import: name}`.

`$import` semantics:

- **Dict value**: replaced with the imported dict. Multiple imports are
supported with ordered precedence; inline keys override imported keys.
- **List entry**: schema-driven behavior for strongly typed lists. If
the snippet schema matches the containing list type, the imported list
is spliced. If the snippet schema matches the list element type, the
imported element is appended. Other schema combinations are rejected.
- **Multi-document YAML**: supports snippets that need an `imports`
header plus a list body.
- **Recursive and scoped**: snippets can import other snippets; import
names are scoped per file.
- **Cycle detection**: circular imports report a clear error.

**Snippet schema validation**

- Every reusable snippet referenced through `imports` must declare a `#
modelopt-schema: ...` preamble.
- Snippets are validated after nested imports are resolved.
- Schema paths are restricted to the `modelopt.` package and may be
Pydantic models, `TypedDict` classes, or explicitly typed container
aliases such as `list[QuantizerCfgEntry]`.
- Untyped list imports are rejected so list append/splice behavior stays
strongly typed.

**Recipe model and directory recipe cleanup**

- `ModelOptRecipeBase` now owns a `metadata: RecipeMetadataConfig`
field.
- `ModelOptPTQRecipe` is the PTQ recipe schema; the overlapping
YAML-specific PTQ config class was removed.
- Directory recipes now use `metadata.yaml` / `metadata.yml` for
top-level metadata fields, plus section files such as `quantize.yaml`.
- Directory recipe loading now delegates import resolution to
`load_config()` instead of manually using raw config loading.

**Config snippet and preset library**

Adds reusable snippets under `modelopt_recipes/configs/`:

- `numerics/`: `fp8`, `nvfp4`, `nvfp4_static`
- `ptq/units/`: `base_disable_all`, `default_disabled_quantizers`,
`w8a8_fp8_fp8`, `w4a4_nvfp4_nvfp4`, `kv_fp8`, `kv_fp8_cast`,
`kv_nvfp4_cast`
- `ptq/presets/`: YAML presets for `FP8_DEFAULT_CFG` and `FP8_KV_CFG`

`FP8_DEFAULT_CFG` and `FP8_KV_CFG` now load from YAML presets via
`load_config()`.

**Recipe migration and naming**

- General PTQ recipes now use shared imports instead of repeating the
same quantizer fragments inline.
- General PTQ recipe paths were renamed to KV-first naming, for example:
  - `general/ptq/fp8_default-fp8_kv` -> `general/ptq/fp8_default-kv_fp8`
- `general/ptq/fp8_default-fp8_cast_kv` ->
`general/ptq/fp8_default-kv_fp8_cast`
- `general/ptq/nvfp4_default-none_kv_gptq` ->
`general/ptq/nvfp4_default-kv_none-gptq`
- `general/ptq/nvfp4_default-nvfp4_cast_kv` ->
`general/ptq/nvfp4_default-kv_nvfp4_cast`
- Example docs and `examples/llm_ptq/hf_ptq.py --recipe` help text were
updated to use the new paths.

**Pre-commit and documentation**

- Recipe validation accepts `$import` entries and handles directory
recipes using `metadata.yaml`.
- The recipe validation hook skips `modelopt_recipes/configs/` because
those files are reusable snippets, not full recipes.
- `docs/source/guides/10_recipes.rst` now documents imports, schema
modelines, list append/splice semantics, built-in snippets, built-in
recipe paths, directory recipes, and the current recipe data model.

#### Backward compatibility

- Existing inline YAML recipes without `$import` continue to load.
- `modelopt.recipe.load_config()` remains public.
- The built-in recipe path renames are user-visible; callers should
update recipe path strings to the KV-first names listed above.

#### Testing

- `pytest tests/unit/recipe/test_loader.py -q` - 90 passed
- `python tools/precommit/check_modelopt_recipes.py ...` for the renamed
built-in PTQ recipes
- `pre-commit run mypy --files
modelopt/onnx/llm_export_utils/quantization_utils.py`
- `python -m py_compile examples/llm_ptq/hf_ptq.py`
- `git diff --check`

### Before your PR is "Ready for review"

- Is this change backward compatible?: Partially. Loader/API behavior is
compatible for existing inline YAML recipes, but built-in recipe path
names were renamed to KV-first paths.
- Did you write any new necessary tests?: Yes.
- Did you update Changelog?: Yes.

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-05-05 17:01:28 -07:00
h-guo18 9d2e6087d1 [Fix]: $HOME in launcher eagle example (#1365)
### What does this PR do?

Type of change: Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

<!-- Details about the change. -->

Launcher example bug raised by @cjluo-nv 

Before fix: task1 in
tools/launcher/examples/Qwen/Qwen3-8B/hf_online_eagle3.yaml fails
Reason: due to `HOME: /tmp` set in container, enroot credentials in
`$HOME/.config/enroot/.crendential` not found
```
GpuFreq=control_disabled
pyxis: importing docker image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc10
Apr 28 13:35:59.491365 2515157 slurmstepd   0x155552c3b780: error: pyxis: child 2515158 failed with error code: 1
Apr 28 13:35:59.491415 2515157 slurmstepd   0x155552c3b780: error: pyxis: failed to import docker image
Apr 28 13:35:59.491433 2515157 slurmstepd   0x155552c3b780: error: pyxis: printing enroot log file:
Apr 28 13:35:59.491453 2515157 slurmstepd   0x155552c3b780: error: pyxis:     [INFO] Querying registry for permission grant
Apr 28 13:35:59.491469 2515157 slurmstepd   0x155552c3b780: error: pyxis:     [INFO] Authenticating with user: <anonymous>
Apr 28 13:35:59.491483 2515157 slurmstepd   0x155552c3b780: error: pyxis:     [INFO] Authentication succeeded
Apr 28 13:35:59.491499 2515157 slurmstepd   0x155552c3b780: error: pyxis:     [INFO] Fetching image manifest list
Apr 28 13:35:59.491512 2515157 slurmstepd   0x155552c3b780: error: pyxis:     [INFO] Fetching image manifest
Apr 28 13:35:59.491524 2515157 slurmstepd   0x155552c3b780: error: pyxis:     [ERROR] URL https://registry-1.docker.io/v2/nvcr.io/nvidia/tensorrt-llm/release/manifests/1.3.0rc10 returned error code: 401 Unauthorized
Apr 28 13:35:59.491564 2515157 slurmstepd   0x155552c3b780: error: pyxis: couldn't start container
Apr 28 13:35:59.491579 2515157 slurmstepd   0x155552c3b780: error: spank: required plugin spank_pyxis.so: task_init() failed with rc=-1
Apr 28 13:35:59.491593 2515157 slurmstepd   0x155552c3b780: error: Failed to invoke spank plugin stack
Apr 28 13:35:59.515523 2515146 slurmstepd   0x155552c3b780: error: pyxis: child 2515240 failed with error code: 1
```

After fix:
```
GpuFreq=control_disabled
pyxis: importing docker image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc10
pyxis: imported docker image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc10
```

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Updated example pipeline to use the standardized dataset example path.
* Removed unnecessary per-task overrides of the process home and cache
directory to simplify environment setup.
* Preserved required model checkpoint environment setting for the
relevant task so model resolution continues to work.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-05-01 16:31:46 -07:00
Chenhan D. YuandClaude Opus 4.6 355c6b7883 fix: PTQ 1GPU, export PP divisibility, hidden states conversations key (#1293)
## Summary
- **megatron_lm_ptq.yaml**: Qwen3-8B PTQ to single GPU for L40 clusters
(TP=1, all tasks)
- **quantize.sh**: Auto-find largest PP dividing model's
`num_hidden_layers` for export step. Qwen3-8B has 36 layers which isn't
divisible by 8, causing `AssertionError` on 8-GPU nodes
- **compute_hidden_states_trtllm.py**: Use `messages` with
`conversations` fallback, matching the HF version. Fixes `KeyError:
'conversations'` when data uses OpenAI `messages` format

## Test plan
- [x] Qwen3-8B PTQ runs on single L40 GPU
- [x] Export PP auto-selects valid divisor (36 layers → PP=6 on 8 GPUs,
PP=4 on 4 GPUs, PP=1 on 1 GPU)
- [x] EAGLE3 offline pipeline reads data with `messages` field

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Dataset input handling now supports multiple field formats for
enhanced compatibility.

* **Bug Fixes**
* Optimized GPU resource allocation during model quantization with
improved pipeline parallelism computation.
* Updated quantization configuration for more efficient resource
utilization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 12:43:58 -07:00
Keval MorabiaandClaude Sonnet 4.6 3d0f0db49e [CI] Replace tox with nox, use nemo:26.04 for megatron tests, and simplify CI workflows (#1286)
### What does this PR do?

Type of change: New feature / infrastructure improvement

Follow-up to #1285 for correct CI test environment for megatron based
tests

Replaces `tox` + `tox-current-env` with `nox` for all test, lint, docs,
and wheel build sessions. The primary motivation was that
`tox-current-env` is incompatible with uv venvs in NGC containers (e.g.
NeMo's `/opt/venv`) — it picks the system Python via
`sys._base_executable` instead of the container's venv Python which has
megatron packages pre-installed.

Key changes:
- **`noxfile.py`** replaces `tox.ini` with GPU, CPU unit,
partial-install, pre-commit, docs, and wheel sessions
- **GPU sessions** use `venv_backend="none"` (run directly in container
env) and `python -m pip/pytest` to avoid PATH mismatches
- **uv** is set as the default venv backend (if available) for CPU
sessions (faster installs)

Also includes CI workflow simplifications:
- **`_pr_gate.yml`** new reusable workflow centralizing file-change
detection + linux-check wait logic (was duplicated across 3 workflow
files)
- **Collapsed pr/non-pr job pairs** into single jobs with conditional
`runs-on` in `gpu_tests.yml`, `example_tests.yml`,
`regression_tests.yml`
- **Collapsed `multi-py` / `multi-torch` / `multi-transformers`** into a
single `multi-version` matrix job in `unit_tests.yml`
- **PR path filtering** for unit test secondary jobs (multi-version,
launcher, partial-install) — skipped if no relevant files changed
- **Fixed schedule/workflow_dispatch skipping** — jobs with `needs:
[pr-gate]` were incorrectly skipped when all pr-gate internal jobs were
skipped; fixed by making the gate job always run
- **multi-version, launcher, partial-install** now also run on
`schedule` / `workflow_dispatch`

### Usage

```bash
python -m pip install nox uv                                                    # install nox and uv (once)
nox -l                                                                          # list all sessions
nox -s gpu_megatron                                                             # run a GPU session (inside container)
nox -s "unit-3.12(torch_211, tf_latest)"                                        # run a specific unit test combination
nox -s "unit-3.12(torch_211, tf_latest)" -R                                     # force-recreate venv (e.g. after dep changes)
COVERAGE_PROCESS_START=pyproject.toml nox -s "unit-3.12(torch_211, tf_latest)"  # with coverage
```

### Testing
- Ran `nox -l` to verify all session names
- Ran `gpu_megatron` session locally inside NeMo container — confirmed
it uses `/opt/venv/bin/python` correctly
- Manually triggered nightly-runs:
- Unit:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608013657
- GPU:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608018763
- Examples:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608017322

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A — CI infrastructure only
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ (added `nox`
and `uv` to `dev-test`, both Apache-2.0)
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — no user-facing changes

### Additional Information
Supersedes the tox-current-env workaround in the parent branch.

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-18 16:56:02 +00:00
Chenhan D. YuandClaude Opus 4.6 76b6fd51a5 fix: DFlash regression tests and vLLM server liveness (#1288)
## Summary
- **hf_online_dflash.yaml**: Add 100K-sample training config with
regression baselines (B200 loss curve),
`MAX_FINAL_LOSS`/`MIN_FINAL_ACC`/`MIN_ACCEPTANCE_LENGTH` thresholds,
vLLM nightly container for DFlash support
- **vllm_smoke_test.sh**: Parse acceptance length from vLLM server log
for regression check; `pip install pandas` workaround for broken nightly
container; capture server output to temp file
- **query.sh**: Detect vLLM server death during startup (PID liveness
check) + 600s timeout to prevent infinite polling that wastes GPU hours;
`pip install pandas` workaround
- Fix empty `environment:` key in DFlash YAML causing nemo_run
`ListParseError`

## Test plan
- [x] E2E pipeline passed on 8x B200 (training + vLLM smoke test + AR
eval)
- [x] Training regression: final loss 3.82 < 5.0, acc 0.20 > 0.15
- [x] vLLM acceptance length: 1.79 >= 1.4 threshold
- [x] AR evaluation: 2.02 overall on MT-Bench (8 categories)
- [x] Server liveness check prevents GPU waste on vLLM crash

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added optional regression validation for vLLM acceptance metrics
* Introduced configurable vLLM server startup timeout (default 600
seconds)

* **Improvements**
* Enhanced logging for vLLM server startup with progress tracking and
waited time reporting
* Faster detection of vLLM server process failures during initialization

* **Configuration Updates**
  * Increased training dataset size and logging granularity
  * Scaled tensor parallelism from 4 to 8 across multiple pipelines
  * Expanded PTQ quantization to multi-step pipeline
  * Added configurable training metric thresholds
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 11:41:17 +05:30