100 Commits
Author SHA1 Message Date
h-guo18andClaude Opus 5 c2aaa44f60 [Speculative Decoding] DFlash2 draft variant (grouped sublayer convolution + candidate selector) (#2216)
### What does this PR do?

Type of change: new feature

Adds **DFlash2** ([blog](https://inco.ai/blog/dflash2/)) as a draft
variant of the existing DFlash mode, selected with
`dflash_architecture_config.projector_type="dflash2"` alongside
`domino`, `dspark` and `lilicorr`.

DFlash2 keeps DFlash's one-pass parallel backbone and adds two
components that recover the acceptance a purely parallel draft loses:

- **Grouped dynamic depthwise convolution** around every attention and
MLP sublayer, giving each block position a view of its predecessors
*inside* the block. Taps do not cross the block boundary, so the draft
stays one forward pass.
- **Low-rank candidate selector** scoring transitions between adjacent
positions' top-k candidates, so serving walks one coherent path instead
of taking an independent argmax per position.

Both start as exact no-ops — the convolution's `base_kernel` is an
identity and `kernel_projection` is zeroed; the selector's
`successor_codebook` is zeroed — so a freshly built DFlash2 draft *is*
its DFlash backbone, and enabling the variant is an extension rather
than a perturbation. This matches the reference implementation
([SpecForge#772](https://github.com/sgl-project/SpecForge/pull/772),
merged) and the way `modeling_lilicorr` already installs this same
convolution class.

**This unblocks a recipe already shipped on `main`.**
`modeling_lilicorr._install_sublayer_convs` imports `DFlashGroupedConv`
from `modeling_dflash2`, so
`modelopt_recipes/general/speculative_decoding/lilicorr_conv.yaml`
raises at model build today and its CHANGELOG entry documents a feature
that cannot run. Landing this makes it runnable.

Module and parameter names match the SGLang/vLLM `DFlash2DraftModel`
loaders. Verified against the released `z-lab/Qwen3.8-27B-DFlash2`
checkpoint: **81 tensors, 21 name patterns, zero difference in either
direction**. The serving side,
[vllm-project/vllm#52816](https://github.com/vllm-project/vllm/pull/52816),
has since merged (`b389ac294`) with no change to the checkpoint
contract.

### Usage

```bash
python examples/speculative_decoding/main.py \
  --config modelopt_recipes/general/speculative_decoding/dflash2.yaml \
  model.model_name_or_path=Qwen/Qwen3-8B \
  data.data_path=<corpus>.jsonl \
  training.output_dir=<out>
```

```yaml
# modelopt_recipes/general/speculative_decoding/dflash2.yaml
dflash:
  dflash_selector_loss_alpha: 1.0      # weight of the candidate-selector CE term
  dflash_architecture_config:
    projector_type: dflash2
    conv_kernel_size: 2                # taps; must not exceed the block size
    conv_group_size: 16                # must divide hidden_size
    selector_rank: 256
    selector_top_k: 16
```

### Testing
<img width="2000" height="1320" alt="image"
src="https://github.com/user-attachments/assets/869004d1-92c6-41a4-a03d-fd824a06255c"
/>

**Unit** — 25 CPU tests in
`tests/unit/torch/speculative/plugins/test_hf_dflash2.py`; the full
`tests/unit/torch/speculative/` suite passes with no regressions. The
ones worth keeping pin invariants that a decreasing loss does not catch:

- the convolution is an exact identity on the **default** construction,
and its taps stay inside the block while a position still sees its
predecessors;
- the block-offset contract shared by the training objective and
`CandidateSelector.greedy_path` — a misaligned objective still
converges;
- the export fields the vLLM loader requires, including the top-level
`block_size` that `DFlash2Exporter` derives the nested copy from;
- which selector factors receive gradient on the first step.
`successor_codebook` starts at zero, so `predecessor_codebook` and
`hidden_projection` take one step to begin moving. That is a warm start,
not a dead branch, and both sides are asserted.

**End-to-end** — trained on Qwen3-8B against a plain DFlash control with
every other argument identical (plot above). Monotonic convergence, no
NaN/divergence, no DDP unused-parameter issues. Note the losses are
**not comparable** across arms: DFlash2's includes the selector CE term.

**Serving (vLLM)** — the exported drafter loads and drafts under the
merged DFlash2 path (`RESOLVED draft architectures:
['DFlash2DraftModel']`). Two notes for anyone reproducing: vLLM sizes
the convolution from `1 + num_speculative_tokens` at runtime rather than
from the checkpoint, so a `block_size=16` drafter is only correct at
`num_speculative_tokens=15`; and at that value the upstream path
currently hits an illegal memory access in `_cache_draft_logits`
([vllm#55279](https://github.com/vllm-project/vllm/issues/55279)),
independent of which checkpoint is used.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — additive. New
`projector_type`, its own registry and exporter, one new config field;
DFlash / Domino / DSpark / LiLiCorr numerics and `state_dict` contents
are untouched.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ —
`modeling_dflash2.py` is adapted from
[SpecForge#772](https://github.com/sgl-project/SpecForge/pull/772) and
carries its MIT notice. No new dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under `0.48.0`.
- Did you get Claude approval on this PR?: ✅ — run on 2026-08-20; all
review threads addressed and resolved.

### Additional Information

Rebased onto current `main`. Two commits from the original branch were
dropped because
[#2342](https://github.com/NVIDIA/Model-Optimizer/pull/2342) landed them
first, with authorship preserved: the no-op sublayer seam in
`modeling_dflash.py`, and the `rope_theta`/`rope_parameters` fix —
`main`'s version of the latter is stricter, so this PR no longer touches
`hf_dflash.py` at all.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added DFlash2 speculative decoding with grouped dynamic convolutions
and low-rank candidate selection.
* Added configurable selector-loss weighting, including an option to
disable it.
  * Added DFlash2 model conversion and export support.
  * Added checkpoints compatible with SGLang and vLLM DFlash2 serving.

* **Documentation**
* Added training recipes and a Qwen3-8B online DFlash2 training
configuration.

* **Tests**
* Added coverage for conversion, training, metrics, gradients, and
export compatibility.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-29 11:43:02 +08:00
Shengliang XuandClaude Opus 5 d0142c9dca Reuse a whole recipe via $import, deprecate recipe_type, and start the published-checkpoint backfill with two aliases (#2376)
### What does this PR do?

**Type of change:** New feature (recipe loading) + one bug fix

Two things, the second built on the first:

1. **A recipe can now reuse another recipe whole.** A top-level
`$import` brings in the imported recipe's entire body; keys given
alongside it override the imported ones. `metadata.recipe_type` becomes
optional and is deprecated along the way.
2. **The deprecated `recipe_type` is swept out of every shipped recipe,
and the checkpoint backfill starts with two published checkpoints
recorded as aliases** that reuse a portable recipe wholesale — the first
users of the alias mechanism — plus a fix to two existing Nemotron NVFP4
recipes.

#### Declaring what kind of recipe a file is

`load_recipe` read `metadata.recipe_type` out of the raw YAML *before*
resolving imports, because it needs the schema class to hand to
`load_config`. That made the field impossible to inherit, so a recipe
reusing another had to restate a line it could only have copied.

It is now optional, and the loader takes the first of these that
answers:

1. a `# modelopt-schema:` comment naming the recipe's schema class,
2. `metadata.recipe_type` — **deprecated**; still read and still
honoured, so a recipe outside this repo keeps working unchanged,
3. the recipe it delegates to via a top-level `$import`.

Whatever a recipe *does* state must be true, in both directions. A
schema comment contradicting a `recipe_type` is rejected, and so is a
recipe importing a different kind of recipe — that used to surface as
whatever pydantic made of, say, an `eagle` section spliced into a PTQ
schema. The concrete recipe classes carry a `RECIPE_TYPE` ClassVar as
the single source of truth.

Only a recipe that another file **imports** needs the schema comment —
that is what `$import` resolution requires to validate the payload. The
sweep here drops `metadata.recipe_type` from all 78 shipped recipes that
carried it and gives the imported ones a `# modelopt-schema:` comment
instead, so nothing in-tree depends on the deprecated field.

A directory recipe's `metadata.yml` resolves its kind the same way —
schema comment first, `recipe_type` as the fallback — it just has no
`$import` to delegate through, since a directory recipe has no body of
its own to hand off. (Follow-up commit, after this PR's initial review:
it originally still required `recipe_type` unconditionally, the one
place the deprecation didn't reach.)

#### Checkpoint aliases

Two checkpoints NVIDIA has published in quantized form use a scheme a
portable recipe already produces, with no checkpoint-specific deviation,
so each is recorded as a thin **alias** (top-level `$import`, overriding
only `metadata`) at its own model-hub path -- the *source* checkpoint's
path, not the published quantized one's:

-
**`models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast`**
delegates to `general/ptq/nvfp4_experts_only_mse-kv_fp8_cast` —
expert-only NVFP4 (MSE static weights, dynamic inputs) with an FP8 KV
cache in cast mode — published as `nvidia/Kimi-K2.6-NVFP4`.
-
**`models/Qwen/Qwen3.5-397B-A17B/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8`**
delegates to the `qwen3_5_moe` architecture recipe
`model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4
(MSE static weights) on the routed experts, ModelOpt-default FP8
elsewhere, FP8 KV cache — published as
`nvidia/Qwen3.5-397B-A17B-NVFP4-V2`.

(Follow-up commit, after this PR's initial review: the Qwen entry
originally lived at `models/nvidia/Qwen3.5-397B-A17B/` -- nvidia is the
*published* checkpoint's org, not Qwen3.5-397B-A17B's own. Moved to
match the source model's actual hub path, same as the Kimi-K2.6 entry
above.)

Editing the base recipe changes every alias that points at it; nothing
is duplicated.

#### One fix

- **The Nemotron-3 Super and Ultra NVFP4 recipes** quantized the MTP
block on the **Megatron-Core** path, where it is a live `model.mtp`
submodule their broad `*mixer.*` patterns matched into, contrary to
their own descriptions. They now disable `mtp.*` explicitly. Hugging
Face runs were unaffected — `NemotronHPreTrainedModel` sets
`_keys_to_ignore_on_load_unexpected = [r"mtp.*"]` and builds no MTP
module.

### Usage

A checkpoint alias resolves through `--recipe` to the recipe it
delegates to:

```bash
python examples/hf_ptq/hf_ptq.py \
    --pyt_ckpt_path <checkpoint> \
    --recipe models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast \
    --export_path <output>
```

A recipe that reuses another whole — the shape the aliases use:

```yaml
imports:
  base: general/ptq/nvfp4_experts_only_mse-kv_fp8_cast

$import: base
metadata:
  description: What this checkpoint uses the base recipe for.
```

### Testing

- **`tests/unit/recipe/test_loader.py`** — 28 new cases covering
whole-recipe reuse with no `metadata` at all; kind resolution from each
of the three sources, from a delegation chain and from a `$import` list;
a delegation cycle failing with `ValueError` rather than recursing;
`peek_declared_schema` including a comment placed below the first YAML
line; `recipe_type` being optional, filled per class, and rejected when
it contradicts; a directory recipe resolving its kind from a schema
comment the same way, rejecting a comment/`recipe_type` disagreement,
and still requiring one or the other; and delegating across kinds being
an error.
- **`tests/unit/recipe/test_recipe_docs.py`** — the
model-specific-recipe check now also covers the two new alias folders,
which must be listed in `ptq.md` like every other
`models/<org>/<model_id>` entry.
- **Recipe validation** (`tools/precommit/check_modelopt_recipes.py`)
and **`pre-commit`** pass on the changed files. The full
`tests/unit/recipe/` suite is left to CI — a broken `transformer_engine`
in the local dev venv keeps the `mtq.quantize`-based cases from running
there.

Not covered: **numerics**. Nothing here asserts accuracy, or that
running one of these recipes reproduces a released checkpoint's weights.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `metadata.recipe_type` is
still read and honoured for recipes outside this repo, the schema
comments are inert for direct loads, and the loader change only relaxes
a check.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — two feature entries, one deprecation, and one bug fix under 0.48.0.
- Did you get Claude approval on this PR?: ❌ — not yet run.



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Recipes can delegate configurations, support checkpoint aliases, and
apply local metadata overrides.
* Recipe types can be inferred from schema declarations or delegated
recipes, with stronger consistency validation.
* Added unquantized KV-cache options, layerwise export, broader operator
calibration, and new PTQ examples.
  * Added checkpoint-specific recipes and MLflow experiment references.

* **Bug Fixes**
* Improved ONNX calibration, FSDP2 export, and fused-MoE quantization
handling.
  * Nemotron-3 recipes keep MTP blocks in BF16.

* **Documentation**
* Expanded guidance for aliases, delegation, schema declarations, and
recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 17:12:31 -07:00
Jenny Chen c4d00f7150 Move Nemotron Nano offline KD example (#2471)
### What does this PR do?

Type of change: Bug fix

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Very short summary of changes only for new features,
backward breaking changes, deprecations, or fixes for critical bugs
present in previous releases. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated the NVIDIA Nemotron offline usage example to use the BF16
source-directory pipeline path.
  * Pipeline configuration remains unchanged.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
2026-09-19 00:13:16 +05:30
yeyu-nvidiaandClaude Opus 5 216f28a6e0 Consolidate speculative-decoding agent skills into one stage/algorithm tree (#2201)
### What does this PR do?

Type of change: documentation

Reorganizes the EAGLE3 agent skills into a single speculative-decoding
skill, then adds algorithm sheets for DFlash, DSpark, and Domino.

**The problem.** The four `eagle3-*` skills each baked the algorithm
into a *stage* of the same draft-model pipeline:

```
skills/eagle3-new-model/    skills/eagle3-review-logs/
skills/eagle3-triage/       skills/eagle3-validate/
```

Adding DFlash would have meant four more near-duplicate skills, since
the stages are shared and only the algorithm differs.

**The change.** One skill dir shaped like `ptq/` (SKILL.md +
references/), split along the two real axes:

```
plugins/modelopt/skills/speculative-decoding/
├── SKILL.md                      # router: stage table x algorithm table
└── references/
    ├── stages/                   # the procedure — algorithm-independent
    │   ├── configure.md          # <- eagle3-new-model
    │   ├── review-logs.md        # <- eagle3-review-logs
    │   ├── triage.md             # <- eagle3-triage
    │   └── validate.md           # <- eagle3-validate
    └── algorithms/               # the data sheet — per-algorithm
        ├── README.md             # contract: 6 required sections
        ├── eagle3.md
        ├── dflash.md
        ├── dspark.md             # DFlash variant — delta only
        └── domino.md             # DFlash variant — delta only
```

Stage docs cite algorithm-sheet sections by heading (*Pipeline tasks*,
*Success markers*, *Quality gate*, *Known failures*, ...), so a new
algorithm means one new file plus a table row — no stage edits. Every
recipe in `modelopt_recipes/general/speculative_decoding/` now has a
sheet.

DSpark and Domino are documented as **DFlash variants**, not separate
pipelines: same `recipe_type: speculative_dflash`, same training script,
same `dflash.*` config namespace, selected by
`dflash_architecture_config.projector_type`. Their sheets carry only the
delta.

Writing the sheets surfaced three things the old EAGLE3-only skills got
wrong or missed:

- **Task counts are not fixed.** The old skills hardcoded "4-step
pipeline, task_0 through task_3". DFlash offline is 2 tasks, DFlash
online is 3, Domino is 2. The stage docs no longer assume a count.
- **`--aux-layers` couples the dump to the draft.** For DFlash the
dump's layer count must equal the draft's `num_hidden_layers`; a
mismatch doesn't error, it silently captures the wrong layers. Recorded
under *Known failures*.
- **In-training AR is meaningless for DSpark and Domino.** Both recipes
pin `estimate_ar: false` / `ar_validate_steps: 0` because eval runs the
DFlash backbone with the new head bypassed. Each sheet says so under
*Quality gate* so nobody reads a backbone-only number as a result.

**Behavior change:** the four `/eagle3-*` slash commands are replaced by
one `/speculative-decoding`. This isn't optional —
`tools/precommit/sync_claude_skills.sh` iterates `.agents/skills/*/` one
level deep and plugin discovery is `skills/<name>/SKILL.md`, so a
directory is either one skill or a container of skills, not both.
`tools/launcher/docs/claude_code.md` is updated accordingly.

### Usage

```
/speculative-decoding
```

Or by description — the skill triggers on EAGLE3 / DFlash / DSpark /
draft model / acceptance rate. For a new model, follow the stages in
order:

```bash
# 1. Configure: copy the closest examples/<Org>/<Model>/hf_<mode>_<algo>.yaml and adapt
cd tools/launcher
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml --dryrun   # preview
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml --yes      # submit
# 2. review-logs -> 3. triage (if anything failed) -> 4. validate
```

### Testing

- `claude plugin validate . --strict` and `claude plugin validate
plugins/modelopt --strict` — both pass
- `pre-commit run --files ...` over all changed files — passes,
including `markdownlint-cli2` and the `sync-claude-skills` symlink hook
(it agrees with the new `.claude/skills/speculative-decoding` symlink)
- Verified the new skill is discovered and its description loads
- Every relative link across the skill tree resolves; every repo path
cited in the sheets exists; no dangling `eagle3-*` reference remains
anywhere in the repo
- Each factual claim in the sheets was checked against its source — the
launcher example YAMLs, the four recipes, `dflash_online_training.sh`,
`vllm_smoke_test.sh`, `check_regression.py`, and
`plugins/hf_{dflash,dspark,domino}.py` — rather than written from memory

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — the four `/eagle3-*` slash
commands become `/speculative-decoding`. Agent tooling only; no library
or API surface is touched. The three YAML comment fixes are
comment-only, no behavior change.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — documentation; covered
by plugin validation and pre-commit
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — agent tooling only, matching how #2025 handled it
- Did you get Claude approval on this PR?: ✅ — `/claude review` run; 2
findings, both fixed in `9b3c568`

### Additional Information

Follows #2025, which moved the skill tree into the installable plugin.

Two stale in-repo comments were found while sourcing the sheets, and are
**fixed in this PR** (`c2696b1`, comment-only):

1. `modelopt_recipes/general/speculative_decoding/dflash.yaml` pointed
`chat_template` at a `chat_templates/` directory under
`modelopt_recipes` that does not exist — templates live per-model beside
each launcher example.
2. Both offline DFlash example YAMLs annotated `--aux-layers dflash`
with "Must match the draft model's num_hidden_layers". `--aux-layers` is
a preset keyword accepting only `eagle`, `dflash`, or an explicit id
list, so it carries no count. The constraint is real but belongs to the
draft depth the preset resolves to: `--num-draft-layers` on the vLLM
dump, and no override at all on the HF/TRT-LLM dumps, which hardcode 5
via `resolve_aux_layers`. This comment had already misled this PR's own
first draft, which is why it's fixed rather than just documented.

Because of (1), this PR now touches `modelopt_recipes/`, which adds
**@NVIDIA/modelopt-recipes-codeowners** to the required reviewers.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added comprehensive speculative-decoding guidance for configuration,
training, validation, troubleshooting, and supported algorithms.
* Added workflow references for DFlash, Domino, DSpark, and EAGLE3,
including quality checks and failure diagnosis.

* **Documentation**
* Generalized experiment-log review and pipeline triage across
algorithms.
* Clarified DFlash draft-depth configuration, resource sizing, task
recovery, and validation.
* Replaced the EAGLE3-specific workflow entry with the broader
speculative-decoding workflow.
* Removed standalone EAGLE3 skill documentation as guidance is now
consolidated under speculative decoding.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 11:09:14 -07:00
yeyu-nvidiaandClaude Opus 4.6 8025a3dc54 specdec(recipe): add MiniMax-M2.7-DFlash streaming multi-node pipeline (#1835)
## Summary

- Add `hf_streaming_dflash_multi_node.yaml` for MiniMax-M2.7 (229B MoE)
streaming DFlash training
- 2 serve replicas (TP=4, whole node) + 2 trainer nodes (4 GPU each)
over NIXL RDMA hidden-state transport
- Capture IDs `[2,17,32,47,62,64]` from `build_target_layer_ids(64, 5)`
+ final layer output
- MiniMax-specific: trust_remote_code, FSDP2 via accelerate config,
mask_token=200054, YaRN rope_scaling factor=48
- Topology matches Kimi-K2.5 large-MoE streaming recipe

Resolves OMNIML-5221

## Test plan

- [ ] Dry-run validation (`uv run launch.py --yaml ... --dry-run`)
- [ ] Server-only smoke on CW-DFW (task_1 with `training.max_steps=1`)
- [ ] Full streaming training run

Signed-off-by: Ye Yu <yeyu@nvidia.com>

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a multi-node launcher configuration for MiniMax-M2.7 streaming
training with speculative decoding.
* Added an end-to-end workflow for conversational data preparation,
distributed training, checkpoint export, and vLLM smoke testing.
* Supports configurable serving and training settings across multiple
nodes.

* **Enhancements**
* Applies Transformers version overrides only to trainer or single-node
environments.
* Preserves the serving environment’s Transformers version on dedicated
serving nodes.
* Uses the launcher’s built-in bfloat16 mixed-precision configuration
for training.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-09-16 12:54:57 -07:00
2d35643452 LiLiCorr training (#2342)
### What does this PR do?

Type of change: new feature

Adds **LiLiCorr**, a candidate-lattice reranker for DFlash drafts, as a
new `projector_type` on the
existing `dflash` mode — plus three DFlash-wide improvements that apply
to every variant, and an
optional composition with DFlash2's grouped convolutions.

A DFlash drafter is trained on per-position marginals rather than on the
joint block distribution,
so its drafted tokens are individually plausible yet jointly incoherent.
LiLiCorr keeps the top-`k`
candidates the backbone already produces at each block position, scores
transitions between adjacent
candidates with a small two-layer transformer, and commits a path
through the lattice greedily.
Serving is unchanged in kind: verify still checks every drafted token
against the target, so the
emitted distribution is untouched and only acceptance length moves.

- Paper: [LiLiCorr: Lightweight Likelihood Correlation of Parallel
Drafts for Speculative Decoding](https://arxiv.org/abs/2608.20530)
(arXiv:2608.20530)
- Blog: https://research.nvidia.com/labs/nemotron/lilicorr/
- **Companion PR — serving support:**
[sgl-project/sglang#37462](https://github.com/sgl-project/sglang/pull/37462)

This PR is the **training** half. It trains the drafters and exports
them; the companion PR above is
what serves the resulting checkpoints, and is what the comparison table
below was measured through.

**What is in the commits**

| | |
| --- | --- |
| LiLiCorr draft variant | `hf_lilicorr.py`, `modeling_lilicorr.py`,
conversion routing, config fields, export |
| Three DFlash-wide features | fp32 master weights for the draft, draft
activation checkpointing, and a DDP hang fix — all default-off or
behaviour-preserving, all applying to `dflash`, `domino`, `dspark` and
`dflash2` alike |
| Optional grouped convolutions | composes LiLiCorr with DFlash2's
`DFlashGroupedConv`; see the dependency note below |
| Two recipes | `lilicorr.yaml` and `lilicorr_conv.yaml` |
| CPU unit tests, CHANGELOG, one launcher example | |

**⚠️ The convolutions depend on the DFlash2 branch, and cannot run until
it merges.**

`modeling_lilicorr.py` imports `DFlashGroupedConv` from
`modeling_dflash2`, which today exists only
on `haoguo/dflash2-support`. The class is **imported rather than copied
on purpose** — it is the only
way the two variants cannot drift apart arithmetically — but the
consequence is that the
convolutional recipe cannot run against `main` as it stands.

So the import is **deferred into `_install_sublayer_convs`** rather than
taken at module scope.
Everything else in this PR, including the plain LiLiCorr reranker, has
no DFlash2 dependency at all
and works on `main` today; an eager import would have made the whole
plugin unimportable for the sake
of one optional feature. Requesting the convolutions without DFlash2
present raises an `ImportError`
naming the two config keys to remove, rather than failing at import
time.

**This PR carries two of @h-guo18's commits, with authorship and
sign-off preserved.** Both are
independent of DFlash2 itself and both are needed here:

- `1419d47e`, the no-op sublayer seam. Without it
`DFlashDecoderLayer.forward` never calls the
wrappers the convolutions install onto, so the modules would be built,
counted and exported while
  computing nothing. It is arithmetically an identity on its own.
- `ba377e7a`, the RoPE-θ fix. On Transformers 5 a config carries both a
top-level `rope_theta` and a
`rope_parameters` dict; the real base lives in the dict while the class
default (10,000 for Qwen3)
stays visible as the flat attribute. Reading the flat field first builds
a draft whose RoPE base is
100× off a Qwen3-8B target's, which trains and exports without
complaint. Both the training-side
enforcement and the exporter's `_get_rope_theta` are affected on `main`
today.

Both are @h-guo18's work and belong to their branches; they are carried
here only so that this PR
stands on its own. **If those branches land first, this PR can be
rebased onto them and the two
commits dropped**, and they can equally be split out now if that is
easier to review.

The same applies to `dflash_fp32_master_weights`, which is also in
flight on
`haoguo/dflash-fp32-master-weights`. The field name is shared
deliberately so that there is only
ever one knob rather than two spellings of it, and both versions default
to off. Whichever lands
first, this PR can be rebased onto it.

### Usage

Train with the shipped recipe:

```python
from modelopt.recipe import load_recipe

config = load_recipe("general/speculative_decoding/lilicorr.yaml")
# Qwen3-8B target, 6 epochs, block size 16 (15 drafted slots, 16 verified),
# DFlash decay objective at gamma 7.0, fp32 master weights for the draft.
```

Or convert directly:

```python
import modelopt.torch.speculative as mtsp

config = {
    "dflash_block_size": 16,
    "dflash_loss_objective": "decay",
    "dflash_loss_decay_factor": 7.0,
    "dflash_fp32_master_weights": True,
    "dflash_lilicorr_w_ce": 0.25,
    "dflash_lilicorr_w_margin": 0.0,
    "dflash_lilicorr_w_pen": 0.25,
    "dflash_architecture_config": {
        "num_hidden_layers": 5,
        "projector_type": "lilicorr",
        "lilicorr_candidate_topk": 8,
        # Optional, and all-or-nothing: adding these two keys wraps every draft
        # sublayer in DFlash2's grouped convolution. Requires the DFlash2 variant.
        # "conv_kernel_size": 2,
        # "conv_group_size": 16,
    },
}
mtsp.convert(model, [("dflash", config)])
```

### Results

Six drafters for a **Qwen3-8B** target, all trained **in ModelOpt on one
matched contract** — the
same corpus, schedule and block geometry for every arm, so no row
carries a training advantage.
Training data is NVIDIA's
[Nemotron Post-Training Dataset
v2](https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2)
with the multilingual split excluded, generated from the target with
**thinking disabled**;
**6 epochs**; block size 16 (15 drafted slots, 16 verified); DFlash
decay objective at gamma 7;
**8 nodes × 8 H100, global batch size 64** (one sequence per device, no
gradient accumulation).

All six were then exported and served through SGLang on a **single H100
80GB**, `tp_size 1`, at
concurrency 1, greedy, `fa3`, mean of two replicates, with the whole
node held exclusive per
benchmark. Speedup is output tokens/s against an autoregressive baseline
measured in the same
allocation.

Cells are `acceptance length / speedup-vs-AR`; **★ fastest, ☆ second
fastest**:

| benchmark | LiLiCorr+conv | LiLiCorr | DSpark | DFlash2 | Domino |
DFlash |
|---|---|---|---|---|---|---|
| gsm8k | ★ 7.715 / 5.26x | ☆ 7.557 / 5.22x | 7.375 / 4.86x | 7.252 /
5.06x | 7.225 / 4.87x | 6.341 / 4.59x |
| math500 | ★ 9.241 / 6.54x | ☆ 9.064 / 6.52x | 9.012 / 6.15x | 8.999 /
6.49x | 8.976 / 6.25x | 7.909 / 5.88x |
| aime25 | ★ 8.285 / 6.03x | ☆ 8.156 / 6.03x | 8.043 / 5.61x | 7.967 /
5.91x | 8.066 / 5.77x | 7.126 / 5.44x |
| humaneval | ★ 7.393 / 4.01x | 7.077 / 3.93x | 7.163 / 3.72x | ☆ 7.081
/ 3.95x | 6.864 / 3.73x | 6.156 / 3.68x |
| mbpp_sanitized | ★ 5.999 / 4.18x | ☆ 5.849 / 4.13x | 5.888 / 3.95x |
5.685 / 4.05x | 5.679 / 3.91x | 5.027 / 3.70x |
| livecodebench | ★ 7.975 / 5.40x | ☆ 7.754 / 5.33x | 7.775 / 5.10x |
7.601 / 5.26x | 7.553 / 5.04x | 6.808 / 4.88x |
| alpaca_eval | ☆ 3.697 / 2.69x | ★ 3.656 / 2.70x | 3.588 / 2.52x |
3.467 / 2.58x | 3.627 / 2.59x | 3.222 / 2.46x |
| mtbench | ★ 4.014 / 2.94x | ☆ 3.939 / 2.93x | 3.957 / 2.78x | 3.748 /
2.80x | 3.948 / 2.84x | 3.478 / 2.67x |

**Against every other approach in the table, LiLiCorr with convolutions
is the fastest on all eight
benchmarks.** Plain LiLiCorr is the fastest on seven of the eight; the
exception is humaneval, a
164-prompt slice, where DFlash2 is ahead by 0.5%.

`DFlash` is the deliberately head-free control; every head clears it by
+7.60% to +21.67% on
acceptance, which is the check that a head actually loaded. Reproducing
the `LiLiCorr+conv` column
additionally needs the DFlash2 variant.

Acceptance length is bit-reproducible under greedy decoding and its
replicate spread here was 0.00%
on every benchmark; throughput has a ~0.2% floor.

### What `dflash_fp32_master_weights` does, and what it is worth

Today the draft is cast to the frozen base model's dtype — bf16 — before
the optimizer is built.
AdamW then allocates its moments with `zeros_like(p)`, so the
**optimizer state becomes bf16 too**.
That is the problem: bf16 has too few mantissa bits to represent the
small updates Adam's second
moment accumulates, so those updates round away and the effective step
size decays on its own,
independently of the learning-rate schedule.

The flag is standard mixed precision instead: the draft's master weights
stay in fp32 while the
matmuls run in bf16. It requires a bf16 autocast around the forward,
which HF `Trainer` supplies
under `TrainingArguments.bf16`. Paths that do not go through the Trainer
— evaluation,
`pseudo_speculative_generate`, a plain `convert()` and forward —
currently need the caller to
supply it, and no shipped recipe exercises those (`estimate_ar: false`,
`do_eval: false`). Making
the draft supply its own autocast is a follow-up, held back from here on
review because it touches
every DFlash variant and wants e2e coverage of the existing recipes.

Compute speed is unchanged. The cost is memory, about 12 bytes per
parameter for the weight plus
Adam's two moments instead of 6, plus a doubled gradient all-reduce
under DDP, since fp32
parameters mean fp32 gradients. Under FSDP2 that second cost is what
`MixedPrecisionPolicy(reduce_dtype=...)` exists to control.

It is worth **7 to 14 percent of acceptance length**, measured at the
end of training on gsm8k, and
it helps every projector type:

| arm | bf16 | fp32 | Δ acceptance length |
| --- | ---: | ---: | ---: |
| LiLiCorr | 6.8670 | 7.5573 | **+10.05%** |
| DFlash2 | 6.7396 | 7.2518 | **+7.60%** |
| Domino | 6.5854 | 7.2252 | **+9.71%** |
| DSpark | 6.4621 | 7.3752 | **+14.13%** |
| DFlash | 5.9030 | 6.3412 | **+7.42%** |

Every arm in the comparison table above was trained with it on, and
**both shipped recipes set it
`true`**, so the documented path gets it.

It defaults to **off**, so no existing DFlash, Domino or DSpark run
changes behaviour. Both shipped
LiLiCorr recipes set it `true`, which is the arithmetic their numbers
were trained with. Flipping
the default is a reasonable follow-up once the autocast above is in.

The draft is drawn in fp32 and, under this flag, kept there; an
unpromoted run rounds the same draw
to the base model's dtype. So the bf16 and fp32 rows of the table above
start from the same
initialization at the precision each trains in, rather than from two
different draws. A unit test
pins that.

The flag also survives a resume. `modify()` runs under `from_pretrained`
with the base model still
on meta and cannot place the draft at all, so `restore_draft_precision`
re-applies the dtype, the
device and the rotary buffer once the weights are loaded and before the
Trainer builds the
optimizer — the last point that can still decide the Adam moment dtype.
It also reloads the draft's
tensors at the dtype they were saved in, since checkpoints store the
draft in fp32 while the base is
bf16 and `dtype="auto"` gives every tensor one dtype.

@h-guo18 has the same field in flight on
`haoguo/dflash-fp32-master-weights`, plus an
HF-format-resume fix this PR does not have. The name is shared
deliberately so there is only ever
one knob; whichever lands first, the other should be dropped rather than
merged.

### Testing

- **257 CPU unit tests pass** across `tests/unit/torch/speculative/`,
including the existing DFlash,
Domino, DSpark and Eagle suites. 48 of them are new and cover LiLiCorr
specifically: conversion
routing, head geometry, the required-field validation, the three-term
objective and its absolute
weights, gradient reach into both the head and the drafter body, and the
export contract.
- Both recipes load and validate through `modelopt.recipe.load_recipe`.
- The three DFlash-wide changes are covered behaviourally: the fp32 flag
is checked on the
optimizer's moment dtypes rather than only on parameters, since the
moments are the point of the
change, and on the initialization described above; activation
checkpointing is asserted to leave
draft gradients bit-identical with the flag on and off; and the rotary
buffer is asserted present
after `modify()` on a real device while still deferred on meta, which is
the case the laziness
  existed for.
- The resume path has its own test: after a `save_pretrained` /
`from_pretrained` round trip,
`restore_draft_precision` is asserted to return the draft to fp32 with
its stored weights intact
and its Adam moments in fp32. Without it the draft comes back in the
base dtype with the flag
  still set, which is the failure it exists to prevent.
- `TestDFlashLazyRotaryEmb` was updated rather than left passing: it
asserted the rotary buffer does
*not* exist after convert, and the DDP fix deliberately changes that on
non-meta devices. The
  replacement pins the refined invariant in both directions.
- The published checkpoints were trained with this arithmetic, verified
rather than assumed: a
fingerprint over draft initialisation, loss and gradients is compared
against the pre-review tree
for both `dflash` and `lilicorr`. Loss and gradients are **bitwise
identical**. Initialisation
moves, by less than bf16 resolution, and that is the single-dtype change
described above.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — every addition is opt-in. The
new `projector_type` is
selected only by config, `dflash_fp32_master_weights` defaults to off,
and the
activation-checkpointing and DDP fixes preserve behaviour. No existing
default changes.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in
`CONTRIBUTING.md`: ✅ — no new dependencies. Four files carry `# Adapted
from
https://github.com/sgl-project/SpecForge/...` headers for the DFlash
backbone and loss they derive
from (Apache-2.0), matching the attribution already on `hf_dflash.py` in
this repo. The two
commits described above are @h-guo18's, cherry-picked with authorship
and sign-off preserved.
- Did you write any new necessary tests?: ✅ — 48 new CPU tests, plus the
updated rotary test.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — will run `/claude review`
once opened.

### Additional Information

The convolutional recipe is the memory worst case: at an 8B target,
combined with fp32 master
weights, it may need `training.gradient_checkpointing: true` to fit on
80 GiB, and it fits without at
4B. Checkpointing is mathematically neutral — same objective, same data
order, same resulting model —
but it trades step time for memory, so a run using it is not
step-time-comparable with one that does
not. The recipe header says so.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added LiLiCorr speculative decoding with candidate-lattice reranking,
configurable objectives, metrics, export support, and optional grouped
convolutions.
* Added FP32 master-weight support with improved mixed-precision
behavior and gradient checkpointing.
* Added LiLiCorr training recipes and a Qwen3-8B launcher configuration.
* **Bug Fixes**
* Improved rotary-embedding configuration handling and corrected DFlash
distributed-training hangs.
* Added validation for invalid LiLiCorr configurations and improved
exported reranking metadata.
* **Documentation**
* Expanded guidance for FP32 master weights, training workflows, and
LiLiCorr configuration.
* **Tests**
* Expanded coverage across training, evaluation, generation, export, and
checkpoint workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: mrusanovsky <mrusanovsky@nvidia.com>
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 00:35:31 +08:00
Shengliang Xu de3eda8a11 Restructure recipes: split per-model_type recipes from model-hub checkpoint recipes (#2219)
### What does this PR do?

**Type of change:** Refactor (recipe-library layout) + documentation —
backward-breaking for saved `--recipe` paths.

Separate the two kinds of built-in Hugging Face recipes that were
previously mixed under `modelopt_recipes/huggingface/`:

- **`huggingface/<model_type>/`** — architecture recipes keyed by the
transformers `model_type`; one recipe covers every checkpoint of that
architecture. **Unchanged.**
- **`models/<org>/<model_id>/`** — a *new top-level tier* for recipes
that mirror one specific published checkpoint, keyed by its **model-hub
path** (as on the Hugging Face Hub, ModelScope, etc.) so the on-disk
path equals the hub path.

Concretely, the model-instance recipes move out of `huggingface/` to the
top level:

- `huggingface/models/mistralai/…`, `huggingface/models/nvidia/…` →
`models/mistralai/…`, `models/nvidia/…`
- `huggingface/step3p5/Step3.5-Flash/…` →
`models/stepfun-ai/Step-3.5-Flash/…` (re-keyed to the canonical HF repo
id
[`stepfun-ai/Step-3.5-Flash`](https://huggingface.co/stepfun-ai/Step-3.5-Flash)
— org `step3p5`→`stepfun-ai`, id `Step3.5-Flash`→`Step-3.5-Flash`)

**Why:** `modelopt_recipes/README.md` already documented a top-level
`models/` tier, but the files lived under `huggingface/models/` and
instance-specific recipes were awkwardly nested under the
per-`model_type` tree. This aligns the filesystem with the documented
layout and makes the instance tier hub-addressable — given a checkpoint
id you can find (or place) its recipe with no lookup table.
`load_recipe` resolves paths directly under `modelopt_recipes/`, so a
top-level `models/` sibling of `general/` and `huggingface/` works
identically.

The move is metadata-only — all recipe YAML content is byte-identical
(`R100` renames). Everything else is updating references (nvidia
launcher YAMLs, `test_loader.py`) and docs: a new `models/README.md`,
plus `huggingface/README.md`, root `README.md`, `ptq.md`, and the
`10_recipes.rst` guide, which no longer describe instances under
`huggingface/`.

### Usage

Recipe paths for the moved checkpoint recipes lose the `huggingface/`
prefix (and Step 3.5 Flash is keyed by its hub id):

```python
from modelopt.recipe import load_recipe

# before
load_recipe("huggingface/models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16")
load_recipe("huggingface/step3p5/Step3.5-Flash/ptq/nvfp4-mlp-only")

# after
load_recipe("models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16")
load_recipe("models/stepfun-ai/Step-3.5-Flash/ptq/nvfp4-mlp-only")
```

The same rename applies to `--recipe …` CLI values and launcher
`QUANT_CFG:` entries. Architecture recipes under
`huggingface/<model_type>/` are unaffected.

### Testing

- **Recipe resolution (torch-free):** parsed every recipe under
`models/` and confirmed all `$import` targets resolve against the recipe
root — 0 dangling across the tier.
- **Docs consistency:** re-ran the
`tests/unit/recipe/test_recipe_docs.py` logic; it now globs both
`huggingface/` and `models/`, and every model dir (incl.
`Step-3.5-Flash`, `Nemotron-3-Nano-4B-BF16`, …) plus every `general/ptq`
recipe is still mentioned in `ptq.md`.
- **Reference sweep:** repo-wide grep confirms no remaining references
to the old paths outside the intentional historical CHANGELOG entries
(released 0.44 / 0.45).
- **pre-commit:** `markdownlint-cli2`, license-insert, and `bandit`
hooks pass on the changed files.
- Note: the full `pytest` suite was not run in my environment (no
`torch`), so `test_recipe_docs.py` / `test_loader.py` should be
exercised in CI.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — `--recipe` / `load_recipe`
paths for the checkpoint-mirror tier change (drop the `huggingface/`
prefix; `step3p5/Step3.5-Flash` → `stepfun-ai/Step-3.5-Flash`).
Documented as a Backward Breaking Change in `CHANGELOG.rst` (0.47); the
only *released* old paths affected shipped in 0.45. A clean break was
chosen over a symlink or loader-alias shim.
- If you copied code from any other sources or added a new PIP
dependency …: N/A
- Did you write any new necessary tests?: ✅ — updated
`test_recipe_docs.py` to also glob the top-level `models/` tier so
instance recipes stay covered by the doc-consistency check.
- Did you update Changelog?: ✅ — added a 0.47 **Backward Breaking
Changes** entry.
- Did you get Claude approval on this PR?: ❌ <!-- run /claude review -->

### Additional Information

Design note: an earlier iteration nested everything under
`huggingface/model_type/` + `huggingface/models/`; the final layout
keeps `huggingface/` flat (per-`model_type`) and lifts instances to a
top-level `models/` tier, matching what `modelopt_recipes/README.md`
already documented. The `Step3p5*` architecture class names (from the
model's `trust_remote_code` modeling code) are unrelated to the recipe
path and are left unchanged.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added checkpoint-specific PTQ recipes for Kimi-K3, Mistral Medium 3.5,
and NVIDIA Nemotron models.
  * Added a Nemotron speculative-decoding warm-start recipe.

* **Documentation**
  * Clarified recipe selection and directory organization.
  * Documented checkpoint naming conventions and updated usage examples.

* **Bug Fixes**
* Updated launcher configurations and examples to reference the new
recipe locations and corrected model names.

* **Tests**
* Improved automatic recipe discovery and validation of documented
recipe paths.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-01 10:22:27 -07:00
Keval MorabiaandClaude Opus 5 d73278808b Bump nemo container requirement to 26.08 for MBridge examples (#2257)
### What does this PR do?

Type of change: Bug fix

Bumps the Megatron-Bridge examples, tests and launcher configs to
`nemo:26.08` and removes the version-gated fallbacks they carried, plus
the fixes needed to make the suites green on that container.

**26.08 bump and shim removal**

- Examples, CI workflows, `noxfile.py` and the `mbridge_*` launcher
configs move to `nemo:26.08`.
- `examples/megatron_bridge/_distillation_provider.py` is deleted —
26.08's Megatron-Bridge ships `convert_to_distillation_provider(...,
distill_submodule=...)` natively, so `distill.py` imports it directly.
- `prune_minitron.py` drops the `AutoBridge.from_hf_config` /
config-only-export probing; `--no_moe_grouped_gemm` is no longer needed
in the MoE pruning tests, and the Qwen3.5-MoE `skipif` is gone (native
MoE expert mappings are in 26.08).
- `_DynamicMambaMixer` targets only the raw `conv1d_weight` /
`conv1d_bias` parameters that replaced the `conv1d` module in
Megatron-Core.

**MambaModel / MambaModelProvider removal**

Megatron-Core has shipped `HybridModel` since 26.06 and `MambaModel` is
a deprecated subclass that shares its `forward`, so `DMRegistry`
resolves those instances to the `HybridModel` registration and the
separate entry is redundant. Same for `MambaModelProvider` vs
`HybridModelProvider` on the bridge side. `MambaMixer` / `MambaLayer` /
`ExtendedRMSNorm` are untouched — the layers still exist. The deprecated
`get_te_mamba_stack_spec` is removed; use `get_te_hybrid_stack_spec`.

**Bug fix: compressed output_layer extra state**

`mtq.compress` converts even a *disabled* `output_layer` into a
`RealQuantLinear` (its weight is left uncompressed, since
`pack_real_quantize_weight` skips disabled quantizers). The guard added
in #2112 excluded `RealQuantLinear` by class, so the layer emitted extra
state and every worker died in `GPTModel.sharded_state_dict`:

```
RuntimeError: Boolean value of Tensor with more than one value is ambiguous
  megatron/core/models/gpt/gpt_model.py:896, in sharded_state_dict
    output_extra_state and output_extra_state.data
```

The guard now keys off whether the weight was actually compressed
(`QTensorWrapper`) instead of the class. This took out all 12
`test_homogeneous_compressed_sharded_state_dict` params, and the crashed
workers poisoned the pool, which surfaced as unrelated timeouts and NCCL
errors in `test_layer_sync_moe_local_experts_amax`,
`test_kv_cache_quant`, `test_kv_cache_amax_sync`,
`test_convert_mcore_te_gpt_model` and
`test_homogeneous_sharded_state_dict_te_spec` — 21 tests in total. The
e2e coverage is `skip_flaky_on_blackwell`, so CI never ran it;
`test_output_layer_extra_state_empty_when_nothing_quantized` now asserts
the contract directly and is not skipped.

**Checkpoint import entry point**

26.08 replaced `examples/conversion/convert_checkpoints.py` with
`scripts/conversion/convert.sh`, so
`tools/launcher/common/megatron_bridge/import/import.sh` and the three
README snippets are retargeted. `import.sh` uses the distributed GPU
backend with `GPUS_PER_NODE` / `TP` / `PP` / `EP` knobs.

**Megatron-LM on nemo:26.06** keeps working: `_get_mamba_conv1d` still
dispatches between the `conv1d` module (26.06 and earlier) and the raw
parameters (26.08+), so `import_mcore_gpt_from_hf` /
`export_mcore_gpt_to_hf` handle NemotronH on both. Only the
Megatron-Bridge examples and Minitron pruning of Mamba/hybrid models
require 26.08.

**Test consolidation**

`test_export_distilled_megatron_to_hf.py` is merged into
`test_distill.py`: `test_distill_llm` becomes
`test_distill_llm_hf_export` and covers the standalone
`--export_iterations all` run on the checkpoints it already produces,
saving one full distillation (~185 s of CI time). The two mamba-named
gpu test files are renamed to `hybrid`.

### Usage

```bash
# HF -> Megatron import, via Megatron-Bridge's 26.08 conversion entry point
bash /opt/Megatron-Bridge/scripts/conversion/convert.sh import \
    --executor local \
    --device gpu \
    --gpus-per-node 8 \
    --hf-model Qwen/Qwen3-8B \
    --megatron-path /tmp/Qwen3-8B-megatron
```

### Testing

All on `nvcr.io/nvidia/nemo:26.08`, 2x RTX 6000 Ada, no timeout
overrides:

- `tests/examples/megatron_bridge`: 16 passed, 1 skipped (28m14s). The
skip is the `gemma3vl` QAD param, now `@pytest.mark.manual` since
`qwen3_5_moe_vl` covers the VLM QAD path.
- `tests/gpu_megatron` (`_extensions`, `distill`, `export`, `opt`,
`peft`, `sparsity`, `speculative`, `utils`): 61 passed, 5 xpassed.
- `tests/gpu_megatron/torch/export` re-run after the conv1d dispatch
change: 27 passed.
- The 21 previously failing/hanging quantization tests: 21 passed (12 +
9).
- `tests/gpu_megatron/torch/{nas,prune}`: verified separately.

`import.sh` equivalence on a toy `qwen3_moe`, comparing all 12 weight
tensors after flattening each dist checkpoint with `dcp_to_torch_save` —
the GPU backend at 1 GPU, `--tp 2`, `--pp 2`, `--ep 2`, and `import.sh`
end-to-end (`GPUS_PER_NODE=2 EP=2`) are all byte-identical to `--device
cpu`.

`nemo:26.06` compatibility was checked directly in that image:
`megatron.core.models.hybrid.HybridModel`, the modelopt hybrid spec and
`hybrid_layer_pattern` are all present, while
`megatron.bridge.models.hybrid` and the bridge's `distill_submodule` are
not. The NemotronH round-trip test failed there before the conv1d
dispatch was restored and the dispatch is back in place; per project
convention the suites themselves only run on 26.08.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ⚠️ Megatron-Bridge examples plus
Minitron pruning of Mamba/hybrid models now require `nemo:26.08`.
Megatron-LM quantization and checkpoint export still run on
`nemo:26.06`.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ —
`test_output_layer_extra_state_empty_when_nothing_quantized` for the
compress fix; existing tests extended for the merged export coverage.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added guidance for importing Hugging Face checkpoints into Megatron
distributed format.
* Expanded distillation workflows to export selected or all checkpoint
iterations.

* **Improvements**
  * Expanded Hybrid model support across Megatron workflows.
  * Updated distributed import tooling with GPU and parallelism options.
  * Updated supported environments and examples to NVIDIA NeMo 26.08.

* **Bug Fixes**
* Corrected output-layer quantization state handling when quantization
is disabled.

* **Documentation**
  * Added compatibility guidance for current and legacy NeMo containers.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 00:31:41 +05:30
h-guo18 5db2682519 [Example]: Calibration-free FP8/NVFP4 PTQ for speculative-decoding drafters (#2027)
### What does this PR do?

Type of change: new example

Adds `examples/speculative_decoding/scripts/quantize_drafter.py`, a CLI
that quantizes an exported speculative-decoding drafter to FP8 or NVFP4
— weight-only or weight+activation — with no calibration data.

It needs no modeling code either. Exported drafters such as
[`nvidia/MiniMax-M3-DSpark`](https://huggingface.co/nvidia/MiniMax-M3-DSpark)
have no importable model class, so each 2-D weight is wrapped in a
throwaway `nn.Linear` under its checkpoint key and ModelOpt's usual
`quantizer_name` patterns select over those names. Works for any drafter
layout (DSpark / DFlash / EAGLE3 / Medusa).

**Formats:** `w4a16_nvfp4`, `nvfp4`, `fp8`, `fp8_pc_pt` — the ModelOpt
formats vLLM's backend can actually serve. AWQ is deliberately not
offered, since `awq_lite` silently degrades to plain RTN without a
`forward_loop`.

**Static activation scales without calibration.** `fp8` and `nvfp4`
normally need an activation amax *measured* on calibration data; a fixed
`input_scale` of 1.0 is applied instead. That works because acceptance
length is governed almost entirely by **clipping**, not resolution:

Sweeping the fixed scale over three decades (same setup as the Testing
section below; bf16 baseline 3.1423):

| `input_scale` | amax | FP8 AL | vs bf16 | NVFP4 AL | vs bf16 |
|---|---|---|---|---|---|
| 0.003 | 1.3 | 2.2204 | -29.34% | 2.2076 | -29.75% |
| 0.01 | 4.5 | 2.6719 | -14.97% | 2.6641 | -15.22% |
| 0.03 | 13.4 | 2.9751 | -5.32% | 2.9259 | -6.89% |
| 0.1 | 44.8 | 3.1013 | -1.31% | 3.0206 | -3.88% |
| 0.2 | 89.6 | 3.1178 | -0.78% | 3.0015 | -4.48% |
| 0.3 | 134.4 | 3.1370 | -0.17% | 3.0222 | -3.82% |
| 0.5 | 224.0 | 3.1268 | -0.50% | 3.0360 | -3.38% |
| **1.0 (default)** | **448.0** | **3.1457** | **+0.11%** | **3.0193** |
**-3.91%** |
| 2.0 | 896.0 | 3.1354 | -0.22% | 3.0172 | -3.98% |
| 4.0 | 1792.0 | 3.1245 | -0.57% | 3.0034 | -4.42% |

Both formats fall off a cliff below ~0.03, where the declared range sits
far under the activations' true magnitude and most of the tensor is
clipped. Both then sit on a flat plateau from ~0.3 to 4.0 **with no
drop-off at the top**, so the scale only has to be big enough. 1.0 is
the middle of that plateau, which is why it is hardcoded rather than
exposed. NVFP4 trails FP8 by a roughly constant 3.5% across the plateau
— that gap is the 4-bit resolution cost, and no choice of scale recovers
it.

Deriving the amax from the weights instead was tried and does not work:
`max|W|` averages 0.79 while a RMSNorm'd activation is O(1) with outlier
channels in the tens, so the range lands 1–2 orders of magnitude low and
clips, measuring -31% to -46% AL.

**Where calibration would go.** All of this sits behind
`resolve_activation_scales()`, the single place deciding where a static
amax comes from. Real calibration slots in ahead of the fixed fallback
with no change to the CLI or the call site, and composes because
`set_static_activation_amax()` skips quantizers that already have an
amax:

```python
if calib_forward_loop is not None:
    mtq.calibrate(root, quant_cfg["algorithm"], forward_loop=calib_forward_loop)
set_static_activation_amax(root)   # fills in what calibration did not reach
```

**Serving a quantized drafter.** Four things had to be written into the
exported checkpoint before vLLM would load one:

- emit `quant_method` (`modelopt_fp4` / `modelopt`) — vLLM reads that
key, ModelOpt writes only `quant_algo`
- emit the exclusion list under `ignore` too — that is the key read from
the flat `quantization_config`; `exclude_modules` alone yields an empty
exclusion set
- add `*<name>` wildcards so exclusions match a runtime's nested module
prefix (`model.fc`) rather than the checkpoint key (`fc`)
- add `*qkv_proj` / `*gate_up_proj` aliases for layers a runtime fuses,
whose names appear in no checkpoint key

Nothing is then needed on the caller side. **This closes the open
question left in the previous revision of this PR: vLLM does read
`quantization_config` off the draft checkpoint.**
`ModelConfig._verify_quantization` fills `quantization` in from
`quant_method` when it is unset, so once the export declares that key —
the first fix above — detection works on its own. Verified on
Nemotron-3.5-Lightning passing nothing: `Detected ModelOpt NVFP4
checkpoint (quant_algo=NVFP4)` → `FlashInferCuteDslNvFp4LinearKernel`,
AL 4.278 against 4.203 measured earlier.

`specdec_bench` also gains a `DSPARK` algorithm, which it did not have:
an exported `Qwen3DSparkModel` would otherwise have to go through
`DFLASH` and be built with vLLM `method="dflash"`. The branch sets
`method="dspark"` and leaves `draft_sample_method` on vLLM's own default
of `greedy`. A target whose fused-collective workspace (sized at
CUDA-graph capture) overflows at large speculative batches can disable
graphs with `--runtime_params '{"engine_args": {"enforce_eager":
true}}'`.

For DFlash-family drafters, `qwen3_dflash.py` builds its fused
context-KV projection by reading `qkv_proj.weight` raw and calling
`F.linear`, which cannot consume a packed weight. Keep those layers in
bf16 with `--exclude '*q_proj*' '*k_proj*' '*v_proj*' '*qkv_proj*'`;
`o_proj` and the MLP — the bulk of the drafter — still quantize. That
exclusion is mandatory, not a tuning choice.

`fc` (the projection from the target's captured layers into the draft)
is the one real knob, and it is a genuine trade rather than a free win —
see the Testing section for both models' numbers. The examples quantize
it; add `'*fc*'` to the exclude list to keep it in bf16.

`embed_tokens`, `markov_head` and `confidence_head` are excluded by
default: they are 2-D so the flat view treats them as GEMMs, but they
are embeddings or a single-output projection. `lm_head` is excluded by
the preset itself — unlike on a base model it is 37% of this drafter's
parameters, so `--quantize_lm_head` is a real lever (~1.9 GiB), but
measure AL first. The flag re-enables both of `lm_head`'s quantizers;
re-enabling only the weight one would ship a W+A checkpoint whose
`lm_head` has no `input_scale` while the config still advertises it as
quantized.

### Usage

```bash
# weight+activation FP8, calibration-free, lossless on both models measured below
python scripts/quantize_drafter.py \
    --drafter_path deepseek-ai/dspark_qwen3_8b_block7 \
    --qformat fp8 \
    --export_path ./dspark-qwen3-8b-fp8 \
    --exclude '*q_proj*' '*k_proj*' '*v_proj*' '*qkv_proj*'

# smallest: weight-only NVFP4
python scripts/quantize_drafter.py \
    --drafter_path nvidia/MiniMax-M3-DSpark \
    --qformat w4a16_nvfp4 \
    --export_path ./MiniMax-M3-DSpark-W4A16
```

Or end to end on Slurm — quantize, then measure AL — via the launcher
examples added here, one per target:

```bash
uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_dspark_ptq_nvfp4.yaml --yes
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/hf_dspark_ptq_nvfp4.yaml --yes
```

Serving one, if you are not going through `specdec_bench`:

```python
speculative_config = {
    "method": "dspark",
    "model": "./dspark-qwen3-8b-fp8",   # quantization is read from its config.json
    "num_speculative_tokens": 7,
}
```

### Testing

Two targets with different architectures, so the conclusions are not one
model's quirk:

* **Qwen3-8B** (dense transformer) +
[`deepseek-ai/dspark_qwen3_8b_block7`](https://huggingface.co/deepseek-ai/dspark_qwen3_8b_block7),
`block_size` 7, TP1.
* **Nemotron-3.5-Lightning-30B-A3B** (hybrid Mamba-MoE) +
[`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark`](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark),
`block_size` 8, TP8, with the mamba engine settings the model card pins
(`mamba_backend=flashinfer`, `mamba_ssm_cache_dtype=float16`, stochastic
SSM-cache rounding).

Both: MT-Bench 80 questions, greedy, one vLLM instance per point.

| recipe | activations | Qwen3-8B AL | vs bf16 | Nemotron-3.5 AL | vs
bf16 |
|---|---|---|---|---|---|
| bf16 baseline | — | 3.1423 | — | 4.3296 | — |
| **`fp8`** | static, `input_scale` 1.0 | **3.1457** | **+0.11%** |
**4.3289** | **-0.02%** |
| `fp8_pc_pt` | dynamic per-token | 3.1228 | -0.62% | 4.3411 | +0.26% |
| `w4a16_nvfp4`, `fc` in bf16 | bf16 (weight-only) | 3.0392 | -3.28% |
4.2899 | -0.92% |
| `w4a16_nvfp4`, `fc` quantized | bf16 (weight-only) | 3.0186 | -3.94% |
4.2334 | -2.22% |
| **`nvfp4`** | static, `input_scale` 1.0 | **3.0193** | **-3.91%** |
**4.2030** | **-2.92%** |

**FP8 weight+activation at the fixed `input_scale` of 1.0 is lossless on
both.** +0.11% and -0.02% are both inside run-to-run noise — the
Nemotron baseline was measured twice under identical settings and the
two runs differ by 0.94% (4.3093 / 4.3499), which sets the resolution of
that column. On the same reading, `fp8` and `fp8_pc_pt` are
indistinguishable on Nemotron; the dynamic variant only pulls ahead on
Qwen3. NVFP4 costs 3-4% on Qwen3 and 2-3% on Nemotron, i.e. the 4-bit
weight resolution is the real price and it is model-dependent but
bounded.

Whether to quantize `fc` is a per-model call rather than a general
recommendation — it buys a few percent of size for an AL cost that
differs by ~2x between these two drafters:

| `fc` bf16 → quantized | Qwen3-8B | Nemotron-3.5 |
|---|---|---|
| checkpoint size | 3.293 → 3.181 GiB (-3.4%) | 1.316 → 1.258 GiB
(-4.4%) |
| AL | 3.0392 → 3.0186 (-0.68%) | 4.2899 → 4.2334 (-1.32%) |

`fc` itself is only 3.5% (Qwen3) / 4.5% (Nemotron) of drafter
parameters; `embed_tokens` is the bulk (26% / 36%) and is excluded by
default.

The Qwen3 `w4a16_nvfp4` rows were measured in a later session than the
rest of that column; the `fc`-in-bf16 run reproduced the original number
to four decimals (3.0392), so the column is internally comparable.

Also validated on `nvidia/MiniMax-M3-DSpark`: `w4a16_nvfp4` runs in 67 s
on CPU, 9.98 GiB (fp32) -> 3.51 GiB; all 43 quantized tensors round-trip
within 0.0952 relative error; the 29 untouched tensors are bit-identical
to `bf16(source)`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (example-only)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ — validated manually as
above. Can add a `tests/examples/speculative_decoding/` test over a
small synthetic drafter if wanted before merge.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (example-only)
- Did you get Claude approval on this PR?: ❌ (not yet run)

### Additional Information

The measurements above are one drafter on one target with one benchmark;
the plateau's location and the ~3.5% NVFP4 gap should be re-measured
before assuming they carry to a different drafter.

Note when reading an exported checkpoint: `input_scale` is `amax/448`
for FP8 but `amax/(6*448)` for NVFP4, so the one fixed amax records as
1.0 in an FP8 checkpoint and 0.1667 in an NVFP4 one. Both mean the same
activation range.

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-08-26 22:04:09 +08:00
h-guo18 2b296b2f62 Support fine-tuning released DFlash/DSpark drafters (causal SWA, attention sink, warm start) (#2149)
# Support fine-tuning released DFlash/DSpark drafters (causal SWA,
attention sink, warm start)

### What does this PR do?

Type of change: New feature + bug fix

Adds what ModelOpt was missing to fine-tune an already-published
DFlash/DSpark draft
model. The concrete target is

[`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark`](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark)
on its hybrid Mamba/attention/MoE base, but every change is generic.

Before this PR that checkpoint could not be trained faithfully — or even
loaded: its
attention-sink tensors were dropped as unexpected keys, its block-causal
attention had no
implementation, and its capture layers were silently overwritten with
ModelOpt's defaults.

**New user-facing options** (all default to today's behavior, so
existing runs are unchanged):

| Option | Values | Purpose |
| --- | --- | --- |
| `dflash_draft_attention` | `bidirectional` (default) / `causal` |
Block-internal attention pattern. `causal` restricts a query at block
position `i` to draft positions `<= i`. |
| `dflash_attention_sink` | `false` (default) / `true` | Learnable
per-head `attention_sink_bias [num_heads]` on every draft layer — one
extra logit appended before the softmax and dropped after, so a head can
put probability mass nowhere instead of being forced to attend inside
its window (the GPT-OSS formulation). |
| `dflash_init_checkpoint` | path | Warm-start the draft from an
exported checkpoint instead of a random init. Any
missing/unexpected/wrong-shaped tensor raises rather than warns. |
| `dflash_architecture_config.target_layer_ids` | list | Which base
layers feed the draft's `fc`. Previously recomputed unconditionally with
no override. |

**Bugs fixed along the way** (each one silently corrupts training rather
than failing):

- The exporter hard-coded `dflash_config.causal: False` and only wrote
it under SWA, so even
a correctly-trained causal draft would be served non-causally. It now
reflects the trained
  setting, and emits `attention_sink_bias` when enabled.
- `_build_generate_swa_mask` returned `None` whenever `swa_window_size`
was unset, which
would have dropped the causal structure at generation time while
training used it.
- `target_layer_ids` was recomputed from the uniform default on every
convert. The released
drafter uses `[1,5,19,29,41,51]`; the default for a 52-layer base is
`[1,11,20,30,39,49]`
— *different layers*. Here it surfaced as a matmul shape error only
because the plane
counts disagreed; with a matching count it would have trained on the
wrong features
  silently.
- The streaming dataset assumed the draft's aux layers all sit below the
base's final layer
(`aux = planes[:-1]`, `target = planes[-1]`). A draft whose top aux id
*is* the final layer
cannot get an extra plane — vLLM captures each layer once — so
`final_aux_is_base_hidden`
now lets the last plane serve both roles. It is derived from the model,
not configured by
  hand.
- DSpark head weights load from either the flat layout ModelOpt exports
(upstream DeepSpec
convention) or the nested `markov_head.` layout the NVIDIA release uses.
Without the remap
the two `[131072, 512]` Markov tables — ~14% of the draft's parameters —
stay randomly
  initialized while everything else warm-starts, with no error.
- `nemotron_h` is enabled in `_FINAL_NORM_TYPE_BY_MODEL_TYPE`: despite
the hybrid stack,
`NemotronHModel.norm_f` is a plain RMSNorm, and without the entry the
offline/streaming
  fake base raises instead of reconstructing the distillation target.

### Usage

```yaml
dflash:
  dflash_init_checkpoint: /path/to/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark
  dflash_draft_attention: causal
  dflash_attention_sink: true
  dflash_swa_window_size: 1024
  dflash_block_size: 8
  dflash_mask_token_id: 990
  dflash_architecture_config:
    target_layer_ids: [1, 5, 19, 29, 41, 51]
```

A full worked example is at

`modelopt_recipes/general/speculative_decoding/dspark_nemotron35_warmstart.yaml`.

### Testing

**Unit tests** — 124 pass (`test_hf_dflash.py`, `test_hf_dspark.py`,
`test_hf_domino.py`,
`test_hf_dflash_offline.py`, `test_modeling_final_norm.py`), 32 of them
new: causal mask
structure (lower-triangular per block, no cross-block leakage, context
visibility
unchanged), the sink math (degenerates to plain attention at `-inf`,
absorbs mass
monotonically, receives gradient), warm-start load/reject paths, Markov
key remapping, and
explicit `target_layer_ids`.

**Checkpoint compatibility** — the released drafter loads with zero
missing/unexpected keys
and zero shape mismatches; all 77 tensors (6 attention sinks and both
Markov tables
included) match bit-exactly, and a training step runs with gradients
reaching the sink and
Markov parameters.

**End-to-end streaming training** — Nemotron-3.5 base served by vLLM (1
node, TP8) feeding
8 trainer GPUs over NIXL; the draft warm-starts from the released
checkpoint and trains with
`causal` + sink + SWA 1024. 128 Daring-Anteater conversations, 20 epochs
(the plot shows the
first 5, where the trend is clearest — the curves flatten after that):

![warm-start training
curves](https://raw.githubusercontent.com/h-guo18/Model-Optimizer/pr-assets/dspark_nemotron35_warmstart_curves.png)

Over the first 5 epochs loss falls **1.85 → 1.36** and train accuracy
rises
**0.25 → 0.49**; across the full 20 epochs they reach **1.21** and
**0.48** (peak 0.54)
before flattening. This validates the pipeline end-to-end — capture
layers, plane split,
mask direction, sink loading and warm-start weights all have to be right
for this curve to
appear. It is *not* a model-quality result: 128 samples over 20 epochs
overfits by
construction, and the corpus is not generated by the base model, so the
absolute numbers are
not meaningful.

### TODO (follow-up)

**A complete, robust checkpoint/config converter.** Both conversions are
handled ad hoc here:

- *Draft config → training config.* The recipe transcribes ~15 fields by
hand from the
drafter's `config.json`. Only the shape-bearing ones
(`num_hidden_layers`,
`num_attention_heads`, `intermediate_size`, `markov_rank`) fail loudly
when mistyped; the
rest — `mask_token_id`, `causal`, `swa_window_size`, `block_size` —
train "successfully" on
a wrong value and only surface later as a mysteriously low acceptance
length. A converter
should derive the whole block from the checkpoint, including its aliases
(`pard_token`,
`dspark_markov_rank`, `dflash_query_causal`, top-level `sliding_window`
/
  `attention_sink_bias`) and duplicated fields.
- *Weight layout.* The `markov_head.` remap is a load-time hook. A
converter should normalize
layouts explicitly, and decide whether export should also emit the
release's aliases so a
round-trip reproduces the original format (today it renames
`architectures` to
  `DFlashDraftModel`).
- *Base config.* Serving this base on vLLM needs its `config.json`
layer-type vocabulary
updated for the transformers-5 path (`mamba` → `linear_attention`,
`attention` →
`full_attention`, plus a matching `hybrid_override_pattern`). That is
done by hand today and
  is not covered by this PR.

### Before your PR is "*Ready for review*"

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added configurable causal or bidirectional attention for DFlash
models.
* Added optional attention sinks, checkpoint warm starts, and explicit
target-layer selection.
* Improved streaming data handling for shared auxiliary and base hidden
states.
* Added Nemotron-3.5 Lightning DSpark warm-start training and serving
recipes.

* **Bug Fixes**
  * Preserved configured attention behavior during model export.
* Prevented warm-start checkpoints from being reapplied during
restoration.
* Improved checkpoint compatibility, validation, and attention-mask
handling.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-08-23 20:49:43 +08:00
Keval MorabiaandClaude Opus 4.8 58ad6edc5f Fix pruned-HF export fallback + add Nemotron-3.5-Lightning launcher examples (#2196)
### What does this PR do?

Type of change: Bug fix + new example

Two related changes for the Megatron-Bridge Minitron prune/quantize
launcher flows:

1. **Fix pruned-HF export crash on containers that reject config-only
save.**
   `#2159` added a config-only HF export path gated only on
`hasattr(AutoBridge, "from_hf_config")`. Some Megatron-Bridge versions
   (e.g. `nemo:26.04`) expose `from_hf_config` but reject a config-only
`save_hf_pretrained` (`ValueError: save_hf_pretrained requires a
pretrained
HuggingFace model`), so `prune_minitron.py` crashed instead of using the
intended dummy-model fallback. Now it attempts the config-only save and
   falls back to the dummy-model path on `ValueError`.

2. **Add Nemotron-3.5-Lightning-30B-A3B launcher examples**
(`mbridge_prune.yaml`,
`mbridge_quantize.yaml`) on `nemo:26.08`. Prune targets 3B active with
an
MMLU gate; quantize runs W4A16 NVFP4 4/6 PTQ via the `w4a16_nvfp4_4o6`
recipe with `tp_size=1` (static-block NVFP4 MSE is unsupported with
TP>1).

### Usage

```shell
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_prune.yaml --yes
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_quantize.yaml --yes
```

### Testing

Verified end-to-end on OCI-HSG:

- **Nano prune (`nemo:26.04`)** — exercises the fallback path:
config-only save
  raised the `ValueError`, the fallback caught it and exported via the
dummy-model path. `mmlu_10pct_bs32 = 0.5196` (gate 0.50) PASS; vLLM gen
PASS.
- **Lightning prune (`nemo:26.08`)** — config-only export path: `score =
0.6000`
  (gate 0.58) PASS, 3.00B active params; vLLM gen PASS.
- **Lightning quantize (`nemo:26.08`)** — recipe PTQ + unified-HF
export;
  MMLU `0.7741` (gate 0.75) PASS.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- launcher example
configs + fallback path exercised by CI prune/quantize jobs -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- fix is for a bug introduced in the same unreleased cycle
(#2159); rest are example configs -->
- Did you get Claude approval on this PR?: ❌ <!-- pending /claude review
-->

### Additional Information

The fallback fix addresses the `mbridge_prune` launcher CI failure
introduced by #2159.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added a pruning workflow for Nemotron-3.5-Lightning-30B-A3B with
calibration, quality scoring, checkpoint export, and multi-GPU
generation.
- Added a four-GPU NVFP4 W4A16 quantization workflow with Hugging Face
conversion and MMLU evaluation.
- **Bug Fixes**
- Improved hybrid model export by falling back to dummy-model export for
supported configuration-only export failures.
- Added clearer logging and handling for supported export failures while
preserving unrelated errors for investigation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-17 22:15:47 +05:30
Jenny Chen 5e887aa7c9 fix(launcher): increase Nemotron 3.5 Lightning QAD parallelism (#2191)
### What does this PR do?

Type of change:  Bug fix

Current Nemotron 3.5 Lightning QAD example can OOM, and also TP=2 is not
supported for static quantization. Increase GPUs, CP and decrease TP=1

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Very short summary of changes only for new features,
backward breaking changes, deprecations, or fixes for critical bugs
present in previous releases. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Configuration**
* Updated the QAD training setup to run across four nodes with revised
tensor and context parallelism settings.
* Documented the 32K sequence-length constraint and removed outdated run
information.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
2026-08-14 12:08:44 -07:00
yueshen2016 71b3d884ee example(launcher): Megatron-Bridge NVFP4 QAD launcher example for Nemotron-3.5-Lightning-30B-A3B (#2142)
## What does this PR do?

Adds `mbridge_qad.yaml`, a launcher example running NVFP4
quantization-aware
distillation for **Nemotron-3.5-Lightning-30B-A3B** through the
Megatron-Bridge
scripts in `examples/megatron_bridge/`, alongside the existing
`mbridge_prune.yaml` / `mbridge_quantize.yaml`.

`megatron_lm_qad.yaml` (#2146) runs the same recipe and the same data
through
Megatron-LM. This is the Megatron-Bridge counterpart: the same

`huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6`
recipe and the same `nvidia/Nemotron-Post-Training-Dataset-v2` chat
data, with the
training hyperparameters from our public-data QAD run.

Four tasks: tokenize the training data, PTQ the student, distill it
against the
frozen BF16 teacher, export to unified HF.

### Why the extra tokenize task

Megatron-LM's finetune path reads an HF parquet shard directly.
Megatron-Bridge
trains from pre-tokenized data, so `distill.py` consumes Megatron
`.bin`/`.idx`
via `--data_paths`. The chat split is therefore tokenized once with
`modelopt.torch.utils.plugins.megatron_preprocess_data`.
`--hf_streaming` avoids
the Arrow cast errors this dataset's nested tool-call fields trigger in
non-streaming mode, and `--append_eod` is omitted because chat rows
already
terminate each conversation via the chat template.

### Details

- Training topology 8 nodes x 4 GPUs, TP=1 PP=1 CP=4 EP=16 -> DP=8;
`gbs` 64 at
`mbs` 1 is 8 gradient-accumulation microbatches. 200 iters x 64 x 32768
= 419M
  training tokens.
- PTQ runs TP=EP=PP=1 across 4 ranks (pure DP), so each rank calibrates
on its own
shard. `--calib_dataset_name` is left unset, selecting the default
public
`cnn_nemotron_v2_mix` (cnn_dailymail +
Nemotron-Post-Training-Dataset-v2).
- Export uses TP=1 (the HF writer does not gather TP shards) and PP=4,
splitting
  52 layers 13/stage.
- Pins `nvcr.io/nvidia/nemo:26.06` like the other `mbridge_*` examples.

## Dependencies

Based on `main`; the PTQ recipe ships in #2146 (merged). No other PR
required.

Nemotron-3.5-Lightning has `tie_word_embeddings: false`, so a correct
quantized
`lm_head` in the exported checkpoint also depends on #2112.

## Testing

The PTQ -> export -> QAD flow and these hyperparameters were run end to
end on
Nemotron-3.5-Lightning (`main` + #2112 + #2113):

- PTQ completed, 6660 quantizers, MTP heads retained (`mtp_num_layers:
1`) with
  all 278 `mtp.*` quantizers disabled by the recipe.
- Export produced a unified-HF checkpoint (18487 keys, including 270 MTP
tensors).
- QAD trained with 900 quantizers through a validation pass at iteration
50.

The YAML itself is validated against the launcher's conventions
(`ntasks_per_node == gpus_per_node` on Slurm, single-line `inline`, no
`args`
alongside `inline`, all `<<global_vars.X>>` resolve, output prefix
matches
`megatron_preprocess_data`'s naming) and by the repo's `validate
launcher YAML
references` pre-commit hook. Topology arithmetic checked: EP divides
world/(TP*PP), `gbs` divisible by DP*mbs.

## Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (new example file only)
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an example workflow for NVFP4 quantization-aware distillation of
NVIDIA Nemotron 3.5 Lightning 30B-A3B.
* Supports dataset tokenization, post-training quantization,
teacher-student distillation, and export of a unified Hugging Face
checkpoint.
* Includes configurable model, dataset, and checkpoint paths,
distributed execution settings, and support for local or Slurm-based
workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: James Shen <yueshen@nvidia.com>
2026-08-13 00:27:55 +00:00
Jenny Chen 4ca7dd8871 Add Nemotron Lightning 3.5 NVFP4 recipe and QAD example (#2146)
### What does this PR do?

Type of change: New example

Add Nemotron Lightning 3.5 NVFP4 recipe and QAD example
Also exclude MTP in default disabled quantizers
### Usage

```python
# uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/megatron_lm_qad.yaml --yes
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, backward breaking changes,
deprecations, or fixes for critical bugs present in previous releases.
-->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added an end-to-end NVFP4 quantization, distillation, and export
workflow for NVIDIA Nemotron 3.5 Lightning 30B-A3B.
- Added a PTQ configuration supporting NVFP4 W4A16 quantization with
optimized scaling and FP8 support for selected components.

- **Bug Fixes**
- Preserved custom model output locations when provided, while retaining
the existing default path.

- **Configuration**
  - Disabled quantization for MTP modules by default.
- Updated an existing Nemotron workflow to use the aggressive NVFP4
quantization profile.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
2026-08-11 18:05:46 +00:00
Jenny Chen f2bfe63183 Nemotron Nano 3 QAD Launcher Example on OSS Nemotron-Post-Training-V2 data (#2134)
### What does this PR do?

Type of change: New example

Add a Nemotron Nano 3 QAD Launcher Example on OSS
Nemotron-Post-Training-V2 data. It performs 4 steps
1. Teacher conversion: Convert the HuggingFace BF16 checkpoint to a
Megatron-Core BF16 checkpoint
2. PTQ: quantize the Megatron-Core checkpoint to
`MAMBA_MOE_NVFP4_AGGRESSIVE_CFG` quant config
3. QAD (Quantization Aware Distillation): distill the BF16 checkpoint to
the PTQ checkpoint on a subset of the Nemotron-Post-Training-V2 `chat`
data. To train on a different subset or load the entire dataset, you may
modify `--finetune-data-split` and `--finetune-data-files` flags.
4. Export: export the QAD checkpoint to HuggingFace format so it is
ready for local inference


All steps use the TE (Transformer Engine) spec, which with the new
TEGroupedMLP per-expert quantizer is approximately 10-15% faster than
the previous local ModelOpt spec (which used SequentialMLP) on
Hybrid-MoE models.

### Usage

```
# Usage from tools/launcher:
source .env-slurm
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/megatron_lm_qad.yaml --yes
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, backward breaking changes,
deprecations, or fixes for critical bugs present in previous releases.
-->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
- Added a launcher configuration for NVFP4 quantization-aware
distillation of the Nemotron 3 Nano 30B-A3B model.
- Added support for selecting training or fine-tuning workflows through
`MLM_TRAIN_SCRIPT`.
  - Improved forwarding of additional training arguments.

- **Updates**
  - Updated the Megatron-LM launcher component to a newer revision.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
2026-08-10 13:52:52 -07:00
Asha AnooshehandClaude Sonnet 4.6 7afbfbc85c Offline-KD QAD example (#1998)
### What does this PR do?

Type of change: new feature

Adds a Nano-v3 launcher example which uses the new MLM offline-logits KD
feature

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌
- Did you get Claude approval on this PR?: ❌ 

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a new offline, five-stage Slurm workflow example for
Nemotron-3-Nano 30B NVFP4 A3B, using cached teacher logits for knowledge
distillation.
* Runs quantization, frozen teacher logits generation, student
fine-tuning from saved logits, checkpoint export to Hugging Face format,
and TensorRT-LLM evaluation.
* Includes ready-to-run configuration for repeatable execution with
tuned parallelism and export settings.
* **Chores**
* Updated local Python pre-commit hooks to execute via `uv` for
consistent dev-time tooling.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-08-05 10:59:36 +02:00
Keval MorabiaandClaude Opus 4.8 93b9e4b176 Add Megatron-Bridge prune & quantize launcher pipelines (#2031)
### What does this PR do?

Type of change: new example + small launcher / modelopt-example features
(backward compatible)

Adds end-to-end ModelOpt **launcher** pipelines for the Megatron-Bridge
flow on Nemotron-3-Nano-30B-A3B, the minimal launcher features to run
them wrapper-free from YAML, and an **in-step accuracy gate** for
Minitron pruning.

**New launcher examples**
(`tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/`):
- `mbridge_prune.yaml` — Minitron prune **with an in-step MMLU gate** →
vLLM sanity gen (2 tasks)
- `mbridge_quantize.yaml` — FP8 quantize → unified-HF export → MMLU gate
on the vLLM backend, which doubles as the deploy sanity check (3 tasks).
Matches the
[tutorial](https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md).

**Prune accuracy gate** (reuses the search's own score — no separate
eval step):
- `modelopt/torch/prune/plugins/mcore_minitron.py` —
`MCoreMinitronSearcher` now stores the exported best `CandidateSubnet`
under `state_dict["best"]` (additive; sits beside the existing
`sorted_layers` key).
- `examples/megatron_bridge/prune_minitron.py` — new
`--score_lower_bound`: reads `pruning_scores["best"].score` and exits
non-zero if the exported model is below the floor. Score-agnostic (any
`--prune_score_func`); rejected with `--prune_export_config` (manual
pruning has no score).

**Launcher (`tools/launcher`)** — run single-node Megatron-Bridge
one-liners directly from YAML:
- `SandboxTask.inline` — a command in the YAML, no `common/**/*.sh`
wrapper (single-line; folded scalar)
- `SandboxTask.reqs` / `reqs_file` — pip-install deps in the container
before the command (shell-safe; on Slurm the install is rank-0-guarded
so multi-rank tasks don't race)
- `SlurmConfig.docker_user` — local-Docker user (e.g. `root`); ignored
on Slurm
- `get_default_env` honors `HF_HOME` / `TRITON_CACHE_DIR` env overrides,
so a non-CI user can point caches at a writable path (the shared
`/cicd/hf-cache` is owned by the CI account)
- reject `args` together with `inline`

**`examples/llm_eval/lm_eval_hf.py`**:
- `--accuracy_lower_bound` — gate on the single requested task's `acc`
(used by the quantize MMLU step; exits non-zero if below)
- drop ModelOpt (hf-only) args for non-`hf` backends, so `--model vllm`
works on a deployable quantized checkpoint

### Usage

```bash
cd tools/launcher
# Prune (in-step MMLU gate) -> vLLM gen
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/mbridge_prune.yaml --yes
# FP8 quantize -> unified-HF export -> MMLU gate (vLLM)
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/mbridge_quantize.yaml --yes
```

### Testing

- Launcher unit tests for `inline`, `reqs`/`reqs_file`, `docker_user`,
the `args`+`inline` guard, and example-resolve; ruff / mypy / bandit
clean. `tests/examples/megatron_bridge/test_prune_minitron.py` now
passes `--score_lower_bound=0.01` to exercise the gate path on the tiny
models.
- **End-to-end on the real Nemotron-3-Nano-30B-A3B (4×B200, OCI-HSG):**
- Prune 30B → 3B-active: `[score_gate] mmlu_10pct = 0.5196 >= 0.45
PASS`; vLLM gen coherent.
- FP8 quantize → unified-HF export (`Detected ModelOpt fp8 checkpoint`)
→ MMLU on the vLLM backend `acc = 0.7077 >= 0.60 PASS`.
- Earlier smoke on **Qwen3-0.6B** in `nemo:26.06` through the same flow.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (new state-dict key is
additive; new CLI args default to off)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- tooling/examples + additive searcher state key -->
- Did you get Claude approval on this PR?: ❌ <!-- pending -->

### Additional Information

- **Container pinning:** saving a pruned Nemotron-H to HF requires
`transformers<5`, so `mbridge_prune.yaml` runs on `nemo:26.04` (26.06
drops it); quantize/export run on `nemo:26.06`.
- **`docker_user: root`** is set on all example tasks — local-Docker
only (ignored on Slurm), needed so downstream tasks can read task_0's
root-owned checkpoints and to read the image's root-only
`/opt/Megatron-Bridge`.
- The quantize MMLU step passes `enforce_eager=True` to vLLM — for a
run-once eval this skips ~17 min of CUDA-graph capture / `torch.compile`
with no accuracy change.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added inline shell commands and per-task Python dependency
installation for launcher workflows.
- Added configurable Docker user selection and preservation of existing
cache environment settings.
- Added evaluation accuracy and pruning score gates that fail workflows
below configured thresholds.
  - Added NVIDIA Nemotron pruning and quantization workflow examples.
  - Improved backend-specific handling of ModelOpt options.

- **Documentation**
- Documented inline commands, dependencies, variable substitution, and
configuration examples.

- **Bug Fixes**
  - Strengthened task execution validation and configuration checks.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-04 10:09:11 +05:30
Frida HouandClaude Opus 4.8 c062b3f829 Add sidecar GPU/CPU memory+utilization monitor for HF PTQ (#2000)
### What does this PR do?

**Type of change:** New feature (developer tool) + new tests

Adds a **standalone, cross-process resource monitor** —
`tools/resource_monitor.py` —
that samples a workload's GPU and CPU usage *from the outside* while it
runs, so
we can verify per-run memory budgets (e.g. the single-GPU layerwise PTQ
target of
≤80 GB GPU / ≤80 GB CPU for OMNIML-4947) without instrumenting
`hf_ptq.py` itself.

The monitor wraps any command, samples at a fixed interval, and on exit
writes a
CSV timeseries plus a peak/mean/min summary:

- **GPU** (per device, via NVML with an `nvidia-smi` fallback): used
memory,
  utilization %, power draw (W), temperature (°C).
- **CPU** (via `psutil`): system total/used/free memory + utilization %,
and the
  monitored process tree's RSS + CPU %.

It is opt-in from `examples/hf_ptq/scripts/huggingface_example.sh` via
`MODELOPT_MEM_MONITOR=1` (off by default → byte-for-byte identical
behavior). When
enabled it wraps the `hf_ptq.py` run and writes the trace/summary to a
**sibling**
`${SAVE_PATH}_mem_monitor/` directory, kept out of the exported
checkpoint that is
uploaded and consumed downstream.

**Files:**
- `tools/resource_monitor.py` — the sidecar (NVML + `nvidia-smi`
fallback; `psutil`).
- `tests/unit/tools/test_resource_monitor.py` — CPU-only unit tests (run
in the `unit` nox lane).
- `examples/hf_ptq/scripts/huggingface_example.sh` — opt-in
`MODELOPT_MEM_MONITOR=1` wrapper.
- `examples/hf_ptq/requirements.txt` — adds `psutil`.
- `pyproject.toml` — adds `psutil` to the `dev-test` extra
(deterministic import in the unit lane).
- `.github/workflows/unit_tests.yml` — adds `tools/resource_monitor.py`
to the unit-test path filters.

#### Why a new tool instead of extending
`modelopt/torch/utils/memory_monitor.py`?

The existing `GPUMemoryMonitor` is a fundamentally different tool and
cannot serve
this use case by extension:

| | `modelopt.torch.utils.memory_monitor.GPUMemoryMonitor` |
`tools/resource_monitor.py` (this PR) |
|---|---|---|
| Scope | **In-process** thread inside the workload | **Cross-process**
— wraps an external command |
| Survives workload OOM/SIGKILL | ❌ dies with the process | ✅ keeps
sampling, still writes the summary |
| Import cost | Pulls `torch` (~19 s) — lives in the workload |
Torch-free (`psutil`+`pynvml`, ~0.03 s) |
| Metrics | GPU device memory only | GPU mem/util/**power/temp** +
**CPU** mem/util + process-tree RSS |
| Output | In-memory / logs | CSV timeseries + peak/mean/min summary |

Merging the two would force `torch` into a standalone sidecar (defeating
the point)
or split the in-process monitor's threading model. A future refactor may
factor out
a **shared torch-free sampling core with two thin frontends**
(in-process + sidecar);
that is tracked as a follow-up rather than blocking this monitoring
harness, which
PR #2008 (single-GPU disk-offload PTQ) depends on.

### Usage

```bash
# Wrap mode (preferred): monitor exits with the workload's return code
python tools/resource_monitor.py --gpus 2,3 --out mem.csv --summary peak.txt -- \
    python hf_ptq.py --pyt_ckpt_path=<model> --qformat=nvfp4 ...

# Opt-in from the HF PTQ example (off by default):
MODELOPT_MEM_MONITOR=1 CUDA_VISIBLE_DEVICES=2,3 CUDA_DEVICE_ORDER=PCI_BUS_ID \
    bash examples/hf_ptq/scripts/huggingface_example.sh <args>
# -> writes ${SAVE_PATH}_mem_monitor/mem_trace.csv and mem_peak.txt
```

### Testing

- **Unit (CPU-only, in the `unit` nox lane):** `pytest
tests/unit/tools/test_resource_monitor.py`
— 11 tests covering `--gpus` parsing (CSV + space-separated, UUID/MIG
rejection),
the disabled/`nvidia-smi` sampling paths (including `[N/A]` → `None` and
the
smi-failure-yields-empty guard), CPU sampling, the accumulator, and
end-to-end
  CSV/summary + exit-code propagation in wrap mode.
- **GPU-validated** on a B200 node (GPUs 2,3): confirmed the `gpu{i}_*`
memory /
utilization / power / temperature columns populate and the summary is
written.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — new tool; the example wrapper
is off unless `MODELOPT_MEM_MONITOR=1`.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — `psutil`
added to `dev-test` + the hf_ptq example requirements;
`pynvml`/`nvidia-ml-py` is an optional runtime dep (graceful
`nvidia-smi` fallback).
- Did you write any new necessary tests?: ✅ —
`tests/unit/tools/test_resource_monitor.py`.
- Did you update Changelog?: N/A — repo-level `tools/` script, not
shipped in the wheel.
- Did you get Claude approval on this PR?: ❌ <!-- run /claude review -->

### Additional Information

Part of **OMNIML-4947** (single-GPU disk-offload PTQ). This is PR 1 of
the stack —
the monitoring harness that PR #2008 (disk-offload layerwise PTQ +
offload-aware
export) builds on.

---------

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-30 06:07:11 +00:00
Frida HouandClaude Opus 4.8 a3ac4759dd tools/mcp: pin mcp<2 to fix unit CI collection (#2026)
### What does this PR do?

**Type of change:** Bug fix (CI)

`mcp==2.0.0` was released and removed the `mcp.server.fastmcp` module
(the 2.0 SDK replaces it with `mcp.server.mcpserver`).
`tools/mcp/pyproject.toml` declared an unpinned `mcp>=1.0`, so CI now
resolves `mcp==2.0.0`, and `tools/mcp/modelopt_mcp/server.py`'s `from
mcp.server.fastmcp import FastMCP` fails at import:

```
ModuleNotFoundError: No module named 'mcp.server.fastmcp'
ERROR collecting tools/mcp/tests/test_bridge.py
```

This breaks the `mcp` unit job — and thus the `unit-pr-required-check`
gate — on **every** PR whose diff touches `pyproject.toml`,
`noxfile.py`, or `.github/workflows/unit_tests.yml` (the changed-files
paths that trigger the `mcp` job).

This PR pins `mcp>=1.0,<2`, keeping the 1.x line that still ships
`mcp.server.fastmcp`. Migrating the server to the mcp 2.0 API
(`mcp.server.mcpserver`) is a larger change tracked separately.

### Usage

```python
# N/A - dependency pin only
```

### Testing

- `uv pip install -e tools/mcp` now resolves an `mcp` 1.x wheel, so
`from mcp.server.fastmcp import FastMCP` imports and
`tools/mcp/tests/test_bridge.py` collects again.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — tightens
an existing dependency's upper bound.
- Did you write any new necessary tests?: N/A — restores collection of
the existing `tools/mcp` tests.
- Did you update Changelog?: N/A — CI/build fix, nothing user-facing in
the wheel.
- Did you get Claude approval on this PR?: ❌

### Additional Information

Unblocks the `unit-pr-required-check` gate for in-flight PRs (surfaced
on #2000). Follow-up: migrate `modelopt_mcp/server.py` to the mcp 2.0
API and relax the pin.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Prevented compatibility issues by limiting the MCP package to
supported version 1.x releases.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-28 21:34:24 -07:00
h-guo18 6105fe84e9 [Examples]: MiniMax-M3 DSpark (#1965)
### What does this PR do?

Type of change: new example + bug fixes

Adds a **MiniMax-M3 DSpark streaming training recipe** under
`tools/launcher/examples/MiniMaxAI/MiniMax-M3/`, plus the four fixes it
needs to actually run. Each fix addresses a failure mode that is silent
or misleading without it:

1. **Gemma-style final norm for the fake base**
(`modeling_final_norm.py`, `modeling_fakebase.py`): M3 uses a
gemma-style final RMSNorm (`(1 + weight)` scale, fp32
multiply-then-cast). Selecting the norm by `model_type` alone picks
plain `rmsnorm` — MiniMax's VL remote code coerces its `model_type`-less
`text_config` to **mixtral** — which silently drops the `+1` and
corrupts the distillation target. New `_FinalGemmaRMSNorm` + selection
by the explicit `use_gemma_norm` config flag (only MiniMax sets it).
2. **Loud failure for un-maskable `answer_only_loss`**
(`hf_streaming_dataset.py`): with a fast tokenizer whose chat template
has no `{% generation %}` tags,
`apply_chat_template(return_assistant_tokens_mask=True)` only warns and
returns an **all-zero mask** — training runs at zero loss on every
sample with no other symptom. Now raises at tokenization with an
actionable message.
3. **`SERVE_BLOCK_SIZE` knob** (`train_eagle_streaming.sh`): nemo_run
exports env values unquoted, so a multi-token
`SERVE_EXTRA_ARGS="--trust-remote-code --block-size 128"` loses
everything after the first token. M3's MSA sparse attention requires KV
block 128 (`ValueError: No common block size for 16` at engine init
otherwise), so `--block-size` gets a dedicated single-token knob.
4. **Relax the speculative_decoding `transformers` pin to `<5.13`**
(match `pyproject.toml`): the old `<5.4` pin downgrades recent vLLM
containers (e.g. transformers 5.12.1, which also provides in-tree
`minimax_m3_vl`) and breaks `vllm serve` (`ALLOWED_LAYER_TYPES` needs
>=5.5.3).

The example itself encodes the validated M3 specifics: generation-tagged
chat template copy (required for `answer_only_loss` — see fix 2), draft
dims + base `rope_theta=5e6` set explicitly (not inherited), mask token
200063 (reserved slot; added tokens end at 200060), `EAGLE_CAPTURE_IDS`
= draft default `target_layer_ids+1` + final layer, `trust_remote_code`
at serve/export, and AWS-EFA NIXL notes (UCX segfaults at agent init on
EFA nodes; LIBFABRIC required there).

### Usage

```bash
cd tools/launcher
export SLURM_HOST=localhost SLURM_ACCOUNT=<account> SLURM_PARTITION=<partition> \
       SLURM_HF_LOCAL=<hf_models_dir> SLURM_JOB_DIR=<experiments_dir> NEMORUN_HOME=$PWD
uv run launch.py --yaml examples/MiniMaxAI/MiniMax-M3/hf_streaming_dspark_multi_node.yaml \
       identity=$HOME/.ssh/id_ecdsa detach=True --yes
```

### Testing

- `tests/unit/torch/speculative/plugins/`: **129 passed** inside a
current vLLM x86_64 nightly container (transformers 5.12.1), including 3
new tests for the assistant-mask guard.
- Generation-tagged template verified against real corpus samples:
`input_ids` identical to the original template, mask covers exactly the
assistant turns (think prefix + content + eos), contiguous, no
user-prompt leak.
- The recipe is exercised end-to-end by a live M3 DSpark training run
(this yaml modulo cluster paths): streaming serve + NIXL transport +
resume all healthy; drafter MT-Bench AL exceeds our Kimi-K2.6 DSpark
reference by ~8k steps.
- `ruff check` / `ruff format` (0.15.20) clean.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (norm selection only changes
models with `use_gemma_norm=True`, previously mis-normed; the guard
turns a silent zero-loss run into an error; `SERVE_BLOCK_SIZE` is
opt-in; the pin relax widens the allowed range)
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ (can add if desired)

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added Gemma-style RMS normalization support for speculative decoding.
- Added MiniMax-M3 multi-node training configuration, including a full
Jinja chat template for tool calls, multimodal content, and thinking
modes.
  - Added optional `SERVE_BLOCK_SIZE` support for vLLM serve launches.
- **Bug Fixes**
- Improved `answer_only_loss` masking validation: now fails fast with
clear errors when required `{% generation %}` markers are missing or
when using a non-fast tokenizer.
- **Compatibility**
- Expanded the supported Transformers version range for speculative
decoding examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-07-26 01:35:35 +00:00
Keval Morabia 01c708e792 Add HybridModel MBridge support for nemo:26.08 (#2005)
### Description

MBridge main (nemo:26.08) will initialize Nemotron-H as HybridModel
instead of MambaModel (subclassed of HybridModel). Also make minimum
nemo container 26.04

### Testing 

Tested Nemotron-3-Nano PTQ with MBridge main (fails otherwise)

Tested locally `tests/gpu_megatron` and `tests/examples/megatron_bridge`
with `nemo:26.06.01` + Mount latest MBridge/Mcore

GH CICD tests will be added with nemo:26.08 release

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added TE HybridModel stack-spec support and enabled Megatron-Bridge
hybrid providers/models for export/import and runtime handling.
* **Updates**
* Dataset packing now oversamples raw text at **16x** and improves the
packed-mode underflow warning.
* Quantization: `--quant_cfg` now defaults to `None` unless explicitly
set (or via `--recipe`).
* Distillation example: validation settings are provided via a dedicated
top-level validation configuration.
* Improved plugin import warnings to report the originating call
location; model stats now support HybridModel.
* **Deprecations**
* Megatron-Bridge / Megatron-LM optimization features now require NeMo
container `nemo:26.04` or newer (`nemo:26.06` recommended).
  * The Mamba stack specification helper is deprecated.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-07-23 18:26:15 +05:30
noeyy-mino f10d5183ec launcher: bump TRT-LLM to 1.3.0rc20, pin vLLM to v0.22.0, fix max_seq… (#1982)
bump TRT-LLM to 1.3.0rc20, pin vLLM to v0.22.0, fix max_seq_len for
Qwen3.5-4B

### What does this PR do?

Type of change: Bug fix, new feature

- Upgrade TRT-LLM container from 1.3.0rc10 to 1.3.0rc20 across all
  Qwen3-8B, Qwen3-30B-A3B, Kimi-K2.5, and gpt-oss-20b launcher configs.
- Replace Kimi-K2.5 aarch64-specific vLLM image (v0.22.0-aarch64) with
the multi-arch v0.22.0 tag, which resolves to amd64/arm64 automatically.
- Fix Qwen3.5-4B throughput_32k runs: raise max_seq_len from 40960 to
  65536 to accommodate outlier prompts (~46.6K tokens) that caused
  VLLMValidationError and aborted the entire benchmark run.
- Fix specdec_bench_mtp_vllm.yaml: remove reference to non-existent
  runtime_params_throughput_32k.yaml; use --max_seq_len 65536 instead

### Usage

```
cd Model-Optimizer/tools/launcher
uv run launch.py --yaml examples/Qwen/Qwen3-8B/megatron_lm_ptq_local.yaml hf_local=/mnt/hf-local --yes
```

### Testing
N/A


- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?:N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information
N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Improvements**
* Updated launcher example pipelines and the default launcher Slurm
container to the newer TensorRT-LLM `1.3.0rc20` image.
* Updated Kimi-K2.5 workflows to use the multi-architecture vLLM
`0.22.0` image (removing prior architecture-specific variants).
* Increased the Qwen3.5-4B long-context benchmark maximum sequence
length to 65,536 tokens and simplified the corresponding throughput
configuration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com>
2026-07-22 06:57:50 -07:00
Chenhan D. YuandClaude Opus 4.8 92296aeac2 fix(mmlu): expert-DP batch sharding for MoE + Nano-30B-A3B PTQ example (#1967)
Combines the NemotronH MoE PTQ launcher example with the MMLU
expert-parallel sharding fix it exercises. The example is the end-to-end
test for the fix — quantize + **MMLU EP=4** + export + vLLM smoke on
Nano-30B-A3B — so they ship together.

## Fix — `megatron_mmlu.py` shards over the expert-DP group for MoE
models

`megatron_mmlu` shards whole batches across the **dense data-parallel
group** and runs `megatron_prefill` per-rank on a disjoint subset.
Correct for dense models, but for MoE with **EP>1** the prefill forward
runs an **expert all-to-all across the EP group**. When EP overlaps the
dense-DP group (e.g. `EP=4,TP=1,PP=1` on 4 GPUs), each rank is on a
different batch, so the all-to-alls desync (uneven batch counts +
differing padded seq-lengths) → trailing ranks block at NCCL
communicator creation until the c10d store times out (600 s). Reproduced
on Nano-30B-A3B at `EP=4`: ranks 2 & 3 wait on rank 0's `ncclUniqueId`.

The fix shards over the **expert-data-parallel** group when `EP>1`, so
EP peers evaluate every batch in lockstep and only true expert-DP
replicas take disjoint batches. Dense models (`EP==1`) are byte-for-byte
unchanged.

## Example —
`examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/megatron_lm_ptq.yaml`

Quantize (EP=4) + MMLU gate (EP=4) → export (PP=4) → vLLM smoke, modeled
on the Super-120B example. Uses `NVFP4_DEFAULT_CFG` (no Nano-30B recipe)
and sets `modelopt_install_path` to the nemo-container venv so the
mounted modelopt overrides the container copy. Passes
`test_examples_resolve.py`.

## Test plan

- [x] MoE MMLU (`EP>1`) completes instead of hanging — validated
end-to-end via this example on Nano-30B-A3B `EP=4` (nmm-sandbox CI).
- [x] `test_examples_resolve.py` green (example parses/resolves).
- [ ] Dense MMLU sharding unchanged (`EP==1` path identical).

Supersedes #1964 (example) and #1966 (fix).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Improved MMLU evaluation for expert-parallel MoE models by aligning
batch sharding with expert-parallel execution for more consistent
results.
* **Bug Fixes**
* Dense (non-MoE) models retain standard data-parallel batch sharding,
while MoE handling is corrected.
* **Documentation**
* Updated the Nemotron-3-Nano BF16 PTQ example to clarify the quantize →
MMLU gate → export flow and why the vLLM smoke is intentionally omitted.
* Adjusted the Llama-3.2-1B-Instruct PTQ/MMLU gate lower-bound threshold
(0.40 → 0.36).
* **Tests**
* Marked the homogeneous compressed sharded state-dict test to skip on
Blackwell to avoid a known flaky issue.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 04:12:53 +00:00
h-guo18 d69d5aab8b [Examples]: Kimi-K2.6/K2.7-Code Dflash/Dspark (#1934)
### What does this PR do?

**Type of change:** New feature (recipe + examples)

Adds the DSpark training recipe and Kimi launcher examples for streaming
speculative-decoding training. Split out of #1849 to keep that PR
focused on the DSpark head implementation.

**Files added (5, all additive — no code changes):**

1. `modelopt_recipes/general/speculative_decoding/dspark.yaml` — the
DSpark training recipe. Selects the DSpark head via
`dflash_architecture_config.projector_type=dspark` (DFlash backbone +
lightweight sequential/Markov head + optional confidence head) and sets
the three-term loss weights (`dflash_ce_loss_alpha` /
`dflash_l1_loss_alpha` / `dflash_confidence_head_alpha`).

2–5.
`tools/launcher/examples/moonshotai/{Kimi-K2.6,Kimi-K2.7-Code}/hf_streaming_{dflash,dspark}_multi_node.yaml`
— four multi-node streaming launcher examples mirroring the existing
Kimi-K2.5 streaming format (`common/eagle3/train_eagle_streaming.sh`).
They carry the Kimi-specific base path, draft dims, capture ids, and
mask token; the DSpark examples build on `dspark.yaml` and override only
the Kimi-specific fields (draft dims, `dflash_block_size=8`, mask
token). K2.7-Code shares the K2.6 architecture (`kimi_k25`, 61 layers),
so the draft config is identical.

### Dependency

**Stacked on #1849.** The recipe uses the config fields
(`dflash_ce_loss_alpha`, `dflash_l1_loss_alpha`,
`dflash_confidence_head_alpha`) added in #1849, so this PR is based on
that branch and must merge **after** it. GitHub will auto-retarget the
base to `main` once #1849 merges.

### Testing

- `dspark.yaml` passes the `validate modelopt recipes` schema check
(against the #1849 config schema).
- The launcher examples ship a placeholder `container:
<vllm-image-with-aux-capture-fix>` and are meant to be adapted per
cluster (image, account, partition), not run verbatim in CI.

- Backward compatible?: ✅ (additive recipe + example files only)

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a DSpark speculative-decoding training recipe with configurable
draft, optional per-position confidence, and Markov head components.
* Added new multi-node streaming training example workflows for
Kimi-K2.6 using DSpark/DFlash, including dataset preparation plus
coordinated serve/train settings.
* Configured answer-only loss options, masking behavior, training
hyperparameters, checkpoint/log cadence, and capture-id wiring with
sensible default runtime timeouts.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-07-10 03:09:48 +00:00
Chenhan D. YuandClaude Opus 4.8 adeca905d4 example(megatron_lm): Llama-3.2-1B-Instruct Megatron QAD launcher example (#1931)
### What does this PR do?

Type of change: new example (+ a launcher unit test)

Adds a public **Megatron QAD** (quantization-aware distillation)
launcher example for `meta-llama/Llama-3.2-1B-Instruct`: PTQ (NVFP4) +
MMLU gate → QAD/SFT train, plus a `common/megatron_lm/train/sft.sh`
wrapper (mirrors `common/megatron_lm/quantize/quantize.sh`).

**Reworked — this no longer changes any modelopt code.** The earlier
`_extra_state` plugin patch was reverted: it was the wrong load path
(the crash is the *main* model load in
`megatron/training/checkpointing.py`, not
`_load_extra_state_from_sharded_checkpoint`) and it regressed
quantizer-state loading (Nemotron `quantize_resume` MMLU 0.70 → 0.31).
The correct fix is a config flag on the QAD train step:

- **`--dist-ckpt-strictness log_all`** → the main model load drops only
the genuinely-missing `decoder.final_layernorm._extra_state` (via
Megatron's own unexpected-key set), instead of the torch_dist DCP
planner hard-raising. Present quantizer `_extra_state` is untouched.

Container `nvcr.io/nvidia/nemo:26.06.00` (ships
`nvidia-resiliency-ext>=0.6.0`).

### Testing
- **Integration (GPU/cluster):** validated end-to-end on nmm-sandbox CI
(CLAB-SC-01, computelab): 2/2 QAD tasks passed — quantize+MMLU gate and
QAD train (the train step loads the quantized ckpt past the former
`Missing key … final_layernorm._extra_state` crash). The Nemotron
`quantize_resume` regression is cleared (MMLU back ≥0.70).
- **Unit (CPU):** adds `tools/launcher/tests/test_examples_resolve.py` —
parses every launcher example and checks each task's
`script`/`_factory_`/`args` shape and that `common/*` scripts exist
(guards config regressions). 53 pass. (It also surfaced two pre-existing
broken example script refs — `common/smoke/hostname.sh`,
`common/eagle3/offline_training.sh` — excluded here, worth a separate
fix.)

### Before your PR is "*Ready for review*"
- Backward compatible?: ✅ N/A (new example + new wrapper + new test; no
code paths changed)
- New tests?: ✅ CPU example-resolution test; GPU validation is the
nmm-sandbox integration run
- Changelog?: N/A (example)

### Additional Information
Follow-ups: OMNIML-5421 (make the Megatron-LM example scripts take CLI
args so these `common/*` wrappers can be dropped for a direct 1-shell
call). Supersedes the reverted modelopt-patch approach.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a launch pipeline configuration to run Quantization-Aware
Distillation (QAD) post-training SFT for the Llama 3.2 1B Instruct
model.
* Added a dedicated launcher script to simplify SFT/QAD runs with
sensible defaults and CI-friendly status handling.
* **Tests**
* Added CPU-only checks that validate all launcher example YAML files
are well-formed and reference existing scripts (where applicable),
preventing common runtime misconfigurations.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 09:20:00 -07:00
Chenhan D. Yu 0185726d62 fix: make MCP Slurm wait poll remote job state (#1900)
## Summary
- run managed Model-Optimizer checkouts via `uv run --with-editable` so
the `modelopt-launcher` console script is available
- persist Slurm submission metadata beside the local experiment and make
`job_status` / `wait_for_experiment` prefer remote `sacct` / `squeue`
state over local `_DONE`
- add the missing `common/smoke/hostname.sh` used by
`examples/smoke/hostname.yaml`

## Why
Detached Slurm submission creates the local `_DONE` marker when the
submit process exits, not when the remote Slurm job completes. This made
MCP `wait_for_experiment` return `done` while Slurm jobs were still
running. Reproduced on both MFA and non-MFA clusters.

## Validation
- `uv run --project tools/mcp ruff check
tools/mcp/modelopt_mcp/bridge.py tools/mcp/tests/test_bridge.py`
- `uv run --project tools/mcp pytest tools/mcp/tests/test_bridge.py`
- `uv run --project tools/launcher pytest
tools/launcher/tests/test_core_extended.py
tools/launcher/tests/test_slurm_config.py`
- `uv run --with-editable tools/launcher modelopt-launcher --yaml
tools/launcher/examples/smoke/hostname.yaml --dryrun --yes`

## Manual smoke evidence
- ptyche smoke completed: experiment `cicd_1783037181`, Slurm job
`2320415`, exit `0:0`
- computelab reproduced the wait bug pre-fix: MCP reported done
immediately while Slurm job `2945347` was still running; job later
completed with exit `0:0`


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an automated hostname smoke test to validate launcher
environment readiness.
* Enhanced Slurm-backed job status reporting by persisting scheduler
submission metadata and exposing Slurm state in responses.

* **Bug Fixes**
* Job status now prioritizes authoritative remote scheduler state (with
graceful fallback to local completion/failure markers when unavailable).
* Improved terminal-state handling so experiment polling won’t stop
early when a local done marker exists but the scheduler is still
running.

* **Tests**
* Updated and expanded Slurm status and waiting behavior tests,
including validation of persisted scheduler metadata.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-07-08 09:19:03 -07:00
h-guo18 bc5bc1ac5f [Feat]: Add Final Norm for vLLM Hidden Extractor (#1846)
### What does this PR do?

**Type of change:** Bug fix

vLLM captures the final-layer hidden state *before* the model's final
norm, but the
offline/streaming distillation path fed it straight into `lm_head`, so
the reconstructed
base logits (the KD target) were computed from un-normed hidden states.

This PR re-applies the base model's final norm before `lm_head` when the
producer declares
a pre-norm capture (`base_hidden_prenorm`), for both DFlash and EAGLE:

- Producer sets `base_hidden_prenorm` (streaming: `True`; offline: from
the dump).
- Consumer (`_maybe_apply_base_final_norm`) re-applies the base final
norm, and **fails loud**
if pre-norm is declared but the model's norm type isn't supported (no
silent corruption).
- `FakeBaseModel` now loads the base final norm (+
`rope_theta`/`rms_norm_eps`); norm type is
gated by an explicit `model_type` allowlist (gpt_oss excluded pending a
matching norm class).

### Testing

`tests/unit/torch/speculative/plugins/test_modeling_final_norm.py`;
DFlash/EAGLE streaming
training verified end-to-end.

- Backward compatible?: ✅ (post-norm captures declare
`base_hidden_prenorm=False` → unchanged)
- New tests: ✅


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Expanded speculative decoding/distillation to reconstruct missing
base-model logits by optionally applying a base model’s final
pre–LM-head normalization when pre-norm hidden states are provided.
* Streaming dataset now emits `base_hidden_prenorm` and can configure
RDMA backends from environment settings.
* **Bug Fixes**
  * Rejects mixed `base_hidden_prenorm` values within a batch.
  * Fails fast on streaming token-length mismatches.
* Streaming dataset loading supports directory inputs by expanding
sorted JSONL shards (and errors if none are found).
* **Tests**
* Added/updated unit and dataset tests for final-norm behavior and the
new batch field.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-07-07 05:07:33 +00:00
Chenhan D. YuandClaude Sonnet 4.6 4b9225bee2 launcher: fix host=None when _factory_ is dropped by nemo_run --yaml path (#1842)
## Summary

- **Root cause:** When `launch.py --yaml <config>` processes a
launcher-format YAML, nemo_run constructs `SlurmConfig` directly from
the YAML dict for inline tasks (`task_0`, `task_1`, ...). It strips the
unrecognised `_factory_` key and only sets known `SlurmConfig` fields.
Since `host` is not in the YAML (it should come from the factory),
`SlurmConfig.host` is left as `None`, crashing paramiko at
`Connection(host=None)`:

  ```
  TypeError: expected str, bytes or os.PathLike object, not NoneType
  ```

- **Fix:** In `SandboxPipeline.__post_init__`, after collecting inline
tasks, detect `host` is falsy (symptom of dropped `_factory_`) and apply
the registered `slurm_factory` as base defaults, overlaying
YAML-specified fields (identified by `value != SlurmConfig dataclass
default`).

- **Scope:** Only triggers when `host` is `None`/`""` and
`slurm_factory` is registered. No-op otherwise. Does not affect `task=@`
/ `pipeline=@` paths.

## Architectural note

Three factory resolution paths exist today:
1. `task=@` / `pipeline=@` — nemo_run's `@run.cli.factory` registry ✓
2. `task_configs` list — `_FACTORY_REGISTRY` via
`create_task_from_yaml()` ✓
3. `--yaml` inline tasks — nemo_run's type system, **bypasses all
factory registries** ✗ (fixed here)

## Test plan
- [ ] 65/65 unit tests pass (`uv run python3 -m pytest tests/ -v`)
- [ ] All pre-commit hooks pass (ruff, mypy, bandit)
- [ ] Manually verified: `SandboxPipeline` with `slurm_config.host=None`
(simulating nemo_run dropping `_factory_`) now gets `host` filled in
from the registered `slurm_factory`

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added support for reusing an existing SSH connection for Slurm-based
launches, reducing repeated login prompts during submissions.
* Added a new minimal smoke-test example for hostname checks on CPU-only
clusters.
* **Bug Fixes**
* Improved handling of optional GPU settings so GPU allocation can be
left unset when not needed.
* Fixed Slurm configuration loading so launch settings are preserved
correctly when using YAML-based workflows.
* **Tests**
* Added coverage for SSH reconnect behavior, Slurm launch overrides, and
optional GPU configuration handling.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-02 13:24:18 -07:00
Jenny ChenandClaude Opus 4.8 9038b71f09 Autoquant and GPTQ in support in Megatron-Core [OMNIML-3151] (#1562)
### What does this PR do?

Type of change: New Feature

Autoquant and GPTQ in support in Megatron-Core

- Add EP support to AutoQuantize
- Register MCore support in AutoQuantize
- Add decoder `output_layer` (lm head) to layerwise hook so that GPTQ
can register all decoder layers & lm head
- Split dataloader helper function out of megatron calibration utils so
that AutoQuantize in Megatron-LM can reuse the same dataloader

### Usage
See https://github.com/NVIDIA/Megatron-LM/pull/4821 for Autoquant usage
in Megatron
```python
# For GPTQ pick a recipe that uses gptq algorithm and run mtq.quantize
# e.g. general/ptq/nvfp4_default-kv_none-gptq
```

### Testing
Tested AutoQuant on Nemotron Nano and Ultra.
Tested GPTQ on Nano 3.
Added unit tests for both AutoQuant and GPTQ

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Summary by CodeRabbit

* **New Features**
* Added lazy Megatron-Core AutoQuant integration with Megatron-specific
quantization hooks and better decoder-layer discovery for layerwise
calibration.
* Improved AutoQuantize for expert-parallel (EP) models, including
consistent per-layer recipe selection across DP/TP/EP.
* Extended quant-layer grouping for NemotronH MCore fused
“local_experts” linear layers.

* **Bug Fixes**
* Prevented division-by-zero when calibration inputs are empty during
Hessian updates.

* **Tests**
* Added/extended unit and GPU coverage for EP AutoQuant, decoder-layer
calibration discovery behavior, and zero-token Hessian no-op.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Jenny Chen <jennifchen@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-01 19:12:06 -07:00
Chenhan D. Yu 248cbf2fd2 OMNIML-5128 Capture Docker experiment id (#1840)
## Summary
- Capture nemo-run experiment_id for Docker submit_job by redirecting
detached launcher output to a side-channel log and tailing it briefly.
- Reuse the launcher output parser for both Docker and Slurm submit
paths.
- Document Docker PID + experiment_id behavior and ignore the MCP local
uv.lock.

## Jira
OMNIML-5128

## Validation
- uv run --project tools/mcp pytest tools/mcp/tests -q


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved Docker job submission to capture `experiment_id` from
launcher output during a brief tail window, while always returning the
detached `pid` (and `stdout_log` diagnostics if capture times out).
* Made Slurm identifier extraction more robust across varying launcher
output formats.
* **Documentation**
* Updated `submit_job` docs/module descriptions to clarify Docker
PID/`experiment_id` and `stdout_log` behavior on timeout.
* **Tests**
* Added/expanded tests for Docker success, timeout/no-id scenarios,
parsing edge cases, and log-side-channel failure handling.
* **Chores**
  * Updated local tooling ignore rules for `.venv/` and `uv.lock`.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-06-29 12:19:07 -07:00
ZhiyuandClaude Opus 4.8 f335459dc0 refactor(examples): rename llm_ptq → hf_ptq (symlink for back-compat) (#1759)
## What does this PR do?

**Type of change:** refactor / deprecation (examples)

Follow-up to #1705 (which consolidated `examples/vlm_ptq` into
`examples/llm_ptq`). Since that example now covers Hugging Face **LLM
and VLM** PTQ, the `llm_ptq` name is a misnomer. This renames the
directory to `examples/hf_ptq` and leaves a relative symlink
`examples/llm_ptq → hf_ptq` so existing paths/commands keep working
during a deprecation window.

Requested by @kevalmorabia97 on #1705 (with the symlink-for-back-compat
approach), targeted for the **same 0.46 release** as the consolidation.

### Changes
- `git mv examples/llm_ptq → examples/hf_ptq` and
`tests/examples/llm_ptq → tests/examples/hf_ptq` (the CI runner maps the
matrix name to both `examples/<name>` and `tests/examples/<name>`).
- Add a tracked back-compat symlink `examples/llm_ptq → hf_ptq`.
- Update CI matrices and all repo **path references** (docs, READMEs,
agent skills, launcher/debugger tools, tests) from `llm_ptq` to
`hf_ptq`.
- Keep Python identifiers / test-util module names
(`run_llm_ptq_command`, `llm_ptq_utils`) — they name the LLM-PTQ task,
not the directory.
- Preserve the CODEOWNERS team slug
(`modelopt-examples-llm_ptq-codeowners`) and historical CHANGELOG
entries; add a CHANGELOG deprecation note.

### Back-compat caveats (inherent to git directory symlinks)
- ✅ Linux/macOS CLI usage and Python `cwd`/pytest resolution work
through the symlink.
- ⚠️ Windows git checkouts don't materialize symlinks by default (low
impact — this example is Linux-only in practice).
- ⚠️ GitHub web doesn't follow directory symlinks, so legacy external
deep-links to `examples/llm_ptq/...` won't navigate in. All **internal**
references are repointed to `hf_ptq`, so the symlink is only for legacy
external/CLI use.

### Usage (unchanged via symlink)
```bash
# New canonical path
cd examples/hf_ptq
scripts/huggingface_example.sh --model <hf_model> --quant fp8

# Old path still works (forwards via symlink)
cd examples/llm_ptq && scripts/huggingface_example.sh --model <hf_model> --quant fp8
```

### Testing
- `bash -n` on moved/edited shell scripts (new path + via symlink).
- `py_compile` on moved/edited Python; test re-export shim repointed to
`examples/hf_ptq/example_utils`.
- Verified git tracks `examples/llm_ptq` as a single symlink (mode
120000), not a duplicated tree (no pre-commit / pytest
double-processing).
- `pre-commit run` on all changed files passes.

### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (relative symlink keeps
`examples/llm_ptq` paths valid; see caveats above)
- Did you write any new necessary tests?: N/A (pure rename; existing
tests moved with the dir)
- Did you update Changelog?: ✅

### Additional Information
Follow-up (later release): remove the `examples/llm_ptq` symlink once
external references have migrated.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* PTQ guidance now directs to the unified Hugging Face PTQ flow,
including VLM quantization via the shared `--vlm` entry point.
* **Documentation**
* Updated README and guide links, references, and command snippets to
use `hf_ptq` (replacing `llm_ptq`).
* Deprecated and consolidated `vlm_ptq` into `hf_ptq`; removed
VILA/NVILA coverage from the Hugging Face PTQ examples.
* **Bug Fixes**
* Improved detection and routing so local/manual setup uses the correct
PTQ source.
* **Tests / Chores**
  * CI and example tests updated to run the `hf_ptq` variants.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-27 08:48:48 +00:00
c248dd5434 [Feat]: Domino support (#1710)
### What does this PR do?

Type of change: New feature

Adds **Domino** speculative decoding: the parallel DFlash draft backbone
plus a lightweight **GRU causal correction head**. The backbone produces
*base* logits for a full draft block in one forward; a GRU over the
block's teacher-forced tokens produces a causal state that is fused with
the backbone hidden state and projected to a vocab-sized logit
correction on the block suffix — injecting the intra-block causal
dependency the parallel backbone lacks. Trained with a dual loss
`(1-λ)*final + λ*base`, where `λ_base` decays linearly 1→0 (curriculum:
learn the parallel backbone first, then the correction).

Reuses the DFlash mode/config/recipe; selected via
`dflash_architecture_config.projector_type=domino` and routed to its own
registry so `HFDominoModel` does not shadow `HFDFlashModel`. Exports in
the z-lab/SpecForge drafter format (`prefix_gru.*` / `embed_proj.*`).

> Note: the inference side (vLLM / AR evaluation) is intentionally
**not** wired up yet — the correction head is not applied in serving. To
be added once the inference path lands.

### Usage

```bash
# Online training (recipe: projector_type=domino)
uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_online_domino.yaml --yes
```

### Testing

CPU unit tests in
`tests/unit/torch/speculative/plugins/test_hf_domino.py` cover
conversion routing, the training forward (dual loss + grads), the λ
schedule, and the export format. Online Qwen3-8B training validated
end-to-end (loss curve below).

<img width="1803" height="809" alt="image"
src="https://github.com/user-attachments/assets/7c9d2001-bd80-4dec-919b-443e61089cca"
/>

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (opt-in via
`projector_type=domino`; DFlash path unchanged)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new
dependency)
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

Reference: SpecForge PR #571 (z-lab); drafter format
`huggingface.co/Huang2020/Qwen3-8B-Domino-b16`.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added Domino speculative-decoding training with a decaying base/final
dual-loss curriculum and Domino-specific lambda scheduling
(training-only; inference wiring not yet included).
  * Added Domino draft-head export support for training checkpoints.

* **Documentation & Configuration**
* Added a Domino speculative-decoding training recipe and an HF Online
Domino launcher configuration for Qwen3-8B.

* **Refactor**
* Updated speculative model conversion/export to route to Domino
variants based on the configured projector type.

* **Tests**
* Added unit tests for Domino conversion, training loss/metrics, lambda
decay behavior, and exporter output layout.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-06-27 07:34:50 +00:00
Keval MorabiaandClaude Sonnet 4.6 33bfa8b1fe CI/Dev env bump (#1818)
### What does this PR do?

Type of change: chore

Bumps CI/dev tooling and test containers.

**Container bumps**
- NeMo test containers → 26.06
- TRT-LLM container → 1.3.0rc19
- transformers max version → 5.12

**Dev tooling bumps**
- ruff bump 0.12.11 → 0.15.18
- mypy 1.17.1 → 2.1.0: enable new defaults (`local_partial_types`,
`strict_bytes`); fix/narrow the errors newly surfaced by mypy 2.0 in 4
modules (rather than blanket-suppressing them); remove 2 stale `# type:
ignore` comments
- pre-commit 4.3.0 → 4.6.0
- sphinx 8.1 → 9.1 + sphinx-rtd-theme 3.0 → 3.1: add `suppress_warnings
= ["ref.python"]` to fix cross-reference ambiguity error new in sphinx
9.x
- trl fix for newly released 1.7 version

**Bug fixes surfaced by the bumps**
- sparsity (weight): make the weight mask DTensor-aware under FSDP. The
transformers→5.12 bump routes the HF Trainer FSDP optimizer-state save
through torch's DTensor-based `get_optimizer_state_dict`, which
triggered `aten.mul.Tensor got mixed torch.Tensor and DTensor` in the
dynamic `weight` getter. The mask is now distributed to the weight's
mesh/placements before masking, cached, and rebuilt only when the
sharding changes (invalidated on `set_mask`). Fixes the `llm_sparsity`
example test.

### Testing

- `pre-commit run --all-files` ✅ (including mypy 2.1.0)
- `nox -s docs` ✅
- `tests/unit/torch/sparsity` + `tests/unit/torch/nas` ✅
- `llm_sparsity` GPU example test (FSDP path) verified in CI

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ✅
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Summary

* **Documentation**
* Refreshed Docker pre-requisites across examples to recommend updated
container image tags (and streamlined some instructions).
* **Bug Fixes**
* Improved sparse weight mask handling for DTensor/FSDP by aligning and
caching distributed masks.
  * Made TensorRT engine byte retrieval return immutable `bytes`.
* Reduced Sphinx cross-reference warnings and tuned Transformers
compatibility warning thresholds.
* **Tests**
  * Increased default unit test timeout on Windows runners.
* **Chores**
* Updated CI workflow container tags and refreshed linting/typing/docs
version pins, plus related mypy configuration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-27 01:00:25 +05:30
Chenhan D. Yu 40936649a3 launcher: move NVIDIA-Nemotron-3-Super-120B YAML from Nemotron-h/ to nvidia/ (#1815)
Moves
`tools/launcher/examples/Nemotron-h/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/specdec_bench_mtp_vllm.yaml`
to
`tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/specdec_bench_mtp_vllm.yaml`
to match the directory convention for NVIDIA-published models (same
family as other models under `nvidia/`).

Also updates the `--yaml` path in the header comment.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated the example launch command in the YAML header to reference the
current NVIDIA example path.
* Simplified the commented command formatting into a single line to
improve readability and copy/paste convenience.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan D. Yu <5185878+ChenhanYu@users.noreply.github.com>
2026-06-24 22:39:31 +00:00
Chenhan D. Yu 7c5741be0c Fix launcher Slurm mounts in installed MCP mode (#1811)
## Summary
- make Slurm ModelOpt source overlay mounts conditional on
source-checkout mode
- force managed-source MCP launches to reinstall `modelopt-launcher`
from the selected checkout so source refs do not reuse a stale cached
package
- add regression coverage for installed-mode Slurm mounts and
managed-source launcher argv construction

## Root cause
PR #1799 correctly stopped packaging `modules/Model-Optimizer/*` when
`modelopt-launcher` runs from an installed package. However,
`build_slurm_executor()` still unconditionally added container mounts
for `code/modules/Model-Optimizer/modelopt` and `modelopt_recipes`. In
installed MCP mode those paths do not exist in the remote package, so
the container runtime fails before the job script starts.

A second issue appeared during validation: the MCP managed-source path
can materialize the right git checkout but still execute a cached
`modelopt-launcher` package with the same version. Adding `uv run
--reinstall-package modelopt-launcher` ensures the selected source ref
is what actually runs.

## Validation
- `uv run pytest tests/test_bridge.py -q` from `tools/mcp`: 51 passed
- `uv run pytest tests/test_slurm_executor.py tests/test_core.py -q`
from `tools/launcher`: 24 passed
- `pre-commit run --files tools/launcher/core.py
tools/launcher/tests/test_slurm_executor.py
tools/mcp/modelopt_mcp/bridge.py tools/mcp/tests/test_bridge.py`: passed
- Live Slurm GPU smoke validated with patched launcher path;
`nvidia-smi` ran successfully and the smoke script completed.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Refactor**
* Restructured container mount assembly for Slurm job execution to
conditionally mount ModelOpt directories based on optional source path
parameter.
* Enhanced launcher command-line generation with package management
improvements.
* Replaced unconditional mount paths with conditional behavior for more
flexible resource utilization.

* **Tests**
* Expanded container mount scenario test coverage for installed and
source execution modes.
* Tightened test assertions for comprehensive mount behavior
verification.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-06-24 13:12:57 -07:00
Jenny ChenandClaude Opus 4.8 e2c4d083d4 [OMNIML-4922] Four over Six PTQ & Updating Nemotron Ultra Example (#1684)
### What does this PR do?

[Four Over Six](https://arxiv.org/pdf/2512.02010) PTQ implementation for
weight-only quantization. Four Over Six was used to produce the Nemotron
3 Ultra NVFP4 checkpoint. Also updates the Ultra PTQ example in the
launcher to use this new 4/6 config
`huggingface/nvidia/Nemotron-3-Ultra-550B-A55B/ptq/ultra-nvfp4-46-max`

### Usage

```bash
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/megatron_lm_ptq.yaml --yes
```

### Testing

- [x] Unit tests pass
- [x] Run launcher example

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added NVFP4 Four‑Over‑Six (4/6) adaptive per‑block weight scaling and
a configurable FP8 normalization option for FP4/FP8 quantization.

* **Documentation**
* Added PTQ recipes/presets and updated config docs to document FP8 max
variants and the Four‑Over‑Six option.

* **Tests**
* Added unit and GPU tests validating 4/6 selection, normalization
threading, scaling behavior, and reconstruction error checks.

* **Chores**
* Updated a PTQ pipeline example to use the NVFP4‑46‑max recipe and
bumped a container image; minor project config tweak.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Jenny Chen <jennifchen@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 15:00:25 +00:00
Chenhan D. Yu 37dbbdac5a Fix ModelOpt MCP Slurm launcher submit (#1799)
## Summary
- fix launcher Slurm task annotation patching so nemo-run CLI resolves
`slurm_factory` correctly for task slots
- harden `modelopt-mcp` submit parsing/status resolution and add
regression coverage for launcher false-positive success cases
- add a minimal `nvidia-smi` smoke YAML/script and fix launcher
packaging so source-backed Slurm jobs package required files recursively

## Validation
- `uv run pytest tests/test_core.py -q`
- `uv run pytest tests/test_bridge.py -q`
- dry-run and live-submit validated through the patched local MCP server
on `cw_dfw`
- interactive smoke job succeeded end-to-end (`nvidia-smi` ran
successfully in-container)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an NVIDIA SMI GPU smoke test (script and minimal Slurm YAML
example) for launcher integration.
* **Bug Fixes**
* Improved detection of fatal launcher errors, including when the
launcher exits with code 0.
* Strengthened Slurm experiment/job identifier parsing and added early
rejection of unsafe experiment IDs, with clearer “unparsed”/failure
behavior.
* Updated sandbox task Slurm config type handling and improved launcher
packaging so examples/common are included consistently.
* **Tests**
* Expanded unit and filesystem-based coverage for parsing/validation,
dry-run fatal stderr handling, and nested experiment directory layouts.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Signed-off-by: Chenhan D. Yu <chenhany@nvidia.com>
2026-06-24 01:05:59 +05:30
Jenny Chen 28b5e26fdb Fix torch import error to remove circular dependency & move Nemotron configs (#1606)
### What does this PR do?

Type of change:  Bug fix
when running Megatron-LM modelopt example `generate.py` a circular
import causes an error. the cause was because it import
modelopt.torch.quantization which imports modelopt/torch -->
```
modelopt/torch/__init__.py:26. The chain is: import modelopt.torch.opt → runs modelopt/torch/__init__.py (parent package) → which pulls in distill
  → distill/mode.py:25 → back into opt before opt.utils is bound.
```

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Clarified configuration comments in Nemotron-3-Super-120B quantization
recipes for improved clarity.

* **Chores**
  * Internal package initialization updates.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
2026-06-23 18:21:17 +00:00
Chenhan D. Yu 985ad4361c launcher: make Slurm memory and user defaults configurable (#1791)
## Summary

- Adds `SlurmConfig.mem` and a `SLURM_MEM` env override for launcher
Slurm jobs.
- Passes configured memory through to `nemo_run.SlurmExecutor` instead
of always using `"0"`.
- Lets `SLURM_USER` provide the launcher default user when local and
cluster usernames differ.
- Adds focused tests for Slurm memory defaults and overrides.

## Test plan

- [x] `uv run pytest tools/launcher/tests/test_slurm_config.py
tools/launcher/tests/test_slurm_executor.py`


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Slurm job memory can now be configured via `SLURM_MEM`, with a default
of `0` when unset.
* Job launch username can now be set via `SLURM_USER`; if omitted, it
falls back to the local login name.

* **Bug Fixes**
* Slurm executor now correctly uses the configured memory value from
Slurm settings, falling back to `0` only when missing/empty.

* **Tests**
* Expanded unit tests to cover default, environment-driven, and executor
parameter memory behavior (including missing/empty/`None` cases).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-06-22 14:39:31 -07:00
Chenhan D. YuandClaude Sonnet 4.6 2a88b6056c launcher: package as modelopt_launcher; mcp: call console script directly (#1766)
## Summary

- `tools/launcher/__init__.py` gains `PACKAGE_DIR`, making the directory
an importable package named `modelopt_launcher` via a `package-dir`
mapping in `pyproject.toml`. **No file moves** — `common/` and
`examples/` stay exactly where they are.
- `pyproject.toml` adds the `packages`/`package-dir`/`package-data`
declarations and a `modelopt-launcher` console script entry point.
- `launch.py` switches imports to
`modelopt_launcher.{core,slurm_config}` (works in both `uv run
launch.py` via editable install and the installed console script), adds
`_has_modelopt_src` to skip packaging modelopt source when running
installed (cluster container already has it), and adds `main()`.
- `bridge.py` deletes the 75-line `_find_launcher_dir()` filesystem
walker and `_launcher_dir_not_found_response()`; simplifies
`_find_launcher_examples_dir()` to 2 strategies (env override → `import
modelopt_launcher`); switches `submit_job` subprocesses from `["uv",
"run", "launch.py"]` with `cwd=launcher_dir` to `["modelopt-launcher"]`
with no `cwd`.
- `tools/mcp/pyproject.toml` declares `modelopt-launcher` as a proper
dependency (dev: editable `../launcher`; published: PyPI), replacing a
lengthy comment explaining why it could not be declared.
- 7 tests for the deleted functions removed; all 38 remaining tests
pass.

## Test plan

- [ ] `cd tools/launcher && uv run python3 -m pytest tests/ -v` — all 65
tests pass
- [ ] `cd tools/mcp && uv run python3 -m pytest tests/ -v` — 38 tests
pass
- [ ] `uv run launch.py --yaml
examples/Qwen/Qwen3-8B/megatron_lm_ptq.yaml --dryrun --yes -v` — dry-run
resolves correctly
- [ ] `modelopt-launcher --yaml
examples/Qwen/Qwen3-8B/megatron_lm_ptq.yaml --dryrun --yes` — console
script works after `pip install -e tools/launcher`
- [ ] `cd tools/mcp && uvx modelopt-mcp` — MCP server starts without
FileNotFoundError

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added a `modelopt-launcher` CLI command as the standardized entry
point for launching optimization jobs.

* **Bug Fixes**
* Improved launcher detection and error reporting when the launcher
isn’t installed.
* Simplified experiment-directory discovery and standardized environment
handling for job submission and log retrieval.

* **Chores**
* Updated launcher packaging and example/resource discovery to work
reliably from installed distributions.
  * Added dev-mode symlink and cleanup safeguards.

* **Tests**
  * Adjusted expectations for draft PR failure behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-22 12:07:44 -07:00
Chenhan D. Yu 12ae5fbb35 [OMNIML-5233] hf_synth.yaml: relative-leaf output_dir contract (#1773)
## What does this PR do?

**Type of change:** Bug fix

**Overview:** Fix the `output_dir` shape for
`tools/launcher/examples/Qwen/Qwen3-8B/hf_synth.yaml` so the downstream
synth_run stage (OMNIML-5233) can submit + resume + publish correctly.
Replaces the prior `/scratchspace/modelopt/qwen3-8b-synth-v1` (which
fails resume + publish) with the canonical **relative-leaf** shape:
`hf-local/modelopt/qwen3-8b-synth-v1`.

The absolute team-folder prefix is **deliberately not committed** — it's
an NVIDIA-internal cluster mount path, and `pensieve-intern`'s
`_scan_for_internal_path_leak` guard refuses any YAML diff that bakes it
(rightly, since this repo is public). The downstream stage that has the
cluster context (synth_run) resolves the prefix via
`mcp__nmm-sandbox__resolve_team_folder` at submit time and injects the
absolute path through `extra_overrides`.

**Why this shape:** the relative leaf `hf-local/modelopt/<id>` carries
the contract between authoring (this stage), submitting (synth_run), and
publishing (`modelopt-storage publish`). Publish promotes from the
team's `hf-local/` tree as a metadata-only op; the leaf shape tells the
operator + downstream what category the artifact lands in, without
leaking the cluster mount.

**Companion MR:** [pensieve-intern
!126](https://gitlab-master.nvidia.com/omniml/integration/pensieve-intern/-/merge_requests/126)
lands the matching SPEC contract in synth_support.md + synth_run.md, AND
elevates the path-contract rule to the engine preamble
(`_ENGINE_AGENT_RULES_PREAMBLE`) so every agent dispatch — across
agent/subprocess/pensieve-artifact tasks — reads it as workflow-wide
common knowledge.

## Usage
Same as before. The shard write location is now resolved at submit time
by synth_run.

## Testing
- [x] One-line YAML change.
- [ ] End-to-end: re-fire OMNIML-5233 synth_run after this + the
pensieve-intern MR merge; expect the agent to submit the slurm array and
stamp success.

## Before your PR is "Ready for review"

- [x] Make sure you read and follow Contributor guidelines and your
commits are signed.
- [x] Is this change backward compatible? **Yes** (Qwen3-8B is the only
consumer; no published checkpoints reference the old path).
- [x] Did you write any new necessary tests? **N/A** — example YAML;
behavior covered by the downstream synth_run agent run.
- [x] Did you add or update any necessary documentation? Covered in the
comment block on the changed lines.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Updated the configured output directory to a more stable,
persist-friendly path so generated artifacts remain available across
repeated runs and re-dispatches.
* Added clearer inline guidance explaining how the resolved absolute
path under the new directory is used to reuse prior shards and to handle
publishing consistently.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-06-19 16:18:46 -07:00
Chenhan D. YuandClaude Sonnet 4.6 fa1d13f85d launcher: add Nemotron-3-Super-120B-A12B-BF16 MTP vLLM specdec bench config (#1714)
Adds a SPEED-bench MTP speculative-decoding YAML for
`NVIDIA-Nemotron-3-Super-120B-A12B-BF16` via vLLM.

Covers two splits:
- `qualitative` — 32 concurrent, 4096 output tokens
- `throughput_32k` — 8 concurrent, 80 requests, 4096 output tokens

Both tasks run `tp_size=4` on a single 4×H100/A100 node.

Part of OMNIML-5095 / OMNIML-5098.

## Test plan
- [ ] `uv run slurm.py --yaml
modules/Model-Optimizer/tools/launcher/examples/Nemotron-h/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/specdec_bench_mtp_vllm.yaml
--dryrun --yes -v`

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Chores**
* Added a new benchmark configuration for evaluating NVIDIA
Nemotron-3-Super-120B model performance using speculative decoding with
vLLM. Configuration includes qualitative and throughput benchmark tasks.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-19 16:18:28 -07:00
93dd08f429 [OMNIML-4760] synth_support (#1696)
Draft PR opened by **pensieve-intern** for
[OMNIML-4760](https://jirasw.nvidia.com/browse/OMNIML-4760).

Stage `synth_support` of Epic `OMNIML-4755`. The agent ran from the SPEC
on the ticket description; review every change before marking ready.

_Always-draft is enforced — the bot never auto-merges._

---

**Agent's self-narration** (stripped from PR diff; surfaced here for
context):

`VERIFICATION_COMMENT.txt`:
```
OMNIML-4760 synth_support verification for qwen3-8b found and fixed one task_0 issue in `tools/launcher/examples/Qwen/Qwen3-8B/hf_offline_eagle3.yaml`.

What changed:
- Updated `task_0` data input from missing `/hf-local/modelopt/Speculative-Decoding-Prompt-Samples` to existing `/hf-local/modelopt/Speculative-Decoding-Dataset-v1-Qwen3-8B/default-openai.jsonl`.
- Left the rest of the 4-task monolithic `hf_offline_eagle3.yaml` unchanged.

Verification:
- `task_0.script` is `common/tensorrt_llm/query.sh`.
- `--model <<global_vars.hf_model>>` resolves to `/hf-local/Qwen/Qwen3-8B`.
- `task_0.slurm_config.container` is `nvcr.io/nvidia/tensorrt-llm/release:1.2.0`.
- Cluster validation on cw_dfw succeeded for the fixed task_0 data path: experiment `cicd_1781221901`, Slurm job `12739762`, remote directory `/lustre/fsw/portfolios/coreai/users/chenhany/experiments/cicd/cicd_1781221901/Qwen3-8B_EAGLE3_offline_task0_verify_jsonl_0`.
- Log evidence included successful SSH tunnel/authentication, TensorRT-LLM 1.2.0 container import, Qwen3-8B server health checks, a successful `/v1/chat/completions` request, and loading `1393367` train examples from the Qwen3-8B JSONL data file.

Next:
- Runner should open the single-file PR for human review because this was a task_0 config fix, not verification-only.
```



_Pollution-strip removed `VERIFICATION_COMMENT.txt` from this commit
(sidecar narration and/or incidental lockfile regeneration are never
part of the agent's intended deliverable)._

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Enhanced TE grouped MoE weight quantization with optional
per-GEMM/per-expert quantizers (environment-controlled) and updated
calibration flow
* Expanded launcher CLI with `--shard-id` and per-run sample limiting
via `--num-samples`
* Added a Qwen3-8B standalone vLLM synthesis job example
(`hf_synth.yaml`)
  * Added Slurm job requeue support
* **Bug Fixes**
* Improved sharded dataset synthesis with idempotent `.done` markers and
smarter shard sizing/capping
* **Tests**
* Added coverage validating TEGrouped vs sequential MoE default amax
behavior and output divergence/accuracy
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Pensieve Intern <pensieve-intern@nvidia.com>
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Pensieve Intern <pensieve-intern@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-18 20:35:10 -07:00
Chenhan D. Yu e0125294f9 launcher: add Qwen3-8B/specdec_bench_dflash_vllm.yaml parent (OMNIML-5057) (#1764)
## Summary

Adds the parent YAML for the Qwen3-8B / DFlash / vLLM SPEED-bench sweep
(Epic OMNIML-5057, 4 cells: t0_d3 / t0_d7 / t1_d3 / t1_d7). Mirror of
[`Qwen/Qwen3.5-4B/specdec_bench_dflash_vllm.yaml`](https://github.com/NVIDIA/Model-Optimizer/blob/main/tools/launcher/examples/Qwen/Qwen3.5-4B/specdec_bench_dflash_vllm.yaml).

### Differences vs Qwen3.5-4B template

- `hf_model`: `/hf-local/Qwen/Qwen3-8B`
- `draft_model`: `/hf-local/modelopt/Qwen3-8B-DFlash-bs8-seq4096-150000`
— an internal NVIDIA checkpoint, not on HF Hub. `cell.md` Step 1's
`PRIVATE_DRAFT_ORGS` branch already covers the expected 404; the cluster
has the file staged by the Epic's `prep_inputs` stage.
- `block_size`: 8 (was 4) — matches the canonical `cell_t0_d7`
draft_length=7.

### Known limitation (acceptable per Shape (2))

Qwen3-8B's `max_position_embeddings = 40960`. The `throughput_32k` split
contains rows whose prompts exceed 40960 input tokens; those rows fail
with the vLLM context-limit assert. Per `cell.md` "Shape (2)" recovery
contract, cells ship qualitative metrics + `null` throughput_32k AL. See
OMNIML-5060 Notes (2026-06-15 correction) for the empirical evidence.

### Reference run

`cicd_1781655951` (OMNIML-5060 pipeline #55054230, cw_dfw):
- qualitative `Average_AL` = **3.4849**
- throughput_32k = `null` (Shape (2) — overlong rows skipped)

### Cascade

3 subprocess cells (OMNIML-5059 / 5061 / 5062) re-run slurm against this
parent YAML at invoke time with per-cell CLI overrides for temperature /
block_size / save_dir / max_seq_len.

## Test plan

- [x] `tools/precommit/check_launcher_yaml.py` passes locally
- [x] YAML successfully drives a real slurm submission
(`cicd_1781655951`) producing valid qualitative SPEED-bench output
- [ ] Once merged: `intern_advance OMNIML-5057` fires the 3 subprocess
cells against this YAML

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Added benchmark configuration file for evaluating speculative decoding
performance with Qwen3-8B model using DFlash acceleration.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-06-18 11:40:00 -07:00
h-guo18 e6790ef7b4 [Examples]: GPT-oss, Qwen3Moe streaming specdec example (#1692)
### What does this PR do?

Type of change: new example

Adds **streaming speculative-decoding examples (EAGLE3 + DFlash)** for
**gpt-oss-20b** and **Qwen3-30B-A3B** to the ModelOpt launcher,
mirroring the existing Qwen3-8B/Kimi examples.

- New yamls:
`tools/launcher/examples/{openai/gpt-oss-20b,Qwen/Qwen3-30B-A3B}/hf_streaming_{eagle3,dflash}_multi_node.yaml`,
plus gpt-oss `chat_template_train.jinja` (generation-tagged, for
`answer_only_loss`).
- `eagle_utils.py`: the streaming path now installs a custom
`data.chat_template` on the tokenizer (the online path already did) —
needed for the tagged template.

### Usage

```bash
cd tools/launcher
export SLURM_HOST=... SLURM_ACCOUNT=... SLURM_HF_LOCAL=... SLURM_JOB_DIR=...
uv run launch.py --yaml examples/openai/gpt-oss-20b/hf_streaming_eagle3_multi_node.yaml --yes
```

### Testing

Pipeline sanity test on **unsynthesized** data (daring-anteater), 1×
H100-80GB, 12k steps. All four train and pass the vLLM acceptance-length
eval:

| Model | Method | Train speed | vLLM AL |
|---|---|---|---|
| Qwen3-30B-A3B | EAGLE3 | 7.12 it/s | **1.74** |
| Qwen3-30B-A3B | DFlash | 2.31 it/s | 1.29 |
| gpt-oss-20b | EAGLE3 | 5.07 it/s | 1.19 |
| gpt-oss-20b | DFlash | 2.01 it/s | 1.14 |

<img width="1300" height="780" alt="image"
src="https://github.com/user-attachments/assets/2be22562-5b77-4a9f-9dc0-6f936a059736"
/>

> Sanity test only, not a quality run. gpt-oss AL is low because it is a
reasoning model (CoT at inference) while daring-anteater has no
reasoning traces and `answer_only_loss` masks all but the final content
— quality runs need synthesized/reasoning data.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added multi-node speculative decoding pipeline configurations for
Qwen3-30B-A3B and gpt-oss-20b with DFlash and EAGLE3 support.
* Introduced chat template training support for improved model
instruction formatting.

* **Enhancements**
* Increased benchmark concurrency from 1 to 32 across Qwen3-8B
configurations for more realistic performance evaluation.
  * Extended training runs from 500 to 2000 steps for Kimi-K2.5 models.
  * Improved chat template handling in speculative decoding workflows.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-06-15 16:49:39 -07:00
yeyu-nvidiaandClaude Opus 4.6 e004d8d90e DFlash speculative decoding for MiniMax-M2.7 (FSDP2): auto mask-token, FSDP2 resume fixes, per-checkpoint draft export (#1621)
## What

Brings up DFlash block-diffusion speculative decoding for large MoE
targets (MiniMax-M2.7, 229B) trained under accelerate FSDP2, and fixes
the regressions that broke checkpoint resume and per-checkpoint draft
export.

## Commits
- **auto-add mask token for DFlash** when the tokenizer lacks one
(resize embeddings, restore dtype).
- **requeue support** in `build_slurm_executor` + **FSDP2
cpu_ram_efficient_loading** for 229B on multi-node.
- **FSDP2 buffer patch** (`fsdp2_buffer_patch.py`): handle non-DTensor
buffers in `fsdp2_load_full_state_dict`, broadcast dtype codes from rank
0, and an FSDP2-safe `clip_grad_norm_`. Required because MiniMax-M2.7
pins transformers 4.57.x (no native `ParallelismConfig`).
- **dtype fix**: use the broadcast dtype (rank 0) rather than the local
meta-device param dtype, so non-leader ranks don't cast bf16 back to
fp32 on resume.
- **restore `DFlashExportCallback`** (this PR's headline): the
Pydantic-recipe refactor (7038dec918) dropped the callback that exported
the draft submodule after each checkpoint save, leaving a stale "export
happens during training via DFlashExportCallback" comment with no
callback. FSDP2 SHARDED_STATE_DICT checkpoints carry no
`model.safetensors`, so without it there is nothing for vLLM /
acceptance-length eval to load. The callback gathers only the ~328 MB
draft submodule across shards via `get_model_state_dict(...,
submodules={dflash_module}, full_state_dict=True, cpu_offload=True)` —
works under SHARDED_STATE_DICT without materializing the 229B base — and
writes `exported-checkpoint-{step}/`.

## Testing
- Resume from FSDP2 sharded checkpoints verified end-to-end (loss/AR
continuity).
- Draft export validated against vLLM: exported drafts load and produce
acceptance-length metrics on MT-Bench across the full checkpoint sweep.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Export draft-submodule weights to dedicated exported checkpoints
during training.
* FSDP2 buffer compatibility and DTensor-aware gradient clipping for
safer distributed loading/training.
  * Detect HF-format checkpoints for smarter resume/load behavior.
  * Auto-add and handle a mask special token for draft workflows.
  * vLLM: disable prefix caching to preserve full prompt hidden states.
  * Add CLI option for answer-only loss and save aligned loss masks.
  * Add a SPEED-Bench config for DFLASH/vLLM benchmarking.

* **Chores**
  * Robust package version fallback to avoid import failures.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-06-15 16:47:50 -07:00
Chenhan D. Yuandpensieve-intern agent 601401b134 [OMNIML-5025] cell_t0_d7 (#1738)
Draft PR opened by **pensieve-intern** for
[OMNIML-5025](https://jirasw.nvidia.com/browse/OMNIML-5025).

Stage `cell_t0_d7` of Epic `OMNIML-5022`. The agent ran from the SPEC on
the ticket description; review every change before marking ready.

_Always-draft is enforced — the bot never auto-merges._

---

**Agent's self-narration** (stripped from PR diff; surfaced here for
context):

`INTERN_ARTIFACTS.json`:
```
{
  "sweep_name": "gemma-4-E4B-it_mtp_vllm_t0_d7",
  "experiment_id": "cicd_1781548540",
  "experiment_dir": "/lustre/fsw/portfolios/coreai/projects/coreai_dlalgo_modelopt/cicd/cicd/cicd_1781548540/",
  "AL_qualitative_overall": 3.2945,
  "AL_qualitative_categories": {
    "coding": 4.8883,
    "humanities": 2.703,
    "math": 4.1139,
    "multilingual": 4.1401,
    "qa": 2.59,
    "rag": 3.9329,
    "reasoning": 3.6251,
    "roleplay": 1.8026,
    "stem": 3.2041,
    "summarization": 2.8275,
    "writing": 2.4118
  },
  "AL_throughput_32k_overall": 3.3803,
  "AL_throughput_32k_categories": {
    "high_entropy": 2.0839,
    "low_entropy": 4.3797,
    "mixed": 3.7143
  }
}
```

`VERIFICATION_COMMENT.txt`:
```
Completed OMNIML-5025 cell_t0_d7 for `google/gemma-4-E4B-it` / MTP / vLLM.

What was done:
- Authored/kept `tools/launcher/common/specdec_bench/_cells/gemma-4-E4B-it_mtp_vllm_t0_d7.yaml` for sweep `gemma-4-E4B-it_mtp_vllm_t0_d7`.
- Updated/kept `tools/launcher/examples/google/gemma-4-E4B-it/specdec_bench_mtp_vllm.yaml` with the Gemma-specific `vllm/vllm-openai:gemma` container after generic v0.22.1 failed startup.
- Verified HF Hub model endpoints for `google/gemma-4-E4B-it` and `google/gemma-4-E4B-it-assistant` returned HTTP 200.
- Wrote `INTERN_ARTIFACTS.json` with parsed AL metrics from the successful cluster run.

Metrics extracted:
- experiment_id: `cicd_1781548540`
- experiment_dir: `/lustre/fsw/portfolios/coreai/projects/coreai_dlalgo_modelopt/cicd/cicd/cicd_1781548540/`
- qualitative Average_AL: `3.2945`
- throughput_32k Average_AL: `3.3803`

PR status:
- PR opened: NONE — this runner prompt says not to commit, push, or create PRs because the runner handles that.

What's next:
- Engine should consume `INTERN_ARTIFACTS.json` and advance the downstream wrap-up/aggregation stage.

Trigger pipeline URL:
- `https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/pipelines/54841359`

Slurm job status:
- submission: completed successfully
- experiment_id: `cicd_1781548540`
- experiment_dir: `/lustre/fsw/portfolios/coreai/projects/coreai_dlalgo_modelopt/cicd/cicd/cicd_1781548540/`
```



_Pollution-strip removed `INTERN_ARTIFACTS.json`,
`INTERN_LEARNING_TICKET.md`, `VERIFICATION_COMMENT.txt` from this commit
(sidecar narration and/or incidental lockfile regeneration are never
part of the agent's intended deliverable)._

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Updated benchmark configuration to use an optimized container image
for Gemma 4 MTP speculative-decoding benchmarks.
* Revised documentation in the configuration to reflect the current
image and its capabilities.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: pensieve-intern agent <noreply@nvidia.com>
2026-06-15 20:09:36 +00:00
Chenhan D. Yu c661366352 launcher: Nemotron-120B specdec_bench — kill stale _cells/ reference (post-PR-#1564) (#1741)
## What

Comment-only change to the Nemotron-3-Super-120B-A12B-BF16 DFlash parent
YAML doc-comment header. Rewrites the example invocation from the
pre-PR-#1564 pattern (with `--runtime_params
common/specdec_bench/_cells/<sweep_name>.yaml`) to the current
post-#1564 pattern (CLI-flag-only overrides, no per-cell file).

No yaml content / config change — only the prose example in the header.

## Why

PR #1564 removed the `tools/launcher/common/specdec_bench/_cells/` and
`_runtime_params/` directories. Cell-specific knobs are now CLI
overrides at slurm-invoke time, not committed files. The cell SPEC in
pensieve-intern (`specdec_bench/specs/cell.md`) is explicit about this —
five times in prose.

But this parent YAML's doc-comment kept advertising the OLD pattern.
When pensieve-intern's agent runner scans `tools/launcher/examples/` for
a reference invocation (good practice — agents should imitate working
examples), it lands on this Nemotron parent and copies the stale
pattern. **Five recent agent dispatches on the gemma-4 Epic OMNIML-5022
(cells OMNIML-5024 / 5025 / 5026 / 5027) authored new
`_cells/<sweep>.yaml` files for this reason, despite the SPEC telling
them not to.**

Prose loses to a concrete checked-in counter-example. The Qwen3.5-4B
reference template that cell.md officially points at
(`tools/launcher/examples/Qwen/Qwen3.5-4B/specdec_bench_mtp_vllm.yaml`)
is clean and shows only the bare `uv run slurm.py --yaml ...` form. This
PR makes Nemotron-120B consistent with that template.

## How surfaced

Diagnosed 2026-06-15 on OMNIML-5025 cell_t0_d7 ([intern-agent job
341631795](https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/341631795)).
The cell agent authored `_cells/gemma-4-E4B-it_mtp_vllm_t0_d7.yaml` —
the engine's diff-shape classifier should have rejected it, but didn't
(tracked separately as OMNIML-5170). Root-causing the agent's behavior
surfaced this stale doc-comment as the source of the pattern.

## Verification

`grep -rn '_cells\|runtime_params common'` across the entire launcher
tree returned only this file. After this PR, the launcher tree carries
zero stale references.

Pairs with NVIDIA/Model-Optimizer#1738 (gemma-4-E4B-it container fix)
and OMNIML-5170 (engine-side classifier hard-reject for `_cells/` paths
— defense in depth).

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated example documentation to clarify how to override per-cell
parameters using CLI flags in Slurm configuration runs.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-06-15 12:58:14 -07:00
Keval MorabiaandClaude Opus 4.8 55a2101e2f Update Nemotron-3 Pruning, Distillation and PTQ results based on new shared calibration loop with seq packing and add tool-calling eval fix (#1660)
### What does this PR do?

Type of change: documentation + minor example-script tweaks

Follow-up to #1601. Originally scoped to add **NVFP4 + QAD**, this PR
was **repurposed** to refresh the [Nemotron-3-Nano-30B-A3B-BF16
tutorial](examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md)
results using the **new shared calibration loop (sequence packing)** and
to **fix tool calling in evaluation**.

- Refreshed the prune → distill → eval → **FP8** results (accuracy +
vLLM throughput tables) with the new calibration loop.
- **Tool-calling eval fix** (`nemo_evaluator.yaml`): GPQA and AIME now
run the Python sandbox tool. The tutorial reports both **with-tools**
and **no-tools** GPQA/AIME and shows `mean ± std_dev`.
- Script tweaks: `quantize.py` calibration now uses sequence packing
(`pack=True`) which leads to slight improvement in PTQ;
`prune_minitron.py` defaults `inference_batch_size` to
`calib_batch_size`.

### Testing

Documentation + small example-script changes; tutorial relative links
resolve and the results tables / figure were verified consistent.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: Yes
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ (will run `/claude review`)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated the main guide and evaluator instructions for prune + distill
+ FP8/NVFP4 quantization, including refreshed vLLM deployment tips,
benchmark/noise presentation, and long-context tool-calling attribution
notes.
* Refreshed README technique examples/links, reordered the model support
matrix rows, and improved pruning overview/support-matrix text.

* **Changes to Examples**
* NAS pruning now documents higher GPU memory usage vs manual pruning;
pruning batching defaults were improved.
* Quantization PTQ calibration uses packed document packing; quantized
checkpoint export messaging was streamlined.
* Updated pruning/distillation/quantization tutorial guidance,
metrics/tables, command parameters, and evaluator YAML settings
(KV-cache dtype, generation defaults, task behavior).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 18:32:13 +00:00
Chenjie LuoandClaude Opus 4.8 d7df14d12a [tools/debugger] Enforce a single relay owner across hosts (#1735)
### What does this PR do?

Type of change: Bug fix

The `tools/debugger` file-based relay assumes a single server but never
enforced it.
Because the relay lives on shared NFS (the repo is often the same
checkout mounted on
multiple hosts), a forgotten `server.sh` on another host kept polling
the same
`.relay/` and could **steal commands** (executing them on the wrong
host), and killing
one server's cleanup could **wipe the active server's markers**.

This adds a `.relay/owner` ownership token (`host:pid:nanos`):

- Each server writes `owner` atomically at startup and **takes over**
instead of
refusing when a stale `server.ready` exists (the old `kill -0 <pid>`
guard was
  host-local and meaningless across hosts).
- The handshake and main loops exit cleanly if `owner` changes
(`[server] Superseded by <id> — exiting.`), so a freshly started server
**evicts**
  any stale one — even on another host.
- `cleanup()` only clears shared markers if we still own them, so a
stepping-down
  server never clobbers its successor's `server.ready`/`owner`.

Also gitignores `tools/debugger/logs/` and documents the `owner` file in
the README.

### Usage

```bash
# Inside the container; a previously-running server elsewhere that shares this
# NFS .relay/ steps down automatically once this one claims ownership:
bash tools/debugger/server.sh
# [server] Note: existing server.ready found (<host:pid:ts>); taking over.
#   (the stale server logs: "[server] Superseded by <id> — exiting.")
```

### Testing

Verified live on computelab: a forgotten `server.sh` on another host was
evicted when
a new server started, after which `client.sh run` executed on the
correct (new) host;
confirmed the stepping-down server's cleanup does not remove the
successor's
`server.ready`/`owner`. `server.sh` passes `bash -n`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ <!-- additive: new .relay/owner
file; client.sh and the wire protocol are unchanged -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- the file-based relay
tool has no test harness; behavior verified manually -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- internal dev tooling, not a shipped feature/API -->
- Did you get Claude approval on this PR?: N/A <!-- can run /claude
review -->

### Additional Information

Scope is limited to `tools/debugger/` (`server.sh`, `README.md`,
`.gitignore`).


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated relay protocol documentation to clarify ownership-based server
coordination.

* **Bug Fixes**
* Improved reliability of multi-server coordination in shared relay
environments.

* **Chores**
  * Updated ignore patterns for logging files.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 18:27:52 +00:00