mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
main
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c2aaa44f60 |
[Speculative Decoding] DFlash2 draft variant (grouped sublayer convolution + candidate selector) (#2216)
### What does this PR do? Type of change: new feature Adds **DFlash2** ([blog](https://inco.ai/blog/dflash2/)) as a draft variant of the existing DFlash mode, selected with `dflash_architecture_config.projector_type="dflash2"` alongside `domino`, `dspark` and `lilicorr`. DFlash2 keeps DFlash's one-pass parallel backbone and adds two components that recover the acceptance a purely parallel draft loses: - **Grouped dynamic depthwise convolution** around every attention and MLP sublayer, giving each block position a view of its predecessors *inside* the block. Taps do not cross the block boundary, so the draft stays one forward pass. - **Low-rank candidate selector** scoring transitions between adjacent positions' top-k candidates, so serving walks one coherent path instead of taking an independent argmax per position. Both start as exact no-ops — the convolution's `base_kernel` is an identity and `kernel_projection` is zeroed; the selector's `successor_codebook` is zeroed — so a freshly built DFlash2 draft *is* its DFlash backbone, and enabling the variant is an extension rather than a perturbation. This matches the reference implementation ([SpecForge#772](https://github.com/sgl-project/SpecForge/pull/772), merged) and the way `modeling_lilicorr` already installs this same convolution class. **This unblocks a recipe already shipped on `main`.** `modeling_lilicorr._install_sublayer_convs` imports `DFlashGroupedConv` from `modeling_dflash2`, so `modelopt_recipes/general/speculative_decoding/lilicorr_conv.yaml` raises at model build today and its CHANGELOG entry documents a feature that cannot run. Landing this makes it runnable. Module and parameter names match the SGLang/vLLM `DFlash2DraftModel` loaders. Verified against the released `z-lab/Qwen3.8-27B-DFlash2` checkpoint: **81 tensors, 21 name patterns, zero difference in either direction**. The serving side, [vllm-project/vllm#52816](https://github.com/vllm-project/vllm/pull/52816), has since merged (`b389ac294`) with no change to the checkpoint contract. ### Usage ```bash python examples/speculative_decoding/main.py \ --config modelopt_recipes/general/speculative_decoding/dflash2.yaml \ model.model_name_or_path=Qwen/Qwen3-8B \ data.data_path=<corpus>.jsonl \ training.output_dir=<out> ``` ```yaml # modelopt_recipes/general/speculative_decoding/dflash2.yaml dflash: dflash_selector_loss_alpha: 1.0 # weight of the candidate-selector CE term dflash_architecture_config: projector_type: dflash2 conv_kernel_size: 2 # taps; must not exceed the block size conv_group_size: 16 # must divide hidden_size selector_rank: 256 selector_top_k: 16 ``` ### Testing <img width="2000" height="1320" alt="image" src="https://github.com/user-attachments/assets/869004d1-92c6-41a4-a03d-fd824a06255c" /> **Unit** — 25 CPU tests in `tests/unit/torch/speculative/plugins/test_hf_dflash2.py`; the full `tests/unit/torch/speculative/` suite passes with no regressions. The ones worth keeping pin invariants that a decreasing loss does not catch: - the convolution is an exact identity on the **default** construction, and its taps stay inside the block while a position still sees its predecessors; - the block-offset contract shared by the training objective and `CandidateSelector.greedy_path` — a misaligned objective still converges; - the export fields the vLLM loader requires, including the top-level `block_size` that `DFlash2Exporter` derives the nested copy from; - which selector factors receive gradient on the first step. `successor_codebook` starts at zero, so `predecessor_codebook` and `hidden_projection` take one step to begin moving. That is a warm start, not a dead branch, and both sides are asserted. **End-to-end** — trained on Qwen3-8B against a plain DFlash control with every other argument identical (plot above). Monotonic convergence, no NaN/divergence, no DDP unused-parameter issues. Note the losses are **not comparable** across arms: DFlash2's includes the selector CE term. **Serving (vLLM)** — the exported drafter loads and drafts under the merged DFlash2 path (`RESOLVED draft architectures: ['DFlash2DraftModel']`). Two notes for anyone reproducing: vLLM sizes the convolution from `1 + num_speculative_tokens` at runtime rather than from the checkpoint, so a `block_size=16` drafter is only correct at `num_speculative_tokens=15`; and at that value the upstream path currently hits an illegal memory access in `_cache_draft_logits` ([vllm#55279](https://github.com/vllm-project/vllm/issues/55279)), independent of which checkpoint is used. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — additive. New `projector_type`, its own registry and exporter, one new config field; DFlash / Domino / DSpark / LiLiCorr numerics and `state_dict` contents are untouched. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — `modeling_dflash2.py` is adapted from [SpecForge#772](https://github.com/sgl-project/SpecForge/pull/772) and carries its MIT notice. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — under `0.48.0`. - Did you get Claude approval on this PR?: ✅ — run on 2026-08-20; all review threads addressed and resolved. ### Additional Information Rebased onto current `main`. Two commits from the original branch were dropped because [#2342](https://github.com/NVIDIA/Model-Optimizer/pull/2342) landed them first, with authorship preserved: the no-op sublayer seam in `modeling_dflash.py`, and the `rope_theta`/`rope_parameters` fix — `main`'s version of the latter is stricter, so this PR no longer touches `hf_dflash.py` at all. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added DFlash2 speculative decoding with grouped dynamic convolutions and low-rank candidate selection. * Added configurable selector-loss weighting, including an option to disable it. * Added DFlash2 model conversion and export support. * Added checkpoints compatible with SGLang and vLLM DFlash2 serving. * **Documentation** * Added training recipes and a Qwen3-8B online DFlash2 training configuration. * **Tests** * Added coverage for conversion, training, metrics, gradients, and export compatibility. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d0142c9dca |
Reuse a whole recipe via $import, deprecate recipe_type, and start the published-checkpoint backfill with two aliases (#2376)
### What does this PR do?
**Type of change:** New feature (recipe loading) + one bug fix
Two things, the second built on the first:
1. **A recipe can now reuse another recipe whole.** A top-level
`$import` brings in the imported recipe's entire body; keys given
alongside it override the imported ones. `metadata.recipe_type` becomes
optional and is deprecated along the way.
2. **The deprecated `recipe_type` is swept out of every shipped recipe,
and the checkpoint backfill starts with two published checkpoints
recorded as aliases** that reuse a portable recipe wholesale — the first
users of the alias mechanism — plus a fix to two existing Nemotron NVFP4
recipes.
#### Declaring what kind of recipe a file is
`load_recipe` read `metadata.recipe_type` out of the raw YAML *before*
resolving imports, because it needs the schema class to hand to
`load_config`. That made the field impossible to inherit, so a recipe
reusing another had to restate a line it could only have copied.
It is now optional, and the loader takes the first of these that
answers:
1. a `# modelopt-schema:` comment naming the recipe's schema class,
2. `metadata.recipe_type` — **deprecated**; still read and still
honoured, so a recipe outside this repo keeps working unchanged,
3. the recipe it delegates to via a top-level `$import`.
Whatever a recipe *does* state must be true, in both directions. A
schema comment contradicting a `recipe_type` is rejected, and so is a
recipe importing a different kind of recipe — that used to surface as
whatever pydantic made of, say, an `eagle` section spliced into a PTQ
schema. The concrete recipe classes carry a `RECIPE_TYPE` ClassVar as
the single source of truth.
Only a recipe that another file **imports** needs the schema comment —
that is what `$import` resolution requires to validate the payload. The
sweep here drops `metadata.recipe_type` from all 78 shipped recipes that
carried it and gives the imported ones a `# modelopt-schema:` comment
instead, so nothing in-tree depends on the deprecated field.
A directory recipe's `metadata.yml` resolves its kind the same way —
schema comment first, `recipe_type` as the fallback — it just has no
`$import` to delegate through, since a directory recipe has no body of
its own to hand off. (Follow-up commit, after this PR's initial review:
it originally still required `recipe_type` unconditionally, the one
place the deprecation didn't reach.)
#### Checkpoint aliases
Two checkpoints NVIDIA has published in quantized form use a scheme a
portable recipe already produces, with no checkpoint-specific deviation,
so each is recorded as a thin **alias** (top-level `$import`, overriding
only `metadata`) at its own model-hub path -- the *source* checkpoint's
path, not the published quantized one's:
-
**`models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast`**
delegates to `general/ptq/nvfp4_experts_only_mse-kv_fp8_cast` —
expert-only NVFP4 (MSE static weights, dynamic inputs) with an FP8 KV
cache in cast mode — published as `nvidia/Kimi-K2.6-NVFP4`.
-
**`models/Qwen/Qwen3.5-397B-A17B/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8`**
delegates to the `qwen3_5_moe` architecture recipe
`model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4
(MSE static weights) on the routed experts, ModelOpt-default FP8
elsewhere, FP8 KV cache — published as
`nvidia/Qwen3.5-397B-A17B-NVFP4-V2`.
(Follow-up commit, after this PR's initial review: the Qwen entry
originally lived at `models/nvidia/Qwen3.5-397B-A17B/` -- nvidia is the
*published* checkpoint's org, not Qwen3.5-397B-A17B's own. Moved to
match the source model's actual hub path, same as the Kimi-K2.6 entry
above.)
Editing the base recipe changes every alias that points at it; nothing
is duplicated.
#### One fix
- **The Nemotron-3 Super and Ultra NVFP4 recipes** quantized the MTP
block on the **Megatron-Core** path, where it is a live `model.mtp`
submodule their broad `*mixer.*` patterns matched into, contrary to
their own descriptions. They now disable `mtp.*` explicitly. Hugging
Face runs were unaffected — `NemotronHPreTrainedModel` sets
`_keys_to_ignore_on_load_unexpected = [r"mtp.*"]` and builds no MTP
module.
### Usage
A checkpoint alias resolves through `--recipe` to the recipe it
delegates to:
```bash
python examples/hf_ptq/hf_ptq.py \
--pyt_ckpt_path <checkpoint> \
--recipe models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast \
--export_path <output>
```
A recipe that reuses another whole — the shape the aliases use:
```yaml
imports:
base: general/ptq/nvfp4_experts_only_mse-kv_fp8_cast
$import: base
metadata:
description: What this checkpoint uses the base recipe for.
```
### Testing
- **`tests/unit/recipe/test_loader.py`** — 28 new cases covering
whole-recipe reuse with no `metadata` at all; kind resolution from each
of the three sources, from a delegation chain and from a `$import` list;
a delegation cycle failing with `ValueError` rather than recursing;
`peek_declared_schema` including a comment placed below the first YAML
line; `recipe_type` being optional, filled per class, and rejected when
it contradicts; a directory recipe resolving its kind from a schema
comment the same way, rejecting a comment/`recipe_type` disagreement,
and still requiring one or the other; and delegating across kinds being
an error.
- **`tests/unit/recipe/test_recipe_docs.py`** — the
model-specific-recipe check now also covers the two new alias folders,
which must be listed in `ptq.md` like every other
`models/<org>/<model_id>` entry.
- **Recipe validation** (`tools/precommit/check_modelopt_recipes.py`)
and **`pre-commit`** pass on the changed files. The full
`tests/unit/recipe/` suite is left to CI — a broken `transformer_engine`
in the local dev venv keeps the `mtq.quantize`-based cases from running
there.
Not covered: **numerics**. Nothing here asserts accuracy, or that
running one of these recipes reproduces a released checkpoint's weights.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `metadata.recipe_type` is
still read and honoured for recipes outside this repo, the schema
comments are inert for direct loads, and the loader change only relaxes
a check.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — two feature entries, one deprecation, and one bug fix under 0.48.0.
- Did you get Claude approval on this PR?: ❌ — not yet run.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Recipes can delegate configurations, support checkpoint aliases, and
apply local metadata overrides.
* Recipe types can be inferred from schema declarations or delegated
recipes, with stronger consistency validation.
* Added unquantized KV-cache options, layerwise export, broader operator
calibration, and new PTQ examples.
* Added checkpoint-specific recipes and MLflow experiment references.
* **Bug Fixes**
* Improved ONNX calibration, FSDP2 export, and fused-MoE quantization
handling.
* Nemotron-3 recipes keep MTP blocks in BF16.
* **Documentation**
* Expanded guidance for aliases, delegation, schema declarations, and
recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
c4d00f7150 |
Move Nemotron Nano offline KD example (#2471)
### What does this PR do? Type of change: Bug fix <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Very short summary of changes only for new features, backward breaking changes, deprecations, or fixes for critical bugs present in previous releases. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the NVIDIA Nemotron offline usage example to use the BF16 source-directory pipeline path. * Pipeline configuration remains unchanged. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> |
||
|
|
216f28a6e0 |
Consolidate speculative-decoding agent skills into one stage/algorithm tree (#2201)
### What does this PR do?
Type of change: documentation
Reorganizes the EAGLE3 agent skills into a single speculative-decoding
skill, then adds algorithm sheets for DFlash, DSpark, and Domino.
**The problem.** The four `eagle3-*` skills each baked the algorithm
into a *stage* of the same draft-model pipeline:
```
skills/eagle3-new-model/ skills/eagle3-review-logs/
skills/eagle3-triage/ skills/eagle3-validate/
```
Adding DFlash would have meant four more near-duplicate skills, since
the stages are shared and only the algorithm differs.
**The change.** One skill dir shaped like `ptq/` (SKILL.md +
references/), split along the two real axes:
```
plugins/modelopt/skills/speculative-decoding/
├── SKILL.md # router: stage table x algorithm table
└── references/
├── stages/ # the procedure — algorithm-independent
│ ├── configure.md # <- eagle3-new-model
│ ├── review-logs.md # <- eagle3-review-logs
│ ├── triage.md # <- eagle3-triage
│ └── validate.md # <- eagle3-validate
└── algorithms/ # the data sheet — per-algorithm
├── README.md # contract: 6 required sections
├── eagle3.md
├── dflash.md
├── dspark.md # DFlash variant — delta only
└── domino.md # DFlash variant — delta only
```
Stage docs cite algorithm-sheet sections by heading (*Pipeline tasks*,
*Success markers*, *Quality gate*, *Known failures*, ...), so a new
algorithm means one new file plus a table row — no stage edits. Every
recipe in `modelopt_recipes/general/speculative_decoding/` now has a
sheet.
DSpark and Domino are documented as **DFlash variants**, not separate
pipelines: same `recipe_type: speculative_dflash`, same training script,
same `dflash.*` config namespace, selected by
`dflash_architecture_config.projector_type`. Their sheets carry only the
delta.
Writing the sheets surfaced three things the old EAGLE3-only skills got
wrong or missed:
- **Task counts are not fixed.** The old skills hardcoded "4-step
pipeline, task_0 through task_3". DFlash offline is 2 tasks, DFlash
online is 3, Domino is 2. The stage docs no longer assume a count.
- **`--aux-layers` couples the dump to the draft.** For DFlash the
dump's layer count must equal the draft's `num_hidden_layers`; a
mismatch doesn't error, it silently captures the wrong layers. Recorded
under *Known failures*.
- **In-training AR is meaningless for DSpark and Domino.** Both recipes
pin `estimate_ar: false` / `ar_validate_steps: 0` because eval runs the
DFlash backbone with the new head bypassed. Each sheet says so under
*Quality gate* so nobody reads a backbone-only number as a result.
**Behavior change:** the four `/eagle3-*` slash commands are replaced by
one `/speculative-decoding`. This isn't optional —
`tools/precommit/sync_claude_skills.sh` iterates `.agents/skills/*/` one
level deep and plugin discovery is `skills/<name>/SKILL.md`, so a
directory is either one skill or a container of skills, not both.
`tools/launcher/docs/claude_code.md` is updated accordingly.
### Usage
```
/speculative-decoding
```
Or by description — the skill triggers on EAGLE3 / DFlash / DSpark /
draft model / acceptance rate. For a new model, follow the stages in
order:
```bash
# 1. Configure: copy the closest examples/<Org>/<Model>/hf_<mode>_<algo>.yaml and adapt
cd tools/launcher
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml --dryrun # preview
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml --yes # submit
# 2. review-logs -> 3. triage (if anything failed) -> 4. validate
```
### Testing
- `claude plugin validate . --strict` and `claude plugin validate
plugins/modelopt --strict` — both pass
- `pre-commit run --files ...` over all changed files — passes,
including `markdownlint-cli2` and the `sync-claude-skills` symlink hook
(it agrees with the new `.claude/skills/speculative-decoding` symlink)
- Verified the new skill is discovered and its description loads
- Every relative link across the skill tree resolves; every repo path
cited in the sheets exists; no dangling `eagle3-*` reference remains
anywhere in the repo
- Each factual claim in the sheets was checked against its source — the
launcher example YAMLs, the four recipes, `dflash_online_training.sh`,
`vllm_smoke_test.sh`, `check_regression.py`, and
`plugins/hf_{dflash,dspark,domino}.py` — rather than written from memory
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ — the four `/eagle3-*` slash
commands become `/speculative-decoding`. Agent tooling only; no library
or API surface is touched. The three YAML comment fixes are
comment-only, no behavior change.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — documentation; covered
by plugin validation and pre-commit
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — agent tooling only, matching how #2025 handled it
- Did you get Claude approval on this PR?: ✅ — `/claude review` run; 2
findings, both fixed in `9b3c568`
### Additional Information
Follows #2025, which moved the skill tree into the installable plugin.
Two stale in-repo comments were found while sourcing the sheets, and are
**fixed in this PR** (`c2696b1`, comment-only):
1. `modelopt_recipes/general/speculative_decoding/dflash.yaml` pointed
`chat_template` at a `chat_templates/` directory under
`modelopt_recipes` that does not exist — templates live per-model beside
each launcher example.
2. Both offline DFlash example YAMLs annotated `--aux-layers dflash`
with "Must match the draft model's num_hidden_layers". `--aux-layers` is
a preset keyword accepting only `eagle`, `dflash`, or an explicit id
list, so it carries no count. The constraint is real but belongs to the
draft depth the preset resolves to: `--num-draft-layers` on the vLLM
dump, and no override at all on the HF/TRT-LLM dumps, which hardcode 5
via `resolve_aux_layers`. This comment had already misled this PR's own
first draft, which is why it's fixed rather than just documented.
Because of (1), this PR now touches `modelopt_recipes/`, which adds
**@NVIDIA/modelopt-recipes-codeowners** to the required reviewers.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added comprehensive speculative-decoding guidance for configuration,
training, validation, troubleshooting, and supported algorithms.
* Added workflow references for DFlash, Domino, DSpark, and EAGLE3,
including quality checks and failure diagnosis.
* **Documentation**
* Generalized experiment-log review and pipeline triage across
algorithms.
* Clarified DFlash draft-depth configuration, resource sizing, task
recovery, and validation.
* Replaced the EAGLE3-specific workflow entry with the broader
speculative-decoding workflow.
* Removed standalone EAGLE3 skill documentation as guidance is now
consolidated under speculative decoding.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
8025a3dc54 |
specdec(recipe): add MiniMax-M2.7-DFlash streaming multi-node pipeline (#1835)
## Summary - Add `hf_streaming_dflash_multi_node.yaml` for MiniMax-M2.7 (229B MoE) streaming DFlash training - 2 serve replicas (TP=4, whole node) + 2 trainer nodes (4 GPU each) over NIXL RDMA hidden-state transport - Capture IDs `[2,17,32,47,62,64]` from `build_target_layer_ids(64, 5)` + final layer output - MiniMax-specific: trust_remote_code, FSDP2 via accelerate config, mask_token=200054, YaRN rope_scaling factor=48 - Topology matches Kimi-K2.5 large-MoE streaming recipe Resolves OMNIML-5221 ## Test plan - [ ] Dry-run validation (`uv run launch.py --yaml ... --dry-run`) - [ ] Server-only smoke on CW-DFW (task_1 with `training.max_steps=1`) - [ ] Full streaming training run Signed-off-by: Ye Yu <yeyu@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a multi-node launcher configuration for MiniMax-M2.7 streaming training with speculative decoding. * Added an end-to-end workflow for conversational data preparation, distributed training, checkpoint export, and vLLM smoke testing. * Supports configurable serving and training settings across multiple nodes. * **Enhancements** * Applies Transformers version overrides only to trainer or single-node environments. * Preserves the serving environment’s Transformers version on dedicated serving nodes. * Uses the launcher’s built-in bfloat16 mixed-precision configuration for training. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
2d35643452 |
LiLiCorr training (#2342)
### What does this PR do? Type of change: new feature Adds **LiLiCorr**, a candidate-lattice reranker for DFlash drafts, as a new `projector_type` on the existing `dflash` mode — plus three DFlash-wide improvements that apply to every variant, and an optional composition with DFlash2's grouped convolutions. A DFlash drafter is trained on per-position marginals rather than on the joint block distribution, so its drafted tokens are individually plausible yet jointly incoherent. LiLiCorr keeps the top-`k` candidates the backbone already produces at each block position, scores transitions between adjacent candidates with a small two-layer transformer, and commits a path through the lattice greedily. Serving is unchanged in kind: verify still checks every drafted token against the target, so the emitted distribution is untouched and only acceptance length moves. - Paper: [LiLiCorr: Lightweight Likelihood Correlation of Parallel Drafts for Speculative Decoding](https://arxiv.org/abs/2608.20530) (arXiv:2608.20530) - Blog: https://research.nvidia.com/labs/nemotron/lilicorr/ - **Companion PR — serving support:** [sgl-project/sglang#37462](https://github.com/sgl-project/sglang/pull/37462) This PR is the **training** half. It trains the drafters and exports them; the companion PR above is what serves the resulting checkpoints, and is what the comparison table below was measured through. **What is in the commits** | | | | --- | --- | | LiLiCorr draft variant | `hf_lilicorr.py`, `modeling_lilicorr.py`, conversion routing, config fields, export | | Three DFlash-wide features | fp32 master weights for the draft, draft activation checkpointing, and a DDP hang fix — all default-off or behaviour-preserving, all applying to `dflash`, `domino`, `dspark` and `dflash2` alike | | Optional grouped convolutions | composes LiLiCorr with DFlash2's `DFlashGroupedConv`; see the dependency note below | | Two recipes | `lilicorr.yaml` and `lilicorr_conv.yaml` | | CPU unit tests, CHANGELOG, one launcher example | | **⚠️ The convolutions depend on the DFlash2 branch, and cannot run until it merges.** `modeling_lilicorr.py` imports `DFlashGroupedConv` from `modeling_dflash2`, which today exists only on `haoguo/dflash2-support`. The class is **imported rather than copied on purpose** — it is the only way the two variants cannot drift apart arithmetically — but the consequence is that the convolutional recipe cannot run against `main` as it stands. So the import is **deferred into `_install_sublayer_convs`** rather than taken at module scope. Everything else in this PR, including the plain LiLiCorr reranker, has no DFlash2 dependency at all and works on `main` today; an eager import would have made the whole plugin unimportable for the sake of one optional feature. Requesting the convolutions without DFlash2 present raises an `ImportError` naming the two config keys to remove, rather than failing at import time. **This PR carries two of @h-guo18's commits, with authorship and sign-off preserved.** Both are independent of DFlash2 itself and both are needed here: - `1419d47e`, the no-op sublayer seam. Without it `DFlashDecoderLayer.forward` never calls the wrappers the convolutions install onto, so the modules would be built, counted and exported while computing nothing. It is arithmetically an identity on its own. - `ba377e7a`, the RoPE-θ fix. On Transformers 5 a config carries both a top-level `rope_theta` and a `rope_parameters` dict; the real base lives in the dict while the class default (10,000 for Qwen3) stays visible as the flat attribute. Reading the flat field first builds a draft whose RoPE base is 100× off a Qwen3-8B target's, which trains and exports without complaint. Both the training-side enforcement and the exporter's `_get_rope_theta` are affected on `main` today. Both are @h-guo18's work and belong to their branches; they are carried here only so that this PR stands on its own. **If those branches land first, this PR can be rebased onto them and the two commits dropped**, and they can equally be split out now if that is easier to review. The same applies to `dflash_fp32_master_weights`, which is also in flight on `haoguo/dflash-fp32-master-weights`. The field name is shared deliberately so that there is only ever one knob rather than two spellings of it, and both versions default to off. Whichever lands first, this PR can be rebased onto it. ### Usage Train with the shipped recipe: ```python from modelopt.recipe import load_recipe config = load_recipe("general/speculative_decoding/lilicorr.yaml") # Qwen3-8B target, 6 epochs, block size 16 (15 drafted slots, 16 verified), # DFlash decay objective at gamma 7.0, fp32 master weights for the draft. ``` Or convert directly: ```python import modelopt.torch.speculative as mtsp config = { "dflash_block_size": 16, "dflash_loss_objective": "decay", "dflash_loss_decay_factor": 7.0, "dflash_fp32_master_weights": True, "dflash_lilicorr_w_ce": 0.25, "dflash_lilicorr_w_margin": 0.0, "dflash_lilicorr_w_pen": 0.25, "dflash_architecture_config": { "num_hidden_layers": 5, "projector_type": "lilicorr", "lilicorr_candidate_topk": 8, # Optional, and all-or-nothing: adding these two keys wraps every draft # sublayer in DFlash2's grouped convolution. Requires the DFlash2 variant. # "conv_kernel_size": 2, # "conv_group_size": 16, }, } mtsp.convert(model, [("dflash", config)]) ``` ### Results Six drafters for a **Qwen3-8B** target, all trained **in ModelOpt on one matched contract** — the same corpus, schedule and block geometry for every arm, so no row carries a training advantage. Training data is NVIDIA's [Nemotron Post-Training Dataset v2](https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2) with the multilingual split excluded, generated from the target with **thinking disabled**; **6 epochs**; block size 16 (15 drafted slots, 16 verified); DFlash decay objective at gamma 7; **8 nodes × 8 H100, global batch size 64** (one sequence per device, no gradient accumulation). All six were then exported and served through SGLang on a **single H100 80GB**, `tp_size 1`, at concurrency 1, greedy, `fa3`, mean of two replicates, with the whole node held exclusive per benchmark. Speedup is output tokens/s against an autoregressive baseline measured in the same allocation. Cells are `acceptance length / speedup-vs-AR`; **★ fastest, ☆ second fastest**: | benchmark | LiLiCorr+conv | LiLiCorr | DSpark | DFlash2 | Domino | DFlash | |---|---|---|---|---|---|---| | gsm8k | ★ 7.715 / 5.26x | ☆ 7.557 / 5.22x | 7.375 / 4.86x | 7.252 / 5.06x | 7.225 / 4.87x | 6.341 / 4.59x | | math500 | ★ 9.241 / 6.54x | ☆ 9.064 / 6.52x | 9.012 / 6.15x | 8.999 / 6.49x | 8.976 / 6.25x | 7.909 / 5.88x | | aime25 | ★ 8.285 / 6.03x | ☆ 8.156 / 6.03x | 8.043 / 5.61x | 7.967 / 5.91x | 8.066 / 5.77x | 7.126 / 5.44x | | humaneval | ★ 7.393 / 4.01x | 7.077 / 3.93x | 7.163 / 3.72x | ☆ 7.081 / 3.95x | 6.864 / 3.73x | 6.156 / 3.68x | | mbpp_sanitized | ★ 5.999 / 4.18x | ☆ 5.849 / 4.13x | 5.888 / 3.95x | 5.685 / 4.05x | 5.679 / 3.91x | 5.027 / 3.70x | | livecodebench | ★ 7.975 / 5.40x | ☆ 7.754 / 5.33x | 7.775 / 5.10x | 7.601 / 5.26x | 7.553 / 5.04x | 6.808 / 4.88x | | alpaca_eval | ☆ 3.697 / 2.69x | ★ 3.656 / 2.70x | 3.588 / 2.52x | 3.467 / 2.58x | 3.627 / 2.59x | 3.222 / 2.46x | | mtbench | ★ 4.014 / 2.94x | ☆ 3.939 / 2.93x | 3.957 / 2.78x | 3.748 / 2.80x | 3.948 / 2.84x | 3.478 / 2.67x | **Against every other approach in the table, LiLiCorr with convolutions is the fastest on all eight benchmarks.** Plain LiLiCorr is the fastest on seven of the eight; the exception is humaneval, a 164-prompt slice, where DFlash2 is ahead by 0.5%. `DFlash` is the deliberately head-free control; every head clears it by +7.60% to +21.67% on acceptance, which is the check that a head actually loaded. Reproducing the `LiLiCorr+conv` column additionally needs the DFlash2 variant. Acceptance length is bit-reproducible under greedy decoding and its replicate spread here was 0.00% on every benchmark; throughput has a ~0.2% floor. ### What `dflash_fp32_master_weights` does, and what it is worth Today the draft is cast to the frozen base model's dtype — bf16 — before the optimizer is built. AdamW then allocates its moments with `zeros_like(p)`, so the **optimizer state becomes bf16 too**. That is the problem: bf16 has too few mantissa bits to represent the small updates Adam's second moment accumulates, so those updates round away and the effective step size decays on its own, independently of the learning-rate schedule. The flag is standard mixed precision instead: the draft's master weights stay in fp32 while the matmuls run in bf16. It requires a bf16 autocast around the forward, which HF `Trainer` supplies under `TrainingArguments.bf16`. Paths that do not go through the Trainer — evaluation, `pseudo_speculative_generate`, a plain `convert()` and forward — currently need the caller to supply it, and no shipped recipe exercises those (`estimate_ar: false`, `do_eval: false`). Making the draft supply its own autocast is a follow-up, held back from here on review because it touches every DFlash variant and wants e2e coverage of the existing recipes. Compute speed is unchanged. The cost is memory, about 12 bytes per parameter for the weight plus Adam's two moments instead of 6, plus a doubled gradient all-reduce under DDP, since fp32 parameters mean fp32 gradients. Under FSDP2 that second cost is what `MixedPrecisionPolicy(reduce_dtype=...)` exists to control. It is worth **7 to 14 percent of acceptance length**, measured at the end of training on gsm8k, and it helps every projector type: | arm | bf16 | fp32 | Δ acceptance length | | --- | ---: | ---: | ---: | | LiLiCorr | 6.8670 | 7.5573 | **+10.05%** | | DFlash2 | 6.7396 | 7.2518 | **+7.60%** | | Domino | 6.5854 | 7.2252 | **+9.71%** | | DSpark | 6.4621 | 7.3752 | **+14.13%** | | DFlash | 5.9030 | 6.3412 | **+7.42%** | Every arm in the comparison table above was trained with it on, and **both shipped recipes set it `true`**, so the documented path gets it. It defaults to **off**, so no existing DFlash, Domino or DSpark run changes behaviour. Both shipped LiLiCorr recipes set it `true`, which is the arithmetic their numbers were trained with. Flipping the default is a reasonable follow-up once the autocast above is in. The draft is drawn in fp32 and, under this flag, kept there; an unpromoted run rounds the same draw to the base model's dtype. So the bf16 and fp32 rows of the table above start from the same initialization at the precision each trains in, rather than from two different draws. A unit test pins that. The flag also survives a resume. `modify()` runs under `from_pretrained` with the base model still on meta and cannot place the draft at all, so `restore_draft_precision` re-applies the dtype, the device and the rotary buffer once the weights are loaded and before the Trainer builds the optimizer — the last point that can still decide the Adam moment dtype. It also reloads the draft's tensors at the dtype they were saved in, since checkpoints store the draft in fp32 while the base is bf16 and `dtype="auto"` gives every tensor one dtype. @h-guo18 has the same field in flight on `haoguo/dflash-fp32-master-weights`, plus an HF-format-resume fix this PR does not have. The name is shared deliberately so there is only ever one knob; whichever lands first, the other should be dropped rather than merged. ### Testing - **257 CPU unit tests pass** across `tests/unit/torch/speculative/`, including the existing DFlash, Domino, DSpark and Eagle suites. 48 of them are new and cover LiLiCorr specifically: conversion routing, head geometry, the required-field validation, the three-term objective and its absolute weights, gradient reach into both the head and the drafter body, and the export contract. - Both recipes load and validate through `modelopt.recipe.load_recipe`. - The three DFlash-wide changes are covered behaviourally: the fp32 flag is checked on the optimizer's moment dtypes rather than only on parameters, since the moments are the point of the change, and on the initialization described above; activation checkpointing is asserted to leave draft gradients bit-identical with the flag on and off; and the rotary buffer is asserted present after `modify()` on a real device while still deferred on meta, which is the case the laziness existed for. - The resume path has its own test: after a `save_pretrained` / `from_pretrained` round trip, `restore_draft_precision` is asserted to return the draft to fp32 with its stored weights intact and its Adam moments in fp32. Without it the draft comes back in the base dtype with the flag still set, which is the failure it exists to prevent. - `TestDFlashLazyRotaryEmb` was updated rather than left passing: it asserted the rotary buffer does *not* exist after convert, and the DDP fix deliberately changes that on non-meta devices. The replacement pins the refined invariant in both directions. - The published checkpoints were trained with this arithmetic, verified rather than assumed: a fingerprint over draft initialisation, loss and gradients is compared against the pre-review tree for both `dflash` and `lilicorr`. Loss and gradients are **bitwise identical**. Initialisation moves, by less than bf16 resolution, and that is the single-dtype change described above. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — every addition is opt-in. The new `projector_type` is selected only by config, `dflash_fp32_master_weights` defaults to off, and the activation-checkpointing and DDP fixes preserve behaviour. No existing default changes. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new dependencies. Four files carry `# Adapted from https://github.com/sgl-project/SpecForge/...` headers for the DFlash backbone and loss they derive from (Apache-2.0), matching the attribution already on `hf_dflash.py` in this repo. The two commits described above are @h-guo18's, cherry-picked with authorship and sign-off preserved. - Did you write any new necessary tests?: ✅ — 48 new CPU tests, plus the updated rotary test. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ — will run `/claude review` once opened. ### Additional Information The convolutional recipe is the memory worst case: at an 8B target, combined with fp32 master weights, it may need `training.gradient_checkpointing: true` to fit on 80 GiB, and it fits without at 4B. Checkpointing is mathematically neutral — same objective, same data order, same resulting model — but it trades step time for memory, so a run using it is not step-time-comparable with one that does not. The recipe header says so. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added LiLiCorr speculative decoding with candidate-lattice reranking, configurable objectives, metrics, export support, and optional grouped convolutions. * Added FP32 master-weight support with improved mixed-precision behavior and gradient checkpointing. * Added LiLiCorr training recipes and a Qwen3-8B launcher configuration. * **Bug Fixes** * Improved rotary-embedding configuration handling and corrected DFlash distributed-training hangs. * Added validation for invalid LiLiCorr configurations and improved exported reranking metadata. * **Documentation** * Expanded guidance for FP32 master weights, training workflows, and LiLiCorr configuration. * **Tests** * Expanded coverage across training, evaluation, generation, export, and checkpoint workflows. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: mrusanovsky <mrusanovsky@nvidia.com> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
de3eda8a11 |
Restructure recipes: split per-model_type recipes from model-hub checkpoint recipes (#2219)
### What does this PR do? **Type of change:** Refactor (recipe-library layout) + documentation — backward-breaking for saved `--recipe` paths. Separate the two kinds of built-in Hugging Face recipes that were previously mixed under `modelopt_recipes/huggingface/`: - **`huggingface/<model_type>/`** — architecture recipes keyed by the transformers `model_type`; one recipe covers every checkpoint of that architecture. **Unchanged.** - **`models/<org>/<model_id>/`** — a *new top-level tier* for recipes that mirror one specific published checkpoint, keyed by its **model-hub path** (as on the Hugging Face Hub, ModelScope, etc.) so the on-disk path equals the hub path. Concretely, the model-instance recipes move out of `huggingface/` to the top level: - `huggingface/models/mistralai/…`, `huggingface/models/nvidia/…` → `models/mistralai/…`, `models/nvidia/…` - `huggingface/step3p5/Step3.5-Flash/…` → `models/stepfun-ai/Step-3.5-Flash/…` (re-keyed to the canonical HF repo id [`stepfun-ai/Step-3.5-Flash`](https://huggingface.co/stepfun-ai/Step-3.5-Flash) — org `step3p5`→`stepfun-ai`, id `Step3.5-Flash`→`Step-3.5-Flash`) **Why:** `modelopt_recipes/README.md` already documented a top-level `models/` tier, but the files lived under `huggingface/models/` and instance-specific recipes were awkwardly nested under the per-`model_type` tree. This aligns the filesystem with the documented layout and makes the instance tier hub-addressable — given a checkpoint id you can find (or place) its recipe with no lookup table. `load_recipe` resolves paths directly under `modelopt_recipes/`, so a top-level `models/` sibling of `general/` and `huggingface/` works identically. The move is metadata-only — all recipe YAML content is byte-identical (`R100` renames). Everything else is updating references (nvidia launcher YAMLs, `test_loader.py`) and docs: a new `models/README.md`, plus `huggingface/README.md`, root `README.md`, `ptq.md`, and the `10_recipes.rst` guide, which no longer describe instances under `huggingface/`. ### Usage Recipe paths for the moved checkpoint recipes lose the `huggingface/` prefix (and Step 3.5 Flash is keyed by its hub id): ```python from modelopt.recipe import load_recipe # before load_recipe("huggingface/models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16") load_recipe("huggingface/step3p5/Step3.5-Flash/ptq/nvfp4-mlp-only") # after load_recipe("models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16") load_recipe("models/stepfun-ai/Step-3.5-Flash/ptq/nvfp4-mlp-only") ``` The same rename applies to `--recipe …` CLI values and launcher `QUANT_CFG:` entries. Architecture recipes under `huggingface/<model_type>/` are unaffected. ### Testing - **Recipe resolution (torch-free):** parsed every recipe under `models/` and confirmed all `$import` targets resolve against the recipe root — 0 dangling across the tier. - **Docs consistency:** re-ran the `tests/unit/recipe/test_recipe_docs.py` logic; it now globs both `huggingface/` and `models/`, and every model dir (incl. `Step-3.5-Flash`, `Nemotron-3-Nano-4B-BF16`, …) plus every `general/ptq` recipe is still mentioned in `ptq.md`. - **Reference sweep:** repo-wide grep confirms no remaining references to the old paths outside the intentional historical CHANGELOG entries (released 0.44 / 0.45). - **pre-commit:** `markdownlint-cli2`, license-insert, and `bandit` hooks pass on the changed files. - Note: the full `pytest` suite was not run in my environment (no `torch`), so `test_recipe_docs.py` / `test_loader.py` should be exercised in CI. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ❌ — `--recipe` / `load_recipe` paths for the checkpoint-mirror tier change (drop the `huggingface/` prefix; `step3p5/Step3.5-Flash` → `stepfun-ai/Step-3.5-Flash`). Documented as a Backward Breaking Change in `CHANGELOG.rst` (0.47); the only *released* old paths affected shipped in 0.45. A clean break was chosen over a symlink or loader-alias shim. - If you copied code from any other sources or added a new PIP dependency …: N/A - Did you write any new necessary tests?: ✅ — updated `test_recipe_docs.py` to also glob the top-level `models/` tier so instance recipes stay covered by the doc-consistency check. - Did you update Changelog?: ✅ — added a 0.47 **Backward Breaking Changes** entry. - Did you get Claude approval on this PR?: ❌ <!-- run /claude review --> ### Additional Information Design note: an earlier iteration nested everything under `huggingface/model_type/` + `huggingface/models/`; the final layout keeps `huggingface/` flat (per-`model_type`) and lifts instances to a top-level `models/` tier, matching what `modelopt_recipes/README.md` already documented. The `Step3p5*` architecture class names (from the model's `trust_remote_code` modeling code) are unrelated to the recipe path and are left unchanged. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added checkpoint-specific PTQ recipes for Kimi-K3, Mistral Medium 3.5, and NVIDIA Nemotron models. * Added a Nemotron speculative-decoding warm-start recipe. * **Documentation** * Clarified recipe selection and directory organization. * Documented checkpoint naming conventions and updated usage examples. * **Bug Fixes** * Updated launcher configurations and examples to reference the new recipe locations and corrected model names. * **Tests** * Improved automatic recipe discovery and validation of documented recipe paths. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
d73278808b |
Bump nemo container requirement to 26.08 for MBridge examples (#2257)
### What does this PR do? Type of change: Bug fix Bumps the Megatron-Bridge examples, tests and launcher configs to `nemo:26.08` and removes the version-gated fallbacks they carried, plus the fixes needed to make the suites green on that container. **26.08 bump and shim removal** - Examples, CI workflows, `noxfile.py` and the `mbridge_*` launcher configs move to `nemo:26.08`. - `examples/megatron_bridge/_distillation_provider.py` is deleted — 26.08's Megatron-Bridge ships `convert_to_distillation_provider(..., distill_submodule=...)` natively, so `distill.py` imports it directly. - `prune_minitron.py` drops the `AutoBridge.from_hf_config` / config-only-export probing; `--no_moe_grouped_gemm` is no longer needed in the MoE pruning tests, and the Qwen3.5-MoE `skipif` is gone (native MoE expert mappings are in 26.08). - `_DynamicMambaMixer` targets only the raw `conv1d_weight` / `conv1d_bias` parameters that replaced the `conv1d` module in Megatron-Core. **MambaModel / MambaModelProvider removal** Megatron-Core has shipped `HybridModel` since 26.06 and `MambaModel` is a deprecated subclass that shares its `forward`, so `DMRegistry` resolves those instances to the `HybridModel` registration and the separate entry is redundant. Same for `MambaModelProvider` vs `HybridModelProvider` on the bridge side. `MambaMixer` / `MambaLayer` / `ExtendedRMSNorm` are untouched — the layers still exist. The deprecated `get_te_mamba_stack_spec` is removed; use `get_te_hybrid_stack_spec`. **Bug fix: compressed output_layer extra state** `mtq.compress` converts even a *disabled* `output_layer` into a `RealQuantLinear` (its weight is left uncompressed, since `pack_real_quantize_weight` skips disabled quantizers). The guard added in #2112 excluded `RealQuantLinear` by class, so the layer emitted extra state and every worker died in `GPTModel.sharded_state_dict`: ``` RuntimeError: Boolean value of Tensor with more than one value is ambiguous megatron/core/models/gpt/gpt_model.py:896, in sharded_state_dict output_extra_state and output_extra_state.data ``` The guard now keys off whether the weight was actually compressed (`QTensorWrapper`) instead of the class. This took out all 12 `test_homogeneous_compressed_sharded_state_dict` params, and the crashed workers poisoned the pool, which surfaced as unrelated timeouts and NCCL errors in `test_layer_sync_moe_local_experts_amax`, `test_kv_cache_quant`, `test_kv_cache_amax_sync`, `test_convert_mcore_te_gpt_model` and `test_homogeneous_sharded_state_dict_te_spec` — 21 tests in total. The e2e coverage is `skip_flaky_on_blackwell`, so CI never ran it; `test_output_layer_extra_state_empty_when_nothing_quantized` now asserts the contract directly and is not skipped. **Checkpoint import entry point** 26.08 replaced `examples/conversion/convert_checkpoints.py` with `scripts/conversion/convert.sh`, so `tools/launcher/common/megatron_bridge/import/import.sh` and the three README snippets are retargeted. `import.sh` uses the distributed GPU backend with `GPUS_PER_NODE` / `TP` / `PP` / `EP` knobs. **Megatron-LM on nemo:26.06** keeps working: `_get_mamba_conv1d` still dispatches between the `conv1d` module (26.06 and earlier) and the raw parameters (26.08+), so `import_mcore_gpt_from_hf` / `export_mcore_gpt_to_hf` handle NemotronH on both. Only the Megatron-Bridge examples and Minitron pruning of Mamba/hybrid models require 26.08. **Test consolidation** `test_export_distilled_megatron_to_hf.py` is merged into `test_distill.py`: `test_distill_llm` becomes `test_distill_llm_hf_export` and covers the standalone `--export_iterations all` run on the checkpoints it already produces, saving one full distillation (~185 s of CI time). The two mamba-named gpu test files are renamed to `hybrid`. ### Usage ```bash # HF -> Megatron import, via Megatron-Bridge's 26.08 conversion entry point bash /opt/Megatron-Bridge/scripts/conversion/convert.sh import \ --executor local \ --device gpu \ --gpus-per-node 8 \ --hf-model Qwen/Qwen3-8B \ --megatron-path /tmp/Qwen3-8B-megatron ``` ### Testing All on `nvcr.io/nvidia/nemo:26.08`, 2x RTX 6000 Ada, no timeout overrides: - `tests/examples/megatron_bridge`: 16 passed, 1 skipped (28m14s). The skip is the `gemma3vl` QAD param, now `@pytest.mark.manual` since `qwen3_5_moe_vl` covers the VLM QAD path. - `tests/gpu_megatron` (`_extensions`, `distill`, `export`, `opt`, `peft`, `sparsity`, `speculative`, `utils`): 61 passed, 5 xpassed. - `tests/gpu_megatron/torch/export` re-run after the conv1d dispatch change: 27 passed. - The 21 previously failing/hanging quantization tests: 21 passed (12 + 9). - `tests/gpu_megatron/torch/{nas,prune}`: verified separately. `import.sh` equivalence on a toy `qwen3_moe`, comparing all 12 weight tensors after flattening each dist checkpoint with `dcp_to_torch_save` — the GPU backend at 1 GPU, `--tp 2`, `--pp 2`, `--ep 2`, and `import.sh` end-to-end (`GPUS_PER_NODE=2 EP=2`) are all byte-identical to `--device cpu`. `nemo:26.06` compatibility was checked directly in that image: `megatron.core.models.hybrid.HybridModel`, the modelopt hybrid spec and `hybrid_layer_pattern` are all present, while `megatron.bridge.models.hybrid` and the bridge's `distill_submodule` are not. The NemotronH round-trip test failed there before the conv1d dispatch was restored and the dispatch is back in place; per project convention the suites themselves only run on 26.08. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ⚠️ Megatron-Bridge examples plus Minitron pruning of Mamba/hybrid models now require `nemo:26.08`. Megatron-LM quantization and checkpoint export still run on `nemo:26.06`. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — `test_output_layer_extra_state_empty_when_nothing_quantized` for the compress fix; existing tests extended for the merged export coverage. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added guidance for importing Hugging Face checkpoints into Megatron distributed format. * Expanded distillation workflows to export selected or all checkpoint iterations. * **Improvements** * Expanded Hybrid model support across Megatron workflows. * Updated distributed import tooling with GPU and parallelism options. * Updated supported environments and examples to NVIDIA NeMo 26.08. * **Bug Fixes** * Corrected output-layer quantization state handling when quantization is disabled. * **Documentation** * Added compatibility guidance for current and legacy NeMo containers. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5db2682519 |
[Example]: Calibration-free FP8/NVFP4 PTQ for speculative-decoding drafters (#2027)
### What does this PR do? Type of change: new example Adds `examples/speculative_decoding/scripts/quantize_drafter.py`, a CLI that quantizes an exported speculative-decoding drafter to FP8 or NVFP4 — weight-only or weight+activation — with no calibration data. It needs no modeling code either. Exported drafters such as [`nvidia/MiniMax-M3-DSpark`](https://huggingface.co/nvidia/MiniMax-M3-DSpark) have no importable model class, so each 2-D weight is wrapped in a throwaway `nn.Linear` under its checkpoint key and ModelOpt's usual `quantizer_name` patterns select over those names. Works for any drafter layout (DSpark / DFlash / EAGLE3 / Medusa). **Formats:** `w4a16_nvfp4`, `nvfp4`, `fp8`, `fp8_pc_pt` — the ModelOpt formats vLLM's backend can actually serve. AWQ is deliberately not offered, since `awq_lite` silently degrades to plain RTN without a `forward_loop`. **Static activation scales without calibration.** `fp8` and `nvfp4` normally need an activation amax *measured* on calibration data; a fixed `input_scale` of 1.0 is applied instead. That works because acceptance length is governed almost entirely by **clipping**, not resolution: Sweeping the fixed scale over three decades (same setup as the Testing section below; bf16 baseline 3.1423): | `input_scale` | amax | FP8 AL | vs bf16 | NVFP4 AL | vs bf16 | |---|---|---|---|---|---| | 0.003 | 1.3 | 2.2204 | -29.34% | 2.2076 | -29.75% | | 0.01 | 4.5 | 2.6719 | -14.97% | 2.6641 | -15.22% | | 0.03 | 13.4 | 2.9751 | -5.32% | 2.9259 | -6.89% | | 0.1 | 44.8 | 3.1013 | -1.31% | 3.0206 | -3.88% | | 0.2 | 89.6 | 3.1178 | -0.78% | 3.0015 | -4.48% | | 0.3 | 134.4 | 3.1370 | -0.17% | 3.0222 | -3.82% | | 0.5 | 224.0 | 3.1268 | -0.50% | 3.0360 | -3.38% | | **1.0 (default)** | **448.0** | **3.1457** | **+0.11%** | **3.0193** | **-3.91%** | | 2.0 | 896.0 | 3.1354 | -0.22% | 3.0172 | -3.98% | | 4.0 | 1792.0 | 3.1245 | -0.57% | 3.0034 | -4.42% | Both formats fall off a cliff below ~0.03, where the declared range sits far under the activations' true magnitude and most of the tensor is clipped. Both then sit on a flat plateau from ~0.3 to 4.0 **with no drop-off at the top**, so the scale only has to be big enough. 1.0 is the middle of that plateau, which is why it is hardcoded rather than exposed. NVFP4 trails FP8 by a roughly constant 3.5% across the plateau — that gap is the 4-bit resolution cost, and no choice of scale recovers it. Deriving the amax from the weights instead was tried and does not work: `max|W|` averages 0.79 while a RMSNorm'd activation is O(1) with outlier channels in the tens, so the range lands 1–2 orders of magnitude low and clips, measuring -31% to -46% AL. **Where calibration would go.** All of this sits behind `resolve_activation_scales()`, the single place deciding where a static amax comes from. Real calibration slots in ahead of the fixed fallback with no change to the CLI or the call site, and composes because `set_static_activation_amax()` skips quantizers that already have an amax: ```python if calib_forward_loop is not None: mtq.calibrate(root, quant_cfg["algorithm"], forward_loop=calib_forward_loop) set_static_activation_amax(root) # fills in what calibration did not reach ``` **Serving a quantized drafter.** Four things had to be written into the exported checkpoint before vLLM would load one: - emit `quant_method` (`modelopt_fp4` / `modelopt`) — vLLM reads that key, ModelOpt writes only `quant_algo` - emit the exclusion list under `ignore` too — that is the key read from the flat `quantization_config`; `exclude_modules` alone yields an empty exclusion set - add `*<name>` wildcards so exclusions match a runtime's nested module prefix (`model.fc`) rather than the checkpoint key (`fc`) - add `*qkv_proj` / `*gate_up_proj` aliases for layers a runtime fuses, whose names appear in no checkpoint key Nothing is then needed on the caller side. **This closes the open question left in the previous revision of this PR: vLLM does read `quantization_config` off the draft checkpoint.** `ModelConfig._verify_quantization` fills `quantization` in from `quant_method` when it is unset, so once the export declares that key — the first fix above — detection works on its own. Verified on Nemotron-3.5-Lightning passing nothing: `Detected ModelOpt NVFP4 checkpoint (quant_algo=NVFP4)` → `FlashInferCuteDslNvFp4LinearKernel`, AL 4.278 against 4.203 measured earlier. `specdec_bench` also gains a `DSPARK` algorithm, which it did not have: an exported `Qwen3DSparkModel` would otherwise have to go through `DFLASH` and be built with vLLM `method="dflash"`. The branch sets `method="dspark"` and leaves `draft_sample_method` on vLLM's own default of `greedy`. A target whose fused-collective workspace (sized at CUDA-graph capture) overflows at large speculative batches can disable graphs with `--runtime_params '{"engine_args": {"enforce_eager": true}}'`. For DFlash-family drafters, `qwen3_dflash.py` builds its fused context-KV projection by reading `qkv_proj.weight` raw and calling `F.linear`, which cannot consume a packed weight. Keep those layers in bf16 with `--exclude '*q_proj*' '*k_proj*' '*v_proj*' '*qkv_proj*'`; `o_proj` and the MLP — the bulk of the drafter — still quantize. That exclusion is mandatory, not a tuning choice. `fc` (the projection from the target's captured layers into the draft) is the one real knob, and it is a genuine trade rather than a free win — see the Testing section for both models' numbers. The examples quantize it; add `'*fc*'` to the exclude list to keep it in bf16. `embed_tokens`, `markov_head` and `confidence_head` are excluded by default: they are 2-D so the flat view treats them as GEMMs, but they are embeddings or a single-output projection. `lm_head` is excluded by the preset itself — unlike on a base model it is 37% of this drafter's parameters, so `--quantize_lm_head` is a real lever (~1.9 GiB), but measure AL first. The flag re-enables both of `lm_head`'s quantizers; re-enabling only the weight one would ship a W+A checkpoint whose `lm_head` has no `input_scale` while the config still advertises it as quantized. ### Usage ```bash # weight+activation FP8, calibration-free, lossless on both models measured below python scripts/quantize_drafter.py \ --drafter_path deepseek-ai/dspark_qwen3_8b_block7 \ --qformat fp8 \ --export_path ./dspark-qwen3-8b-fp8 \ --exclude '*q_proj*' '*k_proj*' '*v_proj*' '*qkv_proj*' # smallest: weight-only NVFP4 python scripts/quantize_drafter.py \ --drafter_path nvidia/MiniMax-M3-DSpark \ --qformat w4a16_nvfp4 \ --export_path ./MiniMax-M3-DSpark-W4A16 ``` Or end to end on Slurm — quantize, then measure AL — via the launcher examples added here, one per target: ```bash uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_dspark_ptq_nvfp4.yaml --yes uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/hf_dspark_ptq_nvfp4.yaml --yes ``` Serving one, if you are not going through `specdec_bench`: ```python speculative_config = { "method": "dspark", "model": "./dspark-qwen3-8b-fp8", # quantization is read from its config.json "num_speculative_tokens": 7, } ``` ### Testing Two targets with different architectures, so the conclusions are not one model's quirk: * **Qwen3-8B** (dense transformer) + [`deepseek-ai/dspark_qwen3_8b_block7`](https://huggingface.co/deepseek-ai/dspark_qwen3_8b_block7), `block_size` 7, TP1. * **Nemotron-3.5-Lightning-30B-A3B** (hybrid Mamba-MoE) + [`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark`](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark), `block_size` 8, TP8, with the mamba engine settings the model card pins (`mamba_backend=flashinfer`, `mamba_ssm_cache_dtype=float16`, stochastic SSM-cache rounding). Both: MT-Bench 80 questions, greedy, one vLLM instance per point. | recipe | activations | Qwen3-8B AL | vs bf16 | Nemotron-3.5 AL | vs bf16 | |---|---|---|---|---|---| | bf16 baseline | — | 3.1423 | — | 4.3296 | — | | **`fp8`** | static, `input_scale` 1.0 | **3.1457** | **+0.11%** | **4.3289** | **-0.02%** | | `fp8_pc_pt` | dynamic per-token | 3.1228 | -0.62% | 4.3411 | +0.26% | | `w4a16_nvfp4`, `fc` in bf16 | bf16 (weight-only) | 3.0392 | -3.28% | 4.2899 | -0.92% | | `w4a16_nvfp4`, `fc` quantized | bf16 (weight-only) | 3.0186 | -3.94% | 4.2334 | -2.22% | | **`nvfp4`** | static, `input_scale` 1.0 | **3.0193** | **-3.91%** | **4.2030** | **-2.92%** | **FP8 weight+activation at the fixed `input_scale` of 1.0 is lossless on both.** +0.11% and -0.02% are both inside run-to-run noise — the Nemotron baseline was measured twice under identical settings and the two runs differ by 0.94% (4.3093 / 4.3499), which sets the resolution of that column. On the same reading, `fp8` and `fp8_pc_pt` are indistinguishable on Nemotron; the dynamic variant only pulls ahead on Qwen3. NVFP4 costs 3-4% on Qwen3 and 2-3% on Nemotron, i.e. the 4-bit weight resolution is the real price and it is model-dependent but bounded. Whether to quantize `fc` is a per-model call rather than a general recommendation — it buys a few percent of size for an AL cost that differs by ~2x between these two drafters: | `fc` bf16 → quantized | Qwen3-8B | Nemotron-3.5 | |---|---|---| | checkpoint size | 3.293 → 3.181 GiB (-3.4%) | 1.316 → 1.258 GiB (-4.4%) | | AL | 3.0392 → 3.0186 (-0.68%) | 4.2899 → 4.2334 (-1.32%) | `fc` itself is only 3.5% (Qwen3) / 4.5% (Nemotron) of drafter parameters; `embed_tokens` is the bulk (26% / 36%) and is excluded by default. The Qwen3 `w4a16_nvfp4` rows were measured in a later session than the rest of that column; the `fc`-in-bf16 run reproduced the original number to four decimals (3.0392), so the column is internally comparable. Also validated on `nvidia/MiniMax-M3-DSpark`: `w4a16_nvfp4` runs in 67 s on CPU, 9.98 GiB (fp32) -> 3.51 GiB; all 43 quantized tensors round-trip within 0.0952 relative error; the 29 untouched tensors are bit-identical to `bf16(source)`. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (example-only) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ❌ — validated manually as above. Can add a `tests/examples/speculative_decoding/` test over a small synthetic drafter if wanted before merge. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (example-only) - Did you get Claude approval on this PR?: ❌ (not yet run) ### Additional Information The measurements above are one drafter on one target with one benchmark; the plateau's location and the ~3.5% NVFP4 gap should be re-measured before assuming they carry to a different drafter. Note when reading an exported checkpoint: `input_scale` is `amax/448` for FP8 but `amax/(6*448)` for NVFP4, so the one fixed amax records as 1.0 in an FP8 checkpoint and 0.1667 in an NVFP4 one. Both mean the same activation range. Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
2b296b2f62 |
Support fine-tuning released DFlash/DSpark drafters (causal SWA, attention sink, warm start) (#2149)
# Support fine-tuning released DFlash/DSpark drafters (causal SWA, attention sink, warm start) ### What does this PR do? Type of change: New feature + bug fix Adds what ModelOpt was missing to fine-tune an already-published DFlash/DSpark draft model. The concrete target is [`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark`](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark) on its hybrid Mamba/attention/MoE base, but every change is generic. Before this PR that checkpoint could not be trained faithfully — or even loaded: its attention-sink tensors were dropped as unexpected keys, its block-causal attention had no implementation, and its capture layers were silently overwritten with ModelOpt's defaults. **New user-facing options** (all default to today's behavior, so existing runs are unchanged): | Option | Values | Purpose | | --- | --- | --- | | `dflash_draft_attention` | `bidirectional` (default) / `causal` | Block-internal attention pattern. `causal` restricts a query at block position `i` to draft positions `<= i`. | | `dflash_attention_sink` | `false` (default) / `true` | Learnable per-head `attention_sink_bias [num_heads]` on every draft layer — one extra logit appended before the softmax and dropped after, so a head can put probability mass nowhere instead of being forced to attend inside its window (the GPT-OSS formulation). | | `dflash_init_checkpoint` | path | Warm-start the draft from an exported checkpoint instead of a random init. Any missing/unexpected/wrong-shaped tensor raises rather than warns. | | `dflash_architecture_config.target_layer_ids` | list | Which base layers feed the draft's `fc`. Previously recomputed unconditionally with no override. | **Bugs fixed along the way** (each one silently corrupts training rather than failing): - The exporter hard-coded `dflash_config.causal: False` and only wrote it under SWA, so even a correctly-trained causal draft would be served non-causally. It now reflects the trained setting, and emits `attention_sink_bias` when enabled. - `_build_generate_swa_mask` returned `None` whenever `swa_window_size` was unset, which would have dropped the causal structure at generation time while training used it. - `target_layer_ids` was recomputed from the uniform default on every convert. The released drafter uses `[1,5,19,29,41,51]`; the default for a 52-layer base is `[1,11,20,30,39,49]` — *different layers*. Here it surfaced as a matmul shape error only because the plane counts disagreed; with a matching count it would have trained on the wrong features silently. - The streaming dataset assumed the draft's aux layers all sit below the base's final layer (`aux = planes[:-1]`, `target = planes[-1]`). A draft whose top aux id *is* the final layer cannot get an extra plane — vLLM captures each layer once — so `final_aux_is_base_hidden` now lets the last plane serve both roles. It is derived from the model, not configured by hand. - DSpark head weights load from either the flat layout ModelOpt exports (upstream DeepSpec convention) or the nested `markov_head.` layout the NVIDIA release uses. Without the remap the two `[131072, 512]` Markov tables — ~14% of the draft's parameters — stay randomly initialized while everything else warm-starts, with no error. - `nemotron_h` is enabled in `_FINAL_NORM_TYPE_BY_MODEL_TYPE`: despite the hybrid stack, `NemotronHModel.norm_f` is a plain RMSNorm, and without the entry the offline/streaming fake base raises instead of reconstructing the distillation target. ### Usage ```yaml dflash: dflash_init_checkpoint: /path/to/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark dflash_draft_attention: causal dflash_attention_sink: true dflash_swa_window_size: 1024 dflash_block_size: 8 dflash_mask_token_id: 990 dflash_architecture_config: target_layer_ids: [1, 5, 19, 29, 41, 51] ``` A full worked example is at `modelopt_recipes/general/speculative_decoding/dspark_nemotron35_warmstart.yaml`. ### Testing **Unit tests** — 124 pass (`test_hf_dflash.py`, `test_hf_dspark.py`, `test_hf_domino.py`, `test_hf_dflash_offline.py`, `test_modeling_final_norm.py`), 32 of them new: causal mask structure (lower-triangular per block, no cross-block leakage, context visibility unchanged), the sink math (degenerates to plain attention at `-inf`, absorbs mass monotonically, receives gradient), warm-start load/reject paths, Markov key remapping, and explicit `target_layer_ids`. **Checkpoint compatibility** — the released drafter loads with zero missing/unexpected keys and zero shape mismatches; all 77 tensors (6 attention sinks and both Markov tables included) match bit-exactly, and a training step runs with gradients reaching the sink and Markov parameters. **End-to-end streaming training** — Nemotron-3.5 base served by vLLM (1 node, TP8) feeding 8 trainer GPUs over NIXL; the draft warm-starts from the released checkpoint and trains with `causal` + sink + SWA 1024. 128 Daring-Anteater conversations, 20 epochs (the plot shows the first 5, where the trend is clearest — the curves flatten after that):  Over the first 5 epochs loss falls **1.85 → 1.36** and train accuracy rises **0.25 → 0.49**; across the full 20 epochs they reach **1.21** and **0.48** (peak 0.54) before flattening. This validates the pipeline end-to-end — capture layers, plane split, mask direction, sink loading and warm-start weights all have to be right for this curve to appear. It is *not* a model-quality result: 128 samples over 20 epochs overfits by construction, and the corpus is not generated by the base model, so the absolute numbers are not meaningful. ### TODO (follow-up) **A complete, robust checkpoint/config converter.** Both conversions are handled ad hoc here: - *Draft config → training config.* The recipe transcribes ~15 fields by hand from the drafter's `config.json`. Only the shape-bearing ones (`num_hidden_layers`, `num_attention_heads`, `intermediate_size`, `markov_rank`) fail loudly when mistyped; the rest — `mask_token_id`, `causal`, `swa_window_size`, `block_size` — train "successfully" on a wrong value and only surface later as a mysteriously low acceptance length. A converter should derive the whole block from the checkpoint, including its aliases (`pard_token`, `dspark_markov_rank`, `dflash_query_causal`, top-level `sliding_window` / `attention_sink_bias`) and duplicated fields. - *Weight layout.* The `markov_head.` remap is a load-time hook. A converter should normalize layouts explicitly, and decide whether export should also emit the release's aliases so a round-trip reproduces the original format (today it renames `architectures` to `DFlashDraftModel`). - *Base config.* Serving this base on vLLM needs its `config.json` layer-type vocabulary updated for the transformers-5 path (`mamba` → `linear_attention`, `attention` → `full_attention`, plus a matching `hybrid_override_pattern`). That is done by hand today and is not covered by this PR. ### Before your PR is "*Ready for review*" - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added configurable causal or bidirectional attention for DFlash models. * Added optional attention sinks, checkpoint warm starts, and explicit target-layer selection. * Improved streaming data handling for shared auxiliary and base hidden states. * Added Nemotron-3.5 Lightning DSpark warm-start training and serving recipes. * **Bug Fixes** * Preserved configured attention behavior during model export. * Prevented warm-start checkpoints from being reapplied during restoration. * Improved checkpoint compatibility, validation, and attention-mask handling. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
58ad6edc5f |
Fix pruned-HF export fallback + add Nemotron-3.5-Lightning launcher examples (#2196)
### What does this PR do? Type of change: Bug fix + new example Two related changes for the Megatron-Bridge Minitron prune/quantize launcher flows: 1. **Fix pruned-HF export crash on containers that reject config-only save.** `#2159` added a config-only HF export path gated only on `hasattr(AutoBridge, "from_hf_config")`. Some Megatron-Bridge versions (e.g. `nemo:26.04`) expose `from_hf_config` but reject a config-only `save_hf_pretrained` (`ValueError: save_hf_pretrained requires a pretrained HuggingFace model`), so `prune_minitron.py` crashed instead of using the intended dummy-model fallback. Now it attempts the config-only save and falls back to the dummy-model path on `ValueError`. 2. **Add Nemotron-3.5-Lightning-30B-A3B launcher examples** (`mbridge_prune.yaml`, `mbridge_quantize.yaml`) on `nemo:26.08`. Prune targets 3B active with an MMLU gate; quantize runs W4A16 NVFP4 4/6 PTQ via the `w4a16_nvfp4_4o6` recipe with `tp_size=1` (static-block NVFP4 MSE is unsupported with TP>1). ### Usage ```shell uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_prune.yaml --yes uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_quantize.yaml --yes ``` ### Testing Verified end-to-end on OCI-HSG: - **Nano prune (`nemo:26.04`)** — exercises the fallback path: config-only save raised the `ValueError`, the fallback caught it and exported via the dummy-model path. `mmlu_10pct_bs32 = 0.5196` (gate 0.50) PASS; vLLM gen PASS. - **Lightning prune (`nemo:26.08`)** — config-only export path: `score = 0.6000` (gate 0.58) PASS, 3.00B active params; vLLM gen PASS. - **Lightning quantize (`nemo:26.08`)** — recipe PTQ + unified-HF export; MMLU `0.7741` (gate 0.75) PASS. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A <!-- launcher example configs + fallback path exercised by CI prune/quantize jobs --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- fix is for a bug introduced in the same unreleased cycle (#2159); rest are example configs --> - Did you get Claude approval on this PR?: ❌ <!-- pending /claude review --> ### Additional Information The fallback fix addresses the `mbridge_prune` launcher CI failure introduced by #2159. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added a pruning workflow for Nemotron-3.5-Lightning-30B-A3B with calibration, quality scoring, checkpoint export, and multi-GPU generation. - Added a four-GPU NVFP4 W4A16 quantization workflow with Hugging Face conversion and MMLU evaluation. - **Bug Fixes** - Improved hybrid model export by falling back to dummy-model export for supported configuration-only export failures. - Added clearer logging and handling for supported export failures while preserving unrelated errors for investigation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
5e887aa7c9 |
fix(launcher): increase Nemotron 3.5 Lightning QAD parallelism (#2191)
### What does this PR do? Type of change: Bug fix Current Nemotron 3.5 Lightning QAD example can OOM, and also TP=2 is not supported for static quantization. Increase GPUs, CP and decrease TP=1 ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Very short summary of changes only for new features, backward breaking changes, deprecations, or fixes for critical bugs present in previous releases. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Configuration** * Updated the QAD training setup to run across four nodes with revised tensor and context parallelism settings. * Documented the 32K sequence-length constraint and removed outdated run information. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> |
||
|
|
71b3d884ee |
example(launcher): Megatron-Bridge NVFP4 QAD launcher example for Nemotron-3.5-Lightning-30B-A3B (#2142)
## What does this PR do? Adds `mbridge_qad.yaml`, a launcher example running NVFP4 quantization-aware distillation for **Nemotron-3.5-Lightning-30B-A3B** through the Megatron-Bridge scripts in `examples/megatron_bridge/`, alongside the existing `mbridge_prune.yaml` / `mbridge_quantize.yaml`. `megatron_lm_qad.yaml` (#2146) runs the same recipe and the same data through Megatron-LM. This is the Megatron-Bridge counterpart: the same `huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6` recipe and the same `nvidia/Nemotron-Post-Training-Dataset-v2` chat data, with the training hyperparameters from our public-data QAD run. Four tasks: tokenize the training data, PTQ the student, distill it against the frozen BF16 teacher, export to unified HF. ### Why the extra tokenize task Megatron-LM's finetune path reads an HF parquet shard directly. Megatron-Bridge trains from pre-tokenized data, so `distill.py` consumes Megatron `.bin`/`.idx` via `--data_paths`. The chat split is therefore tokenized once with `modelopt.torch.utils.plugins.megatron_preprocess_data`. `--hf_streaming` avoids the Arrow cast errors this dataset's nested tool-call fields trigger in non-streaming mode, and `--append_eod` is omitted because chat rows already terminate each conversation via the chat template. ### Details - Training topology 8 nodes x 4 GPUs, TP=1 PP=1 CP=4 EP=16 -> DP=8; `gbs` 64 at `mbs` 1 is 8 gradient-accumulation microbatches. 200 iters x 64 x 32768 = 419M training tokens. - PTQ runs TP=EP=PP=1 across 4 ranks (pure DP), so each rank calibrates on its own shard. `--calib_dataset_name` is left unset, selecting the default public `cnn_nemotron_v2_mix` (cnn_dailymail + Nemotron-Post-Training-Dataset-v2). - Export uses TP=1 (the HF writer does not gather TP shards) and PP=4, splitting 52 layers 13/stage. - Pins `nvcr.io/nvidia/nemo:26.06` like the other `mbridge_*` examples. ## Dependencies Based on `main`; the PTQ recipe ships in #2146 (merged). No other PR required. Nemotron-3.5-Lightning has `tie_word_embeddings: false`, so a correct quantized `lm_head` in the exported checkpoint also depends on #2112. ## Testing The PTQ -> export -> QAD flow and these hyperparameters were run end to end on Nemotron-3.5-Lightning (`main` + #2112 + #2113): - PTQ completed, 6660 quantizers, MTP heads retained (`mtp_num_layers: 1`) with all 278 `mtp.*` quantizers disabled by the recipe. - Export produced a unified-HF checkpoint (18487 keys, including 270 MTP tensors). - QAD trained with 900 quantizers through a validation pass at iteration 50. The YAML itself is validated against the launcher's conventions (`ntasks_per_node == gpus_per_node` on Slurm, single-line `inline`, no `args` alongside `inline`, all `<<global_vars.X>>` resolve, output prefix matches `megatron_preprocess_data`'s naming) and by the repo's `validate launcher YAML references` pre-commit hook. Topology arithmetic checked: EP divides world/(TP*PP), `gbs` divisible by DP*mbs. ## Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (new example file only) - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an example workflow for NVFP4 quantization-aware distillation of NVIDIA Nemotron 3.5 Lightning 30B-A3B. * Supports dataset tokenization, post-training quantization, teacher-student distillation, and export of a unified Hugging Face checkpoint. * Includes configurable model, dataset, and checkpoint paths, distributed execution settings, and support for local or Slurm-based workflows. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: James Shen <yueshen@nvidia.com> |
||
|
|
4ca7dd8871 |
Add Nemotron Lightning 3.5 NVFP4 recipe and QAD example (#2146)
### What does this PR do? Type of change: New example Add Nemotron Lightning 3.5 NVFP4 recipe and QAD example Also exclude MTP in default disabled quantizers ### Usage ```python # uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/megatron_lm_qad.yaml --yes ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, backward breaking changes, deprecations, or fixes for critical bugs present in previous releases. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added an end-to-end NVFP4 quantization, distillation, and export workflow for NVIDIA Nemotron 3.5 Lightning 30B-A3B. - Added a PTQ configuration supporting NVFP4 W4A16 quantization with optimized scaling and FP8 support for selected components. - **Bug Fixes** - Preserved custom model output locations when provided, while retaining the existing default path. - **Configuration** - Disabled quantization for MTP modules by default. - Updated an existing Nemotron workflow to use the aggressive NVFP4 quantization profile. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> |
||
|
|
f2bfe63183 |
Nemotron Nano 3 QAD Launcher Example on OSS Nemotron-Post-Training-V2 data (#2134)
### What does this PR do? Type of change: New example Add a Nemotron Nano 3 QAD Launcher Example on OSS Nemotron-Post-Training-V2 data. It performs 4 steps 1. Teacher conversion: Convert the HuggingFace BF16 checkpoint to a Megatron-Core BF16 checkpoint 2. PTQ: quantize the Megatron-Core checkpoint to `MAMBA_MOE_NVFP4_AGGRESSIVE_CFG` quant config 3. QAD (Quantization Aware Distillation): distill the BF16 checkpoint to the PTQ checkpoint on a subset of the Nemotron-Post-Training-V2 `chat` data. To train on a different subset or load the entire dataset, you may modify `--finetune-data-split` and `--finetune-data-files` flags. 4. Export: export the QAD checkpoint to HuggingFace format so it is ready for local inference All steps use the TE (Transformer Engine) spec, which with the new TEGroupedMLP per-expert quantizer is approximately 10-15% faster than the previous local ModelOpt spec (which used SequentialMLP) on Hybrid-MoE models. ### Usage ``` # Usage from tools/launcher: source .env-slurm uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/megatron_lm_qad.yaml --yes ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, backward breaking changes, deprecations, or fixes for critical bugs present in previous releases. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added a launcher configuration for NVFP4 quantization-aware distillation of the Nemotron 3 Nano 30B-A3B model. - Added support for selecting training or fine-tuning workflows through `MLM_TRAIN_SCRIPT`. - Improved forwarding of additional training arguments. - **Updates** - Updated the Megatron-LM launcher component to a newer revision. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> |
||
|
|
7afbfbc85c |
Offline-KD QAD example (#1998)
### What does this PR do? Type of change: new feature Adds a Nano-v3 launcher example which uses the new MLM offline-logits KD feature ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ - Did you get Claude approval on this PR?: ❌ ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a new offline, five-stage Slurm workflow example for Nemotron-3-Nano 30B NVFP4 A3B, using cached teacher logits for knowledge distillation. * Runs quantization, frozen teacher logits generation, student fine-tuning from saved logits, checkpoint export to Hugging Face format, and TensorRT-LLM evaluation. * Includes ready-to-run configuration for repeatable execution with tuned parallelism and export settings. * **Chores** * Updated local Python pre-commit hooks to execute via `uv` for consistent dev-time tooling. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com> Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
93b9e4b176 |
Add Megatron-Bridge prune & quantize launcher pipelines (#2031)
### What does this PR do? Type of change: new example + small launcher / modelopt-example features (backward compatible) Adds end-to-end ModelOpt **launcher** pipelines for the Megatron-Bridge flow on Nemotron-3-Nano-30B-A3B, the minimal launcher features to run them wrapper-free from YAML, and an **in-step accuracy gate** for Minitron pruning. **New launcher examples** (`tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/`): - `mbridge_prune.yaml` — Minitron prune **with an in-step MMLU gate** → vLLM sanity gen (2 tasks) - `mbridge_quantize.yaml` — FP8 quantize → unified-HF export → MMLU gate on the vLLM backend, which doubles as the deploy sanity check (3 tasks). Matches the [tutorial](https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md). **Prune accuracy gate** (reuses the search's own score — no separate eval step): - `modelopt/torch/prune/plugins/mcore_minitron.py` — `MCoreMinitronSearcher` now stores the exported best `CandidateSubnet` under `state_dict["best"]` (additive; sits beside the existing `sorted_layers` key). - `examples/megatron_bridge/prune_minitron.py` — new `--score_lower_bound`: reads `pruning_scores["best"].score` and exits non-zero if the exported model is below the floor. Score-agnostic (any `--prune_score_func`); rejected with `--prune_export_config` (manual pruning has no score). **Launcher (`tools/launcher`)** — run single-node Megatron-Bridge one-liners directly from YAML: - `SandboxTask.inline` — a command in the YAML, no `common/**/*.sh` wrapper (single-line; folded scalar) - `SandboxTask.reqs` / `reqs_file` — pip-install deps in the container before the command (shell-safe; on Slurm the install is rank-0-guarded so multi-rank tasks don't race) - `SlurmConfig.docker_user` — local-Docker user (e.g. `root`); ignored on Slurm - `get_default_env` honors `HF_HOME` / `TRITON_CACHE_DIR` env overrides, so a non-CI user can point caches at a writable path (the shared `/cicd/hf-cache` is owned by the CI account) - reject `args` together with `inline` **`examples/llm_eval/lm_eval_hf.py`**: - `--accuracy_lower_bound` — gate on the single requested task's `acc` (used by the quantize MMLU step; exits non-zero if below) - drop ModelOpt (hf-only) args for non-`hf` backends, so `--model vllm` works on a deployable quantized checkpoint ### Usage ```bash cd tools/launcher # Prune (in-step MMLU gate) -> vLLM gen uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/mbridge_prune.yaml --yes # FP8 quantize -> unified-HF export -> MMLU gate (vLLM) uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/mbridge_quantize.yaml --yes ``` ### Testing - Launcher unit tests for `inline`, `reqs`/`reqs_file`, `docker_user`, the `args`+`inline` guard, and example-resolve; ruff / mypy / bandit clean. `tests/examples/megatron_bridge/test_prune_minitron.py` now passes `--score_lower_bound=0.01` to exercise the gate path on the tiny models. - **End-to-end on the real Nemotron-3-Nano-30B-A3B (4×B200, OCI-HSG):** - Prune 30B → 3B-active: `[score_gate] mmlu_10pct = 0.5196 >= 0.45 PASS`; vLLM gen coherent. - FP8 quantize → unified-HF export (`Detected ModelOpt fp8 checkpoint`) → MMLU on the vLLM backend `acc = 0.7077 >= 0.60 PASS`. - Earlier smoke on **Qwen3-0.6B** in `nemo:26.06` through the same flow. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (new state-dict key is additive; new CLI args default to off) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- tooling/examples + additive searcher state key --> - Did you get Claude approval on this PR?: ❌ <!-- pending --> ### Additional Information - **Container pinning:** saving a pruned Nemotron-H to HF requires `transformers<5`, so `mbridge_prune.yaml` runs on `nemo:26.04` (26.06 drops it); quantize/export run on `nemo:26.06`. - **`docker_user: root`** is set on all example tasks — local-Docker only (ignored on Slurm), needed so downstream tasks can read task_0's root-owned checkpoints and to read the image's root-only `/opt/Megatron-Bridge`. - The quantize MMLU step passes `enforce_eager=True` to vLLM — for a run-once eval this skips ~17 min of CUDA-graph capture / `torch.compile` with no accuracy change. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added inline shell commands and per-task Python dependency installation for launcher workflows. - Added configurable Docker user selection and preservation of existing cache environment settings. - Added evaluation accuracy and pruning score gates that fail workflows below configured thresholds. - Added NVIDIA Nemotron pruning and quantization workflow examples. - Improved backend-specific handling of ModelOpt options. - **Documentation** - Documented inline commands, dependencies, variable substitution, and configuration examples. - **Bug Fixes** - Strengthened task execution validation and configuration checks. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c062b3f829 |
Add sidecar GPU/CPU memory+utilization monitor for HF PTQ (#2000)
### What does this PR do?
**Type of change:** New feature (developer tool) + new tests
Adds a **standalone, cross-process resource monitor** —
`tools/resource_monitor.py` —
that samples a workload's GPU and CPU usage *from the outside* while it
runs, so
we can verify per-run memory budgets (e.g. the single-GPU layerwise PTQ
target of
≤80 GB GPU / ≤80 GB CPU for OMNIML-4947) without instrumenting
`hf_ptq.py` itself.
The monitor wraps any command, samples at a fixed interval, and on exit
writes a
CSV timeseries plus a peak/mean/min summary:
- **GPU** (per device, via NVML with an `nvidia-smi` fallback): used
memory,
utilization %, power draw (W), temperature (°C).
- **CPU** (via `psutil`): system total/used/free memory + utilization %,
and the
monitored process tree's RSS + CPU %.
It is opt-in from `examples/hf_ptq/scripts/huggingface_example.sh` via
`MODELOPT_MEM_MONITOR=1` (off by default → byte-for-byte identical
behavior). When
enabled it wraps the `hf_ptq.py` run and writes the trace/summary to a
**sibling**
`${SAVE_PATH}_mem_monitor/` directory, kept out of the exported
checkpoint that is
uploaded and consumed downstream.
**Files:**
- `tools/resource_monitor.py` — the sidecar (NVML + `nvidia-smi`
fallback; `psutil`).
- `tests/unit/tools/test_resource_monitor.py` — CPU-only unit tests (run
in the `unit` nox lane).
- `examples/hf_ptq/scripts/huggingface_example.sh` — opt-in
`MODELOPT_MEM_MONITOR=1` wrapper.
- `examples/hf_ptq/requirements.txt` — adds `psutil`.
- `pyproject.toml` — adds `psutil` to the `dev-test` extra
(deterministic import in the unit lane).
- `.github/workflows/unit_tests.yml` — adds `tools/resource_monitor.py`
to the unit-test path filters.
#### Why a new tool instead of extending
`modelopt/torch/utils/memory_monitor.py`?
The existing `GPUMemoryMonitor` is a fundamentally different tool and
cannot serve
this use case by extension:
| | `modelopt.torch.utils.memory_monitor.GPUMemoryMonitor` |
`tools/resource_monitor.py` (this PR) |
|---|---|---|
| Scope | **In-process** thread inside the workload | **Cross-process**
— wraps an external command |
| Survives workload OOM/SIGKILL | ❌ dies with the process | ✅ keeps
sampling, still writes the summary |
| Import cost | Pulls `torch` (~19 s) — lives in the workload |
Torch-free (`psutil`+`pynvml`, ~0.03 s) |
| Metrics | GPU device memory only | GPU mem/util/**power/temp** +
**CPU** mem/util + process-tree RSS |
| Output | In-memory / logs | CSV timeseries + peak/mean/min summary |
Merging the two would force `torch` into a standalone sidecar (defeating
the point)
or split the in-process monitor's threading model. A future refactor may
factor out
a **shared torch-free sampling core with two thin frontends**
(in-process + sidecar);
that is tracked as a follow-up rather than blocking this monitoring
harness, which
PR #2008 (single-GPU disk-offload PTQ) depends on.
### Usage
```bash
# Wrap mode (preferred): monitor exits with the workload's return code
python tools/resource_monitor.py --gpus 2,3 --out mem.csv --summary peak.txt -- \
python hf_ptq.py --pyt_ckpt_path=<model> --qformat=nvfp4 ...
# Opt-in from the HF PTQ example (off by default):
MODELOPT_MEM_MONITOR=1 CUDA_VISIBLE_DEVICES=2,3 CUDA_DEVICE_ORDER=PCI_BUS_ID \
bash examples/hf_ptq/scripts/huggingface_example.sh <args>
# -> writes ${SAVE_PATH}_mem_monitor/mem_trace.csv and mem_peak.txt
```
### Testing
- **Unit (CPU-only, in the `unit` nox lane):** `pytest
tests/unit/tools/test_resource_monitor.py`
— 11 tests covering `--gpus` parsing (CSV + space-separated, UUID/MIG
rejection),
the disabled/`nvidia-smi` sampling paths (including `[N/A]` → `None` and
the
smi-failure-yields-empty guard), CPU sampling, the accumulator, and
end-to-end
CSV/summary + exit-code propagation in wrap mode.
- **GPU-validated** on a B200 node (GPUs 2,3): confirmed the `gpu{i}_*`
memory /
utilization / power / temperature columns populate and the summary is
written.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — new tool; the example wrapper
is off unless `MODELOPT_MEM_MONITOR=1`.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — `psutil`
added to `dev-test` + the hf_ptq example requirements;
`pynvml`/`nvidia-ml-py` is an optional runtime dep (graceful
`nvidia-smi` fallback).
- Did you write any new necessary tests?: ✅ —
`tests/unit/tools/test_resource_monitor.py`.
- Did you update Changelog?: N/A — repo-level `tools/` script, not
shipped in the wheel.
- Did you get Claude approval on this PR?: ❌ <!-- run /claude review -->
### Additional Information
Part of **OMNIML-4947** (single-GPU disk-offload PTQ). This is PR 1 of
the stack —
the monitoring harness that PR #2008 (disk-offload layerwise PTQ +
offload-aware
export) builds on.
---------
Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
a3ac4759dd |
tools/mcp: pin mcp<2 to fix unit CI collection (#2026)
### What does this PR do? **Type of change:** Bug fix (CI) `mcp==2.0.0` was released and removed the `mcp.server.fastmcp` module (the 2.0 SDK replaces it with `mcp.server.mcpserver`). `tools/mcp/pyproject.toml` declared an unpinned `mcp>=1.0`, so CI now resolves `mcp==2.0.0`, and `tools/mcp/modelopt_mcp/server.py`'s `from mcp.server.fastmcp import FastMCP` fails at import: ``` ModuleNotFoundError: No module named 'mcp.server.fastmcp' ERROR collecting tools/mcp/tests/test_bridge.py ``` This breaks the `mcp` unit job — and thus the `unit-pr-required-check` gate — on **every** PR whose diff touches `pyproject.toml`, `noxfile.py`, or `.github/workflows/unit_tests.yml` (the changed-files paths that trigger the `mcp` job). This PR pins `mcp>=1.0,<2`, keeping the 1.x line that still ships `mcp.server.fastmcp`. Migrating the server to the mcp 2.0 API (`mcp.server.mcpserver`) is a larger change tracked separately. ### Usage ```python # N/A - dependency pin only ``` ### Testing - `uv pip install -e tools/mcp` now resolves an `mcp` 1.x wheel, so `from mcp.server.fastmcp import FastMCP` imports and `tools/mcp/tests/test_bridge.py` collects again. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — tightens an existing dependency's upper bound. - Did you write any new necessary tests?: N/A — restores collection of the existing `tools/mcp` tests. - Did you update Changelog?: N/A — CI/build fix, nothing user-facing in the wheel. - Did you get Claude approval on this PR?: ❌ ### Additional Information Unblocks the `unit-pr-required-check` gate for in-flight PRs (surfaced on #2000). Follow-up: migrate `modelopt_mcp/server.py` to the mcp 2.0 API and relax the pin. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Prevented compatibility issues by limiting the MCP package to supported version 1.x releases. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6105fe84e9 |
[Examples]: MiniMax-M3 DSpark (#1965)
### What does this PR do?
Type of change: new example + bug fixes
Adds a **MiniMax-M3 DSpark streaming training recipe** under
`tools/launcher/examples/MiniMaxAI/MiniMax-M3/`, plus the four fixes it
needs to actually run. Each fix addresses a failure mode that is silent
or misleading without it:
1. **Gemma-style final norm for the fake base**
(`modeling_final_norm.py`, `modeling_fakebase.py`): M3 uses a
gemma-style final RMSNorm (`(1 + weight)` scale, fp32
multiply-then-cast). Selecting the norm by `model_type` alone picks
plain `rmsnorm` — MiniMax's VL remote code coerces its `model_type`-less
`text_config` to **mixtral** — which silently drops the `+1` and
corrupts the distillation target. New `_FinalGemmaRMSNorm` + selection
by the explicit `use_gemma_norm` config flag (only MiniMax sets it).
2. **Loud failure for un-maskable `answer_only_loss`**
(`hf_streaming_dataset.py`): with a fast tokenizer whose chat template
has no `{% generation %}` tags,
`apply_chat_template(return_assistant_tokens_mask=True)` only warns and
returns an **all-zero mask** — training runs at zero loss on every
sample with no other symptom. Now raises at tokenization with an
actionable message.
3. **`SERVE_BLOCK_SIZE` knob** (`train_eagle_streaming.sh`): nemo_run
exports env values unquoted, so a multi-token
`SERVE_EXTRA_ARGS="--trust-remote-code --block-size 128"` loses
everything after the first token. M3's MSA sparse attention requires KV
block 128 (`ValueError: No common block size for 16` at engine init
otherwise), so `--block-size` gets a dedicated single-token knob.
4. **Relax the speculative_decoding `transformers` pin to `<5.13`**
(match `pyproject.toml`): the old `<5.4` pin downgrades recent vLLM
containers (e.g. transformers 5.12.1, which also provides in-tree
`minimax_m3_vl`) and breaks `vllm serve` (`ALLOWED_LAYER_TYPES` needs
>=5.5.3).
The example itself encodes the validated M3 specifics: generation-tagged
chat template copy (required for `answer_only_loss` — see fix 2), draft
dims + base `rope_theta=5e6` set explicitly (not inherited), mask token
200063 (reserved slot; added tokens end at 200060), `EAGLE_CAPTURE_IDS`
= draft default `target_layer_ids+1` + final layer, `trust_remote_code`
at serve/export, and AWS-EFA NIXL notes (UCX segfaults at agent init on
EFA nodes; LIBFABRIC required there).
### Usage
```bash
cd tools/launcher
export SLURM_HOST=localhost SLURM_ACCOUNT=<account> SLURM_PARTITION=<partition> \
SLURM_HF_LOCAL=<hf_models_dir> SLURM_JOB_DIR=<experiments_dir> NEMORUN_HOME=$PWD
uv run launch.py --yaml examples/MiniMaxAI/MiniMax-M3/hf_streaming_dspark_multi_node.yaml \
identity=$HOME/.ssh/id_ecdsa detach=True --yes
```
### Testing
- `tests/unit/torch/speculative/plugins/`: **129 passed** inside a
current vLLM x86_64 nightly container (transformers 5.12.1), including 3
new tests for the assistant-mask guard.
- Generation-tagged template verified against real corpus samples:
`input_ids` identical to the original template, mask covers exactly the
assistant turns (think prefix + content + eos), contiguous, no
user-prompt leak.
- The recipe is exercised end-to-end by a live M3 DSpark training run
(this yaml modulo cluster paths): streaming serve + NIXL transport +
resume all healthy; drafter MT-Bench AL exceeds our Kimi-K2.6 DSpark
reference by ~8k steps.
- `ruff check` / `ruff format` (0.15.20) clean.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (norm selection only changes
models with `use_gemma_norm=True`, previously mis-normed; the guard
turns a silent zero-loss run into an error; `SERVE_BLOCK_SIZE` is
opt-in; the pin relax widens the allowed range)
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ (can add if desired)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added Gemma-style RMS normalization support for speculative decoding.
- Added MiniMax-M3 multi-node training configuration, including a full
Jinja chat template for tool calls, multimodal content, and thinking
modes.
- Added optional `SERVE_BLOCK_SIZE` support for vLLM serve launches.
- **Bug Fixes**
- Improved `answer_only_loss` masking validation: now fails fast with
clear errors when required `{% generation %}` markers are missing or
when using a non-fast tokenizer.
- **Compatibility**
- Expanded the supported Transformers version range for speculative
decoding examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
01c708e792 |
Add HybridModel MBridge support for nemo:26.08 (#2005)
### Description MBridge main (nemo:26.08) will initialize Nemotron-H as HybridModel instead of MambaModel (subclassed of HybridModel). Also make minimum nemo container 26.04 ### Testing Tested Nemotron-3-Nano PTQ with MBridge main (fails otherwise) Tested locally `tests/gpu_megatron` and `tests/examples/megatron_bridge` with `nemo:26.06.01` + Mount latest MBridge/Mcore GH CICD tests will be added with nemo:26.08 release <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added TE HybridModel stack-spec support and enabled Megatron-Bridge hybrid providers/models for export/import and runtime handling. * **Updates** * Dataset packing now oversamples raw text at **16x** and improves the packed-mode underflow warning. * Quantization: `--quant_cfg` now defaults to `None` unless explicitly set (or via `--recipe`). * Distillation example: validation settings are provided via a dedicated top-level validation configuration. * Improved plugin import warnings to report the originating call location; model stats now support HybridModel. * **Deprecations** * Megatron-Bridge / Megatron-LM optimization features now require NeMo container `nemo:26.04` or newer (`nemo:26.06` recommended). * The Mamba stack specification helper is deprecated. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
f10d5183ec |
launcher: bump TRT-LLM to 1.3.0rc20, pin vLLM to v0.22.0, fix max_seq… (#1982)
bump TRT-LLM to 1.3.0rc20, pin vLLM to v0.22.0, fix max_seq_len for Qwen3.5-4B ### What does this PR do? Type of change: Bug fix, new feature - Upgrade TRT-LLM container from 1.3.0rc10 to 1.3.0rc20 across all Qwen3-8B, Qwen3-30B-A3B, Kimi-K2.5, and gpt-oss-20b launcher configs. - Replace Kimi-K2.5 aarch64-specific vLLM image (v0.22.0-aarch64) with the multi-arch v0.22.0 tag, which resolves to amd64/arm64 automatically. - Fix Qwen3.5-4B throughput_32k runs: raise max_seq_len from 40960 to 65536 to accommodate outlier prompts (~46.6K tokens) that caused VLLMValidationError and aborted the entire benchmark run. - Fix specdec_bench_mtp_vllm.yaml: remove reference to non-existent runtime_params_throughput_32k.yaml; use --max_seq_len 65536 instead ### Usage ``` cd Model-Optimizer/tools/launcher uv run launch.py --yaml examples/Qwen/Qwen3-8B/megatron_lm_ptq_local.yaml hf_local=/mnt/hf-local --yes ``` ### Testing N/A - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?:N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Improvements** * Updated launcher example pipelines and the default launcher Slurm container to the newer TensorRT-LLM `1.3.0rc20` image. * Updated Kimi-K2.5 workflows to use the multi-architecture vLLM `0.22.0` image (removing prior architecture-specific variants). * Increased the Qwen3.5-4B long-context benchmark maximum sequence length to 65,536 tokens and simplified the corresponding throughput configuration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com> |
||
|
|
92296aeac2 |
fix(mmlu): expert-DP batch sharding for MoE + Nano-30B-A3B PTQ example (#1967)
Combines the NemotronH MoE PTQ launcher example with the MMLU expert-parallel sharding fix it exercises. The example is the end-to-end test for the fix — quantize + **MMLU EP=4** + export + vLLM smoke on Nano-30B-A3B — so they ship together. ## Fix — `megatron_mmlu.py` shards over the expert-DP group for MoE models `megatron_mmlu` shards whole batches across the **dense data-parallel group** and runs `megatron_prefill` per-rank on a disjoint subset. Correct for dense models, but for MoE with **EP>1** the prefill forward runs an **expert all-to-all across the EP group**. When EP overlaps the dense-DP group (e.g. `EP=4,TP=1,PP=1` on 4 GPUs), each rank is on a different batch, so the all-to-alls desync (uneven batch counts + differing padded seq-lengths) → trailing ranks block at NCCL communicator creation until the c10d store times out (600 s). Reproduced on Nano-30B-A3B at `EP=4`: ranks 2 & 3 wait on rank 0's `ncclUniqueId`. The fix shards over the **expert-data-parallel** group when `EP>1`, so EP peers evaluate every batch in lockstep and only true expert-DP replicas take disjoint batches. Dense models (`EP==1`) are byte-for-byte unchanged. ## Example — `examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/megatron_lm_ptq.yaml` Quantize (EP=4) + MMLU gate (EP=4) → export (PP=4) → vLLM smoke, modeled on the Super-120B example. Uses `NVFP4_DEFAULT_CFG` (no Nano-30B recipe) and sets `modelopt_install_path` to the nemo-container venv so the mounted modelopt overrides the container copy. Passes `test_examples_resolve.py`. ## Test plan - [x] MoE MMLU (`EP>1`) completes instead of hanging — validated end-to-end via this example on Nano-30B-A3B `EP=4` (nmm-sandbox CI). - [x] `test_examples_resolve.py` green (example parses/resolves). - [ ] Dense MMLU sharding unchanged (`EP==1` path identical). Supersedes #1964 (example) and #1966 (fix). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Improved MMLU evaluation for expert-parallel MoE models by aligning batch sharding with expert-parallel execution for more consistent results. * **Bug Fixes** * Dense (non-MoE) models retain standard data-parallel batch sharding, while MoE handling is corrected. * **Documentation** * Updated the Nemotron-3-Nano BF16 PTQ example to clarify the quantize → MMLU gate → export flow and why the vLLM smoke is intentionally omitted. * Adjusted the Llama-3.2-1B-Instruct PTQ/MMLU gate lower-bound threshold (0.40 → 0.36). * **Tests** * Marked the homogeneous compressed sharded state-dict test to skip on Blackwell to avoid a known flaky issue. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d69d5aab8b |
[Examples]: Kimi-K2.6/K2.7-Code Dflash/Dspark (#1934)
### What does this PR do? **Type of change:** New feature (recipe + examples) Adds the DSpark training recipe and Kimi launcher examples for streaming speculative-decoding training. Split out of #1849 to keep that PR focused on the DSpark head implementation. **Files added (5, all additive — no code changes):** 1. `modelopt_recipes/general/speculative_decoding/dspark.yaml` — the DSpark training recipe. Selects the DSpark head via `dflash_architecture_config.projector_type=dspark` (DFlash backbone + lightweight sequential/Markov head + optional confidence head) and sets the three-term loss weights (`dflash_ce_loss_alpha` / `dflash_l1_loss_alpha` / `dflash_confidence_head_alpha`). 2–5. `tools/launcher/examples/moonshotai/{Kimi-K2.6,Kimi-K2.7-Code}/hf_streaming_{dflash,dspark}_multi_node.yaml` — four multi-node streaming launcher examples mirroring the existing Kimi-K2.5 streaming format (`common/eagle3/train_eagle_streaming.sh`). They carry the Kimi-specific base path, draft dims, capture ids, and mask token; the DSpark examples build on `dspark.yaml` and override only the Kimi-specific fields (draft dims, `dflash_block_size=8`, mask token). K2.7-Code shares the K2.6 architecture (`kimi_k25`, 61 layers), so the draft config is identical. ### Dependency **Stacked on #1849.** The recipe uses the config fields (`dflash_ce_loss_alpha`, `dflash_l1_loss_alpha`, `dflash_confidence_head_alpha`) added in #1849, so this PR is based on that branch and must merge **after** it. GitHub will auto-retarget the base to `main` once #1849 merges. ### Testing - `dspark.yaml` passes the `validate modelopt recipes` schema check (against the #1849 config schema). - The launcher examples ship a placeholder `container: <vllm-image-with-aux-capture-fix>` and are meant to be adapted per cluster (image, account, partition), not run verbatim in CI. - Backward compatible?: ✅ (additive recipe + example files only) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a DSpark speculative-decoding training recipe with configurable draft, optional per-position confidence, and Markov head components. * Added new multi-node streaming training example workflows for Kimi-K2.6 using DSpark/DFlash, including dataset preparation plus coordinated serve/train settings. * Configured answer-only loss options, masking behavior, training hyperparameters, checkpoint/log cadence, and capture-id wiring with sensible default runtime timeouts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
adeca905d4 |
example(megatron_lm): Llama-3.2-1B-Instruct Megatron QAD launcher example (#1931)
### What does this PR do? Type of change: new example (+ a launcher unit test) Adds a public **Megatron QAD** (quantization-aware distillation) launcher example for `meta-llama/Llama-3.2-1B-Instruct`: PTQ (NVFP4) + MMLU gate → QAD/SFT train, plus a `common/megatron_lm/train/sft.sh` wrapper (mirrors `common/megatron_lm/quantize/quantize.sh`). **Reworked — this no longer changes any modelopt code.** The earlier `_extra_state` plugin patch was reverted: it was the wrong load path (the crash is the *main* model load in `megatron/training/checkpointing.py`, not `_load_extra_state_from_sharded_checkpoint`) and it regressed quantizer-state loading (Nemotron `quantize_resume` MMLU 0.70 → 0.31). The correct fix is a config flag on the QAD train step: - **`--dist-ckpt-strictness log_all`** → the main model load drops only the genuinely-missing `decoder.final_layernorm._extra_state` (via Megatron's own unexpected-key set), instead of the torch_dist DCP planner hard-raising. Present quantizer `_extra_state` is untouched. Container `nvcr.io/nvidia/nemo:26.06.00` (ships `nvidia-resiliency-ext>=0.6.0`). ### Testing - **Integration (GPU/cluster):** validated end-to-end on nmm-sandbox CI (CLAB-SC-01, computelab): 2/2 QAD tasks passed — quantize+MMLU gate and QAD train (the train step loads the quantized ckpt past the former `Missing key … final_layernorm._extra_state` crash). The Nemotron `quantize_resume` regression is cleared (MMLU back ≥0.70). - **Unit (CPU):** adds `tools/launcher/tests/test_examples_resolve.py` — parses every launcher example and checks each task's `script`/`_factory_`/`args` shape and that `common/*` scripts exist (guards config regressions). 53 pass. (It also surfaced two pre-existing broken example script refs — `common/smoke/hostname.sh`, `common/eagle3/offline_training.sh` — excluded here, worth a separate fix.) ### Before your PR is "*Ready for review*" - Backward compatible?: ✅ N/A (new example + new wrapper + new test; no code paths changed) - New tests?: ✅ CPU example-resolution test; GPU validation is the nmm-sandbox integration run - Changelog?: N/A (example) ### Additional Information Follow-ups: OMNIML-5421 (make the Megatron-LM example scripts take CLI args so these `common/*` wrappers can be dropped for a direct 1-shell call). Supersedes the reverted modelopt-patch approach. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a launch pipeline configuration to run Quantization-Aware Distillation (QAD) post-training SFT for the Llama 3.2 1B Instruct model. * Added a dedicated launcher script to simplify SFT/QAD runs with sensible defaults and CI-friendly status handling. * **Tests** * Added CPU-only checks that validate all launcher example YAML files are well-formed and reference existing scripts (where applicable), preventing common runtime misconfigurations. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0185726d62 |
fix: make MCP Slurm wait poll remote job state (#1900)
## Summary - run managed Model-Optimizer checkouts via `uv run --with-editable` so the `modelopt-launcher` console script is available - persist Slurm submission metadata beside the local experiment and make `job_status` / `wait_for_experiment` prefer remote `sacct` / `squeue` state over local `_DONE` - add the missing `common/smoke/hostname.sh` used by `examples/smoke/hostname.yaml` ## Why Detached Slurm submission creates the local `_DONE` marker when the submit process exits, not when the remote Slurm job completes. This made MCP `wait_for_experiment` return `done` while Slurm jobs were still running. Reproduced on both MFA and non-MFA clusters. ## Validation - `uv run --project tools/mcp ruff check tools/mcp/modelopt_mcp/bridge.py tools/mcp/tests/test_bridge.py` - `uv run --project tools/mcp pytest tools/mcp/tests/test_bridge.py` - `uv run --project tools/launcher pytest tools/launcher/tests/test_core_extended.py tools/launcher/tests/test_slurm_config.py` - `uv run --with-editable tools/launcher modelopt-launcher --yaml tools/launcher/examples/smoke/hostname.yaml --dryrun --yes` ## Manual smoke evidence - ptyche smoke completed: experiment `cicd_1783037181`, Slurm job `2320415`, exit `0:0` - computelab reproduced the wait bug pre-fix: MCP reported done immediately while Slurm job `2945347` was still running; job later completed with exit `0:0` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an automated hostname smoke test to validate launcher environment readiness. * Enhanced Slurm-backed job status reporting by persisting scheduler submission metadata and exposing Slurm state in responses. * **Bug Fixes** * Job status now prioritizes authoritative remote scheduler state (with graceful fallback to local completion/failure markers when unavailable). * Improved terminal-state handling so experiment polling won’t stop early when a local done marker exists but the scheduler is still running. * **Tests** * Updated and expanded Slurm status and waiting behavior tests, including validation of persisted scheduler metadata. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
bc5bc1ac5f |
[Feat]: Add Final Norm for vLLM Hidden Extractor (#1846)
### What does this PR do? **Type of change:** Bug fix vLLM captures the final-layer hidden state *before* the model's final norm, but the offline/streaming distillation path fed it straight into `lm_head`, so the reconstructed base logits (the KD target) were computed from un-normed hidden states. This PR re-applies the base model's final norm before `lm_head` when the producer declares a pre-norm capture (`base_hidden_prenorm`), for both DFlash and EAGLE: - Producer sets `base_hidden_prenorm` (streaming: `True`; offline: from the dump). - Consumer (`_maybe_apply_base_final_norm`) re-applies the base final norm, and **fails loud** if pre-norm is declared but the model's norm type isn't supported (no silent corruption). - `FakeBaseModel` now loads the base final norm (+ `rope_theta`/`rms_norm_eps`); norm type is gated by an explicit `model_type` allowlist (gpt_oss excluded pending a matching norm class). ### Testing `tests/unit/torch/speculative/plugins/test_modeling_final_norm.py`; DFlash/EAGLE streaming training verified end-to-end. - Backward compatible?: ✅ (post-norm captures declare `base_hidden_prenorm=False` → unchanged) - New tests: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Expanded speculative decoding/distillation to reconstruct missing base-model logits by optionally applying a base model’s final pre–LM-head normalization when pre-norm hidden states are provided. * Streaming dataset now emits `base_hidden_prenorm` and can configure RDMA backends from environment settings. * **Bug Fixes** * Rejects mixed `base_hidden_prenorm` values within a batch. * Fails fast on streaming token-length mismatches. * Streaming dataset loading supports directory inputs by expanding sorted JSONL shards (and errors if none are found). * **Tests** * Added/updated unit and dataset tests for final-norm behavior and the new batch field. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
4b9225bee2 |
launcher: fix host=None when _factory_ is dropped by nemo_run --yaml path (#1842)
## Summary - **Root cause:** When `launch.py --yaml <config>` processes a launcher-format YAML, nemo_run constructs `SlurmConfig` directly from the YAML dict for inline tasks (`task_0`, `task_1`, ...). It strips the unrecognised `_factory_` key and only sets known `SlurmConfig` fields. Since `host` is not in the YAML (it should come from the factory), `SlurmConfig.host` is left as `None`, crashing paramiko at `Connection(host=None)`: ``` TypeError: expected str, bytes or os.PathLike object, not NoneType ``` - **Fix:** In `SandboxPipeline.__post_init__`, after collecting inline tasks, detect `host` is falsy (symptom of dropped `_factory_`) and apply the registered `slurm_factory` as base defaults, overlaying YAML-specified fields (identified by `value != SlurmConfig dataclass default`). - **Scope:** Only triggers when `host` is `None`/`""` and `slurm_factory` is registered. No-op otherwise. Does not affect `task=@` / `pipeline=@` paths. ## Architectural note Three factory resolution paths exist today: 1. `task=@` / `pipeline=@` — nemo_run's `@run.cli.factory` registry ✓ 2. `task_configs` list — `_FACTORY_REGISTRY` via `create_task_from_yaml()` ✓ 3. `--yaml` inline tasks — nemo_run's type system, **bypasses all factory registries** ✗ (fixed here) ## Test plan - [ ] 65/65 unit tests pass (`uv run python3 -m pytest tests/ -v`) - [ ] All pre-commit hooks pass (ruff, mypy, bandit) - [ ] Manually verified: `SandboxPipeline` with `slurm_config.host=None` (simulating nemo_run dropping `_factory_`) now gets `host` filled in from the registered `slurm_factory` 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for reusing an existing SSH connection for Slurm-based launches, reducing repeated login prompts during submissions. * Added a new minimal smoke-test example for hostname checks on CPU-only clusters. * **Bug Fixes** * Improved handling of optional GPU settings so GPU allocation can be left unset when not needed. * Fixed Slurm configuration loading so launch settings are preserved correctly when using YAML-based workflows. * **Tests** * Added coverage for SSH reconnect behavior, Slurm launch overrides, and optional GPU configuration handling. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
9038b71f09 |
Autoquant and GPTQ in support in Megatron-Core [OMNIML-3151] (#1562)
### What does this PR do? Type of change: New Feature Autoquant and GPTQ in support in Megatron-Core - Add EP support to AutoQuantize - Register MCore support in AutoQuantize - Add decoder `output_layer` (lm head) to layerwise hook so that GPTQ can register all decoder layers & lm head - Split dataloader helper function out of megatron calibration utils so that AutoQuantize in Megatron-LM can reuse the same dataloader ### Usage See https://github.com/NVIDIA/Megatron-LM/pull/4821 for Autoquant usage in Megatron ```python # For GPTQ pick a recipe that uses gptq algorithm and run mtq.quantize # e.g. general/ptq/nvfp4_default-kv_none-gptq ``` ### Testing Tested AutoQuant on Nemotron Nano and Ultra. Tested GPTQ on Nano 3. Added unit tests for both AutoQuant and GPTQ ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Summary by CodeRabbit * **New Features** * Added lazy Megatron-Core AutoQuant integration with Megatron-specific quantization hooks and better decoder-layer discovery for layerwise calibration. * Improved AutoQuantize for expert-parallel (EP) models, including consistent per-layer recipe selection across DP/TP/EP. * Extended quant-layer grouping for NemotronH MCore fused “local_experts” linear layers. * **Bug Fixes** * Prevented division-by-zero when calibration inputs are empty during Hessian updates. * **Tests** * Added/extended unit and GPU coverage for EP AutoQuant, decoder-layer calibration discovery behavior, and zero-token Hessian no-op. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> Signed-off-by: Jenny Chen <jennifchen@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
248cbf2fd2 |
OMNIML-5128 Capture Docker experiment id (#1840)
## Summary - Capture nemo-run experiment_id for Docker submit_job by redirecting detached launcher output to a side-channel log and tailing it briefly. - Reuse the launcher output parser for both Docker and Slurm submit paths. - Document Docker PID + experiment_id behavior and ignore the MCP local uv.lock. ## Jira OMNIML-5128 ## Validation - uv run --project tools/mcp pytest tools/mcp/tests -q <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved Docker job submission to capture `experiment_id` from launcher output during a brief tail window, while always returning the detached `pid` (and `stdout_log` diagnostics if capture times out). * Made Slurm identifier extraction more robust across varying launcher output formats. * **Documentation** * Updated `submit_job` docs/module descriptions to clarify Docker PID/`experiment_id` and `stdout_log` behavior on timeout. * **Tests** * Added/expanded tests for Docker success, timeout/no-id scenarios, parsing edge cases, and log-side-channel failure handling. * **Chores** * Updated local tooling ignore rules for `.venv/` and `uv.lock`. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
f335459dc0 |
refactor(examples): rename llm_ptq → hf_ptq (symlink for back-compat) (#1759)
## What does this PR do? **Type of change:** refactor / deprecation (examples) Follow-up to #1705 (which consolidated `examples/vlm_ptq` into `examples/llm_ptq`). Since that example now covers Hugging Face **LLM and VLM** PTQ, the `llm_ptq` name is a misnomer. This renames the directory to `examples/hf_ptq` and leaves a relative symlink `examples/llm_ptq → hf_ptq` so existing paths/commands keep working during a deprecation window. Requested by @kevalmorabia97 on #1705 (with the symlink-for-back-compat approach), targeted for the **same 0.46 release** as the consolidation. ### Changes - `git mv examples/llm_ptq → examples/hf_ptq` and `tests/examples/llm_ptq → tests/examples/hf_ptq` (the CI runner maps the matrix name to both `examples/<name>` and `tests/examples/<name>`). - Add a tracked back-compat symlink `examples/llm_ptq → hf_ptq`. - Update CI matrices and all repo **path references** (docs, READMEs, agent skills, launcher/debugger tools, tests) from `llm_ptq` to `hf_ptq`. - Keep Python identifiers / test-util module names (`run_llm_ptq_command`, `llm_ptq_utils`) — they name the LLM-PTQ task, not the directory. - Preserve the CODEOWNERS team slug (`modelopt-examples-llm_ptq-codeowners`) and historical CHANGELOG entries; add a CHANGELOG deprecation note. ### Back-compat caveats (inherent to git directory symlinks) - ✅ Linux/macOS CLI usage and Python `cwd`/pytest resolution work through the symlink. - ⚠️ Windows git checkouts don't materialize symlinks by default (low impact — this example is Linux-only in practice). - ⚠️ GitHub web doesn't follow directory symlinks, so legacy external deep-links to `examples/llm_ptq/...` won't navigate in. All **internal** references are repointed to `hf_ptq`, so the symlink is only for legacy external/CLI use. ### Usage (unchanged via symlink) ```bash # New canonical path cd examples/hf_ptq scripts/huggingface_example.sh --model <hf_model> --quant fp8 # Old path still works (forwards via symlink) cd examples/llm_ptq && scripts/huggingface_example.sh --model <hf_model> --quant fp8 ``` ### Testing - `bash -n` on moved/edited shell scripts (new path + via symlink). - `py_compile` on moved/edited Python; test re-export shim repointed to `examples/hf_ptq/example_utils`. - Verified git tracks `examples/llm_ptq` as a single symlink (mode 120000), not a duplicated tree (no pre-commit / pytest double-processing). - `pre-commit run` on all changed files passes. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (relative symlink keeps `examples/llm_ptq` paths valid; see caveats above) - Did you write any new necessary tests?: N/A (pure rename; existing tests moved with the dir) - Did you update Changelog?: ✅ ### Additional Information Follow-up (later release): remove the `examples/llm_ptq` symlink once external references have migrated. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * PTQ guidance now directs to the unified Hugging Face PTQ flow, including VLM quantization via the shared `--vlm` entry point. * **Documentation** * Updated README and guide links, references, and command snippets to use `hf_ptq` (replacing `llm_ptq`). * Deprecated and consolidated `vlm_ptq` into `hf_ptq`; removed VILA/NVILA coverage from the Hugging Face PTQ examples. * **Bug Fixes** * Improved detection and routing so local/manual setup uses the correct PTQ source. * **Tests / Chores** * CI and example tests updated to run the `hf_ptq` variants. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c248dd5434 |
[Feat]: Domino support (#1710)
### What does this PR do? Type of change: New feature Adds **Domino** speculative decoding: the parallel DFlash draft backbone plus a lightweight **GRU causal correction head**. The backbone produces *base* logits for a full draft block in one forward; a GRU over the block's teacher-forced tokens produces a causal state that is fused with the backbone hidden state and projected to a vocab-sized logit correction on the block suffix — injecting the intra-block causal dependency the parallel backbone lacks. Trained with a dual loss `(1-λ)*final + λ*base`, where `λ_base` decays linearly 1→0 (curriculum: learn the parallel backbone first, then the correction). Reuses the DFlash mode/config/recipe; selected via `dflash_architecture_config.projector_type=domino` and routed to its own registry so `HFDominoModel` does not shadow `HFDFlashModel`. Exports in the z-lab/SpecForge drafter format (`prefix_gru.*` / `embed_proj.*`). > Note: the inference side (vLLM / AR evaluation) is intentionally **not** wired up yet — the correction head is not applied in serving. To be added once the inference path lands. ### Usage ```bash # Online training (recipe: projector_type=domino) uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_online_domino.yaml --yes ``` ### Testing CPU unit tests in `tests/unit/torch/speculative/plugins/test_hf_domino.py` cover conversion routing, the training forward (dual loss + grads), the λ schedule, and the export format. Online Qwen3-8B training validated end-to-end (loss curve below). <img width="1803" height="809" alt="image" src="https://github.com/user-attachments/assets/7c9d2001-bd80-4dec-919b-443e61089cca" /> ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (opt-in via `projector_type=domino`; DFlash path unchanged) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new dependency) - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Reference: SpecForge PR #571 (z-lab); drafter format `huggingface.co/Huang2020/Qwen3-8B-Domino-b16`. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added Domino speculative-decoding training with a decaying base/final dual-loss curriculum and Domino-specific lambda scheduling (training-only; inference wiring not yet included). * Added Domino draft-head export support for training checkpoints. * **Documentation & Configuration** * Added a Domino speculative-decoding training recipe and an HF Online Domino launcher configuration for Qwen3-8B. * **Refactor** * Updated speculative model conversion/export to route to Domino variants based on the configured projector type. * **Tests** * Added unit tests for Domino conversion, training loss/metrics, lambda decay behavior, and exporter output layout. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
33bfa8b1fe |
CI/Dev env bump (#1818)
### What does this PR do? Type of change: chore Bumps CI/dev tooling and test containers. **Container bumps** - NeMo test containers → 26.06 - TRT-LLM container → 1.3.0rc19 - transformers max version → 5.12 **Dev tooling bumps** - ruff bump 0.12.11 → 0.15.18 - mypy 1.17.1 → 2.1.0: enable new defaults (`local_partial_types`, `strict_bytes`); fix/narrow the errors newly surfaced by mypy 2.0 in 4 modules (rather than blanket-suppressing them); remove 2 stale `# type: ignore` comments - pre-commit 4.3.0 → 4.6.0 - sphinx 8.1 → 9.1 + sphinx-rtd-theme 3.0 → 3.1: add `suppress_warnings = ["ref.python"]` to fix cross-reference ambiguity error new in sphinx 9.x - trl fix for newly released 1.7 version **Bug fixes surfaced by the bumps** - sparsity (weight): make the weight mask DTensor-aware under FSDP. The transformers→5.12 bump routes the HF Trainer FSDP optimizer-state save through torch's DTensor-based `get_optimizer_state_dict`, which triggered `aten.mul.Tensor got mixed torch.Tensor and DTensor` in the dynamic `weight` getter. The mask is now distributed to the weight's mesh/placements before masking, cached, and rebuilt only when the sharding changes (invalidated on `set_mask`). Fixes the `llm_sparsity` example test. ### Testing - `pre-commit run --all-files` ✅ (including mypy 2.1.0) - `nox -s docs` ✅ - `tests/unit/torch/sparsity` + `tests/unit/torch/nas` ✅ - `llm_sparsity` GPU example test (FSDP path) verified in CI ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Summary * **Documentation** * Refreshed Docker pre-requisites across examples to recommend updated container image tags (and streamlined some instructions). * **Bug Fixes** * Improved sparse weight mask handling for DTensor/FSDP by aligning and caching distributed masks. * Made TensorRT engine byte retrieval return immutable `bytes`. * Reduced Sphinx cross-reference warnings and tuned Transformers compatibility warning thresholds. * **Tests** * Increased default unit test timeout on Windows runners. * **Chores** * Updated CI workflow container tags and refreshed linting/typing/docs version pins, plus related mypy configuration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
40936649a3 |
launcher: move NVIDIA-Nemotron-3-Super-120B YAML from Nemotron-h/ to nvidia/ (#1815)
Moves `tools/launcher/examples/Nemotron-h/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/specdec_bench_mtp_vllm.yaml` to `tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/specdec_bench_mtp_vllm.yaml` to match the directory convention for NVIDIA-published models (same family as other models under `nvidia/`). Also updates the `--yaml` path in the header comment. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the example launch command in the YAML header to reference the current NVIDIA example path. * Simplified the commented command formatting into a single line to improve readability and copy/paste convenience. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan D. Yu <5185878+ChenhanYu@users.noreply.github.com> |
||
|
|
7c5741be0c |
Fix launcher Slurm mounts in installed MCP mode (#1811)
## Summary - make Slurm ModelOpt source overlay mounts conditional on source-checkout mode - force managed-source MCP launches to reinstall `modelopt-launcher` from the selected checkout so source refs do not reuse a stale cached package - add regression coverage for installed-mode Slurm mounts and managed-source launcher argv construction ## Root cause PR #1799 correctly stopped packaging `modules/Model-Optimizer/*` when `modelopt-launcher` runs from an installed package. However, `build_slurm_executor()` still unconditionally added container mounts for `code/modules/Model-Optimizer/modelopt` and `modelopt_recipes`. In installed MCP mode those paths do not exist in the remote package, so the container runtime fails before the job script starts. A second issue appeared during validation: the MCP managed-source path can materialize the right git checkout but still execute a cached `modelopt-launcher` package with the same version. Adding `uv run --reinstall-package modelopt-launcher` ensures the selected source ref is what actually runs. ## Validation - `uv run pytest tests/test_bridge.py -q` from `tools/mcp`: 51 passed - `uv run pytest tests/test_slurm_executor.py tests/test_core.py -q` from `tools/launcher`: 24 passed - `pre-commit run --files tools/launcher/core.py tools/launcher/tests/test_slurm_executor.py tools/mcp/modelopt_mcp/bridge.py tools/mcp/tests/test_bridge.py`: passed - Live Slurm GPU smoke validated with patched launcher path; `nvidia-smi` ran successfully and the smoke script completed. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Refactor** * Restructured container mount assembly for Slurm job execution to conditionally mount ModelOpt directories based on optional source path parameter. * Enhanced launcher command-line generation with package management improvements. * Replaced unconditional mount paths with conditional behavior for more flexible resource utilization. * **Tests** * Expanded container mount scenario test coverage for installed and source execution modes. * Tightened test assertions for comprehensive mount behavior verification. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
e2c4d083d4 |
[OMNIML-4922] Four over Six PTQ & Updating Nemotron Ultra Example (#1684)
### What does this PR do? [Four Over Six](https://arxiv.org/pdf/2512.02010) PTQ implementation for weight-only quantization. Four Over Six was used to produce the Nemotron 3 Ultra NVFP4 checkpoint. Also updates the Ultra PTQ example in the launcher to use this new 4/6 config `huggingface/nvidia/Nemotron-3-Ultra-550B-A55B/ptq/ultra-nvfp4-46-max` ### Usage ```bash uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/megatron_lm_ptq.yaml --yes ``` ### Testing - [x] Unit tests pass - [x] Run launcher example ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added NVFP4 Four‑Over‑Six (4/6) adaptive per‑block weight scaling and a configurable FP8 normalization option for FP4/FP8 quantization. * **Documentation** * Added PTQ recipes/presets and updated config docs to document FP8 max variants and the Four‑Over‑Six option. * **Tests** * Added unit and GPU tests validating 4/6 selection, normalization threading, scaling behavior, and reconstruction error checks. * **Chores** * Updated a PTQ pipeline example to use the NVFP4‑46‑max recipe and bumped a container image; minor project config tweak. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> Signed-off-by: Jenny Chen <jennifchen@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
37dbbdac5a |
Fix ModelOpt MCP Slurm launcher submit (#1799)
## Summary - fix launcher Slurm task annotation patching so nemo-run CLI resolves `slurm_factory` correctly for task slots - harden `modelopt-mcp` submit parsing/status resolution and add regression coverage for launcher false-positive success cases - add a minimal `nvidia-smi` smoke YAML/script and fix launcher packaging so source-backed Slurm jobs package required files recursively ## Validation - `uv run pytest tests/test_core.py -q` - `uv run pytest tests/test_bridge.py -q` - dry-run and live-submit validated through the patched local MCP server on `cw_dfw` - interactive smoke job succeeded end-to-end (`nvidia-smi` ran successfully in-container) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an NVIDIA SMI GPU smoke test (script and minimal Slurm YAML example) for launcher integration. * **Bug Fixes** * Improved detection of fatal launcher errors, including when the launcher exits with code 0. * Strengthened Slurm experiment/job identifier parsing and added early rejection of unsafe experiment IDs, with clearer “unparsed”/failure behavior. * Updated sandbox task Slurm config type handling and improved launcher packaging so examples/common are included consistently. * **Tests** * Expanded unit and filesystem-based coverage for parsing/validation, dry-run fatal stderr handling, and nested experiment directory layouts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Signed-off-by: Chenhan D. Yu <chenhany@nvidia.com> |
||
|
|
28b5e26fdb |
Fix torch import error to remove circular dependency & move Nemotron configs (#1606)
### What does this PR do? Type of change: Bug fix when running Megatron-LM modelopt example `generate.py` a circular import causes an error. the cause was because it import modelopt.torch.quantization which imports modelopt/torch --> ``` modelopt/torch/__init__.py:26. The chain is: import modelopt.torch.opt → runs modelopt/torch/__init__.py (parent package) → which pulls in distill → distill/mode.py:25 → back into opt before opt.utils is bound. ``` ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Clarified configuration comments in Nemotron-3-Super-120B quantization recipes for improved clarity. * **Chores** * Internal package initialization updates. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> |
||
|
|
985ad4361c |
launcher: make Slurm memory and user defaults configurable (#1791)
## Summary - Adds `SlurmConfig.mem` and a `SLURM_MEM` env override for launcher Slurm jobs. - Passes configured memory through to `nemo_run.SlurmExecutor` instead of always using `"0"`. - Lets `SLURM_USER` provide the launcher default user when local and cluster usernames differ. - Adds focused tests for Slurm memory defaults and overrides. ## Test plan - [x] `uv run pytest tools/launcher/tests/test_slurm_config.py tools/launcher/tests/test_slurm_executor.py` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Slurm job memory can now be configured via `SLURM_MEM`, with a default of `0` when unset. * Job launch username can now be set via `SLURM_USER`; if omitted, it falls back to the local login name. * **Bug Fixes** * Slurm executor now correctly uses the configured memory value from Slurm settings, falling back to `0` only when missing/empty. * **Tests** * Expanded unit tests to cover default, environment-driven, and executor parameter memory behavior (including missing/empty/`None` cases). <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
2a88b6056c |
launcher: package as modelopt_launcher; mcp: call console script directly (#1766)
## Summary
- `tools/launcher/__init__.py` gains `PACKAGE_DIR`, making the directory
an importable package named `modelopt_launcher` via a `package-dir`
mapping in `pyproject.toml`. **No file moves** — `common/` and
`examples/` stay exactly where they are.
- `pyproject.toml` adds the `packages`/`package-dir`/`package-data`
declarations and a `modelopt-launcher` console script entry point.
- `launch.py` switches imports to
`modelopt_launcher.{core,slurm_config}` (works in both `uv run
launch.py` via editable install and the installed console script), adds
`_has_modelopt_src` to skip packaging modelopt source when running
installed (cluster container already has it), and adds `main()`.
- `bridge.py` deletes the 75-line `_find_launcher_dir()` filesystem
walker and `_launcher_dir_not_found_response()`; simplifies
`_find_launcher_examples_dir()` to 2 strategies (env override → `import
modelopt_launcher`); switches `submit_job` subprocesses from `["uv",
"run", "launch.py"]` with `cwd=launcher_dir` to `["modelopt-launcher"]`
with no `cwd`.
- `tools/mcp/pyproject.toml` declares `modelopt-launcher` as a proper
dependency (dev: editable `../launcher`; published: PyPI), replacing a
lengthy comment explaining why it could not be declared.
- 7 tests for the deleted functions removed; all 38 remaining tests
pass.
## Test plan
- [ ] `cd tools/launcher && uv run python3 -m pytest tests/ -v` — all 65
tests pass
- [ ] `cd tools/mcp && uv run python3 -m pytest tests/ -v` — 38 tests
pass
- [ ] `uv run launch.py --yaml
examples/Qwen/Qwen3-8B/megatron_lm_ptq.yaml --dryrun --yes -v` — dry-run
resolves correctly
- [ ] `modelopt-launcher --yaml
examples/Qwen/Qwen3-8B/megatron_lm_ptq.yaml --dryrun --yes` — console
script works after `pip install -e tools/launcher`
- [ ] `cd tools/mcp && uvx modelopt-mcp` — MCP server starts without
FileNotFoundError
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added a `modelopt-launcher` CLI command as the standardized entry
point for launching optimization jobs.
* **Bug Fixes**
* Improved launcher detection and error reporting when the launcher
isn’t installed.
* Simplified experiment-directory discovery and standardized environment
handling for job submission and log retrieval.
* **Chores**
* Updated launcher packaging and example/resource discovery to work
reliably from installed distributions.
* Added dev-mode symlink and cleanup safeguards.
* **Tests**
* Adjusted expectations for draft PR failure behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
||
|
|
12ae5fbb35 |
[OMNIML-5233] hf_synth.yaml: relative-leaf output_dir contract (#1773)
## What does this PR do? **Type of change:** Bug fix **Overview:** Fix the `output_dir` shape for `tools/launcher/examples/Qwen/Qwen3-8B/hf_synth.yaml` so the downstream synth_run stage (OMNIML-5233) can submit + resume + publish correctly. Replaces the prior `/scratchspace/modelopt/qwen3-8b-synth-v1` (which fails resume + publish) with the canonical **relative-leaf** shape: `hf-local/modelopt/qwen3-8b-synth-v1`. The absolute team-folder prefix is **deliberately not committed** — it's an NVIDIA-internal cluster mount path, and `pensieve-intern`'s `_scan_for_internal_path_leak` guard refuses any YAML diff that bakes it (rightly, since this repo is public). The downstream stage that has the cluster context (synth_run) resolves the prefix via `mcp__nmm-sandbox__resolve_team_folder` at submit time and injects the absolute path through `extra_overrides`. **Why this shape:** the relative leaf `hf-local/modelopt/<id>` carries the contract between authoring (this stage), submitting (synth_run), and publishing (`modelopt-storage publish`). Publish promotes from the team's `hf-local/` tree as a metadata-only op; the leaf shape tells the operator + downstream what category the artifact lands in, without leaking the cluster mount. **Companion MR:** [pensieve-intern !126](https://gitlab-master.nvidia.com/omniml/integration/pensieve-intern/-/merge_requests/126) lands the matching SPEC contract in synth_support.md + synth_run.md, AND elevates the path-contract rule to the engine preamble (`_ENGINE_AGENT_RULES_PREAMBLE`) so every agent dispatch — across agent/subprocess/pensieve-artifact tasks — reads it as workflow-wide common knowledge. ## Usage Same as before. The shard write location is now resolved at submit time by synth_run. ## Testing - [x] One-line YAML change. - [ ] End-to-end: re-fire OMNIML-5233 synth_run after this + the pensieve-intern MR merge; expect the agent to submit the slurm array and stamp success. ## Before your PR is "Ready for review" - [x] Make sure you read and follow Contributor guidelines and your commits are signed. - [x] Is this change backward compatible? **Yes** (Qwen3-8B is the only consumer; no published checkpoints reference the old path). - [x] Did you write any new necessary tests? **N/A** — example YAML; behavior covered by the downstream synth_run agent run. - [x] Did you add or update any necessary documentation? Covered in the comment block on the changed lines. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated the configured output directory to a more stable, persist-friendly path so generated artifacts remain available across repeated runs and re-dispatches. * Added clearer inline guidance explaining how the resolved absolute path under the new directory is used to reuse prior shards and to handle publishing consistently. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
fa1d13f85d |
launcher: add Nemotron-3-Super-120B-A12B-BF16 MTP vLLM specdec bench config (#1714)
Adds a SPEED-bench MTP speculative-decoding YAML for `NVIDIA-Nemotron-3-Super-120B-A12B-BF16` via vLLM. Covers two splits: - `qualitative` — 32 concurrent, 4096 output tokens - `throughput_32k` — 8 concurrent, 80 requests, 4096 output tokens Both tasks run `tp_size=4` on a single 4×H100/A100 node. Part of OMNIML-5095 / OMNIML-5098. ## Test plan - [ ] `uv run slurm.py --yaml modules/Model-Optimizer/tools/launcher/examples/Nemotron-h/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/specdec_bench_mtp_vllm.yaml --dryrun --yes -v` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Chores** * Added a new benchmark configuration for evaluating NVIDIA Nemotron-3-Super-120B model performance using speculative decoding with vLLM. Configuration includes qualitative and throughput benchmark tasks. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
93dd08f429 |
[OMNIML-4760] synth_support (#1696)
Draft PR opened by **pensieve-intern** for [OMNIML-4760](https://jirasw.nvidia.com/browse/OMNIML-4760). Stage `synth_support` of Epic `OMNIML-4755`. The agent ran from the SPEC on the ticket description; review every change before marking ready. _Always-draft is enforced — the bot never auto-merges._ --- **Agent's self-narration** (stripped from PR diff; surfaced here for context): `VERIFICATION_COMMENT.txt`: ``` OMNIML-4760 synth_support verification for qwen3-8b found and fixed one task_0 issue in `tools/launcher/examples/Qwen/Qwen3-8B/hf_offline_eagle3.yaml`. What changed: - Updated `task_0` data input from missing `/hf-local/modelopt/Speculative-Decoding-Prompt-Samples` to existing `/hf-local/modelopt/Speculative-Decoding-Dataset-v1-Qwen3-8B/default-openai.jsonl`. - Left the rest of the 4-task monolithic `hf_offline_eagle3.yaml` unchanged. Verification: - `task_0.script` is `common/tensorrt_llm/query.sh`. - `--model <<global_vars.hf_model>>` resolves to `/hf-local/Qwen/Qwen3-8B`. - `task_0.slurm_config.container` is `nvcr.io/nvidia/tensorrt-llm/release:1.2.0`. - Cluster validation on cw_dfw succeeded for the fixed task_0 data path: experiment `cicd_1781221901`, Slurm job `12739762`, remote directory `/lustre/fsw/portfolios/coreai/users/chenhany/experiments/cicd/cicd_1781221901/Qwen3-8B_EAGLE3_offline_task0_verify_jsonl_0`. - Log evidence included successful SSH tunnel/authentication, TensorRT-LLM 1.2.0 container import, Qwen3-8B server health checks, a successful `/v1/chat/completions` request, and loading `1393367` train examples from the Qwen3-8B JSONL data file. Next: - Runner should open the single-file PR for human review because this was a task_0 config fix, not verification-only. ``` _Pollution-strip removed `VERIFICATION_COMMENT.txt` from this commit (sidecar narration and/or incidental lockfile regeneration are never part of the agent's intended deliverable)._ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Enhanced TE grouped MoE weight quantization with optional per-GEMM/per-expert quantizers (environment-controlled) and updated calibration flow * Expanded launcher CLI with `--shard-id` and per-run sample limiting via `--num-samples` * Added a Qwen3-8B standalone vLLM synthesis job example (`hf_synth.yaml`) * Added Slurm job requeue support * **Bug Fixes** * Improved sharded dataset synthesis with idempotent `.done` markers and smarter shard sizing/capping * **Tests** * Added coverage validating TEGrouped vs sequential MoE default amax behavior and output divergence/accuracy <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Pensieve Intern <pensieve-intern@nvidia.com> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Pensieve Intern <pensieve-intern@nvidia.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
e0125294f9 |
launcher: add Qwen3-8B/specdec_bench_dflash_vllm.yaml parent (OMNIML-5057) (#1764)
## Summary Adds the parent YAML for the Qwen3-8B / DFlash / vLLM SPEED-bench sweep (Epic OMNIML-5057, 4 cells: t0_d3 / t0_d7 / t1_d3 / t1_d7). Mirror of [`Qwen/Qwen3.5-4B/specdec_bench_dflash_vllm.yaml`](https://github.com/NVIDIA/Model-Optimizer/blob/main/tools/launcher/examples/Qwen/Qwen3.5-4B/specdec_bench_dflash_vllm.yaml). ### Differences vs Qwen3.5-4B template - `hf_model`: `/hf-local/Qwen/Qwen3-8B` - `draft_model`: `/hf-local/modelopt/Qwen3-8B-DFlash-bs8-seq4096-150000` — an internal NVIDIA checkpoint, not on HF Hub. `cell.md` Step 1's `PRIVATE_DRAFT_ORGS` branch already covers the expected 404; the cluster has the file staged by the Epic's `prep_inputs` stage. - `block_size`: 8 (was 4) — matches the canonical `cell_t0_d7` draft_length=7. ### Known limitation (acceptable per Shape (2)) Qwen3-8B's `max_position_embeddings = 40960`. The `throughput_32k` split contains rows whose prompts exceed 40960 input tokens; those rows fail with the vLLM context-limit assert. Per `cell.md` "Shape (2)" recovery contract, cells ship qualitative metrics + `null` throughput_32k AL. See OMNIML-5060 Notes (2026-06-15 correction) for the empirical evidence. ### Reference run `cicd_1781655951` (OMNIML-5060 pipeline #55054230, cw_dfw): - qualitative `Average_AL` = **3.4849** - throughput_32k = `null` (Shape (2) — overlong rows skipped) ### Cascade 3 subprocess cells (OMNIML-5059 / 5061 / 5062) re-run slurm against this parent YAML at invoke time with per-cell CLI overrides for temperature / block_size / save_dir / max_seq_len. ## Test plan - [x] `tools/precommit/check_launcher_yaml.py` passes locally - [x] YAML successfully drives a real slurm submission (`cicd_1781655951`) producing valid qualitative SPEED-bench output - [ ] Once merged: `intern_advance OMNIML-5057` fires the 3 subprocess cells against this YAML 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Added benchmark configuration file for evaluating speculative decoding performance with Qwen3-8B model using DFlash acceleration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
e6790ef7b4 |
[Examples]: GPT-oss, Qwen3Moe streaming specdec example (#1692)
### What does this PR do?
Type of change: new example
Adds **streaming speculative-decoding examples (EAGLE3 + DFlash)** for
**gpt-oss-20b** and **Qwen3-30B-A3B** to the ModelOpt launcher,
mirroring the existing Qwen3-8B/Kimi examples.
- New yamls:
`tools/launcher/examples/{openai/gpt-oss-20b,Qwen/Qwen3-30B-A3B}/hf_streaming_{eagle3,dflash}_multi_node.yaml`,
plus gpt-oss `chat_template_train.jinja` (generation-tagged, for
`answer_only_loss`).
- `eagle_utils.py`: the streaming path now installs a custom
`data.chat_template` on the tokenizer (the online path already did) —
needed for the tagged template.
### Usage
```bash
cd tools/launcher
export SLURM_HOST=... SLURM_ACCOUNT=... SLURM_HF_LOCAL=... SLURM_JOB_DIR=...
uv run launch.py --yaml examples/openai/gpt-oss-20b/hf_streaming_eagle3_multi_node.yaml --yes
```
### Testing
Pipeline sanity test on **unsynthesized** data (daring-anteater), 1×
H100-80GB, 12k steps. All four train and pass the vLLM acceptance-length
eval:
| Model | Method | Train speed | vLLM AL |
|---|---|---|---|
| Qwen3-30B-A3B | EAGLE3 | 7.12 it/s | **1.74** |
| Qwen3-30B-A3B | DFlash | 2.31 it/s | 1.29 |
| gpt-oss-20b | EAGLE3 | 5.07 it/s | 1.19 |
| gpt-oss-20b | DFlash | 2.01 it/s | 1.14 |
<img width="1300" height="780" alt="image"
src="https://github.com/user-attachments/assets/2be22562-5b77-4a9f-9dc0-6f936a059736"
/>
> Sanity test only, not a quality run. gpt-oss AL is low because it is a
reasoning model (CoT at inference) while daring-anteater has no
reasoning traces and `answer_only_loss` masks all but the final content
— quality runs need synthesized/reasoning data.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added multi-node speculative decoding pipeline configurations for
Qwen3-30B-A3B and gpt-oss-20b with DFlash and EAGLE3 support.
* Introduced chat template training support for improved model
instruction formatting.
* **Enhancements**
* Increased benchmark concurrency from 1 to 32 across Qwen3-8B
configurations for more realistic performance evaluation.
* Extended training runs from 500 to 2000 steps for Kimi-K2.5 models.
* Improved chat template handling in speculative decoding workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
e004d8d90e |
DFlash speculative decoding for MiniMax-M2.7 (FSDP2): auto mask-token, FSDP2 resume fixes, per-checkpoint draft export (#1621)
## What
Brings up DFlash block-diffusion speculative decoding for large MoE
targets (MiniMax-M2.7, 229B) trained under accelerate FSDP2, and fixes
the regressions that broke checkpoint resume and per-checkpoint draft
export.
## Commits
- **auto-add mask token for DFlash** when the tokenizer lacks one
(resize embeddings, restore dtype).
- **requeue support** in `build_slurm_executor` + **FSDP2
cpu_ram_efficient_loading** for 229B on multi-node.
- **FSDP2 buffer patch** (`fsdp2_buffer_patch.py`): handle non-DTensor
buffers in `fsdp2_load_full_state_dict`, broadcast dtype codes from rank
0, and an FSDP2-safe `clip_grad_norm_`. Required because MiniMax-M2.7
pins transformers 4.57.x (no native `ParallelismConfig`).
- **dtype fix**: use the broadcast dtype (rank 0) rather than the local
meta-device param dtype, so non-leader ranks don't cast bf16 back to
fp32 on resume.
- **restore `DFlashExportCallback`** (this PR's headline): the
Pydantic-recipe refactor (
|
||
|
|
601401b134 |
[OMNIML-5025] cell_t0_d7 (#1738)
Draft PR opened by **pensieve-intern** for [OMNIML-5025](https://jirasw.nvidia.com/browse/OMNIML-5025). Stage `cell_t0_d7` of Epic `OMNIML-5022`. The agent ran from the SPEC on the ticket description; review every change before marking ready. _Always-draft is enforced — the bot never auto-merges._ --- **Agent's self-narration** (stripped from PR diff; surfaced here for context): `INTERN_ARTIFACTS.json`: ``` { "sweep_name": "gemma-4-E4B-it_mtp_vllm_t0_d7", "experiment_id": "cicd_1781548540", "experiment_dir": "/lustre/fsw/portfolios/coreai/projects/coreai_dlalgo_modelopt/cicd/cicd/cicd_1781548540/", "AL_qualitative_overall": 3.2945, "AL_qualitative_categories": { "coding": 4.8883, "humanities": 2.703, "math": 4.1139, "multilingual": 4.1401, "qa": 2.59, "rag": 3.9329, "reasoning": 3.6251, "roleplay": 1.8026, "stem": 3.2041, "summarization": 2.8275, "writing": 2.4118 }, "AL_throughput_32k_overall": 3.3803, "AL_throughput_32k_categories": { "high_entropy": 2.0839, "low_entropy": 4.3797, "mixed": 3.7143 } } ``` `VERIFICATION_COMMENT.txt`: ``` Completed OMNIML-5025 cell_t0_d7 for `google/gemma-4-E4B-it` / MTP / vLLM. What was done: - Authored/kept `tools/launcher/common/specdec_bench/_cells/gemma-4-E4B-it_mtp_vllm_t0_d7.yaml` for sweep `gemma-4-E4B-it_mtp_vllm_t0_d7`. - Updated/kept `tools/launcher/examples/google/gemma-4-E4B-it/specdec_bench_mtp_vllm.yaml` with the Gemma-specific `vllm/vllm-openai:gemma` container after generic v0.22.1 failed startup. - Verified HF Hub model endpoints for `google/gemma-4-E4B-it` and `google/gemma-4-E4B-it-assistant` returned HTTP 200. - Wrote `INTERN_ARTIFACTS.json` with parsed AL metrics from the successful cluster run. Metrics extracted: - experiment_id: `cicd_1781548540` - experiment_dir: `/lustre/fsw/portfolios/coreai/projects/coreai_dlalgo_modelopt/cicd/cicd/cicd_1781548540/` - qualitative Average_AL: `3.2945` - throughput_32k Average_AL: `3.3803` PR status: - PR opened: NONE — this runner prompt says not to commit, push, or create PRs because the runner handles that. What's next: - Engine should consume `INTERN_ARTIFACTS.json` and advance the downstream wrap-up/aggregation stage. Trigger pipeline URL: - `https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/pipelines/54841359` Slurm job status: - submission: completed successfully - experiment_id: `cicd_1781548540` - experiment_dir: `/lustre/fsw/portfolios/coreai/projects/coreai_dlalgo_modelopt/cicd/cicd/cicd_1781548540/` ``` _Pollution-strip removed `INTERN_ARTIFACTS.json`, `INTERN_LEARNING_TICKET.md`, `VERIFICATION_COMMENT.txt` from this commit (sidecar narration and/or incidental lockfile regeneration are never part of the agent's intended deliverable)._ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated benchmark configuration to use an optimized container image for Gemma 4 MTP speculative-decoding benchmarks. * Revised documentation in the configuration to reflect the current image and its capabilities. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: pensieve-intern agent <noreply@nvidia.com> |
||
|
|
c661366352 |
launcher: Nemotron-120B specdec_bench — kill stale _cells/ reference (post-PR-#1564) (#1741)
## What Comment-only change to the Nemotron-3-Super-120B-A12B-BF16 DFlash parent YAML doc-comment header. Rewrites the example invocation from the pre-PR-#1564 pattern (with `--runtime_params common/specdec_bench/_cells/<sweep_name>.yaml`) to the current post-#1564 pattern (CLI-flag-only overrides, no per-cell file). No yaml content / config change — only the prose example in the header. ## Why PR #1564 removed the `tools/launcher/common/specdec_bench/_cells/` and `_runtime_params/` directories. Cell-specific knobs are now CLI overrides at slurm-invoke time, not committed files. The cell SPEC in pensieve-intern (`specdec_bench/specs/cell.md`) is explicit about this — five times in prose. But this parent YAML's doc-comment kept advertising the OLD pattern. When pensieve-intern's agent runner scans `tools/launcher/examples/` for a reference invocation (good practice — agents should imitate working examples), it lands on this Nemotron parent and copies the stale pattern. **Five recent agent dispatches on the gemma-4 Epic OMNIML-5022 (cells OMNIML-5024 / 5025 / 5026 / 5027) authored new `_cells/<sweep>.yaml` files for this reason, despite the SPEC telling them not to.** Prose loses to a concrete checked-in counter-example. The Qwen3.5-4B reference template that cell.md officially points at (`tools/launcher/examples/Qwen/Qwen3.5-4B/specdec_bench_mtp_vllm.yaml`) is clean and shows only the bare `uv run slurm.py --yaml ...` form. This PR makes Nemotron-120B consistent with that template. ## How surfaced Diagnosed 2026-06-15 on OMNIML-5025 cell_t0_d7 ([intern-agent job 341631795](https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/341631795)). The cell agent authored `_cells/gemma-4-E4B-it_mtp_vllm_t0_d7.yaml` — the engine's diff-shape classifier should have rejected it, but didn't (tracked separately as OMNIML-5170). Root-causing the agent's behavior surfaced this stale doc-comment as the source of the pattern. ## Verification `grep -rn '_cells\|runtime_params common'` across the entire launcher tree returned only this file. After this PR, the launcher tree carries zero stale references. Pairs with NVIDIA/Model-Optimizer#1738 (gemma-4-E4B-it container fix) and OMNIML-5170 (engine-side classifier hard-reject for `_cells/` paths — defense in depth). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated example documentation to clarify how to override per-cell parameters using CLI flags in Slurm configuration runs. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
55a2101e2f |
Update Nemotron-3 Pruning, Distillation and PTQ results based on new shared calibration loop with seq packing and add tool-calling eval fix (#1660)
### What does this PR do? Type of change: documentation + minor example-script tweaks Follow-up to #1601. Originally scoped to add **NVFP4 + QAD**, this PR was **repurposed** to refresh the [Nemotron-3-Nano-30B-A3B-BF16 tutorial](examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md) results using the **new shared calibration loop (sequence packing)** and to **fix tool calling in evaluation**. - Refreshed the prune → distill → eval → **FP8** results (accuracy + vLLM throughput tables) with the new calibration loop. - **Tool-calling eval fix** (`nemo_evaluator.yaml`): GPQA and AIME now run the Python sandbox tool. The tutorial reports both **with-tools** and **no-tools** GPQA/AIME and shows `mean ± std_dev`. - Script tweaks: `quantize.py` calibration now uses sequence packing (`pack=True`) which leads to slight improvement in PTQ; `prune_minitron.py` defaults `inference_batch_size` to `calib_batch_size`. ### Testing Documentation + small example-script changes; tutorial relative links resolve and the results tables / figure were verified consistent. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: Yes - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ (will run `/claude review`) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the main guide and evaluator instructions for prune + distill + FP8/NVFP4 quantization, including refreshed vLLM deployment tips, benchmark/noise presentation, and long-context tool-calling attribution notes. * Refreshed README technique examples/links, reordered the model support matrix rows, and improved pruning overview/support-matrix text. * **Changes to Examples** * NAS pruning now documents higher GPU memory usage vs manual pruning; pruning batching defaults were improved. * Quantization PTQ calibration uses packed document packing; quantized checkpoint export messaging was streamlined. * Updated pruning/distillation/quantization tutorial guidance, metrics/tables, command parameters, and evaluator YAML settings (KV-cache dtype, generation defaults, task behavior). <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
d7df14d12a |
[tools/debugger] Enforce a single relay owner across hosts (#1735)
### What does this PR do? Type of change: Bug fix The `tools/debugger` file-based relay assumes a single server but never enforced it. Because the relay lives on shared NFS (the repo is often the same checkout mounted on multiple hosts), a forgotten `server.sh` on another host kept polling the same `.relay/` and could **steal commands** (executing them on the wrong host), and killing one server's cleanup could **wipe the active server's markers**. This adds a `.relay/owner` ownership token (`host:pid:nanos`): - Each server writes `owner` atomically at startup and **takes over** instead of refusing when a stale `server.ready` exists (the old `kill -0 <pid>` guard was host-local and meaningless across hosts). - The handshake and main loops exit cleanly if `owner` changes (`[server] Superseded by <id> — exiting.`), so a freshly started server **evicts** any stale one — even on another host. - `cleanup()` only clears shared markers if we still own them, so a stepping-down server never clobbers its successor's `server.ready`/`owner`. Also gitignores `tools/debugger/logs/` and documents the `owner` file in the README. ### Usage ```bash # Inside the container; a previously-running server elsewhere that shares this # NFS .relay/ steps down automatically once this one claims ownership: bash tools/debugger/server.sh # [server] Note: existing server.ready found (<host:pid:ts>); taking over. # (the stale server logs: "[server] Superseded by <id> — exiting.") ``` ### Testing Verified live on computelab: a forgotten `server.sh` on another host was evicted when a new server started, after which `client.sh run` executed on the correct (new) host; confirmed the stepping-down server's cleanup does not remove the successor's `server.ready`/`owner`. `server.sh` passes `bash -n`. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ <!-- additive: new .relay/owner file; client.sh and the wire protocol are unchanged --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A <!-- the file-based relay tool has no test harness; behavior verified manually --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- internal dev tooling, not a shipped feature/API --> - Did you get Claude approval on this PR?: N/A <!-- can run /claude review --> ### Additional Information Scope is limited to `tools/debugger/` (`server.sh`, `README.md`, `.gitignore`). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated relay protocol documentation to clarify ownership-based server coordination. * **Bug Fixes** * Improved reliability of multi-server coordination in shared relay environments. * **Chores** * Updated ignore patterns for logging files. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |