Commit Graph
86 Commits
Author SHA1 Message Date
Keval Morabia b6bf6b7997 Update Roadmap Issue link
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-06-23 21:42:37 +05:30
jzh26andh-guo18 48f8d8984f Fix conversation loading logic in UltraChat dataset (#1680)
### What does this PR do?
Type of change: Bug fix.

<!-- Details about the change. -->
Previously, only the first user prompt was extracted from each example,
discarding all subsequent turns. UltraChat stores full multi-turn
conversations in the "messages" field, so switching to that field
preserves the complete dialogue rather than truncating to a single user
message.
### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
 python make_dataset.py -f test_cfg.yaml --full
 python make_dataset.py -f test_cfg.yaml
```yaml
      - name: "ultrachat"
        splits:
          train_gen: 100
          train_sft: 100
```
### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Breaking Changes**
* Removed UltraChat support from the example dataset mixer, so UltraChat
splits can no longer be loaded.

* **Documentation**
* Updated the dataset examples README and the example dataset
configuration to remove UltraChat and adjust split settings for other
datasets.
* Updated the speculative decoding fine-tuning documentation to
reference Daring-Anteater instead of UltraChat.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: jzh26 <226629529+jzh26@users.noreply.github.com>
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-06-22 09:18:35 +00:00
h-guo18 9048d13b86 [Feat]:Support DPace (#1724)
### What does this PR do?

Type of change: New feature

Adds the **D-PACE** (Dynamic Position-Aware Cross-Entropy) loss
objective for DFlash speculative-decoding training
([arXiv:2605.18810](https://arxiv.org/abs/2605.18810)). It replaces the
static exponential position decay with per-position CE weights derived
from the draft's own confidence `q_i = exp(-CE_i)`: smoothed `q̃_i =
(1-α)q_i + α` (Eq.7) and weighted by the suffix-sum of prefix products
`w_j = Σ_{m≥j} ∏_{i≤m} q̃_i` (Eq.8), which directly targets expected
accepted block length and shifts signal toward whichever positions
currently limit acceptance.

Selected via `dflash_loss_objective` — **D-PACE is now the default**
(`dpace`); set `dflash_loss_objective: decay` to restore the previous
static schedule. Smoothing via `dflash_dpace_alpha` (default 0.5).
Weights are detached from the gradient — training-only, ~2.3% overhead,
no architecture or inference change. Mutually exclusive with
`dflash_loss_decay_factor`.

### Usage

```yaml
# DFlash recipe / training config
dflash:
  dflash_loss_objective: dpace   # default: decay
  dflash_dpace_alpha: 0.5        # smoothing in (0, 1]; stable in [0.3, 0.7]
```

### Testing

CPU unit tests in
`tests/unit/torch/speculative/plugins/test_hf_dflash.py`: weights match
the paper closed form, are detached and non-increasing, the α smoothing
floor keeps later weights non-zero, and convert wires/validates the new
fields (rejects bad objective and degenerate α). Training validated on
Qwen3-8B (curve below).

<img width="1803" height="809" alt="image"
src="https://github.com/user-attachments/assets/d34dcd76-9e46-4051-94d4-c880b1987965"
/>

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ⚠️ Behavior change — D-PACE is
now the **default** objective, so DFlash training loss weighting changes
unless you set `dflash_loss_objective=decay` (which reproduces the
previous static-decay behavior).
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new
dependency)
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

Reference: D-PACE, [arXiv:2605.18810](https://arxiv.org/abs/2605.18810).
See `examples/speculative_decoding/doc/dflash.md` for the math and
tuning notes.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added a **D-PACE** training loss objective for DFlash speculative
decoding (`dflash_loss_objective: dpace`), configurable via
`dflash_dpace_alpha` (default `0.5`).
* **Documentation**
* Documented D-PACE’s confidence-derived, dynamically weighted
per-position loss behavior (training-only) and noted that
`dflash_loss_decay_factor` is ignored with D-PACE.
* **Bug Fixes**
* Updated DFlash loss to reuse the precomputed per-token cross-entropy
in the non-KD path.
* **Tests**
* Added unit tests for D-PACE weight correctness, masking, gradient
detachment, monotonicity, smoothing, and config/validation behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-06-20 02:14:31 +00:00
h-guo18 977d34dc3c [Fix](nvbug6304585): specdec README online base-model example should use Instruct model (#1755)
### What does this PR do?

Type of change: Bug fix

The **Training Draft Model with Online/Offline Base Model** examples in
`examples/speculative_decoding/README.md` used
`meta-llama/Llama-3.2-1B`, a
base / pretrained checkpoint that ships **no chat template**. The online
EAGLE3
flow tokenizes conversations through
`tokenizer.apply_chat_template(...)`, so
the data collator fails fast at startup:

```
ValueError: No valid chat template!
```

This PR:

- Switches both README example commands (online and offline) to
  `meta-llama/Llama-3.2-1B-Instruct`, which carries a chat template.
- Makes the collator error message in
`modelopt/torch/utils/plugins/transformers_dataset.py` actionable — it
now
explains the cause (base checkpoints have no chat template) and points
users
  at an Instruct model or a custom `chat_template`.

### Usage

```bash
./launch_train.sh \
    --config ../../modelopt_recipes/general/speculative_decoding/eagle3.yaml \
    model.model_name_or_path=meta-llama/Llama-3.2-1B-Instruct \
    data.data_path=input_conversations/train.jsonl \
    training.output_dir=ckpts/llama-3.2-1b-online
```

### Testing

- Reproduced the original `No valid chat template!` failure with the
base
`Llama-3.2-1B` and confirmed the Instruct variant carries a chat
template.
- Verified the new error message renders correctly when a template is
missing.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌

### Additional Information

Fixes nvbug 6304585: https://nvbugspro.nvidia.com/bug/6304585


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Documentation**
* Updated speculative decoding example training commands to reference
the Llama-3.2-1B-Instruct model.

* **Bug Fixes**
* Enhanced error message when chat template configuration is missing,
providing actionable guidance for resolution.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-06-16 16:42:20 -07:00
h-guo18 e6790ef7b4 [Examples]: GPT-oss, Qwen3Moe streaming specdec example (#1692)
### What does this PR do?

Type of change: new example

Adds **streaming speculative-decoding examples (EAGLE3 + DFlash)** for
**gpt-oss-20b** and **Qwen3-30B-A3B** to the ModelOpt launcher,
mirroring the existing Qwen3-8B/Kimi examples.

- New yamls:
`tools/launcher/examples/{openai/gpt-oss-20b,Qwen/Qwen3-30B-A3B}/hf_streaming_{eagle3,dflash}_multi_node.yaml`,
plus gpt-oss `chat_template_train.jinja` (generation-tagged, for
`answer_only_loss`).
- `eagle_utils.py`: the streaming path now installs a custom
`data.chat_template` on the tokenizer (the online path already did) —
needed for the tagged template.

### Usage

```bash
cd tools/launcher
export SLURM_HOST=... SLURM_ACCOUNT=... SLURM_HF_LOCAL=... SLURM_JOB_DIR=...
uv run launch.py --yaml examples/openai/gpt-oss-20b/hf_streaming_eagle3_multi_node.yaml --yes
```

### Testing

Pipeline sanity test on **unsynthesized** data (daring-anteater), 1×
H100-80GB, 12k steps. All four train and pass the vLLM acceptance-length
eval:

| Model | Method | Train speed | vLLM AL |
|---|---|---|---|
| Qwen3-30B-A3B | EAGLE3 | 7.12 it/s | **1.74** |
| Qwen3-30B-A3B | DFlash | 2.31 it/s | 1.29 |
| gpt-oss-20b | EAGLE3 | 5.07 it/s | 1.19 |
| gpt-oss-20b | DFlash | 2.01 it/s | 1.14 |

<img width="1300" height="780" alt="image"
src="https://github.com/user-attachments/assets/2be22562-5b77-4a9f-9dc0-6f936a059736"
/>

> Sanity test only, not a quality run. gpt-oss AL is low because it is a
reasoning model (CoT at inference) while daring-anteater has no
reasoning traces and `answer_only_loss` masks all but the final content
— quality runs need synthesized/reasoning data.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added multi-node speculative decoding pipeline configurations for
Qwen3-30B-A3B and gpt-oss-20b with DFlash and EAGLE3 support.
* Introduced chat template training support for improved model
instruction formatting.

* **Enhancements**
* Increased benchmark concurrency from 1 to 32 across Qwen3-8B
configurations for more realistic performance evaluation.
  * Extended training runs from 500 to 2000 steps for Kimi-K2.5 models.
  * Improved chat template handling in speculative decoding workflows.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-06-15 16:49:39 -07:00
yeyu-nvidiaandClaude Opus 4.6 e004d8d90e DFlash speculative decoding for MiniMax-M2.7 (FSDP2): auto mask-token, FSDP2 resume fixes, per-checkpoint draft export (#1621)
## What

Brings up DFlash block-diffusion speculative decoding for large MoE
targets (MiniMax-M2.7, 229B) trained under accelerate FSDP2, and fixes
the regressions that broke checkpoint resume and per-checkpoint draft
export.

## Commits
- **auto-add mask token for DFlash** when the tokenizer lacks one
(resize embeddings, restore dtype).
- **requeue support** in `build_slurm_executor` + **FSDP2
cpu_ram_efficient_loading** for 229B on multi-node.
- **FSDP2 buffer patch** (`fsdp2_buffer_patch.py`): handle non-DTensor
buffers in `fsdp2_load_full_state_dict`, broadcast dtype codes from rank
0, and an FSDP2-safe `clip_grad_norm_`. Required because MiniMax-M2.7
pins transformers 4.57.x (no native `ParallelismConfig`).
- **dtype fix**: use the broadcast dtype (rank 0) rather than the local
meta-device param dtype, so non-leader ranks don't cast bf16 back to
fp32 on resume.
- **restore `DFlashExportCallback`** (this PR's headline): the
Pydantic-recipe refactor (7038dec918) dropped the callback that exported
the draft submodule after each checkpoint save, leaving a stale "export
happens during training via DFlashExportCallback" comment with no
callback. FSDP2 SHARDED_STATE_DICT checkpoints carry no
`model.safetensors`, so without it there is nothing for vLLM /
acceptance-length eval to load. The callback gathers only the ~328 MB
draft submodule across shards via `get_model_state_dict(...,
submodules={dflash_module}, full_state_dict=True, cpu_offload=True)` —
works under SHARDED_STATE_DICT without materializing the 229B base — and
writes `exported-checkpoint-{step}/`.

## Testing
- Resume from FSDP2 sharded checkpoints verified end-to-end (loss/AR
continuity).
- Draft export validated against vLLM: exported drafts load and produce
acceptance-length metrics on MT-Bench across the full checkpoint sweep.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Export draft-submodule weights to dedicated exported checkpoints
during training.
* FSDP2 buffer compatibility and DTensor-aware gradient clipping for
safer distributed loading/training.
  * Detect HF-format checkpoints for smarter resume/load behavior.
  * Auto-add and handle a mask special token for draft workflows.
  * vLLM: disable prefix caching to preserve full prompt hidden states.
  * Add CLI option for answer-only loss and save aligned loss masks.
  * Add a SPEED-Bench config for DFLASH/vLLM benchmarking.

* **Chores**
  * Robust package version fallback to avoid import failures.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-06-15 16:47:50 -07:00
h-guo18 46eddab877 [Feat]: Specdec Streaming: RDMA + Multinode (#1611)
### What does this PR do?

Type of change: New feature

Multi-node **streaming** training for speculative decoding (EAGLE3 /
DFlash):
a live `vllm serve` captures the target model's hidden states and moves
them
straight to the trainer over **NIXL RDMA** — no disk round-trip. The
streaming
dataset is map-style — each rank fetches only its own
`DistributedSampler` shard
(concurrency from `dataloader_num_workers`), round-robins across
multiple serve
replicas (`server_urls`), and scales to multi-node DDP. Serve-side
tensor
parallelism (TP>1) is supported: hidden states are replicated across TP
ranks, so
rank 0 alone owns the pool + transfer.

### How

- `RdmaHiddenStatesConnector` — out-of-tree vLLM connector (no vLLM
source edits):
one pre-registered pinned NIXL pool per serve, a ring slot per request,
and a
small HTTP sidecar serving transfer metadata. The trainer RDMA-READs the
slot
  into a per-worker buffer. RDMA is the **only** transport (the earlier
  disk/safetensors path is removed).
- Map-style dataset + multi-node accelerate launch (`--machine_rank`,
optional
  Slurm `--segment` to keep nodes in one NVLink domain).

### Usage

```yaml
data:
  mode: streaming
  streaming_server_url: "http://node0:8000,http://node1:8000"  # round-robin
```

### Validation (Qwen3-8B, oci-nrt H100)
sandbox CI:
https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/337489812

**1. End-to-end convergence — EAGLE3 & DFlash, 5000 steps.** Both
algorithms
converge and export a deployable draft; the DFlash drafts also serve
under vLLM
speculative decoding (8/8 smoke prompts pass).

| algorithm | topology (nodes) | train loss (step 0 → 5000) | vLLM draft
acc-len |
|---|---|---|---|
| EAGLE3 | 2 serve TP=2 + 2 trainer DDP (4) | 37.1 → 8.20 | — |
| DFlash | 1 serve TP=1 + 1 trainer (2)      | 11.7 → 5.56 | 1.11 |
| DFlash | 2 serve TP=2 + 2 trainer DDP (4)  | 10.9 → 5.26 | 1.19 |

<!-- Drag these PNGs in here (GitHub turns them into asset URLs):
eagle3_streaming_loss.png, dflash_streaming_loss_singlenode.png,
dflash_streaming_loss_multinode.png -->

**2. Scalability — 1 → 12 nodes (EAGLE3, 200 steps).** Throughput scales
~23× across the
sweep below. The step-time growth is cross-node DDP all-reduce, not the
streaming path —
RDMA (~0.33 ms/req @ 2 MB, ~47 GB/s host-pinned READ) is never the
bottleneck. Scale
serve + trainer nodes together for near-linear speedup.

| serve / trainer | nodes | step time | samples / step | samples / sec
(global) | acc @ step 200 |
|---|---|---|---|---|---|
| 1 serve / 1 rank (co-located, 1 node 2 GPU) | 1 | 0.23 s | 1 | 4.4 |
[0.141, 0.094, 0.072] |
| 1 serve / 1 rank (cross-node) | 2 | 0.23 s | 1 | 4.3 | [0.137, 0.105,
0.074] |
| 2 serve / 8 ranks | 3 | 0.26 s | 8 | 31.1 | [0.215, 0.126, 0.097] |
| 4 serve / 16 ranks (2 trainer nodes) | 6 | 0.28 s | 16 | 56.5 |
[0.217, 0.148, 0.110] |
| 8 serve / 32 ranks (4 trainer nodes) | 12 | 0.31 s | 32 | 101.9 |
[0.235, 0.165, 0.137] |

**3. Serve-side TP correctness.** TP=1 vs TP=2 draft top-1 accuracy
track
step-for-step (hidden states are replicated across TP ranks).

<img width="910" height="546" alt="serve-tp-acc"
src="https://github.com/user-attachments/assets/73df9214-7ff0-4ab4-bf2f-95842b12cd5f"
/>

### Before your PR is "*Ready for review*"

- Backward compatible?: ❌ — streaming is now RDMA-only; `server_url` →
`server_urls`;
the disk transport (`HS_TRANSPORT`, `streaming_shared_storage_path`) is
removed.
- New tests?: ✅
`tests/unit/torch/speculative/plugins/test_hf_streaming_dataset.py`
  (map-style dataset + mocked RDMA fetch).
- Updated Changelog?: ❌
- Claude approval?: ❌

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-06-10 18:35:02 -07:00
yeyu-nvidiaandClaude Opus 4.6 5bd04c3876 Revert unverified EAGLE3 model examples; keep triage code + baseline (#1623)
### What does this PR do?

Type of change: Revert / cleanup (follow-up to #1417)

Per review feedback (@h-guo18): `main` should be production-ready and
user-facing. Most of the EAGLE3 model example YAMLs added in #1417 are
not yet verified to work end-to-end in modelopt (~80% fail at some
pipeline stage), which is confusing to ship. The agreed plan is to
**land the triage infrastructure now and re-add each model's launcher
YAML in a dedicated follow-up PR once it is verified green**.

**Removed** (unverified, to be re-added per-model once verified):
- Per-model launcher configs (`hf_offline_eagle3.yaml` +
`eagle3_quick_check.yaml`) for: DeepSeek-V3.2, GLM-5, MiniMax-M2.5,
Ministral-3-8B, Ministral-3-14B, Kimi-K2.5, Kimi-K2.5-NVFP4,
GPT-OSS-20B, Qwen3.5-9B, Qwen3.5-27B, Qwen3.5-35B-A3B, Step-3.5-Flash.
- Per-model status docs: `tools/launcher/examples/EAGLE3_TRIAGE.md`,
`examples/speculative_decoding/pipeline/eagle3/eagle3_triage_chart.md`
(volatile status — tracked internally instead).

**Kept** (the durable triage infrastructure from #1417):
- Verified baseline example
`tools/launcher/examples/Qwen/Qwen3-8B/eagle3_quick_check.yaml`.
- Launcher common scripts (vLLM native-extractor dump, etc.) and
`compute_hidden_states_vllm.py`.
- modelopt code fixes: FakeBaseModel VLM detection,
`consolidated.safetensors` load, `use_cache` export templates.
- New-model triage guide (`eagle3_new_model_triage_guide.md`); its
"document results" step now points at the internal tracker rather than
the removed chart.

### Testing

No code paths change — this only removes example YAMLs and two status
docs and edits one doc reference. Pre-commit
(ruff/markdownlint/yaml/license) passes on the kept/edited files.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (removes unverified examples
only; kept infra unchanged)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ❌ (pending)

### Additional Information

Follow-up to #1417. Next step (tracked separately): verify each removed
model end-to-end in modelopt, then re-add its YAML in a dedicated PR.
Note: the nmm-sandbox weekly EAGLE3 CI is being trimmed to the Qwen3-8B
baseline to match.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated EAGLE3 triage guide to streamline the verification workflow;
contributors now record test outcomes (status, experiment IDs, errors,
and applied fixes) in the team's internal triage tracker before
submitting model launcher configurations.

* **Chores**
* Removed legacy EAGLE3 example pipeline configurations and deprecated
triage documentation to reduce maintenance overhead.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-06-03 20:56:45 +00:00
yeyu-nvidiaandClaude Opus 4.6 a7b0a92047 EAGLE3 new model support: pipeline configs, triage docs, and Ministral-3 fixes (#1417)
## Summary

EAGLE3 automation triage work (OKR-30): testing the 4-step EAGLE3
offline pipeline against 12 new model architectures, documenting failure
modes, and fixing issues found.

### Code fixes (modelopt)

| File | Change |
|------|--------|
| `modelopt/torch/speculative/utils.py` | Extend VLM detection in
`load_vlm_or_llm` to check `text_config`/`llm_config` attrs (catches
`mistral3` models) |
| `modelopt/torch/speculative/plugins/modeling_fakebase.py` | Add
`consolidated.safetensors` fallback for checkpoints with incomplete HF
shards |
| `modelopt/torch/export/plugins/hf_spec_configs.py` | Set
`use_cache=True` in EAGLE export templates (fixes strict
`huggingface_hub` validation) |

### Pipeline infrastructure

- `examples/speculative_decoding/pipeline/eagle3/` — pipeline scripts
and configs:
- `offline_training.sh` — training + export with runtime patches for
older container modelopt
- `dump_offline_data_vllm.sh` — vLLM-based hidden state extraction (with
speculators compat patches)
- `dump_offline_data.sh`, `dump_offline_data_hf.sh` — alternative dump
paths
  - 18 quick-fail-check YAMLs for 12 models
  - 4 standalone task1 YAMLs

### Documentation

- `eagle3_triage_chart.md` — model test matrix, triage decision tree,
per-model results, failure catalog
- `eagle3_new_model_triage_guide.md` — step-by-step guide for triaging
new models

### Model test results (as of 2026-05-27)

| Model | task_0 | task_1 | task_2 | task_3 | Blocker |
|-------|--------|--------|--------|--------|---------|
| Qwen3-8B | - | - | - | - | Reference (existing) |
| Kimi-K2.5 | - | - | - | - | Existing (GB200) |
| **Ministral-3-8B** | SKIP | PASS | PASS | FAIL | `use_cache=null` in
export (fixed) |
| Ministral-3-14B | FAIL | - | - | FAIL | vLLM engine init fails |
| Qwen3.5-35B-A3B | TIMEOUT | - | - | - | Data synth too slow |
| gpt-oss-20b | FAIL | - | - | - | Tokenizer `HarmonyError` |
| Step-3.5-Flash | TIMEOUT | - | - | - | Data synth time limit |
| MiniMax-M2.5 | TIMEOUT | - | - | - | `trust_remote_code` needed |
| DeepSeek-V3.2 | no log | - | - | - | May not be mirrored |
| Qwen3.5-9B | - | - | - | - | Not yet run |
| Qwen3.5-27B | - | - | - | - | Not yet run |
| GLM-5 | - | - | - | - | Not yet run |

### Issues found and fixed

| # | Issue | Fix |
|---|-------|-----|
| 1 | `mistral3` model type not detected as VLM | Check
`text_config`/`llm_config` attrs in `load_vlm_or_llm` |
| 2 | Missing HF shard file (Ministral-3-8B) | Fallback to
`consolidated.safetensors` with Mistral native key aliases |
| 3 | `use_cache=null` in exported EAGLE config | Set `use_cache=True`
in export template configs |
| 4 | speculators incompatible with vLLM container | Runtime patches in
`dump_offline_data_vllm.sh` |
| 5 | `offline_training.sh` infra issues | Rewritten with runtime
patches for container modelopt |

## Test plan

- [x] Ministral-3-8B training passes (`cicd_1779829129`)
- [x] Ministral-3-8B export succeeds
- [ ] Ministral-3-8B benchmark passes (`cicd_1779901409` — pending with
all fixes)
- [ ] Dry-run remaining model configs

## Note

GitHub secret scanning alert #6 is a **false positive** —
`Mistral3ForConditionalGeneration` (a HuggingFace model class name in a
YAML comment) was flagged as a "Mistral AI API Key".

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-06-03 12:24:44 -07:00
yeyu-nvidia 651fd223e6 feat: EAGLE3 LoRA co-training improvements (#1607)
Layer-selective LoRA injection for EAGLE3 co-training, optimizer-stable
warmup (LoRA always in the optimizer; warmup gated by a flag), and an
export+merge+lm_eval evaluation script.

Review feedback addressed:
- trust_remote_code is caller-controlled (TRUST_REMOTE_CODE / --trust_remote_code), default False
- eagle_base_lora_start_layer raises ValueError instead of silently injecting zero adapters
- eval_lora.sh validates HF_MODEL_CKPT / EAGLE_CKPT up front
- added unit tests for start-layer injection and validation (8/8 passing, incl. on-cluster GPU run)

Signed-off-by: Ye Yu <yeyu@nvidia.com>
2026-06-03 17:33:12 +00:00
h-guo18 902d36921a [Feat]: Streaming Hidden-states Dataset (#1509)
### What does this PR do?

Type of change: new feature

**Design doc:**
https://gist.github.com/h-guo18/241c94968b0591324c361d97cf995dd0
Jira ticket: https://jirasw.nvidia.com/browse/OMNIML-4341
Sandbox CI:
https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/327911555#L2228

Streaming hidden-states dataset: per-sample activations pulled from a
live `vllm serve` over HTTP, replacing on-disk activation dumps.

Two axes for future extensions:
- **Backend** (`_fetch`): vLLM now; TRT-LLM / SGLang next.
- **Algorithm** (`_format`): Eagle now; distillation / probing next.
- **Sandbox CI**:
https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/merge_requests/169

Shared plumbing — async producer, token-level truncation to
`training_seq_len`, loss-mask alignment, DDP via Accelerate's dispatcher
(rank 0 fetches, broadcasts), circuit breaker, resume — lives in
`StreamingDataset`. First instance: **`EagleVllmStreamingDataset`**.

API: `data.mode ∈ {online, offline, streaming}`; legacy configs
auto-promote.

### Usage

```yaml
data:
  mode: streaming
  data_path: input_conversations/train.jsonl
  streaming_server_url: http://localhost:8000
  streaming_model_name: meta-llama/Llama-3.1-8B-Instruct
training:
  training_seq_len: 4096   # also caps the prompt sent to vllm
```

Requires `vllm serve` with `ExampleHiddenStatesConnector` and
`dataloader_num_workers=0`.

End-to-end Slurm pipeline:
`tools/launcher/examples/Qwen/Qwen3-8B/hf_streaming_eagle3.yaml`.

### Testing

- **Unit**: full-corpus invariant, rank-0-only iter, resume, circuit
breaker, mocked-httpx integration.
- **E2E**: `launch_train.sh` against a stdlib `HTTPServer` mimicking the
connector.
- **Smoke** (Qwen3-8B / 8×H100 / 4096 ultrachat samples, single epoch):
train_loss 32 → 18, MT-Bench AR 1.003 → 1.20.

### TODO before un-drafting

- [ ] Observability counters (filtered, fetch failures, queue depth,
latency).
- [ ] Changelog entry.
- [x] Add test in sandbox.

**Non-goals (v1):** multi-epoch streaming, cross-rank dynamic load
balancing.

### Before your PR is "*Ready for review*"

- Backward compatible: ✅ (legacy configs auto-promote)
- New PIP dep: N/A
- New tests: ✅
- Changelog: ❌ (TODO above)
- Claude approval: ❌


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Streaming training mode with server-backed hidden-state fetching,
deterministic seed control, resume support, and streaming-specific
dataset options (server, model, prefetch, shared storage).

* **Behavior / Bug Fixes**
* Stronger mode validation; offline behavior derived from data mode;
resume handling adjusted to avoid double-skip during streaming runs.

* **Tests**
* End-to-end CI streaming test and expanded unit tests covering
streaming, resume, DDP, determinism, and failure cases.

* **Infrastructure**
* Launcher script and pipeline config for end-to-end streaming training.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1509?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-06-02 14:56:23 -07:00
h-guo18 40a4dd326d [Feat]: Eagle Dry Run Mode (#1566)
### What does this PR do?

Type of change: new feature

Adds `--dry_run` to `examples/speculative_decoding/main.py`: load →
`mtsp.convert` → save, then exit (no `trainer.train()`). With
`FakeBaseModel`, the convert→save→export chain runs in seconds and
produces an exportable EAGLE3 / Medusa / DFlash checkpoint with correct
structure but untrained draft-head weights — useful for end-to-end
plumbing smoke tests on downstream stacks (vLLM, TRT-LLM, SGLang)
without paying for a real training run.

Where it sits among existing EAGLE3 modes:

```
EAGLE3 modes
├── online      base model runs forward in-loop
├── offline     reads pre-dumped hidden states from disk
├── streaming   streams hidden states from a live server in-loop
└── dry-run ★  skip training entirely; convert + save + export   ← NEW (this PR)
```

Companion `FakeBaseModel` fixes so small base checkpoints work:
- Synthesize the weight_map from a single `model.safetensors` when no
sharded index is present (Llama-3.2-1B, Qwen3-0.6B, …).
- Honor `tie_word_embeddings`: reuse `embed_tokens` when `lm_head` is
absent from safetensors.

A new launcher YAML
(`tools/launcher/examples/Qwen/Qwen3-8B/hf_eagle3_dryrun.yaml`) wires
this together as a one-task pipeline.

### Usage

```bash
# Direct
python main.py --dry_run \
  --config modelopt_recipes/general/speculative_decoding/eagle3.yaml \
  model.model_name_or_path=meta-llama/Llama-3.1-8B-Instruct \
  model.use_fake_base_for_offline=true \
  data.offline_data_path=/tmp/dryrun-placeholder \
  training.output_dir=ckpts/dryrun
python scripts/export_hf_checkpoint.py --model_path ckpts/dryrun --export_path export/dryrun

# Launcher
uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_eagle3_dryrun.yaml --yes
```

### Testing

- `tests/unit/torch/speculative/plugins/test_fakebase.py`: 3 new cases —
single-file fallback, tied-embeddings fallback, and the negative case
(missing `lm_head` without tying).
-
`tests/examples/speculative_decoding/test_eagle.py::test_eagle3_dry_run`:
full `launch_train.sh --dry_run → export_hf_checkpoint.py` chain on
`tiny_llama`; asserts exported state_dict has all
`LLAMA_EAGLE_SINGLE_LAYER` required keys.
- Manually verified end-to-end on Llama-3.1-8B-Instruct (sharded),
Llama-3.2-1B-Instruct (single-file + tied), and Qwen3-0.6B
(single-file).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `--dry_run` is opt-in;
`FakeBaseModel` changes are additive fallbacks.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — examples-only addition.
- Did you get Claude approval on this PR?: ❌

### Additional Information

N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a `--dry_run` CLI flag to perform a fast early-exit execution
that saves model artifacts without training.
* Added support for single-file model checkpoint formats in the loader.

* **Bug Fixes**
  * Improved checkpoint-loading error messages and handling.
  * Enhanced tied-embeddings fallback when head weights are absent.

* **Tests**
* Added integration and unit tests covering dry-run behavior and
single-file/tied-embedding loading.

* **Documentation**
  * Added a launcher example for dry-run smoke testing.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1566?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-05-29 15:02:26 -07:00
Shengliang Xu 9d0d97829a chore(lint): modernize typing (PEP 604/585) and enable UP032 (#1537)
### What does this PR do?

Type of change: chore / refactor (no behavior change)

Two small lint-cleanup commits:

**1. `chore(typing): modernize Union/Optional/List to PEP 604 / 585
syntax`** (8 files)

- Replace `X = Union[A, B] # noqa: UP007` with `X: TypeAlias = A | B`
for the six module-level type aliases (`ModelLike`, `Criterion`,
`NodeTarget`, `CalibrationDataType`, `Hparam.Importance` /
`ActiveSlice`). The `TypeAlias` annotation is required so mypy continues
to treat them as aliases under PEP 604.
- Modernize forward-ref unions in `modelopt/onnx/quantization/autotune/`
to full-string forward refs (e.g. `"RegionPattern | None"`).
- Update docstring type tags in
`examples/puzzletron/evaluation/hf_deployable_anymodel.py`.

**2. `chore(lint): remove UP032 ignore and convert .format() to
f-strings`** (10 files)

- Drop `UP032` from `extend-ignore` in `pyproject.toml`.
- Auto-convert 19 `"...".format(...)` calls to f-strings across export
plugins, examples, tests, and tools. One conversion in
`modelopt/torch/utils/plugins/megatron_generate.py` was wrapped manually
to stay under the 100-char limit.

**Intentionally left as-is:**

- `tools/launcher/slurm_config.py` keeps its `# ruff: noqa: UP045` —
nemo_run's CLI parser can't introspect PEP 604 optional annotations.
- `modelopt/torch/puzzletron/*` is **not** touched. The subtree disables
ruff's `UP` family entirely (per-file-ignore `"UP"`) while migration is
in progress, and converting `Optional[X]` to `X | None` there would
silently break runtime introspection in
`block_config._get_dataclass_type` that uses `get_origin(tp) is
typing.Union` (PEP 604 unions return `types.UnionType` from
`get_origin`, not `typing.Union`). Best revisited when puzzletron's lint
carve-out is narrowed.
- `UP038` (`isinstance(x, (int, float))` → `isinstance(x, int | float)`)
— ruff has officially deprecated this rule; PEP 604 in isinstance is
slightly slower and misleads readers about PEP 695 / `Optional`. Ignore
kept.

### Usage

No user-facing API changes.

### Testing

- Pre-commit hooks (ruff check, ruff format, mypy, bandit, license) pass
on both commits.
- Ruff status against `main`: 37 unrelated pre-existing findings
(W291/W293/E501/RUF005/PLR1704); zero new findings introduced by this
PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — Runtime behavior of the six
type aliases changes from a `typing.Union` instance to
`types.UnionType`. Downstream code introspecting via `get_origin(...) is
typing.Union` on these aliases would break, but no in-repo caller does
this on them. (The introspection in
`modelopt/torch/puzzletron/block_config.py` operates on user-supplied
dataclass field types, none of which are these aliases.)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (no behavior change)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — internal style refactor; happy to add a Misc note if reviewers want
one.
- Did you get Claude approval on this PR?: ❌ — not yet.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Style**
* Modernized type annotations across the codebase to use Python 3.10+
union syntax and TypeAlias where appropriate.
* Standardized string formatting to f-strings, improving clarity of
logs, errors, and validation messages.

* **Chores**
  * Updated linting configuration to reflect modern typing/style rules.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1537?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-05-26 16:47:31 -07:00
h-guo18 7038dec918 [1/2Refactor] speculative decoding: use mto config subsystem (#1328)
### What does this PR do?

Type of change: new feature

Port the speculative-decoding example to ModelOpt's recipe/config
subsystem: `model` / `data` / `training` / `<algo>` now load from a
single YAML with Pydantic validation and OmegaConf dotlist overrides.
Adds built-in `eagle3` / `dflash` recipes, drops the redundant
`training.mode` field (inferred from recipe class), and shrinks
`main.py` by ~145 lines (−208 / +63).

JIRA: OMNIML-3859

### Usage

```bash
python main.py --config general/speculative_decoding/eagle3 \
    model.model_name_or_path=meta-llama/Llama-3.2-1B \
    data.data_path=train.jsonl \
    training.output_dir=ckpts/test
```

### Testing

- `pytest tests/unit/recipe/test_loader.py` — new coverage for Eagle /
DFlash YAML loading, dotlist overrides, and field-level validation.
- Smoke-trained both built-in `eagle3` and `dflash` recipes end-to-end.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ — `main.py` CLI switched to
`--config <recipe>` (+ dotlist overrides); the old argparse flags are
removed.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
deps (`pydantic`, `omegaconf` already in core).
- Did you write any new necessary tests?: ✅ —
`tests/unit/recipe/test_loader.py`.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — to be added.

### Additional Information

Follow-up to the `modelopt.recipe` subsystem introduced for PTQ; this PR
extends the same declarative-YAML pattern to speculative decoding
(Eagle3 / DFlash / Medusa).


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added typed speculative-decoding recipe support for EAGLE, DFlash, and
Medusa; CLI dotlist overrides supported for single-file recipes.
* Trainer/config schema extended with speculative-training fields and
draft-vocab cache loading for Eagle.

* **Bug Fixes**
* Offline training no longer mutates model configs; loader enforces
required algorithm sections and prints recipe/config only on the primary
process.
* Reduced noisy per-rank logging by restricting status output to the
primary process.

* **Tests**
* Expanded tests for recipe loading, dotlist overrides, validation
strictness, and error cases.

* **Documentation**
  * Recipe YAMLs updated with metadata and usage notes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-05-16 17:29:24 -07:00
yeyu-nvidiaandClaude Sonnet 4.6 383ab4e224 fix: include medusa in data_module assignment in main.py (#1370)
## Problem
When `training.mode == "medusa"` is used in `main.py`, the `data_module`
variable is never assigned because line 344 only covered `eagle3` and
`dflash` modes. This causes an `UnboundLocalError` when the trainer is
constructed with `**data_module`.

Fixes OMNIML-4147

## Fix
Add `"medusa"` to the `training_args.mode in ("eagle3", "dflash")`
condition so `data_module` is correctly populated for medusa training.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed speculative decoding example to properly handle "medusa" mode
alongside existing "eagle3" and "dflash" modes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-04 15:17:50 +05:30
h-guo18 1ec931c2c7 [2/3][Feat]: Offline DFlash training (#1343)
### What does this PR do?

Type of change: new feature

Part 2 of a 3-PR series splitting #1271:
- **[1/3] #1296**: File reorg + deprecate `ParallelDraft`
- **[2/3] this PR**: Offline DFlash training (depends on #1296)
- **[3/3] #1297**: Extract `HFSpecDecMixin`

Changes:
- Add `dflash_offline` flag to `DFlashConfig` for training from
pre-computed hidden states; deletes base model layers to save memory.
- Add Pydantic validators on `DFlashConfig`:
- `_derive_dflash_offline` — auto-derive `dflash_offline` from
`data_args.offline_data_path` in validation context. Not
user-configurable: any user-supplied value is overridden by the derived
value.
- `_resolve_mask_token_id` — auto-detect `dflash_mask_token_id` from
`tokenizer.mask_token_id`.
  - `_check_mask_token_id` — fail fast if unset after resolution.
- `HFDFlashModel.modify()`: select `num_orig_hidden_layers` when
offline; pick `_base_model_lm_head` device when no base layers present;
drop base-model `layers` module.
- `HFDFlashModel.forward()`: add offline branch — consumes precomputed
`base_model_outputs` via `DFlashBaseModelOutput.from_offline_dict`, and
when `dflash_self_logit_distillation` is enabled with
`base_model_logits` absent, recomputes logits from
`base_model_hidden_states` via `_base_model_lm_head`. Raises a clear
error from the non-training / `pseudo_speculative_generate` paths when
`dflash_offline=True`, since base-model layers have been deleted.
- `DFlashBaseModelOutput` dataclass in `modeling_dflash.py` (with
`from_offline_dict` classmethod) to unify online/offline output shapes.
`aux_hidden_states` is required in `from_offline_dict` so missing keys
fail fast at the entry point rather than deeper in the forward.
- `examples/speculative_decoding/main.py`: replace inline
`mask_token_id` auto-detect with
`DFlashConfig.model_validate(dflash_cfg, context={"tokenizer":
tokenizer, "data_args": data_args})`.

### Silent bug fix — `add_generation_template` → `add_generation_prompt`

The pre-refactor `compute_hidden_states_hf.py` passed
`add_generation_template=False` to `tokenizer.apply_chat_template`. This
kwarg does not exist on HF `apply_chat_template` and was being silently
ignored, so the intended "don't append a generation prompt" behavior was
never actually applied. The new `tokenize_with_loss_mask` helper in
`examples/speculative_decoding/collect_hidden_states/common.py` uses the
correct `add_generation_prompt=False`. **This is a real behavior
change** for anyone re-dumping hidden states: trailing generation
prompts that were previously appended to the tokenized sequences will no
longer be included.


### Testing
- New tests:
- `tests/unit/torch/speculative/plugins/test_hf_dflash_offline.py` — CPU
unit tests for convert path (online keeps base layers, offline deletes
them; `num_orig_hidden_layers` drives `target_layer_ids` in offline
mode) and `DFlashConfig._derive_dflash_offline` validator.
- `TestDFlashOfflineForwardGPU` in
`tests/gpu/torch/speculative/plugins/test_hf_dflash.py` — GPU forward
smoke with precomputed `base_model_outputs`, plus the
`dflash_self_logit_distillation` logit-recompute path.

- training test:
<img width="454" height="317" alt="image"
src="https://github.com/user-attachments/assets/79b92790-4d15-4313-bb9b-f35665b012e6"
/> <img width="456" height="310" alt="image"
src="https://github.com/user-attachments/assets/4558559f-9c35-49ed-b36e-82fbc99eab23"
/>


### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ — additive `dflash_offline`
flag defaulting to `False`; validators fall through when context not
provided.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — see Testing section above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

### TODO (follow-up)

- [x] Update
`examples/speculative_decoding/collect_hidden_states/compute_hidden_states_*.py`
to support DFlash offline data. Current scripts are Eagle-specific —
they hardcode the `[2, N/2, N-3]` aux-layer selection and emit
`{input_ids, hidden_states, aux_hidden_states}`. DFlash offline needs:
- Aux layer indices driven by
`build_target_layer_ids(num_orig_hidden_layers, num_draft_layers)` (or a
configurable list), not the Eagle triplet.
- `base_model_hidden_states` key (last-layer hidden) so
`DFlashBaseModelOutput.from_offline_dict` + the
`dflash_self_logit_distillation` recompute path can consume it.
- Optional `base_model_logits` dump so offline training can skip the
self-distillation logit recomputation when logits are available.

### Additional Information

Base branch is #1296 (file reorg). Retarget to `main` once #1296 merges.



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Offline DFlash speculative-decoding training from precomputed
base-model hidden states
* Answer-only-loss training with persisted loss masks and optional
chat-template support
* Flexible auxiliary-layer selection via CLI and an exposed default
aux-layer helper
* Auto-derived offline flag in config and automatic memory optimization
during offline conversion

* **Documentation**
* Updated guides for offline pipeline, aux-layer selection, and
loss-masking options

* **Tests**
* New unit, GPU, and regression tests covering offline conversion,
training, and config derivation
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-04-25 18:38:35 -07:00
h-guo18 7c80d85751 [1/3][Refactor]: File reorg; deprecate ParallelDraft (#1296)
### What does this PR do?

Type of change: refactoring

Part 1 of a 3-PR series splitting #1271:
- **[1/3] this PR**: File reorg + deprecate `ParallelDraft`
- **[2/3] #1295**: Offline DFlash training
- **[3/3] #1297**: Extract `HFSpecDecMixin`

Changes:
- **File reorg**: `transformers.py` → `hf_eagle.py`; extract
`HFMedusaModel` → `hf_medusa.py`; extract `EagleModule` /
`EagleBaseModelOutput` → `modeling_eagle.py`; extract `DFlashModule` /
`DFlashAttention` / `DFlashDecoderLayer` / `build_target_layer_ids` /
`apply_rotary_pos_emb` → `modeling_dflash.py`.
- **Deprecate `ParallelDraft`**: remove `parallel_draft_step`,
`parallel_draft_heads_num_layers`, and the `ParallelDraft` module from
HF Eagle; remove the `EagleMedusaExporter` branch from
`HFEagleModel.get_exporter()` (the `EagleMedusaExporter` class itself
still lives in `hf_spec_export.py` for Megatron parity).
- **Rename**: `_draft_model_config` → `eagle_config` in export plugin.
- Update imports in `examples/speculative_decoding/` and
`modelopt/torch/speculative/utils.py` to follow the module rename.

### Testing

Validated with existing Eagle and DFlash training scripts (re-run after
`9ae5302729 revert behavior change`).

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ — renames
`modelopt.torch.speculative.plugins.transformers` → `.hf_eagle`; removes
`parallel_draft_step` / `parallel_draft_heads_num_layers` from Eagle
config; renames `_draft_model_config` → `eagle_config` in export plugin.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — pure refactor; existing
tests updated for the rename. `test_hf_spec_rope_export.py` assertions
were also corrected to reflect the actual production path (the old
assertions were masked by `MagicMock` not invoking the
`_draft_model_config` `@property`).
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌

### Additional Information

Breaking changes:
- `modelopt.torch.speculative.plugins.transformers` → `.hf_eagle`
- `parallel_draft_step` / `parallel_draft_heads_num_layers` removed from
Eagle config
- `_draft_model_config` → `eagle_config` in export plugin

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Refactoring**
* Reorganized speculative-decoding plugins into focused modules,
converting the legacy "transformers" entry into a deprecated shim that
re-exports the new plugin surface.
* Consolidated DFlash implementation into a shared modeling component
and introduced a dedicated EAGLE decoder module.

* **New Features**
* Added a Medusa speculative-decoding plugin with configurable heads and
combined-loss training behavior.

* **Chores**
  * Updated pre-commit license-hook exclusion and feature-flag wiring.

* **Tests**
  * Updated export tests to expect rope-scaling fallback semantics.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-04-24 14:46:10 -07:00
yeyu-nvidiaandClaude Opus 4.6 2fef374ded fix: auto-compute dp_replicate_size from world_size (#1302)
## Summary
- When `dp_shard_size < world_size` (e.g., `dp_shard_size=4` on 8 GPUs
across 2 nodes), `ParallelismConfig` raises `total_size (4) does not
match num_processes (8)` because `dp_replicate_size` defaults to 1
- Auto-compute `dp_replicate_size = world_size // (dp_shard_size *
cp_size)` so intra-node FSDP2 sharding + inter-node data-parallel
replication works without manual config
- This enables `dp_shard_size` to be set to per-node GPU count (better
NVLink utilization) while automatically creating replicas across nodes

## Test plan
- [ ] Verify single-node training (dp_shard_size == world_size,
dp_replicate_size == 1) unchanged
- [ ] Verify multi-node with dp_shard_size < world_size creates correct
replica groups
- [ ] Verify existing EAGLE3/DFlash configs still work

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Enhanced parallelism configuration initialization in the speculative
decoding example to better handle distributed training scenarios.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-20 20:39:36 +00:00
Chenhan D. YuandClaude Opus 4.6 355c6b7883 fix: PTQ 1GPU, export PP divisibility, hidden states conversations key (#1293)
## Summary
- **megatron_lm_ptq.yaml**: Qwen3-8B PTQ to single GPU for L40 clusters
(TP=1, all tasks)
- **quantize.sh**: Auto-find largest PP dividing model's
`num_hidden_layers` for export step. Qwen3-8B has 36 layers which isn't
divisible by 8, causing `AssertionError` on 8-GPU nodes
- **compute_hidden_states_trtllm.py**: Use `messages` with
`conversations` fallback, matching the HF version. Fixes `KeyError:
'conversations'` when data uses OpenAI `messages` format

## Test plan
- [x] Qwen3-8B PTQ runs on single L40 GPU
- [x] Export PP auto-selects valid divisor (36 layers → PP=6 on 8 GPUs,
PP=4 on 4 GPUs, PP=1 on 1 GPU)
- [x] EAGLE3 offline pipeline reads data with `messages` field

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Dataset input handling now supports multiple field formats for
enhanced compatibility.

* **Bug Fixes**
* Optimized GPU resource allocation during model quantization with
improved pipeline parallelism computation.
* Updated quantization configuration for more efficient resource
utilization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 12:43:58 -07:00
yeyu-nvidiaandClaude Opus 4.6 07ae8e7128 Add LoRA co-training support for HF EAGLE speculative decoding (#1060)
### What does this PR do?

Type of change: New feature + bug fixes

Adds **LoRA co-training** support for HF EAGLE speculative decoding.
When `eagle_base_lora=True`, HF PEFT LoRA adapters are injected into the
base model and co-trained alongside the EAGLE draft module in a single
online training pass. A preservation loss (KL divergence between the
original frozen base model output and the LoRA-adapted output) prevents
base model drift. LoRA adapter weights are exported in standard peft
format alongside EAGLE draft artifacts.

### Key features

- **LoRA injection**: `peft.inject_adapter_in_model` applied in-place
(no wrapper), keeping the existing `HFEagleModel` structure intact.
- **Preservation loss**: Cross-entropy `H(ref, lora)` — equivalent
gradient to `KL(ref || lora)` since `H(ref)` is constant w.r.t. LoRA
params.
- **Warmup schedule**: `eagle_base_lora_warmup_steps` freezes LoRA for N
steps while the EAGLE head stabilizes, then enables co-training via a
`LoRAWarmupCallback`.
- **Logits detach regularization**: `eagle_base_lora_logits_detach_prob`
stochastically detaches base logits from the EAGLE loss path, preventing
LoRA from degenerating to maximize EAGLE accuracy at the cost of base
model quality.
- **Export**: Standard peft format (`adapter_model.safetensors` +
`adapter_config.json`) alongside EAGLE draft model.
- **Merge script**: `scripts/merge_lora.py` merges LoRA weights into the
base model and restores the original `config.json` (avoids transformers
5.x rewriting `rope_theta` → `rope_parameters` which breaks
vLLM/TRT-LLM).
- **Multinode fix**: `dp_shard_size` now uses `WORLD_SIZE` instead of
local GPU count.

### Config options

```python
mtsp.convert(model, mode=[("eagle", {
    "eagle_base_lora": True,                          # enable LoRA co-training
    "eagle_base_lora_rank": 64,                       # LoRA rank
    "eagle_base_lora_alpha": 16.0,                    # LoRA scaling
    "eagle_base_lora_target_modules": ["q_proj", "k_proj", "v_proj", "o_proj"],
    "eagle_base_lora_preservation_loss_weight": 0.1,  # preservation loss weight
    "eagle_base_lora_warmup_steps": 0,                # freeze LoRA for N steps
    "eagle_base_lora_logits_detach_prob": 0.5,        # detach prob (0=never, 1=always)
})])
```

### Experimental results (Qwen3-8B, checkpoint-60000)

Base model quality preserved across detach_prob sweep (lm_eval: IFEval,
ARC-C, Winogrande — results pending final collection).

**Acceptance rate** (mt_bench, draft_length=3, output_length=4096,
temperature=0):

| detach_prob | vLLM AR | TRT-LLM AR |
|---|---|---|
| baseline (no LoRA) | 2.14 | 2.15 |
| 0.5 | 1.45 | 1.44 |
| 0.8 | **3.06** | **3.01** |
| 0.85 | 2.90 | 2.90 |
| 0.9 | 2.76 | 2.77 |
| 0.95 | 2.51 | 2.58 |
| 0.99 | 2.37 | 2.37 |
| 0.999 | 2.30 | 2.27 |
| 0.9999 | 2.31 | 2.26 |

Best AR at `detach_prob=0.8`: ~40% improvement over baseline.

### Testing

`tests/unit/torch/speculative/plugins/test_hf_speculative_lora.py` (5
tests):
- `test_lora_layers_injected` — LoRA layers present after conversion
- `test_trainable_params` — only `lora_*` and `eagle_module` params are
trainable
- `test_forward_returns_loss` — forward returns non-zero scalar loss
- `test_eagle_offline_incompatible` — `eagle_base_lora=True` +
`eagle_offline=True` raises `ValueError`
- `test_export_lora_artifacts` — export produces standard peft adapter
files

### Bug fixes (included in this PR)

1. **`launch_train.sh` case pattern ordering**: glob
`--eagle_base_lora*` was before specific patterns
(`--eagle_base_lora_rank*`, etc.), silently swallowing LoRA args.
2. **LoRA optimizer exclusion during warmup**: warmup freezing excluded
LoRA from the optimizer entirely; fixed with `add_param_group` in the
callback.
3. **`merge_lora.py` config.json**: `save_pretrained()` with
transformers >=5.x rewrites `rope_theta` → `rope_parameters`, breaking
vLLM positional embeddings. Fixed by copying the original base model
config.
4. **Multinode `dp_shard_size`**: used local GPU count instead of
`WORLD_SIZE`.

### Checklist

- [x] Backward compatible (all new config fields have defaults)
- [x] Uses `peft` via lazy imports (no hard dependency)
- [x] Unit tests added
- [x] Online HF training only (`eagle_offline=True` blocked)

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-16 01:35:15 +00:00
Chenhan D. YuandClaude Opus 4.6 3131195241 add: DFlash block diffusion speculative decoding (#1211)
DFlash (Block Diffusion for Flash Speculative Decoding) predicts an
entire block of tokens in a single forward pass using masked parallel
prediction with KV injection from the target model's hidden states.

Key features:
- Feature fusion (multi-layer hidden states -> FC + RMSNorm)
- KV injection (fused features as K/V in every draft layer with QK-norm)
- Random anchor sampling with bidirectional intra-block attention
- Logit distillation with exponential loss decay (gamma weighting)
- Multi-node DDP training with checkpoint resume
- Export to z-lab compatible HF format
- Online validation (context-dependent ground truth)

Training recipe:
modelopt_recipes/general/speculative_decoding/dflash.yaml
Results: examples/speculative_decoding/doc/dflash_results.md

### ModelOpt Eval (online validation, osl=512)

| Dataset | z-lab | ModelOpt (306K) | Diff |
|---------|-------|-----------------|------|
| gsm8k | 4.10 | **5.19** | **+1.09** |
| MT-Bench | 3.58 | **4.36** | **+0.78** |

### z-lab Official Eval (dflash.benchmark, osl=512)

| Dataset | z-lab | ModelOpt (306K) | Diff |
|---------|-------|-----------------|------|
| gsm8k | **5.00** | 4.08 | -0.92 |
| MT-Bench | **3.28** | 2.99 | -0.29 |

> z-lab model trained with block_size=16. ModelOpt trained with
block_size=8.

## Evaluation Method Impact (gsm8k)

| Eval Method | z-lab checkpoint | ModelOpt (306K) |
|-------------|-----------------|-----------------|
| Fixed GT (ModelOpt eval) | 2.95 | 4.23 |
| Online GT (ModelOpt eval) | 4.10 | **5.19** |
| z-lab official eval | **5.00** | 4.08 |

### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added DFlash speculative decoding mode with parallel block prediction
support.
* Included training launchers and MT-Bench evaluation scripts for DFlash
models.
* Added online acceptance rate validation for improved inference
verification.

* **Documentation**
* DFlash quick start guide with configuration parameters and training
examples.
  * Performance results and benchmarks for DFlash-trained models.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 15:58:39 -07:00
h-guo18 6403389eb0 Feat: Configurable Eagle ROPE scaling during export (#1238)
### What does this PR do?

JIRA ticket: https://jirasw.nvidia.com/browse/OMNIML-3469

Type of change: New feature

Decouple EAGLE training rope configuration from export rope
configuration, enabling separate YaRN rope scaling injection at export
time for long-context inference.

#### Changes

**Configurable export rope scaling (`EagleConfig`)**
- Add `eagle_export_rope_scaling` field to `EagleConfig` with default
YaRN config (`factor=32.0`, `original_max_position_embeddings=2048`)
- Set to `{}` to disable rope scaling injection at export

**Simplified training defaults (`default_config.py`)**
- Change default training rope from `llama3` (theta=500k) to `default`
(theta=10k) — models now train with simple positional embeddings; rope
scaling is applied only at export
- Add `rope_theta` inside `rope_scaling` dict for transformers 5.x
cross-version compatibility

**Move config validation/rewriting into `EagleConfig` (`config.py`)**
- `_derive_eagle_offline`: derives `eagle_offline` from
`data_args.offline_data_path` via validation context, removing manual
assignment in `main.py`
- `_check_rope_scaling_consistency`: rejects configs where
`eagle_export_rope_scaling` is set but training `rope_type` is not
`"default"`
- `_warn_rope_vs_training_seq_len`: warns when
`original_max_position_embeddings` differs from `training_seq_len`

**Export rope injection (`hf_spec_export.py`)**
- Inject `eagle_export_rope_scaling` into the exported HF config when
training rope_type is `"default"`
- Fall back `rope_theta` from `rope_scaling` dict for transformers 5.x
compatibility

**Fix Megatron RotaryEmbedding crash (`megatron_eagle.py`)**
- `dict_to_config()` set `rope_scaling=True` whenever the `rope_scaling`
key existed, even without a `"factor"` — causing `RotaryEmbedding` to
divide by `None`
- Now only enables `rope_scaling` when the dict actually contains a
`"factor"` key

### Usage

Configure in YAML config (or use defaults from `eagle3.yaml`):
```yaml
eagle:
  eagle_export_rope_scaling:
    rope_type: yarn
    factor: 32.0
    original_max_position_embeddings: 2048
```

Set to empty dict to disable export rope injection:
```yaml
eagle:
  eagle_export_rope_scaling: {}
```

### Testing

- New unit tests: `tests/unit/torch/speculative/test_eagle_config.py` —
rope consistency validator, seq_len warning, context-derived
`eagle_offline`
- New unit tests: `tests/unit/torch/export/test_hf_spec_rope_export.py`
— export rope injection, fallback, and empty-config cases

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ (new field has sensible
default; existing configs work unchanged)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ (should be added if merging as a feature)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Add export-time rope-scaling configuration for EAGLE models.

* **Improvements**
* Stronger validation and context-aware reconciliation between training
and export configs.
  * Export now injects rope-scaling and rope-theta when appropriate.
  * Default rope-scaling values updated for EAGLE variants.
  * Model instances now expose export rope-scaling for downstream use.

* **Tests**
* Added unit tests covering rope-scaling export behavior and
configuration validators.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-04-13 16:06:26 -07:00
yeyu-nvidiaandClaude Sonnet 4.6 9050188034 Fix test_collect_hidden_states: use synthetic short conversations (#1234)
## Summary

- `test_collect_hidden_states` was using real daring-anteater
conversations (typically 1000+ tokens) but the tiny test model has
`max_position_embeddings=32`. Both sampled conversations exceeded the
default `--max-seq-len 3072` filter, producing zero `.pt` files and
failing the assertion.
- Added a `tiny_conversations_path` fixture with synthetic short
single-turn conversations that tokenize within
`max_position_embeddings=32`.
- Changed `test_collect_hidden_states` to use this fixture with
`--max-seq-len 32`.
- Added a `None` guard for `tokenizer.chat_template.replace(...)` to
avoid `AttributeError` when the tokenizer has no chat template.

## Test plan
- [ ] `pytest
tests/examples/speculative_decoding/test_eagle_offline_ptq.py::test_collect_hidden_states`
passes
- [ ] CI `speculative_decoding` job passes

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Resolved compatibility issues when tokenizers do not have a chat
template configuration by adding proper error handling.
* Standardized tokenization input extraction logic across different
transformer library versions for consistent behavior.

* **Tests**
* Enhanced test infrastructure with new conversation data fixtures and
improved sequence length validation for speculative decoding examples.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-10 22:54:30 +00:00
yeyu-nvidiaandClaude Sonnet 4.6 c901814ff9 Fix compute_hidden_states_hf.py: handle BatchEncoding from apply_chat_template (#1225)
## Summary

- `apply_chat_template(..., return_tensors="pt")` returns a
`BatchEncoding` in transformers 4.46+, which no longer subclasses `dict`
- The old guard `isinstance(tokenized, dict)` evaluates to `False` for
`BatchEncoding`, so `input_ids` was set to the whole `BatchEncoding`
object
- Calling `.shape[1]` on a `BatchEncoding` triggers
`__getattr__("shape")` → `AttributeError`
- Fix: check `isinstance(tokenized, torch.Tensor)` instead, which
correctly handles both old transformers (plain tensor) and new
transformers (BatchEncoding)

This is causing `test_collect_hidden_states` to fail in the speculative
decoding CI for all open PRs (#1207, #1210, #1221).

## Test plan

- [ ] `torch-pr (speculative_decoding, 26.01)` CI passes
- [ ] Verify fix handles both `torch.Tensor` return (old transformers)
and `BatchEncoding` return (new transformers 4.46+)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-09 19:27:40 +00:00
Keval Morabia 04cd596d79 Add experimental support for transformers>=5.0 + min torch 2.8 (#975)
### What does this PR do?

- Add experimental support for transformers >=5.0 and remove deprecated
usages:
https://github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md
- ⚠️ For accelerate examples that used `--warmup-ratio: float`
(deprecated in 5.x), we now change it to `--warmup-steps: float | int`
which works as ratio if float but only for 5.x. For 4.x, it will error
out if float and prompt user to change back to `--warmup-ratio` or pass
an int absolute step count.
- ⚠️ Unified Hugging Face checkpoint export for quantized checkpoints
may not work for some models with transformers>=5.0 yet as it requires a
lot of fixes (e.g. change in how MoE experts are organized)
- ~Add Workaround for TRT-LLM's import of deprecated transformers
functions so trt-llm based gpu unit tests work fine. Still deployment
for models needs proper fixes directly in TRT-LLM hence llm/vlm ptq
example tests still run with transformers 4.57~
- Everything except PTQ and Export (mainly MoE) should work fine with
transformers>=5.0
- Bump min torch to 2.8 and enable 2.11 cicd testing
- NOTE: Upcoming Nemo:26.04 container comes with transformers 5.3

### Testing
<!-- Mention how have you tested your change if applicable. -->

- [x] CI/CD tests passing
- [x] Manually tested unit tests, gpu tests with transformers 4.56 and
5.4
- [x] Manually tested example tests (except trt-llm container tests)
with transformers 4.56 and 5.4
- [x] 2-gpu nightly CICD tests manually triggered and passing: [gpu
tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867257540),
[example
tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867260643)

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Make remote-code usage opt-in via a configurable --trust_remote_code
flag across examples and tools.

* **Bug Fixes**
* Improve checkpoint/resume detection and related training guidance to
avoid erroneous errors.

* **Refactor**
* Consolidate dtype/config naming, switch warmup settings from ratio →
steps, and unify tokenizer invocation patterns.

* **Documentation**
  * Simplify changelog title and add misc notes for release 0.44.

* **Chores**
* Remove scheduled PR-branch cleanup workflow and relax/remove several
transformers version pins.

* **Tests**
* Adjust test gates, skips, and structures to align with updated deps
and behaviors.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-09 09:59:37 +05:30
yeyu-nvidiaandClaude Sonnet 4.6 cccfded8a9 Add support for offline speculative decoding model PTQ (#883)
## What does this PR do?

**Type of change:**
new feature

**Overview:** 
This PR enables loading in a ModelOpt pretrained offline speculative
decoding model (e.g., EAGLE3) and performs PTQ on it and export.

## Usage
Follow the speculative_decoding examples to train an offline speculative
decoding model first.
Then follow the command below to quantize and export it:

```bash
python hf_ptq.py --pyt_ckpt_path <dir_of_offline_specdec_model> --specdec_offline_dataset <dir_of_dataset>
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Offline speculative decoding workflow: support loading a local dataset
for calibration, generation, and export; new CLI option to specify the
offline dataset.

* **Improvements**
* Export and quantization paths now accept and propagate offline
speculative-decoding inputs.
* Offline data loading honors a sample-size limit and enforces safe
batch sizing for calibration.

* **Bug Fixes**
* Better handling of model/config mismatches and varied batch types in
offline flows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-08 23:56:45 +00:00
Keval Morabia ba4f42df1c Minor fix for example tests
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-08 02:54:18 -07:00
Keval Morabia ebc534d765 Update code-copying guidelines in CONTRIBUTING.md
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-08 02:21:01 -07:00
Chenhan D. Yu 0246041b01 feat(speculative): add vLLM data synthesis pipeline and Nemotron dataset preparation scripts (#1176)
### What does this PR do?

Type of change: New feature, new example, bug fix

Adds a vLLM-based synthetic data generation pipeline for speculative
decoding draft model training, along with dataset preparation scripts
for NVIDIA's Nemotron Post-Training dataset collections.

**Data synthesis pipeline** (`tools/launcher/common/vllm/query.sh` +
`common/query.py`):
- Launch a vLLM server and run multi-turn inference to synthesize
training data from input conversation skeletons
- Fork-safe OpenAI client: reinitializes HTTP connection pool after
`datasets.map()` forks worker processes, preventing 400 errors from
corrupted connections
- Clear Docker `ENTRYPOINT` so vLLM containers (which default to `vllm
serve`) work correctly under NeMo Run's executor
- `--max-tokens` argument to bound generation length
- Local file loading support (`--data /path/to/file.jsonl`)
- Re-raise connection errors so `datasets.map()` halts the shard instead
of silently producing empty rows
- Map `developer` role to `system` (OpenAI format compatibility)

**Multi-turn reasoning trace handling** (`common/query.py`):
- Strip `<think>...</think>` blocks from intermediate assistant turns
before re-feeding to the model; preserve the full trace only on the
final turn

**Nemotron dataset preparation** (`examples/dataset/`):
- `make_nemotron_ptv2_dataset.py` — prepares
[nvidia/Nemotron-Post-Training-Dataset-v2](https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2)
(~3.3M rows generate, ~1.9M rows train)
- `make_nemotron_ptv3_dataset.py` — prepares the [Nemotron PTv3
collection](https://huggingface.co/collections/nvidia/nemotron-post-training-v3)
of 16 datasets (~3.4M rows generate, ~3.9M rows train)
- Both support `generate` mode (strips assistant turns for synthesis
input) and `train` mode (normalizes to clean OpenAI format for SFT)
- `conversation_utils.py` — shared utilities: `strip_assistant_turns`,
`normalize_messages`, `make_augment_fn`, `AugmentationSpec`
- `augmentations.yaml` — 12 language-redirect variants + style/format
hints, cycled across dataset rows
- Scripts live in `examples/dataset/` (not under
`speculative_decoding/`) to signal reusability beyond speculative
decoding

**Bug fixes**:
- `strip_assistant_turns()`: return `{"messages": []}` when no user
turns remain (system-only rows were previously passed through instead of
being filtered)
- `concatenate_datasets()`: guard against empty parts list
- SSH tunnel user precedence: explicit `user` arg now correctly
overrides `slurm_config.user`

### Usage

```bash
# Prepare PTv3 input conversations for synthesis (~3.4M rows):
python examples/dataset/make_nemotron_ptv3_dataset.py --output-dir /tmp/ptv3_gen

# Launch vLLM server + synthesize responses:
bash tools/launcher/common/vllm/query.sh \
    --model /path/to/model \
    --tensor-parallel-size 4 \
    -- \
    --data /tmp/ptv3_gen/default.jsonl \
    --save /tmp/ptv3_responses \
    --num-shards 10 --num-proc 4 --max-tokens 4096

# Prepare PTv2 for direct SFT training (~1.9M rows):
python examples/dataset/make_nemotron_ptv2_dataset.py --mode train --output-dir /tmp/ptv2_train
```

### Testing

Tested end-to-end on an NVIDIA GB10 node (119 GiB GPU memory) with
`vllm/vllm-openai:qwen3_5-cu130` container and `Qwen/Qwen3.5-4B`:
- vLLM server starts correctly with cleared Docker entrypoint
- `datasets.map(num_proc=4)` runs without connection errors (fork-safe
client)
- Multi-turn synthesis produces correct assistant responses with
thinking traces handled

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (data synthesis scripts;
tested manually)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added dataset generation and augmentation capabilities for Nemotron
post-training datasets (v2 and v3)
* Enhanced query functionality with thinking-block filtering and
improved client management for robust parallel processing
* Added support for local dataset file paths alongside HuggingFace Hub
datasets

* **Bug Fixes**
* Fixed SLURM executor user resolution and Docker container entrypoint
configuration
* Improved error handling for connection failures during dataset
synthesis

* **Documentation**
* Updated dataset preparation guide with new generation modes and
augmentation configuration details

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan D. Yu <5185878+ChenhanYu@users.noreply.github.com>
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-04-08 05:22:11 +00:00
h-guo18 82d96a635f [Speculative Decoding] Refactor EAGLE3 training to YAML-based config and recipe system (#1134)
## What does this PR do?

Refactors EAGLE3 training to use a single base YAML config with
OmegaConf dotlist overrides.

**Type of change:** Refactor

## Changes

- Single base config
`modelopt_recipes/speculative_decoding/_base_eagle3.yaml` for all EAGLE3
training; removed per-model child YAMLs.
- `launch_train.sh` accepts `--config <yaml>` plus dotlist overrides
(e.g. `model.model_name_or_path=xxx`).
- Removed `__base__` YAML inheritance logic from `main.py`.
- `dp_shard_size` default changed from `0` sentinel to `None` for
clarity.
- Removed `eagle_config.json` and `fsdp_config.json`; architecture
config is now nested under `eagle.eagle_architecture_config` in YAML.
- `train_eagle3_and_export.sh` now uses base YAML + dotlist instead of
generating a temporary YAML.
- Updated README and tests accordingly.

## Usage

```bash
# Online training
./launch_train.sh \
    --config ../../modelopt_recipes/speculative_decoding/_base_eagle3.yaml \
    model.model_name_or_path=meta-llama/Llama-3.2-1B \
    data.data_path=input_conversations/train.jsonl \
    training.output_dir=ckpts/llama-3.2-1b-online

# Offline training
./launch_train.sh \
    --config ../../modelopt_recipes/speculative_decoding/_base_eagle3.yaml \
    model.model_name_or_path=meta-llama/Llama-3.2-1B \
    data.offline_data_path=$HIDDEN_STATES_DIR \
    training.output_dir=ckpts/llama-3.2-1b-offline
```

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-04-08 01:52:13 +00:00
Keval MorabiaandRinZ27 5dc17dfd15 [Security] Enable torch.load(weights_only=True) for secure checkpoint loading + trust_remote_code fix (#1181)
### What does this PR do?

- Add secure checkpoint loading support using
`torch.serialization.add_safe_globals([cls])`. This also removes 1
existing pickle usage.
- Remove hard-coded `trust_remote_code=True`
- Replaces https://github.com/NVIDIA/Model-Optimizer/pull/1056 by
@RinZ27

### Testing
<!-- Mention how have you tested your change if applicable. -->

CICD tests ran

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information

NVBug: 5999336

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added safe checkpoint save/load helpers and a --trust_remote_code CLI
flag in examples to control remote-code loading.

* **Bug Fixes**
* Checkpoint loading now defaults to safer, weights-only semantics to
reduce arbitrary-code exposure.

* **Documentation**
* CHANGELOG updated with security guidance and opt-in procedure for
unsafe checkpoint loading.

* **Tests**
  * New unit tests validating the safe-load behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: RinZ27 <222222878+RinZ27@users.noreply.github.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: RinZ27 <222222878+RinZ27@users.noreply.github.com>
2026-04-08 00:36:28 +05:30
Keval Morabiaandh-guo18 80d2f02a2d Fix spec dec example tests (#1183)
### What does this PR do?

Type of change: Test fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

- Fix `tests/examples/speculative_decoding` - previously silently
skipped
- Avoid pulling nemotron-post-training-dataset-v2 in tests to reduce
chances of HF loading timeout in CICD
- Make slow and redundant tests manual to speed up CICD

### Testing
<!-- Mention how have you tested your change if applicable. -->

- Tests passing

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Removed git‑LFS install step from CI and deleted an automated
branch‑cleanup workflow
* Trimmed example environment dependencies and relaxed transformers
compatibility; added an optional tokenization dependency

* **Tests**
* Switched tests to generate datasets dynamically and improved fixture
handling
* Standardized PTQ test parameters (explicit calibration dataset) and
refined GPU/test selection

* **Bug Fixes**
* Improved device-awareness and numeric handling in speculative decoding
attention paths
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-04-06 22:23:50 -07:00
h-guo18 7f5fd65003 [Feat]FakeBaseModel for offline eagle; Kimi-K2.5 fixes; (#1052)
### What does this PR do?

Adds `FakeBaseModel` for offline EAGLE training and several Kimi-K2.5
compatibility fixes.

- **New**: `FakeBaseModel` — lightweight model that loads only `lm_head`
and `embed_tokens` from a local checkpoint, avoiding full model weight
loading during offline training. Configured via `FakeBaseArguments` and
integrated into `load_vlm_or_llm`.
- **Fix**: `_find_base_model_parts` — support Kimi-K2.5 VLM layout
(`language_model.model` path)
- **Fix**: offline mode lm_head access and CompressedTensors ignore path
- **Fix**: Kimi-K2.5 decoder `past_key_value`/`past_key_values` argument
mismatch
- **Fix**: `rglob` for `.pt` discovery in nested offline data dirs;
single-node GPU count respects `CUDA_VISIBLE_DEVICES`

Type of change: Bug fix, new feature

### Testing
Tested offline EAGLE training for Kimi-K2.5 end-to-end.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌

### Additional Information
<!-- E.g. related issue. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Lightweight fake-base model support for offline speculative-decoding
training

* **Improvements**
* Added CLI flags: --use_fake_base_for_offline, --trust_remote_code, and
--fsdp
  * Expanded offline .pt discovery to include nested subdirectories
* Better GPU detection with explicit single-node logging; FSDP enabled
only when requested
* Model loading and launch tooling now honor offline and
trust-remote-code flags

* **Bug Fixes**
* Improved compatibility with legacy transformer / Kimi-K2 call
signatures

* **Tests**
* Added tests covering fake-base loading and offline training workflows
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-03-27 16:57:13 -07:00
Keval Morabia d0bf0bef96 Remove deprecated Nemo 2.0 references / examples (#1098)
### What does this PR do?

- Remove `examples/nemo_run` and other deprecated Nemo 2.0 references
- Add Megatron-Bridge example links where missing

<!-- Details about the change. -->

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!---
Mandatory -->
- Did you write any new necessary tests?: N/A <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Deprecations**
* Removed deprecated NeMo 2.0 support and related example flows and
utilities.

* **Documentation**
* Updated docs and examples to emphasize Megatron-Bridge / Megatron-LM
and refreshed technique/deployment guidance and links.

* **New Features**
* Added CLI options for additional parallelism (context/expert
tensor/expert model) in Megatron-Bridge distillation.

* **Chores**
* Removed legacy CI configs and refreshed container image tags across
examples and docs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-24 14:11:13 +05:30
Benjamin Chislettandh-guo18 547d8fe935 [EAGLE] Optimize EAGLE Training (#1044)
### What does this PR do?

Type of change: Optimization

<!-- Details about the change. -->

Changes:
- Precompute base_model_logits.argmax() and base_model_logits.softmax()
instead of recomputing in every call to _eagle_loss
- Calculate per-prediction accuracy on the GPU and synchronize it to the
host after running all TTT steps, to avoid cpu/gpu synchronization
inside the TTT step loop.
- Apply torch.compile to performance-critical training functions:
prepare inputs, eagle forward, and eagle loss calculation. Omitted from
target model in online case for now, as it may not be natively
compatible with all architectures.

### Usage

No changes to external interfaces

### Testing

Ran training commands for benchmarking. Did not do a full training run,
did not validate correctness.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!---
Mandatory -->
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added runtime profiling ranges for speculative decoding to enable
finer performance tracing.

* **Improvements**
* Lazy initialization for rotary embeddings in llama-style decoders for
more reliable startup.
* Speculative decoding now uses base-model predicted tokens and softmax
probabilities for sampling and loss, improving stability and accuracy
reporting.
* Parallelism configuration is now conditional, avoiding unnecessary
setup unless required.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-03-20 18:36:06 -07:00
Benjamin Chislett c76633ac9d [EAGLE] Configurable number of TTT steps (#1042)
### What does this PR do?

Type of change: new CLI option for existing option

<!-- Details about the change. -->

- Added num_ttt_steps CLI flag
- Changed num_ttt_steps default from 4 to 3 for consistency.
Num_spec_tokens == 3 or == 7 are most common in practice, so rounding
down to 3 and allowing users to increment higher on-demand. Will also
improve training efficiency for the OOTB experience.

### Usage

Users can now pass `--num_ttt_steps 7` to `launch_train.sh` when
training an EAGLE3 model for extended speculation lengths.

### Testing

N/A

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added ability to configure train-time-test steps for speculative
decoding training via command-line argument.
  * Updated default train-time-test steps value from 4 to 3.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
2026-03-18 09:32:17 -07:00
Benjamin Chislett 4292505512 Refactor: Clean up EAGLE training dataset preparation (#684)
## What does this PR do?

**Type of change:** Refactor

**Overview:** 
- Consolidate input dataset preparation into `make_dataset.py`
- Read dataset mix spec from a YAML file
- - Can now specify how many samples to take from each split
- - Can no longer easily split a dataset into train/test sections. I
don't think this feature was really useful to begin with. Most datasets
can already be separated into train/val/test at the split level, and
those that can't are usually going to be splitted by the training FW
anyways.
- Add support for a few new dataset types, magpie 300k/500k/1M, nemotron
post-training dataset v2.

## Usage
See README for detailed example

## Testing
Ran it locally on all dataset modes, works successfully and output looks
good. Checked shuffling, conversation IDs, and output contents were all
unique and usable.

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated speculative decoding example documentation with new dataset
references and standardized file paths.

* **New Features**
* Introduced configuration-driven dataset preparation supporting
multiple dataset sources with centralized configuration files.

* **Refactor**
* Simplified dataset preparation workflow with unified tooling and
updated default data paths throughout the training pipeline.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
2026-03-18 09:29:49 -07:00
yeyu-nvidia 5d0e012751 inplement mix hidden_states for eagle3; deprecate eagle1 (#946)
## What does this PR do?

new feature

**Overview:** 
Enable mix hidden_states in eagle3 training. Deprecate eagle1

## Usage
Add --mix_hidden_states True to launch_train.sh

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added --mix_hidden_states option to enable optional hidden-state
mixing during training.
* Added eagle_ttt_steps setting to control speculative multi-step
iterations.

* **Chores**
* Consolidated speculative decoding to EAGLE3 only; legacy Medusa/EAGLE1
paths removed.
* Unified configuration handling so models and plugins accept a single
config object.

* **Tests**
* Updated and expanded tests for hidden-state mixing and EAGLE3-only
scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
2026-03-09 11:14:10 -07:00
skierat dd16a96fda example demonstrating how to train CosmosReason2 Eagle3 (#965)
## What does this PR do?

**Type of change:** New example

**Overview:**
Adds
examples/speculative_decoding/guides/train_eagle_head_cosmos_reason2.ipynb,
a step-by-step Jupyter notebook that walks through the full EAGLE3
draft-head training workflow for nvidia/Cosmos-Reason2-8B. The notebook
covers:

1. Installing dependencies
2. Authenticating with Hugging Face
3. Preparing training data from the Nemotron-Post-Training-Dataset-v2
(chat split) using a curated row-selection mapping
(guides/nemotron_mapping.csv)
4. Inspecting the bundled EAGLE3 config (guides/CR2_eagle_config.json)
tuned for Cosmos-Reason2 (YaRN RoPE, FlexAttention, reduced draft
vocabulary)
5. (Optional) Calibrating the draft vocabulary to 32k tokens for faster
training and inference
6. Launching training via launch_train.sh with FSDP2 multi-GPU support
7. Exporting the checkpoint to HF format and serving with vLLM

Also includes guides/nemotron_mapping.csv and
guides/CR2_eagle_config.json as companion files.

## Usage
Open and run
examples/speculative_decoding/guides/train_eagle_head_cosmos_reason2.ipynb
cell by cell. After training, serve the exported checkpoint with:
```
vllm serve nvidia/Cosmos-Reason2-8B \    --host 0.0.0.0 \    --port 8000 \    --speculative-model export/cosmos-reason2-8b-eagle3 \    --num-speculative-tokens 3 \    --dtype bfloat16
```

## Testing
Tested end-to-end on a 4xB100 GPUs. The exported checkpoint was
validated with specdec_bench

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: Yes (the
notebook is self-documenting)
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

## Additional Information
Cosmos-Reason2-8B requires at least one 80 GB GPU (H100/A100)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a speculative-decoding configuration for draft model and
rotary/attention behavior.
* Added an end-to-end training notebook demonstrating training, export,
and deployment of a speculative-decoding draft head on Cosmos-Reason2.
* Added a data-preparation tool to download, normalize, and convert
Nemotron chat conversations into a standardized conversation format.

* **Documentation**
* Notebook documents environment setup, data prep, training/validation
cadence, export, and deployment steps.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Slawek Kierat <skierat@nvidia.com>
2026-03-05 23:53:49 +00:00
h-guo18 a34d613d3c Feat: Speculatice Decoding export with quantization support (#913)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** 


Main changes:  
- Refactored speculative decoding export logics into `class
EagleExporter` to improve cohesion;

 
- Separated speculative decoding export entrance with quantization
export (`export_hf_checkpoint()`) due to their fundamental differences:
- Quantization export base model's state_dict and config, while
speculative decoding only export drafter's.
- Most of the model-specific logics of quantization export (e.g.
diffusers, vlms) are not needed for speculative decoding export.
- Quantization export produce different format than speculative decoding
checkpoint. (The former produce tokenizer config, generation config,
e.t.c, while the later does not need. )

## Usage
<!-- You can potentially add a usage example below. -->

To export an regular bf16 eagle checkpoint without quantization, the
commands are the same:
```python
python scripts/export_hf_checkpoint.py --model_path <x> --export_path <x>
```

To run PTQ on online-trained eagle checkpoint and export it:
```python
python hf_ptq.py --pyt_ckpt_path <x> --qformat fp8 --export_path <x>
```

The above two commands will produce drafter ckpt for deployment, in the
same foramt.

## Testing
<!-- Mention how have you tested your change if applicable. -->

Tested setting:
- Base model: llama3.1-8b
- Algorithms: eagle
- Export path tested: 
- (Unquantized online ckpt) `python scripts/export_hf_checkpoint.py
--model_path <x> --export_path <x>`
- (PTQ) export `python hf_ptq.py --pyt_ckpt_path <x> --qformat fp8
--export_path <x>`
- Tested deployment on vllm. Got normal AR. 

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
  * Added export functionality for speculative decoding-optimized models
* Support for multiple speculative decoding architectures with
pre-configured deployment templates
* Enhanced model export detection and automatic routing for optimized
models

* **Tests**
  * Updated export validation tests for speculative decoding models

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-03-04 01:46:02 +00:00
Chenhan D. Yu edde087bbd Adding a special list to handle HF models that require (#950)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. --> Bug fix

**Overview:** ?

1. Some models check `layer_types` and our offline overwrites
`num_hidden_layers=0` which will requires passing `layer_types=[]`. It
is also possible that the model will raise unexpected arguments when
passing `layer_types=[]`. So the current WAR is to use a list to handle
by checking the architectures.
2. Fix relative path issue in `launch_train.sh`



## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
  * Enhanced model loading robustness for specific layer configurations.

* **Chores**
* Improved example script path handling and GPU initialization tracking.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-03-03 00:07:10 +00:00
yeyu-nvidia 4a11486f6c Enable multinode training for HF speculative decoding (#922)
## What does this PR do?

**Type of change:** 
New example

**Overview:** 
Modify launch_train.sh script to enable multi-node training.
Provide a slurm template script.

## Usage
Add required fields in slurm.sh.
Then use the command below to submit multi-node job:

```bash
bash slurm.sh
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added multi-node training support with new configuration options for
distributed setups.
* Introduced Slurm batch script for streamlined job submission to
cluster environments.

* **Improvements**
* Enhanced GPU resource management for multi-GPU and distributed
training configurations.
  * Updated speculative decoding model support and validation handling.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
2026-02-26 18:23:43 +00:00
h-guo18 eb99488da1 Fix: restore requires_grad in transformers5 reloading (#907)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** 

Patch transformers 5.x parameter loading to preserve original
`requires_grad` settings.

In transformers v5.x, loading a checkpoint forcibly sets parameters'
requires_grad,
which unintentionally unfreeze frozen parameters (e.g. Base model in
eagle training).

This leads to optimizer initialization error since the restored
optimizer expected more parameter than the checkpoint.

This monkey-patch restores the original`requires_grad` after loading
parameters.

Reference:

https://github.com/huggingface/transformers/blob/v5.0.0.rc1-release/src/transformers/core_model_loading.py#L640

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed model parameter loading in speculative decoding to properly
preserve gradient requirements for each parameter when using HuggingFace
Transformers 5.x, ensuring correct behavior during checkpoint resumption
and model initialization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-02-19 01:25:47 +00:00
h-guo18 b8a4586702 Refactor: Eagle data loading (#668)
## What does this PR do?

**Type of change:** Refactor <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** 
Jira ticket: https://jirasw.nvidia.com/browse/OMNIML-2955

Main changes :

- Consolidate Eagle data loading with @ChenhanYu 's implementation of
`transformers_dataset.py`

- Refactor: baked the following logics from `example/main.py` to
`modelopt/torch` for cleaner example entrance:
  - default config selecting and merging with custom config
  - tokenizer post-processor (chat template and pad_tok_id)
  - d2t loading 
- Implementation refactor: In HF workflow, reuse base modfel's input
hidden states as input_embedding, instead of calculating from input_ids.
This has two main benefits:
    - Easier VLM support, which has various embedding processing logics.
    - Training effieicy.
    
- Deprecating eagle1 from the example. It is still available by setting
custom config.
  
 - Other minor fixes and readme updates.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

Tested that training curves after changes (both online&offline) is
identical with original branch:
<img width="1073" height="634" alt="image"
src="https://github.com/user-attachments/assets/abfd7bea-c82c-48a7-8181-68c5a9e4da8d"
/>


## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added draft vocabulary cache support for EAGLE model training,
enabling runtime vocabulary customization via `--draft_vocab_cache`
parameter
* Introduced new data loading utilities with sharding, streaming, and
tokenization support for large-scale training
  * Added optional `--log_steps` configuration to training launcher

* **Documentation**
* Updated EAGLE configuration guides with draft vocabulary cache setup
instructions and examples

* **Refactor**
* Restructured data pipeline for offline training with improved dataset
handling and batching
* Updated command-line arguments across training scripts (`--input-data`
replaces `--input-file`)

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-02-18 14:04:06 -08:00
yeyu-nvidia a8f5314c93 fix the path change in torch v2.10 for spec dec (#863)
## What does this PR do?

**Type of change:** 
bug fix

**Overview:** 
torch v2.10 changes the path for _SDPAMerger. will need to use the new
path for import

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Updated internal import references to reflect organizational changes
in dependencies.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
2026-02-09 17:56:53 +00:00
yeyu-nvidia 81b67ddf06 Context parallelism for Megatron core models (#818)
## What does this PR do?

New feature

**Overview:** 
This PR implements the context manager which injects attn_mask as
attn_bias to TEDotProductAttention so that we can enable EAGLE training
with arbitrary mask.

## Usage
set CP>1 in
https://github.com/NVIDIA/Megatron-LM/blob/main/examples/post_training/modelopt/finetune.sh

```python
# Add a code snippet demonstrating how to use this
```

## Testing
Tested on DSR1 Llama 8B.
CP1->CP2
38854MB->28050MB
MTbench AL 2.26->2.31

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## New Features
* Context parallelism support added for Eagle speculative decoding with
HuggingFace and Megatron Core models.
* Model checkpoint loading enhanced to enable remote code execution
capabilities when required.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
2026-01-29 19:47:22 +00:00
h-guo18 3036a9ea9f Feat: Context Parallel for Eagle3 Training (#745)
## What does this PR do?

**Type of change:** New Feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

- Supported Context Parallel by patching torch ring attention;
- Require following libirary version for stable cp: 
  - torch2.8.0
  - transformers5.0.0
  - accelrate1.12.0 
 - Move to FSDP2 
- Removed unused arguments in training script (`--multi_gpu`,
`fsdp_wrap_layer`)
 - Bump CI container to `nvcr.io/nvidia/pytorch:25.08-py3`

## Usage
<!-- You can potentially add a usage example below. -->

```bash
./launch_train.sh --model $MODEL \
            --output_dir $OUTPUT_DIR \  
            --data $DATA \
            --num_epochs 0.1 \
            --train_bs 1 \
            --eagle_config eagle_config.json \
            --training_seq_len 1024 \
            --cp_size 2   #newly added
```

## Testing
- SDPA level correctness: tested TTT attention with/without CP, diff <
1%
```
=== Compare context-parallel (CP) outputs and grads with non-CP ===
Forward output comparison (CP vs Non-CP):
  Absolute diff (adiff) cp_out vs out: 0.001953125
  Relative diff (rdiff) cp_out vs out: 0.00182342529296875
WQ (query proj) grad comparison (CP vs Non-CP):
  Absolute diff (adiff) cp_wq_grad vs wq_grad: 0.0078125
  Relative diff (rdiff) cp_wq_grad vs wq_grad: 0.00347900390625
WK (key proj) grad comparison (CP vs Non-CP):
  Absolute diff (adiff) cp_wk_grad vs wk_grad: 0.0078125
  Relative diff (rdiff) cp_wk_grad vs wk_grad: 0.002471923828125
WV (value proj) grad comparison (CP vs Non-CP):
  Absolute diff (adiff) cp_wv_grad vs wv_grad: 0.25
  Relative diff (rdiff) cp_wv_grad vs wv_grad: 0.0069580078125
==============================================================
```

- E2E Training Acc
  (Llama3.1-8B, Unsynthesized magpie)
<img width="911" height="630" alt="image"
src="https://github.com/user-attachments/assets/1ecacc7f-c720-494c-9c1b-b60e7ced7baa"
/>

- Peak Mem Reserved
   (llama3.1-8B, 8xH100, train_length=4k)

    | cp_size | max_memory_allocated(MB) |max_memory_reserved (MB) |
    |----|--------------------------|--------------------------|
    | 1  |         65040.20              |79018.00
    | 2  |           50409.17             |73098.00
    | 4  |              45120.92            |72052.00
    | 8  |              38882.12            |66484.00

- Max Training Length test
  (llama3.1-8B, H100)

  | cp_size               | 6k  | 12k | 24k  | 48k  |
  |--------------------|-----|-----|-----|-----|
  | 1         | ✅ | OOM | OOM | OOM  |
  |2      | ✅  | ✅ | OOM | OOM |
  | 4         | ✅ | ✅ | ✅  | OOM  |
  | 8         | ✅ | ✅ | ✅  | ✅  |

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added context parallelism (CP) and data parallelism shard size
configuration parameters to training arguments.

* **Enhancements**
* Improved TTT attention masking support for speculative decoding
workflows.
* Enhanced training launch script with improved parallelism
configuration handling.

* **Chores**
* Updated core dependencies: torch, transformers, accelerate, and wandb.
  * Added FSDP configuration file for distributed training setup.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-01-24 02:45:50 +00:00
h-guo18 3fd8b804f4 Fix:add vllm dump script (#710)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** ?
Missing one file in last
PR:https://github.com/NVIDIA/Model-Optimizer/pull/689/changes
## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2025-12-19 12:30:04 -08:00
Ben Shoham Ofir 2a51bbd7cb Optimize calibrate_draft_vocab to read only required lines when calib… (#618)
Optimize calibrate_draft_vocab to read only required lines when
calibrate_size is set

## What does this PR do?

**Type of change:** Performance improvement

**Overview:** 
This PR optimizes the
[calibrate_draft_vocab.py](cci:7://file:///Users/obenshoham/PycharmProjects/TensorRT-Model-Optimizer/examples/speculative_decoding/scripts/calibrate_draft_vocab.py:0:0-0:0)
script to improve memory efficiency and I/O performance when using the
`--calibrate_size` parameter. Previously, the script would read all
lines from the data file into memory before slicing to the specified
`calibrate_size`, causing unnecessary resource usage for large datasets.
The optimization uses `itertools.islice` to read only the required
number of lines when `calibrate_size` is specified.

## Usage

The script usage remains unchanged. When using `--calibrate_size`, the
script now only reads the specified number of lines instead of loading
the entire dataset:

```bash
# Only reads first 1000 lines from the dataset (optimized)
python scripts/calibrate_draft_vocab.py \
    --model meta-llama/Llama-3.2-1B-Instruct \
    --data input_conversations/daring-anteater.jsonl \
    --draft_vocab_size 32000 \
    --calibrate_size 1000 \
    --save_dir draft_vocab_cache

Signed-off-by: Ofir Ben Shoham <ofir_benshoham@intuit.com>
2025-12-19 17:47:08 +00:00
h-guo18andyeyu-nvidia bdd10c2dbe Feat: MLA eagle (#689)
## What does this PR do?

**Type of change:** New Feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

- Add MLA Eagle support
- Add new argument "eagle_decoder_type" to switch between llama and
kimik2 eagle;
  - Add patches to load from kimik2 model implementations dynamically;
  - new default config for kimi k2;
- Refactor eagle export to support multilayer/multitype eagle export
concisely;
     - Rename some modules for simplified export logic;
- Other minor improvements;

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

- Tested that kimi k2 thinking works with eagle_type=kimik2:
<img width="1068" height="636" alt="image"
src="https://github.com/user-attachments/assets/5557ef87-c719-4fb1-be18-30435f6b3885"
/>

- Tested that llama 3.2 1b works with eagle_type=llama:
<img width="1066" height="634" alt="image"
src="https://github.com/user-attachments/assets/633c575c-cc79-43af-aed3-0378a303ebc7"
/>



## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Signed-off-by: yeyu-nvidia <yeyu@nvidia.com>
Co-authored-by: yeyu-nvidia <yeyu@nvidia.com>
2025-12-19 00:11:37 +00:00