mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
33bfa8b1fe47d1a86a1a421e24d84a6acd53aad2
86
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b6bf6b7997 |
Update Roadmap Issue link
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
48f8d8984f |
Fix conversation loading logic in UltraChat dataset (#1680)
### What does this PR do?
Type of change: Bug fix.
<!-- Details about the change. -->
Previously, only the first user prompt was extracted from each example,
discarding all subsequent turns. UltraChat stores full multi-turn
conversations in the "messages" field, so switching to that field
preserves the complete dialogue rather than truncating to a single user
message.
### Usage
```python
# Add a code snippet demonstrating how to use this
```
### Testing
python make_dataset.py -f test_cfg.yaml --full
python make_dataset.py -f test_cfg.yaml
```yaml
- name: "ultrachat"
splits:
train_gen: 100
train_sft: 100
```
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Breaking Changes**
* Removed UltraChat support from the example dataset mixer, so UltraChat
splits can no longer be loaded.
* **Documentation**
* Updated the dataset examples README and the example dataset
configuration to remove UltraChat and adjust split settings for other
datasets.
* Updated the speculative decoding fine-tuning documentation to
reference Daring-Anteater instead of UltraChat.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: jzh26 <226629529+jzh26@users.noreply.github.com>
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
9048d13b86 |
[Feat]:Support DPace (#1724)
### What does this PR do? Type of change: New feature Adds the **D-PACE** (Dynamic Position-Aware Cross-Entropy) loss objective for DFlash speculative-decoding training ([arXiv:2605.18810](https://arxiv.org/abs/2605.18810)). It replaces the static exponential position decay with per-position CE weights derived from the draft's own confidence `q_i = exp(-CE_i)`: smoothed `q̃_i = (1-α)q_i + α` (Eq.7) and weighted by the suffix-sum of prefix products `w_j = Σ_{m≥j} ∏_{i≤m} q̃_i` (Eq.8), which directly targets expected accepted block length and shifts signal toward whichever positions currently limit acceptance. Selected via `dflash_loss_objective` — **D-PACE is now the default** (`dpace`); set `dflash_loss_objective: decay` to restore the previous static schedule. Smoothing via `dflash_dpace_alpha` (default 0.5). Weights are detached from the gradient — training-only, ~2.3% overhead, no architecture or inference change. Mutually exclusive with `dflash_loss_decay_factor`. ### Usage ```yaml # DFlash recipe / training config dflash: dflash_loss_objective: dpace # default: decay dflash_dpace_alpha: 0.5 # smoothing in (0, 1]; stable in [0.3, 0.7] ``` ### Testing CPU unit tests in `tests/unit/torch/speculative/plugins/test_hf_dflash.py`: weights match the paper closed form, are detached and non-increasing, the α smoothing floor keeps later weights non-zero, and convert wires/validates the new fields (rejects bad objective and degenerate α). Training validated on Qwen3-8B (curve below). <img width="1803" height="809" alt="image" src="https://github.com/user-attachments/assets/d34dcd76-9e46-4051-94d4-c880b1987965" /> ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ⚠️ Behavior change — D-PACE is now the **default** objective, so DFlash training loss weighting changes unless you set `dflash_loss_objective=decay` (which reproduces the previous static-decay behavior). - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new dependency) - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Reference: D-PACE, [arXiv:2605.18810](https://arxiv.org/abs/2605.18810). See `examples/speculative_decoding/doc/dflash.md` for the math and tuning notes. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added a **D-PACE** training loss objective for DFlash speculative decoding (`dflash_loss_objective: dpace`), configurable via `dflash_dpace_alpha` (default `0.5`). * **Documentation** * Documented D-PACE’s confidence-derived, dynamically weighted per-position loss behavior (training-only) and noted that `dflash_loss_decay_factor` is ignored with D-PACE. * **Bug Fixes** * Updated DFlash loss to reuse the precomputed per-token cross-entropy in the non-KD path. * **Tests** * Added unit tests for D-PACE weight correctness, masking, gradient detachment, monotonicity, smoothing, and config/validation behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
977d34dc3c |
[Fix](nvbug6304585): specdec README online base-model example should use Instruct model (#1755)
### What does this PR do?
Type of change: Bug fix
The **Training Draft Model with Online/Offline Base Model** examples in
`examples/speculative_decoding/README.md` used
`meta-llama/Llama-3.2-1B`, a
base / pretrained checkpoint that ships **no chat template**. The online
EAGLE3
flow tokenizes conversations through
`tokenizer.apply_chat_template(...)`, so
the data collator fails fast at startup:
```
ValueError: No valid chat template!
```
This PR:
- Switches both README example commands (online and offline) to
`meta-llama/Llama-3.2-1B-Instruct`, which carries a chat template.
- Makes the collator error message in
`modelopt/torch/utils/plugins/transformers_dataset.py` actionable — it
now
explains the cause (base checkpoints have no chat template) and points
users
at an Instruct model or a custom `chat_template`.
### Usage
```bash
./launch_train.sh \
--config ../../modelopt_recipes/general/speculative_decoding/eagle3.yaml \
model.model_name_or_path=meta-llama/Llama-3.2-1B-Instruct \
data.data_path=input_conversations/train.jsonl \
training.output_dir=ckpts/llama-3.2-1b-online
```
### Testing
- Reproduced the original `No valid chat template!` failure with the
base
`Llama-3.2-1B` and confirmed the Instruct variant carries a chat
template.
- Verified the new error message renders correctly when a template is
missing.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌
### Additional Information
Fixes nvbug 6304585: https://nvbugspro.nvidia.com/bug/6304585
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **Documentation**
* Updated speculative decoding example training commands to reference
the Llama-3.2-1B-Instruct model.
* **Bug Fixes**
* Enhanced error message when chat template configuration is missing,
providing actionable guidance for resolution.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
e6790ef7b4 |
[Examples]: GPT-oss, Qwen3Moe streaming specdec example (#1692)
### What does this PR do?
Type of change: new example
Adds **streaming speculative-decoding examples (EAGLE3 + DFlash)** for
**gpt-oss-20b** and **Qwen3-30B-A3B** to the ModelOpt launcher,
mirroring the existing Qwen3-8B/Kimi examples.
- New yamls:
`tools/launcher/examples/{openai/gpt-oss-20b,Qwen/Qwen3-30B-A3B}/hf_streaming_{eagle3,dflash}_multi_node.yaml`,
plus gpt-oss `chat_template_train.jinja` (generation-tagged, for
`answer_only_loss`).
- `eagle_utils.py`: the streaming path now installs a custom
`data.chat_template` on the tokenizer (the online path already did) —
needed for the tagged template.
### Usage
```bash
cd tools/launcher
export SLURM_HOST=... SLURM_ACCOUNT=... SLURM_HF_LOCAL=... SLURM_JOB_DIR=...
uv run launch.py --yaml examples/openai/gpt-oss-20b/hf_streaming_eagle3_multi_node.yaml --yes
```
### Testing
Pipeline sanity test on **unsynthesized** data (daring-anteater), 1×
H100-80GB, 12k steps. All four train and pass the vLLM acceptance-length
eval:
| Model | Method | Train speed | vLLM AL |
|---|---|---|---|
| Qwen3-30B-A3B | EAGLE3 | 7.12 it/s | **1.74** |
| Qwen3-30B-A3B | DFlash | 2.31 it/s | 1.29 |
| gpt-oss-20b | EAGLE3 | 5.07 it/s | 1.19 |
| gpt-oss-20b | DFlash | 2.01 it/s | 1.14 |
<img width="1300" height="780" alt="image"
src="https://github.com/user-attachments/assets/2be22562-5b77-4a9f-9dc0-6f936a059736"
/>
> Sanity test only, not a quality run. gpt-oss AL is low because it is a
reasoning model (CoT at inference) while daring-anteater has no
reasoning traces and `answer_only_loss` masks all but the final content
— quality runs need synthesized/reasoning data.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added multi-node speculative decoding pipeline configurations for
Qwen3-30B-A3B and gpt-oss-20b with DFlash and EAGLE3 support.
* Introduced chat template training support for improved model
instruction formatting.
* **Enhancements**
* Increased benchmark concurrency from 1 to 32 across Qwen3-8B
configurations for more realistic performance evaluation.
* Extended training runs from 500 to 2000 steps for Kimi-K2.5 models.
* Improved chat template handling in speculative decoding workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
e004d8d90e |
DFlash speculative decoding for MiniMax-M2.7 (FSDP2): auto mask-token, FSDP2 resume fixes, per-checkpoint draft export (#1621)
## What
Brings up DFlash block-diffusion speculative decoding for large MoE
targets (MiniMax-M2.7, 229B) trained under accelerate FSDP2, and fixes
the regressions that broke checkpoint resume and per-checkpoint draft
export.
## Commits
- **auto-add mask token for DFlash** when the tokenizer lacks one
(resize embeddings, restore dtype).
- **requeue support** in `build_slurm_executor` + **FSDP2
cpu_ram_efficient_loading** for 229B on multi-node.
- **FSDP2 buffer patch** (`fsdp2_buffer_patch.py`): handle non-DTensor
buffers in `fsdp2_load_full_state_dict`, broadcast dtype codes from rank
0, and an FSDP2-safe `clip_grad_norm_`. Required because MiniMax-M2.7
pins transformers 4.57.x (no native `ParallelismConfig`).
- **dtype fix**: use the broadcast dtype (rank 0) rather than the local
meta-device param dtype, so non-leader ranks don't cast bf16 back to
fp32 on resume.
- **restore `DFlashExportCallback`** (this PR's headline): the
Pydantic-recipe refactor (
|
||
|
|
46eddab877 |
[Feat]: Specdec Streaming: RDMA + Multinode (#1611)
### What does this PR do? Type of change: New feature Multi-node **streaming** training for speculative decoding (EAGLE3 / DFlash): a live `vllm serve` captures the target model's hidden states and moves them straight to the trainer over **NIXL RDMA** — no disk round-trip. The streaming dataset is map-style — each rank fetches only its own `DistributedSampler` shard (concurrency from `dataloader_num_workers`), round-robins across multiple serve replicas (`server_urls`), and scales to multi-node DDP. Serve-side tensor parallelism (TP>1) is supported: hidden states are replicated across TP ranks, so rank 0 alone owns the pool + transfer. ### How - `RdmaHiddenStatesConnector` — out-of-tree vLLM connector (no vLLM source edits): one pre-registered pinned NIXL pool per serve, a ring slot per request, and a small HTTP sidecar serving transfer metadata. The trainer RDMA-READs the slot into a per-worker buffer. RDMA is the **only** transport (the earlier disk/safetensors path is removed). - Map-style dataset + multi-node accelerate launch (`--machine_rank`, optional Slurm `--segment` to keep nodes in one NVLink domain). ### Usage ```yaml data: mode: streaming streaming_server_url: "http://node0:8000,http://node1:8000" # round-robin ``` ### Validation (Qwen3-8B, oci-nrt H100) sandbox CI: https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/337489812 **1. End-to-end convergence — EAGLE3 & DFlash, 5000 steps.** Both algorithms converge and export a deployable draft; the DFlash drafts also serve under vLLM speculative decoding (8/8 smoke prompts pass). | algorithm | topology (nodes) | train loss (step 0 → 5000) | vLLM draft acc-len | |---|---|---|---| | EAGLE3 | 2 serve TP=2 + 2 trainer DDP (4) | 37.1 → 8.20 | — | | DFlash | 1 serve TP=1 + 1 trainer (2) | 11.7 → 5.56 | 1.11 | | DFlash | 2 serve TP=2 + 2 trainer DDP (4) | 10.9 → 5.26 | 1.19 | <!-- Drag these PNGs in here (GitHub turns them into asset URLs): eagle3_streaming_loss.png, dflash_streaming_loss_singlenode.png, dflash_streaming_loss_multinode.png --> **2. Scalability — 1 → 12 nodes (EAGLE3, 200 steps).** Throughput scales ~23× across the sweep below. The step-time growth is cross-node DDP all-reduce, not the streaming path — RDMA (~0.33 ms/req @ 2 MB, ~47 GB/s host-pinned READ) is never the bottleneck. Scale serve + trainer nodes together for near-linear speedup. | serve / trainer | nodes | step time | samples / step | samples / sec (global) | acc @ step 200 | |---|---|---|---|---|---| | 1 serve / 1 rank (co-located, 1 node 2 GPU) | 1 | 0.23 s | 1 | 4.4 | [0.141, 0.094, 0.072] | | 1 serve / 1 rank (cross-node) | 2 | 0.23 s | 1 | 4.3 | [0.137, 0.105, 0.074] | | 2 serve / 8 ranks | 3 | 0.26 s | 8 | 31.1 | [0.215, 0.126, 0.097] | | 4 serve / 16 ranks (2 trainer nodes) | 6 | 0.28 s | 16 | 56.5 | [0.217, 0.148, 0.110] | | 8 serve / 32 ranks (4 trainer nodes) | 12 | 0.31 s | 32 | 101.9 | [0.235, 0.165, 0.137] | **3. Serve-side TP correctness.** TP=1 vs TP=2 draft top-1 accuracy track step-for-step (hidden states are replicated across TP ranks). <img width="910" height="546" alt="serve-tp-acc" src="https://github.com/user-attachments/assets/73df9214-7ff0-4ab4-bf2f-95842b12cd5f" /> ### Before your PR is "*Ready for review*" - Backward compatible?: ❌ — streaming is now RDMA-only; `server_url` → `server_urls`; the disk transport (`HS_TRANSPORT`, `streaming_shared_storage_path`) is removed. - New tests?: ✅ `tests/unit/torch/speculative/plugins/test_hf_streaming_dataset.py` (map-style dataset + mocked RDMA fetch). - Updated Changelog?: ❌ - Claude approval?: ❌ --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
5bd04c3876 |
Revert unverified EAGLE3 model examples; keep triage code + baseline (#1623)
### What does this PR do? Type of change: Revert / cleanup (follow-up to #1417) Per review feedback (@h-guo18): `main` should be production-ready and user-facing. Most of the EAGLE3 model example YAMLs added in #1417 are not yet verified to work end-to-end in modelopt (~80% fail at some pipeline stage), which is confusing to ship. The agreed plan is to **land the triage infrastructure now and re-add each model's launcher YAML in a dedicated follow-up PR once it is verified green**. **Removed** (unverified, to be re-added per-model once verified): - Per-model launcher configs (`hf_offline_eagle3.yaml` + `eagle3_quick_check.yaml`) for: DeepSeek-V3.2, GLM-5, MiniMax-M2.5, Ministral-3-8B, Ministral-3-14B, Kimi-K2.5, Kimi-K2.5-NVFP4, GPT-OSS-20B, Qwen3.5-9B, Qwen3.5-27B, Qwen3.5-35B-A3B, Step-3.5-Flash. - Per-model status docs: `tools/launcher/examples/EAGLE3_TRIAGE.md`, `examples/speculative_decoding/pipeline/eagle3/eagle3_triage_chart.md` (volatile status — tracked internally instead). **Kept** (the durable triage infrastructure from #1417): - Verified baseline example `tools/launcher/examples/Qwen/Qwen3-8B/eagle3_quick_check.yaml`. - Launcher common scripts (vLLM native-extractor dump, etc.) and `compute_hidden_states_vllm.py`. - modelopt code fixes: FakeBaseModel VLM detection, `consolidated.safetensors` load, `use_cache` export templates. - New-model triage guide (`eagle3_new_model_triage_guide.md`); its "document results" step now points at the internal tracker rather than the removed chart. ### Testing No code paths change — this only removes example YAMLs and two status docs and edits one doc reference. Pre-commit (ruff/markdownlint/yaml/license) passes on the kept/edited files. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (removes unverified examples only; kept infra unchanged) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: ❌ (pending) ### Additional Information Follow-up to #1417. Next step (tracked separately): verify each removed model end-to-end in modelopt, then re-add its YAML in a dedicated PR. Note: the nmm-sandbox weekly EAGLE3 CI is being trimmed to the Qwen3-8B baseline to match. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated EAGLE3 triage guide to streamline the verification workflow; contributors now record test outcomes (status, experiment IDs, errors, and applied fixes) in the team's internal triage tracker before submitting model launcher configurations. * **Chores** * Removed legacy EAGLE3 example pipeline configurations and deprecated triage documentation to reduce maintenance overhead. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
a7b0a92047 |
EAGLE3 new model support: pipeline configs, triage docs, and Ministral-3 fixes (#1417)
## Summary EAGLE3 automation triage work (OKR-30): testing the 4-step EAGLE3 offline pipeline against 12 new model architectures, documenting failure modes, and fixing issues found. ### Code fixes (modelopt) | File | Change | |------|--------| | `modelopt/torch/speculative/utils.py` | Extend VLM detection in `load_vlm_or_llm` to check `text_config`/`llm_config` attrs (catches `mistral3` models) | | `modelopt/torch/speculative/plugins/modeling_fakebase.py` | Add `consolidated.safetensors` fallback for checkpoints with incomplete HF shards | | `modelopt/torch/export/plugins/hf_spec_configs.py` | Set `use_cache=True` in EAGLE export templates (fixes strict `huggingface_hub` validation) | ### Pipeline infrastructure - `examples/speculative_decoding/pipeline/eagle3/` — pipeline scripts and configs: - `offline_training.sh` — training + export with runtime patches for older container modelopt - `dump_offline_data_vllm.sh` — vLLM-based hidden state extraction (with speculators compat patches) - `dump_offline_data.sh`, `dump_offline_data_hf.sh` — alternative dump paths - 18 quick-fail-check YAMLs for 12 models - 4 standalone task1 YAMLs ### Documentation - `eagle3_triage_chart.md` — model test matrix, triage decision tree, per-model results, failure catalog - `eagle3_new_model_triage_guide.md` — step-by-step guide for triaging new models ### Model test results (as of 2026-05-27) | Model | task_0 | task_1 | task_2 | task_3 | Blocker | |-------|--------|--------|--------|--------|---------| | Qwen3-8B | - | - | - | - | Reference (existing) | | Kimi-K2.5 | - | - | - | - | Existing (GB200) | | **Ministral-3-8B** | SKIP | PASS | PASS | FAIL | `use_cache=null` in export (fixed) | | Ministral-3-14B | FAIL | - | - | FAIL | vLLM engine init fails | | Qwen3.5-35B-A3B | TIMEOUT | - | - | - | Data synth too slow | | gpt-oss-20b | FAIL | - | - | - | Tokenizer `HarmonyError` | | Step-3.5-Flash | TIMEOUT | - | - | - | Data synth time limit | | MiniMax-M2.5 | TIMEOUT | - | - | - | `trust_remote_code` needed | | DeepSeek-V3.2 | no log | - | - | - | May not be mirrored | | Qwen3.5-9B | - | - | - | - | Not yet run | | Qwen3.5-27B | - | - | - | - | Not yet run | | GLM-5 | - | - | - | - | Not yet run | ### Issues found and fixed | # | Issue | Fix | |---|-------|-----| | 1 | `mistral3` model type not detected as VLM | Check `text_config`/`llm_config` attrs in `load_vlm_or_llm` | | 2 | Missing HF shard file (Ministral-3-8B) | Fallback to `consolidated.safetensors` with Mistral native key aliases | | 3 | `use_cache=null` in exported EAGLE config | Set `use_cache=True` in export template configs | | 4 | speculators incompatible with vLLM container | Runtime patches in `dump_offline_data_vllm.sh` | | 5 | `offline_training.sh` infra issues | Rewritten with runtime patches for container modelopt | ## Test plan - [x] Ministral-3-8B training passes (`cicd_1779829129`) - [x] Ministral-3-8B export succeeds - [ ] Ministral-3-8B benchmark passes (`cicd_1779901409` — pending with all fixes) - [ ] Dry-run remaining model configs ## Note GitHub secret scanning alert #6 is a **false positive** — `Mistral3ForConditionalGeneration` (a HuggingFace model class name in a YAML comment) was flagged as a "Mistral AI API Key". 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
651fd223e6 |
feat: EAGLE3 LoRA co-training improvements (#1607)
Layer-selective LoRA injection for EAGLE3 co-training, optimizer-stable warmup (LoRA always in the optimizer; warmup gated by a flag), and an export+merge+lm_eval evaluation script. Review feedback addressed: - trust_remote_code is caller-controlled (TRUST_REMOTE_CODE / --trust_remote_code), default False - eagle_base_lora_start_layer raises ValueError instead of silently injecting zero adapters - eval_lora.sh validates HF_MODEL_CKPT / EAGLE_CKPT up front - added unit tests for start-layer injection and validation (8/8 passing, incl. on-cluster GPU run) Signed-off-by: Ye Yu <yeyu@nvidia.com> |
||
|
|
902d36921a |
[Feat]: Streaming Hidden-states Dataset (#1509)
### What does this PR do? Type of change: new feature **Design doc:** https://gist.github.com/h-guo18/241c94968b0591324c361d97cf995dd0 Jira ticket: https://jirasw.nvidia.com/browse/OMNIML-4341 Sandbox CI: https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/327911555#L2228 Streaming hidden-states dataset: per-sample activations pulled from a live `vllm serve` over HTTP, replacing on-disk activation dumps. Two axes for future extensions: - **Backend** (`_fetch`): vLLM now; TRT-LLM / SGLang next. - **Algorithm** (`_format`): Eagle now; distillation / probing next. - **Sandbox CI**: https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/merge_requests/169 Shared plumbing — async producer, token-level truncation to `training_seq_len`, loss-mask alignment, DDP via Accelerate's dispatcher (rank 0 fetches, broadcasts), circuit breaker, resume — lives in `StreamingDataset`. First instance: **`EagleVllmStreamingDataset`**. API: `data.mode ∈ {online, offline, streaming}`; legacy configs auto-promote. ### Usage ```yaml data: mode: streaming data_path: input_conversations/train.jsonl streaming_server_url: http://localhost:8000 streaming_model_name: meta-llama/Llama-3.1-8B-Instruct training: training_seq_len: 4096 # also caps the prompt sent to vllm ``` Requires `vllm serve` with `ExampleHiddenStatesConnector` and `dataloader_num_workers=0`. End-to-end Slurm pipeline: `tools/launcher/examples/Qwen/Qwen3-8B/hf_streaming_eagle3.yaml`. ### Testing - **Unit**: full-corpus invariant, rank-0-only iter, resume, circuit breaker, mocked-httpx integration. - **E2E**: `launch_train.sh` against a stdlib `HTTPServer` mimicking the connector. - **Smoke** (Qwen3-8B / 8×H100 / 4096 ultrachat samples, single epoch): train_loss 32 → 18, MT-Bench AR 1.003 → 1.20. ### TODO before un-drafting - [ ] Observability counters (filtered, fetch failures, queue depth, latency). - [ ] Changelog entry. - [x] Add test in sandbox. **Non-goals (v1):** multi-epoch streaming, cross-rank dynamic load balancing. ### Before your PR is "*Ready for review*" - Backward compatible: ✅ (legacy configs auto-promote) - New PIP dep: N/A - New tests: ✅ - Changelog: ❌ (TODO above) - Claude approval: ❌ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Streaming training mode with server-backed hidden-state fetching, deterministic seed control, resume support, and streaming-specific dataset options (server, model, prefetch, shared storage). * **Behavior / Bug Fixes** * Stronger mode validation; offline behavior derived from data mode; resume handling adjusted to avoid double-skip during streaming runs. * **Tests** * End-to-end CI streaming test and expanded unit tests covering streaming, resume, DDP, determinism, and failure cases. * **Infrastructure** * Launcher script and pipeline config for end-to-end streaming training. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1509?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
40a4dd326d |
[Feat]: Eagle Dry Run Mode (#1566)
### What does this PR do? Type of change: new feature Adds `--dry_run` to `examples/speculative_decoding/main.py`: load → `mtsp.convert` → save, then exit (no `trainer.train()`). With `FakeBaseModel`, the convert→save→export chain runs in seconds and produces an exportable EAGLE3 / Medusa / DFlash checkpoint with correct structure but untrained draft-head weights — useful for end-to-end plumbing smoke tests on downstream stacks (vLLM, TRT-LLM, SGLang) without paying for a real training run. Where it sits among existing EAGLE3 modes: ``` EAGLE3 modes ├── online base model runs forward in-loop ├── offline reads pre-dumped hidden states from disk ├── streaming streams hidden states from a live server in-loop └── dry-run ★ skip training entirely; convert + save + export ← NEW (this PR) ``` Companion `FakeBaseModel` fixes so small base checkpoints work: - Synthesize the weight_map from a single `model.safetensors` when no sharded index is present (Llama-3.2-1B, Qwen3-0.6B, …). - Honor `tie_word_embeddings`: reuse `embed_tokens` when `lm_head` is absent from safetensors. A new launcher YAML (`tools/launcher/examples/Qwen/Qwen3-8B/hf_eagle3_dryrun.yaml`) wires this together as a one-task pipeline. ### Usage ```bash # Direct python main.py --dry_run \ --config modelopt_recipes/general/speculative_decoding/eagle3.yaml \ model.model_name_or_path=meta-llama/Llama-3.1-8B-Instruct \ model.use_fake_base_for_offline=true \ data.offline_data_path=/tmp/dryrun-placeholder \ training.output_dir=ckpts/dryrun python scripts/export_hf_checkpoint.py --model_path ckpts/dryrun --export_path export/dryrun # Launcher uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_eagle3_dryrun.yaml --yes ``` ### Testing - `tests/unit/torch/speculative/plugins/test_fakebase.py`: 3 new cases — single-file fallback, tied-embeddings fallback, and the negative case (missing `lm_head` without tying). - `tests/examples/speculative_decoding/test_eagle.py::test_eagle3_dry_run`: full `launch_train.sh --dry_run → export_hf_checkpoint.py` chain on `tiny_llama`; asserts exported state_dict has all `LLAMA_EAGLE_SINGLE_LAYER` required keys. - Manually verified end-to-end on Llama-3.1-8B-Instruct (sharded), Llama-3.2-1B-Instruct (single-file + tied), and Qwen3-0.6B (single-file). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `--dry_run` is opt-in; `FakeBaseModel` changes are additive fallbacks. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — examples-only addition. - Did you get Claude approval on this PR?: ❌ ### Additional Information N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a `--dry_run` CLI flag to perform a fast early-exit execution that saves model artifacts without training. * Added support for single-file model checkpoint formats in the loader. * **Bug Fixes** * Improved checkpoint-loading error messages and handling. * Enhanced tied-embeddings fallback when head weights are absent. * **Tests** * Added integration and unit tests covering dry-run behavior and single-file/tied-embedding loading. * **Documentation** * Added a launcher example for dry-run smoke testing. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1566?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
9d0d97829a |
chore(lint): modernize typing (PEP 604/585) and enable UP032 (#1537)
### What does this PR do? Type of change: chore / refactor (no behavior change) Two small lint-cleanup commits: **1. `chore(typing): modernize Union/Optional/List to PEP 604 / 585 syntax`** (8 files) - Replace `X = Union[A, B] # noqa: UP007` with `X: TypeAlias = A | B` for the six module-level type aliases (`ModelLike`, `Criterion`, `NodeTarget`, `CalibrationDataType`, `Hparam.Importance` / `ActiveSlice`). The `TypeAlias` annotation is required so mypy continues to treat them as aliases under PEP 604. - Modernize forward-ref unions in `modelopt/onnx/quantization/autotune/` to full-string forward refs (e.g. `"RegionPattern | None"`). - Update docstring type tags in `examples/puzzletron/evaluation/hf_deployable_anymodel.py`. **2. `chore(lint): remove UP032 ignore and convert .format() to f-strings`** (10 files) - Drop `UP032` from `extend-ignore` in `pyproject.toml`. - Auto-convert 19 `"...".format(...)` calls to f-strings across export plugins, examples, tests, and tools. One conversion in `modelopt/torch/utils/plugins/megatron_generate.py` was wrapped manually to stay under the 100-char limit. **Intentionally left as-is:** - `tools/launcher/slurm_config.py` keeps its `# ruff: noqa: UP045` — nemo_run's CLI parser can't introspect PEP 604 optional annotations. - `modelopt/torch/puzzletron/*` is **not** touched. The subtree disables ruff's `UP` family entirely (per-file-ignore `"UP"`) while migration is in progress, and converting `Optional[X]` to `X | None` there would silently break runtime introspection in `block_config._get_dataclass_type` that uses `get_origin(tp) is typing.Union` (PEP 604 unions return `types.UnionType` from `get_origin`, not `typing.Union`). Best revisited when puzzletron's lint carve-out is narrowed. - `UP038` (`isinstance(x, (int, float))` → `isinstance(x, int | float)`) — ruff has officially deprecated this rule; PEP 604 in isinstance is slightly slower and misleads readers about PEP 695 / `Optional`. Ignore kept. ### Usage No user-facing API changes. ### Testing - Pre-commit hooks (ruff check, ruff format, mypy, bandit, license) pass on both commits. - Ruff status against `main`: 37 unrelated pre-existing findings (W291/W293/E501/RUF005/PLR1704); zero new findings introduced by this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — Runtime behavior of the six type aliases changes from a `typing.Union` instance to `types.UnionType`. Downstream code introspecting via `get_origin(...) is typing.Union` on these aliases would break, but no in-repo caller does this on them. (The introspection in `modelopt/torch/puzzletron/block_config.py` operates on user-supplied dataclass field types, none of which are these aliases.) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (no behavior change) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ — internal style refactor; happy to add a Misc note if reviewers want one. - Did you get Claude approval on this PR?: ❌ — not yet. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Style** * Modernized type annotations across the codebase to use Python 3.10+ union syntax and TypeAlias where appropriate. * Standardized string formatting to f-strings, improving clarity of logs, errors, and validation messages. * **Chores** * Updated linting configuration to reflect modern typing/style rules. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1537?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
7038dec918 |
[1/2Refactor] speculative decoding: use mto config subsystem (#1328)
### What does this PR do?
Type of change: new feature
Port the speculative-decoding example to ModelOpt's recipe/config
subsystem: `model` / `data` / `training` / `<algo>` now load from a
single YAML with Pydantic validation and OmegaConf dotlist overrides.
Adds built-in `eagle3` / `dflash` recipes, drops the redundant
`training.mode` field (inferred from recipe class), and shrinks
`main.py` by ~145 lines (−208 / +63).
JIRA: OMNIML-3859
### Usage
```bash
python main.py --config general/speculative_decoding/eagle3 \
model.model_name_or_path=meta-llama/Llama-3.2-1B \
data.data_path=train.jsonl \
training.output_dir=ckpts/test
```
### Testing
- `pytest tests/unit/recipe/test_loader.py` — new coverage for Eagle /
DFlash YAML loading, dotlist overrides, and field-level validation.
- Smoke-trained both built-in `eagle3` and `dflash` recipes end-to-end.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ❌ — `main.py` CLI switched to
`--config <recipe>` (+ dotlist overrides); the old argparse flags are
removed.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
deps (`pydantic`, `omegaconf` already in core).
- Did you write any new necessary tests?: ✅ —
`tests/unit/recipe/test_loader.py`.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — to be added.
### Additional Information
Follow-up to the `modelopt.recipe` subsystem introduced for PTQ; this PR
extends the same declarative-YAML pattern to speculative decoding
(Eagle3 / DFlash / Medusa).
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added typed speculative-decoding recipe support for EAGLE, DFlash, and
Medusa; CLI dotlist overrides supported for single-file recipes.
* Trainer/config schema extended with speculative-training fields and
draft-vocab cache loading for Eagle.
* **Bug Fixes**
* Offline training no longer mutates model configs; loader enforces
required algorithm sections and prints recipe/config only on the primary
process.
* Reduced noisy per-rank logging by restricting status output to the
primary process.
* **Tests**
* Expanded tests for recipe loading, dotlist overrides, validation
strictness, and error cases.
* **Documentation**
* Recipe YAMLs updated with metadata and usage notes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
383ab4e224 |
fix: include medusa in data_module assignment in main.py (#1370)
## Problem
When `training.mode == "medusa"` is used in `main.py`, the `data_module`
variable is never assigned because line 344 only covered `eagle3` and
`dflash` modes. This causes an `UnboundLocalError` when the trainer is
constructed with `**data_module`.
Fixes OMNIML-4147
## Fix
Add `"medusa"` to the `training_args.mode in ("eagle3", "dflash")`
condition so `data_module` is correctly populated for medusa training.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Fixed speculative decoding example to properly handle "medusa" mode
alongside existing "eagle3" and "dflash" modes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
||
|
|
1ec931c2c7 |
[2/3][Feat]: Offline DFlash training (#1343)
### What does this PR do? Type of change: new feature Part 2 of a 3-PR series splitting #1271: - **[1/3] #1296**: File reorg + deprecate `ParallelDraft` - **[2/3] this PR**: Offline DFlash training (depends on #1296) - **[3/3] #1297**: Extract `HFSpecDecMixin` Changes: - Add `dflash_offline` flag to `DFlashConfig` for training from pre-computed hidden states; deletes base model layers to save memory. - Add Pydantic validators on `DFlashConfig`: - `_derive_dflash_offline` — auto-derive `dflash_offline` from `data_args.offline_data_path` in validation context. Not user-configurable: any user-supplied value is overridden by the derived value. - `_resolve_mask_token_id` — auto-detect `dflash_mask_token_id` from `tokenizer.mask_token_id`. - `_check_mask_token_id` — fail fast if unset after resolution. - `HFDFlashModel.modify()`: select `num_orig_hidden_layers` when offline; pick `_base_model_lm_head` device when no base layers present; drop base-model `layers` module. - `HFDFlashModel.forward()`: add offline branch — consumes precomputed `base_model_outputs` via `DFlashBaseModelOutput.from_offline_dict`, and when `dflash_self_logit_distillation` is enabled with `base_model_logits` absent, recomputes logits from `base_model_hidden_states` via `_base_model_lm_head`. Raises a clear error from the non-training / `pseudo_speculative_generate` paths when `dflash_offline=True`, since base-model layers have been deleted. - `DFlashBaseModelOutput` dataclass in `modeling_dflash.py` (with `from_offline_dict` classmethod) to unify online/offline output shapes. `aux_hidden_states` is required in `from_offline_dict` so missing keys fail fast at the entry point rather than deeper in the forward. - `examples/speculative_decoding/main.py`: replace inline `mask_token_id` auto-detect with `DFlashConfig.model_validate(dflash_cfg, context={"tokenizer": tokenizer, "data_args": data_args})`. ### Silent bug fix — `add_generation_template` → `add_generation_prompt` The pre-refactor `compute_hidden_states_hf.py` passed `add_generation_template=False` to `tokenizer.apply_chat_template`. This kwarg does not exist on HF `apply_chat_template` and was being silently ignored, so the intended "don't append a generation prompt" behavior was never actually applied. The new `tokenize_with_loss_mask` helper in `examples/speculative_decoding/collect_hidden_states/common.py` uses the correct `add_generation_prompt=False`. **This is a real behavior change** for anyone re-dumping hidden states: trailing generation prompts that were previously appended to the tokenized sequences will no longer be included. ### Testing - New tests: - `tests/unit/torch/speculative/plugins/test_hf_dflash_offline.py` — CPU unit tests for convert path (online keeps base layers, offline deletes them; `num_orig_hidden_layers` drives `target_layer_ids` in offline mode) and `DFlashConfig._derive_dflash_offline` validator. - `TestDFlashOfflineForwardGPU` in `tests/gpu/torch/speculative/plugins/test_hf_dflash.py` — GPU forward smoke with precomputed `base_model_outputs`, plus the `dflash_self_logit_distillation` logit-recompute path. - training test: <img width="454" height="317" alt="image" src="https://github.com/user-attachments/assets/79b92790-4d15-4313-bb9b-f35665b012e6" /> <img width="456" height="310" alt="image" src="https://github.com/user-attachments/assets/4558559f-9c35-49ed-b36e-82fbc99eab23" /> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ — additive `dflash_offline` flag defaulting to `False`; validators fall through when context not provided. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — see Testing section above. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ ### TODO (follow-up) - [x] Update `examples/speculative_decoding/collect_hidden_states/compute_hidden_states_*.py` to support DFlash offline data. Current scripts are Eagle-specific — they hardcode the `[2, N/2, N-3]` aux-layer selection and emit `{input_ids, hidden_states, aux_hidden_states}`. DFlash offline needs: - Aux layer indices driven by `build_target_layer_ids(num_orig_hidden_layers, num_draft_layers)` (or a configurable list), not the Eagle triplet. - `base_model_hidden_states` key (last-layer hidden) so `DFlashBaseModelOutput.from_offline_dict` + the `dflash_self_logit_distillation` recompute path can consume it. - Optional `base_model_logits` dump so offline training can skip the self-distillation logit recomputation when logits are available. ### Additional Information Base branch is #1296 (file reorg). Retarget to `main` once #1296 merges. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Offline DFlash speculative-decoding training from precomputed base-model hidden states * Answer-only-loss training with persisted loss masks and optional chat-template support * Flexible auxiliary-layer selection via CLI and an exposed default aux-layer helper * Auto-derived offline flag in config and automatic memory optimization during offline conversion * **Documentation** * Updated guides for offline pipeline, aux-layer selection, and loss-masking options * **Tests** * New unit, GPU, and regression tests covering offline conversion, training, and config derivation <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
7c80d85751 |
[1/3][Refactor]: File reorg; deprecate ParallelDraft (#1296)
### What does this PR do? Type of change: refactoring Part 1 of a 3-PR series splitting #1271: - **[1/3] this PR**: File reorg + deprecate `ParallelDraft` - **[2/3] #1295**: Offline DFlash training - **[3/3] #1297**: Extract `HFSpecDecMixin` Changes: - **File reorg**: `transformers.py` → `hf_eagle.py`; extract `HFMedusaModel` → `hf_medusa.py`; extract `EagleModule` / `EagleBaseModelOutput` → `modeling_eagle.py`; extract `DFlashModule` / `DFlashAttention` / `DFlashDecoderLayer` / `build_target_layer_ids` / `apply_rotary_pos_emb` → `modeling_dflash.py`. - **Deprecate `ParallelDraft`**: remove `parallel_draft_step`, `parallel_draft_heads_num_layers`, and the `ParallelDraft` module from HF Eagle; remove the `EagleMedusaExporter` branch from `HFEagleModel.get_exporter()` (the `EagleMedusaExporter` class itself still lives in `hf_spec_export.py` for Megatron parity). - **Rename**: `_draft_model_config` → `eagle_config` in export plugin. - Update imports in `examples/speculative_decoding/` and `modelopt/torch/speculative/utils.py` to follow the module rename. ### Testing Validated with existing Eagle and DFlash training scripts (re-run after `9ae5302729 revert behavior change`). ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ — renames `modelopt.torch.speculative.plugins.transformers` → `.hf_eagle`; removes `parallel_draft_step` / `parallel_draft_heads_num_layers` from Eagle config; renames `_draft_model_config` → `eagle_config` in export plugin. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — pure refactor; existing tests updated for the rename. `test_hf_spec_rope_export.py` assertions were also corrected to reflect the actual production path (the old assertions were masked by `MagicMock` not invoking the `_draft_model_config` `@property`). - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ ### Additional Information Breaking changes: - `modelopt.torch.speculative.plugins.transformers` → `.hf_eagle` - `parallel_draft_step` / `parallel_draft_heads_num_layers` removed from Eagle config - `_draft_model_config` → `eagle_config` in export plugin <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactoring** * Reorganized speculative-decoding plugins into focused modules, converting the legacy "transformers" entry into a deprecated shim that re-exports the new plugin surface. * Consolidated DFlash implementation into a shared modeling component and introduced a dedicated EAGLE decoder module. * **New Features** * Added a Medusa speculative-decoding plugin with configurable heads and combined-loss training behavior. * **Chores** * Updated pre-commit license-hook exclusion and feature-flag wiring. * **Tests** * Updated export tests to expect rope-scaling fallback semantics. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
2fef374ded |
fix: auto-compute dp_replicate_size from world_size (#1302)
## Summary - When `dp_shard_size < world_size` (e.g., `dp_shard_size=4` on 8 GPUs across 2 nodes), `ParallelismConfig` raises `total_size (4) does not match num_processes (8)` because `dp_replicate_size` defaults to 1 - Auto-compute `dp_replicate_size = world_size // (dp_shard_size * cp_size)` so intra-node FSDP2 sharding + inter-node data-parallel replication works without manual config - This enables `dp_shard_size` to be set to per-node GPU count (better NVLink utilization) while automatically creating replicas across nodes ## Test plan - [ ] Verify single-node training (dp_shard_size == world_size, dp_replicate_size == 1) unchanged - [ ] Verify multi-node with dp_shard_size < world_size creates correct replica groups - [ ] Verify existing EAGLE3/DFlash configs still work 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Enhanced parallelism configuration initialization in the speculative decoding example to better handle distributed training scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
355c6b7883 |
fix: PTQ 1GPU, export PP divisibility, hidden states conversations key (#1293)
## Summary - **megatron_lm_ptq.yaml**: Qwen3-8B PTQ to single GPU for L40 clusters (TP=1, all tasks) - **quantize.sh**: Auto-find largest PP dividing model's `num_hidden_layers` for export step. Qwen3-8B has 36 layers which isn't divisible by 8, causing `AssertionError` on 8-GPU nodes - **compute_hidden_states_trtllm.py**: Use `messages` with `conversations` fallback, matching the HF version. Fixes `KeyError: 'conversations'` when data uses OpenAI `messages` format ## Test plan - [x] Qwen3-8B PTQ runs on single L40 GPU - [x] Export PP auto-selects valid divisor (36 layers → PP=6 on 8 GPUs, PP=4 on 4 GPUs, PP=1 on 1 GPU) - [x] EAGLE3 offline pipeline reads data with `messages` field 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Dataset input handling now supports multiple field formats for enhanced compatibility. * **Bug Fixes** * Optimized GPU resource allocation during model quantization with improved pipeline parallelism computation. * Updated quantization configuration for more efficient resource utilization. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
07ae8e7128 |
Add LoRA co-training support for HF EAGLE speculative decoding (#1060)
### What does this PR do?
Type of change: New feature + bug fixes
Adds **LoRA co-training** support for HF EAGLE speculative decoding.
When `eagle_base_lora=True`, HF PEFT LoRA adapters are injected into the
base model and co-trained alongside the EAGLE draft module in a single
online training pass. A preservation loss (KL divergence between the
original frozen base model output and the LoRA-adapted output) prevents
base model drift. LoRA adapter weights are exported in standard peft
format alongside EAGLE draft artifacts.
### Key features
- **LoRA injection**: `peft.inject_adapter_in_model` applied in-place
(no wrapper), keeping the existing `HFEagleModel` structure intact.
- **Preservation loss**: Cross-entropy `H(ref, lora)` — equivalent
gradient to `KL(ref || lora)` since `H(ref)` is constant w.r.t. LoRA
params.
- **Warmup schedule**: `eagle_base_lora_warmup_steps` freezes LoRA for N
steps while the EAGLE head stabilizes, then enables co-training via a
`LoRAWarmupCallback`.
- **Logits detach regularization**: `eagle_base_lora_logits_detach_prob`
stochastically detaches base logits from the EAGLE loss path, preventing
LoRA from degenerating to maximize EAGLE accuracy at the cost of base
model quality.
- **Export**: Standard peft format (`adapter_model.safetensors` +
`adapter_config.json`) alongside EAGLE draft model.
- **Merge script**: `scripts/merge_lora.py` merges LoRA weights into the
base model and restores the original `config.json` (avoids transformers
5.x rewriting `rope_theta` → `rope_parameters` which breaks
vLLM/TRT-LLM).
- **Multinode fix**: `dp_shard_size` now uses `WORLD_SIZE` instead of
local GPU count.
### Config options
```python
mtsp.convert(model, mode=[("eagle", {
"eagle_base_lora": True, # enable LoRA co-training
"eagle_base_lora_rank": 64, # LoRA rank
"eagle_base_lora_alpha": 16.0, # LoRA scaling
"eagle_base_lora_target_modules": ["q_proj", "k_proj", "v_proj", "o_proj"],
"eagle_base_lora_preservation_loss_weight": 0.1, # preservation loss weight
"eagle_base_lora_warmup_steps": 0, # freeze LoRA for N steps
"eagle_base_lora_logits_detach_prob": 0.5, # detach prob (0=never, 1=always)
})])
```
### Experimental results (Qwen3-8B, checkpoint-60000)
Base model quality preserved across detach_prob sweep (lm_eval: IFEval,
ARC-C, Winogrande — results pending final collection).
**Acceptance rate** (mt_bench, draft_length=3, output_length=4096,
temperature=0):
| detach_prob | vLLM AR | TRT-LLM AR |
|---|---|---|
| baseline (no LoRA) | 2.14 | 2.15 |
| 0.5 | 1.45 | 1.44 |
| 0.8 | **3.06** | **3.01** |
| 0.85 | 2.90 | 2.90 |
| 0.9 | 2.76 | 2.77 |
| 0.95 | 2.51 | 2.58 |
| 0.99 | 2.37 | 2.37 |
| 0.999 | 2.30 | 2.27 |
| 0.9999 | 2.31 | 2.26 |
Best AR at `detach_prob=0.8`: ~40% improvement over baseline.
### Testing
`tests/unit/torch/speculative/plugins/test_hf_speculative_lora.py` (5
tests):
- `test_lora_layers_injected` — LoRA layers present after conversion
- `test_trainable_params` — only `lora_*` and `eagle_module` params are
trainable
- `test_forward_returns_loss` — forward returns non-zero scalar loss
- `test_eagle_offline_incompatible` — `eagle_base_lora=True` +
`eagle_offline=True` raises `ValueError`
- `test_export_lora_artifacts` — export produces standard peft adapter
files
### Bug fixes (included in this PR)
1. **`launch_train.sh` case pattern ordering**: glob
`--eagle_base_lora*` was before specific patterns
(`--eagle_base_lora_rank*`, etc.), silently swallowing LoRA args.
2. **LoRA optimizer exclusion during warmup**: warmup freezing excluded
LoRA from the optimizer entirely; fixed with `add_param_group` in the
callback.
3. **`merge_lora.py` config.json**: `save_pretrained()` with
transformers >=5.x rewrites `rope_theta` → `rope_parameters`, breaking
vLLM positional embeddings. Fixed by copying the original base model
config.
4. **Multinode `dp_shard_size`**: used local GPU count instead of
`WORLD_SIZE`.
### Checklist
- [x] Backward compatible (all new config fields have defaults)
- [x] Uses `peft` via lazy imports (no hard dependency)
- [x] Unit tests added
- [x] Online HF training only (`eagle_offline=True` blocked)
---------
Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
||
|
|
3131195241 |
add: DFlash block diffusion speculative decoding (#1211)
DFlash (Block Diffusion for Flash Speculative Decoding) predicts an entire block of tokens in a single forward pass using masked parallel prediction with KV injection from the target model's hidden states. Key features: - Feature fusion (multi-layer hidden states -> FC + RMSNorm) - KV injection (fused features as K/V in every draft layer with QK-norm) - Random anchor sampling with bidirectional intra-block attention - Logit distillation with exponential loss decay (gamma weighting) - Multi-node DDP training with checkpoint resume - Export to z-lab compatible HF format - Online validation (context-dependent ground truth) Training recipe: modelopt_recipes/general/speculative_decoding/dflash.yaml Results: examples/speculative_decoding/doc/dflash_results.md ### ModelOpt Eval (online validation, osl=512) | Dataset | z-lab | ModelOpt (306K) | Diff | |---------|-------|-----------------|------| | gsm8k | 4.10 | **5.19** | **+1.09** | | MT-Bench | 3.58 | **4.36** | **+0.78** | ### z-lab Official Eval (dflash.benchmark, osl=512) | Dataset | z-lab | ModelOpt (306K) | Diff | |---------|-------|-----------------|------| | gsm8k | **5.00** | 4.08 | -0.92 | | MT-Bench | **3.28** | 2.99 | -0.29 | > z-lab model trained with block_size=16. ModelOpt trained with block_size=8. ## Evaluation Method Impact (gsm8k) | Eval Method | z-lab checkpoint | ModelOpt (306K) | |-------------|-----------------|-----------------| | Fixed GT (ModelOpt eval) | 2.95 | 4.23 | | Online GT (ModelOpt eval) | 4.10 | **5.19** | | z-lab official eval | **5.00** | 4.08 | ### What does this PR do? Type of change: ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added DFlash speculative decoding mode with parallel block prediction support. * Included training launchers and MT-Bench evaluation scripts for DFlash models. * Added online acceptance rate validation for improved inference verification. * **Documentation** * DFlash quick start guide with configuration parameters and training examples. * Performance results and benchmarks for DFlash-trained models. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
6403389eb0 |
Feat: Configurable Eagle ROPE scaling during export (#1238)
### What does this PR do? JIRA ticket: https://jirasw.nvidia.com/browse/OMNIML-3469 Type of change: New feature Decouple EAGLE training rope configuration from export rope configuration, enabling separate YaRN rope scaling injection at export time for long-context inference. #### Changes **Configurable export rope scaling (`EagleConfig`)** - Add `eagle_export_rope_scaling` field to `EagleConfig` with default YaRN config (`factor=32.0`, `original_max_position_embeddings=2048`) - Set to `{}` to disable rope scaling injection at export **Simplified training defaults (`default_config.py`)** - Change default training rope from `llama3` (theta=500k) to `default` (theta=10k) — models now train with simple positional embeddings; rope scaling is applied only at export - Add `rope_theta` inside `rope_scaling` dict for transformers 5.x cross-version compatibility **Move config validation/rewriting into `EagleConfig` (`config.py`)** - `_derive_eagle_offline`: derives `eagle_offline` from `data_args.offline_data_path` via validation context, removing manual assignment in `main.py` - `_check_rope_scaling_consistency`: rejects configs where `eagle_export_rope_scaling` is set but training `rope_type` is not `"default"` - `_warn_rope_vs_training_seq_len`: warns when `original_max_position_embeddings` differs from `training_seq_len` **Export rope injection (`hf_spec_export.py`)** - Inject `eagle_export_rope_scaling` into the exported HF config when training rope_type is `"default"` - Fall back `rope_theta` from `rope_scaling` dict for transformers 5.x compatibility **Fix Megatron RotaryEmbedding crash (`megatron_eagle.py`)** - `dict_to_config()` set `rope_scaling=True` whenever the `rope_scaling` key existed, even without a `"factor"` — causing `RotaryEmbedding` to divide by `None` - Now only enables `rope_scaling` when the dict actually contains a `"factor"` key ### Usage Configure in YAML config (or use defaults from `eagle3.yaml`): ```yaml eagle: eagle_export_rope_scaling: rope_type: yarn factor: 32.0 original_max_position_embeddings: 2048 ``` Set to empty dict to disable export rope injection: ```yaml eagle: eagle_export_rope_scaling: {} ``` ### Testing - New unit tests: `tests/unit/torch/speculative/test_eagle_config.py` — rope consistency validator, seq_len warning, context-derived `eagle_offline` - New unit tests: `tests/unit/torch/export/test_hf_spec_rope_export.py` — export rope injection, fallback, and empty-config cases ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ (new field has sensible default; existing configs work unchanged) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ (should be added if merging as a feature) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Add export-time rope-scaling configuration for EAGLE models. * **Improvements** * Stronger validation and context-aware reconciliation between training and export configs. * Export now injects rope-scaling and rope-theta when appropriate. * Default rope-scaling values updated for EAGLE variants. * Model instances now expose export rope-scaling for downstream use. * **Tests** * Added unit tests covering rope-scaling export behavior and configuration validators. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
9050188034 |
Fix test_collect_hidden_states: use synthetic short conversations (#1234)
## Summary - `test_collect_hidden_states` was using real daring-anteater conversations (typically 1000+ tokens) but the tiny test model has `max_position_embeddings=32`. Both sampled conversations exceeded the default `--max-seq-len 3072` filter, producing zero `.pt` files and failing the assertion. - Added a `tiny_conversations_path` fixture with synthetic short single-turn conversations that tokenize within `max_position_embeddings=32`. - Changed `test_collect_hidden_states` to use this fixture with `--max-seq-len 32`. - Added a `None` guard for `tokenizer.chat_template.replace(...)` to avoid `AttributeError` when the tokenizer has no chat template. ## Test plan - [ ] `pytest tests/examples/speculative_decoding/test_eagle_offline_ptq.py::test_collect_hidden_states` passes - [ ] CI `speculative_decoding` job passes 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Resolved compatibility issues when tokenizers do not have a chat template configuration by adding proper error handling. * Standardized tokenization input extraction logic across different transformer library versions for consistent behavior. * **Tests** * Enhanced test infrastructure with new conversation data fixtures and improved sequence length validation for speculative decoding examples. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
c901814ff9 |
Fix compute_hidden_states_hf.py: handle BatchEncoding from apply_chat_template (#1225)
## Summary
- `apply_chat_template(..., return_tensors="pt")` returns a
`BatchEncoding` in transformers 4.46+, which no longer subclasses `dict`
- The old guard `isinstance(tokenized, dict)` evaluates to `False` for
`BatchEncoding`, so `input_ids` was set to the whole `BatchEncoding`
object
- Calling `.shape[1]` on a `BatchEncoding` triggers
`__getattr__("shape")` → `AttributeError`
- Fix: check `isinstance(tokenized, torch.Tensor)` instead, which
correctly handles both old transformers (plain tensor) and new
transformers (BatchEncoding)
This is causing `test_collect_hidden_states` to fail in the speculative
decoding CI for all open PRs (#1207, #1210, #1221).
## Test plan
- [ ] `torch-pr (speculative_decoding, 26.01)` CI passes
- [ ] Verify fix handles both `torch.Tensor` return (old transformers)
and `BatchEncoding` return (new transformers 4.46+)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
||
|
|
04cd596d79 |
Add experimental support for transformers>=5.0 + min torch 2.8 (#975)
### What does this PR do? - Add experimental support for transformers >=5.0 and remove deprecated usages: https://github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md - ⚠️ For accelerate examples that used `--warmup-ratio: float` (deprecated in 5.x), we now change it to `--warmup-steps: float | int` which works as ratio if float but only for 5.x. For 4.x, it will error out if float and prompt user to change back to `--warmup-ratio` or pass an int absolute step count. - ⚠️ Unified Hugging Face checkpoint export for quantized checkpoints may not work for some models with transformers>=5.0 yet as it requires a lot of fixes (e.g. change in how MoE experts are organized) - ~Add Workaround for TRT-LLM's import of deprecated transformers functions so trt-llm based gpu unit tests work fine. Still deployment for models needs proper fixes directly in TRT-LLM hence llm/vlm ptq example tests still run with transformers 4.57~ - Everything except PTQ and Export (mainly MoE) should work fine with transformers>=5.0 - Bump min torch to 2.8 and enable 2.11 cicd testing - NOTE: Upcoming Nemo:26.04 container comes with transformers 5.3 ### Testing <!-- Mention how have you tested your change if applicable. --> - [x] CI/CD tests passing - [x] Manually tested unit tests, gpu tests with transformers 4.56 and 5.4 - [x] Manually tested example tests (except trt-llm container tests) with transformers 4.56 and 5.4 - [x] 2-gpu nightly CICD tests manually triggered and passing: [gpu tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867257540), [example tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867260643) ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, using `torch.load(..., weights_only=True)`, avoiding `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other source, did you follow IP policy in [CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?: N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Make remote-code usage opt-in via a configurable --trust_remote_code flag across examples and tools. * **Bug Fixes** * Improve checkpoint/resume detection and related training guidance to avoid erroneous errors. * **Refactor** * Consolidate dtype/config naming, switch warmup settings from ratio → steps, and unify tokenizer invocation patterns. * **Documentation** * Simplify changelog title and add misc notes for release 0.44. * **Chores** * Remove scheduled PR-branch cleanup workflow and relax/remove several transformers version pins. * **Tests** * Adjust test gates, skips, and structures to align with updated deps and behaviors. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
cccfded8a9 |
Add support for offline speculative decoding model PTQ (#883)
## What does this PR do? **Type of change:** new feature **Overview:** This PR enables loading in a ModelOpt pretrained offline speculative decoding model (e.g., EAGLE3) and performs PTQ on it and export. ## Usage Follow the speculative_decoding examples to train an offline speculative decoding model first. Then follow the command below to quantize and export it: ```bash python hf_ptq.py --pyt_ckpt_path <dir_of_offline_specdec_model> --specdec_offline_dataset <dir_of_dataset> ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Offline speculative decoding workflow: support loading a local dataset for calibration, generation, and export; new CLI option to specify the offline dataset. * **Improvements** * Export and quantization paths now accept and propagate offline speculative-decoding inputs. * Offline data loading honors a sample-size limit and enforces safe batch sizing for calibration. * **Bug Fixes** * Better handling of model/config mismatches and varied batch types in offline flows. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
ba4f42df1c |
Minor fix for example tests
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
ebc534d765 |
Update code-copying guidelines in CONTRIBUTING.md
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
0246041b01 |
feat(speculative): add vLLM data synthesis pipeline and Nemotron dataset preparation scripts (#1176)
### What does this PR do? Type of change: New feature, new example, bug fix Adds a vLLM-based synthetic data generation pipeline for speculative decoding draft model training, along with dataset preparation scripts for NVIDIA's Nemotron Post-Training dataset collections. **Data synthesis pipeline** (`tools/launcher/common/vllm/query.sh` + `common/query.py`): - Launch a vLLM server and run multi-turn inference to synthesize training data from input conversation skeletons - Fork-safe OpenAI client: reinitializes HTTP connection pool after `datasets.map()` forks worker processes, preventing 400 errors from corrupted connections - Clear Docker `ENTRYPOINT` so vLLM containers (which default to `vllm serve`) work correctly under NeMo Run's executor - `--max-tokens` argument to bound generation length - Local file loading support (`--data /path/to/file.jsonl`) - Re-raise connection errors so `datasets.map()` halts the shard instead of silently producing empty rows - Map `developer` role to `system` (OpenAI format compatibility) **Multi-turn reasoning trace handling** (`common/query.py`): - Strip `<think>...</think>` blocks from intermediate assistant turns before re-feeding to the model; preserve the full trace only on the final turn **Nemotron dataset preparation** (`examples/dataset/`): - `make_nemotron_ptv2_dataset.py` — prepares [nvidia/Nemotron-Post-Training-Dataset-v2](https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2) (~3.3M rows generate, ~1.9M rows train) - `make_nemotron_ptv3_dataset.py` — prepares the [Nemotron PTv3 collection](https://huggingface.co/collections/nvidia/nemotron-post-training-v3) of 16 datasets (~3.4M rows generate, ~3.9M rows train) - Both support `generate` mode (strips assistant turns for synthesis input) and `train` mode (normalizes to clean OpenAI format for SFT) - `conversation_utils.py` — shared utilities: `strip_assistant_turns`, `normalize_messages`, `make_augment_fn`, `AugmentationSpec` - `augmentations.yaml` — 12 language-redirect variants + style/format hints, cycled across dataset rows - Scripts live in `examples/dataset/` (not under `speculative_decoding/`) to signal reusability beyond speculative decoding **Bug fixes**: - `strip_assistant_turns()`: return `{"messages": []}` when no user turns remain (system-only rows were previously passed through instead of being filtered) - `concatenate_datasets()`: guard against empty parts list - SSH tunnel user precedence: explicit `user` arg now correctly overrides `slurm_config.user` ### Usage ```bash # Prepare PTv3 input conversations for synthesis (~3.4M rows): python examples/dataset/make_nemotron_ptv3_dataset.py --output-dir /tmp/ptv3_gen # Launch vLLM server + synthesize responses: bash tools/launcher/common/vllm/query.sh \ --model /path/to/model \ --tensor-parallel-size 4 \ -- \ --data /tmp/ptv3_gen/default.jsonl \ --save /tmp/ptv3_responses \ --num-shards 10 --num-proc 4 --max-tokens 4096 # Prepare PTv2 for direct SFT training (~1.9M rows): python examples/dataset/make_nemotron_ptv2_dataset.py --mode train --output-dir /tmp/ptv2_train ``` ### Testing Tested end-to-end on an NVIDIA GB10 node (119 GiB GPU memory) with `vllm/vllm-openai:qwen3_5-cu130` container and `Qwen/Qwen3.5-4B`: - vLLM server starts correctly with cleared Docker entrypoint - `datasets.map(num_proc=4)` runs without connection errors (fork-safe client) - Multi-turn synthesis produces correct assistant responses with thinking traces handled ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (data synthesis scripts; tested manually) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added dataset generation and augmentation capabilities for Nemotron post-training datasets (v2 and v3) * Enhanced query functionality with thinking-block filtering and improved client management for robust parallel processing * Added support for local dataset file paths alongside HuggingFace Hub datasets * **Bug Fixes** * Fixed SLURM executor user resolution and Docker container entrypoint configuration * Improved error handling for connection failures during dataset synthesis * **Documentation** * Updated dataset preparation guide with new generation modes and augmentation configuration details <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan D. Yu <5185878+ChenhanYu@users.noreply.github.com> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
82d96a635f |
[Speculative Decoding] Refactor EAGLE3 training to YAML-based config and recipe system (#1134)
## What does this PR do?
Refactors EAGLE3 training to use a single base YAML config with
OmegaConf dotlist overrides.
**Type of change:** Refactor
## Changes
- Single base config
`modelopt_recipes/speculative_decoding/_base_eagle3.yaml` for all EAGLE3
training; removed per-model child YAMLs.
- `launch_train.sh` accepts `--config <yaml>` plus dotlist overrides
(e.g. `model.model_name_or_path=xxx`).
- Removed `__base__` YAML inheritance logic from `main.py`.
- `dp_shard_size` default changed from `0` sentinel to `None` for
clarity.
- Removed `eagle_config.json` and `fsdp_config.json`; architecture
config is now nested under `eagle.eagle_architecture_config` in YAML.
- `train_eagle3_and_export.sh` now uses base YAML + dotlist instead of
generating a temporary YAML.
- Updated README and tests accordingly.
## Usage
```bash
# Online training
./launch_train.sh \
--config ../../modelopt_recipes/speculative_decoding/_base_eagle3.yaml \
model.model_name_or_path=meta-llama/Llama-3.2-1B \
data.data_path=input_conversations/train.jsonl \
training.output_dir=ckpts/llama-3.2-1b-online
# Offline training
./launch_train.sh \
--config ../../modelopt_recipes/speculative_decoding/_base_eagle3.yaml \
model.model_name_or_path=meta-llama/Llama-3.2-1B \
data.offline_data_path=$HIDDEN_STATES_DIR \
training.output_dir=ckpts/llama-3.2-1b-offline
```
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
5dc17dfd15 |
[Security] Enable torch.load(weights_only=True) for secure checkpoint loading + trust_remote_code fix (#1181)
### What does this PR do? - Add secure checkpoint loading support using `torch.serialization.add_safe_globals([cls])`. This also removes 1 existing pickle usage. - Remove hard-coded `trust_remote_code=True` - Replaces https://github.com/NVIDIA/Model-Optimizer/pull/1056 by @RinZ27 ### Testing <!-- Mention how have you tested your change if applicable. --> CICD tests ran ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information NVBug: 5999336 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added safe checkpoint save/load helpers and a --trust_remote_code CLI flag in examples to control remote-code loading. * **Bug Fixes** * Checkpoint loading now defaults to safer, weights-only semantics to reduce arbitrary-code exposure. * **Documentation** * CHANGELOG updated with security guidance and opt-in procedure for unsafe checkpoint loading. * **Tests** * New unit tests validating the safe-load behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: RinZ27 <222222878+RinZ27@users.noreply.github.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: RinZ27 <222222878+RinZ27@users.noreply.github.com> |
||
|
|
80d2f02a2d |
Fix spec dec example tests (#1183)
### What does this PR do? Type of change: Test fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> - Fix `tests/examples/speculative_decoding` - previously silently skipped - Avoid pulling nemotron-post-training-dataset-v2 in tests to reduce chances of HF loading timeout in CICD - Make slow and redundant tests manual to speed up CICD ### Testing <!-- Mention how have you tested your change if applicable. --> - Tests passing ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Removed git‑LFS install step from CI and deleted an automated branch‑cleanup workflow * Trimmed example environment dependencies and relaxed transformers compatibility; added an optional tokenization dependency * **Tests** * Switched tests to generate datasets dynamically and improved fixture handling * Standardized PTQ test parameters (explicit calibration dataset) and refined GPU/test selection * **Bug Fixes** * Improved device-awareness and numeric handling in speculative decoding attention paths <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
7f5fd65003 |
[Feat]FakeBaseModel for offline eagle; Kimi-K2.5 fixes; (#1052)
### What does this PR do? Adds `FakeBaseModel` for offline EAGLE training and several Kimi-K2.5 compatibility fixes. - **New**: `FakeBaseModel` — lightweight model that loads only `lm_head` and `embed_tokens` from a local checkpoint, avoiding full model weight loading during offline training. Configured via `FakeBaseArguments` and integrated into `load_vlm_or_llm`. - **Fix**: `_find_base_model_parts` — support Kimi-K2.5 VLM layout (`language_model.model` path) - **Fix**: offline mode lm_head access and CompressedTensors ignore path - **Fix**: Kimi-K2.5 decoder `past_key_value`/`past_key_values` argument mismatch - **Fix**: `rglob` for `.pt` discovery in nested offline data dirs; single-node GPU count respects `CUDA_VISIBLE_DEVICES` Type of change: Bug fix, new feature ### Testing Tested offline EAGLE training for Kimi-K2.5 end-to-end. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ❌ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Lightweight fake-base model support for offline speculative-decoding training * **Improvements** * Added CLI flags: --use_fake_base_for_offline, --trust_remote_code, and --fsdp * Expanded offline .pt discovery to include nested subdirectories * Better GPU detection with explicit single-node logging; FSDP enabled only when requested * Model loading and launch tooling now honor offline and trust-remote-code flags * **Bug Fixes** * Improved compatibility with legacy transformer / Kimi-K2 call signatures * **Tests** * Added tests covering fake-base loading and offline training workflows <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
d0bf0bef96 |
Remove deprecated Nemo 2.0 references / examples (#1098)
### What does this PR do? - Remove `examples/nemo_run` and other deprecated Nemo 2.0 references - Add Megatron-Bridge example links where missing <!-- Details about the change. --> ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!--- Mandatory --> - Did you write any new necessary tests?: N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Deprecations** * Removed deprecated NeMo 2.0 support and related example flows and utilities. * **Documentation** * Updated docs and examples to emphasize Megatron-Bridge / Megatron-LM and refreshed technique/deployment guidance and links. * **New Features** * Added CLI options for additional parallelism (context/expert tensor/expert model) in Megatron-Bridge distillation. * **Chores** * Removed legacy CI configs and refreshed container image tags across examples and docs. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
547d8fe935 |
[EAGLE] Optimize EAGLE Training (#1044)
### What does this PR do? Type of change: Optimization <!-- Details about the change. --> Changes: - Precompute base_model_logits.argmax() and base_model_logits.softmax() instead of recomputing in every call to _eagle_loss - Calculate per-prediction accuracy on the GPU and synchronize it to the host after running all TTT steps, to avoid cpu/gpu synchronization inside the TTT step loop. - Apply torch.compile to performance-critical training functions: prepare inputs, eagle forward, and eagle loss calculation. Omitted from target model in online case for now, as it may not be natively compatible with all architectures. ### Usage No changes to external interfaces ### Testing Ran training commands for benchmarking. Did not do a full training run, did not validate correctness. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!--- Mandatory --> - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added runtime profiling ranges for speculative decoding to enable finer performance tracing. * **Improvements** * Lazy initialization for rotary embeddings in llama-style decoders for more reliable startup. * Speculative decoding now uses base-model predicted tokens and softmax probabilities for sampling and loss, improving stability and accuracy reporting. * Parallelism configuration is now conditional, avoiding unnecessary setup unless required. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Benjamin Chislett <bchislett@nvidia.com> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
c76633ac9d |
[EAGLE] Configurable number of TTT steps (#1042)
### What does this PR do? Type of change: new CLI option for existing option <!-- Details about the change. --> - Added num_ttt_steps CLI flag - Changed num_ttt_steps default from 4 to 3 for consistency. Num_spec_tokens == 3 or == 7 are most common in practice, so rounding down to 3 and allowing users to increment higher on-demand. Will also improve training efficiency for the OOTB experience. ### Usage Users can now pass `--num_ttt_steps 7` to `launch_train.sh` when training an EAGLE3 model for extended speculation lengths. ### Testing N/A ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added ability to configure train-time-test steps for speculative decoding training via command-line argument. * Updated default train-time-test steps value from 4 to 3. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Benjamin Chislett <bchislett@nvidia.com> |
||
|
|
4292505512 |
Refactor: Clean up EAGLE training dataset preparation (#684)
## What does this PR do? **Type of change:** Refactor **Overview:** - Consolidate input dataset preparation into `make_dataset.py` - Read dataset mix spec from a YAML file - - Can now specify how many samples to take from each split - - Can no longer easily split a dataset into train/test sections. I don't think this feature was really useful to begin with. Most datasets can already be separated into train/val/test at the split level, and those that can't are usually going to be splitted by the training FW anyways. - Add support for a few new dataset types, magpie 300k/500k/1M, nemotron post-training dataset v2. ## Usage See README for detailed example ## Testing Ran it locally on all dataset modes, works successfully and output looks good. Checked shuffling, conversation IDs, and output contents were all unique and usable. - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated speculative decoding example documentation with new dataset references and standardized file paths. * **New Features** * Introduced configuration-driven dataset preparation supporting multiple dataset sources with centralized configuration files. * **Refactor** * Simplified dataset preparation workflow with unified tooling and updated default data paths throughout the training pipeline. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Benjamin Chislett <bchislett@nvidia.com> |
||
|
|
5d0e012751 |
inplement mix hidden_states for eagle3; deprecate eagle1 (#946)
## What does this PR do? new feature **Overview:** Enable mix hidden_states in eagle3 training. Deprecate eagle1 ## Usage Add --mix_hidden_states True to launch_train.sh ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added --mix_hidden_states option to enable optional hidden-state mixing during training. * Added eagle_ttt_steps setting to control speculative multi-step iterations. * **Chores** * Consolidated speculative decoding to EAGLE3 only; legacy Medusa/EAGLE1 paths removed. * Unified configuration handling so models and plugins accept a single config object. * **Tests** * Updated and expanded tests for hidden-state mixing and EAGLE3-only scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> |
||
|
|
dd16a96fda |
example demonstrating how to train CosmosReason2 Eagle3 (#965)
## What does this PR do? **Type of change:** New example **Overview:** Adds examples/speculative_decoding/guides/train_eagle_head_cosmos_reason2.ipynb, a step-by-step Jupyter notebook that walks through the full EAGLE3 draft-head training workflow for nvidia/Cosmos-Reason2-8B. The notebook covers: 1. Installing dependencies 2. Authenticating with Hugging Face 3. Preparing training data from the Nemotron-Post-Training-Dataset-v2 (chat split) using a curated row-selection mapping (guides/nemotron_mapping.csv) 4. Inspecting the bundled EAGLE3 config (guides/CR2_eagle_config.json) tuned for Cosmos-Reason2 (YaRN RoPE, FlexAttention, reduced draft vocabulary) 5. (Optional) Calibrating the draft vocabulary to 32k tokens for faster training and inference 6. Launching training via launch_train.sh with FSDP2 multi-GPU support 7. Exporting the checkpoint to HF format and serving with vLLM Also includes guides/nemotron_mapping.csv and guides/CR2_eagle_config.json as companion files. ## Usage Open and run examples/speculative_decoding/guides/train_eagle_head_cosmos_reason2.ipynb cell by cell. After training, serve the exported checkpoint with: ``` vllm serve nvidia/Cosmos-Reason2-8B \ --host 0.0.0.0 \ --port 8000 \ --speculative-model export/cosmos-reason2-8b-eagle3 \ --num-speculative-tokens 3 \ --dtype bfloat16 ``` ## Testing Tested end-to-end on a 4xB100 GPUs. The exported checkpoint was validated with specdec_bench ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: Yes (the notebook is self-documenting) - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No ## Additional Information Cosmos-Reason2-8B requires at least one 80 GB GPU (H100/A100) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a speculative-decoding configuration for draft model and rotary/attention behavior. * Added an end-to-end training notebook demonstrating training, export, and deployment of a speculative-decoding draft head on Cosmos-Reason2. * Added a data-preparation tool to download, normalize, and convert Nemotron chat conversations into a standardized conversation format. * **Documentation** * Notebook documents environment setup, data prep, training/validation cadence, export, and deployment steps. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Slawek Kierat <skierat@nvidia.com> |
||
|
|
a34d613d3c |
Feat: Speculatice Decoding export with quantization support (#913)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Main changes: - Refactored speculative decoding export logics into `class EagleExporter` to improve cohesion; - Separated speculative decoding export entrance with quantization export (`export_hf_checkpoint()`) due to their fundamental differences: - Quantization export base model's state_dict and config, while speculative decoding only export drafter's. - Most of the model-specific logics of quantization export (e.g. diffusers, vlms) are not needed for speculative decoding export. - Quantization export produce different format than speculative decoding checkpoint. (The former produce tokenizer config, generation config, e.t.c, while the later does not need. ) ## Usage <!-- You can potentially add a usage example below. --> To export an regular bf16 eagle checkpoint without quantization, the commands are the same: ```python python scripts/export_hf_checkpoint.py --model_path <x> --export_path <x> ``` To run PTQ on online-trained eagle checkpoint and export it: ```python python hf_ptq.py --pyt_ckpt_path <x> --qformat fp8 --export_path <x> ``` The above two commands will produce drafter ckpt for deployment, in the same foramt. ## Testing <!-- Mention how have you tested your change if applicable. --> Tested setting: - Base model: llama3.1-8b - Algorithms: eagle - Export path tested: - (Unquantized online ckpt) `python scripts/export_hf_checkpoint.py --model_path <x> --export_path <x>` - (PTQ) export `python hf_ptq.py --pyt_ckpt_path <x> --qformat fp8 --export_path <x>` - Tested deployment on vllm. Got normal AR. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added export functionality for speculative decoding-optimized models * Support for multiple speculative decoding architectures with pre-configured deployment templates * Enhanced model export detection and automatic routing for optimized models * **Tests** * Updated export validation tests for speculative decoding models <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
edde087bbd |
Adding a special list to handle HF models that require (#950)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Bug fix **Overview:** ? 1. Some models check `layer_types` and our offline overwrites `num_hidden_layers=0` which will requires passing `layer_types=[]`. It is also possible that the model will raise unexpected arguments when passing `layer_types=[]`. So the current WAR is to use a list to handle by checking the architectures. 2. Fix relative path issue in `launch_train.sh` ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Enhanced model loading robustness for specific layer configurations. * **Chores** * Improved example script path handling and GPU initialization tracking. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
4a11486f6c |
Enable multinode training for HF speculative decoding (#922)
## What does this PR do? **Type of change:** New example **Overview:** Modify launch_train.sh script to enable multi-node training. Provide a slurm template script. ## Usage Add required fields in slurm.sh. Then use the command below to submit multi-node job: ```bash bash slurm.sh ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added multi-node training support with new configuration options for distributed setups. * Introduced Slurm batch script for streamlined job submission to cluster environments. * **Improvements** * Enhanced GPU resource management for multi-GPU and distributed training configurations. * Updated speculative decoding model support and validation handling. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> |
||
|
|
eb99488da1 |
Fix: restore requires_grad in transformers5 reloading (#907)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Patch transformers 5.x parameter loading to preserve original `requires_grad` settings. In transformers v5.x, loading a checkpoint forcibly sets parameters' requires_grad, which unintentionally unfreeze frozen parameters (e.g. Base model in eagle training). This leads to optimizer initialization error since the restored optimizer expected more parameter than the checkpoint. This monkey-patch restores the original`requires_grad` after loading parameters. Reference: https://github.com/huggingface/transformers/blob/v5.0.0.rc1-release/src/transformers/core_model_loading.py#L640 ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed model parameter loading in speculative decoding to properly preserve gradient requirements for each parameter when using HuggingFace Transformers 5.x, ensuring correct behavior during checkpoint resumption and model initialization. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
b8a4586702 |
Refactor: Eagle data loading (#668)
## What does this PR do? **Type of change:** Refactor <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Jira ticket: https://jirasw.nvidia.com/browse/OMNIML-2955 Main changes : - Consolidate Eagle data loading with @ChenhanYu 's implementation of `transformers_dataset.py` - Refactor: baked the following logics from `example/main.py` to `modelopt/torch` for cleaner example entrance: - default config selecting and merging with custom config - tokenizer post-processor (chat template and pad_tok_id) - d2t loading - Implementation refactor: In HF workflow, reuse base modfel's input hidden states as input_embedding, instead of calculating from input_ids. This has two main benefits: - Easier VLM support, which has various embedding processing logics. - Training effieicy. - Deprecating eagle1 from the example. It is still available by setting custom config. - Other minor fixes and readme updates. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> Tested that training curves after changes (both online&offline) is identical with original branch: <img width="1073" height="634" alt="image" src="https://github.com/user-attachments/assets/abfd7bea-c82c-48a7-8181-68c5a9e4da8d" /> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added draft vocabulary cache support for EAGLE model training, enabling runtime vocabulary customization via `--draft_vocab_cache` parameter * Introduced new data loading utilities with sharding, streaming, and tokenization support for large-scale training * Added optional `--log_steps` configuration to training launcher * **Documentation** * Updated EAGLE configuration guides with draft vocabulary cache setup instructions and examples * **Refactor** * Restructured data pipeline for offline training with improved dataset handling and batching * Updated command-line arguments across training scripts (`--input-data` replaces `--input-file`) <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
a8f5314c93 |
fix the path change in torch v2.10 for spec dec (#863)
## What does this PR do? **Type of change:** bug fix **Overview:** torch v2.10 changes the path for _SDPAMerger. will need to use the new path for import ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated internal import references to reflect organizational changes in dependencies. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> |
||
|
|
81b67ddf06 |
Context parallelism for Megatron core models (#818)
## What does this PR do? New feature **Overview:** This PR implements the context manager which injects attn_mask as attn_bias to TEDotProductAttention so that we can enable EAGLE training with arbitrary mask. ## Usage set CP>1 in https://github.com/NVIDIA/Megatron-LM/blob/main/examples/post_training/modelopt/finetune.sh ```python # Add a code snippet demonstrating how to use this ``` ## Testing Tested on DSR1 Llama 8B. CP1->CP2 38854MB->28050MB MTbench AL 2.26->2.31 ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## New Features * Context parallelism support added for Eagle speculative decoding with HuggingFace and Megatron Core models. * Model checkpoint loading enhanced to enable remote code execution capabilities when required. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> |
||
|
|
3036a9ea9f |
Feat: Context Parallel for Eagle3 Training (#745)
## What does this PR do?
**Type of change:** New Feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:**
- Supported Context Parallel by patching torch ring attention;
- Require following libirary version for stable cp:
- torch2.8.0
- transformers5.0.0
- accelrate1.12.0
- Move to FSDP2
- Removed unused arguments in training script (`--multi_gpu`,
`fsdp_wrap_layer`)
- Bump CI container to `nvcr.io/nvidia/pytorch:25.08-py3`
## Usage
<!-- You can potentially add a usage example below. -->
```bash
./launch_train.sh --model $MODEL \
--output_dir $OUTPUT_DIR \
--data $DATA \
--num_epochs 0.1 \
--train_bs 1 \
--eagle_config eagle_config.json \
--training_seq_len 1024 \
--cp_size 2 #newly added
```
## Testing
- SDPA level correctness: tested TTT attention with/without CP, diff <
1%
```
=== Compare context-parallel (CP) outputs and grads with non-CP ===
Forward output comparison (CP vs Non-CP):
Absolute diff (adiff) cp_out vs out: 0.001953125
Relative diff (rdiff) cp_out vs out: 0.00182342529296875
WQ (query proj) grad comparison (CP vs Non-CP):
Absolute diff (adiff) cp_wq_grad vs wq_grad: 0.0078125
Relative diff (rdiff) cp_wq_grad vs wq_grad: 0.00347900390625
WK (key proj) grad comparison (CP vs Non-CP):
Absolute diff (adiff) cp_wk_grad vs wk_grad: 0.0078125
Relative diff (rdiff) cp_wk_grad vs wk_grad: 0.002471923828125
WV (value proj) grad comparison (CP vs Non-CP):
Absolute diff (adiff) cp_wv_grad vs wv_grad: 0.25
Relative diff (rdiff) cp_wv_grad vs wv_grad: 0.0069580078125
==============================================================
```
- E2E Training Acc
(Llama3.1-8B, Unsynthesized magpie)
<img width="911" height="630" alt="image"
src="https://github.com/user-attachments/assets/1ecacc7f-c720-494c-9c1b-b60e7ced7baa"
/>
- Peak Mem Reserved
(llama3.1-8B, 8xH100, train_length=4k)
| cp_size | max_memory_allocated(MB) |max_memory_reserved (MB) |
|----|--------------------------|--------------------------|
| 1 | 65040.20 |79018.00
| 2 | 50409.17 |73098.00
| 4 | 45120.92 |72052.00
| 8 | 38882.12 |66484.00
- Max Training Length test
(llama3.1-8B, H100)
| cp_size | 6k | 12k | 24k | 48k |
|--------------------|-----|-----|-----|-----|
| 1 | ✅ | OOM | OOM | OOM |
|2 | ✅ | ✅ | OOM | OOM |
| 4 | ✅ | ✅ | ✅ | OOM |
| 8 | ✅ | ✅ | ✅ | ✅ |
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added context parallelism (CP) and data parallelism shard size
configuration parameters to training arguments.
* **Enhancements**
* Improved TTT attention masking support for speculative decoding
workflows.
* Enhanced training launch script with improved parallelism
configuration handling.
* **Chores**
* Updated core dependencies: torch, transformers, accelerate, and wandb.
* Added FSDP configuration file for distributed training setup.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
3fd8b804f4 |
Fix:add vllm dump script (#710)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** ? Missing one file in last PR:https://github.com/NVIDIA/Model-Optimizer/pull/689/changes ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
2a51bbd7cb |
Optimize calibrate_draft_vocab to read only required lines when calib… (#618)
Optimize calibrate_draft_vocab to read only required lines when
calibrate_size is set
## What does this PR do?
**Type of change:** Performance improvement
**Overview:**
This PR optimizes the
[calibrate_draft_vocab.py](cci:7://file:///Users/obenshoham/PycharmProjects/TensorRT-Model-Optimizer/examples/speculative_decoding/scripts/calibrate_draft_vocab.py:0:0-0:0)
script to improve memory efficiency and I/O performance when using the
`--calibrate_size` parameter. Previously, the script would read all
lines from the data file into memory before slicing to the specified
`calibrate_size`, causing unnecessary resource usage for large datasets.
The optimization uses `itertools.islice` to read only the required
number of lines when `calibrate_size` is specified.
## Usage
The script usage remains unchanged. When using `--calibrate_size`, the
script now only reads the specified number of lines instead of loading
the entire dataset:
```bash
# Only reads first 1000 lines from the dataset (optimized)
python scripts/calibrate_draft_vocab.py \
--model meta-llama/Llama-3.2-1B-Instruct \
--data input_conversations/daring-anteater.jsonl \
--draft_vocab_size 32000 \
--calibrate_size 1000 \
--save_dir draft_vocab_cache
Signed-off-by: Ofir Ben Shoham <ofir_benshoham@intuit.com>
|
||
|
|
bdd10c2dbe |
Feat: MLA eagle (#689)
## What does this PR do?
**Type of change:** New Feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:**
- Add MLA Eagle support
- Add new argument "eagle_decoder_type" to switch between llama and
kimik2 eagle;
- Add patches to load from kimik2 model implementations dynamically;
- new default config for kimi k2;
- Refactor eagle export to support multilayer/multitype eagle export
concisely;
- Rename some modules for simplified export logic;
- Other minor improvements;
## Usage
<!-- You can potentially add a usage example below. -->
```python
# Add a code snippet demonstrating how to use this
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
- Tested that kimi k2 thinking works with eagle_type=kimik2:
<img width="1068" height="636" alt="image"
src="https://github.com/user-attachments/assets/5557ef87-c719-4fb1-be18-30435f6b3885"
/>
- Tested that llama 3.2 1b works with eagle_type=llama:
<img width="1066" height="634" alt="image"
src="https://github.com/user-attachments/assets/633c575c-cc79-43af-aed3-0378a303ebc7"
/>
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Signed-off-by: yeyu-nvidia <yeyu@nvidia.com>
Co-authored-by: yeyu-nvidia <yeyu@nvidia.com>
|