Commit Graph
29 Commits
Author SHA1 Message Date
ZhiyuandClaude Opus 4.8 f335459dc0 refactor(examples): rename llm_ptq → hf_ptq (symlink for back-compat) (#1759)
## What does this PR do?

**Type of change:** refactor / deprecation (examples)

Follow-up to #1705 (which consolidated `examples/vlm_ptq` into
`examples/llm_ptq`). Since that example now covers Hugging Face **LLM
and VLM** PTQ, the `llm_ptq` name is a misnomer. This renames the
directory to `examples/hf_ptq` and leaves a relative symlink
`examples/llm_ptq → hf_ptq` so existing paths/commands keep working
during a deprecation window.

Requested by @kevalmorabia97 on #1705 (with the symlink-for-back-compat
approach), targeted for the **same 0.46 release** as the consolidation.

### Changes
- `git mv examples/llm_ptq → examples/hf_ptq` and
`tests/examples/llm_ptq → tests/examples/hf_ptq` (the CI runner maps the
matrix name to both `examples/<name>` and `tests/examples/<name>`).
- Add a tracked back-compat symlink `examples/llm_ptq → hf_ptq`.
- Update CI matrices and all repo **path references** (docs, READMEs,
agent skills, launcher/debugger tools, tests) from `llm_ptq` to
`hf_ptq`.
- Keep Python identifiers / test-util module names
(`run_llm_ptq_command`, `llm_ptq_utils`) — they name the LLM-PTQ task,
not the directory.
- Preserve the CODEOWNERS team slug
(`modelopt-examples-llm_ptq-codeowners`) and historical CHANGELOG
entries; add a CHANGELOG deprecation note.

### Back-compat caveats (inherent to git directory symlinks)
- ✅ Linux/macOS CLI usage and Python `cwd`/pytest resolution work
through the symlink.
- ⚠️ Windows git checkouts don't materialize symlinks by default (low
impact — this example is Linux-only in practice).
- ⚠️ GitHub web doesn't follow directory symlinks, so legacy external
deep-links to `examples/llm_ptq/...` won't navigate in. All **internal**
references are repointed to `hf_ptq`, so the symlink is only for legacy
external/CLI use.

### Usage (unchanged via symlink)
```bash
# New canonical path
cd examples/hf_ptq
scripts/huggingface_example.sh --model <hf_model> --quant fp8

# Old path still works (forwards via symlink)
cd examples/llm_ptq && scripts/huggingface_example.sh --model <hf_model> --quant fp8
```

### Testing
- `bash -n` on moved/edited shell scripts (new path + via symlink).
- `py_compile` on moved/edited Python; test re-export shim repointed to
`examples/hf_ptq/example_utils`.
- Verified git tracks `examples/llm_ptq` as a single symlink (mode
120000), not a duplicated tree (no pre-commit / pytest
double-processing).
- `pre-commit run` on all changed files passes.

### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (relative symlink keeps
`examples/llm_ptq` paths valid; see caveats above)
- Did you write any new necessary tests?: N/A (pure rename; existing
tests moved with the dir)
- Did you update Changelog?: ✅

### Additional Information
Follow-up (later release): remove the `examples/llm_ptq` symlink once
external references have migrated.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* PTQ guidance now directs to the unified Hugging Face PTQ flow,
including VLM quantization via the shared `--vlm` entry point.
* **Documentation**
* Updated README and guide links, references, and command snippets to
use `hf_ptq` (replacing `llm_ptq`).
* Deprecated and consolidated `vlm_ptq` into `hf_ptq`; removed
VILA/NVILA coverage from the Hugging Face PTQ examples.
* **Bug Fixes**
* Improved detection and routing so local/manual setup uses the correct
PTQ source.
* **Tests / Chores**
  * CI and example tests updated to run the `hf_ptq` variants.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-27 08:48:48 +00:00
Keval Morabia b6bf6b7997 Update Roadmap Issue link
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-06-23 21:42:37 +05:30
jzh26andh-guo18 48f8d8984f Fix conversation loading logic in UltraChat dataset (#1680)
### What does this PR do?
Type of change: Bug fix.

<!-- Details about the change. -->
Previously, only the first user prompt was extracted from each example,
discarding all subsequent turns. UltraChat stores full multi-turn
conversations in the "messages" field, so switching to that field
preserves the complete dialogue rather than truncating to a single user
message.
### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
 python make_dataset.py -f test_cfg.yaml --full
 python make_dataset.py -f test_cfg.yaml
```yaml
      - name: "ultrachat"
        splits:
          train_gen: 100
          train_sft: 100
```
### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Breaking Changes**
* Removed UltraChat support from the example dataset mixer, so UltraChat
splits can no longer be loaded.

* **Documentation**
* Updated the dataset examples README and the example dataset
configuration to remove UltraChat and adjust split settings for other
datasets.
* Updated the speculative decoding fine-tuning documentation to
reference Daring-Anteater instead of UltraChat.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: jzh26 <226629529+jzh26@users.noreply.github.com>
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-06-22 09:18:35 +00:00
h-guo18 977d34dc3c [Fix](nvbug6304585): specdec README online base-model example should use Instruct model (#1755)
### What does this PR do?

Type of change: Bug fix

The **Training Draft Model with Online/Offline Base Model** examples in
`examples/speculative_decoding/README.md` used
`meta-llama/Llama-3.2-1B`, a
base / pretrained checkpoint that ships **no chat template**. The online
EAGLE3
flow tokenizes conversations through
`tokenizer.apply_chat_template(...)`, so
the data collator fails fast at startup:

```
ValueError: No valid chat template!
```

This PR:

- Switches both README example commands (online and offline) to
  `meta-llama/Llama-3.2-1B-Instruct`, which carries a chat template.
- Makes the collator error message in
`modelopt/torch/utils/plugins/transformers_dataset.py` actionable — it
now
explains the cause (base checkpoints have no chat template) and points
users
  at an Instruct model or a custom `chat_template`.

### Usage

```bash
./launch_train.sh \
    --config ../../modelopt_recipes/general/speculative_decoding/eagle3.yaml \
    model.model_name_or_path=meta-llama/Llama-3.2-1B-Instruct \
    data.data_path=input_conversations/train.jsonl \
    training.output_dir=ckpts/llama-3.2-1b-online
```

### Testing

- Reproduced the original `No valid chat template!` failure with the
base
`Llama-3.2-1B` and confirmed the Instruct variant carries a chat
template.
- Verified the new error message renders correctly when a template is
missing.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌

### Additional Information

Fixes nvbug 6304585: https://nvbugspro.nvidia.com/bug/6304585


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Documentation**
* Updated speculative decoding example training commands to reference
the Llama-3.2-1B-Instruct model.

* **Bug Fixes**
* Enhanced error message when chat template configuration is missing,
providing actionable guidance for resolution.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-06-16 16:42:20 -07:00
h-guo18 46eddab877 [Feat]: Specdec Streaming: RDMA + Multinode (#1611)
### What does this PR do?

Type of change: New feature

Multi-node **streaming** training for speculative decoding (EAGLE3 /
DFlash):
a live `vllm serve` captures the target model's hidden states and moves
them
straight to the trainer over **NIXL RDMA** — no disk round-trip. The
streaming
dataset is map-style — each rank fetches only its own
`DistributedSampler` shard
(concurrency from `dataloader_num_workers`), round-robins across
multiple serve
replicas (`server_urls`), and scales to multi-node DDP. Serve-side
tensor
parallelism (TP>1) is supported: hidden states are replicated across TP
ranks, so
rank 0 alone owns the pool + transfer.

### How

- `RdmaHiddenStatesConnector` — out-of-tree vLLM connector (no vLLM
source edits):
one pre-registered pinned NIXL pool per serve, a ring slot per request,
and a
small HTTP sidecar serving transfer metadata. The trainer RDMA-READs the
slot
  into a per-worker buffer. RDMA is the **only** transport (the earlier
  disk/safetensors path is removed).
- Map-style dataset + multi-node accelerate launch (`--machine_rank`,
optional
  Slurm `--segment` to keep nodes in one NVLink domain).

### Usage

```yaml
data:
  mode: streaming
  streaming_server_url: "http://node0:8000,http://node1:8000"  # round-robin
```

### Validation (Qwen3-8B, oci-nrt H100)
sandbox CI:
https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/337489812

**1. End-to-end convergence — EAGLE3 & DFlash, 5000 steps.** Both
algorithms
converge and export a deployable draft; the DFlash drafts also serve
under vLLM
speculative decoding (8/8 smoke prompts pass).

| algorithm | topology (nodes) | train loss (step 0 → 5000) | vLLM draft
acc-len |
|---|---|---|---|
| EAGLE3 | 2 serve TP=2 + 2 trainer DDP (4) | 37.1 → 8.20 | — |
| DFlash | 1 serve TP=1 + 1 trainer (2)      | 11.7 → 5.56 | 1.11 |
| DFlash | 2 serve TP=2 + 2 trainer DDP (4)  | 10.9 → 5.26 | 1.19 |

<!-- Drag these PNGs in here (GitHub turns them into asset URLs):
eagle3_streaming_loss.png, dflash_streaming_loss_singlenode.png,
dflash_streaming_loss_multinode.png -->

**2. Scalability — 1 → 12 nodes (EAGLE3, 200 steps).** Throughput scales
~23× across the
sweep below. The step-time growth is cross-node DDP all-reduce, not the
streaming path —
RDMA (~0.33 ms/req @ 2 MB, ~47 GB/s host-pinned READ) is never the
bottleneck. Scale
serve + trainer nodes together for near-linear speedup.

| serve / trainer | nodes | step time | samples / step | samples / sec
(global) | acc @ step 200 |
|---|---|---|---|---|---|
| 1 serve / 1 rank (co-located, 1 node 2 GPU) | 1 | 0.23 s | 1 | 4.4 |
[0.141, 0.094, 0.072] |
| 1 serve / 1 rank (cross-node) | 2 | 0.23 s | 1 | 4.3 | [0.137, 0.105,
0.074] |
| 2 serve / 8 ranks | 3 | 0.26 s | 8 | 31.1 | [0.215, 0.126, 0.097] |
| 4 serve / 16 ranks (2 trainer nodes) | 6 | 0.28 s | 16 | 56.5 |
[0.217, 0.148, 0.110] |
| 8 serve / 32 ranks (4 trainer nodes) | 12 | 0.31 s | 32 | 101.9 |
[0.235, 0.165, 0.137] |

**3. Serve-side TP correctness.** TP=1 vs TP=2 draft top-1 accuracy
track
step-for-step (hidden states are replicated across TP ranks).

<img width="910" height="546" alt="serve-tp-acc"
src="https://github.com/user-attachments/assets/73df9214-7ff0-4ab4-bf2f-95842b12cd5f"
/>

### Before your PR is "*Ready for review*"

- Backward compatible?: ❌ — streaming is now RDMA-only; `server_url` →
`server_urls`;
the disk transport (`HS_TRANSPORT`, `streaming_shared_storage_path`) is
removed.
- New tests?: ✅
`tests/unit/torch/speculative/plugins/test_hf_streaming_dataset.py`
  (map-style dataset + mocked RDMA fetch).
- Updated Changelog?: ❌
- Claude approval?: ❌

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-06-10 18:35:02 -07:00
Chenhan D. YuandClaude Opus 4.6 3131195241 add: DFlash block diffusion speculative decoding (#1211)
DFlash (Block Diffusion for Flash Speculative Decoding) predicts an
entire block of tokens in a single forward pass using masked parallel
prediction with KV injection from the target model's hidden states.

Key features:
- Feature fusion (multi-layer hidden states -> FC + RMSNorm)
- KV injection (fused features as K/V in every draft layer with QK-norm)
- Random anchor sampling with bidirectional intra-block attention
- Logit distillation with exponential loss decay (gamma weighting)
- Multi-node DDP training with checkpoint resume
- Export to z-lab compatible HF format
- Online validation (context-dependent ground truth)

Training recipe:
modelopt_recipes/general/speculative_decoding/dflash.yaml
Results: examples/speculative_decoding/doc/dflash_results.md

### ModelOpt Eval (online validation, osl=512)

| Dataset | z-lab | ModelOpt (306K) | Diff |
|---------|-------|-----------------|------|
| gsm8k | 4.10 | **5.19** | **+1.09** |
| MT-Bench | 3.58 | **4.36** | **+0.78** |

### z-lab Official Eval (dflash.benchmark, osl=512)

| Dataset | z-lab | ModelOpt (306K) | Diff |
|---------|-------|-----------------|------|
| gsm8k | **5.00** | 4.08 | -0.92 |
| MT-Bench | **3.28** | 2.99 | -0.29 |

> z-lab model trained with block_size=16. ModelOpt trained with
block_size=8.

## Evaluation Method Impact (gsm8k)

| Eval Method | z-lab checkpoint | ModelOpt (306K) |
|-------------|-----------------|-----------------|
| Fixed GT (ModelOpt eval) | 2.95 | 4.23 |
| Online GT (ModelOpt eval) | 4.10 | **5.19** |
| z-lab official eval | **5.00** | 4.08 |

### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added DFlash speculative decoding mode with parallel block prediction
support.
* Included training launchers and MT-Bench evaluation scripts for DFlash
models.
* Added online acceptance rate validation for improved inference
verification.

* **Documentation**
* DFlash quick start guide with configuration parameters and training
examples.
  * Performance results and benchmarks for DFlash-trained models.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 15:58:39 -07:00
Chenhan D. Yu 0246041b01 feat(speculative): add vLLM data synthesis pipeline and Nemotron dataset preparation scripts (#1176)
### What does this PR do?

Type of change: New feature, new example, bug fix

Adds a vLLM-based synthetic data generation pipeline for speculative
decoding draft model training, along with dataset preparation scripts
for NVIDIA's Nemotron Post-Training dataset collections.

**Data synthesis pipeline** (`tools/launcher/common/vllm/query.sh` +
`common/query.py`):
- Launch a vLLM server and run multi-turn inference to synthesize
training data from input conversation skeletons
- Fork-safe OpenAI client: reinitializes HTTP connection pool after
`datasets.map()` forks worker processes, preventing 400 errors from
corrupted connections
- Clear Docker `ENTRYPOINT` so vLLM containers (which default to `vllm
serve`) work correctly under NeMo Run's executor
- `--max-tokens` argument to bound generation length
- Local file loading support (`--data /path/to/file.jsonl`)
- Re-raise connection errors so `datasets.map()` halts the shard instead
of silently producing empty rows
- Map `developer` role to `system` (OpenAI format compatibility)

**Multi-turn reasoning trace handling** (`common/query.py`):
- Strip `<think>...</think>` blocks from intermediate assistant turns
before re-feeding to the model; preserve the full trace only on the
final turn

**Nemotron dataset preparation** (`examples/dataset/`):
- `make_nemotron_ptv2_dataset.py` — prepares
[nvidia/Nemotron-Post-Training-Dataset-v2](https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2)
(~3.3M rows generate, ~1.9M rows train)
- `make_nemotron_ptv3_dataset.py` — prepares the [Nemotron PTv3
collection](https://huggingface.co/collections/nvidia/nemotron-post-training-v3)
of 16 datasets (~3.4M rows generate, ~3.9M rows train)
- Both support `generate` mode (strips assistant turns for synthesis
input) and `train` mode (normalizes to clean OpenAI format for SFT)
- `conversation_utils.py` — shared utilities: `strip_assistant_turns`,
`normalize_messages`, `make_augment_fn`, `AugmentationSpec`
- `augmentations.yaml` — 12 language-redirect variants + style/format
hints, cycled across dataset rows
- Scripts live in `examples/dataset/` (not under
`speculative_decoding/`) to signal reusability beyond speculative
decoding

**Bug fixes**:
- `strip_assistant_turns()`: return `{"messages": []}` when no user
turns remain (system-only rows were previously passed through instead of
being filtered)
- `concatenate_datasets()`: guard against empty parts list
- SSH tunnel user precedence: explicit `user` arg now correctly
overrides `slurm_config.user`

### Usage

```bash
# Prepare PTv3 input conversations for synthesis (~3.4M rows):
python examples/dataset/make_nemotron_ptv3_dataset.py --output-dir /tmp/ptv3_gen

# Launch vLLM server + synthesize responses:
bash tools/launcher/common/vllm/query.sh \
    --model /path/to/model \
    --tensor-parallel-size 4 \
    -- \
    --data /tmp/ptv3_gen/default.jsonl \
    --save /tmp/ptv3_responses \
    --num-shards 10 --num-proc 4 --max-tokens 4096

# Prepare PTv2 for direct SFT training (~1.9M rows):
python examples/dataset/make_nemotron_ptv2_dataset.py --mode train --output-dir /tmp/ptv2_train
```

### Testing

Tested end-to-end on an NVIDIA GB10 node (119 GiB GPU memory) with
`vllm/vllm-openai:qwen3_5-cu130` container and `Qwen/Qwen3.5-4B`:
- vLLM server starts correctly with cleared Docker entrypoint
- `datasets.map(num_proc=4)` runs without connection errors (fork-safe
client)
- Multi-turn synthesis produces correct assistant responses with
thinking traces handled

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (data synthesis scripts;
tested manually)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added dataset generation and augmentation capabilities for Nemotron
post-training datasets (v2 and v3)
* Enhanced query functionality with thinking-block filtering and
improved client management for robust parallel processing
* Added support for local dataset file paths alongside HuggingFace Hub
datasets

* **Bug Fixes**
* Fixed SLURM executor user resolution and Docker container entrypoint
configuration
* Improved error handling for connection failures during dataset
synthesis

* **Documentation**
* Updated dataset preparation guide with new generation modes and
augmentation configuration details

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan D. Yu <5185878+ChenhanYu@users.noreply.github.com>
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-04-08 05:22:11 +00:00
h-guo18 82d96a635f [Speculative Decoding] Refactor EAGLE3 training to YAML-based config and recipe system (#1134)
## What does this PR do?

Refactors EAGLE3 training to use a single base YAML config with
OmegaConf dotlist overrides.

**Type of change:** Refactor

## Changes

- Single base config
`modelopt_recipes/speculative_decoding/_base_eagle3.yaml` for all EAGLE3
training; removed per-model child YAMLs.
- `launch_train.sh` accepts `--config <yaml>` plus dotlist overrides
(e.g. `model.model_name_or_path=xxx`).
- Removed `__base__` YAML inheritance logic from `main.py`.
- `dp_shard_size` default changed from `0` sentinel to `None` for
clarity.
- Removed `eagle_config.json` and `fsdp_config.json`; architecture
config is now nested under `eagle.eagle_architecture_config` in YAML.
- `train_eagle3_and_export.sh` now uses base YAML + dotlist instead of
generating a temporary YAML.
- Updated README and tests accordingly.

## Usage

```bash
# Online training
./launch_train.sh \
    --config ../../modelopt_recipes/speculative_decoding/_base_eagle3.yaml \
    model.model_name_or_path=meta-llama/Llama-3.2-1B \
    data.data_path=input_conversations/train.jsonl \
    training.output_dir=ckpts/llama-3.2-1b-online

# Offline training
./launch_train.sh \
    --config ../../modelopt_recipes/speculative_decoding/_base_eagle3.yaml \
    model.model_name_or_path=meta-llama/Llama-3.2-1B \
    data.offline_data_path=$HIDDEN_STATES_DIR \
    training.output_dir=ckpts/llama-3.2-1b-offline
```

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-04-08 01:52:13 +00:00
Benjamin Chislett 4292505512 Refactor: Clean up EAGLE training dataset preparation (#684)
## What does this PR do?

**Type of change:** Refactor

**Overview:** 
- Consolidate input dataset preparation into `make_dataset.py`
- Read dataset mix spec from a YAML file
- - Can now specify how many samples to take from each split
- - Can no longer easily split a dataset into train/test sections. I
don't think this feature was really useful to begin with. Most datasets
can already be separated into train/val/test at the split level, and
those that can't are usually going to be splitted by the training FW
anyways.
- Add support for a few new dataset types, magpie 300k/500k/1M, nemotron
post-training dataset v2.

## Usage
See README for detailed example

## Testing
Ran it locally on all dataset modes, works successfully and output looks
good. Checked shuffling, conversation IDs, and output contents were all
unique and usable.

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated speculative decoding example documentation with new dataset
references and standardized file paths.

* **New Features**
* Introduced configuration-driven dataset preparation supporting
multiple dataset sources with centralized configuration files.

* **Refactor**
* Simplified dataset preparation workflow with unified tooling and
updated default data paths throughout the training pipeline.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
2026-03-18 09:29:49 -07:00
h-guo18 b8a4586702 Refactor: Eagle data loading (#668)
## What does this PR do?

**Type of change:** Refactor <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** 
Jira ticket: https://jirasw.nvidia.com/browse/OMNIML-2955

Main changes :

- Consolidate Eagle data loading with @ChenhanYu 's implementation of
`transformers_dataset.py`

- Refactor: baked the following logics from `example/main.py` to
`modelopt/torch` for cleaner example entrance:
  - default config selecting and merging with custom config
  - tokenizer post-processor (chat template and pad_tok_id)
  - d2t loading 
- Implementation refactor: In HF workflow, reuse base modfel's input
hidden states as input_embedding, instead of calculating from input_ids.
This has two main benefits:
    - Easier VLM support, which has various embedding processing logics.
    - Training effieicy.
    
- Deprecating eagle1 from the example. It is still available by setting
custom config.
  
 - Other minor fixes and readme updates.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

Tested that training curves after changes (both online&offline) is
identical with original branch:
<img width="1073" height="634" alt="image"
src="https://github.com/user-attachments/assets/abfd7bea-c82c-48a7-8181-68c5a9e4da8d"
/>


## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added draft vocabulary cache support for EAGLE model training,
enabling runtime vocabulary customization via `--draft_vocab_cache`
parameter
* Introduced new data loading utilities with sharding, streaming, and
tokenization support for large-scale training
  * Added optional `--log_steps` configuration to training launcher

* **Documentation**
* Updated EAGLE configuration guides with draft vocabulary cache setup
instructions and examples

* **Refactor**
* Restructured data pipeline for offline training with improved dataset
handling and batching
* Updated command-line arguments across training scripts (`--input-data`
replaces `--input-file`)

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-02-18 14:04:06 -08:00
h-guo18 3036a9ea9f Feat: Context Parallel for Eagle3 Training (#745)
## What does this PR do?

**Type of change:** New Feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

- Supported Context Parallel by patching torch ring attention;
- Require following libirary version for stable cp: 
  - torch2.8.0
  - transformers5.0.0
  - accelrate1.12.0 
 - Move to FSDP2 
- Removed unused arguments in training script (`--multi_gpu`,
`fsdp_wrap_layer`)
 - Bump CI container to `nvcr.io/nvidia/pytorch:25.08-py3`

## Usage
<!-- You can potentially add a usage example below. -->

```bash
./launch_train.sh --model $MODEL \
            --output_dir $OUTPUT_DIR \  
            --data $DATA \
            --num_epochs 0.1 \
            --train_bs 1 \
            --eagle_config eagle_config.json \
            --training_seq_len 1024 \
            --cp_size 2   #newly added
```

## Testing
- SDPA level correctness: tested TTT attention with/without CP, diff <
1%
```
=== Compare context-parallel (CP) outputs and grads with non-CP ===
Forward output comparison (CP vs Non-CP):
  Absolute diff (adiff) cp_out vs out: 0.001953125
  Relative diff (rdiff) cp_out vs out: 0.00182342529296875
WQ (query proj) grad comparison (CP vs Non-CP):
  Absolute diff (adiff) cp_wq_grad vs wq_grad: 0.0078125
  Relative diff (rdiff) cp_wq_grad vs wq_grad: 0.00347900390625
WK (key proj) grad comparison (CP vs Non-CP):
  Absolute diff (adiff) cp_wk_grad vs wk_grad: 0.0078125
  Relative diff (rdiff) cp_wk_grad vs wk_grad: 0.002471923828125
WV (value proj) grad comparison (CP vs Non-CP):
  Absolute diff (adiff) cp_wv_grad vs wv_grad: 0.25
  Relative diff (rdiff) cp_wv_grad vs wv_grad: 0.0069580078125
==============================================================
```

- E2E Training Acc
  (Llama3.1-8B, Unsynthesized magpie)
<img width="911" height="630" alt="image"
src="https://github.com/user-attachments/assets/1ecacc7f-c720-494c-9c1b-b60e7ced7baa"
/>

- Peak Mem Reserved
   (llama3.1-8B, 8xH100, train_length=4k)

    | cp_size | max_memory_allocated(MB) |max_memory_reserved (MB) |
    |----|--------------------------|--------------------------|
    | 1  |         65040.20              |79018.00
    | 2  |           50409.17             |73098.00
    | 4  |              45120.92            |72052.00
    | 8  |              38882.12            |66484.00

- Max Training Length test
  (llama3.1-8B, H100)

  | cp_size               | 6k  | 12k | 24k  | 48k  |
  |--------------------|-----|-----|-----|-----|
  | 1         | ✅ | OOM | OOM | OOM  |
  |2      | ✅  | ✅ | OOM | OOM |
  | 4         | ✅ | ✅ | ✅  | OOM  |
  | 8         | ✅ | ✅ | ✅  | ✅  |

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added context parallelism (CP) and data parallelism shard size
configuration parameters to training arguments.

* **Enhancements**
* Improved TTT attention masking support for speculative decoding
workflows.
* Enhanced training launch script with improved parallelism
configuration handling.

* **Chores**
* Updated core dependencies: torch, transformers, accelerate, and wandb.
  * Added FSDP configuration file for distributed training setup.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-01-24 02:45:50 +00:00
Keval Morabia 53a2ddebab Product Rename: TensorRT Model Optimizer to Model Optimizer (#583)
- [x] Product Rename: TensorRT Model Optimizer to Model Optimizer
(OMNIML-3033)
- [x] Mention in Latest News section with date on the date of merging
this PR (12/08)

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-12-07 12:33:21 +05:30
h-guo18 bc52b6cf12 Feat: Support VLLM one-model eagle ckpt; Add unit tests; (#573)
## What does this PR do?

**Type of change:** New feature, new tests; <!-- Use one of the
following: Bug fix, new feature, new example, new tests, documentation.
-->

**Overview:** 
- Add conversion scripts for eagle3 llm-compressor style checkpoint
  - Jira Ticket: https://jirasw.nvidia.com/browse/OMNIML-2866

- Add unit tests for `ar_validate.py`, `export_hf_checkpoint.py`, and
`convert_to_vllm_ckpt.py`.

## Usage
<!-- You can potentially add a usage example below. -->

```python
python scripts/convert_to_vllm_ckpt.py --input <eagle3 ckpt> --verifier <base model> --output <path>
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2025-11-19 14:26:39 -08:00
Izzy PuttermanandKeval Morabia 5adb9ba0fc Add Spec dec Bench example (#474)
## What does this PR do?

**Type of change:** New feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** Specdec bench example

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Izzy Putterman <iputterman@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-11-06 21:21:26 +00:00
Keval Morabia 90b1e68fcf Update HF nvidia collection links in docs (#475)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-10-28 22:23:01 +05:30
h-guo18 ff8a1ed126 fix:eagle3 offline (#456)
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2025-10-21 14:01:47 -07:00
h-guo18 557633c986 Fix: supporting gpt-oss HF eagle (#398)
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2025-10-09 01:14:00 +00:00
h-guo18 abed33c3f8 Feat: TRTLLM Dumper for Eagle Offline Training (#404)
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2025-10-07 20:24:14 -07:00
h-guo18 4ff8fc9022 Example: add offline eagle training commands to README (#366)
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2025-09-24 16:41:42 -07:00
Keval Morabia c0590b0255 Deprecate ModelOpt custom docker and directly use TRT-LLM / PyTorch / TRT docker (#346)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-09-23 01:43:25 +05:30
h-guo18 a6fa34cda4 Feat: update eagle3 example; add export (#293)
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2025-09-08 18:02:41 -07:00
Keval Morabia 1ef1d72a1b Code quality improvements - typos, formatting, etc.
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-09-02 19:58:33 +05:30
Keval Morabia 4d1eb0caf5 Major improvement of READMEs and documentation
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-08-30 10:54:28 +05:30
Keval Morabia 4c611e47a6 Update files on GitHub
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-07-31 22:48:09 +05:30
Keval Morabia 7af33d29ce Update for 0.31.0 release 2025-06-05 13:24:07 -07:00
Keval Morabia 7a047435ae Update for 0.29.0 release 2025-05-08 23:43:51 +05:30
Keval Morabia 92f430f6ab Add files for 0.27.0 release 2025-04-03 10:31:41 +05:30
Keval Morabia 2017cd9063 Update for 0.25.0 release 2025-03-03 22:54:22 +05:30
Keval Morabia 73d6af785f Update 0.23.0 - OSS release 2025-01-29 01:45:47 +05:30