mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
c2aaa44f6040658a21a2f5d2213c10ecc6542512
53
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c2aaa44f60 |
[Speculative Decoding] DFlash2 draft variant (grouped sublayer convolution + candidate selector) (#2216)
### What does this PR do? Type of change: new feature Adds **DFlash2** ([blog](https://inco.ai/blog/dflash2/)) as a draft variant of the existing DFlash mode, selected with `dflash_architecture_config.projector_type="dflash2"` alongside `domino`, `dspark` and `lilicorr`. DFlash2 keeps DFlash's one-pass parallel backbone and adds two components that recover the acceptance a purely parallel draft loses: - **Grouped dynamic depthwise convolution** around every attention and MLP sublayer, giving each block position a view of its predecessors *inside* the block. Taps do not cross the block boundary, so the draft stays one forward pass. - **Low-rank candidate selector** scoring transitions between adjacent positions' top-k candidates, so serving walks one coherent path instead of taking an independent argmax per position. Both start as exact no-ops — the convolution's `base_kernel` is an identity and `kernel_projection` is zeroed; the selector's `successor_codebook` is zeroed — so a freshly built DFlash2 draft *is* its DFlash backbone, and enabling the variant is an extension rather than a perturbation. This matches the reference implementation ([SpecForge#772](https://github.com/sgl-project/SpecForge/pull/772), merged) and the way `modeling_lilicorr` already installs this same convolution class. **This unblocks a recipe already shipped on `main`.** `modeling_lilicorr._install_sublayer_convs` imports `DFlashGroupedConv` from `modeling_dflash2`, so `modelopt_recipes/general/speculative_decoding/lilicorr_conv.yaml` raises at model build today and its CHANGELOG entry documents a feature that cannot run. Landing this makes it runnable. Module and parameter names match the SGLang/vLLM `DFlash2DraftModel` loaders. Verified against the released `z-lab/Qwen3.8-27B-DFlash2` checkpoint: **81 tensors, 21 name patterns, zero difference in either direction**. The serving side, [vllm-project/vllm#52816](https://github.com/vllm-project/vllm/pull/52816), has since merged (`b389ac294`) with no change to the checkpoint contract. ### Usage ```bash python examples/speculative_decoding/main.py \ --config modelopt_recipes/general/speculative_decoding/dflash2.yaml \ model.model_name_or_path=Qwen/Qwen3-8B \ data.data_path=<corpus>.jsonl \ training.output_dir=<out> ``` ```yaml # modelopt_recipes/general/speculative_decoding/dflash2.yaml dflash: dflash_selector_loss_alpha: 1.0 # weight of the candidate-selector CE term dflash_architecture_config: projector_type: dflash2 conv_kernel_size: 2 # taps; must not exceed the block size conv_group_size: 16 # must divide hidden_size selector_rank: 256 selector_top_k: 16 ``` ### Testing <img width="2000" height="1320" alt="image" src="https://github.com/user-attachments/assets/869004d1-92c6-41a4-a03d-fd824a06255c" /> **Unit** — 25 CPU tests in `tests/unit/torch/speculative/plugins/test_hf_dflash2.py`; the full `tests/unit/torch/speculative/` suite passes with no regressions. The ones worth keeping pin invariants that a decreasing loss does not catch: - the convolution is an exact identity on the **default** construction, and its taps stay inside the block while a position still sees its predecessors; - the block-offset contract shared by the training objective and `CandidateSelector.greedy_path` — a misaligned objective still converges; - the export fields the vLLM loader requires, including the top-level `block_size` that `DFlash2Exporter` derives the nested copy from; - which selector factors receive gradient on the first step. `successor_codebook` starts at zero, so `predecessor_codebook` and `hidden_projection` take one step to begin moving. That is a warm start, not a dead branch, and both sides are asserted. **End-to-end** — trained on Qwen3-8B against a plain DFlash control with every other argument identical (plot above). Monotonic convergence, no NaN/divergence, no DDP unused-parameter issues. Note the losses are **not comparable** across arms: DFlash2's includes the selector CE term. **Serving (vLLM)** — the exported drafter loads and drafts under the merged DFlash2 path (`RESOLVED draft architectures: ['DFlash2DraftModel']`). Two notes for anyone reproducing: vLLM sizes the convolution from `1 + num_speculative_tokens` at runtime rather than from the checkpoint, so a `block_size=16` drafter is only correct at `num_speculative_tokens=15`; and at that value the upstream path currently hits an illegal memory access in `_cache_draft_logits` ([vllm#55279](https://github.com/vllm-project/vllm/issues/55279)), independent of which checkpoint is used. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — additive. New `projector_type`, its own registry and exporter, one new config field; DFlash / Domino / DSpark / LiLiCorr numerics and `state_dict` contents are untouched. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — `modeling_dflash2.py` is adapted from [SpecForge#772](https://github.com/sgl-project/SpecForge/pull/772) and carries its MIT notice. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — under `0.48.0`. - Did you get Claude approval on this PR?: ✅ — run on 2026-08-20; all review threads addressed and resolved. ### Additional Information Rebased onto current `main`. Two commits from the original branch were dropped because [#2342](https://github.com/NVIDIA/Model-Optimizer/pull/2342) landed them first, with authorship preserved: the no-op sublayer seam in `modeling_dflash.py`, and the `rope_theta`/`rope_parameters` fix — `main`'s version of the latter is stricter, so this PR no longer touches `hf_dflash.py` at all. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added DFlash2 speculative decoding with grouped dynamic convolutions and low-rank candidate selection. * Added configurable selector-loss weighting, including an option to disable it. * Added DFlash2 model conversion and export support. * Added checkpoints compatible with SGLang and vLLM DFlash2 serving. * **Documentation** * Added training recipes and a Qwen3-8B online DFlash2 training configuration. * **Tests** * Added coverage for conversion, training, metrics, gradients, and export compatibility. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
23355eda90 |
fix(deps): declare httpx, unbreaking partial-install (torch) for every PR (#2547)
## Summary `partial-install (torch)` has been failing on **every** PR since 2026-09-24 — including PRs whose branches predate the breakage — and because it is a *collection* error rather than a test failure, it aborts the entire run: ``` ImportError while importing test module '.../tests/unit/torch/speculative/plugins/test_hf_streaming_dataset.py' E ModuleNotFoundError: No module named 'httpx' collected 2243 items / 1 error / 45 skipped !!!!!!!!!!!!!!!!!!!! Interrupted: 1 error during collection !!!!!!!!!!!!!!!!!!!! ``` `unit-pr-required-check` aggregates it, so nothing currently merges on a fresh run. ## What happened **No code changed.** `modelopt/torch/speculative/plugins/hf_streaming_dataset.py` has imported `httpx` at module scope since #1509 (2026-06-02), and `httpx` has never appeared in `pyproject.toml`. It arrived only transitively: `dev-test` → `timm` → `huggingface_hub` → `httpx`. **huggingface_hub 2.0.0**, published **2026-09-24T12:01:21Z**, replaced `httpx<1,>=0.23.0` with the separate **`httpx2<3,>=2.0.0`** distribution. Different package name, so `httpx` stopped being installed and the chain disappeared. The boundary is exact — every run *created* before that timestamp passes, every one after fails: | PR | run created | result | |---|---|---| | #2536 / #2535 | 09-23 22:02 | pass | | #2500 | 09-23 23:45 | pass — **merged 09-24 20:01 on this stale-green result** | | *hub 1.33.0 (still requires httpx)* | *09-24 09:49* | | | **hub 2.0.0 published** | **09-24 12:01** | ← | | #2539 | 09-24 16:57 | fail | | #2544 | 09-24 18:32 | fail | | #2216 | 09-25 11:58 | fail | #2500 merging afterwards is not a counterexample: GitHub does not re-run checks at merge time, so it merged on a result from ~20 hours earlier. That is also why this went unnoticed. ## The changes ### 1. Declare `httpx` in the `hf` extra `httpx` is not incidental to streaming — it is the only transport: - every fetch is HTTP: `POST /v1/completions` to the vLLM serve plus `GET /meta` and `/desc` against the connector's sidecar, all through `httpx.Client`; - there is no non-HTTP path — the base `StreamingDataset._fetch` is an abstract seam and `EagleVllmStreamingDataset._fetch` is its only implementation; - no other HTTP library appears in the module (`requests` / `urllib` / `aiohttp`: zero hits, and `requests` is not declared either); - even the retry predicate is built from it: `_TRANSIENT_FETCH_ERRORS = (httpx.HTTPError, OSError)`. It belongs in `hf` rather than in the core `dependencies`: the same module needs `transformers.trainer_pt_utils` at module scope, so one extra already gates the whole file, and a core install has no use for an HTTP client. The bound matches the 0.x API the code uses — `httpx` has no 1.0 release, and 2.x is a different distribution. This is the part that stops it recurring. `[hf]` currently gets `httpx` only because `datasets` happens to require it — the same accident with a different supplier, one release away from repeating. ### 2. Acquire `httpx` in the test through the existing skip guard The test file already intends to skip where the extra is absent — it has `pytest.importorskip("transformers")` and a comment explaining why, and `transformers` is absent in this job too. It broke only because `import httpx` sat **five lines above** that guard, where a missing module ends collection instead of skipping one file. ## Verification - With everything installed: **18 passed**, no behaviour change. - The import-order property is checked with an AST walk over the module's top-level statements: no `hf`-extra-only import precedes the first `importorskip` (which is now line 41). - A faithful local reproduction was attempted and abandoned honestly: hiding `httpx` locally also breaks `huggingface_hub` 1.28, which `modelopt.torch.opt.plugins.huggingface` imports, so the local failure is not the CI one. CI is the oracle for that half — this PR's own `partial-install (torch)` run is the check that matters. ## Scope Two files, five lines of declaration and four of test import order. Deliberately not folded into any feature PR: it blocks the whole repo, and burying a repo-wide fix inside unrelated work is how these stay invisible. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Optional Hugging Face installations now include `httpx`, supporting features that require HTTP communication without requiring it for all installations. * **Tests** * Hugging Face streaming dataset tests now skip when `httpx` is unavailable, allowing the remaining test suite to be collected and run without it. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
87f7d1432f |
fix(speculative): hold the DFlash draft's fp32 master weights in the optimizer (#2483)
### What does this PR do? Type of change: Bug fix **Follow-up to #2342**, which split this out on review (commit `c67784d9`), and a rethink of how the flag is implemented. `dflash_fp32_master_weights` exists because the DFlash draft is cast to the frozen bf16 target's dtype, so AdamW allocates its moments in bf16 — and bf16 is too coarse to hold them. At `beta2=0.999` a single step changes `v` by at most **0.100%**, while the smallest change bf16 can represent near `v` is **0.164% mean / 0.388% max** (measured): every decrease rounds away, `v` only grows, and the effective step size decays on its own from step 1. #2342 fixed that by **promoting the draft model to fp32**. Everything else followed from giving the model a dtype the rest of it does not have — a bf16 autocast at every entry point, two transformers loader hints so `from_pretrained(dtype="auto")` would not round the draft away, a post-condition check because those hints fail silently, and a doubled DDP gradient all-reduce. **This PR puts the fp32 in the optimizer instead**, where Megatron-LM, DeepSpeed and apex put it. `MasterWeightAdamW` holds an fp32 master copy of each non-fp32 parameter plus fp32 moments in `self.state[p]`, steps on the master, and copies back at the parameter's dtype. The model is never anything but the base dtype, so every one of those follow-on pieces is deleted, gradients stay bf16, and the exported drafter is unchanged. What the placement costs is that wiring the optimizer becomes the training loop's job: `EagleTrainerWithAccLog.create_optimizer` builds it, and `VerifyMasterWeightsCallback` raises at the end of step 1 if the moments are not fp32. **The default flips to `True`** — the flag now changes optimizer memory and optimizer arithmetic and nothing else. Flipping it on the model-promoted implementation turns **25 of 259** unit tests red; flipping it here is **259 passed**. Set it to `False` to reclaim the memory, about 12 bytes per draft parameter instead of 4. <details> <summary>Three drive-by fixes, independent of the above</summary> - `_place_draft` is folded back into `modify()` — it fused the draft's dtype, its device and an eager rotary buffer behind one meta guard. - The module docstring's claim that `DFlashModule` has an `_apply` meta-buffer fix is removed (`grep "def _apply"` matches nothing, and never did). - #2342's field description no longer lists `evaluation` as a broken path — `forward` short-circuits to the base model when `not self.training`, so the draft never runs there. </details> ### Usage No API change. `dflash_fp32_master_weights` now means the *optimizer* holds fp32 master weights rather than the draft model being fp32. ### Testing **1 · The refactor is arithmetically a no-op.** Both implementations run AdamW on an fp32 tensor, so given the same starting values and the same gradients the trajectories are identical — 1000 steps, `weight_decay=0.01`: ``` old fp32 parameter vs new fp32 master : bitwise equal = True (max |diff| 0.0e+00) exp_avg / exp_avg_sq : bitwise equal = True optimizer state dtypes : ['torch.float32'] model parameter dtype : torch.bfloat16 ``` Initial values have to be matched at bf16 first, or the bf16 arm's one-time rounding of the draw shows up as a 2e-4 "difference" that is not arithmetic. With that controlled, the two implementations differ only in their *inputs*: gradient precision (fp32 vs bf16 — torch 2.10 requires `grad.dtype == param.dtype`) and that one-time rounding. **2 · End to end on GPU: the effect survives the refactor.** Qwen3-1.7B base, real corpus, one GPU per arm, three arms — pure bf16 (flag off), the #2342 implementation, and this one — on two algorithms trained independently, sharing seed, data order and initialisation within an algorithm. <img width="2925" height="960" alt="image" src="https://github.com/user-attachments/assets/40d37059-ab8a-4924-b049-85f76b70b156" /> The two fp32 arms sit on top of each other for the whole run while bf16 stays above both, and the old-vs-new gap is 10–23× smaller than the fp32-vs-bf16 effect it has to be compared against. **Acceptance length says the same thing, and settles what the loss could not.** All six drafters at the end of those curves were exported and served under vLLM against the same base, and measured on MT-Bench (80 prompts, 8 categories, greedy, one request at a time, `num_speculative_tokens` = trained `block_size` − 1, every knob but the drafter held fixed): | | pure bf16 | fp32 in model (#2342) | fp32 in optimizer (this PR) | new − old | fp32 − bf16 | |---|---|---|---|---|---| | `dflash` | 1.3068 | 1.3708 | **1.3666** | −0.0042 `t=−0.90` | +0.0619 `t=+13.8` | | `lilicorr` | 1.2536 | 1.2814 | **1.2882** | +0.0068 `t=+1.42` | +0.0312 `t=+9.2` | Paired by prompt, n=80. On both algorithms the new-vs-old 95% CI straddles zero (`dflash` [−0.0134, +0.0051], `lilicorr` [−0.0028, +0.0164]) while fp32-vs-bf16 does not come close to it, and the sign of new-vs-old **flips between the two algorithms** — what a rounding difference looks like, not a bias. This is also the comparison the training loss could not give: all three arms are **exported and served in bf16**, so the old implementation's fp32 draft weights are rounded at export exactly as they would be for deployment, and the "its loss was computed on a more precise forward" caveat below does not apply. `lilicorr` needs [vllm-project/vllm#57934](https://github.com/vllm-project/vllm/pull/57934), applied as an overlay so that both algorithms are measured on one engine build. The right panel is the mechanism, and the one signal that depends on neither the seed nor the choice of loss statistic: Adam's updates to the draft's RMSNorm gains are smaller than the bf16 ULP at 1.0 (0.0078), so in the bf16 arm every one of them rounds away and the gains never move — not one of `dflash`'s 14 in 30000 steps, and two of `lilicorr`'s 20 by 3e-06. Both fp32 arms move all of them, by the same amount. Two results behind the figure rather than in it. **fp32-vs-bf16 grows with the horizon** while old-vs-new does not — on `dflash` −0.129 at 1500 steps → −0.262 at 15000 → −0.341 at 30000, and on `lilicorr` −0.191 → −0.220 → −0.285, against an old-vs-new difference that stays near 0.02 at every horizon and changes sign between them (−0.026 → +0.028 on `lilicorr`). That is what a compounding bias and a rounding difference respectively should look like, and it is the reason the longer runs were worth doing. And **across seeds**, the paired old-vs-new difference at 1500 steps is +0.0003 (n=6) on `dflash` and +0.0643 (n=10) on `lilicorr`, both with a 95% CI straddling zero. <details> <summary>Limits of the above, stated rather than smoothed over</summary> At 5 seeds the `lilicorr` paired difference read +0.1610 ± 0.0557 (t=+2.89, 4/5 seeds in the same direction) — nominally significant, suggesting the new implementation was genuinely worse there. Four further `lilicorr` seeds were run against that pre-declared question; two came back strongly negative and the estimate settled at +0.0643 (95% CI [−0.086, +0.214]). The earlier reading was small-sample noise. At 1500 steps on `lilicorr` that CI is *not* narrower than the fp32-vs-bf16 effect it is being compared against, so the 1500-step sweep alone cannot certify equivalence there — `lilicorr` is still at loss 9.3 and deep in its early transient, and it is the long runs that resolve it. On `dflash` the 1500-step CI (±0.031) is already 4× tighter than the effect (−0.129). One asymmetry the loss comparison cannot separate: the old implementation held the draft weights in fp32 *at forward time*, so its training loss was computed on a more precise forward, while both implementations export bf16. Any residual advantage it appears to have is therefore an upper bound. </details> **3 · Unit tests.** `tests/unit/torch/speculative/` — **259 passed** on **transformers 5.0.0** and **5.3.0**, both ends of the supported `>=5.0,<5.13` (CPU, torch 2.10). `TestDFlashFp32MasterWeights` is rewritten for the new mechanism; the two that would have caught the traps in this design are `test_resume_does_not_round_the_master_back_down` (`Optimizer.load_state_dict` casts float state to its parameter's dtype, so a naive subclass rounds the master and both moments to bf16 on *every* resume, silently, with the loss still falling) and `test_the_callback_refuses_a_loop_that_forgot_the_optimizer`. The rest cover the draft's dtype with the flag either way, that no forward path needs an autocast any more, that plain AdamW really does leave the moments in bf16, and that an fp32 model allocates no redundant master. A sharded FSDP2 `DTensor` keeps an fp32 master and fp32 moments through a step. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ for artifacts, with one intentional default change. The draft's stored dtype goes back to matching the base, as it was before #2342; existing checkpoints load unchanged and the exported drafter is unaffected. The flag now defaults to **`True`** — the measurements above are the reason, and the cost is fp32 master + fp32 moments for the draft only. A training loop that builds its own optimizer instead of using the shipped `create_optimizer` gets plain AdamW and none of this; `VerifyMasterWeightsCallback` makes that fail loudly at step 1 rather than skip the feature quietly. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new dependencies. - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ — draft; will run `/claude review` before marking ready. ### Additional Information **On the 7–14% acceptance-length gain quoted in #2342:** that was measured on the fp32-model arithmetic and is not re-derived here. What is measured above is the like-for-like comparison this PR has to answer — same corpus, same horizon, same serving path, one implementation swapped. **History:** commits 1–3 restore the autocast design as it was split out; commits 4–6 replace it. Happy to squash before review. <details> <summary>Alternatives measured and rejected, so they do not get re-proposed</summary> - **Swapping `p.data` to the master and calling `super().step()`** (reuses all of AdamW, ~20 lines instead of ~50): bit-identical on ordinary parameters over 25 steps, but silently wrong under FSDP2 — assigning `.data` on a `DTensor` parameter updates the wrapper's reported dtype while the local shard keeps the model's, so `p.dtype` reads fp32, `p.data.dtype` reads bf16, and `zeros_like(p)` allocates the moments in bf16 anyway. CPU tests pass either way. - **Narrowing the autocast from `__call__` to `forward`** (while it still existed): turns 10 Domino/DSpark tests red — the variants apply their heads in their own `forward` overrides, outside `DFlashModule.forward`. - **Building the rotary buffer on meta and letting the loader materialise it**: makes RoPE correctness depend on transformers selecting a branch by class-name substring (`"RotaryEmbedding" in module.__class__.__name__`), and the `if not hasattr` guard is then permanently satisfied, so a later `to_empty()` leaves garbage forever — measured `4.56e-41`, i.e. cos=1 / sin=0, no positional encoding at all. - **Building it eagerly in `DFlashModule.__init__`**: lands before the dtype cast, so `Module.to` rounds the RoPE frequencies to bf16 on the default path — measured `0.8659643530845642` → `0.8671875`, loss `3.47230935097` → `3.47114777565`. </details> 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - DFlash now uses FP32 optimizer master weights and Adam moments by default while keeping draft parameters in the base model’s dtype. - Master-weight training preserves optimizer precision when restoring checkpoints. - The feature can be disabled to reduce optimizer memory usage. - Draft models consistently follow the base model’s dtype and device. - **Bug Fixes** - DFlash workflows now support operation without autocast. - Added validation for compatible AdamW-family optimizers and master-weight precision, including resumed training runs. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c5d1065331 |
Modeling Lib: per-model architecture kick off with spec registration (#1828)
### What does this PR do? **Type of change:** New feature — per-model infrastructure (with behavior changes, see below) Starts `modelopt/torch/models/`: somewhere for ModelOpt to keep what it knows about a model, keyed by HF model type (`config.model_type`), one package per type (`<model_type>/specs.py`). Importing the package registers every spec; consumers call `get_spec(model_type)` / `match_moe_block(module, model_type)` and read the fields. The point is the infrastructure, not the migration. Today ModelOpt's per-model knowledge is scattered across if/elif chains in whichever subsystem happened to need it, so the same fact gets restated per subsystem and drifts. A model's MoE block classes, its expert projection naming, which of its norms store `w - 1` — these are facts about the *model*, and several subsystems want them. This PR gives them one home and one lookup. It is a kick-off, so it is deliberately narrow: specs plus the first consumer. **Export is that first consumer**, which is why most of the changed lines are on the export side — not because this is an export refactor. Quantization, speculative decoding and the rest keep their own tables for now; the per-model directory is where their sections land later, and `specs.py` is just the first thing in it. | Data the first consumer moved in | Spec field | Consumers | |---|---|---| | MoE block classes and expert linear naming | `MoESpec` | `get_expert_linear_names`, `get_experts_list`, `is_moe`, `sync_moe_gate_up_amax` | | grouped expert-export support set | `ExportSpec.grouped_expert_export` | `get_experts_list` | | AWQ `pre_quant_scale` fusion rules | `ExportSpec.pqs_fuse_rules` | `fuse_prequant_to_linear` | | weight-plus-one norm class names | `ExportSpec.weight_plus_one_norm_names` | `_layernorm_uses_weight_plus_one` | `ModelSpec` holds each concern as a separate, optional section, and the split is what makes it extensible: **topic sections** hold architecture facts any subsystem can read (`MoESpec` — block classes, expert projection naming, how the experts are stored), **subsystem sections** hold one subsystem's own data and policy (`ExportSpec` — AWQ fusion rules, weight-plus-one norms, whether grouped expert export is validated for this model). A new subsystem adds a section rather than a table. A section is `None` when the model has nothing to say about it, so a dense model carries `moe_spec=None` rather than an empty one. Two further fields record where a model's classes come from at all: `modeling_source` (`transformers` or `remote_code`) and `min_transformers_version`. ### Behavior changes **1. Lookup key: `type(root_model).__name__.lower()` → `config.model_type`.** The old key changes after `quantize` wraps the model, forcing substring matching; `model_type` is stable, so `ExportContext` carries it and lookups are exact. Matching is also tightened to case-insensitive **exact** names against the module's **MRO**, so quantized subclasses still match via their base class without substring false positives. **2. ⚠️ `get_expert_linear_names` raises instead of guessing.** Unmatched MoE blocks previously fell through to `["w1", "w2", "w3"]` (Mixtral naming); they now raise `NotImplementedError` telling you to register a `ModelSpec`. The blast radius depends on the transformers release, because only *iterable* expert layouts consult the spec — a fused expert container is resolved by the structural first-projection check and never asks for naming. Measured against both ends of the support matrix: | transformers | Unregistered MoE families `is_moe` admits | Reach the spec lookup | Actually affected | |---|---|---|---| | 5.14.1 (`tf_latest`) | 19 | **0** — all fused | none | | 4.57.6 (`tf_min`) | 9 | 7 iterable | **2** | On 4.57, five of the seven (`ernie4_5_moe`, `flex_olmo`, `jamba`, `olmoe`, `qwen3_omni_moe`) name their experts `gate_proj`/`up_proj`/`down_proj`, so the `w1` fallback already failed with `AttributeError` on `main` — for them this trades an unhelpful error for one that names the fix. The other two, **`minimax` and `phimoe`**, are Mixtral-derived with `w1`/`w2`/`w3` experts and did export on `main`, so they are now **registered** rather than left to regress. Both only take the per-expert path on transformers 4; on 5 their experts are fused (`MiniMaxExperts`, `PhimoeExperts`) and the structural check resolves them, which their specs decline to answer for. **So no model that exported correctly before this change regresses.** This is backward-incompatible and has a `CHANGELOG` entry. **3. DeepSeek-V3 and V4 are now registered, and newly visible to the MoE path.** Neither was reachable before: `DeepseekV3MoE` (like `DeepseekV2Moe` and `DeepseekV32MoE`) is invisible to `is_moe` — its class name does not end in `SparseMoeBlock` and it calls its router `gate`, so neither the name test nor the structural `router`+`experts` test matches. `DeepseekV4SparseMoeBlock` was detected by name but had no spec to resolve expert naming from. Registering `block_names` is what puts V3 on the MoE path at all, so this is added coverage rather than a fix to existing behavior. `deepseek` stays registered alongside them: it describes the remote-code `DeepseekMoE` block of DeepSeek-MoE/V1, and its spec is the only thing making that block detectable. **4. Fused expert containers are skipped in `get_experts_list` instead of crashing.** transformers 5 replaced several iterable expert `ModuleList`s with a single module holding 3-D parameters, while the specs still describe the transformers 4 iterable layout — a spec cannot tell the two apart, only the module can. The AWQ/SVDQuant resmooth pass in `requantize_resmooth_fused_llm_layers` reached `len(module.experts)` on a module with no `__len__` and died with `TypeError` mid-export. **Mixtral already hits this on `main`**; registering `deepseek_v3` only made an existing bug easier to reach. The skip is scoped to specs that claim iterable experts, so a layout the spec calls unsupported (DBRX, whose per-expert linears live under `experts.mlp`) still fails loudly rather than silently dropping resmoothing. Everything else is behavior-preserving, pinned by tests: - **`sync_moe_gate_up_amax` keeps generic coverage.** Quantization admits MoE blocks *structurally*, so unregistered families (Olmoe, Jamba, MiniMax…) reach it. With no spec it falls back to every declared gate/up naming — what it did pre-registry. Skipping them would leave the fused `gate_up_proj` halves on inconsistent `weight_scale_2`. - **`ExportSpec.grouped_expert_export` matches the legacy `get_experts_list` support set**, except for `deepseek_v3`/`deepseek_v4`, which are new and had no legacy behavior to match. Note `qwen3_5_moe`: legacy keyed off the root class name and `"qwen3_5moeforcausallm"` matched none of its qwen substrings (the `_5` breaks `qwen3moeforcausallm`), so it raised — the spec keeps `False` to match. Enabling it belongs in its own PR. - **`is_moe` consults the model's own spec first**, before the generic name and structural fallbacks, so per-model data always wins. The three checks are or-ed, so this is a no-op; the same reordering is deliberately *not* applied to `get_expert_linear_names`, where the structural fused-experts check must stay first. - **Two values are corrected, not moved.** `gpt_oss` declared `block_names=("GptOssMoE",)`, which matches no real module (transformers names it `GptOssMLP`); naming still resolved via the single-naming shortcut, so this was invisible. Now `("GptOssMLP", "GptOssMoE")`. **5. ⚠️ DBRX expert input amax is now populated, changing its exported scales.** The second corrected value, called out separately because it has a numerical consequence. The legacy branch keyed on `DBRXMoeSparseMoeBlock`, a class transformers does not define — it names the block `DbrxFFN`. So DBRX fell through to the `w1`/`w2`/`w3` default, every `hasattr(experts_mlp, linear_name)` in `_prepare_dbrx_experts` evaluated `False`, and the handler wrote no expert input amax at all. `dbrx/specs.py` declares the real names (`w1_linear`, `w2_linear`, `v1_linear`), so the handler now does what it was written to do. Anyone diffing a re-exported DBRX checkpoint against an older one will see different expert activation scales. That is the fix working, not a regression from the refactor. No `CHANGELOG` entry: DBRX is not listed as a supported export target in `docs/source/deployment/` or `examples/hf_ptq/README.md`. ### Scope One consumer wired up: the unified HF export path. The TRT-LLM builders now live in `modelopt/torch/export/trtllm/` after #2365, and are deprecated. This PR changes one import line there — `is_moe` moved to `modelopt/torch/models/moe.py`, so importing it from `..layer_utils` no longer resolves — and nothing else. Their `is_moe()` calls still pass no `model_type` and fall back to the all-specs search, reproducing today's behavior; their hardcoded tables, including a second copy of the weight-plus-one norm names, stay as they are and migrate whenever that package does. Worth noting that #2365 landed the same boundary from the other direction: the slimmed `export/layer_utils.py` is now documented as "module-shape predicates and MoE quantizer helpers shared by every export backend", which is exactly the set of functions this PR rewires. The Megatron path has its own per-family registry and is untouched; folding it in is a later question, not this PR's. `is_moe` itself moved out of `modelopt/torch/export/layer_utils.py` into `modelopt/torch/models/moe.py` — whether a module is an MoE block is a modeling question, not an export one, and it is the first piece of shared modeling logic to follow the specs into the new package. ### Relationship to #1939 (why a second registry?) Orthogonal layers: #1939's `ExportModuleRegistry` dispatches on **module structure** (*which handler runs?*), this registry resolves **family data** (*what are its values?*). Exporting a `QuantMixtralSparseMoeBlock`, #1939 picks the shared iterable-experts handler (also serving Qwen/DeepSeek/Gemma4); inside, `get_expert_linear_names` resolves the `mixtral` spec to `("w1", "w2", "w3")`. One handler serves many families, one spec serves many handlers. Sharing the matcher machinery is a planned follow-up. ### Usage ```python # modelopt/torch/models/qwen3_moe/specs.py — adding a model needs no engine edits register( ModelSpec( model_type="qwen3_moe", min_transformers_version="4.57", moe_spec=MoESpec( block_names=("Qwen3MoeSparseMoeBlock",), expert_linear_names=("gate_proj", "down_proj", "up_proj"), gate_up_pair=("gate_proj", "up_proj"), ), export_spec=ExportSpec( grouped_expert_export=True, pqs_fuse_rules=( (("Qwen3MoeAttention",), "v_proj", "o_proj"), (("Qwen3MoeMLP",), "up_proj", "down_proj"), ), ), ) ) ``` One layout per model. `block_names` is a tuple, so a model whose MoE appears under several class names is covered as long as they share a layout (`gpt_oss`'s `GptOssMLP`/`GptOssMoE`). `gemma4_text` imports its section from `gemma4` rather than restating it. ### Testing `tests/unit/torch/export/` — **202 passed**; `tests/unit/torch/quantization/` — **903 passed**. The `test_export_diffusers.py` collection error and 6 `test_quant_aware_conversion.py` failures are pre-existing and reproduce on untouched `main`, as does the one failing pre-commit hook (`generate-arguments-md`, no torch in its venv). `tests/unit/torch/models/test_model_specs.py`: registry matching (MRO / quantized classes / model-type scoping); every registered MoE section's `(block_names, expert_linear_names, fused_expert_names, gate_up_pair)` as an exhaustive table that fails if a spec is added without a row; the exhaustive `grouped_expert_export` support set; the structural fused-expert shortcut and the spec's precedence over it; the fused-container skip in `get_experts_list`; the `sync_moe_gate_up_amax` fallback; and legacy-equivalence of the aggregated `pqs_fuse_rules` / `gate_up_pairs` / `weight_plus_one_norm_names`. Mutation-checked — flipping a flag, typo-ing a block name, or swapping the precedence each fail. `tests/unit/torch/models/test_specs_vs_transformers.py` validates the specs against the *installed* transformers rather than a mirrored table: every registered block class must name a real class from its `min_transformers_version` on, and every `remote_code` model must still be absent. It names no model — which specs to check, and which cannot be, both come from the registry. > The run counts above were measured before the follow-up commits in this branch (spec > precedence, the policy split, the `MoESpec`/`MoELayout` merge) and have not been > re-measured; CI on the current head is the authority. **Local per-model-type E2E sweep** (transformers 5.3.0, tiny models from config, run against this branch and base `main`; not checked in): | model_type | real block class | this branch | base `main` | |---|---|---|---| | `qwen3_moe` / `qwen2_moe` / `qwen3_next` | `Qwen*SparseMoeBlock` | `gate/down/up` | same | | `mixtral` | `MixtralSparseMoeBlock` | `w1/w2/w3` | same | | `nemotron_h` | `NemotronHMoE` | `up/down` | same | | `gpt_oss` | `GptOssMLP` | `gate_up/down` | ❌ `w1/w2/w3` | | `deepseek_v3` | `DeepseekV3MoE` | raises | ❌ `w1/w2/w3` | Five types identical to base; both differences are base bugs the registry surfaces. `dbrx` could not be built on transformers 5.3.0 (`DbrxAttentionConfig` missing `rope_theta`). `sync_moe_gate_up_amax` was separately checked against base across registered, unregistered, and no-config models — identical in all cases. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ❌ — `get_expert_linear_names` no longer falls back to `w1/w2/w3`; see "Behavior changes" #2. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no copied code, no new dependency (stdlib only). - Did you write any new necessary tests?: ✅ — `tests/unit/torch/export/test_model_specs.py`. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ — happy to add an entry given the backward-incompatible item. - Did you get Claude approval on this PR?: ✅ — run; its CRITICAL finding on `sync_moe_gate_up_amax` is fixed above. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added model-aware export support across a broad range of MoE and dense architectures, including newer DeepSeek, Gemma, Qwen, Nemotron, Arctic, DBRX, GPT-OSS, Llama, and Mixtral variants. - Improved quantization and export handling for model-specific expert layouts, projection fusion, normalization, and gate/up synchronization. - Added automatic architecture recognition using Hugging Face model metadata. - **Bug Fixes** - Unsupported or ambiguous model layouts now produce explicit errors instead of applying potentially incorrect defaults. - **Documentation** - Added documentation describing model specifications and supported architecture metadata. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> Co-authored-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
5db2682519 |
[Example]: Calibration-free FP8/NVFP4 PTQ for speculative-decoding drafters (#2027)
### What does this PR do? Type of change: new example Adds `examples/speculative_decoding/scripts/quantize_drafter.py`, a CLI that quantizes an exported speculative-decoding drafter to FP8 or NVFP4 — weight-only or weight+activation — with no calibration data. It needs no modeling code either. Exported drafters such as [`nvidia/MiniMax-M3-DSpark`](https://huggingface.co/nvidia/MiniMax-M3-DSpark) have no importable model class, so each 2-D weight is wrapped in a throwaway `nn.Linear` under its checkpoint key and ModelOpt's usual `quantizer_name` patterns select over those names. Works for any drafter layout (DSpark / DFlash / EAGLE3 / Medusa). **Formats:** `w4a16_nvfp4`, `nvfp4`, `fp8`, `fp8_pc_pt` — the ModelOpt formats vLLM's backend can actually serve. AWQ is deliberately not offered, since `awq_lite` silently degrades to plain RTN without a `forward_loop`. **Static activation scales without calibration.** `fp8` and `nvfp4` normally need an activation amax *measured* on calibration data; a fixed `input_scale` of 1.0 is applied instead. That works because acceptance length is governed almost entirely by **clipping**, not resolution: Sweeping the fixed scale over three decades (same setup as the Testing section below; bf16 baseline 3.1423): | `input_scale` | amax | FP8 AL | vs bf16 | NVFP4 AL | vs bf16 | |---|---|---|---|---|---| | 0.003 | 1.3 | 2.2204 | -29.34% | 2.2076 | -29.75% | | 0.01 | 4.5 | 2.6719 | -14.97% | 2.6641 | -15.22% | | 0.03 | 13.4 | 2.9751 | -5.32% | 2.9259 | -6.89% | | 0.1 | 44.8 | 3.1013 | -1.31% | 3.0206 | -3.88% | | 0.2 | 89.6 | 3.1178 | -0.78% | 3.0015 | -4.48% | | 0.3 | 134.4 | 3.1370 | -0.17% | 3.0222 | -3.82% | | 0.5 | 224.0 | 3.1268 | -0.50% | 3.0360 | -3.38% | | **1.0 (default)** | **448.0** | **3.1457** | **+0.11%** | **3.0193** | **-3.91%** | | 2.0 | 896.0 | 3.1354 | -0.22% | 3.0172 | -3.98% | | 4.0 | 1792.0 | 3.1245 | -0.57% | 3.0034 | -4.42% | Both formats fall off a cliff below ~0.03, where the declared range sits far under the activations' true magnitude and most of the tensor is clipped. Both then sit on a flat plateau from ~0.3 to 4.0 **with no drop-off at the top**, so the scale only has to be big enough. 1.0 is the middle of that plateau, which is why it is hardcoded rather than exposed. NVFP4 trails FP8 by a roughly constant 3.5% across the plateau — that gap is the 4-bit resolution cost, and no choice of scale recovers it. Deriving the amax from the weights instead was tried and does not work: `max|W|` averages 0.79 while a RMSNorm'd activation is O(1) with outlier channels in the tens, so the range lands 1–2 orders of magnitude low and clips, measuring -31% to -46% AL. **Where calibration would go.** All of this sits behind `resolve_activation_scales()`, the single place deciding where a static amax comes from. Real calibration slots in ahead of the fixed fallback with no change to the CLI or the call site, and composes because `set_static_activation_amax()` skips quantizers that already have an amax: ```python if calib_forward_loop is not None: mtq.calibrate(root, quant_cfg["algorithm"], forward_loop=calib_forward_loop) set_static_activation_amax(root) # fills in what calibration did not reach ``` **Serving a quantized drafter.** Four things had to be written into the exported checkpoint before vLLM would load one: - emit `quant_method` (`modelopt_fp4` / `modelopt`) — vLLM reads that key, ModelOpt writes only `quant_algo` - emit the exclusion list under `ignore` too — that is the key read from the flat `quantization_config`; `exclude_modules` alone yields an empty exclusion set - add `*<name>` wildcards so exclusions match a runtime's nested module prefix (`model.fc`) rather than the checkpoint key (`fc`) - add `*qkv_proj` / `*gate_up_proj` aliases for layers a runtime fuses, whose names appear in no checkpoint key Nothing is then needed on the caller side. **This closes the open question left in the previous revision of this PR: vLLM does read `quantization_config` off the draft checkpoint.** `ModelConfig._verify_quantization` fills `quantization` in from `quant_method` when it is unset, so once the export declares that key — the first fix above — detection works on its own. Verified on Nemotron-3.5-Lightning passing nothing: `Detected ModelOpt NVFP4 checkpoint (quant_algo=NVFP4)` → `FlashInferCuteDslNvFp4LinearKernel`, AL 4.278 against 4.203 measured earlier. `specdec_bench` also gains a `DSPARK` algorithm, which it did not have: an exported `Qwen3DSparkModel` would otherwise have to go through `DFLASH` and be built with vLLM `method="dflash"`. The branch sets `method="dspark"` and leaves `draft_sample_method` on vLLM's own default of `greedy`. A target whose fused-collective workspace (sized at CUDA-graph capture) overflows at large speculative batches can disable graphs with `--runtime_params '{"engine_args": {"enforce_eager": true}}'`. For DFlash-family drafters, `qwen3_dflash.py` builds its fused context-KV projection by reading `qkv_proj.weight` raw and calling `F.linear`, which cannot consume a packed weight. Keep those layers in bf16 with `--exclude '*q_proj*' '*k_proj*' '*v_proj*' '*qkv_proj*'`; `o_proj` and the MLP — the bulk of the drafter — still quantize. That exclusion is mandatory, not a tuning choice. `fc` (the projection from the target's captured layers into the draft) is the one real knob, and it is a genuine trade rather than a free win — see the Testing section for both models' numbers. The examples quantize it; add `'*fc*'` to the exclude list to keep it in bf16. `embed_tokens`, `markov_head` and `confidence_head` are excluded by default: they are 2-D so the flat view treats them as GEMMs, but they are embeddings or a single-output projection. `lm_head` is excluded by the preset itself — unlike on a base model it is 37% of this drafter's parameters, so `--quantize_lm_head` is a real lever (~1.9 GiB), but measure AL first. The flag re-enables both of `lm_head`'s quantizers; re-enabling only the weight one would ship a W+A checkpoint whose `lm_head` has no `input_scale` while the config still advertises it as quantized. ### Usage ```bash # weight+activation FP8, calibration-free, lossless on both models measured below python scripts/quantize_drafter.py \ --drafter_path deepseek-ai/dspark_qwen3_8b_block7 \ --qformat fp8 \ --export_path ./dspark-qwen3-8b-fp8 \ --exclude '*q_proj*' '*k_proj*' '*v_proj*' '*qkv_proj*' # smallest: weight-only NVFP4 python scripts/quantize_drafter.py \ --drafter_path nvidia/MiniMax-M3-DSpark \ --qformat w4a16_nvfp4 \ --export_path ./MiniMax-M3-DSpark-W4A16 ``` Or end to end on Slurm — quantize, then measure AL — via the launcher examples added here, one per target: ```bash uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_dspark_ptq_nvfp4.yaml --yes uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/hf_dspark_ptq_nvfp4.yaml --yes ``` Serving one, if you are not going through `specdec_bench`: ```python speculative_config = { "method": "dspark", "model": "./dspark-qwen3-8b-fp8", # quantization is read from its config.json "num_speculative_tokens": 7, } ``` ### Testing Two targets with different architectures, so the conclusions are not one model's quirk: * **Qwen3-8B** (dense transformer) + [`deepseek-ai/dspark_qwen3_8b_block7`](https://huggingface.co/deepseek-ai/dspark_qwen3_8b_block7), `block_size` 7, TP1. * **Nemotron-3.5-Lightning-30B-A3B** (hybrid Mamba-MoE) + [`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark`](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark), `block_size` 8, TP8, with the mamba engine settings the model card pins (`mamba_backend=flashinfer`, `mamba_ssm_cache_dtype=float16`, stochastic SSM-cache rounding). Both: MT-Bench 80 questions, greedy, one vLLM instance per point. | recipe | activations | Qwen3-8B AL | vs bf16 | Nemotron-3.5 AL | vs bf16 | |---|---|---|---|---|---| | bf16 baseline | — | 3.1423 | — | 4.3296 | — | | **`fp8`** | static, `input_scale` 1.0 | **3.1457** | **+0.11%** | **4.3289** | **-0.02%** | | `fp8_pc_pt` | dynamic per-token | 3.1228 | -0.62% | 4.3411 | +0.26% | | `w4a16_nvfp4`, `fc` in bf16 | bf16 (weight-only) | 3.0392 | -3.28% | 4.2899 | -0.92% | | `w4a16_nvfp4`, `fc` quantized | bf16 (weight-only) | 3.0186 | -3.94% | 4.2334 | -2.22% | | **`nvfp4`** | static, `input_scale` 1.0 | **3.0193** | **-3.91%** | **4.2030** | **-2.92%** | **FP8 weight+activation at the fixed `input_scale` of 1.0 is lossless on both.** +0.11% and -0.02% are both inside run-to-run noise — the Nemotron baseline was measured twice under identical settings and the two runs differ by 0.94% (4.3093 / 4.3499), which sets the resolution of that column. On the same reading, `fp8` and `fp8_pc_pt` are indistinguishable on Nemotron; the dynamic variant only pulls ahead on Qwen3. NVFP4 costs 3-4% on Qwen3 and 2-3% on Nemotron, i.e. the 4-bit weight resolution is the real price and it is model-dependent but bounded. Whether to quantize `fc` is a per-model call rather than a general recommendation — it buys a few percent of size for an AL cost that differs by ~2x between these two drafters: | `fc` bf16 → quantized | Qwen3-8B | Nemotron-3.5 | |---|---|---| | checkpoint size | 3.293 → 3.181 GiB (-3.4%) | 1.316 → 1.258 GiB (-4.4%) | | AL | 3.0392 → 3.0186 (-0.68%) | 4.2899 → 4.2334 (-1.32%) | `fc` itself is only 3.5% (Qwen3) / 4.5% (Nemotron) of drafter parameters; `embed_tokens` is the bulk (26% / 36%) and is excluded by default. The Qwen3 `w4a16_nvfp4` rows were measured in a later session than the rest of that column; the `fc`-in-bf16 run reproduced the original number to four decimals (3.0392), so the column is internally comparable. Also validated on `nvidia/MiniMax-M3-DSpark`: `w4a16_nvfp4` runs in 67 s on CPU, 9.98 GiB (fp32) -> 3.51 GiB; all 43 quantized tensors round-trip within 0.0952 relative error; the 29 untouched tensors are bit-identical to `bf16(source)`. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (example-only) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ❌ — validated manually as above. Can add a `tests/examples/speculative_decoding/` test over a small synthetic drafter if wanted before merge. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (example-only) - Did you get Claude approval on this PR?: ❌ (not yet run) ### Additional Information The measurements above are one drafter on one target with one benchmark; the plateau's location and the ~3.5% NVFP4 gap should be re-measured before assuming they carry to a different drafter. Note when reading an exported checkpoint: `input_scale` is `amax/448` for FP8 but `amax/(6*448)` for NVFP4, so the one fixed amax records as 1.0 in an FP8 checkpoint and 0.1667 in an NVFP4 one. Both mean the same activation range. Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
2b296b2f62 |
Support fine-tuning released DFlash/DSpark drafters (causal SWA, attention sink, warm start) (#2149)
# Support fine-tuning released DFlash/DSpark drafters (causal SWA, attention sink, warm start) ### What does this PR do? Type of change: New feature + bug fix Adds what ModelOpt was missing to fine-tune an already-published DFlash/DSpark draft model. The concrete target is [`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark`](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark) on its hybrid Mamba/attention/MoE base, but every change is generic. Before this PR that checkpoint could not be trained faithfully — or even loaded: its attention-sink tensors were dropped as unexpected keys, its block-causal attention had no implementation, and its capture layers were silently overwritten with ModelOpt's defaults. **New user-facing options** (all default to today's behavior, so existing runs are unchanged): | Option | Values | Purpose | | --- | --- | --- | | `dflash_draft_attention` | `bidirectional` (default) / `causal` | Block-internal attention pattern. `causal` restricts a query at block position `i` to draft positions `<= i`. | | `dflash_attention_sink` | `false` (default) / `true` | Learnable per-head `attention_sink_bias [num_heads]` on every draft layer — one extra logit appended before the softmax and dropped after, so a head can put probability mass nowhere instead of being forced to attend inside its window (the GPT-OSS formulation). | | `dflash_init_checkpoint` | path | Warm-start the draft from an exported checkpoint instead of a random init. Any missing/unexpected/wrong-shaped tensor raises rather than warns. | | `dflash_architecture_config.target_layer_ids` | list | Which base layers feed the draft's `fc`. Previously recomputed unconditionally with no override. | **Bugs fixed along the way** (each one silently corrupts training rather than failing): - The exporter hard-coded `dflash_config.causal: False` and only wrote it under SWA, so even a correctly-trained causal draft would be served non-causally. It now reflects the trained setting, and emits `attention_sink_bias` when enabled. - `_build_generate_swa_mask` returned `None` whenever `swa_window_size` was unset, which would have dropped the causal structure at generation time while training used it. - `target_layer_ids` was recomputed from the uniform default on every convert. The released drafter uses `[1,5,19,29,41,51]`; the default for a 52-layer base is `[1,11,20,30,39,49]` — *different layers*. Here it surfaced as a matmul shape error only because the plane counts disagreed; with a matching count it would have trained on the wrong features silently. - The streaming dataset assumed the draft's aux layers all sit below the base's final layer (`aux = planes[:-1]`, `target = planes[-1]`). A draft whose top aux id *is* the final layer cannot get an extra plane — vLLM captures each layer once — so `final_aux_is_base_hidden` now lets the last plane serve both roles. It is derived from the model, not configured by hand. - DSpark head weights load from either the flat layout ModelOpt exports (upstream DeepSpec convention) or the nested `markov_head.` layout the NVIDIA release uses. Without the remap the two `[131072, 512]` Markov tables — ~14% of the draft's parameters — stay randomly initialized while everything else warm-starts, with no error. - `nemotron_h` is enabled in `_FINAL_NORM_TYPE_BY_MODEL_TYPE`: despite the hybrid stack, `NemotronHModel.norm_f` is a plain RMSNorm, and without the entry the offline/streaming fake base raises instead of reconstructing the distillation target. ### Usage ```yaml dflash: dflash_init_checkpoint: /path/to/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark dflash_draft_attention: causal dflash_attention_sink: true dflash_swa_window_size: 1024 dflash_block_size: 8 dflash_mask_token_id: 990 dflash_architecture_config: target_layer_ids: [1, 5, 19, 29, 41, 51] ``` A full worked example is at `modelopt_recipes/general/speculative_decoding/dspark_nemotron35_warmstart.yaml`. ### Testing **Unit tests** — 124 pass (`test_hf_dflash.py`, `test_hf_dspark.py`, `test_hf_domino.py`, `test_hf_dflash_offline.py`, `test_modeling_final_norm.py`), 32 of them new: causal mask structure (lower-triangular per block, no cross-block leakage, context visibility unchanged), the sink math (degenerates to plain attention at `-inf`, absorbs mass monotonically, receives gradient), warm-start load/reject paths, Markov key remapping, and explicit `target_layer_ids`. **Checkpoint compatibility** — the released drafter loads with zero missing/unexpected keys and zero shape mismatches; all 77 tensors (6 attention sinks and both Markov tables included) match bit-exactly, and a training step runs with gradients reaching the sink and Markov parameters. **End-to-end streaming training** — Nemotron-3.5 base served by vLLM (1 node, TP8) feeding 8 trainer GPUs over NIXL; the draft warm-starts from the released checkpoint and trains with `causal` + sink + SWA 1024. 128 Daring-Anteater conversations, 20 epochs (the plot shows the first 5, where the trend is clearest — the curves flatten after that):  Over the first 5 epochs loss falls **1.85 → 1.36** and train accuracy rises **0.25 → 0.49**; across the full 20 epochs they reach **1.21** and **0.48** (peak 0.54) before flattening. This validates the pipeline end-to-end — capture layers, plane split, mask direction, sink loading and warm-start weights all have to be right for this curve to appear. It is *not* a model-quality result: 128 samples over 20 epochs overfits by construction, and the corpus is not generated by the base model, so the absolute numbers are not meaningful. ### TODO (follow-up) **A complete, robust checkpoint/config converter.** Both conversions are handled ad hoc here: - *Draft config → training config.* The recipe transcribes ~15 fields by hand from the drafter's `config.json`. Only the shape-bearing ones (`num_hidden_layers`, `num_attention_heads`, `intermediate_size`, `markov_rank`) fail loudly when mistyped; the rest — `mask_token_id`, `causal`, `swa_window_size`, `block_size` — train "successfully" on a wrong value and only surface later as a mysteriously low acceptance length. A converter should derive the whole block from the checkpoint, including its aliases (`pard_token`, `dspark_markov_rank`, `dflash_query_causal`, top-level `sliding_window` / `attention_sink_bias`) and duplicated fields. - *Weight layout.* The `markov_head.` remap is a load-time hook. A converter should normalize layouts explicitly, and decide whether export should also emit the release's aliases so a round-trip reproduces the original format (today it renames `architectures` to `DFlashDraftModel`). - *Base config.* Serving this base on vLLM needs its `config.json` layer-type vocabulary updated for the transformers-5 path (`mamba` → `linear_attention`, `attention` → `full_attention`, plus a matching `hybrid_override_pattern`). That is done by hand today and is not covered by this PR. ### Before your PR is "*Ready for review*" - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added configurable causal or bidirectional attention for DFlash models. * Added optional attention sinks, checkpoint warm starts, and explicit target-layer selection. * Improved streaming data handling for shared auxiliary and base hidden states. * Added Nemotron-3.5 Lightning DSpark warm-start training and serving recipes. * **Bug Fixes** * Preserved configured attention behavior during model export. * Prevented warm-start checkpoints from being reapplied during restoration. * Improved checkpoint compatibility, validation, and attention-mask handling. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
c2070cfd7a |
ci: skip docs preview deploy for fork PRs (#2029)
### What does this PR do? Type of change: Bug fix (CI) The `deploy-preview` job in the `Docs` workflow fails on **every pull request opened from a fork**, which blocks merging for all external contributors. **Root cause.** `deploy-preview` runs `rossjrw/pr-preview-action@v1`, which pushes the built HTML to the `gh-pages` branch. The workflow declares `permissions: contents: write`, but for a `pull_request` event originating from a forked repository GitHub caps the `GITHUB_TOKEN` at **read-only** — the `permissions:` block cannot elevate above that cap. The push is therefore rejected: ``` remote: Permission to NVIDIA/Model-Optimizer.git denied to github-actions[bot]. fatal: unable to access 'https://github.com/NVIDIA/Model-Optimizer.git/': The requested URL returned error: 403 ``` The job's `if:` condition gated on `github.event_name`, `github.event.action` and the `changes` path filter, but never on whether the PR came from a fork — so it always ran and always failed. Because the `changes` filter matches `docs/**`, `modelopt/**` and `.github/workflows/pages.yml`, essentially any substantive fork PR trips this. **Fix.** Restrict `deploy-preview` to PRs whose head branch lives in this repository: ```yaml github.event.pull_request.head.repo.full_name == github.repository ``` A skipped job is not a failed job, so fork PRs are no longer blocked by it. ### Usage N/A — CI-only change. ### Testing Behaviour by scenario: | Scenario | Before | After | | --- | --- | --- | | PR from a branch in this repo | preview deployed | preview deployed (**unchanged**) | | PR from a fork | ❌ fails with 403 | ⏭️ skipped | | Fork deleted (`head.repo` is `null`) | ❌ fails | ⏭️ skipped | - Confirmed against workflow history: recent `Docs` runs on in-repo branches (`main`, `chenjiel/nvfp4-act-headroom`, `mxin/qad-skill`, `haoguo/dspark-ptq-script`) all succeed, while fork-branch runs fail with the 403 above. - `build-docs` was already passing on the affected PRs — only the deploy step failed, so documentation builds are unaffected either way. - YAML parses; `pre-commit run --files .github/workflows/pages.yml` passes. (`yamlfmt` excludes `^.github/workflows/`, so this file is not auto-formatted.) - This PR edits `.github/workflows/pages.yml`, which is itself in the `changes` filter, so it exercises `deploy-preview` on the in-repo path — the preview deploy on this PR passing is a self-check that the unchanged path still works. The `closed` cleanup path is gated by the same condition. That is intentional: a fork PR never deployed a preview directory, so there is nothing to remove. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — workflow-condition change; not unit-testable - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — CI infrastructure, not user-facing - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information Currently blocking #1975 (`add Qwen3-VL support for DFlash training`), which is approved with every other check green and sits at `mergeStateStatus: BLOCKED` solely because of this job. Note that a PR only picks up this fix once its branch contains it, since workflows run from the PR branch's own definitions. A follow-up option, if doc previews for external contributors are wanted: build in the `pull_request` workflow and deploy from a separate `workflow_run`-triggered workflow, which executes in the base-repo context and does get a write token. Deliberately not using `pull_request_target` here — that would run unreviewed PR code with write permissions. Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
6105fe84e9 |
[Examples]: MiniMax-M3 DSpark (#1965)
### What does this PR do?
Type of change: new example + bug fixes
Adds a **MiniMax-M3 DSpark streaming training recipe** under
`tools/launcher/examples/MiniMaxAI/MiniMax-M3/`, plus the four fixes it
needs to actually run. Each fix addresses a failure mode that is silent
or misleading without it:
1. **Gemma-style final norm for the fake base**
(`modeling_final_norm.py`, `modeling_fakebase.py`): M3 uses a
gemma-style final RMSNorm (`(1 + weight)` scale, fp32
multiply-then-cast). Selecting the norm by `model_type` alone picks
plain `rmsnorm` — MiniMax's VL remote code coerces its `model_type`-less
`text_config` to **mixtral** — which silently drops the `+1` and
corrupts the distillation target. New `_FinalGemmaRMSNorm` + selection
by the explicit `use_gemma_norm` config flag (only MiniMax sets it).
2. **Loud failure for un-maskable `answer_only_loss`**
(`hf_streaming_dataset.py`): with a fast tokenizer whose chat template
has no `{% generation %}` tags,
`apply_chat_template(return_assistant_tokens_mask=True)` only warns and
returns an **all-zero mask** — training runs at zero loss on every
sample with no other symptom. Now raises at tokenization with an
actionable message.
3. **`SERVE_BLOCK_SIZE` knob** (`train_eagle_streaming.sh`): nemo_run
exports env values unquoted, so a multi-token
`SERVE_EXTRA_ARGS="--trust-remote-code --block-size 128"` loses
everything after the first token. M3's MSA sparse attention requires KV
block 128 (`ValueError: No common block size for 16` at engine init
otherwise), so `--block-size` gets a dedicated single-token knob.
4. **Relax the speculative_decoding `transformers` pin to `<5.13`**
(match `pyproject.toml`): the old `<5.4` pin downgrades recent vLLM
containers (e.g. transformers 5.12.1, which also provides in-tree
`minimax_m3_vl`) and breaks `vllm serve` (`ALLOWED_LAYER_TYPES` needs
>=5.5.3).
The example itself encodes the validated M3 specifics: generation-tagged
chat template copy (required for `answer_only_loss` — see fix 2), draft
dims + base `rope_theta=5e6` set explicitly (not inherited), mask token
200063 (reserved slot; added tokens end at 200060), `EAGLE_CAPTURE_IDS`
= draft default `target_layer_ids+1` + final layer, `trust_remote_code`
at serve/export, and AWS-EFA NIXL notes (UCX segfaults at agent init on
EFA nodes; LIBFABRIC required there).
### Usage
```bash
cd tools/launcher
export SLURM_HOST=localhost SLURM_ACCOUNT=<account> SLURM_PARTITION=<partition> \
SLURM_HF_LOCAL=<hf_models_dir> SLURM_JOB_DIR=<experiments_dir> NEMORUN_HOME=$PWD
uv run launch.py --yaml examples/MiniMaxAI/MiniMax-M3/hf_streaming_dspark_multi_node.yaml \
identity=$HOME/.ssh/id_ecdsa detach=True --yes
```
### Testing
- `tests/unit/torch/speculative/plugins/`: **129 passed** inside a
current vLLM x86_64 nightly container (transformers 5.12.1), including 3
new tests for the assistant-mask guard.
- Generation-tagged template verified against real corpus samples:
`input_ids` identical to the original template, mask covers exactly the
assistant turns (think prefix + content + eos), contiguous, no
user-prompt leak.
- The recipe is exercised end-to-end by a live M3 DSpark training run
(this yaml modulo cluster paths): streaming serve + NIXL transport +
resume all healthy; drafter MT-Bench AL exceeds our Kimi-K2.6 DSpark
reference by ~8k steps.
- `ruff check` / `ruff format` (0.15.20) clean.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (norm selection only changes
models with `use_gemma_norm=True`, previously mis-normed; the guard
turns a silent zero-loss run into an error; `SERVE_BLOCK_SIZE` is
opt-in; the pin relax widens the allowed range)
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ (can add if desired)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added Gemma-style RMS normalization support for speculative decoding.
- Added MiniMax-M3 multi-node training configuration, including a full
Jinja chat template for tool calls, multimodal content, and thinking
modes.
- Added optional `SERVE_BLOCK_SIZE` support for vLLM serve launches.
- **Bug Fixes**
- Improved `answer_only_loss` masking validation: now fails fast with
clear errors when required `{% generation %}` markers are missing or
when using a non-fast tokenizer.
- **Compatibility**
- Expanded the supported Transformers version range for speculative
decoding examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
7ed19b2906 |
[Feat]: SWA (sliding-window attention) support for DFlash drafter training (#1960)
### What does this PR do?
Type of change: new feature
Adds sliding-window attention (SWA) support to DFlash drafter
**training** in ModelOpt, so a drafter can be trained to match a
target/inference regime that uses a bounded attention window.
Scope (this PR): **all draft layers use non-causal SWA (MiMo-style)** —
each draft query only attends to context positions within
`dflash_swa_window_size` tokens before it, while block-internal
attention stays bidirectional (unchanged). This is the minimal,
config-driven path; block-internal attention is left un-windowed and the
config enforces `window >= block_size`, so a full block always fits
inside the window.
The semantics are aligned end-to-end with the latest vLLM
`_resolve_layer_attention`: export writes `sliding_window` +
`dflash_config.{use_swa, swa_window_size, causal=false}` with
`layer_types` left all `full_attention`, which makes vLLM apply a
non-causal sliding window to every draft layer at inference — so train
and inference match.
Changes:
- `config.py`: new `dflash_swa_window_size` field (None disables) +
validation (`>= dflash_block_size`).
- `dflash_model.py`: base `modify()` reads the field.
- `hf_dflash.py`: `_build_draft_attention_mask` adds the context window
lower-bound; training `forward` passes the window;
`pseudo_speculative_generate` (AR validation) builds the same windowed
mask.
- `hf_spec_export.py`: emits vLLM's SWA fields into the exported draft
config.
### Usage
```yaml
# In the DFlash recipe config:
dflash_swa_window_size: 2048 # each draft query attends to <= 2048 context tokens before it
# None (default) keeps full attention over all context
```
Training then applies the window mask automatically, and `mto.export`
writes the SWA fields so vLLM (with the hybrid-SWA-DFlash support)
applies the same window at inference.
### Testing
<img width="1672" height="1141" alt="image"
src="https://github.com/user-attachments/assets/2ec9441a-6b7f-4666-9432-d10b760bb69d"
/>
- New unit tests in
`tests/unit/torch/speculative/plugins/test_hf_dflash.py`:
- `TestDFlashSwaMask`: context beyond the per-query window is masked;
the windowed mask is a strict subset of full attention; `window <
block_size` is rejected at config validation.
- `TestDFlashExporter::test_export_swa_fields`: exported config carries
`sliding_window` + `dflash_config.{use_swa, swa_window_size, causal}`,
and pre-existing `dflash_config` keys survive the update.
- Full file: 42 passed. Offline/Domino/DSpark plugin unit tests: 31
passed (no regression). `ruff check` clean.
- Smoke: tiny-model training forward with a window runs, loss is finite,
gradients flow, and `pseudo_speculative_generate` works; the windowed
run's grad norm differs from the full-attention run (window is not a
no-op).
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (`dflash_swa_window_size`
defaults to None = existing full-attention behavior)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ (can add if desired)
- Did you get Claude approval on this PR?: ❌ (pending)
### Additional Information
Only all-layer non-causal SWA is implemented; hybrid SWA/full and
causal-SWA (gemma-4 style) are intentionally out of scope and would need
per-layer mask dispatch. Inference requires a vLLM build with DFlash SWA
support.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Summary by CodeRabbit
* **New Features**
* Added optional sliding-window (SWA) attention for DFlash draft layers
during training and generation.
* Extended SWA window masking to Domino and DSpark, ensuring consistent
behavior across training and inference.
* Updated export behavior so SWA/sliding-window settings are included in
exported `config.json` when enabled.
* **Bug Fixes**
* Validate that `dflash_swa_window_size` cannot be smaller than the
DFlash block size.
* **Tests**
* Added unit tests covering SWA mask correctness, exporter fields, and
propagation to DFlash/Domino/DSpark.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
d69d5aab8b |
[Examples]: Kimi-K2.6/K2.7-Code Dflash/Dspark (#1934)
### What does this PR do? **Type of change:** New feature (recipe + examples) Adds the DSpark training recipe and Kimi launcher examples for streaming speculative-decoding training. Split out of #1849 to keep that PR focused on the DSpark head implementation. **Files added (5, all additive — no code changes):** 1. `modelopt_recipes/general/speculative_decoding/dspark.yaml` — the DSpark training recipe. Selects the DSpark head via `dflash_architecture_config.projector_type=dspark` (DFlash backbone + lightweight sequential/Markov head + optional confidence head) and sets the three-term loss weights (`dflash_ce_loss_alpha` / `dflash_l1_loss_alpha` / `dflash_confidence_head_alpha`). 2–5. `tools/launcher/examples/moonshotai/{Kimi-K2.6,Kimi-K2.7-Code}/hf_streaming_{dflash,dspark}_multi_node.yaml` — four multi-node streaming launcher examples mirroring the existing Kimi-K2.5 streaming format (`common/eagle3/train_eagle_streaming.sh`). They carry the Kimi-specific base path, draft dims, capture ids, and mask token; the DSpark examples build on `dspark.yaml` and override only the Kimi-specific fields (draft dims, `dflash_block_size=8`, mask token). K2.7-Code shares the K2.6 architecture (`kimi_k25`, 61 layers), so the draft config is identical. ### Dependency **Stacked on #1849.** The recipe uses the config fields (`dflash_ce_loss_alpha`, `dflash_l1_loss_alpha`, `dflash_confidence_head_alpha`) added in #1849, so this PR is based on that branch and must merge **after** it. GitHub will auto-retarget the base to `main` once #1849 merges. ### Testing - `dspark.yaml` passes the `validate modelopt recipes` schema check (against the #1849 config schema). - The launcher examples ship a placeholder `container: <vllm-image-with-aux-capture-fix>` and are meant to be adapted per cluster (image, account, partition), not run verbatim in CI. - Backward compatible?: ✅ (additive recipe + example files only) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a DSpark speculative-decoding training recipe with configurable draft, optional per-position confidence, and Markov head components. * Added new multi-node streaming training example workflows for Kimi-K2.6 using DSpark/DFlash, including dataset preparation plus coordinated serve/train settings. * Configured answer-only loss options, masking behavior, training hyperparameters, checkpoint/log cadence, and capture-id wiring with sensible default runtime timeouts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
e96d7a21e8 |
[Chore]Dspark license (#1948)
### What does this PR do? Type of change: Chore (license compliance) Adds the required third-party attribution for the DSpark plugin code adapted from [DeepSpec](https://github.com/deepseek-ai/DeepSpec) (MIT), per the "Copying code from other sources" guidance in `CONTRIBUTING.md`: - Fix the header order in `hf_dspark.py` / `modeling_dspark.py`: upstream ref (with commit hash) → DeepSpec MIT → NVIDIA `Apache-2.0 AND MIT`. - Add `Copyright (c) 2026 The DeepSpec Authors` to the MIT section of `LICENSE`. - Exclude both files from the `insert-license` pre-commit hook so it no longer prepends an extra NVIDIA header. ### Usage N/A — no functional or API change. ### Testing `pre-commit run insert-license` on both files → skipped (excluded), files unchanged. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated license notices to include a new third-party attribution. * Adjusted repository checks so license headers are not added to two specific Python modules. * **Documentation** * Refreshed source attribution comments in two modules to point to the latest reference. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
d290839dc9 |
[Feat]: Support Dspark (#1849)
### What does this PR do?
Type of change: New feature
Adds **DSpark** (DeepSeek-AI, *"Confidence-Scheduled Speculative
Decoding with Semi-Autoregressive Generation"*) as a third head in the
DFlash family. DSpark keeps the parallel DFlash backbone for speed and
adds a lightweight sequential **Markov head** that injects the
intra-block causal dependency the parallel backbone lacks (mitigating
suffix acceptance decay), plus an optional **confidence head**. The
Markov head adds a prefix-dependent transition bias `B_k` to the
backbone base logits, inducing a causal block distribution `p_k(x_k |
x_<k) = softmax(U_k + B_k)`.
- Reuses the DFlash mode/pipeline; selected via
`dflash_architecture_config.projector_type=dspark`.
- Three head variants: `vanilla` (memoryless low-rank transition),
`gated` (hidden-gated), `rnn` (GRU-like, full-prefix).
- Trained with a three-term loss: `ce_alpha·CE + l1_alpha·TVD +
conf_alpha·confidence_BCE`.
- Export preserves the upstream DeepSpec submodule names so checkpoints
stay portable.
New files: `plugins/hf_dspark.py`, `plugins/modeling_dspark.py`, recipe
`dspark.yaml`, unit tests. Also touches `config.py` (loss-weight / head
fields), `dflash/conversion.py`, `plugins/__init__.py`, and
`hf_spec_export.py` (export support).
### Usage
```yaml
# modelopt_recipes/general/speculative_decoding/dspark.yaml (key fields)
dflash:
dflash_architecture_config:
projector_type: dspark # select the DSpark head
dflash_ce_loss_alpha: 0.1
dflash_l1_loss_alpha: 0.9 # L1/TVD-dominant (DeepSpec default)
dflash_confidence_head_alpha: 1.0 # 0 disables the confidence head
```
### Testing
`tests/unit/torch/speculative/plugins/test_hf_dspark.py` covers:
convert, all three head variants, confidence-head build/grads, forward
loss + metrics + grads, exported weight-key layout, and that plain
DFlash mode is unaffected. Train / export / AR-generate smoke-validated
end-to-end.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ <!-- additive: new head behind
projector_type=dspark; DFlash/EAGLE modes unchanged -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ / N/A <!-- update before ready-for-review -->
- Did you get Claude approval on this PR?: ❌ / N/A <!-- run /claude
review before ready-for-review -->
### Additional Information
Related: #1846 — DSpark shares the DFlash offline base-logit
reconstruction path (final-norm fix); rebase on top of #1846 once it
merges.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added end-to-end DSpark speculative decoding support, including DSpark
draft-model export and a Markov-style head with selectable variants.
* Introduced DSpark-specific training recipe settings, including
optional confidence scoring and configurable loss weights.
* **Bug Fixes**
* Improved DSpark model conversion routing and validation to prevent
unsupported or incomplete DSpark configurations.
* Updated DSpark export configuration and checkpoint layout to match the
expected DSpark format.
* **Tests**
* Added unit tests covering DSpark conversion, training forward behavior
(including confidence gradients), and DSpark export contents.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
bc5bc1ac5f |
[Feat]: Add Final Norm for vLLM Hidden Extractor (#1846)
### What does this PR do? **Type of change:** Bug fix vLLM captures the final-layer hidden state *before* the model's final norm, but the offline/streaming distillation path fed it straight into `lm_head`, so the reconstructed base logits (the KD target) were computed from un-normed hidden states. This PR re-applies the base model's final norm before `lm_head` when the producer declares a pre-norm capture (`base_hidden_prenorm`), for both DFlash and EAGLE: - Producer sets `base_hidden_prenorm` (streaming: `True`; offline: from the dump). - Consumer (`_maybe_apply_base_final_norm`) re-applies the base final norm, and **fails loud** if pre-norm is declared but the model's norm type isn't supported (no silent corruption). - `FakeBaseModel` now loads the base final norm (+ `rope_theta`/`rms_norm_eps`); norm type is gated by an explicit `model_type` allowlist (gpt_oss excluded pending a matching norm class). ### Testing `tests/unit/torch/speculative/plugins/test_modeling_final_norm.py`; DFlash/EAGLE streaming training verified end-to-end. - Backward compatible?: ✅ (post-norm captures declare `base_hidden_prenorm=False` → unchanged) - New tests: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Expanded speculative decoding/distillation to reconstruct missing base-model logits by optionally applying a base model’s final pre–LM-head normalization when pre-norm hidden states are provided. * Streaming dataset now emits `base_hidden_prenorm` and can configure RDMA backends from environment settings. * **Bug Fixes** * Rejects mixed `base_hidden_prenorm` values within a batch. * Fails fast on streaming token-length mismatches. * Streaming dataset loading supports directory inputs by expanding sorted JSONL shards (and errors if none are found). * **Tests** * Added/updated unit and dataset tests for final-norm behavior and the new batch field. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
c248dd5434 |
[Feat]: Domino support (#1710)
### What does this PR do? Type of change: New feature Adds **Domino** speculative decoding: the parallel DFlash draft backbone plus a lightweight **GRU causal correction head**. The backbone produces *base* logits for a full draft block in one forward; a GRU over the block's teacher-forced tokens produces a causal state that is fused with the backbone hidden state and projected to a vocab-sized logit correction on the block suffix — injecting the intra-block causal dependency the parallel backbone lacks. Trained with a dual loss `(1-λ)*final + λ*base`, where `λ_base` decays linearly 1→0 (curriculum: learn the parallel backbone first, then the correction). Reuses the DFlash mode/config/recipe; selected via `dflash_architecture_config.projector_type=domino` and routed to its own registry so `HFDominoModel` does not shadow `HFDFlashModel`. Exports in the z-lab/SpecForge drafter format (`prefix_gru.*` / `embed_proj.*`). > Note: the inference side (vLLM / AR evaluation) is intentionally **not** wired up yet — the correction head is not applied in serving. To be added once the inference path lands. ### Usage ```bash # Online training (recipe: projector_type=domino) uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_online_domino.yaml --yes ``` ### Testing CPU unit tests in `tests/unit/torch/speculative/plugins/test_hf_domino.py` cover conversion routing, the training forward (dual loss + grads), the λ schedule, and the export format. Online Qwen3-8B training validated end-to-end (loss curve below). <img width="1803" height="809" alt="image" src="https://github.com/user-attachments/assets/7c9d2001-bd80-4dec-919b-443e61089cca" /> ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (opt-in via `projector_type=domino`; DFlash path unchanged) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new dependency) - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Reference: SpecForge PR #571 (z-lab); drafter format `huggingface.co/Huang2020/Qwen3-8B-Domino-b16`. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added Domino speculative-decoding training with a decaying base/final dual-loss curriculum and Domino-specific lambda scheduling (training-only; inference wiring not yet included). * Added Domino draft-head export support for training checkpoints. * **Documentation & Configuration** * Added a Domino speculative-decoding training recipe and an HF Online Domino launcher configuration for Qwen3-8B. * **Refactor** * Updated speculative model conversion/export to route to Domino variants based on the configured projector type. * **Tests** * Added unit tests for Domino conversion, training loss/metrics, lambda decay behavior, and exporter output layout. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
3003589379 |
[Chore]: Add license for Dflash code (#1837)
### What does this PR do? Type of change: Chore / documentation (license compliance) The DFlash implementation in `modelopt/torch/speculative/plugins/` is adapted from [SpecForge](https://github.com/sgl-project/SpecForge) (MIT licensed). This PR adds the required third-party attribution per the [Copying code from other sources](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources) guidance: - **`hf_dflash.py`** ← adapted from `specforge/core/dflash.py` - **`modeling_dflash.py`** ← adapted from `specforge/modeling/draft/dflash.py` Changes: - Add upstream attribution header (source link with commit hash + SpecForge MIT copyright/license notice) above the NVIDIA Apache-2.0 header in both files. - Update `SPDX-License-Identifier` to `Apache-2.0 AND MIT` in both files. - Add `Copyright (c) 2025 sgl-project` to the MIT section under *Third-Party Software Notices* in `LICENSE`. - Exclude both files from the `insert-license` pre-commit hook so the NVIDIA header is not auto-inserted above the upstream header. ### Usage N/A — no functional/API change (comments and license metadata only). ### Testing - `pre-commit run insert-license-py --files <both files>` → reports **Skipped** (files correctly excluded). - Verified `SPDX-License-Identifier: Apache-2.0 AND MIT` is present in both files and the upstream header precedes the NVIDIA header. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ ### Additional Information Upstream pinned at SpecForge commit `8ea5ca6`. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated third-party license notices to include an additional attribution. * Added license/header text to two Python modules for clearer provenance. * Adjusted the pre-commit configuration so those files are skipped by the license insertion check. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
9048d13b86 |
[Feat]:Support DPace (#1724)
### What does this PR do? Type of change: New feature Adds the **D-PACE** (Dynamic Position-Aware Cross-Entropy) loss objective for DFlash speculative-decoding training ([arXiv:2605.18810](https://arxiv.org/abs/2605.18810)). It replaces the static exponential position decay with per-position CE weights derived from the draft's own confidence `q_i = exp(-CE_i)`: smoothed `q̃_i = (1-α)q_i + α` (Eq.7) and weighted by the suffix-sum of prefix products `w_j = Σ_{m≥j} ∏_{i≤m} q̃_i` (Eq.8), which directly targets expected accepted block length and shifts signal toward whichever positions currently limit acceptance. Selected via `dflash_loss_objective` — **D-PACE is now the default** (`dpace`); set `dflash_loss_objective: decay` to restore the previous static schedule. Smoothing via `dflash_dpace_alpha` (default 0.5). Weights are detached from the gradient — training-only, ~2.3% overhead, no architecture or inference change. Mutually exclusive with `dflash_loss_decay_factor`. ### Usage ```yaml # DFlash recipe / training config dflash: dflash_loss_objective: dpace # default: decay dflash_dpace_alpha: 0.5 # smoothing in (0, 1]; stable in [0.3, 0.7] ``` ### Testing CPU unit tests in `tests/unit/torch/speculative/plugins/test_hf_dflash.py`: weights match the paper closed form, are detached and non-increasing, the α smoothing floor keeps later weights non-zero, and convert wires/validates the new fields (rejects bad objective and degenerate α). Training validated on Qwen3-8B (curve below). <img width="1803" height="809" alt="image" src="https://github.com/user-attachments/assets/d34dcd76-9e46-4051-94d4-c880b1987965" /> ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ⚠️ Behavior change — D-PACE is now the **default** objective, so DFlash training loss weighting changes unless you set `dflash_loss_objective=decay` (which reproduces the previous static-decay behavior). - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new dependency) - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Reference: D-PACE, [arXiv:2605.18810](https://arxiv.org/abs/2605.18810). See `examples/speculative_decoding/doc/dflash.md` for the math and tuning notes. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added a **D-PACE** training loss objective for DFlash speculative decoding (`dflash_loss_objective: dpace`), configurable via `dflash_dpace_alpha` (default `0.5`). * **Documentation** * Documented D-PACE’s confidence-derived, dynamically weighted per-position loss behavior (training-only) and noted that `dflash_loss_decay_factor` is ignored with D-PACE. * **Bug Fixes** * Updated DFlash loss to reuse the precomputed per-token cross-entropy in the non-KD path. * **Tests** * Added unit tests for D-PACE weight correctness, masking, gradient detachment, monotonicity, smoothing, and config/validation behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
977d34dc3c |
[Fix](nvbug6304585): specdec README online base-model example should use Instruct model (#1755)
### What does this PR do?
Type of change: Bug fix
The **Training Draft Model with Online/Offline Base Model** examples in
`examples/speculative_decoding/README.md` used
`meta-llama/Llama-3.2-1B`, a
base / pretrained checkpoint that ships **no chat template**. The online
EAGLE3
flow tokenizes conversations through
`tokenizer.apply_chat_template(...)`, so
the data collator fails fast at startup:
```
ValueError: No valid chat template!
```
This PR:
- Switches both README example commands (online and offline) to
`meta-llama/Llama-3.2-1B-Instruct`, which carries a chat template.
- Makes the collator error message in
`modelopt/torch/utils/plugins/transformers_dataset.py` actionable — it
now
explains the cause (base checkpoints have no chat template) and points
users
at an Instruct model or a custom `chat_template`.
### Usage
```bash
./launch_train.sh \
--config ../../modelopt_recipes/general/speculative_decoding/eagle3.yaml \
model.model_name_or_path=meta-llama/Llama-3.2-1B-Instruct \
data.data_path=input_conversations/train.jsonl \
training.output_dir=ckpts/llama-3.2-1b-online
```
### Testing
- Reproduced the original `No valid chat template!` failure with the
base
`Llama-3.2-1B` and confirmed the Instruct variant carries a chat
template.
- Verified the new error message renders correctly when a template is
missing.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌
### Additional Information
Fixes nvbug 6304585: https://nvbugspro.nvidia.com/bug/6304585
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **Documentation**
* Updated speculative decoding example training commands to reference
the Llama-3.2-1B-Instruct model.
* **Bug Fixes**
* Enhanced error message when chat template configuration is missing,
providing actionable guidance for resolution.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
e6790ef7b4 |
[Examples]: GPT-oss, Qwen3Moe streaming specdec example (#1692)
### What does this PR do?
Type of change: new example
Adds **streaming speculative-decoding examples (EAGLE3 + DFlash)** for
**gpt-oss-20b** and **Qwen3-30B-A3B** to the ModelOpt launcher,
mirroring the existing Qwen3-8B/Kimi examples.
- New yamls:
`tools/launcher/examples/{openai/gpt-oss-20b,Qwen/Qwen3-30B-A3B}/hf_streaming_{eagle3,dflash}_multi_node.yaml`,
plus gpt-oss `chat_template_train.jinja` (generation-tagged, for
`answer_only_loss`).
- `eagle_utils.py`: the streaming path now installs a custom
`data.chat_template` on the tokenizer (the online path already did) —
needed for the tagged template.
### Usage
```bash
cd tools/launcher
export SLURM_HOST=... SLURM_ACCOUNT=... SLURM_HF_LOCAL=... SLURM_JOB_DIR=...
uv run launch.py --yaml examples/openai/gpt-oss-20b/hf_streaming_eagle3_multi_node.yaml --yes
```
### Testing
Pipeline sanity test on **unsynthesized** data (daring-anteater), 1×
H100-80GB, 12k steps. All four train and pass the vLLM acceptance-length
eval:
| Model | Method | Train speed | vLLM AL |
|---|---|---|---|
| Qwen3-30B-A3B | EAGLE3 | 7.12 it/s | **1.74** |
| Qwen3-30B-A3B | DFlash | 2.31 it/s | 1.29 |
| gpt-oss-20b | EAGLE3 | 5.07 it/s | 1.19 |
| gpt-oss-20b | DFlash | 2.01 it/s | 1.14 |
<img width="1300" height="780" alt="image"
src="https://github.com/user-attachments/assets/2be22562-5b77-4a9f-9dc0-6f936a059736"
/>
> Sanity test only, not a quality run. gpt-oss AL is low because it is a
reasoning model (CoT at inference) while daring-anteater has no
reasoning traces and `answer_only_loss` masks all but the final content
— quality runs need synthesized/reasoning data.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added multi-node speculative decoding pipeline configurations for
Qwen3-30B-A3B and gpt-oss-20b with DFlash and EAGLE3 support.
* Introduced chat template training support for improved model
instruction formatting.
* **Enhancements**
* Increased benchmark concurrency from 1 to 32 across Qwen3-8B
configurations for more realistic performance evaluation.
* Extended training runs from 500 to 2000 steps for Kimi-K2.5 models.
* Improved chat template handling in speculative decoding workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
cfc823d127 |
[Tests]: Precommit Check for Spec-Dec Recipes (#1527)
### What does this PR do? Type of change: new tests / tooling Adds pre-commit validation for speculative-decoding recipes (the existing `check-modelopt-recipes` hook only ran on PTQ) and for launcher YAML references into the recipe library. - `tools/precommit/check_modelopt_recipes.py`: accept `speculative_eagle` / `speculative_dflash` / `speculative_medusa` in addition to `ptq`, so per-model spec-dec recipes (e.g. `modelopt_recipes/models/Qwen3-8B/dflash.yaml`) get full Pydantic validation via `load_recipe()` at commit time. - `tools/precommit/check_launcher_yaml.py` (new): scans every `tools/launcher/examples/**/*.yaml` for `--config <path>` and `data.chat_template=<path>` references, verifies the resolved files exist, and runs `load_recipe()` on any path under `modelopt_recipes/`. Skips `<<global_vars.x>>` interpolation. `pass_filenames: false` so recipe-side edits also re-validate all launcher references. ### Usage ```bash pre-commit run check-modelopt-recipes --all-files pre-commit run check-launcher-yaml --all-files ``` ### Testing Smoke-tested both hooks manually: | Scenario | Result | |---|---| | spec-dec recipe with `dflash_block_size: not_an_int` | exit 1, Pydantic int_parsing error | | launcher YAML with non-existent `--config` path | exit 1, source file + resolved path reported | | launcher YAML with non-existent `data.chat_template` path | exit 1 | | Current repo state (all valid) | exit 0 | ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ (hooks themselves are the tests) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ pending ### Additional Information Motivated by the per-model recipe migration in #TBD — without these hooks, broken `--config` paths and recipe schema typos surface only at CI or runtime. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Added an automated pre-commit check that validates launcher example YAMLs, reporting parse errors and missing or invalid references. * Expanded recipe validation to cover additional recipe types beyond PTQ, improving detection of invalid recipe formats and metadata. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
46eddab877 |
[Feat]: Specdec Streaming: RDMA + Multinode (#1611)
### What does this PR do? Type of change: New feature Multi-node **streaming** training for speculative decoding (EAGLE3 / DFlash): a live `vllm serve` captures the target model's hidden states and moves them straight to the trainer over **NIXL RDMA** — no disk round-trip. The streaming dataset is map-style — each rank fetches only its own `DistributedSampler` shard (concurrency from `dataloader_num_workers`), round-robins across multiple serve replicas (`server_urls`), and scales to multi-node DDP. Serve-side tensor parallelism (TP>1) is supported: hidden states are replicated across TP ranks, so rank 0 alone owns the pool + transfer. ### How - `RdmaHiddenStatesConnector` — out-of-tree vLLM connector (no vLLM source edits): one pre-registered pinned NIXL pool per serve, a ring slot per request, and a small HTTP sidecar serving transfer metadata. The trainer RDMA-READs the slot into a per-worker buffer. RDMA is the **only** transport (the earlier disk/safetensors path is removed). - Map-style dataset + multi-node accelerate launch (`--machine_rank`, optional Slurm `--segment` to keep nodes in one NVLink domain). ### Usage ```yaml data: mode: streaming streaming_server_url: "http://node0:8000,http://node1:8000" # round-robin ``` ### Validation (Qwen3-8B, oci-nrt H100) sandbox CI: https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/337489812 **1. End-to-end convergence — EAGLE3 & DFlash, 5000 steps.** Both algorithms converge and export a deployable draft; the DFlash drafts also serve under vLLM speculative decoding (8/8 smoke prompts pass). | algorithm | topology (nodes) | train loss (step 0 → 5000) | vLLM draft acc-len | |---|---|---|---| | EAGLE3 | 2 serve TP=2 + 2 trainer DDP (4) | 37.1 → 8.20 | — | | DFlash | 1 serve TP=1 + 1 trainer (2) | 11.7 → 5.56 | 1.11 | | DFlash | 2 serve TP=2 + 2 trainer DDP (4) | 10.9 → 5.26 | 1.19 | <!-- Drag these PNGs in here (GitHub turns them into asset URLs): eagle3_streaming_loss.png, dflash_streaming_loss_singlenode.png, dflash_streaming_loss_multinode.png --> **2. Scalability — 1 → 12 nodes (EAGLE3, 200 steps).** Throughput scales ~23× across the sweep below. The step-time growth is cross-node DDP all-reduce, not the streaming path — RDMA (~0.33 ms/req @ 2 MB, ~47 GB/s host-pinned READ) is never the bottleneck. Scale serve + trainer nodes together for near-linear speedup. | serve / trainer | nodes | step time | samples / step | samples / sec (global) | acc @ step 200 | |---|---|---|---|---|---| | 1 serve / 1 rank (co-located, 1 node 2 GPU) | 1 | 0.23 s | 1 | 4.4 | [0.141, 0.094, 0.072] | | 1 serve / 1 rank (cross-node) | 2 | 0.23 s | 1 | 4.3 | [0.137, 0.105, 0.074] | | 2 serve / 8 ranks | 3 | 0.26 s | 8 | 31.1 | [0.215, 0.126, 0.097] | | 4 serve / 16 ranks (2 trainer nodes) | 6 | 0.28 s | 16 | 56.5 | [0.217, 0.148, 0.110] | | 8 serve / 32 ranks (4 trainer nodes) | 12 | 0.31 s | 32 | 101.9 | [0.235, 0.165, 0.137] | **3. Serve-side TP correctness.** TP=1 vs TP=2 draft top-1 accuracy track step-for-step (hidden states are replicated across TP ranks). <img width="910" height="546" alt="serve-tp-acc" src="https://github.com/user-attachments/assets/73df9214-7ff0-4ab4-bf2f-95842b12cd5f" /> ### Before your PR is "*Ready for review*" - Backward compatible?: ❌ — streaming is now RDMA-only; `server_url` → `server_urls`; the disk transport (`HS_TRANSPORT`, `streaming_shared_storage_path`) is removed. - New tests?: ✅ `tests/unit/torch/speculative/plugins/test_hf_streaming_dataset.py` (map-style dataset + mocked RDMA fetch). - Updated Changelog?: ❌ - Claude approval?: ❌ --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
902d36921a |
[Feat]: Streaming Hidden-states Dataset (#1509)
### What does this PR do? Type of change: new feature **Design doc:** https://gist.github.com/h-guo18/241c94968b0591324c361d97cf995dd0 Jira ticket: https://jirasw.nvidia.com/browse/OMNIML-4341 Sandbox CI: https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/jobs/327911555#L2228 Streaming hidden-states dataset: per-sample activations pulled from a live `vllm serve` over HTTP, replacing on-disk activation dumps. Two axes for future extensions: - **Backend** (`_fetch`): vLLM now; TRT-LLM / SGLang next. - **Algorithm** (`_format`): Eagle now; distillation / probing next. - **Sandbox CI**: https://gitlab-master.nvidia.com/omniml/integration/nmm-sandbox/-/merge_requests/169 Shared plumbing — async producer, token-level truncation to `training_seq_len`, loss-mask alignment, DDP via Accelerate's dispatcher (rank 0 fetches, broadcasts), circuit breaker, resume — lives in `StreamingDataset`. First instance: **`EagleVllmStreamingDataset`**. API: `data.mode ∈ {online, offline, streaming}`; legacy configs auto-promote. ### Usage ```yaml data: mode: streaming data_path: input_conversations/train.jsonl streaming_server_url: http://localhost:8000 streaming_model_name: meta-llama/Llama-3.1-8B-Instruct training: training_seq_len: 4096 # also caps the prompt sent to vllm ``` Requires `vllm serve` with `ExampleHiddenStatesConnector` and `dataloader_num_workers=0`. End-to-end Slurm pipeline: `tools/launcher/examples/Qwen/Qwen3-8B/hf_streaming_eagle3.yaml`. ### Testing - **Unit**: full-corpus invariant, rank-0-only iter, resume, circuit breaker, mocked-httpx integration. - **E2E**: `launch_train.sh` against a stdlib `HTTPServer` mimicking the connector. - **Smoke** (Qwen3-8B / 8×H100 / 4096 ultrachat samples, single epoch): train_loss 32 → 18, MT-Bench AR 1.003 → 1.20. ### TODO before un-drafting - [ ] Observability counters (filtered, fetch failures, queue depth, latency). - [ ] Changelog entry. - [x] Add test in sandbox. **Non-goals (v1):** multi-epoch streaming, cross-rank dynamic load balancing. ### Before your PR is "*Ready for review*" - Backward compatible: ✅ (legacy configs auto-promote) - New PIP dep: N/A - New tests: ✅ - Changelog: ❌ (TODO above) - Claude approval: ❌ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Streaming training mode with server-backed hidden-state fetching, deterministic seed control, resume support, and streaming-specific dataset options (server, model, prefetch, shared storage). * **Behavior / Bug Fixes** * Stronger mode validation; offline behavior derived from data mode; resume handling adjusted to avoid double-skip during streaming runs. * **Tests** * End-to-end CI streaming test and expanded unit tests covering streaming, resume, DDP, determinism, and failure cases. * **Infrastructure** * Launcher script and pipeline config for end-to-end streaming training. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1509?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
40a4dd326d |
[Feat]: Eagle Dry Run Mode (#1566)
### What does this PR do? Type of change: new feature Adds `--dry_run` to `examples/speculative_decoding/main.py`: load → `mtsp.convert` → save, then exit (no `trainer.train()`). With `FakeBaseModel`, the convert→save→export chain runs in seconds and produces an exportable EAGLE3 / Medusa / DFlash checkpoint with correct structure but untrained draft-head weights — useful for end-to-end plumbing smoke tests on downstream stacks (vLLM, TRT-LLM, SGLang) without paying for a real training run. Where it sits among existing EAGLE3 modes: ``` EAGLE3 modes ├── online base model runs forward in-loop ├── offline reads pre-dumped hidden states from disk ├── streaming streams hidden states from a live server in-loop └── dry-run ★ skip training entirely; convert + save + export ← NEW (this PR) ``` Companion `FakeBaseModel` fixes so small base checkpoints work: - Synthesize the weight_map from a single `model.safetensors` when no sharded index is present (Llama-3.2-1B, Qwen3-0.6B, …). - Honor `tie_word_embeddings`: reuse `embed_tokens` when `lm_head` is absent from safetensors. A new launcher YAML (`tools/launcher/examples/Qwen/Qwen3-8B/hf_eagle3_dryrun.yaml`) wires this together as a one-task pipeline. ### Usage ```bash # Direct python main.py --dry_run \ --config modelopt_recipes/general/speculative_decoding/eagle3.yaml \ model.model_name_or_path=meta-llama/Llama-3.1-8B-Instruct \ model.use_fake_base_for_offline=true \ data.offline_data_path=/tmp/dryrun-placeholder \ training.output_dir=ckpts/dryrun python scripts/export_hf_checkpoint.py --model_path ckpts/dryrun --export_path export/dryrun # Launcher uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_eagle3_dryrun.yaml --yes ``` ### Testing - `tests/unit/torch/speculative/plugins/test_fakebase.py`: 3 new cases — single-file fallback, tied-embeddings fallback, and the negative case (missing `lm_head` without tying). - `tests/examples/speculative_decoding/test_eagle.py::test_eagle3_dry_run`: full `launch_train.sh --dry_run → export_hf_checkpoint.py` chain on `tiny_llama`; asserts exported state_dict has all `LLAMA_EAGLE_SINGLE_LAYER` required keys. - Manually verified end-to-end on Llama-3.1-8B-Instruct (sharded), Llama-3.2-1B-Instruct (single-file + tied), and Qwen3-0.6B (single-file). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `--dry_run` is opt-in; `FakeBaseModel` changes are additive fallbacks. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — examples-only addition. - Did you get Claude approval on this PR?: ❌ ### Additional Information N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a `--dry_run` CLI flag to perform a fast early-exit execution that saves model artifacts without training. * Added support for single-file model checkpoint formats in the loader. * **Bug Fixes** * Improved checkpoint-loading error messages and handling. * Enhanced tied-embeddings fallback when head weights are absent. * **Tests** * Added integration and unit tests covering dry-run behavior and single-file/tied-embedding loading. * **Documentation** * Added a launcher example for dry-run smoke testing. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1566?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
f59f3ae1c3 |
[Example]: Dflash-Offline Launcher Example (#1529)
### What does this PR do? Type of change: new example Offline DFlash training launcher example for Qwen3-0.6B. Two-task pipeline: dump base-model hidden states via HF forward, then train DFlash on the dump. - `tools/launcher/examples/Qwen/Qwen3-0.6B/hf_offline_dflash.yaml` — new launcher YAML - `tools/launcher/common/eagle3/dump_offline_data_hf.sh` — new HF-backed dump script (DFlash's `answer_only_loss=true` requires `loss_mask`, which the existing TRT-LLM dump backend does not produce) - `examples/dataset/synthetic_conversations_1k.jsonl` — `conversation_id` added at-source (asserted by `compute_hidden_states_*.py`; previously injected at test-fixture time) - `tests/regression/torch/speculative/test_dflash_offline.py` — `tagged_synth_data_path` fixture removed since the field is now at-source ### Usage ``` uv run launch.py --yaml examples/Qwen/Qwen3-0.6B/hf_offline_dflash.yaml --yes ``` ### Testing End-to-end on CoreWeave Slurm (1×1 GPU): task_0 dump succeeded; task_1 training loss 7.85 → 2.94 over 2 epochs, regression PASSED. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (`conversation_id` is purely additive on the dataset) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (existing `test_dflash_offline.py` covers the workflow) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (new example only) - Did you get Claude approval on this PR?: ❌ ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added launcher script for offline hidden-state processing with HuggingFace backend support and distributed execution capabilities * Added training configuration example for Qwen3-0.6B model with DFlash speculative decoding pipeline * **Tests** * Updated offline hidden-states regression test to improve data handling workflow <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1529?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
7038dec918 |
[1/2Refactor] speculative decoding: use mto config subsystem (#1328)
### What does this PR do?
Type of change: new feature
Port the speculative-decoding example to ModelOpt's recipe/config
subsystem: `model` / `data` / `training` / `<algo>` now load from a
single YAML with Pydantic validation and OmegaConf dotlist overrides.
Adds built-in `eagle3` / `dflash` recipes, drops the redundant
`training.mode` field (inferred from recipe class), and shrinks
`main.py` by ~145 lines (−208 / +63).
JIRA: OMNIML-3859
### Usage
```bash
python main.py --config general/speculative_decoding/eagle3 \
model.model_name_or_path=meta-llama/Llama-3.2-1B \
data.data_path=train.jsonl \
training.output_dir=ckpts/test
```
### Testing
- `pytest tests/unit/recipe/test_loader.py` — new coverage for Eagle /
DFlash YAML loading, dotlist overrides, and field-level validation.
- Smoke-trained both built-in `eagle3` and `dflash` recipes end-to-end.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ❌ — `main.py` CLI switched to
`--config <recipe>` (+ dotlist overrides); the old argparse flags are
removed.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
deps (`pydantic`, `omegaconf` already in core).
- Did you write any new necessary tests?: ✅ —
`tests/unit/recipe/test_loader.py`.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — to be added.
### Additional Information
Follow-up to the `modelopt.recipe` subsystem introduced for PTQ; this PR
extends the same declarative-YAML pattern to speculative decoding
(Eagle3 / DFlash / Medusa).
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added typed speculative-decoding recipe support for EAGLE, DFlash, and
Medusa; CLI dotlist overrides supported for single-file recipes.
* Trainer/config schema extended with speculative-training fields and
draft-vocab cache loading for Eagle.
* **Bug Fixes**
* Offline training no longer mutates model configs; loader enforces
required algorithm sections and prints recipe/config only on the primary
process.
* Reduced noisy per-rank logging by restricting status output to the
primary process.
* **Tests**
* Expanded tests for recipe loading, dotlist overrides, validation
strictness, and error cases.
* **Documentation**
* Recipe YAMLs updated with metadata and usage notes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
7f1f223d8f |
[Feat]: Add dflash in specdec-bench (#1432)
### What does this PR do?
Type of change: new feature
Adds **DFLASH** speculative decoding support to the `specdec_bench`
example for both the SGLang and vLLM backends.
- `run.py`: register `DFLASH` in `--speculative_algorithm` choices.
- `models/sglang.py`: unify the SGLang engine setup so `engine_kwargs`
is built once and shared between the speculative and non-speculative
paths; add a DFLASH branch (default `speculative_num_draft_tokens=8`,
optional `speculative_dflash_draft_window_size`); set
`disable_cuda_graph_padding=True` and
`cuda_graph_max_bs=max_concurrent_requests` to avoid CUDA-graph
bucket-padding mismatches during DFLASH replay; expose
`mamba_scheduler_strategy` passthrough (needed for Qwen3.5).
- `models/vllm.py`: add a DFLASH branch wiring `method="dflash"` with
`speculative_num_draft_tokens` (default 8).
### Usage
```bash
python run.py \
--model_dir <target_model> \
--tokenizer <target_model> \
--draft_model_dir <dflash_draft_model> \
--mtbench mtbench.jsonl \
--engine SGLANG \
--speculative_algorithm DFLASH \
--concurrency 32 \
--tp_size 4 --ep_size 4
```
Use `--engine VLLM` for the vLLM backend.
### Testing
Manually validated end-to-end on Kimi-K2.5-NVFP4 + MTBench with both
SGLang and vLLM backends. `specdec_bench` has no existing test
infrastructure, so no automated tests are added (consistent with how
other algorithms in this example are handled).
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (purely additive — new
algorithm choice; existing algorithms unchanged)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ (specdec_bench is
example-only with no test harness; matches other algorithms)
- Did you update Changelog?: N/A (example-only change)
- Did you get Claude approval on this PR?: ❌
### Additional Information
Requires SGLang and vLLM versions that include DFLASH support.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added DFLASH as a supported speculative decoding algorithm option (CLI
selectable).
* DFLASH support extended to multiple inference backends with
configurable draft-token count and optional draft-window size; CLI now
emits an informational note when some options are ignored.
<!-- review_stack_entry_start -->
[](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1432)
<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
9d2e6087d1 |
[Fix]: $HOME in launcher eagle example (#1365)
### What does this PR do? Type of change: Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> Launcher example bug raised by @cjluo-nv Before fix: task1 in tools/launcher/examples/Qwen/Qwen3-8B/hf_online_eagle3.yaml fails Reason: due to `HOME: /tmp` set in container, enroot credentials in `$HOME/.config/enroot/.crendential` not found ``` GpuFreq=control_disabled pyxis: importing docker image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc10 Apr 28 13:35:59.491365 2515157 slurmstepd 0x155552c3b780: error: pyxis: child 2515158 failed with error code: 1 Apr 28 13:35:59.491415 2515157 slurmstepd 0x155552c3b780: error: pyxis: failed to import docker image Apr 28 13:35:59.491433 2515157 slurmstepd 0x155552c3b780: error: pyxis: printing enroot log file: Apr 28 13:35:59.491453 2515157 slurmstepd 0x155552c3b780: error: pyxis: [INFO] Querying registry for permission grant Apr 28 13:35:59.491469 2515157 slurmstepd 0x155552c3b780: error: pyxis: [INFO] Authenticating with user: <anonymous> Apr 28 13:35:59.491483 2515157 slurmstepd 0x155552c3b780: error: pyxis: [INFO] Authentication succeeded Apr 28 13:35:59.491499 2515157 slurmstepd 0x155552c3b780: error: pyxis: [INFO] Fetching image manifest list Apr 28 13:35:59.491512 2515157 slurmstepd 0x155552c3b780: error: pyxis: [INFO] Fetching image manifest Apr 28 13:35:59.491524 2515157 slurmstepd 0x155552c3b780: error: pyxis: [ERROR] URL https://registry-1.docker.io/v2/nvcr.io/nvidia/tensorrt-llm/release/manifests/1.3.0rc10 returned error code: 401 Unauthorized Apr 28 13:35:59.491564 2515157 slurmstepd 0x155552c3b780: error: pyxis: couldn't start container Apr 28 13:35:59.491579 2515157 slurmstepd 0x155552c3b780: error: spank: required plugin spank_pyxis.so: task_init() failed with rc=-1 Apr 28 13:35:59.491593 2515157 slurmstepd 0x155552c3b780: error: Failed to invoke spank plugin stack Apr 28 13:35:59.515523 2515146 slurmstepd 0x155552c3b780: error: pyxis: child 2515240 failed with error code: 1 ``` After fix: ``` GpuFreq=control_disabled pyxis: importing docker image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc10 pyxis: imported docker image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc10 ``` ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated example pipeline to use the standardized dataset example path. * Removed unnecessary per-task overrides of the process home and cache directory to simplify environment setup. * Preserved required model checkpoint environment setting for the relevant task so model resolution continues to work. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
c07ac215ef |
[Fix]: Relax Dflash Rregression Test Threshold fo 2GPUs (#1373)
### What does this PR do? Type of change: Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> The offline dflash regression test can be runned on 1 or 2 gpus. For 2 gpus, the total steps is half of 1 gpu. This PR relax the failing threshold for 2 gpu tests. ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
1ec931c2c7 |
[2/3][Feat]: Offline DFlash training (#1343)
### What does this PR do? Type of change: new feature Part 2 of a 3-PR series splitting #1271: - **[1/3] #1296**: File reorg + deprecate `ParallelDraft` - **[2/3] this PR**: Offline DFlash training (depends on #1296) - **[3/3] #1297**: Extract `HFSpecDecMixin` Changes: - Add `dflash_offline` flag to `DFlashConfig` for training from pre-computed hidden states; deletes base model layers to save memory. - Add Pydantic validators on `DFlashConfig`: - `_derive_dflash_offline` — auto-derive `dflash_offline` from `data_args.offline_data_path` in validation context. Not user-configurable: any user-supplied value is overridden by the derived value. - `_resolve_mask_token_id` — auto-detect `dflash_mask_token_id` from `tokenizer.mask_token_id`. - `_check_mask_token_id` — fail fast if unset after resolution. - `HFDFlashModel.modify()`: select `num_orig_hidden_layers` when offline; pick `_base_model_lm_head` device when no base layers present; drop base-model `layers` module. - `HFDFlashModel.forward()`: add offline branch — consumes precomputed `base_model_outputs` via `DFlashBaseModelOutput.from_offline_dict`, and when `dflash_self_logit_distillation` is enabled with `base_model_logits` absent, recomputes logits from `base_model_hidden_states` via `_base_model_lm_head`. Raises a clear error from the non-training / `pseudo_speculative_generate` paths when `dflash_offline=True`, since base-model layers have been deleted. - `DFlashBaseModelOutput` dataclass in `modeling_dflash.py` (with `from_offline_dict` classmethod) to unify online/offline output shapes. `aux_hidden_states` is required in `from_offline_dict` so missing keys fail fast at the entry point rather than deeper in the forward. - `examples/speculative_decoding/main.py`: replace inline `mask_token_id` auto-detect with `DFlashConfig.model_validate(dflash_cfg, context={"tokenizer": tokenizer, "data_args": data_args})`. ### Silent bug fix — `add_generation_template` → `add_generation_prompt` The pre-refactor `compute_hidden_states_hf.py` passed `add_generation_template=False` to `tokenizer.apply_chat_template`. This kwarg does not exist on HF `apply_chat_template` and was being silently ignored, so the intended "don't append a generation prompt" behavior was never actually applied. The new `tokenize_with_loss_mask` helper in `examples/speculative_decoding/collect_hidden_states/common.py` uses the correct `add_generation_prompt=False`. **This is a real behavior change** for anyone re-dumping hidden states: trailing generation prompts that were previously appended to the tokenized sequences will no longer be included. ### Testing - New tests: - `tests/unit/torch/speculative/plugins/test_hf_dflash_offline.py` — CPU unit tests for convert path (online keeps base layers, offline deletes them; `num_orig_hidden_layers` drives `target_layer_ids` in offline mode) and `DFlashConfig._derive_dflash_offline` validator. - `TestDFlashOfflineForwardGPU` in `tests/gpu/torch/speculative/plugins/test_hf_dflash.py` — GPU forward smoke with precomputed `base_model_outputs`, plus the `dflash_self_logit_distillation` logit-recompute path. - training test: <img width="454" height="317" alt="image" src="https://github.com/user-attachments/assets/79b92790-4d15-4313-bb9b-f35665b012e6" /> <img width="456" height="310" alt="image" src="https://github.com/user-attachments/assets/4558559f-9c35-49ed-b36e-82fbc99eab23" /> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ — additive `dflash_offline` flag defaulting to `False`; validators fall through when context not provided. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — see Testing section above. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ ### TODO (follow-up) - [x] Update `examples/speculative_decoding/collect_hidden_states/compute_hidden_states_*.py` to support DFlash offline data. Current scripts are Eagle-specific — they hardcode the `[2, N/2, N-3]` aux-layer selection and emit `{input_ids, hidden_states, aux_hidden_states}`. DFlash offline needs: - Aux layer indices driven by `build_target_layer_ids(num_orig_hidden_layers, num_draft_layers)` (or a configurable list), not the Eagle triplet. - `base_model_hidden_states` key (last-layer hidden) so `DFlashBaseModelOutput.from_offline_dict` + the `dflash_self_logit_distillation` recompute path can consume it. - Optional `base_model_logits` dump so offline training can skip the self-distillation logit recomputation when logits are available. ### Additional Information Base branch is #1296 (file reorg). Retarget to `main` once #1296 merges. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Offline DFlash speculative-decoding training from precomputed base-model hidden states * Answer-only-loss training with persisted loss masks and optional chat-template support * Flexible auxiliary-layer selection via CLI and an exposed default aux-layer helper * Auto-derived offline flag in config and automatic memory optimization during offline conversion * **Documentation** * Updated guides for offline pipeline, aux-layer selection, and loss-masking options * **Tests** * New unit, GPU, and regression tests covering offline conversion, training, and config derivation <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
7c80d85751 |
[1/3][Refactor]: File reorg; deprecate ParallelDraft (#1296)
### What does this PR do? Type of change: refactoring Part 1 of a 3-PR series splitting #1271: - **[1/3] this PR**: File reorg + deprecate `ParallelDraft` - **[2/3] #1295**: Offline DFlash training - **[3/3] #1297**: Extract `HFSpecDecMixin` Changes: - **File reorg**: `transformers.py` → `hf_eagle.py`; extract `HFMedusaModel` → `hf_medusa.py`; extract `EagleModule` / `EagleBaseModelOutput` → `modeling_eagle.py`; extract `DFlashModule` / `DFlashAttention` / `DFlashDecoderLayer` / `build_target_layer_ids` / `apply_rotary_pos_emb` → `modeling_dflash.py`. - **Deprecate `ParallelDraft`**: remove `parallel_draft_step`, `parallel_draft_heads_num_layers`, and the `ParallelDraft` module from HF Eagle; remove the `EagleMedusaExporter` branch from `HFEagleModel.get_exporter()` (the `EagleMedusaExporter` class itself still lives in `hf_spec_export.py` for Megatron parity). - **Rename**: `_draft_model_config` → `eagle_config` in export plugin. - Update imports in `examples/speculative_decoding/` and `modelopt/torch/speculative/utils.py` to follow the module rename. ### Testing Validated with existing Eagle and DFlash training scripts (re-run after `9ae5302729 revert behavior change`). ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ — renames `modelopt.torch.speculative.plugins.transformers` → `.hf_eagle`; removes `parallel_draft_step` / `parallel_draft_heads_num_layers` from Eagle config; renames `_draft_model_config` → `eagle_config` in export plugin. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — pure refactor; existing tests updated for the rename. `test_hf_spec_rope_export.py` assertions were also corrected to reflect the actual production path (the old assertions were masked by `MagicMock` not invoking the `_draft_model_config` `@property`). - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ ### Additional Information Breaking changes: - `modelopt.torch.speculative.plugins.transformers` → `.hf_eagle` - `parallel_draft_step` / `parallel_draft_heads_num_layers` removed from Eagle config - `_draft_model_config` → `eagle_config` in export plugin <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactoring** * Reorganized speculative-decoding plugins into focused modules, converting the legacy "transformers" entry into a deprecated shim that re-exports the new plugin surface. * Consolidated DFlash implementation into a shared modeling component and introduced a dedicated EAGLE decoder module. * **New Features** * Added a Medusa speculative-decoding plugin with configurable heads and combined-loss training behavior. * **Chores** * Updated pre-commit license-hook exclusion and feature-flag wiring. * **Tests** * Updated export tests to expect rope-scaling fallback semantics. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
6403389eb0 |
Feat: Configurable Eagle ROPE scaling during export (#1238)
### What does this PR do? JIRA ticket: https://jirasw.nvidia.com/browse/OMNIML-3469 Type of change: New feature Decouple EAGLE training rope configuration from export rope configuration, enabling separate YaRN rope scaling injection at export time for long-context inference. #### Changes **Configurable export rope scaling (`EagleConfig`)** - Add `eagle_export_rope_scaling` field to `EagleConfig` with default YaRN config (`factor=32.0`, `original_max_position_embeddings=2048`) - Set to `{}` to disable rope scaling injection at export **Simplified training defaults (`default_config.py`)** - Change default training rope from `llama3` (theta=500k) to `default` (theta=10k) — models now train with simple positional embeddings; rope scaling is applied only at export - Add `rope_theta` inside `rope_scaling` dict for transformers 5.x cross-version compatibility **Move config validation/rewriting into `EagleConfig` (`config.py`)** - `_derive_eagle_offline`: derives `eagle_offline` from `data_args.offline_data_path` via validation context, removing manual assignment in `main.py` - `_check_rope_scaling_consistency`: rejects configs where `eagle_export_rope_scaling` is set but training `rope_type` is not `"default"` - `_warn_rope_vs_training_seq_len`: warns when `original_max_position_embeddings` differs from `training_seq_len` **Export rope injection (`hf_spec_export.py`)** - Inject `eagle_export_rope_scaling` into the exported HF config when training rope_type is `"default"` - Fall back `rope_theta` from `rope_scaling` dict for transformers 5.x compatibility **Fix Megatron RotaryEmbedding crash (`megatron_eagle.py`)** - `dict_to_config()` set `rope_scaling=True` whenever the `rope_scaling` key existed, even without a `"factor"` — causing `RotaryEmbedding` to divide by `None` - Now only enables `rope_scaling` when the dict actually contains a `"factor"` key ### Usage Configure in YAML config (or use defaults from `eagle3.yaml`): ```yaml eagle: eagle_export_rope_scaling: rope_type: yarn factor: 32.0 original_max_position_embeddings: 2048 ``` Set to empty dict to disable export rope injection: ```yaml eagle: eagle_export_rope_scaling: {} ``` ### Testing - New unit tests: `tests/unit/torch/speculative/test_eagle_config.py` — rope consistency validator, seq_len warning, context-derived `eagle_offline` - New unit tests: `tests/unit/torch/export/test_hf_spec_rope_export.py` — export rope injection, fallback, and empty-config cases ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ (new field has sensible default; existing configs work unchanged) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ (should be added if merging as a feature) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Add export-time rope-scaling configuration for EAGLE models. * **Improvements** * Stronger validation and context-aware reconciliation between training and export configs. * Export now injects rope-scaling and rope-theta when appropriate. * Default rope-scaling values updated for EAGLE variants. * Model instances now expose export rope-scaling for downstream use. * **Tests** * Added unit tests covering rope-scaling export behavior and configuration validators. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
82d96a635f |
[Speculative Decoding] Refactor EAGLE3 training to YAML-based config and recipe system (#1134)
## What does this PR do?
Refactors EAGLE3 training to use a single base YAML config with
OmegaConf dotlist overrides.
**Type of change:** Refactor
## Changes
- Single base config
`modelopt_recipes/speculative_decoding/_base_eagle3.yaml` for all EAGLE3
training; removed per-model child YAMLs.
- `launch_train.sh` accepts `--config <yaml>` plus dotlist overrides
(e.g. `model.model_name_or_path=xxx`).
- Removed `__base__` YAML inheritance logic from `main.py`.
- `dp_shard_size` default changed from `0` sentinel to `None` for
clarity.
- Removed `eagle_config.json` and `fsdp_config.json`; architecture
config is now nested under `eagle.eagle_architecture_config` in YAML.
- `train_eagle3_and_export.sh` now uses base YAML + dotlist instead of
generating a temporary YAML.
- Updated README and tests accordingly.
## Usage
```bash
# Online training
./launch_train.sh \
--config ../../modelopt_recipes/speculative_decoding/_base_eagle3.yaml \
model.model_name_or_path=meta-llama/Llama-3.2-1B \
data.data_path=input_conversations/train.jsonl \
training.output_dir=ckpts/llama-3.2-1b-online
# Offline training
./launch_train.sh \
--config ../../modelopt_recipes/speculative_decoding/_base_eagle3.yaml \
model.model_name_or_path=meta-llama/Llama-3.2-1B \
data.offline_data_path=$HIDDEN_STATES_DIR \
training.output_dir=ckpts/llama-3.2-1b-offline
```
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
7f5fd65003 |
[Feat]FakeBaseModel for offline eagle; Kimi-K2.5 fixes; (#1052)
### What does this PR do? Adds `FakeBaseModel` for offline EAGLE training and several Kimi-K2.5 compatibility fixes. - **New**: `FakeBaseModel` — lightweight model that loads only `lm_head` and `embed_tokens` from a local checkpoint, avoiding full model weight loading during offline training. Configured via `FakeBaseArguments` and integrated into `load_vlm_or_llm`. - **Fix**: `_find_base_model_parts` — support Kimi-K2.5 VLM layout (`language_model.model` path) - **Fix**: offline mode lm_head access and CompressedTensors ignore path - **Fix**: Kimi-K2.5 decoder `past_key_value`/`past_key_values` argument mismatch - **Fix**: `rglob` for `.pt` discovery in nested offline data dirs; single-node GPU count respects `CUDA_VISIBLE_DEVICES` Type of change: Bug fix, new feature ### Testing Tested offline EAGLE training for Kimi-K2.5 end-to-end. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ❌ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Lightweight fake-base model support for offline speculative-decoding training * **Improvements** * Added CLI flags: --use_fake_base_for_offline, --trust_remote_code, and --fsdp * Expanded offline .pt discovery to include nested subdirectories * Better GPU detection with explicit single-node logging; FSDP enabled only when requested * Model loading and launch tooling now honor offline and trust-remote-code flags * **Bug Fixes** * Improved compatibility with legacy transformer / Kimi-K2 call signatures * **Tests** * Added tests covering fake-base loading and offline training workflows <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
a34d613d3c |
Feat: Speculatice Decoding export with quantization support (#913)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Main changes: - Refactored speculative decoding export logics into `class EagleExporter` to improve cohesion; - Separated speculative decoding export entrance with quantization export (`export_hf_checkpoint()`) due to their fundamental differences: - Quantization export base model's state_dict and config, while speculative decoding only export drafter's. - Most of the model-specific logics of quantization export (e.g. diffusers, vlms) are not needed for speculative decoding export. - Quantization export produce different format than speculative decoding checkpoint. (The former produce tokenizer config, generation config, e.t.c, while the later does not need. ) ## Usage <!-- You can potentially add a usage example below. --> To export an regular bf16 eagle checkpoint without quantization, the commands are the same: ```python python scripts/export_hf_checkpoint.py --model_path <x> --export_path <x> ``` To run PTQ on online-trained eagle checkpoint and export it: ```python python hf_ptq.py --pyt_ckpt_path <x> --qformat fp8 --export_path <x> ``` The above two commands will produce drafter ckpt for deployment, in the same foramt. ## Testing <!-- Mention how have you tested your change if applicable. --> Tested setting: - Base model: llama3.1-8b - Algorithms: eagle - Export path tested: - (Unquantized online ckpt) `python scripts/export_hf_checkpoint.py --model_path <x> --export_path <x>` - (PTQ) export `python hf_ptq.py --pyt_ckpt_path <x> --qformat fp8 --export_path <x>` - Tested deployment on vllm. Got normal AR. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added export functionality for speculative decoding-optimized models * Support for multiple speculative decoding architectures with pre-configured deployment templates * Enhanced model export detection and automatic routing for optimized models * **Tests** * Updated export validation tests for speculative decoding models <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
75b5da9b83 |
Fix: quant config error on quantized offline eagle (#925)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** ? ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Refactor** * Enhanced quantization configuration handling for transformer models through improved type validation, ensuring more robust processing of quantized model configurations. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
eb99488da1 |
Fix: restore requires_grad in transformers5 reloading (#907)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Patch transformers 5.x parameter loading to preserve original `requires_grad` settings. In transformers v5.x, loading a checkpoint forcibly sets parameters' requires_grad, which unintentionally unfreeze frozen parameters (e.g. Base model in eagle training). This leads to optimizer initialization error since the restored optimizer expected more parameter than the checkpoint. This monkey-patch restores the original`requires_grad` after loading parameters. Reference: https://github.com/huggingface/transformers/blob/v5.0.0.rc1-release/src/transformers/core_model_loading.py#L640 ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed model parameter loading in speculative decoding to properly preserve gradient requirements for each parameter when using HuggingFace Transformers 5.x, ensuring correct behavior during checkpoint resumption and model initialization. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
b8a4586702 |
Refactor: Eagle data loading (#668)
## What does this PR do? **Type of change:** Refactor <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Jira ticket: https://jirasw.nvidia.com/browse/OMNIML-2955 Main changes : - Consolidate Eagle data loading with @ChenhanYu 's implementation of `transformers_dataset.py` - Refactor: baked the following logics from `example/main.py` to `modelopt/torch` for cleaner example entrance: - default config selecting and merging with custom config - tokenizer post-processor (chat template and pad_tok_id) - d2t loading - Implementation refactor: In HF workflow, reuse base modfel's input hidden states as input_embedding, instead of calculating from input_ids. This has two main benefits: - Easier VLM support, which has various embedding processing logics. - Training effieicy. - Deprecating eagle1 from the example. It is still available by setting custom config. - Other minor fixes and readme updates. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> Tested that training curves after changes (both online&offline) is identical with original branch: <img width="1073" height="634" alt="image" src="https://github.com/user-attachments/assets/abfd7bea-c82c-48a7-8181-68c5a9e4da8d" /> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added draft vocabulary cache support for EAGLE model training, enabling runtime vocabulary customization via `--draft_vocab_cache` parameter * Introduced new data loading utilities with sharding, streaming, and tokenization support for large-scale training * Added optional `--log_steps` configuration to training launcher * **Documentation** * Updated EAGLE configuration guides with draft vocabulary cache setup instructions and examples * **Refactor** * Restructured data pipeline for offline training with improved dataset handling and batching * Updated command-line arguments across training scripts (`--input-data` replaces `--input-file`) <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
3036a9ea9f |
Feat: Context Parallel for Eagle3 Training (#745)
## What does this PR do?
**Type of change:** New Feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:**
- Supported Context Parallel by patching torch ring attention;
- Require following libirary version for stable cp:
- torch2.8.0
- transformers5.0.0
- accelrate1.12.0
- Move to FSDP2
- Removed unused arguments in training script (`--multi_gpu`,
`fsdp_wrap_layer`)
- Bump CI container to `nvcr.io/nvidia/pytorch:25.08-py3`
## Usage
<!-- You can potentially add a usage example below. -->
```bash
./launch_train.sh --model $MODEL \
--output_dir $OUTPUT_DIR \
--data $DATA \
--num_epochs 0.1 \
--train_bs 1 \
--eagle_config eagle_config.json \
--training_seq_len 1024 \
--cp_size 2 #newly added
```
## Testing
- SDPA level correctness: tested TTT attention with/without CP, diff <
1%
```
=== Compare context-parallel (CP) outputs and grads with non-CP ===
Forward output comparison (CP vs Non-CP):
Absolute diff (adiff) cp_out vs out: 0.001953125
Relative diff (rdiff) cp_out vs out: 0.00182342529296875
WQ (query proj) grad comparison (CP vs Non-CP):
Absolute diff (adiff) cp_wq_grad vs wq_grad: 0.0078125
Relative diff (rdiff) cp_wq_grad vs wq_grad: 0.00347900390625
WK (key proj) grad comparison (CP vs Non-CP):
Absolute diff (adiff) cp_wk_grad vs wk_grad: 0.0078125
Relative diff (rdiff) cp_wk_grad vs wk_grad: 0.002471923828125
WV (value proj) grad comparison (CP vs Non-CP):
Absolute diff (adiff) cp_wv_grad vs wv_grad: 0.25
Relative diff (rdiff) cp_wv_grad vs wv_grad: 0.0069580078125
==============================================================
```
- E2E Training Acc
(Llama3.1-8B, Unsynthesized magpie)
<img width="911" height="630" alt="image"
src="https://github.com/user-attachments/assets/1ecacc7f-c720-494c-9c1b-b60e7ced7baa"
/>
- Peak Mem Reserved
(llama3.1-8B, 8xH100, train_length=4k)
| cp_size | max_memory_allocated(MB) |max_memory_reserved (MB) |
|----|--------------------------|--------------------------|
| 1 | 65040.20 |79018.00
| 2 | 50409.17 |73098.00
| 4 | 45120.92 |72052.00
| 8 | 38882.12 |66484.00
- Max Training Length test
(llama3.1-8B, H100)
| cp_size | 6k | 12k | 24k | 48k |
|--------------------|-----|-----|-----|-----|
| 1 | ✅ | OOM | OOM | OOM |
|2 | ✅ | ✅ | OOM | OOM |
| 4 | ✅ | ✅ | ✅ | OOM |
| 8 | ✅ | ✅ | ✅ | ✅ |
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added context parallelism (CP) and data parallelism shard size
configuration parameters to training arguments.
* **Enhancements**
* Improved TTT attention masking support for speculative decoding
workflows.
* Enhanced training launch script with improved parallelism
configuration handling.
* **Chores**
* Updated core dependencies: torch, transformers, accelerate, and wandb.
* Added FSDP configuration file for distributed training setup.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
3fd8b804f4 |
Fix:add vllm dump script (#710)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** ? Missing one file in last PR:https://github.com/NVIDIA/Model-Optimizer/pull/689/changes ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
bdd10c2dbe |
Feat: MLA eagle (#689)
## What does this PR do?
**Type of change:** New Feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:**
- Add MLA Eagle support
- Add new argument "eagle_decoder_type" to switch between llama and
kimik2 eagle;
- Add patches to load from kimik2 model implementations dynamically;
- new default config for kimi k2;
- Refactor eagle export to support multilayer/multitype eagle export
concisely;
- Rename some modules for simplified export logic;
- Other minor improvements;
## Usage
<!-- You can potentially add a usage example below. -->
```python
# Add a code snippet demonstrating how to use this
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
- Tested that kimi k2 thinking works with eagle_type=kimik2:
<img width="1068" height="636" alt="image"
src="https://github.com/user-attachments/assets/5557ef87-c719-4fb1-be18-30435f6b3885"
/>
- Tested that llama 3.2 1b works with eagle_type=llama:
<img width="1066" height="634" alt="image"
src="https://github.com/user-attachments/assets/633c575c-cc79-43af-aed3-0378a303ebc7"
/>
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Signed-off-by: yeyu-nvidia <yeyu@nvidia.com>
Co-authored-by: yeyu-nvidia <yeyu@nvidia.com>
|
||
|
|
5a4242faf4 |
Fix: update file path in eagle3 scripts (#654)
## What does this PR do? **Type of change:** Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Update file path ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
bc52b6cf12 |
Feat: Support VLLM one-model eagle ckpt; Add unit tests; (#573)
## What does this PR do? **Type of change:** New feature, new tests; <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** - Add conversion scripts for eagle3 llm-compressor style checkpoint - Jira Ticket: https://jirasw.nvidia.com/browse/OMNIML-2866 - Add unit tests for `ar_validate.py`, `export_hf_checkpoint.py`, and `convert_to_vllm_ckpt.py`. ## Usage <!-- You can potentially add a usage example below. --> ```python python scripts/convert_to_vllm_ckpt.py --input <eagle3 ckpt> --verifier <base model> --output <path> ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
fd79188d49 |
Refactor: eagle3 training loop & loss mask (#548)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** A few refactors to eagle3 training code: - Put first eagle step into TTT loop; - Support TTT mask beyond 4; - Remove answer-only loss mask in data loading. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
1a89a8842c |
Feat: Eagle3 HF Online - support nemotron and VLMs (#463)
## What does this PR do?
**Type of change:** New feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:**
- Support the nano and nano-VL in eagle3 online mode:
- Added submodule path detection for `base model`, `lm_head`, and
`embeddings` to adapt different base model naming structure;
- Refactored data loading/preprocessing to support VLM;
- Attn backend improvement:
- Added option of `sdpa` in case `flex_attn` doesn't work.
- Added a unified TTT mask function that produce either `BlockMask` for
flex_attn or tensor masks for regular attn.
- Logging improvements:
- Added estimated AR validation during training. This is available for
both online and offline.
- Plot estimated AR and training acc to wandb for better training
visualization;
- Fix: PTQ on speculative decoding model.
## Usage
<!-- You can potentially add a usage example below. -->
For VLM as base model, pass in extra arguments `--vlm_processor
<hf_model_path> --vlm_img_dir <path to images>` in original launching
commands. Other usage unchanged.
E.g.
```bash
./launch_train.sh --model $MODEL \
--output_dir $OUTPUT_DIR \
--data $DATA \
--num_gpu 1 \
--num_epochs 2 \
--train_bs 2 \
--lr 3e-5 \
--eagle_config eagle_config.json \
--training_seq_len 4096 \
--vlm_processor $MODEL \
--vlm_img_dir <path to images>
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
Tested short training with HF Online training on following models:
- `llama-3.2-1b` - data: daring-anteater
- The new nano (Hyrbid LLM) - data: daring-anteater
- The nano-VL - data: Llama-Nemotron-VLM-Dataset-v1/ocr_1
See loss decreasing and AR > 1.
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
ff8a1ed126 |
fix:eagle3 offline (#456)
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
40a7d243c3 |
[Eagle Offline] multinode support for hidden states dumper (#422)
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
557633c986 |
Fix: supporting gpt-oss HF eagle (#398)
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
abed33c3f8 |
Feat: TRTLLM Dumper for Eagle Offline Training (#404)
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
c55bcf0ab8 |
fix: eagle3 quantized base model (#383)
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
615f3c01f1 |
Efficient Eagle3 training with eagle KV cache and flex attention (#350)
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
4ff8fc9022 |
Example: add offline eagle training commands to README (#366)
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
26c203abde |
fix: update tests_eagle.py; disable eval for offline (#358)
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
461980e6df |
update eagle example notebook (#314)
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
a6fa34cda4 |
Feat: update eagle3 example; add export (#293)
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |