Files
Model-Optimizer/examples/speculative_decoding/recipes
h-guo18 7f5fd65003 [Feat]FakeBaseModel for offline eagle; Kimi-K2.5 fixes; (#1052)
### What does this PR do?

Adds `FakeBaseModel` for offline EAGLE training and several Kimi-K2.5
compatibility fixes.

- **New**: `FakeBaseModel` — lightweight model that loads only `lm_head`
and `embed_tokens` from a local checkpoint, avoiding full model weight
loading during offline training. Configured via `FakeBaseArguments` and
integrated into `load_vlm_or_llm`.
- **Fix**: `_find_base_model_parts` — support Kimi-K2.5 VLM layout
(`language_model.model` path)
- **Fix**: offline mode lm_head access and CompressedTensors ignore path
- **Fix**: Kimi-K2.5 decoder `past_key_value`/`past_key_values` argument
mismatch
- **Fix**: `rglob` for `.pt` discovery in nested offline data dirs;
single-node GPU count respects `CUDA_VISIBLE_DEVICES`

Type of change: Bug fix, new feature

### Testing
Tested offline EAGLE training for Kimi-K2.5 end-to-end.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌

### Additional Information
<!-- E.g. related issue. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Lightweight fake-base model support for offline speculative-decoding
training

* **Improvements**
* Added CLI flags: --use_fake_base_for_offline, --trust_remote_code, and
--fsdp
  * Expanded offline .pt discovery to include nested subdirectories
* Better GPU detection with explicit single-node logging; FSDP enabled
only when requested
* Model loading and launch tooling now honor offline and
trust-remote-code flags

* **Bug Fixes**
* Improved compatibility with legacy transformer / Kimi-K2 call
signatures

* **Tests**
* Added tests covering fake-base loading and offline training workflows
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-03-27 16:57:13 -07:00
..