mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Adds `FakeBaseModel` for offline EAGLE training and several Kimi-K2.5 compatibility fixes. - **New**: `FakeBaseModel` — lightweight model that loads only `lm_head` and `embed_tokens` from a local checkpoint, avoiding full model weight loading during offline training. Configured via `FakeBaseArguments` and integrated into `load_vlm_or_llm`. - **Fix**: `_find_base_model_parts` — support Kimi-K2.5 VLM layout (`language_model.model` path) - **Fix**: offline mode lm_head access and CompressedTensors ignore path - **Fix**: Kimi-K2.5 decoder `past_key_value`/`past_key_values` argument mismatch - **Fix**: `rglob` for `.pt` discovery in nested offline data dirs; single-node GPU count respects `CUDA_VISIBLE_DEVICES` Type of change: Bug fix, new feature ### Testing Tested offline EAGLE training for Kimi-K2.5 end-to-end. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ❌ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Lightweight fake-base model support for offline speculative-decoding training * **Improvements** * Added CLI flags: --use_fake_base_for_offline, --trust_remote_code, and --fsdp * Expanded offline .pt discovery to include nested subdirectories * Better GPU detection with explicit single-node logging; FSDP enabled only when requested * Model loading and launch tooling now honor offline and trust-remote-code flags * **Bug Fixes** * Improved compatibility with legacy transformer / Kimi-K2 call signatures * **Tests** * Added tests covering fake-base loading and offline training workflows <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>