Files
Model-Optimizer/examples/pruning/minitron
Daniel Korzekwaandcoderabbitai[bot] 4d19ca1a08 Create a tool for data blend preparation to enable fast experimentation with distillation (#1888)
### What does this PR do?

Create a tool for data blend preparation to enable fast experimentation
with distillation

### Usage

- examples/researcher_guide/README.md (## Prepare token-budgeted data
blends)
- examples/dataset/prepare_data_blend.py

### Testing
- tests/examples/dataset/test_prepare_data_blend.py
-
tests/gpu_megatron/torch/utils/plugins/test_megatron_preprocess_data.py

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ 
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added `max_tokens` support for preprocessing to token-cap outputs with
safe early stopping and distinct capped-run artifacts (including a
matching CLI flag).
* Added a YAML-driven workflow to generate weighted Megatron data blends
with per-source token allocation and reproducible output metadata.

* **Documentation**
* Added a researcher fast-experimentation guide for iterative,
token-limited evaluation and token-budgeted distillation blends.
* Updated dataset-prep and tutorial instructions to recommend
`target_tokens`/token-budgeted subsets.

* **Tests**
* Added CPU/GPU tests covering `max_tokens` stopping behavior, HF
streaming caching, and the data-blend YAML workflow.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-07-13 23:03:52 +02:00
..

Minitron Pruning — End-to-End Tutorials

End-to-end tutorials for Minitron structured pruning followed by knowledge distillation, quantization, evaluation,and vLLM deployment.

Each subdirectory covers a specific source model and target size, including the full data blend, pruning config, distillation hyperparameters, evaluation results, and throughput benchmarks.