Commit Graph
22 Commits
Author SHA1 Message Date
Keval Morabia 53a2ddebab Product Rename: TensorRT Model Optimizer to Model Optimizer (#583)
- [x] Product Rename: TensorRT Model Optimizer to Model Optimizer
(OMNIML-3033)
- [x] Mention in Latest News section with date on the date of merging
this PR (12/08)

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-12-07 12:33:21 +05:30
Keval Morabia be64f6b1d5 Deprecate examples/megatron-lm in favor of M-LM repo examples folder (#567)
- Deprecate `examples/megatron-lm` and move missing readme sections to
https://github.com/NVIDIA/Megatron-LM/tree/main/examples/post_training/modelopt
to keep only one copy of this doc.
- Linked PR: https://github.com/NVIDIA/Megatron-LM/pull/2273

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-11-17 23:41:45 +05:30
Keval Morabia 6a8d6dac26 Enable Yarn RoPE in minitron pruning for gpt-oss support (#530)
## What does this PR do?

**Type of change:** Improve existing feature <!-- Use one of the
following: Bug fix, new feature, new example, new tests, documentation.
-->

**Overview:** GPT-OSS model has Yarn RoPE which adds additional
nn.Embedding modules that need to be enabled in DynamicModule for
Minitron pruning

## Testing
<!-- Mention how have you tested your change if applicable. -->

- gpt-oss-20b pruned using M-LM pruning example and conf scripts.

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-11-11 10:00:23 +05:30
Ali TaghibakhshiandKeval Morabia 41c66bb6d2 Add MoE (e.g. Qwen3-30B-A3B, Mamba hybrid) pruning support in Minitron (#467)
## What does this PR do?

**Type of change:** New feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

- Support pruning `num_moe_experts`, `moe_ffn_hidden_size`, and
`moe_shared_expert_intermediate_size` in `mcore_minitron` pruning

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added Mixture of Experts (MoE) pruning support with new configurable
dimensions for expert count and intermediate sizes
* Extended NAS architecture search capabilities to include MoE model
parameters

* **Documentation**
* Updated support matrix and pruning documentation for MoE-compatible
models
* Clarified available pruning dimensions and parameters for MoE
architectures

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-11-08 00:47:23 +05:30
Keval Morabia 90b1e68fcf Update HF nvidia collection links in docs (#475)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-10-28 22:23:01 +05:30
Keval Morabia 6dffcd0ff1 Add minitron pruning and distillation guidelines in pruning readme (#419)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-10-10 10:06:25 +00:00
Keval Morabia cb44c556d8 Minor pruning update (#397)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-10-02 00:35:48 +05:30
Mohammed Yasin 70abfb416f Fix FLOPs calculation (#388)
Signed-off-by: Mohammed Yasin <32206511+Y-T-G@users.noreply.github.com>
Signed-off-by: Y-T-G <32206511+Y-T-G@users.noreply.github.com>
2025-09-30 21:40:51 +05:30
Keval Morabia b4d6cedb4a Support checkpointing Minitron pruning scores / prune without sorting (#361)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-09-26 01:46:22 +05:30
Keval Morabia c0590b0255 Deprecate ModelOpt custom docker and directly use TRT-LLM / PyTorch / TRT docker (#346)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-09-23 01:43:25 +05:30
Keval Morabia c60baae5fa Add Megatron-LM pruning example link (#344)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-09-19 23:36:42 +05:30
Keval Morabia 4c36abe536 Add llm_ptq PR test, Cleanup dependency installation in modelopt docker (#338)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-09-19 10:48:44 +05:30
Keval Morabia 103b1bb137 Push latest changes
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-09-04 11:02:58 +05:30
Keval Morabia 4d1eb0caf5 Major improvement of READMEs and documentation
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-08-30 10:54:28 +05:30
Keval Morabia dcdc2df084 Push latest changes
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-08-26 01:12:26 +05:30
Keval Morabia cafa7f60ed Update for 0.33.0 release 2025-07-14 23:08:26 +05:30
Keval Morabia 7af33d29ce Update for 0.31.0 release 2025-06-05 13:24:07 -07:00
Keval Morabia 3039f76d6a Fix cifar_resnet.ipynb not rendering in github 2025-05-20 18:37:00 +05:30
Keval Morabia 7a047435ae Update for 0.29.0 release 2025-05-08 23:43:51 +05:30
Keval Morabia 2017cd9063 Update for 0.25.0 release 2025-03-03 22:54:22 +05:30
Keval Morabia 5c9390c215 0.23.1 Release - fix for torch 2.6 2025-02-14 16:15:53 +05:30
Keval Morabia 73d6af785f Update 0.23.0 - OSS release 2025-01-29 01:45:47 +05:30