Deprecate gradnas pruning and bert example (#1427)

### What does this PR do?

Type of change: Deprecation of dead code <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->

Deprecation warning already added in 0.44 as per 1-release deprecation
policy

GradNAS only works for Bert and GPT-J and we dont actively maintain it
or test it. Keeping it creates an expectation that it works plus it adds
one more option for user to choose from. We already have much better
pruning algorithms (Minitron and Puzzletron) for LLM pruning already
hence removing GradNas.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ No but we dont have any users
of this feature either <!--- If ❌, explain why. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Deprecation**
  * GradNAS pruning algorithm deprecated; related examples removed.
* **Documentation**
* Pruning and NAS guides and changelog updated to focus on Minitron and
FastNAS; GradNAS references removed.
* **Chores**
  * Chained-optimizations example and scripts removed.
* Ownership mappings updated for README and examples; license insertion
now applies to a previously excluded example file.
* **Tests**
* Multiple unit tests and test utilities related to GradNAS/transformer
NAS removed.

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1427)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
This commit is contained in:
Keval Morabia
2026-05-12 22:32:22 +05:30
committed by GitHub
co-authored by claude[bot]
parent be85c7fc52
commit 2ce745a92e
36 changed files with 21 additions and 2652 deletions
+5 -9
View File
@@ -8,7 +8,7 @@ Pruning
`ResNet20 on CIFAR-10 Notebook <https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/pruning/cifar_resnet.ipynb>`_
for an end-to-end example of pruning.
ModelOpt provides three main pruning methods (aka ``mode``) - Minitron, FastNAS and GradNAS - via a unified API
ModelOpt provides three main pruning methods (aka ``mode``) - Minitron, Puzzletron, and FastNAS - via a unified API
:meth:`mtp.prune <modelopt.torch.prune.pruning.prune>`. Given a model,
these methods finds the subnet which meets the given deployment constraints (e.g. FLOPs, parameters)
from your provided base model with little to no accuracy degradation (depending on how aggressive is the pruning).
@@ -20,10 +20,9 @@ attention heads of the model. More details on these pruning modes is as follows:
the embedding hidden size, mlp ffn hidden size, transformer attention heads, GQA query groups,
mamba heads and head dimension, and number of layers of the model.
Checkout more details of the algorithm in the `paper <https://arxiv.org/abs/2408.11796>`_.
#. ``puzzletron``: An advanced LLM/VLM pruning method by NVIDIA using Mixed Integer Programming (MIP) based NAS search algorithm.
#. ``fastnas``: A pruning method recommended for Computer Vision models. Given a pretrained model,
FastNAS finds the subnet which maximizes the score function while meeting the given constraints.
#. ``gradnas``: A light-weight pruning method recommended for language models like Hugging Face BERT and GPT-J.
It uses the gradient information to prune the model's linear layers and attention heads to meet the given constraints.
Follow the steps described below to obtain the optimal model satisfying your
requirements using :mod:`mtp<modelopt.torch.prune>`:
@@ -54,10 +53,9 @@ Prerequisites
#. You can provide one search constraint for either ``flops`` or ``params`` by
specifying an upper bound in terms of absolute number (``3e-6``) or a percentage (``"60%"``).
#. You should also specify the pruning algorithm (``mode``), you would like to use. Depending on the
mode, you will need to provide additional ``config`` parameters like ``score_func`` (``fastnas`` mode)
or ``loss_func`` (``gradnas`` mode), ``dataloader``, ``checkpoint``, etc. The most common score function
mode, you will need to provide additional ``config`` parameters like ``score_func`` (``fastnas`` mode),
``data_loader``, ``checkpoint``, etc. The most common score function
is the validation accuracy of the model and is used to rank the sub-nets sampled from the search space.
Loss function is used to run some forward and backward passes on the train dataloader to get the gradients.
#. Please see the API reference of :meth:`mtp.prune() <modelopt.torch.prune.pruning.prune>` for more details.
Below we show an example using :class:`"fastnas" <modelopt.torch.prune.fastnas.FastNASModeDescriptor>`.
@@ -128,9 +126,7 @@ possible network configurations and an optimal configuration is then searched fo
via ``DistributedDataParallel`` in PyTorch.
Currently, the API does not support pruning pytorch Fully Sharded Data Parallel (FSDP) models
so you would need to run pruning on a CPU and then finetune using FSDP. Note that GradNAS is
much much faster than FastNAS (hence feasible on CPU as well) and is recommended for
language models like BERT and GPT-J 6B.
so you would need to run pruning on a CPU and then finetune using FSDP.
Storing the pruned model
+1 -7
View File
@@ -63,7 +63,7 @@ Example usage:
.. note::
The NAS API's are a super-set of the pruning API's. You can use the pruning modes (e.g. ``"fastnas"``, ``"gradnas"``, etc.)
The NAS API's are a super-set of the pruning API's. You can use the pruning modes (e.g. ``"fastnas"``)
here as well.
.. note::
@@ -370,12 +370,6 @@ can be converted into searchable units:
megatron.core.models.gpt.GPTModel
megatron.core.models.mamba.MambaModel
# We convert Hugging Face Attention layers to automatically search over the number of heads
# and MLP hidden size.
# Make sure `config.use_cache` is set to False during pruning.
transformers.models.bert.modeling_bert.BertAttention
transformers.models.gptj.modeling_gptj.GPTJAttention
Generating a search space
^^^^^^^^^^^^^^^^^^^^^^^^^