mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Deprecate gradnas pruning and bert example (#1427)
### What does this PR do? Type of change: Deprecation of dead code <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Deprecation warning already added in 0.44 as per 1-release deprecation policy GradNAS only works for Bert and GPT-J and we dont actively maintain it or test it. Keeping it creates an expectation that it works plus it adds one more option for user to choose from. We already have much better pruning algorithms (Minitron and Puzzletron) for LLM pruning already hence removing GradNas. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ No but we dont have any users of this feature either <!--- If ❌, explain why. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Deprecation** * GradNAS pruning algorithm deprecated; related examples removed. * **Documentation** * Pruning and NAS guides and changelog updated to focus on Minitron and FastNAS; GradNAS references removed. * **Chores** * Chained-optimizations example and scripts removed. * Ownership mappings updated for README and examples; license insertion now applies to a previously excluded example file. * **Tests** * Multiple unit tests and test utilities related to GradNAS/transformer NAS removed. [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1427) <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
This commit is contained in:
co-authored by
claude[bot]
parent
be85c7fc52
commit
2ce745a92e
@@ -8,7 +8,7 @@ Pruning
|
||||
`ResNet20 on CIFAR-10 Notebook <https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/pruning/cifar_resnet.ipynb>`_
|
||||
for an end-to-end example of pruning.
|
||||
|
||||
ModelOpt provides three main pruning methods (aka ``mode``) - Minitron, FastNAS and GradNAS - via a unified API
|
||||
ModelOpt provides three main pruning methods (aka ``mode``) - Minitron, Puzzletron, and FastNAS - via a unified API
|
||||
:meth:`mtp.prune <modelopt.torch.prune.pruning.prune>`. Given a model,
|
||||
these methods finds the subnet which meets the given deployment constraints (e.g. FLOPs, parameters)
|
||||
from your provided base model with little to no accuracy degradation (depending on how aggressive is the pruning).
|
||||
@@ -20,10 +20,9 @@ attention heads of the model. More details on these pruning modes is as follows:
|
||||
the embedding hidden size, mlp ffn hidden size, transformer attention heads, GQA query groups,
|
||||
mamba heads and head dimension, and number of layers of the model.
|
||||
Checkout more details of the algorithm in the `paper <https://arxiv.org/abs/2408.11796>`_.
|
||||
#. ``puzzletron``: An advanced LLM/VLM pruning method by NVIDIA using Mixed Integer Programming (MIP) based NAS search algorithm.
|
||||
#. ``fastnas``: A pruning method recommended for Computer Vision models. Given a pretrained model,
|
||||
FastNAS finds the subnet which maximizes the score function while meeting the given constraints.
|
||||
#. ``gradnas``: A light-weight pruning method recommended for language models like Hugging Face BERT and GPT-J.
|
||||
It uses the gradient information to prune the model's linear layers and attention heads to meet the given constraints.
|
||||
|
||||
Follow the steps described below to obtain the optimal model satisfying your
|
||||
requirements using :mod:`mtp<modelopt.torch.prune>`:
|
||||
@@ -54,10 +53,9 @@ Prerequisites
|
||||
#. You can provide one search constraint for either ``flops`` or ``params`` by
|
||||
specifying an upper bound in terms of absolute number (``3e-6``) or a percentage (``"60%"``).
|
||||
#. You should also specify the pruning algorithm (``mode``), you would like to use. Depending on the
|
||||
mode, you will need to provide additional ``config`` parameters like ``score_func`` (``fastnas`` mode)
|
||||
or ``loss_func`` (``gradnas`` mode), ``dataloader``, ``checkpoint``, etc. The most common score function
|
||||
mode, you will need to provide additional ``config`` parameters like ``score_func`` (``fastnas`` mode),
|
||||
``data_loader``, ``checkpoint``, etc. The most common score function
|
||||
is the validation accuracy of the model and is used to rank the sub-nets sampled from the search space.
|
||||
Loss function is used to run some forward and backward passes on the train dataloader to get the gradients.
|
||||
#. Please see the API reference of :meth:`mtp.prune() <modelopt.torch.prune.pruning.prune>` for more details.
|
||||
|
||||
Below we show an example using :class:`"fastnas" <modelopt.torch.prune.fastnas.FastNASModeDescriptor>`.
|
||||
@@ -128,9 +126,7 @@ possible network configurations and an optimal configuration is then searched fo
|
||||
via ``DistributedDataParallel`` in PyTorch.
|
||||
|
||||
Currently, the API does not support pruning pytorch Fully Sharded Data Parallel (FSDP) models
|
||||
so you would need to run pruning on a CPU and then finetune using FSDP. Note that GradNAS is
|
||||
much much faster than FastNAS (hence feasible on CPU as well) and is recommended for
|
||||
language models like BERT and GPT-J 6B.
|
||||
so you would need to run pruning on a CPU and then finetune using FSDP.
|
||||
|
||||
|
||||
Storing the pruned model
|
||||
|
||||
@@ -63,7 +63,7 @@ Example usage:
|
||||
|
||||
.. note::
|
||||
|
||||
The NAS API's are a super-set of the pruning API's. You can use the pruning modes (e.g. ``"fastnas"``, ``"gradnas"``, etc.)
|
||||
The NAS API's are a super-set of the pruning API's. You can use the pruning modes (e.g. ``"fastnas"``)
|
||||
here as well.
|
||||
|
||||
.. note::
|
||||
@@ -370,12 +370,6 @@ can be converted into searchable units:
|
||||
megatron.core.models.gpt.GPTModel
|
||||
megatron.core.models.mamba.MambaModel
|
||||
|
||||
# We convert Hugging Face Attention layers to automatically search over the number of heads
|
||||
# and MLP hidden size.
|
||||
# Make sure `config.use_cache` is set to False during pruning.
|
||||
transformers.models.bert.modeling_bert.BertAttention
|
||||
transformers.models.gptj.modeling_gptj.GPTJAttention
|
||||
|
||||
Generating a search space
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
|
||||
Reference in New Issue
Block a user