### What does this PR do? Implement puzzletron compression algorithm based on Puzzle paper (https://arxiv.org/abs/2411.19146) <details> <summary> Th list of reviewed and merged MRs that resulted in the feature/puzzletron branch</summary> Merging dkorzekwa/any_model to feature/puzzletron [Add anymodel directories to feature/puzzletron by danielkorzekwa · Pull Request #974 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/974) - merged [Draft: anymodel activation scoring by danielkorzekwa · Pull Request #989 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/989) - merged [Draft: Merge anymodel pruning by danielkorzekwa · Pull Request #990 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/990/) - merged [Draft: Merging anymodel:build_library_and_stats by danielkorzekwa · Pull Request #993 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/993) - merged [Dkorzekwa/any model calc one block scores by danielkorzekwa · Pull Request #994 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/994) - merged [Draft: merge any_model: mip_and_realize_models by danielkorzekwa · Pull Request #995 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/995) - merged [Dkorzekwa/any model other modeqls by danielkorztiekwa · Pull Request #1007 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1007/) - merged PR to 1007: https://github.com/NVIDIA/Model-Optimizer/pull/1039 - merged [Dkorzekwa/anymodel gptoss by danielkorzekwa · Pull Request #1020 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1020) - merged [Merge any_model tutorial by danielkorzekwa · Pull Request #1035 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1035) - merged [Merge mbridge distillation for any_model by danielkorzekwa · Pull Request #1036 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1036) - merged [MR branch for the remaining difference between dkorzekwa/any_model an… by danielkorzekwa · Pull Request #1047 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1047) - merged [Dkorzekwa/decilm hf code cleanup by danielkorzekwa · Pull Request #1071 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1071) - merged [Dkorzekwa/decilm hf code cleanup 2 by danielkorzekwa · Pull Request #1073 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1073) - merged [Dkorzekwa/anymodel subblock stats by danielkorzekwa · Pull Request #1085 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1085) - merged [Dkorzekwa/anymodel subblock stats nodecilm by danielkorzekwa · Pull Request #1102 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1102) - merged [Dkorzekwa/decilm cleanup post subblockstats by danielkorzekwa · Pull Request #1103 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1103) - merged [code clean up by danielkorzekwa · Pull Request #1110 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1110) - merged Merging into main: [Activation hooks redesign (reuse hooks component across both minitron and puzzletron) by danielkorzekwa · Pull Request #1022 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1022) - merged [Dkorzekwa/puzzletron use importance hooks from prune by danielkorzekwa · Pull Request #1115 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1115) - merged </details> <!-- Details about the change. --> ### Usage Puzzletron tutorial: https://github.com/NVIDIA/Model-Optimizer/tree/feature/puzzletron/examples/puzzletron ### Testing The main e2e test for compressing 9 models with Puzzletron: https://github.com/NVIDIA/Model-Optimizer/blob/feature/puzzletron/tests/gpu/torch/puzzletron/test_puzzletron.py 2-gpu nightly tests: - https://github.com/NVIDIA/Model-Optimizer/actions/runs/24468209205/job/71501061203 - https://github.com/NVIDIA/Model-Optimizer/actions/runs/24470214159/job/71508152952 ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Puzzletron: end-to-end heterogeneous pruning & NAS workflow with AnyModel support, example pipelines, deployment and evaluation utilities, and tools for converting/pruning and exporting compressed checkpoints. * **Documentation** * Comprehensive Puzzletron tutorials, model-specific guides, evaluator instructions, example configs, and changelog entry. * **Chores** * CI/workflow updates (extras installation, longer GPU test timeout), pre-commit hook exclusion updated, and CODEOWNERS entries added. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com> Signed-off-by: Liana Mikaelyan <lmikaelyan@nvidia.com> Signed-off-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com> Signed-off-by: Daniel Korzekwa <daniel.korzekwa@gmail.com> Signed-off-by: jrausch <jrausch@nvidia.com> Signed-off-by: root <root@pool0-00848.cm.cluster> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Liana Mikaelyan <lmikaelyan@nvidia.com> Co-authored-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com> Co-authored-by: J Rausch <38429553+j-rausch@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
8.9 KiB
AnyModel Guide
This guide explains how to add support for new models in the Puzzletron pipeline.
Convert model
Convert a HuggingFace model to Puzzletron format.
Step 1: Create Model Descriptor
Extend ModelDescriptor and implement layer_name_predicates() to define regex patterns for grouping weights into subblocks (embeddings, lm_head, block_N_ffn, block_N_attention).
Key points:
- Find weight names on the model's HuggingFace page → click "Files info" to see the safetensors structure with all tensor names (example: Llama-3.1-8B-Instruct)
See example: llama_model_descriptor.py
Step 2: Create Converter
Extend Converter and implement create_block_configs_from_main_config() to create per-layer BlockConfigs from the HuggingFace config.
Key points:
- Import correct HuggingFace config class (e.g.,
MistralConfig,LlamaConfig,Qwen2Config). Find it in the transformers source:github.com/huggingface/transformers/tree/main/src/transformers/models/<model_type>/configuration_<model_type>.py
See example: llama_converter.py
Step 3: Create models/<model_name>/__init__.py
Export descriptor and converter classes:
from models.<model_name>.<model_name>_model_descriptor import MyModelDescriptor
from models.<model_name>.<model_name>_converter import MyConverter
Step 4: Register in models/__init__.py
Add import to trigger factory registration:
from models.<model_name> import *
Usage
from modelopt.torch.puzzletron.anymodel import convert_model
convert_model(
input_dir="path/to/hf_checkpoint",
output_dir="path/to/puzzletron_checkpoint",
converter="model_name",
)
Compress model
Run pruning and compression on a Puzzletron model.
Step 1: Implement ModelDescriptor methods for compression
Add to your ModelDescriptor:
decoder_layer_cls()- return the decoder layer class(es) to patch for heterogeneous config supportblock_config_to_layer_overrides()- map BlockConfig to layer override dict (see details)init_rotary_embedding()- reinitialize rotary embeddings after model loading (see details)input_embedding_name()- return the name of the input embedding layer (see details)output_embedding_name()- return the name of the output embedding layer (see details)layer_block_name()- return the name pattern for decoder layers (see details)final_norm_name()- return the name of the final normalization layer (see details)attn_no_op_post_init()- replace attention sublayers with no-op modulesmlp_no_op_post_init()- replace MLP sublayers with no-op modules
Step 2: Create FFN Layer Descriptor
Extend FFNIntermediateLayerDescriptor to define model-specific paths for FFN pruning hooks (down_proj_name, ffn_prefix_name, linear_weight_names). Derive values from your model's weight names in layer_name_predicates().
See example: llama_model_descriptor.py → LlamaFFNIntermediateLayerDescriptor
Step 3: Configure YAML files
Update the main model config YAML:
- Set
descriptorto match the name used in@ModelDescriptorFactory.register_decorator("your_model_name") - See example: llama_3_1_8b_instruct.yaml
Update pruning YAML files (ffn_pruning.yaml, expert_pruning.yaml, etc.):
- Set
pruning_mixin._target_to the appropriate mixin class - Set
layer_descriptor._target_to your layer descriptor class - Set
hook_classto the activation hook for scoring - Set
target_layerinactivation_hooks_kwargsto the layer name for hook attachment - See examples in configs/llama_3_1_8b_instruct/pruning/
End-to-end example
See test_puzzletron.py for a complete example that runs both convert and compression steps. For container setup and dependencies needed to run this test, see the Puzzletron README environment section.
Advanced Topics
Pruning Configuration
Pruning YAML Structure
Each pruning type has a YAML config with these key fields:
pruning_mixin:
_target_: pruning.<type>_pruning_mixin.<MixinClass>
layer_descriptor:
_target_: models.<model>.<descriptor_class>
hook_class: ${get_object:utils.activation_hooks.hooks.<HookClass>}
activation_hooks_kwargs:
method: <method_name>
target_layer: "<layer.name>" # e.g., "mlp.down_proj", "self_attn.o_proj"
| Field | Description |
|---|---|
pruning_mixin._target_ |
Mixin class that orchestrates this pruning type |
layer_descriptor._target_ |
Model-specific class defining layer paths for hooks |
hook_class |
Activation hook class for importance scoring |
target_layer |
Layer name (relative to decoder block) where hooks attach |
Adding a New Hook Class
-
Implement the hook under
modelopt/torch/prune/importance_hooks/(e.g.base_hooks.pyfor generic hooks,expert_removal_hooks.pyfor MoE expert removal):- Extend an existing hook base class (e.g.,
RemoveExpertsIndependentHookinexpert_removal_hooks.py) - Implement required methods (e.g.,
get_router_logits_and_routed_experts)
- Extend an existing hook base class (e.g.,
-
Register the hook in the appropriate pruning mixin's
supported_hooks():For FFN pruning (
pruning/ffn_intermediate_pruning_mixin.py):def supported_hooks(self) -> List[Type[ActivationsHook]]: return [IndependentChannelContributionHook, IterativeChannelContributionHook, YourNewHook]For expert removal (
pruning/expert_removal_pruning_mixin.py):def supported_hooks(self) -> List[Type[ActivationsHook]]: return [RankedChoiceVotingHook, ..., YourNewHook] -
Reference in YAML:
hook_class: ${get_object:utils.activation_hooks.hooks.YourNewHook}
Pruning Types Reference
| Type | Mixin | Example Hooks |
|---|---|---|
| FFN intermediate | FFNIntermediatePruningMixIn |
IterativeChannelContributionHook, IndependentChannelContributionHook |
| Expert removal | ExpertRemovalPruningMixIn |
NemotronHRemoveExpertsIndependentHook, Qwen3VLRemoveExpertsIndependentHook |
| KV heads | KVHeadsPruningMixIn |
IndependentKvHeadContributionHook |
Implementing block_config_to_layer_overrides
Maps Puzzletron's BlockConfig fields to HuggingFace config attribute names. Only override attributes that change during pruning:
| BlockConfig Field | HuggingFace Attribute (check config.json) |
|---|---|
attention.num_key_value_heads |
num_key_value_heads |
ffn.intermediate_size |
intermediate_size |
ffn.moe.num_local_experts |
num_experts or n_routed_experts (model-specific) |
ffn.moe.expert_intermediate_dim |
moe_intermediate_size |
Tip: Check the model's config.json for exact attribute names - they vary between models.
See examples: qwen3_vl, nemotron_h
Implementing path-based methods
These methods return paths derived from the model's weight names:
input_embedding_name(),output_embedding_name(),layer_block_name(),final_norm_name()
Find them on the model's HuggingFace page → "Files info" → safetensors structure (example: Llama-3.1-8B-Instruct).
See example: llama_model_descriptor.py
Implementing init_rotary_embedding
Rotary embeddings are computed modules (not saved weights). After model sharding, they need re-initialization on the correct device/dtype.
Look in github.com/huggingface/transformers/tree/main/src/transformers/models/<model_type>/modeling_<model_type>.py for:
class.*Rotary— the rotary embedding class name and constructor argumentsself.rotary_emb— the attribute path
See example: llama_model_descriptor.py