mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Implement puzzletron compression algorithm based on Puzzle paper (https://arxiv.org/abs/2411.19146) <details> <summary> Th list of reviewed and merged MRs that resulted in the feature/puzzletron branch</summary> Merging dkorzekwa/any_model to feature/puzzletron [Add anymodel directories to feature/puzzletron by danielkorzekwa · Pull Request #974 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/974) - merged [Draft: anymodel activation scoring by danielkorzekwa · Pull Request #989 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/989) - merged [Draft: Merge anymodel pruning by danielkorzekwa · Pull Request #990 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/990/) - merged [Draft: Merging anymodel:build_library_and_stats by danielkorzekwa · Pull Request #993 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/993) - merged [Dkorzekwa/any model calc one block scores by danielkorzekwa · Pull Request #994 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/994) - merged [Draft: merge any_model: mip_and_realize_models by danielkorzekwa · Pull Request #995 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/995) - merged [Dkorzekwa/any model other modeqls by danielkorztiekwa · Pull Request #1007 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1007/) - merged PR to 1007: https://github.com/NVIDIA/Model-Optimizer/pull/1039 - merged [Dkorzekwa/anymodel gptoss by danielkorzekwa · Pull Request #1020 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1020) - merged [Merge any_model tutorial by danielkorzekwa · Pull Request #1035 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1035) - merged [Merge mbridge distillation for any_model by danielkorzekwa · Pull Request #1036 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1036) - merged [MR branch for the remaining difference between dkorzekwa/any_model an… by danielkorzekwa · Pull Request #1047 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1047) - merged [Dkorzekwa/decilm hf code cleanup by danielkorzekwa · Pull Request #1071 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1071) - merged [Dkorzekwa/decilm hf code cleanup 2 by danielkorzekwa · Pull Request #1073 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1073) - merged [Dkorzekwa/anymodel subblock stats by danielkorzekwa · Pull Request #1085 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1085) - merged [Dkorzekwa/anymodel subblock stats nodecilm by danielkorzekwa · Pull Request #1102 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1102) - merged [Dkorzekwa/decilm cleanup post subblockstats by danielkorzekwa · Pull Request #1103 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1103) - merged [code clean up by danielkorzekwa · Pull Request #1110 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1110) - merged Merging into main: [Activation hooks redesign (reuse hooks component across both minitron and puzzletron) by danielkorzekwa · Pull Request #1022 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1022) - merged [Dkorzekwa/puzzletron use importance hooks from prune by danielkorzekwa · Pull Request #1115 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1115) - merged </details> <!-- Details about the change. --> ### Usage Puzzletron tutorial: https://github.com/NVIDIA/Model-Optimizer/tree/feature/puzzletron/examples/puzzletron ### Testing The main e2e test for compressing 9 models with Puzzletron: https://github.com/NVIDIA/Model-Optimizer/blob/feature/puzzletron/tests/gpu/torch/puzzletron/test_puzzletron.py 2-gpu nightly tests: - https://github.com/NVIDIA/Model-Optimizer/actions/runs/24468209205/job/71501061203 - https://github.com/NVIDIA/Model-Optimizer/actions/runs/24470214159/job/71508152952 ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Puzzletron: end-to-end heterogeneous pruning & NAS workflow with AnyModel support, example pipelines, deployment and evaluation utilities, and tools for converting/pruning and exporting compressed checkpoints. * **Documentation** * Comprehensive Puzzletron tutorials, model-specific guides, evaluator instructions, example configs, and changelog entry. * **Chores** * CI/workflow updates (extras installation, longer GPU test timeout), pre-commit hook exclusion updated, and CODEOWNERS entries added. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com> Signed-off-by: Liana Mikaelyan <lmikaelyan@nvidia.com> Signed-off-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com> Signed-off-by: Daniel Korzekwa <daniel.korzekwa@gmail.com> Signed-off-by: jrausch <jrausch@nvidia.com> Signed-off-by: root <root@pool0-00848.cm.cluster> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Liana Mikaelyan <lmikaelyan@nvidia.com> Co-authored-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com> Co-authored-by: J Rausch <38429553+j-rausch@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
63 lines
3.2 KiB
Plaintext
63 lines
3.2 KiB
Plaintext
# GitHub Teams defined at https://github.com/orgs/NVIDIA/teams/modelopt-devs/teams
|
|
|
|
# Configuration files
|
|
.github @NVIDIA/modelopt-setup-codeowners
|
|
.gitlab @NVIDIA/modelopt-setup-codeowners
|
|
.pre-commit-config.yaml @NVIDIA/modelopt-setup-codeowners
|
|
CONTRIBUTING.md @NVIDIA/modelopt-setup-codeowners
|
|
LICENSE @NVIDIA/modelopt-setup-codeowners
|
|
LICENSE_HEADER @NVIDIA/modelopt-setup-codeowners
|
|
pyproject.toml @NVIDIA/modelopt-setup-codeowners
|
|
SECURITY.md @NVIDIA/modelopt-setup-codeowners
|
|
tox.ini @NVIDIA/modelopt-setup-codeowners
|
|
uv.lock @NVIDIA/modelopt-setup-codeowners
|
|
|
|
# Library
|
|
modelopt/deploy @NVIDIA/modelopt-deploy-codeowners
|
|
modelopt/onnx @NVIDIA/modelopt-onnx-codeowners
|
|
modelopt/onnx/autocast @NVIDIA/modelopt-onnx-autocast-codeowners
|
|
modelopt/torch @NVIDIA/modelopt-torch-codeowners
|
|
modelopt/torch/_deploy @NVIDIA/modelopt-torch-deploy-codeowners
|
|
modelopt/torch/distill @NVIDIA/modelopt-torch-distill-codeowners
|
|
modelopt/torch/export @NVIDIA/modelopt-torch-export-codeowners
|
|
modelopt/torch/nas @NVIDIA/modelopt-torch-nas-prune-codeowners
|
|
modelopt/torch/opt @NVIDIA/modelopt-torch-opt-codeowners
|
|
modelopt/torch/peft @NVIDIA/modelopt-torch-peft-codeowners
|
|
modelopt/torch/prune @NVIDIA/modelopt-torch-nas-prune-codeowners
|
|
modelopt/torch/puzzletron @NVIDIA/modelopt-torch-puzzletron-codeowners
|
|
modelopt/torch/quantization @NVIDIA/modelopt-torch-quantization-codeowners
|
|
modelopt/torch/sparsity @NVIDIA/modelopt-torch-sparsity-codeowners
|
|
modelopt/torch/speculative @NVIDIA/modelopt-torch-speculative-codeowners
|
|
modelopt/torch/trace @NVIDIA/modelopt-torch-nas-prune-codeowners
|
|
modelopt/torch/utils @NVIDIA/modelopt-torch-utils-codeowners
|
|
modelopt_recipes @NVIDIA/modelopt-recipes-codeowners
|
|
|
|
# Examples
|
|
/README.md @NVIDIA/modelopt-examples-codeowners
|
|
/examples @NVIDIA/modelopt-examples-codeowners
|
|
/examples/chained_optimizations @NVIDIA/modelopt-torch-nas-prune-codeowners
|
|
/examples/cnn_qat @NVIDIA/modelopt-examples-cnn_qat-codeowners
|
|
/examples/deepseek @NVIDIA/modelopt-deploy-codeowners
|
|
/examples/diffusers @NVIDIA/modelopt-examples-diffusers-codeowners
|
|
/examples/gpt-oss @NVIDIA/modelopt-examples-gpt-oss-codeowners
|
|
/examples/llm_autodeploy @NVIDIA/modelopt-deploy-codeowners
|
|
/examples/llm_distill @NVIDIA/modelopt-torch-distill-codeowners
|
|
/examples/llm_eval @NVIDIA/modelopt-examples-llm_ptq-codeowners
|
|
/examples/llm_ptq @NVIDIA/modelopt-examples-llm_ptq-codeowners
|
|
/examples/llm_qat @NVIDIA/modelopt-examples-llm_qat-codeowners
|
|
/examples/llm_sparsity @NVIDIA/modelopt-torch-sparsity-codeowners
|
|
/examples/megatron_bridge @NVIDIA/modelopt-examples-megatron-codeowners
|
|
/examples/model_hub @NVIDIA/modelopt-examples-model_hub-codeowners
|
|
/examples/onnx_ptq @NVIDIA/modelopt-onnx-codeowners
|
|
/examples/pruning @NVIDIA/modelopt-torch-nas-prune-codeowners
|
|
/examples/puzzletron @NVIDIA/modelopt-torch-puzzletron-codeowners
|
|
/examples/specdec_bench @NVIDIA/modelopt-torch-speculative-codeowners
|
|
/examples/speculative_decoding @NVIDIA/modelopt-torch-speculative-codeowners
|
|
/examples/torch_onnx @NVIDIA/modelopt-onnx-codeowners
|
|
/examples/vlm_ptq @NVIDIA/modelopt-examples-vlm-codeowners
|
|
/examples/vllm_serve @NVIDIA/modelopt-examples-llm_ptq-codeowners
|
|
/examples/windows @NVIDIA/modelopt-windows-codeowners
|
|
|
|
# Requirements files are owned by the setup team regardless of location
|
|
requirements*.txt @NVIDIA/modelopt-setup-codeowners
|