mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Implement puzzletron compression algorithm based on Puzzle paper (https://arxiv.org/abs/2411.19146) <details> <summary> Th list of reviewed and merged MRs that resulted in the feature/puzzletron branch</summary> Merging dkorzekwa/any_model to feature/puzzletron [Add anymodel directories to feature/puzzletron by danielkorzekwa · Pull Request #974 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/974) - merged [Draft: anymodel activation scoring by danielkorzekwa · Pull Request #989 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/989) - merged [Draft: Merge anymodel pruning by danielkorzekwa · Pull Request #990 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/990/) - merged [Draft: Merging anymodel:build_library_and_stats by danielkorzekwa · Pull Request #993 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/993) - merged [Dkorzekwa/any model calc one block scores by danielkorzekwa · Pull Request #994 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/994) - merged [Draft: merge any_model: mip_and_realize_models by danielkorzekwa · Pull Request #995 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/995) - merged [Dkorzekwa/any model other modeqls by danielkorztiekwa · Pull Request #1007 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1007/) - merged PR to 1007: https://github.com/NVIDIA/Model-Optimizer/pull/1039 - merged [Dkorzekwa/anymodel gptoss by danielkorzekwa · Pull Request #1020 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1020) - merged [Merge any_model tutorial by danielkorzekwa · Pull Request #1035 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1035) - merged [Merge mbridge distillation for any_model by danielkorzekwa · Pull Request #1036 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1036) - merged [MR branch for the remaining difference between dkorzekwa/any_model an… by danielkorzekwa · Pull Request #1047 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1047) - merged [Dkorzekwa/decilm hf code cleanup by danielkorzekwa · Pull Request #1071 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1071) - merged [Dkorzekwa/decilm hf code cleanup 2 by danielkorzekwa · Pull Request #1073 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1073) - merged [Dkorzekwa/anymodel subblock stats by danielkorzekwa · Pull Request #1085 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1085) - merged [Dkorzekwa/anymodel subblock stats nodecilm by danielkorzekwa · Pull Request #1102 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1102) - merged [Dkorzekwa/decilm cleanup post subblockstats by danielkorzekwa · Pull Request #1103 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1103) - merged [code clean up by danielkorzekwa · Pull Request #1110 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1110) - merged Merging into main: [Activation hooks redesign (reuse hooks component across both minitron and puzzletron) by danielkorzekwa · Pull Request #1022 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1022) - merged [Dkorzekwa/puzzletron use importance hooks from prune by danielkorzekwa · Pull Request #1115 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1115) - merged </details> <!-- Details about the change. --> ### Usage Puzzletron tutorial: https://github.com/NVIDIA/Model-Optimizer/tree/feature/puzzletron/examples/puzzletron ### Testing The main e2e test for compressing 9 models with Puzzletron: https://github.com/NVIDIA/Model-Optimizer/blob/feature/puzzletron/tests/gpu/torch/puzzletron/test_puzzletron.py 2-gpu nightly tests: - https://github.com/NVIDIA/Model-Optimizer/actions/runs/24468209205/job/71501061203 - https://github.com/NVIDIA/Model-Optimizer/actions/runs/24470214159/job/71508152952 ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Puzzletron: end-to-end heterogeneous pruning & NAS workflow with AnyModel support, example pipelines, deployment and evaluation utilities, and tools for converting/pruning and exporting compressed checkpoints. * **Documentation** * Comprehensive Puzzletron tutorials, model-specific guides, evaluator instructions, example configs, and changelog entry. * **Chores** * CI/workflow updates (extras installation, longer GPU test timeout), pre-commit hook exclusion updated, and CODEOWNERS entries added. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com> Signed-off-by: Liana Mikaelyan <lmikaelyan@nvidia.com> Signed-off-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com> Signed-off-by: Daniel Korzekwa <daniel.korzekwa@gmail.com> Signed-off-by: jrausch <jrausch@nvidia.com> Signed-off-by: root <root@pool0-00848.cm.cluster> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Liana Mikaelyan <lmikaelyan@nvidia.com> Co-authored-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com> Co-authored-by: J Rausch <38429553+j-rausch@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
214 lines
8.0 KiB
Python
214 lines
8.0 KiB
Python
# SPDX-FileCopyrightText: Copyright (c) 2024 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
|
# SPDX-License-Identifier: Apache-2.0
|
|
#
|
|
# Licensed under the Apache License, Version 2.0 (the "License");
|
|
# you may not use this file except in compliance with the License.
|
|
# You may obtain a copy of the License at
|
|
#
|
|
# http://www.apache.org/licenses/LICENSE-2.0
|
|
#
|
|
# Unless required by applicable law or agreed to in writing, software
|
|
# distributed under the License is distributed on an "AS IS" BASIS,
|
|
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
# See the License for the specific language governing permissions and
|
|
# limitations under the License.
|
|
|
|
"""Conf.py file for Sphinx documentation."""
|
|
|
|
# Configuration file for the Sphinx documentation builder.
|
|
#
|
|
# This file only contains a selection of the most common options. For a full
|
|
# list see the documentation:
|
|
# https://www.sphinx-doc.org/en/master/usage/configuration.html
|
|
|
|
# -- Path setup --------------------------------------------------------------
|
|
|
|
# If extensions (or modules to document with autodoc) are in another directory,
|
|
# add these directories to sys.path here. If the directory is relative to the
|
|
# documentation root, use os.path.abspath to make it absolute, like shown here.
|
|
#
|
|
# import os
|
|
# import sys
|
|
# sys.path.insert(0, os.path.abspath('.'))
|
|
|
|
import contextlib
|
|
import os
|
|
import sys
|
|
|
|
import sphinx.application
|
|
from docutils import nodes
|
|
from docutils.nodes import Element
|
|
from sphinx.writers.html5 import HTML5Translator
|
|
|
|
from modelopt import __version__
|
|
|
|
sys.path.insert(0, os.path.abspath("../../"))
|
|
sys.path.append(os.path.abspath("./_ext"))
|
|
|
|
# Pre-import modelopt.torch so it is cached in sys.modules before Sphinx applies
|
|
# autodoc_mock_imports. Mocking triton/tensorrt_llm at the Sphinx level can break
|
|
# transitive imports (transformers, transformer_engine, …) and cause modelopt.torch
|
|
# to fail inside autosummary. Importing here — while the real packages are still on
|
|
# sys.path — avoids that problem entirely.
|
|
with contextlib.suppress(Exception):
|
|
import modelopt.torch # noqa: F401
|
|
|
|
# -- Project information -----------------------------------------------------
|
|
|
|
project = "Model Optimizer" # pylint: disable=C0103
|
|
copyright = "2023-2025, NVIDIA Corporation" # pylint: disable=C0103
|
|
author = "NVIDIA Corporation" # pylint: disable=C0103
|
|
version = __version__
|
|
|
|
# -- General configuration ---------------------------------------------------
|
|
|
|
# Add any Sphinx extension module names here, as strings. They can be
|
|
# extensions coming with Sphinx (named 'sphinx.ext.*') or your custom
|
|
# ones.
|
|
extensions = [
|
|
"sphinx.ext.autodoc",
|
|
"sphinx.ext.autosummary",
|
|
"sphinx.ext.githubpages",
|
|
"sphinx.ext.napoleon", # Support for NumPy and Google style docstrings
|
|
"sphinxarg.ext", # for command-line help documentation
|
|
"sphinx_copybutton", # line numbers getting copied so cannot use `:linenos:`
|
|
"sphinx_inline_tabs",
|
|
"sphinx_togglebutton",
|
|
"sphinxcontrib.autodoc_pydantic",
|
|
"modelopt_autodoc_pydantic",
|
|
]
|
|
|
|
# Only show copybutton for python code-blocks
|
|
copybutton_selector = ", ".join(
|
|
[
|
|
"div.highlight-python pre",
|
|
"div.highlight-ipython3 pre",
|
|
"div.highlight-bash pre",
|
|
"div.highlight-shell pre",
|
|
]
|
|
)
|
|
|
|
# List of patterns, relative to source directory, that match files and
|
|
# directories to ignore when looking for source files.
|
|
# This pattern also affects html_static_path and html_extra_path.
|
|
exclude_patterns = []
|
|
templates_path = ["_templates"]
|
|
|
|
|
|
# -- Options for HTML output -------------------------------------------------
|
|
|
|
# The theme to use for HTML and HTML Help pages. See the documentation for
|
|
# a list of builtin themes.
|
|
#
|
|
html_theme = "sphinx_rtd_theme"
|
|
html_theme_options = {
|
|
"style_external_links": True,
|
|
}
|
|
|
|
# Add any paths that contain custom static files (such as style sheets) here,
|
|
# relative to this directory. They are copied after the builtin static files,
|
|
# so a file named "default.css" will overwrite the builtin "default.css".
|
|
html_static_path = ["_static"]
|
|
|
|
html_title = f"Model Optimizer {version}"
|
|
html_css_files = ["custom.css"]
|
|
html_permalinks_icon = "#" # default icon not rendering properly
|
|
|
|
# TODO: left here as reference for future
|
|
# You can put all these files in the `_static` folder and then activate them as shown below
|
|
# html_favicon = "_static/Nvidia_Symbol.png"
|
|
# html_theme_options = {
|
|
# "light_logo": "Nvidia_Light.png",
|
|
# "dark_logo": "Nvidia_Dark.png",
|
|
# }
|
|
|
|
|
|
# Mock imports for autodoc
|
|
autodoc_mock_imports = ["mpi4py", "tensorrt_llm", "triton"]
|
|
|
|
autosummary_generate = True
|
|
autosummary_imported_members = False
|
|
|
|
# Only consider members from __all__ if available
|
|
autosummary_ignore_module_all = False
|
|
|
|
# Disable docstring inheritance
|
|
autodoc_inherit_docstrings = False
|
|
|
|
# show inheritance
|
|
autodoc_default_options = {"show-inheritance": True}
|
|
|
|
# Omit type package prefixes and type annotations as much as possible to reduce verbosity
|
|
add_module_names = False
|
|
python_use_unqualified_type_names = True
|
|
|
|
# Automatically extract typehints when specified and place them in
|
|
# descriptions of the relevant function/method.
|
|
autodoc_typehints = "description"
|
|
|
|
# Don't show class signature with the class' name.
|
|
autodoc_class_signature = "separated"
|
|
|
|
# Order of autodoc
|
|
# NOTE: summary table on the top of each page does not follow this order so we should also set __all__ in sorted order
|
|
autodoc_member_order = "alphabetical" # can also use `bysource` or `groupwise` to sort members
|
|
|
|
|
|
# autodoc_pydantic model settings
|
|
autodoc_pydantic_model_show_config_summary = False
|
|
autodoc_pydantic_model_show_validator_summary = False
|
|
autodoc_pydantic_model_show_validator_members = False
|
|
autodoc_pydantic_model_show_field_summary = False
|
|
autodoc_pydantic_model_show_json = True # we overwrite this to show the schema or default or both
|
|
autodoc_pydantic_model_signature_prefix = "ModeloptConfig"
|
|
autodoc_pydantic_model_modelopt_show_default_dict = True # show default inside json
|
|
autodoc_pydantic_model_modelopt_show_json_schema = False # hide json schema
|
|
|
|
# autodoc_pydantic field settings
|
|
autodoc_pydantic_field_swap_name_and_alias = True
|
|
autodoc_pydantic_field_doc_policy = "description"
|
|
autodoc_pydantic_field_show_alias = False
|
|
autodoc_pydantic_field_list_validators = False
|
|
autodoc_pydantic_field_show_default = False
|
|
|
|
|
|
class PatchedHTMLTranslator(HTML5Translator):
|
|
"""Open all external links in a new tab. Ref: https://stackoverflow.com/a/61669375 ."""
|
|
|
|
def visit_reference(self, node: Element) -> None:
|
|
"""Visit a reference node."""
|
|
atts = {"class": "reference"}
|
|
if node.get("internal") or "refuri" not in node:
|
|
atts["class"] += " internal"
|
|
else:
|
|
atts["class"] += " external"
|
|
# ---------------------------------------------------------
|
|
# Customize behavior (open in new tab, secure linking site)
|
|
atts["target"] = "_blank"
|
|
atts["rel"] = "noopener noreferrer"
|
|
# ---------------------------------------------------------
|
|
if "refuri" in node:
|
|
atts["href"] = node["refuri"] or "#"
|
|
if self.settings.cloak_email_addresses and atts["href"].startswith("mailto:"):
|
|
atts["href"] = self.cloak_mailto(atts["href"])
|
|
self.in_mailto = True
|
|
else:
|
|
assert "refid" in node, 'References must have "refuri" or "refid" attribute.'
|
|
atts["href"] = "#" + node["refid"]
|
|
if not isinstance(node.parent, nodes.TextElement):
|
|
assert len(node) == 1 and isinstance(node[0], nodes.image)
|
|
atts["class"] += " image-reference"
|
|
if "reftitle" in node:
|
|
atts["title"] = node["reftitle"]
|
|
if "target" in node:
|
|
atts["target"] = node["target"]
|
|
self.body.append(self.starttag(node, "a", "", **atts))
|
|
|
|
if node.get("secnumber"):
|
|
self.body.append(("%s" + self.secnumber_suffix) % ".".join(map(str, node["secnumber"])))
|
|
|
|
|
|
def setup(app: sphinx.application.Sphinx) -> None:
|
|
"""Setup according to the Sphinx extension API."""
|
|
app.set_translator("html", PatchedHTMLTranslator)
|