mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: new feature Splits the calibrated ONNX graph utilities into four cohesive capability owners: - `graph_indexing.py` owns read-only graph indexing and pattern matching. - `graph_selection.py` owns calibrated placement decisions and the runtime probes needed to make those decisions; `get_extended_model_outputs` stays with selection for that reason. - `graph_rewrites.py` owns graph transformations that are independent of Q/DQ policy. - `qdq_graph.py` owns graph-level Q/DQ analysis and policy, including policy-specific rewrites such as `remove_partial_input_qdq`; `qdq_utils.py` remains the lower-level Q/DQ node, tensor, and format utility layer. This four-way split is the intended long-term layout. All production and test callers now import the owning modules directly, and the former `graph_utils.py` catch-all module is removed. The extraction preserves the existing graph-selection and rewrite algorithms; all 43 moved functions were verified to have identical ASTs before and after the split. ### Usage ```python # N/A: behavior-preserving internal refactor. ``` ### Testing - `pytest -q tests/unit/onnx/quantization` - 376 passed, 13 intentionally xfailed - Focused tests after rebasing onto latest `main` - 47 passed, 6 intentionally xfailed - GPU ONNX quantization suite, excluding the existing AutoTune integration test - 46 passed, 3 skipped - Broader ONNX unit suite - 658 passed, 1 skipped, 13 intentionally xfailed - One ONNX Runtime `CopyTensorAsync is not implemented` failure was reproduced unchanged on the base revision. - Review-feedback validation - Public `insert_matmul_casts` star-import assertion passed. - `pytest -q tests/unit/onnx/quantization/test_graph_selection.py tests/unit/onnx/test_partitioning.py`: 42 passed. - Changed-file pre-commit hooks passed, including Ruff, formatting, mypy, Bandit, RST checks, and license checks. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: No -- direct imports from `modelopt.onnx.quantization.graph_utils` must migrate to the new capability modules. No shim is retained because removing the catch-all path is the purpose of this ownership split; the package-level quantization entry points are unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: Yes - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: Yes -- the backward-breaking section contains a symbol-level migration map. - Did you get Claude approval on this PR?: N/A ### Additional Information Follows #2457, which added focused characterization coverage for the calibrated ONNX quantization paths. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added expanded ONNX graph analysis for tensor relationships, pattern detection, FP8 attention matching, and quantization candidate selection. - Added safer Q/DQ transformations, custom-operator casting, mixed-precision configuration, and redundant-cast cleanup. - Improved quantization support for INT4, INT8, FP8, convolution, attention, and weight-selection scenarios. - **Refactor** - Reorganized ONNX quantization utilities into specialized components. - Removed the legacy graph utility module; related functionality is now available through dedicated modules. - **Tests** - Updated coverage for graph selection, partitioning, insertion points, and legacy-module removal. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>