Files
Model-Optimizer/tests
Ajinkya Rasane 17ef5b6c39 Split calibrated ONNX graph capabilities (#2468)
### What does this PR do?

Type of change: new feature

Splits the calibrated ONNX graph utilities into four cohesive capability
owners:

- `graph_indexing.py` owns read-only graph indexing and pattern
matching.
- `graph_selection.py` owns calibrated placement decisions and the
runtime probes needed to make those decisions;
`get_extended_model_outputs` stays with selection for that reason.
- `graph_rewrites.py` owns graph transformations that are independent of
Q/DQ policy.
- `qdq_graph.py` owns graph-level Q/DQ analysis and policy, including
policy-specific rewrites such as `remove_partial_input_qdq`;
`qdq_utils.py` remains the lower-level Q/DQ node, tensor, and format
utility layer.

This four-way split is the intended long-term layout. All production and
test callers now import the owning modules directly, and the former
`graph_utils.py` catch-all module is removed. The extraction preserves
the existing graph-selection and rewrite algorithms; all 43 moved
functions were verified to have identical ASTs before and after the
split.

### Usage

```python
# N/A: behavior-preserving internal refactor.
```

### Testing

- `pytest -q tests/unit/onnx/quantization`
  - 376 passed, 13 intentionally xfailed
- Focused tests after rebasing onto latest `main`
  - 47 passed, 6 intentionally xfailed
- GPU ONNX quantization suite, excluding the existing AutoTune
integration test
  - 46 passed, 3 skipped
- Broader ONNX unit suite
  - 658 passed, 1 skipped, 13 intentionally xfailed
- One ONNX Runtime `CopyTensorAsync is not implemented` failure was
reproduced unchanged on the base revision.
- Review-feedback validation
  - Public `insert_matmul_casts` star-import assertion passed.
- `pytest -q tests/unit/onnx/quantization/test_graph_selection.py
tests/unit/onnx/test_partitioning.py`: 42 passed.
- Changed-file pre-commit hooks passed, including Ruff, formatting,
mypy, Bandit, RST checks, and license checks.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: No -- direct imports from
`modelopt.onnx.quantization.graph_utils` must migrate to the new
capability modules. No shim is retained because removing the catch-all
path is the purpose of this ownership split; the package-level
quantization entry points are unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: Yes
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
Yes -- the backward-breaking section contains a symbol-level migration
map.
- Did you get Claude approval on this PR?: N/A

### Additional Information

Follows #2457, which added focused characterization coverage for the
calibrated ONNX quantization paths.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added expanded ONNX graph analysis for tensor relationships, pattern
detection, FP8 attention matching, and quantization candidate selection.
- Added safer Q/DQ transformations, custom-operator casting,
mixed-precision configuration, and redundant-cast cleanup.
- Improved quantization support for INT4, INT8, FP8, convolution,
attention, and weight-selection scenarios.

- **Refactor**
  - Reorganized ONNX quantization utilities into specialized components.
- Removed the legacy graph utility module; related functionality is now
available through dedicated modules.

- **Tests**
- Updated coverage for graph selection, partitioning, insertion points,
and legacy-module removal.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2026-09-18 10:01:14 -04:00
..