mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
picker.py -- singleton-name collision fix
=========================================
Regression: when ``blocks=`` was set and an unmatched node's name happened to
equal a configured group name, ``_assign_groups``'s
``groups.setdefault(key, []).append(node_name)`` silently merged the unmatched
node into that group instead of keeping it as an isolated singleton. On a
graph like ``scores={"g_n1": 0.01, "g": 10.0}, blocks={"g": [r"^g_"]}``, the
standalone node ``g`` (score 10.0) would land inside the ``g`` group's member
list alongside ``g_n1``, inflating the group's aggregated score and altering
which groups the coverage / threshold picker selected.
Fix: prefix unmatched-node singleton keys with a sentinel (``"\0singleton:"``)
that cannot collide with any user-supplied group name. Group ranking still
receives correct (name, members) pairs; the sentinel is internal to the
``groups`` dict and does not leak into the return value (the picker returns
member node names, not group keys).
Added ``test_unmatched_node_name_collision_with_group_name_stays_isolated``
in ``test_picker.py`` as a regression witness.
utils.py -- host ``validate_file_size``
=======================================
``modelopt.onnx.quantization.sensitivity.__main__`` imported ``validate_file_size``
from ``modelopt.onnx.quantization.__main__``, pulling the main quantize CLI's
full import graph (``autotune.utils``, ``quantize``, ...) into the sensitivity
CLI just to reuse a 15-line size check. If the main CLI later grows a heavier
import, the sensitivity CLI breaks silently.
Moved ``validate_file_size`` to ``modelopt/onnx/utils.py`` alongside the other
low-level ONNX helpers. Both CLIs now import from ``modelopt.onnx.utils``. No
behavior change.
docs -- coverage-mode wording + sample output
=============================================
The ``_onnx_quantization.rst`` guide and ``examples/onnx_ptq/README.md`` still
described coverage mode as "excludes the largest node set whose cumulative
sensitivity score stays at or below ``coverage * total_mass``" -- the same
phrasing the ``suggest_exclusion`` docstring was corrected away from in
an earlier round. Synced to the rank-prefix wording ("walks targets in
descending score order and accumulates them until the next one would push
cumulative sensitivity above ``coverage * total_mass``").
The rendered-ranking example in ``_onnx_quantization.rst`` also could not
be produced by ``_render_ranked_table()``: it mixed ``~0`` string values
(``_render_ranked_table`` uses ``{value:.3f}``, never ``~0``) with a
visible ``Gemm 0 <-- lowest impact`` row (zero-score rows are hidden by
default) and a ``1 target(s) hidden`` footer. Regenerated the sample to
match actual CLI output: drop the ``~0`` rows, move the lowest-impact
marker to the last visible non-zero row (MatMul 0.015), and update the
hidden-count footer to 4.
test_cli.py -- coverage for previously untested CLI helpers
===========================================================
New ``tests/unit/onnx/quantization/sensitivity/test_cli.py`` covers the CLI
functions that landed in earlier rounds but had no unit tests:
* ``_default_output_json``: absolute-path and relative-path derivation.
* ``_validate_calibration_dir``: empty dir raises, under-limit dir passes,
per-file cap trips, aggregate cap trips.
* ``_render_ranked_table``: no scores + no failed, no scores + all failed,
normal ranking with hidden zeros, ``show_zero_scores=True``, all-zero
footer, failed footer.
* ``main()``: synthetic-calibration happy path, default output-JSON path,
oversize ONNX rejected before ``score()`` is called, calibration-directory
path validated before ``score()`` runs.
17 tests total, all pure Python (``score`` monkeypatched, no GPU / no real
ONNX quantization). Wall-clock < 1 s.
test_quantize_api.py -- hoist local imports
===========================================
Moved ``onnx_graphsurgeon as gs`` and ``modelopt.onnx.utils.save_onnx``
from inside ``test_quantize_honors_nodes_to_quantize_allowlist`` to
module scope. Neither is optional or circular.
Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
Co-Authored-By: modelopt-fix-agent-bot (Opus 4.7) <noreply@anthropic.com>