feat(quantization): fail fast when a quant config matches no weight quantizer (#2203)

### What does this PR do?

Type of change: New feature (fail-fast guard; behavior change on a
previously silent path)

A `quant_cfg` whose module patterns don't match the model is not an
error to `set_quantizer_by_cfg` — every pattern simply matches nothing.
The run then calibrates, exports, and hands back a checkpoint that is
silently unquantized:

```json
{"quantization": {"quant_algo": null, "kv_cache_quant_algo": "FP8", "quantized_layers": {}}}
```

Nothing in the run says so. It has only ever been caught by someone
reading the exported `hf_quant_config.json` afterwards — most recently
on Step-3.7 ([NVBug 6518665](https://nvbugspro.nvidia.com/bug/6518665),
after a full 8×B200 calibration), and before that on MiniMax-M3, where
fused-expert detection skipped the experts and an experts-only recipe
matched nothing.

`mtq.quantize` now compares the config's intent against the outcome and
raises **before calibration**:

```
RuntimeError: The quantization config asks for weight quantization but no weight quantizer was
enabled, so nothing would be quantized (3 quantizer(s) inserted). These patterns matched no
weight quantizer:
  *.experts.*weight_quantizer
Either the patterns do not match this architecture's module names (check the model-specific
recipes under modelopt_recipes/huggingface/<model_type>/), or the modules holding the weights
were never converted to quantized modules (an unsupported custom module, e.g. a
trust_remote_code MoE layout).
```

Scoped to avoid false positives:

- **Only configs that ask for weight quantization** are checked (an
entry with `enable` and `weight_quantizer` in its pattern), so
activation-only and KV-cache-only configs are unaffected.
- **Intent is read from each pattern's final entry**, since `quant_cfg`
entries apply in order: a pattern that is enabled and then disabled
later asks for nothing by the end.
- **Configs refining an already-quantized model** (weight quantizers
enabled by an earlier `mtq.quantize`) are left alone.

Matching goes through `conversion._match_quantizer` — the same matcher
`set_quantizer_by_cfg` used to apply the config — so "did this pattern
match anything?" is answered exactly as the applying code would. A local
`fnmatch` diverges on the two cases that matcher handles:
`SequentialQuantizer` modules (W4A8-style list-valued `cfg`) and
fused-experts names (`..._weight_quantizers.0` normalizing to
`..._weight_quantizer`).

### Usage

No API change. A config that would previously have produced an
unquantized checkpoint now raises:

```python
mtq.quantize(model, {"quant_cfg": [
    {"quantizer_name": "*", "enable": False},
    {"quantizer_name": "*.experts.*weight_quantizer", "cfg": {"num_bits": 8, "axis": 0}},
]}, forward_loop)   # RuntimeError if the model has no `experts` modules
```

### Testing

Seven tests in `tests/unit/torch/quantization/test_quantize_cpu.py`, one
per branch of the guard: patterns matching nothing raise; an
activation-only config still runs; weight patterns disabled by a later
entry still run; enabled-then-retracted patterns still run;
`SequentialQuantizer` (list-valued `cfg`) and fused-experts quantizer
names count as matched; and the already-quantized refinement path is
exercised. Each was checked to be non-vacuous by removing the
corresponding branch and confirming exactly that test fails.

**One existing test changed.**
`tests/gpu/torch/export/test_fsdp2_export.py` parametrized over
`NVFP4_MLP_ONLY_CFG`, but its `SmallQKVModel` has no MLP — so that case
ran the FSDP2 paths against an *unquantized* model, and the new guard
reported it (4 GPU failures on the first CI run, all `quant_config6`;
`NVFP4_OMLP_ONLY_CFG` passed because that model does have `o_proj`). The
parametrization is dropped with a comment; `NVFP4_OMLP_ONLY_CFG` keeps
the scoped-recipe coverage. **If reviewers would rather not change that
test's meaning, the alternative is to downgrade the guard to a warning —
flagging it explicitly as a decision.**

I also swept every shipped `mtq.*_CFG` against `SmallQKVModel`: only the
four MLP/experts-scoped configs raise, and the other three
(`NVFP4_EXPERTS_ONLY_CFG`, `MXFP4_MLP_WEIGHT_ONLY_CFG`,
`NVFP4_MLP_WEIGHT_ONLY_CFG`) are used elsewhere only against real MoE
models (Qwen3-MoE, gpt-oss), so no other test is affected.

Ran locally after rebasing onto current `main` (torch 2.11, transformers
5.5.4): `tests/unit/torch/quantization` + `tests/unit/recipe` — 1271
passed, 7 skipped. Full `tests/unit` (minus onnx, and puzzletron which
needs `hydra`): 2676 passed, with 4 pre-existing
`test_quant_aware_conversion.py` failures that reproduce unchanged on
clean `main`. GPU tests were not run locally (no suitable GPU); the
FSDP2 change above is reasoned from the CI failure, not re-run.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — deliberately. A config that
previously produced a `quant_algo: null` checkpoint now raises. Any such
run was already not doing what it claimed; the three scoping rules above
keep intentional non-weight quantization working.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ (Backward Breaking Changes)
- Did you get Claude approval on this PR?: ❌

### Additional Information

Pairs with #2202 (PTQ support for Step-3.7 MoE checkpoints), which fixes
the specific model that motivated this. Independent branches; either can
merge first.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Quantization now detects enabled weight-quantization patterns that do
not apply to any model weights and reports a clear validation error
before calibration.
- Broad wildcard patterns and nested quantizers are now handled
correctly.
- Overlapping patterns respect the final matching setting, including
later disabling rules.
- Existing quantized models can be refined using the parsed
configuration.
- Activation-only and explicitly disabled weight-quantization
configurations remain supported.
- Pipeline-parallel stages without targeted weights can bypass this
validation when configured to do so.

- **Documentation**
- Documented the process-wide override for bypassing unmatched
weight-quantizer validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Zhiyu
2026-09-11 13:26:35 -07:00
committed by GitHub
co-authored by Claude Opus 5
parent bd90a5ed51
commit c37a6948db
4 changed files with 274 additions and 4 deletions
+1
View File
@@ -87,6 +87,7 @@ Changelog
- Remove in-trainer quantization via ``QuantizationArguments.quant_cfg`` / ``--quant_cfg`` (deprecated in 0.45); use ``--recipe``. New recipes ``general/ptq/mxfp4_mlp_weight_only`` and ``general/ptq/nvfp4_mlp_weight_only`` replace ``MXFP4_MLP_WEIGHT_ONLY_CFG`` / ``NVFP4_MLP_WEIGHT_ONLY_CFG`` in the ``examples/gpt-oss`` QAT flow.
- Remove the ``QuantizationArgumentsWithConfig`` alias in ``modelopt.torch.quantization.plugins.transformers_trainer`` (deprecated in 0.45). Use ``QuantizationArguments``.
- Transformer Engine ``TEGroupedLinear`` (fused MoE experts) now uses **per-expert** weight quantization (one ``amax`` per expert) instead of a single shared ``amax``, so ModelOpt checkpoints containing quantized ``TEGroupedLinear`` modules saved before 0.47 are **not compatible** with 0.47. Re-run PTQ to regenerate compatible checkpoints.
- ``mtq.quantize`` now raises when a config asks for weight quantization but none of its weight-quantizer patterns match the model, instead of calibrating and exporting a silently unquantized checkpoint (``"quant_algo": null``). Configs that quantize activations or the KV cache only are unaffected, as are patterns that match and are then disabled by a later entry. If this fires, use the recipe for that architecture under ``modelopt_recipes/huggingface/<model_type>/`` or fix the module patterns. Set ``MODELOPT_SKIP_WEIGHT_QUANT_CHECK=1`` to disable the check process-wide, e.g. for a pipeline-parallel rank whose local stage legitimately has none of the targeted modules.
**Deprecations**
+83 -2
View File
@@ -147,6 +147,85 @@ def postprocess_amax(model: nn.Module, key: str, post_process_fn) -> nn.Module:
return model
_SKIP_WEIGHT_QUANT_CHECK_ENV = "MODELOPT_SKIP_WEIGHT_QUANT_CHECK"
def _check_weight_quantization_took_effect(model: nn.Module, config: QuantizeConfig) -> None:
"""Raise when a config asks for weight quantization but no weight quantizer is enabled.
A config whose module patterns do not match the model is not an error to
:func:`set_quantizer_by_cfg` — every pattern simply matches nothing — so the run
proceeds through calibration and export and produces a checkpoint that is silently
unquantized (``"quant_algo": null`` with an empty ``quantized_layers``). That has
bitten several MoE architectures whose module naming differs from the wildcards in
the general recipes, and it is only noticed when someone reads the exported config.
By the time this runs, :func:`set_quantizer_by_cfg` (or the ``apply_mode`` conversion
that calls it) has already applied ``config`` to ``model``, so each quantizer's
``is_enabled`` *is* the true outcome of that application — checking it directly cannot
diverge from what the config actually did. An earlier version of this check instead
re-derived "did this pattern match anything?" via a separate matcher call, which missed
the case of two *different* overlapping patterns (e.g. ``*weight_quantizer`` enabling
something a later, broader ``*`` then disables): the narrower pattern registered as
"matched" even though the quantizer it matched ended up disabled.
A config that never asks for weight quantization (activation-only or KV-cache-only)
must not raise, so the check first looks at the config's own intent — via each
pattern's *final* entry, since entries apply in order and the last one for a pattern
wins — before looking at the model at all.
"""
if os.environ.get(_SKIP_WEIGHT_QUANT_CHECK_ENV) == "1":
return
# Later entries override earlier ones, so only each pattern's final state states intent.
# A pattern naming ``weight_quantizer`` explicitly (the common case, e.g.
# ``*weight_quantizer``, ``*.experts.*weight_quantizer``) is caught by the substring
# check. A broad wildcard that never mentions "weight" -- a bare ``"*"`` catch-all, or
# ``"*_quantizer"`` -- can still match weight quantizers at runtime, so it must count
# too, or a config built only from patterns like that would never trip the guard
# regardless of what the model contains. ``fnmatch`` against the literal probe string
# ``"weight_quantizer"`` catches those (a pattern matching that bare name is, by
# construction, asking for one) without replacing the substring check: the probe alone
# would miss ``*.experts.*weight_quantizer`` (there is no ``.experts.`` in the probe
# string), which is what recognizes model-scoped patterns like the Step / MoE recipes use.
last_entry_per_pattern = {entry.quantizer_name: entry for entry in config.quant_cfg}
weight_patterns = [
pattern
for pattern, entry in last_entry_per_pattern.items()
if entry.enable
and ("weight_quantizer" in pattern or fnmatch.fnmatch("weight_quantizer", pattern))
]
if not weight_patterns:
return
# `SequentialQuantizer.is_enabled` delegates to its first member, so a list-valued `cfg`'s
# quantizers are already covered here without naming `SequentialQuantizer` explicitly:
# `named_modules()` recurses into the container and yields those children too, individually,
# named `...weight_quantizer.0` / `.1` (the substring match below still applies to them).
if any(
module.is_enabled
for name, module in model.named_modules()
if isinstance(module, TensorQuantizer) and "weight_quantizer" in name
):
return
patterns = "\n ".join(sorted(weight_patterns))
raise RuntimeError(
"The quantization config asks for weight quantization but no weight quantizer is "
f"enabled, so nothing would be quantized. These patterns asked for it:\n {patterns}\n"
"Either the patterns do not match this architecture's module names (check the "
"model-specific recipes under modelopt_recipes/huggingface/<model_type>/), or the "
"modules holding the weights were never converted to quantized modules (an "
"unsupported custom module, e.g. a trust_remote_code MoE layout).\n"
"Under pipeline parallelism, a rank whose local stage genuinely has none of the "
"targeted modules (e.g. a pure-attention stage under an experts-only recipe) hits "
"this too, while other ranks proceed into calibration -- a collective hang, not "
f"just a wrong per-rank verdict. Set {_SKIP_WEIGHT_QUANT_CHECK_ENV}=1 to bypass this "
"check in that situation -- note this is process-global, so it silences the check "
"on every rank, not only the one with the legitimately empty stage."
)
def quantize(
model: nn.Module,
config: dict[str, Any | QuantizeConfig],
@@ -244,12 +323,14 @@ def quantize(
Returns: A pytorch model which has been quantized and calibrated.
"""
quantize_config = QuantizeConfig(**dict(config))
if not is_quantized(model):
model = apply_mode(model, mode=[("quantize", dict(config))], registry=QuantizeModeRegistry)
else:
# Already quantized, so lets apply the quant_cfg from the config
quant_cfg = QuantizeConfig(**dict(config)).quant_cfg
set_quantizer_by_cfg(model, quant_cfg)
set_quantizer_by_cfg(model, quantize_config.quant_cfg)
# Fail before calibration rather than after exporting an unquantized checkpoint.
_check_weight_quantization_took_effect(model, quantize_config)
return calibrate(model, config.get("algorithm"), forward_loop=forward_loop)
+6 -2
View File
@@ -273,7 +273,9 @@ def test_fsdp2_weight_update_context_for_export(dist_workers):
# mtq.W4A8_AWQ_BETA_CFG, #TODO: Fix unit test for this case
# mtq.FP8_2D_BLOCKWISE_WEIGHT_ONLY_CFG, #TODO: Fix unit test for this case
mtq.W4A8_MXFP4_FP8_CFG,
mtq.NVFP4_MLP_ONLY_CFG,
# NVFP4_MLP_ONLY_CFG is omitted: SmallQKVModel has no MLP, so that config matched
# nothing and the parametrization exercised an unquantized model. NVFP4_OMLP_ONLY_CFG
# covers the same scoped-recipe shape here because it also selects `o_proj`.
mtq.NVFP4_OMLP_ONLY_CFG,
],
)
@@ -293,7 +295,9 @@ def test_fsdp2_weight_update_context_for_fuse_layers(dist_workers, quant_config,
# mtq.W4A8_AWQ_BETA_CFG, #TODO: Fix unit test for this case
# mtq.FP8_2D_BLOCKWISE_WEIGHT_ONLY_CFG, #TODO: Fix unit test for this case
mtq.W4A8_MXFP4_FP8_CFG,
mtq.NVFP4_MLP_ONLY_CFG,
# NVFP4_MLP_ONLY_CFG is omitted: SmallQKVModel has no MLP, so that config matched
# nothing and the parametrization exercised an unquantized model. NVFP4_OMLP_ONLY_CFG
# covers the same scoped-recipe shape here because it also selects `o_proj`.
mtq.NVFP4_OMLP_ONLY_CFG,
],
)
@@ -16,6 +16,8 @@
"""High-level tests for quantization."""
import copy
import os
from unittest import mock
import pytest
import torch
@@ -441,6 +443,188 @@ def test_enable_only_entry_preserves_attributes():
assert module.axis == 0, "axis should be preserved by enable-only entry"
def test_weight_patterns_matching_nothing_raise():
"""A config whose weight patterns match no module must fail, not quantize nothing.
Otherwise calibration and export run to completion and produce a checkpoint that is
silently unquantized (``"quant_algo": null``).
"""
model = SimpleLinear()
config = {
"quant_cfg": [
{"quantizer_name": "*", "enable": False},
# No module in this model is named `experts`.
{"quantizer_name": "*.experts.*weight_quantizer", "cfg": {"num_bits": 8, "axis": 0}},
],
"algorithm": "max",
}
with pytest.raises(RuntimeError, match="no weight quantizer is enabled"):
mtq.quantize(model, config, lambda m: m(m.get_input()))
def test_config_without_weight_quantization_is_allowed():
"""Activation-only configs quantize no weight on purpose and must still run."""
model = SimpleLinear()
config = {
"quant_cfg": [
{"quantizer_name": "*", "enable": False},
{"quantizer_name": "*input_quantizer", "cfg": {"num_bits": 8, "axis": None}},
],
"algorithm": "max",
}
model = mtq.quantize(model, config, lambda m: m(m.get_input()))
for name, module in model.named_modules():
if name.endswith("weight_quantizer"):
assert not module.is_enabled
def test_weight_quantizers_disabled_by_a_later_entry_are_allowed():
"""Patterns that match and are then switched off are a choice, not a mismatch."""
model = SimpleLinear()
config = {
"quant_cfg": [
{"quantizer_name": "*weight_quantizer", "cfg": {"num_bits": 4, "axis": 0}},
{"quantizer_name": "*weight_quantizer", "enable": False},
],
"algorithm": "max",
}
model = mtq.quantize(model, config, lambda m: m(m.get_input()))
for name, module in model.named_modules():
if name.endswith("weight_quantizer"):
assert not module.is_enabled
def test_sequential_weight_quantizers_do_not_trip_the_guard():
"""List-valued `cfg` builds `SequentialQuantizer`s, but the guard only ever looks at
`TensorQuantizer` instances -- this pins that it still doesn't wrongly raise for them.
`SequentialQuantizer` (itself an `nn.Sequential`) is not special-cased: its `TensorQuantizer`
children are reachable directly via `named_modules()`, individually named
`...weight_quantizer.0` / `.1`, so the substring match already finds them.
"""
model = SimpleLinear()
calib_data = [model.get_input() for _ in range(2)]
quantize_model_and_forward(model, copy.deepcopy(WINT4INT8_CFG), calib_data)
for name, module in model.named_modules():
if name.endswith("weight_quantizer"):
assert isinstance(module, SequentialQuantizer)
def test_fused_experts_quantizer_names_do_not_trip_the_guard():
"""Fused-experts quantizers are named `..._weight_quantizers.N`, plural and indexed, but
still contain `weight_quantizer` as a substring and so are still read by the guard.
"""
model = SimpleLinear()
mtq.quantize(model, mtq.INT8_DEFAULT_CFG, lambda m: m(m.get_input()))
# Rename as the fused-experts path does; the config's `*weight_quantizer` must still match.
linear = model.net[0]
linear.add_module("gate_up_proj_weight_quantizers", torch.nn.ModuleList([TensorQuantizer()]))
config = {
"quant_cfg": [
{"quantizer_name": "*", "enable": False},
{"quantizer_name": "*gate_up_proj_weight_quantizer", "cfg": {"num_bits": 8, "axis": 0}},
],
"algorithm": "max",
}
mtq.quantize(model, config, lambda m: m(m.get_input()))
def test_refining_an_already_quantized_model_does_not_raise():
"""A second config whose own patterns match nothing still sees the earlier weight
quantizers as enabled, so this is refining an already-quantized model, not a no-op run.
"""
model = SimpleLinear()
model = mtq.quantize(model, mtq.INT8_DEFAULT_CFG, lambda m: m(m.get_input()))
assert any(m.is_enabled for n, m in model.named_modules() if n.endswith("weight_quantizer"))
refinement = {
"quant_cfg": [
{"quantizer_name": "*.experts.*weight_quantizer", "cfg": {"num_bits": 4, "axis": 0}},
],
"algorithm": None,
}
mtq.quantize(model, refinement)
# The earlier weight quantization is untouched.
assert any(m.is_enabled for n, m in model.named_modules() if n.endswith("weight_quantizer"))
def test_weight_patterns_enabled_then_retracted_do_not_raise():
"""An unmatched pattern that a later entry disables asks for nothing by the end."""
model = SimpleLinear()
config = {
"quant_cfg": [
{"quantizer_name": "*", "enable": False},
{"quantizer_name": "*.missing.*weight_quantizer", "cfg": {"num_bits": 8, "axis": 0}},
{"quantizer_name": "*.missing.*weight_quantizer", "enable": False},
],
"algorithm": "max",
}
model = mtq.quantize(model, config, lambda m: m(m.get_input()))
for name, module in model.named_modules():
if name.endswith("weight_quantizer"):
assert not module.is_enabled
def test_bare_wildcard_pattern_matching_nothing_raises():
"""A bare `"*"` (or another pattern never mentioning "weight") still expresses weight
intent if it would match a `weight_quantizer` name -- and must still raise if nothing in
the model actually has one, exactly like an explicit `*weight_quantizer` pattern would.
A model with zero quantizable modules (e.g. `nn.Module()`, no Linear/Conv anywhere) is
the degenerate case where this matters: nothing in the config's own text says "weight",
so a substring-only check would silently return without ever looking at the model.
"""
model = torch.nn.Module() # no quantizable submodules at all
config = {"quant_cfg": [{"quantizer_name": "*", "cfg": {"num_bits": 8, "axis": 0}}]}
with pytest.raises(RuntimeError, match="no weight quantizer is enabled"):
mtq.quantize(model, config)
def test_overlapping_patterns_disabled_by_a_broader_later_one_still_raise():
"""A narrower pattern "matching" is not enough -- the final enabled state is what counts.
`*weight_quantizer` enables real quantizers, but the later, broader `*` disables
everything again; the config's net effect is still "nothing quantized" and must raise.
A check that asked "did any weight pattern match something" instead of "is anything
actually enabled" would miss this, since the narrower pattern did match.
"""
model = SimpleLinear()
config = {
"quant_cfg": [
{"quantizer_name": "*weight_quantizer", "cfg": {"num_bits": 8, "axis": 0}},
{"quantizer_name": "*", "enable": False},
],
"algorithm": "max",
}
with pytest.raises(RuntimeError, match="no weight quantizer is enabled"):
mtq.quantize(model, config, lambda m: m(m.get_input()))
def test_skip_weight_quant_check_env_var_bypasses_the_guard():
"""Documented escape hatch for pipeline-parallel ranks whose local stage legitimately
has none of the targeted modules (e.g. a pure-attention stage under an experts-only
recipe): raising there while other ranks proceed into calibration is a collective hang,
not just a wrong per-rank verdict.
"""
model = SimpleLinear()
config = {
"quant_cfg": [
{"quantizer_name": "*", "enable": False},
{"quantizer_name": "*.experts.*weight_quantizer", "cfg": {"num_bits": 8, "axis": 0}},
],
"algorithm": "max",
}
with mock.patch.dict(os.environ, {"MODELOPT_SKIP_WEIGHT_QUANT_CHECK": "1"}):
mtq.quantize(model, config, lambda m: m(m.get_input()))
def test_atomicity_later_cfg_entry_does_not_inherit_earlier():
"""When two cfg-bearing entries match the same quantizer, the second fully replaces the first."""
model = SimpleLinear()