Deprecate the single-format quantization CLI flags in favour of --recipe (#2426)

### What does this PR do?

Type of change: deprecation

Deprecates the single-format quantization CLI flags in favour of
`--recipe`. Passing one now emits a `DeprecationWarning`; nothing else
changes.

| script | flags |
|---|---|
| `examples/hf_ptq` | `--qformat`, `--kv_cache_qformat` |
| `examples/megatron_bridge/quantize.py` | `--quant_cfg`,
`--kv_cache_quant`, `--weight_only` |
| `examples/torch_onnx` | `--qformat` |

`--recipe` was already authoritative over all six — silently on
`hf_ptq`, and with a runtime warning on `megatron_bridge` — and
`modelopt/recipe/presets.py` already records the intent in a comment:
*"the long-term direction is to retire `--qformat` /
`--kv_cache_qformat` in favour of `--recipe`"*. This makes that a real
deprecation.

A recipe carries the quantization config, the calibration algorithm and
the KV-cache setting in one file, so they cannot drift apart the way
separate flags can. That drift is not hypothetical: the preset path
applies no MTP exclusion while the recipe unit
`default_disabled_quantizers` disables `mtp.*`, so the same model
quantizes differently depending on which entry point was used.

#### The warning fires only when a flag is actually passed

`RecipeSupersededAction` is an `argparse.Action`, and argparse invokes
an action only for options present on the command line — never for a
default. That matters because several of these default to a *quantizing*
value (`--qformat fp8`, `--kv_cache_qformat fp8_cast`); warning on the
defaults would fire on every run, including runs that correctly use
`--recipe` and never mention the flag.

`examples/speculative_decoding/scripts/quantize_drafter.py` keeps
`--qformat` undeprecated: it has no `--recipe`, so there would be
nothing to migrate to.

### Usage

```bash
# deprecated
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> --qformat nvfp4 --kv_cache_qformat fp8_cast

# replacement
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> \
    --recipe general/ptq/nvfp4_experts_only-kv_fp8_cast
```

### Testing

Three tests in `tests/examples/hf_ptq/test_hf_ptq_args.py`, all passing:

- passing `--qformat` / `--kv_cache_qformat` raises `DeprecationWarning`
and still parses the value;
- omitting them raises nothing and leaves the defaults (`fp8`,
`fp8_cast`) untouched;
- the action stays wired to both flags, so a future edit cannot drop it
while leaving the help text.

Defaults and parsed values were diffed against `main` and are unchanged
— the action stores exactly what `store` / `store_true` would have.
`ruff` findings are at parity with `main` on every changed file.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — the flags still work, they
only warn.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 Deprecations.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Draft: the removal release for these flags is not decided here, only the
deprecation.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Deprecations**
* Legacy quantization CLI options now issue visible `FutureWarning`
messages only when explicitly provided.
* Use `--recipe` instead of deprecated options in Hugging Face PTQ,
Megatron-Bridge, and torch-to-ONNX workflows.
* Existing option values, defaults, and parsing behavior remain
unchanged.
* Weight AutoQuantize recipes without an explicit `kv_cache` setting
continue to use `--kv_cache_qformat` as a fallback.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
This commit is contained in:
Shengliang Xu
2026-09-14 13:32:28 -07:00
committed by GitHub
parent f70991f36e
commit 3c87751903
7 changed files with 192 additions and 10 deletions
+56 -1
View File
@@ -13,6 +13,7 @@
# See the License for the specific language governing permissions and
# limitations under the License.
import argparse
import getpass
import importlib
import json
@@ -27,7 +28,7 @@ from _test_utils.torch.transformers_models import get_tiny_qwen3
from modelopt.recipe import load_recipe
from modelopt.recipe.config import AutoQuantizeConfig, AutoQuantizeConstraints
from modelopt.recipe.presets import QUANT_CFG_CHOICES
from modelopt.recipe.presets import QUANT_CFG_CHOICES, RecipeSupersededAction
from modelopt.torch.quantization import tensor_quant
from modelopt.torch.quantization.config import QuantizeConfig
@@ -817,3 +818,57 @@ def test_untracked_runs_write_no_experiment_json(monkeypatch, example_utils, tmp
pass
assert not (tmp_path / ".experiment.json").exists()
# --- flags that --recipe supersedes -------------------------------------------------------------
@pytest.mark.parametrize(
("flag", "value"),
[("--qformat", "nvfp4"), ("--kv_cache_qformat", "nvfp4")],
)
def test_recipe_superseded_flag_warns_when_passed(monkeypatch, flag, value):
"""Passing one of these must say so; they are slated for removal in favour of --recipe."""
with pytest.warns(FutureWarning, match=f"{flag} is deprecated"):
_, args = _parse_hf_ptq_args(
monkeypatch, "--pyt_ckpt_path", "/models/Qwen3-0.6B", flag, value
)
assert getattr(args, flag.lstrip("-")) == value
def test_recipe_superseded_flags_are_silent_when_defaulted(monkeypatch, recwarn):
"""The defaults quantize (--qformat fp8, --kv_cache_qformat fp8_cast), so warning on every
run -- including runs that correctly pass --recipe -- would be pure noise. argparse only
invokes an action for options actually present, which is what keeps this quiet."""
_, args = _parse_hf_ptq_args(monkeypatch, "--pyt_ckpt_path", "/models/Qwen3-0.6B")
deprecations = [w for w in recwarn if issubclass(w.category, FutureWarning)]
assert not [w for w in deprecations if "is deprecated" in str(w.message)]
# and the defaults themselves are untouched by the deprecation wiring
assert args.qformat == "fp8"
assert args.kv_cache_qformat == "fp8_cast"
def test_recipe_superseded_action_is_wired_to_both_flags(monkeypatch):
"""Guards against a future edit dropping the action while leaving the help text.
Introspects the parser hf_ptq actually builds rather than its source text, so reordering
keyword arguments or reflowing the call does not fail the test while the wiring is intact.
"""
hf_ptq = _import_hf_ptq(monkeypatch)
monkeypatch.setattr(sys, "argv", ["hf_ptq.py", "--pyt_ckpt_path", "/models/Qwen3-0.6B"])
built = {}
real_parse_args = argparse.ArgumentParser.parse_args
def capture(self, *args, **kwargs):
built.setdefault("parser", self)
return real_parse_args(self, *args, **kwargs)
monkeypatch.setattr(argparse.ArgumentParser, "parse_args", capture)
hf_ptq.parse_args()
by_dest = {action.dest: action for action in built["parser"]._actions}
for dest in ("qformat", "kv_cache_qformat"):
assert isinstance(by_dest[dest], RecipeSupersededAction), (
f"--{dest} lost its deprecation action"
)
+61
View File
@@ -21,10 +21,16 @@ preset YAML would otherwise break ``import modelopt.recipe.presets`` (and every
PTQ example).
"""
import argparse
import subprocess
import sys
import textwrap
import pytest
import modelopt.torch.quantization as mtq
from modelopt.recipe import load_recipe, presets
from modelopt.recipe.presets import RecipeSupersededAction
from modelopt.torch.opt.config_loader import BUILTIN_CONFIG_ROOT
from modelopt.torch.quantization.config import QuantizeConfig
@@ -91,3 +97,58 @@ def test_mlp_weight_only_recipe_matches_its_mtq_cfg(recipe_name, cfg_name):
recipe_cfg = load_recipe(recipe_name).quantize.model_dump(exclude_unset=True)
mtq_cfg = QuantizeConfig(**getattr(mtq, cfg_name)).model_dump(exclude_unset=True)
assert recipe_cfg == mtq_cfg
# --- RecipeSupersededAction: the flags --recipe replaces ----------------------------------------
def _one_flag_parser(**kwargs):
parser = argparse.ArgumentParser()
parser.add_argument("--weight_only", action=RecipeSupersededAction, **kwargs)
return parser
def test_store_true_style_flag_defaults_without_warning():
"""``nargs=0`` flags default to False and stay silent -- argparse skips absent options."""
args = _one_flag_parser(nargs=0, const=True, default=False).parse_args([])
assert args.weight_only is False
def test_store_true_style_flag_stores_const_not_an_empty_list():
"""A ``nargs=0`` flag is handed ``[]``, so the action has to store ``const`` instead.
``--weight_only`` on megatron_bridge is the only caller of this branch; storing the empty list
would leave a falsy value and silently turn weight-only quantization off.
"""
parser = _one_flag_parser(nargs=0, const=True, default=False)
with pytest.warns(FutureWarning, match="--weight_only is deprecated"):
args = parser.parse_args(["--weight_only"])
assert args.weight_only is True
def test_deprecation_reaches_stderr_under_the_real_default_filters():
"""The warning has to reach an actual CLI user, not just a test run.
pytest enables every warning, so a category CPython suppresses looks healthy here and says
nothing in production. argparse invokes the action from its own module, so a
``DeprecationWarning`` would be dropped by the default ``ignore::DeprecationWarning`` filter --
the flag would keep working with nothing said. A subprocess is the only honest check: it uses
the interpreter's real filters rather than a reconstruction of them.
"""
script = textwrap.dedent(
"""
import argparse
from modelopt.recipe.presets import RecipeSupersededAction
parser = argparse.ArgumentParser()
parser.add_argument("--qformat", action=RecipeSupersededAction)
parser.parse_args(["--qformat", "nvfp4"])
"""
)
proc = subprocess.run([sys.executable, "-c", script], capture_output=True, text=True)
assert proc.returncode == 0, proc.stderr
assert "--qformat is deprecated" in proc.stderr, (
"the deprecation is filtered out under Python's default filters, so a CLI user would "
f"never see it; stderr was: {proc.stderr!r}"
)