From 3c87751903124deeb3cb5aaf17b19816cfd9d3de Mon Sep 17 00:00:00 2001 From: Shengliang Xu <106840466+shengliangxu@users.noreply.github.com> Date: Mon, 14 Sep 2026 13:32:28 -0700 Subject: [PATCH] Deprecate the single-format quantization CLI flags in favour of --recipe (#2426) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ### What does this PR do? Type of change: deprecation Deprecates the single-format quantization CLI flags in favour of `--recipe`. Passing one now emits a `DeprecationWarning`; nothing else changes. | script | flags | |---|---| | `examples/hf_ptq` | `--qformat`, `--kv_cache_qformat` | | `examples/megatron_bridge/quantize.py` | `--quant_cfg`, `--kv_cache_quant`, `--weight_only` | | `examples/torch_onnx` | `--qformat` | `--recipe` was already authoritative over all six — silently on `hf_ptq`, and with a runtime warning on `megatron_bridge` — and `modelopt/recipe/presets.py` already records the intent in a comment: *"the long-term direction is to retire `--qformat` / `--kv_cache_qformat` in favour of `--recipe`"*. This makes that a real deprecation. A recipe carries the quantization config, the calibration algorithm and the KV-cache setting in one file, so they cannot drift apart the way separate flags can. That drift is not hypothetical: the preset path applies no MTP exclusion while the recipe unit `default_disabled_quantizers` disables `mtp.*`, so the same model quantizes differently depending on which entry point was used. #### The warning fires only when a flag is actually passed `RecipeSupersededAction` is an `argparse.Action`, and argparse invokes an action only for options present on the command line — never for a default. That matters because several of these default to a *quantizing* value (`--qformat fp8`, `--kv_cache_qformat fp8_cast`); warning on the defaults would fire on every run, including runs that correctly use `--recipe` and never mention the flag. `examples/speculative_decoding/scripts/quantize_drafter.py` keeps `--qformat` undeprecated: it has no `--recipe`, so there would be nothing to migrate to. ### Usage ```bash # deprecated python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path --qformat nvfp4 --kv_cache_qformat fp8_cast # replacement python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path \ --recipe general/ptq/nvfp4_experts_only-kv_fp8_cast ``` ### Testing Three tests in `tests/examples/hf_ptq/test_hf_ptq_args.py`, all passing: - passing `--qformat` / `--kv_cache_qformat` raises `DeprecationWarning` and still parses the value; - omitting them raises nothing and leaves the defaults (`fp8`, `fp8_cast`) untouched; - the action stays wired to both flags, so a future edit cannot drop it while leaving the help text. Defaults and parsed values were diffed against `main` and are unchanged — the action stores exactly what `store` / `store_true` would have. `ruff` findings are at parity with `main` on every changed file. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — the flags still work, they only warn. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — under 0.48.0 Deprecations. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information Draft: the removal release for these flags is not decided here, only the deprecation. ## Summary by CodeRabbit * **Deprecations** * Legacy quantization CLI options now issue visible `FutureWarning` messages only when explicitly provided. * Use `--recipe` instead of deprecated options in Hugging Face PTQ, Megatron-Bridge, and torch-to-ONNX workflows. * Existing option values, defaults, and parsing behavior remain unchanged. * Weight AutoQuantize recipes without an explicit `kv_cache` setting continue to use `--kv_cache_qformat` as a fallback. --------- Signed-off-by: Shengliang Xu --- CHANGELOG.rst | 2 + examples/hf_ptq/hf_ptq.py | 13 +++-- examples/megatron_bridge/quantize.py | 25 +++++++-- examples/torch_onnx/torch_quant_to_onnx.py | 8 ++- modelopt/recipe/presets.py | 36 +++++++++++++ tests/examples/hf_ptq/test_hf_ptq_args.py | 57 +++++++++++++++++++- tests/unit/recipe/test_presets.py | 61 ++++++++++++++++++++++ 7 files changed, 192 insertions(+), 10 deletions(-) diff --git a/CHANGELOG.rst b/CHANGELOG.rst index 01a0fbe1b..f0e1e67dd 100755 --- a/CHANGELOG.rst +++ b/CHANGELOG.rst @@ -21,6 +21,8 @@ Changelog **Deprecations** +- The single-format quantization CLI flags are deprecated in favour of ``--recipe`` and will be removed in a future release; passing one now emits a ``FutureWarning``. ``examples/hf_ptq``: ``--qformat`` and ``--kv_cache_qformat``. ``examples/megatron_bridge/quantize.py``: ``--quant_cfg``, ``--kv_cache_quant`` and ``--weight_only``. ``examples/torch_onnx/torch_quant_to_onnx.py``: ``--qformat``. A recipe carries the quantization config, the calibration algorithm and the KV-cache setting in one file, so they cannot drift apart the way separate flags can -- and ``--recipe`` already took precedence over all six, silently on ``hf_ptq`` and with a warning on ``megatron_bridge`` -- with one gap the recipe closes rather than inherits: a weight AutoQuantize recipe that omits ``kv_cache`` still falls back to ``--kv_cache_qformat``, so set ``kv_cache`` in the recipe when migrating. Use a recipe from ``modelopt_recipes/general/ptq/`` or a model-specific one under ``modelopt_recipes/models/``. The warning fires only when a flag is passed explicitly: ``--qformat`` defaults to ``fp8`` and ``--kv_cache_qformat`` to ``fp8_cast``, so warning on the defaults would fire on every run, including runs that correctly use ``--recipe``. ``examples/speculative_decoding/scripts/quantize_drafter.py`` keeps ``--qformat`` undeprecated: it has no ``--recipe`` alternative yet. + - The TensorRT-LLM checkpoint export format is deprecated and will be removed in 0.49.0: ``export_tensorrt_llm_checkpoint`` and ``torch_to_tensorrt_llm_checkpoint`` now emit a ``DeprecationWarning`` on use. Use ``export_hf_checkpoint``, which exports a unified Hugging Face checkpoint deployable on TensorRT-LLM, vLLM and SGLang. Its implementation moved to ``modelopt.torch.export.trtllm``, so import those two functions from there and the ``ModelConfig`` dataclasses from ``modelopt.torch.export.trtllm.model_config``; both functions remain importable from ``modelopt.torch.export`` for this release only. **Bug Fixes** diff --git a/examples/hf_ptq/hf_ptq.py b/examples/hf_ptq/hf_ptq.py index 94ca3b39b..2c71a0b4d 100755 --- a/examples/hf_ptq/hf_ptq.py +++ b/examples/hf_ptq/hf_ptq.py @@ -71,7 +71,12 @@ import modelopt.torch.opt as mto import modelopt.torch.quantization as mtq import modelopt.torch.sparsity as mts from modelopt.recipe import ModelOptAutoQuantizeRecipe, ModelOptPTQRecipe, load_recipe -from modelopt.recipe.presets import KV_CACHE_NONE, KV_QUANT_CFG_CHOICES, QUANT_CFG_CHOICES +from modelopt.recipe.presets import ( + KV_CACHE_NONE, + KV_QUANT_CFG_CHOICES, + QUANT_CFG_CHOICES, + RecipeSupersededAction, +) from modelopt.torch.export import ( export_hf_checkpoint, export_hf_vllm_fq_checkpoint, @@ -1523,8 +1528,9 @@ def parse_args() -> argparse.Namespace: parser.add_argument("--device", default="cuda") parser.add_argument( "--qformat", + action=RecipeSupersededAction, help="Quantization format for single-format PTQ. For mixed-precision search, use an " - "AutoQuantize recipe via --recipe.", + "AutoQuantize recipe via --recipe. (deprecated: use --recipe)", default="fp8", ) parser.add_argument( @@ -1586,6 +1592,7 @@ def parse_args() -> argparse.Namespace: ) parser.add_argument( "--kv_cache_qformat", + action=RecipeSupersededAction, required=False, default="fp8_cast", choices=[KV_CACHE_NONE, *KV_QUANT_CFG_CHOICES], @@ -1596,7 +1603,7 @@ def parse_args() -> argparse.Namespace: "calibration; all other formats (fp8, nvfp4, ...) use data-driven calibration. " "With --recipe, the source depends on the recipe type: a PTQ recipe is " "authoritative for KV cache and ignores this flag; an AutoQuantize recipe " - "falls back to this flag unless it sets an explicit kv_cache field." + "falls back to this flag unless it sets an explicit kv_cache field. (deprecated: use --recipe)" ), ) parser.add_argument( diff --git a/examples/megatron_bridge/quantize.py b/examples/megatron_bridge/quantize.py index 6b7477228..cc63214aa 100644 --- a/examples/megatron_bridge/quantize.py +++ b/examples/megatron_bridge/quantize.py @@ -68,7 +68,12 @@ from transformers import AutoProcessor import modelopt.torch.quantization as mtq import modelopt.torch.utils.distributed as dist from modelopt.recipe import ModelOptPTQRecipe, load_recipe -from modelopt.recipe.presets import KV_CACHE_NONE, KV_QUANT_CFG_CHOICES, QUANT_CFG_CHOICES +from modelopt.recipe.presets import ( + KV_CACHE_NONE, + KV_QUANT_CFG_CHOICES, + QUANT_CFG_CHOICES, + RecipeSupersededAction, +) from modelopt.torch.utils import print_args, print_rank_0, warn_rank_0 from modelopt.torch.utils.dataset_utils import get_supported_datasets from modelopt.torch.utils.plugins.mbridge import ( @@ -137,9 +142,11 @@ def get_args() -> argparse.Namespace: ) parser.add_argument( "--quant_cfg", + action=RecipeSupersededAction, type=str, default=None, help=( + "(deprecated: use --recipe) " f"Quantization config. Preset names: {', '.join(QUANT_CFG_CHOICES)}. " "You can also pass any full config name exposed by modelopt (e.g. FP8_DEFAULT_CFG). " "Ignored when --recipe is set." @@ -147,15 +154,25 @@ def get_args() -> argparse.Namespace: ) parser.add_argument( "--kv_cache_quant", + action=RecipeSupersededAction, type=str, default=KV_CACHE_NONE, choices=[KV_CACHE_NONE, *KV_QUANT_CFG_CHOICES], - help="KV-cache quantization config to apply on top of --quant_cfg. Ignored when --recipe is set.", + help=( + "(deprecated: use --recipe) KV-cache quantization config to apply on top of " + "--quant_cfg. Ignored when --recipe is set." + ), ) parser.add_argument( "--weight_only", - action="store_true", - help="Disable input (activation) quantization, i.e. weight-only quantization.", + action=RecipeSupersededAction, + nargs=0, + const=True, + default=False, + help=( + "(deprecated: use --recipe) Disable input (activation) quantization, i.e. " + "weight-only quantization." + ), ) parser.add_argument( "--compress", diff --git a/examples/torch_onnx/torch_quant_to_onnx.py b/examples/torch_onnx/torch_quant_to_onnx.py index 7450f1242..cb27a56e5 100644 --- a/examples/torch_onnx/torch_quant_to_onnx.py +++ b/examples/torch_onnx/torch_quant_to_onnx.py @@ -35,7 +35,7 @@ from evaluation import evaluate import modelopt.torch.quantization as mtq from modelopt.recipe import ModelOptAutoQuantizeRecipe, ModelOptPTQRecipe, load_recipe -from modelopt.recipe.presets import QUANT_CFG_CHOICES +from modelopt.recipe.presets import QUANT_CFG_CHOICES, RecipeSupersededAction from modelopt.torch.quantization.nn import TensorQuantizer from modelopt.torch.quantization.plugins.custom import CUSTOM_POST_CONVERSION_PLUGINS @@ -517,9 +517,13 @@ def main(): ) parser.add_argument( "--qformat", + action=RecipeSupersededAction, choices=["fp8", "mxfp8", "int8", "nvfp4", "int4_awq", "auto"], default="mxfp8", - help="Quantization format to apply when --recipe is not provided. Default is MXFP8.", + help=( + "(deprecated: use --recipe) Quantization format to apply when --recipe is not " + "provided. Default is MXFP8." + ), ) parser.add_argument( "--recipe", diff --git a/modelopt/recipe/presets.py b/modelopt/recipe/presets.py index 90f36b8f8..6f7037338 100644 --- a/modelopt/recipe/presets.py +++ b/modelopt/recipe/presets.py @@ -32,6 +32,8 @@ that mutate a returned config must deepcopy it first (this mirrors how the ``mtq.*_CFG`` module constants — themselves eagerly-loaded shared dicts — are used). """ +import argparse +import warnings from typing import Any from modelopt.torch.opt.config_loader import BUILTIN_CONFIG_ROOT, load_config @@ -43,6 +45,7 @@ __all__ = [ "KV_QUANT_PRESET_DIR", "MODEL_QUANT_PRESET_DIR", "QUANT_CFG_CHOICES", + "RecipeSupersededAction", "load_quant_cfg_choices", ] @@ -95,3 +98,36 @@ KV_QUANT_CFG_CHOICES: dict[str, dict[str, Any]] = load_quant_cfg_choices(KV_QUAN assert KV_CACHE_NONE not in KV_QUANT_CFG_CHOICES, ( f"KV_CACHE_NONE sentinel {KV_CACHE_NONE!r} collides with a KV preset; rename the preset." ) + + +class RecipeSupersededAction(argparse.Action): + """``argparse`` action for a CLI flag that ``--recipe`` replaces. + + Warns only when the flag is actually passed: argparse invokes an action for options present on + the command line, never for a default. That distinction matters here because several of these + flags default to a *quantizing* value -- ``--qformat fp8``, ``--kv_cache_qformat fp8_cast`` -- + so warning unconditionally would fire on every run, including runs that correctly use + ``--recipe`` and never mention the deprecated flag. + + Handles both value-taking flags and ``store_true`` ones; for the latter pass + ``nargs=0, const=True``. + + The warning is a ``FutureWarning``, not a ``DeprecationWarning``. Python ignores + ``DeprecationWarning`` by default everywhere except ``__main__``, and argparse calls this action + from its own module, so the attributed frame is ``argparse`` and the default filters would drop + it -- the flag would go on working with nothing said, which defeats the point. ``FutureWarning`` + is the category Python documents for deprecations aimed at end users, and it is shown by + default. (Test suites enable all warnings, so this is invisible in tests either way.) + """ + + def __call__(self, parser, namespace, values, option_string=None): + """Warn that this flag is deprecated, then store the value as usual.""" + warnings.warn( + f"{option_string} is deprecated and will be removed in a future release. Use " + "--recipe with a YAML recipe instead: a recipe carries the quantization config, the " + "calibration algorithm and the KV-cache setting together, so they cannot drift apart. " + "See modelopt_recipes/general/ptq/ and modelopt.recipe.", + FutureWarning, + stacklevel=2, + ) + setattr(namespace, self.dest, self.const if self.nargs == 0 else values) diff --git a/tests/examples/hf_ptq/test_hf_ptq_args.py b/tests/examples/hf_ptq/test_hf_ptq_args.py index c3e04f253..15d760316 100644 --- a/tests/examples/hf_ptq/test_hf_ptq_args.py +++ b/tests/examples/hf_ptq/test_hf_ptq_args.py @@ -13,6 +13,7 @@ # See the License for the specific language governing permissions and # limitations under the License. +import argparse import getpass import importlib import json @@ -27,7 +28,7 @@ from _test_utils.torch.transformers_models import get_tiny_qwen3 from modelopt.recipe import load_recipe from modelopt.recipe.config import AutoQuantizeConfig, AutoQuantizeConstraints -from modelopt.recipe.presets import QUANT_CFG_CHOICES +from modelopt.recipe.presets import QUANT_CFG_CHOICES, RecipeSupersededAction from modelopt.torch.quantization import tensor_quant from modelopt.torch.quantization.config import QuantizeConfig @@ -817,3 +818,57 @@ def test_untracked_runs_write_no_experiment_json(monkeypatch, example_utils, tmp pass assert not (tmp_path / ".experiment.json").exists() + + +# --- flags that --recipe supersedes ------------------------------------------------------------- + + +@pytest.mark.parametrize( + ("flag", "value"), + [("--qformat", "nvfp4"), ("--kv_cache_qformat", "nvfp4")], +) +def test_recipe_superseded_flag_warns_when_passed(monkeypatch, flag, value): + """Passing one of these must say so; they are slated for removal in favour of --recipe.""" + with pytest.warns(FutureWarning, match=f"{flag} is deprecated"): + _, args = _parse_hf_ptq_args( + monkeypatch, "--pyt_ckpt_path", "/models/Qwen3-0.6B", flag, value + ) + assert getattr(args, flag.lstrip("-")) == value + + +def test_recipe_superseded_flags_are_silent_when_defaulted(monkeypatch, recwarn): + """The defaults quantize (--qformat fp8, --kv_cache_qformat fp8_cast), so warning on every + run -- including runs that correctly pass --recipe -- would be pure noise. argparse only + invokes an action for options actually present, which is what keeps this quiet.""" + _, args = _parse_hf_ptq_args(monkeypatch, "--pyt_ckpt_path", "/models/Qwen3-0.6B") + deprecations = [w for w in recwarn if issubclass(w.category, FutureWarning)] + assert not [w for w in deprecations if "is deprecated" in str(w.message)] + # and the defaults themselves are untouched by the deprecation wiring + assert args.qformat == "fp8" + assert args.kv_cache_qformat == "fp8_cast" + + +def test_recipe_superseded_action_is_wired_to_both_flags(monkeypatch): + """Guards against a future edit dropping the action while leaving the help text. + + Introspects the parser hf_ptq actually builds rather than its source text, so reordering + keyword arguments or reflowing the call does not fail the test while the wiring is intact. + """ + hf_ptq = _import_hf_ptq(monkeypatch) + monkeypatch.setattr(sys, "argv", ["hf_ptq.py", "--pyt_ckpt_path", "/models/Qwen3-0.6B"]) + + built = {} + real_parse_args = argparse.ArgumentParser.parse_args + + def capture(self, *args, **kwargs): + built.setdefault("parser", self) + return real_parse_args(self, *args, **kwargs) + + monkeypatch.setattr(argparse.ArgumentParser, "parse_args", capture) + hf_ptq.parse_args() + + by_dest = {action.dest: action for action in built["parser"]._actions} + for dest in ("qformat", "kv_cache_qformat"): + assert isinstance(by_dest[dest], RecipeSupersededAction), ( + f"--{dest} lost its deprecation action" + ) diff --git a/tests/unit/recipe/test_presets.py b/tests/unit/recipe/test_presets.py index 89645773f..af0f40b8d 100644 --- a/tests/unit/recipe/test_presets.py +++ b/tests/unit/recipe/test_presets.py @@ -21,10 +21,16 @@ preset YAML would otherwise break ``import modelopt.recipe.presets`` (and every PTQ example). """ +import argparse +import subprocess +import sys +import textwrap + import pytest import modelopt.torch.quantization as mtq from modelopt.recipe import load_recipe, presets +from modelopt.recipe.presets import RecipeSupersededAction from modelopt.torch.opt.config_loader import BUILTIN_CONFIG_ROOT from modelopt.torch.quantization.config import QuantizeConfig @@ -91,3 +97,58 @@ def test_mlp_weight_only_recipe_matches_its_mtq_cfg(recipe_name, cfg_name): recipe_cfg = load_recipe(recipe_name).quantize.model_dump(exclude_unset=True) mtq_cfg = QuantizeConfig(**getattr(mtq, cfg_name)).model_dump(exclude_unset=True) assert recipe_cfg == mtq_cfg + + +# --- RecipeSupersededAction: the flags --recipe replaces ---------------------------------------- + + +def _one_flag_parser(**kwargs): + parser = argparse.ArgumentParser() + parser.add_argument("--weight_only", action=RecipeSupersededAction, **kwargs) + return parser + + +def test_store_true_style_flag_defaults_without_warning(): + """``nargs=0`` flags default to False and stay silent -- argparse skips absent options.""" + args = _one_flag_parser(nargs=0, const=True, default=False).parse_args([]) + assert args.weight_only is False + + +def test_store_true_style_flag_stores_const_not_an_empty_list(): + """A ``nargs=0`` flag is handed ``[]``, so the action has to store ``const`` instead. + + ``--weight_only`` on megatron_bridge is the only caller of this branch; storing the empty list + would leave a falsy value and silently turn weight-only quantization off. + """ + parser = _one_flag_parser(nargs=0, const=True, default=False) + with pytest.warns(FutureWarning, match="--weight_only is deprecated"): + args = parser.parse_args(["--weight_only"]) + assert args.weight_only is True + + +def test_deprecation_reaches_stderr_under_the_real_default_filters(): + """The warning has to reach an actual CLI user, not just a test run. + + pytest enables every warning, so a category CPython suppresses looks healthy here and says + nothing in production. argparse invokes the action from its own module, so a + ``DeprecationWarning`` would be dropped by the default ``ignore::DeprecationWarning`` filter -- + the flag would keep working with nothing said. A subprocess is the only honest check: it uses + the interpreter's real filters rather than a reconstruction of them. + """ + script = textwrap.dedent( + """ + import argparse + from modelopt.recipe.presets import RecipeSupersededAction + + parser = argparse.ArgumentParser() + parser.add_argument("--qformat", action=RecipeSupersededAction) + parser.parse_args(["--qformat", "nvfp4"]) + """ + ) + proc = subprocess.run([sys.executable, "-c", script], capture_output=True, text=True) + + assert proc.returncode == 0, proc.stderr + assert "--qformat is deprecated" in proc.stderr, ( + "the deprecation is filtered out under Python's default filters, so a CLI user would " + f"never see it; stderr was: {proc.stderr!r}" + )