Deprecate the single-format quantization CLI flags in favour of --recipe (#2426)

### What does this PR do?

Type of change: deprecation

Deprecates the single-format quantization CLI flags in favour of
`--recipe`. Passing one now emits a `DeprecationWarning`; nothing else
changes.

| script | flags |
|---|---|
| `examples/hf_ptq` | `--qformat`, `--kv_cache_qformat` |
| `examples/megatron_bridge/quantize.py` | `--quant_cfg`,
`--kv_cache_quant`, `--weight_only` |
| `examples/torch_onnx` | `--qformat` |

`--recipe` was already authoritative over all six — silently on
`hf_ptq`, and with a runtime warning on `megatron_bridge` — and
`modelopt/recipe/presets.py` already records the intent in a comment:
*"the long-term direction is to retire `--qformat` /
`--kv_cache_qformat` in favour of `--recipe`"*. This makes that a real
deprecation.

A recipe carries the quantization config, the calibration algorithm and
the KV-cache setting in one file, so they cannot drift apart the way
separate flags can. That drift is not hypothetical: the preset path
applies no MTP exclusion while the recipe unit
`default_disabled_quantizers` disables `mtp.*`, so the same model
quantizes differently depending on which entry point was used.

#### The warning fires only when a flag is actually passed

`RecipeSupersededAction` is an `argparse.Action`, and argparse invokes
an action only for options present on the command line — never for a
default. That matters because several of these default to a *quantizing*
value (`--qformat fp8`, `--kv_cache_qformat fp8_cast`); warning on the
defaults would fire on every run, including runs that correctly use
`--recipe` and never mention the flag.

`examples/speculative_decoding/scripts/quantize_drafter.py` keeps
`--qformat` undeprecated: it has no `--recipe`, so there would be
nothing to migrate to.

### Usage

```bash
# deprecated
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> --qformat nvfp4 --kv_cache_qformat fp8_cast

# replacement
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> \
    --recipe general/ptq/nvfp4_experts_only-kv_fp8_cast
```

### Testing

Three tests in `tests/examples/hf_ptq/test_hf_ptq_args.py`, all passing:

- passing `--qformat` / `--kv_cache_qformat` raises `DeprecationWarning`
and still parses the value;
- omitting them raises nothing and leaves the defaults (`fp8`,
`fp8_cast`) untouched;
- the action stays wired to both flags, so a future edit cannot drop it
while leaving the help text.

Defaults and parsed values were diffed against `main` and are unchanged
— the action stores exactly what `store` / `store_true` would have.
`ruff` findings are at parity with `main` on every changed file.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — the flags still work, they
only warn.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 Deprecations.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Draft: the removal release for these flags is not decided here, only the
deprecation.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Deprecations**
* Legacy quantization CLI options now issue visible `FutureWarning`
messages only when explicitly provided.
* Use `--recipe` instead of deprecated options in Hugging Face PTQ,
Megatron-Bridge, and torch-to-ONNX workflows.
* Existing option values, defaults, and parsing behavior remain
unchanged.
* Weight AutoQuantize recipes without an explicit `kv_cache` setting
continue to use `--kv_cache_qformat` as a fallback.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
This commit is contained in:
Shengliang Xu
2026-09-14 13:32:28 -07:00
committed by GitHub
parent f70991f36e
commit 3c87751903
7 changed files with 192 additions and 10 deletions
+10 -3
View File
@@ -71,7 +71,12 @@ import modelopt.torch.opt as mto
import modelopt.torch.quantization as mtq
import modelopt.torch.sparsity as mts
from modelopt.recipe import ModelOptAutoQuantizeRecipe, ModelOptPTQRecipe, load_recipe
from modelopt.recipe.presets import KV_CACHE_NONE, KV_QUANT_CFG_CHOICES, QUANT_CFG_CHOICES
from modelopt.recipe.presets import (
KV_CACHE_NONE,
KV_QUANT_CFG_CHOICES,
QUANT_CFG_CHOICES,
RecipeSupersededAction,
)
from modelopt.torch.export import (
export_hf_checkpoint,
export_hf_vllm_fq_checkpoint,
@@ -1523,8 +1528,9 @@ def parse_args() -> argparse.Namespace:
parser.add_argument("--device", default="cuda")
parser.add_argument(
"--qformat",
action=RecipeSupersededAction,
help="Quantization format for single-format PTQ. For mixed-precision search, use an "
"AutoQuantize recipe via --recipe.",
"AutoQuantize recipe via --recipe. (deprecated: use --recipe)",
default="fp8",
)
parser.add_argument(
@@ -1586,6 +1592,7 @@ def parse_args() -> argparse.Namespace:
)
parser.add_argument(
"--kv_cache_qformat",
action=RecipeSupersededAction,
required=False,
default="fp8_cast",
choices=[KV_CACHE_NONE, *KV_QUANT_CFG_CHOICES],
@@ -1596,7 +1603,7 @@ def parse_args() -> argparse.Namespace:
"calibration; all other formats (fp8, nvfp4, ...) use data-driven calibration. "
"With --recipe, the source depends on the recipe type: a PTQ recipe is "
"authoritative for KV cache and ignores this flag; an AutoQuantize recipe "
"falls back to this flag unless it sets an explicit kv_cache field."
"falls back to this flag unless it sets an explicit kv_cache field. (deprecated: use --recipe)"
),
)
parser.add_argument(
+21 -4
View File
@@ -68,7 +68,12 @@ from transformers import AutoProcessor
import modelopt.torch.quantization as mtq
import modelopt.torch.utils.distributed as dist
from modelopt.recipe import ModelOptPTQRecipe, load_recipe
from modelopt.recipe.presets import KV_CACHE_NONE, KV_QUANT_CFG_CHOICES, QUANT_CFG_CHOICES
from modelopt.recipe.presets import (
KV_CACHE_NONE,
KV_QUANT_CFG_CHOICES,
QUANT_CFG_CHOICES,
RecipeSupersededAction,
)
from modelopt.torch.utils import print_args, print_rank_0, warn_rank_0
from modelopt.torch.utils.dataset_utils import get_supported_datasets
from modelopt.torch.utils.plugins.mbridge import (
@@ -137,9 +142,11 @@ def get_args() -> argparse.Namespace:
)
parser.add_argument(
"--quant_cfg",
action=RecipeSupersededAction,
type=str,
default=None,
help=(
"(deprecated: use --recipe) "
f"Quantization config. Preset names: {', '.join(QUANT_CFG_CHOICES)}. "
"You can also pass any full config name exposed by modelopt (e.g. FP8_DEFAULT_CFG). "
"Ignored when --recipe is set."
@@ -147,15 +154,25 @@ def get_args() -> argparse.Namespace:
)
parser.add_argument(
"--kv_cache_quant",
action=RecipeSupersededAction,
type=str,
default=KV_CACHE_NONE,
choices=[KV_CACHE_NONE, *KV_QUANT_CFG_CHOICES],
help="KV-cache quantization config to apply on top of --quant_cfg. Ignored when --recipe is set.",
help=(
"(deprecated: use --recipe) KV-cache quantization config to apply on top of "
"--quant_cfg. Ignored when --recipe is set."
),
)
parser.add_argument(
"--weight_only",
action="store_true",
help="Disable input (activation) quantization, i.e. weight-only quantization.",
action=RecipeSupersededAction,
nargs=0,
const=True,
default=False,
help=(
"(deprecated: use --recipe) Disable input (activation) quantization, i.e. "
"weight-only quantization."
),
)
parser.add_argument(
"--compress",
+6 -2
View File
@@ -35,7 +35,7 @@ from evaluation import evaluate
import modelopt.torch.quantization as mtq
from modelopt.recipe import ModelOptAutoQuantizeRecipe, ModelOptPTQRecipe, load_recipe
from modelopt.recipe.presets import QUANT_CFG_CHOICES
from modelopt.recipe.presets import QUANT_CFG_CHOICES, RecipeSupersededAction
from modelopt.torch.quantization.nn import TensorQuantizer
from modelopt.torch.quantization.plugins.custom import CUSTOM_POST_CONVERSION_PLUGINS
@@ -517,9 +517,13 @@ def main():
)
parser.add_argument(
"--qformat",
action=RecipeSupersededAction,
choices=["fp8", "mxfp8", "int8", "nvfp4", "int4_awq", "auto"],
default="mxfp8",
help="Quantization format to apply when --recipe is not provided. Default is MXFP8.",
help=(
"(deprecated: use --recipe) Quantization format to apply when --recipe is not "
"provided. Default is MXFP8."
),
)
parser.add_argument(
"--recipe",