mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Deprecate the single-format quantization CLI flags in favour of --recipe (#2426)
### What does this PR do?
Type of change: deprecation
Deprecates the single-format quantization CLI flags in favour of
`--recipe`. Passing one now emits a `DeprecationWarning`; nothing else
changes.
| script | flags |
|---|---|
| `examples/hf_ptq` | `--qformat`, `--kv_cache_qformat` |
| `examples/megatron_bridge/quantize.py` | `--quant_cfg`,
`--kv_cache_quant`, `--weight_only` |
| `examples/torch_onnx` | `--qformat` |
`--recipe` was already authoritative over all six — silently on
`hf_ptq`, and with a runtime warning on `megatron_bridge` — and
`modelopt/recipe/presets.py` already records the intent in a comment:
*"the long-term direction is to retire `--qformat` /
`--kv_cache_qformat` in favour of `--recipe`"*. This makes that a real
deprecation.
A recipe carries the quantization config, the calibration algorithm and
the KV-cache setting in one file, so they cannot drift apart the way
separate flags can. That drift is not hypothetical: the preset path
applies no MTP exclusion while the recipe unit
`default_disabled_quantizers` disables `mtp.*`, so the same model
quantizes differently depending on which entry point was used.
#### The warning fires only when a flag is actually passed
`RecipeSupersededAction` is an `argparse.Action`, and argparse invokes
an action only for options present on the command line — never for a
default. That matters because several of these default to a *quantizing*
value (`--qformat fp8`, `--kv_cache_qformat fp8_cast`); warning on the
defaults would fire on every run, including runs that correctly use
`--recipe` and never mention the flag.
`examples/speculative_decoding/scripts/quantize_drafter.py` keeps
`--qformat` undeprecated: it has no `--recipe`, so there would be
nothing to migrate to.
### Usage
```bash
# deprecated
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> --qformat nvfp4 --kv_cache_qformat fp8_cast
# replacement
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> \
--recipe general/ptq/nvfp4_experts_only-kv_fp8_cast
```
### Testing
Three tests in `tests/examples/hf_ptq/test_hf_ptq_args.py`, all passing:
- passing `--qformat` / `--kv_cache_qformat` raises `DeprecationWarning`
and still parses the value;
- omitting them raises nothing and leaves the defaults (`fp8`,
`fp8_cast`) untouched;
- the action stays wired to both flags, so a future edit cannot drop it
while leaving the help text.
Defaults and parsed values were diffed against `main` and are unchanged
— the action stores exactly what `store` / `store_true` would have.
`ruff` findings are at parity with `main` on every changed file.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — the flags still work, they
only warn.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 Deprecations.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
Draft: the removal release for these flags is not decided here, only the
deprecation.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Deprecations**
* Legacy quantization CLI options now issue visible `FutureWarning`
messages only when explicitly provided.
* Use `--recipe` instead of deprecated options in Hugging Face PTQ,
Megatron-Bridge, and torch-to-ONNX workflows.
* Existing option values, defaults, and parsing behavior remain
unchanged.
* Weight AutoQuantize recipes without an explicit `kv_cache` setting
continue to use `--kv_cache_qformat` as a fallback.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
This commit is contained in:
@@ -71,7 +71,12 @@ import modelopt.torch.opt as mto
|
||||
import modelopt.torch.quantization as mtq
|
||||
import modelopt.torch.sparsity as mts
|
||||
from modelopt.recipe import ModelOptAutoQuantizeRecipe, ModelOptPTQRecipe, load_recipe
|
||||
from modelopt.recipe.presets import KV_CACHE_NONE, KV_QUANT_CFG_CHOICES, QUANT_CFG_CHOICES
|
||||
from modelopt.recipe.presets import (
|
||||
KV_CACHE_NONE,
|
||||
KV_QUANT_CFG_CHOICES,
|
||||
QUANT_CFG_CHOICES,
|
||||
RecipeSupersededAction,
|
||||
)
|
||||
from modelopt.torch.export import (
|
||||
export_hf_checkpoint,
|
||||
export_hf_vllm_fq_checkpoint,
|
||||
@@ -1523,8 +1528,9 @@ def parse_args() -> argparse.Namespace:
|
||||
parser.add_argument("--device", default="cuda")
|
||||
parser.add_argument(
|
||||
"--qformat",
|
||||
action=RecipeSupersededAction,
|
||||
help="Quantization format for single-format PTQ. For mixed-precision search, use an "
|
||||
"AutoQuantize recipe via --recipe.",
|
||||
"AutoQuantize recipe via --recipe. (deprecated: use --recipe)",
|
||||
default="fp8",
|
||||
)
|
||||
parser.add_argument(
|
||||
@@ -1586,6 +1592,7 @@ def parse_args() -> argparse.Namespace:
|
||||
)
|
||||
parser.add_argument(
|
||||
"--kv_cache_qformat",
|
||||
action=RecipeSupersededAction,
|
||||
required=False,
|
||||
default="fp8_cast",
|
||||
choices=[KV_CACHE_NONE, *KV_QUANT_CFG_CHOICES],
|
||||
@@ -1596,7 +1603,7 @@ def parse_args() -> argparse.Namespace:
|
||||
"calibration; all other formats (fp8, nvfp4, ...) use data-driven calibration. "
|
||||
"With --recipe, the source depends on the recipe type: a PTQ recipe is "
|
||||
"authoritative for KV cache and ignores this flag; an AutoQuantize recipe "
|
||||
"falls back to this flag unless it sets an explicit kv_cache field."
|
||||
"falls back to this flag unless it sets an explicit kv_cache field. (deprecated: use --recipe)"
|
||||
),
|
||||
)
|
||||
parser.add_argument(
|
||||
|
||||
@@ -68,7 +68,12 @@ from transformers import AutoProcessor
|
||||
import modelopt.torch.quantization as mtq
|
||||
import modelopt.torch.utils.distributed as dist
|
||||
from modelopt.recipe import ModelOptPTQRecipe, load_recipe
|
||||
from modelopt.recipe.presets import KV_CACHE_NONE, KV_QUANT_CFG_CHOICES, QUANT_CFG_CHOICES
|
||||
from modelopt.recipe.presets import (
|
||||
KV_CACHE_NONE,
|
||||
KV_QUANT_CFG_CHOICES,
|
||||
QUANT_CFG_CHOICES,
|
||||
RecipeSupersededAction,
|
||||
)
|
||||
from modelopt.torch.utils import print_args, print_rank_0, warn_rank_0
|
||||
from modelopt.torch.utils.dataset_utils import get_supported_datasets
|
||||
from modelopt.torch.utils.plugins.mbridge import (
|
||||
@@ -137,9 +142,11 @@ def get_args() -> argparse.Namespace:
|
||||
)
|
||||
parser.add_argument(
|
||||
"--quant_cfg",
|
||||
action=RecipeSupersededAction,
|
||||
type=str,
|
||||
default=None,
|
||||
help=(
|
||||
"(deprecated: use --recipe) "
|
||||
f"Quantization config. Preset names: {', '.join(QUANT_CFG_CHOICES)}. "
|
||||
"You can also pass any full config name exposed by modelopt (e.g. FP8_DEFAULT_CFG). "
|
||||
"Ignored when --recipe is set."
|
||||
@@ -147,15 +154,25 @@ def get_args() -> argparse.Namespace:
|
||||
)
|
||||
parser.add_argument(
|
||||
"--kv_cache_quant",
|
||||
action=RecipeSupersededAction,
|
||||
type=str,
|
||||
default=KV_CACHE_NONE,
|
||||
choices=[KV_CACHE_NONE, *KV_QUANT_CFG_CHOICES],
|
||||
help="KV-cache quantization config to apply on top of --quant_cfg. Ignored when --recipe is set.",
|
||||
help=(
|
||||
"(deprecated: use --recipe) KV-cache quantization config to apply on top of "
|
||||
"--quant_cfg. Ignored when --recipe is set."
|
||||
),
|
||||
)
|
||||
parser.add_argument(
|
||||
"--weight_only",
|
||||
action="store_true",
|
||||
help="Disable input (activation) quantization, i.e. weight-only quantization.",
|
||||
action=RecipeSupersededAction,
|
||||
nargs=0,
|
||||
const=True,
|
||||
default=False,
|
||||
help=(
|
||||
"(deprecated: use --recipe) Disable input (activation) quantization, i.e. "
|
||||
"weight-only quantization."
|
||||
),
|
||||
)
|
||||
parser.add_argument(
|
||||
"--compress",
|
||||
|
||||
@@ -35,7 +35,7 @@ from evaluation import evaluate
|
||||
|
||||
import modelopt.torch.quantization as mtq
|
||||
from modelopt.recipe import ModelOptAutoQuantizeRecipe, ModelOptPTQRecipe, load_recipe
|
||||
from modelopt.recipe.presets import QUANT_CFG_CHOICES
|
||||
from modelopt.recipe.presets import QUANT_CFG_CHOICES, RecipeSupersededAction
|
||||
from modelopt.torch.quantization.nn import TensorQuantizer
|
||||
from modelopt.torch.quantization.plugins.custom import CUSTOM_POST_CONVERSION_PLUGINS
|
||||
|
||||
@@ -517,9 +517,13 @@ def main():
|
||||
)
|
||||
parser.add_argument(
|
||||
"--qformat",
|
||||
action=RecipeSupersededAction,
|
||||
choices=["fp8", "mxfp8", "int8", "nvfp4", "int4_awq", "auto"],
|
||||
default="mxfp8",
|
||||
help="Quantization format to apply when --recipe is not provided. Default is MXFP8.",
|
||||
help=(
|
||||
"(deprecated: use --recipe) Quantization format to apply when --recipe is not "
|
||||
"provided. Default is MXFP8."
|
||||
),
|
||||
)
|
||||
parser.add_argument(
|
||||
"--recipe",
|
||||
|
||||
Reference in New Issue
Block a user