mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Deprecate the single-format quantization CLI flags in favour of --recipe (#2426)
### What does this PR do?
Type of change: deprecation
Deprecates the single-format quantization CLI flags in favour of
`--recipe`. Passing one now emits a `DeprecationWarning`; nothing else
changes.
| script | flags |
|---|---|
| `examples/hf_ptq` | `--qformat`, `--kv_cache_qformat` |
| `examples/megatron_bridge/quantize.py` | `--quant_cfg`,
`--kv_cache_quant`, `--weight_only` |
| `examples/torch_onnx` | `--qformat` |
`--recipe` was already authoritative over all six — silently on
`hf_ptq`, and with a runtime warning on `megatron_bridge` — and
`modelopt/recipe/presets.py` already records the intent in a comment:
*"the long-term direction is to retire `--qformat` /
`--kv_cache_qformat` in favour of `--recipe`"*. This makes that a real
deprecation.
A recipe carries the quantization config, the calibration algorithm and
the KV-cache setting in one file, so they cannot drift apart the way
separate flags can. That drift is not hypothetical: the preset path
applies no MTP exclusion while the recipe unit
`default_disabled_quantizers` disables `mtp.*`, so the same model
quantizes differently depending on which entry point was used.
#### The warning fires only when a flag is actually passed
`RecipeSupersededAction` is an `argparse.Action`, and argparse invokes
an action only for options present on the command line — never for a
default. That matters because several of these default to a *quantizing*
value (`--qformat fp8`, `--kv_cache_qformat fp8_cast`); warning on the
defaults would fire on every run, including runs that correctly use
`--recipe` and never mention the flag.
`examples/speculative_decoding/scripts/quantize_drafter.py` keeps
`--qformat` undeprecated: it has no `--recipe`, so there would be
nothing to migrate to.
### Usage
```bash
# deprecated
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> --qformat nvfp4 --kv_cache_qformat fp8_cast
# replacement
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> \
--recipe general/ptq/nvfp4_experts_only-kv_fp8_cast
```
### Testing
Three tests in `tests/examples/hf_ptq/test_hf_ptq_args.py`, all passing:
- passing `--qformat` / `--kv_cache_qformat` raises `DeprecationWarning`
and still parses the value;
- omitting them raises nothing and leaves the defaults (`fp8`,
`fp8_cast`) untouched;
- the action stays wired to both flags, so a future edit cannot drop it
while leaving the help text.
Defaults and parsed values were diffed against `main` and are unchanged
— the action stores exactly what `store` / `store_true` would have.
`ruff` findings are at parity with `main` on every changed file.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — the flags still work, they
only warn.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 Deprecations.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
Draft: the removal release for these flags is not decided here, only the
deprecation.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Deprecations**
* Legacy quantization CLI options now issue visible `FutureWarning`
messages only when explicitly provided.
* Use `--recipe` instead of deprecated options in Hugging Face PTQ,
Megatron-Bridge, and torch-to-ONNX workflows.
* Existing option values, defaults, and parsing behavior remain
unchanged.
* Weight AutoQuantize recipes without an explicit `kv_cache` setting
continue to use `--kv_cache_qformat` as a fallback.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
This commit is contained in:
@@ -13,6 +13,7 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
import argparse
|
||||
import getpass
|
||||
import importlib
|
||||
import json
|
||||
@@ -27,7 +28,7 @@ from _test_utils.torch.transformers_models import get_tiny_qwen3
|
||||
|
||||
from modelopt.recipe import load_recipe
|
||||
from modelopt.recipe.config import AutoQuantizeConfig, AutoQuantizeConstraints
|
||||
from modelopt.recipe.presets import QUANT_CFG_CHOICES
|
||||
from modelopt.recipe.presets import QUANT_CFG_CHOICES, RecipeSupersededAction
|
||||
from modelopt.torch.quantization import tensor_quant
|
||||
from modelopt.torch.quantization.config import QuantizeConfig
|
||||
|
||||
@@ -817,3 +818,57 @@ def test_untracked_runs_write_no_experiment_json(monkeypatch, example_utils, tmp
|
||||
pass
|
||||
|
||||
assert not (tmp_path / ".experiment.json").exists()
|
||||
|
||||
|
||||
# --- flags that --recipe supersedes -------------------------------------------------------------
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("flag", "value"),
|
||||
[("--qformat", "nvfp4"), ("--kv_cache_qformat", "nvfp4")],
|
||||
)
|
||||
def test_recipe_superseded_flag_warns_when_passed(monkeypatch, flag, value):
|
||||
"""Passing one of these must say so; they are slated for removal in favour of --recipe."""
|
||||
with pytest.warns(FutureWarning, match=f"{flag} is deprecated"):
|
||||
_, args = _parse_hf_ptq_args(
|
||||
monkeypatch, "--pyt_ckpt_path", "/models/Qwen3-0.6B", flag, value
|
||||
)
|
||||
assert getattr(args, flag.lstrip("-")) == value
|
||||
|
||||
|
||||
def test_recipe_superseded_flags_are_silent_when_defaulted(monkeypatch, recwarn):
|
||||
"""The defaults quantize (--qformat fp8, --kv_cache_qformat fp8_cast), so warning on every
|
||||
run -- including runs that correctly pass --recipe -- would be pure noise. argparse only
|
||||
invokes an action for options actually present, which is what keeps this quiet."""
|
||||
_, args = _parse_hf_ptq_args(monkeypatch, "--pyt_ckpt_path", "/models/Qwen3-0.6B")
|
||||
deprecations = [w for w in recwarn if issubclass(w.category, FutureWarning)]
|
||||
assert not [w for w in deprecations if "is deprecated" in str(w.message)]
|
||||
# and the defaults themselves are untouched by the deprecation wiring
|
||||
assert args.qformat == "fp8"
|
||||
assert args.kv_cache_qformat == "fp8_cast"
|
||||
|
||||
|
||||
def test_recipe_superseded_action_is_wired_to_both_flags(monkeypatch):
|
||||
"""Guards against a future edit dropping the action while leaving the help text.
|
||||
|
||||
Introspects the parser hf_ptq actually builds rather than its source text, so reordering
|
||||
keyword arguments or reflowing the call does not fail the test while the wiring is intact.
|
||||
"""
|
||||
hf_ptq = _import_hf_ptq(monkeypatch)
|
||||
monkeypatch.setattr(sys, "argv", ["hf_ptq.py", "--pyt_ckpt_path", "/models/Qwen3-0.6B"])
|
||||
|
||||
built = {}
|
||||
real_parse_args = argparse.ArgumentParser.parse_args
|
||||
|
||||
def capture(self, *args, **kwargs):
|
||||
built.setdefault("parser", self)
|
||||
return real_parse_args(self, *args, **kwargs)
|
||||
|
||||
monkeypatch.setattr(argparse.ArgumentParser, "parse_args", capture)
|
||||
hf_ptq.parse_args()
|
||||
|
||||
by_dest = {action.dest: action for action in built["parser"]._actions}
|
||||
for dest in ("qformat", "kv_cache_qformat"):
|
||||
assert isinstance(by_dest[dest], RecipeSupersededAction), (
|
||||
f"--{dest} lost its deprecation action"
|
||||
)
|
||||
|
||||
@@ -21,10 +21,16 @@ preset YAML would otherwise break ``import modelopt.recipe.presets`` (and every
|
||||
PTQ example).
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import subprocess
|
||||
import sys
|
||||
import textwrap
|
||||
|
||||
import pytest
|
||||
|
||||
import modelopt.torch.quantization as mtq
|
||||
from modelopt.recipe import load_recipe, presets
|
||||
from modelopt.recipe.presets import RecipeSupersededAction
|
||||
from modelopt.torch.opt.config_loader import BUILTIN_CONFIG_ROOT
|
||||
from modelopt.torch.quantization.config import QuantizeConfig
|
||||
|
||||
@@ -91,3 +97,58 @@ def test_mlp_weight_only_recipe_matches_its_mtq_cfg(recipe_name, cfg_name):
|
||||
recipe_cfg = load_recipe(recipe_name).quantize.model_dump(exclude_unset=True)
|
||||
mtq_cfg = QuantizeConfig(**getattr(mtq, cfg_name)).model_dump(exclude_unset=True)
|
||||
assert recipe_cfg == mtq_cfg
|
||||
|
||||
|
||||
# --- RecipeSupersededAction: the flags --recipe replaces ----------------------------------------
|
||||
|
||||
|
||||
def _one_flag_parser(**kwargs):
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--weight_only", action=RecipeSupersededAction, **kwargs)
|
||||
return parser
|
||||
|
||||
|
||||
def test_store_true_style_flag_defaults_without_warning():
|
||||
"""``nargs=0`` flags default to False and stay silent -- argparse skips absent options."""
|
||||
args = _one_flag_parser(nargs=0, const=True, default=False).parse_args([])
|
||||
assert args.weight_only is False
|
||||
|
||||
|
||||
def test_store_true_style_flag_stores_const_not_an_empty_list():
|
||||
"""A ``nargs=0`` flag is handed ``[]``, so the action has to store ``const`` instead.
|
||||
|
||||
``--weight_only`` on megatron_bridge is the only caller of this branch; storing the empty list
|
||||
would leave a falsy value and silently turn weight-only quantization off.
|
||||
"""
|
||||
parser = _one_flag_parser(nargs=0, const=True, default=False)
|
||||
with pytest.warns(FutureWarning, match="--weight_only is deprecated"):
|
||||
args = parser.parse_args(["--weight_only"])
|
||||
assert args.weight_only is True
|
||||
|
||||
|
||||
def test_deprecation_reaches_stderr_under_the_real_default_filters():
|
||||
"""The warning has to reach an actual CLI user, not just a test run.
|
||||
|
||||
pytest enables every warning, so a category CPython suppresses looks healthy here and says
|
||||
nothing in production. argparse invokes the action from its own module, so a
|
||||
``DeprecationWarning`` would be dropped by the default ``ignore::DeprecationWarning`` filter --
|
||||
the flag would keep working with nothing said. A subprocess is the only honest check: it uses
|
||||
the interpreter's real filters rather than a reconstruction of them.
|
||||
"""
|
||||
script = textwrap.dedent(
|
||||
"""
|
||||
import argparse
|
||||
from modelopt.recipe.presets import RecipeSupersededAction
|
||||
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--qformat", action=RecipeSupersededAction)
|
||||
parser.parse_args(["--qformat", "nvfp4"])
|
||||
"""
|
||||
)
|
||||
proc = subprocess.run([sys.executable, "-c", script], capture_output=True, text=True)
|
||||
|
||||
assert proc.returncode == 0, proc.stderr
|
||||
assert "--qformat is deprecated" in proc.stderr, (
|
||||
"the deprecation is filtered out under Python's default filters, so a CLI user would "
|
||||
f"never see it; stderr was: {proc.stderr!r}"
|
||||
)
|
||||
|
||||
Reference in New Issue
Block a user