mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do?
Type of change: Backward breaking change (deprecation removal)
Ahead of the 0.47 code freeze, this removes every deprecation still
outstanding from the previous two releases (0.45 and 0.46). Two are
intentionally left in place: the **Python 3.10** drop and the
**transformers 4.x** drop
| Deprecation | Marked in | Replacement |
| --- | --- | --- |
| `--auto_quantize_bits` / `_method` / `_score_size` / `_cost_model` /
`_active_moe_expert_ratio` | 0.46 | AutoQuantize `--recipe` |
| `examples/llm_ptq` symlink + `examples/vlm_ptq/` forwarder | 0.46 |
`examples/hf_ptq` (`--vlm` for VLMs) |
| `QuantizationArgumentsWithConfig` alias | 0.45 |
`QuantizationArguments` |
| `QFORMAT_ALIASES` short names | 0.45 | canonical preset basenames |
| `layerwise` bool + flat `layerwise_checkpoint_dir` | 0.45 | nested
`layerwise: {enable, checkpoint_dir}` |
| in-trainer `quant_cfg` / `--quant_cfg` | 0.45 | `--recipe` |
#### Two things worth a closer look
**1. The `use_sequential` alias goes too.** It is the pre-#1251 alias on
`QuantizeAlgorithmConfig.layerwise` and only ever carried a bool. Once
the bool form is rejected it cannot accept a valid value, so keeping it
would only produce a differently-worded validation error. Note the
direction is breaking either way (`extra="forbid"`): a pre-0.45
`modelopt_state` carrying `use_sequential: True` or a top-level
`layerwise_checkpoint_dir` now fails validation instead of being
migrated.
**2. Removing in-trainer `--quant_cfg` required two new recipes.** The
`examples/gpt-oss` QAT flow ran on `--quant_cfg
MXFP4_MLP_WEIGHT_ONLY_CFG` and no `general/ptq/` recipe covered it. This
PR adds `general/ptq/mxfp4_mlp_weight_only` and
`general/ptq/nvfp4_mlp_weight_only`, verified to `model_dump` identical
to `mtq.MXFP4_MLP_WEIGHT_ONLY_CFG` / `mtq.NVFP4_MLP_WEIGHT_ONLY_CFG`,
and migrates the gpt-oss README, both SFT configs, `sft.py` and
`tests/examples/gpt-oss/test_gpt_oss_qat.py`. `examples/llm_qat` was
already recipe-only.
### Usage
```bash
# AutoQuantize: --auto_quantize_* flags -> an AutoQuantize recipe
scripts/huggingface_example.sh --model $HF_PATH \
--recipe general/auto_quantize/nvfp4_fp8_at_5p4bits --calib_batch_size 4
# --qformat / --quant_cfg: short name -> canonical preset basename
# int8_sq -> int8_smoothquant nvfp4_mse -> nvfp4_w4a4_weight_mse_fp8_sweep
# int8_wo -> int8_weight_only nvfp4_local_hessian -> nvfp4_w4a4_weight_local_hessian
# w4a8_awq -> w4a8_awq_beta fp8_pb_wo -> fp8_2d_blockwise_weight_only
# nvfp4_awq -> nvfp4_awq_lite fp8_pc_pt -> fp8_per_channel_per_token
scripts/huggingface_example.sh --model $HF_PATH --quant int8_smoothquant
# VLM PTQ: examples/vlm_ptq -> examples/hf_ptq with --vlm
scripts/huggingface_example.sh --model $HF_PATH --quant fp8 --vlm
# gpt-oss QAT: --quant_cfg <CFG name> -> --recipe <recipe path>
accelerate launch --config_file configs/zero3.yaml sft.py \
--config configs/sft_full.yaml --model_name_or_path openai/gpt-oss-20b \
--recipe general/ptq/mxfp4_mlp_weight_only --output_dir gpt-oss-20b-qat
```
```python
# Layerwise calibration: bool / flat key -> nested LayerwiseConfig
quant_cfg["algorithm"] = {"method": "gptq", "layerwise": {"enable": True, "checkpoint_dir": "/path"}}
```
### Testing
- `tests/unit/recipe` (229 passed),
`tests/unit/torch/quantization/test_config_validation.py` (79 passed),
`tests/examples/hf_ptq/test_hf_ptq_args.py` (23 passed).
- Verified the two new recipes `model_dump` identical to the `mtq.*_CFG`
constants they replace.
- `ruff check modelopt/ examples/ tests/` clean; `ruff format --check`
clean on all changed Python files.
- GPU suites
(`tests/gpu/torch/export/test_unified_hf_export_and_check_safetensors.py`,
`test_accelerate_gpu.py`, `test_gptq.py`) had their preset / layerwise
literals updated but were not run locally — relying on CI.
- `examples/llm_qat/ARGUMENTS.md` is hand-edited to match what the
`generate-arguments-md` hook emits; the generator could not run locally
(missing `transformers` package metadata in this environment).
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ — that is the point of the PR:
it removes shims deprecated in 0.45/0.46. Callers must move to the
replacements in the table above. Additionally, a pre-0.45
`modelopt_state` carrying `use_sequential` or a top-level
`layerwise_checkpoint_dir` will now fail config validation.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — existing tests migrated to
the surviving APIs;
`TestLayerwiseNestedConfig::test_legacy_forms_rejected` pins that the
bool form, the `use_sequential` alias and the flat checkpoint-dir key
are all rejected. Tests covering the removed shims were deleted.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌
### Additional Information
Follow-up: the transformers 4.x drop deprecated in 0.46 is still
outstanding and will need its own PR.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## New Features
- Added MXFP4 and NVFP4 weight-only quantization recipes for MLP and MoE
layers.
- Added shared layer exclusions for more accurate effective-bits
calculations.
## Improvements
- Updated PTQ, QAT, GPT-OSS, deployment, and quantization-format
examples with current recipe names and configuration formats.
- Standardized layerwise settings under nested configuration fields.
## Breaking Changes
- Removed deprecated AutoQuantize options, `quant_cfg` usage, format
aliases, legacy layerwise settings, and compatibility example paths.
- Recipe-based and nested configuration forms are now required.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
186 lines
7.0 KiB
Bash
186 lines
7.0 KiB
Bash
#!/bin/bash
|
|
|
|
# SPDX-FileCopyrightText: Copyright (c) 2024 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
|
# SPDX-License-Identifier: Apache-2.0
|
|
#
|
|
# Licensed under the Apache License, Version 2.0 (the "License");
|
|
# you may not use this file except in compliance with the License.
|
|
# You may obtain a copy of the License at
|
|
#
|
|
# http://www.apache.org/licenses/LICENSE-2.0
|
|
#
|
|
# Unless required by applicable law or agreed to in writing, software
|
|
# distributed under the License is distributed on an "AS IS" BASIS,
|
|
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
# See the License for the specific language governing permissions and
|
|
# limitations under the License.
|
|
|
|
# Define a function to parse command-line options
|
|
parse_options() {
|
|
# Default values
|
|
MODEL_PATH=""
|
|
QFORMAT=""
|
|
RECIPE=""
|
|
KV_CACHE_QUANT=""
|
|
TP=1
|
|
PP=1
|
|
SPARSITY_FMT="dense"
|
|
LM_EVAL_TASKS="mmlu,gsm8k"
|
|
LM_EVAL_LIMIT=
|
|
SIMPLE_EVAL_TASKS="mmlu"
|
|
MMLU_LIMIT=
|
|
|
|
TASKS="quant"
|
|
|
|
TRUST_REMOTE_CODE=false
|
|
KV_CACHE_FREE_GPU_MEMORY_FRACTION=0.8
|
|
VERBOSE=true
|
|
USE_SEQ_DEVICE_MAP=false
|
|
CAST_MXFP4_TO_NVFP4=false
|
|
VLM=false
|
|
CALIB_WITH_IMAGES=false
|
|
|
|
# Parse command-line options
|
|
ARGS=$(getopt -o "" -l "model:,quant:,recipe:,kv_cache_quant:,tp:,pp:,sparsity:,awq_block_size:,calib:,calib_batch_size:,input:,output:,batch:,tasks:,lm_eval_tasks:,lm_eval_limit:,simple_eval_tasks:,simple_eval_limit:,mmlu_limit:,trust_remote_code,use_seq_device_map,gpu_max_mem_percentage:,kv_cache_free_gpu_memory_fraction:,low_memory_mode,no-verbose,calib_dataset:,calib_seq:,auto_quantize_checkpoint:,moe_calib_experts_ratio:,cast_mxfp4_to_nvfp4,vlm,calib_with_images" -n "$0" -- "$@")
|
|
|
|
eval set -- "$ARGS"
|
|
while true; do
|
|
case "$1" in
|
|
--model ) MODEL_PATH="$2"; shift 2;;
|
|
--quant ) QFORMAT="$2"; shift 2;;
|
|
--recipe ) RECIPE="$2"; shift 2;;
|
|
--kv_cache_quant ) KV_CACHE_QUANT="$2"; shift 2;;
|
|
--tp ) TP="$2"; shift 2;;
|
|
--pp ) PP="$2"; shift 2;;
|
|
--sparsity ) SPARSITY_FMT="$2"; shift 2;;
|
|
--awq_block_size ) AWQ_BLOCK_SIZE="$2"; shift 2;;
|
|
--calib ) CALIB_SIZE="$2"; shift 2;;
|
|
--calib_batch_size ) CALIB_BATCH_SIZE="$2"; shift 2;;
|
|
--input ) BUILD_MAX_INPUT_LEN="$2"; shift 2;;
|
|
--output ) BUILD_MAX_OUTPUT_LEN="$2"; shift 2;;
|
|
--batch ) BUILD_MAX_BATCH_SIZE="$2"; shift 2;;
|
|
--tasks ) TASKS="$2"; shift 2;;
|
|
--lm_eval_tasks ) LM_EVAL_TASKS="$2"; shift 2;;
|
|
--lm_eval_limit ) LM_EVAL_LIMIT="$2"; shift 2;;
|
|
--simple_eval_tasks ) SIMPLE_EVAL_TASKS="$2"; shift 2;;
|
|
--simple_eval_limit ) SIMPLE_EVAL_LIMIT="$2"; shift 2;;
|
|
--mmlu_limit ) MMLU_LIMIT="$2"; shift 2;;
|
|
--trust_remote_code ) TRUST_REMOTE_CODE=true; shift;;
|
|
--use_seq_device_map ) USE_SEQ_DEVICE_MAP=true; shift;;
|
|
--gpu_max_mem_percentage ) GPU_MAX_MEM_PERCENTAGE="$2"; shift 2;;
|
|
--kv_cache_free_gpu_memory_fraction ) KV_CACHE_FREE_GPU_MEMORY_FRACTION="$2"; shift 2;;
|
|
--no-verbose ) VERBOSE=false; shift;;
|
|
--low_memory_mode ) LOW_MEMORY_MODE=true; shift;;
|
|
--calib_dataset ) CALIB_DATASET="$2"; shift 2;;
|
|
--calib_seq ) CALIB_SEQ="$2"; shift 2;;
|
|
--auto_quantize_checkpoint ) AUTO_QUANTIZE_CHECKPOINT="$2"; shift 2;;
|
|
--moe_calib_experts_ratio ) MOE_CALIB_EXPERTS_RATIO="$2"; shift 2;;
|
|
--cast_mxfp4_to_nvfp4 ) CAST_MXFP4_TO_NVFP4=true; shift;;
|
|
--vlm ) VLM=true; shift;;
|
|
--calib_with_images ) CALIB_WITH_IMAGES=true; shift;;
|
|
-- ) shift; break ;;
|
|
* ) break ;;
|
|
esac
|
|
done
|
|
|
|
DEFAULT_CALIB_SIZE=512
|
|
DEFAULT_CALIB_SEQ=512
|
|
DEFAULT_CALIB_BATCH_SIZE=0
|
|
DEFAULT_BUILD_MAX_INPUT_LEN=4096
|
|
DEFAULT_BUILD_MAX_OUTPUT_LEN=1024
|
|
DEFAULT_BUILD_MAX_BATCH_SIZE=2
|
|
|
|
if [ -z "$CALIB_SIZE" ]; then
|
|
CALIB_SIZE=$DEFAULT_CALIB_SIZE
|
|
fi
|
|
if [ -z "$CALIB_SEQ" ]; then
|
|
CALIB_SEQ=$DEFAULT_CALIB_SEQ
|
|
fi
|
|
if [ -z "$CALIB_BATCH_SIZE" ]; then
|
|
CALIB_BATCH_SIZE=$DEFAULT_CALIB_BATCH_SIZE
|
|
fi
|
|
if [ -z "$BUILD_MAX_INPUT_LEN" ]; then
|
|
BUILD_MAX_INPUT_LEN=$DEFAULT_BUILD_MAX_INPUT_LEN
|
|
fi
|
|
if [ -z "$BUILD_MAX_OUTPUT_LEN" ]; then
|
|
BUILD_MAX_OUTPUT_LEN=$DEFAULT_BUILD_MAX_OUTPUT_LEN
|
|
fi
|
|
if [ -z "$BUILD_MAX_BATCH_SIZE" ]; then
|
|
BUILD_MAX_BATCH_SIZE=$DEFAULT_BUILD_MAX_BATCH_SIZE
|
|
fi
|
|
|
|
# Verify required options are provided
|
|
if [ -z "$MODEL_PATH" ] || [ -z "$TASKS" ] || ([ -z "$QFORMAT" ] && [ -z "$RECIPE" ]); then
|
|
echo "Usage: $0 --model=<MODEL_PATH> (--quant=<QFORMAT> | --recipe=<RECIPE>) --tasks=<TASK,...>"
|
|
echo "Optional args: --sparsity=<SPARSITY_FMT> --awq_block_size=<AWQ_BLOCK_SIZE> --calib=<CALIB_SIZE>"
|
|
exit 1
|
|
fi
|
|
|
|
# --quant and --recipe are mutually exclusive: --recipe is a full PTQ spec, while
|
|
# --quant selects a built-in qformat preset. Pick exactly one.
|
|
if [ -n "$QFORMAT" ] && [ -n "$RECIPE" ]; then
|
|
echo "Cannot specify both --quant and --recipe; pick one." >&2
|
|
exit 1
|
|
fi
|
|
|
|
VALID_TASKS=("quant" "mmlu" "lm_eval" "livecodebench" "simple_eval")
|
|
|
|
for task in $(echo "$TASKS" | tr ',' ' '); do
|
|
is_valid_task=false
|
|
for valid_task in "${VALID_TASKS[@]}"; do
|
|
if [[ "$valid_task" == "$task" ]]; then
|
|
is_valid_task=true
|
|
break
|
|
fi
|
|
done
|
|
if [ "$is_valid_task" = false ]; then
|
|
echo "task $task is not valid"
|
|
VALID_TASKS_STRING=$(IFS=','; echo "${VALID_TASKS[*]}")
|
|
echo "Allowed tasks are: $VALID_TASKS_STRING"
|
|
exit 1
|
|
fi
|
|
done
|
|
|
|
# Make sparsity and int4 quantization mutually exclusive as it does not brings speedup
|
|
if [[ "$SPARSITY_FMT" = "sparsegpt" || "$SPARSITY_FMT" = "sparse_magnitude" ]]; then
|
|
if [[ "$QFORMAT" == *"awq"* ]]; then
|
|
echo "Sparsity is not compatible with 'awq' quantization for TRT-LLM deployment."
|
|
exit 1 # Exit script with an error
|
|
fi
|
|
fi
|
|
|
|
# Now you can use the variables $GPU, $MODEL, and $TASKS in your script
|
|
echo "================="
|
|
echo "model: $MODEL_PATH"
|
|
echo "quant: $QFORMAT"
|
|
echo "recipe: $RECIPE"
|
|
echo "tp (TensorRT-LLM Checkpoint only): $TP"
|
|
echo "pp (TensorRT-LLM Checkpoint only): $PP"
|
|
echo "sparsity: $SPARSITY_FMT"
|
|
echo "awq_block_size: $AWQ_BLOCK_SIZE"
|
|
echo "calib: $CALIB_SIZE"
|
|
echo "calib_batch_size: $CALIB_BATCH_SIZE"
|
|
echo "input: $BUILD_MAX_INPUT_LEN"
|
|
echo "output: $BUILD_MAX_OUTPUT_LEN"
|
|
echo "batch: $BUILD_MAX_BATCH_SIZE"
|
|
echo "tasks: $TASKS"
|
|
echo "lm_eval_tasks: $LM_EVAL_TASKS"
|
|
echo "lm_eval_limit: $LM_EVAL_LIMIT"
|
|
echo "simple_eval_tasks: $SIMPLE_EVAL_TASKS"
|
|
echo "simple_eval_limit: $SIMPLE_EVAL_LIMIT"
|
|
echo "mmlu_limit: $MMLU_LIMIT"
|
|
echo "num_sample: $NUM_SAMPLES"
|
|
echo "use_seq_device_map: $USE_SEQ_DEVICE_MAP"
|
|
echo "gpu_max_mem_percentage: $GPU_MAX_MEM_PERCENTAGE"
|
|
echo "kv_cache_free_gpu_memory_fraction: $KV_CACHE_FREE_GPU_MEMORY_FRACTION"
|
|
echo "low_memory_mode: $LOW_MEMORY_MODE"
|
|
echo "calib_dataset: $CALIB_DATASET"
|
|
echo "calib_seq: $CALIB_SEQ"
|
|
echo "auto_quantize_checkpoint: $AUTO_QUANTIZE_CHECKPOINT"
|
|
echo "moe_calib_experts_ratio: $MOE_CALIB_EXPERTS_RATIO"
|
|
echo "cast_mxfp4_to_nvfp4: $CAST_MXFP4_TO_NVFP4"
|
|
echo "vlm: $VLM"
|
|
echo "calib_with_images: $CALIB_WITH_IMAGES"
|
|
echo "================="
|
|
}
|