Files
Wei-Ming Chen 1d392999b4 [OMNIML-5570] 2/2 Compose GEMM and KV-cache AutoQuant workflows (#2273)
### What does this PR do?

Type of change: new feature.

Follow-up to merged #2272. Adds composition of existing GEMM
quantization with KV-cache AutoQuantize:

- fixed FP8 GEMM PTQ followed by mixed-KV AutoQuantize;
- gradient-based NVFP4/FP8 GEMM AutoQuantize followed by independent
mixed-KV AutoQuantize;
- an optional `kv_auto_quantize` recipe stage with independent method,
constraints, candidates, and checkpoint path;
- ordered `hf_ptq.py` orchestration that keeps selected
weight/activation QDQ active while its calibration state remains frozen
during KV candidate calibration;
- fail-closed validation when a preceding stage leaves actual K/V
quantizers enabled; and
- unified export of a uniform-weight or mixed-weight checkpoint together
with the selected per-layer KV map.

The KV search still uses the public `mtq.auto_quantize(...,
constraints={"cost_model": "kv_cache", ...})` API from #2272. On a
converted model, the API preserves existing non-KV quantizers and
requires K/V to be disabled before search. Fresh-model behavior is
unchanged and starts from a deny-all quantizer baseline.

#### Why a follow-up field instead of a generic stage list?

This PR deliberately supports the two composition forms required by
`hf_ptq.py` without replacing the stable recipe schema. Existing recipes
already express a fixed `quantize` baseline plus one primary
`auto_quantize` search. A generic ordered `stages` list would require a
broader recipe/API migration, indexed checkpoint semantics, and
compatibility rules for arbitrary stage sequences. There is not yet a
demonstrated third search stage that justifies that surface-area change.

The two searches are not combined inside `mtq.auto_quantize`: each
invocation owns one search domain, constraint model, scoring method, and
resumable checkpoint. Their ordering and independent checkpoint paths
are orchestration concerns, while candidate calibration, scoring,
selection, and state application remain in the shared public API. A
general stage pipeline can be considered separately if more than this
one optional KV follow-up is needed.

Both solvers and scoring protocols are unchanged. The KV checkpoint
compatibility signature additionally fingerprints the preceding
quantizer configuration and calibrated state. Unsupported uniform-weight
plus mixed-KV exports record `kv_cache_deployment_supported: false` in
both ModelOpt and converted HF metadata.

### Usage

Fixed FP8 GEMM PTQ followed by KV AutoQuantize:

```bash
python examples/hf_ptq/hf_ptq.py \
  --pyt_ckpt_path Qwen/Qwen3-8B \
  --recipe general/auto_quantize/fp8_ptq_then_kv_fp8_nvfp4_cast_kl_div_at_5p4bits \
  --kv_auto_quantize_checkpoint /path/to/kv_autoquant.pth \
  --export_path /path/to/qwen3-8b-fp8-and-mixed-kv
```

Weight AutoQuantize followed by KV AutoQuantize:

```bash
python examples/hf_ptq/hf_ptq.py \
  --pyt_ckpt_path Qwen/Qwen3-8B \
  --recipe general/auto_quantize/nvfp4_fp8_gradient_then_kv_fp8_nvfp4_cast_kl_div_at_5p4bits \
  --auto_quantize_checkpoint /path/to/weight_autoquant.pth \
  --kv_auto_quantize_checkpoint /path/to/kv_autoquant.pth \
  --export_path /path/to/qwen3-8b-autoquant-and-mixed-kv
```

KV checkpoint resume requires identical preceding non-K/V quantizer
configuration and calibrated state. If rerunning the preceding stage
changes that state, use a new KV checkpoint path to recompute
sensitivities; configuration identity alone is insufficient to reuse the
scores safely.

### Testing

- Latest changed-area validation: 126 tests passed across `hf_ptq.py`
orchestration, KV checkpoint compatibility, export metadata, and HF
configuration conversion.
- A broader local run had 604 passes, one skip, and six failures: two
socket-binding failures under the sandbox and four local Transformers
API incompatibilities. This is not a full-suite pass.
- The fixed-PTQ→KV recipe executes end to end on a tiny offline Qwen
fixture.
- Public API coverage verifies that composed KV search preserves
preceding weight quantization and rejects enabled K/V state.
- Changed-file pre-commit hooks passed; the isolated recipe validator
also passed after dependency bootstrap.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅ (0.48.0 composition feature and KV
checkpoint flag deprecation)
- Did you get Claude approval on this PR?: ❌

### Additional Information

- This follow-up targets `main`, which contains merged #2272.
- `--auto_quantize_checkpoint` and `--kv_auto_quantize_checkpoint` are
intentionally separate because KV sensitivities depend on the preceding
GEMM state.
- Uniform-weight plus mixed-KV exports are for artifact inspection until
the runtime's uniform-weight ModelOpt configuration consumes
`kv_cache_quantized_layers`. Export emits an actionable warning and
records `kv_cache_deployment_supported: false`; this marker does not
itself add runtime support.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added staged post-training quantization workflows for weights and KV
caches, including dedicated KV-cache checkpoints.
- Added FP8/NVFP4 recipes with configurable bit constraints and scoring.
  - KV-cache quantization now supports pre-quantized models.

- **Bug Fixes**
- Mixed weight and KV-cache quantization now exports with a warning
instead of failing.
- Improved validation and checkpoint compatibility for staged
configurations.
- Added safeguards for configurations without enabled weight quantizers.

- **Documentation**
- Clarified staged KV-cache workflows, checkpoint options, configuration
behavior, and unsupported deployment combinations.
- Documented deprecated legacy quantization options and their
replacement behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-09-29 00:15:40 +00:00

1596 lines
62 KiB
Python
Executable File

# SPDX-FileCopyrightText: Copyright (c) 2024 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import argparse
import os
import random
import time
import warnings
from pathlib import Path
from typing import Any
import numpy as np
import torch
from accelerate.hooks import remove_hook_from_module
from autoquant_utils import (
_FSDP2_KV_AUTOQUANT_ERROR,
_quantize_config_explicitly_enables_kv,
_recipe_is_auto_quantize,
_recipe_is_kv_auto_quantize,
auto_quantize,
)
from cast_mxfp4_to_nvfp4 import apply_to_model as apply_cast_mxfp4_to_nvfp4
from example_utils import (
HF_PTQ,
_prepare_quant_cfg,
_resolve_model_path,
build_quant_cfg,
cleanup_distributed,
copy_custom_model_files,
create_vlm_calibration_loop,
get_model,
get_processor,
get_tokenizer,
is_enc_dec,
is_nemotron_vl,
mlflow_run,
recipe_layerwise_blocks,
run_nemotron_vl_preview,
save_processor_config,
save_source_config,
setup_distributed_args,
validate_fsdp2_supported,
)
from torch.utils.data import DataLoader
from transformers import (
AutoConfig,
AutoModelForCausalLM,
AutoProcessor,
PreTrainedTokenizer,
PreTrainedTokenizerBase,
PreTrainedTokenizerFast,
ProcessorMixin,
WhisperProcessor,
)
import modelopt.torch.opt as mto
import modelopt.torch.quantization as mtq
import modelopt.torch.sparsity as mts
from modelopt.recipe import ModelOptAutoQuantizeRecipe, ModelOptPTQRecipe, load_recipe
from modelopt.recipe.presets import (
KV_CACHE_NONE,
KV_QUANT_CFG_CHOICES,
QUANT_CFG_CHOICES,
RecipeSupersededAction,
)
from modelopt.torch.export import (
export_hf_checkpoint,
export_hf_vllm_fq_checkpoint,
export_speculative_decoding,
get_model_type,
has_spec_opt,
save_expert_token_count_table,
)
from modelopt.torch.export.layerwise_export import LayerwiseExporter
from modelopt.torch.export.model_utils import get_language_model_from_vl, is_multimodal_model
from modelopt.torch.export.trtllm import export_tensorrt_llm_checkpoint
from modelopt.torch.quantization.config import need_calibration
from modelopt.torch.quantization.plugins.accelerate import init_quantized_weights
from modelopt.torch.quantization.utils import is_quantized
from modelopt.torch.speculative.eagle.utils import (
EagleOfflineDataCollator,
OfflineSupervisedDataset,
)
from modelopt.torch.utils import print_rank_0
from modelopt.torch.utils.dataset_utils import (
create_forward_loop,
get_dataset_dataloader,
get_max_batch_size,
get_supported_datasets,
)
from modelopt.torch.utils.memory_monitor import launch_memory_monitor
from modelopt.torch.utils.mlflow import add_mlflow_args, resolve_mlflow_args
from modelopt.torch.utils.plugins.model_load_utils import parallel_load_and_prepare_fsdp2
from modelopt.torch.utils.speech_dataset_utils import get_speech_dataset_dataloader
from modelopt.torch.utils.vlm_dataset_utils import get_vlm_dataset_dataloader
RAND_SEED = 1234
mto.enable_huggingface_checkpointing()
def extract_and_prepare_language_model_from_vl(full_model):
"""Extract language model from VL model and disable quantization for non-language components.
Args:
full_model: The full VLM model
Returns:
tuple: (language_model, model_type) or (None, None) if not a VLM
"""
language_model_lineage = get_language_model_from_vl(full_model, strict=True)
if language_model_lineage is not None:
language_model = language_model_lineage.pop(-1)
ancestors = language_model_lineage
# Apply disabled quant to all modules that are not part of language_model
# This excludes them during HF export
disabled_quant_cfg = {
"quant_cfg": [{"quantizer_name": "*", "enable": False}],
"algorithm": "max",
}
memo = set(ancestors) | {language_model}
for ancestor in ancestors:
for _, module in ancestor.named_children():
if module not in memo:
mtq.quantize(module, disabled_quant_cfg, forward_loop=None)
memo.add(module)
model_type = get_model_type(language_model)
return language_model, model_type
return None, None
class _DeviceDataLoader:
"""Wrapper around a DataLoader that moves each batch to a target device."""
def __init__(self, dataloader: DataLoader, device: torch.device):
self.dataloader = dataloader
self.device = device
def __iter__(self):
for batch in self.dataloader:
yield _move_batch_to_device(batch, self.device)
def __len__(self):
return len(self.dataloader)
def _move_batch_to_device(batch: dict, device: torch.device) -> dict:
"""Recursively move all tensors in a batch dict to the given device."""
def _to_device(value):
if isinstance(value, torch.Tensor):
return value.to(device)
if isinstance(value, dict):
return {k: _to_device(v) for k, v in value.items()}
return value
return {k: _to_device(v) for k, v in batch.items()}
def make_calib_dataloader(
args: argparse.Namespace,
language_model: torch.nn.Module,
processor: ProcessorMixin | None,
tokenizer: PreTrainedTokenizerBase | None,
device: torch.device,
model_type: str | None,
autoquant_gradient_recipe: bool = False,
) -> tuple[DataLoader | _DeviceDataLoader, str | None]:
calib_dataloader = None
first_text_speech_dataset = None
if args.specdec_offline_dataset is not None:
offline_data_path = Path(args.specdec_offline_dataset)
dumped_files = sorted(str(p) for p in offline_data_path.glob("*.pt"))
if not dumped_files:
raise ValueError(f"No .pt files found in {args.specdec_offline_dataset}")
if args.calib_size[0] > 0:
dumped_files = dumped_files[: args.calib_size[0]]
dataset = OfflineSupervisedDataset(dumped_files)
collator = EagleOfflineDataCollator(train_len=args.calib_seq)
raw_loader = DataLoader(
dataset,
batch_size=args.batch_size,
shuffle=False,
collate_fn=collator,
)
# Wrap to move batches to the target device; device-transfer logic is kept
# out of the data collator to avoid interference with dataloader prefetching.
calib_dataloader = _DeviceDataLoader(raw_loader, device)
elif args.calib_with_images:
# VLM image-text calibration path: assume Nemotron VLM dataset by default.
assert processor is not None, (
"Please provide a processor (e.g., AutoProcessor) for image calibration."
)
assert len(args.calib_size) == 1, (
"Image calibration currently supports a single dataset. "
"Please pass --calib_size with one value (e.g., --calib_size 256)."
)
calib_dataloader = get_vlm_dataset_dataloader(
dataset_name="nemotron_vlm_dataset_v2",
processor=processor,
batch_size=args.batch_size,
num_samples=args.calib_size[0],
device=device,
require_image=True,
subsets=["sparsetables", "plotqa_cot", "wiki_en"],
shuffle_buffer_size=10_000,
seed=42,
use_media_shards=True,
max_shards=1,
)
elif model_type == "whisper":
assert processor is not None and isinstance(processor, WhisperProcessor), (
"The AutoProcessor must be set."
)
assert len(args.calib_size) == 1, (
"whisper only supports one dataset for calibration, can extend this in the future"
)
calib_dataloader, first_text_speech_dataset = get_speech_dataset_dataloader(
dataset_name=args.dataset[0] if args.dataset else "peoples_speech",
processor=processor,
batch_size=args.batch_size,
num_samples=args.calib_size[0],
device=device,
dtype=language_model.dtype,
)
else:
assert tokenizer is not None and isinstance(
tokenizer, PreTrainedTokenizer | PreTrainedTokenizerFast
), "The PreTrainedTokenizer must be set"
# Labels are only needed for gradient-based auto_quantize
include_labels = autoquant_gradient_recipe
calib_dataloader = get_dataset_dataloader(
dataset_name=args.dataset,
tokenizer=tokenizer,
batch_size=args.batch_size,
num_samples=args.calib_size,
max_sample_length=args.calib_seq,
device=device,
include_labels=include_labels,
distributed=args.use_fsdp2,
sampler_kwargs=(
{"num_replicas": args.dist_state.world_size, "rank": args.dist_state.rank}
if args.use_fsdp2
else None
),
)
return calib_dataloader, first_text_speech_dataset
def _resolve_kv_auto_quantize_checkpoint(args: argparse.Namespace) -> str | None:
"""Resolve a KV-primary checkpoint with a one-release legacy fallback."""
if args.kv_auto_quantize_checkpoint is not None:
return args.kv_auto_quantize_checkpoint
if args.auto_quantize_checkpoint is not None:
warnings.warn(
"Using --auto_quantize_checkpoint for a KV-cache search is deprecated; use "
"--kv_auto_quantize_checkpoint instead.",
FutureWarning,
)
return args.auto_quantize_checkpoint
return None
def _validate_recipe_calibration(args: argparse.Namespace, recipe) -> None:
"""Require image-text calibration when a PTQ recipe enables visual input quantizers."""
if not isinstance(recipe, ModelOptPTQRecipe) or args.calib_with_images:
return
visual_input_entry = next(
(
entry
for entry in reversed(recipe.quantize.quant_cfg)
if entry.quantizer_name == "*visual.*input_quantizer"
),
None,
)
if visual_input_entry is not None and visual_input_entry.enable:
raise ValueError(
"Vision Encoder quantization recipes require --calib_with_images so visual input "
"quantizers receive activation calibration data."
)
def load_model(args: argparse.Namespace):
# If low memory mode is enabled, we compress the model while loading the HF checkpoint.
calibration_only = False
if args.use_fsdp2 and _recipe_is_kv_auto_quantize(args.recipe):
raise NotImplementedError(_FSDP2_KV_AUTOQUANT_ERROR)
if args.use_fsdp2:
hf_config = AutoConfig.from_pretrained(
args.pyt_ckpt_path, trust_remote_code=args.trust_remote_code
)
validate_fsdp2_supported(args, hf_config)
full_model = parallel_load_and_prepare_fsdp2(
args.pyt_ckpt_path,
args.dist_state.device,
args.dist_state.rank,
args.dist_state.world_size,
trust_remote_code=args.trust_remote_code,
cpu_offload=args.cpu_offload,
attn_implementation=args.attn_implementation,
hf_config=hf_config,
)
elif args.specdec_offline_dataset is not None or not args.low_memory_mode:
full_model = get_model(
args.pyt_ckpt_path,
args.dist_state.device,
gpu_mem_percentage=args.gpu_max_mem_percentage,
trust_remote_code=args.trust_remote_code,
use_seq_device_map=args.use_seq_device_map,
attn_implementation=args.attn_implementation,
offload_folder=args.offload_folder,
max_cpu_memory_gb=args.max_cpu_memory_gb,
max_gpu_memory_gb=args.max_gpu_memory_gb,
)
else:
assert args.qformat in QUANT_CFG_CHOICES, (
f"Quantization format is not supported for low memory mode. Supported formats: {list(QUANT_CFG_CHOICES)}"
)
quant_cfg = QUANT_CFG_CHOICES[args.qformat]
if args.kv_cache_qformat != KV_CACHE_NONE:
quant_cfg = mtq.utils.update_quant_cfg_with_kv_cache_quant(
quant_cfg,
KV_QUANT_CFG_CHOICES[args.kv_cache_qformat]["quant_cfg"],
)
# Do not use real quant GEMM so the calibration can be more accurate.
with init_quantized_weights(
quant_cfg, gpu_mem_percentage=args.gpu_max_mem_percentage, quant_gemm=False
):
model_kwargs = {"trust_remote_code": args.trust_remote_code}
if args.attn_implementation is not None:
model_kwargs["attn_implementation"] = args.attn_implementation
full_model = AutoModelForCausalLM.from_pretrained(
args.pyt_ckpt_path,
**model_kwargs,
)
calibration_only = True
model_type = get_model_type(full_model)
if args.use_fsdp2:
device = args.dist_state.device
else:
device = full_model.device
if hasattr(full_model, "model"):
device = full_model.model.device
processor = None
tokenizer = None
language_model = full_model
default_padding_side = None
default_pad_token = None
is_nemotron_vl_model = is_nemotron_vl(full_model)
# Default to image-text calibration for VLM models. Skip for the AutoQuantize recipe path, whose
# text-only path does not support image-text calibration yet (auto_quantize() would raise);
# auto-enabling it here would make Nemotron-VL AutoQuantize fail unconditionally.
if (
is_nemotron_vl_model
and not args.calib_with_images
and not _recipe_is_auto_quantize(args.recipe)
):
print("Nemotron VL model detected. Enabling image-text calibration by default.")
args.calib_with_images = True
if model_type == "whisper":
processor = get_processor(
args.pyt_ckpt_path,
model_type,
trust_remote_code=args.trust_remote_code,
)
elif args.calib_with_images:
# For VLM image calibration, we need an AutoProcessor to build multimodal inputs.
processor = AutoProcessor.from_pretrained(
args.pyt_ckpt_path,
trust_remote_code=args.trust_remote_code,
padding_side="left",
)
if hasattr(processor, "tokenizer") and processor.tokenizer is not None:
tokenizer = processor.tokenizer
else:
tokenizer = get_tokenizer(args.pyt_ckpt_path, trust_remote_code=args.trust_remote_code)
default_pad_token = tokenizer.pad_token
# Some Nemotron tokenizers may not define pad_token by default; but we use padding=True during calibration.
if tokenizer.pad_token is None:
tokenizer.pad_token = tokenizer.eos_token
assert tokenizer.pad_token is not None, f"Pad token for {args.pyt_ckpt_path} cannot be set!"
default_padding_side = tokenizer.padding_side
tokenizer.padding_side = "left"
# Plain PTQ quantizes only the language model. Recipes keep the complete VLM so their
# quantizer rules can target vision and language components in one state.
if args.recipe is None:
extracted_lm, extracted_model_type = extract_and_prepare_language_model_from_vl(
full_model
)
if extracted_lm is not None:
language_model = extracted_lm
model_type = extracted_model_type
else:
if args.specdec_offline_dataset is not None:
language_model = full_model
else:
if args.dataset is None:
args.dataset = ["cnn_nemotron_v2_mix"]
warnings.warn(
"No dataset specified. Defaulting to the 'cnn_nemotron_v2_mix' combo "
"(cnn_dailymail + nemotron-post-training-dataset-v2)."
)
# Adjust calib_size to match dataset length by extending or truncating as needed
args.calib_size = (args.calib_size + [args.calib_size[-1]] * len(args.dataset))[
: len(args.dataset)
]
# Plain PTQ quantizes only the extracted language model. The recipe path keeps the outer
# CausalLM so recipes / search can see the Qwen3.5/3.6-MoE VLM lm_head; extracting here
# would leave modelopt state on the ancestors and make auto_quantize() fail with
# "multiple modelopt states".
if args.recipe is None:
extracted_lm, extracted_model_type = extract_and_prepare_language_model_from_vl(
full_model
)
if extracted_lm is not None:
language_model = extracted_lm
model_type = extracted_model_type
tokenizer = get_tokenizer(args.pyt_ckpt_path, trust_remote_code=args.trust_remote_code)
default_padding_side = tokenizer.padding_side
default_pad_token = tokenizer.pad_token
# Left padding usually provides better calibration result.
tokenizer.padding_side = "left"
return (
full_model,
language_model,
model_type,
calibration_only,
processor,
tokenizer,
default_padding_side,
default_pad_token,
device,
)
def sparsity_main(
args: argparse.Namespace,
full_model: torch.nn.Module,
tokenizer: PreTrainedTokenizerBase | None,
device: torch.device,
):
if args.batch_size == 0:
# Sparse algorithm takes more GPU memory so we reduce the batch_size by 4.
args.batch_size = max(get_max_batch_size(full_model) // 4, 1)
args.batch_size = min(args.batch_size, sum(args.calib_size))
print(f"Use calib batch_size {args.batch_size}")
# Different calibration datasets are also available, e.g., "pile" and "wikipedia"
# Please also check the docstring for the datasets available
assert tokenizer is not None and isinstance(
tokenizer, PreTrainedTokenizer | PreTrainedTokenizerFast
), "The PreTrainedTokenizer must be set"
calib_dataloader = get_dataset_dataloader(
dataset_name=args.dataset,
tokenizer=tokenizer,
batch_size=args.batch_size,
num_samples=args.calib_size,
max_sample_length=args.calib_seq,
device=device,
)
full_model = mts.sparsify(
full_model,
args.sparsity_fmt,
config={"data_loader": calib_dataloader, "collect_func": lambda x: x},
)
mts.export(full_model)
def mono_quantize(
args: argparse.Namespace,
quant_cfg: dict[str, Any],
full_model: torch.nn.Module,
language_model: torch.nn.Module,
model_type: str | None,
calibration_only: bool,
calib_dataloader: DataLoader,
is_nemotron_vl_model: bool,
):
"""Plain quantization of the selected model target to one quantization configuration."""
model_is_already_quantized = is_quantized(language_model)
if "awq" in args.qformat:
print(
"\n####\nAWQ calibration could take longer than other calibration methods. "
"Consider reducing calib_size to reduce calibration time.\n####\n"
)
if not model_is_already_quantized or calibration_only:
# quantize the model
use_calibration = need_calibration(quant_cfg)
if not use_calibration:
warnings.warn("Dynamic quantization. Calibration skipped.")
calibrate_loop = None
if use_calibration:
# Image calibration batches contain multimodal kwargs (for example pixel_values).
# They must be consumed by the complete VLM even when only a nested component is the
# quantization target; the full forward still exercises that component's quantizers.
if args.calib_with_images:
calibrate_loop = create_vlm_calibration_loop(full_model, calib_dataloader)
else:
calibrate_loop = create_forward_loop(
dataloader=calib_dataloader,
allowed_non_tensor_keys={"base_model_outputs"}
if args.specdec_offline_dataset is not None
else None,
)
if args.layerwise_export:
LayerwiseExporter(full_model, args.export_path)
if calibration_only:
language_model = mtq.calibrate(
language_model, quant_cfg["algorithm"], forward_loop=calibrate_loop
)
else:
language_model = mtq.quantize(language_model, quant_cfg, forward_loop=calibrate_loop)
# For VL models, update full_model to use the quantized language model
if is_nemotron_vl_model and language_model is not full_model:
language_model_lineage = get_language_model_from_vl(full_model)
if language_model_lineage is not None:
print("Updating full_model with quantized language_model...")
language_model_lineage[-2].language_model = language_model
else:
warnings.warn("Skipping quantization: model is already quantized.")
def _run_auto_quantize_recipe(
args: argparse.Namespace,
recipe: ModelOptAutoQuantizeRecipe,
full_model: torch.nn.Module,
language_model: torch.nn.Module,
model_type: str | None,
calibration_only: bool,
calib_dataloader: DataLoader,
is_nemotron_vl_model: bool,
) -> None:
"""Run the recipe's fixed PTQ, weight search, and KV search in order."""
primary = recipe.auto_quantize
followup_kv = recipe.kv_auto_quantize
primary_is_kv = primary.constraints.cost_model == "kv_cache"
fixed_quantize_config = recipe.quantize
if fixed_quantize_config is not None and (primary_is_kv or followup_kv is not None):
if _quantize_config_explicitly_enables_kv(fixed_quantize_config.model_dump()):
raise ValueError(
"The fixed quantize stage explicitly enables K/V quantizers before KV-cache "
"AutoQuantize. Disable them in the fixed stage."
)
if primary_is_kv and fixed_quantize_config is not None:
quant_cfg = _prepare_quant_cfg(args, fixed_quantize_config.model_dump(), full_model)
mono_quantize(
args,
quant_cfg,
full_model,
language_model,
model_type,
calibration_only,
calib_dataloader,
is_nemotron_vl_model,
)
fixed_quantize_config = None
auto_quantize(
args,
full_model,
calib_dataloader,
aq_config=primary,
full_model=full_model,
fixed_quantize_config=fixed_quantize_config,
allow_uniform_kv=followup_kv is None,
checkpoint=(
_resolve_kv_auto_quantize_checkpoint(args)
if primary_is_kv
else args.auto_quantize_checkpoint
),
)
if followup_kv is not None:
auto_quantize(
args,
full_model,
calib_dataloader,
aq_config=followup_kv,
full_model=full_model,
allow_uniform_kv=False,
# The weight search owns --auto_quantize_checkpoint, so a follow-up KV search must
# never use the KV-primary legacy fallback and collide with the weight state.
checkpoint=args.kv_auto_quantize_checkpoint,
)
def export_quantized(
args: argparse.Namespace,
full_model: torch.nn.Module,
language_model: torch.nn.Module,
model_type: str | None,
tokenizer: PreTrainedTokenizerBase | None,
default_padding_side,
default_pad_token,
):
# Not inference_mode: the FSDP2 path gathers full params in this context and
# inference tensors break the subsequent state_dict() -> param.detach().
with torch.no_grad():
if model_type is None:
print(f"Unknown model type {type(language_model).__name__}. Continue exporting...")
model_type = f"unknown:{type(language_model).__name__}"
export_path = args.export_path
# Early exit for speculative decoding checkpoints
# No tokenizer saving needed for spec ckpts
if has_spec_opt(full_model):
export_speculative_decoding(full_model, export_dir=export_path)
args.checkpoint_exported = True
print(f"Quantized speculative decoding checkpoint exported to: {export_path}")
return
if is_multimodal_model(full_model):
# Per-layer export writes its own config.json with quantization_config, which
# the source config would replace; it never writes a processor config.
if not args.layerwise_export:
save_source_config(args, export_path)
save_processor_config(args, export_path)
start_time = time.time()
is_tensorrt_llm_export = (
model_type in ["t5", "bart", "whisper"]
or args.sparsity_fmt != "dense"
or "int8_smoothquant" in args.qformat
)
if is_tensorrt_llm_export:
if (
args.inference_tensor_parallel != 1 or args.inference_pipeline_parallel != 1
) and args.qformat == "nvfp4_svdquant":
raise NotImplementedError("Svdquant does not support multiple GPUs yet.")
warnings.warn(
"Still exporting TensorRT-LLM checkpoints for models not supported by the TensorRT-LLM torch runtime."
)
# Move meta tensor back to device before exporting.
remove_hook_from_module(language_model, recurse=True)
export_tensorrt_llm_checkpoint(
language_model,
model_type,
export_dir=export_path,
inference_tensor_parallel=args.inference_tensor_parallel,
inference_pipeline_parallel=args.inference_pipeline_parallel,
)
# Copy custom model files (Python files and JSON configs) for TensorRT-LLM export
# TRT-LLM checkpoints are rank<N>.safetensors plus their own config; nothing
# there reads an off-index sidecar, and the exclude_modules seeding that gives
# one meaning happens only inside export_hf_checkpoint.
copy_custom_model_files(
args.pyt_ckpt_path,
export_path,
args.trust_remote_code,
copy_off_index_weights=False,
)
else:
# Check arguments for unified_hf export format and set to default if unsupported arguments are provided
assert args.sparsity_fmt == "dense", (
f"Sparsity format {args.sparsity_fmt} not supported by unified export api."
)
if args.inference_tensor_parallel != 1 or args.inference_pipeline_parallel != 1:
warnings.warn(
"Unified HF export format does not specify inference tensor parallel or pipeline parallel. "
"They will be set at deployment time."
)
# Load any missing weights from non-standard safetensors (handled in get_model for non-low-memory mode)
# Store the MTP layer prefixes on the model for later exclusion from quantization
if args.vllm_fakequant_export:
# save_pretrained inside the exporter writes model-backed state only; weights
# the loader could not place (an MTP head, an auxiliary tower) are carried over
# as an extra shard afterward -- see _carry_over_unplaced_weights.
export_hf_vllm_fq_checkpoint(
full_model, export_dir=export_path, inplace_mem_efficient=True
)
else:
# Weights the loader could not place (an MTP head, an auxiliary tower) are
# carried over by the exporter from the keys recorded at load time; nothing
# architecture-specific is needed here.
export_hf_checkpoint(
full_model,
export_dir=export_path,
)
if args.qformat == "w4a16_nvfp4":
warnings.warn(
"TensorRT-LLM and SGLang do not support this format. "
"vLLM deployment support is in progress."
)
# Restore default padding and export the tokenizer as well.
if tokenizer is not None:
tokenizer.padding_side = default_padding_side
if default_pad_token is not None:
tokenizer.pad_token = default_pad_token
if args.dist_state.is_main:
tokenizer.save_pretrained(export_path)
# Copy custom model files (Python files and JSON configs) if trust_remote_code is used.
# This must run AFTER tokenizer.save_pretrained() so original tokenizer files
# from the source checkpoint take precedence over regenerated ones (which may
# differ in format due to newer transformers versions).
if args.dist_state.is_main:
exclude_files = None if is_tensorrt_llm_export else {"generation_config.json"}
copy_custom_model_files(
args.pyt_ckpt_path,
export_path,
args.trust_remote_code,
exclude_files=exclude_files,
copy_off_index_weights=not is_tensorrt_llm_export,
)
args.checkpoint_exported = True
end_time = time.time()
print_rank_0(
f"Quantized model exported to: {export_path}. Total time used {end_time - start_time}s"
)
def pre_quantize(
args: argparse.Namespace,
full_model: torch.nn.Module,
model_type: str | None,
tokenizer: PreTrainedTokenizerBase | None,
calib_dataloader: DataLoader | None,
is_nemotron_vl_model: bool,
):
"""
Processing before the quantization.
Currently we run one round of generation for a sample prompt, to be compared with
post-quantize generation.
"""
# Offline specdec models skip pre-quantize preview (no tokenizer or standard dataloader)
if args.specdec_offline_dataset is not None:
return None, None, None
# Only run single sample for preview
assert calib_dataloader is not None, "calib_dataloader is required for pre-quantize preview"
batch = next(iter(calib_dataloader))
input_key = "input_features" if model_type == "whisper" else "input_ids"
preview_input_ids = batch[input_key][0:1]
# Pass attention_mask to generate(): HF cannot infer it when pad_token == eos_token.
preview_attention_mask = batch["attention_mask"][0:1] if "attention_mask" in batch else None
# Generate preview before quantization
if args.skip_generate:
generated_ids_before_ptq = None
elif model_type == "deepseek":
# DeepSeek generation may go OOM, so we skip it
generated_ids_before_ptq = None
elif model_type == "nemotron_h":
# NemotronH (SSM/Mamba hybrid) modeling code does not work with accelerate's big model inference
# when multiple GPUs are used. So we skip generation for NemotronH models. The issue presents in
# the remote code and also in transformers library integration code from v5.3
generated_ids_before_ptq = None
elif is_nemotron_vl_model and tokenizer is not None:
generated_ids_before_ptq = run_nemotron_vl_preview(
full_model,
tokenizer,
preview_input_ids,
args.pyt_ckpt_path,
"before quantization",
allow_fallback=False,
trust_remote_code=args.trust_remote_code,
)
else:
generated_ids_before_ptq = full_model.generate(
preview_input_ids,
attention_mask=preview_attention_mask,
max_new_tokens=100,
)
return preview_input_ids, preview_attention_mask, generated_ids_before_ptq
def post_quantize(
args: argparse.Namespace,
full_model: torch.nn.Module,
language_model: torch.nn.Module,
model_type: str | None,
tokenizer: PreTrainedTokenizerBase | None,
processor: ProcessorMixin | None,
preview_input_ids,
preview_attention_mask,
generated_ids_before_ptq,
is_nemotron_vl_model,
first_text_speech_dataset,
default_padding_side,
default_pad_token,
calib_dataloader: DataLoader,
):
"""
Processing after the quantization, then export.
For offline speculative decoding models, skip generation comparison and proceed
directly to export. For standard models, run one round of generation using the
quantized model for a sample prompt and compare it with pre-quantize generation.
"""
# Early exit for offline speculative decoding: skip generation comparison and export directly.
# The model's get_dummy_inputs() provides the right input format for the export forward pass.
if args.specdec_offline_dataset is not None:
export_quantized(
args,
full_model,
language_model,
model_type,
tokenizer,
default_padding_side,
default_pad_token,
)
return
if args.verbose and args.dist_state.is_main:
try:
mtq.print_quant_summary(full_model, args.export_path)
save_expert_token_count_table(full_model, args.export_path)
except Exception as e:
print(f"Error saving quant summary: {e}")
print("Continuing with generation...")
# Run some samples
torch.cuda.empty_cache()
generated_ids_after_ptq = None
if generated_ids_before_ptq is None:
pass
elif model_type != "llama4" and not is_nemotron_vl_model:
# Our fake quantizer may not be fully compatible with torch.compile.
# This is a best-effort sanity check: e.g. a `device_map="auto"` load that offloads
# part of the model to CPU (seen on unified-memory single-GPU hosts) can make a
# quantized layer run on CPU, which some kernels (e.g. NVFP4 dynamic block
# quantization) don't support. Don't let that discard the completed calibration.
try:
generated_ids_after_ptq = full_model.generate(
preview_input_ids,
attention_mask=preview_attention_mask,
max_new_tokens=100,
)
except Exception as e:
warnings.warn(f"Post-quantization generation sanity check failed, skipping it: {e}")
elif is_nemotron_vl_model and tokenizer is not None:
generated_ids_after_ptq = run_nemotron_vl_preview(
full_model,
tokenizer,
preview_input_ids,
args.pyt_ckpt_path,
"after quantization",
allow_fallback=False,
trust_remote_code=args.trust_remote_code,
)
else:
warnings.warn(
"Llama4 Maverick generation after quantization has a bug. Skipping generation sample."
)
def input_decode(input_ids):
if processor is not None and isinstance(processor, WhisperProcessor):
return first_text_speech_dataset
elif tokenizer is not None:
return tokenizer.batch_decode(input_ids)
else:
raise ValueError("The processor or tokenizer must be set")
def output_decode(generated_ids, input_shape):
# Some `.generate()` returns a ModelOutput dataclass (e.g. DiffusionGemma);
# unwrap to the token tensor so downstream slicing works uniformly.
if hasattr(generated_ids, "sequences"):
generated_ids = generated_ids.sequences
if is_enc_dec(model_type):
if processor is not None and isinstance(processor, WhisperProcessor):
return processor.tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
elif tokenizer is not None:
return tokenizer.batch_decode(generated_ids, skip_special_tokens=True)
elif tokenizer is not None:
return tokenizer.batch_decode(generated_ids[:, input_shape:])
else:
raise ValueError("The processor or tokenizer must be set")
if generated_ids_after_ptq is not None:
print("--------")
if is_nemotron_vl_model:
# For Nemotron VL models, generated_ids are text strings from model.chat()
print("Nemotron VL model text-only generation results:")
print(f"Text response before quantization: {generated_ids_before_ptq}")
print("--------")
print(f"Text response after quantization: {generated_ids_after_ptq}")
print("--------")
print("Note: Additional VL tests with images were run separately above")
else:
# For regular LLMs, generated_ids are token tensors that need decoding
print(f"example test input: {input_decode(preview_input_ids)}")
print("--------")
print(
f"example outputs before ptq: {output_decode(generated_ids_before_ptq, preview_input_ids.shape[1])}"
)
print("--------")
print(
f"example outputs after ptq: {output_decode(generated_ids_after_ptq, preview_input_ids.shape[1])}"
)
export_quantized(
args,
full_model,
language_model,
model_type,
tokenizer,
default_padding_side,
default_pad_token,
)
def quantize_main(
args: argparse.Namespace,
full_model: torch.nn.Module,
language_model: torch.nn.Module,
model_type: str | None,
calibration_only: bool,
processor: ProcessorMixin | None,
tokenizer: PreTrainedTokenizerBase | None,
default_padding_side,
default_pad_token,
device: torch.device,
):
# Load the recipe up front so we can detect layerwise calibration before batch-size probing.
recipe = None
if args.recipe is not None:
print(f"Use recipe {args.recipe} for quantization")
recipe = load_recipe(args.recipe)
if not isinstance(recipe, ModelOptPTQRecipe | ModelOptAutoQuantizeRecipe):
raise TypeError(
f"Expected PTQ or AutoQuantize recipe, but got {type(recipe).__name__} "
f"from {args.recipe}"
)
_validate_recipe_calibration(args, recipe)
# AutoQuantize is recipe-driven: everything downstream reads the resolved AutoQuantizeConfig.
if isinstance(recipe, ModelOptAutoQuantizeRecipe):
aq_config = recipe.auto_quantize
else:
aq_config = None
layerwise_cfgs = recipe_layerwise_blocks(recipe)
is_layerwise = any(cfg.get("enable", False) for cfg in layerwise_cfgs)
# The value is a placeholder, replaced with --export_path below; presence is the switch.
args.layerwise_export = any(
cfg.get("export_dir") is not None and cfg.get("enable", False) for cfg in layerwise_cfgs
)
if not args.layerwise_export and any(
cfg.get("export_dir") is not None for cfg in layerwise_cfgs
):
warnings.warn(
"layerwise.export_dir is set but layerwise.enable is not, so there is no "
"per-layer pass to write the shards: the whole-model export runs instead, which "
"holds the full state dict in host memory."
)
if args.layerwise_export:
if isinstance(recipe, ModelOptAutoQuantizeRecipe):
# Only the mono-quantize path retargets export_dir and runs the refusals;
# auto_quantize would export to the placeholder and skip the real export.
raise NotImplementedError(
"layerwise.export_dir is not supported with an AutoQuantize recipe; "
"use a PTQ recipe, or drop export_dir and export afterwards."
)
if not args.skip_generate:
print("Layerwise export: forcing --skip_generate, the model is left in export form.")
args.skip_generate = True
if args.batch_size == 0:
# For VL models with image-text calibration, skip automatic batch size detection
# since get_max_batch_size can't handle multimodal inputs
if args.calib_with_images:
print("Image-text calibration enabled. Using default batch_size=1 for calibration.")
args.batch_size = 1
# Speculative decoding offline model dost not support get_max_batch_size() because of
# the customized dataloader, so we set batch_size to 1 to avoid OOM.
elif args.specdec_offline_dataset is not None:
print(
"Offline speculative decoding calibration enabled. Using default batch_size=1 for calibration."
)
args.batch_size = 1
# Layerwise calibration processes one layer at a time; auto batch-size probing runs a
# full-model forward which defeats the point and can OOM on very large models.
elif is_layerwise:
print("Layerwise calibration enabled. Using default batch_size=1 for calibration.")
args.batch_size = 1
else:
# Calibration/sparsification will actually take much more memory than regular inference
# due to intermediate tensors for fake quantization. Setting sample_memory_usage_ratio
# to 2 to avoid OOM for AWQ/SmoothQuant fake quantization as it will take more memory than inference.
sample_memory_usage_ratio = (
2 if "awq" in args.qformat or "smoothquant" in args.qformat else 1.1
)
# Whisper model expects mel-spectrogram input features of length 3000
# Whisper model needs input of shape (batch_size, num_mel_bins, 3000)
# As the encoder of Whisper doesn't have embedding layer, input dtype has to be float
# For non-Whisper models (language models), sample_input will be set up inside get_max_batch_size()
if model_type == "whisper":
max_sample_length = 3000
num_mel_bins = language_model.config.num_mel_bins
sample_input_single_batch = (
torch.ones([1, num_mel_bins, max_sample_length], dtype=language_model.dtype).to(
language_model.device
)
* 100
)
else:
sample_input_single_batch = None
run_auto_quant = aq_config is not None
args.batch_size = get_max_batch_size(
language_model,
max_sample_length=args.calib_seq,
sample_memory_usage_ratio=sample_memory_usage_ratio if not run_auto_quant else 1.0,
sample_input_single_batch=sample_input_single_batch,
enable_grad=run_auto_quant,
)
args.batch_size = min(args.batch_size, sum(args.calib_size))
print(f"Use calib batch_size {args.batch_size}")
calib_dataloader, first_text_speech_dataset = make_calib_dataloader(
args,
language_model,
processor,
tokenizer,
device,
model_type,
autoquant_gradient_recipe=(
aq_config is not None and aq_config.auto_quantize_method == "gradient"
),
)
# Detect if this is a Nemotron VL model using architecture-based detection
is_nemotron_vl_model = is_nemotron_vl(full_model)
preview_input_ids, preview_attention_mask, generated_ids_before_ptq = pre_quantize(
args, full_model, model_type, tokenizer, calib_dataloader, is_nemotron_vl_model
)
if aq_config is not None:
assert isinstance(recipe, ModelOptAutoQuantizeRecipe)
_run_auto_quantize_recipe(
args,
recipe,
full_model,
language_model,
model_type,
calibration_only,
calib_dataloader,
is_nemotron_vl_model,
)
else:
# mono quantization
if recipe is not None:
quant_cfg = recipe.quantize.model_dump()
else:
assert len(args.qformat.split(",")) == 1, (
"Plain quantization supports only one quantization format."
)
assert args.qformat in QUANT_CFG_CHOICES, (
f"Unsupported quantization format: {args.qformat}, choices are: {list(QUANT_CFG_CHOICES)}"
)
quant_cfg = QUANT_CFG_CHOICES[args.qformat]
quant_cfg = build_quant_cfg(
quant_cfg,
args.awq_block_size,
args.moe_calib_experts_ratio,
)
enable_quant_kv_cache = args.kv_cache_qformat != KV_CACHE_NONE
print(f"{'Enable' if enable_quant_kv_cache else 'Disable'} KV cache quantization")
# Check if any bmm_quantizer is in the quant_cfg. If so, we need to enable the bmm_quantizer.
if enable_quant_kv_cache:
quant_cfg = mtq.update_quant_cfg_with_kv_cache_quant(
quant_cfg,
KV_QUANT_CFG_CHOICES[args.kv_cache_qformat]["quant_cfg"],
)
quant_cfg = _prepare_quant_cfg(args, quant_cfg, full_model)
if quant_cfg:
mono_quantize(
args,
quant_cfg,
full_model,
language_model,
model_type,
calibration_only,
calib_dataloader,
is_nemotron_vl_model,
)
else:
assert model_type != "dbrx", f"Does not support export {model_type} without quantizaton"
print(f"qformat: {args.qformat}. No quantization applied, export {device} model")
# If asked, run the closed-form MXFP4 -> NVFP4 cast: read the source MXFP4
# *_scales tensors and pin each NVFP4 weight quantizer's scale_2 to 2^m.
# Runs after calibration (max_calibrate has already promoted weight quantizers
# to NVFP4StaticQuantizer with a data-derived ``_global_amax``); we just
# override that scalar with the closed-form value before export.
if args.cast_mxfp4_to_nvfp4:
# The cast reads the source MXFP4 ``*_scales``/``*_blocks`` tensors from a local
# checkpoint directory. ``--pyt_ckpt_path`` may be a HF Hub ID (e.g.
# ``openai/gpt-oss-20b``); resolve it to the local snapshot dir that load_model's
# ``from_pretrained`` already populated so the cast works with the documented command.
source_ckpt_dir = _resolve_model_path(args.pyt_ckpt_path, args.trust_remote_code)
apply_cast_mxfp4_to_nvfp4(language_model, source_ckpt_dir)
post_quantize(
args,
full_model,
language_model,
model_type,
tokenizer,
processor,
preview_input_ids,
preview_attention_mask,
generated_ids_before_ptq,
is_nemotron_vl_model,
first_text_speech_dataset,
default_padding_side,
default_pad_token,
calib_dataloader,
)
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument(
"--pyt_ckpt_path",
"--model",
help=(
"Model name or path to the PyTorch checkpoint to be quantized. "
"Can be a local path or a Huggingface model name."
),
required=True,
)
parser.add_argument(
"--recipe",
help=(
"PTQ or AutoQuantize recipe YAML file or name without suffix (e.g. "
"general/ptq/nvfp4_default-kv_fp8_cast, general/auto_quantize/nvfp4_fp8_at_4p8bits). "
"KV cache behavior depends on the recipe type: PTQ recipes configure it in quant_cfg "
"and ignore --kv_cache_qformat; weight AutoQuantize recipes use their kv_cache setting "
"or fall back to --kv_cache_qformat; KV-cache AutoQuantize recipes select per-layer K/V "
"formats from candidate_formats and ignore --kv_cache_qformat."
),
default=None,
)
parser.add_argument("--device", default="cuda")
parser.add_argument(
"--qformat",
action=RecipeSupersededAction,
help="Quantization format for single-format PTQ. For mixed-precision search, use an "
"AutoQuantize recipe via --recipe. (deprecated: use --recipe)",
default="fp8",
)
parser.add_argument(
"--batch_size",
help="Batch size for calibration. Default to 0 as we calculate max batch size on-the-fly",
type=int,
default=0,
)
parser.add_argument(
"--calib_size",
help=(
"Number of samples for calibration. If a comma separated list of values is provided, "
"each value will be used as the calibration size for the corresponding dataset. "
"This argument will be parsed and converted as a list of ints."
),
type=str,
default="1024",
)
parser.add_argument(
"--calib_seq",
help="Maximum sequence length for calibration.",
type=int,
default=512,
)
parser.add_argument("--export_path", default="exported_model")
parser.add_argument(
"--dataset",
help=(
f"name of a dataset, or a comma separated list of datasets. "
f"dataset choices are {get_supported_datasets()}"
),
type=str,
default=None,
)
parser.add_argument(
"--specdec_offline_dataset",
help=(
"If set, the model is a speculative decoding model,"
"which uses offline dataset for calibration. "
),
default=None,
)
parser.add_argument(
"--calib_with_images",
action="store_true",
help=(
"Calibrate with image-text pairs (for VLMs). "
"This uses nemotron_vlm_dataset_v2 with default subsets (sparsetables, plotqa_cot, wiki_en)."
),
)
parser.add_argument("--inference_tensor_parallel", type=int, default=1)
parser.add_argument("--inference_pipeline_parallel", type=int, default=1)
parser.add_argument("--awq_block_size", default=0, type=int)
parser.add_argument(
"--sparsity_fmt",
help="Sparsity format.",
default="dense",
choices=["dense", "sparsegpt"],
)
parser.add_argument(
"--kv_cache_qformat",
action=RecipeSupersededAction,
required=False,
default="fp8_cast",
choices=[KV_CACHE_NONE, *KV_QUANT_CFG_CHOICES],
help=(
"Specify KV cache quantization format. Default: fp8_cast. "
"Formats whose preset pins use_constant_amax on the KV bmm quantizer "
"(e.g. fp8_cast, nvfp4_cast) set the amax to FP8 range without data-driven "
"calibration; all other formats (fp8, nvfp4, ...) use data-driven calibration. "
"With --recipe, the source depends on the recipe type: a PTQ recipe is "
"authoritative for KV cache and ignores this flag; an AutoQuantize recipe "
"falls back to this flag unless it sets an explicit kv_cache field. (deprecated: use --recipe)"
),
)
parser.add_argument(
"--export_fmt",
required=False,
default="hf",
choices=["tensorrt_llm", "hf"],
help="Deprecated. Please avoid using this argument.",
)
parser.add_argument(
"--trust_remote_code",
help="Set trust_remote_code for Huggingface models and tokenizers",
default=False,
action="store_true",
)
parser.add_argument(
"--gpu_max_mem_percentage",
help=(
"Specify the percentage of available GPU memory to use for loading the model when "
"device_map is set to sequential. "
"By default, 80%% of the available GPU memory is used."
),
type=float,
default=0.8,
)
parser.add_argument(
"--use_seq_device_map",
help=(
"Use device_map=sequential to load the model onto GPUs. This ensures the model is loaded "
"utilizing the percentage of available GPU memory as specified by the value passed with gpu_max_mem flag."
"Helpful in cases where device_map=auto loads model unevenly on GPUs causing GPU OOM during quantization."
),
default=False,
action="store_true",
)
parser.add_argument(
"--use_fsdp2",
action="store_true",
help=(
"Run calibration under PyTorch FSDP2 (requires torchrun); takes precedence over "
"--use_seq_device_map. v1: standard causal-LM only (no VILA / pack-quantized / "
"speculative / auto-quantize / sparsity / VLM / MTP)."
),
)
parser.add_argument(
"--cpu_offload",
action="store_true",
help="With --use_fsdp2, keep decoder shards on CPU between forwards (frees GPU memory, adds PCIe traffic).",
)
parser.add_argument(
"--verbose",
help="Print verbose output (e.g. quantization summary). Disable by --no-verbose.",
default=True,
action=argparse.BooleanOptionalAction,
)
parser.add_argument(
"--skip_generate",
help=(
"Skip pre/post-quantization preview calls that invoke model.generate(). "
"Note: this does not skip calibration or batch-size probing. "
"For very large models, pair with --batch_size 1 to avoid max-batch probing."
),
default=False,
action="store_true",
)
parser.add_argument(
"--low_memory_mode",
help=(
"Use low memory mode for quantization."
"This is an experimental feature and may not work for all quantization formats."
),
default=False,
action="store_true",
)
parser.add_argument(
"--attn_implementation",
help=(
"Specify the attention implementation to use. "
"This arg will be passed to the HF model loading if specified."
),
default=None,
type=str,
)
parser.add_argument(
"--auto_quantize_checkpoint",
type=str,
default=None,
help=(
"Path to checkpoint file for saving/restoring weight AutoQuantize search state "
"(sensitivity scores, costs, etc.). Used with an AutoQuantize --recipe."
),
)
parser.add_argument(
"--kv_auto_quantize_checkpoint",
type=str,
default=None,
help=(
"Path for saving/restoring any KV-cache AutoQuantize search checkpoint. Use a new "
"path whenever the preceding weight/activation quantization stage changes. "
"KV-primary recipes temporarily accept --auto_quantize_checkpoint as a deprecated "
"fallback."
),
)
parser.add_argument(
"--moe_calib_experts_ratio",
type=float,
default=None,
help=(
"Fraction of experts to calibrate during forward pass (ratio in (0.0, 1.0]). "
"Only used for MOE models; used to reduce the number of experts calibrated during the forward pass. "
"Does not impact non-MOE models."
),
)
parser.add_argument(
"--vllm_fakequant_export",
default=False,
action="store_true",
help="Export as vLLM fake-quant checkpoint (produces vllm_fq_modelopt_state.pth "
"for use with vllm_serve_fakequant.py).",
)
parser.add_argument(
"--cast_mxfp4_to_nvfp4",
action="store_true",
default=False,
help=(
"After calibration, override NVFP4 weight quantizers' global_amax with "
"the closed-form value derived from the source MXFP4 *_scales. "
"Per-block _amax is computed from the loaded BF16 weights (data-derived). "
"Use when --pyt_ckpt_path points at an MXFP4 HF checkpoint (e.g. "
"openai/gpt-oss-20b) and the target qformat is NVFP4-family."
),
)
parser.add_argument(
"--offload_folder",
type=str,
default=None,
help=(
"Path to a local folder for disk-offloaded model weights. "
"When set, activates disk-offload mode: model weights that exceed the GPU+CPU "
"budgets are streamed from disk during calibration and export. "
"Pair with --max_cpu_memory_gb to cap CPU RAM usage. "
"Incompatible with --low_memory_mode and --use_seq_device_map."
),
)
parser.add_argument(
"--max_cpu_memory_gb",
type=float,
default=None,
help=(
"Maximum CPU RAM budget in GiB for disk-offload model loading. "
"Only effective when --offload_folder is set. "
"Weights beyond this limit are streamed from disk."
),
)
parser.add_argument(
"--max_gpu_memory_gb",
type=float,
default=None,
help=(
"Maximum GPU memory budget per device in GiB for disk-offload model loading. "
"Only effective when --offload_folder is set. "
"Defaults to 80%% of available GPU memory when not specified."
),
)
add_mlflow_args(parser, HF_PTQ)
args = parser.parse_args()
# Flipped by export_quantized once a checkpoint is actually on disk. The MLflow pointer
# is gated on it rather than on --export_path existing, which proves nothing.
args.checkpoint_exported = False
resolve_mlflow_args(args, parser, HF_PTQ)
if args.moe_calib_experts_ratio is not None and not (0.0 < args.moe_calib_experts_ratio <= 1.0):
parser.error("--moe_calib_experts_ratio must be in the range (0.0, 1.0].")
if args.specdec_offline_dataset is not None and args.sparsity_fmt != "dense":
parser.error("--specdec_offline_dataset is only supported with --sparsity_fmt dense (PTQ).")
if args.specdec_offline_dataset is not None and args.low_memory_mode:
parser.error("--specdec_offline_dataset is not compatible with --low_memory_mode.")
# The low-memory loader pre-instruments quantizers from --qformat/--kv_cache_qformat
# via init_quantized_weights(), so it cannot honor a --recipe (which is authoritative
# for the quant layout in quantize_main). Reject the combination rather than silently
# instrumenting a layout that diverges from the recipe.
if args.low_memory_mode and args.recipe is not None:
parser.error(
"--low_memory_mode does not support --recipe; the low-memory loader initializes "
"quantizers from --qformat/--kv_cache_qformat."
)
if args.use_fsdp2 and args.use_seq_device_map:
warnings.warn("--use_seq_device_map is ignored when --use_fsdp2 is set.")
args.use_seq_device_map = False
if args.use_fsdp2 and os.environ.get("RANK") is None:
parser.error("--use_fsdp2 requires launching with torchrun")
if args.cpu_offload and not args.use_fsdp2:
parser.error("--cpu_offload requires --use_fsdp2")
if args.use_fsdp2 and args.sparsity_fmt != "dense":
parser.error(f"--use_fsdp2 does not support --sparsity_fmt {args.sparsity_fmt}.")
if args.use_fsdp2 and args.vllm_fakequant_export:
parser.error("--use_fsdp2 does not support --vllm_fakequant_export.")
if args.use_fsdp2 and args.cast_mxfp4_to_nvfp4:
parser.error("--use_fsdp2 does not support --cast_mxfp4_to_nvfp4.")
if args.offload_folder is not None and args.low_memory_mode:
parser.error("--offload_folder (disk-offload) is not compatible with --low_memory_mode.")
if args.offload_folder is not None and args.use_seq_device_map:
parser.error(
"--offload_folder (disk-offload) is not compatible with --use_seq_device_map; "
"device_map=auto is used for disk-offload to let accelerate place layers across "
"GPU, CPU, and disk."
)
if args.offload_folder is not None and args.device == "cpu":
parser.error(
"--offload_folder (disk-offload) is not compatible with --device cpu; "
"device_map=cpu makes accelerate ignore the memory budgets and offload folder, "
"loading the whole model into RAM."
)
if args.offload_folder is None and (
args.max_cpu_memory_gb is not None or args.max_gpu_memory_gb is not None
):
parser.error(
"--max_cpu_memory_gb/--max_gpu_memory_gb only apply to disk-offload loading; "
"pass --offload_folder to enable it."
)
return args
# Derived state and the tracking settings themselves; everything else argparse parsed is a
# parameter of the run. Deriving the list means a new flag is tracked without touching this.
def main(args: argparse.Namespace):
if not torch.cuda.is_available():
raise OSError("GPU is required for inference.")
random.seed(RAND_SEED)
np.random.seed(RAND_SEED)
setup_distributed_args(args)
try:
# Entered inside the try: opening the run is fatal by design, and skipping
# cleanup_distributed would leave the other ranks blocked on the first collective
# until the NCCL timeout.
with mlflow_run(args):
# launch a memory monitor to read the currently used GPU memory.
launch_memory_monitor()
# Force eager execution for all model types.
torch.compiler.set_stance("force_eager")
(
full_model,
language_model,
model_type,
calibration_only,
processor,
tokenizer,
default_padding_side,
default_pad_token,
device,
) = load_model(args)
if args.sparsity_fmt != "dense":
# Sparse
sparsity_main(args, full_model, tokenizer, device)
else:
# Quantize
quantize_main(
args,
full_model,
language_model,
model_type,
calibration_only,
processor,
tokenizer,
default_padding_side,
default_pad_token,
device,
)
finally:
cleanup_distributed(args)
if __name__ == "__main__":
args = parse_args()
if args.export_fmt != "hf":
warnings.warn("Deprecated. --export_fmt forced to hf.")
args.dataset = args.dataset.split(",") if isinstance(args.dataset, str) else args.dataset
args.calib_size = [int(num_sample) for num_sample in args.calib_size.split(",")]
if args.specdec_offline_dataset is not None and len(args.calib_size) != 1:
raise ValueError(
"--specdec_offline_dataset expects a single --calib value, not a comma-separated list."
)
if args.cast_mxfp4_to_nvfp4:
qformats = [q.strip() for q in args.qformat.split(",")]
if not all("nvfp4" in q for q in qformats):
raise ValueError(
"--cast_mxfp4_to_nvfp4 requires NVFP4-family --qformat values "
f"(got {args.qformat!r}). Use e.g. --qformat nvfp4 or nvfp4_mlp_only."
)
main(args)