mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
[OMNIML-5899] Export IQ checkpoints from HF and Megatron (#2447)
## Summary - add IQ format metadata and packed-weight export - support Hugging Face and TP=1 Megatron export paths - reject fused-MoE IQ export until a deployment loader owns its packed layout - document the shaped `uint8` weight contract and the fused-expert boundary - add Hugging Face, Megatron, metadata, and fused-expert export tests ## PR split This work is split into four focused PRs. Each PR targets `main` and owns a disjoint file set: 1. **Kernel** — [#2448: Add CUDA kernels for IQ packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448) 2. **Quantization** — [#2446: Add IQ quantization codecs and backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446) 3. **Export** — [#2447: Export IQ checkpoints from HF and Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447) 4. **Recipes** — [#2449: Add IQ post-training quantization recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449) The required merge order is #2448, #2446, #2447, then #2449. ## Scope This PR owns only export code, deployment documentation, and export tests. It targets `main` and should merge after #2448 and #2446. It does not contain kernel, codec/backend, or recipe files. ## Deployment consumer boundary Dense weights and individually named expert weights use the documented shaped `uint8` contract. Megatron fused-MoE IQ export is intentionally rejected with `NotImplementedError`: its payload would have shape `[num_experts, out_features, in_features // 256, payload_bytes]`, and no deployment loader in this stack currently owns that layout. Support should be enabled only with a loader integration test. ## Test coverage - [Hugging Face packed-weight export](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/unit/torch/export/test_export_weight.py) - [quantization metadata](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/unit/torch/export/test_get_quantization.py) - [Megatron unified export and fused-MoE rejection](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/gpu_megatron/torch/export/test_unified_export_megatron.py) ## Validation - all pre-commit hooks pass for the changed files - 89 focused Hugging Face export, metadata, and fused-expert tests pass locally - direct checks cover both fused-MoE export entry points for IQ1_S and IQ2_XS - Megatron GPU execution remains delegated to GPU CI - restricted-term scan passes <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for IQ1_S and IQ2_XS GGML quantization formats in unified Hugging Face and Megatron exports. * Added quantization metadata, tensor-shape recovery, packing details, and IQ2_XS size documentation. * Added validation for required block sizes and tensor parallelism settings. * **Limitations** * Fused-MoE and GPT-OSS IQ expert packing are not supported. * IQ exports require standard `weight` attributes in Hugging Face models. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Chenjie Luo
Claude Opus 5
parent
a7166965e3
commit
ed5c5ed369
@@ -50,6 +50,54 @@ The unified HF export API supports the following quantization formats:
|
||||
4. NVFP4_AWQ - NVIDIA 4-bit floating point with AWQ optimization
|
||||
5. INT4_AWQ - 4-bit integer with AWQ optimization
|
||||
6. W4A8_AWQ - 4-bit weights and 8-bit activations with AWQ optimization
|
||||
7. IQ1_S - 1-bit codebook quantization using the GGML block layout
|
||||
8. IQ2_XS - 2-bit codebook quantization using the GGML block layout
|
||||
|
||||
.. note::
|
||||
GGML has no equivalent for ModelOpt's per-tensor FP8 weight-and-activation format. In particular,
|
||||
GGML does not define a first-class FP8 tensor type with the corresponding per-tensor weight and
|
||||
activation scale semantics. Converting a ModelOpt FP8 checkpoint to GGUF therefore requires
|
||||
conversion to another GGML-supported tensor type rather than a lossless FP8 encoding.
|
||||
|
||||
IQ weight representation
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
For IQ1_S and IQ2_XS, unified export replaces each floating-point ``<module>.weight`` with a
|
||||
``uint8`` tensor containing byte-exact GGML blocks. Its shape is
|
||||
``[*logical_shape[:-1], logical_shape[-1] // 256, payload_bytes]``, where ``payload_bytes`` is 50
|
||||
for IQ1_S and 74 for IQ2_XS. No separate shape tensor is stored: a loader recovers the logical
|
||||
shape as ``[*weight.shape[:-2], weight.shape[-2] * 256]``. This is unambiguous because IQ export
|
||||
requires the logical last dimension to be divisible by 256.
|
||||
|
||||
.. note::
|
||||
Megatron IQ export currently requires tensor and pipeline model parallel sizes of 1. Packing
|
||||
happens during export, so a tensor-parallel shard would be packed as if it were a whole
|
||||
weight, and a pipeline stage holding no IQ layer would not reach the same rejection as its
|
||||
peers. Expert parallelism is supported, assuming every expert uses the same format.
|
||||
|
||||
.. warning::
|
||||
Megatron fused-MoE IQ export is not currently supported. Its packed tensor would require the
|
||||
deployment consumer to understand
|
||||
``[num_experts, out_features, in_features // 256, payload_bytes]`` rather than the ordinary HF
|
||||
fused-expert order. The exporter raises ``NotImplementedError`` until a deployment loader owns
|
||||
this layout and is covered by an integration test. Dense and individually named expert weights
|
||||
continue to use the representation above.
|
||||
|
||||
The generated configuration records ``quant_method: modelopt``, ``packing: ggml``, the 256-value
|
||||
block size, and the payload byte count. IQ payloads are not represented as compressed-tensors
|
||||
integer ``weights`` groups because all scales and indices are embedded in each packed block.
|
||||
|
||||
Each 74-byte IQ2_XS block represents 256 logical weights:
|
||||
|
||||
* bytes 0--1 are the little-endian FP16 super-block scale ``d``;
|
||||
* bytes 2--65 are 32 little-endian ``uint16`` codes, one per group of eight weights. Each code
|
||||
contains a 9-bit codebook index and seven stored sign bits; the eighth sign bit is derived from
|
||||
parity; and
|
||||
* bytes 66--73 contain sixteen 4-bit local-scale codes, packed two per byte. Each local scale is
|
||||
shared by two adjacent eight-weight groups.
|
||||
|
||||
The canonical 512-by-8 IQ2_XS codebook is part of the implementation rather than the checkpoint.
|
||||
The complete block therefore costs ``74 * 8 / 256 = 2.3125`` bits per logical weight.
|
||||
|
||||
Minimum Framework Versions
|
||||
--------------------------
|
||||
|
||||
@@ -19,6 +19,15 @@ import warnings
|
||||
from collections import defaultdict
|
||||
from typing import Any
|
||||
|
||||
from modelopt.torch.quantization.ggml import (
|
||||
IQ1_S_BLOCK_BYTES,
|
||||
IQ1_S_BLOCK_SIZE,
|
||||
IQ1_S_EFFECTIVE_BITS,
|
||||
IQ2_XS_BLOCK_BYTES,
|
||||
IQ2_XS_BLOCK_SIZE,
|
||||
IQ2_XS_EFFECTIVE_BITS,
|
||||
)
|
||||
|
||||
|
||||
def _quant_algo_to_group_config(quant_algo: str, group_size: int | None = None) -> dict[str, Any]:
|
||||
"""Map a per-layer quant_algo string to compressed-tensors config group details.
|
||||
@@ -29,7 +38,8 @@ def _quant_algo_to_group_config(quant_algo: str, group_size: int | None = None)
|
||||
|
||||
Returns:
|
||||
Dictionary with ``input_activations`` and ``weights`` entries suitable for
|
||||
a compressed-tensors ``config_groups`` entry.
|
||||
a compressed-tensors ``config_groups`` entry, or ModelOpt-owned metadata for
|
||||
self-contained IQ payloads.
|
||||
"""
|
||||
if quant_algo == "FP8":
|
||||
return {
|
||||
@@ -117,6 +127,26 @@ def _quant_algo_to_group_config(quant_algo: str, group_size: int | None = None)
|
||||
},
|
||||
"weights": {"dynamic": False, "num_bits": 8, "type": "float", "group_size": gs},
|
||||
}
|
||||
elif quant_algo in ("IQ1_S", "IQ2_XS"):
|
||||
if quant_algo == "IQ1_S":
|
||||
block_size = IQ1_S_BLOCK_SIZE
|
||||
payload_bytes = IQ1_S_BLOCK_BYTES
|
||||
effective_bits = IQ1_S_EFFECTIVE_BITS
|
||||
else:
|
||||
block_size = IQ2_XS_BLOCK_SIZE
|
||||
payload_bytes = IQ2_XS_BLOCK_BYTES
|
||||
effective_bits = IQ2_XS_EFFECTIVE_BITS
|
||||
if group_size not in (None, block_size):
|
||||
raise ValueError(f"{quant_algo} requires group size {block_size}, got {group_size}")
|
||||
# IQ payloads are self-contained blocks, not compressed-tensors integer groups.
|
||||
# Keep their format marker outside a ``weights`` quantization scheme.
|
||||
return {
|
||||
"quant_algo": quant_algo,
|
||||
"effective_bits": effective_bits,
|
||||
"group_size": block_size,
|
||||
"packing": "ggml",
|
||||
"block_payload_bytes": payload_bytes,
|
||||
}
|
||||
else:
|
||||
warnings.warn(
|
||||
f"Unsupported quantization algorithm '{quant_algo}' in "
|
||||
@@ -209,6 +239,13 @@ def convert_hf_quant_config_format(input_config: dict[str, Any]) -> dict[str, An
|
||||
"targets": ["Linear"],
|
||||
}
|
||||
new_config["config_groups"] = {"group_0": config_group_details}
|
||||
elif quant_algo_value in ("IQ1_S", "IQ2_XS"):
|
||||
# Forward the caller's group size so a mismatched one is rejected rather than rewritten
|
||||
# to the format's block size.
|
||||
iq_metadata = _quant_algo_to_group_config(
|
||||
quant_algo_value, original_quantization_details.get("group_size")
|
||||
)
|
||||
new_config.update(iq_metadata)
|
||||
elif quant_algo_value == "NVFP4_SVD":
|
||||
# NVFP4 + SVDQuant: NVFP4 weights/activations plus an AWQ-style
|
||||
# pre_quant_scale and a low-rank residual (svdquant_lora_a/b) stored as
|
||||
|
||||
@@ -36,11 +36,21 @@ QUANTIZATION_NVFP4_AWQ = "nvfp4_awq"
|
||||
QUANTIZATION_FP8_PB_REAL = "fp8_pb_real"
|
||||
QUANTIZATION_FP8_PB_WO = "fp8_pb_wo"
|
||||
QUANTIZATION_FP8_PC_PT = "fp8_pc_pt"
|
||||
QUANTIZATION_IQ1_S = "iq1_s"
|
||||
QUANTIZATION_IQ2_XS = "iq2_xs"
|
||||
|
||||
# Formats whose scales are purely per-module, so export never merges them across the q/k/v
|
||||
# and gate/up groups that share an input. Every other format unifies input_amax (and, for
|
||||
# NVFP4, weight_scale_2) across such a group, which only a whole-model forward can discover.
|
||||
FUSION_FREE_FORMATS = frozenset({QUANTIZATION_FP8, QUANTIZATION_NONE, QUANTIZATION_FP8_PB_REAL})
|
||||
FUSION_FREE_FORMATS = frozenset(
|
||||
{
|
||||
QUANTIZATION_FP8,
|
||||
QUANTIZATION_IQ1_S,
|
||||
QUANTIZATION_IQ2_XS,
|
||||
QUANTIZATION_NONE,
|
||||
QUANTIZATION_FP8_PB_REAL,
|
||||
}
|
||||
)
|
||||
|
||||
KV_CACHE_FP8 = "FP8"
|
||||
KV_CACHE_FP8_K_NVFP4_V = "FP8_K_NVFP4_V"
|
||||
|
||||
@@ -26,6 +26,14 @@ import torch.nn as nn
|
||||
|
||||
from modelopt import __version__
|
||||
from modelopt.torch.models import get_spec, list_all_possible
|
||||
from modelopt.torch.quantization.ggml import (
|
||||
IQ1_S_BLOCK_BYTES,
|
||||
IQ1_S_BLOCK_SIZE,
|
||||
IQ1_S_EFFECTIVE_BITS,
|
||||
IQ2_XS_BLOCK_BYTES,
|
||||
IQ2_XS_BLOCK_SIZE,
|
||||
IQ2_XS_EFFECTIVE_BITS,
|
||||
)
|
||||
from modelopt.torch.quantization.model_calib import (
|
||||
enable_stats_collection,
|
||||
finish_stats_collection,
|
||||
@@ -62,6 +70,8 @@ from .quant_format import (
|
||||
QUANTIZATION_INT4_AWQ,
|
||||
QUANTIZATION_INT8_SQ,
|
||||
QUANTIZATION_INT8_WO,
|
||||
QUANTIZATION_IQ1_S,
|
||||
QUANTIZATION_IQ2_XS,
|
||||
QUANTIZATION_MXFP4,
|
||||
QUANTIZATION_MXFP8,
|
||||
QUANTIZATION_NONE,
|
||||
@@ -440,6 +450,36 @@ def get_weight_block_size(module: nn.Module, weight_name: str = "weight") -> int
|
||||
return 0
|
||||
|
||||
|
||||
def uses_iq_quantization(module) -> bool:
|
||||
"""Whether any weight quantizer in ``module`` or its children targets an IQ format.
|
||||
|
||||
``get_quantization_format`` returns the *first* non-``NONE`` format it finds, so in a
|
||||
mixed-format model IQ layers sitting behind, say, an FP8 layer are invisible to it. Callers
|
||||
that must reject IQ specifically need to see every layer.
|
||||
|
||||
This reads ``num_bits`` directly rather than resolving each layer's full format, so an
|
||||
unrelated unsupported quantizer elsewhere in the model cannot turn the check into an error.
|
||||
|
||||
Known gap, shared with ``get_quantization_format``: ``weight_attr_names`` yields nothing for
|
||||
a TEGroupedLinear, whose parameters are ``weight0..N`` while its quantizer is a single
|
||||
``GroupedQuantizer`` under ``weight_quantizer``. Neither function sees such a module, so an
|
||||
experts-only IQ model reports no format at all -- not just here. Closing it belongs in
|
||||
``weight_attr_names``, where it affects every format, rather than in this helper.
|
||||
"""
|
||||
for weight_name in weight_attr_names(module):
|
||||
weight_quantizer = representative_weight_quantizer(module, weight_name)
|
||||
# getattr: a SequentialQuantizer has is_enabled but no num_bits, and is never IQ --
|
||||
# IQ is a single quantizer with backend="ggml".
|
||||
if (
|
||||
weight_quantizer is not None
|
||||
and weight_quantizer.is_enabled
|
||||
and getattr(weight_quantizer, "num_bits", None)
|
||||
in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS)
|
||||
):
|
||||
return True
|
||||
return any(uses_iq_quantization(child) for _, child in module.named_children())
|
||||
|
||||
|
||||
def get_quantization_format(module) -> str | None:
|
||||
"""Gets the quantization string.
|
||||
|
||||
@@ -474,6 +514,24 @@ def get_quantization_format(module) -> str | None:
|
||||
return QUANTIZATION_W4A8_AWQ
|
||||
|
||||
# Handle individual num_bits cases
|
||||
if weight_quantizer.num_bits in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
if weight_quantizer.backend != "ggml":
|
||||
raise ValueError("IQ formats require the built-in 'ggml' quantization backend")
|
||||
# Both exporters return before collecting input_scale and before the pre_quant_scale
|
||||
# handling below, so an enabled activation quantizer would be dropped without a trace
|
||||
# and the checkpoint would load as weight-only. Refuse instead.
|
||||
if input_quantizer is not None and input_quantizer.is_enabled:
|
||||
raise NotImplementedError(
|
||||
"IQ1_S/IQ2_XS export is weight-only, but this layer has an enabled input "
|
||||
"quantizer. The GGML block payload carries no activation scale, so the "
|
||||
"activation quantization would be silently lost."
|
||||
)
|
||||
if input_quantizer is not None and hasattr(input_quantizer, "_pre_quant_scale"):
|
||||
raise NotImplementedError(
|
||||
"IQ1_S/IQ2_XS export does not support an AWQ-style pre_quant_scale."
|
||||
)
|
||||
return weight_quantizer.num_bits
|
||||
|
||||
if weight_quantizer.num_bits == 4:
|
||||
assert len(weight_quantizer.block_sizes) > 0 and weight_quantizer.block_sizes[-1] > 0, (
|
||||
"Invalid block_sizes for INT4 quantizer"
|
||||
@@ -722,6 +780,26 @@ def process_layer_quant_config(layer_config_dict):
|
||||
"quant_algo": "MXFP8",
|
||||
"group_size": block_size_value,
|
||||
}
|
||||
elif v in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
if v == QUANTIZATION_IQ1_S:
|
||||
block_size = IQ1_S_BLOCK_SIZE
|
||||
payload_bytes = IQ1_S_BLOCK_BYTES
|
||||
effective_bits = IQ1_S_EFFECTIVE_BITS
|
||||
else:
|
||||
block_size = IQ2_XS_BLOCK_SIZE
|
||||
payload_bytes = IQ2_XS_BLOCK_BYTES
|
||||
effective_bits = IQ2_XS_EFFECTIVE_BITS
|
||||
if block_size_value != block_size:
|
||||
raise ValueError(
|
||||
f"{v.upper()} requires block size {block_size}, got {block_size_value}"
|
||||
)
|
||||
layer_config = {
|
||||
"quant_algo": v.upper(),
|
||||
"group_size": block_size,
|
||||
"effective_bits": effective_bits,
|
||||
"block_payload_bytes": payload_bytes,
|
||||
"packing": "ggml",
|
||||
}
|
||||
else:
|
||||
layer_config = {"quant_algo": v}
|
||||
|
||||
|
||||
@@ -62,6 +62,7 @@ except ImportError:
|
||||
from modelopt.torch.opt.conversion import ModeloptStateManager, modelopt_state
|
||||
from modelopt.torch.opt.plugins.huggingface import _MODELOPT_STATE_SAVE_NAME
|
||||
from modelopt.torch.quantization import set_quantizer_by_cfg_context
|
||||
from modelopt.torch.quantization.ggml import quantize_iq1_s, quantize_iq2_xs
|
||||
from modelopt.torch.quantization.nn import SequentialQuantizer, TensorQuantizer
|
||||
from modelopt.torch.quantization.qtensor import MXFP8QTensor, NVFP4QTensor
|
||||
from modelopt.torch.quantization.qtensor.base_qtensor import QTensorWrapper
|
||||
@@ -97,6 +98,8 @@ from .quant_format import (
|
||||
QUANTIZATION_FP8,
|
||||
QUANTIZATION_FP8_PB_REAL,
|
||||
QUANTIZATION_FP8_PC_PT,
|
||||
QUANTIZATION_IQ1_S,
|
||||
QUANTIZATION_IQ2_XS,
|
||||
QUANTIZATION_MXFP8,
|
||||
QUANTIZATION_NONE,
|
||||
QUANTIZATION_NVFP4,
|
||||
@@ -621,6 +624,20 @@ def _export_quantized_weight(
|
||||
"which dispatches to the streaming writer that materialises weights layer-by-layer."
|
||||
)
|
||||
|
||||
if quantization_format in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
if weight_name != "weight":
|
||||
raise NotImplementedError(
|
||||
"IQ unified export currently supports modules with a standard 'weight' "
|
||||
f"attribute, got {weight_name!r} on {type(sub_module).__name__}"
|
||||
)
|
||||
quantize_iq = (
|
||||
quantize_iq1_s if quantization_format == QUANTIZATION_IQ1_S else quantize_iq2_xs
|
||||
)
|
||||
packed_weight, _ = quantize_iq(weight.to(dtype))
|
||||
setattr(sub_module, weight_name, nn.Parameter(packed_weight, requires_grad=False))
|
||||
maybe_clear_cuda_cache()
|
||||
return
|
||||
|
||||
weight_quantizer: TensorQuantizer | SequentialQuantizer = getattr(
|
||||
sub_module, quantizer_attrs.weight_quantizer
|
||||
)
|
||||
|
||||
@@ -35,6 +35,7 @@ from safetensors import safe_open
|
||||
from safetensors.torch import save_file
|
||||
|
||||
from modelopt import __version__
|
||||
from modelopt.torch.quantization.ggml import quantize_iq1_s, quantize_iq2_xs
|
||||
from modelopt.torch.quantization.nn.modules.tensor_quantizer import GroupedQuantizer
|
||||
from modelopt.torch.utils import import_plugin, warn_rank_0
|
||||
|
||||
@@ -61,6 +62,8 @@ from .quant_format import (
|
||||
QUANTIZATION_FP8,
|
||||
QUANTIZATION_FP8_PB_REAL,
|
||||
QUANTIZATION_FP8_PB_WO,
|
||||
QUANTIZATION_IQ1_S,
|
||||
QUANTIZATION_IQ2_XS,
|
||||
QUANTIZATION_NONE,
|
||||
QUANTIZATION_NVFP4,
|
||||
QUANTIZATION_W4A16_NVFP4,
|
||||
@@ -75,6 +78,7 @@ from .quant_utils import (
|
||||
get_weight_scaling_factor_2,
|
||||
process_layer_quant_config,
|
||||
to_quantized_weight,
|
||||
uses_iq_quantization,
|
||||
)
|
||||
|
||||
with import_plugin("transformers", verbose=False):
|
||||
@@ -94,6 +98,7 @@ with import_plugin("megatron"):
|
||||
get_pipeline_model_parallel_rank,
|
||||
get_pipeline_model_parallel_world_size,
|
||||
get_tensor_model_parallel_rank,
|
||||
get_tensor_model_parallel_world_size,
|
||||
)
|
||||
from megatron.core.ssm.mamba_layer import MambaLayer
|
||||
from megatron.core.transformer.identity_op import IdentityOp
|
||||
@@ -312,10 +317,28 @@ class GPTModelExporter:
|
||||
is_last_stage_main_rank = pp_rank == pp_size - 1 and tp_rank == 0
|
||||
is_writer_rank = self._is_sidecar_writer_rank(is_last_stage_main_rank)
|
||||
|
||||
quantization_format = self._get_quantization_format(self.model)
|
||||
if self._any_rank_uses_iq_quantization():
|
||||
# Both sizes below are identical on every rank, and the IQ flag is agreed across
|
||||
# ranks, so these raise everywhere or nowhere. Raising on only a subset would strand
|
||||
# the rest in the collectives further down.
|
||||
if get_tensor_model_parallel_world_size() != 1:
|
||||
raise NotImplementedError(
|
||||
"Megatron IQ1_S/IQ2_XS unified export currently requires tensor model "
|
||||
"parallel size 1"
|
||||
)
|
||||
# Requiring PP=1 is also what makes the per-expert fused-MoE rejection safe: with
|
||||
# every rank holding the same layers, that check runs on all of them rather than
|
||||
# only the stages that happen to own an MoE block.
|
||||
if pp_size != 1:
|
||||
raise NotImplementedError(
|
||||
"Megatron IQ1_S/IQ2_XS unified export currently requires pipeline model "
|
||||
"parallel size 1"
|
||||
)
|
||||
|
||||
# Main export process
|
||||
layer_state_dicts = self.layer_state_dicts
|
||||
|
||||
quantization_format = self._get_quantization_format(self.model)
|
||||
quantization = None
|
||||
if quantization_format in (
|
||||
QUANTIZATION_FP8_PB_REAL,
|
||||
@@ -328,6 +351,8 @@ class GPTModelExporter:
|
||||
quantization = "NVFP4"
|
||||
elif quantization_format == QUANTIZATION_W4A16_NVFP4:
|
||||
quantization = "W4A16_NVFP4"
|
||||
elif quantization_format in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
quantization = quantization_format.upper()
|
||||
|
||||
if is_last_stage_main_rank:
|
||||
if is_writer_rank:
|
||||
@@ -1031,6 +1056,7 @@ class GPTModelExporter:
|
||||
module: torch.nn.Module,
|
||||
dtype: torch.dtype = torch.float16,
|
||||
name_to_value: dict[str, torch.Tensor] | None = None,
|
||||
keep_weight_device: bool = False,
|
||||
) -> dict[str, torch.Tensor]:
|
||||
"""Get the weight and bias of the module.
|
||||
|
||||
@@ -1039,6 +1065,7 @@ class GPTModelExporter:
|
||||
dtype: The data type of the weight and bias.
|
||||
name_to_value: The dictionary to store the weight and bias. A new dict is created
|
||||
if not provided.
|
||||
keep_weight_device: Keep the weight on its current device instead of moving it to CPU.
|
||||
|
||||
Returns:
|
||||
The dictionary containing the weight and bias.
|
||||
@@ -1049,7 +1076,9 @@ class GPTModelExporter:
|
||||
# layers whose weight is a placeholder) so callers can use "weight" in name_to_value
|
||||
# as a reliable guard without re-inspecting module.weight.
|
||||
if hasattr(module, "weight") and module.weight is not None and module.weight.numel() > 0:
|
||||
weight = module.weight.to(dtype).cpu()
|
||||
weight = module.weight.to(dtype)
|
||||
if not keep_weight_device:
|
||||
weight = weight.cpu()
|
||||
name_to_value["weight"] = weight
|
||||
|
||||
if hasattr(module, "bias") and module.bias is not None and module.bias.numel() > 0:
|
||||
@@ -1086,13 +1115,21 @@ class GPTModelExporter:
|
||||
self._record_excluded_module(prefix)
|
||||
block_size = get_weight_block_size(module)
|
||||
|
||||
name_to_value = self._get_weight_bias(module, dtype, name_to_value)
|
||||
is_iq = qformat in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS)
|
||||
name_to_value = self._get_weight_bias(
|
||||
module, dtype, name_to_value, keep_weight_device=is_iq
|
||||
)
|
||||
|
||||
if "weight" not in name_to_value:
|
||||
return name_to_value, qformat, block_size
|
||||
|
||||
if qformat == QUANTIZATION_NONE:
|
||||
return name_to_value, qformat, block_size
|
||||
# IQ formats derive all block metadata directly from the weight and do not use amax or
|
||||
# separately exported scaling tensors. Keep the weight on-device until it can be packed
|
||||
# along its contraction axis, so the CUDA packer can be used.
|
||||
if is_iq:
|
||||
return name_to_value, qformat, block_size
|
||||
# Getting the weight scales
|
||||
weight_scale = get_weight_scaling_factor(module)
|
||||
weight_scale_2 = get_weight_scaling_factor_2(module)
|
||||
@@ -1112,6 +1149,23 @@ class GPTModelExporter:
|
||||
|
||||
return name_to_value, qformat, block_size
|
||||
|
||||
def _any_rank_uses_iq_quantization(self) -> bool:
|
||||
"""Whether any rank's local stage holds an IQ layer.
|
||||
|
||||
Two reasons this is not ``self._get_quantization_format(self.model) in (...)``. That
|
||||
returns only the first non-NONE format in the tree, so a mixed-format model whose IQ
|
||||
layers follow, say, an FP8 one would slip past the caller's guard and pack TP-sharded
|
||||
weights as whole ones. And the scan is rank-local: under pipeline parallelism a stage
|
||||
holding no IQ layer would skip the raise and then block in the next collective while its
|
||||
peers exit. Agree across ranks first, mirroring ``_gather_exclude_modules``.
|
||||
"""
|
||||
local_uses_iq = uses_iq_quantization(self.model)
|
||||
if not torch.distributed.is_initialized():
|
||||
return local_uses_iq
|
||||
per_rank = [None] * torch.distributed.get_world_size()
|
||||
torch.distributed.all_gather_object(per_rank, local_uses_iq)
|
||||
return any(per_rank)
|
||||
|
||||
def _get_quantization_format(self, module: torch.nn.Module):
|
||||
return get_quantization_format(module)
|
||||
|
||||
@@ -1128,6 +1182,40 @@ class GPTModelExporter:
|
||||
|
||||
return weight_scale, weight_scale_2
|
||||
|
||||
@staticmethod
|
||||
def _pack_iq_weight(weight: torch.Tensor, qformat: str) -> torch.Tensor:
|
||||
"""Pack one ``[out, in]`` weight and return its CPU payload."""
|
||||
quantize_iq = quantize_iq1_s if qformat == QUANTIZATION_IQ1_S else quantize_iq2_xs
|
||||
packed_weight, _ = quantize_iq(weight)
|
||||
return packed_weight.detach().cpu()
|
||||
|
||||
@classmethod
|
||||
def _get_iq_weight_state(
|
||||
cls, weight_key: str, weight: torch.Tensor, qformat: str
|
||||
) -> dict[str, torch.Tensor]:
|
||||
"""Pack one ``[out, in]`` weight into the IQ checkpoint representation."""
|
||||
return {weight_key: cls._pack_iq_weight(weight, qformat)}
|
||||
|
||||
@staticmethod
|
||||
def _reject_unsupported_fused_iq_export(qformat: str) -> None:
|
||||
"""Reject fused-expert IQ payloads until a deployment loader owns their layout.
|
||||
|
||||
Raised from inside the per-expert loops, so it only runs on ranks that own an expert.
|
||||
The guards in ``save_pretrained`` are what make that safe: IQ export requires PP=1 and
|
||||
TP=1, so every rank holds the same layers and reaches the same loops, and expert
|
||||
parallelism shards a set of experts quantized alike -- so every rank arrives here with
|
||||
the same ``qformat`` and they raise together rather than stranding each other in a
|
||||
collective.
|
||||
|
||||
The one gap left is a rank holding no local expert at all, which needs expert-parallel
|
||||
size to exceed the expert count. Worth revisiting if that becomes a supported topology.
|
||||
"""
|
||||
if qformat in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
raise NotImplementedError(
|
||||
"Fused-MoE IQ export requires a deployment loader that supports "
|
||||
"[num_experts, out_features, in_features // 256, payload_bytes]"
|
||||
)
|
||||
|
||||
def _record_layer_quant_config(self, prefix: str, qformat: str | None, block_size: int | None):
|
||||
"""Record per-HF-layer quantization metadata for mixed precision exports."""
|
||||
if qformat in (None, QUANTIZATION_NONE):
|
||||
@@ -1192,7 +1280,9 @@ class GPTModelExporter:
|
||||
weight = weight + 1.0
|
||||
weight_scale, weight_scale_2 = self._get_weight_scales(name_to_value, qformat)
|
||||
|
||||
if weight_scale is None:
|
||||
if qformat in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
self._state_dict.update(self._get_iq_weight_state(prefix + "weight", weight, qformat))
|
||||
elif weight_scale is None:
|
||||
self._state_dict[prefix + "weight"] = weight
|
||||
else:
|
||||
self._state_dict[prefix + "weight"] = to_quantized_weight(
|
||||
@@ -1237,7 +1327,14 @@ class GPTModelExporter:
|
||||
gate_proj_weight = weight[:ffn_hidden_size, :]
|
||||
up_proj_weight = weight[ffn_hidden_size:, :]
|
||||
|
||||
if weight_scale is None:
|
||||
if qformat in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
self._state_dict.update(
|
||||
self._get_iq_weight_state(gate_proj_prefix + "weight", gate_proj_weight, qformat)
|
||||
)
|
||||
self._state_dict.update(
|
||||
self._get_iq_weight_state(up_proj_prefix + "weight", up_proj_weight, qformat)
|
||||
)
|
||||
elif weight_scale is None:
|
||||
self._state_dict[gate_proj_prefix + "weight"] = gate_proj_weight
|
||||
self._state_dict[up_proj_prefix + "weight"] = up_proj_weight
|
||||
else:
|
||||
@@ -1403,7 +1500,9 @@ class GPTModelExporter:
|
||||
name_to_value.pop("weight", None)
|
||||
seen_qformat, seen_block_size = qformat, block_size
|
||||
|
||||
weight = state_dict[weight_key].to(self.dtype).cpu()
|
||||
weight = state_dict[weight_key].to(self.dtype)
|
||||
if qformat not in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
weight = weight.cpu()
|
||||
weight_scale_cpu = (
|
||||
weight_scale.detach().cpu().clone() if weight_scale is not None else None
|
||||
)
|
||||
@@ -1434,7 +1533,13 @@ class GPTModelExporter:
|
||||
]
|
||||
|
||||
for shard_prefix, shard_weight, shard_scale in shards:
|
||||
if shard_scale is None:
|
||||
if qformat in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
local_expert_state.update(
|
||||
self._get_iq_weight_state(
|
||||
shard_prefix + "weight", shard_weight, qformat
|
||||
)
|
||||
)
|
||||
elif shard_scale is None:
|
||||
local_expert_state[shard_prefix + "weight"] = shard_weight
|
||||
else:
|
||||
local_expert_state[shard_prefix + "weight"] = to_quantized_weight(
|
||||
@@ -1597,7 +1702,10 @@ class GPTModelExporter:
|
||||
proj_weights = [_take(weight, s, hidden_size, g) for s, g in zip(slices, gated)]
|
||||
proj_keys = [p + "weight" for p in prefixes]
|
||||
|
||||
if weight_scale is None:
|
||||
if qformat in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
for key, weight in zip(proj_keys, proj_weights):
|
||||
self._state_dict.update(self._get_iq_weight_state(key, weight, qformat))
|
||||
elif weight_scale is None:
|
||||
for key, weight in zip(proj_keys, proj_weights):
|
||||
self._state_dict[key] = weight
|
||||
else:
|
||||
@@ -1712,7 +1820,15 @@ class GPTModelExporter:
|
||||
proj_keys = [p + "weight" for p in proj_prefixes]
|
||||
weight_scale, weight_scale_2 = self._get_weight_scales(name_to_value, qformat)
|
||||
|
||||
if weight_scale is None:
|
||||
if qformat in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
for proj_prefix, proj_weight in zip(proj_prefixes, proj_weights):
|
||||
if proj_prefix in keep_bf16:
|
||||
self._state_dict[proj_prefix + "weight"] = proj_weight.cpu()
|
||||
else:
|
||||
self._state_dict.update(
|
||||
self._get_iq_weight_state(proj_prefix + "weight", proj_weight, qformat)
|
||||
)
|
||||
elif weight_scale is None:
|
||||
for key, proj_weight in zip(proj_keys, proj_weights):
|
||||
self._state_dict[key] = proj_weight
|
||||
else:
|
||||
@@ -1808,6 +1924,7 @@ class GPTModelExporter:
|
||||
name_to_value, qformat, block_size = self._get_quantized_state(
|
||||
getattr(expert, layer_type), self.dtype, prefix=prefix
|
||||
)
|
||||
self._reject_unsupported_fused_iq_export(qformat)
|
||||
weight = name_to_value.pop("weight")
|
||||
weight_scale, weight_scale_2 = self._get_weight_scales(name_to_value, qformat)
|
||||
input_scale = (
|
||||
@@ -1877,6 +1994,7 @@ class GPTModelExporter:
|
||||
name_to_value, qformat, block_size = self._get_quantized_state(
|
||||
getattr(expert, layer_type), self.dtype, prefix=prefix
|
||||
)
|
||||
self._reject_unsupported_fused_iq_export(qformat)
|
||||
weight = name_to_value.pop("weight")
|
||||
bias = name_to_value.pop("bias", None)
|
||||
weight_scale, weight_scale_2 = self._get_weight_scales(name_to_value, qformat)
|
||||
|
||||
@@ -46,6 +46,9 @@ import modelopt.torch.quantization as mtq
|
||||
import modelopt.torch.speculative as mtsp
|
||||
from modelopt.torch.export import KV_CACHE_FP8, export_mcore_gpt_to_hf, import_mcore_gpt_from_hf
|
||||
from modelopt.torch.export.unified_export_megatron import GPTModelExporter
|
||||
from modelopt.torch.quantization.config import QuantizerAttributeConfig
|
||||
from modelopt.torch.quantization.ggml import dequantize_iq1_s, dequantize_iq2_xs, quantize_iq2_xs
|
||||
from modelopt.torch.quantization.nn import TensorQuantizer
|
||||
from modelopt.torch.speculative.eagle.default_config import default_eagle_config
|
||||
from modelopt.torch.speculative.plugins.megatron_eagle import _DynamicEagleGPTModel
|
||||
from modelopt.torch.speculative.plugins.megatron_medusa import _DynamicMedusaGPTModel
|
||||
@@ -86,6 +89,270 @@ def _verify_model_quant_config(
|
||||
assert quant_config_dict["kv_cache_quant_algo"] == KV_CACHE_FP8
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("qformat", "payload_bytes", "dequantize"),
|
||||
[("iq1_s", 50, dequantize_iq1_s), ("iq2_xs", 74, dequantize_iq2_xs)],
|
||||
)
|
||||
def test_megatron_name_remapping_exports_iq_payload(qformat, payload_bytes, dequantize):
|
||||
"""Megatron export writes the same scale-free IQ representation as HF export."""
|
||||
linear = torch.nn.Linear(256, 2, bias=False, dtype=torch.bfloat16)
|
||||
linear.weight_quantizer = TensorQuantizer(
|
||||
QuantizerAttributeConfig(
|
||||
num_bits=qformat,
|
||||
block_sizes={-1: 256},
|
||||
backend="ggml",
|
||||
)
|
||||
)
|
||||
exporter = object.__new__(GPTModelExporter)
|
||||
exporter.dtype = torch.bfloat16
|
||||
exporter._state_dict = {}
|
||||
exporter.exclude_modules = []
|
||||
exporter.layer_config_dict = {}
|
||||
|
||||
exporter._name_remapping(linear, "model.layers.0.mlp.down_proj.")
|
||||
|
||||
packed_key = "model.layers.0.mlp.down_proj.weight"
|
||||
assert exporter._state_dict[packed_key].shape == (2, 1, payload_bytes)
|
||||
assert exporter._state_dict[packed_key].dtype == torch.uint8
|
||||
logical_shape = torch.tensor(
|
||||
[
|
||||
*exporter._state_dict[packed_key].shape[:-2],
|
||||
exporter._state_dict[packed_key].shape[-2] * 256,
|
||||
]
|
||||
)
|
||||
reconstructed = dequantize(
|
||||
exporter._state_dict[packed_key],
|
||||
logical_shape,
|
||||
dtype=torch.bfloat16,
|
||||
)
|
||||
torch.testing.assert_close(reconstructed, linear.weight_quantizer(linear.weight))
|
||||
assert exporter.layer_config_dict == {
|
||||
"model.layers.0.mlp.down_proj.quantization": qformat,
|
||||
"model.layers.0.mlp.down_proj.awq_block_size": 256,
|
||||
}
|
||||
|
||||
|
||||
def _make_iq_experts(qformat, layer_type, *, bias=False):
|
||||
experts = torch.nn.ModuleList()
|
||||
generator = torch.Generator().manual_seed(1234)
|
||||
for _ in range(2):
|
||||
expert = torch.nn.Module()
|
||||
linear = torch.nn.Linear(256, 4, bias=bias, dtype=torch.bfloat16)
|
||||
with torch.no_grad():
|
||||
linear.weight.copy_(torch.randn(linear.weight.shape, generator=generator))
|
||||
if linear.bias is not None:
|
||||
linear.bias.copy_(torch.randn(linear.bias.shape, generator=generator))
|
||||
linear.weight_quantizer = TensorQuantizer(
|
||||
QuantizerAttributeConfig(
|
||||
num_bits=qformat,
|
||||
block_sizes={-1: 256},
|
||||
backend="ggml",
|
||||
)
|
||||
)
|
||||
expert.add_module(layer_type, linear)
|
||||
experts.append(expert)
|
||||
return experts
|
||||
|
||||
|
||||
def _make_iq_exporter():
|
||||
exporter = object.__new__(GPTModelExporter)
|
||||
exporter.dtype = torch.bfloat16
|
||||
exporter._state_dict = {}
|
||||
exporter.exclude_modules = []
|
||||
exporter.layer_config_dict = {}
|
||||
return exporter
|
||||
|
||||
|
||||
def _make_iq_weight(rows):
|
||||
return torch.linspace(-1, 1, rows * 256, dtype=torch.float32).reshape(rows, 256).bfloat16()
|
||||
|
||||
|
||||
def _assert_iq2_payload_matches(packed, logical_weight):
|
||||
expected, _ = quantize_iq2_xs(logical_weight)
|
||||
torch.testing.assert_close(packed, expected.cpu(), rtol=0, atol=0)
|
||||
|
||||
|
||||
def test_megatron_gated_mlp_slicing_exports_iq_payloads():
|
||||
weight = _make_iq_weight(8)
|
||||
module = SimpleNamespace(config=SimpleNamespace(ffn_hidden_size=4))
|
||||
exporter = _make_iq_exporter()
|
||||
exporter._get_quantized_state = lambda *a, **k: ({"weight": weight}, "iq2_xs", 256)
|
||||
|
||||
exporter._gated_mlp_slicing(module, "model.layers.0.mlp.")
|
||||
|
||||
_assert_iq2_payload_matches(
|
||||
exporter._state_dict["model.layers.0.mlp.gate_proj.weight"], weight[:4]
|
||||
)
|
||||
_assert_iq2_payload_matches(
|
||||
exporter._state_dict["model.layers.0.mlp.up_proj.weight"], weight[4:]
|
||||
)
|
||||
|
||||
|
||||
def test_megatron_grouped_mlp_slicing_exports_iq_payloads():
|
||||
weight = _make_iq_weight(8)
|
||||
module = SimpleNamespace(
|
||||
num_gemms=1,
|
||||
weight0=weight,
|
||||
local_expert_indices=[0],
|
||||
state_dict=lambda: {"weight0": weight},
|
||||
)
|
||||
exporter = _make_iq_exporter()
|
||||
exporter._get_quantized_state = lambda *a, **k: (
|
||||
{"weight": module.weight},
|
||||
"iq2_xs",
|
||||
256,
|
||||
)
|
||||
|
||||
exporter._grouped_mlp_slicing(
|
||||
module,
|
||||
"model.layers.0.mlp.experts.{}",
|
||||
gate_proj_name="gate_proj",
|
||||
up_proj_name="up_proj",
|
||||
)
|
||||
|
||||
_assert_iq2_payload_matches(
|
||||
exporter._state_dict["model.layers.0.mlp.experts.0.gate_proj.weight"], weight[:4]
|
||||
)
|
||||
_assert_iq2_payload_matches(
|
||||
exporter._state_dict["model.layers.0.mlp.experts.0.up_proj.weight"], weight[4:]
|
||||
)
|
||||
|
||||
|
||||
def test_megatron_qkv_slicing_exports_iq_payloads():
|
||||
weight = _make_iq_weight(8)
|
||||
module = SimpleNamespace(
|
||||
config=SimpleNamespace(
|
||||
hidden_size=256,
|
||||
num_query_groups=1,
|
||||
num_attention_heads=2,
|
||||
kv_channels=2,
|
||||
attention_output_gate=False,
|
||||
)
|
||||
)
|
||||
exporter = _make_iq_exporter()
|
||||
exporter._get_quantized_state = lambda *a, **k: ({"weight": weight}, "iq2_xs", 256)
|
||||
|
||||
exporter._qkv_slicing(module, "model.layers.0.self_attn.")
|
||||
|
||||
reshaped = weight.reshape(4, 2, 256)
|
||||
expected = {
|
||||
"q_proj": reshaped[:2].reshape(4, 256),
|
||||
"k_proj": reshaped[2].reshape(2, 256),
|
||||
"v_proj": reshaped[3].reshape(2, 256),
|
||||
}
|
||||
for projection, logical_weight in expected.items():
|
||||
_assert_iq2_payload_matches(
|
||||
exporter._state_dict[f"model.layers.0.self_attn.{projection}.weight"],
|
||||
logical_weight,
|
||||
)
|
||||
|
||||
|
||||
def test_megatron_gated_delta_net_slicing_exports_iq_payloads():
|
||||
weight = _make_iq_weight(12)
|
||||
module = SimpleNamespace(
|
||||
in_proj=object(),
|
||||
in_proj_split_names=("query", "key", "value", "z", "beta", "alpha"),
|
||||
in_proj_split_sections=(2, 2, 2, 2, 2, 2),
|
||||
)
|
||||
exporter = _make_iq_exporter()
|
||||
exporter._get_quantized_state = lambda *a, **k: ({"weight": weight}, "iq2_xs", 256)
|
||||
|
||||
exporter._gated_delta_net_slicing(module, "model.layers.0.mixer.")
|
||||
|
||||
_assert_iq2_payload_matches(
|
||||
exporter._state_dict["model.layers.0.mixer.in_proj_qkv.weight"], weight[:6]
|
||||
)
|
||||
_assert_iq2_payload_matches(
|
||||
exporter._state_dict["model.layers.0.mixer.in_proj_z.weight"], weight[6:8]
|
||||
)
|
||||
torch.testing.assert_close(
|
||||
exporter._state_dict["model.layers.0.mixer.in_proj_b.weight"], weight[8:10]
|
||||
)
|
||||
torch.testing.assert_close(
|
||||
exporter._state_dict["model.layers.0.mixer.in_proj_a.weight"], weight[10:]
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("qformat", ["iq1_s", "iq2_xs"])
|
||||
def test_megatron_packed_experts_reject_iq_without_deployment_loader(qformat):
|
||||
experts = _make_iq_experts(qformat, "linear_fc2")
|
||||
exporter = _make_iq_exporter()
|
||||
|
||||
with pytest.raises(NotImplementedError, match="Fused-MoE IQ export requires"):
|
||||
exporter._pack_name_remapping(
|
||||
experts,
|
||||
"model.layers.0.mlp.experts.down_proj",
|
||||
layer_type="linear_fc2",
|
||||
)
|
||||
assert exporter._state_dict == {}
|
||||
|
||||
|
||||
def test_megatron_gpt_oss_packed_experts_reject_iq_without_deployment_loader():
|
||||
experts = _make_iq_experts("iq2_xs", "linear_fc1", bias=True)
|
||||
exporter = _make_iq_exporter()
|
||||
|
||||
with pytest.raises(NotImplementedError, match="Fused-MoE IQ export requires"):
|
||||
exporter._pack_name_remapping_gpt_oss(
|
||||
experts,
|
||||
"model.layers.0.mlp.experts.gate_up_proj",
|
||||
layer_type="linear_fc1",
|
||||
)
|
||||
assert exporter._state_dict == {}
|
||||
|
||||
|
||||
def test_megatron_iq_export_rejects_tensor_parallelism():
|
||||
"""IQ packing is intentionally limited to complete TP=1 weights."""
|
||||
linear = torch.nn.Linear(256, 2, bias=False, dtype=torch.bfloat16)
|
||||
linear.weight_quantizer = TensorQuantizer(
|
||||
QuantizerAttributeConfig(
|
||||
num_bits="iq2_xs",
|
||||
block_sizes={-1: 256},
|
||||
backend="ggml",
|
||||
)
|
||||
)
|
||||
exporter = object.__new__(GPTModelExporter)
|
||||
exporter.model = torch.nn.Sequential(linear)
|
||||
|
||||
with (
|
||||
patch.object(exporter, "_is_sidecar_writer_rank", return_value=False),
|
||||
patch.object(uem, "get_pipeline_model_parallel_rank", return_value=0),
|
||||
patch.object(uem, "get_pipeline_model_parallel_world_size", return_value=1),
|
||||
patch.object(uem, "get_tensor_model_parallel_rank", return_value=0),
|
||||
patch.object(uem, "get_tensor_model_parallel_world_size", return_value=2),
|
||||
pytest.raises(NotImplementedError, match="tensor model parallel size 1"),
|
||||
):
|
||||
exporter.save_pretrained("unused", "unused")
|
||||
|
||||
|
||||
def test_megatron_iq_export_rejects_pipeline_parallelism():
|
||||
"""IQ packing requires PP=1 so the fused-MoE rejection reaches every rank.
|
||||
|
||||
The rejection raises from inside the per-expert loops, so a stage owning no expert would
|
||||
skip it and block in save_pretrained's collectives while its peers exit. PP=1 removes the
|
||||
divergence rather than trying to detect it.
|
||||
"""
|
||||
linear = torch.nn.Linear(256, 2, bias=False, dtype=torch.bfloat16)
|
||||
linear.weight_quantizer = TensorQuantizer(
|
||||
QuantizerAttributeConfig(
|
||||
num_bits="iq2_xs",
|
||||
block_sizes={-1: 256},
|
||||
backend="ggml",
|
||||
)
|
||||
)
|
||||
exporter = object.__new__(GPTModelExporter)
|
||||
exporter.model = torch.nn.Sequential(linear)
|
||||
|
||||
with (
|
||||
patch.object(exporter, "_is_sidecar_writer_rank", return_value=False),
|
||||
patch.object(uem, "get_pipeline_model_parallel_rank", return_value=0),
|
||||
patch.object(uem, "get_pipeline_model_parallel_world_size", return_value=2),
|
||||
patch.object(uem, "get_tensor_model_parallel_rank", return_value=0),
|
||||
patch.object(uem, "get_tensor_model_parallel_world_size", return_value=1),
|
||||
pytest.raises(NotImplementedError, match="pipeline model parallel size 1"),
|
||||
):
|
||||
exporter.save_pretrained("unused", "unused")
|
||||
|
||||
|
||||
def _test_unified_export_megatron(
|
||||
tmp_path,
|
||||
model_type,
|
||||
|
||||
@@ -20,10 +20,13 @@ import torch.nn as nn
|
||||
from _test_utils.torch.export.utils import ToyModel, partial_fp8_config, partial_w4a8_config
|
||||
|
||||
import modelopt.torch.quantization as mtq
|
||||
from modelopt.torch.export.quant_utils import postprocess_state_dict
|
||||
from modelopt.torch.export.unified_export_hf import (
|
||||
_export_quantized_weight,
|
||||
_process_quantized_modules,
|
||||
)
|
||||
from modelopt.torch.quantization.config import QuantizerAttributeConfig
|
||||
from modelopt.torch.quantization.nn import TensorQuantizer
|
||||
from modelopt.torch.quantization.utils import quantizer_attr_names
|
||||
|
||||
|
||||
@@ -102,6 +105,27 @@ def test_export_per_block_quantized_weight():
|
||||
assert not hasattr(model.linears[2], quantizer_attrs.output_scale)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(("num_bits", "payload_bytes"), [("iq1_s", 50), ("iq2_xs", 74)])
|
||||
def test_export_iq_payload_as_weight(num_bits, payload_bytes):
|
||||
linear = nn.Linear(256, 4, bias=False, dtype=torch.bfloat16)
|
||||
linear.weight_quantizer = TensorQuantizer(
|
||||
QuantizerAttributeConfig(
|
||||
num_bits=num_bits,
|
||||
block_sizes={-1: 256},
|
||||
backend="ggml",
|
||||
)
|
||||
)
|
||||
|
||||
_export_quantized_weight(linear, torch.bfloat16)
|
||||
state_dict = postprocess_state_dict(linear.state_dict(), maxbound=448, quantization=None)
|
||||
|
||||
assert isinstance(linear.weight, nn.Parameter)
|
||||
assert state_dict["weight"].shape == (4, 1, payload_bytes)
|
||||
assert state_dict["weight"].dtype == torch.uint8
|
||||
assert "packed_weights" not in state_dict
|
||||
assert "weight_shape" not in state_dict
|
||||
|
||||
|
||||
class QuantMoELinear(nn.Module):
|
||||
def __init__(self):
|
||||
super().__init__()
|
||||
|
||||
@@ -33,6 +33,8 @@ from modelopt.torch.export.quant_format import (
|
||||
KV_CACHE_FP8_K_NVFP4_V,
|
||||
KV_CACHE_NVFP4,
|
||||
QUANTIZATION_FP8,
|
||||
QUANTIZATION_IQ1_S,
|
||||
QUANTIZATION_IQ2_XS,
|
||||
QUANTIZATION_NVFP4,
|
||||
QUANTIZATION_W4A8_AWQ,
|
||||
)
|
||||
@@ -42,8 +44,14 @@ from modelopt.torch.export.quant_utils import (
|
||||
get_quant_config,
|
||||
get_quantization_format,
|
||||
postprocess_state_dict,
|
||||
process_layer_quant_config,
|
||||
uses_iq_quantization,
|
||||
)
|
||||
from modelopt.torch.quantization.nn import (
|
||||
NVFP4StaticQuantizer,
|
||||
SequentialQuantizer,
|
||||
TensorQuantizer,
|
||||
)
|
||||
from modelopt.torch.quantization.nn import NVFP4StaticQuantizer, TensorQuantizer
|
||||
|
||||
|
||||
class _FakeAttention(torch.nn.Module):
|
||||
@@ -53,6 +61,165 @@ class _FakeAttention(torch.nn.Module):
|
||||
self.v_bmm_quantizer = TensorQuantizer()
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("num_bits", "quantization_format", "payload_bytes", "effective_bits"),
|
||||
[
|
||||
("iq1_s", QUANTIZATION_IQ1_S, 50, 1.5625),
|
||||
("iq2_xs", QUANTIZATION_IQ2_XS, 74, 2.3125),
|
||||
],
|
||||
)
|
||||
def test_iq_quantization_config(num_bits, quantization_format, payload_bytes, effective_bits):
|
||||
model = torch.nn.Sequential(torch.nn.Linear(256, 256, bias=False))
|
||||
mtq.quantize(
|
||||
model,
|
||||
{
|
||||
"quant_cfg": [
|
||||
{"quantizer_name": "*", "enable": False},
|
||||
{
|
||||
"quantizer_name": "*weight_quantizer",
|
||||
"cfg": {
|
||||
"num_bits": num_bits,
|
||||
"block_sizes": {-1: 256},
|
||||
"backend": "ggml",
|
||||
},
|
||||
},
|
||||
],
|
||||
"algorithm": None,
|
||||
},
|
||||
)
|
||||
|
||||
assert get_quantization_format(model) == quantization_format
|
||||
config = get_quant_config(model)
|
||||
assert config["quantization"]["quant_algo"] == num_bits.upper()
|
||||
assert config["quantization"]["block_payload_bytes"] == payload_bytes
|
||||
assert config["quantization"]["effective_bits"] == effective_bits
|
||||
hf_config = convert_hf_quant_config_format(config)
|
||||
assert "config_groups" not in hf_config
|
||||
assert hf_config["group_size"] == 256
|
||||
assert hf_config["effective_bits"] == effective_bits
|
||||
assert hf_config["packing"] == "ggml"
|
||||
assert hf_config["block_payload_bytes"] == payload_bytes
|
||||
|
||||
|
||||
def _quantize_sequential(layer_cfgs):
|
||||
"""Quantize a two-Linear model, one quantizer config per layer."""
|
||||
model = torch.nn.Sequential(
|
||||
torch.nn.Linear(256, 256, bias=False), torch.nn.Linear(256, 256, bias=False)
|
||||
)
|
||||
mtq.quantize(
|
||||
model,
|
||||
{
|
||||
"quant_cfg": [{"quantizer_name": "*", "enable": False}, *layer_cfgs],
|
||||
"algorithm": None,
|
||||
},
|
||||
)
|
||||
return model
|
||||
|
||||
|
||||
_IQ_WEIGHT_CFG = {"num_bits": "iq1_s", "block_sizes": {-1: 256}, "backend": "ggml"}
|
||||
|
||||
|
||||
def test_uses_iq_quantization_sees_iq_behind_another_format():
|
||||
"""get_quantization_format stops at the first format, so the TP guard cannot rely on it."""
|
||||
model = _quantize_sequential(
|
||||
[
|
||||
{"quantizer_name": "0.weight_quantizer", "cfg": {"num_bits": (4, 3)}},
|
||||
{"quantizer_name": "1.weight_quantizer", "cfg": _IQ_WEIGHT_CFG},
|
||||
]
|
||||
)
|
||||
|
||||
assert get_quantization_format(model) == QUANTIZATION_FP8
|
||||
assert uses_iq_quantization(model)
|
||||
|
||||
|
||||
def test_uses_iq_quantization_false_without_iq_layers():
|
||||
model = _quantize_sequential(
|
||||
[{"quantizer_name": "*weight_quantizer", "cfg": {"num_bits": (4, 3)}}]
|
||||
)
|
||||
|
||||
assert not uses_iq_quantization(model)
|
||||
|
||||
|
||||
def test_uses_iq_quantization_tolerates_sequential_quantizer():
|
||||
"""A SequentialQuantizer has is_enabled but no num_bits, and is never IQ.
|
||||
|
||||
save_pretrained calls this on every Megatron export, so reading num_bits directly would
|
||||
raise AttributeError on a W4A8_AWQ model before any format dispatch.
|
||||
"""
|
||||
layer = torch.nn.Linear(256, 256, bias=False)
|
||||
layer.weight_quantizer = SequentialQuantizer(TensorQuantizer(), TensorQuantizer())
|
||||
assert not hasattr(layer.weight_quantizer, "num_bits")
|
||||
|
||||
assert not uses_iq_quantization(torch.nn.Sequential(layer))
|
||||
|
||||
|
||||
def test_iq_export_rejects_enabled_input_quantizer():
|
||||
"""IQ payloads carry no activation scale, so W-IQ + A-FP8 must not export as weight-only."""
|
||||
model = _quantize_sequential(
|
||||
[
|
||||
{"quantizer_name": "*weight_quantizer", "cfg": _IQ_WEIGHT_CFG},
|
||||
{"quantizer_name": "*input_quantizer", "cfg": {"num_bits": (4, 3)}},
|
||||
]
|
||||
)
|
||||
|
||||
with pytest.raises(NotImplementedError, match="weight-only"):
|
||||
get_quantization_format(model)
|
||||
|
||||
|
||||
def test_iq_hf_config_rejects_mismatched_group_size():
|
||||
"""A uniformly-IQ config must validate group_size, not silently rewrite it to the block size.
|
||||
|
||||
The MIXED_PRECISION branch already forwards the per-layer group size; this covers the
|
||||
top-level branch, which did not.
|
||||
"""
|
||||
with pytest.raises(ValueError, match="IQ2_XS requires group size 256, got 128"):
|
||||
convert_hf_quant_config_format(
|
||||
{
|
||||
"quantization": {
|
||||
"quant_algo": "IQ2_XS",
|
||||
"group_size": 128,
|
||||
"effective_bits": 2.3125,
|
||||
"packing": "ggml",
|
||||
"block_payload_bytes": 74,
|
||||
}
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
def test_mixed_iq_config_group_does_not_claim_integer_weight_schema():
|
||||
converted = convert_hf_quant_config_format(
|
||||
{
|
||||
"quantization": {
|
||||
"quant_algo": "MIXED_PRECISION",
|
||||
"quantized_layers": {
|
||||
"model.layers.0.mlp.down_proj": {
|
||||
"quant_algo": "IQ2_XS",
|
||||
"group_size": 256,
|
||||
"effective_bits": 2.3125,
|
||||
"packing": "ggml",
|
||||
"block_payload_bytes": 74,
|
||||
}
|
||||
},
|
||||
}
|
||||
}
|
||||
)
|
||||
|
||||
group = converted["config_groups"]["group_0"]
|
||||
assert "weights" not in group
|
||||
assert group["quant_algo"] == "IQ2_XS"
|
||||
assert group["packing"] == "ggml"
|
||||
|
||||
|
||||
def test_iq_quantization_config_rejects_mismatched_block_size():
|
||||
with pytest.raises(ValueError, match="IQ2_XS requires block size 256, got 128"):
|
||||
process_layer_quant_config(
|
||||
{
|
||||
"model.layers.0.mlp.down_proj.quantization": "iq2_xs",
|
||||
"model.layers.0.mlp.down_proj.awq_block_size": 128,
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
class _FakeKVCacheQuantizer(torch.nn.Module):
|
||||
"""Minimal FP8 KV cache quantizer for scaling-factor tests."""
|
||||
|
||||
|
||||
Reference in New Issue
Block a user