mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
[NVBug: 6563509] Drop Phi-3-vision / Phi-4-multimodal PTQ support (#2115)
### What does this PR do?
Type of change: Deprecation
Resolves [NVBug 6563509](https://nvbugspro.nvidia.com/bug/6563509),
where
`hf_ptq.py` on Phi-4-multimodal-instruct died with
`RuntimeError: Tensor.item() cannot be called on meta tensors`.
The crash is real but not fixable on our side, and it is not the reason
the model
is unusable. Phi-4-multimodal's bundled remote code predates
Transformers v5 and
does not load on **any** version in our supported range
(`transformers>=4.57,<5.15`):
| Blocker | Where |
|---|---|
| `peft.get_peft_model` reads `prepare_inputs_for_generation`, gone
since transformers 4.52 dropped `GenerationMixin` from `PreTrainedModel`
| `modeling_phi4mm.py:1959` |
| `_tied_weights_keys` declared as a list; Transformers 5.x calls
`.keys()` on it in `post_init` | `modeling_phi4mm.py:1937` |
| `int(torch.tensor(...))` in `__init__`, which cannot run on a meta
device — the reported crash | `speech_conformer_encoder.py:1435` |
The model card pins `transformers==4.48.2` / `peft==0.13.2`, so there is
no
overlap with our floor and nothing on our side can bridge it. The model
is
therefore dropped rather than worked around.
**Phi-3-vision is dropped alongside it because it is the older,
superseded model
in the same family** — with its successor unsupportable there is no
reason to
keep carrying the predecessor. This is a product-scope call, not a
separate
compatibility finding: Phi-3-vision shares the list-valued
`_tied_weights_keys`
defect (`modeling_phi3_v.py:1214`) and so is likewise broken on
Transformers 5.x,
but it does **not** hit the `peft` blocker, and it was not re-verified
on 4.57.
Per the 0.46 changelog we have already bumped the floor to 4.57 and
noted that
"Transformers 4.x support will be dropped in a future release", so any
remaining
window closes on its own. Same reasoning already applied to VILA / NVILA
in this
release.
**Removed**
- the support-matrix row in `examples/hf_ptq/README.md`
- `"Phi4MMForCausalLM": "phi4mm"` from `MODEL_NAME_TO_TYPE`
- the multimodal-detection heuristics that only ever matched these two —
`vision_lora`, `audio_processor`, `embd_layer.image_embd_layer`, and the
`phi4mm` model-type check — in both `is_multimodal_model` and
`_is_multimodal_config`
- the `Phi3Image` / `PhiImage` exclusions in `is_embedding`
- the phi4mm input-mode warning in `hf_ptq.py`
- `modelopt_recipes/huggingface/phi4mm/` and its references in
`modelopt_recipes/ptq.md`
**Not changed:** the device-map sizing path (meta-device skeleton,
`infer_auto_device_map`, and the `--gpu_max_mem_percentage` cap) keeps
its
original behavior. That cap is wanted exactly where it already fires —
when the
model is already offloading to CPU, where it costs little and the
headroom is
required. With the affected checkpoints removed, there is no supported
model
that trips the meta-device build, so there is nothing to work around
here.
Text-only **Phi-3/Phi-4** and **Phi-3.5-MoE** are natively supported by
transformers and are untouched.
### Testing
On H200, `nvcr.io/nvidia/tensorrt-llm/release` (torch 2.12, transformers
5.5.4),
against the real checkpoint:
- **Version matrix** (vanilla transformers, no modelopt) — Phi-4-MM
loads at
4.48.2 / 4.49.0 / 4.50.0 / 4.51.3 and fails at 4.53.3 / 4.56.2 / 4.57.1
(`AttributeError: 'Phi4MMModel' object has no attribute
'prepare_inputs_for_generation'`) and at 5.5.4 (meta-init, then
tied-keys).
This is what establishes that no supported version works.
- `tests/examples/hf_ptq/test_example_utils.py` — 28 passed.
- **Sweep**: `tests/examples/hf_ptq` + `tests/unit/torch/export` —
failure set
identical to the pre-change tree (GPU/model-dependent `test_vlm_ptq`,
plus
`test_quant_aware_conversion` scoped-mapping tests), so none are
introduced
here.
- `pre-commit` clean on all changed files, including recipe validation.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ — PTQ for Phi-3-vision and
Phi-4-multimodal is removed, along with the `huggingface/phi4mm/ptq/*`
recipes. Phi-4-multimodal is already unloadable on every supported
transformers
version, so no working workflow regresses; Phi-3-vision is a deliberate
scope
removal as its superseded predecessor.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — this is a deletion; the
existing
`test_get_model_*` / `test_resolve_init_config_*` tests are unchanged
and still
pass.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
Two related references were left in place deliberately; say the word and
I'll
fold them in:
- `tests/examples/hf_ptq/test_deploy.py` still deploys the
already-published
`nvidia/Phi-4-multimodal-instruct-{NVFP4,FP8}` checkpoints. Those
artifacts
exist and serve fine; this PR only removes the ability to *produce*
them.
- `examples/torch_onnx/README.md` still lists Phi-4-multimodal-instruct.
That
is a separate ONNX pipeline that does not go through `get_model()` and
was not
tested here.
Earlier revisions of this branch also reworked the device-map sizing so
the
meta-tensor crash could not occur. That was reverted in 701180ed6: the
guard is
correct as written, and every alternative either changed behavior for
models that
fit today or moved the guard somewhere it does not belong, for a crash
that only
ever affected the checkpoints this PR removes.
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
f2bfe63183
commit
9220fac053
+2
-1
@@ -44,7 +44,7 @@ Changelog
|
||||
- **Deduplicate the modules shared at source** in the quantized export step: ``_export_quantized_weight`` and ``_export_fused_experts`` now alias bit-identical packed ``weight`` / ``weight_scale`` / ``weight_scale_2`` buffers across modules sharing a source weight ``data_ptr()`` so the downstream ``postprocess_state_dict`` dedup catches them (~42% storage reduction on ``nvfp4_experts_only`` for tied 26B MoE checkpoints).
|
||||
- New ``sync_tied_input_amax`` helper max-merges per-side ``input_quantizer.amax`` across tied modules before export so single-backbone consumers that load one ``input_scale`` per parameter don't clip either side.
|
||||
- The exported state_dict is also **reordered (decoder keys win instead of encoder)** so canonical-side keys per HF's ``_tied_weights_keys`` declaration win the data_ptr dedup; gated to the DiffusionGemma model class in ``_reorder_canonical_first``, no-op for every other model.
|
||||
- New DiffusionGemma model-specific recipe under ``modelopt_recipes/huggingface/diffusion_gemma/ptq/`` (``nvfp4_experts_only.yaml`` + its ``disabled_quantizers.yaml`` unit) adds the ``*self_conditioning*`` exclude on top of the standard default, leaving the shared ``default_disabled_quantizers`` unit clean for non-diffusion models — pattern matches the existing ``phi4mm`` / ``nemotron_vl`` model-specific recipes.
|
||||
- New DiffusionGemma model-specific recipe under ``modelopt_recipes/huggingface/diffusion_gemma/ptq/`` (``nvfp4_experts_only.yaml`` + its ``disabled_quantizers.yaml`` unit) adds the ``*self_conditioning*`` exclude on top of the standard default, leaving the shared ``default_disabled_quantizers`` unit clean for non-diffusion models — pattern matches the existing ``nemotron_vl`` model-specific recipes.
|
||||
- ``hf_ptq.py`` also unwraps ``ModelOutput`` dataclasses from ``.generate()`` so the preview decode works on diffusion models. Non-tied models see no behavioral change.
|
||||
- Add Torch-TensorRT FP8 deployment example for HuggingFace ViT (``examples/torch_trt/``): ``torch_tensorrt_ptq.py`` covers ``mtq.quantize`` → ``torch_tensorrt.compile(ir="dynamo")``, and ``torch_tensorrt_accuracy.py`` reports the compiled model's ImageNet-1k top-1/top-5 accuracy via the ``onnx_ptq`` ``evaluate`` harness (the unquantized baseline is Torch-TensorRT-compiled too, for an apples-to-apples comparison). Ships a ViT-tuned FP8 PTQ recipe under ``modelopt_recipes/huggingface/vit/ptq/`` (``fp8.yaml``) composed from the shared ``modelopt_recipes/configs/`` units: it quantizes the encoder Linears, patch-embed ``nn.Conv2d``, ``classifier``, and per-block LayerNorm inputs plus the attention Q/K/V BMMs and softmax. Verified on ``google/vit-base-patch16-224`` (ImageNet-1k 50k validation): FP8 stays within 0.13 pp Top-1 of the FP16 baseline.
|
||||
- Add **AutoQuantize recipe** support: ``mtq.auto_quantize`` can be driven declaratively from a YAML recipe (``RecipeType.AUTO_QUANTIZE`` / ``AutoQuantizeConfig``) specifying candidate formats, the ``effective_bits`` target, cost model (incl. ``active_moe`` and ``excluded_module_name_patterns``), scoring method, and disabled layers. Adds an ``effective_bits`` cost-model override on ``QuantizeConfig`` / ``QuantizerAttributeConfig`` (block-scale-accurate NVFP4 = 4.5 via ``configs/numerics/nvfp4``). Shipped recipes live under ``modelopt_recipes/general/auto_quantize/`` and model-specific ones under ``modelopt_recipes/huggingface/<model>/auto_quantize/``.
|
||||
@@ -87,6 +87,7 @@ Changelog
|
||||
- Remove the deprecated ``examples/llm_autodeploy`` example (deprecated in 0.45). Use TensorRT-LLM's `AutoDeploy <https://github.com/NVIDIA/TensorRT-LLM/tree/main/examples/auto_deploy>`_ directly together with ModelOpt PTQ in ``examples/hf_ptq``.
|
||||
- Remove the deprecated ``examples/llm_qad`` Megatron-LM QAD example (deprecated in 0.45). Use the `megatron_bridge QAD example <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/megatron_bridge#quantization-aware-distillation-qad>`_ instead, which provides a simpler Python-based interface and better model coverage.
|
||||
- Dropped VILA / NVILA vision-language model support in ``examples/hf_ptq``. VILA's modeling code requires ``transformers<=4.50.0``, which conflicts with ModelOpt's minimum supported ``transformers`` version. The VILA-specific bootstrap (repo clone, ``requirements-vila.txt``) and loading paths in ``example_utils.py`` have been removed.
|
||||
- Dropped **Phi-4-multimodal** PTQ support in ``examples/hf_ptq`` (NVBug 6563509). Its bundled remote code predates Transformers v5 and no longer loads on any version ModelOpt supports (``transformers>=4.57``): it requires ``transformers<4.52`` because it reaches ``prepare_inputs_for_generation`` through ``peft``, which needs ``PreTrainedModel`` to still inherit ``GenerationMixin``, and it declares ``_tied_weights_keys`` as a list, which Transformers 5.x rejects. **Phi-3-vision** is dropped alongside it: it is the older, superseded model in the same family, so with its successor unsupportable there is no reason to keep carrying the predecessor. (Phi-3-vision shares the list-valued ``_tied_weights_keys`` defect and so is likewise broken on Transformers 5.x, though it does not hit the ``peft`` blocker.) The support-matrix row, the ``phi4mm`` model type, the multimodal-detection heuristics that only ever matched these two (``vision_lora`` / ``audio_processor`` / ``embd_layer.image_embd_layer``), the ``Phi3Image`` / ``PhiImage`` embedding-export exclusions, and the ``modelopt_recipes/huggingface/phi4mm/`` recipes have been removed. Text-only Phi-3/Phi-4 and Phi-3.5-MoE are unaffected.
|
||||
|
||||
**Deprecations**
|
||||
|
||||
|
||||
@@ -119,7 +119,6 @@ Please reference our [framework scripts](#framework-scripts) and our [docs](http
|
||||
| Whisper<sup>9</sup> | ✅ | ❌ | ❌ | ❌ | - |
|
||||
| Nemotron-3 | ✅ | ❌ | ❌ | ❌ | ✅ |
|
||||
| Llava (VLM)<sup>11</sup> | ✅ | ✅<sup>12</sup> | ✅ | ✅ | - |
|
||||
| Phi-3-vision, Phi-4-multimodal (VLM)<sup>11</sup> | ✅ | ✅<sup>12</sup> | ✅ | ✅ | ✅ |
|
||||
| Qwen2, 2.5-VL (VLM)<sup>11</sup> | ✅ | ✅<sup>12</sup> | ✅ | ✅ | ✅ |
|
||||
| Gemma 3 (VLM)<sup>11</sup> | ✅ | - | - | - | - |
|
||||
| Nemotron VL (VLM)<sup>11,13</sup> | ✅ | - | - | - | ✅ |
|
||||
|
||||
@@ -177,12 +177,6 @@ def _is_multimodal_config(config):
|
||||
"""Check if a config indicates a multimodal model (config-only version of is_multimodal_model)."""
|
||||
return (
|
||||
hasattr(config, "vision_config") # Standard vision config (e.g., Qwen2.5-VL)
|
||||
or getattr(config, "model_type", "") == "phi4mm" # Phi-4 multimodal
|
||||
or hasattr(config, "vision_lora") # Vision LoRA configurations
|
||||
or hasattr(config, "audio_processor") # Audio processing capabilities
|
||||
or (
|
||||
hasattr(config, "embd_layer") and hasattr(config.embd_layer, "image_embd_layer")
|
||||
) # Image embedding layers
|
||||
or getattr(config, "is_encoder_decoder", False) # Encoder-decoder VL models
|
||||
or any( # Architecture-based detection for custom VL models (e.g., Nemotron-Parse)
|
||||
"conditionalgeneration" in arch.lower() for arch in getattr(config, "architectures", [])
|
||||
|
||||
@@ -705,9 +705,6 @@ def load_model(args: argparse.Namespace):
|
||||
# Left padding usually provides better calibration result.
|
||||
tokenizer.padding_side = "left"
|
||||
|
||||
if model_type == "phi4mm":
|
||||
warnings.warn("Please set the default input_mode to InputMode.LANGUAGE before quantizing.")
|
||||
|
||||
return (
|
||||
full_model,
|
||||
language_model,
|
||||
|
||||
@@ -222,12 +222,7 @@ def is_conv(module: nn.Module) -> bool:
|
||||
def is_embedding(module: nn.Module) -> bool:
|
||||
"""Returns whether the module is an embedding layer."""
|
||||
module_type_name = type(module).__name__
|
||||
return (
|
||||
"Embedding" in module_type_name
|
||||
and "Rotary" not in module_type_name
|
||||
and "PhiImage" not in module_type_name
|
||||
and "Phi3Image" not in module_type_name
|
||||
)
|
||||
return "Embedding" in module_type_name and "Rotary" not in module_type_name
|
||||
|
||||
|
||||
def build_embedding_config(module: nn.Module, normalization_constant: float = 1) -> EmbeddingConfig:
|
||||
|
||||
@@ -44,7 +44,6 @@ MODEL_NAME_TO_TYPE = {
|
||||
"phi3small": "phi3small",
|
||||
"phi3": "phi3",
|
||||
"PhiMoEForCausalLM": "phi3",
|
||||
"Phi4MMForCausalLM": "phi4mm",
|
||||
"phi": "phi",
|
||||
"TLGv4ForCausalLM": "phi",
|
||||
"MixtralForCausalLM": "llama",
|
||||
@@ -88,10 +87,6 @@ def is_multimodal_model(model):
|
||||
This function detects various multimodal model architectures by checking for:
|
||||
- Standard vision configurations (vision_config)
|
||||
- Language model attributes (language_model)
|
||||
- Specific multimodal model types (phi4mm)
|
||||
- Vision LoRA configurations
|
||||
- Audio processing capabilities
|
||||
- Image embedding layers
|
||||
- Nemotron-Parse conditional generation models
|
||||
|
||||
Args:
|
||||
@@ -104,10 +99,6 @@ def is_multimodal_model(model):
|
||||
>>> model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-VL-7B-Instruct")
|
||||
>>> is_multimodal_model(model)
|
||||
True
|
||||
|
||||
>>> model = AutoModelForCausalLM.from_pretrained("microsoft/Phi-4-multimodal-instruct")
|
||||
>>> is_multimodal_model(model)
|
||||
True
|
||||
"""
|
||||
config = model.config
|
||||
|
||||
@@ -118,12 +109,6 @@ def is_multimodal_model(model):
|
||||
return (
|
||||
hasattr(config, "vision_config") # Standard vision config (e.g., Qwen2.5-VL)
|
||||
or hasattr(model, "language_model") # Language model attribute (e.g., LLaVA)
|
||||
or getattr(config, "model_type", "") == "phi4mm" # Phi-4 multimodal
|
||||
or hasattr(config, "vision_lora") # Vision LoRA configurations
|
||||
or hasattr(config, "audio_processor") # Audio processing capabilities
|
||||
or (
|
||||
hasattr(config, "embd_layer") and hasattr(config.embd_layer, "image_embd_layer")
|
||||
) # Image embedding layers
|
||||
or is_nemotron_parse # Nemotron-Parse conditional generation model
|
||||
)
|
||||
|
||||
|
||||
@@ -1,13 +0,0 @@
|
||||
# Phi-4-Multimodal PTQ recipes
|
||||
|
||||
Phi-4-Multimodal is a multimodal model. Quantization should be applied only to
|
||||
the language model; the speech, audio, image, and vision branches are kept in
|
||||
full precision to avoid accuracy regressions on those modalities.
|
||||
|
||||
| File | What's model-specific |
|
||||
|------|-----------------------|
|
||||
| `disabled_quantizers.yaml` | Reusable unit (`QuantizerCfgListConfig`). Merges the standard `default_disabled_quantizers` exclusions with Phi-4-MM ones (`*speech*`, `*audio*`, `*image*`, `*vision*`). Imported by recipes below as the single `disabled_quantizers` slot so they don't pull in two disabled-quantizer sets. |
|
||||
| `nvfp4-kv_fp8_cast.yaml` | NVFP4 W4A4 model quantization + FP8 KV-cache cast (constant amax, no KV calibration). Identical numerics to the general `nvfp4` preset / `kv_fp8_cast` unit; what makes it model-specific is that it imports `disabled_quantizers.yaml` from this folder to skip the non-language branches. |
|
||||
|
||||
Additional `<qformat>-kv_fp8_cast.yaml` recipes can be generated for other formats
|
||||
if needed; only `nvfp4-kv_fp8_cast.yaml` is shipped by default.
|
||||
@@ -1,34 +0,0 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# QuantizerCfgList snippet of disabled quantizers for Phi-4-Multimodal.
|
||||
# Splices in the standard `default_disabled_quantizers` exclusions and appends
|
||||
# Phi-4-MM-specific ones so that only the language model is quantized;
|
||||
# speech/audio/image/vision branches are skipped. Recipes that import this
|
||||
# should NOT also import `default_disabled_quantizers`.
|
||||
|
||||
# modelopt-schema: modelopt.torch.quantization.config.QuantizerCfgListConfig
|
||||
imports:
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
---
|
||||
- $import: default_disabled_quantizers
|
||||
- quantizer_name: '*speech*'
|
||||
enable: false
|
||||
- quantizer_name: '*audio*'
|
||||
enable: false
|
||||
- quantizer_name: '*image*'
|
||||
enable: false
|
||||
- quantizer_name: '*vision*'
|
||||
enable: false
|
||||
@@ -1,36 +0,0 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# Phi-4-Multimodal-specific PTQ recipe for the `nvfp4` quantization format.
|
||||
# Equivalent to the general `nvfp4` preset with quantization disabled
|
||||
# on non-language branches.
|
||||
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
w4a4_nvfp4_nvfp4: configs/ptq/units/w4a4_nvfp4_nvfp4
|
||||
disabled_quantizers: huggingface/phi4mm/ptq/disabled_quantizers
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: 'Phi-4-Multimodal PTQ recipe (nvfp4): same numerics as the general nvfp4 preset, applied to the language model only (speech, audio, image,
|
||||
and vision branches are skipped).'
|
||||
quantize:
|
||||
algorithm: max
|
||||
quant_cfg:
|
||||
- $import: base_disable_all
|
||||
- $import: w4a4_nvfp4_nvfp4
|
||||
- $import: kv_fp8_cast
|
||||
- $import: disabled_quantizers
|
||||
@@ -234,7 +234,7 @@ that baseline. The deviations come in four kinds:
|
||||
|------|-------------------------------------|----------|
|
||||
| **Architecture-aware `quant_cfg`** | Per-sub-module format choices a single wildcard scheme can't express | `minimax_m3_vl`, `qwen3_5`, `qwen3_5_moe`, `vit`, `nemotron_llama` |
|
||||
| **Algorithm override** | Same numerics & scope, but the *calibration algorithm* is tweaked because the default breaks or regresses | `gemma`, `gemma4`, `mpt` |
|
||||
| **Extra exclusions** | Adds disabled-quantizer patterns so non-language branches stay full precision | `nemotron_vl`, `phi4mm`, `diffusion_gemma` |
|
||||
| **Extra exclusions** | Adds disabled-quantizer patterns so non-language branches stay full precision | `nemotron_vl`, `diffusion_gemma` |
|
||||
| **Checkpoint mirror** | A mixed-precision map reproducing one published checkpoint exactly | `models/nvidia/Nemotron-3-*`, `models/nvidia/Mistral-Medium-3.5-128B-NVFP4` |
|
||||
|
||||
The numerics and standard exclusions are still inherited from `configs/`
|
||||
@@ -314,7 +314,7 @@ These quantize the **same layers** as the general recipes; only the
|
||||
*Why special:* identical scope/numerics to a general scheme, but a general
|
||||
recipe's default algorithm would overflow or regress here.
|
||||
|
||||
### Extra exclusions — `nemotron_vl`, `phi4mm`, `diffusion_gemma`
|
||||
### Extra exclusions — `nemotron_vl`, `diffusion_gemma`
|
||||
|
||||
Each of these is **numerically identical** to a general recipe. What makes them
|
||||
special is a model-local `disabled_quantizers.yaml` unit that *extends* the
|
||||
@@ -324,8 +324,6 @@ standard exclusions so a model-specific branch stays in full precision:
|
||||
`nvfp4_default-kv_fp8_cast` numerics, adding `*vision*`, `*image*`, `*radio*`,
|
||||
`*visual*`, `*encoder*`, `*model_encoder*` so only the language decoder is
|
||||
quantized.
|
||||
- **`phi4mm`** (Phi-4-Multimodal) — general `nvfp4_default-kv_fp8_cast`
|
||||
numerics, adding `*speech*`, `*audio*`, `*image*`, `*vision*`.
|
||||
- **`diffusion_gemma`** (block-diffusion encoder-decoder text LLM on a Gemma4
|
||||
MoE backbone) — general `nvfp4_experts_only-kv_fp8_cast` numerics, adding
|
||||
`*self_conditioning*`: the self-conditioning network is text-only and never
|
||||
|
||||
Reference in New Issue
Block a user