mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Fail fast on non-finite AutoQuantize output gradients (#2432)
### What does this PR do? Type of change: Bug fix Fail fast when AutoQuantize receives non-finite output gradients, before accumulating sensitivity scores. The error names the affected module and suggests checking the model, data, and loss; it also identifies cuDNN SDPA backward on fully masked rows as one possible cause and gives an explicit retry workaround. Unlike the earlier revision, this does not disable cuDNN or change any attention backend settings. Invalid gradients are not zeroed or ignored, and there is no automatic retry. ### Usage No API or recipe changes. For the reproduced cuDNN failure, the caller can explicitly set `torch.backends.cuda.enable_cudnn_sdp(False)` before a fresh AutoQuantize run. ### Testing - AutoQuantize unit suite: **110 passed**. Coverage includes NaN and positive/negative infinity, module diagnostics, preventing invalid score accumulation, model-state cleanup, and unchanged SDPA backend settings. - Real-model E2E on **four GB300 GPUs**, Qwen/Qwen3.6-35B-A3B, main `8025a3dc5481129aa21fef99cb13a879e1b5847e` plus this patch, batch size 8, 512 calibration samples, and `w4a16_nvfp4_fp8_at_6p0bits-active_moe.yaml`: - Default backend: the new diagnostic fired at `model.language_model.layers.39.self_attn.q_proj`; cuDNN remained enabled and no quantized model was exported. Expected-error check passed. - Explicit cuDNN-disabled fresh run: both 64-batch calibration passes, all 16 scoring batches, optimization at **5.99 effective bits**, and checkpoint export completed with exit code 0. Verified all three indexed safetensors shards and quantization configuration; the index contains 93,563 tensors. - All applicable pre-commit checks and `git diff --check` passed. ### Before your PR is "*Ready for review*" Contributor guidelines and security guidance followed; commits are signed and signed off. - Is this change backward compatible?: Yes; finite-gradient behavior and backend settings are unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A; neither added. - Did you write any new necessary tests?: Yes. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: Yes, 0.48.0 bug fixes. - Did you get Claude approval on this PR?: No; awaiting review. ### Additional Information This improves error reporting rather than fixing the upstream cuDNN kernel. Blackwell-specificity is not established. Checkpoint deployment/reload was not tested. A calibration-only checkpoint-resume attempt completed scoring but encountered a separate `candidate_stats` KeyError. That issue is outside this patch; the successful export validation above used a fresh run without search-state resume. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * AutoQuantize now fails fast when output gradients contain non-finite values, with an actionable error identifying the affected module. * Attention backend settings are preserved and restored after successful runs and failures. * Model state is restored when sensitivity scoring encounters an error during setup or execution. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
This commit is contained in:
@@ -40,6 +40,7 @@ Changelog
|
||||
|
||||
- Fix ``examples/megatron_bridge/export_quantized_megatron_to_hf.py`` storing the MoE router at Megatron's ``moe_router_dtype``, which is a routing *compute* dtype, not a storage one. The router now exports at the export ``dtype`` like every other unquantized weight, matching what ``hf_ptq.py`` and the released NVFP4 checkpoints contain; pass ``moe_router_dtype`` to ``export_mcore_gpt_to_hf`` explicitly if you want the old fp32 storage.
|
||||
- Fix unified Megatron export writing a second, unreferenced copy of the vocab embedding when a model with MTP layers is exported with pipeline parallelism. The duplicate was never loaded but inflated the checkpoint by the size of the embedding (about 1 GB for Qwen3.6-35B-A3B); re-export to reclaim the space.
|
||||
- Fail fast on non-finite AutoQuantize output gradients with an actionable error before accumulating sensitivity scores, without changing attention backend settings.
|
||||
- Fix ONNX INT8 entropy calibration failing or producing invalid quantization parameters for FP16 activations.
|
||||
- Fix ``--use_fsdp2`` HuggingFace checkpoint export gathering the whole model onto rank 0, which made export the dominant phase of a PTQ run and could exhaust host memory on large models. The model is now split into per-decoder-layer units dealt round-robin across ranks; each rank gathers every unit but keeps, packs, and writes only the ones it owns, so a rank buffers roughly ``model / world_size`` instead of the whole checkpoint, and rank 0 writes the combined index. Export configurations that cannot be split this way now raise instead of producing a mismatched checkpoint: FSDP2 combined with another DTensor parallelism (for example FSDP2 + tensor parallel on a 2-D mesh; HSDP is supported), models whose decoder layers cannot be discovered, a decoder layer object reused across layers, and a module that holds the decoder layers while owning parameters of its own.
|
||||
- Speed up ``mtq.quantize`` on FSDP2-sharded fused-MoE models. Promoting static-block weight quantizers gathered each expert's slice of the fused weight across ranks even though only quantizer state is read, adding a collective per expert to calibration.
|
||||
|
||||
@@ -1604,6 +1604,16 @@ class _AutoQuantizeGradientScoringSession(_AutoQuantizeBackwardScoringSession):
|
||||
|
||||
def _accumulate_scores(self, module, invocation_diffs, grad_output) -> None:
|
||||
"""Accumulate scores for the invocation that produced ``grad_output``."""
|
||||
if not torch.isfinite(grad_output).all():
|
||||
module_name = next(
|
||||
name for name, child in self.model.named_modules() if child is module
|
||||
)
|
||||
raise RuntimeError(
|
||||
f"AutoQuantize: Non-finite output gradients in module '{module_name or '<root>'}'. "
|
||||
"Cannot compute reliable sensitivity scores. Check the model, data, and loss. "
|
||||
"cuDNN SDPA backward on fully masked attention rows is one possible cause; "
|
||||
"try torch.backends.cuda.enable_cudnn_sdp(False) before rerunning auto_quantize."
|
||||
)
|
||||
for hparam, output_diffs in invocation_diffs.items():
|
||||
for recipe, output_diff in output_diffs.items():
|
||||
importance = hparam._importance_dict[recipe][module]
|
||||
|
||||
@@ -16,6 +16,7 @@
|
||||
import copy
|
||||
import io
|
||||
import warnings
|
||||
from contextlib import nullcontext
|
||||
from types import SimpleNamespace
|
||||
|
||||
import pytest
|
||||
@@ -1135,6 +1136,75 @@ def test_gradient_scoring_restores_model_after_failure():
|
||||
assert hparam.active == hparam.original
|
||||
|
||||
|
||||
@pytest.mark.parametrize("cudnn_enabled", [False, True])
|
||||
@pytest.mark.parametrize("fail", [False, True])
|
||||
def test_backward_scoring_session_preserves_sdpa_backends(cudnn_enabled, fail):
|
||||
original = torch.backends.cuda.cudnn_sdp_enabled()
|
||||
others = (
|
||||
torch.backends.cuda.flash_sdp_enabled(),
|
||||
torch.backends.cuda.mem_efficient_sdp_enabled(),
|
||||
torch.backends.cuda.math_sdp_enabled(),
|
||||
)
|
||||
try:
|
||||
torch.backends.cuda.enable_cudnn_sdp(cudnn_enabled)
|
||||
with (
|
||||
pytest.raises(RuntimeError, match="scoring failed") if fail else nullcontext(),
|
||||
_AutoQuantizeGradientScoringSession(torch.nn.Identity(), [], lambda *_: False),
|
||||
):
|
||||
assert torch.backends.cuda.cudnn_sdp_enabled() == cudnn_enabled
|
||||
assert others == (
|
||||
torch.backends.cuda.flash_sdp_enabled(),
|
||||
torch.backends.cuda.mem_efficient_sdp_enabled(),
|
||||
torch.backends.cuda.math_sdp_enabled(),
|
||||
)
|
||||
if fail:
|
||||
raise RuntimeError("scoring failed")
|
||||
finally:
|
||||
restored = torch.backends.cuda.cudnn_sdp_enabled()
|
||||
torch.backends.cuda.enable_cudnn_sdp(original)
|
||||
assert restored == cudnn_enabled
|
||||
|
||||
|
||||
@pytest.mark.parametrize("bad_gradient", [float("nan"), float("inf"), -float("inf")])
|
||||
def test_auto_quantize_fails_fast_on_nonfinite_gradients(bad_gradient):
|
||||
model = SimpleLinear()
|
||||
original_cudnn = torch.backends.cuda.cudnn_sdp_enabled()
|
||||
|
||||
def loss_func(output, _data):
|
||||
output.register_hook(lambda grad: torch.full_like(grad, bad_gradient))
|
||||
return output.sum()
|
||||
|
||||
with pytest.raises(RuntimeError, match="Non-finite output gradients in module") as exc_info:
|
||||
mtq.auto_quantize(
|
||||
model,
|
||||
constraints={"effective_bits": 12.0},
|
||||
quantization_formats=[mtq.INT8_DEFAULT_CFG],
|
||||
data_loader=[model.get_input()],
|
||||
forward_step=lambda model, batch: model(batch),
|
||||
loss_func=loss_func,
|
||||
num_calib_steps=1,
|
||||
num_score_steps=1,
|
||||
)
|
||||
|
||||
assert "torch.backends.cuda.enable_cudnn_sdp(False)" in str(exc_info.value)
|
||||
assert torch.backends.cuda.cudnn_sdp_enabled() == original_cudnn
|
||||
assert any(
|
||||
f"module '{name}'" in str(exc_info.value)
|
||||
for name, module in model.named_modules()
|
||||
if getattr(module, "_hparams_for_scoring", [])
|
||||
)
|
||||
for name, module in model.named_modules():
|
||||
for hparam in getattr(module, "_hparams_for_scoring", []):
|
||||
assert "forward" not in module.__dict__
|
||||
assert hparam.active == hparam.original
|
||||
for recipe in hparam.choices:
|
||||
importance = hparam._importance_dict[recipe][module]
|
||||
if f"module '{name}'" in str(exc_info.value):
|
||||
assert importance is None
|
||||
else:
|
||||
assert importance is None or torch.isfinite(importance).all()
|
||||
|
||||
|
||||
def test_backward_scoring_session_restores_partial_setup():
|
||||
model = torch.nn.Sequential(torch.nn.Linear(4, 4))
|
||||
score_module = model[0]
|
||||
|
||||
Reference in New Issue
Block a user