[NVBUG: 5612606] Clear GPU cache for large models layer quantization during export (#497)

## What does this PR do?

**Type of change:** Bug fix

**Overview:** ?

For large models like llama4 maverick, the stacked weights to fp8
conversion might hit OOM. This change aim to fix that.

---------

Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
This commit is contained in:
Chenjie Luo
2025-11-04 18:46:58 +00:00
committed by GitHub
parent 5f0ef3b310
commit 9a85e4922b
+3
View File
@@ -37,6 +37,7 @@ from modelopt.torch.quantization.utils import (
quantizer_attr_names,
weight_attr_names,
)
from modelopt.torch.utils import clear_cuda_cache
from ..quantization.nn import SequentialQuantizer, TensorQuantizer
from .model_config import (
@@ -763,6 +764,8 @@ def to_quantized_weight(
if weight.dim() == 3:
# for MOE stacked weights
# Clear GPU cache to avoid pontential GPU OOM issues for large models.
clear_cuda_cache()
return (weight / weights_scaling_factor.unsqueeze(-1)).to(torch.float8_e4m3fn)
return (weight / weights_scaling_factor).to(torch.float8_e4m3fn)