Frida HouandClaude Opus 5 1c4cde7788 Fix: Release per-layer expert weights in layerwise export under offload (#2466)
### What does this PR do?

Type of change: Bug fix

Layerwise export leaks one layer's worth of quantized tensors per layer
on offloaded models, and dies of OOM partway through a large MoE. Two
things accumulate, both because the offload window cannot reclaim what
the export pass adds inside it.

**1. Per-expert holder modules.** `_export_fused_experts` splits a fused
MoE experts module into per-expert holders and attaches them to the live
model:

```python
proj = nn.Module();  proj.weight = wrapper.weight   # packed U8
expert.add_module(proj_name, proj)
module.add_module(str(idx), expert)
```

They are plain `nn.Module`s built *inside* the weight-access window, so
they carry no accelerate `_hf_hook`.
`weight_access_and_writeback_context` closes by iterating the modules it
collected at entry and calling `hook.post_forward()` through each one's
own offload hook — the holders satisfy neither condition, so nothing
returns them to meta.

**2. Scale buffers.** `_export_quantized_weight` registers
`weight_scale` / `weight_scale_2` / `input_scale` on the layer's
pre-existing, hooked sub-modules, and `AlignDevicesHook.post_forward`
runs with `offload_buffers=False`:

```
after post_forward: {'weight': 'meta', 'weight_scale': 'cpu'}
```

The packed weight goes back to meta; the scales do not. For
`_QuantMoELinear` models this half is the whole leak on its own —
`_reconstruct_fused_moe_linear` restacks every expert's scales into one
`register_buffer` on the hooked wrapper.

A whole-model export never notices: one pass, write the state dict,
exit. Layerwise runs the same pass once per decoder layer, so every
finished layer stays resident.

### Measurements

Through unmodified `examples/hf_ptq/hf_ptq.py` on a Qwen3.5-MoE-shaped
model (10 layers, 64 experts), printing `torch.cuda.memory_allocated()`
after each exported layer:

| placement | per layer | over 9 layers |
| --- | --- | --- |
| offload | +0.052 GiB | 0.579 → 1.047 GiB |
| resident | -0.135 GiB | falls, as designed |

+0.052 GiB is exactly one layer's quantized experts: packed U8 100.7M/2
= 0.047 GiB plus FP8 block scales 100.7M/16 = 0.006 GiB.

At scale it is fatal rather than wasteful. Qwen/Qwen3.8-2.4T-A95B (92
layers, 512 experts) leaks 12.9 GB packed + 1.6 GB scales per layer, so
92 layers want 1.33 TB that no budget on a 283 GB card or 952 GB host
absorbs. The run died of CUDA OOM at layer 16/92 with
`--max_gpu_memory_gb 240`, and at `--max_gpu_memory_gb 30` leaked the
same 15 GB/layer onto the host instead.

### The fix

`_release_exported_tensors` (`model_utils.py`) is a context manager that
snapshots each sub-module's child-module and buffer names on entry, and
on exit drops whatever appeared. Persisting happens *inside* the block,
so "release only once it is on disk" is structural rather than a
comment.

Both packing sites use it: `LayerwiseExporter.export_layer` and the
offload decoder loop in `_export_transformers_checkpoint_streaming`.

The streaming writer had solved the same leak inline with a heuristic —
null every CUDA buffer, and every CUDA parameter on a hook-less module —
and that block is deleted in favour of the shared helper. Keying on
*what the pass added* rather than on device and hook presence drops two
assumptions that only held for a terminal, offloaded export: it no
longer nulls buffers the layer already had, nor parameters of
sub-modules accelerate simply did not hook. That is also what makes it
safe for the layerwise path, where resident models are supported and the
model outlives the export.

Deliberately out of scope: the FSDP2 per-unit loop in
`collect_export_tensors` keeps its per-unit holders. That predates this
PR, this PR does not touch that loop, and closing it needs its own
change and its own test.

### Usage

No API change. Existing layerwise export under offload simply stops
growing:

```bash
python examples/hf_ptq/hf_ptq.py \
    --pyt_ckpt_path Qwen/Qwen3.8-2.4T-A95B \
    --qformat nvfp4 \
    --export_path /path/to/export \
    --max_gpu_memory_gb 240
```

### Testing

- `tests/unit/torch/export` and `tests/unit/torch/quantization` — 1246
passed
- `tests/gpu/torch/export/test_layerwise_export.py` +
`test_offload_export.py` — 35 passed
- `cuda_alloc` over 10 offloaded layers: 0.526 → 0.518 GiB (-0.008), was
+0.468
- Exported checkpoint byte-identical to the unfixed run: 7841 tensors, 0
mismatches, max abs diff 0.0; `hf_quant_config.json` / `config.json` /
index identical
- Full Qwen3.8-2.4T-A95B PTQ then completed all 92 layers: peak GPU 117
GB of 283, peak host RSS 71 GB of 952, flat across 50 consecutive layers
at 23-26 s/layer

Coverage added: the existing `test_export_creates_per_expert_submodules`
now runs the export inside the context manager and asserts the holders
are gone on exit, and a new test in `test_offload_export.py` pins the
`offload_buffers=False` behaviour the buffer half exists for —
export-registered scales dropped, pre-existing buffers untouched.

`tests/gpu/torch/export/test_fsdp2_export.py` reports 34 failures in my
environment. They are **pre-existing and unrelated**: the same 34 fail
identically on this branch and on the merge-base (`2b1f33d0ef`), with
byte-identical failure sets and runtimes within 3 s. All 34 are `Failed:
Timeout (>120.0s)` from hung NCCL collectives, with zero assertion
failures.

The GPU export suites and the whole-model measurements above were run at
`2999d7cfb7`. The two commits since — swapping the holder marker for a
child-name diff, and moving the helper to `model_utils` — are covered by
the unit suites; a GPU re-run before merge is worthwhile.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — one new offload test for
the buffer half, and the existing fused-experts export test now covers
holder release
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — `layerwise.export_dir` is new in the unreleased 0.48.0, so this
bug was introduced and fixed within the same cycle
- Did you get Claude approval on this PR?: ✅

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 14:43:12 -07:00
2026-09-10 17:30:37 +00:00
2026-09-10 17:30:37 +00:00
…

Banner image

NVIDIA Model Optimizer

Documentation version license

Documentation | Roadmap | Announcement Blogs


NVIDIA Model Optimizer (referred to as Model Optimizer, or ModelOpt) is a library comprising state-of-the-art model optimization techniques including quantization, pruning, Neural Architecture Search (NAS), distillation, speculative decoding and sparsity to accelerate models.

[Input] Model Optimizer currently supports inputs of a Hugging Face, PyTorch or ONNX model.

[Optimize] Model Optimizer provides Python APIs for users to easily compose the above model optimization techniques and export an optimized quantized checkpoint. Model Optimizer is also integrated with NVIDIA Megatron-Bridge, Megatron-LM and Hugging Face Accelerate for training required inference optimization techniques.

[Export for deployment] Seamlessly integrated within the NVIDIA AI software ecosystem, the quantized checkpoint generated from Model Optimizer is ready for deployment in downstream inference frameworks like SGLang, TensorRT-LLM, TensorRT, or vLLM. The unified Hugging Face export API now supports both transformers and diffusers models.

Latest News

Previous News

Install

To install stable release packages for Model Optimizer with pip from PyPI:

pip install -U nvidia-modelopt[all]

Model Optimizer will download and install additional third-party open source software projects. Review the license terms of these open source projects before use.

To install from source in editable mode with all development dependencies or to use the latest features, run:

# Clone the Model Optimizer repository
git clone git@github.com:NVIDIA/Model-Optimizer.git
cd Model-Optimizer

pip install -e .[dev]

You can also directly use NVIDIA container images, which have Model Optimizer pre-installed:

  • nvcr.io/nvidia/pytorch:<version>-py3
  • nvcr.io/nvidia/nemo:<version>
  • nvcr.io/nvidia/tensorrt-llm/release:<version>

Before pulling and using the container images, please review their respective license terms. Make sure to upgrade Model Optimizer to the latest version as described above. Visit our installation guide for more fine-grained control on installed dependencies or for alternative docker images and environment variables to setup.

Techniques

Technique Description Examples Docs
Post Training Quantization Compress model size by 2x-4x, speeding up inference while preserving model quality! [HF LLMs / VLMs] [Megatron-Bridge LLMs / VLMs] [Diffusers] [ONNX] [Windows] [docs]
Quantization Aware Training / Distillation Refine accuracy of quantized models even further with a few training steps! [Hugging Face] [Megatron-Bridge] [docs]
Pruning Reduce your model parameters or memory footprint and accelerate inference by removing unnecessary weights! [General] [Megatron-Bridge]
Distillation Reduce deployment model size by teaching small models to behave like larger models! [Hugging Face] [Megatron-Bridge] [Megatron-LM] [docs]
Speculative Decoding Train draft modules to predict extra tokens during inference! [Hugging Face] [Megatron-LM] [docs]
Sparsity Efficiently compress your model by storing only its non-zero parameter values and their locations [Hugging Face] [docs]

Pre-Quantized Checkpoints

Resources

Model Support Matrix

Model Type Support Matrix
LLM / VLM Quantization View Support Matrix
Diffusers Quantization View Support Matrix
ONNX Quantization View Support Matrix
Windows Quantization View Support Matrix
Quantization Aware Training View Support Matrix
Pruning View Support Matrix
Distillation View Support Matrix
Speculative Decoding View Support Matrix

Deprecation Policy

Model Optimizer follows a structured approach to managing deprecated features:

  • Communication: Deprecation notices are documented in the Changelog. Deprecated items include source code statements indicating deprecation timing, with runtime warnings issued upon use.
  • Migration Period: Since Model Optimizer is still pre-1.0, we provide a 1-release (~1-month) migration period after deprecation. During this window, deprecated features continue functioning while issuing warnings.
  • Scope: The policy addresses both complete deprecations (entire APIs removed) and partial ones (specific parameters removed while methods remain).
  • Removal: Following the migration period, deprecated elements are removed in alignment with semantic versioning standards, potentially including breaking changes in minor version updates while Model Optimizer remains in 0.x.

Citation

If you use NVIDIA Model Optimizer in your research, please cite it as follows:

@misc{nvidia-modelopt,
  author       = {{NVIDIA Corporation}},
  title        = {{NVIDIA Model Optimizer}},
  howpublished = {\url{https://github.com/NVIDIA/Model-Optimizer}},
  year         = {2024--2026},
  note         = {GitHub repository}
}

Contributing

Model Optimizer is now open source! We welcome any feedback, feature requests and PRs. Please read our Contributing guidelines for details on how to contribute to this project.

AI Agents

ModelOpt's agent skills can be installed from this repository and used in any workspace.

Claude Code

claude plugin marketplace add https://github.com/NVIDIA/Model-Optimizer.git
claude plugin install modelopt@modelopt

Codex

codex plugin marketplace add https://github.com/NVIDIA/Model-Optimizer.git

Then open /plugins, select the modelopt marketplace, and install modelopt. Contributors can also use the skills directly from a checkout. See the agent tooling notes.

Top Contributors

Contributors

Happy optimizing!

S
Description
GitHub Trending: NVIDIA/Model-Optimizer
Readme Multiple Licenses
1.2 GiB
Languages
Python 97.3%
Shell 1.3%
Cuda 0.8%
Jinja 0.3%
C++ 0.2%