Files
Model-Optimizer/modelopt/torch
Keval MorabiaandClaude Opus 5 f13a7962aa Fix KV-cache scales dropped on Qwen Megatron-Core HF export (#2332)
### What does this PR do?

Type of change: Bug fix

Megatron-Core → HuggingFace export silently dropped KV-cache
quantization for every Qwen
architecture. A checkpoint calibrated with an FP8 (or NVFP4) KV cache
exported with
`kv_cache_quant_algo` unset, so the served model used an unquantized KV
cache while the recipe
and the Megatron checkpoint both said otherwise. Nothing warned.

`_GPTModelExporter` only emits KV-cache state for layers whose
architecture mapping defines a
`core_attention` rule:

```python
# modelopt/torch/export/unified_export_megatron.py
if hasattr(layer.self_attention, "core_attention") and "core_attention" in self.rules:
    self.rules["core_attention"](layer.self_attention.core_attention, layer_id, is_mtp=is_mtp)
```

`SelfAttentionScaling` was wired in `mcore_llama.py` and
`mcore_nemotron.py` but never in
`mcore_qwen.py`, so `_self_attention_scaling` never ran for Qwen:
`self.kv_cache_dtype` stayed
unset and `_gather_kv_cache_dtype()` returned `None`.

Adding the rule to `qwen3_causal_lm_export` and
`qwen25_causal_lm_export` covers all six Qwen
architectures — `qwen3vl_causal_lm_export` and
`qwen3_5_vl_causal_lm_export` derive from
`qwen3_causal_lm_export` through `with_language_model_prefix`, which
rewrites the prefix to
`model.language_model.layers.{}.self_attn.` automatically. GatedDeltaNet
linear-attention layers
have no `core_attention` submodule, so the existing `hasattr` guard
skips them.

**Known remaining gap, not addressed here:**
`deepseek_causal_lm_export`,
`gptoss_causal_lm_export` and `llama4_causal_lm_export` are missing the
same rule. I could not
validate those end to end, and DeepSeek's MLA uses different KV
projection names, so they need
their own change rather than a copy of this one.

### Usage

No API or flag change. The mapping now resolves for every Qwen
architecture:

```python
from modelopt.torch.export.plugins.mcore_common import all_mcore_hf_export_mapping

rule = all_mcore_hf_export_mapping["Qwen3_5MoeForConditionalGeneration"]["core_attention"]
print(rule.target_name_or_prefix)   # model.language_model.layers.{}.self_attn.
```

### Testing

New
`tests/gpu_megatron/torch/export/plugins/test_mcore_export_mappings.py`
(9 cases). It needs
no GPU but imports `mcore_common`, so it sits beside
`test_moe_layout_choice.py`, which is the
same shape. Confirmed the guard actually fires — reverting only
`mcore_qwen.py` gives
**6 failed / 3 passed** (the Llama and Nemotron controls pass either
way); with the fix,
**9 passed**.

End to end on a GB200 node in `nvcr.io/nvidia/nemo:26.08`: quantized
`Qwen/Qwen3.5-0.8B` with a
W4A16-NVFP4 MLP / FP8-attention / `kv_fp8_cast` recipe via
`examples/megatron_bridge/quantize.py`, then
`export_quantized_megatron_to_hf.py`.

| exported `hf_quant_config.json` | before | after | released
`nvidia/Qwen3.6-35B-A3B-NVFP4` |
| --- | --- | --- | --- |
| `kv_cache_quant_algo` | `None` | `FP8` | `FP8` |
| `k_scale` / `v_scale` tensors | 0 | 0 | 0 |

The absent scale tensors are correct for `kv_fp8_cast`:
`use_constant_amax` pins the amax to the
E4M3 maxbound, so `export_amax()` returns nothing for
`get_scaling_factor` and the runtime uses
the implicit 1.0 scale, while `get_kv_cache_dtype` still reports FP8
from `num_bits`. The
released checkpoint has exactly this shape, which is what the "after"
column was checked
against.

The rest of the exported layer map is unchanged by this PR and was
spot-checked against the
released checkpoint: NVFP4 W4A16 on the MLP projections, FP8 on
`linear_attn.{in_proj_qkv,in_proj_z,out_proj}` and
`self_attn.{q,k,v,o}_proj`, with
`in_proj_a` / `in_proj_b` / `conv1d` / `mtp.*` excluded.

`pre-commit run --files ...` passes on all changed files.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ <!-- Previously exported
checkpoints still load; re-export to gain the KV-cache field. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!-- 0.47.0 → Bug Fixes -->
- Did you get Claude approval on this PR?: ❌ <!-- Not run; happy to
trigger /claude review. -->

### Additional Information

Found while reproducing the `nvidia/Qwen3.6-35B-A3B-NVFP4` recipe
through
`examples/megatron_bridge/` rather than `examples/hf_ptq/`. Labeled
`cherry-pick-0.47.0`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Bug Fixes**
- Fixed Qwen checkpoint exports so calibrated FP8/NVFP4 KV-cache scales
are preserved.
- Exported checkpoints now retain the correct KV-cache quantization
settings, preventing unintentionally unquantized KV-cache serving.
- Improved KV-cache quantization mapping support across Qwen, Llama,
Nemotron, and Qwen VLM exports.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-04 17:57:54 +05:30
..
2026-06-27 01:00:25 +05:30
2026-08-11 11:51:02 -07:00
2026-06-27 01:00:25 +05:30
2026-06-27 01:00:25 +05:30