Files
Model-Optimizer/examples
Keval Morabia 5dde396bdf Fix vLLM 0.24+ compatibility: registry TypeError and MoE RoutedExperts port (#2054)
### What does this PR do?

Type of change: Bug fix

vLLM 0.24 (as shipped in `nemo:26.08`) reworked the fused-MoE layer,
which broke ModelOpt in two ways:

1. **`FusedMoE` became a factory function** returning a `MoERunner`
pipeline. Registering it put a plain function into
`QuantModuleRegistry`, so *every* later registry lookup raised
`TypeError: issubclass() arg 2 must be a class, ...` — taking down large
parts of `tests/gpu_megatron` (TE, Megatron chaining, MSE calibrator)
that have nothing to do with vLLM.
2. **Expert weights moved onto a `RoutedExperts` submodule** of
`MoERunner`, and `UnquantizedFusedMoEMethod` moved modules, so the MoE
fakequant path had no valid registration target.

Changes:

- `_DMRegistryCls.register` now asserts keys are `nn.Module` subclasses,
so a future upstream change fails at the registration site instead of as
a confusing `TypeError` at lookup.
- Register `RoutedExperts` (vLLM >= 0.24) while keeping the `FusedMoE` /
`SharedFusedMoE` registrations for older releases; both go through the
same `_QuantFusedMoEBase`. `MoERunner` calls `forward_modular` /
`forward_monolithic` directly (`RoutedExperts.forward` raises by
design), so those are hooked instead of `forward`.
- The fused-MoE kernel patch now covers `experts.triton_moe` in addition
to `fused_moe`. 0.24's launcher binds the kernel names at import time,
so patching only the defining module would leave fakequant **silently
inactive**.
- `UnquantizedFusedMoEMethod` is resolved from either module layout.
- `examples/vllm_serve/vllm_reload_utils.py`: quantizer module paths are
now `mlp.experts.routed_experts.*`, so HF→vLLM expert key mapping
inserts a matching `.routed_experts` infix when that layout is present.
- CI `gpu_vllm` now runs on **two** containers: `v0.24.0` (first release
with the FusedMoE-factory / `RoutedExperts` layout — `v0.24.1` was never
released, `nemo:26.08` ships a `0.24.1.dev0` build of the same layout)
and `v0.20.0`, which keeps the legacy `FusedMoE`/`SharedFusedMoE`
branches covered.

Test fixes for the newer vLLM (not product bugs):

- The tiny Llama fixture used `hidden_size=32 / 16 heads` →
`head_dim=2`, which `FLEX_ATTENTION` (the only backend available in this
image) rejects with `NYI: embedding dimension ... must be at least 16`,
killing the engine core at warmup. Now `head_dim=64`, matching the
Qwen3-MoE fixture.
- The FlashInfer metadata-builder stub used `causal=False`, which in
0.24 forces the FI-native path (`all_uses_trtllm = causal and ...`)
requiring workspace buffers and real wrapper planning. Keep it on the
all-TRTLLM path it was originally exercising; the stashed `_modelopt_*`
fields are path-independent.

### Usage

No API change — existing `mtq.quantize` / `examples/vllm_serve` flows
work unmodified on both old and new vLLM.

### Testing

All runs in the `nemo:26.08.rc3` container (vLLM
`0.24.1.dev0+gee0da84ab`, the same 0.24.1 the CI job now pins).

**Suites**

- `tests/gpu_vllm`: **73 passed, 1 skipped** (was 70 passed / 3 failed).
- `tests/gpu_megatron`: all pass (previously ~120 failures, all from the
registry `TypeError` — TE, Megatron chaining and MSE-calibrator tests
that never touch vLLM).
- `tests/unit/torch/opt/test_dynamic.py`: 2 passed, including the new
`test_register_rejects_non_module_classes` (rejects a factory function
and a non-`nn.Module` class, and asserts no partial registration).
- `pre-commit` clean on all touched files.

**MoE fakequant verified by module-tree probe, not just by test
assertions**

After `mtq.quantize(..., NVFP4_DEFAULT_CFG)` inside the vLLM worker,
every weight-owning module was enumerated on tiny DeepSeek-V3 (MLA +
routed MoE + shared experts) and tiny Qwen3-MoE:

```
model.layers.0.mlp.experts.routed_experts [QuantRoutedExperts]
  w13_input_quantizer=3.484  w13_weight_quantizer=0.0840
  w2_input_quantizer=0.1060  w2_weight_quantizer=0.0845
model.layers.0.mlp.shared_experts.gate_up_proj [QuantMergedColumnParallelLinear]  ✅
model.layers.0.mlp.shared_experts.down_proj    [QuantRowParallelLinear]           ✅
```

Weight amax being populated (not just input amax) means the `B is
self.w13_weight` identity check and the Parameter-swap weight-fakequant
branch actually execute through 0.24's kernel path — i.e.
`forward_modular`/`forward_monolithic` really are the live entry points
and the `experts.triton_moe` patch target is the one that fires.
Unquantized modules were only the expected ones: embeddings, RMSNorms,
MoE router `gate`, `lm_head`.

Registration parity vs. older vLLM: Row/Column/MergedColumn/QKV
`ParallelLinear` and all four attention types (`Attention`,
`CrossAttention`, `EncoderOnlyAttention`, `MLAAttention`) register
unchanged; `FusedMoE` → `RoutedExperts`; `SharedFusedMoE` has no
counterpart because the `shared_fused_moe` module no longer exists in
0.24 — shared experts are now a plain MLP whose linears we already
quantize (confirmed above).

**Known gaps (pre-existing, not regressions from this PR)**

- `DeepSeekV2FusedQkvAProjLinear` is not quantized: it subclasses
`MergedColumnParallelLinear` but overrides `forward`, so the registry's
shared-forward rule declines it. Pre-0.24 the equivalent (`q_a_proj` /
`kv_a_proj_with_mqa`) were `ReplicatedLinear`, which ModelOpt never
quantized — effective coverage is unchanged.
- MoE fakequant hooks only the Triton expert kernels;
FlashInfer/CUTLASS/DeepGEMM MoE backends bypass them (why the fixtures
pin `moe_backend="triton"`).
- `_setup` still requires a plain `UnquantizedFusedMoEMethod`; a
`FusedMoEModularMethod` swap (some DP/all2all configs) still asserts.

**Not covered by tests**

- `examples/vllm_serve/vllm_reload_utils.py` — the expert key mapping is
now asserted in `test_tiny_qwen3_moe_quantize` against the quantizer
module paths of a booted MoE model, so a stale infix fails loudly
instead of silently serving uncalibrated experts. The rest of the reload
path is still inspection-only. Note the registry key moved
`vllm_FusedMoE` → `vllm_RoutedExperts` and quantizer paths gained
`.routed_experts`, so a `modelopt_state` saved under an older vLLM will
not restore onto 0.24 as-is.
- The legacy `FusedMoE`/`SharedFusedMoE` branches are covered by the
second CI entry; the `v0.20.0` job is green on this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — all new paths are
feature-detected; older vLLM keeps the `FusedMoE` registration.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — existing
`tests/gpu_vllm` coverage exercises the new registration path.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌

### Additional Information

Found while bumping the Megatron test environment from `nemo:26.06` to
`nemo:26.08.rc3`.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added compatibility for newer vLLM MoE implementations, module
layouts, and routed-expert configurations.
* Improved model reload support for models using routed-expert
submodules.
* **Bug Fixes**
* Improved detection and patching of vLLM MoE execution paths across
supported configurations.
* Registry validation now rejects invalid module registrations without
partially applying changes.
* **Tests**
* Expanded GPU coverage for vLLM 0.24.0, dynamic module validation, and
causal attention metadata paths.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-08-04 09:37:06 -07:00
..
2026-06-08 22:11:37 +00:00
2026-06-27 01:00:25 +05:30
2026-06-27 01:00:25 +05:30