mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: Bug fix vLLM 0.24 (as shipped in `nemo:26.08`) reworked the fused-MoE layer, which broke ModelOpt in two ways: 1. **`FusedMoE` became a factory function** returning a `MoERunner` pipeline. Registering it put a plain function into `QuantModuleRegistry`, so *every* later registry lookup raised `TypeError: issubclass() arg 2 must be a class, ...` — taking down large parts of `tests/gpu_megatron` (TE, Megatron chaining, MSE calibrator) that have nothing to do with vLLM. 2. **Expert weights moved onto a `RoutedExperts` submodule** of `MoERunner`, and `UnquantizedFusedMoEMethod` moved modules, so the MoE fakequant path had no valid registration target. Changes: - `_DMRegistryCls.register` now asserts keys are `nn.Module` subclasses, so a future upstream change fails at the registration site instead of as a confusing `TypeError` at lookup. - Register `RoutedExperts` (vLLM >= 0.24) while keeping the `FusedMoE` / `SharedFusedMoE` registrations for older releases; both go through the same `_QuantFusedMoEBase`. `MoERunner` calls `forward_modular` / `forward_monolithic` directly (`RoutedExperts.forward` raises by design), so those are hooked instead of `forward`. - The fused-MoE kernel patch now covers `experts.triton_moe` in addition to `fused_moe`. 0.24's launcher binds the kernel names at import time, so patching only the defining module would leave fakequant **silently inactive**. - `UnquantizedFusedMoEMethod` is resolved from either module layout. - `examples/vllm_serve/vllm_reload_utils.py`: quantizer module paths are now `mlp.experts.routed_experts.*`, so HF→vLLM expert key mapping inserts a matching `.routed_experts` infix when that layout is present. - CI `gpu_vllm` now runs on **two** containers: `v0.24.0` (first release with the FusedMoE-factory / `RoutedExperts` layout — `v0.24.1` was never released, `nemo:26.08` ships a `0.24.1.dev0` build of the same layout) and `v0.20.0`, which keeps the legacy `FusedMoE`/`SharedFusedMoE` branches covered. Test fixes for the newer vLLM (not product bugs): - The tiny Llama fixture used `hidden_size=32 / 16 heads` → `head_dim=2`, which `FLEX_ATTENTION` (the only backend available in this image) rejects with `NYI: embedding dimension ... must be at least 16`, killing the engine core at warmup. Now `head_dim=64`, matching the Qwen3-MoE fixture. - The FlashInfer metadata-builder stub used `causal=False`, which in 0.24 forces the FI-native path (`all_uses_trtllm = causal and ...`) requiring workspace buffers and real wrapper planning. Keep it on the all-TRTLLM path it was originally exercising; the stashed `_modelopt_*` fields are path-independent. ### Usage No API change — existing `mtq.quantize` / `examples/vllm_serve` flows work unmodified on both old and new vLLM. ### Testing All runs in the `nemo:26.08.rc3` container (vLLM `0.24.1.dev0+gee0da84ab`, the same 0.24.1 the CI job now pins). **Suites** - `tests/gpu_vllm`: **73 passed, 1 skipped** (was 70 passed / 3 failed). - `tests/gpu_megatron`: all pass (previously ~120 failures, all from the registry `TypeError` — TE, Megatron chaining and MSE-calibrator tests that never touch vLLM). - `tests/unit/torch/opt/test_dynamic.py`: 2 passed, including the new `test_register_rejects_non_module_classes` (rejects a factory function and a non-`nn.Module` class, and asserts no partial registration). - `pre-commit` clean on all touched files. **MoE fakequant verified by module-tree probe, not just by test assertions** After `mtq.quantize(..., NVFP4_DEFAULT_CFG)` inside the vLLM worker, every weight-owning module was enumerated on tiny DeepSeek-V3 (MLA + routed MoE + shared experts) and tiny Qwen3-MoE: ``` model.layers.0.mlp.experts.routed_experts [QuantRoutedExperts] w13_input_quantizer=3.484 w13_weight_quantizer=0.0840 w2_input_quantizer=0.1060 w2_weight_quantizer=0.0845 model.layers.0.mlp.shared_experts.gate_up_proj [QuantMergedColumnParallelLinear] ✅ model.layers.0.mlp.shared_experts.down_proj [QuantRowParallelLinear] ✅ ``` Weight amax being populated (not just input amax) means the `B is self.w13_weight` identity check and the Parameter-swap weight-fakequant branch actually execute through 0.24's kernel path — i.e. `forward_modular`/`forward_monolithic` really are the live entry points and the `experts.triton_moe` patch target is the one that fires. Unquantized modules were only the expected ones: embeddings, RMSNorms, MoE router `gate`, `lm_head`. Registration parity vs. older vLLM: Row/Column/MergedColumn/QKV `ParallelLinear` and all four attention types (`Attention`, `CrossAttention`, `EncoderOnlyAttention`, `MLAAttention`) register unchanged; `FusedMoE` → `RoutedExperts`; `SharedFusedMoE` has no counterpart because the `shared_fused_moe` module no longer exists in 0.24 — shared experts are now a plain MLP whose linears we already quantize (confirmed above). **Known gaps (pre-existing, not regressions from this PR)** - `DeepSeekV2FusedQkvAProjLinear` is not quantized: it subclasses `MergedColumnParallelLinear` but overrides `forward`, so the registry's shared-forward rule declines it. Pre-0.24 the equivalent (`q_a_proj` / `kv_a_proj_with_mqa`) were `ReplicatedLinear`, which ModelOpt never quantized — effective coverage is unchanged. - MoE fakequant hooks only the Triton expert kernels; FlashInfer/CUTLASS/DeepGEMM MoE backends bypass them (why the fixtures pin `moe_backend="triton"`). - `_setup` still requires a plain `UnquantizedFusedMoEMethod`; a `FusedMoEModularMethod` swap (some DP/all2all configs) still asserts. **Not covered by tests** - `examples/vllm_serve/vllm_reload_utils.py` — the expert key mapping is now asserted in `test_tiny_qwen3_moe_quantize` against the quantizer module paths of a booted MoE model, so a stale infix fails loudly instead of silently serving uncalibrated experts. The rest of the reload path is still inspection-only. Note the registry key moved `vllm_FusedMoE` → `vllm_RoutedExperts` and quantizer paths gained `.routed_experts`, so a `modelopt_state` saved under an older vLLM will not restore onto 0.24 as-is. - The legacy `FusedMoE`/`SharedFusedMoE` branches are covered by the second CI entry; the `v0.20.0` job is green on this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — all new paths are feature-detected; older vLLM keeps the `FusedMoE` registration. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — existing `tests/gpu_vllm` coverage exercises the new registration path. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ ### Additional Information Found while bumping the Megatron test environment from `nemo:26.06` to `nemo:26.08.rc3`. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added compatibility for newer vLLM MoE implementations, module layouts, and routed-expert configurations. * Improved model reload support for models using routed-expert submodules. * **Bug Fixes** * Improved detection and patching of vLLM MoE execution paths across supported configurations. * Registry validation now rejects invalid module registrations without partially applying changes. * **Tests** * Expanded GPU coverage for vLLM 0.24.0, dynamic module validation, and causal attention metadata paths. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>