mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? **Type of change:** New feature — per-model infrastructure (with behavior changes, see below) Starts `modelopt/torch/models/`: somewhere for ModelOpt to keep what it knows about a model, keyed by HF model type (`config.model_type`), one package per type (`<model_type>/specs.py`). Importing the package registers every spec; consumers call `get_spec(model_type)` / `match_moe_block(module, model_type)` and read the fields. The point is the infrastructure, not the migration. Today ModelOpt's per-model knowledge is scattered across if/elif chains in whichever subsystem happened to need it, so the same fact gets restated per subsystem and drifts. A model's MoE block classes, its expert projection naming, which of its norms store `w - 1` — these are facts about the *model*, and several subsystems want them. This PR gives them one home and one lookup. It is a kick-off, so it is deliberately narrow: specs plus the first consumer. **Export is that first consumer**, which is why most of the changed lines are on the export side — not because this is an export refactor. Quantization, speculative decoding and the rest keep their own tables for now; the per-model directory is where their sections land later, and `specs.py` is just the first thing in it. | Data the first consumer moved in | Spec field | Consumers | |---|---|---| | MoE block classes and expert linear naming | `MoESpec` | `get_expert_linear_names`, `get_experts_list`, `is_moe`, `sync_moe_gate_up_amax` | | grouped expert-export support set | `ExportSpec.grouped_expert_export` | `get_experts_list` | | AWQ `pre_quant_scale` fusion rules | `ExportSpec.pqs_fuse_rules` | `fuse_prequant_to_linear` | | weight-plus-one norm class names | `ExportSpec.weight_plus_one_norm_names` | `_layernorm_uses_weight_plus_one` | `ModelSpec` holds each concern as a separate, optional section, and the split is what makes it extensible: **topic sections** hold architecture facts any subsystem can read (`MoESpec` — block classes, expert projection naming, how the experts are stored), **subsystem sections** hold one subsystem's own data and policy (`ExportSpec` — AWQ fusion rules, weight-plus-one norms, whether grouped expert export is validated for this model). A new subsystem adds a section rather than a table. A section is `None` when the model has nothing to say about it, so a dense model carries `moe_spec=None` rather than an empty one. Two further fields record where a model's classes come from at all: `modeling_source` (`transformers` or `remote_code`) and `min_transformers_version`. ### Behavior changes **1. Lookup key: `type(root_model).__name__.lower()` → `config.model_type`.** The old key changes after `quantize` wraps the model, forcing substring matching; `model_type` is stable, so `ExportContext` carries it and lookups are exact. Matching is also tightened to case-insensitive **exact** names against the module's **MRO**, so quantized subclasses still match via their base class without substring false positives. **2. ⚠️ `get_expert_linear_names` raises instead of guessing.** Unmatched MoE blocks previously fell through to `["w1", "w2", "w3"]` (Mixtral naming); they now raise `NotImplementedError` telling you to register a `ModelSpec`. The blast radius depends on the transformers release, because only *iterable* expert layouts consult the spec — a fused expert container is resolved by the structural first-projection check and never asks for naming. Measured against both ends of the support matrix: | transformers | Unregistered MoE families `is_moe` admits | Reach the spec lookup | Actually affected | |---|---|---|---| | 5.14.1 (`tf_latest`) | 19 | **0** — all fused | none | | 4.57.6 (`tf_min`) | 9 | 7 iterable | **2** | On 4.57, five of the seven (`ernie4_5_moe`, `flex_olmo`, `jamba`, `olmoe`, `qwen3_omni_moe`) name their experts `gate_proj`/`up_proj`/`down_proj`, so the `w1` fallback already failed with `AttributeError` on `main` — for them this trades an unhelpful error for one that names the fix. The other two, **`minimax` and `phimoe`**, are Mixtral-derived with `w1`/`w2`/`w3` experts and did export on `main`, so they are now **registered** rather than left to regress. Both only take the per-expert path on transformers 4; on 5 their experts are fused (`MiniMaxExperts`, `PhimoeExperts`) and the structural check resolves them, which their specs decline to answer for. **So no model that exported correctly before this change regresses.** This is backward-incompatible and has a `CHANGELOG` entry. **3. DeepSeek-V3 and V4 are now registered, and newly visible to the MoE path.** Neither was reachable before: `DeepseekV3MoE` (like `DeepseekV2Moe` and `DeepseekV32MoE`) is invisible to `is_moe` — its class name does not end in `SparseMoeBlock` and it calls its router `gate`, so neither the name test nor the structural `router`+`experts` test matches. `DeepseekV4SparseMoeBlock` was detected by name but had no spec to resolve expert naming from. Registering `block_names` is what puts V3 on the MoE path at all, so this is added coverage rather than a fix to existing behavior. `deepseek` stays registered alongside them: it describes the remote-code `DeepseekMoE` block of DeepSeek-MoE/V1, and its spec is the only thing making that block detectable. **4. Fused expert containers are skipped in `get_experts_list` instead of crashing.** transformers 5 replaced several iterable expert `ModuleList`s with a single module holding 3-D parameters, while the specs still describe the transformers 4 iterable layout — a spec cannot tell the two apart, only the module can. The AWQ/SVDQuant resmooth pass in `requantize_resmooth_fused_llm_layers` reached `len(module.experts)` on a module with no `__len__` and died with `TypeError` mid-export. **Mixtral already hits this on `main`**; registering `deepseek_v3` only made an existing bug easier to reach. The skip is scoped to specs that claim iterable experts, so a layout the spec calls unsupported (DBRX, whose per-expert linears live under `experts.mlp`) still fails loudly rather than silently dropping resmoothing. Everything else is behavior-preserving, pinned by tests: - **`sync_moe_gate_up_amax` keeps generic coverage.** Quantization admits MoE blocks *structurally*, so unregistered families (Olmoe, Jamba, MiniMax…) reach it. With no spec it falls back to every declared gate/up naming — what it did pre-registry. Skipping them would leave the fused `gate_up_proj` halves on inconsistent `weight_scale_2`. - **`ExportSpec.grouped_expert_export` matches the legacy `get_experts_list` support set**, except for `deepseek_v3`/`deepseek_v4`, which are new and had no legacy behavior to match. Note `qwen3_5_moe`: legacy keyed off the root class name and `"qwen3_5moeforcausallm"` matched none of its qwen substrings (the `_5` breaks `qwen3moeforcausallm`), so it raised — the spec keeps `False` to match. Enabling it belongs in its own PR. - **`is_moe` consults the model's own spec first**, before the generic name and structural fallbacks, so per-model data always wins. The three checks are or-ed, so this is a no-op; the same reordering is deliberately *not* applied to `get_expert_linear_names`, where the structural fused-experts check must stay first. - **Two values are corrected, not moved.** `gpt_oss` declared `block_names=("GptOssMoE",)`, which matches no real module (transformers names it `GptOssMLP`); naming still resolved via the single-naming shortcut, so this was invisible. Now `("GptOssMLP", "GptOssMoE")`. **5. ⚠️ DBRX expert input amax is now populated, changing its exported scales.** The second corrected value, called out separately because it has a numerical consequence. The legacy branch keyed on `DBRXMoeSparseMoeBlock`, a class transformers does not define — it names the block `DbrxFFN`. So DBRX fell through to the `w1`/`w2`/`w3` default, every `hasattr(experts_mlp, linear_name)` in `_prepare_dbrx_experts` evaluated `False`, and the handler wrote no expert input amax at all. `dbrx/specs.py` declares the real names (`w1_linear`, `w2_linear`, `v1_linear`), so the handler now does what it was written to do. Anyone diffing a re-exported DBRX checkpoint against an older one will see different expert activation scales. That is the fix working, not a regression from the refactor. No `CHANGELOG` entry: DBRX is not listed as a supported export target in `docs/source/deployment/` or `examples/hf_ptq/README.md`. ### Scope One consumer wired up: the unified HF export path. The TRT-LLM builders now live in `modelopt/torch/export/trtllm/` after #2365, and are deprecated. This PR changes one import line there — `is_moe` moved to `modelopt/torch/models/moe.py`, so importing it from `..layer_utils` no longer resolves — and nothing else. Their `is_moe()` calls still pass no `model_type` and fall back to the all-specs search, reproducing today's behavior; their hardcoded tables, including a second copy of the weight-plus-one norm names, stay as they are and migrate whenever that package does. Worth noting that #2365 landed the same boundary from the other direction: the slimmed `export/layer_utils.py` is now documented as "module-shape predicates and MoE quantizer helpers shared by every export backend", which is exactly the set of functions this PR rewires. The Megatron path has its own per-family registry and is untouched; folding it in is a later question, not this PR's. `is_moe` itself moved out of `modelopt/torch/export/layer_utils.py` into `modelopt/torch/models/moe.py` — whether a module is an MoE block is a modeling question, not an export one, and it is the first piece of shared modeling logic to follow the specs into the new package. ### Relationship to #1939 (why a second registry?) Orthogonal layers: #1939's `ExportModuleRegistry` dispatches on **module structure** (*which handler runs?*), this registry resolves **family data** (*what are its values?*). Exporting a `QuantMixtralSparseMoeBlock`, #1939 picks the shared iterable-experts handler (also serving Qwen/DeepSeek/Gemma4); inside, `get_expert_linear_names` resolves the `mixtral` spec to `("w1", "w2", "w3")`. One handler serves many families, one spec serves many handlers. Sharing the matcher machinery is a planned follow-up. ### Usage ```python # modelopt/torch/models/qwen3_moe/specs.py — adding a model needs no engine edits register( ModelSpec( model_type="qwen3_moe", min_transformers_version="4.57", moe_spec=MoESpec( block_names=("Qwen3MoeSparseMoeBlock",), expert_linear_names=("gate_proj", "down_proj", "up_proj"), gate_up_pair=("gate_proj", "up_proj"), ), export_spec=ExportSpec( grouped_expert_export=True, pqs_fuse_rules=( (("Qwen3MoeAttention",), "v_proj", "o_proj"), (("Qwen3MoeMLP",), "up_proj", "down_proj"), ), ), ) ) ``` One layout per model. `block_names` is a tuple, so a model whose MoE appears under several class names is covered as long as they share a layout (`gpt_oss`'s `GptOssMLP`/`GptOssMoE`). `gemma4_text` imports its section from `gemma4` rather than restating it. ### Testing `tests/unit/torch/export/` — **202 passed**; `tests/unit/torch/quantization/` — **903 passed**. The `test_export_diffusers.py` collection error and 6 `test_quant_aware_conversion.py` failures are pre-existing and reproduce on untouched `main`, as does the one failing pre-commit hook (`generate-arguments-md`, no torch in its venv). `tests/unit/torch/models/test_model_specs.py`: registry matching (MRO / quantized classes / model-type scoping); every registered MoE section's `(block_names, expert_linear_names, fused_expert_names, gate_up_pair)` as an exhaustive table that fails if a spec is added without a row; the exhaustive `grouped_expert_export` support set; the structural fused-expert shortcut and the spec's precedence over it; the fused-container skip in `get_experts_list`; the `sync_moe_gate_up_amax` fallback; and legacy-equivalence of the aggregated `pqs_fuse_rules` / `gate_up_pairs` / `weight_plus_one_norm_names`. Mutation-checked — flipping a flag, typo-ing a block name, or swapping the precedence each fail. `tests/unit/torch/models/test_specs_vs_transformers.py` validates the specs against the *installed* transformers rather than a mirrored table: every registered block class must name a real class from its `min_transformers_version` on, and every `remote_code` model must still be absent. It names no model — which specs to check, and which cannot be, both come from the registry. > The run counts above were measured before the follow-up commits in this branch (spec > precedence, the policy split, the `MoESpec`/`MoELayout` merge) and have not been > re-measured; CI on the current head is the authority. **Local per-model-type E2E sweep** (transformers 5.3.0, tiny models from config, run against this branch and base `main`; not checked in): | model_type | real block class | this branch | base `main` | |---|---|---|---| | `qwen3_moe` / `qwen2_moe` / `qwen3_next` | `Qwen*SparseMoeBlock` | `gate/down/up` | same | | `mixtral` | `MixtralSparseMoeBlock` | `w1/w2/w3` | same | | `nemotron_h` | `NemotronHMoE` | `up/down` | same | | `gpt_oss` | `GptOssMLP` | `gate_up/down` | ❌ `w1/w2/w3` | | `deepseek_v3` | `DeepseekV3MoE` | raises | ❌ `w1/w2/w3` | Five types identical to base; both differences are base bugs the registry surfaces. `dbrx` could not be built on transformers 5.3.0 (`DbrxAttentionConfig` missing `rope_theta`). `sync_moe_gate_up_amax` was separately checked against base across registered, unregistered, and no-config models — identical in all cases. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ❌ — `get_expert_linear_names` no longer falls back to `w1/w2/w3`; see "Behavior changes" #2. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no copied code, no new dependency (stdlib only). - Did you write any new necessary tests?: ✅ — `tests/unit/torch/export/test_model_specs.py`. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ — happy to add an entry given the backward-incompatible item. - Did you get Claude approval on this PR?: ✅ — run; its CRITICAL finding on `sync_moe_gate_up_amax` is fixed above. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added model-aware export support across a broad range of MoE and dense architectures, including newer DeepSeek, Gemma, Qwen, Nemotron, Arctic, DBRX, GPT-OSS, Llama, and Mixtral variants. - Improved quantization and export handling for model-specific expert layouts, projection fusion, normalization, and gate/up synchronization. - Added automatic architecture recognition using Hugging Face model metadata. - **Bug Fixes** - Unsupported or ambiguous model layouts now produce explicit errors instead of applying potentially incorrect defaults. - **Documentation** - Added documentation describing model specifications and supported architecture metadata. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> Co-authored-by: Shengliang Xu <shengliangx@nvidia.com>