Files
Model-Optimizer/modelopt
h-guo18andShengliang Xu c5d1065331 Modeling Lib: per-model architecture kick off with spec registration (#1828)
### What does this PR do?

**Type of change:** New feature — per-model infrastructure (with
behavior changes, see below)

Starts `modelopt/torch/models/`: somewhere for ModelOpt to keep what it
knows about a
model, keyed by HF model type (`config.model_type`), one package per
type
(`<model_type>/specs.py`). Importing the package registers every spec;
consumers call
`get_spec(model_type)` / `match_moe_block(module, model_type)` and read
the fields.

The point is the infrastructure, not the migration. Today ModelOpt's
per-model knowledge
is scattered across if/elif chains in whichever subsystem happened to
need it, so the same
fact gets restated per subsystem and drifts. A model's MoE block
classes, its expert
projection naming, which of its norms store `w - 1` — these are facts
about the *model*,
and several subsystems want them. This PR gives them one home and one
lookup.

It is a kick-off, so it is deliberately narrow: specs plus the first
consumer. **Export is
that first consumer**, which is why most of the changed lines are on the
export side —
not because this is an export refactor. Quantization, speculative
decoding and the rest
keep their own tables for now; the per-model directory is where their
sections land later,
and `specs.py` is just the first thing in it.

| Data the first consumer moved in | Spec field | Consumers |
|---|---|---|
| MoE block classes and expert linear naming | `MoESpec` |
`get_expert_linear_names`, `get_experts_list`, `is_moe`,
`sync_moe_gate_up_amax` |
| grouped expert-export support set | `ExportSpec.grouped_expert_export`
| `get_experts_list` |
| AWQ `pre_quant_scale` fusion rules | `ExportSpec.pqs_fuse_rules` |
`fuse_prequant_to_linear` |
| weight-plus-one norm class names |
`ExportSpec.weight_plus_one_norm_names` |
`_layernorm_uses_weight_plus_one` |

`ModelSpec` holds each concern as a separate, optional section, and the
split is what
makes it extensible: **topic sections** hold architecture facts any
subsystem can read
(`MoESpec` — block classes, expert projection naming, how the experts
are stored),
**subsystem sections** hold one subsystem's own data and policy
(`ExportSpec` — AWQ fusion
rules, weight-plus-one norms, whether grouped expert export is validated
for this model).
A new subsystem adds a section rather than a table.

A section is `None` when the model has nothing to say about it, so a
dense model carries
`moe_spec=None` rather than an empty one. Two further fields record
where a model's
classes come from at all: `modeling_source` (`transformers` or
`remote_code`) and
`min_transformers_version`.

### Behavior changes

**1. Lookup key: `type(root_model).__name__.lower()` →
`config.model_type`.** The old key
changes after `quantize` wraps the model, forcing substring matching;
`model_type` is
stable, so `ExportContext` carries it and lookups are exact. Matching is
also tightened to
case-insensitive **exact** names against the module's **MRO**, so
quantized subclasses
still match via their base class without substring false positives.

**2. ⚠️ `get_expert_linear_names` raises instead of guessing.**
Unmatched MoE blocks
previously fell through to `["w1", "w2", "w3"]` (Mixtral naming); they
now raise
`NotImplementedError` telling you to register a `ModelSpec`.

The blast radius depends on the transformers release, because only
*iterable* expert
layouts consult the spec — a fused expert container is resolved by the
structural
first-projection check and never asks for naming. Measured against both
ends of the
support matrix:

| transformers | Unregistered MoE families `is_moe` admits | Reach the
spec lookup | Actually affected |
|---|---|---|---|
| 5.14.1 (`tf_latest`) | 19 | **0** — all fused | none |
| 4.57.6 (`tf_min`) | 9 | 7 iterable | **2** |

On 4.57, five of the seven (`ernie4_5_moe`, `flex_olmo`, `jamba`,
`olmoe`,
`qwen3_omni_moe`) name their experts `gate_proj`/`up_proj`/`down_proj`,
so the `w1`
fallback already failed with `AttributeError` on `main` — for them this
trades an
unhelpful error for one that names the fix.

The other two, **`minimax` and `phimoe`**, are Mixtral-derived with
`w1`/`w2`/`w3`
experts and did export on `main`, so they are now **registered** rather
than left to
regress. Both only take the per-expert path on transformers 4; on 5
their experts are
fused (`MiniMaxExperts`, `PhimoeExperts`) and the structural check
resolves them, which
their specs decline to answer for. **So no model that exported correctly
before this
change regresses.**

This is backward-incompatible and has a `CHANGELOG` entry.

**3. DeepSeek-V3 and V4 are now registered, and newly visible to the MoE
path.** Neither
was reachable before: `DeepseekV3MoE` (like `DeepseekV2Moe` and
`DeepseekV32MoE`) is
invisible to `is_moe` — its class name does not end in `SparseMoeBlock`
and it calls its
router `gate`, so neither the name test nor the structural
`router`+`experts` test
matches. `DeepseekV4SparseMoeBlock` was detected by name but had no spec
to resolve expert
naming from. Registering `block_names` is what puts V3 on the MoE path
at all, so this is
added coverage rather than a fix to existing behavior.

`deepseek` stays registered alongside them: it describes the remote-code
`DeepseekMoE`
block of DeepSeek-MoE/V1, and its spec is the only thing making that
block detectable.

**4. Fused expert containers are skipped in `get_experts_list` instead
of crashing.**
transformers 5 replaced several iterable expert `ModuleList`s with a
single module holding
3-D parameters, while the specs still describe the transformers 4
iterable layout — a spec
cannot tell the two apart, only the module can. The AWQ/SVDQuant
resmooth pass in
`requantize_resmooth_fused_llm_layers` reached `len(module.experts)` on
a module with no
`__len__` and died with `TypeError` mid-export. **Mixtral already hits
this on `main`**;
registering `deepseek_v3` only made an existing bug easier to reach. The
skip is scoped to
specs that claim iterable experts, so a layout the spec calls
unsupported (DBRX, whose
per-expert linears live under `experts.mlp`) still fails loudly rather
than silently
dropping resmoothing.

Everything else is behavior-preserving, pinned by tests:

- **`sync_moe_gate_up_amax` keeps generic coverage.** Quantization
admits MoE blocks
*structurally*, so unregistered families (Olmoe, Jamba, MiniMax…) reach
it. With no
spec it falls back to every declared gate/up naming — what it did
pre-registry.
Skipping them would leave the fused `gate_up_proj` halves on
inconsistent
  `weight_scale_2`.
- **`ExportSpec.grouped_expert_export` matches the legacy
`get_experts_list` support set**,
except for `deepseek_v3`/`deepseek_v4`, which are new and had no legacy
behavior to
  match. Note
`qwen3_5_moe`: legacy keyed off the root class name and
`"qwen3_5moeforcausallm"`
matched none of its qwen substrings (the `_5` breaks
`qwen3moeforcausallm`), so it
raised — the spec keeps `False` to match. Enabling it belongs in its own
PR.
- **`is_moe` consults the model's own spec first**, before the generic
name and structural
fallbacks, so per-model data always wins. The three checks are or-ed, so
this is a
no-op; the same reordering is deliberately *not* applied to
`get_expert_linear_names`,
  where the structural fused-experts check must stay first.
- **Two values are corrected, not moved.** `gpt_oss` declared
`block_names=("GptOssMoE",)`, which matches no real module (transformers
names it
`GptOssMLP`); naming still resolved via the single-naming shortcut, so
this was
  invisible. Now `("GptOssMLP", "GptOssMoE")`.

**5. ⚠️ DBRX expert input amax is now populated, changing its exported
scales.** The
second corrected value, called out separately because it has a numerical
consequence. The
legacy branch keyed on `DBRXMoeSparseMoeBlock`, a class transformers
does not define — it
names the block `DbrxFFN`. So DBRX fell through to the `w1`/`w2`/`w3`
default, every
`hasattr(experts_mlp, linear_name)` in `_prepare_dbrx_experts` evaluated
`False`, and the
handler wrote no expert input amax at all. `dbrx/specs.py` declares the
real names
(`w1_linear`, `w2_linear`, `v1_linear`), so the handler now does what it
was written to
do.

Anyone diffing a re-exported DBRX checkpoint against an older one will
see different
expert activation scales. That is the fix working, not a regression from
the refactor. No
`CHANGELOG` entry: DBRX is not listed as a supported export target in
`docs/source/deployment/` or `examples/hf_ptq/README.md`.

### Scope

One consumer wired up: the unified HF export path.

The TRT-LLM builders now live in `modelopt/torch/export/trtllm/` after
#2365, and are
deprecated. This PR changes one import line there — `is_moe` moved to
`modelopt/torch/models/moe.py`, so importing it from `..layer_utils` no
longer resolves —
and nothing else. Their `is_moe()` calls still pass no `model_type` and
fall back to the
all-specs search, reproducing today's behavior; their hardcoded tables,
including a second
copy of the weight-plus-one norm names, stay as they are and migrate
whenever that package
does.

Worth noting that #2365 landed the same boundary from the other
direction: the slimmed
`export/layer_utils.py` is now documented as "module-shape predicates
and MoE quantizer
helpers shared by every export backend", which is exactly the set of
functions this PR
rewires. The Megatron path has its own per-family registry and is
untouched; folding it in
is a later question, not this PR's.

`is_moe` itself moved out of `modelopt/torch/export/layer_utils.py` into
`modelopt/torch/models/moe.py` — whether a module is an MoE block is a
modeling question,
not an export one, and it is the first piece of shared modeling logic to
follow the specs
into the new package.

### Relationship to #1939 (why a second registry?)

Orthogonal layers: #1939's `ExportModuleRegistry` dispatches on **module
structure**
(*which handler runs?*), this registry resolves **family data** (*what
are its values?*).
Exporting a `QuantMixtralSparseMoeBlock`, #1939 picks the shared
iterable-experts handler
(also serving Qwen/DeepSeek/Gemma4); inside, `get_expert_linear_names`
resolves the
`mixtral` spec to `("w1", "w2", "w3")`. One handler serves many
families, one spec serves
many handlers. Sharing the matcher machinery is a planned follow-up.

### Usage

```python
# modelopt/torch/models/qwen3_moe/specs.py — adding a model needs no engine edits
register(
    ModelSpec(
        model_type="qwen3_moe",
        min_transformers_version="4.57",
        moe_spec=MoESpec(
            block_names=("Qwen3MoeSparseMoeBlock",),
            expert_linear_names=("gate_proj", "down_proj", "up_proj"),
            gate_up_pair=("gate_proj", "up_proj"),
        ),
        export_spec=ExportSpec(
            grouped_expert_export=True,
            pqs_fuse_rules=(
                (("Qwen3MoeAttention",), "v_proj", "o_proj"),
                (("Qwen3MoeMLP",), "up_proj", "down_proj"),
            ),
        ),
    )
)
```

One layout per model. `block_names` is a tuple, so a model whose MoE
appears under several
class names is covered as long as they share a layout (`gpt_oss`'s
`GptOssMLP`/`GptOssMoE`).
`gemma4_text` imports its section from `gemma4` rather than restating
it.

### Testing

`tests/unit/torch/export/` — **202 passed**;
`tests/unit/torch/quantization/` — **903
passed**. The `test_export_diffusers.py` collection error and 6
`test_quant_aware_conversion.py` failures are pre-existing and reproduce
on untouched
`main`, as does the one failing pre-commit hook
(`generate-arguments-md`, no torch in its
venv).

`tests/unit/torch/models/test_model_specs.py`: registry matching (MRO /
quantized classes
/ model-type scoping); every registered MoE section's `(block_names,
expert_linear_names,
fused_expert_names, gate_up_pair)` as an exhaustive table that fails if
a spec is added
without a row; the exhaustive `grouped_expert_export` support set; the
structural
fused-expert shortcut and the spec's precedence over it; the
fused-container skip in
`get_experts_list`; the `sync_moe_gate_up_amax` fallback; and
legacy-equivalence of the
aggregated `pqs_fuse_rules` / `gate_up_pairs` /
`weight_plus_one_norm_names`.
Mutation-checked — flipping a flag, typo-ing a block name, or swapping
the precedence each
fail.

`tests/unit/torch/models/test_specs_vs_transformers.py` validates the
specs against the
*installed* transformers rather than a mirrored table: every registered
block class must
name a real class from its `min_transformers_version` on, and every
`remote_code` model
must still be absent. It names no model — which specs to check, and
which cannot be, both
come from the registry.

> The run counts above were measured before the follow-up commits in
this branch (spec
> precedence, the policy split, the `MoESpec`/`MoELayout` merge) and
have not been
> re-measured; CI on the current head is the authority.

**Local per-model-type E2E sweep** (transformers 5.3.0, tiny models from
config, run
against this branch and base `main`; not checked in):

| model_type | real block class | this branch | base `main` |
|---|---|---|---|
| `qwen3_moe` / `qwen2_moe` / `qwen3_next` | `Qwen*SparseMoeBlock` |
`gate/down/up` | same |
| `mixtral` | `MixtralSparseMoeBlock` | `w1/w2/w3` | same |
| `nemotron_h` | `NemotronHMoE` | `up/down` | same |
| `gpt_oss` | `GptOssMLP` | `gate_up/down` | ❌ `w1/w2/w3` |
| `deepseek_v3` | `DeepseekV3MoE` | raises | ❌ `w1/w2/w3` |

Five types identical to base; both differences are base bugs the
registry surfaces.
`dbrx` could not be built on transformers 5.3.0 (`DbrxAttentionConfig`
missing
`rope_theta`). `sync_moe_gate_up_amax` was separately checked against
base across
registered, unregistered, and no-config models — identical in all cases.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — `get_expert_linear_names` no
longer falls back to `w1/w2/w3`; see "Behavior changes" #2.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no
copied code, no new dependency (stdlib only).
- Did you write any new necessary tests?: ✅ —
`tests/unit/torch/export/test_model_specs.py`.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — happy to add an entry given the backward-incompatible item.
- Did you get Claude approval on this PR?: ✅ — run; its CRITICAL finding
on `sync_moe_gate_up_amax` is fixed above.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added model-aware export support across a broad range of MoE and dense
architectures, including newer DeepSeek, Gemma, Qwen, Nemotron, Arctic,
DBRX, GPT-OSS, Llama, and Mixtral variants.
- Improved quantization and export handling for model-specific expert
layouts, projection fusion, normalization, and gate/up synchronization.
- Added automatic architecture recognition using Hugging Face model
metadata.

- **Bug Fixes**
- Unsupported or ambiguous model layouts now produce explicit errors
instead of applying potentially incorrect defaults.

- **Documentation**
- Added documentation describing model specifications and supported
architecture metadata.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-11 12:12:54 -07:00
..