Files
Model-Optimizer/tests
Shengliang Xu f1abc75626 Carry unplaced checkpoint weights using the loader's accounting, replacing MTP name-matching (#2427)
### What does this PR do?

Type of change: bug fix + new tests

Replaces `hf_ptq`'s name-based MTP detection with the Transformers
loader's own accounting of what it could not place.

#### The problem

`load_mtp_weights` found MTP weights by name — `"mtp" in key` plus
config-derived layer indices — backed by a support matrix of three
storage conventions (GLM-5.1 inlined, GLM-4.7 standalone file,
Qwen3-Next tail shard). Every architecture that spells it differently is
a silent miss, and a miss means a checkpoint exported without a
component its own config still advertises. The failure is quiet on both
sides: Transformers drops unexpected keys, and vLLM's weight loading is
pull-based, so a missing MTP produces no warning at all.

#### The fix

`from_pretrained(..., output_loading_info=True)` reports
`unexpected_keys` — *"keys that are found in the checkpoints, but not
expected in the model's architecture"* — which is exactly the carry-over
set, derived structurally rather than by naming, and already accounting
for on-the-fly key conversion that a set re-derived afterwards would
have to replay. Record those keys at load; carry them at export. Weights
the loader *did* place go through the normal export path unchanged.

#### Two mechanisms, disjoint by construction

Measured against a real `from_pretrained`, an off-index file reports
**nothing**: the loader opens only shards named in
`model.safetensors.index.json`, so it never saw those tensors to call
them unexpected. Those are sidecars, not weights — untouched by
quantization and absent from the export — so they are **copied
verbatim**, which costs no host memory, preserves the bytes and file
layout, and leaves the filename a consumer looks for where it was.

| on-disk layout | mechanism |
|---|---|
| inlined layer past `num_hidden_layers` (GLM-5.1, DeepSeek-V3) |
`extra_state_dict` |
| indexed `mtp.*` tail shard (Qwen3-Next) | `extra_state_dict` |
| standalone off-index file (GLM-4.7) | **copied whole** |

A tensor is never both copied and carried; that would export it twice,
and a test asserts it.

#### Removed as redundant

`load_mtp_weights`, `mtp_layer_prefixes_from_checkpoint`,
`get_inlined_mtp_prefixes`, `_load_tensors_matching`,
`_apply_to_model_state_dict`, `_keys_to_prefixes`, `_add_mtp_exclusions`
and its three call sites, the pre-quantization `enable: False` entries
`hf_ptq` appended to the recipe's `quant_cfg`, and the dead
`_mtp_layer_prefixes` fallback in `_get_num_nextn_predict_layers`.

#### Two deliberate behavioural changes

**MTP now follows the recipe** instead of being force-excluded by the
script — which is what `examples/megatron_bridge` already does; it has
no MTP-specific code at all. Recipes importing
`configs/ptq/units/default_disabled_quantizers` still disable `mtp.*`,
so their behaviour is unchanged; a recipe omitting that unit will now
quantize an MTP the model actually built.

**`quantization_config.ignore` can no longer claim a layer is
unquantized that the export in fact quantized.** That contradiction came
from `_add_mtp_exclusions` firing off a model attribute with no
cross-check against quantizer state.

### Usage

No API change for callers of `export_hf_checkpoint`. Within
`examples/hf_ptq`, model loading now goes through a wrapper that records
the loader's accounting:

```python
model, loading_info = auto_class.from_pretrained(ckpt_path, output_loading_info=True, **kwargs)
record_unplaced_source_keys(model, ckpt_path, loading_info.get("unexpected_keys"))
```

### Testing

`tests/examples/hf_ptq/test_carry_over_layouts.py` — 8 tests driving a
**real** `from_pretrained` against a tiny model, covering each of the
three conventions above plus an auxiliary (non-MTP) tower, two layouts
at once, a checkpoint with nothing stray, and that indexed shards are
never copied. CPU-only: the mechanism is bookkeeping during load, so a
GPU adds nothing; the export side already has GPU coverage in
`tests/gpu/torch/export/test_export_carry_over.py`.

The six `load_mtp_weights` tests are replaced by three on the recording
path, and the `get_model` test doubles now model `output_loading_info`
the way Transformers does.

All passing: 8 layout tests, 82 in the surrounding `examples/hf_ptq`
suite. `ruff` findings at parity with `main` on every changed file.

Files named like a main weight shard are excluded from the off-index set
whatever the index says — a fixture with an empty `weight_map` would
otherwise have made the source weights look like sidecars and copied
them into an export beside the quantized ones.

### The algorithm: which weights get carried, and how

Two disjoint sets of source weights reach the export without passing
through quantization. They are distinguished by **what the loader did
with the file**, and that difference decides both how each is found and
how each is moved.

**Set 1 — unplaced weights.** The loader opened the file and read the
tensor, but the model had no parameter for it, so Transformers reports
it in `unexpected_keys`. An MTP head the recipe did not quantize is the
common case. Moved as **tensors**: located in whichever shard holds
them, read, and merged into the exporter's `extra_state_dict`.

**Set 2 — off-index sidecars.** The index never names the file, so the
loader never opened it and never had the chance to call anything
unexpected. GLM-4.7 keeps its MTP head in a standalone `mtp.safetensors`
exactly this way. Moved as **files**: copied byte for byte, so no host
memory is spent re-serialising tensors the export does not otherwise
touch.

#### The index is not an inventory of the checkpoint

This is the part that is easy to get wrong, and it cost a silent
data-loss bug during review.

`model.safetensors.index.json` selects which **files** the loader opens
— not which **tensors** it sees. Within a file it opens, Transformers
enumerates every tensor present and reports the unexpected ones.
Verified by experiment against transformers 5.3.0:

| case | reported in `unexpected_keys`? |
|---|---|
| key absent from the index, in a shard the index names for *other*
tensors | **yes** |
| key in a file the index never names (`mtp.safetensors`) | **no** — the
file is never opened |

So a tensor missing from `weight_map` but sitting inside a main shard is
**set 1, not set 2**. An MTP head stored that way is reported, recorded
— and was then silently dropped, because resolution went through
`weight_map`, which by construction has no entry for it. The
`--vllm_fakequant_export` guard shared that lookup, so the check written
to refuse exports that drop weights stayed silent in exactly the case it
existed for.

`locate_source_keys` now resolves through the index first (free for
everything it lists) and header-scans the shards only for the leftovers
— names, never tensor data — warning when a key is in no file at all.
The carry and the guard share it, so they cannot disagree again.

#### Flow

1. **At load.** `record_unplaced_source_keys` stores Transformers' own
`unexpected_keys` on the model (`_modelopt_unplaced_source_keys`) plus
the resolved local checkpoint path. The question asked is "does the
model have a parameter for this key", never "is this an MTP head" — so
the mechanism is architecture-agnostic.
2. **At export, before dispatch.** `read_unplaced_weights` resolves each
recorded key to its shard and reads the tensors, merging them into
`extra_state_dict`. An explicitly passed `extra_state_dict` wins on a
name clash: a caller naming a tensor is more specific than our
inference.
3. **Rank behaviour.** Only the rank that writes `extra_state_dict`
reads the bytes — the FSDP2 writer emits it from rank 0 alone, so a full
read on every rank would be host memory spent and discarded (a
DeepSeek-V3-class MTP head is 10 GB+ in bf16). The **key list** is still
resolved on every rank, because `get_quant_config` runs per rank and the
configs must agree.
4. **Recording what was written.** `export_hf_checkpoint` records
`_modelopt_carried_over_names` — the union of carried tensors and the
off-index sidecars' tensor names — before `get_quant_config` runs,
because that is the first point that knows what was *written* rather
than what was merely unplaced.
5. **Exclusions.** Both sets must reach `quantization_config.ignore`, or
a deployment framework reads the top-level `quant_algo` and tries to
load an original-precision weight as a quantized one (the NVBug 5718750
class). `seed_carried_over_exclusions` is the single path for this,
called by `get_quant_config` after its per-layer pass and again by the
layerwise exporter from `finalize()` — which snapshots its config during
`bind()`, while calibration is still running, so it cannot see the
carried set any earlier.

#### What is deliberately excluded

- **Files that re-ship indexed weights.** Mistral's
`consolidated.safetensors` is a second full copy of the model, and
PEFT's `adapter_model.safetensors` is an adapter. Both are off-index,
and copying either would put unquantized weights beside the quantized
ones — vLLM's mistral load-format looks for `consolidated.safetensors`
by name, so it is not inert. Caught by name for the known conventions
and by tensor-name overlap for the rest.
- **Files named like a main weight shard**, whatever the index says, so
a broken or partial index cannot make the real weights look like
sidecars.
- **Symlinks are *not* excluded.** A Hugging Face snapshot stores every
file as a symlink into a sibling `blobs/`, so refusing links would drop
the sidecar of every hub-downloaded checkpoint.
`resolve_checkpoint_file` checks where the link *lands* — regular file,
inside the checkpoint dir or its blob root — rather than whether it is a
link.
- **Buffers Transformers recomputes.** `*.inv_freq` is skipped by the
fake-quant guard even when a shard provides it: older
Llama/Mistral-lineage conversions do list it in the index, and refusing
an export over it would reject checkpoints that export correctly today.

### How a carried, never-quantized MTP head reaches
`quantization_config.ignore`

Raised in review: `_add_mtp_exclusions` is gone, and an unplaced weight
has no module, so
`get_quant_config` walks right past it. That was a real gap, not just a
documentation one —
fixed here.

The export writes a carried MTP head in its original precision. If it is
absent from
`exclude_modules`, a deployment framework reads the top-level
`quant_algo` and tries to load
`eh_proj` as an FP8/NVFP4 weight — the same class of failure as NVBug
5718750, where a
`transformers>=5.0` MoE router was written in BF16 but never excluded.

The two cases now differ only in *why* the module is invisible to the
quantizer walk:

| | why invisible | handled by |
|---|---|---|
| MoE router (tf≥5.0) | module exists, never gets a quantizer |
`_get_unquantized_moe_router_names` |
| carried weight | no module at all in the live model |
`_get_carried_over_module_names` |

`_get_carried_over_module_names` reads the keys the loader recorded as
unplaced
(`_modelopt_unplaced_source_keys`), strips the trailing parameter name —
a state-dict key is
`<module path>.<parameter>` — and dedupes. Those names are seeded into
`layer_config_dict` as
`QUANTIZATION_NONE`, exactly as the router pass does, so they flow
through
`process_layer_quant_config` into `exclude_modules`, which
`convert_hf_config` emits as
`quantization_config.ignore`.

An MTP head the recipe *does* quantize is unaffected: it has a module,
is loaded normally, is
never in the unplaced set, and is reported as quantized. The
backward-breaking note above still
holds — MTP follows the recipe instead of being force-excluded — but a
head that ends up carried
rather than quantized is no longer silently missing from `ignore`.

Covered by `test_carried_over_weights_are_excluded_from_quantization`
and
`test_carried_over_module_names_strip_parameter_and_dedup`.

`--vllm_fakequant_export` does not carry unplaced weights; it now raises
rather than writing a
checkpoint quietly missing them.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — MTP layers now follow the
recipe rather than being force-excluded, and
`quantization_config.ignore` no longer lists layers the export may have
quantized (carried, never-quantized weights are still listed -- see
below). Shipped recipes are unaffected; see the Changelog entry.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 Backward Breaking Changes.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Draft: the behavioural change to MTP quantization is the part most worth
a second opinion — it aligns `hf_ptq` with `megatron_bridge`, which
special-cases nothing.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Quantization supports enabled operators outside transformer layers,
including language-model heads.
* Exports preserve checkpoint weights not loaded into the quantized
model.
* Additional safetensors sidecar files are copied unchanged into
exported checkpoints.
* Unquantized auxiliary components, such as vision layers, remain
available in exported models.

* **Behavior Changes**
  * Local recipe files take precedence over built-in recipes.
* Legacy architecture-specific recipe paths remain supported with
warnings.
  * MTP-specific export exclusions are no longer applied.

* **Bug Fixes**
  * Preserved weights are no longer incorrectly reported as unquantized.
* Incomplete exports are rejected when source weights cannot be placed
or preserved.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-21 14:00:34 -07:00
..