Commit Graph
497 Commits
Author SHA1 Message Date
Chenjie LuoandClaude Opus 5 a74054ab2b Let callers add MLflow tags to a fakequant serve's run (#2364)
### What does this PR do?

Type of change: new feature

The quantization run records what this library can see — the model, the
checkpoint, the vLLM and ModelOpt versions — but nothing about the
harness that launched it. A downstream tool that wants its own revision,
a sweep id, or a ticket number on the run has no way to put it there
today:

- `_run_tags()` returns a fixed dict
- `quant_config` (which becomes the run's params) is a hardcoded set of
`QUANT_*` variables
- MLflow itself has no environment variable for arbitrary tags

`MODELOPT_MLFLOW_EXTRA_TAGS` takes comma-separated `key=value` pairs and
merges them into the run's tags.

Two details worth a reviewer's attention:

**It joins `MLFLOW_ENV_VARS`.** A Ray-backed serve receives only the
variables named there, and the tracker runs in the rank-0 worker —
omitting it would make the feature silently do nothing under Ray.

**Caller tags are merged first**, so the library's own keys (`tool`,
`model`, `checkpoint_path`, `vllm_version`) are written over them and
keep describing the run truthfully whatever a caller sends.

`key=value` rather than JSON, learned from a live run: the variable
reaches the worker through a shell `export VAR="..."`, and JSON's own
double quotes terminate that quoting —

```
export MODELOPT_MLFLOW_EXTRA_TAGS_732b_DEPLOYMENT="{"internal_version": "4d8c"}"
```

arrived as `{`. A quote-free format survives verbatim and needs no
`json` import or exception handling. Splitting on the first `=` keeps
values that contain one, such as a URL with a query string.

### Usage

```bash
export MODELOPT_MLFLOW_EXTRA_TAGS="modelopt_internal_version=49fa29d5,sweep=kv-study"
python3 vllm_serve_fakequant.py "$MODEL" --mlflow https://your-mlflow-server/ ...
```

### Testing

Unit-level, over the helper: unset and empty variable, one and several
pairs, surrounding whitespace, an empty value, an entry with no `=`, a
trailing comma, and a value containing `=`. None raise; malformed
entries warn and are skipped.

End to end on a real fakequant serve (Nemotron-3-Nano-30B-A3B BF16,
`NVFP4_DEFAULT_CFG`, TP=8, Ray executor, vLLM 0.15, SLURM):

```
modelopt_internal_version  '49fa29d5'
modelopt_version           '0.47.0rc0.post32+gd38ed5ead'
git_sha                    'd38ed5ead'
quant_cfg                  'NVFP4_DEFAULT_CFG'
```

The tag was written by the `RayWorkerWrapper` process, which exercises
the whole path — env var → shell export → `--container-env` → raylet →
Ray actor → `_run_tags` — and confirms the `MLFLOW_ENV_VARS` entry is
doing its job. Also verified that the emitted payload survives a shell
export round-trip unchanged.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ <!--- Additive; with the
variable unset the tags are exactly as before. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ <!--- Verified manually as
above; there is no existing test module for vllm_mlflow_utils. Happy to
add one if you would like it. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <!--- Small additive feature in an example; tell me if it warrants an
entry. -->
- Did you get Claude approval on this PR?: ❌

### Additional Information

Consumed by Model-Optimizer-Internal MR !141/!147, which sets the
variable so a fakequant eval records the same harness commit on both its
quantization run and its evaluation-score run.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 20:29:16 +00:00
Shengliang Xu d19925e446 simple refactor(export): split TensorRT-LLM-only code into modelopt/torch/export/trtllm (#2365)
### What does this PR do?

Type of change: refactor.

**The TensorRT-LLM checkpoint export format is deprecated.** Per
`docs/source/deployment/1_tensorrt_llm.rst`: *"The
`export_tensorrt_llm_checkpoint` API will be deprecated in future
releases. Users are encouraged to transition to the unified HF export
API, which provides enhanced functionality and flexibility for exporting
models to multiple inference frameworks including TensorRT-LLM, vLLM,
and SGLang."*

That deprecated code was not sitting off to one side — it was
**interleaved with the export path we actually want to grow.**
`modelopt/torch/export` mixed the deprecated TensorRT-LLM checkpoint
logic with the framework-agnostic HF/Megatron export code, in the same
modules:

- `layer_utils.py` was 1,986 lines, of which ~1,600 were TensorRT-LLM
`build_*_config` builders. The HF path imports this module for five
small predicates (`is_moe`, `is_quantlinear`, …) and dragged the whole
deprecated builder set in with them.
- `model_config.py` held the TensorRT-LLM `ModelConfig` dataclasses
*and* the `QUANTIZATION_*` / `KV_CACHE_*` constants that every backend
needs, so all of HF export imported the deprecated checkpoint schema to
get a format name string.
- `quant_utils.py` carried two helpers whose only caller is the
deprecated `postprocess.py`.

**This PR isolates the deprecated format so it stops polluting the
HuggingFace export path.** Everything reachable only from
`export_tensorrt_llm_checkpoint` now lives under
`modelopt/torch/export/trtllm/`, and the dependency is **one-way**:
`trtllm/` reaches into the parent through `quant_format`, `quant_utils`
and `layer_utils`, and **no implementation module in the parent imports
`trtllm/`.** The single exception is the deprecation re-export in
`modelopt/torch/export/__init__.py` described below, which is scheduled
for deletion in 0.49.0.

That one-way edge is the property worth protecting in review. It means
the deprecated format can be evolved, frozen, or eventually removed
without touching HF export, and HF export can no longer accidentally
grow a dependency on it.

### Deprecation handling

The format has carried a deprecation notice in the deployment docs since
`bc546943b4` (2025-10-08, first shipped in 0.39.0) — about 11 months.
But the deprecation policy in `README.md` also specifies *how* a
deprecation is communicated: a changelog entry, a source statement of
timing, and a runtime warning on use. **None of those existed**; only
one docs page ever said anything. So 0.48.0 is the first release that
gives users a signal they can act on, and this PR treats it as the
*start* of the migration period rather than the end:

- Both entry points now emit a `DeprecationWarning` naming 0.48.0 and
the 0.49.0 removal.
- `export_tensorrt_llm_checkpoint` and
`torch_to_tensorrt_llm_checkpoint` **remain importable from
`modelopt.torch.export`** for this release only, so existing callers
keep working *and* actually receive the warning. Removing the path in
the same release that first warns would mean callers hit `ImportError`
and never see it.
- The 0.49.0 removal date is stated in all four channels the policy
names: the runtime warning, the source (`.. deprecated:: 0.48.0` plus a
comment), the changelog, and the deployment doc.

The deeper module paths (`modelopt.torch.export.model_config_export`,
`modelopt.torch.export.model_config`) are **not** forwarded. Neither
appeared in a docs example, and `model_config.py` never declared
`__all__`, so by the `__all__` convention in `CONTRIBUTING.md` they were
never part of the public surface.

Eight modules had no non-TRT-LLM importer and moved whole:
`model_config_export`, `model_config_utils`, `postprocess`,
`distribute`, `tensorrt_llm_utils`, `tensorrt_llm_type`,
`hf_config_map`, `mcore_config_map`.

Three were genuinely mixed and were split by call-graph analysis rather
than by file:

| module | stayed shared (HF path) | moved to `trtllm/` (deprecated) |
|---|---|---|
| `model_config.py` | `QUANTIZATION_*`, `KV_CACHE_*`,
`FUSION_FREE_FORMATS` → new leaf module `quant_format.py` | the
`ModelConfig` dataclasses + `LINEAR_*`/`LAYERNORM_*` checkpoint-layout
constants |
| `layer_utils.py` | 9 module-shape predicates and MoE quantizer helpers
(`is_moe`, `is_quantlinear`, `get_experts_list`,
`sync_moe_gate_up_amax`, …) | the 39 `build_*_config` builders and
enc/dec helpers |
| `quant_utils.py` | everything else | `get_scaling_factor_from_weight`,
`resmooth_and_get_scale` (only caller is `trtllm/postprocess.py`) |

`adjust_attn_amax_values` was deliberately left in the shared
`quant_utils.py`: it has no production caller at all (only a test), so
"used only by TRT-LLM export" is not demonstrable for it.

Nothing was added or removed. `export_tensorrt_llm_checkpoint` behaves
exactly as before, just from a new import path and with a warning
attached.

### Usage

```python
# Deprecated TensorRT-LLM checkpoint export — new home, and warns on call
from modelopt.torch.export.trtllm import (
    export_tensorrt_llm_checkpoint,
    torch_to_tensorrt_llm_checkpoint,
)
from modelopt.torch.export.trtllm.model_config import ModelConfig

# The pre-0.48 path still works for one release, and warns — removed in 0.49.0
from modelopt.torch.export import export_tensorrt_llm_checkpoint

# Shared format constants — new home, still re-exported from the top level
from modelopt.torch.export.quant_format import QUANTIZATION_NVFP4, KV_CACHE_FP8
from modelopt.torch.export import QUANTIZATION_NVFP4  # still works

# The recommended path — unchanged
from modelopt.torch.export import export_hf_checkpoint, get_model_type
```

### Testing

- `pre-commit` on all changed files: passes (ruff, ruff-format,
**mypy**, bandit, markdownlint). mypy caught one implicit re-export of
`is_layernorm`, now imported from the shared module directly.
- `tests/unit/torch/export`: **189 passed**. With the new `trtllm/` test
dir: **193 passed**.
- Full `tests/unit/torch`: **2367 passed, 0 export failures**. The 45
failures are pre-existing environment issues — a deepspeed circular
import and a read-only HF cache — confirmed by reading their error text,
not assumed.
- `pytest tests/gpu/torch/export --collect-only`: 172 items, no
collection error.
- In-repo consumers updated and re-verified by an AST scan that imports
every `modelopt.torch.export*` module referenced anywhere in the tree
and checks each imported name still resolves: `hf_ptq.py`,
`export_trtllm_ckpt.py`, `deepseek_v3/ptq.py`, the AutoQuantize
notebook, `hf_ptq/README.md`, 2 docs pages, 4 tests.
- **Deprecation contract is covered by committed tests** (3 new, in the
`trtllm/` test dir): the pre-0.48 top-level import still resolves to the
same objects, `torch_to_tensorrt_llm_checkpoint` warns *at call time*
rather than on first `next()` (it returns a generator, so a naive
`warnings.warn` in the body would fire late or never), and one
`export_tensorrt_llm_checkpoint` call emits exactly one warning rather
than two. The first of these makes closing the migration window early a
test failure rather than a silent regression. `pyproject.toml` sets no
`filterwarnings = error`, so no suite fails on the new warning.
- **After merging `main`** (4 commits, incl. a 180-line rewrite of
`unified_export_megatron.py` that touches a file this PR also edits):
merged with no conflicts, then re-verified rather than trusted — import
scan clean across 24 export modules, `ruff check` clean repo-wide, 193
export tests passing, GPU collection still clean.

**Not run: the GPU suites** (`tests/gpu/torch/export`,
`tests/gpu_trtllm`) — no GPU in my environment.
`tests/gpu/torch/export/test_export.py` had its imports retargeted, so
it is the one most worth a GPU run before merge.

### Reviewer note: the deprecated path has no test coverage

Worth knowing before reviewing. **No test in the repo — including
`tests/examples/` — calls `export_tensorrt_llm_checkpoint`,
`torch_to_tensorrt_llm_checkpoint`, any `build_*_config`,
`convert_to_tensorrt_llm_config`, or `postprocess_model_config`.** So
~4,600 moved lines have no direct tests, and this refactor is validated
by import-graph reasoning, lint and mypy rather than by tests exercising
the moved code.

Given the format is deprecated and scheduled for removal in 0.49.0, **no
new coverage is planned for the conversion path itself** — writing fresh
tests for an API being removed next release isn't a good use of effort.
The gap is documented so reviewers can weigh the risk, not as a TODO.
(The deprecation *mechanism* is tested; see Testing.)

One caveat on how the gap was established: a runtime check showing all
12 `trtllm` modules in `sys.modules` after the export suites is *not*
evidence of coverage — importing any submodule runs
`trtllm/__init__.py`, which star-imports `model_config_export` and pulls
in the rest. Real line coverage could not be measured (`coverage`'s
tracer is incompatible with this venv's torch build: `ValueError: module
functions cannot set METH_CLASS or METH_STATIC`, on both the C tracer
and `sysmon`). The claim rests on a call-site audit generated from the
actual public symbols of those modules.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ for the public API —
`export_tensorrt_llm_checkpoint` and `torch_to_tensorrt_llm_checkpoint`
remain importable from `modelopt.torch.export` through the 0.49.0
migration period, now with a `DeprecationWarning`. The undocumented
submodule paths `modelopt.torch.export.model_config_export` and
`.model_config` did move; see **Usage**.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
code or dependencies; existing code relocated.
- Did you write any new necessary tests?: ✅ — 3 tests covering the
deprecation contract (old import path, call-time warning, exactly-one
warning). One existing test also moved to mirror the source split. No
new coverage for the deprecated conversion path itself; see the note
above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 **Deprecations**, covering both the runtime warning and
the new import location.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Git detected the moves, so the diff stays reviewable: 8 files show as
pure renames (100%), the three split files as rename/copy at 94–99%
similarity, and only `layer_utils.py` as a 79% rewrite — expected, since
it shed 1,616 lines to `trtllm/`.

`examples/hf_ptq/hf_ptq.py` and
`examples/llm_sparsity/weight_sparsity/export_trtllm_ckpt.py` still call
the deprecated API, so those examples now print the warning. That is the
intended nudge, but happy to silence or migrate them if preferred. They
import from the new `.trtllm` path already, so they need no change at
0.49.0.

Two incidental changes, easy to revert if unwanted:
- `modelopt/torch/export/layer_utils.py` mode `100755 → 100644` (it was
needlessly executable).
- The new test is named `test_trtllm_quant_utils.py`, not
`test_quant_utils.py`: these directories have no `__init__.py`, so
pytest derives the module name from the bare filename and the shorter
name fails collection with `import file mismatch` against the existing
`test_quant_utils.py` one level up.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added shared quantization and KV-cache format definitions for export
workflows.
- Added expanded TensorRT-LLM export support, including broader model
architecture and quantization handling.
- Added distributed export utilities for coordinating checkpoint data
across processes.

- **Deprecation**
- TensorRT-LLM checkpoint export now emits a warning and is scheduled
for removal in version 0.49.0.
- Use the documented export module and save optimized model state
explicitly when needed.

- **Documentation**
- Updated guides and examples with new import paths and deprecation
guidance.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-10 10:44:04 -07:00
yeyu-nvidiaandClaude Opus 4.8 9d0df45849 specdec: config_overrides for nested text_config checkpoints + load VLM-capable bases in merge_lora (#2289)
### What does this PR do?

Type of change: New feature + bug fix

Two related gaps, both hit while enabling EAGLE3 on a checkpoint whose
config nests its text dims.

**1. `config_overrides` for checkpoints whose `text_config` dims don't
propagate.**
Some multimodal checkpoints carry the real text-tower dims only under
`config.text_config`, leaving the parent fields `None`.
`from_pretrained` then builds a text tower with the wrong shape.
`load_vlm_or_llm` gains an optional `config_overrides` dict applied to
*both* the parent config and its `text_config` before instantiation, and
the three entrypoints that load checkpoints — `ar_validate.py`,
`export_hf_checkpoint.py`, `merge_lora.py` — get a `--config_overrides`
passthrough. `main.py` threads it from `ModelArguments`.

**2. `merge_lora.py` could not merge into any VLM base.**
It loaded via `AutoModelForCausalLM`, which cannot load architectures
absent from the CausalLM Auto map — every VLM base failed. It now goes
through `load_vlm_or_llm`, which routes VLMs to
`AutoModelForVision2Seq`/`AutoModelForImageTextToText` and plain LLMs to
`AutoModelForCausalLM` with the same `dtype`/`device_map`, so LLM
behavior is byte-for-byte unchanged.

Also adds an optional `transformers_cosmos3` import so `cosmos3_omni` is
registered with `AutoConfig` before use, and dispatches that
`model_type` to its model class directly — that plugin registers only a
*config*, never a model under `Auto*`, so `AutoModelForCausalLM` raised
`KeyError('cosmos3_omni')` regardless of imports. The import is wrapped
in `contextlib.suppress(ImportError)`, so it is a no-op when the plugin
isn't installed.

### Usage

```bash
# Checkpoint whose real dims live under config.text_config
python examples/speculative_decoding/scripts/ar_validate.py \
    --model_path <ckpt> --trust_remote_code \
    --config_overrides '{"num_hidden_layers": 36, "intermediate_size": 12288, "num_key_value_heads": 8}'

# Same flag on export and merge
python examples/speculative_decoding/scripts/export_hf_checkpoint.py \
    --model_path <ckpt> --export_path <out> --config_overrides '{"num_hidden_layers": 36}'
python examples/speculative_decoding/scripts/merge_lora.py \
    --base_model_path <base> --exported_lora_dir <out> --output_path <merged> \
    --config_overrides '{"num_hidden_layers": 36}'
```

```python
model = load_vlm_or_llm(path, config_overrides={"num_hidden_layers": 36})  # default None
```

### Testing

Exercised end-to-end on a Cosmos3-Nano (16B, 36-layer text tower) EAGLE3
LoRA run:

- **Training** — the base loads with all 36 text layers and correct
dims; two 4-epoch co-training runs completed (46,816 steps each).
- **Export + merge** — produced `adapter_model.safetensors` and a merged
base. Verified correct by per-layer weight diff: a `start_layer=18` run
changed **exactly** layers 18-35, with layers 0-17 bit-identical to the
base.
- **AR validation** — `--config_overrides` loads the trained checkpoint;
80/80 MT-Bench samples, AR 3.42.
- **Regression check** — `merge_lora` via `load_vlm_or_llm` produces a
base loadable by `lm_eval`; ifeval/arc_challenge/winogrande all ran to
completion.

No local unit-test run: `nvidia-modelopt` isn't installed in my
checkout, so `tests/conftest.py` fails to import. Relying on CI.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `config_overrides` defaults
to `None`; the `merge_lora` loader swap keeps the same class, dtype and
device_map for plain LLMs.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependency; `transformers_cosmos3` is an optional import guarded by
`contextlib.suppress`.
- Did you write any new necessary tests?: ❌ — exercising these paths
needs a checkpoint with a nested `text_config`, which the unit suite has
no fixture for. Happy to add one if a reviewer can point me at a small
suitable model.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — can add a *Speculative Decoding* entry for the `merge_lora` VLM fix
if you consider it changelog-worthy.
- Did you get Claude approval on this PR?: ❌ — not yet run.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added JSON-based model configuration overrides across speculative
decoding, training, validation, export, and LoRA workflows.
  - Overrides can update primary model and text configuration settings.
- Expanded support for vision-language models and Cosmos3 Omni
checkpoints.

- **Bug Fixes**
- Improved configuration handling for offline loading and
checkpoint-based initialization.
- Restored draft-model precision during checkpoint loading and model
conversion.
- Added validation for malformed, unsupported, and non-finite override
values.
- Standardized configuration override guidance across command-line
workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-09-10 10:06:35 -07:00
yeyu-nvidiaandClaude Opus 4.8 279d510616 fix(specdec): correct resume and bound staging in the vLLM hidden-state dump (#2080)
### What does this PR do?

**Type of change:** Bug fix

Fixes two issues in the vLLM offline hidden-state dump

(`examples/speculative_decoding/collect_hidden_states/compute_hidden_states_vllm.py`).
Both are invisible on small dumps and only bite at scale, which is why
they survived until
now — they were found while dumping ~194k conversations for a MiniMax-M3
draft.

**1. Resume silently re-processed already-finished work.**

`keep_conversation` skips conversations whose `.pt` already exists, but
that predicate reads
**on-disk state**, which is not part of the fingerprint `datasets`
computes for `filter()`
(it hashes the function and the dataset). With a persistent HF cache
reused across a resumed
or requeued run, the cached *"keep everything"* result from an earlier
run — computed when
few or no `.pt` files existed — is replayed. The run then re-generates
and **overwrites**
conversations it had already completed, and reports `Removed 0
conversations due to existing
output files` while doing so.

Observed on a 194k-conversation dump: ~62k `.pt` rewritten over a
two-hour window with the
total output count completely flat.

Fix: pass `load_from_cache_file=False` so the filter re-checks the disk
on every run.

**2. Staging exhausted `/dev/shm` partway through large dumps.**

The script generated the **entire** dataset before saving anything. The
KV connector stages
each conversation's hidden states under its `shared_storage_path`
(`/dev/shm`, i.e. RAM, by
default) and they are only freed by `cleanup_hidden_states()` in the
save loop — so every
conversation stayed staged simultaneously. On a large dump this exhausts
the space and the
connector starts failing writes:

```
Hidden-states write failed for req_id=...:
  SafetensorError('Error while serializing: I/O error: No space left on device (os error 28)')
```

Fix: generate and save in chunks of `--save-chunk-size` (default 256),
so at most one chunk
is staged at a time. As a side benefit the dump becomes **incrementally
durable** — an
interrupted run (walltime limit, node failure) keeps its finished
conversations and the
resume path above continues from them, instead of losing the whole run's
work.

### Testing

- Reproduced both failures on a 194k-conversation MiniMax-M3 dump (8-way
DP, TP8), and
confirmed both fixes on the same workload: after the change the output
count advanced
monotonically across requeues (123k → 194k) with no rewrites, and
`/dev/shm` stayed bounded
  through completion.
- `pre-commit run --files ...` passes (ruff check/format, mypy, bandit,
license, rst checks).
- Behavior is unchanged for a fresh single-shot dump other than the
chunked generate calls;
  the default `--save-chunk-size 256` is the only new knob.

### Additional Information

Extracted from #1749, which is otherwise superseded by the streaming
DFlash/DSpark path — these
two fixes are model-agnostic and apply to any offline dump, so they are
worth landing on their
own.

### Before your PR is "*Ready for review*"

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No — the failure modes are
multi-process/at-scale (datasets cache reuse across runs, connector RAM
staging) and are not reproducible in the unit-test harness.
- **Did you add or update any necessary documentation?**: Yes —
CHANGELOG entry.
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added chunked hidden-state generation for large vLLM offline runs.
* Added a configurable save-chunk size, defaulting to 256 conversations.
  * Enabled incremental saving and resumption of hidden-state outputs.

* **Bug Fixes**
* Improved resume filtering to accurately detect existing output files.
  * Reduced memory usage by saving and releasing each generated chunk.
* Ensured temporary files are cleaned up after interrupted or skipped
saves.
  * Added atomic output-file replacement to prevent incomplete results.
* Added validation to prevent invalid conversation IDs from creating
unsafe paths.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-09 20:08:42 +00:00
yeyu-nvidia 56af187565 ar_validate: fail loudly when every sample fails (#2288)
### What does this PR do?

Type of change: Bug fix

`validate_ar()` catches per-sample exceptions, prints a `WARNING`, and
returns whatever succeeded. When *every* sample failed it returned an
empty list, and the reporting block was guarded by `if results and
accelerator.is_main_process:` — so the script printed no results and
exited **0**. A run where 100% of samples failed was indistinguishable
from a successful one.

This bit us on a real run: an EAGLE3 checkpoint loaded with
`device_map="auto"` was sharded across 8 GPUs, every one of the 80
samples died with `Expected all tensors to be on the same device`, and
the job still exited 0 with no AR number anywhere in the log — the
wrapper stamped it PASS.

Now it raises, so the caller sees a non-zero exit. Any previously
"passing" run that printed no AR number was never meaningful.

### Usage

No API change. Existing invocations are unaffected when at least one
sample succeeds:

```bash
python examples/speculative_decoding/scripts/ar_validate.py \
    --model_path <ckpt> --steps 3 --osl 1024 --num_samples 80
```

### Testing

Reproduced the silent-pass on a Cosmos3-Nano EAGLE3 checkpoint (80/80
samples failing): before this change the job exited 0 and stamped PASS;
after it, the job exits non-zero with the sample failures visible.
Confirmed the normal path is unchanged by a subsequent run that
completed 80/80 and printed AR 3.42.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — only affects the
all-samples-failed case, which previously produced no output and a
misleading exit 0.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ — the failure path requires
a model that errors during AR validation; the existing suite has no
harness for that.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — behavior fix in an example script, not a released-feature change.
- Did you get Claude approval on this PR?: ❌ — not yet run.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
  * Improved validation error handling when all samples fail.
  * Validation now rejects non-positive sample counts before processing.
* Empty validation results are clearly distinguished from cases where
all samples fail.
* Error messages report the actual number of validation samples
attempted, capped at the available dataset size.
  * Empty validation results are no longer reported as successful.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
2026-09-09 12:40:55 -07:00
Keval MorabiaandClaude Opus 5 4956213d67 Unblock Qwen3.5/3.6 QAD: Megatron export, calibration, and distillation fixes (#2334)
### What does this PR do?

Type of change: Bug fix

Everything that blocked running QAD on a quantized Qwen3.5 / Qwen3.6 MoE
checkpoint: two
Megatron-Core → HuggingFace export bugs that make it unservable (§1–2),
the dead code the first
leaves behind (§3), a no-op flag (§4), a multi-GPU calibration deadlock
(§5), and four
distillation / data-prep bugs that stopped QAD itself from running (§6).

#### 1. Routed experts were exported packed, and vLLM cannot load that

```
AttributeError: Layer language_model.model.layers.23.mlp.experts has no parameter
  'w2_weight_weight_scale_2' for checkpoint weight ...experts.down_proj_weight_scale_2
```

`mcore_qwen35vl.py` used `GroupedMLPPacking`, mirroring the **BF16
upstream** checkpoint, which
really is packed. But that mapping is only used for **quantized**
export, and vLLM's quantized MoE
loader needs per-expert scales — both released NVFP4 checkpoints
(`Qwen3.6-35B-A3B-NVFP4` via
hf_ptq, `Nemotron-3.5-Lightning-30B-A3B-NVFP4` via Megatron-LM) are
per-expert.

`_grouped_mlp_slicing` gains `gate_proj_name` / `up_proj_name` to split
each expert's fused gate+up
and slice its per-block `weight_scale`; `GroupedGatedMLPSlicing` wires
it up. The Megatron
checkpoint layout is unchanged, so affected checkpoints need only a
**re-export**.

`_verify_exported_keys` is relaxed to match: exported modules now
contribute their ancestor
prefixes, so expanding one source module into many is not reported as
~82 dropped tensors. A module
genuinely absent still has nothing beneath its prefix and is still
caught.

#### 2. A quantized `output_layer` (`lm_head`) could not be checkpointed

`GPTModel.sharded_state_dict` drops `output_layer._extra_state` and
asserts it is empty. ModelOpt
keeps quantizer state there, so saving raised and — since that method
also backs the load plan —
loading silently restored the layer **unquantized**.

`keep_gpt_output_layer_extra_state()` retains it, applied from
`megatron_replace_quant_module_hook` so **every** Megatron model gets it
(Megatron-LM and NeMo
users included, neither of whom can import `mbridge`, which needs
`megatron.bridge`). It matches
the upstream body by AST before replacing it and self-disables
otherwise.

[NVIDIA/Megatron-LM#7086](https://github.com/NVIDIA/Megatron-LM/pull/7086)
is **closed, not
merged**: nemo:26.10 migrates `GPTModel` to `HybridModel`, whose
`sharded_state_dict` has no
pop-and-assert, so this side keeps the workaround.

Not cosmetic: `lm_head` is 248320×2048 = 509M params, **34.6% of
per-token weight traffic** on a
model with ~2.9B active params.

#### 3. Cleanup

Nothing maps `GroupedMLPPacking` once qwen3_5 is switched over; it is
removed with
`_grouped_mlp_packing` and the `quantize=` / `record_quant_config=`
parameters that existed only to
serve it. Llama-4's `PackNameRemapping` is unaffected.

Two smaller review-driven fixes: the gated-split shape checks raise
`ValueError` rather than
`assert` (stripped under `-O`), and per-expert quant metadata is
recorded for
`local_expert_indices` rather than every global id, fixing
non-contiguous EP.

#### 4. Remove the no-op `--moe_calib_experts_ratio` from the Megatron
quantize example

`examples/megatron_bridge/quantize.py` accepted the flag and threaded it
into the `mtq` config, but
`_moe_calib_experts_ratio` exists only in `plugins/huggingface.py` (9
refs) and never in
`plugins/megatron.py` (0); `mode.py:247` only assigns it to modules
already exposing the attribute.
On a Megatron MoE model it was accepted and silently ignored — a trap,
since on a 256-expert model
it reads like a major quality lever. `hf_ptq.py` keeps it, where it
works.

#### 5. Fix multi-GPU image-text (VLM) calibration deadlocking

VLM calibration hung for 30 minutes and died on a gloo timeout whenever
`world_size > 1`, with no
error until the timeout fired.

`NemotronTarPlusJsonlIterable` split its budget with truncating
division, so the stream supplied
fewer samples than requested (1024 over 3 subsets → 341×3 = **1023**).
`_ShardedIterable` gives
rank *r* items *r, r+W, r+2W…*, so a stream that is not a multiple of
`world_size` leaves the
trailing rank one short — it exits the forward loop early and the others
block on the next
collective. The arithmetic predicts both observed hangs exactly: 1024 →
stall at **255/256**,
512 (yielding 510) → **127/128**.

Fixed both ends: subset budgets are distributed with `divmod` so they
sum exactly, and
`_ShardedIterable` truncates every rank to `floor(len / world)` — which
also covers `num_samples`
not being divisible by `world_size`, as the first fix alone does not.

Verified on Qwen3.6-35B-A3B (EP=4, `nemotron_vlm_dataset_v2`, 1024
samples): the configuration that
hung twice now completes 256/256 and exports. Unit tests cover both
fixes and fail without them.

#### 6. Fix the distillation path so QAD can actually run

Four independent bugs, all hit while running QAD end to end on
Qwen3.6-35B-A3B. Each blocks a
different configuration, and together they made every sequence length
OOM or abort.

- **Context parallel aborts.** The DDP config derived
`average_in_collective` from `--sft` alone,
but context parallel also needs per-token loss reduction, so any
`--cp_size > 1` run died on
  `Cannot average in collective when calculating per-token loss`.
- **`TopKLogitsKLLoss` was not memory-efficient.** Despite documenting
"without gathering full
logits", it cast the *whole* vocabulary to FP32 before selecting the
top-k, allocating two
`[seq, vocab]` tensors — 30.3 GiB each at seq 32768 on this model's 248k
vocab. Reducing before
the cast is equivalent: widening is exact and temperature scaling is
monotonic, so the selected
  entries and the loss are unchanged.
- **MTP cross-entropy ran when it had nothing to recover.**
`skip_lm_loss` exempts the MTP heads
unconditionally, so their CE materialised another FP32 `[seq, vocab]`
tensor even when the MTP
head is excluded from quantization — as it is in every recipe here (775
of 906
`exclude_modules`, zero MTP `weight_scale` tensors exported). It is now
skipped **only** when the
model is quantized and MTP is left out of it; plain distillation such as
pruning recovery still
trains the MTP head. `test_mtp_excluded_from_quantization` pins all four
cases.
- **One bad record deadlocked data prep.** `megatron_preprocess_data`
re-raised chat-template
failures out of a pool worker, stalling the whole job until it timed out
— three malformed
records cost a multi-hour tokenization run. They are now skipped with a
warning, matching the
  existing handling of malformed JSONL a few lines above.

Also exposes `--logit_kl_topk`, which `DistillationConfig` has supported
for a while but the
example never passed through; `test_qad` now exercises that path.

§4, §5 and §6 are independent of §1–3; happy to split them out if
reviewers prefer.

### Usage

No API change. Exported names now match the released checkpoints:

```
model.language_model.layers.0.mlp.experts.<E>.{gate,up,down}_proj.{weight,weight_scale,weight_scale_2}
lm_head.{weight,weight_scale,weight_scale_2}
```

### Testing

- `test_mcore_export_mappings.py` — qwen3_5 mappings emit per-expert
rules. Verified these fail
without the fix (2 failed / 11 passed), with `Qwen3MoeForCausalLM` /
`NemotronHForCausalLM` as
  controls.
- `test_unified_export_megatron.py` — the gate/up split, per-block scale
slicing, the 0-dim scalar
fallback, and both directions of the `_verify_exported_keys` relaxation.
- `test_megatron.py::TestKeepGptOutputLayerExtraState` — 15 cases:
payload detection, no-op second
call, warn-and-skip on an unrecognised `sharded_state_dict`, and
`test_patches_stock_megatron_core`
which installs a replica of the real pre-fix upstream body (verified
against `be08ce5b1~1`) so the
  patched path is exercised whichever megatron-core is installed.
- `test_qad.py` — CI caught that its reference comparison still assumed
packed experts; fixed.

**End to end on `Qwen/Qwen3.6-35B-A3B` (35B MoE, 256 experts), 4×GB200,
nemo:26.08:**

| | before | after |
| --- | --- | --- |
| export self-check | `Export dropped 82 tensor(s)` | passes |
| expert tensors | `mlp.experts.gate_up_proj` (packed) |
`mlp.experts.<E>.{gate,up,down}_proj` |
| **vLLM v0.28.0 load** | **`AttributeError`, engine never starts** |
**`Loading weights took 25.61 s`** |
| **NEL eval (GPQA-D, MMMU-Pro)** | **FAILED** | **SUCCESS** |

### Results these fixes unblocked

The export fix is what made a Megatron-produced NVFP4 MoE checkpoint
servable at all, so it enabled
a full PTQ study on Qwen3.6-35B-A3B. Accuracy deltas are against a BF16
baseline measured on the
same harness, from **paired** per-question tests:

| recipe | throughput vs BF16 | GPQA-D | SciCode ×8 | MMMU-Pro | IFBench
|
| --- | --- | --- | --- | --- | --- |
| **W4A16** (weight-only) | **0.64–0.86×** — *slower* | −0.06 | −0.15 |
+0.48 | −0.44 |
| **W4A4** | 8/12 shapes faster | −0.60 | −0.70 | −1.48 (p=0.019) |
−0.53 |
| **W4A4 + 4-bit `lm_head`** | **9/12 shapes**, up to **1.30×** |
**+0.03** (p=0.96) | −0.81 | **−1.16** (p=0.016) | −1.65 (ns) |

Repeats: GPQA-D is `pass@1[avg-of-16]`; SciCode is 8 pooled runs per
recipe; MMMU-Pro is 3 runs per
side and IFBench 2–3 for BF16 and the last row, 1 elsewhere. AA-LCR
(68.33 → 71.33, p=0.25, 3 runs
per side) and τ²-Telecom (94.25 → 94.25, 3 runs per side) are on par; at
100 questions and 114
tasks they cannot resolve below ~5 pp and ~3 pp, so they carry no claim
either way.

#### QAD status (what §6 unblocked)

With the §6 fixes in place, QAD runs end to end on this model: 32 nodes,
`TP=1 PP=1 CP=1 EP=8`,
seq 32768, gbs 512, ~38 s/iter, 124 GB/GPU peak. First accuracy read,
MMMU-Pro at iteration 50
(0.84 B tokens), 3 runs per side, paired per-question:

| | MMMU-Pro | vs BF16 |
| --- | --- | --- |
| BF16 | 74.55 | — |
| W4A4 + 4-bit `lm_head` (PTQ) | 73.39 | **−1.16, p=0.016** |
| + QAD, iteration 50 | 73.78 | −0.77, p=0.089 (ns) |

The PTQ deficit that motivated this work is no longer statistically
significant after 50 QAD
iterations. The improvement itself (+0.39 vs PTQ) is **not** significant
at p=0.41, and 50
iterations is 10% of the planned budget, so this is a direction rather
than a result. A full
six-benchmark sweep at iterations 50 and 300 is running; these numbers
will be superseded.

Two findings worth flagging beyond this PR:

- **Weight-only NVFP4 is slower than BF16 on Blackwell.** W4A16 leaves
activations in BF16, so vLLM
cannot use the FP4 tensor cores and falls back to
`MarlinNvFp4LinearKernel` / `'MARLIN'` MoE.
W4A4 selects `FLASHINFER_TRTLLM` + `FlashInferCuteDslNvFp4LinearKernel`
and beats W4A16 in
**12/12** shapes. The Marlin line count tracks the recipe exactly (one
W4A16 layer ⇒ one Marlin
  line ⇒ zero once `lm_head` is W4A4).
- **The only accuracy cost is multimodal**: **−1.2 pp on MMMU-Pro** for
the fastest recipe,
confirmed over 3 runs per side (p=0.016). GPQA-D, SciCode, IFBench,
AA-LCR and τ²-Telecom show no
  significant regression.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ <!-- Megatron checkpoints
unaffected; re-export to gain the loadable layout. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!-- 0.47.0 → Bug Fixes; includes the removed flag, since passing it
now errors instead of being ignored -->
- Did you get Claude approval on this PR?: ✅ <!-- Reviewed; all findings
addressed, threads resolved. -->

### Additional Information

Both export bugs were found while reproducing
`nvidia/Qwen3.6-35B-A3B-NVFP4` through
`examples/megatron_bridge/`. Follow-up to #2332. Upstream counterpart

[NVIDIA/Megatron-LM#7086](https://github.com/NVIDIA/Megatron-LM/pull/7086)
is closed — see §2.
Labeled `cherry-pick-0.47.0`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-09 23:19:56 +05:30
2d35643452 LiLiCorr training (#2342)
### What does this PR do?

Type of change: new feature

Adds **LiLiCorr**, a candidate-lattice reranker for DFlash drafts, as a
new `projector_type` on the
existing `dflash` mode — plus three DFlash-wide improvements that apply
to every variant, and an
optional composition with DFlash2's grouped convolutions.

A DFlash drafter is trained on per-position marginals rather than on the
joint block distribution,
so its drafted tokens are individually plausible yet jointly incoherent.
LiLiCorr keeps the top-`k`
candidates the backbone already produces at each block position, scores
transitions between adjacent
candidates with a small two-layer transformer, and commits a path
through the lattice greedily.
Serving is unchanged in kind: verify still checks every drafted token
against the target, so the
emitted distribution is untouched and only acceptance length moves.

- Paper: [LiLiCorr: Lightweight Likelihood Correlation of Parallel
Drafts for Speculative Decoding](https://arxiv.org/abs/2608.20530)
(arXiv:2608.20530)
- Blog: https://research.nvidia.com/labs/nemotron/lilicorr/
- **Companion PR — serving support:**
[sgl-project/sglang#37462](https://github.com/sgl-project/sglang/pull/37462)

This PR is the **training** half. It trains the drafters and exports
them; the companion PR above is
what serves the resulting checkpoints, and is what the comparison table
below was measured through.

**What is in the commits**

| | |
| --- | --- |
| LiLiCorr draft variant | `hf_lilicorr.py`, `modeling_lilicorr.py`,
conversion routing, config fields, export |
| Three DFlash-wide features | fp32 master weights for the draft, draft
activation checkpointing, and a DDP hang fix — all default-off or
behaviour-preserving, all applying to `dflash`, `domino`, `dspark` and
`dflash2` alike |
| Optional grouped convolutions | composes LiLiCorr with DFlash2's
`DFlashGroupedConv`; see the dependency note below |
| Two recipes | `lilicorr.yaml` and `lilicorr_conv.yaml` |
| CPU unit tests, CHANGELOG, one launcher example | |

**⚠️ The convolutions depend on the DFlash2 branch, and cannot run until
it merges.**

`modeling_lilicorr.py` imports `DFlashGroupedConv` from
`modeling_dflash2`, which today exists only
on `haoguo/dflash2-support`. The class is **imported rather than copied
on purpose** — it is the only
way the two variants cannot drift apart arithmetically — but the
consequence is that the
convolutional recipe cannot run against `main` as it stands.

So the import is **deferred into `_install_sublayer_convs`** rather than
taken at module scope.
Everything else in this PR, including the plain LiLiCorr reranker, has
no DFlash2 dependency at all
and works on `main` today; an eager import would have made the whole
plugin unimportable for the sake
of one optional feature. Requesting the convolutions without DFlash2
present raises an `ImportError`
naming the two config keys to remove, rather than failing at import
time.

**This PR carries two of @h-guo18's commits, with authorship and
sign-off preserved.** Both are
independent of DFlash2 itself and both are needed here:

- `1419d47e`, the no-op sublayer seam. Without it
`DFlashDecoderLayer.forward` never calls the
wrappers the convolutions install onto, so the modules would be built,
counted and exported while
  computing nothing. It is arithmetically an identity on its own.
- `ba377e7a`, the RoPE-θ fix. On Transformers 5 a config carries both a
top-level `rope_theta` and a
`rope_parameters` dict; the real base lives in the dict while the class
default (10,000 for Qwen3)
stays visible as the flat attribute. Reading the flat field first builds
a draft whose RoPE base is
100× off a Qwen3-8B target's, which trains and exports without
complaint. Both the training-side
enforcement and the exporter's `_get_rope_theta` are affected on `main`
today.

Both are @h-guo18's work and belong to their branches; they are carried
here only so that this PR
stands on its own. **If those branches land first, this PR can be
rebased onto them and the two
commits dropped**, and they can equally be split out now if that is
easier to review.

The same applies to `dflash_fp32_master_weights`, which is also in
flight on
`haoguo/dflash-fp32-master-weights`. The field name is shared
deliberately so that there is only
ever one knob rather than two spellings of it, and both versions default
to off. Whichever lands
first, this PR can be rebased onto it.

### Usage

Train with the shipped recipe:

```python
from modelopt.recipe import load_recipe

config = load_recipe("general/speculative_decoding/lilicorr.yaml")
# Qwen3-8B target, 6 epochs, block size 16 (15 drafted slots, 16 verified),
# DFlash decay objective at gamma 7.0, fp32 master weights for the draft.
```

Or convert directly:

```python
import modelopt.torch.speculative as mtsp

config = {
    "dflash_block_size": 16,
    "dflash_loss_objective": "decay",
    "dflash_loss_decay_factor": 7.0,
    "dflash_fp32_master_weights": True,
    "dflash_lilicorr_w_ce": 0.25,
    "dflash_lilicorr_w_margin": 0.0,
    "dflash_lilicorr_w_pen": 0.25,
    "dflash_architecture_config": {
        "num_hidden_layers": 5,
        "projector_type": "lilicorr",
        "lilicorr_candidate_topk": 8,
        # Optional, and all-or-nothing: adding these two keys wraps every draft
        # sublayer in DFlash2's grouped convolution. Requires the DFlash2 variant.
        # "conv_kernel_size": 2,
        # "conv_group_size": 16,
    },
}
mtsp.convert(model, [("dflash", config)])
```

### Results

Six drafters for a **Qwen3-8B** target, all trained **in ModelOpt on one
matched contract** — the
same corpus, schedule and block geometry for every arm, so no row
carries a training advantage.
Training data is NVIDIA's
[Nemotron Post-Training Dataset
v2](https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2)
with the multilingual split excluded, generated from the target with
**thinking disabled**;
**6 epochs**; block size 16 (15 drafted slots, 16 verified); DFlash
decay objective at gamma 7;
**8 nodes × 8 H100, global batch size 64** (one sequence per device, no
gradient accumulation).

All six were then exported and served through SGLang on a **single H100
80GB**, `tp_size 1`, at
concurrency 1, greedy, `fa3`, mean of two replicates, with the whole
node held exclusive per
benchmark. Speedup is output tokens/s against an autoregressive baseline
measured in the same
allocation.

Cells are `acceptance length / speedup-vs-AR`; **★ fastest, ☆ second
fastest**:

| benchmark | LiLiCorr+conv | LiLiCorr | DSpark | DFlash2 | Domino |
DFlash |
|---|---|---|---|---|---|---|
| gsm8k | ★ 7.715 / 5.26x | ☆ 7.557 / 5.22x | 7.375 / 4.86x | 7.252 /
5.06x | 7.225 / 4.87x | 6.341 / 4.59x |
| math500 | ★ 9.241 / 6.54x | ☆ 9.064 / 6.52x | 9.012 / 6.15x | 8.999 /
6.49x | 8.976 / 6.25x | 7.909 / 5.88x |
| aime25 | ★ 8.285 / 6.03x | ☆ 8.156 / 6.03x | 8.043 / 5.61x | 7.967 /
5.91x | 8.066 / 5.77x | 7.126 / 5.44x |
| humaneval | ★ 7.393 / 4.01x | 7.077 / 3.93x | 7.163 / 3.72x | ☆ 7.081
/ 3.95x | 6.864 / 3.73x | 6.156 / 3.68x |
| mbpp_sanitized | ★ 5.999 / 4.18x | ☆ 5.849 / 4.13x | 5.888 / 3.95x |
5.685 / 4.05x | 5.679 / 3.91x | 5.027 / 3.70x |
| livecodebench | ★ 7.975 / 5.40x | ☆ 7.754 / 5.33x | 7.775 / 5.10x |
7.601 / 5.26x | 7.553 / 5.04x | 6.808 / 4.88x |
| alpaca_eval | ☆ 3.697 / 2.69x | ★ 3.656 / 2.70x | 3.588 / 2.52x |
3.467 / 2.58x | 3.627 / 2.59x | 3.222 / 2.46x |
| mtbench | ★ 4.014 / 2.94x | ☆ 3.939 / 2.93x | 3.957 / 2.78x | 3.748 /
2.80x | 3.948 / 2.84x | 3.478 / 2.67x |

**Against every other approach in the table, LiLiCorr with convolutions
is the fastest on all eight
benchmarks.** Plain LiLiCorr is the fastest on seven of the eight; the
exception is humaneval, a
164-prompt slice, where DFlash2 is ahead by 0.5%.

`DFlash` is the deliberately head-free control; every head clears it by
+7.60% to +21.67% on
acceptance, which is the check that a head actually loaded. Reproducing
the `LiLiCorr+conv` column
additionally needs the DFlash2 variant.

Acceptance length is bit-reproducible under greedy decoding and its
replicate spread here was 0.00%
on every benchmark; throughput has a ~0.2% floor.

### What `dflash_fp32_master_weights` does, and what it is worth

Today the draft is cast to the frozen base model's dtype — bf16 — before
the optimizer is built.
AdamW then allocates its moments with `zeros_like(p)`, so the
**optimizer state becomes bf16 too**.
That is the problem: bf16 has too few mantissa bits to represent the
small updates Adam's second
moment accumulates, so those updates round away and the effective step
size decays on its own,
independently of the learning-rate schedule.

The flag is standard mixed precision instead: the draft's master weights
stay in fp32 while the
matmuls run in bf16. It requires a bf16 autocast around the forward,
which HF `Trainer` supplies
under `TrainingArguments.bf16`. Paths that do not go through the Trainer
— evaluation,
`pseudo_speculative_generate`, a plain `convert()` and forward —
currently need the caller to
supply it, and no shipped recipe exercises those (`estimate_ar: false`,
`do_eval: false`). Making
the draft supply its own autocast is a follow-up, held back from here on
review because it touches
every DFlash variant and wants e2e coverage of the existing recipes.

Compute speed is unchanged. The cost is memory, about 12 bytes per
parameter for the weight plus
Adam's two moments instead of 6, plus a doubled gradient all-reduce
under DDP, since fp32
parameters mean fp32 gradients. Under FSDP2 that second cost is what
`MixedPrecisionPolicy(reduce_dtype=...)` exists to control.

It is worth **7 to 14 percent of acceptance length**, measured at the
end of training on gsm8k, and
it helps every projector type:

| arm | bf16 | fp32 | Δ acceptance length |
| --- | ---: | ---: | ---: |
| LiLiCorr | 6.8670 | 7.5573 | **+10.05%** |
| DFlash2 | 6.7396 | 7.2518 | **+7.60%** |
| Domino | 6.5854 | 7.2252 | **+9.71%** |
| DSpark | 6.4621 | 7.3752 | **+14.13%** |
| DFlash | 5.9030 | 6.3412 | **+7.42%** |

Every arm in the comparison table above was trained with it on, and
**both shipped recipes set it
`true`**, so the documented path gets it.

It defaults to **off**, so no existing DFlash, Domino or DSpark run
changes behaviour. Both shipped
LiLiCorr recipes set it `true`, which is the arithmetic their numbers
were trained with. Flipping
the default is a reasonable follow-up once the autocast above is in.

The draft is drawn in fp32 and, under this flag, kept there; an
unpromoted run rounds the same draw
to the base model's dtype. So the bf16 and fp32 rows of the table above
start from the same
initialization at the precision each trains in, rather than from two
different draws. A unit test
pins that.

The flag also survives a resume. `modify()` runs under `from_pretrained`
with the base model still
on meta and cannot place the draft at all, so `restore_draft_precision`
re-applies the dtype, the
device and the rotary buffer once the weights are loaded and before the
Trainer builds the
optimizer — the last point that can still decide the Adam moment dtype.
It also reloads the draft's
tensors at the dtype they were saved in, since checkpoints store the
draft in fp32 while the base is
bf16 and `dtype="auto"` gives every tensor one dtype.

@h-guo18 has the same field in flight on
`haoguo/dflash-fp32-master-weights`, plus an
HF-format-resume fix this PR does not have. The name is shared
deliberately so there is only ever
one knob; whichever lands first, the other should be dropped rather than
merged.

### Testing

- **257 CPU unit tests pass** across `tests/unit/torch/speculative/`,
including the existing DFlash,
Domino, DSpark and Eagle suites. 48 of them are new and cover LiLiCorr
specifically: conversion
routing, head geometry, the required-field validation, the three-term
objective and its absolute
weights, gradient reach into both the head and the drafter body, and the
export contract.
- Both recipes load and validate through `modelopt.recipe.load_recipe`.
- The three DFlash-wide changes are covered behaviourally: the fp32 flag
is checked on the
optimizer's moment dtypes rather than only on parameters, since the
moments are the point of the
change, and on the initialization described above; activation
checkpointing is asserted to leave
draft gradients bit-identical with the flag on and off; and the rotary
buffer is asserted present
after `modify()` on a real device while still deferred on meta, which is
the case the laziness
  existed for.
- The resume path has its own test: after a `save_pretrained` /
`from_pretrained` round trip,
`restore_draft_precision` is asserted to return the draft to fp32 with
its stored weights intact
and its Adam moments in fp32. Without it the draft comes back in the
base dtype with the flag
  still set, which is the failure it exists to prevent.
- `TestDFlashLazyRotaryEmb` was updated rather than left passing: it
asserted the rotary buffer does
*not* exist after convert, and the DDP fix deliberately changes that on
non-meta devices. The
  replacement pins the refined invariant in both directions.
- The published checkpoints were trained with this arithmetic, verified
rather than assumed: a
fingerprint over draft initialisation, loss and gradients is compared
against the pre-review tree
for both `dflash` and `lilicorr`. Loss and gradients are **bitwise
identical**. Initialisation
moves, by less than bf16 resolution, and that is the single-dtype change
described above.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — every addition is opt-in. The
new `projector_type` is
selected only by config, `dflash_fp32_master_weights` defaults to off,
and the
activation-checkpointing and DDP fixes preserve behaviour. No existing
default changes.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in
`CONTRIBUTING.md`: ✅ — no new dependencies. Four files carry `# Adapted
from
https://github.com/sgl-project/SpecForge/...` headers for the DFlash
backbone and loss they derive
from (Apache-2.0), matching the attribution already on `hf_dflash.py` in
this repo. The two
commits described above are @h-guo18's, cherry-picked with authorship
and sign-off preserved.
- Did you write any new necessary tests?: ✅ — 48 new CPU tests, plus the
updated rotary test.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — will run `/claude review`
once opened.

### Additional Information

The convolutional recipe is the memory worst case: at an 8B target,
combined with fp32 master
weights, it may need `training.gradient_checkpointing: true` to fit on
80 GiB, and it fits without at
4B. Checkpointing is mathematically neutral — same objective, same data
order, same resulting model —
but it trades step time for memory, so a run using it is not
step-time-comparable with one that does
not. The recipe header says so.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added LiLiCorr speculative decoding with candidate-lattice reranking,
configurable objectives, metrics, export support, and optional grouped
convolutions.
* Added FP32 master-weight support with improved mixed-precision
behavior and gradient checkpointing.
* Added LiLiCorr training recipes and a Qwen3-8B launcher configuration.
* **Bug Fixes**
* Improved rotary-embedding configuration handling and corrected DFlash
distributed-training hangs.
* Added validation for invalid LiLiCorr configurations and improved
exported reranking metadata.
* **Documentation**
* Expanded guidance for FP32 master weights, training workflows, and
LiLiCorr configuration.
* **Tests**
* Expanded coverage across training, evaluation, generation, export, and
checkpoint workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: mrusanovsky <mrusanovsky@nvidia.com>
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 00:35:31 +08:00
haoxiz-nvidia acdf330414 Add TensorRT-RTX ABI EP support for ONNX quantization (#2262)
### What does this PR do?

Type of change: new feature

Adds opt-in support for using the standalone TensorRT-RTX ABI Execution
Provider during ModelOpt ONNX quantization.

Users select the ABI backend with:

`--calibration_eps=NvTensorRtRtx --trt_rtx_backend=abi`

When selected, ModelOpt imports and registers the installed TensorRT-RTX
ABI provider before creating the ONNX Runtime inference session. The
backend selection is propagated through INT8, FP8, and INT4 AWQ
calibration paths, including the Windows GenAI LLM quantization example.

The existing `--calibration_eps=NvTensorRtRtx` behavior remains backward
compatible. The `legacy` backend is still the default and continues to
use TensorRT-RTX libraries supplied through `PATH`.

For Windows x64 with Python 3.11 or newer, the ONNX dependencies now
include:

- `onnxruntime-gpu~=1.26.0`
- `onnxruntime-ep-nv-tensorrt-rtx-cu13==0.4.0`

Keeping `onnxruntime-gpu` allows users to select either CUDA EP or
TensorRT-RTX ABI EP for calibration. Windows-on-Arm source-build
instructions are intentionally out of scope and will be documented
separately.

### Usage

```powershell
python -m modelopt.onnx.quantization `
  --onnx_path="C:\path\to\Llama-3.2-3B-Instruct\model.onnx" `
  --model_id="C:\path\to\Llama-3.2-3B-Instruct\config.json" `
  --quantize_mode=int8 `
  --output_path="C:\path\to\int8_abi\model.onnx" `
  --calibration_eps=NvTensorRtRtx `
  --trt_rtx_backend=abi `
  --use_external_data_format `
  --high_precision_dtype=fp32 `
  --log_level=INFO

### Testing
unit test have been added

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅
- Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ 
- Did you get Claude approval on this PR?: pending



<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit

- **New Features**
  - Added optional TensorRT-RTX ABI backend support for ONNX calibration on Windows ARM64.
  - Added `legacy` and `abi` backend selection to quantization APIs and command-line tools; `legacy` remains the default.
  - Added validation for unsupported backends and incompatible TensorRT plugin configurations.
  - Updated Windows ARM64 installation support and platform-specific package configuration.

- **Documentation**
  - Updated Windows installation guidance, Python compatibility requirements, ARM64 setup, and verification steps.
  - Documented the new TensorRT-RTX backend command-line option.

- **Tests**
  - Added coverage for ABI provider registration, backend validation, and compatibility checks.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Haoxi Zhang <haoxiz@nvidia.com>
2026-09-09 04:04:46 +00:00
Frida HouandClaude Opus 5 0688761ce9 feat(export): support multimodal and MTP models in layerwise export (#2303)
### What does this PR do?

Type of change: New feature

**Layerwise export now supports multimodal and MTP models.** Both were
refused outright, and
both were refused for the same reason: `finalize()` was called from
inside
`layerwise_calibrate`, which is the wrong scope for it.

**1. Calibration does not know which model the checkpoint describes.**
It only sees the
module it was handed. A VLM calibrates its *language model*, so the
shards, the exclusions
and `config.json` all came out describing that submodel rather than the
whole VLM. Moving the
call out lets the caller root the exporter at the parent — and without
the key prefixing,
tower collection or ambient parent handle an earlier attempt needed,
because the decoder
layers are the same objects from either root.

**2. Calibration runs before things the export needs exist.** Orphaned
MTP weights are loaded
*after* calibration, by which point every shard had already been
written, so they could not
be passed at all. After the move they are an ordinary argument to
`finalize()`, with no
staging attribute stashed on the model.

### How it works

The exporter is created by whoever owns the export and **announced on
the model** that
`mtq.quantize` is given. Calibration picks it up, binds it, and drives
it per layer; the
export that follows reads it back and finishes the checkpoint:

```python
LayerwiseExporter(full_model, export_path).announce(language_model)
mtq.quantize(language_model, quant_cfg, forward_loop=loop)
...
getattr(full_model, LAYERWISE_EXPORTER_ATTR).finalize(extra_state_dict=mtp_state_dict)
```

Calibration and export are handed *different* models, so `announce()`
publishes the exporter
on each end separately: the caller announces on the model being
calibrated, and `bind()`
announces on the export root. Neither side has to know where the other
looked, and the lookup
stays an O(1) `getattr` rather than a `named_modules()` scan — worth
avoiding at roughly
1.65 µs/module, or ~500 ms on a Kimi-K3-sized model. For a non-VLM both
roots are the same
object and the second announcement is a no-op. `finalize()` clears every
attachment it
recorded, so the module graph does not retain a live exporter
afterwards.

`mtq.quantize` and `mtq.calibrate` are **unchanged** — a layerwise-only
feature does not
belong in the public quantization API. The attribute follows
`_mtp_layer_prefixes`, which
crosses the same calibration→export boundary the same way
(`hf_ptq.py:538` sets it,
`unified_export_hf.py:870` reads it back).

Construction is inert: `__init__` records only the export root and the
directory, because the
caller builds it before `mtq.quantize`, when there are no quantizers yet
to validate or read
a config from. `bind()` does that, called from calibration after
quantizer insertion and
before any layer is converted — the only window where both hold, and the
same instant the
exporter used to be constructed, so unsupported models still fail in
seconds rather than
hours. Only the calibration pass that sets `export_dir` drives the
exporter: a list-form
algorithm runs one pass per entry, and an earlier one must not convert
layers a later one
still has to calibrate.

### Usage

Nothing changes for a plain layerwise-export recipe:
`layerwise.export_dir` still drives it.
Pre-attaching an exporter is the opt-in for the two cases that need it —
a checkpoint whose
root is wider than the calibrated model, and orphaned tensors to merge
at the end.

The one behaviour change for a config-only caller is that `mtq.quantize`
now writes the layer
shards but no longer finishes the checkpoint. Both exit paths warn with
what is still owed,
and `LayerwiseConfig.export_dir`'s description has been corrected — it
previously promised "a
complete, loadable checkpoint when the last layer lands" and still
listed multimodal and MTP
as raising `NotImplementedError`.

### Testing

`tests/gpu/torch/export/test_layerwise_export.py` — **29 passed**.
Beyond the 24 inherited
from #2136, five new ones, each with a negative control confirming it
fails without its fix:

- orphaned MTP tensors reach the tail shard *and* the index
- an exporter rooted at the parent widens the checkpoint's namespace
- the config-only path announces an exporter that can be finished, and
finalize clears it
- only the pass that sets `export_dir` drives the exporter
- an exporter whose root holds a different number of layers is refused
at `bind()`

Full suites: `tests/gpu/torch/export` + `tests/gpu/torch/quantization`
**1012 passed / 55
skipped**, `tests/unit` **3318 passed / 15 skipped**, pre-commit clean.
Both suites also
report failures in `test_implicit_gemm.py` (FP4 conv kernels),
`test_triton_fa_p_qdq.py`,
`test_autocast_quantize_int8` and `test_engine_builder.py` collection;
all reproduce unchanged
on `main` and none touch the paths in this diff.

Measured against the whole-model exporter on a tiny Gemma3-VL, towers
prepared exactly as
`hf_ptq` does:

```
keys: baseline=80  layerwise=80   only-baseline=[]  only-layerwise=[]
differing values: 0
vision tower present: True     VLM namespace: True
config.json is the VLM: True   hf_quant_config match: True
exclude_modules: ['language_model.lm_head', 'vision_tower.vision_model*']   (both sides)
```

#### End-to-end through `hf_ptq.py`

Same FP8 recipe both sides; the baseline drops `layerwise.export_dir`
and is exported by
`main`, so the diff isolates this PR. Every tensor matches in key,
dtype, shape and value,
and `config.json` / `hf_quant_config.json` match too.

| Model | Covers | Keys | Differing |
|---|---|---|---|
| Qwen3-VL-8B-Instruct | multimodal | 1254 = 1254 | 0 |
| GLM-4.7-Flash | MoE + MTP | 28119 = 28119 | 0 |

The VLM checkpoint keeps the vision tower unquantized (351
`model.visual.*` keys, no
`weight_scale` among them) while the language model is FP8. The MTP run
reports 212 orphaned
tensors; all 212 land in `model-tail.safetensors` and in the index, with
`model.layers.47*` in
`exclude_modules`.

**Not yet validated:** an accelerate-offloaded run, and a serving canary
on the exported
checkpoints.

### Refusals

`export_dir` without `enable`, and an exporting algorithm entry with no
calibration method,
are both refused before calibration starts — neither reaches the
per-layer pass, so both
would otherwise export nothing. The early gate is a heuristic on the
recipe, so `hf_ptq` also
raises a plain `RuntimeError` at export time if calibration turned out
not to have run; that
backstop, not the gate, is what makes the failure legible on paths the
recipe check cannot
predict.

`bind()` requires the layers calibration will drive and refuses a root
that discovers a
different number of them. Only the count is checked here: `export_layer`
already rejects a
reordering or a substituted module on its first call, and a length
difference is the one
mismatch it structurally cannot catch — every call would pass and
`_write_index` would then
open a shard that was never written, at the very end of the run.

Orphan tensors are merged into the tail with no collision check,
matching the whole-model path
(`unified_export_hf.py:1623`). `load_mtp_weights` returns exactly the
keys absent from
`model.state_dict()`, so a collision with an exported tensor is not
reachable through the only
producer, and a guard would only make the two export paths diverge.

### Why not reuse `export_hf_checkpoint`

It was the first idea and it is the most expensive one. Its transformers
path is whole-model
at every step — `_prepare_moe_inputs`,
`requantize_resmooth_fused_llm_layers` (which runs a
dummy forward that would fail on already-converted layers),
`_process_quantized_modules`, a
full `model.state_dict()` in host RAM, then `save_pretrained` rewriting
shards already on disk
— and it raises outright under `has_accelerate_offload`.
`save_pretrained(state_dict={})` is
not an escape either: safetensors' shared-storage check fires on MoE
even with an empty dict.

The natural consolidation target is the **streaming** exporter, which is
already most of
`finalize()`: 122 lines vs 74, sharing `decoder_owned_ids`,
`enable_weight_access_and_writeback`, `_dispatch_export_handler`,
`_reconstruct_fused_moe_linear`, `_add_mtp_exclusions`,
`_postprocess_single_tensor`,
`requires_weight_materialization` and `save_non_weight_artifacts`.
Folding them together needs
roughly four knobs: skip the whole-model prep, skip layers already
written, seed the index
with the existing shards, and inject the quant config. That is a
separate change and
deliberately not in this one.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ —
`mtq.quantize`/`mtq.calibrate` signatures are
unchanged, and a recipe that only sets `layerwise.export_dir` behaves as
before. The one
behaviour change is that `mtq.quantize` no longer finishes the
checkpoint on its own:
callers must now call `finalize()` on the exporter, which calibration
leaves on the model.
- If you copied code from any other sources or added a new PIP
dependency, did you follow
  guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ❌ — pending.
- Did you get Claude approval on this PR?: ❌ — the last review's
findings are all addressed;
  needs a re-run.

### Additional Information

Follow-ups this enables: #2259 (MTP) reduces to close to nothing, and
the multimodal work in
#2218 no longer needs `export_parent`, the key prefixing, or the tower
collection.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
- Layerwise export now supports clearer control over export locations
and calibrated layer handling.
- Export workflows provide improved support for resuming, sharded
checkpoints, mixture-of-experts models, and nested model namespaces.

- **Bug Fixes**
  - Improved handling of exported checkpoint shards and extra tensors.
- Added clearer warnings when exports require completion before loading.

- **Documentation**
- Clarified that layerwise exports write shards during calibration and
require an explicit finalization step.
- Documented that the in-memory model is not suitable for inference
after layerwise export.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 20:48:47 +00:00
Chenjie LuoandClaude Opus 5 19de0075cb Forward kv_cache_free_gpu_memory_fraction to the lm_eval TensorRT-LLM engine (NVBug 6701763) (#2300)
### What does this PR do?

Type of change: Bug fix

`scripts/huggingface_example.sh --kv_cache_free_gpu_memory_fraction` has
no effect on the `lm_eval` task: the value is parsed by `parser.sh`,
printed, and then dropped.

lm-eval's built-in `trtllm` backend
(`lm_eval.models.trtllm_causallms.TRTLLM.__init__`, which this example
switched to in #2066) accepts `**kwargs`, but builds
`KvCacheConfig(enable_block_reuse=False)` and passes `LLM(...)` a fixed
set of keys — `kwargs` is never merged in. So an extra `--model_args`
entry is accepted by the CLI and silently discarded, and the KV cache is
sized from TensorRT-LLM's default `free_gpu_memory_fraction=0.9`. There
is no way to fix this from the caller: `--model_args` only yields
scalars, so a `KvCacheConfig` object cannot be passed in either.

On a GH200 that means ~119.6 GiB of KV cache (`119.55 / 0.9 ≈ 132.8 GiB
free`), leaving 87.8 MiB free, and `prompt_logprobs` deserialization
then OOMs asking for 2.82 GiB.

`examples/llm_eval/lm_eval_trtllm.py` already exists to patch this
backend (its `_parse_logprobs` misaligns TensorRT-LLM's
`prompt_logprobs` by one). It now also injects the fraction into the
`KvCacheConfig` the backend builds, defaulting to 0.8 — the same default
`parser.sh` declares, and below TensorRT-LLM's 0.9.
`huggingface_example.sh` passes the parsed value through in
`--model_args`.

Scoped deliberately to the `lm_eval` path: the `quant` smoke test and
`mmlu` go through `modelopt.deploy.llm.LLM` (0.7, hardcoded) and
`simple_eval`/`livecodebench` through `trtllm-serve` (0.9); those are
left as they are.

### Usage

```bash
# Via the example script (parser.sh default 0.8)
scripts/huggingface_example.sh --model $HF_PATH --quant fp8 --tp 1 \
    --tasks quant,lm_eval --lm_eval_tasks mmlu --lm_eval_limit 50 \
    --kv_cache_free_gpu_memory_fraction 0.5
```

```bash
# Standalone, via lm-eval's --model_args
python lm_eval_trtllm.py --model trtllm \
    --model_args model=<ckpt>,tokenizer=<tok>,max_input_len=4096,kv_cache_free_gpu_memory_fraction=0.5 \
    --tasks mmlu --batch_size 8
```

### Testing

- `pytest tests/examples/llm_eval/test_lm_eval_trtllm.py` — 21 passed
(lm-eval 0.4.12, no GPU).
- The new tests instantiate the **real** upstream `TRTLLM.__init__`
through `create_from_arg_obj`, with `tensorrt_llm` and the tokenizer
stubbed, and assert the engine receives
`KvCacheConfig(enable_block_reuse=False, free_gpu_memory_fraction=0.5)`;
that an unset key still yields 0.8 rather than 0.9; and that the patch
does not outlive the constructor. Reverting the fix fails 3 of them.
- Tripwire test asserts upstream still neither declares nor forwards the
argument, so this shim gets deleted rather than silently kept once
lm-eval fixes it.
- `pre-commit run --files <changed>` clean (ruff, mypy, bandit,
markdownlint); `bash -n` on the modified script.
- Not run: the GPU end-to-end
`tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8`, which
exercises `lm_eval` through the modified script — no GPU in this
environment.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — the `lm_eval` KV cache goes
from TensorRT-LLM's 0.9 to 0.8, which is strictly more conservative;
`parser.sh`'s declared default is unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

NVBug 6701763. The 0.9 default on this path arrived with #2066 and was
documented as a known limitation in `examples/llm_eval/README.md` ("the
KV cache uses 90% of free GPU memory rather than 70%"); that note is
replaced by the working knob.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Fixed the TensorRT-LLM evaluation workflow so
`kv_cache_free_gpu_memory_fraction` is correctly passed to the backend.
- The setting now defaults to `0.8`, providing more predictable GPU
memory allocation for KV-cache usage.

- **Documentation**
- Updated the TensorRT-LLM evaluation example and usage guidance to
describe the KV-cache memory setting and its default behavior.
- Updated the Hugging Face example to pass the configured KV-cache
memory fraction.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 13:06:51 -07:00
Ajinkya RasaneandCodex 2aefe08f20 [OMNIML-5613] Quantize ResNet residual adds in torch ONNX example (#2024)
### What does this PR do?

Type of change: Bug fix

Adds recipe-backed FP8 and INT8 residual quantization for timm ResNet
models in the torch ONNX example:

- Adds FP8 and INT8 PTQ recipes that enable a shortcut quantizer
immediately before each residual `Add`.
- Inserts shortcut quantizers after ModelOpt module conversion so
recipes configure and calibrate them in the normal quantization pass.
- Adds `--recipe` support for PTQ and AutoQuantize recipes and renames
`--quantize_mode` to `--qformat`.
- Verifies all 16 ResNet-50 residual additions have shortcut Q/DQ
immediately before the `Add`.

### ResNet support scope

ResNet and other convolutional architectures are supported only with FP8
and INT8. AutoQuantize, MXFP8, NVFP4, and INT4_AWQ are not supported for
ResNet because TensorRT has limited convolution kernel support.
Transformer architectures containing individual Conv2d layers continue
to use format-specific Conv overrides.

### Usage

```bash
python examples/torch_onnx/torch_quant_to_onnx.py \
    --timm_model_name=resnet50 \
    --recipe=timm/resnet/ptq/fp8 \
    --onnx_save_path=resnet50.onnx
```

Use `timm/resnet/ptq/int8` for INT8. Without `--recipe`, `--qformat`
selects a built-in quantization preset.

### Testing

- All configured pre-commit hooks passed, including recipe schema and
license validation.
- Focused AutoQuantize recipe mapping regression passed.
- FP8/INT8 recipe export coverage verifies all 16 ResNet-50 shortcut
Q/DQ pairs.
- TensorRT engine builds passed for FP8 and INT8 on Ada.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update the changelog?: ✅

> 🤖 _Generated by Codex (AI agent)._

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <noreply@openai.com>
2026-09-08 15:50:00 -04:00
Keval MorabiaandClaude Opus 5 58eafdf172 Fix the llm_eval README commands that no longer run as written (#2358)
### What does this PR do?

Type of change: Documentation (plus one small example-script fix)

An audit of `examples/llm_eval/README.md` against the current scripts
(nvbug 6701343) found several documented commands that no longer run as
written:

- **T5 / seq2seq.** `--model hf-seq2seq` is not a registered lm-eval
backend in any version this example supports — the string does not
appear in the 0.4.12 or 0.4.13 wheels, so the command fails at model
lookup. `HFLM` detects encoder-decoder models from `config.json`, so the
example now uses `--model hf` and mentions `backend=seq2seq` as the
override for checkpoints lm-eval cannot classify. No ModelOpt-side
change was needed: encoder-decoder calibration already works (verified
below).
- **auto_quantize format list.** `FP8_DEFAULT_CFG|NVFP4_DEFAULT_CFG` was
shown as a literal value in both README locations, but each
comma-separated entry is resolved with `getattr(mtq, ...)` and that name
does not exist. Now shows a valid list, spells out the choices, and
names the placeholder consistently with the surrounding block.
- **`vllm serve`.** A missing line continuation meant `--port` ran as a
separate shell command.
- **MMLU setup.** Dropped a stray `cd ..` left over from the 0.11
examples release. It leaves `examples/llm_eval`, where both `mmlu.py`
and its default `--data_dir data/mmlu` live;
`hf_ptq/scripts/huggingface_example.sh` correctly stays put throughout
its MMLU flow, so the README was the only thing out of step.
- **`run_simple_eval.sh`.** Documented the optional fifth argument
(`--examples`), which `huggingface_example.sh` already passes as
`$SIMPLE_EVAL_LIMIT`.

Two changes beyond the docs:

- **`quantization_utils.py`:** under `auto_quantize`, a `quant_cfg`
string was iterated character by character, so a single format failed
with the baffling `AttributeError: module 'modelopt.torch.quantization'
has no attribute 'F'`. Normalized `str -> list` at the point the list is
consumed, which covers both `mmlu.py` and `lm_eval_hf.py` rather than
one caller. This also honors the existing `str | list[str]` annotation.
- **`requirements.txt`:** added the missing `openai`. `modeling.py`
imports it unconditionally and `lm_eval[api]` supplies only `tiktoken`,
so every documented `mmlu.py` command died with `ModuleNotFoundError` on
a clean install of the stated requirements.

Note on the filed report: its item 3 claimed `mmlu.py` fails to split
the comma-separated config list. That does not reproduce — `mmlu.py`
uses `fire`, which already parses `A,B,NONE` into a tuple, and the
unmodified script completes `auto_quantize` fine. Applying the suggested
`quant_cfg.split(",")` would have *broken* the documented command with
`AttributeError: 'tuple' object has no attribute 'split'`. The
`quantization_utils.py` change above addresses the real adjacent defect
instead. Pushback recorded on the bug.

### Usage

No new API or flag. The corrected commands:

```bash
# T5 / encoder-decoder (was: --model hf-seq2seq, which does not exist)
python lm_eval_hf.py --model hf --model_args pretrained=t5-small \
    --quant_cfg FP8_DEFAULT_CFG --tasks <comma separated tasks> --batch_size 4

# auto_quantize search list (was: W4A8_AWQ_BETA_CFG,FP8_DEFAULT_CFG|NVFP4_DEFAULT_CFG,NONE)
python mmlu.py --model_name causal --model_path <model> \
    --quant_cfg W4A8_AWQ_BETA_CFG,FP8_DEFAULT_CFG,NONE --auto_quantize_bits 4.8 --batch_size 4

# simple evals, optional 5th arg
bash run_simple_eval.sh <model> <evals> <max_tokens> <port> [num examples per eval]
```

### Testing

Ran on 2x RTX 6000 Ada with a tiny Qwen3 and a locally synthesized MMLU
tree (no download):

- **`mmlu.py --auto_quantize_bits` with the documented comma-separated
list** — completes quantization on both the unpatched and patched
script, confirming the reported item 3 is a false positive. Probed
`fire` directly: bare, quoted and `--flag=value` forms all yield
`('W4A8_AWQ_BETA_CFG', 'FP8_DEFAULT_CFG', 'NONE')`.
- **`mmlu.py --auto_quantize_bits` with a single format** — proved the
new guard fires by reverting it: without the change the run dies with
`AttributeError: module 'modelopt.torch.quantization' has no attribute
'F'`; with it, the run reaches a legitimate domain assertion
(`effective_bits 4.8` cannot be below FP8's 8 bits).
- **Encoder-decoder calibration** — quantized a T5 with
`FP8_DEFAULT_CFG` through `quantize_model` and confirmed encoder,
decoder and cross-attention (`EncDecAttention`) layers all calibrate
with real amax values. This is what settled keeping the T5 example
rather than deleting it.
- **`vllm serve` snippet** — parsed the fixed block with `bash`;
`--quantization`, `--port` and `--tensor-parallel-size` now all belong
to one command.
- **`run_simple_eval.sh`** — confirmed the 4-arg form is unchanged and
the 5-arg form emits `--examples 16`.
- **Lint** — `ruff-check`, `ruff-format`, `markdownlint-cli2`, `typos`,
`bandit`, `mypy`, `requirements-txt-fixer`, `mixed-line-ending` all
pass.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — added
`openai` to `examples/llm_eval/requirements.txt`; it is Apache 2.0
(permissive), so no codeowners exception is needed. It is not a new
runtime dependency of the library, and `run_simple_eval.sh` already `pip
install`s it.
- Did you write any new necessary tests?: N/A — docs plus a two-line
defensive normalization in an example util. `mmlu.py` cannot be imported
without `openai`/`rwkv`/`tiktoken`, so a hermetic unit test would need
more stub scaffolding than the line it guards; verified by direct
execution instead, as above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — examples-only documentation cleanup, not a feature, breaking
change, deprecation, or a critical bug from a previous release.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Fixes nvbug 6701343 / OMNIML-5806. Item 3 of the filed report is a false
positive; pushback and evidence are recorded in a comment on the bug.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
- Auto-quantization now supports comma-separated format configurations.
  - Added an optional example-limit setting for Simple Evals.
  - Added OpenAI support for LLM evaluation examples.

- **Documentation**
  - Clarified encoder-decoder model usage with `lm_eval`.
- Added instructions for running MMLU from the evaluation examples
directory.
  - Corrected the vLLM command formatting.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-09 00:44:23 +05:30
Ajinkya RasaneandCodex 5c123ce183 [OMNIML-5563] Add PETR ONNX PTQ and accuracy evaluation example (#2180)
### What does this PR do?

Type of change: new example, example simplification, and
backward-breaking example migration

Adds end-to-end PETRv1/PETRv2 ONNX PTQ and reduces PETR/FAR3D to one
shared workflow:

- quantizes the shared VoVNet image backbone/encoder to INT8 or FP8;
- runs both the selected historical and current PETRv2 six-camera sweeps
through the same precision-matched TensorRT backbone engine using
distinct execution contexts during accuracy evaluation;
- keeps the PETR head and FAR3D decoder in their exported mixed
FP16/FP32 precision;
- reuses one NPZ calibration format, VoVNet exclusion helper,
quantization entry point, and TensorRT runner;
- does not change generic Model Optimizer calibration behavior or its
public CLI.

### Container boundary

Both examples use two targets from one Dockerfile, with no virtual
environments:

- `evaluator`: a digest-pinned `nvcr.io/nvidia/pytorch:22.06-py3` base
with the legacy PyTorch 1.13.1/OpenMMLab stack for source setup,
metadata generation, ONNX export, direct PyTorch calibration capture,
and final accuracy evaluation;
- `modelopt`: a digest-pinned `nvcr.io/nvidia/pytorch:26.07-py3` base
for Model Optimizer, ONNX Runtime CUDA, AutoCast, INT8/FP8 quantization,
and TensorRT engine builds.

Both targets use TensorRT `11.1.0.106`. Engines are built and evaluated
on the same GPU architecture. Final metrics remain in the evaluator
because they import the legacy model-framework postprocessing and
dataset code; only artifacts cross the container boundary through the
shared workspace.

PETR is used without patches. FAR3D applies only the official
`patch/far3d.patch` from the pinned NVIDIA DL4AGX revision. This PR
carries no patch files.

### Evaluator dependencies

The dependencies intentionally installed without transitive dependencies
are listed in `requirements-evaluator-nodeps.txt`. Their pins rely on
runtime packages supplied by the digest-pinned PyTorch 22.06 evaluator
base.

`lyft-dataset-sdk` is required only by mmdet3d's eager dataset import;
neither PETR nor FAR3D uses Lyft data. `flash-attn` remains in the main
evaluator requirements because its compiled installation uses the
evaluator build step rather than the intentionally dependency-free
legacy package step.

Fresh setup and dependency approval is requested for the final reduced
dependency set.

### Reproducible PETR metadata

The documented workflow mounts raw nuScenes read-only and creates a
writable dataset view using symlinks. It then runs the pinned
mmdetection3d converter and a temporary, untracked copy of PETR's pinned
sweep generator configured only for the validation prefix and writable
dataset root.

A clean run generated both metadata files with 6,019 validation records.
The referenced camera, lidar, and sweep paths are absolute and
resolvable through the writable dataset view.

### Example-local utilities

The per-batch NPZ streaming and TensorRT runtime utilities remain
example-local because they execute in the legacy evaluator, where Model
Optimizer is not installed. The core `CalibrationDataProvider` consumes
one in-memory mapping of stacked arrays and does not provide this
streamed per-file workflow.

### Validation

- Focused CPU tests: 10 passed.
- Broader ONNX quantization CPU tests: 326 passed.
- All applicable pre-commit and documentation checks, plus `git diff
--check`, passed.
- Rebuilt both Docker targets and verified their exact dependency
versions, imports, TensorRT `11.1.0.106`, GPU runtime initialization,
and absence of virtual environments.
- Generated both PETR metadata files from a clean writable dataset view
and verified 6,019 validation records plus resolvable data paths.
- PETRv1 passed a one-sample TensorRT regression smoke.
- PETRv2 passed FP16, INT8, and FP8 TensorRT smokes and full
6,019-sample validation. Both the selected historical and current sweeps
are computed by the matching backbone engine; accuracy evaluation no
longer extracts image features with PyTorch.
- FAR3D passed a recurrent two-frame TensorRT smoke covering plugin
loading and recurrent state.

TensorRT `11.1.0.106` mAP follows. PETRv2 was remeasured after
correcting its temporal feature path; the PETRv1 and FAR3D numerical
paths are unchanged.

| Pipeline | FP16 | INT8 | FP8 |
| --- | ---: | ---: | ---: |
| PETRv1: 1 backbone pass + fixed typed mixed FP16/FP32 head | 0.3778 |
0.3707 | 0.3756 |
| PETRv2: 2 serial backbone passes + fixed typed mixed FP16/FP32 head |
0.4102 | 0.3982 | 0.4084 |
| FAR3D: 1 encoder pass + fixed mixed FP16/FP32 decoder | 0.241 | 0.235
| 0.239 |

Normalized engine-only performance improvement over each matching FP16
pipeline:

| Pipeline | INT8 speedup | FP8 speedup |
| --- | ---: | ---: |
| PETRv1 | 1.49x | 1.29x |
| PETRv2 | 1.51x | 1.30x |
| FAR3D | 1.69x | 1.40x |

Performance was measured with TensorRT `11.1.0.106` on an NVIDIA RTX
6000 Ada Generation GPU using five interleaved trials per engine
component. Each component uses the median `trtexec`-reported GPU Compute
Time with data transfers disabled and CUDA Graphs enabled. Component
times are summed before normalization: PETRv1 uses one backbone pass
plus its fixed head, PETRv2 uses two serial backbone passes plus its
fixed head with no temporal cache assumed, and FAR3D uses one encoder
pass plus its fixed decoder. Absolute latency values are intentionally
not published.

Adapted files retain exact public-source references and upstream
notices, and the top-level license attribution is updated.

- Is this change backward compatible?: ❌
- Did you write the necessary tests?: ✅
- Did you update the changelog?: ✅

> 🤖 _Generated by Codex (AI agent)._

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
2026-09-08 17:32:56 +00:00
haoxiz-nvidia 4773f72f8a Docs: Add WOA documentation (#2264)
### What does this PR do?

Add WoA env setup guide. Includes build instruction of pyarrow, which
used by datatsets

### Usage

N/A

### Testing
N/A

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Added comprehensive Windows on Arm installation guidance, including
prerequisites, environment setup, dependency installation, verification,
and troubleshooting.
- Documented experimental Windows ARM64 support, supported quantization
formats, native dependency requirements, and Support Matrix details.
  - Expanded supported Windows Python versions through 3.13.
- Expanded TensorRT-RTX guidance for calibration, deployment, provider
setup, and standalone plugin usage.
- Clarified PyArrow requirements and local build instructions for
Windows ARM64.
- Added links to dedicated Windows on Arm installation resources and
shared TensorRT-RTX documentation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Haoxi Zhang <haoxiz@nvidia.com>
2026-09-08 15:04:04 +05:30
Keval MorabiaandClaude Opus 5 61757c9781 Support quantized Qwen3-VL / Qwen3.5-VL (dense + MoE) export from Megatron-Bridge and verify exported checkpoints (#2276)
### What does this PR do?

Type of change: Bug fix + new feature

Enables quantized **Qwen3-VL** and **Qwen3.5-VL** (dense and MoE) →
unified HuggingFace export from Megatron-Bridge, and fixes the bugs
found along the way (ten from testing, plus a further round from
review). Most of them produced a valid-looking checkpoint and a green
test run, so the PR also makes the export path verify its own output.

Review is easiest commit-by-commit — each of the eleven commits is
self-contained and independently green.

#### Two blockers

1. **The exporter rejected the Megatron-Bridge VLM wrapper.**
`GPTModelExporter` only unwrapped MCore's `LLaVAModel`, so
`Qwen3VLModel` raised `ValueError: Input to GPTModelExport must be a
megatron.core.models.GPTModel!`. It now unwraps any wrapper exposing
`.language_model`.
2. **A VLM QAD checkpoint couldn't be loaded back.** `distill.py` passes
`distill_submodule="language_model"`, so the checkpoint holds only the
language model and the load died on `KeyError:
vision_model.patch_embed.proj.weight`. The loader now reads the
checkpoint metadata and targets `.language_model` when there are no
vision weights.

#### Four silent-corruption bugs

3. **VLM QAD discarded all ModelOpt state** (shipped in 0.46).
`ModeloptStateManager` requires state on the **root** of whatever gets
checkpointed. `quantize.py` quantizes the VLM root, so PTQ anchors it
there — but QAD checkpoints only `language_model`, orphaning it. The
saved `modelopt_state_dict` was literally `[]`; the `*_quantizer._amax`
tensors were still present but got dropped on load
(`dist_ckpt_strictness="assume_ok_unexpected"`), and the export came out
plain BF16 with no `hf_quant_config.json`.
4. **Fused grouped-GEMM MoE experts were omitted entirely.** The MoE
dispatch had no `else`, so an architecture without an
`experts.linear_fc1` rule exported *zero routed experts*. This hit
**`Qwen3MoeForCausalLM`** — a registered, supported architecture with no
export test — not just VLMs. A tiny Qwen3-MoE exported 37 of 45 tensors,
exit 0, no warning.
5. **Qwen3.5's GatedDeltaNet output norm was off by exactly 1.0.**
Megatron stores that gamma zero-centered, HF centers it on 1. Correct
names, correct shapes, wrong values — invisible to any structural check.
Megatron-Bridge's importer confirms the convention
(`RMSNorm2ZeroCenteredRMSNormMapping`).
6. **The disabled-quantizer patterns silently no-op on Megatron paths.**
They are written against HuggingFace module names. `*mixer.conv1d*`
matches only because MCore and HF happen to agree on "mixer" for Mamba;
`*linear_attn.conv1d*` never matched (Megatron calls it
`self_attention.conv1d`), so the conv1d was calibrated.
`*linear_attn.in_proj_a/b*` **cannot** match at all — Megatron fuses all
six GDN sections behind one quantizer — so the alpha/beta gates the
recipe wants in BF16 were exported in FP8.

#### Four more bugs, found only by running real checkpoints

The tiny fixtures could not reach these; each came from a real model or
a real quant format.

7. **Routed experts were written in a layout no real Qwen3.5 checkpoint
uses.** Real Qwen3.5 stores experts packed as `[num_experts, out, in]`;
the mapping emitted per-expert names, so every routed expert was
dropped. The fixture actively hid this: transformers *unpacks* experts
on `save_pretrained`, so the saved reference agreed with the wrong
output. Fixed with a `transpose` kwarg on `_pack_name_remapping` plus a
`GroupedMLPPacking` rule, so fused `TEGroupedMLP` reaches the same
packed tensors — which is also what lets Qwen3.5 keep grouped GEMM
(**22.1 GB/GPU vs 38.9 GB/GPU** on a 20-layer, 256-expert model).
8. **`_grouped_mlp_packing` was broken for NVFP4.** It max-merged
`weight_scale`, but NVFP4 needs each expert's per-block scales *stacked*
with only the global `weight_scale_2` merged; it also dequantized packed
`uint8` against per-block scales, and passed `block_size=None`.
`weight_scale_2` is never populated in an FP8 run, so the whole branch
was dead code under FP8-only testing. `_grouped_mlp_slicing` gained
`quantize=False` so packing can quantize once over the stack, matching
`_pack_name_remapping`.
9. **`_mtp_prefix` corrupted every VLM's MTP tensor names.** It did
`prefix.replace("model", "mtp")` uncounted, so
`model.language_model.layers.{}` became `mtp.language_mtp.layers.0.*` —
tensors present and correctly valued, under names nothing loads.
LLM-only prefixes contain one occurrence, so this was invisible until a
VLM with MTP was exported.
10. **`load_multimodal_components` rejected HF repo ids.** `quantize.py
--hf_model_name_or_path Qwen/Qwen3.5-0.8B` worked, but the documented
export step failed with *"It should be a directory"*. Its sibling in the
same file already resolved repo ids via `snapshot_download`; now it does
too. This affected **every** VLM export.

`Qwen3_5ForConditionalGeneration` (dense Qwen3.5-VL) is now registered
for export and vision passthrough, which bugs 9 and 10 were blocking.

#### New: Qwen3.5-VL

`GatedDeltaNetSlicing` splits Megatron's fused `in_proj` (`[query, key,
value, z, beta, alpha]`) into HF's `in_proj_qkv` / `_z` / `_b` / `_a`,
taking sizes from the module's own `in_proj_split_sections` so TP
sharding falls out. Widening coverage to Qwen3.5's *gated
full-attention* layers then exposed a further split bug: gated attention
packs a per-head output gate beside each query head, so `_qkv_slicing`
split 192 rows as 96/48/48 instead of 128/32/32. It now derives the
group stride from `config.attention_output_gate`, matching
Megatron-Bridge's `split_qkv_weights`. The non-gated path is unchanged.

#### New: the export path verifies itself

- `assert_exported_checkpoint_matches` compares an exported checkpoint
against the model it came from — key set, shapes (accounting for NVFP4
`uint8` packing), safetensors index consistency, and values — replacing
existence-only assertions in all three export tests.
- `GPTModelExporter.save_pretrained` now raises if the export dropped
tensors the source checkpoint has, so *user* runs on architectures CI
never sees are protected too, not just tiny models.
- Loading a checkpoint whose quantizer tensors have no restorable state
now raises instead of silently loading unquantized.
- `assert_has_modelopt_state` replaces `rglob("modelopt_state")`, which
passes on an empty state; `assert_no_quantizers_matching` fails on
future HF↔Megatron name drift.

The mapping is also table-driven now: vision-tower prefixes live in
`all_mcore_hf_vision_passthrough_mapping` and
`with_language_model_prefix` is shared, so adding a VLM no longer means
editing `unified_export_megatron.py`. Five call sites that answered "is
this a VLM" three different ways now share `get_language_model` /
`is_vlm_config`.

### Usage

```bash
# Dense VLM (Qwen3-VL) -- no extra flags
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-VL-8B-Instruct \
    --quant_cfg nvfp4 --tp_size 2 \
    --export_megatron_path /tmp/Qwen3-VL-8B-NVFP4-megatron

torchrun --nproc_per_node 2 export_quantized_megatron_to_hf.py \
    --hf_model_name_or_path Qwen/Qwen3-VL-8B-Instruct \
    --megatron_path /tmp/Qwen3-VL-8B-NVFP4-megatron \
    --pp_size 2 --export_unified_hf_path /tmp/Qwen3-VL-8B-NVFP4-hf

# Gated MoE (Qwen3.5-VL, Qwen3-MoE) -- no extra flags either. The scripts derive the
# expert layout from the model config, so quantize / distill / export all agree.
# --no_moe_grouped_gemm forces SequentialMLP if you want it explicitly.
```

### Testing

All in `nvcr.io/nvidia/nemo:26.08` on 2x RTX 6000 Ada.

| Suite | Result | Time |
|---|---|---|
| `tests/examples/megatron_bridge/` (full) | 18 passed | 27m58 |
| `tests/gpu_megatron/torch/export/` | 38 passed | 2m13 |
| `tests/unit/torch/export/` | 186 passed | 1.5s |
| pre-commit (ruff, ruff format, mypy, bandit) | clean | — |
| `tests/examples/megatron_bridge/test_quantize_export.py` on **2 GPUs**
(`pp_size=2`) | 3 passed | 5m |

The export leg of `test_quantize_and_export` now scales with `num_gpus`
like its quantize leg
already did. Previously it was hardcoded to one process, so the
collective checkpoint load ran at
PP=1 on both the 1-GPU PR runner and the 2-GPU nightly — which is how a
guard that raised on only
some pipeline stages (and therefore hung the job) reached review. The
dense `qwen3` case was dropped
in exchange: `qwen3_moe` already covers the non-VLM script path,
`qwen3vl` covers a dense decoder,
and that case was the one exceeding the 300s cap in CI.

#### Model coverage

`tests/gpu_megatron` runs in-process and is cheap, so it owns
per-architecture **mapping**
correctness. The example tests spawn `torchrun` per step and are ~50x
slower per case, so they
cover **script wiring** only — CLI flags, recipe resolution, and
checkpoint hand-off between steps.

| Suite | Models |
|---|---|
| `test_unified_export_megatron` | llama, nemotron, nemotron_h, qwen3vl,
qwen3_moe, qwen3_5_moe_vl x {none, FP8, NVFP4, +/-KV} x {grouped GEMM,
SequentialMLP} + eagle / medusa / MTP (29 params) |
| `test_megatron_importer` | nemotron_h, llama export->import round-trip
|
| `test_moe_layout_choice` | per-architecture grouped-GEMM exportability
(6 architectures) |
| `test_distill_megatron` | KD loss mechanics |

| Model | prune | quantize+export | QAD | distill+export |
|---|:--:|:--:|:--:|:--:|
| qwen3 | Y | Y | Y | Y |
| qwen3_moe | - | **Y (new)** | - | - |
| qwen3vl | - | **Y (moved from QAD)** | - | - |
| nemotron_h | Y | **Y (new)** | - | - |
| qwen3_5_vl | - | - | - | Y |
| qwen3_5_moe_vl | Y | **Y (new, both expert layouts)** | Y | - |
| deepseek_v3 | Y | - | - | - |
| gemma3vl | Y | - | ~~manual~~ removed | - |

QAD's unique property is that ModelOpt state survives distillation,
which needs one LLM and one
VLM rather than one case per architecture. Moving the rest to
quantize+export drops a `torchrun`
launch each: QAD went from 3 CI cases to 2 while quantize+export went
from 1 to 4, adding two
architectures for about a minute.

#### Real-model validation

Tiny fixtures cannot catch layout or scale bugs that only appear at real
dimensions, so the export
path was run end-to-end on released checkpoints. This is where bugs 7-10
came from.

| Model | Run | Result |
|---|---|---|
| Nemotron-3.5-Lightning-30B-A3B | NVFP4 4o6 PTQ → export → MMLU |
**0.7825 ± 0.0105** (gate 0.75) |
| Nemotron-3.5-Lightning-30B-A3B | Minitron pruning | 22.28B/3.00B
active, **0.5944** (gate 0.58) |
| Qwen3.5-0.8B (dense VLM) | FP8 PTQ → export → MMLU | BF16 0.4895 →
**0.4832** (±0.0127) |
| Qwen3.5-35B-A3B, half-depth (20 layers, 256 experts) | FP8 + NVFP4 PTQ
→ export | keys + shapes + **values** match reference |
| Qwen3.5-35B-A3B, full | FP8 PTQ | OOM on 2x48GB (see below) |

The half-depth model keeps real weights, real dims and all 256 experts.
Both expert layouts produce
identical key sets, and all exports pass
`assert_exported_checkpoint_matches(..., check_values=True)`
— every tensor, including all 20 x 256 experts, dequantizes to within
tolerance of the BF16
reference, so a transposed or mis-ordered expert stack would fail. NVFP4
lands in the correct packed
layout (`gate_up_proj [256, 1024, 1024]` U8, `weight_scale [256, 1024,
128]` E4M3,
`weight_scale_2 []` F32). Its *accuracy* is not meaningful — truncating
to 20 of 40 layers leaves a
chance-level model (BF16 0.2322, FP8 0.2538) — so it validates
correctness, not quality.

**Re-validated on the final code.** The numbers above were first taken
mid-review; since then the
NVFP4 block-scale merge changed on both packed paths, the vision-tower
download became two-stage,
and an expert-layout load guard was added. Both gating runs were
therefore repeated end to end:
Nemotron went 0.7748 → **0.7825 ± 0.0105** and Qwen3.5-0.8B went 0.4678
→ **0.4832 ± 0.0127**, with
the rest of the Nemotron pipeline reproducing exactly (3519 quantizers,
69GB checkpoint, 21GB
export). Both deltas are inside their own stderr, so the claim is that
the rework costs no accuracy
— not that it improved it. The Nemotron export also runs at `--pp_size
2`, exercising the new
collective layout guard on a real 30B MoE across pipeline stages.

Two limitations worth stating plainly:

- **No quantized accuracy number for a full-size MoE.** The full 35B
OOMs at 47.37 GiB while
*constructing* the model on 2x48GB, with grouped GEMM already enabled,
so no calibration knob
  helps. Needs more GPUs than this setup has.
- **vLLM cannot yet serve packed FP8 Qwen3.5 experts.** `vllm
0.24.1.dev0` builds its fused expert
mapping weight-only, rewriting `experts.down_proj_input_scale` to
`w2_weight_input_scale` while the
parameter it registers is `w2_input_scale`. This is upstream and
independent of how the checkpoint
is produced — both of our export paths fail it identically. The 0.8B
numbers above are unaffected
(dense), and the packed exports are verified against the reference
checkpoint instead.

#### Guard verification

Each new guard was made to fire, not just to compile:

| Guard | Verification |
|---|---|
| Export self-check | Disabled the MoE guard, re-exported Qwen3-MoE -
independently reported all 24 dropped tensors. No false positives across
llama, nemotron, qwen3, qwen3-moe, qwen3vl, qwen3.5-vl, deepseek_v3
incl. eagle / medusa / MTP |
| Dropped-state raise | Deleted `modelopt_state` from a checkpoint with
50 quantizer tensors - raised instead of loading unquantized |
| NVFP4 value check | Flipped a `q_proj` - failed at `max_rel_err=1.74`
against a 0.3 threshold |
| Zero-centered gamma | Reproduced the off-by-1.0 on a good export -
caught as "not bit-exact" |
| Exclusion guard | Asserts no calibrated quantizer matches `conv1d` /
`mlp.router` / `output_layer` |

Exported artifacts are validated, not just their existence: 0 missing
keys vs reference, vision
tower bitwise-identical, dequantized weights within FP8 E4M3 error
(<=4.6%). The
`in_proj_a`/`in_proj_b` check is load-bearing - swapped alpha/beta would
still match on shape but
show ~100% error.

Also ran a tiny-Qwen3 **LLM** control through both steps to confirm the
exporter changes are a
no-op off the VLM path.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — the scripts now derive the
MoE expert layout from the model config, building SequentialMLP only for
architectures with no `experts.linear_fc1` rule, and the exporter raises
rather than dropping experts it has no rule for. Those runs previously
"succeeded" while writing a checkpoint containing no expert weights, so
no working behaviour is removed. `--no_moe_grouped_gemm` forces
SequentialMLP explicitly.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅ — approved (round 10: 0
CRITICAL, 0 IMPORTANT, 0 new suggestions); CodeRabbit approved earlier

### Additional Information

**MoE expert layout is now chosen automatically.** Only Nemotron-H can
export fused grouped-GEMM experts, so every other MoE architecture would
otherwise need `--no_moe_grouped_gemm` on all four scripts or hit a wall
at export. The scripts derive the layout from the model config — grouped
GEMM unless it would not be exportable — so they agree without threading
a flag. This changes MoE activation scales from one shared scale to
per-expert for the affected architectures.

Known gaps, unchanged by this PR:

- **Gated MoE still cannot use fused grouped GEMM.**
`_grouped_mlp_slicing` emits one weight per expert with no gate/up split
— its only prior caller, Nemotron-H, is non-gated, so every other MoE
architecture is built as `SequentialMLP` (see below). Adding that split
would restore the faster layout, but it needs a deliberate call on
activation-scale semantics: grouped GEMM keeps **one shared** activation
scale across experts while `SequentialMLP` has **per-expert** scales, so
the two are not numerically equivalent. It also needs EP>1 coverage.
- **Qwen3.5's alpha/beta gates share Megatron's fused `in_proj`
quantizer,** so they can only be kept in BF16 at export, not excluded by
name. Full fidelity needs per-section quantizers on the fused
projection.
- **Anchoring ModelOpt state on `.language_model`** (which would let
`quantize.py` quantize the language model directly and drop its
name-based non-LM disabling) needs a coordinated Megatron-Bridge change:
`save_sharded_modelopt_state` is ModelOpt code, but the restore the
Bridge path uses is Bridge's own and unconditionally restores onto the
root.
- **Gemma3-VL** remains Megatron-checkpoint only (`OMNIML-5366`).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
  * Added Muse Glimmer AutoQuantize and Alpamayo QAD workflows.
* Added streaming Kimi-K3 conversion and NVFP4 activation headroom
calibration.
  * Added SFT-masked distillation for Megatron-Bridge.
* Added unified Hugging Face export for quantized Qwen3-VL and
Qwen3.5-VL checkpoints.
* MoE expert layouts are selected automatically, with an option to force
sequential experts.

* **Bug Fixes**
* Improved export validation for tensor coverage, MoE mappings,
quantizer state, and NVFP4 scales.
  * Fixed Qwen3.5-VL GatedDeltaNet export handling.
  * Preserved visual-model weights exactly during export.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-02 13:29:59 +00:00
21b95adabb Add FP8 Vision Encoder quantization for Qwen3-VL and Qwen3.5 (#2083)
### What does this PR do?

Type of change: New feature

Adds opt-in FP8 Vision Encoder quantization recipes for Qwen3-VL and
dense Qwen3.5:

- `fp8_vision-kv_none`: FP8 Vision Encoder Linears, with the LLM and KV
cache kept in high precision.
- `fp8_vision_lm-kv_fp8_cast`: FP8 Vision Encoder and LLM Linears, with
FP8 KV-cache cast.

Patch embedding and vision-attention BMM operands remain in high
precision. With `--calib_with_images`, calibration batches now pass
through the complete VLM so multimodal inputs exercise the selected
quantizers. This fixes image-text calibration for non-Nemotron VLMs and
may change language-model activation ranges and output scales for
existing commands.

### Usage

```bash
# Vision Encoder only
python examples/hf_ptq/hf_ptq.py \
  --pyt_ckpt_path <Qwen3-VL-checkpoint> \
  --recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none \
  --calib_with_images \
  --calib_size 512 \
  --skip_generate \
  --export_path <output-checkpoint>

# Vision Encoder + LLM + KV cache
python examples/hf_ptq/hf_ptq.py \
  --pyt_ckpt_path <Qwen3-VL-checkpoint> \
  --recipe huggingface/qwen3_vl/ptq/fp8_vision_lm-kv_fp8_cast \
  --calib_with_images \
  --calib_size 512 \
  --skip_generate \
  --export_path <output-checkpoint>
```

For dense Qwen3.5, replace `qwen3_vl` with `qwen3_5` in the recipe path.

### Testing

- Validated recipe selection, image calibration, and GPU
calibration/export for Qwen3-VL and Qwen3.5, in both vision-only and
joint configurations.
- Consolidated test run after rebasing: 364 passed, 77 skipped.
- Ruff, recipe validation, and `git diff --check` passed.
- Transformers 4.57 compatibility verified for Qwen3-VL; Qwen3.5 tests
capability-skip when the required Transformers classes are unavailable.

Deployment evidence with Qwen3-VL-2B on RTX PRO 6000 BSE, eight fixed
frames and a BF16 LLM:

| Configuration | Accuracy mean | Vision Encoder kernel time |
Full-request GPU kernel time |
|---|---:|---:|---:|
| BF16 | 48.75 | 25.45 ms | 47.43 ms |
| Standard FP8 | 48.42 | 19.35 ms (**24.0% faster**) | 41.39 ms (**12.7%
faster**) |

The accuracy mean covers MMMU, RealWorldQA, Video-MMMU, MVBench, and
Video-MME. Serving reached 7.7% lower end-to-end latency and 7.9% higher
throughput at concurrency 16.

Qwen3-VL-2B accuracy was evaluated through vLLM on B300 with
`--enforce-eager`. Both checkpoints used the same judge-free tasks, Qwen
sampling preset, seed, and task parameters.

| Benchmark | BF16 | VE-only FP8 | Delta |
|---|---:|---:|---:|
| MMMU validation | 45.33 | 45.22 | -0.11 pt |
| RealWorldQA | 64.97 | 65.10 | +0.13 pt |
| Video-MMMU | 31.56 | 31.11 | -0.45 pt |
| MVBench | 51.40 | 50.10 | -1.30 pt |
| Video-MME | 50.48 | 50.59 | +0.11 pt |
| **Unweighted mean** | **48.75** | **48.42** | **-0.33 pt** |

Runtime support for quantized Vision Encoder Linears is separate from
this ModelOpt checkpoint-generation change.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — `--calib_with_images` now
performs the intended complete VLM forward and may change language-model
calibration scales. Recipe-based VLM PTQ also scopes recipe rules to the
complete model. Both changes are documented in the changelog.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — will run `/claude review`
after opening the PR.

### Additional Information


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added FP8 vision quantization recipes for Qwen3-VL and Qwen3.5.
* Added Muse Glimmer AutoQuantize, Alpamayo QAD, streaming Kimi-K3
conversion, layerwise checkpoint export, and NVFP4 calibration/export
workflows.
  * Added ONNX FP16 conversion support for excluding selected nodes.
* **Bug Fixes**
* Improved multimodal calibration, ONNX scale handling, and NVFP4 CPU
compatibility checks.
* **Documentation**
* Expanded guidance for vision quantization, calibration, precision, and
conversion workflows.
* **Breaking Changes**
* Removed deprecated PTQ and evaluation interfaces and raised the
minimum supported Megatron container version.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: mpariente <mpariente@nvidia.com>
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Shengliang Xu <106840466+shengliangxu@users.noreply.github.com>
2026-09-01 16:30:28 -07:00
Shengliang Xu de3eda8a11 Restructure recipes: split per-model_type recipes from model-hub checkpoint recipes (#2219)
### What does this PR do?

**Type of change:** Refactor (recipe-library layout) + documentation —
backward-breaking for saved `--recipe` paths.

Separate the two kinds of built-in Hugging Face recipes that were
previously mixed under `modelopt_recipes/huggingface/`:

- **`huggingface/<model_type>/`** — architecture recipes keyed by the
transformers `model_type`; one recipe covers every checkpoint of that
architecture. **Unchanged.**
- **`models/<org>/<model_id>/`** — a *new top-level tier* for recipes
that mirror one specific published checkpoint, keyed by its **model-hub
path** (as on the Hugging Face Hub, ModelScope, etc.) so the on-disk
path equals the hub path.

Concretely, the model-instance recipes move out of `huggingface/` to the
top level:

- `huggingface/models/mistralai/…`, `huggingface/models/nvidia/…` →
`models/mistralai/…`, `models/nvidia/…`
- `huggingface/step3p5/Step3.5-Flash/…` →
`models/stepfun-ai/Step-3.5-Flash/…` (re-keyed to the canonical HF repo
id
[`stepfun-ai/Step-3.5-Flash`](https://huggingface.co/stepfun-ai/Step-3.5-Flash)
— org `step3p5`→`stepfun-ai`, id `Step3.5-Flash`→`Step-3.5-Flash`)

**Why:** `modelopt_recipes/README.md` already documented a top-level
`models/` tier, but the files lived under `huggingface/models/` and
instance-specific recipes were awkwardly nested under the
per-`model_type` tree. This aligns the filesystem with the documented
layout and makes the instance tier hub-addressable — given a checkpoint
id you can find (or place) its recipe with no lookup table.
`load_recipe` resolves paths directly under `modelopt_recipes/`, so a
top-level `models/` sibling of `general/` and `huggingface/` works
identically.

The move is metadata-only — all recipe YAML content is byte-identical
(`R100` renames). Everything else is updating references (nvidia
launcher YAMLs, `test_loader.py`) and docs: a new `models/README.md`,
plus `huggingface/README.md`, root `README.md`, `ptq.md`, and the
`10_recipes.rst` guide, which no longer describe instances under
`huggingface/`.

### Usage

Recipe paths for the moved checkpoint recipes lose the `huggingface/`
prefix (and Step 3.5 Flash is keyed by its hub id):

```python
from modelopt.recipe import load_recipe

# before
load_recipe("huggingface/models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16")
load_recipe("huggingface/step3p5/Step3.5-Flash/ptq/nvfp4-mlp-only")

# after
load_recipe("models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16")
load_recipe("models/stepfun-ai/Step-3.5-Flash/ptq/nvfp4-mlp-only")
```

The same rename applies to `--recipe …` CLI values and launcher
`QUANT_CFG:` entries. Architecture recipes under
`huggingface/<model_type>/` are unaffected.

### Testing

- **Recipe resolution (torch-free):** parsed every recipe under
`models/` and confirmed all `$import` targets resolve against the recipe
root — 0 dangling across the tier.
- **Docs consistency:** re-ran the
`tests/unit/recipe/test_recipe_docs.py` logic; it now globs both
`huggingface/` and `models/`, and every model dir (incl.
`Step-3.5-Flash`, `Nemotron-3-Nano-4B-BF16`, …) plus every `general/ptq`
recipe is still mentioned in `ptq.md`.
- **Reference sweep:** repo-wide grep confirms no remaining references
to the old paths outside the intentional historical CHANGELOG entries
(released 0.44 / 0.45).
- **pre-commit:** `markdownlint-cli2`, license-insert, and `bandit`
hooks pass on the changed files.
- Note: the full `pytest` suite was not run in my environment (no
`torch`), so `test_recipe_docs.py` / `test_loader.py` should be
exercised in CI.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — `--recipe` / `load_recipe`
paths for the checkpoint-mirror tier change (drop the `huggingface/`
prefix; `step3p5/Step3.5-Flash` → `stepfun-ai/Step-3.5-Flash`).
Documented as a Backward Breaking Change in `CHANGELOG.rst` (0.47); the
only *released* old paths affected shipped in 0.45. A clean break was
chosen over a symlink or loader-alias shim.
- If you copied code from any other sources or added a new PIP
dependency …: N/A
- Did you write any new necessary tests?: ✅ — updated
`test_recipe_docs.py` to also glob the top-level `models/` tier so
instance recipes stay covered by the doc-consistency check.
- Did you update Changelog?: ✅ — added a 0.47 **Backward Breaking
Changes** entry.
- Did you get Claude approval on this PR?: ❌ <!-- run /claude review -->

### Additional Information

Design note: an earlier iteration nested everything under
`huggingface/model_type/` + `huggingface/models/`; the final layout
keeps `huggingface/` flat (per-`model_type`) and lifts instances to a
top-level `models/` tier, matching what `modelopt_recipes/README.md`
already documented. The `Step3p5*` architecture class names (from the
model's `trust_remote_code` modeling code) are unrelated to the recipe
path and are left unchanged.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added checkpoint-specific PTQ recipes for Kimi-K3, Mistral Medium 3.5,
and NVIDIA Nemotron models.
  * Added a Nemotron speculative-decoding warm-start recipe.

* **Documentation**
  * Clarified recipe selection and directory organization.
  * Documented checkpoint naming conventions and updated usage examples.

* **Bug Fixes**
* Updated launcher configurations and examples to reference the new
recipe locations and corrected model names.

* **Tests**
* Improved automatic recipe discovery and validation of documented
recipe paths.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-01 10:22:27 -07:00
Keval MorabiaandClaude Opus 5 8810eb5e31 Add ModelOpt recipe for DeepSeek-V4-Pro-0813 NVFP4 and --recipe to its PTQ script (#2287)
### What does this PR do?

Type of change: new feature

The quantization config for `nvidia/DeepSeek-V4-Pro-0813-NVFP4` existed
only as Python inside `_build_nvfp4_experts_cfg()`, so the released
checkpoint had **no entry in `modelopt_recipes/`** and could not be
looked up by name the way every other published model can.
`modelopt_recipes/README.md` states the goal directly — a recipe is
*"the single, version-controlled source of truth for how a model is
optimized … expressed as data instead of code"* — and this model was the
exception.

This adds:

-
`modelopt_recipes/huggingface/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only.yaml`,
composed from the existing `configs/ptq/units/base_disable_all` and
`configs/numerics/nvfp4` units.
- An optional `--recipe` flag on `examples/deepseek/deepseek_v4/ptq.py`.

It follows **`examples/kimi/kimi_k3`**, the closest precedent: a very
large MoE whose source already ships MXFP4 routed experts, converted via
`--cast_mxfp4_to_nvfp4` rather than through `examples/hf_ptq`, and
already wired to `--recipe` with a published YAML.

### Usage

```sh
torchrun --nproc-per-node 8 deepseek_v4/ptq.py \
    --model_path  <mp8_checkpoint> \
    --config      <DeepSeek-V4-Pro-0813>/inference/config.json \
    --calib_size  512 \
    --calib_seq   4096 \
    --output_path <amax_dump> \
    --recipe huggingface/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only
```

Omitting `--recipe` keeps the previous behaviour exactly.

### Testing

- `load_recipe` resolves the YAML and yields `num_bits (2, 1)` with
`block_sizes {-1: 16, type: dynamic, scale_bits: (4, 3)}` — identical to
the hardcoded config.
- **Equivalence checked behaviourally**, not by eyeballing dicts: both
configs were resolved against representative quantizer names using
last-match-wins, and agree on all of them.

  | quantizer | hardcoded | recipe |
  | --- | --- | --- |
| `...ffn.experts.17.w1_weight_quantizer` | enabled, NVFP4 | enabled,
NVFP4 |
| `...ffn.experts.17.w2_input_quantizer` | enabled, NVFP4 | enabled,
NVFP4 |
  | `...ffn.shared_experts.w1_weight_quantizer` | disabled | disabled |
  | `...attn.wq_weight_quantizer` | disabled | disabled |
  | `mtp.0.ffn.experts.2.w1_weight_quantizer` | disabled | disabled |
  | `lm_head_weight_quantizer` | disabled | disabled |

- `mtq.quantize` documents `algorithm` as a string **or** a dict keyed
on `method`, so the recipe's `{'method': 'max'}` needs no translation.
- `pre-commit` clean, including `validate modelopt recipes`.

No GPU run: this changes config plumbing only, and the default path is
byte-identical to before.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `--recipe` is optional and
defaults to `None`; without it `_build_nvfp4_experts_cfg()` is used
exactly as before.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies; `modelopt.recipe` is already a first-party import.
- Did you write any new necessary tests?: N/A — no new logic;
equivalence to the existing config is the property that matters and is
documented above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — recipe library addition; recent recipe/example PRs add no entry.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

The recipe covers the **quant config only**. `--calib_seq` — the setting
that mattered most for this checkpoint, since the 512 default does not
cover long-context activation ranges — is a dataloader argument rather
than part of the `mtq` config, so it stays on the CLI. Worth knowing if
the recipe is ever treated as a complete reproduction of the released
checkpoint: it is not, on its own.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added post-training quantization support for DeepSeek-V4-Pro-0813
routed experts using NVFP4.
* Added an optional recipe path for selecting equivalent quantization
settings.
* Preserved source formats for shared experts, attention, embeddings,
output layers, and MTP components.

* **Bug Fixes**
* Improved validation for missing or malformed quantization
configurations.
* Added safeguards against unsupported formats, scopes, algorithms, and
enabled MTP quantizers.

* **Documentation**
* Documented checkpoint conversion behavior, calibration requirements,
and supported quantization workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 21:50:15 +00:00
Frida HouandClaude Opus 5 029c67f27e feat(export): export each decoder layer as layerwise calibration finishes it (#2136)
### What does this PR do?

Type of change: new feature

Layerwise calibration can already resume, but only through a
full-precision scratch checkpoint, and a completed run still pays for a
second whole-model export pass over it.

`layerwise.export_dir` writes each decoder layer to its own quantized
shard as soon as calibration finishes with it, so the directory is a
complete, loadable checkpoint when the last layer lands and
`export_hf_checkpoint()` is skipped. The shards *are* the resume
artifact: a restarted run reuses layers already on disk instead of
recalibrating and re-exporting them, so no full-precision copy of the
model accumulates. The resume directory beside it holds only the current
boundary's cached activations and the per-layer output shapes.

Setting the config field is the whole switch — no CLI flag. `hf_ptq.py`
rewrites its value to `--export_path`, and derives the resume directory
(`<export_path>.layerwise_resume`) when you haven't chosen one.

One shard per layer is what makes resume safe: shards are written whole
and named from the layer index, so a crash can lose the layer in flight
but never corrupt an earlier one, and a re-run overwrites in place.

Because a resumed run never recalibrates the layers it skipped, **the
in-memory model is not valid for inference afterwards**; the field
implies `--skip_generate`.

Includes a pre-existing `main` fix this depends on: `_is_layerwise` used
`getattr` on an algorithm that YAML parses as a **dict**, so it answered
`False` for every layerwise recipe in the repo and the batch-size probe
it gates was never skipped. **Behaviour change:** `--batch_size 0` now
yields `batch_size=1` for layerwise recipes, as its comment intends.
Detection also now scans every algorithm entry rather than the first, so
a list-form recipe whose `layerwise` block is not first is recognised as
layerwise — same batch-size consequence. Nothing else on the non-fused
paths changes: `FUSION_FREE_FORMATS` is the exact set the inline list
held, `save_non_weight_artifacts` is a lift of the streaming exporter's
own block, and the calibration-loop changes are gated on an exporter
being present.

**Refused before calibration starts**, since each would otherwise
produce a silently different checkpoint rather than fail:

| Refused | Why |
|---|---|
| AWQ / SVDQuant | need pre-quant-scale steps that are still whole-model
|
| Weight-tied quantized modules | `sync_tied_input_amax` merges amaxes
across a partner that may be uncalibrated or already written |
| Multi-process (FSDP2) | every rank would write the same shards |
| Multimodal (VLM) | calibration runs on the extracted language model |
| MTP models | exclusions applied after calibration has written
everything |
| AutoQuantize recipes | only the mono-quantize path retargets
`export_dir` |
| Spec-dec, `--vllm_fakequant_export`, non-dense sparsity,
`int8_smoothquant`, encoder-decoder `model_type` | each routes to a
second exporter that would overwrite `--export_path` |
| `export_dir` on more than one algorithm entry, or on any but the last
| export finalizes shards as calibration walks the layers, so a later
pass would change the model after its checkpoint was written |

Shards are also bound to the run that produced them
(`.layerwise_export.json`: model class, layer count, formats, KV-cache
format, and a digest of the resolved quant config), so one run's
manifest cannot finalize another's shards. Source weights are not
digested — that would mean reading the whole model — so
differently-trained weights at the same path compare equal.

### Why a separate exporter

Three reuse paths were considered before adding one:

- **Extend `_StreamingShardWriter`.** It buffers by `max_shard_size`
into `__shard_part_*` temp names and renames to canonical names only in
`finalize()`, once the shard count is known. The resume invariant needs
the opposite: a stable `model-layer-00007.safetensors` committed when
layer 7 finishes, so "shard exists" means "layer done" across a restart.
Forcing a per-layer flush still leaves temp names, finalize-time
renaming, and an in-memory `_key_to_part` — every method would change.
- **Keep the layerwise checkpoint and run the streaming exporter at the
end.** This works, and it is why the pitch above is *not* durability:
that already exists. What it leaves is a second whole-model pass owed
*after* calibration finishes — itself needing a GPU session — where
per-layer export makes the last calibrated layer also the last exported
one. Scratch size only separates them for weight-mutating calibrators:
`save_layer_state` is off under per-layer export, but with
`calib_mutates_weights: false` (the shipped recipe) the checkpoint holds
just amax buffers either way.
- **Factor a shared per-module writer around `ExportContext`.** The
right long-term shape, but it touches all three existing export paths;
doing it here makes this change larger, not smaller.

One deliberate divergence from `_StreamingShardWriter` worth knowing
about: it clones tensors that share storage, this path lets `save_file`
raise instead. No model was found where the clone fires, and copying
unattributed aliases can hide a real bug rather than surface it. If a
checkpoint ever trips it, that is information we want.

Happy to take a different call on this — flagging it for maintainer
sign-off rather than assuming it.

### Usage

```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --export_path <out> \
    --recipe modelopt_recipes/general/ptq/nvfp4_experts_only-kv_fp8_layerwise_export.yaml
```

Interrupt and rerun the same command: calibration resumes from the last
committed layer, finished shards are reused, and a run that had already
finished every layer only re-runs `finalize()`.

```yaml
quantize:
  algorithm:
    method: max
    layerwise:
      enable: true
      calib_mutates_weights: false
      export_dir: /tmp/modelopt_layerwise_export   # presence is the switch; value replaced with --export_path
      # checkpoint_dir omitted -> derived as <export_path>.layerwise_resume
```

### Testing

Each row exports the same calibration two ways — per-layer, and
whole-model via `export_hf_checkpoint()` — and compares them **tensor
for tensor and config for config**.

The 35B row was re-run on the current head, against a baseline built
from a `main` worktree rather than from this branch, so it covers both
"per-layer differs from whole-model" and "this branch broke the shared
whole-model path". The other rows date from earlier heads; the code they
exercise is unchanged, but they are not fresh runs.

| Model | Config | Result |
|---|---|---|
| Qwen3.6-35B-A3B (40 layers, 256 fused experts) | NVFP4 W4A4
experts-only + FP8 KV | 123,513 tensors, 0 mismatched; `config.json`,
`hf_quant_config.json`, `generation_config.json` all identical |
| Qwen3-30B-A3B (48 layers, 128 per-expert linears) | NVFP4 experts
(`nvfp4_static` weights) + `mse`, offload | 74,163 tensors, 0 mismatched
|
| Qwen3-30B-A3B | same, `SIGKILL` after 25/48 layers, then resumed |
74,163 tensors, 0 mismatched **vs the uninterrupted run** |
| Llama-3.1-8B-Instruct | FP8 dense + FP8 KV, resident | 803 tensors, 0
mismatched |

**Refusals verified on real checkpoints**, each writing **zero shards**
and never reaching calibration — the "refused before calibration starts"
claim above, demonstrated rather than asserted: multimodal and MTP
(Qwen3.6-35B, the MTP case on a text-only view since the multimodal gate
fires first), tied embeddings (Qwen3-0.6B), and multi-process (2-rank
`torchrun --use_fsdp2`, Llama-3.1-8B).

**Served, not just compared.** Under vLLM 0.27.1 (Marlin NVFP4 kernels,
SM 8.9): the 30B checkpoint exported three ways — whole-model,
per-layer, per-layer-resumed-after-a-kill — and the 8B exported both
ways all load and produce **identical greedy generations, 4/4 prompts**
within each model.

**Index integrity** on every checkpoint above: each `weight_map` key
resolves to the shard actually holding it; 0 missing, 0 extra, 0
mis-routed. Tensor equality alone never exercises that, and it is the
one artifact per-layer export builds differently.

Resume state stays bounded: **332 KB beside 22 GB** of shards on the
35B, **396 KB beside 19 GB** on the 30B — the committed boundary's
activations only, not one set per layer.

Not covered: the `trust_remote_code` `*.py` copy path.
Nemotron-Nano-12B-v2-Base fails with a CUDA illegal memory access on
these cards, on the whole-model baseline too, so it is an environment
limit rather than a result.

**Comparing the configs is new, and it caught a real bug.**
`get_quant_config` reports on the quantizer modules, which
`export_layer` replaces as it goes, so reading it in `finalize()`
described a model with no quantizers left: the checkpoint advertised
`quant_algo: null` while its weights were packed NVFP4, and under the
shipped experts-only recipe `hf_quant_config.json` was not written at
all. It is snapshotted in `__init__` now, beside the kv-cache format
already captured there — which is why that one field was correct while
the rest were not. Uniform FP8 and NVFP4 hid it because their configs
survive the conversion; only a mixed model loses its algo, and mixed is
what every shipped layerwise-export recipe is. Reverting the fix fails
`test_moe_export_matches` and passes the ten uniform-format cases,
matching what the 35B shows.

**24 GPU tests** in `tests/gpu/torch/export/test_layerwise_export.py`.
The equivalence oracle is a cross-product: {FP8, NVFP4, NVFP4 +
`get_qdq_activations_from_prev_layer`, mixed FP8/NVFP4, KV-cache} ×
{fresh, resumed-after-interruption}, each compared tensor-for-tensor
against `export_hf_checkpoint`. Plus MoE export; resume fail-fast;
resume artifacts replaced and pruned; complete-manifest finalize-only;
shards-without-manifest refusal; shards-from-a-different-run refusal
(format and module selection); identity-without-shards does not block a
rerun; export-does-not-mutate-the-model; index routes every key to the
shard holding it; AWQ refusal (from config, and after calibration);
export-without-`checkpoint_dir`.

**Unit tests** in `tests/examples/hf_ptq/test_example_utils.py` cover
the list-valued `algorithm` shapes: which entry owns export, whose
`checkpoint_dir` is derived, per-entry resume bases, both ambiguity
refusals, and the recipe shapes `recipe_layerwise_blocks` normalizes
(dict, list order, config object, and the empty cases).

`tests/gpu/torch/export/` 150 passed / 2 skipped (pre-existing env
skips) · `tests/unit/recipe` 284 · `tests/unit/torch/export` 186 ·
`test_layerwise_calibrate` 33 · `test_example_utils` 42 · pre-commit
clean.

Also verified: the exported directory reloads through
`AutoModelForCausalLM` and runs a forward.

Not a speed win: per-layer export was slower than the streaming export
in one offload pairing (271s vs 208s, the per-layer fusion probe),
though those runs shared GPUs so the magnitude is not cleanly measured.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `export_dir` defaults to
`None`; existing paths unchanged when unset, except the batch-size
change noted above.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — not yet; draft.

### Additional Information

**Pre-existing bug found on the way, not fixed here.** Layerwise
calibration leaves `self_attn.o_proj`'s input amax at `0.0` on every
layer but the last, so a full-NVFP4 layerwise model cannot be exported
by *any* path. `get_qdq_activations_from_prev_layer=True` avoids it,
pinning the cause to the pre-`calib_func` capture pass — which also
explains why only the last layer, the one that skips it, is correct.
That combination now works with per-layer export (it asserted on layer 0
until review caught it). Hidden until now because the shipped NVFP4
layerwise recipes are experts-only; the NVFP4 tests here exclude
`o_proj` for the same reason. Deserves its own issue.

---------

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-30 11:50:29 -07:00
Rohan Joshi 424f47b871 Add QAD example for Alpamayo (#2271)
### What does this PR do?

Add `examples/alpamayo/qad.py` which distills the quantized Alpamayo
checkpoint's VLM against the FP16 VLM of the original with ModelOpt's
QADTrainer, and shards student and teacher with FSDP2 for multi-GPU
runs. Only the VLM is trained; the action expert stays frozen.


Type of change: new example

<!-- Details about the change. -->

### Usage

```
torchrun --standalone --nproc_per_node 8 qad.py \
    --student_ckpt ./alpamayo-auto \
    --output_dir ./alpamayo-auto-qad \
    --parquet ./train_clips.parquet \
    --max_steps 500 --fsdp2 --grad_ckpt --export
```

### Testing
Tested end-to-end on public Alpamayo-1 checkpoint

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Very short summary of changes only for new features,
backward breaking changes, deprecations, or fixes for critical bugs
present in previous releases. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added a quantization-aware distillation workflow for AlpamayoR1
models.
- Supports prompt-only or rollout-inclusive distillation, FSDP2
training, dataset slicing, revision pinning, and checkpoint resumption.
- Supports exporting trained models as complete, reloadable AlpamayoR1
checkpoints.
- Added optional vision-parameter freezing, trajectory-history fusion,
gradient checkpointing, evaluation, and synchronized training cadence.

- **Documentation**
- Expanded the Alpamayo guide with setup, training, dataset, and export
instructions.
  - Clarified sensitivity-based quantization behavior.
  - Added a version 0.47 quantization changelog entry.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Rohan Joshi <rohjoshi@nvidia.com>
2026-08-28 17:19:51 +00:00
Keval MorabiaandClaude Opus 5 6a2ae5a25b Fix the llm_eval timeout: reachable MMLU mirror + no pipe deadlock (#2270)
### What does this PR do?

Type of change: Bug fix

`tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8` has been
failing with `Failed: Timeout (>900.0s) from pytest-timeout` on
unrelated branches (runs 33027056246 and the one for `ad83a428`, while
the 2026-08-24 nightly passed). It is not the test being slow — it is
the harness deadlocking, and the deadlock also destroys the diagnostics
that would explain the underlying kill.

**Mechanism.** The traceback shows `self = <Popen: returncode: -9 args:
['scripts/huggingface_example.sh', ...]>` while still blocked in
`stdout.read()`. The launcher was SIGKILLed (nothing in pytest sends
SIGKILL — pytest-timeout raises in the main thread, and the test's
`finally` `pkill` sends SIGTERM and only runs afterwards — so an OOM
kill is the likely source). But `subprocess.run(..., stdout=PIPE,
stderr=STDOUT)` waits for **EOF on the pipe**, not for the process, and
a surviving grandchild (the TRT-LLM serve/build worker) still holds the
write end. EOF never arrives, so the test blocks until the 900 s alarm.
Because the pipe is never drained, **every line of child output is
discarded**, which is why the CI log says nothing about what the script
was doing when it died.

Reduced to a self-contained reproducer:

```python
script = "sleep 300 & echo 'launcher output'; sleep 0.3; kill -9 $$"
subprocess.run(["bash", "-c", script], stdout=PIPE, stderr=STDOUT, text=True, timeout=20)
# -> TimeoutExpired: still blocked in communicate() after 20.0s, output lost
```

**Fix.** `_run_capturing` now starts the command in its own session,
drains its output on a reader thread (so logs stream as they arrive
instead of being buffered until the end), waits on the *process*, and
kills the process group if descendants still hold the pipe after a 30 s
grace period. A killed launcher now fails in seconds with its logs
intact instead of silently burning the test's whole timeout.

This does not fix whatever kills the script; it makes it diagnosable.
Worth noting separately: `test_qwen3_eval_fp8` took **749.10 s against
its 900 s mark** on the last green nightly, so it is fragile regardless
and may want its work trimmed or its budget raised once the logs show
where the time goes.

### Usage

```python
# unchanged public API
run_example_command(cmd_parts, example_path="llm_eval")
```

### Testing

Verified against the reproducer above and on the normal paths:

| scenario | before | after |
| --- | --- | --- |
| launcher SIGKILLed, survivor holds the pipe | blocks indefinitely (900
s in CI) | `rc=-9` in 3.5 s, `'launcher output'` captured |
| the surviving descendant | keeps running | killed with the process
group (stopped ticking, 20 -> 20 bytes) |
| normal exit | ok | `rc=0`, stdout and stderr interleaved in order |
| non-zero exit | ok | `rc=3`, output captured |

The example-test suites that use this helper run through the same code
path; `tests/examples/megatron_bridge` (16 passed, 1 skipped) exercised
it on nemo:26.08 in the branch this was extracted from.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `_run_capturing` keeps its
`(returncode, output)` contract; only the buffering strategy changed.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ — this is test
infrastructure; the scenario needs a process that outlives a SIGKILLed
parent, which is awkward to assert in CI. Verified manually with the
reproducer above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — test-infrastructure fix.
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved command execution reliability with real-time output capture.
* Ensured lingering child processes are cleaned up after commands exit
or are interrupted.
  * Added warnings when forced cleanup may truncate output.
  * Prevented hangs when descendant processes keep output streams open.

* **Documentation**
* Updated MMLU setup instructions to use the Hugging Face dataset
repository.
  * Improved Windows instructions by explicitly using `curl.exe`.

* **Examples**
* Improved MMLU downloads with retries, separate timeouts, resume
support, and automatic temporary-file cleanup.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---

### Update: the second half of the failure

With the streaming fix in place, the next CI run showed the *actual*
cause, which the old code had been hiding. That run blocked in
`process.wait()` with `<Popen: returncode: None ...>` — child alive and
working, not the previous dead-child pipe deadlock — and the now-visible
output was:

```
--2026-08-27 18:09:20--  (try: 4)  https://people.eecs.berkeley.edu/~hendrycks/data.tar
Connecting to people.eecs.berkeley.edu ...|128.32.139.28|:443... failed: Connection timed out.
Retrying.
--2026-08-27 18:11:40--  (try: 5)  ...
```

`huggingface_example.sh` downloads the MMLU tarball from
`people.eecs.berkeley.edu`, that host stopped answering around
2026-08-25, and wget's default retry policy (20 tries, ~2 min per
connect timeout) consumed the whole 900 s budget. Not runner-specific:
the URL also times out from a developer workstation, and the nightlies
flipped 08-24 ✅ / 08-25 ✅ / **08-26 ❌ / 08-27 ❌**, matching the outage.

So this PR now carries both halves of the same failure:

1. the harness no longer deadlocks and no longer swallows the logs
(`985809cc2d`), and
2. the MMLU data comes from HuggingFace's copy of the same tarball, with
bounded retries (`40f1d89154`).

The mirror is byte-for-byte the same dataset in the same layout the
script already expects — verified by running the exact download/extract
commands:

```
https://huggingface.co/datasets/cais/mmlu/resolve/main/data.tar  ->  HTTP 200, 166 MB
data/mmlu/{dev,test,val}/  ->  57 subject CSVs each, plus auxiliary_train/
```

`wget --timeout=20 --tries=3` plus an explicit error means the next
dataset-host outage fails in about a minute with "Could not download the
MMLU test data. Set MMLU_DATA_PATH to a local copy." instead of silently
eating a test's timeout. The same URL is updated in
`examples/llm_eval/README.md` so a manual run does not hit the dead host
either.

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 21:38:25 +05:30
Zhiyu 5500999d0b Add Kimi-K3 NVFP4 experts and FP8-PB attention recipe (#2206)
### What does this PR do?

Type of change: new example

Adds the calibration-free conversion pipeline and checkpoint-mirror PTQ
recipe used for `nvidia/Kimi-K3-NVFP4`:

- streams the 96-shard Kimi-K3 checkpoint without loading the 2.8T
model;
- casts the source MXFP4 routed experts to NVFP4 with expert
`input_scale=1.0`;
- quantizes the selected KDA and MLA attention weights to 128x128 block
FP8;
- leaves shared/latent experts, routers, convolutions, norms, the vision
tower, `lm_head`, and KV cache unquantized;
- emits mixed-precision Hugging Face metadata for deployment; and
- adds an exact recipe under
`modelopt_recipes/huggingface/models/moonshotai/Kimi-K3/` that directly
configures the streaming converter.

It also fixes `NVFP4QTensor.quantize()` probing CUDA/Blackwell
capability before checking whether the tensor is on CUDA and whether the
optional TensorRT-LLM fast path was requested. That probe broke the
converter's supported CPU path on hosts without a compatible GPU.

### Usage

```bash
python examples/kimi/kimi_k3/quantize_to_nvfp4.py \
    --source_ckpt /models/moonshotai/Kimi-K3 \
    --output_ckpt /models/Kimi-K3-NVFP4 \
    --recipe huggingface/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention \
    --jobs 8
```

The conversion requires no calibration dataset, forward pass, or GPU.
Multi-node shard conversion is also supported through `--rank`,
`--world_size`, and `--run_id`.

### Testing

```bash
uv run --frozen --extra dev python -m pytest -q \
    tests/unit/torch/quantization/test_nvfp4_tensor.py \
    tests/unit/recipe/test_kimi_k3_recipe.py \
    tests/unit/recipe/test_recipe_docs.py \
    tests/unit/torch/export/test_shard_cast_utils.py \
    tests/examples/kimi/test_kimi_k3_quantize_to_nvfp4.py
```

Result: 34 passed.

All pre-commit hooks pass for the changed files, including recipe
validation, Ruff, mypy, Bandit, YAML formatting, and markdownlint.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: N/A

### Additional Information

The resulting checkpoint and model card are available at
https://huggingface.co/nvidia/Kimi-K3-NVFP4.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added calibration-free Kimi-K3 MXFP4-to-NVFP4 conversion with optional
FP8 attention quantization and distributed processing.
- Added NVFP4 activation calibration, weight-only quantization recipes,
grouped-expert quantization, compiled quantization options, SFT-masked
distillation, and MLflow tracking.

- **Documentation**
- Updated quantization terminology, recipe catalogs, checkpoint
guidance, and Kimi-K3 conversion instructions.

- **Bug Fixes**
- Improved CPU NVFP4 behavior, tied-weight export handling, EAGLE-3
training compatibility, and checkpoint export reliability.
- Removed deprecated configuration options and legacy evaluation
examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
2026-08-28 02:33:14 +00:00
Keval MorabiaandClaude Opus 5 7ff81dd795 Add day-0 verbosity gate + harden release skills from a live run (#2254)
### What does this PR do?

Type of change: Bug fix, new feature, documentation

Day-0 release treats verbosity as a hard gate, but nothing in the skill
measured it — `gate_compare.py` only does accuracy. A run could complete
Step 5 and report a publish recommendation with a gate silently
unmeasured. This adds the missing gate and folds in fixes for problems
that cost GPU-h on a live release run.

**New: `gate_verbosity.py` + Step 5b.** Pure `evaluate_verbosity()` plus
an artifact harvester, matching `gate_compare.py`'s conventions and
`failure_class` vocabulary.

Two details it encodes, both of which produced wrong verdicts before
they were understood:

- **Read `response_stats.avg_completion_tokens`.** The
`reasoning.*_tokens` fields are always `0`, which reads as "the harness
never captured tokens" and pushes you to word counts. Only the
reasoning/content *split* is missing, not the total. Words disagree with
the gate: one task read **+6.10% FAIL** in words and **+1.32% PASS** in
tokens.
- **Two run-hygiene filters.** Pooling a mismatched reasoning-effort run
reported **+43.06% FAIL** on a task that is **+1.24% PASS** matched; one
truncated run (n=200 vs 294) reported **+9.58% FAIL** on a task that is
**+1.71% PASS** without it. Tasks with no common sample count are
reported `not_comparable` rather than as a delta.

**Skill hardening**, each from a specific failure:

- **Step 2b canary** — poll ceiling must exceed load time (a 50 min poll
against a 51 min load failed a checkpoint that serves fine), print an
explicit `RESULT:` on every path (a fall-through exits 0 and reads as
PASS), log to shared storage, and canary the **as-exported** artifact
rather than a copy modified to make it work.
- **Step 4 config parity** — assert the candidate config differs from
the baseline's in nothing but checkpoint path and served-model name. A
mismatched `parallelism` was worth ~2 pp, enough to invert the sign of a
delta, and cost four re-runs.
- **Statistical power** — re-running does not guarantee fresh samples:
with a warm NEL response cache two runs came back bit-identical to 16
digits.
- **Step 6 closeout** — verify the published path against the evaluated
one by inode, and prefix rejected sibling exports.
- **Size gate** — growth is blocking by default and waived only when the
validation summary's
`source_precision` shows an already-sub-8-bit source (which cannot
shrink further under a
4-bit recipe) and the growth is within what that explains.
`source_precision` is now a
recorded field in the ptq validation table, so the waiver is reachable
from the normal
pipeline, and `SIZE_NOT_REDUCED` has a triage row pointing at declaring
it.
- **`ptq.py`** — `--calib_seq` matters more than `--calib_size`, and
`--mse_calibrate` is a no-op under `--cast_mxfp4_to_nvfp4` (it tunes
weight quantizers only, and the cast overwrites `weight_scale`).
- **`.gitignore workspaces/`** — the skills create scratch directories
inside the repo; nothing excluded them.

### Usage

```bash
python "$SKILL_DIR/scripts/gate_verbosity.py" \
    --baseline <baseline_eval_root> --candidate <candidate_eval_root> \
    --glob 'eval_*' --threshold 0.05
```

Exit codes match the sibling gates: `0` pass, `1` the gate ran and
failed, `2` the gate could not
read its input (wrong root, `--glob` matched nothing, everything
excluded). Prints per-task tokens,
delta, `within_threshold`, `sample_count`, run counts,
`dropped_mismatched_runs`, any
`truncated_comparison`, `not_comparable`, `harvest_diagnostics`, and a
`max_abs_delta` summary.

### Testing

- 8 new unit tests in `test_gates.py`, one per real failure mode
(two-sided threshold, partial-run filtering, unequal sample counts,
short-output warning, one-sided tasks, empty input). Full suite: **36
passed**, no GPU or network.
- `gate_verbosity.py` validated end-to-end against a real day-0 run's
artifacts: reproduces the hand-computed result (`max_abs_delta =
0.0171`, pass) and correctly marks the two unequal-sample tasks
`not_comparable`.
- `pre-commit run --files <changed>` clean, including ruff, mypy,
bandit, markdownlint, and the `.claude/skills` symlink sync.
- Verified `max_sample_length` is a real `get_dataset_dataloader`
parameter with default 512, matching `--calib_seq`'s default, so
existing `ptq()` callers are unaffected.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `--calib_seq` defaults to the
pre-existing 512; `gate_ptq.py` reclassifies size growth from
`QUANT_COVERAGE_FAILURE` to `SIZE_NOT_REDUCED`, which is a more precise
class for an already-4-bit source and is covered by a new test asserting
a real coverage failure still outranks it.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependencies, stdlib only.
- Did you write any new necessary tests?: ✅ — 8 new tests for the new
gate.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

The measured figures come from a completed day-0 NVFP4 release
qualification. Model-specific results were removed from the general
skills where the rule stands on its own; two references were kept
deliberately — a model card citation illustrating per-scenario sampling,
and a model-specific vLLM MoE kernel crash where the model name *is* the
evidence.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added configurable calibration sequence-length limits for DeepSeek-V4
quantization.
* Added automated release checks for output verbosity, evaluation
comparability, serving readiness, and configuration parity.

* **Bug Fixes**
* Improved quantization size-ratio reporting by distinguishing
explainable growth from blocking failures.
* Clarified handling of deployment memory-access errors and infeasible
evaluations.

* **Documentation**
* Expanded guidance for calibration, remote execution, workspace
management, evaluation setup, deployment troubleshooting, and
statistical reliability.
* Added task-specific guidance for SciCode and GDPVal feasibility
checks.

* **Chores**
  * Excluded workspace session directories from version control.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 12:52:09 -07:00
Keval MorabiaandClaude Opus 5 d73278808b Bump nemo container requirement to 26.08 for MBridge examples (#2257)
### What does this PR do?

Type of change: Bug fix

Bumps the Megatron-Bridge examples, tests and launcher configs to
`nemo:26.08` and removes the version-gated fallbacks they carried, plus
the fixes needed to make the suites green on that container.

**26.08 bump and shim removal**

- Examples, CI workflows, `noxfile.py` and the `mbridge_*` launcher
configs move to `nemo:26.08`.
- `examples/megatron_bridge/_distillation_provider.py` is deleted —
26.08's Megatron-Bridge ships `convert_to_distillation_provider(...,
distill_submodule=...)` natively, so `distill.py` imports it directly.
- `prune_minitron.py` drops the `AutoBridge.from_hf_config` /
config-only-export probing; `--no_moe_grouped_gemm` is no longer needed
in the MoE pruning tests, and the Qwen3.5-MoE `skipif` is gone (native
MoE expert mappings are in 26.08).
- `_DynamicMambaMixer` targets only the raw `conv1d_weight` /
`conv1d_bias` parameters that replaced the `conv1d` module in
Megatron-Core.

**MambaModel / MambaModelProvider removal**

Megatron-Core has shipped `HybridModel` since 26.06 and `MambaModel` is
a deprecated subclass that shares its `forward`, so `DMRegistry`
resolves those instances to the `HybridModel` registration and the
separate entry is redundant. Same for `MambaModelProvider` vs
`HybridModelProvider` on the bridge side. `MambaMixer` / `MambaLayer` /
`ExtendedRMSNorm` are untouched — the layers still exist. The deprecated
`get_te_mamba_stack_spec` is removed; use `get_te_hybrid_stack_spec`.

**Bug fix: compressed output_layer extra state**

`mtq.compress` converts even a *disabled* `output_layer` into a
`RealQuantLinear` (its weight is left uncompressed, since
`pack_real_quantize_weight` skips disabled quantizers). The guard added
in #2112 excluded `RealQuantLinear` by class, so the layer emitted extra
state and every worker died in `GPTModel.sharded_state_dict`:

```
RuntimeError: Boolean value of Tensor with more than one value is ambiguous
  megatron/core/models/gpt/gpt_model.py:896, in sharded_state_dict
    output_extra_state and output_extra_state.data
```

The guard now keys off whether the weight was actually compressed
(`QTensorWrapper`) instead of the class. This took out all 12
`test_homogeneous_compressed_sharded_state_dict` params, and the crashed
workers poisoned the pool, which surfaced as unrelated timeouts and NCCL
errors in `test_layer_sync_moe_local_experts_amax`,
`test_kv_cache_quant`, `test_kv_cache_amax_sync`,
`test_convert_mcore_te_gpt_model` and
`test_homogeneous_sharded_state_dict_te_spec` — 21 tests in total. The
e2e coverage is `skip_flaky_on_blackwell`, so CI never ran it;
`test_output_layer_extra_state_empty_when_nothing_quantized` now asserts
the contract directly and is not skipped.

**Checkpoint import entry point**

26.08 replaced `examples/conversion/convert_checkpoints.py` with
`scripts/conversion/convert.sh`, so
`tools/launcher/common/megatron_bridge/import/import.sh` and the three
README snippets are retargeted. `import.sh` uses the distributed GPU
backend with `GPUS_PER_NODE` / `TP` / `PP` / `EP` knobs.

**Megatron-LM on nemo:26.06** keeps working: `_get_mamba_conv1d` still
dispatches between the `conv1d` module (26.06 and earlier) and the raw
parameters (26.08+), so `import_mcore_gpt_from_hf` /
`export_mcore_gpt_to_hf` handle NemotronH on both. Only the
Megatron-Bridge examples and Minitron pruning of Mamba/hybrid models
require 26.08.

**Test consolidation**

`test_export_distilled_megatron_to_hf.py` is merged into
`test_distill.py`: `test_distill_llm` becomes
`test_distill_llm_hf_export` and covers the standalone
`--export_iterations all` run on the checkpoints it already produces,
saving one full distillation (~185 s of CI time). The two mamba-named
gpu test files are renamed to `hybrid`.

### Usage

```bash
# HF -> Megatron import, via Megatron-Bridge's 26.08 conversion entry point
bash /opt/Megatron-Bridge/scripts/conversion/convert.sh import \
    --executor local \
    --device gpu \
    --gpus-per-node 8 \
    --hf-model Qwen/Qwen3-8B \
    --megatron-path /tmp/Qwen3-8B-megatron
```

### Testing

All on `nvcr.io/nvidia/nemo:26.08`, 2x RTX 6000 Ada, no timeout
overrides:

- `tests/examples/megatron_bridge`: 16 passed, 1 skipped (28m14s). The
skip is the `gemma3vl` QAD param, now `@pytest.mark.manual` since
`qwen3_5_moe_vl` covers the VLM QAD path.
- `tests/gpu_megatron` (`_extensions`, `distill`, `export`, `opt`,
`peft`, `sparsity`, `speculative`, `utils`): 61 passed, 5 xpassed.
- `tests/gpu_megatron/torch/export` re-run after the conv1d dispatch
change: 27 passed.
- The 21 previously failing/hanging quantization tests: 21 passed (12 +
9).
- `tests/gpu_megatron/torch/{nas,prune}`: verified separately.

`import.sh` equivalence on a toy `qwen3_moe`, comparing all 12 weight
tensors after flattening each dist checkpoint with `dcp_to_torch_save` —
the GPU backend at 1 GPU, `--tp 2`, `--pp 2`, `--ep 2`, and `import.sh`
end-to-end (`GPUS_PER_NODE=2 EP=2`) are all byte-identical to `--device
cpu`.

`nemo:26.06` compatibility was checked directly in that image:
`megatron.core.models.hybrid.HybridModel`, the modelopt hybrid spec and
`hybrid_layer_pattern` are all present, while
`megatron.bridge.models.hybrid` and the bridge's `distill_submodule` are
not. The NemotronH round-trip test failed there before the conv1d
dispatch was restored and the dispatch is back in place; per project
convention the suites themselves only run on 26.08.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ⚠️ Megatron-Bridge examples plus
Minitron pruning of Mamba/hybrid models now require `nemo:26.08`.
Megatron-LM quantization and checkpoint export still run on
`nemo:26.06`.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ —
`test_output_layer_extra_state_empty_when_nothing_quantized` for the
compress fix; existing tests extended for the merged export coverage.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added guidance for importing Hugging Face checkpoints into Megatron
distributed format.
* Expanded distillation workflows to export selected or all checkpoint
iterations.

* **Improvements**
  * Expanded Hybrid model support across Megatron workflows.
  * Updated distributed import tooling with GPU and parallelism options.
  * Updated supported environments and examples to NVIDIA NeMo 26.08.

* **Bug Fixes**
* Corrected output-layer quantization state handling when quantization
is disabled.

* **Documentation**
  * Added compatibility guidance for current and legacy NeMo containers.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 00:31:41 +05:30
Keval MorabiaandClaude Opus 5 449a39922b Pin nemo_automodel below 0.6 for the fastgen example (#2260)
### What does this PR do?

Type of change: Bug fix

`nemo_automodel` 0.6.0 removed
`nemo_automodel.recipes.diffusion.train.is_main_process` without a
replacement (it was a three-line rank-zero predicate in 0.5.0, and 0.6.0
defines no equivalent anywhere in the package).
`examples/diffusers/fastgen/dmd2_recipe.py` imports it, so the example's
import guard fires and **every** test in `tests/examples/diffusers/`
errors at collection:

```
ImportError: cannot import name 'is_main_process' from 'nemo_automodel.recipes.diffusion.train'
  tests/examples/diffusers/fastgen/test_resume_dataloader.py
E   ImportError: The DMD2 fastgen example requires `nemo_automodel`. ...
collected 42 items / 1 error
```

The requirement was `>=0.4.0,<1.0`, so CI picked 0.6.0 as soon as it was
published and the `onnx (diffusers)` job started failing on every PR
(e.g. runs 33020467654, 33019418460, 33010815298, 33007613265,
33006944292 — all unrelated branches). Capping at `<0.6` restores the
tested range.

Every other `nemo_automodel` symbol the example imports still exists in
0.6.0 (`_diffusers.auto_diffusion_pipeline.NeMoAutoDiffusionPipeline`,
`recipes.diffusion.train.TrainDiffusionRecipe`, and the four
`components.datasets.diffusion.*` helpers), so `is_main_process` is the
only blocker; the alternative is defining that predicate locally and
widening the cap again, which is worth doing separately if the example
is meant to track 0.6.

### Usage

```bash
pip install -r examples/diffusers/fastgen/requirements.txt
```

### Testing

Reproduced the break by diffing the published wheels: `is_main_process`
is defined at `nemo_automodel/recipes/diffusion/train.py:692` in 0.5.0
and absent from 0.6.0 (`grep -rn "def is_main_process"` over the
unpacked 0.6.0 wheel returns nothing). Confirmed the remaining imported
symbols are all still present in 0.6.0.

CI on this PR exercises the fix directly: the `onnx (diffusers)` job
installs from this requirements file and is the job that has been
failing.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — existing
dependency, tightened bound.
- Did you write any new necessary tests?: N/A — the existing
`tests/examples/diffusers/` suite is what this unblocks.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — dependency-pin fix for a break introduced and fixed within the
same unreleased cycle.
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
  * Fixed dependency compatibility for the FastGen diffusion example.
* Prevented installation of versions that could cause the example to
fail at startup.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 05:59:02 +05:30
h-guo18 5db2682519 [Example]: Calibration-free FP8/NVFP4 PTQ for speculative-decoding drafters (#2027)
### What does this PR do?

Type of change: new example

Adds `examples/speculative_decoding/scripts/quantize_drafter.py`, a CLI
that quantizes an exported speculative-decoding drafter to FP8 or NVFP4
— weight-only or weight+activation — with no calibration data.

It needs no modeling code either. Exported drafters such as
[`nvidia/MiniMax-M3-DSpark`](https://huggingface.co/nvidia/MiniMax-M3-DSpark)
have no importable model class, so each 2-D weight is wrapped in a
throwaway `nn.Linear` under its checkpoint key and ModelOpt's usual
`quantizer_name` patterns select over those names. Works for any drafter
layout (DSpark / DFlash / EAGLE3 / Medusa).

**Formats:** `w4a16_nvfp4`, `nvfp4`, `fp8`, `fp8_pc_pt` — the ModelOpt
formats vLLM's backend can actually serve. AWQ is deliberately not
offered, since `awq_lite` silently degrades to plain RTN without a
`forward_loop`.

**Static activation scales without calibration.** `fp8` and `nvfp4`
normally need an activation amax *measured* on calibration data; a fixed
`input_scale` of 1.0 is applied instead. That works because acceptance
length is governed almost entirely by **clipping**, not resolution:

Sweeping the fixed scale over three decades (same setup as the Testing
section below; bf16 baseline 3.1423):

| `input_scale` | amax | FP8 AL | vs bf16 | NVFP4 AL | vs bf16 |
|---|---|---|---|---|---|
| 0.003 | 1.3 | 2.2204 | -29.34% | 2.2076 | -29.75% |
| 0.01 | 4.5 | 2.6719 | -14.97% | 2.6641 | -15.22% |
| 0.03 | 13.4 | 2.9751 | -5.32% | 2.9259 | -6.89% |
| 0.1 | 44.8 | 3.1013 | -1.31% | 3.0206 | -3.88% |
| 0.2 | 89.6 | 3.1178 | -0.78% | 3.0015 | -4.48% |
| 0.3 | 134.4 | 3.1370 | -0.17% | 3.0222 | -3.82% |
| 0.5 | 224.0 | 3.1268 | -0.50% | 3.0360 | -3.38% |
| **1.0 (default)** | **448.0** | **3.1457** | **+0.11%** | **3.0193** |
**-3.91%** |
| 2.0 | 896.0 | 3.1354 | -0.22% | 3.0172 | -3.98% |
| 4.0 | 1792.0 | 3.1245 | -0.57% | 3.0034 | -4.42% |

Both formats fall off a cliff below ~0.03, where the declared range sits
far under the activations' true magnitude and most of the tensor is
clipped. Both then sit on a flat plateau from ~0.3 to 4.0 **with no
drop-off at the top**, so the scale only has to be big enough. 1.0 is
the middle of that plateau, which is why it is hardcoded rather than
exposed. NVFP4 trails FP8 by a roughly constant 3.5% across the plateau
— that gap is the 4-bit resolution cost, and no choice of scale recovers
it.

Deriving the amax from the weights instead was tried and does not work:
`max|W|` averages 0.79 while a RMSNorm'd activation is O(1) with outlier
channels in the tens, so the range lands 1–2 orders of magnitude low and
clips, measuring -31% to -46% AL.

**Where calibration would go.** All of this sits behind
`resolve_activation_scales()`, the single place deciding where a static
amax comes from. Real calibration slots in ahead of the fixed fallback
with no change to the CLI or the call site, and composes because
`set_static_activation_amax()` skips quantizers that already have an
amax:

```python
if calib_forward_loop is not None:
    mtq.calibrate(root, quant_cfg["algorithm"], forward_loop=calib_forward_loop)
set_static_activation_amax(root)   # fills in what calibration did not reach
```

**Serving a quantized drafter.** Four things had to be written into the
exported checkpoint before vLLM would load one:

- emit `quant_method` (`modelopt_fp4` / `modelopt`) — vLLM reads that
key, ModelOpt writes only `quant_algo`
- emit the exclusion list under `ignore` too — that is the key read from
the flat `quantization_config`; `exclude_modules` alone yields an empty
exclusion set
- add `*<name>` wildcards so exclusions match a runtime's nested module
prefix (`model.fc`) rather than the checkpoint key (`fc`)
- add `*qkv_proj` / `*gate_up_proj` aliases for layers a runtime fuses,
whose names appear in no checkpoint key

Nothing is then needed on the caller side. **This closes the open
question left in the previous revision of this PR: vLLM does read
`quantization_config` off the draft checkpoint.**
`ModelConfig._verify_quantization` fills `quantization` in from
`quant_method` when it is unset, so once the export declares that key —
the first fix above — detection works on its own. Verified on
Nemotron-3.5-Lightning passing nothing: `Detected ModelOpt NVFP4
checkpoint (quant_algo=NVFP4)` → `FlashInferCuteDslNvFp4LinearKernel`,
AL 4.278 against 4.203 measured earlier.

`specdec_bench` also gains a `DSPARK` algorithm, which it did not have:
an exported `Qwen3DSparkModel` would otherwise have to go through
`DFLASH` and be built with vLLM `method="dflash"`. The branch sets
`method="dspark"` and leaves `draft_sample_method` on vLLM's own default
of `greedy`. A target whose fused-collective workspace (sized at
CUDA-graph capture) overflows at large speculative batches can disable
graphs with `--runtime_params '{"engine_args": {"enforce_eager":
true}}'`.

For DFlash-family drafters, `qwen3_dflash.py` builds its fused
context-KV projection by reading `qkv_proj.weight` raw and calling
`F.linear`, which cannot consume a packed weight. Keep those layers in
bf16 with `--exclude '*q_proj*' '*k_proj*' '*v_proj*' '*qkv_proj*'`;
`o_proj` and the MLP — the bulk of the drafter — still quantize. That
exclusion is mandatory, not a tuning choice.

`fc` (the projection from the target's captured layers into the draft)
is the one real knob, and it is a genuine trade rather than a free win —
see the Testing section for both models' numbers. The examples quantize
it; add `'*fc*'` to the exclude list to keep it in bf16.

`embed_tokens`, `markov_head` and `confidence_head` are excluded by
default: they are 2-D so the flat view treats them as GEMMs, but they
are embeddings or a single-output projection. `lm_head` is excluded by
the preset itself — unlike on a base model it is 37% of this drafter's
parameters, so `--quantize_lm_head` is a real lever (~1.9 GiB), but
measure AL first. The flag re-enables both of `lm_head`'s quantizers;
re-enabling only the weight one would ship a W+A checkpoint whose
`lm_head` has no `input_scale` while the config still advertises it as
quantized.

### Usage

```bash
# weight+activation FP8, calibration-free, lossless on both models measured below
python scripts/quantize_drafter.py \
    --drafter_path deepseek-ai/dspark_qwen3_8b_block7 \
    --qformat fp8 \
    --export_path ./dspark-qwen3-8b-fp8 \
    --exclude '*q_proj*' '*k_proj*' '*v_proj*' '*qkv_proj*'

# smallest: weight-only NVFP4
python scripts/quantize_drafter.py \
    --drafter_path nvidia/MiniMax-M3-DSpark \
    --qformat w4a16_nvfp4 \
    --export_path ./MiniMax-M3-DSpark-W4A16
```

Or end to end on Slurm — quantize, then measure AL — via the launcher
examples added here, one per target:

```bash
uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_dspark_ptq_nvfp4.yaml --yes
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/hf_dspark_ptq_nvfp4.yaml --yes
```

Serving one, if you are not going through `specdec_bench`:

```python
speculative_config = {
    "method": "dspark",
    "model": "./dspark-qwen3-8b-fp8",   # quantization is read from its config.json
    "num_speculative_tokens": 7,
}
```

### Testing

Two targets with different architectures, so the conclusions are not one
model's quirk:

* **Qwen3-8B** (dense transformer) +
[`deepseek-ai/dspark_qwen3_8b_block7`](https://huggingface.co/deepseek-ai/dspark_qwen3_8b_block7),
`block_size` 7, TP1.
* **Nemotron-3.5-Lightning-30B-A3B** (hybrid Mamba-MoE) +
[`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark`](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark),
`block_size` 8, TP8, with the mamba engine settings the model card pins
(`mamba_backend=flashinfer`, `mamba_ssm_cache_dtype=float16`, stochastic
SSM-cache rounding).

Both: MT-Bench 80 questions, greedy, one vLLM instance per point.

| recipe | activations | Qwen3-8B AL | vs bf16 | Nemotron-3.5 AL | vs
bf16 |
|---|---|---|---|---|---|
| bf16 baseline | — | 3.1423 | — | 4.3296 | — |
| **`fp8`** | static, `input_scale` 1.0 | **3.1457** | **+0.11%** |
**4.3289** | **-0.02%** |
| `fp8_pc_pt` | dynamic per-token | 3.1228 | -0.62% | 4.3411 | +0.26% |
| `w4a16_nvfp4`, `fc` in bf16 | bf16 (weight-only) | 3.0392 | -3.28% |
4.2899 | -0.92% |
| `w4a16_nvfp4`, `fc` quantized | bf16 (weight-only) | 3.0186 | -3.94% |
4.2334 | -2.22% |
| **`nvfp4`** | static, `input_scale` 1.0 | **3.0193** | **-3.91%** |
**4.2030** | **-2.92%** |

**FP8 weight+activation at the fixed `input_scale` of 1.0 is lossless on
both.** +0.11% and -0.02% are both inside run-to-run noise — the
Nemotron baseline was measured twice under identical settings and the
two runs differ by 0.94% (4.3093 / 4.3499), which sets the resolution of
that column. On the same reading, `fp8` and `fp8_pc_pt` are
indistinguishable on Nemotron; the dynamic variant only pulls ahead on
Qwen3. NVFP4 costs 3-4% on Qwen3 and 2-3% on Nemotron, i.e. the 4-bit
weight resolution is the real price and it is model-dependent but
bounded.

Whether to quantize `fc` is a per-model call rather than a general
recommendation — it buys a few percent of size for an AL cost that
differs by ~2x between these two drafters:

| `fc` bf16 → quantized | Qwen3-8B | Nemotron-3.5 |
|---|---|---|
| checkpoint size | 3.293 → 3.181 GiB (-3.4%) | 1.316 → 1.258 GiB
(-4.4%) |
| AL | 3.0392 → 3.0186 (-0.68%) | 4.2899 → 4.2334 (-1.32%) |

`fc` itself is only 3.5% (Qwen3) / 4.5% (Nemotron) of drafter
parameters; `embed_tokens` is the bulk (26% / 36%) and is excluded by
default.

The Qwen3 `w4a16_nvfp4` rows were measured in a later session than the
rest of that column; the `fc`-in-bf16 run reproduced the original number
to four decimals (3.0392), so the column is internally comparable.

Also validated on `nvidia/MiniMax-M3-DSpark`: `w4a16_nvfp4` runs in 67 s
on CPU, 9.98 GiB (fp32) -> 3.51 GiB; all 43 quantized tensors round-trip
within 0.0952 relative error; the 29 untouched tensors are bit-identical
to `bf16(source)`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (example-only)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ — validated manually as
above. Can add a `tests/examples/speculative_decoding/` test over a
small synthetic drafter if wanted before merge.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (example-only)
- Did you get Claude approval on this PR?: ❌ (not yet run)

### Additional Information

The measurements above are one drafter on one target with one benchmark;
the plateau's location and the ~3.5% NVFP4 gap should be re-measured
before assuming they carry to a different drafter.

Note when reading an exported checkpoint: `input_scale` is `amax/448`
for FP8 but `amax/(6*448)` for NVFP4, so the one fixed amax records as
1.0 in an FP8 checkpoint and 0.1667 in an NVFP4 one. Both mean the same
activation range.

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-08-26 22:04:09 +08:00
h-guo18 2b296b2f62 Support fine-tuning released DFlash/DSpark drafters (causal SWA, attention sink, warm start) (#2149)
# Support fine-tuning released DFlash/DSpark drafters (causal SWA,
attention sink, warm start)

### What does this PR do?

Type of change: New feature + bug fix

Adds what ModelOpt was missing to fine-tune an already-published
DFlash/DSpark draft
model. The concrete target is

[`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark`](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark)
on its hybrid Mamba/attention/MoE base, but every change is generic.

Before this PR that checkpoint could not be trained faithfully — or even
loaded: its
attention-sink tensors were dropped as unexpected keys, its block-causal
attention had no
implementation, and its capture layers were silently overwritten with
ModelOpt's defaults.

**New user-facing options** (all default to today's behavior, so
existing runs are unchanged):

| Option | Values | Purpose |
| --- | --- | --- |
| `dflash_draft_attention` | `bidirectional` (default) / `causal` |
Block-internal attention pattern. `causal` restricts a query at block
position `i` to draft positions `<= i`. |
| `dflash_attention_sink` | `false` (default) / `true` | Learnable
per-head `attention_sink_bias [num_heads]` on every draft layer — one
extra logit appended before the softmax and dropped after, so a head can
put probability mass nowhere instead of being forced to attend inside
its window (the GPT-OSS formulation). |
| `dflash_init_checkpoint` | path | Warm-start the draft from an
exported checkpoint instead of a random init. Any
missing/unexpected/wrong-shaped tensor raises rather than warns. |
| `dflash_architecture_config.target_layer_ids` | list | Which base
layers feed the draft's `fc`. Previously recomputed unconditionally with
no override. |

**Bugs fixed along the way** (each one silently corrupts training rather
than failing):

- The exporter hard-coded `dflash_config.causal: False` and only wrote
it under SWA, so even
a correctly-trained causal draft would be served non-causally. It now
reflects the trained
  setting, and emits `attention_sink_bias` when enabled.
- `_build_generate_swa_mask` returned `None` whenever `swa_window_size`
was unset, which
would have dropped the causal structure at generation time while
training used it.
- `target_layer_ids` was recomputed from the uniform default on every
convert. The released
drafter uses `[1,5,19,29,41,51]`; the default for a 52-layer base is
`[1,11,20,30,39,49]`
— *different layers*. Here it surfaced as a matmul shape error only
because the plane
counts disagreed; with a matching count it would have trained on the
wrong features
  silently.
- The streaming dataset assumed the draft's aux layers all sit below the
base's final layer
(`aux = planes[:-1]`, `target = planes[-1]`). A draft whose top aux id
*is* the final layer
cannot get an extra plane — vLLM captures each layer once — so
`final_aux_is_base_hidden`
now lets the last plane serve both roles. It is derived from the model,
not configured by
  hand.
- DSpark head weights load from either the flat layout ModelOpt exports
(upstream DeepSpec
convention) or the nested `markov_head.` layout the NVIDIA release uses.
Without the remap
the two `[131072, 512]` Markov tables — ~14% of the draft's parameters —
stay randomly
  initialized while everything else warm-starts, with no error.
- `nemotron_h` is enabled in `_FINAL_NORM_TYPE_BY_MODEL_TYPE`: despite
the hybrid stack,
`NemotronHModel.norm_f` is a plain RMSNorm, and without the entry the
offline/streaming
  fake base raises instead of reconstructing the distillation target.

### Usage

```yaml
dflash:
  dflash_init_checkpoint: /path/to/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark
  dflash_draft_attention: causal
  dflash_attention_sink: true
  dflash_swa_window_size: 1024
  dflash_block_size: 8
  dflash_mask_token_id: 990
  dflash_architecture_config:
    target_layer_ids: [1, 5, 19, 29, 41, 51]
```

A full worked example is at

`modelopt_recipes/general/speculative_decoding/dspark_nemotron35_warmstart.yaml`.

### Testing

**Unit tests** — 124 pass (`test_hf_dflash.py`, `test_hf_dspark.py`,
`test_hf_domino.py`,
`test_hf_dflash_offline.py`, `test_modeling_final_norm.py`), 32 of them
new: causal mask
structure (lower-triangular per block, no cross-block leakage, context
visibility
unchanged), the sink math (degenerates to plain attention at `-inf`,
absorbs mass
monotonically, receives gradient), warm-start load/reject paths, Markov
key remapping, and
explicit `target_layer_ids`.

**Checkpoint compatibility** — the released drafter loads with zero
missing/unexpected keys
and zero shape mismatches; all 77 tensors (6 attention sinks and both
Markov tables
included) match bit-exactly, and a training step runs with gradients
reaching the sink and
Markov parameters.

**End-to-end streaming training** — Nemotron-3.5 base served by vLLM (1
node, TP8) feeding
8 trainer GPUs over NIXL; the draft warm-starts from the released
checkpoint and trains with
`causal` + sink + SWA 1024. 128 Daring-Anteater conversations, 20 epochs
(the plot shows the
first 5, where the trend is clearest — the curves flatten after that):

![warm-start training
curves](https://raw.githubusercontent.com/h-guo18/Model-Optimizer/pr-assets/dspark_nemotron35_warmstart_curves.png)

Over the first 5 epochs loss falls **1.85 → 1.36** and train accuracy
rises
**0.25 → 0.49**; across the full 20 epochs they reach **1.21** and
**0.48** (peak 0.54)
before flattening. This validates the pipeline end-to-end — capture
layers, plane split,
mask direction, sink loading and warm-start weights all have to be right
for this curve to
appear. It is *not* a model-quality result: 128 samples over 20 epochs
overfits by
construction, and the corpus is not generated by the base model, so the
absolute numbers are
not meaningful.

### TODO (follow-up)

**A complete, robust checkpoint/config converter.** Both conversions are
handled ad hoc here:

- *Draft config → training config.* The recipe transcribes ~15 fields by
hand from the
drafter's `config.json`. Only the shape-bearing ones
(`num_hidden_layers`,
`num_attention_heads`, `intermediate_size`, `markov_rank`) fail loudly
when mistyped; the
rest — `mask_token_id`, `causal`, `swa_window_size`, `block_size` —
train "successfully" on
a wrong value and only surface later as a mysteriously low acceptance
length. A converter
should derive the whole block from the checkpoint, including its aliases
(`pard_token`,
`dspark_markov_rank`, `dflash_query_causal`, top-level `sliding_window`
/
  `attention_sink_bias`) and duplicated fields.
- *Weight layout.* The `markov_head.` remap is a load-time hook. A
converter should normalize
layouts explicitly, and decide whether export should also emit the
release's aliases so a
round-trip reproduces the original format (today it renames
`architectures` to
  `DFlashDraftModel`).
- *Base config.* Serving this base on vLLM needs its `config.json`
layer-type vocabulary
updated for the transformers-5 path (`mamba` → `linear_attention`,
`attention` →
`full_attention`, plus a matching `hybrid_override_pattern`). That is
done by hand today and
  is not covered by this PR.

### Before your PR is "*Ready for review*"

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added configurable causal or bidirectional attention for DFlash
models.
* Added optional attention sinks, checkpoint warm starts, and explicit
target-layer selection.
* Improved streaming data handling for shared auxiliary and base hidden
states.
* Added Nemotron-3.5 Lightning DSpark warm-start training and serving
recipes.

* **Bug Fixes**
  * Preserved configured attention behavior during model export.
* Prevented warm-start checkpoints from being reapplied during
restoration.
* Improved checkpoint compatibility, validation, and attention-mask
handling.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-08-23 20:49:43 +08:00
Keval MorabiaandClaude Opus 5 fbcdc16c2d Remove deprecations marked in 0.45 and 0.46 (#2182)
### What does this PR do?

Type of change: Backward breaking change (deprecation removal)

Ahead of the 0.47 code freeze, this removes every deprecation still
outstanding from the previous two releases (0.45 and 0.46). Two are
intentionally left in place: the **Python 3.10** drop and the
**transformers 4.x** drop

| Deprecation | Marked in | Replacement |
| --- | --- | --- |
| `--auto_quantize_bits` / `_method` / `_score_size` / `_cost_model` /
`_active_moe_expert_ratio` | 0.46 | AutoQuantize `--recipe` |
| `examples/llm_ptq` symlink + `examples/vlm_ptq/` forwarder | 0.46 |
`examples/hf_ptq` (`--vlm` for VLMs) |
| `QuantizationArgumentsWithConfig` alias | 0.45 |
`QuantizationArguments` |
| `QFORMAT_ALIASES` short names | 0.45 | canonical preset basenames |
| `layerwise` bool + flat `layerwise_checkpoint_dir` | 0.45 | nested
`layerwise: {enable, checkpoint_dir}` |
| in-trainer `quant_cfg` / `--quant_cfg` | 0.45 | `--recipe` |

#### Two things worth a closer look

**1. The `use_sequential` alias goes too.** It is the pre-#1251 alias on
`QuantizeAlgorithmConfig.layerwise` and only ever carried a bool. Once
the bool form is rejected it cannot accept a valid value, so keeping it
would only produce a differently-worded validation error. Note the
direction is breaking either way (`extra="forbid"`): a pre-0.45
`modelopt_state` carrying `use_sequential: True` or a top-level
`layerwise_checkpoint_dir` now fails validation instead of being
migrated.

**2. Removing in-trainer `--quant_cfg` required two new recipes.** The
`examples/gpt-oss` QAT flow ran on `--quant_cfg
MXFP4_MLP_WEIGHT_ONLY_CFG` and no `general/ptq/` recipe covered it. This
PR adds `general/ptq/mxfp4_mlp_weight_only` and
`general/ptq/nvfp4_mlp_weight_only`, verified to `model_dump` identical
to `mtq.MXFP4_MLP_WEIGHT_ONLY_CFG` / `mtq.NVFP4_MLP_WEIGHT_ONLY_CFG`,
and migrates the gpt-oss README, both SFT configs, `sft.py` and
`tests/examples/gpt-oss/test_gpt_oss_qat.py`. `examples/llm_qat` was
already recipe-only.

### Usage

```bash
# AutoQuantize: --auto_quantize_* flags -> an AutoQuantize recipe
scripts/huggingface_example.sh --model $HF_PATH \
  --recipe general/auto_quantize/nvfp4_fp8_at_5p4bits --calib_batch_size 4

# --qformat / --quant_cfg: short name -> canonical preset basename
#   int8_sq -> int8_smoothquant                nvfp4_mse           -> nvfp4_w4a4_weight_mse_fp8_sweep
#   int8_wo -> int8_weight_only                nvfp4_local_hessian -> nvfp4_w4a4_weight_local_hessian
#   w4a8_awq -> w4a8_awq_beta                  fp8_pb_wo           -> fp8_2d_blockwise_weight_only
#   nvfp4_awq -> nvfp4_awq_lite                fp8_pc_pt           -> fp8_per_channel_per_token
scripts/huggingface_example.sh --model $HF_PATH --quant int8_smoothquant

# VLM PTQ: examples/vlm_ptq -> examples/hf_ptq with --vlm
scripts/huggingface_example.sh --model $HF_PATH --quant fp8 --vlm

# gpt-oss QAT: --quant_cfg <CFG name> -> --recipe <recipe path>
accelerate launch --config_file configs/zero3.yaml sft.py \
  --config configs/sft_full.yaml --model_name_or_path openai/gpt-oss-20b \
  --recipe general/ptq/mxfp4_mlp_weight_only --output_dir gpt-oss-20b-qat
```

```python
# Layerwise calibration: bool / flat key -> nested LayerwiseConfig
quant_cfg["algorithm"] = {"method": "gptq", "layerwise": {"enable": True, "checkpoint_dir": "/path"}}
```

### Testing

- `tests/unit/recipe` (229 passed),
`tests/unit/torch/quantization/test_config_validation.py` (79 passed),
`tests/examples/hf_ptq/test_hf_ptq_args.py` (23 passed).
- Verified the two new recipes `model_dump` identical to the `mtq.*_CFG`
constants they replace.
- `ruff check modelopt/ examples/ tests/` clean; `ruff format --check`
clean on all changed Python files.
- GPU suites
(`tests/gpu/torch/export/test_unified_hf_export_and_check_safetensors.py`,
`test_accelerate_gpu.py`, `test_gptq.py`) had their preset / layerwise
literals updated but were not run locally — relying on CI.
- `examples/llm_qat/ARGUMENTS.md` is hand-edited to match what the
`generate-arguments-md` hook emits; the generator could not run locally
(missing `transformers` package metadata in this environment).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — that is the point of the PR:
it removes shims deprecated in 0.45/0.46. Callers must move to the
replacements in the table above. Additionally, a pre-0.45
`modelopt_state` carrying `use_sequential` or a top-level
`layerwise_checkpoint_dir` will now fail config validation.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — existing tests migrated to
the surviving APIs;
`TestLayerwiseNestedConfig::test_legacy_forms_rejected` pins that the
bool form, the `use_sequential` alias and the flat checkpoint-dir key
are all rejected. Tests covering the removed shims were deleted.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

Follow-up: the transformers 4.x drop deprecated in 0.46 is still
outstanding and will need its own PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## New Features
- Added MXFP4 and NVFP4 weight-only quantization recipes for MLP and MoE
layers.
- Added shared layer exclusions for more accurate effective-bits
calculations.

## Improvements
- Updated PTQ, QAT, GPT-OSS, deployment, and quantization-format
examples with current recipe names and configuration formats.
- Standardized layerwise settings under nested configuration fields.

## Breaking Changes
- Removed deprecated AutoQuantize options, `quant_cfg` usage, format
aliases, legacy layerwise settings, and compatibility example paths.
- Recipe-based and nested configuration forms are now required.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 11:43:47 +05:30
Keval MorabiaandClaude Opus 4.8 58ad6edc5f Fix pruned-HF export fallback + add Nemotron-3.5-Lightning launcher examples (#2196)
### What does this PR do?

Type of change: Bug fix + new example

Two related changes for the Megatron-Bridge Minitron prune/quantize
launcher flows:

1. **Fix pruned-HF export crash on containers that reject config-only
save.**
   `#2159` added a config-only HF export path gated only on
`hasattr(AutoBridge, "from_hf_config")`. Some Megatron-Bridge versions
   (e.g. `nemo:26.04`) expose `from_hf_config` but reject a config-only
`save_hf_pretrained` (`ValueError: save_hf_pretrained requires a
pretrained
HuggingFace model`), so `prune_minitron.py` crashed instead of using the
intended dummy-model fallback. Now it attempts the config-only save and
   falls back to the dummy-model path on `ValueError`.

2. **Add Nemotron-3.5-Lightning-30B-A3B launcher examples**
(`mbridge_prune.yaml`,
`mbridge_quantize.yaml`) on `nemo:26.08`. Prune targets 3B active with
an
MMLU gate; quantize runs W4A16 NVFP4 4/6 PTQ via the `w4a16_nvfp4_4o6`
recipe with `tp_size=1` (static-block NVFP4 MSE is unsupported with
TP>1).

### Usage

```shell
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_prune.yaml --yes
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_quantize.yaml --yes
```

### Testing

Verified end-to-end on OCI-HSG:

- **Nano prune (`nemo:26.04`)** — exercises the fallback path:
config-only save
  raised the `ValueError`, the fallback caught it and exported via the
dummy-model path. `mmlu_10pct_bs32 = 0.5196` (gate 0.50) PASS; vLLM gen
PASS.
- **Lightning prune (`nemo:26.08`)** — config-only export path: `score =
0.6000`
  (gate 0.58) PASS, 3.00B active params; vLLM gen PASS.
- **Lightning quantize (`nemo:26.08`)** — recipe PTQ + unified-HF
export;
  MMLU `0.7741` (gate 0.75) PASS.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- launcher example
configs + fallback path exercised by CI prune/quantize jobs -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- fix is for a bug introduced in the same unreleased cycle
(#2159); rest are example configs -->
- Did you get Claude approval on this PR?: ❌ <!-- pending /claude review
-->

### Additional Information

The fallback fix addresses the `mbridge_prune` launcher CI failure
introduced by #2159.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added a pruning workflow for Nemotron-3.5-Lightning-30B-A3B with
calibration, quality scoring, checkpoint export, and multi-GPU
generation.
- Added a four-GPU NVFP4 W4A16 quantization workflow with Hugging Face
conversion and MMLU evaluation.
- **Bug Fixes**
- Improved hybrid model export by falling back to dummy-model export for
supported configuration-only export failures.
- Added clearer logging and handling for supported export failures while
preserving unrelated errors for investigation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-17 22:15:47 +05:30
skierat c4129b6e03 Add Cosmos3 Nano DFlash multimodal training recipe (#2053)
### What does this PR do?

  Type of change: new example

Adds an end-to-end Cosmos3 Nano DFlash training recipe for multimodal
speculative decoding.

- Adds a notebook that prepares data, launches synthetic generation in
Slurm, trains a DFlash draft model, exports it, and provides a vLLM
smoke-test command.
- Adds PAI-Understanding, VQA v2, and multilingual prompt sharding and
distributed-generation helpers.
- Adds an atomic, multimodal-safe merge and conservative deduplication
flow.
- Extends the VLM data collator to handle structured image/video
messages, configurable visual bounds, and fixed DFlash sequence lengths.
- Hardens generation launch scripts and preserves truncated generated
responses.

  ### Usage

  ```bash
  cd examples/speculative_decoding/recipes

  export MODEL_PATH=/path/to/cosmos3-nano
export
PLAIN_TEXT_INPUT=/path/to/nemotron-chat-or-approved-user-data.jsonl

  jupyter lab train_dflash_cosmos3_nano.ipynb

  Run the notebook in order:

  1. Configure paths.
2. Prepare prompts on a CPU-only node and generate target completions in
a Slurm GPU allocation.
  3. Merge the four required sources and submit training.
  4. Export a saved checkpoint and run the vLLM deployment smoke test.

  ### Testing

- jq empty
examples/speculative_decoding/recipes/train_dflash_cosmos3_nano.ipynb
  - bash -n on the modified launch, worker, and recipe shell scripts.
- Ran a two-step Cosmos3 Nano DFlash Slurm smoke job; it completed and
wrote modelopt_state.pth.
- Not run: pytest
tests/unit/torch/speculative/plugins/test_hf_speculative_offline.py
(pytest is unavailable in the current environment).

  ### Before your PR is "Ready for review"

  - Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in CONTRIBUTING.md?: N/A
  - Did you write any new necessary tests?: ✅
  - Did you update CHANGELOG.rst?: N/A
  - Did you get Claude approval on this PR?: N/A

  ### Additional Information

Security follow-up required before marking ready: the notebook hardcodes
model.trust_remote_code=true and --trust_remote_code. Either
parameterize
this with a default of false, or obtain and document a security
exception.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added end-to-end multimodal workflows for dataset preparation,
distributed generation, result merging, training, export, and deployment
testing.
* Added support for image and video inputs, multiple dataset formats,
resumable JSONL generation, configurable serving, and parallel
processing.
* Added configurable prompt, media, token, sequence, temperature, and
tensor-parallel settings.
* **Bug Fixes**
* Improved truncated-response handling, assistant-label processing,
validation, health checks, cleanup, deduplication, media resolution,
atomic outputs, and failure reporting.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Slawomir Kierat <skierat@nvidia.com>
2026-08-14 21:20:17 +08:00
Jenny Chen a57fb44d46 Preserve HF PTQ checkpoint sidecar files [NV BUG 6491822] (#2060)
### What does this PR do?

Type of change: Bug fix

In `hf_ptq.py` when exporting a PTQ checkpoint, it would drop some files
from the original BF16 checkpoint because it uses a whitelist pattern to
allow certain files. However that is brittle and can drop files such as
reasoning parsers.

Now we make hf_ptq.py match Megatron-Core export behavior by copying all
non-safe tensor files, but filter only allowed non-safetensor files for
more safety.


### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved Hugging Face checkpoint handling to preserve eligible sidecar
files while excluding weights, indexes, stale quantization metadata, and
unsupported artifacts.
* Preserved existing export files and applied consistent file filtering.
  * Improved snapshot resolution when remote code is disabled.
* Ensured unified exports handle generation configuration files
correctly.

* **Tests**
* Added coverage for sidecar copying, exclusions, existing-file
preservation, supported file patterns, and snapshot downloads.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
2026-08-13 20:09:39 -07:00
yueshen2016 686da8d893 feat(megatron-bridge): SFT-masked data support in distillation (#2113)
## What does this PR do ?

**Type of change:** New feature

**Overview:** Adds SFT-masked data support to the Megatron-Bridge
distillation example, so a
model can be distilled on prompt/response pairs with the loss masked to
the response.

Today `examples/megatron_bridge/distill.py` only consumes
pretraining-style data — `GPTDataset`
over pre-tokenized blends with `NullTokenizer` — so the loss is computed
over every token. When
distilling an instruction-tuned model it is usually preferable to train
on prompt/response pairs
and mask the loss to the response, matching how the model was
fine-tuned.

## Usage

```bash
python examples/megatron_bridge/distill.py \
  --teacher_hf_path <teacher> --student_hf_path <student> \
  --sft --sft_dataset_root /path/to/data \
  ...
```

where `/path/to/data` holds `training.jsonl` / `validation.jsonl` of
records:

```json
{"input": "<prompt>", "output": "<response>"}
```

## How it works

Switches the data path to Bridge's `FinetuningDatasetConfig` (NeMo-style
`GPTSFTDataset`):

* `prompt_template="{input}{output}"` tokenizes input+output verbatim —
adjacent placeholders,
  no separator — so the text is fed exactly as provided
* `label_key="output"` with `answer_only_loss=True` masks the loss to
the response
  (`answer_start_idx == len(context_ids)`)
* `truncation_field="input"` truncates the context when a pair exceeds
`seq_length`

Two supporting changes, both scoped to `--sft`:

* **Tokenizer.** SFT reads raw text, so it uses the model's real
HuggingFace tokenizer. The
pretraining path consumes pre-tokenized data and keeps `NullTokenizer`.
* **Loss reduction.** A response-only mask requires per-token loss to
combine correctly across
context-parallel ranks, so `calculate_per_token_loss` is enabled and
`average_in_collective`
  is disabled. Both are untouched on the pretraining path.

## Testing

Used for quantization-aware distillation of Nemotron-Nano-3 (W4A16
NVFP4) at `seq_length=32768`
with CP>1: 200 iterations, logits-distillation loss `3.37e-2 -> 1.91e-2`
monotonically, router
`seq_load_balancing_loss` steady, and the resulting checkpoint exports
and serves correctly.

Opt-in: without `--sft` the existing mock/blend data path is unchanged.

## Before your PR is "Ready for review"

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes — purely additive and
opt-in behind `--sft`.
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added supervised fine-tuning (SFT) support for distillation workflows.
* Added configuration and validation for SFT dataset locations and
supported inputs.
* Added raw prompt/completion JSONL datasets with response-only loss
masking.
* Added truncation, end-of-sequence handling, and student-tokenizer
support without automatic chat templates or BOS tokens.
  * Added matching student and teacher vocabulary validation.
  * Preserved existing mock and GPT dataset modes for non-SFT runs.

* **Documentation**
* Documented required filenames, record format, tokenizer behavior, and
completion-only loss masking.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: James Shen <yueshen@nvidia.com>
2026-08-13 01:10:29 +00:00
Chenjie Luo b96841db3e Add optional MLflow tracking to the vLLM fake-quant server (#2120)
### What does this PR do?

Type of change: new feature

Wires `examples/vllm_serve/vllm_serve_fakequant.py` up to
`modelopt.torch.utils.mlflow` via `--mlflow <tracking-uri>`, the same
way #2023 did for `hf_ptq.py`, so a fake-quant serve records **what it
actually quantized** and an evaluation of that endpoint can be traced
back to a recipe. Without the flag, behavior is unchanged — every hook
is gated on it.

Three design points worth review:

1. **The run is recorded in the vLLM worker, not the launcher.**
`vllm_serve_fakequant.py` is the API-server frontend; the engine and its
workers are separate processes whose stdout it never sees, so a run
opened there would capture none of the calibration. The launcher instead
only settles the tracking configuration — validating the URI, naming the
experiment, recording the command the user actually typed — and
publishes it through the environment, which is how every other setting
in this example (`QUANT_CFG`, `RECIPE_PATH`, …) already reaches the
workers. Global rank 0 opens the run, so a TP-8 serve produces one run.

2. **The run covers load-through-warm-up, not the server's lifetime.**
It opens *before the weights load*, so an unreachable server or a
missing token fails in seconds rather than after a load and a full
calibration, and it closes `FINISHED` once the model is quantized and
warmed up. A run that stayed open for the serving lifetime would never
close cleanly on SIGTERM.

3. **`recipe/quant_cfg.yaml` is only written on the preset path.** With
`RECIPE_PATH`, `get_quant_config` returns the recipe's `quantize`
section unchanged and `resolved_recipe.yaml` already carries it. With
`QUANT_CFG`/`KV_QUANT_CFG` it is the *only* record of what ran: the
params carry the preset names, while the config reaching `mtq.quantize`
is those two deep-copied, merged, and — for an MLA model — extended at
runtime with `*kv_c_bmm_quantizer` / `*k_pe_bmm_quantizer` by inspecting
the loaded model.

Uploaded artifacts:

| Artifact | Contents |
| --- | --- |
| `command.txt` | The launcher's invocation, copy-pasteable, credentials
masked |
| `version.txt` | The ModelOpt version that ran |
| `recipe/resolved_recipe.yaml` | `RECIPE_PATH` with its `$import`s
expanded |
| `recipe/quant_cfg.yaml` | Merged `QUANT_CFG`/`KV_QUANT_CFG` + MLA
fixup (preset path only) |
| `logs/<script>.log` | The rank-0 worker's stdout/stderr, including a
crash traceback |
| `summary/quant_summary.txt` | The per-quantizer summary |

Plus the quantization *and* serving settings as searchable params, and
`user` / `hostname` / `modelopt_version` / `git_sha` / `vllm_version`
tags. The `checkpoint_path` tag matches the one `hf_ptq.py` sets, so a
checkpoint's PTQ run and every serve of it join up.

Two small library additions, both consumed by the new example module:

- `command_text(argv=None)` — records another process's invocation,
since a spawned worker's own `sys.argv` is vLLM plumbing rather than
anything a user typed.
- `MlflowRunLogger.log_text()` — uploads a value settled midway through
a run, so a crash during calibration still keeps the config that caused
it.

The example `Dockerfile` installs the `mlflow` extra; the client remains
optional and is imported only once tracking is enabled.

### Usage

```bash
RECIPE_PATH=<recipe.yaml> python vllm_serve_fakequant.py <model_path> -tp 8 \
  --host 0.0.0.0 --port 8000 \
  --mlflow https://<your-mlflow-server>/
```

```
[mlflow] tracking to https://<your-mlflow-server>, experiment $USER/vllm_serve_fakequant/<model>-<recipe>
(Worker_TP0) [mlflow] run: https://<your-mlflow-server>/#/experiments/19/runs/1c6679448f25...
```

`--mlflow-experiment` / `--mlflow-run-name` override the defaults.
`$MLFLOW_TRACKING_URI` enables tracking on its own and is best-effort;
an explicit `--mlflow` overrides it and fails loudly.

> This is the **quantization** tracking server. It is unrelated to any
server an evaluation harness exports its scores to — NeMo Evaluator
Launcher has its own `export.mlflow.tracking_uri`. The README calls this
out.

### Testing

**Unit — 87 passing**
(`tests/examples/vllm_serve/test_vllm_mlflow_utils.py`, 33 new;
`tests/unit/torch/utils/test_mlflow.py`, +5). `vllm_mlflow_utils`
deliberately imports no vLLM, so the whole launcher→worker handover is
covered without a GPU, a server, or the mlflow client.

**End to end on aws-cmh** (4× GB300, `simple_evals.gpqa_diamond`,
Nemotron-3.5-Lightning-30B-A3B-BF16 fake-quantized with
`general/ptq/nvfp4_mlp_only-kv_fp8_cast`): run `FINISHED` in 261.5 s,
opened by `Worker_TP0` only, all artifacts present and verified by
content — `command.txt` held the launcher's invocation rather than the
worker's spawn argv, and `resolved_recipe.yaml` was 6797 B against 1845
B of source. 104 quantizers enabled (92 NVFP4 dynamic block-16 expert
weight/input with calibrated amax, 12 FP8 KV bmm). The eval then ran to
completion against the served endpoint, 22/22 requests HTTP 200.

Two bugs the hardware run caught, both fixed here with regression tests:

- `--mlflow_run_name` was rejected. vLLM's
`FlexibleArgumentParser.parse_args` rewrites **every** `--foo_bar` to
`--foo-bar` before matching, so a flag registered only under the
underscored spelling is unreachable from its CLI. Both spellings are now
registered. A unit test on a plain `ArgumentParser` could not have
caught this.
- `recipe/quant_cfg.yaml` uploaded a Python `repr` blob under a `.yaml`
name: a recipe's `quantize` is a `QuantizeConfig`, `yaml.safe_dump`
raises `RepresenterError` on it, and the old JSON fallback stringified
the object. `_dump_yaml` now unwraps pydantic via
`model_dump(mode="json")` and raises otherwise, with the caller
downgrading that to a warning so a bad config cannot take down a serve.

**Known coverage gap:** the preset (`QUANT_CFG`/`KV_QUANT_CFG`) path —
the only one that now writes `recipe/quant_cfg.yaml` — is covered by
unit test but has not been exercised on hardware; the canary used
`RECIPE_PATH`. Likewise the case where `$MLFLOW_TRACKING_URI` is present
*inside* the deployment container and `--mlflow` overrides it is
unit-tested only: NeMo Evaluator Launcher forwards only declared env
vars, so the eval server's URI never entered the container in the
canary.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — new optional flags only; no
`--mlflow` means no behavior change.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependency. Uses the existing optional `nvidia-modelopt[mlflow]` extra
(`mlflow-skinny`, Apache-2.0) added in #2023; the example `Dockerfile`
now installs it. No code copied from other sources.
- Did you write any new necessary tests?: ✅ — 38 new tests.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — 0.47 Misc.
- Did you get Claude approval on this PR?: ❌ — `/claude review` not yet
run.

### Additional Information

Follows #2023, which added `MlflowRunLogger` and the `hf_ptq.py`
integration.

Note for anyone tracking from an OCI cluster:
`mlflow-modelopt.nvidia.com` is unreachable from oci-nrt and oci-hsg.
TCP 443 completes and the connection is then reset on the first
application byte, regardless of SNI or protocol, one RTT away — the PDX
PaaS ingress appears to apply a source-IP policy, and the OCI clusters
egress from Oracle-owned addresses (`155.248.190.0`, `168.110.199.1`)
rather than NVIDIA's. gcp-nrt, aws-cmh and cw-dfw all reach it. This is
an infrastructure matter, not a property of this change, but it
determines where the feature is usable today.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added optional MLflow tracking for vLLM fake-quantization serving
runs.
* Records serving, quantization, worker, and invocation metadata,
including configuration and summary artifacts.
* Supports tracking URI, credentials, environment, and command-line
configuration.
  * Added command and text artifact logging for active MLflow runs.
* **Documentation**
* Documented setup, configuration, recorded artifacts, lifecycle, and
fallback behavior.
  * Updated the example container to include MLflow support.
* **Tests**
* Added comprehensive coverage for tracking configuration, logging,
failures, and disabled tracking.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-08-13 00:00:21 +00:00
ZhiyuandClaude Opus 5 6261f854aa docs: rebuild the unified HF deployment support matrix from the deploy test suite (NVBug 6550792) (#2087)
### What does this PR do?

Type of change: documentation

Fixes [NVBug 6550792](https://nvbugspro.nvidia.com/bug/6550792) /
OMNIML-5693.

The **Unified HF Checkpoint Deployment Model Support Matrix** listed 9
model families and **no VLMs**, while
`tests/examples/hf_ptq/test_deploy.py` declares deployment cases for ~80
checkpoints across TRT-LLM, vLLM, and SGLang — including `Qwen2.5-VL`,
`Qwen3-VL-235B`, and `Nemotron-3-Nano-Omni`. QA (the filer) could not
use the doc to scope testing, and users could not tell what is actually
covered.

Filing also surfaced that the matrix lived in **three places that had
drifted apart**: only the `.rst` listed Qwen3-VL, only the README listed
Qwen3.5 MoE, and the skill reference had neither.

#### Changes

1. **Rebuilt the matrix in `docs/source/deployment/3_unified_hf.rst`**
from `test_deploy.py`, split into language models,
vision-language/multimodal, speculative decoding drafters, and
diffusion.

2. **Stated plainly what the matrix is and is not.** Review established
that the original "CI-validated" framing claimed more than the suite
substantiates, so a *What this matrix is based on* section now leads
with two limits:
- The suite is marked `release` and collects only under `--run-release`,
which **no workflow passes** — these are declared cases, not PR-gated
coverage.
- Each case is a **load-and-generate smoke check on the text path**: no
accuracy, no image/audio input, no diffusion output, no verification
that speculative decoding engages.

The legend follows from that: ✅ = declared in the suite, ⚠ = expected to
work but not a suite entry (or an entry that does not exercise the
feature the row names), `-` = not in the suite. Sections that would
otherwise over-read carry their own qualifiers — VLM rows are labelled
text-only smoke coverage, and Medusa and Wan 2.2 are ⚠ with the reason
stated.

3. **Removed the two duplicate copies**, replacing them with links, so
there is one table to maintain.

4. **Fixed stale prose**: the deployment tabs still claimed FP8-only
support on vLLM v0.6.5 and a source build of SGLang main from Jan 2025,
both contradicting the version table above them. The TRT-LLM floor moves
to v1.2.0, qualified as the oldest version stated rather than the oldest
that works.

5. **Dropped the Phi series** from the deployment matrix, following
#2115 (NVBug 6563509) and confirmation that Phi-4 is being deprecated.

### Usage

N/A — documentation only.

### Testing

- `docutils` parse of the modified `.rst`: no warnings or errors from
the new content; all 5 tables parse with every cell in the correct
column.
- Cell contents cross-checked against `test_deploy.py` by AST-parsing
the `ModelDeployerList(...)` calls rather than by eye; the scope caveats
were each verified against `tests/_test_utils/deploy_utils.py`.
- `pre-commit run --files …` passes; `build-docs` green.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — documentation only
- Did you get Claude approval on this PR?: ❌ — not yet run

### Additional Information

**Two known follow-ups, neither in scope here:**

1. **Nothing enforces that the doc matrix tracks `test_deploy.py`.**
Consolidating to one copy removes the three-way drift but not the
doc-vs-test drift; a generator plus a CI check would close it.
2. **The release deployment suite does not run in CI.** Wiring it into
per-backend release CI is what would let ✅ mean "verified to pass"
rather than "declared". That needs GPU capacity across three backends
and should be tracked on its own.

**For the filer (@Kenny Kang):** the ✅ cells are the scope the release
deploy suite declares, and `test_deploy.py` carries the checkpoint, TP
size, and minimum SM version per entry — but please read the legend
first, since those cases are not currently executed by CI.

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 01:37:07 +05:30
Keval Morabia bee497de03 Fix EAGLE3 offline dump skipping all conversations on newer transformers (#2172)
### What does this PR do?

Type of change: Bug fix

**Fix EAGLE3 offline hidden-state dump silently skipping every
conversation on newer `transformers`.**

`tokenizer.apply_chat_template(...)` returns a **`BatchEncoding`** (dict
of `input_ids` + `attention_mask`) on `transformers>=5` rather than a
`list[int]`, so `len(input_ids)` evaluated to **2** (the number of dict
fields), tripping the `num_input_tokens <= 10` "too short" filter for
**every** conversation. The dump wrote **zero `.pt` files** and offline
EAGLE3 training aborted with `No .pt files found`.

The token-id extraction is consolidated into
`modelopt.torch.speculative.utils.get_conversation_input_ids`, which
normalizes the result to a flat `list[int]` (unwrapping `BatchEncoding`
/ 2-D tensor / batch-wrapped list, asserting the shape so a future
`transformers` change fails loudly instead of silently). It is called
from all three offline-dump entry points that shared the bug:

-
`examples/speculative_decoding/collect_hidden_states/compute_hidden_states_trtllm.py`
-
`examples/speculative_decoding/collect_hidden_states/send_conversations_for_hiddens.py`
- `examples/speculative_decoding/scripts/send_conversation_vllm.py`

(the two `send_conversation*` scripts additionally indexed/`decode()`d
the `BatchEncoding`). Also fixes the `add_generation_template` ->
`add_generation_prompt` typo at each site.

### Testing

- `tests/unit/torch/speculative/test_speculative_utils.py` — asserts the
helper returns the exact token-id sequence of the rendered chat prompt,
and pins every `apply_chat_template` return shape (`BatchEncoding`, 2-D
tensor, batch-wrapped list, plain list) to a flat `list[int]` via
deterministic stubs, so the fixed branch is covered regardless of the
installed `transformers` version.
- **End-to-end on ComputeLab (H100, TRT-LLM 1.3.0rc20):** reran the
exact dump on the 100 conversations that previously failed. Before:
0/100 (0 `.pt` files). After: **97/100** (97 `.pt` files; the 3 skips
are genuinely `> max_seq_len`).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
(`tests/unit/torch/speculative/test_speculative_utils.py`)
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: 🔄 `/claude review` run;
findings addressed, re-review pending

### Additional Information

Surfaced by an nmm-sandbox CI run where `Qwen3-8B_EAGLE3_offline` failed
after the container bump to `tensorrt-llm/release:1.3.0rc20`; the
auto-blame heuristic mis-attributed it to an unrelated MLflow commit.
`compute_hidden_states_vllm.py` is unaffected (it routes through
`common.tokenize_with_loss_mask`, which passes `return_dict=True`).

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-08-12 18:29:56 +00:00
Sepehr Sameni 10145db53f Add link to puzzletron_v2 (#1996)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added a top-level Puzzletron overview with a link to the step-by-step
algorithm tutorial.
* Added a reference to the experimental Puzzletron branch for advanced,
production-scale usage.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Sepehr Sameni <ssameni@nvidia.com>
2026-08-12 15:37:43 +00:00
Keval MorabiaandClaude Opus 5 f99523279a Minitron pruning fixes for Nemotron-3.5-Lightning-30B-A3B and Deepseek (#2159)
### What does this PR do?

Type of change: Bug fix + new feature

Two model families that could not be pruned end-to-end now can:

- **Nemotron-3.5-Lightning**
(`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16`) — a native
`NemotronHForCausalLM` that ships without remote code and carries MTP
heads. Fixes a calibration crash and HF-export failures on the modern
Megatron-Bridge / transformers stack.
- **DeepSeek-V3** — fixes an MLA Q-LoRA crash during calibration, and
adds a `candidate_filter` search option to `mcore_minitron` so its
MoE-FFN dimensions stay prunable while remaining representable in HF.

Also makes a rank-local failure under pipeline parallelism fail fast
instead of stalling.

#### 1. Nemotron Lightning: prune + HF export
(`examples/megatron_bridge/prune_minitron.py`)

1. **MTP calibration crash.** On newer Megatron-LM, `mtp_process` is
derived from the hybrid *pattern*, not from `mtp_num_layers`. Setting
only `mtp_num_layers=0` in the calibration provider overrides was
insufficient: the provider's `finalize()` re-appended the MTP suffix to
`hybrid_layer_pattern` (because `mtp_hybrid_override_pattern` was still
set and `mtp_use_repeated_layer=True`), so `mtp_process=True` while
`mtp_num_layers=0` and the calibration forward hit `assert
self.config.mtp_num_layers > 0`. Fix: also clear
`mtp_hybrid_override_pattern` in the calibration overrides so MTP is
fully disabled (MTP heads are dropped from the pruned model, as before).

2. **HF export via a config-only bridge (hybrid models only).** The old
export built a *dummy* HF model to obtain the bridge, then streamed
weights. This breaks on native NemotronH because (a) native
`NemotronHConfig` makes `hybrid_override_pattern` a read-only property,
and (b) transformers 5.12 saves the input embedding under a different
key than the bridge mapping expects (`backbone.embedding` vs
`backbone.embeddings`); the mismatch made `build_conversion_tasks` drop
the embedding task on its owning rank, leaving an owner-less PP
placeholder that crashed `save_hf_weights` with `Object must exist on at
least one PP rank`. Fix: stream weights through a **config-only** bridge
(`AutoBridge.from_hf_config(hf_cfg).save_hf_pretrained(...)`), available
since Megatron-Bridge 0.5.0 (nemo:26.06). A config-only bridge has
`hf_keys=None`, so the embedding task is never dropped, no dummy model
is built, and the output uses the canonical HF key names.

This is **restricted to hybrid providers**, which are the only models
that need it; non-hybrids keep the dummy-model path that CI has always
exercised.

Writing the source artifacts is now rank-0-only. Every rank used to
write the source `config.json`, which races with the pruned
`config.json` that `save_hf_pretrained` writes from rank 0 alone: a late
write from another rank leaves a checkpoint whose config does not match
its weights. This reproduced intermittently on both Qwen3 and NemotronH
before the fix, and 3/3 clean runs after.

`save_hf_pretrained` takes no `trust_remote_code` argument — it reads
the flag **off the bridge** to fetch the source checkpoint's artifacts,
and `from_hf_config` cannot infer it because
`AutoConfig.from_pretrained` consumes the kwarg rather than storing it
on the config. So the flag is set explicitly on the bridge instance;
otherwise remote-code models would silently lose it.

3. **Config write-back correctness:**
- `hybrid_override_pattern` is only written for older remote-code
configs that lack `layer_types`; native configs carry the cadence in
`layer_types` (read-only `hybrid_override_pattern` is skipped).
- `n_shared_experts` is preserved (a fixed count) instead of being
re-derived by `moe_shared_expert_intermediate_size //
moe_ffn_hidden_size`, which is DeepSeek-style logic that would corrupt
NemotronH's count.

Non-hybrids, VLMs, and Megatron-Bridge builds without config-only export
keep the dummy-model path, with a `warn_rank_0` when a hybrid has to
fall back. The README's `transformers<5` workaround is **removed**: it
existed because the dummy-model path broke on transformers 5, and the
config-only path handles NemotronH on every supported container.

#### 2. `candidate_filter` for `mcore_minitron`
(`modelopt/torch/prune/plugins/mcore_minitron.py`)

DeepSeek-style MoE configs have no explicit shared-expert-size field:
they size the shared expert as `n_shared_experts *
moe_intermediate_size`, where `moe_intermediate_size` is the (also
prunable) **routed** expert size. So only candidates with
`moe_shared_expert_intermediate_size % moe_ffn_hidden_size == 0` can be
written back to HF at all.

Candidates come from a Cartesian `product()` of independent per-hparam
choice lists, so no per-hparam restriction can express a constraint
*between* two hparams. New optional `candidate_filter` search-config key
(default `None`, so existing behaviour is unchanged): a callable that
rejects candidate configs before the metric computation, making the
search cheaper rather than more expensive. It receives **every**
supported hparam, with non-searched ones filled in from the model
config, so a filter still works when one of its hparams was skipped or
had a single choice.

Rejected candidates are not cached, so — like `score_func`, whose cached
scores are reused without re-validation — the filter is assumed
unchanged when resuming from a `checkpoint`.

`prune_minitron.py` wires this up for DeepSeek-style configs, so
**both** `moe_ffn_hidden_size` and `moe_shared_expert_intermediate_size`
stay prunable (the search then only picks shared sizes that are a
multiple of the routed one). A `--prune_export_config` that violates the
constraint never reaches the filter, so the export path now raises
`ValueError` instead of writing a checkpoint whose config disagrees with
its weights.

#### 3. MLA Q-LoRA pruning
(`modelopt/torch/prune/plugins/mcore_minitron.py`)

Pruning any MLA model with `q_lora_rank` set died during calibration
with `AttributeError: 'tuple' object has no attribute 'view'`.

`hidden_size` importance estimation blanket-patches every
`TELayerNormColumnParallelLinear` with `return_layernorm_output=True` to
capture post-layernorm activations. When `q_lora_rank` is set, MCore
builds `linear_q_up_proj` as a `TELayerNormColumnParallelLinear` — the
Q-LoRA layernorm is fused into it, which is why `q_layernorm` is
`IdentityOp` — so it was patched too, even though its layernorm is over
the **latent rank**, not `hidden_size`. TE then returns `((out, ln_out),
bias)` and MCore's `q, _ = self.linear_q_up_proj(...)` leaves `q` a
tuple.

Isolated by probing the module before and after dynamic conversion:

| Setup | `linear_q_up_proj` returns | Forward |
| --- | --- | --- |
| Before conversion | `tuple(Tensor, NoneType)` | — |
| After conversion, no hooks | `tuple(Tensor, NoneType)` | OK |
| After conversion **+ importance hooks** | `tuple(tuple(Tensor,
Tensor), NoneType)` | AttributeError |

So conversion is innocent; registering the importance hooks is the
trigger. Fix: exclude MLA's Q/KV up-projections from both the patch and
unpatch loops. `test_mcore_mla_pruning` did not catch this because it
builds MLA without `q_lora_rank`, where MCore uses a plain
`linear_q_proj` and nothing is patched.

#### 4. Fail fast instead of stalling on a rank-local error under PP
(`modelopt/torch/utils/distributed.py`)

A rank raising inside a distributed entrypoint left the whole job
stalled until the process group timed out, with **no diagnostic output
at all**: the failing rank blocked in `cleanup()`'s barrier while its
peers blocked in `recv_from_prev_pipeline_rank_`, and Python only prints
a traceback once the enclosing `finally` returns. A crash on one rank
was indistinguishable from a slow job.

- `dist.cleanup()` skips the barrier when unwinding from an exception.
- New `dist.abort()` prints the traceback, flushes and exits
immediately. Skipping the barrier alone is **not** enough — a stack dump
showed the failing rank then blocking in `destroy_process_group` for the
same reason — so the error path must not tear the process group down at
all. `SystemExit` is re-raised rather than aborted, so an intentional
exit (e.g. the `--score_lower_bound` gate) keeps its exit code and
prints no traceback. Kept out of `cleanup()` so no library caller gets a
surprise process exit.
- Called from the entrypoints that wrap `main()` in `try/finally`: the
five `examples/megatron_bridge` scripts.

Measured on a 2-GPU PP run whose rank 0 raises during calibration: **10
min timeout kill with no visible error → 31s, exit 1, real traceback.**
This is a latent, pre-existing issue (the `try/finally` predates this
PR); it only surfaces on a failing PP run, which is why CI never hit it.

### Usage

```bash
torchrun --nproc_per_node 4 examples/megatron_bridge/prune_minitron.py \
    --hf_model_name_or_path nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 \
    --pp_size 4 \
    --prune_target_active_params 3e9 \
    --output_hf_path /path/to/Nemotron-3.5-Lightning-30B-A3B-Pruned-A3.0B
```

### Testing

- **End-to-end on nemo:26.08.rc6** (4× GB300, transformers 5.12.1,
Megatron-Bridge with config-only export): pruning + export complete
(`EXIT=0`, "Saved pruned model … Done!"). The exported checkpoint has
canonical **plural** `backbone.embeddings.weight` keys, **0 MTP
tensors**, and a config that reloads correctly (`num_hidden_layers=52`
from `layers_block_type`, `n_shared_experts=1`,
`num_nextn_predict_layers=0`, pruned `hidden_size`/`mamba_*`/MoE dims,
reconstructed `hybrid_override_pattern`).

  <details>
<summary>Pruning search log (<code>--prune_target_active_params
3e9</code>)</summary>

  ```text
Top 10 Candidates with Scores

┏━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━┓
┃ # ┃ export_config ┃ active_params ┃ params ┃ score ┃

┡━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━┩
│ 1 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56,
'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 23.49B │ 0.5406
│
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size':
3584} │ │ │ │
│ 2 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 56,
'mamba_head_dim': 48, 'num_moe_experts': 96, │ 3.00B │ 20.09B │ 0.2427 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 3 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48,
'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 21.61B │ 0.2643
│
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 4 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48,
'mamba_head_dim': 64, 'num_moe_experts': 96, │ 3.00B │ 19.28B │ 0.4552 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size':
3712} │ │ │ │
│ 5 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64,
'mamba_head_dim': 64, 'num_moe_experts': 104, │ 3.00B │ 22.28B │ 0.5860
│
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 6 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48,
'mamba_head_dim': 48, 'num_moe_experts': 96, │ 3.00B │ 21.99B │ 0.2294 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size':
3328} │ │ │ │
│ 7 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 56,
'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 23.68B │ 0.5231
│
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 8 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56,
'mamba_head_dim': 56, 'num_moe_experts': 96, │ 3.00B │ 21.81B │ 0.5042 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size':
3584} │ │ │ │
│ 9 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48,
'mamba_head_dim': 56, 'num_moe_experts': 96, │ 3.00B │ 20.09B │ 0.2462 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 10 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64,
'mamba_head_dim': 64, 'num_moe_experts': 96, │ 3.00B │ 20.70B │ 0.5685 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │

└────┴───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┴───────────────┴────────┴────────┘


╭────────────────────────────────────────────────────────────────────────
Best Subnet
─────────────────────────────────────────────────────────────────────────╮
│ export_config {'num_layers': 52, 'hidden_size': 2304,
'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 104,
'moe_ffn_hidden_size': 1856, │
│ 'moe_shared_expert_intermediate_size': 3072} │
│ active_params 3.00B │
│ params 22.28B │
│ score 0.5860 │

╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

╭────────────────────────────────────────────────────── Pruned Model
Stats ───────────────────────────────────────────────────────╮
│ Total Parameters 22.28B │
│ Active Parameters 3.00B │
│ Memory (BF16, seq_length=8192, batch_size=8) weights: 42489.7 MB,
kv_cache: 384.0 MB, mamba_state: 190.5 MB, Total: 43064.2 MB │

╰─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
  ```

  </details>

- **`tests/examples/megatron_bridge/test_prune_minitron.py`** —
`nemotron_h` now exports to HF and reloads (previously it stopped at a
Megatron checkpoint, since the dummy-model path needed
`transformers<5`), plus an `n_shared_experts` config assertion; the dead
`megatron_format` branch is gone. It runs on the CI container:
**verified on nemo:26.06.01 (transformers 5.8.1) and nemo:26.08.rc6.**
-
**`tests/gpu_megatron/torch/prune/plugins/test_mcore_mamba_minitron_pruning.py`**
— the `nas_memory_mb` search test now passes a `candidate_filter` and
asserts the exact number of rejected candidates (256 of the 512-combo
grid) plus the surviving candidates' validity; its `expected_top_k`
goldens are regenerated accordingly. Because
`moe_shared_expert_intermediate_size` is in that test's skip list, this
also covers the model-config fallback for hparams that are not in the
search space.

Verified on 2 GPUs, on both the CI container (nemo:26.06.01) and
nemo:26.08.rc6:

| Test | Result |
| --- | --- |
| `test_prune_minitron[qwen3]` | PASSED on 26.06.01 and 26.08.rc6 |
| `test_prune_minitron[deepseek_v3]` | PASSED (52s) — MLA Q-LoRA +
`candidate_filter` end-to-end |
| `test_prune_minitron[nemotron_h]` | PASSED on 26.06.01 (58s) and
26.08.rc6 (61s) |
| `test_mcore_mamba_hybrid_pruning_nas_memory_mb` | PASSED |
| `test_mcore_mamba_hybrid_pruning_nas_params` | PASSED (unchanged
sibling, run to check the regenerated goldens did not disturb it) |
| 2-GPU PP run failing on rank 0 | fails in 31s with a real traceback
(was a 10 min stall) |


### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `candidate_filter` defaults
to `None` (existing searches unchanged), and the config-only export is
limited to hybrid providers on nemo:26.08+, so dense / MoE / VLM exports
keep the path they use today.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ✅

### Additional Information

Enables the Prune + Distill workflow for
`NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16` (native, no-remote-code
`NemotronHForCausalLM` with MTP heads). Pruning-time MTP support was
scoped and intentionally deferred — MTP heads are dropped and can be
re-derived via a short SFT with `mtp_num_layers=1` on the
pruned+distilled model.

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 21:05:56 +05:30
sugunav14andClaude Opus 5 a21173a4d5 Bug fix: 6542481 (#2064)
### What does this PR do?

Type of change: Bug fix

Fixes `AssertionError: Model already has modelopt state!` when exporting
a QLoRA checkpoint
(NVBug 6542481). The QLoRA output is adapter-only, so `from_pretrained`
resolves the quantized
base model and already restores the ModelOpt state; `export.py` then
restored a second time.

Fixing that exposed two more breakages on the same path, also fixed
here:

- `_restore_qtensor_wrappers` missed every module — PEFT renames the
compressed linears to
`<name>.base_layer`, so no weight got re-wrapped and the packed NVFP4
weight hit a shape error.
- `postprocess_state_dict` dropped `weight_scale_2` (missing from the
QLoRA rename map), leaving
  the exported checkpoint impossible to dequantize.

### Usage

No API change — `examples/llm_qat/export.py --pyt_ckpt_path <qlora_ckpt>
--export_path <out>`
now completes on the documented quantize → train → export flow.

### Testing

Reproduced in the reported environment (TRT-LLM 1.3.0rc22, transformers
5.5.4, NVFP4).

- Added the missing export step to `test_qwen3_qlora_nvfp4` and a unit
test for the QLoRA
  `base_layer` rename; both fail without the fix.
- Exported base model is byte-identical to a plain PTQ export;
dequantized NVFP4 weights match
  the bf16 original (worst rel. error 0.10).
- No regressions: `tests/gpu/torch/export/test_export.py` (49 passed),
save/load plugin tests.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ❌ — can add if wanted
- Did you get Claude approval on this PR?: ❌ — not run yet

### Additional Information

Fixes NVBug 6542481.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved QLoRA checkpoint export and restoration across supported
model configurations.
* Preserved secondary weight-scale information and other deployment
tensors in exported checkpoints.
* Corrected handling of quantized base-layer weights after adapter
reparenting.
* Prevented duplicate state restoration when checkpoints already include
the required model state.
* Removed internal adapter prefixes and quantizer details from exported
state data.

* **Tests**
* Added validation for packed weights, quantization metadata, required
scales, and removal of embedded adapter layers.
* Added regression coverage for QLoRA state processing and quantized
weight restoration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 11:51:02 -07:00
Chenjie LuoandClaude Opus 5 9220fac053 [NVBug: 6563509] Drop Phi-3-vision / Phi-4-multimodal PTQ support (#2115)
### What does this PR do?

Type of change: Deprecation

Resolves [NVBug 6563509](https://nvbugspro.nvidia.com/bug/6563509),
where
`hf_ptq.py` on Phi-4-multimodal-instruct died with
`RuntimeError: Tensor.item() cannot be called on meta tensors`.

The crash is real but not fixable on our side, and it is not the reason
the model
is unusable. Phi-4-multimodal's bundled remote code predates
Transformers v5 and
does not load on **any** version in our supported range
(`transformers>=4.57,<5.15`):

| Blocker | Where |
|---|---|
| `peft.get_peft_model` reads `prepare_inputs_for_generation`, gone
since transformers 4.52 dropped `GenerationMixin` from `PreTrainedModel`
| `modeling_phi4mm.py:1959` |
| `_tied_weights_keys` declared as a list; Transformers 5.x calls
`.keys()` on it in `post_init` | `modeling_phi4mm.py:1937` |
| `int(torch.tensor(...))` in `__init__`, which cannot run on a meta
device — the reported crash | `speech_conformer_encoder.py:1435` |

The model card pins `transformers==4.48.2` / `peft==0.13.2`, so there is
no
overlap with our floor and nothing on our side can bridge it. The model
is
therefore dropped rather than worked around.

**Phi-3-vision is dropped alongside it because it is the older,
superseded model
in the same family** — with its successor unsupportable there is no
reason to
keep carrying the predecessor. This is a product-scope call, not a
separate
compatibility finding: Phi-3-vision shares the list-valued
`_tied_weights_keys`
defect (`modeling_phi3_v.py:1214`) and so is likewise broken on
Transformers 5.x,
but it does **not** hit the `peft` blocker, and it was not re-verified
on 4.57.
Per the 0.46 changelog we have already bumped the floor to 4.57 and
noted that
"Transformers 4.x support will be dropped in a future release", so any
remaining
window closes on its own. Same reasoning already applied to VILA / NVILA
in this
release.

**Removed**

- the support-matrix row in `examples/hf_ptq/README.md`
- `"Phi4MMForCausalLM": "phi4mm"` from `MODEL_NAME_TO_TYPE`
- the multimodal-detection heuristics that only ever matched these two —
`vision_lora`, `audio_processor`, `embd_layer.image_embd_layer`, and the
  `phi4mm` model-type check — in both `is_multimodal_model` and
  `_is_multimodal_config`
- the `Phi3Image` / `PhiImage` exclusions in `is_embedding`
- the phi4mm input-mode warning in `hf_ptq.py`
- `modelopt_recipes/huggingface/phi4mm/` and its references in
  `modelopt_recipes/ptq.md`

**Not changed:** the device-map sizing path (meta-device skeleton,
`infer_auto_device_map`, and the `--gpu_max_mem_percentage` cap) keeps
its
original behavior. That cap is wanted exactly where it already fires —
when the
model is already offloading to CPU, where it costs little and the
headroom is
required. With the affected checkpoints removed, there is no supported
model
that trips the meta-device build, so there is nothing to work around
here.

Text-only **Phi-3/Phi-4** and **Phi-3.5-MoE** are natively supported by
transformers and are untouched.

### Testing

On H200, `nvcr.io/nvidia/tensorrt-llm/release` (torch 2.12, transformers
5.5.4),
against the real checkpoint:

- **Version matrix** (vanilla transformers, no modelopt) — Phi-4-MM
loads at
4.48.2 / 4.49.0 / 4.50.0 / 4.51.3 and fails at 4.53.3 / 4.56.2 / 4.57.1
  (`AttributeError: 'Phi4MMModel' object has no attribute
'prepare_inputs_for_generation'`) and at 5.5.4 (meta-init, then
tied-keys).
  This is what establishes that no supported version works.
- `tests/examples/hf_ptq/test_example_utils.py` — 28 passed.
- **Sweep**: `tests/examples/hf_ptq` + `tests/unit/torch/export` —
failure set
identical to the pre-change tree (GPU/model-dependent `test_vlm_ptq`,
plus
`test_quant_aware_conversion` scoped-mapping tests), so none are
introduced
  here.
- `pre-commit` clean on all changed files, including recipe validation.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — PTQ for Phi-3-vision and
  Phi-4-multimodal is removed, along with the `huggingface/phi4mm/ptq/*`
recipes. Phi-4-multimodal is already unloadable on every supported
transformers
version, so no working workflow regresses; Phi-3-vision is a deliberate
scope
  removal as its superseded predecessor.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — this is a deletion; the
existing
`test_get_model_*` / `test_resolve_init_config_*` tests are unchanged
and still
  pass.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Two related references were left in place deliberately; say the word and
I'll
fold them in:

- `tests/examples/hf_ptq/test_deploy.py` still deploys the
already-published
`nvidia/Phi-4-multimodal-instruct-{NVFP4,FP8}` checkpoints. Those
artifacts
exist and serve fine; this PR only removes the ability to *produce*
them.
- `examples/torch_onnx/README.md` still lists Phi-4-multimodal-instruct.
That
is a separate ONNX pipeline that does not go through `get_model()` and
was not
  tested here.

Earlier revisions of this branch also reworked the device-map sizing so
the
meta-tensor crash could not occur. That was reverted in 701180ed6: the
guard is
correct as written, and every alternative either changed behavior for
models that
fit today or moved the guard somewhere it does not belong, for a crash
that only
ever affected the checkpoints this PR removes.

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 23:22:23 +00:00
Chenjie LuoandClaude Opus 5 9b8caf623a Use lm-eval 0.4.12's built-in trtllm backend, deprecate lm_eval_tensorrt_llm.py (#2066)
### What does this PR do?

Type of change: documentation / example update (with a behaviour fix)

lm-evaluation-harness **0.4.12** is the first release that ships a
TensorRT-LLM backend
(`lm_eval.models.trtllm_causallms`, registered as `trtllm`) — it is
absent in 0.4.10 and
0.4.11. This example no longer maintains its own, so:

- Pin `lm_eval[api,ifeval]>=0.4.12,<0.5` (the 0.5.0.dev line drops the
file) and bump
  `lm_eval_hf.py`'s version guard to match.
- **Delete** `examples/llm_eval/lm_eval_tensorrt_llm.py` (the `trt-llm`
model). Replace
`python lm_eval_tensorrt_llm.py --model trt-llm --model_args
tokenizer=<tok>,checkpoint_dir=<ckpt>`
with `python lm_eval_trtllm.py --model trtllm --model_args
model=<ckpt>,tokenizer=<tok>`.
- Add `examples/llm_eval/lm_eval_trtllm.py`, whose entire content is one
corrected
`_parse_logprobs` plus `cli_evaluate()` (see below). `lm_eval_hf.py`
stays HF-only.
- `examples/hf_ptq/scripts/huggingface_example.sh` and the docs use the
upstream backend.
`parser.sh` gains `--input` (`BUILD_MAX_INPUT_LEN`, default 4096) — it
already *echoed*
that variable but never parsed or defaulted it, so it printed empty on
every run.

#### Why `lm_eval_trtllm.py` exists: an upstream off-by-one

TensorRT-LLM aligns `prompt_logprobs` to the *next* token.
`executor/base_worker.py`:

```python
# Pass prompt_token_ids with an offset of 1 for correct mapping to the context logits
prompt_token_ids = generation_result._generation_request.prompt_token_ids[1:] + first_generation_token
```

So entry `i` is the distribution that predicted `tokens[i + 1]`, and
`_topk_logprobs`
appends that token's id when it is not in the top-k. lm-eval's
`_parse_logprobs` instead
reads `prompt_logprobs[i][tokens[i]]` and applies its own shift on top,
which raises
`KeyError` on the **first request of every loglikelihood task**
(hellaswag, mmlu, arc, ...):

```
File ".../lm_eval/models/trtllm_causallms.py", line 324, in _parse_logprobs
    current_token_logprob = prompt_logprob[tokens[i]]
KeyError: 6503
```

Probed against TRT-LLM 1.3.0rc23 with a 14-token prompt for
`prompt_logprobs` 0, 1 and 2:
`tokens[i]` is missing at **every** position, `tokens[i+1]` is present
at every position.
Only `generate_until` tasks work unpatched. **This wants an upstream
issue against
EleutherAI/lm-evaluation-harness.**

The override also fails loudly rather than quietly: it checks
`prompt_logprobs` covers
every prompt token and raises on a missing token, instead of skipping
the term and
silently inflating the reported accuracy.

#### Defaults that must be set explicitly

`TRTLLM.__init__` accepts `**kwargs` but forwards only a fixed set to
the **TensorRT-LLM
`LLM` API**, so extra `--model_args` aimed at the engine are silently
dropped. (lm-eval's
own named parameters — `max_gen_toks`, `batch_size`, `truncation_side`,
... — are honored
normally.) Two engine defaults are unsafe for few-shot eval:

- `tensor_parallel_size` defaults to **1** (the deleted wrapper used
every visible GPU).
- `max_input_len` defaults to **2048**, and longer prompts are silently
left-truncated —
  5-shot MMLU/gsm8k prompts exceed that.

### Usage

```bash
python lm_eval_trtllm.py --model trtllm \
    --model_args model=<quantized checkpoint dir>,tokenizer=<HF model folder>,tensor_parallel_size=<tp>,max_batch_size=<bs>,max_input_len=4096,max_output_len=512 \
    --tasks hellaswag,gsm8k \
    --batch_size <bs>
```

Flat arguments (no `run` subcommand) are what 0.4.12's
`HarnessCLI.parse_args` inserts
`run` for automatically (`_cli/harness.py:48-51`); this is the exact
command form used for
the results below.

### Testing

**Unit** — `tests/examples/llm_eval/test_lm_eval_trtllm.py`, no GPU and
no `tensorrt_llm`
install: stubs the response object and pins the `i-1` alignment, the
`rank != 1` →
`is_greedy` rule, the `ctxlen=0` edge, and both `RuntimeError` paths.
Mutation-checked —
dropping the `-1` shift is caught by 5/5 cases, ignoring `ctxlen` by
4/5. A sixth test is a
**tripwire**: it asserts lm-eval's own implementation is still
misaligned, so a future
0.4.x that fixes the bug fails the test and says to delete this file
rather than being
silently re-broken by the override.

**End to end** — `nvidia/Qwen3.5-122B-A10B-NVFP4` (NVFP4 MoE, 256
experts) on **4x B300**,
TRT-LLM 1.3.0rc23, lm-eval 0.4.12, `--limit 32`:

| run | hellaswag acc | hellaswag acc_norm | gsm8k flexible | gsm8k
strict |
|---|---|---|---|---|
| deleted impl (`trt-llm`), tp=4 | 0.7188 | 0.7812 | 0.8438 | 0.7812 |
| `lm_eval_trtllm.py`, tp=1 | 0.7188 | 0.7812 | 0.8438 | 0.8125 |
| `lm_eval_trtllm.py`, tp=2 | 0.7188 | 0.7812 | 0.9062 | 0.8125 |
| `lm_eval_trtllm.py`, tp=4 | 0.7188 | 0.7812 | 0.8750 | 0.8438 |

- hellaswag (the loglikelihood path this PR fixes) is **identical at
every tp and identical
to the deleted implementation** — the alignment fix is exact, not
approximate.
- gsm8k varies by 1–2 samples out of 32 (generation path: upstream uses
native `stop=`
sequences and per-request `SamplingParams`; the old wrapper used
beam-search-of-1 with
  post-hoc string truncation).
- Without the override, every hellaswag run above dies with the
`KeyError`.
- Re-verified at tp=4 after the code moved out of `lm_eval_hf.py` into
`lm_eval_trtllm.py`.

Note: NVFP4 fused-MoE has no CUTLASS tactic on Hopper (`No supported MoE
GEMM tactic
remains after replacing unsupported NO_SMEM epilogues.`), so this had to
be validated on
Blackwell.

### Feature parity notes

Gained from upstream: `loglikelihood_rolling` (was
`NotImplementedError`), pipeline
parallelism, `add_bos_token` auto-detection, prompt truncation,
per-request sampling params,
`prompt_logprobs` instead of full-vocab context logits (much lower
memory), thinking-tag
handling, `batch_size=auto`.

Not reachable through the upstream backend (were set by
`modelopt.deploy.llm.LLM`):
`enable_attention_dp` for MoE, `CudaGraphConfig`,
`enable_chunked_prefill`,
`moe_expert_parallel_size=1`, and `free_gpu_memory_fraction=0.7` with a
capped
`kv_cache.max_tokens` — upstream uses the TRT-LLM default 0.9 (observed
allocating 218 GiB
of paged KV cache on B300), so OOM risk is higher on smaller GPUs. This
is documented in
`examples/llm_eval/README.md`, and `huggingface_example.sh` honours a
preset `LM_EVAL_TP`
so users can lower the tensor-parallel size without editing the script.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — `lm_eval_tensorrt_llm.py` is
removed and the CLI changes (`--model trt-llm` → `trtllm`,
`checkpoint_dir=` → `model=`). Migration command is in the README and
CHANGELOG.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependency; existing `lm_eval` pin tightened.
- Did you write any new necessary tests?: ✅ —
`tests/examples/llm_eval/test_lm_eval_trtllm.py` (6 cases, no GPU).
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.47 *Deprecations*.
- Did you get Claude approval on this PR?: ✅ — reviewed, feedback
addressed in `dcedd37b4` and `622b97c26`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added TensorRT-LLM evaluation through lm-evaluation-harness’s `trtllm`
backend.
* Added configurable input/output lengths, batching, tensor parallelism,
and build input length.
* Improved prompt log-probability alignment for more accurate evaluation
results.

* **Documentation**
* Updated evaluation instructions, truncation guidance, backend
limitations, and configuration examples.

* **Deprecations**
  * Removed the legacy TensorRT-LLM evaluation script and entry point.

* **Updates**
  * lm-evaluation-harness now requires versions 0.4.12 through 0.4.x.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 22:36:55 +00:00
kinjalpatel27 bd3798a794 [6562078]: fix calibration for vLLM 0.26.0 (#2093)
### What does this PR do?

Type of change: Bug fix

Fixes calibration failures when running fake-quantization with vLLM
0.26.0.
Two root causes are addressed:

1. **`finish_requests` must be called explicitly after `add_requests`.**
In vLLM 0.26.0 the scheduler calls `finish_requests` *before*
`add_requests`
inside `execute_model`, so request IDs are not registered yet and
cleanup
never runs. The calibration loop now calls `finish_requests` directly
after
each batch using `dataclasses.replace`. Wrapped in `try/finally` with an
inner
`try/except` so it always runs and never masks the original exception. A
warning
is emitted when `finish_requests` is absent so the regression is
self-diagnosing
   on future vLLM API changes.

2. **`NewRequestData` gained a `prefill_token_ids` field.**
vLLM 0.26.0 added this required argument; the calibration helper now
passes it.

Additional:
- Dockerfile updated to vLLM 0.26.0 with `USER vllm` (non-root).
- README updated to include vLLM 0.26.0 in tested versions.

### Usage

```bash
cd examples/vllm_serve
QUANT_CFG=FP8_DEFAULT_CFG QUANT_CALIB_SIZE=8 CALIB_BATCH_SIZE=1 \
  python3 vllm_serve_fakequant.py Qwen/Qwen1.5-MoE-A2.7B-Chat -tp 1 \
  --host 0.0.0.0 --port 8000
```

### Testing

Tested end-to-end FQ calibration with vLLM 0.26.0 using the Docker image
built
from `examples/vllm_serve/Dockerfile`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (examples change only)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

Changes are confined to `examples/vllm_serve/` and do not affect the
core library.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Summary by CodeRabbit

* **New Features**
  * Added compatibility with vLLM 0.26.0 in the serving example.
* Improved calibration request handling and cleanup after model
execution.
  * Enabled the serving container to run with a non-root user.

* **Bug Fixes**
* Calibration cleanup failures no longer obscure the original model
execution error.
  * Added warnings when calibration cleanup cannot be completed.

* **Documentation**
* Updated the serving example documentation to list vLLM 0.26.0 among
tested versions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-08-07 14:53:11 -07:00
Ajinkya RasaneandCodex ccf44ea67a [OMNIML-5562] Add FAR3D ONNX PTQ and accuracy evaluation example (#2012)
### What does this PR do?

Type of change: new example

Adds an end-to-end FAR3D ONNX PTQ example under
`examples/onnx_ptq/far3d`. The example prepares Argoverse 2 validation
metadata and calibration batches, quantizes the image encoder to INT8,
builds TensorRT engines, and evaluates 3D object detection accuracy.

It also provides a reproducible FAR3D runtime image, preserves
accuracy-sensitive encoder layers in high precision, and supports
temporal decoder state during evaluation.

### Usage

```bash
python prepare_metadata.py /path/to/av2
python prepare_calibration.py /path/to/far3d.py far3d_calibration
python quantize.py far3d.encoder.onnx far3d_calibration
python evaluate.py /path/to/far3d.py \
  far3d.encoder.int8.engine far3d.decoder.fp16.engine
```

### Testing

- Ran all configured pre-commit hooks on the changed files.
- Ran synthetic calibration-reader and graph-exclusion tests.
- Built the documented FAR3D runtime image and verified TensorRT 10.11,
Argoverse 2 imports, and TensorRT engine execution.
- Ran the complete workflow on the Argoverse 2 validation split using
500 calibration batches and 23,522 evaluation frames. The INT8 encoder
and FP16 decoder produced 0.238 mAP.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information

- [NVIDIA DL4AGX FAR3D TensorRT
reference](https://github.com/NVIDIA/DL4AGX/tree/master/AV-Solutions/far3d-trt)

> 🤖 _Generated by Codex (AI agent)._


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an end-to-end FAR3D ONNX post-training quantization workflow for
Argoverse 2, including calibration batch generation, INT8/FP8
quantization, TensorRT engine inference, and mAP evaluation.
* Added a dedicated FAR3D example Docker environment plus detailed
README instructions.
* Added FAR3D evaluation, calibration preparation, metadata preparation,
and quantization scripts.
* **Bug Fixes**
* Improved handling when flash-attention is unavailable, with clearer
error messaging.
* **Chores**
* Updated pre-commit exclusions and refreshed third-party license
attribution.
  * Added an Experimental changelog entry for the FAR3D example.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
2026-08-07 05:43:42 +00:00
Keval MorabiaandClaude Opus 5 22b6a148b0 Fix EAGLE-3 context-parallel training and re-enable its tests (#2086)
### What does this PR do?

Type of change: Bug fix

**EAGLE-3 context-parallel training (`--cp_size > 1`) is fixed, and its
tests run again.** CP has been broken since `accelerate` 1.13, and the
tests never caught it: the guard compared `Version("2.10.0a0")` against
`Version("2.10.0")`, which is False on every NGC alpha torch build, so
`test_llama_eagle3[cp_size=2]` has never actually run in CI.

Five fixes:

- **`main.py`** — rebuild the FSDP2 plugin accelerate requires for
`cp_size > 1`. The `--fsdp full_shard --fsdp_config` launcher flags that
used to supply it were dropped from `launch_train.sh`, so CP could not
start at all. Also pass the CP degree to the draft model.
- **`modeling_eagle.py`** — apply the draft model's first input norm
inside `layers[0]`'s own forward, where FSDP2 has actually unsharded its
weights, and only stash the input embeds on the path whose pre-hook
consumes them.
- **`hf_eagle.py`** — skip the dense eagle attention mask under CP
(causal masking comes from `is_causal`, TTT masking from the
ring-attention patch), and warn that padded positions are therefore
unmasked. Also stop `(eagle_loss or 0)` replacing a `0.0` loss tensor
with a plain `int`, which detached the graph.
- **`eagle_utils.py`** — key TTT-mask injection off the backward call's
`grad_out` kwarg, since newer torch omits `attn_bias` on the forward
call, silently disabling TTT masking.
- **`utils.py`** — CUDNN-only SDPA under CP; the `MATH` backend
decomposes SDPA and breaks on DTensors. Scoped to `cp_size > 1`, since
this context manager wraps every training forward and CPU has no cudnn
backend.

**Drops the `speculative_decoding` 26.01 container override.** It was
added when the lane ran 25.06 and spec-dec needed something *newer* — a
floor. Later bumps moved the default past it, so it had silently become
a ceiling holding spec-dec on a 6-month-old image.

### Testing

Ran `tests/examples/speculative_decoding` in
`nvcr.io/nvidia/pytorch:26.07-py3` on 2 GPUs, reproducing the CI install
steps (`pip uninstall -y nvidia-modelopt`, `pip install -e
".[hf,dev-test]"`, example requirements): **16 passed, 2 skipped** — the
2 skipped being pre-existing `--run-manual` tests. All four
`test_llama_eagle3` cases pass, including both `cp_size=2` ones.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — the existing `cp_size=2`
tests are re-enabled
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — not yet run

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 21:01:02 +05:30
Frida HouandClaude Sonnet 4.6 6a8102591e Single gpu disk offload PTQ for DSR1/Ultra (#2008)
### What does this PR do?

Type of change: New feature, bug fix, new tests

Enables single-GPU PTQ for models too large to fit in VRAM (e.g.
Nemotron-Ultra-550B at 1.1 TB BF16, DeepSeek-R1 at 642 GB BF16) by
adding accelerate disk/CPU offload support to the HF PTQ example and
fixing the export path to correctly handle offloaded models.

**G1 Offload-aware unified HF export
(`modelopt/torch/export/unified_export_hf.py`)**

The existing `_export_transformers_checkpoint` removed accelerate hooks
before materializing weights, silently writing meta tensors (empty
weights) to the checkpoint. Fix:

- `_has_accelerate_offload(model)`detects any disk/CPU-offload
accelerate hook in the model tree.
- `_process_quantized_modules_offloaded(model, dtype)` new export path
for offloaded models: materializes one decoder layer at a time via
`enable_weight_access_and_writeback`, dispatches export handlers inside
the context window, and snapshots the layer state dict before hooks
re-offload the weights. A second pass collects non-decoder modules that
are also disk-offloaded (embed, norm, lm_head) to avoid meta tensors in
the returned state dict. Hooks are removed only after the full state
dict is assembled.
- Meta-tensor guard in `_export_quantized_weight` raises `RuntimeError`
on meta input instead of silently corrupting the checkpoint.

**G2 Disk-offload CLI (`examples/hf_ptq/hf_ptq.py`,
`example_utils.py`)**

Three new arguments to `hf_ptq.py`:
- `--offload_folder PATH`  enable accelerate disk offload; shards spill
here.
- `--max_gpu_memory_gb N` VRAM budget for the accelerate device map.
- `--max_cpu_memory_gb N`  CPU RAM budget for the accelerate device map.

Validation: `--offload_folder` is incompatible with `--low_memory_mode`
and `--use_seq_device_map`.


**G3 Streaming shard writer for 80 GB CPU RAM
(`modelopt/torch/export/unified_export_hf.py`)**

The G1 path accumulated the entire quantized state dict in CPU RAM
before writing
(~764 GiB for Ultra 550B), blocking the 80 GB target.

New streaming path writes shard files layer-by-layer. Peak memory = 1
decoder layer +
1 shard buffer instead of the full checkpoint:

| Model | Old peak CPU RAM | New peak CPU RAM |
|-------|-----------------|-----------------|
| Ultra NemotronH 550B | ~764 GiB | ~57 GB |
| DeepSeek-R1 | ~630 GiB | ~55 GB |

Key pieces:

- `_StreamingShardWriter(export_dir, max_shard_size)` buffers tensors up
to `max_shard_size` bytes, flushes to numbered temp files
(`__shard_part_NNNNN.safetensors`), renames to canonical shard names at
`finalize()`, writes
`model.safetensors.index.json`. Single-shard exports produce
`model.safetensors` with no index file.
- `_postprocess_single_tensor(key, value, ...)` per-tensor extraction of
`postprocess_state_dict` logic (KV amax scale, skip/rename, squeeze) for
streaming use.
- `_parse_shard_size(size)` converts `"10GB"` / `"500MB"` strings to
bytes.
- `_export_transformers_checkpoint_streaming(model, dtype, export_dir,
max_shard_size)` streams decoder layers via
`enable_weight_access_and_writeback`, applies per-tensor postprocessing
+ name reversal + tied-alias filter, writes shard files directly.
Non-decoder offloaded modules and GPU-resident tensors are handled in
separate passes.
- `export_hf_checkpoint` dispatches to the streaming path when
`_has_accelerate_offload(model)` is true; `hf_quant_config.json`,
quant-config name reversal, and `config.json` update are shared between
paths.

`export_hf_checkpoint` accepts a new `max_shard_size` parameter (default
`"10GB"`) that controls the shard size for both paths.


**Supporting changes**

- `modelopt/torch/quantization/plugins/huggingface.py` 
`get_nemotron_h_decoder_layers` now checks both `model.backbone.layers`
(remote-code variant) and `model.model.layers` (native HF variant),
fixing layer discovery for NemotronH when loaded without
`trust_remote_code`.
-
`modelopt_recipes/general/ptq/nvfp4_experts_only-kv_fp8_layerwise_offload.yaml` 
new recipe combining NVFP4 W4A4 on MoE experts, FP8 KV cache, and
layerwise calibration with `calib_mutates_weights: false` (required for
disk-offload compatibility).
- `example_utils.py`  `_FP8BF16Fallback` shim: dequantizes block-scaled
FP8 expert weights to BF16 for calibration forward passes when the
`kernels` package is unavailable (e.g. DSR1 on nodes without finegrained
FP8 kernel support).

### Usage

```python
# Single-GPU PTQ for a model too large to fit in VRAM, using disk offload
python examples/hf_ptq/hf_ptq.py \
    --pyt_ckpt_path /path/to/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 \
    --recipe general/ptq/nvfp4_experts_only-kv_fp8_layerwise_offload \
    --export_path /path/to/output \
    --offload_folder /path/to/offload \
    --max_gpu_memory_gb 170 \
    --max_cpu_memory_gb 500 \
    --trust_remote_code \
    --calib_size 8 --batch_size 1 --skip_generate
```
### Testing

**Unit tests** (`tests/unit/torch/export/test_offload_export.py`, 7
tests, CPU-only):
- `_has_accelerate_offload` detection (true/false/nested-module cases)
- `_export_quantized_weight` meta-tensor guard (raises on meta, passes
on real)
- `_process_quantized_modules_offloaded` with disk-offloaded embed +
GPU-resident decoder layer: verifies no meta tensor in returned state
dict

**GPU integration tests**
(`tests/gpu/torch/export/test_offload_export.py`, 2 tests):
- Tiny 2-layer LLaMA with CPU offload: FP8 quantization + export,
asserts no meta tensors and valid `hf_quant_config.json`
- Same with layerwise FP8 (`calib_mutates_weights=False`): disk-offload
path end-to-end

## End-to-end validation

Verified with DSR1 that the non-layerwise path provide identical
checkpoint before and after this change, also the layerwise with cpu
off-load path produce same identical checkpoint (with same max
calibration setting).

Two production-scale checkpoints were quantized end-to-end using the new
disk-offload PTQ path on a single GB200 GPU (189 GiB VRAM).

### DeepSeek-R1 (671B, MoE)

| | |
|---|---|
| **Checkpoint** | `DeepseekV3ForCausalLM`, 671B params, 61 decoder
layers |
| **Input size** | 642 GB BF16 |
| **Recipe** | `nvfp4_experts_only-kv_fp8_layerwise_offload` |
| **`--max_gpu_memory_gb`** | 80 |
| **`--max_cpu_memory_gb`** | 80 |
| **`--calib_size` / `--batch_size`** | 8 / 1 |
| **`--trust_remote_code`** | no (built-in transformers) |
| **Wall-clock** | 40 min 12 s (load ~14 min, calib ~12 min, export ~14
min) |
| **Peak GPU memory** | 88.9 GB |
| **Peak process RSS** | 376 GB |
| **Output** | 40 shards x ~10 GB = 403 GB (~37% compression) |

<img width="1783" height="2532" alt="image"
src="https://github.com/user-attachments/assets/fe515035-c880-453f-af1c-2d98395c9197"
/>


### NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 (550B, NemotronH MoE + Mamba)

| | |
|---|---|
| **Checkpoint** | `NemotronHForCausalLM`, ~550B params, 108 decoder
layers |
| **Input size** | ~1.1 TB BF16 |
| **Recipe** | `nvfp4_experts_only-kv_fp8_layerwise_offload` |
| **`--max_gpu_memory_gb` / `--max_cpu_memory_gb`** | 170 / 500 and **80
/ 80** |
| **`--calib_size` / `--batch_size`** | 8 / 1 |
| **`--trust_remote_code`** | yes (`NemotronHForCausalLM`) |

| | 170 GB GPU / 500 GB CPU | **80 GB GPU / 80 GB CPU** |
|---|---|---|
| **Wall-clock** | 41 min 13 s | **47 min 16 s** |
| **Peak GPU memory** | 165.8 GB  | **76.7 GB** |
| **Peak RSS (load)** | 789 GB transient | **345 GB transient** |
| **Steady-state RSS** | ~454-496 GB | **~50 GB** |
| **Output** | 34 shards x ~11 GB = 365 GB | 34 shards x ~11 GB = 365 GB
|

<img width="1783" height="2532" alt="image"
src="https://github.com/user-attachments/assets/da4771b8-b522-4549-8e40-7f975bd6f9b1"
/>


### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅ 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added disk/CPU/GPU memory-limited offload model loading.
* Added offload-aware streaming Hugging Face checkpoint export with
sharded output.
  * Added an NVFP4 expert-only PTQ recipe with FP8 KV-cache support.
  * Improved Nemotron-H model layout support.
* **Bug Fixes**
  * Improved DeepSeek bundled-code selection based on remote-code trust.
* Strengthened handling of meta/offloaded weights, tied-weight
deduplication, and export post-processing.
* **Tests**
  * Added coverage for offload exports and DeepSeek loading behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Fridah-nv <fridah@nvidia.com>
Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-05 14:44:29 -07:00
realAsma fed1980b29 Pin Accelerate below 1.14 for llm_qat (#2067)
### What does this PR do?

Type of change: Bug fix

Pins Accelerate below 1.14 for the `llm_qat` example. Accelerate 1.14
introduced an FSDP2 regression for models whose input embeddings and
output head share a parameter; the shared weight can be assigned to two
FSDP groups and training fails before the first step.

The pin is example-local. The project-wide Hugging Face dependency range
and unrelated examples remain unchanged.

### Usage

No usage change. Installing the `llm_qat` requirements now resolves
Accelerate to the existing supported range below 1.14.

### Testing

- Reproduced the duplicate shared-parameter FSDP2 failure with
Accelerate 1.14.0 on two ranks.
- Verified Accelerate 1.13.0 completes a distillation training step with
the otherwise-identical environment and a clean `main` source tree.
- Verified the combined requirements resolve to
`accelerate>=1.0.0,<1.14` and select 1.13.0.
- `pre-commit run --files examples/llm_qat/requirements.txt`
- `git diff --check`

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — dependency-only change
validated by an exact two-rank A/B run.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ — draft PR; automated
review is pending.

### Additional Information

This is a scoped compatibility pin while the upstream Accelerate
regression remains unresolved.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Constrained the Accelerate dependency to versions below 1.14 to
improve compatibility for the LLM question-answering example.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-08-05 09:19:22 +05:30
2d4be28818 fix(hf_ptq): use no_grad instead of inference_mode in export_quantized (NVBug 6537702) (#2047)
### What does this PR do?

Type of change: Bug fix

Fixes [NVBug 6537702](https://nvbugspro.nvidia.com/bug/6537702) /
[OMNIML-5658](https://jirasw.nvidia.com/browse/OMNIML-5658) — multi-node
FSDP2 PTQ export fails on all ranks:

```
hf_ptq.py:910 export_quantized -> export_hf_checkpoint
unified_export_hf.py:1446 _export_transformers_checkpoint -> get_model_state_dict
modelopt/torch/opt/_hooks.py:88 _get_model_state_dict_with_dm_check
torch/distributed/checkpoint/state_dict.py:481 _get_model_state_dict
torch/nn/modules/module.py:2160 _save_to_state_dict
    destination[prefix + name] = param if keep_vars else param.detach()
RuntimeError: Cannot set version_counter for inference tensor
```

**Root cause.** `export_quantized` wrapped its whole body in
`torch.inference_mode()`. On the FSDP2 path (`--use_fsdp2`),
`get_model_state_dict(full_state_dict=True)` gathers the full params
*inside* that context, so the gathered tensors are inference tensors.
Inference tensors have no version counter, so the subsequent
`state_dict()` → `param.detach()` raises.

**Fix.** Use `torch.no_grad()` for the export context. It still disables
autograd, but the gathered params stay normal tensors with an intact
version counter, so `detach()` works. FSDP2-only failure — the non-FSDP2
path never hit it because its params already exist outside the context.

The one-line fix is originally by @shengliangx (`b0e4328` on
`shengliangx/distributed-unified`); this PR retargets it to the
post-rename `examples/hf_ptq/` path and adds a changelog entry and a
regression guard.

### Usage

```bash
# 2 nodes x 8 GB200, previously failed at export on every rank
torchrun --nnodes=2 --node_rank=0 --master_addr=$MASTER --master_port=6000 --nproc_per_node=8 \
  hf_ptq.py --model Llama-3.1-8B-Instruct --dataset cnn_dailymail \
  --recipe general/ptq/fp8_default-kv_fp8 --batch_size 8 --calib_size 512 \
  --export_path ./Llama-3.1-8B-Instruct-fp8_default-kv_fp8 --use_fsdp2
```

### Testing

- End-to-end on 2 nodes by @shengliangx on the original branch: dense
Qwen3-8B and Qwen3-30B-A3B (MoE) FSDP2 PTQ fp8 checkpoints export
successfully.
- Added `tests/examples/hf_ptq/test_export_quantized_context.py`, a
CPU-only guard asserting `export_quantized` enters `torch.no_grad()` and
not `torch.inference_mode()`. A functional regression test would need a
2-node FSDP2 job, which CI does not run, so this encodes the invariant
instead.
- `pre-commit run --files` clean on all three changed files.

Reporter (Kenny Kang, GPU SWQA) still needs to confirm on the original
2x8 GB200 Llama-3.1-8B repro.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

Keyword `Committed_ModelOpt_0.46.0` on the bug — should land for 0.46.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Fixed multi-node quantized model exports to prevent runtime errors
when gathering and detaching parameters.
  * Improved compatibility with FSDP2 during Hugging Face PTQ exports.

* **Tests**
* Added coverage to verify the export process uses the compatible
gradient context.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:57:24 +00:00
Keval Morabia 5dde396bdf Fix vLLM 0.24+ compatibility: registry TypeError and MoE RoutedExperts port (#2054)
### What does this PR do?

Type of change: Bug fix

vLLM 0.24 (as shipped in `nemo:26.08`) reworked the fused-MoE layer,
which broke ModelOpt in two ways:

1. **`FusedMoE` became a factory function** returning a `MoERunner`
pipeline. Registering it put a plain function into
`QuantModuleRegistry`, so *every* later registry lookup raised
`TypeError: issubclass() arg 2 must be a class, ...` — taking down large
parts of `tests/gpu_megatron` (TE, Megatron chaining, MSE calibrator)
that have nothing to do with vLLM.
2. **Expert weights moved onto a `RoutedExperts` submodule** of
`MoERunner`, and `UnquantizedFusedMoEMethod` moved modules, so the MoE
fakequant path had no valid registration target.

Changes:

- `_DMRegistryCls.register` now asserts keys are `nn.Module` subclasses,
so a future upstream change fails at the registration site instead of as
a confusing `TypeError` at lookup.
- Register `RoutedExperts` (vLLM >= 0.24) while keeping the `FusedMoE` /
`SharedFusedMoE` registrations for older releases; both go through the
same `_QuantFusedMoEBase`. `MoERunner` calls `forward_modular` /
`forward_monolithic` directly (`RoutedExperts.forward` raises by
design), so those are hooked instead of `forward`.
- The fused-MoE kernel patch now covers `experts.triton_moe` in addition
to `fused_moe`. 0.24's launcher binds the kernel names at import time,
so patching only the defining module would leave fakequant **silently
inactive**.
- `UnquantizedFusedMoEMethod` is resolved from either module layout.
- `examples/vllm_serve/vllm_reload_utils.py`: quantizer module paths are
now `mlp.experts.routed_experts.*`, so HF→vLLM expert key mapping
inserts a matching `.routed_experts` infix when that layout is present.
- CI `gpu_vllm` now runs on **two** containers: `v0.24.0` (first release
with the FusedMoE-factory / `RoutedExperts` layout — `v0.24.1` was never
released, `nemo:26.08` ships a `0.24.1.dev0` build of the same layout)
and `v0.20.0`, which keeps the legacy `FusedMoE`/`SharedFusedMoE`
branches covered.

Test fixes for the newer vLLM (not product bugs):

- The tiny Llama fixture used `hidden_size=32 / 16 heads` →
`head_dim=2`, which `FLEX_ATTENTION` (the only backend available in this
image) rejects with `NYI: embedding dimension ... must be at least 16`,
killing the engine core at warmup. Now `head_dim=64`, matching the
Qwen3-MoE fixture.
- The FlashInfer metadata-builder stub used `causal=False`, which in
0.24 forces the FI-native path (`all_uses_trtllm = causal and ...`)
requiring workspace buffers and real wrapper planning. Keep it on the
all-TRTLLM path it was originally exercising; the stashed `_modelopt_*`
fields are path-independent.

### Usage

No API change — existing `mtq.quantize` / `examples/vllm_serve` flows
work unmodified on both old and new vLLM.

### Testing

All runs in the `nemo:26.08.rc3` container (vLLM
`0.24.1.dev0+gee0da84ab`, the same 0.24.1 the CI job now pins).

**Suites**

- `tests/gpu_vllm`: **73 passed, 1 skipped** (was 70 passed / 3 failed).
- `tests/gpu_megatron`: all pass (previously ~120 failures, all from the
registry `TypeError` — TE, Megatron chaining and MSE-calibrator tests
that never touch vLLM).
- `tests/unit/torch/opt/test_dynamic.py`: 2 passed, including the new
`test_register_rejects_non_module_classes` (rejects a factory function
and a non-`nn.Module` class, and asserts no partial registration).
- `pre-commit` clean on all touched files.

**MoE fakequant verified by module-tree probe, not just by test
assertions**

After `mtq.quantize(..., NVFP4_DEFAULT_CFG)` inside the vLLM worker,
every weight-owning module was enumerated on tiny DeepSeek-V3 (MLA +
routed MoE + shared experts) and tiny Qwen3-MoE:

```
model.layers.0.mlp.experts.routed_experts [QuantRoutedExperts]
  w13_input_quantizer=3.484  w13_weight_quantizer=0.0840
  w2_input_quantizer=0.1060  w2_weight_quantizer=0.0845
model.layers.0.mlp.shared_experts.gate_up_proj [QuantMergedColumnParallelLinear]  ✅
model.layers.0.mlp.shared_experts.down_proj    [QuantRowParallelLinear]           ✅
```

Weight amax being populated (not just input amax) means the `B is
self.w13_weight` identity check and the Parameter-swap weight-fakequant
branch actually execute through 0.24's kernel path — i.e.
`forward_modular`/`forward_monolithic` really are the live entry points
and the `experts.triton_moe` patch target is the one that fires.
Unquantized modules were only the expected ones: embeddings, RMSNorms,
MoE router `gate`, `lm_head`.

Registration parity vs. older vLLM: Row/Column/MergedColumn/QKV
`ParallelLinear` and all four attention types (`Attention`,
`CrossAttention`, `EncoderOnlyAttention`, `MLAAttention`) register
unchanged; `FusedMoE` → `RoutedExperts`; `SharedFusedMoE` has no
counterpart because the `shared_fused_moe` module no longer exists in
0.24 — shared experts are now a plain MLP whose linears we already
quantize (confirmed above).

**Known gaps (pre-existing, not regressions from this PR)**

- `DeepSeekV2FusedQkvAProjLinear` is not quantized: it subclasses
`MergedColumnParallelLinear` but overrides `forward`, so the registry's
shared-forward rule declines it. Pre-0.24 the equivalent (`q_a_proj` /
`kv_a_proj_with_mqa`) were `ReplicatedLinear`, which ModelOpt never
quantized — effective coverage is unchanged.
- MoE fakequant hooks only the Triton expert kernels;
FlashInfer/CUTLASS/DeepGEMM MoE backends bypass them (why the fixtures
pin `moe_backend="triton"`).
- `_setup` still requires a plain `UnquantizedFusedMoEMethod`; a
`FusedMoEModularMethod` swap (some DP/all2all configs) still asserts.

**Not covered by tests**

- `examples/vllm_serve/vllm_reload_utils.py` — the expert key mapping is
now asserted in `test_tiny_qwen3_moe_quantize` against the quantizer
module paths of a booted MoE model, so a stale infix fails loudly
instead of silently serving uncalibrated experts. The rest of the reload
path is still inspection-only. Note the registry key moved
`vllm_FusedMoE` → `vllm_RoutedExperts` and quantizer paths gained
`.routed_experts`, so a `modelopt_state` saved under an older vLLM will
not restore onto 0.24 as-is.
- The legacy `FusedMoE`/`SharedFusedMoE` branches are covered by the
second CI entry; the `v0.20.0` job is green on this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — all new paths are
feature-detected; older vLLM keeps the `FusedMoE` registration.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — existing
`tests/gpu_vllm` coverage exercises the new registration path.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌

### Additional Information

Found while bumping the Megatron test environment from `nemo:26.06` to
`nemo:26.08.rc3`.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added compatibility for newer vLLM MoE implementations, module
layouts, and routed-expert configurations.
* Improved model reload support for models using routed-expert
submodules.
* **Bug Fixes**
* Improved detection and patching of vLLM MoE execution paths across
supported configurations.
* Registry validation now rejects invalid module registrations without
partially applying changes.
* **Tests**
* Expanded GPU coverage for vLLM 0.24.0, dynamic module validation, and
causal attention metadata paths.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-08-04 09:37:06 -07:00
Keval MorabiaandClaude Opus 5 9e3425de05 Fix saving pruned Nemotron-3-Nano hybrid_override_pattern with MTP or Pipe symbols (#2061)
Saving pruned Nemotron-3-Nano (with MTP) to HF format raised an
assertion which is fixed here

Tested on nemo:26.04 with transformers 4.57

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved exported hybrid layer patterns for pruned models by removing
MTP and pipeline-parallel markers.
* Ensured exported configurations accurately represent the model’s main
layers.

* **Tests**
* Added coverage for pruning models with an MTP prediction layer and
hybrid override patterns.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:15:11 +05:30
Chenjie LuoandClaude Opus 5 77dbeb1872 Add optional MLflow tracking to hf_ptq.py (#2023)
### What does this PR do?

Type of change: new feature

Adds `modelopt.torch.utils.mlflow.MlflowRunLogger`, a reusable helper
for recording a script run on an MLflow tracking server, and wires
`examples/hf_ptq/hf_ptq.py` up to it via `--mlflow <tracking-uri>` so a
PTQ run can be reproduced from its MLflow entry alone. Without the flag,
behavior is unchanged — every hook is gated on it.

The logger lives in the library rather than the example so other scripts
can record runs the same way: it takes a tracking URI, an experiment
name and an explicit `enabled` flag, with params, tags and artifacts
passed in. `hf_ptq.py` supplies only the PTQ-specific pieces (its
params, the resolved recipe, the quantization summaries). `mlflow` is an
optional dependency, imported only once tracking is enabled, so it is
not a new requirement for the library.

The run is opened **before the model loads**, so a bad URI or an
unreachable server fails in seconds rather than after hours of
calibration. The invocation and the recipe are uploaded at that point
too, which keeps a crashed run useful: it is still recorded, with status
`FAILED` and its log attached.

Uploaded artifacts:

| Artifact | Contents |
| --- | --- |
| `command.txt` | The full invocation, copy-pasteable |
| `version.txt` | The ModelOpt version that ran (also a searchable tag)
|
| `recipe/resolved_recipe.yaml` | The `--recipe` with `$import`s
expanded |
| `logs/hf_ptq.log` | Everything the run printed, including a crash
traceback |
| `summary/quant_summary.txt` | Per-quantizer summary (unless
`--no-verbose`) |
| `summary/moe.html` | Per-expert calibration token counts, when the run
produces them |

Plus model / format / calibration settings as searchable params, and
`user` / `hostname` / `modelopt_version` / `git_sha` tags.

Three design points worth review:

1. **The recipe is uploaded resolved, not verbatim.** A recipe may be a
directory or use `$import`s, so the source file is not self-contained.
For
`huggingface/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_at_6p0bits-active_moe`
the source is 2,230 B / 58 lines against 7,563 B / 308 lines resolved —
the raw file records under 30% of what actually ran.
2. **`hf_ptq.py` has no logging framework** (bare `print()`), so the log
is produced by teeing stdout/stderr. Handlers that libraries bound to
`sys.stderr` at import time are re-pointed at the tee for the run's
duration and handed back afterwards; without that, `transformers` /
`huggingface_hub` warnings reach the console but never the log. Native
(C-level) output is still not captured — documented in the README.
3. **The recipe upload lives in the caller, not the library.** That
keeps `modelopt.recipe` out of `modelopt.torch.utils`, which would
otherwise risk a `modelopt.torch.utils` → `modelopt.recipe` →
`modelopt.torch.quantization` → `modelopt.torch.utils` import cycle.
4. **MLflow failures never fail the quantization.** Startup validation
is fatal by design (it is before any GPU work); the end-of-run upload is
best-effort.

Only the main rank uploads, so `--use_fsdp2` runs produce a single run.

### Usage

```bash
python hf_ptq.py \
  --pyt_ckpt_path <huggingface_model_card> \
  --recipe general/ptq/nvfp4_default-kv_fp8_cast \
  --export_path <quantized_ckpt_path> \
  --mlflow https://<your-mlflow-server>/
```

```
[mlflow] experiment: $USER/hf_ptq/<checkpoint basename>-<recipe name>
[mlflow] run: https://<your-mlflow-server>/#/experiments/13/runs/c243352e...
```

`--mlflow_experiment` and `--mlflow_run_name` override the defaults
(`$USER/hf_ptq/<basename>-<recipe name or --qformat>`, and the UTC start
time). Passing `--mlflow` with no value uses `$MLFLOW_TRACKING_URI`.
Authentication uses MLflow's own env vars.

### Testing

**Unit** — 51 tests in `tests/unit/torch/utils/test_mlflow.py` for the
library, plus 13 in `tests/examples/hf_ptq/test_hf_ptq_args.py` for the
hf_ptq wiring. CPU-only, no network and no `mlflow` dependency (driven
against a stub module). Covers experiment-name derivation and
sanitization, URI accept/reject, tee pass-through, the pre-bound-handler
redirect, artifact renaming, skipping absent optional outputs, the
disabled path, and `version.txt`. 85 tests pass together with the
existing `test_hf_ptq_args.py` / `test_example_utils.py`.

**Hardware** — real PTQ runs against a live MLflow server:

| Run | Result |
| --- | --- |
| Qwen3-0.6B, NVFP4 PTQ, 1×B200 | `FINISHED`, all artifacts, sane
post-quant generations |
| Qwen3.6-35B-A3B MoE, AutoQuantize
`w4a16_nvfp4_fp8_at_6p0bits-active_moe`, 2×B200 | `FINISHED` in 63 min,
search hit `effective bits: 6.00`; 106 KB log capturing every per-layer
decision, 4.4 MB quant summary |
| Qwen3.6-35B-A3B, plain NVFP4 PTQ, 2×B200 | `FINISHED` |
| Qwen3-0.6B re-run after the library move, 1×H200 | `FINISHED`, all
five artifacts including `version.txt` |
| Run **without** `--mlflow` after the review fixes | exactly 1
`[load_recipe]` line and 0 `[mlflow]` lines, confirming the untracked
path is untouched |
| Two runs sharing one `--export_path`, second crashed early | second
run uploads **no** summary — the first run's 124 KB file on disk is
correctly not attributed to it, and its traceback is in the log |
| Crash mid-run (gated HF dataset) | `FAILED` recorded with log +
traceback attached, summaries correctly absent |
| Malformed URI | Rejected by `argparse` with a `Did you mean
https://…?` hint |
| Unreachable host | Fails in 9.9 s total, before any model load |
| No `--mlflow` | Exit 0, no MLflow output, unchanged export |

**Coverage gap, stated plainly:** `summary/moe.html` is verified only
against a synthetic file (unit test + a real upload). It could not be
produced naturally — `expert_token_count` buffers live on
`_QuantSparseSequentialMoe`, while Qwen3.5/3.6 experts take the fused
`_QuantFusedExperts` path, so no such file is written for these models
regardless of `--moe_calib_experts_ratio`. The uploader's conditional is
correct; the branch simply had no natural input available here.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — new optional flags only; no
`--mlflow` means no behavior change.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — adds
`mlflow` as an optional extra in `pyproject.toml`
(`nvidia-modelopt[mlflow]`, folded into `all`) and to
`examples/hf_ptq/requirements.txt`. Apache-2.0 (permissive). Imported
lazily, so it is not required to install or import ModelOpt. No code
copied from other sources.
- Did you write any new necessary tests?: ✅ — 29 new unit tests.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — 0.47 New Features.
- Did you get Claude approval on this PR?: ❌ — `/claude review` not yet
run. A self-review was done first and its six findings are fixed in the
third commit (the notable one: gathering the MLflow inputs re-read the
recipe on *every* run, including without `--mlflow`).

### Additional Information

The one deliberate coverage gap is `summary/moe.html`, described under
Testing: no model available here takes the sparse-sequential MoE path
that writes it, so it is covered by unit test and a synthetic upload
rather than a natural one. The uploader treats it as an optional output
and skips it when absent, which is exercised by test.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 11:25:58 +05:30
Keval MorabiaandClaude Opus 4.8 93b9e4b176 Add Megatron-Bridge prune & quantize launcher pipelines (#2031)
### What does this PR do?

Type of change: new example + small launcher / modelopt-example features
(backward compatible)

Adds end-to-end ModelOpt **launcher** pipelines for the Megatron-Bridge
flow on Nemotron-3-Nano-30B-A3B, the minimal launcher features to run
them wrapper-free from YAML, and an **in-step accuracy gate** for
Minitron pruning.

**New launcher examples**
(`tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/`):
- `mbridge_prune.yaml` — Minitron prune **with an in-step MMLU gate** →
vLLM sanity gen (2 tasks)
- `mbridge_quantize.yaml` — FP8 quantize → unified-HF export → MMLU gate
on the vLLM backend, which doubles as the deploy sanity check (3 tasks).
Matches the
[tutorial](https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md).

**Prune accuracy gate** (reuses the search's own score — no separate
eval step):
- `modelopt/torch/prune/plugins/mcore_minitron.py` —
`MCoreMinitronSearcher` now stores the exported best `CandidateSubnet`
under `state_dict["best"]` (additive; sits beside the existing
`sorted_layers` key).
- `examples/megatron_bridge/prune_minitron.py` — new
`--score_lower_bound`: reads `pruning_scores["best"].score` and exits
non-zero if the exported model is below the floor. Score-agnostic (any
`--prune_score_func`); rejected with `--prune_export_config` (manual
pruning has no score).

**Launcher (`tools/launcher`)** — run single-node Megatron-Bridge
one-liners directly from YAML:
- `SandboxTask.inline` — a command in the YAML, no `common/**/*.sh`
wrapper (single-line; folded scalar)
- `SandboxTask.reqs` / `reqs_file` — pip-install deps in the container
before the command (shell-safe; on Slurm the install is rank-0-guarded
so multi-rank tasks don't race)
- `SlurmConfig.docker_user` — local-Docker user (e.g. `root`); ignored
on Slurm
- `get_default_env` honors `HF_HOME` / `TRITON_CACHE_DIR` env overrides,
so a non-CI user can point caches at a writable path (the shared
`/cicd/hf-cache` is owned by the CI account)
- reject `args` together with `inline`

**`examples/llm_eval/lm_eval_hf.py`**:
- `--accuracy_lower_bound` — gate on the single requested task's `acc`
(used by the quantize MMLU step; exits non-zero if below)
- drop ModelOpt (hf-only) args for non-`hf` backends, so `--model vllm`
works on a deployable quantized checkpoint

### Usage

```bash
cd tools/launcher
# Prune (in-step MMLU gate) -> vLLM gen
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/mbridge_prune.yaml --yes
# FP8 quantize -> unified-HF export -> MMLU gate (vLLM)
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/mbridge_quantize.yaml --yes
```

### Testing

- Launcher unit tests for `inline`, `reqs`/`reqs_file`, `docker_user`,
the `args`+`inline` guard, and example-resolve; ruff / mypy / bandit
clean. `tests/examples/megatron_bridge/test_prune_minitron.py` now
passes `--score_lower_bound=0.01` to exercise the gate path on the tiny
models.
- **End-to-end on the real Nemotron-3-Nano-30B-A3B (4×B200, OCI-HSG):**
- Prune 30B → 3B-active: `[score_gate] mmlu_10pct = 0.5196 >= 0.45
PASS`; vLLM gen coherent.
- FP8 quantize → unified-HF export (`Detected ModelOpt fp8 checkpoint`)
→ MMLU on the vLLM backend `acc = 0.7077 >= 0.60 PASS`.
- Earlier smoke on **Qwen3-0.6B** in `nemo:26.06` through the same flow.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (new state-dict key is
additive; new CLI args default to off)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- tooling/examples + additive searcher state key -->
- Did you get Claude approval on this PR?: ❌ <!-- pending -->

### Additional Information

- **Container pinning:** saving a pruned Nemotron-H to HF requires
`transformers<5`, so `mbridge_prune.yaml` runs on `nemo:26.04` (26.06
drops it); quantize/export run on `nemo:26.06`.
- **`docker_user: root`** is set on all example tasks — local-Docker
only (ignored on Slurm), needed so downstream tasks can read task_0's
root-owned checkpoints and to read the image's root-only
`/opt/Megatron-Bridge`.
- The quantize MMLU step passes `enforce_eager=True` to vLLM — for a
run-once eval this skips ~17 min of CUDA-graph capture / `torch.compile`
with no accuracy change.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added inline shell commands and per-task Python dependency
installation for launcher workflows.
- Added configurable Docker user selection and preservation of existing
cache environment settings.
- Added evaluation accuracy and pruning score gates that fail workflows
below configured thresholds.
  - Added NVIDIA Nemotron pruning and quantization workflow examples.
  - Improved backend-specific handling of ModelOpt options.

- **Documentation**
- Documented inline commands, dependencies, variable substitution, and
configuration examples.

- **Bug Fixes**
  - Strengthened task execution validation and configuration checks.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-04 10:09:11 +05:30