mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
a74054ab2b56d2a3e45a7600e112bd334abd35ee
497
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a74054ab2b |
Let callers add MLflow tags to a fakequant serve's run (#2364)
### What does this PR do?
Type of change: new feature
The quantization run records what this library can see — the model, the
checkpoint, the vLLM and ModelOpt versions — but nothing about the
harness that launched it. A downstream tool that wants its own revision,
a sweep id, or a ticket number on the run has no way to put it there
today:
- `_run_tags()` returns a fixed dict
- `quant_config` (which becomes the run's params) is a hardcoded set of
`QUANT_*` variables
- MLflow itself has no environment variable for arbitrary tags
`MODELOPT_MLFLOW_EXTRA_TAGS` takes comma-separated `key=value` pairs and
merges them into the run's tags.
Two details worth a reviewer's attention:
**It joins `MLFLOW_ENV_VARS`.** A Ray-backed serve receives only the
variables named there, and the tracker runs in the rank-0 worker —
omitting it would make the feature silently do nothing under Ray.
**Caller tags are merged first**, so the library's own keys (`tool`,
`model`, `checkpoint_path`, `vllm_version`) are written over them and
keep describing the run truthfully whatever a caller sends.
`key=value` rather than JSON, learned from a live run: the variable
reaches the worker through a shell `export VAR="..."`, and JSON's own
double quotes terminate that quoting —
```
export MODELOPT_MLFLOW_EXTRA_TAGS_732b_DEPLOYMENT="{"internal_version": "4d8c"}"
```
arrived as `{`. A quote-free format survives verbatim and needs no
`json` import or exception handling. Splitting on the first `=` keeps
values that contain one, such as a URL with a query string.
### Usage
```bash
export MODELOPT_MLFLOW_EXTRA_TAGS="modelopt_internal_version=49fa29d5,sweep=kv-study"
python3 vllm_serve_fakequant.py "$MODEL" --mlflow https://your-mlflow-server/ ...
```
### Testing
Unit-level, over the helper: unset and empty variable, one and several
pairs, surrounding whitespace, an empty value, an entry with no `=`, a
trailing comma, and a value containing `=`. None raise; malformed
entries warn and are skipped.
End to end on a real fakequant serve (Nemotron-3-Nano-30B-A3B BF16,
`NVFP4_DEFAULT_CFG`, TP=8, Ray executor, vLLM 0.15, SLURM):
```
modelopt_internal_version '49fa29d5'
modelopt_version '0.47.0rc0.post32+gd38ed5ead'
git_sha 'd38ed5ead'
quant_cfg 'NVFP4_DEFAULT_CFG'
```
The tag was written by the `RayWorkerWrapper` process, which exercises
the whole path — env var → shell export → `--container-env` → raylet →
Ray actor → `_run_tags` — and confirms the `MLFLOW_ENV_VARS` entry is
doing its job. Also verified that the emitted payload survives a shell
export round-trip unchanged.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ <!--- Additive; with the
variable unset the tags are exactly as before. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ <!--- Verified manually as
above; there is no existing test module for vllm_mlflow_utils. Happy to
add one if you would like it. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <!--- Small additive feature in an example; tell me if it warrants an
entry. -->
- Did you get Claude approval on this PR?: ❌
### Additional Information
Consumed by Model-Optimizer-Internal MR !141/!147, which sets the
variable so a fakequant eval records the same harness commit on both its
quantization run and its evaluation-score run.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
d19925e446 |
simple refactor(export): split TensorRT-LLM-only code into modelopt/torch/export/trtllm (#2365)
### What does this PR do?
Type of change: refactor.
**The TensorRT-LLM checkpoint export format is deprecated.** Per
`docs/source/deployment/1_tensorrt_llm.rst`: *"The
`export_tensorrt_llm_checkpoint` API will be deprecated in future
releases. Users are encouraged to transition to the unified HF export
API, which provides enhanced functionality and flexibility for exporting
models to multiple inference frameworks including TensorRT-LLM, vLLM,
and SGLang."*
That deprecated code was not sitting off to one side — it was
**interleaved with the export path we actually want to grow.**
`modelopt/torch/export` mixed the deprecated TensorRT-LLM checkpoint
logic with the framework-agnostic HF/Megatron export code, in the same
modules:
- `layer_utils.py` was 1,986 lines, of which ~1,600 were TensorRT-LLM
`build_*_config` builders. The HF path imports this module for five
small predicates (`is_moe`, `is_quantlinear`, …) and dragged the whole
deprecated builder set in with them.
- `model_config.py` held the TensorRT-LLM `ModelConfig` dataclasses
*and* the `QUANTIZATION_*` / `KV_CACHE_*` constants that every backend
needs, so all of HF export imported the deprecated checkpoint schema to
get a format name string.
- `quant_utils.py` carried two helpers whose only caller is the
deprecated `postprocess.py`.
**This PR isolates the deprecated format so it stops polluting the
HuggingFace export path.** Everything reachable only from
`export_tensorrt_llm_checkpoint` now lives under
`modelopt/torch/export/trtllm/`, and the dependency is **one-way**:
`trtllm/` reaches into the parent through `quant_format`, `quant_utils`
and `layer_utils`, and **no implementation module in the parent imports
`trtllm/`.** The single exception is the deprecation re-export in
`modelopt/torch/export/__init__.py` described below, which is scheduled
for deletion in 0.49.0.
That one-way edge is the property worth protecting in review. It means
the deprecated format can be evolved, frozen, or eventually removed
without touching HF export, and HF export can no longer accidentally
grow a dependency on it.
### Deprecation handling
The format has carried a deprecation notice in the deployment docs since
`bc546943b4` (2025-10-08, first shipped in 0.39.0) — about 11 months.
But the deprecation policy in `README.md` also specifies *how* a
deprecation is communicated: a changelog entry, a source statement of
timing, and a runtime warning on use. **None of those existed**; only
one docs page ever said anything. So 0.48.0 is the first release that
gives users a signal they can act on, and this PR treats it as the
*start* of the migration period rather than the end:
- Both entry points now emit a `DeprecationWarning` naming 0.48.0 and
the 0.49.0 removal.
- `export_tensorrt_llm_checkpoint` and
`torch_to_tensorrt_llm_checkpoint` **remain importable from
`modelopt.torch.export`** for this release only, so existing callers
keep working *and* actually receive the warning. Removing the path in
the same release that first warns would mean callers hit `ImportError`
and never see it.
- The 0.49.0 removal date is stated in all four channels the policy
names: the runtime warning, the source (`.. deprecated:: 0.48.0` plus a
comment), the changelog, and the deployment doc.
The deeper module paths (`modelopt.torch.export.model_config_export`,
`modelopt.torch.export.model_config`) are **not** forwarded. Neither
appeared in a docs example, and `model_config.py` never declared
`__all__`, so by the `__all__` convention in `CONTRIBUTING.md` they were
never part of the public surface.
Eight modules had no non-TRT-LLM importer and moved whole:
`model_config_export`, `model_config_utils`, `postprocess`,
`distribute`, `tensorrt_llm_utils`, `tensorrt_llm_type`,
`hf_config_map`, `mcore_config_map`.
Three were genuinely mixed and were split by call-graph analysis rather
than by file:
| module | stayed shared (HF path) | moved to `trtllm/` (deprecated) |
|---|---|---|
| `model_config.py` | `QUANTIZATION_*`, `KV_CACHE_*`,
`FUSION_FREE_FORMATS` → new leaf module `quant_format.py` | the
`ModelConfig` dataclasses + `LINEAR_*`/`LAYERNORM_*` checkpoint-layout
constants |
| `layer_utils.py` | 9 module-shape predicates and MoE quantizer helpers
(`is_moe`, `is_quantlinear`, `get_experts_list`,
`sync_moe_gate_up_amax`, …) | the 39 `build_*_config` builders and
enc/dec helpers |
| `quant_utils.py` | everything else | `get_scaling_factor_from_weight`,
`resmooth_and_get_scale` (only caller is `trtllm/postprocess.py`) |
`adjust_attn_amax_values` was deliberately left in the shared
`quant_utils.py`: it has no production caller at all (only a test), so
"used only by TRT-LLM export" is not demonstrable for it.
Nothing was added or removed. `export_tensorrt_llm_checkpoint` behaves
exactly as before, just from a new import path and with a warning
attached.
### Usage
```python
# Deprecated TensorRT-LLM checkpoint export — new home, and warns on call
from modelopt.torch.export.trtllm import (
export_tensorrt_llm_checkpoint,
torch_to_tensorrt_llm_checkpoint,
)
from modelopt.torch.export.trtllm.model_config import ModelConfig
# The pre-0.48 path still works for one release, and warns — removed in 0.49.0
from modelopt.torch.export import export_tensorrt_llm_checkpoint
# Shared format constants — new home, still re-exported from the top level
from modelopt.torch.export.quant_format import QUANTIZATION_NVFP4, KV_CACHE_FP8
from modelopt.torch.export import QUANTIZATION_NVFP4 # still works
# The recommended path — unchanged
from modelopt.torch.export import export_hf_checkpoint, get_model_type
```
### Testing
- `pre-commit` on all changed files: passes (ruff, ruff-format,
**mypy**, bandit, markdownlint). mypy caught one implicit re-export of
`is_layernorm`, now imported from the shared module directly.
- `tests/unit/torch/export`: **189 passed**. With the new `trtllm/` test
dir: **193 passed**.
- Full `tests/unit/torch`: **2367 passed, 0 export failures**. The 45
failures are pre-existing environment issues — a deepspeed circular
import and a read-only HF cache — confirmed by reading their error text,
not assumed.
- `pytest tests/gpu/torch/export --collect-only`: 172 items, no
collection error.
- In-repo consumers updated and re-verified by an AST scan that imports
every `modelopt.torch.export*` module referenced anywhere in the tree
and checks each imported name still resolves: `hf_ptq.py`,
`export_trtllm_ckpt.py`, `deepseek_v3/ptq.py`, the AutoQuantize
notebook, `hf_ptq/README.md`, 2 docs pages, 4 tests.
- **Deprecation contract is covered by committed tests** (3 new, in the
`trtllm/` test dir): the pre-0.48 top-level import still resolves to the
same objects, `torch_to_tensorrt_llm_checkpoint` warns *at call time*
rather than on first `next()` (it returns a generator, so a naive
`warnings.warn` in the body would fire late or never), and one
`export_tensorrt_llm_checkpoint` call emits exactly one warning rather
than two. The first of these makes closing the migration window early a
test failure rather than a silent regression. `pyproject.toml` sets no
`filterwarnings = error`, so no suite fails on the new warning.
- **After merging `main`** (4 commits, incl. a 180-line rewrite of
`unified_export_megatron.py` that touches a file this PR also edits):
merged with no conflicts, then re-verified rather than trusted — import
scan clean across 24 export modules, `ruff check` clean repo-wide, 193
export tests passing, GPU collection still clean.
**Not run: the GPU suites** (`tests/gpu/torch/export`,
`tests/gpu_trtllm`) — no GPU in my environment.
`tests/gpu/torch/export/test_export.py` had its imports retargeted, so
it is the one most worth a GPU run before merge.
### Reviewer note: the deprecated path has no test coverage
Worth knowing before reviewing. **No test in the repo — including
`tests/examples/` — calls `export_tensorrt_llm_checkpoint`,
`torch_to_tensorrt_llm_checkpoint`, any `build_*_config`,
`convert_to_tensorrt_llm_config`, or `postprocess_model_config`.** So
~4,600 moved lines have no direct tests, and this refactor is validated
by import-graph reasoning, lint and mypy rather than by tests exercising
the moved code.
Given the format is deprecated and scheduled for removal in 0.49.0, **no
new coverage is planned for the conversion path itself** — writing fresh
tests for an API being removed next release isn't a good use of effort.
The gap is documented so reviewers can weigh the risk, not as a TODO.
(The deprecation *mechanism* is tested; see Testing.)
One caveat on how the gap was established: a runtime check showing all
12 `trtllm` modules in `sys.modules` after the export suites is *not*
evidence of coverage — importing any submodule runs
`trtllm/__init__.py`, which star-imports `model_config_export` and pulls
in the rest. Real line coverage could not be measured (`coverage`'s
tracer is incompatible with this venv's torch build: `ValueError: module
functions cannot set METH_CLASS or METH_STATIC`, on both the C tracer
and `sysmon`). The claim rests on a call-site audit generated from the
actual public symbols of those modules.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ for the public API —
`export_tensorrt_llm_checkpoint` and `torch_to_tensorrt_llm_checkpoint`
remain importable from `modelopt.torch.export` through the 0.49.0
migration period, now with a `DeprecationWarning`. The undocumented
submodule paths `modelopt.torch.export.model_config_export` and
`.model_config` did move; see **Usage**.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
code or dependencies; existing code relocated.
- Did you write any new necessary tests?: ✅ — 3 tests covering the
deprecation contract (old import path, call-time warning, exactly-one
warning). One existing test also moved to mirror the source split. No
new coverage for the deprecated conversion path itself; see the note
above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 **Deprecations**, covering both the runtime warning and
the new import location.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
Git detected the moves, so the diff stays reviewable: 8 files show as
pure renames (100%), the three split files as rename/copy at 94–99%
similarity, and only `layer_utils.py` as a 79% rewrite — expected, since
it shed 1,616 lines to `trtllm/`.
`examples/hf_ptq/hf_ptq.py` and
`examples/llm_sparsity/weight_sparsity/export_trtllm_ckpt.py` still call
the deprecated API, so those examples now print the warning. That is the
intended nudge, but happy to silence or migrate them if preferred. They
import from the new `.trtllm` path already, so they need no change at
0.49.0.
Two incidental changes, easy to revert if unwanted:
- `modelopt/torch/export/layer_utils.py` mode `100755 → 100644` (it was
needlessly executable).
- The new test is named `test_trtllm_quant_utils.py`, not
`test_quant_utils.py`: these directories have no `__init__.py`, so
pytest derives the module name from the bare filename and the shorter
name fails collection with `import file mismatch` against the existing
`test_quant_utils.py` one level up.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added shared quantization and KV-cache format definitions for export
workflows.
- Added expanded TensorRT-LLM export support, including broader model
architecture and quantization handling.
- Added distributed export utilities for coordinating checkpoint data
across processes.
- **Deprecation**
- TensorRT-LLM checkpoint export now emits a warning and is scheduled
for removal in version 0.49.0.
- Use the documented export module and save optimized model state
explicitly when needed.
- **Documentation**
- Updated guides and examples with new import paths and deprecation
guidance.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
9d0df45849 |
specdec: config_overrides for nested text_config checkpoints + load VLM-capable bases in merge_lora (#2289)
### What does this PR do?
Type of change: New feature + bug fix
Two related gaps, both hit while enabling EAGLE3 on a checkpoint whose
config nests its text dims.
**1. `config_overrides` for checkpoints whose `text_config` dims don't
propagate.**
Some multimodal checkpoints carry the real text-tower dims only under
`config.text_config`, leaving the parent fields `None`.
`from_pretrained` then builds a text tower with the wrong shape.
`load_vlm_or_llm` gains an optional `config_overrides` dict applied to
*both* the parent config and its `text_config` before instantiation, and
the three entrypoints that load checkpoints — `ar_validate.py`,
`export_hf_checkpoint.py`, `merge_lora.py` — get a `--config_overrides`
passthrough. `main.py` threads it from `ModelArguments`.
**2. `merge_lora.py` could not merge into any VLM base.**
It loaded via `AutoModelForCausalLM`, which cannot load architectures
absent from the CausalLM Auto map — every VLM base failed. It now goes
through `load_vlm_or_llm`, which routes VLMs to
`AutoModelForVision2Seq`/`AutoModelForImageTextToText` and plain LLMs to
`AutoModelForCausalLM` with the same `dtype`/`device_map`, so LLM
behavior is byte-for-byte unchanged.
Also adds an optional `transformers_cosmos3` import so `cosmos3_omni` is
registered with `AutoConfig` before use, and dispatches that
`model_type` to its model class directly — that plugin registers only a
*config*, never a model under `Auto*`, so `AutoModelForCausalLM` raised
`KeyError('cosmos3_omni')` regardless of imports. The import is wrapped
in `contextlib.suppress(ImportError)`, so it is a no-op when the plugin
isn't installed.
### Usage
```bash
# Checkpoint whose real dims live under config.text_config
python examples/speculative_decoding/scripts/ar_validate.py \
--model_path <ckpt> --trust_remote_code \
--config_overrides '{"num_hidden_layers": 36, "intermediate_size": 12288, "num_key_value_heads": 8}'
# Same flag on export and merge
python examples/speculative_decoding/scripts/export_hf_checkpoint.py \
--model_path <ckpt> --export_path <out> --config_overrides '{"num_hidden_layers": 36}'
python examples/speculative_decoding/scripts/merge_lora.py \
--base_model_path <base> --exported_lora_dir <out> --output_path <merged> \
--config_overrides '{"num_hidden_layers": 36}'
```
```python
model = load_vlm_or_llm(path, config_overrides={"num_hidden_layers": 36}) # default None
```
### Testing
Exercised end-to-end on a Cosmos3-Nano (16B, 36-layer text tower) EAGLE3
LoRA run:
- **Training** — the base loads with all 36 text layers and correct
dims; two 4-epoch co-training runs completed (46,816 steps each).
- **Export + merge** — produced `adapter_model.safetensors` and a merged
base. Verified correct by per-layer weight diff: a `start_layer=18` run
changed **exactly** layers 18-35, with layers 0-17 bit-identical to the
base.
- **AR validation** — `--config_overrides` loads the trained checkpoint;
80/80 MT-Bench samples, AR 3.42.
- **Regression check** — `merge_lora` via `load_vlm_or_llm` produces a
base loadable by `lm_eval`; ifeval/arc_challenge/winogrande all ran to
completion.
No local unit-test run: `nvidia-modelopt` isn't installed in my
checkout, so `tests/conftest.py` fails to import. Relying on CI.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `config_overrides` defaults
to `None`; the `merge_lora` loader swap keeps the same class, dtype and
device_map for plain LLMs.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependency; `transformers_cosmos3` is an optional import guarded by
`contextlib.suppress`.
- Did you write any new necessary tests?: ❌ — exercising these paths
needs a checkpoint with a nested `text_config`, which the unit suite has
no fixture for. Happy to add one if a reviewer can point me at a small
suitable model.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — can add a *Speculative Decoding* entry for the `merge_lora` VLM fix
if you consider it changelog-worthy.
- Did you get Claude approval on this PR?: ❌ — not yet run.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added JSON-based model configuration overrides across speculative
decoding, training, validation, export, and LoRA workflows.
- Overrides can update primary model and text configuration settings.
- Expanded support for vision-language models and Cosmos3 Omni
checkpoints.
- **Bug Fixes**
- Improved configuration handling for offline loading and
checkpoint-based initialization.
- Restored draft-model precision during checkpoint loading and model
conversion.
- Added validation for malformed, unsupported, and non-finite override
values.
- Standardized configuration override guidance across command-line
workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
279d510616 |
fix(specdec): correct resume and bound staging in the vLLM hidden-state dump (#2080)
### What does this PR do?
**Type of change:** Bug fix
Fixes two issues in the vLLM offline hidden-state dump
(`examples/speculative_decoding/collect_hidden_states/compute_hidden_states_vllm.py`).
Both are invisible on small dumps and only bite at scale, which is why
they survived until
now — they were found while dumping ~194k conversations for a MiniMax-M3
draft.
**1. Resume silently re-processed already-finished work.**
`keep_conversation` skips conversations whose `.pt` already exists, but
that predicate reads
**on-disk state**, which is not part of the fingerprint `datasets`
computes for `filter()`
(it hashes the function and the dataset). With a persistent HF cache
reused across a resumed
or requeued run, the cached *"keep everything"* result from an earlier
run — computed when
few or no `.pt` files existed — is replayed. The run then re-generates
and **overwrites**
conversations it had already completed, and reports `Removed 0
conversations due to existing
output files` while doing so.
Observed on a 194k-conversation dump: ~62k `.pt` rewritten over a
two-hour window with the
total output count completely flat.
Fix: pass `load_from_cache_file=False` so the filter re-checks the disk
on every run.
**2. Staging exhausted `/dev/shm` partway through large dumps.**
The script generated the **entire** dataset before saving anything. The
KV connector stages
each conversation's hidden states under its `shared_storage_path`
(`/dev/shm`, i.e. RAM, by
default) and they are only freed by `cleanup_hidden_states()` in the
save loop — so every
conversation stayed staged simultaneously. On a large dump this exhausts
the space and the
connector starts failing writes:
```
Hidden-states write failed for req_id=...:
SafetensorError('Error while serializing: I/O error: No space left on device (os error 28)')
```
Fix: generate and save in chunks of `--save-chunk-size` (default 256),
so at most one chunk
is staged at a time. As a side benefit the dump becomes **incrementally
durable** — an
interrupted run (walltime limit, node failure) keeps its finished
conversations and the
resume path above continues from them, instead of losing the whole run's
work.
### Testing
- Reproduced both failures on a 194k-conversation MiniMax-M3 dump (8-way
DP, TP8), and
confirmed both fixes on the same workload: after the change the output
count advanced
monotonically across requeues (123k → 194k) with no rewrites, and
`/dev/shm` stayed bounded
through completion.
- `pre-commit run --files ...` passes (ruff check/format, mypy, bandit,
license, rst checks).
- Behavior is unchanged for a fresh single-shot dump other than the
chunked generate calls;
the default `--save-chunk-size 256` is the only new knob.
### Additional Information
Extracted from #1749, which is otherwise superseded by the streaming
DFlash/DSpark path — these
two fixes are model-agnostic and apply to any offline dump, so they are
worth landing on their
own.
### Before your PR is "*Ready for review*"
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No — the failure modes are
multi-process/at-scale (datasets cache reuse across runs, connector RAM
staging) and are not reproducible in the unit-test harness.
- **Did you add or update any necessary documentation?**: Yes —
CHANGELOG entry.
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added chunked hidden-state generation for large vLLM offline runs.
* Added a configurable save-chunk size, defaulting to 256 conversations.
* Enabled incremental saving and resumption of hidden-state outputs.
* **Bug Fixes**
* Improved resume filtering to accurately detect existing output files.
* Reduced memory usage by saving and releasing each generated chunk.
* Ensured temporary files are cleaned up after interrupted or skipped
saves.
* Added atomic output-file replacement to prevent incomplete results.
* Added validation to prevent invalid conversation IDs from creating
unsafe paths.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
56af187565 |
ar_validate: fail loudly when every sample fails (#2288)
### What does this PR do?
Type of change: Bug fix
`validate_ar()` catches per-sample exceptions, prints a `WARNING`, and
returns whatever succeeded. When *every* sample failed it returned an
empty list, and the reporting block was guarded by `if results and
accelerator.is_main_process:` — so the script printed no results and
exited **0**. A run where 100% of samples failed was indistinguishable
from a successful one.
This bit us on a real run: an EAGLE3 checkpoint loaded with
`device_map="auto"` was sharded across 8 GPUs, every one of the 80
samples died with `Expected all tensors to be on the same device`, and
the job still exited 0 with no AR number anywhere in the log — the
wrapper stamped it PASS.
Now it raises, so the caller sees a non-zero exit. Any previously
"passing" run that printed no AR number was never meaningful.
### Usage
No API change. Existing invocations are unaffected when at least one
sample succeeds:
```bash
python examples/speculative_decoding/scripts/ar_validate.py \
--model_path <ckpt> --steps 3 --osl 1024 --num_samples 80
```
### Testing
Reproduced the silent-pass on a Cosmos3-Nano EAGLE3 checkpoint (80/80
samples failing): before this change the job exited 0 and stamped PASS;
after it, the job exits non-zero with the sample failures visible.
Confirmed the normal path is unchanged by a subsequent run that
completed 80/80 and printed AR 3.42.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — only affects the
all-samples-failed case, which previously produced no output and a
misleading exit 0.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ — the failure path requires
a model that errors during AR validation; the existing suite has no
harness for that.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — behavior fix in an example script, not a released-feature change.
- Did you get Claude approval on this PR?: ❌ — not yet run.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved validation error handling when all samples fail.
* Validation now rejects non-positive sample counts before processing.
* Empty validation results are clearly distinguished from cases where
all samples fail.
* Error messages report the actual number of validation samples
attempted, capped at the available dataset size.
* Empty validation results are no longer reported as successful.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Ye Yu <yeyu@nvidia.com>
|
||
|
|
4956213d67 |
Unblock Qwen3.5/3.6 QAD: Megatron export, calibration, and distillation fixes (#2334)
### What does this PR do? Type of change: Bug fix Everything that blocked running QAD on a quantized Qwen3.5 / Qwen3.6 MoE checkpoint: two Megatron-Core → HuggingFace export bugs that make it unservable (§1–2), the dead code the first leaves behind (§3), a no-op flag (§4), a multi-GPU calibration deadlock (§5), and four distillation / data-prep bugs that stopped QAD itself from running (§6). #### 1. Routed experts were exported packed, and vLLM cannot load that ``` AttributeError: Layer language_model.model.layers.23.mlp.experts has no parameter 'w2_weight_weight_scale_2' for checkpoint weight ...experts.down_proj_weight_scale_2 ``` `mcore_qwen35vl.py` used `GroupedMLPPacking`, mirroring the **BF16 upstream** checkpoint, which really is packed. But that mapping is only used for **quantized** export, and vLLM's quantized MoE loader needs per-expert scales — both released NVFP4 checkpoints (`Qwen3.6-35B-A3B-NVFP4` via hf_ptq, `Nemotron-3.5-Lightning-30B-A3B-NVFP4` via Megatron-LM) are per-expert. `_grouped_mlp_slicing` gains `gate_proj_name` / `up_proj_name` to split each expert's fused gate+up and slice its per-block `weight_scale`; `GroupedGatedMLPSlicing` wires it up. The Megatron checkpoint layout is unchanged, so affected checkpoints need only a **re-export**. `_verify_exported_keys` is relaxed to match: exported modules now contribute their ancestor prefixes, so expanding one source module into many is not reported as ~82 dropped tensors. A module genuinely absent still has nothing beneath its prefix and is still caught. #### 2. A quantized `output_layer` (`lm_head`) could not be checkpointed `GPTModel.sharded_state_dict` drops `output_layer._extra_state` and asserts it is empty. ModelOpt keeps quantizer state there, so saving raised and — since that method also backs the load plan — loading silently restored the layer **unquantized**. `keep_gpt_output_layer_extra_state()` retains it, applied from `megatron_replace_quant_module_hook` so **every** Megatron model gets it (Megatron-LM and NeMo users included, neither of whom can import `mbridge`, which needs `megatron.bridge`). It matches the upstream body by AST before replacing it and self-disables otherwise. [NVIDIA/Megatron-LM#7086](https://github.com/NVIDIA/Megatron-LM/pull/7086) is **closed, not merged**: nemo:26.10 migrates `GPTModel` to `HybridModel`, whose `sharded_state_dict` has no pop-and-assert, so this side keeps the workaround. Not cosmetic: `lm_head` is 248320×2048 = 509M params, **34.6% of per-token weight traffic** on a model with ~2.9B active params. #### 3. Cleanup Nothing maps `GroupedMLPPacking` once qwen3_5 is switched over; it is removed with `_grouped_mlp_packing` and the `quantize=` / `record_quant_config=` parameters that existed only to serve it. Llama-4's `PackNameRemapping` is unaffected. Two smaller review-driven fixes: the gated-split shape checks raise `ValueError` rather than `assert` (stripped under `-O`), and per-expert quant metadata is recorded for `local_expert_indices` rather than every global id, fixing non-contiguous EP. #### 4. Remove the no-op `--moe_calib_experts_ratio` from the Megatron quantize example `examples/megatron_bridge/quantize.py` accepted the flag and threaded it into the `mtq` config, but `_moe_calib_experts_ratio` exists only in `plugins/huggingface.py` (9 refs) and never in `plugins/megatron.py` (0); `mode.py:247` only assigns it to modules already exposing the attribute. On a Megatron MoE model it was accepted and silently ignored — a trap, since on a 256-expert model it reads like a major quality lever. `hf_ptq.py` keeps it, where it works. #### 5. Fix multi-GPU image-text (VLM) calibration deadlocking VLM calibration hung for 30 minutes and died on a gloo timeout whenever `world_size > 1`, with no error until the timeout fired. `NemotronTarPlusJsonlIterable` split its budget with truncating division, so the stream supplied fewer samples than requested (1024 over 3 subsets → 341×3 = **1023**). `_ShardedIterable` gives rank *r* items *r, r+W, r+2W…*, so a stream that is not a multiple of `world_size` leaves the trailing rank one short — it exits the forward loop early and the others block on the next collective. The arithmetic predicts both observed hangs exactly: 1024 → stall at **255/256**, 512 (yielding 510) → **127/128**. Fixed both ends: subset budgets are distributed with `divmod` so they sum exactly, and `_ShardedIterable` truncates every rank to `floor(len / world)` — which also covers `num_samples` not being divisible by `world_size`, as the first fix alone does not. Verified on Qwen3.6-35B-A3B (EP=4, `nemotron_vlm_dataset_v2`, 1024 samples): the configuration that hung twice now completes 256/256 and exports. Unit tests cover both fixes and fail without them. #### 6. Fix the distillation path so QAD can actually run Four independent bugs, all hit while running QAD end to end on Qwen3.6-35B-A3B. Each blocks a different configuration, and together they made every sequence length OOM or abort. - **Context parallel aborts.** The DDP config derived `average_in_collective` from `--sft` alone, but context parallel also needs per-token loss reduction, so any `--cp_size > 1` run died on `Cannot average in collective when calculating per-token loss`. - **`TopKLogitsKLLoss` was not memory-efficient.** Despite documenting "without gathering full logits", it cast the *whole* vocabulary to FP32 before selecting the top-k, allocating two `[seq, vocab]` tensors — 30.3 GiB each at seq 32768 on this model's 248k vocab. Reducing before the cast is equivalent: widening is exact and temperature scaling is monotonic, so the selected entries and the loss are unchanged. - **MTP cross-entropy ran when it had nothing to recover.** `skip_lm_loss` exempts the MTP heads unconditionally, so their CE materialised another FP32 `[seq, vocab]` tensor even when the MTP head is excluded from quantization — as it is in every recipe here (775 of 906 `exclude_modules`, zero MTP `weight_scale` tensors exported). It is now skipped **only** when the model is quantized and MTP is left out of it; plain distillation such as pruning recovery still trains the MTP head. `test_mtp_excluded_from_quantization` pins all four cases. - **One bad record deadlocked data prep.** `megatron_preprocess_data` re-raised chat-template failures out of a pool worker, stalling the whole job until it timed out — three malformed records cost a multi-hour tokenization run. They are now skipped with a warning, matching the existing handling of malformed JSONL a few lines above. Also exposes `--logit_kl_topk`, which `DistillationConfig` has supported for a while but the example never passed through; `test_qad` now exercises that path. §4, §5 and §6 are independent of §1–3; happy to split them out if reviewers prefer. ### Usage No API change. Exported names now match the released checkpoints: ``` model.language_model.layers.0.mlp.experts.<E>.{gate,up,down}_proj.{weight,weight_scale,weight_scale_2} lm_head.{weight,weight_scale,weight_scale_2} ``` ### Testing - `test_mcore_export_mappings.py` — qwen3_5 mappings emit per-expert rules. Verified these fail without the fix (2 failed / 11 passed), with `Qwen3MoeForCausalLM` / `NemotronHForCausalLM` as controls. - `test_unified_export_megatron.py` — the gate/up split, per-block scale slicing, the 0-dim scalar fallback, and both directions of the `_verify_exported_keys` relaxation. - `test_megatron.py::TestKeepGptOutputLayerExtraState` — 15 cases: payload detection, no-op second call, warn-and-skip on an unrecognised `sharded_state_dict`, and `test_patches_stock_megatron_core` which installs a replica of the real pre-fix upstream body (verified against `be08ce5b1~1`) so the patched path is exercised whichever megatron-core is installed. - `test_qad.py` — CI caught that its reference comparison still assumed packed experts; fixed. **End to end on `Qwen/Qwen3.6-35B-A3B` (35B MoE, 256 experts), 4×GB200, nemo:26.08:** | | before | after | | --- | --- | --- | | export self-check | `Export dropped 82 tensor(s)` | passes | | expert tensors | `mlp.experts.gate_up_proj` (packed) | `mlp.experts.<E>.{gate,up,down}_proj` | | **vLLM v0.28.0 load** | **`AttributeError`, engine never starts** | **`Loading weights took 25.61 s`** | | **NEL eval (GPQA-D, MMMU-Pro)** | **FAILED** | **SUCCESS** | ### Results these fixes unblocked The export fix is what made a Megatron-produced NVFP4 MoE checkpoint servable at all, so it enabled a full PTQ study on Qwen3.6-35B-A3B. Accuracy deltas are against a BF16 baseline measured on the same harness, from **paired** per-question tests: | recipe | throughput vs BF16 | GPQA-D | SciCode ×8 | MMMU-Pro | IFBench | | --- | --- | --- | --- | --- | --- | | **W4A16** (weight-only) | **0.64–0.86×** — *slower* | −0.06 | −0.15 | +0.48 | −0.44 | | **W4A4** | 8/12 shapes faster | −0.60 | −0.70 | −1.48 (p=0.019) | −0.53 | | **W4A4 + 4-bit `lm_head`** | **9/12 shapes**, up to **1.30×** | **+0.03** (p=0.96) | −0.81 | **−1.16** (p=0.016) | −1.65 (ns) | Repeats: GPQA-D is `pass@1[avg-of-16]`; SciCode is 8 pooled runs per recipe; MMMU-Pro is 3 runs per side and IFBench 2–3 for BF16 and the last row, 1 elsewhere. AA-LCR (68.33 → 71.33, p=0.25, 3 runs per side) and τ²-Telecom (94.25 → 94.25, 3 runs per side) are on par; at 100 questions and 114 tasks they cannot resolve below ~5 pp and ~3 pp, so they carry no claim either way. #### QAD status (what §6 unblocked) With the §6 fixes in place, QAD runs end to end on this model: 32 nodes, `TP=1 PP=1 CP=1 EP=8`, seq 32768, gbs 512, ~38 s/iter, 124 GB/GPU peak. First accuracy read, MMMU-Pro at iteration 50 (0.84 B tokens), 3 runs per side, paired per-question: | | MMMU-Pro | vs BF16 | | --- | --- | --- | | BF16 | 74.55 | — | | W4A4 + 4-bit `lm_head` (PTQ) | 73.39 | **−1.16, p=0.016** | | + QAD, iteration 50 | 73.78 | −0.77, p=0.089 (ns) | The PTQ deficit that motivated this work is no longer statistically significant after 50 QAD iterations. The improvement itself (+0.39 vs PTQ) is **not** significant at p=0.41, and 50 iterations is 10% of the planned budget, so this is a direction rather than a result. A full six-benchmark sweep at iterations 50 and 300 is running; these numbers will be superseded. Two findings worth flagging beyond this PR: - **Weight-only NVFP4 is slower than BF16 on Blackwell.** W4A16 leaves activations in BF16, so vLLM cannot use the FP4 tensor cores and falls back to `MarlinNvFp4LinearKernel` / `'MARLIN'` MoE. W4A4 selects `FLASHINFER_TRTLLM` + `FlashInferCuteDslNvFp4LinearKernel` and beats W4A16 in **12/12** shapes. The Marlin line count tracks the recipe exactly (one W4A16 layer ⇒ one Marlin line ⇒ zero once `lm_head` is W4A4). - **The only accuracy cost is multimodal**: **−1.2 pp on MMMU-Pro** for the fastest recipe, confirmed over 3 runs per side (p=0.016). GPQA-D, SciCode, IFBench, AA-LCR and τ²-Telecom show no significant regression. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ <!-- Megatron checkpoints unaffected; re-export to gain the loadable layout. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!-- 0.47.0 → Bug Fixes; includes the removed flag, since passing it now errors instead of being ignored --> - Did you get Claude approval on this PR?: ✅ <!-- Reviewed; all findings addressed, threads resolved. --> ### Additional Information Both export bugs were found while reproducing `nvidia/Qwen3.6-35B-A3B-NVFP4` through `examples/megatron_bridge/`. Follow-up to #2332. Upstream counterpart [NVIDIA/Megatron-LM#7086](https://github.com/NVIDIA/Megatron-LM/pull/7086) is closed — see §2. Labeled `cherry-pick-0.47.0`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2d35643452 |
LiLiCorr training (#2342)
### What does this PR do? Type of change: new feature Adds **LiLiCorr**, a candidate-lattice reranker for DFlash drafts, as a new `projector_type` on the existing `dflash` mode — plus three DFlash-wide improvements that apply to every variant, and an optional composition with DFlash2's grouped convolutions. A DFlash drafter is trained on per-position marginals rather than on the joint block distribution, so its drafted tokens are individually plausible yet jointly incoherent. LiLiCorr keeps the top-`k` candidates the backbone already produces at each block position, scores transitions between adjacent candidates with a small two-layer transformer, and commits a path through the lattice greedily. Serving is unchanged in kind: verify still checks every drafted token against the target, so the emitted distribution is untouched and only acceptance length moves. - Paper: [LiLiCorr: Lightweight Likelihood Correlation of Parallel Drafts for Speculative Decoding](https://arxiv.org/abs/2608.20530) (arXiv:2608.20530) - Blog: https://research.nvidia.com/labs/nemotron/lilicorr/ - **Companion PR — serving support:** [sgl-project/sglang#37462](https://github.com/sgl-project/sglang/pull/37462) This PR is the **training** half. It trains the drafters and exports them; the companion PR above is what serves the resulting checkpoints, and is what the comparison table below was measured through. **What is in the commits** | | | | --- | --- | | LiLiCorr draft variant | `hf_lilicorr.py`, `modeling_lilicorr.py`, conversion routing, config fields, export | | Three DFlash-wide features | fp32 master weights for the draft, draft activation checkpointing, and a DDP hang fix — all default-off or behaviour-preserving, all applying to `dflash`, `domino`, `dspark` and `dflash2` alike | | Optional grouped convolutions | composes LiLiCorr with DFlash2's `DFlashGroupedConv`; see the dependency note below | | Two recipes | `lilicorr.yaml` and `lilicorr_conv.yaml` | | CPU unit tests, CHANGELOG, one launcher example | | **⚠️ The convolutions depend on the DFlash2 branch, and cannot run until it merges.** `modeling_lilicorr.py` imports `DFlashGroupedConv` from `modeling_dflash2`, which today exists only on `haoguo/dflash2-support`. The class is **imported rather than copied on purpose** — it is the only way the two variants cannot drift apart arithmetically — but the consequence is that the convolutional recipe cannot run against `main` as it stands. So the import is **deferred into `_install_sublayer_convs`** rather than taken at module scope. Everything else in this PR, including the plain LiLiCorr reranker, has no DFlash2 dependency at all and works on `main` today; an eager import would have made the whole plugin unimportable for the sake of one optional feature. Requesting the convolutions without DFlash2 present raises an `ImportError` naming the two config keys to remove, rather than failing at import time. **This PR carries two of @h-guo18's commits, with authorship and sign-off preserved.** Both are independent of DFlash2 itself and both are needed here: - `1419d47e`, the no-op sublayer seam. Without it `DFlashDecoderLayer.forward` never calls the wrappers the convolutions install onto, so the modules would be built, counted and exported while computing nothing. It is arithmetically an identity on its own. - `ba377e7a`, the RoPE-θ fix. On Transformers 5 a config carries both a top-level `rope_theta` and a `rope_parameters` dict; the real base lives in the dict while the class default (10,000 for Qwen3) stays visible as the flat attribute. Reading the flat field first builds a draft whose RoPE base is 100× off a Qwen3-8B target's, which trains and exports without complaint. Both the training-side enforcement and the exporter's `_get_rope_theta` are affected on `main` today. Both are @h-guo18's work and belong to their branches; they are carried here only so that this PR stands on its own. **If those branches land first, this PR can be rebased onto them and the two commits dropped**, and they can equally be split out now if that is easier to review. The same applies to `dflash_fp32_master_weights`, which is also in flight on `haoguo/dflash-fp32-master-weights`. The field name is shared deliberately so that there is only ever one knob rather than two spellings of it, and both versions default to off. Whichever lands first, this PR can be rebased onto it. ### Usage Train with the shipped recipe: ```python from modelopt.recipe import load_recipe config = load_recipe("general/speculative_decoding/lilicorr.yaml") # Qwen3-8B target, 6 epochs, block size 16 (15 drafted slots, 16 verified), # DFlash decay objective at gamma 7.0, fp32 master weights for the draft. ``` Or convert directly: ```python import modelopt.torch.speculative as mtsp config = { "dflash_block_size": 16, "dflash_loss_objective": "decay", "dflash_loss_decay_factor": 7.0, "dflash_fp32_master_weights": True, "dflash_lilicorr_w_ce": 0.25, "dflash_lilicorr_w_margin": 0.0, "dflash_lilicorr_w_pen": 0.25, "dflash_architecture_config": { "num_hidden_layers": 5, "projector_type": "lilicorr", "lilicorr_candidate_topk": 8, # Optional, and all-or-nothing: adding these two keys wraps every draft # sublayer in DFlash2's grouped convolution. Requires the DFlash2 variant. # "conv_kernel_size": 2, # "conv_group_size": 16, }, } mtsp.convert(model, [("dflash", config)]) ``` ### Results Six drafters for a **Qwen3-8B** target, all trained **in ModelOpt on one matched contract** — the same corpus, schedule and block geometry for every arm, so no row carries a training advantage. Training data is NVIDIA's [Nemotron Post-Training Dataset v2](https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2) with the multilingual split excluded, generated from the target with **thinking disabled**; **6 epochs**; block size 16 (15 drafted slots, 16 verified); DFlash decay objective at gamma 7; **8 nodes × 8 H100, global batch size 64** (one sequence per device, no gradient accumulation). All six were then exported and served through SGLang on a **single H100 80GB**, `tp_size 1`, at concurrency 1, greedy, `fa3`, mean of two replicates, with the whole node held exclusive per benchmark. Speedup is output tokens/s against an autoregressive baseline measured in the same allocation. Cells are `acceptance length / speedup-vs-AR`; **★ fastest, ☆ second fastest**: | benchmark | LiLiCorr+conv | LiLiCorr | DSpark | DFlash2 | Domino | DFlash | |---|---|---|---|---|---|---| | gsm8k | ★ 7.715 / 5.26x | ☆ 7.557 / 5.22x | 7.375 / 4.86x | 7.252 / 5.06x | 7.225 / 4.87x | 6.341 / 4.59x | | math500 | ★ 9.241 / 6.54x | ☆ 9.064 / 6.52x | 9.012 / 6.15x | 8.999 / 6.49x | 8.976 / 6.25x | 7.909 / 5.88x | | aime25 | ★ 8.285 / 6.03x | ☆ 8.156 / 6.03x | 8.043 / 5.61x | 7.967 / 5.91x | 8.066 / 5.77x | 7.126 / 5.44x | | humaneval | ★ 7.393 / 4.01x | 7.077 / 3.93x | 7.163 / 3.72x | ☆ 7.081 / 3.95x | 6.864 / 3.73x | 6.156 / 3.68x | | mbpp_sanitized | ★ 5.999 / 4.18x | ☆ 5.849 / 4.13x | 5.888 / 3.95x | 5.685 / 4.05x | 5.679 / 3.91x | 5.027 / 3.70x | | livecodebench | ★ 7.975 / 5.40x | ☆ 7.754 / 5.33x | 7.775 / 5.10x | 7.601 / 5.26x | 7.553 / 5.04x | 6.808 / 4.88x | | alpaca_eval | ☆ 3.697 / 2.69x | ★ 3.656 / 2.70x | 3.588 / 2.52x | 3.467 / 2.58x | 3.627 / 2.59x | 3.222 / 2.46x | | mtbench | ★ 4.014 / 2.94x | ☆ 3.939 / 2.93x | 3.957 / 2.78x | 3.748 / 2.80x | 3.948 / 2.84x | 3.478 / 2.67x | **Against every other approach in the table, LiLiCorr with convolutions is the fastest on all eight benchmarks.** Plain LiLiCorr is the fastest on seven of the eight; the exception is humaneval, a 164-prompt slice, where DFlash2 is ahead by 0.5%. `DFlash` is the deliberately head-free control; every head clears it by +7.60% to +21.67% on acceptance, which is the check that a head actually loaded. Reproducing the `LiLiCorr+conv` column additionally needs the DFlash2 variant. Acceptance length is bit-reproducible under greedy decoding and its replicate spread here was 0.00% on every benchmark; throughput has a ~0.2% floor. ### What `dflash_fp32_master_weights` does, and what it is worth Today the draft is cast to the frozen base model's dtype — bf16 — before the optimizer is built. AdamW then allocates its moments with `zeros_like(p)`, so the **optimizer state becomes bf16 too**. That is the problem: bf16 has too few mantissa bits to represent the small updates Adam's second moment accumulates, so those updates round away and the effective step size decays on its own, independently of the learning-rate schedule. The flag is standard mixed precision instead: the draft's master weights stay in fp32 while the matmuls run in bf16. It requires a bf16 autocast around the forward, which HF `Trainer` supplies under `TrainingArguments.bf16`. Paths that do not go through the Trainer — evaluation, `pseudo_speculative_generate`, a plain `convert()` and forward — currently need the caller to supply it, and no shipped recipe exercises those (`estimate_ar: false`, `do_eval: false`). Making the draft supply its own autocast is a follow-up, held back from here on review because it touches every DFlash variant and wants e2e coverage of the existing recipes. Compute speed is unchanged. The cost is memory, about 12 bytes per parameter for the weight plus Adam's two moments instead of 6, plus a doubled gradient all-reduce under DDP, since fp32 parameters mean fp32 gradients. Under FSDP2 that second cost is what `MixedPrecisionPolicy(reduce_dtype=...)` exists to control. It is worth **7 to 14 percent of acceptance length**, measured at the end of training on gsm8k, and it helps every projector type: | arm | bf16 | fp32 | Δ acceptance length | | --- | ---: | ---: | ---: | | LiLiCorr | 6.8670 | 7.5573 | **+10.05%** | | DFlash2 | 6.7396 | 7.2518 | **+7.60%** | | Domino | 6.5854 | 7.2252 | **+9.71%** | | DSpark | 6.4621 | 7.3752 | **+14.13%** | | DFlash | 5.9030 | 6.3412 | **+7.42%** | Every arm in the comparison table above was trained with it on, and **both shipped recipes set it `true`**, so the documented path gets it. It defaults to **off**, so no existing DFlash, Domino or DSpark run changes behaviour. Both shipped LiLiCorr recipes set it `true`, which is the arithmetic their numbers were trained with. Flipping the default is a reasonable follow-up once the autocast above is in. The draft is drawn in fp32 and, under this flag, kept there; an unpromoted run rounds the same draw to the base model's dtype. So the bf16 and fp32 rows of the table above start from the same initialization at the precision each trains in, rather than from two different draws. A unit test pins that. The flag also survives a resume. `modify()` runs under `from_pretrained` with the base model still on meta and cannot place the draft at all, so `restore_draft_precision` re-applies the dtype, the device and the rotary buffer once the weights are loaded and before the Trainer builds the optimizer — the last point that can still decide the Adam moment dtype. It also reloads the draft's tensors at the dtype they were saved in, since checkpoints store the draft in fp32 while the base is bf16 and `dtype="auto"` gives every tensor one dtype. @h-guo18 has the same field in flight on `haoguo/dflash-fp32-master-weights`, plus an HF-format-resume fix this PR does not have. The name is shared deliberately so there is only ever one knob; whichever lands first, the other should be dropped rather than merged. ### Testing - **257 CPU unit tests pass** across `tests/unit/torch/speculative/`, including the existing DFlash, Domino, DSpark and Eagle suites. 48 of them are new and cover LiLiCorr specifically: conversion routing, head geometry, the required-field validation, the three-term objective and its absolute weights, gradient reach into both the head and the drafter body, and the export contract. - Both recipes load and validate through `modelopt.recipe.load_recipe`. - The three DFlash-wide changes are covered behaviourally: the fp32 flag is checked on the optimizer's moment dtypes rather than only on parameters, since the moments are the point of the change, and on the initialization described above; activation checkpointing is asserted to leave draft gradients bit-identical with the flag on and off; and the rotary buffer is asserted present after `modify()` on a real device while still deferred on meta, which is the case the laziness existed for. - The resume path has its own test: after a `save_pretrained` / `from_pretrained` round trip, `restore_draft_precision` is asserted to return the draft to fp32 with its stored weights intact and its Adam moments in fp32. Without it the draft comes back in the base dtype with the flag still set, which is the failure it exists to prevent. - `TestDFlashLazyRotaryEmb` was updated rather than left passing: it asserted the rotary buffer does *not* exist after convert, and the DDP fix deliberately changes that on non-meta devices. The replacement pins the refined invariant in both directions. - The published checkpoints were trained with this arithmetic, verified rather than assumed: a fingerprint over draft initialisation, loss and gradients is compared against the pre-review tree for both `dflash` and `lilicorr`. Loss and gradients are **bitwise identical**. Initialisation moves, by less than bf16 resolution, and that is the single-dtype change described above. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — every addition is opt-in. The new `projector_type` is selected only by config, `dflash_fp32_master_weights` defaults to off, and the activation-checkpointing and DDP fixes preserve behaviour. No existing default changes. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new dependencies. Four files carry `# Adapted from https://github.com/sgl-project/SpecForge/...` headers for the DFlash backbone and loss they derive from (Apache-2.0), matching the attribution already on `hf_dflash.py` in this repo. The two commits described above are @h-guo18's, cherry-picked with authorship and sign-off preserved. - Did you write any new necessary tests?: ✅ — 48 new CPU tests, plus the updated rotary test. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ — will run `/claude review` once opened. ### Additional Information The convolutional recipe is the memory worst case: at an 8B target, combined with fp32 master weights, it may need `training.gradient_checkpointing: true` to fit on 80 GiB, and it fits without at 4B. Checkpointing is mathematically neutral — same objective, same data order, same resulting model — but it trades step time for memory, so a run using it is not step-time-comparable with one that does not. The recipe header says so. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added LiLiCorr speculative decoding with candidate-lattice reranking, configurable objectives, metrics, export support, and optional grouped convolutions. * Added FP32 master-weight support with improved mixed-precision behavior and gradient checkpointing. * Added LiLiCorr training recipes and a Qwen3-8B launcher configuration. * **Bug Fixes** * Improved rotary-embedding configuration handling and corrected DFlash distributed-training hangs. * Added validation for invalid LiLiCorr configurations and improved exported reranking metadata. * **Documentation** * Expanded guidance for FP32 master weights, training workflows, and LiLiCorr configuration. * **Tests** * Expanded coverage across training, evaluation, generation, export, and checkpoint workflows. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: mrusanovsky <mrusanovsky@nvidia.com> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
acdf330414 |
Add TensorRT-RTX ABI EP support for ONNX quantization (#2262)
### What does this PR do? Type of change: new feature Adds opt-in support for using the standalone TensorRT-RTX ABI Execution Provider during ModelOpt ONNX quantization. Users select the ABI backend with: `--calibration_eps=NvTensorRtRtx --trt_rtx_backend=abi` When selected, ModelOpt imports and registers the installed TensorRT-RTX ABI provider before creating the ONNX Runtime inference session. The backend selection is propagated through INT8, FP8, and INT4 AWQ calibration paths, including the Windows GenAI LLM quantization example. The existing `--calibration_eps=NvTensorRtRtx` behavior remains backward compatible. The `legacy` backend is still the default and continues to use TensorRT-RTX libraries supplied through `PATH`. For Windows x64 with Python 3.11 or newer, the ONNX dependencies now include: - `onnxruntime-gpu~=1.26.0` - `onnxruntime-ep-nv-tensorrt-rtx-cu13==0.4.0` Keeping `onnxruntime-gpu` allows users to select either CUDA EP or TensorRT-RTX ABI EP for calibration. Windows-on-Arm source-build instructions are intentionally out of scope and will be documented separately. ### Usage ```powershell python -m modelopt.onnx.quantization ` --onnx_path="C:\path\to\Llama-3.2-3B-Instruct\model.onnx" ` --model_id="C:\path\to\Llama-3.2-3B-Instruct\config.json" ` --quantize_mode=int8 ` --output_path="C:\path\to\int8_abi\model.onnx" ` --calibration_eps=NvTensorRtRtx ` --trt_rtx_backend=abi ` --use_external_data_format ` --high_precision_dtype=fp32 ` --log_level=INFO ### Testing unit test have been added ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: pending <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added optional TensorRT-RTX ABI backend support for ONNX calibration on Windows ARM64. - Added `legacy` and `abi` backend selection to quantization APIs and command-line tools; `legacy` remains the default. - Added validation for unsupported backends and incompatible TensorRT plugin configurations. - Updated Windows ARM64 installation support and platform-specific package configuration. - **Documentation** - Updated Windows installation guidance, Python compatibility requirements, ARM64 setup, and verification steps. - Documented the new TensorRT-RTX backend command-line option. - **Tests** - Added coverage for ABI provider registration, backend validation, and compatibility checks. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Haoxi Zhang <haoxiz@nvidia.com> |
||
|
|
0688761ce9 |
feat(export): support multimodal and MTP models in layerwise export (#2303)
### What does this PR do? Type of change: New feature **Layerwise export now supports multimodal and MTP models.** Both were refused outright, and both were refused for the same reason: `finalize()` was called from inside `layerwise_calibrate`, which is the wrong scope for it. **1. Calibration does not know which model the checkpoint describes.** It only sees the module it was handed. A VLM calibrates its *language model*, so the shards, the exclusions and `config.json` all came out describing that submodel rather than the whole VLM. Moving the call out lets the caller root the exporter at the parent — and without the key prefixing, tower collection or ambient parent handle an earlier attempt needed, because the decoder layers are the same objects from either root. **2. Calibration runs before things the export needs exist.** Orphaned MTP weights are loaded *after* calibration, by which point every shard had already been written, so they could not be passed at all. After the move they are an ordinary argument to `finalize()`, with no staging attribute stashed on the model. ### How it works The exporter is created by whoever owns the export and **announced on the model** that `mtq.quantize` is given. Calibration picks it up, binds it, and drives it per layer; the export that follows reads it back and finishes the checkpoint: ```python LayerwiseExporter(full_model, export_path).announce(language_model) mtq.quantize(language_model, quant_cfg, forward_loop=loop) ... getattr(full_model, LAYERWISE_EXPORTER_ATTR).finalize(extra_state_dict=mtp_state_dict) ``` Calibration and export are handed *different* models, so `announce()` publishes the exporter on each end separately: the caller announces on the model being calibrated, and `bind()` announces on the export root. Neither side has to know where the other looked, and the lookup stays an O(1) `getattr` rather than a `named_modules()` scan — worth avoiding at roughly 1.65 µs/module, or ~500 ms on a Kimi-K3-sized model. For a non-VLM both roots are the same object and the second announcement is a no-op. `finalize()` clears every attachment it recorded, so the module graph does not retain a live exporter afterwards. `mtq.quantize` and `mtq.calibrate` are **unchanged** — a layerwise-only feature does not belong in the public quantization API. The attribute follows `_mtp_layer_prefixes`, which crosses the same calibration→export boundary the same way (`hf_ptq.py:538` sets it, `unified_export_hf.py:870` reads it back). Construction is inert: `__init__` records only the export root and the directory, because the caller builds it before `mtq.quantize`, when there are no quantizers yet to validate or read a config from. `bind()` does that, called from calibration after quantizer insertion and before any layer is converted — the only window where both hold, and the same instant the exporter used to be constructed, so unsupported models still fail in seconds rather than hours. Only the calibration pass that sets `export_dir` drives the exporter: a list-form algorithm runs one pass per entry, and an earlier one must not convert layers a later one still has to calibrate. ### Usage Nothing changes for a plain layerwise-export recipe: `layerwise.export_dir` still drives it. Pre-attaching an exporter is the opt-in for the two cases that need it — a checkpoint whose root is wider than the calibrated model, and orphaned tensors to merge at the end. The one behaviour change for a config-only caller is that `mtq.quantize` now writes the layer shards but no longer finishes the checkpoint. Both exit paths warn with what is still owed, and `LayerwiseConfig.export_dir`'s description has been corrected — it previously promised "a complete, loadable checkpoint when the last layer lands" and still listed multimodal and MTP as raising `NotImplementedError`. ### Testing `tests/gpu/torch/export/test_layerwise_export.py` — **29 passed**. Beyond the 24 inherited from #2136, five new ones, each with a negative control confirming it fails without its fix: - orphaned MTP tensors reach the tail shard *and* the index - an exporter rooted at the parent widens the checkpoint's namespace - the config-only path announces an exporter that can be finished, and finalize clears it - only the pass that sets `export_dir` drives the exporter - an exporter whose root holds a different number of layers is refused at `bind()` Full suites: `tests/gpu/torch/export` + `tests/gpu/torch/quantization` **1012 passed / 55 skipped**, `tests/unit` **3318 passed / 15 skipped**, pre-commit clean. Both suites also report failures in `test_implicit_gemm.py` (FP4 conv kernels), `test_triton_fa_p_qdq.py`, `test_autocast_quantize_int8` and `test_engine_builder.py` collection; all reproduce unchanged on `main` and none touch the paths in this diff. Measured against the whole-model exporter on a tiny Gemma3-VL, towers prepared exactly as `hf_ptq` does: ``` keys: baseline=80 layerwise=80 only-baseline=[] only-layerwise=[] differing values: 0 vision tower present: True VLM namespace: True config.json is the VLM: True hf_quant_config match: True exclude_modules: ['language_model.lm_head', 'vision_tower.vision_model*'] (both sides) ``` #### End-to-end through `hf_ptq.py` Same FP8 recipe both sides; the baseline drops `layerwise.export_dir` and is exported by `main`, so the diff isolates this PR. Every tensor matches in key, dtype, shape and value, and `config.json` / `hf_quant_config.json` match too. | Model | Covers | Keys | Differing | |---|---|---|---| | Qwen3-VL-8B-Instruct | multimodal | 1254 = 1254 | 0 | | GLM-4.7-Flash | MoE + MTP | 28119 = 28119 | 0 | The VLM checkpoint keeps the vision tower unquantized (351 `model.visual.*` keys, no `weight_scale` among them) while the language model is FP8. The MTP run reports 212 orphaned tensors; all 212 land in `model-tail.safetensors` and in the index, with `model.layers.47*` in `exclude_modules`. **Not yet validated:** an accelerate-offloaded run, and a serving canary on the exported checkpoints. ### Refusals `export_dir` without `enable`, and an exporting algorithm entry with no calibration method, are both refused before calibration starts — neither reaches the per-layer pass, so both would otherwise export nothing. The early gate is a heuristic on the recipe, so `hf_ptq` also raises a plain `RuntimeError` at export time if calibration turned out not to have run; that backstop, not the gate, is what makes the failure legible on paths the recipe check cannot predict. `bind()` requires the layers calibration will drive and refuses a root that discovers a different number of them. Only the count is checked here: `export_layer` already rejects a reordering or a substituted module on its first call, and a length difference is the one mismatch it structurally cannot catch — every call would pass and `_write_index` would then open a shard that was never written, at the very end of the run. Orphan tensors are merged into the tail with no collision check, matching the whole-model path (`unified_export_hf.py:1623`). `load_mtp_weights` returns exactly the keys absent from `model.state_dict()`, so a collision with an exported tensor is not reachable through the only producer, and a guard would only make the two export paths diverge. ### Why not reuse `export_hf_checkpoint` It was the first idea and it is the most expensive one. Its transformers path is whole-model at every step — `_prepare_moe_inputs`, `requantize_resmooth_fused_llm_layers` (which runs a dummy forward that would fail on already-converted layers), `_process_quantized_modules`, a full `model.state_dict()` in host RAM, then `save_pretrained` rewriting shards already on disk — and it raises outright under `has_accelerate_offload`. `save_pretrained(state_dict={})` is not an escape either: safetensors' shared-storage check fires on MoE even with an empty dict. The natural consolidation target is the **streaming** exporter, which is already most of `finalize()`: 122 lines vs 74, sharing `decoder_owned_ids`, `enable_weight_access_and_writeback`, `_dispatch_export_handler`, `_reconstruct_fused_moe_linear`, `_add_mtp_exclusions`, `_postprocess_single_tensor`, `requires_weight_materialization` and `save_non_weight_artifacts`. Folding them together needs roughly four knobs: skip the whole-model prep, skip layers already written, seed the index with the existing shards, and inject the quant config. That is a separate change and deliberately not in this one. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `mtq.quantize`/`mtq.calibrate` signatures are unchanged, and a recipe that only sets `layerwise.export_dir` behaves as before. The one behaviour change is that `mtq.quantize` no longer finishes the checkpoint on its own: callers must now call `finalize()` on the exporter, which calibration leaves on the model. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ❌ — pending. - Did you get Claude approval on this PR?: ❌ — the last review's findings are all addressed; needs a re-run. ### Additional Information Follow-ups this enables: #2259 (MTP) reduces to close to nothing, and the multimodal work in #2218 no longer needs `export_parent`, the key prefixing, or the tower collection. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Layerwise export now supports clearer control over export locations and calibrated layer handling. - Export workflows provide improved support for resuming, sharded checkpoints, mixture-of-experts models, and nested model namespaces. - **Bug Fixes** - Improved handling of exported checkpoint shards and extra tensors. - Added clearer warnings when exports require completion before loading. - **Documentation** - Clarified that layerwise exports write shards during calibration and require an explicit finalization step. - Documented that the in-memory model is not suitable for inference after layerwise export. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
19de0075cb |
Forward kv_cache_free_gpu_memory_fraction to the lm_eval TensorRT-LLM engine (NVBug 6701763) (#2300)
### What does this PR do? Type of change: Bug fix `scripts/huggingface_example.sh --kv_cache_free_gpu_memory_fraction` has no effect on the `lm_eval` task: the value is parsed by `parser.sh`, printed, and then dropped. lm-eval's built-in `trtllm` backend (`lm_eval.models.trtllm_causallms.TRTLLM.__init__`, which this example switched to in #2066) accepts `**kwargs`, but builds `KvCacheConfig(enable_block_reuse=False)` and passes `LLM(...)` a fixed set of keys — `kwargs` is never merged in. So an extra `--model_args` entry is accepted by the CLI and silently discarded, and the KV cache is sized from TensorRT-LLM's default `free_gpu_memory_fraction=0.9`. There is no way to fix this from the caller: `--model_args` only yields scalars, so a `KvCacheConfig` object cannot be passed in either. On a GH200 that means ~119.6 GiB of KV cache (`119.55 / 0.9 ≈ 132.8 GiB free`), leaving 87.8 MiB free, and `prompt_logprobs` deserialization then OOMs asking for 2.82 GiB. `examples/llm_eval/lm_eval_trtllm.py` already exists to patch this backend (its `_parse_logprobs` misaligns TensorRT-LLM's `prompt_logprobs` by one). It now also injects the fraction into the `KvCacheConfig` the backend builds, defaulting to 0.8 — the same default `parser.sh` declares, and below TensorRT-LLM's 0.9. `huggingface_example.sh` passes the parsed value through in `--model_args`. Scoped deliberately to the `lm_eval` path: the `quant` smoke test and `mmlu` go through `modelopt.deploy.llm.LLM` (0.7, hardcoded) and `simple_eval`/`livecodebench` through `trtllm-serve` (0.9); those are left as they are. ### Usage ```bash # Via the example script (parser.sh default 0.8) scripts/huggingface_example.sh --model $HF_PATH --quant fp8 --tp 1 \ --tasks quant,lm_eval --lm_eval_tasks mmlu --lm_eval_limit 50 \ --kv_cache_free_gpu_memory_fraction 0.5 ``` ```bash # Standalone, via lm-eval's --model_args python lm_eval_trtllm.py --model trtllm \ --model_args model=<ckpt>,tokenizer=<tok>,max_input_len=4096,kv_cache_free_gpu_memory_fraction=0.5 \ --tasks mmlu --batch_size 8 ``` ### Testing - `pytest tests/examples/llm_eval/test_lm_eval_trtllm.py` — 21 passed (lm-eval 0.4.12, no GPU). - The new tests instantiate the **real** upstream `TRTLLM.__init__` through `create_from_arg_obj`, with `tensorrt_llm` and the tokenizer stubbed, and assert the engine receives `KvCacheConfig(enable_block_reuse=False, free_gpu_memory_fraction=0.5)`; that an unset key still yields 0.8 rather than 0.9; and that the patch does not outlive the constructor. Reverting the fix fails 3 of them. - Tripwire test asserts upstream still neither declares nor forwards the argument, so this shim gets deleted rather than silently kept once lm-eval fixes it. - `pre-commit run --files <changed>` clean (ruff, mypy, bandit, markdownlint); `bash -n` on the modified script. - Not run: the GPU end-to-end `tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8`, which exercises `lm_eval` through the modified script — no GPU in this environment. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — the `lm_eval` KV cache goes from TensorRT-LLM's 0.9 to 0.8, which is strictly more conservative; `parser.sh`'s declared default is unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information NVBug 6701763. The 0.9 default on this path arrived with #2066 and was documented as a known limitation in `examples/llm_eval/README.md` ("the KV cache uses 90% of free GPU memory rather than 70%"); that note is replaced by the working knob. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Fixed the TensorRT-LLM evaluation workflow so `kv_cache_free_gpu_memory_fraction` is correctly passed to the backend. - The setting now defaults to `0.8`, providing more predictable GPU memory allocation for KV-cache usage. - **Documentation** - Updated the TensorRT-LLM evaluation example and usage guidance to describe the KV-cache memory setting and its default behavior. - Updated the Hugging Face example to pass the configured KV-cache memory fraction. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2aefe08f20 |
[OMNIML-5613] Quantize ResNet residual adds in torch ONNX example (#2024)
### What does this PR do?
Type of change: Bug fix
Adds recipe-backed FP8 and INT8 residual quantization for timm ResNet
models in the torch ONNX example:
- Adds FP8 and INT8 PTQ recipes that enable a shortcut quantizer
immediately before each residual `Add`.
- Inserts shortcut quantizers after ModelOpt module conversion so
recipes configure and calibrate them in the normal quantization pass.
- Adds `--recipe` support for PTQ and AutoQuantize recipes and renames
`--quantize_mode` to `--qformat`.
- Verifies all 16 ResNet-50 residual additions have shortcut Q/DQ
immediately before the `Add`.
### ResNet support scope
ResNet and other convolutional architectures are supported only with FP8
and INT8. AutoQuantize, MXFP8, NVFP4, and INT4_AWQ are not supported for
ResNet because TensorRT has limited convolution kernel support.
Transformer architectures containing individual Conv2d layers continue
to use format-specific Conv overrides.
### Usage
```bash
python examples/torch_onnx/torch_quant_to_onnx.py \
--timm_model_name=resnet50 \
--recipe=timm/resnet/ptq/fp8 \
--onnx_save_path=resnet50.onnx
```
Use `timm/resnet/ptq/int8` for INT8. Without `--recipe`, `--qformat`
selects a built-in quantization preset.
### Testing
- All configured pre-commit hooks passed, including recipe schema and
license validation.
- Focused AutoQuantize recipe mapping regression passed.
- FP8/INT8 recipe export coverage verifies all 16 ResNet-50 shortcut
Q/DQ pairs.
- TensorRT engine builds passed for FP8 and INT8 on Ada.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update the changelog?: ✅
> 🤖 _Generated by Codex (AI agent)._
---------
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <noreply@openai.com>
|
||
|
|
58eafdf172 |
Fix the llm_eval README commands that no longer run as written (#2358)
### What does this PR do?
Type of change: Documentation (plus one small example-script fix)
An audit of `examples/llm_eval/README.md` against the current scripts
(nvbug 6701343) found several documented commands that no longer run as
written:
- **T5 / seq2seq.** `--model hf-seq2seq` is not a registered lm-eval
backend in any version this example supports — the string does not
appear in the 0.4.12 or 0.4.13 wheels, so the command fails at model
lookup. `HFLM` detects encoder-decoder models from `config.json`, so the
example now uses `--model hf` and mentions `backend=seq2seq` as the
override for checkpoints lm-eval cannot classify. No ModelOpt-side
change was needed: encoder-decoder calibration already works (verified
below).
- **auto_quantize format list.** `FP8_DEFAULT_CFG|NVFP4_DEFAULT_CFG` was
shown as a literal value in both README locations, but each
comma-separated entry is resolved with `getattr(mtq, ...)` and that name
does not exist. Now shows a valid list, spells out the choices, and
names the placeholder consistently with the surrounding block.
- **`vllm serve`.** A missing line continuation meant `--port` ran as a
separate shell command.
- **MMLU setup.** Dropped a stray `cd ..` left over from the 0.11
examples release. It leaves `examples/llm_eval`, where both `mmlu.py`
and its default `--data_dir data/mmlu` live;
`hf_ptq/scripts/huggingface_example.sh` correctly stays put throughout
its MMLU flow, so the README was the only thing out of step.
- **`run_simple_eval.sh`.** Documented the optional fifth argument
(`--examples`), which `huggingface_example.sh` already passes as
`$SIMPLE_EVAL_LIMIT`.
Two changes beyond the docs:
- **`quantization_utils.py`:** under `auto_quantize`, a `quant_cfg`
string was iterated character by character, so a single format failed
with the baffling `AttributeError: module 'modelopt.torch.quantization'
has no attribute 'F'`. Normalized `str -> list` at the point the list is
consumed, which covers both `mmlu.py` and `lm_eval_hf.py` rather than
one caller. This also honors the existing `str | list[str]` annotation.
- **`requirements.txt`:** added the missing `openai`. `modeling.py`
imports it unconditionally and `lm_eval[api]` supplies only `tiktoken`,
so every documented `mmlu.py` command died with `ModuleNotFoundError` on
a clean install of the stated requirements.
Note on the filed report: its item 3 claimed `mmlu.py` fails to split
the comma-separated config list. That does not reproduce — `mmlu.py`
uses `fire`, which already parses `A,B,NONE` into a tuple, and the
unmodified script completes `auto_quantize` fine. Applying the suggested
`quant_cfg.split(",")` would have *broken* the documented command with
`AttributeError: 'tuple' object has no attribute 'split'`. The
`quantization_utils.py` change above addresses the real adjacent defect
instead. Pushback recorded on the bug.
### Usage
No new API or flag. The corrected commands:
```bash
# T5 / encoder-decoder (was: --model hf-seq2seq, which does not exist)
python lm_eval_hf.py --model hf --model_args pretrained=t5-small \
--quant_cfg FP8_DEFAULT_CFG --tasks <comma separated tasks> --batch_size 4
# auto_quantize search list (was: W4A8_AWQ_BETA_CFG,FP8_DEFAULT_CFG|NVFP4_DEFAULT_CFG,NONE)
python mmlu.py --model_name causal --model_path <model> \
--quant_cfg W4A8_AWQ_BETA_CFG,FP8_DEFAULT_CFG,NONE --auto_quantize_bits 4.8 --batch_size 4
# simple evals, optional 5th arg
bash run_simple_eval.sh <model> <evals> <max_tokens> <port> [num examples per eval]
```
### Testing
Ran on 2x RTX 6000 Ada with a tiny Qwen3 and a locally synthesized MMLU
tree (no download):
- **`mmlu.py --auto_quantize_bits` with the documented comma-separated
list** — completes quantization on both the unpatched and patched
script, confirming the reported item 3 is a false positive. Probed
`fire` directly: bare, quoted and `--flag=value` forms all yield
`('W4A8_AWQ_BETA_CFG', 'FP8_DEFAULT_CFG', 'NONE')`.
- **`mmlu.py --auto_quantize_bits` with a single format** — proved the
new guard fires by reverting it: without the change the run dies with
`AttributeError: module 'modelopt.torch.quantization' has no attribute
'F'`; with it, the run reaches a legitimate domain assertion
(`effective_bits 4.8` cannot be below FP8's 8 bits).
- **Encoder-decoder calibration** — quantized a T5 with
`FP8_DEFAULT_CFG` through `quantize_model` and confirmed encoder,
decoder and cross-attention (`EncDecAttention`) layers all calibrate
with real amax values. This is what settled keeping the T5 example
rather than deleting it.
- **`vllm serve` snippet** — parsed the fixed block with `bash`;
`--quantization`, `--port` and `--tensor-parallel-size` now all belong
to one command.
- **`run_simple_eval.sh`** — confirmed the 4-arg form is unchanged and
the 5-arg form emits `--examples 16`.
- **Lint** — `ruff-check`, `ruff-format`, `markdownlint-cli2`, `typos`,
`bandit`, `mypy`, `requirements-txt-fixer`, `mixed-line-ending` all
pass.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — added
`openai` to `examples/llm_eval/requirements.txt`; it is Apache 2.0
(permissive), so no codeowners exception is needed. It is not a new
runtime dependency of the library, and `run_simple_eval.sh` already `pip
install`s it.
- Did you write any new necessary tests?: N/A — docs plus a two-line
defensive normalization in an example util. `mmlu.py` cannot be imported
without `openai`/`rwkv`/`tiktoken`, so a hermetic unit test would need
more stub scaffolding than the line it guards; verified by direct
execution instead, as above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — examples-only documentation cleanup, not a feature, breaking
change, deprecation, or a critical bug from a previous release.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
Fixes nvbug 6701343 / OMNIML-5806. Item 3 of the filed report is a false
positive; pushback and evidence are recorded in a comment on the bug.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Auto-quantization now supports comma-separated format configurations.
- Added an optional example-limit setting for Simple Evals.
- Added OpenAI support for LLM evaluation examples.
- **Documentation**
- Clarified encoder-decoder model usage with `lm_eval`.
- Added instructions for running MMLU from the evaluation examples
directory.
- Corrected the vLLM command formatting.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
5c123ce183 |
[OMNIML-5563] Add PETR ONNX PTQ and accuracy evaluation example (#2180)
### What does this PR do? Type of change: new example, example simplification, and backward-breaking example migration Adds end-to-end PETRv1/PETRv2 ONNX PTQ and reduces PETR/FAR3D to one shared workflow: - quantizes the shared VoVNet image backbone/encoder to INT8 or FP8; - runs both the selected historical and current PETRv2 six-camera sweeps through the same precision-matched TensorRT backbone engine using distinct execution contexts during accuracy evaluation; - keeps the PETR head and FAR3D decoder in their exported mixed FP16/FP32 precision; - reuses one NPZ calibration format, VoVNet exclusion helper, quantization entry point, and TensorRT runner; - does not change generic Model Optimizer calibration behavior or its public CLI. ### Container boundary Both examples use two targets from one Dockerfile, with no virtual environments: - `evaluator`: a digest-pinned `nvcr.io/nvidia/pytorch:22.06-py3` base with the legacy PyTorch 1.13.1/OpenMMLab stack for source setup, metadata generation, ONNX export, direct PyTorch calibration capture, and final accuracy evaluation; - `modelopt`: a digest-pinned `nvcr.io/nvidia/pytorch:26.07-py3` base for Model Optimizer, ONNX Runtime CUDA, AutoCast, INT8/FP8 quantization, and TensorRT engine builds. Both targets use TensorRT `11.1.0.106`. Engines are built and evaluated on the same GPU architecture. Final metrics remain in the evaluator because they import the legacy model-framework postprocessing and dataset code; only artifacts cross the container boundary through the shared workspace. PETR is used without patches. FAR3D applies only the official `patch/far3d.patch` from the pinned NVIDIA DL4AGX revision. This PR carries no patch files. ### Evaluator dependencies The dependencies intentionally installed without transitive dependencies are listed in `requirements-evaluator-nodeps.txt`. Their pins rely on runtime packages supplied by the digest-pinned PyTorch 22.06 evaluator base. `lyft-dataset-sdk` is required only by mmdet3d's eager dataset import; neither PETR nor FAR3D uses Lyft data. `flash-attn` remains in the main evaluator requirements because its compiled installation uses the evaluator build step rather than the intentionally dependency-free legacy package step. Fresh setup and dependency approval is requested for the final reduced dependency set. ### Reproducible PETR metadata The documented workflow mounts raw nuScenes read-only and creates a writable dataset view using symlinks. It then runs the pinned mmdetection3d converter and a temporary, untracked copy of PETR's pinned sweep generator configured only for the validation prefix and writable dataset root. A clean run generated both metadata files with 6,019 validation records. The referenced camera, lidar, and sweep paths are absolute and resolvable through the writable dataset view. ### Example-local utilities The per-batch NPZ streaming and TensorRT runtime utilities remain example-local because they execute in the legacy evaluator, where Model Optimizer is not installed. The core `CalibrationDataProvider` consumes one in-memory mapping of stacked arrays and does not provide this streamed per-file workflow. ### Validation - Focused CPU tests: 10 passed. - Broader ONNX quantization CPU tests: 326 passed. - All applicable pre-commit and documentation checks, plus `git diff --check`, passed. - Rebuilt both Docker targets and verified their exact dependency versions, imports, TensorRT `11.1.0.106`, GPU runtime initialization, and absence of virtual environments. - Generated both PETR metadata files from a clean writable dataset view and verified 6,019 validation records plus resolvable data paths. - PETRv1 passed a one-sample TensorRT regression smoke. - PETRv2 passed FP16, INT8, and FP8 TensorRT smokes and full 6,019-sample validation. Both the selected historical and current sweeps are computed by the matching backbone engine; accuracy evaluation no longer extracts image features with PyTorch. - FAR3D passed a recurrent two-frame TensorRT smoke covering plugin loading and recurrent state. TensorRT `11.1.0.106` mAP follows. PETRv2 was remeasured after correcting its temporal feature path; the PETRv1 and FAR3D numerical paths are unchanged. | Pipeline | FP16 | INT8 | FP8 | | --- | ---: | ---: | ---: | | PETRv1: 1 backbone pass + fixed typed mixed FP16/FP32 head | 0.3778 | 0.3707 | 0.3756 | | PETRv2: 2 serial backbone passes + fixed typed mixed FP16/FP32 head | 0.4102 | 0.3982 | 0.4084 | | FAR3D: 1 encoder pass + fixed mixed FP16/FP32 decoder | 0.241 | 0.235 | 0.239 | Normalized engine-only performance improvement over each matching FP16 pipeline: | Pipeline | INT8 speedup | FP8 speedup | | --- | ---: | ---: | | PETRv1 | 1.49x | 1.29x | | PETRv2 | 1.51x | 1.30x | | FAR3D | 1.69x | 1.40x | Performance was measured with TensorRT `11.1.0.106` on an NVIDIA RTX 6000 Ada Generation GPU using five interleaved trials per engine component. Each component uses the median `trtexec`-reported GPU Compute Time with data transfers disabled and CUDA Graphs enabled. Component times are summed before normalization: PETRv1 uses one backbone pass plus its fixed head, PETRv2 uses two serial backbone passes plus its fixed head with no temporal cache assumed, and FAR3D uses one encoder pass plus its fixed decoder. Absolute latency values are intentionally not published. Adapted files retain exact public-source references and upstream notices, and the top-level license attribution is updated. - Is this change backward compatible?: ❌ - Did you write the necessary tests?: ✅ - Did you update the changelog?: ✅ > 🤖 _Generated by Codex (AI agent)._ --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com> |
||
|
|
4773f72f8a |
Docs: Add WOA documentation (#2264)
### What does this PR do? Add WoA env setup guide. Includes build instruction of pyarrow, which used by datatsets ### Usage N/A ### Testing N/A ### Before your PR is "*Ready for review*" - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Documentation** - Added comprehensive Windows on Arm installation guidance, including prerequisites, environment setup, dependency installation, verification, and troubleshooting. - Documented experimental Windows ARM64 support, supported quantization formats, native dependency requirements, and Support Matrix details. - Expanded supported Windows Python versions through 3.13. - Expanded TensorRT-RTX guidance for calibration, deployment, provider setup, and standalone plugin usage. - Clarified PyArrow requirements and local build instructions for Windows ARM64. - Added links to dedicated Windows on Arm installation resources and shared TensorRT-RTX documentation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Haoxi Zhang <haoxiz@nvidia.com> |
||
|
|
61757c9781 |
Support quantized Qwen3-VL / Qwen3.5-VL (dense + MoE) export from Megatron-Bridge and verify exported checkpoints (#2276)
### What does this PR do?
Type of change: Bug fix + new feature
Enables quantized **Qwen3-VL** and **Qwen3.5-VL** (dense and MoE) →
unified HuggingFace export from Megatron-Bridge, and fixes the bugs
found along the way (ten from testing, plus a further round from
review). Most of them produced a valid-looking checkpoint and a green
test run, so the PR also makes the export path verify its own output.
Review is easiest commit-by-commit — each of the eleven commits is
self-contained and independently green.
#### Two blockers
1. **The exporter rejected the Megatron-Bridge VLM wrapper.**
`GPTModelExporter` only unwrapped MCore's `LLaVAModel`, so
`Qwen3VLModel` raised `ValueError: Input to GPTModelExport must be a
megatron.core.models.GPTModel!`. It now unwraps any wrapper exposing
`.language_model`.
2. **A VLM QAD checkpoint couldn't be loaded back.** `distill.py` passes
`distill_submodule="language_model"`, so the checkpoint holds only the
language model and the load died on `KeyError:
vision_model.patch_embed.proj.weight`. The loader now reads the
checkpoint metadata and targets `.language_model` when there are no
vision weights.
#### Four silent-corruption bugs
3. **VLM QAD discarded all ModelOpt state** (shipped in 0.46).
`ModeloptStateManager` requires state on the **root** of whatever gets
checkpointed. `quantize.py` quantizes the VLM root, so PTQ anchors it
there — but QAD checkpoints only `language_model`, orphaning it. The
saved `modelopt_state_dict` was literally `[]`; the `*_quantizer._amax`
tensors were still present but got dropped on load
(`dist_ckpt_strictness="assume_ok_unexpected"`), and the export came out
plain BF16 with no `hf_quant_config.json`.
4. **Fused grouped-GEMM MoE experts were omitted entirely.** The MoE
dispatch had no `else`, so an architecture without an
`experts.linear_fc1` rule exported *zero routed experts*. This hit
**`Qwen3MoeForCausalLM`** — a registered, supported architecture with no
export test — not just VLMs. A tiny Qwen3-MoE exported 37 of 45 tensors,
exit 0, no warning.
5. **Qwen3.5's GatedDeltaNet output norm was off by exactly 1.0.**
Megatron stores that gamma zero-centered, HF centers it on 1. Correct
names, correct shapes, wrong values — invisible to any structural check.
Megatron-Bridge's importer confirms the convention
(`RMSNorm2ZeroCenteredRMSNormMapping`).
6. **The disabled-quantizer patterns silently no-op on Megatron paths.**
They are written against HuggingFace module names. `*mixer.conv1d*`
matches only because MCore and HF happen to agree on "mixer" for Mamba;
`*linear_attn.conv1d*` never matched (Megatron calls it
`self_attention.conv1d`), so the conv1d was calibrated.
`*linear_attn.in_proj_a/b*` **cannot** match at all — Megatron fuses all
six GDN sections behind one quantizer — so the alpha/beta gates the
recipe wants in BF16 were exported in FP8.
#### Four more bugs, found only by running real checkpoints
The tiny fixtures could not reach these; each came from a real model or
a real quant format.
7. **Routed experts were written in a layout no real Qwen3.5 checkpoint
uses.** Real Qwen3.5 stores experts packed as `[num_experts, out, in]`;
the mapping emitted per-expert names, so every routed expert was
dropped. The fixture actively hid this: transformers *unpacks* experts
on `save_pretrained`, so the saved reference agreed with the wrong
output. Fixed with a `transpose` kwarg on `_pack_name_remapping` plus a
`GroupedMLPPacking` rule, so fused `TEGroupedMLP` reaches the same
packed tensors — which is also what lets Qwen3.5 keep grouped GEMM
(**22.1 GB/GPU vs 38.9 GB/GPU** on a 20-layer, 256-expert model).
8. **`_grouped_mlp_packing` was broken for NVFP4.** It max-merged
`weight_scale`, but NVFP4 needs each expert's per-block scales *stacked*
with only the global `weight_scale_2` merged; it also dequantized packed
`uint8` against per-block scales, and passed `block_size=None`.
`weight_scale_2` is never populated in an FP8 run, so the whole branch
was dead code under FP8-only testing. `_grouped_mlp_slicing` gained
`quantize=False` so packing can quantize once over the stack, matching
`_pack_name_remapping`.
9. **`_mtp_prefix` corrupted every VLM's MTP tensor names.** It did
`prefix.replace("model", "mtp")` uncounted, so
`model.language_model.layers.{}` became `mtp.language_mtp.layers.0.*` —
tensors present and correctly valued, under names nothing loads.
LLM-only prefixes contain one occurrence, so this was invisible until a
VLM with MTP was exported.
10. **`load_multimodal_components` rejected HF repo ids.** `quantize.py
--hf_model_name_or_path Qwen/Qwen3.5-0.8B` worked, but the documented
export step failed with *"It should be a directory"*. Its sibling in the
same file already resolved repo ids via `snapshot_download`; now it does
too. This affected **every** VLM export.
`Qwen3_5ForConditionalGeneration` (dense Qwen3.5-VL) is now registered
for export and vision passthrough, which bugs 9 and 10 were blocking.
#### New: Qwen3.5-VL
`GatedDeltaNetSlicing` splits Megatron's fused `in_proj` (`[query, key,
value, z, beta, alpha]`) into HF's `in_proj_qkv` / `_z` / `_b` / `_a`,
taking sizes from the module's own `in_proj_split_sections` so TP
sharding falls out. Widening coverage to Qwen3.5's *gated
full-attention* layers then exposed a further split bug: gated attention
packs a per-head output gate beside each query head, so `_qkv_slicing`
split 192 rows as 96/48/48 instead of 128/32/32. It now derives the
group stride from `config.attention_output_gate`, matching
Megatron-Bridge's `split_qkv_weights`. The non-gated path is unchanged.
#### New: the export path verifies itself
- `assert_exported_checkpoint_matches` compares an exported checkpoint
against the model it came from — key set, shapes (accounting for NVFP4
`uint8` packing), safetensors index consistency, and values — replacing
existence-only assertions in all three export tests.
- `GPTModelExporter.save_pretrained` now raises if the export dropped
tensors the source checkpoint has, so *user* runs on architectures CI
never sees are protected too, not just tiny models.
- Loading a checkpoint whose quantizer tensors have no restorable state
now raises instead of silently loading unquantized.
- `assert_has_modelopt_state` replaces `rglob("modelopt_state")`, which
passes on an empty state; `assert_no_quantizers_matching` fails on
future HF↔Megatron name drift.
The mapping is also table-driven now: vision-tower prefixes live in
`all_mcore_hf_vision_passthrough_mapping` and
`with_language_model_prefix` is shared, so adding a VLM no longer means
editing `unified_export_megatron.py`. Five call sites that answered "is
this a VLM" three different ways now share `get_language_model` /
`is_vlm_config`.
### Usage
```bash
# Dense VLM (Qwen3-VL) -- no extra flags
torchrun --nproc_per_node 2 quantize.py \
--hf_model_name_or_path Qwen/Qwen3-VL-8B-Instruct \
--quant_cfg nvfp4 --tp_size 2 \
--export_megatron_path /tmp/Qwen3-VL-8B-NVFP4-megatron
torchrun --nproc_per_node 2 export_quantized_megatron_to_hf.py \
--hf_model_name_or_path Qwen/Qwen3-VL-8B-Instruct \
--megatron_path /tmp/Qwen3-VL-8B-NVFP4-megatron \
--pp_size 2 --export_unified_hf_path /tmp/Qwen3-VL-8B-NVFP4-hf
# Gated MoE (Qwen3.5-VL, Qwen3-MoE) -- no extra flags either. The scripts derive the
# expert layout from the model config, so quantize / distill / export all agree.
# --no_moe_grouped_gemm forces SequentialMLP if you want it explicitly.
```
### Testing
All in `nvcr.io/nvidia/nemo:26.08` on 2x RTX 6000 Ada.
| Suite | Result | Time |
|---|---|---|
| `tests/examples/megatron_bridge/` (full) | 18 passed | 27m58 |
| `tests/gpu_megatron/torch/export/` | 38 passed | 2m13 |
| `tests/unit/torch/export/` | 186 passed | 1.5s |
| pre-commit (ruff, ruff format, mypy, bandit) | clean | — |
| `tests/examples/megatron_bridge/test_quantize_export.py` on **2 GPUs**
(`pp_size=2`) | 3 passed | 5m |
The export leg of `test_quantize_and_export` now scales with `num_gpus`
like its quantize leg
already did. Previously it was hardcoded to one process, so the
collective checkpoint load ran at
PP=1 on both the 1-GPU PR runner and the 2-GPU nightly — which is how a
guard that raised on only
some pipeline stages (and therefore hung the job) reached review. The
dense `qwen3` case was dropped
in exchange: `qwen3_moe` already covers the non-VLM script path,
`qwen3vl` covers a dense decoder,
and that case was the one exceeding the 300s cap in CI.
#### Model coverage
`tests/gpu_megatron` runs in-process and is cheap, so it owns
per-architecture **mapping**
correctness. The example tests spawn `torchrun` per step and are ~50x
slower per case, so they
cover **script wiring** only — CLI flags, recipe resolution, and
checkpoint hand-off between steps.
| Suite | Models |
|---|---|
| `test_unified_export_megatron` | llama, nemotron, nemotron_h, qwen3vl,
qwen3_moe, qwen3_5_moe_vl x {none, FP8, NVFP4, +/-KV} x {grouped GEMM,
SequentialMLP} + eagle / medusa / MTP (29 params) |
| `test_megatron_importer` | nemotron_h, llama export->import round-trip
|
| `test_moe_layout_choice` | per-architecture grouped-GEMM exportability
(6 architectures) |
| `test_distill_megatron` | KD loss mechanics |
| Model | prune | quantize+export | QAD | distill+export |
|---|:--:|:--:|:--:|:--:|
| qwen3 | Y | Y | Y | Y |
| qwen3_moe | - | **Y (new)** | - | - |
| qwen3vl | - | **Y (moved from QAD)** | - | - |
| nemotron_h | Y | **Y (new)** | - | - |
| qwen3_5_vl | - | - | - | Y |
| qwen3_5_moe_vl | Y | **Y (new, both expert layouts)** | Y | - |
| deepseek_v3 | Y | - | - | - |
| gemma3vl | Y | - | ~~manual~~ removed | - |
QAD's unique property is that ModelOpt state survives distillation,
which needs one LLM and one
VLM rather than one case per architecture. Moving the rest to
quantize+export drops a `torchrun`
launch each: QAD went from 3 CI cases to 2 while quantize+export went
from 1 to 4, adding two
architectures for about a minute.
#### Real-model validation
Tiny fixtures cannot catch layout or scale bugs that only appear at real
dimensions, so the export
path was run end-to-end on released checkpoints. This is where bugs 7-10
came from.
| Model | Run | Result |
|---|---|---|
| Nemotron-3.5-Lightning-30B-A3B | NVFP4 4o6 PTQ → export → MMLU |
**0.7825 ± 0.0105** (gate 0.75) |
| Nemotron-3.5-Lightning-30B-A3B | Minitron pruning | 22.28B/3.00B
active, **0.5944** (gate 0.58) |
| Qwen3.5-0.8B (dense VLM) | FP8 PTQ → export → MMLU | BF16 0.4895 →
**0.4832** (±0.0127) |
| Qwen3.5-35B-A3B, half-depth (20 layers, 256 experts) | FP8 + NVFP4 PTQ
→ export | keys + shapes + **values** match reference |
| Qwen3.5-35B-A3B, full | FP8 PTQ | OOM on 2x48GB (see below) |
The half-depth model keeps real weights, real dims and all 256 experts.
Both expert layouts produce
identical key sets, and all exports pass
`assert_exported_checkpoint_matches(..., check_values=True)`
— every tensor, including all 20 x 256 experts, dequantizes to within
tolerance of the BF16
reference, so a transposed or mis-ordered expert stack would fail. NVFP4
lands in the correct packed
layout (`gate_up_proj [256, 1024, 1024]` U8, `weight_scale [256, 1024,
128]` E4M3,
`weight_scale_2 []` F32). Its *accuracy* is not meaningful — truncating
to 20 of 40 layers leaves a
chance-level model (BF16 0.2322, FP8 0.2538) — so it validates
correctness, not quality.
**Re-validated on the final code.** The numbers above were first taken
mid-review; since then the
NVFP4 block-scale merge changed on both packed paths, the vision-tower
download became two-stage,
and an expert-layout load guard was added. Both gating runs were
therefore repeated end to end:
Nemotron went 0.7748 → **0.7825 ± 0.0105** and Qwen3.5-0.8B went 0.4678
→ **0.4832 ± 0.0127**, with
the rest of the Nemotron pipeline reproducing exactly (3519 quantizers,
69GB checkpoint, 21GB
export). Both deltas are inside their own stderr, so the claim is that
the rework costs no accuracy
— not that it improved it. The Nemotron export also runs at `--pp_size
2`, exercising the new
collective layout guard on a real 30B MoE across pipeline stages.
Two limitations worth stating plainly:
- **No quantized accuracy number for a full-size MoE.** The full 35B
OOMs at 47.37 GiB while
*constructing* the model on 2x48GB, with grouped GEMM already enabled,
so no calibration knob
helps. Needs more GPUs than this setup has.
- **vLLM cannot yet serve packed FP8 Qwen3.5 experts.** `vllm
0.24.1.dev0` builds its fused expert
mapping weight-only, rewriting `experts.down_proj_input_scale` to
`w2_weight_input_scale` while the
parameter it registers is `w2_input_scale`. This is upstream and
independent of how the checkpoint
is produced — both of our export paths fail it identically. The 0.8B
numbers above are unaffected
(dense), and the packed exports are verified against the reference
checkpoint instead.
#### Guard verification
Each new guard was made to fire, not just to compile:
| Guard | Verification |
|---|---|
| Export self-check | Disabled the MoE guard, re-exported Qwen3-MoE -
independently reported all 24 dropped tensors. No false positives across
llama, nemotron, qwen3, qwen3-moe, qwen3vl, qwen3.5-vl, deepseek_v3
incl. eagle / medusa / MTP |
| Dropped-state raise | Deleted `modelopt_state` from a checkpoint with
50 quantizer tensors - raised instead of loading unquantized |
| NVFP4 value check | Flipped a `q_proj` - failed at `max_rel_err=1.74`
against a 0.3 threshold |
| Zero-centered gamma | Reproduced the off-by-1.0 on a good export -
caught as "not bit-exact" |
| Exclusion guard | Asserts no calibrated quantizer matches `conv1d` /
`mlp.router` / `output_layer` |
Exported artifacts are validated, not just their existence: 0 missing
keys vs reference, vision
tower bitwise-identical, dequantized weights within FP8 E4M3 error
(<=4.6%). The
`in_proj_a`/`in_proj_b` check is load-bearing - swapped alpha/beta would
still match on shape but
show ~100% error.
Also ran a tiny-Qwen3 **LLM** control through both steps to confirm the
exporter changes are a
no-op off the VLM path.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — the scripts now derive the
MoE expert layout from the model config, building SequentialMLP only for
architectures with no `experts.linear_fc1` rule, and the exporter raises
rather than dropping experts it has no rule for. Those runs previously
"succeeded" while writing a checkpoint containing no expert weights, so
no working behaviour is removed. `--no_moe_grouped_gemm` forces
SequentialMLP explicitly.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅ — approved (round 10: 0
CRITICAL, 0 IMPORTANT, 0 new suggestions); CodeRabbit approved earlier
### Additional Information
**MoE expert layout is now chosen automatically.** Only Nemotron-H can
export fused grouped-GEMM experts, so every other MoE architecture would
otherwise need `--no_moe_grouped_gemm` on all four scripts or hit a wall
at export. The scripts derive the layout from the model config — grouped
GEMM unless it would not be exportable — so they agree without threading
a flag. This changes MoE activation scales from one shared scale to
per-expert for the affected architectures.
Known gaps, unchanged by this PR:
- **Gated MoE still cannot use fused grouped GEMM.**
`_grouped_mlp_slicing` emits one weight per expert with no gate/up split
— its only prior caller, Nemotron-H, is non-gated, so every other MoE
architecture is built as `SequentialMLP` (see below). Adding that split
would restore the faster layout, but it needs a deliberate call on
activation-scale semantics: grouped GEMM keeps **one shared** activation
scale across experts while `SequentialMLP` has **per-expert** scales, so
the two are not numerically equivalent. It also needs EP>1 coverage.
- **Qwen3.5's alpha/beta gates share Megatron's fused `in_proj`
quantizer,** so they can only be kept in BF16 at export, not excluded by
name. Full fidelity needs per-section quantizers on the fused
projection.
- **Anchoring ModelOpt state on `.language_model`** (which would let
`quantize.py` quantize the language model directly and drop its
name-based non-LM disabling) needs a coordinated Megatron-Bridge change:
`save_sharded_modelopt_state` is ModelOpt code, but the restore the
Bridge path uses is Bridge's own and unconditionally restores onto the
root.
- **Gemma3-VL** remains Megatron-checkpoint only (`OMNIML-5366`).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added Muse Glimmer AutoQuantize and Alpamayo QAD workflows.
* Added streaming Kimi-K3 conversion and NVFP4 activation headroom
calibration.
* Added SFT-masked distillation for Megatron-Bridge.
* Added unified Hugging Face export for quantized Qwen3-VL and
Qwen3.5-VL checkpoints.
* MoE expert layouts are selected automatically, with an option to force
sequential experts.
* **Bug Fixes**
* Improved export validation for tensor coverage, MoE mappings,
quantizer state, and NVFP4 scales.
* Fixed Qwen3.5-VL GatedDeltaNet export handling.
* Preserved visual-model weights exactly during export.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
21b95adabb |
Add FP8 Vision Encoder quantization for Qwen3-VL and Qwen3.5 (#2083)
### What does this PR do? Type of change: New feature Adds opt-in FP8 Vision Encoder quantization recipes for Qwen3-VL and dense Qwen3.5: - `fp8_vision-kv_none`: FP8 Vision Encoder Linears, with the LLM and KV cache kept in high precision. - `fp8_vision_lm-kv_fp8_cast`: FP8 Vision Encoder and LLM Linears, with FP8 KV-cache cast. Patch embedding and vision-attention BMM operands remain in high precision. With `--calib_with_images`, calibration batches now pass through the complete VLM so multimodal inputs exercise the selected quantizers. This fixes image-text calibration for non-Nemotron VLMs and may change language-model activation ranges and output scales for existing commands. ### Usage ```bash # Vision Encoder only python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path <Qwen3-VL-checkpoint> \ --recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none \ --calib_with_images \ --calib_size 512 \ --skip_generate \ --export_path <output-checkpoint> # Vision Encoder + LLM + KV cache python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path <Qwen3-VL-checkpoint> \ --recipe huggingface/qwen3_vl/ptq/fp8_vision_lm-kv_fp8_cast \ --calib_with_images \ --calib_size 512 \ --skip_generate \ --export_path <output-checkpoint> ``` For dense Qwen3.5, replace `qwen3_vl` with `qwen3_5` in the recipe path. ### Testing - Validated recipe selection, image calibration, and GPU calibration/export for Qwen3-VL and Qwen3.5, in both vision-only and joint configurations. - Consolidated test run after rebasing: 364 passed, 77 skipped. - Ruff, recipe validation, and `git diff --check` passed. - Transformers 4.57 compatibility verified for Qwen3-VL; Qwen3.5 tests capability-skip when the required Transformers classes are unavailable. Deployment evidence with Qwen3-VL-2B on RTX PRO 6000 BSE, eight fixed frames and a BF16 LLM: | Configuration | Accuracy mean | Vision Encoder kernel time | Full-request GPU kernel time | |---|---:|---:|---:| | BF16 | 48.75 | 25.45 ms | 47.43 ms | | Standard FP8 | 48.42 | 19.35 ms (**24.0% faster**) | 41.39 ms (**12.7% faster**) | The accuracy mean covers MMMU, RealWorldQA, Video-MMMU, MVBench, and Video-MME. Serving reached 7.7% lower end-to-end latency and 7.9% higher throughput at concurrency 16. Qwen3-VL-2B accuracy was evaluated through vLLM on B300 with `--enforce-eager`. Both checkpoints used the same judge-free tasks, Qwen sampling preset, seed, and task parameters. | Benchmark | BF16 | VE-only FP8 | Delta | |---|---:|---:|---:| | MMMU validation | 45.33 | 45.22 | -0.11 pt | | RealWorldQA | 64.97 | 65.10 | +0.13 pt | | Video-MMMU | 31.56 | 31.11 | -0.45 pt | | MVBench | 51.40 | 50.10 | -1.30 pt | | Video-MME | 50.48 | 50.59 | +0.11 pt | | **Unweighted mean** | **48.75** | **48.42** | **-0.33 pt** | Runtime support for quantized Vision Encoder Linears is separate from this ModelOpt checkpoint-generation change. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ❌ — `--calib_with_images` now performs the intended complete VLM forward and may change language-model calibration scales. Recipe-based VLM PTQ also scopes recipe rules to the complete model. Both changes are documented in the changelog. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ — will run `/claude review` after opening the PR. ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added FP8 vision quantization recipes for Qwen3-VL and Qwen3.5. * Added Muse Glimmer AutoQuantize, Alpamayo QAD, streaming Kimi-K3 conversion, layerwise checkpoint export, and NVFP4 calibration/export workflows. * Added ONNX FP16 conversion support for excluding selected nodes. * **Bug Fixes** * Improved multimodal calibration, ONNX scale handling, and NVFP4 CPU compatibility checks. * **Documentation** * Expanded guidance for vision quantization, calibration, precision, and conversion workflows. * **Breaking Changes** * Removed deprecated PTQ and evaluation interfaces and raised the minimum supported Megatron container version. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: mpariente <mpariente@nvidia.com> Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> Co-authored-by: Shengliang Xu <shengliangx@nvidia.com> Co-authored-by: Shengliang Xu <106840466+shengliangxu@users.noreply.github.com> |
||
|
|
de3eda8a11 |
Restructure recipes: split per-model_type recipes from model-hub checkpoint recipes (#2219)
### What does this PR do? **Type of change:** Refactor (recipe-library layout) + documentation — backward-breaking for saved `--recipe` paths. Separate the two kinds of built-in Hugging Face recipes that were previously mixed under `modelopt_recipes/huggingface/`: - **`huggingface/<model_type>/`** — architecture recipes keyed by the transformers `model_type`; one recipe covers every checkpoint of that architecture. **Unchanged.** - **`models/<org>/<model_id>/`** — a *new top-level tier* for recipes that mirror one specific published checkpoint, keyed by its **model-hub path** (as on the Hugging Face Hub, ModelScope, etc.) so the on-disk path equals the hub path. Concretely, the model-instance recipes move out of `huggingface/` to the top level: - `huggingface/models/mistralai/…`, `huggingface/models/nvidia/…` → `models/mistralai/…`, `models/nvidia/…` - `huggingface/step3p5/Step3.5-Flash/…` → `models/stepfun-ai/Step-3.5-Flash/…` (re-keyed to the canonical HF repo id [`stepfun-ai/Step-3.5-Flash`](https://huggingface.co/stepfun-ai/Step-3.5-Flash) — org `step3p5`→`stepfun-ai`, id `Step3.5-Flash`→`Step-3.5-Flash`) **Why:** `modelopt_recipes/README.md` already documented a top-level `models/` tier, but the files lived under `huggingface/models/` and instance-specific recipes were awkwardly nested under the per-`model_type` tree. This aligns the filesystem with the documented layout and makes the instance tier hub-addressable — given a checkpoint id you can find (or place) its recipe with no lookup table. `load_recipe` resolves paths directly under `modelopt_recipes/`, so a top-level `models/` sibling of `general/` and `huggingface/` works identically. The move is metadata-only — all recipe YAML content is byte-identical (`R100` renames). Everything else is updating references (nvidia launcher YAMLs, `test_loader.py`) and docs: a new `models/README.md`, plus `huggingface/README.md`, root `README.md`, `ptq.md`, and the `10_recipes.rst` guide, which no longer describe instances under `huggingface/`. ### Usage Recipe paths for the moved checkpoint recipes lose the `huggingface/` prefix (and Step 3.5 Flash is keyed by its hub id): ```python from modelopt.recipe import load_recipe # before load_recipe("huggingface/models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16") load_recipe("huggingface/step3p5/Step3.5-Flash/ptq/nvfp4-mlp-only") # after load_recipe("models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16") load_recipe("models/stepfun-ai/Step-3.5-Flash/ptq/nvfp4-mlp-only") ``` The same rename applies to `--recipe …` CLI values and launcher `QUANT_CFG:` entries. Architecture recipes under `huggingface/<model_type>/` are unaffected. ### Testing - **Recipe resolution (torch-free):** parsed every recipe under `models/` and confirmed all `$import` targets resolve against the recipe root — 0 dangling across the tier. - **Docs consistency:** re-ran the `tests/unit/recipe/test_recipe_docs.py` logic; it now globs both `huggingface/` and `models/`, and every model dir (incl. `Step-3.5-Flash`, `Nemotron-3-Nano-4B-BF16`, …) plus every `general/ptq` recipe is still mentioned in `ptq.md`. - **Reference sweep:** repo-wide grep confirms no remaining references to the old paths outside the intentional historical CHANGELOG entries (released 0.44 / 0.45). - **pre-commit:** `markdownlint-cli2`, license-insert, and `bandit` hooks pass on the changed files. - Note: the full `pytest` suite was not run in my environment (no `torch`), so `test_recipe_docs.py` / `test_loader.py` should be exercised in CI. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ❌ — `--recipe` / `load_recipe` paths for the checkpoint-mirror tier change (drop the `huggingface/` prefix; `step3p5/Step3.5-Flash` → `stepfun-ai/Step-3.5-Flash`). Documented as a Backward Breaking Change in `CHANGELOG.rst` (0.47); the only *released* old paths affected shipped in 0.45. A clean break was chosen over a symlink or loader-alias shim. - If you copied code from any other sources or added a new PIP dependency …: N/A - Did you write any new necessary tests?: ✅ — updated `test_recipe_docs.py` to also glob the top-level `models/` tier so instance recipes stay covered by the doc-consistency check. - Did you update Changelog?: ✅ — added a 0.47 **Backward Breaking Changes** entry. - Did you get Claude approval on this PR?: ❌ <!-- run /claude review --> ### Additional Information Design note: an earlier iteration nested everything under `huggingface/model_type/` + `huggingface/models/`; the final layout keeps `huggingface/` flat (per-`model_type`) and lifts instances to a top-level `models/` tier, matching what `modelopt_recipes/README.md` already documented. The `Step3p5*` architecture class names (from the model's `trust_remote_code` modeling code) are unrelated to the recipe path and are left unchanged. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added checkpoint-specific PTQ recipes for Kimi-K3, Mistral Medium 3.5, and NVIDIA Nemotron models. * Added a Nemotron speculative-decoding warm-start recipe. * **Documentation** * Clarified recipe selection and directory organization. * Documented checkpoint naming conventions and updated usage examples. * **Bug Fixes** * Updated launcher configurations and examples to reference the new recipe locations and corrected model names. * **Tests** * Improved automatic recipe discovery and validation of documented recipe paths. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
8810eb5e31 |
Add ModelOpt recipe for DeepSeek-V4-Pro-0813 NVFP4 and --recipe to its PTQ script (#2287)
### What does this PR do?
Type of change: new feature
The quantization config for `nvidia/DeepSeek-V4-Pro-0813-NVFP4` existed
only as Python inside `_build_nvfp4_experts_cfg()`, so the released
checkpoint had **no entry in `modelopt_recipes/`** and could not be
looked up by name the way every other published model can.
`modelopt_recipes/README.md` states the goal directly — a recipe is
*"the single, version-controlled source of truth for how a model is
optimized … expressed as data instead of code"* — and this model was the
exception.
This adds:
-
`modelopt_recipes/huggingface/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only.yaml`,
composed from the existing `configs/ptq/units/base_disable_all` and
`configs/numerics/nvfp4` units.
- An optional `--recipe` flag on `examples/deepseek/deepseek_v4/ptq.py`.
It follows **`examples/kimi/kimi_k3`**, the closest precedent: a very
large MoE whose source already ships MXFP4 routed experts, converted via
`--cast_mxfp4_to_nvfp4` rather than through `examples/hf_ptq`, and
already wired to `--recipe` with a published YAML.
### Usage
```sh
torchrun --nproc-per-node 8 deepseek_v4/ptq.py \
--model_path <mp8_checkpoint> \
--config <DeepSeek-V4-Pro-0813>/inference/config.json \
--calib_size 512 \
--calib_seq 4096 \
--output_path <amax_dump> \
--recipe huggingface/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only
```
Omitting `--recipe` keeps the previous behaviour exactly.
### Testing
- `load_recipe` resolves the YAML and yields `num_bits (2, 1)` with
`block_sizes {-1: 16, type: dynamic, scale_bits: (4, 3)}` — identical to
the hardcoded config.
- **Equivalence checked behaviourally**, not by eyeballing dicts: both
configs were resolved against representative quantizer names using
last-match-wins, and agree on all of them.
| quantizer | hardcoded | recipe |
| --- | --- | --- |
| `...ffn.experts.17.w1_weight_quantizer` | enabled, NVFP4 | enabled,
NVFP4 |
| `...ffn.experts.17.w2_input_quantizer` | enabled, NVFP4 | enabled,
NVFP4 |
| `...ffn.shared_experts.w1_weight_quantizer` | disabled | disabled |
| `...attn.wq_weight_quantizer` | disabled | disabled |
| `mtp.0.ffn.experts.2.w1_weight_quantizer` | disabled | disabled |
| `lm_head_weight_quantizer` | disabled | disabled |
- `mtq.quantize` documents `algorithm` as a string **or** a dict keyed
on `method`, so the recipe's `{'method': 'max'}` needs no translation.
- `pre-commit` clean, including `validate modelopt recipes`.
No GPU run: this changes config plumbing only, and the default path is
byte-identical to before.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `--recipe` is optional and
defaults to `None`; without it `_build_nvfp4_experts_cfg()` is used
exactly as before.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies; `modelopt.recipe` is already a first-party import.
- Did you write any new necessary tests?: N/A — no new logic;
equivalence to the existing config is the property that matters and is
documented above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — recipe library addition; recent recipe/example PRs add no entry.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
The recipe covers the **quant config only**. `--calib_seq` — the setting
that mattered most for this checkpoint, since the 512 default does not
cover long-context activation ranges — is a dataloader argument rather
than part of the `mtq` config, so it stays on the CLI. Worth knowing if
the recipe is ever treated as a complete reproduction of the released
checkpoint: it is not, on its own.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added post-training quantization support for DeepSeek-V4-Pro-0813
routed experts using NVFP4.
* Added an optional recipe path for selecting equivalent quantization
settings.
* Preserved source formats for shared experts, attention, embeddings,
output layers, and MTP components.
* **Bug Fixes**
* Improved validation for missing or malformed quantization
configurations.
* Added safeguards against unsupported formats, scopes, algorithms, and
enabled MTP quantizers.
* **Documentation**
* Documented checkpoint conversion behavior, calibration requirements,
and supported quantization workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
029c67f27e |
feat(export): export each decoder layer as layerwise calibration finishes it (#2136)
### What does this PR do?
Type of change: new feature
Layerwise calibration can already resume, but only through a
full-precision scratch checkpoint, and a completed run still pays for a
second whole-model export pass over it.
`layerwise.export_dir` writes each decoder layer to its own quantized
shard as soon as calibration finishes with it, so the directory is a
complete, loadable checkpoint when the last layer lands and
`export_hf_checkpoint()` is skipped. The shards *are* the resume
artifact: a restarted run reuses layers already on disk instead of
recalibrating and re-exporting them, so no full-precision copy of the
model accumulates. The resume directory beside it holds only the current
boundary's cached activations and the per-layer output shapes.
Setting the config field is the whole switch — no CLI flag. `hf_ptq.py`
rewrites its value to `--export_path`, and derives the resume directory
(`<export_path>.layerwise_resume`) when you haven't chosen one.
One shard per layer is what makes resume safe: shards are written whole
and named from the layer index, so a crash can lose the layer in flight
but never corrupt an earlier one, and a re-run overwrites in place.
Because a resumed run never recalibrates the layers it skipped, **the
in-memory model is not valid for inference afterwards**; the field
implies `--skip_generate`.
Includes a pre-existing `main` fix this depends on: `_is_layerwise` used
`getattr` on an algorithm that YAML parses as a **dict**, so it answered
`False` for every layerwise recipe in the repo and the batch-size probe
it gates was never skipped. **Behaviour change:** `--batch_size 0` now
yields `batch_size=1` for layerwise recipes, as its comment intends.
Detection also now scans every algorithm entry rather than the first, so
a list-form recipe whose `layerwise` block is not first is recognised as
layerwise — same batch-size consequence. Nothing else on the non-fused
paths changes: `FUSION_FREE_FORMATS` is the exact set the inline list
held, `save_non_weight_artifacts` is a lift of the streaming exporter's
own block, and the calibration-loop changes are gated on an exporter
being present.
**Refused before calibration starts**, since each would otherwise
produce a silently different checkpoint rather than fail:
| Refused | Why |
|---|---|
| AWQ / SVDQuant | need pre-quant-scale steps that are still whole-model
|
| Weight-tied quantized modules | `sync_tied_input_amax` merges amaxes
across a partner that may be uncalibrated or already written |
| Multi-process (FSDP2) | every rank would write the same shards |
| Multimodal (VLM) | calibration runs on the extracted language model |
| MTP models | exclusions applied after calibration has written
everything |
| AutoQuantize recipes | only the mono-quantize path retargets
`export_dir` |
| Spec-dec, `--vllm_fakequant_export`, non-dense sparsity,
`int8_smoothquant`, encoder-decoder `model_type` | each routes to a
second exporter that would overwrite `--export_path` |
| `export_dir` on more than one algorithm entry, or on any but the last
| export finalizes shards as calibration walks the layers, so a later
pass would change the model after its checkpoint was written |
Shards are also bound to the run that produced them
(`.layerwise_export.json`: model class, layer count, formats, KV-cache
format, and a digest of the resolved quant config), so one run's
manifest cannot finalize another's shards. Source weights are not
digested — that would mean reading the whole model — so
differently-trained weights at the same path compare equal.
### Why a separate exporter
Three reuse paths were considered before adding one:
- **Extend `_StreamingShardWriter`.** It buffers by `max_shard_size`
into `__shard_part_*` temp names and renames to canonical names only in
`finalize()`, once the shard count is known. The resume invariant needs
the opposite: a stable `model-layer-00007.safetensors` committed when
layer 7 finishes, so "shard exists" means "layer done" across a restart.
Forcing a per-layer flush still leaves temp names, finalize-time
renaming, and an in-memory `_key_to_part` — every method would change.
- **Keep the layerwise checkpoint and run the streaming exporter at the
end.** This works, and it is why the pitch above is *not* durability:
that already exists. What it leaves is a second whole-model pass owed
*after* calibration finishes — itself needing a GPU session — where
per-layer export makes the last calibrated layer also the last exported
one. Scratch size only separates them for weight-mutating calibrators:
`save_layer_state` is off under per-layer export, but with
`calib_mutates_weights: false` (the shipped recipe) the checkpoint holds
just amax buffers either way.
- **Factor a shared per-module writer around `ExportContext`.** The
right long-term shape, but it touches all three existing export paths;
doing it here makes this change larger, not smaller.
One deliberate divergence from `_StreamingShardWriter` worth knowing
about: it clones tensors that share storage, this path lets `save_file`
raise instead. No model was found where the clone fires, and copying
unattributed aliases can hide a real bug rather than surface it. If a
checkpoint ever trips it, that is information we want.
Happy to take a different call on this — flagging it for maintainer
sign-off rather than assuming it.
### Usage
```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --export_path <out> \
--recipe modelopt_recipes/general/ptq/nvfp4_experts_only-kv_fp8_layerwise_export.yaml
```
Interrupt and rerun the same command: calibration resumes from the last
committed layer, finished shards are reused, and a run that had already
finished every layer only re-runs `finalize()`.
```yaml
quantize:
algorithm:
method: max
layerwise:
enable: true
calib_mutates_weights: false
export_dir: /tmp/modelopt_layerwise_export # presence is the switch; value replaced with --export_path
# checkpoint_dir omitted -> derived as <export_path>.layerwise_resume
```
### Testing
Each row exports the same calibration two ways — per-layer, and
whole-model via `export_hf_checkpoint()` — and compares them **tensor
for tensor and config for config**.
The 35B row was re-run on the current head, against a baseline built
from a `main` worktree rather than from this branch, so it covers both
"per-layer differs from whole-model" and "this branch broke the shared
whole-model path". The other rows date from earlier heads; the code they
exercise is unchanged, but they are not fresh runs.
| Model | Config | Result |
|---|---|---|
| Qwen3.6-35B-A3B (40 layers, 256 fused experts) | NVFP4 W4A4
experts-only + FP8 KV | 123,513 tensors, 0 mismatched; `config.json`,
`hf_quant_config.json`, `generation_config.json` all identical |
| Qwen3-30B-A3B (48 layers, 128 per-expert linears) | NVFP4 experts
(`nvfp4_static` weights) + `mse`, offload | 74,163 tensors, 0 mismatched
|
| Qwen3-30B-A3B | same, `SIGKILL` after 25/48 layers, then resumed |
74,163 tensors, 0 mismatched **vs the uninterrupted run** |
| Llama-3.1-8B-Instruct | FP8 dense + FP8 KV, resident | 803 tensors, 0
mismatched |
**Refusals verified on real checkpoints**, each writing **zero shards**
and never reaching calibration — the "refused before calibration starts"
claim above, demonstrated rather than asserted: multimodal and MTP
(Qwen3.6-35B, the MTP case on a text-only view since the multimodal gate
fires first), tied embeddings (Qwen3-0.6B), and multi-process (2-rank
`torchrun --use_fsdp2`, Llama-3.1-8B).
**Served, not just compared.** Under vLLM 0.27.1 (Marlin NVFP4 kernels,
SM 8.9): the 30B checkpoint exported three ways — whole-model,
per-layer, per-layer-resumed-after-a-kill — and the 8B exported both
ways all load and produce **identical greedy generations, 4/4 prompts**
within each model.
**Index integrity** on every checkpoint above: each `weight_map` key
resolves to the shard actually holding it; 0 missing, 0 extra, 0
mis-routed. Tensor equality alone never exercises that, and it is the
one artifact per-layer export builds differently.
Resume state stays bounded: **332 KB beside 22 GB** of shards on the
35B, **396 KB beside 19 GB** on the 30B — the committed boundary's
activations only, not one set per layer.
Not covered: the `trust_remote_code` `*.py` copy path.
Nemotron-Nano-12B-v2-Base fails with a CUDA illegal memory access on
these cards, on the whole-model baseline too, so it is an environment
limit rather than a result.
**Comparing the configs is new, and it caught a real bug.**
`get_quant_config` reports on the quantizer modules, which
`export_layer` replaces as it goes, so reading it in `finalize()`
described a model with no quantizers left: the checkpoint advertised
`quant_algo: null` while its weights were packed NVFP4, and under the
shipped experts-only recipe `hf_quant_config.json` was not written at
all. It is snapshotted in `__init__` now, beside the kv-cache format
already captured there — which is why that one field was correct while
the rest were not. Uniform FP8 and NVFP4 hid it because their configs
survive the conversion; only a mixed model loses its algo, and mixed is
what every shipped layerwise-export recipe is. Reverting the fix fails
`test_moe_export_matches` and passes the ten uniform-format cases,
matching what the 35B shows.
**24 GPU tests** in `tests/gpu/torch/export/test_layerwise_export.py`.
The equivalence oracle is a cross-product: {FP8, NVFP4, NVFP4 +
`get_qdq_activations_from_prev_layer`, mixed FP8/NVFP4, KV-cache} ×
{fresh, resumed-after-interruption}, each compared tensor-for-tensor
against `export_hf_checkpoint`. Plus MoE export; resume fail-fast;
resume artifacts replaced and pruned; complete-manifest finalize-only;
shards-without-manifest refusal; shards-from-a-different-run refusal
(format and module selection); identity-without-shards does not block a
rerun; export-does-not-mutate-the-model; index routes every key to the
shard holding it; AWQ refusal (from config, and after calibration);
export-without-`checkpoint_dir`.
**Unit tests** in `tests/examples/hf_ptq/test_example_utils.py` cover
the list-valued `algorithm` shapes: which entry owns export, whose
`checkpoint_dir` is derived, per-entry resume bases, both ambiguity
refusals, and the recipe shapes `recipe_layerwise_blocks` normalizes
(dict, list order, config object, and the empty cases).
`tests/gpu/torch/export/` 150 passed / 2 skipped (pre-existing env
skips) · `tests/unit/recipe` 284 · `tests/unit/torch/export` 186 ·
`test_layerwise_calibrate` 33 · `test_example_utils` 42 · pre-commit
clean.
Also verified: the exported directory reloads through
`AutoModelForCausalLM` and runs a forward.
Not a speed win: per-layer export was slower than the streaming export
in one offload pairing (271s vs 208s, the per-layer fusion probe),
though those runs shared GPUs so the magnitude is not cleanly measured.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `export_dir` defaults to
`None`; existing paths unchanged when unset, except the batch-size
change noted above.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — not yet; draft.
### Additional Information
**Pre-existing bug found on the way, not fixed here.** Layerwise
calibration leaves `self_attn.o_proj`'s input amax at `0.0` on every
layer but the last, so a full-NVFP4 layerwise model cannot be exported
by *any* path. `get_qdq_activations_from_prev_layer=True` avoids it,
pinning the cause to the pre-`calib_func` capture pass — which also
explains why only the last layer, the one that skips it, is correct.
That combination now works with per-layer export (it asserted on layer 0
until review caught it). Hidden until now because the shipped NVFP4
layerwise recipes are experts-only; the NVFP4 tests here exclude
`o_proj` for the same reason. Deserves its own issue.
---------
Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
424f47b871 |
Add QAD example for Alpamayo (#2271)
### What does this PR do?
Add `examples/alpamayo/qad.py` which distills the quantized Alpamayo
checkpoint's VLM against the FP16 VLM of the original with ModelOpt's
QADTrainer, and shards student and teacher with FSDP2 for multi-GPU
runs. Only the VLM is trained; the action expert stays frozen.
Type of change: new example
<!-- Details about the change. -->
### Usage
```
torchrun --standalone --nproc_per_node 8 qad.py \
--student_ckpt ./alpamayo-auto \
--output_dir ./alpamayo-auto-qad \
--parquet ./train_clips.parquet \
--max_steps 500 --fsdp2 --grad_ckpt --export
```
### Testing
Tested end-to-end on public Alpamayo-1 checkpoint
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Very short summary of changes only for new features,
backward breaking changes, deprecations, or fixes for critical bugs
present in previous releases. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added a quantization-aware distillation workflow for AlpamayoR1
models.
- Supports prompt-only or rollout-inclusive distillation, FSDP2
training, dataset slicing, revision pinning, and checkpoint resumption.
- Supports exporting trained models as complete, reloadable AlpamayoR1
checkpoints.
- Added optional vision-parameter freezing, trajectory-history fusion,
gradient checkpointing, evaluation, and synchronized training cadence.
- **Documentation**
- Expanded the Alpamayo guide with setup, training, dataset, and export
instructions.
- Clarified sensitivity-based quantization behavior.
- Added a version 0.47 quantization changelog entry.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Rohan Joshi <rohjoshi@nvidia.com>
|
||
|
|
6a2ae5a25b |
Fix the llm_eval timeout: reachable MMLU mirror + no pipe deadlock (#2270)
### What does this PR do? Type of change: Bug fix `tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8` has been failing with `Failed: Timeout (>900.0s) from pytest-timeout` on unrelated branches (runs 33027056246 and the one for `ad83a428`, while the 2026-08-24 nightly passed). It is not the test being slow — it is the harness deadlocking, and the deadlock also destroys the diagnostics that would explain the underlying kill. **Mechanism.** The traceback shows `self = <Popen: returncode: -9 args: ['scripts/huggingface_example.sh', ...]>` while still blocked in `stdout.read()`. The launcher was SIGKILLed (nothing in pytest sends SIGKILL — pytest-timeout raises in the main thread, and the test's `finally` `pkill` sends SIGTERM and only runs afterwards — so an OOM kill is the likely source). But `subprocess.run(..., stdout=PIPE, stderr=STDOUT)` waits for **EOF on the pipe**, not for the process, and a surviving grandchild (the TRT-LLM serve/build worker) still holds the write end. EOF never arrives, so the test blocks until the 900 s alarm. Because the pipe is never drained, **every line of child output is discarded**, which is why the CI log says nothing about what the script was doing when it died. Reduced to a self-contained reproducer: ```python script = "sleep 300 & echo 'launcher output'; sleep 0.3; kill -9 $$" subprocess.run(["bash", "-c", script], stdout=PIPE, stderr=STDOUT, text=True, timeout=20) # -> TimeoutExpired: still blocked in communicate() after 20.0s, output lost ``` **Fix.** `_run_capturing` now starts the command in its own session, drains its output on a reader thread (so logs stream as they arrive instead of being buffered until the end), waits on the *process*, and kills the process group if descendants still hold the pipe after a 30 s grace period. A killed launcher now fails in seconds with its logs intact instead of silently burning the test's whole timeout. This does not fix whatever kills the script; it makes it diagnosable. Worth noting separately: `test_qwen3_eval_fp8` took **749.10 s against its 900 s mark** on the last green nightly, so it is fragile regardless and may want its work trimmed or its budget raised once the logs show where the time goes. ### Usage ```python # unchanged public API run_example_command(cmd_parts, example_path="llm_eval") ``` ### Testing Verified against the reproducer above and on the normal paths: | scenario | before | after | | --- | --- | --- | | launcher SIGKILLed, survivor holds the pipe | blocks indefinitely (900 s in CI) | `rc=-9` in 3.5 s, `'launcher output'` captured | | the surviving descendant | keeps running | killed with the process group (stopped ticking, 20 -> 20 bytes) | | normal exit | ok | `rc=0`, stdout and stderr interleaved in order | | non-zero exit | ok | `rc=3`, output captured | The example-test suites that use this helper run through the same code path; `tests/examples/megatron_bridge` (16 passed, 1 skipped) exercised it on nemo:26.08 in the branch this was extracted from. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `_run_capturing` keeps its `(returncode, output)` contract; only the buffering strategy changed. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ❌ — this is test infrastructure; the scenario needs a process that outlives a SIGKILLed parent, which is awkward to assert in CI. Verified manually with the reproducer above. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — test-infrastructure fix. - Did you get Claude approval on this PR?: ❌ 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved command execution reliability with real-time output capture. * Ensured lingering child processes are cleaned up after commands exit or are interrupted. * Added warnings when forced cleanup may truncate output. * Prevented hangs when descendant processes keep output streams open. * **Documentation** * Updated MMLU setup instructions to use the Hugging Face dataset repository. * Improved Windows instructions by explicitly using `curl.exe`. * **Examples** * Improved MMLU downloads with retries, separate timeouts, resume support, and automatic temporary-file cleanup. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --- ### Update: the second half of the failure With the streaming fix in place, the next CI run showed the *actual* cause, which the old code had been hiding. That run blocked in `process.wait()` with `<Popen: returncode: None ...>` — child alive and working, not the previous dead-child pipe deadlock — and the now-visible output was: ``` --2026-08-27 18:09:20-- (try: 4) https://people.eecs.berkeley.edu/~hendrycks/data.tar Connecting to people.eecs.berkeley.edu ...|128.32.139.28|:443... failed: Connection timed out. Retrying. --2026-08-27 18:11:40-- (try: 5) ... ``` `huggingface_example.sh` downloads the MMLU tarball from `people.eecs.berkeley.edu`, that host stopped answering around 2026-08-25, and wget's default retry policy (20 tries, ~2 min per connect timeout) consumed the whole 900 s budget. Not runner-specific: the URL also times out from a developer workstation, and the nightlies flipped 08-24 ✅ / 08-25 ✅ / **08-26 ❌ / 08-27 ❌**, matching the outage. So this PR now carries both halves of the same failure: 1. the harness no longer deadlocks and no longer swallows the logs (`985809cc2d`), and 2. the MMLU data comes from HuggingFace's copy of the same tarball, with bounded retries (`40f1d89154`). The mirror is byte-for-byte the same dataset in the same layout the script already expects — verified by running the exact download/extract commands: ``` https://huggingface.co/datasets/cais/mmlu/resolve/main/data.tar -> HTTP 200, 166 MB data/mmlu/{dev,test,val}/ -> 57 subject CSVs each, plus auxiliary_train/ ``` `wget --timeout=20 --tries=3` plus an explicit error means the next dataset-host outage fails in about a minute with "Could not download the MMLU test data. Set MMLU_DATA_PATH to a local copy." instead of silently eating a test's timeout. The same URL is updated in `examples/llm_eval/README.md` so a manual run does not hit the dead host either. --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5500999d0b |
Add Kimi-K3 NVFP4 experts and FP8-PB attention recipe (#2206)
### What does this PR do?
Type of change: new example
Adds the calibration-free conversion pipeline and checkpoint-mirror PTQ
recipe used for `nvidia/Kimi-K3-NVFP4`:
- streams the 96-shard Kimi-K3 checkpoint without loading the 2.8T
model;
- casts the source MXFP4 routed experts to NVFP4 with expert
`input_scale=1.0`;
- quantizes the selected KDA and MLA attention weights to 128x128 block
FP8;
- leaves shared/latent experts, routers, convolutions, norms, the vision
tower, `lm_head`, and KV cache unquantized;
- emits mixed-precision Hugging Face metadata for deployment; and
- adds an exact recipe under
`modelopt_recipes/huggingface/models/moonshotai/Kimi-K3/` that directly
configures the streaming converter.
It also fixes `NVFP4QTensor.quantize()` probing CUDA/Blackwell
capability before checking whether the tensor is on CUDA and whether the
optional TensorRT-LLM fast path was requested. That probe broke the
converter's supported CPU path on hosts without a compatible GPU.
### Usage
```bash
python examples/kimi/kimi_k3/quantize_to_nvfp4.py \
--source_ckpt /models/moonshotai/Kimi-K3 \
--output_ckpt /models/Kimi-K3-NVFP4 \
--recipe huggingface/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention \
--jobs 8
```
The conversion requires no calibration dataset, forward pass, or GPU.
Multi-node shard conversion is also supported through `--rank`,
`--world_size`, and `--run_id`.
### Testing
```bash
uv run --frozen --extra dev python -m pytest -q \
tests/unit/torch/quantization/test_nvfp4_tensor.py \
tests/unit/recipe/test_kimi_k3_recipe.py \
tests/unit/recipe/test_recipe_docs.py \
tests/unit/torch/export/test_shard_cast_utils.py \
tests/examples/kimi/test_kimi_k3_quantize_to_nvfp4.py
```
Result: 34 passed.
All pre-commit hooks pass for the changed files, including recipe
validation, Ruff, mypy, Bandit, YAML formatting, and markdownlint.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: N/A
### Additional Information
The resulting checkpoint and model card are available at
https://huggingface.co/nvidia/Kimi-K3-NVFP4.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added calibration-free Kimi-K3 MXFP4-to-NVFP4 conversion with optional
FP8 attention quantization and distributed processing.
- Added NVFP4 activation calibration, weight-only quantization recipes,
grouped-expert quantization, compiled quantization options, SFT-masked
distillation, and MLflow tracking.
- **Documentation**
- Updated quantization terminology, recipe catalogs, checkpoint
guidance, and Kimi-K3 conversion instructions.
- **Bug Fixes**
- Improved CPU NVFP4 behavior, tied-weight export handling, EAGLE-3
training compatibility, and checkpoint export reliability.
- Removed deprecated configuration options and legacy evaluation
examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
|
||
|
|
7ff81dd795 |
Add day-0 verbosity gate + harden release skills from a live run (#2254)
### What does this PR do?
Type of change: Bug fix, new feature, documentation
Day-0 release treats verbosity as a hard gate, but nothing in the skill
measured it — `gate_compare.py` only does accuracy. A run could complete
Step 5 and report a publish recommendation with a gate silently
unmeasured. This adds the missing gate and folds in fixes for problems
that cost GPU-h on a live release run.
**New: `gate_verbosity.py` + Step 5b.** Pure `evaluate_verbosity()` plus
an artifact harvester, matching `gate_compare.py`'s conventions and
`failure_class` vocabulary.
Two details it encodes, both of which produced wrong verdicts before
they were understood:
- **Read `response_stats.avg_completion_tokens`.** The
`reasoning.*_tokens` fields are always `0`, which reads as "the harness
never captured tokens" and pushes you to word counts. Only the
reasoning/content *split* is missing, not the total. Words disagree with
the gate: one task read **+6.10% FAIL** in words and **+1.32% PASS** in
tokens.
- **Two run-hygiene filters.** Pooling a mismatched reasoning-effort run
reported **+43.06% FAIL** on a task that is **+1.24% PASS** matched; one
truncated run (n=200 vs 294) reported **+9.58% FAIL** on a task that is
**+1.71% PASS** without it. Tasks with no common sample count are
reported `not_comparable` rather than as a delta.
**Skill hardening**, each from a specific failure:
- **Step 2b canary** — poll ceiling must exceed load time (a 50 min poll
against a 51 min load failed a checkpoint that serves fine), print an
explicit `RESULT:` on every path (a fall-through exits 0 and reads as
PASS), log to shared storage, and canary the **as-exported** artifact
rather than a copy modified to make it work.
- **Step 4 config parity** — assert the candidate config differs from
the baseline's in nothing but checkpoint path and served-model name. A
mismatched `parallelism` was worth ~2 pp, enough to invert the sign of a
delta, and cost four re-runs.
- **Statistical power** — re-running does not guarantee fresh samples:
with a warm NEL response cache two runs came back bit-identical to 16
digits.
- **Step 6 closeout** — verify the published path against the evaluated
one by inode, and prefix rejected sibling exports.
- **Size gate** — growth is blocking by default and waived only when the
validation summary's
`source_precision` shows an already-sub-8-bit source (which cannot
shrink further under a
4-bit recipe) and the growth is within what that explains.
`source_precision` is now a
recorded field in the ptq validation table, so the waiver is reachable
from the normal
pipeline, and `SIZE_NOT_REDUCED` has a triage row pointing at declaring
it.
- **`ptq.py`** — `--calib_seq` matters more than `--calib_size`, and
`--mse_calibrate` is a no-op under `--cast_mxfp4_to_nvfp4` (it tunes
weight quantizers only, and the cast overwrites `weight_scale`).
- **`.gitignore workspaces/`** — the skills create scratch directories
inside the repo; nothing excluded them.
### Usage
```bash
python "$SKILL_DIR/scripts/gate_verbosity.py" \
--baseline <baseline_eval_root> --candidate <candidate_eval_root> \
--glob 'eval_*' --threshold 0.05
```
Exit codes match the sibling gates: `0` pass, `1` the gate ran and
failed, `2` the gate could not
read its input (wrong root, `--glob` matched nothing, everything
excluded). Prints per-task tokens,
delta, `within_threshold`, `sample_count`, run counts,
`dropped_mismatched_runs`, any
`truncated_comparison`, `not_comparable`, `harvest_diagnostics`, and a
`max_abs_delta` summary.
### Testing
- 8 new unit tests in `test_gates.py`, one per real failure mode
(two-sided threshold, partial-run filtering, unequal sample counts,
short-output warning, one-sided tasks, empty input). Full suite: **36
passed**, no GPU or network.
- `gate_verbosity.py` validated end-to-end against a real day-0 run's
artifacts: reproduces the hand-computed result (`max_abs_delta =
0.0171`, pass) and correctly marks the two unequal-sample tasks
`not_comparable`.
- `pre-commit run --files <changed>` clean, including ruff, mypy,
bandit, markdownlint, and the `.claude/skills` symlink sync.
- Verified `max_sample_length` is a real `get_dataset_dataloader`
parameter with default 512, matching `--calib_seq`'s default, so
existing `ptq()` callers are unaffected.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `--calib_seq` defaults to the
pre-existing 512; `gate_ptq.py` reclassifies size growth from
`QUANT_COVERAGE_FAILURE` to `SIZE_NOT_REDUCED`, which is a more precise
class for an already-4-bit source and is covered by a new test asserting
a real coverage failure still outranks it.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependencies, stdlib only.
- Did you write any new necessary tests?: ✅ — 8 new tests for the new
gate.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
The measured figures come from a completed day-0 NVFP4 release
qualification. Model-specific results were removed from the general
skills where the rule stands on its own; two references were kept
deliberately — a model card citation illustrating per-scenario sampling,
and a model-specific vLLM MoE kernel crash where the model name *is* the
evidence.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added configurable calibration sequence-length limits for DeepSeek-V4
quantization.
* Added automated release checks for output verbosity, evaluation
comparability, serving readiness, and configuration parity.
* **Bug Fixes**
* Improved quantization size-ratio reporting by distinguishing
explainable growth from blocking failures.
* Clarified handling of deployment memory-access errors and infeasible
evaluations.
* **Documentation**
* Expanded guidance for calibration, remote execution, workspace
management, evaluation setup, deployment troubleshooting, and
statistical reliability.
* Added task-specific guidance for SciCode and GDPVal feasibility
checks.
* **Chores**
* Excluded workspace session directories from version control.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
d73278808b |
Bump nemo container requirement to 26.08 for MBridge examples (#2257)
### What does this PR do? Type of change: Bug fix Bumps the Megatron-Bridge examples, tests and launcher configs to `nemo:26.08` and removes the version-gated fallbacks they carried, plus the fixes needed to make the suites green on that container. **26.08 bump and shim removal** - Examples, CI workflows, `noxfile.py` and the `mbridge_*` launcher configs move to `nemo:26.08`. - `examples/megatron_bridge/_distillation_provider.py` is deleted — 26.08's Megatron-Bridge ships `convert_to_distillation_provider(..., distill_submodule=...)` natively, so `distill.py` imports it directly. - `prune_minitron.py` drops the `AutoBridge.from_hf_config` / config-only-export probing; `--no_moe_grouped_gemm` is no longer needed in the MoE pruning tests, and the Qwen3.5-MoE `skipif` is gone (native MoE expert mappings are in 26.08). - `_DynamicMambaMixer` targets only the raw `conv1d_weight` / `conv1d_bias` parameters that replaced the `conv1d` module in Megatron-Core. **MambaModel / MambaModelProvider removal** Megatron-Core has shipped `HybridModel` since 26.06 and `MambaModel` is a deprecated subclass that shares its `forward`, so `DMRegistry` resolves those instances to the `HybridModel` registration and the separate entry is redundant. Same for `MambaModelProvider` vs `HybridModelProvider` on the bridge side. `MambaMixer` / `MambaLayer` / `ExtendedRMSNorm` are untouched — the layers still exist. The deprecated `get_te_mamba_stack_spec` is removed; use `get_te_hybrid_stack_spec`. **Bug fix: compressed output_layer extra state** `mtq.compress` converts even a *disabled* `output_layer` into a `RealQuantLinear` (its weight is left uncompressed, since `pack_real_quantize_weight` skips disabled quantizers). The guard added in #2112 excluded `RealQuantLinear` by class, so the layer emitted extra state and every worker died in `GPTModel.sharded_state_dict`: ``` RuntimeError: Boolean value of Tensor with more than one value is ambiguous megatron/core/models/gpt/gpt_model.py:896, in sharded_state_dict output_extra_state and output_extra_state.data ``` The guard now keys off whether the weight was actually compressed (`QTensorWrapper`) instead of the class. This took out all 12 `test_homogeneous_compressed_sharded_state_dict` params, and the crashed workers poisoned the pool, which surfaced as unrelated timeouts and NCCL errors in `test_layer_sync_moe_local_experts_amax`, `test_kv_cache_quant`, `test_kv_cache_amax_sync`, `test_convert_mcore_te_gpt_model` and `test_homogeneous_sharded_state_dict_te_spec` — 21 tests in total. The e2e coverage is `skip_flaky_on_blackwell`, so CI never ran it; `test_output_layer_extra_state_empty_when_nothing_quantized` now asserts the contract directly and is not skipped. **Checkpoint import entry point** 26.08 replaced `examples/conversion/convert_checkpoints.py` with `scripts/conversion/convert.sh`, so `tools/launcher/common/megatron_bridge/import/import.sh` and the three README snippets are retargeted. `import.sh` uses the distributed GPU backend with `GPUS_PER_NODE` / `TP` / `PP` / `EP` knobs. **Megatron-LM on nemo:26.06** keeps working: `_get_mamba_conv1d` still dispatches between the `conv1d` module (26.06 and earlier) and the raw parameters (26.08+), so `import_mcore_gpt_from_hf` / `export_mcore_gpt_to_hf` handle NemotronH on both. Only the Megatron-Bridge examples and Minitron pruning of Mamba/hybrid models require 26.08. **Test consolidation** `test_export_distilled_megatron_to_hf.py` is merged into `test_distill.py`: `test_distill_llm` becomes `test_distill_llm_hf_export` and covers the standalone `--export_iterations all` run on the checkpoints it already produces, saving one full distillation (~185 s of CI time). The two mamba-named gpu test files are renamed to `hybrid`. ### Usage ```bash # HF -> Megatron import, via Megatron-Bridge's 26.08 conversion entry point bash /opt/Megatron-Bridge/scripts/conversion/convert.sh import \ --executor local \ --device gpu \ --gpus-per-node 8 \ --hf-model Qwen/Qwen3-8B \ --megatron-path /tmp/Qwen3-8B-megatron ``` ### Testing All on `nvcr.io/nvidia/nemo:26.08`, 2x RTX 6000 Ada, no timeout overrides: - `tests/examples/megatron_bridge`: 16 passed, 1 skipped (28m14s). The skip is the `gemma3vl` QAD param, now `@pytest.mark.manual` since `qwen3_5_moe_vl` covers the VLM QAD path. - `tests/gpu_megatron` (`_extensions`, `distill`, `export`, `opt`, `peft`, `sparsity`, `speculative`, `utils`): 61 passed, 5 xpassed. - `tests/gpu_megatron/torch/export` re-run after the conv1d dispatch change: 27 passed. - The 21 previously failing/hanging quantization tests: 21 passed (12 + 9). - `tests/gpu_megatron/torch/{nas,prune}`: verified separately. `import.sh` equivalence on a toy `qwen3_moe`, comparing all 12 weight tensors after flattening each dist checkpoint with `dcp_to_torch_save` — the GPU backend at 1 GPU, `--tp 2`, `--pp 2`, `--ep 2`, and `import.sh` end-to-end (`GPUS_PER_NODE=2 EP=2`) are all byte-identical to `--device cpu`. `nemo:26.06` compatibility was checked directly in that image: `megatron.core.models.hybrid.HybridModel`, the modelopt hybrid spec and `hybrid_layer_pattern` are all present, while `megatron.bridge.models.hybrid` and the bridge's `distill_submodule` are not. The NemotronH round-trip test failed there before the conv1d dispatch was restored and the dispatch is back in place; per project convention the suites themselves only run on 26.08. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ⚠️ Megatron-Bridge examples plus Minitron pruning of Mamba/hybrid models now require `nemo:26.08`. Megatron-LM quantization and checkpoint export still run on `nemo:26.06`. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — `test_output_layer_extra_state_empty_when_nothing_quantized` for the compress fix; existing tests extended for the merged export coverage. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added guidance for importing Hugging Face checkpoints into Megatron distributed format. * Expanded distillation workflows to export selected or all checkpoint iterations. * **Improvements** * Expanded Hybrid model support across Megatron workflows. * Updated distributed import tooling with GPU and parallelism options. * Updated supported environments and examples to NVIDIA NeMo 26.08. * **Bug Fixes** * Corrected output-layer quantization state handling when quantization is disabled. * **Documentation** * Added compatibility guidance for current and legacy NeMo containers. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
449a39922b |
Pin nemo_automodel below 0.6 for the fastgen example (#2260)
### What does this PR do? Type of change: Bug fix `nemo_automodel` 0.6.0 removed `nemo_automodel.recipes.diffusion.train.is_main_process` without a replacement (it was a three-line rank-zero predicate in 0.5.0, and 0.6.0 defines no equivalent anywhere in the package). `examples/diffusers/fastgen/dmd2_recipe.py` imports it, so the example's import guard fires and **every** test in `tests/examples/diffusers/` errors at collection: ``` ImportError: cannot import name 'is_main_process' from 'nemo_automodel.recipes.diffusion.train' tests/examples/diffusers/fastgen/test_resume_dataloader.py E ImportError: The DMD2 fastgen example requires `nemo_automodel`. ... collected 42 items / 1 error ``` The requirement was `>=0.4.0,<1.0`, so CI picked 0.6.0 as soon as it was published and the `onnx (diffusers)` job started failing on every PR (e.g. runs 33020467654, 33019418460, 33010815298, 33007613265, 33006944292 — all unrelated branches). Capping at `<0.6` restores the tested range. Every other `nemo_automodel` symbol the example imports still exists in 0.6.0 (`_diffusers.auto_diffusion_pipeline.NeMoAutoDiffusionPipeline`, `recipes.diffusion.train.TrainDiffusionRecipe`, and the four `components.datasets.diffusion.*` helpers), so `is_main_process` is the only blocker; the alternative is defining that predicate locally and widening the cap again, which is worth doing separately if the example is meant to track 0.6. ### Usage ```bash pip install -r examples/diffusers/fastgen/requirements.txt ``` ### Testing Reproduced the break by diffing the published wheels: `is_main_process` is defined at `nemo_automodel/recipes/diffusion/train.py:692` in 0.5.0 and absent from 0.6.0 (`grep -rn "def is_main_process"` over the unpacked 0.6.0 wheel returns nothing). Confirmed the remaining imported symbols are all still present in 0.6.0. CI on this PR exercises the fix directly: the `onnx (diffusers)` job installs from this requirements file and is the job that has been failing. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — existing dependency, tightened bound. - Did you write any new necessary tests?: N/A — the existing `tests/examples/diffusers/` suite is what this unblocks. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — dependency-pin fix for a break introduced and fixed within the same unreleased cycle. - Did you get Claude approval on this PR?: ❌ 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed dependency compatibility for the FastGen diffusion example. * Prevented installation of versions that could cause the example to fail at startup. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5db2682519 |
[Example]: Calibration-free FP8/NVFP4 PTQ for speculative-decoding drafters (#2027)
### What does this PR do? Type of change: new example Adds `examples/speculative_decoding/scripts/quantize_drafter.py`, a CLI that quantizes an exported speculative-decoding drafter to FP8 or NVFP4 — weight-only or weight+activation — with no calibration data. It needs no modeling code either. Exported drafters such as [`nvidia/MiniMax-M3-DSpark`](https://huggingface.co/nvidia/MiniMax-M3-DSpark) have no importable model class, so each 2-D weight is wrapped in a throwaway `nn.Linear` under its checkpoint key and ModelOpt's usual `quantizer_name` patterns select over those names. Works for any drafter layout (DSpark / DFlash / EAGLE3 / Medusa). **Formats:** `w4a16_nvfp4`, `nvfp4`, `fp8`, `fp8_pc_pt` — the ModelOpt formats vLLM's backend can actually serve. AWQ is deliberately not offered, since `awq_lite` silently degrades to plain RTN without a `forward_loop`. **Static activation scales without calibration.** `fp8` and `nvfp4` normally need an activation amax *measured* on calibration data; a fixed `input_scale` of 1.0 is applied instead. That works because acceptance length is governed almost entirely by **clipping**, not resolution: Sweeping the fixed scale over three decades (same setup as the Testing section below; bf16 baseline 3.1423): | `input_scale` | amax | FP8 AL | vs bf16 | NVFP4 AL | vs bf16 | |---|---|---|---|---|---| | 0.003 | 1.3 | 2.2204 | -29.34% | 2.2076 | -29.75% | | 0.01 | 4.5 | 2.6719 | -14.97% | 2.6641 | -15.22% | | 0.03 | 13.4 | 2.9751 | -5.32% | 2.9259 | -6.89% | | 0.1 | 44.8 | 3.1013 | -1.31% | 3.0206 | -3.88% | | 0.2 | 89.6 | 3.1178 | -0.78% | 3.0015 | -4.48% | | 0.3 | 134.4 | 3.1370 | -0.17% | 3.0222 | -3.82% | | 0.5 | 224.0 | 3.1268 | -0.50% | 3.0360 | -3.38% | | **1.0 (default)** | **448.0** | **3.1457** | **+0.11%** | **3.0193** | **-3.91%** | | 2.0 | 896.0 | 3.1354 | -0.22% | 3.0172 | -3.98% | | 4.0 | 1792.0 | 3.1245 | -0.57% | 3.0034 | -4.42% | Both formats fall off a cliff below ~0.03, where the declared range sits far under the activations' true magnitude and most of the tensor is clipped. Both then sit on a flat plateau from ~0.3 to 4.0 **with no drop-off at the top**, so the scale only has to be big enough. 1.0 is the middle of that plateau, which is why it is hardcoded rather than exposed. NVFP4 trails FP8 by a roughly constant 3.5% across the plateau — that gap is the 4-bit resolution cost, and no choice of scale recovers it. Deriving the amax from the weights instead was tried and does not work: `max|W|` averages 0.79 while a RMSNorm'd activation is O(1) with outlier channels in the tens, so the range lands 1–2 orders of magnitude low and clips, measuring -31% to -46% AL. **Where calibration would go.** All of this sits behind `resolve_activation_scales()`, the single place deciding where a static amax comes from. Real calibration slots in ahead of the fixed fallback with no change to the CLI or the call site, and composes because `set_static_activation_amax()` skips quantizers that already have an amax: ```python if calib_forward_loop is not None: mtq.calibrate(root, quant_cfg["algorithm"], forward_loop=calib_forward_loop) set_static_activation_amax(root) # fills in what calibration did not reach ``` **Serving a quantized drafter.** Four things had to be written into the exported checkpoint before vLLM would load one: - emit `quant_method` (`modelopt_fp4` / `modelopt`) — vLLM reads that key, ModelOpt writes only `quant_algo` - emit the exclusion list under `ignore` too — that is the key read from the flat `quantization_config`; `exclude_modules` alone yields an empty exclusion set - add `*<name>` wildcards so exclusions match a runtime's nested module prefix (`model.fc`) rather than the checkpoint key (`fc`) - add `*qkv_proj` / `*gate_up_proj` aliases for layers a runtime fuses, whose names appear in no checkpoint key Nothing is then needed on the caller side. **This closes the open question left in the previous revision of this PR: vLLM does read `quantization_config` off the draft checkpoint.** `ModelConfig._verify_quantization` fills `quantization` in from `quant_method` when it is unset, so once the export declares that key — the first fix above — detection works on its own. Verified on Nemotron-3.5-Lightning passing nothing: `Detected ModelOpt NVFP4 checkpoint (quant_algo=NVFP4)` → `FlashInferCuteDslNvFp4LinearKernel`, AL 4.278 against 4.203 measured earlier. `specdec_bench` also gains a `DSPARK` algorithm, which it did not have: an exported `Qwen3DSparkModel` would otherwise have to go through `DFLASH` and be built with vLLM `method="dflash"`. The branch sets `method="dspark"` and leaves `draft_sample_method` on vLLM's own default of `greedy`. A target whose fused-collective workspace (sized at CUDA-graph capture) overflows at large speculative batches can disable graphs with `--runtime_params '{"engine_args": {"enforce_eager": true}}'`. For DFlash-family drafters, `qwen3_dflash.py` builds its fused context-KV projection by reading `qkv_proj.weight` raw and calling `F.linear`, which cannot consume a packed weight. Keep those layers in bf16 with `--exclude '*q_proj*' '*k_proj*' '*v_proj*' '*qkv_proj*'`; `o_proj` and the MLP — the bulk of the drafter — still quantize. That exclusion is mandatory, not a tuning choice. `fc` (the projection from the target's captured layers into the draft) is the one real knob, and it is a genuine trade rather than a free win — see the Testing section for both models' numbers. The examples quantize it; add `'*fc*'` to the exclude list to keep it in bf16. `embed_tokens`, `markov_head` and `confidence_head` are excluded by default: they are 2-D so the flat view treats them as GEMMs, but they are embeddings or a single-output projection. `lm_head` is excluded by the preset itself — unlike on a base model it is 37% of this drafter's parameters, so `--quantize_lm_head` is a real lever (~1.9 GiB), but measure AL first. The flag re-enables both of `lm_head`'s quantizers; re-enabling only the weight one would ship a W+A checkpoint whose `lm_head` has no `input_scale` while the config still advertises it as quantized. ### Usage ```bash # weight+activation FP8, calibration-free, lossless on both models measured below python scripts/quantize_drafter.py \ --drafter_path deepseek-ai/dspark_qwen3_8b_block7 \ --qformat fp8 \ --export_path ./dspark-qwen3-8b-fp8 \ --exclude '*q_proj*' '*k_proj*' '*v_proj*' '*qkv_proj*' # smallest: weight-only NVFP4 python scripts/quantize_drafter.py \ --drafter_path nvidia/MiniMax-M3-DSpark \ --qformat w4a16_nvfp4 \ --export_path ./MiniMax-M3-DSpark-W4A16 ``` Or end to end on Slurm — quantize, then measure AL — via the launcher examples added here, one per target: ```bash uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_dspark_ptq_nvfp4.yaml --yes uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/hf_dspark_ptq_nvfp4.yaml --yes ``` Serving one, if you are not going through `specdec_bench`: ```python speculative_config = { "method": "dspark", "model": "./dspark-qwen3-8b-fp8", # quantization is read from its config.json "num_speculative_tokens": 7, } ``` ### Testing Two targets with different architectures, so the conclusions are not one model's quirk: * **Qwen3-8B** (dense transformer) + [`deepseek-ai/dspark_qwen3_8b_block7`](https://huggingface.co/deepseek-ai/dspark_qwen3_8b_block7), `block_size` 7, TP1. * **Nemotron-3.5-Lightning-30B-A3B** (hybrid Mamba-MoE) + [`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark`](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark), `block_size` 8, TP8, with the mamba engine settings the model card pins (`mamba_backend=flashinfer`, `mamba_ssm_cache_dtype=float16`, stochastic SSM-cache rounding). Both: MT-Bench 80 questions, greedy, one vLLM instance per point. | recipe | activations | Qwen3-8B AL | vs bf16 | Nemotron-3.5 AL | vs bf16 | |---|---|---|---|---|---| | bf16 baseline | — | 3.1423 | — | 4.3296 | — | | **`fp8`** | static, `input_scale` 1.0 | **3.1457** | **+0.11%** | **4.3289** | **-0.02%** | | `fp8_pc_pt` | dynamic per-token | 3.1228 | -0.62% | 4.3411 | +0.26% | | `w4a16_nvfp4`, `fc` in bf16 | bf16 (weight-only) | 3.0392 | -3.28% | 4.2899 | -0.92% | | `w4a16_nvfp4`, `fc` quantized | bf16 (weight-only) | 3.0186 | -3.94% | 4.2334 | -2.22% | | **`nvfp4`** | static, `input_scale` 1.0 | **3.0193** | **-3.91%** | **4.2030** | **-2.92%** | **FP8 weight+activation at the fixed `input_scale` of 1.0 is lossless on both.** +0.11% and -0.02% are both inside run-to-run noise — the Nemotron baseline was measured twice under identical settings and the two runs differ by 0.94% (4.3093 / 4.3499), which sets the resolution of that column. On the same reading, `fp8` and `fp8_pc_pt` are indistinguishable on Nemotron; the dynamic variant only pulls ahead on Qwen3. NVFP4 costs 3-4% on Qwen3 and 2-3% on Nemotron, i.e. the 4-bit weight resolution is the real price and it is model-dependent but bounded. Whether to quantize `fc` is a per-model call rather than a general recommendation — it buys a few percent of size for an AL cost that differs by ~2x between these two drafters: | `fc` bf16 → quantized | Qwen3-8B | Nemotron-3.5 | |---|---|---| | checkpoint size | 3.293 → 3.181 GiB (-3.4%) | 1.316 → 1.258 GiB (-4.4%) | | AL | 3.0392 → 3.0186 (-0.68%) | 4.2899 → 4.2334 (-1.32%) | `fc` itself is only 3.5% (Qwen3) / 4.5% (Nemotron) of drafter parameters; `embed_tokens` is the bulk (26% / 36%) and is excluded by default. The Qwen3 `w4a16_nvfp4` rows were measured in a later session than the rest of that column; the `fc`-in-bf16 run reproduced the original number to four decimals (3.0392), so the column is internally comparable. Also validated on `nvidia/MiniMax-M3-DSpark`: `w4a16_nvfp4` runs in 67 s on CPU, 9.98 GiB (fp32) -> 3.51 GiB; all 43 quantized tensors round-trip within 0.0952 relative error; the 29 untouched tensors are bit-identical to `bf16(source)`. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (example-only) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ❌ — validated manually as above. Can add a `tests/examples/speculative_decoding/` test over a small synthetic drafter if wanted before merge. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (example-only) - Did you get Claude approval on this PR?: ❌ (not yet run) ### Additional Information The measurements above are one drafter on one target with one benchmark; the plateau's location and the ~3.5% NVFP4 gap should be re-measured before assuming they carry to a different drafter. Note when reading an exported checkpoint: `input_scale` is `amax/448` for FP8 but `amax/(6*448)` for NVFP4, so the one fixed amax records as 1.0 in an FP8 checkpoint and 0.1667 in an NVFP4 one. Both mean the same activation range. Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
2b296b2f62 |
Support fine-tuning released DFlash/DSpark drafters (causal SWA, attention sink, warm start) (#2149)
# Support fine-tuning released DFlash/DSpark drafters (causal SWA, attention sink, warm start) ### What does this PR do? Type of change: New feature + bug fix Adds what ModelOpt was missing to fine-tune an already-published DFlash/DSpark draft model. The concrete target is [`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark`](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark) on its hybrid Mamba/attention/MoE base, but every change is generic. Before this PR that checkpoint could not be trained faithfully — or even loaded: its attention-sink tensors were dropped as unexpected keys, its block-causal attention had no implementation, and its capture layers were silently overwritten with ModelOpt's defaults. **New user-facing options** (all default to today's behavior, so existing runs are unchanged): | Option | Values | Purpose | | --- | --- | --- | | `dflash_draft_attention` | `bidirectional` (default) / `causal` | Block-internal attention pattern. `causal` restricts a query at block position `i` to draft positions `<= i`. | | `dflash_attention_sink` | `false` (default) / `true` | Learnable per-head `attention_sink_bias [num_heads]` on every draft layer — one extra logit appended before the softmax and dropped after, so a head can put probability mass nowhere instead of being forced to attend inside its window (the GPT-OSS formulation). | | `dflash_init_checkpoint` | path | Warm-start the draft from an exported checkpoint instead of a random init. Any missing/unexpected/wrong-shaped tensor raises rather than warns. | | `dflash_architecture_config.target_layer_ids` | list | Which base layers feed the draft's `fc`. Previously recomputed unconditionally with no override. | **Bugs fixed along the way** (each one silently corrupts training rather than failing): - The exporter hard-coded `dflash_config.causal: False` and only wrote it under SWA, so even a correctly-trained causal draft would be served non-causally. It now reflects the trained setting, and emits `attention_sink_bias` when enabled. - `_build_generate_swa_mask` returned `None` whenever `swa_window_size` was unset, which would have dropped the causal structure at generation time while training used it. - `target_layer_ids` was recomputed from the uniform default on every convert. The released drafter uses `[1,5,19,29,41,51]`; the default for a 52-layer base is `[1,11,20,30,39,49]` — *different layers*. Here it surfaced as a matmul shape error only because the plane counts disagreed; with a matching count it would have trained on the wrong features silently. - The streaming dataset assumed the draft's aux layers all sit below the base's final layer (`aux = planes[:-1]`, `target = planes[-1]`). A draft whose top aux id *is* the final layer cannot get an extra plane — vLLM captures each layer once — so `final_aux_is_base_hidden` now lets the last plane serve both roles. It is derived from the model, not configured by hand. - DSpark head weights load from either the flat layout ModelOpt exports (upstream DeepSpec convention) or the nested `markov_head.` layout the NVIDIA release uses. Without the remap the two `[131072, 512]` Markov tables — ~14% of the draft's parameters — stay randomly initialized while everything else warm-starts, with no error. - `nemotron_h` is enabled in `_FINAL_NORM_TYPE_BY_MODEL_TYPE`: despite the hybrid stack, `NemotronHModel.norm_f` is a plain RMSNorm, and without the entry the offline/streaming fake base raises instead of reconstructing the distillation target. ### Usage ```yaml dflash: dflash_init_checkpoint: /path/to/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark dflash_draft_attention: causal dflash_attention_sink: true dflash_swa_window_size: 1024 dflash_block_size: 8 dflash_mask_token_id: 990 dflash_architecture_config: target_layer_ids: [1, 5, 19, 29, 41, 51] ``` A full worked example is at `modelopt_recipes/general/speculative_decoding/dspark_nemotron35_warmstart.yaml`. ### Testing **Unit tests** — 124 pass (`test_hf_dflash.py`, `test_hf_dspark.py`, `test_hf_domino.py`, `test_hf_dflash_offline.py`, `test_modeling_final_norm.py`), 32 of them new: causal mask structure (lower-triangular per block, no cross-block leakage, context visibility unchanged), the sink math (degenerates to plain attention at `-inf`, absorbs mass monotonically, receives gradient), warm-start load/reject paths, Markov key remapping, and explicit `target_layer_ids`. **Checkpoint compatibility** — the released drafter loads with zero missing/unexpected keys and zero shape mismatches; all 77 tensors (6 attention sinks and both Markov tables included) match bit-exactly, and a training step runs with gradients reaching the sink and Markov parameters. **End-to-end streaming training** — Nemotron-3.5 base served by vLLM (1 node, TP8) feeding 8 trainer GPUs over NIXL; the draft warm-starts from the released checkpoint and trains with `causal` + sink + SWA 1024. 128 Daring-Anteater conversations, 20 epochs (the plot shows the first 5, where the trend is clearest — the curves flatten after that):  Over the first 5 epochs loss falls **1.85 → 1.36** and train accuracy rises **0.25 → 0.49**; across the full 20 epochs they reach **1.21** and **0.48** (peak 0.54) before flattening. This validates the pipeline end-to-end — capture layers, plane split, mask direction, sink loading and warm-start weights all have to be right for this curve to appear. It is *not* a model-quality result: 128 samples over 20 epochs overfits by construction, and the corpus is not generated by the base model, so the absolute numbers are not meaningful. ### TODO (follow-up) **A complete, robust checkpoint/config converter.** Both conversions are handled ad hoc here: - *Draft config → training config.* The recipe transcribes ~15 fields by hand from the drafter's `config.json`. Only the shape-bearing ones (`num_hidden_layers`, `num_attention_heads`, `intermediate_size`, `markov_rank`) fail loudly when mistyped; the rest — `mask_token_id`, `causal`, `swa_window_size`, `block_size` — train "successfully" on a wrong value and only surface later as a mysteriously low acceptance length. A converter should derive the whole block from the checkpoint, including its aliases (`pard_token`, `dspark_markov_rank`, `dflash_query_causal`, top-level `sliding_window` / `attention_sink_bias`) and duplicated fields. - *Weight layout.* The `markov_head.` remap is a load-time hook. A converter should normalize layouts explicitly, and decide whether export should also emit the release's aliases so a round-trip reproduces the original format (today it renames `architectures` to `DFlashDraftModel`). - *Base config.* Serving this base on vLLM needs its `config.json` layer-type vocabulary updated for the transformers-5 path (`mamba` → `linear_attention`, `attention` → `full_attention`, plus a matching `hybrid_override_pattern`). That is done by hand today and is not covered by this PR. ### Before your PR is "*Ready for review*" - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added configurable causal or bidirectional attention for DFlash models. * Added optional attention sinks, checkpoint warm starts, and explicit target-layer selection. * Improved streaming data handling for shared auxiliary and base hidden states. * Added Nemotron-3.5 Lightning DSpark warm-start training and serving recipes. * **Bug Fixes** * Preserved configured attention behavior during model export. * Prevented warm-start checkpoints from being reapplied during restoration. * Improved checkpoint compatibility, validation, and attention-mask handling. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
fbcdc16c2d |
Remove deprecations marked in 0.45 and 0.46 (#2182)
### What does this PR do?
Type of change: Backward breaking change (deprecation removal)
Ahead of the 0.47 code freeze, this removes every deprecation still
outstanding from the previous two releases (0.45 and 0.46). Two are
intentionally left in place: the **Python 3.10** drop and the
**transformers 4.x** drop
| Deprecation | Marked in | Replacement |
| --- | --- | --- |
| `--auto_quantize_bits` / `_method` / `_score_size` / `_cost_model` /
`_active_moe_expert_ratio` | 0.46 | AutoQuantize `--recipe` |
| `examples/llm_ptq` symlink + `examples/vlm_ptq/` forwarder | 0.46 |
`examples/hf_ptq` (`--vlm` for VLMs) |
| `QuantizationArgumentsWithConfig` alias | 0.45 |
`QuantizationArguments` |
| `QFORMAT_ALIASES` short names | 0.45 | canonical preset basenames |
| `layerwise` bool + flat `layerwise_checkpoint_dir` | 0.45 | nested
`layerwise: {enable, checkpoint_dir}` |
| in-trainer `quant_cfg` / `--quant_cfg` | 0.45 | `--recipe` |
#### Two things worth a closer look
**1. The `use_sequential` alias goes too.** It is the pre-#1251 alias on
`QuantizeAlgorithmConfig.layerwise` and only ever carried a bool. Once
the bool form is rejected it cannot accept a valid value, so keeping it
would only produce a differently-worded validation error. Note the
direction is breaking either way (`extra="forbid"`): a pre-0.45
`modelopt_state` carrying `use_sequential: True` or a top-level
`layerwise_checkpoint_dir` now fails validation instead of being
migrated.
**2. Removing in-trainer `--quant_cfg` required two new recipes.** The
`examples/gpt-oss` QAT flow ran on `--quant_cfg
MXFP4_MLP_WEIGHT_ONLY_CFG` and no `general/ptq/` recipe covered it. This
PR adds `general/ptq/mxfp4_mlp_weight_only` and
`general/ptq/nvfp4_mlp_weight_only`, verified to `model_dump` identical
to `mtq.MXFP4_MLP_WEIGHT_ONLY_CFG` / `mtq.NVFP4_MLP_WEIGHT_ONLY_CFG`,
and migrates the gpt-oss README, both SFT configs, `sft.py` and
`tests/examples/gpt-oss/test_gpt_oss_qat.py`. `examples/llm_qat` was
already recipe-only.
### Usage
```bash
# AutoQuantize: --auto_quantize_* flags -> an AutoQuantize recipe
scripts/huggingface_example.sh --model $HF_PATH \
--recipe general/auto_quantize/nvfp4_fp8_at_5p4bits --calib_batch_size 4
# --qformat / --quant_cfg: short name -> canonical preset basename
# int8_sq -> int8_smoothquant nvfp4_mse -> nvfp4_w4a4_weight_mse_fp8_sweep
# int8_wo -> int8_weight_only nvfp4_local_hessian -> nvfp4_w4a4_weight_local_hessian
# w4a8_awq -> w4a8_awq_beta fp8_pb_wo -> fp8_2d_blockwise_weight_only
# nvfp4_awq -> nvfp4_awq_lite fp8_pc_pt -> fp8_per_channel_per_token
scripts/huggingface_example.sh --model $HF_PATH --quant int8_smoothquant
# VLM PTQ: examples/vlm_ptq -> examples/hf_ptq with --vlm
scripts/huggingface_example.sh --model $HF_PATH --quant fp8 --vlm
# gpt-oss QAT: --quant_cfg <CFG name> -> --recipe <recipe path>
accelerate launch --config_file configs/zero3.yaml sft.py \
--config configs/sft_full.yaml --model_name_or_path openai/gpt-oss-20b \
--recipe general/ptq/mxfp4_mlp_weight_only --output_dir gpt-oss-20b-qat
```
```python
# Layerwise calibration: bool / flat key -> nested LayerwiseConfig
quant_cfg["algorithm"] = {"method": "gptq", "layerwise": {"enable": True, "checkpoint_dir": "/path"}}
```
### Testing
- `tests/unit/recipe` (229 passed),
`tests/unit/torch/quantization/test_config_validation.py` (79 passed),
`tests/examples/hf_ptq/test_hf_ptq_args.py` (23 passed).
- Verified the two new recipes `model_dump` identical to the `mtq.*_CFG`
constants they replace.
- `ruff check modelopt/ examples/ tests/` clean; `ruff format --check`
clean on all changed Python files.
- GPU suites
(`tests/gpu/torch/export/test_unified_hf_export_and_check_safetensors.py`,
`test_accelerate_gpu.py`, `test_gptq.py`) had their preset / layerwise
literals updated but were not run locally — relying on CI.
- `examples/llm_qat/ARGUMENTS.md` is hand-edited to match what the
`generate-arguments-md` hook emits; the generator could not run locally
(missing `transformers` package metadata in this environment).
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ — that is the point of the PR:
it removes shims deprecated in 0.45/0.46. Callers must move to the
replacements in the table above. Additionally, a pre-0.45
`modelopt_state` carrying `use_sequential` or a top-level
`layerwise_checkpoint_dir` will now fail config validation.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — existing tests migrated to
the surviving APIs;
`TestLayerwiseNestedConfig::test_legacy_forms_rejected` pins that the
bool form, the `use_sequential` alias and the flat checkpoint-dir key
are all rejected. Tests covering the removed shims were deleted.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌
### Additional Information
Follow-up: the transformers 4.x drop deprecated in 0.46 is still
outstanding and will need its own PR.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## New Features
- Added MXFP4 and NVFP4 weight-only quantization recipes for MLP and MoE
layers.
- Added shared layer exclusions for more accurate effective-bits
calculations.
## Improvements
- Updated PTQ, QAT, GPT-OSS, deployment, and quantization-format
examples with current recipe names and configuration formats.
- Standardized layerwise settings under nested configuration fields.
## Breaking Changes
- Removed deprecated AutoQuantize options, `quant_cfg` usage, format
aliases, legacy layerwise settings, and compatibility example paths.
- Recipe-based and nested configuration forms are now required.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
58ad6edc5f |
Fix pruned-HF export fallback + add Nemotron-3.5-Lightning launcher examples (#2196)
### What does this PR do? Type of change: Bug fix + new example Two related changes for the Megatron-Bridge Minitron prune/quantize launcher flows: 1. **Fix pruned-HF export crash on containers that reject config-only save.** `#2159` added a config-only HF export path gated only on `hasattr(AutoBridge, "from_hf_config")`. Some Megatron-Bridge versions (e.g. `nemo:26.04`) expose `from_hf_config` but reject a config-only `save_hf_pretrained` (`ValueError: save_hf_pretrained requires a pretrained HuggingFace model`), so `prune_minitron.py` crashed instead of using the intended dummy-model fallback. Now it attempts the config-only save and falls back to the dummy-model path on `ValueError`. 2. **Add Nemotron-3.5-Lightning-30B-A3B launcher examples** (`mbridge_prune.yaml`, `mbridge_quantize.yaml`) on `nemo:26.08`. Prune targets 3B active with an MMLU gate; quantize runs W4A16 NVFP4 4/6 PTQ via the `w4a16_nvfp4_4o6` recipe with `tp_size=1` (static-block NVFP4 MSE is unsupported with TP>1). ### Usage ```shell uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_prune.yaml --yes uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_quantize.yaml --yes ``` ### Testing Verified end-to-end on OCI-HSG: - **Nano prune (`nemo:26.04`)** — exercises the fallback path: config-only save raised the `ValueError`, the fallback caught it and exported via the dummy-model path. `mmlu_10pct_bs32 = 0.5196` (gate 0.50) PASS; vLLM gen PASS. - **Lightning prune (`nemo:26.08`)** — config-only export path: `score = 0.6000` (gate 0.58) PASS, 3.00B active params; vLLM gen PASS. - **Lightning quantize (`nemo:26.08`)** — recipe PTQ + unified-HF export; MMLU `0.7741` (gate 0.75) PASS. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A <!-- launcher example configs + fallback path exercised by CI prune/quantize jobs --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- fix is for a bug introduced in the same unreleased cycle (#2159); rest are example configs --> - Did you get Claude approval on this PR?: ❌ <!-- pending /claude review --> ### Additional Information The fallback fix addresses the `mbridge_prune` launcher CI failure introduced by #2159. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added a pruning workflow for Nemotron-3.5-Lightning-30B-A3B with calibration, quality scoring, checkpoint export, and multi-GPU generation. - Added a four-GPU NVFP4 W4A16 quantization workflow with Hugging Face conversion and MMLU evaluation. - **Bug Fixes** - Improved hybrid model export by falling back to dummy-model export for supported configuration-only export failures. - Added clearer logging and handling for supported export failures while preserving unrelated errors for investigation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c4129b6e03 |
Add Cosmos3 Nano DFlash multimodal training recipe (#2053)
### What does this PR do? Type of change: new example Adds an end-to-end Cosmos3 Nano DFlash training recipe for multimodal speculative decoding. - Adds a notebook that prepares data, launches synthetic generation in Slurm, trains a DFlash draft model, exports it, and provides a vLLM smoke-test command. - Adds PAI-Understanding, VQA v2, and multilingual prompt sharding and distributed-generation helpers. - Adds an atomic, multimodal-safe merge and conservative deduplication flow. - Extends the VLM data collator to handle structured image/video messages, configurable visual bounds, and fixed DFlash sequence lengths. - Hardens generation launch scripts and preserves truncated generated responses. ### Usage ```bash cd examples/speculative_decoding/recipes export MODEL_PATH=/path/to/cosmos3-nano export PLAIN_TEXT_INPUT=/path/to/nemotron-chat-or-approved-user-data.jsonl jupyter lab train_dflash_cosmos3_nano.ipynb Run the notebook in order: 1. Configure paths. 2. Prepare prompts on a CPU-only node and generate target completions in a Slurm GPU allocation. 3. Merge the four required sources and submit training. 4. Export a saved checkpoint and run the vLLM deployment smoke test. ### Testing - jq empty examples/speculative_decoding/recipes/train_dflash_cosmos3_nano.ipynb - bash -n on the modified launch, worker, and recipe shell scripts. - Ran a two-step Cosmos3 Nano DFlash Slurm smoke job; it completed and wrote modelopt_state.pth. - Not run: pytest tests/unit/torch/speculative/plugins/test_hf_speculative_offline.py (pytest is unavailable in the current environment). ### Before your PR is "Ready for review" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md?: N/A - Did you write any new necessary tests?: ✅ - Did you update CHANGELOG.rst?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information Security follow-up required before marking ready: the notebook hardcodes model.trust_remote_code=true and --trust_remote_code. Either parameterize this with a default of false, or obtain and document a security exception. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added end-to-end multimodal workflows for dataset preparation, distributed generation, result merging, training, export, and deployment testing. * Added support for image and video inputs, multiple dataset formats, resumable JSONL generation, configurable serving, and parallel processing. * Added configurable prompt, media, token, sequence, temperature, and tensor-parallel settings. * **Bug Fixes** * Improved truncated-response handling, assistant-label processing, validation, health checks, cleanup, deduplication, media resolution, atomic outputs, and failure reporting. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Slawomir Kierat <skierat@nvidia.com> |
||
|
|
a57fb44d46 |
Preserve HF PTQ checkpoint sidecar files [NV BUG 6491822] (#2060)
### What does this PR do? Type of change: Bug fix In `hf_ptq.py` when exporting a PTQ checkpoint, it would drop some files from the original BF16 checkpoint because it uses a whitelist pattern to allow certain files. However that is brittle and can drop files such as reasoning parsers. Now we make hf_ptq.py match Megatron-Core export behavior by copying all non-safe tensor files, but filter only allowed non-safetensor files for more safety. ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved Hugging Face checkpoint handling to preserve eligible sidecar files while excluding weights, indexes, stale quantization metadata, and unsupported artifacts. * Preserved existing export files and applied consistent file filtering. * Improved snapshot resolution when remote code is disabled. * Ensured unified exports handle generation configuration files correctly. * **Tests** * Added coverage for sidecar copying, exclusions, existing-file preservation, supported file patterns, and snapshot downloads. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> |
||
|
|
686da8d893 |
feat(megatron-bridge): SFT-masked data support in distillation (#2113)
## What does this PR do ?
**Type of change:** New feature
**Overview:** Adds SFT-masked data support to the Megatron-Bridge
distillation example, so a
model can be distilled on prompt/response pairs with the loss masked to
the response.
Today `examples/megatron_bridge/distill.py` only consumes
pretraining-style data — `GPTDataset`
over pre-tokenized blends with `NullTokenizer` — so the loss is computed
over every token. When
distilling an instruction-tuned model it is usually preferable to train
on prompt/response pairs
and mask the loss to the response, matching how the model was
fine-tuned.
## Usage
```bash
python examples/megatron_bridge/distill.py \
--teacher_hf_path <teacher> --student_hf_path <student> \
--sft --sft_dataset_root /path/to/data \
...
```
where `/path/to/data` holds `training.jsonl` / `validation.jsonl` of
records:
```json
{"input": "<prompt>", "output": "<response>"}
```
## How it works
Switches the data path to Bridge's `FinetuningDatasetConfig` (NeMo-style
`GPTSFTDataset`):
* `prompt_template="{input}{output}"` tokenizes input+output verbatim —
adjacent placeholders,
no separator — so the text is fed exactly as provided
* `label_key="output"` with `answer_only_loss=True` masks the loss to
the response
(`answer_start_idx == len(context_ids)`)
* `truncation_field="input"` truncates the context when a pair exceeds
`seq_length`
Two supporting changes, both scoped to `--sft`:
* **Tokenizer.** SFT reads raw text, so it uses the model's real
HuggingFace tokenizer. The
pretraining path consumes pre-tokenized data and keeps `NullTokenizer`.
* **Loss reduction.** A response-only mask requires per-token loss to
combine correctly across
context-parallel ranks, so `calculate_per_token_loss` is enabled and
`average_in_collective`
is disabled. Both are untouched on the pretraining path.
## Testing
Used for quantization-aware distillation of Nemotron-Nano-3 (W4A16
NVFP4) at `seq_length=32768`
with CP>1: 200 iterations, logits-distillation loss `3.37e-2 -> 1.91e-2`
monotonically, router
`seq_load_balancing_loss` steady, and the resulting checkpoint exports
and serves correctly.
Opt-in: without `--sft` the existing mock/blend data path is unchanged.
## Before your PR is "Ready for review"
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes — purely additive and
opt-in behind `--sft`.
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added supervised fine-tuning (SFT) support for distillation workflows.
* Added configuration and validation for SFT dataset locations and
supported inputs.
* Added raw prompt/completion JSONL datasets with response-only loss
masking.
* Added truncation, end-of-sequence handling, and student-tokenizer
support without automatic chat templates or BOS tokens.
* Added matching student and teacher vocabulary validation.
* Preserved existing mock and GPT dataset modes for non-SFT runs.
* **Documentation**
* Documented required filenames, record format, tokenizer behavior, and
completion-only loss masking.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: James Shen <yueshen@nvidia.com>
|
||
|
|
b96841db3e |
Add optional MLflow tracking to the vLLM fake-quant server (#2120)
### What does this PR do? Type of change: new feature Wires `examples/vllm_serve/vllm_serve_fakequant.py` up to `modelopt.torch.utils.mlflow` via `--mlflow <tracking-uri>`, the same way #2023 did for `hf_ptq.py`, so a fake-quant serve records **what it actually quantized** and an evaluation of that endpoint can be traced back to a recipe. Without the flag, behavior is unchanged — every hook is gated on it. Three design points worth review: 1. **The run is recorded in the vLLM worker, not the launcher.** `vllm_serve_fakequant.py` is the API-server frontend; the engine and its workers are separate processes whose stdout it never sees, so a run opened there would capture none of the calibration. The launcher instead only settles the tracking configuration — validating the URI, naming the experiment, recording the command the user actually typed — and publishes it through the environment, which is how every other setting in this example (`QUANT_CFG`, `RECIPE_PATH`, …) already reaches the workers. Global rank 0 opens the run, so a TP-8 serve produces one run. 2. **The run covers load-through-warm-up, not the server's lifetime.** It opens *before the weights load*, so an unreachable server or a missing token fails in seconds rather than after a load and a full calibration, and it closes `FINISHED` once the model is quantized and warmed up. A run that stayed open for the serving lifetime would never close cleanly on SIGTERM. 3. **`recipe/quant_cfg.yaml` is only written on the preset path.** With `RECIPE_PATH`, `get_quant_config` returns the recipe's `quantize` section unchanged and `resolved_recipe.yaml` already carries it. With `QUANT_CFG`/`KV_QUANT_CFG` it is the *only* record of what ran: the params carry the preset names, while the config reaching `mtq.quantize` is those two deep-copied, merged, and — for an MLA model — extended at runtime with `*kv_c_bmm_quantizer` / `*k_pe_bmm_quantizer` by inspecting the loaded model. Uploaded artifacts: | Artifact | Contents | | --- | --- | | `command.txt` | The launcher's invocation, copy-pasteable, credentials masked | | `version.txt` | The ModelOpt version that ran | | `recipe/resolved_recipe.yaml` | `RECIPE_PATH` with its `$import`s expanded | | `recipe/quant_cfg.yaml` | Merged `QUANT_CFG`/`KV_QUANT_CFG` + MLA fixup (preset path only) | | `logs/<script>.log` | The rank-0 worker's stdout/stderr, including a crash traceback | | `summary/quant_summary.txt` | The per-quantizer summary | Plus the quantization *and* serving settings as searchable params, and `user` / `hostname` / `modelopt_version` / `git_sha` / `vllm_version` tags. The `checkpoint_path` tag matches the one `hf_ptq.py` sets, so a checkpoint's PTQ run and every serve of it join up. Two small library additions, both consumed by the new example module: - `command_text(argv=None)` — records another process's invocation, since a spawned worker's own `sys.argv` is vLLM plumbing rather than anything a user typed. - `MlflowRunLogger.log_text()` — uploads a value settled midway through a run, so a crash during calibration still keeps the config that caused it. The example `Dockerfile` installs the `mlflow` extra; the client remains optional and is imported only once tracking is enabled. ### Usage ```bash RECIPE_PATH=<recipe.yaml> python vllm_serve_fakequant.py <model_path> -tp 8 \ --host 0.0.0.0 --port 8000 \ --mlflow https://<your-mlflow-server>/ ``` ``` [mlflow] tracking to https://<your-mlflow-server>, experiment $USER/vllm_serve_fakequant/<model>-<recipe> (Worker_TP0) [mlflow] run: https://<your-mlflow-server>/#/experiments/19/runs/1c6679448f25... ``` `--mlflow-experiment` / `--mlflow-run-name` override the defaults. `$MLFLOW_TRACKING_URI` enables tracking on its own and is best-effort; an explicit `--mlflow` overrides it and fails loudly. > This is the **quantization** tracking server. It is unrelated to any server an evaluation harness exports its scores to — NeMo Evaluator Launcher has its own `export.mlflow.tracking_uri`. The README calls this out. ### Testing **Unit — 87 passing** (`tests/examples/vllm_serve/test_vllm_mlflow_utils.py`, 33 new; `tests/unit/torch/utils/test_mlflow.py`, +5). `vllm_mlflow_utils` deliberately imports no vLLM, so the whole launcher→worker handover is covered without a GPU, a server, or the mlflow client. **End to end on aws-cmh** (4× GB300, `simple_evals.gpqa_diamond`, Nemotron-3.5-Lightning-30B-A3B-BF16 fake-quantized with `general/ptq/nvfp4_mlp_only-kv_fp8_cast`): run `FINISHED` in 261.5 s, opened by `Worker_TP0` only, all artifacts present and verified by content — `command.txt` held the launcher's invocation rather than the worker's spawn argv, and `resolved_recipe.yaml` was 6797 B against 1845 B of source. 104 quantizers enabled (92 NVFP4 dynamic block-16 expert weight/input with calibrated amax, 12 FP8 KV bmm). The eval then ran to completion against the served endpoint, 22/22 requests HTTP 200. Two bugs the hardware run caught, both fixed here with regression tests: - `--mlflow_run_name` was rejected. vLLM's `FlexibleArgumentParser.parse_args` rewrites **every** `--foo_bar` to `--foo-bar` before matching, so a flag registered only under the underscored spelling is unreachable from its CLI. Both spellings are now registered. A unit test on a plain `ArgumentParser` could not have caught this. - `recipe/quant_cfg.yaml` uploaded a Python `repr` blob under a `.yaml` name: a recipe's `quantize` is a `QuantizeConfig`, `yaml.safe_dump` raises `RepresenterError` on it, and the old JSON fallback stringified the object. `_dump_yaml` now unwraps pydantic via `model_dump(mode="json")` and raises otherwise, with the caller downgrading that to a warning so a bad config cannot take down a serve. **Known coverage gap:** the preset (`QUANT_CFG`/`KV_QUANT_CFG`) path — the only one that now writes `recipe/quant_cfg.yaml` — is covered by unit test but has not been exercised on hardware; the canary used `RECIPE_PATH`. Likewise the case where `$MLFLOW_TRACKING_URI` is present *inside* the deployment container and `--mlflow` overrides it is unit-tested only: NeMo Evaluator Launcher forwards only declared env vars, so the eval server's URI never entered the container in the canary. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — new optional flags only; no `--mlflow` means no behavior change. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new dependency. Uses the existing optional `nvidia-modelopt[mlflow]` extra (`mlflow-skinny`, Apache-2.0) added in #2023; the example `Dockerfile` now installs it. No code copied from other sources. - Did you write any new necessary tests?: ✅ — 38 new tests. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — 0.47 Misc. - Did you get Claude approval on this PR?: ❌ — `/claude review` not yet run. ### Additional Information Follows #2023, which added `MlflowRunLogger` and the `hf_ptq.py` integration. Note for anyone tracking from an OCI cluster: `mlflow-modelopt.nvidia.com` is unreachable from oci-nrt and oci-hsg. TCP 443 completes and the connection is then reset on the first application byte, regardless of SNI or protocol, one RTT away — the PDX PaaS ingress appears to apply a source-IP policy, and the OCI clusters egress from Oracle-owned addresses (`155.248.190.0`, `168.110.199.1`) rather than NVIDIA's. gcp-nrt, aws-cmh and cw-dfw all reach it. This is an infrastructure matter, not a property of this change, but it determines where the feature is usable today. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added optional MLflow tracking for vLLM fake-quantization serving runs. * Records serving, quantization, worker, and invocation metadata, including configuration and summary artifacts. * Supports tracking URI, credentials, environment, and command-line configuration. * Added command and text artifact logging for active MLflow runs. * **Documentation** * Documented setup, configuration, recorded artifacts, lifecycle, and fallback behavior. * Updated the example container to include MLflow support. * **Tests** * Added comprehensive coverage for tracking configuration, logging, failures, and disabled tracking. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
6261f854aa |
docs: rebuild the unified HF deployment support matrix from the deploy test suite (NVBug 6550792) (#2087)
### What does this PR do? Type of change: documentation Fixes [NVBug 6550792](https://nvbugspro.nvidia.com/bug/6550792) / OMNIML-5693. The **Unified HF Checkpoint Deployment Model Support Matrix** listed 9 model families and **no VLMs**, while `tests/examples/hf_ptq/test_deploy.py` declares deployment cases for ~80 checkpoints across TRT-LLM, vLLM, and SGLang — including `Qwen2.5-VL`, `Qwen3-VL-235B`, and `Nemotron-3-Nano-Omni`. QA (the filer) could not use the doc to scope testing, and users could not tell what is actually covered. Filing also surfaced that the matrix lived in **three places that had drifted apart**: only the `.rst` listed Qwen3-VL, only the README listed Qwen3.5 MoE, and the skill reference had neither. #### Changes 1. **Rebuilt the matrix in `docs/source/deployment/3_unified_hf.rst`** from `test_deploy.py`, split into language models, vision-language/multimodal, speculative decoding drafters, and diffusion. 2. **Stated plainly what the matrix is and is not.** Review established that the original "CI-validated" framing claimed more than the suite substantiates, so a *What this matrix is based on* section now leads with two limits: - The suite is marked `release` and collects only under `--run-release`, which **no workflow passes** — these are declared cases, not PR-gated coverage. - Each case is a **load-and-generate smoke check on the text path**: no accuracy, no image/audio input, no diffusion output, no verification that speculative decoding engages. The legend follows from that: ✅ = declared in the suite, ⚠ = expected to work but not a suite entry (or an entry that does not exercise the feature the row names), `-` = not in the suite. Sections that would otherwise over-read carry their own qualifiers — VLM rows are labelled text-only smoke coverage, and Medusa and Wan 2.2 are ⚠ with the reason stated. 3. **Removed the two duplicate copies**, replacing them with links, so there is one table to maintain. 4. **Fixed stale prose**: the deployment tabs still claimed FP8-only support on vLLM v0.6.5 and a source build of SGLang main from Jan 2025, both contradicting the version table above them. The TRT-LLM floor moves to v1.2.0, qualified as the oldest version stated rather than the oldest that works. 5. **Dropped the Phi series** from the deployment matrix, following #2115 (NVBug 6563509) and confirmation that Phi-4 is being deprecated. ### Usage N/A — documentation only. ### Testing - `docutils` parse of the modified `.rst`: no warnings or errors from the new content; all 5 tables parse with every cell in the correct column. - Cell contents cross-checked against `test_deploy.py` by AST-parsing the `ModelDeployerList(...)` calls rather than by eye; the scope caveats were each verified against `tests/_test_utils/deploy_utils.py`. - `pre-commit run --files …` passes; `build-docs` green. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — documentation only - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information **Two known follow-ups, neither in scope here:** 1. **Nothing enforces that the doc matrix tracks `test_deploy.py`.** Consolidating to one copy removes the three-way drift but not the doc-vs-test drift; a generator plus a CI check would close it. 2. **The release deployment suite does not run in CI.** Wiring it into per-backend release CI is what would let ✅ mean "verified to pass" rather than "declared". That needs GPU capacity across three backends and should be tracked on its own. **For the filer (@Kenny Kang):** the ✅ cells are the scope the release deploy suite declares, and `test_deploy.py` carries the checkpoint, TP size, and minimum SM version per entry — but please read the legend first, since those cases are not currently executed by CI. --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bee497de03 |
Fix EAGLE3 offline dump skipping all conversations on newer transformers (#2172)
### What does this PR do? Type of change: Bug fix **Fix EAGLE3 offline hidden-state dump silently skipping every conversation on newer `transformers`.** `tokenizer.apply_chat_template(...)` returns a **`BatchEncoding`** (dict of `input_ids` + `attention_mask`) on `transformers>=5` rather than a `list[int]`, so `len(input_ids)` evaluated to **2** (the number of dict fields), tripping the `num_input_tokens <= 10` "too short" filter for **every** conversation. The dump wrote **zero `.pt` files** and offline EAGLE3 training aborted with `No .pt files found`. The token-id extraction is consolidated into `modelopt.torch.speculative.utils.get_conversation_input_ids`, which normalizes the result to a flat `list[int]` (unwrapping `BatchEncoding` / 2-D tensor / batch-wrapped list, asserting the shape so a future `transformers` change fails loudly instead of silently). It is called from all three offline-dump entry points that shared the bug: - `examples/speculative_decoding/collect_hidden_states/compute_hidden_states_trtllm.py` - `examples/speculative_decoding/collect_hidden_states/send_conversations_for_hiddens.py` - `examples/speculative_decoding/scripts/send_conversation_vllm.py` (the two `send_conversation*` scripts additionally indexed/`decode()`d the `BatchEncoding`). Also fixes the `add_generation_template` -> `add_generation_prompt` typo at each site. ### Testing - `tests/unit/torch/speculative/test_speculative_utils.py` — asserts the helper returns the exact token-id sequence of the rendered chat prompt, and pins every `apply_chat_template` return shape (`BatchEncoding`, 2-D tensor, batch-wrapped list, plain list) to a flat `list[int]` via deterministic stubs, so the fixed branch is covered regardless of the installed `transformers` version. - **End-to-end on ComputeLab (H100, TRT-LLM 1.3.0rc20):** reran the exact dump on the 100 conversations that previously failed. Before: 0/100 (0 `.pt` files). After: **97/100** (97 `.pt` files; the 3 skips are genuinely `> max_seq_len`). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ (`tests/unit/torch/speculative/test_speculative_utils.py`) - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: 🔄 `/claude review` run; findings addressed, re-review pending ### Additional Information Surfaced by an nmm-sandbox CI run where `Qwen3-8B_EAGLE3_offline` failed after the container bump to `tensorrt-llm/release:1.3.0rc20`; the auto-blame heuristic mis-attributed it to an unrelated MLflow commit. `compute_hidden_states_vllm.py` is unaffected (it routes through `common.tokenize_with_loss_mask`, which passes `return_dict=True`). --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
10145db53f |
Add link to puzzletron_v2 (#1996)
<!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added a top-level Puzzletron overview with a link to the step-by-step algorithm tutorial. * Added a reference to the experimental Puzzletron branch for advanced, production-scale usage. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Sepehr Sameni <ssameni@nvidia.com> |
||
|
|
f99523279a |
Minitron pruning fixes for Nemotron-3.5-Lightning-30B-A3B and Deepseek (#2159)
### What does this PR do?
Type of change: Bug fix + new feature
Two model families that could not be pruned end-to-end now can:
- **Nemotron-3.5-Lightning**
(`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16`) — a native
`NemotronHForCausalLM` that ships without remote code and carries MTP
heads. Fixes a calibration crash and HF-export failures on the modern
Megatron-Bridge / transformers stack.
- **DeepSeek-V3** — fixes an MLA Q-LoRA crash during calibration, and
adds a `candidate_filter` search option to `mcore_minitron` so its
MoE-FFN dimensions stay prunable while remaining representable in HF.
Also makes a rank-local failure under pipeline parallelism fail fast
instead of stalling.
#### 1. Nemotron Lightning: prune + HF export
(`examples/megatron_bridge/prune_minitron.py`)
1. **MTP calibration crash.** On newer Megatron-LM, `mtp_process` is
derived from the hybrid *pattern*, not from `mtp_num_layers`. Setting
only `mtp_num_layers=0` in the calibration provider overrides was
insufficient: the provider's `finalize()` re-appended the MTP suffix to
`hybrid_layer_pattern` (because `mtp_hybrid_override_pattern` was still
set and `mtp_use_repeated_layer=True`), so `mtp_process=True` while
`mtp_num_layers=0` and the calibration forward hit `assert
self.config.mtp_num_layers > 0`. Fix: also clear
`mtp_hybrid_override_pattern` in the calibration overrides so MTP is
fully disabled (MTP heads are dropped from the pruned model, as before).
2. **HF export via a config-only bridge (hybrid models only).** The old
export built a *dummy* HF model to obtain the bridge, then streamed
weights. This breaks on native NemotronH because (a) native
`NemotronHConfig` makes `hybrid_override_pattern` a read-only property,
and (b) transformers 5.12 saves the input embedding under a different
key than the bridge mapping expects (`backbone.embedding` vs
`backbone.embeddings`); the mismatch made `build_conversion_tasks` drop
the embedding task on its owning rank, leaving an owner-less PP
placeholder that crashed `save_hf_weights` with `Object must exist on at
least one PP rank`. Fix: stream weights through a **config-only** bridge
(`AutoBridge.from_hf_config(hf_cfg).save_hf_pretrained(...)`), available
since Megatron-Bridge 0.5.0 (nemo:26.06). A config-only bridge has
`hf_keys=None`, so the embedding task is never dropped, no dummy model
is built, and the output uses the canonical HF key names.
This is **restricted to hybrid providers**, which are the only models
that need it; non-hybrids keep the dummy-model path that CI has always
exercised.
Writing the source artifacts is now rank-0-only. Every rank used to
write the source `config.json`, which races with the pruned
`config.json` that `save_hf_pretrained` writes from rank 0 alone: a late
write from another rank leaves a checkpoint whose config does not match
its weights. This reproduced intermittently on both Qwen3 and NemotronH
before the fix, and 3/3 clean runs after.
`save_hf_pretrained` takes no `trust_remote_code` argument — it reads
the flag **off the bridge** to fetch the source checkpoint's artifacts,
and `from_hf_config` cannot infer it because
`AutoConfig.from_pretrained` consumes the kwarg rather than storing it
on the config. So the flag is set explicitly on the bridge instance;
otherwise remote-code models would silently lose it.
3. **Config write-back correctness:**
- `hybrid_override_pattern` is only written for older remote-code
configs that lack `layer_types`; native configs carry the cadence in
`layer_types` (read-only `hybrid_override_pattern` is skipped).
- `n_shared_experts` is preserved (a fixed count) instead of being
re-derived by `moe_shared_expert_intermediate_size //
moe_ffn_hidden_size`, which is DeepSeek-style logic that would corrupt
NemotronH's count.
Non-hybrids, VLMs, and Megatron-Bridge builds without config-only export
keep the dummy-model path, with a `warn_rank_0` when a hybrid has to
fall back. The README's `transformers<5` workaround is **removed**: it
existed because the dummy-model path broke on transformers 5, and the
config-only path handles NemotronH on every supported container.
#### 2. `candidate_filter` for `mcore_minitron`
(`modelopt/torch/prune/plugins/mcore_minitron.py`)
DeepSeek-style MoE configs have no explicit shared-expert-size field:
they size the shared expert as `n_shared_experts *
moe_intermediate_size`, where `moe_intermediate_size` is the (also
prunable) **routed** expert size. So only candidates with
`moe_shared_expert_intermediate_size % moe_ffn_hidden_size == 0` can be
written back to HF at all.
Candidates come from a Cartesian `product()` of independent per-hparam
choice lists, so no per-hparam restriction can express a constraint
*between* two hparams. New optional `candidate_filter` search-config key
(default `None`, so existing behaviour is unchanged): a callable that
rejects candidate configs before the metric computation, making the
search cheaper rather than more expensive. It receives **every**
supported hparam, with non-searched ones filled in from the model
config, so a filter still works when one of its hparams was skipped or
had a single choice.
Rejected candidates are not cached, so — like `score_func`, whose cached
scores are reused without re-validation — the filter is assumed
unchanged when resuming from a `checkpoint`.
`prune_minitron.py` wires this up for DeepSeek-style configs, so
**both** `moe_ffn_hidden_size` and `moe_shared_expert_intermediate_size`
stay prunable (the search then only picks shared sizes that are a
multiple of the routed one). A `--prune_export_config` that violates the
constraint never reaches the filter, so the export path now raises
`ValueError` instead of writing a checkpoint whose config disagrees with
its weights.
#### 3. MLA Q-LoRA pruning
(`modelopt/torch/prune/plugins/mcore_minitron.py`)
Pruning any MLA model with `q_lora_rank` set died during calibration
with `AttributeError: 'tuple' object has no attribute 'view'`.
`hidden_size` importance estimation blanket-patches every
`TELayerNormColumnParallelLinear` with `return_layernorm_output=True` to
capture post-layernorm activations. When `q_lora_rank` is set, MCore
builds `linear_q_up_proj` as a `TELayerNormColumnParallelLinear` — the
Q-LoRA layernorm is fused into it, which is why `q_layernorm` is
`IdentityOp` — so it was patched too, even though its layernorm is over
the **latent rank**, not `hidden_size`. TE then returns `((out, ln_out),
bias)` and MCore's `q, _ = self.linear_q_up_proj(...)` leaves `q` a
tuple.
Isolated by probing the module before and after dynamic conversion:
| Setup | `linear_q_up_proj` returns | Forward |
| --- | --- | --- |
| Before conversion | `tuple(Tensor, NoneType)` | — |
| After conversion, no hooks | `tuple(Tensor, NoneType)` | OK |
| After conversion **+ importance hooks** | `tuple(tuple(Tensor,
Tensor), NoneType)` | AttributeError |
So conversion is innocent; registering the importance hooks is the
trigger. Fix: exclude MLA's Q/KV up-projections from both the patch and
unpatch loops. `test_mcore_mla_pruning` did not catch this because it
builds MLA without `q_lora_rank`, where MCore uses a plain
`linear_q_proj` and nothing is patched.
#### 4. Fail fast instead of stalling on a rank-local error under PP
(`modelopt/torch/utils/distributed.py`)
A rank raising inside a distributed entrypoint left the whole job
stalled until the process group timed out, with **no diagnostic output
at all**: the failing rank blocked in `cleanup()`'s barrier while its
peers blocked in `recv_from_prev_pipeline_rank_`, and Python only prints
a traceback once the enclosing `finally` returns. A crash on one rank
was indistinguishable from a slow job.
- `dist.cleanup()` skips the barrier when unwinding from an exception.
- New `dist.abort()` prints the traceback, flushes and exits
immediately. Skipping the barrier alone is **not** enough — a stack dump
showed the failing rank then blocking in `destroy_process_group` for the
same reason — so the error path must not tear the process group down at
all. `SystemExit` is re-raised rather than aborted, so an intentional
exit (e.g. the `--score_lower_bound` gate) keeps its exit code and
prints no traceback. Kept out of `cleanup()` so no library caller gets a
surprise process exit.
- Called from the entrypoints that wrap `main()` in `try/finally`: the
five `examples/megatron_bridge` scripts.
Measured on a 2-GPU PP run whose rank 0 raises during calibration: **10
min timeout kill with no visible error → 31s, exit 1, real traceback.**
This is a latent, pre-existing issue (the `try/finally` predates this
PR); it only surfaces on a failing PP run, which is why CI never hit it.
### Usage
```bash
torchrun --nproc_per_node 4 examples/megatron_bridge/prune_minitron.py \
--hf_model_name_or_path nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 \
--pp_size 4 \
--prune_target_active_params 3e9 \
--output_hf_path /path/to/Nemotron-3.5-Lightning-30B-A3B-Pruned-A3.0B
```
### Testing
- **End-to-end on nemo:26.08.rc6** (4× GB300, transformers 5.12.1,
Megatron-Bridge with config-only export): pruning + export complete
(`EXIT=0`, "Saved pruned model … Done!"). The exported checkpoint has
canonical **plural** `backbone.embeddings.weight` keys, **0 MTP
tensors**, and a config that reloads correctly (`num_hidden_layers=52`
from `layers_block_type`, `n_shared_experts=1`,
`num_nextn_predict_layers=0`, pruned `hidden_size`/`mamba_*`/MoE dims,
reconstructed `hybrid_override_pattern`).
<details>
<summary>Pruning search log (<code>--prune_target_active_params
3e9</code>)</summary>
```text
Top 10 Candidates with Scores
┏━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━┓
┃ # ┃ export_config ┃ active_params ┃ params ┃ score ┃
┡━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━┩
│ 1 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56,
'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 23.49B │ 0.5406
│
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size':
3584} │ │ │ │
│ 2 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 56,
'mamba_head_dim': 48, 'num_moe_experts': 96, │ 3.00B │ 20.09B │ 0.2427 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 3 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48,
'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 21.61B │ 0.2643
│
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 4 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48,
'mamba_head_dim': 64, 'num_moe_experts': 96, │ 3.00B │ 19.28B │ 0.4552 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size':
3712} │ │ │ │
│ 5 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64,
'mamba_head_dim': 64, 'num_moe_experts': 104, │ 3.00B │ 22.28B │ 0.5860
│
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 6 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48,
'mamba_head_dim': 48, 'num_moe_experts': 96, │ 3.00B │ 21.99B │ 0.2294 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size':
3328} │ │ │ │
│ 7 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 56,
'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 23.68B │ 0.5231
│
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 8 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56,
'mamba_head_dim': 56, 'num_moe_experts': 96, │ 3.00B │ 21.81B │ 0.5042 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size':
3584} │ │ │ │
│ 9 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48,
'mamba_head_dim': 56, 'num_moe_experts': 96, │ 3.00B │ 20.09B │ 0.2462 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 10 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64,
'mamba_head_dim': 64, 'num_moe_experts': 96, │ 3.00B │ 20.70B │ 0.5685 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
└────┴───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┴───────────────┴────────┴────────┘
╭────────────────────────────────────────────────────────────────────────
Best Subnet
─────────────────────────────────────────────────────────────────────────╮
│ export_config {'num_layers': 52, 'hidden_size': 2304,
'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 104,
'moe_ffn_hidden_size': 1856, │
│ 'moe_shared_expert_intermediate_size': 3072} │
│ active_params 3.00B │
│ params 22.28B │
│ score 0.5860 │
╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
╭────────────────────────────────────────────────────── Pruned Model
Stats ───────────────────────────────────────────────────────╮
│ Total Parameters 22.28B │
│ Active Parameters 3.00B │
│ Memory (BF16, seq_length=8192, batch_size=8) weights: 42489.7 MB,
kv_cache: 384.0 MB, mamba_state: 190.5 MB, Total: 43064.2 MB │
╰─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
```
</details>
- **`tests/examples/megatron_bridge/test_prune_minitron.py`** —
`nemotron_h` now exports to HF and reloads (previously it stopped at a
Megatron checkpoint, since the dummy-model path needed
`transformers<5`), plus an `n_shared_experts` config assertion; the dead
`megatron_format` branch is gone. It runs on the CI container:
**verified on nemo:26.06.01 (transformers 5.8.1) and nemo:26.08.rc6.**
-
**`tests/gpu_megatron/torch/prune/plugins/test_mcore_mamba_minitron_pruning.py`**
— the `nas_memory_mb` search test now passes a `candidate_filter` and
asserts the exact number of rejected candidates (256 of the 512-combo
grid) plus the surviving candidates' validity; its `expected_top_k`
goldens are regenerated accordingly. Because
`moe_shared_expert_intermediate_size` is in that test's skip list, this
also covers the model-config fallback for hparams that are not in the
search space.
Verified on 2 GPUs, on both the CI container (nemo:26.06.01) and
nemo:26.08.rc6:
| Test | Result |
| --- | --- |
| `test_prune_minitron[qwen3]` | PASSED on 26.06.01 and 26.08.rc6 |
| `test_prune_minitron[deepseek_v3]` | PASSED (52s) — MLA Q-LoRA +
`candidate_filter` end-to-end |
| `test_prune_minitron[nemotron_h]` | PASSED on 26.06.01 (58s) and
26.08.rc6 (61s) |
| `test_mcore_mamba_hybrid_pruning_nas_memory_mb` | PASSED |
| `test_mcore_mamba_hybrid_pruning_nas_params` | PASSED (unchanged
sibling, run to check the regenerated goldens did not disturb it) |
| 2-GPU PP run failing on rank 0 | fails in 31s with a real traceback
(was a 10 min stall) |
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `candidate_filter` defaults
to `None` (existing searches unchanged), and the config-only export is
limited to hybrid providers on nemo:26.08+, so dense / MoE / VLM exports
keep the path they use today.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ✅
### Additional Information
Enables the Prune + Distill workflow for
`NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16` (native, no-remote-code
`NemotronHForCausalLM` with MTP heads). Pruning-time MTP support was
scoped and intentionally deferred — MTP heads are dropped and can be
re-derived via a short SFT with `mtp_num_layers=1` on the
pruned+distilled model.
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
a21173a4d5 |
Bug fix: 6542481 (#2064)
### What does this PR do? Type of change: Bug fix Fixes `AssertionError: Model already has modelopt state!` when exporting a QLoRA checkpoint (NVBug 6542481). The QLoRA output is adapter-only, so `from_pretrained` resolves the quantized base model and already restores the ModelOpt state; `export.py` then restored a second time. Fixing that exposed two more breakages on the same path, also fixed here: - `_restore_qtensor_wrappers` missed every module — PEFT renames the compressed linears to `<name>.base_layer`, so no weight got re-wrapped and the packed NVFP4 weight hit a shape error. - `postprocess_state_dict` dropped `weight_scale_2` (missing from the QLoRA rename map), leaving the exported checkpoint impossible to dequantize. ### Usage No API change — `examples/llm_qat/export.py --pyt_ckpt_path <qlora_ckpt> --export_path <out>` now completes on the documented quantize → train → export flow. ### Testing Reproduced in the reported environment (TRT-LLM 1.3.0rc22, transformers 5.5.4, NVFP4). - Added the missing export step to `test_qwen3_qlora_nvfp4` and a unit test for the QLoRA `base_layer` rename; both fail without the fix. - Exported base model is byte-identical to a plain PTQ export; dequantized NVFP4 weights match the bf16 original (worst rel. error 0.10). - No regressions: `tests/gpu/torch/export/test_export.py` (49 passed), save/load plugin tests. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ❌ — can add if wanted - Did you get Claude approval on this PR?: ❌ — not run yet ### Additional Information Fixes NVBug 6542481. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved QLoRA checkpoint export and restoration across supported model configurations. * Preserved secondary weight-scale information and other deployment tensors in exported checkpoints. * Corrected handling of quantized base-layer weights after adapter reparenting. * Prevented duplicate state restoration when checkpoints already include the required model state. * Removed internal adapter prefixes and quantizer details from exported state data. * **Tests** * Added validation for packed weights, quantization metadata, required scales, and removal of embedded adapter layers. * Added regression coverage for QLoRA state processing and quantized weight restoration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9220fac053 |
[NVBug: 6563509] Drop Phi-3-vision / Phi-4-multimodal PTQ support (#2115)
### What does this PR do?
Type of change: Deprecation
Resolves [NVBug 6563509](https://nvbugspro.nvidia.com/bug/6563509),
where
`hf_ptq.py` on Phi-4-multimodal-instruct died with
`RuntimeError: Tensor.item() cannot be called on meta tensors`.
The crash is real but not fixable on our side, and it is not the reason
the model
is unusable. Phi-4-multimodal's bundled remote code predates
Transformers v5 and
does not load on **any** version in our supported range
(`transformers>=4.57,<5.15`):
| Blocker | Where |
|---|---|
| `peft.get_peft_model` reads `prepare_inputs_for_generation`, gone
since transformers 4.52 dropped `GenerationMixin` from `PreTrainedModel`
| `modeling_phi4mm.py:1959` |
| `_tied_weights_keys` declared as a list; Transformers 5.x calls
`.keys()` on it in `post_init` | `modeling_phi4mm.py:1937` |
| `int(torch.tensor(...))` in `__init__`, which cannot run on a meta
device — the reported crash | `speech_conformer_encoder.py:1435` |
The model card pins `transformers==4.48.2` / `peft==0.13.2`, so there is
no
overlap with our floor and nothing on our side can bridge it. The model
is
therefore dropped rather than worked around.
**Phi-3-vision is dropped alongside it because it is the older,
superseded model
in the same family** — with its successor unsupportable there is no
reason to
keep carrying the predecessor. This is a product-scope call, not a
separate
compatibility finding: Phi-3-vision shares the list-valued
`_tied_weights_keys`
defect (`modeling_phi3_v.py:1214`) and so is likewise broken on
Transformers 5.x,
but it does **not** hit the `peft` blocker, and it was not re-verified
on 4.57.
Per the 0.46 changelog we have already bumped the floor to 4.57 and
noted that
"Transformers 4.x support will be dropped in a future release", so any
remaining
window closes on its own. Same reasoning already applied to VILA / NVILA
in this
release.
**Removed**
- the support-matrix row in `examples/hf_ptq/README.md`
- `"Phi4MMForCausalLM": "phi4mm"` from `MODEL_NAME_TO_TYPE`
- the multimodal-detection heuristics that only ever matched these two —
`vision_lora`, `audio_processor`, `embd_layer.image_embd_layer`, and the
`phi4mm` model-type check — in both `is_multimodal_model` and
`_is_multimodal_config`
- the `Phi3Image` / `PhiImage` exclusions in `is_embedding`
- the phi4mm input-mode warning in `hf_ptq.py`
- `modelopt_recipes/huggingface/phi4mm/` and its references in
`modelopt_recipes/ptq.md`
**Not changed:** the device-map sizing path (meta-device skeleton,
`infer_auto_device_map`, and the `--gpu_max_mem_percentage` cap) keeps
its
original behavior. That cap is wanted exactly where it already fires —
when the
model is already offloading to CPU, where it costs little and the
headroom is
required. With the affected checkpoints removed, there is no supported
model
that trips the meta-device build, so there is nothing to work around
here.
Text-only **Phi-3/Phi-4** and **Phi-3.5-MoE** are natively supported by
transformers and are untouched.
### Testing
On H200, `nvcr.io/nvidia/tensorrt-llm/release` (torch 2.12, transformers
5.5.4),
against the real checkpoint:
- **Version matrix** (vanilla transformers, no modelopt) — Phi-4-MM
loads at
4.48.2 / 4.49.0 / 4.50.0 / 4.51.3 and fails at 4.53.3 / 4.56.2 / 4.57.1
(`AttributeError: 'Phi4MMModel' object has no attribute
'prepare_inputs_for_generation'`) and at 5.5.4 (meta-init, then
tied-keys).
This is what establishes that no supported version works.
- `tests/examples/hf_ptq/test_example_utils.py` — 28 passed.
- **Sweep**: `tests/examples/hf_ptq` + `tests/unit/torch/export` —
failure set
identical to the pre-change tree (GPU/model-dependent `test_vlm_ptq`,
plus
`test_quant_aware_conversion` scoped-mapping tests), so none are
introduced
here.
- `pre-commit` clean on all changed files, including recipe validation.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ — PTQ for Phi-3-vision and
Phi-4-multimodal is removed, along with the `huggingface/phi4mm/ptq/*`
recipes. Phi-4-multimodal is already unloadable on every supported
transformers
version, so no working workflow regresses; Phi-3-vision is a deliberate
scope
removal as its superseded predecessor.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — this is a deletion; the
existing
`test_get_model_*` / `test_resolve_init_config_*` tests are unchanged
and still
pass.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
Two related references were left in place deliberately; say the word and
I'll
fold them in:
- `tests/examples/hf_ptq/test_deploy.py` still deploys the
already-published
`nvidia/Phi-4-multimodal-instruct-{NVFP4,FP8}` checkpoints. Those
artifacts
exist and serve fine; this PR only removes the ability to *produce*
them.
- `examples/torch_onnx/README.md` still lists Phi-4-multimodal-instruct.
That
is a separate ONNX pipeline that does not go through `get_model()` and
was not
tested here.
Earlier revisions of this branch also reworked the device-map sizing so
the
meta-tensor crash could not occur. That was reverted in
|
||
|
|
9b8caf623a |
Use lm-eval 0.4.12's built-in trtllm backend, deprecate lm_eval_tensorrt_llm.py (#2066)
### What does this PR do?
Type of change: documentation / example update (with a behaviour fix)
lm-evaluation-harness **0.4.12** is the first release that ships a
TensorRT-LLM backend
(`lm_eval.models.trtllm_causallms`, registered as `trtllm`) — it is
absent in 0.4.10 and
0.4.11. This example no longer maintains its own, so:
- Pin `lm_eval[api,ifeval]>=0.4.12,<0.5` (the 0.5.0.dev line drops the
file) and bump
`lm_eval_hf.py`'s version guard to match.
- **Delete** `examples/llm_eval/lm_eval_tensorrt_llm.py` (the `trt-llm`
model). Replace
`python lm_eval_tensorrt_llm.py --model trt-llm --model_args
tokenizer=<tok>,checkpoint_dir=<ckpt>`
with `python lm_eval_trtllm.py --model trtllm --model_args
model=<ckpt>,tokenizer=<tok>`.
- Add `examples/llm_eval/lm_eval_trtllm.py`, whose entire content is one
corrected
`_parse_logprobs` plus `cli_evaluate()` (see below). `lm_eval_hf.py`
stays HF-only.
- `examples/hf_ptq/scripts/huggingface_example.sh` and the docs use the
upstream backend.
`parser.sh` gains `--input` (`BUILD_MAX_INPUT_LEN`, default 4096) — it
already *echoed*
that variable but never parsed or defaulted it, so it printed empty on
every run.
#### Why `lm_eval_trtllm.py` exists: an upstream off-by-one
TensorRT-LLM aligns `prompt_logprobs` to the *next* token.
`executor/base_worker.py`:
```python
# Pass prompt_token_ids with an offset of 1 for correct mapping to the context logits
prompt_token_ids = generation_result._generation_request.prompt_token_ids[1:] + first_generation_token
```
So entry `i` is the distribution that predicted `tokens[i + 1]`, and
`_topk_logprobs`
appends that token's id when it is not in the top-k. lm-eval's
`_parse_logprobs` instead
reads `prompt_logprobs[i][tokens[i]]` and applies its own shift on top,
which raises
`KeyError` on the **first request of every loglikelihood task**
(hellaswag, mmlu, arc, ...):
```
File ".../lm_eval/models/trtllm_causallms.py", line 324, in _parse_logprobs
current_token_logprob = prompt_logprob[tokens[i]]
KeyError: 6503
```
Probed against TRT-LLM 1.3.0rc23 with a 14-token prompt for
`prompt_logprobs` 0, 1 and 2:
`tokens[i]` is missing at **every** position, `tokens[i+1]` is present
at every position.
Only `generate_until` tasks work unpatched. **This wants an upstream
issue against
EleutherAI/lm-evaluation-harness.**
The override also fails loudly rather than quietly: it checks
`prompt_logprobs` covers
every prompt token and raises on a missing token, instead of skipping
the term and
silently inflating the reported accuracy.
#### Defaults that must be set explicitly
`TRTLLM.__init__` accepts `**kwargs` but forwards only a fixed set to
the **TensorRT-LLM
`LLM` API**, so extra `--model_args` aimed at the engine are silently
dropped. (lm-eval's
own named parameters — `max_gen_toks`, `batch_size`, `truncation_side`,
... — are honored
normally.) Two engine defaults are unsafe for few-shot eval:
- `tensor_parallel_size` defaults to **1** (the deleted wrapper used
every visible GPU).
- `max_input_len` defaults to **2048**, and longer prompts are silently
left-truncated —
5-shot MMLU/gsm8k prompts exceed that.
### Usage
```bash
python lm_eval_trtllm.py --model trtllm \
--model_args model=<quantized checkpoint dir>,tokenizer=<HF model folder>,tensor_parallel_size=<tp>,max_batch_size=<bs>,max_input_len=4096,max_output_len=512 \
--tasks hellaswag,gsm8k \
--batch_size <bs>
```
Flat arguments (no `run` subcommand) are what 0.4.12's
`HarnessCLI.parse_args` inserts
`run` for automatically (`_cli/harness.py:48-51`); this is the exact
command form used for
the results below.
### Testing
**Unit** — `tests/examples/llm_eval/test_lm_eval_trtllm.py`, no GPU and
no `tensorrt_llm`
install: stubs the response object and pins the `i-1` alignment, the
`rank != 1` →
`is_greedy` rule, the `ctxlen=0` edge, and both `RuntimeError` paths.
Mutation-checked —
dropping the `-1` shift is caught by 5/5 cases, ignoring `ctxlen` by
4/5. A sixth test is a
**tripwire**: it asserts lm-eval's own implementation is still
misaligned, so a future
0.4.x that fixes the bug fails the test and says to delete this file
rather than being
silently re-broken by the override.
**End to end** — `nvidia/Qwen3.5-122B-A10B-NVFP4` (NVFP4 MoE, 256
experts) on **4x B300**,
TRT-LLM 1.3.0rc23, lm-eval 0.4.12, `--limit 32`:
| run | hellaswag acc | hellaswag acc_norm | gsm8k flexible | gsm8k
strict |
|---|---|---|---|---|
| deleted impl (`trt-llm`), tp=4 | 0.7188 | 0.7812 | 0.8438 | 0.7812 |
| `lm_eval_trtllm.py`, tp=1 | 0.7188 | 0.7812 | 0.8438 | 0.8125 |
| `lm_eval_trtllm.py`, tp=2 | 0.7188 | 0.7812 | 0.9062 | 0.8125 |
| `lm_eval_trtllm.py`, tp=4 | 0.7188 | 0.7812 | 0.8750 | 0.8438 |
- hellaswag (the loglikelihood path this PR fixes) is **identical at
every tp and identical
to the deleted implementation** — the alignment fix is exact, not
approximate.
- gsm8k varies by 1–2 samples out of 32 (generation path: upstream uses
native `stop=`
sequences and per-request `SamplingParams`; the old wrapper used
beam-search-of-1 with
post-hoc string truncation).
- Without the override, every hellaswag run above dies with the
`KeyError`.
- Re-verified at tp=4 after the code moved out of `lm_eval_hf.py` into
`lm_eval_trtllm.py`.
Note: NVFP4 fused-MoE has no CUTLASS tactic on Hopper (`No supported MoE
GEMM tactic
remains after replacing unsupported NO_SMEM epilogues.`), so this had to
be validated on
Blackwell.
### Feature parity notes
Gained from upstream: `loglikelihood_rolling` (was
`NotImplementedError`), pipeline
parallelism, `add_bos_token` auto-detection, prompt truncation,
per-request sampling params,
`prompt_logprobs` instead of full-vocab context logits (much lower
memory), thinking-tag
handling, `batch_size=auto`.
Not reachable through the upstream backend (were set by
`modelopt.deploy.llm.LLM`):
`enable_attention_dp` for MoE, `CudaGraphConfig`,
`enable_chunked_prefill`,
`moe_expert_parallel_size=1`, and `free_gpu_memory_fraction=0.7` with a
capped
`kv_cache.max_tokens` — upstream uses the TRT-LLM default 0.9 (observed
allocating 218 GiB
of paged KV cache on B300), so OOM risk is higher on smaller GPUs. This
is documented in
`examples/llm_eval/README.md`, and `huggingface_example.sh` honours a
preset `LM_EVAL_TP`
so users can lower the tensor-parallel size without editing the script.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ — `lm_eval_tensorrt_llm.py` is
removed and the CLI changes (`--model trt-llm` → `trtllm`,
`checkpoint_dir=` → `model=`). Migration command is in the README and
CHANGELOG.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependency; existing `lm_eval` pin tightened.
- Did you write any new necessary tests?: ✅ —
`tests/examples/llm_eval/test_lm_eval_trtllm.py` (6 cases, no GPU).
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.47 *Deprecations*.
- Did you get Claude approval on this PR?: ✅ — reviewed, feedback
addressed in `dcedd37b4` and `622b97c26`.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added TensorRT-LLM evaluation through lm-evaluation-harness’s `trtllm`
backend.
* Added configurable input/output lengths, batching, tensor parallelism,
and build input length.
* Improved prompt log-probability alignment for more accurate evaluation
results.
* **Documentation**
* Updated evaluation instructions, truncation guidance, backend
limitations, and configuration examples.
* **Deprecations**
* Removed the legacy TensorRT-LLM evaluation script and entry point.
* **Updates**
* lm-evaluation-harness now requires versions 0.4.12 through 0.4.x.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
bd3798a794 |
[6562078]: fix calibration for vLLM 0.26.0 (#2093)
### What does this PR do? Type of change: Bug fix Fixes calibration failures when running fake-quantization with vLLM 0.26.0. Two root causes are addressed: 1. **`finish_requests` must be called explicitly after `add_requests`.** In vLLM 0.26.0 the scheduler calls `finish_requests` *before* `add_requests` inside `execute_model`, so request IDs are not registered yet and cleanup never runs. The calibration loop now calls `finish_requests` directly after each batch using `dataclasses.replace`. Wrapped in `try/finally` with an inner `try/except` so it always runs and never masks the original exception. A warning is emitted when `finish_requests` is absent so the regression is self-diagnosing on future vLLM API changes. 2. **`NewRequestData` gained a `prefill_token_ids` field.** vLLM 0.26.0 added this required argument; the calibration helper now passes it. Additional: - Dockerfile updated to vLLM 0.26.0 with `USER vllm` (non-root). - README updated to include vLLM 0.26.0 in tested versions. ### Usage ```bash cd examples/vllm_serve QUANT_CFG=FP8_DEFAULT_CFG QUANT_CALIB_SIZE=8 CALIB_BATCH_SIZE=1 \ python3 vllm_serve_fakequant.py Qwen/Qwen1.5-MoE-A2.7B-Chat -tp 1 \ --host 0.0.0.0 --port 8000 ``` ### Testing Tested end-to-end FQ calibration with vLLM 0.26.0 using the Docker image built from `examples/vllm_serve/Dockerfile`. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (examples change only) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information Changes are confined to `examples/vllm_serve/` and do not affect the core library. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Summary by CodeRabbit * **New Features** * Added compatibility with vLLM 0.26.0 in the serving example. * Improved calibration request handling and cleanup after model execution. * Enabled the serving container to run with a non-root user. * **Bug Fixes** * Calibration cleanup failures no longer obscure the original model execution error. * Added warnings when calibration cleanup cannot be completed. * **Documentation** * Updated the serving example documentation to list vLLM 0.26.0 among tested versions. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> |
||
|
|
ccf44ea67a |
[OMNIML-5562] Add FAR3D ONNX PTQ and accuracy evaluation example (#2012)
### What does this PR do? Type of change: new example Adds an end-to-end FAR3D ONNX PTQ example under `examples/onnx_ptq/far3d`. The example prepares Argoverse 2 validation metadata and calibration batches, quantizes the image encoder to INT8, builds TensorRT engines, and evaluates 3D object detection accuracy. It also provides a reproducible FAR3D runtime image, preserves accuracy-sensitive encoder layers in high precision, and supports temporal decoder state during evaluation. ### Usage ```bash python prepare_metadata.py /path/to/av2 python prepare_calibration.py /path/to/far3d.py far3d_calibration python quantize.py far3d.encoder.onnx far3d_calibration python evaluate.py /path/to/far3d.py \ far3d.encoder.int8.engine far3d.decoder.fp16.engine ``` ### Testing - Ran all configured pre-commit hooks on the changed files. - Ran synthetic calibration-reader and graph-exclusion tests. - Built the documented FAR3D runtime image and verified TensorRT 10.11, Argoverse 2 imports, and TensorRT engine execution. - Ran the complete workflow on the Argoverse 2 validation split using 500 calibration batches and 23,522 evaluation frames. The INT8 encoder and FP16 decoder produced 0.238 mAP. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information - [NVIDIA DL4AGX FAR3D TensorRT reference](https://github.com/NVIDIA/DL4AGX/tree/master/AV-Solutions/far3d-trt) > 🤖 _Generated by Codex (AI agent)._ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an end-to-end FAR3D ONNX post-training quantization workflow for Argoverse 2, including calibration batch generation, INT8/FP8 quantization, TensorRT engine inference, and mAP evaluation. * Added a dedicated FAR3D example Docker environment plus detailed README instructions. * Added FAR3D evaluation, calibration preparation, metadata preparation, and quantization scripts. * **Bug Fixes** * Improved handling when flash-attention is unavailable, with clearer error messaging. * **Chores** * Updated pre-commit exclusions and refreshed third-party license attribution. * Added an Experimental changelog entry for the FAR3D example. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com> |
||
|
|
22b6a148b0 |
Fix EAGLE-3 context-parallel training and re-enable its tests (#2086)
### What does this PR do?
Type of change: Bug fix
**EAGLE-3 context-parallel training (`--cp_size > 1`) is fixed, and its
tests run again.** CP has been broken since `accelerate` 1.13, and the
tests never caught it: the guard compared `Version("2.10.0a0")` against
`Version("2.10.0")`, which is False on every NGC alpha torch build, so
`test_llama_eagle3[cp_size=2]` has never actually run in CI.
Five fixes:
- **`main.py`** — rebuild the FSDP2 plugin accelerate requires for
`cp_size > 1`. The `--fsdp full_shard --fsdp_config` launcher flags that
used to supply it were dropped from `launch_train.sh`, so CP could not
start at all. Also pass the CP degree to the draft model.
- **`modeling_eagle.py`** — apply the draft model's first input norm
inside `layers[0]`'s own forward, where FSDP2 has actually unsharded its
weights, and only stash the input embeds on the path whose pre-hook
consumes them.
- **`hf_eagle.py`** — skip the dense eagle attention mask under CP
(causal masking comes from `is_causal`, TTT masking from the
ring-attention patch), and warn that padded positions are therefore
unmasked. Also stop `(eagle_loss or 0)` replacing a `0.0` loss tensor
with a plain `int`, which detached the graph.
- **`eagle_utils.py`** — key TTT-mask injection off the backward call's
`grad_out` kwarg, since newer torch omits `attn_bias` on the forward
call, silently disabling TTT masking.
- **`utils.py`** — CUDNN-only SDPA under CP; the `MATH` backend
decomposes SDPA and breaks on DTensors. Scoped to `cp_size > 1`, since
this context manager wraps every training forward and CPU has no cudnn
backend.
**Drops the `speculative_decoding` 26.01 container override.** It was
added when the lane ran 25.06 and spec-dec needed something *newer* — a
floor. Later bumps moved the default past it, so it had silently become
a ceiling holding spec-dec on a 6-month-old image.
### Testing
Ran `tests/examples/speculative_decoding` in
`nvcr.io/nvidia/pytorch:26.07-py3` on 2 GPUs, reproducing the CI install
steps (`pip uninstall -y nvidia-modelopt`, `pip install -e
".[hf,dev-test]"`, example requirements): **16 passed, 2 skipped** — the
2 skipped being pre-existing `--run-manual` tests. All four
`test_llama_eagle3` cases pass, including both `cp_size=2` ones.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — the existing `cp_size=2`
tests are re-enabled
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — not yet run
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
6a8102591e |
Single gpu disk offload PTQ for DSR1/Ultra (#2008)
### What does this PR do? Type of change: New feature, bug fix, new tests Enables single-GPU PTQ for models too large to fit in VRAM (e.g. Nemotron-Ultra-550B at 1.1 TB BF16, DeepSeek-R1 at 642 GB BF16) by adding accelerate disk/CPU offload support to the HF PTQ example and fixing the export path to correctly handle offloaded models. **G1 Offload-aware unified HF export (`modelopt/torch/export/unified_export_hf.py`)** The existing `_export_transformers_checkpoint` removed accelerate hooks before materializing weights, silently writing meta tensors (empty weights) to the checkpoint. Fix: - `_has_accelerate_offload(model)`detects any disk/CPU-offload accelerate hook in the model tree. - `_process_quantized_modules_offloaded(model, dtype)` new export path for offloaded models: materializes one decoder layer at a time via `enable_weight_access_and_writeback`, dispatches export handlers inside the context window, and snapshots the layer state dict before hooks re-offload the weights. A second pass collects non-decoder modules that are also disk-offloaded (embed, norm, lm_head) to avoid meta tensors in the returned state dict. Hooks are removed only after the full state dict is assembled. - Meta-tensor guard in `_export_quantized_weight` raises `RuntimeError` on meta input instead of silently corrupting the checkpoint. **G2 Disk-offload CLI (`examples/hf_ptq/hf_ptq.py`, `example_utils.py`)** Three new arguments to `hf_ptq.py`: - `--offload_folder PATH` enable accelerate disk offload; shards spill here. - `--max_gpu_memory_gb N` VRAM budget for the accelerate device map. - `--max_cpu_memory_gb N` CPU RAM budget for the accelerate device map. Validation: `--offload_folder` is incompatible with `--low_memory_mode` and `--use_seq_device_map`. **G3 Streaming shard writer for 80 GB CPU RAM (`modelopt/torch/export/unified_export_hf.py`)** The G1 path accumulated the entire quantized state dict in CPU RAM before writing (~764 GiB for Ultra 550B), blocking the 80 GB target. New streaming path writes shard files layer-by-layer. Peak memory = 1 decoder layer + 1 shard buffer instead of the full checkpoint: | Model | Old peak CPU RAM | New peak CPU RAM | |-------|-----------------|-----------------| | Ultra NemotronH 550B | ~764 GiB | ~57 GB | | DeepSeek-R1 | ~630 GiB | ~55 GB | Key pieces: - `_StreamingShardWriter(export_dir, max_shard_size)` buffers tensors up to `max_shard_size` bytes, flushes to numbered temp files (`__shard_part_NNNNN.safetensors`), renames to canonical shard names at `finalize()`, writes `model.safetensors.index.json`. Single-shard exports produce `model.safetensors` with no index file. - `_postprocess_single_tensor(key, value, ...)` per-tensor extraction of `postprocess_state_dict` logic (KV amax scale, skip/rename, squeeze) for streaming use. - `_parse_shard_size(size)` converts `"10GB"` / `"500MB"` strings to bytes. - `_export_transformers_checkpoint_streaming(model, dtype, export_dir, max_shard_size)` streams decoder layers via `enable_weight_access_and_writeback`, applies per-tensor postprocessing + name reversal + tied-alias filter, writes shard files directly. Non-decoder offloaded modules and GPU-resident tensors are handled in separate passes. - `export_hf_checkpoint` dispatches to the streaming path when `_has_accelerate_offload(model)` is true; `hf_quant_config.json`, quant-config name reversal, and `config.json` update are shared between paths. `export_hf_checkpoint` accepts a new `max_shard_size` parameter (default `"10GB"`) that controls the shard size for both paths. **Supporting changes** - `modelopt/torch/quantization/plugins/huggingface.py` `get_nemotron_h_decoder_layers` now checks both `model.backbone.layers` (remote-code variant) and `model.model.layers` (native HF variant), fixing layer discovery for NemotronH when loaded without `trust_remote_code`. - `modelopt_recipes/general/ptq/nvfp4_experts_only-kv_fp8_layerwise_offload.yaml` new recipe combining NVFP4 W4A4 on MoE experts, FP8 KV cache, and layerwise calibration with `calib_mutates_weights: false` (required for disk-offload compatibility). - `example_utils.py` `_FP8BF16Fallback` shim: dequantizes block-scaled FP8 expert weights to BF16 for calibration forward passes when the `kernels` package is unavailable (e.g. DSR1 on nodes without finegrained FP8 kernel support). ### Usage ```python # Single-GPU PTQ for a model too large to fit in VRAM, using disk offload python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path /path/to/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 \ --recipe general/ptq/nvfp4_experts_only-kv_fp8_layerwise_offload \ --export_path /path/to/output \ --offload_folder /path/to/offload \ --max_gpu_memory_gb 170 \ --max_cpu_memory_gb 500 \ --trust_remote_code \ --calib_size 8 --batch_size 1 --skip_generate ``` ### Testing **Unit tests** (`tests/unit/torch/export/test_offload_export.py`, 7 tests, CPU-only): - `_has_accelerate_offload` detection (true/false/nested-module cases) - `_export_quantized_weight` meta-tensor guard (raises on meta, passes on real) - `_process_quantized_modules_offloaded` with disk-offloaded embed + GPU-resident decoder layer: verifies no meta tensor in returned state dict **GPU integration tests** (`tests/gpu/torch/export/test_offload_export.py`, 2 tests): - Tiny 2-layer LLaMA with CPU offload: FP8 quantization + export, asserts no meta tensors and valid `hf_quant_config.json` - Same with layerwise FP8 (`calib_mutates_weights=False`): disk-offload path end-to-end ## End-to-end validation Verified with DSR1 that the non-layerwise path provide identical checkpoint before and after this change, also the layerwise with cpu off-load path produce same identical checkpoint (with same max calibration setting). Two production-scale checkpoints were quantized end-to-end using the new disk-offload PTQ path on a single GB200 GPU (189 GiB VRAM). ### DeepSeek-R1 (671B, MoE) | | | |---|---| | **Checkpoint** | `DeepseekV3ForCausalLM`, 671B params, 61 decoder layers | | **Input size** | 642 GB BF16 | | **Recipe** | `nvfp4_experts_only-kv_fp8_layerwise_offload` | | **`--max_gpu_memory_gb`** | 80 | | **`--max_cpu_memory_gb`** | 80 | | **`--calib_size` / `--batch_size`** | 8 / 1 | | **`--trust_remote_code`** | no (built-in transformers) | | **Wall-clock** | 40 min 12 s (load ~14 min, calib ~12 min, export ~14 min) | | **Peak GPU memory** | 88.9 GB | | **Peak process RSS** | 376 GB | | **Output** | 40 shards x ~10 GB = 403 GB (~37% compression) | <img width="1783" height="2532" alt="image" src="https://github.com/user-attachments/assets/fe515035-c880-453f-af1c-2d98395c9197" /> ### NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 (550B, NemotronH MoE + Mamba) | | | |---|---| | **Checkpoint** | `NemotronHForCausalLM`, ~550B params, 108 decoder layers | | **Input size** | ~1.1 TB BF16 | | **Recipe** | `nvfp4_experts_only-kv_fp8_layerwise_offload` | | **`--max_gpu_memory_gb` / `--max_cpu_memory_gb`** | 170 / 500 and **80 / 80** | | **`--calib_size` / `--batch_size`** | 8 / 1 | | **`--trust_remote_code`** | yes (`NemotronHForCausalLM`) | | | 170 GB GPU / 500 GB CPU | **80 GB GPU / 80 GB CPU** | |---|---|---| | **Wall-clock** | 41 min 13 s | **47 min 16 s** | | **Peak GPU memory** | 165.8 GB | **76.7 GB** | | **Peak RSS (load)** | 789 GB transient | **345 GB transient** | | **Steady-state RSS** | ~454-496 GB | **~50 GB** | | **Output** | 34 shards x ~11 GB = 365 GB | 34 shards x ~11 GB = 365 GB | <img width="1783" height="2532" alt="image" src="https://github.com/user-attachments/assets/da4771b8-b522-4549-8e40-7f975bd6f9b1" /> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added disk/CPU/GPU memory-limited offload model loading. * Added offload-aware streaming Hugging Face checkpoint export with sharded output. * Added an NVFP4 expert-only PTQ recipe with FP8 KV-cache support. * Improved Nemotron-H model layout support. * **Bug Fixes** * Improved DeepSeek bundled-code selection based on remote-code trust. * Strengthened handling of meta/offloaded weights, tied-weight deduplication, and export post-processing. * **Tests** * Added coverage for offload exports and DeepSeek loading behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Fridah-nv <fridah@nvidia.com> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
fed1980b29 |
Pin Accelerate below 1.14 for llm_qat (#2067)
### What does this PR do? Type of change: Bug fix Pins Accelerate below 1.14 for the `llm_qat` example. Accelerate 1.14 introduced an FSDP2 regression for models whose input embeddings and output head share a parameter; the shared weight can be assigned to two FSDP groups and training fails before the first step. The pin is example-local. The project-wide Hugging Face dependency range and unrelated examples remain unchanged. ### Usage No usage change. Installing the `llm_qat` requirements now resolves Accelerate to the existing supported range below 1.14. ### Testing - Reproduced the duplicate shared-parameter FSDP2 failure with Accelerate 1.14.0 on two ranks. - Verified Accelerate 1.13.0 completes a distillation training step with the otherwise-identical environment and a clean `main` source tree. - Verified the combined requirements resolve to `accelerate>=1.0.0,<1.14` and select 1.13.0. - `pre-commit run --files examples/llm_qat/requirements.txt` - `git diff --check` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — dependency-only change validated by an exact two-rank A/B run. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ — draft PR; automated review is pending. ### Additional Information This is a scoped compatibility pin while the upstream Accelerate regression remains unresolved. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Constrained the Accelerate dependency to versions below 1.14 to improve compatibility for the LLM question-answering example. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
2d4be28818 |
fix(hf_ptq): use no_grad instead of inference_mode in export_quantized (NVBug 6537702) (#2047)
### What does this PR do? Type of change: Bug fix Fixes [NVBug 6537702](https://nvbugspro.nvidia.com/bug/6537702) / [OMNIML-5658](https://jirasw.nvidia.com/browse/OMNIML-5658) — multi-node FSDP2 PTQ export fails on all ranks: ``` hf_ptq.py:910 export_quantized -> export_hf_checkpoint unified_export_hf.py:1446 _export_transformers_checkpoint -> get_model_state_dict modelopt/torch/opt/_hooks.py:88 _get_model_state_dict_with_dm_check torch/distributed/checkpoint/state_dict.py:481 _get_model_state_dict torch/nn/modules/module.py:2160 _save_to_state_dict destination[prefix + name] = param if keep_vars else param.detach() RuntimeError: Cannot set version_counter for inference tensor ``` **Root cause.** `export_quantized` wrapped its whole body in `torch.inference_mode()`. On the FSDP2 path (`--use_fsdp2`), `get_model_state_dict(full_state_dict=True)` gathers the full params *inside* that context, so the gathered tensors are inference tensors. Inference tensors have no version counter, so the subsequent `state_dict()` → `param.detach()` raises. **Fix.** Use `torch.no_grad()` for the export context. It still disables autograd, but the gathered params stay normal tensors with an intact version counter, so `detach()` works. FSDP2-only failure — the non-FSDP2 path never hit it because its params already exist outside the context. The one-line fix is originally by @shengliangx (`b0e4328` on `shengliangx/distributed-unified`); this PR retargets it to the post-rename `examples/hf_ptq/` path and adds a changelog entry and a regression guard. ### Usage ```bash # 2 nodes x 8 GB200, previously failed at export on every rank torchrun --nnodes=2 --node_rank=0 --master_addr=$MASTER --master_port=6000 --nproc_per_node=8 \ hf_ptq.py --model Llama-3.1-8B-Instruct --dataset cnn_dailymail \ --recipe general/ptq/fp8_default-kv_fp8 --batch_size 8 --calib_size 512 \ --export_path ./Llama-3.1-8B-Instruct-fp8_default-kv_fp8 --use_fsdp2 ``` ### Testing - End-to-end on 2 nodes by @shengliangx on the original branch: dense Qwen3-8B and Qwen3-30B-A3B (MoE) FSDP2 PTQ fp8 checkpoints export successfully. - Added `tests/examples/hf_ptq/test_export_quantized_context.py`, a CPU-only guard asserting `export_quantized` enters `torch.no_grad()` and not `torch.inference_mode()`. A functional regression test would need a 2-node FSDP2 job, which CI does not run, so this encodes the invariant instead. - `pre-commit run --files` clean on all three changed files. Reporter (Kenny Kang, GPU SWQA) still needs to confirm on the original 2x8 GB200 Llama-3.1-8B repro. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Keyword `Committed_ModelOpt_0.46.0` on the bug — should land for 0.46. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed multi-node quantized model exports to prevent runtime errors when gathering and detaching parameters. * Improved compatibility with FSDP2 during Hugging Face PTQ exports. * **Tests** * Added coverage to verify the export process uses the compatible gradient context. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Shengliang Xu <shengliangx@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5dde396bdf |
Fix vLLM 0.24+ compatibility: registry TypeError and MoE RoutedExperts port (#2054)
### What does this PR do? Type of change: Bug fix vLLM 0.24 (as shipped in `nemo:26.08`) reworked the fused-MoE layer, which broke ModelOpt in two ways: 1. **`FusedMoE` became a factory function** returning a `MoERunner` pipeline. Registering it put a plain function into `QuantModuleRegistry`, so *every* later registry lookup raised `TypeError: issubclass() arg 2 must be a class, ...` — taking down large parts of `tests/gpu_megatron` (TE, Megatron chaining, MSE calibrator) that have nothing to do with vLLM. 2. **Expert weights moved onto a `RoutedExperts` submodule** of `MoERunner`, and `UnquantizedFusedMoEMethod` moved modules, so the MoE fakequant path had no valid registration target. Changes: - `_DMRegistryCls.register` now asserts keys are `nn.Module` subclasses, so a future upstream change fails at the registration site instead of as a confusing `TypeError` at lookup. - Register `RoutedExperts` (vLLM >= 0.24) while keeping the `FusedMoE` / `SharedFusedMoE` registrations for older releases; both go through the same `_QuantFusedMoEBase`. `MoERunner` calls `forward_modular` / `forward_monolithic` directly (`RoutedExperts.forward` raises by design), so those are hooked instead of `forward`. - The fused-MoE kernel patch now covers `experts.triton_moe` in addition to `fused_moe`. 0.24's launcher binds the kernel names at import time, so patching only the defining module would leave fakequant **silently inactive**. - `UnquantizedFusedMoEMethod` is resolved from either module layout. - `examples/vllm_serve/vllm_reload_utils.py`: quantizer module paths are now `mlp.experts.routed_experts.*`, so HF→vLLM expert key mapping inserts a matching `.routed_experts` infix when that layout is present. - CI `gpu_vllm` now runs on **two** containers: `v0.24.0` (first release with the FusedMoE-factory / `RoutedExperts` layout — `v0.24.1` was never released, `nemo:26.08` ships a `0.24.1.dev0` build of the same layout) and `v0.20.0`, which keeps the legacy `FusedMoE`/`SharedFusedMoE` branches covered. Test fixes for the newer vLLM (not product bugs): - The tiny Llama fixture used `hidden_size=32 / 16 heads` → `head_dim=2`, which `FLEX_ATTENTION` (the only backend available in this image) rejects with `NYI: embedding dimension ... must be at least 16`, killing the engine core at warmup. Now `head_dim=64`, matching the Qwen3-MoE fixture. - The FlashInfer metadata-builder stub used `causal=False`, which in 0.24 forces the FI-native path (`all_uses_trtllm = causal and ...`) requiring workspace buffers and real wrapper planning. Keep it on the all-TRTLLM path it was originally exercising; the stashed `_modelopt_*` fields are path-independent. ### Usage No API change — existing `mtq.quantize` / `examples/vllm_serve` flows work unmodified on both old and new vLLM. ### Testing All runs in the `nemo:26.08.rc3` container (vLLM `0.24.1.dev0+gee0da84ab`, the same 0.24.1 the CI job now pins). **Suites** - `tests/gpu_vllm`: **73 passed, 1 skipped** (was 70 passed / 3 failed). - `tests/gpu_megatron`: all pass (previously ~120 failures, all from the registry `TypeError` — TE, Megatron chaining and MSE-calibrator tests that never touch vLLM). - `tests/unit/torch/opt/test_dynamic.py`: 2 passed, including the new `test_register_rejects_non_module_classes` (rejects a factory function and a non-`nn.Module` class, and asserts no partial registration). - `pre-commit` clean on all touched files. **MoE fakequant verified by module-tree probe, not just by test assertions** After `mtq.quantize(..., NVFP4_DEFAULT_CFG)` inside the vLLM worker, every weight-owning module was enumerated on tiny DeepSeek-V3 (MLA + routed MoE + shared experts) and tiny Qwen3-MoE: ``` model.layers.0.mlp.experts.routed_experts [QuantRoutedExperts] w13_input_quantizer=3.484 w13_weight_quantizer=0.0840 w2_input_quantizer=0.1060 w2_weight_quantizer=0.0845 model.layers.0.mlp.shared_experts.gate_up_proj [QuantMergedColumnParallelLinear] ✅ model.layers.0.mlp.shared_experts.down_proj [QuantRowParallelLinear] ✅ ``` Weight amax being populated (not just input amax) means the `B is self.w13_weight` identity check and the Parameter-swap weight-fakequant branch actually execute through 0.24's kernel path — i.e. `forward_modular`/`forward_monolithic` really are the live entry points and the `experts.triton_moe` patch target is the one that fires. Unquantized modules were only the expected ones: embeddings, RMSNorms, MoE router `gate`, `lm_head`. Registration parity vs. older vLLM: Row/Column/MergedColumn/QKV `ParallelLinear` and all four attention types (`Attention`, `CrossAttention`, `EncoderOnlyAttention`, `MLAAttention`) register unchanged; `FusedMoE` → `RoutedExperts`; `SharedFusedMoE` has no counterpart because the `shared_fused_moe` module no longer exists in 0.24 — shared experts are now a plain MLP whose linears we already quantize (confirmed above). **Known gaps (pre-existing, not regressions from this PR)** - `DeepSeekV2FusedQkvAProjLinear` is not quantized: it subclasses `MergedColumnParallelLinear` but overrides `forward`, so the registry's shared-forward rule declines it. Pre-0.24 the equivalent (`q_a_proj` / `kv_a_proj_with_mqa`) were `ReplicatedLinear`, which ModelOpt never quantized — effective coverage is unchanged. - MoE fakequant hooks only the Triton expert kernels; FlashInfer/CUTLASS/DeepGEMM MoE backends bypass them (why the fixtures pin `moe_backend="triton"`). - `_setup` still requires a plain `UnquantizedFusedMoEMethod`; a `FusedMoEModularMethod` swap (some DP/all2all configs) still asserts. **Not covered by tests** - `examples/vllm_serve/vllm_reload_utils.py` — the expert key mapping is now asserted in `test_tiny_qwen3_moe_quantize` against the quantizer module paths of a booted MoE model, so a stale infix fails loudly instead of silently serving uncalibrated experts. The rest of the reload path is still inspection-only. Note the registry key moved `vllm_FusedMoE` → `vllm_RoutedExperts` and quantizer paths gained `.routed_experts`, so a `modelopt_state` saved under an older vLLM will not restore onto 0.24 as-is. - The legacy `FusedMoE`/`SharedFusedMoE` branches are covered by the second CI entry; the `v0.20.0` job is green on this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — all new paths are feature-detected; older vLLM keeps the `FusedMoE` registration. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — existing `tests/gpu_vllm` coverage exercises the new registration path. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ ### Additional Information Found while bumping the Megatron test environment from `nemo:26.06` to `nemo:26.08.rc3`. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added compatibility for newer vLLM MoE implementations, module layouts, and routed-expert configurations. * Improved model reload support for models using routed-expert submodules. * **Bug Fixes** * Improved detection and patching of vLLM MoE execution paths across supported configurations. * Registry validation now rejects invalid module registrations without partially applying changes. * **Tests** * Expanded GPU coverage for vLLM 0.24.0, dynamic module validation, and causal attention metadata paths. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
9e3425de05 |
Fix saving pruned Nemotron-3-Nano hybrid_override_pattern with MTP or Pipe symbols (#2061)
Saving pruned Nemotron-3-Nano (with MTP) to HF format raised an assertion which is fixed here Tested on nemo:26.04 with transformers 4.57 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved exported hybrid layer patterns for pruned models by removing MTP and pipeline-parallel markers. * Ensured exported configurations accurately represent the model’s main layers. * **Tests** * Added coverage for pruning models with an MTP prediction layer and hybrid override patterns. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
77dbeb1872 |
Add optional MLflow tracking to hf_ptq.py (#2023)
### What does this PR do? Type of change: new feature Adds `modelopt.torch.utils.mlflow.MlflowRunLogger`, a reusable helper for recording a script run on an MLflow tracking server, and wires `examples/hf_ptq/hf_ptq.py` up to it via `--mlflow <tracking-uri>` so a PTQ run can be reproduced from its MLflow entry alone. Without the flag, behavior is unchanged — every hook is gated on it. The logger lives in the library rather than the example so other scripts can record runs the same way: it takes a tracking URI, an experiment name and an explicit `enabled` flag, with params, tags and artifacts passed in. `hf_ptq.py` supplies only the PTQ-specific pieces (its params, the resolved recipe, the quantization summaries). `mlflow` is an optional dependency, imported only once tracking is enabled, so it is not a new requirement for the library. The run is opened **before the model loads**, so a bad URI or an unreachable server fails in seconds rather than after hours of calibration. The invocation and the recipe are uploaded at that point too, which keeps a crashed run useful: it is still recorded, with status `FAILED` and its log attached. Uploaded artifacts: | Artifact | Contents | | --- | --- | | `command.txt` | The full invocation, copy-pasteable | | `version.txt` | The ModelOpt version that ran (also a searchable tag) | | `recipe/resolved_recipe.yaml` | The `--recipe` with `$import`s expanded | | `logs/hf_ptq.log` | Everything the run printed, including a crash traceback | | `summary/quant_summary.txt` | Per-quantizer summary (unless `--no-verbose`) | | `summary/moe.html` | Per-expert calibration token counts, when the run produces them | Plus model / format / calibration settings as searchable params, and `user` / `hostname` / `modelopt_version` / `git_sha` tags. Three design points worth review: 1. **The recipe is uploaded resolved, not verbatim.** A recipe may be a directory or use `$import`s, so the source file is not self-contained. For `huggingface/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_at_6p0bits-active_moe` the source is 2,230 B / 58 lines against 7,563 B / 308 lines resolved — the raw file records under 30% of what actually ran. 2. **`hf_ptq.py` has no logging framework** (bare `print()`), so the log is produced by teeing stdout/stderr. Handlers that libraries bound to `sys.stderr` at import time are re-pointed at the tee for the run's duration and handed back afterwards; without that, `transformers` / `huggingface_hub` warnings reach the console but never the log. Native (C-level) output is still not captured — documented in the README. 3. **The recipe upload lives in the caller, not the library.** That keeps `modelopt.recipe` out of `modelopt.torch.utils`, which would otherwise risk a `modelopt.torch.utils` → `modelopt.recipe` → `modelopt.torch.quantization` → `modelopt.torch.utils` import cycle. 4. **MLflow failures never fail the quantization.** Startup validation is fatal by design (it is before any GPU work); the end-of-run upload is best-effort. Only the main rank uploads, so `--use_fsdp2` runs produce a single run. ### Usage ```bash python hf_ptq.py \ --pyt_ckpt_path <huggingface_model_card> \ --recipe general/ptq/nvfp4_default-kv_fp8_cast \ --export_path <quantized_ckpt_path> \ --mlflow https://<your-mlflow-server>/ ``` ``` [mlflow] experiment: $USER/hf_ptq/<checkpoint basename>-<recipe name> [mlflow] run: https://<your-mlflow-server>/#/experiments/13/runs/c243352e... ``` `--mlflow_experiment` and `--mlflow_run_name` override the defaults (`$USER/hf_ptq/<basename>-<recipe name or --qformat>`, and the UTC start time). Passing `--mlflow` with no value uses `$MLFLOW_TRACKING_URI`. Authentication uses MLflow's own env vars. ### Testing **Unit** — 51 tests in `tests/unit/torch/utils/test_mlflow.py` for the library, plus 13 in `tests/examples/hf_ptq/test_hf_ptq_args.py` for the hf_ptq wiring. CPU-only, no network and no `mlflow` dependency (driven against a stub module). Covers experiment-name derivation and sanitization, URI accept/reject, tee pass-through, the pre-bound-handler redirect, artifact renaming, skipping absent optional outputs, the disabled path, and `version.txt`. 85 tests pass together with the existing `test_hf_ptq_args.py` / `test_example_utils.py`. **Hardware** — real PTQ runs against a live MLflow server: | Run | Result | | --- | --- | | Qwen3-0.6B, NVFP4 PTQ, 1×B200 | `FINISHED`, all artifacts, sane post-quant generations | | Qwen3.6-35B-A3B MoE, AutoQuantize `w4a16_nvfp4_fp8_at_6p0bits-active_moe`, 2×B200 | `FINISHED` in 63 min, search hit `effective bits: 6.00`; 106 KB log capturing every per-layer decision, 4.4 MB quant summary | | Qwen3.6-35B-A3B, plain NVFP4 PTQ, 2×B200 | `FINISHED` | | Qwen3-0.6B re-run after the library move, 1×H200 | `FINISHED`, all five artifacts including `version.txt` | | Run **without** `--mlflow` after the review fixes | exactly 1 `[load_recipe]` line and 0 `[mlflow]` lines, confirming the untracked path is untouched | | Two runs sharing one `--export_path`, second crashed early | second run uploads **no** summary — the first run's 124 KB file on disk is correctly not attributed to it, and its traceback is in the log | | Crash mid-run (gated HF dataset) | `FAILED` recorded with log + traceback attached, summaries correctly absent | | Malformed URI | Rejected by `argparse` with a `Did you mean https://…?` hint | | Unreachable host | Fails in 9.9 s total, before any model load | | No `--mlflow` | Exit 0, no MLflow output, unchanged export | **Coverage gap, stated plainly:** `summary/moe.html` is verified only against a synthetic file (unit test + a real upload). It could not be produced naturally — `expert_token_count` buffers live on `_QuantSparseSequentialMoe`, while Qwen3.5/3.6 experts take the fused `_QuantFusedExperts` path, so no such file is written for these models regardless of `--moe_calib_experts_ratio`. The uploader's conditional is correct; the branch simply had no natural input available here. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — new optional flags only; no `--mlflow` means no behavior change. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — adds `mlflow` as an optional extra in `pyproject.toml` (`nvidia-modelopt[mlflow]`, folded into `all`) and to `examples/hf_ptq/requirements.txt`. Apache-2.0 (permissive). Imported lazily, so it is not required to install or import ModelOpt. No code copied from other sources. - Did you write any new necessary tests?: ✅ — 29 new unit tests. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — 0.47 New Features. - Did you get Claude approval on this PR?: ❌ — `/claude review` not yet run. A self-review was done first and its six findings are fixed in the third commit (the notable one: gathering the MLflow inputs re-read the recipe on *every* run, including without `--mlflow`). ### Additional Information The one deliberate coverage gap is `summary/moe.html`, described under Testing: no model available here takes the sparse-sequential MoE path that writes it, so it is covered by unit test and a synthetic upload rather than a natural one. The uploader treats it as an optional output and skips it when absent, which is exercised by test. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
93b9e4b176 |
Add Megatron-Bridge prune & quantize launcher pipelines (#2031)
### What does this PR do? Type of change: new example + small launcher / modelopt-example features (backward compatible) Adds end-to-end ModelOpt **launcher** pipelines for the Megatron-Bridge flow on Nemotron-3-Nano-30B-A3B, the minimal launcher features to run them wrapper-free from YAML, and an **in-step accuracy gate** for Minitron pruning. **New launcher examples** (`tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/`): - `mbridge_prune.yaml` — Minitron prune **with an in-step MMLU gate** → vLLM sanity gen (2 tasks) - `mbridge_quantize.yaml` — FP8 quantize → unified-HF export → MMLU gate on the vLLM backend, which doubles as the deploy sanity check (3 tasks). Matches the [tutorial](https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md). **Prune accuracy gate** (reuses the search's own score — no separate eval step): - `modelopt/torch/prune/plugins/mcore_minitron.py` — `MCoreMinitronSearcher` now stores the exported best `CandidateSubnet` under `state_dict["best"]` (additive; sits beside the existing `sorted_layers` key). - `examples/megatron_bridge/prune_minitron.py` — new `--score_lower_bound`: reads `pruning_scores["best"].score` and exits non-zero if the exported model is below the floor. Score-agnostic (any `--prune_score_func`); rejected with `--prune_export_config` (manual pruning has no score). **Launcher (`tools/launcher`)** — run single-node Megatron-Bridge one-liners directly from YAML: - `SandboxTask.inline` — a command in the YAML, no `common/**/*.sh` wrapper (single-line; folded scalar) - `SandboxTask.reqs` / `reqs_file` — pip-install deps in the container before the command (shell-safe; on Slurm the install is rank-0-guarded so multi-rank tasks don't race) - `SlurmConfig.docker_user` — local-Docker user (e.g. `root`); ignored on Slurm - `get_default_env` honors `HF_HOME` / `TRITON_CACHE_DIR` env overrides, so a non-CI user can point caches at a writable path (the shared `/cicd/hf-cache` is owned by the CI account) - reject `args` together with `inline` **`examples/llm_eval/lm_eval_hf.py`**: - `--accuracy_lower_bound` — gate on the single requested task's `acc` (used by the quantize MMLU step; exits non-zero if below) - drop ModelOpt (hf-only) args for non-`hf` backends, so `--model vllm` works on a deployable quantized checkpoint ### Usage ```bash cd tools/launcher # Prune (in-step MMLU gate) -> vLLM gen uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/mbridge_prune.yaml --yes # FP8 quantize -> unified-HF export -> MMLU gate (vLLM) uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/mbridge_quantize.yaml --yes ``` ### Testing - Launcher unit tests for `inline`, `reqs`/`reqs_file`, `docker_user`, the `args`+`inline` guard, and example-resolve; ruff / mypy / bandit clean. `tests/examples/megatron_bridge/test_prune_minitron.py` now passes `--score_lower_bound=0.01` to exercise the gate path on the tiny models. - **End-to-end on the real Nemotron-3-Nano-30B-A3B (4×B200, OCI-HSG):** - Prune 30B → 3B-active: `[score_gate] mmlu_10pct = 0.5196 >= 0.45 PASS`; vLLM gen coherent. - FP8 quantize → unified-HF export (`Detected ModelOpt fp8 checkpoint`) → MMLU on the vLLM backend `acc = 0.7077 >= 0.60 PASS`. - Earlier smoke on **Qwen3-0.6B** in `nemo:26.06` through the same flow. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (new state-dict key is additive; new CLI args default to off) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- tooling/examples + additive searcher state key --> - Did you get Claude approval on this PR?: ❌ <!-- pending --> ### Additional Information - **Container pinning:** saving a pruned Nemotron-H to HF requires `transformers<5`, so `mbridge_prune.yaml` runs on `nemo:26.04` (26.06 drops it); quantize/export run on `nemo:26.06`. - **`docker_user: root`** is set on all example tasks — local-Docker only (ignored on Slurm), needed so downstream tasks can read task_0's root-owned checkpoints and to read the image's root-only `/opt/Megatron-Bridge`. - The quantize MMLU step passes `enforce_eager=True` to vLLM — for a run-once eval this skips ~17 min of CUDA-graph capture / `torch.compile` with no accuracy change. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added inline shell commands and per-task Python dependency installation for launcher workflows. - Added configurable Docker user selection and preservation of existing cache environment settings. - Added evaluation accuracy and pruning score gates that fail workflows below configured thresholds. - Added NVIDIA Nemotron pruning and quantization workflow examples. - Improved backend-specific handling of ModelOpt options. - **Documentation** - Documented inline commands, dependencies, variable substitution, and configuration examples. - **Bug Fixes** - Strengthened task execution validation and configuration checks. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |