mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
main
522
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fadbf74d31 |
Make the QAT/QAD guide the central place for concepts, background, and framework selection (#2590)
### What does this PR do? Make the QAT/QAD guide the central place for concepts, background, and framework selection. Have the Hugging Face and Megatron Bridge tutorials link back to it instead of repeating explanations of QAT and QAD, keeping the tutorials focused on setup and execution. In main QAT/QAD guide make links to all relevant blogposts. Note: MBridge example doc is out of scope for this MR. ### Testing Doc changes only, manual check. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - Did you write any new necessary tests?: N/A docs changes only ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Expanded the QAT/QAD guide with workflows, use cases, and a comparison, including QAD’s use of a frozen BF16 teacher and logit-level loss to recover accuracy after quantization. * Updated README and quick-start navigation to link to the combined QAT/QAD guide; the previous standalone QAT guide now redirects readers there. * Reorganized the LLM QAT tutorial: recipe guidance is now part of the end-to-end example, while trainer examples and Python quantize-and-fine-tune guidance are in Advanced Topics. The tutorial also notes Triton accelerated kernels. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com> |
||
|
|
4f62d418f4 |
Use TensorRT optimization level 0 in Torch ONNX example tests (#2619)
### What does this PR do?
Type of change: Bug fix
The Torch ONNX example tests can exceed their 300-second deadline while
building a ResNet50 INT8 TensorRT engine at optimization level 4. Add
`--trt_builder_optimization_level` to the vision example and select
level 0 in the existing integration tests. The example and helper retain
level 4 by default. Quantization, ONNX export, residual Q/DQ assertions,
engine execution, and test timeout limits are unchanged.
Document the build-time versus inference-performance tradeoff in the
example README.
### Usage
```bash
cd examples/torch_onnx
python torch_quant_to_onnx.py \
--timm_model_name resnet50 \
--recipe timm/resnet/ptq/int8 \
--onnx_save_path resnet50.int8.onnx \
--calibration_data_size 1 --no_pretrained \
--trt_build --trt_builder_optimization_level 0
```
### Testing
Validation used the TensorRT 26.05 container, TensorRT 10.16.1.11, and
PyTorch 2.13.0, with the existing 300-second per-test deadline.
- RTX 6000 Ada: **8 passed**, covering FP8 and INT8 on ViT, Swin,
SwinV2, and ResNet50. ResNet50 INT8 passed in 96.65 seconds.
- RTX PRO 6000 Blackwell Max-Q: **20 passed, 3 existing skips**,
covering the complete test file. ResNet50 INT8 passed in 81.27 seconds;
the baseline timed out at 300 seconds.
```bash
# RTX 6000 Ada: supported FP8/INT8 cases
python -m pytest tests/examples/torch_onnx/test_torch_quant_to_onnx.py \
-k '(fp8 or int8) and not mxfp8' --cov
# RTX PRO 6000 Blackwell: complete example test file
python -m pytest tests/examples/torch_onnx/test_torch_quant_to_onnx.py --cov
```
On each GPU, levels 4 and 0 used the same exported ResNet50 INT8 graph
with TensorRT 10.16.1.11:
- RTX 6000 Ada: TensorRT-reported engine build time decreased from 122.6
seconds at level 4 to 27.7 seconds at level 0. Both builds and inference
runs succeeded.
- RTX PRO 6000 Blackwell Max-Q: the original level-4 test hit its
300-second deadline during the engine build; level 0 built that saved
graph in 10.9 seconds and completed inference successfully.
All pre-commit checks passed for the changed files. The existing
integration tests exercise the real engine build; no redundant mocked
tests were added.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — the existing integration
tests were updated to exercise level 0; no new test cases were needed.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — minor example/CI fix; library behavior and example defaults are
preserved.
- Did you get Claude approval on this PR?: N/A — not requested for this
focused change.
### Additional Information
Example timeout: [ResNet50 INT8 CI
failure](https://github.com/NVIDIA/Model-Optimizer/actions/runs/36560176214/job/109381161169).
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added a configurable TensorRT builder optimization level for engine
builds, with a default of 4 and support for values from 0 to 5.
* Documented that level 0 can speed up builds, while lower optimization
levels may reduce inference performance.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
|
||
|
|
9c1cf80f1b |
test(megatron_bridge): cover context parallelism in the VLM QAD test (#2592)
### What does this PR do? Type of change: new tests Runs the VLM case of `test_qad` under context parallelism, so QAD on a Qwen3-VL model is covered on the path that until now could not run at all. `Qwen3VLMultimodalRotaryEmbedding` CP-shards its own embedding, so the batch has to hand it full-length `position_ids`. Megatron-Bridge's `get_batch` was sharding them too, leaving the rotary embedding at `seq / cp**2` against hidden states at `seq / cp`. The fix is upstream in [NVIDIA-NeMo/Megatron-Bridge#6243](https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/6243); this PR is the coverage that would have caught it. The VLM case moves from tensor to context parallelism and the PTQ step is sized to the same TP, so QAD still loads a matching checkpoint. The LLM case is unchanged (`tp_size=num_gpus, cp_size=1`), and `test_distill_vlm` still covers TP for a VLM, so nothing loses coverage. ### Usage ```bash # Unchanged: --cp_size is already a distill.py flag. On a container carrying Megatron-Bridge#6243 # it now works for VLMs, where it previously died in the rotary embedding. python examples/megatron_bridge/distill.py --cp_size 2 --tp_size 1 ... ``` ### Testing On 2x RTX 6000 Ada, in `nemo:26.08` with Megatron-Bridge#6243 on `PYTHONPATH`: - `test_qad[qwen3_5_moe_vl]` at `--tp_size 1 --cp_size 2` — FP8 PTQ, QAD across 2 CP ranks, export; quantizers survive and the vision tower is byte-identical. **1 passed (183 s).** Without the upstream fix the same run dies with `AttributeError: 'NoneType' object has no attribute 'ndim'` in `rope.py:175`. - `test_qad[qwen3]`, the unchanged LLM path — **1 passed (194 s).** - Gate check: on today's `nemo:26.08` (no #6243) the probe resolves `False` and the VLM case stays at `cp_size=1`, byte-identical to current CI; with #6243 it resolves `True` and runs at `cp_size=num_gpus`. On a 1-GPU runner it degenerates to today's config either way. - `pre-commit run --files ...` clean (ruff check, ruff format, mypy, bandit, markdownlint). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — existing tests extended rather than new ones added. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — test coverage and one doc line; no feature, break, deprecation, or fix for a released bug. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information - Depends on [NVIDIA-NeMo/Megatron-Bridge#6243](https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/6243). Safe to merge before it lands: the gate keeps the VLM case at `cp_size=1` until a container ships the fix, at which point the coverage switches on by itself. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the Qwen3.6 QAD instructions to keep tensor and pipeline parallelism set to 1, while allowing context parallelism to increase for longer sequences with the `nemo:26.10` container. * **Tests** * QAD validation now selects parallelism settings based on whether the Megatron-Bridge context-parallel fix is available, and reports when multi-GPU VLM coverage is reduced. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
333ace1bc9 |
Fix multi-turn synthetic generation and mark incomplete outputs (#2571)
### What does this PR do?
Type of change: Bug fix
Fixes synthetic conversation generation that assumed alternating user
and assistant messages. That assumption skipped user turns in prompt
skeletons and mishandled
leading system messages.
• Preserve conversation history. Regenerate every user turn while
retaining system messages and generated reasoning for subsequent
requests.
• Expose generation controls. Support model-specific request parameters,
configurable timeouts, and server-managed response budgets.
• Handle failures explicitly. Reject empty final answers and unsupported
tool calls. Failed conversations remain retryable without duplicating
saved output.
• Identify incomplete outputs. Mark length- and repetition-stopped
conversations as truncated, preserve stop metadata, and stop generating
follow-up turns.
### Usage
Run from the repository root against a compatible Qwen server with
reasoning parsing enabled:
python examples/speculative_decoding/scripts/server_generate.py \
--data_path input_conversations/train.jsonl \
--output_path synthetic/train.jsonl \
--url http://localhost:8000/v1 \
--model model \
--max_tokens 0 \
--request_timeout 3600 \
--extra_body
'{"chat_template_kwargs":{"enable_thinking":true,"preserve_thinking":true}}'
The model name must match the server’s configured name. Filter truncated
conversations before training.
### Testing
Focused regression tests: 15 passed.
The tests execute the command-line entry point using the real OpenAI
client library with mocked HTTP transport.
Coverage includes multi-turn generation, system prompts, reasoning
preservation, request parameters, failure recovery, resume
deduplication, truncation, and invalid
responses.
python -m pytest \
--confcutdir=tests/examples/speculative_decoding \
tests/examples/speculative_decoding/test_server_generate.py -q
The isolated test configuration avoids an unrelated parent configuration
import failure. All applicable pre-commit checks passed for the
generator, tests, and
documentation.
### Before your PR is "Ready for review"
• Is this change backward compatible?: ✅ Existing valid inputs,
defaults, conversation output structure, and resume behavior remain
supported. Invalid inputs and
failed requests now raise errors instead of being silently accepted.
• Copied code or new PIP dependencies?: N/A. No new third-party code or
dependencies were added.
• Did you write any new necessary tests?: ✅ Added focused command-line
regression tests.
• Did you update Changelog?: N/A. These are example-script correctness
fixes, not critical released library fixes.
• Did you get Claude approval on this PR?: ❌ Not yet obtained.
### Additional Information
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Data generation supports `conversations` and `messages` inputs,
preserves reasoning content, and accepts additional chat settings and
configurable request timeouts.
* Failed conversations are recorded separately, with options to retry
failures or exit when errors occur. Resume behavior distinguishes
retryable failures from rejected inputs.
* Outputs identify conversations truncated by length or repetition
limits.
* **Documentation**
* Updated data preparation guides with generation setup, input formats,
failure handling, resuming, and training guidance.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
|
||
|
|
ad9ea97a4b |
Fix vLLM compilation guard for models without marker (#2518)
### What does this PR do?
Type of change: Bug fix
Makes the vLLM `disable_compilation` context manager support inner model
implementations that do not predefine a `do_not_compile` attribute,
including GLM-5.3. The context manager now installs the marker
temporarily and removes it afterward, while preserving and restoring
existing marker values for other vLLM models.
Adds regression coverage for both supported wrapper layouts:
`model.model` and `model.language_model.model`.
### Usage
```python
with disable_compilation(model):
mtq.quantize(model, quant_cfg, forward_loop=calibrate_loop)
```
No caller changes are required.
### Testing
- Ran `tests/gpu_vllm/torch/quantization/test_vllm_dynamic_modules.py`:
24 passed with vLLM 0.28.
- Ran pre-commit on both changed files: all applicable hooks passed.
- Installed this branch into `vllm/vllm-openai:glm53-flash` on OCI-JHB
and served the GLM-5.3-Flash BF16 checkpoint with
`QUANT_CFG=NVFP4_DEFAULT_CFG`, TP=4, eager mode, and BF16 KV cache.
- GLM passed the previous `do_not_compile` failure point, inserted 1,700
quantizers, enabled 456 weight quantizers, reached a healthy API server,
and returned a relevant manual prompt response.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — integration compatibility fix; no user-facing API change.
- Did you get Claude approval on this PR?: N/A
### Additional Information
Validated against GLM-5.3-Flash using ModelOpt commit
`869b64fcee0b20be323663449b00e8c52940a289`.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Bug Fixes**
- Compilation settings are now handled across supported nested model
configurations and restored after calibration, including when errors
occur.
- Calibration inputs correctly exclude padding when an attention mask is
provided and reject empty sequences.
- vLLM warmup reserves the required cache space for supported tail-cache
configurations.
- Serving startup supports an alternate vLLM launcher import path when
the OpenAI entrypoint is unavailable.
- **Compatibility**
- The vLLM serving example now defaults to vLLM 0.30.0 and documents
tested support for Nemotron 3 Nano hybrid attention/Mamba serving on
vLLM 0.26.0 and 0.30.0.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
|
||
|
|
1d392999b4 |
[OMNIML-5570] 2/2 Compose GEMM and KV-cache AutoQuant workflows (#2273)
### What does this PR do? Type of change: new feature. Follow-up to merged #2272. Adds composition of existing GEMM quantization with KV-cache AutoQuantize: - fixed FP8 GEMM PTQ followed by mixed-KV AutoQuantize; - gradient-based NVFP4/FP8 GEMM AutoQuantize followed by independent mixed-KV AutoQuantize; - an optional `kv_auto_quantize` recipe stage with independent method, constraints, candidates, and checkpoint path; - ordered `hf_ptq.py` orchestration that keeps selected weight/activation QDQ active while its calibration state remains frozen during KV candidate calibration; - fail-closed validation when a preceding stage leaves actual K/V quantizers enabled; and - unified export of a uniform-weight or mixed-weight checkpoint together with the selected per-layer KV map. The KV search still uses the public `mtq.auto_quantize(..., constraints={"cost_model": "kv_cache", ...})` API from #2272. On a converted model, the API preserves existing non-KV quantizers and requires K/V to be disabled before search. Fresh-model behavior is unchanged and starts from a deny-all quantizer baseline. #### Why a follow-up field instead of a generic stage list? This PR deliberately supports the two composition forms required by `hf_ptq.py` without replacing the stable recipe schema. Existing recipes already express a fixed `quantize` baseline plus one primary `auto_quantize` search. A generic ordered `stages` list would require a broader recipe/API migration, indexed checkpoint semantics, and compatibility rules for arbitrary stage sequences. There is not yet a demonstrated third search stage that justifies that surface-area change. The two searches are not combined inside `mtq.auto_quantize`: each invocation owns one search domain, constraint model, scoring method, and resumable checkpoint. Their ordering and independent checkpoint paths are orchestration concerns, while candidate calibration, scoring, selection, and state application remain in the shared public API. A general stage pipeline can be considered separately if more than this one optional KV follow-up is needed. Both solvers and scoring protocols are unchanged. The KV checkpoint compatibility signature additionally fingerprints the preceding quantizer configuration and calibrated state. Unsupported uniform-weight plus mixed-KV exports record `kv_cache_deployment_supported: false` in both ModelOpt and converted HF metadata. ### Usage Fixed FP8 GEMM PTQ followed by KV AutoQuantize: ```bash python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path Qwen/Qwen3-8B \ --recipe general/auto_quantize/fp8_ptq_then_kv_fp8_nvfp4_cast_kl_div_at_5p4bits \ --kv_auto_quantize_checkpoint /path/to/kv_autoquant.pth \ --export_path /path/to/qwen3-8b-fp8-and-mixed-kv ``` Weight AutoQuantize followed by KV AutoQuantize: ```bash python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path Qwen/Qwen3-8B \ --recipe general/auto_quantize/nvfp4_fp8_gradient_then_kv_fp8_nvfp4_cast_kl_div_at_5p4bits \ --auto_quantize_checkpoint /path/to/weight_autoquant.pth \ --kv_auto_quantize_checkpoint /path/to/kv_autoquant.pth \ --export_path /path/to/qwen3-8b-autoquant-and-mixed-kv ``` KV checkpoint resume requires identical preceding non-K/V quantizer configuration and calibrated state. If rerunning the preceding stage changes that state, use a new KV checkpoint path to recompute sensitivities; configuration identity alone is insufficient to reuse the scores safely. ### Testing - Latest changed-area validation: 126 tests passed across `hf_ptq.py` orchestration, KV checkpoint compatibility, export metadata, and HF configuration conversion. - A broader local run had 604 passes, one skip, and six failures: two socket-binding failures under the sandbox and four local Transformers API incompatibilities. This is not a full-suite pass. - The fixed-PTQ→KV recipe executes end to end on a tiny offline Qwen fixture. - Public API coverage verifies that composed KV search preserves preceding weight quantization and rejects enabled K/V state. - Changed-file pre-commit hooks passed; the isolated recipe validator also passed after dependency bootstrap. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ (0.48.0 composition feature and KV checkpoint flag deprecation) - Did you get Claude approval on this PR?: ❌ ### Additional Information - This follow-up targets `main`, which contains merged #2272. - `--auto_quantize_checkpoint` and `--kv_auto_quantize_checkpoint` are intentionally separate because KV sensitivities depend on the preceding GEMM state. - Uniform-weight plus mixed-KV exports are for artifact inspection until the runtime's uniform-weight ModelOpt configuration consumes `kv_cache_quantized_layers`. Export emits an actionable warning and records `kv_cache_deployment_supported: false`; this marker does not itself add runtime support. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added staged post-training quantization workflows for weights and KV caches, including dedicated KV-cache checkpoints. - Added FP8/NVFP4 recipes with configurable bit constraints and scoring. - KV-cache quantization now supports pre-quantized models. - **Bug Fixes** - Mixed weight and KV-cache quantization now exports with a warning instead of failing. - Improved validation and checkpoint compatibility for staged configurations. - Added safeguards for configurations without enabled weight quantizers. - **Documentation** - Clarified staged KV-cache workflows, checkpoint options, configuration behavior, and unsupported deployment combinations. - Documented deprecated legacy quantization options and their replacement behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
0058a15537 |
[2/2] Track every Megatron-Bridge script with MLflow (#2514)
### What does this PR do? Type of change: new feature **[2/2] of a split. Based on #2544 — merge that first; this PR's diff is only the Megatron-Bridge half.** #2477 added MLflow tracking to `examples/megatron_bridge/quantize.py`. It was one of five scripts in that directory that write a checkpoint; the other four recorded nothing, so the provenance chain stopped at the PTQ checkpoint and a deployed model could not be traced back to the run that produced it. All five now take the same `--mlflow` / `--mlflow_experiment` / `--mlflow_run_name` flags, and **each declares what it records as a `Tool` beside its own flags** — the shared `mlflow_utils.py` knows none of them: | Script | Records | | --- | --- | | `prune_minitron.py` | command, arguments, log, `prune_score` metric, pointer | | `quantize.py` (#2477, moved onto the shared `Tool` in #2544) | + resolved recipe, quantizer summary | | `distill.py` | + Megatron-Bridge's per-iteration metrics and resolved config | | `export_quantized_megatron_to_hf.py` | command, arguments, log, pointer | | `export_distilled_megatron_to_hf.py` | same, one pointer per exported checkpoint | Each writes `.experiment.json` into the checkpoint it produced, and each tags what it consumed, so `prune → quantize → distill → export` is walkable both from disk and by tag query on the server. **`distill.py` opens the run and Megatron-Bridge joins it.** Its `LoggerConfig` records per-iteration metrics and the full resolved config — which a wrapper around `main()` cannot see — but nothing of `distill.py`'s own arguments and no invocation. Megatron-Bridge takes `mlflow.active_run()` when one exists, applies the tags and logs into it, so `distill_run()` opens the run on the rank Megatron-Bridge looks at (the **last** one) and the two share it. Its early exit is handled explicitly: `train()` leaves through `sys.exit(0)` on `--exit_interval`, which a blanket handler would record as `FAILED`. **The library pieces that exist for that shared run land here with their first caller**, rather than in [1/2] where they would have none: `split_tracking_credentials`, so a URI handed to something which *records* it carries no credential; `log_active_run_experiment_json`, for pointing a checkpoint at a run this process did not open; and `MlflowRunLogger._reattach`, because a co-owner can end the run first — Megatron-Bridge does, as `KILLED`, when SIGTERM arrives mid-training. Two of Megatron-Bridge's defaults are deliberately not inherited: **checkpoint artifact upload stays off** unless `--mlflow_log_checkpoints` (it pushes the whole checkpoint over HTTP after every save), and **an untracked run passes no `mlflow_*` fields at all**, since they landed in Megatron-Bridge 0.6 and sending them unconditionally would break an untracked run on an older one. ### Usage ```bash # Any of the five, same flags: torchrun --nproc_per_node 8 prune_minitron.py ... --mlflow https://<server>/ torchrun --nproc_per_node 8 quantize.py ... --mlflow https://<server>/ torchrun --nproc_per_node 8 distill.py ... --mlflow https://<server>/ torchrun --nproc_per_node 8 export_quantized_megatron_to_hf.py ... --mlflow https://<server>/ # Each checkpoint names the run that wrote it: cat /output/qad/checkpoints/.experiment.json ``` Experiments default to `$USER/megatron_bridge_{prune,quantize,distill,export,distill_export}/<model basename>-<variant>`. ### Testing - Real runs on a toy Qwen3 in one MLflow experiment covering all five Megatron-Bridge scripts and `hf_ptq` — prune, quantize, QAD distillation, quantized export, BF16 distillation, distilled export, HF PTQ — each closing `FINISHED` with the invocation, its arguments as params, its log, and a matching `.experiment.json` on disk. The chain tags line up: each stage's `source_checkpoint_path` is the previous stage's `checkpoint_path`. - `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08`, the only lane that runs it: **76 passed**. Plus the three suites from #2544: **195 pass**. - `pre-commit run --files <changed>`: all hooks pass. - Each fix from the review rounds has a test that fails with the fix reverted: the resumed run, the foreign active run, the percent-decoded credential, the credential that cannot be moved, the rank-dependent `LoggerConfig`, the exit-callback guard, and the `iter_*` join. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: several rounds; re-requested on this head. ### Additional Information Split from a single ~1150-line PR at review's request; #2544 carries the library consolidation this builds on, and this branch is based on it. Earlier review threads here show as outdated after the rebases — they are all resolved and their fixes are in this branch. One known gap, stated in the README rather than implied: `distill.py --hf_export_path` writes a second HuggingFace checkpoint from rank 0, which is not the rank that owns the run, so it carries no pointer yet. For the same reason the uploaded `logs/distill.log` holds the last rank's output — `print_rank_0` keeps the script's own lines on rank 0 — which the README now says outright; carrying rank 0's log into a run owned by another rank needs cross-rank upload and is a follow-up. Two defects found on shared-run paths during review, both verified against the installed Megatron-Bridge 0.6 rather than its docs. Megatron-Bridge ends the run it shares with `distill.py` as `KILLED` from its SIGTERM handler (`train.py:1413`) and then leaves through `sys.exit()` (`train.py:805`), i.e. before `distill_run`'s `finally` — and MLflow's fluent calls resolve their target by *opening* a run when none is active, so a preempted distillation's log and metrics went to a second, empty run and its `KILLED` status was overwritten. Separately, an unreachable server disabled our logger but `logger_kwargs` still handed Megatron-Bridge the same URI, and `state.py` calls `set_experiment` unguarded from inside the training loop — so a best-effort `$MLFLOW_TRACKING_URI` aborted the training instead of degrading to untracked. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4eb86524f0 |
[1/2] One MLflow tracking core behind a Tool record (#2544)
### What does this PR do? Type of change: refactor (no functional change) **[1/2] of a split. Merge this first; #2514 is [2/2] and is based on this branch.** Three example scripts had each reimplemented the same MLflow wiring: the flags, the `$USER/<tool>/<model>-<variant>` experiment convention, the params/tags/artifacts a run uploads, and the open/close dance with its status. The copies had already drifted — only `hf_ptq` wrote a provenance pointer, only `vllm_serve` republished the resolved URI — and every new tracked script meant another copy. What a script records is now one declarative `Tool` record, **declared in the script itself, beside the flags it reads**: ```python # examples/megatron_bridge/quantize.py QUANTIZE = Tool( name="megatron_bridge_quantize", tracks="Track this run on an MLflow server, uploading the command, the resolved recipe, ...", variant_help="recipe name, or --quant_cfg if no --recipe", variant=lambda args: Path(args.recipe).stem if args.recipe else (args.quant_cfg or "none"), model=lambda args: args.hf_model_name_or_path, checkpoint=lambda args: args.export_megatron_path, texts=lambda args: resolved_recipe_texts(args.recipe), outputs=lambda args: {"summary/quant_summary.txt": Path(args.export_megatron_path) / ".quant_summary.txt"}, ) ``` `tracked_run` takes that record and runs the whole thing, so a script adds tracking in three lines: `add_mlflow_args(parser, TOOL)`, `resolve_mlflow_args(args, parser, TOOL)`, and `with mlflow_run(args, TOOL):`. The shared module knows no script's flags. `examples/hf_ptq`, `examples/vllm_serve` and `examples/megatron_bridge/quantize.py` move onto it. Three helpers fall away as redundant (`track_run`, `checkpoint_run_tags`, and `hf_ptq`'s two flag pass-throughs). ### Usage No user-facing change. The flags, their spellings and the experiment naming are exactly as before; a script author now writes a `Tool` instead of four functions. ### Testing - `tests/unit/torch/utils/test_mlflow.py`, `tests/examples/hf_ptq/test_hf_ptq_args.py`, `tests/examples/vllm_serve/test_vllm_mlflow_utils.py` — **179 pass**. - `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08` (the only lane that runs it), which drives `quantize.py` for real: **34 passed**, locally and in this PR's `megatron` lane. - `pre-commit run --files <changed>`: all hooks pass. - The four suites shared four copies of a stand-in for the `mlflow` module, which had drifted — one recorded artifacts as a list, another as a dict, a third made `log_artifact` a no-op, so a test asserting on an upload asserted nothing. They now share one `tests/_test_utils/mlflow.py`, which also emulates the fluent API's habit of opening a run when none is active. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `track_run` and `checkpoint_run_tags` are removed, but neither shipped in a release (0.47.0's `__all__` is `MlflowRunLogger`, `command_text`, `current_user`, `default_experiment_name`, `validate_tracking_uri`, all unchanged here). Three deliberate behaviour changes, each in shared code and each tested: - The `source_checkpoint_path` tag resolves to an absolute path where it recorded the raw argument, which a chain of runs needs to join on the pair. `run_tags` is shared, so this applies to every script that writes the tag — `hf_ptq` **and** `megatron_bridge/quantize.py`, for a local `--hf_model_name_or_path`. A source that names no directory, such as a Hub `org/name` id, is still recorded as given. - `MlflowRunLogger.track()` — which *did* ship in 0.47.0 — records a block ending in `SystemExit(0)` as `FINISHED` where it recorded `FAILED`, since a script that ends by calling `sys.exit()` rather than returning has still finished. - `.experiment.json`'s `tracking_uri` and the `run_url` built from it drop a trailing `/` from the tracking URI, so the link is `https://host/#/...` rather than `https://host//#/...`. Only reachable by constructing `MlflowRunLogger` directly; every CLI path already stripped the slash in `resolve_tracking_uri`. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — no user-visible change; the entry is in [2/2]. - Did you get Claude approval on this PR?: several rounds; re-requested on this head. ### Additional Information Split out of #2514. This half is the enabling refactor with no behaviour change; #2514 is the feature it unlocks and is based on this branch. At ~605 changed lines of core logic it is over the ~500 guideline; the owner accepted a two-PR split rather than three, and everything #2514 alone consumes — `split_tracking_credentials`, `log_active_run_experiment_json`, `MlflowRunLogger._reattach` — lands there rather than here. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d16dad1c20 |
docs(llm_distill): document KDTrainer-based example flow (#2524)
## Summary - Update `examples/llm_distill/README.md` to describe the current `main.py` flow, which uses `KDTrainer` (from `modelopt.torch.distill.plugins.huggingface`) instead of `mtd.convert()` / `DistillationModel` wrapping. - Document that `KDTrainer` only supports logit-level distillation today, and that hidden-state/intermediate-layer KD still requires `mtd.convert()` + `DistillationModel` until `KDTrainer` gains that support. ## Test plan - [x] Reviewed rendered README diff for accuracy against `modelopt/torch/distill/plugins/huggingface.py` and `examples/llm_distill/main.py` - [x] `pre-commit` hooks (markdownlint-cli2, etc.) passed on commit 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the Hugging Face getting-started example to use `KDTrainer` with the standard training and model-saving workflow, including an example of combining it with `SFTTrainer`. * Clarified that `KDTrainer` supports logit-level distillation; hidden-state distillation uses a separate approach. * Explained that KD loss and evaluation cross-entropy are reported separately, with weighted CE/KD loss combination unsupported. * Added guidance on distributed training options, including the FSDP2 requirement and alternatives to default DataParallel. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
f2f0d6958e |
Add MLflow tracking flags to megatron_bridge quantize.py (#2477)
### What does this PR do?
Type of change: new feature
`examples/megatron_bridge/quantize.py` gains the MLflow tracking flags
`examples/hf_ptq/hf_ptq.py` already has: `--mlflow <tracking-uri>`
(MLflow's own `$MLFLOW_TRACKING_URI` is honoured too),
`--mlflow_experiment` and `--mlflow_run_name`. Only the master rank
opens a run, so a `torchrun` launch produces one run carrying the
invocation, every command-line argument as a searchable param, the
resolved `--recipe` (with `$import`s expanded), that rank's log and the
quantizer summary. Once `bridge.save_megatron_model` returns,
`.experiment.json` is written into `--export_megatron_path`, so a
Megatron checkpoint found on disk names the run that produced it; a run
that fails is still recorded as `FAILED` with its traceback.
Rather than copy the wiring a third time, the part `hf_ptq` and
`vllm_serve` had each duplicated moves into
`modelopt.torch.utils.mlflow`:
- `add_mlflow_args(parser, tool, tracks=, variant_help=)` — the three
flags, registered under both the `--mlflow_x` and `--mlflow-x` spellings
(vLLM's `FlexibleArgumentParser` only matches the dashed one).
- `resolve_tracking_uri(uri, parser)` → `(uri, required)` — the flag
overrides the environment and is fatal when the URI is unusable; a URI
inferred from `$MLFLOW_TRACKING_URI` warns and continues untracked,
since that variable is commonly exported for unrelated tooling.
- `resolve_mlflow_args(args, parser, tool, model, variant)` — the same,
settled onto `args`, plus the default experiment name.
- `EXPERIMENT_JSON`, `MlflowRunLogger.log_experiment_json()` and
`drop_experiment_json()` — the checkpoint→run provenance pointer,
previously private to `hf_ptq`.
Both existing callers now delegate to those, keeping their own help
wording and variant naming, so the three scripts share one convention
instead of three copies (`example_utils.py` and `vllm_mlflow_utils.py`
each lose ~60 lines). Their flags and defaults are unchanged; the only
user-visible difference is that `hf_ptq`'s ignored-URI warning gains the
`$` the vLLM one already had (`Ignoring $MLFLOW_TRACKING_URI, continuing
untracked`), so one shared message serves both.
One behaviour change reaches `hf_ptq` through the shared helper, and it
is a fix: when tracking was inferred from `$MLFLOW_TRACKING_URI` and the
run never opened (unreachable server, or `mlflow` not installed), it
used to leave the previous run's `.experiment.json` beside a freshly
exported checkpoint. `log_experiment_json` now drops the pointer when it
has no run to record, so after a completed export the file is this run's
or absent.
The new example-side code lives in
`examples/megatron_bridge/mlflow_utils.py`, which deliberately imports
no Megatron, so the whole flag-to-artifact path is testable without the
Megatron container (the same split
`examples/vllm_serve/vllm_mlflow_utils.py` uses).
### Usage
```bash
torchrun --nproc_per_node 2 quantize.py \
--hf_model_name_or_path Qwen/Qwen3-8B \
--recipe general/ptq/nvfp4_default-kv_fp8 \
--tp_size 2 \
--export_megatron_path /tmp/Qwen3-8B-NVFP4-megatron \
--mlflow https://<your-mlflow-server>/
# The checkpoint then names the run that produced it:
cat /tmp/Qwen3-8B-NVFP4-megatron/.experiment.json
```
The experiment defaults to `$USER/megatron_bridge_quantize/<model
basename>-<recipe name, or --quant_cfg>`.
### Testing
- `tests/examples/megatron_bridge/test_mlflow_utils.py` — 20 new tests
covering the flags (both spellings, env-vs-flag precedence, the
fatal/best-effort split), the params/tags/artifacts a run records, rank
gating, and the `.experiment.json` lifecycle. The last one guards the
seam with `quantize.py` as text, since that script needs Megatron to
import.
- `tests/unit/torch/utils/test_mlflow.py` — 13 new tests for the
extracted library API; suite at **75 passed**.
- Full `tests/examples/megatron_bridge` suite in
`nvcr.io/nvidia/nemo:26.08` on an RTX 6000 Ada: **37 passed (26m)**,
including the three `test_quantize_export` cases that drive the real
`quantize.py`, plus QAD, distill and prune.
- Regression proof for the refactor:
`tests/examples/hf_ptq/test_hf_ptq_args.py` **47 passed** and
`tests/examples/vllm_serve/test_vllm_mlflow_utils.py` **32 passed**,
unchanged apart from one renamed constant reference.
- Both new guards were shown to fire: mutating the `checkpoint_exported`
gate and removing `with mlflow_run(args):` each failed exactly one test.
- End-to-end tracked run in `nvcr.io/nvidia/nemo:26.08` (tiny Qwen3-MoE,
`general/ptq/fp8_default-kv_fp8`, 1 GPU) against an internal MLflow
server: run `47d4ccd7cd9e48269e7248868347ccd0` under experiment
`$USER/megatron_bridge_quantize/mbridge-ptq-validation` closed
`FINISHED` carrying `command.txt`, `version.txt`, `experiment.json`,
`recipe/resolved_recipe.yaml`, `logs/quantize.log` and
`summary/quant_summary.txt`; all 19 CLI arguments plus `world_size`
logged as params with no `mlflow_*` leakage, the
`model`/`checkpoint_path`/`source_checkpoint_path` tags set, and
`.experiment.json` written into the Megatron checkpoint beside
`iter_0000000/`.
- `pre-commit run --files <changed>`: all hooks pass (ruff, mypy,
bandit, markdownlint).
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under *Megatron Framework (M-LM / M-Bridge)*.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
`mlflow` stays an optional dependency, imported only once tracking is
enabled, so an untracked run behaves exactly as before.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added optional MLflow tracking for Megatron-Bridge quantization runs.
- Configure tracking with `--mlflow` or `MLFLOW_TRACKING_URI`, with
customizable experiment and run names.
- Records searchable parameters, resolved recipes, quantization
summaries, logs, and checkpoint provenance.
- Captures successful and failed runs and cleans up stale checkpoint
metadata when appropriate.
- **Documentation**
- Added setup instructions and usage examples covering artifacts,
naming, checkpoint metadata, validation, and authentication.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
87f7d1432f |
fix(speculative): hold the DFlash draft's fp32 master weights in the optimizer (#2483)
### What does this PR do? Type of change: Bug fix **Follow-up to #2342**, which split this out on review (commit `c67784d9`), and a rethink of how the flag is implemented. `dflash_fp32_master_weights` exists because the DFlash draft is cast to the frozen bf16 target's dtype, so AdamW allocates its moments in bf16 — and bf16 is too coarse to hold them. At `beta2=0.999` a single step changes `v` by at most **0.100%**, while the smallest change bf16 can represent near `v` is **0.164% mean / 0.388% max** (measured): every decrease rounds away, `v` only grows, and the effective step size decays on its own from step 1. #2342 fixed that by **promoting the draft model to fp32**. Everything else followed from giving the model a dtype the rest of it does not have — a bf16 autocast at every entry point, two transformers loader hints so `from_pretrained(dtype="auto")` would not round the draft away, a post-condition check because those hints fail silently, and a doubled DDP gradient all-reduce. **This PR puts the fp32 in the optimizer instead**, where Megatron-LM, DeepSpeed and apex put it. `MasterWeightAdamW` holds an fp32 master copy of each non-fp32 parameter plus fp32 moments in `self.state[p]`, steps on the master, and copies back at the parameter's dtype. The model is never anything but the base dtype, so every one of those follow-on pieces is deleted, gradients stay bf16, and the exported drafter is unchanged. What the placement costs is that wiring the optimizer becomes the training loop's job: `EagleTrainerWithAccLog.create_optimizer` builds it, and `VerifyMasterWeightsCallback` raises at the end of step 1 if the moments are not fp32. **The default flips to `True`** — the flag now changes optimizer memory and optimizer arithmetic and nothing else. Flipping it on the model-promoted implementation turns **25 of 259** unit tests red; flipping it here is **259 passed**. Set it to `False` to reclaim the memory, about 12 bytes per draft parameter instead of 4. <details> <summary>Three drive-by fixes, independent of the above</summary> - `_place_draft` is folded back into `modify()` — it fused the draft's dtype, its device and an eager rotary buffer behind one meta guard. - The module docstring's claim that `DFlashModule` has an `_apply` meta-buffer fix is removed (`grep "def _apply"` matches nothing, and never did). - #2342's field description no longer lists `evaluation` as a broken path — `forward` short-circuits to the base model when `not self.training`, so the draft never runs there. </details> ### Usage No API change. `dflash_fp32_master_weights` now means the *optimizer* holds fp32 master weights rather than the draft model being fp32. ### Testing **1 · The refactor is arithmetically a no-op.** Both implementations run AdamW on an fp32 tensor, so given the same starting values and the same gradients the trajectories are identical — 1000 steps, `weight_decay=0.01`: ``` old fp32 parameter vs new fp32 master : bitwise equal = True (max |diff| 0.0e+00) exp_avg / exp_avg_sq : bitwise equal = True optimizer state dtypes : ['torch.float32'] model parameter dtype : torch.bfloat16 ``` Initial values have to be matched at bf16 first, or the bf16 arm's one-time rounding of the draw shows up as a 2e-4 "difference" that is not arithmetic. With that controlled, the two implementations differ only in their *inputs*: gradient precision (fp32 vs bf16 — torch 2.10 requires `grad.dtype == param.dtype`) and that one-time rounding. **2 · End to end on GPU: the effect survives the refactor.** Qwen3-1.7B base, real corpus, one GPU per arm, three arms — pure bf16 (flag off), the #2342 implementation, and this one — on two algorithms trained independently, sharing seed, data order and initialisation within an algorithm. <img width="2925" height="960" alt="image" src="https://github.com/user-attachments/assets/40d37059-ab8a-4924-b049-85f76b70b156" /> The two fp32 arms sit on top of each other for the whole run while bf16 stays above both, and the old-vs-new gap is 10–23× smaller than the fp32-vs-bf16 effect it has to be compared against. **Acceptance length says the same thing, and settles what the loss could not.** All six drafters at the end of those curves were exported and served under vLLM against the same base, and measured on MT-Bench (80 prompts, 8 categories, greedy, one request at a time, `num_speculative_tokens` = trained `block_size` − 1, every knob but the drafter held fixed): | | pure bf16 | fp32 in model (#2342) | fp32 in optimizer (this PR) | new − old | fp32 − bf16 | |---|---|---|---|---|---| | `dflash` | 1.3068 | 1.3708 | **1.3666** | −0.0042 `t=−0.90` | +0.0619 `t=+13.8` | | `lilicorr` | 1.2536 | 1.2814 | **1.2882** | +0.0068 `t=+1.42` | +0.0312 `t=+9.2` | Paired by prompt, n=80. On both algorithms the new-vs-old 95% CI straddles zero (`dflash` [−0.0134, +0.0051], `lilicorr` [−0.0028, +0.0164]) while fp32-vs-bf16 does not come close to it, and the sign of new-vs-old **flips between the two algorithms** — what a rounding difference looks like, not a bias. This is also the comparison the training loss could not give: all three arms are **exported and served in bf16**, so the old implementation's fp32 draft weights are rounded at export exactly as they would be for deployment, and the "its loss was computed on a more precise forward" caveat below does not apply. `lilicorr` needs [vllm-project/vllm#57934](https://github.com/vllm-project/vllm/pull/57934), applied as an overlay so that both algorithms are measured on one engine build. The right panel is the mechanism, and the one signal that depends on neither the seed nor the choice of loss statistic: Adam's updates to the draft's RMSNorm gains are smaller than the bf16 ULP at 1.0 (0.0078), so in the bf16 arm every one of them rounds away and the gains never move — not one of `dflash`'s 14 in 30000 steps, and two of `lilicorr`'s 20 by 3e-06. Both fp32 arms move all of them, by the same amount. Two results behind the figure rather than in it. **fp32-vs-bf16 grows with the horizon** while old-vs-new does not — on `dflash` −0.129 at 1500 steps → −0.262 at 15000 → −0.341 at 30000, and on `lilicorr` −0.191 → −0.220 → −0.285, against an old-vs-new difference that stays near 0.02 at every horizon and changes sign between them (−0.026 → +0.028 on `lilicorr`). That is what a compounding bias and a rounding difference respectively should look like, and it is the reason the longer runs were worth doing. And **across seeds**, the paired old-vs-new difference at 1500 steps is +0.0003 (n=6) on `dflash` and +0.0643 (n=10) on `lilicorr`, both with a 95% CI straddling zero. <details> <summary>Limits of the above, stated rather than smoothed over</summary> At 5 seeds the `lilicorr` paired difference read +0.1610 ± 0.0557 (t=+2.89, 4/5 seeds in the same direction) — nominally significant, suggesting the new implementation was genuinely worse there. Four further `lilicorr` seeds were run against that pre-declared question; two came back strongly negative and the estimate settled at +0.0643 (95% CI [−0.086, +0.214]). The earlier reading was small-sample noise. At 1500 steps on `lilicorr` that CI is *not* narrower than the fp32-vs-bf16 effect it is being compared against, so the 1500-step sweep alone cannot certify equivalence there — `lilicorr` is still at loss 9.3 and deep in its early transient, and it is the long runs that resolve it. On `dflash` the 1500-step CI (±0.031) is already 4× tighter than the effect (−0.129). One asymmetry the loss comparison cannot separate: the old implementation held the draft weights in fp32 *at forward time*, so its training loss was computed on a more precise forward, while both implementations export bf16. Any residual advantage it appears to have is therefore an upper bound. </details> **3 · Unit tests.** `tests/unit/torch/speculative/` — **259 passed** on **transformers 5.0.0** and **5.3.0**, both ends of the supported `>=5.0,<5.13` (CPU, torch 2.10). `TestDFlashFp32MasterWeights` is rewritten for the new mechanism; the two that would have caught the traps in this design are `test_resume_does_not_round_the_master_back_down` (`Optimizer.load_state_dict` casts float state to its parameter's dtype, so a naive subclass rounds the master and both moments to bf16 on *every* resume, silently, with the loss still falling) and `test_the_callback_refuses_a_loop_that_forgot_the_optimizer`. The rest cover the draft's dtype with the flag either way, that no forward path needs an autocast any more, that plain AdamW really does leave the moments in bf16, and that an fp32 model allocates no redundant master. A sharded FSDP2 `DTensor` keeps an fp32 master and fp32 moments through a step. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ for artifacts, with one intentional default change. The draft's stored dtype goes back to matching the base, as it was before #2342; existing checkpoints load unchanged and the exported drafter is unaffected. The flag now defaults to **`True`** — the measurements above are the reason, and the cost is fp32 master + fp32 moments for the draft only. A training loop that builds its own optimizer instead of using the shipped `create_optimizer` gets plain AdamW and none of this; `VerifyMasterWeightsCallback` makes that fail loudly at step 1 rather than skip the feature quietly. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new dependencies. - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ — draft; will run `/claude review` before marking ready. ### Additional Information **On the 7–14% acceptance-length gain quoted in #2342:** that was measured on the fp32-model arithmetic and is not re-derived here. What is measured above is the like-for-like comparison this PR has to answer — same corpus, same horizon, same serving path, one implementation swapped. **History:** commits 1–3 restore the autocast design as it was split out; commits 4–6 replace it. Happy to squash before review. <details> <summary>Alternatives measured and rejected, so they do not get re-proposed</summary> - **Swapping `p.data` to the master and calling `super().step()`** (reuses all of AdamW, ~20 lines instead of ~50): bit-identical on ordinary parameters over 25 steps, but silently wrong under FSDP2 — assigning `.data` on a `DTensor` parameter updates the wrapper's reported dtype while the local shard keeps the model's, so `p.dtype` reads fp32, `p.data.dtype` reads bf16, and `zeros_like(p)` allocates the moments in bf16 anyway. CPU tests pass either way. - **Narrowing the autocast from `__call__` to `forward`** (while it still existed): turns 10 Domino/DSpark tests red — the variants apply their heads in their own `forward` overrides, outside `DFlashModule.forward`. - **Building the rotary buffer on meta and letting the loader materialise it**: makes RoPE correctness depend on transformers selecting a branch by class-name substring (`"RotaryEmbedding" in module.__class__.__name__`), and the `if not hasattr` guard is then permanently satisfied, so a later `to_empty()` leaves garbage forever — measured `4.56e-41`, i.e. cos=1 / sin=0, no positional encoding at all. - **Building it eagerly in `DFlashModule.__init__`**: lands before the dtype cast, so `Module.to` rounds the RoPE frequencies to bf16 on the default path — measured `0.8659643530845642` → `0.8671875`, loss `3.47230935097` → `3.47114777565`. </details> 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - DFlash now uses FP32 optimizer master weights and Adam moments by default while keeping draft parameters in the base model’s dtype. - Master-weight training preserves optimizer precision when restoring checkpoints. - The feature can be disabled to reduce optimizer memory usage. - Draft models consistently follow the base model’s dtype and device. - **Bug Fixes** - DFlash workflows now support operation without autocast. - Added validation for compatible AdamW-family optimizers and master-weight precision, including resumed training runs. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f1abc75626 |
Carry unplaced checkpoint weights using the loader's accounting, replacing MTP name-matching (#2427)
### What does this PR do?
Type of change: bug fix + new tests
Replaces `hf_ptq`'s name-based MTP detection with the Transformers
loader's own accounting of what it could not place.
#### The problem
`load_mtp_weights` found MTP weights by name — `"mtp" in key` plus
config-derived layer indices — backed by a support matrix of three
storage conventions (GLM-5.1 inlined, GLM-4.7 standalone file,
Qwen3-Next tail shard). Every architecture that spells it differently is
a silent miss, and a miss means a checkpoint exported without a
component its own config still advertises. The failure is quiet on both
sides: Transformers drops unexpected keys, and vLLM's weight loading is
pull-based, so a missing MTP produces no warning at all.
#### The fix
`from_pretrained(..., output_loading_info=True)` reports
`unexpected_keys` — *"keys that are found in the checkpoints, but not
expected in the model's architecture"* — which is exactly the carry-over
set, derived structurally rather than by naming, and already accounting
for on-the-fly key conversion that a set re-derived afterwards would
have to replay. Record those keys at load; carry them at export. Weights
the loader *did* place go through the normal export path unchanged.
#### Two mechanisms, disjoint by construction
Measured against a real `from_pretrained`, an off-index file reports
**nothing**: the loader opens only shards named in
`model.safetensors.index.json`, so it never saw those tensors to call
them unexpected. Those are sidecars, not weights — untouched by
quantization and absent from the export — so they are **copied
verbatim**, which costs no host memory, preserves the bytes and file
layout, and leaves the filename a consumer looks for where it was.
| on-disk layout | mechanism |
|---|---|
| inlined layer past `num_hidden_layers` (GLM-5.1, DeepSeek-V3) |
`extra_state_dict` |
| indexed `mtp.*` tail shard (Qwen3-Next) | `extra_state_dict` |
| standalone off-index file (GLM-4.7) | **copied whole** |
A tensor is never both copied and carried; that would export it twice,
and a test asserts it.
#### Removed as redundant
`load_mtp_weights`, `mtp_layer_prefixes_from_checkpoint`,
`get_inlined_mtp_prefixes`, `_load_tensors_matching`,
`_apply_to_model_state_dict`, `_keys_to_prefixes`, `_add_mtp_exclusions`
and its three call sites, the pre-quantization `enable: False` entries
`hf_ptq` appended to the recipe's `quant_cfg`, and the dead
`_mtp_layer_prefixes` fallback in `_get_num_nextn_predict_layers`.
#### Two deliberate behavioural changes
**MTP now follows the recipe** instead of being force-excluded by the
script — which is what `examples/megatron_bridge` already does; it has
no MTP-specific code at all. Recipes importing
`configs/ptq/units/default_disabled_quantizers` still disable `mtp.*`,
so their behaviour is unchanged; a recipe omitting that unit will now
quantize an MTP the model actually built.
**`quantization_config.ignore` can no longer claim a layer is
unquantized that the export in fact quantized.** That contradiction came
from `_add_mtp_exclusions` firing off a model attribute with no
cross-check against quantizer state.
### Usage
No API change for callers of `export_hf_checkpoint`. Within
`examples/hf_ptq`, model loading now goes through a wrapper that records
the loader's accounting:
```python
model, loading_info = auto_class.from_pretrained(ckpt_path, output_loading_info=True, **kwargs)
record_unplaced_source_keys(model, ckpt_path, loading_info.get("unexpected_keys"))
```
### Testing
`tests/examples/hf_ptq/test_carry_over_layouts.py` — 8 tests driving a
**real** `from_pretrained` against a tiny model, covering each of the
three conventions above plus an auxiliary (non-MTP) tower, two layouts
at once, a checkpoint with nothing stray, and that indexed shards are
never copied. CPU-only: the mechanism is bookkeeping during load, so a
GPU adds nothing; the export side already has GPU coverage in
`tests/gpu/torch/export/test_export_carry_over.py`.
The six `load_mtp_weights` tests are replaced by three on the recording
path, and the `get_model` test doubles now model `output_loading_info`
the way Transformers does.
All passing: 8 layout tests, 82 in the surrounding `examples/hf_ptq`
suite. `ruff` findings at parity with `main` on every changed file.
Files named like a main weight shard are excluded from the off-index set
whatever the index says — a fixture with an empty `weight_map` would
otherwise have made the source weights look like sidecars and copied
them into an export beside the quantized ones.
### The algorithm: which weights get carried, and how
Two disjoint sets of source weights reach the export without passing
through quantization. They are distinguished by **what the loader did
with the file**, and that difference decides both how each is found and
how each is moved.
**Set 1 — unplaced weights.** The loader opened the file and read the
tensor, but the model had no parameter for it, so Transformers reports
it in `unexpected_keys`. An MTP head the recipe did not quantize is the
common case. Moved as **tensors**: located in whichever shard holds
them, read, and merged into the exporter's `extra_state_dict`.
**Set 2 — off-index sidecars.** The index never names the file, so the
loader never opened it and never had the chance to call anything
unexpected. GLM-4.7 keeps its MTP head in a standalone `mtp.safetensors`
exactly this way. Moved as **files**: copied byte for byte, so no host
memory is spent re-serialising tensors the export does not otherwise
touch.
#### The index is not an inventory of the checkpoint
This is the part that is easy to get wrong, and it cost a silent
data-loss bug during review.
`model.safetensors.index.json` selects which **files** the loader opens
— not which **tensors** it sees. Within a file it opens, Transformers
enumerates every tensor present and reports the unexpected ones.
Verified by experiment against transformers 5.3.0:
| case | reported in `unexpected_keys`? |
|---|---|
| key absent from the index, in a shard the index names for *other*
tensors | **yes** |
| key in a file the index never names (`mtp.safetensors`) | **no** — the
file is never opened |
So a tensor missing from `weight_map` but sitting inside a main shard is
**set 1, not set 2**. An MTP head stored that way is reported, recorded
— and was then silently dropped, because resolution went through
`weight_map`, which by construction has no entry for it. The
`--vllm_fakequant_export` guard shared that lookup, so the check written
to refuse exports that drop weights stayed silent in exactly the case it
existed for.
`locate_source_keys` now resolves through the index first (free for
everything it lists) and header-scans the shards only for the leftovers
— names, never tensor data — warning when a key is in no file at all.
The carry and the guard share it, so they cannot disagree again.
#### Flow
1. **At load.** `record_unplaced_source_keys` stores Transformers' own
`unexpected_keys` on the model (`_modelopt_unplaced_source_keys`) plus
the resolved local checkpoint path. The question asked is "does the
model have a parameter for this key", never "is this an MTP head" — so
the mechanism is architecture-agnostic.
2. **At export, before dispatch.** `read_unplaced_weights` resolves each
recorded key to its shard and reads the tensors, merging them into
`extra_state_dict`. An explicitly passed `extra_state_dict` wins on a
name clash: a caller naming a tensor is more specific than our
inference.
3. **Rank behaviour.** Only the rank that writes `extra_state_dict`
reads the bytes — the FSDP2 writer emits it from rank 0 alone, so a full
read on every rank would be host memory spent and discarded (a
DeepSeek-V3-class MTP head is 10 GB+ in bf16). The **key list** is still
resolved on every rank, because `get_quant_config` runs per rank and the
configs must agree.
4. **Recording what was written.** `export_hf_checkpoint` records
`_modelopt_carried_over_names` — the union of carried tensors and the
off-index sidecars' tensor names — before `get_quant_config` runs,
because that is the first point that knows what was *written* rather
than what was merely unplaced.
5. **Exclusions.** Both sets must reach `quantization_config.ignore`, or
a deployment framework reads the top-level `quant_algo` and tries to
load an original-precision weight as a quantized one (the NVBug 5718750
class). `seed_carried_over_exclusions` is the single path for this,
called by `get_quant_config` after its per-layer pass and again by the
layerwise exporter from `finalize()` — which snapshots its config during
`bind()`, while calibration is still running, so it cannot see the
carried set any earlier.
#### What is deliberately excluded
- **Files that re-ship indexed weights.** Mistral's
`consolidated.safetensors` is a second full copy of the model, and
PEFT's `adapter_model.safetensors` is an adapter. Both are off-index,
and copying either would put unquantized weights beside the quantized
ones — vLLM's mistral load-format looks for `consolidated.safetensors`
by name, so it is not inert. Caught by name for the known conventions
and by tensor-name overlap for the rest.
- **Files named like a main weight shard**, whatever the index says, so
a broken or partial index cannot make the real weights look like
sidecars.
- **Symlinks are *not* excluded.** A Hugging Face snapshot stores every
file as a symlink into a sibling `blobs/`, so refusing links would drop
the sidecar of every hub-downloaded checkpoint.
`resolve_checkpoint_file` checks where the link *lands* — regular file,
inside the checkpoint dir or its blob root — rather than whether it is a
link.
- **Buffers Transformers recomputes.** `*.inv_freq` is skipped by the
fake-quant guard even when a shard provides it: older
Llama/Mistral-lineage conversions do list it in the index, and refusing
an export over it would reject checkpoints that export correctly today.
### How a carried, never-quantized MTP head reaches
`quantization_config.ignore`
Raised in review: `_add_mtp_exclusions` is gone, and an unplaced weight
has no module, so
`get_quant_config` walks right past it. That was a real gap, not just a
documentation one —
fixed here.
The export writes a carried MTP head in its original precision. If it is
absent from
`exclude_modules`, a deployment framework reads the top-level
`quant_algo` and tries to load
`eh_proj` as an FP8/NVFP4 weight — the same class of failure as NVBug
5718750, where a
`transformers>=5.0` MoE router was written in BF16 but never excluded.
The two cases now differ only in *why* the module is invisible to the
quantizer walk:
| | why invisible | handled by |
|---|---|---|
| MoE router (tf≥5.0) | module exists, never gets a quantizer |
`_get_unquantized_moe_router_names` |
| carried weight | no module at all in the live model |
`_get_carried_over_module_names` |
`_get_carried_over_module_names` reads the keys the loader recorded as
unplaced
(`_modelopt_unplaced_source_keys`), strips the trailing parameter name —
a state-dict key is
`<module path>.<parameter>` — and dedupes. Those names are seeded into
`layer_config_dict` as
`QUANTIZATION_NONE`, exactly as the router pass does, so they flow
through
`process_layer_quant_config` into `exclude_modules`, which
`convert_hf_config` emits as
`quantization_config.ignore`.
An MTP head the recipe *does* quantize is unaffected: it has a module,
is loaded normally, is
never in the unplaced set, and is reported as quantized. The
backward-breaking note above still
holds — MTP follows the recipe instead of being force-excluded — but a
head that ends up carried
rather than quantized is no longer silently missing from `ignore`.
Covered by `test_carried_over_weights_are_excluded_from_quantization`
and
`test_carried_over_module_names_strip_parameter_and_dedup`.
`--vllm_fakequant_export` does not carry unplaced weights; it now raises
rather than writing a
checkpoint quietly missing them.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ — MTP layers now follow the
recipe rather than being force-excluded, and
`quantization_config.ignore` no longer lists layers the export may have
quantized (carried, never-quantized weights are still listed -- see
below). Shipped recipes are unaffected; see the Changelog entry.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 Backward Breaking Changes.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
Draft: the behavioural change to MTP quantization is the part most worth
a second opinion — it aligns `hf_ptq` with `megatron_bridge`, which
special-cases nothing.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Quantization supports enabled operators outside transformer layers,
including language-model heads.
* Exports preserve checkpoint weights not loaded into the quantized
model.
* Additional safetensors sidecar files are copied unchanged into
exported checkpoints.
* Unquantized auxiliary components, such as vision layers, remain
available in exported models.
* **Behavior Changes**
* Local recipe files take precedence over built-in recipes.
* Legacy architecture-specific recipe paths remain supported with
warnings.
* MTP-specific export exclusions are no longer applied.
* **Bug Fixes**
* Preserved weights are no longer incorrectly reported as unquantized.
* Incomplete exports are rejected when source weights cannot be placed
or preserved.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
0a8e70804a |
Fix hf_ptq.py discarding completed PTQ run on sanity-generate() failure (#2480)
### What does this PR do? Type of change: Bug fix `post_quantize()` in `examples/hf_ptq/hf_ptq.py` ran the optional post-quantization sanity-check `full_model.generate()` unguarded, directly before `export_quantized()`. Any exception raised there aborted the whole run and discarded a completed calibration without exporting a checkpoint. Root cause (traced from [NVBug 6752977](https://nvbugspro.nvidia.com/bug/6752977), DGX Spark GB10 / DeepSeek-R1-Distill-Llama-8B / NVFP4): `get_model()` loads with `device_map="auto"`, relying on `accelerate`'s `infer_auto_device_map`/`get_max_memory()` to decide GPU vs. CPU placement. On DGX Spark's unified-memory single-GPU host, that memory probe under-reports GPU capacity, so part of even an 8B model can land on CPU — and the existing fallback shrinks the GPU budget further (`* gpu_mem_percentage`), compounding it. Calibration survives this because it never invokes the real fake-quant kernel, but the post-PTQ sanity `generate()` does, and NVFP4's dynamic-block-quantize op (`modelopt/torch/quantization/tensor_quant.py`) hard-asserts `amax.is_cuda` with no CPU fallback, so any CPU-offloaded layer crashes there — after ~5.8 hours of calibration, before export. This PR does not attempt to fix the underlying `device_map`/memory-probing behavior (unverified without the actual hardware/logs, which weren't reachable from this environment). Instead it makes the failure mode safe: a failure in the optional sanity check now only skips that check and warns, and export always proceeds, regardless of why `generate()` failed. ### Usage No new API. Behavior change only: `examples/hf_ptq/hf_ptq.py` now completes export even if the post-quantization sanity `generate()` call raises. ### Testing - Added `tests/examples/hf_ptq/test_hf_ptq_args.py::test_post_quantize_export_survives_a_failed_sanity_generate`, which drives `post_quantize()` with a `full_model.generate()` that raises and asserts `export_quantized()` still runs. - Ran `pytest tests/examples/hf_ptq/test_hf_ptq_args.py` (48 passed). - Ran `pre-commit` on the changed files (`ruff-format` reformatted line wrapping only). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ <!-- pending: run `/claude review` --> ### Additional Information Fixes NVBug 6752977. Linked JIRA: OMNIML-5932. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Quantized checkpoint export now continues when the optional post-quantization generation check fails. - A warning is shown when the generation check cannot complete. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
d23030f91d |
[1/n] Adds skip-softmax calibration through the vLLM serving path (#1992)
### What does this PR do? Type of change: new feature Calibrates skip-softmax thresholds through the vLLM V1 execution path for FlashAttention and FlashInfer. Calibration measures the paged KV-cache path used at serving time, aggregates raw skipped/total tile counts across tensor-parallel head shards, fits separate prefill and decode curves, and exports the existing `sparse_attention_config` checkpoint schema. This uses raw counts rather than averaging per-rank sparsity ratios because TP ranks can contribute different tile populations; summing numerators and denominators before division preserves the global tile-weighted result. The vLLM adapter lives in `plugins/sparse_attn_calibration.py` rather than `SparseAttentionStatsManager`: the latter records module-local ratios for the HF calibration flow and has no aligned cross-process merge contract, while this path must merge per-sample raw counts from vLLM workers. Fitting and export still reuse `DynamicThresholdCalibrator` and the canonical conversion helpers so the model and checkpoint schema do not fork. Skip decisions depend on tile geometry. The common Triton launch boundary fixes the KV tile at 128 tokens and the prefill query tile at 128 tokens, including for direct kernel callers. Single-query decode can use a 16x128 compute tile without changing its skip decision. Measurement bypasses autotuning; serving still tunes warp and pipeline-stage counts while keeping the decision geometry fixed. ### Usage ```bash python examples/vllm_serve/calibrate_sparse_attn.py <CKPT> \ --prompts_file prompts.txt \ --target_sparse_ratio 0.7 \ --fit_logspace \ --tensor_parallel_size 4 \ --decode_tokens 32 \ --update_checkpoint_config ``` Calibration supports tensor parallelism and requires pipeline-parallel and data-parallel sizes of 1. It always writes `sparse_attention_config.json`; `--update_checkpoint_config` also merges the result into `<CKPT>/config.json`. ### Testing Latest revision `1e969cb380` (rebased onto main `02b58eb146`, 2026-09-17): - Calibration/count-fitting unit tests: **33 passed** (`test_sparse_attn_calibration.py` and `test_calibrator_fitting.py`). - Paged and contiguous calibration GPU suite: **33 passed** (`test_paged_calibrate.py` and `test_triton_fa_calibrate.py`), including NHD/HND equivalence, partial query tiles, decode counts, and malformed-cache rejection. Run with `CUDA_VISIBLE_DEVICES=1` on an RTX A6000; local GPU 0 was unavailable. - Calibration CLI tests: **21 passed** (`tests/examples/vllm_serve/test_calibrate_sparse_attn.py`). - `pre-commit run --files <four changed files>`: passed, including Ruff, mypy, and Bandit. - The new regression tests reproduced the skipped-counter truncation and missing cache-boundary checks before the fix. Calibration arithmetic and the 20-point threshold grid are unchanged. Historical validation from earlier revisions (not rerun end-to-end for this update): - `PYTHONPATH="$PWD" pytest -q tests/examples/vllm_serve/test_calibrate_sparse_attn.py tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_calibration.py` — 37 passed. - `PYTHONPATH="$PWD" pytest -q tests/gpu_vllm/torch/sparsity/attention_sparsity/test_vllm_calibration.py tests/gpu_vllm/torch/sparsity/attention_sparsity/test_sparse_attn_worker.py` — 65 passed, including kv-first, blocks-first, and packed FlashAttention cache layouts. - `PYTHONPATH="$PWD" pytest -q tests/gpu_vllm/torch/sparsity/attention_sparsity/test_vllm_runtime.py tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_config.py` — 33 passed. - `PYTHONPATH="$PWD" pytest -q tests/gpu/torch/kernels/sparsity/attention/test_paged_calibrate.py tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_calibrate.py tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_skip_softmax.py` — 31 passed, 1 skipped because the GPU lacks enough shared memory for the fp32 tile. - `pre-commit run --files <changed files>` — passed. - Historical end-to-end Nemotron 3 Ultra (GCP job `558552`), TP4, FA4, 48 RULER prompts, and 20 threshold trials: completed `0:0` with prefill `(a, b) = (9.9104, 10.8881)`, respectively +0.147% and -0.066% versus the matching 20-point reference `(9.8958, 10.8953)`. The supplied legacy fit `(14.47, 10.91)` used a different threshold grid; its `b` differs by only -0.201%, while `a` retains the known grid-weighting shift. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ Active skip-softmax fixes the calibrated decision geometry (serving still tunes warp/stage counts), and sparse-only vLLM installs fail fast for unsupported DCP, DBO/ubatching, speculative decoding, and FULL mixed-batch graphs instead of installing silently. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no copied code or new dependency. - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Pipeline parallelism is rejected during calibration because the current count-merging contract aligns records across tensor-parallel head shards, not across pipeline stages with disjoint attention layers. The unrelated HF padded-query behavior change was removed from this PR so it can be reviewed independently with its own compatibility test. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added vLLM skip-softmax calibration for paged attention, including prefill/decode support and checkpoint configuration generation. - Added Muse Glimmer AutoQuantize, Alpamayo QAD, streaming Kimi-K3 conversion, and NVFP4 activation headroom calibration recipes. - Added calibration statistics aggregation, phase-specific fitting, threshold validation, and preservation of existing sparse-attention settings. - **Bug Fixes** - Improved NVFP4 CPU/ONNX scale validation and clamping. - Added clearer handling for unsupported quantization, cache, CUDA graph, and engine configurations. - Standardized serving and calibration tile behavior. - **Documentation** - Expanded vLLM serving guidance, calibration instructions, compatibility requirements, and sparse-attention limitations. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kai Xu <kaix@nvidia.com> |
||
|
|
cf1f48fa0f |
[5565357] Fix SDXL NVFP4 export and performance (#2336)
### What does this PR do?
Type of change: Bug fix
Adds a compact SDXL and SDXL-Turbo mixed-precision FP4 recipe:
- block-16 NVFP4 for non-QKV Linear/GEMM layers;
- FP8 for Conv2d layers;
- high-precision Q/K/V projection Linears to preserve TensorRT
horizontal fusion;
- optional FP8 MHA quantization.
For SDXL FP4 export, Conv2d quantizers export directly through the
shared FP8 custom-op path. The previous `generate_fp8_scales` plus
`convert_zp_fp8` INT8 zero-point workaround is removed. The graph then
uses the existing FP8 Q/DQ normalization and `NVFP4QuantExporter`
lowering, with opset 23 for FLOAT4 support. Flux FP8 export also saves
the graph returned by its RoPE weight conversion.
This PR also changes shared exporter behavior:
- `_fp8_quantize` refreshes ONNX shape/type inference after applying the
custom FP8 operator's uint8 output metadata, affecting all FP8 ONNX
exports through this symbolic.
- `_quantized_sdpa` derives `disable_fp8_mha` from the live Q/K/V
quantizer state instead of a restored private module flag.
Other model recipe configurations remain unchanged.
### Usage
```bash
python quantize.py \
--model sdxl-1.0 \
--model-dtype Half \
--trt-high-precision-dtype Half \
--format fp4 \
--block-size 16 \
--batch-size 2 \
--calib-size 128 \
--n-steps 20 \
--quantized-torch-ckpt-save-path ./sdxl-fp4 \
--onnx-dir ./onnx-sdxl-fp4
```
### Testing
- CPU-only focused and generic NVFP4 exporter tests: 44 passed in 4.35
seconds.
- Focused Flux returned-graph save test: 1 passed.
- Required Linux unit CI at `034fe23ec` passed with the `all` dependency
set, including `tests/unit/examples/test_diffusers_fp4.py`.
- Latest changed-file pre-commit checks: all passed.
- TensorRT 10.14 on a B200 GPU:
- 302 native block-scaled NVFP4 GEMM tactics;
- 38 native FP8 Conv tactics;
- no FP4 Q/K/V projections;
- all 11 FP16 Q/K/V projection-fusion groups preserved;
- three alternating batch-2 profiles measured 18.614 ms FP4 versus
20.028 ms FP16 median UNet latency, a 7.06% reduction.
- FP8 SDXL/SD3 ONNX-to-TensorRT end-to-end runs were not executed
because they require explicit approval. The existing end-to-end test
matrix now includes SD3 FP8 alongside SDXL FP8.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ — no public API or CLI flags
change; the shared changes preserve the intended FP8 export and
attention behavior.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — the shared NVFP4 opset, FP8 shape-inference, and Diffusers
attention-policy changes are recorded under bug fixes.
- Did you get Claude approval on this PR?: N/A
### Additional Information
Tracking: [5565357]
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added SDXL support for mixed NVFP4/FP8 quantization, including
convolution and softmax handling.
- Added an SDXL quantization preset for streamlined post-training
quantization workflows.
- Expanded FP4 ONNX export support to Flux and SDXL, with improved
FP4/FP8 graph processing and export reliability.
- Added automatic quantization policy and format restoration from
checkpoints.
- **Documentation**
- Documented SDXL layer behavior, optional FP8 attention quantization,
and Blackwell/TensorRT requirements for FP4 and FP8 deployment.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
> 🤖 _Generated by Codex (AI agent)._
---------
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
|
||
|
|
02b58eb146 |
Fix vLLM fakequant calibration for hybrid attention models (#2414)
### What does this PR do? Type of change: Bug fix Fix fakequant calibration for hybrid attention/Mamba models, including NVIDIA Nemotron-3-Nano, on vLLM 0.26 and 0.28. The manual calibration scheduler path previously submitted requests with empty KV-cache block tables. Hybrid models require scheduler-compatible cache state during prefill; on current vLLM releases the empty tables caused the Mamba state to use the reserved null block and calibration activations became NaN. Request cleanup also no longer matched the vLLM 0.28 execution lifecycle, which could leave request-scoped state in the persistent batch. This PR: - Allocates non-null scratch blocks for every KV-cache group using the vLLM warmup reservation policy. - Supports both the vLLM 0.28 reservation helper and the equivalent vLLM 0.26 calculation. - Passes newly allocated blocks through `new_block_ids_to_zero` when that scheduler field is available. - Validates that the calibration batch fits in the configured cache and reports how to reduce calibration demand if it does not. - Cleans up calibration requests through a zero-token scheduler step on current vLLM, with a direct cleanup fallback for older runners. - Updates the example Dockerfile to default to vLLM 0.28.0 while retaining vLLM 0.26.0 through `VLLM_VERSION`. - Documents the validated Nemotron-3-Nano NVFP4 KV-cache workflow and clarifies that reducing `--max-num-batched-tokens` is not required. ### Usage Build the default vLLM 0.28.0 image: ```bash docker build -f examples/vllm_serve/Dockerfile \ -t vllm-modelopt:v0.28.0 . ``` Build with vLLM 0.26.0: ```bash docker build --build-arg VLLM_VERSION=0.26.0 \ -f examples/vllm_serve/Dockerfile \ -t vllm-modelopt:v0.26.0 . ``` Calibrate and serve Nemotron-3-Nano with NVFP4 KV-cache fakequant: ```bash KV_QUANT_CFG=NVFP4_KV_CFG QUANT_CALIB_SIZE=512 \ python examples/vllm_serve/vllm_serve_fakequant.py \ <nemotron3_nano_model_path> \ --trust-remote-code --enforce-eager -tp 8 \ --max-model-len 8192 --host 0.0.0.0 --port 8000 ``` ### Testing Validated on omniml-a0 with `NVIDIA-Nemotron-3-Nano-30B-A3B-BF16`, tensor parallel size 8, `NVFP4_KV_CFG`, `QUANT_CALIB_SIZE=512`, and `--max-model-len 8192`. No `--max-num-batched-tokens` override was used. - vLLM 0.28.0: - All 512 calibration samples completed. - No NaNs or cache-cleanup warnings were observed. - The server started and `/health` passed. - An OpenAI-compatible completion request returned coherent generated text. - vLLM 0.26.0: - Repeated the same 512-sample TP8 calibration with the official `vllm/vllm-openai:v0.26.0` image. - No NaNs were observed. - The server started, passed `/health`, and returned coherent generated text. - Docker: - Built and verified the updated vLLM 0.28.0 image. - Focused tests: - `tests/examples/vllm_serve/test_vllm_mlflow_utils.py`: 32 passed. - Cleanup failure, missing legacy API, and legacy fallback tests: 5 passed on both vLLM 0.26.0 and 0.28.0. - Repository hooks: - Targeted pre-commit hooks for every changed Python, Markdown, and Docker file: passed. - `git diff --check`: passed. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — added focused coverage for fail-closed cleanup, exception chaining, and the legacy cleanup fallback; the full regression was also validated end to end. - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information The change is quantization-format agnostic. It corrects the calibration scheduler and cache lifecycle rather than special-casing `NVFP4_KV_CFG` or using an NVFP4 cast path. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added support for configuring the vLLM version through `VLLM_VERSION`, with vLLM 0.28.0 as the default. - Added calibration and serving guidance for hybrid attention/Mamba models, including Nemotron 3 Nano with NVFP4 KV-cache fake quantization. - **Bug Fixes** - Improved calibration block handling across supported vLLM versions. - Improved calibration cleanup to preserve original errors and provide reliable fallback behavior when standard cleanup is unavailable. - **Documentation** - Documented tested versions, direct installation commands, ModelOpt setup, serving options, and guidance to avoid NaNs during batched serving. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> |
||
|
|
835c041c58 |
fix(specdec): resolve the eagle aux-layer preset in the vLLM hidden-state dump (#2410)
### What does this PR do? Type of change: Bug fix Fixes `nvbugs/6753684`, filed against #2080 by the ModelOpt QA Sentinel. The vLLM offline hidden-state dump rejected `--aux-layers eagle` — **the flag's own default** — so the documented invocation aborted before writing any state: ``` File "collect_hidden_states/compute_hidden_states_vllm.py", line 76, in _resolve_aux_layers_standalone ids = sorted({int(t) for t in aux_layers.split(',') if t.strip()}) ValueError: invalid literal for int() with base 10: 'eagle' ``` **Root cause.** `compute_hidden_states_vllm.py` runs in a stock vLLM container, where importing `modelopt.torch` fails (the full init chain pulls in omegaconf and friends). It therefore carries `_resolve_aux_layers_standalone`, a local copy of the preset logic in `common.resolve_aux_layers`. That copy implemented the `dflash` preset and explicit id lists, but never `eagle` — while `add_aux_layers_args` defaults to `eagle`. The HF and TRT-LLM dumps call the shared helper and were unaffected; only the vLLM path forked, and nothing compared the fork against its source. This PR resolves `eagle` inline, mirroring `hf_eagle.default_eagle_aux_layer_ids`. It also fixes a second defect the bug exposes: the function already had a message naming the accepted values, but it was unreachable, because `int()` raised first. An unrecognised preset now reports what it accepts instead of surfacing the raw `int()` error — which is what made the original failure opaque. ### Usage The previously-broken documented invocation now works: ```bash cd examples/speculative_decoding python collect_hidden_states/compute_hidden_states_vllm.py \ --model Qwen/Qwen2.5-0.5B-Instruct \ --input-data ../dataset/synthetic_conversations_1k.jsonl \ --output-dir /tmp/hs_vllm \ --max-seq-len 512 --tp 1 ``` `--aux-layers dflash` and explicit lists such as `--aux-layers 2,5,8` are unchanged. ### Testing Added `tests/unit/examples/test_vllm_hidden_states_aux_layers.py`, which pins the standalone copy to the shared implementation it mirrors: - `eagle` matches `hf_eagle.default_eagle_aux_layer_ids` across layer counts 4, 6, 8, 12, 24, 28, 32, 36, 48, 52, 61, 80 — deliberately including counts small enough that the `max(0, ...)` clamps collapse ids together. - A named regression case for `nvbugs/6753684`. - `dflash` and explicit-list behaviour unchanged. - Unknown specs (`bogus`, `EAGLE3`, `eagle3`, empty) raise the actionable message. - Out-of-range ids still rejected. Divergence here is silent — the dump would write plausible-looking hidden states from the *wrong* layers, surfacing much later as a poor acceptance rate. Hence pinning to the reference rather than asserting hardcoded lists alone. All 20 assertions verified and every pre-commit hook passes (`ruff`, `mypy`, `bandit`, RST lint, license headers). One caveat worth stating plainly: **pytest could not be run locally.** `tests/unit/conftest.py` imports `modelopt.torch.utils.distributed`, which needs `CPUOffloadPolicy` from `torch.distributed.fsdp` — absent in this machine's torch. Each assertion was executed directly against the real module instead, but CI is the first genuine pytest run. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — strictly widens accepted input; `dflash` and explicit lists behave identically. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — bug fix for a defect present in a previous release. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information The underlying fragility is the duplicated implementation, not this one missing branch. The function's own `TODO: drop this once common.resolve_aux_layers is decoupled from the heavy modelopt.torch import chain` is the real fix; the new test narrows the gap but does not close it. Worth tracking separately if the vLLM dump is expected to keep pace with new presets. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed `--aux-layers eagle` for vLLM offline hidden-state collection. * Added support for the documented `eagle` preset alongside `dflash` and explicit layer IDs. * Improved invalid-option errors to clearly list accepted formats. * Rejects `dflash` configurations when the target model has too few layers. * Continues rejecting layer IDs outside the model’s available range. * **Documentation** * Added a v0.48.0 changelog entry for the fix. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
a448ba9757 |
Add end-to-end W4A4 NVFP4 + QAD tutorial for Qwen3.6-35B-A3B (#2411)
### What does this PR do? Type of change: new example + bug fix <img width="2085" height="1239" alt="image" src="https://github.com/user-attachments/assets/b9ced215-ce8c-4dbe-be74-a75c1c4714b3" /> Adds an end-to-end **W4A4 NVFP4 + Quantization-Aware Distillation** tutorial for [Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) at `examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B/`. It complements the existing Nemotron-3-Nano tutorial (pruning + distillation + FP8). Here the model is unpruned and the technique under test is **W4A4** — aggressive enough that PTQ alone leaves a measurable accuracy gap, which is what QAD exists to close. **Why W4A4 rather than weight-only NVFP4:** W4A16 measured *slower than BF16* in 10 of 12 shapes, because a BF16 activation forces vLLM onto the Marlin dequant fallback and never reaches the Blackwell FP4 tensor cores. W4A4 beats BF16 in 9 of 12 shapes (up to 1.30x) and shrinks the checkpoint 67 GiB -> 22 GiB (3.1x). **What the study found:** only 2 of 6 benchmarks show a statistically significant PTQ deficit, so those are the only two QAD can recover. IFBench is recovered to parity with BF16 (-2.62 pp -> -0.29 pp, gain of +2.33 pp, p=0.036); MMMU-Pro recovers ~40% and retains a significant gap. The other four are lossless under W4A4 to begin with. Also added: - `modelopt_recipes/model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore.yaml` — the PTQ recipe used as the QAD student, usable via `--recipe`. - `data_blend.yaml` — the token-budgeted blend config for the distillation data. - `eval_configs/*.yaml` — one NeMo Evaluator config per benchmark. tau2-bench is separate because it needs `--enable-auto-tool-choice --tool-call-parser qwen3_coder` and `deployment.command` is global to a config. **Two export fixes found while producing these checkpoints** (both change library/example behaviour, both have changelog entries under 0.48.0 Bug Fixes): - `unified_export_megatron.py` — MCore builds `embedding` on the MTP stage as well as the first, so gating export on `hasattr(model, "embedding")` wrote a **second, unreferenced copy of the vocab embedding** whenever an MTP model was exported with PP > 1. The index mapped the key to the later shard, so the extra copy never loaded but still shipped — ~1 GB for this model. Now gated on `model.pre_process`, MCore's own "this rank owns the input embedding" flag. - `export_quantized_megatron_to_hf.py` — stopped passing Megatron's `moe_router_dtype` as the router's *storage* dtype. It is a routing *compute* dtype; the parameter is bf16 in a bf16 model, so the export was widening bf16 to fp32. All 21,495,808 router values in the exported checkpoint have their low 16 bits zero, and vLLM builds the gate at the model dtype and rounds on load, so the dropped bytes carried no information. `export_mcore_gpt_to_hf` still accepts the override. ### Usage ```bash # 1. PTQ (2 GB200 nodes for EP=8) srun ... python examples/megatron_bridge/quantize.py \ --hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \ --recipe model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore \ --tp_size 1 --ep_size 8 --pp_size 1 \ --calib_dataset_name cnn_nemotron_v2_mix --calib_num_samples 1024 --calib_batch_size 1 \ --seq_length 8192 --skip_generate \ --export_megatron_path /path/to/qwen36_w4a4_megatron # 2. QAD (32 nodes x 4 GB200) python -u examples/megatron_bridge/distill.py \ --teacher_hf_path Qwen/Qwen3.6-35B-A3B --student_hf_path Qwen/Qwen3.6-35B-A3B \ --student_megatron_path /path/to/qwen36_w4a4_megatron \ --tp_size 1 --pp_size 1 --cp_size 1 --ep_size 8 \ --seq_length 32768 --mbs 1 --gbs 512 --train_iters 500 \ --lr 1e-5 --min_lr 1e-6 --lr_warmup_iters 50 --logit_kl_topk 4096 \ --recompute_granularity full --recompute_method uniform --recompute_num_layers 1 \ --no_async_save --eval_iters 0 --save_interval 50 \ --data_paths "${DATA_BLEND}" --output_dir /path/to/qad_output ``` ### Testing **Library changes.** `tests/gpu_megatron/torch/export/test_unified_export_megatron.py` gains `test_unified_export_megatron_pp2_mtp_no_duplicate_tensors`: it exports a PP=2 model built with `mtp_num_layers=1` and asserts no tensor lands in more than one shard. Verified to **fail without the fix**: ``` AssertionError: tensors written to more than one shard: {'model.embed_tokens.weight': ('model-00001-of-00002.safetensors', 'model-00002-of-00002.safetensors')} ``` The pre-existing `..._pp2_mtp_metadata_matches_shards` test cannot catch this — it fakes `_get_mtp_state_dict` on a model with no real MTP, so the last stage never builds an embedding. Ran the whole `tests/gpu_megatron/torch/export/` suite with and without the fix: identical failure sets (3 failures both ways, all `qwen3_5_moe_vl_*` from a local `ImportError: FLA is not installed`), 58 passed with vs 56 without — the +2 being the new test's two workers. `tests/unit/recipe` passes 368/368 after the recipe path move. `model.pre_process` is always present: `GPTModelExporter.__init__` raises unless the model is `GPTModel` or `HybridModel`, and both set it unconditionally. Both export fixes were also applied to the real 23 GB checkpoints and re-validated end to end: every retained tensor md5-identical, index/shard integrity re-checked, and a **full GPQA re-evaluation of the fixed checkpoint** scored 83.49 vs 84.25 before (paired per-question t-test over the same 198 questions x 16 repeats: -0.76 pp, p=0.21, not significant). **Numbers in the tutorial** come from real runs, not estimates: - **253 evaluation runs** across BF16, the published W4A16 checkpoint, W4A4 PTQ, and QAD at 50 / 300 / 500 iterations — 8 repeats per benchmark (3 for tau2-bench; GPQA is one `num_repeats: 16` run). - The published `nvidia/Qwen3.6-35B-A3B-NVFP4` checkpoint was re-evaluated under this same harness (36 runs) rather than quoted from its card, so the W4A16 row is same-harness. - Every figure and results-table value is generated from the collected `results.yml` files by a script, and I verified the README table cell-by-cell against that data after each edit. - Throughput rows were cross-checked against the recorded AIPerf sweeps; the QAD wall-clock figures against the two jobs' Slurm records (`03:34:49` + `02:09:49`). - All CLI flags in the tutorial were verified to exist in `quantize.py` / `distill.py` / `export_quantized_megatron_to_hf.py`, and `cnn_nemotron_v2_mix` against `dataset_utils.py`. The tutorial also records the non-obvious constraints found the hard way: QAD on this model requires `TP=PP=CP=1` (TP breaks quantizer `_amax` dist-checkpoint sharding, PP starves Qwen3-VL's M-RoPE of `position_ids`, CP hits a rope shard mismatch), EP must match the PTQ checkpoint, and `--logit_kl_topk` is mandatory at 32K because the dense `[seq, vocab]` fp32 logits are 30.31 GiB per tensor on a 248,320-token vocabulary. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ <!-- gpu_megatron PP=2+MTP export dedup test --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!-- 0.48.0: Megatron Framework + two Bug Fixes --> - Did you get Claude approval on this PR?: ✅ <!-- not yet run --> ### Additional Information Changelog entries are filed under **0.48.0**; the `cherry-pick-0.47.0` label has been removed. Rebased onto `main` after #2328 renamed `modelopt_recipes/huggingface` to `model_type` (it is now a compatibility symlink), so the recipe moved to `model_type/qwen3_6_moe/ptq/`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an end-to-end Qwen3.6-35B-A3B tutorial for W4A4 NVFP4 quantization and quantization-aware distillation. * Added checkpoint export, accuracy evaluation, and vLLM throughput benchmarking workflows. * Added evaluation configurations for AA-LCR, GPQA, IFBench, MMMU-Pro, SciCode, and tau2 Telecom. * Added a token-budgeted supervised fine-tuning data configuration. * Added a Megatron-Core NVFP4/FP8 quantization recipe for Qwen3.6-MoE. * **Documentation** * Added benchmark results, deployment guidance, hardware requirements, reproduction steps, HTTPS endpoint guidance, announcement filters, and tutorial links. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c7ed23a103 |
Rename modelopt_recipes/huggingface to model_type with backward-compat alias (#2328)
### What does this PR do?
Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.
Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.
- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
`$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.
### Usage
```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
--recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none
# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
--recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```
```python
from modelopt.recipe import load_recipe
load_recipe("model_type/vit/ptq/fp8") # canonical
load_recipe("huggingface/vit/ptq/fp8") # deprecated alias, resolves to the same recipe
```
### Testing
- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
`model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
the loader alias — including a recipe that pulls internal `$import`s.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.
### Additional Information
The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.
- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.
- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
- Local recipe files now take precedence over built-in recipes.
- Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
3c87751903 |
Deprecate the single-format quantization CLI flags in favour of --recipe (#2426)
### What does this PR do?
Type of change: deprecation
Deprecates the single-format quantization CLI flags in favour of
`--recipe`. Passing one now emits a `DeprecationWarning`; nothing else
changes.
| script | flags |
|---|---|
| `examples/hf_ptq` | `--qformat`, `--kv_cache_qformat` |
| `examples/megatron_bridge/quantize.py` | `--quant_cfg`,
`--kv_cache_quant`, `--weight_only` |
| `examples/torch_onnx` | `--qformat` |
`--recipe` was already authoritative over all six — silently on
`hf_ptq`, and with a runtime warning on `megatron_bridge` — and
`modelopt/recipe/presets.py` already records the intent in a comment:
*"the long-term direction is to retire `--qformat` /
`--kv_cache_qformat` in favour of `--recipe`"*. This makes that a real
deprecation.
A recipe carries the quantization config, the calibration algorithm and
the KV-cache setting in one file, so they cannot drift apart the way
separate flags can. That drift is not hypothetical: the preset path
applies no MTP exclusion while the recipe unit
`default_disabled_quantizers` disables `mtp.*`, so the same model
quantizes differently depending on which entry point was used.
#### The warning fires only when a flag is actually passed
`RecipeSupersededAction` is an `argparse.Action`, and argparse invokes
an action only for options present on the command line — never for a
default. That matters because several of these default to a *quantizing*
value (`--qformat fp8`, `--kv_cache_qformat fp8_cast`); warning on the
defaults would fire on every run, including runs that correctly use
`--recipe` and never mention the flag.
`examples/speculative_decoding/scripts/quantize_drafter.py` keeps
`--qformat` undeprecated: it has no `--recipe`, so there would be
nothing to migrate to.
### Usage
```bash
# deprecated
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> --qformat nvfp4 --kv_cache_qformat fp8_cast
# replacement
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> \
--recipe general/ptq/nvfp4_experts_only-kv_fp8_cast
```
### Testing
Three tests in `tests/examples/hf_ptq/test_hf_ptq_args.py`, all passing:
- passing `--qformat` / `--kv_cache_qformat` raises `DeprecationWarning`
and still parses the value;
- omitting them raises nothing and leaves the defaults (`fp8`,
`fp8_cast`) untouched;
- the action stays wired to both flags, so a future edit cannot drop it
while leaving the help text.
Defaults and parsed values were diffed against `main` and are unchanged
— the action stores exactly what `store` / `store_true` would have.
`ruff` findings are at parity with `main` on every changed file.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — the flags still work, they
only warn.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 Deprecations.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
Draft: the removal release for these flags is not decided here, only the
deprecation.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Deprecations**
* Legacy quantization CLI options now issue visible `FutureWarning`
messages only when explicitly provided.
* Use `--recipe` instead of deprecated options in Hugging Face PTQ,
Megatron-Bridge, and torch-to-ONNX workflows.
* Existing option values, defaults, and parsing behavior remain
unchanged.
* Weight AutoQuantize recipes without an explicit `kv_cache` setting
continue to use `--kv_cache_qformat` as a fallback.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
5b1f7e86cc |
[6701308][OMNIML-5805] Correct ONNX PTQ documentation contracts (#2413)
### What does this PR do?
Type of change: documentation
Align the ONNX PTQ README, guide, and executable example with the
implemented contracts:
- use the canonical `--calibration_data_path` CLI option;
- load `.npy` calibration data before passing it to the Python API;
- document the supported Autotune modes and calibration methods;
- correct the minimum opsets to INT8 19, FP8 19, and INT4 21; and
- describe the no-data fallback as random calibration inputs.
This also removes an inaccurate source comment without changing runtime
behavior.
### Usage
```bash
python -m modelopt.onnx.quantization \
--onnx_path=model.onnx \
--quantize_mode=int8 \
--calibration_data_path=calib.npy \
--output_path=model.quant.onnx
```
### Testing
- `pre-commit run --files docs/source/guides/_onnx_quantization.rst
examples/onnx_ptq/README.md modelopt/onnx/quantization/quantize.py
tests/examples/test_onnx_ptq.sh`
- `bash -n tests/examples/test_onnx_ptq.sh`
- `CUDA_VISIBLE_DEVICES="" python -m pytest -o addopts="" -p
no:cacheprovider --confcutdir=tests/unit/onnx/quantization -q
tests/unit/onnx/quantization/test_autotune_quantization_integration.py`
(4 passed)
- `nox -s docs` (passed; Sphinx built 881 HTML files)
- Focused before/after contract probe covering the documented CLI
option, API data type, Autotune modes and methods, opset minimums, and
random-input wording
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A
### Additional Information
Tracking: [6701308]
> 🤖 _Generated by Codex (AI agent)._
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Documentation**
- Clarified that random calibration inputs are used when no calibration
dataset is provided.
- Updated ONNX post-training quantization examples with minimum opset
requirements and the `calibration_data_path` argument.
- Clarified Autotune support for FP8 and INT8 calibration methods using
`max` or `entropy`.
- **Tests**
- Updated quantization command examples to use the current calibration
data path option.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
|
||
|
|
bd90a5ed51 |
[OMNIML-5774] Add BEVFormer ONNX PTQ and evaluation example (#2208)
### What does this PR do? Type of change: new example Adds an end-to-end BEVFormer-tiny ONNX PTQ example under `examples/onnx_ptq/bevformer` with: - The [BEVFormer Dockerfile](https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile) from pinned NVIDIA DL4AGX commit `9f7b291`, rather than a second container definition in this repository. - A mounted Model Optimizer checkout, the official DL4AGX TensorRT 10 source patch, and TensorRT plugin compilation at container runtime. - Ordered temporal calibration-data generation that propagates `prev_bev`, resets state at scene boundaries, computes CAN bus deltas, writes the exact requested sample count, and publishes output atomically. - INT8 and FP8 quantization with BEVFormer-specific calibration defaults, custom-plugin handling, `MatMul` exclusions, and FP16 fallback policy. - Strongly typed FP16, INT8, and FP8 TensorRT engine generation plus nuScenes evaluation instructions. - Shared temporary-ONNX-copy handling for BEVFormer and VoVNet quantization so shape inference cannot mutate the source model. - Focused CPU-only tests for temporal state, exact-count cleanup, quantization defaults, plugin configuration, and source-model preservation. The ONNX PTQ index continues to document the shared PETR/FAR3D containers separately and links to the BEVFormer guide. ### Usage Clone the pinned DL4AGX revision and build its BEVFormer image: ```bash git clone https://github.com/NVIDIA/DL4AGX.git /path/to/DL4AGX git -C /path/to/DL4AGX checkout --detach \ 9f7b29104c253d5bc68334e7b83b3eecb72d4572 docker build \ --build-arg TORCH_CUDA_ARCH_LIST=8.9 \ --file /path/to/DL4AGX/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile \ --tag modelopt-onnx-bevformer \ /path/to/DL4AGX/AV-Solutions/bevformer-int8-eq ``` The example command targets compute capability 8.9; use the deployment GPU's compute capability for another architecture. After exporting the model and generating temporal calibration data, quantize it with: ```bash python /opt/Model-Optimizer/examples/onnx_ptq/bevformer/quantize.py \ --onnx=/artifacts/bevformer_tiny_epoch_24_cp2_op13.onnx \ --calibration-dir=/artifacts/calibration \ --trt-plugins=/workspace/BEVFormer_tensorrt/TensorRT/lib/libtensorrt_ops.so \ --quantization-mode=fp8 \ --output=/artifacts/bevformer_tiny_epoch_24_cp2_op13.fp8.onnx ``` See `examples/onnx_ptq/bevformer/README.md` for dataset setup, plugin compilation, export, FP16 feedback-engine creation, temporal calibration, INT8/FP8 engine builds, and evaluation. ### Validation #### Current revision: three-sample smoke validation The Dockerfile at the pinned DL4AGX revision built successfully, and the current Model Optimizer checkout was mounted into it for the workflow below. The runtime audit confirmed TensorRT 10.14.1.48, CUDA 13.1, Torch 2.9, the TensorRT/CUDA/CPU execution providers, and plugin linkage. A fresh three-frame A/B/B smoke passed without calculating partial NDS or mAP: - Fresh Torch 2.9 export, AutoCast, and strongly typed FP16, INT8, and FP8 engine builds passed. - Temporal calibration published exactly three batches with `use_prev_bev=[0, 0, 1]`, zero state at both scene starts, recurrent feedback on the third frame, and the expected CAN bus deltas. - The source ONNX hash remained unchanged. INT8 contained 136 Q/DQ pairs with INT8 zero points; FP8 contained 127 Q/DQ pairs with FP8 zero points. - Every engine produced finite outputs and 300 non-empty decoded detections for each frame; maximum scores ranged from 0.9489 to 0.9608. CPU-only validation passed the 16-test focused example set and the full ONNX session (`649 passed, 1 skipped`). Full-diff pre-commit and the warning-as-error documentation build also passed. #### Historical full accuracy reference The following results were collected previously with TensorRT 10.14.1.48 on an NVIDIA RTX 6000 Ada Generation GPU, using 600 ordered calibration samples and all 6,019 nuScenes validation samples. They are reference results, not a completed full evaluation of the current revision. | Precision | NDS | mAP | | :-- | --: | --: | | FP16 | 0.3546 | 0.2515 | | INT8/FP16 | 0.3512 | 0.2505 | | FP8/FP16 | 0.3526 | 0.2489 | #### Performance Following the PETR and FAR3D convention, performance is reported only as speedup normalized to FP16. The archived full-validation engines were benchmarked with five interleaved TensorRT 10.14 trials on the same NVIDIA RTX 6000 Ada Generation GPU. Each trial used the `trtexec` GPU Compute Time median with data transfers disabled, CUDA Graphs enabled, spin-wait, a 1-second warmup, a 10-second measurement window, and one inference stream. | Precision | Speedup vs. FP16 | | :-- | --: | | FP16 | 1.00x | | INT8/FP16 | 1.87x | | FP8/FP16 | 1.18x | TODO: Investigate why FP8 delivers less speedup than INT8 for BEVFormer-tiny. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors). - Is this change backward compatible?: ✅ - If you copied code from another source or added a new PIP dependency, did you follow the guidance in `CONTRIBUTING.md`?: ✅ - Did you write the necessary tests?: ✅ - Did you update the [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ ### Additional information Reference workflow: [NVIDIA DL4AGX BEVFormer INT8 example](https://github.com/NVIDIA/DL4AGX/tree/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq). > 🤖 _Generated by Codex (AI agent)._ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an end-to-end BEVFormer 3D detection workflow for ONNX post-training quantization, including temporal calibration, INT8/FP8 quantization, TensorRT engine generation, and nuScenes evaluation. * Added command-line tools for preparing calibration data and quantizing BEVFormer models. * **Bug Fixes** * Prevented source ONNX models from being overwritten during quantization and improved temporary model cleanup. * Fixed FP8 export for BF16 models during real-weight compression. * **Documentation** * Added setup, usage, compatibility, and performance guidance for the BEVFormer workflow. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com> |
||
|
|
5d2d5a5d15 |
deprecate trtllm-build in weight_sparsity (#2371)
### What does this PR do?
Type of change: export PTS/finetuned model to Hugging Face checkpoint,
then replace trtllm-build with trtllm-serve
Renamed export_trtllm_ckpt.py to export_hf_ckpt.py.
Replaced the legacy export_tensorrt_llm_checkpoint() flow with
export_hf_checkpoint().
Fix bug: 5823190
<!-- Details about the change. -->
### Usage
```
python examples/llm_sparsity/weight_sparsity/hf_pts.py --model_name_or_path Llama-3.1-8B-Instruct --device cuda --model_max_length 1024 --dtype fp16 --sparsity_fmt sparsegpt --calib_size 128 --output_dir Llama-3.1-8B-Instruct_pts
python examples/llm_sparsity/weight_sparsity/export_hf_ckpt.py --model_name_or_path Llama-3.1-8B-Instruct --model_max_length 1024 --dtype fp16 --modelopt_restore_path Llama-3.1-8B-Instruct_pts/pts_modelopt_state.pth --output_dir Llama-3.1-8B-Instruct_pts/trtllm/ckpt_pts
trtllm-serve Llama-3.1-8B-Instruct_pts/trtllm/ckpt_pts \
--tp_size 1 \
--pp_size 1 \
--host 0.0.0.0 \
--port 8000
```
### Testing
PTS and SAT tested
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A
### Additional Information
N/A
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Updated sparsity example instructions to export Hugging Face
checkpoints and serve models with `trtllm-serve`.
* Documented tensor and pipeline parallelism, host and port settings,
and the OpenAI-compatible chat completions endpoint.
* Corrected the PTS model restoration path.
* **Bug Fixes**
* Model export now saves the tokenizer alongside the checkpoint.
* Model length configuration is interpreted as an integer.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Noey Yang <174223378+noeyy-mino@users.noreply.github.com>
|
||
|
|
59d93af064 |
[OMNIML-5570, OMNIML-5569] 1/2 Add layer-wise KV-cache AutoQuant with forward KL (#2272)
### What does this PR do?
Type of change: new feature.
Adds standalone layer-wise KV-cache AutoQuantize through the existing
public
`mtq.auto_quantize` API:
- dispatches KV search with
`constraints={"effective_bits": ..., "cost_model": "kv_cache"}` and
forward-KL
sensitivity;
- selects one supported K/V format for every eligible causal-attention
layer;
- supports persistent/exportable FP8 K/V, NVFP4 K/V, and FP8-K/NVFP4-V
candidates;
- solves a K/V-width- and scale-storage-aware additive recipe with the
existing
PuLP-backed constrained solver;
- uses `BaseSearcher` lifecycle and safe checkpoint restore/save
machinery;
- preserves existing non-KV execution while isolating K/V candidate
calibration;
- returns standard AutoQuantize state that can be re-solved at another
KV budget;
- produces a complete KV-only replay config that disables every non-KV
quantizer;
- saves JSON-safe sensitivity metadata and the exact selected layer
mapping; and
- invokes the public API from `examples/hf_ptq/hf_ptq.py` through a
standalone
calibration-free recipe.
The implementation is architecture-driven. Plain and
conditional-generation Qwen
causal attention is supported, VLM vision attention is excluded through
the existing
language-model extraction boundary, hybrid full-attention mixers are
discovered through
their paired K/V quantizers, and nonattention/Mamba modules remain
outside the search.
Ambiguous language-model roots, unsupported distributed execution,
structural
algorithms, invalid storage declarations, nonpersistent scales, and
unsupported K/V
pairs fail closed.
KV-only unified HF exports leave weight-quantization fields unset.
Uniform all-FP8 or
all-NVFP4 selections retain their legacy KV scheme while also carrying
the complete
`kv_cache_quantized_layers` map and schema version; genuinely
layer-mixed selections use
the KV-side `MIXED_PRECISION` marker plus the same map. This keeps
weight-loader metadata
accurate and prevents disabled vision attention from making uniform
language-model KV
quantization appear partially quantized.
GEMM PTQ/AutoQuantize followed by KV AutoQuantize is intentionally
excluded and proposed
separately in stacked PR #2273.
### Why KV search has a dedicated backend
The user-facing entry point remains `mtq.auto_quantize`; no separate
public KV search API
is introduced. `AutoQuantizeKVSearcher` extends `BaseSearcher` and
reuses its reset,
checkpoint load/save, and search lifecycle, along with existing Pydantic
configuration,
calibration, safe checkpoint I/O, and PuLP-backed selection utilities.
The backend remains KV-specific because a decision owns paired K/V
quantizers on one
attention layer, its cost depends on separate K/V widths and data/scale
storage, BF16 is a
scoring reference but not a deployable solver choice, and the
optimization objective is
additive isolated forward KL under a KV-storage constraint. These
contracts do not match
the weight-domain hparam grouping, parameter-count cost, or
threshold-selection behavior
of the existing weight AutoQuant searchers. Keeping the specialization
behind the shared
API avoids changing established weight-search solver and scoring
behavior.
### Usage
```bash
python examples/hf_ptq/hf_ptq.py \
--pyt_ckpt_path Qwen/Qwen3.8-27B \
--recipe general/auto_quantize/kv_fp8_nvfp4_cast_kl_div_at_5p4bits \
--auto_quantize_checkpoint /path/to/kv_autoquant.pth \
--export_path /path/to/qwen3.8-27b-mixed-kv
```
The search checkpoint is compatible only with the same model,
eligible-layer geometry,
candidate configurations, and scoring setup. Use a distinct checkpoint
path after any of
those inputs change.
KV-cache AutoQuantize rejects `--use_fsdp2` before model loading because
its sensitivity
scoring, selection, and checkpoint writes are single-process. Existing
weight
AutoQuantize retains its previous experimental FSDP2 warning and
behavior.
### Testing
- Focused coverage exercises candidate validation/calibration, paired
K/V scoring and
storage accounting, solving, checkpoint resume, failure atomicity,
disabled layers,
fresh-model replay, Qwen/VLM/hybrid boundaries, JSON-safe reports, and
unified export.
- Uniform FP8/NVFP4 KV-only exports retain the legacy KV scheme and
complete layer map
without claiming a weight algorithm; disabled VLM vision attention is
excluded from
causal-KV eligibility.
- The shipped recipe runs end to end on a tiny offline Qwen fixture and
preserves
exportable scale state.
- After merging current `main`: 432 focused recipe/KV/export/hf_ptq
tests passed, with one
unrelated optional-dependency skip; changed-file pre-commit hooks
passed.
### Deployment gate
The producer schema is covered here. Runtime consumption of
`kv_cache_quantized_layers` is tracked in vLLM PR
https://github.com/vllm-project/vllm/pull/52813. Do not treat a produced
checkpoint as
runtime-supported until that consumer lands and the target K/V kernels
are available.
### Before your PR is "Ready for review"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow
guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅
### Additional information
- This is split from the combined ground-truth implementation in draft
PR #2211 to reduce
review scope; composition is isolated in #2273.
- The standalone core tree contains no composed GEMM→KV recipe schema or
orchestration.
- No model-name checks, checkpoint-specific layer lists, campaign data
contracts, cluster
launch logic, or runtime-kernel implementations are included.
---------
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
|
||
|
|
d69e93a72b |
Record the MLflow run that produced a checkpoint in .experiment.json (#2374)
### What does this PR do?
Type of change: new feature
A tracked `hf_ptq` run already tags itself with the checkpoint it writes
(`checkpoint_path`), so a run can be followed to its output. The reverse
was missing: given a checkpoint on disk, there was no way to find the
run that quantized it without searching the tracking server by path.
A tracked run now writes `.experiment.json` into `--export_path` naming
the experiment, the MLflow run id and the run URL, and uploads the same
bytes as the `experiment.json` artifact so a downloaded artifact set is
self-describing. `MlflowRunLogger` gains a `run_info` property carrying
that identity, with the tracking URI credential-masked the way `run_url`
already was.
Two deliberate behaviours:
- **Written from a `finally`**, so a run that crashes after export still
leaves the pointer behind.
- **Skipped when the export directory is absent** — a run that exported
nothing has nowhere to put it, and creating the directory would suggest
a checkpoint that does not exist. The artifact is still uploaded in that
case, so a failed run is traceable from the server side.
A failed local write warns and continues rather than failing the job,
consistent with the rest of the MLflow path. Only the main rank writes,
since the logger is inert on other ranks.
### Usage
```bash
python hf_ptq.py --pyt_ckpt_path Qwen/Qwen3.5-0.8B --qformat fp8 \
--export_path /tmp/qwen35-fp8 --mlflow https://<your-mlflow-server>
```
```console
$ cat /tmp/qwen35-fp8/.experiment.json
{
"tracking_uri": "https://<your-mlflow-server>",
"experiment_name": "alice/hf_ptq/Qwen3.5-0.8B-fp8",
"experiment_id": "36",
"run_id": "7bec239a3a154970b062f3024a5ff20e",
"run_name": "20260910-175422",
"run_url": "https://<your-mlflow-server>/#/experiments/36/runs/7bec239a3a154970b062f3024a5ff20e"
}
```
```python
# checkpoint -> run
import json, mlflow
info = json.load(open("/tmp/qwen35-fp8/.experiment.json"))
mlflow.set_tracking_uri(info["tracking_uri"])
run = mlflow.get_run(info["run_id"])
```
### Testing
**Unit** — `tests/unit/torch/utils/test_mlflow.py` (61 passed):
`run_info` contents before/after the run opens, the defaulted run name
being reported rather than left blank, and credential masking of the
tracking URI.
**Example** — `tests/examples/hf_ptq/test_hf_ptq_args.py` (27 passed):
the file landing in the checkpoint and on the server with identical
content, the failed-run path, the no-export path, and untracked runs
writing nothing.
**Real runs**, 1x H200, `Qwen3.5-0.8B` FP8 PTQ,
`tensorrt-llm/release:1.3.0rc26`:
- Against a local MLflow server — checkpoint copy and uploaded artifact
byte-identical; artifacts on the run were `command.txt`,
`experiment.json`, `logs/hf_ptq.log`, `summary/quant_summary.txt`,
`version.txt`.
- Against the internal `mlflow-modelopt` server (experiment
`chenjiel/hf_ptq/Qwen3.5-0.8B-fp8`, run
`7bec239a3a154970b062f3024a5ff20e`) — same result, confirming artifact
upload against a real backend. Reading `.experiment.json` back and
calling `mlflow.get_run(run_id)` resolved to `FINISHED` with
`checkpoint_path` pointing at the export directory.
- Crash path exercised for real when a first attempt died on a gated
calibration dataset: no export directory created, `experiment.json`
still uploaded, run closed `FAILED`.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — new entry under `*Misc*` in the open 0.48.0 section, matching where
the MLflow entries sit in 0.47.0.
- Did you get Claude approval on this PR?: ❌
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Exported checkpoints now record experiment and run traceability
metadata in `.experiment.json`.
* Checkpoint metadata is uploaded with opened MLflow runs, including
runs where export fails.
* Active MLflow run details—including identifiers, resolved run name,
URL, and tracking server—are available with credentials redacted.
* **Bug Fixes**
* Improved handling of failed, untracked, and pre-existing exports to
prevent inherited metadata pointers.
* **Documentation**
* Updated MLflow integration guidance and changelog information for
checkpoint metadata.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
a74054ab2b |
Let callers add MLflow tags to a fakequant serve's run (#2364)
### What does this PR do?
Type of change: new feature
The quantization run records what this library can see — the model, the
checkpoint, the vLLM and ModelOpt versions — but nothing about the
harness that launched it. A downstream tool that wants its own revision,
a sweep id, or a ticket number on the run has no way to put it there
today:
- `_run_tags()` returns a fixed dict
- `quant_config` (which becomes the run's params) is a hardcoded set of
`QUANT_*` variables
- MLflow itself has no environment variable for arbitrary tags
`MODELOPT_MLFLOW_EXTRA_TAGS` takes comma-separated `key=value` pairs and
merges them into the run's tags.
Two details worth a reviewer's attention:
**It joins `MLFLOW_ENV_VARS`.** A Ray-backed serve receives only the
variables named there, and the tracker runs in the rank-0 worker —
omitting it would make the feature silently do nothing under Ray.
**Caller tags are merged first**, so the library's own keys (`tool`,
`model`, `checkpoint_path`, `vllm_version`) are written over them and
keep describing the run truthfully whatever a caller sends.
`key=value` rather than JSON, learned from a live run: the variable
reaches the worker through a shell `export VAR="..."`, and JSON's own
double quotes terminate that quoting —
```
export MODELOPT_MLFLOW_EXTRA_TAGS_732b_DEPLOYMENT="{"internal_version": "4d8c"}"
```
arrived as `{`. A quote-free format survives verbatim and needs no
`json` import or exception handling. Splitting on the first `=` keeps
values that contain one, such as a URL with a query string.
### Usage
```bash
export MODELOPT_MLFLOW_EXTRA_TAGS="modelopt_internal_version=49fa29d5,sweep=kv-study"
python3 vllm_serve_fakequant.py "$MODEL" --mlflow https://your-mlflow-server/ ...
```
### Testing
Unit-level, over the helper: unset and empty variable, one and several
pairs, surrounding whitespace, an empty value, an entry with no `=`, a
trailing comma, and a value containing `=`. None raise; malformed
entries warn and are skipped.
End to end on a real fakequant serve (Nemotron-3-Nano-30B-A3B BF16,
`NVFP4_DEFAULT_CFG`, TP=8, Ray executor, vLLM 0.15, SLURM):
```
modelopt_internal_version '49fa29d5'
modelopt_version '0.47.0rc0.post32+gd38ed5ead'
git_sha 'd38ed5ead'
quant_cfg 'NVFP4_DEFAULT_CFG'
```
The tag was written by the `RayWorkerWrapper` process, which exercises
the whole path — env var → shell export → `--container-env` → raylet →
Ray actor → `_run_tags` — and confirms the `MLFLOW_ENV_VARS` entry is
doing its job. Also verified that the emitted payload survives a shell
export round-trip unchanged.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ <!--- Additive; with the
variable unset the tags are exactly as before. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ <!--- Verified manually as
above; there is no existing test module for vllm_mlflow_utils. Happy to
add one if you would like it. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <!--- Small additive feature in an example; tell me if it warrants an
entry. -->
- Did you get Claude approval on this PR?: ❌
### Additional Information
Consumed by Model-Optimizer-Internal MR !141/!147, which sets the
variable so a fakequant eval records the same harness commit on both its
quantization run and its evaluation-score run.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
d19925e446 |
simple refactor(export): split TensorRT-LLM-only code into modelopt/torch/export/trtllm (#2365)
### What does this PR do?
Type of change: refactor.
**The TensorRT-LLM checkpoint export format is deprecated.** Per
`docs/source/deployment/1_tensorrt_llm.rst`: *"The
`export_tensorrt_llm_checkpoint` API will be deprecated in future
releases. Users are encouraged to transition to the unified HF export
API, which provides enhanced functionality and flexibility for exporting
models to multiple inference frameworks including TensorRT-LLM, vLLM,
and SGLang."*
That deprecated code was not sitting off to one side — it was
**interleaved with the export path we actually want to grow.**
`modelopt/torch/export` mixed the deprecated TensorRT-LLM checkpoint
logic with the framework-agnostic HF/Megatron export code, in the same
modules:
- `layer_utils.py` was 1,986 lines, of which ~1,600 were TensorRT-LLM
`build_*_config` builders. The HF path imports this module for five
small predicates (`is_moe`, `is_quantlinear`, …) and dragged the whole
deprecated builder set in with them.
- `model_config.py` held the TensorRT-LLM `ModelConfig` dataclasses
*and* the `QUANTIZATION_*` / `KV_CACHE_*` constants that every backend
needs, so all of HF export imported the deprecated checkpoint schema to
get a format name string.
- `quant_utils.py` carried two helpers whose only caller is the
deprecated `postprocess.py`.
**This PR isolates the deprecated format so it stops polluting the
HuggingFace export path.** Everything reachable only from
`export_tensorrt_llm_checkpoint` now lives under
`modelopt/torch/export/trtllm/`, and the dependency is **one-way**:
`trtllm/` reaches into the parent through `quant_format`, `quant_utils`
and `layer_utils`, and **no implementation module in the parent imports
`trtllm/`.** The single exception is the deprecation re-export in
`modelopt/torch/export/__init__.py` described below, which is scheduled
for deletion in 0.49.0.
That one-way edge is the property worth protecting in review. It means
the deprecated format can be evolved, frozen, or eventually removed
without touching HF export, and HF export can no longer accidentally
grow a dependency on it.
### Deprecation handling
The format has carried a deprecation notice in the deployment docs since
`bc546943b4` (2025-10-08, first shipped in 0.39.0) — about 11 months.
But the deprecation policy in `README.md` also specifies *how* a
deprecation is communicated: a changelog entry, a source statement of
timing, and a runtime warning on use. **None of those existed**; only
one docs page ever said anything. So 0.48.0 is the first release that
gives users a signal they can act on, and this PR treats it as the
*start* of the migration period rather than the end:
- Both entry points now emit a `DeprecationWarning` naming 0.48.0 and
the 0.49.0 removal.
- `export_tensorrt_llm_checkpoint` and
`torch_to_tensorrt_llm_checkpoint` **remain importable from
`modelopt.torch.export`** for this release only, so existing callers
keep working *and* actually receive the warning. Removing the path in
the same release that first warns would mean callers hit `ImportError`
and never see it.
- The 0.49.0 removal date is stated in all four channels the policy
names: the runtime warning, the source (`.. deprecated:: 0.48.0` plus a
comment), the changelog, and the deployment doc.
The deeper module paths (`modelopt.torch.export.model_config_export`,
`modelopt.torch.export.model_config`) are **not** forwarded. Neither
appeared in a docs example, and `model_config.py` never declared
`__all__`, so by the `__all__` convention in `CONTRIBUTING.md` they were
never part of the public surface.
Eight modules had no non-TRT-LLM importer and moved whole:
`model_config_export`, `model_config_utils`, `postprocess`,
`distribute`, `tensorrt_llm_utils`, `tensorrt_llm_type`,
`hf_config_map`, `mcore_config_map`.
Three were genuinely mixed and were split by call-graph analysis rather
than by file:
| module | stayed shared (HF path) | moved to `trtllm/` (deprecated) |
|---|---|---|
| `model_config.py` | `QUANTIZATION_*`, `KV_CACHE_*`,
`FUSION_FREE_FORMATS` → new leaf module `quant_format.py` | the
`ModelConfig` dataclasses + `LINEAR_*`/`LAYERNORM_*` checkpoint-layout
constants |
| `layer_utils.py` | 9 module-shape predicates and MoE quantizer helpers
(`is_moe`, `is_quantlinear`, `get_experts_list`,
`sync_moe_gate_up_amax`, …) | the 39 `build_*_config` builders and
enc/dec helpers |
| `quant_utils.py` | everything else | `get_scaling_factor_from_weight`,
`resmooth_and_get_scale` (only caller is `trtllm/postprocess.py`) |
`adjust_attn_amax_values` was deliberately left in the shared
`quant_utils.py`: it has no production caller at all (only a test), so
"used only by TRT-LLM export" is not demonstrable for it.
Nothing was added or removed. `export_tensorrt_llm_checkpoint` behaves
exactly as before, just from a new import path and with a warning
attached.
### Usage
```python
# Deprecated TensorRT-LLM checkpoint export — new home, and warns on call
from modelopt.torch.export.trtllm import (
export_tensorrt_llm_checkpoint,
torch_to_tensorrt_llm_checkpoint,
)
from modelopt.torch.export.trtllm.model_config import ModelConfig
# The pre-0.48 path still works for one release, and warns — removed in 0.49.0
from modelopt.torch.export import export_tensorrt_llm_checkpoint
# Shared format constants — new home, still re-exported from the top level
from modelopt.torch.export.quant_format import QUANTIZATION_NVFP4, KV_CACHE_FP8
from modelopt.torch.export import QUANTIZATION_NVFP4 # still works
# The recommended path — unchanged
from modelopt.torch.export import export_hf_checkpoint, get_model_type
```
### Testing
- `pre-commit` on all changed files: passes (ruff, ruff-format,
**mypy**, bandit, markdownlint). mypy caught one implicit re-export of
`is_layernorm`, now imported from the shared module directly.
- `tests/unit/torch/export`: **189 passed**. With the new `trtllm/` test
dir: **193 passed**.
- Full `tests/unit/torch`: **2367 passed, 0 export failures**. The 45
failures are pre-existing environment issues — a deepspeed circular
import and a read-only HF cache — confirmed by reading their error text,
not assumed.
- `pytest tests/gpu/torch/export --collect-only`: 172 items, no
collection error.
- In-repo consumers updated and re-verified by an AST scan that imports
every `modelopt.torch.export*` module referenced anywhere in the tree
and checks each imported name still resolves: `hf_ptq.py`,
`export_trtllm_ckpt.py`, `deepseek_v3/ptq.py`, the AutoQuantize
notebook, `hf_ptq/README.md`, 2 docs pages, 4 tests.
- **Deprecation contract is covered by committed tests** (3 new, in the
`trtllm/` test dir): the pre-0.48 top-level import still resolves to the
same objects, `torch_to_tensorrt_llm_checkpoint` warns *at call time*
rather than on first `next()` (it returns a generator, so a naive
`warnings.warn` in the body would fire late or never), and one
`export_tensorrt_llm_checkpoint` call emits exactly one warning rather
than two. The first of these makes closing the migration window early a
test failure rather than a silent regression. `pyproject.toml` sets no
`filterwarnings = error`, so no suite fails on the new warning.
- **After merging `main`** (4 commits, incl. a 180-line rewrite of
`unified_export_megatron.py` that touches a file this PR also edits):
merged with no conflicts, then re-verified rather than trusted — import
scan clean across 24 export modules, `ruff check` clean repo-wide, 193
export tests passing, GPU collection still clean.
**Not run: the GPU suites** (`tests/gpu/torch/export`,
`tests/gpu_trtllm`) — no GPU in my environment.
`tests/gpu/torch/export/test_export.py` had its imports retargeted, so
it is the one most worth a GPU run before merge.
### Reviewer note: the deprecated path has no test coverage
Worth knowing before reviewing. **No test in the repo — including
`tests/examples/` — calls `export_tensorrt_llm_checkpoint`,
`torch_to_tensorrt_llm_checkpoint`, any `build_*_config`,
`convert_to_tensorrt_llm_config`, or `postprocess_model_config`.** So
~4,600 moved lines have no direct tests, and this refactor is validated
by import-graph reasoning, lint and mypy rather than by tests exercising
the moved code.
Given the format is deprecated and scheduled for removal in 0.49.0, **no
new coverage is planned for the conversion path itself** — writing fresh
tests for an API being removed next release isn't a good use of effort.
The gap is documented so reviewers can weigh the risk, not as a TODO.
(The deprecation *mechanism* is tested; see Testing.)
One caveat on how the gap was established: a runtime check showing all
12 `trtllm` modules in `sys.modules` after the export suites is *not*
evidence of coverage — importing any submodule runs
`trtllm/__init__.py`, which star-imports `model_config_export` and pulls
in the rest. Real line coverage could not be measured (`coverage`'s
tracer is incompatible with this venv's torch build: `ValueError: module
functions cannot set METH_CLASS or METH_STATIC`, on both the C tracer
and `sysmon`). The claim rests on a call-site audit generated from the
actual public symbols of those modules.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ for the public API —
`export_tensorrt_llm_checkpoint` and `torch_to_tensorrt_llm_checkpoint`
remain importable from `modelopt.torch.export` through the 0.49.0
migration period, now with a `DeprecationWarning`. The undocumented
submodule paths `modelopt.torch.export.model_config_export` and
`.model_config` did move; see **Usage**.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
code or dependencies; existing code relocated.
- Did you write any new necessary tests?: ✅ — 3 tests covering the
deprecation contract (old import path, call-time warning, exactly-one
warning). One existing test also moved to mirror the source split. No
new coverage for the deprecated conversion path itself; see the note
above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 **Deprecations**, covering both the runtime warning and
the new import location.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
Git detected the moves, so the diff stays reviewable: 8 files show as
pure renames (100%), the three split files as rename/copy at 94–99%
similarity, and only `layer_utils.py` as a 79% rewrite — expected, since
it shed 1,616 lines to `trtllm/`.
`examples/hf_ptq/hf_ptq.py` and
`examples/llm_sparsity/weight_sparsity/export_trtllm_ckpt.py` still call
the deprecated API, so those examples now print the warning. That is the
intended nudge, but happy to silence or migrate them if preferred. They
import from the new `.trtllm` path already, so they need no change at
0.49.0.
Two incidental changes, easy to revert if unwanted:
- `modelopt/torch/export/layer_utils.py` mode `100755 → 100644` (it was
needlessly executable).
- The new test is named `test_trtllm_quant_utils.py`, not
`test_quant_utils.py`: these directories have no `__init__.py`, so
pytest derives the module name from the bare filename and the shorter
name fails collection with `import file mismatch` against the existing
`test_quant_utils.py` one level up.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added shared quantization and KV-cache format definitions for export
workflows.
- Added expanded TensorRT-LLM export support, including broader model
architecture and quantization handling.
- Added distributed export utilities for coordinating checkpoint data
across processes.
- **Deprecation**
- TensorRT-LLM checkpoint export now emits a warning and is scheduled
for removal in version 0.49.0.
- Use the documented export module and save optimized model state
explicitly when needed.
- **Documentation**
- Updated guides and examples with new import paths and deprecation
guidance.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
9d0df45849 |
specdec: config_overrides for nested text_config checkpoints + load VLM-capable bases in merge_lora (#2289)
### What does this PR do?
Type of change: New feature + bug fix
Two related gaps, both hit while enabling EAGLE3 on a checkpoint whose
config nests its text dims.
**1. `config_overrides` for checkpoints whose `text_config` dims don't
propagate.**
Some multimodal checkpoints carry the real text-tower dims only under
`config.text_config`, leaving the parent fields `None`.
`from_pretrained` then builds a text tower with the wrong shape.
`load_vlm_or_llm` gains an optional `config_overrides` dict applied to
*both* the parent config and its `text_config` before instantiation, and
the three entrypoints that load checkpoints — `ar_validate.py`,
`export_hf_checkpoint.py`, `merge_lora.py` — get a `--config_overrides`
passthrough. `main.py` threads it from `ModelArguments`.
**2. `merge_lora.py` could not merge into any VLM base.**
It loaded via `AutoModelForCausalLM`, which cannot load architectures
absent from the CausalLM Auto map — every VLM base failed. It now goes
through `load_vlm_or_llm`, which routes VLMs to
`AutoModelForVision2Seq`/`AutoModelForImageTextToText` and plain LLMs to
`AutoModelForCausalLM` with the same `dtype`/`device_map`, so LLM
behavior is byte-for-byte unchanged.
Also adds an optional `transformers_cosmos3` import so `cosmos3_omni` is
registered with `AutoConfig` before use, and dispatches that
`model_type` to its model class directly — that plugin registers only a
*config*, never a model under `Auto*`, so `AutoModelForCausalLM` raised
`KeyError('cosmos3_omni')` regardless of imports. The import is wrapped
in `contextlib.suppress(ImportError)`, so it is a no-op when the plugin
isn't installed.
### Usage
```bash
# Checkpoint whose real dims live under config.text_config
python examples/speculative_decoding/scripts/ar_validate.py \
--model_path <ckpt> --trust_remote_code \
--config_overrides '{"num_hidden_layers": 36, "intermediate_size": 12288, "num_key_value_heads": 8}'
# Same flag on export and merge
python examples/speculative_decoding/scripts/export_hf_checkpoint.py \
--model_path <ckpt> --export_path <out> --config_overrides '{"num_hidden_layers": 36}'
python examples/speculative_decoding/scripts/merge_lora.py \
--base_model_path <base> --exported_lora_dir <out> --output_path <merged> \
--config_overrides '{"num_hidden_layers": 36}'
```
```python
model = load_vlm_or_llm(path, config_overrides={"num_hidden_layers": 36}) # default None
```
### Testing
Exercised end-to-end on a Cosmos3-Nano (16B, 36-layer text tower) EAGLE3
LoRA run:
- **Training** — the base loads with all 36 text layers and correct
dims; two 4-epoch co-training runs completed (46,816 steps each).
- **Export + merge** — produced `adapter_model.safetensors` and a merged
base. Verified correct by per-layer weight diff: a `start_layer=18` run
changed **exactly** layers 18-35, with layers 0-17 bit-identical to the
base.
- **AR validation** — `--config_overrides` loads the trained checkpoint;
80/80 MT-Bench samples, AR 3.42.
- **Regression check** — `merge_lora` via `load_vlm_or_llm` produces a
base loadable by `lm_eval`; ifeval/arc_challenge/winogrande all ran to
completion.
No local unit-test run: `nvidia-modelopt` isn't installed in my
checkout, so `tests/conftest.py` fails to import. Relying on CI.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `config_overrides` defaults
to `None`; the `merge_lora` loader swap keeps the same class, dtype and
device_map for plain LLMs.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependency; `transformers_cosmos3` is an optional import guarded by
`contextlib.suppress`.
- Did you write any new necessary tests?: ❌ — exercising these paths
needs a checkpoint with a nested `text_config`, which the unit suite has
no fixture for. Happy to add one if a reviewer can point me at a small
suitable model.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — can add a *Speculative Decoding* entry for the `merge_lora` VLM fix
if you consider it changelog-worthy.
- Did you get Claude approval on this PR?: ❌ — not yet run.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added JSON-based model configuration overrides across speculative
decoding, training, validation, export, and LoRA workflows.
- Overrides can update primary model and text configuration settings.
- Expanded support for vision-language models and Cosmos3 Omni
checkpoints.
- **Bug Fixes**
- Improved configuration handling for offline loading and
checkpoint-based initialization.
- Restored draft-model precision during checkpoint loading and model
conversion.
- Added validation for malformed, unsupported, and non-finite override
values.
- Standardized configuration override guidance across command-line
workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
279d510616 |
fix(specdec): correct resume and bound staging in the vLLM hidden-state dump (#2080)
### What does this PR do?
**Type of change:** Bug fix
Fixes two issues in the vLLM offline hidden-state dump
(`examples/speculative_decoding/collect_hidden_states/compute_hidden_states_vllm.py`).
Both are invisible on small dumps and only bite at scale, which is why
they survived until
now — they were found while dumping ~194k conversations for a MiniMax-M3
draft.
**1. Resume silently re-processed already-finished work.**
`keep_conversation` skips conversations whose `.pt` already exists, but
that predicate reads
**on-disk state**, which is not part of the fingerprint `datasets`
computes for `filter()`
(it hashes the function and the dataset). With a persistent HF cache
reused across a resumed
or requeued run, the cached *"keep everything"* result from an earlier
run — computed when
few or no `.pt` files existed — is replayed. The run then re-generates
and **overwrites**
conversations it had already completed, and reports `Removed 0
conversations due to existing
output files` while doing so.
Observed on a 194k-conversation dump: ~62k `.pt` rewritten over a
two-hour window with the
total output count completely flat.
Fix: pass `load_from_cache_file=False` so the filter re-checks the disk
on every run.
**2. Staging exhausted `/dev/shm` partway through large dumps.**
The script generated the **entire** dataset before saving anything. The
KV connector stages
each conversation's hidden states under its `shared_storage_path`
(`/dev/shm`, i.e. RAM, by
default) and they are only freed by `cleanup_hidden_states()` in the
save loop — so every
conversation stayed staged simultaneously. On a large dump this exhausts
the space and the
connector starts failing writes:
```
Hidden-states write failed for req_id=...:
SafetensorError('Error while serializing: I/O error: No space left on device (os error 28)')
```
Fix: generate and save in chunks of `--save-chunk-size` (default 256),
so at most one chunk
is staged at a time. As a side benefit the dump becomes **incrementally
durable** — an
interrupted run (walltime limit, node failure) keeps its finished
conversations and the
resume path above continues from them, instead of losing the whole run's
work.
### Testing
- Reproduced both failures on a 194k-conversation MiniMax-M3 dump (8-way
DP, TP8), and
confirmed both fixes on the same workload: after the change the output
count advanced
monotonically across requeues (123k → 194k) with no rewrites, and
`/dev/shm` stayed bounded
through completion.
- `pre-commit run --files ...` passes (ruff check/format, mypy, bandit,
license, rst checks).
- Behavior is unchanged for a fresh single-shot dump other than the
chunked generate calls;
the default `--save-chunk-size 256` is the only new knob.
### Additional Information
Extracted from #1749, which is otherwise superseded by the streaming
DFlash/DSpark path — these
two fixes are model-agnostic and apply to any offline dump, so they are
worth landing on their
own.
### Before your PR is "*Ready for review*"
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No — the failure modes are
multi-process/at-scale (datasets cache reuse across runs, connector RAM
staging) and are not reproducible in the unit-test harness.
- **Did you add or update any necessary documentation?**: Yes —
CHANGELOG entry.
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added chunked hidden-state generation for large vLLM offline runs.
* Added a configurable save-chunk size, defaulting to 256 conversations.
* Enabled incremental saving and resumption of hidden-state outputs.
* **Bug Fixes**
* Improved resume filtering to accurately detect existing output files.
* Reduced memory usage by saving and releasing each generated chunk.
* Ensured temporary files are cleaned up after interrupted or skipped
saves.
* Added atomic output-file replacement to prevent incomplete results.
* Added validation to prevent invalid conversation IDs from creating
unsafe paths.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
56af187565 |
ar_validate: fail loudly when every sample fails (#2288)
### What does this PR do?
Type of change: Bug fix
`validate_ar()` catches per-sample exceptions, prints a `WARNING`, and
returns whatever succeeded. When *every* sample failed it returned an
empty list, and the reporting block was guarded by `if results and
accelerator.is_main_process:` — so the script printed no results and
exited **0**. A run where 100% of samples failed was indistinguishable
from a successful one.
This bit us on a real run: an EAGLE3 checkpoint loaded with
`device_map="auto"` was sharded across 8 GPUs, every one of the 80
samples died with `Expected all tensors to be on the same device`, and
the job still exited 0 with no AR number anywhere in the log — the
wrapper stamped it PASS.
Now it raises, so the caller sees a non-zero exit. Any previously
"passing" run that printed no AR number was never meaningful.
### Usage
No API change. Existing invocations are unaffected when at least one
sample succeeds:
```bash
python examples/speculative_decoding/scripts/ar_validate.py \
--model_path <ckpt> --steps 3 --osl 1024 --num_samples 80
```
### Testing
Reproduced the silent-pass on a Cosmos3-Nano EAGLE3 checkpoint (80/80
samples failing): before this change the job exited 0 and stamped PASS;
after it, the job exits non-zero with the sample failures visible.
Confirmed the normal path is unchanged by a subsequent run that
completed 80/80 and printed AR 3.42.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — only affects the
all-samples-failed case, which previously produced no output and a
misleading exit 0.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ — the failure path requires
a model that errors during AR validation; the existing suite has no
harness for that.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — behavior fix in an example script, not a released-feature change.
- Did you get Claude approval on this PR?: ❌ — not yet run.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved validation error handling when all samples fail.
* Validation now rejects non-positive sample counts before processing.
* Empty validation results are clearly distinguished from cases where
all samples fail.
* Error messages report the actual number of validation samples
attempted, capped at the available dataset size.
* Empty validation results are no longer reported as successful.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Ye Yu <yeyu@nvidia.com>
|
||
|
|
4956213d67 |
Unblock Qwen3.5/3.6 QAD: Megatron export, calibration, and distillation fixes (#2334)
### What does this PR do? Type of change: Bug fix Everything that blocked running QAD on a quantized Qwen3.5 / Qwen3.6 MoE checkpoint: two Megatron-Core → HuggingFace export bugs that make it unservable (§1–2), the dead code the first leaves behind (§3), a no-op flag (§4), a multi-GPU calibration deadlock (§5), and four distillation / data-prep bugs that stopped QAD itself from running (§6). #### 1. Routed experts were exported packed, and vLLM cannot load that ``` AttributeError: Layer language_model.model.layers.23.mlp.experts has no parameter 'w2_weight_weight_scale_2' for checkpoint weight ...experts.down_proj_weight_scale_2 ``` `mcore_qwen35vl.py` used `GroupedMLPPacking`, mirroring the **BF16 upstream** checkpoint, which really is packed. But that mapping is only used for **quantized** export, and vLLM's quantized MoE loader needs per-expert scales — both released NVFP4 checkpoints (`Qwen3.6-35B-A3B-NVFP4` via hf_ptq, `Nemotron-3.5-Lightning-30B-A3B-NVFP4` via Megatron-LM) are per-expert. `_grouped_mlp_slicing` gains `gate_proj_name` / `up_proj_name` to split each expert's fused gate+up and slice its per-block `weight_scale`; `GroupedGatedMLPSlicing` wires it up. The Megatron checkpoint layout is unchanged, so affected checkpoints need only a **re-export**. `_verify_exported_keys` is relaxed to match: exported modules now contribute their ancestor prefixes, so expanding one source module into many is not reported as ~82 dropped tensors. A module genuinely absent still has nothing beneath its prefix and is still caught. #### 2. A quantized `output_layer` (`lm_head`) could not be checkpointed `GPTModel.sharded_state_dict` drops `output_layer._extra_state` and asserts it is empty. ModelOpt keeps quantizer state there, so saving raised and — since that method also backs the load plan — loading silently restored the layer **unquantized**. `keep_gpt_output_layer_extra_state()` retains it, applied from `megatron_replace_quant_module_hook` so **every** Megatron model gets it (Megatron-LM and NeMo users included, neither of whom can import `mbridge`, which needs `megatron.bridge`). It matches the upstream body by AST before replacing it and self-disables otherwise. [NVIDIA/Megatron-LM#7086](https://github.com/NVIDIA/Megatron-LM/pull/7086) is **closed, not merged**: nemo:26.10 migrates `GPTModel` to `HybridModel`, whose `sharded_state_dict` has no pop-and-assert, so this side keeps the workaround. Not cosmetic: `lm_head` is 248320×2048 = 509M params, **34.6% of per-token weight traffic** on a model with ~2.9B active params. #### 3. Cleanup Nothing maps `GroupedMLPPacking` once qwen3_5 is switched over; it is removed with `_grouped_mlp_packing` and the `quantize=` / `record_quant_config=` parameters that existed only to serve it. Llama-4's `PackNameRemapping` is unaffected. Two smaller review-driven fixes: the gated-split shape checks raise `ValueError` rather than `assert` (stripped under `-O`), and per-expert quant metadata is recorded for `local_expert_indices` rather than every global id, fixing non-contiguous EP. #### 4. Remove the no-op `--moe_calib_experts_ratio` from the Megatron quantize example `examples/megatron_bridge/quantize.py` accepted the flag and threaded it into the `mtq` config, but `_moe_calib_experts_ratio` exists only in `plugins/huggingface.py` (9 refs) and never in `plugins/megatron.py` (0); `mode.py:247` only assigns it to modules already exposing the attribute. On a Megatron MoE model it was accepted and silently ignored — a trap, since on a 256-expert model it reads like a major quality lever. `hf_ptq.py` keeps it, where it works. #### 5. Fix multi-GPU image-text (VLM) calibration deadlocking VLM calibration hung for 30 minutes and died on a gloo timeout whenever `world_size > 1`, with no error until the timeout fired. `NemotronTarPlusJsonlIterable` split its budget with truncating division, so the stream supplied fewer samples than requested (1024 over 3 subsets → 341×3 = **1023**). `_ShardedIterable` gives rank *r* items *r, r+W, r+2W…*, so a stream that is not a multiple of `world_size` leaves the trailing rank one short — it exits the forward loop early and the others block on the next collective. The arithmetic predicts both observed hangs exactly: 1024 → stall at **255/256**, 512 (yielding 510) → **127/128**. Fixed both ends: subset budgets are distributed with `divmod` so they sum exactly, and `_ShardedIterable` truncates every rank to `floor(len / world)` — which also covers `num_samples` not being divisible by `world_size`, as the first fix alone does not. Verified on Qwen3.6-35B-A3B (EP=4, `nemotron_vlm_dataset_v2`, 1024 samples): the configuration that hung twice now completes 256/256 and exports. Unit tests cover both fixes and fail without them. #### 6. Fix the distillation path so QAD can actually run Four independent bugs, all hit while running QAD end to end on Qwen3.6-35B-A3B. Each blocks a different configuration, and together they made every sequence length OOM or abort. - **Context parallel aborts.** The DDP config derived `average_in_collective` from `--sft` alone, but context parallel also needs per-token loss reduction, so any `--cp_size > 1` run died on `Cannot average in collective when calculating per-token loss`. - **`TopKLogitsKLLoss` was not memory-efficient.** Despite documenting "without gathering full logits", it cast the *whole* vocabulary to FP32 before selecting the top-k, allocating two `[seq, vocab]` tensors — 30.3 GiB each at seq 32768 on this model's 248k vocab. Reducing before the cast is equivalent: widening is exact and temperature scaling is monotonic, so the selected entries and the loss are unchanged. - **MTP cross-entropy ran when it had nothing to recover.** `skip_lm_loss` exempts the MTP heads unconditionally, so their CE materialised another FP32 `[seq, vocab]` tensor even when the MTP head is excluded from quantization — as it is in every recipe here (775 of 906 `exclude_modules`, zero MTP `weight_scale` tensors exported). It is now skipped **only** when the model is quantized and MTP is left out of it; plain distillation such as pruning recovery still trains the MTP head. `test_mtp_excluded_from_quantization` pins all four cases. - **One bad record deadlocked data prep.** `megatron_preprocess_data` re-raised chat-template failures out of a pool worker, stalling the whole job until it timed out — three malformed records cost a multi-hour tokenization run. They are now skipped with a warning, matching the existing handling of malformed JSONL a few lines above. Also exposes `--logit_kl_topk`, which `DistillationConfig` has supported for a while but the example never passed through; `test_qad` now exercises that path. §4, §5 and §6 are independent of §1–3; happy to split them out if reviewers prefer. ### Usage No API change. Exported names now match the released checkpoints: ``` model.language_model.layers.0.mlp.experts.<E>.{gate,up,down}_proj.{weight,weight_scale,weight_scale_2} lm_head.{weight,weight_scale,weight_scale_2} ``` ### Testing - `test_mcore_export_mappings.py` — qwen3_5 mappings emit per-expert rules. Verified these fail without the fix (2 failed / 11 passed), with `Qwen3MoeForCausalLM` / `NemotronHForCausalLM` as controls. - `test_unified_export_megatron.py` — the gate/up split, per-block scale slicing, the 0-dim scalar fallback, and both directions of the `_verify_exported_keys` relaxation. - `test_megatron.py::TestKeepGptOutputLayerExtraState` — 15 cases: payload detection, no-op second call, warn-and-skip on an unrecognised `sharded_state_dict`, and `test_patches_stock_megatron_core` which installs a replica of the real pre-fix upstream body (verified against `be08ce5b1~1`) so the patched path is exercised whichever megatron-core is installed. - `test_qad.py` — CI caught that its reference comparison still assumed packed experts; fixed. **End to end on `Qwen/Qwen3.6-35B-A3B` (35B MoE, 256 experts), 4×GB200, nemo:26.08:** | | before | after | | --- | --- | --- | | export self-check | `Export dropped 82 tensor(s)` | passes | | expert tensors | `mlp.experts.gate_up_proj` (packed) | `mlp.experts.<E>.{gate,up,down}_proj` | | **vLLM v0.28.0 load** | **`AttributeError`, engine never starts** | **`Loading weights took 25.61 s`** | | **NEL eval (GPQA-D, MMMU-Pro)** | **FAILED** | **SUCCESS** | ### Results these fixes unblocked The export fix is what made a Megatron-produced NVFP4 MoE checkpoint servable at all, so it enabled a full PTQ study on Qwen3.6-35B-A3B. Accuracy deltas are against a BF16 baseline measured on the same harness, from **paired** per-question tests: | recipe | throughput vs BF16 | GPQA-D | SciCode ×8 | MMMU-Pro | IFBench | | --- | --- | --- | --- | --- | --- | | **W4A16** (weight-only) | **0.64–0.86×** — *slower* | −0.06 | −0.15 | +0.48 | −0.44 | | **W4A4** | 8/12 shapes faster | −0.60 | −0.70 | −1.48 (p=0.019) | −0.53 | | **W4A4 + 4-bit `lm_head`** | **9/12 shapes**, up to **1.30×** | **+0.03** (p=0.96) | −0.81 | **−1.16** (p=0.016) | −1.65 (ns) | Repeats: GPQA-D is `pass@1[avg-of-16]`; SciCode is 8 pooled runs per recipe; MMMU-Pro is 3 runs per side and IFBench 2–3 for BF16 and the last row, 1 elsewhere. AA-LCR (68.33 → 71.33, p=0.25, 3 runs per side) and τ²-Telecom (94.25 → 94.25, 3 runs per side) are on par; at 100 questions and 114 tasks they cannot resolve below ~5 pp and ~3 pp, so they carry no claim either way. #### QAD status (what §6 unblocked) With the §6 fixes in place, QAD runs end to end on this model: 32 nodes, `TP=1 PP=1 CP=1 EP=8`, seq 32768, gbs 512, ~38 s/iter, 124 GB/GPU peak. First accuracy read, MMMU-Pro at iteration 50 (0.84 B tokens), 3 runs per side, paired per-question: | | MMMU-Pro | vs BF16 | | --- | --- | --- | | BF16 | 74.55 | — | | W4A4 + 4-bit `lm_head` (PTQ) | 73.39 | **−1.16, p=0.016** | | + QAD, iteration 50 | 73.78 | −0.77, p=0.089 (ns) | The PTQ deficit that motivated this work is no longer statistically significant after 50 QAD iterations. The improvement itself (+0.39 vs PTQ) is **not** significant at p=0.41, and 50 iterations is 10% of the planned budget, so this is a direction rather than a result. A full six-benchmark sweep at iterations 50 and 300 is running; these numbers will be superseded. Two findings worth flagging beyond this PR: - **Weight-only NVFP4 is slower than BF16 on Blackwell.** W4A16 leaves activations in BF16, so vLLM cannot use the FP4 tensor cores and falls back to `MarlinNvFp4LinearKernel` / `'MARLIN'` MoE. W4A4 selects `FLASHINFER_TRTLLM` + `FlashInferCuteDslNvFp4LinearKernel` and beats W4A16 in **12/12** shapes. The Marlin line count tracks the recipe exactly (one W4A16 layer ⇒ one Marlin line ⇒ zero once `lm_head` is W4A4). - **The only accuracy cost is multimodal**: **−1.2 pp on MMMU-Pro** for the fastest recipe, confirmed over 3 runs per side (p=0.016). GPQA-D, SciCode, IFBench, AA-LCR and τ²-Telecom show no significant regression. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ <!-- Megatron checkpoints unaffected; re-export to gain the loadable layout. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!-- 0.47.0 → Bug Fixes; includes the removed flag, since passing it now errors instead of being ignored --> - Did you get Claude approval on this PR?: ✅ <!-- Reviewed; all findings addressed, threads resolved. --> ### Additional Information Both export bugs were found while reproducing `nvidia/Qwen3.6-35B-A3B-NVFP4` through `examples/megatron_bridge/`. Follow-up to #2332. Upstream counterpart [NVIDIA/Megatron-LM#7086](https://github.com/NVIDIA/Megatron-LM/pull/7086) is closed — see §2. Labeled `cherry-pick-0.47.0`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2d35643452 |
LiLiCorr training (#2342)
### What does this PR do? Type of change: new feature Adds **LiLiCorr**, a candidate-lattice reranker for DFlash drafts, as a new `projector_type` on the existing `dflash` mode — plus three DFlash-wide improvements that apply to every variant, and an optional composition with DFlash2's grouped convolutions. A DFlash drafter is trained on per-position marginals rather than on the joint block distribution, so its drafted tokens are individually plausible yet jointly incoherent. LiLiCorr keeps the top-`k` candidates the backbone already produces at each block position, scores transitions between adjacent candidates with a small two-layer transformer, and commits a path through the lattice greedily. Serving is unchanged in kind: verify still checks every drafted token against the target, so the emitted distribution is untouched and only acceptance length moves. - Paper: [LiLiCorr: Lightweight Likelihood Correlation of Parallel Drafts for Speculative Decoding](https://arxiv.org/abs/2608.20530) (arXiv:2608.20530) - Blog: https://research.nvidia.com/labs/nemotron/lilicorr/ - **Companion PR — serving support:** [sgl-project/sglang#37462](https://github.com/sgl-project/sglang/pull/37462) This PR is the **training** half. It trains the drafters and exports them; the companion PR above is what serves the resulting checkpoints, and is what the comparison table below was measured through. **What is in the commits** | | | | --- | --- | | LiLiCorr draft variant | `hf_lilicorr.py`, `modeling_lilicorr.py`, conversion routing, config fields, export | | Three DFlash-wide features | fp32 master weights for the draft, draft activation checkpointing, and a DDP hang fix — all default-off or behaviour-preserving, all applying to `dflash`, `domino`, `dspark` and `dflash2` alike | | Optional grouped convolutions | composes LiLiCorr with DFlash2's `DFlashGroupedConv`; see the dependency note below | | Two recipes | `lilicorr.yaml` and `lilicorr_conv.yaml` | | CPU unit tests, CHANGELOG, one launcher example | | **⚠️ The convolutions depend on the DFlash2 branch, and cannot run until it merges.** `modeling_lilicorr.py` imports `DFlashGroupedConv` from `modeling_dflash2`, which today exists only on `haoguo/dflash2-support`. The class is **imported rather than copied on purpose** — it is the only way the two variants cannot drift apart arithmetically — but the consequence is that the convolutional recipe cannot run against `main` as it stands. So the import is **deferred into `_install_sublayer_convs`** rather than taken at module scope. Everything else in this PR, including the plain LiLiCorr reranker, has no DFlash2 dependency at all and works on `main` today; an eager import would have made the whole plugin unimportable for the sake of one optional feature. Requesting the convolutions without DFlash2 present raises an `ImportError` naming the two config keys to remove, rather than failing at import time. **This PR carries two of @h-guo18's commits, with authorship and sign-off preserved.** Both are independent of DFlash2 itself and both are needed here: - `1419d47e`, the no-op sublayer seam. Without it `DFlashDecoderLayer.forward` never calls the wrappers the convolutions install onto, so the modules would be built, counted and exported while computing nothing. It is arithmetically an identity on its own. - `ba377e7a`, the RoPE-θ fix. On Transformers 5 a config carries both a top-level `rope_theta` and a `rope_parameters` dict; the real base lives in the dict while the class default (10,000 for Qwen3) stays visible as the flat attribute. Reading the flat field first builds a draft whose RoPE base is 100× off a Qwen3-8B target's, which trains and exports without complaint. Both the training-side enforcement and the exporter's `_get_rope_theta` are affected on `main` today. Both are @h-guo18's work and belong to their branches; they are carried here only so that this PR stands on its own. **If those branches land first, this PR can be rebased onto them and the two commits dropped**, and they can equally be split out now if that is easier to review. The same applies to `dflash_fp32_master_weights`, which is also in flight on `haoguo/dflash-fp32-master-weights`. The field name is shared deliberately so that there is only ever one knob rather than two spellings of it, and both versions default to off. Whichever lands first, this PR can be rebased onto it. ### Usage Train with the shipped recipe: ```python from modelopt.recipe import load_recipe config = load_recipe("general/speculative_decoding/lilicorr.yaml") # Qwen3-8B target, 6 epochs, block size 16 (15 drafted slots, 16 verified), # DFlash decay objective at gamma 7.0, fp32 master weights for the draft. ``` Or convert directly: ```python import modelopt.torch.speculative as mtsp config = { "dflash_block_size": 16, "dflash_loss_objective": "decay", "dflash_loss_decay_factor": 7.0, "dflash_fp32_master_weights": True, "dflash_lilicorr_w_ce": 0.25, "dflash_lilicorr_w_margin": 0.0, "dflash_lilicorr_w_pen": 0.25, "dflash_architecture_config": { "num_hidden_layers": 5, "projector_type": "lilicorr", "lilicorr_candidate_topk": 8, # Optional, and all-or-nothing: adding these two keys wraps every draft # sublayer in DFlash2's grouped convolution. Requires the DFlash2 variant. # "conv_kernel_size": 2, # "conv_group_size": 16, }, } mtsp.convert(model, [("dflash", config)]) ``` ### Results Six drafters for a **Qwen3-8B** target, all trained **in ModelOpt on one matched contract** — the same corpus, schedule and block geometry for every arm, so no row carries a training advantage. Training data is NVIDIA's [Nemotron Post-Training Dataset v2](https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2) with the multilingual split excluded, generated from the target with **thinking disabled**; **6 epochs**; block size 16 (15 drafted slots, 16 verified); DFlash decay objective at gamma 7; **8 nodes × 8 H100, global batch size 64** (one sequence per device, no gradient accumulation). All six were then exported and served through SGLang on a **single H100 80GB**, `tp_size 1`, at concurrency 1, greedy, `fa3`, mean of two replicates, with the whole node held exclusive per benchmark. Speedup is output tokens/s against an autoregressive baseline measured in the same allocation. Cells are `acceptance length / speedup-vs-AR`; **★ fastest, ☆ second fastest**: | benchmark | LiLiCorr+conv | LiLiCorr | DSpark | DFlash2 | Domino | DFlash | |---|---|---|---|---|---|---| | gsm8k | ★ 7.715 / 5.26x | ☆ 7.557 / 5.22x | 7.375 / 4.86x | 7.252 / 5.06x | 7.225 / 4.87x | 6.341 / 4.59x | | math500 | ★ 9.241 / 6.54x | ☆ 9.064 / 6.52x | 9.012 / 6.15x | 8.999 / 6.49x | 8.976 / 6.25x | 7.909 / 5.88x | | aime25 | ★ 8.285 / 6.03x | ☆ 8.156 / 6.03x | 8.043 / 5.61x | 7.967 / 5.91x | 8.066 / 5.77x | 7.126 / 5.44x | | humaneval | ★ 7.393 / 4.01x | 7.077 / 3.93x | 7.163 / 3.72x | ☆ 7.081 / 3.95x | 6.864 / 3.73x | 6.156 / 3.68x | | mbpp_sanitized | ★ 5.999 / 4.18x | ☆ 5.849 / 4.13x | 5.888 / 3.95x | 5.685 / 4.05x | 5.679 / 3.91x | 5.027 / 3.70x | | livecodebench | ★ 7.975 / 5.40x | ☆ 7.754 / 5.33x | 7.775 / 5.10x | 7.601 / 5.26x | 7.553 / 5.04x | 6.808 / 4.88x | | alpaca_eval | ☆ 3.697 / 2.69x | ★ 3.656 / 2.70x | 3.588 / 2.52x | 3.467 / 2.58x | 3.627 / 2.59x | 3.222 / 2.46x | | mtbench | ★ 4.014 / 2.94x | ☆ 3.939 / 2.93x | 3.957 / 2.78x | 3.748 / 2.80x | 3.948 / 2.84x | 3.478 / 2.67x | **Against every other approach in the table, LiLiCorr with convolutions is the fastest on all eight benchmarks.** Plain LiLiCorr is the fastest on seven of the eight; the exception is humaneval, a 164-prompt slice, where DFlash2 is ahead by 0.5%. `DFlash` is the deliberately head-free control; every head clears it by +7.60% to +21.67% on acceptance, which is the check that a head actually loaded. Reproducing the `LiLiCorr+conv` column additionally needs the DFlash2 variant. Acceptance length is bit-reproducible under greedy decoding and its replicate spread here was 0.00% on every benchmark; throughput has a ~0.2% floor. ### What `dflash_fp32_master_weights` does, and what it is worth Today the draft is cast to the frozen base model's dtype — bf16 — before the optimizer is built. AdamW then allocates its moments with `zeros_like(p)`, so the **optimizer state becomes bf16 too**. That is the problem: bf16 has too few mantissa bits to represent the small updates Adam's second moment accumulates, so those updates round away and the effective step size decays on its own, independently of the learning-rate schedule. The flag is standard mixed precision instead: the draft's master weights stay in fp32 while the matmuls run in bf16. It requires a bf16 autocast around the forward, which HF `Trainer` supplies under `TrainingArguments.bf16`. Paths that do not go through the Trainer — evaluation, `pseudo_speculative_generate`, a plain `convert()` and forward — currently need the caller to supply it, and no shipped recipe exercises those (`estimate_ar: false`, `do_eval: false`). Making the draft supply its own autocast is a follow-up, held back from here on review because it touches every DFlash variant and wants e2e coverage of the existing recipes. Compute speed is unchanged. The cost is memory, about 12 bytes per parameter for the weight plus Adam's two moments instead of 6, plus a doubled gradient all-reduce under DDP, since fp32 parameters mean fp32 gradients. Under FSDP2 that second cost is what `MixedPrecisionPolicy(reduce_dtype=...)` exists to control. It is worth **7 to 14 percent of acceptance length**, measured at the end of training on gsm8k, and it helps every projector type: | arm | bf16 | fp32 | Δ acceptance length | | --- | ---: | ---: | ---: | | LiLiCorr | 6.8670 | 7.5573 | **+10.05%** | | DFlash2 | 6.7396 | 7.2518 | **+7.60%** | | Domino | 6.5854 | 7.2252 | **+9.71%** | | DSpark | 6.4621 | 7.3752 | **+14.13%** | | DFlash | 5.9030 | 6.3412 | **+7.42%** | Every arm in the comparison table above was trained with it on, and **both shipped recipes set it `true`**, so the documented path gets it. It defaults to **off**, so no existing DFlash, Domino or DSpark run changes behaviour. Both shipped LiLiCorr recipes set it `true`, which is the arithmetic their numbers were trained with. Flipping the default is a reasonable follow-up once the autocast above is in. The draft is drawn in fp32 and, under this flag, kept there; an unpromoted run rounds the same draw to the base model's dtype. So the bf16 and fp32 rows of the table above start from the same initialization at the precision each trains in, rather than from two different draws. A unit test pins that. The flag also survives a resume. `modify()` runs under `from_pretrained` with the base model still on meta and cannot place the draft at all, so `restore_draft_precision` re-applies the dtype, the device and the rotary buffer once the weights are loaded and before the Trainer builds the optimizer — the last point that can still decide the Adam moment dtype. It also reloads the draft's tensors at the dtype they were saved in, since checkpoints store the draft in fp32 while the base is bf16 and `dtype="auto"` gives every tensor one dtype. @h-guo18 has the same field in flight on `haoguo/dflash-fp32-master-weights`, plus an HF-format-resume fix this PR does not have. The name is shared deliberately so there is only ever one knob; whichever lands first, the other should be dropped rather than merged. ### Testing - **257 CPU unit tests pass** across `tests/unit/torch/speculative/`, including the existing DFlash, Domino, DSpark and Eagle suites. 48 of them are new and cover LiLiCorr specifically: conversion routing, head geometry, the required-field validation, the three-term objective and its absolute weights, gradient reach into both the head and the drafter body, and the export contract. - Both recipes load and validate through `modelopt.recipe.load_recipe`. - The three DFlash-wide changes are covered behaviourally: the fp32 flag is checked on the optimizer's moment dtypes rather than only on parameters, since the moments are the point of the change, and on the initialization described above; activation checkpointing is asserted to leave draft gradients bit-identical with the flag on and off; and the rotary buffer is asserted present after `modify()` on a real device while still deferred on meta, which is the case the laziness existed for. - The resume path has its own test: after a `save_pretrained` / `from_pretrained` round trip, `restore_draft_precision` is asserted to return the draft to fp32 with its stored weights intact and its Adam moments in fp32. Without it the draft comes back in the base dtype with the flag still set, which is the failure it exists to prevent. - `TestDFlashLazyRotaryEmb` was updated rather than left passing: it asserted the rotary buffer does *not* exist after convert, and the DDP fix deliberately changes that on non-meta devices. The replacement pins the refined invariant in both directions. - The published checkpoints were trained with this arithmetic, verified rather than assumed: a fingerprint over draft initialisation, loss and gradients is compared against the pre-review tree for both `dflash` and `lilicorr`. Loss and gradients are **bitwise identical**. Initialisation moves, by less than bf16 resolution, and that is the single-dtype change described above. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — every addition is opt-in. The new `projector_type` is selected only by config, `dflash_fp32_master_weights` defaults to off, and the activation-checkpointing and DDP fixes preserve behaviour. No existing default changes. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new dependencies. Four files carry `# Adapted from https://github.com/sgl-project/SpecForge/...` headers for the DFlash backbone and loss they derive from (Apache-2.0), matching the attribution already on `hf_dflash.py` in this repo. The two commits described above are @h-guo18's, cherry-picked with authorship and sign-off preserved. - Did you write any new necessary tests?: ✅ — 48 new CPU tests, plus the updated rotary test. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ — will run `/claude review` once opened. ### Additional Information The convolutional recipe is the memory worst case: at an 8B target, combined with fp32 master weights, it may need `training.gradient_checkpointing: true` to fit on 80 GiB, and it fits without at 4B. Checkpointing is mathematically neutral — same objective, same data order, same resulting model — but it trades step time for memory, so a run using it is not step-time-comparable with one that does not. The recipe header says so. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added LiLiCorr speculative decoding with candidate-lattice reranking, configurable objectives, metrics, export support, and optional grouped convolutions. * Added FP32 master-weight support with improved mixed-precision behavior and gradient checkpointing. * Added LiLiCorr training recipes and a Qwen3-8B launcher configuration. * **Bug Fixes** * Improved rotary-embedding configuration handling and corrected DFlash distributed-training hangs. * Added validation for invalid LiLiCorr configurations and improved exported reranking metadata. * **Documentation** * Expanded guidance for FP32 master weights, training workflows, and LiLiCorr configuration. * **Tests** * Expanded coverage across training, evaluation, generation, export, and checkpoint workflows. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: mrusanovsky <mrusanovsky@nvidia.com> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
acdf330414 |
Add TensorRT-RTX ABI EP support for ONNX quantization (#2262)
### What does this PR do? Type of change: new feature Adds opt-in support for using the standalone TensorRT-RTX ABI Execution Provider during ModelOpt ONNX quantization. Users select the ABI backend with: `--calibration_eps=NvTensorRtRtx --trt_rtx_backend=abi` When selected, ModelOpt imports and registers the installed TensorRT-RTX ABI provider before creating the ONNX Runtime inference session. The backend selection is propagated through INT8, FP8, and INT4 AWQ calibration paths, including the Windows GenAI LLM quantization example. The existing `--calibration_eps=NvTensorRtRtx` behavior remains backward compatible. The `legacy` backend is still the default and continues to use TensorRT-RTX libraries supplied through `PATH`. For Windows x64 with Python 3.11 or newer, the ONNX dependencies now include: - `onnxruntime-gpu~=1.26.0` - `onnxruntime-ep-nv-tensorrt-rtx-cu13==0.4.0` Keeping `onnxruntime-gpu` allows users to select either CUDA EP or TensorRT-RTX ABI EP for calibration. Windows-on-Arm source-build instructions are intentionally out of scope and will be documented separately. ### Usage ```powershell python -m modelopt.onnx.quantization ` --onnx_path="C:\path\to\Llama-3.2-3B-Instruct\model.onnx" ` --model_id="C:\path\to\Llama-3.2-3B-Instruct\config.json" ` --quantize_mode=int8 ` --output_path="C:\path\to\int8_abi\model.onnx" ` --calibration_eps=NvTensorRtRtx ` --trt_rtx_backend=abi ` --use_external_data_format ` --high_precision_dtype=fp32 ` --log_level=INFO ### Testing unit test have been added ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: pending <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added optional TensorRT-RTX ABI backend support for ONNX calibration on Windows ARM64. - Added `legacy` and `abi` backend selection to quantization APIs and command-line tools; `legacy` remains the default. - Added validation for unsupported backends and incompatible TensorRT plugin configurations. - Updated Windows ARM64 installation support and platform-specific package configuration. - **Documentation** - Updated Windows installation guidance, Python compatibility requirements, ARM64 setup, and verification steps. - Documented the new TensorRT-RTX backend command-line option. - **Tests** - Added coverage for ABI provider registration, backend validation, and compatibility checks. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Haoxi Zhang <haoxiz@nvidia.com> |
||
|
|
0688761ce9 |
feat(export): support multimodal and MTP models in layerwise export (#2303)
### What does this PR do? Type of change: New feature **Layerwise export now supports multimodal and MTP models.** Both were refused outright, and both were refused for the same reason: `finalize()` was called from inside `layerwise_calibrate`, which is the wrong scope for it. **1. Calibration does not know which model the checkpoint describes.** It only sees the module it was handed. A VLM calibrates its *language model*, so the shards, the exclusions and `config.json` all came out describing that submodel rather than the whole VLM. Moving the call out lets the caller root the exporter at the parent — and without the key prefixing, tower collection or ambient parent handle an earlier attempt needed, because the decoder layers are the same objects from either root. **2. Calibration runs before things the export needs exist.** Orphaned MTP weights are loaded *after* calibration, by which point every shard had already been written, so they could not be passed at all. After the move they are an ordinary argument to `finalize()`, with no staging attribute stashed on the model. ### How it works The exporter is created by whoever owns the export and **announced on the model** that `mtq.quantize` is given. Calibration picks it up, binds it, and drives it per layer; the export that follows reads it back and finishes the checkpoint: ```python LayerwiseExporter(full_model, export_path).announce(language_model) mtq.quantize(language_model, quant_cfg, forward_loop=loop) ... getattr(full_model, LAYERWISE_EXPORTER_ATTR).finalize(extra_state_dict=mtp_state_dict) ``` Calibration and export are handed *different* models, so `announce()` publishes the exporter on each end separately: the caller announces on the model being calibrated, and `bind()` announces on the export root. Neither side has to know where the other looked, and the lookup stays an O(1) `getattr` rather than a `named_modules()` scan — worth avoiding at roughly 1.65 µs/module, or ~500 ms on a Kimi-K3-sized model. For a non-VLM both roots are the same object and the second announcement is a no-op. `finalize()` clears every attachment it recorded, so the module graph does not retain a live exporter afterwards. `mtq.quantize` and `mtq.calibrate` are **unchanged** — a layerwise-only feature does not belong in the public quantization API. The attribute follows `_mtp_layer_prefixes`, which crosses the same calibration→export boundary the same way (`hf_ptq.py:538` sets it, `unified_export_hf.py:870` reads it back). Construction is inert: `__init__` records only the export root and the directory, because the caller builds it before `mtq.quantize`, when there are no quantizers yet to validate or read a config from. `bind()` does that, called from calibration after quantizer insertion and before any layer is converted — the only window where both hold, and the same instant the exporter used to be constructed, so unsupported models still fail in seconds rather than hours. Only the calibration pass that sets `export_dir` drives the exporter: a list-form algorithm runs one pass per entry, and an earlier one must not convert layers a later one still has to calibrate. ### Usage Nothing changes for a plain layerwise-export recipe: `layerwise.export_dir` still drives it. Pre-attaching an exporter is the opt-in for the two cases that need it — a checkpoint whose root is wider than the calibrated model, and orphaned tensors to merge at the end. The one behaviour change for a config-only caller is that `mtq.quantize` now writes the layer shards but no longer finishes the checkpoint. Both exit paths warn with what is still owed, and `LayerwiseConfig.export_dir`'s description has been corrected — it previously promised "a complete, loadable checkpoint when the last layer lands" and still listed multimodal and MTP as raising `NotImplementedError`. ### Testing `tests/gpu/torch/export/test_layerwise_export.py` — **29 passed**. Beyond the 24 inherited from #2136, five new ones, each with a negative control confirming it fails without its fix: - orphaned MTP tensors reach the tail shard *and* the index - an exporter rooted at the parent widens the checkpoint's namespace - the config-only path announces an exporter that can be finished, and finalize clears it - only the pass that sets `export_dir` drives the exporter - an exporter whose root holds a different number of layers is refused at `bind()` Full suites: `tests/gpu/torch/export` + `tests/gpu/torch/quantization` **1012 passed / 55 skipped**, `tests/unit` **3318 passed / 15 skipped**, pre-commit clean. Both suites also report failures in `test_implicit_gemm.py` (FP4 conv kernels), `test_triton_fa_p_qdq.py`, `test_autocast_quantize_int8` and `test_engine_builder.py` collection; all reproduce unchanged on `main` and none touch the paths in this diff. Measured against the whole-model exporter on a tiny Gemma3-VL, towers prepared exactly as `hf_ptq` does: ``` keys: baseline=80 layerwise=80 only-baseline=[] only-layerwise=[] differing values: 0 vision tower present: True VLM namespace: True config.json is the VLM: True hf_quant_config match: True exclude_modules: ['language_model.lm_head', 'vision_tower.vision_model*'] (both sides) ``` #### End-to-end through `hf_ptq.py` Same FP8 recipe both sides; the baseline drops `layerwise.export_dir` and is exported by `main`, so the diff isolates this PR. Every tensor matches in key, dtype, shape and value, and `config.json` / `hf_quant_config.json` match too. | Model | Covers | Keys | Differing | |---|---|---|---| | Qwen3-VL-8B-Instruct | multimodal | 1254 = 1254 | 0 | | GLM-4.7-Flash | MoE + MTP | 28119 = 28119 | 0 | The VLM checkpoint keeps the vision tower unquantized (351 `model.visual.*` keys, no `weight_scale` among them) while the language model is FP8. The MTP run reports 212 orphaned tensors; all 212 land in `model-tail.safetensors` and in the index, with `model.layers.47*` in `exclude_modules`. **Not yet validated:** an accelerate-offloaded run, and a serving canary on the exported checkpoints. ### Refusals `export_dir` without `enable`, and an exporting algorithm entry with no calibration method, are both refused before calibration starts — neither reaches the per-layer pass, so both would otherwise export nothing. The early gate is a heuristic on the recipe, so `hf_ptq` also raises a plain `RuntimeError` at export time if calibration turned out not to have run; that backstop, not the gate, is what makes the failure legible on paths the recipe check cannot predict. `bind()` requires the layers calibration will drive and refuses a root that discovers a different number of them. Only the count is checked here: `export_layer` already rejects a reordering or a substituted module on its first call, and a length difference is the one mismatch it structurally cannot catch — every call would pass and `_write_index` would then open a shard that was never written, at the very end of the run. Orphan tensors are merged into the tail with no collision check, matching the whole-model path (`unified_export_hf.py:1623`). `load_mtp_weights` returns exactly the keys absent from `model.state_dict()`, so a collision with an exported tensor is not reachable through the only producer, and a guard would only make the two export paths diverge. ### Why not reuse `export_hf_checkpoint` It was the first idea and it is the most expensive one. Its transformers path is whole-model at every step — `_prepare_moe_inputs`, `requantize_resmooth_fused_llm_layers` (which runs a dummy forward that would fail on already-converted layers), `_process_quantized_modules`, a full `model.state_dict()` in host RAM, then `save_pretrained` rewriting shards already on disk — and it raises outright under `has_accelerate_offload`. `save_pretrained(state_dict={})` is not an escape either: safetensors' shared-storage check fires on MoE even with an empty dict. The natural consolidation target is the **streaming** exporter, which is already most of `finalize()`: 122 lines vs 74, sharing `decoder_owned_ids`, `enable_weight_access_and_writeback`, `_dispatch_export_handler`, `_reconstruct_fused_moe_linear`, `_add_mtp_exclusions`, `_postprocess_single_tensor`, `requires_weight_materialization` and `save_non_weight_artifacts`. Folding them together needs roughly four knobs: skip the whole-model prep, skip layers already written, seed the index with the existing shards, and inject the quant config. That is a separate change and deliberately not in this one. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `mtq.quantize`/`mtq.calibrate` signatures are unchanged, and a recipe that only sets `layerwise.export_dir` behaves as before. The one behaviour change is that `mtq.quantize` no longer finishes the checkpoint on its own: callers must now call `finalize()` on the exporter, which calibration leaves on the model. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ❌ — pending. - Did you get Claude approval on this PR?: ❌ — the last review's findings are all addressed; needs a re-run. ### Additional Information Follow-ups this enables: #2259 (MTP) reduces to close to nothing, and the multimodal work in #2218 no longer needs `export_parent`, the key prefixing, or the tower collection. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Layerwise export now supports clearer control over export locations and calibrated layer handling. - Export workflows provide improved support for resuming, sharded checkpoints, mixture-of-experts models, and nested model namespaces. - **Bug Fixes** - Improved handling of exported checkpoint shards and extra tensors. - Added clearer warnings when exports require completion before loading. - **Documentation** - Clarified that layerwise exports write shards during calibration and require an explicit finalization step. - Documented that the in-memory model is not suitable for inference after layerwise export. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
19de0075cb |
Forward kv_cache_free_gpu_memory_fraction to the lm_eval TensorRT-LLM engine (NVBug 6701763) (#2300)
### What does this PR do? Type of change: Bug fix `scripts/huggingface_example.sh --kv_cache_free_gpu_memory_fraction` has no effect on the `lm_eval` task: the value is parsed by `parser.sh`, printed, and then dropped. lm-eval's built-in `trtllm` backend (`lm_eval.models.trtllm_causallms.TRTLLM.__init__`, which this example switched to in #2066) accepts `**kwargs`, but builds `KvCacheConfig(enable_block_reuse=False)` and passes `LLM(...)` a fixed set of keys — `kwargs` is never merged in. So an extra `--model_args` entry is accepted by the CLI and silently discarded, and the KV cache is sized from TensorRT-LLM's default `free_gpu_memory_fraction=0.9`. There is no way to fix this from the caller: `--model_args` only yields scalars, so a `KvCacheConfig` object cannot be passed in either. On a GH200 that means ~119.6 GiB of KV cache (`119.55 / 0.9 ≈ 132.8 GiB free`), leaving 87.8 MiB free, and `prompt_logprobs` deserialization then OOMs asking for 2.82 GiB. `examples/llm_eval/lm_eval_trtllm.py` already exists to patch this backend (its `_parse_logprobs` misaligns TensorRT-LLM's `prompt_logprobs` by one). It now also injects the fraction into the `KvCacheConfig` the backend builds, defaulting to 0.8 — the same default `parser.sh` declares, and below TensorRT-LLM's 0.9. `huggingface_example.sh` passes the parsed value through in `--model_args`. Scoped deliberately to the `lm_eval` path: the `quant` smoke test and `mmlu` go through `modelopt.deploy.llm.LLM` (0.7, hardcoded) and `simple_eval`/`livecodebench` through `trtllm-serve` (0.9); those are left as they are. ### Usage ```bash # Via the example script (parser.sh default 0.8) scripts/huggingface_example.sh --model $HF_PATH --quant fp8 --tp 1 \ --tasks quant,lm_eval --lm_eval_tasks mmlu --lm_eval_limit 50 \ --kv_cache_free_gpu_memory_fraction 0.5 ``` ```bash # Standalone, via lm-eval's --model_args python lm_eval_trtllm.py --model trtllm \ --model_args model=<ckpt>,tokenizer=<tok>,max_input_len=4096,kv_cache_free_gpu_memory_fraction=0.5 \ --tasks mmlu --batch_size 8 ``` ### Testing - `pytest tests/examples/llm_eval/test_lm_eval_trtllm.py` — 21 passed (lm-eval 0.4.12, no GPU). - The new tests instantiate the **real** upstream `TRTLLM.__init__` through `create_from_arg_obj`, with `tensorrt_llm` and the tokenizer stubbed, and assert the engine receives `KvCacheConfig(enable_block_reuse=False, free_gpu_memory_fraction=0.5)`; that an unset key still yields 0.8 rather than 0.9; and that the patch does not outlive the constructor. Reverting the fix fails 3 of them. - Tripwire test asserts upstream still neither declares nor forwards the argument, so this shim gets deleted rather than silently kept once lm-eval fixes it. - `pre-commit run --files <changed>` clean (ruff, mypy, bandit, markdownlint); `bash -n` on the modified script. - Not run: the GPU end-to-end `tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8`, which exercises `lm_eval` through the modified script — no GPU in this environment. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — the `lm_eval` KV cache goes from TensorRT-LLM's 0.9 to 0.8, which is strictly more conservative; `parser.sh`'s declared default is unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information NVBug 6701763. The 0.9 default on this path arrived with #2066 and was documented as a known limitation in `examples/llm_eval/README.md` ("the KV cache uses 90% of free GPU memory rather than 70%"); that note is replaced by the working knob. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Fixed the TensorRT-LLM evaluation workflow so `kv_cache_free_gpu_memory_fraction` is correctly passed to the backend. - The setting now defaults to `0.8`, providing more predictable GPU memory allocation for KV-cache usage. - **Documentation** - Updated the TensorRT-LLM evaluation example and usage guidance to describe the KV-cache memory setting and its default behavior. - Updated the Hugging Face example to pass the configured KV-cache memory fraction. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2aefe08f20 |
[OMNIML-5613] Quantize ResNet residual adds in torch ONNX example (#2024)
### What does this PR do?
Type of change: Bug fix
Adds recipe-backed FP8 and INT8 residual quantization for timm ResNet
models in the torch ONNX example:
- Adds FP8 and INT8 PTQ recipes that enable a shortcut quantizer
immediately before each residual `Add`.
- Inserts shortcut quantizers after ModelOpt module conversion so
recipes configure and calibrate them in the normal quantization pass.
- Adds `--recipe` support for PTQ and AutoQuantize recipes and renames
`--quantize_mode` to `--qformat`.
- Verifies all 16 ResNet-50 residual additions have shortcut Q/DQ
immediately before the `Add`.
### ResNet support scope
ResNet and other convolutional architectures are supported only with FP8
and INT8. AutoQuantize, MXFP8, NVFP4, and INT4_AWQ are not supported for
ResNet because TensorRT has limited convolution kernel support.
Transformer architectures containing individual Conv2d layers continue
to use format-specific Conv overrides.
### Usage
```bash
python examples/torch_onnx/torch_quant_to_onnx.py \
--timm_model_name=resnet50 \
--recipe=timm/resnet/ptq/fp8 \
--onnx_save_path=resnet50.onnx
```
Use `timm/resnet/ptq/int8` for INT8. Without `--recipe`, `--qformat`
selects a built-in quantization preset.
### Testing
- All configured pre-commit hooks passed, including recipe schema and
license validation.
- Focused AutoQuantize recipe mapping regression passed.
- FP8/INT8 recipe export coverage verifies all 16 ResNet-50 shortcut
Q/DQ pairs.
- TensorRT engine builds passed for FP8 and INT8 on Ada.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update the changelog?: ✅
> 🤖 _Generated by Codex (AI agent)._
---------
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <noreply@openai.com>
|
||
|
|
58eafdf172 |
Fix the llm_eval README commands that no longer run as written (#2358)
### What does this PR do?
Type of change: Documentation (plus one small example-script fix)
An audit of `examples/llm_eval/README.md` against the current scripts
(nvbug 6701343) found several documented commands that no longer run as
written:
- **T5 / seq2seq.** `--model hf-seq2seq` is not a registered lm-eval
backend in any version this example supports — the string does not
appear in the 0.4.12 or 0.4.13 wheels, so the command fails at model
lookup. `HFLM` detects encoder-decoder models from `config.json`, so the
example now uses `--model hf` and mentions `backend=seq2seq` as the
override for checkpoints lm-eval cannot classify. No ModelOpt-side
change was needed: encoder-decoder calibration already works (verified
below).
- **auto_quantize format list.** `FP8_DEFAULT_CFG|NVFP4_DEFAULT_CFG` was
shown as a literal value in both README locations, but each
comma-separated entry is resolved with `getattr(mtq, ...)` and that name
does not exist. Now shows a valid list, spells out the choices, and
names the placeholder consistently with the surrounding block.
- **`vllm serve`.** A missing line continuation meant `--port` ran as a
separate shell command.
- **MMLU setup.** Dropped a stray `cd ..` left over from the 0.11
examples release. It leaves `examples/llm_eval`, where both `mmlu.py`
and its default `--data_dir data/mmlu` live;
`hf_ptq/scripts/huggingface_example.sh` correctly stays put throughout
its MMLU flow, so the README was the only thing out of step.
- **`run_simple_eval.sh`.** Documented the optional fifth argument
(`--examples`), which `huggingface_example.sh` already passes as
`$SIMPLE_EVAL_LIMIT`.
Two changes beyond the docs:
- **`quantization_utils.py`:** under `auto_quantize`, a `quant_cfg`
string was iterated character by character, so a single format failed
with the baffling `AttributeError: module 'modelopt.torch.quantization'
has no attribute 'F'`. Normalized `str -> list` at the point the list is
consumed, which covers both `mmlu.py` and `lm_eval_hf.py` rather than
one caller. This also honors the existing `str | list[str]` annotation.
- **`requirements.txt`:** added the missing `openai`. `modeling.py`
imports it unconditionally and `lm_eval[api]` supplies only `tiktoken`,
so every documented `mmlu.py` command died with `ModuleNotFoundError` on
a clean install of the stated requirements.
Note on the filed report: its item 3 claimed `mmlu.py` fails to split
the comma-separated config list. That does not reproduce — `mmlu.py`
uses `fire`, which already parses `A,B,NONE` into a tuple, and the
unmodified script completes `auto_quantize` fine. Applying the suggested
`quant_cfg.split(",")` would have *broken* the documented command with
`AttributeError: 'tuple' object has no attribute 'split'`. The
`quantization_utils.py` change above addresses the real adjacent defect
instead. Pushback recorded on the bug.
### Usage
No new API or flag. The corrected commands:
```bash
# T5 / encoder-decoder (was: --model hf-seq2seq, which does not exist)
python lm_eval_hf.py --model hf --model_args pretrained=t5-small \
--quant_cfg FP8_DEFAULT_CFG --tasks <comma separated tasks> --batch_size 4
# auto_quantize search list (was: W4A8_AWQ_BETA_CFG,FP8_DEFAULT_CFG|NVFP4_DEFAULT_CFG,NONE)
python mmlu.py --model_name causal --model_path <model> \
--quant_cfg W4A8_AWQ_BETA_CFG,FP8_DEFAULT_CFG,NONE --auto_quantize_bits 4.8 --batch_size 4
# simple evals, optional 5th arg
bash run_simple_eval.sh <model> <evals> <max_tokens> <port> [num examples per eval]
```
### Testing
Ran on 2x RTX 6000 Ada with a tiny Qwen3 and a locally synthesized MMLU
tree (no download):
- **`mmlu.py --auto_quantize_bits` with the documented comma-separated
list** — completes quantization on both the unpatched and patched
script, confirming the reported item 3 is a false positive. Probed
`fire` directly: bare, quoted and `--flag=value` forms all yield
`('W4A8_AWQ_BETA_CFG', 'FP8_DEFAULT_CFG', 'NONE')`.
- **`mmlu.py --auto_quantize_bits` with a single format** — proved the
new guard fires by reverting it: without the change the run dies with
`AttributeError: module 'modelopt.torch.quantization' has no attribute
'F'`; with it, the run reaches a legitimate domain assertion
(`effective_bits 4.8` cannot be below FP8's 8 bits).
- **Encoder-decoder calibration** — quantized a T5 with
`FP8_DEFAULT_CFG` through `quantize_model` and confirmed encoder,
decoder and cross-attention (`EncDecAttention`) layers all calibrate
with real amax values. This is what settled keeping the T5 example
rather than deleting it.
- **`vllm serve` snippet** — parsed the fixed block with `bash`;
`--quantization`, `--port` and `--tensor-parallel-size` now all belong
to one command.
- **`run_simple_eval.sh`** — confirmed the 4-arg form is unchanged and
the 5-arg form emits `--examples 16`.
- **Lint** — `ruff-check`, `ruff-format`, `markdownlint-cli2`, `typos`,
`bandit`, `mypy`, `requirements-txt-fixer`, `mixed-line-ending` all
pass.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — added
`openai` to `examples/llm_eval/requirements.txt`; it is Apache 2.0
(permissive), so no codeowners exception is needed. It is not a new
runtime dependency of the library, and `run_simple_eval.sh` already `pip
install`s it.
- Did you write any new necessary tests?: N/A — docs plus a two-line
defensive normalization in an example util. `mmlu.py` cannot be imported
without `openai`/`rwkv`/`tiktoken`, so a hermetic unit test would need
more stub scaffolding than the line it guards; verified by direct
execution instead, as above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — examples-only documentation cleanup, not a feature, breaking
change, deprecation, or a critical bug from a previous release.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
Fixes nvbug 6701343 / OMNIML-5806. Item 3 of the filed report is a false
positive; pushback and evidence are recorded in a comment on the bug.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Auto-quantization now supports comma-separated format configurations.
- Added an optional example-limit setting for Simple Evals.
- Added OpenAI support for LLM evaluation examples.
- **Documentation**
- Clarified encoder-decoder model usage with `lm_eval`.
- Added instructions for running MMLU from the evaluation examples
directory.
- Corrected the vLLM command formatting.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
5c123ce183 |
[OMNIML-5563] Add PETR ONNX PTQ and accuracy evaluation example (#2180)
### What does this PR do? Type of change: new example, example simplification, and backward-breaking example migration Adds end-to-end PETRv1/PETRv2 ONNX PTQ and reduces PETR/FAR3D to one shared workflow: - quantizes the shared VoVNet image backbone/encoder to INT8 or FP8; - runs both the selected historical and current PETRv2 six-camera sweeps through the same precision-matched TensorRT backbone engine using distinct execution contexts during accuracy evaluation; - keeps the PETR head and FAR3D decoder in their exported mixed FP16/FP32 precision; - reuses one NPZ calibration format, VoVNet exclusion helper, quantization entry point, and TensorRT runner; - does not change generic Model Optimizer calibration behavior or its public CLI. ### Container boundary Both examples use two targets from one Dockerfile, with no virtual environments: - `evaluator`: a digest-pinned `nvcr.io/nvidia/pytorch:22.06-py3` base with the legacy PyTorch 1.13.1/OpenMMLab stack for source setup, metadata generation, ONNX export, direct PyTorch calibration capture, and final accuracy evaluation; - `modelopt`: a digest-pinned `nvcr.io/nvidia/pytorch:26.07-py3` base for Model Optimizer, ONNX Runtime CUDA, AutoCast, INT8/FP8 quantization, and TensorRT engine builds. Both targets use TensorRT `11.1.0.106`. Engines are built and evaluated on the same GPU architecture. Final metrics remain in the evaluator because they import the legacy model-framework postprocessing and dataset code; only artifacts cross the container boundary through the shared workspace. PETR is used without patches. FAR3D applies only the official `patch/far3d.patch` from the pinned NVIDIA DL4AGX revision. This PR carries no patch files. ### Evaluator dependencies The dependencies intentionally installed without transitive dependencies are listed in `requirements-evaluator-nodeps.txt`. Their pins rely on runtime packages supplied by the digest-pinned PyTorch 22.06 evaluator base. `lyft-dataset-sdk` is required only by mmdet3d's eager dataset import; neither PETR nor FAR3D uses Lyft data. `flash-attn` remains in the main evaluator requirements because its compiled installation uses the evaluator build step rather than the intentionally dependency-free legacy package step. Fresh setup and dependency approval is requested for the final reduced dependency set. ### Reproducible PETR metadata The documented workflow mounts raw nuScenes read-only and creates a writable dataset view using symlinks. It then runs the pinned mmdetection3d converter and a temporary, untracked copy of PETR's pinned sweep generator configured only for the validation prefix and writable dataset root. A clean run generated both metadata files with 6,019 validation records. The referenced camera, lidar, and sweep paths are absolute and resolvable through the writable dataset view. ### Example-local utilities The per-batch NPZ streaming and TensorRT runtime utilities remain example-local because they execute in the legacy evaluator, where Model Optimizer is not installed. The core `CalibrationDataProvider` consumes one in-memory mapping of stacked arrays and does not provide this streamed per-file workflow. ### Validation - Focused CPU tests: 10 passed. - Broader ONNX quantization CPU tests: 326 passed. - All applicable pre-commit and documentation checks, plus `git diff --check`, passed. - Rebuilt both Docker targets and verified their exact dependency versions, imports, TensorRT `11.1.0.106`, GPU runtime initialization, and absence of virtual environments. - Generated both PETR metadata files from a clean writable dataset view and verified 6,019 validation records plus resolvable data paths. - PETRv1 passed a one-sample TensorRT regression smoke. - PETRv2 passed FP16, INT8, and FP8 TensorRT smokes and full 6,019-sample validation. Both the selected historical and current sweeps are computed by the matching backbone engine; accuracy evaluation no longer extracts image features with PyTorch. - FAR3D passed a recurrent two-frame TensorRT smoke covering plugin loading and recurrent state. TensorRT `11.1.0.106` mAP follows. PETRv2 was remeasured after correcting its temporal feature path; the PETRv1 and FAR3D numerical paths are unchanged. | Pipeline | FP16 | INT8 | FP8 | | --- | ---: | ---: | ---: | | PETRv1: 1 backbone pass + fixed typed mixed FP16/FP32 head | 0.3778 | 0.3707 | 0.3756 | | PETRv2: 2 serial backbone passes + fixed typed mixed FP16/FP32 head | 0.4102 | 0.3982 | 0.4084 | | FAR3D: 1 encoder pass + fixed mixed FP16/FP32 decoder | 0.241 | 0.235 | 0.239 | Normalized engine-only performance improvement over each matching FP16 pipeline: | Pipeline | INT8 speedup | FP8 speedup | | --- | ---: | ---: | | PETRv1 | 1.49x | 1.29x | | PETRv2 | 1.51x | 1.30x | | FAR3D | 1.69x | 1.40x | Performance was measured with TensorRT `11.1.0.106` on an NVIDIA RTX 6000 Ada Generation GPU using five interleaved trials per engine component. Each component uses the median `trtexec`-reported GPU Compute Time with data transfers disabled and CUDA Graphs enabled. Component times are summed before normalization: PETRv1 uses one backbone pass plus its fixed head, PETRv2 uses two serial backbone passes plus its fixed head with no temporal cache assumed, and FAR3D uses one encoder pass plus its fixed decoder. Absolute latency values are intentionally not published. Adapted files retain exact public-source references and upstream notices, and the top-level license attribution is updated. - Is this change backward compatible?: ❌ - Did you write the necessary tests?: ✅ - Did you update the changelog?: ✅ > 🤖 _Generated by Codex (AI agent)._ --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com> |
||
|
|
4773f72f8a |
Docs: Add WOA documentation (#2264)
### What does this PR do? Add WoA env setup guide. Includes build instruction of pyarrow, which used by datatsets ### Usage N/A ### Testing N/A ### Before your PR is "*Ready for review*" - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Documentation** - Added comprehensive Windows on Arm installation guidance, including prerequisites, environment setup, dependency installation, verification, and troubleshooting. - Documented experimental Windows ARM64 support, supported quantization formats, native dependency requirements, and Support Matrix details. - Expanded supported Windows Python versions through 3.13. - Expanded TensorRT-RTX guidance for calibration, deployment, provider setup, and standalone plugin usage. - Clarified PyArrow requirements and local build instructions for Windows ARM64. - Added links to dedicated Windows on Arm installation resources and shared TensorRT-RTX documentation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Haoxi Zhang <haoxiz@nvidia.com> |
||
|
|
61757c9781 |
Support quantized Qwen3-VL / Qwen3.5-VL (dense + MoE) export from Megatron-Bridge and verify exported checkpoints (#2276)
### What does this PR do?
Type of change: Bug fix + new feature
Enables quantized **Qwen3-VL** and **Qwen3.5-VL** (dense and MoE) →
unified HuggingFace export from Megatron-Bridge, and fixes the bugs
found along the way (ten from testing, plus a further round from
review). Most of them produced a valid-looking checkpoint and a green
test run, so the PR also makes the export path verify its own output.
Review is easiest commit-by-commit — each of the eleven commits is
self-contained and independently green.
#### Two blockers
1. **The exporter rejected the Megatron-Bridge VLM wrapper.**
`GPTModelExporter` only unwrapped MCore's `LLaVAModel`, so
`Qwen3VLModel` raised `ValueError: Input to GPTModelExport must be a
megatron.core.models.GPTModel!`. It now unwraps any wrapper exposing
`.language_model`.
2. **A VLM QAD checkpoint couldn't be loaded back.** `distill.py` passes
`distill_submodule="language_model"`, so the checkpoint holds only the
language model and the load died on `KeyError:
vision_model.patch_embed.proj.weight`. The loader now reads the
checkpoint metadata and targets `.language_model` when there are no
vision weights.
#### Four silent-corruption bugs
3. **VLM QAD discarded all ModelOpt state** (shipped in 0.46).
`ModeloptStateManager` requires state on the **root** of whatever gets
checkpointed. `quantize.py` quantizes the VLM root, so PTQ anchors it
there — but QAD checkpoints only `language_model`, orphaning it. The
saved `modelopt_state_dict` was literally `[]`; the `*_quantizer._amax`
tensors were still present but got dropped on load
(`dist_ckpt_strictness="assume_ok_unexpected"`), and the export came out
plain BF16 with no `hf_quant_config.json`.
4. **Fused grouped-GEMM MoE experts were omitted entirely.** The MoE
dispatch had no `else`, so an architecture without an
`experts.linear_fc1` rule exported *zero routed experts*. This hit
**`Qwen3MoeForCausalLM`** — a registered, supported architecture with no
export test — not just VLMs. A tiny Qwen3-MoE exported 37 of 45 tensors,
exit 0, no warning.
5. **Qwen3.5's GatedDeltaNet output norm was off by exactly 1.0.**
Megatron stores that gamma zero-centered, HF centers it on 1. Correct
names, correct shapes, wrong values — invisible to any structural check.
Megatron-Bridge's importer confirms the convention
(`RMSNorm2ZeroCenteredRMSNormMapping`).
6. **The disabled-quantizer patterns silently no-op on Megatron paths.**
They are written against HuggingFace module names. `*mixer.conv1d*`
matches only because MCore and HF happen to agree on "mixer" for Mamba;
`*linear_attn.conv1d*` never matched (Megatron calls it
`self_attention.conv1d`), so the conv1d was calibrated.
`*linear_attn.in_proj_a/b*` **cannot** match at all — Megatron fuses all
six GDN sections behind one quantizer — so the alpha/beta gates the
recipe wants in BF16 were exported in FP8.
#### Four more bugs, found only by running real checkpoints
The tiny fixtures could not reach these; each came from a real model or
a real quant format.
7. **Routed experts were written in a layout no real Qwen3.5 checkpoint
uses.** Real Qwen3.5 stores experts packed as `[num_experts, out, in]`;
the mapping emitted per-expert names, so every routed expert was
dropped. The fixture actively hid this: transformers *unpacks* experts
on `save_pretrained`, so the saved reference agreed with the wrong
output. Fixed with a `transpose` kwarg on `_pack_name_remapping` plus a
`GroupedMLPPacking` rule, so fused `TEGroupedMLP` reaches the same
packed tensors — which is also what lets Qwen3.5 keep grouped GEMM
(**22.1 GB/GPU vs 38.9 GB/GPU** on a 20-layer, 256-expert model).
8. **`_grouped_mlp_packing` was broken for NVFP4.** It max-merged
`weight_scale`, but NVFP4 needs each expert's per-block scales *stacked*
with only the global `weight_scale_2` merged; it also dequantized packed
`uint8` against per-block scales, and passed `block_size=None`.
`weight_scale_2` is never populated in an FP8 run, so the whole branch
was dead code under FP8-only testing. `_grouped_mlp_slicing` gained
`quantize=False` so packing can quantize once over the stack, matching
`_pack_name_remapping`.
9. **`_mtp_prefix` corrupted every VLM's MTP tensor names.** It did
`prefix.replace("model", "mtp")` uncounted, so
`model.language_model.layers.{}` became `mtp.language_mtp.layers.0.*` —
tensors present and correctly valued, under names nothing loads.
LLM-only prefixes contain one occurrence, so this was invisible until a
VLM with MTP was exported.
10. **`load_multimodal_components` rejected HF repo ids.** `quantize.py
--hf_model_name_or_path Qwen/Qwen3.5-0.8B` worked, but the documented
export step failed with *"It should be a directory"*. Its sibling in the
same file already resolved repo ids via `snapshot_download`; now it does
too. This affected **every** VLM export.
`Qwen3_5ForConditionalGeneration` (dense Qwen3.5-VL) is now registered
for export and vision passthrough, which bugs 9 and 10 were blocking.
#### New: Qwen3.5-VL
`GatedDeltaNetSlicing` splits Megatron's fused `in_proj` (`[query, key,
value, z, beta, alpha]`) into HF's `in_proj_qkv` / `_z` / `_b` / `_a`,
taking sizes from the module's own `in_proj_split_sections` so TP
sharding falls out. Widening coverage to Qwen3.5's *gated
full-attention* layers then exposed a further split bug: gated attention
packs a per-head output gate beside each query head, so `_qkv_slicing`
split 192 rows as 96/48/48 instead of 128/32/32. It now derives the
group stride from `config.attention_output_gate`, matching
Megatron-Bridge's `split_qkv_weights`. The non-gated path is unchanged.
#### New: the export path verifies itself
- `assert_exported_checkpoint_matches` compares an exported checkpoint
against the model it came from — key set, shapes (accounting for NVFP4
`uint8` packing), safetensors index consistency, and values — replacing
existence-only assertions in all three export tests.
- `GPTModelExporter.save_pretrained` now raises if the export dropped
tensors the source checkpoint has, so *user* runs on architectures CI
never sees are protected too, not just tiny models.
- Loading a checkpoint whose quantizer tensors have no restorable state
now raises instead of silently loading unquantized.
- `assert_has_modelopt_state` replaces `rglob("modelopt_state")`, which
passes on an empty state; `assert_no_quantizers_matching` fails on
future HF↔Megatron name drift.
The mapping is also table-driven now: vision-tower prefixes live in
`all_mcore_hf_vision_passthrough_mapping` and
`with_language_model_prefix` is shared, so adding a VLM no longer means
editing `unified_export_megatron.py`. Five call sites that answered "is
this a VLM" three different ways now share `get_language_model` /
`is_vlm_config`.
### Usage
```bash
# Dense VLM (Qwen3-VL) -- no extra flags
torchrun --nproc_per_node 2 quantize.py \
--hf_model_name_or_path Qwen/Qwen3-VL-8B-Instruct \
--quant_cfg nvfp4 --tp_size 2 \
--export_megatron_path /tmp/Qwen3-VL-8B-NVFP4-megatron
torchrun --nproc_per_node 2 export_quantized_megatron_to_hf.py \
--hf_model_name_or_path Qwen/Qwen3-VL-8B-Instruct \
--megatron_path /tmp/Qwen3-VL-8B-NVFP4-megatron \
--pp_size 2 --export_unified_hf_path /tmp/Qwen3-VL-8B-NVFP4-hf
# Gated MoE (Qwen3.5-VL, Qwen3-MoE) -- no extra flags either. The scripts derive the
# expert layout from the model config, so quantize / distill / export all agree.
# --no_moe_grouped_gemm forces SequentialMLP if you want it explicitly.
```
### Testing
All in `nvcr.io/nvidia/nemo:26.08` on 2x RTX 6000 Ada.
| Suite | Result | Time |
|---|---|---|
| `tests/examples/megatron_bridge/` (full) | 18 passed | 27m58 |
| `tests/gpu_megatron/torch/export/` | 38 passed | 2m13 |
| `tests/unit/torch/export/` | 186 passed | 1.5s |
| pre-commit (ruff, ruff format, mypy, bandit) | clean | — |
| `tests/examples/megatron_bridge/test_quantize_export.py` on **2 GPUs**
(`pp_size=2`) | 3 passed | 5m |
The export leg of `test_quantize_and_export` now scales with `num_gpus`
like its quantize leg
already did. Previously it was hardcoded to one process, so the
collective checkpoint load ran at
PP=1 on both the 1-GPU PR runner and the 2-GPU nightly — which is how a
guard that raised on only
some pipeline stages (and therefore hung the job) reached review. The
dense `qwen3` case was dropped
in exchange: `qwen3_moe` already covers the non-VLM script path,
`qwen3vl` covers a dense decoder,
and that case was the one exceeding the 300s cap in CI.
#### Model coverage
`tests/gpu_megatron` runs in-process and is cheap, so it owns
per-architecture **mapping**
correctness. The example tests spawn `torchrun` per step and are ~50x
slower per case, so they
cover **script wiring** only — CLI flags, recipe resolution, and
checkpoint hand-off between steps.
| Suite | Models |
|---|---|
| `test_unified_export_megatron` | llama, nemotron, nemotron_h, qwen3vl,
qwen3_moe, qwen3_5_moe_vl x {none, FP8, NVFP4, +/-KV} x {grouped GEMM,
SequentialMLP} + eagle / medusa / MTP (29 params) |
| `test_megatron_importer` | nemotron_h, llama export->import round-trip
|
| `test_moe_layout_choice` | per-architecture grouped-GEMM exportability
(6 architectures) |
| `test_distill_megatron` | KD loss mechanics |
| Model | prune | quantize+export | QAD | distill+export |
|---|:--:|:--:|:--:|:--:|
| qwen3 | Y | Y | Y | Y |
| qwen3_moe | - | **Y (new)** | - | - |
| qwen3vl | - | **Y (moved from QAD)** | - | - |
| nemotron_h | Y | **Y (new)** | - | - |
| qwen3_5_vl | - | - | - | Y |
| qwen3_5_moe_vl | Y | **Y (new, both expert layouts)** | Y | - |
| deepseek_v3 | Y | - | - | - |
| gemma3vl | Y | - | ~~manual~~ removed | - |
QAD's unique property is that ModelOpt state survives distillation,
which needs one LLM and one
VLM rather than one case per architecture. Moving the rest to
quantize+export drops a `torchrun`
launch each: QAD went from 3 CI cases to 2 while quantize+export went
from 1 to 4, adding two
architectures for about a minute.
#### Real-model validation
Tiny fixtures cannot catch layout or scale bugs that only appear at real
dimensions, so the export
path was run end-to-end on released checkpoints. This is where bugs 7-10
came from.
| Model | Run | Result |
|---|---|---|
| Nemotron-3.5-Lightning-30B-A3B | NVFP4 4o6 PTQ → export → MMLU |
**0.7825 ± 0.0105** (gate 0.75) |
| Nemotron-3.5-Lightning-30B-A3B | Minitron pruning | 22.28B/3.00B
active, **0.5944** (gate 0.58) |
| Qwen3.5-0.8B (dense VLM) | FP8 PTQ → export → MMLU | BF16 0.4895 →
**0.4832** (±0.0127) |
| Qwen3.5-35B-A3B, half-depth (20 layers, 256 experts) | FP8 + NVFP4 PTQ
→ export | keys + shapes + **values** match reference |
| Qwen3.5-35B-A3B, full | FP8 PTQ | OOM on 2x48GB (see below) |
The half-depth model keeps real weights, real dims and all 256 experts.
Both expert layouts produce
identical key sets, and all exports pass
`assert_exported_checkpoint_matches(..., check_values=True)`
— every tensor, including all 20 x 256 experts, dequantizes to within
tolerance of the BF16
reference, so a transposed or mis-ordered expert stack would fail. NVFP4
lands in the correct packed
layout (`gate_up_proj [256, 1024, 1024]` U8, `weight_scale [256, 1024,
128]` E4M3,
`weight_scale_2 []` F32). Its *accuracy* is not meaningful — truncating
to 20 of 40 layers leaves a
chance-level model (BF16 0.2322, FP8 0.2538) — so it validates
correctness, not quality.
**Re-validated on the final code.** The numbers above were first taken
mid-review; since then the
NVFP4 block-scale merge changed on both packed paths, the vision-tower
download became two-stage,
and an expert-layout load guard was added. Both gating runs were
therefore repeated end to end:
Nemotron went 0.7748 → **0.7825 ± 0.0105** and Qwen3.5-0.8B went 0.4678
→ **0.4832 ± 0.0127**, with
the rest of the Nemotron pipeline reproducing exactly (3519 quantizers,
69GB checkpoint, 21GB
export). Both deltas are inside their own stderr, so the claim is that
the rework costs no accuracy
— not that it improved it. The Nemotron export also runs at `--pp_size
2`, exercising the new
collective layout guard on a real 30B MoE across pipeline stages.
Two limitations worth stating plainly:
- **No quantized accuracy number for a full-size MoE.** The full 35B
OOMs at 47.37 GiB while
*constructing* the model on 2x48GB, with grouped GEMM already enabled,
so no calibration knob
helps. Needs more GPUs than this setup has.
- **vLLM cannot yet serve packed FP8 Qwen3.5 experts.** `vllm
0.24.1.dev0` builds its fused expert
mapping weight-only, rewriting `experts.down_proj_input_scale` to
`w2_weight_input_scale` while the
parameter it registers is `w2_input_scale`. This is upstream and
independent of how the checkpoint
is produced — both of our export paths fail it identically. The 0.8B
numbers above are unaffected
(dense), and the packed exports are verified against the reference
checkpoint instead.
#### Guard verification
Each new guard was made to fire, not just to compile:
| Guard | Verification |
|---|---|
| Export self-check | Disabled the MoE guard, re-exported Qwen3-MoE -
independently reported all 24 dropped tensors. No false positives across
llama, nemotron, qwen3, qwen3-moe, qwen3vl, qwen3.5-vl, deepseek_v3
incl. eagle / medusa / MTP |
| Dropped-state raise | Deleted `modelopt_state` from a checkpoint with
50 quantizer tensors - raised instead of loading unquantized |
| NVFP4 value check | Flipped a `q_proj` - failed at `max_rel_err=1.74`
against a 0.3 threshold |
| Zero-centered gamma | Reproduced the off-by-1.0 on a good export -
caught as "not bit-exact" |
| Exclusion guard | Asserts no calibrated quantizer matches `conv1d` /
`mlp.router` / `output_layer` |
Exported artifacts are validated, not just their existence: 0 missing
keys vs reference, vision
tower bitwise-identical, dequantized weights within FP8 E4M3 error
(<=4.6%). The
`in_proj_a`/`in_proj_b` check is load-bearing - swapped alpha/beta would
still match on shape but
show ~100% error.
Also ran a tiny-Qwen3 **LLM** control through both steps to confirm the
exporter changes are a
no-op off the VLM path.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — the scripts now derive the
MoE expert layout from the model config, building SequentialMLP only for
architectures with no `experts.linear_fc1` rule, and the exporter raises
rather than dropping experts it has no rule for. Those runs previously
"succeeded" while writing a checkpoint containing no expert weights, so
no working behaviour is removed. `--no_moe_grouped_gemm` forces
SequentialMLP explicitly.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅ — approved (round 10: 0
CRITICAL, 0 IMPORTANT, 0 new suggestions); CodeRabbit approved earlier
### Additional Information
**MoE expert layout is now chosen automatically.** Only Nemotron-H can
export fused grouped-GEMM experts, so every other MoE architecture would
otherwise need `--no_moe_grouped_gemm` on all four scripts or hit a wall
at export. The scripts derive the layout from the model config — grouped
GEMM unless it would not be exportable — so they agree without threading
a flag. This changes MoE activation scales from one shared scale to
per-expert for the affected architectures.
Known gaps, unchanged by this PR:
- **Gated MoE still cannot use fused grouped GEMM.**
`_grouped_mlp_slicing` emits one weight per expert with no gate/up split
— its only prior caller, Nemotron-H, is non-gated, so every other MoE
architecture is built as `SequentialMLP` (see below). Adding that split
would restore the faster layout, but it needs a deliberate call on
activation-scale semantics: grouped GEMM keeps **one shared** activation
scale across experts while `SequentialMLP` has **per-expert** scales, so
the two are not numerically equivalent. It also needs EP>1 coverage.
- **Qwen3.5's alpha/beta gates share Megatron's fused `in_proj`
quantizer,** so they can only be kept in BF16 at export, not excluded by
name. Full fidelity needs per-section quantizers on the fused
projection.
- **Anchoring ModelOpt state on `.language_model`** (which would let
`quantize.py` quantize the language model directly and drop its
name-based non-LM disabling) needs a coordinated Megatron-Bridge change:
`save_sharded_modelopt_state` is ModelOpt code, but the restore the
Bridge path uses is Bridge's own and unconditionally restores onto the
root.
- **Gemma3-VL** remains Megatron-checkpoint only (`OMNIML-5366`).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added Muse Glimmer AutoQuantize and Alpamayo QAD workflows.
* Added streaming Kimi-K3 conversion and NVFP4 activation headroom
calibration.
* Added SFT-masked distillation for Megatron-Bridge.
* Added unified Hugging Face export for quantized Qwen3-VL and
Qwen3.5-VL checkpoints.
* MoE expert layouts are selected automatically, with an option to force
sequential experts.
* **Bug Fixes**
* Improved export validation for tensor coverage, MoE mappings,
quantizer state, and NVFP4 scales.
* Fixed Qwen3.5-VL GatedDeltaNet export handling.
* Preserved visual-model weights exactly during export.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
21b95adabb |
Add FP8 Vision Encoder quantization for Qwen3-VL and Qwen3.5 (#2083)
### What does this PR do? Type of change: New feature Adds opt-in FP8 Vision Encoder quantization recipes for Qwen3-VL and dense Qwen3.5: - `fp8_vision-kv_none`: FP8 Vision Encoder Linears, with the LLM and KV cache kept in high precision. - `fp8_vision_lm-kv_fp8_cast`: FP8 Vision Encoder and LLM Linears, with FP8 KV-cache cast. Patch embedding and vision-attention BMM operands remain in high precision. With `--calib_with_images`, calibration batches now pass through the complete VLM so multimodal inputs exercise the selected quantizers. This fixes image-text calibration for non-Nemotron VLMs and may change language-model activation ranges and output scales for existing commands. ### Usage ```bash # Vision Encoder only python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path <Qwen3-VL-checkpoint> \ --recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none \ --calib_with_images \ --calib_size 512 \ --skip_generate \ --export_path <output-checkpoint> # Vision Encoder + LLM + KV cache python examples/hf_ptq/hf_ptq.py \ --pyt_ckpt_path <Qwen3-VL-checkpoint> \ --recipe huggingface/qwen3_vl/ptq/fp8_vision_lm-kv_fp8_cast \ --calib_with_images \ --calib_size 512 \ --skip_generate \ --export_path <output-checkpoint> ``` For dense Qwen3.5, replace `qwen3_vl` with `qwen3_5` in the recipe path. ### Testing - Validated recipe selection, image calibration, and GPU calibration/export for Qwen3-VL and Qwen3.5, in both vision-only and joint configurations. - Consolidated test run after rebasing: 364 passed, 77 skipped. - Ruff, recipe validation, and `git diff --check` passed. - Transformers 4.57 compatibility verified for Qwen3-VL; Qwen3.5 tests capability-skip when the required Transformers classes are unavailable. Deployment evidence with Qwen3-VL-2B on RTX PRO 6000 BSE, eight fixed frames and a BF16 LLM: | Configuration | Accuracy mean | Vision Encoder kernel time | Full-request GPU kernel time | |---|---:|---:|---:| | BF16 | 48.75 | 25.45 ms | 47.43 ms | | Standard FP8 | 48.42 | 19.35 ms (**24.0% faster**) | 41.39 ms (**12.7% faster**) | The accuracy mean covers MMMU, RealWorldQA, Video-MMMU, MVBench, and Video-MME. Serving reached 7.7% lower end-to-end latency and 7.9% higher throughput at concurrency 16. Qwen3-VL-2B accuracy was evaluated through vLLM on B300 with `--enforce-eager`. Both checkpoints used the same judge-free tasks, Qwen sampling preset, seed, and task parameters. | Benchmark | BF16 | VE-only FP8 | Delta | |---|---:|---:|---:| | MMMU validation | 45.33 | 45.22 | -0.11 pt | | RealWorldQA | 64.97 | 65.10 | +0.13 pt | | Video-MMMU | 31.56 | 31.11 | -0.45 pt | | MVBench | 51.40 | 50.10 | -1.30 pt | | Video-MME | 50.48 | 50.59 | +0.11 pt | | **Unweighted mean** | **48.75** | **48.42** | **-0.33 pt** | Runtime support for quantized Vision Encoder Linears is separate from this ModelOpt checkpoint-generation change. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ❌ — `--calib_with_images` now performs the intended complete VLM forward and may change language-model calibration scales. Recipe-based VLM PTQ also scopes recipe rules to the complete model. Both changes are documented in the changelog. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ — will run `/claude review` after opening the PR. ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added FP8 vision quantization recipes for Qwen3-VL and Qwen3.5. * Added Muse Glimmer AutoQuantize, Alpamayo QAD, streaming Kimi-K3 conversion, layerwise checkpoint export, and NVFP4 calibration/export workflows. * Added ONNX FP16 conversion support for excluding selected nodes. * **Bug Fixes** * Improved multimodal calibration, ONNX scale handling, and NVFP4 CPU compatibility checks. * **Documentation** * Expanded guidance for vision quantization, calibration, precision, and conversion workflows. * **Breaking Changes** * Removed deprecated PTQ and evaluation interfaces and raised the minimum supported Megatron container version. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: mpariente <mpariente@nvidia.com> Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> Co-authored-by: Shengliang Xu <shengliangx@nvidia.com> Co-authored-by: Shengliang Xu <106840466+shengliangxu@users.noreply.github.com> |
||
|
|
de3eda8a11 |
Restructure recipes: split per-model_type recipes from model-hub checkpoint recipes (#2219)
### What does this PR do? **Type of change:** Refactor (recipe-library layout) + documentation — backward-breaking for saved `--recipe` paths. Separate the two kinds of built-in Hugging Face recipes that were previously mixed under `modelopt_recipes/huggingface/`: - **`huggingface/<model_type>/`** — architecture recipes keyed by the transformers `model_type`; one recipe covers every checkpoint of that architecture. **Unchanged.** - **`models/<org>/<model_id>/`** — a *new top-level tier* for recipes that mirror one specific published checkpoint, keyed by its **model-hub path** (as on the Hugging Face Hub, ModelScope, etc.) so the on-disk path equals the hub path. Concretely, the model-instance recipes move out of `huggingface/` to the top level: - `huggingface/models/mistralai/…`, `huggingface/models/nvidia/…` → `models/mistralai/…`, `models/nvidia/…` - `huggingface/step3p5/Step3.5-Flash/…` → `models/stepfun-ai/Step-3.5-Flash/…` (re-keyed to the canonical HF repo id [`stepfun-ai/Step-3.5-Flash`](https://huggingface.co/stepfun-ai/Step-3.5-Flash) — org `step3p5`→`stepfun-ai`, id `Step3.5-Flash`→`Step-3.5-Flash`) **Why:** `modelopt_recipes/README.md` already documented a top-level `models/` tier, but the files lived under `huggingface/models/` and instance-specific recipes were awkwardly nested under the per-`model_type` tree. This aligns the filesystem with the documented layout and makes the instance tier hub-addressable — given a checkpoint id you can find (or place) its recipe with no lookup table. `load_recipe` resolves paths directly under `modelopt_recipes/`, so a top-level `models/` sibling of `general/` and `huggingface/` works identically. The move is metadata-only — all recipe YAML content is byte-identical (`R100` renames). Everything else is updating references (nvidia launcher YAMLs, `test_loader.py`) and docs: a new `models/README.md`, plus `huggingface/README.md`, root `README.md`, `ptq.md`, and the `10_recipes.rst` guide, which no longer describe instances under `huggingface/`. ### Usage Recipe paths for the moved checkpoint recipes lose the `huggingface/` prefix (and Step 3.5 Flash is keyed by its hub id): ```python from modelopt.recipe import load_recipe # before load_recipe("huggingface/models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16") load_recipe("huggingface/step3p5/Step3.5-Flash/ptq/nvfp4-mlp-only") # after load_recipe("models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16") load_recipe("models/stepfun-ai/Step-3.5-Flash/ptq/nvfp4-mlp-only") ``` The same rename applies to `--recipe …` CLI values and launcher `QUANT_CFG:` entries. Architecture recipes under `huggingface/<model_type>/` are unaffected. ### Testing - **Recipe resolution (torch-free):** parsed every recipe under `models/` and confirmed all `$import` targets resolve against the recipe root — 0 dangling across the tier. - **Docs consistency:** re-ran the `tests/unit/recipe/test_recipe_docs.py` logic; it now globs both `huggingface/` and `models/`, and every model dir (incl. `Step-3.5-Flash`, `Nemotron-3-Nano-4B-BF16`, …) plus every `general/ptq` recipe is still mentioned in `ptq.md`. - **Reference sweep:** repo-wide grep confirms no remaining references to the old paths outside the intentional historical CHANGELOG entries (released 0.44 / 0.45). - **pre-commit:** `markdownlint-cli2`, license-insert, and `bandit` hooks pass on the changed files. - Note: the full `pytest` suite was not run in my environment (no `torch`), so `test_recipe_docs.py` / `test_loader.py` should be exercised in CI. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ❌ — `--recipe` / `load_recipe` paths for the checkpoint-mirror tier change (drop the `huggingface/` prefix; `step3p5/Step3.5-Flash` → `stepfun-ai/Step-3.5-Flash`). Documented as a Backward Breaking Change in `CHANGELOG.rst` (0.47); the only *released* old paths affected shipped in 0.45. A clean break was chosen over a symlink or loader-alias shim. - If you copied code from any other sources or added a new PIP dependency …: N/A - Did you write any new necessary tests?: ✅ — updated `test_recipe_docs.py` to also glob the top-level `models/` tier so instance recipes stay covered by the doc-consistency check. - Did you update Changelog?: ✅ — added a 0.47 **Backward Breaking Changes** entry. - Did you get Claude approval on this PR?: ❌ <!-- run /claude review --> ### Additional Information Design note: an earlier iteration nested everything under `huggingface/model_type/` + `huggingface/models/`; the final layout keeps `huggingface/` flat (per-`model_type`) and lifts instances to a top-level `models/` tier, matching what `modelopt_recipes/README.md` already documented. The `Step3p5*` architecture class names (from the model's `trust_remote_code` modeling code) are unrelated to the recipe path and are left unchanged. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added checkpoint-specific PTQ recipes for Kimi-K3, Mistral Medium 3.5, and NVIDIA Nemotron models. * Added a Nemotron speculative-decoding warm-start recipe. * **Documentation** * Clarified recipe selection and directory organization. * Documented checkpoint naming conventions and updated usage examples. * **Bug Fixes** * Updated launcher configurations and examples to reference the new recipe locations and corrected model names. * **Tests** * Improved automatic recipe discovery and validation of documented recipe paths. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
8810eb5e31 |
Add ModelOpt recipe for DeepSeek-V4-Pro-0813 NVFP4 and --recipe to its PTQ script (#2287)
### What does this PR do?
Type of change: new feature
The quantization config for `nvidia/DeepSeek-V4-Pro-0813-NVFP4` existed
only as Python inside `_build_nvfp4_experts_cfg()`, so the released
checkpoint had **no entry in `modelopt_recipes/`** and could not be
looked up by name the way every other published model can.
`modelopt_recipes/README.md` states the goal directly — a recipe is
*"the single, version-controlled source of truth for how a model is
optimized … expressed as data instead of code"* — and this model was the
exception.
This adds:
-
`modelopt_recipes/huggingface/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only.yaml`,
composed from the existing `configs/ptq/units/base_disable_all` and
`configs/numerics/nvfp4` units.
- An optional `--recipe` flag on `examples/deepseek/deepseek_v4/ptq.py`.
It follows **`examples/kimi/kimi_k3`**, the closest precedent: a very
large MoE whose source already ships MXFP4 routed experts, converted via
`--cast_mxfp4_to_nvfp4` rather than through `examples/hf_ptq`, and
already wired to `--recipe` with a published YAML.
### Usage
```sh
torchrun --nproc-per-node 8 deepseek_v4/ptq.py \
--model_path <mp8_checkpoint> \
--config <DeepSeek-V4-Pro-0813>/inference/config.json \
--calib_size 512 \
--calib_seq 4096 \
--output_path <amax_dump> \
--recipe huggingface/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only
```
Omitting `--recipe` keeps the previous behaviour exactly.
### Testing
- `load_recipe` resolves the YAML and yields `num_bits (2, 1)` with
`block_sizes {-1: 16, type: dynamic, scale_bits: (4, 3)}` — identical to
the hardcoded config.
- **Equivalence checked behaviourally**, not by eyeballing dicts: both
configs were resolved against representative quantizer names using
last-match-wins, and agree on all of them.
| quantizer | hardcoded | recipe |
| --- | --- | --- |
| `...ffn.experts.17.w1_weight_quantizer` | enabled, NVFP4 | enabled,
NVFP4 |
| `...ffn.experts.17.w2_input_quantizer` | enabled, NVFP4 | enabled,
NVFP4 |
| `...ffn.shared_experts.w1_weight_quantizer` | disabled | disabled |
| `...attn.wq_weight_quantizer` | disabled | disabled |
| `mtp.0.ffn.experts.2.w1_weight_quantizer` | disabled | disabled |
| `lm_head_weight_quantizer` | disabled | disabled |
- `mtq.quantize` documents `algorithm` as a string **or** a dict keyed
on `method`, so the recipe's `{'method': 'max'}` needs no translation.
- `pre-commit` clean, including `validate modelopt recipes`.
No GPU run: this changes config plumbing only, and the default path is
byte-identical to before.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `--recipe` is optional and
defaults to `None`; without it `_build_nvfp4_experts_cfg()` is used
exactly as before.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies; `modelopt.recipe` is already a first-party import.
- Did you write any new necessary tests?: N/A — no new logic;
equivalence to the existing config is the property that matters and is
documented above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — recipe library addition; recent recipe/example PRs add no entry.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
The recipe covers the **quant config only**. `--calib_seq` — the setting
that mattered most for this checkpoint, since the 512 default does not
cover long-context activation ranges — is a dataloader argument rather
than part of the `mtq` config, so it stays on the CLI. Worth knowing if
the recipe is ever treated as a complete reproduction of the released
checkpoint: it is not, on its own.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added post-training quantization support for DeepSeek-V4-Pro-0813
routed experts using NVFP4.
* Added an optional recipe path for selecting equivalent quantization
settings.
* Preserved source formats for shared experts, attention, embeddings,
output layers, and MTP components.
* **Bug Fixes**
* Improved validation for missing or malformed quantization
configurations.
* Added safeguards against unsupported formats, scopes, algorithms, and
enabled MTP quantizers.
* **Documentation**
* Documented checkpoint conversion behavior, calibration requirements,
and supported quantization workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
029c67f27e |
feat(export): export each decoder layer as layerwise calibration finishes it (#2136)
### What does this PR do?
Type of change: new feature
Layerwise calibration can already resume, but only through a
full-precision scratch checkpoint, and a completed run still pays for a
second whole-model export pass over it.
`layerwise.export_dir` writes each decoder layer to its own quantized
shard as soon as calibration finishes with it, so the directory is a
complete, loadable checkpoint when the last layer lands and
`export_hf_checkpoint()` is skipped. The shards *are* the resume
artifact: a restarted run reuses layers already on disk instead of
recalibrating and re-exporting them, so no full-precision copy of the
model accumulates. The resume directory beside it holds only the current
boundary's cached activations and the per-layer output shapes.
Setting the config field is the whole switch — no CLI flag. `hf_ptq.py`
rewrites its value to `--export_path`, and derives the resume directory
(`<export_path>.layerwise_resume`) when you haven't chosen one.
One shard per layer is what makes resume safe: shards are written whole
and named from the layer index, so a crash can lose the layer in flight
but never corrupt an earlier one, and a re-run overwrites in place.
Because a resumed run never recalibrates the layers it skipped, **the
in-memory model is not valid for inference afterwards**; the field
implies `--skip_generate`.
Includes a pre-existing `main` fix this depends on: `_is_layerwise` used
`getattr` on an algorithm that YAML parses as a **dict**, so it answered
`False` for every layerwise recipe in the repo and the batch-size probe
it gates was never skipped. **Behaviour change:** `--batch_size 0` now
yields `batch_size=1` for layerwise recipes, as its comment intends.
Detection also now scans every algorithm entry rather than the first, so
a list-form recipe whose `layerwise` block is not first is recognised as
layerwise — same batch-size consequence. Nothing else on the non-fused
paths changes: `FUSION_FREE_FORMATS` is the exact set the inline list
held, `save_non_weight_artifacts` is a lift of the streaming exporter's
own block, and the calibration-loop changes are gated on an exporter
being present.
**Refused before calibration starts**, since each would otherwise
produce a silently different checkpoint rather than fail:
| Refused | Why |
|---|---|
| AWQ / SVDQuant | need pre-quant-scale steps that are still whole-model
|
| Weight-tied quantized modules | `sync_tied_input_amax` merges amaxes
across a partner that may be uncalibrated or already written |
| Multi-process (FSDP2) | every rank would write the same shards |
| Multimodal (VLM) | calibration runs on the extracted language model |
| MTP models | exclusions applied after calibration has written
everything |
| AutoQuantize recipes | only the mono-quantize path retargets
`export_dir` |
| Spec-dec, `--vllm_fakequant_export`, non-dense sparsity,
`int8_smoothquant`, encoder-decoder `model_type` | each routes to a
second exporter that would overwrite `--export_path` |
| `export_dir` on more than one algorithm entry, or on any but the last
| export finalizes shards as calibration walks the layers, so a later
pass would change the model after its checkpoint was written |
Shards are also bound to the run that produced them
(`.layerwise_export.json`: model class, layer count, formats, KV-cache
format, and a digest of the resolved quant config), so one run's
manifest cannot finalize another's shards. Source weights are not
digested — that would mean reading the whole model — so
differently-trained weights at the same path compare equal.
### Why a separate exporter
Three reuse paths were considered before adding one:
- **Extend `_StreamingShardWriter`.** It buffers by `max_shard_size`
into `__shard_part_*` temp names and renames to canonical names only in
`finalize()`, once the shard count is known. The resume invariant needs
the opposite: a stable `model-layer-00007.safetensors` committed when
layer 7 finishes, so "shard exists" means "layer done" across a restart.
Forcing a per-layer flush still leaves temp names, finalize-time
renaming, and an in-memory `_key_to_part` — every method would change.
- **Keep the layerwise checkpoint and run the streaming exporter at the
end.** This works, and it is why the pitch above is *not* durability:
that already exists. What it leaves is a second whole-model pass owed
*after* calibration finishes — itself needing a GPU session — where
per-layer export makes the last calibrated layer also the last exported
one. Scratch size only separates them for weight-mutating calibrators:
`save_layer_state` is off under per-layer export, but with
`calib_mutates_weights: false` (the shipped recipe) the checkpoint holds
just amax buffers either way.
- **Factor a shared per-module writer around `ExportContext`.** The
right long-term shape, but it touches all three existing export paths;
doing it here makes this change larger, not smaller.
One deliberate divergence from `_StreamingShardWriter` worth knowing
about: it clones tensors that share storage, this path lets `save_file`
raise instead. No model was found where the clone fires, and copying
unattributed aliases can hide a real bug rather than surface it. If a
checkpoint ever trips it, that is information we want.
Happy to take a different call on this — flagging it for maintainer
sign-off rather than assuming it.
### Usage
```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --export_path <out> \
--recipe modelopt_recipes/general/ptq/nvfp4_experts_only-kv_fp8_layerwise_export.yaml
```
Interrupt and rerun the same command: calibration resumes from the last
committed layer, finished shards are reused, and a run that had already
finished every layer only re-runs `finalize()`.
```yaml
quantize:
algorithm:
method: max
layerwise:
enable: true
calib_mutates_weights: false
export_dir: /tmp/modelopt_layerwise_export # presence is the switch; value replaced with --export_path
# checkpoint_dir omitted -> derived as <export_path>.layerwise_resume
```
### Testing
Each row exports the same calibration two ways — per-layer, and
whole-model via `export_hf_checkpoint()` — and compares them **tensor
for tensor and config for config**.
The 35B row was re-run on the current head, against a baseline built
from a `main` worktree rather than from this branch, so it covers both
"per-layer differs from whole-model" and "this branch broke the shared
whole-model path". The other rows date from earlier heads; the code they
exercise is unchanged, but they are not fresh runs.
| Model | Config | Result |
|---|---|---|
| Qwen3.6-35B-A3B (40 layers, 256 fused experts) | NVFP4 W4A4
experts-only + FP8 KV | 123,513 tensors, 0 mismatched; `config.json`,
`hf_quant_config.json`, `generation_config.json` all identical |
| Qwen3-30B-A3B (48 layers, 128 per-expert linears) | NVFP4 experts
(`nvfp4_static` weights) + `mse`, offload | 74,163 tensors, 0 mismatched
|
| Qwen3-30B-A3B | same, `SIGKILL` after 25/48 layers, then resumed |
74,163 tensors, 0 mismatched **vs the uninterrupted run** |
| Llama-3.1-8B-Instruct | FP8 dense + FP8 KV, resident | 803 tensors, 0
mismatched |
**Refusals verified on real checkpoints**, each writing **zero shards**
and never reaching calibration — the "refused before calibration starts"
claim above, demonstrated rather than asserted: multimodal and MTP
(Qwen3.6-35B, the MTP case on a text-only view since the multimodal gate
fires first), tied embeddings (Qwen3-0.6B), and multi-process (2-rank
`torchrun --use_fsdp2`, Llama-3.1-8B).
**Served, not just compared.** Under vLLM 0.27.1 (Marlin NVFP4 kernels,
SM 8.9): the 30B checkpoint exported three ways — whole-model,
per-layer, per-layer-resumed-after-a-kill — and the 8B exported both
ways all load and produce **identical greedy generations, 4/4 prompts**
within each model.
**Index integrity** on every checkpoint above: each `weight_map` key
resolves to the shard actually holding it; 0 missing, 0 extra, 0
mis-routed. Tensor equality alone never exercises that, and it is the
one artifact per-layer export builds differently.
Resume state stays bounded: **332 KB beside 22 GB** of shards on the
35B, **396 KB beside 19 GB** on the 30B — the committed boundary's
activations only, not one set per layer.
Not covered: the `trust_remote_code` `*.py` copy path.
Nemotron-Nano-12B-v2-Base fails with a CUDA illegal memory access on
these cards, on the whole-model baseline too, so it is an environment
limit rather than a result.
**Comparing the configs is new, and it caught a real bug.**
`get_quant_config` reports on the quantizer modules, which
`export_layer` replaces as it goes, so reading it in `finalize()`
described a model with no quantizers left: the checkpoint advertised
`quant_algo: null` while its weights were packed NVFP4, and under the
shipped experts-only recipe `hf_quant_config.json` was not written at
all. It is snapshotted in `__init__` now, beside the kv-cache format
already captured there — which is why that one field was correct while
the rest were not. Uniform FP8 and NVFP4 hid it because their configs
survive the conversion; only a mixed model loses its algo, and mixed is
what every shipped layerwise-export recipe is. Reverting the fix fails
`test_moe_export_matches` and passes the ten uniform-format cases,
matching what the 35B shows.
**24 GPU tests** in `tests/gpu/torch/export/test_layerwise_export.py`.
The equivalence oracle is a cross-product: {FP8, NVFP4, NVFP4 +
`get_qdq_activations_from_prev_layer`, mixed FP8/NVFP4, KV-cache} ×
{fresh, resumed-after-interruption}, each compared tensor-for-tensor
against `export_hf_checkpoint`. Plus MoE export; resume fail-fast;
resume artifacts replaced and pruned; complete-manifest finalize-only;
shards-without-manifest refusal; shards-from-a-different-run refusal
(format and module selection); identity-without-shards does not block a
rerun; export-does-not-mutate-the-model; index routes every key to the
shard holding it; AWQ refusal (from config, and after calibration);
export-without-`checkpoint_dir`.
**Unit tests** in `tests/examples/hf_ptq/test_example_utils.py` cover
the list-valued `algorithm` shapes: which entry owns export, whose
`checkpoint_dir` is derived, per-entry resume bases, both ambiguity
refusals, and the recipe shapes `recipe_layerwise_blocks` normalizes
(dict, list order, config object, and the empty cases).
`tests/gpu/torch/export/` 150 passed / 2 skipped (pre-existing env
skips) · `tests/unit/recipe` 284 · `tests/unit/torch/export` 186 ·
`test_layerwise_calibrate` 33 · `test_example_utils` 42 · pre-commit
clean.
Also verified: the exported directory reloads through
`AutoModelForCausalLM` and runs a forward.
Not a speed win: per-layer export was slower than the streaming export
in one offload pairing (271s vs 208s, the per-layer fusion probe),
though those runs shared GPUs so the magnitude is not cleanly measured.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `export_dir` defaults to
`None`; existing paths unchanged when unset, except the batch-size
change noted above.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — not yet; draft.
### Additional Information
**Pre-existing bug found on the way, not fixed here.** Layerwise
calibration leaves `self_attn.o_proj`'s input amax at `0.0` on every
layer but the last, so a full-NVFP4 layerwise model cannot be exported
by *any* path. `get_qdq_activations_from_prev_layer=True` avoids it,
pinning the cause to the pre-`calib_func` capture pass — which also
explains why only the last layer, the one that skips it, is correct.
That combination now works with per-layer export (it asserted on layer 0
until review caught it). Hidden until now because the shipped NVFP4
layerwise recipes are experts-only; the NVFP4 tests here exclude
`o_proj` for the same reason. Deserves its own issue.
---------
Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
424f47b871 |
Add QAD example for Alpamayo (#2271)
### What does this PR do?
Add `examples/alpamayo/qad.py` which distills the quantized Alpamayo
checkpoint's VLM against the FP16 VLM of the original with ModelOpt's
QADTrainer, and shards student and teacher with FSDP2 for multi-GPU
runs. Only the VLM is trained; the action expert stays frozen.
Type of change: new example
<!-- Details about the change. -->
### Usage
```
torchrun --standalone --nproc_per_node 8 qad.py \
--student_ckpt ./alpamayo-auto \
--output_dir ./alpamayo-auto-qad \
--parquet ./train_clips.parquet \
--max_steps 500 --fsdp2 --grad_ckpt --export
```
### Testing
Tested end-to-end on public Alpamayo-1 checkpoint
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Very short summary of changes only for new features,
backward breaking changes, deprecations, or fixes for critical bugs
present in previous releases. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added a quantization-aware distillation workflow for AlpamayoR1
models.
- Supports prompt-only or rollout-inclusive distillation, FSDP2
training, dataset slicing, revision pinning, and checkpoint resumption.
- Supports exporting trained models as complete, reloadable AlpamayoR1
checkpoints.
- Added optional vision-parameter freezing, trajectory-history fusion,
gradient checkpointing, evaluation, and synchronized training cadence.
- **Documentation**
- Expanded the Alpamayo guide with setup, training, dataset, and export
instructions.
- Clarified sensitivity-based quantization behavior.
- Added a version 0.47 quantization changelog entry.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Rohan Joshi <rohjoshi@nvidia.com>
|
||
|
|
6a2ae5a25b |
Fix the llm_eval timeout: reachable MMLU mirror + no pipe deadlock (#2270)
### What does this PR do? Type of change: Bug fix `tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8` has been failing with `Failed: Timeout (>900.0s) from pytest-timeout` on unrelated branches (runs 33027056246 and the one for `ad83a428`, while the 2026-08-24 nightly passed). It is not the test being slow — it is the harness deadlocking, and the deadlock also destroys the diagnostics that would explain the underlying kill. **Mechanism.** The traceback shows `self = <Popen: returncode: -9 args: ['scripts/huggingface_example.sh', ...]>` while still blocked in `stdout.read()`. The launcher was SIGKILLed (nothing in pytest sends SIGKILL — pytest-timeout raises in the main thread, and the test's `finally` `pkill` sends SIGTERM and only runs afterwards — so an OOM kill is the likely source). But `subprocess.run(..., stdout=PIPE, stderr=STDOUT)` waits for **EOF on the pipe**, not for the process, and a surviving grandchild (the TRT-LLM serve/build worker) still holds the write end. EOF never arrives, so the test blocks until the 900 s alarm. Because the pipe is never drained, **every line of child output is discarded**, which is why the CI log says nothing about what the script was doing when it died. Reduced to a self-contained reproducer: ```python script = "sleep 300 & echo 'launcher output'; sleep 0.3; kill -9 $$" subprocess.run(["bash", "-c", script], stdout=PIPE, stderr=STDOUT, text=True, timeout=20) # -> TimeoutExpired: still blocked in communicate() after 20.0s, output lost ``` **Fix.** `_run_capturing` now starts the command in its own session, drains its output on a reader thread (so logs stream as they arrive instead of being buffered until the end), waits on the *process*, and kills the process group if descendants still hold the pipe after a 30 s grace period. A killed launcher now fails in seconds with its logs intact instead of silently burning the test's whole timeout. This does not fix whatever kills the script; it makes it diagnosable. Worth noting separately: `test_qwen3_eval_fp8` took **749.10 s against its 900 s mark** on the last green nightly, so it is fragile regardless and may want its work trimmed or its budget raised once the logs show where the time goes. ### Usage ```python # unchanged public API run_example_command(cmd_parts, example_path="llm_eval") ``` ### Testing Verified against the reproducer above and on the normal paths: | scenario | before | after | | --- | --- | --- | | launcher SIGKILLed, survivor holds the pipe | blocks indefinitely (900 s in CI) | `rc=-9` in 3.5 s, `'launcher output'` captured | | the surviving descendant | keeps running | killed with the process group (stopped ticking, 20 -> 20 bytes) | | normal exit | ok | `rc=0`, stdout and stderr interleaved in order | | non-zero exit | ok | `rc=3`, output captured | The example-test suites that use this helper run through the same code path; `tests/examples/megatron_bridge` (16 passed, 1 skipped) exercised it on nemo:26.08 in the branch this was extracted from. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `_run_capturing` keeps its `(returncode, output)` contract; only the buffering strategy changed. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ❌ — this is test infrastructure; the scenario needs a process that outlives a SIGKILLed parent, which is awkward to assert in CI. Verified manually with the reproducer above. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — test-infrastructure fix. - Did you get Claude approval on this PR?: ❌ 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved command execution reliability with real-time output capture. * Ensured lingering child processes are cleaned up after commands exit or are interrupted. * Added warnings when forced cleanup may truncate output. * Prevented hangs when descendant processes keep output streams open. * **Documentation** * Updated MMLU setup instructions to use the Hugging Face dataset repository. * Improved Windows instructions by explicitly using `curl.exe`. * **Examples** * Improved MMLU downloads with retries, separate timeouts, resume support, and automatic temporary-file cleanup. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --- ### Update: the second half of the failure With the streaming fix in place, the next CI run showed the *actual* cause, which the old code had been hiding. That run blocked in `process.wait()` with `<Popen: returncode: None ...>` — child alive and working, not the previous dead-child pipe deadlock — and the now-visible output was: ``` --2026-08-27 18:09:20-- (try: 4) https://people.eecs.berkeley.edu/~hendrycks/data.tar Connecting to people.eecs.berkeley.edu ...|128.32.139.28|:443... failed: Connection timed out. Retrying. --2026-08-27 18:11:40-- (try: 5) ... ``` `huggingface_example.sh` downloads the MMLU tarball from `people.eecs.berkeley.edu`, that host stopped answering around 2026-08-25, and wget's default retry policy (20 tries, ~2 min per connect timeout) consumed the whole 900 s budget. Not runner-specific: the URL also times out from a developer workstation, and the nightlies flipped 08-24 ✅ / 08-25 ✅ / **08-26 ❌ / 08-27 ❌**, matching the outage. So this PR now carries both halves of the same failure: 1. the harness no longer deadlocks and no longer swallows the logs (`985809cc2d`), and 2. the MMLU data comes from HuggingFace's copy of the same tarball, with bounded retries (`40f1d89154`). The mirror is byte-for-byte the same dataset in the same layout the script already expects — verified by running the exact download/extract commands: ``` https://huggingface.co/datasets/cais/mmlu/resolve/main/data.tar -> HTTP 200, 166 MB data/mmlu/{dev,test,val}/ -> 57 subject CSVs each, plus auxiliary_train/ ``` `wget --timeout=20 --tries=3` plus an explicit error means the next dataset-host outage fails in about a minute with "Could not download the MMLU test data. Set MMLU_DATA_PATH to a local copy." instead of silently eating a test's timeout. The same URL is updated in `examples/llm_eval/README.md` so a manual run does not hit the dead host either. --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5500999d0b |
Add Kimi-K3 NVFP4 experts and FP8-PB attention recipe (#2206)
### What does this PR do?
Type of change: new example
Adds the calibration-free conversion pipeline and checkpoint-mirror PTQ
recipe used for `nvidia/Kimi-K3-NVFP4`:
- streams the 96-shard Kimi-K3 checkpoint without loading the 2.8T
model;
- casts the source MXFP4 routed experts to NVFP4 with expert
`input_scale=1.0`;
- quantizes the selected KDA and MLA attention weights to 128x128 block
FP8;
- leaves shared/latent experts, routers, convolutions, norms, the vision
tower, `lm_head`, and KV cache unquantized;
- emits mixed-precision Hugging Face metadata for deployment; and
- adds an exact recipe under
`modelopt_recipes/huggingface/models/moonshotai/Kimi-K3/` that directly
configures the streaming converter.
It also fixes `NVFP4QTensor.quantize()` probing CUDA/Blackwell
capability before checking whether the tensor is on CUDA and whether the
optional TensorRT-LLM fast path was requested. That probe broke the
converter's supported CPU path on hosts without a compatible GPU.
### Usage
```bash
python examples/kimi/kimi_k3/quantize_to_nvfp4.py \
--source_ckpt /models/moonshotai/Kimi-K3 \
--output_ckpt /models/Kimi-K3-NVFP4 \
--recipe huggingface/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention \
--jobs 8
```
The conversion requires no calibration dataset, forward pass, or GPU.
Multi-node shard conversion is also supported through `--rank`,
`--world_size`, and `--run_id`.
### Testing
```bash
uv run --frozen --extra dev python -m pytest -q \
tests/unit/torch/quantization/test_nvfp4_tensor.py \
tests/unit/recipe/test_kimi_k3_recipe.py \
tests/unit/recipe/test_recipe_docs.py \
tests/unit/torch/export/test_shard_cast_utils.py \
tests/examples/kimi/test_kimi_k3_quantize_to_nvfp4.py
```
Result: 34 passed.
All pre-commit hooks pass for the changed files, including recipe
validation, Ruff, mypy, Bandit, YAML formatting, and markdownlint.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: N/A
### Additional Information
The resulting checkpoint and model card are available at
https://huggingface.co/nvidia/Kimi-K3-NVFP4.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added calibration-free Kimi-K3 MXFP4-to-NVFP4 conversion with optional
FP8 attention quantization and distributed processing.
- Added NVFP4 activation calibration, weight-only quantization recipes,
grouped-expert quantization, compiled quantization options, SFT-masked
distillation, and MLflow tracking.
- **Documentation**
- Updated quantization terminology, recipe catalogs, checkpoint
guidance, and Kimi-K3 conversion instructions.
- **Bug Fixes**
- Improved CPU NVFP4 behavior, tied-weight export handling, EAGLE-3
training compatibility, and checkpoint export reliability.
- Removed deprecated configuration options and legacy evaluation
examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
|
||
|
|
7ff81dd795 |
Add day-0 verbosity gate + harden release skills from a live run (#2254)
### What does this PR do?
Type of change: Bug fix, new feature, documentation
Day-0 release treats verbosity as a hard gate, but nothing in the skill
measured it — `gate_compare.py` only does accuracy. A run could complete
Step 5 and report a publish recommendation with a gate silently
unmeasured. This adds the missing gate and folds in fixes for problems
that cost GPU-h on a live release run.
**New: `gate_verbosity.py` + Step 5b.** Pure `evaluate_verbosity()` plus
an artifact harvester, matching `gate_compare.py`'s conventions and
`failure_class` vocabulary.
Two details it encodes, both of which produced wrong verdicts before
they were understood:
- **Read `response_stats.avg_completion_tokens`.** The
`reasoning.*_tokens` fields are always `0`, which reads as "the harness
never captured tokens" and pushes you to word counts. Only the
reasoning/content *split* is missing, not the total. Words disagree with
the gate: one task read **+6.10% FAIL** in words and **+1.32% PASS** in
tokens.
- **Two run-hygiene filters.** Pooling a mismatched reasoning-effort run
reported **+43.06% FAIL** on a task that is **+1.24% PASS** matched; one
truncated run (n=200 vs 294) reported **+9.58% FAIL** on a task that is
**+1.71% PASS** without it. Tasks with no common sample count are
reported `not_comparable` rather than as a delta.
**Skill hardening**, each from a specific failure:
- **Step 2b canary** — poll ceiling must exceed load time (a 50 min poll
against a 51 min load failed a checkpoint that serves fine), print an
explicit `RESULT:` on every path (a fall-through exits 0 and reads as
PASS), log to shared storage, and canary the **as-exported** artifact
rather than a copy modified to make it work.
- **Step 4 config parity** — assert the candidate config differs from
the baseline's in nothing but checkpoint path and served-model name. A
mismatched `parallelism` was worth ~2 pp, enough to invert the sign of a
delta, and cost four re-runs.
- **Statistical power** — re-running does not guarantee fresh samples:
with a warm NEL response cache two runs came back bit-identical to 16
digits.
- **Step 6 closeout** — verify the published path against the evaluated
one by inode, and prefix rejected sibling exports.
- **Size gate** — growth is blocking by default and waived only when the
validation summary's
`source_precision` shows an already-sub-8-bit source (which cannot
shrink further under a
4-bit recipe) and the growth is within what that explains.
`source_precision` is now a
recorded field in the ptq validation table, so the waiver is reachable
from the normal
pipeline, and `SIZE_NOT_REDUCED` has a triage row pointing at declaring
it.
- **`ptq.py`** — `--calib_seq` matters more than `--calib_size`, and
`--mse_calibrate` is a no-op under `--cast_mxfp4_to_nvfp4` (it tunes
weight quantizers only, and the cast overwrites `weight_scale`).
- **`.gitignore workspaces/`** — the skills create scratch directories
inside the repo; nothing excluded them.
### Usage
```bash
python "$SKILL_DIR/scripts/gate_verbosity.py" \
--baseline <baseline_eval_root> --candidate <candidate_eval_root> \
--glob 'eval_*' --threshold 0.05
```
Exit codes match the sibling gates: `0` pass, `1` the gate ran and
failed, `2` the gate could not
read its input (wrong root, `--glob` matched nothing, everything
excluded). Prints per-task tokens,
delta, `within_threshold`, `sample_count`, run counts,
`dropped_mismatched_runs`, any
`truncated_comparison`, `not_comparable`, `harvest_diagnostics`, and a
`max_abs_delta` summary.
### Testing
- 8 new unit tests in `test_gates.py`, one per real failure mode
(two-sided threshold, partial-run filtering, unequal sample counts,
short-output warning, one-sided tasks, empty input). Full suite: **36
passed**, no GPU or network.
- `gate_verbosity.py` validated end-to-end against a real day-0 run's
artifacts: reproduces the hand-computed result (`max_abs_delta =
0.0171`, pass) and correctly marks the two unequal-sample tasks
`not_comparable`.
- `pre-commit run --files <changed>` clean, including ruff, mypy,
bandit, markdownlint, and the `.claude/skills` symlink sync.
- Verified `max_sample_length` is a real `get_dataset_dataloader`
parameter with default 512, matching `--calib_seq`'s default, so
existing `ptq()` callers are unaffected.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `--calib_seq` defaults to the
pre-existing 512; `gate_ptq.py` reclassifies size growth from
`QUANT_COVERAGE_FAILURE` to `SIZE_NOT_REDUCED`, which is a more precise
class for an already-4-bit source and is covered by a new test asserting
a real coverage failure still outranks it.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependencies, stdlib only.
- Did you write any new necessary tests?: ✅ — 8 new tests for the new
gate.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
The measured figures come from a completed day-0 NVFP4 release
qualification. Model-specific results were removed from the general
skills where the rule stands on its own; two references were kept
deliberately — a model card citation illustrating per-scenario sampling,
and a model-specific vLLM MoE kernel crash where the model name *is* the
evidence.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added configurable calibration sequence-length limits for DeepSeek-V4
quantization.
* Added automated release checks for output verbosity, evaluation
comparability, serving readiness, and configuration parity.
* **Bug Fixes**
* Improved quantization size-ratio reporting by distinguishing
explainable growth from blocking failures.
* Clarified handling of deployment memory-access errors and infeasible
evaluations.
* **Documentation**
* Expanded guidance for calibration, remote execution, workspace
management, evaluation setup, deployment troubleshooting, and
statistical reliability.
* Added task-specific guidance for SciCode and GDPVal feasibility
checks.
* **Chores**
* Excluded workspace session directories from version control.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
d73278808b |
Bump nemo container requirement to 26.08 for MBridge examples (#2257)
### What does this PR do? Type of change: Bug fix Bumps the Megatron-Bridge examples, tests and launcher configs to `nemo:26.08` and removes the version-gated fallbacks they carried, plus the fixes needed to make the suites green on that container. **26.08 bump and shim removal** - Examples, CI workflows, `noxfile.py` and the `mbridge_*` launcher configs move to `nemo:26.08`. - `examples/megatron_bridge/_distillation_provider.py` is deleted — 26.08's Megatron-Bridge ships `convert_to_distillation_provider(..., distill_submodule=...)` natively, so `distill.py` imports it directly. - `prune_minitron.py` drops the `AutoBridge.from_hf_config` / config-only-export probing; `--no_moe_grouped_gemm` is no longer needed in the MoE pruning tests, and the Qwen3.5-MoE `skipif` is gone (native MoE expert mappings are in 26.08). - `_DynamicMambaMixer` targets only the raw `conv1d_weight` / `conv1d_bias` parameters that replaced the `conv1d` module in Megatron-Core. **MambaModel / MambaModelProvider removal** Megatron-Core has shipped `HybridModel` since 26.06 and `MambaModel` is a deprecated subclass that shares its `forward`, so `DMRegistry` resolves those instances to the `HybridModel` registration and the separate entry is redundant. Same for `MambaModelProvider` vs `HybridModelProvider` on the bridge side. `MambaMixer` / `MambaLayer` / `ExtendedRMSNorm` are untouched — the layers still exist. The deprecated `get_te_mamba_stack_spec` is removed; use `get_te_hybrid_stack_spec`. **Bug fix: compressed output_layer extra state** `mtq.compress` converts even a *disabled* `output_layer` into a `RealQuantLinear` (its weight is left uncompressed, since `pack_real_quantize_weight` skips disabled quantizers). The guard added in #2112 excluded `RealQuantLinear` by class, so the layer emitted extra state and every worker died in `GPTModel.sharded_state_dict`: ``` RuntimeError: Boolean value of Tensor with more than one value is ambiguous megatron/core/models/gpt/gpt_model.py:896, in sharded_state_dict output_extra_state and output_extra_state.data ``` The guard now keys off whether the weight was actually compressed (`QTensorWrapper`) instead of the class. This took out all 12 `test_homogeneous_compressed_sharded_state_dict` params, and the crashed workers poisoned the pool, which surfaced as unrelated timeouts and NCCL errors in `test_layer_sync_moe_local_experts_amax`, `test_kv_cache_quant`, `test_kv_cache_amax_sync`, `test_convert_mcore_te_gpt_model` and `test_homogeneous_sharded_state_dict_te_spec` — 21 tests in total. The e2e coverage is `skip_flaky_on_blackwell`, so CI never ran it; `test_output_layer_extra_state_empty_when_nothing_quantized` now asserts the contract directly and is not skipped. **Checkpoint import entry point** 26.08 replaced `examples/conversion/convert_checkpoints.py` with `scripts/conversion/convert.sh`, so `tools/launcher/common/megatron_bridge/import/import.sh` and the three README snippets are retargeted. `import.sh` uses the distributed GPU backend with `GPUS_PER_NODE` / `TP` / `PP` / `EP` knobs. **Megatron-LM on nemo:26.06** keeps working: `_get_mamba_conv1d` still dispatches between the `conv1d` module (26.06 and earlier) and the raw parameters (26.08+), so `import_mcore_gpt_from_hf` / `export_mcore_gpt_to_hf` handle NemotronH on both. Only the Megatron-Bridge examples and Minitron pruning of Mamba/hybrid models require 26.08. **Test consolidation** `test_export_distilled_megatron_to_hf.py` is merged into `test_distill.py`: `test_distill_llm` becomes `test_distill_llm_hf_export` and covers the standalone `--export_iterations all` run on the checkpoints it already produces, saving one full distillation (~185 s of CI time). The two mamba-named gpu test files are renamed to `hybrid`. ### Usage ```bash # HF -> Megatron import, via Megatron-Bridge's 26.08 conversion entry point bash /opt/Megatron-Bridge/scripts/conversion/convert.sh import \ --executor local \ --device gpu \ --gpus-per-node 8 \ --hf-model Qwen/Qwen3-8B \ --megatron-path /tmp/Qwen3-8B-megatron ``` ### Testing All on `nvcr.io/nvidia/nemo:26.08`, 2x RTX 6000 Ada, no timeout overrides: - `tests/examples/megatron_bridge`: 16 passed, 1 skipped (28m14s). The skip is the `gemma3vl` QAD param, now `@pytest.mark.manual` since `qwen3_5_moe_vl` covers the VLM QAD path. - `tests/gpu_megatron` (`_extensions`, `distill`, `export`, `opt`, `peft`, `sparsity`, `speculative`, `utils`): 61 passed, 5 xpassed. - `tests/gpu_megatron/torch/export` re-run after the conv1d dispatch change: 27 passed. - The 21 previously failing/hanging quantization tests: 21 passed (12 + 9). - `tests/gpu_megatron/torch/{nas,prune}`: verified separately. `import.sh` equivalence on a toy `qwen3_moe`, comparing all 12 weight tensors after flattening each dist checkpoint with `dcp_to_torch_save` — the GPU backend at 1 GPU, `--tp 2`, `--pp 2`, `--ep 2`, and `import.sh` end-to-end (`GPUS_PER_NODE=2 EP=2`) are all byte-identical to `--device cpu`. `nemo:26.06` compatibility was checked directly in that image: `megatron.core.models.hybrid.HybridModel`, the modelopt hybrid spec and `hybrid_layer_pattern` are all present, while `megatron.bridge.models.hybrid` and the bridge's `distill_submodule` are not. The NemotronH round-trip test failed there before the conv1d dispatch was restored and the dispatch is back in place; per project convention the suites themselves only run on 26.08. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ⚠️ Megatron-Bridge examples plus Minitron pruning of Mamba/hybrid models now require `nemo:26.08`. Megatron-LM quantization and checkpoint export still run on `nemo:26.06`. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — `test_output_layer_extra_state_empty_when_nothing_quantized` for the compress fix; existing tests extended for the merged export coverage. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added guidance for importing Hugging Face checkpoints into Megatron distributed format. * Expanded distillation workflows to export selected or all checkpoint iterations. * **Improvements** * Expanded Hybrid model support across Megatron workflows. * Updated distributed import tooling with GPU and parallelism options. * Updated supported environments and examples to NVIDIA NeMo 26.08. * **Bug Fixes** * Corrected output-layer quantization state handling when quantization is disabled. * **Documentation** * Added compatibility guidance for current and legacy NeMo containers. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
449a39922b |
Pin nemo_automodel below 0.6 for the fastgen example (#2260)
### What does this PR do? Type of change: Bug fix `nemo_automodel` 0.6.0 removed `nemo_automodel.recipes.diffusion.train.is_main_process` without a replacement (it was a three-line rank-zero predicate in 0.5.0, and 0.6.0 defines no equivalent anywhere in the package). `examples/diffusers/fastgen/dmd2_recipe.py` imports it, so the example's import guard fires and **every** test in `tests/examples/diffusers/` errors at collection: ``` ImportError: cannot import name 'is_main_process' from 'nemo_automodel.recipes.diffusion.train' tests/examples/diffusers/fastgen/test_resume_dataloader.py E ImportError: The DMD2 fastgen example requires `nemo_automodel`. ... collected 42 items / 1 error ``` The requirement was `>=0.4.0,<1.0`, so CI picked 0.6.0 as soon as it was published and the `onnx (diffusers)` job started failing on every PR (e.g. runs 33020467654, 33019418460, 33010815298, 33007613265, 33006944292 — all unrelated branches). Capping at `<0.6` restores the tested range. Every other `nemo_automodel` symbol the example imports still exists in 0.6.0 (`_diffusers.auto_diffusion_pipeline.NeMoAutoDiffusionPipeline`, `recipes.diffusion.train.TrainDiffusionRecipe`, and the four `components.datasets.diffusion.*` helpers), so `is_main_process` is the only blocker; the alternative is defining that predicate locally and widening the cap again, which is worth doing separately if the example is meant to track 0.6. ### Usage ```bash pip install -r examples/diffusers/fastgen/requirements.txt ``` ### Testing Reproduced the break by diffing the published wheels: `is_main_process` is defined at `nemo_automodel/recipes/diffusion/train.py:692` in 0.5.0 and absent from 0.6.0 (`grep -rn "def is_main_process"` over the unpacked 0.6.0 wheel returns nothing). Confirmed the remaining imported symbols are all still present in 0.6.0. CI on this PR exercises the fix directly: the `onnx (diffusers)` job installs from this requirements file and is the job that has been failing. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — existing dependency, tightened bound. - Did you write any new necessary tests?: N/A — the existing `tests/examples/diffusers/` suite is what this unblocks. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — dependency-pin fix for a break introduced and fixed within the same unreleased cycle. - Did you get Claude approval on this PR?: ❌ 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed dependency compatibility for the FastGen diffusion example. * Prevented installation of versions that could cause the example to fail at startup. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |