Commit Graph
710 Commits
Author SHA1 Message Date
Keval MorabiaandClaude Sonnet 4.6 e56682e34a docs: update installation pages with legal-approved license notices (#1322)
## Summary

- Replaces the old pip license notice ("Please review the license terms
of ModelOpt and any dependencies before use") with the Legal-approved
wording: "Model Optimizer will download and install additional
third-party open source software projects. Review the license terms of
these open source projects before use."
- Adds a generic container license review notice ("Before pulling and
using the container images, please review their respective license
terms.") to the Linux installation doc (Docker tab) and README.
- Adds a `.. note::` with the pip notice to the Windows installation
page (covers both standalone and Olive child pages).
- Expands the README container section to explicitly list all four
recommended NVIDIA container images (`pytorch`, `nemo`, `tensorrt-llm`,
`tensorrt`).

## Test plan

- [x] Verify rendered docs look correct (`nox -s docs`)
- [x] Confirm legal notices appear in Linux, Windows, and README install
sections

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated installation guides with explicit references to supported
NVIDIA container images (PyTorch, NeMo, TensorRT-LLM and variants),
clarified pre-installed Model Optimizer in some images, and added notes
to review each container’s license terms; clarified conditional
environment setup wording and local install license guidance.
* **Chores**
  * Updated project license header year.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-22 22:49:23 +05:30
kinjalpatel27 0678136335 Fix vLLM fakequant MoE megatron export bug (#1305)
### What does this PR do?

Type of change: Bug fix                     
Fixes two bugs in the vLLM + Megatron-Core MoE export path and cleans up
the related weight-collection helper:
    
1. **`_QuantFusedMoEBase` (vllm.py)**: The weight-quantizer path in
`_invoke_fused_moe_quantized_function` was temporarily mutating
`self.w13_weight` / `self.w2_weight` to the quantized tensor, then
restoring them via `finally`. This exposed a stale quantized tensor on
`self` between the mutation and the kernel call. Fixed by computing the
quantized weight directly into a local `B` without touching `self.*`
attributes.
2. **`GPTModelExporter` / `VllmFqGPTModelExporter`
(unified_export_megatron.py / vllm_fakequant_megatron.py)**:
`expert_bias` (present in grouped MoE layers) was silently dropped
during export because the bias collection ran after the early-return on
missing `weight`. Extracted a `_get_weight_bias` helper that collects
weight, bias, and expert_bias together, so bias/expert_bias are captured
even when weight is absent or zero-element.

### Usage

```python
# No API change; export pipelines pick this up automatically.
# export_mcore_gpt_to_hf_vllm_fq / export_mcore_gpt_to_hf now correctly                   
# export expert_bias for grouped-MoE checkpoints.    
```

### Testing
Step 1 — Quantize (run from Megatron-LM
examples/post_training/modelopt):
```
HF_MODEL_CKPT=<path/to/hf/weights> MLM_MODEL_SAVE=<quant-ckpt-name> \
bash quantize.sh <hf-model-id> NVFP4_DEFAULT_CFG
```
Step 2 — Export for vLLM fakequant:
```
MLM_EXTRA_ARGS=--export-vllm-fq \ 
HF_MODEL_CKPT=<path/to/hf/weights> \ 
MLM_MODEL_CKPT=<quant-ckpt-name> \ 
EXPORT_DIR=<export-dir> \ 
bash export.sh <hf-model-id> 
```
Step 3 — Serve (run from examples/vllm_serve):
```
 QUANT_CFG=NVFP4_DEFAULT_CFG \ 
 QUANT_FILE_PATH=<export-dir>/quantizer_state.pth \ 
 python3 vllm_serve_fakequant.py <export-dir> \ 
 -tp 1 --served-model-name <model-name> \ 
 --host 0.0.0.0 --port 8000 \ 
--trust-remote-code --enforce-eager \ 
 --disable-custom-all-reduce \ 
--gpu-memory-utilization 0.8          
```
### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Refactor**
* Centralized weight/bias/expert-bias extraction and export to a single
helper for consistent handling.
* Standardized quantized-weight flow to temporarily swap and restore
parameter tensors during computation.

* **Bug Fixes**
* Prevented missing or incorrect weight/bias exports by unifying
extraction logic.
* Broadened checkpoint key matching to preserve more quantizer state
during reloads.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-04-22 09:24:25 -07:00
nv-samcheng c417e6f4d9 Exclude small-k and small-n Matmul nodes from Int8 quantization (#1256)
### What does this PR do?
Exclude small-dimension MatMul nodes from INT8 quantization. MatMuls
with N or K < 16 cannot efficiently use INT8, causing performance
regressions.



### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved quantization exclusions so MatMul/Gemm ops with derived K<16
or N<16 are skipped, honoring Gemm transB, using inferred and
runtime-determined shapes, and avoiding duplicate outputs.

* **Tests**
* Expanded unit tests to cover constant, inferred, and runtime-derived
shapes, Gemm transB behavior, small-dimension edge cases, and output
deduplication.

* **Documentation**
* Added changelog entry documenting the new small-dimension exclusion
thresholds and transB handling.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: samcheng <samcheng@nvidia.com>
2026-04-22 15:47:43 +05:30
Keval Morabia 785d3a2df6 [CI] Bump test containers to latest (#1299)
- Use latest containers for testing in CICD

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Bumped TensorRT-LLM Docker images to 1.3.0rc12 in example and GPU test
workflows.
  * Updated PyTorch container image from 26.01 to 26.03 for GPU tests.
* Captured uv lock upgrade output to a temp file, inlined it into PR
bodies, and adjusted workflow heredoc/templating and step behavior.

* **Documentation**
* Clarified an inline comment and simplified a warning message for an
ONNX quantization extension.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-21 23:44:12 -04:00
Keval Morabia c51c1762b3 fix: prevent gh-pages repo bloat from doc preview artifacts (#1309)
### What does this PR do?

Type of change: Bug fix

Fixes gh-pages branch bloat that grew from ~26 MB to ~441 MB in four
weeks (nvbug 6099503). Three compounding causes were identified and
addressed:

1. **Sphinx `.doctrees/` cache published to gh-pages** — `sphinx-build`
was writing its build cache inside `build/html/` which was then uploaded
verbatim. Accounts for ~3.3 GB uncompressed across history.
2. **`JamesIves/github-pages-deploy-action` appending a commit on every
push** — main-site files accumulated forever with `single-commit: false`
(default).
3. **PR preview deploying on every `synchronize` event for all PRs** —
`rossjrw/pr-preview-action` re-deployed the full site for every push to
any PR regardless of whether docs changed (e.g. PR #1128 triggered 64
preview deploys × ~11 MB each).

Changes:
- Pass `-d /tmp/doctrees` to `sphinx-build` so `.doctrees/` is never
written into `build/html/`
- Add `paths: [docs/**, modelopt/**]` filter to `pull_request` trigger
so the docs workflow only runs on PRs that touch docs or source code
- Set `single-commit: true` on the deploy action so main-site pushes
squash into one commit
- Deduplicate docs build: `deploy-preview` now downloads the artifact
from `build-docs` instead of running a second `sphinx-build`
- Set `retention-days: 1` on the artifact since it is only needed for
the duration of the workflow run

The one-time cleanup (force-push squashed orphan to gh-pages) was
already applied separately — repo is now ~59 MB for a full clone vs ~441
MB before.

### Usage

N/A — CI/workflow change only.

### Testing

- Workflow logic reviewed manually.
- The one-time cleanup was verified: `git rev-list --objects
--disk-usage origin/gh-pages` now reports ~28 MB; full clone is ~59 MB.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information

nvbug 6099503

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Optimized documentation build and deployment workflow in CI/CD
pipeline.
* Improved pull request documentation preview handling with faster build
timeouts and refined artifact management.
* Enhanced GitHub Pages deployment configuration for better consistency.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-21 20:19:08 +00:00
sychen52 5ffb8487d9 add gptq fused kernel (#1291)
### What does this PR do?

Add gptq fused kernel to improve speed.

### Usage

check unittest

### Testing
added a unittest

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Fused GPTQ backend for faster blockwise weight updates, toggleable via
a new "fused" option.
  * Shared NVFP4 quantization primitives exposed for reuse.

* **Refactor**
* Consolidated FP4 scale/quantization logic into reusable utilities and
centralized Hessian inversion handling.

* **Tests**
* Expanded GPU tests comparing fused vs unfused GPTQ, added
Triton-availability gating and a local benchmark entrypoint.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
2026-04-21 10:50:04 -07:00
yeyu-nvidiaandClaude Opus 4.6 2fef374ded fix: auto-compute dp_replicate_size from world_size (#1302)
## Summary
- When `dp_shard_size < world_size` (e.g., `dp_shard_size=4` on 8 GPUs
across 2 nodes), `ParallelismConfig` raises `total_size (4) does not
match num_processes (8)` because `dp_replicate_size` defaults to 1
- Auto-compute `dp_replicate_size = world_size // (dp_shard_size *
cp_size)` so intra-node FSDP2 sharding + inter-node data-parallel
replication works without manual config
- This enables `dp_shard_size` to be set to per-node GPU count (better
NVLink utilization) while automatically creating replicas across nodes

## Test plan
- [ ] Verify single-node training (dp_shard_size == world_size,
dp_replicate_size == 1) unchanged
- [ ] Verify multi-node with dp_shard_size < world_size creates correct
replica groups
- [ ] Verify existing EAGLE3/DFlash configs still work

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Enhanced parallelism configuration initialization in the speculative
decoding example to better handle distributed training scenarios.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-20 20:39:36 +00:00
Chenhan D. YuandClaude Opus 4.6 355c6b7883 fix: PTQ 1GPU, export PP divisibility, hidden states conversations key (#1293)
## Summary
- **megatron_lm_ptq.yaml**: Qwen3-8B PTQ to single GPU for L40 clusters
(TP=1, all tasks)
- **quantize.sh**: Auto-find largest PP dividing model's
`num_hidden_layers` for export step. Qwen3-8B has 36 layers which isn't
divisible by 8, causing `AssertionError` on 8-GPU nodes
- **compute_hidden_states_trtllm.py**: Use `messages` with
`conversations` fallback, matching the HF version. Fixes `KeyError:
'conversations'` when data uses OpenAI `messages` format

## Test plan
- [x] Qwen3-8B PTQ runs on single L40 GPU
- [x] Export PP auto-selects valid divisor (36 layers → PP=6 on 8 GPUs,
PP=4 on 4 GPUs, PP=1 on 1 GPU)
- [x] EAGLE3 offline pipeline reads data with `messages` field

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Dataset input handling now supports multiple field formats for
enhanced compatibility.

* **Bug Fixes**
* Optimized GPU resource allocation during model quantization with
improved pipeline parallelism computation.
* Updated quantization configuration for more efficient resource
utilization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
0.45.0dev
2026-04-20 12:43:58 -07:00
yeyu-nvidiaandClaude Opus 4.6 289a239ca5 fix: use data_dir for directory paths in ShardedDataset (#1301)
## Summary
- `datasets`' `resolve_pattern` only matches entries with
`type=="file"`, so passing a bare directory path as `data_files` to
`load_dataset` results in `FileNotFoundError` even when the directory
exists on disk
- Detect directory paths in `ShardedDataset._load_dataset()` and pass
them via `data_dir` instead of `data_files`

## Reproduction
```python
from datasets import load_dataset
# This fails with FileNotFoundError:
load_dataset("json", data_files="/path/to/data_directory")
# This works:
load_dataset("json", data_dir="/path/to/data_directory")
```

## Test plan
- [ ] Verify existing EAGLE3/DFlash training pipelines that pass
directory paths work
- [ ] Verify file path and glob patterns still work (falls through to
`data_files`)
- [ ] Verify `data_files=None` (no data_files arg) still works

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Bug Fixes

* Fixed an issue with dataset loading that prevented proper handling of
directory-based data sources. Directories are now correctly detected and
processed during dataset initialization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-20 19:10:02 +00:00
Frida HouandKeval Morabia 97d153118e [minor] Add custom calibration backend registry (#1281)
### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a public backend-specific calibrator registration API to support
FP8 scale-sweep calibration, allowing backends to supply custom
calibrators used during FP8 tuning.

* **Tests**
* Added unit tests confirming registry insertion/overwrite, that
registered calibrators are invoked when FP8 scale-sweep is enabled, are
not invoked when disabled, and that calibration falls back to defaults
when no backend is registered.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-20 09:24:17 -07:00
kinjalpatel27 010b220dc0 vLLM fakequant export update for AWQ checkpoint (#1242)
### What does this PR do?

Type of change: Bug

Enables end-to-end AWQ checkpoint export and reload in the vLLM
fake-quant serving path (`MODELOPT_STATE_PATH`). Previously, the
`input_quantizer` was using incorrect `pre_quant_scale` especially with
grouped quantizers like `qkv_proj`, using simply the first
`input_quantizer.pre_quant_scale`. This MR adds
`_resmooth_experts_for_export` that non-mutatively averages
`pre_quant_scale` across MoE experts and unifies input `_amax`, required
because vLLM uses a single input quantizer per expert group. Adds
`merge_amax_tensors_for_group` (element-wise max for same-shape, `cat`
for GQA, scalar-max fallback) replacing the scalar-collapsing
`torch.stack().max()` that dropped per-channel `_amax` structure.

### Usage

```python
# Export AWQ checkpoint from HF model
  from modelopt.torch.export.plugins.vllm_fakequant_hf import export_hf_vllm_fq_checkpoint
  export_hf_vllm_fq_checkpoint(model, export_dir="./awq_vllm_checkpoint")      
```

### Testing
**Step 1 — Export the quantized checkpoint:**
  ```bash                    
  python examples/llm_ptq/hf_ptq.py \
--pyt_ckpt_path <MODEL_PATH> \
    --recipe <AWQ_RECIPE> \
--calib_size 512 \
--export_path <EXPORT_DIR> \
    --vllm_fakequant_export
```
  This produces `<EXPORT_DIR>/vllm_fq_modelopt_state.pth` with the averaged per-expert                                                                                                                           
  pre_quant_scale and unified _amax now included.                                                                                                                                                              
   

 Step 2 — Serve via vLLM fakequant worker:                                                                                                                                                                    
```bash
  MODELOPT_STATE_PATH=<EXPORT_DIR>/vllm_fq_modelopt_state.pth \
python examples/vllm_serve/vllm_serve_fakequant.py \
      <EXPORT_DIR> --tensor-parallel-size <TP>   
```

Tested for quantization configurations:
```
FP8_DEFAULT_CFG
FP8_DEFAULT_CFG (input_q disabled)
INT8_SMOOTHQUANT_CFG
INT8_WEIGHT_ONLY_CFG
NVFP4_DEFAULT_CFG
NVFP4_AWQ_LITE_CFG
INT4_AWQ_CFG
NVFP4_AWQ_CFG
NVFP4_DEFAULT_CFG (input_q disabled)
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A 
- Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A 

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit

* **New Features**
  * Added Nemotron-style MoE export support and group-aware AWQ resmoothing with optional requantization during export.
  * Improved handling for shared-input / expert groups and tensor-parallel sharding of pre-quantization scales.

* **Bug Fixes**
  * Removed AWQ reload limitation from known issues; improved checkpoint validation and safer save/load behavior.
  * Better detection and handling of enabled weight-quantizers and clearer warnings for mismatched checkpoint keys.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-04-19 22:12:45 -07:00
jingyu-ml 26ae8da517 [2/3] Implicit Gemm NVFP4 (#1227)
### What does this PR do?

Type of change: new feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

- Add Conv3D implicit GEMM kernel with BF16 WMMA tensor cores and fused
NVFP4 activation quantization for video diffusion VAE layers
- Integrate into _QuantConv3d via QuantModuleRegistry — automatically
dispatched when NVFP4 quantization is applied to nn.Conv3d
- Move kernel from `experimental/conv/ to modelopt/torch/kernels/conv/`;
move tests to `tests/gpu/torch/quantization/kernels/`

### Testing
<!-- Mention how have you tested your change if applicable. -->

- Added test cases to measure the difference between cuDNN and our CUDA
implicit GEMM kernel
- Added an NVFP4 fake quantization test using CUDA code

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Per-backbone quantization/export in a single run with per-backbone
checkpoints and backbone-aware quant filters
* Configurable NVFP4 block-size via CLI/config; improved NVFP4 Conv3D
inference path and Wan 2.2 quantization support
* **Bug Fixes**
* Video-model calibration now respects extra params and forces video
decoding during calibration
* **Documentation**
* Added comprehensive Conv3D implicit‑GEMM kernel documentation; removed
experimental Conv3D prototype docs/benchmark
* **Tests**
* New Wan 2.2 quantization/export tests and expanded Conv3D/FP4 kernel
test coverage
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
0.44.0rc0
2026-04-19 12:20:14 +05:30
kaix-nv c20f9c411d Add a standalone monitor skill for persistent job tracking (#1252)
### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->
Add a standalone monitor skill for persistent job tracking across
sessions, and integrate it with PTQ, evaluation, and deployment skills.

Problem: Each skill had ad-hoc inline monitoring (squeue polling, nel
status checks) that didn't survive session restarts and couldn't track
multiple jobs. Users had to manually ask "check status" every time.

Solution: A centralized monitor skill with:
- Job registry (.claude/active_jobs.json): single source of truth for
all active jobs
- Durable recurring cron: polls every 15 min, survives session restarts,
self-cleans when all jobs complete
- User-initiated mode: works in new conversations by reading the
registry
- Aggregated reporting: "2 of 4 completed" instead of per-job noise

### Usage
After any skill submits a job, the monitor skill automatically:

1. Registers the job in .claude/active_jobs.json
2. Sets up a durable cron to poll status every 15 minutes

User can also trigger manually:
User: "check my eval status" → reads registry, reports current state
User: "is the PTQ done?"            → finds job, checks status
User: "what jobs are running?"      → lists all registered jobs

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added monitor skill for tracking SLURM jobs, NEL evaluations, and
launcher experiments with persistent job registry.

* **Documentation**
* Updated deployment, evaluation, and PTQ documentation to use the new
monitor skill.
  * Simplified diagnostic and troubleshooting instructions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kai Xu <kaix@nvidia.com>
2026-04-19 04:44:17 +00:00
github-actions[bot]andgithub-actions[bot] e9a49890f1 [chore]: weekly bump of uv.lock on main (2026-04-18) (#1292)
## Summary
Automated weekly update of uv.lock file for nSpect Scanning:
- `uv.lock` — upgraded all transitive dependencies to latest compatible
versions

Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-04-18 17:22:23 +00:00
Keval MorabiaandClaude Sonnet 4.6 3d0f0db49e [CI] Replace tox with nox, use nemo:26.04 for megatron tests, and simplify CI workflows (#1286)
### What does this PR do?

Type of change: New feature / infrastructure improvement

Follow-up to #1285 for correct CI test environment for megatron based
tests

Replaces `tox` + `tox-current-env` with `nox` for all test, lint, docs,
and wheel build sessions. The primary motivation was that
`tox-current-env` is incompatible with uv venvs in NGC containers (e.g.
NeMo's `/opt/venv`) — it picks the system Python via
`sys._base_executable` instead of the container's venv Python which has
megatron packages pre-installed.

Key changes:
- **`noxfile.py`** replaces `tox.ini` with GPU, CPU unit,
partial-install, pre-commit, docs, and wheel sessions
- **GPU sessions** use `venv_backend="none"` (run directly in container
env) and `python -m pip/pytest` to avoid PATH mismatches
- **uv** is set as the default venv backend (if available) for CPU
sessions (faster installs)

Also includes CI workflow simplifications:
- **`_pr_gate.yml`** new reusable workflow centralizing file-change
detection + linux-check wait logic (was duplicated across 3 workflow
files)
- **Collapsed pr/non-pr job pairs** into single jobs with conditional
`runs-on` in `gpu_tests.yml`, `example_tests.yml`,
`regression_tests.yml`
- **Collapsed `multi-py` / `multi-torch` / `multi-transformers`** into a
single `multi-version` matrix job in `unit_tests.yml`
- **PR path filtering** for unit test secondary jobs (multi-version,
launcher, partial-install) — skipped if no relevant files changed
- **Fixed schedule/workflow_dispatch skipping** — jobs with `needs:
[pr-gate]` were incorrectly skipped when all pr-gate internal jobs were
skipped; fixed by making the gate job always run
- **multi-version, launcher, partial-install** now also run on
`schedule` / `workflow_dispatch`

### Usage

```bash
python -m pip install nox uv                                                    # install nox and uv (once)
nox -l                                                                          # list all sessions
nox -s gpu_megatron                                                             # run a GPU session (inside container)
nox -s "unit-3.12(torch_211, tf_latest)"                                        # run a specific unit test combination
nox -s "unit-3.12(torch_211, tf_latest)" -R                                     # force-recreate venv (e.g. after dep changes)
COVERAGE_PROCESS_START=pyproject.toml nox -s "unit-3.12(torch_211, tf_latest)"  # with coverage
```

### Testing
- Ran `nox -l` to verify all session names
- Ran `gpu_megatron` session locally inside NeMo container — confirmed
it uses `/opt/venv/bin/python` correctly
- Manually triggered nightly-runs:
- Unit:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608013657
- GPU:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608018763
- Examples:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608017322

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A — CI infrastructure only
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ (added `nox`
and `uv` to `dev-test`, both Apache-2.0)
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — no user-facing changes

### Additional Information
Supersedes the tox-current-env workaround in the parent branch.

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-18 16:56:02 +00:00
Keval Morabia 2b315eda47 Replace mip package with pulp (#663)
Replace mip package with more popular pulp package for puzzle mip
solving. Both use the CBC solver under the hood

## Testing

- Results very close for Qwen3-8B and Nemotron-Nano-12B-v2

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Simplified GPU test environment setup by removing unnecessary system
dependency installation
* Updated internal optimization solver dependencies in the puzzletron
module

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-18 21:19:36 +05:30
Farid AdilazuardaandKeval Morabia 2004779a67 Update README.md for DMS (fix cd experimental/DMS to cd Model-Optimizer/experimental/DMS) (#879)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** ?

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated DMS installation instructions to reflect the repository
structure and correct directory navigation during setup.
* Clarified the setup steps so users follow the accurate directory
change before running installation commands.
* Small wording improvements to reduce confusion during the installation
process.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Farid Adilazuarda <42537562+faridlazuarda@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-18 14:55:42 +00:00
760c980727 Add ResNet50 support for torch_onnx quantization workflow (#1263)
## Summary
- Add end-to-end ResNet50 support in the torch_onnx quantization → ONNX
export → TRT engine pipeline
- Fix multiple Conv2d-related export issues that blocked Conv2d-heavy
models from working with FP8/INT8/MXFP8/NVFP4/auto quantization modes
- Fix `configure_linear_module_onnx_quantizers` to handle all modules
with block quantization (not just `nn.Linear`), fixing NVFP4/MXFP8
export for models with quantized non-Linear modules
- Add `--trt_build` flag to `torch_quant_to_onnx.py` and simplify test
infrastructure

### Files Changed
- `modelopt/torch/_deploy/utils/torch_onnx.py` — Disable FP8 Conv2d
weight quantizers and autocast during ONNX export
- `modelopt/torch/quantization/export_onnx.py` — Fix
`configure_linear_module_onnx_quantizers` for all module types with
block quantization
- `examples/torch_onnx/torch_quant_to_onnx.py` — Add `--trt_build` flag,
calibration for FP8 override quantizers, Conv2d→FP8 override for auto
mode, filter_func updates
- `examples/torch_onnx/README.md` — Add ResNet50 to supported models
table
- `tests/examples/torch_onnx/test_torch_quant_to_onnx.py` — Add ResNet50
test entry, simplify using `--trt_build`
- `tests/_test_utils/torch/vision_models.py` — Add ResNet50 to timm
model registry

### Quantization modes passing
- ✅ FP8, INT8, MXFP8, NVFP4, Auto (all 5 modes pass export + TRT build)
- INT4_AWQ excluded (pre-existing limitation for all models)

## Test plan
- [x] All 5 resnet50 test modes pass: `pytest
tests/examples/torch_onnx/test_torch_quant_to_onnx.py -k resnet50` (5/5
passed)
- [x] Full regression: 18 passed, 2 failed (pre-existing swinv2_tiny
fp8/int8 failures)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added ResNet50 to supported ONNX export vision models with FP8, INT8,
MXFP8, and NVFP4 support.
  * Optional TensorRT engine build after export via a new CLI flag.

* **Improvements**
* Enhanced quantization calibration and export flows for FP8/INT8
models, including broader block-quantization support across module types
and safer export handling.
  * Tests updated to include ResNet50 in the model matrix.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Signed-off-by: ajrasane <arasane@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-18 14:30:45 +00:00
bkartal-dev 92622a9aa6 Add nvfp4_mse and nvfp4_local_hessian options to the ptq script. (#1113)
### What does this PR do?

Type of change: Bugfix

<!-- Details about the change. -->

Add newly added quant configs to the example PTQ script.

### Testing

I have locally run auto_quantize with these two quant_configs, and
obtained successfully exported HF artifacts.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added support for three new quantization formats: nvfp4_mse,
nvfp4_local_hessian, and nvfp4_experts_only, expanding available export
options when using auto-quantize.

* **Bug Fixes / UX**
* Updated the invalid-quantization error message to include the newly
accepted format identifiers.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Bilal Kartal <bkartal@nvidia.com>
Signed-off-by: bkartal-dev <bkartal@nvidia.com>
2026-04-18 07:28:49 +00:00
jingyu-ml feec81ad2b Add the Skip softmax for diffusion (#1166)
### What does this PR do?

Type of change: new feature, new example <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->

<!-- Details about the change. -->

## Summary

- Add skip-softmax sparse attention (BLASST) for diffusion models via
dedicated Triton kernels — an inference kernel with tile skipping and a
calibration kernel with vectorized multi-threshold sparsity measurement
- Add `triton_skip_softmax` method with exponential model calibration
(`scale_factor = a * exp(b * sparsity)`) and log-space fitting for
diffusion models
- Add Triton kernel backends for diffusers and LTX attention dispatch
- Fix calibration to skip RULER dataset generation when user provides
their own `forward_loop` (required for non-LLM models)

## Changes

### Triton kernels (`modelopt/torch/kernels/triton_fa.py`)
- **`_attn_fwd`**: Forward kernel with optional tile skipping — tiles
whose max attention score is far below the running softmax max are
skipped entirely (no V load, no softmax, no accumulation). Runtime
sparsity measurement via atomic counters.
- **`_attn_fwd_calibrate`**: Calibration kernel that computes full
attention while measuring how many tiles would be skipped at each of N
thresholds simultaneously. Uses per-program output buffers (zero atomic
contention) and vectorized multi-threshold comparison.
- **`attention()`** / **`attention_calibrate()`**: Python wrappers for
inference and calibration kernels.

### Kernel backends
(`modelopt/torch/sparsity/attention_sparsity/kernels/`)
- **`diffusers_triton_attention.py`**: Registers `modelopt_triton`
backend in diffusers' attention dispatch. Handles [B, S, H, D] → varlen
layout conversion, calibration/inference mode switching, thread-local
configuration, and counter accumulation.
- **`ltx_triton_attention.py`**: Patches `ltx_core.Attention` modules
for Triton dispatch with the same calibration/inference modes.

### Method
(`modelopt/torch/sparsity/attention_sparsity/methods/triton_skip_softmax.py`)
- `TritonSkipSoftmaxMethod`: Context managers for calibration (→
calibration kernel) and inference (→ forward kernel with tile skipping).
Three threshold priority levels: raw threshold > calibrated scale_factor
> static threshold.

### Calibration
(`modelopt/torch/sparsity/attention_sparsity/calibration/`)
- **`calibrator.py`**: `DynamicThresholdCalibrator` with `fit_logspace`
option — fits exponential model in log space (minimizes relative error)
for diffusion models where scale_factors span many orders of magnitude.
Records observed sparsity range for extrapolation warnings.
- **`calibrate.py`**: Skips RULER dataset when `forward_loop` is
provided; passes `fit_logspace` through from config.

### Config & conversion
- **`config.py`**: `CalibrationConfig.fit_logspace` field (default
False, recommended True for diffusion models).
`skip_softmax_raw_threshold` field for direct threshold mode.
- **`conversion.py`**: Auto-registers diffusers/LTX Triton backends on
`sparsify()`. Updated summary display.

### Example
- **`wan22_skip_softmax.py`**: End-to-end example for WAN 2.2 5B/14B
with baseline, raw-threshold, and calibrated modes. Supports runtime
sparsity reporting.

## Threshold modes

| Mode | How it works | Use case |
|------|-------------|----------|
| **Raw threshold** (`--raw-threshold -0.7`) | Passed directly to kernel
as `skip_threshold_log2` | Quick testing, sweeps |
| **Calibrated** (`--calibrate --target-sparsity 0.5`) | `scale_factor =
a * exp(b * target)`, then `threshold = scale_factor / seq_k` at runtime
| Production use with seqlen adaptation |
| **Static** (default `skip_softmax_threshold=0.1`) | `log2(lambda) *
sm_scale` | Fallback |

## Usage

```bash
# Fixed raw threshold (no calibration)
python examples/diffusers/sparsity/wan22_skip_softmax.py \
    --model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
    --raw-threshold -0.7 \
    --prompt "A cat playing piano" --output out.mp4

# With calibration (log-space fit for diffusion models)
python examples/diffusers/sparsity/wan22_skip_softmax.py \
    --model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
    --calibrate --target-sparsity 0.5 \
    --prompt "A cat playing piano" --output out.mp4

# Dense baseline for comparison
python examples/diffusers/sparsity/wan22_skip_softmax.py \
    --model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
    --baseline \
    --prompt "A cat playing piano" --output baseline.mp4
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added skip-softmax sparse attention support for Diffusers models,
enabling efficient video generation
* Added support for both eager and Triton attention backends for sparse
attention
* Added new example script for Wan 2.2 text-to-video generation with
sparse attention optimization

* **Documentation**
* Updated documentation with sparse attention configuration guide and
usage examples

* **Tests**
* Added comprehensive unit tests for kernel backend registration and
skip-softmax functionality
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-04-18 06:54:19 +00:00
Chenhan D. YuandClaude Opus 4.6 76b6fd51a5 fix: DFlash regression tests and vLLM server liveness (#1288)
## Summary
- **hf_online_dflash.yaml**: Add 100K-sample training config with
regression baselines (B200 loss curve),
`MAX_FINAL_LOSS`/`MIN_FINAL_ACC`/`MIN_ACCEPTANCE_LENGTH` thresholds,
vLLM nightly container for DFlash support
- **vllm_smoke_test.sh**: Parse acceptance length from vLLM server log
for regression check; `pip install pandas` workaround for broken nightly
container; capture server output to temp file
- **query.sh**: Detect vLLM server death during startup (PID liveness
check) + 600s timeout to prevent infinite polling that wastes GPU hours;
`pip install pandas` workaround
- Fix empty `environment:` key in DFlash YAML causing nemo_run
`ListParseError`

## Test plan
- [x] E2E pipeline passed on 8x B200 (training + vLLM smoke test + AR
eval)
- [x] Training regression: final loss 3.82 < 5.0, acc 0.20 > 0.15
- [x] vLLM acceptance length: 1.79 >= 1.4 threshold
- [x] AR evaluation: 2.02 overall on MT-Bench (8 categories)
- [x] Server liveness check prevents GPU waste on vLLM crash

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added optional regression validation for vLLM acceptance metrics
* Introduced configurable vLLM server startup timeout (default 600
seconds)

* **Improvements**
* Enhanced logging for vLLM server startup with progress tracking and
waited time reporting
* Faster detection of vLLM server process failures during initialization

* **Configuration Updates**
  * Increased training dataset size and logging granularity
  * Scaled tensor parallelism from 4 to 8 across multiple pipelines
  * Expanded PTQ quantization to multi-step pipeline
  * Added configurable training metric thresholds
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 11:41:17 +05:30
realAsma 2d868d3f1f Performant layerwise calibration for large models (#1251)
## Summary

Adds **performant layerwise calibration** for quantizing large models
(e.g. DeepSeek-R1 671B) that don't fit entirely on GPU. ([Example
commands](#example-commands))

1. **Performant calibration for large models** — Each decoder layer is
moved from CPU/disk to GPU (accelerate) or unsharded (FSDP2) **only
once** and kept on GPU for the entire calibration step. Previously,
every calibration batch triggered weight transfer for every layer —
O(num_batches) weight movements per layer. Now it is O(1) per layer.
This also means you can **increase batch size** since only one layer's
weights occupy GPU at a time — e.g. DeepSeek-R1 on a single node
(8×80GB) with `batch_size=16` and `gpu_max_mem_percentage=0.5`.
2. **Checkpoint save/resume** — Saves progress after each layer, so jobs
that exceed cluster time limits (e.g. 4-hour Slurm windows for 100+
layer MoE models) can resume from the last completed layer.
3. **Rename** `sequential_calibrate` → `layerwise_calibrate` for
clarity.

### Design details

The existing layerwise state machine (skip/run/capture) already
processes one layer at a time, but skip-mode layers still kept their
parameters in the ModuleList — so frameworks transferred all weights
every forward pass. This PR adds:
- **`_SkipLayer`**: replaces fully-calibrated layers with a
parameter-free dummy in the ModuleList, so framework hooks have nothing
to transfer
- **`persistent_materialization`**: keeps the active layer on GPU for
the entire calibration step, avoiding repeated offload/reload cycles

Checkpoint save is per-layer; restore is bulk — quantizer state and
weights for layers 0..K-1 are restored once at the end of calibration,
keeping the hot path fast.

### Example commands

**Qwen3-8B** (NVFP4+GPTQ, single GPU):
```bash
python hf_ptq.py \
    --pyt_ckpt_path Qwen/Qwen3-8B \
    --recipe nvfp4_gptq_sequential.yaml \
    --calib_size 64 \
    --batch_size 16 \
    --dataset cnn_dailymail \
    --export_path outputs/qwen3_8b_nvfp4_gptq_seq \
    --gpu_max_mem_percentage 0.5 \
    --use_seq_device_map \
    --vllm_fakequant_export
```

**DeepSeek-R1** (NVFP4 experts-only + FP8 KV, 8×80GB):
```bash
python hf_ptq.py \
    --model unsloth/DeepSeek-R1-0528-BF16 \
    --recipe ../../modelopt_recipes/general/ptq/nvfp4_experts_only-fp8_kv.yaml \
    --dataset cnn_dailymail \
    --batch_size 16 \
    --calib_size 64 \
    --calib_seq 512 \
    --gpu_max_mem_percentage 0.5 \
    --use_seq_device_map \
    --trust_remote_code \
    --export_path output/DeepSeek-R1-BF16-nvfp4-experts-only-fp8-kv \
    --vllm_fakequant_export
```

### Example: NVFP4+GPTQ layerwise calibration on Qwen3-8B (36 layers,
single GPU — 20 GB peak)

**Initial run** (killed after layer 11):
```
Layerwise calibration: Found 36 transformer layers
Calibrating layer 1/36 | capture: [1]
Computing Hessians for 7 linear layers...
GPTQ time: 51.39s
Calibrating layer 2/36 | run: [1] | capture: [2]
Checkpoint: saved layer 0
GPTQ time: 50.06s
Calibrating layer 3/36 | skip: 1 | run: [2] | capture: [3]
Checkpoint: saved layer 1
...
Calibrating layer 12/36 | skip: 10 | run: [11] | capture: [12]
Checkpoint: saved layer 10
<killed>
```

**Resumed run** (picks up from layer 11, finishes all 36):
```
Layerwise calibration: Found 36 transformer layers
Checkpoint: resuming layerwise calibration from layer 11/36
Calibrating layer 12 (resumed)
GPTQ time: 51.45s
Calibrating layer 13/36 | skip: 11 | run: [12] | capture: [13]
Checkpoint: saved layer 11
...
Calibrating layer 36/36 | skip: 34 | run: [35] | capture: [36]
Checkpoint: saved layer 34
GPTQ time: 50.33s
Checkpoint: saved layer 35 (final)
Checkpoint: restored 11 previously calibrated layers
Layerwise calibration completed
Quantized model exported to: outputs/qwen3_8b_nvfp4_gptq_seq
GPU 0: Peak memory usage = 20.42 GB
```

## TODO
- [ ] Update CHANGELOG

## Test plan
- `tests/unit/torch/quantization/test_layerwise_calibrate.py` — unit
tests for skip/swap/restore
- `tests/unit/torch/quantization/test_sequential_checkpoint.py` —
checkpoint save/resume correctness
- `tests/gpu/torch/quantization/plugins/test_accelerate_gpu.py` —
CPU-offloaded layerwise + GPTQ + checkpoint resume
- `tests/gpu/torch/quantization/test_fsdp2.py` — FSDP2 layerwise
calibration

### Verified
- [x] Qwen3-8B: layerwise calibration + checkpoint save/restore +
fakequantized checkpoint export + vLLM serve
- [x] DeepSeek-R1: checkpoint resume tested
- [x] DeepSeek-R1: fakequantized checkpoint export verified

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-04-18 00:32:34 +00:00
sugunav14 dc7ad66b71 GPTQ vector (#1223)
### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added backend-specific GPTQ helper registration to allow
backend-tailored GPTQ behavior.

* **Bug Fixes**
* Prevented KV-cache state from leaking across repeated per-layer
forwards during calibration.

* **Tests**
* Added GPU-focused tests validating GPTQ combined with vector
quantization, including accuracy and end-to-end comparisons.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com>
2026-04-17 15:54:39 -07:00
Keval MorabiaandClaude Sonnet 4.6 e4b054bf32 Fix and Speedup megatron_mmlu by >10x via prefill scoring and global batching (#1280)
### What does this PR do?

Type of change: new feature + bug fix

Two improvements to Megatron inference utilities:

**1. Pipeline Parallel (PP) correctness fixes**

PP inference was producing garbage output (MMLU ~0.24, random chance).
Two root causes:

- `megatron_generate` / `megatron_prefill` used
`get_forward_backward_func()` (the training pipeline scheduler), which
is not designed for inference. Rewrote both functions to use explicit
P2P communication via `recv_from_prev_pipeline_rank_` /
`send_to_next_pipeline_rank`, matching the `run_mcore_inference`
pattern.
- `import_mcore_gpt_from_hf` loads HF weights into stage 0's embedding
but never updates the output_layer on the last PP stage when
`share_embeddings_and_output_weights=True`. At model init,
`setup_embeddings_and_output_layer()` all-reduces from stage 0 to sync
the output layer; after importing HF weights that all-reduce is stale.
Fix: call `model.setup_embeddings_and_output_layer()` again after
import.

**2. `megatron_mmlu` speedup (~6x)**

Replaces the `megatron_mmlu` implementation with a significantly faster
approach that matches how `lm-evaluation-harness` scores multiple-choice
questions.

**Before:** autoregressive generation (`megatron_generate`, `osl=2`) per
example, 114 separate `load_dataset` calls, batch_size=1 — 260s for 5%
data.

**After:** single prefill forward pass + argmax over {A,B,C,D} logits, 2
`load_dataset` calls, configurable batch_size — 18s for 5% data (~6x
faster).

### Changes

**PP fixes:**
- `megatron_generate` / `megatron_prefill`: replace
`get_forward_backward_func` with explicit P2P
(`recv_from_prev_pipeline_rank_` / `send_to_next_pipeline_rank`)
- `import_mcore_gpt_from_hf`: call
`model.setup_embeddings_and_output_layer()` after HF weight import when
PP>1 and `share_embeddings_and_output_weights=True`
- `megatron_prefill`: add `skip_return_logits` param and VLM support
(needed for PP non-last stages)

**MMLU speedup:**
- **Log-likelihood scoring**: replace `megatron_generate` with
`megatron_prefill` — one forward pass per batch, no autoregressive
decode loop
- **Global batching**: collect all examples across all subjects, sort by
descending sequence length, run in `batch_size` chunks
- **2 dataset loads** instead of 114: use `load_dataset("cais/mmlu",
"all")` with per-subject grouping; skip dev load when `few_shots=0`
- **`percentage` → `fraction`** parameter rename for clarity
- **tqdm progress bar** (rank-0 only)

### Testing

- `test_megatron_generate_and_mmlu` parametrized over `tp` and `pp`.
Accuracy assertion: `0.36 < score < 0.39`. Manually checked generated
text is coherent.
- Re-ran M-Bridge Minitron MMLU based pruning for Nano v2 9B -> 7B and
all top 10 candidate's MMLU numbers are ballpark similar as before

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — `percentage` parameter
renamed to `fraction`; `enable_kv_cache` removed from `megatron_mmlu`
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — existing test updated and
parametrized for TP+PP
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

🤖 Generated with [Claude Code](https://claude.ai/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved pipeline-parallel generation and MMLU evaluation reliability;
fixed output-layer synchronization in shared-embedding + pipeline
setups.

* **New Features**
* MMLU scoring now uses batched prefill logit scoring for faster,
batched evaluation.

* **Behavior Changes**
* Default MMLU sampling increased from 5% to 10%; calibration batch
sizing adjusted and related CLI/help text updated.

* **Tests**
* Distributed tests cover tensor- and pipeline-parallel modes and
tighten MMLU validation ranges.

* **Documentation**
* Updated pruning example and benchmark timing to reflect new sampling
and speedup.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-17 19:20:54 +00:00
Keval Morabia 4e33368dbe Temporarily disable latest pulp and mcore until we fix its nvidia-resiliency-ext dependency (#1285)
- `megatron-core==0.17.0` released yesterday which requires nightly
version of `nvidia-resiliency-ext` for an import. Pre-installed version
in DLFW Pytorch container is `nvidia-resiliency-ext==0.5.0`
  - Temporarily pin `mcore<0.17.0` to unblock PR from merging. 
- Pin `pulp<4.0` as it has some breaking changes and release imminent

Correct fix is to just use `nemo:26.04` container instead of PyTorch
container for megatron-based tests since it always has correct
combination of all packages needed for the megatron ecosystem - Done in
#1286

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-17 16:31:19 +05:30
Keval MorabiaandClaude Sonnet 4.6 7e82a5cb03 [Serialization]: remove explicit weights_only default from safe_load to allow user to bypass if needed (#1279)
## Summary

- Remove the `kwargs.setdefault("weights_only", True)` call from
`safe_load`, deferring to torch's built-in default (which is `True` for
torch>=2.6)
- This allows users to override via the
`TORCH_FORCE_NO_WEIGHTS_ONLY_LOAD=1` env var when they trust a
checkpoint but hit `pickle.UnpicklingError`
- Add a test that verifies the default fails on unsafe objects and the
env var bypass works

## Test plan

- [x] `python -m pytest tests/unit/torch/utils/test_serialization.py -v`

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Serialization utility now respects PyTorch's default behavior and
environment-variable configuration instead of forcibly enforcing
parameter overrides, providing greater configuration flexibility.

* **Tests**
* Added test coverage validating environment-variable override
functionality and default behavior in the serialization utility.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-17 11:54:35 +05:30
Keval Morabia 3162ff003f Update 0.43 release date in CHANGELOG.rst (#1277)
As title

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
  * Updated the release date for version 0.43 in the changelog.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-17 11:23:19 +05:30
Keval Morabia d073d8d8e6 Update codecov.yml (#1278)
Dont allow more than 1% overall project coverage drop per PR. 2% was too
much for such a large codebase

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Updated code coverage enforcement thresholds for pull requests to
maintain stricter quality standards.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-17 09:47:56 +05:30
Hrishith Thadicherla fe8c5178c7 Removed version fixes for torch transformers in windows ptq example requirements (#1275)
### What does this PR do?

Type of change: Bug fix

Removed version fixes for torch and transformers



### Testing
Tested quantization with a couple of models . Working as expected.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Relaxed dependency specs: removed strict pin for torch to allow latest
compatible installs, and constrained transformers to <5.0.0 for broader
compatibility and easier updates.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com>
Signed-off-by: Hrishith Thadicherla <99313418+hthadicherla@users.noreply.github.com>
2026-04-17 09:16:48 +05:30
Chenjie Luo 04fcf24227 Fix LLM deploy test failure by defaulting expert parallelism to 1 (#1273)
### What does this PR do?

Type of change: Bug fix

Fixes TRT-LLM DeepEP kernel failures during LLM deployment on
unsupported GPUs (e.g. Blackwell SM 12.0) by defaulting expert
parallelism (`ep`) to 1 instead of auto-setting it to the GPU count for
MoE models.

Previously, when the model config contained expert-related keys, `ep`
was automatically set to `torch.cuda.device_count()`, which triggered
DeepEP kernel failures on GPUs that don't support it. Now `ep` defaults
to 1 while still enabling attention data parallelism for MoE models.
Expert parallelism can be enabled explicitly by the caller when the
environment is known to support it.

### Testing

- [x] Verified that the `llm_ptq` test passes with this fix on Blackwell
GPUs.
- [x] 2-gpu CI test triggered:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24495054531/job/71588037727

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-16 16:18:38 -07:00
kaix-nv 6ded36bcbb Add dep check for ptq and runtime check for evaluation/deployment (#1240)
### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->
PTQ: model-specific dependency support
- Add EXTRA_PIP_DEPS support to the launcher's `ptq.sh` so models
requiring extra pip packages (e.g., `mamba-ssm` for hybrid Mamba
architectures like Nemotron) can install them automatically before
running PTQ. Also updates the PTQ skill with a new Step 2.5 for
detecting model-specific dependencies.

Container registry auth checks
- Add new section 6 covering auth detection for enroot/pyxis, Docker,
and Singularity/Apptainer. Includes credential locations, how to add
them, and common failure modes.
- Add Step 7.5 with NEL default image table, DockerHub-first strategy
with NGC fallback, and build-config CLI note.
- Add auth check before remote SLURM deployment.

### Usage

Set EXTRA_PIP_DEPS in the launcher YAML's environment section:
```
task_0:
  script: common/hf/ptq.sh
  args:
    - --repo nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
    - --local-dir /hf-local/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
    - --
    - --quant nvfp4
    - --tasks quant
  environment:
    - EXTRA_PIP_DEPS: "mamba-ssm causal-conv1d"
```

### Testing
<!-- Mention how have you tested your change if applicable. -->
Tested end-to-end: NVFP4 quantization of
`NVIDIA-Nemotron-3-Nano-30B-A3B-BF16` on a B200 cluster via the
launcher. Job succeeded: mamba-ssm installed automatically, calibration
completed (512 samples, 84s), checkpoint exported (18 GB, 2 safetensor
shards).

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **Documentation**
* Added container registry authentication verification workflow for
SLURM deployments, including credential checks, verification commands,
common failure symptoms, and remediation guidance.
* Required credential validation before SLURM job submission and added
SLURM-only verification steps with image fallback recommendations.
* New dependency-checking step for models that use
remote/trust_remote_code, plus guidance for resolving extra package
requirements and tightened build-config guidance.
* Updated PTQ launcher documentation to reference the new wrapper
script.

* **New Features**
* Support for specifying extra pip dependencies during model processing
via an environment variable.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kai Xu <kaix@nvidia.com>
2026-04-16 18:21:32 +00:00
github-actions[bot]andgithub-actions[bot] 0a4908d185 [chore]: weekly bump of uv.lock on main (2026-04-15) (#1266)
## Summary
Automated weekly update of uv.lock file for nSpect Scanning:
- `uv.lock` — upgraded all transitive dependencies to latest compatible
versions

Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-04-16 10:19:33 -07:00
kinjalpatel27 f238d93d86 vLLM fakequant fold weight_quantizer for megatron export (#1246)
### What does this PR do?

Type of change: Bug fix
During Megatron→vLLM fakequant export
(`export_mcore_gpt_to_hf_vllm_fq`), the `weight_quantizer` is now
applied as fake-quantization (quantize + dequantize) directly into the
exported weight tensor, and its amax is no longer saved to
`quantizer_state.pth`. On reload, if `weight_quantizer` keys are absent
from the checkpoint (because they were folded at export time), the
corresponding quantizer modules are disabled.
This change is useful especially when amax across experts are not synced
for `weight_quantizer`, this allows the `weight_quantizer` to keep them
different for better accuracy.

### Usage
```python                                                                                                                                                                                                    
# Unchanged — export API is the same                                                                                                                                                                         
export_mcore_gpt_to_hf_vllm_fq(model, pretrained_model_name_or_path=..., export_dir=...)
```
 
### Testing
Step 1 — Quantize (run from Megatron-LM
`examples/post_training/modelopt`):
  ```bash
HF_MODEL_CKPT=<path/to/hf/weights> MLM_MODEL_SAVE=<quant-ckpt-name> \
bash quantize.sh <hf-model-id> NVFP4_DEFAULT_CFG
```  

Step 2 — Export for vLLM fakequant:                                                                                                                                                                          
```bash  
MLM_EXTRA_ARGS=--export-vllm-fq \ 
HF_MODEL_CKPT=<path/to/hf/weights> \ 
MLM_MODEL_CKPT=<quant-ckpt-name> \ 
EXPORT_DIR=<export-dir> \ 
bash export.sh <hf-model-id> 
```

Step 3 — Serve (run from examples/vllm_serve):                                                                                                                                                               
```bash
 QUANT_CFG=NVFP4_DEFAULT_CFG \ 
 QUANT_FILE_PATH=<export-dir>/quantizer_state.pth \ 
 python3 vllm_serve_fakequant.py <export-dir> \ 
 -tp 1 --served-model-name <model-name> \ 
 --host 0.0.0.0 --port 8000 \ 
--trust-remote-code --enforce-eager \ 
 --disable-custom-all-reduce \ 
--gpu-memory-utilization 0.8          
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A 
- Did you write any new necessary tests?: N/A 
- Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A 

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit

* **Bug Fixes**
  * Better handling when loading checkpoints: missing weight-quantizer entries are validated and corresponding modules are disabled to avoid load failures.

* **Improvements**
  * Export now folds enabled weight quantizers into exported weights when present and omits internal weight-quantizer tensors from the exported state to produce cleaner exports.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-04-16 09:42:36 -07:00
Zhiyu 9f8188db64 [1/N] Polish deployment skills - Add a debug loop for unsupported models (#1236)
### What does this PR do?

Type of change: Skills update

Add a debug loop guide for deploying unsupported models to the
deployment skill. When deploying models not in the validated support
matrix (e.g., newly quantized VLMs or models with new architectures like
Devstral/ministral3), the inference framework (vLLM, SGLang, TRT-LLM)
often fails during model init or weight loading.

This PR adds:
- `references/unsupported-models.md` — a 5-step iterative debug
workflow: **run → read error → diagnose → patch framework source →
re-run**
- A short pointer in `SKILL.md` under "Unsupported Models" (keeps
SKILL.md concise, matching the PTQ skill's pattern)

The guide covers five common error categories with real-world examples:
- **Weight key mismatches** (e.g.,
[vllm#39406](https://github.com/vllm-project/vllm/pull/39406))
- **Quantized/unquantized layer confusion** (e.g.,
[sglang#18937](https://github.com/sgl-project/sglang/pull/18937))
- **Missing architecture support** (e.g., `ministral3` not handled in
vLLM's `mistral3.py`)
- **Transformers version mismatches**
- **Kernel-level issues** (escalate to framework team)

Motivated by deploying a Devstral-Small-2-24B NVFP4 checkpoint on vLLM,
where vLLM's `mistral3.py` didn't handle `ministral3` as a text backbone
model type.

### Testing

Validated end-to-end: NVFP4 quantization of Devstral-Small-2-24B → vLLM
deployment on B100 GPUs with the debug loop (3 iterations to get the
server running).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A (documentation only)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (skill documentation)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added a deployment guide for unsupported models with an iterative "run
→ read error → diagnose → patch → re-run" troubleshooting workflow,
common failure categories, escalation criteria, and practical
remediation tips.
* Added post-quantization validation guidance and a lightweight script
to verify which layers are quantized vs excluded, plus recommendations
for addressing unexpected layers and MoE/VLM naming gaps.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
2026-04-16 09:03:50 +00:00
Chenjie Luo d45219b390 Fix debugger server failing to detect editable-installed modelopt (#1270)
## Summary
- Removed `PYTHONPATH="" python -I` override in `check_modelopt_local()`
so the PYTHONPATH validation uses the actual environment instead of an
isolated one
- Moved the workdir log line earlier in `server.sh` for better debugging
visibility

## Test plan
- [x] Start `server.sh` inside a Docker container and verify it
correctly detects editable-installed modelopt
- [x] Confirm the workdir is logged before the modelopt check runs

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Server startup now displays the configured work directory earlier in
the initialization process, providing improved visibility of the active
directory during server launch.
* Simplified the modelopt validation check during server initialization
while maintaining the same validation behavior.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-15 22:53:46 -07:00
yeyu-nvidiaandClaude Opus 4.6 07ae8e7128 Add LoRA co-training support for HF EAGLE speculative decoding (#1060)
### What does this PR do?

Type of change: New feature + bug fixes

Adds **LoRA co-training** support for HF EAGLE speculative decoding.
When `eagle_base_lora=True`, HF PEFT LoRA adapters are injected into the
base model and co-trained alongside the EAGLE draft module in a single
online training pass. A preservation loss (KL divergence between the
original frozen base model output and the LoRA-adapted output) prevents
base model drift. LoRA adapter weights are exported in standard peft
format alongside EAGLE draft artifacts.

### Key features

- **LoRA injection**: `peft.inject_adapter_in_model` applied in-place
(no wrapper), keeping the existing `HFEagleModel` structure intact.
- **Preservation loss**: Cross-entropy `H(ref, lora)` — equivalent
gradient to `KL(ref || lora)` since `H(ref)` is constant w.r.t. LoRA
params.
- **Warmup schedule**: `eagle_base_lora_warmup_steps` freezes LoRA for N
steps while the EAGLE head stabilizes, then enables co-training via a
`LoRAWarmupCallback`.
- **Logits detach regularization**: `eagle_base_lora_logits_detach_prob`
stochastically detaches base logits from the EAGLE loss path, preventing
LoRA from degenerating to maximize EAGLE accuracy at the cost of base
model quality.
- **Export**: Standard peft format (`adapter_model.safetensors` +
`adapter_config.json`) alongside EAGLE draft model.
- **Merge script**: `scripts/merge_lora.py` merges LoRA weights into the
base model and restores the original `config.json` (avoids transformers
5.x rewriting `rope_theta` → `rope_parameters` which breaks
vLLM/TRT-LLM).
- **Multinode fix**: `dp_shard_size` now uses `WORLD_SIZE` instead of
local GPU count.

### Config options

```python
mtsp.convert(model, mode=[("eagle", {
    "eagle_base_lora": True,                          # enable LoRA co-training
    "eagle_base_lora_rank": 64,                       # LoRA rank
    "eagle_base_lora_alpha": 16.0,                    # LoRA scaling
    "eagle_base_lora_target_modules": ["q_proj", "k_proj", "v_proj", "o_proj"],
    "eagle_base_lora_preservation_loss_weight": 0.1,  # preservation loss weight
    "eagle_base_lora_warmup_steps": 0,                # freeze LoRA for N steps
    "eagle_base_lora_logits_detach_prob": 0.5,        # detach prob (0=never, 1=always)
})])
```

### Experimental results (Qwen3-8B, checkpoint-60000)

Base model quality preserved across detach_prob sweep (lm_eval: IFEval,
ARC-C, Winogrande — results pending final collection).

**Acceptance rate** (mt_bench, draft_length=3, output_length=4096,
temperature=0):

| detach_prob | vLLM AR | TRT-LLM AR |
|---|---|---|
| baseline (no LoRA) | 2.14 | 2.15 |
| 0.5 | 1.45 | 1.44 |
| 0.8 | **3.06** | **3.01** |
| 0.85 | 2.90 | 2.90 |
| 0.9 | 2.76 | 2.77 |
| 0.95 | 2.51 | 2.58 |
| 0.99 | 2.37 | 2.37 |
| 0.999 | 2.30 | 2.27 |
| 0.9999 | 2.31 | 2.26 |

Best AR at `detach_prob=0.8`: ~40% improvement over baseline.

### Testing

`tests/unit/torch/speculative/plugins/test_hf_speculative_lora.py` (5
tests):
- `test_lora_layers_injected` — LoRA layers present after conversion
- `test_trainable_params` — only `lora_*` and `eagle_module` params are
trainable
- `test_forward_returns_loss` — forward returns non-zero scalar loss
- `test_eagle_offline_incompatible` — `eagle_base_lora=True` +
`eagle_offline=True` raises `ValueError`
- `test_export_lora_artifacts` — export produces standard peft adapter
files

### Bug fixes (included in this PR)

1. **`launch_train.sh` case pattern ordering**: glob
`--eagle_base_lora*` was before specific patterns
(`--eagle_base_lora_rank*`, etc.), silently swallowing LoRA args.
2. **LoRA optimizer exclusion during warmup**: warmup freezing excluded
LoRA from the optimizer entirely; fixed with `add_param_group` in the
callback.
3. **`merge_lora.py` config.json**: `save_pretrained()` with
transformers >=5.x rewrites `rope_theta` → `rope_parameters`, breaking
vLLM positional embeddings. Fixed by copying the original base model
config.
4. **Multinode `dp_shard_size`**: used local GPU count instead of
`WORLD_SIZE`.

### Checklist

- [x] Backward compatible (all new config fields have defaults)
- [x] Uses `peft` via lazy imports (no hard dependency)
- [x] Unit tests added
- [x] Online HF training only (`eagle_offline=True` blocked)

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-16 01:35:15 +00:00
361f7e391b Merge puzzletron compression algorithm (#1121)
### What does this PR do?

Implement puzzletron compression algorithm based on Puzzle paper
(https://arxiv.org/abs/2411.19146)

<details>
<summary> Th list of reviewed and merged MRs that resulted in the
feature/puzzletron branch</summary>

Merging dkorzekwa/any_model to feature/puzzletron

[Add anymodel directories to feature/puzzletron by danielkorzekwa · Pull
Request #974 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/974)
- merged

[Draft: anymodel activation scoring by danielkorzekwa · Pull Request
#989 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/989)
- merged

[Draft: Merge anymodel pruning by danielkorzekwa · Pull Request #990 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/990/)
- merged

[Draft: Merging anymodel:build_library_and_stats by danielkorzekwa ·
Pull Request #993 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/993)
- merged

[Dkorzekwa/any model calc one block scores by danielkorzekwa · Pull
Request #994 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/994)
- merged

[Draft: merge any_model: mip_and_realize_models by danielkorzekwa · Pull
Request #995 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/995)
- merged

[Dkorzekwa/any model other modeqls by danielkorztiekwa · Pull Request
#1007 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1007/)
- merged

PR to 1007: https://github.com/NVIDIA/Model-Optimizer/pull/1039 - merged

[Dkorzekwa/anymodel gptoss by danielkorzekwa · Pull Request #1020 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1020)
- merged

[Merge any_model tutorial by danielkorzekwa · Pull Request #1035 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1035)
- merged

[Merge mbridge distillation for any_model by danielkorzekwa · Pull
Request #1036 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1036)
- merged

[MR branch for the remaining difference between dkorzekwa/any_model an…
by danielkorzekwa · Pull Request #1047 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1047)
- merged

[Dkorzekwa/decilm hf code cleanup by danielkorzekwa · Pull Request #1071
·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1071)
- merged

[Dkorzekwa/decilm hf code cleanup 2 by danielkorzekwa · Pull Request
#1073 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1073)
- merged

[Dkorzekwa/anymodel subblock stats by danielkorzekwa · Pull Request
#1085 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1085)
- merged

[Dkorzekwa/anymodel subblock stats nodecilm by danielkorzekwa · Pull
Request #1102 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1102)
- merged

[Dkorzekwa/decilm cleanup post subblockstats by danielkorzekwa · Pull
Request #1103 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1103)
- merged

[code clean up by danielkorzekwa · Pull Request #1110 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1110)
- merged

Merging into main:

[Activation hooks redesign (reuse hooks component across both minitron
and puzzletron) by danielkorzekwa · Pull Request #1022 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1022)
- merged

[Dkorzekwa/puzzletron use importance hooks from prune by danielkorzekwa
· Pull Request #1115 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1115)
- merged

</details>

<!-- Details about the change. -->

### Usage

Puzzletron tutorial:

https://github.com/NVIDIA/Model-Optimizer/tree/feature/puzzletron/examples/puzzletron

### Testing
The main e2e test for compressing 9 models with Puzzletron:

https://github.com/NVIDIA/Model-Optimizer/blob/feature/puzzletron/tests/gpu/torch/puzzletron/test_puzzletron.py

2-gpu nightly tests: 

-
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24468209205/job/71501061203
-
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24470214159/job/71508152952

### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added Puzzletron: end-to-end heterogeneous pruning & NAS workflow with
AnyModel support, example pipelines, deployment and evaluation
utilities, and tools for converting/pruning and exporting compressed
checkpoints.

* **Documentation**
* Comprehensive Puzzletron tutorials, model-specific guides, evaluator
instructions, example configs, and changelog entry.

* **Chores**
* CI/workflow updates (extras installation, longer GPU test timeout),
pre-commit hook exclusion updated, and CODEOWNERS entries added.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com>
Signed-off-by: Liana Mikaelyan <lmikaelyan@nvidia.com>
Signed-off-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com>
Signed-off-by: Daniel Korzekwa <daniel.korzekwa@gmail.com>
Signed-off-by: jrausch <jrausch@nvidia.com>
Signed-off-by: root <root@pool0-00848.cm.cluster>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Liana Mikaelyan <lmikaelyan@nvidia.com>
Co-authored-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com>
Co-authored-by: J Rausch <38429553+j-rausch@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-16 00:48:11 +05:30
Gwena Cunhaanddmoodie dec2952992 [6034518] Downgrade TRT support for remote autotuning in Autotune from 10.16 to 10.15 (#1259)
### What does this PR do?

Type of change: Bug fix

Remote autotuning is supported in TensorRT from version 10.15, but fails
with Autotune as it's checking for 10.16+. This PR fixes that check and
updates documentation accordingly.

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
See bug 6034518.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!---
Mandatory -->
- Did you write any new necessary tests?: N/A <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added a Remote Autotuning guide for TensorRT 10.15+ with CLI examples;
updated examples to require `--safe --skipInference`.

* **Updates**
* Lowered TensorRT minimum requirement for remote autotuning from 10.16
to 10.15.
  * Clarified CLI help text for trtexec/autotune arguments.

* **Bug Fixes**
* trtexec-based autotuning now verifies the trtexec executable version
when checking compatibility.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
Signed-off-by: dmoodie <dmoodie@nvidia.com>
Co-authored-by: dmoodie <dmoodie@nvidia.com>
2026-04-15 14:31:41 -04:00
Chenjie Luo 7c8557158d Add job cancellation support to the debugger command relay (#1262)
## Summary

- Add a `cancel` subcommand to the client that terminates the currently
running command on the server
- Server now runs commands in the background with PID tracking, enabling
cancellation mid-execution
- Client-side timeouts automatically cancel the running command on the
server (previously the server process was left running)
- Hardened against race conditions through 4 rounds of adversarial
review (15 fixes total)

### Key changes

**server.sh:**
- Commands run in background with PID tracked in `$RELAY_DIR/running`
(atomic tmp+mv write)
- Cancel detection loop checks for `$RELAY_DIR/cancel` file with cmd_id
verification
- SIGTERM with 5s grace period, then SIGKILL escalation for stuck
processes
- `.exit` file written before `running` marker removed (ordering
guarantee)
- `set -e`-safe: `wait` uses `|| exit_code=$?` pattern; cleanup trap
fully guarded
- Stale cancel files cleared at command start; mismatched/empty signals
rejected
- Command file read into memory and removed before execution (eliminates
TOCTOU with client timeout)

**client.sh:**
- New `cancel` subcommand: writes target cmd_id to cancel file, waits
for server acknowledgment (30s timeout)
- `run` timeout now sends targeted cancel signal (verifies cmd_id match
to avoid killing wrong command)
- `run` timeout cleans up orphaned result files
- `status` shows currently running command
- `flush` rejects if a command is currently running (prevents state
corruption)
- Exit code validated as numeric before use

### Protocol additions

```
.relay/
├── running    # server writes cmd_id:pid while executing (atomic)
├── cancel     # client writes target cmd_id to request cancellation
```

## Test plan

- [ ] Start server in Docker, handshake from host
- [ ] Run a command (`client.sh run "sleep 30"`), cancel it (`client.sh
cancel`), verify exit code 130
- [ ] Run a command with short timeout (`--timeout 5 run "sleep 30"`),
verify auto-cancel
- [ ] Run a command that exits non-zero, verify server stays alive
- [ ] Run `status` during execution, verify it shows the running command
- [ ] Attempt `flush` during execution, verify it is rejected
- [ ] Cancel when nothing is running, verify clean message

🤖 Generated with [Claude Code](https://claude.com/claude-code)

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (bash scripts for dev
tooling, tested manually)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (internal tooling)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added a comprehensive debug skill guide and protocol reference with
quick CLI examples and a new “Cancelling Commands” section.

* **New Features**
* Client-side `cancel` command to terminate the currently running remote
command.
  * Status now reports active command (or `(idle)`).

* **Improvements**
* Stronger startup/validation guidance, safer shutdown/cleanup,
deterministic cancel exit semantics (130), and auto-cancel on
client-side timeout.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-15 16:49:42 +00:00
Chenjie LuoandClaude Opus 4.6 952a62bf65 Fix missing attention_mask in calibration dataloader (#1261)
## Summary
- When `include_labels=False` (the default for PTQ calibration),
`get_dataset_dataloader` was discarding the `attention_mask` produced by
the tokenizer and only returning `input_ids`.
- Without `attention_mask`, HuggingFace models create a full causal
mask, causing padding tokens to participate in attention during
calibration and skewing quantization statistics.
- This fix includes `attention_mask` alongside `input_ids` so the model
correctly ignores padding tokens during calibration forward passes.

## Details
In `modelopt/torch/utils/dataset_utils.py`, the tokenizer call at line
387 with `padding=True` produces both `input_ids` and `attention_mask`.
The `include_labels=True` path (line 406) already preserves the full
`batch_encoded` dict including `attention_mask`. However, the
`include_labels=False` path was only keeping `input_ids` "for backward
compatibility."

During the calibration forward loop (`_forward_loop` →
`_process_batch`), the batch dict is unpacked as `**kwargs` into
`model.forward()`. Without `attention_mask`, HF models default to
attending to all positions including padding, which pollutes calibration
statistics.

**Practical impact**: With `batch_size=1` there is no padding so the bug
is invisible. With larger batch sizes and variable-length samples,
shorter sequences get padded and the effect grows.

## Test plan
- [x] Existing unit tests pass
(`tests/unit/torch/utils/test_dataset_utils.py`)
- [x] Pre-commit hooks pass
- [ ] Verify PTQ accuracy with batch_size > 1 on a padded calibration
dataset (GPU required)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 00:04:27 -07:00
Shengliang Xu c9b11559f1 Normalize .yml to .yaml in modelopt_recipes (#1260)
### What does this PR do?

Type of change: Chore

Standardize YAML file extensions in `modelopt_recipes/` to `.yaml` for
consistency. The existing recipes used a mix of `.yml` (PTQ recipes) and
`.yaml` (speculative decoding, model-specific recipes).

#### Changes

**Renamed files:**
- `modelopt_recipes/general/ptq/fp8_default-fp8_kv.yml` → `.yaml`
- `modelopt_recipes/general/ptq/nvfp4_default-fp8_kv.yml` → `.yaml`
- `modelopt_recipes/general/ptq/nvfp4_experts_only-fp8_kv.yml` → `.yaml`
- `modelopt_recipes/general/ptq/nvfp4_mlp_only-fp8_kv.yml` → `.yaml`
- `modelopt_recipes/general/ptq/nvfp4_omlp_only-fp8_kv.yml` → `.yaml`

**New pre-commit hook:** `normalize-yaml-ext`
- `tools/precommit/normalize_yaml_ext.py` — auto-renames `.yml` to
`.yaml`
  for any staged file under `modelopt_recipes/`. Runs before recipe
  validation so future contributions are caught automatically.

**Updated references:**
- `tests/unit/recipe/test_loader.py` — built-in recipe paths updated to
  `.yaml`

Note: `load_recipe()` and `load_config()` probe both `.yml` and `.yaml`
suffixes, so callers using paths without extensions (e.g.,
`load_recipe("general/ptq/fp8_default-fp8_kv")`) are unaffected.

### Testing

Existing recipe loader tests pass with updated paths.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Standardized recipe file extensions and added an automated pre-commit
normalization hook to enforce the convention.
* **Tests**
* Updated unit tests to reference the new recipe filename convention and
ensure consistency with configuration loading.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-04-14 17:49:51 -07:00
kinjalpatel27 1619421383 Added support for MoE for vllm >= 0.14.0rc1 (#1162)
### What does this PR do?

Type of change: Bug fix
`_QuantFusedMoEBase.forward()` previously replaced
`vllm_fused_moe_package.invoke_fused_moe_kernel`, which was replaced
starting in vLLM v0.14.0rc1,

There are two paths for FusedMoE forward:
```
Path 1 (Modular — standard CUDA path):                                                                                                                                                                       
    FusedMoE.forward()                                                                                                                                                                                         
      → self.runner.forward()                                                                                                                                                                                  
        → TritonExperts.apply()                                                                                                                                                                                
          → invoke_fused_moe_triton_kernel()  ← called twice (w1, w2)                                                                                                                                          
                                                                                                                                                                                                               
  Path 2 (legacy):                      
    inplace_fused_experts / outplace_fused_experts                                                                                                                                                             
      → fused_experts_impl()                       
        → dispatch_fused_moe_kernel()              
          → invoke_fused_moe_triton_kernel()                                                           
            or invoke_fused_moe_wna16_triton_kernel()                                                  
            or invoke_fused_moe_wna16_cuda_kernel()
```
This caused an `AttributeError` / assertion failure for any MoE model
quantized with vLLM ≥ v0.14.0rc1.

The fix refactors the kernel-patching logic into a `_patch_moe_kernel()`
context manager that probes for both attribute names (the two names are
mutually exclusive across vLLM versions — confirmed by inspecting every
release from v0.10.0 to v0.19.1).
    
### Usage

NA

### Testing
```
docker run --gpus all -it --shm-size=160GB --network host --rm -v <modelopt path>:/home/modelopt \
vllm/vllm-openai:v0.15.0 bash -c "cd /home/modelopt && pip install . && pip install datasets && \
  QUANT_CFG=NVFP4_DEFAULT_CFG python3 /home/modelopt/examples/vllm_serve/vllm_serve_fakequant.py \
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 -tp 1 --served-model-name NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 \ 
  --host 0.0.0.0 --port 8001 --trust-remote-code --disable-custom-all-reduce \
--gpu-memory-utilization 0.8" 
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?:  N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Refactor**
* Ensures quantized expert weights are correctly used by the fused-MoE
execution path so inference uses the intended quantized tensors.
* Replaces fragile manual swapping of the runtime kernel with a safer,
context-managed swap that reliably caches and restores the original.
* Adds runtime detection and selection among available fused-MoE kernel
entrypoints to support multiple variants.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-04-14 16:40:35 -07:00
Chenhan D. YuandClaude Opus 4.6 3131195241 add: DFlash block diffusion speculative decoding (#1211)
DFlash (Block Diffusion for Flash Speculative Decoding) predicts an
entire block of tokens in a single forward pass using masked parallel
prediction with KV injection from the target model's hidden states.

Key features:
- Feature fusion (multi-layer hidden states -> FC + RMSNorm)
- KV injection (fused features as K/V in every draft layer with QK-norm)
- Random anchor sampling with bidirectional intra-block attention
- Logit distillation with exponential loss decay (gamma weighting)
- Multi-node DDP training with checkpoint resume
- Export to z-lab compatible HF format
- Online validation (context-dependent ground truth)

Training recipe:
modelopt_recipes/general/speculative_decoding/dflash.yaml
Results: examples/speculative_decoding/doc/dflash_results.md

### ModelOpt Eval (online validation, osl=512)

| Dataset | z-lab | ModelOpt (306K) | Diff |
|---------|-------|-----------------|------|
| gsm8k | 4.10 | **5.19** | **+1.09** |
| MT-Bench | 3.58 | **4.36** | **+0.78** |

### z-lab Official Eval (dflash.benchmark, osl=512)

| Dataset | z-lab | ModelOpt (306K) | Diff |
|---------|-------|-----------------|------|
| gsm8k | **5.00** | 4.08 | -0.92 |
| MT-Bench | **3.28** | 2.99 | -0.29 |

> z-lab model trained with block_size=16. ModelOpt trained with
block_size=8.

## Evaluation Method Impact (gsm8k)

| Eval Method | z-lab checkpoint | ModelOpt (306K) |
|-------------|-----------------|-----------------|
| Fixed GT (ModelOpt eval) | 2.95 | 4.23 |
| Online GT (ModelOpt eval) | 4.10 | **5.19** |
| z-lab official eval | **5.00** | 4.08 |

### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added DFlash speculative decoding mode with parallel block prediction
support.
* Included training launchers and MT-Bench evaluation scripts for DFlash
models.
* Added online acceptance rate validation for improved inference
verification.

* **Documentation**
* DFlash quick start guide with configuration parameters and training
examples.
  * Performance results and benchmarks for DFlash-trained models.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 15:58:39 -07:00
kinjalpatel27 73be81037d vLLM fakequant: add recipe-based quantization support (#1233)
### What does this PR do?

Type of change: example update

This PR adds recipe-based quantization support to the vLLM fakequant
example.


### Testing
```
docker run --gpus all -it --shm-size=160GB --network host --rm --entrypoint bash -v <modelopt>:/home/modelopt vllm/vllm-openai:v0.15.0 -c "cd /home/modelopt && pip install . && pip install datasets && RECIPE_PATH=/home/modelopt/modelopt_recipes/general/ptq/nvfp4_mlp_only-fp8_kv.yml python3 /home/modelopt/examples/vllm_serve/vllm_serve_fakequant.py Qwen/Qwen3-0.6B -tp 1 --served-model-name Qwen3-0.6B --host 0.0.0.0 --port 8001 --trust-remote-code --disable-custom-all-reduce --gpu-memory-utilization 0.8"
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added `RECIPE_PATH` environment variable support enabling users to
specify ModelOpt PTQ recipe YAML files for quantization configuration in
vLLM serving.

* **Documentation**
* Updated examples and documentation to support recipe-driven
quantization configuration, aligning export workflow with recipe-based
setup.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-04-14 13:37:02 -07:00
Shengliang Xu b6c6ec342c use typed quantize config instead of a raw dict (#1249)
### What does this PR do?

But fix:

Use typed QuantizeConfig instead using raw dict for formal typed
ModelOpt configs.

The dict typing was accidental.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Quantization recipe configuration is now implemented with a
strongly-typed, structured schema that enforces type safety and provides
enhanced validation with comprehensive error detection capabilities.

* **Tests**
* Updated recipe loading tests to correctly validate quantization
configurations when recipes are loaded from directories, fully
supporting the new structured object-based configuration format.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-04-13 18:24:22 -07:00
Ajinkya RasaneandClaude Opus 4.6 202c3d3894 Add SwinTransformer support for torch_onnx quantization workflow (#1235)
## Summary
- Enable end-to-end quantize → ONNX export → TRT engine pipeline for
SwinTransformer models (v1 and v2) across FP8, INT8, MXFP8, NVFP4, and
auto precision modes
- Add Conv2d quantization overrides for TRT compatibility (TRT only
supports FP8/INT8 for convolutions)
- Fix FP8 LayerNorm type mismatch in TRT stronglyTyped mode by adding
`LayerNormalization` to `change_casts_to_fp16`
- Fix `cast_initializer_to_dtype` crash when node has no initializer
inputs
- Simplify `download_example_onnx.py` to a single `--timm_model_name`
(required) flag, removing redundant `--vit` and `--llama` flags
- Add vision model support matrix to README (ViT, Swin, SwinV2)
- Rewrite tests: parametrize over (ViT, Swin, SwinV2) × (fp8, int8,
mxfp8, nvfp4, auto) with TRT engine build verification

## Test plan
- [ ] `python -m pytest
tests/examples/torch_onnx/test_torch_quant_to_onnx.py -v` — 15 tests (3
models × 5 modes), all pass
- [ ] Verified Swin accuracy on ImageNet-1k across all precisions (FP8:
81.29%, INT8: 81.12%, MXFP8: 81.32%, NVFP4: 80.79%, Auto: 80.84% TRT
top-1 vs 81.37% base)
- [ ] INT4_AWQ deferred (TODO in test file) — requires INT4 exporter
changes for non-MatMul/Gemm consumer patterns

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* ONNX export supports arbitrary timm vision models with auto device
selection and new CLI options (--timm_model_name, --model_kwargs,
--no_pretrained); batch-size/input sizing is now model-generic.

* **Bug Fixes**
  * Expanded FP16/BF16 cast handling to additional ONNX ops.
* Disabled inplace ReLU before auto-quantization to avoid incorrect
transforms.
* Conv2d quantization overrides added for improved TensorRT
compatibility.
  * Safer handling when initializers are missing during dtype casting.

* **Documentation**
* README updated with supported models table, quantization mappings, and
example CLI usage.

* **Tests**
* Tests expanded to multiple architectures/quant modes and now verify
TensorRT engine build.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 04:48:29 +05:30
h-guo18 6403389eb0 Feat: Configurable Eagle ROPE scaling during export (#1238)
### What does this PR do?

JIRA ticket: https://jirasw.nvidia.com/browse/OMNIML-3469

Type of change: New feature

Decouple EAGLE training rope configuration from export rope
configuration, enabling separate YaRN rope scaling injection at export
time for long-context inference.

#### Changes

**Configurable export rope scaling (`EagleConfig`)**
- Add `eagle_export_rope_scaling` field to `EagleConfig` with default
YaRN config (`factor=32.0`, `original_max_position_embeddings=2048`)
- Set to `{}` to disable rope scaling injection at export

**Simplified training defaults (`default_config.py`)**
- Change default training rope from `llama3` (theta=500k) to `default`
(theta=10k) — models now train with simple positional embeddings; rope
scaling is applied only at export
- Add `rope_theta` inside `rope_scaling` dict for transformers 5.x
cross-version compatibility

**Move config validation/rewriting into `EagleConfig` (`config.py`)**
- `_derive_eagle_offline`: derives `eagle_offline` from
`data_args.offline_data_path` via validation context, removing manual
assignment in `main.py`
- `_check_rope_scaling_consistency`: rejects configs where
`eagle_export_rope_scaling` is set but training `rope_type` is not
`"default"`
- `_warn_rope_vs_training_seq_len`: warns when
`original_max_position_embeddings` differs from `training_seq_len`

**Export rope injection (`hf_spec_export.py`)**
- Inject `eagle_export_rope_scaling` into the exported HF config when
training rope_type is `"default"`
- Fall back `rope_theta` from `rope_scaling` dict for transformers 5.x
compatibility

**Fix Megatron RotaryEmbedding crash (`megatron_eagle.py`)**
- `dict_to_config()` set `rope_scaling=True` whenever the `rope_scaling`
key existed, even without a `"factor"` — causing `RotaryEmbedding` to
divide by `None`
- Now only enables `rope_scaling` when the dict actually contains a
`"factor"` key

### Usage

Configure in YAML config (or use defaults from `eagle3.yaml`):
```yaml
eagle:
  eagle_export_rope_scaling:
    rope_type: yarn
    factor: 32.0
    original_max_position_embeddings: 2048
```

Set to empty dict to disable export rope injection:
```yaml
eagle:
  eagle_export_rope_scaling: {}
```

### Testing

- New unit tests: `tests/unit/torch/speculative/test_eagle_config.py` —
rope consistency validator, seq_len warning, context-derived
`eagle_offline`
- New unit tests: `tests/unit/torch/export/test_hf_spec_rope_export.py`
— export rope injection, fallback, and empty-config cases

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ (new field has sensible
default; existing configs work unchanged)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ (should be added if merging as a feature)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Add export-time rope-scaling configuration for EAGLE models.

* **Improvements**
* Stronger validation and context-aware reconciliation between training
and export configs.
  * Export now injects rope-scaling and rope-theta when appropriate.
  * Default rope-scaling values updated for EAGLE variants.
  * Model instances now expose export rope-scaling for downstream use.

* **Tests**
* Added unit tests covering rope-scaling export behavior and
configuration validators.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-04-13 16:06:26 -07:00
Shengliang Xu 14b78aed1f recipes doc (#1165)
### What does this PR do?

Added an extensive guide for the ModelOpt recipe system: recipe
structure, YAML schema (quantization-focused), built-in discovery and
path resolution, floating-point shorthand conversion, example usage
(Python/CLI), authoring guidance, repository layout, and future
directions.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added comprehensive guide for ModelOpt recipes, introducing
declarative YAML-based optimization specifications. Documents recipe
structure, configuration loading, path resolution, and the three-layer
system architecture. Includes built-in recipe discovery conventions and
detailed instructions for authoring custom recipes with practical
examples.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-04-13 13:56:34 -07:00
Keval Morabia 0b42c143dd Update LICENSE and SPDX-License-Identifier as per OSRB guidance (#1244)
Update LICENSE and SPDX-License-Identifier as per OSRB guidance

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated contribution guidelines with expanded license compliance
instructions and SPDX identifier guidance.
* Extended LICENSE file with new "Third-Party Software Notices" section
documenting Apache 2.0, MIT, and BSD 3-Clause licensed components.

* **Chores**
* Updated SPDX license identifiers across multiple files to reflect dual
and triple licensing (Apache 2.0 with MIT and/or BSD 3-Clause).

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-13 22:32:54 +05:30
github-actions[bot]andgithub-actions[bot] 5523505b06 [chore]: weekly bump of uv.lock on main (2026-04-13) (#1243)
## Summary
Automated weekly update of uv.lock file for nSpect Scanning:
- `uv.lock` — upgraded all transitive dependencies to latest compatible
versions

Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-04-13 10:54:59 +00:00