Commit Graph
7 Commits
Author SHA1 Message Date
jingyu-ml 26ae8da517 [2/3] Implicit Gemm NVFP4 (#1227)
### What does this PR do?

Type of change: new feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

- Add Conv3D implicit GEMM kernel with BF16 WMMA tensor cores and fused
NVFP4 activation quantization for video diffusion VAE layers
- Integrate into _QuantConv3d via QuantModuleRegistry — automatically
dispatched when NVFP4 quantization is applied to nn.Conv3d
- Move kernel from `experimental/conv/ to modelopt/torch/kernels/conv/`;
move tests to `tests/gpu/torch/quantization/kernels/`

### Testing
<!-- Mention how have you tested your change if applicable. -->

- Added test cases to measure the difference between cuDNN and our CUDA
implicit GEMM kernel
- Added an NVFP4 fake quantization test using CUDA code

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Per-backbone quantization/export in a single run with per-backbone
checkpoints and backbone-aware quant filters
* Configurable NVFP4 block-size via CLI/config; improved NVFP4 Conv3D
inference path and Wan 2.2 quantization support
* **Bug Fixes**
* Video-model calibration now respects extra params and forces video
decoding during calibration
* **Documentation**
* Added comprehensive Conv3D implicit‑GEMM kernel documentation; removed
experimental Conv3D prototype docs/benchmark
* **Tests**
* New Wan 2.2 quantization/export tests and expanded Conv3D/FP4 kernel
test coverage
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-04-19 12:20:14 +05:30
Farid AdilazuardaandKeval Morabia 2004779a67 Update README.md for DMS (fix cd experimental/DMS to cd Model-Optimizer/experimental/DMS) (#879)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** ?

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated DMS installation instructions to reflect the repository
structure and correct directory navigation during setup.
* Clarified the setup steps so users follow the accurate directory
change before running installation commands.
* Small wording improvements to reduce confusion during the installation
process.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Farid Adilazuarda <42537562+faridlazuarda@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-18 14:55:42 +00:00
Keval Morabia bd80265123 Remove experimental/dms/uv.lock
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-08 03:02:35 -07:00
jingyu-ml 6f32d242fc Implicit Gemm NVFP4 on Conv3D (#886)
## What does this PR do?

**Type of change:** new feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

Experimental Conv3D implicit-GEMM CUDA kernel with optional NVFP4-style
(E2M1 + FP8 E4M3 scale) fake quantization for activations.

It is intended for research/prototyping and quantization-accuracy
experiments only, not production deployment.
The implementation runs as a JIT-compiled PyTorch extension, mirrors
conv3d output shape, and provides a quantized and non-quantized path to
compare numerical behavior.

There is currently no real quantized production kernel integration in
the formal ModelOpt export/compress/runtime stack; this path is kept in
experimental/ for fake-quant accuracy validation and benchmarking.

## Usage
<!-- You can potentially add a usage example below. -->

```python
import torch

from experimental.conv.implicit_gemm_cuda import conv3d_implicit_gemm_cuda
from modelopt.torch.quantization.tensor_quant import dynamic_block_quantize_op

x = torch.randn(1, 128, 21, 60, 106, device="cuda")
w = torch.randn(512, 128, 3, 3, 3, device="cuda")
block_size = 128

# Without FP4 activation quantization (drop-in-style Conv3D call)
out = conv3d_implicit_gemm_cuda(x, w, stride=(1, 1, 1), padding=(1, 1, 1))

# Optional FP4 block quantization of weights along the GEMM K dimension.
# The kernel's A-tile (activations) is quantized along K = Cin*kD*kH*kW,
# so weights must be flattened to [Cout, K] before quantizing to match.
Cout, Cin = w.shape[:2]
K = Cin * w.shape[2] * w.shape[3] * w.shape[4]
w_flat = w.reshape(Cout, K)
w_q_flat = dynamic_block_quantize_op(
    w_flat,
    block_size,
    w_flat.abs().max().unsqueeze(0),
    4,  # num_bits
    2,  # exponent_bits
    8,  # scale_num_bits
    4,  # scale_exponent_bits
)
w_q = w_q_flat.reshape_as(w)

# With FP4 activation fake quantization
out_q = conv3d_implicit_gemm_cuda(
    x,
    w_q,
    stride=(1, 1, 1),
    padding=(1, 1, 1),
    act_amax=x.abs().max().unsqueeze(0),
    quant_act=True,
    fp4_block_size=block_size,  # 128 or 256
)
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added experimental Conv3D implementation with implicit GEMM
acceleration and optional FP4 quantization support
* Added benchmarking tool to compare 3D convolution performance across
implementations
  * Enhanced quantization framework integration for Conv3D operations

* **Documentation**
* Added comprehensive guide for experimental Conv3D prototype, including
supported scenarios, API reference, and current limitations
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-03-14 04:01:41 +00:00
Keval Morabia d780fa593f Add uv.lock for experimental/dms (#964)
For nspect scanning, we need `uv.lock` with each `pyproject.toml`

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
  * Updated dependency management and lock configuration settings.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-03 18:39:01 +05:30
kstaniszewsknvandcoderabbitai[bot] 549cea8c98 Add Dynamic Memory Sparsification (DMS) training and inference implementation (#877)
## What does this PR do?

**Type of change:**  new feature

**Overview:** Training and inference code for Dynamic Memory
Sparsification (DMS) - method from NeurIPS 2025 paper [Inference-Time
Hyper-Scaling with KV Cache
Compression](https://neurips.cc/virtual/2025/loc/san-diego/poster/119605)

## Usage
Detailed in `experimental/dms/README.md` and
`experimental/dms/ARCHITECTURE.md`


## Testing
DMS tests in `experimental/dms/tests` covering:
* prefill
* generation
* gradient propagation
* chunked prefill

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No, DMS is currently experimental feature with description in
`experimental/dms`

## Additional Information
A minimal, optimized implementation of the DMS algorithm for KV-cache
compression, as described in:

> **Inference-Time Hyper-Scaling with KV Cache Compression**  
> Adrian Łańcucki, Konrad Staniszewski, Piotr Nawrot, Edoardo M. Ponti  
> Paper:
[https://arxiv.org/abs/2506.05345](https://arxiv.org/abs/2506.05345)
> NeurIPS:
[https://neurips.cc/virtual/2025/loc/san-diego/poster/119605](https://neurips.cc/virtual/2025/loc/san-diego/poster/119605)

Inference-time scaling trades efficiency for improved reasoning by
generating longer sequences. In Transformer LLMs, generation cost is
often bottlenecked by the size of the key-value (KV) cache. DMS
addresses this by learning a KV cache eviction policy that compresses
the cache while preserving accuracy.

## How it works

DMS learns a per-head eviction policy that determines which KV cache
entries to keep during generation. Rather than immediately discarding
tokens, DMS delays eviction decisions, implicitly merging
representations and preserving critical information. During training,
the compression ratio is gradually increased from 1× to a target value
(e.g., 8×), using knowledge distillation to match the outputs of an
uncompressed teacher model.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Introduces Dynamic Memory Sparsification (DMS), an algorithm for
efficient LLM inference and training with adaptive attention gating.
* Adds DMS-enabled Qwen3 models with memory-efficient KV cache
management and paged block-based storage.
* Includes student-teacher distillation training infrastructure with
noise scheduling and compression ratio control.
* Provides configuration system and training/evaluation scripts for DMS
adaptation.

* **Documentation**
  * Added architecture guide, README, and example inference notebook.

* **Tests**
* Added comprehensive test suite for chunked prefill, cache management,
and prefill/inference validation.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Konrad Staniszewski <kstaniszewsk@nvidia.com>
Signed-off-by: kstaniszewsknv <kstaniszewsk@nvidia.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-02-10 23:09:10 +01:00
kaix-nv e53ca61b71 Add contribution guidelines for experimental features (#867)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** ?

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added comprehensive guide for experimental optimization technique
development, including recommended structure, testing conventions,
licensing requirements, and graduation path to production.

* **New Features**
* Introduced experimental package with templates and utilities for
implementing research-stage optimization techniques. Includes
configuration framework and example code patterns. Emits stability
warnings to indicate unstable APIs.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Kai Xu <kaix@nvidia.com>
2026-02-07 00:31:30 +00:00