Files
Model-Optimizer/experimental
jingyu-ml 26ae8da517 [2/3] Implicit Gemm NVFP4 (#1227)
### What does this PR do?

Type of change: new feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

- Add Conv3D implicit GEMM kernel with BF16 WMMA tensor cores and fused
NVFP4 activation quantization for video diffusion VAE layers
- Integrate into _QuantConv3d via QuantModuleRegistry — automatically
dispatched when NVFP4 quantization is applied to nn.Conv3d
- Move kernel from `experimental/conv/ to modelopt/torch/kernels/conv/`;
move tests to `tests/gpu/torch/quantization/kernels/`

### Testing
<!-- Mention how have you tested your change if applicable. -->

- Added test cases to measure the difference between cuDNN and our CUDA
implicit GEMM kernel
- Added an NVFP4 fake quantization test using CUDA code

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Per-backbone quantization/export in a single run with per-backbone
checkpoints and backbone-aware quant filters
* Configurable NVFP4 block-size via CLI/config; improved NVFP4 Conv3D
inference path and Wan 2.2 quantization support
* **Bug Fixes**
* Video-model calibration now respects extra params and forces video
decoding during calibration
* **Documentation**
* Added comprehensive Conv3D implicit‑GEMM kernel documentation; removed
experimental Conv3D prototype docs/benchmark
* **Tests**
* New Wan 2.2 quantization/export tests and expanded Conv3D/FP4 kernel
test coverage
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-04-19 12:20:14 +05:30
..

Experimental Optimization Techniques

Experimental optimization algorithms and research prototypes under active development.

Purpose

For new optimization techniques (quantization, pruning, sparsity, etc.) that are:

  • Novel or research-stage algorithms
  • Not yet production-ready
  • May have unstable APIs

⚠️ Warning: Experimental features are not guaranteed to work across releases. APIs may change or features may be removed without notice. Use at your own risk.

Requirements

Each experimental technique must include:

  • README.md - Explains what the technique does, how to use it, current status, model support, and references
  • Working code - Clear, readable implementation
  • Comprehensive tests - Good test coverage demonstrating correctness
  • Detailed documentation - Clear docs on usage, APIs, and behavior
  • Example - Demonstrating usage
  • Model support list - Which models/frameworks are supported
  • Deployment info - Supported deployment frameworks (TensorRT-LLM, vLLM, SGLang, etc.) and whether custom kernels are required
  • requirements.txt - Additional dependencies beyond base modelopt
  • License headers - Apache 2.0 headers on all Python files

Example Structures

Organize your code however makes sense. Here are some examples:

Simple flat structure:

experimental/my_technique/
├── README.md
├── requirements.txt
├── my_technique.py
├── test_my_technique.py
└── example.py

Package structure:

experimental/my_technique/
├── README.md
├── requirements.txt
├── my_technique/
│   ├── __init__.py
│   ├── core.py
│   └── config.py
├── tests/
│   └── test_core.py
└── examples/
    └── example_usage.py

Quality Standards

Experimental code must meet quality standards:

  • Comprehensive test coverage required
  • Clear documentation required
  • Pass all pre-commit checks

PR Guidelines

Keep PRs focused and reviewable:

  • Split large features: Break complex techniques into multiple PRs if needed
  • Reasonable scope: PRs with tens of thousands of lines are difficult to review
  • Incremental development: Consider submitting core functionality first, then enhancements
  • If your technique is large, discuss the implementation plan in an issue first

Example Documentation Template

Your technique's README.md should include:

# Your Technique Name

Brief description of the optimization technique.

## Model Support

| Model/Framework | Supported | Notes |
|-----------------|-----------|-------|
| LLMs (Llama, GPT, etc.) | ✅ | Tested on Llama 3.1 |
| Diffusion Models | ❌ | Not yet supported |
| Vision Models | ✅ | Experimental |

## Deployment

| Framework | Supported | Notes |
|-----------|-----------|-------|
| TensorRT-LLM | ✅ | Requires custom kernel |
| vLLM | ❌ | Not yet supported |
| SGLang | ✅ | Uses standard ops |

## Usage

\`\`\`python
from experimental.my_technique import my_optimize
...
\`\`\`

## Status

Current state: Prototype

Known issues:
- Issue 1
- Issue 2

## References

- [Paper](link)
- [Code repository](link)
- [Project page](link)
- [Related work](link)

Path to Production

When a technique is ready for production (proven effective, stable API, full tests, comprehensive docs), it can be promoted to the main modelopt package.

Contributors: Open an issue proposing graduation with evidence of effectiveness and stability.

Users: If you find an experimental feature valuable, open a GitHub issue requesting promotion to production. User demand is a key signal for production readiness.

Questions?

Open a GitHub issue with [experimental] prefix.