## What does this PR do?
**Type of change:** new feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:**
Experimental Conv3D implicit-GEMM CUDA kernel with optional NVFP4-style
(E2M1 + FP8 E4M3 scale) fake quantization for activations.
It is intended for research/prototyping and quantization-accuracy
experiments only, not production deployment.
The implementation runs as a JIT-compiled PyTorch extension, mirrors
conv3d output shape, and provides a quantized and non-quantized path to
compare numerical behavior.
There is currently no real quantized production kernel integration in
the formal ModelOpt export/compress/runtime stack; this path is kept in
experimental/ for fake-quant accuracy validation and benchmarking.
## Usage
<!-- You can potentially add a usage example below. -->
```python
import torch
from experimental.conv.implicit_gemm_cuda import conv3d_implicit_gemm_cuda
from modelopt.torch.quantization.tensor_quant import dynamic_block_quantize_op
x = torch.randn(1, 128, 21, 60, 106, device="cuda")
w = torch.randn(512, 128, 3, 3, 3, device="cuda")
block_size = 128
# Without FP4 activation quantization (drop-in-style Conv3D call)
out = conv3d_implicit_gemm_cuda(x, w, stride=(1, 1, 1), padding=(1, 1, 1))
# Optional FP4 block quantization of weights along the GEMM K dimension.
# The kernel's A-tile (activations) is quantized along K = Cin*kD*kH*kW,
# so weights must be flattened to [Cout, K] before quantizing to match.
Cout, Cin = w.shape[:2]
K = Cin * w.shape[2] * w.shape[3] * w.shape[4]
w_flat = w.reshape(Cout, K)
w_q_flat = dynamic_block_quantize_op(
w_flat,
block_size,
w_flat.abs().max().unsqueeze(0),
4, # num_bits
2, # exponent_bits
8, # scale_num_bits
4, # scale_exponent_bits
)
w_q = w_q_flat.reshape_as(w)
# With FP4 activation fake quantization
out_q = conv3d_implicit_gemm_cuda(
x,
w_q,
stride=(1, 1, 1),
padding=(1, 1, 1),
act_amax=x.abs().max().unsqueeze(0),
quant_act=True,
fp4_block_size=block_size, # 128 or 256
)
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added experimental Conv3D implementation with implicit GEMM
acceleration and optional FP4 quantization support
* Added benchmarking tool to compare 3D convolution performance across
implementations
* Enhanced quantization framework integration for Conv3D operations
* **Documentation**
* Added comprehensive guide for experimental Conv3D prototype, including
supported scenarios, API reference, and current limitations
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Experimental Optimization Techniques
Experimental optimization algorithms and research prototypes under active development.
Purpose
For new optimization techniques (quantization, pruning, sparsity, etc.) that are:
- Novel or research-stage algorithms
- Not yet production-ready
- May have unstable APIs
⚠️ Warning: Experimental features are not guaranteed to work across releases. APIs may change or features may be removed without notice. Use at your own risk.
Requirements
Each experimental technique must include:
- README.md - Explains what the technique does, how to use it, current status, model support, and references
- Working code - Clear, readable implementation
- Comprehensive tests - Good test coverage demonstrating correctness
- Detailed documentation - Clear docs on usage, APIs, and behavior
- Example - Demonstrating usage
- Model support list - Which models/frameworks are supported
- Deployment info - Supported deployment frameworks (TensorRT-LLM, vLLM, SGLang, etc.) and whether custom kernels are required
- requirements.txt - Additional dependencies beyond base modelopt
- License headers - Apache 2.0 headers on all Python files
Example Structures
Organize your code however makes sense. Here are some examples:
Simple flat structure:
experimental/my_technique/
├── README.md
├── requirements.txt
├── my_technique.py
├── test_my_technique.py
└── example.py
Package structure:
experimental/my_technique/
├── README.md
├── requirements.txt
├── my_technique/
│ ├── __init__.py
│ ├── core.py
│ └── config.py
├── tests/
│ └── test_core.py
└── examples/
└── example_usage.py
Quality Standards
Experimental code must meet quality standards:
- Comprehensive test coverage required
- Clear documentation required
- Pass all pre-commit checks
PR Guidelines
Keep PRs focused and reviewable:
- Split large features: Break complex techniques into multiple PRs if needed
- Reasonable scope: PRs with tens of thousands of lines are difficult to review
- Incremental development: Consider submitting core functionality first, then enhancements
- If your technique is large, discuss the implementation plan in an issue first
Example Documentation Template
Your technique's README.md should include:
# Your Technique Name
Brief description of the optimization technique.
## Model Support
| Model/Framework | Supported | Notes |
|-----------------|-----------|-------|
| LLMs (Llama, GPT, etc.) | ✅ | Tested on Llama 3.1 |
| Diffusion Models | ❌ | Not yet supported |
| Vision Models | ✅ | Experimental |
## Deployment
| Framework | Supported | Notes |
|-----------|-----------|-------|
| TensorRT-LLM | ✅ | Requires custom kernel |
| vLLM | ❌ | Not yet supported |
| SGLang | ✅ | Uses standard ops |
## Usage
\`\`\`python
from experimental.my_technique import my_optimize
...
\`\`\`
## Status
Current state: Prototype
Known issues:
- Issue 1
- Issue 2
## References
- [Paper](link)
- [Code repository](link)
- [Project page](link)
- [Related work](link)
Path to Production
When a technique is ready for production (proven effective, stable API, full tests, comprehensive docs), it can be promoted to the main modelopt package.
Contributors: Open an issue proposing graduation with evidence of effectiveness and stability.
Users: If you find an experimental feature valuable, open a GitHub issue requesting promotion to production. User demand is a key signal for production readiness.
Questions?
Open a GitHub issue with [experimental] prefix.