Files
Model-Optimizer/docs/source
Chenjie LuoandClaude Opus 4.8 56c4af2333 feat(recipes): add kv_fp8_cast variants for partial-NVFP4 and weight-only PTQ recipes (#1652)
### What does this PR do?

Type of change: new feature (recipes)

Several `general/ptq` recipe families shipped a data-driven FP8 KV-cache
(`-kv_fp8`) variant but lacked the constant-amax `kv_fp8_cast` companion
that `fp8_default` and `nvfp4_default` already have. This PR adds the
missing cast variants so every KV-quantizing (and the weight-only)
family offers the calibration-free FP8 KV-cache option:

- `general/ptq/nvfp4_experts_only-kv_fp8_cast`
- `general/ptq/nvfp4_mlp_only-kv_fp8_cast`
- `general/ptq/nvfp4_omlp_only-kv_fp8_cast`
- `general/ptq/nvfp4_weight_only-kv_fp8_cast`

Each new recipe composes the exact same model-quant config as its
existing sibling and swaps the `kv_fp8` unit for the shared
`kv_fp8_cast` unit (constant-amax FP8 KV cache; no KV calibration
forward pass). The docs guide table/tree and the changelog are updated
to match.

### Usage

```bash
python examples/llm_ptq/hf_ptq.py \
    --pyt_ckpt_path <model> \
    --recipe general/ptq/nvfp4_mlp_only-kv_fp8_cast
```

### Testing

Extended the built-in PTQ smoke test
`tests/unit/recipe/test_loader.py::test_load_recipe_all_builtins` with
the four new recipe paths; all four load into a valid
`ModelOptPTQRecipe` with a populated `quantize` section.

```
$ python -m pytest tests/unit/recipe/test_loader.py tests/unit/recipe/test_presets.py -q
180 passed
```

`pre-commit` (including the `validate modelopt recipes` hook) passes on
all changed files.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (additive — only new recipe
files)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ (extended the builtin recipe
smoke test)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ (not yet)

### Additional Information

The two weight-only families were discussed for scope;
`nvfp4_weight_only` is included (it already names a KV mode, `kv_fp16`),
while `int4_blockwise_weight_only` is intentionally left untouched since
it carries no `-kv_` composition.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added four new NVFP4 PTQ (Post-Training Quantization) recipe variants:
experts-only, MLP-only, OMLP-only, and weight-only configurations.
* All new recipes include FP8 KV-cache cast mode support for improved
inference performance.

* **Documentation**
* Updated built-in recipes guide with new NVFP4 recipe options and
repository layout.

* **Tests**
  * Expanded recipe loader test coverage for new recipe configurations.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 16:13:25 -07:00
..
2025-06-05 13:24:07 -07:00
2025-03-03 23:02:22 +05:30