Files
Zhehao-Hu 051d6adb20 Require PTQ recipe guidance before selecting quantization candidates (#2478)
### What does this PR do?

Type of change: documentation

Updates `quant-recipe-search` to require reading
`modelopt_recipes/ptq.md` before selecting initial or subsequent
quantization candidates.

The skill uses the guide to inform quantization scope, KV-cache scheme,
and calibration method. It requires inspecting candidate YAMLs and
supporting configs, citing relevant guide sections, and explaining
deviations. Existing coverage, runtime compatibility, and evaluation
checks remain required.

### Usage

Invoke `quant-recipe-search` as usual. The skill reads the guide from
the ModelOpt source checkout.

### Testing

- Passed a local Codex skill smoke test recommending an initial NVFP4
candidate for Qwen/Qwen3-8B targeting high-concurrency inference on
Blackwell.
- The test prompt named the skill without explicitly requesting that the
agent read ptq.md.
- Verified successful reads of the updated skill, ptq.md, the selected
recipe YAML, and supporting configs before the final recommendation.
- The agent selected general/ptq/nvfp4_mlp_only-kv_fp8_cast and cited
relevant guide sections. It explicitly left coverage, accuracy, and
throughput pending validation.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ✅

### Additional Information
This change places the reading requirement in the shared recipe-search
skill so consumers receive it directly. Consumers that pin the ModelOpt
plugin to a commit must update their pin to adopt it.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Updated the quantization recipe search workflow to require reviewing
the PTQ reference documentation before selecting a recipe.
- Clarified that available quantization schemes and model-specific
exceptions should be considered during recipe evaluation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Zhehao Hu <zhehaoh@nvidia.com>
2026-09-21 14:58:38 -07:00
..