mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: documentation Updates `quant-recipe-search` to require reading `modelopt_recipes/ptq.md` before selecting initial or subsequent quantization candidates. The skill uses the guide to inform quantization scope, KV-cache scheme, and calibration method. It requires inspecting candidate YAMLs and supporting configs, citing relevant guide sections, and explaining deviations. Existing coverage, runtime compatibility, and evaluation checks remain required. ### Usage Invoke `quant-recipe-search` as usual. The skill reads the guide from the ModelOpt source checkout. ### Testing - Passed a local Codex skill smoke test recommending an initial NVFP4 candidate for Qwen/Qwen3-8B targeting high-concurrency inference on Blackwell. - The test prompt named the skill without explicitly requesting that the agent read ptq.md. - Verified successful reads of the updated skill, ptq.md, the selected recipe YAML, and supporting configs before the final recommendation. - The agent selected general/ptq/nvfp4_mlp_only-kv_fp8_cast and cited relevant guide sections. It explicitly left coverage, accuracy, and throughput pending validation. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ ### Additional Information This change places the reading requirement in the shared recipe-search skill so consumers receive it directly. Consumers that pin the ModelOpt plugin to a commit must update their pin to adopt it. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Documentation** - Updated the quantization recipe search workflow to require reviewing the PTQ reference documentation before selecting a recipe. - Clarified that available quantization schemes and model-specific exceptions should be considered during recipe evaluation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Zhehao Hu <zhehaoh@nvidia.com>