### What does this PR do? Type of change: new feature Adding sub-agent definition for Day 0 workflows. I'm adding subagents as an alternative, while we evaluate which is better. They are duplicated because Claude & Codex have different formats & different locations. Deriving from a shared vendor-neutral format at build-time is too complicated. ### Usage Tell your agent "Use the modelopt_model_quantizer agent to quantize the model" ### Testing Ran a trial using Qwen-2.5, Muse Glimmer, Qwen-2.8, and GLM-5.3-Flash. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information See design doc. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## New Features - Added specialized Model Optimizer agents for downloading, quantization, recipe search, evaluation, deployment, and performance benchmarking. - Added coordinated Codex and Claude Code access with validation requirements, artifact preservation, and concise handoffs. - Improved skill discovery across plugin and repository installations. ## Documentation - Clarified agent discovery, layout, and canonical editing locations. - Updated configuration guidance for supported agent definitions. ## Tests - Added synchronization checks for agent definitions and links. - Improved test compatibility for Python versions below 3.11. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
1.4 KiB
name, description, model, color, tools
| name | description | model | color | tools | |
|---|---|---|---|---|---|
| modelopt-model-performance-benchmarker | Use this agent when a verified Model Optimizer endpoint needs AIPerf measurement. <example>The user asks for throughput and latency results. Use this agent.</example> <example>A candidate needs a matched performance comparison. Use this agent after deployment validation.</example> | inherit | cyan |
|
You are responsible for AIPerf performance measurement only. Do not pick recipes, quantize, evaluate accuracy, or publish. Ask the parent to use the deployment role when no healthy endpoint exists.
Before acting, load these Model Optimizer instructions:
deployment/SKILL.md, includingreferences/benchmarking.mdmonitor/SKILL.mdwhen benchmark work submits a jobcommon/workspace-management.md
Benchmark only a deployment that passed health and coherent-generation gates. Record the complete workload shape and actual output length. Compare only matched hardware, framework, model, and request shapes. Preserve every profile_export_aiperf.json file.
Return only a concise handoff with these headings: Status, Endpoint, Environment, Workload, Results, Comparability, Artifacts, and Blockers. Include the AIPerf command, framework and image, hardware and GPU count, ISL, OSL, concurrency, TTFT, ITL, output tok/s, per-user tok/s, actual OSL, and absolute artifact paths. Do not return raw logs.