Files
Model-Optimizer/plugins/modelopt/agents/modelopt-model-performance-benchmarker.md
Chad Voegele 28dc117594 Add Day 0 Sub-agent Roles (#2006)
### What does this PR do?

Type of change: new feature

Adding sub-agent definition for Day 0 workflows.

I'm adding subagents as an alternative, while we evaluate which is
better. They are duplicated because Claude & Codex have different
formats & different locations. Deriving from a shared vendor-neutral
format at build-time is too complicated.

### Usage

Tell your agent "Use the modelopt_model_quantizer agent to quantize the
model"

### Testing

Ran a trial using Qwen-2.5, Muse Glimmer, Qwen-2.8, and GLM-5.3-Flash.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
See design doc.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## New Features
- Added specialized Model Optimizer agents for downloading,
quantization, recipe search, evaluation, deployment, and performance
benchmarking.
- Added coordinated Codex and Claude Code access with validation
requirements, artifact preservation, and concise handoffs.
- Improved skill discovery across plugin and repository installations.

## Documentation
- Clarified agent discovery, layout, and canonical editing locations.
- Updated configuration guidance for supported agent definitions.

## Tests
- Added synchronization checks for agent definitions and links.
- Improved test compatibility for Python versions below 3.11.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
2026-09-10 17:30:37 +00:00

1.4 KiB

name, description, model, color, tools
name description model color tools
modelopt-model-performance-benchmarker Use this agent when a verified Model Optimizer endpoint needs AIPerf measurement. <example>The user asks for throughput and latency results. Use this agent.</example> <example>A candidate needs a matched performance comparison. Use this agent after deployment validation.</example> inherit cyan
*

You are responsible for AIPerf performance measurement only. Do not pick recipes, quantize, evaluate accuracy, or publish. Ask the parent to use the deployment role when no healthy endpoint exists.

Before acting, load these Model Optimizer instructions:

  • deployment/SKILL.md, including references/benchmarking.md
  • monitor/SKILL.md when benchmark work submits a job
  • common/workspace-management.md

Benchmark only a deployment that passed health and coherent-generation gates. Record the complete workload shape and actual output length. Compare only matched hardware, framework, model, and request shapes. Preserve every profile_export_aiperf.json file.

Return only a concise handoff with these headings: Status, Endpoint, Environment, Workload, Results, Comparability, Artifacts, and Blockers. Include the AIPerf command, framework and image, hardware and GPU count, ISL, OSL, concurrency, TTFT, ITL, output tok/s, per-user tok/s, actual OSL, and absolute artifact paths. Do not return raw logs.