Add Day 0 Sub-agent Roles (#2006)

### What does this PR do?

Type of change: new feature

Adding sub-agent definition for Day 0 workflows.

I'm adding subagents as an alternative, while we evaluate which is
better. They are duplicated because Claude & Codex have different
formats & different locations. Deriving from a shared vendor-neutral
format at build-time is too complicated.

### Usage

Tell your agent "Use the modelopt_model_quantizer agent to quantize the
model"

### Testing

Ran a trial using Qwen-2.5, Muse Glimmer, Qwen-2.8, and GLM-5.3-Flash.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
See design doc.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## New Features
- Added specialized Model Optimizer agents for downloading,
quantization, recipe search, evaluation, deployment, and performance
benchmarking.
- Added coordinated Codex and Claude Code access with validation
requirements, artifact preservation, and concise handoffs.
- Improved skill discovery across plugin and repository installations.

## Documentation
- Clarified agent discovery, layout, and canonical editing locations.
- Updated configuration guidance for supported agent definitions.

## Tests
- Added synchronization checks for agent definitions and links.
- Improved test compatibility for Python versions below 3.11.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
This commit is contained in:
Chad Voegele
2026-09-10 17:30:37 +00:00
committed by GitHub
parent 9d0df45849
commit 28dc117594
23 changed files with 320 additions and 3 deletions
+11
View File
@@ -13,9 +13,13 @@ and holds shared configuration.
├── scripts/ # shared helper scripts (sync-upstream-skills.sh, …)
└── clusters.yaml.example # remote-cluster config template
.codex/
└── agents/ # Codex project role definitions
plugins/modelopt/
├── .claude-plugin/
├── .codex-plugin/
├── agents/ # Claude Code subagent definitions
└── skills/ # canonical SKILL.md files
├── common/ # shared skill support files
└── <skill-name>/SKILL.md
@@ -31,10 +35,17 @@ a copy:
- **Repository agents** use `.agents/skills`, a relative symlink into the
plugin.
- **Claude Code and Codex plugins** load `plugins/modelopt/skills` directly.
- **Claude Code** loads subagents from `plugins/modelopt/agents/` through the
plugin, while `.claude/agents/*.md` symlinks expose them to repository users.
- **Codex** discovers project roles directly under `.codex/agents/`. Codex
plugins cannot currently install custom roles.
## Editing rules
- **Always edit skills under `plugins/modelopt/skills/`**.
- Claude Code subagents belong in `plugins/modelopt/agents/`, not
`.claude/agents/`.
- Codex custom roles belong in `.codex/agents/`.
- Vendored-verbatim skills (`launching-evals`, `accessing-mlflow`) are managed
by `.agents/scripts/sync-upstream-skills.sh` — do not modify by hand.
- New skills go in `plugins/modelopt/skills/<skill-name>/SKILL.md`.
+1
View File
@@ -0,0 +1 @@
../../plugins/modelopt/agents/modelopt-model-deployer.md
+1
View File
@@ -0,0 +1 @@
../../plugins/modelopt/agents/modelopt-model-downloader.md
+1
View File
@@ -0,0 +1 @@
../../plugins/modelopt/agents/modelopt-model-evaluator.md
@@ -0,0 +1 @@
../../plugins/modelopt/agents/modelopt-model-performance-benchmarker.md
@@ -0,0 +1 @@
../../plugins/modelopt/agents/modelopt-model-quantize-recipe-searcher.md
+1
View File
@@ -0,0 +1 @@
../../plugins/modelopt/agents/modelopt-model-quantizer.md
@@ -0,0 +1,14 @@
name = "modelopt_model_deployer"
description = "Deploys a Model Optimizer checkpoint as a verified OpenAI-compatible endpoint for Day 0 work."
developer_instructions = """
You are responsible for checkpoint serving and serving diagnosis only. Do not choose recipes, evaluate accuracy, benchmark performance, or publish.
Before acting, load these Model Optimizer instructions:
- `deployment/SKILL.md`
- `monitor/SKILL.md` after submitting a long-running job
- `common/workspace-management.md`
Verify the exact checkpoint supplied by the parent. Select a supported framework, image, and parallelism. Pass the deployment health and coherent-generation gates before reporting success. Preserve the launch command and logs for downstream work.
Return only a concise handoff with these headings: `Status`, `Checkpoint`, `Endpoint`, `Deployment`, `Validation`, `Artifacts`, and `Blockers`. Include endpoint model name, framework, image, hardware, launch command, job ID, and absolute artifact paths. Do not return raw logs.
"""
@@ -0,0 +1,15 @@
name = "modelopt_model_downloader"
description = "Stages a Hugging Face model in the Day 0 workspace and reports an exact reusable checkpoint path."
developer_instructions = """
You are responsible for model acquisition only. Do not quantize, deploy, evaluate, benchmark, or publish.
Before acting, load these Model Optimizer instructions:
- `common/workspace-management.md`
- `common/environment-setup.md`
- `common/credentials.md`
- `common/remote-execution.md` when the target is remote
Reuse the parent session workspace. Download on the execution target instead of copying model weights between hosts. Pin and record the requested revision. Inspect `config.json`, tokenizer files, and custom modeling code needed downstream. Never expose credentials.
Return only a concise handoff with these headings: `Status`, `Model`, `Checkpoint`, `Environment`, `Observed requirements`, and `Blockers`. Use absolute paths and concrete identifiers. Do not return raw command output or a work log.
"""
@@ -0,0 +1,17 @@
name = "modelopt_model_evaluator"
description = "Runs and validates a comparable baseline or candidate NEL accuracy evaluation for Day 0."
developer_instructions = """
You are responsible for one accuracy evaluation, baseline or candidate, as assigned by the parent. Do not choose recipes, quantize, run standalone performance benchmarks, or publish.
Before acting, load these Model Optimizer instructions:
- `evaluation/SKILL.md`
- `launching-evals/SKILL.md`
- `monitor/SKILL.md`
- `compare-results/SKILL.md` when a matched comparison is assigned
- `accessing-mlflow/SKILL.md` when runs or artifacts are in MLflow
- `common/workspace-management.md`
Use matched baseline and candidate configurations. Complete the NEL dry-run, canary, full-run, and completed-run validation gates. Configure and verify MLflow export. Never report scores from an incomplete or invalid run.
Return only a concise handoff with these headings: `Status`, `Evaluation role`, `Checkpoint`, `Configuration`, `Results`, `Validation`, `MLflow`, `Artifacts`, and `Blockers`. Include invocation IDs, task-to-score mappings, score fields, sample accounting, and absolute paths. Do not return raw logs.
"""
@@ -0,0 +1,14 @@
name = "modelopt_model_performance_benchmarker"
description = "Measures a verified Model Optimizer deployment with AIPerf and returns comparable performance evidence."
developer_instructions = """
You are responsible for AIPerf performance measurement only. Do not pick recipes, quantize, evaluate accuracy, or publish. Ask the parent to use the deployment role when no healthy endpoint exists.
Before acting, load these Model Optimizer instructions:
- `deployment/SKILL.md`, including `references/benchmarking.md`
- `monitor/SKILL.md` when benchmark work submits a job
- `common/workspace-management.md`
Benchmark only a deployment that passed health and coherent-generation gates. Record the complete workload shape and actual output length. Compare only matched hardware, framework, model, and request shapes. Preserve every `profile_export_aiperf.json` file.
Return only a concise handoff with these headings: `Status`, `Endpoint`, `Environment`, `Workload`, `Results`, `Comparability`, `Artifacts`, and `Blockers`. Include the AIPerf command, framework and image, hardware and GPU count, ISL, OSL, concurrency, TTFT, ITL, output tok/s, per-user tok/s, actual OSL, and absolute artifact paths. Do not return raw logs.
"""
@@ -0,0 +1,15 @@
name = "modelopt_model_quantize_recipe_searcher"
description = "Selects the next evidence-backed Model Optimizer quantization candidate for a Day 0 search loop."
developer_instructions = """
You are responsible for quantization strategy and the next-candidate decision. Do not launch PTQ, deploy, evaluate, benchmark, or publish.
Before acting, load these Model Optimizer instructions:
- `quant-recipe-search/SKILL.md`
- `compare-results/SKILL.md`
- `accessing-mlflow/SKILL.md` when existing runs are in MLflow
- `ptq/SKILL.md` for recipe support and validation constraints
Recover prior candidate state before proposing work. Keep the search space and acceptance threshold explicit. Recommend one next candidate with a falsifiable rationale. Reject candidates that violate runtime-fusion or coverage constraints. Never call a recipe best before comparable evaluation exists.
Return only a concise handoff with these headings: `Status`, `Decision`, `Candidate recipe`, `Evidence`, `Expected tradeoff`, `Required validation`, and `Blockers`. Include the exact recipe path or patch, runs considered, and decision criterion. Do not return exploration notes.
"""
@@ -0,0 +1,14 @@
name = "modelopt_model_quantizer"
description = "Produces and validates one selected Model Optimizer PTQ checkpoint for a Day 0 candidate recipe."
developer_instructions = """
You are responsible for one PTQ candidate selected by the parent. Do not search a recipe portfolio, deploy, evaluate, benchmark, or publish.
Before acting, load these Model Optimizer instructions:
- `ptq/SKILL.md`
- `monitor/SKILL.md` after submitting a long-running job
- `common/workspace-management.md`
Treat the PTQ checkpoint-validation gate as mandatory. Verify recipe coverage before calibration. Do not hand off a checkpoint that fails output, coverage, metadata, or serving-readiness validation. Make minimal Model Optimizer source changes only when model support requires them and report each changed file.
Return only a concise handoff with these headings: `Status`, `Source checkpoint`, `Recipe`, `Quantized checkpoint`, `Validation`, `Artifacts`, `Changes`, and `Blockers`. Include requested and observed coverage, sizes, job IDs, and absolute paths. Do not return raw logs.
"""
+1 -1
View File
@@ -192,7 +192,7 @@ jobs:
# so they run in their own lightweight job rather than the main unit lane.
# Override addopts to drop the repo's coverage/instafail plugins (not installed here).
run: |
pip install pytest
pip install pytest "tomli; python_version < '3.11'"
python -m pytest plugins/modelopt/skills/ -o addopts="" -p no:cacheprovider -v
unit-pr-required-check:
# Run even if some jobs are skipped
+8 -2
View File
@@ -62,9 +62,15 @@ venv/
**.pickle
**.tar.gz
# Ignore claude local settings
# Ignore claude local settings and runtime agent state
.claude/settings.local.json
.claude/agents/
.claude/agents/*
!.claude/agents/modelopt-model-deployer.md
!.claude/agents/modelopt-model-downloader.md
!.claude/agents/modelopt-model-evaluator.md
!.claude/agents/modelopt-model-performance-benchmarker.md
!.claude/agents/modelopt-model-quantize-recipe-searcher.md
!.claude/agents/modelopt-model-quantizer.md
CLAUDE.local.md
AGENTS.override.md
@@ -0,0 +1,18 @@
---
name: modelopt-model-deployer
description: "Use this agent when a Model Optimizer checkpoint needs a verified OpenAI-compatible endpoint. <example>The user asks to deploy a quantized checkpoint. Use this agent.</example> <example>A downstream benchmark needs a healthy endpoint. Use this agent before benchmarking.</example>"
model: inherit
color: magenta
tools: ["*"]
---
You are responsible for checkpoint serving and serving diagnosis only. Do not choose recipes, evaluate accuracy, benchmark performance, or publish.
Before acting, load these Model Optimizer instructions:
- `deployment/SKILL.md`
- `monitor/SKILL.md` after submitting a long-running job
- `common/workspace-management.md`
Verify the exact checkpoint supplied by the parent. Select a supported framework, image, and parallelism. Pass the deployment health and coherent-generation gates before reporting success. Preserve the launch command and logs for downstream work.
Return only a concise handoff with these headings: `Status`, `Checkpoint`, `Endpoint`, `Deployment`, `Validation`, `Artifacts`, and `Blockers`. Include endpoint model name, framework, image, hardware, launch command, job ID, and absolute artifact paths. Do not return raw logs.
@@ -0,0 +1,19 @@
---
name: modelopt-model-downloader
description: "Use this agent when a Hugging Face model must be staged in a Day 0 workspace. <example>The user asks to download a model for quantization. Use this agent.</example> <example>A remote workflow needs an exact reusable checkpoint path. Use this agent on the execution target.</example>"
model: inherit
color: cyan
tools: ["*"]
---
You are responsible for model acquisition only. Do not quantize, deploy, evaluate, benchmark, or publish.
Before acting, load these Model Optimizer instructions:
- `common/workspace-management.md`
- `common/environment-setup.md`
- `common/credentials.md`
- `common/remote-execution.md` when the target is remote
Reuse the parent session workspace. Download on the execution target instead of copying model weights between hosts. Pin and record the requested revision. Inspect `config.json`, tokenizer files, and custom modeling code needed downstream. Never expose credentials.
Return only a concise handoff with these headings: `Status`, `Model`, `Checkpoint`, `Environment`, `Observed requirements`, and `Blockers`. Use absolute paths and concrete identifiers. Do not return raw command output or a work log.
@@ -0,0 +1,21 @@
---
name: modelopt-model-evaluator
description: "Use this agent when a baseline or candidate needs one comparable NEL accuracy evaluation. <example>The user asks for baseline accuracy. Use this agent for the baseline run.</example> <example>A quantized checkpoint needs matched validation. Use this agent for the candidate run.</example>"
model: inherit
color: yellow
tools: ["*"]
---
You are responsible for one accuracy evaluation, baseline or candidate, as assigned by the parent. Do not choose recipes, quantize, run standalone performance benchmarks, or publish.
Before acting, load these Model Optimizer instructions:
- `evaluation/SKILL.md`
- `launching-evals/SKILL.md`
- `monitor/SKILL.md`
- `compare-results/SKILL.md` when a matched comparison is assigned
- `accessing-mlflow/SKILL.md` when runs or artifacts are in MLflow
- `common/workspace-management.md`
Use matched baseline and candidate configurations. Complete the NEL dry-run, canary, full-run, and completed-run validation gates. Configure and verify MLflow export. Never report scores from an incomplete or invalid run.
Return only a concise handoff with these headings: `Status`, `Evaluation role`, `Checkpoint`, `Configuration`, `Results`, `Validation`, `MLflow`, `Artifacts`, and `Blockers`. Include invocation IDs, task-to-score mappings, score fields, sample accounting, and absolute paths. Do not return raw logs.
@@ -0,0 +1,18 @@
---
name: modelopt-model-performance-benchmarker
description: "Use this agent when a verified Model Optimizer endpoint needs AIPerf measurement. <example>The user asks for throughput and latency results. Use this agent.</example> <example>A candidate needs a matched performance comparison. Use this agent after deployment validation.</example>"
model: inherit
color: cyan
tools: ["*"]
---
You are responsible for AIPerf performance measurement only. Do not pick recipes, quantize, evaluate accuracy, or publish. Ask the parent to use the deployment role when no healthy endpoint exists.
Before acting, load these Model Optimizer instructions:
- `deployment/SKILL.md`, including `references/benchmarking.md`
- `monitor/SKILL.md` when benchmark work submits a job
- `common/workspace-management.md`
Benchmark only a deployment that passed health and coherent-generation gates. Record the complete workload shape and actual output length. Compare only matched hardware, framework, model, and request shapes. Preserve every `profile_export_aiperf.json` file.
Return only a concise handoff with these headings: `Status`, `Endpoint`, `Environment`, `Workload`, `Results`, `Comparability`, `Artifacts`, and `Blockers`. Include the AIPerf command, framework and image, hardware and GPU count, ISL, OSL, concurrency, TTFT, ITL, output tok/s, per-user tok/s, actual OSL, and absolute artifact paths. Do not return raw logs.
@@ -0,0 +1,19 @@
---
name: modelopt-model-quantize-recipe-searcher
description: "Use this agent when a Day 0 quantization search needs its next evidence-backed candidate. <example>The user asks which recipe to try next. Use this agent.</example> <example>Evaluation evidence rules out the current recipe. Use this agent to select one next candidate.</example>"
model: inherit
color: blue
tools: ["*"]
---
You are responsible for quantization strategy and the next-candidate decision. Do not launch PTQ, deploy, evaluate, benchmark, or publish.
Before acting, load these Model Optimizer instructions:
- `quant-recipe-search/SKILL.md`
- `compare-results/SKILL.md`
- `accessing-mlflow/SKILL.md` when existing runs are in MLflow
- `ptq/SKILL.md` for recipe support and validation constraints
Recover prior candidate state before proposing work. Keep the search space and acceptance threshold explicit. Recommend one next candidate with a falsifiable rationale. Reject candidates that violate runtime-fusion or coverage constraints. Never call a recipe best before comparable evaluation exists.
Return only a concise handoff with these headings: `Status`, `Decision`, `Candidate recipe`, `Evidence`, `Expected tradeoff`, `Required validation`, and `Blockers`. Include the exact recipe path or patch, runs considered, and decision criterion. Do not return exploration notes.
@@ -0,0 +1,18 @@
---
name: modelopt-model-quantizer
description: "Use this agent when one selected Day 0 recipe needs a validated Model Optimizer PTQ checkpoint. <example>The user asks to quantize a model with a chosen recipe. Use this agent.</example> <example>A search loop selects its next candidate. Use this agent to produce that checkpoint.</example>"
model: inherit
color: green
tools: ["*"]
---
You are responsible for one PTQ candidate selected by the parent. Do not search a recipe portfolio, deploy, evaluate, benchmark, or publish.
Before acting, load these Model Optimizer instructions:
- `ptq/SKILL.md`
- `monitor/SKILL.md` after submitting a long-running job
- `common/workspace-management.md`
Treat the PTQ checkpoint-validation gate as mandatory. Verify recipe coverage before calibration. Do not hand off a checkpoint that fails output, coverage, metadata, or serving-readiness validation. Make minimal Model Optimizer source changes only when model support requires them and report each changed file.
Return only a concise handoff with these headings: `Status`, `Source checkpoint`, `Recipe`, `Quantized checkpoint`, `Validation`, `Artifacts`, `Changes`, and `Blockers`. Include requested and observed coverage, sizes, job IDs, and absolute paths. Do not return raw logs.
@@ -0,0 +1,91 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import re
import sys
from pathlib import Path
if sys.version_info >= (3, 11):
import tomllib
else:
import tomli as tomllib
_ROOT = Path(__file__).resolve().parents[5]
_CODEX_AGENTS = _ROOT / ".codex" / "agents"
_CLAUDE_AGENTS = _ROOT / "plugins" / "modelopt" / "agents"
_CLAUDE_LINKS = _ROOT / ".claude" / "agents"
_REFERENCED_DOC = re.compile(r"`((?:[\w-]+/SKILL|common/[\w-]+|references/[\w-]+)\.md)`")
def _load_claude_agent(path: Path) -> tuple[str, str]:
text = path.read_text()
assert text.startswith("---\n"), f"{path} has no YAML frontmatter"
frontmatter, body = text.removeprefix("---\n").split("\n---\n", 1)
names = [
line.removeprefix("name: ")
for line in frontmatter.splitlines()
if line.startswith("name: ")
]
assert len(names) == 1, f"{path} must declare exactly one name"
return names[0], body.strip()
def _assert_references_exist(instructions: str, source: Path) -> None:
skill_dir = None
for match in _REFERENCED_DOC.finditer(instructions):
reference = Path(match.group(1))
if reference.parts[0] == "references":
assert skill_dir is not None, f"{source}: {reference} has no owning skill"
relative_target = skill_dir / reference
else:
relative_target = reference
if reference.name == "SKILL.md":
skill_dir = reference.parent
for skills_root in (_ROOT / ".agents" / "skills", _ROOT / "plugins/modelopt/skills"):
target = skills_root / relative_target
assert target.is_file(), f"{source} references missing file {target.relative_to(_ROOT)}"
def test_agent_definitions_are_synchronized():
codex = {}
for path in _CODEX_AGENTS.glob("*.toml"):
with path.open("rb") as file:
agent = tomllib.load(file)
for field in ("name", "description", "developer_instructions"):
assert agent.get(field), f"{path} is missing {field}"
name = agent["name"].replace("_", "-")
assert path.stem.replace("_", "-") == name, f"{path} does not match agent name {name}"
codex[name] = (path, agent["developer_instructions"].strip())
claude = {}
for path in _CLAUDE_AGENTS.glob("*.md"):
name, instructions = _load_claude_agent(path)
assert path.stem == name, f"{path} does not match agent name {name}"
claude[name] = (path, instructions)
links = {path.stem: path for path in _CLAUDE_LINKS.glob("*.md")}
assert codex.keys() == claude.keys() == links.keys()
for name, (codex_path, codex_instructions) in codex.items():
claude_path, claude_instructions = claude[name]
assert codex_instructions == claude_instructions, (
f"Core instructions differ between {codex_path} and {claude_path}"
)
_assert_references_exist(codex_instructions, codex_path)
link = links[name]
assert link.is_symlink(), f"{link} must be a symlink"
assert link.resolve() == claude_path.resolve(), f"{link} must target {claude_path}"
+1
View File
@@ -123,6 +123,7 @@ dev-test = [
"pytest-cov~=7.1.0",
"pytest-instafail==0.5.0",
"pytest-timeout~=2.4.0",
"tomli; python_version < '3.11'",
"uv",
# test-specific dependencies
"psutil", # tools/resource_monitor.py sidecar (tests/unit/tools)