### What does this PR do? Type of change: new feature Packages the existing ModelOpt agent skills as installable Codex and Claude plugins: - Adds a repo-scoped Codex marketplace and Claude-compatible marketplace. - Adds the canonical `plugins/modelopt/` plugin tree and manifests. - Moves the skill tree into the plugin and keeps `.agents/skills` as a compatibility symlink. - Adds a minimal `common` placeholder skill required by Codex validation. - Documents installation from this repository. ### Usage ```bash codex plugin marketplace add NVIDIA/Model-Optimizer ``` Then open `/plugins`, select the `modelopt` marketplace, and install `modelopt`. For Claude Code: ```bash claude plugin marketplace add https://github.com/NVIDIA/Model-Optimizer.git claude plugin install modelopt@modelopt ``` ### Testing - Codex plugin validator - `claude plugin validate . --strict` - `claude plugin validate plugins/modelopt --strict` 1. Install the marketplace plugin with Codex and Claude from an unrelated temporary workspace. 2. Exercise packaged evaluation helpers, a day-0 gate, and the shared remote helper from that workspace. 3. Run `uv run --frozen --extra dev python -m pytest -q plugins/modelopt/skills/day0-release/tests/test_gates.py plugins/modelopt/skills/benchmark-model-kernels/tests`. 4. Run pre-commit hooks for all changed files. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — added a plugin-path validator; existing focused skill tests and installed-plugin smoke tests pass. - Did you update Changelog?: N/A — agent tooling and distribution only. - Did you get Claude approval on this PR?: N/A ### Additional Information Skills remain available through `.agents/skills`; bundled helpers are packaged under the plugin and resolved from `$SKILL_DIR` so installed workflows do not depend on the current workspace. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added installable ModelOpt plugins for Claude Code and Codex. * Added skills for PTQ, deployment, evaluation, monitoring, debugging, benchmarking, MLflow access, EAGLE3 workflows, and release management. * Added deployment helpers, evaluation recipes, checkpoint validation, and release-gating tools. * **Documentation** * Expanded setup, credential, SLURM, benchmarking, deployment, evaluation, troubleshooting, and workspace guidance. * Added installation instructions and updated agent-skill discovery guidance. * **Maintenance** * Updated skill references and compatibility links for reliable use across supported plugin environments. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
3.2 KiB
name, description, user_invocable
| name | description | user_invocable |
|---|---|---|
| eagle3-review-logs | Review EAGLE3 pipeline experiment logs from the launcher's experiments/ directory. Summarizes pass/fail status for all 4 tasks, diagnoses failures with root causes and fixes, and flags warnings. Use when the user asks to review job logs, check experiment results, or diagnose why a specific task failed. | true |
Review EAGLE3 Experiment Logs
Analyze output logs from an EAGLE3 pipeline run launched via launch.py or slurm.py.
Step 0 — Find experiment logs
Locate the experiment directory. The default is experiments/ relative to the launcher root,
or wherever --job-dir was pointed.
ls -td experiments/cicd/cicd_* | head -10
If no experiments exist, ask the user for the directory.
Step 1 — Read all task logs
Each experiment has one subdirectory per task (0–3). Log filenames vary by launch mode
(Slurm writes sbatch_*.out, local Docker writes *.log), so match log files generally and
read the tail of each in a single Bash call — errors surface at the end:
find experiments/<exp_id>/ -type f \( -name '*.out' -o -name '*.log' \) | sort | while read -r f; do
echo "=== $f ==="; tail -200 "$f"; echo
done
Step 2 — Analyze
For each task log, check:
- Exit / cancellation:
DUE TO TIME LIMIT,FAILED, signal (e.g.,signal 15) - Python exceptions / tracebacks: last exception is usually the root cause
- CUDA errors: OOM, NCCL timeout
- Slurm state: COMPLETED, FAILED, TIMEOUT, OUT_OF_MEMORY
- Success indicators: "Saved N samples", "Successfully processed N conversations", training loss line, AR output
Step 3 — Produce report
Output a structured markdown report:
Summary
- Overall status: PASSED / FAILED / MIXED / PARTIAL
- Task breakdown: e.g., task_0 TIMEOUT, task_1 FAIL, task_2 skipped, task_3 skipped
Task Results
For each task (0–3):
Task N — <name>: PASS / FAIL / TIMEOUT
- Key output: (e.g., "3277/3295 samples generated" or "Script not found")
- Error (if failed): quoted error message, max 10 lines
- Root cause: one-line diagnosis
- Suggested fix: actionable step
Warnings
Non-fatal issues worth noting (near-OOM, tokenizer warnings, slow throughput).
Step 4 — Suggest next steps
Based on results:
-
If a task failed due to a known issue, suggest the fix and how to re-run from that task:
uv run launch.py --yaml examples/<Org>/<Model>/hf_offline_eagle3.yaml \ pipeline.task_0.skip=true \ --yes -
If the failure pattern looks new, suggest capturing it in the team's internal triage tracker, and use
/eagle3-triagefor a deeper diagnosis. -
If all tasks passed, suggest running
/eagle3-validateto confirm AR meets threshold.
Known benign patterns (do NOT mark as failures)
| Pattern | Explanation |
|---|---|
| vLLM server exit code 143 | SIGTERM — server was killed after queries completed. Expected. |
CANCELLED AT ... DUE TO TASK FAILURE after exit code: 0 |
Slurm cleanup of worker nodes after main task succeeded. |
destroy_process_group() was not called |
Benign PyTorch shutdown warning. |
tokenizer class ... not equal to the registered tokenizer class |
Harmless tokenizer mismatch warning. |