### What does this PR do? Type of change: new feature Packages the existing ModelOpt agent skills as installable Codex and Claude plugins: - Adds a repo-scoped Codex marketplace and Claude-compatible marketplace. - Adds the canonical `plugins/modelopt/` plugin tree and manifests. - Moves the skill tree into the plugin and keeps `.agents/skills` as a compatibility symlink. - Adds a minimal `common` placeholder skill required by Codex validation. - Documents installation from this repository. ### Usage ```bash codex plugin marketplace add NVIDIA/Model-Optimizer ``` Then open `/plugins`, select the `modelopt` marketplace, and install `modelopt`. For Claude Code: ```bash claude plugin marketplace add https://github.com/NVIDIA/Model-Optimizer.git claude plugin install modelopt@modelopt ``` ### Testing - Codex plugin validator - `claude plugin validate . --strict` - `claude plugin validate plugins/modelopt --strict` 1. Install the marketplace plugin with Codex and Claude from an unrelated temporary workspace. 2. Exercise packaged evaluation helpers, a day-0 gate, and the shared remote helper from that workspace. 3. Run `uv run --frozen --extra dev python -m pytest -q plugins/modelopt/skills/day0-release/tests/test_gates.py plugins/modelopt/skills/benchmark-model-kernels/tests`. 4. Run pre-commit hooks for all changed files. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — added a plugin-path validator; existing focused skill tests and installed-plugin smoke tests pass. - Did you update Changelog?: N/A — agent tooling and distribution only. - Did you get Claude approval on this PR?: N/A ### Additional Information Skills remain available through `.agents/skills`; bundled helpers are packaged under the plugin and resolved from `$SKILL_DIR` so installed workflows do not depend on the current workspace. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added installable ModelOpt plugins for Claude Code and Codex. * Added skills for PTQ, deployment, evaluation, monitoring, debugging, benchmarking, MLflow access, EAGLE3 workflows, and release management. * Added deployment helpers, evaluation recipes, checkpoint validation, and release-gating tools. * **Documentation** * Expanded setup, credential, SLURM, benchmarking, deployment, evaluation, troubleshooting, and workspace guidance. * Added installation instructions and updated agent-skill discovery guidance. * **Maintenance** * Updated skill references and compatibility links for reliable use across supported plugin environments. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
2.9 KiB
Environment Setup
Common detection for all ModelOpt skills. After this, you know what's available.
Env-1. Get ModelOpt source
ls examples/hf_ptq/hf_ptq.py 2>/dev/null && echo "Source found"
If not found: git clone https://github.com/NVIDIA/Model-Optimizer.git && cd Model-Optimizer
If found, ensure the source is up to date:
git pull origin main
If previous runs left patches in modelopt/ (from 4C unlisted model work), check whether they should be kept. Reset only if starting a completely new task: git checkout main.
Env-2. Local or remote?
- User explicitly requests local or remote → follow the user's choice
- User doesn't specify → check for cluster config:
cat ~/.config/modelopt/clusters.yaml 2>/dev/null || cat .agents/clusters.yaml 2>/dev/null || cat .claude/clusters.yaml 2>/dev/null
If a cluster config exists with content → use the remote cluster (do not fall back to local even if local GPUs are available — the cluster config indicates the user's preferred execution environment). Otherwise → local execution.
If the cluster config contains multiple clusters and the user did not name the target cluster, ask which cluster to use before calling remote_load_cluster. Do not silently fall back to default_cluster in multi-cluster configs; different clusters can have different filesystems, GPU types, auth paths, and SSH setup.
For remote, connect:
source "$SKILL_DIR/remote_exec.sh"
remote_load_cluster <cluster_name>
remote_check_ssh
remote_detect_env # sets REMOTE_ENV_TYPE = slurm / docker / bare
If remote but no config, ask user for: hostname, SSH username, SSH key path, remote workdir. Create ~/.config/modelopt/clusters.yaml (see remote-execution.md for format).
Env-3. What compute is available?
Run on the target machine (local, or via remote_run if remote):
which srun sbatch 2>/dev/null && echo "SLURM"
docker info 2>/dev/null | grep -qi nvidia && echo "Docker+GPU"
nvidia-smi --query-gpu=name,memory.total --format=csv,noheader 2>/dev/null
Also check:
ls tools/launcher/launch.py 2>/dev/null && echo "Launcher available"
No GPU detected?
- If local with no GPU and no cluster config → ask the user: "No local GPU detected. Do you have a remote machine or cluster with GPUs? If so, I'll need connection details (hostname, SSH username, key path, remote workdir) to run there."
- If user provides remote info → create
clusters.yaml, go back to Env-2 - If user has no GPU anywhere → stop: this task requires a CUDA GPU
Summary
After this, you should know:
- ModelOpt source location
- Local or remote (+ cluster config if remote)
- SLURM / Docker+GPU / bare GPU
- Launcher availability
- GPU model and count
Return to the skill's SKILL.md for the execution path based on these results.
Multi-user / Slack bot
If MODELOPT_WORKSPACE_ROOT is set, read workspace-management.md before proceeding.