Files
Model-Optimizer/plugins/modelopt/skills/common/environment-setup.md
T
Chad Voegele d3fe8117ff Add ModelOpt agent plugin marketplace (#2025)
### What does this PR do?

Type of change: new feature

Packages the existing ModelOpt agent skills as installable Codex and
Claude plugins:

- Adds a repo-scoped Codex marketplace and Claude-compatible
marketplace.
- Adds the canonical `plugins/modelopt/` plugin tree and manifests.
- Moves the skill tree into the plugin and keeps `.agents/skills` as a
compatibility symlink.
- Adds a minimal `common` placeholder skill required by Codex
validation.
- Documents installation from this repository.

### Usage

```bash
codex plugin marketplace add NVIDIA/Model-Optimizer
```

Then open `/plugins`, select the `modelopt` marketplace, and install
`modelopt`.

For Claude Code:

```bash
claude plugin marketplace add https://github.com/NVIDIA/Model-Optimizer.git
claude plugin install modelopt@modelopt
```

### Testing

- Codex plugin validator
- `claude plugin validate . --strict`
- `claude plugin validate plugins/modelopt --strict`

1. Install the marketplace plugin with Codex and Claude from an
unrelated temporary workspace.
2. Exercise packaged evaluation helpers, a day-0 gate, and the shared
remote helper from that workspace.
3. Run `uv run --frozen --extra dev python -m pytest -q
plugins/modelopt/skills/day0-release/tests/test_gates.py
plugins/modelopt/skills/benchmark-model-kernels/tests`.
4. Run pre-commit hooks for all changed files.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — added a plugin-path
validator; existing focused skill tests and installed-plugin smoke tests
pass.
- Did you update Changelog?: N/A — agent tooling and distribution only.
- Did you get Claude approval on this PR?: N/A

### Additional Information

Skills remain available through `.agents/skills`; bundled helpers are
packaged under the plugin and resolved from `$SKILL_DIR` so installed
workflows do not depend on the current workspace.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added installable ModelOpt plugins for Claude Code and Codex.
* Added skills for PTQ, deployment, evaluation, monitoring, debugging,
benchmarking, MLflow access, EAGLE3 workflows, and release management.
* Added deployment helpers, evaluation recipes, checkpoint validation,
and release-gating tools.
* **Documentation**
* Expanded setup, credential, SLURM, benchmarking, deployment,
evaluation, troubleshooting, and workspace guidance.
* Added installation instructions and updated agent-skill discovery
guidance.
* **Maintenance**
* Updated skill references and compatibility links for reliable use
across supported plugin environments.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
2026-08-12 11:33:17 -05:00

2.9 KiB

Environment Setup

Common detection for all ModelOpt skills. After this, you know what's available.

Env-1. Get ModelOpt source

ls examples/hf_ptq/hf_ptq.py 2>/dev/null && echo "Source found"

If not found: git clone https://github.com/NVIDIA/Model-Optimizer.git && cd Model-Optimizer

If found, ensure the source is up to date:

git pull origin main

If previous runs left patches in modelopt/ (from 4C unlisted model work), check whether they should be kept. Reset only if starting a completely new task: git checkout main.

Env-2. Local or remote?

  1. User explicitly requests local or remote → follow the user's choice
  2. User doesn't specify → check for cluster config:
cat ~/.config/modelopt/clusters.yaml 2>/dev/null || cat .agents/clusters.yaml 2>/dev/null || cat .claude/clusters.yaml 2>/dev/null

If a cluster config exists with content → use the remote cluster (do not fall back to local even if local GPUs are available — the cluster config indicates the user's preferred execution environment). Otherwise → local execution.

If the cluster config contains multiple clusters and the user did not name the target cluster, ask which cluster to use before calling remote_load_cluster. Do not silently fall back to default_cluster in multi-cluster configs; different clusters can have different filesystems, GPU types, auth paths, and SSH setup.

For remote, connect:

source "$SKILL_DIR/remote_exec.sh"
remote_load_cluster <cluster_name>
remote_check_ssh
remote_detect_env    # sets REMOTE_ENV_TYPE = slurm / docker / bare

If remote but no config, ask user for: hostname, SSH username, SSH key path, remote workdir. Create ~/.config/modelopt/clusters.yaml (see remote-execution.md for format).

Env-3. What compute is available?

Run on the target machine (local, or via remote_run if remote):

which srun sbatch 2>/dev/null && echo "SLURM"
docker info 2>/dev/null | grep -qi nvidia && echo "Docker+GPU"
nvidia-smi --query-gpu=name,memory.total --format=csv,noheader 2>/dev/null

Also check:

ls tools/launcher/launch.py 2>/dev/null && echo "Launcher available"

No GPU detected?

  • If local with no GPU and no cluster config → ask the user: "No local GPU detected. Do you have a remote machine or cluster with GPUs? If so, I'll need connection details (hostname, SSH username, key path, remote workdir) to run there."
  • If user provides remote info → create clusters.yaml, go back to Env-2
  • If user has no GPU anywhere → stop: this task requires a CUDA GPU

Summary

After this, you should know:

  • ModelOpt source location
  • Local or remote (+ cluster config if remote)
  • SLURM / Docker+GPU / bare GPU
  • Launcher availability
  • GPU model and count

Return to the skill's SKILL.md for the execution path based on these results.

Multi-user / Slack bot

If MODELOPT_WORKSPACE_ROOT is set, read workspace-management.md before proceeding.