Add ModelOpt agent plugin marketplace (#2025)

### What does this PR do?

Type of change: new feature

Packages the existing ModelOpt agent skills as installable Codex and
Claude plugins:

- Adds a repo-scoped Codex marketplace and Claude-compatible
marketplace.
- Adds the canonical `plugins/modelopt/` plugin tree and manifests.
- Moves the skill tree into the plugin and keeps `.agents/skills` as a
compatibility symlink.
- Adds a minimal `common` placeholder skill required by Codex
validation.
- Documents installation from this repository.

### Usage

```bash
codex plugin marketplace add NVIDIA/Model-Optimizer
```

Then open `/plugins`, select the `modelopt` marketplace, and install
`modelopt`.

For Claude Code:

```bash
claude plugin marketplace add https://github.com/NVIDIA/Model-Optimizer.git
claude plugin install modelopt@modelopt
```

### Testing

- Codex plugin validator
- `claude plugin validate . --strict`
- `claude plugin validate plugins/modelopt --strict`

1. Install the marketplace plugin with Codex and Claude from an
unrelated temporary workspace.
2. Exercise packaged evaluation helpers, a day-0 gate, and the shared
remote helper from that workspace.
3. Run `uv run --frozen --extra dev python -m pytest -q
plugins/modelopt/skills/day0-release/tests/test_gates.py
plugins/modelopt/skills/benchmark-model-kernels/tests`.
4. Run pre-commit hooks for all changed files.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — added a plugin-path
validator; existing focused skill tests and installed-plugin smoke tests
pass.
- Did you update Changelog?: N/A — agent tooling and distribution only.
- Did you get Claude approval on this PR?: N/A

### Additional Information

Skills remain available through `.agents/skills`; bundled helpers are
packaged under the plugin and resolved from `$SKILL_DIR` so installed
workflows do not depend on the current workspace.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added installable ModelOpt plugins for Claude Code and Codex.
* Added skills for PTQ, deployment, evaluation, monitoring, debugging,
benchmarking, MLflow access, EAGLE3 workflows, and release management.
* Added deployment helpers, evaluation recipes, checkpoint validation,
and release-gating tools.
* **Documentation**
* Expanded setup, credential, SLURM, benchmarking, deployment,
evaluation, troubleshooting, and workspace guidance.
* Added installation instructions and updated agent-skill discovery
guidance.
* **Maintenance**
* Updated skill references and compatibility links for reliable use
across supported plugin environments.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
This commit is contained in:
Chad Voegele
2026-08-12 11:33:17 -05:00
committed by GitHub
parent 1b3fd43015
commit d3fe8117ff
103 changed files with 281 additions and 122 deletions
+20 -22
View File
@@ -1,46 +1,44 @@
# `.agents/` — agent-agnostic source of truth
# `.agents/` — agent compatibility and shared config
This directory is the canonical location for assets shared by AI coding agents
working in this repository (Claude Code, Codex, Cursor, …).
This directory exposes the ModelOpt plugin skills to repository-local agents
and holds shared configuration.
## Layout
```text
.agents/
├── skills/ # SKILL.md files (canonical)
│ └── <skill-name>/SKILL.md
├── skills → ../plugins/modelopt/skills
├── plugins/
│ └── marketplace.json # Codex marketplace
├── scripts/ # shared helper scripts (sync-upstream-skills.sh, …)
└── clusters.yaml.example # remote-cluster config template
plugins/modelopt/
├── .claude-plugin/
├── .codex-plugin/
└── skills/ # canonical SKILL.md files
├── common/ # shared skill support files
└── <skill-name>/SKILL.md
```
## Why this exists
Different agents look for skills/config in vendor-specific directories. Rather
than maintaining N copies that drift out of sync, **`.agents/` is the single
source of truth** — each agent's guidance or install mechanism points here
directly.
## How each agent finds these
Each agent points at `.agents/` through whatever mechanism it supports — never
a copy:
- **Claude Code** only auto-discovers skills under `.claude/skills/`, so
`.claude/` holds relative in-repo symlinks back into `.agents/`:
`.claude/skills → ../.agents/skills`, `.claude/scripts → ../.agents/scripts`,
and `.claude/clusters.yaml.example → ../.agents/clusters.yaml.example`. These
follow the same committed-symlink pattern already used elsewhere in this repo
(e.g. `CLAUDE.md`, `tools/launcher/modules/Model-Optimizer`).
- **Future agents** (Codex, Cursor, …) add their own symlink or config pointing
at `.agents/`.
`.claude/skills/` holds relative symlinks into `.agents/skills/`.
- **Repository agents** use `.agents/skills`, a relative symlink into the
plugin.
- **Claude Code and Codex plugins** load `plugins/modelopt/skills` directly.
## Editing rules
- **Always edit files under `.agents/`**.
- **Always edit skills under `plugins/modelopt/skills/`**.
- Vendored-verbatim skills (`launching-evals`, `accessing-mlflow`) are managed
by `.agents/scripts/sync-upstream-skills.sh` — do not modify by hand.
- New skills go in `.agents/skills/<skill-name>/SKILL.md` following the
conventions of existing skills (e.g. `.agents/skills/monitor/SKILL.md`).
- New skills go in `plugins/modelopt/skills/<skill-name>/SKILL.md`.
- Shared support files go in `plugins/modelopt/skills/common/`.
## Project-level cluster config
+5
View File
@@ -8,6 +8,11 @@ of the always-loaded agent instructions.
Update `AGENTS.md` for repository-wide agent instructions. `CLAUDE.md` is
symlinked to `AGENTS.md`, so changes there apply to both Codex and Claude Code.
## Installable Skills
The `modelopt` plugin packages the repository skills for use from any
workspace. Installation commands are in the [README](../README.md#ai-agents).
## Local Overrides
For private local instructions, use the tool-specific override file:
+20
View File
@@ -0,0 +1,20 @@
{
"name": "modelopt",
"interface": {
"displayName": "NVIDIA Model Optimizer"
},
"plugins": [
{
"name": "modelopt",
"source": {
"source": "local",
"path": "./plugins/modelopt"
},
"policy": {
"installation": "AVAILABLE",
"authentication": "ON_INSTALL"
},
"category": "Productivity"
}
]
}
+3 -2
View File
@@ -26,13 +26,14 @@
#
# Requires: gh, base64, awk. Run from the repo root.
#
# The script overwrites .agents/skills/<skill>/ with upstream contents and
# The script overwrites plugins/modelopt/skills/<skill>/ through the
# .agents/skills compatibility symlink and
# re-applies our provenance lines into each SKILL.md frontmatter. If you have
# local changes to a vendored skill, they will be lost — that is expected,
# since vendored-verbatim skills should not be modified locally.
#
# Note: .claude/skills/ (and other agent-specific skill dirs) are symlinks to
# .agents/skills/ — see .agents/README.md.
# plugins/modelopt/skills/ — see .agents/README.md.
set -euo pipefail
+1
View File
@@ -0,0 +1 @@
../plugins/modelopt/skills
+21
View File
@@ -0,0 +1,21 @@
{
"$schema": "https://json.schemastore.org/claude-code-marketplace.json",
"name": "modelopt",
"version": "0.1.0",
"description": "Model Optimizer agent plugins.",
"owner": {
"name": "NVIDIA Corporation"
},
"plugins": [
{
"name": "modelopt",
"source": "./plugins/modelopt",
"description": "Skills for Model Optimizer development, quantization, deployment, and evaluation.",
"version": "0.1.0",
"author": {
"name": "NVIDIA Corporation"
},
"category": "development"
}
]
}
+1
View File
@@ -0,0 +1 @@
../../.agents/skills/benchmark-model-kernels
+3 -1
View File
@@ -15,6 +15,7 @@ on:
- "tools/mcp/**"
- "tools/resource_monitor.py"
- ".agents/skills/**"
- "plugins/modelopt/skills/**"
schedule:
- cron: "0 0 * * *" # Nightly
workflow_dispatch:
@@ -55,6 +56,7 @@ jobs:
tools/mcp/**
tools/resource_monitor.py
.agents/skills/**
plugins/modelopt/skills/**
linux:
runs-on: ubuntu-latest
timeout-minutes: 15
@@ -190,7 +192,7 @@ jobs:
# Override addopts to drop the repo's coverage/instafail plugins (not installed here).
run: |
pip install pytest
python -m pytest .agents/skills/ -o addopts="" -p no:cacheprovider -v
python -m pytest plugins/modelopt/skills/ -o addopts="" -p no:cacheprovider -v
unit-pr-required-check:
# Run even if some jobs are skipped
if: ${{ github.event_name == 'pull_request' && always() }}
+2 -2
View File
@@ -14,5 +14,5 @@ config:
# Vendored upstream skills — kept byte-identical to upstream via
# .agents/scripts/sync-upstream-skills.sh; do not reformat.
ignores:
- ".agents/skills/launching-evals/**"
- ".agents/skills/accessing-mlflow/**"
- "plugins/modelopt/skills/launching-evals/**"
- "plugins/modelopt/skills/accessing-mlflow/**"
+2 -2
View File
@@ -70,10 +70,10 @@ repos:
exclude: ^modelopt_recipes/configs/
- id: sync-claude-skills
name: sync .claude/skills/ symlinks from .agents/skills/
name: sync .claude/skills/ symlinks from plugin skills
entry: bash tools/precommit/sync_claude_skills.sh
language: system
files: ^\.agents/skills/
files: ^plugins/modelopt/skills/
pass_filenames: false
- id: check-launcher-yaml
+4 -6
View File
@@ -7,12 +7,10 @@ These instructions apply to AI-assisted work in this repository.
- Start with `README.md` for project overview and install.
- Use `modelopt/` for source, `tests/` for focused test coverage, and
`examples/` or `docs/` for usage patterns.
- **Agent skills and shared config live under `.agents/`** — the canonical,
agent-agnostic source of truth (`.agents/skills/<name>/SKILL.md`,
`.agents/scripts/`, `.agents/clusters.yaml.example`). Claude Code's
`.claude/skills`, `.claude/scripts`, and `.claude/clusters.yaml.example` are
relative symlinks into `.agents/`. Always edit files under `.agents/`, not the
symlink path. See `.agents/README.md` for the convention.
- **Agent skills live under `plugins/modelopt/skills/`**, the installable
plugin's canonical skill tree. `.agents/skills` and `.claude/skills` expose
those skills through relative symlinks. Shared agent config and scripts
remain under `.agents/`. See `.agents/README.md` for the convention.
## Coding guidelines
+19 -1
View File
@@ -170,7 +170,25 @@ Please read our [Contributing](./CONTRIBUTING.md) guidelines for details on how
## AI Agents
For AI-assisted development setup, see the [agent tooling notes](./.agents/TOOLING.md).
ModelOpt's agent skills can be installed from this repository and used in any
workspace.
### Claude Code
```bash
claude plugin marketplace add https://github.com/NVIDIA/Model-Optimizer.git
claude plugin install modelopt@modelopt
```
### Codex
```bash
codex plugin marketplace add https://github.com/NVIDIA/Model-Optimizer.git
```
Then open `/plugins`, select the `modelopt` marketplace, and install `modelopt`.
Contributors can also use the skills directly from a checkout. See the
[agent tooling notes](./.agents/TOOLING.md).
### Top Contributors
@@ -0,0 +1,30 @@
{
"$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
"name": "modelopt",
"displayName": "ModelOpt",
"version": "0.1.0",
"description": "Skills for Model Optimizer development, quantization, deployment, and evaluation.",
"author": {
"name": "NVIDIA Corporation"
},
"homepage": "https://github.com/NVIDIA/Model-Optimizer",
"repository": "https://github.com/NVIDIA/Model-Optimizer",
"license": "Apache-2.0",
"keywords": [
"modelopt",
"quantization",
"evaluation",
"deployment",
"llm"
],
"mcpServers": {
"modelopt": {
"command": "uvx",
"args": [
"--from",
"git+https://github.com/NVIDIA/Model-Optimizer.git#subdirectory=tools/mcp",
"modelopt-mcp"
]
}
}
}
@@ -0,0 +1,38 @@
{
"name": "modelopt",
"version": "0.1.0",
"description": "Skills for Model Optimizer development, quantization, deployment, and evaluation.",
"author": {
"name": "NVIDIA Corporation",
"url": "https://github.com/NVIDIA/Model-Optimizer"
},
"homepage": "https://github.com/NVIDIA/Model-Optimizer",
"repository": "https://github.com/NVIDIA/Model-Optimizer",
"license": "Apache-2.0",
"keywords": [
"modelopt",
"quantization",
"evaluation",
"deployment",
"llm"
],
"skills": "./skills/",
"mcpServers": "./.mcp.json",
"interface": {
"displayName": "ModelOpt",
"shortDescription": "Optimize, deploy, and evaluate models.",
"longDescription": "Provides Model Optimizer workflows for post-training quantization, deployment, evaluation, result comparison, and release validation.",
"developerName": "NVIDIA",
"category": "Developer Tools",
"capabilities": [
"Interactive",
"Write"
],
"defaultPrompt": [
"Quantize this model with ModelOpt.",
"Deploy and evaluate this checkpoint.",
"Compare the baseline and quantized results."
],
"brandColor": "#76B900"
}
}
+10
View File
@@ -0,0 +1,10 @@
{
"modelopt": {
"command": "uvx",
"args": [
"--from",
"git+https://github.com/NVIDIA/Model-Optimizer.git#subdirectory=tools/mcp",
"modelopt-mcp"
]
}
}
@@ -41,7 +41,7 @@ order:
GPU needed:
```bash
python .agents/skills/benchmark-model-kernels/scripts/benchmark_model.py <model> \
python "$SKILL_DIR/scripts/benchmark_model.py" <model> \
--tp <tp> --ep <ep> --ms <m1> <m2> ... --print_only
```
@@ -63,7 +63,7 @@ order:
```bash
CUDA_VISIBLE_DEVICES=<gpu-index> \
python .agents/skills/benchmark-model-kernels/scripts/benchmark_model.py <model> \
python "$SKILL_DIR/scripts/benchmark_model.py" <model> \
--tp <tp> --ep <ep> --ms <m1> <m2> ... \
--flashinfer_repo <flashinfer-repo> --workdir <workdir>
```
@@ -112,7 +112,7 @@ missing shape:
```bash
CUDA_VISIBLE_DEVICES=<gpu-index> \
python .agents/skills/benchmark-model-kernels/scripts/benchmark_via_builtin.py \
python "$SKILL_DIR/scripts/benchmark_via_builtin.py" \
--flashinfer_repo <flashinfer-repo> --ms <m1> <m2> ... \
--nks <n>,<k>,<name> --workdir <workdir>
```
+10
View File
@@ -0,0 +1,10 @@
---
name: common
description: Shared ModelOpt support files. Use only when another ModelOpt skill directs you here.
---
# Shared ModelOpt Support
This skill is a placeholder required for Codex plugin validation.
Read only the file named by the calling ModelOpt skill.
@@ -34,13 +34,13 @@ If the cluster config contains multiple clusters and the user did not name the t
For remote, connect:
```bash
source .agents/skills/common/remote_exec.sh
source "$SKILL_DIR/remote_exec.sh"
remote_load_cluster <cluster_name>
remote_check_ssh
remote_detect_env # sets REMOTE_ENV_TYPE = slurm / docker / bare
```
If remote but no config, ask user for: hostname, SSH username, SSH key path, remote workdir. Create `~/.config/modelopt/clusters.yaml` (see `skills/common/remote-execution.md` for format).
If remote but no config, ask user for: hostname, SSH username, SSH key path, remote workdir. Create `~/.config/modelopt/clusters.yaml` (see `remote-execution.md` for format).
## Env-3. What compute is available?
@@ -79,4 +79,4 @@ Return to the skill's SKILL.md for the execution path based on these results.
## Multi-user / Slack bot
If `MODELOPT_WORKSPACE_ROOT` is set, read `skills/common/workspace-management.md` before proceeding.
If `MODELOPT_WORKSPACE_ROOT` is set, read `workspace-management.md` before proceeding.
@@ -46,7 +46,7 @@ See `.agents/clusters.yaml.example` for a fully annotated example with multiple
## 2. Connect and Establish Persistent Session
```bash
source .agents/skills/common/remote_exec.sh
source "$SKILL_DIR/remote_exec.sh"
remote_load_cluster <cluster_name> # or omit name to use default_cluster
remote_check_ssh # validates connectivity + starts persistent session
```
@@ -153,6 +153,6 @@ remote_sync_from <remote_output_subdir> /local/output/
## Reference Files
- **`skills/common/remote_exec.sh`** — Full utility library (session, run, sync, SLURM, Docker helpers)
- **`remote_exec.sh`** — Full utility library (session, run, sync, SLURM, Docker helpers)
- **`.agents/clusters.yaml`** — Active cluster configuration (canonical; `.claude/clusters.yaml` also accepted for back-compat)
- **`.agents/clusters.yaml.example`** — Annotated example config
@@ -17,7 +17,7 @@
# remote_exec.sh — Remote execution utility for ModelOpt agent skills
#
# Usage:
# source .agents/skills/common/remote_exec.sh
# source "$SKILL_DIR/remote_exec.sh"
# remote_load_cluster <cluster_name> # or: remote_load_cluster (uses default)
# remote_check_ssh
# remote_detect_env # detect SLURM vs Docker vs bare metal
@@ -53,7 +53,7 @@ srun \
### Container registry credentials (pyxis)
If `srun --container-image` uses an image from a private registry (e.g., `nvcr.io/nvidia/...`), pyxis/enroot needs registry credentials on the cluster in `~/.config/enroot/.credentials`. See `skills/common/credentials.md` for the NGC / Docker / HF token setup. Without this, `srun` fails with `401 Unauthorized` when the compute node pulls.
If `srun --container-image` uses an image from a private registry (e.g., `nvcr.io/nvidia/...`), pyxis/enroot needs registry credentials on the cluster in `~/.config/enroot/.credentials`. See `credentials.md` for the NGC / Docker / HF token setup. Without this, `srun` fails with `401 Unauthorized` when the compute node pulls.
Submit and capture the job ID:
@@ -29,10 +29,9 @@ change is being measured, typically a further quantized version of the baseline.
before comparing scores. If not, validate logs, server health,
judge/code-execution status, sample accounting, and reasoning parsing before
computing deltas.
5. For each task, use the canonical score field from the matching
`.agents/skills/evaluation/recipes/tasks/<task>.md` Score Extraction
section.
6. Read and perform `.agents/skills/evaluation/references/run-validation.md`
5. For each task, use the canonical score field from the matching evaluation
skill task recipe, `recipes/tasks/<task>.md`, under **Score Extraction**.
6. Use the evaluation skill's `references/run-validation.md` to perform the
**External Baseline Sanity Check**. Record each source URL, protocol
difference, and task status before applying the candidate-delta gate. A
failed baseline blocks a success verdict; correct and rerun it first. If no
@@ -80,8 +79,9 @@ If any item differs, either rerun with matched settings or label the result as
not an apples-to-apples quantization comparison.
These checks compare the baseline and candidate to each other. The external
baseline check in `evaluation/references/run-validation.md` separately tests
whether the baseline's absolute score is credible; both guards must be reported.
baseline check in the evaluation skill's `references/run-validation.md`
separately tests whether the baseline's absolute score is credible; both guards
must be reported.
## Report Format
@@ -27,8 +27,10 @@ Resolve these before starting (ask the user for anything missing):
- **Model** — HF handle or checkpoint path.
- **Recipe / qformat** — e.g. `nvfp4`, `fp8`, or a recipe path. One candidate for v1.
- **Cluster / launcher** — from `clusters.yaml` (see `skills/common/environment-setup.md`).
- **Eval set** — defaults to the AA suite (`evaluation/recipes/tasks/aa/`).
- **Cluster / launcher** — from `clusters.yaml` (see the common skill's
`environment-setup.md`).
- **Eval set** — defaults to the evaluation skill's AA suite
(`recipes/tasks/aa/`).
- **Threshold** — max accuracy drop; default `0.01` (1%).
## The chain
@@ -60,8 +62,8 @@ progress:
### Step 1 — Setup gate
Confirm credentials (`skills/common/credentials.md`) and cluster reachability
(`skills/common/remote-execution.md`). If either fails, stop with
Use the common skill's `credentials.md` and `remote-execution.md` to confirm
credentials and cluster reachability. If either fails, stop with
`SYSTEMIC` — do not start PTQ.
### Step 2 — PTQ
@@ -70,9 +72,9 @@ Invoke the **ptq** skill to produce the quantized checkpoint. Then gate:
```bash
# The ptq skill's post-PTQ validation produces a validation-summary JSON (size
# ratio + layer-precision counts + metadata diffs; see
# ptq/references/checkpoint-validation.md). v1 gates on that summary:
python .agents/skills/day0-release/scripts/gate_ptq.py --summary <validation-summary.json>
# ratio + layer-precision counts + metadata diffs; see the ptq skill's
# references/checkpoint-validation.md). v1 gates on that summary:
python "$SKILL_DIR/scripts/gate_ptq.py" --summary <validation-summary.json>
# add `--recipe <qformat>` to override the recipe recorded in the summary
```
@@ -100,7 +102,7 @@ the working command back into NEL's `deployment.command` and resume the eval. If
the checkpoint genuinely can't serve, `POINT_INFEASIBLE`. Gate:
```bash
python .agents/skills/day0-release/scripts/gate_run.py --run <run-summary.json>
python "$SKILL_DIR/scripts/gate_run.py" --run <run-summary.json>
```
A `pass: false` here means the run is incomplete or invalid (judge/parse error,
@@ -118,7 +120,7 @@ baseline.
After recording the external status, produce per-task deltas and run:
```bash
python .agents/skills/day0-release/scripts/gate_compare.py \
python "$SKILL_DIR/scripts/gate_compare.py" \
--baseline <baseline_scores.json> --candidate <candidate_scores.json> \
--threshold 0.01
```
@@ -18,7 +18,7 @@
These are deterministic — no GPU, cluster, or network. They test the pure
decision functions that the gates rest on. Run with:
python -m pytest .agents/skills/day0-release/tests/test_gates.py
python -m pytest "$SKILL_DIR/tests/test_gates.py"
"""
import sys
@@ -10,26 +10,26 @@ Serve a model checkpoint as an OpenAI-compatible inference endpoint. Supports vL
## Quick Start
Prefer `scripts/deploy.sh` for standard local deployments — it handles quant detection, health checks, and server lifecycle. Use the raw framework commands in Step 4 when you need flags the script doesn't support, or for remote deployment.
Prefer `$SKILL_DIR/scripts/deploy.sh` for standard local deployments — it handles quant detection, health checks, and server lifecycle. Use the raw framework commands in Step 4 when you need flags the script doesn't support, or for remote deployment.
```bash
# Start vLLM server with a ModelOpt checkpoint
scripts/deploy.sh start --model ./qwen3-0.6b-fp8
"$SKILL_DIR/scripts/deploy.sh" start --model ./qwen3-0.6b-fp8
# Start with SGLang and tensor parallelism
scripts/deploy.sh start --model ./llama-70b-nvfp4 --framework sglang --tp 4
"$SKILL_DIR/scripts/deploy.sh" start --model ./llama-70b-nvfp4 --framework sglang --tp 4
# Start from HuggingFace hub
scripts/deploy.sh start --model nvidia/Llama-3.1-8B-Instruct-FP8
"$SKILL_DIR/scripts/deploy.sh" start --model nvidia/Llama-3.1-8B-Instruct-FP8
# Test the API
scripts/deploy.sh test
"$SKILL_DIR/scripts/deploy.sh" test
# Check status
scripts/deploy.sh status
"$SKILL_DIR/scripts/deploy.sh" status
# Stop
scripts/deploy.sh stop
"$SKILL_DIR/scripts/deploy.sh" stop
```
The script handles: GPU detection, quantization flag auto-detection (FP8 vs FP4), server lifecycle (start/stop/restart/status), health check polling, and API testing.
@@ -38,7 +38,7 @@ The script handles: GPU detection, quantization flag auto-detection (FP8 vs FP4)
### 0. Check workspace (multi-user / Slack bot)
If `MODELOPT_WORKSPACE_ROOT` is set, read `skills/common/workspace-management.md`. Before creating a new workspace, check the current session for existing model workspaces — especially if deploying a checkpoint from a prior PTQ run:
If `MODELOPT_WORKSPACE_ROOT` is set, use the common skill's `workspace-management.md`. Before creating a new workspace, check the current session for existing model workspaces — especially if deploying a checkpoint from a prior PTQ run:
```bash
ls "$MODELOPT_WORKSPACE_ROOT/<session_id>/" 2>/dev/null
@@ -79,7 +79,7 @@ Check the support matrix in `references/support-matrix.md` to confirm the model
### 3. Check the environment
Read `skills/common/environment-setup.md` for GPU detection, local vs remote, and SLURM/Docker/bare metal detection. After completing it you should know: GPU model/count, local or remote, and execution environment.
Use the common skill's `environment-setup.md` for GPU detection, local vs remote, and SLURM/Docker/bare metal detection. After completing it you should know: GPU model/count, local or remote, and execution environment.
Then check the **deployment framework** is installed:
@@ -211,12 +211,13 @@ token shapes, and how to read `profile_export_aiperf.json`.
If a cluster config exists (`~/.config/modelopt/clusters.yaml`, `.agents/clusters.yaml`, or `.claude/clusters.yaml`), or the user mentions running on a remote machine:
0. **Check container registry auth** — before submitting any SLURM job with a container image, verify credentials exist on the cluster per `skills/common/slurm-setup.md` section 6. If credentials are missing for the image's registry, ask the user to fix auth or switch to an image on an authenticated registry (e.g., NGC). **Do not submit until auth is confirmed.**
0. **Check container registry auth** — before submitting any SLURM job with a container image, verify credentials exist on the cluster per the common skill's `slurm-setup.md` section 6. If credentials are missing for the image's registry, ask the user to fix auth or switch to an image on an authenticated registry (e.g., NGC). **Do not submit until auth is confirmed.**
1. **Source remote utilities:**
1. **Source remote utilities:** Load the common skill, then resolve
`remote_exec.sh` from that skill's root.
```bash
source .agents/skills/common/remote_exec.sh
source "<common-skill-dir>/remote_exec.sh"
remote_load_cluster
remote_check_ssh
remote_detect_env
@@ -232,7 +233,7 @@ If a cluster config exists (`~/.config/modelopt/clusters.yaml`, `.agents/cluster
3. **Deploy based on remote environment:**
- **SLURM** — see `skills/common/slurm-setup.md` for job script templates (container setup, account/partition discovery). The server command inside the container is the same as Step 4 (e.g., `python -m vllm.entrypoints.openai.api_server --model <path> --quantization modelopt`). After submitting, register the job and set up monitoring per the **monitor skill**. Get the node hostname from `squeue -j $JOBID -o %N`.
- **SLURM** — see the common skill's `slurm-setup.md` for job script templates (container setup, account/partition discovery). The server command inside the container is the same as Step 4 (e.g., `python -m vllm.entrypoints.openai.api_server --model <path> --quantization modelopt`). After submitting, register the job and set up monitoring per the **monitor skill**. Get the node hostname from `squeue -j $JOBID -o %N`.
- **Bare metal / Docker** — use `remote_run` to start the server directly:
@@ -27,7 +27,7 @@
"files": [],
"expected_behavior": [
"Checks for cluster config at ~/.config/modelopt/clusters.yaml, .agents/clusters.yaml, or .claude/clusters.yaml",
"Sources .agents/skills/common/remote_exec.sh",
"Sources remote_exec.sh from the loaded common skill root",
"Calls remote_load_cluster, remote_check_ssh, remote_detect_env",
"Checks if checkpoint is already on remote (e.g., from prior PTQ run) before syncing; only syncs if local",
"For SLURM: writes a job script with srun --container-image and --container-mounts on srun line (not #SBATCH)",
@@ -12,7 +12,7 @@ Guide the user through creating NEL YAML configs, running evaluations, and monit
### Workspace integration
If `MODELOPT_WORKSPACE_ROOT` is set, read `skills/common/workspace-management.md` and reuse existing workspaces (this skill is usually the final stage of PTQ → Deploy → Eval; carry any deployment-time patches into `deployment.command`).
If `MODELOPT_WORKSPACE_ROOT` is set, use the common skill's `workspace-management.md` and reuse existing workspaces (this skill is usually the final stage of PTQ → Deploy → Eval; carry any deployment-time patches into `deployment.command`).
### Workflow
@@ -41,7 +41,7 @@ overrides, and `services`/`benchmarks`/`cluster`/`output` schema. If the user as
for one, do **not** add it to a 0.2.6 `evaluation.tasks` list — instead:
1. Read **`references/nel-next.md`** (shared: venv, schema, AWS creds, architecture, timeout strategy, MLflow, run flow) + the per-benchmark recipe `recipes/tasks/aa_next/{terminal_bench_2_1,swebench_verified}.md`; start from `recipes/examples/example_eval_next.yaml`.
2. Isolated nel-next venv: `.agents/scripts/nel-next.sh --setup-only` (keeps 0.2.6 `nel` untouched).
2. Isolated nel-next venv: `"$SKILL_DIR/scripts/nel-next.sh" --setup-only` (keeps 0.2.6 `nel` untouched).
3. Run **`modelopttools:eval-config`** (Step 3b) to write the AWS-sandbox creds + harbor infra rows (`${NEL_NEXT_EVAL_IMAGE}`, `${HARBOR_*_ECR_REPOSITORY}`) into `.env`; always include the `output.export_config.mlflow` block.
4. Dry-run → canary → full (`nel-next.sh eval run`), then **push to MLflow** — SLURM doesn't auto-export, so run `nel-next.sh mlflow-push -r <run_id> -c <cfg>` after (config-driven; see `references/nel-next.md`).
@@ -63,7 +63,7 @@ GDPVal:
self-contained file.
3. Prerequisite — the Apptainer SIF. **If your site provides one, use it**
(NVIDIA-internal: `modelopttools:eval-config` Step 3c); otherwise set
`GDPVAL_SIF_DIR` in `.env` and build with `.agents/scripts/gdpval-sif.sh`
`GDPVAL_SIF_DIR` in `.env` and build with `"$SKILL_DIR/scripts/gdpval-sif.sh"`
(build-if-absent, no cross-cluster copy). Either way the mounted dir must contain
the file `GDPVAL_CONTAINER_PATH` names (template: `python-3.13.gdpval.sif`) — a
name mismatch passes NEL's `test -d` check and the agent then silently runs
@@ -81,7 +81,7 @@ GDPVal:
Run `nel --version`; if missing, instruct `pip install nemo-evaluator-launcher`. If user has an existing config, skip to Step 8 (optionally review for `???` and quantization flags first).
**Set up `.env` now (not Step 8).** The working `.env` lives at the **workspace root** — the directory you run `nel` from — matching `modelopttools:eval-config`'s convention; do **not** create it under the skill dir. (NEL does not discover `.env` by path: it reads secrets from the shell env via the `host:` prefix after you `source`, so the location is purely *which file you source* before `nel run`. Keeping the single `.env` at the workspace root avoids a stale duplicate under the symlinked, shared `.agents/` skill tree.) For judge-scored / user-sim tasks (HLE, AA-LCR, Tau2), seed it from the template if absent — the template ships under the skill dir, the working `.env` does not: `[ -f .env ] || cp .agents/skills/evaluation/recipes/env.example .env`. Then try `modelopttools:eval-config` (if available) to fill the judge `model_id`/`url` rows (user adds the secret key). Needed before Step 5, which substitutes those values into task `<VAR>` placeholders.
**Set up `.env` now (not Step 8).** The working `.env` lives at the **workspace root** — the directory you run `nel` from — matching `modelopttools:eval-config`'s convention; do **not** create it under the skill dir. (NEL does not discover `.env` by path: it reads secrets from the shell env via the `host:` prefix after you `source`, so the location is purely *which file you source* before `nel run`. Keeping the single `.env` at the workspace root avoids a stale duplicate under the symlinked, shared `.agents/` skill tree.) For judge-scored / user-sim tasks (HLE, AA-LCR, Tau2), seed it from the template if absent — the template ships under the skill dir, the working `.env` does not: `[ -f .env ] || cp "$SKILL_DIR/recipes/env.example" .env`. Then try `modelopttools:eval-config` (if available) to fill the judge `model_id`/`url` rows (user adds the secret key). Needed before Step 5, which substitutes those values into task `<VAR>` placeholders.
**Secret safety — never open `.env` with Read/Write/Edit.** The harness mirrors later edits of any agent-opened file into the transcript, so touching `.env` leaks the keys the user adds afterward. Use shell only (`cp` to create, `source` to load — neither echoes); edit `env.example`, never `.env`; leave value entry to the user / `modelopttools:eval-config`.
@@ -378,7 +378,7 @@ Public images → submit without preflight. Private/restricted → check credent
ssh <host> "grep -E '^\s*machine\s+' ~/.config/enroot/.credentials 2>/dev/null"
```
Add credentials per `skills/common/slurm-setup.md` §6 if missing. If you can't add, switch to a compatible public image (e.g. `nvcr.io/nvidia/vllm:<YY.MM>-py3` — check catalog.ngc.nvidia.com). **Do not retry more than once** after an auth failure.
Add credentials per the common skill's `slurm-setup.md` §6 if missing. If you can't add, switch to a compatible public image (e.g. `nvcr.io/nvidia/vllm:<YY.MM>-py3` — check catalog.ngc.nvidia.com). **Do not retry more than once** after an auth failure.
---
@@ -390,7 +390,7 @@ Run directly when the user asked to launch; otherwise ask before submitting.
```bash
# .env lives at the workspace root (where you run nel); the template ships under the skill dir
[ -f .env ] || cp .agents/skills/evaluation/recipes/env.example .env # create only if Step 1 didn't
[ -f .env ] || cp "$SKILL_DIR/recipes/env.example" .env # create only if Step 1 didn't
set -a && source .env && set +a
# If pre_cmd/post_cmd in config (review pre_cmd first — runs arbitrary commands):
@@ -2,7 +2,7 @@
#
# Copy this file to your workspace root (the dir you run `nel` from) — NOT into
# the skill dir — and fill in the keys you need:
# cp .agents/skills/evaluation/recipes/env.example .env
# cp "$SKILL_DIR/recipes/env.example" .env
# # Edit .env with your keys
# set -a && source .env && set +a
#
@@ -45,7 +45,7 @@ NEMO_EVALUATOR_TRUST_PRE_CMD=1
# TAVILY_API_KEY=
# GDPVal (nemo_gym) — persistent Apptainer SIF cache dir on the TARGET cluster's
# shared FS (a path, not a secret). .agents/scripts/gdpval-sif.sh builds the SIF
# shared FS (a path, not a secret). $SKILL_DIR/scripts/gdpval-sif.sh builds the SIF
# here if absent and reuses it otherwise; the config bind-mounts this dir at
# /gdpval/sif. Convention: a per-user .cache dir.
# GDPVAL_SIF_DIR=<shared-fs>/<user>/.cache/gdpval/sif
@@ -3,9 +3,9 @@
# Benchmark shown = Terminal-Bench 2.1; swap the `benchmarks:` block per the recipe.
#
# Run via the isolated nel-next venv:
# .agents/scripts/nel-next.sh --setup-only
# "$SKILL_DIR/scripts/nel-next.sh" --setup-only
# set -a && source .env && set +a # HF_TOKEN, AWS_*, NEL_NEXT_EVAL_IMAGE, HARBOR_*_ECR_REPOSITORY (from modelopttools:eval-config)
# .agents/scripts/nel-next.sh eval run recipes/examples/example_eval_next.yaml --dry-run
# "$SKILL_DIR/scripts/nel-next.sh" eval run "$SKILL_DIR/recipes/examples/example_eval_next.yaml" --dry-run
# ... --submit -O benchmarks.0.max_problems=2 -O benchmarks.0.repeats=1 -O benchmarks.0.max_concurrent=2 # canary
# ... --submit # full
# Internal harbor infra (eval_image + ECR) comes from .env via ${VAR}; all blocks
@@ -23,7 +23,7 @@
# Before running: `.env` needs HF_TOKEN, INFERENCE_API_KEY, TAVILY_API_KEY,
# INFERENCE_JUDGE_URL, GDPVAL_SIF_DIR, NEMO_EVALUATOR_TRUST_PRE_CMD=1; and the
# SIF must exist (prefer a site-provided one, else
# `srun -p cpu -t 01:00:00 --pty .agents/scripts/gdpval-sif.sh`).
# `srun -p cpu -t 01:00:00 --pty "$SKILL_DIR/scripts/gdpval-sif.sh"`).
#
# nel run --config example_gym_gdpval.yaml --env-file .env
#
@@ -20,7 +20,7 @@ Steps 1–9 apply — but with the branch differences below.
- **Standalone** — one gym eval per config. Never add GDPVal to a multi-task
`evaluation.tasks` list, and never add other tasks to a GDPVal config.
- **Apptainer SIF sandbox** — prefer a site-provided SIF; otherwise
`.agents/scripts/gdpval-sif.sh` builds one into `$GDPVAL_SIF_DIR` (build-if-absent,
`$SKILL_DIR/scripts/gdpval-sif.sh` builds one into `$GDPVAL_SIF_DIR` (build-if-absent,
never copied between clusters). Missing/misnamed → **silent** unsandboxed exec.
- **Thinking mode is mandatory** — non-thinking loses ~86% of pairwise judgements.
Serve with the model's `--reasoning-parser` and force it on via the adapter's
@@ -34,7 +34,7 @@ Step 3c has the path.) Otherwise build it on the target cluster — never copy a
between clusters:
```bash
srun -p cpu -t 01:00:00 --pty .agents/scripts/gdpval-sif.sh # uses $GDPVAL_SIF_DIR
srun -p cpu -t 01:00:00 --pty "$SKILL_DIR/scripts/gdpval-sif.sh" # uses $GDPVAL_SIF_DIR
```
`gdpval-sif.sh` is idempotent (flock-guarded, atomic): it builds from `gdpval.def` at
@@ -160,7 +160,7 @@ mount source, and `raise ValueError` listing the missing ones **before** any
validation, and the run then silently degrades. Guard with the verify-only mode:
```bash
.agents/scripts/gdpval-sif.sh --check # uses $GDPVAL_SIF_DIR; exit 1 + lists what IS there
"$SKILL_DIR/scripts/gdpval-sif.sh" --check # uses $GDPVAL_SIF_DIR; exit 1 + lists what IS there
```
Keep `GDPVAL_SIF_NAME` / the helper's default in sync with the config's
@@ -12,7 +12,7 @@ Start configs from `recipes/examples/example_eval_next.yaml`.
| | default (SKILL Steps 1–9) | nel-next |
|---|---|---|
| package | `nemo-evaluator-launcher` 0.2.6 | `nemo-evaluator[harbor]` 0.4.x |
| env | the skill's normal env | **separate venv** (`.agents/scripts/nel-next.sh`) |
| env | the skill's normal env | **separate venv** (`$SKILL_DIR/scripts/nel-next.sh`) |
| CLI | `nel run --config X.yaml` | `nel eval run X.yaml [--submit]` |
| overrides | `-o ++a.b.c=v` | `-O a.b.c=v` |
| canary limiter | `++…limit_samples=N` | `-O benchmarks.0.max_problems=N` (NOT `--max-problems`, which is `--bench`-only) |
@@ -23,8 +23,8 @@ Start configs from `recipes/examples/example_eval_next.yaml`.
Installing 0.4.x into the 0.2.6 env clobbers `nel`, so it lives in its own venv:
```bash
.agents/scripts/nel-next.sh --setup-only # one-time, ~1-2 min (needs `uv`)
.agents/scripts/nel-next.sh eval run <cfg> --dry-run | --submit | …
"$SKILL_DIR/scripts/nel-next.sh" --setup-only # one-time, ~1-2 min (needs `uv`)
"$SKILL_DIR/scripts/nel-next.sh" eval run <cfg> --dry-run | --submit | …
```
Default install is a git build from `github.com/NVIDIA-NeMo/Evaluator` via `NEL_NEXT_ORIGIN`
@@ -197,12 +197,12 @@ with its own `run_id`, copying the shared `services:` block.
## Run (dry-run → canary → full) → push to MLflow
```bash
set -a && source .env && set +a; NEL=.agents/scripts/nel-next.sh
$NEL eval run <cfg>.yaml --dry-run # validate/render (no SSH)
$NEL eval run <cfg>.yaml --submit -O benchmarks.0.max_problems=2 -O benchmarks.0.repeats=1 -O benchmarks.0.max_concurrent=2 # canary
$NEL eval run <cfg>.yaml --submit # full
$NEL eval {status|logs -f|report -f markdown|merge} -r <run_id> # lifecycle
$NEL mlflow-push -r <run_id> -c <cfg>.yaml # post-run: push merged bundle(s) to MLflow
set -a && source .env && set +a; NEL="$SKILL_DIR/scripts/nel-next.sh"
"$NEL" eval run <cfg>.yaml --dry-run # validate/render (no SSH)
"$NEL" eval run <cfg>.yaml --submit -O benchmarks.0.max_problems=2 -O benchmarks.0.repeats=1 -O benchmarks.0.max_concurrent=2 # canary
"$NEL" eval run <cfg>.yaml --submit # full
"$NEL" eval {status|logs -f|report -f markdown|merge} -r <run_id> # lifecycle
"$NEL" mlflow-push -r <run_id> -c <cfg>.yaml # post-run: push merged bundle(s) to MLflow
```
`eval run` on a slurm cluster scp's the sbatch + redacted `.secrets.env` and
@@ -22,7 +22,7 @@
# run reuses the built SIF instantly.
#
# Usage:
# .agents/scripts/gdpval-sif.sh [<sif-dir-or-file>] [--commit <sha>] [--force|--check]
# "$SKILL_DIR/scripts/gdpval-sif.sh" [<sif-dir-or-file>] [--commit <sha>] [--force|--check]
# <sif-dir-or-file> Persistent path on the target cluster's shared FS.
# DEFAULTS to $GDPVAL_SIF_DIR (from .env) when omitted. A
# directory -> <dir>/$GDPVAL_SIF_NAME (default python-3.13.gdpval.sif,
@@ -45,7 +45,7 @@
# build support, plus network egress to GitHub/base image. Run on a node that has
# it — a login node, or (preferred for the ~30-min build) the CPU partition:
# srun -p cpu -t 01:00:00 --pty \
# .agents/scripts/gdpval-sif.sh /lustre/<...>/gdpval/sif
# "$SKILL_DIR/scripts/gdpval-sif.sh" /lustre/<...>/gdpval/sif
#
# Env overrides: GDPVAL_GYM_COMMIT, GDPVAL_SIF_NAME, APPTAINER_BIN.
set -euo pipefail
@@ -24,10 +24,10 @@
# `uvx` environment (uv resolves + caches + reuses it) and forwards to its `nel`.
#
# Usage (source .env FIRST so the config's ${VAR}s resolve; this never reads secrets):
# .agents/scripts/nel-next.sh --setup-only|--which|--version
# .agents/scripts/nel-next.sh eval run <config.yaml> [--dry-run|--submit] [-O k=v ...]
# .agents/scripts/nel-next.sh eval {status|logs|report|merge|resume|stop} -r <run_id>
# .agents/scripts/nel-next.sh mlflow-push -r <run_id> -c <config.yaml> [-- -o k=v ...]
# "$SKILL_DIR/scripts/nel-next.sh" --setup-only|--which|--version
# "$SKILL_DIR/scripts/nel-next.sh" eval run <config.yaml> [--dry-run|--submit] [-O k=v ...]
# "$SKILL_DIR/scripts/nel-next.sh" eval {status|logs|report|merge|resume|stop} -r <run_id>
# "$SKILL_DIR/scripts/nel-next.sh" mlflow-push -r <run_id> -c <config.yaml> [-- -o k=v ...]
# Post-run: SLURM doesn't auto-export. Pulls the merged bundle(s) + pushes to MLflow
# using the config's export_config.mlflow (resolves ${MLFLOW_TRACKING_URI}, forces
# emit_traces=false to avoid the per-sample hang). Run after `source .env`.
@@ -18,7 +18,8 @@ optimization. Use this skill for each selected recipe's PTQ run.
## Step 1 — Environment
Read `skills/common/environment-setup.md` and `skills/common/workspace-management.md`. After completing them you should know:
Use the common skill's `environment-setup.md` and `workspace-management.md`.
After completing them you should know:
- ModelOpt source is available
- Local or remote (+ cluster config if remote)
@@ -128,7 +129,8 @@ python examples/hf_ptq/hf_ptq.py \
Run `--help` for all options.
For remote: use `remote_run` from `remote_exec.sh` (see `skills/common/remote-execution.md`).
For remote, use `remote_run` from the common skill's `remote_exec.sh`; see its
`remote-execution.md`.
### 4B — Launcher: supported model on SLURM or local Docker
@@ -148,7 +150,8 @@ The launcher blocks and tails logs until the job completes. If the launcher fail
Follow `references/unsupported-models.md`. It walks through investigating the model, patching ModelOpt if needed, and running `hf_ptq.py`. Run manually (like 4A) for easier monitoring and debugging.
For SLURM, see `skills/common/slurm-setup.md` and `references/slurm-setup-ptq.md`.
For SLURM, see the common skill's `slurm-setup.md` and this skill's
`references/slurm-setup-ptq.md`.
### Monitoring
@@ -190,20 +193,20 @@ Report the gate result before moving on. Follow the canonical report format and
- **Model-specific dependencies**: Models with `trust_remote_code` may import packages not in the container (e.g., `mamba-ssm` for hybrid Mamba models). See Step 2.5. Use `EXTRA_PIP_DEPS` env var with the launcher, or install manually before running `hf_ptq.py`
- **Transformers version**: New models may need a newer version of transformers than what's installed. Check `config.json` for `transformers_version`. In containers, beware of `PIP_CONSTRAINT` blocking upgrades — see `references/slurm-setup-ptq.md` for workarounds
- **Gated datasets**: Some calibration datasets require HF authentication. Set `HF_TOKEN` in the job environment. Use `--dataset cnn_dailymail` only as the constrained-environment fallback described in Step 4, not as the preferred calibration set
- **NFS root_squash + Docker**: See `skills/common/slurm-setup.md` section 5
- **NFS root_squash + Docker**: See the common skill's `slurm-setup.md` section 5
## References
| Reference | When to read |
| --- | --- |
| `skills/common/environment-setup.md` | Step 1: always |
| `skills/common/workspace-management.md` | Step 1: always |
| common skill: `environment-setup.md` | Step 1: always |
| common skill: `workspace-management.md` | Step 1: always |
| `references/launcher-guide.md` | Step 4B only (launcher path) |
| `tools/launcher/CLAUDE.md` | Step 4B only, if you need more launcher detail |
| `references/unsupported-models.md` | Step 4C only (unlisted model) |
| `references/checkpoint-validation.md` | Step 5: mandatory post-PTQ gate before deployment/evaluation |
| `skills/common/remote-execution.md` | Step 4A/4C only, if target is remote |
| `skills/common/slurm-setup.md` | Step 4A/4C only, if using SLURM manually (not launcher) |
| common skill: `remote-execution.md` | Step 4A/4C only, if target is remote |
| common skill: `slurm-setup.md` | Step 4A/4C only, if using SLURM manually (not launcher) |
| `references/slurm-setup-ptq.md` | Step 4A/4C only, PTQ-specific SLURM (container, GPU sizing, FSDP2) |
| `examples/hf_ptq/README.md` | Step 3: support matrix, CLI flags, accuracy |
| `modelopt/torch/quantization/config.py` | Step 3: format definitions |
@@ -1,7 +1,7 @@
# SLURM Setup for PTQ
PTQ-specific SLURM details. For generic SLURM patterns (account discovery, job template,
monitoring), see `skills/common/slurm-setup.md`.
monitoring), see the common skill's `slurm-setup.md`.
---
@@ -73,7 +73,7 @@ FSDP2* section of `examples/hf_ptq/README.md`.
Sizing guidance specific to this path: when the per-rank decoder shard approaches GPU capacity (200B+ at low rank count), either add more nodes (more ranks → smaller shard per rank) or add `--cpu_offload`. Layer detection is automatic; no YAML config needed.
Use the multi-node template from `skills/common/slurm-setup.md` section 4 as the job script wrapper.
Use the multi-node template from the common skill's `slurm-setup.md` section 4 as the job script wrapper.
---
@@ -82,6 +82,6 @@ Use the multi-node template from `skills/common/slurm-setup.md` section 4 as the
Before the full calibration run, submit a smoke test with `--calib_size 4` and `--time=00:30:00`.
This catches script errors cheaply before using GPU quota on a real run.
See `skills/common/slurm-setup.md` section 2 for the smoke test partition pattern.
See the common skill's `slurm-setup.md` section 2 for the smoke test partition pattern.
Only submit the full calibration job after the smoke test exits cleanly.
@@ -6,7 +6,7 @@ Follow the investigation steps below to determine if `hf_ptq.py` works or if pat
## Step A — Download the model and locate the source
**Download first.** Follow `skills/common/workspace-management.md` to set up local and remote workspaces, sync ModelOpt source, and download the model on the target machine. This avoids downloading twice and gives access to README, custom modeling code, and tokenizer config.
**Download first.** Follow the common skill's `workspace-management.md` to set up local and remote workspaces, sync ModelOpt source, and download the model on the target machine. This avoids downloading twice and gives access to README, custom modeling code, and tokenizer config.
After download, inspect the model files on the target machine (use `remote_run` if remote):
@@ -20,8 +20,8 @@ Before constructing commands, read:
- `examples/megatron_bridge/README.md`, especially PTQ, data preparation, QAD,
export, and Slurm usage
- `examples/megatron_bridge/{quantize.py,distill.py}` via `--help`
- `skills/common/{environment-setup,workspace-management,slurm-setup}.md`; also
`skills/common/remote-execution.md` for remote Slurm
- the common skill's `environment-setup.md`, `workspace-management.md`, and
`slurm-setup.md`; also its `remote-execution.md` for remote Slurm
Treat the example README and `--help` output as authoritative for mutable flags,
commands, containers, and checkpoint formats. This skill supports Slurm only.

Some files were not shown because too many files have changed in this diff Show More