mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
[SKILL.md Chore] Make .agents/ the canonical agent-skills location (#1362)
### What does this PR do?
Type of change: documentation / repo housekeeping
Centralize agent-shared assets under **`.agents/`** as the single,
agent-agnostic source of truth, so the same `SKILL.md` (plus shared
scripts and cluster config) works across every coding agent without
maintaining N copies that drift out of sync. Claude Code discovers
skills only under `.claude/skills/`, so `.claude/` holds **relative
in-repo symlinks** into `.agents/` for back-compat.
```bash
repo-root/
├── .agents/ ← canonical source of truth
│ ├── README.md
│ ├── clusters.yaml.example
│ ├── scripts/
│ │ └── sync-upstream-skills.sh
│ └── skills/
│ ├── accessing-mlflow/ compare-results/ debug/
│ ├── deployment/ eagle3-new-model/ eagle3-review-logs/
│ ├── eagle3-triage/ eagle3-validate/ evaluation/
│ ├── launching-evals/ monitor/ ptq/
│ ├── quant-recipe-search/ release-cherry-pick/ common/
│
├── .claude/ ← back-compat (relative symlinks)
│ ├── clusters.yaml.example → ../.agents/clusters.yaml.example
│ ├── scripts → ../.agents/scripts
│ └── skills → ../.agents/skills
│
└── (future agents — add a symlink/config, no copies)
├── .codex/skills → ../.agents/skills
└── .cursor/skills → ../.agents/skills
```
### Why symlinks (and not "just point each agent's config at
`.agents/`")
Claude Code **only** auto-discovers project skills under
`.claude/skills/` — there is no setting/env var to redirect discovery to
an arbitrary path, and the plugin route would require committing a
`.claude/settings.json` + marketplace manifest, add a first-open
workspace-trust gate (breaks headless/CI runs), and namespace every
skill (`/ptq` → `/<plugin>:ptq`). A single relative in-repo symlink is
the smallest change that keeps `.agents/` canonical while satisfying
Claude Code's discovery requirement. This repo already commits relative
symlinks (`CLAUDE.md`, `tools/launcher/modules/Model-Optimizer`). See
the discussion thread for the full comparison.
### Changes
- Move `.claude/{skills,scripts,clusters.yaml.example}` → `.agents/`
(git renames preserve history).
- Add `.agents/README.md` documenting the convention and per-agent
wiring.
- Re-add `.claude/skills`, `.claude/scripts`,
`.claude/clusters.yaml.example` as relative symlinks into `.agents/`.
- Update internal path references and lint/sync config from `.claude/`
to `.agents/` (upstream provenance paths and `.claude/clusters.yaml`
back-compat lookups left intact).
- **Merged latest `main`** and folded in skills added there since branch
time — `compare-results`, `eagle3-new-model`, `eagle3-triage`,
`eagle3-review-logs`, `eagle3-validate`, `quant-recipe-search`, and new
`evaluation` recipes/tasks/references — into `.agents/skills/`.
- `main` converted `CLAUDE.md` into a symlink to a new agent-agnostic
`AGENTS.md`; this PR keeps that and moves the "skills live in
`.agents/`" guidance into `AGENTS.md`.
### Testing
- `.claude/skills` symlink resolves to all 15 skills; `ls
.claude/skills` and `ls .agents/skills` match.
- `pre-commit run check-symlinks --all-files` and `markdownlint-cli2
--all-files` pass.
- `bash -n` clean on `sync-upstream-skills.sh` and `remote_exec.sh`.
- `.claude/skills`, `.claude/scripts`, `.claude/clusters.yaml.example`,
and `CLAUDE.md` are all recorded as git symlinks (mode `120000`).
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).
- Is this change backward compatible?: ✅ `.claude/skills/`,
`.claude/scripts/`, and `.claude/clusters.yaml.example` continue to
resolve to the same content via symlinks; Claude Code auto-discovery is
unaffected; `remote_exec.sh` still accepts `.claude/clusters.yaml`.
- Did you write any new necessary tests?: N/A — directory move with
symlinks; verified as listed above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — repo housekeeping only, no API/feature/bugfix change.
### Additional Information
- A few vendored skill files still carry internal details (Slurm account
names, lustre paths, internal `:5005` GitLab registry advice in
`launching-evals/`) worth scrubbing in a follow-up.
---------
Signed-off-by: Seonghee Lee <seongheel@nvidia.com>
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Shengliang Xu <106840466+shengliangxu@users.noreply.github.com>
Co-authored-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Shengliang Xu
Zhiyu Cheng
Claude Opus 4.8
parent
54ce4e09d8
commit
bdb793a42d
@@ -0,0 +1,53 @@
|
||||
# `.agents/` — agent-agnostic source of truth
|
||||
|
||||
This directory is the canonical location for assets shared by AI coding agents
|
||||
working in this repository (Claude Code, Codex, Cursor, …).
|
||||
|
||||
## Layout
|
||||
|
||||
```text
|
||||
.agents/
|
||||
├── skills/ # SKILL.md files (canonical)
|
||||
│ └── <skill-name>/SKILL.md
|
||||
├── scripts/ # shared helper scripts (sync-upstream-skills.sh, …)
|
||||
└── clusters.yaml.example # remote-cluster config template
|
||||
```
|
||||
|
||||
## Why this exists
|
||||
|
||||
Different agents look for skills/config in vendor-specific directories. Rather
|
||||
than maintaining N copies that drift out of sync, **`.agents/` is the single
|
||||
source of truth** — each agent's guidance or install mechanism points here
|
||||
directly.
|
||||
|
||||
## How each agent finds these
|
||||
|
||||
Each agent points at `.agents/` through whatever mechanism it supports — never
|
||||
a copy:
|
||||
|
||||
- **Claude Code** only auto-discovers skills under `.claude/skills/`, so
|
||||
`.claude/` holds relative in-repo symlinks back into `.agents/`:
|
||||
`.claude/skills → ../.agents/skills`, `.claude/scripts → ../.agents/scripts`,
|
||||
and `.claude/clusters.yaml.example → ../.agents/clusters.yaml.example`. These
|
||||
follow the same committed-symlink pattern already used elsewhere in this repo
|
||||
(e.g. `CLAUDE.md`, `tools/launcher/modules/Model-Optimizer`).
|
||||
- **Future agents** (Codex, Cursor, …) add their own symlink or config pointing
|
||||
at `.agents/`.
|
||||
|
||||
## Editing rules
|
||||
|
||||
- **Always edit files under `.agents/`**.
|
||||
- Vendored-verbatim skills (`launching-evals`, `accessing-mlflow`) are managed
|
||||
by `.agents/scripts/sync-upstream-skills.sh` — do not modify by hand.
|
||||
- New skills go in `.agents/skills/<skill-name>/SKILL.md` following the
|
||||
conventions of existing skills (e.g. `.agents/skills/monitor/SKILL.md`).
|
||||
|
||||
## Project-level cluster config
|
||||
|
||||
The remote-execution skills look for a `clusters.yaml` at, in order:
|
||||
|
||||
1. `~/.config/modelopt/clusters.yaml` (user-level, recommended)
|
||||
2. `<repo-root>/.agents/clusters.yaml` (project-level, canonical)
|
||||
3. `<repo-root>/.claude/clusters.yaml` (project-level, back-compat)
|
||||
|
||||
See `clusters.yaml.example` for the schema.
|
||||
@@ -0,0 +1,19 @@
|
||||
# ModelOpt Remote Cluster Configuration
|
||||
# Copy to ~/.config/modelopt/clusters.yaml (user-level, recommended)
|
||||
# or .agents/clusters.yaml (project-level, can be committed).
|
||||
# .claude/clusters.yaml is also accepted for back-compat.
|
||||
|
||||
clusters:
|
||||
# GPU workstation or SLURM login node
|
||||
my-cluster:
|
||||
login_node: cluster-login.example.com
|
||||
user: myusername
|
||||
ssh_key: ~/.ssh/id_rsa
|
||||
# ssh_proxy: "socat - PROXY:localhost:%h:%p,proxyport=3128" # optional
|
||||
workspace: /path/to/remote/workdir
|
||||
gpu_type: H100 # used for quantization format recommendation
|
||||
# slurm:
|
||||
# default_account: my_account
|
||||
# default_partition: batch_short
|
||||
|
||||
default_cluster: my-cluster
|
||||
@@ -21,15 +21,18 @@
|
||||
# NOT managed by this script — update it manually when pulling upstream changes.
|
||||
#
|
||||
# Usage:
|
||||
# .claude/scripts/sync-upstream-skills.sh # re-vendor at the pinned SHA
|
||||
# UPSTREAM_SHA=<sha> .claude/scripts/sync-upstream-skills.sh # bump to a new SHA
|
||||
# .agents/scripts/sync-upstream-skills.sh # re-vendor at the pinned SHA
|
||||
# UPSTREAM_SHA=<sha> .agents/scripts/sync-upstream-skills.sh # bump to a new SHA
|
||||
#
|
||||
# Requires: gh, base64, awk. Run from the repo root.
|
||||
#
|
||||
# The script overwrites .claude/skills/<skill>/ with upstream contents and
|
||||
# The script overwrites .agents/skills/<skill>/ with upstream contents and
|
||||
# re-applies our provenance lines into each SKILL.md frontmatter. If you have
|
||||
# local changes to a vendored skill, they will be lost — that is expected,
|
||||
# since vendored-verbatim skills should not be modified locally.
|
||||
#
|
||||
# Note: .claude/skills/ (and other agent-specific skill dirs) are symlinks to
|
||||
# .agents/skills/ — see .agents/README.md.
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
@@ -40,7 +43,7 @@ SHORT_SHA="${SHA:0:7}"
|
||||
|
||||
UPSTREAM_REPO="NVIDIA-NeMo/Evaluator"
|
||||
UPSTREAM_BASE="packages/nemo-evaluator-launcher/.claude/skills"
|
||||
DEST_BASE=".claude/skills"
|
||||
DEST_BASE=".agents/skills"
|
||||
|
||||
if [[ ! -d "$DEST_BASE" ]]; then
|
||||
echo "error: run from the repo root (expected $DEST_BASE/ to exist)" >&2
|
||||
@@ -116,7 +119,7 @@ inject_provenance() {
|
||||
print "license: Apache-2.0"
|
||||
print "# Vendored verbatim from NVIDIA NeMo Evaluator (commit " short ")"
|
||||
print "# https://github.com/NVIDIA-NeMo/Evaluator/tree/" sha "/packages/nemo-evaluator-launcher/.claude/skills/" skill
|
||||
print "# To re-sync: .claude/scripts/sync-upstream-skills.sh"
|
||||
print "# To re-sync: .agents/scripts/sync-upstream-skills.sh"
|
||||
if (extra != "") {
|
||||
n = split(extra, lines, "\\|")
|
||||
for (i = 1; i <= n; i++) print "# " lines[i]
|
||||
+1
-1
@@ -34,7 +34,7 @@ If the cluster config contains multiple clusters and the user did not name the t
|
||||
For remote, connect:
|
||||
|
||||
```bash
|
||||
source .claude/skills/common/remote_exec.sh
|
||||
source .agents/skills/common/remote_exec.sh
|
||||
remote_load_cluster <cluster_name>
|
||||
remote_check_ssh
|
||||
remote_detect_env # sets REMOTE_ENV_TYPE = slurm / docker / bare
|
||||
+7
-6
@@ -9,8 +9,9 @@ Read this when Claude Code runs on a different machine than the target GPU clust
|
||||
Config locations (checked in order, first found wins):
|
||||
|
||||
1. `~/.config/modelopt/clusters.yaml` — user-level (not committed, recommended)
|
||||
2. `.claude/clusters.yaml` — project-level (can be committed for shared defaults)
|
||||
3. Interactive input — if neither file exists, ask the user (see SKILL.md Step 0) and write `~/.config/modelopt/clusters.yaml` before proceeding
|
||||
2. `.agents/clusters.yaml` — project-level, canonical (can be committed for shared defaults)
|
||||
3. `.claude/clusters.yaml` — project-level, back-compat
|
||||
4. Interactive input — if no file exists, ask the user (see SKILL.md Step 0) and write `~/.config/modelopt/clusters.yaml` before proceeding
|
||||
|
||||
```yaml
|
||||
clusters:
|
||||
@@ -38,14 +39,14 @@ rsync -av /path/to/local/checkpoint <cluster-login>:<cluster-workspace>/<session
|
||||
|
||||
Use the `workspace` path from your cluster config as the destination root, and keep staged checkpoints under the session/model directory. Compute nodes on a given cluster share the same storage as its login node, so once staged, the path works everywhere on that cluster.
|
||||
|
||||
See `.claude/clusters.yaml.example` for a fully annotated example with multiple cluster types.
|
||||
See `.agents/clusters.yaml.example` for a fully annotated example with multiple cluster types.
|
||||
|
||||
---
|
||||
|
||||
## 2. Connect and Establish Persistent Session
|
||||
|
||||
```bash
|
||||
source .claude/skills/common/remote_exec.sh
|
||||
source .agents/skills/common/remote_exec.sh
|
||||
remote_load_cluster <cluster_name> # or omit name to use default_cluster
|
||||
remote_check_ssh # validates connectivity + starts persistent session
|
||||
```
|
||||
@@ -153,5 +154,5 @@ remote_sync_from <remote_output_subdir> /local/output/
|
||||
## Reference Files
|
||||
|
||||
- **`skills/common/remote_exec.sh`** — Full utility library (session, run, sync, SLURM, Docker helpers)
|
||||
- **`.claude/clusters.yaml`** — Active cluster configuration
|
||||
- **`.claude/clusters.yaml.example`** — Annotated example config
|
||||
- **`.agents/clusters.yaml`** — Active cluster configuration (canonical; `.claude/clusters.yaml` also accepted for back-compat)
|
||||
- **`.agents/clusters.yaml.example`** — Annotated example config
|
||||
@@ -17,7 +17,7 @@
|
||||
# remote_exec.sh — Remote execution utility for ModelOpt agent skills
|
||||
#
|
||||
# Usage:
|
||||
# source .claude/skills/common/remote_exec.sh
|
||||
# source .agents/skills/common/remote_exec.sh
|
||||
# remote_load_cluster <cluster_name> # or: remote_load_cluster (uses default)
|
||||
# remote_check_ssh
|
||||
# remote_detect_env # detect SLURM vs Docker vs bare metal
|
||||
@@ -41,12 +41,17 @@
|
||||
# ── Helpers ──────────────────────────────────────────────────────────────────
|
||||
|
||||
_remote_config_file() {
|
||||
# Find clusters.yaml: user-level > project-level
|
||||
# Find clusters.yaml: user-level > project-level.
|
||||
# Project-level is checked at .agents/clusters.yaml (canonical) and then
|
||||
# .claude/clusters.yaml (back-compat).
|
||||
local user_config="${HOME}/.config/modelopt/clusters.yaml"
|
||||
local project_config
|
||||
# Walk up from pwd looking for .claude/clusters.yaml
|
||||
local dir="$PWD"
|
||||
while [[ "$dir" != "/" ]]; do
|
||||
if [[ -f "$dir/.agents/clusters.yaml" ]]; then
|
||||
project_config="$dir/.agents/clusters.yaml"
|
||||
break
|
||||
fi
|
||||
if [[ -f "$dir/.claude/clusters.yaml" ]]; then
|
||||
project_config="$dir/.claude/clusters.yaml"
|
||||
break
|
||||
@@ -196,7 +201,7 @@ remote_load_cluster() {
|
||||
if [[ -z "$config_file" ]]; then
|
||||
echo "ERROR: No clusters.yaml found. Provide cluster info interactively or create one." >&2
|
||||
echo " User config: ~/.config/modelopt/clusters.yaml" >&2
|
||||
echo " Project config: .claude/clusters.yaml" >&2
|
||||
echo " Project config: .agents/clusters.yaml (or .claude/clusters.yaml)" >&2
|
||||
return 1
|
||||
fi
|
||||
|
||||
@@ -215,7 +215,7 @@ which docker 2>/dev/null && echo "RUNTIME=docker"
|
||||
|
||||
| Runtime | Typical clusters | SLURM integration |
|
||||
| --- | --- | --- |
|
||||
| **enroot/pyxis** | NVIDIA internal (DGX Cloud, EOS, Selene, GCP-NRT) | `srun --container-image` |
|
||||
| **enroot/pyxis** | HPC clusters with container runtime (e.g. DGX Cloud and similar Slurm + container setups) | `srun --container-image` |
|
||||
| **Docker** | Bare-metal / on-prem with GPU | `docker run` inside job script |
|
||||
|
||||
### Step 2: Check credentials for the image's registry
|
||||
@@ -30,7 +30,7 @@ change is being measured, typically a further quantized version of the baseline.
|
||||
judge/code-execution status, sample accounting, and reasoning parsing before
|
||||
computing deltas.
|
||||
5. For each task, use the canonical score field from the matching
|
||||
`.claude/skills/evaluation/recipes/tasks/<task>.md` Score Extraction
|
||||
`.agents/skills/evaluation/recipes/tasks/<task>.md` Score Extraction
|
||||
section.
|
||||
6. Compute exact deltas outside the chat context when there are multiple tasks
|
||||
or repeated runs.
|
||||
@@ -192,7 +192,7 @@ If a cluster config exists (`~/.config/modelopt/clusters.yaml` or `.claude/clust
|
||||
1. **Source remote utilities:**
|
||||
|
||||
```bash
|
||||
source .claude/skills/common/remote_exec.sh
|
||||
source .agents/skills/common/remote_exec.sh
|
||||
remote_load_cluster
|
||||
remote_check_ssh
|
||||
remote_detect_env
|
||||
+1
-1
@@ -27,7 +27,7 @@
|
||||
"files": [],
|
||||
"expected_behavior": [
|
||||
"Checks for cluster config at ~/.config/modelopt/clusters.yaml or .claude/clusters.yaml",
|
||||
"Sources .claude/skills/common/remote_exec.sh",
|
||||
"Sources .agents/skills/common/remote_exec.sh",
|
||||
"Calls remote_load_cluster, remote_check_ssh, remote_detect_env",
|
||||
"Checks if checkpoint is already on remote (e.g., from prior PTQ run) before syncing; only syncs if local",
|
||||
"For SLURM: writes a job script with srun --container-image and --container-mounts on srun line (not #SBATCH)",
|
||||
@@ -62,9 +62,9 @@ The complete evaluation workflow is divided into the following steps you should
|
||||
# Key Facts
|
||||
|
||||
- Benchmark-specific info learned during launching/analyzing evals should be added to `references/benchmarks/`
|
||||
- **PPP** = Slurm account (the `account` field in cluster_config.yaml). When the user says "change PPP to X", update the account value (e.g., `coreai_dlalgo_compeval` → `coreai_dlalgo_llm`).
|
||||
- **PPP** = Slurm account / project portfolio code (the `account` field in cluster_config.yaml). When the user says "change PPP to X", update the account value (e.g., `<old_account>` → `<new_account>`).
|
||||
- **Slurm job pairs**: NEL (nemo-evaluator-launcher) submits paired Slurm jobs — a RUNNING job + a PENDING restart job (for when the 4h walltime expires). Never cancel the pending restart jobs — they are expected and necessary.
|
||||
- **HF cache requirement**: For configs with `HF_HUB_OFFLINE=1`, models must be pre-downloaded to the HF cache on each cluster before launching. **Before running a model on a new cluster, always ask the user if the model is already cached there.** If not, on the cluster login node: `python3 -m venv hf_cli && source hf_cli/bin/activate && pip install huggingface_hub` then `HF_HOME=/lustre/fsw/portfolios/coreai/users/<username>/cache/huggingface hf download <model>`. Without this, vLLM will fail with `LocalEntryNotFoundError`.
|
||||
- **HF cache requirement**: For configs with `HF_HUB_OFFLINE=1`, models must be pre-downloaded to the HF cache on each cluster before launching. **Before running a model on a new cluster, always ask the user if the model is already cached there.** If not, on the cluster login node: `python3 -m venv hf_cli && source hf_cli/bin/activate && pip install huggingface_hub` then `HF_HOME=<your_hf_cache_path> hf download <model>` (on lustre-style HPC clusters this is typically under `/lustre/.../<group>/users/<username>/cache/huggingface`). Without this, vLLM will fail with `LocalEntryNotFoundError`.
|
||||
- **`data_parallel_size` is per node**: `dp_size=1` with `num_nodes=8` means 8 model instances total (one per node), load-balanced by haproxy. Do NOT interpret `dp_size` as the global replica count.
|
||||
- **`payload_modifier` interceptor**: The `params_to_remove` list (e.g. `[max_tokens, max_completion_tokens]`) strips those fields from the outgoing payload, intentionally lifting output length limits so reasoning models can think as long as they need.
|
||||
- **Auto-export git workaround**: The export container (`python:3.12-slim`) lacks `git`. When installing the launcher from a git URL, set `auto_export.launcher_install_cmd` to install git first (e.g., `apt-get update -qq && apt-get install -qq -y git && pip install "nemo-evaluator-launcher[all] @ git+...#subdirectory=packages/nemo-evaluator-launcher"`).
|
||||
+2
-2
@@ -70,7 +70,7 @@ tail -200 $LOGS/client-*.log
|
||||
- **CUDA OOM**: Increase `deployment.tensor_parallel_size` to shard across more GPUs. For multi-node: increase `execution.num_nodes` and set `deployment.pipeline_parallel_size`. As last resort: add `--max-model-len <lower_value>` to `deployment.extra_args`. Do NOT quantize as a first fix — scale compute instead.
|
||||
- **Missing model/checkpoint**: `FileNotFoundError` or `RepositoryNotFoundError` or `GatedRepoError: 403` — verify `deployment.checkpoint_path` or `deployment.hf_model_handle`. For gated models, set `HF_TOKEN` via `deployment.env_vars`.
|
||||
- **Bad `extra_args`**: `unrecognized arguments` or `unexpected keyword argument` — check flags against deployment engine version. Some flags change between versions (e.g., `--rope-scaling` removed in vLLM > 0.11.0).
|
||||
- **Image pull failure**: `manifest not found` or `pyxis: child 1 failed` — verify image tag exists. Drop `:5005` from GitLab container registry URLs.
|
||||
- **Image pull failure**: `manifest not found` or `pyxis: child 1 failed` — verify image tag exists. If the image is on an on-prem GitLab registry, drop the registry port suffix (e.g. `:5005`) from the URL.
|
||||
- **GPU driver mismatch**: `CUDA driver version is insufficient` — use an older container image matching the host CUDA driver.
|
||||
- **Health check timeout / connection refused**: Server didn't start — check server logs first. Increase `execution.endpoint_readiness_timeout` (seconds). SLURM default: `null` (falls back to walltime).
|
||||
- **Server crashed mid-eval**: `Connection reset by peer` — check server logs for OOM. Reduce `parallelism` (concurrent requests). Check SLURM logs for preemption or walltime exceeded.
|
||||
@@ -80,7 +80,7 @@ tail -200 $LOGS/client-*.log
|
||||
- **Config validation**: `MissingMandatoryValue` (unfilled `???`), `ValidationError` (type mismatch), `ScannerError` (invalid YAML) — run `--dry-run` to catch these upfront.
|
||||
- **Walltime exceeded**: `CANCELLED DUE TO TIME LIMIT` — NEL submits paired restart jobs that automatically resume when walltime expires, so this is often expected behavior, not a failure. Only increase `execution.walltime` if the evaluation isn't making progress across restarts.
|
||||
- **Preemption**: `CANCELLED DUE TO PREEMPTION` — the paired restart job should automatically resume. If it doesn't, use non-preemptible partition, or re-run.
|
||||
- **Container not found**: Applies to both `deployment.image` and task-level eval container. Drop `:5005` from GitLab registry URLs.
|
||||
- **Container not found**: Applies to both `deployment.image` and task-level eval container. For on-prem GitLab registries, drop the registry port suffix (e.g. `:5005`) from the URL.
|
||||
- Troubleshooting docs: list files with WebFetch `https://api.github.com/repos/NVIDIA-NeMo/Evaluator/contents/docs/troubleshooting`, then fetch relevant ones from `https://raw.githubusercontent.com/NVIDIA-NeMo/Evaluator/main/docs/troubleshooting/<file>`
|
||||
|
||||
**Fix Slurm invalid account/partition:**
|
||||
+1
-1
@@ -191,7 +191,7 @@ Do not reimplement workflows that existing skills own:
|
||||
| Compute baseline-vs-candidate deltas | `compare-results` |
|
||||
|
||||
Before launching PTQ in a ModelOpt repo, read the current PTQ skill from
|
||||
`.claude/skills/ptq/SKILL.md`; recipe paths and validation gates can change.
|
||||
`.agents/skills/ptq/SKILL.md`; recipe paths and validation gates can change.
|
||||
|
||||
## ModelOpt Starting Points
|
||||
|
||||
@@ -1,18 +0,0 @@
|
||||
# ModelOpt Remote Cluster Configuration
|
||||
# Copy to ~/.config/modelopt/clusters.yaml (user-level, recommended)
|
||||
# or .claude/clusters.yaml (project-level, can be committed).
|
||||
|
||||
clusters:
|
||||
# GPU workstation or SLURM login node
|
||||
my-cluster:
|
||||
login_node: cluster-login.example.com
|
||||
user: myusername
|
||||
ssh_key: ~/.ssh/id_rsa
|
||||
# ssh_proxy: "socat - PROXY:localhost:%h:%p,proxyport=3128" # optional
|
||||
workspace: /path/to/remote/workdir
|
||||
gpu_type: H100 # used for quantization format recommendation
|
||||
# slurm:
|
||||
# default_account: my_account
|
||||
# default_partition: batch_short
|
||||
|
||||
default_cluster: my-cluster
|
||||
Symlink
+1
@@ -0,0 +1 @@
|
||||
../.agents/clusters.yaml.example
|
||||
Symlink
+1
@@ -0,0 +1 @@
|
||||
../.agents/scripts
|
||||
Symlink
+1
@@ -0,0 +1 @@
|
||||
../.agents/skills
|
||||
@@ -12,7 +12,7 @@ config:
|
||||
MD059: false # no-hard-tabs
|
||||
|
||||
# Vendored upstream skills — kept byte-identical to upstream via
|
||||
# .claude/scripts/sync-upstream-skills.sh; do not reformat.
|
||||
# .agents/scripts/sync-upstream-skills.sh; do not reformat.
|
||||
ignores:
|
||||
- ".claude/skills/launching-evals/**"
|
||||
- ".claude/skills/accessing-mlflow/**"
|
||||
- ".agents/skills/launching-evals/**"
|
||||
- ".agents/skills/accessing-mlflow/**"
|
||||
|
||||
@@ -7,6 +7,12 @@ These instructions apply to AI-assisted work in this repository.
|
||||
- Start with `README.md` for project overview and install.
|
||||
- Use `modelopt/` for source, `tests/` for focused test coverage, and
|
||||
`examples/` or `docs/` for usage patterns.
|
||||
- **Agent skills and shared config live under `.agents/`** — the canonical,
|
||||
agent-agnostic source of truth (`.agents/skills/<name>/SKILL.md`,
|
||||
`.agents/scripts/`, `.agents/clusters.yaml.example`). Claude Code's
|
||||
`.claude/skills`, `.claude/scripts`, and `.claude/clusters.yaml.example` are
|
||||
relative symlinks into `.agents/`. Always edit files under `.agents/`, not the
|
||||
symlink path. See `.agents/README.md` for the convention.
|
||||
|
||||
## Coding guidelines
|
||||
|
||||
|
||||
Reference in New Issue
Block a user