mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
eval skill: auto-detect predefined per-cluster execution configs (#1599)
### What does this PR do? Type of change: documentation Some NEL installs ship ready-made per-cluster execution configs as an `internal/slurm/<cluster>` group (via the optional `nemo_evaluator_launcher_internal` package). When present, the matching one pre-fills the cluster's `hostname` / `partition` / `gres` (and node-exclusivity), so the user only sets `account` / `output_dir` / `walltime` instead of hand-entering the hostname. - **`SKILL.md` Step 4** — adds a "check FIRST" step: run a discovery snippet that lists the available `cluster → hostname` pairs **from the installed package at runtime**; on a hostname match, use `defaults: - execution: internal/slurm/<cluster>` (replacing `slurm/default`) and drop the now-redundant `execution.hostname`. If the package isn't installed or nothing matches, fall back to `slurm/default` and fill the fields manually. - **`example_eval.yaml`** — a short, name-free comment on the `defaults` block pointing to that check. **Discovery-based by design (no internal data committed):** cluster names, hostnames, and accounts are read from the install at runtime and are **not** hardcoded here. The only internal references in the repo are the package name `nemo_evaluator_launcher_internal` and the generic `internal/slurm/<cluster>` group pattern. External users without the package degrade gracefully (import fails → `slurm/default`). ### Usage N/A — documentation / skill guidance only. ### Testing Ran the discovery snippet verbatim (lists `cluster → hostname` from the installed package) and confirmed `internal/slurm/<cluster>` resolves via a `--dry-run`. `pre-commit run` passes (markdownlint + YAML format). Grep-verified no cluster names/hostnames are hardcoded in the changed files. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (falls back to `slurm/default`) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (documentation) - Did you update Changelog?: N/A (skill docs) - Did you get Claude approval on this PR?: ❌ (pending) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added Step 4 guidance for discovering and using optional internal per-cluster SLURM execution configs: how to list available internal configs, select the matching cluster by switching the default execution, remove redundant hostname entries, and verify with a --dry-run. * Clarified that slurm/default is cluster-agnostic, recommend using internal per-cluster configs when present, and explained fallback behavior to manual hostname/account/output_dir entry. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
651fd223e6
commit
88fd7ff958
@@ -197,6 +197,17 @@ Reasoning models: prefer reasoning mode (highest scores). For lower variance / c
|
||||
|
||||
### Step 4 — Fill remaining ??? values
|
||||
|
||||
**Predefined per-cluster execution config (check FIRST).** Some installs ship `internal/slurm/<cluster>` execution groups (optional `nemo_evaluator_launcher_internal` pkg) that pre-fill hostname/partition/gres — leaving only account/output_dir/walltime. Discover at runtime (nothing cluster-specific hardcoded):
|
||||
|
||||
```bash
|
||||
python3 -c 'import nemo_evaluator_launcher_internal' 2>/dev/null && \
|
||||
PKG=$(python3 -c 'import nemo_evaluator_launcher_internal as m,os;print(os.path.dirname(m.__file__))') && \
|
||||
for f in "$PKG"/configs/execution/internal/slurm/*.yaml; do \
|
||||
echo "$(basename "$f" .yaml) -> $(grep -E '^hostname:' "$f" | awk '{print $2}')"; done
|
||||
```
|
||||
|
||||
Hostname match → set `defaults: - execution: internal/slurm/<cluster>`, drop the redundant `execution.hostname` (keep account/output_dir/walltime), verify with `--dry-run`. Else keep `slurm/default` and fill hostname/account/output_dir manually.
|
||||
|
||||
- Find every `???` left. Ask the user only for what can't be inferred (SLURM hostname/account/output_dir, MLflow tracking URI, etc.). Don't propose defaults; let them give plain text.
|
||||
- **`parallelism`** — size it yourself from the run shape (total requests = `dataset_size × repeats` vs GPU serving capacity), and set `--max-num-seqs` to match. Read `references/parallelism.md` for the decision rule and worked examples; only ask the user if a non-GPU cap (e.g. judge rate limit) is unknown.
|
||||
- Ask about other defaults they may want to change (partition, walltime, MLflow tags).
|
||||
|
||||
@@ -41,6 +41,8 @@
|
||||
# parallelism.
|
||||
# nel run --config ... -o ++evaluation.nemo_evaluator_config.config.params.limit_samples=2
|
||||
defaults:
|
||||
# slurm/default works anywhere; if your install ships a predefined
|
||||
# internal/slurm/<cluster> config, prefer it (see SKILL.md Step 4).
|
||||
- execution: slurm/default
|
||||
- deployment: vllm
|
||||
- _self_
|
||||
|
||||
Reference in New Issue
Block a user