Add concise AgentX benchmark skill (#2573)

### What does this PR do?

  Type of change: Documentation.

Adds a concise AgentX skill covering harness installation, automatic
dataset downloads, benchmark execution, and result reporting. Reuses the
existing deployment skill
  and adds a Claude discovery link.

  ### Usage

Use run-agentx to benchmark my deployed model with a concurrency sweep.

  ### Testing

  • Skill structure and metadata validation passed.
  • Shell syntax checks passed.
• Benchmark arguments parsed and produced a valid configuration using
the pinned harness.
  • All applicable pre-commit checks passed.
  • No GPU benchmark was launched.

  ### Before your PR is "Ready for review"

Contributor guidelines and security practices were reviewed. The commit
is signed and signed off.

  • Is this change backward compatible?: ✅
• If you copied code from other sources or added a new PIP dependency,
did you follow guidance in CONTRIBUTING.md?: N/A. No copied
implementation or project dependency
    changes.

• Did you write any new necessary tests?: N/A. Documentation changes
were validated as described above.
  • Did you update Changelog?: N/A. Skill documentation only.
  • Did you get Claude approval on this PR?: ❌ Not run.

  ### Additional Information


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added guidance for configuring and running SemiAnalysis AgentX serving
benchmarks, including endpoint, model, tokenizer, context limits,
caching, and dataset setup.
* Documented using a pinned benchmark harness in a separate client
environment and running each concurrency level in a fresh artifact
directory with a fixed seed.
* Expanded reporting guidance to cover overlapping requests, cache and
preemption metrics, errors, unfinished requests, warmup failures, and
submission validity.
* Clarified that missing server-reported usage makes cache-hit data
unknown, invalid or missing submission validity should be flagged, and
smoke runs are not benchmark results.
* Directed AgentX agentic workloads from the optional AIPerf guidance to
the AgentX benchmark instructions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
This commit is contained in:
sychen52
2026-09-29 00:33:53 +00:00
committed by GitHub
parent 1d392999b4
commit be74001256
2 changed files with 91 additions and 0 deletions
+1
View File
@@ -0,0 +1 @@
../../.agents/skills/run-agentx
@@ -0,0 +1,90 @@
---
name: run-agentx
description: Run the AgentX agentic serving benchmark with the SemiAnalysis harness. Use for AgentX setup, dataset downloads, concurrency sweeps, or serving latency and throughput comparisons.
license: Apache-2.0
---
# Run AgentX
## Setup
Resolve the endpoint URL, served model name, matching tokenizer, and requested
concurrency from the task. Reuse the model checkpoint. If serving is needed, follow
[deployment](../deployment/SKILL.md) to download missing weights and launch it.
Enable prefix caching, streaming usage, and cached-token reporting. For vLLM, use
`--enable-prefix-caching --enable-prompt-tokens-details`. Verify the context limit
fits the traces with the model's tokenizer; never silently truncate requests.
Install the [AgentX fork](https://github.com/SemiAnalysisAI/agentx-harness) in a
separate client environment. The pinned revision supplies the scenario and dataset
loader used below; installing the upstream `aiperf` package is insufficient.
```bash
set -euo pipefail
export AGENTX_WORKDIR="${AGENTX_WORKDIR:-$PWD/agentx-work}"
python3.12 -m venv "$AGENTX_WORKDIR/venv"
source "$AGENTX_WORKDIR/venv/bin/activate"
python -m pip install 'aiperf @ git+https://github.com/SemiAnalysisAI/agentx-harness.git@56a0cf70f4c0359454ee4bd15a17770b541a3e3e'
```
Set `AGENTX_DATASET` to an accepted, date-pinned with-subagents alias from the
[pinned AgentX tutorial](https://github.com/SemiAnalysisAI/agentx-harness/blob/56a0cf70f4c0359454ee4bd15a17770b541a3e3e/docs/tutorials/agentx-mvp.md).
The public loader downloads and caches the traces and prompt reconstruction assets.
Record the resolved dataset revision; dated aliases can move. Prefer the deployed
checkpoint's local tokenizer. If it requires custom code, add
`--tokenizer-trust-remote-code` only after the user explicitly trusts its repository.
## Choose concurrency
Sweep concurrency upward until KV-cache usage nears capacity, then refine around
that point. Resweep when the model, hardware, or cache budget changes. Publish the
full sweep and repeat promising points before claiming an advantage.
## Run
In each shell, set `AGENTX_WORKDIR` to the setup directory and set `AGENTX_DATASET`,
`AGENTX_URL`, `AGENTX_MODEL`, `AGENTX_TOKENIZER`, `AGENTX_MAX_CONTEXT_LENGTH`,
and `AGENTX_CONCURRENCY` from the deployment and sweep. Use the server's base URL
and actual context limit. Fix `AGENTX_SEED` across comparisons.
First check endpoint health and repeat a long prompt to verify nonzero cached-token
usage. Keep the scenario's default trajectory window and duration. Run each
concurrency separately with a fresh artifact directory:
```bash
set -euo pipefail
source "${AGENTX_WORKDIR:?}/venv/bin/activate"
export HF_HOME="${HF_HOME:-$AGENTX_WORKDIR/hf-cache}"
export AIPERF_DATASET_MMAP_CACHE_DIR="$AGENTX_WORKDIR/dataset-cache"
mkdir -p "$AGENTX_WORKDIR/results"
AGENTX_RUN_DIR=$(mktemp -d "$AGENTX_WORKDIR/results/run-XXXXXX")
python -m pip freeze > "$AGENTX_RUN_DIR/client-packages.txt"
aiperf profile \
--scenario inferencex-agentx-mvp \
--url "${AGENTX_URL:?}" --model "${AGENTX_MODEL:?}" \
--tokenizer "${AGENTX_TOKENIZER:?}" --endpoint-type chat \
--public-dataset "${AGENTX_DATASET:?}" \
--max-context-length "${AGENTX_MAX_CONTEXT_LENGTH:?}" \
--concurrency "${AGENTX_CONCURRENCY:?}" \
--random-seed "${AGENTX_SEED:?}" \
--streaming --use-server-token-count --extra-inputs ignore_eos:true \
--artifact-dir "$AGENTX_RUN_DIR" --ui simple \
2>&1 | tee "$AGENTX_RUN_DIR/client.log"
```
## Compare and report
Use the [benchmarking guide](../deployment/references/benchmarking.md#3-run-a-sweep)
for sweep isolation, performance metrics, and comparison controls. For AgentX:
- Concurrency counts session trees. Report actual overlapping HTTP requests
separately because child sessions can overlap.
- Hold KV-cache bytes fixed when comparing KV formats. Keep the dataset, seed,
context filter, and warmup fixed, and start a fresh server per point.
- Report token-weighted cache hits from server-reported usage, cache usage, and
preemptions. Missing usage means cache hits are unknown.
- Report `metadata.submission_valid` and any `metadata.submission_invalid_reasons`
from the result export. Flag invalid or missing validity instead of treating the
run as a valid comparison.
- Include request errors, unfinished requests, and warmup failures.
Synthetic prompts measure serving performance, not model quality; shortened
smoke runs are not benchmark results.