From be740012568eb2ff97af1c08f7a283c5edd766c2 Mon Sep 17 00:00:00 2001 From: sychen52 <41452870+sychen52@users.noreply.github.com> Date: Mon, 28 Sep 2026 17:33:53 -0700 Subject: [PATCH] Add concise AgentX benchmark skill (#2573) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ### What does this PR do? Type of change: Documentation. Adds a concise AgentX skill covering harness installation, automatic dataset downloads, benchmark execution, and result reporting. Reuses the existing deployment skill and adds a Claude discovery link. ### Usage Use run-agentx to benchmark my deployed model with a concurrency sweep. ### Testing • Skill structure and metadata validation passed. • Shell syntax checks passed. • Benchmark arguments parsed and produced a valid configuration using the pinned harness. • All applicable pre-commit checks passed. • No GPU benchmark was launched. ### Before your PR is "Ready for review" Contributor guidelines and security practices were reviewed. The commit is signed and signed off. • Is this change backward compatible?: ✅ • If you copied code from other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md?: N/A. No copied implementation or project dependency changes. • Did you write any new necessary tests?: N/A. Documentation changes were validated as described above. • Did you update Changelog?: N/A. Skill documentation only. • Did you get Claude approval on this PR?: ❌ Not run. ### Additional Information ## Summary by CodeRabbit * **Documentation** * Added guidance for configuring and running SemiAnalysis AgentX serving benchmarks, including endpoint, model, tokenizer, context limits, caching, and dataset setup. * Documented using a pinned benchmark harness in a separate client environment and running each concurrency level in a fresh artifact directory with a fixed seed. * Expanded reporting guidance to cover overlapping requests, cache and preemption metrics, errors, unfinished requests, warmup failures, and submission validity. * Clarified that missing server-reported usage makes cache-hit data unknown, invalid or missing submission validity should be flagged, and smoke runs are not benchmark results. * Directed AgentX agentic workloads from the optional AIPerf guidance to the AgentX benchmark instructions. Signed-off-by: Shiyang Chen --- .claude/skills/run-agentx | 1 + plugins/modelopt/skills/run-agentx/SKILL.md | 90 +++++++++++++++++++++ 2 files changed, 91 insertions(+) create mode 120000 .claude/skills/run-agentx create mode 100644 plugins/modelopt/skills/run-agentx/SKILL.md diff --git a/.claude/skills/run-agentx b/.claude/skills/run-agentx new file mode 120000 index 000000000..e45ee744e --- /dev/null +++ b/.claude/skills/run-agentx @@ -0,0 +1 @@ +../../.agents/skills/run-agentx \ No newline at end of file diff --git a/plugins/modelopt/skills/run-agentx/SKILL.md b/plugins/modelopt/skills/run-agentx/SKILL.md new file mode 100644 index 000000000..2f4097b90 --- /dev/null +++ b/plugins/modelopt/skills/run-agentx/SKILL.md @@ -0,0 +1,90 @@ +--- +name: run-agentx +description: Run the AgentX agentic serving benchmark with the SemiAnalysis harness. Use for AgentX setup, dataset downloads, concurrency sweeps, or serving latency and throughput comparisons. +license: Apache-2.0 +--- + +# Run AgentX + +## Setup + +Resolve the endpoint URL, served model name, matching tokenizer, and requested +concurrency from the task. Reuse the model checkpoint. If serving is needed, follow +[deployment](../deployment/SKILL.md) to download missing weights and launch it. +Enable prefix caching, streaming usage, and cached-token reporting. For vLLM, use +`--enable-prefix-caching --enable-prompt-tokens-details`. Verify the context limit +fits the traces with the model's tokenizer; never silently truncate requests. + +Install the [AgentX fork](https://github.com/SemiAnalysisAI/agentx-harness) in a +separate client environment. The pinned revision supplies the scenario and dataset +loader used below; installing the upstream `aiperf` package is insufficient. + +```bash +set -euo pipefail +export AGENTX_WORKDIR="${AGENTX_WORKDIR:-$PWD/agentx-work}" +python3.12 -m venv "$AGENTX_WORKDIR/venv" +source "$AGENTX_WORKDIR/venv/bin/activate" +python -m pip install 'aiperf @ git+https://github.com/SemiAnalysisAI/agentx-harness.git@56a0cf70f4c0359454ee4bd15a17770b541a3e3e' +``` + +Set `AGENTX_DATASET` to an accepted, date-pinned with-subagents alias from the +[pinned AgentX tutorial](https://github.com/SemiAnalysisAI/agentx-harness/blob/56a0cf70f4c0359454ee4bd15a17770b541a3e3e/docs/tutorials/agentx-mvp.md). +The public loader downloads and caches the traces and prompt reconstruction assets. +Record the resolved dataset revision; dated aliases can move. Prefer the deployed +checkpoint's local tokenizer. If it requires custom code, add +`--tokenizer-trust-remote-code` only after the user explicitly trusts its repository. + +## Choose concurrency + +Sweep concurrency upward until KV-cache usage nears capacity, then refine around +that point. Resweep when the model, hardware, or cache budget changes. Publish the +full sweep and repeat promising points before claiming an advantage. + +## Run + +In each shell, set `AGENTX_WORKDIR` to the setup directory and set `AGENTX_DATASET`, +`AGENTX_URL`, `AGENTX_MODEL`, `AGENTX_TOKENIZER`, `AGENTX_MAX_CONTEXT_LENGTH`, +and `AGENTX_CONCURRENCY` from the deployment and sweep. Use the server's base URL +and actual context limit. Fix `AGENTX_SEED` across comparisons. +First check endpoint health and repeat a long prompt to verify nonzero cached-token +usage. Keep the scenario's default trajectory window and duration. Run each +concurrency separately with a fresh artifact directory: + +```bash +set -euo pipefail +source "${AGENTX_WORKDIR:?}/venv/bin/activate" +export HF_HOME="${HF_HOME:-$AGENTX_WORKDIR/hf-cache}" +export AIPERF_DATASET_MMAP_CACHE_DIR="$AGENTX_WORKDIR/dataset-cache" +mkdir -p "$AGENTX_WORKDIR/results" +AGENTX_RUN_DIR=$(mktemp -d "$AGENTX_WORKDIR/results/run-XXXXXX") +python -m pip freeze > "$AGENTX_RUN_DIR/client-packages.txt" +aiperf profile \ + --scenario inferencex-agentx-mvp \ + --url "${AGENTX_URL:?}" --model "${AGENTX_MODEL:?}" \ + --tokenizer "${AGENTX_TOKENIZER:?}" --endpoint-type chat \ + --public-dataset "${AGENTX_DATASET:?}" \ + --max-context-length "${AGENTX_MAX_CONTEXT_LENGTH:?}" \ + --concurrency "${AGENTX_CONCURRENCY:?}" \ + --random-seed "${AGENTX_SEED:?}" \ + --streaming --use-server-token-count --extra-inputs ignore_eos:true \ + --artifact-dir "$AGENTX_RUN_DIR" --ui simple \ + 2>&1 | tee "$AGENTX_RUN_DIR/client.log" +``` + +## Compare and report + +Use the [benchmarking guide](../deployment/references/benchmarking.md#3-run-a-sweep) +for sweep isolation, performance metrics, and comparison controls. For AgentX: + +- Concurrency counts session trees. Report actual overlapping HTTP requests + separately because child sessions can overlap. +- Hold KV-cache bytes fixed when comparing KV formats. Keep the dataset, seed, + context filter, and warmup fixed, and start a fresh server per point. +- Report token-weighted cache hits from server-reported usage, cache usage, and + preemptions. Missing usage means cache hits are unknown. +- Report `metadata.submission_valid` and any `metadata.submission_invalid_reasons` + from the result export. Flag invalid or missing validity instead of treating the + run as a valid comparison. +- Include request errors, unfinished requests, and warmup failures. + Synthetic prompts measure serving performance, not model quality; shortened + smoke runs are not benchmark results.