mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Add concise AgentX benchmark skill (#2573)
### What does this PR do? Type of change: Documentation. Adds a concise AgentX skill covering harness installation, automatic dataset downloads, benchmark execution, and result reporting. Reuses the existing deployment skill and adds a Claude discovery link. ### Usage Use run-agentx to benchmark my deployed model with a concurrency sweep. ### Testing • Skill structure and metadata validation passed. • Shell syntax checks passed. • Benchmark arguments parsed and produced a valid configuration using the pinned harness. • All applicable pre-commit checks passed. • No GPU benchmark was launched. ### Before your PR is "Ready for review" Contributor guidelines and security practices were reviewed. The commit is signed and signed off. • Is this change backward compatible?: ✅ • If you copied code from other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md?: N/A. No copied implementation or project dependency changes. • Did you write any new necessary tests?: N/A. Documentation changes were validated as described above. • Did you update Changelog?: N/A. Skill documentation only. • Did you get Claude approval on this PR?: ❌ Not run. ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added guidance for configuring and running SemiAnalysis AgentX serving benchmarks, including endpoint, model, tokenizer, context limits, caching, and dataset setup. * Documented using a pinned benchmark harness in a separate client environment and running each concurrency level in a fresh artifact directory with a fixed seed. * Expanded reporting guidance to cover overlapping requests, cache and preemption metrics, errors, unfinished requests, warmup failures, and submission validity. * Clarified that missing server-reported usage makes cache-hit data unknown, invalid or missing submission validity should be flagged, and smoke runs are not benchmark results. * Directed AgentX agentic workloads from the optional AIPerf guidance to the AgentX benchmark instructions. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
This commit is contained in:
Symlink
+1
@@ -0,0 +1 @@
|
||||
../../.agents/skills/run-agentx
|
||||
@@ -0,0 +1,90 @@
|
||||
---
|
||||
name: run-agentx
|
||||
description: Run the AgentX agentic serving benchmark with the SemiAnalysis harness. Use for AgentX setup, dataset downloads, concurrency sweeps, or serving latency and throughput comparisons.
|
||||
license: Apache-2.0
|
||||
---
|
||||
|
||||
# Run AgentX
|
||||
|
||||
## Setup
|
||||
|
||||
Resolve the endpoint URL, served model name, matching tokenizer, and requested
|
||||
concurrency from the task. Reuse the model checkpoint. If serving is needed, follow
|
||||
[deployment](../deployment/SKILL.md) to download missing weights and launch it.
|
||||
Enable prefix caching, streaming usage, and cached-token reporting. For vLLM, use
|
||||
`--enable-prefix-caching --enable-prompt-tokens-details`. Verify the context limit
|
||||
fits the traces with the model's tokenizer; never silently truncate requests.
|
||||
|
||||
Install the [AgentX fork](https://github.com/SemiAnalysisAI/agentx-harness) in a
|
||||
separate client environment. The pinned revision supplies the scenario and dataset
|
||||
loader used below; installing the upstream `aiperf` package is insufficient.
|
||||
|
||||
```bash
|
||||
set -euo pipefail
|
||||
export AGENTX_WORKDIR="${AGENTX_WORKDIR:-$PWD/agentx-work}"
|
||||
python3.12 -m venv "$AGENTX_WORKDIR/venv"
|
||||
source "$AGENTX_WORKDIR/venv/bin/activate"
|
||||
python -m pip install 'aiperf @ git+https://github.com/SemiAnalysisAI/agentx-harness.git@56a0cf70f4c0359454ee4bd15a17770b541a3e3e'
|
||||
```
|
||||
|
||||
Set `AGENTX_DATASET` to an accepted, date-pinned with-subagents alias from the
|
||||
[pinned AgentX tutorial](https://github.com/SemiAnalysisAI/agentx-harness/blob/56a0cf70f4c0359454ee4bd15a17770b541a3e3e/docs/tutorials/agentx-mvp.md).
|
||||
The public loader downloads and caches the traces and prompt reconstruction assets.
|
||||
Record the resolved dataset revision; dated aliases can move. Prefer the deployed
|
||||
checkpoint's local tokenizer. If it requires custom code, add
|
||||
`--tokenizer-trust-remote-code` only after the user explicitly trusts its repository.
|
||||
|
||||
## Choose concurrency
|
||||
|
||||
Sweep concurrency upward until KV-cache usage nears capacity, then refine around
|
||||
that point. Resweep when the model, hardware, or cache budget changes. Publish the
|
||||
full sweep and repeat promising points before claiming an advantage.
|
||||
|
||||
## Run
|
||||
|
||||
In each shell, set `AGENTX_WORKDIR` to the setup directory and set `AGENTX_DATASET`,
|
||||
`AGENTX_URL`, `AGENTX_MODEL`, `AGENTX_TOKENIZER`, `AGENTX_MAX_CONTEXT_LENGTH`,
|
||||
and `AGENTX_CONCURRENCY` from the deployment and sweep. Use the server's base URL
|
||||
and actual context limit. Fix `AGENTX_SEED` across comparisons.
|
||||
First check endpoint health and repeat a long prompt to verify nonzero cached-token
|
||||
usage. Keep the scenario's default trajectory window and duration. Run each
|
||||
concurrency separately with a fresh artifact directory:
|
||||
|
||||
```bash
|
||||
set -euo pipefail
|
||||
source "${AGENTX_WORKDIR:?}/venv/bin/activate"
|
||||
export HF_HOME="${HF_HOME:-$AGENTX_WORKDIR/hf-cache}"
|
||||
export AIPERF_DATASET_MMAP_CACHE_DIR="$AGENTX_WORKDIR/dataset-cache"
|
||||
mkdir -p "$AGENTX_WORKDIR/results"
|
||||
AGENTX_RUN_DIR=$(mktemp -d "$AGENTX_WORKDIR/results/run-XXXXXX")
|
||||
python -m pip freeze > "$AGENTX_RUN_DIR/client-packages.txt"
|
||||
aiperf profile \
|
||||
--scenario inferencex-agentx-mvp \
|
||||
--url "${AGENTX_URL:?}" --model "${AGENTX_MODEL:?}" \
|
||||
--tokenizer "${AGENTX_TOKENIZER:?}" --endpoint-type chat \
|
||||
--public-dataset "${AGENTX_DATASET:?}" \
|
||||
--max-context-length "${AGENTX_MAX_CONTEXT_LENGTH:?}" \
|
||||
--concurrency "${AGENTX_CONCURRENCY:?}" \
|
||||
--random-seed "${AGENTX_SEED:?}" \
|
||||
--streaming --use-server-token-count --extra-inputs ignore_eos:true \
|
||||
--artifact-dir "$AGENTX_RUN_DIR" --ui simple \
|
||||
2>&1 | tee "$AGENTX_RUN_DIR/client.log"
|
||||
```
|
||||
|
||||
## Compare and report
|
||||
|
||||
Use the [benchmarking guide](../deployment/references/benchmarking.md#3-run-a-sweep)
|
||||
for sweep isolation, performance metrics, and comparison controls. For AgentX:
|
||||
|
||||
- Concurrency counts session trees. Report actual overlapping HTTP requests
|
||||
separately because child sessions can overlap.
|
||||
- Hold KV-cache bytes fixed when comparing KV formats. Keep the dataset, seed,
|
||||
context filter, and warmup fixed, and start a fresh server per point.
|
||||
- Report token-weighted cache hits from server-reported usage, cache usage, and
|
||||
preemptions. Missing usage means cache hits are unknown.
|
||||
- Report `metadata.submission_valid` and any `metadata.submission_invalid_reasons`
|
||||
from the result export. Flag invalid or missing validity instead of treating the
|
||||
run as a valid comparison.
|
||||
- Include request errors, unfinished requests, and warmup failures.
|
||||
Synthetic prompts measure serving performance, not model quality; shortened
|
||||
smoke runs are not benchmark results.
|
||||
Reference in New Issue
Block a user