Files
sychen52 be74001256 Add concise AgentX benchmark skill (#2573)
### What does this PR do?

  Type of change: Documentation.

Adds a concise AgentX skill covering harness installation, automatic
dataset downloads, benchmark execution, and result reporting. Reuses the
existing deployment skill
  and adds a Claude discovery link.

  ### Usage

Use run-agentx to benchmark my deployed model with a concurrency sweep.

  ### Testing

  • Skill structure and metadata validation passed.
  • Shell syntax checks passed.
• Benchmark arguments parsed and produced a valid configuration using
the pinned harness.
  • All applicable pre-commit checks passed.
  • No GPU benchmark was launched.

  ### Before your PR is "Ready for review"

Contributor guidelines and security practices were reviewed. The commit
is signed and signed off.

  • Is this change backward compatible?: ✅
• If you copied code from other sources or added a new PIP dependency,
did you follow guidance in CONTRIBUTING.md?: N/A. No copied
implementation or project dependency
    changes.

• Did you write any new necessary tests?: N/A. Documentation changes
were validated as described above.
  • Did you update Changelog?: N/A. Skill documentation only.
  • Did you get Claude approval on this PR?: ❌ Not run.

  ### Additional Information


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added guidance for configuring and running SemiAnalysis AgentX serving
benchmarks, including endpoint, model, tokenizer, context limits,
caching, and dataset setup.
* Documented using a pinned benchmark harness in a separate client
environment and running each concurrency level in a fresh artifact
directory with a fixed seed.
* Expanded reporting guidance to cover overlapping requests, cache and
preemption metrics, errors, unfinished requests, warmup failures, and
submission validity.
* Clarified that missing server-reported usage makes cache-hit data
unknown, invalid or missing submission validity should be flagged, and
smoke runs are not benchmark results.
* Directed AgentX agentic workloads from the optional AIPerf guidance to
the AgentX benchmark instructions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
2026-09-29 00:33:53 +00:00
..
2026-09-10 17:30:37 +00:00