mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: Documentation. Adds a concise AgentX skill covering harness installation, automatic dataset downloads, benchmark execution, and result reporting. Reuses the existing deployment skill and adds a Claude discovery link. ### Usage Use run-agentx to benchmark my deployed model with a concurrency sweep. ### Testing • Skill structure and metadata validation passed. • Shell syntax checks passed. • Benchmark arguments parsed and produced a valid configuration using the pinned harness. • All applicable pre-commit checks passed. • No GPU benchmark was launched. ### Before your PR is "Ready for review" Contributor guidelines and security practices were reviewed. The commit is signed and signed off. • Is this change backward compatible?: ✅ • If you copied code from other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md?: N/A. No copied implementation or project dependency changes. • Did you write any new necessary tests?: N/A. Documentation changes were validated as described above. • Did you update Changelog?: N/A. Skill documentation only. • Did you get Claude approval on this PR?: ❌ Not run. ### Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added guidance for configuring and running SemiAnalysis AgentX serving benchmarks, including endpoint, model, tokenizer, context limits, caching, and dataset setup. * Documented using a pinned benchmark harness in a separate client environment and running each concurrency level in a fresh artifact directory with a fixed seed. * Expanded reporting guidance to cover overlapping requests, cache and preemption metrics, errors, unfinished requests, warmup failures, and submission validity. * Clarified that missing server-reported usage makes cache-hit data unknown, invalid or missing submission validity should be flagged, and smoke runs are not benchmark results. * Directed AgentX agentic workloads from the optional AIPerf guidance to the AgentX benchmark instructions. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Shiyang Chen <shiychen@nvidia.com>