mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-06 09:22:34 +08:00
* feat(server): add gateway capacity metrics and optional HPA Expose per-replica supervisor sessions, pending relay capacity, relay rejections and claim latency, and outbound peer request outcomes and latency. Use bounded labels and Prometheus histograms for the new latency metrics while preserving existing summary metrics. Add an optional Helm HPA with external-database and resource validation, conservative scale-down defaults, and support for custom metrics. Keep certificate hook pods outside gateway workload selectors. Document per-pod scraping, scaling limits, upgrade behavior, and PostgreSQL connection sizing. Part of #3528 Signed-off-by: Emilien Macchi <emacchi@redhat.com> * feat(server): unify routing metrics and clarify replica capacity Combine local relay setup and outbound peer requests in one counter, labeled by operation, route, target, outcome, and gRPC status. Preserve peer latency metrics. Rename the rejection reason from global_capacity to replica_capacity to reflect the per-replica relay budget. Keep capacity limits unchanged. Update tests and documentation. Signed-off-by: divesh <dgude@nvidia.com> * fix(server): align routed request metric labels and outcomes Count a local relay as successful only when its supervisor claims it, the same event that answers a peer relay on the owner, and rename the outcomes to local_error and remote_error so they say where an attempt failed. Label the peer latency histogram by operation, like the routed request counter, and rename the target label to relay_kind so it does not read as the Prometheus scrape target. Rename RelayCapacity.global to per_replica, drop the per-sandbox relay capacity gauge, which no per-sandbox series can pair with, describe the latency histogram as peer-only, and restore the note that unavailable spikes are expected during rollouts. Part of #3528 Signed-off-by: Emilien Macchi <emacchi@redhat.com> --------- Signed-off-by: Emilien Macchi <emacchi@redhat.com> Signed-off-by: divesh <dgude@nvidia.com> Co-authored-by: divesh <dgude@nvidia.com>