mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-02 07:34:45 +08:00
Expose per-replica supervisor sessions, pending relay capacity, relay rejections and claim latency, and outbound peer request outcomes and latency. Use bounded labels and Prometheus histograms for the new latency metrics while preserving existing summary metrics. Add an optional Helm HPA with external-database and resource validation, conservative scale-down defaults, and support for custom metrics. Keep certificate hook pods outside gateway workload selectors. Document per-pod scraping, scaling limits, upgrade behavior, and PostgreSQL connection sizing. Part of #3528 Signed-off-by: Emilien Macchi <emacchi@redhat.com>