mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-02 07:34:45 +08:00
Count a local relay as successful only when its supervisor claims it, the same event that answers a peer relay on the owner, and rename the outcomes to local_error and remote_error so they say where an attempt failed. Label the peer latency histogram by operation, like the routed request counter, and rename the target label to relay_kind so it does not read as the Prometheus scrape target. Rename RelayCapacity.global to per_replica, drop the per-sandbox relay capacity gauge, which no per-sandbox series can pair with, describe the latency histogram as peer-only, and restore the note that unavailable spikes are expected during rollouts. Part of #3528 Signed-off-by: Emilien Macchi <emacchi@redhat.com>