Files
Emilien Macchianddivesh 8b3cc3fdc0 feat(server): add gateway capacity metrics and optional HPA (#3978)
* feat(server): add gateway capacity metrics and optional HPA

Expose per-replica supervisor sessions, pending relay capacity, relay
rejections and claim latency, and outbound peer request outcomes and
latency. Use bounded labels and Prometheus histograms for the new latency
metrics while preserving existing summary metrics.

Add an optional Helm HPA with external-database and resource validation,
conservative scale-down defaults, and support for custom metrics. Keep
certificate hook pods outside gateway workload selectors.

Document per-pod scraping, scaling limits, upgrade behavior, and
PostgreSQL connection sizing.

Part of #3528

Signed-off-by: Emilien Macchi <emacchi@redhat.com>

* feat(server): unify routing metrics and clarify replica capacity

Combine local relay setup and outbound peer requests in one counter, labeled by operation, route, target, outcome, and gRPC status. Preserve peer latency metrics.

Rename the rejection reason from global_capacity to replica_capacity to reflect the per-replica relay budget. Keep capacity limits unchanged.

Update tests and documentation.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): align routed request metric labels and outcomes

Count a local relay as successful only when its supervisor claims it, the
same event that answers a peer relay on the owner, and rename the outcomes
to local_error and remote_error so they say where an attempt failed. Label
the peer latency histogram by operation, like the routed request counter,
and rename the target label to relay_kind so it does not read as the
Prometheus scrape target.

Rename RelayCapacity.global to per_replica, drop the per-sandbox relay
capacity gauge, which no per-sandbox series can pair with, describe the
latency histogram as peer-only, and restore the note that unavailable
spikes are expected during rollouts.

Part of #3528

Signed-off-by: Emilien Macchi <emacchi@redhat.com>

---------

Signed-off-by: Emilien Macchi <emacchi@redhat.com>
Signed-off-by: divesh <dgude@nvidia.com>
Co-authored-by: divesh <dgude@nvidia.com>
2026-10-05 15:38:40 +00:00
..