Files
Piotr Mlocek 9b5bcdd7a2 fix(ha): keep sandboxes ready across gateway pod rolls (#4321)
* fix(ha): keep sandboxes ready across gateway pod rolls

A gateway pod roll left sandboxes not ready long enough for clients to
fail with "sandbox is not ready", which made the HA e2e test
sandbox_file_sync_survives_gateway_pod_rolls flaky. Two causes:

- The supervisor never reset its reconnect backoff after an accepted
  session, so each gateway restart over a sandbox's lifetime doubled the
  delay until every reconnect waited the full 30s maximum. Reset it
  after an accepted session.
- A stopping gateway replica redirected its supervisors and then
  immediately demoted their sandboxes to Provisioning, before the
  redirected supervisor published its replacement session. Wait up to
  5s for the replacement before demoting; the existing disconnect path
  keeps the sandbox Ready once a replacement exists.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ha): wait for handoff only when a redirect went out

Address review feedback on the shutdown handoff:

- Wait for a replacement session only when the stopping replica queued a
  SessionRedirect. Single-replica gateways, supervisors without redirect
  support, and full outbound channels no longer pay the grace period.
- Close the session stream in both directions before cleanup so a
  supervisor without a redirect sees EOF and reconnects at once.
- Bound the replacement wait with timeout_at and the shared session-wait
  backoff, so the grace is a hard ceiling.
- Name the supervisor session shutdown timeout and assert at compile
  time that it leaves room after the grace period.
- Demote the sandbox when session setup fails after the owner publish,
  so a failed takeover cannot leave it Ready with no session.
- Move the supervisor reconnect delays into ReconnectBackoff so both
  loop branches reset the same way, and test reconnect sequences.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ha): refresh gateway membership before shutdown redirects

A stopping replica redirected its supervisors using its cached ring,
which the membership worker refreshes only every 10s. During a rolling
update the replacement for the previous pod is often only seconds old,
so the cached ring missed it and still named the departed pod. The
redirect then went nowhere or to a dead pod, whose 10s connect timeout
outlasted the handoff grace, and the sandbox was demoted mid-roll.

Re-read live membership, bounded to 2s, just before session shutdown
computes redirects.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* fix(ha): hold ownership through the shutdown handoff

The stopping replica released its owner record before waiting for the
replacement session. A reconciliation sweep on another replica in that
window saw no owner, and with the Kubernetes driver leaving readiness
to the gateway it demoted the sandbox to Provisioning anyway. Keep the
record until a different session supersedes it or the grace expires,
then release it only if it is still ours.

Also from review:
- Keep polling the owner record through store errors until the deadline.
- Skip the handoff wait for sandboxes that are not Ready.
- Re-read membership in the membership worker's shutdown branch and
  await the worker before session shutdown, so the worker stays the only
  writer of the ring and peer map.
- Share the remove, release, and demote steps of failed session setup.
- Test redirect queuing directly, including unsupported supervisors and
  a full outbound channel.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
2026-10-08 05:18:10 +00:00
..