Files
OpenShell/proto
0a770d9173 feat(kubernetes): support HA gateway rebalancing (#1868)
* feat(kubernetes): support HA gateway rebalancing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(server): cache peer connections, tokens, and owner lookups

Every forwarded relay rebuilt its setup from scratch: an owner lookup, a
blocking read of the peer token, a TLS connect to the owning replica, and
a TokenReview plus Pod GET on the receiving side. Sandbox service routing
does this per HTTP request, so the apiserver calls scaled with traffic.

Cache all of it on ServerState:

- peer channels pooled per endpoint, so relays multiplex over one
  connection instead of redialing
- peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit
  cannot accept an expired token
- owner records for 3s against a 45s ownership TTL, still freshness
  checked before use

Entries are evicted when a relay fails. Also raise HTTP/2
max_concurrent_streams to 1024, since pooling funnels every relay between
two replicas onto one connection and hyper's default of 200 sits below
the 256 pending-relay budget.

Signed-off-by: divesh <dgude@nvidia.com>

* perf(server): pool upstream connections for sandbox services

Each HTTP request to a sandbox service opened its own supervisor relay,
paying a new TCP connection and HTTP/1 handshake every time. Worse, it
counted against the 32 in-flight relay cap, so a service handling more
than 32 concurrent requests failed outright.

Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is
safe because the pool only returns a connection hyper reports as ready,
and HTTP/1 cannot start a request until the previous body has drained.
Upgrades are never pooled since they take the connection over, and a
failed send evicts that endpoint. Pruning is bounded per key, with the
full sweep limited to once per 30s.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): address HA gateway review findings (#3449)

- Let a gateway own supervisor sessions without a peer endpoint. Requiring
  one whenever the store is PostgreSQL broke every single-instance
  PostgreSQL deployment, because no sandbox supervisor could connect.
  A cross-replica request to an owner that advertises no endpoint now fails
  immediately naming the cause, instead of retrying until the wait timeout.
- Close a supervisor session on heartbeat only when another replica owns it,
  or after renewals fail for the ownership TTL. A database error no longer
  drops every session heartbeating during an outage.
- Clamp owner record ages at zero so a skewed or corrupt stored timestamp
  cannot produce a negative age.
- Bound the cross-object advisory lock with a lock timeout, so a stuck holder
  fails instead of blocking every mutation in the fleet.
- Refuse to start when a peer endpoint is configured on a multi-replica
  backend but peer authentication is unavailable, and warn when a
  multi-replica backend has no peer endpoint at all.
- Reject a plaintext peer endpoint when the gateway serves TLS.
- Skip the sandbox watch poller on single-replica backends, where the local
  update bus already sees every write.
- Rate-limit the peer owner cache sweep so an insert no longer scans the
  whole map under the lock.
- Retry GET and HEAD on a pooled upstream the sandbox closed, instead of
  returning 502, and drop an emptied endpoint from the pool right away.
- Document the gateway peer environment variables and the post-rollout
  ownership skew operators should expect.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): harden HA supervisor ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: divesh <dgude@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: divesh <dgude@nvidia.com>
Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com>
2026-09-21 21:22:12 +00:00
..

Protobuf API conventions

This directory defines OpenShell's gRPC contracts. These conventions are the source of truth for public protobuf API design. Apply them to new APIs and when changing existing APIs; generated SDK naming follows from these definitions.

Entity references

  • Use name for the primary resource targeted by an RPC.
  • Use the resource's role for a related resource reference, such as sandbox, provider, service, or rule.
  • Do not append _name to a canonical entity reference. Its value is already the canonical name. Descriptive values, local map keys, implementation labels, and configured registration names that are not entity references may retain the suffix. Examples include display_name, file_name, driver_name, runtime_class_name, rule_name, and middleware_name.
  • Public callers reference entities by canonical name. Keep immutable IDs at authentication, persistence, compute-driver, and other internal boundaries.

For example:

message GetSandboxRequest {
  openshell.datamodel.v1.WorkspaceSelector workspace_scope = 2;
  string name = 1;
}

message ExposeServiceRequest {
  openshell.datamodel.v1.WorkspaceSelector workspace_scope = 5;
  string sandbox = 1;
  string name = 2;
}

GetSandboxRequest.name identifies the RPC's primary resource. ExposeServiceRequest.sandbox identifies a related sandbox while name identifies the service being exposed.

Workspace scope

  • Public workspace-scoped requests declare workspace_scope before every other field. Existing wire field numbers do not need to follow declaration order.
  • Type workspace_scope as openshell.datamodel.v1.WorkspaceSelector.
  • Requests operating in one workspace require a non-empty canonical WorkspaceSelector.workspace. The default workspace is an explicit name, not an omitted-value fallback.
  • Accept WorkspaceSelector.all_workspaces only on collection-list RPCs that explicitly document and authorize cross-workspace access. The supported public collections are sandboxes, sandbox templates, providers, and services.
  • Requests whose primary resource is a workspace use name, not workspace_scope.
  • Document every request that permits an omitted selector. Current exceptions are platform provider-profile scope and authenticated sandbox bootstrap.

Field and message design

  • Prefer a dedicated request and response message for each RPC, including requests that are currently empty.
  • Use google.protobuf.Timestamp for absolute time and google.protobuf.Duration for elapsed time.
  • Use optional presence when omitted and explicitly empty values have different meanings. Do not infer presence from a protobuf scalar's default value.
  • Keep public request fields in semantic reading order. Field numbers preserve wire identity and may therefore differ from declaration order.
  • Document authorization-sensitive selector variants and omission semantics on the containing request.

Schema evolution

  • Treat field numbers and fully qualified message names as durable wire identities.
  • When removing or renaming a field, reserve its old number and source name. Do not reuse either for a different meaning.
  • Review changes against both the public descriptor closure and durable stored protobuf closure described in the gateway architecture.
  • Regenerate Rust, Python, Go, and TypeScript bindings after contract changes. Run mise run pre-commit, the affected SDK checks, and relevant server tests before submitting the change.