mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-10 03:32:47 +08:00
* feat(kubernetes): support HA gateway rebalancing Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(server): cache peer connections, tokens, and owner lookups Every forwarded relay rebuilt its setup from scratch: an owner lookup, a blocking read of the peer token, a TLS connect to the owning replica, and a TokenReview plus Pod GET on the receiving side. Sandbox service routing does this per HTTP request, so the apiserver calls scaled with traffic. Cache all of it on ServerState: - peer channels pooled per endpoint, so relays multiplex over one connection instead of redialing - peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit cannot accept an expired token - owner records for 3s against a 45s ownership TTL, still freshness checked before use Entries are evicted when a relay fails. Also raise HTTP/2 max_concurrent_streams to 1024, since pooling funnels every relay between two replicas onto one connection and hyper's default of 200 sits below the 256 pending-relay budget. Signed-off-by: divesh <dgude@nvidia.com> * perf(server): pool upstream connections for sandbox services Each HTTP request to a sandbox service opened its own supervisor relay, paying a new TCP connection and HTTP/1 handshake every time. Worse, it counted against the 32 in-flight relay cap, so a service handling more than 32 concurrent requests failed outright. Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is safe because the pool only returns a connection hyper reports as ready, and HTTP/1 cannot start a request until the previous body has drained. Upgrades are never pooled since they take the connection over, and a failed send evicts that endpoint. Pruning is bounded per key, with the full sweep limited to once per 30s. Signed-off-by: divesh <dgude@nvidia.com> * fix(server): address HA gateway review findings (#3449) - Let a gateway own supervisor sessions without a peer endpoint. Requiring one whenever the store is PostgreSQL broke every single-instance PostgreSQL deployment, because no sandbox supervisor could connect. A cross-replica request to an owner that advertises no endpoint now fails immediately naming the cause, instead of retrying until the wait timeout. - Close a supervisor session on heartbeat only when another replica owns it, or after renewals fail for the ownership TTL. A database error no longer drops every session heartbeating during an outage. - Clamp owner record ages at zero so a skewed or corrupt stored timestamp cannot produce a negative age. - Bound the cross-object advisory lock with a lock timeout, so a stuck holder fails instead of blocking every mutation in the fleet. - Refuse to start when a peer endpoint is configured on a multi-replica backend but peer authentication is unavailable, and warn when a multi-replica backend has no peer endpoint at all. - Reject a plaintext peer endpoint when the gateway serves TLS. - Skip the sandbox watch poller on single-replica backends, where the local update bus already sees every write. - Rate-limit the peer owner cache sweep so an insert no longer scans the whole map under the lock. - Retry GET and HEAD on a pooled upstream the sandbox closed, instead of returning 502, and drop an emptied endpoint from the pool right away. - Document the gateway peer environment variables and the post-rollout ownership skew operators should expect. Signed-off-by: divesh <dgude@nvidia.com> * fix(server): harden HA supervisor ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: divesh <dgude@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: divesh <dgude@nvidia.com> Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com>
Protobuf API conventions
This directory defines OpenShell's gRPC contracts. These conventions are the source of truth for public protobuf API design. Apply them to new APIs and when changing existing APIs; generated SDK naming follows from these definitions.
Entity references
- Use
namefor the primary resource targeted by an RPC. - Use the resource's role for a related resource reference, such as
sandbox,provider,service, orrule. - Do not append
_nameto a canonical entity reference. Its value is already the canonical name. Descriptive values, local map keys, implementation labels, and configured registration names that are not entity references may retain the suffix. Examples includedisplay_name,file_name,driver_name,runtime_class_name,rule_name, andmiddleware_name. - Public callers reference entities by canonical name. Keep immutable IDs at authentication, persistence, compute-driver, and other internal boundaries.
For example:
message GetSandboxRequest {
openshell.datamodel.v1.WorkspaceSelector workspace_scope = 2;
string name = 1;
}
message ExposeServiceRequest {
openshell.datamodel.v1.WorkspaceSelector workspace_scope = 5;
string sandbox = 1;
string name = 2;
}
GetSandboxRequest.name identifies the RPC's primary resource.
ExposeServiceRequest.sandbox identifies a related sandbox while name
identifies the service being exposed.
Workspace scope
- Public workspace-scoped requests declare
workspace_scopebefore every other field. Existing wire field numbers do not need to follow declaration order. - Type
workspace_scopeasopenshell.datamodel.v1.WorkspaceSelector. - Requests operating in one workspace require a non-empty canonical
WorkspaceSelector.workspace. Thedefaultworkspace is an explicit name, not an omitted-value fallback. - Accept
WorkspaceSelector.all_workspacesonly on collection-list RPCs that explicitly document and authorize cross-workspace access. The supported public collections are sandboxes, sandbox templates, providers, and services. - Requests whose primary resource is a workspace use
name, notworkspace_scope. - Document every request that permits an omitted selector. Current exceptions are platform provider-profile scope and authenticated sandbox bootstrap.
Field and message design
- Prefer a dedicated request and response message for each RPC, including requests that are currently empty.
- Use
google.protobuf.Timestampfor absolute time andgoogle.protobuf.Durationfor elapsed time. - Use optional presence when omitted and explicitly empty values have different meanings. Do not infer presence from a protobuf scalar's default value.
- Keep public request fields in semantic reading order. Field numbers preserve wire identity and may therefore differ from declaration order.
- Document authorization-sensitive selector variants and omission semantics on the containing request.
Schema evolution
- Treat field numbers and fully qualified message names as durable wire identities.
- When removing or renaming a field, reserve its old number and source name. Do not reuse either for a different meaning.
- Review changes against both the public descriptor closure and durable stored protobuf closure described in the gateway architecture.
- Regenerate Rust, Python, Go, and TypeScript bindings after contract changes.
Run
mise run pre-commit, the affected SDK checks, and relevant server tests before submitting the change.