* feat(kubernetes): support HA gateway rebalancing Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(server): cache peer connections, tokens, and owner lookups Every forwarded relay rebuilt its setup from scratch: an owner lookup, a blocking read of the peer token, a TLS connect to the owning replica, and a TokenReview plus Pod GET on the receiving side. Sandbox service routing does this per HTTP request, so the apiserver calls scaled with traffic. Cache all of it on ServerState: - peer channels pooled per endpoint, so relays multiplex over one connection instead of redialing - peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit cannot accept an expired token - owner records for 3s against a 45s ownership TTL, still freshness checked before use Entries are evicted when a relay fails. Also raise HTTP/2 max_concurrent_streams to 1024, since pooling funnels every relay between two replicas onto one connection and hyper's default of 200 sits below the 256 pending-relay budget. Signed-off-by: divesh <dgude@nvidia.com> * perf(server): pool upstream connections for sandbox services Each HTTP request to a sandbox service opened its own supervisor relay, paying a new TCP connection and HTTP/1 handshake every time. Worse, it counted against the 32 in-flight relay cap, so a service handling more than 32 concurrent requests failed outright. Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is safe because the pool only returns a connection hyper reports as ready, and HTTP/1 cannot start a request until the previous body has drained. Upgrades are never pooled since they take the connection over, and a failed send evicts that endpoint. Pruning is bounded per key, with the full sweep limited to once per 30s. Signed-off-by: divesh <dgude@nvidia.com> * fix(server): address HA gateway review findings (#3449) - Let a gateway own supervisor sessions without a peer endpoint. Requiring one whenever the store is PostgreSQL broke every single-instance PostgreSQL deployment, because no sandbox supervisor could connect. A cross-replica request to an owner that advertises no endpoint now fails immediately naming the cause, instead of retrying until the wait timeout. - Close a supervisor session on heartbeat only when another replica owns it, or after renewals fail for the ownership TTL. A database error no longer drops every session heartbeating during an outage. - Clamp owner record ages at zero so a skewed or corrupt stored timestamp cannot produce a negative age. - Bound the cross-object advisory lock with a lock timeout, so a stuck holder fails instead of blocking every mutation in the fleet. - Refuse to start when a peer endpoint is configured on a multi-replica backend but peer authentication is unavailable, and warn when a multi-replica backend has no peer endpoint at all. - Reject a plaintext peer endpoint when the gateway serves TLS. - Skip the sandbox watch poller on single-replica backends, where the local update bus already sees every write. - Rate-limit the peer owner cache sweep so an insert no longer scans the whole map under the lock. - Retry GET and HEAD on a pooled upstream the sandbox closed, instead of returning 502, and drop an emptied endpoint from the pool right away. - Document the gateway peer environment variables and the post-rollout ownership skew operators should expect. Signed-off-by: divesh <dgude@nvidia.com> * fix(server): harden HA supervisor ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: divesh <dgude@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: divesh <dgude@nvidia.com> Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com>
openshell-sdk
openshell-sdk is the shared async Rust client for OpenShell gateways. It owns
gRPC channel setup, TLS, OIDC refresh, and the Cloudflare Access tunnel so the
CLI, the TUI, and language bindings share one client implementation. Callers
pass an explicit bearer token; the SDK does no filesystem access and no
gateway-name resolution.
Two layers
OpenShellClient— the curated, sandbox-focused surface: health, sandbox CRUD, reusable sandbox template CRUD, readiness/deletion waits, and non-streaming exec.raw— direct access to the generated tonic clients for RPCs the curated surface doesn't yet cover (providers, policy, logs, settings, SSH, forwarding).
Auth and refresh
The curated surface drives OIDC refresh automatically: proactively before a
request and reactively on Unauthenticated. Refreshes are single-flight, so
only one is in flight at a time.
The plain raw_grpc accessor does not refresh; it returns a client bound to
the current token. When a refresher is wired, use raw_grpc_fresh to refresh before the call, and
force_refresh to recover after a raw RPC returns Unauthenticated.
The SDK consumes a Refresh trait that the caller implements; it does not run
the OIDC browser flow itself. Its non-interactive refresh-token exchange also
accepts scopes that identity providers may require to select the API resource
for the refreshed access token.
Transport modes
- Plaintext (local development)
- Server-authenticated TLS (system roots, or a pinned private CA via
ca_cert) - OIDC bearer over HTTPS (gateways behind an OAuth2/OIDC IdP)
- Cloudflare Access tunnel (hosted gateways)
- Insecure TLS (development/debug; certificate verification disabled)
mTLS (client certificates) is not supported.
Public surface
OpenShellClient::connect(ClientConfig) returns a connected client exposing
health, create_sandbox, get_sandbox, list_sandboxes, list_all_sandboxes, delete_sandbox,
create_sandbox_from_template, create_sandbox_template,
get_sandbox_template, list_sandbox_templates, delete_sandbox_template,
list_sandboxes_all_workspaces, list_sandbox_templates_all_workspaces,
wait_ready, wait_deleted, and exec. Curated types (SandboxSpec,
SandboxRef, Health, ListOptions, SandboxTemplateListOptions,
ExecOptions, SandboxPhase) use SDK-shaped enums rather than raw proto
integers where practical. Reusable template resources are exposed as
SandboxWorkloadTemplate proto aliases so callers can populate the full
portable workload shape and driver config. Failures map to a typed SdkError
with a discriminable kind.
Set SandboxSpec::service_exposures to register named or unnamed loopback HTTP
services during creation. Each ServiceExposure contains a service name and a
target port; an empty name selects the unnamed endpoint. The returned
SandboxRef::service_urls map contains each routed URL under the same name.
Curated calls without a workspace argument explicitly select the default
workspace. Cross-workspace listing uses the separate *_all_workspaces
methods and requires Platform Admin access.
For an accepted sandbox deletion, pass its original ID to wait_deleted so a
same-name replacement does not extend the wait. Both the default and
workspace-scoped clients accept the optional third argument; pass None to wait
for name absence instead.
let deletion = client.delete_sandbox(name, openshell_sdk::DeleteOptions::default()).await?;
if deletion.outcome == openshell_sdk::DeletionOutcome::Accepted {
client.wait_deleted(
name,
std::time::Duration::from_secs(60),
deletion.sandbox_id.as_deref(),
).await?;
}
Curated list_* methods return a lazy Pager<T>. Each next_page() call
issues at most one RPC and returns a Page<T> with its opaque continuation
token. The explicit list_all_* conveniences exhaust that pager; page_size
always controls one gateway request, and page_token resumes a saved traversal.
let mut pages = client.list_sandboxes(ListOptions {
page_size: 100,
..Default::default()
});
while let Some(page) = pages.next_page().await? {
for sandbox in page.items {
println!("{}", sandbox.name);
}
}
use openshell_sdk::{
ClientConfig, OpenShellClient, SandboxTemplateCreateSpec,
SandboxWorkloadConfig, SandboxWorkloadTemplate, SandboxWorkloadTemplateSpec,
};
# async fn run() -> Result<(), openshell_sdk::SdkError> {
let client = OpenShellClient::connect(ClientConfig::new("http://127.0.0.1:8080")).await?;
client
.create_sandbox_template(SandboxWorkloadTemplate {
metadata: Some(openshell_sdk::raw::proto::datamodel::v1::ObjectMeta {
name: "python".to_string(),
..Default::default()
}),
spec: Some(SandboxWorkloadTemplateSpec {
workload: Some(SandboxWorkloadConfig {
image: "ghcr.io/nvidia/openshell-community/sandboxes/python:latest".to_string(),
..Default::default()
}),
..Default::default()
}),
})
.await?;
let _sandbox = client
.create_sandbox_from_template(SandboxTemplateCreateSpec {
template_name: "python".to_string(),
policy: Some(openshell_sdk::raw::proto::SandboxPolicy {
version: 1,
..Default::default()
}),
..Default::default()
})
.await?;
# Ok(())
# }
Wait for a provider change
Provider attach, detach, and update responses include a ProviderMutationReceipt: a saved record identifying the exact change requested for one sandbox. Pass that record to provider_readiness::wait_for_provider to wait until the current sandbox runtime confirms it applied the change. Detach completes with Revoked; attach and update complete with Ready.
use std::time::Duration;
use openshell_sdk::{OpenShellClient, raw::ProviderMutationReceipt};
use openshell_sdk::provider_readiness::{
ProviderWaitOutcome, wait_for_provider,
};
async fn wait_for_change(
client: &OpenShellClient,
change: &ProviderMutationReceipt,
) -> Result<(), Box<dyn std::error::Error>> {
let mut grpc = client.raw_grpc_fresh().await?;
let result = wait_for_provider(&mut grpc, change, Duration::from_secs(30)).await?;
match result.outcome {
ProviderWaitOutcome::Complete => println!("The sandbox applied the change."),
ProviderWaitOutcome::TimedOut => println!("Still waiting; check the same change again."),
ProviderWaitOutcome::Terminal => println!("The change failed, was withheld, or was replaced."),
}
Ok(())
}
The result preserves the last known status when the deadline expires. A later change cannot satisfy a wait for the original request. provider_status queries once, and wait_for_provider_until accepts a shared deadline for waiting on the sandboxes selected by one provider update. Status responses contain configuration identities and safe reason categories, without credentials or raw installation errors.
For ordinary static credentials, launch a new client after update readiness to receive the updated reference. Existing processes keep their revision-scoped references; a successful wait does not retarget them or prove that the old upstream key can be retired. After detach completes, retained references cannot resolve and new processes do not receive them.
The status's operation field is the common operation's historical outcome, keyed by the receipt ID. These helpers complete from the live provider state and its matching evidence; a historical applied operation cannot override a disconnected, expired, or superseded live result.
These helpers use the raw client's authentication slot. They do not perform OIDC refresh themselves. Follow the raw-client refresh guidance above if a request returns Unauthenticated, then resume waiting for the same change ID.
Modules
| Module | Purpose |
|---|---|
client |
High-level OpenShellClient and the curated sandbox surface. |
config |
ClientConfig, AuthConfig. |
transport |
Channel construction, TLS resolution, request interceptors. |
auth |
EdgeAuthInterceptor for bearer-token attachment. |
oidc |
OIDC token handling at the transport layer. |
refresh |
Refresh trait and single-flight refresh coalescing. |
edge_tunnel |
Cloudflare Access tunnel dialer. |
error |
SdkError taxonomy. |
pagination |
Lazy Pager<T> and response Page<T>. |
types |
Curated request/response types and proto conversions. |
raw |
Escape hatch re-exporting the generated tonic clients. |
provider_readiness |
Check and wait for an exact provider change to take effect. |
Notes
- Async-only. Tonic is async-native; callers needing a blocking call can wrap with their own runtime.
- The curated surface will grow as more RPCs graduate from
raw.