Files
OpenShell/crates/openshell-sdk
0a770d9173 feat(kubernetes): support HA gateway rebalancing (#1868)
* feat(kubernetes): support HA gateway rebalancing

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(server): cache peer connections, tokens, and owner lookups

Every forwarded relay rebuilt its setup from scratch: an owner lookup, a
blocking read of the peer token, a TLS connect to the owning replica, and
a TokenReview plus Pod GET on the receiving side. Sandbox service routing
does this per HTTP request, so the apiserver calls scaled with traffic.

Cache all of it on ServerState:

- peer channels pooled per endpoint, so relays multiplex over one
  connection instead of redialing
- peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit
  cannot accept an expired token
- owner records for 3s against a 45s ownership TTL, still freshness
  checked before use

Entries are evicted when a relay fails. Also raise HTTP/2
max_concurrent_streams to 1024, since pooling funnels every relay between
two replicas onto one connection and hyper's default of 200 sits below
the 256 pending-relay budget.

Signed-off-by: divesh <dgude@nvidia.com>

* perf(server): pool upstream connections for sandbox services

Each HTTP request to a sandbox service opened its own supervisor relay,
paying a new TCP connection and HTTP/1 handshake every time. Worse, it
counted against the 32 in-flight relay cap, so a service handling more
than 32 concurrent requests failed outright.

Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is
safe because the pool only returns a connection hyper reports as ready,
and HTTP/1 cannot start a request until the previous body has drained.
Upgrades are never pooled since they take the connection over, and a
failed send evicts that endpoint. Pruning is bounded per key, with the
full sweep limited to once per 30s.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): address HA gateway review findings (#3449)

- Let a gateway own supervisor sessions without a peer endpoint. Requiring
  one whenever the store is PostgreSQL broke every single-instance
  PostgreSQL deployment, because no sandbox supervisor could connect.
  A cross-replica request to an owner that advertises no endpoint now fails
  immediately naming the cause, instead of retrying until the wait timeout.
- Close a supervisor session on heartbeat only when another replica owns it,
  or after renewals fail for the ownership TTL. A database error no longer
  drops every session heartbeating during an outage.
- Clamp owner record ages at zero so a skewed or corrupt stored timestamp
  cannot produce a negative age.
- Bound the cross-object advisory lock with a lock timeout, so a stuck holder
  fails instead of blocking every mutation in the fleet.
- Refuse to start when a peer endpoint is configured on a multi-replica
  backend but peer authentication is unavailable, and warn when a
  multi-replica backend has no peer endpoint at all.
- Reject a plaintext peer endpoint when the gateway serves TLS.
- Skip the sandbox watch poller on single-replica backends, where the local
  update bus already sees every write.
- Rate-limit the peer owner cache sweep so an insert no longer scans the
  whole map under the lock.
- Retry GET and HEAD on a pooled upstream the sandbox closed, instead of
  returning 502, and drop an emptied endpoint from the pool right away.
- Document the gateway peer environment variables and the post-rollout
  ownership skew operators should expect.

Signed-off-by: divesh <dgude@nvidia.com>

* fix(server): harden HA supervisor ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: divesh <dgude@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: divesh <dgude@nvidia.com>
Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com>
2026-09-21 21:22:12 +00:00
..

openshell-sdk

openshell-sdk is the shared async Rust client for OpenShell gateways. It owns gRPC channel setup, TLS, OIDC refresh, and the Cloudflare Access tunnel so the CLI, the TUI, and language bindings share one client implementation. Callers pass an explicit bearer token; the SDK does no filesystem access and no gateway-name resolution.

Two layers

  • OpenShellClient — the curated, sandbox-focused surface: health, sandbox CRUD, reusable sandbox template CRUD, readiness/deletion waits, and non-streaming exec.
  • raw — direct access to the generated tonic clients for RPCs the curated surface doesn't yet cover (providers, policy, logs, settings, SSH, forwarding).

Auth and refresh

The curated surface drives OIDC refresh automatically: proactively before a request and reactively on Unauthenticated. Refreshes are single-flight, so only one is in flight at a time.

The plain raw_grpc accessor does not refresh; it returns a client bound to the current token. When a refresher is wired, use raw_grpc_fresh to refresh before the call, and force_refresh to recover after a raw RPC returns Unauthenticated.

The SDK consumes a Refresh trait that the caller implements; it does not run the OIDC browser flow itself. Its non-interactive refresh-token exchange also accepts scopes that identity providers may require to select the API resource for the refreshed access token.

Transport modes

  • Plaintext (local development)
  • Server-authenticated TLS (system roots, or a pinned private CA via ca_cert)
  • OIDC bearer over HTTPS (gateways behind an OAuth2/OIDC IdP)
  • Cloudflare Access tunnel (hosted gateways)
  • Insecure TLS (development/debug; certificate verification disabled)

mTLS (client certificates) is not supported.

Public surface

OpenShellClient::connect(ClientConfig) returns a connected client exposing health, create_sandbox, get_sandbox, list_sandboxes, list_all_sandboxes, delete_sandbox, create_sandbox_from_template, create_sandbox_template, get_sandbox_template, list_sandbox_templates, delete_sandbox_template, list_sandboxes_all_workspaces, list_sandbox_templates_all_workspaces, wait_ready, wait_deleted, and exec. Curated types (SandboxSpec, SandboxRef, Health, ListOptions, SandboxTemplateListOptions, ExecOptions, SandboxPhase) use SDK-shaped enums rather than raw proto integers where practical. Reusable template resources are exposed as SandboxWorkloadTemplate proto aliases so callers can populate the full portable workload shape and driver config. Failures map to a typed SdkError with a discriminable kind.

Set SandboxSpec::service_exposures to register named or unnamed loopback HTTP services during creation. Each ServiceExposure contains a service name and a target port; an empty name selects the unnamed endpoint. The returned SandboxRef::service_urls map contains each routed URL under the same name.

Curated calls without a workspace argument explicitly select the default workspace. Cross-workspace listing uses the separate *_all_workspaces methods and requires Platform Admin access.

For an accepted sandbox deletion, pass its original ID to wait_deleted so a same-name replacement does not extend the wait. Both the default and workspace-scoped clients accept the optional third argument; pass None to wait for name absence instead.

let deletion = client.delete_sandbox(name, openshell_sdk::DeleteOptions::default()).await?;
if deletion.outcome == openshell_sdk::DeletionOutcome::Accepted {
    client.wait_deleted(
        name,
        std::time::Duration::from_secs(60),
        deletion.sandbox_id.as_deref(),
    ).await?;
}

Curated list_* methods return a lazy Pager<T>. Each next_page() call issues at most one RPC and returns a Page<T> with its opaque continuation token. The explicit list_all_* conveniences exhaust that pager; page_size always controls one gateway request, and page_token resumes a saved traversal.

let mut pages = client.list_sandboxes(ListOptions {
    page_size: 100,
    ..Default::default()
});
while let Some(page) = pages.next_page().await? {
    for sandbox in page.items {
        println!("{}", sandbox.name);
    }
}
use openshell_sdk::{
    ClientConfig, OpenShellClient, SandboxTemplateCreateSpec,
    SandboxWorkloadConfig, SandboxWorkloadTemplate, SandboxWorkloadTemplateSpec,
};

# async fn run() -> Result<(), openshell_sdk::SdkError> {
let client = OpenShellClient::connect(ClientConfig::new("http://127.0.0.1:8080")).await?;
client
    .create_sandbox_template(SandboxWorkloadTemplate {
        metadata: Some(openshell_sdk::raw::proto::datamodel::v1::ObjectMeta {
            name: "python".to_string(),
            ..Default::default()
        }),
        spec: Some(SandboxWorkloadTemplateSpec {
            workload: Some(SandboxWorkloadConfig {
                image: "ghcr.io/nvidia/openshell-community/sandboxes/python:latest".to_string(),
                ..Default::default()
            }),
            ..Default::default()
        }),
    })
    .await?;

let _sandbox = client
    .create_sandbox_from_template(SandboxTemplateCreateSpec {
        template_name: "python".to_string(),
        policy: Some(openshell_sdk::raw::proto::SandboxPolicy {
            version: 1,
            ..Default::default()
        }),
        ..Default::default()
    })
    .await?;
# Ok(())
# }

Wait for a provider change

Provider attach, detach, and update responses include a ProviderMutationReceipt: a saved record identifying the exact change requested for one sandbox. Pass that record to provider_readiness::wait_for_provider to wait until the current sandbox runtime confirms it applied the change. Detach completes with Revoked; attach and update complete with Ready.

use std::time::Duration;
use openshell_sdk::{OpenShellClient, raw::ProviderMutationReceipt};
use openshell_sdk::provider_readiness::{
    ProviderWaitOutcome, wait_for_provider,
};

async fn wait_for_change(
    client: &OpenShellClient,
    change: &ProviderMutationReceipt,
) -> Result<(), Box<dyn std::error::Error>> {
    let mut grpc = client.raw_grpc_fresh().await?;
    let result = wait_for_provider(&mut grpc, change, Duration::from_secs(30)).await?;
    match result.outcome {
        ProviderWaitOutcome::Complete => println!("The sandbox applied the change."),
        ProviderWaitOutcome::TimedOut => println!("Still waiting; check the same change again."),
        ProviderWaitOutcome::Terminal => println!("The change failed, was withheld, or was replaced."),
    }
    Ok(())
}

The result preserves the last known status when the deadline expires. A later change cannot satisfy a wait for the original request. provider_status queries once, and wait_for_provider_until accepts a shared deadline for waiting on the sandboxes selected by one provider update. Status responses contain configuration identities and safe reason categories, without credentials or raw installation errors.

For ordinary static credentials, launch a new client after update readiness to receive the updated reference. Existing processes keep their revision-scoped references; a successful wait does not retarget them or prove that the old upstream key can be retired. After detach completes, retained references cannot resolve and new processes do not receive them.

The status's operation field is the common operation's historical outcome, keyed by the receipt ID. These helpers complete from the live provider state and its matching evidence; a historical applied operation cannot override a disconnected, expired, or superseded live result.

These helpers use the raw client's authentication slot. They do not perform OIDC refresh themselves. Follow the raw-client refresh guidance above if a request returns Unauthenticated, then resume waiting for the same change ID.

Modules

Module Purpose
client High-level OpenShellClient and the curated sandbox surface.
config ClientConfig, AuthConfig.
transport Channel construction, TLS resolution, request interceptors.
auth EdgeAuthInterceptor for bearer-token attachment.
oidc OIDC token handling at the transport layer.
refresh Refresh trait and single-flight refresh coalescing.
edge_tunnel Cloudflare Access tunnel dialer.
error SdkError taxonomy.
pagination Lazy Pager<T> and response Page<T>.
types Curated request/response types and proto conversions.
raw Escape hatch re-exporting the generated tonic clients.
provider_readiness Check and wait for an exact provider change to take effect.

Notes

  • Async-only. Tonic is async-native; callers needing a blocking call can wrap with their own runtime.
  • The curated surface will grow as more RPCs graduate from raw.