mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-10 11:44:45 +08:00
dev
39
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4c1b16a4a1 |
feat(server): add operator-only provider credential retrieval (#4357)
* feat(server): add operator-only provider credential retrieval Authenticate operators exclusively through verified direct gateway mTLS. Export selected runtime credentials with freshness validation and all-or-nothing delivery. Coordinate refresh across callers and replicas, preserve rotated grants across RPC cancellation, and fence credential delivery against reconfiguration. Add generated bindings, security and regression tests, and operator documentation. Signed-off-by: Seth Jennings <sjenning@redhat.com> * feat(sdk): add curated operator credential retrieval Expose provider credential clients in Python and TypeScript with explicit workspace and key selection, freshness options, optional expiration, and preserved gateway errors. Reuse existing mTLS transports without introducing retries or secret caching. Add generated-wire and error-mapping tests, SDK package validation, and operator usage documentation. Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(server): honor recovery state for queued automatic refreshes Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
8b3cc3fdc0 |
feat(server): add gateway capacity metrics and optional HPA (#3978)
* feat(server): add gateway capacity metrics and optional HPA Expose per-replica supervisor sessions, pending relay capacity, relay rejections and claim latency, and outbound peer request outcomes and latency. Use bounded labels and Prometheus histograms for the new latency metrics while preserving existing summary metrics. Add an optional Helm HPA with external-database and resource validation, conservative scale-down defaults, and support for custom metrics. Keep certificate hook pods outside gateway workload selectors. Document per-pod scraping, scaling limits, upgrade behavior, and PostgreSQL connection sizing. Part of #3528 Signed-off-by: Emilien Macchi <emacchi@redhat.com> * feat(server): unify routing metrics and clarify replica capacity Combine local relay setup and outbound peer requests in one counter, labeled by operation, route, target, outcome, and gRPC status. Preserve peer latency metrics. Rename the rejection reason from global_capacity to replica_capacity to reflect the per-replica relay budget. Keep capacity limits unchanged. Update tests and documentation. Signed-off-by: divesh <dgude@nvidia.com> * fix(server): align routed request metric labels and outcomes Count a local relay as successful only when its supervisor claims it, the same event that answers a peer relay on the owner, and rename the outcomes to local_error and remote_error so they say where an attempt failed. Label the peer latency histogram by operation, like the routed request counter, and rename the target label to relay_kind so it does not read as the Prometheus scrape target. Rename RelayCapacity.global to per_replica, drop the per-sandbox relay capacity gauge, which no per-sandbox series can pair with, describe the latency histogram as peer-only, and restore the note that unavailable spikes are expected during rollouts. Part of #3528 Signed-off-by: Emilien Macchi <emacchi@redhat.com> --------- Signed-off-by: Emilien Macchi <emacchi@redhat.com> Signed-off-by: divesh <dgude@nvidia.com> Co-authored-by: divesh <dgude@nvidia.com> |
||
|
|
0a770d9173 |
feat(kubernetes): support HA gateway rebalancing (#1868)
* feat(kubernetes): support HA gateway rebalancing Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(server): cache peer connections, tokens, and owner lookups Every forwarded relay rebuilt its setup from scratch: an owner lookup, a blocking read of the peer token, a TLS connect to the owning replica, and a TokenReview plus Pod GET on the receiving side. Sandbox service routing does this per HTTP request, so the apiserver calls scaled with traffic. Cache all of it on ServerState: - peer channels pooled per endpoint, so relays multiplex over one connection instead of redialing - peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit cannot accept an expired token - owner records for 3s against a 45s ownership TTL, still freshness checked before use Entries are evicted when a relay fails. Also raise HTTP/2 max_concurrent_streams to 1024, since pooling funnels every relay between two replicas onto one connection and hyper's default of 200 sits below the 256 pending-relay budget. Signed-off-by: divesh <dgude@nvidia.com> * perf(server): pool upstream connections for sandbox services Each HTTP request to a sandbox service opened its own supervisor relay, paying a new TCP connection and HTTP/1 handshake every time. Worse, it counted against the 32 in-flight relay cap, so a service handling more than 32 concurrent requests failed outright. Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is safe because the pool only returns a connection hyper reports as ready, and HTTP/1 cannot start a request until the previous body has drained. Upgrades are never pooled since they take the connection over, and a failed send evicts that endpoint. Pruning is bounded per key, with the full sweep limited to once per 30s. Signed-off-by: divesh <dgude@nvidia.com> * fix(server): address HA gateway review findings (#3449) - Let a gateway own supervisor sessions without a peer endpoint. Requiring one whenever the store is PostgreSQL broke every single-instance PostgreSQL deployment, because no sandbox supervisor could connect. A cross-replica request to an owner that advertises no endpoint now fails immediately naming the cause, instead of retrying until the wait timeout. - Close a supervisor session on heartbeat only when another replica owns it, or after renewals fail for the ownership TTL. A database error no longer drops every session heartbeating during an outage. - Clamp owner record ages at zero so a skewed or corrupt stored timestamp cannot produce a negative age. - Bound the cross-object advisory lock with a lock timeout, so a stuck holder fails instead of blocking every mutation in the fleet. - Refuse to start when a peer endpoint is configured on a multi-replica backend but peer authentication is unavailable, and warn when a multi-replica backend has no peer endpoint at all. - Reject a plaintext peer endpoint when the gateway serves TLS. - Skip the sandbox watch poller on single-replica backends, where the local update bus already sees every write. - Rate-limit the peer owner cache sweep so an insert no longer scans the whole map under the lock. - Retry GET and HEAD on a pooled upstream the sandbox closed, instead of returning 502, and drop an emptied endpoint from the pool right away. - Document the gateway peer environment variables and the post-rollout ownership skew operators should expect. Signed-off-by: divesh <dgude@nvidia.com> * fix(server): harden HA supervisor ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: divesh <dgude@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: divesh <dgude@nvidia.com> Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com> |
||
|
|
1d010f4187 |
feat(sandbox): validate configuration before workload activation (#3259)
* feat(sandbox): validate configuration before workload activation Closes #3145 Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): bound startup failures and preserve activation history Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): enforce deadlines on startup RPC attempts Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): isolate provider auto-create policy fixture Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): remove unnecessary fixture string delimiters Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): capture startup logs before ephemeral cleanup Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): preserve credential revocation after rebase Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): retain provider revision helper for endpoint reports Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): import provider object trait in production Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(supervisor): preserve admission across boundary extraction Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): use existing boundary discovery imports Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): distinguish workload and supervisor containers Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): reconcile admission schema with timestamp migration Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): reconcile admission with typed deletion schema Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): reconcile admission with mutation request IDs Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(supervisor): adapt local startup fixture to admission state Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(supervisor): deduplicate startup quarantine diagnostics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(tui): show configuration blockers in sandbox notes Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(sandbox): expire provisioning repair attempts after five minutes Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(supervisor): reconcile admission with provider readiness Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(tui): separate configuration summaries from full diagnostics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(tui): shorten invalid configuration note Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): refresh schema inventory after rebase Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
9708ba9999 |
feat(providers): report applied sandbox provider changes (#3391)
* feat(providers): report applied sandbox provider changes Record exact provider mutation targets in shared configuration operations. Require authenticated evidence that credentials, effective policy, and the workload launch environment have been installed before reporting readiness. Add bounded CLI and Rust SDK status and wait support, preserving ordinary revision-scoped references for existing processes. Verify new-client rotation and acknowledged detach revocation without external-stable resolver changes. Signed-off-by: Shiju <shiju@nvidia.com> * fix(providers): align readiness times with protobuf contracts Represent readiness receipts, status, and operation times with Timestamp and report intervals with Duration. Reserve the scalar field tags, update all consumers and generated bindings, and preserve timestamp presence and nanosecond identity through storage and client validation. Qualify both empty-map constructors in the Linux boundary test so its module compiles while retaining the explicit default required by Clippy. Signed-off-by: Shiju <shiju@nvidia.com> * fix(cli): preserve provider mutation storage uncertainty Recognize the gateway's exact structured storage-uncertainty reason for provider attach, detach, and update. Explain that the change may already be saved and must be reconciled before retrying, without exposing server messages or metadata. Preserve uncertainty ahead of generic retry hints. Exercise saved mutations through the CLI and verify single submission, redaction, missing receipt handling, and untrusted error-detail rejection. Document the recovery guidance for users and the public CLI skill. Signed-off-by: Shiju <shiju@nvidia.com> * fix(cli): explain denied provider profile lookups Report exact and alias profile lookup denials with fixed permission and workspace guidance. Keep backend details redacted and stop before provider mutations. Cover denied create and update calls through the CLI. Verify the complete provider list independently in the cross-workspace OIDC regression, extracting its JSON object from surrounding startup diagnostics. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
c502be9fd7 |
feat(api): return typed deletion outcomes with explicit missing-target semantics (#3317)
* feat(api): return typed deletion outcomes with explicit missing-target semantics Implement phase 2 of #3051 across the public gateway API and first-party SDKs. Preserve asynchronous sandbox deletion and observed resource identities, reserve legacy wire fields, and document the coordinated migration. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(sdk): bind deletion waits to sandbox identity Track accepted deletions by original sandbox identity in Rust and TypeScript. Preserve name-only waits and cover replacement races and lookup failures. Merge the phase-one cleanup fix and adapt its interceptor regressions to typed deletion outcomes. Refs #3051. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * test(e2e): accept asynchronous stopped sandbox deletion Allow either completed or accepted deletion output, then continue polling for actual sandbox absence. This preserves the lifecycle assertion under the phase-two deletion contract. Refs #3051. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
fd3fd9cf74 |
feat(sandbox): explain failed calls to external tool servers (#3207)
Show configured tool server addresses and their last observed connection results together in sandbox status. Keep sandbox lifecycle readiness separate so an external connection failure does not mark the sandbox unready. Expose direct endpoint records through the CLI and SDKs, with plain-language failure explanations and gateway acceptance times. Keep observation tracking, runtime reporting, and gateway validation in dedicated endpoint status modules. Preserve bounded reporting, request attribution, retry ordering, and configuration and supervisor authority checks. Clear obsolete observations while retaining the configured addresses, and document the distinction between an observed HTTP response, current availability, and tool success. Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
48c449d8c8 |
chore(deps): replace ring with AWS-LC (#3243)
* chore(deps): replace ring with AWS-LC Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(lint): address warnings after dependency upgrades Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(tls): limit provider initialization to reqwest clients Signed-off-by: Simon Scatton <sscatton@nvidia.com> --------- Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
118b250f01 |
feat(sandbox): support rootfs tar as --from source for VM driver (#2863)
* feat(sandbox): support rootfs tar as --from source for VM driver Accept flat rootfs tar archives (.tar, .tar.gz, .tgz) via the --from flag for VM-backed gateways. The CLI detects the archive extension, validates that the gateway uses the VM compute driver, and passes the tar path through driver_config. The VM driver copies the tar into its staging area and feeds it into the existing rootfs extraction and ext4 disk creation pipeline, skipping the container image pull/export steps. Closes #2175 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): validate rootfs tar path at the VM driver boundary The rootfs_tar_path field in driver_config was passed from the API caller directly to tokio::fs::copy without validation. An authenticated user bypassing the CLI could supply arbitrary host paths (e.g. /dev/zero for disk exhaustion, or readable host files for data exfiltration). Introduce a trusted staging directory that the VM driver creates on startup and advertises via GetCapabilities. The CLI now copies the tar into the staging directory before creating the sandbox, and the driver validates that the received path is a regular file inside the staging root and within a configurable size limit (default 10 GiB) before any I/O. New VmDriverConfig options: - rootfs_tar_staging_dir: override the staging directory (default: <state_dir>/rootfs-tar-staging) - rootfs_tar_max_bytes: override the size limit (default: 10 GiB) Addresses GATOR-28b5152e-01. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): request-scoped staging, size pre-check, and cleanup for rootfs tar Tighten the rootfs tar staging flow to address the remaining GATOR-01 obligations: - Request-scoped staging: the CLI creates a unique per-request subdirectory (req-<pid>) under the staging root instead of placing files directly in the shared directory. The driver enforces that the tar path is at depth 2 (staging_root/<subdir>/<file>), preventing cross-request path selection. - Size pre-check: the driver advertises rootfs_tar_max_bytes via GetCapabilities. The CLI reads this limit and rejects oversized files before copying, avoiding disk exhaustion in the staging directory. - Cleanup: the driver removes the request staging subdirectory after consuming the tar (on cache hit, copy success, or copy failure), ensuring staged data does not persist beyond the request. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): restore rootfs-tar sandboxes from persisted image identity On restore or restart, the one-shot staged tar archive has already been cleaned up. Reading the persisted image identity from the sandbox state directory and resolving the cached disk path directly avoids re-accessing the deleted staging path. Addresses GATOR-168b9210-01. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(cli): use random staging dirs and enforce byte limit during rootfs tar copy Replace PID-based request staging directories with tempfile-generated random names to prevent collisions and make paths unpredictable. Replace bare tokio::fs::copy with a streaming copy loop that enforces the advertised max_bytes limit during transfer, closing the TOCTOU gap between the pre-copy size check and the actual copy. Signed-off-by: Philippe Martin <phmartin@nvidia.com> Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): issue rootfs tar staging slots from the gateway A caller could name any host path in `driver_config.vm.rootfs_tar_path`, which the privileged VM driver then read. The CLI-side locality check did not apply to direct API requests. The gateway now owns staging. `BeginRootfsTarStaging` allocates a request-scoped directory under the driver-advertised staging root and returns an opaque single-use token; `CreateSandbox` carries the token, and the gateway substitutes the path it allocated before dispatching to the driver. `template.driver_config.<driver>.rootfs_tar_path` is rejected outright in request validation, so a caller-supplied path never reaches privileged I/O. Tokens are bound to the issuing workspace and subject, consumed once, and expire after 30 minutes. Outstanding slots are capped per caller and overall, so one caller can neither exhaust the staging filesystem nor starve others. An RAII guard reclaims the directory on every failure path after consumption, and an age-gated sweep runs at startup and on each reconcile pass for directories whose driver died before its own cleanup. The token is stripped from the public sandbox before persistence: the stored copy is returned verbatim by GetSandbox, ListSandboxes and WatchSandbox to every member of the workspace. Also fixes two defects this exposed: - The CLI wrote `rootfs_tar_path` at the top level of `driver_config`, but the gateway forwards only `driver_config.<driver_name>`, silently dropping unmatched keys. The archive never reached the VM driver, so the documented `--from ./rootfs.tar` flow did not work at all. Config is now nested under `vm` and deep-merged, so a caller's existing VM settings survive instead of being clobbered by a shallow extend. - Staging previously required `GetGatewayInfo`, which is restricted to `platform_admin`, making the feature unusable for ordinary users on any RBAC-enabled gateway. The new RPC matches CreateSandbox at `sandbox:write` / `workspace_role: user`. `compute_driver.proto` is unchanged; the gateway reads the staging root from the capabilities it already stores. Refs #2175 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): derive rootfs tar cache identity from archive contents The prepared-disk cache key combined the archive's full path with an mtime truncated to seconds, then mapped punctuation to `-`. Distinct paths such as `/tmp/a/b.tar` and `/tmp/a-b.tar` collapsed onto the same key and reused each other's disk, a rewrite within the same second kept stale contents, and a long path could exceed filesystem component limits. Identity is now a SHA-256 of the archive contents. This is also what makes the cache work at all now that the gateway allocates a fresh staging directory per request: a path-derived key would miss on every create. The archive is hashed, the cache checked, and only on a miss copied — so a hit skips writing a multi-gigabyte file. The copy is hashed as it is written and rejected if the digest differs from the first pass, which closes the window where the source changes during staging rather than approximating it with a re-stat. Refs #2175 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): decompress gzip rootfs tar archives during staging `--from` accepts `.tar.gz` and `.tgz`, but the driver staged whatever bytes it was given and the guest image-prep VM extracts the staged file with a plain `tar -xpf`. Compressed sources therefore depended on the guest tar auto-detecting gzip, and the prepared disk was sized from the compressed length, which is far too small for the expanded rootfs. Staging now detects gzip from the archive's magic bytes -- the driver only ever sees a gateway-issued staging path, never the caller's file name -- and writes an uncompressed tar. The digest still covers the source bytes, so the "archive changed while staging" check is unaffected, and expansion is bounded by `rootfs_tar_max_bytes` so a compression bomb cannot fill the host disk. `extract_rootfs_archive_to` sniffs gzip as well, so the host-side extraction path matches. Adds unit coverage for gzip staging, bounded expansion, and gzip extraction, plus an e2e sandbox created from a gzip-compressed export. Signed-off-by: Philippe Martin <phmartin@redhat.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> Signed-off-by: Philippe Martin <phmartin@nvidia.com> |
||
|
|
74960ebfae |
feat(server): add sandbox templates (#2833)
* feat(server): add sandbox workload templates Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(go-sdk): add sandbox workload template support Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(rust-sdk): add sandbox workload template support Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(python-sdk): add sandbox workload template support Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(typescript-sdk): add sandbox workload template support Signed-off-by: Gordon Sim <gsim@redhat.com> * docs(agents): document sandbox workload templates Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(cli): support default GPU requests in sandbox templates Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(server): cap sandbox templates per workspace Signed-off-by: Gordon Sim <gsim@redhat.com> * docs(architecture): document sandbox workload template boundaries Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(cli+sdk): expose sandbox workload template provenance Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(cli): include sandbox template annotations in output Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(sandbox): add label selectors to template listing Signed-off-by: Gordon Sim <gsim@redhat.com> * test(e2e): cover sandbox template failure paths Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(server): preserve command and ttl when creating sandbox from template Signed-off-by: Gordon Sim <gsim@redhat.com> * test(server): add field coverage test for template merge Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(go-sdk): add pagination support to fake client Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(cli): warn on env vars that looks like secrets Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(sdk-ts): propagate sandbox workspace through lifecycle calls Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(docs): update workspace management docs Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(ts-sdk): support command and tty when creating from template Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(go-sdk): support command and tty when creating from template Signed-off-by: Gordon Sim <gsim@redhat.com> * test(python-sdk): verify command and tty handling when creating from template Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(go-sdk): update docs and ClientInterface Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(server): validate sandbox create specs before I/O Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(go-sdk): guard empty DNS-1123 label validation Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(python-sdk): allow empty template builder mappings Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(cli): align template GPU JSON default output Signed-off-by: Gordon Sim <gsim@redhat.com> --------- Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
69a05ebb3b | fix(sandbox): complete successful main processes (#2884) | ||
|
|
40d1b48666 |
feat(provider): support for SPIFFE backed token exchange (#1970)
* feat(provider): add ability to request token exchange instead of client credentials as OAuth grant_type Signed-off-by: Gordon Sim <gsim@redhat.com> * test(proxy): add further tests for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * test(provider): add runnable example for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * test(e2e): cover Podman token exchange grants Signed-off-by: Gordon Sim <gsim@redhat.com> * refactor(oauth): extract duplicated functionality from server and supervisor Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(provider): evict nearest-to-expiry entry from intermediate token cache Signed-off-by: Gordon Sim <gsim@redhat.com> * doc(supervisor): add podman example for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(provider): withhold token-exchange subject credentials Signed-off-by: Gordon Sim <gsim@redhat.com> --------- Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
ef296806f5 |
feat(sandbox): add canonical main process (#2726)
* feat(sandbox): add canonical main process Closes #2710 Persist and supervise one canonical workload per sandbox, attach sandbox connect to its retained session, and make every unexpected main-process exit terminal. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sandbox): simplify canonical main process contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): preserve legacy VM main compatibility Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): preserve main status across driver updates Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): satisfy macOS process lint Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): gate Linux exit acknowledgement publisher Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): simplify main process plumbing Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(supervisor): make controlling tty ioctl portable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(supervisor): initialize canonical process environment Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(supervisor): optimize retained main session Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sandbox): detach main session on ctrl-c Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): use explicit main detach keys Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(dev): atomically stage Docker supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(test): align Docker main environment assertion Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sdk): expose canonical main process fields Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
c4b500a7de |
feat(helm): cert-manager external issuer + OpenShift passthrough Route (#2468)
* fix(core): trust public root CAs alongside the sandbox mTLS CA The supervisor gRPC client only trusted the CA configured via OPENSHELL_TLS_CA, since tonic ClientTlsConfig starts with an empty root store unless with_native_roots()/with_webpki_roots() is also enabled. Deployments where the gateway server certificate is issued by a public CA (e.g. cert-manager against an ACME issuer) caused every supervisor connection to fail the TLS handshake with "UnknownCA", since the sandbox mTLS CA and the server cert issuer were no longer the same. Enable both native and webpki roots in addition to the configured CA. tonic root store is a union of all configured sources, so this does not weaken verification for existing self-signed deployments. webpki-roots (compiled in) is enabled alongside native-roots since the supervisor binary may run in minimal sandbox images without a populated system CA bundle. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * feat(helm): support external cert-manager issuers and OpenShift Route passthrough Add certManager.serverIssuerRef/clientIssuerRef so the gateway and mTLS client certificates can be issued by a real Issuer/ClusterIssuer (e.g. ACME) instead of only the chart built-in self-signed CA. Add openshiftRoute template for exposing the gateway via a TLS passthrough Route so the gateway keeps terminating its own TLS/mTLS. The server Certificate excludes internal-only SANs (cluster-local, localhost, loopback) when an external issuer is configured, since ACME issuers reject those per CA/Browser Forum baseline requirements. A template-time fail guard catches the misconfiguration at helm install time rather than asynchronously at cert-manager issuance time. Includes Helm unittest coverage for both issuerRef overrides and Route rendering, plus a CI values overlay for lint coverage. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs: document cert-manager external issuer and OpenShift Route Update managing-certificates.mdx with the serverIssuerRef workflow and install-time validation behavior. Add a production section to the OpenShift guide covering passthrough Route with a real certificate. Regenerate Helm README for new certManager and openshiftRoute values. Sync debug-openshell-cluster skill with new troubleshooting steps for ACME issuance failures and supervisor UnknownCA from mismatched CAs. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(helm,core): address PR review feedback on cert-manager external issuer Addresses all five blocking review items from #2468: 1. Remove .with_native_roots() from supervisor gRPC client -- the supervisor runs inside the user-selected sandbox image, so the image CA bundle is not operator-controlled. Keep .with_webpki_roots() (compiled-in, not user-controlled) alongside the configured CA. 2. Fail at render time when serverIssuerRef.name is set but clientCaFromServerTlsSecret is still true. Add negative Helm test. 3. Remove clientIssuerRef -- changing only clientIssuerRef breaks both directions because trust bundles are not modeled separately. Change serverIssuerRef.kind default from ClusterIssuer to Issuer. 4. Add server.oidc.issuer and server.oidc.audience to the documented OpenShift production Helm command. Add Access Control prerequisite. 5. Fail at render time when openshiftRoute.enabled and disableTls are both true. Add negative Helm test. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(drivers): strip GATEWAY_TLS_SERVER_NAME from Docker and Podman env Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(helm,drivers): guard default clientCaSecretName and add env-strip tests Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * feat(tls): SNI-based dual certificate for internal and external server TLS Split the gateway server certificate into two: an internal cert issued by the chart's own CA (for supervisor connections via cluster-local SANs) and an external cert issued by an operator-configured Issuer such as ACME/Let's Encrypt (for CLI and Route access via public SANs). The gateway uses SNI-based certificate selection: connections whose SNI hostname matches external_server_names receive the external cert; all others (including those with no SNI) receive the internal cert. Security improvement: remove .with_webpki_roots() from the supervisor gRPC client so supervisors trust only the chart CA, closing a MITM vector via publicly-trusted certificates in user-supplied container images. Key changes: - Add DualCertResolver with SNI-based cert selection and full test coverage - Add external_cert_path, external_key_path, external_server_names to TlsConfig - Validate partial external cert config (error on cert-without-key or vice versa) - Validate empty external_server_names when external cert is configured - Split cert-manager templates into internal + external Certificate resources - Add Helm guards for misconfigured external issuer (empty serverDnsNames, internal-only SANs with external issuer, conflicting clientCaFromServerTlsSecret) - Update gateway-config.mdx, managing-certificates.mdx, openshift.mdx docs - Update debug-openshell-cluster skill for dual-cert troubleshooting Signed-off-by: Pi Agent <agent@openshell.local> * fix(drivers): strip GATEWAY_TLS_SERVER_NAME in VM driver and correct comments Add the same GATEWAY_TLS_SERVER_NAME environment stripping to the VM compute driver that Docker, Podman, and Kubernetes drivers already perform. Without this, a sandbox user on the VM driver could override the TLS server name the supervisor verifies. Fix stale comments in Docker and Podman drivers that referenced 'with WebPKI roots trusted' — WebPKI roots are explicitly not trusted after the tls-webpki-roots removal. Use tls-ring instead of bare channel for tonic in openshell-core so the TLS API (ClientTlsConfig, Endpoint::tls_config) is available without pulling in any root certificate store. Signed-off-by: Pi Agent <agent@openshell.local> * fix(tls,helm): wildcard SNI matching and Route host validation Add RFC 6125 single-level wildcard matching to DualCertResolver so external_server_names entries like *.example.com correctly match SNI hostnames like gw.example.com. Previously only exact matches worked, silently falling back to the internal cert for wildcard configurations. Add a Helm fail guard in route.yaml that rejects openshiftRoute.host values not listed in certManager.serverDnsNames when an external issuer is configured — catches cert/route hostname mismatches at install time instead of at TLS connect time. Quote the host field in route.yaml for robustness. Signed-off-by: Pi Agent <agent@openshell.local> * fix(helm): address blocking review items — client-CA guard, wildcard Route, serverIssuerRef gate 1. Remove the obsolete guard rejecting serverIssuerRef + clientCaFromServerTlsSecret=true. The internal server certificate is always signed by the chart CA (the same CA that signs the client cert), so clientCaFromServerTlsSecret=true is correct — its filtered ca.crt is exactly the right trust anchor. The old workaround (mounting openshell-ca-tls directly) unnecessarily exposed the CA private key to the gateway container. Remove the client-CA overrides from docs, CI overlay, and production examples. 2. Route host validation now supports wildcard certificates per RFC 6125: single-level wildcards like *.example.com match gateway.example.com but not deep.sub.example.com. Require an explicit openshiftRoute.host when an external issuer is configured — without one, OpenShift generates a hostname absent from serverDnsNames. 3. Reject serverIssuerRef.name when certManager.enabled is false — the external certificate, its Secret mount, and the gateway TLS config all require cert-manager to be enabled. Validated on ROSA (dev.dyee.p3) with branch-built images: - Fresh install with letsencrypt-prod ClusterIssuer - SNI dual-cert: external hostname served Let's Encrypt cert - Supervisor mTLS via internal cert path: ConnectSupervisor accepted - Client CA volume: filtered ca.crt from internal server secret (no key) - CLI connected via Route + OIDC Helm tests: 81 pass across 7 suites. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> Signed-off-by: Pi Agent <agent@openshell.local> |
||
|
|
0f8fad23c4 |
feat(sandbox): add stop and start operations (#2653)
* feat(sandbox): add suspend and resume operations Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(server): preserve lifecycle work after cancellation Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(server): reconcile ambiguous lifecycle outcomes Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(server): complete suspended session cleanup Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(vm): preserve suspension state on resume failure Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(server): retry retained lifecycle transitions Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(server): clean sessions after suspend reconciliation Signed-off-by: Seth Jennings <sjenning@redhat.com> * test(sandbox): cover deleting suspended sandbox Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(kubernetes): preserve progressing sandbox suspension Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(kubernetes): bound suspend status polling Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(kubernetes): detect legacy sandbox suspension Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(tui): render suspended sandbox phases Signed-off-by: Seth Jennings <sjenning@redhat.com> * refactor(sandbox): rename suspend and resume lifecycle Signed-off-by: Seth Jennings <sjenning@redhat.com> * perf(server): clean stopped sessions on transition Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(kubernetes): fail fast on rejected stop Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(compute): fence stale restart lifecycle events Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
9c019a93f5 |
Wire authorization into workspace model (#2445)
* feat(auth): implement RFC 0011 Phase 2 workspace authorization Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): address PR review feedback on workspace authorization - Docker e2e: add --health-port and switch readiness probe from `openshell status` to `curl /healthz`, fixing a false-positive readiness check in OIDC mode where the CLI exited 0 without actually contacting the gateway - ListWorkspaces: move membership filtering from post-query N+1 lookups into a SQL EXISTS subquery so pagination applies to the visible set, not the global ordering. Add generic list_with_membership to the persistence layer. - Descriptor validator: reject role/scope fields on unauthenticated and sandbox auth modes, and allow-list workspace_role as user/admin and global_role as platform_admin to catch typos at startup Signed-off-by: Derek Carr <decarr@redhat.com> * fix(server): use authed request in delete telemetry test The workspace authorization added by the Phase 2 auth changes requires a Principal on every delete request. The delete-telemetry test was still using a bare Request::new, so extract_principal failed before the handler could acquire the delete gate, causing a 5-second timeout flake. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): address gator review findings for workspace authorization - Inject unauthenticated-local-dev principal in no-auth gateway mode so handlers that call extract_principal() always find one. - Cap label-selector membership query at MAX_PAGE_SIZE instead of u32::MAX to bound the in-memory read. - Authorize workspace membership before resolving workspace existence in all sandbox RPCs to prevent workspace-name enumeration by non-members. - Remove dead_code allow on AuthorizedWorkspace.workspace now that callers use the normalized name from the authz result. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): close workspace-name oracle and label-selector truncation Swap authorize-before-resolve ordering in 27 handlers across provider.rs, service.rs, policy.rs, and workspace.rs to prevent CWE-203 workspace-name enumeration by non-members. Add combined membership+label SQL query (list_with_membership_and_selector) to both persistence backends so ListWorkspaces with label selectors no longer silently drops results beyond the first page of membership matches. Signed-off-by: Derek Carr <decarr@redhat.com> * test(auth): add non-member rejection and membership+label persistence tests Add comprehensive test coverage for workspace authorization changes: - Non-member rejection tests across all 44 workspace-scoped handlers (sandbox, provider, service, policy, workspace, inference) verifying PERMISSION_DENIED is returned instead of NOT_FOUND to prevent CWE-203 workspace-name oracle - Persistence test for list_with_membership_and_selector verifying SQL-level membership EXISTS + label filtering, multiple predicates, no-match cases, and pagination Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): format merged import line in sandbox tests Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): address gator re-review findings on workspace authorization - Fix TUI unconditionally setting providers_v2_enabled after provider refresh; read the actual gateway setting via GetGatewayConfig at startup instead - Fix SQLite json_extract with dotted label keys (e.g. example.com/env) by quoting the key in the JSON path - Add authed_request wrappers to upstream OCI identity tests that were missing a principal after rebase - Add test proving GetGatewayConfig is accessible without Platform Admin - Add test for dotted/prefixed Kubernetes-style label key filtering Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): address second gator re-review findings - Loosen GetGatewayConfig from platform_admin to scope-only so workspace users can discover providers_v2_enabled during sandbox creation with inferred-provider commands; update proto descriptor, descriptor validation, and RFC 0011 access table - Add validate_label_selector to handle_list_workspaces and escape single quotes in SQLite json_extract interpolation (CWE-89 defense-in-depth) - Re-fetch providers_v2_enabled after TUI gateway switch so the new gateway's capability is reflected - Add e2e test for workspace user with inferred-provider command - Add persistence test for adversarial label keys with SQL injection attempts - Add handler test for invalid label selector rejection in ListWorkspaces Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): address third gator review findings - Cap label selector pairs at 64 (CWE-400) to bound SQLite dynamic SQL - Add SCOPE_ONLY_METHODS allowlist for scope-without-role RPCs (CWE-863) - Normalize ID-based data-plane handlers to return NOT_FOUND for unauthorized sandboxes, closing the cross-workspace oracle (CWE-203) - Fix TUI provider profile cache lookup key mismatch for legacy providers with empty profile_workspace - Add whoami to CLI skill reference command tree - Update TUI skill doc with workspace, provider, and settings coverage - Document scope/workspace orthogonality on GetGatewayConfig proto Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): extend CWE-203 normalization to policy.rs sandbox handlers GetSandboxConfig and GetSandboxLogs in policy.rs had the same fetch-before-authorize pattern that leaked cross-workspace sandbox existence. Promote fetch_and_authorize_sandbox to pub(super) and use it from both sandbox.rs and policy.rs handlers. Signed-off-by: Derek Carr <decarr@redhat.com> * test(auth): update assertions for CWE-203 sandbox ID normalization Cross-workspace sandbox access via ID-based handlers now returns NOT_FOUND instead of PERMISSION_DENIED to prevent existence inference. Update the unit test and OIDC e2e assertion to match. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(auth): narrow CWE-203 error mapping and correct whoami output formats Only remap PERMISSION_DENIED to NOT_FOUND in fetch_and_authorize_sandbox and RevokeSshSession, letting INTERNAL and UNAUTHENTICATED propagate as-is. Fix whoami --output format values in cli-reference.md to match the actual CLI (table/json/yaml, not text/json). Signed-off-by: Derek Carr <decarr@redhat.com> * fix(ci): share network namespace with Keycloak in containerized CI In GitHub Actions job containers, Docker port publishing lands on the host, not inside the job container. Detect this environment and attach Keycloak to the job container's network namespace instead, with hardened defaults (cap-drop ALL, no-new-privileges, loopback-only listener). Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
5952a5a23f |
feat(workspace): add workspace resource model with scoping, membershi… (#2243)
* feat(workspace): implement workspace model (Phase 1 of RFC 0011) Implements workspace and membership model providing hard isolation boundaries for multi-player OpenShell deployments. Workspace CRUD with Kubernetes-style Terminating phase for graceful deletion. All resources scoped by workspace via ObjectMeta. Membership RPCs for workspace access control. Persistence migration shifts name uniqueness to (object_type, workspace, name). Provider profiles support platform and workspace scoping. Service routing uses workspace-prefixed DNS labels. Inference routes renamed and workspace-scoped with DeleteInferenceRoute RPC. Python SDK with WorkspaceClient, two-method list pattern (workspace-scoped and for_all_workspaces), and workspace parameter on all methods. CLI workspace flags, TUI workspace cycling. K8s driver filters unmanaged CRs and uses delete preconditions. Podman driver uses immutable container IDs. Label serialization fixed across all put_if call sites. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(cli): delegate sandbox upload command to existing upload function The standalone `sandbox upload` command reimplemented upload logic inline with two bugs: it used `Path::exists()` which follows symlinks (rejecting dangling symlinks), and it ran git-aware filtering on symlink sources. The `run::sandbox_upload()` function already handles both cases correctly via `sandbox_upload_plan()`. Replace the inline logic with a call to the existing function. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(e2e): shorten sandbox names and fix test compatibility Shorten the sandbox name in initial_sparse_policy_is_acknowledged_as_loaded from 'e2e-2159-sparse-enrich' (22 chars) to 'e2e-sparse-enrich' (17 chars) to comply with MAX_ROUTABLE_NAME_LEN (19 chars). Also capture stderr in create_keep_with_args so future sandbox creation failures include the actual CLI error instead of reporting empty output. Signed-off-by: Derek Carr <decarr@redhat.com> * test(workspace): add test coverage for workspace CRUD and persistence isolation Add unit tests for workspace create happy path, get round-trip, get not-found, get empty-name rejection, already-exists error, and resolve_workspace not-found. Add persistence test proving cross-workspace name uniqueness (same name in different workspaces produces separate records). Add workspace name max-length boundary tests. Fix e2e harness to include stderr in name-parse-failure error path. Align Python e2e test_workspace_crud with try/finally pattern. Document provider profile catalog workspace scoping gap in RFC 0011. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(examples): update examples for workspace model compatibility Shorten sandbox names in demo scripts to fit the 19-character MAX_ROUTABLE_NAME_LEN limit: policy-demo prefix to pd-, multi-agent notepad derives a short SANDBOX_TAG from the run ID, governance interceptor uses gs-PID-RANDOM. Update vscode-remote-sandbox.md SSH host aliases from openshell-{name} to openshell-{name}.{workspace} format. Signed-off-by: Derek Carr <decarr@redhat.com> * feat(sdk): add workspace-scoped client and workspace CRUD Add WorkspaceScopedClient modeled after kube::Api::namespaced — captures workspace once and injects it into every sandbox request. Add workspace CRUD methods (create, get, list, delete) and list_sandboxes_all_workspaces on OpenShellClient. Extend SandboxRef with workspace field and add WorkspaceRef type. Include mock tests for all new operations. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(lint): resolve clippy warnings in workspace test assertions Signed-off-by: Derek Carr <decarr@redhat.com> * fix(docs): convert indented code blocks to fenced in RFC 0011 Signed-off-by: Derek Carr <decarr@redhat.com> * fix(lint): resolve clippy warnings and apply cargo fmt across workspace Auto-format with cargo fmt and fix clippy warnings exposed by the reformat: unnecessary qualifications, map_unwrap_or, identical match arms, unused variable prefix, dead code annotations, and let-unit-value in e2e harness. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(workspace): address workspace scoping issues from review - Add workspace field to settings JSON output (CLI) - Skip Podman containers missing workspace label instead of defaulting to empty string, matching K8s driver behavior - Add resource_version to list_by_scope SELECT in both SQLite and Postgres backends, with regression test - Gate PolicyLocalContext proposal/lookup routes on workspace readiness, returning 503 when workspace is not yet discovered - Block sandbox and provider creation in TUI all-workspaces mode - Clear workspace vectors in TUI reset_sandbox_state Signed-off-by: Derek Carr <decarr@redhat.com> * fix(workspace): make provider profile catalog workspace-aware Thread workspace through snapshot_catalog so the EffectiveProviderProfileCatalog enforces workspace boundaries on both read and write paths. UserProviderProfileSource now loads platform-scoped profiles (workspace "") plus the target workspace's profiles, preventing cross-workspace duplicate profile ID collisions that previously caused global catalog failures. Update RFC 0011 to reflect catalog scoping is implemented in Phase 1 rather than deferred to future work. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(persistence): include workspace column in atomic policy revision INSERT put_policy_revision_atomic omitted the workspace column from the INSERT into the objects table in both SQLite and Postgres backends, causing atomically-written policy revisions to lose their workspace association. Add workspace field to AtomicPolicyRevisionWrite and thread it through both backend INSERT statements, matching the non-atomic put_policy_revision path which already included it. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proxy): skip ancestor walk when socket owner is the entrypoint collect_ancestor_identities walked the entire process tree above the entrypoint when the connecting process was the entrypoint itself, SHA256-hashing every ancestor binary (IDE, shell, container runtime). On dev machines with large binaries in the ancestor chain this exceeded the 30-second test timeout. When start_pid == stop_pid there are no intermediate ancestors to verify, so return an empty list immediately. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(workspace): make provider profile catalog scope-aware Allow the same profile ID at platform and workspace scopes by introducing layered catalog entries where workspace profiles shadow platform profiles. Add source and scope fields to the ProviderProfile proto and CLI output. Migrate List/Get handlers to the catalog, fixing divergence with runtime profile resolution. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(e2e): align podman e2e labels with centralized driver constants The podman driver moved its container labels to the centralized openshell.ai/ prefix, but the e2e test harness and cleanup script still referenced the old openshell.sandbox-* keys, causing the local_driver_token_restart test to fail on container lookup. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(e2e): align python profile isolation test with scope-aware catalog Platform profiles are now visible in workspace listings as fallbacks per the layered catalog design. Update the assertion to match. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(workspace): honor profile_workspace in runtime profile resolution Runtime profile lookups now consult provider.profile_workspace via get_type_profile_for_scope. Providers created with --global-profile (profile_workspace="") resolve to the platform profile even when a workspace profile shadows the same ID. All 6 runtime call sites updated; type-only call sites remain scope-agnostic. Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
21aaa89522 |
feat(gateway): add elevated gateway info (#2202)
* feat(cli)!: fold gateway metadata into list BREAKING CHANGE: openshell gateway info no longer shows local gateway registration metadata. Use openshell gateway list or openshell gateway list -o json for local registration details. Signed-off-by: Evan Lezar <elezar@nvidia.com> * feat(gateway): add elevated gateway info Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
48a7d09e82 |
feat(providers): support profile updates (#1914)
* feat(providers): support profile updates Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(providers): require CAS for profile updates Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(providers): serialize profile update validation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(providers): serialize profile import invariants Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(providers): require profile update target id Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
ff028ce0df |
feat(server): support TLS certificate hot-reload (#1870)
* feat(server): support TLS certificate hot-reload Signed-off-by: Yuedong Wu <dwcn22@outlook.com> * refactor(server): extract shared TLS test utilities Signed-off-by: Yuedong Wu <dwcn22@outlook.com> --------- Signed-off-by: Yuedong Wu <dwcn22@outlook.com> |
||
|
|
d5b79e5ba1 |
refactor(server): deduplicate test helpers and grpc utilities (#1708)
Remove three groups of copy-pasted code in openshell-server: 1. grpc/mod.rs had a private current_time_ms() wrapper identical to the one already exported from persistence/mod.rs. Remove the duplicate and update the three grpc sub-modules (policy, sandbox, service) to import directly from crate::persistence. 2. test_store() was repeated verbatim in seven #[cfg(test)] blocks. Promote a single canonical version to persistence/mod.rs (cfg-gated) and replace all copies with crate::persistence::test_store() calls or a thin Arc wrapper in supervisor_session. 3. grpc_client_mtls() and build_tls_root() were copy-pasted across edge_tunnel_auth.rs and multiplex_tls_integration.rs. Move both into the existing tests/common/mod.rs shared module and import from there. |
||
|
|
2bdc968ed9 |
fix(gateway): make readiness health checks dependency-aware (#1328)
* feat(gateway): add readiness probe metrics and test-only store close Emit Prometheus readiness metrics for database probes (healthy gauge and outcome-labeled latency histogram) with coverage in health HTTP tests. Restrict Store::close behind test support cfg to prevent accidental runtime pool shutdown under live traffic. Signed-off-by: Adrien Langou <alangou@nvidia.com> * test(e2e): add simple e2e test with kubernetes to test /readyz Signed-off-by: Adrien Langou <alangou@nvidia.com> --------- Signed-off-by: Adrien Langou <alangou@nvidia.com> |
||
|
|
a3b16c18ab | feat(auth): per-sandbox authentication to gateway (#1404) | ||
|
|
5620c8b229 |
refactor: deduplicate repeated patterns across crates (#1499)
Remove ~280 lines of duplicated code across 30 files in 5 areas: - centered_rect: consolidate 5 identical TUI layout helpers into a single pub fn in openshell-tui/src/ui/mod.rs - server test helpers: replace ~100 inline Store::connect() calls with local test_store() helpers; deduplicate test_server_state() in grpc/service.rs to use the shared test_support version - rogue PKI: extract 20-line rogue CA+client cert generation block (duplicated in two integration tests) into generate_rogue_pki() in tests/common/mod.rs - provider tests: replace 8 identical 28-line test modules with a single macro_rules! test_discovers_env_credential! invocation - label constants: centralize openshell.ai/ container label keys in openshell-core::driver_utils; update Docker and Kubernetes drivers to import from there instead of redefining them locally |
||
|
|
10af3e609f |
refactor: deduplicate shared test helpers (#1399)
Extract four categories of copy-pasted test code into single canonical locations: - POLICY_OBJECT_TYPE / DRAFT_CHUNK_OBJECT_TYPE constants that were defined identically in both persistence/sqlite.rs and persistence/postgres.rs are now owned by persistence/mod.rs. - test_server_state() was duplicated across grpc/provider.rs, grpc/policy.rs, and grpc/sandbox.rs. A shared grpc::test_support module in grpc/mod.rs now owns the single implementation; each submodule imports it. - The TestOpenShell stub (full OpenShell trait impl), install_rustls_provider(), PkiBundle, generate_pki(), and start_test_server() were copied across up to five openshell-server integration test files. A new tests/common/mod.rs module owns them; each test file uses mod common. - StubResponse, unique_socket_path(), and spawn_podman_stub() were duplicated between openshell-driver-podman/src/driver.rs and openshell-driver-podman/src/grpc.rs. A new src/test_utils.rs (cfg(test)-gated) owns the shared helpers. |
||
|
|
d255cdd9c9 |
feat(providers): add credential refresh foundation (#1349)
* feat(providers): add credential refresh foundation * feat(providers): mint oauth refresh token credentials * fix(providers): allow delegated refresh bootstrap * fix(providers): guard refresh credential modes * fix(providers): refine refresh lifecycle UX * fix(cli): restore provider test gateway response import * fix(server): clean auth endpoint test qualifications * fix(providers): tighten refresh authorization and collisions * chore(providers): trim bundled v2 profiles * fix(providers): resolve refresh rebase fallout * fix(cli): accept rfc3339 credential expiry * test(providers): update inferred claude provider type * test(providers): avoid removed outlook default profile * test(providers): isolate attach limit fixtures |
||
|
|
442b0b6b92 |
feat(exec): add bidirectional streaming for interactive TTY sessions (#1331)
* feat(exec): add bidirectional streaming for interactive TTY sessions The existing ExecSandbox RPC sends stdin upfront in the request body, making interactive programs (bash, top, vim) unusable with --tty. This adds an ExecSandboxInteractive bidirectional streaming RPC that forwards live keystrokes, terminal resize events, and EOF to the sandbox. - Add ExecSandboxInteractive RPC, ExecSandboxInput and ExecSandboxWindowResize messages to the proto definition - Implement server handler with relay bridging, PTY allocation via russh channel.split(), and timeout support - Implement CLI client with raw mode, spawn_blocking stdin reader, SIGWINCH resize forwarding, and process::exit for clean shutdown - Route to interactive path only when --tty is explicitly passed Signed-off-by: Florent Benoit <fbenoit@redhat.com> * fix(exec): address interactive TTY streaming review feedback Server-side (sandbox.rs): - Close SSH channel after EOF so PTY-backed programs (bash, vim, top) terminate on client disconnect instead of leaking the session - Break output loop when tx.send() fails (client gone) to tear down the streaming pipeline promptly - Default zero cols/rows to 80x24 before request_pty, matching the proto "0 = use default" contract and preventing 1x1 terminals for non-CLI clients Client-side (run.rs): - Replace spawn_blocking stdin reader with a detached std::thread so the tokio runtime shutdown does not hang on a thread blocked in stdin.read(). This removes the need for std::process::exit(), allowing save_last_sandbox() in the caller to execute normally - Switch from futures::channel::mpsc with try_send() to tokio::sync::mpsc with blocking_send() for proper backpressure — try_send silently stopped forwarding all input when the channel filled during paste or slow network - Add TaskGuard (abort-on-drop) for the resize task to ensure cleanup on all return paths including early errors --------- Signed-off-by: Florent Benoit <fbenoit@redhat.com> |
||
|
|
c94cddbfb8 |
feat(server): separate HTTPS from mTLS authentication (#1351)
Make --tls-client-ca optional and make client certificates always optional when a CA is configured. This decouples HTTPS encryption from mTLS authentication, allowing mTLS and OIDC bearer tokens to coexist as parallel authentication mechanisms. When --tls-client-ca is provided, client certificates are validated against the CA when presented but never required. Clients may connect with or without a certificate — authentication is handled at the application layer (e.g. OIDC). Two TLS modes are now supported: - HTTPS with optional mTLS (--tls-client-ca provided) - HTTPS-only (--tls-client-ca omitted) The --disable-gateway-auth flag is preserved for backward compatibility but is now a no-op. The allow_unauthenticated field has been removed from TlsConfig. The Helm chart conditionally includes the client-ca volume and env var based on whether clientCaSecretName is configured. |
||
|
|
0797fefa44 | feat(gateway): add local-domain service routing (#1101) | ||
|
|
5abc36c461 | feat(relay): route forwarding through ForwardTcp (#1029) | ||
|
|
1d3b741ee3 |
feat(providers): support sandbox provider attach lifecycle (#1242)
* feat(providers): support sandbox provider attach lifecycle Closes #1171 Adds sandbox provider list, attach, and detach API/CLI support while keeping provider policy and credential resolution derived from current sandbox attachments. * fix(providers): refresh sandbox provider credentials Adds provider environment revisions and generation-scoped sandbox credential snapshots so future SSH and exec launches pick up provider attach, detach, and credential updates without mutating already-running processes. Also blocks provider deletion while attached to prevent stale sandbox provider references. * fix(providers): serialize sandbox object mutations * test(providers): cover sandbox provider attach lifecycle * test(providers): accept versioned credential placeholders |
||
|
|
cdb1de59ba |
feat(providers): add custom profile registry (#1170)
Add custom profile registry. Allow attaching custom profiles at sandbox start. |
||
|
|
043bde279a |
feat(providers): add profile-backed policy composition (#1037)
Foundation for providers v2. Add provider profiles and provider profile composition with user policies. |
||
|
|
04e48d585c |
feat(server): add request-ID middleware for request correlation (#1082)
Add a UUID-based request-ID middleware using tower-http's request-id feature. Each inbound request receives a unique x-request-id header (or preserves a client-supplied one), which is recorded in the tracing span and propagated to the response. This enables operators to correlate log lines across the middleware stack for a single request under concurrent load, and lets clients reference specific requests in bug reports. Signed-off-by: sauagarwa <sauagarw@redhat.com> |
||
|
|
24724742a8 | ci(rust): enforce -D warnings on clippy (#1008) | ||
|
|
a6d45528c1 | feat(server,sandbox): supervisor-initiated SSH connect and exec over gRPC-multiplexed relay (#867) | ||
|
|
bbcaed2ea7 | refactor(proto): rename UpdateSettings to UpdateConfig for consistency with read path (#515) | ||
|
|
a831a8921b |
feat(settings): gateway-to-sandbox runtime settings channel (#474)
* feat(gateway/sandbox): add global and sandbox runtime settings flow |
||
|
|
fbd93a4632 | refactor: rename navigator- crate prefix to openshell- (#277) |